RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation
Source
arxiv.org
Author
Yuan Si, Simeng Han, Daming Li, Jialu Zhang
Date
Why it matters
How you render retrieved memory can swing accuracy by tens of points at a matched token budgetA cap on how many tokens a task, session, or agent run may consume — the practical control on both cost and how long an agent will grind.Full definition → — three models score 0% on formal typed ledger packets but 45-53% on the same facts as natural-language entries.
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
token budget — A cap on how many tokens a task, session, or agent run may consume — the practical control on both cost and how long an agent will grind.