
Anyone building an memory layer is likely reaching for summarization by default; this shows summarization silently drops the preference signal personalization depends on, and that plain retrieval can outperform it.
“Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons.”
“On 2,994 questions from 276 personas, seven frontier LLMs and up to seven memory configurations remain far from saturated: no full-context configuration exceeds approximately 0.47 overall accuracy or approximately 39% on multiple-choice questions.”
“Attribution analysis shows that summary-based memory often preserves factual details while losing the preference signals needed for personalization; simple retrieval can therefore outperform purpose-built memory systems, with the gap widening after shocks.”
articleSetoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous DataLingyang Zeng, Guangze Chen, Kaichen Yu, Zhicheng Pan, Siyang Weng, Zirui Hu, Xiangyun Du, Hailin He, Rong Zhang, Chengcheng Yang, Kai Huang, Xuan Zhou
articleWhen Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM AgentsShweta Mishra, Shashank Mishra
articleMemArenaJiadong Zhang, Xiaosong MaChecking sign-in…
Loading comments…