
Self-graded episodic memory in agents can inflate confidence in bad past actions, causing agents to repeat their own mistakes; LUCID's answer-free de-inflation approach lifts BIRD execution accuracy from 52.4% (no memory) and 54.0% (self-graded memory) to 56.9%, giving a concrete fix for a subtle but consequential failure mode in memory design.
“Incorrect episodes receive inflated rewards; thus, the agent preferentially reuses the very mistakes it has most confident in.”
“a usable signal must track truth *and* decorrelate its error from the memory bias”
articleWhen Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM AgentsShweta Mishra, Shashank Mishra
articleVoice Memory for Agentic Speech RecognitionChao-Han Huck Yang, Zih-Ching Chen, Piotr Zelasko, Zhehuai Chen, Jagadeesh Balam, Boris Ginsburg
articleFinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM AgentsBen Wang, Kang Zhou, Lifan Guo, Feng Chen, Chi ZhangChecking sign-in…
Loading comments…