AgentMemBench
arxiv.org- Category
- Developer Tools
- Pricing
- Open Source
- Type
- ARTICLE
- Added
- Aug 4, 2026
About
A reproducible benchmark comparing five long-term memory strategies for conversational AI agents (in-context windowing, external key-value store, graph-based episodic memory, compression summarisation, web-augmented memory) across three datasets. Findings show external key-value retrieval dominates on quality metrics, especially for long-range recall where other methods nearly fail, at the cost of higher memory footprint.
What it can do
Compare long-term memory strategies for conversational AI agents across multiple datasets
Conversational AI agent memory strategy implementations → Quality metrics comparison (e.g., recall performance, memory footprint)
Why it made the leaderboard
AgentMemBench shows external key-value retrieval substantially outperforms in-context windowing, graph-based episodic memory, summarization, and web-augmented approaches on long-range recall tasks, at the cost of memory footprint — a concrete tradeoff to weigh when designing agent memory.
Tags
Media

Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.