Vibeleaderboard
Index / article

AgentMemBench

arxiv.org
Visit arxiv.org
Category
Developer Tools
Pricing
Open Source
Type
ARTICLE
Added
Aug 4, 2026

About

A reproducible benchmark comparing five long-term memory strategies for conversational AI agents (in-context windowing, external key-value store, graph-based episodic memory, compression summarisation, web-augmented memory) across three datasets. Findings show external key-value retrieval dominates on quality metrics, especially for long-range recall where other methods nearly fail, at the cost of higher memory footprint.

What it can do

  • Compare long-term memory strategies for conversational AI agents across multiple datasets

    Conversational AI agent memory strategy implementationsQuality metrics comparison (e.g., recall performance, memory footprint)

Why it made the leaderboard

AgentMemBench shows external key-value retrieval substantially outperforms in-context windowing, graph-based episodic memory, summarization, and web-augmented approaches on long-range recall tasks, at the cost of memory footprint — a concrete tradeoff to weigh when designing agent memory.

Tags

benchmarkai-agentslong-term-memoryllm-evaluationretrievalmemgpthipporag

Media

AgentMemBench

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.