
If you are building agents that must remember a project across many sessions, this quantifies how badly off-the-shelf vector, graph and file-based memory degrade once recall has to be joined against structured domain data — best system under 33% accuracy. It gives you a reason and a template to evaluate memory on your own domain rather than trusting generic conversational-recall scores.
“The strongest system achieves only 32.4% answer accuracy under a deployment-realistic ingestion scope, and remains below 60% under oracle-filtered ingestion or a stronger probe agent.”
“Analysis shows that current general-purpose memory systems often retrieve topically relevant context but store project knowledge as incomplete or fragmented facts.”
Checking sign-in…
Loading comments…