The Recall Trap: A Recall-Maximizing Retriever Configuration Reduces Issue Resolution in Fixed-Budget Code Context
Source
Alexander Adkins, Teimuraz Trapaidze
Author
Alexander Adkins, Teimuraz Trapaidze
Date
Terms in this piece · Glossary
SWE-bench — The standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Why it matters
Retrieval for code agents is routinely tuned against recall@k, and this is direct evidence that optimizing that metric can cost you resolved issues when the context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → is fixed.