Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory
Source
arxiv.org
Author
Mustafa Arslan
Date
Why it matters
A concrete answer to tool-registry bloat: route tools by retrieval over a compact signature instead of re-prefilling schemas, and treat KV-cache splicing as bounded by RoPE phase drift rather than by effort.
Terms in this piece · Glossary
MCP — The Model Context Protocol — an open standard that lets any AI assistant plug into any tool or data source without custom integration code.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.