Cartograph: Federated Tool Discovery with Operator-Attested Retrieval for AI Agents
Source
Justice Owusu Agyemang, Michael Agyare, Kwame Opuni-Boachie Obour Agyekum, Kwame Agyeman-Prempeh Agyekum, Francisca Adoma Acheampong, Jerry John Kponyo
Author
Justice Owusu Agyemang, Michael Agyare, Kwame Opuni-Boachie Obour Agyekum, Kwame Agyeman-Prempeh Agyekum, Francisca Adoma Acheampong, Jerry John Kponyo
Date
Key takeaways · AI-distilled
Cartograph is a federated MCPThe Model Context Protocol — an open standard that lets any AI assistant plug into any tool or data source without custom integration code.Full definition → proxy that swaps full tool-catalog loading for progressive disclosure: on a 22-server, 374-tool deployment the AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → sees three proxy tools instead of 374 definitions.
A top-5 discovery exchange used 475 tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → versus 42,450 under full-catalog accounting, and on a 49-query author-built benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → recall at 5 was 0.816 against 0.592 for a keyword baseline.
Tool descriptions are operator-attested capability cards signed with Ed25519 under the deploying operator's control, so ranking does not rely on publisher-written copy.
Its Rift analysis flagged 49 clusters of easily confused tools, four rated high risk, and the gateway added about 5ms mean latency (0.8%) over direct stdio MCP calls.
Terms in this piece · Glossary
MCP — The Model Context Protocol — an open standard that lets any AI assistant plug into any tool or data source without custom integration code.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Cartograph collapses MCP tool discovery from loading every tool definition to three proxy tools and beats keyword search on retrieval accuracy (0.816 vs 0.592 recall@5) on a 374-tool, 22-server deployment.