Turning My Obsidian Vault Into a Local AI Engineer — Filip Makraduli, Superlinked
Source
youtube.com
Author
AI Engineer
Date
Why it matters
Routing document work to local models via MCP cuts context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → cost and keeps data in-house. But if the agent can still read raw files first, instructions alone leave a privacy gap, so access must be separated.
Key takeaways · AI-distilled
Makraduli's demo turns a scanned NDA into a redacted Markdown document before a coding AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → uses it: small open models on team-controlled GPUs do the OCR and redaction, and the agent receives only the smaller processed artifact.
The setup splits into a host, an MCP server exposing extraction, summarization, question answering and redaction tools, and GPU worker pools; Superlinked's inference layer handles model loading, batching and workloads sharing the cluster.
He ties these choices to embeddingA list of numbers representing a piece of text's meaning, so that similar meanings end up numerically close and can be searched.Full definition →, rerankingA second pass that re-scores retrieved candidates by reading each one against the query, fixing the ordering that fast vector search got approximately right.Full definition → and task-specific LoRAs, and to routing work between models already warm on a GPU, so different document jobs can share one private cluster.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
embedding — A list of numbers representing a piece of text's meaning, so that similar meanings end up numerically close and can be searched.
reranking — A second pass that re-scores retrieved candidates by reading each one against the query, fixing the ordering that fast vector search got approximately right.