Continual Learning: How AI Agents Get Better With Every Use | Arjun Karanam, Trajectory
Source
youtube.com
Author
Sequoia Capital
Date
Why it matters
Agents that do not retain lessons between runs repeat the same mistakes. This talk lays out an approach to continual learning from agent trajectories, which bears on how you design feedback loops and evals.
Key takeaways · AI-distilled
Arjun Karanam says thumbs-up/down feedback is too noisy, since coding-AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → users tend to accept everything. The useful signal is corrective behavior: edits, undos and retries, which products should both prompt for and capture.
Karanam warns that tool calls returning only 'done' confuse agents and leave nothing to learn from. He recommends tool responses that say what was actually read or written.
Trajectory sorts feedback by where it belongs. A fact, like a company being delisted, goes to the agent harnessThe scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.Full definition →'s context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → and not the weights. A tool call that keeps failing is likely true for everyone, so it goes into the model. One user's preferences stay in context.
Karanam's evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → advice: build evals from real traffic, make every task replayable ('roll outable'), which he calls a big infrastructure challenge, and grade through the real production harness rather than a variant of it.
On privacy, Karanam says you can avoid training on customer data directly. Instead, sample its distributions, generate synthetic data, and check that the synthetic data matches the real distribution.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.