Beyond Static Intelligence: Evaluating Continual Learning — Parth Asawa, UC Berkeley
- Source
- AI Engineer
- Author
- AI Engineer
- Date
- benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
- AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
- context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
It supplies a way to measure whether an actually learns across tasks, and finds the elaborate -management systems losing to in-context learning.
“Every leaderboard you have seen was built by asking a model to do one task, wiping its memory, and asking it another.”
AI Engineer
“run a system with state, then run the identical system reset between every single instance, and take the difference.”
Parth Asawa
“Cumulative reward cannot show you this, because a stronger base model can post a higher total while learning less than a weaker one that genuinely improves.”
AI Engineer
“The headline result is uncomfortable: plain in context learning tops the leaderboard, beating the more elaborate context management systems on reward, on gain, and on cost.”
AI Engineer
“standard benchmarks are deliberately independent and therefore offer nothing to improve on, which is why chaining existing benchmarks together does not work.”
Parth Asawa
videoDistill the LLM, Don't Serve It: Search & Personalization at DoorDash — Raghav Saboo, DoorDash
videoTeaching LLMs to Speak Spotify — Yves Raimond & Jacqueline Wood, Spotify
videoWhy LLM Recommenders Will Be AI's Biggest Consumer App — Devansh Tandon, Meta
videoWorld Models Need Causality, Not Pretty Pixels — Christopher Manning, Moonlake AI
Checking sign-in…
Loading comments…

