← All IntelClip / AI AgentsIn-band memory writing improves with model generation
From Claude for Long-Horizon Tasks — Lance Martin, Anthropic · ≈11:19
“So, the key point I'm making here is that Claude has gotten much better at this in-band memory writing across model generations.”
AI Engineer
“higher capacity models have a better sense of like what abstraction to save to memory that'll be useful later. Like they're not just writing a specific fact. They're writing how do I how does this generalize to future sessions?”
AI Engineer
“This is a task that basically asked the model to perform a sequential question answering with a SQL database and it can write memory in between each step. And what you see is basically the performance improves across models.”
AI Engineer
“So, the memories it writes are pretty crappy. It's kind of um it's kind of uh tactical notes. It's it's not very strategic. And the game progress is quite limited.”
AI Engineer
What’s in it
- Backed by two data points — Claude plays Pokemon notes going from tactical to strategic across generations, and a sequential SQL QA task on the open-source Continual Learning Bench.
Clip transcript
system, that is basically a memory directory. That's really it. Now, this is showing some work on Claude plays Pokémon with Claude's on it 3.5. And here's the key point. When Sonnet 3.5 is given this access to a memory directory, and it can write memory, {quote} in-band, as it progresses through this game, it's not very good. So, the memories it writes are pretty crappy. It's kind of um it's kind of uh tactical notes. It's it's not very strategic. And the game progress is quite limited. But with more recent models like this is looking at 4 6, the notes are much more strategic and game progress is much further. So, the key point I'm making here is that Claude has gotten much better at this in-band memory writing across model generations. And this is another way to show that same result. So, this is a benchmark that I ran called Continual Learning Bench. It's an open-source benchmark. I took one of the tasks. This is a task that basically asked the model to perform a sequential question answering with a SQL database and it can write memory in between each step. And what you see is basically the performance improves across models. Um and so what this is kind of showing is that models get natively better at this in-band memory writing with respect to model capability. >> [snorts] >> And some of the most interesting things I found from this are that the main differentiation between There we go. Um the main differentiation between like a lower capacity model and a high capacity model is kind of this distillation step. And so, basically, higher capacity models have a better sense of like what abstraction to save to memory that'll be useful later. Like they're not just writing a specific fact. They're writing how do I how does this generalize to future sessions? That's kind of the key difference that I found that higher capacity models kind of have when they're writing memory. So, this is a very important thing to keep in mind that models are getting better better and better at this kind of in-band memory writing across model generations.
Comments
Checking sign-in…
Loading comments…