A technical progression from a Python list to graph-vector hybrid memory, with the tradeoffs at each step. The reference to reach for when "just stuff it in the context window" stops scaling.
Key takeaways Β· AI-distilled
Accuracy drops over 30% when the relevant fact sits in the middle of a long context windowThe maximum amount of text a model can consider at once β its working memory for the current conversation or task.Full definition β rather than near either end. Buying a bigger window does not buy reliable recall β it just moves where things get lost.
A conversation list evicts strictly by age, so the user's name from turn one disappears before yesterday's throwaway joke. The flaw is not size, it is that chronology is standing in for importance.
Flat markdown memory is genuinely good at four facts and useless at two thousand. Once the file outgrows the context window your only retrieval tool is keyword matching, which misses anything phrased differently.
Consolidation is the piece people skip: repeated specific events should distill into a general rule, so noticing across dozens of sessions that users prefer executive summaries becomes a standing instruction rather than dozens of replayed episodes.
Terms in this piece Β· Glossary
context window β The maximum amount of text a model can consider at once β its working memory for the current conversation or task.
Key quotes
βThe "memory" you feel when chatting with ChatGPT is an illusion created by re-sending the entire conversation history with every request.β
βAccuracy drops over 30% when relevant information sits in the middle of a long context.β
βMemory isn't about cramming more text into the prompt. It's about structuring what the agent remembers so it can find what matters.β
βStrip away the frameworks and an agent is a loop: perceive, think, act.β