An LLM has no memory. Every turn, the entire conversation is re-sent, and "remembering" is just a decision about what to put in that payload. The context window makes that decision urgent: past some length you physically cannot send everything, so something has to be dropped, and the four standard strategies differ only in what they drop. The right question is not "how do I give my agent memory" — it is "which facts will still be reachable at turn 200, and at what cost per turn".
The experiment: a long conversation in which facts are stated at random turns ("my flight is on the 14th", "I'm allergic to shellfish", "the project code is BLUEJAY"). At turn T the user refers back to one of them, chosen at random from everything said so far. Each strategy gets the same token budget, and the measurement is simply: was the needed fact in the context? 4000 conversations per point:
The shapes are the lesson, and they are all different. Full history is perfect until it hits the budget and then falls off a cliff — and note that its cost per turn is rising the whole time, so it is getting more expensive right up to the moment it breaks. Sliding window has a flat, honest failure: it remembers exactly the last W turns and nothing else, forever. Summarization degrades gracefully but compounds — each round of compression loses a little more, so a fact stated at turn 3 has been through many summaries by turn 200. Retrieval is the only one whose recall does not depend on when the fact was said, which is why it is the strategy that survives long conversations. Note what it costs, though: while everything still fits in the window, retrieval is behind the others, because its recall is capped by the store's search quality. It does not win by being better — it wins by not degrading.
Recall is only half a decision. The same four strategies, priced by how many tokens each sends on every single turn of the conversation:
Full history's cost is linear in turns, so the total cost of a conversation is quadratic in its length — every turn re-sends everything before it. That is the real reason nobody ships it, ahead of the window limit. The other three are flat in the conversation length, which is what makes them viable, and they differ mainly in what they spend the flat budget on.
In practice you combine them, because they are not really competitors:
The design decision that matters most is the one this simulation cannot make for you: what gets written to the store at all. Writing every turn makes retrieval noisy and expensive; writing only what an extraction step judges to be a durable fact keeps the store small and sharp, at the cost of an extra call per turn and a new failure mode — a fact the extractor did not think was important is a fact you no longer have. Everything on this page assumes the fact made it into the store; deciding that is the actual engineering.
Second opinion (taught here — these corroborate): LangGraph · memory · Lost in the Middle · Anthropic · building effective agents.