Memory: What to Keep

Lesson 12 / 16 · updated 2026-10-01 · 8 min


ONELINE

Memory is not storage, it is a forgetting policy. The design question is never what to save — it is what to throw away, and when.

METAPHOR

A notebook, a diary, and a textbook. Three memories that answer three different kinds of question.

The model is stateless. Every request starts from nothing; the only reason a conversation feels continuous is that you re-send the history each time.

That works until the history outgrows the budget — which, per Chapter 10, an agent’s does quickly. Then you need somewhere else to put things, and a policy for what goes where.

Three tiers

Memory tiersFlowchart with 3 labelled stages: L1 WORKING - the context window This turn's focus. Under 50ms. Wiped every session.; L2 EPISODIC - vector store What happened before. 100-300ms. 'What did we decide last Tuesday?'; L3 SEMANTIC - graph or database Durable facts and rules. Over 500ms. 'This user is on the enterprise plan.'. Connections: L1 WORKING - the context window This turn's focus. Under 50ms. Wiped every session. leads, consolidate: summarise
what mattered, to L2 EPISODIC - vector store What happened before. 100-300ms. 'What did we decide last Tuesday?'; L2 EPISODIC - vector store What happened before. 100-300ms. 'What did we decide last Tuesday?' leads, consolidate: extract
stable facts, to L3 SEMANTIC - graph or database Durable facts and rules. Over 500ms. 'This user is on the enterprise plan.'; L3 SEMANTIC - graph or database Durable facts and rules. Over 500ms. 'This user is on the enterprise plan.' leads, retrieve on demand, to L1 WORKING - the context window This turn's focus. Under 50ms. Wiped every session.; L2 EPISODIC - vector store What happened before. 100-300ms. 'What did we decide last Tuesday?' leads, retrieve on demand, to L1 WORKING - the context window This turn's focus. Under 50ms. Wiped every session..
consolidate:summarisewhat matteredconsolidate: extractstable factsretrieve on demandretrieve on demandL1 WORKING - the context windowThis turn's focus. Under 50ms.Wiped every session.L2 EPISODIC - vector storeWhat happened before. 100-300ms.'What did we decide last Tuesday?'L3 SEMANTIC - graph or databaseDurable facts and rules. Over 500ms.'This user is on the enterprise plan.'
Three tiers, three latencies, three purposes. Information flows up by consolidation and down by retrieval.
Tier What it is Where it lives Latency
L1 Working This turn’s active focus Context window / KV cache <50 ms
L2 Episodic What happened before Vector store 100–300 ms
L3 Semantic Durable facts and rules Graph or database >500 ms

L1 is the context window. Chapter 6 governs it: the recent turns, the system prompt, whatever was just retrieved. It is fast because it is already resident in GPU memory, and it is gone when the session ends.

L2 is episodic — experiences, in order. “What did we decide about the pricing page last Tuesday?” Stored as embedded summaries, retrieved by similarity. This is ordinary retrieval (Part II) pointed at the conversation history instead of a document corpus.

L3 is semantic — facts stripped of their episode. “This user is on the enterprise plan and prefers metric units.” Not a memory of a conversation; a distilled fact. Best stored as structured data or a graph, because you want to query it exactly rather than approximately.

ASIDE

A good test of which tier something belongs in: if you would answer with “you told me on Tuesday”, it is L2. If you would answer “that’s just true about you”, it is L3.

Consolidation: the part that is actually hard

Information moves upward by being compressed and generalised, which is exactly where systems go wrong.

From L1 to L2, at the end of a session: summarise what happened, keeping decisions, constraints, and unresolved threads; drop the redundant tool output and the false starts.

From L2 to L3, over time: notice that a user has mentioned a deadline in three separate conversations and promote “this user’s fiscal year ends in March” to a durable fact.

Two failure modes sit on either side of this, and you will meet both.

WATCHOUT

Remembering too much is the more common failure. Every stray remark becomes a “fact”, the store fills with contradictions, and retrieval surfaces a mention of a competitor from eight months ago as though it were current preference. Memory quality degrades as volume grows.

Remembering too little is the more visible one: the agent asks the same question every session and feels broken.

Bias toward forgetting. Store facts that are stable, consequential, and likely to recur — not everything that was said.

Facts change, and old ones do not politely delete themselves

The hardest problem in long-lived memory is contradiction. A user said they worked at Company A. Now they say Company B. Both are in the store. Which wins?

You need three things:

  • Timestamps on every memory, so recency is knowable.
  • An update path, not just an append path. Systems that only ever insert accumulate contradictions until retrieval becomes a coin flip. When a new fact contradicts an old one, the old one should be superseded or deleted.
  • A precedence rule. Structured L3 facts generally beat fuzzy L2 recall, because L3 was written deliberately and L2 was inferred.

A practical consolidation step asks the model, for each candidate memory: is this new, does it update an existing memory, or does it contradict one? Then it writes, updates, or deletes accordingly. Treating memory as a mutable store rather than an append-only log is the single most important design choice here.

WATCHOUT

Consolidation is also an attack surface. If a tool result or a pasted document can reach the consolidation step, then whatever it asserts can be promoted to a durable fact — and durable facts are loaded unconditionally on every future session.

This is prompt injection with a persistence bug attached. Every other failure in this book ends when the request does; a poisoned memory is still there next week, shaping answers, with nothing in the logs to explain why.

Consolidate from your own turns, not from retrieved content. Record where each fact came from. And make L3 inspectable by the user whose memory it is, because they are the only reviewer who can tell a wrong fact from a right one.

Memory is a privacy surface

The moment you persist information about users across sessions, you have built a system with obligations.

WATCHOUT

Scope every memory to its owner and enforce that at query time, not in the prompt. A memory store shared across tenants without a hard filter will eventually surface one customer’s information to another — and it will do so fluently, with no error anywhere in your logs.

You also need deletion to actually work. “Forget everything about me” must remove the L2 entries, the L3 facts, and anything derived from them. Design for this on day one; retrofitting deletion into a memory system that only ever appended is genuinely painful.

This is one of the strongest arguments for keeping durable memory structured rather than embedded in free text. You can delete a row. Finding every paraphrase of a fact scattered across thousands of embedded summaries is much harder.

What this looks like in practice

For most systems, this is enough:

  1. Keep the last N turns verbatim in context (L1).
  2. When the session gets long, compact the older turns into a summary that stays in context (Chapter 6).
  3. At session end, write a summary to a vector store (L2).
  4. Extract durable, stable facts into a structured store (L3).
  5. On a new session, load L3 facts unconditionally — they are small — and retrieve from L2 only when the current query suggests history is relevant.

Step 5 matters: L3 is cheap and always relevant, L2 is expensive and only sometimes relevant. Retrieving episodic memory on every turn is a common way to spend a lot of tokens filling context with old conversations that have nothing to do with the question.

What to carry forward

Three tiers, distinguished by how durable and how structured the information is. Consolidation upward is compression, and it is where quality is won or lost. Prefer forgetting. Make memory mutable so facts can be corrected. Treat the store as a privacy surface from the first commit.

That closes Part III. You can now build a system that answers, retrieves, and acts. Part IV is about the three things that decide whether it survives contact with real users.

RECALL

  1. Name the three memory tiers, with an example question each one answers.
  2. What is consolidation, and why is it the hard part?
  3. Which failure is more common — remembering too much or too little — and why is the other more visible?
  4. A user changes jobs. Describe what a well-designed memory system does with the old fact.
  5. Why prefer structured storage over embedded text for durable facts? Give two reasons.
  6. Why load L3 on every session but retrieve L2 only conditionally?