Skip to main content

Command Palette

Search for a command to run...

Long-term memory is compression, not retrieval

Updated
•8 min read•View as Markdown
Long-term memory is compression, not retrieval

Epistemic status: A design note. The motivating failure is from a long-form narrative agent I ran. One ICLR 2026 workshop paper supplies the measurement. Scope is at the end.

Related: Chen, Epistemic Memory Failures in Long-Form Narrative Agents: A Deployment Study (ICLR 2026 Workshop on Memory for LLM-Based Agentic Systems).


In a long narrative agent I ran, Alice already knew she was being followed. She learned it two chapters back. She changed her route because of it. Then the next scene opened with her noticing the same tail as if it were a fresh clue. She asked her companion what it meant. She reasoned it out from scratch.

The prose was fluent. The story world did not collapse. Nobody invented a city or a corpse. She had forgotten that she already knew.

That is the failure that made me distrust "just retrieve more history." The facts were still true. The next action was built on the wrong slice of a true past.

The disk is full of true noise

The usual fix is a bigger window, a vector store, another retrieval pass. Treat memory as an extra disk. Write everything. Fetch the nearest chunks. Paste them into the prompt.

I have run agents long enough for that story to stop working. The facts are often already in context. The model still does not know which three of them should govern the next step.

An agent's life is a pile of state changes. A place it visited. An event it triggered. A relationship it broke. For the next scene, maybe three facts matter. The other 99% can be perfectly true and still be interference. True history is not free context. Most of it is noise with a high relevance score.

Similarity search ranks what looks like the current scene. That is not the same job as ranking what must still bind it.

Take one stretch of the same story. Alice is being watched. She decides to detour. She fights with a companion. She picks up a new clue.

Compress for the task, and you keep: she is heading to the old station; the clue points at the watcher.

Compress for the relationship, and you keep: trust with the companion has broken; the next exchange has to be guarded.

Compress for who knows what, and you keep: she already knows she is being watched; she must not play shock.

The model can write any of those scenes. What it lacks is a rule that says which cut of the past is the one to use right now.

A viewpoint is a compression rule

I use viewpoint here as an engineering word: the use-constraint that decides which facts survive into the next prompt, rather than a mood or an opinion the model should perform.

Same history. Different constraint. Different memory. The job of long-term memory is not to store more. It is to compress a huge true record under an explicit rule, until only the facts that should bind the next action remain.

Skip that rule and retrieval still returns truth. The next scene still drifts. The model treats a known fact as scenery, or as a new puzzle, because nothing told it how the fact is allowed to be used.

A retrieved sentence is a candidate. A compressed packet is an instruction: keep this, discard that, and here is how the kept fact may be used.

What I measured

I wrote this up as Epistemic Memory Failures in Long-Form Narrative Agents: A Deployment Study, at the ICLR 2026 Workshop on Memory for LLM-Based Agentic Systems.

The run was one deployment: about three months, ninety chapters, more than 180,000 tokens. The story world could stay consistent. The facts could be true. Characters still re-asked, rediscovered, or re-derived things they had already learned. Recency-based context was dropping mid-chapter keys. I called that known-information forgetting.

This fails differently from the usual hallucination. The world did not break. A fact was not invented. What failed was the thin record of who already knows what.

Take Alice. The fact is the same one from the opening: organization X is tracking her.

The live record says she does not know it. Bob, her companion, does. No reveal has happened. Unchecked, the model writes a fluent scene in which she confronts him with the tail. The prose is fine. The step is illegal. She has not learned it yet. Left alone, that scene becomes the next chapter, and later chapters treat the leak as history.

Now take the other frame, the one I actually saw at scale. She already knows. Unchecked, a fluent draft lets her find the same tail again. She is shocked. She investigates. She spends the scene earning a fact she already has. The next action is delayed, and the "discovery" hardens into a second origin story.

Key Facts Injection (KFI) is the cheap repair: pull the important facts from episode memory and put them back with an explicit use tag. A bare line like "organization X is tracking Alice" reads like scenery, or like a new lead. Prefer a tagged packet:

Alice already knows that organization X is tracking her route. This should guide her next action, not be rediscovered as new information.

After that packet, she already knows. The fact has to guide the next action. It must not be rediscovered. In that deployment, this cut the forgetting events by 73%.

The method is a tagged packet: the fact, who already holds it, and how it is allowed to bind the next step. That is compression under a who-knows-what constraint, not a better search hit.

Two stages, because the model will fill gaps

A single model call will invent that packet if you let it. The compression has to be a harness: control flow around the model, so the lens and the facts are not the same job.

A cheap, fast model is allowed to pick the compression direction. It is not allowed to mint the facts. It can say: keep continuity of who knows what; do not rediscover known facts; keep only what changes the next scene.

A stronger model then reads that lens and extracts facts from traceable memory. The packet that goes forward has to point at a recorded state, not at a fluent paraphrase the small model made up.

An external checker decides whether the packet may be written. Models can be swapped. The control flow stays.

That split matters because the quiet failure is not only omission or invention. It is wrong compression. Every retained sentence can be true, and the weights can still be wrong. At a crisis, the compressor drops "she already knows she is being watched" and keeps "she ate an apple this morning." Both are correct. One of them is fatal. The next step is deduced from a true, unfocused slice, and the plot logic dies.

The use-constraint is what names the discard rule. Which memories must be thrown out this frame. Which ones must come back with a high-priority tag.

The warehouse loop is the wrong picture

A lot of retrieval stacks still do: write, retrieve, concatenate, generate. For a long-running agent that is a brittle loop. A wrong splice pollutes the record the next call will treat as true.

The picture I want is a runtime pipeline. Capture the live record. Retrieve under a constraint. Compress under a use-constraint. Ground the facts in sources. Check them. Inject the tagged packet. Generate the next action. Keep an audit trail that can roll a bad write back.

The domains change. A novel. A multi-day research log. A long support thread. A coding agent that has to live with last week's decisions. The pinch is the same. Facts pile up. Context is scarce. A bad splice is not a one-off error. It becomes the world the next call sees.

So the test is not how many gigabytes of history the store holds. Before the next action, can the system hand the model a small packet of key facts that are sourced, checkable, and tagged with who already knows them?

If it can, the agent has memory. If it can only search, it has a disk.

Scope and limits

The 73% figure is from that one narrative deployment, not from a public bake-off. I have not compared tagged injection against equal-token retrieval or against someone else's memory stack. The two-stage pipeline is a design I use. It is not a model horse-race.

I would drop the compression framing if naive retrieval on long windows stopped producing these rediscovery traces. I would also drop it if ordinary episode summaries kept who-knows-what well enough that an explicit tagged packet sat unused.

Until then: store the history if you must. Compress it under a viewpoint before the next step.

M

The cheap-model-picks-the-lens, strong-model-extracts-under-it split is the part I hadn't seen articulated this cleanly before. One thing I'm not sure the two-stage design handles: your three example compressions (task, relationship, who-knows-what) each read like they're picking a single dominant constraint per scene, but a lot of real next-actions need more than one satisfied at once - Alice has to act on the tail without confronting Bob about knowing, while their trust is still broken. If the cheap model is choosing one lens per compression pass, does the axis that didn't get picked just silently fall back into the same rediscovery risk for that turn, or does the packet actually carry more than one constraint at a time?

9
99 sunny5d ago

Agreed. The concrete wiring for viewpoint compression and KFI injection lives in my GitHub project ai-bios-narrative-agent under sunnyspot114514. It has the state baseline, dynamic prompt compression, labeled memory injection, and validator-gated commits in working form. Happy to walk through any part if useful.

9
99 sunny5d ago

The three cuts in the essay are pedagogical. Same history under one named constraint at a time so the difference is easy to see. In the runtime the packet is not a single lens winner. It is a small set of tagged facts, and each fact keeps its own use rule. For the Alice scene you mentioned the cheap stage should return an active constraint set, not one enum. Keep who knows what on the tail. Keep the relationship guard with Bob. Keep the task move that the surveillance forces. The strong stage then extracts only sourced facts that satisfy that set. Soft axes that do not bind this turn can wait. Hard constraints that make the next action legal or illegal cannot. If who knows what drops out while Alice still has to move under the tail, you get the rediscovery failure again. That is why the validator sits after the packet. Even a partial compression pass is not allowed to commit a draft that rediscovers a known fact or leaks a fact Alice does not yet hold. So the short answer is yes, the packet can and should carry more than one constraint when the next action needs them jointly. The cheap model is choosing which constraints are live for this turn, not crowning one dominant story beat.