Skip to content

When to Forget, What to Keep: Four Papers That Rethink Agent Memory

#llm-agents #agent-memory #context-management #retrieval #context-compaction #long-horizon-agents

The context window is a budget, not a buffer ​

Watch a coding agent work a real repository task and you'll see the problem within ten minutes. It inspects files, follows imports, tries an edit, runs tests, gets a failure, and backtracks. Two hours in, the context window is full of exploration that no longer matters: dead code paths, outdated diffs, the fourth variant of a function that the second one replaced. The window didn't overflow, once. It just quietly degraded.

Most systems treat this as plumbing. Truncate the oldest messages when the window fills. Summarize everything into a digest. Retrieve a few chunks and hope. That works until it doesn't, and it fails in ways that are hard to debug because the problem is what the agent no longer knows, not what it knows wrong.

Four new papers from the October arXiv roundup attack this at four different points in the pipeline. AutoCompact trains the agent to decide when to compact. Mem++ refuses to distill documents at write time. CMP shows that memory utility scores are silently biased by retrieval. SourceLearn treats repeated access to a source as a chance to master it. Different problems, one through-line: memory decisions should be explicit, and they should be measured.

Four papers, four intervention points ​

PaperIntervention pointCore moveHeadline resultWhat it means in practice
AutoCompactCompaction decisionTrain when and what to compact as part of the agent policy, then RL with task success+9.2 points SWE-bench Verified, +5.0 points SWE-PolyBench VerifiedCompaction is a learned policy, not an overflow trigger
Mem++Write and read pathsStore documents whole with dates and authors; select at read time+8.0 to +13.1 points OrgMemBench, 2.6 above RAG with gpt-4.1-miniPoint-in-time questions need the full record, not extracted facts
CMPRetrievalReserve context slots for propensity-sampled memories; inverse propensity weightingAUC 0.54 to 0.66 on LongMemEvalWithout randomized retrieval, most memory utility scores are unidentifiable
SourceLearnPersistent source modelSelf-directed revisits plus task-guided updates, rebuilt from the sourceBest in 13 of 15 settings, up to +22.6 over Hybrid RAGRepeated use of a source should compound into competence

None of these is a storage trick. They all change who decides and when.

AutoCompact: compaction as a learned policy, not a fire drill ​

The standard approach to context overflow is reactive: when the agent hits the limit, compact everything into a summary and continue. AutoCompact starts from a different observation. Earlier exploration goes stale long before the window fills, and the right time to compact depends on the task. Sometimes you should drop a dead search path at turn 40. Sometimes you shouldn't drop that one code snippet at turn 200, because it's the key to the bug.

So AutoCompact trains the agent to make compaction decisions as part of its policy. The data pipeline is the clever part. They run the base agent on coding tasks, and a judge reviews each compaction decision, the summary produced, and the actions that followed it. Flawed decisions get replaced with corrected ones, and the trajectory continues from the corrected state so the model sees what good compaction looks like downstream. That yields supervised fine-tuning data. Then they jointly optimize coding and compaction with reinforcement learning, using task success as the reward. Coding behavior and compaction behavior improve together.

The results hold across inference budgets: +9.2 points absolute on SWE-bench Verified and +5.0 on SWE-PolyBench Verified, a harder multi-repo benchmark. The gain appears whether the agent runs in a 256K window that never overflows or a 16K window whose overflow triggers fallback compaction. SWE-bench Verified grades on real GitHub issues with hidden tests, so that 9.2-point jump means measurably more tasks resolved with the same base model. The fact that the numbers hold at 256K matters too: the gain is not about escaping overflow. It's about deciding that some context is dead weight and acting on it.

I've watched agents burn tokens re-reading files they already explored. This is the first system I've seen that treats that waste as a learnable decision rather than an engineering accident.

Quick Take: The direction of travel is to decide less at write time and more at read time, because the future question is unknown when the memory is stored.

Mem++: stop distilling at write time ​

Mem++ comes from the organizational setting, and the motivating example is sharp. Many authors record decisions across documents over months. A revised decision arrives as a new document, not an edit to the old one. Someone asks: what was the policy in March? The document record contains the answer, but most memory systems never make it reachable.

Write-time distillation is the culprit. Most systems compress each document into facts, notes, or graph edges at ingest. That fixes what can be answered before anyone asks. A fact extractor that drops the date an old decision was superseded has made the version question unanswerable forever. Mem++ takes the opposite trade. It stores every document whole, with its date and author, and calls no generative model at write time. Writes are cheap, and nothing is decided about the future at ingest. At read time, it retrieves only documents dated up to the moment the question asks about and fuses lexical and semantic rankings. The answering model makes the choice between versions.

On OrgMemBench it beats the strongest memory baseline by 8.0 to 13.1 points across two answering models. With gpt-4.1-mini it posts the best overall score, 2.6 points above plain RAG. It also gets the best average LLM-judge score on LoCoMo and lands second on LongMemEval-S, behind only its own entity-graph variant.

The cost is real: read-time selection on a full document store is more expensive per query than a pre-built graph. For organizational memory, the record is the product. Version-aware answers are the whole game.

CMP: your memory scores are systematically biased ​

This is the paper that made me recheck my assumptions. Memory-augmented agents often estimate each memory's utility by its effect on task performance, then use those estimates to decide what to retain and what to evict. The Causal Memory Policy paper (CMP) shows the estimates are built on sand.

The failure is a retrieval-level positivity violation. Utility estimates rely entirely on retrieved memories. If a memory is never retrieved, you observe identical outcomes whether it's in the store or not. Its utility is unidentified, and no diagnostic that inspects memory operations alone can see the gap, because from the store's perspective nothing happened.

The scale of the problem is worse than I expected. On LongMemEval, 54% of the memories the system actually required were unidentified. On LoCoMo, 67%. The failure persists in a deployed memory system, so this is not a benchmark artifact.

Key Numbers

  • 54% of required memories on LongMemEval have unidentifiable utility. 67% on LoCoMo.
  • 0.54 to 0.66 AUC: what causal retrieval intervention buys. The before number is barely better than a coin flip; the after number is a solid improvement with plenty of room left.
  • 0.78 AUC for per-query utility, and still no aggregation that predicts value on unseen queries. Past usefulness does not imply future usefulness.

The fix is to intervene on retrieval itself. Reserve a fixed number of context slots for memories sampled with known propensities, then estimate each memory's utility with self-normalized inverse propensity weighting under a balanced assignment design. The paper proves unbiasedness and exact variance of the estimator, and derives the optimal decision rule under irreversible operations like deletion. The implementation cost is one number: a known sampling probability per memory. That's a small price for turning coin-flip discrimination into 0.66 AUC.

The humbling part is the ceiling. Per-query utility reaches 0.78 AUC on the query it was estimated for, but no aggregation available to a retention policy predicts a memory's value on unseen queries. If you delete memories based on how useful they were, you are betting that usefulness transfers, and this evidence says it does not.

SourceLearn: build competence in the source itself ​

The last paper reframes a problem you may not have noticed you had. Agents that repeatedly use the same persistent source, a codebase, an API spec, a regulatory corpus, treat each contact as fresh retrieval. Search, fetch chunks, answer, done. Nothing compounds. The agents get faster at finding the right pages, but never better at understanding the source.

SourceLearn's "source learning" is the idea that repeated use should develop source-specific competence: how the source is structured, how its passages should be interpreted, how its knowledge applies to tasks. That competence lives in a persistent source model. Two mechanisms refine it. Self-directed learning identifies what remains incompletely understood and revisits the source on its own. Task-guided learning uses downstream task failures to expose representational gaps and recurring needs. Both produce learning signals, but the actual updates are always reconstructed from the authoritative source. The model can't drift from the source because it re-derives everything from it.

The numbers are strong: best performance in 13 of 15 settings across five benchmarks and three LLM backends, with gains up to 22.6 points over Hybrid RAG. That 22.6-point gap is the distance between a lookup system and a specialist. The three-backend consistency matters for a different reason: the result is not a quirk of one model's weights.

Where the four approaches fit together ​

Looking at the four papers as a stack makes the pattern clear. Each one claims a decision point in the memory pipeline and shows that the default choice there is wrong.

Compaction is a policy, not an overflow handler. Storage should not destroy information at write time. Retrieval must be randomized to make memory evaluation honest. Repeated access should compound into understanding. A production agent memory system probably needs all four layers, and they stack exactly as the diagram shows: a named decision at every handoff, each with a measurable effect on task success.

Common Pitfalls ​

  • Compact only on overflow. If your agent compacts when the window fills, you're making a capacity decision, not a quality decision. AutoCompact's results suggest the right moment is often earlier, and the timing is task-dependent. A fixed threshold throws away the signal you need.
  • Distill organizational documents at write time. Extract facts and notes from a decision record and you've made the version history unanswerable. Mem++ shows read-time selection on whole documents preserves what write-time distillation destroys. If your queries can ask about states of the record, don't compress the record.
  • Trust utility scores for deletion without checking identification. If a memory was never retrieved, its utility score is unidentified, and 54 to 67% of required memories fall in that bucket on standard benchmarks. You cannot diagnose this by inspecting memory operations. Randomize at least a fraction of retrieval slots, or stop deleting on past utility.
  • Tune compaction separately from coding behavior. AutoCompact's improvement came from joint optimization: RL on task success shaped both coding and compaction together. Optimize compaction as an isolated summarization module and you get summaries that read well and help nothing.
  • Build one flat memory store for all sources. SourceLearn's edge comes from source-specific structure. A generic experience-memory that retrieves chunks across everything never develops the kind of competence that produces 22.6-point gains over Hybrid RAG.

One thing to remember ​

Every one of these papers is about who decides, not where things are stored. AutoCompact gives the decision to the policy. Mem++ gives it to the answering model at read time. CMP makes the retrieval design acknowledge its own bias. SourceLearn lets the agent decide what about a source it does not yet understand. If you take one design rule from this batch: a memory system is a sequence of decisions, and every default choice, especially the invisible ones, deserves a measurement.

The Bottom Line ​

  • If you're building a long-horizon coding agent in a 16K to 64K window, adopt a learned compaction policy like AutoCompact instead of truncation or overflow-triggered summarization, because timing compaction as part of the policy is what produced 9.2 absolute points on SWE-bench Verified.
  • If you're building memory for organizational documents where version history matters, skip write-time distillation entirely and store documents whole with dates and authors, because read-time selection beat the strongest memory baseline by 8.0 to 13.1 points on OrgMemBench.
  • One thing to watch: propensity-aware retrieval is about to become table stakes. Expect memory frameworks within six months to expose randomized retrieval slots or propensity logging as first-class features, because the CMP results show that without intervention, more than half of your stored memories have unidentifiable value.