iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
For long-running agent work, the pattern that current sources support is not to keep enlarging the prompt with every earlier step. It is to keep selected information outside the active context, store it in a form the agent can read or search later, and bring back only what the current step needs. The published results behind this pattern are specific to the systems and tasks they tested. They show that selective storage and retrieval can be useful design choices. They do not show that persistent memory universally beats a larger context window.
This article is a report on published design work and vendor documentation. It does not describe a controlled test run by the author on a particular agent, and none of the figures below are personal benchmark results.
Why carrying raw history stops working
An agent doing a multi-hour task accumulates far more than conversation. It collects tool outputs, file reads, failed attempts, user corrections, and intermediate conclusions. If every model call receives all of it, three things happen. Cost and latency grow with the history. Early decisions and constraints get buried under later, less important material. And the model has to sort through logs that are mostly irrelevant to the step it is currently taking.
A larger context window raises the ceiling on how much material fits, but it does not decide what matters. Putting more history into the window is still a bet that the useful parts will be noticed in the middle of the noise.
#1 Best Overall
Context, compaction, and durable memory are different things
These three terms are often used interchangeably, and that confusion leads to poor design choices. They describe different places where information can live.
| Mechanism | What it is | What persists after the call | Main risk |
|---|---|---|---|
| Context window | The material the model can use during one inference step | Nothing beyond what is re-sent in the next call | Cost, dilution by irrelevant material, and hard truncation at the limit |
| Compaction | Summarizing a running session near the context limit, then continuing from the summary | The summary, which replaces earlier detail | Losing details whose importance only becomes clear later |
| Persistent memory | Notes or structured records stored outside the prompt and retrieved when relevant | Selected notes, facts, or skills kept in an external store | Stale, missing, or wrongly retrieved entries |
Compaction and persistent memory can be combined. Compaction keeps the live session moving. Persistent memory carries what must survive beyond that session. In both cases, anything the agent uses still enters the context at the moment of use. Memory changes which information gets there, not whether the model needs it in order to act.
The simplest version: persistent notes
The most basic form of memory is a set of notes the agent writes and rereads. Anthropic’s engineering article on context management describes this technique as follows: “Structured note-taking, or agentic memory, is a technique where the agent regularly writes notes persisted to memory outside of the context window.” The same article describes the pattern as useful for tracking progress, decisions, and dependencies across long tasks. It also notes that Anthropic provides a file-based memory tool on its developer platform. Confirm current availability and terms in Anthropic’s developer documentation before relying on it, since product details change.
Recommended Free Tools
Rank #2
A workable notes pattern
- Define the fields each note must carry: the goal, decisions made and the reason for each, open tasks, dependencies between them, and where relevant files or outputs live.
- Write notes at checkpoints, such as after a decision, a completed subtask, or a constraint discovered in the environment. Writing after every tool call usually produces clutter.
- Store the notes outside the prompt, in a file or database the agent can read back.
- At the start of each new step or session, load the notes that match the task. Some setups load everything, while others let the agent search for the relevant entries.
- When a decision changes, mark the older note as superseded instead of appending a contradiction and hoping the agent resolves it.
The weak point is quality. A note that records a wrong assumption will carry that error forward just as faithfully as a correct one. The pattern works only as well as the agent’s judgment about what to write.
Structured memory: gist plus lookup, and knowledge-centric stores
More elaborate approaches go beyond notes. They try to separate what the agent knows from the raw material that produced it, and they retrieve units of knowledge rather than whole transcripts.
ReadAgent: compressed gists with access to the original text
ReadAgent, described by Google DeepMind researchers in 2024, partitions a long document into episodes, creates a concise gist memory for each, and retrieves the original passages when more detail is needed. The design keeps compression but does not rely on it alone, so the agent can return to source text when a summary is too lossy. In the paper’s evaluations on QuALITY, NarrativeQA, and QMSum, the approach extended effective context length by 3 to 20 times. That figure applies to those long-document reading tasks. It is not a general measure of agent capability.
Rank #3
PlugMem: facts and skills instead of transcripts
Microsoft Research describes PlugMem as a system that converts agent interactions into structured facts and reusable skills, then retrieves and distills the knowledge relevant to the current task. Its authors report that PlugMem outperformed generic retrieval methods and task-specific memory designs across three benchmarks, while using significantly less memory-token budget. The published description does not give a single numeric improvement figure, so none should be inferred from it. These are results from the authors’ evaluations, and PlugMem is a research system rather than a shipped product.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →The authors frame the motivation this way: “It seems counterintuitive: giving AI agents more memory can make them less effective.” Read that as the premise of their design rather than a general law. The idea is that unfiltered memory adds noise, and that organizing memory into reusable units and retrieving only task-relevant ones is the remedy.
Comparing the approaches
The options differ on five practical axes. The table below uses those axes. Where a published source does not address a cell, the cell says so.
| Approach | What is stored | When it is retrieved | Update behavior | Fidelity and provenance | Typical failure |
|---|---|---|---|---|---|
| Larger context window | Raw transcript and tool output | Always present in each call | Only appended | Complete until truncated | Dilution and higher cost |
| Compaction | Summary of the session | Always present after summarizing | Replaced by new summaries | Original detail usually lost | Early details dropped |
| Persistent notes | Decisions, goals, dependencies, open tasks | Loaded or searched at task start or checkpoint | Appended or marked superseded | Depends on what the agent wrote | Stale or incorrect notes |
| Gist plus lookup (ReadAgent) | Episode gists plus original passages | Gist first, original text on demand | Not stated in the reported evaluations | Original passages retrievable | Gist misses detail the lookup never triggers |
| Knowledge-centric (PlugMem) | Structured facts and reusable skills | Task-relevant units retrieved and distilled | Not stated in the published description | Source context depends on the design | Irrelevant or missed retrieval |
No single row wins across tasks. The right choice depends on what the agent must do, how long the work runs, and whether exact wording or only the outcome of past steps matters.
Where memory fails
Storing information is the easy part. The harder questions concern whether the right information comes back at the right moment and whether it is still true.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
- Stale facts. A stored decision that was later reversed can steer the agent wrongly if the update was never recorded.
- Missed retrieval. A relevant note that is not found does nothing. Retrieval quality is as important as storage.
- Retrieval in the wrong situation. A fact can be correct and still be applied to a case it does not fit.
- Over-compression. Anthropic’s article warns that aggressive compaction can discard details whose importance emerges later, which is why it recommends preserving critical decisions and unresolved work.
- Similarity is not causality. A 2026 AMA-Bench paper argues that dialogue-only memory evaluations miss continuous agent-environment trajectories. It reports that similarity-based retrieval can weaken the capture of causal and objective information. This is the paper’s finding rather than an uncontested conclusion across the field.
- Open lifecycle questions. An AAAI Symposium Series review identifies separating memory types and managing memory across an agent’s lifetime as open problems. Vector databases are a common implementation for long-term memory, but the review treats lifecycle management as unsolved.
- Privacy and retention. Persisted task data and user information raise questions about how long it is kept, who can read it, and how it is removed. The sources cited here leave these largely open.
Deciding between a longer window and persistent memory
Use a larger context window when the work fits in one session, the material needed is known in advance, and exact wording matters for each step. Long single documents are a common example.
Use persistent memory when the work spans sessions, when the agent must recall earlier decisions and their reasons, or when most of the raw history is irrelevant to the current step.
Combine them when a single session is long and cross-session continuity also matters. Compaction or a summary can keep the live session moving, while notes or a structured store carry decisions and dependencies forward.
How to evaluate a memory setup on real work
A chat-style recall test does not tell you much about an agent that edits files, calls tools, and makes decisions over hours. The AMA-Bench authors make this point directly, arguing that the evaluation should resemble continuous trajectories of states, actions, observations, and tool outputs.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall- Run the agent on tasks that match its real workflow, including tool outputs and interruptions, not only question-and-answer exchanges.
- Check whether goals, decisions, dependencies, and the reasons for decisions survive the memory process.
- Measure whether the correct note is retrieved at the step where it matters, and whether irrelevant notes are pulled in.
- Introduce a deliberate change, such as a reversed decision, and confirm the agent uses the updated version.
- Compare against a baseline with the same model and the same task, measuring useful information delivered per token rather than total memory size.
Only a matched comparison of this kind can show whether a memory design helps your agent. Published benchmark gains indicate what the authors tested, not what your workflow will produce.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

