Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A large language model does not remember your users. Each call receives a bounded input, made up of instructions, the current conversation and whatever else your code places in the request. It returns an output and keeps nothing that your application did not store somewhere else. When an assistant appears to recall last week’s conversation, the application retrieved that information and put it back into the prompt. Continuity is something you build around the model, not a property the model has.

Why an assistant forgets between sessions

Each request to a model is an independent computation. The model reads what is in the request, produces a response and moves on. Nothing from the previous call is held inside the model unless your code sends it again. A chat window may display earlier messages on the screen, but the model only knows what the next request contains.

This is why an assistant can seem to forget a preference you stated yesterday, even though the product “remembers” your account. The conversation log, the user profile and any saved facts live in the application’s storage. The question for a builder is therefore not whether the model has memory, but which state your system persists, how it finds that state later and how it places it in front of the model at the right moment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the model actually sees on one call

Everything the model can use on a given call must fit inside its context window for that request. In practice a request is assembled from several parts: system instructions, any retrieved or persisted memory, the recent turns of the conversation and the user’s current input. Each part consumes tokens, and each token has a cost in money and latency. Memory design is largely the discipline of deciding which parts earn a place in that limited input.

How does memory work across calls? The lifecycle

AWS Prescriptive Guidance describes an agent that retrieves recent and long-term state, places that memory into the prompt, generates an output and stores new information for future tasks. Broken into steps, the lifecycle looks like this:

  1. Decide what to retain. Not every message deserves storage. Facts the user stated, outcomes of completed tasks and values that changed are typical candidates. Decide this explicitly, because storing everything produces noise that is later retrieved by mistake.
  2. Index it. Store the material in a form that can be searched later. That may be raw transcripts, summaries, structured records or a semantic index, and it is common to keep more than one.
  3. Retrieve for the new request. Combine recent state with a search over longer history, using the current user input to decide what is relevant.
  4. Read or interpret the results. Retrieved items are often fragments, dates or old values. The application or model must interpret them, including which one is current. LongMemEval treats this reading stage as a distinct part of long-term memory design.
  5. Inject into the prompt. AWS states that “the memory context is embedded into the LLM prompt, allowing the agent to reason based on both current inputs and prior knowledge.” That embedding step is the point where memory becomes model input.
  6. Update state after the response. Write new facts, record what happened and mark superseded values as replaced rather than leaving old and new versions side by side.

Each step can fail independently. A system can store the right fact and still never retrieve it, or retrieve it and present an outdated version. Testing should cover each stage, not just the final answer.

Where each kind of state lives

AWS gives illustrative service mappings for these roles. They show one architecture, and equivalent components from other vendors or self-hosted software can fill the same roles.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
State or role Purpose in the lifecycle Services named as examples in AWS guidance
Recent state Fast access to the current session or task DynamoDB, Redis or Bedrock context
Structured long-term memory Facts, relationships and values that change over time Aurora, DynamoDB or Neptune
Semantic retrieval Finding past material by meaning rather than exact wording OpenSearch or Pinecone
Transcripts and files Raw record for audit, reprocessing and verification S3
Orchestration Running the retrieve, generate and store steps in order Lambda or Step Functions
Reasoning Generating the response from the assembled prompt Bedrock

Keeping raw transcripts alongside derived memory is useful beyond storage. When a derived summary or extracted fact is wrong, the original text is the only reliable way to check it.

Which memory approach fits your product?

Compare approaches on the requirements that matter for your product: retrieval quality, handling of updates, latency, token and storage cost, auditability, user control and the risk of mixing unrelated people or domains. The four approaches below are not mutually exclusive, and most production designs combine them.

Auto-injected curated layers

This approach adds metadata, explicitly saved facts, recent summaries and the current conversation to every request. Microsoft’s architecture guidance says it can make continuity feel seamless. Its drawbacks are that it adds token cost to every call, gives users less control over what is used, and risks mixing unrelated contexts or carrying hallucinated summaries forward. It suits products where a small, stable set of facts is needed on almost every turn.

On-demand retrieval

Here the application searches stored history or structured memory only when a request needs it. This avoids injecting the full history on every call. Its quality depends entirely on indexing and retrieval surfacing the right evidence. If the relevant item is not retrieved, the model will not use it, and the failure looks like forgetfulness. LongMemEval frames long-term memory as indexing, retrieval and reading, and it evaluates errors beyond simple text recall.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Structured or extracted memory

This approach stores selected facts, relationships, task outcomes or changing state in a form the application can inspect and update. Compared with an undifferentiated transcript, structured records make it possible to replace a value when a user corrects it, to check which fact was current on a given date and to show the user what is stored. Microsoft Research’s 2026 work and the LongMemEval benchmark both motivate testing update handling, temporal reasoning and consolidation, rather than assuming a transcript is sufficient.

Full-context replay or summaries

Replaying full history gives the model everything, which makes it a useful reference baseline, but it can consume a large share of the context window. Summaries compress history and reduce cost, but they lose detail. Microsoft’s guidance warns that summaries can produce hallucinated memories, so a summary should be treated as a derived claim that can be checked against the transcript.

How to test whether memory works

LongMemEval, published at the International Conference on Learning Representations (ICLR) 2025, names five abilities that a long-term memory system should be able to show. They give a useful starting taxonomy for a test set, though not a complete production checklist:

  • Information extraction: pulling the right facts out of earlier interactions.
  • Multi-session reasoning: combining facts that appeared in different sessions.
  • Temporal reasoning: knowing when something was true and what came first.
  • Knowledge updates: using the newest value after a correction or change.
  • Abstention: declining to answer when the stored evidence is missing.

Build the test set from the questions your users actually ask, and include cases for each of the following:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Whether the relevant stored item is retrieved at all.
  • Whether a corrected or superseded value is ignored in favour of the current one.
  • Whether the system abstains when no evidence exists, rather than inventing a memory.
  • Whether information from one user, account or domain is excluded from another’s responses.
  • What the memory layer adds in tokens and latency per request, measured under realistic history lengths.

What the published figures do and do not show

Benchmark numbers describe the tested systems, datasets and comparisons. They are not guarantees for another application, and vendor-reported results are not independent validation.

Figure Source and date What it measured Limit on how it applies
500 curated questions ICLR 2025, LongMemEval Size of the benchmark Benchmark questions, not your traffic
30% accuracy drop on memorizing information across sustained interactions ICLR 2025, LongMemEval abstract Evaluated commercial chat assistants and long-context LLMs The benchmark’s reported finding for its evaluation, not a universal loss for all LLMs
97.2% retention precision and 58% store reduction Microsoft Research, May 2026 Deduplication-based consolidation on a VSCode issue-tracking dataset One dataset and one method
86.3% LLM-judge accuracy on LoCoMo and 87.4% on LongMemEval Microsoft Research, 2026, for its Memora system Memora’s evaluation on those benchmarks Publisher-reported, and scored by an LLM judge
Up to 98% fewer context tokens than full-context inference Microsoft Research, 2026, for Memora Token use compared with full-context inference on the tested system “Up to” figure for the tested system and comparisons only
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Where newer research is pointing

Microsoft Research’s May 2026 paper, “Human-Inspired Memory Architecture for LLM Agents,” describes six mechanisms: sleep-phase consolidation, interference-based forgetting, engram maturation, reconsolidation upon retrieval, entity knowledge graphs and hybrid multi-cue retrieval. These are one group’s reported methods and evaluations. They are useful as design vocabulary, not as settled practice.

Its companion work on Memora describes separating rich memory content from lightweight retrieval abstractions and cue anchors, with iterative, policy-guided retrieval. The direction is consistent across these sources: keep detailed content available, retrieve through compact cues and treat consolidation and forgetting as explicit design decisions.

Why does my assistant forget or mix things up?

Match the symptom to the most likely cause before changing models or prompts:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • It forgets facts from earlier sessions. Nothing was persisted, or the update step was skipped after the response. Check the stored records directly.
  • It repeats a value the user already changed. The system has no supersession rule, so old and new values both reach the prompt. Store changes as replacements with timestamps.
  • It brings in details that belong to someone else. Retrieval is not scoped by user, account or tenant. This is the context-mixing risk Microsoft’s guidance describes for auto-injected layers.
  • It states a memory that does not exist. A summary or extracted fact may be hallucinated. Verify it against the raw transcript and make the assistant abstain when no evidence is retrieved.
  • Costs rise as conversations get longer. The application injects too much history on every call. Move stable facts into a short curated layer and retrieve the rest on demand.
  • The stored fact exists but is not used. The retrieval step missed it. Test the index with the same queries your users send.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.