iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
None of Mem0, Zep, LangChain memory, and Letta is an evidence-backed universal winner. They remember different things: extracted conversational facts, a temporal knowledge graph, thread or cross-thread application state, or editable files owned by an agent. Choose by the kind of information your agent must retain and update—not by the word “memory” or a score from a benchmark that did not test all four systems on equal terms.
What does “memory” mean in this comparison?
An agent can appear to remember because its current prompt contains earlier messages, because an application retrieved a saved fact, or because a runtime maintains persistent state. Those mechanisms have different scopes and responsibilities. Before selecting a product, decide whether you need continuity within one conversation, useful facts carried between conversations, relationships and facts that change over time, or an agent that can manage its own persistent memory.
- Thread state: information retained so an agent can continue a particular conversation.
- Durable user or conversation facts: selected information extracted from interactions and made available later.
- Temporal relationships: entities, facts, and connections represented with attention to how they change.
- Agent-owned memory: persistent content that belongs to a stateful agent and can be inspected or edited as part of its workflow.
These categories can overlap in a deployed application, but they are not interchangeable implementations.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11How the four approaches differ
| Option | What it retains and how | Where it fits | Main implementation responsibility |
|---|---|---|---|
| Mem0 | A dedicated memory layer that extracts, consolidates, and retrieves salient information from conversations. Its paper also describes a graph-based variant for relationships among conversational elements. | Adding extracted, durable conversational information to an existing application. | Check what is extracted, how corrections and updates behave in your workload, whether retrieval is relevant, and whether hosting and data controls fit your requirements. |
| Zep | The default documentation describes v3 as a governed context layer built around temporal Context Graphs and a Context Lake that manages and serves graphs. A separate v2 Memory API page describes a session-based interface that builds a user-level knowledge graph. | Applications where evolving facts, relationships, provenance, or enterprise governance are central. | Confirm the documentation and API version you are implementing; assess governance and retrieval in your own deployment. |
| LangChain memory | Current LangChain documentation distinguishes short-term, thread-level agent state persisted by a checkpointer from long-term, user-specific or application-level data across threads and sessions. | Teams already using LangChain or LangGraph that want framework-native state and persistence choices. | Choose what to persist and how to scope it; design cross-session storage and retrieval when long-term recall is required. |
| Letta | Its memory documentation describes MemFS, a git-backed memory filesystem an agent can inspect and edit, shared across that agent’s conversations. Optional “dreaming” uses background subagents to review recent conversations and update memory. | Teams that want memory to be an inspectable, editable part of a stateful agent runtime. | Adopt and operate Letta’s agent-centric memory workflow rather than treating it as a drop-in persistence component for any stack. |
Mem0: extracted memory for an existing application
Mem0 is the clearest fit when you want a dedicated memory layer to identify useful information in conversations and make it available to an application later. Its paper presents a memory-centric architecture that dynamically extracts, consolidates, and retrieves salient conversational information, and also describes a graph-based variant. The paper’s evaluation results are author-reported and specific to its tested setup; they do not establish independent superiority for every application. See the Mem0 introduction and the Mem0 paper.
#1 Best Overall
For a real selection, test whether the information your users care about is actually retained, whether corrections replace or qualify earlier facts as intended, and whether the result can be retrieved with the right context. Also review the service’s hosting and data controls against your application’s requirements; the documented architecture alone does not establish how it will behave on your traffic.
Zep: temporal context, with an important version distinction
Zep’s default documentation identifies v3 and describes a governed enterprise context layer using temporal Context Graphs for entities, relationships, and facts, alongside a Context Lake that manages and serves those graphs. That architecture is relevant when a flat collection of saved preferences is not enough—for example, when an application needs to reason about connected facts that may change over time. These are product-documentation descriptions, not an independent assessment of retrieval quality or governance in a particular deployment. See the Zep v3 overview.
Rank #2
Zep also publishes a separate v2 Memory API page. That page describes sending chat messages by session, building a user-level knowledge graph, and retrieving relevant context from recent messages that can come from any session belonging to that user. Treat those details as v2-specific; they should not be assumed to describe the current v3 interface. See the Zep v2 Memory API.
Recommended Free Tools
LangChain memory: thread persistence is not cross-session recall
“LangChain Memory” is ambiguous unless you specify which layer you mean. Current LangChain documentation describes short-term memory as thread-level state: an agent’s state can include conversation history, and a checkpointer persists it so the thread can resume. For production use, the documentation recommends a database-backed checkpointer. It describes long-term memory separately, for user-specific or application-level information across threads and sessions. See LangChain’s short-term memory documentation.
Rank #3
This framework-native approach is a natural starting point when your application already uses LangChain or LangGraph and the immediate need is reliable continuity within a thread. It does not by itself answer what should be remembered across separate conversations: the developer still chooses the stored information, its scope, and how cross-session retrieval works. It is therefore misleading to treat older BaseMemory classes as the whole current LangChain memory model.
Letta: memory belongs to the stateful agent
Letta takes an agent-centric approach. Its current memory documentation describes MemFS as a git-backed filesystem that an agent can inspect and edit, with memory shared across that agent’s conversations. It also documents optional “dreaming”: background subagents review recent conversations, consolidate useful lessons, and update memory. See Letta’s Memory & dreaming documentation.
Rank #4
This model is worth evaluating if inspectable, editable, versioned memory is part of how you want agents to work. The trade-off is architectural commitment: Letta presents memory as part of its agent runtime and workflow, not as a generic persistence plug-in for arbitrary existing agents.
What published benchmark results can—and cannot—tell you
The available paper results are useful evidence about particular experiments, not a matched four-way test of current products. Their datasets, models, baselines, and methods differ, so the figures below should not be compared as if they came from one leaderboard.
Zep paper results
In its paper dated 2025-01-20, Zep’s authors report 94.8% DMR accuracy with gpt-4-turbo, compared with 93.4% for MemGPT as reported by the MemGPT team. The paper also reports 94.4% for its full-conversation baseline and 78.6% for its conversation-summary baseline under that model. Using gpt-4o-mini, it reports 98.2% for Zep and 98.0% for full conversation. The authors describe DMR as limited, including single-turn fact-retrieval questions and ambiguous question wording. These are the paper’s experiments, not a neutral comparison of the four options in this article. See the Zep paper.
Mem0 paper results
In its paper dated 2025-04-28, Mem0’s authors report a 26% relative improvement in their LLM-as-a-Judge metric over OpenAI and around a 2% higher overall score for the graph-memory variant than for its base configuration on the paper’s LOCOMO evaluation. The paper also reports 91% lower p95 latency and more than 90% token-cost savings versus its full-context method. Each figure belongs to the paper’s own setup and comparison; it cannot be directly ranked against the Zep paper’s results, which use different methods and baselines. See the Mem0 paper.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to choose for your workload
Use the architecture that matches the information you need to keep and the degree of control you want over it. These are fit-based recommendations inferred from the documented approaches, not claims from hands-on testing.
- Choose LangGraph state first if your main requirement is thread continuity inside a LangChain/LangGraph application and you want to control the persisted state.
- Evaluate Mem0 if you want a dedicated layer to extract and retrieve durable conversational information for an existing application.
- Evaluate Zep if changing facts, graph relationships, provenance, and governed enterprise context are central requirements.
- Evaluate Letta if you want to use an agent runtime where editable, inspectable memory is a first-class part of agent state.
Run a matched test before choosing on performance
A meaningful “which remembers better?” test should use representative tasks and hold the important conditions constant. Include more than a clean fact lookup: exercise updates, errors, deletion, and operational constraints.
- Define the memory task. Write down whether success means continuing a thread, recalling a stable user fact, tracking a changing relationship, or exposing agent-editable memory.
- Prepare the same interaction history. Include facts that become stale, explicit corrections, multi-hop questions, and separate sessions where relevant.
- Keep evaluation conditions comparable. Use the same model, retrieval budget, question set, privacy constraints, and latency and cost measurement rules for each approach.
- Score the outcomes that matter. Check whether the right fact is retrieved, whether an outdated answer is avoided, how ingestion and retrieval affect latency, and what ongoing model and storage costs look like.
- Validate operations and governance. Check data boundaries, permissions, audit or provenance needs, deletion behavior, deployment choices, and who owns persistence in your design.
No independent, matched benchmark across all four current offerings is established by the cited papers and documentation. A test built around your own task set is the evidence needed to turn architectural fit into a workload-specific decision.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

