Skip to main content

Best AI Agent Memory Frameworks in 2026: Mem0, Zep, LangChain, Letta Compared

Emily Winks, Data Governance Expert, Atlan
Data Governance Expert
Updated:
|
Published:
35 min read

Key takeaways

  • Every LongMemEval and LoCoMo score here is a vendor measuring its own system. None are independent.
  • Choose by use case: Mem0 for personalization, Zep for temporal reasoning, LangMem for LangChain, Letta for long-running.
  • All 8 frameworks lack enterprise governance: no glossary, lineage, or entity resolution — a context layer fills the gap.

What are the best AI agent memory frameworks in 2026?

AI agent memory frameworks give agents persistent context across sessions — storing conversation history, user preferences, and learned facts so agents don't restart from zero every time. We evaluated 8 frameworks on memory architecture, persistence model, multi-agent coordination, self-hosting support, enterprise authentication, and the benchmark figures each vendor publishes. Read those figures carefully: every score in this category was produced by the vendor selling the system, on its own test configuration, and no third party has reproduced any of them.

Core components

  • Mem0 - open-source, 65,589 GitHub stars, best for chatbot and personal assistant memory
  • Zep - production-grade, hybrid vector+graph, strong for long-running agent sessions
  • LangChain Memory - modular, integrates with LangChain ecosystem, multiple memory strategies
  • Letta (MemGPT) - tiered memory architecture, self-editing memory, strong for complex agents

Is your data ready for AI agents?

AI agent memory frameworks are the infrastructure layer that lets AI agents persist, retrieve, and reason over information across sessions. They answer the questions: Does this agent know who it’s talking to? Does it remember what was decided last week? Does it know what it has already learned? The leading frameworks — Mem0, Zep (Graphiti), LangGraph (LangMem), Letta (formerly MemGPT), and LangChain’s ConversationBufferMemory — each take a different architectural approach: vector-based semantic recall, temporal knowledge graphs, checkpoint-based persistence, or self-editing memory blocks.

Agent Memory Framework Picker


Give it your setup and constraints. It narrows the field to the frameworks that fit, and says why the rest do not. Read the skill.

Paste into a new chat

Use the skill at https://atlan.com/skills/agent-memory-framework-picker.md to pick an agent memory framework for our setup. Ask me for whatever it needs.

Run once in a terminal

curl -fsSL --create-dirs \
  -o ~/.agents/skills/agent-memory-framework-picker/SKILL.md \
  https://atlan.com/skills/agent-memory-framework-picker.md

For an agent

curl -fsSL https://atlan.com/skills/agent-memory-framework-picker.md

The field has matured considerably. Three memory scopes have become standard: episodic (specific past interactions), semantic (facts and preferences), and procedural (learned behaviors and rules). Two delivery models dominate: managed cloud services and self-hosted open source. But the gap between frameworks has widened too, and the published scores are a poor guide to it. Every LongMemEval and LoCoMo figure in this category was produced by the vendor selling the system, and the enterprise governance requirements that most frameworks still haven’t addressed are becoming impossible to ignore.

This comparison evaluates 8 frameworks on architecture, persistence model, multi-agent coordination, self-hosting support, enterprise auth, and benchmark accuracy.


Quick Facts

Frameworks reviewed 8
Top star count ~100K (LangChain ecosystem)
Highest funded Mem0 ($24M across seed and Series A)
Independently verified benchmarks 0 of 8
Self-hosted options 7 of 8
SOC 2 on an enterprise tier 2 (Mem0, Zep)

At a glance: all 8 frameworks compared

Framework Architecture Persistence Multi-agent Self-host Enterprise Auth Pricing (entry)
Mem0 Hybrid (vector + graph + KV) Yes Partial (scoped) Yes Yes (Enterprise tier) Free / $19/mo
Zep / Graphiti Temporal knowledge graph Yes Partial Yes (OSS) Yes (Enterprise) Free / usage-based
LangChain / LangMem Modular (pluggable backends) Yes (via LangGraph) Via LangGraph Yes Via LangSmith/Azure Free (OSS)
Letta / MemGPT OS-inspired tiered (core/archival/recall) Yes Yes (native) Yes (OSS) Partial Free (OSS) / usage-based
MS Semantic Kernel / Kernel Memory RAG pipeline + Vector Store Yes Partial Yes Yes (Azure IAM) Free (OSS) / Azure costs
Cognee Poly-store (graph + vector + relational) Yes Partial Yes (local-first) No Free (OSS)
Supermemory Memory API (cloud + OSS) Yes Partial Yes Not confirmed Free tier
Redis Agent Memory Server In-memory + vector search Yes No native Yes Via Redis Cloud Free (OSS) / ~$0.07/GB/hr


What makes the best AI agent memory framework?

The right framework depends on what kind of memory your agent actually needs. An agent that must remember a user’s tone preferences needs different infrastructure than an agent tracking how a customer relationship has changed over six months of interactions.

Memory architecture matters more than star count. A vector-only store retrieves semantically similar facts but cannot model how facts change over time. A temporal knowledge graph tracks validity windows: when a fact was true, when it was superseded. No published benchmark settles which architecture wins, because every score in this category is a vendor’s measurement of its own system. Before evaluating features, know which problem you’re solving.

Evaluation criteria used in this comparison


Architecture — vector-only vs. hybrid vs. temporal knowledge graph. This determines what kinds of queries your agent can handle accurately, not just how many facts it can store.

Persistence model — how memory survives across sessions. Whether it’s managed cloud (someone else’s infrastructure) or a self-hosted storage backend with your own data residency requirements matters for regulated industries especially.

Multi-agent coordination — whether agents can share a memory pool without polluting each other’s state. Scoped memory (per user, per session, per agent) is the standard approach; true native multi-agent coordination is rarer.

Self-hosting support — open source availability, data residency requirements, and dependency footprint. For teams with air-gapped requirements or GDPR constraints, this is often the first filter.

Enterprise auth — SSO, RBAC, and audit logging. What comes built into the framework versus what you must configure through your cloud provider is a meaningful operational distinction.

Benchmark accuracy — LongMemEval and LoCoMo scores where published, read with the provenance attached. In Zep’s own paper, Zep scored 71.2% on LongMemEval with GPT-4o against a 60.2% full-context baseline; that paper never evaluates Mem0 [arXiv 2501.13956]. Mem0 self-reports 94.4% on LongMemEval for its managed platform. The two figures come from different papers, different models and different judges, so they do not form a head-to-head.


The 8 best AI agent memory frameworks at a glance

  1. Mem0: best managed, drop-in memory API for personalization agents
  2. Zep / Graphiti: best for agents that reason about how facts change over time
  3. LangChain / LangMem: best for teams already running on LangChain/LangGraph
  4. Letta / MemGPT: best for long-running agents that need OS-level memory management
  5. Microsoft Semantic Kernel / Kernel Memory: best for Azure-native enterprise shops
  6. Cognee: best for local-first, privacy-critical deployments with graph reasoning
  7. Supermemory: best for coding agents (Claude Code, OpenCode integrations)
  8. Redis Agent Memory Server: best as a low-latency storage backend for teams already running Redis

1. Mem0: in practice

Best managed memory API for personalization agents

Mem0 gives agents a three-tier memory system: user, session, and agent scopes, backed by a hybrid store combining vectors, graph relationships, and key-value lookups. When facts conflict, Mem0 self-edits rather than appending duplicates, keeping memory lean. At 65,589 GitHub stars and $24M raised across seed and Series A, it has the largest developer community of any standalone memory framework. One caveat that shapes every other line in this section: Mem0 has removed graph memory from the open-source SDK, and it is now a paid Platform feature only.

Official site: mem0.ai | GitHub: mem0ai/mem0 (65,589 stars, 18 Sep 2026) | Docs: docs.mem0.ai

Pros


  • Largest community of any standalone memory tool (65,589 stars, 18 Sep 2026)
  • Self-editing memory eliminates duplicate entries without manual deduplication logic
  • Mem0 describes itself as SOC 2 and HIPAA ready; SSO, audit logs and on-prem are Enterprise-tier
  • First-party MCP server at mcp.mem0.ai, plus an OpenAI-compatible API surface
  • 27 documented vector-store backends, including Chroma, Qdrant, Pinecone and Weaviate

Cons


  • Graph memory is paywalled at $249/mo, and Mem0 removed it from the open-source SDK entirely
  • No temporal fact modeling. Memories are timestamped at creation, but there is no validity window or fact supersession
  • Multi-agent shared memory requires custom implementation; it is not native
  • The 92.5% LoCoMo and 94.4% LongMemEval scores Mem0 publishes describe the managed platform, which Mem0 says carries proprietary optimisations the SDK does not

Key Capabilities


Mem0 stores memories across three isolated scopes: user-level (preferences and history), session-level (current conversation context), and agent-level (agent-specific knowledge). The self-editing model resolves conflicts on write — when a user corrects a preference, Mem0 updates the existing record rather than creating a duplicate. Graph memory adds relationship modeling on top of the vector and KV stores, and it now runs only on the Platform: Mem0’s own docs state that graph memory is a Platform feature, and the external graph drivers (Neo4j, Memgraph, Kuzu, Apache AGE, Neptune) came out of the open-source SDK. REST API plus Python and TypeScript SDKs cover most integration paths.

Mem0 supports multi-LLM backends including OpenAI, Anthropic, Gemini, and Groq. Its first-party MCP server makes it accessible from Claude Code and similar agentic environments. Enterprise tier adds on-prem deployment, SSO, a dedicated SLA, and audit logs.

The framework does its stated job well: personalization memory for consumer-facing agents and B2B copilots. Its limit is the absence of a temporal model. Memories are stored and retrieved, not modeled as time-bounded facts that can be superseded. For agents that need to reason about how things changed, that is a meaningful gap.

Pricing


  • Hobby: free, 10K add requests and 1K retrievals per month
  • Starter: $19/month, 50K add requests, no graph memory
  • Pro: $249/month, 500K add and 50K retrieval requests, plus graph memory and Dream memory consolidation
  • Enterprise: custom, and the only tier Mem0 describes as unlimited. Adds on-prem deployment, SSO and audit logs

Source: mem0.ai/pricing, read 18 September 2026. The units are add requests, not stored memories.


2. Zep / Graphiti

Best for agents that need to reason about changing facts over time

Zep stores every fact as a knowledge graph node with a validity window. “Kendra loves Adidas shoes (as of March 2026)” is not just a stored string, it is a fact with a temporal bound. When new information contradicts old, Graphiti invalidates the old without discarding the historical record. In Zep’s own paper, Zep scored 71.2% on LongMemEval with GPT-4o against a 60.2% full-context baseline, and 63.8% against a 55.4% baseline on GPT-4o mini. That paper evaluates Zep against full context, not against another vendor.

Official site: getzep.com | GitHub: getzep/graphiti (30,985 stars, 18 Sep 2026) | Docs: help.getzep.com

Pros


  • Best temporal reasoning of any reviewed framework, purpose-built for “how did this fact change over time”
  • P95 retrieval latency ~300ms with no LLM calls at query time (hybrid semantic + BM25 + graph traversal)
  • Graphiti open-source for self-hosting on Neo4j, FalkorDB or Amazon Neptune; SOC 2 Type II, HIPAA BAA and the EU DPA are Enterprise-tier items, not general Zep Cloud properties
  • Can integrate structured business data (JSON objects) alongside conversation history
  • Repositioned as a context engineering platform (v3 SDK, 2025), signals broader scope than session memory

Cons


  • Zep’s managed service no longer runs Graphiti. Per Graphiti’s README, Zep production uses a proprietary Context Graph Engine, so the OSS engine and the hosted product have diverged
  • No constitutional layer. Zep stores whatever is ingested with no validation that referenced entities are authoritative or governance-restricted
  • Credit-based pricing with no published credits-per-operation rate, so a true cost comparison against request-priced rivals is not computable from vendor sources
  • Still an interaction and business data memory store. The temporal graph tracks ingested facts, not live enterprise data estate governance state

Key Capabilities


Graphiti is Zep’s open-source temporal knowledge graph framework. It stores facts as nodes with start and end validity windows, with entity resolution that tracks the same entity across both unstructured conversation data and structured business records. Hybrid retrieval combines semantic embeddings, BM25 keyword search, and direct graph traversal, without requiring LLM inference at query time.

The distinction between Zep’s context graph approach and a simple vector store matters most for agents handling temporal queries: “How did this customer’s behavior change after the pricing update?” or “What was the revenue metric before the finance team revised the calculation?” Flat vector stores retrieve the most recent or most similar entry. Temporal graphs retrieve the fact that was valid at the time being queried.

Zep also integrates structured JSON business data objects alongside conversation history, meaning agents can incorporate operational data from CRM exports, transaction logs, and external data sources into the same memory graph.

Pricing


Credit-based billing. Per Zep’s pricing page, read 18 September 2026: a free tier at 10,000 credits a month across 2 projects, Flex at $125/month with 50,000 credits included and $25 per additional 10,000, Flex Plus at $375/month, and Enterprise on consultation. Context Graph and temporal memory are included on every paid tier. SOC 2 Type II, the HIPAA BAA and the EU DPA sit on Enterprise only.


3. LangChain / LangMem

Best for teams already committed to the LangChain ecosystem

LangChain’s LangMem SDK adds three memory types to LangGraph agents: episodic (past interactions), semantic (facts and preferences), and procedural (agents rewriting their own system instructions based on feedback). If your team already runs LangChain, LangMem is the path of least resistance. If you’re not on LangChain, the ecosystem coupling cost is high.

Official site: langchain.com | GitHub: langchain-ai/langchain (~100K stars) | LangMem docs: langchain-ai.github.io/langmem

Pros


  • Already in the stack for most LangChain teams, zero new dependency to add long-term memory
  • Procedural memory is architecturally unique: agents update their own operating instructions based on user feedback
  • Pluggable storage backends (any vector DB, MongoDB, Postgres via pgvector, etc.)
  • Largest AI framework community by contributor count (~100K LangChain stars)
  • MIT license; free to run

Cons


  • Tightly coupled to LangChain/LangGraph — standalone use is impractical; adds framework lock-in
  • No built-in temporal reasoning; no fact validity windows
  • Graph memory not native — requires external integration
  • No managed memory hosting — your team runs its own infrastructure
  • LangChain API churn (memory APIs changed across v0.1, v0.2, v0.3) creates real maintenance burden

Key Capabilities


LangMem supports three memory types built on top of LangGraph’s persistent StateGraph store layer. Episodic memory records specific past interactions and can distill them into few-shot examples. Semantic memory stores general facts about users or the world. Procedural memory, the genuinely novel capability, allows agents to update their own system prompt instructions based on accumulated user feedback. Agents learn what works and modify their own operating rules.

Memory is namespaced by user_id, team_id, or app_id, preventing cross-contamination between users and sessions. Background memory extraction runs after conversations, extracting and updating memories without blocking agent execution. Storage backends are pluggable: any store that implements the LangGraph store interface works, including MongoDB, Postgres via pgvector, and in-memory stores for prototyping.

The lock-in cost is real. LangMem is tightly bound to LangChain’s data structures and abstractions. If your team is not already on LangGraph, adopting LangMem means adopting LangGraph too. There is no managed memory hosting: your team configures and operates the storage backend.

Pricing


LangMem SDK: free (MIT). LangSmith (observability and tracing): free tier, $39/mo Developer, $259/mo Plus, Enterprise custom. LangGraph Platform (managed deployment) has separate pricing.


4. Letta / memgpt: in practice

Best for long-running agents that actively manage their own memory

Letta (formerly MemGPT, UC Berkeley) treats agents as active memory managers, not passive recipients. Its current documentation describes four tiers ordered by scale and importance: memory blocks, files, archival memory and external RAG. Agents decide what to keep close, what to archive and what to search. Read this section with one date in mind: on 16 March 2026 Letta began sunsetting its server-side memory tools, templates, identities and MCP integrations, moving to git-backed memory files and client-side orchestration.

Official site: letta.com | GitHub: letta-ai/letta (24,786 stars, 18 Sep 2026) | Docs: docs.letta.com

Pros


  • Most architecturally distinctive approach: agents as active participants in their own memory management, not passive recipients
  • Full retrieval depth (graph + temporal) available even at the free self-hosted tier, no paywall
  • Complete agent platform with state management, tool calling, and multi-agent coordination built in
  • Strong academic research foundation (MemGPT paper, UC Berkeley, 2023)
  • Letta Code launched in December 2025 and became the flagship in March 2026
  • Letta’s own test puts its filesystem agent at 74.0% on LoCoMo with GPT-4o mini (Letta, Aug 2025)

Cons


  • Not a drop-in memory component. Adopting Letta means adopting its full agent runtime
  • The March 2026 pivot deprecates server-side memory tools, templates, identities and MCP integrations, so a hosted tiered-memory deployment is building on a path Letta is walking away from
  • Smaller community than Mem0 or LangChain (24,786 stars against Mem0’s 65,589)
  • Agent-manages-own-memory paradigm requires careful design to avoid runaway context drift
  • No native enterprise governance layer; no business glossary, lineage, or policy enforcement

Key Capabilities


Letta’s documented hierarchy runs four tiers, ordered by scale of data and importance rather than by a storage metaphor: memory blocks, files, archival memory and external RAG. The RAM-and-disk framing belongs to the MemGPT paper, not to Letta’s current docs. Memory blocks are always in-context and are edited with memory_rethink, memory_replace and memory_insert. Archival memory is an external store the agent queries explicitly with archival_memory_search and writes with archival_memory_insert.

Agents use explicit memory management function calls to move information between tiers, deciding what is important enough to keep in-context versus what gets archived. This is a genuinely different paradigm: the agent is not just a consumer of retrieved context, it is an active curator of its own knowledge base.

Multi-agent coordination is native: Letta agents can call sub-agents and pass state between them. All four retrieval strategies, including graph and temporal, are available at every tier, including the self-hosted free version. The Agent Development Environment (ADE) provides visual tooling for inspecting and debugging agent memory state.

Pricing


Per Letta’s pricing docs, read 18 September 2026:

  • Free: $0
  • Pro: $20/month, up to 20 stateful agents
  • API plan: $20/month plus usage; BYOK supported
  • Teams Pro: $20/seat. Enterprise: custom
  • Self-hosted: free. Both letta-ai/letta and letta-ai/letta-code are Apache-2.0

5. Microsoft semantic kernel / kernel memory

Best for Azure-native enterprise development teams

Microsoft Semantic Kernel and Kernel Memory form the memory backbone for Azure-native AI agents. Kernel Memory handles ingestion, chunking, embedding, and retrieval as a standalone microservice. Vector Store connectors link to Azure AI Search, Qdrant, Redis, and more. With 27K+ GitHub stars and tight Microsoft 365 / Copilot integration, this is the default choice for .NET enterprise shops, provided you’re already running Azure.

Official site: learn.microsoft.com/semantic-kernel | GitHub: microsoft/semantic-kernel (~27K stars) | Docs: learn.microsoft.com/semantic-kernel/concepts/vector-store-connectors

Pros


  • Natural fit for Azure / Microsoft 365 / Copilot organizations — no new cloud relationship required
  • Enterprise-grade access control via Azure IAM out of the box
  • Multi-language SDK (C#, Python, Java) for .NET enterprise development teams
  • Azure Monitor integration provides audit logging within the Azure ecosystem
  • Kernel Memory provides a production-ready RAG pipeline, not just a vector store wrapper

Cons


  • Azure ecosystem lock-in is significant; non-Azure deployments are possible but not the primary use case
  • Memory architecture is document and RAG-centric, not conversation or agent-centric — better for knowledge retrieval than stateful agent memory
  • ISemanticTextMemory deprecated in October 2025; teams on older codebases face migration burden
  • No temporal reasoning; no fact validity windows; no graph-based memory
  • Memory governance is only as strong as your Azure configuration — it is not built into the memory layer itself

Key Capabilities


Semantic Kernel is an orchestration framework with Vector Store abstractions that connect to Azure AI Search, Qdrant, Chroma, Pinecone, Redis, and other backends. Kernel Memory is a separate standalone microservice that handles the full ingestion pipeline: OCR, document chunking, embedding generation, and indexing, exposing it as a callable function within Semantic Kernel.

In October 2025, Microsoft merged Semantic Kernel and AutoGen into the unified Microsoft Agent Framework (MAF). Vector Store abstractions replaced the older ISemanticTextMemory across all new documentation. Azure AI Foundry integration deepened for enterprise RAG pipelines in Q1 2026.

For teams inside the Microsoft ecosystem, the auth story is genuinely strong: Azure Active Directory SSO, RBAC via Azure IAM, and audit logging via Azure Monitor come without additional configuration. For everyone outside that ecosystem, the lock-in cost is high and the absence of temporal or graph memory means the framework is better suited to document retrieval than evolving agent memory state.

Pricing


Open source (MIT). Costs come from Azure services consumed: Azure OpenAI, Azure AI Search, Azure Blob Storage, billed at standard Azure rates. No separate Semantic Kernel licensing cost.


6. cognee

Best for local-first, privacy-critical deployments with graph reasoning

Cognee ingests documents, code and app data into a knowledge graph you can query and improve over time, combining vector search with multiple graph database backends (Neo4j, FalkorDB, KuzuDB, NetworkX) and relational metadata in a poly-store design. It runs completely offline via Ollama, no cloud dependency required. The Memify Pipeline runs background enrichment continuously, adding semantic associations and pruning stale data without manual curation.

Official site: cognee.ai | GitHub: topoteretes/cognee (30,809 stars, 18 Sep 2026) | Docs: docs.cognee.ai

Pros


  • Poly-store flexibility. Swap the graph DB, vector DB, or relational layer independently without changing the API
  • Simplest onboarding of any graph-capable tool. The v1.0 API is four operations: .remember, .recall, .improve and .forget
  • 100% local deployment, and Cognee’s own words for it are “free, forever”. Runs on commodity hardware via Ollama
  • Cognee reports 5M+ SDK runs per month and names Bayer as a production customer in its own case study
  • Background Memify Pipeline reduces manual knowledge curation burden

Cons


  • Time Awareness (temporal) feature is new and less proven than Zep’s temporal knowledge graph
  • The v1.0 API rename means add, cognify, search and memify are now documented as legacy, so older code samples and tutorials describe a superseded surface
  • Cognee’s docs publish no SOC 2 or HIPAA certification, so a regulated-industry deployment needs its own compliance review
  • Enterprise is BYOC with 6, 12 and 24-month engagement packages including a proprietary runtime, so the fully-open framing has a commercial ceiling

Key Capabilities


Cognee’s API makes graph-backed memory more accessible than any other tool in this comparison. Cognee v1.0 exposes four operations: .remember ingests and builds, .recall queries via vector similarity, graph traversal or both, .improve promotes session memory into the permanent graph, and .forget removes it. The older add, cognify, search and memify calls are still documented, under legacy operations.

The poly-store architecture means you can run Neo4j for complex graph queries, swap to FalkorDB for performance characteristics, or use NetworkX for in-process development, without rewriting application code. The relational layer (SQLite or Postgres) holds metadata and lightweight structured state.

Memify Pipeline runs background enrichment on existing knowledge, cleaning stale relationships, adding semantic associations between new and existing data, and weighting frequently-accessed facts. Time Awareness, added in 2025, captures and reconciles temporal context, though this feature is newer and less battle-tested than Zep’s temporal graph.

For teams with strict data residency requirements or air-gapped environments, Cognee’s fully local deployment is a genuine differentiator, and Cognee states the boundary itself: run the full memory engine locally or on your own stack, free, forever. The trade-off is the absence of published compliance certifications.

Pricing


Token-metered, not seat- or request-metered, per Cognee’s pricing page, read 18 September 2026. Self-hosted open source is free forever. Managed: Free at 1M tokens, Standard at $1.00 per 1M tokens plus $5 per extra workspace, and Enterprise as BYOC with 6, 12 or 24-month engagement packages.


7. supermemory

Best for coding agents and MCP-native integrations

Supermemory provides a single memory API covering fact extraction, user profile building, contradiction resolution, and selective forgetting. It claims benchmark leadership on LongMemEval, LoCoMo, and ConvoMem, though these claims are self-reported as of late 2025 and have not been independently verified. Its MCP server and plugins for Claude Code and OpenCode make it the most purpose-fit option for coding agent memory workflows in 2026.

Official site: supermemory.ai | GitHub: supermemoryai/supermemory | Docs: docs.supermemory.ai

Pros


  • MCP-native: purpose-built integrations with Claude Code, OpenCode, and OpenClaw
  • Explicit forgetting mechanism — handles memory expiration, a feature most frameworks omit
  • Self-reported benchmark leadership across LongMemEval, LoCoMo, and ConvoMem (third-party verification pending)
  • Open source plus managed cloud options
  • Browser extension for personal knowledge management alongside agent use

Cons


  • Benchmark claims are self-reported; independent third-party verification has not been published at time of writing
  • Younger product with fewer enterprise production deployments
  • Supermemory’s docs publish no SOC 2 or HIPAA certification
  • Smaller adoption signal than Mem0, Zep, or LangChain

Key Capabilities


Supermemory wraps memory management: extraction, profile building, contradiction resolution, and forgetting, behind a single API surface. The explicit forgetting mechanism is genuinely notable. Most frameworks handle addition and deduplication but not deletion by design. Supermemory treats memory expiration as a first-class operation, not an edge case.

MCP server integration enables native memory access from Claude Code and OpenCode without custom integration work. This is a specific advantage for coding agent workflows where memory context (what files were touched, what the user prefers, what was already tried) needs to persist across sessions without developer tooling overhead.

The browser extension adds personal knowledge management on top of agent memory, useful for teams that want a unified memory surface across their tools, though enterprise governance is not addressed.

Pricing


Free tier available. Pro and Enterprise tiers; specific pricing not publicly disclosed.


8. Redis agent memory server

Best as a low-latency storage backend for teams already running Redis

Redis Agent Memory Server separates working memory (current session, sub-millisecond in-memory retrieval) from long-term memory (cross-session vector search via RediSearch VSS). Redis is 20+ years of production-proven infrastructure. If your team already runs Redis, adding agent memory is an infrastructure extension rather than a new dependency. But Redis is the plumbing, not the memory framework — and that distinction matters for scoping what you’re actually buying.

Official site: redis.io/redis-for-ai | GitHub: redis/agent-memory-server (~1K stars) | Docs: redis.io/docs

Pros


  • Sub-millisecond in-memory latency for working memory — fastest retrieval of any option reviewed
  • Battle-tested infrastructure with 20+ years of production reliability
  • Works as a storage backend for Mem0, LangMem, and Kong AI Gateway — composable with existing stacks
  • Flexible deployment: Redis Cloud (managed) or Redis Stack (self-hosted)

Cons


  • Not a memory framework — Redis is infrastructure; memory management logic (extraction, deduplication, summarization, graph reasoning) must come from a layer above it (Mem0, LangMem, etc.)
  • No graph memory; no temporal fact modeling
  • In-memory storage is bounded by Redis cluster size; can be expensive at scale for long-term memory workloads
  • No built-in memory management logic at all

Key Capabilities


Redis Agent Memory Server operates on two tiers. Working memory stores current session events in Redis in-memory store, retrieval is sub-millisecond, making it useful for within-session context where latency matters. Long-term memory uses Redis vector search (RediSearch VSS) for cross-session persistence via semantic similarity retrieval.

Redis integrates natively with LangChain, LangGraph, LiteLLM, Mem0, and Kong AI Gateway, making it composable as a backend beneath a full memory framework rather than a standalone memory solution. If your team uses Mem0 or LangMem and needs a self-hosted storage backend with deterministic latency characteristics, Redis is the natural choice.

The important constraint: Redis provides no memory management logic. It stores and retrieves. The extraction, deduplication, summarization, and reasoning must come from a framework layer above it. Evaluate Redis as infrastructure, not as a memory framework.

Pricing


Redis Cloud: free tier (30MB), paid plans starting at ~$0.07/GB/hr. Redis Stack: free (self-hosted).


What none of these do: the shared enterprise governance gap

Every framework reviewed solves the same problem: giving AI agents the ability to remember what happened in their interactions, and optionally enriching that with user preferences or structured facts ingested into the memory store. This is genuinely useful for chatbots, coding assistants, and personal productivity agents.

The evaluation surfaced a consistent pattern across all 8 tools. Not one is designed for what enterprise data agents actually need.

Business glossary. No tool connects stored memories to governed business term definitions. When an agent stores “revenue was $8.4M in Q4,” there is no mechanism to attach which revenue definition was used, pre-returns or post-returns, which calculation methodology, who certified it. Facts are stored as strings or embeddings, not as semantically governed assertions tied to authoritative definitions.

Data lineage. No tool tracks where the data underlying a stored memory came from, through what transformations it passed, or how fresh it is. Memory is stored based on what an agent received in context, but the provenance of that context (which table, which pipeline, which model) is invisible. Audit-traceable AI reasoning requires lineage. None of these frameworks provide it.

Governance policy enforcement. Zep has SOC 2. Azure IAM exists in Semantic Kernel. But none of them prevent an agent from retrieving governance-restricted information across user boundaries, enforcing data retention policies on memory contents, or applying GDPR deletion requirements to facts stored in the memory pool.

Multi-platform entity resolution. Zep and Cognee both perform entity resolution within ingested data. This is not the same as resolving that account_id in Salesforce, org_id in Stripe, and tenant_id in Zendesk are the same company. Memory tools operate on what you give them; they do not connect to the live enterprise data estate to understand cross-system entity identity.

Certified asset status. No tool distinguishes between an agent’s recalled fact and a certified, board-approved metric definition. All stored memories are epistemically equivalent — there is no quality tier, no endorsement mechanism, no concept of authoritative versus unverified knowledge.

Regulatory memory governance. GDPR, CCPA, HIPAA, and SOX apply to data used by AI agents, including data stored in memory. Most frameworks treat memory as a technical cache, not as a governed data asset subject to deletion schedules and retention policies. 76% of organizations report governance frameworks lag AI adoption — and most memory frameworks are not built to close that gap.

Cross-agent institutional memory with governance. In multi-agent systems where dozens of agents write to a shared pool, without governance, memory becomes an append-only store polluted by inconsistent assertions. None of the frameworks reviewed provide a mechanism to resolve conflicts between memories from different agents, apply trust levels, or mark one agent’s assertion as authoritative over another’s.

The tools reviewed are built for the same use case: chatbot and personal assistant personalization. Enterprise data agents have a structurally different problem. They need to understand the data estate they are operating on: what the data means, where it came from, who owns it, whether it is trustworthy, and under what rules it can be used. That problem requires infrastructure that connects to the data estate itself, not infrastructure that stores conversation context alongside it.

The two categories are complementary, not competitive. Recognizing the distinction is the first step to scoping your enterprise AI memory layer investment correctly.


How to choose an AI agent memory framework

Before selecting a framework, answer three questions. What data sources will your agents operate on? What happens when your agent produces a wrong answer — how do you trace it? Does your team have capacity to run and maintain the memory layer infrastructure, or do you need managed cloud?

Those answers eliminate most options before you evaluate features.

Decision framework


If you need… Consider… Why
Fastest path to personalization memory Mem0 Managed, drop-in, largest community, self-editing memory
Temporal reasoning (“how did this change over time?”) Zep / Graphiti Validity windows on every node and edge, so an agent can answer what was true at a past point
Memory for an existing LangChain stack LangChain / LangMem Zero new dependency, procedural memory available
Long-running agents with unlimited persistent memory Letta OS-inspired tiered memory, full retrieval depth on all tiers
Azure / Microsoft 365 enterprise deployment Microsoft Semantic Kernel Azure IAM, .NET SDK, Copilot Studio integration
Fully local deployment, graph reasoning, privacy-first Cognee No cloud dependency, poly-store flexibility, 6-line setup
Coding agents (Claude Code, OpenCode) Supermemory MCP-native, explicit forgetting, coding agent plugins
Low-latency backend for existing Redis infrastructure Redis Agent Memory Server Sub-ms working memory, composable with Mem0/LangMem

By team type


Individual developers / prototyping: Mem0 (managed, free tier to start) or Cognee (local, zero cloud cost).

Teams on LangChain: LangMem is the natural extension. Evaluate Zep if your agents need to answer temporal queries.

Azure enterprise shops: Semantic Kernel and Kernel Memory is the default path; evaluate whether Azure AI Search meets your retrieval requirements before adding another vector DB.

Research and long-horizon agents: Letta’s tiered memory with full retrieval depth at every tier, including self-hosted free.

Coding agent workflows: Supermemory via MCP. Redis as a low-latency backend if working memory retrieval speed is the constraint.


Atlan’s context layer: what enterprise data agents need that memory frameworks don’t provide

Atlan’s context layer is not a memory framework. It is a governed metadata layer designed to ground enterprise data agents in authoritative business context. It provides the five components the frameworks above do not: a semantic layer with governed metric definitions, cross-system entity resolution, operational playbooks, data lineage and provenance, and decision memory via active metadata. It is designed for agents operating across multi-platform data estates, not agents remembering conversations.

What it does that memory frameworks don’t:

Business glossary integration. Agents query governed metric definitions, not raw schema names. “Revenue” routes to the certified board-level definition, not the first column named revenue in a database schema.

Cross-system entity resolution. Atlan maps entity identity across Salesforce, Snowflake, Databricks, and operational systems simultaneously. An agent asking about a customer gets a resolved entity, not a per-system fact.

Data lineage. Every answer is traceable to the source table, pipeline, and transformation that produced it. Agents can cite provenance; compliance teams can audit reasoning. Snowflake’s published research found that adding a context layer to data agents delivered a 20% accuracy improvement and 39% reduction in tool calls — the attribution traces to governed context, not expanded memory.

Governance policy enforcement. Data access policies, certification status, and retention rules are enforced at the context layer, before an agent retrieves restricted information.

Active metadata. The institutional history of how data assets have been used, queried, and modified — not conversation logs, but the history of the data estate itself.

What it does not replace. Atlan does not provide conversation memory, user preference storage, or session persistence. For those capabilities, the tools reviewed above apply. The context layer and a memory framework are designed to work together, not compete.

See how the agent context layer fits into enterprise AI architecture and what context layer enterprise AI means for governed data agents.


FAQs about AI agent memory frameworks

1. What is the best AI agent memory framework in 2026?


There is no single best framework; the right choice depends on your use case. For managed, drop-in personalization memory, Mem0 leads on community size. For temporal reasoning, Zep models when each fact was true and when it stopped being true. Teams on LangChain should evaluate LangMem first. Long-running agents benefit from Letta’s tiered memory model. If your use case is coding agents, Supermemory’s MCP integrations are worth evaluating.

2. How does Mem0 compare to Zep for AI agent memory?


Mem0 is broader and easier to adopt. Zep stores fact validity windows rather than timestamped snapshots, so it answers questions about change over time. The published numbers do not settle it. Zep self-reports 90.2% on LongMemEval with gpt-5.4 as reader and judge; Mem0 self-reports 94.4% for its managed platform and does not publish its test configuration. Different papers, different models, different judges, and no third party has reproduced either. Mem0 wins on community size and managed cloud polish. Zep’s SOC 2 Type II and HIPAA BAA sit on its Enterprise tier.

3. What is the difference between Mem0 and LangMem?


Mem0 is a standalone managed service; LangMem is a sub-package of LangChain. Mem0 works with any agent stack via REST API. LangMem is tightly coupled to LangChain/LangGraph — practical adoption requires adopting those frameworks too. Mem0 provides managed cloud infrastructure; LangMem requires your team to run its own storage backend. LangMem’s procedural memory (agents rewriting their own instructions) has no equivalent in Mem0.

4. Does LangChain have built-in long-term memory for AI agents?


Yes, via the LangMem SDK (launched early 2025) and LangGraph’s persistent store layer. LangMem supports three memory types: episodic (past interactions), semantic (facts and preferences), and procedural (agents updating their own system instructions). The SDK is free and open source. Long-term memory requires LangGraph’s StateGraph — it does not work with older non-LangGraph chains.

5. How does Letta (formerly MemGPT) handle agent memory?


Letta documents four tiers, ordered by scale of data and importance rather than by a storage metaphor: memory blocks, files, archival memory and external RAG. Agents do not passively receive context; they explicitly call memory management functions to move information between tiers. The three-tier RAM-and-disk framing belongs to the original MemGPT paper, not to Letta’s current documentation.

6. What is a temporal knowledge graph and why does it matter for AI agents?


A temporal knowledge graph stores facts as nodes with validity windows — a fact is true “from X until Y,” not just stored at a timestamp. When new information contradicts an existing fact, the old fact is invalidated but preserved, maintaining historical state. For agents tracking how business relationships, customer behavior, or data values change over time, temporal graphs outperform flat vector stores by significant margins.

7. What is the difference between short-term and long-term memory in AI agents?


Short-term (working) memory is the agent’s current session context — everything in the active prompt window. It is fast but bounded by the context limit and is lost when the session ends. Long-term memory persists across sessions via external storage (vector DB, graph DB, or key-value store). Most memory frameworks bridge this gap: storing what matters from short-term into retrievable long-term storage. The design choices around what to store and how to retrieve it drive most of the performance differences between frameworks.

8. Can multiple AI agents share the same memory pool?


Yes, but multi-agent shared memory requires careful design to prevent contamination across agent sessions. Mem0 uses scoped memory (user/agent/session isolation). LangGraph supports shared state across agents in a graph. Letta has native multi-agent coordination. Redis provides a low-latency shared backend. The harder problem is governance: without conflict resolution and authority rules, shared memory pools degrade as agents write inconsistent facts about the same entities.

9. What is context engineering for AI agents?


Context engineering is the practice of deliberately designing what information an agent receives before it reasons — rather than relying solely on what it retrieves at query time. Zep rebranded its v3 SDK as a “Context Engineering Platform” to signal this shift. Broader definitions include structured context injection, dynamic retrieval based on query type, and prompt engineering for agent grounding. The field is evolving quickly; definitions vary significantly across vendors.

10. Why do AI agents forget things between sessions, and how do you fix it?


Agents forget between sessions because LLM context windows are stateless — each new session starts with no memory of previous ones. The fix is a persistent external memory store that writes important facts at session end and injects them at session start. All 8 frameworks reviewed solve this problem. The differences between them show up at scale: temporal accuracy, multi-agent coordination, compliance posture, and whether the framework can be grounded in your actual data estate rather than just session history.


External citations: Zep arXiv preprint 2501.13956 (vendor research, Zep staff) | Zep research page | Zep pricing | Mem0 research | Mem0 $24M seed and Series A | Mem0 pricing | Letta benchmarking post | Letta’s next phase | Cognee pricing | AI governance gap: Galileo.ai | Snowflake context layer experiment

Share this article

signoff-panel-logo

Atlan is the Context Layer for AI. It translates business knowledge, including data definitions, working procedures, and governance policies, into context AI can actually use. This knowledge lives in a single Enterprise Data Graph that every team and AI agent can reach.

In Atlan's AI Labs benchmark, adding this context improved AI's text-to-SQL accuracy by 38%.

Atlan is recognized as a Leader across multiple Gartner reports and Forrester Waves, and is trusted by over 400 enterprises representing $10T+ in market cap, including Mastercard, Workday, General Motors, CME Group, HubSpot, FOX, Virgin Media O2, and Elastic.

Bridge the context gap.
Ship AI that works.