Brief descriptions of the six baseline categories compared in Mem0 evaluation
The Mem0 paper organizes system comparisons into six baseline categories. This page summarizes the core idea of each.
Two-tier short-term and long-term memory. Generates summaries per session as short-term memory; converts single-turn dialogues into observations stored in long-term memory with references to specific conversation turns. Introduces temporal event graphs to track cross-session causal events.
Mimics human reading of long texts through a three-stage process:
Three-part pipeline:
Borrows from operating system tiered memory management:
Agentic memory: Each note carries keywords, context descriptions, and tags. When new notes are generated, related notes are found via semantic embeddings and linked, and existing notes are updated to integrate new knowledge.
Memory component provided by the LangChain ecosystem. Experiments use
Treats entire conversation history as a document corpus, using fixed chunk sizes (128 to 8192) and retrieving top-k (k=1 or k=2):
Directly passes complete conversation (~26k tokens) as context to LLM. Achieves highest J score (72.90%) but p95 latency reaches 17 seconds, with token costs far exceeding other methods.
Uses
Commercial temporal knowledge graph architecture. Retains timestamp information, supports time-sensitive retrieval. Limitations:
Six Baseline Categories Overview
| Category | Representative Methods | Characteristics |
|---|---|---|
| Existing LOCOMO baselines | LoCoMo, ReadAgent, MemoryBank, MemGPT, A-Mem | Various mainstream memory architectures |
| Open-source memory solutions | LangMem (Hot Path) | Memory component in LangChain ecosystem |
| RAG | Fixed-length chunking + vector retrieval | Exhaustive search over chunk_size and k combinations |
| Full-context | Directly concatenated into LLM | No retrieval; upper-bound reference |
| Proprietary systems | OpenAI ChatGPT memory | Commercial memory feature |
| Memory platforms | Zep | Commercial-grade graph memory management |
Established LOCOMO Baselines
LoCoMo
Two-tier short-term and long-term memory. Generates summaries per session as short-term memory; converts single-turn dialogues into observations stored in long-term memory with references to specific conversation turns. Introduces temporal event graphs to track cross-session causal events.
ReadAgent
Mimics human reading of long texts through a three-stage process:
- Episode Pagination: Segments at natural boundaries
- Memory Gisting: Compresses each segment into key points
- Interactive Lookup: Locates by key points during queries and recalls original segments
MemoryBank
Three-part pipeline:
- Memory Storage: Detailed conversation logs + hierarchical event summaries + user profiles
- Memory Retrieval: Dual-tower dense retrieval
- Memory Updating: Simulates human forgetting curves; reinforced when recalled, decayed when not
MemGPT
Borrows from operating system tiered memory management:
- Main Context: Similar to RAM, limited by window size
- External Context: Similar to disk, unlimited capacity
- Function calls: LLM actively decides paging and retrieval
A-Mem
Agentic memory: Each note carries keywords, context descriptions, and tags. When new notes are generated, related notes are found via semantic embeddings and linked, and existing notes are updated to integrate new knowledge.
Open-Source Solutions
LangMem (Hot Path)
Memory component provided by the LangChain ecosystem. Experiments use gpt-4o-mini as the main model and text-embedding-small-3 for embeddings. Retrieval latency is notably high (p50 ~18 seconds), limiting interactive applications.
RAG Baseline
Treats entire conversation history as a document corpus, using fixed chunk sizes (128 to 8192) and retrieving top-k (k=1 or k=2):
- When k > 2, total retrieval approaches entire conversation, losing selectivity
- Best configuration: k=2, chunk=256, Overall J ≈ 60.97%
Full-Context
Directly passes complete conversation (~26k tokens) as context to LLM. Achieves highest J score (72.90%) but p95 latency reaches 17 seconds, with token costs far exceeding other methods.
OpenAI ChatGPT Memory
Uses gpt-4o-mini, generating timestamped memories by injecting prompts through the ChatGPT interface, then answering questions with full context. No external retrieval API; evaluation effectively gives it preferential treatment.
Zep (Memory Platform)
Commercial temporal knowledge graph architecture. Retains timestamp information, supports time-sensitive retrieval. Limitations:
- Graph nodes replicate full summaries, causing token usage up to 600k+
- Graph construction involves multiple asynchronous LLM calls; empirical testing shows hours before stable availability