A two-stage incremental memory pipeline for long-term conversational memory
Mem0 is an incremental memory pipeline that runs in real time alongside conversations. It extracts facts from message pairs, writes them into a searchable memory store, and recalls relevant entries on demand when answering—achieving near-full-context answer quality with only ~7k tokens per conversation.
Key Results
| Metric | Value |
|---|---|
| Overall J score (LOCOMO) | 66.88% |
| Search latency p50 | 0.148s |
| Token usage per conversation | ~7k (vs. Full-Context's ~26k) |
| Performance gap vs. Full-Context | ~4–5 percentage points |
Two Stages
Extraction
Constructs full context (summary + recent messages + current pair) and calls LLM to extract candidate facts
Update
Retrieves top-s similar memories, decides ADD/UPDATE/DELETE/NOOP via LLM tool call
Core Design Principles
- Incremental: Updates run in real time with the conversation flow; no periodic batch processing
- No classifier: Uses LLM reasoning directly instead of training a separate operation classifier
- Natural language storage: Facts stored as natural language rather than structured triples, preserving nuance
- Asynchronous summary: Conversation summaries refreshed independently without blocking the main pipeline
Reading Path
- Motivation: Why existing methods fall short
- Architecture: Extraction and update stages
- Graph Memory: Mem0g variant supporting relational reasoning
- Benchmarks: LOCOMO evaluation results
- Latency & Cost: Production-ready performance data