Performance of the Mem0 family in retrieval latency, response latency, and token usage
One of Mem0's core selling points is maintaining answer quality close to Full-Context while reducing latency and cost to production-ready levels. This page summarizes key data.
Time taken for a single memory/text chunk retrieval.
Mem0's retrieval completes almost instantly; even with graph queries introduced, Mem0g still outperforms all existing solutions.
Total time including retrieval + LLM generation.
Zep causes cross-node redundancy explosion because it caches full summaries at each node and stores facts on edges.
RAG's best performance across different chunk sizes and k values is ~61% J, still below Mem0's 66.88%.
Conclusion: Compressing conversations into compact facts is more precise than retrieving raw text chunks.
Mem0's memory construction typically completes within 1 minute, and memories are immediately retrievable after writing. In contrast, Zep's graph construction involves multiple asynchronous LLM calls; empirical testing shows it takes hours to stabilize, making it difficult to support real-time applications.
Search Latency
Time taken for a single memory/text chunk retrieval.
| Method | p50 (s) | p95 (s) |
|---|---|---|
| A-Mem | 0.668 | 1.485 |
| LangMem | 17.99 | 59.82 |
| Zep | 0.513 | 0.778 |
| Mem0 | 0.148 | 0.200 |
| Mem0g | 0.476 | 0.657 |
End-to-End Response Latency
Total time including retrieval + LLM generation.
| Method | p50 (s) | p95 (s) |
|---|---|---|
| Full-Context | 9.870 | 17.117 |
| Best RAG (k=2, 8k chunk) | 2.312 | 9.942 |
| LangMem | 18.53 | 60.40 |
| Zep | 1.292 | 2.926 |
| Mem0 | 0.708 | 1.440 |
| Mem0g | 1.091 | 2.590 |
- Mem0's p95 is ~92% lower than Full-Context
- Mem0g's p95 is ~85% lower than Full-Context
Token Usage (Per Conversation)
| System | Token Count | Relative to Full-Context |
|---|---|---|
| Full-Context (original) | ~26k | 1.0x |
| Mem0 | ~7k | 0.27x |
| Mem0g | ~14k | 0.54x |
| Zep Graph | 600k+ | 23x |
RAG Comparison
RAG's best performance across different chunk sizes and k values is ~61% J, still below Mem0's 66.88%.
| Scenario | Overall J |
|---|---|
| RAG (k=1, chunk=256) | 50.15% |
| RAG (k=2, chunk=256) | 60.97% |
| RAG (k=2, chunk=8192) | 60.53% |
| Mem0 | 66.88% |