Skip to main content
Evaluation

Latency & Cost

Performance of the Mem0 family in retrieval latency, response latency, and token usage

One of Mem0's core selling points is maintaining answer quality close to Full-Context while reducing latency and cost to production-ready levels. This page summarizes key data.

Search Latency

Time taken for a single memory/text chunk retrieval.
Methodp50 (s)p95 (s)
A-Mem0.6681.485
LangMem17.9959.82
Zep0.5130.778
Mem00.1480.200
Mem0g0.4760.657
Mem0's retrieval completes almost instantly; even with graph queries introduced, Mem0g still outperforms all existing solutions.

End-to-End Response Latency

Total time including retrieval + LLM generation.
Methodp50 (s)p95 (s)
Full-Context9.87017.117
Best RAG (k=2, 8k chunk)2.3129.942
LangMem18.5360.40
Zep1.2922.926
Mem00.7081.440
Mem0g1.0912.590
  • Mem0's p95 is ~92% lower than Full-Context
  • Mem0g's p95 is ~85% lower than Full-Context

Token Usage (Per Conversation)

SystemToken CountRelative to Full-Context
Full-Context (original)~26k1.0x
Mem0~7k0.27x
Mem0g~14k0.54x
Zep Graph600k+23x
Zep causes cross-node redundancy explosion because it caches full summaries at each node and stores facts on edges.

RAG Comparison

RAG's best performance across different chunk sizes and k values is ~61% J, still below Mem0's 66.88%.
ScenarioOverall J
RAG (k=1, chunk=256)50.15%
RAG (k=2, chunk=256)60.97%
RAG (k=2, chunk=8192)60.53%
Mem066.88%
Conclusion: Compressing conversations into compact facts is more precise than retrieving raw text chunks.

Runtime Availability

Mem0's memory construction typically completes within 1 minute, and memories are immediately retrievable after writing. In contrast, Zep's graph construction involves multiple asynchronous LLM calls; empirical testing shows it takes hours to stabilize, making it difficult to support real-time applications.
For specific baselines compared against Mem0: Baseline Reference; for prompts used in generation: Prompt Reference.