Skip to main content
Evaluation

LOCOMO Benchmark Results

Performance of Mem0 and Mem0g across four question types

On the LOCOMO long-term conversational memory benchmark, the Mem0 family leads or nearly leads across four question types (single-hop, multi-hop, temporal, open-domain).

LOCOMO Dataset

AttributeValue
Number of conversations10
Turns per conversation~600
Tokens per conversation~26,000
Average questions per conversation200
Question categoriesSingle-hop, multi-hop, temporal, open-domain
The original dataset also includes adversarial questions, excluded due to lack of ground truth.

Evaluation Metrics

Three metrics are used simultaneously:
MetricTypeDescription
F1Lexical overlapTraditional QA metric, only considers word coverage
BLEU-1 (B1)Lexical overlapTraditional machine translation metric
LLM-as-a-Judge (J)Semantic correctnessScored by another LLM; primary reference metric
F1 and BLEU-1 only consider word overlap and tend to overestimate answers with "factual errors"; J is a metric closer to human judgment. All J scores are mean ± standard deviation over 10 independent runs.

Performance Across Four Question Types

Summarized by J score (numbers are percentages, ↑ higher is better).
MethodSingle-Hop JMulti-Hop JOpen-Domain JTemporal J
A-Mem*39.7918.8554.0549.91
LangMem62.2347.9271.1223.43
Zep61.7041.3576.6049.31
OpenAI63.7942.9262.2921.71
Mem067.1351.1572.9355.51
Mem0g65.7147.1975.7158.13

Interpretation by Question Type

Single-Hop: Answer lies within a single conversation turn. Mem0's dense natural language memories are sufficiently precise; graph structure slightly degrades performance. Multi-Hop: Requires synthesizing information across multiple turns. Mem0 clearly leads, demonstrating that natural language memories suffice for complex retrieval and aggregation. Open-Domain: Requires combining external knowledge. Zep has a slight edge here; Mem0g follows closely. Temporal: Requires reconstructing event chronology. Mem0g significantly leads, validating the value of graph structure for temporal reasoning.

Overall Gap vs. Baselines

In overall J score:
MethodOverall J
Best RAG60.97%
Full-Context72.90%
OpenAI52.90%
Zep65.99%
Mem066.88%
Mem0g68.44%
Full-Context still leads by ~4–5 percentage points, but at the cost of 17-second p95 latency. The Mem0 family achieves the best trade-off in the accuracy-latency-cost triangle.
For latency and token cost details, see Latency & Cost.