Performance of Mem0 and Mem0g across four question types
On the LOCOMO long-term conversational memory benchmark, the Mem0 family leads or nearly leads across four question types (single-hop, multi-hop, temporal, open-domain).
The original dataset also includes adversarial questions, excluded due to lack of ground truth.
Three metrics are used simultaneously:
Summarized by J score (numbers are percentages, ↑ higher is better).
Single-Hop: Answer lies within a single conversation turn. Mem0's dense natural language memories are sufficiently precise; graph structure slightly degrades performance.
Multi-Hop: Requires synthesizing information across multiple turns. Mem0 clearly leads, demonstrating that natural language memories suffice for complex retrieval and aggregation.
Open-Domain: Requires combining external knowledge. Zep has a slight edge here; Mem0g follows closely.
Temporal: Requires reconstructing event chronology. Mem0g significantly leads, validating the value of graph structure for temporal reasoning.
In overall J score:
Full-Context still leads by ~4–5 percentage points, but at the cost of 17-second p95 latency. The Mem0 family achieves the best trade-off in the accuracy-latency-cost triangle.
LOCOMO Dataset
| Attribute | Value |
|---|---|
| Number of conversations | 10 |
| Turns per conversation | ~600 |
| Tokens per conversation | ~26,000 |
| Average questions per conversation | 200 |
| Question categories | Single-hop, multi-hop, temporal, open-domain |
Evaluation Metrics
Three metrics are used simultaneously:
| Metric | Type | Description |
|---|---|---|
| F1 | Lexical overlap | Traditional QA metric, only considers word coverage |
| BLEU-1 (B1) | Lexical overlap | Traditional machine translation metric |
| LLM-as-a-Judge (J) | Semantic correctness | Scored by another LLM; primary reference metric |
F1 and BLEU-1 only consider word overlap and tend to overestimate answers with "factual errors"; J is a metric closer to human judgment. All J scores are mean ± standard deviation over 10 independent runs.
Performance Across Four Question Types
Summarized by J score (numbers are percentages, ↑ higher is better).
| Method | Single-Hop J | Multi-Hop J | Open-Domain J | Temporal J |
|---|---|---|---|---|
| A-Mem* | 39.79 | 18.85 | 54.05 | 49.91 |
| LangMem | 62.23 | 47.92 | 71.12 | 23.43 |
| Zep | 61.70 | 41.35 | 76.60 | 49.31 |
| OpenAI | 63.79 | 42.92 | 62.29 | 21.71 |
| Mem0 | 67.13 | 51.15 | 72.93 | 55.51 |
| Mem0g | 65.71 | 47.19 | 75.71 | 58.13 |
Interpretation by Question Type
Single-Hop: Answer lies within a single conversation turn. Mem0's dense natural language memories are sufficiently precise; graph structure slightly degrades performance.
Multi-Hop: Requires synthesizing information across multiple turns. Mem0 clearly leads, demonstrating that natural language memories suffice for complex retrieval and aggregation.
Open-Domain: Requires combining external knowledge. Zep has a slight edge here; Mem0g follows closely.
Temporal: Requires reconstructing event chronology. Mem0g significantly leads, validating the value of graph structure for temporal reasoning.
Overall Gap vs. Baselines
In overall J score:
| Method | Overall J |
|---|---|
| Best RAG | 60.97% |
| Full-Context | 72.90% |
| OpenAI | 52.90% |
| Zep | 65.99% |
| Mem0 | 66.88% |
| Mem0g | 68.44% |