Skip to main content
Overview

Mem0 Overview

A two-stage incremental memory pipeline for long-term conversational memory

Mem0 is an incremental memory pipeline that runs in real time alongside conversations. It extracts facts from message pairs, writes them into a searchable memory store, and recalls relevant entries on demand when answering—achieving near-full-context answer quality with only ~7k tokens per conversation.

Key Results

MetricValue
Overall J score (LOCOMO)66.88%
Search latency p500.148s
Token usage per conversation~7k (vs. Full-Context's ~26k)
Performance gap vs. Full-Context~4–5 percentage points

Two Stages

Extraction

Constructs full context (summary + recent messages + current pair) and calls LLM to extract candidate facts

Update

Retrieves top-s similar memories, decides ADD/UPDATE/DELETE/NOOP via LLM tool call

Core Design Principles

  • Incremental: Updates run in real time with the conversation flow; no periodic batch processing
  • No classifier: Uses LLM reasoning directly instead of training a separate operation classifier
  • Natural language storage: Facts stored as natural language rather than structured triples, preserving nuance
  • Asynchronous summary: Conversation summaries refreshed independently without blocking the main pipeline

Reading Path

  1. Motivation: Why existing methods fall short
  2. Architecture: Extraction and update stages
  3. Graph Memory: Mem0g variant supporting relational reasoning
  4. Benchmarks: LOCOMO evaluation results
  5. Latency & Cost: Production-ready performance data
Start with Motivation to understand the design rationale before diving into Architecture.
Mem0 Overview - LLM Memory and Context Engineering