Skip to main content
Introduction

Research Background

Limitations of LLM fixed context windows and the motivation for external memory/context adaptation

Large language models have made significant progress in generating fluent responses, but they remain constrained by fixed context windows. Real-world conversations often span days or sessions, with topics interleaving across unrelated dialogues; simply expanding the context window cannot fundamentally solve this problem.

Three Limitations of Fixed Context

LimitationManifestationConsequence
Length capEven 128k–10M tokens cannot cover week/month-long conversationsKey facts are truncated or diluted
Attention decayAttention weights drop for distant tokensEarly settings are ignored
Topic switchingUsers switch back and forth between multiple topicsRelevant facts get buried in irrelevant content

Two Complementary Technical Approaches

An Intuitive Example

A user mentions in an initial conversation that they are vegetarian and do not eat dairy. If the system has no persistent memory, asking for dinner suggestions weeks later might result in chicken recommendations, conflicting with established preferences.
  • Without memory: The system relies only on the current session context; early preferences are "forgotten"
  • With memory: The system recalls dietary preferences from the memory store and provides compliant suggestions

Problem Types Addressed

  • Single-hop: Answers given directly within one conversation turn
  • Multi-hop: Requires synthesizing information across multiple conversation turns
  • Temporal: Requires reasoning about event order based on timestamps
  • Open-domain: Requires combining external knowledge to answer
After understanding the background, dive into specific methods: Mem0 Overview or ACE Overview.