Two core pain points of existing methods: brevity bias and context collapse
Context adaptation has become the mainstream paradigm for building LLM applications: modify inputs, not weights. However, existing methods suffer from two recurring deficiencies.
Many prompt optimizers tend to compress context into brief, generic instructions, sacrificing domain specificity.
When an LLM is asked to fully rewrite accumulated context each time, as context grows longer, the model tends to compress it into short summaries, causing drastic information loss.
An empirical observation on AppWorld:
It collapses at the next step: Accuracy drops even below the no-adaptation baseline (63.7).
The paper's core claim: Context should be a comprehensive and evolving Playbook, not a brief summary.
Deficiency 1: Brevity Bias
Many prompt optimizers tend to compress context into brief, generic instructions, sacrificing domain specificity.
- Iterative optimization repeatedly produces similar generic instructions (e.g., "Create unit tests to ensure methods behave as expected")
- Shrinks the search space and allows the same errors to propagate through iterations
- Significantly harmful for multi-step agents, program synthesis, and knowledge-intensive reasoning
Deficiency 2: Context Collapse
When an LLM is asked to fully rewrite accumulated context each time, as context grows longer, the model tends to compress it into short summaries, causing drastic information loss.
An empirical observation on AppWorld:
| Step | Token Count | Accuracy |
|---|---|---|
| Step 60 | 18,282 | 66.7 |
| Step 61 | 122 | 57.1 |
ACE's Stance
The paper's core claim: Context should be a comprehensive and evolving Playbook, not a brief summary.
| Traditional Stance | ACE Stance |
|---|---|
| Shorter context is better | Context should contain sufficient domain details |
| Full rewrite | Incremental delta updates |
| Model self-compression | Model autonomously selects relevant portions |
Why LLMs Can Handle Long Context
- Modern long-context LLMs can handle hundreds of thousands of tokens
- Server-side KV cache reuse, compression, and offloading reduce the amortized cost of long context
- LLMs naturally excel at autonomously extracting relevant portions from long context; humans, conversely, need concise summaries
Design Rationale for Delta Updates and Grow-and-Refine
- Delta updates: Replace full rewrites with minimal edit units to prevent collapse
- Grow-and-Refine: Add first, then deduplicate, keeping context scannable and non-redundant