Agentic Context Engineering: Treating context as an evolving Playbook
ACE (Agentic Context Engineering) treats LLM context as a continuously evolving Playbook through three roles—Generator, Reflector, and Curator—accumulating, refining, and organizing strategies and domain insights rather than compressing them into brief summaries.
On the latest AppWorld leaderboard, ReAct + ACE based on open-source model DeepSeek-V3.1 ties with top commercial model IBM CUGA (based on GPT-4.1) in overall score, and surpasses IBM CUGA on the harder test-challenge subset.
ACE is especially suited for two types of applications:
Key Results
| Scenario | ACE Improvement |
|---|---|
| AppWorld agent benchmark | Average +10.6% |
| Finance benchmarks (FiNER / Formula) | Average +8.6% |
| No-GT-label scenarios | Still improves +14.8% (AppWorld) |
| Adaptation latency | Average reduction 86.9% |
| Token cost (vs Dynamic Cheatsheet) | Reduction 83.6% |
Three Components
Generator
Generates reasoning traces for new problems, marking useful/misleading strategies
Reflector
Distills reusable experiences from traces, supporting multi-round refinement
Curator
Merges experiences into incremental deltas using non-LLM logic
Two Key Innovations
Delta Updates
Edits only relevant bullets, avoiding full rewrites
Grow-and-Refine
Expands first, then deduplicates to prevent context collapse
Target Problems
ACE is especially suited for two types of applications:
- Agents: Multi-turn reasoning, tool use, environment interaction; cross-episode strategy reuse
- Domain-specific tasks: Require specialized concepts and rules (e.g., XBRL financial analysis)
Reading Path
- Motivation & Background: Why existing methods fall short
- ACE Framework: How the three roles collaborate
- Delta Updates: How incremental evolution is achieved
- Grow-and-Refine: How to avoid context bloat
- Benchmark Results: Performance on real benchmarks