Four categories of context adaptation baselines compared against ACE
ACE compares against four categories of context adaptation methods in experiments, covering few-shot, prompt optimizers, and adaptive memory.
No context engineering; evaluated directly with official prompts. On AppWorld, uses the official ReAct implementation as a common starting point for all methods.
ACE is positioned as broad context adaptation, encompassing agent memory, system prompts, factual evidence, etc.; while other agent memory systems (AgentFly, AWM, A-MEM, Agentic Plan Caching) focus primarily on memory mechanisms.
Baseline Overview
| Method | Type | Requires GT |
|---|---|---|
| Base LLM | No adaptation | – |
| ICL | Few-shot / Many-shot examples | Yes |
| MIPROv2 | Prompt optimizer | Yes |
| GEPA | Reflective prompt optimizer | Yes |
| Dynamic Cheatsheet (DC) | Test-time adaptive memory | Optional |
Base LLM
No context engineering; evaluated directly with official prompts. On AppWorld, uses the official ReAct implementation as a common starting point for all methods.
In-Context Learning (ICL)
- Provides task examples in the prompt
- Includes all training samples that fit in context; otherwise fills the window
- Few-shot or many-shot
MIPROv2
- Prompt optimizer provided by DSPy
- Jointly optimizes system instructions and examples via Bayesian optimization
- Experiments use
auto="heavy"to maximize optimization effort
GEPA (Genetic-Pareto)
- Sample-efficient optimizer based on reflective prompt evolution
- Collects execution traces (reasoning, tool calls, intermediate outputs)
- Uses natural language reflection to diagnose errors and propose prompt updates
- Uses genetic Pareto search to retain frontier high-performance prompts
- Achieves 35× fewer rollouts and 10-20% higher accuracy compared to RL methods like GRPO
Dynamic Cheatsheet (DC)
- Test-time adaptive external memory
- Accumulates reusable strategies and code snippets
- Can self-update without GT labels
- Experiments use cumulative mode (DC-CU)
Relationship to Agent Memory Systems
ACE is positioned as broad context adaptation, encompassing agent memory, system prompts, factual evidence, etc.; while other agent memory systems (AgentFly, AWM, A-MEM, Agentic Plan Caching) focus primarily on memory mechanisms.
| System | Focus | Relationship to ACE |
|---|---|---|
| AgentFly | Continuously evolving memory + RL | Similar but focuses on training |
| AWM | Workflow reuse | Similar but only concerns workflows |
| A-MEM | Zettelkasten-style memory | ACE's bullet concept inspired by it |
| Agentic Plan Caching | Plan template caching | Similar but only addresses efficiency |