Skip to main content
Benchmarks

Ablation Study

Individual contributions of Reflector, multi-round refinement, and Offline warmup

The paper conducts three ablation experiments on AppWorld to validate ACE's key design choices.

Ablation Dimensions

Reflector

Distills experiences from traces, independent of Curator

Multi-Round Refinement

Reflector iteratively refines experiences

Offline Warmup

Preheats Playbook with training set before Online phase

AppWorld Offline Ablation

VariantGTTest-NormalTest-ChallengeAverage
TGCSGCTGCSGC
ReAct (baseline)–63.742.941.521.642.4
ACE without Reflector or multi-round✓70.855.455.938.155.1
ACE without multi-round✓72.060.754.939.656.8
Full ACE✓76.264.357.339.659.4

AppWorld Online Ablation

VariantGTTest-NormalTest-ChallengeAverage
TGCSGCTGCSGC
ACE Online✗67.951.861.443.256.1
ACE Online + Offline warmup✗69.653.666.048.959.5

Three Conclusions

1. Reflector is the core of role separation design From 55.1% → 56.8% (+1.7), separating evaluation and distillation from generation is the first layer of significant gains. 2. Multi-round refinement further improves performance From 56.8% → 59.4% (+2.6), 5 rounds of Reflector iteration significantly strengthen experience distillation quality. 3. Offline warmup yields substantial gains on challenge split Online goes from 56.1% → 59.5%, with test-challenge TGC rising from 61.4% → 66.0% (+4.6) and SGC from 43.2% → 48.9% (+5.7).
Offline warmup provides an initial Playbook already containing strategies, allowing Online adaptation to start from a higher baseline.

Design Trade-offs

ComponentValue Retained
ReflectorClear responsibility for "experience distillation"
Multi-round refinementHigher-quality experience entries
Offline warmupReduces Online learning curve
For latency and cost ablations, see Cost & Latency Analysis.