Skip to main content
Benchmarks

AppWorld Agent Results

ACE's performance on multi-turn tool use and environment interaction tasks

AppWorld is an autonomous agent evaluation benchmark covering API usage, code generation, and environment interaction. Tasks are divided into normal and challenge difficulty levels; as of submission, the best system on the official leaderboard averages only 60.3%.

Evaluation Metrics

MetricMeaning
TGCTask Goal Completion rate
SGCScenario Goal Completion rate
Evaluated separately on test-normal and test-challenge splits.

Main Results

Base model: DeepSeek-V3.1, base framework: ReAct.
MethodGTTest-NormalTest-ChallengeAverage
TGCSGCTGCSGC
ReAct (baseline)–63.742.941.521.642.4
Offline
ReAct + ICL✓64.346.446.027.346.0
ReAct + GEPA✓64.944.646.030.246.4
ReAct + ACE✓76.264.357.339.659.4
ReAct + ACE✗75.064.354.435.257.2
Online
ReAct + DC (CU)✗65.558.952.330.851.9
ReAct + ACE✗69.653.666.048.959.5

Three Key Findings

1. ACE significantly leads in offline scenarios ReAct + ACE improves over ReAct + ICL / ReAct + GEPA by 12.3% and 11.9% respectively, demonstrating that structured evolving context outperforms fixed examples or single-optimization instruction prompts. 2. Effective even without GT labels ReAct + ACE (no GT) still improves over baseline by 14.8%. ACE leverages execution feedback (whether code runs successfully) as signals for Reflector and Curator. 3. Small models match top commercial systems
  • Leaderboard (2025-09-20): IBM CUGA (GPT-4.1 driven) averages 60.3%
  • ReAct + ACE (DeepSeek-V3.1) averages 59.4%
  • On test-challenge, ACE Online + Offline warmup surpasses IBM CUGA by 8.4% TGC / 0.7% SGC

Why Greater Improvement on Challenge Split

  • Complex tasks benefit more from strategy reuse
  • ACE's Playbook accumulates long-term experience to help reduce common failure patterns
  • Differences on simple tasks are overshadowed by base capabilities

Comparison with GEPA / DC

MethodLimitationACE's Improvement
GEPASingle genetic optimizationContinuous evolution
DCFull rewrite causes collapseDelta + Grow-and-Refine
ICLFixed examplesContext can accumulate
For finance domain performance, see FiNER / Formula Results.