Skip to main content
Analysis

Cost & Latency Analysis

How delta updates significantly reduce adaptation latency and cost

Delta incremental updates and non-LLM merging give ACE significant advantages over GEPA / Dynamic Cheatsheet in both adaptation latency and token cost.

Offline Scenario (AppWorld)

Compared to GEPA:
MetricGEPAACEChange
Adaptation Latency (seconds)53,8989,517-82.3%
Number of Rollouts1,434357-75.1%

Online Scenario (FiNER)

Compared to Dynamic Cheatsheet (DC):
MetricDCACEChange
Adaptation Latency (seconds)65,1045,503-91.5%
Token Cost ($)17.72.9-83.6%

Sources of Latency Gains

1

Delta Instead of Full Rewrite

Generates only a small set of candidate bullets each time; model output size is small.
2

Non-LLM Logic Merging

Merging is just id matching + counter updates; no LLM calls needed.
3

Parallel Merging

Multiple deltas can be merged in parallel for batch adaptation.
4

Fewer Rollouts

A single rollout accumulates multiple experiences; no need for repeated trial-and-error.

Long Context Does Not Equal High Serving Cost

Although ACE-generated Playbooks are typically longer than GEPA prompts, inference cost during serving does not increase linearly. Reasons:
MechanismEffect
KV Cache ReuseIdentical prefixes avoid repeated prefill
KV Cache CompressionSmaller memory footprint
KV Cache OffloadCan be loaded from disk/remote storage
Modern serving stacks have made system-level optimizations for long-context workloads, making ACE's long Playbooks feasible in production environments.

When Savings Are Greatest

ScenarioSignificant Savings
Offline optimization~75% fewer rollouts vs GEPA
Online adaptation~90% less latency vs DC
Domain tasks requiring long-term evolutionHigh efficiency of delta accumulation
Learn about specific baselines ACE compares against: Baseline Reference; also see Known Limitations.