ACE's performance on FiNER and Formula financial benchmarks
Financial analysis relies heavily on domain concepts and rules. ACE achieves significant improvements on FiNER and Formula, two XBRL benchmarks, by accumulating Playbooks containing specific rules.
Metric is Accuracy.
Base model: DeepSeek-V3.1.
1. Especially pronounced gains on domain tasks
ACE (with GT) leads ICL / MIPROv2 / GEPA by an average of 10.9%. Structured evolving context is particularly suited for tasks requiring precise concepts (financial concepts, XBRL rules).
2. Online still outperforms DC
ACE Online leads DC by an average of 6.2%, demonstrating that long-term accumulated domain strategies are more valuable than simple memory caching.
3. Feedback signal quality is critical
Two Datasets
| Dataset | Task Type | Description |
|---|---|---|
| FiNER | Token classification | Annotates 139 fine-grained entity types from XBRL documents |
| Formula | Numerical reasoning | Extracts values from XBRL and computes to answer financial queries |
Main Results
Base model: DeepSeek-V3.1.
| Method | GT | FiNER | Formula | Average |
|---|---|---|---|---|
| Base LLM | – | 70.7 | 67.5 | 69.1 |
| Offline | ||||
| ICL | ✓ | 72.3 | 67.0 | 69.6 |
| MIPROv2 | ✓ | 72.4 | 69.5 | 70.9 |
| GEPA | ✓ | 73.5 | 71.5 | 72.5 |
| ACE | ✓ | 78.3 | 85.5 | 81.9 |
| ACE | ✗ | 71.1 | 83.0 | 77.1 |
| Online | ||||
| DC (CU) | ✓ | 74.2 | 69.5 | 71.8 |
| DC (CU) | ✗ | 68.3 | 62.5 | 65.4 |
| ACE | ✓ | 76.7 | 76.5 | 76.6 |
| ACE | ✗ | 67.3 | 78.5 | 72.9 |
Key Observations
1. Especially pronounced gains on domain tasks
ACE (with GT) leads ICL / MIPROv2 / GEPA by an average of 10.9%. Structured evolving context is particularly suited for tasks requiring precise concepts (financial concepts, XBRL rules).
2. Online still outperforms DC
ACE Online leads DC by an average of 6.2%, demonstrating that long-term accumulated domain strategies are more valuable than simple memory caching.
3. Feedback signal quality is critical
When ACE Yields Maximum Gains
| Condition | ACE Gain |
|---|---|
| Many domain concepts, fine-grained rules | High |
| Reliable execution feedback available (code execution, formula matching) | High |
| GT labels available | Higher |
| No GT and no execution signals | Degradation risk |
Comparison with Agent Benchmark
| Scenario | Agent (AppWorld) | Finance (FiNER + Formula) |
|---|---|---|
| Feedback source | Code execution success/failure | Formula matching |
| GT dependency | Weak dependency | Strong dependency |
| Playbook evolution | Cross-episode strategy accumulation | Domain rule library |