ACE's dependencies on Reflector strength, feedback quality, and applicability scope
ACE is not a universal solution. The following conditions weaken its effectiveness or make it unsuitable.
ACE assumes the Reflector can distill meaningful experiences from traces.
Feedback sources include:
Playbooks emphasize detailed domain specifics, but for some tasks this is redundant.
ACE's context is human-readable, naturally supporting:
Depends on Strong Reflector
ACE assumes the Reflector can distill meaningful experiences from traces.
- Weak Reflector → Playbook becomes noise or even misleading
- For some domains, any model struggles to distill useful insights → results are naturally mediocre
Depends on Reliable Feedback
Feedback sources include:
| Feedback Source | Strength |
|---|---|
| Ground-truth labels | Strong |
| Code execution success/failure | Strong |
| Formula matching | Strong |
| Environment rewards | Medium |
| No external signals | Weak |
Not All Tasks Need Long Playbooks
Playbooks emphasize detailed domain specifics, but for some tasks this is redundant.
| Task Type | Suitable for ACE |
|---|---|
| HotPotQA-style retrieval QA | Concise instructions suffice; Playbook would be redundant |
| Game of 24-style fixed-rule games | One rule suffices; no need for long context |
| Multi-step agent tasks | Suitable |
| Domain-specific reasoning (finance, law) | Suitable |
| Tool-use intensive | Suitable |
Relationship to Unlearning
ACE's context is human-readable, naturally supporting:
- Selective forgetting for privacy/legal reasons
- Deletion of outdated/incorrect information identified by domain experts
When to Choose Fine-Tuning Over ACE
- Long-term stable rules where training cost can be amortized
- Context cost exceeds weight training cost
- Requires implicit knowledge injection across tasks/prompts