Skip to main content
Analysis

Known Limitations

ACE's dependencies on Reflector strength, feedback quality, and applicability scope

ACE is not a universal solution. The following conditions weaken its effectiveness or make it unsuitable.

Depends on Strong Reflector

ACE assumes the Reflector can distill meaningful experiences from traces.
  • Weak Reflector → Playbook becomes noise or even misleading
  • For some domains, any model struggles to distill useful insights → results are naturally mediocre
This shares the same root cause as Dynamic Cheatsheet's limitation: memory quality depends on the model itself.

Depends on Reliable Feedback

When neither GT nor reliable execution signals are available, both ACE and DC may degrade. Context can be polluted by false signals.
Feedback sources include:
Feedback SourceStrength
Ground-truth labelsStrong
Code execution success/failureStrong
Formula matchingStrong
Environment rewardsMedium
No external signalsWeak

Not All Tasks Need Long Playbooks

Playbooks emphasize detailed domain specifics, but for some tasks this is redundant.
Task TypeSuitable for ACE
HotPotQA-style retrieval QAConcise instructions suffice; Playbook would be redundant
Game of 24-style fixed-rule gamesOne rule suffices; no need for long context
Multi-step agent tasksSuitable
Domain-specific reasoning (finance, law)Suitable
Tool-use intensiveSuitable

Relationship to Unlearning

ACE's context is human-readable, naturally supporting:
  • Selective forgetting for privacy/legal reasons
  • Deletion of outdated/incorrect information identified by domain experts
This is also a unique advantage over weight adaptation (fine-tuning).

When to Choose Fine-Tuning Over ACE

  • Long-term stable rules where training cost can be amortized
  • Context cost exceeds weight training cost
  • Requires implicit knowledge injection across tasks/prompts
ACE's core value is cheap, interpretable, instantly updatable—not replacing all weight updates.
Next: Compare Mem0 and ACE: Mem0 vs ACE.