Collaboration among Generator, Reflector, and Curator roles
ACE borrows the agentic design from Dynamic Cheatsheet, decomposing context adaptation into three roles.
Context is stored as a structured bullet list, with each bullet containing:
ACE supports both modes simultaneously:
The two modes can be combined: Offline warmup → Online adaptation for further improvement.
Three Roles at a Glance
| Role | Input | Output |
|---|---|---|
| Generator | Query + current Playbook | Reasoning trace (including strategies and errors) |
| Reflector | Trace + execution feedback | Specific experiences (insights) |
| Curator | Experiences | Delta context entries |
Workflow
Responsibilities of Each Role
Generator
- Faces new problems, produces complete reasoning traces
- Explicitly annotates which bullets are useful and which are misleading
- Feedback guides the Reflector
Reflector
- Extracts specific experiences from traces (successful strategies, failure patterns)
- Supports multi-round iteration (default max 5 rounds)
- Separated from Curator to avoid mixing "evaluation + organization" in one model
Curator
- Merges experiences into compact delta entries
- Uses deterministic non-LLM logic for merging, supporting parallelism
- Ensures existing knowledge is not erased
Playbook Composition
Context is stored as a structured bullet list, with each bullet containing:
| Field | Content |
|---|---|
| id | Unique identifier |
| helpful counter | Number of times marked useful |
| harmful counter | Number of times marked misleading |
| content | A small reusable unit of strategy, concept, or failure pattern |
Bullets are similar to memory entries in Dynamic Cheatsheet and A-MEM but add helpful/harmful counters to support subsequent deduplication and filtering.
Why Separate the Roles
- Generator focuses on reasoning; model attention is not diverted
- Reflector focuses on induction; avoids losing focus from "generating while summarizing"
- Curator focuses on merging; uses non-LLM logic to ensure determinism and parallelism
Offline vs Online Modes
ACE supports both modes simultaneously:
| Mode | Scenario | Description |
|---|---|---|
| Offline | Optimizing system prompts | Iteratively improves Playbook using training set |
| Online | Test-time memory | Predicts first, then updates per sample |
Key Implementation Settings
| Item | Value |
|---|---|
| Shared model for three roles | DeepSeek-V3.1 in non-thinking mode |
| Batch size | 1 (one delta generated per sample) |
| Reflector iteration limit | 5 |
| Offline epoch limit | 5 |