Skip to main content
Analysis

Baseline Reference

Four categories of context adaptation baselines compared against ACE

ACE compares against four categories of context adaptation methods in experiments, covering few-shot, prompt optimizers, and adaptive memory.

Baseline Overview

MethodTypeRequires GT
Base LLMNo adaptation–
ICLFew-shot / Many-shot examplesYes
MIPROv2Prompt optimizerYes
GEPAReflective prompt optimizerYes
Dynamic Cheatsheet (DC)Test-time adaptive memoryOptional

Base LLM

No context engineering; evaluated directly with official prompts. On AppWorld, uses the official ReAct implementation as a common starting point for all methods.

In-Context Learning (ICL)

  • Provides task examples in the prompt
  • Includes all training samples that fit in context; otherwise fills the window
  • Few-shot or many-shot

MIPROv2

  • Prompt optimizer provided by DSPy
  • Jointly optimizes system instructions and examples via Bayesian optimization
  • Experiments use auto="heavy" to maximize optimization effort

GEPA (Genetic-Pareto)

  • Sample-efficient optimizer based on reflective prompt evolution
  • Collects execution traces (reasoning, tool calls, intermediate outputs)
  • Uses natural language reflection to diagnose errors and propose prompt updates
  • Uses genetic Pareto search to retain frontier high-performance prompts
  • Achieves 35× fewer rollouts and 10-20% higher accuracy compared to RL methods like GRPO

Dynamic Cheatsheet (DC)

  • Test-time adaptive external memory
  • Accumulates reusable strategies and code snippets
  • Can self-update without GT labels
  • Experiments use cumulative mode (DC-CU)
DC carries the risk of context collapse from full rewrites; see specific cases in ACE Motivation.

Relationship to Agent Memory Systems

ACE is positioned as broad context adaptation, encompassing agent memory, system prompts, factual evidence, etc.; while other agent memory systems (AgentFly, AWM, A-MEM, Agentic Plan Caching) focus primarily on memory mechanisms.
SystemFocusRelationship to ACE
AgentFlyContinuously evolving memory + RLSimilar but focuses on training
AWMWorkflow reuseSimilar but only concerns workflows
A-MEMZettelkasten-style memoryACE's bullet concept inspired by it
Agentic Plan CachingPlan template cachingSimilar but only addresses efficiency
Learn about ACE's boundaries and failure modes: Known Limitations.