Thinking of ACE? We Can Do It with Fewer Tokens

This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).

IBM Research compares ALTK-Evolve and ACE (Agentic Context Engineering), two systems that let LLM agents learn from their own past trajectories without weight updates. Both systems avoid compressing learned lessons into summaries, instead maintaining countable guidelines or playbooks. The key difference is delivery: ACE injects its full playbook on every inference step, while ALTK-Evolve selectively retrieves only the guidelines relevant to the current task and model capability. On the AppWorld benchmark with 168 tasks, ALTK-Evolve achieves 89.3 TGC vs ACE's 80.4 on DeepSeek-V3.2 at ~40% of ACE's token cost (263K vs 634K tokens/task). On the weaker gpt-oss-120b model, accuracy is roughly tied (56.0 vs 54.8 TGC) but ALTK-Evolve uses about one-seventh the tokens (116K vs 777K). The post argues that calibrated delivery — matching the volume of injected guidelines to what a given model can actually absorb — is the key to both efficiency and accuracy gains.

•8m read time•From huggingface.co
Post cover image
Table of contents
What we agree onWhere we differWhy it mattersSame lessons, different deliveryLinked artifacts / referencesMethod notes

Questions this post answers

How does ALTK-Evolve compare to ACE (Agentic Context Engineering) on token cost and accuracy for LLM agents?

On AppWorld with DeepSeek-V3.2, ALTK-Evolve scores 89.3 TGC / 80.4 SGC versus ACE's 80.4 / 73.2, using 263K tokens per task versus ACE's 634K, about 40% of the cost. On gpt-oss-120b, ALTK-Evolve edges ACE 56.0 to 54.8 TGC while using only 116K tokens versus ACE's 777K, roughly one-seventh the cost. The difference comes from delivery: ACE injects its full playbook every step, while ALTK-Evolve selectively retrieves guidelines per task. Developers weighing agent memory approaches can track cost-versus-accuracy comparisons like this one on daily.dev.

What is the difference between ACE's playbook approach and ALTK-Evolve's guideline retrieval for agent memory?

ACE consolidates lessons into one comprehensive playbook via a Generator-Reflector-Curator loop and injects the whole thing at every inference step regardless of model or task. ALTK-Evolve clusters and merges near-duplicate lessons support-conservingly into typed, retrievable guidelines, then delivers a small fixed core plus a per-task selected subset, or the full set only when a model has the headroom to use it. Teams designing agent memory pipelines can compare consolidation and delivery strategies discussed on daily.dev.

Does giving an LLM agent more context or memory always improve its task accuracy?

No. On gpt-oss-120b, a weaker model, injecting ACE's full playbook helped on easy and medium AppWorld tasks but selective, per-task retrieval of guidelines won on hard tasks and on the aggregate score, because a large context can overwhelm a weaker model rather than help it. On the stronger DeepSeek-V3.2 model, more delivered lessons kept helping instead of crowding each other out. Engineers tuning how much context to feed an agent can follow benchmark findings like these on daily.dev.

Share this post