A benchmark comparing token and cost outcomes with and without RTK (Rust Token Killer), a terminal-output-compressing plugin for AI coding agents, tested with Claude Code on Fable 5.0 and OpenCode with DeepSeek V4 Pro on Terminal-Bench 2.1. Results show mixed to negative effects: Fable's apparent savings depended almost entirely on a single task, while DeepSeek got 17% more expensive on average due to extra agent turns triggered by RTK's rewrites. The analysis argues RTK's self-reported

•7m read time•From quesma.com
Post cover image
Table of contents
How RTK worksTesting RTK on Terminal-Bench 2.1The first chart was promisingOne task made the difference in the whole benchmarkrtk gain is useless as a cost metricRTK bugs can bite youTerminal output is a small share of the billExtra turns can erase the savingsRTK does not make AI coding cheaper

Questions this post answers

Does RTK actually reduce the cost of running Claude Code or OpenCode agents?

Not reliably. Testing on Terminal-Bench 2.1 found Fable 5.0 with Claude Code was only 1% more expensive on a task-weighted basis, with almost all apparent savings coming from a single task (winning-avg-corewars), while OpenCode with DeepSeek V4 Pro cost 17% more on average with RTK enabled, mainly because RTK's output rewriting triggered more agent turns. Anyone deciding whether to adopt RTK for agent cost control can track benchmark writeups like this on daily.dev.

Why does RTK's reported token savings metric (rtk gain) not match actual cost savings?

rtk gain measures raw minus filtered command output in bytes divided by four, not billed tokens, so it can massively overstate savings. In one test, two head -1 calls accounted for 69% of a comparison's reported savings by comparing limited reads against a whole file that was never going to be returned, making an actually more expensive attempt look optimized. Developers weighing AI-agent cost metrics can follow methodology critiques like this via daily.dev.

Can a bug in an AI coding agent tool like RTK cause runaway costs?

Yes. A DeepSeek git-multibranch attempt using RTK 0.45.0 hit an unsupported find flag that got rewritten incorrectly on every retry, producing 339 consecutive errors before timing out and costing about 9 times as much as the matching baseline attempt that also passed. The bug was fixed in RTK 0.46.0, after the benchmark runs. Teams relying on agent tooling plugins can watch for gotchas like this by following coverage on daily.dev.

Share this post