Skip to content

I tried to build a "token optimization stack" for coding agents. Here's why I killed it.

7.5 relevance
Score Breakdown
technical depth
8
novelty
8
actionability
7
community
6
strategic
6
personal
9

Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.

Deep technical post on token optimization for coding agents, directly relevant to AI/ML and developer tools.

AI/ML dev.to
I tried to build a "token optimization stack" for coding agents. Here's why I killed it.
Summary

A developer built a token optimization stack for coding agents (Graphify, Serena, LeanCTX, Caveman) to reduce Claude Code costs, but abandoned it after discovering that a claimed 97% savings metric masked silent failures. Two of five initial tools (Headroom and LiteLLM) failed under real benchmarks: Headroom required a complex proxy setup despite promising "no behavioral changes," and LiteLLM only offered load-balancing, not the complexity-based per-task routing needed. The pilot benchmark—31 tasks on claude-haiku-4-5 across 2 stack versions—cost $5.60 in API spend alone, and scaling to a rigorous 4,800-run evaluation across SWE-bench and Multi-SWE-bench would have been financially unsustainable.

Author

Shreyash

More from Shreyash →