Spotify cut Claude Code token usage by 90%. Here's what I got when I tested it.
A cheap model does the reading and the boilerplate. Claude keeps the thinking. I rebuilt the setup for plain Claude Code, then ran it on 4 real tasks, with and without.
Based on the setup a Spotify product manager published in September 2026. Theirs runs on Portal and Gemini Flash. This one needs only the claude CLI and uses Haiku. Not affiliated with Spotify.
code-writethe part that pays
Haiku writes predictable code (tests, type stubs, config) from a spec plus a reference file, straight to disk. Claude never reads the reference or the output.
bulk-readnever called
Sends whole files to Haiku in one call and returns dense bullets with line numbers. The files never enter Claude's context.
block-big-readsnever fired
A PreToolUse hook that stops any full read over 350 lines and points Claude to the two scripts. Targeted reads pass.
2 skills + benchmark
Skills that tell Claude when to call each script, and the 4 scenarios with the runner, so you can reproduce every number below.
# clone it, install it, restart Claude Code
git clone https://github.com/ToolMonsters/claude-code-routing
cd claude-code-routing && ./install.sh
- Copies the scripts, the hook and the skills into
~/.claude, and adds the hook to~/.claude/settings.json. Your existing settings are kept, and a backup is written first. - Settings:
BLOCK_MIN_LINES(default 350) ·CHEAP_MODEL(defaulthaiku) ·CLAUDE_BINifclaudeis not on your PATH. - No API key, no extra platform. The cheap model is called through the Claude CLI you already have.
Claude Code 2.1.270, Opus 5, on psf/requests. Each task ran once without the setup and once with it. Total cost includes the Haiku calls.
| Task | Without | With | What happened |
|---|---|---|---|
| S1 Inventory of 110 classes and functions across 3 files | $0.75 | $0.71 | 110/110 both. Neither run opened a file. |
| S2 Every raise and except in 2 files | $0.74 | $0.85 | Same answer. Without: 2 full file reads. With: ranged reads instead. |
| S3 Write tests matching a 3,094-line test file | $0.96 | $1.22 | Worse. 17 tests down to 11. 90s up to 243s. |
| S4 Write a type stub for a 1,184-line module | $0.82 | $0.59 | Cheaper and better. 51 of 57 functions covered vs 42. 54s up to 164s. |
One run per cell. Treat small gaps as noise.
- The hook never fired. Without the setup, Claude read big files in full in 2 of 4 tasks. With it installed, Claude avoided full reads on its own, so there was nothing left to block.
- The cheap reader saved nothing. Claude never called it, and on S2 the run with the setup cost more.
- The cheap writer paid, on boilerplate. On S4 Claude never opened the module, Haiku wrote the stub, and it came out 28% cheaper and more complete. On S3, which needed judgment about an existing style, it cost more and did worse.
- The 90% is a real number. It measures file tokens that stop entering Claude's context. In these 4 tasks the reading side barely moved the bill. The writing side is where the money was, and it costs time.
- A ranged read (offset and limit) passes the hook whatever its size. On S2 Claude read 545 lines at once through that door.
- Haiku calls run as separate
claude -pprocesses, so they do not show in the session's/cost. The scripts print their cost. - Never route debugging, architecture or anything subtle to the cheap model. Spotify saw the same: their worker missed a thread-safety bug Claude caught in seconds.
benchmark/run.sh opus clones psf/requests and reruns the 8 sessions. Budget about $6 to $7 with Opus.Everyone copies the reader.
The writer is the part that pays.
Your next customers are already around you.
We build GTM systems that turn your audience, network and market signals into sales conversations. toolmonsters.com
