These 33 Lines Cut Claude Code Token Usage by 90%

summarized

TLDR

Spotify's open-source Shunt plugin cuts Claude Code token usage on file reads by 82-94% using a 33-line bash hook that blocks large file reads and delegates them to a cheap Gemini Flash model. The real saving is on money, not just tokens: about 86% off the bill for a single read, and over 91% across a 20-turn session when cached re-reads are included. But the mechanism has sharp limits — it can't delegate edits, the cheap reader misses bugs Claude catches instantly, and JetBrains' structurally identical experiment found it made runs 7.6% more expensive at low reasoning effort while quality didn't move.

Key points

Shunt uses a pre-tool-call hook to block reads of files over 350 lines and delegate them to a cheaper model.

The hook is 33 lines of bash and requires no vendor, proxy, or service sitting in front of traffic.

Spotify's benchmark shows 82-94% token reduction on file reads, with 90% being the mean of three scenarios.

The cheap reader missed a thread safety bug that Claude spotted in seconds, per Spotify's own testing.

JetBrains' structurally identical RTK experiment found 7.6% higher cost per task at low reasoning effort with no quality improvement.

Tools mentioned

Techniques

  • pre-tool-call hook
  • file size threshold gating
  • sub-agent delegation
  • bulk reader skill
  • mode-based routing
Transcript (captions)

0:00 Watch what happens when Claude code tries to open this file. 602 lines. It never gets there. Something stops the read before the file opens, then hands the job to a cheaper model instead.

0:12 Spotify built that gate, and the post about it traveled fast with one number stuck to it. 90% off Claude code token usage. That's the claim published earlier this month, and it's real. The

0:24 plugin behind it is open source, and it's been on GitHub since July. One correction first, because it changes what you believe. That post is written in the first person. Its own title says,

0:36 "My token usage." It's by Dmitri Masmanov, a principal product manager, not the engineering fleet. But the receipts are better than the post. There's a benchmark table in the repo

0:47 that the blog rounds off. The plugin prints its own counter, too. Remember that line, because we'll come back to it. Four scenarios on a Java mono repo. Three are file reads, and the

0:59 repo puts them at 82%, 94, and 94. So, 90 isn't a result. It's the mean of those three, and the repo's own line reads 82 to 94. The fourth scenario, code generation, has a dash in that

1:14 column. So, here's the plan. First, we'll take the machine apart, then we'll put the money back in. By the end, you'll know which half to steal and which half to leave. Start

1:25 with why it exists. A lot of what a coding agent does is fetching, not thinking. You ask one question about one function, and it opens whatever files it needs to answer. Everyone lands in the

1:37 context window, which is just a transcript the model carries. Think of it as notes spread across a desk. The desk is what you're paying for. Spotify's first fix was a list of

1:47 routing rules in a markdown file. Use the cheap model on big files, it said. Please. Those rules were advisory, so Claude could ignore them, and every project needed its own copy. Version two

2:00 stopped asking. Written rules are a suggestion, a block is not. Claude code lets you register a hook that runs before any tool call. And a hook is just an ordinary program on your machine.

2:11 Think of it as a turnstile instead of a sign. Shunt registers two of them. Here's how it works. One sits on the read tool. Claude code writes the pending call into that program as text

2:22 on standard input. Which file? What offset? What limit? The hook pulls those three fields out and reads nothing else. Then it runs four tests in order. And the first three exist to let the call

2:35 through. If the model already asked for a slice, an offset, or a limit, it goes straight past. It knows what it wants. If the path is empty or the file is missing,

2:45 past again, so the read tool reports its own error. Then it counts the lines. At or under 350, the default, the read is allowed because below that the detour costs more than it

2:57 saves. Otherwise, it prints a refusal and exits clean. The refusal names the numbers. This file is 602 lines, threshold 350. Then it names the fix. Use the bulk

3:10 reader skill, and if you need exact content for editing, reread with an offset for just the section you need. So, let's walk through one read. Claude calls read on that 602-line file with no

3:22 offset. The hook counts the lines, compares them against 350, and prints the block instead. Claude doesn't see one byte of the file. The read never runs.

3:34 From there, the reason string is handed back to the model as a result of its own attempt, so the instruction arrives at the moment of the attempt instead of sitting in a config file hoping to be

3:44 remembered. The whole gate is 33 lines of bash. You can read it in a minute. No service, no proxy, no vendor, and nothing sitting in front of your traffic. A second hook covers the

3:56 obvious hole because you could just run cat or head or tail. It lets a pipe through on purpose since a grep is already targeted and a redirect, too, because that output never enters the

4:08 conversation. Now the other half. Once the read is refused, the model calls a skill, which is a markdown file naming one script and its arguments. Claude doesn't have to assemble a pipeline out

4:19 of pros. That script builds its message on disk, not in memory. For each file, it streams an opening tag, cats the file in, closes the tag, and puts the question on the end.

4:31 It also checks the payload will physically fit because the request travels as one command line argument, and Linux caps those. Then a single call goes out to Spotify's

4:42 internal gateway addressed to a mode by name. A mode is basically a saved agent, some instructions, a model, a temperature. This one runs on a Gemini flash tier at a temperature of not point

4:54 two. Its entire instruction is this. You are a precise code analyst. Read the provided files and answer the question concisely. Output structured bullets only. No greetings, no pros, no

5:07 preambles. Lead every bullet with the exact name, type, or line number. After that, the answer comes back and gets checked three times, and the third check is the interesting one. It reads the

5:18 mode name out of the response because a stale mode doesn't fail loudly. It answers anyway with no instructions at all. So, if that field is empty, shunt throws the answer away.

5:29 What comes back is bullets, and only the bullets enter Claude's context. The corpus never does. Then the plugin prints that one line to your terminal, counting the input tokens it just

5:40 delegated. And on the benchmark that gave Spotify its best number, you can see what that buys. A source file and its tests 7 and 1/2 thousand lines together cost Claude about 76,000 tokens

5:52 to read directly. The bullets cost around 4,000, which is 94 and 1/2% fewer tokens. And I want to be exact about which tokens. They're tokens that would have entered Claude's context on one

6:05 file read. They're not dollars and they're not your bill, which raises a question worth answering before the arithmetic answers it for you. Would you let a cheaper model decide

6:15 what your expensive one is allowed to see? Hold your answer because a token isn't money until you name the model. Claude's top tier cost $5 per million tokens in.

6:26 The flash tier doing the reading cost 30 cents. That ratio, about 16 to 1, is the whole economic engine. Notice what the ratio doesn't say. Those 76,000 tokens don't vanish.

6:39 Somebody still reads every one of them just at 30 cents a million and then you pay again for the answer coming back. So, let's follow the money instead of the counter. This is list price

6:49 arithmetic, not anyone's invoice. Without the plugin that read cost about 38 cents. With it, you pay 2 cents into Claude, a couple more for the cheap read, and

7:00 about a penny for the answer. Call it 5 and 1/2 cents. That's 86% off the money against 94 and 1/2% off the tokens. Both are true, both measure the same event, and neither of them is a

7:13 saving because a file you read isn't billed once. It sits in the transcript and gets sent again on every turn after. Claude code caches that prefix, so the resends bill at a 10th of the price,

7:25 which sounds like nothing until you count the turns. Say you're 20 turns into that session. At list prices, the plain version has cost $1.14 and 2/3 of that is cached

7:36 re-reads of a file you already paid for. With the plugin, it's under 10 cents all in. 91% better than the one-shot number. So, the post undersells its own

7:47 mechanism. And the same sum is why you should be careful. Your bill scales with turns. Anything that adds turns eats the saving. And a 10 to 30-second detour that returns a summary the model

7:59 distrusts does exactly that. Which is why the best parts of the post are its limits. You can't delegate an edit because those summaries carry no reliable line numbers. And in Spotify's

8:11 own testing, the cheap reader missed a thread safety bug that Claude spotted in seconds. That's the line I trust most because it cost them something to say. So, two things so far. The block is real

8:23 and it's tiny. And the money is smaller than the headline. Which leaves the obvious question, if it's that good, why isn't it in the box? Half of it is, and here's the strange

8:33 part. The other half of Spotify's own system has no enforcement, either. Buried in the repo, under known limitations, sits a line that didn't make the blog post. No enforcement for

8:44 code writer. Read that next to the thesis. Written rules are a suggestion. A block is not. And the code-writing half relies on Claude noticing a description and choosing to use it. Both

8:56 sentences are Spotify's. Only one was in the post. Meanwhile, Claude code already ships a cheap reader. It's called explore. A sub-agent whose job is going and looking so the main

9:08 session doesn't have to. For a long time, it ran on the smallest model in the family. Then in July, it stopped. A release note says the built-in explore

9:18 agent now inherits the main session's model, capped at the top tier, instead of running on the small one. The docs still said otherwise weeks later. So, the in-the-box answer got more expensive

9:29 about 2 months before this post shipped. And that reframes the whole thing. Spotify didn't build a cheaper model. It built a place where the cheap path isn't optional. The Hacker News thread went

9:40 straight at that. One reader, Cave Tech, wrote that you can already use hooks to force sub agents, so the rest of the stack is unnecessary. Fair. But, the stack is where the enforcement lives.

9:52 Others went at the accounting. That reading didn't get cheaper, it moved to a different service with its own budget. "This isn't cutting Claude code token usage by 90%," wrote another. "It's

10:03 cutting file read token usage." Sharpest of the lot was about price. "Saving 90% of input tokens isn't saving 90% of tokens," wrote a reader called Ric Dobby, "because output costs far more.

10:16 And routing on file size tells you nothing about how hard the code is." Somebody called it a language model Bloom filter. A cheap index saying this might be the code you want. Another

10:27 reader noted that a Bloom filter has no false negatives, and a model does. Two days later, somebody else in that thread shipped the same idea for Cursor. All of which is opinion. And opinion is where

10:39 this whole category usually stops. So, here's the part I care about. Six weeks before the Spotify post, JetBrains ran the experiment on a structurally identical hook.

10:49 Their tool was called RTK. And the setup was serious. The same 86 task suite in both arms, the agent pinned to one version, four paired runs, 425 build trials, and the endpoints written down

11:03 before any run started. Advertised saving 60 to 90%. Measured result, 7.6% more expensive per task at low reasoning effort. At high effort, it was a coin flip. Quality didn't move either way.

11:19 The reason matters. Those runs used about 14% more turns and 14% more cache reads, while the one class of traffic the tool actually shrinks barely moved. More steps means the model rereads

11:31 everything so far more times. Meanwhile, the tool's own dashboard reported 96 million tokens saved. 96 million on runs that cost more. It counted the full raw output as its imaginary alternative,

11:45 estimated tokens by dividing characters by four, and couldn't see most of the context anyway. That's the sentence I'd put on the wall of this whole category, and it's theirs, not mine.

11:56 A tool's self-reported savings are a claim about its counterfactual, not about your bill. Now, go back to that line the plug-in prints. It counts what was sent, not

12:06 what was saved. So, why does Spotify's version survive that autopsy? Because of one number in the same post. JetBrains sorted nearly 2 million characters of real tool output into three piles. A

12:17 fifth of it their tool could compress. Nearly half was shell output it had no rule for. And 34% was file reading and search, the one pile their hook couldn't touch. That pile is exactly what

12:29 Spotify's gate catches. So, here's where I land. Take the hook. It's a few dozen lines of bash. It needs no vendor, and it's the only part of this that changes the model's behavior whether or not the

12:41 model agrees. If your agent reads big files on a big repo, that block pays for the afternoon it cost you. Leave the round trip, at least for now. That 10 to 30-second wait is real. The summaries

12:53 aren't exact enough to edit from, and I wouldn't put the cheap model anywhere near a change. Spotify says both of those out loud, which is most of why I believe the rest.

13:03 And that leaves the question the thread is really arguing about. One reader described the version that finally worked for them. The cheap model is only allowed to point, never to

13:13 decide. If that's all it does, is that still delegation, or have you just built yourself an index?

Frontier News · by Hyperjump Technology