The KV Cache Layer That Makes LLMs 10x Faster? (LMCache)

summarized

TLDR

The KV cache is the dominant cost in long-context LLM serving, and the built-in prefix caching in frameworks like vLLM stops being effective once the working set exceeds GPU memory. LMCache solves this by treating the KV cache as a multi-tier storage system—GPU memory, system RAM, local flash, and remote object storage—and moving it out of the inference engine into a shared process. The result is a 79% reduction in time-to-first-token and 264% higher input throughput on long shared contexts, but the technique only pays off above roughly 250,000 tokens of sustained working set; below that, it adds overhead and can degrade throughput by 10–30%.

Key points

In long agent sessions, 93–97% of tokens are identical across turns, yet most serving stacks recompute the full KV cache each time. Anthropic and OpenAI both price cached reads at 1/10 of normal input tokens, confirming the waste is real and measurable.

LMCache moves the KV cache out of the inference engine into a separate process with multiple storage tiers, enabling cache sharing across replicas. On an 8-GPU setup, mean time-to-first-token dropped from 3.98s to 0.29s, and decoding throughput nearly quadrupled.

The technique only helps when the sustained working set exceeds ~250,000 tokens. Below that, the built-in prefix cache is faster, and LMCache can reduce throughput by 10–30% due to lookup and transfer overhead.

Tools mentioned

Techniques

  • KV cache
  • prefix caching
  • multi-tier cache storage
  • cache sharing across replicas
  • Cache Blend
Transcript (captions)

0:00 Here's a coding agent late in a session, 115,000 tokens go into the model on this turn. 93 to 97% of them are word-for-word identical to what it sent one turn ago. Your graphics card reads

0:12 all of them again anyway. Every token from scratch as though it had never seen them. Your agent isn't broken and neither is the model. Serving stacks work this way and the waste lands on

0:22 your bill. There's an open-source project built to stop it sitting at 11,300 stars today. It's called LM Cache and Nvidia, AMD, and CoreWeave have all put money into the company that

0:34 maintains it. Before any of that lands, you need one word. The word is prefill. You send a prompt and the model doesn't answer token by token. It reads your whole input in one pass first. That read

0:45 is the prefill and it's where your time to first token goes. While it reads, the model writes two tensors for every token, a key and a value. That pair is the KV Cache. Keep the cache and

0:56 generation carries on. Drop it and your next request pays for that whole read again. Same words, same answer, full price. So, the shape of the problem is simple.

1:05 Reading is expensive, the notes are reusable, and most stacks throw the notes away as soon as memory gets tight. What follows is who fixed that and the exact point where the fix starts costing

1:15 you instead. How much waste are we talking about? A benchmark team replayed 739 anonymized Anthropic Claude code conversations against a live server on a pair of MDI cards. The median trace

1:28 opens around 20,000 tokens of input and ends near 115,000. Across the turns of one conversation, that median trace could serve 97% of its input from cache because the only thing

1:39 changing is the newest tool call. Reread it all every turn and their write-up puts it plainly. You waste 95% of the compute. You're not imagining the bill either. Anthropic prices a cache read at

1:51 1/10 of a normal input token. OpenAI charges the same 1/10 across its whole flagship line. Skipping the read knocks 90% off, which is two labs putting a number on what that reading was worth to

2:02 them. Which raises the obvious question about your own stack. The surprise is that it already does this, up to a point. vLLM ships automatic prefix caching, and it's on by default. It

2:13 hashes every block of your prompt, hangs on to the matching KV blocks, and skips the prefill for anything it recognizes. That's the feels instant on the second query effect, and it costs you nothing.

2:23 Hugging Face LLM ships the same trait under the name Radix Attention. So, that's the built-in fix, and it works. Now, the ceiling, which is where this whole story lives. Those cash blocks sit

2:33 in graphics card memory, the expensive fast stuff, and there isn't much of it to go around. One agent holding 100,000 tokens of context needs about 12 GB of card memory on the model those agent

2:44 traces ran against, purely for its notes. Multiply that by everyone hitting your service at once, and the card fills up. Then, the eviction policy starts deleting, and what goes is whatever

2:54 hasn't been touched for the longest, usually a conversation somebody is about to come back to. Your cache is also trapped inside one process. Restart the engine, and it's gone.

3:03 Route the next turn to a different replica, and it may as well not exist. Same user, same conversation, full price again, because the notes were sitting on a machine the load balancer didn't pick.

3:14 An engineer at Google ran the test that shows where the line falls, and published it with the LLM Cache team. Eight H100 cards, a 70 billion parameter Llama model, five shared prompt lengths,

3:25 three storage layouts head-to-head. Its first result is the one people skip past. When the whole working set fitted inside card memory, adding a slower tier bought nothing at all. The built-in

3:35 cache was already the right answer, and everything bolted on top was overhead. Step back, because two things are true at once. Skipping that read is worth 90% of an

3:44 input token, and the free fix stops helping the moment your context outgrows the card. So, what do do with a cache too valuable to recompute and too big to keep? In 2024, a group of researchers at

3:56 the University of Chicago started asking exactly that. What they built is LM Cache, and the idea sounds almost boring. Stop treating the KV cache as scratch memory and treat it as storage

4:06 with levels. Card memory on top, answering in microseconds. Under it, ordinary system memory, around 100 microseconds. Under that, local flash in milliseconds. At the bottom, remote

4:18 object storage, slower and enormously bigger at every step down. The bet is that fetching your old notes off one of those tiers beats computing them again. On that same Google run, past a certain

4:28 prompt length, it isn't close. At 5,000 tokens of shared context, time to first token drops 18%. At 10,000, 44. At 50,000, 68. At 100,000 tokens of shared context, it falls 79% and input

4:45 throughput climbs 264% on identical hardware. Read that as a curve, not a headline. The longer your shared context, the more this layer is worth. And below a few

4:55 thousand tokens, it's worth nothing at all. Every number in this video bends around that one shape. So, that's the storage side. The other half of the problem is who gets to see the cache. In

5:05 April, the project shipped a rebuild that moves the cache out of the engine and into its own process. Before that, eight workers on one machine each kept a private cache and none of them could see

5:15 the others. Identical context computed eight separate times. On a 235 billion parameter mixture of experts model across eight cards, mean time to first token went from 3.98

5:26 seconds down to 0.29. Tail latency improved more than 10 times, and decoding got nearly four times faster on top. 13 times on the mean, which is where the 10x headline

5:37 comes from. The mechanism is unglamorous. Eight private caches became one shared pool. No new kernel, no better model, just a cache that stopped being private. One of the project's

5:48 users hit a stranger version of the same result. The research paper anonymizes them, though LM Cash and CoreWeave have publicly documented the setup for the AI company Cohere. Prefill was eating them

5:59 alive, so they moved the KV cache onto remote object storage over the network, off the machine. Received wisdom says that has to lose because storage bandwidth runs an order

6:10 of magnitude under memory bandwidth. Instead, it came back 22 to 32% faster than a full prefill. The paper writes that up as an assumption the industry had been carrying and had never

6:20 rechecked. Remote object throughput went from about 100 megabytes a second to nearly a gigabyte, and that was enough to flip it. Prefix caching has one more blind spot. It only fires when the reuse

6:32 sits right at the front of your prompt. Retrieval systems break that constantly, stitching documents together in whatever order search returned them. A technique called Cache Blend reuses

6:42 blocks from anywhere in the prompt and recomputes only the tokens shifts for gains between 2.2 and 3.3 times lower time to first token. Which brings us to the part a marketing

6:53 page would leave out. To the credit of the people building this, they published it themselves on their own blog, raw numbers attached, in a post that spends most of its length on a test their

7:03 software wins. That same benchmark also ran a controlled sweep, fixed cache hit rates from 0 to 100% same model, same two cards, three configurations. At a 75% hit rate, the built-in cache managed

7:17 3,061 tokens a second. The dedicated cache layer managed 1,956. The post summarizes the shortfall as 10 to 17% on that row, it's near a third. Either way, it lost on a test built to

7:32 make caching look good. The explanation is this whole video in one sentence. Reuse was plentiful, but the card was under no memory pressure, so the extra tier never got reached and charged for

7:43 itself on every request anyway. Cache key handling, lookups, transfer checks, connector work, all paid in full for a slower tier the system had no reason to touch.

7:53 If you want a rule of thumb, that's it. A cache tier you don't need is a tax. That same report located the crossover on its own hardware, and it's a specific number, somewhere around 250 to 300,000

8:06 tokens of sustained working set. Underneath it, use what your engine already gives you. Above it, the layer pays off non-linearly. The research paper arrives at the same place from

8:16 another direction. Over a 32 gigabit link, fetching a cache only beat recomputing it once input past 256,000 tokens. Give the machine 64 gigabits and

8:27 loading one at every link they tried. So, your network speed decides whether any of this helps, which makes it an infrastructure question more than a model question. Would you replumb a

8:36 cluster for a speed up that only shows up past a quarter of a million tokens? Sharper edges exist. Python randomizes its hash seed on every process start, so two workers hash the same prompt into

8:47 two different cache keys, and every lookup misses. Bit identical prompts, zero hits, no error in the logs. The fix is a single environment variable, and the benchmark

8:57 team's guess is that a missing one causes most of the deployment failures they see. Not exotic, just silent and expensive for as long as it runs. One finding here should change how you write

9:06 prompts. Truncating a conversation to fit inside a context window cuts your hit rate roughly in half. The research team measured it on one customer's real traffic, 85% down to 45.

9:18 Slide the window forward, trim the front of a chat, reorder a config key, and every cache prefix behind it stops matching. So, the cheapest optimization available to you costs nothing and needs

9:28 no new software. Stop rewriting the beginning of your prompts. An LLM cache isn't unopposed. Nvidia's own serving framework ships a competing block manager, lists it first among the

9:39 offloading options in its docs, and wires LM cache in as one of the alternatives beside it. That's the same Nvidia whose venture arm wrote a check to the LM cache team. Both of those

9:48 things are true, and neither is a scandal. It does raise a question worth sitting with. Why did the rest of the industry converge on this project anyway when the hardware vendors shipped their

9:57 own? Because being the layer beats being the feature. It became a PyTorch ecosystem project in October 2025. Google built its tiered cache on Kubernetes with it. Nvidia's serving

10:09 framework wires it in through a dedicated connector. Two weeks after that AMD benchmark went up, the company behind the project raised $20 million on top of a 4 and 1/2

10:19 million seed. AMD Ventures, Coreweave, and Nvidia's venture arm all took part. Two chip makers and the cloud that rents their output funding one caching layer. That tells you what they think the

10:29 bottleneck is, and it isn't the accelerator. Junchao Jiang, a Chicago professor who co-founded the project and now runs the company built on it, describes the KEV cache as a whole new

10:39 class of data rather than a byproduct of inference. The commit log backs the framing up. 283 people have contributed. More than 1,500 commits landed in the past year, 33 of them in the past week.

10:51 So, here's the call. If you self-host on vLLM or edgy lang, and your sustained working set runs above a couple of hundred thousand tokens, install it. Agent platforms,

11:01 long document analysis, retrieval systems, that describes almost all of your traffic. The ceiling on that decision is a 79% cut in time to first token, and nearly

11:11 four times the input throughput. I'd take that deal at 10 times the operational hassle it actually costs. If your prompts are short, and your cache fits comfortably on the card, don't.

11:21 You'll surrender somewhere between a tenth and a third of your throughput to a tier you won't reach, and the benchmark proving it came from the project's own team on the project's own

11:30 blog. That last one is the receipt that convinces me. The Google curve, a remote cache beating a local prefill, and a vendor benchmark where the vendor loses. A project that publishes the test it

11:41 fails is one I'd put production traffic on. No company is the villain here. The habit is measuring inference in tokens per second on a fresh prompt when hardly any agent request is fresh

11:53 anymore. Your traffic is mostly the same context arriving again slightly longer each time. So, here's my bet with a date on it. By the end of 2027, cached input tokens will be free on at least one

12:05 major serving platform because charging for work you didn't do gets harder to defend every quarter. The company behind LM Cash already prices them at zero on its own service

12:14 and calls that permanent rather than promotional, which leaves the question this layer opens up. If your model's notes about a conversation now outlive the conversation, get written to disk,

12:25 and get shared between machines you don't own, whose data is that, exactly?

Frontier News · by Hyperjump Technology