Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
MIT's superposition analysis shows that LLM loss scales as roughly 1/width, meaning doubling a model's width halves crowding error but more than doubles cost. For many coding and writing tasks, test-time compute can match or beat larger models at a fraction of the cost, making it the cheaper lever to pull before scaling up.
Key points
MIT research finds that superposition of features in LLMs creates a noise floor that scales as 1/width.
Doubling model width cuts the superposition noise in half, but costs more than double to train and run.
A 2024 Berkeley and Google DeepMind study found that test-time compute lets a small model match one 14x its size on easy problems.
For a real coding task, a smaller model with docs and test feedback can keep pace with a larger one-shot model.
Hard new problems still require larger models; test-time compute helps most on problems the small model already has a chance at.
Tools mentioned
Techniques
- superposition
- test-time compute
- weight decay
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
Watch a model answer a question it should get right. It's summarizing a short function and it says with complete confidence that the code retries three times when a request fails. Now look at
the function. There's no retry anywhere in it. This isn't a hard prompt and it isn't a weak model. It's a capable one describing a detail the function doesn't contain. And when that happens, the fix
we reach for is usually the same. Get a bigger model. How much does a bigger model actually fix? It's a sharper question than it sounds. and there's a second one hiding behind it. Is size
even the best lever you've got? A result out of MIT gives the first half a surprisingly exact answer. Bigger really does help in a way you can put a number on. But that same number shows why it
helps a little less every time you pay for it. And it points straight at a lever most developers rarely reach for. Start with why bigger works at all. Because it really does. Since about
2020, labs have planned around scaling laws. Feed a model more parameters and more data. and its error drops on a smooth predictable curve. That error has a name, the loss, which just measures
how wrong the model's predictions are on average. Lower loss, better model. This one curve is why companies raised billions to build something bigger. It has paid off for years, but a smooth
curve isn't an explanation. For years, the scaling laws were basically an observation. We could measure that loss falls as models grow without really knowing why it falls so cleanly. That's
the gap a team at MIT set out to close. Their paper first went up in May 2025. So, this isn't breaking news. Think of it as a lens on the curve all those billions are chasing. Their answer has a
name, superposition, and it comes with a built-in ceiling. I'll define it properly in a moment. The short version, a model represents far more ideas than it has room for, and that crowding
leaves a floor of noise it can't fully remove. Making the model wider lowers the floor, but slowly on a curve you can write down. One caveat up front. This is about the loss, not hallucinations
directly and not a hard cap on intelligence. So superposition start with how a model stores anything at all. It turns each word into a list of numbers, a vector that points somewhere
in an internal space. That space has a fixed number of directions. And that count is basically the model's width. A small model might have a few thousand of them. Every idea the model knows has to
live as some direction in that space. Now the squeeze. A model needs to represent far more distinct features than it has directions to spare. Not a few thousand ideas, but many times more
than the few thousand directions it's working with. It's like having a few thousand shelves and tens of thousands of labeled boxes to store. You can't give each box its own shelf. There
simply isn't the room. So, the model does something clever. It stores several features along the same shared directions as vectors that don't quite line up. That trick has a name,
superposition, which just means packing in more things than you have room for. The idea traces back to Anthropic's toy models work from 2022. Two features sharing space aren't perfectly separate,
so they bleed into each other a little. That bleed is called overlap. And that overlap is the whole story. Every time two features share space, reading one back picks up a faint smear of the
other. across thousands of packed features. Those smears add up to a steady hum of noise on everything the model represents. That hum is a floor under the loss. Even a perfectly trained
model keeps a little error purely from the crowding. It's working through a low hum of interference the whole time. Now the payoff and it's the useful part. Give the model more directions to work
with a wider model and the packed vectors have more room to spread apart. As they spread, their overlap shrink the math in the paper is clean. The noise falls as roughly one over the width.
Double the width and you roughly have the error that comes from crowding. That right there is the scaling law falling straight out of geometry. And this isn't just a toy. The team checked real open
models families like OPT, Quinn 2.5, and Pythia spanning about 100 million up to 70 billion parameters. The overlaps really do shrink as one over the width. The loss really does track it. When they
fit the exponent, they get about 0.91 close to the clean value of one. The theory predicts the geometry holds in real language models. One knob controls all of this weight decay, the training
setting that pushes weights towards zero. Turn it up and the model gets picky. It stores the common features cleanly and drops the rare ones. Turn it down and the model cra with heavy
overlap. Real models run in that crammed regime. So that's why bigger works. Now, the part that should change how you spend money. Look closely at that curve, though, because one over the width is a
slow road. To cut the crowding error in half, you double the width. To have it again, you double it again. But doubling a model's width more than doubles what it costs to train and to run. So, each
real step down in error cost you more than the last one did. You're paying more and more for less and less. This is where the title needs a caveat because the careful version matters more. The
paper measures loss, the model's average prediction error. Loss tracks quality, but it isn't the same thing as a hallucination rate. The work also doesn't prove a hard ceiling on how
smart models can get. It explains one specific cost that grows as you scale using a toy model and a check on real ones. That's a strong result. It just isn't the end of intelligence. So, does
money fix it? Partly, and that's the fair answer. A bigger model really is a better model, and the loss really does fall. But you're buying linear gains at exponential prices and that trade gets
worse the higher you climb. Brute scale is the most expensive lever in the room which raises the obvious question. What are the cheaper ones? Consider a different move. Instead of pouring money
into a wider model, spend it while the model is answering. The width is locked in once training ends, but how hard the model works on your specific question is up to you. That answer time compute is a
lever most people leave sitting on the table. This one is called test time compute. And the idea is simple. Give the model room to work before it commits to an answer. Let it draft several
attempts instead of one. Let it check those attempts and revise them. You're trading a bit more time and money per question for a better answer without touching the model size at all. There
are two main ways to spend it. One, generate several candidate answers, then score them with a separate checker, a verifier, and keep the best. Two, let the model read its own answer and
rewrite it in passes, fixing what's weak. Search and pick or revise and improve. Both spend compute at answer time instead of in the parameters. How well does that actually work? A 2024
study from Berkeley and Google DeepMind put numbers on it. Spent well testime compute beat the naive approach by more than four times for the same budget. And on easier problems, a small model given
room to think matched a model 14 times its size at the same total compute. Same hardware bill. Very different answer. But the same study is careful and so am I. This edge shows up on easy and medium
problems where the small model already had a real shot. On the hardest questions, the ones the small model just can't touch. More thinking doesn't rescue it. There, the bigger model still
wins. Test time. Compute is a lever, not a magic wand. So, put both levers on the same bench. This is a test any developer can run, and it's the practical heart of it. Take one real coding task against a
library the model doesn't know. Cold. Config. A is the reflex. Hand the whole thing to a big frontier model. One shot, no help. Config B uses a smaller, cheaper model, but you give it the
library's real docs and the output from the failing tests, and you let it iterate. Look at what config B actually is. The docs are context, the right information in front of the model at the
moment it needs it. The failing tests are a verifier, a cheap automatic check on whether the code is correct. Read, try, check against the tests, fix, repeat. That's test time compute applied
by hand on a laptop. So which one wins? That's the experiment and it gets reported straight on correctness, on dollars, and on wall clock time. Config B reads the docs, runs the tests,
watches them fail, and edits until they pass or until it gives up. The comparison is logged and shown on screen. Whichever way it lands, the measured result belongs right here from
the actual run with no thumb on the scale. And this part holds no matter who wins. If the big model wins outright, then some problems really are worth the premium, and now you know which ones. If
the small setup keeps pace, you just matched a far bigger model for a fraction of the cost using context and a test loop you already had. Either way, the cheaper lever earned its place next
to raw size. So, where does this leave you the next time a model gets something wrong? Take a side. For most everyday coding and writing tasks, before you pay for a bigger model, spend the compute
you already control. Give the model the right context. Let it check its own work against something real. That path is cheaper. It's faster to try and often it closes the gap on its own. Now steelman
the other side because it's real. Some problems are truly hard and new. The kind where a small model had no real chance. Those still reward the bigger model and the research agrees. And the
MIT limit we started with is about loss, not a wall on what models can eventually do. Bigger still matters. It just isn't the only move anymore or the first one. Come back to that invented retry line
from the very start. A bigger model gets that kind of thing wrong less often and it charges you every single time it answers. But a smaller model that can open the real file and check its own
claim against the code catches the same mistake for far less. The crowding the MID team measured doesn't fully go away. What changes is how cleverly you work around it. So that's the shift worth
keeping. Size still buys quality, but on a curve that charges more for every step for a reason we can finally see. The bigger lever most days is how you use the model, not how large it is. If
breakdowns like this are useful to you, subscribe. One clear explainer like this every week.