Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Google's Gemini 3.7 Flash beats Claude Sonnet 5 on coding benchmarks at $0.75 per million tokens — half the price of competitors — but that price expires in January 2027. The catch is that the cheap model is a Flash-tier workhorse, while Google's actual flagship (Gemini 3.5 Pro) has missed three deadlines. This is a 140-day promotion, not a permanent price cut, and the real game is getting developers to spend more tokens on agent loops.
Key points
- Gemini 3.7 Flash scores 1588 on Code Arena's web development board, ahead of Sonnet 5 and other top models.
- Google's own flagship Gemini 3.5 Pro has missed three deadlines and hasn't shipped since February.
- The $0.75/million token price doubles to $1.50 on January 1, 2027 — a pre-announced promotion, not a permanent cut.
- Flash models are being released on a 23-day cadence, suggesting an agent-focused strategy where tokens are consumed in loops.
- Industry-wide, token consumption is growing faster than prices fall: agentic tasks can use 15 to 1000x more tokens than simple chats.
- Anthropic also ran a promotion and then made it permanent; Google's countdown is still running.
- On OpenRouter, American models' share of traffic dropped from 70% to 30% in 12 months, and Google's Flash model isn't even in the top 10.
- Despite Flash's strong coding scores, Google loses on many other benchmarks — 10 of 20 rows on its own comparison card go to competitors.
Tools mentioned
Techniques
- agentic token loops
- tiered pricing with expiry
- monthly model release cadence
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
On the 13th of August, Google shipped its best coding model of the year and half the price on the way out. 23 days earlier, it had shipped the previous one. That model is now sitting at half
price as well. Gemini 3.7 Flash 75 cents a million tokens in, 375 out on Google's own page. That is under half of what Anthropic charges for Claude 5 and under half of what OpenAI charges for Terra.
And the coding numbers move hard on Frontier Code. Google's card goes from 34.4 to 43.6. On Deep's SWE, the Long Horizon software engineering test, it climbs from 49% to 65.3. And on Code
Arena's web development board, it lands at 1588 ahead of Sonnet 5, Terra, and Meta's Muse Spark. Google's claim in its own words is that it beats comparable anthropic and open AI models across nine
separate benchmarks. But the same day Google published that, a second Google story ran and it points the other way entirely. Its actual flagship Gemini 3.5 Pro missed a deadline for the third
time. It was promised on stage in May. So, the cheap model is the news and the expensive one is still a rumor. That combination is the whole story here. And the price you just heard has an expiry
date printed underneath it on Google's own documentation. Start with the failure because on the 13th of August while the launch post went up, Forbes ran a piece under the headline Gemini
3.5 Pro delay continues. Google announced that model on stage at its developer conference in May for a roll out the following month. It missed June, it missed July. Forbes counts three
missed deadlines and reports the causes as coding and reliability work, senior researchers leaving, and the possibility that the pre-training run has to be restarted from the foundation. Google
has confirmed none of that. What Google has confirmed is the wait. Sundar Pichai told the July earnings call that 3.5 Pro was currently in testing and that the team was already building the next
generation of models. 3 weeks later, it was still in testing. The last ProClass model Google actually shipped was 3.1 Pro on the 19th of February. That is 175 days without a new flagship. And inside
that window, three Flash models went out the door instead, which shows up on the public scoreboard in a way that is difficult to look away from. On Code Arena's web development board, Google's
own 3.1 Pro preview sits at rank 46 with 1447. Its Flash sits at rank 8 with 1588 on a score the board still marks preliminary, 141 ELO points in favor of the model
that cost 75 instead of $2. Google's cheapest model is beating Google's most expensive one at the job this video is about, which turns the delay into a decision. And Pichai stated the decision
out loud on that same call. Picking up pace and releasing models almost at a monthly cadence, he said, is part of our road map as we are building Gemini 4 as well. The calendar backs him. 3.5 flash
on the 19th of May. 3.6 flash on the 21st of July. 3.7 flash on the 13th of August. three models at one tier in 86 days, each one landing cheaper on the output line than the last. And Tulsi
Dashi, who leads product for the Gemini model, described the upgrade in terms that give the strategy away. It better adapts to roadblocks, she said. Clarifies intent when needed and follows
instructions with greater fidelity. That is a description of an agent, not a chatbot. Which matters because the flash tier is where agent tokens actually get spent in loops that run all day rather
than the one big question you ask once. Google's model APIs went from 16 billion tokens a minute to 22 billion in a single quarter. That is about 37% more in 3 months. So, is this a price war?
Open Google's pricing page and read the sell instead of the headline. 75 cents through December 31st, 2026. $150 starting January 1st, 2027. The output line doubles on exactly the same date,
which makes this 140 days of half price, not a price cut, a promotion with a printed end date, and Google printed the end date itself, which is more than most launch posts bother to do. And it is not
only Google. Deepseek moves to peak and off- peak billing on the 16th of August, which raises the cached input line agents live on by about 12 times during peak hours. Meta lists a contributor
tier at 10 cents in and 20 cents out, roughly 16 times below its standard rate in exchange for permission to train future Meta models on your traffic. Anthropic ran the same play and then did
something nobody else did. Sonnet 5 launched at $2 in and 10 out as introductory pricing through the 31st of August with 3 and 15 scheduled for September. Anthropic's own documentation
now says that increase will not occur. The promotion became the price. One Lab took its countdown off. Google's is still running. Which brings us to the part of the title that matters most and
the part that is hardest to answer with a leaderboard. If the sticker price keeps falling, why do the bills keep climbing? Because the token counts multiplied faster than the prices fell.
Anthropic's own research puts a multi- aent system at roughly 15 times the tokens of a simple chat. Work out of Microsoft and Stanford's digital economy lab puts agentic tasks near a thousand
times a chat message. two figures, two orders of magnitude apart, pointing the same direction. So, a price cut at the workhorse tier is not charity aimed at the small developer. It is an invitation
to spend more tokens, and the invitation works, which is why Google's own meter is climbing while its own prices fall. For a person who is not buying tokens by the billion, though, the picture is
better than it sounds. Flashclass models are free in Google Studio, rate limited, but free. And in May, Google split its $250 Ultra plan into a $100 tier and a $200 one. The bottom of this market is
getting cheaper quickly. The top of it is not, and the free tier still has a price. Google's own pricing table carries a row that says prompts on the free tier are used to improve our
products, and prompts on the paid tier are not. Cheap has a currency here, and sometimes the currency is your traffic. So, is Google behind Open AI, Anthropic, and the Chinese Labs or in front of
them? Both. And the board is precise about which is which. At the top of Code Arena's web development ranking, Claude Opus 5 max at 1691, Kim K 3AX at 1674, QN 3.8 max at 1669. Anthropic first,
then two Chinese labs, and four more entries before Google appears at rank 8 with a flash. On tokens actually consumed, it reads worse. The share of open router traffic going to American
models fell from about 70% to about 30 in the 12 months to June. And Google is not in that platform's 10 mostus models at all. And on Google's own comparison card, count the rows, 20 benchmarks, and
10 of them go to somebody else. Terra takes deep sf. Both versions of terminal bench and OS world. Sonnet 5 takes agents last exam. 3.7 Flash even loses to the model it replaces at reading
charts. So, here is the call. At the workhorse tier, the tier where most agent tokens are actually spent, Google is not behind anybody this week. And the thing that got it there is a 23-day
release cycle rather than a benchmark table. For the builder shipping product against a wall clock, 3.7 flash is the model I would reach for today, and I would still reach for it at double the
money, which according to Google's own pricing page is what it becomes on the 1st of January. The concession is the frontier, and it is a real one. If your work needs the best model on Earth, it
is Opus 5. And two of the three names at the top of that board are Chinese. Google's answer at that tier is a model it has now failed to ship three times. The villain in this story is not Google,
which printed its expiry date where anyone could read it. It is the idea that a launch day price is a price. My bet on the record, Google does not let that January doubling stand and the
workhorse price stays at or under a dollar past New Year. So ask yourself what you are buying at 75 cents a model or 140day trial of