Gemini 3.7 Flash vs DeepSeek V4 Pro (0813): Who Actually Wins Per Dollar?

summarized

TLDR

DeepSeek V4 Pro is currently the cheaper model by a wide margin, but on Sunday its price spikes up to 11x, collapsing the gap with Google's Gemini 3.7 Flash. Per-dollar winner depends on when you run and what you run: DeepSeek still wins on cost per finished task at all hours, but Gemini is faster and more stable for real-time products. The real takeaway is that sticker price shopping ignores token burn rates, cache discounts, and rate card expiry dates—both companies are renting you a number, not selling you one.

Key points

  • DeepSeek V4 Pro's current pricing is 43.5¢ per million input tokens, but on Sunday peak rates jump to $1.32 (3x) and the cache hit line—critical for agent loops—rises 12x, making the 'cheap' model reprice overnight.
  • Google Gemini 3.7 Flash has an introductory rate of 75¢ input / $3.75 output, with a published expiry on January 1st when rates double, and a cache discount of 90% (vs DeepSeek's 99.2%) with no storage fee on implicit caching.
  • Blended billing (70% cache, 20% fresh, 10% output) puts DeepSeek at 18¢/M tokens before the price change and ~35¢ off-peak after it, while Gemini sits at 58¢—the gap narrowing from 3.25x to 1.7x.
  • On the Artificial Analysis Intelligence Index, DeepSeek scores 53 and Gemini 56; cost per completed task is ~25¢ for DeepSeek vs 40¢ for Gemini, so DeepSeek still wins on pure cost even at peak rates.
  • Terminal Bench 2.1 scores diverge: DeepSeek claims 87.9 but independent testing at Artificial Analysis gets 79 (nine points lower), while Google reports 85.8—both vendor numbers, the gap likely due to DeepSeek's own agent harness.
  • Gemini outputs 340 tokens/second vs DeepSeek's 83, finishing a 10K token generation in 39 seconds vs 2 minutes; but DeepSeek has a 384K output token limit per call (6x Gemini's 64K) and shorter first-token latency (~2s vs ~10s).
  • DeepSeek caps concurrent requests at 500 vs Gemini's 2,500, and its weights are open-source (MIT license, 1.4M downloads), allowing provider switching if prices change again.
  • Gemini offers a batch tier at half price (~29¢ blended), undercutting DeepSeek's peak rates, and its pricing holds through December—making it more predictable for quarterly budgets.

Tools mentioned

Techniques

  • Cache discount pricing
  • Peak / off-peak billing
  • Blended billing (70/20/10 mix)
  • Agent harness benchmarking
Transcript (captions)

0:00 Two of the cheapest capable models on Earth shipped a day apart this week. One stops being cheap on Sunday. DeepSeek V4 Pro went generally available on August 12th at 43.5 cents per million output,

0:12 87 cents. A cash hit, a third of a cent. Frontier capability priced like a rounding error. A day later Google shipped Gemini 3.7 Flash. 75 cents in, 375 out. That is an introductory rate

0:26 and Google's own pricing page prints the expiry date on it. So on the sticker this comparison ends before it starts. DeepSeek is cheaper on all three lines, but the sticker is not the price. On

0:37 Sunday DeepSeek's rates move by up to 1,100% which means the answer to who wins per dollar has a shelf life measured in days. It expires this weekend. Per dollar has two halves

0:48 anyway. What a token costs and how many tokens it burns to finish your job. One of those is published on a pricing page. The other has to be measured and a third-party measured it. Those two

0:59 numbers do not pick the same winner which is the whole reason this video is longer than a tweet. So start with the money because the money moves first and it moves on a deadline that lands this

1:08 weekend. That 43.5 cents is a cash miss. The price of the first time the model sees your prompt. In an agent loop most tokens are not misses. An agent replays the whole conversation every turn. Same

1:20 system prompt, same tools, same files over and over. DeepSeek charges a third of a cent for that repeat. A discount of 99.2% OpenRouter measures the real hit rate on

1:31 V4 Pro between 89 and 92% which drags the average input token down to about 5 cents per million, about a ninth of list. Google discounts cached input too by 90%

1:43 rather than 99. 7.5 cents per million. Explicit caches also carry a storage fee. 50 cents per million tokens per hour running while your context sits there waiting to be reused. To be fair

1:55 to Google implicit caching is automatic and charges no storage. The meter only runs on caches you declare and hold. But on a cash token, DeepSeek is not a little cheaper. It is about 21 times

2:07 cheaper. Artificial analysis blends these on a 721 mix. 70% cash hits, 20% fresh input, 10% output, which is the closest published proxy for a real agent bill. DeepSeek V4 Pro lands at 18 cents

2:21 per million. Gemini lands at 58. Three and a quarter times the price for the same blended million, which makes this a route, and it was one until DeepSeek published a second rate card with a date

2:31 on it. At 1600 UTC on Sunday, August 16th, DeepSeek switches to peak and off-peak billing. With off-peak set at half the peak rate. The stated reason is capacity, pushing teams to schedule work

2:44 away from congested hours. Peak runs 1:00 to 4:00 and 6:00 to 10:00 UTC. 7 hours a day at the high rate, 17 at the low one, on a window shaped around Beijing working hours rather than yours.

2:56 Cash miss input goes from 43 and 1/2 cents to a dollar 32 at peak, about three times. Output goes from 87 cents to 396, four and 1/2 times. And the cash hit, the line an agent loop actually

3:09 lives on, goes from a third of a cent to four and four tenths, 12 times on the one number that mattered most. That is where the 1,100% headline comes from, and it is the cheapest token in the

3:20 comparison turning into the most repriced one. Off-peak is not a discount off today's price, which is the part the coverage keeps getting backwards. Off-peak cash

3:29 hits still cost six times what they cost right now. Therefore, the floor moved, not just the ceiling. Rebuild the blended bill on the new card. Off-peak DeepSeek comes to about 35 cents against

3:40 Gemini's 58. Still cheaper, but the gap collapsed from three and a quarter times to 1.7. At peak, DeepSeek lands at 69 cents and Gemini stays at 58. For 7 hours a day, the cheap Chinese model is

3:53 the more expensive token. Artificial analysis already list the new rates, so a tracker you read yesterday and one you read tomorrow will hand you opposite verdicts on the same model. And that

4:03 drift has history. Tracker spent months quoting $1.74 for V4 Pro, the April preview rate, after DeepSeek made a 75% promotional cut permanent in May. Four times wrong for a quarter. So, schedule

4:16 everything off-peak and take the win. That is the obvious move, and it is where most stop. The question they skip is what the cheaper token costs you everywhere else. Price per token is a

4:26 denominator. The numerator is whether the model finishes the task and how many tokens it burns getting there. On the Artificial Analysis Intelligence Index, nine evaluations, one harness, both

4:38 models, DeepSeek scores 53 and Gemini 56. Three points is not a landslide, and cost per completed task is the sharper number. There, DeepSeek runs about 25 cents to Gemini's 40. So, even after

4:50 Sunday, at peak rates, on the strictest per dollar measure published, DeepSeek is still the cheaper finished task. Hold on to that, because it is the strongest thing in DeepSeek's favor, and any

5:00 verdict has to survive it. But the same index run says something less flattering. Terminal Bench 2.1 is the agentic coding benchmark both camps are quoting. DeepSeek's own chart puts V4

5:11 Pro at 87.9, up from 72 on the preview. Artificial Analysis ran the same benchmark, same version, and got 79. Nearly nine points of daylight. Google's card puts Gemini at 85.8 on the same

5:25 test, also a vendor number, measured in Google's own setup, and it takes the same discount. The reason for DeepSeek's gap is printed in its own footnote. 87.9 was measured inside DeepSeek's agent

5:36 harness, at maximum reasoning effort, and they open-sourced that harness the same day, MIT licensed. So, the number is reproducible. It is just It's a score for a model.

5:46 It is a score for a model plus a scaffold. And if you run V4 Pro inside your own loop, 79 is what you budget for. Split by workload and it separates. On web development, Gemini posts 1,588

5:59 ELO on Code Arena, 50 points up on the last flash, and 43.6% on Frontier Code. On long horizon software engineering, the two claims nearly meet. 65.3 for Gemini, 62.7 for DeepSeek. Both are

6:14 vendor figures, and the two labs are not quoting the same version of the test. Speed cuts the other way from what the prices suggest. Gemini pushes about 340 output tokens a second. DeepSeek pushes

6:25 83. Four times the throughput measured on the same bench by the same people. DeepSeek answers sooner, under 2 seconds to first token against Gemini's nearly 10, because Gemini spends those first

6:36 seconds thinking, and thinking tokens bill as output. On a reasoning model, latency and cost are the same lever pulled twice. But on a 10,000 token generation, Gemini finishes in 39

6:47 seconds to DeepSeek's 2 minutes. Therefore, if a human is waiting at the keyboard, the cheaper token is buying you three times the wall clock. If it is a nightly batch, that cost is zero.

6:57 Shape matters, too. Both hold a million tokens of context, but DeepSeek will emit 384,000 output tokens in a single call against Gemini's 64,000. Six times the ceiling, and for some jobs, that

7:10 decides it before price enters the room. There is a ceiling that costs money in the other direction. DeepSeek caps V4 Pro at 500 concurrent requests against 2,500 on flash. If you are fanning out

7:21 an agent fleet, that limit is a bill of its own. And the weights are open on the same MIT license as the harness, 1.4 million downloads last month. 1.6 trillion parameters, 49 billion active,

7:34 trained on over 32 trillion tokens. If the API price moves again, you can move providers. Google has the lever DeepSeek lacks in the other direction. A batch tier and a flex tier each half the

7:45 standard rate. That puts Gemini's blended bill at 29 cents, cheaper than DeepSeek at peak, and three index points smarter. Which raises a model neither launch post

7:55 wants to discuss. DeepSeek's own V4 flash sits one point behind Pro on that same index at roughly a third of the price. By OpenRouter's June figures, flash already carried about 70% of

8:06 DeepSeek's agentic tokens. Google's price has an expiry date, too, and it belongs in the ledger. On the 1st of January, Gemini 3.7 flash doubles, $1.50 in, $7.50 out. Both of these

8:19 companies are renting you a number, not selling you one. So, here is the call. On raw cost per finished task, DeepSeek V4 Pro still wins at every hour of the day, even after Sunday, and it is still

8:31 the wrong default for most of you watching. It wins by about 1.6 times at peak and three times off-peak. That margin was nearly seven times before the new rate card. A number that moves that

8:42 far on days' notice is a rental, and you cannot build a quarterly budget on a rental you do not control. Gemini 3.7 flash wins the decision for anyone shipping product on a wall clock. Three

8:53 index points ahead, four times the throughput, a published price that holds through December, and a batch tier that undercuts DeepSeek's peak outright. The concession is real, and there are two of

9:04 them. If your workload is a cash-heavy loop you can schedule off-peak, V4 Pro is the cheaper token, and I would take it at twice Sunday's price. 384,000 output

9:14 tokens in one call has no Gemini equivalent at any tier. And DeepSeek shipped its harness open while both labs were still quoting their own charts. That is more of the measurement handed

9:24 over than anyone else gave you this week, and it deserves saying out loud. If you are optimizing purely for cost per coding task, buy neither flagship. V4 flash sits one index point behind Pro

9:35 for a third of the money. That is the knee of the curve, and the knee is where the recommendation lives. The villain is not a company. It is sticker price shopping, reading the pricing page as

9:45 though it were the bill. When the bill is tokens burned times the rate at the hour you ran it times which harness the benchmark was run in, which leaves the question worth sitting with. Deep Seek

9:55 built its name on a price no competitor could match, then raised one line of it 12-fold with days of notice. If the cheapest capable model on Earth can reprice over a weekend, what are you

10:05 actually building on?

Frontier News · by Hyperjump Technology