Google Built a Future Predicting Model. So Why Ban Us From Using It?

summarized

TLDR

Times FM3, Google's new time series forecasting model, achieves roughly 36% lower error than a naive baseline by reading multiple data series at once and predicting the entire forecast horizon in a single pass. However, this result is flattening: nine agentic forecasting systems that use a language model to dynamically choose a method now rank above it on the broadest public leaderboard, suggesting that the field may be moving past foundation-model-style forecasters toward systems that select the right small model for each task. The weights ship under a strict non-commercial license, while the cheaper, open Kronos 2 model is within 1.4 skill points and is a more practical choice for most commercial users.

Key points

Times FM3 forecasts the entire requested horizon in a single forward pass without autoregressive decoding.

Google's model achieves an 87.2% average win rate on the FevBench but ranks only 10th on the broader GiftEval leaderboard.

The two top-ranked systems on GiftEval are agentic forecasting pipelines, not pure foundation models.

Times FM3's weights are released under a non-commercial license that prohibits production use and commercial decision-making.

An independent study found foundation models beat classical methods only on 15 of 30 datasets, with classical methods overtaking them with as little as 2% of training data.

Kronos 2, Apache 2.0 licensed and four times faster per series, loses to Times FM3 by just 1.4 skill points on FevBench.

Tools mentioned

Techniques

  • patch-based tokenization for time series
  • bivariate attention across time and variate dimensions
  • single-pass prediction of entire forecast horizon
  • quantile regression output (10th-90th percentiles)
  • look-ahead encoding of known future covariates
Transcript (captions)

0:00 How much ice cream will this chain sell next month? Here's the ordinary answer. Take last month's shape and push it forward. Now, the same month from a model that also got the promotion

0:09 calendar. Same history, same store, and it jumps on every promotion day. That second line is Google's newest model, and it drew the whole month in one shot, not day by day, one pass. It's called

0:21 Times FM3, and Google put it on the internet on the 31st of August, 2026. 330 million parameters. That's roughly a thousandth the size of the models you type questions into. It doesn't know a

0:34 single word. It hasn't read a sentence. Ask it anything at all and you'll get silence back. What it read was a trillion numbers pulled off a clock. Sales, sensors, page views, electricity

0:45 load. And on the hardest public forecasting test I know of, it's now the best model that has ever been submitted. So, how it actually works, what it's really for, and where it loses. That's

0:56 the next 20 minutes because there's a second leaderboard on that one. Times FM3 comes 10th. Nine systems sit above it and the one in first place was also built at Google by a different team with

1:07 a different idea. The part that works is strange though, so start there. Forecasting is the oldest boring problem in software. How many units next quarter? How much power at 7 p.m.? How

1:18 many servers on Black Friday? For decades, the answer was a statistical model you fit to one series by hand, one at a time. A thousand products meant a thousand models. Then somebody asked the

1:29 obvious question. Language models learned from every sentence ever written. Could a forecaster learn from every series ever recorded? Google's answer arrived in February 2024 as a

1:41 paper with four authors and a name, a decoderonly foundation model for time series forecasting. The idea underneath it is the one worth understanding, and it's smaller than you'd think. A

1:51 language model reads text as tokens, roughly a word, sometimes a piece of one. It stacks them up, looks back across all of them, and guesses what comes next. Times FM does the identical

2:02 thing, except its token isn't a word. Its token is a patch, 32 consecutive numbers off your series, glued together, and treated as one unit. 32 hourly readings is a day and a bit. 32 daily

2:14 readings is a month. So, one token is roughly one chunk of shape. A dip, a spike, a weekend. That's the whole trick. Chop the past into shapes. Learn from a trillion points which shapes tend

2:26 to follow which other shapes, which is why it can forecast a business it hasn't seen. It isn't recognizing your company. It's recognizing that your Tuesday looks like 10 million other Tuesdays.

2:37 Everything after that is engineering. And this year, the engineering changed twice. Here's what changed first. Every version of this model up to last year could only look at one series at a time,

2:46 which is a strange way to run a business. Ice cream sales don't move alone. They move with cones, with syrup, with foot traffic, with the weather, with whatever you put on sale. Google's

2:56 own writeup says it plainly. Most real world forecasting problems are inherently multivariate. Multiple series and outside features jointly deciding what happens next. So version three

3:06 reads a grid. Time runs one way. Your different series run the other. And attention runs in both directions. Across time, it's strictly one way. A token can only look backwards at its own

3:17 series. That's not politeness. That's a leak prevention rule. The model must never see a value it wouldn't have on the day. Across series, it's wide open. At any given moment, every series can

3:28 look at every other series. That's how a promotion on cones informs the forecast for cones and syrup. Those two attention passes alternate layer after layer, 20 layers deep. Temporal, variate,

3:39 temporal, variate. And your inputs now come in three flavors, which is the part that matters for actual work. Targets are the things you want predicted. You can hand it several at once, and it

3:49 forecasts them together. Past coariantss are things you measured but can't know ahead. Yesterday's foot traffic, last week's weather that actually happened. Past future coariantss are the good

3:59 ones. Things you already know about the future. Your promotion calendar, public holidays, a weather forecast, the schedule you yourself control. That third flavor gets special treatment

4:09 inside the model. A normal token holds one patch. A known future token holds its own patch stitched to the next two. Google calls it a look ahead. And the phrasing in the post is nice. Each token

4:20 concatenates the current patch with future patches, letting the model peak at upcoming known signals, which is the whole ice cream demo from a minute ago. The blue line jumps on promotion days

4:30 because the promotion days were legally visible to it. The red line couldn't see them, so it drew a normal week. On that example, Google's chart shows about a 20% sales bump anticipated on each

4:40 promotion day, anticipated, not discovered afterwards. Quick check on where we are. patches for tokens, a two-way grid of attention, three kinds of input. That's the reading half. The

4:51 writing half is where the title comes from. Older versions wrote the future the way a language model writes a sentence. One patch, then the next, each one conditioned on the last, which has

5:01 the same three problems it has in text. It's slow because you can't parallelize a loop. It compounds error because a mistake at step one poisons step 9, and it costs more. Version 3 refuses to do

5:13 that. Before it runs, it staples blank placeholder tokens onto the end of your data, one for every chunk of future you ask for. Then it runs once through those 20 alternating layers. Every blank gets

5:24 filled at the same time. Each one seeing all the others. That's what the title means. The whole horizon in a single forward pass. Not the whole future, the whole horizon you requested, but all of

5:35 it at once with no loop. And in that same pass, your known future coariants stay visible while the targets stay masked. The model can see the promotion, it can't see the sales. That asymmetry

5:46 is the entire design. One more thing comes out and most coverage skips it. It doesn't hand you a number. For every future step, it emits nine numbers. The 10th percentile through the 90th, a

5:57 floor, a middle, a ceiling, and the shape in between. So that is the whole machine reading half and writing half. That's not decoration either. A retailer stocking to the median goes out of stock

6:08 half the time. Stalking to the 90th percentile is a different, much more expensive, much more sensible business decision. So, the accurate one-s sentence description isn't an AI that

6:19 predicts the future. It's a quantile regression model with a transformer inside it. Which brings us to the comparison everybody wants. You've heard this thing called a foundation model,

6:28 and it does share a backbone with the chat bots, but line them up and almost nothing else matches. A language model has a vocabulary. 50,000 100,000 discrete symbols and every step it picks

6:40 one. There is a right answer and a wrong answer. Times FM has no vocabulary. Its inputs are real numbers and its outputs are real numbers. Nothing is chosen. Something is estimated. A language model

6:52 samples. Turn the temperature up and it gets creative. This thing is deterministic. Same series in, same forecast out every time. A language model is auto reggressive by nature.

7:02 Take that away and it can't write. This model had auto regression taken away on purpose and got better. Size is where the gap gets silly. A frontier language model is hundreds of billions of

7:13 parameters. This is 330 million. Small enough that the weights fit comfortably on a laptop. And the training data isn't the internet. It's public forecasting data sets. Wikipedia page views up to

7:24 late 2023 and the top Google Trends queries up to the end of 2022, plus a large pile of synthetic series. Read that list again because it tells you what the model actually knows. It knows

7:36 what human attention looks like through a week and a year. It has no idea what your product is. There's a deeper reason it can't be a language model. And it's about data, not architecture. Text is

7:46 effectively infinite. There's another book, another forum, another decade of archive. A business series is not. A company that's four years old has 48 monthly data points and that is the

7:57 entire universe. You cannot brute force 48 numbers with a 100 billion parameters. There's nothing there to overwhelm the noise with scale. The thing that made language models work is

8:08 exactly the thing forecasting can't have. So the sensible question is whether all this machinery is worth it. And this is where the numbers stop being press release numbers and start being

8:17 interesting. Three public benchmarks matter here. The first is Fev Bench from Amazon's forecasting group. 100 real forecasting tasks across seven domains, 46 of which come with covariants. On its

8:29 public leaderboard, Times FM3 sits first with an average win rate of 87.2% and a skill score of 48.7. Second place is Kronos 2 from Amazon at 82.1 and 47.3. Third and fourth are two

8:44 more open models, roughly three and four skill points further back. go head-to-head on that board and Times FM3 beats Kronos 2 on 69% of the 100 tasks. Not a route, a clear, repeatable edge.

8:58 Against its own predecessor from 2025, it wins 88% of them. And against seasonal naive, the dumbest possible forecast, next Tuesday looks like last Tuesday. It wins 100 out of 100 every

9:10 task. Hold on to that number because it comes back. The second benchmark is called time. 50 fresh data sets, 98 tasks built specifically so that no preprint model could have seen the

9:22 answers. There everything gets scored against seasonal naive on purpose. So seasonal naive is exactly 1.0 times FM 3 scores 0.640 Kronos 2 0.662

9:36 its own previous version 0.669 669. And here's the family history in one column. Version 1, back in 2024, 0.788. Version 2, 0.718. Version 2.5, 0.669.

9:52 Version 3, 0.640. That is a real curve going the right way over 2 years and 7 months with 10 times the training data. It is also, and we'll come back to this, a curve that is

10:04 flattening. So step back, the model works. The architecture is clever, the benchmarks are real, and the wins are consistent. What does 0.640 actually buy you? It means that averaged across the

10:16 whole benchmark, its error is 36% lower than the dumbest forecast in the room. Sit with that for a second. A trillion training points, 330 million parameters, 2 and 1/2 years of a Google research

10:28 team, and the headline result is roughly a third better than same as last Tuesday. Now 36% is not nothing. In a supply chain, a third off your forecast error is millions of dollars and a lot

10:39 of unsold stock. Anyone who's run inventory will take it. But it's also not what a breakthrough looks like. And the person who said so said it first in March, months before this model shipped.

10:50 A forecaster writing under the name Shako published an essay in March of 2026, arguing against time series foundation models as a whole. His background is the reason it landed. term

11:01 structure forecast at the Federal Reserve, then supply chain at Amazon, then four years at Stripe on cohort and financial forecasting. His argument was arithmetic. On FEV bench, he wrote, "The

11:12 strongest models reduce error against seasonal naive by about a third." That's not nothing, but it's also not what a major breakthrough looks like. And then the sentence that has aged unusually

11:22 well. If you bring millions of parameters trained on millions of series and you beat a decent seasonal baseline by roughly a third, the natural conclusion is that the amount of

11:31 learnable structure is fairly limited. Under 6 months later, the best model ever submitted landed at 36%. He called the ceiling before the ceiling was measured. His explanation is worth

11:42 hearing even if you disagree with it. The depth of structure to learn from language or audio or video, he argues, simply doesn't exist in a time series. There's a second sharper version of the

11:52 same point. Benchmarks are made of series somebody chose to collect, clean, and publish, which means the wild ones aren't in there at all. So, part of what a pre-printed model buys, you might not

12:03 be deep insight at all. It might be a very good prior against forecasting something absurd, a learned refusal to forecast a 10,000% spike because it has seen what normal looks like. That's a

12:13 real product. It's just a different product from understanding your business. And the evidence backs the modest reading. A break even study out of Carls Institute of Technology in July

12:23 compared foundation models against classical methods across 30 data sets at every training size. Foundation models won outright on 15 of 30. On six, classical methods overtook them with as

12:34 little as 2% of the training data. On the other nine, break even landed anywhere from 24 samples to over 8,000. They did find one rule that holds up. If you have under 700 training points and

12:46 there's real seasonality, use the foundation model zero shot and don't bother fine-tuning. Fine-tuning short series with low rank adapters actively made things worse, which is the real

12:57 shape of this technology. It's a superb default when you have little data, no time, and a lot of series. It's not a replacement for knowing your domain. So, does that make times FM 3 a foundation

13:08 model or an extremely well-trained curve fitter carrying a very good prior? Google would say the first. The leaderboard is about to say something more awkward than either. The third

13:18 benchmark is gift eval run by Salesforce and it's the broadest of the three. 127 systems on the board statistical baselines and deep learning and everything since Google's post is

13:29 carefully worded about it. Times FM 3 it says is the top ranked model among all pre-trained foundation models. Read that clause again because the qualifier is doing real work. At this point you have

13:41 two boards saying two different things. Open the actual one and times FM3 is 10th. Its average rank is 26.0. Nine entries are above it. Every one of those nine is tagged the same way.

13:53 Agentic. They aren't forecasting models. There are systems where a language model looks at your data, decides what kind of series it's dealing with, picks a method, fits it, checks it, and hands

14:03 you a forecast. First place is a system called Stride with Synapse at an average rank of 13.6. six, nearly twice as good a rank as Times FM3. Second is a forecasting agent from LG AI research

14:16 and stride is built by Google Cloud AI research. Google is beating Google and the thing doing the beating isn't a model, it's a procedure. Now go back to March. The forecaster who called that

14:26 ceiling didn't stop there. He spent the back half of that essay describing what should replace foundation models. His proposal was an agent per account. Something that looks at the data

14:36 pipelines, forms an explicit hypothesis about the structure, picks a small model that encodes it, fits it, and explains why. He wrote that 174 days before the gift of alboard filled up with exactly

14:48 that, nine of them, above Google's best. He had a reason to, and it's the strongest idea in this whole story. He built a forecasting system for about 800,000 series and the whole thing came

14:59 to roughly 4 million parameters, mostly small local ones inside an explicit structure you can read. His summary of the alternative is blunt. You're burning millions of parameters to try to learn

15:10 the optimal 10 parameter structural representation. There's a version of that critique that's just cynicism. And this isn't it because he names the case where the big model wins. If you need a

15:20 halfdeent forecast from an API and don't have time to build your own system, these models are fine, which is most people. Most people do not have four years and a research budget. That's the

15:30 actual market. And times FM3 is very good at serving it. There's a catch though, and it's the reason a trade headline about this launch read, "Google's new forecasting model beats

15:40 everyone. You can't use it at work. Every previous version shipped under Apache 2.0, free for anything, including making money." The two closest rivals, Kronos 2 and Toto, still do. Version 3

15:52 doesn't. The weights ship under something called the Times FM non-commercial license version 1.0, and it's stricter than that name suggests. The code stays open. The weights are for

16:03 testing, evaluation, and research not tied to commercial gain, academic work, internal benchmarking, experiments on your own data. All fine, but the license spells out what isn't. any revenue

16:15 generating activity, any interaction with end users or production systems, and training or distilling another model for commercial use taken at face value. It goes further than don't ship it.

16:25 Internal benchmarking is allowed only if the results aren't used in commercial decision-m, so a company can't legally use it to decide anything. You also can't redistribute it at all. The grant

16:36 is revokable. A commercial license exists in theory at Google's sole discretion, possibly for a fee or a revenue share. Your forecast at least are yours. The license is explicit that

16:47 outputs aren't derivatives. It's the weights that are on a leash. And this isn't accidental positioning. Google says the BigQuery integration is landing in the coming weeks. And today that

16:56 command runs on version 2.5. So the pattern is clean and it's worth naming without being cynical about it. The best weights go behind a restriction. The paid path through the data warehouse

17:06 stays wide open. That's a business model, not a betrayal. One more thing a careful viewer should know. There's no Times FM 3 paper. The model card cites the original 2023 architecture paper.

17:18 And everything we know about its benchmark results comes from a launch post and public leaderboards. Those leaderboards are third party and reproducible, which is a lot. But

17:27 there's no task level table, no ablation, nothing showing which of the two new ideas, the variate attention or the single pass decode earned the win. There's also a limit sitting in the

17:37 config file that isn't in the announcement. The VA attention is capped at 32 series at once. Fine for a product line, not a warehouse. And the single pass for all its elegance doesn't make

17:49 it fast. On Febench, it takes about 3.7 seconds per 100 series. Kronos 2 takes 0.8. Tyrex 2 takes 0.27. So, the best model on that board is also roughly four times slower than the runnerup and 13

18:03 times slower than third place. One pass is about error, not about latency. That leaves the practical question. Here's what you can do about any of this on Monday morning in order of how little

18:14 work it is. If your data already lives in BigQuery, you're one line of SQL away. The AI.cast command is generally available. It runs times FM 2.5 today and you pay for it like any other query

18:26 at the standard analysis rate with a monthly free tier that covers small jobs outright. The same model is wired into Aloy DB so a forecast can happen next to your rows instead of after an export.

18:38 Register the endpoint. Call the forecast function. Done. If you want weights you own, pip install xfm gets you version 2.5 under Apache 2.0 200 million parameters, a 16,000 point context

18:51 window and no lawyer required. If your problem has coariates, and most real ones do, Kronos 2 is the pick. 12 million parameters, Apache 2.0 multivariat, and second on the same

19:04 boards FM3 tops. And if you want to see what all this is like before committing to anything, version 3 is one download away legally for evaluation. Point it at your own history and compare it to your

19:15 current forecast. That comparison is exactly what the license permits. Which brings me to what I actually think, and I'm not going to hedge it. Kronos 2 wins for 95% of people watching this. It's

19:26 nearly three times smaller. four times faster per series, Apache 2.0, and it loses to Times FM 3 by 1.4 skill points on the benchmark Google chose to lead with. 1.4 points is real. It is not

19:39 worth a license that forbids you from using the result to make a decision. The 5% it doesn't win for research groups, benchmark authors, anyone whose job is measuring the frontier rather than

19:50 shipping on it for them. Times FM three is the new reference point and the license doesn't bite. And I'd switch back the day Google puts the weights back under an open license or the day it

20:00 lands in Big Query because on the accuracy question, Google is right. It wins 69% of head-to-heads against the best rival anyone has shipped and 100 out of 100 against the baseline. That's

20:12 earned. The thing I'd argue against isn't Google. It's the sentence beats everyone. That sentence describes a rank on one board under one filter. And the same week it was written, nine Agentic

20:23 systems were sitting above this model on a different board with 127 entries in it. Both facts are true. The best pre-trained forecaster on Earth is now Google's. And the best forecasting

20:35 system on the broadest public board is a language model with a strategy, not a forecaster with weights, which is the thing worth chewing on. For 2 years, the field's answer to how do we forecast

20:45 better was train a bigger model on more series. That answer is now producing a third of an improvement over same as last Tuesday and flattening. The other answer, the one the leaderboard is

20:55 filling up with, is to stop building the forecaster and start building the thing that chooses the forecaster. So, the open question isn't whether times FM3 is good. It measurably is. It's whether

21:06 Google just shipped the best version of an idea that's already being replaced. And what does it say? That the team replacing it works down the

Frontier News · by Hyperjump Technology