Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Times FM3, Google's new time series forecasting model, achieves roughly 36% lower error than a naive baseline by reading multiple data series at once and predicting the entire forecast horizon in a single pass. However, this result is flattening: nine agentic forecasting systems that use a language model to dynamically choose a method now rank above it on the broadest public leaderboard, suggesting that the field may be moving past foundation-model-style forecasters toward systems that select the right small model for each task. The weights ship under a strict non-commercial license, while the cheaper, open Kronos 2 model is within 1.4 skill points and is a more practical choice for most commercial users.
Key points
Times FM3 forecasts the entire requested horizon in a single forward pass without autoregressive decoding.
Google's model achieves an 87.2% average win rate on the FevBench but ranks only 10th on the broader GiftEval leaderboard.
The two top-ranked systems on GiftEval are agentic forecasting pipelines, not pure foundation models.
Times FM3's weights are released under a non-commercial license that prohibits production use and commercial decision-making.
An independent study found foundation models beat classical methods only on 15 of 30 datasets, with classical methods overtaking them with as little as 2% of training data.
Kronos 2, Apache 2.0 licensed and four times faster per series, loses to Times FM3 by just 1.4 skill points on FevBench.
Tools mentioned
Techniques
- patch-based tokenization for time series
- bivariate attention across time and variate dimensions
- single-pass prediction of entire forecast horizon
- quantile regression output (10th-90th percentiles)
- look-ahead encoding of known future covariates
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
How much ice cream will this chain sell next month? Here's the ordinary answer. Take last month's shape and push it forward. Now, the same month from a model that also got the promotion
calendar. Same history, same store, and it jumps on every promotion day. That second line is Google's newest model, and it drew the whole month in one shot, not day by day, one pass. It's called
Times FM3, and Google put it on the internet on the 31st of August, 2026. 330 million parameters. That's roughly a thousandth the size of the models you type questions into. It doesn't know a
single word. It hasn't read a sentence. Ask it anything at all and you'll get silence back. What it read was a trillion numbers pulled off a clock. Sales, sensors, page views, electricity
load. And on the hardest public forecasting test I know of, it's now the best model that has ever been submitted. So, how it actually works, what it's really for, and where it loses. That's
the next 20 minutes because there's a second leaderboard on that one. Times FM3 comes 10th. Nine systems sit above it and the one in first place was also built at Google by a different team with
a different idea. The part that works is strange though, so start there. Forecasting is the oldest boring problem in software. How many units next quarter? How much power at 7 p.m.? How
many servers on Black Friday? For decades, the answer was a statistical model you fit to one series by hand, one at a time. A thousand products meant a thousand models. Then somebody asked the
obvious question. Language models learned from every sentence ever written. Could a forecaster learn from every series ever recorded? Google's answer arrived in February 2024 as a
paper with four authors and a name, a decoderonly foundation model for time series forecasting. The idea underneath it is the one worth understanding, and it's smaller than you'd think. A
language model reads text as tokens, roughly a word, sometimes a piece of one. It stacks them up, looks back across all of them, and guesses what comes next. Times FM does the identical
thing, except its token isn't a word. Its token is a patch, 32 consecutive numbers off your series, glued together, and treated as one unit. 32 hourly readings is a day and a bit. 32 daily
readings is a month. So, one token is roughly one chunk of shape. A dip, a spike, a weekend. That's the whole trick. Chop the past into shapes. Learn from a trillion points which shapes tend
to follow which other shapes, which is why it can forecast a business it hasn't seen. It isn't recognizing your company. It's recognizing that your Tuesday looks like 10 million other Tuesdays.
Everything after that is engineering. And this year, the engineering changed twice. Here's what changed first. Every version of this model up to last year could only look at one series at a time,
which is a strange way to run a business. Ice cream sales don't move alone. They move with cones, with syrup, with foot traffic, with the weather, with whatever you put on sale. Google's
own writeup says it plainly. Most real world forecasting problems are inherently multivariate. Multiple series and outside features jointly deciding what happens next. So version three
reads a grid. Time runs one way. Your different series run the other. And attention runs in both directions. Across time, it's strictly one way. A token can only look backwards at its own
series. That's not politeness. That's a leak prevention rule. The model must never see a value it wouldn't have on the day. Across series, it's wide open. At any given moment, every series can
look at every other series. That's how a promotion on cones informs the forecast for cones and syrup. Those two attention passes alternate layer after layer, 20 layers deep. Temporal, variate,
temporal, variate. And your inputs now come in three flavors, which is the part that matters for actual work. Targets are the things you want predicted. You can hand it several at once, and it
forecasts them together. Past coariantss are things you measured but can't know ahead. Yesterday's foot traffic, last week's weather that actually happened. Past future coariantss are the good
ones. Things you already know about the future. Your promotion calendar, public holidays, a weather forecast, the schedule you yourself control. That third flavor gets special treatment
inside the model. A normal token holds one patch. A known future token holds its own patch stitched to the next two. Google calls it a look ahead. And the phrasing in the post is nice. Each token
concatenates the current patch with future patches, letting the model peak at upcoming known signals, which is the whole ice cream demo from a minute ago. The blue line jumps on promotion days
because the promotion days were legally visible to it. The red line couldn't see them, so it drew a normal week. On that example, Google's chart shows about a 20% sales bump anticipated on each
promotion day, anticipated, not discovered afterwards. Quick check on where we are. patches for tokens, a two-way grid of attention, three kinds of input. That's the reading half. The
writing half is where the title comes from. Older versions wrote the future the way a language model writes a sentence. One patch, then the next, each one conditioned on the last, which has
the same three problems it has in text. It's slow because you can't parallelize a loop. It compounds error because a mistake at step one poisons step 9, and it costs more. Version 3 refuses to do
that. Before it runs, it staples blank placeholder tokens onto the end of your data, one for every chunk of future you ask for. Then it runs once through those 20 alternating layers. Every blank gets
filled at the same time. Each one seeing all the others. That's what the title means. The whole horizon in a single forward pass. Not the whole future, the whole horizon you requested, but all of
it at once with no loop. And in that same pass, your known future coariants stay visible while the targets stay masked. The model can see the promotion, it can't see the sales. That asymmetry
is the entire design. One more thing comes out and most coverage skips it. It doesn't hand you a number. For every future step, it emits nine numbers. The 10th percentile through the 90th, a
floor, a middle, a ceiling, and the shape in between. So that is the whole machine reading half and writing half. That's not decoration either. A retailer stocking to the median goes out of stock
half the time. Stalking to the 90th percentile is a different, much more expensive, much more sensible business decision. So, the accurate one-s sentence description isn't an AI that
predicts the future. It's a quantile regression model with a transformer inside it. Which brings us to the comparison everybody wants. You've heard this thing called a foundation model,
and it does share a backbone with the chat bots, but line them up and almost nothing else matches. A language model has a vocabulary. 50,000 100,000 discrete symbols and every step it picks
one. There is a right answer and a wrong answer. Times FM has no vocabulary. Its inputs are real numbers and its outputs are real numbers. Nothing is chosen. Something is estimated. A language model
samples. Turn the temperature up and it gets creative. This thing is deterministic. Same series in, same forecast out every time. A language model is auto reggressive by nature.
Take that away and it can't write. This model had auto regression taken away on purpose and got better. Size is where the gap gets silly. A frontier language model is hundreds of billions of
parameters. This is 330 million. Small enough that the weights fit comfortably on a laptop. And the training data isn't the internet. It's public forecasting data sets. Wikipedia page views up to
late 2023 and the top Google Trends queries up to the end of 2022, plus a large pile of synthetic series. Read that list again because it tells you what the model actually knows. It knows
what human attention looks like through a week and a year. It has no idea what your product is. There's a deeper reason it can't be a language model. And it's about data, not architecture. Text is
effectively infinite. There's another book, another forum, another decade of archive. A business series is not. A company that's four years old has 48 monthly data points and that is the
entire universe. You cannot brute force 48 numbers with a 100 billion parameters. There's nothing there to overwhelm the noise with scale. The thing that made language models work is
exactly the thing forecasting can't have. So the sensible question is whether all this machinery is worth it. And this is where the numbers stop being press release numbers and start being
interesting. Three public benchmarks matter here. The first is Fev Bench from Amazon's forecasting group. 100 real forecasting tasks across seven domains, 46 of which come with covariants. On its
public leaderboard, Times FM3 sits first with an average win rate of 87.2% and a skill score of 48.7. Second place is Kronos 2 from Amazon at 82.1 and 47.3. Third and fourth are two
more open models, roughly three and four skill points further back. go head-to-head on that board and Times FM3 beats Kronos 2 on 69% of the 100 tasks. Not a route, a clear, repeatable edge.
Against its own predecessor from 2025, it wins 88% of them. And against seasonal naive, the dumbest possible forecast, next Tuesday looks like last Tuesday. It wins 100 out of 100 every
task. Hold on to that number because it comes back. The second benchmark is called time. 50 fresh data sets, 98 tasks built specifically so that no preprint model could have seen the
answers. There everything gets scored against seasonal naive on purpose. So seasonal naive is exactly 1.0 times FM 3 scores 0.640 Kronos 2 0.662
its own previous version 0.669 669. And here's the family history in one column. Version 1, back in 2024, 0.788. Version 2, 0.718. Version 2.5, 0.669.
Version 3, 0.640. That is a real curve going the right way over 2 years and 7 months with 10 times the training data. It is also, and we'll come back to this, a curve that is
flattening. So step back, the model works. The architecture is clever, the benchmarks are real, and the wins are consistent. What does 0.640 actually buy you? It means that averaged across the
whole benchmark, its error is 36% lower than the dumbest forecast in the room. Sit with that for a second. A trillion training points, 330 million parameters, 2 and 1/2 years of a Google research
team, and the headline result is roughly a third better than same as last Tuesday. Now 36% is not nothing. In a supply chain, a third off your forecast error is millions of dollars and a lot
of unsold stock. Anyone who's run inventory will take it. But it's also not what a breakthrough looks like. And the person who said so said it first in March, months before this model shipped.
A forecaster writing under the name Shako published an essay in March of 2026, arguing against time series foundation models as a whole. His background is the reason it landed. term
structure forecast at the Federal Reserve, then supply chain at Amazon, then four years at Stripe on cohort and financial forecasting. His argument was arithmetic. On FEV bench, he wrote, "The
strongest models reduce error against seasonal naive by about a third." That's not nothing, but it's also not what a major breakthrough looks like. And then the sentence that has aged unusually
well. If you bring millions of parameters trained on millions of series and you beat a decent seasonal baseline by roughly a third, the natural conclusion is that the amount of
learnable structure is fairly limited. Under 6 months later, the best model ever submitted landed at 36%. He called the ceiling before the ceiling was measured. His explanation is worth
hearing even if you disagree with it. The depth of structure to learn from language or audio or video, he argues, simply doesn't exist in a time series. There's a second sharper version of the
same point. Benchmarks are made of series somebody chose to collect, clean, and publish, which means the wild ones aren't in there at all. So, part of what a pre-printed model buys, you might not
be deep insight at all. It might be a very good prior against forecasting something absurd, a learned refusal to forecast a 10,000% spike because it has seen what normal looks like. That's a
real product. It's just a different product from understanding your business. And the evidence backs the modest reading. A break even study out of Carls Institute of Technology in July
compared foundation models against classical methods across 30 data sets at every training size. Foundation models won outright on 15 of 30. On six, classical methods overtook them with as
little as 2% of the training data. On the other nine, break even landed anywhere from 24 samples to over 8,000. They did find one rule that holds up. If you have under 700 training points and
there's real seasonality, use the foundation model zero shot and don't bother fine-tuning. Fine-tuning short series with low rank adapters actively made things worse, which is the real
shape of this technology. It's a superb default when you have little data, no time, and a lot of series. It's not a replacement for knowing your domain. So, does that make times FM 3 a foundation
model or an extremely well-trained curve fitter carrying a very good prior? Google would say the first. The leaderboard is about to say something more awkward than either. The third
benchmark is gift eval run by Salesforce and it's the broadest of the three. 127 systems on the board statistical baselines and deep learning and everything since Google's post is
carefully worded about it. Times FM 3 it says is the top ranked model among all pre-trained foundation models. Read that clause again because the qualifier is doing real work. At this point you have
two boards saying two different things. Open the actual one and times FM3 is 10th. Its average rank is 26.0. Nine entries are above it. Every one of those nine is tagged the same way.
Agentic. They aren't forecasting models. There are systems where a language model looks at your data, decides what kind of series it's dealing with, picks a method, fits it, checks it, and hands
you a forecast. First place is a system called Stride with Synapse at an average rank of 13.6. six, nearly twice as good a rank as Times FM3. Second is a forecasting agent from LG AI research
and stride is built by Google Cloud AI research. Google is beating Google and the thing doing the beating isn't a model, it's a procedure. Now go back to March. The forecaster who called that
ceiling didn't stop there. He spent the back half of that essay describing what should replace foundation models. His proposal was an agent per account. Something that looks at the data
pipelines, forms an explicit hypothesis about the structure, picks a small model that encodes it, fits it, and explains why. He wrote that 174 days before the gift of alboard filled up with exactly
that, nine of them, above Google's best. He had a reason to, and it's the strongest idea in this whole story. He built a forecasting system for about 800,000 series and the whole thing came
to roughly 4 million parameters, mostly small local ones inside an explicit structure you can read. His summary of the alternative is blunt. You're burning millions of parameters to try to learn
the optimal 10 parameter structural representation. There's a version of that critique that's just cynicism. And this isn't it because he names the case where the big model wins. If you need a
halfdeent forecast from an API and don't have time to build your own system, these models are fine, which is most people. Most people do not have four years and a research budget. That's the
actual market. And times FM3 is very good at serving it. There's a catch though, and it's the reason a trade headline about this launch read, "Google's new forecasting model beats
everyone. You can't use it at work. Every previous version shipped under Apache 2.0, free for anything, including making money." The two closest rivals, Kronos 2 and Toto, still do. Version 3
doesn't. The weights ship under something called the Times FM non-commercial license version 1.0, and it's stricter than that name suggests. The code stays open. The weights are for
testing, evaluation, and research not tied to commercial gain, academic work, internal benchmarking, experiments on your own data. All fine, but the license spells out what isn't. any revenue
generating activity, any interaction with end users or production systems, and training or distilling another model for commercial use taken at face value. It goes further than don't ship it.
Internal benchmarking is allowed only if the results aren't used in commercial decision-m, so a company can't legally use it to decide anything. You also can't redistribute it at all. The grant
is revokable. A commercial license exists in theory at Google's sole discretion, possibly for a fee or a revenue share. Your forecast at least are yours. The license is explicit that
outputs aren't derivatives. It's the weights that are on a leash. And this isn't accidental positioning. Google says the BigQuery integration is landing in the coming weeks. And today that
command runs on version 2.5. So the pattern is clean and it's worth naming without being cynical about it. The best weights go behind a restriction. The paid path through the data warehouse
stays wide open. That's a business model, not a betrayal. One more thing a careful viewer should know. There's no Times FM 3 paper. The model card cites the original 2023 architecture paper.
And everything we know about its benchmark results comes from a launch post and public leaderboards. Those leaderboards are third party and reproducible, which is a lot. But
there's no task level table, no ablation, nothing showing which of the two new ideas, the variate attention or the single pass decode earned the win. There's also a limit sitting in the
config file that isn't in the announcement. The VA attention is capped at 32 series at once. Fine for a product line, not a warehouse. And the single pass for all its elegance doesn't make
it fast. On Febench, it takes about 3.7 seconds per 100 series. Kronos 2 takes 0.8. Tyrex 2 takes 0.27. So, the best model on that board is also roughly four times slower than the runnerup and 13
times slower than third place. One pass is about error, not about latency. That leaves the practical question. Here's what you can do about any of this on Monday morning in order of how little
work it is. If your data already lives in BigQuery, you're one line of SQL away. The AI.cast command is generally available. It runs times FM 2.5 today and you pay for it like any other query
at the standard analysis rate with a monthly free tier that covers small jobs outright. The same model is wired into Aloy DB so a forecast can happen next to your rows instead of after an export.
Register the endpoint. Call the forecast function. Done. If you want weights you own, pip install xfm gets you version 2.5 under Apache 2.0 200 million parameters, a 16,000 point context
window and no lawyer required. If your problem has coariates, and most real ones do, Kronos 2 is the pick. 12 million parameters, Apache 2.0 multivariat, and second on the same
boards FM3 tops. And if you want to see what all this is like before committing to anything, version 3 is one download away legally for evaluation. Point it at your own history and compare it to your
current forecast. That comparison is exactly what the license permits. Which brings me to what I actually think, and I'm not going to hedge it. Kronos 2 wins for 95% of people watching this. It's
nearly three times smaller. four times faster per series, Apache 2.0, and it loses to Times FM 3 by 1.4 skill points on the benchmark Google chose to lead with. 1.4 points is real. It is not
worth a license that forbids you from using the result to make a decision. The 5% it doesn't win for research groups, benchmark authors, anyone whose job is measuring the frontier rather than
shipping on it for them. Times FM three is the new reference point and the license doesn't bite. And I'd switch back the day Google puts the weights back under an open license or the day it
lands in Big Query because on the accuracy question, Google is right. It wins 69% of head-to-heads against the best rival anyone has shipped and 100 out of 100 against the baseline. That's
earned. The thing I'd argue against isn't Google. It's the sentence beats everyone. That sentence describes a rank on one board under one filter. And the same week it was written, nine Agentic
systems were sitting above this model on a different board with 127 entries in it. Both facts are true. The best pre-trained forecaster on Earth is now Google's. And the best forecasting
system on the broadest public board is a language model with a strategy, not a forecaster with weights, which is the thing worth chewing on. For 2 years, the field's answer to how do we forecast
better was train a bigger model on more series. That answer is now producing a third of an improvement over same as last Tuesday and flattening. The other answer, the one the leaderboard is
filling up with, is to stop building the forecaster and start building the thing that chooses the forecaster. So, the open question isn't whether times FM3 is good. It measurably is. It's whether
Google just shipped the best version of an idea that's already being replaced. And what does it say? That the team replacing it works down the