The Bitter Lesson in Demand Forecasting

What a new study on Alibaba data says about the value of forecasting models and why the answer has less to do with the model than you’d expect.

Alibaba spent years building its own demand forecast. A purpose-built deep learning system, refined by experienced teams, running in production across 70,000 to 100,000 items. In a new study by Alibaba’s forecasting team together with researchers from Cornell University and the University of Chicago, that system was put up against three openly available pretrained forecasting models, models that had never seen Alibaba’s data.

One of them won. Over a full fiscal year, 3.5 points better forecast accuracy on domestic products and 5.4 points on cross-border ones. In every single month. In every single product category.

If you own demand planning, that number is worth pausing on. But not for the reason it first suggests.

The number that matters is 1

That’s how many additional inputs the winning setup used. One. A forward-looking signal describing how strongly an upcoming promotion is expected to lift demand. Not 45 carefully engineered features.

The authors’ ablation studies are blunt on this point:

  • Replace the intensity signal with a plain promotion flag “there’s a campaign on this date” and the model loses to the incumbent system.
  • On a public retail dataset, the model with one well-chosen input reached an R² of 0.667. The best method documented in the accompanying textbook, using all 45 features, reached 0.585. Add all 45 features on top of the single input, and accuracy drops to 0.619.
  • Two of the three pretrained models tested lost to the incumbent anyway. Size alone explains nothing.

And the winning signal didn’t come from some new technique. It came from a classical statistical decomposition model that has been around for years. The authors call it distillation: the established model isn’t replaced, it supplies the input that lets the pretrained model do its work. In the study’s strongest configuration, the two together reached 0.718.

The benchmarks change faster than your forecasts.

By late August, the ranking was already out of date

The study is dated July 2026. On 31 August, Google released TimesFM-3, the successor to precisely the model that performed worst at Alibaba. The study had named three reasons for that: the architecture, the lack of native support for accompanying variables, and error accumulation across the long 93-day horizon.

TimesFM-3 addresses all three. It processes multiple series jointly, takes known future events such as planned promotions or weather forecasts directly as input, and produces the entire forecast horizon in a single pass rather than step by step. Google’s own illustration is an ice cream sales forecast that factors in the planned promotion calendar and anticipates the sales lift on each campaign day.

It is the same mechanism the Alibaba study describes. And it is the second lap of the same race inside two months.

Which is exactly the point. The model layer turns over quarterly. Build your planning architecture around one specific model today, and you’ve built around something that won’t be the best choice in six months. What does not turn over quarterly is the question of which knowledge from your planning process makes it into the forecast at all and what happens to the result afterwards.

How does this effect planning teams?

The decisive knowledge rarely sits in the sales history. It sits with your planners.

The grid operator knows when the substation changeover is scheduled. The food manufacturer knows the retailer’s promotion calendar three months out, secondary placements included. The logistics provider knows which major account reliably pulls volume forward before year-end, and that the maintenance window in week 34 costs two days of capacity. The industrial supplier knows which tender gets decided in autumn.

In most planning processes, that knowledge ends up as a manual override on the system forecast, in a spreadsheet column, or nowhere at all. The study shows how much leverage there is in supplying it in a form a model can actually compute with — and that a raw yes/no flag doesn’t cut it. You need the intensity, not just the date.

A second finding matters just as much in production. Learning jointly across many items only helped when those items belonged together. Across 44 highly heterogeneous items, accuracy collapsed from 0.667 to 0.033. Alibaba deliberately switched this capability off in production because a clean grouping wasn’t available. Pool a heterogeneous assortment without grouping it, and you make the forecast worse, not better.

What it means for us at prognotix

The authors say it themselves: better forecasts do not automatically become better inventory decisions. Three points of accuracy are not a P&L result. Whether they turn into less working capital, fewer stockouts or fewer expedited shipments is decided in the replenishment logic downstream: In safety stocks, reorder points, lot sizes and capacity limits.

That translation is a discipline of its own. It is also why raw model accuracy carries less and less competitive weight.

Therefore we don’t build our own forecasting models to compete on a benchmark leaderboard. We work where the study locates the leverage.

Faster to a first forecast you can rely on. Our Autopilot produces the initial forecast fully automatically, so getting started doesn’t depend on a months-long modelling project. It’s exactly the barrier pretrained models lower in principle, made usable for companies that don’t want to assign a data science team to it.

From forecast to decision. Our optimization modules for inventory and warehouse pick up where the forecast ends. That’s where the measurable contribution is created.

Your planning knowledge as an input. Promotion calendars, maintenance windows, tender dates, customer behaviour, captured in a structured way instead of corrected after the fact.

One question we now ask in every conversation with a planning team:

Finding the decisive knowledge is key.

How do known future events reach your forecast today?

Manually, as a structured input, or not at all. The answer usually says more about the improvement potential than any model comparison does.

Sources: Sui, Xiao, Xin, Yang, Huang, Cao (2026), “A Bitter Lesson for Retail Demand Forecasting: Evidence from Fine-Tuning Foundation Models”, SSRN, July 2026. Google Research, “TimesFM-3: A zero-shot foundation model for multivariate forecasting”, 31 August 2026.