Forecast Error Is Data: How a Startup Learns From Its Own Misses

A startup revises its forecast every month and almost never checks how the last one did. The forecasts that get better are the ones with a track record — and a track record takes one saved copy and ten minutes a month to keep.

Ask a founder how accurate their forecast is and you will usually get a feeling rather than a number. "Pretty close on revenue, we always overspend on hiring." That feeling is not worthless, but it is the only instrument the company has, and it is built from memory. The reason is mechanical. Every month the model is updated with the new actuals, the assumptions are nudged, and the previous version is overwritten. The forecast is rebuilt constantly and graded never, so the company has no way to learn whether it is optimistic, which lines it gets wrong, or whether its misses are random.

In an earlier post I argued that runway is a model with ranges, not a single number. A range is only honest if it is sized from evidence, and the evidence is the model's own past errors. This post is about collecting that evidence. It is the cheapest improvement in startup finance, because it needs no new tool and no new data — only the discipline of keeping the old forecast.

Freeze the forecast before the month starts

The mechanism is simple. At the start of each month, save a dated, read-only copy of the forecast: the revenue, the cost lines, the closing cash. Do not edit it afterwards, however wrong it turns out to be, because the whole point is to preserve what you believed at the time. When the month closes, put the actuals next to the frozen copy and record the difference for each line.

Do this for more than one horizon. The error on next month's number tells you how well you understand the business today. The error on the number three months out tells you how well you understand the plan — the hires that slip, the deals that close late, the costs that step up. The two are different skills, and a company can be good at the first and poor at the second without knowing it. After two or three quarters of saved copies you have something the company has never had before: a record of how its own judgement performs.

Bias and noise are different problems

The first thing the record shows is whether the errors lean. Add up the signed misses — over and under — across several months. If revenue is over-forecast in seven months of eight, that is bias: a systematic assumption is wrong, and it can be fixed. Maybe sales cycles are longer than the plan says, or the pipeline is converting later than it used to. Bias is good news in a sense, because it points at a specific belief to change.

If the misses are about the same size but scatter in both directions, that is noise: the business is genuinely uncertain on that line, and no tweak to the assumptions will make it go away. Noise cannot be corrected, but it can be sized. The typical size of the miss is exactly the width the scenario range should have. A founder who knows that monthly collections routinely land within a certain band of forecast can draw a runway range with a straight face, instead of a hopeful one.

It matters to separate the two because they call for opposite responses. Bias calls for changing the forecast. Noise calls for changing the buffer. Treating noise as bias — adjusting the assumption after every miss — produces a model that chases last month and is no more accurate for it.

Find the two lines that cause most of the miss

Total error is a poor guide, because lines cancel each other out. A company can hit its cash forecast to within a few percent while revenue timing is wrong in one direction and a delayed hire is wrong in the other. Score every line separately and the pattern is almost always concentrated: two or three lines produce most of the absolute error, and the rest are close to rounding.

The usual suspects are the same across startups. Revenue timing — not the amount, but the month it is recognised or collected. Hiring start dates — the plan assumes a person on the first of the month, reality delivers them six weeks later (the hiring plan is the budget, so this one moves everything). Usage-based costs such as cloud and API spend, which scale with product behaviour that finance does not see. And annual or irregular payments that were missing from the model altogether. Once the two biggest lines are known, the forecasting effort can go where it pays, instead of being spread evenly across forty rows.

A forecast scorecard in four columns

For each line, each month: the frozen forecast, the actual, the signed difference, and the same difference for the three-month-ahead copy. Add one row at the bottom for closing cash. After a quarter, read down the signed column for lean (bias) and the size of the differences for spread (noise). If you can only do one thing, do closing cash and your two largest lines.

Where AI helps, and where it should not

This is a natural place to bring in a language model, and it is useful in two ways. It can draft the variance commentary — the plain-English paragraph on why a line missed — from the ledger entries and the notes behind the numbers, which is the part nobody enjoys writing. And it can scan the record for patterns a person would take months to notice, such as a line that has been over-forecast every month since a pricing change.

What it should not do is produce or grade the score itself. The scorecard is arithmetic on the closed books: a frozen number, an actual and a difference. That calculation should be deterministic and traceable to the ledger, for the same reason I gave in AI should read the ledger, not write it. A model that explains the miss is a helpful analyst. A model that decides how big the miss was is marking its own homework, and the first time its figure disagrees with the books, nobody will trust any of it.

What to do with the number

A track record earns its keep when it changes a decision. Three places it usually does. First, the runway range: size the downside case from the observed error rather than from a round-number guess, so that "we have fourteen months, probably" becomes a statement with a width. Second, the investor update: telling investors how your forecasts have actually performed, and what you changed because of it, is one of the more credible things a small team can say, and it costs nothing to say once you have the record. Third, the plan itself: if hiring dates are late every time, move them; if collections run slower than terms, plan for the slower number.

None of this requires an elaborate system. It requires keeping the old forecast, comparing it to what happened, and being willing to read the result without defending it. A forecast that is occasionally wrong is not a failure of finance — that is what forecasts are. A forecast that is wrong in the same direction for a year, and nobody noticed because nobody kept the old copy, is.