Explainer
How live win probability models actually work
Every live betting screen shows a win probability. Very few explain where the number comes from, or where it should not be trusted.
The question a live model is answering
A live win probability is one specific claim: given everything true about this game right now — the score, the time left, who has the ball or who is at bat, and the relative strength of the two teams — what share of games that looked exactly like this ended with the home side winning? It is a frequency statement about a population of similar game states, not a prediction about this particular game.
That framing matters, because it tells you what would count as being wrong. A model that says 80 percent and loses is not wrong. A model that says 80 percent across a thousand games and wins 640 of them is badly wrong. You can only judge one of these things, and it takes a season of records to do it.
What actually goes into the number
Game state
The dominant input, by a distance. In football that means score margin, time remaining, down, distance, field position and possession. In baseball it is inning, outs, the base configuration and the score. These are not soft factors — a first-and-goal at the two is a different world from a first-and-ten at your own twelve, and the base-out state alone explains most of the run-scoring variance in an inning.
Time and possessions remaining
Time is the mechanism that converts a lead into a win. Ten points with twelve minutes left is a different bet from ten points with two minutes left, not because the lead changed but because the number of remaining chances collapsed. Good football models count possessions, not seconds. Baseball models count outs. Anything that measures only the clock will misprice the endgame.
Team strength
A pre-game rating — a power number, a market-implied spread — anchors the model before the game starts and keeps mattering as it goes, though its weight decays. Early in a game, the prior carries most of the load. By the fourth quarter, the scoreboard has swamped it. A model that keeps too much prior late will be stubborn about upsets; one that keeps too little will overreact to a lucky score.
Situational modifiers
Pitcher fatigue and bullpen quality in baseball. Timeouts remaining and two-minute-drill ability in football. These are second-order but not trivial; a tiring starter with nobody warming is worth real probability.
Why two models disagree on the same game
Three reasons, almost always. First, different priors: one model rates the home team a point better than the other, and that gap persists through the whole game. Second, different decay schedules for how quickly the prior gives way to the score. Third, different training data — a model built on college football will misprice the NFL, because fourth-quarter comebacks behave differently when both offences are competent.
A five-point disagreement between reasonable models is normal. If you see thirty, at least one of them is broken or is answering a different question — for example, pricing a spread cover rather than an outright win.
Where every live model is least trustworthy
- Rare states. Onside kicks, extra innings with the automatic runner, fake punts. There is not enough history to estimate these well, and models tend to fall back on smoothed guesses.
- Injuries. A model reading a data feed does not know the starting quarterback just limped off. For a few minutes the market is better informed than the model.
- Blowouts. At 97 percent, small errors in the tail are relatively enormous, and teams stop playing to win — they play to end the game. Model assumptions about behaviour break down here.
- Weather and conditions. A twenty-mile-an-hour crosswind changes both scoring and variance. Most feeds do not encode it.
- The last two minutes. Discrete, high-leverage, and enormously sensitive to timeouts and clock rules. Small state errors become large probability errors.
Reading a probability alongside a price
The number is only useful next to the market. Convert the two-way live price into a no-vig probability, then compare. If the model says 62 percent and the de-vigged market says 58, that four-point gap is your claimed edge — before you ask whether the model deserves that much confidence.
The honest version of this comparison includes a band. A model that outputs 62 percent with a plus-or-minus five band is saying the truth is somewhere between 57 and 67. If the market sits inside that band, there is no bet, however tempting the point estimate looks. That is the whole argument of the companion piece on when not to bet.
What a good live model feels like in use
It moves smoothly on ordinary plays and sharply on genuinely decisive ones. It does not lurch on a five-yard gain in the second quarter. It agrees with the market most of the time — if yours disagrees constantly, the likeliest explanation is a bug, not an edge. And it tells you when it does not know. A live number without an uncertainty band is a marketing asset, not a tool.