2026-08-17 · NRFI Edge

NRFI model accuracy: a 2,134-game backtest

The nightly replay has kept going since this post first ran. Through September 3, 2026, GET /api/record (pending 0) grades 2,134 regular-season games across 160 slates. League NRFI rate 1,059/2,134 (49.6%). Brier score is mean squared error between the stated probability and the 0/1 first-inning result. Lower is more honest.

ModelTop pickTop 3Every gameBrier
The Full Count98/160 (61.3%)275/476 (57.8%)1,125/2,040 (55.1%)0.2717
The Heater95/160 (59.4%)256/476 (53.8%)1,122/2,132 (52.6%)0.2977

Those are the two models a pass unlocks. The five-model table below is the published August 13 study (1,819 games, 139 slates). That is still the only window where Basic, Poisson, and Log5 were scored the same way. How to read the three measures, and how the five models are built, is in methodology.

Three days ago we published a backtest over 625 games and 46 slates. That was seven weeks of baseball, and it was not enough.

Forty-six slates is a small enough sample that a model can finish on top by getting lucky twice in August. So we re-ran the whole thing: every regular season game from opening day through August 13. 1,819 games across 139 slates. Spring training and the All-Star exhibition are filtered out. Each day, each model named its highest-probability NRFI play, and we graded it against what actually happened in the first inning.

The league-wide NRFI rate over that span was 49.6%. That is the number every model below has to beat to be worth anything.

The numbers

ModelTop pick per dayTop-3 poolEvery gameCalibration (Brier)
The Full Count85/139 (61.2%)230/413 (55.7%)962/1741 (55.3%)0.2737
The Heater80/139 (57.6%)217/413 (52.5%)948/1817 (52.2%)0.2977
Basic76/139 (54.7%)221/413 (53.5%)953/1817 (52.4%)0.3014
Poisson76/139 (54.7%)206/413 (49.9%)910/1741 (52.3%)0.2735
Log575/139 (54.0%)206/413 (49.9%)903/1741 (51.9%)0.3762

We were wrong three days ago

The seven-week study had The Heater on top. Over a full season it is clearly second. Our own headline model changed when we looked at four times as much data, which is precisely the thing we keep telling you to be suspicious of when somebody shows you a good month.

We are not going to quietly swap the numbers and hope nobody noticed. The earlier post was honest about its sample and the sample was too small. This one is bigger and it disagrees.

Why The Full Count is the default

Ranking and calibration are different jobs, and models are usually good at one or the other.

Log5 is the clean example of the split: it ranks respectably but its Brier score of 0.3762 is the worst in the test by a distance, because it prints confident numbers like 88% and 92% that do not mean what they appear to mean. Poisson is the mirror image: the best-calibrated model we have (0.2735), and dead last at actually picking a game, finishing at the base rate.

The Full Count is the only model that does both. Its 0.2737 Brier is within 0.0002 of Poisson’s, statistically the same calibration, while ranking six points better on top picks. When a model both sorts games correctly and reports probabilities you can take at face value, that is the one that should load by default.

The part most sites would leave out

Look at the “Every game” column again.

The Full Count hits 61.2% on its top pick but 55.3% across the entire board. That gap is the most important thing on this page. The edge is concentrated in a handful of games per slate, and it thins out fast as you work down the list.

Here is what that means in practice. NRFI markets are typically priced between -110 and -130. At -110 you need 52.4% to break even, at -120 you need 54.5%, and at -130 you need 56.5%. So a 55.3% board-wide hit rate clears the cheaper prices and loses money at -130 or worse.

Betting every game we publish is not a strategy. The top of the board is where the model has something to say.

It is not a straight line

The Full Count’s top pick by month: 73% in April, then 52%, 63%, 57%, and 62%.

April was not skill and May was not a broken model. That is a coin weighted to about 61% doing what a coin weighted to 61% does over thirty flips. Any month of that sequence, screenshotted on its own, would tell you a story that the other five months contradict.

Basic makes the same point from the other direction. It finishes third on top picks, genuinely respectable, while carrying the second-worst calibration in the test. It is a rate average with no opinion about anything, and it will still produce weeks that look like genius.

Check it yourself

Every number above came from replaying the season through the same code that builds tonight’s board, using only data that existed before each game started. No lookahead, no dropped cold streaks.

You do not have to take our word for any of it. Pick any past date on the board and the leans render next to what actually happened.

The league rate those models are scored against is how often the first inning is scoreless. Park and weather, the two context multipliers inside The Full Count, are in NRFI park factors and Does weather change the first inning.

Unlock the board

Statistics and research, not betting advice. 21+. Gambling problem? Call 1-800-GAMBLER.