How we measure

Methodology

Last updated August 17, 2026

Why this page exists. We publish hit rates, so we owe you the method behind them. Everything below describes how the numbers were produced and where they break down. You can reproduce the production models’ results on the board itself — that is the point.

1. The data

SourceMLB Stats API (public)
WindowOpening Day 2026 through August 13, 2026
Sample1,819 completed regular-season games across 139 slates
ExcludedSpring training, the All-Star exhibition, postponed and suspended games, and any game without a completed first inning
League NRFI rate49.6% over the same window

Every input is public. We do not use paid feeds, private data, or proprietary tracking. The advantage we claim is in how the public data is combined, not in having data nobody else can get.

The window above is the published study. Since September 2026 the same replay also runs every night on the board’s own server, extending the sample through the last completed slate, and the home page reads that tally live (/api/record, public). The two models a pass unlocks are scored there by the definitions in section 3. The table in section 4 stays fixed at the August 13 study so that the five-model comparison, including the three models that lost, is reproducible as printed.

2. How the backtest is constructed

This is the part worth checking us on, because it is where most published “hit rates” quietly fall apart.

The replay is point-in-time. For each historical game, every model is re-run using only information that existed before that game started: the starting pitchers as announced, each pitcher’s prior starts, each team’s prior completed games, and the park and weather context for that date. No result from the game being scored, and no game played after it, is visible to the model producing its number.

That constraint is what makes the figures meaningful. A model allowed to see the season it is being scored against will produce impressive numbers that mean nothing.

3. What each published number means

These are not interchangeable, and the difference between them is large. When we quote a figure anywhere on this site, it is one of these.

MeasureDefinition
Top pick Each slate, the single game with the model’s highest NRFI probability. Scored right if no run was scored in the first inning. One result per slate.
Top 3 The three highest-probability games per slate, scored the same way. Three results per slate.
Every game Every game the model could price, scored against the model’s own call — NRFI or YRFI — not just its favourites.

Games the model cannot price for lack of data are excluded rather than counted as wrong, which is why the “every game” denominators below are smaller than 1,819.

4. Results

All five models, same window, same rules. Nothing omitted, including the models that lost. The board today carries the two winners — The Full Count and The Heater; the other three stay published here because results you cannot see do not count as transparency.

ModelTop pickTop 3Every gameBrier
The Full Count85/139 (61.2%)230/413 (55.7%)962/1741 (55.3%)0.2737
The Heater80/139 (57.6%)217/413 (52.5%)948/1817 (52.2%)0.2977
Basic76/139 (54.7%)221/413 (53.5%)953/1817 (52.4%)0.3014
Poisson76/139 (54.7%)206/413 (49.9%)910/1741 (52.3%)0.2735
Log575/139 (54.0%)206/413 (49.9%)903/1741 (51.9%)0.3762

League base rate over the same window: 49.6%. That is the number every column has to beat to mean anything.

Why calibration is in the table

Brier score measures whether a stated probability is honest — lower is better. A model can rank games well while its numbers mean nothing. Log5 ranks respectably and has by far the worst calibration in the test, because it prints confident-looking figures that do not correspond to real frequencies. Poisson is the reverse: best calibrated, and it picks at roughly the base rate.

The Full Count is the only model that does both jobs. Its calibration is within 0.0002 of Poisson’s while ranking six points better on top picks. That combination, not the headline percentage, is why it loads by default.

5. What these numbers do not mean

A hit rate is not a return. Sportsbooks and prediction markets price NRFI with a margin, and you have to clear it before a winning percentage becomes a winning position.

PriceBreakevenvs 55.3% board-wide
−11052.4%clears by 2.9 pts
−11553.5%clears by 1.8 pts
−12054.5%clears by 0.7 pts
−13056.5%short by 1.3 pts
−14058.3%short by 3.1 pts

Betting the whole board is not a strategy. The edge is concentrated in the top few games of each slate and thins out quickly below them. Anyone quoting our 61.2% as though it applies to every game we publish is misreading it, and so is anyone selling you that reading.

The 61.2% figure also assumes a discipline most people do not have: exactly one position per slate, on the model’s top-rated game, every day of a season, through the losing stretches. With 139 slates and a true rate near 61%, runs of four and five straight misses are expected, not evidence that anything broke.

6. Known limitations

7. Live results are not backtest results

Everything above is historical replay. The board also grades itself in public, day by day: each finished game shows whether the NRFI actually held, each model’s call is marked right or wrong, and the day’s record is tallied on screen. That running record is the honest test, and it is visible whether it flatters us or not.

8. Check it yourself

You do not have to take any of this on trust. Open the board, choose any past date, and each model’s calls for that day render beside what actually happened in the first inning, with the day’s hit rate tallied. Step back through the season and you are re-running the same study by hand.

If your count disagrees with ours, we want to hear it: [email protected].

9. What we publish, and what we don’t

We publish the measurement in full: the source, the window, the exclusions, the definition behind every number, the results for models that lost as well as the one that won, and the limitations above. That is what substantiates a performance claim, and none of it depends on trusting us.

We describe what each model considers — a pitcher’s first-inning history against his recent form, the quality of the offence he faces at the top of the order, handedness, park, and weather — but we do not publish how those pieces are weighted against each other. The weighting is the product.

We think that is the right line. How we measured is something you should be able to audit and argue with. What we measured with is what a pass buys.