hermes football: market-grade win probabilities for NFL and college football
Football is the hardest market in American sports to out-predict. The Hermes football models were built to match it honestly: stacked ensembles over play-by-play efficiency, Elo ratings, and market consensus, evaluated only walk-forward, with every claim tested against the closing line.
What the football models do
hermes-nfl-1.1 and hermes-ncaaf-1.1 produce the win probabilities shown on NFL and college football matchups. Like the rest of the Hermes family, they are pregame winner models: they do not score props, totals, or live situations, and analysis copy stays in the scorecard layer.
Both models share one architecture. A stats-only ensemble learns from team strength and form with no market input at all. A market-aware ensemble adds sportsbook consensus. A residual model learns where outcomes systematically deviate from market prices. A calibrated meta-blend combines them, weighted by what actually worked out-of-sample.
The data foundation
The NFL side is built on 24 seasons of play-by-play data — over 6,200 games reduced to leakage-safe pregame features: offensive and defensive EPA per play split by pass and rush, success and explosive-play rates, early-down efficiency, turnover rates, a 538-style Elo rating with margin-of-victory updates and preseason regression, announced-starter quarterback continuity, rest, travel, and weather.
The college side covers 20 seasons and more than 33,000 games, with division-aware Elo initialization (an FCS program does not start with an FBS rating), rolling scoring form, conference context, and per-book betting lines reconstructed from 2006 onward. Every feature is computed strictly from information available before kickoff, and the feature builders emit audit files that verify it.
Market history came from two places: free public archives for closing lines (a much longer record than most published studies use), and a dedicated historical snapshot backfill from The Odds API for verified opening and pre-kick boards from 2020 through 2025 — the dataset behind the line-movement study below.
- NFL: EPA/play efficiency splits, QB continuity, Elo, rest, weather — 60 pregame features
- NCAAF: division-aware Elo, scoring form, conference context — 23 pregame features
- 10 market features per league: spread, no-vig moneyline probability, cross-book dispersion, open-to-close movement
- Walk-forward evaluation only: train through season N−1, test season N, repeated across 5–7 seasons
Honest baselines, honest results
Most published football prediction accuracy numbers do not survive contact with two questions: was the evaluation walk-forward, and was it compared to the betting market? Papers claiming 85–95% accuracy invariably leak in-game information. The real bar is the closing line, which picks NFL winners at roughly 66–67% and college winners at roughly 74–76%.
Against that bar, on pooled 2019–2025 walk-forward tests, the NFL stack picks winners at 66.8–67.0% against the true closing line — beating or matching it in six of seven seasons — and at 68.7% versus the verified pre-kick board on 2024–25 holdout games, where it also wins on AUC and Brier score. The college model ties the closing market on calibration (Brier 0.1702 vs 0.1700) while covering the thousands of games each season that have no betting line at all.
| Tier (NFL, walk-forward 2019–2025) | AUC | Accuracy | Brier |
|---|---|---|---|
| Closing-line baseline | 0.7269 | 66.3% | 0.2105 |
| Elo-only baseline | 0.6928 | 63.7% | 0.2233 |
| Stats-only ensemble (no odds input) | 0.7042 | 63.6% | 0.2190 |
| Residual model | 0.7209 | 66.8% | 0.2121 |
| Meta-blend (2021–25) | 0.7177 | 67.0% | 0.2131 |
The line-movement discovery
The most important result did not come from beating the close. It came from the open. The stats-only NFL model — which never sees a betting line — predicts how lines move between the opening board and kickoff with a correlation of 0.68. Placebo controls (a constant signal, and a shuffled model) score near zero, so the effect is not a statistical artifact. Adding the model to the opening line lifts closing-line prediction R² from 0.72 to 0.85.
When the model disagreed with the opening line by at least two points of probability, the line subsequently moved toward the model 70% of the time, capturing an average of 5.2 probability points of closing-line value. Consistently obtaining better-than-closing prices is the standard professional definition of real predictive edge — and it is a much more honest claim than pretending any model beats the closing line outright.
What this means on WhoWins.ai
Every upcoming NFL and college football matchup now carries a Hermes probability, generated the moment odds are available and refreshed as they change. The model components stay inspectable: each prediction records the stats-only view, the market view, and the blended result, so users can see when the model agrees with the market and when it is leaning against it.
College football coverage includes the full FBS slate — including the games too small for deep market attention, where a strength-based model has the most to add.
Keep reading
