Models
How good are the house prediction models, reported the way we'd want a rival to report theirs
Three models sit behind the predictions on this site. elo_v1 is a pure Elo rating system -- transparent, well-understood math that updates a team's rating after every game based on the result and the opponent's strength. elo_epa_blend_v1 keeps that same Elo-based win probability but blends the projected margin 0.6 parts Elo to 0.4 parts a ridge-regression-adjusted EPA rating, an attempt to fold in-season efficiency into a system that otherwise only sees final scores. fitted_v1 is the current house model and the one selected by default: a ridge regression fit over a twenty-feature vector -- Elo, adjusted EPA on both sides, success rate, explosiveness, tempo, returning production, recruiting, preseason SP+ -- with its own win probability rather than a borrowed Elo one. It is also the only one of the three that projects a full season.
Default does not mean best at everything, and the table below is the reason we say so. On 2025, fitted_v1 predicted the final margin more accurately than either Elo model -- but it picked against the spread worse than both of them, and its win-probability calibration was a wash. Predicting a score and beating a betting line are different problems, and a model can get measurably better at the first while getting no better, or slightly worse, at the second. Which number matters depends on the question you came with.
Every number below comes from a walk-forward backtest: for each game, the model only ever sees data that would have been available before kickoff, then is scored against what actually happened. No season retroactively benefits from information the model didn't have at the time -- that is the entire point of testing this way, and it is a stricter bar than a model that is simply fit to the whole season at once and graded on the same data.
We are publishing this page because a prediction model that only shows you when it's right isn't worth trusting. Straight-up win prediction is meaningfully better than a coin flip; against-the-spread performance -- the harder test, since the market has already priced in most of what a public model knows -- lands close to a coin flip most seasons. That is not a bug in the presentation. It is the honest result, and it is also roughly what the published literature on public point-spread models would predict.
Backtest Accuracy
One row per model per season, at the base edge threshold (no minimum-conviction filter). Lower is better for MAE, RMSE, and both Brier columns. CFBD Brier is CFBD's own published win-probability model, shown as an external benchmark, not a house metric.
| Model | Season | Games | MAE | RMSE | ATS W-L-P | ATS Hit Rate | Brier | CFBD Brier |
|---|---|---|---|---|---|---|---|---|
| Elo (v1) | 2025 | 3829 | 15.9 pts | 20.2 pts | 792-780-25 | 50.4% | 0.182 | 0.157 |
| Elo (v1) | 2024 | 3799 | 15.6 pts | 20.0 pts | 791-729-31 | 52.0% | 0.198 | 0.177 |
| Elo (v1) | 2023 | 3724 | 16.0 pts | 20.5 pts | 704-682-25 | 50.8% | 0.184 | 0.161 |
| Elo (v1) | 2022 | 3705 | 16.7 pts | 21.4 pts | 699-741-19 | 48.5% | 0.186 | 0.154 |
| Elo (v1) | 2021 | 2454 | 16.3 pts | 20.8 pts | 426-447-13 | 48.8% | 0.185 | 0.164 |
| Elo (v1) | 2020 | 1125 | 15.2 pts | 19.4 pts | 278-280-9 | 49.8% | 0.189 | 0.169 |
| Elo (v1) | 2019 | 1623 | 14.9 pts | 19.0 pts | 0-0-0 | — | 0.173 | 0.147 |
| Elo (v1) | 2018 | 1556 | 15.7 pts | 19.7 pts | 0-0-0 | — | 0.179 | 0.156 |
| Elo (v1) | 2017 | 1551 | 14.9 pts | 19.1 pts | 0-0-0 | — | 0.186 | 0.151 |
| Elo (v1) | 2016 | 1549 | 14.7 pts | 19.0 pts | 0-0-0 | — | 0.186 | 0.168 |
| Elo (v1) | 2015 | 1538 | 15.0 pts | 19.2 pts | 0-0-0 | — | 0.171 | 0.147 |
| Elo + EPA blend (v1) | 2025 | 3829 | 15.7 pts | 20.0 pts | 780-792-25 | 49.6% | 0.182 | 0.157 |
| Elo + EPA blend (v1) | 2024 | 3799 | 15.5 pts | 19.8 pts | 778-742-31 | 51.2% | 0.198 | 0.177 |
| Elo + EPA blend (v1) | 2023 | 3724 | 15.8 pts | 20.2 pts | 693-693-25 | 50.0% | 0.184 | 0.161 |
| Elo + EPA blend (v1) | 2022 | 3705 | 16.6 pts | 21.3 pts | 709-731-19 | 49.2% | 0.186 | 0.154 |
| Elo + EPA blend (v1) | 2021 | 2454 | 16.2 pts | 20.7 pts | 413-460-13 | 47.3% | 0.185 | 0.164 |
| Elo + EPA blend (v1) | 2020 | 1125 | 15.2 pts | 19.3 pts | 280-278-9 | 50.2% | 0.189 | 0.169 |
| Elo + EPA blend (v1) | 2019 | 1623 | 14.6 pts | 18.5 pts | 0-0-0 | — | 0.173 | 0.147 |
| Elo + EPA blend (v1) | 2018 | 1556 | 15.4 pts | 19.5 pts | 0-0-0 | — | 0.179 | 0.156 |
| Elo + EPA blend (v1) | 2017 | 1551 | 14.8 pts | 19.0 pts | 0-0-0 | — | 0.186 | 0.151 |
| Elo + EPA blend (v1) | 2016 | 1549 | 14.7 pts | 18.8 pts | 0-0-0 | — | 0.186 | 0.168 |
| Elo + EPA blend (v1) | 2015 | 1538 | 15.0 pts | 19.0 pts | 0-0-0 | — | 0.171 | 0.147 |
| Fitted ridge (v1) | 2025 | 3829 | 14.7 pts | 18.5 pts | 749-823-25 | 47.6% | 0.181 | 0.157 |
| Fitted ridge (v1) | 2024 | 3799 | 14.3 pts | 18.2 pts | 736-784-31 | 48.4% | 0.197 | 0.177 |
| Fitted ridge (v1) | 2023 | 3724 | 14.6 pts | 18.6 pts | 688-698-25 | 49.6% | 0.179 | 0.161 |
| Fitted ridge (v1) | 2022 | 3705 | 15.8 pts | 20.2 pts | 708-732-19 | 49.2% | 0.183 | 0.154 |
| Fitted ridge (v1) | 2021 | 2454 | 15.7 pts | 20.0 pts | 429-444-13 | 49.1% | 0.184 | 0.164 |
| Fitted ridge (v1) | 2020 | 1125 | 15.1 pts | 19.2 pts | 290-268-9 | 52.0% | 0.183 | 0.169 |
| Fitted ridge (v1) | 2019 | 1623 | 14.1 pts | 17.8 pts | 0-0-0 | — | 0.165 | 0.147 |
| Fitted ridge (v1) | 2018 | 1556 | 14.9 pts | 18.7 pts | 0-0-0 | — | 0.175 | 0.156 |
Higher-conviction splits -- the same backtest filtered to games where the model's edge over the market exceeds a minimum threshold -- also exist in the warehouse but are omitted here for readability. Those thresholds are what power the Edge Board's scored slate at game time.
Accuracy Over Time
Against-the-spread hit rate by season, one line per model, against a coin-flip reference.