Model calibration audit

For every bucket of predicted probability, what did the model actually hit? A well-calibrated model has actual win rate ≈ predicted prob. A negative gap means the model is overconfident in that bucket (dangerous); positive gap means it's underpredicting (safe). Rows turn red when |gap| > 5pp AND N ≥ 10.

Calibration cohort by sport (click to filter):
All sports MLB N=3,254NCAAF N=960TENNIS_WTA N=668TENNIS_ATP N=567MMA_MIXED_MARTIAL_ARTS N=522LALIGA N=219AMERICANFOOTBALL_NFL N=187SERIEA N=171EPL N=169NHL N=162LIGUE1 N=147NFL_PRESEASON N=145BUNDESLIGA N=128UCL N=71CRICKET_IPL N=27NCAAB N=11NBA N=9

Combined "all sports" is rarely meaningful — sports differ in market efficiency, signal availability, and base rates. Use the chips to drill into a single sport. N<50 (red) means the calibration is brittle; N≥200 (green) is trustworthy.

Filter: window=90d · sport=tennis_wta

Overall: N = 668 · mean predicted 67.8% · actual win rate 68.0% · gap +0.2pp · Brier 0.202 · log-loss 0.585
Calibration by predicted-probability bucket.
Predicted-prob bucket N Mean predicted Actual win rate Gap (actual − predicted) Brier
50-55% 113 52.5% 57.5% +5.1pp 0.244
55-60% 142 56.8% 56.3% -0.5pp 0.245
60-65% 53 62.0% 50.9% -11.1pp 0.262
65-70% 65 67.5% 70.8% +3.3pp 0.208
70-75% 29 73.0% 75.9% +2.9pp 0.181
75-80% 168 75.8% 73.2% -2.6pp 0.198
80-90% 71 88.0% 91.5% +3.5pp 0.079
90%+ 27 93.3% 96.3% +3.0pp 0.035