Model Calibration · v3.0 Walk-Forward 3243 OOS fights
Expected Calibration Error
0.158
Average gap between predicted and actual win rates. Lower = better. Miscalibrated
Brier Score
0.1888
0.25 = coin flip
Log Loss
0.5599
0.693 = coin flip
Accuracy
76.8%
walk-forward
Buckets
7
confidence tiers
Reliability by Confidence Bucket Predicted Actual
50-55%
P 52.5%
A 67.9%
+15.4pp
55-60%
P 57.1%
A 90.0%
+32.9pp
60-65%
P 62.2%
A 90.7%
+28.5pp
65-70%
P 66.5%
A 82.1%
+15.7pp
70-75%
P 74.1%
A 89.0%
+15.0pp
75-80%
P 76.3%
A 92.0%
+15.6pp
80-85%
P 84.2%
A 97.8%
+13.6pp
Calibration reliability by confidence bucket
Bucket Fights Predicted average Actual win rate Delta percentage points
50-55% 2053 52.5% 67.9% +15.4
55-60% 50 57.1% 90.0% +32.9
60-65% 54 62.2% 90.7% +28.5
65-70% 56 66.5% 82.1% +15.7
70-75% 164 74.1% 89.0% +15.0
75-80% 634 76.3% 92.0% +15.6
80-85% 232 84.2% 97.8% +13.6
Model is under-confident — its picks win more often than advertised. Real edge is larger than displayed confidence.
Predicted vs Actual Diagonal = calibrated
50-55%: predicted 52.5%, actual 67.9%, n=2053 55-60%: predicted 57.1%, actual 90.0%, n=50 60-65%: predicted 62.2%, actual 90.7%, n=54 65-70%: predicted 66.5%, actual 82.1%, n=56 70-75%: predicted 74.1%, actual 89.0%, n=164 75-80%: predicted 76.3%, actual 92.0%, n=634 80-85%: predicted 84.2%, actual 97.8%, n=232
Exact numbers table
Predicted Range Fights Avg Predicted Actual Win Rate Delta Status
50-55% 2053 52.5% 67.9% +15.4% Under-confident (wins MORE than predicted — extra value)
55-60% 50 57.1% 90.0% +32.9% Under-confident (wins MORE than predicted — extra value)
60-65% 54 62.2% 90.7% +28.5% Under-confident (wins MORE than predicted — extra value)
65-70% 56 66.5% 82.1% +15.7% Under-confident (wins MORE than predicted — extra value)
70-75% 164 74.1% 89.0% +15.0% Under-confident (wins MORE than predicted — extra value)
75-80% 634 76.3% 92.0% +15.6% Under-confident (wins MORE than predicted — extra value)
80-85% 232 84.2% 97.8% +13.6% Under-confident (wins MORE than predicted — extra value)
Walk-forward = compact per-fight validation rows (3243 fights). True out-of-sample, zero lookahead.
Method-call reliability
Actual Method Sample Method Calls Correct
DEC 1607 48.4%
KO/TKO 1054 60.5%
SUB 582 23.7%
When a fight actually ends by submission, our full read (winner + method) matched it only ~24% of the time; by KO ~61%, by decision ~48% (n=3,243). Submissions are our hardest outcome to project - treat the finish-method line as directional. Method rows are descriptive reliability checks and do not alter the win-probability model.
Empirical Hit Rates by Public Zone 208 scored
Zone Scored Record Expected Hit Rate 95% Wilson CI
LOCK 30 21/30 79.2% 70.0% 52.1%-83.3%
STRONG 5 4/5 65.6% 80.0% 37.6%-96.4%
SOLID 14 11/14 61.9% 78.6% 52.4%-92.4%
LEAN 18 13/18 61.8% 72.2% 49.1%-87.5%
COIN FLIP 141 84/141 52.3% 59.6% 51.3%-67.3%
These rows are read-only audit data from resolved prediction-log entries. Empty zones stay visible so missing sample areas are obvious.
How to Read These Metrics
Brier Score — measures probability calibration. 0.25 = coin flip, lower = better. Excellent
Log Loss — information-theoretic quality. 0.693 = coin flip, lower = better. Excellent
ECE (Expected Calibration Error) — avg gap between predicted and actual rates. Lower = better calibrated. Miscalibrated
Live Tracking Data (208 predictions scored — click to collapse)
208
Scored
64%
Accuracy
0.234
Brier
Calibration by Confidence Bucket
Predicted Range Count Avg Predicted Actual Win Rate Delta Status
50%-55% 134 52.2% 59.7% +7.5% Under-confident
55%-60% 15 57.4% 53.3% -4.1% Well calibrated
60%-65% 21 61.9% 81.0% +19.0% Under-confident
65%-70% 7 65.4% 85.7% +20.3% Under-confident
70%-75% 3 72.1% 66.7% -5.4% Over-confident
75%-80% 13 75.6% 69.2% -6.4% Over-confident
80%-85% 6 81.8% 66.7% -15.1% Over-confident
85%-90% 9 85.0% 77.8% -7.2% Over-confident
Rolling Accuracy (10-fight window)
Fight 201
50%
Fight 202
50%
Fight 203
50%
Fight 204
50%
Fight 205
50%
Fight 206
60%
Fight 207
60%
Fight 208
60%