Transparency · Reliability
Calibration Metrics
How well model probabilities match real-world outcomes — verified bucket-by-bucket.
Model Calibration · v3.0 Walk-Forward
3243 OOS fights
Expected Calibration Error
0.158
Average gap between predicted and actual win rates.
Lower = better.
Miscalibrated
Brier Score
0.1888
0.25 = coin flip
Log Loss
0.5599
0.693 = coin flip
Accuracy
76.8%
walk-forward
Buckets
7
confidence tiers
Reliability by Confidence Bucket
Predicted
Actual
50-55%
+15.4pp
55-60%
+32.9pp
60-65%
+28.5pp
65-70%
+15.7pp
70-75%
+15.0pp
75-80%
+15.6pp
80-85%
+13.6pp
| Bucket | Fights | Predicted average | Actual win rate | Delta percentage points |
|---|---|---|---|---|
| 50-55% | 2053 | 52.5% | 67.9% | +15.4 |
| 55-60% | 50 | 57.1% | 90.0% | +32.9 |
| 60-65% | 54 | 62.2% | 90.7% | +28.5 |
| 65-70% | 56 | 66.5% | 82.1% | +15.7 |
| 70-75% | 164 | 74.1% | 89.0% | +15.0 |
| 75-80% | 634 | 76.3% | 92.0% | +15.6 |
| 80-85% | 232 | 84.2% | 97.8% | +13.6 |
Model is under-confident — its picks win more often than advertised. Real edge is larger than displayed confidence.
Predicted vs Actual
Diagonal = calibrated
▸ Exact numbers table
| Predicted Range | Fights | Avg Predicted | Actual Win Rate | Delta | Status |
|---|---|---|---|---|---|
| 50-55% | 2053 | 52.5% | 67.9% | +15.4% | Under-confident (wins MORE than predicted — extra value) |
| 55-60% | 50 | 57.1% | 90.0% | +32.9% | Under-confident (wins MORE than predicted — extra value) |
| 60-65% | 54 | 62.2% | 90.7% | +28.5% | Under-confident (wins MORE than predicted — extra value) |
| 65-70% | 56 | 66.5% | 82.1% | +15.7% | Under-confident (wins MORE than predicted — extra value) |
| 70-75% | 164 | 74.1% | 89.0% | +15.0% | Under-confident (wins MORE than predicted — extra value) |
| 75-80% | 634 | 76.3% | 92.0% | +15.6% | Under-confident (wins MORE than predicted — extra value) |
| 80-85% | 232 | 84.2% | 97.8% | +13.6% | Under-confident (wins MORE than predicted — extra value) |
Walk-forward = compact per-fight validation rows (3243 fights). True out-of-sample, zero lookahead.
Method-call reliability
| Actual Method | Sample | Method Calls Correct |
|---|---|---|
| DEC | 1607 | 48.4% |
| KO/TKO | 1054 | 60.5% |
| SUB | 582 | 23.7% |
When a fight actually ends by submission, our full read (winner + method) matched it only ~24% of the time; by KO ~61%, by decision ~48% (n=3,243). Submissions are our hardest outcome to project - treat the finish-method line as directional.
Method rows are descriptive reliability checks and do not alter the win-probability model.
Empirical Hit Rates by Public Zone
208 scored
| Zone | Scored | Record | Expected | Hit Rate | 95% Wilson CI |
|---|---|---|---|---|---|
| LOCK | 30 | 21/30 | 79.2% | 70.0% | 52.1%-83.3% |
| STRONG | 5 | 4/5 | 65.6% | 80.0% | 37.6%-96.4% |
| SOLID | 14 | 11/14 | 61.9% | 78.6% | 52.4%-92.4% |
| LEAN | 18 | 13/18 | 61.8% | 72.2% | 49.1%-87.5% |
| COIN FLIP | 141 | 84/141 | 52.3% | 59.6% | 51.3%-67.3% |
These rows are read-only audit data from resolved prediction-log entries. Empty zones stay visible so missing sample areas are obvious.
How to Read These Metrics
Brier Score — measures probability calibration. 0.25 = coin flip, lower = better.
Excellent
Log Loss — information-theoretic quality. 0.693 = coin flip, lower = better.
Excellent
ECE (Expected Calibration Error) — avg gap between predicted and actual rates. Lower = better calibrated.
Miscalibrated
Live Tracking Data (208 predictions scored — click to collapse)
208
Scored
64%
Accuracy
0.234
Brier
Calibration by Confidence Bucket
| Predicted Range | Count | Avg Predicted | Actual Win Rate | Delta | Status |
|---|---|---|---|---|---|
| 50%-55% | 134 | 52.2% | 59.7% | +7.5% | Under-confident |
| 55%-60% | 15 | 57.4% | 53.3% | -4.1% | Well calibrated |
| 60%-65% | 21 | 61.9% | 81.0% | +19.0% | Under-confident |
| 65%-70% | 7 | 65.4% | 85.7% | +20.3% | Under-confident |
| 70%-75% | 3 | 72.1% | 66.7% | -5.4% | Over-confident |
| 75%-80% | 13 | 75.6% | 69.2% | -6.4% | Over-confident |
| 80%-85% | 6 | 81.8% | 66.7% | -15.1% | Over-confident |
| 85%-90% | 9 | 85.0% | 77.8% | -7.2% | Over-confident |
Rolling Accuracy (10-fight window)
Fight 201
50%
Fight 202
50%
Fight 203
50%
Fight 204
50%
Fight 205
50%
Fight 206
60%
Fight 207
60%
Fight 208
60%