Proof & records

Calibration Metrics

How well model probabilities match real-world outcomes — verified bucket by bucket.

Graded Public Record
58.9%
Graded Public Hit Rate
297
Graded Public Picks
Every bucket below is computed from graded public picks scored against real results. The walk-forward backtest is research and is labelled that way on the Accuracy page.
Calibration reliability status
Status Detail
Withdrawn Walk-forward backtest buckets are research and are not used as calibration proof.
Empirical Hit Rates by Public Zone 297 scored
Public zone hit rates
Zone Scored Record Expected Hit Rate 95% Wilson CI
LOCK 32 22/32 79.2% 68.8% 51.4%-82.0%
STRONG 7 6/7 65.4% 85.7% 48.7%-97.4%
SOLID 29 20/29 61.8% 69.0% 50.8%-82.7%
LEAN 30 15/30 61.1% 50.0% 33.2%-66.8%
COIN FLIP 199 112/199 52.2% 56.3% 49.3%-63.0%
These rows are read-only audit data from resolved prediction-log entries. Empty zones stay visible so missing sample areas are obvious.
How to Read These Metrics
Brier Score — measures probability calibration. 0.25 = coin flip, lower = better. Benchmark unavailable
Log Loss — information-theoretic quality. 0.693 = coin flip, lower = better. Benchmark unavailable
ECE (Expected Calibration Error) — avg gap between predicted and actual rates. Lower = better calibrated. Benchmark unavailable
Live Tracking Data 297 predictions scored
297
Scored
59%
Accuracy
0.242
Brier
Calibration by Confidence Bucket
Live calibration buckets
Predicted Range Count Avg Predicted Actual Win Rate Delta Status
50%-55% 190 52.0% 55.8% +3.7pp Well calibrated
55%-60% 22 57.5% 54.5% -3.0pp Well calibrated
60%-65% 44 61.8% 63.6% +1.8pp Well calibrated
65%-70% 8 65.4% 75.0% +9.6pp Under-confident
70%-75% 3 72.1% 66.7% -5.4pp Over-confident
75%-80% 15 75.8% 66.7% -9.2pp Over-confident
80%-85% 5 81.5% 60.0% -21.5pp Over-confident
85%-90% 10 85.0% 80.0% -5.0pp Well calibrated
Rolling Accuracy (10-fight window)
Fight 290
40%
Fight 291
40%
Fight 292
50%
Fight 293
60%
Fight 294
60%
Fight 295
60%
Fight 296
50%
Fight 297
50%