Skip to content
XSPORT

Model ops

did the model get better? — eval leaderboards, calibration and paper backtests

Calibration — reliability curves

loading reliability tables

Bins mirror metrics.reliability_table exactly: left-edge label, count, mean predicted winner probability, observed winner rate; empty bins are kept (count 0). Positive gap = under-confident, negative = over-confident. Bins below 0.33 (3-way) / 0.50 (2-way) are structurally empty — the winner probability is a max-prob.