When a venue says 70%, how often does it happen? Reliability curves and Expected Calibration Error per venue, computed from resolved markets. Venue transparency report →
Scores the probability each market assigned ~24h before it resolved, over events whose complete outcome book was captured. Lower ECE is better.
| Predicted | Predicted mean | Realized rate | Pairs | |
|---|---|---|---|---|
| 0-10% | 3.6% | 1.7% | 231 | |
| 10-20% | 14.3% | 14.3% | 91 | |
| 20-30% | 24.3% | 16.2% | 76 | |
| 30-40% | 34.5% | 29.8% | 33 | |
| 40-50% | 45.8% | 44.3% | 46 | |
| 50-60% | 53.6% | 57.1% | 52 | |
| 60-70% | 65.0% | 74.6% | 43 | |
| 70-80% | 73.8% | 88.7% | 29 | |
| 80-90% | 84.9% | 84.3% | 46 | |
| 90-100% | 96.2% | 97.0% | 103 |
| Predicted | Predicted mean | Realized rate | Pairs | |
|---|---|---|---|---|
| 0-10% | 4.0% | 3.2% | 190 | |
| 10-20% | 14.1% | 16.0% | 66 | |
| 20-30% | 24.9% | 16.8% | 71 | |
| 30-40% | 35.0% | 28.3% | 56 | |
| 40-50% | 46.3% | 50.9% | 119 | |
| 50-60% | 53.4% | 48.9% | 118 | |
| 60-70% | 64.8% | 73.2% | 51 | |
| 70-80% | 75.1% | 82.8% | 41 | |
| 80-90% | 85.8% | 83.9% | 31 | |
| 90-100% | 95.8% | 96.8% | 94 |
| Predicted | Predicted mean | Realized rate | Pairs | |
|---|---|---|---|---|
| 0-10% | 1.1% | 8.0% | 6,217 | |
| 10-20% | 15.3% | 10.9% | 321 | |
| 20-30% | 25.2% | 21.9% | 289 | |
| 30-40% | 35.2% | 32.0% | 286 | |
| 40-50% | 44.9% | 39.3% | 483 | |
| 50-60% | 54.6% | 55.4% | 472 | |
| 60-70% | 64.2% | 64.2% | 258 | |
| 70-80% | 74.4% | 68.5% | 229 | |
| 80-90% | 84.2% | 69.0% | 225 | |
| 90-100% | 94.5% | 66.0% | 323 |
Not yet scoreable: ForecastEx (no provider-confirmed resolution history yet) · Futuur (no provider-confirmed resolution history yet) · Gemini (only 8 scorable events (need 30)) · Limitless (no provider-confirmed resolution history yet) · Metaculus (no provider-confirmed resolution history yet) · Myriad (no provider-confirmed resolution history yet) · PredictIt (no provider-confirmed resolution history yet) · Robinhood (Rothera) (no provider-confirmed resolution history yet) · Smarkets (no provider-confirmed resolution history yet)
Calibration compares the probability the market assigned to each outcome ~24h before resolution against the realised result, over resolved markets with provider-confirmed outcomes and ONE captured price snapshot in the 20-28h window before resolution. That is a single pre-resolution observation, NOT a claim of continuous 24h history: a market with no snapshot that far back (most same-day markets) is excluded entirely rather than scored on a nearer capture. Reliability bins group (outcome, probability) pairs; a well-calibrated venue's realised win rate tracks its predicted probability. calibrationError is the event-weighted mean gap (ECE, lower is better): each resolved event contributes one unit regardless of outcome count, preventing large outcome catalogs from dominating venue comparisons. Plausibly complete books are normalized to remove overround; independent ladders remain raw. An event is scored only when the t-24h snapshot captured the market's COMPLETE outcome book: a partial snapshot would make inclusion depend on whether the eventual winner happened to be captured, which biases the result. Markets with no capture in that window and partial-book snapshots are both excluded, and the per-venue counts are published in `excluded` — the exclusion rate differs sharply by venue, so compare sampleSize alongside calibrationError.
Scores the venue-reported final pre-resolution price over resolved two-outcome markets. Lead time varies, so this series is not comparable with the 24h series above.
Basis: venue-reported final pre-resolution price.
| Year | Scored events | Calibration error (ECE) | |
|---|---|---|---|
| 2017 | 10,732 | 8.72pp | |
| 2018 | 22,445 | 9.19pp | |
| 2019 | 21,480 | 2.05pp | |
| 2020 | 8,384 | 0.52pp | |
| 2021 | 14,113 | 1.03pp | |
| 2022 | 14,396 | 2.42pp | |
| 2023 | 14,160 | 1.19pp | |
| 2024 | 12,857 | 1.03pp | |
| 2025 | 14,460 | 1.44pp | |
| 2026 | 9,021 | 1.52pp |
| Predicted | Predicted mean | Realized rate | Pairs | |
|---|---|---|---|---|
| 0-10% | 5.7% | 4.4% | 8,897 | |
| 10-20% | 14.7% | 11.8% | 15,527 | |
| 20-30% | 25.0% | 21.9% | 21,662 | |
| 30-40% | 35.1% | 30.8% | 37,178 | |
| 40-50% | 44.4% | 41.6% | 49,079 | |
| 50-60% | 53.5% | 55.6% | 63,575 | |
| 60-70% | 63.9% | 68.0% | 39,285 | |
| 70-80% | 74.0% | 77.3% | 22,766 | |
| 80-90% | 84.3% | 87.1% | 15,712 | |
| 90-100% | 93.7% | 95.1% | 10,421 |
The final-price series scores the probability the venue itself reported as the market's last pre-resolution price (persisted at ingest) against the realised result, over resolved two-outcome markets whose winner is one of the outcomes and whose both outcome prices are present. Lead time before resolution is unknown and not uniform, so these figures are NOT comparable with the t-24h series above and are reported as a separate basis. Plausibly complete books (price sum 95-125) are normalized to remove overround; others are read raw. Currently published only for venues whose persisted resolved prices were verified to be genuine trading prices rather than settled 0/100 values (Futuur, verified 2026-08-18 across a monotone 2016-2026 reliability curve). Per-year rows below the minimum sample are withheld.
Scores the last probability CoinRithm captured before each market resolved against the realised result. With a median lead of minutes this is a near-final price integrity check across venues, not a forecast-skill measure.
Basis: our own last pre-resolution capture. Lead time is published per venue; not comparable with the 24-hour series.
| Year | Scored events | Calibration error (ECE) | |
|---|---|---|---|
| 2026 | 12,647 | 0.41pp |
| Predicted | Predicted mean | Realized rate | Pairs | |
|---|---|---|---|---|
| 0-10% | 1.0% | 0.8% | 8,913 | |
| 10-20% | 14.9% | 13.7% | 1,022 | |
| 20-30% | 25.1% | 25.9% | 953 | |
| 30-40% | 35.2% | 35.2% | 1,005 | |
| 40-50% | 45.6% | 46.1% | 1,329 | |
| 50-60% | 54.1% | 53.6% | 1,393 | |
| 60-70% | 64.7% | 64.7% | 1,000 | |
| 70-80% | 74.8% | 73.6% | 953 | |
| 80-90% | 84.9% | 86.4% | 990 | |
| 90-100% | 98.9% | 99.2% | 8,347 |
Basis: our own last pre-resolution capture. Lead time is published per venue; not comparable with the 24-hour series.
| Year | Scored events | Calibration error (ECE) | |
|---|---|---|---|
| 2026 | 1,237 | 1.68pp |
| Predicted | Predicted mean | Realized rate | Pairs | |
|---|---|---|---|---|
| 0-10% | 1.8% | 0.9% | 1,709 | |
| 10-20% | 14.1% | 7.4% | 172 | |
| 20-30% | 24.6% | 26.3% | 91 | |
| 30-40% | 34.4% | 33.0% | 76 | |
| 40-50% | 46.0% | 43.3% | 93 | |
| 50-60% | 53.6% | 55.5% | 93 | |
| 60-70% | 65.1% | 68.7% | 55 | |
| 70-80% | 74.8% | 74.9% | 66 | |
| 80-90% | 85.5% | 92.8% | 108 | |
| 90-100% | 98.0% | 99.2% | 885 |
Basis: our own last pre-resolution capture. Lead time is published per venue; not comparable with the 24-hour series.
| Year | Scored events | Calibration error (ECE) | |
|---|---|---|---|
| 2026 | 43,476 | 2.22pp |
| Predicted | Predicted mean | Realized rate | Pairs | |
|---|---|---|---|---|
| 0-10% | 2.4% | 1.1% | 105,872 | |
| 10-20% | 14.5% | 9.6% | 7,167 | |
| 20-30% | 24.7% | 19.1% | 5,665 | |
| 30-40% | 34.7% | 31.1% | 5,966 | |
| 40-50% | 44.6% | 43.5% | 6,448 | |
| 50-60% | 54.5% | 55.2% | 6,425 | |
| 60-70% | 64.6% | 67.8% | 5,910 | |
| 70-80% | 74.6% | 79.4% | 5,224 | |
| 80-90% | 84.9% | 89.0% | 6,121 | |
| 90-100% | 97.2% | 96.6% | 22,710 |
Basis: our own last pre-resolution capture. Lead time is published per venue; not comparable with the 24-hour series.
| Year | Scored events | Calibration error (ECE) | |
|---|---|---|---|
| 2026 | 341 | 6.52pp |
| Predicted | Predicted mean | Realized rate | Pairs | |
|---|---|---|---|---|
| 0-10% | 3.9% | 1.4% | 147 | |
| 10-20% | 15.2% | 11.9% | 59 | |
| 20-30% | 25.3% | 18.4% | 38 | |
| 30-40% | 34.2% | 40.0% | 30 | |
| 40-50% | 46.7% | 26.6% | 64 | |
| 50-60% | 53.0% | 71.4% | 70 | |
| 60-70% | 65.6% | 62.1% | 29 | |
| 70-80% | 74.5% | 79.0% | 38 | |
| 80-90% | 84.7% | 88.1% | 59 | |
| 90-100% | 96.1% | 98.7% | 148 |
Basis: our own last pre-resolution capture. Lead time is published per venue; not comparable with the 24-hour series.
| Year | Scored events | Calibration error (ECE) | |
|---|---|---|---|
| 2026 | 287 | 7.37pp |
| Predicted | Predicted mean | Realized rate | Pairs | |
|---|---|---|---|---|
| 0-10% | 4.6% | 0.0% | 146 | |
| 10-20% | 14.3% | 7.4% | 93 | |
| 20-30% | 24.1% | 6.5% | 72 | |
| 30-40% | 35.5% | 25.3% | 62 | |
| 40-50% | 45.1% | 36.1% | 51 | |
| 50-60% | 55.0% | 49.7% | 63 | |
| 60-70% | 64.5% | 72.1% | 55 | |
| 70-80% | 76.0% | 84.0% | 64 | |
| 80-90% | 85.5% | 80.2% | 70 | |
| 90-100% | 95.4% | 90.2% | 127 |
The own-capture series scores the LAST probability CoinRithm itself captured strictly before the provider-reported resolution time (from the durable frozen timeline) against the realised result. Lead time varies and is typically SHORT — per-venue median and p90 lead seconds are published — so with a median lead of minutes this is a near-final-price integrity check over our own observations, NOT a forecast-skill measure, and is not comparable with the t-24h series. An event is scored only when the captured point holds the market's complete outcome book (the same complete-book rule as the t-24h series: a partial capture makes inclusion depend on whether the winner was captured). Plausibly complete books (price sum 95-125) are normalized to remove overround; others are read raw. Events are unit-weighted regardless of outcome count. Coverage grows with forward capture and spans venues the t-24h series undersamples. Per-year rows below the minimum sample are withheld.