Every prediction was published before kickoff and settled after the fact. We took the whole archive — wins and losses mixed together — and checked the thing that matters most: does our confidence actually mean anything. Below is what held up, where our weak spot is, and why we don't show ROI yet.
207 settled predictions are the ones whose match result is already known. Two bots: a football bot (a testing ground on national-team matches, 108 picks) and a KHL hockey bot (99 picks). A prediction is locked in the moment it's published and never edited afterward. Every number below comes with a 95% confidence interval — the range the true value likely falls in. The wider the interval, the less data behind it, and the more cautiously you should trust it.
| Segment | Hit | Win rate | 95% interval | Rating |
|---|---|---|---|---|
| Both bots, full track | 140/207 | 67.6% | 61.0 – 73.6% | solid |
| Football, full track | 79/108 | 73.1% | 64.1 – 80.6% | solid |
| KHL hockey, full track | 61/99 | 61.6% | 51.8 – 70.6% | solid |
Anyone can write "95% confidence." There's only one way to check it: group predictions by their stated confidence and see what share actually hit. If the scale is honest, the actual line rises along with confidence. Ours does. That's not something you can fake after the fact.
The line climbs left to right — the higher the confidence, the more often the pick hits. But across 108 matches, each individual bucket still carries a wide margin of error (thin bars). We show this openly instead of hiding it: the direction is clear, and precision builds with every season.
The top bucket states 94 on average, and actually hits 89.7% — the model slightly overrates itself. The dip at 70–74 (61.5%) is real; it's our known weak spot.
On top of the statistics runs an AI review: it reads current lineups, injuries, and motivation, then renders a verdict — confirm the pick, lower confidence, or cancel it. The honest question is: does it add value, or is it just burning tokens? On football, the difference is measurable.
| AI verdict · football | Hit | Win rate | 95% interval | Rating |
|---|---|---|---|---|
| Confirmed the pick | 38/50 | 76.0% | 62.6 – 85.7% | solid |
| AI didn't trigger (pure statistics) | 29/41 | 70.7% | 55.5 – 82.4% | solid |
When AI confirms the statistical signal, hit rate goes up by +5 points. Not huge, but consistent: the intervals still overlap at the edges, so this is more a directional signal than a proven fact. We keep measuring.
A piece about trust has to start with its own rules. Here they are, no polish.
Every pick is published before kickoff, with a date and time, and is never edited after that. No "we knew it all along" after the fact.
The whole archive runs in chronological order — wins and losses mixed together. A losing pick stays a losing pick. We don't quietly clean up the cold streaks.
Before this season we weren't recording the odds at the moment a pick was published — which means an honest ROI calculation isn't possible, and we're not going to fabricate one. Starting this new season, the odds are locked in together with the prediction, and ROI will appear.
A high hit rate on heavy favorites at low odds doesn't by itself mean you're making money. That's why we lead with calibration and treat the hit rate as supporting evidence.
Starting in September, the track record becomes fully verifiable — and measurement only runs forward, with no after-the-fact edits.