The same win rate means less at some tiers than others because the underlying sample's spread differs by tier: Should I Dodge's own calibration puts the standard deviation of its win-probability output at 4.40 points at Diamond (the tightest) and 12.23 at Challenger (the widest, nearly three times as noisy). A gap that looks decisive at Challenger can be well within ordinary noise elsewhere.
The same champion, a different number
Pull up a champion's win rate at Iron and again at Challenger and they're rarely the same figure. Part of that is real — champions that reward mechanical execution or long-term planning perform differently when the average player can actually pull them off — but part of it is something less obvious: the reliability of the number itself changes with tier, because the size and shape of the sample changes with tier.
What the model's own calibration shows
Should I Dodge's win estimate is checked against a large offline simulation — 400,000 simulated lobbies per tier — that measures how spread out the model's output actually is at each tier. The result is a standard deviation per tier, in percentage points, and it isn't flat:
- Iron: 7.59
- Bronze: 6.07
- Silver: 5.85
- Gold: 5.45
- Platinum: 5.00
- Emerald: 4.53
- Diamond: 4.40
- Master: 4.77
- Grandmaster: 7.50
- Challenger: 12.23
- All tiers combined: 4.78
That list doesn't move in a straight line from "low rank, high spread" to "high rank, low spread" — it narrows through the middle of the ladder and then widens sharply again at the very top. Diamond is the tightest tier on the list, and Challenger is nearly three times wider. This isn't a different formula at different tiers; it's the same calculation run against real pick data at each tier.
Why the wide tiers are wide
Three tiers stand out as unusually wide: Iron, Grandmaster, and Challenger. At the top of the ladder the mechanism is fairly clear — Challenger is a few hundred players per region, not the hundreds of thousands playing Gold or Silver, so its champion win-rate samples are thinner and noisier. A champion that's a genuine, stable 52% pick can easily show up as 48% or 56% in a given data pull at a tier where only a few hundred games decide the number, purely from sampling variance, with no underlying change in the champion at all. See why that makes Challenger tier lists the least reliable ones people quote most. Grandmaster sits just below Challenger, in a similarly small, high-turnover population, and lands at a similarly wide 7.50 — plausibly noisy for the same kind of reason.
Iron is the harder case. Its 7.59 standard deviation sits almost exactly where Grandmaster's does — a genuinely striking coincidence, given the two tiers sit at opposite ends of the ladder in skill. But Iron isn't a small population the way Challenger and Grandmaster are, so the thin-sample explanation that works at the top of the ladder doesn't obviously carry over to the bottom. The calibration data doesn't establish why Iron is wide, only that it is — it would be a mistake to assume both ends are noisy for the same underlying reason when the data only supports that claim for one of them.
What the calibration also doesn't tell us is why the tightest tier is Diamond rather than, say, Silver or Gold, which have far larger player populations. If population size alone drove the spread, those high-population mid tiers would be the tightest on the list, and they aren't — Diamond is, at 4.40, tighter than Gold's 5.45. One plausible explanation is that champion win rates are simply more dispersed at lower ranks generally, independent of sample size, so the same formula produces a wider spread of lobby totals there — but that's a hypothesis, not something this simulation established.
The same principle shows up with partial lobbies
Tier isn't the only place this shows up. When champion select isn't finished yet — some picks still open, or a few recognised picks unmatched — the app widens its uncertainty rather than narrowing it, for the same underlying reason: fewer known picks is thinner evidence, not a tighter reading. It would be tempting to assume six known picks produce a more confident number just because there's less left to guess at, but the unfilled slots still represent real champions who haven't been chosen yet, and treating them as exactly average understates how much they could still swing the result. So a six-champion reading gets wider bands than a ten-champion one, never narrower — the same "less data, more caution" logic that separates Diamond's tight distribution from Challenger's wide one.
Why not use one spread for every tier
It would be simpler to calibrate the model once, against the whole player base, and use that single spread everywhere. But a single flat spread would be wrong at both ends of the ladder in opposite directions: it would flag ordinary Challenger noise as an extreme reading, and undersell how unusual a genuinely lopsided Diamond lobby actually is. Calibrating per tier costs more up front — simulating hundreds of thousands of lobbies at each of eleven tier scopes instead of one — but it means a verdict means the same thing regardless of which tier produced it: ordinary readings look ordinary, and genuinely unusual ones stand out, at every tier.
What this means in practice
A reading of, say, 58% means something different depending on which tier it came from. In Diamond, where the typical spread is about 4.4 points, 58% is a noticeably one-sided lobby — well outside the tier's normal middle band. In Challenger, where the typical spread is over 12 points, that same 58% is barely outside the ordinary range — the kind of reading you'd see regularly just from the noise baked into a small, elite population. Should I Dodge's verdict bands account for this directly, judging a reading against its own tier's distribution rather than a single fixed scale, so "ordinary" and "extreme" mean what they should at every tier instead of applying a Diamond-shaped yardstick to a Challenger lobby, or vice versa.
The practical takeaway: don't compare a raw percentage across tiers as if it means the same thing everywhere, and don't read the same number as equally alarming regardless of where you're queuing. For a concrete way to tell a real gap from noise at a given tier, see this app's own published noise floor. See what a win rate hides for the sample-size issues that drive this in the first place.