Challenger's calibrated standard deviation is 12.23 percentage points — nearly three times Diamond's 4.40, the tightest tier on the board — because its few-hundred-player population per region produces win-rate samples far noisier than Diamond's, and summing ten noisy terms compounds that spread rather than averaging it away. The tier treated as the strongest signal is, by this app's own calibration, the least reliable one.

The assumption worth checking

"What's good in Challenger" gets treated as the strongest possible evidence for what's actually strong — the players are the best, so their picks and their win rates must be the most trustworthy signal on the ladder. Should I Dodge's own calibration data says the opposite: the tightest tier on the entire board sits just one rank below the widest one, by a factor of nearly three.

This isn't a claim that Challenger players are worse at the game, or that their picks are somehow fake. It's a claim about the data those win rates are built from, and it comes straight out of how a win rate is actually measured: as a proportion of games won out of games played.

Why a smaller sample means a noisier number

A win rate estimated from a sample of games carries sampling error the same way any measured proportion does — its standard error shrinks with the square root of how many games it's built from, not with the count itself. Halving the number of games behind an estimate doesn't halve its precision, it grows the error by roughly 41%; cutting the sample by a factor of a hundred grows the error by a factor of ten. Challenger isn't a smaller sample than Diamond by a modest margin — it's a few hundred players per region against the hundreds of thousands who reach Diamond, so the number of games behind any one champion's Challenger win rate is orders of magnitude smaller than the number behind its Diamond figure. The square-root relationship means that gap in sample size shows up as a real, not marginal, gap in how noisy each resulting percentage is.

A champion that's a genuinely stable 52% pick can easily post 47% or 57% in a given Challenger data pull purely from which few hundred games happened to get played that week, with nothing about the champion itself having changed. The same champion's Diamond number, built from vastly more games, barely moves for the same underlying reason — there's simply more data averaging the noise out.

To put the order of magnitude on it rather than just the direction: a proportion estimated from a sample around 50,000 games carries a standard error in the neighborhood of 0.2 percentage points; the same formula applied to a sample two orders of magnitude smaller, around 500 games, gives a standard error in the neighborhood of 2.2 points — roughly ten times larger, for a hundred-fold drop in sample size. Those two sample sizes are illustrative round numbers to show the scaling, not measured figures for any specific champion at either tier, but the ten-fold gap in reliability for a hundred-fold gap in games played is exactly the shape of the problem Challenger's much smaller player base creates.

Why that compounds instead of averaging out

It would be reasonable to guess that ten noisy numbers, summed together, average their noise away rather than making the total noisier. That's not how variance works when you add independent quantities together: variances add, which means standard deviations grow with the square root of how many noisy terms you're summing, not shrink. Should I Dodge's win estimate sums one log-odds term per champion across all ten lobby slots — see how the calculation works — so each of those ten terms carries whatever extra sampling noise its own tier's win-rate data has. At Challenger, where each individual term is noisier to begin with, that noise doesn't cancel across the ten picks; it adds up into a lobby-level spread that's wider in exactly the way the calibration measured. A single noisy input would be a curiosity. Ten of them, summed, is a systematically wider distribution for the whole tier — which is exactly what 12.23 against Diamond's 4.40 describes.

There's a second layer to this that's easy to miss. The calibration doesn't just feed the model each champion's win rate — it draws champions into each simulated lobby in proportion to their real pick rate within their role, so the simulation also needs an estimate of how often each champion gets picked at that tier. That pick-rate estimate is built from the same small Challenger population as the win-rate estimate is, so it carries its own version of the same sampling noise. At Challenger, both the number a champion is scored with and how often it shows up in a simulated lobby at all are less certain than they are at a tier with a larger population behind both figures — two noisy inputs compounding into the total, not one.

Iron and Grandmaster, briefly

Grandmaster's standard deviation, 7.50, sits close to Iron's, 7.59 — a near-exact match between the tier one step below Challenger and the tier at the very bottom of the ladder. Grandmaster's width plausibly comes from the same mechanism as Challenger's, just to a lesser degree: it's a larger population than Challenger but still a small, high-churn one compared to the ranks below it. Iron's cause is murkier — Iron is not a small population, so the thin-sample explanation that fits the top of the ladder doesn't mechanically transfer to the bottom, and the calibration doesn't establish why Iron is wide, only that it is. That open question, and the rest of the shape of the ladder, is covered in why tier affects win rates — this article's focus is specifically the top of the ladder, where the sample-size mechanism is the clearest.

Why this matters beyond the calibration numbers

This isn't only a fact about how Should I Dodge's own bands are calibrated — it's a fact about the underlying win-rate data that any Challenger-tier tier list is drawing from too. "This is what the best players in the world are picking" carries real weight when a trend is durable across patches. It carries much less when it's one data pull's worth of a few hundred players' games, and the two can look identical on the page. A single Challenger snapshot showing a champion at 55% doesn't distinguish between "this is a genuinely strong pick at the highest level" and "this is what the noise happened to look like this week" — and the smaller the population behind the number, the more of the gap between those two explanations the noise alone can account for.

The counterintuitive part

Put plainly: the tier everyone treats as the gold-standard read on what's strong is, by this app's own calibration, the single least reliable tier to read a win rate from. A 58% reading at Challenger is unremarkable — see what a given reading means at your own rank for the full boundary table — while the same number at Diamond would be a genuinely rare result. The same holds one level down for champion win rates directly: a Challenger number that looks decisive is frequently well within the noise floor that a much larger, calmer sample at Diamond or Master simply doesn't have. None of this means Challenger data is useless, or that pro and Challenger trends never matter — some do. It means the instinct to treat the smallest, most elite sample as automatically the most authoritative one has the relationship backwards, and it's worth checking a Challenger-driven claim against a lower, larger-sample tier before trusting it as settled.