To tell a real win-rate gap from noise, divide the gap in percentage points by the tier's standard deviation — 4.40 at Diamond, 5.45 at Gold, 12.23 at Challenger, from this app's own 400,000-lobby-per-tier calibration — and treat anything under about half a standard deviation as indistinguishable from chance. A typical 1.3-point gap between adjacent tier-list rows clears barely a quarter of that bar at Gold.

Tier lists read like ground truth

A tier list puts champions in rows. Row three looks better than row five. Nothing about the format communicates uncertainty — there's no error bar next to a 51.2% win rate, no asterisk saying "this ordering would probably reshuffle if we pulled the data again tomorrow." It reads as settled fact, ranked to the tenth of a percentage point.

Underneath that presentation is a noisier reality than the format admits. Should I Dodge's own win estimate is built from the same kind of aggregate win-rate data every tier list uses, and calibrating it required actually measuring how much noise that data carries at each rank. The result is a number — a standard deviation, per tier — that tells you how much spread to expect from chance alone. Published in full, it gives you a way to ask "is this gap real?" about any tier list you're looking at, not just this app's own numbers.

How the noise floor was measured

The calibration ran offline against a third-party champion tier list, patch 16.14, across all 11 tier scopes, fetched 2026-07-25. For each tier, 400,000 ten-champion lobbies were simulated — 4.4 million lobbies in total — drawing champions in proportion to their real pick rate within their role, then scoring each lobby through the exact formula the live app uses: clamp each win rate to [30%, 70%], convert to log-odds, sum allies and subtract enemies, apply a damping factor of 1.0, and pass the result through a sigmoid to get a probability. See how the calculation works for the full mechanics of that formula.

The output distribution turned out to be almost exactly normal at every tier — skew and excess kurtosis both close to zero, and a Kolmogorov–Smirnov statistic under 0.012 against a fitted normal curve at every single tier. That matters practically: when a distribution is this cleanly normal, its entire spread is described by one number, the standard deviation, and the familiar rule that the middle 38.3% of readings fall within half a standard deviation of the mean, the next 24.2% on each side fall between half and one and a half standard deviations out, and the outer 6.7% on each tail sit beyond that, applies exactly.

The full table

Standard deviation of the win-probability output, in percentage points, per tier:

Tier Standard deviation (points)
Iron7.59
Bronze6.07
Silver5.85
Gold5.45
Platinum5.00
Emerald4.53
Diamond4.40
Master4.77
Grandmaster7.50
Challenger12.23
All tiers combined4.78

Diamond is the tightest tier on the board; Challenger is nearly three times wider. Why the shape curves the way it does — narrowing through the middle of the ladder, then widening sharply at both ends — is its own subject, covered in why tier affects win rates and, for Challenger specifically, in why the highest-elo data is the least trustworthy. This article is about what to do with the numbers once you have them.

Turning a win-rate gap into a noise-width

Here's the useful part. Should I Dodge's damping factor is 1.0, which has a specific consequence: dropped into an otherwise perfectly neutral lobby, a single champion's clamped win rate comes back out as the win probability, unchanged. A 51.8% win rate produces a 51.80% team win probability. That's not a finding, it's an identity — sigmoid and log-odds are inverse functions of each other, so applying one after the other returns the input. But the identity is exactly what makes it possible to compare a champion's win-rate edge against the tier's noise floor on the same scale: a champion's win-rate gap over another, in percentage points, is approximately how many percentage points that swap would move an otherwise-neutral lobby.

Divide that gap by the tier's standard deviation and you get the gap expressed in noise-widths — how many standard deviations of the tier's own ordinary spread the gap actually spans. A 1.8-point win-rate edge is 0.33 noise-widths at Gold (1.8 ÷ 5.45), 0.41 at Diamond (1.8 ÷ 4.40), and just 0.15 at Challenger (1.8 ÷ 12.23). None of those clears even half a standard deviation — the threshold this app itself uses before calling a lobby reading anything other than neutral. A gap that size, at any of those tiers, is comfortably inside the noise floor.

A worked example

Say two junglers sit a row apart on a Gold jungle tier list: 51.6% and 50.3%. That's a realistic-sized gap for two adjacent rows, not a cherry-picked small one — a 1.3 point difference. Divide it by Gold's 5.45 standard deviation and it comes to about 0.24 noise-widths, under a quarter of the 0.5-sigma threshold this app itself requires before it will call a lobby reading anything other than Neutral. Two champions separated by a gap that small aren't shown to be different performers by this data; the tier list is just as consistent with them being identical and the order coming out however that week's sample happened to fall.

Should I Dodge's own tier-list pages don't apply a pick-rate cutoff before ranking — that gate exists elsewhere, on the homepage's much narrower "best picks" shortlist (a 5% pick-rate floor applied in ChampionPickScorer), not on the full ranked list. Every eligible champion in a role gets a row here, which is exactly why these lists run long: /tier-list/mid/gold carries 55 rows, /tier-list/top/gold 54, /tier-list/support/iron 50, /tier-list/adc/challenger 45. Fifty-odd champions packed into a win-rate range that spans only a handful of points once you exclude the extremes means adjacent rows are typically separated by less than the 1.3-point example above, not more — the ungated design that produces so many rows is exactly what compresses the gap between any two of them.

What this means for reading a tier list

Run a gap that size, or smaller, through the noise-widths calculation for the tier you actually care about, and a lot of adjacent-row orderings on a full-length list like this turn out to be well within a fraction of a standard deviation of each other — differences the underlying data can't actually support at the precision the ranking implies.

This doesn't mean tier lists are worthless, and it doesn't mean every gap is noise. A champion sitting ten or fifteen points above the field — the kind of outlier that gets hotfixed — is a real, large signal by this same yardstick, several noise-widths out regardless of tier. The point is narrower and more useful than "don't trust tier lists": don't treat the exact row order as meaningful when the underlying gap is small, and do treat a large, durable gap as exactly the strong signal it looks like. The calibration above is what lets you tell which one you're looking at, rather than reading every row difference as equally real.

Should I Dodge's own tier list pages rank champions by the same underlying win-rate data this calibration was built from — the ranking itself doesn't display a noise-width, but nothing stops you running the same division by hand for any two rows you're actually deciding between. And a champion's win rate is itself a compressed number with its own caveats before you even get to this stage — see what a win rate actually hides for the sample-size and pick-rate issues that sit underneath every figure this article uses.