| Win-model recalibration |
rating gaps behave at 0.652 of nominal |
k1 |
| Newcomer mixture |
new accounts split ≈50/50 novice vs restart; separable in ten games at 2.2% false-positive |
k2 |
| Style persistence |
a stable personal accuracy constant, τ = 0.23 ≈ 250 Elo |
k1 |
| Tilt, population line |
next game after a loss: −0.085pp [−0.105, −0.065], n = 142,090 |
tilt-1 |
| After-wins trait |
no population effect; a real minority trait, τ = 2.2pp |
metrics-1 |
| Slump cleanliness |
47% of twenty-game slumps show at-or-above-baseline play (replicated) |
r-luck |
| Skill-dimension lines |
opening, small-mistake, blunder, defence; per-dimension trait ladder |
ins-1/2 |
| Late-middlegame line |
ply 41–60 precision as a personal trait, KS 0.0177 |
ins-3 |
| Night-owl trait |
population clock is flat; late-night decline is personal (τ = 2.3pp, ~7%) |
ins-3 |
| Ambush check |
honest absence; the null holds at nominal |
ins-3 |
| Colour line |
validated-rare: fires 0.27%, median 3.2pp |
colour-y2 |
| Conversion line |
validated-rare, heavy-tailed (ν = 3): 10.7× enrichment over null |
ins-3b |
| Progress factors |
one of five behaviours survives its negative control: quality trend, +0.86 Elo/SD |
factors-1 |
| Style-twin validity |
self-match at 28–50× chance, rising with games |
twin-gates |
| Improvement Verdict, calibration |
coverage 89.1% / 89.1% in-band; constant-truth false-fire 0.25% / 0.00%; two independent passes |
iv2b-4/5 |
| Graded verdict tiers (95 / 80) |
coverage 87.9% (target 90, band 87 to 93); false-fire 0.17% and 4.25% against ceilings 3% and 12% |
iv2c |
| Form verdict |
recent-vs-baseline; constant-truth false-fire 1.0% (95 tier) |
iv2c |