Arcifact Verify

The Gate Ledger

Every claim below was tested against a pre-registered gate. Passes, failures, and in-flight work are listed in full. Receipt records attach at launch.

Validated

Claim Key figure Gate
Win-model recalibration rating gaps behave at 0.652 of nominal k1
Newcomer mixture new accounts split ≈50/50 novice vs restart; separable in ten games at 2.2% false-positive k2
Style persistence a stable personal accuracy constant, τ = 0.23 ≈ 250 Elo k1
Tilt, population line next game after a loss: −0.085pp [−0.105, −0.065], n = 142,090 tilt-1
After-wins trait no population effect; a real minority trait, τ = 2.2pp metrics-1
Slump cleanliness 47% of twenty-game slumps show at-or-above-baseline play (replicated) r-luck
Skill-dimension lines opening, small-mistake, blunder, defence; per-dimension trait ladder ins-1/2
Late-middlegame line ply 41–60 precision as a personal trait, KS 0.0177 ins-3
Night-owl trait population clock is flat; late-night decline is personal (τ = 2.3pp, ~7%) ins-3
Ambush check honest absence; the null holds at nominal ins-3
Colour line validated-rare: fires 0.27%, median 3.2pp colour-y2
Conversion line validated-rare, heavy-tailed (ν = 3): 10.7× enrichment over null ins-3b
Progress factors one of five behaviours survives its negative control: quality trend, +0.86 Elo/SD factors-1
Style-twin validity self-match at 28–50× chance, rising with games twin-gates
Improvement Verdict, calibration coverage 89.1% / 89.1% in-band; constant-truth false-fire 0.25% / 0.00%; two independent passes iv2b-4/5
Graded verdict tiers (95 / 80) coverage 87.9% (target 90, band 87 to 93); false-fire 0.17% and 4.25% against ceilings 3% and 12% iv2c
Form verdict recent-vs-baseline; constant-truth false-fire 1.0% (95 tier) iv2c

Published failures

Finding What it means
Volume, opposition strength, post-loss grinding, colour share all failed negative controls as improvement factors; the association runs from rating to behaviour, not the reverse
Deep-endgame line confounded by selection: who reaches deep endgames is itself stylistic
Breadth line instrument invalid as built: opening choice arrives in session streaks, so naive intervals understate variance. Returns with corrected intervals
Evidence-selected drift model retired. With results-only data at year scale, model selection cannot detect real drift
Threshold revision (v2 decision rule) failed its registered gates: power at +150 reached 15% against a 20% minimum, and the spoken-rate prediction band was exceeded. The shipped rule is retained; the full budget-versus-power curve is published
The endpoint question below ±150 Elo retired. Any model free enough to follow real change carries about ±90 Elo of irreducible width. A structural limit of the data, stated here

In flight

Programme Status
Certified population mixture + cohort error control registered
Two-channel fusion registered; the next major cycle
Cross-platform transfer gates registered
Extended condition fields registered

A registered programme is a sealed pre-registration with a dated hash. It ships when its gate passes.