Follow-up toe0666a5, which showed Pattern alone (10.78%) beats the full rack (6.93%). That left two open questions: is a SMALL rack of good guns better than Pattern alone, and does the selector add value on a good rack (rather than only on the bloated one)? Both are now answered: NO and NO. 6 arms x 7 runs x 7 rounds, one frozen binary from CLEAN HEADe0666a5(built via `git archive`, source verified byte-identical to the clean tree), rack knobs only, 8 concurrent battles, real server-side hit rate vs the real DrussGT, exact two-sided permutation test on per-run rates. arm guns (selector active?) real % dmg/run p vs onlyPattern onlyPattern Pattern, NO selection 10.36 264 -- lean8 HeadOn,Linear,Circular,Accel,Pattern,GF,KNN,WallBounce 6.31 146 0.0169 lean6 lean8 - HeadOn 8.83 212 0.0262 pairPC Pattern + Circular 8.23 185 0.0460 pairPK Pattern + KNN 9.80 264 0.3998 pairPL Pattern + Linear 8.23 200 0.0035 The control replicates the prior run (10.36% vs 10.78% before; same binary tree, different build path). THE MECHANISM, from the per-arm selected-gun mix - the virtual signal keeps ranking the WRONG guns first, even on a two-gun rack: lean8: HeadOn 46.2% of ticks at 2.0% REAL; Pattern only 14.4% (12.6% real) lean6: Pattern 29.2% (10.4% real) vs KNN 25.3% (8.1%) and Linear 15.2% (8.0%) pairPC: Circular 66.8% (6.9% real) vs Pattern 33.2% (11.2% real) - over-picks Circular pairPL: Linear 57.7% (6.0% real) vs Pattern 42.3% (11.1% real) - over-picks Linear pairPK: Pattern 86.4% - ties ONLY because the selector happens to pick Pattern most of the time; it is numerically lower with identical dmg/run So the failure is NOT rack size. Pruning does not fix it; the ranking is wrong. VERDICT: ship `onlyPattern` - Pattern alone with selection bypassed - at 10.36% real and 264 dmg/run, vs lean8 6.31%/146 and the prior full rack 6.93%/159. This DIRECTLY CONTRADICTS the standing user directive to keep virtual-fitness selection, so it is recorded here plainly rather than quietly acted on: disable the selector (`TR_RACK_<every gun but PATTERN>=off`) pending a better fitness signal. The mechanism itself is left intact and functional so it can be re-enabled with one env var, and so it can be fixed rather than discarded. REMAINING CAVEAT: ONE ADVERSARY. All of this is vs DrussGT. Pattern as the default must be re-checked against other bots first - that is the next job. Extends the reusable harness (tools/ab/which_gun_arm_env.sh now has lean8/lean6/ pairPC/pairPK/pairPL; which_gun_analyze.py is parameterised by WHICHGUN_OUT and compares against both `full` and `onlyPattern`).
7.6 KiB
The gun selector is currently NEGATIVE value
Measured. The best single gun beats the full rack, and not by a little.
| arm | runs | shots | hits | real % | dmg/run | per-run range | exact two-sided p vs full |
|---|---|---|---|---|---|---|---|
full (shipped rack) |
7 | 3898 | 270 | 6.93 | 159 | 2.54–10.37 | — |
onlyPattern |
7 | 4582 | 494 | 10.78 | 287 | 9.14–11.85 | 0.0012 |
onlyKNN |
7 | 4033 | 207 | 5.13 | 119 | 4.15–6.47 | 0.1340 |
onlyLinear |
7 | 3215 | 105 | 3.27 | 65 | 1.85–4.69 | 0.0082 |
onlyGF |
7 | 3193 | 72 | 2.25 | 45 | 1.52–3.47 | 0.0012 |
Firing Pattern alone: +3.85 pp pooled hit rate, +80% damage per run, p = 0.0012.
It also fires MORE shots (4582 vs 3898), so it dominates on rate and volume.
This is not "any single gun wins" — full beats onlyLinear, onlyGF and onlyKNN.
It is specifically "Pattern alone beats the rack".
Follow-up: a small rack of GOOD guns STILL loses to Pattern alone
The obvious objection to the result above — "the rack is bloated, so of course it loses; prune it" — was
tested directly. Same protocol (one frozen binary, rack knobs only, real DrussGT, server-side real hit
rate), 6 arms × 7 runs × 7 rounds. lean8/lean6 are the offline audit's recommended racks, with the
selector ACTIVE; onlyPattern is the no-selection control.
| arm | guns (selector active?) | runs | shots | hits | real % | dmg/run | per-run range | exact two-sided p vs onlyPattern |
|---|---|---|---|---|---|---|---|---|
onlyPattern (control) |
Pattern, NO selection | 7 | 4325 | 448 | 10.36 | 264 | 8.96–11.47 | — |
lean8 |
HeadOn, Linear, Circular, Accel, Pattern, GF, KNN, WallBounce | 7 | 3711 | 234 | 6.31 | 146 | 2.64–12.99 | 0.0169 |
lean6 |
lean8 − HeadOn | 7 | 4113 | 363 | 8.83 | 212 | 7.35–10.87 | 0.0262 |
pairPC |
Pattern + Circular | 7 | 3862 | 318 | 8.23 | 185 | 3.36–11.64 | 0.0460 |
pairPK |
Pattern + KNN | 7 | 4571 | 448 | 9.80 | 264 | 7.90–12.76 | 0.3998 |
pairPL |
Pattern + Linear | 7 | 4109 | 338 | 8.23 | 200 | 6.14–9.90 | 0.0035 |
No small rack beats Pattern alone. Every selector-active arm is worse or (in one case) statistically
indistinguishable:
lean8is far worse — 6.31% vs 10.36%, p = 0.0169. Pruning to the audit's recommended good guns did NOT rescue the selector.lean6(HeadOn removed) is still worse — 8.83% vs 10.36%, p = 0.0262. Removing the single most-over-selected gun helps (+2.5 pp over lean8) but still does not beat no-selection.pairPKis the only near-tie — 9.80% vs 10.36%, p = 0.3998, identical dmg/run. It ties only because the selector happens to pickPattern86.4% of the time on a 2-gun rack; it does not beat Pattern.
Verdict: the selector is negative value on a good rack too — DISABLE it
lean6 and lean8 (selector active) are both significantly WORSE than Pattern alone. This contradicts
the standing directive to keep the virtual-fitness selection mechanism, so it is stated plainly: the
selection apparatus should be disabled (TR_RACK_<every gun but PATTERN>=off) pending a better fitness
signal. Keeping it costs ~1.5–4 pp of real hit rate on a good rack.
Why the pruned racks still lose: the same mis-ranking, visible in the mix
The selected-gun mix (liveness proof) shows the virtual signal still over-selects the wrong guns:
lean8:HeadOn46.2% of ticks → only 2.0% real (34/1670);WallBounce21.0% → 6.5% real (35/540);Patternonly 14.4% → 12.6% real (65/514). HeadOn and WallBounce crowd out the best gun.lean6(HeadOn gone):Pattern29.2% → 10.4% real (127/1216) — the best real gun — whileKNN25.3% → 8.1%,Linear15.2% → 8.0%,Accel13.8% → 8.7%. The rack is still diluted.pairPL: the selector givesLinear57.7% of ticks → 6.0% real (103/1712), andPatternonly 42.3% → 11.1% real. Over-selecting Linear is whypairPLis the worst 2-gun arm (p = 0.0035).
The failure is not "the rack is too big". It is that the virtual fitness signal ranks the wrong guns first, on a big rack and on a small one alike.
Ship recommendation
onlyPattern — Pattern alone, selector bypassed. Measured: 10.36% real hit rate, 264 damage/run
(vs lean8 6.31%/146 and full 6.93%/159). This replicates the original finding (10.78%/287).
Method note: one frozen binary built from clean e0666a5 via git archive (other agents had
common_libs/guns/* dirty and have since committed further changes; this build predates them), sha256
df8d4f2e…. 6 arms × 7 runs × 7 rounds, 8 concurrent bridge battles. Liveness verified for every arm from
gun_stats.jsonl; the [rack] active=… line confirmed each arm's membership at process start.
Method
One frozen binary built from clean HEAD via git archive (other agents had common_libs/guns/*
dirty), sha256 f02481d8…. Arms selected with the rack knobs only (TR_RACK_<GUN>=off), no source
edits. 5 arms × 7 runs × 7 rounds, 8 concurrent bridge battles, real DrussGT, judged ONLY on
server-side real hit rate from the events sidecar. Exact two-sided permutation test on per-run rates
(C(14,7)=3432 splits). Each arm's liveness verified from the selected-gun mix.
Harness preserved at tools/ab/which_gun_*.sh and tools/ab/which_gun_analyze.py.
Why: the virtual fitness signal mis-ranks guns vs real outcomes
From the full arm's own selection mix and per-gun real rates:
HeadOnis massively over-selected — 31.4% of ticks, the most real shots (1070), but only 4.5% real. It alone drags the rack down.Patternhas the best virtual rank and near-best real rate (11.9%, real rank 2), yet is selected only 22.6% of the time.Linear's apparent strength was SELECTION BIAS. Conditional on being selected it looked like 15.2% (n=33); its unconditional rate (onlyLinear) is 3.27%. Every earlier per-gun "real rate" in this repo is conditional on selection and is therefore confounded — this experiment is the clean measurement.
What this does NOT yet settle
- One adversary. Everything here is vs DrussGT.
Patternshould be re-checked against other bots before it becomes the default on the strength of this alone. (Supporting evidence: an offline audit foundPatternis the only gun competitive in every distance/speed bucket.) - Whether a SMALL good rack beats
Patternalone — SETTLED (follow-up above): NO. A 6-arm follow-up (lean8,lean6,pairPC,pairPK,pairPL) found every selector-active rack worse or tied;lean68.83% andlean86.31% both lose toPatternalone 10.36% (p = 0.026 / 0.017). The selector is negative value on a good rack too. - The user's directive was to KEEP the virtual-fitness selection mechanism — now directly tested. The follow-up tested the selector on small, good racks instead of assuming either answer; it lost there too, so the directive conflicts with the evidence and disabling is recommended pending a better signal.
Prior context: three failed selection-side attempts
| attempt | result |
|---|---|
| hysteresis (commit to incumbent) | 7.02% → 5.10%, p=0.002 |
| commitment (remove the random draw) | 7.17% → 4.44%, p=0.0012 |
| arrival-accuracy tie-break (rank by path, narrow by point) | 7.08%, p=0.88 — null |
So the per-tick random draw is load-bearing on three independent measurements, and no attempt to "smarten" the tied band has helped. This experiment shows the problem is one level up: which guns are in the rack, and the fact that the virtual signal ranks them wrongly.