The gun selector is NEGATIVE value: Pattern alone beats the full rack (p=0.0012)

MEASURED against the real DrussGT, one frozen binary built from clean HEAD, rack
knobs only (no source edits), 5 arms x 7 runs x 7 rounds, 8 concurrent battles,
judged ONLY on server-side real hit rate from the events sidecar, exact
two-sided permutation test on per-run rates.

  arm             runs  shots  hits  real %  dmg/run  p vs full
  full (shipped)     7   3898   270    6.93     159      --
  onlyPattern        7   4582   494   10.78     287      0.0012  <- BETTER
  onlyKNN            7   4033   207    5.13     119      0.1340
  onlyLinear         7   3215   105    3.27      65      0.0082
  onlyGF             7   3193    72    2.25      45      0.0012

Firing Pattern ALONE gives +3.85pp pooled hit rate and +80% damage per run, and
it fires MORE shots (4582 vs 3898) - it dominates on rate and volume. This is not
"any single gun wins" (full beats Linear, GF and KNN); it is specifically
"Pattern alone beats the rack".

WHY - the virtual fitness signal mis-ranks guns against real outcomes:
- HeadOn is massively over-selected: 31.4% of ticks, the most real shots (1070),
  but only 4.5% REAL. It alone drags the rack down.
- Pattern has the best virtual rank and near-best real rate (11.9%, rank 2), yet
  is selected only 22.6% of the time.
- Linear's apparent strength was SELECTION BIAS: conditional on being selected it
  looked like 15.2% (n=33), but its UNCONDITIONAL rate (onlyLinear) is 3.27%.
  Every earlier per-gun "real rate" in this repo is conditional on selection and
  is therefore confounded. This experiment is the clean measurement.

NOT YET SETTLED (do not overclaim):
- ONE ADVERSARY. All of this is vs DrussGT. Pattern must be re-checked against
  other bots before it becomes the default on this evidence alone.
- Whether a SMALL rack of good guns beats Pattern alone. The selector is negative
  value on the CURRENT bloated rack; that does not prove it is negative value on
  a rack of only good guns. That is the next experiment and it decides whether
  the selection apparatus is fixed or disabled.
- The user's standing directive is to KEEP virtual-fitness selection. This
  measurement conflicts with it, so the next step tests the selector on a small
  good rack rather than assuming either answer.

Context - three prior selection-side attempts all failed: hysteresis (7.02% ->
5.10%, p=0.002), commitment (7.17% -> 4.44%, p=0.0012), arrival-accuracy
tie-break (7.08%, p=0.88 null). The per-tick random draw is load-bearing on
three independent measurements. This experiment locates the real problem one
level up: which guns are in the rack, and that the virtual signal ranks them
wrongly.

Preserves the reusable harness (tools/ab/which_gun_run_one.sh,
which_gun_arm_env.sh, which_gun_analyze.py) and the full writeup
(docs/selector_negative_value.md).
This commit is contained in:
2026-09-22 01:21:14 +02:00
parent 0ede6d12ec
commit e0666a562d
4 changed files with 318 additions and 0 deletions
+64
View File
@@ -0,0 +1,64 @@
# The gun selector is currently NEGATIVE value
**Measured. The best single gun beats the full rack, and not by a little.**
| arm | runs | shots | hits | real % | dmg/run | per-run range | exact two-sided p vs `full` |
|---|---:|---:|---:|---:|---:|---|---:|
| `full` (shipped rack) | 7 | 3898 | 270 | **6.93** | 159 | 2.54–10.37 | — |
| **`onlyPattern`** | 7 | 4582 | 494 | **10.78** | **287** | 9.14–11.85 | **0.0012** |
| `onlyKNN` | 7 | 4033 | 207 | 5.13 | 119 | 4.15–6.47 | 0.1340 |
| `onlyLinear` | 7 | 3215 | 105 | 3.27 | 65 | 1.85–4.69 | 0.0082 |
| `onlyGF` | 7 | 3193 | 72 | 2.25 | 45 | 1.52–3.47 | 0.0012 |
**Firing `Pattern` alone: +3.85 pp pooled hit rate, +80% damage per run, p = 0.0012.**
It also fires MORE shots (4582 vs 3898), so it dominates on rate *and* volume.
This is **not** "any single gun wins" — `full` beats `onlyLinear`, `onlyGF` and `onlyKNN`.
It is specifically **"Pattern alone beats the rack"**.
## Method
One frozen binary built from clean `HEAD` via `git archive` (other agents had `common_libs/guns/*`
dirty), sha256 `f02481d8…`. Arms selected with the rack knobs only (`TR_RACK_<GUN>=off`), no source
edits. 5 arms × 7 runs × 7 rounds, 8 concurrent bridge battles, real DrussGT, judged ONLY on
server-side real hit rate from the events sidecar. Exact two-sided permutation test on per-run rates
(C(14,7)=3432 splits). Each arm's liveness verified from the selected-gun mix.
Harness preserved at `tools/ab/which_gun_*.sh` and `tools/ab/which_gun_analyze.py`.
## Why: the virtual fitness signal mis-ranks guns vs real outcomes
From the `full` arm's own selection mix and per-gun real rates:
- **`HeadOn` is massively over-selected** — **31.4% of ticks**, the most real shots (1070), but only
**4.5% real**. It alone drags the rack down.
- **`Pattern`** has the best *virtual* rank and near-best *real* rate (11.9%, real rank 2), yet is
selected only **22.6%** of the time.
- **`Linear`'s apparent strength was SELECTION BIAS.** Conditional on being selected it looked like
15.2% (n=33); its *unconditional* rate (`onlyLinear`) is **3.27%**. Every earlier per-gun "real
rate" in this repo is conditional on selection and is therefore confounded — this experiment is
the clean measurement.
## What this does NOT yet settle
- **One adversary.** Everything here is vs DrussGT. `Pattern` should be re-checked against other
bots before it becomes the default on the strength of this alone. (Supporting evidence: an offline
audit found `Pattern` is the only gun competitive in *every* distance/speed bucket.)
- **Whether a SMALL good rack beats `Pattern` alone.** The selector is negative value on the current
bloated rack; that does not prove it is negative value on a rack of only good guns. That is the
next experiment, and it decides whether the selection apparatus gets fixed or disabled.
- **The user's directive was to KEEP the virtual-fitness selection mechanism.** This measurement
conflicts with that directive, so the next step is to test the selector on a small, good rack
rather than to assume either answer.
## Prior context: three failed selection-side attempts
| attempt | result |
|---|---|
| hysteresis (commit to incumbent) | 7.02% → 5.10%, p=0.002 |
| commitment (remove the random draw) | 7.17% → 4.44%, p=0.0012 |
| arrival-accuracy tie-break (rank by path, narrow by point) | 7.08%, p=0.88 — null |
So the per-tick random draw is load-bearing on three independent measurements, and no attempt to
"smarten" the tied band has helped. This experiment shows the problem is one level up: **which guns
are in the rack, and the fact that the virtual signal ranks them wrongly.**