Files
SirRoboGarage/docs/selector_negative_value.md
T
SirStone e0666a562d The gun selector is NEGATIVE value: Pattern alone beats the full rack (p=0.0012)
MEASURED against the real DrussGT, one frozen binary built from clean HEAD, rack
knobs only (no source edits), 5 arms x 7 runs x 7 rounds, 8 concurrent battles,
judged ONLY on server-side real hit rate from the events sidecar, exact
two-sided permutation test on per-run rates.

  arm             runs  shots  hits  real %  dmg/run  p vs full
  full (shipped)     7   3898   270    6.93     159      --
  onlyPattern        7   4582   494   10.78     287      0.0012  <- BETTER
  onlyKNN            7   4033   207    5.13     119      0.1340
  onlyLinear         7   3215   105    3.27      65      0.0082
  onlyGF             7   3193    72    2.25      45      0.0012

Firing Pattern ALONE gives +3.85pp pooled hit rate and +80% damage per run, and
it fires MORE shots (4582 vs 3898) - it dominates on rate and volume. This is not
"any single gun wins" (full beats Linear, GF and KNN); it is specifically
"Pattern alone beats the rack".

WHY - the virtual fitness signal mis-ranks guns against real outcomes:
- HeadOn is massively over-selected: 31.4% of ticks, the most real shots (1070),
  but only 4.5% REAL. It alone drags the rack down.
- Pattern has the best virtual rank and near-best real rate (11.9%, rank 2), yet
  is selected only 22.6% of the time.
- Linear's apparent strength was SELECTION BIAS: conditional on being selected it
  looked like 15.2% (n=33), but its UNCONDITIONAL rate (onlyLinear) is 3.27%.
  Every earlier per-gun "real rate" in this repo is conditional on selection and
  is therefore confounded. This experiment is the clean measurement.

NOT YET SETTLED (do not overclaim):
- ONE ADVERSARY. All of this is vs DrussGT. Pattern must be re-checked against
  other bots before it becomes the default on this evidence alone.
- Whether a SMALL rack of good guns beats Pattern alone. The selector is negative
  value on the CURRENT bloated rack; that does not prove it is negative value on
  a rack of only good guns. That is the next experiment and it decides whether
  the selection apparatus is fixed or disabled.
- The user's standing directive is to KEEP virtual-fitness selection. This
  measurement conflicts with it, so the next step tests the selector on a small
  good rack rather than assuming either answer.

Context - three prior selection-side attempts all failed: hysteresis (7.02% ->
5.10%, p=0.002), commitment (7.17% -> 4.44%, p=0.0012), arrival-accuracy
tie-break (7.08%, p=0.88 null). The per-tick random draw is load-bearing on
three independent measurements. This experiment locates the real problem one
level up: which guns are in the rack, and that the virtual signal ranks them
wrongly.

Preserves the reusable harness (tools/ab/which_gun_run_one.sh,
which_gun_arm_env.sh, which_gun_analyze.py) and the full writeup
(docs/selector_negative_value.md).
2026-09-22 01:21:14 +02:00

65 lines
3.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# The gun selector is currently NEGATIVE value
**Measured. The best single gun beats the full rack, and not by a little.**
| arm | runs | shots | hits | real % | dmg/run | per-run range | exact two-sided p vs `full` |
|---|---:|---:|---:|---:|---:|---|---:|
| `full` (shipped rack) | 7 | 3898 | 270 | **6.93** | 159 | 2.54–10.37 | — |
| **`onlyPattern`** | 7 | 4582 | 494 | **10.78** | **287** | 9.14–11.85 | **0.0012** |
| `onlyKNN` | 7 | 4033 | 207 | 5.13 | 119 | 4.15–6.47 | 0.1340 |
| `onlyLinear` | 7 | 3215 | 105 | 3.27 | 65 | 1.85–4.69 | 0.0082 |
| `onlyGF` | 7 | 3193 | 72 | 2.25 | 45 | 1.52–3.47 | 0.0012 |
**Firing `Pattern` alone: +3.85 pp pooled hit rate, +80% damage per run, p = 0.0012.**
It also fires MORE shots (4582 vs 3898), so it dominates on rate *and* volume.
This is **not** "any single gun wins" — `full` beats `onlyLinear`, `onlyGF` and `onlyKNN`.
It is specifically **"Pattern alone beats the rack"**.
## Method
One frozen binary built from clean `HEAD` via `git archive` (other agents had `common_libs/guns/*`
dirty), sha256 `f02481d8…`. Arms selected with the rack knobs only (`TR_RACK_<GUN>=off`), no source
edits. 5 arms × 7 runs × 7 rounds, 8 concurrent bridge battles, real DrussGT, judged ONLY on
server-side real hit rate from the events sidecar. Exact two-sided permutation test on per-run rates
(C(14,7)=3432 splits). Each arm's liveness verified from the selected-gun mix.
Harness preserved at `tools/ab/which_gun_*.sh` and `tools/ab/which_gun_analyze.py`.
## Why: the virtual fitness signal mis-ranks guns vs real outcomes
From the `full` arm's own selection mix and per-gun real rates:
- **`HeadOn` is massively over-selected** — **31.4% of ticks**, the most real shots (1070), but only
**4.5% real**. It alone drags the rack down.
- **`Pattern`** has the best *virtual* rank and near-best *real* rate (11.9%, real rank 2), yet is
selected only **22.6%** of the time.
- **`Linear`'s apparent strength was SELECTION BIAS.** Conditional on being selected it looked like
15.2% (n=33); its *unconditional* rate (`onlyLinear`) is **3.27%**. Every earlier per-gun "real
rate" in this repo is conditional on selection and is therefore confounded — this experiment is
the clean measurement.
## What this does NOT yet settle
- **One adversary.** Everything here is vs DrussGT. `Pattern` should be re-checked against other
bots before it becomes the default on the strength of this alone. (Supporting evidence: an offline
audit found `Pattern` is the only gun competitive in *every* distance/speed bucket.)
- **Whether a SMALL good rack beats `Pattern` alone.** The selector is negative value on the current
bloated rack; that does not prove it is negative value on a rack of only good guns. That is the
next experiment, and it decides whether the selection apparatus gets fixed or disabled.
- **The user's directive was to KEEP the virtual-fitness selection mechanism.** This measurement
conflicts with that directive, so the next step is to test the selector on a small, good rack
rather than to assume either answer.
## Prior context: three failed selection-side attempts
| attempt | result |
|---|---|
| hysteresis (commit to incumbent) | 7.02% → 5.10%, p=0.002 |
| commitment (remove the random draw) | 7.17% → 4.44%, p=0.0012 |
| arrival-accuracy tie-break (rank by path, narrow by point) | 7.08%, p=0.88 — null |
So the per-tick random draw is load-bearing on three independent measurements, and no attempt to
"smarten" the tied band has helped. This experiment shows the problem is one level up: **which guns
are in the rack, and the fact that the virtual signal ranks them wrongly.**