Files
SirRoboGarage/docs/selector_negative_value.md
T
SirStone a54ae6a162 SETTLED: no small rack beats Pattern alone; the selector is negative value on a GOOD rack
Follow-up to e0666a5, which showed Pattern alone (10.78%) beats the full rack
(6.93%). That left two open questions: is a SMALL rack of good guns better than
Pattern alone, and does the selector add value on a good rack (rather than only
on the bloated one)? Both are now answered: NO and NO.

6 arms x 7 runs x 7 rounds, one frozen binary from CLEAN HEAD e0666a5 (built via
`git archive`, source verified byte-identical to the clean tree), rack knobs
only, 8 concurrent battles, real server-side hit rate vs the real DrussGT, exact
two-sided permutation test on per-run rates.

  arm          guns (selector active?)                    real %  dmg/run  p vs onlyPattern
  onlyPattern  Pattern, NO selection                       10.36    264     --
  lean8        HeadOn,Linear,Circular,Accel,Pattern,GF,KNN,WallBounce  6.31  146  0.0169
  lean6        lean8 - HeadOn                               8.83    212     0.0262
  pairPC       Pattern + Circular                           8.23    185     0.0460
  pairPK       Pattern + KNN                                9.80    264     0.3998
  pairPL       Pattern + Linear                             8.23    200     0.0035

The control replicates the prior run (10.36% vs 10.78% before; same binary tree,
different build path).

THE MECHANISM, from the per-arm selected-gun mix - the virtual signal keeps
ranking the WRONG guns first, even on a two-gun rack:
  lean8: HeadOn 46.2% of ticks at 2.0% REAL; Pattern only 14.4% (12.6% real)
  lean6: Pattern 29.2% (10.4% real) vs KNN 25.3% (8.1%) and Linear 15.2% (8.0%)
  pairPC: Circular 66.8% (6.9% real) vs Pattern 33.2% (11.2% real) - over-picks Circular
  pairPL: Linear 57.7% (6.0% real) vs Pattern 42.3% (11.1% real) - over-picks Linear
  pairPK: Pattern 86.4% - ties ONLY because the selector happens to pick Pattern
          most of the time; it is numerically lower with identical dmg/run
So the failure is NOT rack size. Pruning does not fix it; the ranking is wrong.

VERDICT: ship `onlyPattern` - Pattern alone with selection bypassed - at 10.36%
real and 264 dmg/run, vs lean8 6.31%/146 and the prior full rack 6.93%/159.
This DIRECTLY CONTRADICTS the standing user directive to keep virtual-fitness
selection, so it is recorded here plainly rather than quietly acted on: disable
the selector (`TR_RACK_<every gun but PATTERN>=off`) pending a better fitness
signal. The mechanism itself is left intact and functional so it can be re-enabled
with one env var, and so it can be fixed rather than discarded.

REMAINING CAVEAT: ONE ADVERSARY. All of this is vs DrussGT. Pattern as the default
must be re-checked against other bots first - that is the next job.

Extends the reusable harness (tools/ab/which_gun_arm_env.sh now has lean8/lean6/
pairPC/pairPK/pairPL; which_gun_analyze.py is parameterised by WHICHGUN_OUT and
compares against both `full` and `onlyPattern`).
2026-09-22 01:35:51 +02:00

124 lines
7.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# The gun selector is currently NEGATIVE value
**Measured. The best single gun beats the full rack, and not by a little.**
| arm | runs | shots | hits | real % | dmg/run | per-run range | exact two-sided p vs `full` |
|---|---:|---:|---:|---:|---:|---|---:|
| `full` (shipped rack) | 7 | 3898 | 270 | **6.93** | 159 | 2.54–10.37 | — |
| **`onlyPattern`** | 7 | 4582 | 494 | **10.78** | **287** | 9.14–11.85 | **0.0012** |
| `onlyKNN` | 7 | 4033 | 207 | 5.13 | 119 | 4.15–6.47 | 0.1340 |
| `onlyLinear` | 7 | 3215 | 105 | 3.27 | 65 | 1.85–4.69 | 0.0082 |
| `onlyGF` | 7 | 3193 | 72 | 2.25 | 45 | 1.52–3.47 | 0.0012 |
**Firing `Pattern` alone: +3.85 pp pooled hit rate, +80% damage per run, p = 0.0012.**
It also fires MORE shots (4582 vs 3898), so it dominates on rate *and* volume.
This is **not** "any single gun wins" — `full` beats `onlyLinear`, `onlyGF` and `onlyKNN`.
It is specifically **"Pattern alone beats the rack"**.
## Follow-up: a small rack of GOOD guns STILL loses to `Pattern` alone
The obvious objection to the result above — "the rack is bloated, so of course it loses; prune it" — was
tested directly. Same protocol (one frozen binary, rack knobs only, real DrussGT, server-side real hit
rate), **6 arms × 7 runs × 7 rounds**. `lean8`/`lean6` are the offline audit's recommended racks, with the
selector ACTIVE; `onlyPattern` is the no-selection control.
| arm | guns (selector active?) | runs | shots | hits | real % | dmg/run | per-run range | exact two-sided p vs `onlyPattern` |
|---|---|---:|---:|---:|---:|---:|---|---:|
| **`onlyPattern`** (control) | Pattern, NO selection | 7 | 4325 | 448 | **10.36** | **264** | 8.96–11.47 | — |
| `lean8` | HeadOn, Linear, Circular, Accel, Pattern, GF, KNN, WallBounce | 7 | 3711 | 234 | 6.31 | 146 | 2.64–12.99 | 0.0169 |
| `lean6` | lean8 − HeadOn | 7 | 4113 | 363 | 8.83 | 212 | 7.35–10.87 | 0.0262 |
| `pairPC` | Pattern + Circular | 7 | 3862 | 318 | 8.23 | 185 | 3.36–11.64 | 0.0460 |
| `pairPK` | Pattern + KNN | 7 | 4571 | 448 | 9.80 | 264 | 7.90–12.76 | 0.3998 |
| `pairPL` | Pattern + Linear | 7 | 4109 | 338 | 8.23 | 200 | 6.14–9.90 | 0.0035 |
**No small rack beats `Pattern` alone.** Every selector-active arm is worse or (in one case) statistically
indistinguishable:
- **`lean8` is far worse** — 6.31% vs 10.36%, **p = 0.0169**. Pruning to the audit's recommended good guns
did NOT rescue the selector.
- **`lean6` (HeadOn removed) is still worse** — 8.83% vs 10.36%, **p = 0.0262**. Removing the single
most-over-selected gun helps (+2.5 pp over lean8) but still does not beat no-selection.
- **`pairPK` is the only near-tie** — 9.80% vs 10.36%, p = 0.3998, identical dmg/run. It ties only because
the selector happens to pick `Pattern` **86.4%** of the time on a 2-gun rack; it does not *beat* Pattern.
### Verdict: the selector is negative value on a good rack too — DISABLE it
`lean6` and `lean8` (selector active) are both significantly WORSE than `Pattern` alone. This contradicts
the standing directive to keep the virtual-fitness selection mechanism, so it is stated plainly: **the
selection apparatus should be disabled (`TR_RACK_<every gun but PATTERN>=off`) pending a better fitness
signal.** Keeping it costs ~1.5–4 pp of real hit rate on a good rack.
### Why the pruned racks still lose: the same mis-ranking, visible in the mix
The selected-gun mix (liveness proof) shows the virtual signal still over-selects the wrong guns:
- **`lean8`**: `HeadOn` **46.2%** of ticks → only **2.0% real** (34/1670); `WallBounce` **21.0%** → 6.5%
real (35/540); `Pattern` only **14.4%** → **12.6% real** (65/514). HeadOn and WallBounce crowd out the
best gun.
- **`lean6`** (HeadOn gone): `Pattern` **29.2%** → **10.4% real** (127/1216) — the best real gun — while
`KNN` 25.3% → 8.1%, `Linear` 15.2% → 8.0%, `Accel` 13.8% → 8.7%. The rack is still diluted.
- **`pairPL`**: the selector gives `Linear` **57.7%** of ticks → **6.0% real** (103/1712), and `Pattern`
only 42.3% → **11.1% real**. Over-selecting Linear is why `pairPL` is the *worst* 2-gun arm (p = 0.0035).
The failure is not "the rack is too big". It is that the **virtual fitness signal ranks the wrong guns
first**, on a big rack and on a small one alike.
### Ship recommendation
**`onlyPattern`** — `Pattern` alone, selector bypassed. Measured: **10.36% real hit rate, 264 damage/run**
(vs `lean8` 6.31%/146 and `full` 6.93%/159). This replicates the original finding (10.78%/287).
Method note: one frozen binary built from clean `e0666a5` via `git archive` (other agents had
`common_libs/guns/*` dirty and have since committed further changes; this build predates them), sha256
`df8d4f2e…`. 6 arms × 7 runs × 7 rounds, 8 concurrent bridge battles. Liveness verified for every arm from
`gun_stats.jsonl`; the `[rack] active=…` line confirmed each arm's membership at process start.
## Method
One frozen binary built from clean `HEAD` via `git archive` (other agents had `common_libs/guns/*`
dirty), sha256 `f02481d8…`. Arms selected with the rack knobs only (`TR_RACK_<GUN>=off`), no source
edits. 5 arms × 7 runs × 7 rounds, 8 concurrent bridge battles, real DrussGT, judged ONLY on
server-side real hit rate from the events sidecar. Exact two-sided permutation test on per-run rates
(C(14,7)=3432 splits). Each arm's liveness verified from the selected-gun mix.
Harness preserved at `tools/ab/which_gun_*.sh` and `tools/ab/which_gun_analyze.py`.
## Why: the virtual fitness signal mis-ranks guns vs real outcomes
From the `full` arm's own selection mix and per-gun real rates:
- **`HeadOn` is massively over-selected** — **31.4% of ticks**, the most real shots (1070), but only
**4.5% real**. It alone drags the rack down.
- **`Pattern`** has the best *virtual* rank and near-best *real* rate (11.9%, real rank 2), yet is
selected only **22.6%** of the time.
- **`Linear`'s apparent strength was SELECTION BIAS.** Conditional on being selected it looked like
15.2% (n=33); its *unconditional* rate (`onlyLinear`) is **3.27%**. Every earlier per-gun "real
rate" in this repo is conditional on selection and is therefore confounded — this experiment is
the clean measurement.
## What this does NOT yet settle
- **One adversary.** Everything here is vs DrussGT. `Pattern` should be re-checked against other
bots before it becomes the default on the strength of this alone. (Supporting evidence: an offline
audit found `Pattern` is the only gun competitive in *every* distance/speed bucket.)
- **Whether a SMALL good rack beats `Pattern` alone — SETTLED (follow-up above): NO.** A 6-arm
follow-up (`lean8`, `lean6`, `pairPC`, `pairPK`, `pairPL`) found every selector-active rack worse or tied;
`lean6` 8.83% and `lean8` 6.31% both lose to `Pattern` alone 10.36% (p = 0.026 / 0.017). The selector is
negative value on a good rack too.
- **The user's directive was to KEEP the virtual-fitness selection mechanism — now directly tested.** The
follow-up tested the selector on small, good racks instead of assuming either answer; it lost there too,
so the directive conflicts with the evidence and disabling is recommended pending a better signal.
## Prior context: three failed selection-side attempts
| attempt | result |
|---|---|
| hysteresis (commit to incumbent) | 7.02% → 5.10%, p=0.002 |
| commitment (remove the random draw) | 7.17% → 4.44%, p=0.0012 |
| arrival-accuracy tie-break (rank by path, narrow by point) | 7.08%, p=0.88 — null |
So the per-tick random draw is load-bearing on three independent measurements, and no attempt to
"smarten" the tied band has helped. This experiment shows the problem is one level up: **which guns
are in the rack, and the fact that the virtual signal ranks them wrongly.**