SETTLED: no small rack beats Pattern alone; the selector is negative value on a GOOD rack

Follow-up to e0666a5, which showed Pattern alone (10.78%) beats the full rack
(6.93%). That left two open questions: is a SMALL rack of good guns better than
Pattern alone, and does the selector add value on a good rack (rather than only
on the bloated one)? Both are now answered: NO and NO.

6 arms x 7 runs x 7 rounds, one frozen binary from CLEAN HEAD e0666a5 (built via
`git archive`, source verified byte-identical to the clean tree), rack knobs
only, 8 concurrent battles, real server-side hit rate vs the real DrussGT, exact
two-sided permutation test on per-run rates.

  arm          guns (selector active?)                    real %  dmg/run  p vs onlyPattern
  onlyPattern  Pattern, NO selection                       10.36    264     --
  lean8        HeadOn,Linear,Circular,Accel,Pattern,GF,KNN,WallBounce  6.31  146  0.0169
  lean6        lean8 - HeadOn                               8.83    212     0.0262
  pairPC       Pattern + Circular                           8.23    185     0.0460
  pairPK       Pattern + KNN                                9.80    264     0.3998
  pairPL       Pattern + Linear                             8.23    200     0.0035

The control replicates the prior run (10.36% vs 10.78% before; same binary tree,
different build path).

THE MECHANISM, from the per-arm selected-gun mix - the virtual signal keeps
ranking the WRONG guns first, even on a two-gun rack:
  lean8: HeadOn 46.2% of ticks at 2.0% REAL; Pattern only 14.4% (12.6% real)
  lean6: Pattern 29.2% (10.4% real) vs KNN 25.3% (8.1%) and Linear 15.2% (8.0%)
  pairPC: Circular 66.8% (6.9% real) vs Pattern 33.2% (11.2% real) - over-picks Circular
  pairPL: Linear 57.7% (6.0% real) vs Pattern 42.3% (11.1% real) - over-picks Linear
  pairPK: Pattern 86.4% - ties ONLY because the selector happens to pick Pattern
          most of the time; it is numerically lower with identical dmg/run
So the failure is NOT rack size. Pruning does not fix it; the ranking is wrong.

VERDICT: ship `onlyPattern` - Pattern alone with selection bypassed - at 10.36%
real and 264 dmg/run, vs lean8 6.31%/146 and the prior full rack 6.93%/159.
This DIRECTLY CONTRADICTS the standing user directive to keep virtual-fitness
selection, so it is recorded here plainly rather than quietly acted on: disable
the selector (`TR_RACK_<every gun but PATTERN>=off`) pending a better fitness
signal. The mechanism itself is left intact and functional so it can be re-enabled
with one env var, and so it can be fixed rather than discarded.

REMAINING CAVEAT: ONE ADVERSARY. All of this is vs DrussGT. Pattern as the default
must be re-checked against other bots first - that is the next job.

Extends the reusable harness (tools/ab/which_gun_arm_env.sh now has lean8/lean6/
pairPC/pairPK/pairPL; which_gun_analyze.py is parameterised by WHICHGUN_OUT and
compares against both `full` and `onlyPattern`).
This commit is contained in:
2026-09-22 01:35:51 +02:00
parent 4657fe715e
commit a54ae6a162
3 changed files with 94 additions and 22 deletions
+65 -6
View File
@@ -16,6 +16,64 @@ It also fires MORE shots (4582 vs 3898), so it dominates on rate *and* volume.
This is **not** "any single gun wins" — `full` beats `onlyLinear`, `onlyGF` and `onlyKNN`.
It is specifically **"Pattern alone beats the rack"**.
## Follow-up: a small rack of GOOD guns STILL loses to `Pattern` alone
The obvious objection to the result above — "the rack is bloated, so of course it loses; prune it" — was
tested directly. Same protocol (one frozen binary, rack knobs only, real DrussGT, server-side real hit
rate), **6 arms × 7 runs × 7 rounds**. `lean8`/`lean6` are the offline audit's recommended racks, with the
selector ACTIVE; `onlyPattern` is the no-selection control.
| arm | guns (selector active?) | runs | shots | hits | real % | dmg/run | per-run range | exact two-sided p vs `onlyPattern` |
|---|---|---:|---:|---:|---:|---:|---|---:|
| **`onlyPattern`** (control) | Pattern, NO selection | 7 | 4325 | 448 | **10.36** | **264** | 8.96–11.47 | — |
| `lean8` | HeadOn, Linear, Circular, Accel, Pattern, GF, KNN, WallBounce | 7 | 3711 | 234 | 6.31 | 146 | 2.64–12.99 | 0.0169 |
| `lean6` | lean8 − HeadOn | 7 | 4113 | 363 | 8.83 | 212 | 7.35–10.87 | 0.0262 |
| `pairPC` | Pattern + Circular | 7 | 3862 | 318 | 8.23 | 185 | 3.36–11.64 | 0.0460 |
| `pairPK` | Pattern + KNN | 7 | 4571 | 448 | 9.80 | 264 | 7.90–12.76 | 0.3998 |
| `pairPL` | Pattern + Linear | 7 | 4109 | 338 | 8.23 | 200 | 6.14–9.90 | 0.0035 |
**No small rack beats `Pattern` alone.** Every selector-active arm is worse or (in one case) statistically
indistinguishable:
- **`lean8` is far worse** — 6.31% vs 10.36%, **p = 0.0169**. Pruning to the audit's recommended good guns
did NOT rescue the selector.
- **`lean6` (HeadOn removed) is still worse** — 8.83% vs 10.36%, **p = 0.0262**. Removing the single
most-over-selected gun helps (+2.5 pp over lean8) but still does not beat no-selection.
- **`pairPK` is the only near-tie** — 9.80% vs 10.36%, p = 0.3998, identical dmg/run. It ties only because
the selector happens to pick `Pattern` **86.4%** of the time on a 2-gun rack; it does not *beat* Pattern.
### Verdict: the selector is negative value on a good rack too — DISABLE it
`lean6` and `lean8` (selector active) are both significantly WORSE than `Pattern` alone. This contradicts
the standing directive to keep the virtual-fitness selection mechanism, so it is stated plainly: **the
selection apparatus should be disabled (`TR_RACK_<every gun but PATTERN>=off`) pending a better fitness
signal.** Keeping it costs ~1.5–4 pp of real hit rate on a good rack.
### Why the pruned racks still lose: the same mis-ranking, visible in the mix
The selected-gun mix (liveness proof) shows the virtual signal still over-selects the wrong guns:
- **`lean8`**: `HeadOn` **46.2%** of ticks → only **2.0% real** (34/1670); `WallBounce` **21.0%** → 6.5%
real (35/540); `Pattern` only **14.4%** → **12.6% real** (65/514). HeadOn and WallBounce crowd out the
best gun.
- **`lean6`** (HeadOn gone): `Pattern` **29.2%** → **10.4% real** (127/1216) — the best real gun — while
`KNN` 25.3% → 8.1%, `Linear` 15.2% → 8.0%, `Accel` 13.8% → 8.7%. The rack is still diluted.
- **`pairPL`**: the selector gives `Linear` **57.7%** of ticks → **6.0% real** (103/1712), and `Pattern`
only 42.3% → **11.1% real**. Over-selecting Linear is why `pairPL` is the *worst* 2-gun arm (p = 0.0035).
The failure is not "the rack is too big". It is that the **virtual fitness signal ranks the wrong guns
first**, on a big rack and on a small one alike.
### Ship recommendation
**`onlyPattern`** — `Pattern` alone, selector bypassed. Measured: **10.36% real hit rate, 264 damage/run**
(vs `lean8` 6.31%/146 and `full` 6.93%/159). This replicates the original finding (10.78%/287).
Method note: one frozen binary built from clean `e0666a5` via `git archive` (other agents had
`common_libs/guns/*` dirty and have since committed further changes; this build predates them), sha256
`df8d4f2e…`. 6 arms × 7 runs × 7 rounds, 8 concurrent bridge battles. Liveness verified for every arm from
`gun_stats.jsonl`; the `[rack] active=…` line confirmed each arm's membership at process start.
## Method
One frozen binary built from clean `HEAD` via `git archive` (other agents had `common_libs/guns/*`
@@ -44,12 +102,13 @@ From the `full` arm's own selection mix and per-gun real rates:
- **One adversary.** Everything here is vs DrussGT. `Pattern` should be re-checked against other
bots before it becomes the default on the strength of this alone. (Supporting evidence: an offline
audit found `Pattern` is the only gun competitive in *every* distance/speed bucket.)
- **Whether a SMALL good rack beats `Pattern` alone.** The selector is negative value on the current
bloated rack; that does not prove it is negative value on a rack of only good guns. That is the
next experiment, and it decides whether the selection apparatus gets fixed or disabled.
- **The user's directive was to KEEP the virtual-fitness selection mechanism.** This measurement
conflicts with that directive, so the next step is to test the selector on a small, good rack
rather than to assume either answer.
- **Whether a SMALL good rack beats `Pattern` alone — SETTLED (follow-up above): NO.** A 6-arm
follow-up (`lean8`, `lean6`, `pairPC`, `pairPK`, `pairPL`) found every selector-active rack worse or tied;
`lean6` 8.83% and `lean8` 6.31% both lose to `Pattern` alone 10.36% (p = 0.026 / 0.017). The selector is
negative value on a good rack too.
- **The user's directive was to KEEP the virtual-fitness selection mechanism — now directly tested.** The
follow-up tested the selector on small, good racks instead of assuming either answer; it lost there too,
so the directive conflicts with the evidence and disabling is recommended pending a better signal.
## Prior context: three failed selection-side attempts
+15 -8
View File
@@ -7,8 +7,10 @@ Also reads gun_stats.jsonl to prove each arm was LIVE (selected-gun mix).
"""
import json, glob, os, re, sys, collections, itertools, math
OUT = "/tmp/whichgun"
ARMS = ["full", "onlyLinear", "onlyPattern", "onlyGF", "onlyKNN"]
OUT = os.environ.get("WHICHGUN_OUT", "/tmp/whichgun")
ARMS = ["full", "onlyLinear", "onlyPattern", "onlyGF", "onlyKNN",
"lean8", "lean6", "pairPC", "pairPK", "pairPL"]
BASELINES = ["full", "onlyPattern"]
MB_POWERS = {1.0, 1.5, 2.0, 3.0}
@@ -40,7 +42,10 @@ def ev_stats(path, subject_fired):
line = line.strip()
if not line:
continue
o = json.loads(line)
try:
o = json.loads(line)
except json.JSONDecodeError:
continue # partial line from an in-progress battle
t = o.get("type")
if t == "fire":
fires[o["owner"]] += 1
@@ -178,15 +183,17 @@ def main():
pr = " ".join(f"r{p['run']}={p['rate']:.2f}" for p in res[a]["per"])
print(f" {a:<12} {pr}")
if "full" in res and len(arms) > 1:
for base in BASELINES:
if base not in res or len(arms) < 2:
continue
print("\n" + "=" * 96)
print("EXACT TWO-SIDED PERMUTATION TEST vs `full` (per-run rates)")
print(f"EXACT TWO-SIDED PERMUTATION TEST vs `{base}` (per-run rates)")
print("=" * 96)
for a in arms:
if a == "full":
if a == base:
continue
obs, p = perm_test_exact(res["full"], res[a])
print(f" full vs {a:<12}: mean-rate diff = {obs:+.2f} pp, exact two-sided p = {p:.4f}")
obs, p = perm_test_exact(res[base], res[a])
print(f" {base} vs {a:<12}: mean-rate diff = {obs:+.2f} pp, exact two-sided p = {p:.4f}")
print("\n" + "=" * 96)
print("LIVENESS — selected-gun mix (sum of per-tick `selected` over all rounds)")
+14 -8
View File
@@ -1,5 +1,5 @@
# emits `TR_RACK_<GUN>=off ...` for every gun except the one to keep.
# usage: arm_env.sh <full|onlyLinear|onlyPattern|onlyGF|onlyKNN>
# emits `TR_RACK_<GUN>=off ...` for every gun NOT in the arm.
# usage: arm_env.sh <full|onlyLinear|onlyPattern|onlyGF|onlyKNN|lean8|lean6|pairPC|pairPK|pairPL>
arm="$1"
GUNS="HEADON LINEAR TSETLIN CIRCULAR GUESSFACTOR PATTERN WALLBOUNCE ACCEL STOPSHOT DISPLACE AVGLEAD DECAYGF KNN TMSELECT"
case "$arm" in
@@ -8,13 +8,19 @@ case "$arm" in
onlyPattern) keep="PATTERN";;
onlyGF) keep="GUESSFACTOR";;
onlyKNN) keep="KNN";;
# ── small-good-rack follow-up ──────────────────────────────────────────────
lean8) keep="HEADON LINEAR CIRCULAR ACCEL PATTERN GUESSFACTOR KNN WALLBOUNCE";;
lean6) keep="LINEAR CIRCULAR ACCEL PATTERN GUESSFACTOR KNN";;
pairPC) keep="PATTERN CIRCULAR";;
pairPK) keep="PATTERN KNN";;
pairPL) keep="PATTERN LINEAR";;
*) echo "unknown arm $arm" >&2; exit 1;;
esac
out=""
if [ -n "$keep" ]; then
for g in $GUNS; do
[ "$g" = "$keep" ] && continue
out="$out TR_RACK_${g}=off"
done
fi
for g in $GUNS; do
case " $keep " in
*" $g "*) continue;;
esac
out="$out TR_RACK_${g}=off"
done
echo "$out"