20 KiB
Gun campaign — ledger
Goal (owner's mandate, 2026-09-26 overnight): the movement campaign produced
a replicated champion (TR_MOVEMENT=strafe, docs/movement_campaign.md) which a
parallel job is shipping. "Continue until you found an amazing movement. When
found do the same over for a gun." This file is the GUN campaign's single source
of truth: later jobs append a ## Batch N section and never edit an earlier
one (a wrong earlier number gets a correction line, not a rewrite). This file is
the gun successor to docs/movement_campaign.md; read that first, then this.
0. The question
The shipped rack admits Pattern only (onlyPattern). That decision was
taken because the virtual-fitness selector measured negative value at every
rack size tested, and Pattern is the best single gun by the DrussGT hit-rate
table (docs/selector_negative_value.md, docs/gun_rack_analysis.md, commit
e0666a5). But every one of those measurements — like every pre-campaign
movement claim — is DrussGT-heavy, and the movement campaign's standing
lesson is that a one-opponent result is not a result.
So Batch 1 asks the direct question, on a frozen 15-opponent panel:
Does any rack gun, or any small gun configuration, beat the shipped
Patternon damage/run AND round wins across 15 opponents — or isonlyPatternCONFIRMED rather than merely assumed?
A null is a successful outcome here: it upgrades onlyPattern from an
assumption to a measured verdict across the panel, which is itself valuable.
1. Protocol (identical to the movement campaign; the instrument is reused verbatim)
tools/ab/tournament_run.sh + tools/ab/tournament_analyze.py +
tools/ab/panel_movement.txt (frozen panel), reused unchanged. The unit of
evidence is the number of opponents, not the number of runs.
| Element | Rule |
|---|---|
| Subject | ONE frozen binary built from git archive HEAD (tournament_run.sh records commit + binary sha256 in session.json) |
| Arms | env dicts only — no per-arm rebuild, ever; the arm file is a committed file (tools/ab/arms_gun_b1.txt) |
| Panel | the frozen tools/ab/panel_movement.txt (15 opponents). Adding/removing an opponent starts a new batch number |
| Pairing | per opponent: average the arm's runs, subtract the reference's average for that same opponent → one delta per opponent; then aggregate |
| Isolation | per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir |
| Serialization | one battle fleet at a time via --wait-arena; never a broad pkill robocode_shim |
| Liveness | every declared env token must appear verbatim in OUR bot's own raw-env report, else the run is excluded and named |
| Never shipped | this is a measurement campaign: git status clean, shipped defaults untouched, .gitignore untouched |
Movement is PINNED in every arm: TR_MOVEMENT=strafe
Reason (hard rule). A movement-default flip is landing from job j120 during
this run, so the default engine can change underneath us. Every arm therefore
sets TR_MOVEMENT=strafe explicitly. This (a) removes movement as a
confound, (b) makes all arms share exactly one movement, so every delta is a pure
gun delta, and (c) satisfies the analyzer's liveness rule for every arm
(each declared token must appear verbatim; an undeclared leaked TR_MOVEMENT
is fatal, but here every arm declares it). The reference arm is therefore the
shipped gun rack under the pinned movement, not under the (mutable) default.
2. Pre-registered decision rules (fixed BEFORE Batch 1 ran)
- Primary metrics: damage/run and ROUND WINS. Secondary/explanation only: damage taken/run, incoming hit rate, mean distance. Never hit rate alone — that trap has inverted six verdicts in this project.
- BETTER than the reference iff one primary metric is up with a
cross-opponent sign test p < 0.05 while the other does not go down;
the mirror image for WORSE. Anything else is NOT DISTINGUISHABLE (a
real answer, not a failure). Both the strict reading (the other metric's mean
delta
>= 0) and the substantive reading (not detectably down: not significant and smaller than that metric's MDE) are printed. - A verdict must survive the between-opponent spread: the pooled mean delta is reported with the SD across opponents, its SE, a 95% CI, and the MDE (α=0.05 two-sided, 80% power). An effect smaller than the MDE is reported as not detectable — never as absent, never as a win.
- Somewhere to stop: if no arm beats the shipped
patternby rule 2 in Batch 1 and no arm shows a>= +MDEdamage gain with p<0.10, then the gun stage's first phase is closed with "the shippedonlyPatternrack is the best gun configuration we have measured across the panel" — that is a successful outcome, and the campaign moves to a named next axis rather than inventing more arms. See What would make us stop. - No promotion off a single metric, a single opponent, or a single run. A change that wins damage by losing wins (or vice-versa) is not a win.
- Every batch is shot with a pre-registered prediction stated in its section before the battles finish; a prediction that turns out wrong is recorded as wrong.
3. Stage 0 — what we already know (given, not re-derived)
| fact | value | source |
|---|---|---|
| shipped rack | onlyPattern (id 5), all other guns off |
common_libs/gun_harness/selector.nim |
| why | selector measured negative value at every rack size (16, lean8, lean6, pairPC/PK/PL) and 10/10 adversaries | docs/selector_negative_value.md, docs/gun_rack_analysis.md |
| Pattern hit rate vs DrussGT | ~10.8% | given |
| KNN / Linear / Circular / WallBounce / GF vs DrussGT | 5.6 / 3.0 / 2.9 / 2.7 / 2.1 % | given |
| BitBrain when idle | == Pattern (30 runs/arm: 97/210 vs 97/210 wins) | docs/bitbrain_vs_tmhorizon_ab.md |
| BitBrain across 32 legacy opponents | neutral: -3.6 dmg/run, +0.02 wins/run, p=0.53/0.91 | docs/gauntlet_bitbrain_vs_pattern.md |
| lead amplitude | dead axis — full gain sweep {0, 0.25-1.0, 1.0, 1.25, 1.5, 2.0} run, nothing beats Pattern | docs/bitbrain_campaign.md |
| offline prediction checks | veto-only — may kill a design, never select one | docs/offline_harness_trust.md |
| spinner-specific claim | UNTESTED: no constant-turn spinner exists in the legacy roster (a note for a later job, not built here unless nearly free) | docs/gauntlet_bitbrain_vs_pattern.md |
Movement context (why the pin): the movement champion is TR_MOVEMENT=strafe
at its current defaults (range 325, tol 25, tilt 15/0.10); it replicated its win
over the shipped tfil in four independent sessions (Batch 4: 52.6% vs 38.5%
round-win rate). Movement is held constant at that champion for every gun arm.
Batch 1 — does anything beat the shipped Pattern?
Design. One frozen binary, six env-only arms, one frozen panel
(tools/ab/panel_movement.txt, 15 opponents: 5 dodger, 3 pattern, 1
wall-follower, 1 corner-camper, 1 spinner, 1 rammer, 1 brawler, 2 aggressive
megas), 3 runs × 3 rounds per (opponent, arm) = 270 battles. Arm file:
tools/ab/arms_gun_b1.txt. Reference: pattern.
| # | arm | env | what it isolates |
|---|---|---|---|
| 1 | pattern |
TR_MOVEMENT=strafe |
the arm to beat (shipped onlyPattern rack) |
| 2 | bitbrain |
+ TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay |
BitBrain-only, learned gain (panel replication of the 32-opp gauntlet at the pinned-movement standard) |
| 3 | tmhorizon |
+ TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both |
TMHorizon-only corrector (predicted to lose; never panel-tested) |
| 4 | knn |
+ TR_RACK_PATTERN=off TR_RACK_KNN=both |
best non-Pattern single gun (never panel-tested) |
| 5 | rack_pk |
+ TR_RACK_KNN=both |
Pattern + KNN, selector ON (different-family hedge) |
| 6 | rack_pt |
+ TR_RACK_TMHORIZON=both |
Pattern + TMHorizon, selector ON (corrector-family hedge) |
(+ = the pinned TR_MOVEMENT=strafe is present in every arm; see §1.)
Pre-registered prediction (written BEFORE the battles finished):
- No arm beats
patternon BOTH primaries, and the point estimates sit inside the MDE.onlyPatternis confirmed across the panel. bitbrainis a wash vspattern(replicating the 32-opponent gauntlet: ±3.6 dmg/run, ≈0 wins) — the DrussGT-only gain-config penalty does not carry.tmhorizonis WORSE thanpattern(its own offline work predicts it needs ~80% side accuracy and can reach ~60%); no damage win, no win win.knnis WORSE thanpatternon damage (its DrussGT hit rate is roughly half Pattern's) and not better on wins.- The two small racks do not beat
pattern;rack_pklands closer to Pattern thanknn-alone does (the selector at least partially hedges back), but the selector's poor ranking keeps it below the reference. - If any surprise exists, it is a rack arm on damage without wins — the
ring-mover mirror-image trap — and it will not be read as a win.
Outcome — direct answer
MEASURED. Session /tmp/ab/j121_g1, commit 1d8143a15e4a038c39dfc2153ff518fa003b8612
(the pre-registration commit), frozen binary sha256 aa49a45fec20…, 15 opponents
× 6 arms × 3 runs × 3 rounds = 270 battles, 0 failed, 0 never started, 0
liveness exclusions. Every arm declared TR_MOVEMENT=strafe, verified in each
bot's own raw-env report; the reference rack line reads PATTERN for pattern
and the overridden rack reads BITBRAIN / TMHORIZON / KNN for the others.
DIRECT ANSWER (MEASURED, n=15 opponents): NO — nothing beats the shipped
Patternacross the panel on damage/run AND round wins. Every arm exceptknnis NOT DISTINGUISHABLE frompatternunder both the strict and the substantive reading of the pre-registered rule. The best nominal challenger,tmhorizon, is +0.13 wins/run (MDE 0.36) and +8.7 dmg/run (MDE 13.8) — an effect ~3× smaller than the design can detect — andbitbrainis+0.11wins /+10.3dmg, also inside the MDE.knnis detectably WORSE on both primaries. The shippedonlyPatternrack is therefore CONFIRMED rather than merely assumed, at the resolution of this panel. That is a real result, not a failure.
Pooled dashboard (all valid runs — explanation only, NOT the verdict)
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|---|---|---|---|---|---|---|---|---|
pattern (REF) |
45 | 99.6 | 155.9 | 1.51 | 68/135 | 50.4% | 12.84% | 435 |
bitbrain |
45 | 109.9 | 151.2 | 1.62 | 73/135 | 54.1% | 12.87% | 427 |
tmhorizon |
45 | 108.3 | 153.7 | 1.64 | 74/135 | 54.8% | 12.53% | 428 |
knn |
45 | 60.3 | 175.7 | 0.89 | 40/135 | 29.6% | 13.31% | 430 |
rack_pk |
45 | 107.3 | 161.3 | 1.42 | 64/135 | 47.4% | 13.65% | 432 |
rack_pt |
45 | 105.3 | 149.9 | 1.64 | 74/135 | 54.8% | 12.90% | 424 |
Cross-opponent aggregation (the verdict layer, damage + wins)
| arm | metric | mean Δ | spread (SD) | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|---|---|---|---|---|---|---|---|---|---|
bitbrain |
damage | +10.34 | 16.65 | [+1.12, +19.57] | 11/15 | 0.1185 | 0.03247 | 0.05006 | 12.05 |
bitbrain |
wins | +0.11 | 0.48 | [-0.16, +0.38] | 6/9 | 0.5078 | 0.4922 | 0.4764 | 0.35 |
tmhorizon |
damage | +8.71 | 19.10 | [-1.86, +19.29] | 11/15 | 0.1185 | 0.09918 | 0.1183 | 13.82 |
tmhorizon |
wins | +0.13 | 0.50 | [-0.14, +0.41] | 5/9 | 1.0 | 0.4062 | 0.3118 | 0.36 |
knn |
damage | −39.33 | 23.15 | [−52.15, −26.51] | 0/15 | 6.1e-5 | 6.1e-5 | 0.0007 | 16.75 |
knn |
wins | −0.62 | 0.59 | [−0.95, −0.30] | 1/11 | 0.01172 | 0.001953 | 0.004948 | 0.43 |
rack_pk |
damage | +7.68 | 32.09 | [−10.10, +25.45] | 9/15 | 0.6072 | 0.423 | 0.5895 | 23.21 |
rack_pk |
wins | −0.09 | 0.60 | [−0.42, +0.24] | 6/11 | 1.0 | 0.6738 | 0.5932 | 0.43 |
rack_pt |
damage | +5.69 | 16.50 | [−3.45, +14.83] | 10/15 | 0.3018 | 0.2026 | 0.2681 | 11.93 |
rack_pt |
wins | +0.13 | 0.41 | [−0.10, +0.36] | 6/10 | 0.7539 | 0.3242 | 0.1997 | 0.30 |
The pre-registered verdict (verbatim)
| rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) |
|---|---|---|---|---|---|---|---|
| 1 | tmhorizon |
+0.13 | +8.7 | 5/9 p=1 | 11/15 p=0.1185 | not distinguishable | not distinguishable |
| 2 | rack_pt |
+0.13 | +5.7 | 6/10 p=0.7539 | 10/15 p=0.3018 | not distinguishable | not distinguishable |
| 3 | bitbrain |
+0.11 | +10.3 | 6/9 p=0.5078 | 11/15 p=0.1185 | not distinguishable | not distinguishable |
| 4 | rack_pk |
−0.09 | +7.7 | 6/11 p=1 | 9/15 p=0.6072 | not distinguishable | not distinguishable |
| 5 | knn |
−0.62 | −39.3 | 1/11 p=0.01172 | 0/15 p=6.104e-05 | WORSE | not distinguishable |
Reference pattern: 99.6 dmg/run, 1.51 wins/run, 12.84% incoming, 435 px.
Reading (MEASURED, with the mechanism)
knnis a clean, large loss and is dropped. It deals −39 dmg/run and loses −0.62 wins/run, positive on 0/15 opponents on damage and 1/11 on wins (p=6e-5 / p=0.012). Its incoming hit rate is no worse (13.31% vs 12.84%) — it simply cannot aim: the same movement, a worse gun.- The two small racks do not beat
pattern.rack_pk(Pattern+KNN, selector ON) is not the average of Pattern and KNN: it recovers most of KNN's damage loss (+7.7 vs KNN-alone's −39.3) but still lands nominally below Pattern on wins (−0.09).rack_pt(Pattern+TMHorizon) is statistically indistinguishable fromtmhorizon-alone — adding Pattern to the rack changes nothing, i.e. the selector picks the corrector essentially always. The selector's 16-gun negative verdict reproduces at a 2-gun rack: it does not manufacture a win it did not have. - The only signal is a damage-side hint on the two Pattern-lead correctors
(
bitbrain+10.3,tmhorizon+8.7, both 11/15 opponents). It is below the MDE (12.0 / 13.8) and not significant by the pre-registered sign test (p=0.1185), so it is not a win and not a null. Note the sign-flip test on the mean is nominally significant forbitbrain(p=0.032) while the sign test is not — the positive deltas are larger than the negatives — but with 5 arms compared this is not compelling after multiplicity, and the effect is under the MDE. This is exactly the "damage without wins" pattern this project keeps paying for (theringmover: +31 dmg/run and fewer wins); here the win deltas (+0.11/+0.13) are themselves sub-MDE. - The reference is stable:
patternunderTR_MOVEMENT=strafemeasures 99.6 dmg/run and 50.4% round wins here, matching the movement campaign'sstrafereference in Batch 4 (103.5 dmg/run, 52.6%) — same movement, same panel, different job.
Pre-registered predictions — scorecard (an honest count)
| # | prediction | outcome |
|---|---|---|
| 1 | no arm beats pattern on both primaries, point estimates inside the MDE; onlyPattern confirmed |
CORRECT |
| 2 | bitbrain is a wash vs pattern (gauntlet: −3.6 dmg, ≈0 wins) |
CORRECT on the verdict, WRONG on the point estimate: +10.3 dmg here vs −3.6 in the 32-opponent gauntlet; both inside their MDEs. Direction differs, verdict (wash) holds |
| 3 | tmhorizon is WORSE than pattern (its own docs predict it loses) |
WRONG: it is nominally better on both primaries (+8.7 dmg, +0.13 wins), though not distinguishable |
| 4 | knn is WORSE than pattern on damage, not better on wins |
CORRECT (detectably on both) |
| 5 | the small racks do not beat pattern; rack_pk lands closer to Pattern than knn-alone |
CORRECT |
| 6 | any surprise is a rack arm on damage without wins | PARTLY CORRECT: the nominal damage edge is on the single-gun correctors and the racks; every one of them is win-flat |
Batch 2 — wider-panel calibration of the damage-side hint
The Batch-1 answer is decisive on the primary question (nothing beats pattern)
but the two corrector arms carry a sub-MDE damage hint that rule 3 forbids
calling either way. Resolving it needs more opponents, not more runs, so this
batch re-runs exactly those two arms plus the reference on a wider panel.
Design. One frozen binary, three env-only arms, panel
tools/ab/panel_gun_b2.txt = the j117 32-opponent legacy roster minus Aurora
(inert in j117: one fire in 6 battles — it cannot reveal a gun regression)
plus SpinBot = 33 opponents (a superset of Batch 1's 15). 3 runs × 3
rounds = 297 battles, --conc 6, --wait-arena. Arm file:
tools/ab/arms_gun_b2.txt. Reference: pattern. Movement pinned
TR_MOVEMENT=strafe in every arm (same rationale as Batch 1).
| # | arm | env | role |
|---|---|---|---|
| 1 | pattern |
TR_MOVEMENT=strafe |
reference |
| 2 | bitbrain |
+ TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay |
largest Batch-1 damage edge (+10.3) |
| 3 | tmhorizon |
+ TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both |
second Batch-1 damage edge (+8.7) |
Pre-registered prediction (written BEFORE the battles finished):
- The
+10dmg/run hint does not survive the wider panel:bitbrainandtmhorizondamage deltas fall toward zero and stay inside the new (smaller) MDE; the per-opponent sign test remains non-significant. Rationale: both arms come from the same Pattern-lead-correction base, and the prior 32-opponent gauntlet measuredbitbrainat −3.6 dmg/run on a largely overlapping roster; +10 on 15 opponents is plausibly a small-panel fluctuation. - Round wins remain flat for both arms — if anything they regress toward zero.
- If instead the damage hint survives at a significant cross-opponent sign
test with wins not down, it is recorded as the first gun to beat
patternby rule 2 — and the campaign immediately looks for what the corrector is exploiting. That would be a genuine up-set, and it would not be spun as a null.
Outcome — Batch 2
(filled in by the Batch-2 results commit)
What to try next
(Rewritten after Batch 1 — these are recommendations, not results. Filled in the results commit below.)
What would make us stop
- Stop the first phase once a batch's best arm cannot beat
patternbeyond the MDE, or when a gun arm's damage gain is bought with a detectable win loss. At that point "onlyPatternis the measured optimum of this rack" is the conclusion, not a failure (§2 rule 4). - Stop a single batch early only for a contract violation (arena not free, liveness FAIL, non-zero exit rate) — never because the numbers look boring.
- Do not invent more arms on the same axis once two consecutive batches fail
to improve on
patternbeyond the MDE; move to a named new axis instead (candidate list in "What to try next").
How to run a batch (exact commands)
# 0. wait for the arena (j120's movement run may still be fighting)
TOURNAMENT_NIMCACHE=/tmp/nc_j121 \
tools/ab/tournament_run.sh \
--arms tools/ab/arms_gun_b1.txt \
--panel tools/ab/panel_movement.txt \
--runs 3 --rounds 3 --conc 6 --wait-arena 45 \
--reference pattern \
--outdir /tmp/ab/j121_g1
# 1. the paired per-opponent table, sign tests, MDE and the pre-registered verdict
python3 tools/ab/tournament_analyze.py /tmp/ab/j121_g1 --reference pattern
--reference may be ANY arm of the session: re-analyzing an old session with a
different reference is a free pairwise comparison with no battles.
Session log (outdirs are in /tmp and are NOT committed)
| session | commit | battles | arms | verdict |
|---|---|---|---|---|
/tmp/ab/j121_g1 |
(recorded at run time) | 270 | pattern, bitbrain, tmhorizon, knn, rack_pk, rack_pt | (Batch 1 outcome below) |