Files
SirRoboGarage/docs/gun_campaign.md
T

29 KiB
Raw Blame History

Gun campaign — ledger

Goal (owner's mandate, 2026-09-26 overnight): the movement campaign produced a replicated champion (TR_MOVEMENT=strafe, docs/movement_campaign.md) which a parallel job is shipping. "Continue until you found an amazing movement. When found do the same over for a gun." This file is the GUN campaign's single source of truth: later jobs append a ## Batch N section and never edit an earlier one (a wrong earlier number gets a correction line, not a rewrite). This file is the gun successor to docs/movement_campaign.md; read that first, then this.


0. The question

The shipped rack admits Pattern only (onlyPattern). That decision was taken because the virtual-fitness selector measured negative value at every rack size tested, and Pattern is the best single gun by the DrussGT hit-rate table (docs/selector_negative_value.md, docs/gun_rack_analysis.md, commit e0666a5). But every one of those measurements — like every pre-campaign movement claim — is DrussGT-heavy, and the movement campaign's standing lesson is that a one-opponent result is not a result.

So Batch 1 asks the direct question, on a frozen 15-opponent panel:

Does any rack gun, or any small gun configuration, beat the shipped Pattern on damage/run AND round wins across 15 opponents — or is onlyPattern CONFIRMED rather than merely assumed?

A null is a successful outcome here: it upgrades onlyPattern from an assumption to a measured verdict across the panel, which is itself valuable.


1. Protocol (identical to the movement campaign; the instrument is reused verbatim)

tools/ab/tournament_run.sh + tools/ab/tournament_analyze.py + tools/ab/panel_movement.txt (frozen panel), reused unchanged. The unit of evidence is the number of opponents, not the number of runs.

Element Rule
Subject ONE frozen binary built from git archive HEAD (tournament_run.sh records commit + binary sha256 in session.json)
Arms env dicts only — no per-arm rebuild, ever; the arm file is a committed file (tools/ab/arms_gun_b1.txt)
Panel the frozen tools/ab/panel_movement.txt (15 opponents). Adding/removing an opponent starts a new batch number
Pairing per opponent: average the arm's runs, subtract the reference's average for that same opponent → one delta per opponent; then aggregate
Isolation per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir
Serialization one battle fleet at a time via --wait-arena; never a broad pkill robocode_shim
Liveness every declared env token must appear verbatim in OUR bot's own raw-env report, else the run is excluded and named
Never shipped this is a measurement campaign: git status clean, shipped defaults untouched, .gitignore untouched

Movement-default flip note (MEASURED). During this job a parallel job (j120, commit 3fd6db9) flipped the shipped TR_MOVEMENT default from tfil to strafe. Batch 1 built before the flip, Batch 2 after. Every arm in both batches declares TR_MOVEMENT=strafe, so the effective movement is identical and the two batches are directly comparable — the pin did its job. The reference rack lines in the boot reports read PATTERN for pattern, and the active 1v1 rack reads BITBRAIN / TMHORIZON / KNN for the overridden arms.

Movement is PINNED in every arm: TR_MOVEMENT=strafe

Reason (hard rule). A movement-default flip is landing from job j120 during this run, so the default engine can change underneath us. Every arm therefore sets TR_MOVEMENT=strafe explicitly. This (a) removes movement as a confound, (b) makes all arms share exactly one movement, so every delta is a pure gun delta, and (c) satisfies the analyzer's liveness rule for every arm (each declared token must appear verbatim; an undeclared leaked TR_MOVEMENT is fatal, but here every arm declares it). The reference arm is therefore the shipped gun rack under the pinned movement, not under the (mutable) default.


2. Pre-registered decision rules (fixed BEFORE Batch 1 ran)

  1. Primary metrics: damage/run and ROUND WINS. Secondary/explanation only: damage taken/run, incoming hit rate, mean distance. Never hit rate alone — that trap has inverted six verdicts in this project.
  2. BETTER than the reference iff one primary metric is up with a cross-opponent sign test p < 0.05 while the other does not go down; the mirror image for WORSE. Anything else is NOT DISTINGUISHABLE (a real answer, not a failure). Both the strict reading (the other metric's mean delta >= 0) and the substantive reading (not detectably down: not significant and smaller than that metric's MDE) are printed.
  3. A verdict must survive the between-opponent spread: the pooled mean delta is reported with the SD across opponents, its SE, a 95% CI, and the MDE (α=0.05 two-sided, 80% power). An effect smaller than the MDE is reported as not detectable — never as absent, never as a win.
  4. Somewhere to stop: if no arm beats the shipped pattern by rule 2 in Batch 1 and no arm shows a >= +MDE damage gain with p<0.10, then the gun stage's first phase is closed with "the shipped onlyPattern rack is the best gun configuration we have measured across the panel" — that is a successful outcome, and the campaign moves to a named next axis rather than inventing more arms. See What would make us stop.
  5. No promotion off a single metric, a single opponent, or a single run. A change that wins damage by losing wins (or vice-versa) is not a win.
  6. Every batch is shot with a pre-registered prediction stated in its section before the battles finish; a prediction that turns out wrong is recorded as wrong.

3. Stage 0 — what we already know (given, not re-derived)

fact value source
shipped rack onlyPattern (id 5), all other guns off common_libs/gun_harness/selector.nim
why selector measured negative value at every rack size (16, lean8, lean6, pairPC/PK/PL) and 10/10 adversaries docs/selector_negative_value.md, docs/gun_rack_analysis.md
Pattern hit rate vs DrussGT ~10.8% given
KNN / Linear / Circular / WallBounce / GF vs DrussGT 5.6 / 3.0 / 2.9 / 2.7 / 2.1 % given
BitBrain when idle == Pattern (30 runs/arm: 97/210 vs 97/210 wins) docs/bitbrain_vs_tmhorizon_ab.md
BitBrain across 32 legacy opponents neutral: -3.6 dmg/run, +0.02 wins/run, p=0.53/0.91 docs/gauntlet_bitbrain_vs_pattern.md
lead amplitude dead axis — full gain sweep {0, 0.25-1.0, 1.0, 1.25, 1.5, 2.0} run, nothing beats Pattern docs/bitbrain_campaign.md
offline prediction checks veto-only — may kill a design, never select one docs/offline_harness_trust.md
spinner-specific claim UNTESTED: no constant-turn spinner exists in the legacy roster (a note for a later job, not built here unless nearly free) docs/gauntlet_bitbrain_vs_pattern.md

Movement context (why the pin): the movement champion is TR_MOVEMENT=strafe at its current defaults (range 325, tol 25, tilt 15/0.10); it replicated its win over the shipped tfil in four independent sessions (Batch 4: 52.6% vs 38.5% round-win rate). Movement is held constant at that champion for every gun arm.


Batch 1 — does anything beat the shipped Pattern?

Design. One frozen binary, six env-only arms, one frozen panel (tools/ab/panel_movement.txt, 15 opponents: 5 dodger, 3 pattern, 1 wall-follower, 1 corner-camper, 1 spinner, 1 rammer, 1 brawler, 2 aggressive megas), 3 runs × 3 rounds per (opponent, arm) = 270 battles. Arm file: tools/ab/arms_gun_b1.txt. Reference: pattern.

# arm env what it isolates
1 pattern TR_MOVEMENT=strafe the arm to beat (shipped onlyPattern rack)
2 bitbrain + TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay BitBrain-only, learned gain (panel replication of the 32-opp gauntlet at the pinned-movement standard)
3 tmhorizon + TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both TMHorizon-only corrector (predicted to lose; never panel-tested)
4 knn + TR_RACK_PATTERN=off TR_RACK_KNN=both best non-Pattern single gun (never panel-tested)
5 rack_pk + TR_RACK_KNN=both Pattern + KNN, selector ON (different-family hedge)
6 rack_pt + TR_RACK_TMHORIZON=both Pattern + TMHorizon, selector ON (corrector-family hedge)

(+ = the pinned TR_MOVEMENT=strafe is present in every arm; see §1.)

Pre-registered prediction (written BEFORE the battles finished):

  1. No arm beats pattern on BOTH primaries, and the point estimates sit inside the MDE. onlyPattern is confirmed across the panel.
  2. bitbrain is a wash vs pattern (replicating the 32-opponent gauntlet: ±3.6 dmg/run, ≈0 wins) — the DrussGT-only gain-config penalty does not carry.
  3. tmhorizon is WORSE than pattern (its own offline work predicts it needs ~80% side accuracy and can reach ~60%); no damage win, no win win.
  4. knn is WORSE than pattern on damage (its DrussGT hit rate is roughly half Pattern's) and not better on wins.
  5. The two small racks do not beat pattern; rack_pk lands closer to Pattern than knn-alone does (the selector at least partially hedges back), but the selector's poor ranking keeps it below the reference.
  6. If any surprise exists, it is a rack arm on damage without wins — the ring-mover mirror-image trap — and it will not be read as a win.

Outcome — direct answer

MEASURED. Session /tmp/ab/j121_g1, commit 1d8143a15e4a038c39dfc2153ff518fa003b8612 (the pre-registration commit), frozen binary sha256 aa49a45fec20…, 15 opponents × 6 arms × 3 runs × 3 rounds = 270 battles, 0 failed, 0 never started, 0 liveness exclusions. Every arm declared TR_MOVEMENT=strafe, verified in each bot's own raw-env report; the reference rack line reads PATTERN for pattern and the overridden rack reads BITBRAIN / TMHORIZON / KNN for the others.

DIRECT ANSWER (MEASURED, n=15 opponents): NO — nothing beats the shipped Pattern across the panel on damage/run AND round wins. Every arm except knn is NOT DISTINGUISHABLE from pattern under both the strict and the substantive reading of the pre-registered rule. The best nominal challenger, tmhorizon, is +0.13 wins/run (MDE 0.36) and +8.7 dmg/run (MDE 13.8) — an effect ~3× smaller than the design can detect — and bitbrain is +0.11 wins / +10.3 dmg, also inside the MDE. knn is detectably WORSE on both primaries. The shipped onlyPattern rack is therefore CONFIRMED rather than merely assumed, at the resolution of this panel. That is a real result, not a failure.

Pooled dashboard (all valid runs — explanation only, NOT the verdict)

arm runs dmg/run dmg taken/run wins/run round wins win rate incoming hit rate mean distance
pattern (REF) 45 99.6 155.9 1.51 68/135 50.4% 12.84% 435
bitbrain 45 109.9 151.2 1.62 73/135 54.1% 12.87% 427
tmhorizon 45 108.3 153.7 1.64 74/135 54.8% 12.53% 428
knn 45 60.3 175.7 0.89 40/135 29.6% 13.31% 430
rack_pk 45 107.3 161.3 1.42 64/135 47.4% 13.65% 432
rack_pt 45 105.3 149.9 1.64 74/135 54.8% 12.90% 424

Cross-opponent aggregation (the verdict layer, damage + wins)

arm metric mean Δ spread (SD) 95% CI sign test (wins/n) p(sign) p(sign-flip) Wilcoxon p MDE
bitbrain damage +10.34 16.65 [+1.12, +19.57] 11/15 0.1185 0.03247 0.05006 12.05
bitbrain wins +0.11 0.48 [-0.16, +0.38] 6/9 0.5078 0.4922 0.4764 0.35
tmhorizon damage +8.71 19.10 [-1.86, +19.29] 11/15 0.1185 0.09918 0.1183 13.82
tmhorizon wins +0.13 0.50 [-0.14, +0.41] 5/9 1.0 0.4062 0.3118 0.36
knn damage −39.33 23.15 [−52.15, −26.51] 0/15 6.1e-5 6.1e-5 0.0007 16.75
knn wins −0.62 0.59 [−0.95, −0.30] 1/11 0.01172 0.001953 0.004948 0.43
rack_pk damage +7.68 32.09 [−10.10, +25.45] 9/15 0.6072 0.423 0.5895 23.21
rack_pk wins −0.09 0.60 [−0.42, +0.24] 6/11 1.0 0.6738 0.5932 0.43
rack_pt damage +5.69 16.50 [−3.45, +14.83] 10/15 0.3018 0.2026 0.2681 11.93
rack_pt wins +0.13 0.41 [−0.10, +0.36] 6/10 0.7539 0.3242 0.1997 0.30

The pre-registered verdict (verbatim)

rank arm Δwins/run Δdmg/run sign test wins sign test dmg verdict (strict) verdict (substantive)
1 tmhorizon +0.13 +8.7 5/9 p=1 11/15 p=0.1185 not distinguishable not distinguishable
2 rack_pt +0.13 +5.7 6/10 p=0.7539 10/15 p=0.3018 not distinguishable not distinguishable
3 bitbrain +0.11 +10.3 6/9 p=0.5078 11/15 p=0.1185 not distinguishable not distinguishable
4 rack_pk −0.09 +7.7 6/11 p=1 9/15 p=0.6072 not distinguishable not distinguishable
5 knn −0.62 −39.3 1/11 p=0.01172 0/15 p=6.104e-05 WORSE not distinguishable

Reference pattern: 99.6 dmg/run, 1.51 wins/run, 12.84% incoming, 435 px.

Reading (MEASURED, with the mechanism)

  • knn is a clean, large loss and is dropped. It deals −39 dmg/run and loses −0.62 wins/run, positive on 0/15 opponents on damage and 1/11 on wins (p=6e-5 / p=0.012). Its incoming hit rate is no worse (13.31% vs 12.84%) — it simply cannot aim: the same movement, a worse gun.
  • The two small racks do not beat pattern. rack_pk (Pattern+KNN, selector ON) is not the average of Pattern and KNN: it recovers most of KNN's damage loss (+7.7 vs KNN-alone's −39.3) but still lands nominally below Pattern on wins (−0.09). rack_pt (Pattern+TMHorizon) is statistically indistinguishable from tmhorizon-alone — adding Pattern to the rack changes nothing, i.e. the selector picks the corrector essentially always. The selector's 16-gun negative verdict reproduces at a 2-gun rack: it does not manufacture a win it did not have.
  • The only signal is a damage-side hint on the two Pattern-lead correctors (bitbrain +10.3, tmhorizon +8.7, both 11/15 opponents). It is below the MDE (12.0 / 13.8) and not significant by the pre-registered sign test (p=0.1185), so it is not a win and not a null. Note the sign-flip test on the mean is nominally significant for bitbrain (p=0.032) while the sign test is not — the positive deltas are larger than the negatives — but with 5 arms compared this is not compelling after multiplicity, and the effect is under the MDE. This is exactly the "damage without wins" pattern this project keeps paying for (the ring mover: +31 dmg/run and fewer wins); here the win deltas (+0.11/+0.13) are themselves sub-MDE.
  • The reference is stable: pattern under TR_MOVEMENT=strafe measures 99.6 dmg/run and 50.4% round wins here, matching the movement campaign's strafe reference in Batch 4 (103.5 dmg/run, 52.6%) — same movement, same panel, different job.

Pre-registered predictions — scorecard (an honest count)

# prediction outcome
1 no arm beats pattern on both primaries, point estimates inside the MDE; onlyPattern confirmed CORRECT
2 bitbrain is a wash vs pattern (gauntlet: −3.6 dmg, ≈0 wins) CORRECT on the verdict, WRONG on the point estimate: +10.3 dmg here vs −3.6 in the 32-opponent gauntlet; both inside their MDEs. Direction differs, verdict (wash) holds
3 tmhorizon is WORSE than pattern (its own docs predict it loses) WRONG: it is nominally better on both primaries (+8.7 dmg, +0.13 wins), though not distinguishable
4 knn is WORSE than pattern on damage, not better on wins CORRECT (detectably on both)
5 the small racks do not beat pattern; rack_pk lands closer to Pattern than knn-alone CORRECT
6 any surprise is a rack arm on damage without wins PARTLY CORRECT: the nominal damage edge is on the single-gun correctors and the racks; every one of them is win-flat

Batch 2 — wider-panel calibration of the damage-side hint

The Batch-1 answer is decisive on the primary question (nothing beats pattern) but the two corrector arms carry a sub-MDE damage hint that rule 3 forbids calling either way. Resolving it needs more opponents, not more runs, so this batch re-runs exactly those two arms plus the reference on a wider panel.

Design. One frozen binary, three env-only arms, panel tools/ab/panel_gun_b2.txt = the j117 32-opponent legacy roster minus Aurora (inert in j117: one fire in 6 battles — it cannot reveal a gun regression) plus SpinBot = 33 opponents (a superset of Batch 1's 15). 3 runs × 3 rounds = 297 battles, --conc 6, --wait-arena. Arm file: tools/ab/arms_gun_b2.txt. Reference: pattern. Movement pinned TR_MOVEMENT=strafe in every arm (same rationale as Batch 1).

# arm env role
1 pattern TR_MOVEMENT=strafe reference
2 bitbrain + TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay largest Batch-1 damage edge (+10.3)
3 tmhorizon + TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both second Batch-1 damage edge (+8.7)

Pre-registered prediction (written BEFORE the battles finished):

  1. The +10 dmg/run hint does not survive the wider panel: bitbrain and tmhorizon damage deltas fall toward zero and stay inside the new (smaller) MDE; the per-opponent sign test remains non-significant. Rationale: both arms come from the same Pattern-lead-correction base, and the prior 32-opponent gauntlet measured bitbrain at −3.6 dmg/run on a largely overlapping roster; +10 on 15 opponents is plausibly a small-panel fluctuation.
  2. Round wins remain flat for both arms — if anything they regress toward zero.
  3. If instead the damage hint survives at a significant cross-opponent sign test with wins not down, it is recorded as the first gun to beat pattern by rule 2 — and the campaign immediately looks for what the corrector is exploiting. That would be a genuine up-set, and it would not be spun as a null.

Outcome — Batch 2

MEASURED. Session /tmp/ab/j121_g2, commit c343c00aaf800bc05ce9bf1b277054c11f5f577c, frozen binary sha256 d267ab78f6ce…, 33 opponents × 3 arms × 3 runs × 3 rounds = 297 battles, 0 failed, 0 never started, 0 liveness exclusions.

A movement-default flip landed between the two batches (commit 3fd6db9, TR_MOVEMENT default tfil → strafe; the only source change). Because every arm in both batches pins TR_MOVEMENT=strafe, the effective movement is identical in Batch 1 and Batch 2 — the batches are directly comparable and the gun comparison is unconfounded. This is exactly what the pin was for.

DIRECT ANSWER (MEASURED, n=33 opponents): the Batch-1 damage hint did NOT survive. bitbrain collapses to +1.8 dmg/run (19/33 opponents, sign test p=0.49) with +0.11 wins/run — a wash. tmhorizon reverses sign and is detectably WORSE on damage: −10.4 dmg/run (10/33 opponents, sign test p=0.035, sign-flip p=0.0047, 95% CI [−17.2, −3.6]) at unchanged wins (−0.03). The only open signal in Batch 1 was a small-panel fluctuation, and the wider panel resolves it against the correctors. Combined with Batch 1: nothing beats the shipped Pattern on damage/run AND round wins; onlyPattern is CONFIRMED on 15 opponents and again on 33.

Pooled dashboard (all valid runs — explanation only, NOT the verdict)

arm runs dmg/run dmg taken/run wins/run round wins win rate incoming hit rate mean distance
pattern (REF) 99 131.2 126.4 1.90 188/297 63.3% 12.93% 417
bitbrain 99 133.0 126.7 2.01 199/297 67.0% 12.91% 417
tmhorizon 99 120.8 133.2 1.87 185/297 62.3% 13.04% 422

Cross-opponent aggregation (the verdict layer, damage + wins)

arm metric mean Δ spread (SD) 95% CI sign test (wins/n) p(sign) p(sign-flip) Wilcoxon p MDE
bitbrain damage +1.80 18.95 [−4.67, +8.26] 19/33 0.4869 0.5914 0.5675 9.24
bitbrain wins +0.11 0.48 [−0.05, +0.28] 12/19 0.3593 0.2436 0.5716 0.24
tmhorizon damage −10.40 20.01 [−17.22, −3.57] 10/33 0.03508 0.00472 0.006975 9.76
tmhorizon wins −0.03 0.59 [−0.23, +0.17] 9/19 1.0 0.8459 0.4804 0.29

The pre-registered verdict (verbatim)

rank arm Δwins/run Δdmg/run sign test wins sign test dmg verdict (strict) verdict (substantive)
1 bitbrain +0.11 +1.8 12/19 p=0.3593 19/33 p=0.4869 not distinguishable not distinguishable
2 tmhorizon −0.03 −10.4 9/19 p=1 10/33 p=0.03508 WORSE WORSE

Reference pattern: 131.2 dmg/run, 1.90 wins/run, 12.93% incoming, 417 px.

Reading

  • BitBrain is a wash — now measured three times. vs the shipped Pattern: −3.6 dmg/run on the 32-opponent legacy gauntlet (with tfil movement, docs/gauntlet_bitbrain_vs_pattern.md), +10.3 on the 15-opponent Batch-1 panel, +1.8 here on 33 opponents. All three sit inside their own MDEs; the pooled evidence is a zero. The Batch-1 +10 was a 15-opponent fluctuation.
  • TMHorizon is the one arm the wider panel separates — and it loses. Its Batch-1 sign (+8.7) does not merely fail to replicate, it inverts (−10.4, 10/33, p=0.035). The damage loss is broad (24/33 opponents negative, worst Aristocles −61, Jen −54, WallAvoider −41), so it is not one outlier. Since TMHorizon was already predicted to lose by its own design docs, this is the first confirming live panel measurement of that prediction.
  • Wins are flat for both arms (+0.11 / −0.03, MDEs 0.24 / 0.29): the correctors change damage at most, never survival. In this harness round wins are survival wins, so neither is a win.
  • The corrector family is closed as a source of a pattern-beating gun at the resolutions tested.

Pre-registered predictions — scorecard

# prediction outcome
1 the +10 dmg hint does not survive the wider panel; deltas fall inside the new MDE; sign test non-significant CORRECT (bitbrain +1.8, p=0.49; tmhorizon even reversed to −10.4, p=0.035)
2 wins remain flat CORRECT (+0.11 / −0.03)
3 if the hint survived at p<0.05 with wins not down, it is the first gun to beat pattern not invoked

Final direct answer for the phase (MEASURED). Across two frozen panels (15 and 33 opponents), a single frozen-binary-per-session design, 567 battles, six distinct gun configurations (five challengers + the reference), and movement pinned identically in every arm: no rack gun and no small rack beats the shipped Pattern on damage/run AND round wins. Every candidate is either not distinguishable (bitbrain, tmhorizon, rack_pk, rack_pt at n=15), a detectable loss (knn; tmhorizon at n=33), or a sub-MDE damage-only wobble (bitbrain). The shipped onlyPattern rack is CONFIRMED, not merely assumed — the gun axis's first phase ends in a measured optimum of the tested design space.


What to try next (rewritten AFTER Batches 1–2 — recommendations, not results)

The primary question is answered: nothing beats the shipped pattern across the frozen panel, and onlyPattern is CONFIRMED. Ranked by value per battle:

  1. The damage-side hint is now CLOSED. Batch 2 (33 opponents) killed it: bitbrain is a wash (+1.8 dmg/run, p=0.49) and tmhorizon is detectably worse (−10.4 dmg/run, p=0.035). Do not spend another batch on the Pattern-lead correctors; three independent measurements of bitbrain vs Pattern now average ~0 (−3.6 / +10.3 / +1.8 dmg/run, all inside their MDEs).
  2. Lead INFORMATION remains the only named gun axis (given: amplitude is dead). But the two live implementations of it (BitBrain's gain, TMHorizon's shift) are both measured negative/neutral across panels. A future axis would need a different information source, not another knob on these two — e.g. a new base prediction, or a gun-agnostic ensemble. Do not re-open the correctors.
  3. The selector is confirmed negative at a 2-gun rack (rack_pt == tmhorizon; rack_pk below pattern on wins). Do not re-open it with a larger rack: the 16-gun negative already stands and Batch 1 shows the mechanism does not manufacture a win.
  4. Drop knn (detectably worse on both primaries, 0/15 on damage).
  5. Never chase damage without wins. This harness's round wins are survival wins, so the gun's job is to kill; a damage gain that does not raise the win rate is the ring-mover trap in a new costume (and tmhorizon's n=33 result is a negative damage signal with flat wins — also not a win).
  6. Spinner-specific claim (owner's) remains UNTESTED. The legacy roster has no constant-turn spinner; SpinBot is a periodic circle-mover and is in both panels (Batch 1: bitbrain +25.3 dmg / +0 wins, tmhorizon +15.2 / +0 wins; on the 33-panel SpinBot's deltas are in the per-opponent tables). Building a true constant-turn spinner opponent is a nearly-free follow-up and is the one part of the owner's claim this campaign has not touched.

What would make us stop

PHASE STATUS (updated after Batch 2): STOPPED — the condition fired. Across Batches 1–2 no arm beats pattern beyond the MDE, and no arm shows a >= +MDE damage gain with p<0.10 (bitbrain +1.8 vs MDE 9.24; tmhorizon actually negative). The gun stage's first phase is therefore closed with "the shipped onlyPattern rack is the best gun configuration we have measured across the panels (15 and 33 opponents)" — a successful outcome, not a failure. Further gun work must open a named new axis (see What to try next), not another arm on the same axis.

  • Stop the first phase once a batch's best arm cannot beat pattern beyond the MDE, or when a gun arm's damage gain is bought with a detectable win loss. At that point "onlyPattern is the measured optimum of this rack" is the conclusion, not a failure (§2 rule 4). This fired after Batch 2.
  • Stop a single batch early only for a contract violation (arena not free, liveness FAIL, non-zero exit rate) — never because the numbers look boring.
  • Do not invent more arms on the same axis once two consecutive batches fail to improve on pattern beyond the MDE; move to a named new axis instead (candidate list in "What to try next").

How to run a batch (exact commands)

# 0. wait for the arena (other campaign jobs may be fighting)
TOURNAMENT_NIMCACHE=/tmp/nc_j121 \
tools/ab/tournament_run.sh \
    --arms    tools/ab/arms_gun_b1.txt \
    --panel   tools/ab/panel_movement.txt \
    --runs 3 --rounds 3 --conc 6 --wait-arena 45 \
    --reference pattern \
    --outdir /tmp/ab/j121_g1

# Batch 2 (the wider-panel calibration) was the same command with
#   --arms tools/ab/arms_gun_b2.txt --panel tools/ab/panel_gun_b2.txt \
#   --outdir /tmp/ab/j121_g2 --wait-arena 50

# 1. the paired per-opponent table, sign tests, MDE and the pre-registered verdict
python3 tools/ab/tournament_analyze.py /tmp/ab/j121_g1 --reference pattern

--reference may be ANY arm of the session: re-analyzing an old session with a different reference is a free pairwise comparison with no battles.

Session log (outdirs are in /tmp and are NOT committed)

session commit battles arms verdict
/tmp/ab/j121_g1 1d8143a 270 (0 failed, 0 never started, 0 excluded) pattern, bitbrain, tmhorizon, knn, rack_pk, rack_pt nothing beats the shipped pattern; all arms not distinguishable except knn = WORSE; onlyPattern CONFIRMED
/tmp/ab/j121_g2 c343c00 297 (0 failed, 0 never started, 0 excluded) pattern, bitbrain, tmhorizon the Batch-1 damage hint does not survive: bitbrain +1.8 dmg/run p=0.49 (wash); tmhorizon −10.4 dmg/run p=0.035 (WORSE); wins flat