29 KiB
Gun campaign — ledger
Goal (owner's mandate, 2026-09-26 overnight): the movement campaign produced
a replicated champion (TR_MOVEMENT=strafe, docs/movement_campaign.md) which a
parallel job is shipping. "Continue until you found an amazing movement. When
found do the same over for a gun." This file is the GUN campaign's single source
of truth: later jobs append a ## Batch N section and never edit an earlier
one (a wrong earlier number gets a correction line, not a rewrite). This file is
the gun successor to docs/movement_campaign.md; read that first, then this.
0. The question
The shipped rack admits Pattern only (onlyPattern). That decision was
taken because the virtual-fitness selector measured negative value at every
rack size tested, and Pattern is the best single gun by the DrussGT hit-rate
table (docs/selector_negative_value.md, docs/gun_rack_analysis.md, commit
e0666a5). But every one of those measurements — like every pre-campaign
movement claim — is DrussGT-heavy, and the movement campaign's standing
lesson is that a one-opponent result is not a result.
So Batch 1 asks the direct question, on a frozen 15-opponent panel:
Does any rack gun, or any small gun configuration, beat the shipped
Patternon damage/run AND round wins across 15 opponents — or isonlyPatternCONFIRMED rather than merely assumed?
A null is a successful outcome here: it upgrades onlyPattern from an
assumption to a measured verdict across the panel, which is itself valuable.
1. Protocol (identical to the movement campaign; the instrument is reused verbatim)
tools/ab/tournament_run.sh + tools/ab/tournament_analyze.py +
tools/ab/panel_movement.txt (frozen panel), reused unchanged. The unit of
evidence is the number of opponents, not the number of runs.
| Element | Rule |
|---|---|
| Subject | ONE frozen binary built from git archive HEAD (tournament_run.sh records commit + binary sha256 in session.json) |
| Arms | env dicts only — no per-arm rebuild, ever; the arm file is a committed file (tools/ab/arms_gun_b1.txt) |
| Panel | the frozen tools/ab/panel_movement.txt (15 opponents). Adding/removing an opponent starts a new batch number |
| Pairing | per opponent: average the arm's runs, subtract the reference's average for that same opponent → one delta per opponent; then aggregate |
| Isolation | per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir |
| Serialization | one battle fleet at a time via --wait-arena; never a broad pkill robocode_shim |
| Liveness | every declared env token must appear verbatim in OUR bot's own raw-env report, else the run is excluded and named |
| Never shipped | this is a measurement campaign: git status clean, shipped defaults untouched, .gitignore untouched |
Movement-default flip note (MEASURED). During this job a parallel job (j120, commit
3fd6db9) flipped the shippedTR_MOVEMENTdefault fromtfiltostrafe. Batch 1 built before the flip, Batch 2 after. Every arm in both batches declaresTR_MOVEMENT=strafe, so the effective movement is identical and the two batches are directly comparable — the pin did its job. The reference rack lines in the boot reports readPATTERNforpattern, and the active 1v1 rack readsBITBRAIN/TMHORIZON/KNNfor the overridden arms.
Movement is PINNED in every arm: TR_MOVEMENT=strafe
Reason (hard rule). A movement-default flip is landing from job j120 during
this run, so the default engine can change underneath us. Every arm therefore
sets TR_MOVEMENT=strafe explicitly. This (a) removes movement as a
confound, (b) makes all arms share exactly one movement, so every delta is a pure
gun delta, and (c) satisfies the analyzer's liveness rule for every arm
(each declared token must appear verbatim; an undeclared leaked TR_MOVEMENT
is fatal, but here every arm declares it). The reference arm is therefore the
shipped gun rack under the pinned movement, not under the (mutable) default.
2. Pre-registered decision rules (fixed BEFORE Batch 1 ran)
- Primary metrics: damage/run and ROUND WINS. Secondary/explanation only: damage taken/run, incoming hit rate, mean distance. Never hit rate alone — that trap has inverted six verdicts in this project.
- BETTER than the reference iff one primary metric is up with a
cross-opponent sign test p < 0.05 while the other does not go down;
the mirror image for WORSE. Anything else is NOT DISTINGUISHABLE (a
real answer, not a failure). Both the strict reading (the other metric's mean
delta
>= 0) and the substantive reading (not detectably down: not significant and smaller than that metric's MDE) are printed. - A verdict must survive the between-opponent spread: the pooled mean delta is reported with the SD across opponents, its SE, a 95% CI, and the MDE (α=0.05 two-sided, 80% power). An effect smaller than the MDE is reported as not detectable — never as absent, never as a win.
- Somewhere to stop: if no arm beats the shipped
patternby rule 2 in Batch 1 and no arm shows a>= +MDEdamage gain with p<0.10, then the gun stage's first phase is closed with "the shippedonlyPatternrack is the best gun configuration we have measured across the panel" — that is a successful outcome, and the campaign moves to a named next axis rather than inventing more arms. See What would make us stop. - No promotion off a single metric, a single opponent, or a single run. A change that wins damage by losing wins (or vice-versa) is not a win.
- Every batch is shot with a pre-registered prediction stated in its section before the battles finish; a prediction that turns out wrong is recorded as wrong.
3. Stage 0 — what we already know (given, not re-derived)
| fact | value | source |
|---|---|---|
| shipped rack | onlyPattern (id 5), all other guns off |
common_libs/gun_harness/selector.nim |
| why | selector measured negative value at every rack size (16, lean8, lean6, pairPC/PK/PL) and 10/10 adversaries | docs/selector_negative_value.md, docs/gun_rack_analysis.md |
| Pattern hit rate vs DrussGT | ~10.8% | given |
| KNN / Linear / Circular / WallBounce / GF vs DrussGT | 5.6 / 3.0 / 2.9 / 2.7 / 2.1 % | given |
| BitBrain when idle | == Pattern (30 runs/arm: 97/210 vs 97/210 wins) | docs/bitbrain_vs_tmhorizon_ab.md |
| BitBrain across 32 legacy opponents | neutral: -3.6 dmg/run, +0.02 wins/run, p=0.53/0.91 | docs/gauntlet_bitbrain_vs_pattern.md |
| lead amplitude | dead axis — full gain sweep {0, 0.25-1.0, 1.0, 1.25, 1.5, 2.0} run, nothing beats Pattern | docs/bitbrain_campaign.md |
| offline prediction checks | veto-only — may kill a design, never select one | docs/offline_harness_trust.md |
| spinner-specific claim | UNTESTED: no constant-turn spinner exists in the legacy roster (a note for a later job, not built here unless nearly free) | docs/gauntlet_bitbrain_vs_pattern.md |
Movement context (why the pin): the movement champion is TR_MOVEMENT=strafe
at its current defaults (range 325, tol 25, tilt 15/0.10); it replicated its win
over the shipped tfil in four independent sessions (Batch 4: 52.6% vs 38.5%
round-win rate). Movement is held constant at that champion for every gun arm.
Batch 1 — does anything beat the shipped Pattern?
Design. One frozen binary, six env-only arms, one frozen panel
(tools/ab/panel_movement.txt, 15 opponents: 5 dodger, 3 pattern, 1
wall-follower, 1 corner-camper, 1 spinner, 1 rammer, 1 brawler, 2 aggressive
megas), 3 runs × 3 rounds per (opponent, arm) = 270 battles. Arm file:
tools/ab/arms_gun_b1.txt. Reference: pattern.
| # | arm | env | what it isolates |
|---|---|---|---|
| 1 | pattern |
TR_MOVEMENT=strafe |
the arm to beat (shipped onlyPattern rack) |
| 2 | bitbrain |
+ TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay |
BitBrain-only, learned gain (panel replication of the 32-opp gauntlet at the pinned-movement standard) |
| 3 | tmhorizon |
+ TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both |
TMHorizon-only corrector (predicted to lose; never panel-tested) |
| 4 | knn |
+ TR_RACK_PATTERN=off TR_RACK_KNN=both |
best non-Pattern single gun (never panel-tested) |
| 5 | rack_pk |
+ TR_RACK_KNN=both |
Pattern + KNN, selector ON (different-family hedge) |
| 6 | rack_pt |
+ TR_RACK_TMHORIZON=both |
Pattern + TMHorizon, selector ON (corrector-family hedge) |
(+ = the pinned TR_MOVEMENT=strafe is present in every arm; see §1.)
Pre-registered prediction (written BEFORE the battles finished):
- No arm beats
patternon BOTH primaries, and the point estimates sit inside the MDE.onlyPatternis confirmed across the panel. bitbrainis a wash vspattern(replicating the 32-opponent gauntlet: ±3.6 dmg/run, ≈0 wins) — the DrussGT-only gain-config penalty does not carry.tmhorizonis WORSE thanpattern(its own offline work predicts it needs ~80% side accuracy and can reach ~60%); no damage win, no win win.knnis WORSE thanpatternon damage (its DrussGT hit rate is roughly half Pattern's) and not better on wins.- The two small racks do not beat
pattern;rack_pklands closer to Pattern thanknn-alone does (the selector at least partially hedges back), but the selector's poor ranking keeps it below the reference. - If any surprise exists, it is a rack arm on damage without wins — the
ring-mover mirror-image trap — and it will not be read as a win.
Outcome — direct answer
MEASURED. Session /tmp/ab/j121_g1, commit 1d8143a15e4a038c39dfc2153ff518fa003b8612
(the pre-registration commit), frozen binary sha256 aa49a45fec20…, 15 opponents
× 6 arms × 3 runs × 3 rounds = 270 battles, 0 failed, 0 never started, 0
liveness exclusions. Every arm declared TR_MOVEMENT=strafe, verified in each
bot's own raw-env report; the reference rack line reads PATTERN for pattern
and the overridden rack reads BITBRAIN / TMHORIZON / KNN for the others.
DIRECT ANSWER (MEASURED, n=15 opponents): NO — nothing beats the shipped
Patternacross the panel on damage/run AND round wins. Every arm exceptknnis NOT DISTINGUISHABLE frompatternunder both the strict and the substantive reading of the pre-registered rule. The best nominal challenger,tmhorizon, is +0.13 wins/run (MDE 0.36) and +8.7 dmg/run (MDE 13.8) — an effect ~3× smaller than the design can detect — andbitbrainis+0.11wins /+10.3dmg, also inside the MDE.knnis detectably WORSE on both primaries. The shippedonlyPatternrack is therefore CONFIRMED rather than merely assumed, at the resolution of this panel. That is a real result, not a failure.
Pooled dashboard (all valid runs — explanation only, NOT the verdict)
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|---|---|---|---|---|---|---|---|---|
pattern (REF) |
45 | 99.6 | 155.9 | 1.51 | 68/135 | 50.4% | 12.84% | 435 |
bitbrain |
45 | 109.9 | 151.2 | 1.62 | 73/135 | 54.1% | 12.87% | 427 |
tmhorizon |
45 | 108.3 | 153.7 | 1.64 | 74/135 | 54.8% | 12.53% | 428 |
knn |
45 | 60.3 | 175.7 | 0.89 | 40/135 | 29.6% | 13.31% | 430 |
rack_pk |
45 | 107.3 | 161.3 | 1.42 | 64/135 | 47.4% | 13.65% | 432 |
rack_pt |
45 | 105.3 | 149.9 | 1.64 | 74/135 | 54.8% | 12.90% | 424 |
Cross-opponent aggregation (the verdict layer, damage + wins)
| arm | metric | mean Δ | spread (SD) | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|---|---|---|---|---|---|---|---|---|---|
bitbrain |
damage | +10.34 | 16.65 | [+1.12, +19.57] | 11/15 | 0.1185 | 0.03247 | 0.05006 | 12.05 |
bitbrain |
wins | +0.11 | 0.48 | [-0.16, +0.38] | 6/9 | 0.5078 | 0.4922 | 0.4764 | 0.35 |
tmhorizon |
damage | +8.71 | 19.10 | [-1.86, +19.29] | 11/15 | 0.1185 | 0.09918 | 0.1183 | 13.82 |
tmhorizon |
wins | +0.13 | 0.50 | [-0.14, +0.41] | 5/9 | 1.0 | 0.4062 | 0.3118 | 0.36 |
knn |
damage | −39.33 | 23.15 | [−52.15, −26.51] | 0/15 | 6.1e-5 | 6.1e-5 | 0.0007 | 16.75 |
knn |
wins | −0.62 | 0.59 | [−0.95, −0.30] | 1/11 | 0.01172 | 0.001953 | 0.004948 | 0.43 |
rack_pk |
damage | +7.68 | 32.09 | [−10.10, +25.45] | 9/15 | 0.6072 | 0.423 | 0.5895 | 23.21 |
rack_pk |
wins | −0.09 | 0.60 | [−0.42, +0.24] | 6/11 | 1.0 | 0.6738 | 0.5932 | 0.43 |
rack_pt |
damage | +5.69 | 16.50 | [−3.45, +14.83] | 10/15 | 0.3018 | 0.2026 | 0.2681 | 11.93 |
rack_pt |
wins | +0.13 | 0.41 | [−0.10, +0.36] | 6/10 | 0.7539 | 0.3242 | 0.1997 | 0.30 |
The pre-registered verdict (verbatim)
| rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) |
|---|---|---|---|---|---|---|---|
| 1 | tmhorizon |
+0.13 | +8.7 | 5/9 p=1 | 11/15 p=0.1185 | not distinguishable | not distinguishable |
| 2 | rack_pt |
+0.13 | +5.7 | 6/10 p=0.7539 | 10/15 p=0.3018 | not distinguishable | not distinguishable |
| 3 | bitbrain |
+0.11 | +10.3 | 6/9 p=0.5078 | 11/15 p=0.1185 | not distinguishable | not distinguishable |
| 4 | rack_pk |
−0.09 | +7.7 | 6/11 p=1 | 9/15 p=0.6072 | not distinguishable | not distinguishable |
| 5 | knn |
−0.62 | −39.3 | 1/11 p=0.01172 | 0/15 p=6.104e-05 | WORSE | not distinguishable |
Reference pattern: 99.6 dmg/run, 1.51 wins/run, 12.84% incoming, 435 px.
Reading (MEASURED, with the mechanism)
knnis a clean, large loss and is dropped. It deals −39 dmg/run and loses −0.62 wins/run, positive on 0/15 opponents on damage and 1/11 on wins (p=6e-5 / p=0.012). Its incoming hit rate is no worse (13.31% vs 12.84%) — it simply cannot aim: the same movement, a worse gun.- The two small racks do not beat
pattern.rack_pk(Pattern+KNN, selector ON) is not the average of Pattern and KNN: it recovers most of KNN's damage loss (+7.7 vs KNN-alone's −39.3) but still lands nominally below Pattern on wins (−0.09).rack_pt(Pattern+TMHorizon) is statistically indistinguishable fromtmhorizon-alone — adding Pattern to the rack changes nothing, i.e. the selector picks the corrector essentially always. The selector's 16-gun negative verdict reproduces at a 2-gun rack: it does not manufacture a win it did not have. - The only signal is a damage-side hint on the two Pattern-lead correctors
(
bitbrain+10.3,tmhorizon+8.7, both 11/15 opponents). It is below the MDE (12.0 / 13.8) and not significant by the pre-registered sign test (p=0.1185), so it is not a win and not a null. Note the sign-flip test on the mean is nominally significant forbitbrain(p=0.032) while the sign test is not — the positive deltas are larger than the negatives — but with 5 arms compared this is not compelling after multiplicity, and the effect is under the MDE. This is exactly the "damage without wins" pattern this project keeps paying for (theringmover: +31 dmg/run and fewer wins); here the win deltas (+0.11/+0.13) are themselves sub-MDE. - The reference is stable:
patternunderTR_MOVEMENT=strafemeasures 99.6 dmg/run and 50.4% round wins here, matching the movement campaign'sstrafereference in Batch 4 (103.5 dmg/run, 52.6%) — same movement, same panel, different job.
Pre-registered predictions — scorecard (an honest count)
| # | prediction | outcome |
|---|---|---|
| 1 | no arm beats pattern on both primaries, point estimates inside the MDE; onlyPattern confirmed |
CORRECT |
| 2 | bitbrain is a wash vs pattern (gauntlet: −3.6 dmg, ≈0 wins) |
CORRECT on the verdict, WRONG on the point estimate: +10.3 dmg here vs −3.6 in the 32-opponent gauntlet; both inside their MDEs. Direction differs, verdict (wash) holds |
| 3 | tmhorizon is WORSE than pattern (its own docs predict it loses) |
WRONG: it is nominally better on both primaries (+8.7 dmg, +0.13 wins), though not distinguishable |
| 4 | knn is WORSE than pattern on damage, not better on wins |
CORRECT (detectably on both) |
| 5 | the small racks do not beat pattern; rack_pk lands closer to Pattern than knn-alone |
CORRECT |
| 6 | any surprise is a rack arm on damage without wins | PARTLY CORRECT: the nominal damage edge is on the single-gun correctors and the racks; every one of them is win-flat |
Batch 2 — wider-panel calibration of the damage-side hint
The Batch-1 answer is decisive on the primary question (nothing beats pattern)
but the two corrector arms carry a sub-MDE damage hint that rule 3 forbids
calling either way. Resolving it needs more opponents, not more runs, so this
batch re-runs exactly those two arms plus the reference on a wider panel.
Design. One frozen binary, three env-only arms, panel
tools/ab/panel_gun_b2.txt = the j117 32-opponent legacy roster minus Aurora
(inert in j117: one fire in 6 battles — it cannot reveal a gun regression)
plus SpinBot = 33 opponents (a superset of Batch 1's 15). 3 runs × 3
rounds = 297 battles, --conc 6, --wait-arena. Arm file:
tools/ab/arms_gun_b2.txt. Reference: pattern. Movement pinned
TR_MOVEMENT=strafe in every arm (same rationale as Batch 1).
| # | arm | env | role |
|---|---|---|---|
| 1 | pattern |
TR_MOVEMENT=strafe |
reference |
| 2 | bitbrain |
+ TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay |
largest Batch-1 damage edge (+10.3) |
| 3 | tmhorizon |
+ TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both |
second Batch-1 damage edge (+8.7) |
Pre-registered prediction (written BEFORE the battles finished):
- The
+10dmg/run hint does not survive the wider panel:bitbrainandtmhorizondamage deltas fall toward zero and stay inside the new (smaller) MDE; the per-opponent sign test remains non-significant. Rationale: both arms come from the same Pattern-lead-correction base, and the prior 32-opponent gauntlet measuredbitbrainat −3.6 dmg/run on a largely overlapping roster; +10 on 15 opponents is plausibly a small-panel fluctuation. - Round wins remain flat for both arms — if anything they regress toward zero.
- If instead the damage hint survives at a significant cross-opponent sign
test with wins not down, it is recorded as the first gun to beat
patternby rule 2 — and the campaign immediately looks for what the corrector is exploiting. That would be a genuine up-set, and it would not be spun as a null.
Outcome — Batch 2
MEASURED. Session /tmp/ab/j121_g2, commit
c343c00aaf800bc05ce9bf1b277054c11f5f577c, frozen binary sha256 d267ab78f6ce…,
33 opponents × 3 arms × 3 runs × 3 rounds = 297 battles, 0 failed, 0 never
started, 0 liveness exclusions.
A movement-default flip landed between the two batches (commit 3fd6db9,
TR_MOVEMENT default tfil → strafe; the only source change). Because
every arm in both batches pins TR_MOVEMENT=strafe, the effective
movement is identical in Batch 1 and Batch 2 — the batches are directly
comparable and the gun comparison is unconfounded. This is exactly what the pin
was for.
DIRECT ANSWER (MEASURED, n=33 opponents): the Batch-1 damage hint did NOT survive.
bitbraincollapses to +1.8 dmg/run (19/33 opponents, sign test p=0.49) with +0.11 wins/run — a wash.tmhorizonreverses sign and is detectably WORSE on damage: −10.4 dmg/run (10/33 opponents, sign test p=0.035, sign-flip p=0.0047, 95% CI [−17.2, −3.6]) at unchanged wins (−0.03). The only open signal in Batch 1 was a small-panel fluctuation, and the wider panel resolves it against the correctors. Combined with Batch 1: nothing beats the shippedPatternon damage/run AND round wins;onlyPatternis CONFIRMED on 15 opponents and again on 33.
Pooled dashboard (all valid runs — explanation only, NOT the verdict)
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|---|---|---|---|---|---|---|---|---|
pattern (REF) |
99 | 131.2 | 126.4 | 1.90 | 188/297 | 63.3% | 12.93% | 417 |
bitbrain |
99 | 133.0 | 126.7 | 2.01 | 199/297 | 67.0% | 12.91% | 417 |
tmhorizon |
99 | 120.8 | 133.2 | 1.87 | 185/297 | 62.3% | 13.04% | 422 |
Cross-opponent aggregation (the verdict layer, damage + wins)
| arm | metric | mean Δ | spread (SD) | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|---|---|---|---|---|---|---|---|---|---|
bitbrain |
damage | +1.80 | 18.95 | [−4.67, +8.26] | 19/33 | 0.4869 | 0.5914 | 0.5675 | 9.24 |
bitbrain |
wins | +0.11 | 0.48 | [−0.05, +0.28] | 12/19 | 0.3593 | 0.2436 | 0.5716 | 0.24 |
tmhorizon |
damage | −10.40 | 20.01 | [−17.22, −3.57] | 10/33 | 0.03508 | 0.00472 | 0.006975 | 9.76 |
tmhorizon |
wins | −0.03 | 0.59 | [−0.23, +0.17] | 9/19 | 1.0 | 0.8459 | 0.4804 | 0.29 |
The pre-registered verdict (verbatim)
| rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) |
|---|---|---|---|---|---|---|---|
| 1 | bitbrain |
+0.11 | +1.8 | 12/19 p=0.3593 | 19/33 p=0.4869 | not distinguishable | not distinguishable |
| 2 | tmhorizon |
−0.03 | −10.4 | 9/19 p=1 | 10/33 p=0.03508 | WORSE | WORSE |
Reference pattern: 131.2 dmg/run, 1.90 wins/run, 12.93% incoming, 417 px.
Reading
- BitBrain is a wash — now measured three times. vs the shipped Pattern:
−3.6 dmg/run on the 32-opponent legacy gauntlet (with
tfilmovement,docs/gauntlet_bitbrain_vs_pattern.md), +10.3 on the 15-opponent Batch-1 panel, +1.8 here on 33 opponents. All three sit inside their own MDEs; the pooled evidence is a zero. The Batch-1+10was a 15-opponent fluctuation. - TMHorizon is the one arm the wider panel separates — and it loses. Its
Batch-1 sign (+8.7) does not merely fail to replicate, it inverts
(−10.4, 10/33, p=0.035). The damage loss is broad (24/33 opponents negative,
worst
Aristocles−61,Jen−54,WallAvoider−41), so it is not one outlier. Since TMHorizon was already predicted to lose by its own design docs, this is the first confirming live panel measurement of that prediction. - Wins are flat for both arms (+0.11 / −0.03, MDEs 0.24 / 0.29): the correctors change damage at most, never survival. In this harness round wins are survival wins, so neither is a win.
- The corrector family is closed as a source of a
pattern-beating gun at the resolutions tested.
Pre-registered predictions — scorecard
| # | prediction | outcome |
|---|---|---|
| 1 | the +10 dmg hint does not survive the wider panel; deltas fall inside the new MDE; sign test non-significant | CORRECT (bitbrain +1.8, p=0.49; tmhorizon even reversed to −10.4, p=0.035) |
| 2 | wins remain flat | CORRECT (+0.11 / −0.03) |
| 3 | if the hint survived at p<0.05 with wins not down, it is the first gun to beat pattern |
not invoked |
Final direct answer for the phase (MEASURED). Across two frozen panels
(15 and 33 opponents), a single frozen-binary-per-session design, 567 battles,
seven distinct gun configurations, and movement pinned identically in every arm:
no rack gun and no small rack beats the shipped Pattern on damage/run AND
round wins. Every candidate is either not distinguishable (bitbrain,
tmhorizon, rack_pk, rack_pt at n=15), a detectable loss
(knn; tmhorizon at n=33), or a sub-MDE damage-only wobble
(bitbrain). The shipped onlyPattern rack is CONFIRMED, not merely
assumed — the gun axis's first phase ends in a measured optimum of the tested
design space.
What to try next (rewritten AFTER Batches 1–2 — recommendations, not results)
The primary question is answered: nothing beats the shipped pattern across
the frozen panel, and onlyPattern is CONFIRMED. Ranked by value per battle:
- The damage-side hint is now CLOSED. Batch 2 (33 opponents) killed it:
bitbrainis a wash (+1.8 dmg/run, p=0.49) andtmhorizonis detectably worse (−10.4 dmg/run, p=0.035). Do not spend another batch on the Pattern-lead correctors; three independent measurements ofbitbrainvs Pattern now average ~0 (−3.6 / +10.3 / +1.8 dmg/run, all inside their MDEs). - Lead INFORMATION remains the only named gun axis (given: amplitude is dead). But the two live implementations of it (BitBrain's gain, TMHorizon's shift) are both measured negative/neutral across panels. A future axis would need a different information source, not another knob on these two — e.g. a new base prediction, or a gun-agnostic ensemble. Do not re-open the correctors.
- The selector is confirmed negative at a 2-gun rack (
rack_pt==tmhorizon;rack_pkbelowpatternon wins). Do not re-open it with a larger rack: the 16-gun negative already stands and Batch 1 shows the mechanism does not manufacture a win. - Drop
knn(detectably worse on both primaries, 0/15 on damage). - Never chase damage without wins. This harness's round wins are survival
wins, so the gun's job is to kill; a damage gain that does not raise the win
rate is the
ring-mover trap in a new costume (andtmhorizon's n=33 result is a negative damage signal with flat wins — also not a win). - Spinner-specific claim (owner's) remains UNTESTED. The legacy roster has
no constant-turn spinner; SpinBot is a periodic circle-mover and is in
both panels (Batch 1:
bitbrain+25.3 dmg / +0 wins,tmhorizon+15.2 / +0 wins; on the 33-panel SpinBot's deltas are in the per-opponent tables). Building a true constant-turn spinner opponent is a nearly-free follow-up and is the one part of the owner's claim this campaign has not touched.
What would make us stop
PHASE STATUS (updated after Batch 2): STOPPED — the condition fired. Across
Batches 1–2 no arm beats pattern beyond the MDE, and no arm shows a
>= +MDE damage gain with p<0.10 (bitbrain +1.8 vs MDE 9.24; tmhorizon
actually negative). The gun stage's first phase is therefore closed with
"the shipped onlyPattern rack is the best gun configuration we have measured
across the panels (15 and 33 opponents)" — a successful outcome, not a failure.
Further gun work must open a named new axis (see What to try next), not
another arm on the same axis.
- Stop the first phase once a batch's best arm cannot beat
patternbeyond the MDE, or when a gun arm's damage gain is bought with a detectable win loss. At that point "onlyPatternis the measured optimum of this rack" is the conclusion, not a failure (§2 rule 4). This fired after Batch 2. - Stop a single batch early only for a contract violation (arena not free, liveness FAIL, non-zero exit rate) — never because the numbers look boring.
- Do not invent more arms on the same axis once two consecutive batches fail
to improve on
patternbeyond the MDE; move to a named new axis instead (candidate list in "What to try next").
How to run a batch (exact commands)
# 0. wait for the arena (other campaign jobs may be fighting)
TOURNAMENT_NIMCACHE=/tmp/nc_j121 \
tools/ab/tournament_run.sh \
--arms tools/ab/arms_gun_b1.txt \
--panel tools/ab/panel_movement.txt \
--runs 3 --rounds 3 --conc 6 --wait-arena 45 \
--reference pattern \
--outdir /tmp/ab/j121_g1
# Batch 2 (the wider-panel calibration) was the same command with
# --arms tools/ab/arms_gun_b2.txt --panel tools/ab/panel_gun_b2.txt \
# --outdir /tmp/ab/j121_g2 --wait-arena 50
# 1. the paired per-opponent table, sign tests, MDE and the pre-registered verdict
python3 tools/ab/tournament_analyze.py /tmp/ab/j121_g1 --reference pattern
--reference may be ANY arm of the session: re-analyzing an old session with a
different reference is a free pairwise comparison with no battles.
Session log (outdirs are in /tmp and are NOT committed)
| session | commit | battles | arms | verdict |
|---|---|---|---|---|
/tmp/ab/j121_g1 |
1d8143a |
270 (0 failed, 0 never started, 0 excluded) | pattern, bitbrain, tmhorizon, knn, rack_pk, rack_pt | nothing beats the shipped pattern; all arms not distinguishable except knn = WORSE; onlyPattern CONFIRMED |
/tmp/ab/j121_g2 |
c343c00 |
297 (0 failed, 0 never started, 0 excluded) | pattern, bitbrain, tmhorizon | the Batch-1 damage hint does not survive: bitbrain +1.8 dmg/run p=0.49 (wash); tmhorizon −10.4 dmg/run p=0.035 (WORSE); wins flat |