61 KiB
Gun campaign — ledger
Goal (owner's mandate, 2026-09-26 overnight): the movement campaign produced
a replicated champion (TR_MOVEMENT=strafe, docs/movement_campaign.md) which a
parallel job is shipping. "Continue until you found an amazing movement. When
found do the same over for a gun." This file is the GUN campaign's single source
of truth: later jobs append a ## Batch N section and never edit an earlier
one (a wrong earlier number gets a correction line, not a rewrite). This file is
the gun successor to docs/movement_campaign.md; read that first, then this.
OUTCOME (read this first)
The shipped gun is UNCHANGED: onlyPattern — the rack admits Pattern only
(common_libs/gun_harness/selector.nim, commit e0666a5). Nothing beat it.
The gun design space explored here is CLOSED on evidence, not on belief.
Headline (MEASURED). Four frozen-binary tournament sessions on the frozen
15- and 33-opponent panels: 1,233 battles (phase 1: 270 + 297; phase 2:
270 + 396), 0 failed, 0 never-started, 0 liveness exclusions. Each session is
ONE binary built from git archive HEAD; arms differ only by env; movement is
pinned TR_MOVEMENT=strafe in every arm, so every delta is a pure gun delta.
With the earlier 16-gun selector campaign (≈660 battles) the gun ledger now
stands on ≈1,900 measured battles. What was detectably worse: knn
(−39.3 dmg/run, 0/15 opponents; −0.62 wins/run, p=0.012) and, on the wider
33-opponent panel, TMHorizon (−10.4 dmg/run, sign p=0.035). What was
not distinguishable (inside the MDE, no verdict): bitbrain,
tmhorizon@15, rack_pk, rack_pt, len6, len16, depth100, the two radial
controls.
CLOSED axes (one-line reason each)
| axis | reason it is closed |
|---|---|
| other single guns | knn detectably worse; bitbrain a wash measured 3× (−3.6 / +10.3 / +1.8 dmg/run, all inside their MDEs); TMHorizon worse on 33 (docs/gauntlet_bitbrain_vs_pattern.md, Batches 1–2) |
| 2-gun selector racks | rack_pk nominally below pattern on wins; rack_pt == tmhorizon-alone → the selector just picks the corrector (## Batch 1) |
| selector mechanics | the virtual-fitness selector is negative value at every rack size tested (16, lean8, lean6, pairs); it never manufactured a win (docs/selector_negative_value.md) |
| lead amplitude | a full gain sweep (1.0/1.5/2.0/3.0) makes Pattern strictly worse at every range band — the lever is lead information, not amplitude (docs/bitbrain_campaign.md §0.3.2) |
| Pattern's own match-length / history params | len6's single-session +0.49 wins (p=0.039 on 15) reversed to a wash on 33 (−0.04, p=1); len16/depth100 flat; two batches, no replication (## Phase 2) |
radial knobs (TR_PATTERN_RAD_*) |
bearing-invariant by construction (job j99 proved bmPath is a structural no-op and scale/offset cannot change the aim bearing); in Batch 2 rad_offset is nominally negative |
THE ONE OPEN AXIS
A genuinely new source of lead INFORMATION — not another knob on
BitBrain/TMHorizon (both measured neutral→negative), not amplitude (dead), not
radial (bearing-invariant). The standing mechanism is docs/bitbrain_campaign.md
§0.3.3: at 450+ px Pattern's own lead correlation with the required lead is only
0.165, so it is adding variance to a nearly uninformative signal.
Honest caveat (MEASURED, docs/state_window_gate.md). The best surviving
idea from the state line — a single wave-relative state at Q=4 — does carry
signal at coarse resolution: it predicts the miss-offset bin on held-out
battles at 0.4094 (vs 0.2348 majority), but those bins are 36/42/60 px wide =
4.58/5.34/7.63° at 450 px, and its implied hit probability is still the base
rate (~0.09) — i.e. it is informative about which side the miss falls on,
not about whether the shot hits. The live hit window is
atan(18/450) = 2.29° half-width; the measured signal is at ~4.6–7.6°, not at
the ~2° the 450 px hit window needs. A temporal window does not fix this (the
window gate is negative; long contexts never recur). So the open axis is a new
information source that resolves the arrival offset to better than ~2°, which
does not exist yet.
Cost estimate for trying it (INFERRED). Build and offline-gate the new
information source first — require it to beat Pattern's 0.165 lead correlation
and, ideally, predict the arrival offset to <2.29° on held-out shots (days of
work; the offline instruments exist, e.g. common_libs/tests/state_window_gate.py).
Only then spend a live batch: one frozen binary, 6 env-only arms × 15 opponents ×
3 runs × 3 rounds = 270 battles ≈ 1–2 h wall at conc 6 with --wait-arena
(exact command in How to run a batch). Do not open with a live batch.
What would change our mind
A named new information source that (1) offline predicts the arrival offset to
better than the ~2.29° hit half-window at 450 px on held-out battles and
(2) in a live frozen-panel batch beats pattern on wins/run with the winning
metric's CI excluding 0 and sign-flip p<0.05 while damage is not detectably
down. A damage-only gain without wins is the ring-mover trap and does not
count.
What changed tonight (2026-09-26)
- SHIPPED: nothing in the gun. The rack is still
onlyPattern. - NOT shipped, and why: every candidate (other gun, small rack, selector
mechanic, match-length/history parameter, radial knob) failed to beat
patternbeyond the MDE on the frozen panels. Phase 2's one signal (len6, +0.49 wins/run on 15 opponents) did not replicate on 33 opponents. The only shipped change of the night is the movement default (tfil -> strafe, seedocs/movement_campaign.md); the gun was left untouched. - Revert: no gun revert is needed (no gun default changed). Movement:
TR_MOVEMENT=tfil. - Reproduce the key gun evidence (one command + the analyzer) — phase 2,
Batch 2, the session that killed
len6:TOURNAMENT_NIMCACHE=/tmp/nc_j123 \ tools/ab/tournament_run.sh \ --arms tools/ab/arms_gun_b4.txt \ --panel tools/ab/panel_gun_b2.txt \ --runs 3 --rounds 3 --conc 6 --wait-arena 45 \ --reference pattern \ --outdir /tmp/ab/j123_b2 python3 tools/ab/tournament_analyze.py /tmp/ab/j123_b2 --reference pattern
0. The question
The shipped rack admits Pattern only (onlyPattern). That decision was
taken because the virtual-fitness selector measured negative value at every
rack size tested, and Pattern is the best single gun by the DrussGT hit-rate
table (docs/selector_negative_value.md, docs/gun_rack_analysis.md, commit
e0666a5). But every one of those measurements — like every pre-campaign
movement claim — is DrussGT-heavy, and the movement campaign's standing
lesson is that a one-opponent result is not a result.
So Batch 1 asks the direct question, on a frozen 15-opponent panel:
Does any rack gun, or any small gun configuration, beat the shipped
Patternon damage/run AND round wins across 15 opponents — or isonlyPatternCONFIRMED rather than merely assumed?
A null is a successful outcome here: it upgrades onlyPattern from an
assumption to a measured verdict across the panel, which is itself valuable.
1. Protocol (identical to the movement campaign; the instrument is reused verbatim)
tools/ab/tournament_run.sh + tools/ab/tournament_analyze.py +
tools/ab/panel_movement.txt (frozen panel), reused unchanged. The unit of
evidence is the number of opponents, not the number of runs.
| Element | Rule |
|---|---|
| Subject | ONE frozen binary built from git archive HEAD (tournament_run.sh records commit + binary sha256 in session.json) |
| Arms | env dicts only — no per-arm rebuild, ever; the arm file is a committed file (tools/ab/arms_gun_b1.txt) |
| Panel | the frozen tools/ab/panel_movement.txt (15 opponents). Adding/removing an opponent starts a new batch number |
| Pairing | per opponent: average the arm's runs, subtract the reference's average for that same opponent → one delta per opponent; then aggregate |
| Isolation | per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir |
| Serialization | one battle fleet at a time via --wait-arena; never a broad pkill robocode_shim |
| Liveness | every declared env token must appear verbatim in OUR bot's own raw-env report, else the run is excluded and named |
| Never shipped | this is a measurement campaign: git status clean, shipped defaults untouched, .gitignore untouched |
Movement-default flip note (MEASURED). During this job a parallel job (j120, commit
3fd6db9) flipped the shippedTR_MOVEMENTdefault fromtfiltostrafe. Batch 1 built before the flip, Batch 2 after. Every arm in both batches declaresTR_MOVEMENT=strafe, so the effective movement is identical and the two batches are directly comparable — the pin did its job. The reference rack lines in the boot reports readPATTERNforpattern, and the active 1v1 rack readsBITBRAIN/TMHORIZON/KNNfor the overridden arms.
Movement is PINNED in every arm: TR_MOVEMENT=strafe
Reason (hard rule). A movement-default flip is landing from job j120 during
this run, so the default engine can change underneath us. Every arm therefore
sets TR_MOVEMENT=strafe explicitly. This (a) removes movement as a
confound, (b) makes all arms share exactly one movement, so every delta is a pure
gun delta, and (c) satisfies the analyzer's liveness rule for every arm
(each declared token must appear verbatim; an undeclared leaked TR_MOVEMENT
is fatal, but here every arm declares it). The reference arm is therefore the
shipped gun rack under the pinned movement, not under the (mutable) default.
2. Pre-registered decision rules (fixed BEFORE Batch 1 ran)
- Primary metrics: damage/run and ROUND WINS. Secondary/explanation only: damage taken/run, incoming hit rate, mean distance. Never hit rate alone — that trap has inverted six verdicts in this project.
- BETTER than the reference iff one primary metric is up with a
cross-opponent sign test p < 0.05 while the other does not go down;
the mirror image for WORSE. Anything else is NOT DISTINGUISHABLE (a
real answer, not a failure). Both the strict reading (the other metric's mean
delta
>= 0) and the substantive reading (not detectably down: not significant and smaller than that metric's MDE) are printed. - A verdict must survive the between-opponent spread: the pooled mean delta is reported with the SD across opponents, its SE, a 95% CI, and the MDE (α=0.05 two-sided, 80% power). An effect smaller than the MDE is reported as not detectable — never as absent, never as a win.
- Somewhere to stop: if no arm beats the shipped
patternby rule 2 in Batch 1 and no arm shows a>= +MDEdamage gain with p<0.10, then the gun stage's first phase is closed with "the shippedonlyPatternrack is the best gun configuration we have measured across the panel" — that is a successful outcome, and the campaign moves to a named next axis rather than inventing more arms. See What would make us stop. - No promotion off a single metric, a single opponent, or a single run. A change that wins damage by losing wins (or vice-versa) is not a win.
- Every batch is shot with a pre-registered prediction stated in its section before the battles finish; a prediction that turns out wrong is recorded as wrong.
3. Stage 0 — what we already know (given, not re-derived)
| fact | value | source |
|---|---|---|
| shipped rack | onlyPattern (id 5), all other guns off |
common_libs/gun_harness/selector.nim |
| why | selector measured negative value at every rack size (16, lean8, lean6, pairPC/PK/PL) and 10/10 adversaries | docs/selector_negative_value.md, docs/gun_rack_analysis.md |
| Pattern hit rate vs DrussGT | ~10.8% | given |
| KNN / Linear / Circular / WallBounce / GF vs DrussGT | 5.6 / 3.0 / 2.9 / 2.7 / 2.1 % | given |
| BitBrain when idle | == Pattern (30 runs/arm: 97/210 vs 97/210 wins) | docs/bitbrain_vs_tmhorizon_ab.md |
| BitBrain across 32 legacy opponents | neutral: -3.6 dmg/run, +0.02 wins/run, p=0.53/0.91 | docs/gauntlet_bitbrain_vs_pattern.md |
| lead amplitude | dead axis — full gain sweep {0, 0.25-1.0, 1.0, 1.25, 1.5, 2.0} run, nothing beats Pattern | docs/bitbrain_campaign.md |
| offline prediction checks | veto-only — may kill a design, never select one | docs/offline_harness_trust.md |
| spinner-specific claim | UNTESTED: no constant-turn spinner exists in the legacy roster (a note for a later job, not built here unless nearly free) | docs/gauntlet_bitbrain_vs_pattern.md |
Movement context (why the pin): the movement champion is TR_MOVEMENT=strafe
at its current defaults (range 325, tol 25, tilt 15/0.10); it replicated its win
over the shipped tfil in four independent sessions (Batch 4: 52.6% vs 38.5%
round-win rate). Movement is held constant at that champion for every gun arm.
Batch 1 — does anything beat the shipped Pattern?
Design. One frozen binary, six env-only arms, one frozen panel
(tools/ab/panel_movement.txt, 15 opponents: 5 dodger, 3 pattern, 1
wall-follower, 1 corner-camper, 1 spinner, 1 rammer, 1 brawler, 2 aggressive
megas), 3 runs × 3 rounds per (opponent, arm) = 270 battles. Arm file:
tools/ab/arms_gun_b1.txt. Reference: pattern.
| # | arm | env | what it isolates |
|---|---|---|---|
| 1 | pattern |
TR_MOVEMENT=strafe |
the arm to beat (shipped onlyPattern rack) |
| 2 | bitbrain |
+ TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay |
BitBrain-only, learned gain (panel replication of the 32-opp gauntlet at the pinned-movement standard) |
| 3 | tmhorizon |
+ TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both |
TMHorizon-only corrector (predicted to lose; never panel-tested) |
| 4 | knn |
+ TR_RACK_PATTERN=off TR_RACK_KNN=both |
best non-Pattern single gun (never panel-tested) |
| 5 | rack_pk |
+ TR_RACK_KNN=both |
Pattern + KNN, selector ON (different-family hedge) |
| 6 | rack_pt |
+ TR_RACK_TMHORIZON=both |
Pattern + TMHorizon, selector ON (corrector-family hedge) |
(+ = the pinned TR_MOVEMENT=strafe is present in every arm; see §1.)
Pre-registered prediction (written BEFORE the battles finished):
- No arm beats
patternon BOTH primaries, and the point estimates sit inside the MDE.onlyPatternis confirmed across the panel. bitbrainis a wash vspattern(replicating the 32-opponent gauntlet: ±3.6 dmg/run, ≈0 wins) — the DrussGT-only gain-config penalty does not carry.tmhorizonis WORSE thanpattern(its own offline work predicts it needs ~80% side accuracy and can reach ~60%); no damage win, no win win.knnis WORSE thanpatternon damage (its DrussGT hit rate is roughly half Pattern's) and not better on wins.- The two small racks do not beat
pattern;rack_pklands closer to Pattern thanknn-alone does (the selector at least partially hedges back), but the selector's poor ranking keeps it below the reference. - If any surprise exists, it is a rack arm on damage without wins — the
ring-mover mirror-image trap — and it will not be read as a win.
Outcome — direct answer
MEASURED. Session /tmp/ab/j121_g1, commit 1d8143a15e4a038c39dfc2153ff518fa003b8612
(the pre-registration commit), frozen binary sha256 aa49a45fec20…, 15 opponents
× 6 arms × 3 runs × 3 rounds = 270 battles, 0 failed, 0 never started, 0
liveness exclusions. Every arm declared TR_MOVEMENT=strafe, verified in each
bot's own raw-env report; the reference rack line reads PATTERN for pattern
and the overridden rack reads BITBRAIN / TMHORIZON / KNN for the others.
DIRECT ANSWER (MEASURED, n=15 opponents): NO — nothing beats the shipped
Patternacross the panel on damage/run AND round wins. Every arm exceptknnis NOT DISTINGUISHABLE frompatternunder both the strict and the substantive reading of the pre-registered rule. The best nominal challenger,tmhorizon, is +0.13 wins/run (MDE 0.36) and +8.7 dmg/run (MDE 13.8) — an effect ~3× smaller than the design can detect — andbitbrainis+0.11wins /+10.3dmg, also inside the MDE.knnis detectably WORSE on both primaries. The shippedonlyPatternrack is therefore CONFIRMED rather than merely assumed, at the resolution of this panel. That is a real result, not a failure.
Pooled dashboard (all valid runs — explanation only, NOT the verdict)
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|---|---|---|---|---|---|---|---|---|
pattern (REF) |
45 | 99.6 | 155.9 | 1.51 | 68/135 | 50.4% | 12.84% | 435 |
bitbrain |
45 | 109.9 | 151.2 | 1.62 | 73/135 | 54.1% | 12.87% | 427 |
tmhorizon |
45 | 108.3 | 153.7 | 1.64 | 74/135 | 54.8% | 12.53% | 428 |
knn |
45 | 60.3 | 175.7 | 0.89 | 40/135 | 29.6% | 13.31% | 430 |
rack_pk |
45 | 107.3 | 161.3 | 1.42 | 64/135 | 47.4% | 13.65% | 432 |
rack_pt |
45 | 105.3 | 149.9 | 1.64 | 74/135 | 54.8% | 12.90% | 424 |
Cross-opponent aggregation (the verdict layer, damage + wins)
| arm | metric | mean Δ | spread (SD) | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|---|---|---|---|---|---|---|---|---|---|
bitbrain |
damage | +10.34 | 16.65 | [+1.12, +19.57] | 11/15 | 0.1185 | 0.03247 | 0.05006 | 12.05 |
bitbrain |
wins | +0.11 | 0.48 | [-0.16, +0.38] | 6/9 | 0.5078 | 0.4922 | 0.4764 | 0.35 |
tmhorizon |
damage | +8.71 | 19.10 | [-1.86, +19.29] | 11/15 | 0.1185 | 0.09918 | 0.1183 | 13.82 |
tmhorizon |
wins | +0.13 | 0.50 | [-0.14, +0.41] | 5/9 | 1.0 | 0.4062 | 0.3118 | 0.36 |
knn |
damage | −39.33 | 23.15 | [−52.15, −26.51] | 0/15 | 6.1e-5 | 6.1e-5 | 0.0007 | 16.75 |
knn |
wins | −0.62 | 0.59 | [−0.95, −0.30] | 1/11 | 0.01172 | 0.001953 | 0.004948 | 0.43 |
rack_pk |
damage | +7.68 | 32.09 | [−10.10, +25.45] | 9/15 | 0.6072 | 0.423 | 0.5895 | 23.21 |
rack_pk |
wins | −0.09 | 0.60 | [−0.42, +0.24] | 6/11 | 1.0 | 0.6738 | 0.5932 | 0.43 |
rack_pt |
damage | +5.69 | 16.50 | [−3.45, +14.83] | 10/15 | 0.3018 | 0.2026 | 0.2681 | 11.93 |
rack_pt |
wins | +0.13 | 0.41 | [−0.10, +0.36] | 6/10 | 0.7539 | 0.3242 | 0.1997 | 0.30 |
The pre-registered verdict (verbatim)
| rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) |
|---|---|---|---|---|---|---|---|
| 1 | tmhorizon |
+0.13 | +8.7 | 5/9 p=1 | 11/15 p=0.1185 | not distinguishable | not distinguishable |
| 2 | rack_pt |
+0.13 | +5.7 | 6/10 p=0.7539 | 10/15 p=0.3018 | not distinguishable | not distinguishable |
| 3 | bitbrain |
+0.11 | +10.3 | 6/9 p=0.5078 | 11/15 p=0.1185 | not distinguishable | not distinguishable |
| 4 | rack_pk |
−0.09 | +7.7 | 6/11 p=1 | 9/15 p=0.6072 | not distinguishable | not distinguishable |
| 5 | knn |
−0.62 | −39.3 | 1/11 p=0.01172 | 0/15 p=6.104e-05 | WORSE | not distinguishable |
Reference pattern: 99.6 dmg/run, 1.51 wins/run, 12.84% incoming, 435 px.
Reading (MEASURED, with the mechanism)
knnis a clean, large loss and is dropped. It deals −39 dmg/run and loses −0.62 wins/run, positive on 0/15 opponents on damage and 1/11 on wins (p=6e-5 / p=0.012). Its incoming hit rate is no worse (13.31% vs 12.84%) — it simply cannot aim: the same movement, a worse gun.- The two small racks do not beat
pattern.rack_pk(Pattern+KNN, selector ON) is not the average of Pattern and KNN: it recovers most of KNN's damage loss (+7.7 vs KNN-alone's −39.3) but still lands nominally below Pattern on wins (−0.09).rack_pt(Pattern+TMHorizon) is statistically indistinguishable fromtmhorizon-alone — adding Pattern to the rack changes nothing, i.e. the selector picks the corrector essentially always. The selector's 16-gun negative verdict reproduces at a 2-gun rack: it does not manufacture a win it did not have. - The only signal is a damage-side hint on the two Pattern-lead correctors
(
bitbrain+10.3,tmhorizon+8.7, both 11/15 opponents). It is below the MDE (12.0 / 13.8) and not significant by the pre-registered sign test (p=0.1185), so it is not a win and not a null. Note the sign-flip test on the mean is nominally significant forbitbrain(p=0.032) while the sign test is not — the positive deltas are larger than the negatives — but with 5 arms compared this is not compelling after multiplicity, and the effect is under the MDE. This is exactly the "damage without wins" pattern this project keeps paying for (theringmover: +31 dmg/run and fewer wins); here the win deltas (+0.11/+0.13) are themselves sub-MDE. - The reference is stable:
patternunderTR_MOVEMENT=strafemeasures 99.6 dmg/run and 50.4% round wins here, matching the movement campaign'sstrafereference in Batch 4 (103.5 dmg/run, 52.6%) — same movement, same panel, different job.
Pre-registered predictions — scorecard (an honest count)
| # | prediction | outcome |
|---|---|---|
| 1 | no arm beats pattern on both primaries, point estimates inside the MDE; onlyPattern confirmed |
CORRECT |
| 2 | bitbrain is a wash vs pattern (gauntlet: −3.6 dmg, ≈0 wins) |
CORRECT on the verdict, WRONG on the point estimate: +10.3 dmg here vs −3.6 in the 32-opponent gauntlet; both inside their MDEs. Direction differs, verdict (wash) holds |
| 3 | tmhorizon is WORSE than pattern (its own docs predict it loses) |
WRONG: it is nominally better on both primaries (+8.7 dmg, +0.13 wins), though not distinguishable |
| 4 | knn is WORSE than pattern on damage, not better on wins |
CORRECT (detectably on both) |
| 5 | the small racks do not beat pattern; rack_pk lands closer to Pattern than knn-alone |
CORRECT |
| 6 | any surprise is a rack arm on damage without wins | PARTLY CORRECT: the nominal damage edge is on the single-gun correctors and the racks; every one of them is win-flat |
Batch 2 — wider-panel calibration of the damage-side hint
The Batch-1 answer is decisive on the primary question (nothing beats pattern)
but the two corrector arms carry a sub-MDE damage hint that rule 3 forbids
calling either way. Resolving it needs more opponents, not more runs, so this
batch re-runs exactly those two arms plus the reference on a wider panel.
Design. One frozen binary, three env-only arms, panel
tools/ab/panel_gun_b2.txt = the j117 32-opponent legacy roster minus Aurora
(inert in j117: one fire in 6 battles — it cannot reveal a gun regression)
plus SpinBot = 33 opponents (a superset of Batch 1's 15). 3 runs × 3
rounds = 297 battles, --conc 6, --wait-arena. Arm file:
tools/ab/arms_gun_b2.txt. Reference: pattern. Movement pinned
TR_MOVEMENT=strafe in every arm (same rationale as Batch 1).
| # | arm | env | role |
|---|---|---|---|
| 1 | pattern |
TR_MOVEMENT=strafe |
reference |
| 2 | bitbrain |
+ TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay |
largest Batch-1 damage edge (+10.3) |
| 3 | tmhorizon |
+ TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both |
second Batch-1 damage edge (+8.7) |
Pre-registered prediction (written BEFORE the battles finished):
- The
+10dmg/run hint does not survive the wider panel:bitbrainandtmhorizondamage deltas fall toward zero and stay inside the new (smaller) MDE; the per-opponent sign test remains non-significant. Rationale: both arms come from the same Pattern-lead-correction base, and the prior 32-opponent gauntlet measuredbitbrainat −3.6 dmg/run on a largely overlapping roster; +10 on 15 opponents is plausibly a small-panel fluctuation. - Round wins remain flat for both arms — if anything they regress toward zero.
- If instead the damage hint survives at a significant cross-opponent sign
test with wins not down, it is recorded as the first gun to beat
patternby rule 2 — and the campaign immediately looks for what the corrector is exploiting. That would be a genuine up-set, and it would not be spun as a null.
Outcome — Batch 2
MEASURED. Session /tmp/ab/j121_g2, commit
c343c00aaf800bc05ce9bf1b277054c11f5f577c, frozen binary sha256 d267ab78f6ce…,
33 opponents × 3 arms × 3 runs × 3 rounds = 297 battles, 0 failed, 0 never
started, 0 liveness exclusions.
A movement-default flip landed between the two batches (commit 3fd6db9,
TR_MOVEMENT default tfil → strafe; the only source change). Because
every arm in both batches pins TR_MOVEMENT=strafe, the effective
movement is identical in Batch 1 and Batch 2 — the batches are directly
comparable and the gun comparison is unconfounded. This is exactly what the pin
was for.
DIRECT ANSWER (MEASURED, n=33 opponents): the Batch-1 damage hint did NOT survive.
bitbraincollapses to +1.8 dmg/run (19/33 opponents, sign test p=0.49) with +0.11 wins/run — a wash.tmhorizonreverses sign and is detectably WORSE on damage: −10.4 dmg/run (10/33 opponents, sign test p=0.035, sign-flip p=0.0047, 95% CI [−17.2, −3.6]) at unchanged wins (−0.03). The only open signal in Batch 1 was a small-panel fluctuation, and the wider panel resolves it against the correctors. Combined with Batch 1: nothing beats the shippedPatternon damage/run AND round wins;onlyPatternis CONFIRMED on 15 opponents and again on 33.
Pooled dashboard (all valid runs — explanation only, NOT the verdict)
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|---|---|---|---|---|---|---|---|---|
pattern (REF) |
99 | 131.2 | 126.4 | 1.90 | 188/297 | 63.3% | 12.93% | 417 |
bitbrain |
99 | 133.0 | 126.7 | 2.01 | 199/297 | 67.0% | 12.91% | 417 |
tmhorizon |
99 | 120.8 | 133.2 | 1.87 | 185/297 | 62.3% | 13.04% | 422 |
Cross-opponent aggregation (the verdict layer, damage + wins)
| arm | metric | mean Δ | spread (SD) | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|---|---|---|---|---|---|---|---|---|---|
bitbrain |
damage | +1.80 | 18.95 | [−4.67, +8.26] | 19/33 | 0.4869 | 0.5914 | 0.5675 | 9.24 |
bitbrain |
wins | +0.11 | 0.48 | [−0.05, +0.28] | 12/19 | 0.3593 | 0.2436 | 0.5716 | 0.24 |
tmhorizon |
damage | −10.40 | 20.01 | [−17.22, −3.57] | 10/33 | 0.03508 | 0.00472 | 0.006975 | 9.76 |
tmhorizon |
wins | −0.03 | 0.59 | [−0.23, +0.17] | 9/19 | 1.0 | 0.8459 | 0.4804 | 0.29 |
The pre-registered verdict (verbatim)
| rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) |
|---|---|---|---|---|---|---|---|
| 1 | bitbrain |
+0.11 | +1.8 | 12/19 p=0.3593 | 19/33 p=0.4869 | not distinguishable | not distinguishable |
| 2 | tmhorizon |
−0.03 | −10.4 | 9/19 p=1 | 10/33 p=0.03508 | WORSE | WORSE |
Reference pattern: 131.2 dmg/run, 1.90 wins/run, 12.93% incoming, 417 px.
Reading
- BitBrain is a wash — now measured three times. vs the shipped Pattern:
−3.6 dmg/run on the 32-opponent legacy gauntlet (with
tfilmovement,docs/gauntlet_bitbrain_vs_pattern.md), +10.3 on the 15-opponent Batch-1 panel, +1.8 here on 33 opponents. All three sit inside their own MDEs; the pooled evidence is a zero. The Batch-1+10was a 15-opponent fluctuation. - TMHorizon is the one arm the wider panel separates — and it loses. Its
Batch-1 sign (+8.7) does not merely fail to replicate, it inverts
(−10.4, 10/33, p=0.035). The damage loss is broad (24/33 opponents negative,
worst
Aristocles−61,Jen−54,WallAvoider−41), so it is not one outlier. Since TMHorizon was already predicted to lose by its own design docs, this is the first confirming live panel measurement of that prediction. - Wins are flat for both arms (+0.11 / −0.03, MDEs 0.24 / 0.29): the correctors change damage at most, never survival. In this harness round wins are survival wins, so neither is a win.
- The corrector family is closed as a source of a
pattern-beating gun at the resolutions tested.
Pre-registered predictions — scorecard
| # | prediction | outcome |
|---|---|---|
| 1 | the +10 dmg hint does not survive the wider panel; deltas fall inside the new MDE; sign test non-significant | CORRECT (bitbrain +1.8, p=0.49; tmhorizon even reversed to −10.4, p=0.035) |
| 2 | wins remain flat | CORRECT (+0.11 / −0.03) |
| 3 | if the hint survived at p<0.05 with wins not down, it is the first gun to beat pattern |
not invoked |
Final direct answer for the phase (MEASURED). Across two frozen panels
(15 and 33 opponents), a single frozen-binary-per-session design, 567 battles,
six distinct gun configurations (five challengers + the reference), and movement
pinned identically in every arm:
no rack gun and no small rack beats the shipped Pattern on damage/run AND
round wins. Every candidate is either not distinguishable (bitbrain,
tmhorizon, rack_pk, rack_pt at n=15), a detectable loss
(knn; tmhorizon at n=33), or a sub-MDE damage-only wobble
(bitbrain). The shipped onlyPattern rack is CONFIRMED, not merely
assumed — the gun axis's first phase ends in a measured optimum of the tested
design space.
What to try next (rewritten AFTER Batches 1–2 — recommendations, not results)
The primary question is answered: nothing beats the shipped pattern across
the frozen panel, and onlyPattern is CONFIRMED. Ranked by value per battle:
- The damage-side hint is now CLOSED. Batch 2 (33 opponents) killed it:
bitbrainis a wash (+1.8 dmg/run, p=0.49) andtmhorizonis detectably worse (−10.4 dmg/run, p=0.035). Do not spend another batch on the Pattern-lead correctors; three independent measurements ofbitbrainvs Pattern now average ~0 (−3.6 / +10.3 / +1.8 dmg/run, all inside their MDEs). - Lead INFORMATION remains the only named gun axis (given: amplitude is dead). But the two live implementations of it (BitBrain's gain, TMHorizon's shift) are both measured negative/neutral across panels. A future axis would need a different information source, not another knob on these two — e.g. a new base prediction, or a gun-agnostic ensemble. Do not re-open the correctors.
- The selector is confirmed negative at a 2-gun rack (
rack_pt==tmhorizon;rack_pkbelowpatternon wins). Do not re-open it with a larger rack: the 16-gun negative already stands and Batch 1 shows the mechanism does not manufacture a win. - Drop
knn(detectably worse on both primaries, 0/15 on damage). - Never chase damage without wins. This harness's round wins are survival
wins, so the gun's job is to kill; a damage gain that does not raise the win
rate is the
ring-mover trap in a new costume (andtmhorizon's n=33 result is a negative damage signal with flat wins — also not a win). - Spinner-specific claim (owner's) remains UNTESTED. The legacy roster has
no constant-turn spinner; SpinBot is a periodic circle-mover and is in
both panels (Batch 1:
bitbrain+25.3 dmg / +0 wins,tmhorizon+15.2 / +0 wins; on the 33-panel SpinBot's deltas are in the per-opponent tables). Building a true constant-turn spinner opponent is a nearly-free follow-up and is the one part of the owner's claim this campaign has not touched.
What would make us stop
PHASE STATUS (updated after Batch 2): STOPPED — the condition fired. Across
Batches 1–2 no arm beats pattern beyond the MDE, and no arm shows a
>= +MDE damage gain with p<0.10 (bitbrain +1.8 vs MDE 9.24; tmhorizon
actually negative). The gun stage's first phase is therefore closed with
"the shipped onlyPattern rack is the best gun configuration we have measured
across the panels (15 and 33 opponents)" — a successful outcome, not a failure.
Further gun work must open a named new axis (see What to try next), not
another arm on the same axis.
- Stop the first phase once a batch's best arm cannot beat
patternbeyond the MDE, or when a gun arm's damage gain is bought with a detectable win loss. At that point "onlyPatternis the measured optimum of this rack" is the conclusion, not a failure (§2 rule 4). This fired after Batch 2. - Stop a single batch early only for a contract violation (arena not free, liveness FAIL, non-zero exit rate) — never because the numbers look boring.
- Do not invent more arms on the same axis once two consecutive batches fail
to improve on
patternbeyond the MDE; move to a named new axis instead (candidate list in "What to try next").
How to run a batch (exact commands)
# 0. wait for the arena (other campaign jobs may be fighting)
TOURNAMENT_NIMCACHE=/tmp/nc_j121 \
tools/ab/tournament_run.sh \
--arms tools/ab/arms_gun_b1.txt \
--panel tools/ab/panel_movement.txt \
--runs 3 --rounds 3 --conc 6 --wait-arena 45 \
--reference pattern \
--outdir /tmp/ab/j121_g1
# Batch 2 (the wider-panel calibration) was the same command with
# --arms tools/ab/arms_gun_b2.txt --panel tools/ab/panel_gun_b2.txt \
# --outdir /tmp/ab/j121_g2 --wait-arena 50
# 1. the paired per-opponent table, sign tests, MDE and the pre-registered verdict
python3 tools/ab/tournament_analyze.py /tmp/ab/j121_g1 --reference pattern
--reference may be ANY arm of the session: re-analyzing an old session with a
different reference is a free pairwise comparison with no battles.
Session log (outdirs are in /tmp and are NOT committed)
| session | commit | battles | arms | verdict |
|---|---|---|---|---|
/tmp/ab/j121_g1 |
1d8143a |
270 (0 failed, 0 never started, 0 excluded) | pattern, bitbrain, tmhorizon, knn, rack_pk, rack_pt | nothing beats the shipped pattern; all arms not distinguishable except knn = WORSE; onlyPattern CONFIRMED |
/tmp/ab/j121_g2 |
c343c00 |
297 (0 failed, 0 never started, 0 excluded) | pattern, bitbrain, tmhorizon | the Batch-1 damage hint does not survive: bitbrain +1.8 dmg/run p=0.49 (wash); tmhorizon −10.4 dmg/run p=0.035 (WORSE); wins flat |
Phase 2: tuning the incumbent
Owner's mandate (phase 2): the phase-1 campaign closed the "other gun / other
rack" design space (nothing beats the shipped Pattern; the correctors are a
wash or a loss). What remains OPEN is the incumbent's own tuning: Pattern's
match-length / history parameters had NEVER been swept. Phase 2 asks the direct
question:
Can the shipped
Patternbe improved by tuning its own match parameters — and if so, by how much on damage/run and round wins?
The lead-amplitude axis is already known dead (docs/bitbrain_campaign.md), and
the radial knobs are known non-winners (job j99: bmPath structural no-op; live
+0.28 pp p=0.62 for scale 0.98, −0.42 pp p=0.46 for offset −20). Those are
therefore controls here, not candidates.
Task A — exposing Pattern's match-shape parameters (what and why)
MEASURED (code read). common_libs/guns/pattern_matcher.nim had exactly two
tunable knobs, both RADIAL (TR_PATTERN_RAD_SCALE, TR_PATTERN_RAD_OFFSET), and
neither can change the lead bearing (job j99 proved the bearing is untouched), so
neither was ever the "lead information" axis. The parameters that actually
control how the pattern is matched were compile-time constants:
| parameter | code | what it controls | exposed as |
|---|---|---|---|
| match-key length | PatternLen = 10 |
the length of the movement segment compared (the search key); also how far after the match the replay starts | TR_PATTERN_LEN (int, default 10) |
| search depth | implicit HistorySize = 500 |
how far back the best-match scan may reach (scanEnd was always count−PatternLen−1) |
TR_PATTERN_DEPTH (int, default 500 = full buffer) |
| similarity/search radius | does not exist | the search always takes the single lowest-cost match; there is no acceptance threshold or radius to expose | not exposed — nothing to expose |
| history buffer capacity | HistorySize = 500 |
the fixed array[HistorySize] backing store |
runtime depth limit only; the buffer ceiling cannot be raised at runtime (see below) |
Why env and not -d:. The phase-1 instrument (tools/ab/tournament_run.sh)
builds ONE frozen binary from git archive HEAD and every arm differs only by its
env dict. A {.intdefine.} knob would need one binary per arm, which the
instrument forbids. Both new knobs are therefore resolved lazily from the
environment on the first predict, exactly like the radial knobs, and default to
the pre-knob constants. setMatchParams(patternLen, histDepth) is the explicit
offline/unit-test twin that writes the same fields.
What could NOT be exposed cheaply (MEASURED). HistorySize sizes four fixed
array[HistorySize(+1)] fields in the gun object. Raising it above 500 at
runtime is impossible without a heap buffer; the runtime TR_PATTERN_DEPTH knob
therefore lowers the effective search depth within the existing 500-entry
buffer. A depth above 500 is clamped to 500, and a non-positive or unparsable
value falls back to the shipped default. PatternLen is clamped to
1..HistorySize.
Task A — default parity (MEASURED, byte-for-byte)
- Baseline-vs-new parity dump. A throwaway harness replayed 800 ticks of
tools/fixtures/drussgt_vs_crazy.jsonlthrough onePatternMatcherGunat all four power bins (3200 predictions) and printed every point at full precision. Compiled once against aHEADworktree (/tmp/j123_base, before the change) and once against the modified tree: the two dumps are byte-identical (diff -qclean). The shipped default path is unchanged. - The knobs are live and self-falling-back.
TR_PATTERN_LEN=6/16andTR_PATTERN_DEPTH=100/20each change the dump;TR_PATTERN_LEN=bananareproduces the default dump exactly. - Existing guards pass (no new guard tests, per the "cut ceremony" rule):
common_libs/tests/test_pattern_radial_offset.nim(6 checks) andcommon_libs/tests/test_gun_harness.nim(all checks) pass; the env-report guardtest_env_report.nimpasses withTR_PATTERN_LEN/TR_PATTERN_DEPTHregistered inknownEnvNames()and the effective-values report. - Clean-archive compile is exercised by
tournament_run.shitself, which builds the frozen binary fromgit archive HEAD.
Task B — Batch 1 (pre-registered BEFORE any battle)
Design. One frozen binary, six env-only arms, the FROZEN 15-opponent panel
tools/ab/panel_movement.txt, 3 runs × 3 rounds per (opponent, arm) = 270
battles, --conc 6, --wait-arena. Arm file: tools/ab/arms_gun_b3.txt.
Reference: pattern. Movement pinned TR_MOVEMENT=strafe in every arm.
| # | arm | env over the pin | what it isolates |
|---|---|---|---|
| 1 | pattern |
(none) | the arm to beat (shipped onlyPattern rack) |
| 2 | len6 |
TR_PATTERN_LEN=6 |
shorter movement segment compared |
| 3 | len16 |
TR_PATTERN_LEN=16 |
longer movement segment compared |
| 4 | depth100 |
TR_PATTERN_DEPTH=100 |
shallower history / match search |
| 5 | rad_offset |
TR_PATTERN_RAD_OFFSET=-20 |
control: aim 20 px short |
| 6 | rad_scale |
TR_PATTERN_RAD_SCALE=0.95 |
control: scale the aim distance |
Pre-registered prediction (written BEFORE the battles finished):
- No arm beats
patternon both primaries. The incumbent is already tuned — the phase-2 null. The point estimates sit inside the MDE. len16is WORSE or flat on damage: a 16-tick key matches rarely in a ~500-tick buffer, so the gun falls back to the linear forecast (the weaker base) more often; wins flat.len6is flat on damage (more matches but a noisier replay) and flat on wins; possibly a sub-MDE wobble in either direction.depth100is not distinguishable frompattern— the full buffer already contains the useful candidates; a sub-MDE damage wobble is possible.- The two radial controls are not distinguishable and, per job j99, cannot change the real aim bearing; any damage delta is the fire-gate channel only.
- If any surprise exists, it is a damage-only wobble without wins — the
ring-mover trap — and will not be read as a win.
Pre-registered Batch-2 trigger (task rule). A Batch-1 arm is promoted to a
higher-power Batch 2 only if it beats the reference pattern by the
campaign rule 2 (one primary up at cross-opponent sign-test p<0.05 while the
other does not go down) AND its damage CI excludes 0 AND the sign-flip
permutation test gives p<0.05. Otherwise Batch 1 is the answer.
Outcome — Batch 1
MEASURED. Session /tmp/ab/j123_b1, commit
2a98aba91b0ec023f580cd2485cafba2404d275d, frozen binary sha256 287d292fb7d1…,
15 opponents × 6 arms × 3 runs × 3 rounds = 270 battles, 0 failed, 0 never
started, 0 liveness exclusions. Arena serialization held: the parallel job
j124 (/tmp/ab/j124_spinner.run.log) detected this session's processes at
04:08–04:18 and waited; no foreign battle ran concurrently. Every arm
declared TR_MOVEMENT=strafe; the new TR_PATTERN_LEN / TR_PATTERN_DEPTH
tokens appear verbatim in each overriding arm's own raw-env report and the
reference arm shows the shipped defaults 10 / 500.
Pooled dashboard (all valid runs — explanation only, NOT the verdict)
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|---|---|---|---|---|---|---|---|---|
pattern (REF) |
45 | 101.1 | 161.0 | 1.27 | 57/135 | 42.2% | 13.27% | 428 |
len6 |
45 | 110.8 | 151.2 | 1.76 | 79/135 | 58.5% | 12.28% | 424 |
len16 |
45 | 103.3 | 158.8 | 1.53 | 69/135 | 51.1% | 13.07% | 427 |
depth100 |
45 | 99.9 | 154.5 | 1.53 | 69/135 | 51.1% | 12.77% | 429 |
rad_offset |
45 | 107.9 | 157.2 | 1.69 | 76/135 | 56.3% | 13.21% | 431 |
rad_scale |
45 | 105.9 | 157.1 | 1.38 | 62/135 | 45.9% | 12.62% | 429 |
Per-opponent paired deltas (arm − pattern), damage and wins
| opponent | style | Δdmg_len6 | Δwins_len6 | Δdmg_len16 | Δwins_len16 | Δdmg_depth100 | Δwins_depth100 | Δdmg_rad_offset | Δwins_rad_offset | Δdmg_rad_scale | Δwins_rad_scale |
|---|---|---|---|---|---|---|---|---|---|---|---|
| DrussGT | dodger | +8.4 | +0.67 | -21.7 | -0.33 | +3.6 | +0.67 | +10.5 | +1.33 | +18.7 | +0.67 |
| Diamond | dodger | -22.3 | +0.00 | -18.1 | -0.33 | -54.7 | -0.33 | -30.9 | +0.00 | -17.0 | -0.33 |
| Dookious | dodger | -6.6 | +1.33 | +7.6 | +1.00 | -0.8 | +0.67 | +2.6 | +0.67 | -15.6 | -0.33 |
| GresSuffurd | dodger | +32.3 | +0.00 | +31.9 | +0.00 | -6.6 | +0.00 | +28.4 | +0.67 | -0.9 | +0.00 |
| CassiusClay | dodger | +15.0 | +0.67 | +6.8 | +1.00 | +10.0 | +0.33 | +12.5 | +0.67 | +40.4 | +1.00 |
| RetroGirl | pattern | +48.0 | +1.33 | +20.3 | +0.67 | +59.9 | +1.00 | +4.1 | +0.00 | +55.5 | +0.33 |
| TripHammer | pattern | +6.1 | +0.00 | +0.1 | +0.33 | -21.6 | +0.00 | -9.3 | +0.00 | +4.7 | +0.33 |
| Coriantumr | pattern | +1.9 | -0.33 | -3.9 | +0.33 | +2.0 | +1.00 | -5.3 | -0.33 | -10.3 | -0.33 |
| WallAvoider | wallfollower | +28.5 | +1.00 | -23.5 | +0.00 | +41.0 | +0.00 | +23.6 | +1.00 | -24.0 | +0.00 |
| HawkOnFire | cornercamper | -8.9 | +0.67 | -9.5 | +0.00 | -9.3 | -0.33 | -12.0 | +0.33 | -5.0 | +0.33 |
| SpinBot | spinner | +5.7 | +0.00 | +13.3 | +0.00 | -13.5 | +0.00 | -1.7 | +0.00 | -0.3 | +0.00 |
| DiamondStealer | rammer | -16.1 | +0.00 | -9.8 | +0.00 | -14.6 | +0.00 | +31.8 | +1.33 | +1.1 | -0.67 |
| BlitzBat | brawler | +9.1 | +0.00 | +17.8 | +0.33 | -6.3 | +0.33 | +9.6 | +0.33 | +9.2 | -0.67 |
| YersiniaPestis | aggressive | +21.0 | +1.33 | +2.6 | +0.67 | -18.6 | +0.33 | +27.4 | +0.67 | +3.0 | +0.67 |
| Ascendant | aggressive | +22.4 | +0.67 | +18.7 | +0.33 | +10.9 | +0.33 | +9.3 | -0.33 | +12.6 | +0.67 |
Cross-opponent aggregation (the verdict layer, damage + wins)
| arm | metric | mean Δ | spread (SD) | SE | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|---|---|---|---|---|---|---|---|---|---|---|
len6 |
damage | +9.63 | 18.97 | 4.90 | [−0.88, +20.13] | 11/15 | 0.1185 | 0.06976 | 0.09384 | 13.72 |
len6 |
wins | +0.49 | 0.58 | 0.15 | [+0.17, +0.81] | 8/9 | 0.03906 | 0.007812 | 0.01269 | 0.42 |
len16 |
damage | +2.18 | 16.62 | 4.29 | [−7.02, +11.38] | 9/15 | 0.6072 | 0.6142 | 0.712 | 12.02 |
len16 |
wins | +0.27 | 0.42 | 0.11 | [+0.03, +0.50] | 8/10 | 0.1094 | 0.04688 | 0.01852 | 0.30 |
depth100 |
damage | −1.24 | 26.52 | 6.85 | [−15.93, +13.45] | 6/15 | 0.6072 | 0.8677 | 0.5137 | 19.19 |
depth100 |
wins | +0.27 | 0.42 | 0.11 | [+0.03, +0.50] | 8/10 | 0.1094 | 0.04688 | 0.046 | 0.30 |
rad_offset |
damage | +6.71 | 17.18 | 4.44 | [−2.80, +16.23] | 10/15 | 0.3018 | 0.1522 | 0.1323 | 12.43 |
rad_offset |
wins | +0.42 | 0.54 | 0.14 | [+0.12, +0.72] | 9/11 | 0.06543 | 0.01465 | 0.01424 | 0.39 |
rad_scale |
damage | +4.80 | 21.10 | 5.45 | [−6.89, +16.49] | 8/15 | 1 | 0.4095 | 0.6293 | 15.27 |
rad_scale |
wins | +0.11 | 0.51 | 0.13 | [−0.17, +0.40] | 7/12 | 0.7744 | 0.5107 | 0.289 | 0.37 |
The pre-registered verdict (verbatim)
| rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) |
|---|---|---|---|---|---|---|---|
| 1 | len6 |
+0.49 | +9.6 | 8/9 p=0.03906 | 11/15 p=0.1185 | BETTER | BETTER |
| 2 | rad_offset |
+0.42 | +6.7 | 9/11 p=0.06543 | 10/15 p=0.3018 | not distinguishable | not distinguishable |
| 3 | len16 |
+0.27 | +2.2 | 8/10 p=0.1094 | 9/15 p=0.6072 | not distinguishable | not distinguishable |
| 4 | depth100 |
+0.27 | −1.2 | 8/10 p=0.1094 | 6/15 p=0.6072 | not distinguishable | not distinguishable |
| 5 | rad_scale |
+0.11 | +4.8 | 7/12 p=0.7744 | 8/15 p=1 | not distinguishable | not distinguishable |
Reference pattern: 101.1 dmg/run, 1.27 wins/run, 13.27% incoming, 428 px.
Reading and the single red flag
len6is the only arm that beatspatternby the campaign rule — on the WINS primary: +0.49 wins/run, sign test 8/9 p=0.0391, sign-flip p=0.0078, wins CI [+0.17, +0.81] excluding 0 — while damage is nominally up (+9.6, but inside the 13.7 MDE and its CI includes 0). A 6-tick match key is broadly positive (10/15 opponents non-negative, onlyCoriantumr−0.33 negative).- RED FLAG (MEASURED): all five arms are nominally positive on wins
(+0.49 / +0.27 / +0.27 / +0.42 / +0.11), including
rad_offset, which by job j99's exact structural proof cannot change the aim bearing and whose live effect was measured as a tie (+0.28 pp, p=0.62). The referencepatternalso measures 42.2% wins here vs 50.4% in j121 on the same 15-panel (its damage is unchanged at 101–99.6), so the leading explanation for the all-arms-positive pattern is a session-level low reference, not five independent gun wins. The within-session paired deltas remain valid, butlen6's win is exactly the kind of single-session effect that must replicate on a second, wider panel. - The two shape knobs behave plausibly.
len16(fewer, more specific matches → more linear fallback) is damage-flat;depth100(shallower search) is damage-flat at a larger MDE. Neither is a detectable win.
Pre-registered predictions — scorecard
| # | prediction | outcome |
|---|---|---|
| 1 | no arm beats pattern on BOTH primaries |
CORRECT (len6 beats on wins only; damage CI includes 0) |
| 2 | len16 flat/worse on damage, wins flat |
PARTLY WRONG: damage flat (+2.2), wins nominally +0.27 |
| 3 | len6 flat on damage and wins |
WRONG: wins +0.49 significant; damage +9.6 sub-MDE |
| 4 | depth100 not distinguishable |
CORRECT (damage −1.2), though wins nominally +0.27 |
| 5 | radial controls not distinguishable / bearing-invariant | CORRECT: rad_scale flat; rad_offset nominal +0.42 wins but sign p=0.065 (not significant) |
| 6 | any surprise is damage-only without wins | WRONG in direction: the surprise is wins-with-nominal-damage |
Batch 2 — higher-power confirmation of len6 (+ the structural control)
Why we run it (task trigger, with one recorded deviation). The task's Batch-2
trigger is: an arm beats the reference with the winning primary's CI excluding
0 AND sign-flip p<0.05. len6 satisfies it on wins (CI [+0.17, +0.81];
p(sign-flip)=0.0078). Deviation, recorded: the Batch-1 pre-registration added a
stricter clause — "its damage CI excludes 0" — which len6's damage
([−0.88, +20.13]) does not meet. We run Batch 2 anyway because (a) the task rule
keys on the winning primary and (b) the all-arms-positive red flag means a
replication is the only way to tell a real len6 effect from a session-level
reference anomaly. This is a deviation from the letter of my own pre-registration,
not from the task's criterion.
Design. One frozen binary, four env-only arms, the wider panel
tools/ab/panel_gun_b2.txt (33 opponents, superset of Batch 1's 15),
3 runs × 3 rounds = 396 battles, --conc 6, --wait-arena. Arm file:
tools/ab/arms_gun_b4.txt. Reference: pattern. Movement pinned
TR_MOVEMENT=strafe in every arm.
| # | arm | env over the pin | role |
|---|---|---|---|
| 1 | pattern |
(none) | reference |
| 2 | len6 |
TR_PATTERN_LEN=6 |
the candidate |
| 3 | len16 |
TR_PATTERN_LEN=16 |
dose-response on key length |
| 4 | rad_offset |
TR_PATTERN_RAD_OFFSET=-20 |
structural positive control (bearing-invariant) |
Pre-registered prediction (written BEFORE the battles finished):
len6's +0.49 wins does NOT replicate. On 33 opponents its wins delta collapses toward the MDE and the sign test (or the sign-flip test) fails. Rationale: all five Batch-1 arms were positive on wins including the bearing-invariantrad_offset, and the reference was 8 points low on win rate versus j121's same-panel measurement — a single-session low reference, not alen6effect.rad_offsetis not a detectable win; if it, too, replicates a wins gain, the wins axis is being moved by something other than the aim bearing and the whole Batch-1 wins signal is an artifact.len16stays within its MDE on both primaries.- Damage stays within the MDE for every arm.
- If
len6DOES replicate (winning-metric CI excludes 0, sign-flip p<0.05, damage not detectably down), it is recorded as the campaign's champion candidate — and the shipped-default flip remains a separate, explicit decision that is NOT taken in this job.
Outcome — Batch 2
MEASURED. Session /tmp/ab/j123_b2, commit
1b59b6581075cf4019773dd74e6db9eecf2fcb02, frozen binary sha256 8155c39fbc6a…,
33 opponents × 4 arms × 3 runs × 3 rounds = 396 battles, 0 failed, 0 never
started, 0 liveness exclusions. Arena serialization held: the job waited at
the start of the run for j124's spinner fleet (/tmp/ab/j124_spinner) and only
started once it was gone; no foreign battle ran concurrently with the measurement.
Every arm declared TR_MOVEMENT=strafe; TR_PATTERN_LEN / TR_PATTERN_RAD_OFFSET
appear verbatim in each overriding arm's raw-env report.
Pooled dashboard (all valid runs — explanation only, NOT the verdict)
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|---|---|---|---|---|---|---|---|---|
pattern (REF) |
99 | 129.5 | 124.4 | 2.11 | 209/297 | 70.4% | 12.54% | 428 |
len6 |
99 | 131.4 | 121.0 | 2.07 | 205/297 | 69.0% | 12.04% | 423 |
len16 |
99 | 125.8 | 121.4 | 2.04 | 202/297 | 68.0% | 12.36% | 426 |
rad_offset |
99 | 126.9 | 132.4 | 1.91 | 189/297 | 63.6% | 13.03% | 425 |
Per-opponent paired deltas (arm − pattern), damage and wins
| opponent | style | Δdmg_len6 | Δwins_len6 | Δdmg_len16 | Δwins_len16 | Δdmg_rad_offset | Δwins_rad_offset |
|---|---|---|---|---|---|---|---|
| Aristocles | dodger | +0.7 | +0.00 | -12.7 | +0.00 | -5.3 | +0.00 |
| Ascendant | other | +13.8 | -0.67 | -23.1 | -1.33 | +0.7 | -1.00 |
| BlitzBat | regular | -2.6 | +0.00 | +4.4 | +0.33 | -12.0 | +0.33 |
| BrokenSword | other | -17.7 | +0.00 | +6.0 | +0.33 | -15.9 | +0.33 |
| CassiusClay | dodger | +3.7 | +0.67 | -7.5 | +0.00 | +32.4 | -0.33 |
| Cigaret | dodger | -26.3 | -2.00 | -5.1 | -0.33 | -29.1 | -1.67 |
| CigaretBH | dodger | +17.6 | +0.33 | +1.1 | +0.33 | +18.2 | +0.00 |
| Coriantumr | other | -3.9 | +0.00 | -14.1 | -0.33 | +3.9 | +0.00 |
| Diamond | dodger | -5.2 | -0.33 | -8.8 | -0.67 | -13.7 | -0.67 |
| DiamondHawk | other | +3.2 | -0.33 | -3.9 | +0.00 | -6.6 | -0.33 |
| DiamondStealer | regular | +29.5 | +1.00 | +27.4 | +0.33 | +16.5 | +0.33 |
| Dookious | dodger | +16.6 | +0.33 | +14.5 | +0.33 | -2.1 | +0.00 |
| DrussGT | other | -13.2 | +0.00 | +20.7 | +0.33 | -15.9 | -0.33 |
| FloodMini | regular | +1.0 | +0.00 | +12.0 | +0.00 | +20.7 | +0.00 |
| GresSuffurd | dodger | +10.8 | +0.33 | -11.1 | +0.00 | +4.8 | +0.00 |
| HawkOnFire | regular | +24.9 | +0.67 | +26.3 | -0.33 | -2.6 | +0.67 |
| Jen | dodger | -36.3 | +0.00 | -27.0 | +0.00 | -16.1 | -0.33 |
| Komarious | dodger | +3.3 | +0.33 | +1.2 | +0.00 | -11.4 | +0.33 |
| KurtWaveSurfer | dodger | -18.9 | -0.67 | -35.9 | +0.00 | -7.8 | +0.00 |
| LightningBug | other | -15.3 | +0.00 | -10.7 | +0.00 | -20.7 | +0.00 |
| LionWWSVMvoid | dodger | +19.1 | +0.00 | -1.3 | +0.00 | -1.7 | +0.00 |
| Lukious | dodger | +17.6 | +0.00 | +1.8 | +0.33 | +19.6 | +0.33 |
| PatternRobot | regular | +21.3 | +0.00 | +6.7 | +0.00 | -1.3 | +0.00 |
| Phoenix | other | -8.2 | +0.33 | -29.5 | -0.33 | -4.4 | -0.67 |
| RetroGirl | other | -29.6 | -0.33 | -6.7 | +0.00 | -17.0 | -1.67 |
| RougeDC | dodger | +7.5 | +0.00 | -1.9 | -0.33 | +2.3 | +0.00 |
| Shadow | other | +3.1 | -0.67 | -10.9 | -0.67 | -10.4 | -0.67 |
| SpinBot | spinner | -14.7 | +0.00 | -9.5 | +0.00 | -3.5 | +0.00 |
| TripHammer | other | +0.9 | -0.33 | -7.6 | -0.33 | -16.8 | -1.00 |
| WallAvoider | regular | +35.7 | -0.33 | +3.3 | +0.33 | -5.6 | -1.00 |
| WaveSurferGF | dodger | +12.5 | +0.33 | -16.5 | -0.67 | +26.3 | +0.33 |
| WaveSurferPG | dodger | -1.2 | -0.67 | +0.3 | +0.33 | -22.7 | +0.00 |
| YersiniaPestis | other | +10.8 | +0.67 | -4.7 | +0.00 | +8.6 | +0.33 |
Cross-opponent aggregation (the verdict layer, damage + wins)
| arm | metric | mean Δ | spread (SD) | SE | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|---|---|---|---|---|---|---|---|---|---|---|
len6 |
damage | +1.83 | 17.09 | 2.98 | [−4.00, +7.67] | 20/33 | 0.2962 | 0.5417 | 0.4859 | 8.34 |
len6 |
wins | −0.04 | 0.54 | 0.09 | [−0.22, +0.14] | 10/20 | 1 | 0.7631 | 0.8806 | 0.26 |
len16 |
damage | −3.72 | 14.45 | 2.52 | [−8.65, +1.22] | 13/33 | 0.2962 | 0.1494 | 0.09657 | 7.05 |
len16 |
wins | −0.07 | 0.38 | 0.07 | [−0.20, +0.06] | 9/19 | 1 | 0.379 | 0.5303 | 0.19 |
rad_offset |
damage | −2.69 | 14.72 | 2.56 | [−7.72, +2.33] | 11/33 | 0.08014 | 0.2997 | 0.1952 | 7.18 |
rad_offset |
wins | −0.20 | 0.56 | 0.10 | [−0.39, −0.01] | 8/20 | 0.5034 | 0.05926 | 0.05875 | 0.28 |
The pre-registered verdict (verbatim)
| rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) |
|---|---|---|---|---|---|---|---|
| 1 | len6 |
−0.04 | +1.8 | 10/20 p=1 | 20/33 p=0.2962 | not distinguishable | not distinguishable |
| 2 | len16 |
−0.07 | −3.7 | 9/19 p=1 | 13/33 p=0.2962 | not distinguishable | not distinguishable |
| 3 | rad_offset |
−0.20 | −2.7 | 8/20 p=0.5034 | 11/33 p=0.08014 | not distinguishable | not distinguishable |
Reference pattern: 129.5 dmg/run, 2.11 wins/run, 12.54% incoming, 428 px.
Reading
len6's Batch-1 wins win did not replicate. On 33 opponents it collapses from +0.49 wins/run to −0.04 wins/run (10/20, p=1) and from +9.6 to +1.8 dmg/run (20/33, p=0.30, CI [−4.00, +7.67], MDE 8.34). Both primaries are squarely inside the MDE. Verdict: NOT DISTINGUISHABLE.- The Batch-1 red flag is explained. All five Batch-1 arms were wins-positive
including the bearing-invariant
rad_offset; hererad_offsetturns negative on both primaries (−0.20 wins, −2.7 dmg) and the other two arms are wins-flat. The Batch-1 effect was a session-level low reference:patternmeasured 42.2% win rate on 15 opponents in Batch 1, versus 70.4% here on its 33-opponent superset and 63.3% in j121 Batch 2 on the same 33-panel. A single-session paired comparison is internally valid, but a wins-only signal in one session is not evidence until it replicates — which is exactly what Batch 2 tested. - The dose-response is flat and slightly negative for longer keys.
len16is −3.7 dmg/run and −0.07 wins/run (damage sign test p=0.30; if anything it drifts toward the linear fallback as predicted in Batch 1). - No shipped default was touched; this job is a measurement.
Pre-registered predictions — scorecard
| # | prediction | outcome |
|---|---|---|
| 1 | len6's +0.49 wins does not replicate; delta collapses toward the MDE |
CORRECT (+0.49 → −0.04 wins, p=1; damage +9.6 → +1.8) |
| 2 | rad_offset is not a detectable win |
CORRECT (it is nominally negative: −0.20 wins, −2.7 dmg) |
| 3 | len16 stays within its MDE on both primaries |
CORRECT (−3.7 dmg vs MDE 7.05; −0.07 wins) |
| 4 | damage stays within the MDE for every arm | CORRECT |
| 5 | if len6 replicates it is the champion candidate |
not invoked |
Phase 2 — final direct answer
MEASURED (two sessions, 15 then 33 opponents, 666 battles, 6 frozen-binary arms + the reference): NO — the shipped
Patterncannot be improved by tuning its own match-length or history parameters at the resolution we can measure. The incumbent is already tuned on this axis.
- The one Batch-1 signal (
len6, +0.49 wins/run on 15 opponents, sign p=0.039, sign-flip p=0.008) reversed to a wash on 33 opponents (−0.04 wins/run, p=1; +1.8 dmg/run, p=0.30). It was a session-level low reference, not a gun effect. - On damage/run and round wins, the best Batch-2 challenger is inside its MDE
on both primaries (len6: +1.8 dmg/run vs MDE 8.34, −0.04 wins/run vs MDE
0.26). Longer keys (
len16) and a shallower search (depth100) do not help. - The two pre-existing radial knobs remain what j99 measured: bearing-invariant
and not winners; in Batch 2
rad_offsetis nominally negative. - No shipped default changed. Exposing the knobs (
TR_PATTERN_LEN,TR_PATTERN_DEPTH) does not change the default path — proven byte-identical before any battle.
(MEASURED vs INFERRED. The knob exposure, the byte-identity parity dump, the 270- and 396-battle sessions, every per-opponent delta, the CIs, MDEs, sign tests and sign-flip tests above are all MEASURED. The "incumbent is already tuned" reading is a MEASURED null at these resolutions; any claim that no setting whatsoever could ever help would be INFERRED and is deliberately not made.)
What to try next (rewritten AFTER Phase 2)
The phase-1 conclusion stands and is now joined by the phase-2 one:
- The match-shape axis is CLOSED. Match-key lengths 6 and 16 and a
shallow (100-tick) history search do not beat the shipped defaults (10,
full 500). Do not spend another batch on
TR_PATTERN_LEN/TR_PATTERN_DEPTHunless a specific mechanism is named; the MDE is ~8 dmg/run and ~0.26 wins/run, so only a very large effect could even be seen. TR_MOVEMENT=strafe+ defaultPatternis the measured optimum of the tested design space, and Pattern is now measured-tuned on its own parameters as well.- The named open axes remain the ones phase 1 named: a different information source for the lead (not another knob on BitBrain/TMHorizon, both measured negative/neutral), and the owner's spinner-specific claim (j124 is now on it). The lead-amplitude axis and the radial axis stay dead.
- Session-level reference variance is large and must be respected. The same
frozen reference on the same movement measured 42.2% wins (Batch 1, 15
opponents) and 70.4% wins (Batch 2, 33 opponents, superset). Any wins-only
signal from one session must be replicated before it is believed — this
job's Batch-1
len6is the second campaign example (after j121'sbitbraindamage hint) of a single-session wobble dying on replication.
Session log (Phase 2 rows appended)
| session | commit | battles | arms | verdict |
|---|---|---|---|---|
/tmp/ab/j123_b1 |
2a98aba |
270 (0 failed, 0 never started, 0 excluded) | pattern, len6, len16, depth100, rad_offset, rad_scale | len6 the only rule-2 win (wins +0.49, sign p=0.039); red flag: all 5 arms wins-positive incl. the bearing-invariant rad_offset |
/tmp/ab/j123_b2 |
1b59b65 |
396 (0 failed, 0 never started, 0 excluded) | pattern, len6, len16, rad_offset | len6 does not replicate (−0.04 wins, p=1; +1.8 dmg, p=0.30); all arms not distinguishable; the incumbent is already tuned |