1156 lines
73 KiB
Markdown
1156 lines
73 KiB
Markdown
# Gun campaign — ledger
|
||
|
||
**Goal (owner's mandate, 2026-09-26 overnight):** the movement campaign produced
|
||
a replicated champion (`TR_MOVEMENT=strafe`, `docs/movement_campaign.md`) which a
|
||
parallel job is shipping. *"Continue until you found an amazing movement. When
|
||
found do the same over for a gun."* This file is the GUN campaign's single source
|
||
of truth: later jobs **append** a `## Batch N` section and never edit an earlier
|
||
one (a wrong earlier number gets a correction line, not a rewrite). This file is
|
||
the gun successor to `docs/movement_campaign.md`; read that first, then this.
|
||
|
||
---
|
||
|
||
## OUTCOME (read this first)
|
||
|
||
**The shipped gun is UNCHANGED: `onlyPattern`** — the rack admits `Pattern` only
|
||
(`common_libs/gun_harness/selector.nim`, commit `e0666a5`). **Nothing beat it.**
|
||
The gun design space explored here is **CLOSED on evidence**, not on belief.
|
||
|
||
**Headline (MEASURED).** Four frozen-binary tournament sessions on the frozen
|
||
**15- and 33-opponent panels**: **1,233 battles** (phase 1: 270 + 297; phase 2:
|
||
270 + 396), **0 failed, 0 never-started, 0 liveness exclusions**. Each session is
|
||
ONE binary built from `git archive HEAD`; arms differ only by env; movement is
|
||
pinned `TR_MOVEMENT=strafe` in every arm, so every delta is a pure gun delta.
|
||
With the earlier 16-gun selector campaign (≈660 battles) the gun ledger now
|
||
stands on **≈1,900 measured battles**. What was **detectably worse**: `knn`
|
||
(−39.3 dmg/run, 0/15 opponents; −0.62 wins/run, p=0.012) and, on the wider
|
||
33-opponent panel, `TMHorizon` (−10.4 dmg/run, sign p=0.035). What was
|
||
**not distinguishable** (inside the MDE, no verdict): `bitbrain`,
|
||
`tmhorizon`@15, `rack_pk`, `rack_pt`, `len6`, `len16`, `depth100`, the two radial
|
||
controls.
|
||
|
||
### CLOSED axes (one-line reason each)
|
||
|
||
| axis | reason it is closed |
|
||
|---|---|
|
||
| other single guns | `knn` detectably worse; `bitbrain` a wash measured 3× (−3.6 / +10.3 / +1.8 dmg/run, all inside their MDEs); `TMHorizon` worse on 33 (`docs/gauntlet_bitbrain_vs_pattern.md`, Batches 1–2) |
|
||
| 2-gun selector racks | `rack_pk` nominally below `pattern` on wins; `rack_pt` == `tmhorizon`-alone → the selector just picks the corrector (`## Batch 1`) |
|
||
| selector mechanics | the virtual-fitness selector is negative value at every rack size tested (16, lean8, lean6, pairs); it never manufactured a win (`docs/selector_negative_value.md`) |
|
||
| **gun allocation (share)** | **forced `TR_RACK_SHARE` schedules (50/50, 70/30; PB/PT/PK) never beat `pattern` by the pre-registered rule; the only detectable effect is negative (50/50 KNN, −11 dmg/run, p=0.042); the observed ~66/34 split was already right (`## Allocation`) |
|
||
| lead amplitude | a full gain sweep (1.0/1.5/2.0/3.0) makes Pattern strictly worse at every range band — the lever is lead *information*, not amplitude (`docs/bitbrain_campaign.md` §0.3.2) |
|
||
| Pattern's own match-length / history params | `len6`'s single-session +0.49 wins (p=0.039 on 15) reversed to a wash on 33 (−0.04, p=1); `len16`/`depth100` flat; two batches, no replication (`## Phase 2`) |
|
||
| radial knobs (`TR_PATTERN_RAD_*`) | bearing-invariant **by construction** (job j99 proved `bmPath` is a structural no-op and scale/offset cannot change the aim bearing); in Batch 2 `rad_offset` is nominally *negative* |
|
||
|
||
### THE ONE OPEN AXIS
|
||
|
||
**A genuinely new source of lead INFORMATION** — not another knob on
|
||
BitBrain/TMHorizon (both measured neutral→negative), not amplitude (dead), not
|
||
radial (bearing-invariant). The standing mechanism is `docs/bitbrain_campaign.md`
|
||
§0.3.3: at 450+ px Pattern's own lead correlation with the required lead is only
|
||
**0.165**, so it is adding variance to a nearly uninformative signal.
|
||
|
||
**Honest caveat (MEASURED, `docs/state_window_gate.md`).** The best surviving
|
||
idea from the state line — a *single* wave-relative state at Q=4 — **does carry
|
||
signal at coarse resolution**: it predicts the miss-offset *bin* on held-out
|
||
battles at 0.4094 (vs 0.2348 majority), but those bins are **36/42/60 px wide =
|
||
4.58/5.34/7.63° at 450 px**, and its implied hit probability is **still the base
|
||
rate (~0.09)** — i.e. it is informative about **which side** the miss falls on,
|
||
not about **whether** the shot hits. The live hit window is
|
||
`atan(18/450) = 2.29°` half-width; **the measured signal is at ~4.6–7.6°, not at
|
||
the ~2° the 450 px hit window needs.** A temporal window does not fix this (the
|
||
window gate is negative; long contexts never recur). So the open axis is a *new
|
||
information source that resolves the arrival offset to better than ~2°*, which
|
||
does not exist yet.
|
||
|
||
**Cost estimate for trying it (INFERRED).** Build and *offline-gate* the new
|
||
information source first — require it to beat Pattern's 0.165 lead correlation
|
||
and, ideally, predict the arrival offset to <2.29° on held-out shots (days of
|
||
work; the offline instruments exist, e.g. `common_libs/tests/state_window_gate.py`).
|
||
Only then spend a live batch: one frozen binary, 6 env-only arms × 15 opponents ×
|
||
3 runs × 3 rounds = **270 battles ≈ 1–2 h wall** at conc 6 with `--wait-arena`
|
||
(exact command in *How to run a batch*). Do **not** open with a live batch.
|
||
|
||
### What would change our mind
|
||
|
||
A named new information source that (1) offline predicts the arrival offset to
|
||
better than the **~2.29° hit half-window at 450 px** on held-out battles **and**
|
||
(2) in a live frozen-panel batch beats `pattern` on **wins/run** with the winning
|
||
metric's CI excluding 0 and sign-flip p<0.05 while **damage is not detectably
|
||
down**. A damage-only gain without wins is the `ring`-mover trap and does **not**
|
||
count.
|
||
|
||
---
|
||
|
||
## What changed tonight (2026-09-26)
|
||
|
||
* **SHIPPED:** nothing in the gun. The rack is still `onlyPattern`.
|
||
* **NOT shipped, and why:** every candidate (other gun, small rack, selector
|
||
mechanic, match-length/history parameter, radial knob) failed to beat `pattern`
|
||
beyond the MDE on the frozen panels. Phase 2's one signal (`len6`, +0.49
|
||
wins/run on 15 opponents) **did not replicate** on 33 opponents. The only
|
||
shipped change of the night is the **movement** default (`tfil -> strafe`, see
|
||
`docs/movement_campaign.md`); the gun was left untouched.
|
||
* **Revert:** no gun revert is needed (no gun default changed). Movement:
|
||
`TR_MOVEMENT=tfil`.
|
||
* **Reproduce the key gun evidence (one command + the analyzer)** — phase 2,
|
||
Batch 2, the session that killed `len6`:
|
||
```sh
|
||
TOURNAMENT_NIMCACHE=/tmp/nc_j123 \
|
||
tools/ab/tournament_run.sh \
|
||
--arms tools/ab/arms_gun_b4.txt \
|
||
--panel tools/ab/panel_gun_b2.txt \
|
||
--runs 3 --rounds 3 --conc 6 --wait-arena 45 \
|
||
--reference pattern \
|
||
--outdir /tmp/ab/j123_b2
|
||
python3 tools/ab/tournament_analyze.py /tmp/ab/j123_b2 --reference pattern
|
||
```
|
||
|
||
---
|
||
|
||
## 0. The question
|
||
|
||
The shipped rack admits **`Pattern` only** (`onlyPattern`). That decision was
|
||
taken because the virtual-fitness selector measured **negative value** at every
|
||
rack size tested, and Pattern is the best single gun by the DrussGT hit-rate
|
||
table (`docs/selector_negative_value.md`, `docs/gun_rack_analysis.md`, commit
|
||
`e0666a5`). But every one of those measurements — like every pre-campaign
|
||
movement claim — is **DrussGT-heavy**, and the movement campaign's standing
|
||
lesson is that a one-opponent result is not a result.
|
||
|
||
So Batch 1 asks the direct question, on a frozen 15-opponent panel:
|
||
|
||
> **Does any rack gun, or any small gun configuration, beat the shipped
|
||
> `Pattern` on damage/run AND round wins across 15 opponents — or is
|
||
> `onlyPattern` CONFIRMED rather than merely assumed?**
|
||
|
||
A **null is a successful outcome here**: it upgrades `onlyPattern` from an
|
||
assumption to a measured verdict across the panel, which is itself valuable.
|
||
|
||
---
|
||
|
||
## 1. Protocol (identical to the movement campaign; the instrument is reused verbatim)
|
||
|
||
`tools/ab/tournament_run.sh` + `tools/ab/tournament_analyze.py` +
|
||
`tools/ab/panel_movement.txt` (frozen panel), reused **unchanged**. The unit of
|
||
evidence is the **number of opponents**, not the number of runs.
|
||
|
||
| Element | Rule |
|
||
|---|---|
|
||
| Subject | ONE frozen binary built from `git archive HEAD` (`tournament_run.sh` records commit + binary sha256 in `session.json`) |
|
||
| Arms | env dicts only — **no per-arm rebuild, ever**; the arm file is a committed file (`tools/ab/arms_gun_b1.txt`) |
|
||
| Panel | the **frozen** `tools/ab/panel_movement.txt` (15 opponents). Adding/removing an opponent starts a new batch number |
|
||
| Pairing | per opponent: average the arm's runs, subtract the reference's average for that same opponent → one delta per opponent; then aggregate |
|
||
| Isolation | per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir |
|
||
| Serialization | **one battle fleet at a time** via `--wait-arena`; never a broad `pkill robocode_shim` |
|
||
| Liveness | every declared env token must appear verbatim in OUR bot's own raw-env report, else the run is excluded and named |
|
||
| Never shipped | this is a measurement campaign: `git status` clean, shipped defaults untouched, `.gitignore` untouched |
|
||
|
||
> **Movement-default flip note (MEASURED).** During this job a parallel job
|
||
> (j120, commit `3fd6db9`) flipped the shipped `TR_MOVEMENT` default from `tfil`
|
||
> to `strafe`. Batch 1 built before the flip, Batch 2 after. **Every arm in both
|
||
> batches declares `TR_MOVEMENT=strafe`**, so the effective movement is identical
|
||
> and the two batches are directly comparable — the pin did its job. The reference
|
||
> rack lines in the boot reports read `PATTERN` for `pattern`, and the active 1v1
|
||
> rack reads `BITBRAIN` / `TMHORIZON` / `KNN` for the overridden arms.
|
||
|
||
### Movement is PINNED in every arm: `TR_MOVEMENT=strafe`
|
||
|
||
**Reason (hard rule).** A movement-default flip is landing from job j120 during
|
||
this run, so the default engine can change underneath us. Every arm therefore
|
||
sets **`TR_MOVEMENT=strafe` explicitly**. This (a) removes movement as a
|
||
confound, (b) makes all arms share exactly one movement, so every delta is a pure
|
||
**gun** delta, and (c) satisfies the analyzer's liveness rule for every arm
|
||
(each declared token must appear verbatim; an *undeclared* leaked `TR_MOVEMENT`
|
||
is fatal, but here every arm declares it). The reference arm is therefore the
|
||
shipped **gun** rack under the pinned movement, not under the (mutable) default.
|
||
|
||
---
|
||
|
||
## 2. Pre-registered decision rules (fixed BEFORE Batch 1 ran)
|
||
|
||
1. **Primary metrics:** **damage/run** and **ROUND WINS**. Secondary/explanation
|
||
only: damage taken/run, incoming hit rate, mean distance. **Never hit rate
|
||
alone** — that trap has inverted six verdicts in this project.
|
||
2. **BETTER than the reference** iff one primary metric is up with a
|
||
cross-opponent **sign test p < 0.05** while the other does **not** go down;
|
||
the mirror image for **WORSE**. Anything else is **NOT DISTINGUISHABLE** (a
|
||
real answer, not a failure). Both the strict reading (the other metric's mean
|
||
delta `>= 0`) and the substantive reading (not *detectably* down: not
|
||
significant **and** smaller than that metric's MDE) are printed.
|
||
3. **A verdict must survive the between-opponent spread:** the pooled mean delta
|
||
is reported with the SD across opponents, its SE, a 95% CI, and the **MDE**
|
||
(α=0.05 two-sided, 80% power). An effect smaller than the MDE is reported as
|
||
*not detectable* — never as *absent*, never as a win.
|
||
4. **Somewhere to stop:** if no arm beats the shipped `pattern` by rule 2 in
|
||
Batch 1 **and** no arm shows a `>= +MDE` damage gain with p<0.10, then the
|
||
gun stage's first phase is closed with *"the shipped `onlyPattern` rack is the
|
||
best gun configuration we have measured across the panel"* — that is a
|
||
**successful** outcome, and the campaign moves to a named next axis rather
|
||
than inventing more arms. See *What would make us stop*.
|
||
5. **No promotion off a single metric, a single opponent, or a single run.**
|
||
A change that wins damage by losing wins (or vice-versa) is not a win.
|
||
6. Every batch is shot with a **pre-registered prediction** stated in its section
|
||
*before* the battles finish; a prediction that turns out wrong is recorded as
|
||
wrong.
|
||
|
||
---
|
||
|
||
## 3. Stage 0 — what we already know (given, not re-derived)
|
||
|
||
| fact | value | source |
|
||
|---|---|---|
|
||
| shipped rack | `onlyPattern` (id 5), all other guns `off` | `common_libs/gun_harness/selector.nim` |
|
||
| why | selector measured negative value at every rack size (16, lean8, lean6, pairPC/PK/PL) and 10/10 adversaries | `docs/selector_negative_value.md`, `docs/gun_rack_analysis.md` |
|
||
| Pattern hit rate vs DrussGT | ~10.8% | given |
|
||
| KNN / Linear / Circular / WallBounce / GF vs DrussGT | 5.6 / 3.0 / 2.9 / 2.7 / 2.1 % | given |
|
||
| BitBrain when idle | == Pattern (30 runs/arm: 97/210 vs 97/210 wins) | `docs/bitbrain_vs_tmhorizon_ab.md` |
|
||
| BitBrain across 32 legacy opponents | neutral: **-3.6 dmg/run**, +0.02 wins/run, p=0.53/0.91 | `docs/gauntlet_bitbrain_vs_pattern.md` |
|
||
| lead **amplitude** | dead axis — full gain sweep {0, 0.25-1.0, 1.0, 1.25, 1.5, 2.0} run, nothing beats Pattern | `docs/bitbrain_campaign.md` |
|
||
| offline prediction checks | **veto-only** — may kill a design, never select one | `docs/offline_harness_trust.md` |
|
||
| spinner-specific claim | **UNTESTED**: no constant-turn spinner exists in the legacy roster (a note for a later job, not built here unless nearly free) | `docs/gauntlet_bitbrain_vs_pattern.md` |
|
||
|
||
**Movement context (why the pin):** the movement champion is `TR_MOVEMENT=strafe`
|
||
at its current defaults (range 325, tol 25, tilt 15/0.10); it replicated its win
|
||
over the shipped `tfil` in four independent sessions (Batch 4: 52.6% vs 38.5%
|
||
round-win rate). Movement is held constant at that champion for every gun arm.
|
||
|
||
---
|
||
|
||
## Batch 1 — does anything beat the shipped `Pattern`?
|
||
|
||
**Design.** One frozen binary, six env-only arms, one frozen panel
|
||
(`tools/ab/panel_movement.txt`, 15 opponents: 5 dodger, 3 pattern, 1
|
||
wall-follower, 1 corner-camper, 1 spinner, 1 rammer, 1 brawler, 2 aggressive
|
||
megas), **3 runs × 3 rounds per (opponent, arm) = 270 battles**. Arm file:
|
||
`tools/ab/arms_gun_b1.txt`. Reference: `pattern`.
|
||
|
||
| # | arm | env | what it isolates |
|
||
|---|---|---|---|
|
||
| 1 | `pattern` | `TR_MOVEMENT=strafe` | the arm to beat (shipped `onlyPattern` rack) |
|
||
| 2 | `bitbrain` | `+ TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay` | BitBrain-only, learned gain (panel replication of the 32-opp gauntlet at the pinned-movement standard) |
|
||
| 3 | `tmhorizon` | `+ TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both` | TMHorizon-only corrector (predicted to lose; never panel-tested) |
|
||
| 4 | `knn` | `+ TR_RACK_PATTERN=off TR_RACK_KNN=both` | best non-Pattern single gun (never panel-tested) |
|
||
| 5 | `rack_pk` | `+ TR_RACK_KNN=both` | Pattern + KNN, **selector ON** (different-family hedge) |
|
||
| 6 | `rack_pt` | `+ TR_RACK_TMHORIZON=both` | Pattern + TMHorizon, **selector ON** (corrector-family hedge) |
|
||
|
||
*(`+` = the pinned `TR_MOVEMENT=strafe` is present in every arm; see §1.)*
|
||
|
||
**Pre-registered prediction (written BEFORE the battles finished):**
|
||
1. **No arm beats `pattern` on BOTH primaries**, and the point estimates sit
|
||
inside the MDE. `onlyPattern` is confirmed across the panel.
|
||
2. `bitbrain` is a **wash** vs `pattern` (replicating the 32-opponent gauntlet:
|
||
±3.6 dmg/run, ≈0 wins) — the DrussGT-only gain-config penalty does not carry.
|
||
3. `tmhorizon` is **WORSE** than `pattern` (its own offline work predicts it
|
||
needs ~80% side accuracy and can reach ~60%); no damage win, no win win.
|
||
4. `knn` is **WORSE** than `pattern` on damage (its DrussGT hit rate is roughly
|
||
half Pattern's) and not better on wins.
|
||
5. The two small racks **do not beat `pattern`**; `rack_pk` lands closer to
|
||
Pattern than `knn`-alone does (the selector at least partially hedges back),
|
||
but the selector's poor ranking keeps it below the reference.
|
||
6. If any surprise exists, it is a **rack arm on damage without wins** — the
|
||
`ring`-mover mirror-image trap — and it will not be read as a win.
|
||
|
||
### Outcome — direct answer
|
||
|
||
**MEASURED.** Session `/tmp/ab/j121_g1`, commit `1d8143a15e4a038c39dfc2153ff518fa003b8612`
|
||
(the pre-registration commit), frozen binary sha256 `aa49a45fec20…`, 15 opponents
|
||
× 6 arms × 3 runs × 3 rounds = **270 battles, 0 failed, 0 never started, 0
|
||
liveness exclusions**. Every arm declared `TR_MOVEMENT=strafe`, verified in each
|
||
bot's own raw-env report; the reference rack line reads `PATTERN` for `pattern`
|
||
and the overridden rack reads `BITBRAIN` / `TMHORIZON` / `KNN` for the others.
|
||
|
||
> **DIRECT ANSWER (MEASURED, n=15 opponents): NO — nothing beats the shipped
|
||
> `Pattern` across the panel on damage/run AND round wins.** Every arm except
|
||
> `knn` is **NOT DISTINGUISHABLE** from `pattern` under both the strict and the
|
||
> substantive reading of the pre-registered rule. The best nominal challenger,
|
||
> `tmhorizon`, is **+0.13 wins/run** (MDE 0.36) and **+8.7 dmg/run** (MDE 13.8)
|
||
> — an effect ~3× smaller than the design can detect — and `bitbrain` is
|
||
> `+0.11` wins / `+10.3` dmg, also inside the MDE. `knn` is **detectably WORSE**
|
||
> on both primaries. The shipped `onlyPattern` rack is therefore **CONFIRMED**
|
||
> rather than merely assumed, at the resolution of this panel. That is a real
|
||
> result, not a failure.
|
||
|
||
#### Pooled dashboard (all valid runs — explanation only, NOT the verdict)
|
||
|
||
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|
||
|---|---:|---:|---:|---:|---:|---:|---:|---:|
|
||
| `pattern` (REF) | 45 | 99.6 | 155.9 | 1.51 | 68/135 | 50.4% | 12.84% | 435 |
|
||
| `bitbrain` | 45 | 109.9 | 151.2 | 1.62 | 73/135 | 54.1% | 12.87% | 427 |
|
||
| `tmhorizon` | 45 | 108.3 | 153.7 | **1.64** | **74/135** | **54.8%** | **12.53%** | 428 |
|
||
| `knn` | 45 | 60.3 | 175.7 | 0.89 | 40/135 | 29.6% | 13.31% | 430 |
|
||
| `rack_pk` | 45 | 107.3 | 161.3 | 1.42 | 64/135 | 47.4% | 13.65% | 432 |
|
||
| `rack_pt` | 45 | 105.3 | 149.9 | 1.64 | 74/135 | 54.8% | 12.90% | 424 |
|
||
|
||
#### Cross-opponent aggregation (the verdict layer, damage + wins)
|
||
|
||
| arm | metric | mean Δ | spread (SD) | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|
||
|---|---|---:|---:|---|---:|---:|---:|---:|---:|
|
||
| `bitbrain` | damage | +10.34 | 16.65 | [+1.12, +19.57] | 11/15 | 0.1185 | 0.03247 | 0.05006 | 12.05 |
|
||
| `bitbrain` | wins | +0.11 | 0.48 | [-0.16, +0.38] | 6/9 | 0.5078 | 0.4922 | 0.4764 | 0.35 |
|
||
| `tmhorizon` | damage | +8.71 | 19.10 | [-1.86, +19.29] | 11/15 | 0.1185 | 0.09918 | 0.1183 | 13.82 |
|
||
| `tmhorizon` | wins | +0.13 | 0.50 | [-0.14, +0.41] | 5/9 | 1.0 | 0.4062 | 0.3118 | 0.36 |
|
||
| `knn` | damage | −39.33 | 23.15 | [−52.15, −26.51] | 0/15 | 6.1e-5 | 6.1e-5 | 0.0007 | 16.75 |
|
||
| `knn` | wins | −0.62 | 0.59 | [−0.95, −0.30] | 1/11 | 0.01172 | 0.001953 | 0.004948 | 0.43 |
|
||
| `rack_pk` | damage | +7.68 | 32.09 | [−10.10, +25.45] | 9/15 | 0.6072 | 0.423 | 0.5895 | 23.21 |
|
||
| `rack_pk` | wins | −0.09 | 0.60 | [−0.42, +0.24] | 6/11 | 1.0 | 0.6738 | 0.5932 | 0.43 |
|
||
| `rack_pt` | damage | +5.69 | 16.50 | [−3.45, +14.83] | 10/15 | 0.3018 | 0.2026 | 0.2681 | 11.93 |
|
||
| `rack_pt` | wins | +0.13 | 0.41 | [−0.10, +0.36] | 6/10 | 0.7539 | 0.3242 | 0.1997 | 0.30 |
|
||
|
||
#### The pre-registered verdict (verbatim)
|
||
|
||
| rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) |
|
||
|---:|---|---:|---:|---|---|---|---|
|
||
| 1 | `tmhorizon` | +0.13 | +8.7 | 5/9 p=1 | 11/15 p=0.1185 | **not distinguishable** | **not distinguishable** |
|
||
| 2 | `rack_pt` | +0.13 | +5.7 | 6/10 p=0.7539 | 10/15 p=0.3018 | **not distinguishable** | **not distinguishable** |
|
||
| 3 | `bitbrain` | +0.11 | +10.3 | 6/9 p=0.5078 | 11/15 p=0.1185 | **not distinguishable** | **not distinguishable** |
|
||
| 4 | `rack_pk` | −0.09 | +7.7 | 6/11 p=1 | 9/15 p=0.6072 | **not distinguishable** | **not distinguishable** |
|
||
| 5 | `knn` | −0.62 | −39.3 | 1/11 p=0.01172 | 0/15 p=6.104e-05 | **WORSE** | **not distinguishable** |
|
||
|
||
Reference `pattern`: 99.6 dmg/run, 1.51 wins/run, 12.84% incoming, 435 px.
|
||
|
||
#### Reading (MEASURED, with the mechanism)
|
||
|
||
* **`knn` is a clean, large loss and is dropped.** It deals **−39 dmg/run** and
|
||
loses **−0.62 wins/run**, positive on **0/15** opponents on damage and 1/11 on
|
||
wins (p=6e-5 / p=0.012). Its incoming hit rate is *no worse* (13.31% vs
|
||
12.84%) — it simply cannot aim: the same movement, a worse gun.
|
||
* **The two small racks do not beat `pattern`.** `rack_pk` (Pattern+KNN,
|
||
selector ON) is *not* the average of Pattern and KNN: it recovers most of KNN's
|
||
damage loss (+7.7 vs KNN-alone's −39.3) but still lands **nominally below**
|
||
Pattern on wins (−0.09). `rack_pt` (Pattern+TMHorizon) is statistically
|
||
indistinguishable from `tmhorizon`-alone — adding Pattern to the rack changes
|
||
nothing, i.e. the selector picks the corrector essentially always. **The
|
||
selector's 16-gun negative verdict reproduces at a 2-gun rack**: it does not
|
||
manufacture a win it did not have.
|
||
* **The only signal is a damage-side hint on the two Pattern-lead correctors**
|
||
(`bitbrain` +10.3, `tmhorizon` +8.7, both 11/15 opponents). It is *below the
|
||
MDE* (12.0 / 13.8) and *not significant by the pre-registered sign test*
|
||
(p=0.1185), so it is **not a win and not a null**. Note the sign-flip test on
|
||
the *mean* is nominally significant for `bitbrain` (p=0.032) while the sign
|
||
test is not — the positive deltas are larger than the negatives — but with 5
|
||
arms compared this is not compelling after multiplicity, and the effect is
|
||
under the MDE. **This is exactly the "damage without wins" pattern this
|
||
project keeps paying for** (the `ring` mover: +31 dmg/run and *fewer* wins);
|
||
here the win deltas (+0.11/+0.13) are themselves sub-MDE.
|
||
* **The reference is stable:** `pattern` under `TR_MOVEMENT=strafe` measures
|
||
99.6 dmg/run and 50.4% round wins here, matching the movement campaign's
|
||
`strafe` reference in Batch 4 (103.5 dmg/run, 52.6%) — same movement, same
|
||
panel, different job.
|
||
|
||
#### Pre-registered predictions — scorecard (an honest count)
|
||
|
||
| # | prediction | outcome |
|
||
|---|---|---|
|
||
| 1 | no arm beats `pattern` on both primaries, point estimates inside the MDE; `onlyPattern` confirmed | **CORRECT** |
|
||
| 2 | `bitbrain` is a wash vs `pattern` (gauntlet: −3.6 dmg, ≈0 wins) | **CORRECT on the verdict, WRONG on the point estimate**: +10.3 dmg here vs −3.6 in the 32-opponent gauntlet; both inside their MDEs. Direction differs, verdict (wash) holds |
|
||
| 3 | `tmhorizon` is WORSE than `pattern` (its own docs predict it loses) | **WRONG**: it is nominally *better* on both primaries (+8.7 dmg, +0.13 wins), though not distinguishable |
|
||
| 4 | `knn` is WORSE than `pattern` on damage, not better on wins | **CORRECT** (detectably on both) |
|
||
| 5 | the small racks do not beat `pattern`; `rack_pk` lands closer to Pattern than `knn`-alone | **CORRECT** |
|
||
| 6 | any surprise is a rack arm on damage without wins | **PARTLY CORRECT**: the nominal damage edge is on the single-gun correctors and the racks; every one of them is win-flat |
|
||
|
||
---
|
||
|
||
## Batch 2 — wider-panel calibration of the damage-side hint
|
||
|
||
The Batch-1 answer is decisive on the primary question (nothing beats `pattern`)
|
||
but the two corrector arms carry a sub-MDE damage hint that rule 3 forbids
|
||
calling either way. Resolving it needs **more opponents, not more runs**, so this
|
||
batch re-runs exactly those two arms plus the reference on a wider panel.
|
||
|
||
**Design.** One frozen binary, three env-only arms, panel
|
||
`tools/ab/panel_gun_b2.txt` = the j117 32-opponent legacy roster **minus Aurora**
|
||
(inert in j117: one fire in 6 battles — it cannot reveal a gun regression)
|
||
**plus SpinBot** = **33 opponents** (a superset of Batch 1's 15). **3 runs × 3
|
||
rounds = 297 battles**, `--conc 6`, `--wait-arena`. Arm file:
|
||
`tools/ab/arms_gun_b2.txt`. Reference: `pattern`. Movement pinned
|
||
`TR_MOVEMENT=strafe` in every arm (same rationale as Batch 1).
|
||
|
||
| # | arm | env | role |
|
||
|---|---|---|---|
|
||
| 1 | `pattern` | `TR_MOVEMENT=strafe` | reference |
|
||
| 2 | `bitbrain` | `+ TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay` | largest Batch-1 damage edge (+10.3) |
|
||
| 3 | `tmhorizon` | `+ TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both` | second Batch-1 damage edge (+8.7) |
|
||
|
||
**Pre-registered prediction (written BEFORE the battles finished):**
|
||
1. The `+10` dmg/run hint **does not survive** the wider panel: `bitbrain` and
|
||
`tmhorizon` damage deltas fall toward zero and stay inside the new (smaller)
|
||
MDE; the per-opponent sign test remains non-significant. Rationale: both arms
|
||
come from the same Pattern-lead-correction base, and the prior 32-opponent
|
||
gauntlet measured `bitbrain` at **−3.6 dmg/run** on a largely overlapping
|
||
roster; +10 on 15 opponents is plausibly a small-panel fluctuation.
|
||
2. Round wins remain **flat** for both arms — if anything they regress toward
|
||
zero.
|
||
3. If instead the damage hint **survives** at a significant cross-opponent sign
|
||
test with wins not down, it is recorded as the first gun to **beat** `pattern`
|
||
by rule 2 — and the campaign immediately looks for what the corrector is
|
||
exploiting. That would be a genuine up-set, and it would **not** be spun as a
|
||
null.
|
||
|
||
### Outcome — Batch 2
|
||
|
||
**MEASURED.** Session `/tmp/ab/j121_g2`, commit
|
||
`c343c00aaf800bc05ce9bf1b277054c11f5f577c`, frozen binary sha256 `d267ab78f6ce…`,
|
||
**33 opponents × 3 arms × 3 runs × 3 rounds = 297 battles, 0 failed, 0 never
|
||
started, 0 liveness exclusions**.
|
||
|
||
**A movement-default flip landed between the two batches** (commit `3fd6db9`,
|
||
`TR_MOVEMENT` default `tfil` → `strafe`; the only source change). Because
|
||
**every** arm in **both** batches pins `TR_MOVEMENT=strafe`, the effective
|
||
movement is identical in Batch 1 and Batch 2 — the batches are directly
|
||
comparable and the gun comparison is unconfounded. This is exactly what the pin
|
||
was for.
|
||
|
||
> **DIRECT ANSWER (MEASURED, n=33 opponents): the Batch-1 damage hint did NOT
|
||
> survive.** `bitbrain` collapses to **+1.8 dmg/run** (19/33 opponents, sign test
|
||
> p=0.49) with **+0.11 wins/run** — a wash. `tmhorizon` **reverses sign and is
|
||
> detectably WORSE on damage**: **−10.4 dmg/run** (10/33 opponents, sign test
|
||
> p=0.035, sign-flip p=0.0047, 95% CI [−17.2, −3.6]) at unchanged wins (−0.03).
|
||
> The only open signal in Batch 1 was a **small-panel fluctuation**, and the
|
||
> wider panel resolves it **against** the correctors. Combined with Batch 1:
|
||
> **nothing beats the shipped `Pattern` on damage/run AND round wins**;
|
||
> `onlyPattern` is CONFIRMED on 15 opponents and again on 33.
|
||
|
||
#### Pooled dashboard (all valid runs — explanation only, NOT the verdict)
|
||
|
||
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|
||
|---|---:|---:|---:|---:|---:|---:|---:|---:|
|
||
| `pattern` (REF) | 99 | 131.2 | 126.4 | 1.90 | 188/297 | 63.3% | 12.93% | 417 |
|
||
| `bitbrain` | 99 | 133.0 | 126.7 | **2.01** | **199/297** | **67.0%** | 12.91% | 417 |
|
||
| `tmhorizon` | 99 | 120.8 | 133.2 | 1.87 | 185/297 | 62.3% | 13.04% | 422 |
|
||
|
||
#### Cross-opponent aggregation (the verdict layer, damage + wins)
|
||
|
||
| arm | metric | mean Δ | spread (SD) | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|
||
|---|---|---:|---:|---|---:|---:|---:|---:|---:|
|
||
| `bitbrain` | damage | +1.80 | 18.95 | [−4.67, +8.26] | 19/33 | 0.4869 | 0.5914 | 0.5675 | 9.24 |
|
||
| `bitbrain` | wins | +0.11 | 0.48 | [−0.05, +0.28] | 12/19 | 0.3593 | 0.2436 | 0.5716 | 0.24 |
|
||
| `tmhorizon` | damage | **−10.40** | 20.01 | **[−17.22, −3.57]** | 10/33 | **0.03508** | **0.00472** | **0.006975** | 9.76 |
|
||
| `tmhorizon` | wins | −0.03 | 0.59 | [−0.23, +0.17] | 9/19 | 1.0 | 0.8459 | 0.4804 | 0.29 |
|
||
|
||
#### The pre-registered verdict (verbatim)
|
||
|
||
| rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) |
|
||
|---:|---|---:|---:|---|---|---|---|
|
||
| 1 | `bitbrain` | +0.11 | +1.8 | 12/19 p=0.3593 | 19/33 p=0.4869 | **not distinguishable** | **not distinguishable** |
|
||
| 2 | `tmhorizon` | −0.03 | −10.4 | 9/19 p=1 | 10/33 p=0.03508 | **WORSE** | **WORSE** |
|
||
|
||
Reference `pattern`: 131.2 dmg/run, 1.90 wins/run, 12.93% incoming, 417 px.
|
||
|
||
#### Reading
|
||
|
||
* **BitBrain is a wash — now measured three times.** vs the shipped Pattern:
|
||
−3.6 dmg/run on the 32-opponent legacy gauntlet (with `tfil` movement,
|
||
`docs/gauntlet_bitbrain_vs_pattern.md`), **+10.3** on the 15-opponent Batch-1
|
||
panel, **+1.8** here on 33 opponents. All three sit inside their own MDEs; the
|
||
pooled evidence is a zero. The Batch-1 `+10` was a 15-opponent fluctuation.
|
||
* **TMHorizon is the one arm the wider panel separates — and it loses.** Its
|
||
Batch-1 sign (+8.7) does not merely fail to replicate, it **inverts**
|
||
(−10.4, 10/33, p=0.035). The damage loss is broad (24/33 opponents negative,
|
||
worst `Aristocles` −61, `Jen` −54, `WallAvoider` −41), so it is not one
|
||
outlier. Since TMHorizon was already predicted to lose by its own design docs,
|
||
this is the first *confirming* live panel measurement of that prediction.
|
||
* **Wins are flat for both arms** (+0.11 / −0.03, MDEs 0.24 / 0.29): the
|
||
correctors change damage at most, never survival. In this harness round wins
|
||
are survival wins, so neither is a win.
|
||
* **The corrector family is closed** as a source of a `pattern`-beating gun at
|
||
the resolutions tested.
|
||
|
||
#### Pre-registered predictions — scorecard
|
||
|
||
| # | prediction | outcome |
|
||
|---|---|---|
|
||
| 1 | the +10 dmg hint does not survive the wider panel; deltas fall inside the new MDE; sign test non-significant | **CORRECT** (`bitbrain` +1.8, p=0.49; `tmhorizon` even reversed to −10.4, p=0.035) |
|
||
| 2 | wins remain flat | **CORRECT** (+0.11 / −0.03) |
|
||
| 3 | if the hint survived at p<0.05 with wins not down, it is the first gun to beat `pattern` | **not invoked** |
|
||
|
||
**Final direct answer for the phase (MEASURED).** Across two frozen panels
|
||
(15 and 33 opponents), a single frozen-binary-per-session design, 567 battles,
|
||
six distinct gun configurations (five challengers + the reference), and movement
|
||
pinned identically in every arm:
|
||
**no rack gun and no small rack beats the shipped `Pattern` on damage/run AND
|
||
round wins.** Every candidate is either not distinguishable (`bitbrain`,
|
||
`tmhorizon`, `rack_pk`, `rack_pt` at n=15), a detectable loss
|
||
(`knn`; `tmhorizon` at n=33), or a sub-MDE damage-only wobble
|
||
(`bitbrain`). The shipped **`onlyPattern` rack is CONFIRMED**, not merely
|
||
assumed — the gun axis's first phase ends in a measured optimum of the tested
|
||
design space.
|
||
|
||
|
||
---
|
||
|
||
## What to try next (rewritten AFTER Batches 1–2 — recommendations, not results)
|
||
|
||
**The primary question is answered:** nothing beats the shipped `pattern` across
|
||
the frozen panel, and `onlyPattern` is CONFIRMED. Ranked by value per battle:
|
||
|
||
1. **The damage-side hint is now CLOSED.** Batch 2 (33 opponents) killed it:
|
||
`bitbrain` is a wash (+1.8 dmg/run, p=0.49) and `tmhorizon` is detectably
|
||
**worse** (−10.4 dmg/run, p=0.035). Do **not** spend another batch on the
|
||
Pattern-lead correctors; three independent measurements of `bitbrain` vs
|
||
Pattern now average ~0 (−3.6 / +10.3 / +1.8 dmg/run, all inside their MDEs).
|
||
2. **Lead INFORMATION remains the only named gun axis** (given: amplitude is
|
||
dead). But the two live implementations of it (BitBrain's gain, TMHorizon's
|
||
shift) are both measured negative/neutral across panels. A future axis would
|
||
need a *different* information source, not another knob on these two — e.g. a
|
||
new base prediction, or a gun-agnostic ensemble. Do not re-open the correctors.
|
||
3. **The selector is confirmed negative at a 2-gun rack** (`rack_pt` ==
|
||
`tmhorizon`; `rack_pk` below `pattern` on wins). Do not re-open it with a
|
||
larger rack: the 16-gun negative already stands and Batch 1 shows the
|
||
mechanism does not manufacture a win.
|
||
4. **Drop `knn`** (detectably worse on both primaries, 0/15 on damage).
|
||
5. **Never chase damage without wins.** This harness's round wins are survival
|
||
wins, so the gun's job is to *kill*; a damage gain that does not raise the win
|
||
rate is the `ring`-mover trap in a new costume (and `tmhorizon`'s n=33 result
|
||
is a *negative* damage signal with flat wins — also not a win).
|
||
6. **Spinner-specific claim (owner's) remains UNTESTED.** The legacy roster has
|
||
**no constant-turn spinner**; SpinBot is a periodic circle-mover and is in
|
||
both panels (Batch 1: `bitbrain` +25.3 dmg / +0 wins, `tmhorizon` +15.2 / +0
|
||
wins; on the 33-panel SpinBot's deltas are in the per-opponent tables).
|
||
Building a true constant-turn spinner opponent is a nearly-free follow-up and
|
||
is the one part of the owner's claim this campaign has not touched.
|
||
|
||
## What would make us stop
|
||
|
||
**PHASE STATUS (updated after Batch 2): STOPPED — the condition fired.** Across
|
||
Batches 1–2 no arm beats `pattern` beyond the MDE, and no arm shows a
|
||
`>= +MDE` damage gain with p<0.10 (`bitbrain` +1.8 vs MDE 9.24; `tmhorizon`
|
||
actually negative). The gun stage's first phase is therefore **closed** with
|
||
*"the shipped `onlyPattern` rack is the best gun configuration we have measured
|
||
across the panels (15 and 33 opponents)"* — a successful outcome, not a failure.
|
||
Further gun work must open a **named new axis** (see *What to try next*), not
|
||
another arm on the same axis.
|
||
|
||
* **Stop the first phase** once a batch's best arm cannot beat `pattern` beyond
|
||
the MDE, or when a gun arm's damage gain is bought with a detectable win loss.
|
||
At that point *"`onlyPattern` is the measured optimum of this rack"* is the
|
||
conclusion, not a failure (§2 rule 4). **This fired after Batch 2.**
|
||
* **Stop a single batch early** only for a contract violation (arena not free,
|
||
liveness FAIL, non-zero exit rate) — never because the numbers look boring.
|
||
* **Do not invent more arms on the same axis** once two consecutive batches fail
|
||
to improve on `pattern` beyond the MDE; move to a *named* new axis instead
|
||
(candidate list in "What to try next").
|
||
|
||
---
|
||
|
||
## How to run a batch (exact commands)
|
||
|
||
```sh
|
||
# 0. wait for the arena (other campaign jobs may be fighting)
|
||
TOURNAMENT_NIMCACHE=/tmp/nc_j121 \
|
||
tools/ab/tournament_run.sh \
|
||
--arms tools/ab/arms_gun_b1.txt \
|
||
--panel tools/ab/panel_movement.txt \
|
||
--runs 3 --rounds 3 --conc 6 --wait-arena 45 \
|
||
--reference pattern \
|
||
--outdir /tmp/ab/j121_g1
|
||
|
||
# Batch 2 (the wider-panel calibration) was the same command with
|
||
# --arms tools/ab/arms_gun_b2.txt --panel tools/ab/panel_gun_b2.txt \
|
||
# --outdir /tmp/ab/j121_g2 --wait-arena 50
|
||
|
||
# 1. the paired per-opponent table, sign tests, MDE and the pre-registered verdict
|
||
python3 tools/ab/tournament_analyze.py /tmp/ab/j121_g1 --reference pattern
|
||
```
|
||
|
||
`--reference` may be ANY arm of the session: re-analyzing an old session with a
|
||
different reference is a free pairwise comparison with no battles.
|
||
|
||
## Session log (outdirs are in `/tmp` and are NOT committed)
|
||
|
||
| session | commit | battles | arms | verdict |
|
||
|---|---|---:|---|---|
|
||
| `/tmp/ab/j121_g1` | `1d8143a` | 270 (0 failed, 0 never started, 0 excluded) | pattern, bitbrain, tmhorizon, knn, rack_pk, rack_pt | **nothing beats the shipped `pattern`**; all arms not distinguishable except `knn` = WORSE; `onlyPattern` CONFIRMED |
|
||
| `/tmp/ab/j121_g2` | `c343c00` | 297 (0 failed, 0 never started, 0 excluded) | pattern, bitbrain, tmhorizon | **the Batch-1 damage hint does not survive**: `bitbrain` +1.8 dmg/run p=0.49 (wash); `tmhorizon` −10.4 dmg/run p=0.035 (WORSE); wins flat |
|
||
|
||
---
|
||
|
||
# Phase 2: tuning the incumbent
|
||
|
||
**Owner's mandate (phase 2):** the phase-1 campaign closed the "other gun / other
|
||
rack" design space (nothing beats the shipped `Pattern`; the correctors are a
|
||
wash or a loss). What remains OPEN is the **incumbent's own tuning**: Pattern's
|
||
match-length / history parameters had NEVER been swept. Phase 2 asks the direct
|
||
question:
|
||
|
||
> **Can the shipped `Pattern` be improved by tuning its own match parameters —
|
||
> and if so, by how much on damage/run and round wins?**
|
||
|
||
The lead-amplitude axis is already known dead (`docs/bitbrain_campaign.md`), and
|
||
the radial knobs are known non-winners (job j99: `bmPath` structural no-op; live
|
||
+0.28 pp p=0.62 for scale 0.98, −0.42 pp p=0.46 for offset −20). Those are
|
||
therefore **controls** here, not candidates.
|
||
|
||
## Task A — exposing Pattern's match-shape parameters (what and why)
|
||
|
||
**MEASURED (code read).** `common_libs/guns/pattern_matcher.nim` had exactly two
|
||
tunable knobs, both RADIAL (`TR_PATTERN_RAD_SCALE`, `TR_PATTERN_RAD_OFFSET`), and
|
||
neither can change the lead bearing (job j99 proved the bearing is untouched), so
|
||
neither was ever the "lead information" axis. The parameters that actually
|
||
control **how the pattern is matched** were compile-time constants:
|
||
|
||
| parameter | code | what it controls | exposed as |
|
||
|---|---|---|---|
|
||
| match-key length | `PatternLen = 10` | the length of the movement segment compared (the search key); also how far after the match the replay starts | `TR_PATTERN_LEN` (int, default 10) |
|
||
| search depth | implicit `HistorySize = 500` | how far back the best-match scan may reach (`scanEnd` was always `count−PatternLen−1`) | `TR_PATTERN_DEPTH` (int, default 500 = full buffer) |
|
||
| similarity/search radius | **does not exist** | the search always takes the single lowest-cost match; there is no acceptance threshold or radius to expose | **not exposed — nothing to expose** |
|
||
| history buffer capacity | `HistorySize = 500` | the fixed `array[HistorySize]` backing store | runtime depth limit only; the buffer ceiling cannot be raised at runtime (see below) |
|
||
|
||
**Why env and not `-d:`.** The phase-1 instrument (`tools/ab/tournament_run.sh`)
|
||
builds ONE frozen binary from `git archive HEAD` and every arm differs only by its
|
||
env dict. A `{.intdefine.}` knob would need one binary per arm, which the
|
||
instrument forbids. Both new knobs are therefore resolved lazily from the
|
||
environment on the first `predict`, exactly like the radial knobs, and default to
|
||
the pre-knob constants. `setMatchParams(patternLen, histDepth)` is the explicit
|
||
offline/unit-test twin that writes the same fields.
|
||
|
||
**What could NOT be exposed cheaply (MEASURED).** `HistorySize` sizes four fixed
|
||
`array[HistorySize(+1)]` fields in the gun object. Raising it above 500 at
|
||
runtime is impossible without a heap buffer; the runtime `TR_PATTERN_DEPTH` knob
|
||
therefore *lowers* the effective search depth within the existing 500-entry
|
||
buffer. A depth above 500 is clamped to 500, and a non-positive or unparsable
|
||
value falls back to the shipped default. `PatternLen` is clamped to
|
||
`1..HistorySize`.
|
||
|
||
## Task A — default parity (MEASURED, byte-for-byte)
|
||
|
||
* **Baseline-vs-new parity dump.** A throwaway harness replayed 800 ticks of
|
||
`tools/fixtures/drussgt_vs_crazy.jsonl` through one `PatternMatcherGun` at all
|
||
four power bins (3200 predictions) and printed every point at full precision.
|
||
Compiled once against a `HEAD` worktree (`/tmp/j123_base`, before the change)
|
||
and once against the modified tree: **the two dumps are byte-identical**
|
||
(`diff -q` clean). The shipped default path is unchanged.
|
||
* **The knobs are live and self-falling-back.** `TR_PATTERN_LEN=6/16` and
|
||
`TR_PATTERN_DEPTH=100/20` each change the dump; `TR_PATTERN_LEN=banana`
|
||
reproduces the default dump exactly.
|
||
* **Existing guards pass** (no new guard tests, per the "cut ceremony" rule):
|
||
`common_libs/tests/test_pattern_radial_offset.nim` (6 checks) and
|
||
`common_libs/tests/test_gun_harness.nim` (all checks) pass; the env-report
|
||
guard `test_env_report.nim` passes with `TR_PATTERN_LEN` / `TR_PATTERN_DEPTH`
|
||
registered in `knownEnvNames()` and the effective-values report.
|
||
* **Clean-archive compile** is exercised by `tournament_run.sh` itself, which
|
||
builds the frozen binary from `git archive HEAD`.
|
||
|
||
## Task B — Batch 1 (pre-registered BEFORE any battle)
|
||
|
||
**Design.** One frozen binary, six env-only arms, the FROZEN 15-opponent panel
|
||
`tools/ab/panel_movement.txt`, **3 runs × 3 rounds per (opponent, arm) = 270
|
||
battles**, `--conc 6`, `--wait-arena`. Arm file: `tools/ab/arms_gun_b3.txt`.
|
||
Reference: `pattern`. Movement pinned `TR_MOVEMENT=strafe` in every arm.
|
||
|
||
| # | arm | env over the pin | what it isolates |
|
||
|---|---|---|---|
|
||
| 1 | `pattern` | (none) | the arm to beat (shipped `onlyPattern` rack) |
|
||
| 2 | `len6` | `TR_PATTERN_LEN=6` | shorter movement segment compared |
|
||
| 3 | `len16` | `TR_PATTERN_LEN=16` | longer movement segment compared |
|
||
| 4 | `depth100` | `TR_PATTERN_DEPTH=100` | shallower history / match search |
|
||
| 5 | `rad_offset` | `TR_PATTERN_RAD_OFFSET=-20` | control: aim 20 px short |
|
||
| 6 | `rad_scale` | `TR_PATTERN_RAD_SCALE=0.95` | control: scale the aim distance |
|
||
|
||
**Pre-registered prediction (written BEFORE the battles finished):**
|
||
1. **No arm beats `pattern` on both primaries.** The incumbent is already tuned —
|
||
the phase-2 null. The point estimates sit inside the MDE.
|
||
2. `len16` is **WORSE or flat** on damage: a 16-tick key matches rarely in a
|
||
~500-tick buffer, so the gun falls back to the linear forecast (the weaker
|
||
base) more often; wins flat.
|
||
3. `len6` is **flat** on damage (more matches but a noisier replay) and flat on
|
||
wins; possibly a sub-MDE wobble in either direction.
|
||
4. `depth100` is **not distinguishable** from `pattern` — the full buffer already
|
||
contains the useful candidates; a sub-MDE damage wobble is possible.
|
||
5. The two radial controls are **not distinguishable** and, per job j99, cannot
|
||
change the real aim bearing; any damage delta is the fire-gate channel only.
|
||
6. If any surprise exists, it is a **damage-only wobble without wins** — the
|
||
`ring`-mover trap — and will not be read as a win.
|
||
|
||
**Pre-registered Batch-2 trigger (task rule).** A Batch-1 arm is promoted to a
|
||
higher-power Batch 2 **only if** it beats the reference `pattern` by the
|
||
campaign rule 2 (one primary up at cross-opponent sign-test p<0.05 while the
|
||
other does not go down) **AND** its damage CI excludes 0 **AND** the sign-flip
|
||
permutation test gives p<0.05. Otherwise Batch 1 is the answer.
|
||
|
||
|
||
### Outcome — Batch 1
|
||
|
||
**MEASURED.** Session `/tmp/ab/j123_b1`, commit
|
||
`2a98aba91b0ec023f580cd2485cafba2404d275d`, frozen binary sha256 `287d292fb7d1…`,
|
||
**15 opponents × 6 arms × 3 runs × 3 rounds = 270 battles, 0 failed, 0 never
|
||
started, 0 liveness exclusions**. Arena serialization held: the parallel job
|
||
j124 (`/tmp/ab/j124_spinner.run.log`) detected this session's processes at
|
||
04:08–04:18 and **waited**; no foreign battle ran concurrently. Every arm
|
||
declared `TR_MOVEMENT=strafe`; the new `TR_PATTERN_LEN` / `TR_PATTERN_DEPTH`
|
||
tokens appear verbatim in each overriding arm's own raw-env report and the
|
||
reference arm shows the shipped defaults 10 / 500.
|
||
|
||
#### Pooled dashboard (all valid runs — explanation only, NOT the verdict)
|
||
|
||
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|
||
|---|---:|---:|---:|---:|---:|---:|---:|---:|
|
||
| `pattern` (REF) | 45 | 101.1 | 161.0 | 1.27 | 57/135 | 42.2% | 13.27% | 428 |
|
||
| `len6` | 45 | 110.8 | 151.2 | **1.76** | **79/135** | **58.5%** | **12.28%** | 424 |
|
||
| `len16` | 45 | 103.3 | 158.8 | 1.53 | 69/135 | 51.1% | 13.07% | 427 |
|
||
| `depth100` | 45 | 99.9 | 154.5 | 1.53 | 69/135 | 51.1% | 12.77% | 429 |
|
||
| `rad_offset` | 45 | 107.9 | 157.2 | 1.69 | 76/135 | 56.3% | 13.21% | 431 |
|
||
| `rad_scale` | 45 | 105.9 | 157.1 | 1.38 | 62/135 | 45.9% | 12.62% | 429 |
|
||
|
||
#### Per-opponent paired deltas (arm − `pattern`), damage and wins
|
||
|
||
| opponent | style | Δdmg_len6 | Δwins_len6 | Δdmg_len16 | Δwins_len16 | Δdmg_depth100 | Δwins_depth100 | Δdmg_rad_offset | Δwins_rad_offset | Δdmg_rad_scale | Δwins_rad_scale |
|
||
|---|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
||
| DrussGT | dodger | +8.4 | +0.67 | -21.7 | -0.33 | +3.6 | +0.67 | +10.5 | +1.33 | +18.7 | +0.67 |
|
||
| Diamond | dodger | -22.3 | +0.00 | -18.1 | -0.33 | -54.7 | -0.33 | -30.9 | +0.00 | -17.0 | -0.33 |
|
||
| Dookious | dodger | -6.6 | +1.33 | +7.6 | +1.00 | -0.8 | +0.67 | +2.6 | +0.67 | -15.6 | -0.33 |
|
||
| GresSuffurd | dodger | +32.3 | +0.00 | +31.9 | +0.00 | -6.6 | +0.00 | +28.4 | +0.67 | -0.9 | +0.00 |
|
||
| CassiusClay | dodger | +15.0 | +0.67 | +6.8 | +1.00 | +10.0 | +0.33 | +12.5 | +0.67 | +40.4 | +1.00 |
|
||
| RetroGirl | pattern | +48.0 | +1.33 | +20.3 | +0.67 | +59.9 | +1.00 | +4.1 | +0.00 | +55.5 | +0.33 |
|
||
| TripHammer | pattern | +6.1 | +0.00 | +0.1 | +0.33 | -21.6 | +0.00 | -9.3 | +0.00 | +4.7 | +0.33 |
|
||
| Coriantumr | pattern | +1.9 | -0.33 | -3.9 | +0.33 | +2.0 | +1.00 | -5.3 | -0.33 | -10.3 | -0.33 |
|
||
| WallAvoider | wallfollower | +28.5 | +1.00 | -23.5 | +0.00 | +41.0 | +0.00 | +23.6 | +1.00 | -24.0 | +0.00 |
|
||
| HawkOnFire | cornercamper | -8.9 | +0.67 | -9.5 | +0.00 | -9.3 | -0.33 | -12.0 | +0.33 | -5.0 | +0.33 |
|
||
| SpinBot | spinner | +5.7 | +0.00 | +13.3 | +0.00 | -13.5 | +0.00 | -1.7 | +0.00 | -0.3 | +0.00 |
|
||
| DiamondStealer | rammer | -16.1 | +0.00 | -9.8 | +0.00 | -14.6 | +0.00 | +31.8 | +1.33 | +1.1 | -0.67 |
|
||
| BlitzBat | brawler | +9.1 | +0.00 | +17.8 | +0.33 | -6.3 | +0.33 | +9.6 | +0.33 | +9.2 | -0.67 |
|
||
| YersiniaPestis | aggressive | +21.0 | +1.33 | +2.6 | +0.67 | -18.6 | +0.33 | +27.4 | +0.67 | +3.0 | +0.67 |
|
||
| Ascendant | aggressive | +22.4 | +0.67 | +18.7 | +0.33 | +10.9 | +0.33 | +9.3 | -0.33 | +12.6 | +0.67 |
|
||
|
||
#### Cross-opponent aggregation (the verdict layer, damage + wins)
|
||
|
||
| arm | metric | mean Δ | spread (SD) | SE | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|
||
|---|---|---:|---:|---:|---|---:|---:|---:|---:|---:|
|
||
| `len6` | damage | +9.63 | 18.97 | 4.90 | [−0.88, +20.13] | 11/15 | 0.1185 | 0.06976 | 0.09384 | 13.72 |
|
||
| `len6` | wins | **+0.49** | 0.58 | 0.15 | **[+0.17, +0.81]** | 8/9 | **0.03906** | **0.007812** | **0.01269** | 0.42 |
|
||
| `len16` | damage | +2.18 | 16.62 | 4.29 | [−7.02, +11.38] | 9/15 | 0.6072 | 0.6142 | 0.712 | 12.02 |
|
||
| `len16` | wins | +0.27 | 0.42 | 0.11 | [+0.03, +0.50] | 8/10 | 0.1094 | 0.04688 | 0.01852 | 0.30 |
|
||
| `depth100` | damage | −1.24 | 26.52 | 6.85 | [−15.93, +13.45] | 6/15 | 0.6072 | 0.8677 | 0.5137 | 19.19 |
|
||
| `depth100` | wins | +0.27 | 0.42 | 0.11 | [+0.03, +0.50] | 8/10 | 0.1094 | 0.04688 | 0.046 | 0.30 |
|
||
| `rad_offset` | damage | +6.71 | 17.18 | 4.44 | [−2.80, +16.23] | 10/15 | 0.3018 | 0.1522 | 0.1323 | 12.43 |
|
||
| `rad_offset` | wins | +0.42 | 0.54 | 0.14 | [+0.12, +0.72] | 9/11 | 0.06543 | 0.01465 | 0.01424 | 0.39 |
|
||
| `rad_scale` | damage | +4.80 | 21.10 | 5.45 | [−6.89, +16.49] | 8/15 | 1 | 0.4095 | 0.6293 | 15.27 |
|
||
| `rad_scale` | wins | +0.11 | 0.51 | 0.13 | [−0.17, +0.40] | 7/12 | 0.7744 | 0.5107 | 0.289 | 0.37 |
|
||
|
||
#### The pre-registered verdict (verbatim)
|
||
|
||
| rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) |
|
||
|---:|---|---:|---:|---|---|---|---|
|
||
| 1 | `len6` | +0.49 | +9.6 | 8/9 p=0.03906 | 11/15 p=0.1185 | **BETTER** | **BETTER** |
|
||
| 2 | `rad_offset` | +0.42 | +6.7 | 9/11 p=0.06543 | 10/15 p=0.3018 | **not distinguishable** | **not distinguishable** |
|
||
| 3 | `len16` | +0.27 | +2.2 | 8/10 p=0.1094 | 9/15 p=0.6072 | **not distinguishable** | **not distinguishable** |
|
||
| 4 | `depth100` | +0.27 | −1.2 | 8/10 p=0.1094 | 6/15 p=0.6072 | **not distinguishable** | **not distinguishable** |
|
||
| 5 | `rad_scale` | +0.11 | +4.8 | 7/12 p=0.7744 | 8/15 p=1 | **not distinguishable** | **not distinguishable** |
|
||
|
||
Reference `pattern`: 101.1 dmg/run, 1.27 wins/run, 13.27% incoming, 428 px.
|
||
|
||
#### Reading and the single red flag
|
||
|
||
* **`len6` is the only arm that beats `pattern` by the campaign rule** — on the
|
||
WINS primary: +0.49 wins/run, sign test 8/9 p=0.0391, sign-flip p=0.0078,
|
||
wins CI [+0.17, +0.81] excluding 0 — while damage is nominally up (+9.6, but
|
||
inside the 13.7 MDE and its CI includes 0). A 6-tick match key is broadly
|
||
positive (10/15 opponents non-negative, only `Coriantumr` −0.33 negative).
|
||
* **RED FLAG (MEASURED): all five arms are nominally positive on wins**
|
||
(+0.49 / +0.27 / +0.27 / +0.42 / +0.11), including `rad_offset`, which by job
|
||
j99's exact structural proof **cannot change the aim bearing** and whose live
|
||
effect was measured as a tie (+0.28 pp, p=0.62). The reference `pattern` also
|
||
measures **42.2% wins here vs 50.4% in j121 on the same 15-panel** (its damage
|
||
is unchanged at 101–99.6), so the leading explanation for the all-arms-positive
|
||
pattern is a **session-level low reference**, not five independent gun wins.
|
||
The within-session paired deltas remain valid, but `len6`'s win is exactly the
|
||
kind of single-session effect that must replicate on a second, wider panel.
|
||
* **The two shape knobs behave plausibly.** `len16` (fewer, more specific
|
||
matches → more linear fallback) is damage-flat; `depth100` (shallower search)
|
||
is damage-flat at a larger MDE. Neither is a detectable win.
|
||
|
||
#### Pre-registered predictions — scorecard
|
||
|
||
| # | prediction | outcome |
|
||
|---|---|---|
|
||
| 1 | no arm beats `pattern` on BOTH primaries | **CORRECT** (`len6` beats on wins only; damage CI includes 0) |
|
||
| 2 | `len16` flat/worse on damage, wins flat | **PARTLY WRONG**: damage flat (+2.2), wins nominally +0.27 |
|
||
| 3 | `len6` flat on damage and wins | **WRONG**: wins +0.49 significant; damage +9.6 sub-MDE |
|
||
| 4 | `depth100` not distinguishable | **CORRECT** (damage −1.2), though wins nominally +0.27 |
|
||
| 5 | radial controls not distinguishable / bearing-invariant | **CORRECT**: `rad_scale` flat; `rad_offset` nominal +0.42 wins but sign p=0.065 (not significant) |
|
||
| 6 | any surprise is damage-only without wins | **WRONG in direction**: the surprise is wins-with-nominal-damage |
|
||
|
||
## Batch 2 — higher-power confirmation of `len6` (+ the structural control)
|
||
|
||
**Why we run it (task trigger, with one recorded deviation).** The task's Batch-2
|
||
trigger is: an arm beats the reference with **the winning primary's CI excluding
|
||
0 AND sign-flip p<0.05**. `len6` satisfies it on wins (CI [+0.17, +0.81];
|
||
p(sign-flip)=0.0078). *Deviation, recorded:* the Batch-1 pre-registration added a
|
||
stricter clause — "its **damage** CI excludes 0" — which `len6`'s damage
|
||
([−0.88, +20.13]) does not meet. We run Batch 2 anyway because (a) the task rule
|
||
keys on the **winning** primary and (b) the all-arms-positive red flag means a
|
||
replication is the only way to tell a real `len6` effect from a session-level
|
||
reference anomaly. This is a deviation from the letter of my own pre-registration,
|
||
not from the task's criterion.
|
||
|
||
**Design.** One frozen binary, four env-only arms, the wider panel
|
||
`tools/ab/panel_gun_b2.txt` (**33 opponents**, superset of Batch 1's 15),
|
||
**3 runs × 3 rounds = 396 battles**, `--conc 6`, `--wait-arena`. Arm file:
|
||
`tools/ab/arms_gun_b4.txt`. Reference: `pattern`. Movement pinned
|
||
`TR_MOVEMENT=strafe` in every arm.
|
||
|
||
| # | arm | env over the pin | role |
|
||
|---|---|---|---|
|
||
| 1 | `pattern` | (none) | reference |
|
||
| 2 | `len6` | `TR_PATTERN_LEN=6` | the candidate |
|
||
| 3 | `len16` | `TR_PATTERN_LEN=16` | dose-response on key length |
|
||
| 4 | `rad_offset` | `TR_PATTERN_RAD_OFFSET=-20` | **structural positive control** (bearing-invariant) |
|
||
|
||
**Pre-registered prediction (written BEFORE the battles finished):**
|
||
1. **`len6`'s +0.49 wins does NOT replicate.** On 33 opponents its wins delta
|
||
collapses toward the MDE and the sign test (or the sign-flip test) fails.
|
||
Rationale: all five Batch-1 arms were positive on wins including the
|
||
bearing-invariant `rad_offset`, and the reference was 8 points low on win rate
|
||
versus j121's same-panel measurement — a single-session low reference, not a
|
||
`len6` effect.
|
||
2. **`rad_offset` is not a detectable win**; if it, too, replicates a wins gain,
|
||
the wins axis is being moved by something other than the aim bearing and the
|
||
whole Batch-1 wins signal is an artifact.
|
||
3. `len16` stays within its MDE on both primaries.
|
||
4. Damage stays within the MDE for every arm.
|
||
5. **If `len6` DOES replicate** (winning-metric CI excludes 0, sign-flip
|
||
p<0.05, damage not detectably down), it is recorded as the campaign's
|
||
**champion candidate** — and the shipped-default flip remains a separate,
|
||
explicit decision that is **NOT** taken in this job.
|
||
|
||
### Outcome — Batch 2
|
||
|
||
**MEASURED.** Session `/tmp/ab/j123_b2`, commit
|
||
`1b59b6581075cf4019773dd74e6db9eecf2fcb02`, frozen binary sha256 `8155c39fbc6a…`,
|
||
**33 opponents × 4 arms × 3 runs × 3 rounds = 396 battles, 0 failed, 0 never
|
||
started, 0 liveness exclusions**. Arena serialization held: the job waited at
|
||
the start of the run for j124's spinner fleet (`/tmp/ab/j124_spinner`) and only
|
||
started once it was gone; no foreign battle ran concurrently with the measurement.
|
||
Every arm declared `TR_MOVEMENT=strafe`; `TR_PATTERN_LEN` / `TR_PATTERN_RAD_OFFSET`
|
||
appear verbatim in each overriding arm's raw-env report.
|
||
|
||
#### Pooled dashboard (all valid runs — explanation only, NOT the verdict)
|
||
|
||
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|
||
|---|---:|---:|---:|---:|---:|---:|---:|---:|
|
||
| `pattern` (REF) | 99 | 129.5 | 124.4 | 2.11 | 209/297 | 70.4% | 12.54% | 428 |
|
||
| `len6` | 99 | 131.4 | 121.0 | 2.07 | 205/297 | 69.0% | 12.04% | 423 |
|
||
| `len16` | 99 | 125.8 | 121.4 | 2.04 | 202/297 | 68.0% | 12.36% | 426 |
|
||
| `rad_offset` | 99 | 126.9 | 132.4 | 1.91 | 189/297 | 63.6% | 13.03% | 425 |
|
||
|
||
#### Per-opponent paired deltas (arm − `pattern`), damage and wins
|
||
|
||
| opponent | style | Δdmg_len6 | Δwins_len6 | Δdmg_len16 | Δwins_len16 | Δdmg_rad_offset | Δwins_rad_offset |
|
||
|---|---|---:|---:|---:|---:|---:|---:|
|
||
| Aristocles | dodger | +0.7 | +0.00 | -12.7 | +0.00 | -5.3 | +0.00 |
|
||
| Ascendant | other | +13.8 | -0.67 | -23.1 | -1.33 | +0.7 | -1.00 |
|
||
| BlitzBat | regular | -2.6 | +0.00 | +4.4 | +0.33 | -12.0 | +0.33 |
|
||
| BrokenSword | other | -17.7 | +0.00 | +6.0 | +0.33 | -15.9 | +0.33 |
|
||
| CassiusClay | dodger | +3.7 | +0.67 | -7.5 | +0.00 | +32.4 | -0.33 |
|
||
| Cigaret | dodger | -26.3 | -2.00 | -5.1 | -0.33 | -29.1 | -1.67 |
|
||
| CigaretBH | dodger | +17.6 | +0.33 | +1.1 | +0.33 | +18.2 | +0.00 |
|
||
| Coriantumr | other | -3.9 | +0.00 | -14.1 | -0.33 | +3.9 | +0.00 |
|
||
| Diamond | dodger | -5.2 | -0.33 | -8.8 | -0.67 | -13.7 | -0.67 |
|
||
| DiamondHawk | other | +3.2 | -0.33 | -3.9 | +0.00 | -6.6 | -0.33 |
|
||
| DiamondStealer | regular | +29.5 | +1.00 | +27.4 | +0.33 | +16.5 | +0.33 |
|
||
| Dookious | dodger | +16.6 | +0.33 | +14.5 | +0.33 | -2.1 | +0.00 |
|
||
| DrussGT | other | -13.2 | +0.00 | +20.7 | +0.33 | -15.9 | -0.33 |
|
||
| FloodMini | regular | +1.0 | +0.00 | +12.0 | +0.00 | +20.7 | +0.00 |
|
||
| GresSuffurd | dodger | +10.8 | +0.33 | -11.1 | +0.00 | +4.8 | +0.00 |
|
||
| HawkOnFire | regular | +24.9 | +0.67 | +26.3 | -0.33 | -2.6 | +0.67 |
|
||
| Jen | dodger | -36.3 | +0.00 | -27.0 | +0.00 | -16.1 | -0.33 |
|
||
| Komarious | dodger | +3.3 | +0.33 | +1.2 | +0.00 | -11.4 | +0.33 |
|
||
| KurtWaveSurfer | dodger | -18.9 | -0.67 | -35.9 | +0.00 | -7.8 | +0.00 |
|
||
| LightningBug | other | -15.3 | +0.00 | -10.7 | +0.00 | -20.7 | +0.00 |
|
||
| LionWWSVMvoid | dodger | +19.1 | +0.00 | -1.3 | +0.00 | -1.7 | +0.00 |
|
||
| Lukious | dodger | +17.6 | +0.00 | +1.8 | +0.33 | +19.6 | +0.33 |
|
||
| PatternRobot | regular | +21.3 | +0.00 | +6.7 | +0.00 | -1.3 | +0.00 |
|
||
| Phoenix | other | -8.2 | +0.33 | -29.5 | -0.33 | -4.4 | -0.67 |
|
||
| RetroGirl | other | -29.6 | -0.33 | -6.7 | +0.00 | -17.0 | -1.67 |
|
||
| RougeDC | dodger | +7.5 | +0.00 | -1.9 | -0.33 | +2.3 | +0.00 |
|
||
| Shadow | other | +3.1 | -0.67 | -10.9 | -0.67 | -10.4 | -0.67 |
|
||
| SpinBot | spinner | -14.7 | +0.00 | -9.5 | +0.00 | -3.5 | +0.00 |
|
||
| TripHammer | other | +0.9 | -0.33 | -7.6 | -0.33 | -16.8 | -1.00 |
|
||
| WallAvoider | regular | +35.7 | -0.33 | +3.3 | +0.33 | -5.6 | -1.00 |
|
||
| WaveSurferGF | dodger | +12.5 | +0.33 | -16.5 | -0.67 | +26.3 | +0.33 |
|
||
| WaveSurferPG | dodger | -1.2 | -0.67 | +0.3 | +0.33 | -22.7 | +0.00 |
|
||
| YersiniaPestis | other | +10.8 | +0.67 | -4.7 | +0.00 | +8.6 | +0.33 |
|
||
|
||
#### Cross-opponent aggregation (the verdict layer, damage + wins)
|
||
|
||
| arm | metric | mean Δ | spread (SD) | SE | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|
||
|---|---|---:|---:|---:|---|---:|---:|---:|---:|---:|
|
||
| `len6` | damage | +1.83 | 17.09 | 2.98 | [−4.00, +7.67] | 20/33 | 0.2962 | 0.5417 | 0.4859 | 8.34 |
|
||
| `len6` | wins | −0.04 | 0.54 | 0.09 | [−0.22, +0.14] | 10/20 | 1 | 0.7631 | 0.8806 | 0.26 |
|
||
| `len16` | damage | −3.72 | 14.45 | 2.52 | [−8.65, +1.22] | 13/33 | 0.2962 | 0.1494 | 0.09657 | 7.05 |
|
||
| `len16` | wins | −0.07 | 0.38 | 0.07 | [−0.20, +0.06] | 9/19 | 1 | 0.379 | 0.5303 | 0.19 |
|
||
| `rad_offset` | damage | −2.69 | 14.72 | 2.56 | [−7.72, +2.33] | 11/33 | 0.08014 | 0.2997 | 0.1952 | 7.18 |
|
||
| `rad_offset` | wins | −0.20 | 0.56 | 0.10 | [−0.39, −0.01] | 8/20 | 0.5034 | 0.05926 | 0.05875 | 0.28 |
|
||
|
||
#### The pre-registered verdict (verbatim)
|
||
|
||
| rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) |
|
||
|---:|---|---:|---:|---|---|---|---|
|
||
| 1 | `len6` | −0.04 | +1.8 | 10/20 p=1 | 20/33 p=0.2962 | **not distinguishable** | **not distinguishable** |
|
||
| 2 | `len16` | −0.07 | −3.7 | 9/19 p=1 | 13/33 p=0.2962 | **not distinguishable** | **not distinguishable** |
|
||
| 3 | `rad_offset` | −0.20 | −2.7 | 8/20 p=0.5034 | 11/33 p=0.08014 | **not distinguishable** | **not distinguishable** |
|
||
|
||
Reference `pattern`: 129.5 dmg/run, 2.11 wins/run, 12.54% incoming, 428 px.
|
||
|
||
#### Reading
|
||
|
||
* **`len6`'s Batch-1 wins win did not replicate.** On 33 opponents it collapses
|
||
from +0.49 wins/run to **−0.04 wins/run** (10/20, p=1) and from +9.6 to
|
||
**+1.8 dmg/run** (20/33, p=0.30, CI [−4.00, +7.67], MDE 8.34). Both primaries
|
||
are squarely inside the MDE. Verdict: **NOT DISTINGUISHABLE**.
|
||
* **The Batch-1 red flag is explained.** All five Batch-1 arms were wins-positive
|
||
including the bearing-invariant `rad_offset`; here `rad_offset` turns
|
||
**negative** on both primaries (−0.20 wins, −2.7 dmg) and the other two arms
|
||
are wins-flat. The Batch-1 effect was a **session-level low reference**:
|
||
`pattern` measured 42.2% win rate on 15 opponents in Batch 1, versus **70.4%**
|
||
here on its 33-opponent superset and 63.3% in j121 Batch 2 on the same 33-panel.
|
||
A single-session paired comparison is internally valid, but a wins-only signal
|
||
in one session is not evidence until it replicates — which is exactly what
|
||
Batch 2 tested.
|
||
* **The dose-response is flat and slightly negative for longer keys.**
|
||
`len16` is −3.7 dmg/run and −0.07 wins/run (damage sign test p=0.30; if
|
||
anything it drifts toward the linear fallback as predicted in Batch 1).
|
||
* **No shipped default was touched**; this job is a measurement.
|
||
|
||
#### Pre-registered predictions — scorecard
|
||
|
||
| # | prediction | outcome |
|
||
|---|---|---|
|
||
| 1 | `len6`'s +0.49 wins does not replicate; delta collapses toward the MDE | **CORRECT** (+0.49 → −0.04 wins, p=1; damage +9.6 → +1.8) |
|
||
| 2 | `rad_offset` is not a detectable win | **CORRECT** (it is nominally *negative*: −0.20 wins, −2.7 dmg) |
|
||
| 3 | `len16` stays within its MDE on both primaries | **CORRECT** (−3.7 dmg vs MDE 7.05; −0.07 wins) |
|
||
| 4 | damage stays within the MDE for every arm | **CORRECT** |
|
||
| 5 | if `len6` replicates it is the champion candidate | **not invoked** |
|
||
|
||
## Phase 2 — final direct answer
|
||
|
||
> **MEASURED (two sessions, 15 then 33 opponents, 666 battles, 6 frozen-binary
|
||
> arms + the reference): NO — the shipped `Pattern` cannot be improved by tuning
|
||
> its own match-length or history parameters at the resolution we can measure.**
|
||
> The incumbent is already tuned on this axis.
|
||
|
||
* The one Batch-1 signal (`len6`, +0.49 wins/run on 15 opponents, sign p=0.039,
|
||
sign-flip p=0.008) **reversed to a wash** on 33 opponents (−0.04 wins/run,
|
||
p=1; +1.8 dmg/run, p=0.30). It was a session-level low reference, not a gun
|
||
effect.
|
||
* On damage/run and round wins, the best Batch-2 challenger is **inside its MDE
|
||
on both primaries** (len6: +1.8 dmg/run vs MDE 8.34, −0.04 wins/run vs MDE
|
||
0.26). Longer keys (`len16`) and a shallower search (`depth100`) do not help.
|
||
* The two pre-existing radial knobs remain what j99 measured: bearing-invariant
|
||
and not winners; in Batch 2 `rad_offset` is nominally *negative*.
|
||
* **No shipped default changed.** Exposing the knobs (`TR_PATTERN_LEN`,
|
||
`TR_PATTERN_DEPTH`) does not change the default path — proven byte-identical
|
||
before any battle.
|
||
|
||
*(MEASURED vs INFERRED. The knob exposure, the byte-identity parity dump, the
|
||
270- and 396-battle sessions, every per-opponent delta, the CIs, MDEs, sign
|
||
tests and sign-flip tests above are all MEASURED. The "incumbent is already
|
||
tuned" reading is a MEASURED null at these resolutions; any claim that no
|
||
setting whatsoever could ever help would be INFERRED and is deliberately not
|
||
made.)*
|
||
|
||
## What to try next (rewritten AFTER Phase 2)
|
||
|
||
The phase-1 conclusion stands and is now joined by the phase-2 one:
|
||
|
||
1. **The match-shape axis is CLOSED.** Match-key lengths 6 and 16 and a
|
||
shallow (100-tick) history search do not beat the shipped defaults (10,
|
||
full 500). Do not spend another batch on `TR_PATTERN_LEN` / `TR_PATTERN_DEPTH`
|
||
unless a specific mechanism is named; the MDE is ~8 dmg/run and ~0.26
|
||
wins/run, so only a very large effect could even be seen.
|
||
2. **`TR_MOVEMENT=strafe` + default `Pattern` is the measured optimum of the
|
||
tested design space**, and Pattern is now measured-tuned on its own
|
||
parameters as well.
|
||
3. **The named open axes remain the ones phase 1 named**: a *different*
|
||
information source for the lead (not another knob on BitBrain/TMHorizon,
|
||
both measured negative/neutral), and the owner's **spinner-specific claim**
|
||
(j124 is now on it). The lead-amplitude axis and the radial axis stay dead.
|
||
4. **Session-level reference variance is large and must be respected.** The same
|
||
frozen reference on the same movement measured 42.2% wins (Batch 1, 15
|
||
opponents) and 70.4% wins (Batch 2, 33 opponents, superset). Any wins-only
|
||
signal from one session must be replicated before it is believed — this
|
||
job's Batch-1 `len6` is the second campaign example (after j121's `bitbrain`
|
||
damage hint) of a single-session wobble dying on replication.
|
||
|
||
## Session log (Phase 2 rows appended)
|
||
|
||
| session | commit | battles | arms | verdict |
|
||
|---|---|---:|---|---|
|
||
| `/tmp/ab/j123_b1` | `2a98aba` | 270 (0 failed, 0 never started, 0 excluded) | pattern, len6, len16, depth100, rad_offset, rad_scale | `len6` the only rule-2 win (wins +0.49, sign p=0.039); **red flag: all 5 arms wins-positive incl. the bearing-invariant `rad_offset`** |
|
||
| `/tmp/ab/j123_b2` | `1b59b65` | 396 (0 failed, 0 never started, 0 excluded) | pattern, len6, len16, rad_offset | **`len6` does not replicate** (−0.04 wins, p=1; +1.8 dmg, p=0.30); all arms not distinguishable; the incumbent is already tuned |
|
||
|
||
|
||
---
|
||
|
||
## Allocation: is the selector leaving value on the table?
|
||
|
||
**Question (owner-spotted, 2026-09-26).** The selector's own `/tmp/gun_stats.jsonl`
|
||
(10,608 rounds, 2-gun racks, selector ON) gives Pattern ~66% of selection ticks and
|
||
that *looks* justified by per-shot hit rate. But `GUN_SELECTOR_FLOOR` is a FITNESS
|
||
floor (only ignores a gun below 25% of the best), **not a SHARE floor**: nothing
|
||
guarantees the second gun any share of the shots. The ~66/34 split is an OUTCOME of
|
||
ranking + hysteresis, not a POLICY. **Nobody has ever tested whether a deliberately
|
||
different split beats `pattern` alone.** This batch builds the mechanism
|
||
(`TR_RACK_SHARE`, default-off) and asks directly:
|
||
|
||
> **Is gun ALLOCATION a lever?**
|
||
|
||
### Pre-registration (written BEFORE the batch)
|
||
|
||
* **Mechanism.** `TR_RACK_SHARE=pattern:50,bitbrain:50` (relative weights,
|
||
case-insensitive `RackGunNames`) replaces `chooseFromFit`'s ranking with a
|
||
deterministic **deficit round-robin** over the named **ADMITTED** guns. Each
|
||
allocation is held for `GUN_SELECTOR_DWELL` ticks (turret convergence) so the
|
||
share is over dwell epochs; selected-tick counts follow the weights.
|
||
Deterministic (not randomized) because the question is whether a *specified*
|
||
split is better, the schedule is auditable, and it needs no seed.
|
||
* **Default-OFF parity.** Unset `TR_RACK_SHARE` leaves the weights empty and
|
||
`selectGun` takes the unchanged ranking path. Parity proven before fighting:
|
||
`test_gun_harness 39`, `test_rack_membership 48`, `test_selector_tiebreak 19`,
|
||
`test_env_report 25` — all green at their exact shipped counts.
|
||
* **Panel.** `tools/ab/panel_movement.txt` — the frozen 15-opponent movement panel
|
||
(never edited). Unit of evidence = number of opponents.
|
||
* **Arms** (`tools/ab/arms_alloc.txt`, 6 arms): `pattern` (reference, shipped);
|
||
`sel_pb` (selector ON, Pattern+BitBrain, shipped policy = the ~66/34 OUTCOME);
|
||
`share_5050_pb`; `share_7030_pb`; `share_5050_pt` (TMHorizon); `share_5050_pk`
|
||
(KNN). Movement pinned `TR_MOVEMENT=strafe` in every arm.
|
||
* **Judge on damage/run and ROUND WINS** (standing rule); hit rate and distance
|
||
are explanation. Per-opponent deltas vs the reference; CI, sign test, sign-flip
|
||
permutation, MDE.
|
||
* **Decision rule (pre-registered).**
|
||
1. If a **forced-share arm beats `pattern`** (CI excluding 0 **and** sign-flip
|
||
p<0.05) → **ALLOCATION IS A LEVER**; the selector was leaving value on the
|
||
table and the next job tunes the share.
|
||
2. If **no forced-share arm beats `pattern`** → the ~66/34 allocation was
|
||
already right; the rack question is genuinely CLOSED.
|
||
3. Either way, report whether the forced arms differ from `sel_pb` — that
|
||
isolates `policy` from `gun choice`.
|
||
* **Allocation evidence.** Per arm, the **applied share** (from each run's
|
||
`gun_stats.jsonl` `selected` counts) and **each gun's per-shot hit rate**
|
||
(realShots/realHits, same battles). Per-shot hit rate is the right signal HERE
|
||
because the movement is identical inside a battle.
|
||
|
||
### RESULTS
|
||
|
||
**Session** `/tmp/ab/j127_alloc2`, commit `2a3a62a`, frozen binary sha256 `1b663dbbd365…`,
|
||
15 opponents × 6 arms × 3 runs × 3 rounds = **270 battles, 0 failed, 0 never started,
|
||
0 liveness exclusions** (827 s wall). Analyzer: `tools/ab/tournament_analyze.py`.
|
||
|
||
**Mechanism bug found and fixed before the ledger run (MEASURED).** The first
|
||
session (`/tmp/ab/j127_alloc`) exposed that `bot.tick` RESETS to 0 each round
|
||
while `tracker.currentSince` persists, so the dwell hold `(tick - currentSince) <
|
||
DWELL` went negative and LOCKED the round-1-end gun for every later round
|
||
(per-round applied share was `[50, 100, 100]` instead of `[50, 50, 50]`). The
|
||
share path now requires `tick >= currentSince` (commit `2a3a62a`); the shipped
|
||
ranking path is untouched. **The same latent lock exists in the shipped
|
||
selector's hysteresis for any multi-gun rack** — the `sel_pb` arm's per-round
|
||
share is likewise bimodal (e.g. `[98, 100]`, `[91, 0, 51]`), which is why its
|
||
aggregate ~69/31 is itself a round-boundary artifact, not a stable policy. This
|
||
is a MEASURED property of the shipped mechanism, not fixed here (no shipping).
|
||
|
||
**Applied share and per-gun per-shot hit rate (MEASURED, per-run `gun_stats.jsonl`).**
|
||
All 45 runs per forced arm print `TR_RACK_SHARE active`; 0 errors, 0 unrecognized.
|
||
The forced shares land on target (50.1/49.9, 70.0/30.0). Per-shot hit rate is
|
||
`realHits/realShots` summed over the same battles.
|
||
|
||
| arm | applied share (`selected` ticks) | shots | per-gun per-shot hit |
|
||
|---|---|---:|---|
|
||
| `pattern` | Pattern 100.0% | 7016 | Pattern **15.94%** |
|
||
| `sel_pb` (shipped policy) | Pattern 69.4% / BitBrain 30.6% | 4470 / 2024 | Pattern **19.08%** / BitBrain 12.85% |
|
||
| `share_5050_pb` | Pattern 50.1% / BitBrain 49.9% | 3034 / 3042 | Pattern **17.77%** / BitBrain 14.14% |
|
||
| `share_7030_pb` | Pattern 70.0% / BitBrain 30.0% | 4405 / 1877 | Pattern **17.68%** / BitBrain 16.36% |
|
||
| `share_5050_pt` | Pattern 50.1% / TMHorizon 49.9% | 3404 / 3407 | Pattern **18.18%** / TMHorizon 15.59% |
|
||
| `share_5050_pk` | Pattern 50.1% / KNN 49.9% | 3164 / 2633 | Pattern **17.98%** / KNN 12.53% |
|
||
|
||
**Per-opponent results (Δdmg/run / Δwins/run, arm − `pattern`).**
|
||
|
||
| opponent | `sel_pb` | `share_5050_pb` | `share_7030_pb` | `share_5050_pt` | `share_5050_pk` |
|
||
|---|---:|---:|---:|---:|---:|
|
||
| DrussGT | +26.5 / +0.67 | -36.8 / -0.67 | -10.3 / -0.33 | -1.1 / -0.33 | -5.1 / +0.33 |
|
||
| Diamond | +7.6 / +0.33 | -4.1 / -0.33 | -15.0 / +0.00 | +12.2 / -0.33 | -6.3 / -0.33 |
|
||
| Dookious | +5.0 / +0.00 | -12.6 / -0.33 | -13.7 / -0.33 | -12.9 / -0.67 | -13.9 / +0.00 |
|
||
| GresSuffurd | +29.7 / +0.33 | -3.0 / -0.33 | +16.5 / -0.67 | +23.6 / +0.00 | -11.4 / -0.67 |
|
||
| CassiusClay | +25.7 / +0.33 | +33.0 / +0.33 | +21.4 / +0.67 | +29.5 / +0.67 | -7.5 / +0.00 |
|
||
| RetroGirl | +13.5 / +1.33 | -19.6 / +1.00 | +7.7 / +0.67 | +73.7 / +2.00 | +20.4 / +0.67 |
|
||
| TripHammer | +12.9 / +0.33 | +1.1 / +0.33 | +4.8 / +0.67 | -6.8 / +0.33 | +1.2 / +0.00 |
|
||
| Coriantumr | -1.9 / -0.33 | -25.7 / -0.67 | -16.9 / +0.33 | -13.3 / -1.33 | -44.9 / -1.67 |
|
||
| WallAvoider | -31.8 / +1.00 | +28.9 / +1.33 | +9.0 / +1.00 | -8.2 / +1.00 | -27.7 / -0.33 |
|
||
| HawkOnFire | +0.2 / -1.33 | +4.7 / -0.33 | +37.6 / +0.00 | -9.3 / -0.33 | -41.2 / -2.00 |
|
||
| SpinBot | -14.7 / +0.00 | +12.7 / +0.00 | +6.0 / +0.00 | +16.0 / +0.00 | -15.3 / +0.00 |
|
||
| DiamondStealer | +27.4 / +0.00 | +16.0 / +0.67 | -24.1 / -0.33 | -4.3 / +0.00 | +8.9 / +0.67 |
|
||
| BlitzBat | +3.5 / +0.33 | +26.3 / +0.00 | +10.2 / +0.33 | +11.1 / +0.33 | +10.6 / +0.00 |
|
||
| YersiniaPestis | -18.5 / -1.33 | -15.9 / -0.67 | -11.9 / -0.67 | +4.4 / -0.33 | -0.3 / -1.00 |
|
||
| Ascendant | -0.0 / +0.00 | -5.3 / -0.33 | +31.4 / +1.67 | -6.5 / +0.33 | -33.1 / -0.33 |
|
||
|
||
**Cross-opponent tests/CI/MDE vs `pattern` (the verdict layer).** Primary metrics
|
||
only; `p` is the exact sign-flip permutation on the per-opponent mean.
|
||
|
||
| arm | metric | mean Δ | 95% CI | sign test | p(sign-flip) | MDE | verdict |
|
||
|---|---|---:|---|---:|---:|---:|---|
|
||
| `sel_pb` | damage | +5.68 | [-4.29, +15.64] | 10/15 | 0.2393 | 13.01 | not distinguishable |
|
||
| `sel_pb` | wins | +0.11 | [-0.29, +0.51] | 8/11 | 0.6387 | 0.52 | not distinguishable |
|
||
| `share_5050_pb` | damage | -0.02 | [-11.41, +11.36] | 7/15 | 0.9968 | 14.87 | not distinguishable |
|
||
| `share_5050_pb` | wins | -0.00 | [-0.34, +0.34] | 5/13 | 1 | 0.45 | not distinguishable |
|
||
| `share_7030_pb` | damage | +3.50 | [-6.73, +13.74] | 9/15 | 0.4725 | 13.37 | not distinguishable |
|
||
| `share_7030_pb` | wins | +0.20 | [-0.16, +0.56] | 7/12 | 0.3193 | 0.47 | not distinguishable |
|
||
| `share_5050_pt` | damage | +7.21 | [-5.37, +19.80] | 7/15 | 0.2582 | 16.44 | not distinguishable |
|
||
| `share_5050_pt` | wins | +0.09 | [-0.34, +0.52] | 6/12 | 0.7539 | 0.56 | not distinguishable |
|
||
| `share_5050_pk` | damage | -11.03 | **[-21.54, -0.52]** | 4/15 | **0.04205** | 13.72 | **WORSE (damage)** |
|
||
| `share_5050_pk` | wins | -0.31 | [-0.73, +0.11] | 3/10 | 0.1777 | 0.55 | not distinguishable |
|
||
|
||
**Policy isolation — forced arms vs `sel_pb` (MEASURED, `--reference sel_pb`).**
|
||
This is the arm that holds gun choice fixed and asks whether the POLICY (ranking
|
||
vs forced schedule) changes anything. `share_7030_pb` deliberately reproduces
|
||
`sel_pb`'s applied share (70.0/30.0 vs 69.4/30.6), so it is the cleanest test.
|
||
|
||
| arm | metric | mean Δ | 95% CI | sign test | p(sign-flip) | MDE |
|
||
|---|---|---:|---|---:|---:|---:|
|
||
| `share_7030_pb` | damage | -2.17 | [-16.90, +12.55] | 6/15 | 0.7542 | 19.23 |
|
||
| `share_7030_pb` | wins | +0.09 | [-0.34, +0.52] | 6/12 | 0.7432 | 0.56 |
|
||
| `share_5050_pb` | damage | -5.70 | [-21.93, +10.54] | 6/15 | 0.4659 | 21.20 |
|
||
| `share_5050_pb` | wins | -0.11 | [-0.44, +0.22] | 4/12 | 0.5771 | 0.43 |
|
||
| `share_5050_pt` | damage | +1.54 | [-12.09, +15.16] | 7/15 | 0.8182 | 17.80 |
|
||
| `share_5050_pt` | wins | -0.02 | [-0.37, +0.33] | 5/10 | 1 | 0.46 |
|
||
| `share_5050_pk` | damage | -16.71 | **[-27.92, -5.50]** | 4/15 | **0.00769** | 14.64 |
|
||
| `share_5050_pk` | wins | -0.42 | **[-0.73, -0.11]** | 2/13 | **0.01733** | 0.40 |
|
||
|
||
Pooled dashboard (descriptive, not the verdict): `pattern` 101.9 dmg/1.47 wins;
|
||
`sel_pb` 107.5/1.58; `share_5050_pb` 101.8/1.47; `share_7030_pb` 105.4/1.67;
|
||
`share_5050_pt` 109.1/1.56; `share_5050_pk` 90.8/1.16.
|
||
|
||
### VERDICT
|
||
|
||
> **DIRECT ANSWER (MEASURED): NO — gun ALLOCATION is not a lever on the frozen
|
||
> 15-opponent panel at this resolution. No forced-share arm beats `pattern`
|
||
> alone by the pre-registered rule (CI excluding 0 AND sign-flip p<0.05). The
|
||
> selector's ~66/34 split was already right within the MDE; the rack question is
|
||
> genuinely CLOSED.**
|
||
|
||
* The best forced arm is `share_7030_pb` (applied 70/30, i.e. re-creating the
|
||
observed ~66/34 split deliberately): **+0.20 wins/run, +3.5 dmg/run vs
|
||
`pattern`, both inside their MDEs** (0.47 wins, 13.4 dmg), sign-flip p=0.32 /
|
||
0.47. A nominal, non-significant lean toward the incumbent's own split.
|
||
* The deliberate 50/50 arms gain nothing (`share_5050_pb` ≈ 0/0) or cost damage
|
||
(`share_5050_pt` +7.2 dmg but +0.09 wins, not distinguishable).
|
||
* The only DETECTABLE allocation effect is NEGATIVE: forcing 50/50 with KNN is
|
||
detectably WORSE on damage vs `pattern` (−11.0, CI excludes 0, p=0.042) and
|
||
detectably worse than `sel_pb` on BOTH wins and damage (−0.42 wins, p=0.017;
|
||
−16.7 dmg, p=0.008). Allocation can HURT; it cannot help here.
|
||
* **`policy` vs `gun choice` (MEASURED).** Forcing the selector's own applied
|
||
share (`share_7030_pb` ≈ `sel_pb`) is not distinguishable on either primary,
|
||
so at the same allocation the ranking-vs-forced *policy* makes no measurable
|
||
difference. `sel_pb` itself is nominally positive (+0.11 wins, +5.7 dmg) but
|
||
not distinguishable — this session's reference is a low one (48.9% wins),
|
||
consistent with the campaign's known session-level reference variance.
|
||
* **Caveat that strengthens the finding (MEASURED).** The applied share of the
|
||
*shipped* selector is not a stable policy at all: its round-boundary dwell
|
||
lock makes per-round shares bimodal (a gun can hold 100% of a round). The
|
||
forced arms, with the fix, are the first clean allocation measurement — and
|
||
they show no upside to deviating from Pattern-dominant.
|
||
|
||
**MEASURED:** the allocator code + default-OFF parity (39/48/19/25), the
|
||
round-boundary bug and its fix, all 4 forced arms' applied shares (exactly
|
||
50/50 and 70/30) and per-gun per-shot hit rates, the 270-battle session, every
|
||
per-opponent delta, CI, MDE, sign and sign-flip test, and the policy-isolation
|
||
comparison. **INFERRED:** that no share schedule whatsoever could ever help
|
||
(this batch bounds only the shares it tested, at this MDE and panel).
|