gun j123 Task A+B prereg: expose Pattern's TR_PATTERN_LEN/TR_PATTERN_DEPTH (default parity) and pre-register the 6-arm match-parameter sweep

This commit is contained in:
2026-09-26 04:01:51 +02:00
parent 467e07a6a4
commit 2a98aba91b
4 changed files with 231 additions and 10 deletions
+106
View File
@@ -464,3 +464,109 @@ different reference is a free pairwise comparison with no battles.
|---|---|---:|---|---|
| `/tmp/ab/j121_g1` | `1d8143a` | 270 (0 failed, 0 never started, 0 excluded) | pattern, bitbrain, tmhorizon, knn, rack_pk, rack_pt | **nothing beats the shipped `pattern`**; all arms not distinguishable except `knn` = WORSE; `onlyPattern` CONFIRMED |
| `/tmp/ab/j121_g2` | `c343c00` | 297 (0 failed, 0 never started, 0 excluded) | pattern, bitbrain, tmhorizon | **the Batch-1 damage hint does not survive**: `bitbrain` +1.8 dmg/run p=0.49 (wash); `tmhorizon` −10.4 dmg/run p=0.035 (WORSE); wins flat |
---
# Phase 2: tuning the incumbent
**Owner's mandate (phase 2):** the phase-1 campaign closed the "other gun / other
rack" design space (nothing beats the shipped `Pattern`; the correctors are a
wash or a loss). What remains OPEN is the **incumbent's own tuning**: Pattern's
match-length / history parameters had NEVER been swept. Phase 2 asks the direct
question:
> **Can the shipped `Pattern` be improved by tuning its own match parameters —
> and if so, by how much on damage/run and round wins?**
The lead-amplitude axis is already known dead (`docs/bitbrain_campaign.md`), and
the radial knobs are known non-winners (job j99: `bmPath` structural no-op; live
+0.28 pp p=0.62 for scale 0.98, −0.42 pp p=0.46 for offset −20). Those are
therefore **controls** here, not candidates.
## Task A — exposing Pattern's match-shape parameters (what and why)
**MEASURED (code read).** `common_libs/guns/pattern_matcher.nim` had exactly two
tunable knobs, both RADIAL (`TR_PATTERN_RAD_SCALE`, `TR_PATTERN_RAD_OFFSET`), and
neither can change the lead bearing (job j99 proved the bearing is untouched), so
neither was ever the "lead information" axis. The parameters that actually
control **how the pattern is matched** were compile-time constants:
| parameter | code | what it controls | exposed as |
|---|---|---|---|
| match-key length | `PatternLen = 10` | the length of the movement segment compared (the search key); also how far after the match the replay starts | `TR_PATTERN_LEN` (int, default 10) |
| search depth | implicit `HistorySize = 500` | how far back the best-match scan may reach (`scanEnd` was always `count−PatternLen−1`) | `TR_PATTERN_DEPTH` (int, default 500 = full buffer) |
| similarity/search radius | **does not exist** | the search always takes the single lowest-cost match; there is no acceptance threshold or radius to expose | **not exposed — nothing to expose** |
| history buffer capacity | `HistorySize = 500` | the fixed `array[HistorySize]` backing store | runtime depth limit only; the buffer ceiling cannot be raised at runtime (see below) |
**Why env and not `-d:`.** The phase-1 instrument (`tools/ab/tournament_run.sh`)
builds ONE frozen binary from `git archive HEAD` and every arm differs only by its
env dict. A `{.intdefine.}` knob would need one binary per arm, which the
instrument forbids. Both new knobs are therefore resolved lazily from the
environment on the first `predict`, exactly like the radial knobs, and default to
the pre-knob constants. `setMatchParams(patternLen, histDepth)` is the explicit
offline/unit-test twin that writes the same fields.
**What could NOT be exposed cheaply (MEASURED).** `HistorySize` sizes four fixed
`array[HistorySize(+1)]` fields in the gun object. Raising it above 500 at
runtime is impossible without a heap buffer; the runtime `TR_PATTERN_DEPTH` knob
therefore *lowers* the effective search depth within the existing 500-entry
buffer. A depth above 500 is clamped to 500, and a non-positive or unparsable
value falls back to the shipped default. `PatternLen` is clamped to
`1..HistorySize`.
## Task A — default parity (MEASURED, byte-for-byte)
* **Baseline-vs-new parity dump.** A throwaway harness replayed 800 ticks of
`tools/fixtures/drussgt_vs_crazy.jsonl` through one `PatternMatcherGun` at all
four power bins (3200 predictions) and printed every point at full precision.
Compiled once against a `HEAD` worktree (`/tmp/j123_base`, before the change)
and once against the modified tree: **the two dumps are byte-identical**
(`diff -q` clean). The shipped default path is unchanged.
* **The knobs are live and self-falling-back.** `TR_PATTERN_LEN=6/16` and
`TR_PATTERN_DEPTH=100/20` each change the dump; `TR_PATTERN_LEN=banana`
reproduces the default dump exactly.
* **Existing guards pass** (no new guard tests, per the "cut ceremony" rule):
`common_libs/tests/test_pattern_radial_offset.nim` (6 checks) and
`common_libs/tests/test_gun_harness.nim` (all checks) pass; the env-report
guard `test_env_report.nim` passes with `TR_PATTERN_LEN` / `TR_PATTERN_DEPTH`
registered in `knownEnvNames()` and the effective-values report.
* **Clean-archive compile** is exercised by `tournament_run.sh` itself, which
builds the frozen binary from `git archive HEAD`.
## Task B — Batch 1 (pre-registered BEFORE any battle)
**Design.** One frozen binary, six env-only arms, the FROZEN 15-opponent panel
`tools/ab/panel_movement.txt`, **3 runs × 3 rounds per (opponent, arm) = 270
battles**, `--conc 6`, `--wait-arena`. Arm file: `tools/ab/arms_gun_b3.txt`.
Reference: `pattern`. Movement pinned `TR_MOVEMENT=strafe` in every arm.
| # | arm | env over the pin | what it isolates |
|---|---|---|---|
| 1 | `pattern` | (none) | the arm to beat (shipped `onlyPattern` rack) |
| 2 | `len6` | `TR_PATTERN_LEN=6` | shorter movement segment compared |
| 3 | `len16` | `TR_PATTERN_LEN=16` | longer movement segment compared |
| 4 | `depth100` | `TR_PATTERN_DEPTH=100` | shallower history / match search |
| 5 | `rad_offset` | `TR_PATTERN_RAD_OFFSET=-20` | control: aim 20 px short |
| 6 | `rad_scale` | `TR_PATTERN_RAD_SCALE=0.95` | control: scale the aim distance |
**Pre-registered prediction (written BEFORE the battles finished):**
1. **No arm beats `pattern` on both primaries.** The incumbent is already tuned —
the phase-2 null. The point estimates sit inside the MDE.
2. `len16` is **WORSE or flat** on damage: a 16-tick key matches rarely in a
~500-tick buffer, so the gun falls back to the linear forecast (the weaker
base) more often; wins flat.
3. `len6` is **flat** on damage (more matches but a noisier replay) and flat on
wins; possibly a sub-MDE wobble in either direction.
4. `depth100` is **not distinguishable** from `pattern` — the full buffer already
contains the useful candidates; a sub-MDE damage wobble is possible.
5. The two radial controls are **not distinguishable** and, per job j99, cannot
change the real aim bearing; any damage delta is the fire-gate channel only.
6. If any surprise exists, it is a **damage-only wobble without wins** — the
`ring`-mover trap — and will not be read as a win.
**Pre-registered Batch-2 trigger (task rule).** A Batch-1 arm is promoted to a
higher-power Batch 2 **only if** it beats the reference `pattern` by the
campaign rule 2 (one primary up at cross-opponent sign-test p<0.05 while the
other does not go down) **AND** its damage CI excludes 0 **AND** the sign-flip
permutation test gives p<0.05. Otherwise Batch 1 is the answer.