diff --git a/docs/gun_campaign.md b/docs/gun_campaign.md index 32f3dce..36c424e 100644 --- a/docs/gun_campaign.md +++ b/docs/gun_campaign.md @@ -36,6 +36,7 @@ controls. | other single guns | `knn` detectably worse; `bitbrain` a wash measured 3× (−3.6 / +10.3 / +1.8 dmg/run, all inside their MDEs); `TMHorizon` worse on 33 (`docs/gauntlet_bitbrain_vs_pattern.md`, Batches 1–2) | | 2-gun selector racks | `rack_pk` nominally below `pattern` on wins; `rack_pt` == `tmhorizon`-alone → the selector just picks the corrector (`## Batch 1`) | | selector mechanics | the virtual-fitness selector is negative value at every rack size tested (16, lean8, lean6, pairs); it never manufactured a win (`docs/selector_negative_value.md`) | +| **gun allocation (share)** | **forced `TR_RACK_SHARE` schedules (50/50, 70/30; PB/PT/PK) never beat `pattern` by the pre-registered rule; the only detectable effect is negative (50/50 KNN, −11 dmg/run, p=0.042); the observed ~66/34 split was already right (`## Allocation`) | | lead amplitude | a full gain sweep (1.0/1.5/2.0/3.0) makes Pattern strictly worse at every range band — the lever is lead *information*, not amplitude (`docs/bitbrain_campaign.md` §0.3.2) | | Pattern's own match-length / history params | `len6`'s single-session +0.49 wins (p=0.039 on 15) reversed to a wash on 33 (−0.04, p=1); `len16`/`depth100` flat; two batches, no replication (`## Phase 2`) | | radial knobs (`TR_PATTERN_RAD_*`) | bearing-invariant **by construction** (job j99 proved `bmPath` is a structural no-op and scale/offset cannot change the aim bearing); in Batch 2 `rad_offset` is nominally *negative* | @@ -1030,4 +1031,125 @@ different split beats `pattern` alone.** This batch builds the mechanism ### RESULTS -_(filled in after the batch)_ +**Session** `/tmp/ab/j127_alloc2`, commit `2a3a62a`, frozen binary sha256 `1b663dbbd365…`, +15 opponents × 6 arms × 3 runs × 3 rounds = **270 battles, 0 failed, 0 never started, +0 liveness exclusions** (827 s wall). Analyzer: `tools/ab/tournament_analyze.py`. + +**Mechanism bug found and fixed before the ledger run (MEASURED).** The first +session (`/tmp/ab/j127_alloc`) exposed that `bot.tick` RESETS to 0 each round +while `tracker.currentSince` persists, so the dwell hold `(tick - currentSince) < +DWELL` went negative and LOCKED the round-1-end gun for every later round +(per-round applied share was `[50, 100, 100]` instead of `[50, 50, 50]`). The +share path now requires `tick >= currentSince` (commit `2a3a62a`); the shipped +ranking path is untouched. **The same latent lock exists in the shipped +selector's hysteresis for any multi-gun rack** — the `sel_pb` arm's per-round +share is likewise bimodal (e.g. `[98, 100]`, `[91, 0, 51]`), which is why its +aggregate ~69/31 is itself a round-boundary artifact, not a stable policy. This +is a MEASURED property of the shipped mechanism, not fixed here (no shipping). + +**Applied share and per-gun per-shot hit rate (MEASURED, per-run `gun_stats.jsonl`).** +All 45 runs per forced arm print `TR_RACK_SHARE active`; 0 errors, 0 unrecognized. +The forced shares land on target (50.1/49.9, 70.0/30.0). Per-shot hit rate is +`realHits/realShots` summed over the same battles. + +| arm | applied share (`selected` ticks) | shots | per-gun per-shot hit | +|---|---|---:|---| +| `pattern` | Pattern 100.0% | 7016 | Pattern **15.94%** | +| `sel_pb` (shipped policy) | Pattern 69.4% / BitBrain 30.6% | 4470 / 2024 | Pattern **19.08%** / BitBrain 12.85% | +| `share_5050_pb` | Pattern 50.1% / BitBrain 49.9% | 3034 / 3042 | Pattern **17.77%** / BitBrain 14.14% | +| `share_7030_pb` | Pattern 70.0% / BitBrain 30.0% | 4405 / 1877 | Pattern **17.68%** / BitBrain 16.36% | +| `share_5050_pt` | Pattern 50.1% / TMHorizon 49.9% | 3404 / 3407 | Pattern **18.18%** / TMHorizon 15.59% | +| `share_5050_pk` | Pattern 50.1% / KNN 49.9% | 3164 / 2633 | Pattern **17.98%** / KNN 12.53% | + +**Per-opponent results (Δdmg/run / Δwins/run, arm − `pattern`).** + +| opponent | `sel_pb` | `share_5050_pb` | `share_7030_pb` | `share_5050_pt` | `share_5050_pk` | +|---|---:|---:|---:|---:|---:| +| DrussGT | +26.5 / +0.67 | -36.8 / -0.67 | -10.3 / -0.33 | -1.1 / -0.33 | -5.1 / +0.33 | +| Diamond | +7.6 / +0.33 | -4.1 / -0.33 | -15.0 / +0.00 | +12.2 / -0.33 | -6.3 / -0.33 | +| Dookious | +5.0 / +0.00 | -12.6 / -0.33 | -13.7 / -0.33 | -12.9 / -0.67 | -13.9 / +0.00 | +| GresSuffurd | +29.7 / +0.33 | -3.0 / -0.33 | +16.5 / -0.67 | +23.6 / +0.00 | -11.4 / -0.67 | +| CassiusClay | +25.7 / +0.33 | +33.0 / +0.33 | +21.4 / +0.67 | +29.5 / +0.67 | -7.5 / +0.00 | +| RetroGirl | +13.5 / +1.33 | -19.6 / +1.00 | +7.7 / +0.67 | +73.7 / +2.00 | +20.4 / +0.67 | +| TripHammer | +12.9 / +0.33 | +1.1 / +0.33 | +4.8 / +0.67 | -6.8 / +0.33 | +1.2 / +0.00 | +| Coriantumr | -1.9 / -0.33 | -25.7 / -0.67 | -16.9 / +0.33 | -13.3 / -1.33 | -44.9 / -1.67 | +| WallAvoider | -31.8 / +1.00 | +28.9 / +1.33 | +9.0 / +1.00 | -8.2 / +1.00 | -27.7 / -0.33 | +| HawkOnFire | +0.2 / -1.33 | +4.7 / -0.33 | +37.6 / +0.00 | -9.3 / -0.33 | -41.2 / -2.00 | +| SpinBot | -14.7 / +0.00 | +12.7 / +0.00 | +6.0 / +0.00 | +16.0 / +0.00 | -15.3 / +0.00 | +| DiamondStealer | +27.4 / +0.00 | +16.0 / +0.67 | -24.1 / -0.33 | -4.3 / +0.00 | +8.9 / +0.67 | +| BlitzBat | +3.5 / +0.33 | +26.3 / +0.00 | +10.2 / +0.33 | +11.1 / +0.33 | +10.6 / +0.00 | +| YersiniaPestis | -18.5 / -1.33 | -15.9 / -0.67 | -11.9 / -0.67 | +4.4 / -0.33 | -0.3 / -1.00 | +| Ascendant | -0.0 / +0.00 | -5.3 / -0.33 | +31.4 / +1.67 | -6.5 / +0.33 | -33.1 / -0.33 | + +**Cross-opponent tests/CI/MDE vs `pattern` (the verdict layer).** Primary metrics +only; `p` is the exact sign-flip permutation on the per-opponent mean. + +| arm | metric | mean Δ | 95% CI | sign test | p(sign-flip) | MDE | verdict | +|---|---|---:|---|---:|---:|---:|---| +| `sel_pb` | damage | +5.68 | [-4.29, +15.64] | 10/15 | 0.2393 | 13.01 | not distinguishable | +| `sel_pb` | wins | +0.11 | [-0.29, +0.51] | 8/11 | 0.6387 | 0.52 | not distinguishable | +| `share_5050_pb` | damage | -0.02 | [-11.41, +11.36] | 7/15 | 0.9968 | 14.87 | not distinguishable | +| `share_5050_pb` | wins | -0.00 | [-0.34, +0.34] | 5/13 | 1 | 0.45 | not distinguishable | +| `share_7030_pb` | damage | +3.50 | [-6.73, +13.74] | 9/15 | 0.4725 | 13.37 | not distinguishable | +| `share_7030_pb` | wins | +0.20 | [-0.16, +0.56] | 7/12 | 0.3193 | 0.47 | not distinguishable | +| `share_5050_pt` | damage | +7.21 | [-5.37, +19.80] | 7/15 | 0.2582 | 16.44 | not distinguishable | +| `share_5050_pt` | wins | +0.09 | [-0.34, +0.52] | 6/12 | 0.7539 | 0.56 | not distinguishable | +| `share_5050_pk` | damage | -11.03 | **[-21.54, -0.52]** | 4/15 | **0.04205** | 13.72 | **WORSE (damage)** | +| `share_5050_pk` | wins | -0.31 | [-0.73, +0.11] | 3/10 | 0.1777 | 0.55 | not distinguishable | + +**Policy isolation — forced arms vs `sel_pb` (MEASURED, `--reference sel_pb`).** +This is the arm that holds gun choice fixed and asks whether the POLICY (ranking +vs forced schedule) changes anything. `share_7030_pb` deliberately reproduces +`sel_pb`'s applied share (70.0/30.0 vs 69.4/30.6), so it is the cleanest test. + +| arm | metric | mean Δ | 95% CI | sign test | p(sign-flip) | MDE | +|---|---|---:|---|---:|---:|---:| +| `share_7030_pb` | damage | -2.17 | [-16.90, +12.55] | 6/15 | 0.7542 | 19.23 | +| `share_7030_pb` | wins | +0.09 | [-0.34, +0.52] | 6/12 | 0.7432 | 0.56 | +| `share_5050_pb` | damage | -5.70 | [-21.93, +10.54] | 6/15 | 0.4659 | 21.20 | +| `share_5050_pb` | wins | -0.11 | [-0.44, +0.22] | 4/12 | 0.5771 | 0.43 | +| `share_5050_pt` | damage | +1.54 | [-12.09, +15.16] | 7/15 | 0.8182 | 17.80 | +| `share_5050_pt` | wins | -0.02 | [-0.37, +0.33] | 5/10 | 1 | 0.46 | +| `share_5050_pk` | damage | -16.71 | **[-27.92, -5.50]** | 4/15 | **0.00769** | 14.64 | +| `share_5050_pk` | wins | -0.42 | **[-0.73, -0.11]** | 2/13 | **0.01733** | 0.40 | + +Pooled dashboard (descriptive, not the verdict): `pattern` 101.9 dmg/1.47 wins; +`sel_pb` 107.5/1.58; `share_5050_pb` 101.8/1.47; `share_7030_pb` 105.4/1.67; +`share_5050_pt` 109.1/1.56; `share_5050_pk` 90.8/1.16. + +### VERDICT + +> **DIRECT ANSWER (MEASURED): NO — gun ALLOCATION is not a lever on the frozen +> 15-opponent panel at this resolution. No forced-share arm beats `pattern` +> alone by the pre-registered rule (CI excluding 0 AND sign-flip p<0.05). The +> selector's ~66/34 split was already right within the MDE; the rack question is +> genuinely CLOSED.** + +* The best forced arm is `share_7030_pb` (applied 70/30, i.e. re-creating the + observed ~66/34 split deliberately): **+0.20 wins/run, +3.5 dmg/run vs + `pattern`, both inside their MDEs** (0.47 wins, 13.4 dmg), sign-flip p=0.32 / + 0.47. A nominal, non-significant lean toward the incumbent's own split. +* The deliberate 50/50 arms gain nothing (`share_5050_pb` ≈ 0/0) or cost damage + (`share_5050_pt` +7.2 dmg but +0.09 wins, not distinguishable). +* The only DETECTABLE allocation effect is NEGATIVE: forcing 50/50 with KNN is + detectably WORSE on damage vs `pattern` (−11.0, CI excludes 0, p=0.042) and + detectably worse than `sel_pb` on BOTH wins and damage (−0.42 wins, p=0.017; + −16.7 dmg, p=0.008). Allocation can HURT; it cannot help here. +* **`policy` vs `gun choice` (MEASURED).** Forcing the selector's own applied + share (`share_7030_pb` ≈ `sel_pb`) is not distinguishable on either primary, + so at the same allocation the ranking-vs-forced *policy* makes no measurable + difference. `sel_pb` itself is nominally positive (+0.11 wins, +5.7 dmg) but + not distinguishable — this session's reference is a low one (48.9% wins), + consistent with the campaign's known session-level reference variance. +* **Caveat that strengthens the finding (MEASURED).** The applied share of the + *shipped* selector is not a stable policy at all: its round-boundary dwell + lock makes per-round shares bimodal (a gun can hold 100% of a round). The + forced arms, with the fix, are the first clean allocation measurement — and + they show no upside to deviating from Pattern-dominant. + +**MEASURED:** the allocator code + default-OFF parity (39/48/19/25), the +round-boundary bug and its fix, all 4 forced arms' applied shares (exactly +50/50 and 70/30) and per-gun per-shot hit rates, the 270-battle session, every +per-opponent delta, CI, MDE, sign and sign-flip test, and the policy-isolation +comparison. **INFERRED:** that no share schedule whatsoever could ever help +(this batch bounds only the shares it tested, at this MDE and panel).