diff --git a/docs/bitbrain_campaign.md b/docs/bitbrain_campaign.md index 8b465d2..eba4eab 100644 --- a/docs/bitbrain_campaign.md +++ b/docs/bitbrain_campaign.md @@ -418,4 +418,211 @@ here** — the live gate is a separate phase. | D7 | Constant sub-unity lead gain at all ranges | **DEAD — measured** (§1.1: gain < 1 hurts below 300) | -## Phase 2 — *(unclaimed; append below)* +## Phase 2: live lead gains — is a gain ABOVE 1.0 better? *(owner: job j101, committed)* + +> **CORRECTION NOTICE — READ FIRST. Phase 1's offline per-band gain table +> `[1,1,1,0,0]` is LIVE-REFUTED by `140fe25` (`docs/headon_longrange_live.md`). +> Do not act on it.** HeadOn — which *is* gain 0 at every range — scored +> **14 dmg/run vs Pattern's 279**, won **0 of 105 rounds**, and hit **0.4% vs +> Pattern's 9.2%** at 450+. Zero lead above 300 px is a live catastrophe, not a +> +2.16 pp improvement. Phase 1's §1.4 claim ("a per-band lead-gain rule beats +> Pattern offline") is demoted to an offline-only observation that live killed. +> +> **New standing rule: offline is VETO-ONLY** (`docs/offline_harness_trust.md`, +> `e40c849`). It may reject a clearly broken design; it may **never select a +> winner**. Every number in this phase is LIVE. No offline number is cited here +> as evidence of a live win. + +### 2.0 The hypothesis (a hypothesis, not a fact) + +The offline ruler has now been wrong **twice, both times preferring LESS lead** +than reality: Phase 0 said the static gun beats Pattern at long range, and +Phase 1 said gain 0 above 300 px. If the ruler systematically under-values lead, +then it will also have **under-rated gains ABOVE 1.0** — and those had never +been tested live. The hypothesis of this phase is therefore: *Pattern's full +lead is under-shot at long range live, so scaling it up (`gain > 1`) beats +Pattern.* **[INFERRED]** — a reasoned guess from two ruler failures, not a +measurement. + +### 2.1 Method — every arm is pure env on ONE frozen binary + +Task A added `TR_BITBRAIN_GAINS` (comma-separated candidate set, commit +`2747ebd`): unset reproduces the shipped candidate set `[0,.25,.5,.75,1.0]` +exactly, and **exactly ONE value is a FIXED gain with no learning**. BitBrain's +base prediction *is* Pattern (`tmh.pattern.predict`), and the correction is a +multiplicative gain on Pattern's lead over the line of sight, applied only at +range >= 300 px (`BB_GAIN_BAND_MIN`) — the long bands, where the hypothesis +lives. So every arm swaps the admitted rack gun (Pattern off, BitBrain on) and +changes only the lead gain. + +**Session [MEASURED]:** `tools/ab/ab_run.sh --arms tools/ab/arms_leadgain.txt +--runs 7 --outdir /tmp/ab_leadgain --conc 8 --rounds 7`; **commit +`2747ebd`**, frozen binary sha256 `3aa2da14…`; 6 arms x 7 runs x 7 rounds = 42 +battles, 294 rounds, real DrussGT, **42 ok / 0 failed**. Liveness **OK 7/7 runs +for every arm** (boot report shows the arm env verbatim). + +| arm | env (beyond `TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both`) | what it tests | +|---|---|---| +| `control` | (none; shipped Pattern-only rack) | reference | +| `g100` | `TR_BITBRAIN_GAINS=1.0` | **validity check**: fixed gain 1.0 == identity | +| `glo` | `TR_BITBRAIN_GAINS=0.25,0.5,0.75,1.0` | learner restricted to <= 1 | +| `ghi` | `TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0` | learner allowed ABOVE 1 (**hypothesis**) | +| `gfix150` | `TR_BITBRAIN_GAINS=1.5` | fixed 1.5, no learning | +| `gfix125` | `TR_BITBRAIN_GAINS=1.25` | fixed 1.25, no learning | + +### 2.2 The `g100` VALIDITY CHECK — the plumbing is sound [MEASURED] + +Gain 1.0 is the identity, so `g100` **must** be statistically indistinguishable +from `control`. It is: + +| metric | control | g100 | diff | perm p (exact 7v7) | MDE | +|---|---:|---:|---:|---:|---:| +| dmg/run | 259 | 267 | **-8.3** | **0.6492** | 41.4 | +| round wins | 16/49 | 19/49 | **-0.43** | **0.6247** | 1.42 | +| hit rate ALL | 9.9% | 10.0% | **+0.04 pp** | **0.9225** | 1.16 | +| hit rate 450+ | 8.5% | 8.1% | **-0.47 pp** | **0.4656** | 2.01 | + +The `g100` bot emitted **ZERO `[bb]` lines in 7/7 runs** (`bbLog` is only +reached when a non-1.0 gain is applied), i.e. it provably applied no +correction at all — a built-in placebo. **The validity check PASSES: the other +arms are interpretable.** + +### 2.3 Primary result — damage/run and round wins (live decides) [MEASURED] + +| arm | runs | dmg/run | dmgtk/run | round wins | win% | shots/run | +|---|---:|---:|---:|---:|---:|---:| +| `control` | 7 | **259** | 211 | **16/49** | 32.7 | 780 | +| `g100` | 7 | 267 | 200 | 19/49 | 38.8 | 764 | +| `glo` | 7 | 261 | 215 | 19/49 | 38.8 | 783 | +| `ghi` | 7 | **273** | 202 | **22/49** | **44.9** | 796 | +| `gfix150` | 7 | **165** | 235 | **4/49** | **8.2** | 726 | +| `gfix125` | 7 | 225 | 213 | 14/49 | 28.6 | 753 | + +Per-run damage (never just the mean): +`control` 290 281 248 247 222 290 236 · `g100` 265 307 236 228 234 329 273 · +`glo` 215 282 260 282 288 221 282 · `ghi` 271 243 215 297 322 283 282 · +`gfix150` 169 134 164 174 189 152 175 · `gfix125` 174 294 209 204 242 198 252. + +vs `control` (exact permutation, per-run): + +| metric | arm | diff (arm - control) | perm p | MW p | MDE | +|---|---|---:|---:|---:|---:| +| dmg/run | g100 | -8.3 | 0.6492 | 1.0000 | 41.4 | +| dmg/run | glo | -2.5 | 0.8671 | 1.0000 | 41.4 | +| dmg/run | **ghi** | **+14.3** | **0.4091** | 0.5229 | 41.4 | +| dmg/run | **gfix150** | **-93.8** | **0.0006** | 0.0022 | 41.4 | +| dmg/run | gfix125 | -34.2 | 0.0874 | 0.1599 | 41.4 | +| round wins | g100 | -0.43 | 0.6247 | 0.4769 | 1.42 | +| round wins | glo | -0.43 | 0.6329 | 0.5799 | 1.42 | +| round wins | ghi | -0.86 | 0.2756 | 0.2097 | 1.42 | +| round wins | **gfix150** | **+1.71** | **0.0093** | 0.0079 | 1.42 | +| round wins | gfix125 | +0.29 | 0.8042 | 0.8928 | 1.42 | + +**MDE stated plainly: at 7 runs/arm the test only sees large effects** — 41.4 +dmg/run (16% of the control mean) and 1.42 round wins (62% of 2.3). `ghi`'s ++14 dmg/run is a third of the MDE: a live signal smaller than the MDE is NOT a +demonstrated effect. `gfix150`'s -94 dmg/run is 2.3x MDE and is decisive. + +### 2.4 HIT RATE BY RANGE BAND — the load-bearing view [MEASURED] + +`tools/ab/ab_range_bands.py /tmp/ab_leadgain --reference control`. Band = range +at the fire tick. The claim is range-specific; a whole-battle number is not +enough. + +| band (px) | control | g100 | glo | ghi | gfix150 | gfix125 | +|---|---:|---:|---:|---:|---:|---:| +| 0-100 | 1/2 50.0% | 1/2 50.0% | 2/3 66.7% | 0/0 - | 0/1 0.0% | 2/5 40.0% | +| 100-200 | 5/30 16.7% | 3/26 11.5% | 9/33 27.3% | 3/18 16.7% | 3/16 18.8% | 5/19 26.3% | +| 200-300 | 19/98 19.4% | 20/110 18.2% | 16/94 17.0% | 10/99 10.1% | 19/100 19.0% | 15/96 15.6% | +| **300-450** | **237/2081 11.4%** | 248/2012 12.3% | 237/2018 11.7% | 222/2004 11.1% | **145/1946 7.5%** | **183/1925 9.5%** | +| **450+** | **267/3125 8.5%** | 249/3075 8.1% | 256/3211 8.0% | **314/3325 9.4%** | **183/2917 6.3%** | 240/3094 7.8% | +| ALL | 529/5336 9.9% | 521/5225 10.0% | 520/5359 9.7% | **549/5446 10.1%** | **350/4980 7.0%** | 445/5139 8.7% | + +Per-band permutation test on per-run band rates (arm - control), 7v7 exact: + +| band | arm | d(pp) | p | MDE(pp) | +|---|---|---:|---:|---:| +| 300-450 | gfix150 | **-3.85** | **0.0006** | 1.87 | +| 450+ | gfix150 | **-2.25** | **0.0023** | 2.01 | +| 300-450 | gfix125 | **-1.83** | **0.0221** | 1.87 | +| 450+ | gfix125 | -0.81 | 0.1737 | 2.01 | +| 450+ | **ghi** | **+0.83** | **0.3473** | 2.01 | +| 300-450 | ghi | -0.37 | 0.7191 | 1.87 | +| 450+ | glo | -0.56 | 0.3502 | 2.01 | +| 300-450 | glo | +0.24 | 0.8071 | 1.87 | +| ALL | g100 | +0.04 | 0.9225 | 1.16 | + +**The fixed gains above 1.0 clearly LOSE in the exact bands where they are +applied** (300+ px): gfix150 -3.85 pp / -2.25 pp at 2x the MDE, gfix125 -1.83 pp +at 300-450. The learner allowed above 1.0 (`ghi`) is the only arm whose 450+ hit +rate is above control (+0.83 pp) — but **p = 0.35, well inside the MDE**. + +### 2.5 Applied-gain evidence (the knob really moved the gun) [MEASURED] + +Boot report, verbatim: `[env] TR_BITBRAIN_GAINS = 1.0,1.25,1.5,2.0 (source: +env)` (`ghi` run1); `= 1.5` (`gfix150`); `= 1.0` (`g100`). Liveness OK 7/7 for +every arm. The change-gated `[bb]` line (now carries BOTH the applied gain and +the resulting angular shift) shows what each arm actually did: + +| arm | runs w/ `[bb]` | lines | gains applied | shift min/max (deg) | +|---|---:|---:|---|---| +| `g100` | 0/7 | 0 | none (identity) | - | +| `glo` | 7/7 | 247 | 0.25 x105, 0.50 x77, 0.75 x65 | -21.66 / +20.68 | +| `ghi` | 5/7 | 20 | 1.25 x10, 1.50 x6, 2.00 x4 | -23.64 / +8.48 | +| `gfix150` | 7/7 | 1396 | 1.50 (fixed) | -14.36 / +14.58 | +| `gfix125` | 7/7 | 1409 | 1.25 (fixed) | -7.01 / +7.12 | + +Two readings. (a) The learner in `ghi` **did explore above 1.0** (every logged +non-1.0 gain was > 1), but it moved off 1.0 only rarely — the candidate hit +rates are near-tied, so it mostly sat at Pattern. (b) The `glo` learner applied +sub-unity gains constantly and was still neutral at long range (450+ -0.56 pp, +p = 0.35) — sub-unity *fractional* gain is not the same lever as the HeadOn +kill: it shrinks Pattern's lead without removing it. + +### 2.6 Verdict — DIRECT ANSWER + +**[MEASURED] NO — live gains above 1.0 do not beat Pattern.** + +* The decisive arms are the fixed ones: `gfix150` loses **-94 dmg/run + (p = 0.0006, 165 vs 259)** and **-1.71 round wins for control (p = 0.009, + 4/49 vs 16/49)**, and loses the long-range hit rate by 2-4 pp at 2x MDE. + `gfix125` is directionally worse too (-34 dmg/run, p = 0.087; -1.83 pp at + 300-450, p = 0.022). A fixed gain > 1 at 300+ px is **harmful**. +* The hypothesis arm `ghi` (learner allowed above 1) is **directionally + positive but not significant**: +14.3 dmg/run (p = 0.41), +0.86 wins (p = + 0.28), +0.83 pp at 450+ (p = 0.35) — all inside the 7-run MDE. This is NOT + evidence of a win. +* The Phase-1 `[1,1,1,0,0]` table's opposite direction (less lead) was already + refuted live by `140fe25`, and the `glo` arm here confirms the constrained + learner is neutral, not a win. + +**KILL the gain axis — on this evidence, in BOTH directions.** The live gain +sweep is now complete across `gain in {0 (140fe25), 0.25-1.0 (glo), 1.0 (g100), +1.25, 1.5 (fixed), 2.0 (learner)}`: **nothing beats Pattern**, and both extremes +(0 and 1.5) are measurably worse. The offline ruler that ranked these gains is +dead (§2.0 notice). **Do not spend more live runs on the gain of Pattern's +existing lead.** + +**What the campaign should try next [INFERRED].** The gain axis is amplitude; +the law measured in Phase 0 §0.3.3 is that the lever is lead **information** +(correlation 0.165 at 450+), not amplitude. Recommend, in order: + +1. **A better base predictor at long range** (Phase 0 D2/D3): raise the lead + correlation with longer / multi-length pattern keys, per-distance tables, or + k-NN over movement signatures. Offline may *veto* a broken arm; only a live + A/B may select one. +2. **A causal predictability gate** (Phase 0 D4): fall back to a low-variance + aim only when the match quality is provably poor — attacks the same failure + mode as the HeadOn idea without the live catastrophe HeadOn demonstrated. +3. **Long-range power policy** (Phase 0 D5) — a shorter horizon is a different, + already-partly-live lever on the same long-range hit rate. + +Because 7 runs can only see effects larger than ~16% of the mean, a next step +should be chosen for **plausible large effect**, not for a sub-MDE gain slope. + +**Reproduce:** `tools/ab/ab_run.sh --arms tools/ab/arms_leadgain.txt --runs 7 +--outdir /tmp/ab_leadgain --conc 8 --rounds 7`; +`python3 tools/ab/ab_analyze.py /tmp/ab_leadgain --reference control`; +`python3 tools/ab/ab_range_bands.py /tmp/ab_leadgain --reference control`. +Liveness from `/run.bot.stdout.log` (`[env]`), applied gain from the +same file (`[bb]`). Fixtures: `/tmp/ab_leadgain` (ephemeral, as is the corpus).