Campaign phase 2 ledger: live lead-gain sweep (gains above 1.0 do NOT beat Pattern)
6 arms x 7 runs x 7 rounds vs real DrussGT on one frozen binary (2747ebd). Validity check PASSES: g100 (fixed gain 1.0) is statistically indistinguishable from control (dmg p=0.65, wins p=0.62, ALL hit rate +0.04pp p=0.92; zero [bb] lines = provably no correction). Fixed gains above 1.0 LOSE at 300+: gfix150 -94 dmg/run (p=0.0006), 4/49 vs 16/49 wins (p=0.009), -3.85pp at 300-450 (p=0.0006) and -2.25pp at 450+ (p=0.0023). The hypothesis arm ghi (learner allowed above 1) is directionally positive but inside the MDE (+14.3 dmg/run p=0.41; +0.83pp at 450+ p=0.35). Kill the gain axis in both directions. Includes the mandatory correction notice: Phase 1's [1,1,1,0,0] is LIVE-REFUTED by140fe25, and the standing rule that offline is veto-only / live decides.
This commit is contained in:
+208
-1
@@ -418,4 +418,211 @@ here** — the live gate is a separate phase.
|
||||
| D7 | Constant sub-unity lead gain at all ranges | **DEAD — measured** (§1.1: gain < 1 hurts below 300) |
|
||||
|
||||
|
||||
## Phase 2 — *(unclaimed; append below)*
|
||||
## Phase 2: live lead gains — is a gain ABOVE 1.0 better? *(owner: job j101, committed)*
|
||||
|
||||
> **CORRECTION NOTICE — READ FIRST. Phase 1's offline per-band gain table
|
||||
> `[1,1,1,0,0]` is LIVE-REFUTED by `140fe25` (`docs/headon_longrange_live.md`).
|
||||
> Do not act on it.** HeadOn — which *is* gain 0 at every range — scored
|
||||
> **14 dmg/run vs Pattern's 279**, won **0 of 105 rounds**, and hit **0.4% vs
|
||||
> Pattern's 9.2%** at 450+. Zero lead above 300 px is a live catastrophe, not a
|
||||
> +2.16 pp improvement. Phase 1's §1.4 claim ("a per-band lead-gain rule beats
|
||||
> Pattern offline") is demoted to an offline-only observation that live killed.
|
||||
>
|
||||
> **New standing rule: offline is VETO-ONLY** (`docs/offline_harness_trust.md`,
|
||||
> `e40c849`). It may reject a clearly broken design; it may **never select a
|
||||
> winner**. Every number in this phase is LIVE. No offline number is cited here
|
||||
> as evidence of a live win.
|
||||
|
||||
### 2.0 The hypothesis (a hypothesis, not a fact)
|
||||
|
||||
The offline ruler has now been wrong **twice, both times preferring LESS lead**
|
||||
than reality: Phase 0 said the static gun beats Pattern at long range, and
|
||||
Phase 1 said gain 0 above 300 px. If the ruler systematically under-values lead,
|
||||
then it will also have **under-rated gains ABOVE 1.0** — and those had never
|
||||
been tested live. The hypothesis of this phase is therefore: *Pattern's full
|
||||
lead is under-shot at long range live, so scaling it up (`gain > 1`) beats
|
||||
Pattern.* **[INFERRED]** — a reasoned guess from two ruler failures, not a
|
||||
measurement.
|
||||
|
||||
### 2.1 Method — every arm is pure env on ONE frozen binary
|
||||
|
||||
Task A added `TR_BITBRAIN_GAINS` (comma-separated candidate set, commit
|
||||
`2747ebd`): unset reproduces the shipped candidate set `[0,.25,.5,.75,1.0]`
|
||||
exactly, and **exactly ONE value is a FIXED gain with no learning**. BitBrain's
|
||||
base prediction *is* Pattern (`tmh.pattern.predict`), and the correction is a
|
||||
multiplicative gain on Pattern's lead over the line of sight, applied only at
|
||||
range >= 300 px (`BB_GAIN_BAND_MIN`) — the long bands, where the hypothesis
|
||||
lives. So every arm swaps the admitted rack gun (Pattern off, BitBrain on) and
|
||||
changes only the lead gain.
|
||||
|
||||
**Session [MEASURED]:** `tools/ab/ab_run.sh --arms tools/ab/arms_leadgain.txt
|
||||
--runs 7 --outdir /tmp/ab_leadgain --conc 8 --rounds 7`; **commit
|
||||
`2747ebd`**, frozen binary sha256 `3aa2da14…`; 6 arms x 7 runs x 7 rounds = 42
|
||||
battles, 294 rounds, real DrussGT, **42 ok / 0 failed**. Liveness **OK 7/7 runs
|
||||
for every arm** (boot report shows the arm env verbatim).
|
||||
|
||||
| arm | env (beyond `TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both`) | what it tests |
|
||||
|---|---|---|
|
||||
| `control` | (none; shipped Pattern-only rack) | reference |
|
||||
| `g100` | `TR_BITBRAIN_GAINS=1.0` | **validity check**: fixed gain 1.0 == identity |
|
||||
| `glo` | `TR_BITBRAIN_GAINS=0.25,0.5,0.75,1.0` | learner restricted to <= 1 |
|
||||
| `ghi` | `TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0` | learner allowed ABOVE 1 (**hypothesis**) |
|
||||
| `gfix150` | `TR_BITBRAIN_GAINS=1.5` | fixed 1.5, no learning |
|
||||
| `gfix125` | `TR_BITBRAIN_GAINS=1.25` | fixed 1.25, no learning |
|
||||
|
||||
### 2.2 The `g100` VALIDITY CHECK — the plumbing is sound [MEASURED]
|
||||
|
||||
Gain 1.0 is the identity, so `g100` **must** be statistically indistinguishable
|
||||
from `control`. It is:
|
||||
|
||||
| metric | control | g100 | diff | perm p (exact 7v7) | MDE |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| dmg/run | 259 | 267 | **-8.3** | **0.6492** | 41.4 |
|
||||
| round wins | 16/49 | 19/49 | **-0.43** | **0.6247** | 1.42 |
|
||||
| hit rate ALL | 9.9% | 10.0% | **+0.04 pp** | **0.9225** | 1.16 |
|
||||
| hit rate 450+ | 8.5% | 8.1% | **-0.47 pp** | **0.4656** | 2.01 |
|
||||
|
||||
The `g100` bot emitted **ZERO `[bb]` lines in 7/7 runs** (`bbLog` is only
|
||||
reached when a non-1.0 gain is applied), i.e. it provably applied no
|
||||
correction at all — a built-in placebo. **The validity check PASSES: the other
|
||||
arms are interpretable.**
|
||||
|
||||
### 2.3 Primary result — damage/run and round wins (live decides) [MEASURED]
|
||||
|
||||
| arm | runs | dmg/run | dmgtk/run | round wins | win% | shots/run |
|
||||
|---|---:|---:|---:|---:|---:|---:|
|
||||
| `control` | 7 | **259** | 211 | **16/49** | 32.7 | 780 |
|
||||
| `g100` | 7 | 267 | 200 | 19/49 | 38.8 | 764 |
|
||||
| `glo` | 7 | 261 | 215 | 19/49 | 38.8 | 783 |
|
||||
| `ghi` | 7 | **273** | 202 | **22/49** | **44.9** | 796 |
|
||||
| `gfix150` | 7 | **165** | 235 | **4/49** | **8.2** | 726 |
|
||||
| `gfix125` | 7 | 225 | 213 | 14/49 | 28.6 | 753 |
|
||||
|
||||
Per-run damage (never just the mean):
|
||||
`control` 290 281 248 247 222 290 236 · `g100` 265 307 236 228 234 329 273 ·
|
||||
`glo` 215 282 260 282 288 221 282 · `ghi` 271 243 215 297 322 283 282 ·
|
||||
`gfix150` 169 134 164 174 189 152 175 · `gfix125` 174 294 209 204 242 198 252.
|
||||
|
||||
vs `control` (exact permutation, per-run):
|
||||
|
||||
| metric | arm | diff (arm - control) | perm p | MW p | MDE |
|
||||
|---|---|---:|---:|---:|---:|
|
||||
| dmg/run | g100 | -8.3 | 0.6492 | 1.0000 | 41.4 |
|
||||
| dmg/run | glo | -2.5 | 0.8671 | 1.0000 | 41.4 |
|
||||
| dmg/run | **ghi** | **+14.3** | **0.4091** | 0.5229 | 41.4 |
|
||||
| dmg/run | **gfix150** | **-93.8** | **0.0006** | 0.0022 | 41.4 |
|
||||
| dmg/run | gfix125 | -34.2 | 0.0874 | 0.1599 | 41.4 |
|
||||
| round wins | g100 | -0.43 | 0.6247 | 0.4769 | 1.42 |
|
||||
| round wins | glo | -0.43 | 0.6329 | 0.5799 | 1.42 |
|
||||
| round wins | ghi | -0.86 | 0.2756 | 0.2097 | 1.42 |
|
||||
| round wins | **gfix150** | **+1.71** | **0.0093** | 0.0079 | 1.42 |
|
||||
| round wins | gfix125 | +0.29 | 0.8042 | 0.8928 | 1.42 |
|
||||
|
||||
**MDE stated plainly: at 7 runs/arm the test only sees large effects** — 41.4
|
||||
dmg/run (16% of the control mean) and 1.42 round wins (62% of 2.3). `ghi`'s
|
||||
+14 dmg/run is a third of the MDE: a live signal smaller than the MDE is NOT a
|
||||
demonstrated effect. `gfix150`'s -94 dmg/run is 2.3x MDE and is decisive.
|
||||
|
||||
### 2.4 HIT RATE BY RANGE BAND — the load-bearing view [MEASURED]
|
||||
|
||||
`tools/ab/ab_range_bands.py /tmp/ab_leadgain --reference control`. Band = range
|
||||
at the fire tick. The claim is range-specific; a whole-battle number is not
|
||||
enough.
|
||||
|
||||
| band (px) | control | g100 | glo | ghi | gfix150 | gfix125 |
|
||||
|---|---:|---:|---:|---:|---:|---:|
|
||||
| 0-100 | 1/2 50.0% | 1/2 50.0% | 2/3 66.7% | 0/0 - | 0/1 0.0% | 2/5 40.0% |
|
||||
| 100-200 | 5/30 16.7% | 3/26 11.5% | 9/33 27.3% | 3/18 16.7% | 3/16 18.8% | 5/19 26.3% |
|
||||
| 200-300 | 19/98 19.4% | 20/110 18.2% | 16/94 17.0% | 10/99 10.1% | 19/100 19.0% | 15/96 15.6% |
|
||||
| **300-450** | **237/2081 11.4%** | 248/2012 12.3% | 237/2018 11.7% | 222/2004 11.1% | **145/1946 7.5%** | **183/1925 9.5%** |
|
||||
| **450+** | **267/3125 8.5%** | 249/3075 8.1% | 256/3211 8.0% | **314/3325 9.4%** | **183/2917 6.3%** | 240/3094 7.8% |
|
||||
| ALL | 529/5336 9.9% | 521/5225 10.0% | 520/5359 9.7% | **549/5446 10.1%** | **350/4980 7.0%** | 445/5139 8.7% |
|
||||
|
||||
Per-band permutation test on per-run band rates (arm - control), 7v7 exact:
|
||||
|
||||
| band | arm | d(pp) | p | MDE(pp) |
|
||||
|---|---|---:|---:|---:|
|
||||
| 300-450 | gfix150 | **-3.85** | **0.0006** | 1.87 |
|
||||
| 450+ | gfix150 | **-2.25** | **0.0023** | 2.01 |
|
||||
| 300-450 | gfix125 | **-1.83** | **0.0221** | 1.87 |
|
||||
| 450+ | gfix125 | -0.81 | 0.1737 | 2.01 |
|
||||
| 450+ | **ghi** | **+0.83** | **0.3473** | 2.01 |
|
||||
| 300-450 | ghi | -0.37 | 0.7191 | 1.87 |
|
||||
| 450+ | glo | -0.56 | 0.3502 | 2.01 |
|
||||
| 300-450 | glo | +0.24 | 0.8071 | 1.87 |
|
||||
| ALL | g100 | +0.04 | 0.9225 | 1.16 |
|
||||
|
||||
**The fixed gains above 1.0 clearly LOSE in the exact bands where they are
|
||||
applied** (300+ px): gfix150 -3.85 pp / -2.25 pp at 2x the MDE, gfix125 -1.83 pp
|
||||
at 300-450. The learner allowed above 1.0 (`ghi`) is the only arm whose 450+ hit
|
||||
rate is above control (+0.83 pp) — but **p = 0.35, well inside the MDE**.
|
||||
|
||||
### 2.5 Applied-gain evidence (the knob really moved the gun) [MEASURED]
|
||||
|
||||
Boot report, verbatim: `[env] TR_BITBRAIN_GAINS = 1.0,1.25,1.5,2.0 (source:
|
||||
env)` (`ghi` run1); `= 1.5` (`gfix150`); `= 1.0` (`g100`). Liveness OK 7/7 for
|
||||
every arm. The change-gated `[bb]` line (now carries BOTH the applied gain and
|
||||
the resulting angular shift) shows what each arm actually did:
|
||||
|
||||
| arm | runs w/ `[bb]` | lines | gains applied | shift min/max (deg) |
|
||||
|---|---:|---:|---|---|
|
||||
| `g100` | 0/7 | 0 | none (identity) | - |
|
||||
| `glo` | 7/7 | 247 | 0.25 x105, 0.50 x77, 0.75 x65 | -21.66 / +20.68 |
|
||||
| `ghi` | 5/7 | 20 | 1.25 x10, 1.50 x6, 2.00 x4 | -23.64 / +8.48 |
|
||||
| `gfix150` | 7/7 | 1396 | 1.50 (fixed) | -14.36 / +14.58 |
|
||||
| `gfix125` | 7/7 | 1409 | 1.25 (fixed) | -7.01 / +7.12 |
|
||||
|
||||
Two readings. (a) The learner in `ghi` **did explore above 1.0** (every logged
|
||||
non-1.0 gain was > 1), but it moved off 1.0 only rarely — the candidate hit
|
||||
rates are near-tied, so it mostly sat at Pattern. (b) The `glo` learner applied
|
||||
sub-unity gains constantly and was still neutral at long range (450+ -0.56 pp,
|
||||
p = 0.35) — sub-unity *fractional* gain is not the same lever as the HeadOn
|
||||
kill: it shrinks Pattern's lead without removing it.
|
||||
|
||||
### 2.6 Verdict — DIRECT ANSWER
|
||||
|
||||
**[MEASURED] NO — live gains above 1.0 do not beat Pattern.**
|
||||
|
||||
* The decisive arms are the fixed ones: `gfix150` loses **-94 dmg/run
|
||||
(p = 0.0006, 165 vs 259)** and **-1.71 round wins for control (p = 0.009,
|
||||
4/49 vs 16/49)**, and loses the long-range hit rate by 2-4 pp at 2x MDE.
|
||||
`gfix125` is directionally worse too (-34 dmg/run, p = 0.087; -1.83 pp at
|
||||
300-450, p = 0.022). A fixed gain > 1 at 300+ px is **harmful**.
|
||||
* The hypothesis arm `ghi` (learner allowed above 1) is **directionally
|
||||
positive but not significant**: +14.3 dmg/run (p = 0.41), +0.86 wins (p =
|
||||
0.28), +0.83 pp at 450+ (p = 0.35) — all inside the 7-run MDE. This is NOT
|
||||
evidence of a win.
|
||||
* The Phase-1 `[1,1,1,0,0]` table's opposite direction (less lead) was already
|
||||
refuted live by `140fe25`, and the `glo` arm here confirms the constrained
|
||||
learner is neutral, not a win.
|
||||
|
||||
**KILL the gain axis — on this evidence, in BOTH directions.** The live gain
|
||||
sweep is now complete across `gain in {0 (140fe25), 0.25-1.0 (glo), 1.0 (g100),
|
||||
1.25, 1.5 (fixed), 2.0 (learner)}`: **nothing beats Pattern**, and both extremes
|
||||
(0 and 1.5) are measurably worse. The offline ruler that ranked these gains is
|
||||
dead (§2.0 notice). **Do not spend more live runs on the gain of Pattern's
|
||||
existing lead.**
|
||||
|
||||
**What the campaign should try next [INFERRED].** The gain axis is amplitude;
|
||||
the law measured in Phase 0 §0.3.3 is that the lever is lead **information**
|
||||
(correlation 0.165 at 450+), not amplitude. Recommend, in order:
|
||||
|
||||
1. **A better base predictor at long range** (Phase 0 D2/D3): raise the lead
|
||||
correlation with longer / multi-length pattern keys, per-distance tables, or
|
||||
k-NN over movement signatures. Offline may *veto* a broken arm; only a live
|
||||
A/B may select one.
|
||||
2. **A causal predictability gate** (Phase 0 D4): fall back to a low-variance
|
||||
aim only when the match quality is provably poor — attacks the same failure
|
||||
mode as the HeadOn idea without the live catastrophe HeadOn demonstrated.
|
||||
3. **Long-range power policy** (Phase 0 D5) — a shorter horizon is a different,
|
||||
already-partly-live lever on the same long-range hit rate.
|
||||
|
||||
Because 7 runs can only see effects larger than ~16% of the mean, a next step
|
||||
should be chosen for **plausible large effect**, not for a sub-MDE gain slope.
|
||||
|
||||
**Reproduce:** `tools/ab/ab_run.sh --arms tools/ab/arms_leadgain.txt --runs 7
|
||||
--outdir /tmp/ab_leadgain --conc 8 --rounds 7`;
|
||||
`python3 tools/ab/ab_analyze.py /tmp/ab_leadgain --reference control`;
|
||||
`python3 tools/ab/ab_range_bands.py /tmp/ab_leadgain --reference control`.
|
||||
Liveness from `<arm>/run<N>.bot.stdout.log` (`[env]`), applied gain from the
|
||||
same file (`[bb]`). Fixtures: `/tmp/ab_leadgain` (ephemeral, as is the corpus).
|
||||
|
||||
Reference in New Issue
Block a user