BitBrain campaign phase 1: the missing gain sweep and BitBrain as a lead-gain corrector

Task A — the gain region Phase 0 never covered (gain < 1). Extend the
prediction-quality ruler with gain 0.25/0.50/0.75 arms and a fixed causal
per-band arm. Full 70-run result: the hitProxy-argmax curve is
[1.00, 1.00, 1.00, 0.00, 0.00] — Pattern below 300 px, HeadOn above —
worth +0.13 pp at 300-450 and +2.16 pp at 450+ (0.0767 -> 0.0984). Lead
correlation is identical for every g>0 (Pearson is scale-invariant), so a
shrinking gain adds no lead information; and the least-squares optimum
[1,1,1,1,0.25,0.25] diverges from the hitProxy optimum because Pattern's
lead errors are bimodal.

Task B — rebuild guns/bitbrain_gun.nim as a lead-gain corrector:
aim = LOS + gain*(patternAim - LOS), gain learned online per range band by
ranking candidate gains on the hit-probability proxy (the observed lead
label via tmhObservedAt), gated to range >= 300 px. ADE+SBC output removed.
Offline (70 runs): Pattern below 300 px, +0.37 pp at 300-450, +1.90 pp at
450+ (hitProxy 0.0957 vs 0.0767), matching the fixed rule to within 0.27 pp
at 450+ and exceeding it at 300-450. ~0.0007 ms/tick marginal (old gun
~0.114 ms/tick). Default off; rack membership, env report and guard tests
(bitbrain 32, registration 13, rack 48, tm_pattern 20, env_report) unchanged
and green.

Ledger: docs/bitbrain_campaign.md Phase 1, including the three on-file
negatives and the causal-shippability note. No live claim.
This commit is contained in:
2026-09-25 00:10:19 +02:00
parent 140fe2519a
commit c305ef4212
4 changed files with 553 additions and 313 deletions
+164 -1
View File
@@ -253,6 +253,169 @@ Every D-item must end in a live A/B before any phase verdict.
---
## Phase 1 — *(unclaimed; append below)*
## Phase 1: the missing gain sweep *(owner: overnight job j100, committed)*
### Three negatives are on file (the morning reader must see these)
1. **BitBrain as previously shipped was statistically identical to Pattern
live.** 30 runs/arm, MDE 24.4 dmg/run: `docs/bitbrain_gun_verdict.md`,
commit `d93ce44`.
2. **Gun-mixing does not disrupt DrussGT.** TMHorizon+BitBrain mixed into the
rack produced no dodge disruption: `docs/gun_mix_disruption.md`, commit
`32a5e72`.
3. **BitBrain added no measurable aim information offline.** 450+: 16.200 deg
vs Pattern 16.193, hitProxy 0.0767 vs 0.0767 (`a82c864`, §0.3.5 above).
These bound the plausible upside: the previous BitBrain output — an ADDITIVE
angular shift — was information-free, so Phase 1 changes the output shape, not
the learning rate.
### 1.1 The missing sweep: Pattern x gain in [0.00, 1.00] **[MEASURED]**
`common_libs/tests/run_prediction_quality.nim` now carries sub-unity gain arms
(gain 0.0 is HeadOn, gain 1.0 is Pattern). Full 70-run sweep, 14 arms,
3 598 428 tick-bins, `common_libs/tests/prediction_quality_results.txt`.
Cells are `mean|err| deg [hitProxy]`; **hitProxy is the objective** (fraction of
tick-bins within `atan(18/range)`, the ruler's proxy for hit probability).
| band | |req| deg | g=0.00 | g=0.25 | g=0.50 | g=0.75 | g=1.00 | best hpx | dHpx | best |err| |
|---|---|---|---|---|---|---|---|---|
| 0–100 | 19.62 | 19.62 [.342] | 16.98 [.391] | 14.58 [.469] | 12.36 [.640] | **10.56 [.699]** | 1.00 | +.0000 | 1.00 |
| 100–200 | 19.98 | 19.98 [.172] | 17.79 [.166] | 16.10 [.167] | 14.97 [.234] | **14.75 [.342]** | 1.00 | +.0000 | 1.00 |
| 200–300 | 17.34 | 17.34 [.133] | 15.95 [.128] | 15.32 [.137] | 15.53 [.139] | **16.61 [.185]** | 1.00 | +.0000 | 0.50 |
| 300–450 | 14.61 | **14.61 [.105]** | 14.13 [.102] | 14.49 [.100] | 15.64 [.095] | 17.53 [.104] | 0.00 | +.0013 | 0.25 |
| 450+ | 12.33 | **12.33 [.098]** | 12.26 [.093] | 12.94 [.088] | 14.28 [.081] | 16.19 [.077] | 0.00 | +.0216 | 0.25 |
**The optimal gain curve (hitProxy-argmax per band) is**
**`[1.00, 1.00, 1.00, 0.00, 0.00]`** — use Pattern's full lead below 300 px and
**aim at the enemy's current position (zero lead, HeadOn) at 300+ px**. The
implied hit-probability gains vs Pattern are +0.00 pp (0–300), +0.13 pp
(300–450) and **+2.16 pp at 450+ (0.0767 -> 0.0984, a 28 % relative rise)**.
Weighted by tick-bin share (450+ is 64.2 %, 300–450 is 31.1 %) the whole-corpus
proxy rises ~+1.4 pp.
**[CAUSAL-SHIPPABILITY, important]** The rule only needs the *range*, which is
known at fire time, so applying `[1,1,1,0,0]` is causal and needs **no learning
at all** — a per-band table is shippable as a constant, exactly like the
Pattern radial-offset knob. What is *not* causal is the **estimation** of the
table from the same runs (it is in-sample here); a shipped table would be fitted
offline on past battles or learned online, which is what BitBrain does. The
table's value is robust to that caveat because the winning entries are the two
extremes (full Pattern lead, or zero lead), not a knife-edge intermediate value.
### 1.2 Shrinking the gain does NOT add lead information **[MEASURED]**
| band | corr @ g=0.25 | g=0.50 | g=0.75 | g=1.00 |
|---|---|---|---|---|
| 0–100 | 0.774 | 0.774 | 0.774 | 0.774 |
| 100–200 | 0.612 | 0.612 | 0.612 | 0.612 |
| 200–300 | 0.457 | 0.457 | 0.457 | 0.457 |
| 300–450 | 0.266 | 0.266 | 0.266 | 0.266 |
| 450+ | 0.165 | 0.165 | 0.165 | 0.165 |
Every `g > 0` column is **identical**: Pearson correlation is invariant under
positive scaling. A fractional gain therefore buys nothing on the
lead-information axis — it only shrinks the magnitude of an uninformative signal
toward the low-variance static aim. This confirms §0.3.3's reading and is the
mechanism behind the whole curve.
**The MSE-optimal gain and the hitProxy-optimal gain diverge [MEASURED, key].**
The `best |err|` column is the least-squares optimum `[1.00, 1.00, 0.50, 0.25,
0.25]`. At 200–300, MSE wants ~0.5 but hitProxy wants 1.0: Pattern's lead errors
are **bimodal** (it either nails the lead or is far off), so shrinking every
sample trades many small-within-tolerance hits for a smaller tail. Any *learned*
corrector that minimises squared error will therefore under-perform at mid range
— a fact this phase measured the hard way (an MSE-gain BitBrain dropped the
200–300 proxy from 0.190 to 0.131).
### 1.3 BitBrain rebuilt as a lead-gain corrector **[MEASURED]**
`common_libs/guns/bitbrain_gun.nim` is rewritten. The ADE+SBC network is gone
from the gun; the output is now a multiplicative gain on Pattern's lead,
`aim = LOS + gain * (patternAim - LOS)`, with `gain` chosen from a fixed
candidate set `{0, 0.25, 0.5, 0.75, 1.0}`.
* **Label path** (unchanged): at fire time we remember Pattern's lead and the
target's angular half-width `atan(18/range)`; `h = round(dist/speed)` ticks
later `tmhObservedAt` returns the enemy's observed bearing from the firing
position, giving `requiredLead = observedBearing - LOS`.
* **Learning rule** (changed): for every resolved sample we score *each*
candidate by whether `|gain*baseLead - requiredLead| <= atan(18/range)` and keep
the hit counts per range band; the band's gain is the **argmax hit rate** — the
hit-probability proxy itself, not squared error. This directly fixes the
bimodality failure above.
* **State / gate**: the state is the range band (causally known). The correction
is applied only for range >= 300 px (`BB_GAIN_BAND_MIN = 3`), below which
Pattern's lead is informative.
* **Stats are battle-scale**: a round boundary keeps the counts; a new battle or
target change (`resetLearning`/`targetChanged`) wipes them.
Offline 70-run result (same ruler, same run as §1.1):
| band | Pattern hpx | fixed rule [1,1,1,0,0] | BitBrain hpx | fixed−Pat | BB−Pat |
|---|---|---|---|---|---|
| 0–100 | 0.6993 | 0.6993 | 0.6993 | +0.0000 | +0.0000 |
| 100–200 | 0.3418 | 0.3418 | 0.3418 | +0.0000 | +0.0000 |
| 200–300 | 0.1850 | 0.1850 | 0.1850 | +0.0000 | +0.0000 |
| 300–450 | 0.1036 | 0.1049 | **0.1073** | +0.0013 | **+0.0037** |
| 450+ | 0.0767 | **0.0984** | 0.0957 | +0.0216 | **+0.0190** |
BitBrain's effective point estimates match the fixed table to within 0.27 pp at
450+ and actually exceed it at 300–450. The internal log shows why it is not
exactly equal: at long range the candidate hit rates are near-tied, so the
argmax flips between 0.00 and 0.25 (450+) / 0.00 and 0.75 (300–450); the
aggregate still lands on the right side. A fixed table is more stable; the
learned version needs no table and adapts per battle.
**Cost [MEASURED].** `/tmp/bench_bb` (200k ticks, 4 power bins) — Pattern
0.0224 ms/tick, BitBrain 0.0230 ms/tick, i.e. a **marginal ~0.0007 ms/tick**
(0.7 us/tick), versus the ~0.114 ms/tick of the old ADE+SBC gun and the
13.16 ms/tick budget. The whole 14-arm offline sweep now runs in 271 s
(0.0754 ms per tick-bin vs Phase 0's 0.1131 for 10 arms).
**Guard tests stay green [MEASURED].** `test_bitbrain` 32 checks,
`test_bitbrain_registration` 13, `test_rack_membership` 48,
`test_tm_pattern_registration` 20, `test_env_report` all pass. `env_report.nim`
is unchanged: the legacy BitBrain knobs are still resolved and reported. The
ruler still validates (recorded hits 1.360 deg vs misses 16.597 deg,
separation 12.20x deg / 13.34x px; oracle 0.000000 deg).
### 1.4 Verdict
**[MEASURED] Yes — a per-band lead-gain rule beats Pattern offline**, but only in
the two long bands: `[1,1,1,0,0]` gives +0.13 pp at 300–450 and **+2.16 pp at
450+** (a ~28 % relative rise in the hit-proxy, against Phase 0's realistic
causal ceiling of HeadOn's ~0.10, which it reaches). Below 300 px Pattern is
already optimal and the rule is a no-op.
**[MEASURED] Yes — BitBrain-as-gain-corrector matches that rule** (+1.90 pp at
450+, +0.37 pp at 300–450 over Pattern; within 0.27 pp of the fixed table at
450+, above it at 300–450), while being **causally learnable online** and
~163x cheaper per tick than the old gun.
**[INFERRED / honest caveat] BitBrain is not required to capture the gain** — a
constant `[1,1,1,0,0]` table would do it. BitBrain's value is that it discovers
the table per battle without one; its cost is cold-start (it begins at gain 1.0,
explaining the 0.27 pp gap at 450+) and the near-tie jitter in its point
estimates. So the honest ship decision is: **the per-band rule is the thing to
test live, and it can be shipped either as a constant table or as this online
learner** — the learner is redundant if a table is acceptable, and preferable
only if the optimum is expected to drift per enemy. The highest-value live
experiment is unchanged and stronger: a **range-gated Pattern<->HeadOn switch at
~300 px vs pure Pattern** (Phase 0's D1), left-running. **No live win is claimed
here** — the live gate is a separate phase.
### 1.5 Designs after Phase 1
| # | design | status |
|---|---|---|
| D1 | Live A/B: Pattern<->HeadOn switch at ~300 px vs Pattern | **TODO (highest value, live)** — offline now says +2.16 pp at 450+ |
| D2 | Raise Pattern lead *correlation* at long range (longer/multi-length keys, per-distance tables, k-NN) | TODO (offline-searchable) |
| D3 | Match-quality-conditional gain (keep full lead when the pattern match is good, shrink when bad) — attacks the bimodality that makes the MSE gain diverge | TODO (offline-searchable) |
| D4 | A causal predictability gate falling back to HeadOn when the future is unpredictable | TODO |
| D5 | Long-range power policy (already partly live) | TODO (offline proxy only) |
| D6 | Lead-gain sweep gains >= 1.0 | **DEAD — measured** (§0.3.2) |
| D7 | Constant sub-unity lead gain at all ranges | **DEAD — measured** (§1.1: gain < 1 hurts below 300) |
## Phase 2 — *(unclaimed; append below)*