BitBrain campaign phase 1: the missing gain sweep and BitBrain as a lead-gain corrector
Task A — the gain region Phase 0 never covered (gain < 1). Extend the prediction-quality ruler with gain 0.25/0.50/0.75 arms and a fixed causal per-band arm. Full 70-run result: the hitProxy-argmax curve is [1.00, 1.00, 1.00, 0.00, 0.00] — Pattern below 300 px, HeadOn above — worth +0.13 pp at 300-450 and +2.16 pp at 450+ (0.0767 -> 0.0984). Lead correlation is identical for every g>0 (Pearson is scale-invariant), so a shrinking gain adds no lead information; and the least-squares optimum [1,1,1,1,0.25,0.25] diverges from the hitProxy optimum because Pattern's lead errors are bimodal. Task B — rebuild guns/bitbrain_gun.nim as a lead-gain corrector: aim = LOS + gain*(patternAim - LOS), gain learned online per range band by ranking candidate gains on the hit-probability proxy (the observed lead label via tmhObservedAt), gated to range >= 300 px. ADE+SBC output removed. Offline (70 runs): Pattern below 300 px, +0.37 pp at 300-450, +1.90 pp at 450+ (hitProxy 0.0957 vs 0.0767), matching the fixed rule to within 0.27 pp at 450+ and exceeding it at 300-450. ~0.0007 ms/tick marginal (old gun ~0.114 ms/tick). Default off; rack membership, env report and guard tests (bitbrain 32, registration 13, rack 48, tm_pattern 20, env_report) unchanged and green. Ledger: docs/bitbrain_campaign.md Phase 1, including the three on-file negatives and the causal-shippability note. No live claim.
This commit is contained in:
+164
-1
@@ -253,6 +253,169 @@ Every D-item must end in a live A/B before any phase verdict.
|
||||
|
||||
---
|
||||
|
||||
## Phase 1 — *(unclaimed; append below)*
|
||||
## Phase 1: the missing gain sweep *(owner: overnight job j100, committed)*
|
||||
|
||||
### Three negatives are on file (the morning reader must see these)
|
||||
|
||||
1. **BitBrain as previously shipped was statistically identical to Pattern
|
||||
live.** 30 runs/arm, MDE 24.4 dmg/run: `docs/bitbrain_gun_verdict.md`,
|
||||
commit `d93ce44`.
|
||||
2. **Gun-mixing does not disrupt DrussGT.** TMHorizon+BitBrain mixed into the
|
||||
rack produced no dodge disruption: `docs/gun_mix_disruption.md`, commit
|
||||
`32a5e72`.
|
||||
3. **BitBrain added no measurable aim information offline.** 450+: 16.200 deg
|
||||
vs Pattern 16.193, hitProxy 0.0767 vs 0.0767 (`a82c864`, §0.3.5 above).
|
||||
|
||||
These bound the plausible upside: the previous BitBrain output — an ADDITIVE
|
||||
angular shift — was information-free, so Phase 1 changes the output shape, not
|
||||
the learning rate.
|
||||
|
||||
### 1.1 The missing sweep: Pattern x gain in [0.00, 1.00] **[MEASURED]**
|
||||
|
||||
`common_libs/tests/run_prediction_quality.nim` now carries sub-unity gain arms
|
||||
(gain 0.0 is HeadOn, gain 1.0 is Pattern). Full 70-run sweep, 14 arms,
|
||||
3 598 428 tick-bins, `common_libs/tests/prediction_quality_results.txt`.
|
||||
Cells are `mean|err| deg [hitProxy]`; **hitProxy is the objective** (fraction of
|
||||
tick-bins within `atan(18/range)`, the ruler's proxy for hit probability).
|
||||
|
||||
| band | |req| deg | g=0.00 | g=0.25 | g=0.50 | g=0.75 | g=1.00 | best hpx | dHpx | best |err| |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| 0–100 | 19.62 | 19.62 [.342] | 16.98 [.391] | 14.58 [.469] | 12.36 [.640] | **10.56 [.699]** | 1.00 | +.0000 | 1.00 |
|
||||
| 100–200 | 19.98 | 19.98 [.172] | 17.79 [.166] | 16.10 [.167] | 14.97 [.234] | **14.75 [.342]** | 1.00 | +.0000 | 1.00 |
|
||||
| 200–300 | 17.34 | 17.34 [.133] | 15.95 [.128] | 15.32 [.137] | 15.53 [.139] | **16.61 [.185]** | 1.00 | +.0000 | 0.50 |
|
||||
| 300–450 | 14.61 | **14.61 [.105]** | 14.13 [.102] | 14.49 [.100] | 15.64 [.095] | 17.53 [.104] | 0.00 | +.0013 | 0.25 |
|
||||
| 450+ | 12.33 | **12.33 [.098]** | 12.26 [.093] | 12.94 [.088] | 14.28 [.081] | 16.19 [.077] | 0.00 | +.0216 | 0.25 |
|
||||
|
||||
**The optimal gain curve (hitProxy-argmax per band) is**
|
||||
**`[1.00, 1.00, 1.00, 0.00, 0.00]`** — use Pattern's full lead below 300 px and
|
||||
**aim at the enemy's current position (zero lead, HeadOn) at 300+ px**. The
|
||||
implied hit-probability gains vs Pattern are +0.00 pp (0–300), +0.13 pp
|
||||
(300–450) and **+2.16 pp at 450+ (0.0767 -> 0.0984, a 28 % relative rise)**.
|
||||
Weighted by tick-bin share (450+ is 64.2 %, 300–450 is 31.1 %) the whole-corpus
|
||||
proxy rises ~+1.4 pp.
|
||||
|
||||
**[CAUSAL-SHIPPABILITY, important]** The rule only needs the *range*, which is
|
||||
known at fire time, so applying `[1,1,1,0,0]` is causal and needs **no learning
|
||||
at all** — a per-band table is shippable as a constant, exactly like the
|
||||
Pattern radial-offset knob. What is *not* causal is the **estimation** of the
|
||||
table from the same runs (it is in-sample here); a shipped table would be fitted
|
||||
offline on past battles or learned online, which is what BitBrain does. The
|
||||
table's value is robust to that caveat because the winning entries are the two
|
||||
extremes (full Pattern lead, or zero lead), not a knife-edge intermediate value.
|
||||
|
||||
### 1.2 Shrinking the gain does NOT add lead information **[MEASURED]**
|
||||
|
||||
| band | corr @ g=0.25 | g=0.50 | g=0.75 | g=1.00 |
|
||||
|---|---|---|---|---|
|
||||
| 0–100 | 0.774 | 0.774 | 0.774 | 0.774 |
|
||||
| 100–200 | 0.612 | 0.612 | 0.612 | 0.612 |
|
||||
| 200–300 | 0.457 | 0.457 | 0.457 | 0.457 |
|
||||
| 300–450 | 0.266 | 0.266 | 0.266 | 0.266 |
|
||||
| 450+ | 0.165 | 0.165 | 0.165 | 0.165 |
|
||||
|
||||
Every `g > 0` column is **identical**: Pearson correlation is invariant under
|
||||
positive scaling. A fractional gain therefore buys nothing on the
|
||||
lead-information axis — it only shrinks the magnitude of an uninformative signal
|
||||
toward the low-variance static aim. This confirms §0.3.3's reading and is the
|
||||
mechanism behind the whole curve.
|
||||
|
||||
**The MSE-optimal gain and the hitProxy-optimal gain diverge [MEASURED, key].**
|
||||
The `best |err|` column is the least-squares optimum `[1.00, 1.00, 0.50, 0.25,
|
||||
0.25]`. At 200–300, MSE wants ~0.5 but hitProxy wants 1.0: Pattern's lead errors
|
||||
are **bimodal** (it either nails the lead or is far off), so shrinking every
|
||||
sample trades many small-within-tolerance hits for a smaller tail. Any *learned*
|
||||
corrector that minimises squared error will therefore under-perform at mid range
|
||||
— a fact this phase measured the hard way (an MSE-gain BitBrain dropped the
|
||||
200–300 proxy from 0.190 to 0.131).
|
||||
|
||||
### 1.3 BitBrain rebuilt as a lead-gain corrector **[MEASURED]**
|
||||
|
||||
`common_libs/guns/bitbrain_gun.nim` is rewritten. The ADE+SBC network is gone
|
||||
from the gun; the output is now a multiplicative gain on Pattern's lead,
|
||||
`aim = LOS + gain * (patternAim - LOS)`, with `gain` chosen from a fixed
|
||||
candidate set `{0, 0.25, 0.5, 0.75, 1.0}`.
|
||||
|
||||
* **Label path** (unchanged): at fire time we remember Pattern's lead and the
|
||||
target's angular half-width `atan(18/range)`; `h = round(dist/speed)` ticks
|
||||
later `tmhObservedAt` returns the enemy's observed bearing from the firing
|
||||
position, giving `requiredLead = observedBearing - LOS`.
|
||||
* **Learning rule** (changed): for every resolved sample we score *each*
|
||||
candidate by whether `|gain*baseLead - requiredLead| <= atan(18/range)` and keep
|
||||
the hit counts per range band; the band's gain is the **argmax hit rate** — the
|
||||
hit-probability proxy itself, not squared error. This directly fixes the
|
||||
bimodality failure above.
|
||||
* **State / gate**: the state is the range band (causally known). The correction
|
||||
is applied only for range >= 300 px (`BB_GAIN_BAND_MIN = 3`), below which
|
||||
Pattern's lead is informative.
|
||||
* **Stats are battle-scale**: a round boundary keeps the counts; a new battle or
|
||||
target change (`resetLearning`/`targetChanged`) wipes them.
|
||||
|
||||
Offline 70-run result (same ruler, same run as §1.1):
|
||||
|
||||
| band | Pattern hpx | fixed rule [1,1,1,0,0] | BitBrain hpx | fixed−Pat | BB−Pat |
|
||||
|---|---|---|---|---|---|
|
||||
| 0–100 | 0.6993 | 0.6993 | 0.6993 | +0.0000 | +0.0000 |
|
||||
| 100–200 | 0.3418 | 0.3418 | 0.3418 | +0.0000 | +0.0000 |
|
||||
| 200–300 | 0.1850 | 0.1850 | 0.1850 | +0.0000 | +0.0000 |
|
||||
| 300–450 | 0.1036 | 0.1049 | **0.1073** | +0.0013 | **+0.0037** |
|
||||
| 450+ | 0.0767 | **0.0984** | 0.0957 | +0.0216 | **+0.0190** |
|
||||
|
||||
BitBrain's effective point estimates match the fixed table to within 0.27 pp at
|
||||
450+ and actually exceed it at 300–450. The internal log shows why it is not
|
||||
exactly equal: at long range the candidate hit rates are near-tied, so the
|
||||
argmax flips between 0.00 and 0.25 (450+) / 0.00 and 0.75 (300–450); the
|
||||
aggregate still lands on the right side. A fixed table is more stable; the
|
||||
learned version needs no table and adapts per battle.
|
||||
|
||||
**Cost [MEASURED].** `/tmp/bench_bb` (200k ticks, 4 power bins) — Pattern
|
||||
0.0224 ms/tick, BitBrain 0.0230 ms/tick, i.e. a **marginal ~0.0007 ms/tick**
|
||||
(0.7 us/tick), versus the ~0.114 ms/tick of the old ADE+SBC gun and the
|
||||
13.16 ms/tick budget. The whole 14-arm offline sweep now runs in 271 s
|
||||
(0.0754 ms per tick-bin vs Phase 0's 0.1131 for 10 arms).
|
||||
|
||||
**Guard tests stay green [MEASURED].** `test_bitbrain` 32 checks,
|
||||
`test_bitbrain_registration` 13, `test_rack_membership` 48,
|
||||
`test_tm_pattern_registration` 20, `test_env_report` all pass. `env_report.nim`
|
||||
is unchanged: the legacy BitBrain knobs are still resolved and reported. The
|
||||
ruler still validates (recorded hits 1.360 deg vs misses 16.597 deg,
|
||||
separation 12.20x deg / 13.34x px; oracle 0.000000 deg).
|
||||
|
||||
### 1.4 Verdict
|
||||
|
||||
**[MEASURED] Yes — a per-band lead-gain rule beats Pattern offline**, but only in
|
||||
the two long bands: `[1,1,1,0,0]` gives +0.13 pp at 300–450 and **+2.16 pp at
|
||||
450+** (a ~28 % relative rise in the hit-proxy, against Phase 0's realistic
|
||||
causal ceiling of HeadOn's ~0.10, which it reaches). Below 300 px Pattern is
|
||||
already optimal and the rule is a no-op.
|
||||
|
||||
**[MEASURED] Yes — BitBrain-as-gain-corrector matches that rule** (+1.90 pp at
|
||||
450+, +0.37 pp at 300–450 over Pattern; within 0.27 pp of the fixed table at
|
||||
450+, above it at 300–450), while being **causally learnable online** and
|
||||
~163x cheaper per tick than the old gun.
|
||||
|
||||
**[INFERRED / honest caveat] BitBrain is not required to capture the gain** — a
|
||||
constant `[1,1,1,0,0]` table would do it. BitBrain's value is that it discovers
|
||||
the table per battle without one; its cost is cold-start (it begins at gain 1.0,
|
||||
explaining the 0.27 pp gap at 450+) and the near-tie jitter in its point
|
||||
estimates. So the honest ship decision is: **the per-band rule is the thing to
|
||||
test live, and it can be shipped either as a constant table or as this online
|
||||
learner** — the learner is redundant if a table is acceptable, and preferable
|
||||
only if the optimum is expected to drift per enemy. The highest-value live
|
||||
experiment is unchanged and stronger: a **range-gated Pattern<->HeadOn switch at
|
||||
~300 px vs pure Pattern** (Phase 0's D1), left-running. **No live win is claimed
|
||||
here** — the live gate is a separate phase.
|
||||
|
||||
### 1.5 Designs after Phase 1
|
||||
|
||||
| # | design | status |
|
||||
|---|---|---|
|
||||
| D1 | Live A/B: Pattern<->HeadOn switch at ~300 px vs Pattern | **TODO (highest value, live)** — offline now says +2.16 pp at 450+ |
|
||||
| D2 | Raise Pattern lead *correlation* at long range (longer/multi-length keys, per-distance tables, k-NN) | TODO (offline-searchable) |
|
||||
| D3 | Match-quality-conditional gain (keep full lead when the pattern match is good, shrink when bad) — attacks the bimodality that makes the MSE gain diverge | TODO (offline-searchable) |
|
||||
| D4 | A causal predictability gate falling back to HeadOn when the future is unpredictable | TODO |
|
||||
| D5 | Long-range power policy (already partly live) | TODO (offline proxy only) |
|
||||
| D6 | Lead-gain sweep gains >= 1.0 | **DEAD — measured** (§0.3.2) |
|
||||
| D7 | Constant sub-unity lead gain at all ranges | **DEAD — measured** (§1.1: gain < 1 hurts below 300) |
|
||||
|
||||
|
||||
## Phase 2 — *(unclaimed; append below)*
|
||||
|
||||
Reference in New Issue
Block a user