c305ef4212
Task A — the gain region Phase 0 never covered (gain < 1). Extend the prediction-quality ruler with gain 0.25/0.50/0.75 arms and a fixed causal per-band arm. Full 70-run result: the hitProxy-argmax curve is [1.00, 1.00, 1.00, 0.00, 0.00] — Pattern below 300 px, HeadOn above — worth +0.13 pp at 300-450 and +2.16 pp at 450+ (0.0767 -> 0.0984). Lead correlation is identical for every g>0 (Pearson is scale-invariant), so a shrinking gain adds no lead information; and the least-squares optimum [1,1,1,1,0.25,0.25] diverges from the hitProxy optimum because Pattern's lead errors are bimodal. Task B — rebuild guns/bitbrain_gun.nim as a lead-gain corrector: aim = LOS + gain*(patternAim - LOS), gain learned online per range band by ranking candidate gains on the hit-probability proxy (the observed lead label via tmhObservedAt), gated to range >= 300 px. ADE+SBC output removed. Offline (70 runs): Pattern below 300 px, +0.37 pp at 300-450, +1.90 pp at 450+ (hitProxy 0.0957 vs 0.0767), matching the fixed rule to within 0.27 pp at 450+ and exceeding it at 300-450. ~0.0007 ms/tick marginal (old gun ~0.114 ms/tick). Default off; rack membership, env report and guard tests (bitbrain 32, registration 13, rack 48, tm_pattern 20, env_report) unchanged and green. Ledger: docs/bitbrain_campaign.md Phase 1, including the three on-file negatives and the causal-shippability note. No live claim.
422 lines
24 KiB
Markdown
422 lines
24 KiB
Markdown
# BitBrain campaign ledger
|
||
|
||
**Goal:** make ModularBot's gun **beat Pattern live** against the real DrussGT.
|
||
The user has granted full freedom over the gun ("change input, output, every
|
||
knob of it") and accepts it may fail — the deliverable is that the attempt is
|
||
visible and evidence-backed.
|
||
|
||
**THE FINAL VERDICT IS ALWAYS LIVE.** Everything in this file except the
|
||
`## Phase N` verdict lines is offline, open-loop, on a *fixed recorded enemy
|
||
trajectory*. Per `docs/offline_harness_trust.md` (commit `e40c849`) the offline
|
||
harness is trustworthy for exactly one thing: **per-gun single-tick prediction
|
||
quality on a fixed enemy trajectory** — and it is *never* trustworthy for
|
||
closed-loop questions (movement, range, round length, adaptation, gun
|
||
selection, damage, wins, survival). No offline number here is a win/damage
|
||
claim, and no phase may be called a success without a live A/B
|
||
(`tools/ab/ab_run.sh`, server-side event hit rate, left-running).
|
||
|
||
Every claim below is tagged **[MEASURED]** (a command in §0 reproduces it) or
|
||
**[INFERRED]** (reasoning from measured facts).
|
||
|
||
---
|
||
|
||
## Phase 0 — BUILD THE RULER AND ESTABLISH THE BAR *(owner: overnight job, committed)*
|
||
|
||
### 0.1 The ruler
|
||
|
||
`common_libs/gun_harness/prediction_quality.nim` + `common_libs/tests/run_prediction_quality.nim`.
|
||
|
||
At each recorded tick the shooter sits at `O = (selfX, selfY)`. For a bullet of
|
||
speed `v` the **true interception point** is the first fractional time `t > 0`
|
||
at which the enemy's ACTUAL recorded track reaches distance `v*t` from `O`
|
||
(linear interpolation between recorded ticks). A bullet fired along the bearing
|
||
to `E(t)` coincides with the enemy at `t`. Angular error is
|
||
`wrap180(bearing(O→pred) − bearing(O→E(t)))` in **degrees**; every tick is
|
||
scored for the four power bins (speeds 17/15.5/14/11), all bands share that
|
||
horizon set. Per range band we report `mean|err|`, RMSE, mean signed err and the
|
||
hit-probability proxy `mean(|err| ≤ atan(18/range))`.
|
||
|
||
The integer-tick solve from `analyze_lead_capture_by_range.py` (commit
|
||
`f91e121`) is kept as `interceptBearingQuant` and reported as `OracleQuant`; the
|
||
ruler ships the **continuous** solve because it separates recorded hits from
|
||
misses slightly better and removes the coarse solve's own overshoot
|
||
(§0.3.6). `--ruler quant` selects the integer solve.
|
||
|
||
**Data:** the recorded live-vs-real-DrussGT corpus `/tmp/tfil_ab2/out`
|
||
(70 battles / 490 rounds / 899 607 ticks + `.events.jsonl` + `.rounds.json`),
|
||
**verified present before use**. It lives in `/tmp` and is therefore ephemeral;
|
||
if a later job finds it gone, regenerate it with the A/B harness
|
||
(`tools/ab/ab_run.sh`, which sets `TR_RECORD_WORLDSTATE` so ModularBot appends
|
||
per-tick world state) and point `--corpus` at the new output root. Layout:
|
||
`<root>/<arm>/runN.jsonl` + `runN.events.jsonl` + `runN.jsonl.rounds.json`.
|
||
149 MB of JSONL is converted once per run into a compact float32 `.qcache`
|
||
(keyed on source mtime+size) and ALL measurement is taken from the cache, so two
|
||
runs are byte-identical. See §0.4 for speed.
|
||
|
||
### 0.2 Validation — the ruler must pass ALL of these **[MEASURED]**
|
||
|
||
Run: `nim c -d:release --nimcache:/tmp/nc_j98 -r common_libs/tests/run_prediction_quality.nim`
|
||
|
||
**1. Recorded HITS separate from recorded MISSES** (our ACTUAL server-fired
|
||
bearings, scored against the SAME interception solve):
|
||
|
||
| ruler | hits n | hits mean\|err\| | misses n | misses mean\|err\| | separation |
|
||
|---|---|---|---|---|---|
|
||
| continuous | 5480 | **1.360° / 10.5 px** | 48304 | **16.597° / 140.3 px** | **12.20× deg / 13.34× px** |
|
||
| integer | 5480 | 1.478° / 11.4 px | 48304 | 16.724° / 141.3 px | 11.32× / 12.43× |
|
||
|
||
(The earlier validated run quoted 11.59× / 11.6 px on hits; reproduced and
|
||
improved.) The continuous ruler is shipped because it separates better.
|
||
|
||
**2. Perfect oracle scores 0.** Max `|err|` over all 3 598 428 tick-bins =
|
||
**0.000000°**. OK.
|
||
|
||
**3. A static line-of-sight gun is far from the predictor on learnable motion.**
|
||
On a synthetic constant-velocity and a seeded random-walk trajectory the
|
||
ordering is exactly as physics demands: HeadOn (zero lead) is the worst, Pattern
|
||
and naive-linear are near-zero, and the lead-gain arms overshoot monotonically.
|
||
On the real DrussGT corpus the static gun is *not* worst — see §0.3.4, this is a
|
||
genuine property of the corpus, not a harness defect.
|
||
|
||
**4. Determinism.** Two full 70-run sweeps, stdout diffed with the two wall-time
|
||
lines excluded: **byte-identical**. (The only difference between the two raw
|
||
outputs is `wall time 406.88s` vs `402.41s` and the derived ms-per-tick-bin.)
|
||
**[MEASURED]**
|
||
|
||
**5. A real bug was found and fixed by this validation.** The ruler's
|
||
`wrap180` used Nim's float `mod`, which keeps the dividend's sign (C `fmod`), so
|
||
`(x+180) mod 360 − 180` returned `x−360` instead of the wrapped equivalent for
|
||
`x < −180`. This inflated the negative tail of every error (maxAbs read ~360°
|
||
instead of ~180°) and made HeadOn's mean error disagree with `mean|required|`.
|
||
After the fix HeadOn's `mean|err|` equals `mean|required lead|` to the last
|
||
digit at every band (see the `HO |err| / HO |req| / Pat|req|` columns in the
|
||
fixture). **A wrong ruler is worse than no ruler; this was the most important
|
||
10 minutes of the phase.**
|
||
|
||
### 0.3 THE BAR — per-band numbers (70 runs, 3 598 428 tick-bins) **[MEASURED]**
|
||
|
||
Format: `mean|err| deg` and, in brackets, `hitProxy`. `hitProxy` is the fraction
|
||
of tick-bins aimed within `atan(18/range)` of the true interception point.
|
||
|
||
| band | Pattern | naive-linear | TMHorizon | BitBrain | HeadOn (static) | Oracle |
|
||
|---|---|---|---|---|---|---|
|
||
| 0–100 | **10.56** [0.699] | 17.92 [0.651] | 10.68 [0.696] | 11.12 [0.681] | 19.62 [0.342] | 0.00 [1.000] |
|
||
| 100–200 | **14.75** [0.342] | 14.63 [0.388] | 14.84 [0.342] | 15.11 [0.324] | 19.98 [0.172] | 0.00 [1.000] |
|
||
| 200–300 | **16.61** [0.185] | 17.57 [0.192] | 16.64 [0.179] | 16.84 [0.175] | 17.34 [0.133] | 0.00 [1.000] |
|
||
| 300–450 | 17.53 [0.104] | 20.98 [0.100] | 17.57 [0.100] | 17.58 [0.103] | **14.61** [0.105] | 0.00 [1.000] |
|
||
| 450+ | 16.19 [0.077] | 22.86 [0.054] | 16.20 [0.076] | 16.20 [0.077] | **12.33** [0.098] | 0.00 [1.000] |
|
||
|
||
`n`: 4 423 / 24 908 / 74 215 / 1 119 777 / 2 311 323. Skipped (no valid
|
||
interception): 63 782 tick-bins (≈1.7 %).
|
||
|
||
**0.3.1 DIRECT ANSWER — the gap between Pattern and the oracle ceiling:**
|
||
|
||
| band | Pattern hitProxy | Oracle hitProxy | headroom (pp) |
|
||
|---|---|---|---|
|
||
| 0–100 | 0.6993 | 1.0000 | **+30.07** |
|
||
| 100–200 | 0.3418 | 1.0000 | **+65.82** |
|
||
| 200–300 | 0.1850 | 1.0000 | **+81.50** |
|
||
| 300–450 | 0.1036 | 1.0000 | **+89.64** |
|
||
| 450+ | 0.0767 | 1.0000 | **+92.33** |
|
||
|
||
**[INFERRED, important]** The oracle is *non-causal*: it aims with perfect
|
||
knowledge of the enemy's future, so its 100 % is a definition, not an
|
||
achievement, and the 92 pp at 450+ is an UPPER bound that contains both "a
|
||
better predictor could get this" and "this is physically unknowable". The
|
||
**realistic** causal bound measured today is the best arm at 450+: **HeadOn at
|
||
9.8 %**, barely above Pattern's 7.7 %. So the campaign is playing for a few
|
||
percentage points at long range, not for 92 pp. The honest target statement is
|
||
"raise the 450+ proxy from 7.7 % toward the ~10 % causal band", not "toward
|
||
100 %".
|
||
|
||
**0.3.2 The lead-gain sweep on Pattern is a DEAD END [MEASURED].** Multiply
|
||
Pattern's angular lead over LOS by a constant, per band:
|
||
|
||
| band | gain 1.0 | gain 1.5 | gain 2.0 | gain 3.0 |
|
||
|---|---|---|---|---|
|
||
| 0–100 | **10.56** | 13.76 | 20.11 | 34.77 |
|
||
| 100–200 | **14.75** | 19.64 | 26.78 | 42.97 |
|
||
| 200–300 | **16.61** | 22.26 | 29.47 | 45.23 |
|
||
| 300–450 | **17.53** | 23.28 | 29.96 | 44.27 |
|
||
| 450+ | **16.19** | 21.25 | 27.00 | 39.27 |
|
||
|
||
Gain 1.0 wins at EVERY band. Scaling Pattern's lead up makes it strictly worse.
|
||
This is the single most important negative result of Phase 0 and it should stop
|
||
any later job from "just adding more lead".
|
||
|
||
**0.3.3 The naive-linear / capture tension, resolved [MEASURED].** Capture slope
|
||
= regression of the arm's own lead on the required lead (job-95's statistic);
|
||
corr = Pearson correlation of the arm's lead with the required lead. **corr is
|
||
the informative number; a large slope on an uncorrelated lead is just amplified
|
||
noise.**
|
||
|
||
| band | mean\|req\| | Pattern cap / corr | naive-linear cap / corr | TMHorizon | BitBrain |
|
||
|---|---|---|---|---|---|
|
||
| 100–200 | 19.98 | 0.553 / 0.612 | 0.592 / 0.523 | 0.560 / 0.614 | 0.570 / 0.611 |
|
||
| 300–450 | 14.61 | 0.278 / 0.266 | 0.476 / 0.324 | 0.273 / 0.262 | 0.280 / 0.266 |
|
||
| 450+ | 12.33 | 0.175 / 0.165 | 0.310 / 0.178 | 0.174 / 0.164 | 0.175 / 0.165 |
|
||
|
||
Yes — on this corpus the naive-linear predictor applies **~1.8× more lead** than
|
||
Pattern at 450+ (0.310 vs 0.175; job-95 measured ~2×). Job-95's "we under-lead"
|
||
reading is confirmed. **But** the two arms carry almost the same lead
|
||
*information* (corr 0.178 vs 0.165), so the extra amplitude buys nothing and
|
||
costs angular accuracy: naive-linear's `mean|err|` is 22.86° vs Pattern's
|
||
16.19° at 450+. **Conclusion: the campaign's lever is lead INFORMATION
|
||
(correlation), not lead RESPONSE (capture slope).** Capturing more of an
|
||
uninformative lead is worse than capturing little of it — which is also exactly
|
||
why the gain sweep fails.
|
||
|
||
**0.3.4 The surprise: at long range, static line-of-sight beats Pattern.**
|
||
HeadOn (aim at the enemy's current position) has `mean|err|` 14.61°/12.33° and
|
||
`hitProxy` 0.105/0.098 at 300–450/450+, both better than Pattern's
|
||
17.53°/16.19° and 0.104/0.077. **[INFERRED]** At 450+ the required lead
|
||
(`mean|req|` = 12.3°) is essentially unpredictable from the past (Pattern
|
||
corr 0.165), so Pattern's predicted lead is mostly variance added to a nearly
|
||
uninformative signal; a zero-lead aim has error = `|required lead|`, which is
|
||
smaller. Consistent with the live record: the live bot's own applied lead
|
||
capture was 0.135 at 450+ (job-95), i.e. the live bot was already nearly
|
||
zero-lead and hit 9.14 % there; Pattern's offline proxy is 7.7 %.
|
||
|
||
**[INFERRED / CAVEAT]** The corpus is open-loop: DrussGT's recorded dodge was a
|
||
reaction to the LIVE bot's (near-zero-lead) bullets. Replaying Pattern on that
|
||
trajectory cannot show what DrussGT would do against Pattern's bullets. This
|
||
makes a **live A/B of HeadOn vs Pattern at long range the highest-value cheap
|
||
experiment in the campaign** (see §0.6). No offline claim that "HeadOn
|
||
beats Pattern" is permitted — only the live A/B decides.
|
||
|
||
**0.3.5 BitBrain, as shipped, is Pattern [MEASURED].** BitBrain's base is
|
||
Pattern and its ADE/SBC corrector changes almost nothing: 450+ `mean|err|`
|
||
16.200° vs Pattern 16.193°, `hitProxy` 0.0767 vs 0.0767. TMHorizon likewise
|
||
(16.199° / 0.0757). The corrector is currently **adding no measurable aim
|
||
information** on this corpus. That is the thing Phase 1 must change.
|
||
|
||
**0.3.6 Ruler resolution is NOT the limiter [MEASURED].** Aiming at the
|
||
integer-tick solve instead of the exact intercept costs only 0.29–1.10° of mean
|
||
error (`OracleQuant` column). So the "maybe the oracle only reaches 35 % because
|
||
the solve is coarse" worry is dead: the coarse/fine difference is ≈0.3° at long
|
||
range, far below the target tolerance (1.93° at 450+). Whatever caps the score,
|
||
it is the enemy's unpredictability, not the ruler.
|
||
|
||
### 0.4 Speed **[MEASURED]**
|
||
|
||
Full 70-run / 899 607-tick / 3 598 428 tick-bin sweep, 10 arms:
|
||
**406.9 s wall**, i.e. `0.1131 ms per tick-bin` over 10 arms,
|
||
**≈ 0.045 s per gun per 1000 ticks** (1000 ticks × 4 power bins).
|
||
BitBrain is the dominant cost (its ADE pass runs on every `predict` call);
|
||
Pattern/TMHorizon cache their per-tick work. A single-arm Pattern-only sweep is
|
||
several times cheaper. The binary cache (§0.1) is what makes repeat sweeps
|
||
affordable: without it every run re-parses 149 MB of JSONL.
|
||
|
||
### 0.5 How to reproduce **[MEASURED]**
|
||
|
||
```
|
||
nim c -d:release --nimcache:/tmp/nc_j98 -o:/tmp/bbq_run \
|
||
common_libs/tests/run_prediction_quality.nim
|
||
/tmp/bbq_run --corpus /tmp/tfil_ab2/out # full bar, ~7 min
|
||
/tmp/bbq_run --corpus /tmp/tfil_ab2/out --limit 10 # fast subset
|
||
/tmp/bbq_run --corpus /tmp/tfil_ab2/out --ruler quant # integer-tick solve
|
||
```
|
||
Verbatim full output: `common_libs/tests/prediction_quality_results.txt`.
|
||
Determinism: two consecutive full runs are byte-identical except the two
|
||
wall-time lines.
|
||
|
||
**Clean-checkout proof [MEASURED]:** `git archive HEAD | tar -x -C /tmp/bbq_clean`
|
||
then, from `/tmp/bbq_clean`,
|
||
`nim c -d:release --nimcache:/tmp/nc_j98 -o:bbq_run common_libs/tests/run_prediction_quality.nim`
|
||
builds, and `./bbq_run --corpus /tmp/tfil_ab2/out --limit 3` runs and prints the
|
||
same tables (separation 13.68× px on the 3-run subset). The committed harness is
|
||
self-contained; only the corpus is external.
|
||
|
||
### 0.6 Designs still to try (seed for later phases)
|
||
|
||
Ordered by expected value per unit of effort. Phase 0 has already killed one.
|
||
|
||
| # | design | why it is worth trying | status |
|
||
|---|---|---|---|
|
||
| D1 | **Live A/B: HeadOn at 450+ vs Pattern** (distance-gated switch, or HeadOn-only control) | Offline says a static gun beats Pattern at long range; the cheapest possible test of the campaign's central premise | **TODO (highest value, live)** |
|
||
| D2 | **Pattern variants that raise lead CORRELATION at long range**: longer keys / multi-length keys, per-distance learned pattern tables, different match weighting, k-NN over movement signatures | The lever is corr (0.165 at 450+), not gain; the ruler measures exactly this | TODO (offline-searchable) |
|
||
| D3 | **Supervise BitBrain with the ruler's own labels** — per-tick bearing error to the true intercept, trained on N−1 runs, evaluated on a held-out run | BitBrain's corrector currently adds nothing (0.3.5); the ruler gives it a real target. Must hold out runs or it is overfitting | TODO (offline-searchable) |
|
||
| D4 | **A causal "predictability" gate**: at each tick estimate whether the future is predictable (e.g. recent pattern-match score, reversal entropy) and fall back to HeadOn/low-variance aim when it is not | Directly attacks the 0.3.4 failure mode without needing a better long-range predictor | TODO |
|
||
| D5 | Power policy at long range (already partly done live): lower power = faster bullet = less lead error | Shortens the horizon the predictor must extrapolate; affects hit rate, live-only verdict | TODO (offline proxy only) |
|
||
| D6 | Lead-gain sweep 1.0/1.5/2.0/3.0 | **DEAD — measured.** Gain 1.0 wins at every band (§0.3.2) | **KILLED** |
|
||
|
||
Every D-item must end in a live A/B before any phase verdict.
|
||
|
||
### 0.7 What would make us quit
|
||
|
||
> If (a) no causal design raises the 450+ `hitProxy` above the static-gun
|
||
> reference (~0.10) on held-out runs by a margin larger than the run-to-run
|
||
> spread, **and** (b) the live A/B of the best such design shows no hit-rate or
|
||
> damage gain over Pattern with the left-running liveness check satisfied, then
|
||
> the campaign stops and we ship the simpler gun. We do not keep tuning an
|
||
> offline proxy that has stopped predicting live outcomes.
|
||
|
||
---
|
||
|
||
## Phase 1: the missing gain sweep *(owner: overnight job j100, committed)*
|
||
|
||
### Three negatives are on file (the morning reader must see these)
|
||
|
||
1. **BitBrain as previously shipped was statistically identical to Pattern
|
||
live.** 30 runs/arm, MDE 24.4 dmg/run: `docs/bitbrain_gun_verdict.md`,
|
||
commit `d93ce44`.
|
||
2. **Gun-mixing does not disrupt DrussGT.** TMHorizon+BitBrain mixed into the
|
||
rack produced no dodge disruption: `docs/gun_mix_disruption.md`, commit
|
||
`32a5e72`.
|
||
3. **BitBrain added no measurable aim information offline.** 450+: 16.200 deg
|
||
vs Pattern 16.193, hitProxy 0.0767 vs 0.0767 (`a82c864`, §0.3.5 above).
|
||
|
||
These bound the plausible upside: the previous BitBrain output — an ADDITIVE
|
||
angular shift — was information-free, so Phase 1 changes the output shape, not
|
||
the learning rate.
|
||
|
||
### 1.1 The missing sweep: Pattern x gain in [0.00, 1.00] **[MEASURED]**
|
||
|
||
`common_libs/tests/run_prediction_quality.nim` now carries sub-unity gain arms
|
||
(gain 0.0 is HeadOn, gain 1.0 is Pattern). Full 70-run sweep, 14 arms,
|
||
3 598 428 tick-bins, `common_libs/tests/prediction_quality_results.txt`.
|
||
Cells are `mean|err| deg [hitProxy]`; **hitProxy is the objective** (fraction of
|
||
tick-bins within `atan(18/range)`, the ruler's proxy for hit probability).
|
||
|
||
| band | |req| deg | g=0.00 | g=0.25 | g=0.50 | g=0.75 | g=1.00 | best hpx | dHpx | best |err| |
|
||
|---|---|---|---|---|---|---|---|---|
|
||
| 0–100 | 19.62 | 19.62 [.342] | 16.98 [.391] | 14.58 [.469] | 12.36 [.640] | **10.56 [.699]** | 1.00 | +.0000 | 1.00 |
|
||
| 100–200 | 19.98 | 19.98 [.172] | 17.79 [.166] | 16.10 [.167] | 14.97 [.234] | **14.75 [.342]** | 1.00 | +.0000 | 1.00 |
|
||
| 200–300 | 17.34 | 17.34 [.133] | 15.95 [.128] | 15.32 [.137] | 15.53 [.139] | **16.61 [.185]** | 1.00 | +.0000 | 0.50 |
|
||
| 300–450 | 14.61 | **14.61 [.105]** | 14.13 [.102] | 14.49 [.100] | 15.64 [.095] | 17.53 [.104] | 0.00 | +.0013 | 0.25 |
|
||
| 450+ | 12.33 | **12.33 [.098]** | 12.26 [.093] | 12.94 [.088] | 14.28 [.081] | 16.19 [.077] | 0.00 | +.0216 | 0.25 |
|
||
|
||
**The optimal gain curve (hitProxy-argmax per band) is**
|
||
**`[1.00, 1.00, 1.00, 0.00, 0.00]`** — use Pattern's full lead below 300 px and
|
||
**aim at the enemy's current position (zero lead, HeadOn) at 300+ px**. The
|
||
implied hit-probability gains vs Pattern are +0.00 pp (0–300), +0.13 pp
|
||
(300–450) and **+2.16 pp at 450+ (0.0767 -> 0.0984, a 28 % relative rise)**.
|
||
Weighted by tick-bin share (450+ is 64.2 %, 300–450 is 31.1 %) the whole-corpus
|
||
proxy rises ~+1.4 pp.
|
||
|
||
**[CAUSAL-SHIPPABILITY, important]** The rule only needs the *range*, which is
|
||
known at fire time, so applying `[1,1,1,0,0]` is causal and needs **no learning
|
||
at all** — a per-band table is shippable as a constant, exactly like the
|
||
Pattern radial-offset knob. What is *not* causal is the **estimation** of the
|
||
table from the same runs (it is in-sample here); a shipped table would be fitted
|
||
offline on past battles or learned online, which is what BitBrain does. The
|
||
table's value is robust to that caveat because the winning entries are the two
|
||
extremes (full Pattern lead, or zero lead), not a knife-edge intermediate value.
|
||
|
||
### 1.2 Shrinking the gain does NOT add lead information **[MEASURED]**
|
||
|
||
| band | corr @ g=0.25 | g=0.50 | g=0.75 | g=1.00 |
|
||
|---|---|---|---|---|
|
||
| 0–100 | 0.774 | 0.774 | 0.774 | 0.774 |
|
||
| 100–200 | 0.612 | 0.612 | 0.612 | 0.612 |
|
||
| 200–300 | 0.457 | 0.457 | 0.457 | 0.457 |
|
||
| 300–450 | 0.266 | 0.266 | 0.266 | 0.266 |
|
||
| 450+ | 0.165 | 0.165 | 0.165 | 0.165 |
|
||
|
||
Every `g > 0` column is **identical**: Pearson correlation is invariant under
|
||
positive scaling. A fractional gain therefore buys nothing on the
|
||
lead-information axis — it only shrinks the magnitude of an uninformative signal
|
||
toward the low-variance static aim. This confirms §0.3.3's reading and is the
|
||
mechanism behind the whole curve.
|
||
|
||
**The MSE-optimal gain and the hitProxy-optimal gain diverge [MEASURED, key].**
|
||
The `best |err|` column is the least-squares optimum `[1.00, 1.00, 0.50, 0.25,
|
||
0.25]`. At 200–300, MSE wants ~0.5 but hitProxy wants 1.0: Pattern's lead errors
|
||
are **bimodal** (it either nails the lead or is far off), so shrinking every
|
||
sample trades many small-within-tolerance hits for a smaller tail. Any *learned*
|
||
corrector that minimises squared error will therefore under-perform at mid range
|
||
— a fact this phase measured the hard way (an MSE-gain BitBrain dropped the
|
||
200–300 proxy from 0.190 to 0.131).
|
||
|
||
### 1.3 BitBrain rebuilt as a lead-gain corrector **[MEASURED]**
|
||
|
||
`common_libs/guns/bitbrain_gun.nim` is rewritten. The ADE+SBC network is gone
|
||
from the gun; the output is now a multiplicative gain on Pattern's lead,
|
||
`aim = LOS + gain * (patternAim - LOS)`, with `gain` chosen from a fixed
|
||
candidate set `{0, 0.25, 0.5, 0.75, 1.0}`.
|
||
|
||
* **Label path** (unchanged): at fire time we remember Pattern's lead and the
|
||
target's angular half-width `atan(18/range)`; `h = round(dist/speed)` ticks
|
||
later `tmhObservedAt` returns the enemy's observed bearing from the firing
|
||
position, giving `requiredLead = observedBearing - LOS`.
|
||
* **Learning rule** (changed): for every resolved sample we score *each*
|
||
candidate by whether `|gain*baseLead - requiredLead| <= atan(18/range)` and keep
|
||
the hit counts per range band; the band's gain is the **argmax hit rate** — the
|
||
hit-probability proxy itself, not squared error. This directly fixes the
|
||
bimodality failure above.
|
||
* **State / gate**: the state is the range band (causally known). The correction
|
||
is applied only for range >= 300 px (`BB_GAIN_BAND_MIN = 3`), below which
|
||
Pattern's lead is informative.
|
||
* **Stats are battle-scale**: a round boundary keeps the counts; a new battle or
|
||
target change (`resetLearning`/`targetChanged`) wipes them.
|
||
|
||
Offline 70-run result (same ruler, same run as §1.1):
|
||
|
||
| band | Pattern hpx | fixed rule [1,1,1,0,0] | BitBrain hpx | fixed−Pat | BB−Pat |
|
||
|---|---|---|---|---|---|
|
||
| 0–100 | 0.6993 | 0.6993 | 0.6993 | +0.0000 | +0.0000 |
|
||
| 100–200 | 0.3418 | 0.3418 | 0.3418 | +0.0000 | +0.0000 |
|
||
| 200–300 | 0.1850 | 0.1850 | 0.1850 | +0.0000 | +0.0000 |
|
||
| 300–450 | 0.1036 | 0.1049 | **0.1073** | +0.0013 | **+0.0037** |
|
||
| 450+ | 0.0767 | **0.0984** | 0.0957 | +0.0216 | **+0.0190** |
|
||
|
||
BitBrain's effective point estimates match the fixed table to within 0.27 pp at
|
||
450+ and actually exceed it at 300–450. The internal log shows why it is not
|
||
exactly equal: at long range the candidate hit rates are near-tied, so the
|
||
argmax flips between 0.00 and 0.25 (450+) / 0.00 and 0.75 (300–450); the
|
||
aggregate still lands on the right side. A fixed table is more stable; the
|
||
learned version needs no table and adapts per battle.
|
||
|
||
**Cost [MEASURED].** `/tmp/bench_bb` (200k ticks, 4 power bins) — Pattern
|
||
0.0224 ms/tick, BitBrain 0.0230 ms/tick, i.e. a **marginal ~0.0007 ms/tick**
|
||
(0.7 us/tick), versus the ~0.114 ms/tick of the old ADE+SBC gun and the
|
||
13.16 ms/tick budget. The whole 14-arm offline sweep now runs in 271 s
|
||
(0.0754 ms per tick-bin vs Phase 0's 0.1131 for 10 arms).
|
||
|
||
**Guard tests stay green [MEASURED].** `test_bitbrain` 32 checks,
|
||
`test_bitbrain_registration` 13, `test_rack_membership` 48,
|
||
`test_tm_pattern_registration` 20, `test_env_report` all pass. `env_report.nim`
|
||
is unchanged: the legacy BitBrain knobs are still resolved and reported. The
|
||
ruler still validates (recorded hits 1.360 deg vs misses 16.597 deg,
|
||
separation 12.20x deg / 13.34x px; oracle 0.000000 deg).
|
||
|
||
### 1.4 Verdict
|
||
|
||
**[MEASURED] Yes — a per-band lead-gain rule beats Pattern offline**, but only in
|
||
the two long bands: `[1,1,1,0,0]` gives +0.13 pp at 300–450 and **+2.16 pp at
|
||
450+** (a ~28 % relative rise in the hit-proxy, against Phase 0's realistic
|
||
causal ceiling of HeadOn's ~0.10, which it reaches). Below 300 px Pattern is
|
||
already optimal and the rule is a no-op.
|
||
|
||
**[MEASURED] Yes — BitBrain-as-gain-corrector matches that rule** (+1.90 pp at
|
||
450+, +0.37 pp at 300–450 over Pattern; within 0.27 pp of the fixed table at
|
||
450+, above it at 300–450), while being **causally learnable online** and
|
||
~163x cheaper per tick than the old gun.
|
||
|
||
**[INFERRED / honest caveat] BitBrain is not required to capture the gain** — a
|
||
constant `[1,1,1,0,0]` table would do it. BitBrain's value is that it discovers
|
||
the table per battle without one; its cost is cold-start (it begins at gain 1.0,
|
||
explaining the 0.27 pp gap at 450+) and the near-tie jitter in its point
|
||
estimates. So the honest ship decision is: **the per-band rule is the thing to
|
||
test live, and it can be shipped either as a constant table or as this online
|
||
learner** — the learner is redundant if a table is acceptable, and preferable
|
||
only if the optimum is expected to drift per enemy. The highest-value live
|
||
experiment is unchanged and stronger: a **range-gated Pattern<->HeadOn switch at
|
||
~300 px vs pure Pattern** (Phase 0's D1), left-running. **No live win is claimed
|
||
here** — the live gate is a separate phase.
|
||
|
||
### 1.5 Designs after Phase 1
|
||
|
||
| # | design | status |
|
||
|---|---|---|
|
||
| D1 | Live A/B: Pattern<->HeadOn switch at ~300 px vs Pattern | **TODO (highest value, live)** — offline now says +2.16 pp at 450+ |
|
||
| D2 | Raise Pattern lead *correlation* at long range (longer/multi-length keys, per-distance tables, k-NN) | TODO (offline-searchable) |
|
||
| D3 | Match-quality-conditional gain (keep full lead when the pattern match is good, shrink when bad) — attacks the bimodality that makes the MSE gain diverge | TODO (offline-searchable) |
|
||
| D4 | A causal predictability gate falling back to HeadOn when the future is unpredictable | TODO |
|
||
| D5 | Long-range power policy (already partly live) | TODO (offline proxy only) |
|
||
| D6 | Lead-gain sweep gains >= 1.0 | **DEAD — measured** (§0.3.2) |
|
||
| D7 | Constant sub-unity lead gain at all ranges | **DEAD — measured** (§1.1: gain < 1 hurts below 300) |
|
||
|
||
|
||
## Phase 2 — *(unclaimed; append below)*
|