Task A — the gain region Phase 0 never covered (gain < 1). Extend the prediction-quality ruler with gain 0.25/0.50/0.75 arms and a fixed causal per-band arm. Full 70-run result: the hitProxy-argmax curve is [1.00, 1.00, 1.00, 0.00, 0.00] — Pattern below 300 px, HeadOn above — worth +0.13 pp at 300-450 and +2.16 pp at 450+ (0.0767 -> 0.0984). Lead correlation is identical for every g>0 (Pearson is scale-invariant), so a shrinking gain adds no lead information; and the least-squares optimum [1,1,1,1,0.25,0.25] diverges from the hitProxy optimum because Pattern's lead errors are bimodal. Task B — rebuild guns/bitbrain_gun.nim as a lead-gain corrector: aim = LOS + gain*(patternAim - LOS), gain learned online per range band by ranking candidate gains on the hit-probability proxy (the observed lead label via tmhObservedAt), gated to range >= 300 px. ADE+SBC output removed. Offline (70 runs): Pattern below 300 px, +0.37 pp at 300-450, +1.90 pp at 450+ (hitProxy 0.0957 vs 0.0767), matching the fixed rule to within 0.27 pp at 450+ and exceeding it at 300-450. ~0.0007 ms/tick marginal (old gun ~0.114 ms/tick). Default off; rack membership, env report and guard tests (bitbrain 32, registration 13, rack 48, tm_pattern 20, env_report) unchanged and green. Ledger: docs/bitbrain_campaign.md Phase 1, including the three on-file negatives and the causal-shippability note. No live claim.
24 KiB
BitBrain campaign ledger
Goal: make ModularBot's gun beat Pattern live against the real DrussGT. The user has granted full freedom over the gun ("change input, output, every knob of it") and accepts it may fail — the deliverable is that the attempt is visible and evidence-backed.
THE FINAL VERDICT IS ALWAYS LIVE. Everything in this file except the
## Phase N verdict lines is offline, open-loop, on a fixed recorded enemy
trajectory. Per docs/offline_harness_trust.md (commit e40c849) the offline
harness is trustworthy for exactly one thing: per-gun single-tick prediction
quality on a fixed enemy trajectory — and it is never trustworthy for
closed-loop questions (movement, range, round length, adaptation, gun
selection, damage, wins, survival). No offline number here is a win/damage
claim, and no phase may be called a success without a live A/B
(tools/ab/ab_run.sh, server-side event hit rate, left-running).
Every claim below is tagged [MEASURED] (a command in §0 reproduces it) or [INFERRED] (reasoning from measured facts).
Phase 0 — BUILD THE RULER AND ESTABLISH THE BAR (owner: overnight job, committed)
0.1 The ruler
common_libs/gun_harness/prediction_quality.nim + common_libs/tests/run_prediction_quality.nim.
At each recorded tick the shooter sits at O = (selfX, selfY). For a bullet of
speed v the true interception point is the first fractional time t > 0
at which the enemy's ACTUAL recorded track reaches distance v*t from O
(linear interpolation between recorded ticks). A bullet fired along the bearing
to E(t) coincides with the enemy at t. Angular error is
wrap180(bearing(O→pred) − bearing(O→E(t))) in degrees; every tick is
scored for the four power bins (speeds 17/15.5/14/11), all bands share that
horizon set. Per range band we report mean|err|, RMSE, mean signed err and the
hit-probability proxy mean(|err| ≤ atan(18/range)).
The integer-tick solve from analyze_lead_capture_by_range.py (commit
f91e121) is kept as interceptBearingQuant and reported as OracleQuant; the
ruler ships the continuous solve because it separates recorded hits from
misses slightly better and removes the coarse solve's own overshoot
(§0.3.6). --ruler quant selects the integer solve.
Data: the recorded live-vs-real-DrussGT corpus /tmp/tfil_ab2/out
(70 battles / 490 rounds / 899 607 ticks + .events.jsonl + .rounds.json),
verified present before use. It lives in /tmp and is therefore ephemeral;
if a later job finds it gone, regenerate it with the A/B harness
(tools/ab/ab_run.sh, which sets TR_RECORD_WORLDSTATE so ModularBot appends
per-tick world state) and point --corpus at the new output root. Layout:
<root>/<arm>/runN.jsonl + runN.events.jsonl + runN.jsonl.rounds.json.
149 MB of JSONL is converted once per run into a compact float32 .qcache
(keyed on source mtime+size) and ALL measurement is taken from the cache, so two
runs are byte-identical. See §0.4 for speed.
0.2 Validation — the ruler must pass ALL of these [MEASURED]
Run: nim c -d:release --nimcache:/tmp/nc_j98 -r common_libs/tests/run_prediction_quality.nim
1. Recorded HITS separate from recorded MISSES (our ACTUAL server-fired bearings, scored against the SAME interception solve):
| ruler | hits n | hits mean|err| | misses n | misses mean|err| | separation |
|---|---|---|---|---|---|
| continuous | 5480 | 1.360° / 10.5 px | 48304 | 16.597° / 140.3 px | 12.20× deg / 13.34× px |
| integer | 5480 | 1.478° / 11.4 px | 48304 | 16.724° / 141.3 px | 11.32× / 12.43× |
(The earlier validated run quoted 11.59× / 11.6 px on hits; reproduced and improved.) The continuous ruler is shipped because it separates better.
2. Perfect oracle scores 0. Max |err| over all 3 598 428 tick-bins =
0.000000°. OK.
3. A static line-of-sight gun is far from the predictor on learnable motion. On a synthetic constant-velocity and a seeded random-walk trajectory the ordering is exactly as physics demands: HeadOn (zero lead) is the worst, Pattern and naive-linear are near-zero, and the lead-gain arms overshoot monotonically. On the real DrussGT corpus the static gun is not worst — see §0.3.4, this is a genuine property of the corpus, not a harness defect.
4. Determinism. Two full 70-run sweeps, stdout diffed with the two wall-time
lines excluded: byte-identical. (The only difference between the two raw
outputs is wall time 406.88s vs 402.41s and the derived ms-per-tick-bin.)
[MEASURED]
5. A real bug was found and fixed by this validation. The ruler's
wrap180 used Nim's float mod, which keeps the dividend's sign (C fmod), so
(x+180) mod 360 − 180 returned x−360 instead of the wrapped equivalent for
x < −180. This inflated the negative tail of every error (maxAbs read ~360°
instead of ~180°) and made HeadOn's mean error disagree with mean|required|.
After the fix HeadOn's mean|err| equals mean|required lead| to the last
digit at every band (see the HO |err| / HO |req| / Pat|req| columns in the
fixture). A wrong ruler is worse than no ruler; this was the most important
10 minutes of the phase.
0.3 THE BAR — per-band numbers (70 runs, 3 598 428 tick-bins) [MEASURED]
Format: mean|err| deg and, in brackets, hitProxy. hitProxy is the fraction
of tick-bins aimed within atan(18/range) of the true interception point.
| band | Pattern | naive-linear | TMHorizon | BitBrain | HeadOn (static) | Oracle |
|---|---|---|---|---|---|---|
| 0–100 | 10.56 [0.699] | 17.92 [0.651] | 10.68 [0.696] | 11.12 [0.681] | 19.62 [0.342] | 0.00 [1.000] |
| 100–200 | 14.75 [0.342] | 14.63 [0.388] | 14.84 [0.342] | 15.11 [0.324] | 19.98 [0.172] | 0.00 [1.000] |
| 200–300 | 16.61 [0.185] | 17.57 [0.192] | 16.64 [0.179] | 16.84 [0.175] | 17.34 [0.133] | 0.00 [1.000] |
| 300–450 | 17.53 [0.104] | 20.98 [0.100] | 17.57 [0.100] | 17.58 [0.103] | 14.61 [0.105] | 0.00 [1.000] |
| 450+ | 16.19 [0.077] | 22.86 [0.054] | 16.20 [0.076] | 16.20 [0.077] | 12.33 [0.098] | 0.00 [1.000] |
n: 4 423 / 24 908 / 74 215 / 1 119 777 / 2 311 323. Skipped (no valid
interception): 63 782 tick-bins (≈1.7 %).
0.3.1 DIRECT ANSWER — the gap between Pattern and the oracle ceiling:
| band | Pattern hitProxy | Oracle hitProxy | headroom (pp) |
|---|---|---|---|
| 0–100 | 0.6993 | 1.0000 | +30.07 |
| 100–200 | 0.3418 | 1.0000 | +65.82 |
| 200–300 | 0.1850 | 1.0000 | +81.50 |
| 300–450 | 0.1036 | 1.0000 | +89.64 |
| 450+ | 0.0767 | 1.0000 | +92.33 |
[INFERRED, important] The oracle is non-causal: it aims with perfect knowledge of the enemy's future, so its 100 % is a definition, not an achievement, and the 92 pp at 450+ is an UPPER bound that contains both "a better predictor could get this" and "this is physically unknowable". The realistic causal bound measured today is the best arm at 450+: HeadOn at 9.8 %, barely above Pattern's 7.7 %. So the campaign is playing for a few percentage points at long range, not for 92 pp. The honest target statement is "raise the 450+ proxy from 7.7 % toward the ~10 % causal band", not "toward 100 %".
0.3.2 The lead-gain sweep on Pattern is a DEAD END [MEASURED]. Multiply Pattern's angular lead over LOS by a constant, per band:
| band | gain 1.0 | gain 1.5 | gain 2.0 | gain 3.0 |
|---|---|---|---|---|
| 0–100 | 10.56 | 13.76 | 20.11 | 34.77 |
| 100–200 | 14.75 | 19.64 | 26.78 | 42.97 |
| 200–300 | 16.61 | 22.26 | 29.47 | 45.23 |
| 300–450 | 17.53 | 23.28 | 29.96 | 44.27 |
| 450+ | 16.19 | 21.25 | 27.00 | 39.27 |
Gain 1.0 wins at EVERY band. Scaling Pattern's lead up makes it strictly worse. This is the single most important negative result of Phase 0 and it should stop any later job from "just adding more lead".
0.3.3 The naive-linear / capture tension, resolved [MEASURED]. Capture slope = regression of the arm's own lead on the required lead (job-95's statistic); corr = Pearson correlation of the arm's lead with the required lead. corr is the informative number; a large slope on an uncorrelated lead is just amplified noise.
| band | mean|req| | Pattern cap / corr | naive-linear cap / corr | TMHorizon | BitBrain |
|---|---|---|---|---|---|
| 100–200 | 19.98 | 0.553 / 0.612 | 0.592 / 0.523 | 0.560 / 0.614 | 0.570 / 0.611 |
| 300–450 | 14.61 | 0.278 / 0.266 | 0.476 / 0.324 | 0.273 / 0.262 | 0.280 / 0.266 |
| 450+ | 12.33 | 0.175 / 0.165 | 0.310 / 0.178 | 0.174 / 0.164 | 0.175 / 0.165 |
Yes — on this corpus the naive-linear predictor applies ~1.8× more lead than
Pattern at 450+ (0.310 vs 0.175; job-95 measured ~2×). Job-95's "we under-lead"
reading is confirmed. But the two arms carry almost the same lead
information (corr 0.178 vs 0.165), so the extra amplitude buys nothing and
costs angular accuracy: naive-linear's mean|err| is 22.86° vs Pattern's
16.19° at 450+. Conclusion: the campaign's lever is lead INFORMATION
(correlation), not lead RESPONSE (capture slope). Capturing more of an
uninformative lead is worse than capturing little of it — which is also exactly
why the gain sweep fails.
0.3.4 The surprise: at long range, static line-of-sight beats Pattern.
HeadOn (aim at the enemy's current position) has mean|err| 14.61°/12.33° and
hitProxy 0.105/0.098 at 300–450/450+, both better than Pattern's
17.53°/16.19° and 0.104/0.077. [INFERRED] At 450+ the required lead
(mean|req| = 12.3°) is essentially unpredictable from the past (Pattern
corr 0.165), so Pattern's predicted lead is mostly variance added to a nearly
uninformative signal; a zero-lead aim has error = |required lead|, which is
smaller. Consistent with the live record: the live bot's own applied lead
capture was 0.135 at 450+ (job-95), i.e. the live bot was already nearly
zero-lead and hit 9.14 % there; Pattern's offline proxy is 7.7 %.
[INFERRED / CAVEAT] The corpus is open-loop: DrussGT's recorded dodge was a reaction to the LIVE bot's (near-zero-lead) bullets. Replaying Pattern on that trajectory cannot show what DrussGT would do against Pattern's bullets. This makes a live A/B of HeadOn vs Pattern at long range the highest-value cheap experiment in the campaign (see §0.6). No offline claim that "HeadOn beats Pattern" is permitted — only the live A/B decides.
0.3.5 BitBrain, as shipped, is Pattern [MEASURED]. BitBrain's base is
Pattern and its ADE/SBC corrector changes almost nothing: 450+ mean|err|
16.200° vs Pattern 16.193°, hitProxy 0.0767 vs 0.0767. TMHorizon likewise
(16.199° / 0.0757). The corrector is currently adding no measurable aim
information on this corpus. That is the thing Phase 1 must change.
0.3.6 Ruler resolution is NOT the limiter [MEASURED]. Aiming at the
integer-tick solve instead of the exact intercept costs only 0.29–1.10° of mean
error (OracleQuant column). So the "maybe the oracle only reaches 35 % because
the solve is coarse" worry is dead: the coarse/fine difference is ≈0.3° at long
range, far below the target tolerance (1.93° at 450+). Whatever caps the score,
it is the enemy's unpredictability, not the ruler.
0.4 Speed [MEASURED]
Full 70-run / 899 607-tick / 3 598 428 tick-bin sweep, 10 arms:
406.9 s wall, i.e. 0.1131 ms per tick-bin over 10 arms,
≈ 0.045 s per gun per 1000 ticks (1000 ticks × 4 power bins).
BitBrain is the dominant cost (its ADE pass runs on every predict call);
Pattern/TMHorizon cache their per-tick work. A single-arm Pattern-only sweep is
several times cheaper. The binary cache (§0.1) is what makes repeat sweeps
affordable: without it every run re-parses 149 MB of JSONL.
0.5 How to reproduce [MEASURED]
nim c -d:release --nimcache:/tmp/nc_j98 -o:/tmp/bbq_run \
common_libs/tests/run_prediction_quality.nim
/tmp/bbq_run --corpus /tmp/tfil_ab2/out # full bar, ~7 min
/tmp/bbq_run --corpus /tmp/tfil_ab2/out --limit 10 # fast subset
/tmp/bbq_run --corpus /tmp/tfil_ab2/out --ruler quant # integer-tick solve
Verbatim full output: common_libs/tests/prediction_quality_results.txt.
Determinism: two consecutive full runs are byte-identical except the two
wall-time lines.
Clean-checkout proof [MEASURED]: git archive HEAD | tar -x -C /tmp/bbq_clean
then, from /tmp/bbq_clean,
nim c -d:release --nimcache:/tmp/nc_j98 -o:bbq_run common_libs/tests/run_prediction_quality.nim
builds, and ./bbq_run --corpus /tmp/tfil_ab2/out --limit 3 runs and prints the
same tables (separation 13.68× px on the 3-run subset). The committed harness is
self-contained; only the corpus is external.
0.6 Designs still to try (seed for later phases)
Ordered by expected value per unit of effort. Phase 0 has already killed one.
| # | design | why it is worth trying | status |
|---|---|---|---|
| D1 | Live A/B: HeadOn at 450+ vs Pattern (distance-gated switch, or HeadOn-only control) | Offline says a static gun beats Pattern at long range; the cheapest possible test of the campaign's central premise | TODO (highest value, live) |
| D2 | Pattern variants that raise lead CORRELATION at long range: longer keys / multi-length keys, per-distance learned pattern tables, different match weighting, k-NN over movement signatures | The lever is corr (0.165 at 450+), not gain; the ruler measures exactly this | TODO (offline-searchable) |
| D3 | Supervise BitBrain with the ruler's own labels — per-tick bearing error to the true intercept, trained on N−1 runs, evaluated on a held-out run | BitBrain's corrector currently adds nothing (0.3.5); the ruler gives it a real target. Must hold out runs or it is overfitting | TODO (offline-searchable) |
| D4 | A causal "predictability" gate: at each tick estimate whether the future is predictable (e.g. recent pattern-match score, reversal entropy) and fall back to HeadOn/low-variance aim when it is not | Directly attacks the 0.3.4 failure mode without needing a better long-range predictor | TODO |
| D5 | Power policy at long range (already partly done live): lower power = faster bullet = less lead error | Shortens the horizon the predictor must extrapolate; affects hit rate, live-only verdict | TODO (offline proxy only) |
| D6 | Lead-gain sweep 1.0/1.5/2.0/3.0 | DEAD — measured. Gain 1.0 wins at every band (§0.3.2) | KILLED |
Every D-item must end in a live A/B before any phase verdict.
0.7 What would make us quit
If (a) no causal design raises the 450+
hitProxyabove the static-gun reference (~0.10) on held-out runs by a margin larger than the run-to-run spread, and (b) the live A/B of the best such design shows no hit-rate or damage gain over Pattern with the left-running liveness check satisfied, then the campaign stops and we ship the simpler gun. We do not keep tuning an offline proxy that has stopped predicting live outcomes.
Phase 1: the missing gain sweep (owner: overnight job j100, committed)
Three negatives are on file (the morning reader must see these)
- BitBrain as previously shipped was statistically identical to Pattern
live. 30 runs/arm, MDE 24.4 dmg/run:
docs/bitbrain_gun_verdict.md, commitd93ce44. - Gun-mixing does not disrupt DrussGT. TMHorizon+BitBrain mixed into the
rack produced no dodge disruption:
docs/gun_mix_disruption.md, commit32a5e72. - BitBrain added no measurable aim information offline. 450+: 16.200 deg
vs Pattern 16.193, hitProxy 0.0767 vs 0.0767 (
a82c864, §0.3.5 above).
These bound the plausible upside: the previous BitBrain output — an ADDITIVE angular shift — was information-free, so Phase 1 changes the output shape, not the learning rate.
1.1 The missing sweep: Pattern x gain in [0.00, 1.00] [MEASURED]
common_libs/tests/run_prediction_quality.nim now carries sub-unity gain arms
(gain 0.0 is HeadOn, gain 1.0 is Pattern). Full 70-run sweep, 14 arms,
3 598 428 tick-bins, common_libs/tests/prediction_quality_results.txt.
Cells are mean|err| deg [hitProxy]; hitProxy is the objective (fraction of
tick-bins within atan(18/range), the ruler's proxy for hit probability).
| band | |req| deg | g=0.00 | g=0.25 | g=0.50 | g=0.75 | g=1.00 | best hpx | dHpx | best |err| | |---|---|---|---|---|---|---|---|---| | 0–100 | 19.62 | 19.62 [.342] | 16.98 [.391] | 14.58 [.469] | 12.36 [.640] | 10.56 [.699] | 1.00 | +.0000 | 1.00 | | 100–200 | 19.98 | 19.98 [.172] | 17.79 [.166] | 16.10 [.167] | 14.97 [.234] | 14.75 [.342] | 1.00 | +.0000 | 1.00 | | 200–300 | 17.34 | 17.34 [.133] | 15.95 [.128] | 15.32 [.137] | 15.53 [.139] | 16.61 [.185] | 1.00 | +.0000 | 0.50 | | 300–450 | 14.61 | 14.61 [.105] | 14.13 [.102] | 14.49 [.100] | 15.64 [.095] | 17.53 [.104] | 0.00 | +.0013 | 0.25 | | 450+ | 12.33 | 12.33 [.098] | 12.26 [.093] | 12.94 [.088] | 14.28 [.081] | 16.19 [.077] | 0.00 | +.0216 | 0.25 |
The optimal gain curve (hitProxy-argmax per band) is
[1.00, 1.00, 1.00, 0.00, 0.00] — use Pattern's full lead below 300 px and
aim at the enemy's current position (zero lead, HeadOn) at 300+ px. The
implied hit-probability gains vs Pattern are +0.00 pp (0–300), +0.13 pp
(300–450) and +2.16 pp at 450+ (0.0767 -> 0.0984, a 28 % relative rise).
Weighted by tick-bin share (450+ is 64.2 %, 300–450 is 31.1 %) the whole-corpus
proxy rises ~+1.4 pp.
[CAUSAL-SHIPPABILITY, important] The rule only needs the range, which is
known at fire time, so applying [1,1,1,0,0] is causal and needs no learning
at all — a per-band table is shippable as a constant, exactly like the
Pattern radial-offset knob. What is not causal is the estimation of the
table from the same runs (it is in-sample here); a shipped table would be fitted
offline on past battles or learned online, which is what BitBrain does. The
table's value is robust to that caveat because the winning entries are the two
extremes (full Pattern lead, or zero lead), not a knife-edge intermediate value.
1.2 Shrinking the gain does NOT add lead information [MEASURED]
| band | corr @ g=0.25 | g=0.50 | g=0.75 | g=1.00 |
|---|---|---|---|---|
| 0–100 | 0.774 | 0.774 | 0.774 | 0.774 |
| 100–200 | 0.612 | 0.612 | 0.612 | 0.612 |
| 200–300 | 0.457 | 0.457 | 0.457 | 0.457 |
| 300–450 | 0.266 | 0.266 | 0.266 | 0.266 |
| 450+ | 0.165 | 0.165 | 0.165 | 0.165 |
Every g > 0 column is identical: Pearson correlation is invariant under
positive scaling. A fractional gain therefore buys nothing on the
lead-information axis — it only shrinks the magnitude of an uninformative signal
toward the low-variance static aim. This confirms §0.3.3's reading and is the
mechanism behind the whole curve.
The MSE-optimal gain and the hitProxy-optimal gain diverge [MEASURED, key].
The best |err| column is the least-squares optimum [1.00, 1.00, 0.50, 0.25, 0.25]. At 200–300, MSE wants ~0.5 but hitProxy wants 1.0: Pattern's lead errors
are bimodal (it either nails the lead or is far off), so shrinking every
sample trades many small-within-tolerance hits for a smaller tail. Any learned
corrector that minimises squared error will therefore under-perform at mid range
— a fact this phase measured the hard way (an MSE-gain BitBrain dropped the
200–300 proxy from 0.190 to 0.131).
1.3 BitBrain rebuilt as a lead-gain corrector [MEASURED]
common_libs/guns/bitbrain_gun.nim is rewritten. The ADE+SBC network is gone
from the gun; the output is now a multiplicative gain on Pattern's lead,
aim = LOS + gain * (patternAim - LOS), with gain chosen from a fixed
candidate set {0, 0.25, 0.5, 0.75, 1.0}.
- Label path (unchanged): at fire time we remember Pattern's lead and the
target's angular half-width
atan(18/range);h = round(dist/speed)ticks latertmhObservedAtreturns the enemy's observed bearing from the firing position, givingrequiredLead = observedBearing - LOS. - Learning rule (changed): for every resolved sample we score each
candidate by whether
|gain*baseLead - requiredLead| <= atan(18/range)and keep the hit counts per range band; the band's gain is the argmax hit rate — the hit-probability proxy itself, not squared error. This directly fixes the bimodality failure above. - State / gate: the state is the range band (causally known). The correction
is applied only for range >= 300 px (
BB_GAIN_BAND_MIN = 3), below which Pattern's lead is informative. - Stats are battle-scale: a round boundary keeps the counts; a new battle or
target change (
resetLearning/targetChanged) wipes them.
Offline 70-run result (same ruler, same run as §1.1):
| band | Pattern hpx | fixed rule [1,1,1,0,0] | BitBrain hpx | fixed−Pat | BB−Pat |
|---|---|---|---|---|---|
| 0–100 | 0.6993 | 0.6993 | 0.6993 | +0.0000 | +0.0000 |
| 100–200 | 0.3418 | 0.3418 | 0.3418 | +0.0000 | +0.0000 |
| 200–300 | 0.1850 | 0.1850 | 0.1850 | +0.0000 | +0.0000 |
| 300–450 | 0.1036 | 0.1049 | 0.1073 | +0.0013 | +0.0037 |
| 450+ | 0.0767 | 0.0984 | 0.0957 | +0.0216 | +0.0190 |
BitBrain's effective point estimates match the fixed table to within 0.27 pp at 450+ and actually exceed it at 300–450. The internal log shows why it is not exactly equal: at long range the candidate hit rates are near-tied, so the argmax flips between 0.00 and 0.25 (450+) / 0.00 and 0.75 (300–450); the aggregate still lands on the right side. A fixed table is more stable; the learned version needs no table and adapts per battle.
Cost [MEASURED]. /tmp/bench_bb (200k ticks, 4 power bins) — Pattern
0.0224 ms/tick, BitBrain 0.0230 ms/tick, i.e. a marginal ~0.0007 ms/tick
(0.7 us/tick), versus the ~0.114 ms/tick of the old ADE+SBC gun and the
13.16 ms/tick budget. The whole 14-arm offline sweep now runs in 271 s
(0.0754 ms per tick-bin vs Phase 0's 0.1131 for 10 arms).
Guard tests stay green [MEASURED]. test_bitbrain 32 checks,
test_bitbrain_registration 13, test_rack_membership 48,
test_tm_pattern_registration 20, test_env_report all pass. env_report.nim
is unchanged: the legacy BitBrain knobs are still resolved and reported. The
ruler still validates (recorded hits 1.360 deg vs misses 16.597 deg,
separation 12.20x deg / 13.34x px; oracle 0.000000 deg).
1.4 Verdict
[MEASURED] Yes — a per-band lead-gain rule beats Pattern offline, but only in
the two long bands: [1,1,1,0,0] gives +0.13 pp at 300–450 and +2.16 pp at
450+ (a ~28 % relative rise in the hit-proxy, against Phase 0's realistic
causal ceiling of HeadOn's ~0.10, which it reaches). Below 300 px Pattern is
already optimal and the rule is a no-op.
[MEASURED] Yes — BitBrain-as-gain-corrector matches that rule (+1.90 pp at 450+, +0.37 pp at 300–450 over Pattern; within 0.27 pp of the fixed table at 450+, above it at 300–450), while being causally learnable online and ~163x cheaper per tick than the old gun.
[INFERRED / honest caveat] BitBrain is not required to capture the gain — a
constant [1,1,1,0,0] table would do it. BitBrain's value is that it discovers
the table per battle without one; its cost is cold-start (it begins at gain 1.0,
explaining the 0.27 pp gap at 450+) and the near-tie jitter in its point
estimates. So the honest ship decision is: the per-band rule is the thing to
test live, and it can be shipped either as a constant table or as this online
learner — the learner is redundant if a table is acceptable, and preferable
only if the optimum is expected to drift per enemy. The highest-value live
experiment is unchanged and stronger: a range-gated Pattern<->HeadOn switch at
~300 px vs pure Pattern (Phase 0's D1), left-running. No live win is claimed
here — the live gate is a separate phase.
1.5 Designs after Phase 1
| # | design | status |
|---|---|---|
| D1 | Live A/B: Pattern<->HeadOn switch at ~300 px vs Pattern | TODO (highest value, live) — offline now says +2.16 pp at 450+ |
| D2 | Raise Pattern lead correlation at long range (longer/multi-length keys, per-distance tables, k-NN) | TODO (offline-searchable) |
| D3 | Match-quality-conditional gain (keep full lead when the pattern match is good, shrink when bad) — attacks the bimodality that makes the MSE gain diverge | TODO (offline-searchable) |
| D4 | A causal predictability gate falling back to HeadOn when the future is unpredictable | TODO |
| D5 | Long-range power policy (already partly live) | TODO (offline proxy only) |
| D6 | Lead-gain sweep gains >= 1.0 | DEAD — measured (§0.3.2) |
| D7 | Constant sub-unity lead gain at all ranges | DEAD — measured (§1.1: gain < 1 hurts below 300) |