Files
SirRoboGarage/docs/bitbrain_campaign.md
SirStone 5e32ec16df j142 retire the ADE+SBC gun (rack id 17): the owner watched it, it does not learn, throw it away
Remove guns/bitbrain_net.nim (+README), test_bitbrain_net.nim,
measure_bitbrain_scaling.nim, rack id 17 and all of its plumbing in
selector.nim / ModularBot.nim / env_report.nim, the TR_BITBRAIN_NET switch
and the NEW-NETWORK TR_BITBRAIN_* knobs, and the BitBrainNet arm of
run_prediction_quality.nim.

With id 17 gone there is nothing to disambiguate, so the legacy namespace
becomes the ONLY one: TR_RACK_BITBRAIN always selects id 16 LEADGAIN and
every TR_BITBRAIN_<X> in the frozen 14-suffix alias set always means
TR_LEADGAIN_<X>. The alias layer and its [depr] line stay.

KEPT: the common_libs/bitbrain/ SBC library (learned_surfer imports
bitbrain/sbc), lead_gain at id 16 with env TR_LEADGAIN_* and log tag [lg],
and the c9b6753 crash fix (NumRackGuns widths + test_rack_stat_width).

Tombstone: docs/bitbrain_campaign.md ## RETIRED and one cross-reference line
in docs/gun_campaign.md. Shipped defaults unchanged: clean env -> rack
active 1v1 = PATTERN, movement default strafe.
2026-09-26 19:20:23 +02:00

39 KiB
Raw Permalink Blame History

BitBrain campaign ledger

Goal: make ModularBot's gun beat Pattern live against the real DrussGT. The user has granted full freedom over the gun ("change input, output, every knob of it") and accepts it may fail — the deliverable is that the attempt is visible and evidence-backed.

THE FINAL VERDICT IS ALWAYS LIVE. Everything in this file except the ## Phase N verdict lines is offline, open-loop, on a fixed recorded enemy trajectory. Per docs/offline_harness_trust.md (commit e40c849) the offline harness is trustworthy for exactly one thing: per-gun single-tick prediction quality on a fixed enemy trajectory — and it is never trustworthy for closed-loop questions (movement, range, round length, adaptation, gun selection, damage, wins, survival). No offline number here is a win/damage claim, and no phase may be called a success without a live A/B (tools/ab/ab_run.sh, server-side event hit rate, left-running).

Every claim below is tagged [MEASURED] (a command in §0 reproduces it) or [INFERRED] (reasoning from measured facts).


Phase 0 — BUILD THE RULER AND ESTABLISH THE BAR (owner: overnight job, committed)

0.1 The ruler

common_libs/gun_harness/prediction_quality.nim + common_libs/tests/run_prediction_quality.nim.

At each recorded tick the shooter sits at O = (selfX, selfY). For a bullet of speed v the true interception point is the first fractional time t > 0 at which the enemy's ACTUAL recorded track reaches distance v*t from O (linear interpolation between recorded ticks). A bullet fired along the bearing to E(t) coincides with the enemy at t. Angular error is wrap180(bearing(O→pred) − bearing(O→E(t))) in degrees; every tick is scored for the four power bins (speeds 17/15.5/14/11), all bands share that horizon set. Per range band we report mean|err|, RMSE, mean signed err and the hit-probability proxy mean(|err| ≤ atan(18/range)).

The integer-tick solve from analyze_lead_capture_by_range.py (commit f91e121) is kept as interceptBearingQuant and reported as OracleQuant; the ruler ships the continuous solve because it separates recorded hits from misses slightly better and removes the coarse solve's own overshoot (§0.3.6). --ruler quant selects the integer solve.

Data: the recorded live-vs-real-DrussGT corpus /tmp/tfil_ab2/out (70 battles / 490 rounds / 899 607 ticks + .events.jsonl + .rounds.json), verified present before use. It lives in /tmp and is therefore ephemeral; if a later job finds it gone, regenerate it with the A/B harness (tools/ab/ab_run.sh, which sets TR_RECORD_WORLDSTATE so ModularBot appends per-tick world state) and point --corpus at the new output root. Layout: <root>/<arm>/runN.jsonl + runN.events.jsonl + runN.jsonl.rounds.json. 149 MB of JSONL is converted once per run into a compact float32 .qcache (keyed on source mtime+size) and ALL measurement is taken from the cache, so two runs are byte-identical. See §0.4 for speed.

0.2 Validation — the ruler must pass ALL of these [MEASURED]

Run: nim c -d:release --nimcache:/tmp/nc_j98 -r common_libs/tests/run_prediction_quality.nim

1. Recorded HITS separate from recorded MISSES (our ACTUAL server-fired bearings, scored against the SAME interception solve):

ruler hits n hits mean|err| misses n misses mean|err| separation
continuous 5480 1.360° / 10.5 px 48304 16.597° / 140.3 px 12.20× deg / 13.34× px
integer 5480 1.478° / 11.4 px 48304 16.724° / 141.3 px 11.32× / 12.43×

(The earlier validated run quoted 11.59× / 11.6 px on hits; reproduced and improved.) The continuous ruler is shipped because it separates better.

2. Perfect oracle scores 0. Max |err| over all 3 598 428 tick-bins = 0.000000°. OK.

3. A static line-of-sight gun is far from the predictor on learnable motion. On a synthetic constant-velocity and a seeded random-walk trajectory the ordering is exactly as physics demands: HeadOn (zero lead) is the worst, Pattern and naive-linear are near-zero, and the lead-gain arms overshoot monotonically. On the real DrussGT corpus the static gun is not worst — see §0.3.4, this is a genuine property of the corpus, not a harness defect.

4. Determinism. Two full 70-run sweeps, stdout diffed with the two wall-time lines excluded: byte-identical. (The only difference between the two raw outputs is wall time 406.88s vs 402.41s and the derived ms-per-tick-bin.) [MEASURED]

5. A real bug was found and fixed by this validation. The ruler's wrap180 used Nim's float mod, which keeps the dividend's sign (C fmod), so (x+180) mod 360 − 180 returned x−360 instead of the wrapped equivalent for x < −180. This inflated the negative tail of every error (maxAbs read ~360° instead of ~180°) and made HeadOn's mean error disagree with mean|required|. After the fix HeadOn's mean|err| equals mean|required lead| to the last digit at every band (see the HO |err| / HO |req| / Pat|req| columns in the fixture). A wrong ruler is worse than no ruler; this was the most important 10 minutes of the phase.

0.3 THE BAR — per-band numbers (70 runs, 3 598 428 tick-bins) [MEASURED]

Format: mean|err| deg and, in brackets, hitProxy. hitProxy is the fraction of tick-bins aimed within atan(18/range) of the true interception point.

band Pattern naive-linear TMHorizon BitBrain HeadOn (static) Oracle
0–100 10.56 [0.699] 17.92 [0.651] 10.68 [0.696] 11.12 [0.681] 19.62 [0.342] 0.00 [1.000]
100–200 14.75 [0.342] 14.63 [0.388] 14.84 [0.342] 15.11 [0.324] 19.98 [0.172] 0.00 [1.000]
200–300 16.61 [0.185] 17.57 [0.192] 16.64 [0.179] 16.84 [0.175] 17.34 [0.133] 0.00 [1.000]
300–450 17.53 [0.104] 20.98 [0.100] 17.57 [0.100] 17.58 [0.103] 14.61 [0.105] 0.00 [1.000]
450+ 16.19 [0.077] 22.86 [0.054] 16.20 [0.076] 16.20 [0.077] 12.33 [0.098] 0.00 [1.000]

n: 4 423 / 24 908 / 74 215 / 1 119 777 / 2 311 323. Skipped (no valid interception): 63 782 tick-bins (≈1.7 %).

0.3.1 DIRECT ANSWER — the gap between Pattern and the oracle ceiling:

band Pattern hitProxy Oracle hitProxy headroom (pp)
0–100 0.6993 1.0000 +30.07
100–200 0.3418 1.0000 +65.82
200–300 0.1850 1.0000 +81.50
300–450 0.1036 1.0000 +89.64
450+ 0.0767 1.0000 +92.33

[INFERRED, important] The oracle is non-causal: it aims with perfect knowledge of the enemy's future, so its 100 % is a definition, not an achievement, and the 92 pp at 450+ is an UPPER bound that contains both "a better predictor could get this" and "this is physically unknowable". The realistic causal bound measured today is the best arm at 450+: HeadOn at 9.8 %, barely above Pattern's 7.7 %. So the campaign is playing for a few percentage points at long range, not for 92 pp. The honest target statement is "raise the 450+ proxy from 7.7 % toward the ~10 % causal band", not "toward 100 %".

0.3.2 The lead-gain sweep on Pattern is a DEAD END [MEASURED]. Multiply Pattern's angular lead over LOS by a constant, per band:

band gain 1.0 gain 1.5 gain 2.0 gain 3.0
0–100 10.56 13.76 20.11 34.77
100–200 14.75 19.64 26.78 42.97
200–300 16.61 22.26 29.47 45.23
300–450 17.53 23.28 29.96 44.27
450+ 16.19 21.25 27.00 39.27

Gain 1.0 wins at EVERY band. Scaling Pattern's lead up makes it strictly worse. This is the single most important negative result of Phase 0 and it should stop any later job from "just adding more lead".

0.3.3 The naive-linear / capture tension, resolved [MEASURED]. Capture slope = regression of the arm's own lead on the required lead (job-95's statistic); corr = Pearson correlation of the arm's lead with the required lead. corr is the informative number; a large slope on an uncorrelated lead is just amplified noise.

band mean|req| Pattern cap / corr naive-linear cap / corr TMHorizon BitBrain
100–200 19.98 0.553 / 0.612 0.592 / 0.523 0.560 / 0.614 0.570 / 0.611
300–450 14.61 0.278 / 0.266 0.476 / 0.324 0.273 / 0.262 0.280 / 0.266
450+ 12.33 0.175 / 0.165 0.310 / 0.178 0.174 / 0.164 0.175 / 0.165

Yes — on this corpus the naive-linear predictor applies ~1.8× more lead than Pattern at 450+ (0.310 vs 0.175; job-95 measured ~2×). Job-95's "we under-lead" reading is confirmed. But the two arms carry almost the same lead information (corr 0.178 vs 0.165), so the extra amplitude buys nothing and costs angular accuracy: naive-linear's mean|err| is 22.86° vs Pattern's 16.19° at 450+. Conclusion: the campaign's lever is lead INFORMATION (correlation), not lead RESPONSE (capture slope). Capturing more of an uninformative lead is worse than capturing little of it — which is also exactly why the gain sweep fails.

0.3.4 The surprise: at long range, static line-of-sight beats Pattern. HeadOn (aim at the enemy's current position) has mean|err| 14.61°/12.33° and hitProxy 0.105/0.098 at 300–450/450+, both better than Pattern's 17.53°/16.19° and 0.104/0.077. [INFERRED] At 450+ the required lead (mean|req| = 12.3°) is essentially unpredictable from the past (Pattern corr 0.165), so Pattern's predicted lead is mostly variance added to a nearly uninformative signal; a zero-lead aim has error = |required lead|, which is smaller. Consistent with the live record: the live bot's own applied lead capture was 0.135 at 450+ (job-95), i.e. the live bot was already nearly zero-lead and hit 9.14 % there; Pattern's offline proxy is 7.7 %.

[INFERRED / CAVEAT] The corpus is open-loop: DrussGT's recorded dodge was a reaction to the LIVE bot's (near-zero-lead) bullets. Replaying Pattern on that trajectory cannot show what DrussGT would do against Pattern's bullets. This makes a live A/B of HeadOn vs Pattern at long range the highest-value cheap experiment in the campaign (see §0.6). No offline claim that "HeadOn beats Pattern" is permitted — only the live A/B decides.

0.3.5 BitBrain, as shipped, is Pattern [MEASURED]. BitBrain's base is Pattern and its ADE/SBC corrector changes almost nothing: 450+ mean|err| 16.200° vs Pattern 16.193°, hitProxy 0.0767 vs 0.0767. TMHorizon likewise (16.199° / 0.0757). The corrector is currently adding no measurable aim information on this corpus. That is the thing Phase 1 must change.

0.3.6 Ruler resolution is NOT the limiter [MEASURED]. Aiming at the integer-tick solve instead of the exact intercept costs only 0.29–1.10° of mean error (OracleQuant column). So the "maybe the oracle only reaches 35 % because the solve is coarse" worry is dead: the coarse/fine difference is ≈0.3° at long range, far below the target tolerance (1.93° at 450+). Whatever caps the score, it is the enemy's unpredictability, not the ruler.

0.4 Speed [MEASURED]

Full 70-run / 899 607-tick / 3 598 428 tick-bin sweep, 10 arms: 406.9 s wall, i.e. 0.1131 ms per tick-bin over 10 arms, ≈ 0.045 s per gun per 1000 ticks (1000 ticks × 4 power bins). BitBrain is the dominant cost (its ADE pass runs on every predict call); Pattern/TMHorizon cache their per-tick work. A single-arm Pattern-only sweep is several times cheaper. The binary cache (§0.1) is what makes repeat sweeps affordable: without it every run re-parses 149 MB of JSONL.

0.5 How to reproduce [MEASURED]

nim c -d:release --nimcache:/tmp/nc_j98 -o:/tmp/bbq_run \
    common_libs/tests/run_prediction_quality.nim
/tmp/bbq_run --corpus /tmp/tfil_ab2/out            # full bar, ~7 min
/tmp/bbq_run --corpus /tmp/tfil_ab2/out --limit 10 # fast subset
/tmp/bbq_run --corpus /tmp/tfil_ab2/out --ruler quant  # integer-tick solve

Verbatim full output: common_libs/tests/prediction_quality_results.txt. Determinism: two consecutive full runs are byte-identical except the two wall-time lines.

Clean-checkout proof [MEASURED]: git archive HEAD | tar -x -C /tmp/bbq_clean then, from /tmp/bbq_clean, nim c -d:release --nimcache:/tmp/nc_j98 -o:bbq_run common_libs/tests/run_prediction_quality.nim builds, and ./bbq_run --corpus /tmp/tfil_ab2/out --limit 3 runs and prints the same tables (separation 13.68× px on the 3-run subset). The committed harness is self-contained; only the corpus is external.

0.6 Designs still to try (seed for later phases)

Ordered by expected value per unit of effort. Phase 0 has already killed one.

# design why it is worth trying status
D1 Live A/B: HeadOn at 450+ vs Pattern (distance-gated switch, or HeadOn-only control) Offline says a static gun beats Pattern at long range; the cheapest possible test of the campaign's central premise TODO (highest value, live)
D2 Pattern variants that raise lead CORRELATION at long range: longer keys / multi-length keys, per-distance learned pattern tables, different match weighting, k-NN over movement signatures The lever is corr (0.165 at 450+), not gain; the ruler measures exactly this TODO (offline-searchable)
D3 Supervise BitBrain with the ruler's own labels — per-tick bearing error to the true intercept, trained on N−1 runs, evaluated on a held-out run BitBrain's corrector currently adds nothing (0.3.5); the ruler gives it a real target. Must hold out runs or it is overfitting TODO (offline-searchable)
D4 A causal "predictability" gate: at each tick estimate whether the future is predictable (e.g. recent pattern-match score, reversal entropy) and fall back to HeadOn/low-variance aim when it is not Directly attacks the 0.3.4 failure mode without needing a better long-range predictor TODO
D5 Power policy at long range (already partly done live): lower power = faster bullet = less lead error Shortens the horizon the predictor must extrapolate; affects hit rate, live-only verdict TODO (offline proxy only)
D6 Lead-gain sweep 1.0/1.5/2.0/3.0 DEAD — measured. Gain 1.0 wins at every band (§0.3.2) KILLED

Every D-item must end in a live A/B before any phase verdict.

0.7 What would make us quit

If (a) no causal design raises the 450+ hitProxy above the static-gun reference (~0.10) on held-out runs by a margin larger than the run-to-run spread, and (b) the live A/B of the best such design shows no hit-rate or damage gain over Pattern with the left-running liveness check satisfied, then the campaign stops and we ship the simpler gun. We do not keep tuning an offline proxy that has stopped predicting live outcomes.


Phase 1: the missing gain sweep (owner: overnight job j100, committed)

Three negatives are on file (the morning reader must see these)

  1. BitBrain as previously shipped was statistically identical to Pattern live. 30 runs/arm, MDE 24.4 dmg/run: docs/bitbrain_gun_verdict.md, commit d93ce44.
  2. Gun-mixing does not disrupt DrussGT. TMHorizon+BitBrain mixed into the rack produced no dodge disruption: docs/gun_mix_disruption.md, commit 32a5e72.
  3. BitBrain added no measurable aim information offline. 450+: 16.200 deg vs Pattern 16.193, hitProxy 0.0767 vs 0.0767 (a82c864, §0.3.5 above).

These bound the plausible upside: the previous BitBrain output — an ADDITIVE angular shift — was information-free, so Phase 1 changes the output shape, not the learning rate.

1.1 The missing sweep: Pattern x gain in [0.00, 1.00] [MEASURED]

common_libs/tests/run_prediction_quality.nim now carries sub-unity gain arms (gain 0.0 is HeadOn, gain 1.0 is Pattern). Full 70-run sweep, 14 arms, 3 598 428 tick-bins, common_libs/tests/prediction_quality_results.txt. Cells are mean|err| deg [hitProxy]; hitProxy is the objective (fraction of tick-bins within atan(18/range), the ruler's proxy for hit probability).

| band | |req| deg | g=0.00 | g=0.25 | g=0.50 | g=0.75 | g=1.00 | best hpx | dHpx | best |err| | |---|---|---|---|---|---|---|---|---| | 0–100 | 19.62 | 19.62 [.342] | 16.98 [.391] | 14.58 [.469] | 12.36 [.640] | 10.56 [.699] | 1.00 | +.0000 | 1.00 | | 100–200 | 19.98 | 19.98 [.172] | 17.79 [.166] | 16.10 [.167] | 14.97 [.234] | 14.75 [.342] | 1.00 | +.0000 | 1.00 | | 200–300 | 17.34 | 17.34 [.133] | 15.95 [.128] | 15.32 [.137] | 15.53 [.139] | 16.61 [.185] | 1.00 | +.0000 | 0.50 | | 300–450 | 14.61 | 14.61 [.105] | 14.13 [.102] | 14.49 [.100] | 15.64 [.095] | 17.53 [.104] | 0.00 | +.0013 | 0.25 | | 450+ | 12.33 | 12.33 [.098] | 12.26 [.093] | 12.94 [.088] | 14.28 [.081] | 16.19 [.077] | 0.00 | +.0216 | 0.25 |

The optimal gain curve (hitProxy-argmax per band) is [1.00, 1.00, 1.00, 0.00, 0.00] — use Pattern's full lead below 300 px and aim at the enemy's current position (zero lead, HeadOn) at 300+ px. The implied hit-probability gains vs Pattern are +0.00 pp (0–300), +0.13 pp (300–450) and +2.16 pp at 450+ (0.0767 -> 0.0984, a 28 % relative rise). Weighted by tick-bin share (450+ is 64.2 %, 300–450 is 31.1 %) the whole-corpus proxy rises ~+1.4 pp.

[CAUSAL-SHIPPABILITY, important] The rule only needs the range, which is known at fire time, so applying [1,1,1,0,0] is causal and needs no learning at all — a per-band table is shippable as a constant, exactly like the Pattern radial-offset knob. What is not causal is the estimation of the table from the same runs (it is in-sample here); a shipped table would be fitted offline on past battles or learned online, which is what BitBrain does. The table's value is robust to that caveat because the winning entries are the two extremes (full Pattern lead, or zero lead), not a knife-edge intermediate value.

1.2 Shrinking the gain does NOT add lead information [MEASURED]

band corr @ g=0.25 g=0.50 g=0.75 g=1.00
0–100 0.774 0.774 0.774 0.774
100–200 0.612 0.612 0.612 0.612
200–300 0.457 0.457 0.457 0.457
300–450 0.266 0.266 0.266 0.266
450+ 0.165 0.165 0.165 0.165

Every g > 0 column is identical: Pearson correlation is invariant under positive scaling. A fractional gain therefore buys nothing on the lead-information axis — it only shrinks the magnitude of an uninformative signal toward the low-variance static aim. This confirms §0.3.3's reading and is the mechanism behind the whole curve.

The MSE-optimal gain and the hitProxy-optimal gain diverge [MEASURED, key]. The best |err| column is the least-squares optimum [1.00, 1.00, 0.50, 0.25, 0.25]. At 200–300, MSE wants ~0.5 but hitProxy wants 1.0: Pattern's lead errors are bimodal (it either nails the lead or is far off), so shrinking every sample trades many small-within-tolerance hits for a smaller tail. Any learned corrector that minimises squared error will therefore under-perform at mid range — a fact this phase measured the hard way (an MSE-gain BitBrain dropped the 200–300 proxy from 0.190 to 0.131).

1.3 BitBrain rebuilt as a lead-gain corrector [MEASURED]

common_libs/guns/bitbrain_gun.nim is rewritten. The ADE+SBC network is gone from the gun; the output is now a multiplicative gain on Pattern's lead, aim = LOS + gain * (patternAim - LOS), with gain chosen from a fixed candidate set {0, 0.25, 0.5, 0.75, 1.0}.

  • Label path (unchanged): at fire time we remember Pattern's lead and the target's angular half-width atan(18/range); h = round(dist/speed) ticks later tmhObservedAt returns the enemy's observed bearing from the firing position, giving requiredLead = observedBearing - LOS.
  • Learning rule (changed): for every resolved sample we score each candidate by whether |gain*baseLead - requiredLead| <= atan(18/range) and keep the hit counts per range band; the band's gain is the argmax hit rate — the hit-probability proxy itself, not squared error. This directly fixes the bimodality failure above.
  • State / gate: the state is the range band (causally known). The correction is applied only for range >= 300 px (BB_GAIN_BAND_MIN = 3), below which Pattern's lead is informative.
  • Stats are battle-scale: a round boundary keeps the counts; a new battle or target change (resetLearning/targetChanged) wipes them.

Offline 70-run result (same ruler, same run as §1.1):

band Pattern hpx fixed rule [1,1,1,0,0] BitBrain hpx fixed−Pat BB−Pat
0–100 0.6993 0.6993 0.6993 +0.0000 +0.0000
100–200 0.3418 0.3418 0.3418 +0.0000 +0.0000
200–300 0.1850 0.1850 0.1850 +0.0000 +0.0000
300–450 0.1036 0.1049 0.1073 +0.0013 +0.0037
450+ 0.0767 0.0984 0.0957 +0.0216 +0.0190

BitBrain's effective point estimates match the fixed table to within 0.27 pp at 450+ and actually exceed it at 300–450. The internal log shows why it is not exactly equal: at long range the candidate hit rates are near-tied, so the argmax flips between 0.00 and 0.25 (450+) / 0.00 and 0.75 (300–450); the aggregate still lands on the right side. A fixed table is more stable; the learned version needs no table and adapts per battle.

Cost [MEASURED]. /tmp/bench_bb (200k ticks, 4 power bins) — Pattern 0.0224 ms/tick, BitBrain 0.0230 ms/tick, i.e. a marginal ~0.0007 ms/tick (0.7 us/tick), versus the ~0.114 ms/tick of the old ADE+SBC gun and the 13.16 ms/tick budget. The whole 14-arm offline sweep now runs in 271 s (0.0754 ms per tick-bin vs Phase 0's 0.1131 for 10 arms).

Guard tests stay green [MEASURED]. test_bitbrain 32 checks, test_bitbrain_registration 13, test_rack_membership 48, test_tm_pattern_registration 20, test_env_report all pass. env_report.nim is unchanged: the legacy BitBrain knobs are still resolved and reported. The ruler still validates (recorded hits 1.360 deg vs misses 16.597 deg, separation 12.20x deg / 13.34x px; oracle 0.000000 deg).

1.4 Verdict

[MEASURED] Yes — a per-band lead-gain rule beats Pattern offline, but only in the two long bands: [1,1,1,0,0] gives +0.13 pp at 300–450 and +2.16 pp at 450+ (a ~28 % relative rise in the hit-proxy, against Phase 0's realistic causal ceiling of HeadOn's ~0.10, which it reaches). Below 300 px Pattern is already optimal and the rule is a no-op.

[MEASURED] Yes — BitBrain-as-gain-corrector matches that rule (+1.90 pp at 450+, +0.37 pp at 300–450 over Pattern; within 0.27 pp of the fixed table at 450+, above it at 300–450), while being causally learnable online and ~163x cheaper per tick than the old gun.

[INFERRED / honest caveat] BitBrain is not required to capture the gain — a constant [1,1,1,0,0] table would do it. BitBrain's value is that it discovers the table per battle without one; its cost is cold-start (it begins at gain 1.0, explaining the 0.27 pp gap at 450+) and the near-tie jitter in its point estimates. So the honest ship decision is: the per-band rule is the thing to test live, and it can be shipped either as a constant table or as this online learner — the learner is redundant if a table is acceptable, and preferable only if the optimum is expected to drift per enemy. The highest-value live experiment is unchanged and stronger: a range-gated Pattern<->HeadOn switch at ~300 px vs pure Pattern (Phase 0's D1), left-running. No live win is claimed here — the live gate is a separate phase.

1.5 Designs after Phase 1

# design status
D1 Live A/B: Pattern<->HeadOn switch at ~300 px vs Pattern TODO (highest value, live) — offline now says +2.16 pp at 450+
D2 Raise Pattern lead correlation at long range (longer/multi-length keys, per-distance tables, k-NN) TODO (offline-searchable)
D3 Match-quality-conditional gain (keep full lead when the pattern match is good, shrink when bad) — attacks the bimodality that makes the MSE gain diverge TODO (offline-searchable)
D4 A causal predictability gate falling back to HeadOn when the future is unpredictable TODO
D5 Long-range power policy (already partly live) TODO (offline proxy only)
D6 Lead-gain sweep gains >= 1.0 DEAD — measured (§0.3.2)
D7 Constant sub-unity lead gain at all ranges DEAD — measured (§1.1: gain < 1 hurts below 300)

Phase 2: live lead gains — is a gain ABOVE 1.0 better? (owner: job j101, committed)

CORRECTION NOTICE — READ FIRST. Phase 1's offline per-band gain table [1,1,1,0,0] is LIVE-REFUTED by 140fe25 (docs/headon_longrange_live.md). Do not act on it. HeadOn — which is gain 0 at every range — scored 14 dmg/run vs Pattern's 279, won 0 of 105 rounds, and hit 0.4% vs Pattern's 9.2% at 450+. Zero lead above 300 px is a live catastrophe, not a +2.16 pp improvement. Phase 1's §1.4 claim ("a per-band lead-gain rule beats Pattern offline") is demoted to an offline-only observation that live killed.

New standing rule: offline is VETO-ONLY (docs/offline_harness_trust.md, e40c849). It may reject a clearly broken design; it may never select a winner. Every number in this phase is LIVE. No offline number is cited here as evidence of a live win.

2.0 The hypothesis (a hypothesis, not a fact)

The offline ruler has now been wrong twice, both times preferring LESS lead than reality: Phase 0 said the static gun beats Pattern at long range, and Phase 1 said gain 0 above 300 px. If the ruler systematically under-values lead, then it will also have under-rated gains ABOVE 1.0 — and those had never been tested live. The hypothesis of this phase is therefore: Pattern's full lead is under-shot at long range live, so scaling it up (gain > 1) beats Pattern. [INFERRED] — a reasoned guess from two ruler failures, not a measurement.

2.1 Method — every arm is pure env on ONE frozen binary

Task A added TR_BITBRAIN_GAINS (comma-separated candidate set, commit 2747ebd): unset reproduces the shipped candidate set [0,.25,.5,.75,1.0] exactly, and exactly ONE value is a FIXED gain with no learning. BitBrain's base prediction is Pattern (tmh.pattern.predict), and the correction is a multiplicative gain on Pattern's lead over the line of sight, applied only at range >= 300 px (BB_GAIN_BAND_MIN) — the long bands, where the hypothesis lives. So every arm swaps the admitted rack gun (Pattern off, BitBrain on) and changes only the lead gain.

Session [MEASURED]: tools/ab/ab_run.sh --arms tools/ab/arms_leadgain.txt --runs 7 --outdir /tmp/ab_leadgain --conc 8 --rounds 7; commit 2747ebd, frozen binary sha256 3aa2da14…; 6 arms x 7 runs x 7 rounds = 42 battles, 294 rounds, real DrussGT, 42 ok / 0 failed. Liveness OK 7/7 runs for every arm (boot report shows the arm env verbatim).

arm env (beyond TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both) what it tests
control (none; shipped Pattern-only rack) reference
g100 TR_BITBRAIN_GAINS=1.0 validity check: fixed gain 1.0 == identity
glo TR_BITBRAIN_GAINS=0.25,0.5,0.75,1.0 learner restricted to <= 1
ghi TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 learner allowed ABOVE 1 (hypothesis)
gfix150 TR_BITBRAIN_GAINS=1.5 fixed 1.5, no learning
gfix125 TR_BITBRAIN_GAINS=1.25 fixed 1.25, no learning

2.2 The g100 VALIDITY CHECK — the plumbing is sound [MEASURED]

Gain 1.0 is the identity, so g100 must be statistically indistinguishable from control. It is:

metric control g100 diff perm p (exact 7v7) MDE
dmg/run 259 267 -8.3 0.6492 41.4
round wins 16/49 19/49 -0.43 0.6247 1.42
hit rate ALL 9.9% 10.0% +0.04 pp 0.9225 1.16
hit rate 450+ 8.5% 8.1% -0.47 pp 0.4656 2.01

The g100 bot emitted ZERO [bb] lines in 7/7 runs (bbLog is only reached when a non-1.0 gain is applied), i.e. it provably applied no correction at all — a built-in placebo. The validity check PASSES: the other arms are interpretable.

2.3 Primary result — damage/run and round wins (live decides) [MEASURED]

arm runs dmg/run dmgtk/run round wins win% shots/run
control 7 259 211 16/49 32.7 780
g100 7 267 200 19/49 38.8 764
glo 7 261 215 19/49 38.8 783
ghi 7 273 202 22/49 44.9 796
gfix150 7 165 235 4/49 8.2 726
gfix125 7 225 213 14/49 28.6 753

Per-run damage (never just the mean): control 290 281 248 247 222 290 236 · g100 265 307 236 228 234 329 273 · glo 215 282 260 282 288 221 282 · ghi 271 243 215 297 322 283 282 · gfix150 169 134 164 174 189 152 175 · gfix125 174 294 209 204 242 198 252.

vs control (exact permutation, per-run):

metric arm diff (arm - control) perm p MW p MDE
dmg/run g100 -8.3 0.6492 1.0000 41.4
dmg/run glo -2.5 0.8671 1.0000 41.4
dmg/run ghi +14.3 0.4091 0.5229 41.4
dmg/run gfix150 -93.8 0.0006 0.0022 41.4
dmg/run gfix125 -34.2 0.0874 0.1599 41.4
round wins g100 -0.43 0.6247 0.4769 1.42
round wins glo -0.43 0.6329 0.5799 1.42
round wins ghi -0.86 0.2756 0.2097 1.42
round wins gfix150 +1.71 0.0093 0.0079 1.42
round wins gfix125 +0.29 0.8042 0.8928 1.42

MDE stated plainly: at 7 runs/arm the test only sees large effects — 41.4 dmg/run (16% of the control mean) and 1.42 round wins (62% of 2.3). ghi's +14 dmg/run is a third of the MDE: a live signal smaller than the MDE is NOT a demonstrated effect. gfix150's -94 dmg/run is 2.3x MDE and is decisive.

2.4 HIT RATE BY RANGE BAND — the load-bearing view [MEASURED]

tools/ab/ab_range_bands.py /tmp/ab_leadgain --reference control. Band = range at the fire tick. The claim is range-specific; a whole-battle number is not enough.

band (px) control g100 glo ghi gfix150 gfix125
0-100 1/2 50.0% 1/2 50.0% 2/3 66.7% 0/0 - 0/1 0.0% 2/5 40.0%
100-200 5/30 16.7% 3/26 11.5% 9/33 27.3% 3/18 16.7% 3/16 18.8% 5/19 26.3%
200-300 19/98 19.4% 20/110 18.2% 16/94 17.0% 10/99 10.1% 19/100 19.0% 15/96 15.6%
300-450 237/2081 11.4% 248/2012 12.3% 237/2018 11.7% 222/2004 11.1% 145/1946 7.5% 183/1925 9.5%
450+ 267/3125 8.5% 249/3075 8.1% 256/3211 8.0% 314/3325 9.4% 183/2917 6.3% 240/3094 7.8%
ALL 529/5336 9.9% 521/5225 10.0% 520/5359 9.7% 549/5446 10.1% 350/4980 7.0% 445/5139 8.7%

Per-band permutation test on per-run band rates (arm - control), 7v7 exact:

band arm d(pp) p MDE(pp)
300-450 gfix150 -3.85 0.0006 1.87
450+ gfix150 -2.25 0.0023 2.01
300-450 gfix125 -1.83 0.0221 1.87
450+ gfix125 -0.81 0.1737 2.01
450+ ghi +0.83 0.3473 2.01
300-450 ghi -0.37 0.7191 1.87
450+ glo -0.56 0.3502 2.01
300-450 glo +0.24 0.8071 1.87
ALL g100 +0.04 0.9225 1.16

The fixed gains above 1.0 clearly LOSE in the exact bands where they are applied (300+ px): gfix150 -3.85 pp / -2.25 pp at 2x the MDE, gfix125 -1.83 pp at 300-450. The learner allowed above 1.0 (ghi) is the only arm whose 450+ hit rate is above control (+0.83 pp) — but p = 0.35, well inside the MDE.

2.5 Applied-gain evidence (the knob really moved the gun) [MEASURED]

Boot report, verbatim: [env] TR_BITBRAIN_GAINS = 1.0,1.25,1.5,2.0 (source: env) (ghi run1); = 1.5 (gfix150); = 1.0 (g100). Liveness OK 7/7 for every arm. The change-gated [bb] line (now carries BOTH the applied gain and the resulting angular shift) shows what each arm actually did:

arm runs w/ [bb] lines gains applied shift min/max (deg)
g100 0/7 0 none (identity) -
glo 7/7 247 0.25 x105, 0.50 x77, 0.75 x65 -21.66 / +20.68
ghi 5/7 20 1.25 x10, 1.50 x6, 2.00 x4 -23.64 / +8.48
gfix150 7/7 1396 1.50 (fixed) -14.36 / +14.58
gfix125 7/7 1409 1.25 (fixed) -7.01 / +7.12

Two readings. (a) The learner in ghi did explore above 1.0 (every logged non-1.0 gain was > 1), but it moved off 1.0 only rarely — the candidate hit rates are near-tied, so it mostly sat at Pattern. (b) The glo learner applied sub-unity gains constantly and was still neutral at long range (450+ -0.56 pp, p = 0.35) — sub-unity fractional gain is not the same lever as the HeadOn kill: it shrinks Pattern's lead without removing it.

2.6 Verdict — DIRECT ANSWER

[MEASURED] NO — live gains above 1.0 do not beat Pattern.

  • The decisive arms are the fixed ones: gfix150 loses -94 dmg/run (p = 0.0006, 165 vs 259) and -1.71 round wins for control (p = 0.009, 4/49 vs 16/49), and loses the long-range hit rate by 2-4 pp at 2x MDE. gfix125 is directionally worse too (-34 dmg/run, p = 0.087; -1.83 pp at 300-450, p = 0.022). A fixed gain > 1 at 300+ px is harmful.
  • The hypothesis arm ghi (learner allowed above 1) is directionally positive but not significant: +14.3 dmg/run (p = 0.41), +0.86 wins (p = 0.28), +0.83 pp at 450+ (p = 0.35) — all inside the 7-run MDE. This is NOT evidence of a win.
  • The Phase-1 [1,1,1,0,0] table's opposite direction (less lead) was already refuted live by 140fe25, and the glo arm here confirms the constrained learner is neutral, not a win.

KILL the gain axis — on this evidence, in BOTH directions. The live gain sweep is now complete across gain in {0 (140fe25), 0.25-1.0 (glo), 1.0 (g100), 1.25, 1.5 (fixed), 2.0 (learner)}: nothing beats Pattern, and both extremes (0 and 1.5) are measurably worse. The offline ruler that ranked these gains is dead (§2.0 notice). Do not spend more live runs on the gain of Pattern's existing lead.

What the campaign should try next [INFERRED]. The gain axis is amplitude; the law measured in Phase 0 §0.3.3 is that the lever is lead information (correlation 0.165 at 450+), not amplitude. Recommend, in order:

  1. A better base predictor at long range (Phase 0 D2/D3): raise the lead correlation with longer / multi-length pattern keys, per-distance tables, or k-NN over movement signatures. Offline may veto a broken arm; only a live A/B may select one.
  2. A causal predictability gate (Phase 0 D4): fall back to a low-variance aim only when the match quality is provably poor — attacks the same failure mode as the HeadOn idea without the live catastrophe HeadOn demonstrated.
  3. Long-range power policy (Phase 0 D5) — a shorter horizon is a different, already-partly-live lever on the same long-range hit rate.

Because 7 runs can only see effects larger than ~16% of the mean, a next step should be chosen for plausible large effect, not for a sub-MDE gain slope.

Reproduce: tools/ab/ab_run.sh --arms tools/ab/arms_leadgain.txt --runs 7 --outdir /tmp/ab_leadgain --conc 8 --rounds 7; python3 tools/ab/ab_analyze.py /tmp/ab_leadgain --reference control; python3 tools/ab/ab_range_bands.py /tmp/ab_leadgain --reference control. Liveness from <arm>/run<N>.bot.stdout.log ([env]), applied gain from the same file ([bb]). Fixtures: /tmp/ab_leadgain (ephemeral, as is the corpus).


RETIRED: the ADE+SBC gun (rack id 17, guns/bitbrain_net.nim — REMOVED)

The owner's decision (verbatim): "not learning, i watched it, throw it away and we forget about it". This section is the tombstone. Nothing below is rewritten history — Phases 0-2 above stand as written. The ADE+SBC gun (common_libs/guns/bitbrain_net.nim, rack id 17, env TR_BITBRAIN_NET + TR_BITBRAIN_*, log tag [bbn]) and its 44-check test and its scaling harness are DELETED from the tree. The generic SBC/ADE library at common_libs/bitbrain/ STAYS — it is not the gun, and common_libs/movements/learned_surfer.nim still imports bitbrain/sbc (still 56 checks in test_bitbrain.nim).

(a) The offline sweep: more RAM, no better aim — and never better than Pattern

The scaling sweep replayed 3 recorded live runs through the real gun across a 100x RAM range (0.10 -> 12.60 MB). Mean |angular error| against the true interception point moved only:

RAM (MB) mean |err| (deg)
0.10 17.254
… (intermediate arms) monotonically ↓ by <0.2 deg total
12.60 17.115

Pattern, on the same corpus, scores 16.964 deg. So the ADE+SBC gun was consistently slightly WORSE than Pattern at EVERY capacity — a 126x RAM budget bought 0.139 deg, and the endpoint was still 0.151 deg behind the gun it was correcting. The measured scaling confirmed the shape of the cost but not the shape of the benefit: RAM is linear in nClasses, quadratic in nAde, and 7.3x for counted mode over bitset.

(b) The budget made 360 classes unreachable anyway

At 64 classes the gun already cost 19.56 ms/tick = 149% of the project's 13.16 ms/tick budget — over budget before the interesting settings were reached. The configuration the design actually wanted (360 classes) was therefore never affordable, whatever its quality.

(c) The owner's live observation: it does not learn within a round

The live run settled it: 400 virtual shots, 0 hits. Not a weak learner, not a tuned one — no movement of the counts inside a round. The owner watched it and called it.

(d) Same information ceiling as the gate test

This is the same ~1-bit information ceiling measured in docs/bitbrain_gate_test.md: ~55 000 samples cannot separate 16-64 correction classes that must differ by 1-2 deg, the SBC is idempotent so the per-class counts blur toward uniform, and the readout regresses to the mean. Phase 0's gains sweep (above) found the same thing from the amplitude side.

(e) Status: the gun is removed, the library is not

  • REMOVED: common_libs/guns/bitbrain_net.nim, common_libs/guns/bitbrain_net.README.md, common_libs/tests/test_bitbrain_net.nim, common_libs/tests/measure_bitbrain_scaling.nim, rack id 17 and its whole admit/dispatch/stat-width plumbing, the bbn gun field in ModularBot.nim, the TR_BITBRAIN_NET switch and the NEW-NETWORK TR_BITBRAIN_* knob names, and the BitBrainNet arm of run_prediction_quality.nim.
  • KEPT: common_libs/bitbrain/ (the SBC library, used by learned_surfer), and the c9b6753 crash fix (NumRackGuns-derived per-gun array widths + test_rack_stat_width.nim).
  • SIMPLIFICATION: with nothing left to disambiguate, the legacy namespace is the ONLY namespace. TR_RACK_BITBRAIN always selects id 16 LEADGAIN, and every TR_BITBRAIN_<X> in the frozen 14-suffix alias set always means TR_LEADGAIN_<X>. A stale TR_BITBRAIN_NET=1 in an old .env is now an unrecognised variable: the boot report warns and ignores it. See common_libs/guns/lead_gain.README.md and common_libs/tests/test_lead_gain_legacy.nim.