Files
SirRoboGarage/docs/bitbrain_campaign.md
T
SirStone a82c864c60 bitbrain campaign phase 0: offline prediction-quality ruler and the bar
New harness (common_libs/gun_harness/prediction_quality.nim +
common_libs/tests/run_prediction_quality.nim): per-gun single-tick aim error in
degrees against the true continuous interception point on the recorded
live-vs-real-DrussGT corpus (/tmp/tfil_ab2/out, 70 runs, 899607 ticks), per
range band, with the hit-probability proxy mean(|err|<=atan(18/range)).
Validated: recorded hits separate from misses 13.34x px (reference 11.59x),
perfect-oracle max |err| = 0, correct ordering on synthetic ground truth, two
full runs byte-identical. Fixed a wrap180 bug (Nim float mod keeps the dividend
sign) that inflated the negative error tail.

Bar (mean|err| deg [hitProxy] at 450+): Pattern 16.19 [0.077], naive-linear
22.86 [0.054], TMHorizon 16.20 [0.076], BitBrain 16.20 [0.077], static HeadOn
12.33 [0.098], oracle 0 [1.0]. Lead-gain sweep on Pattern is a dead end (1.0
wins every band). Naive-linear applies ~1.8x Pattern's lead but carries no more
lead information (corr 0.178 vs 0.165) and is strictly worse. Ledger:
docs/bitbrain_campaign.md. All verdicts remain live-only.
2026-09-24 23:40:56 +02:00

259 lines
15 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# BitBrain campaign ledger
**Goal:** make ModularBot's gun **beat Pattern live** against the real DrussGT.
The user has granted full freedom over the gun ("change input, output, every
knob of it") and accepts it may fail — the deliverable is that the attempt is
visible and evidence-backed.
**THE FINAL VERDICT IS ALWAYS LIVE.** Everything in this file except the
`## Phase N` verdict lines is offline, open-loop, on a *fixed recorded enemy
trajectory*. Per `docs/offline_harness_trust.md` (commit `e40c849`) the offline
harness is trustworthy for exactly one thing: **per-gun single-tick prediction
quality on a fixed enemy trajectory** — and it is *never* trustworthy for
closed-loop questions (movement, range, round length, adaptation, gun
selection, damage, wins, survival). No offline number here is a win/damage
claim, and no phase may be called a success without a live A/B
(`tools/ab/ab_run.sh`, server-side event hit rate, left-running).
Every claim below is tagged **[MEASURED]** (a command in §0 reproduces it) or
**[INFERRED]** (reasoning from measured facts).
---
## Phase 0 — BUILD THE RULER AND ESTABLISH THE BAR *(owner: overnight job, committed)*
### 0.1 The ruler
`common_libs/gun_harness/prediction_quality.nim` + `common_libs/tests/run_prediction_quality.nim`.
At each recorded tick the shooter sits at `O = (selfX, selfY)`. For a bullet of
speed `v` the **true interception point** is the first fractional time `t > 0`
at which the enemy's ACTUAL recorded track reaches distance `v*t` from `O`
(linear interpolation between recorded ticks). A bullet fired along the bearing
to `E(t)` coincides with the enemy at `t`. Angular error is
`wrap180(bearing(O→pred) − bearing(O→E(t)))` in **degrees**; every tick is
scored for the four power bins (speeds 17/15.5/14/11), all bands share that
horizon set. Per range band we report `mean|err|`, RMSE, mean signed err and the
hit-probability proxy `mean(|err| ≤ atan(18/range))`.
The integer-tick solve from `analyze_lead_capture_by_range.py` (commit
`f91e121`) is kept as `interceptBearingQuant` and reported as `OracleQuant`; the
ruler ships the **continuous** solve because it separates recorded hits from
misses slightly better and removes the coarse solve's own overshoot
(§0.3.6). `--ruler quant` selects the integer solve.
**Data:** the recorded live-vs-real-DrussGT corpus `/tmp/tfil_ab2/out`
(70 battles / 490 rounds / 899 607 ticks + `.events.jsonl` + `.rounds.json`),
**verified present before use**. It lives in `/tmp` and is therefore ephemeral;
if a later job finds it gone, regenerate it with the A/B harness
(`tools/ab/ab_run.sh`, which sets `TR_RECORD_WORLDSTATE` so ModularBot appends
per-tick world state) and point `--corpus` at the new output root. Layout:
`<root>/<arm>/runN.jsonl` + `runN.events.jsonl` + `runN.jsonl.rounds.json`.
149 MB of JSONL is converted once per run into a compact float32 `.qcache`
(keyed on source mtime+size) and ALL measurement is taken from the cache, so two
runs are byte-identical. See §0.4 for speed.
### 0.2 Validation — the ruler must pass ALL of these **[MEASURED]**
Run: `nim c -d:release --nimcache:/tmp/nc_j98 -r common_libs/tests/run_prediction_quality.nim`
**1. Recorded HITS separate from recorded MISSES** (our ACTUAL server-fired
bearings, scored against the SAME interception solve):
| ruler | hits n | hits mean\|err\| | misses n | misses mean\|err\| | separation |
|---|---|---|---|---|---|
| continuous | 5480 | **1.360° / 10.5 px** | 48304 | **16.597° / 140.3 px** | **12.20× deg / 13.34× px** |
| integer | 5480 | 1.478° / 11.4 px | 48304 | 16.724° / 141.3 px | 11.32× / 12.43× |
(The earlier validated run quoted 11.59× / 11.6 px on hits; reproduced and
improved.) The continuous ruler is shipped because it separates better.
**2. Perfect oracle scores 0.** Max `|err|` over all 3 598 428 tick-bins =
**0.000000°**. OK.
**3. A static line-of-sight gun is far from the predictor on learnable motion.**
On a synthetic constant-velocity and a seeded random-walk trajectory the
ordering is exactly as physics demands: HeadOn (zero lead) is the worst, Pattern
and naive-linear are near-zero, and the lead-gain arms overshoot monotonically.
On the real DrussGT corpus the static gun is *not* worst — see §0.3.4, this is a
genuine property of the corpus, not a harness defect.
**4. Determinism.** Two full 70-run sweeps, stdout diffed with the two wall-time
lines excluded: **byte-identical**. (The only difference between the two raw
outputs is `wall time 406.88s` vs `402.41s` and the derived ms-per-tick-bin.)
**[MEASURED]**
**5. A real bug was found and fixed by this validation.** The ruler's
`wrap180` used Nim's float `mod`, which keeps the dividend's sign (C `fmod`), so
`(x+180) mod 360 − 180` returned `x−360` instead of the wrapped equivalent for
`x < −180`. This inflated the negative tail of every error (maxAbs read ~360°
instead of ~180°) and made HeadOn's mean error disagree with `mean|required|`.
After the fix HeadOn's `mean|err|` equals `mean|required lead|` to the last
digit at every band (see the `HO |err| / HO |req| / Pat|req|` columns in the
fixture). **A wrong ruler is worse than no ruler; this was the most important
10 minutes of the phase.**
### 0.3 THE BAR — per-band numbers (70 runs, 3 598 428 tick-bins) **[MEASURED]**
Format: `mean|err| deg` and, in brackets, `hitProxy`. `hitProxy` is the fraction
of tick-bins aimed within `atan(18/range)` of the true interception point.
| band | Pattern | naive-linear | TMHorizon | BitBrain | HeadOn (static) | Oracle |
|---|---|---|---|---|---|---|
| 0–100 | **10.56** [0.699] | 17.92 [0.651] | 10.68 [0.696] | 11.12 [0.681] | 19.62 [0.342] | 0.00 [1.000] |
| 100–200 | **14.75** [0.342] | 14.63 [0.388] | 14.84 [0.342] | 15.11 [0.324] | 19.98 [0.172] | 0.00 [1.000] |
| 200–300 | **16.61** [0.185] | 17.57 [0.192] | 16.64 [0.179] | 16.84 [0.175] | 17.34 [0.133] | 0.00 [1.000] |
| 300–450 | 17.53 [0.104] | 20.98 [0.100] | 17.57 [0.100] | 17.58 [0.103] | **14.61** [0.105] | 0.00 [1.000] |
| 450+ | 16.19 [0.077] | 22.86 [0.054] | 16.20 [0.076] | 16.20 [0.077] | **12.33** [0.098] | 0.00 [1.000] |
`n`: 4 423 / 24 908 / 74 215 / 1 119 777 / 2 311 323. Skipped (no valid
interception): 63 782 tick-bins (≈1.7 %).
**0.3.1 DIRECT ANSWER — the gap between Pattern and the oracle ceiling:**
| band | Pattern hitProxy | Oracle hitProxy | headroom (pp) |
|---|---|---|---|
| 0–100 | 0.6993 | 1.0000 | **+30.07** |
| 100–200 | 0.3418 | 1.0000 | **+65.82** |
| 200–300 | 0.1850 | 1.0000 | **+81.50** |
| 300–450 | 0.1036 | 1.0000 | **+89.64** |
| 450+ | 0.0767 | 1.0000 | **+92.33** |
**[INFERRED, important]** The oracle is *non-causal*: it aims with perfect
knowledge of the enemy's future, so its 100 % is a definition, not an
achievement, and the 92 pp at 450+ is an UPPER bound that contains both "a
better predictor could get this" and "this is physically unknowable". The
**realistic** causal bound measured today is the best arm at 450+: **HeadOn at
9.8 %**, barely above Pattern's 7.7 %. So the campaign is playing for a few
percentage points at long range, not for 92 pp. The honest target statement is
"raise the 450+ proxy from 7.7 % toward the ~10 % causal band", not "toward
100 %".
**0.3.2 The lead-gain sweep on Pattern is a DEAD END [MEASURED].** Multiply
Pattern's angular lead over LOS by a constant, per band:
| band | gain 1.0 | gain 1.5 | gain 2.0 | gain 3.0 |
|---|---|---|---|---|
| 0–100 | **10.56** | 13.76 | 20.11 | 34.77 |
| 100–200 | **14.75** | 19.64 | 26.78 | 42.97 |
| 200–300 | **16.61** | 22.26 | 29.47 | 45.23 |
| 300–450 | **17.53** | 23.28 | 29.96 | 44.27 |
| 450+ | **16.19** | 21.25 | 27.00 | 39.27 |
Gain 1.0 wins at EVERY band. Scaling Pattern's lead up makes it strictly worse.
This is the single most important negative result of Phase 0 and it should stop
any later job from "just adding more lead".
**0.3.3 The naive-linear / capture tension, resolved [MEASURED].** Capture slope
= regression of the arm's own lead on the required lead (job-95's statistic);
corr = Pearson correlation of the arm's lead with the required lead. **corr is
the informative number; a large slope on an uncorrelated lead is just amplified
noise.**
| band | mean\|req\| | Pattern cap / corr | naive-linear cap / corr | TMHorizon | BitBrain |
|---|---|---|---|---|---|
| 100–200 | 19.98 | 0.553 / 0.612 | 0.592 / 0.523 | 0.560 / 0.614 | 0.570 / 0.611 |
| 300–450 | 14.61 | 0.278 / 0.266 | 0.476 / 0.324 | 0.273 / 0.262 | 0.280 / 0.266 |
| 450+ | 12.33 | 0.175 / 0.165 | 0.310 / 0.178 | 0.174 / 0.164 | 0.175 / 0.165 |
Yes — on this corpus the naive-linear predictor applies **~1.8× more lead** than
Pattern at 450+ (0.310 vs 0.175; job-95 measured ~2×). Job-95's "we under-lead"
reading is confirmed. **But** the two arms carry almost the same lead
*information* (corr 0.178 vs 0.165), so the extra amplitude buys nothing and
costs angular accuracy: naive-linear's `mean|err|` is 22.86° vs Pattern's
16.19° at 450+. **Conclusion: the campaign's lever is lead INFORMATION
(correlation), not lead RESPONSE (capture slope).** Capturing more of an
uninformative lead is worse than capturing little of it — which is also exactly
why the gain sweep fails.
**0.3.4 The surprise: at long range, static line-of-sight beats Pattern.**
HeadOn (aim at the enemy's current position) has `mean|err|` 14.61°/12.33° and
`hitProxy` 0.105/0.098 at 300–450/450+, both better than Pattern's
17.53°/16.19° and 0.104/0.077. **[INFERRED]** At 450+ the required lead
(`mean|req|` = 12.3°) is essentially unpredictable from the past (Pattern
corr 0.165), so Pattern's predicted lead is mostly variance added to a nearly
uninformative signal; a zero-lead aim has error = `|required lead|`, which is
smaller. Consistent with the live record: the live bot's own applied lead
capture was 0.135 at 450+ (job-95), i.e. the live bot was already nearly
zero-lead and hit 9.14 % there; Pattern's offline proxy is 7.7 %.
**[INFERRED / CAVEAT]** The corpus is open-loop: DrussGT's recorded dodge was a
reaction to the LIVE bot's (near-zero-lead) bullets. Replaying Pattern on that
trajectory cannot show what DrussGT would do against Pattern's bullets. This
makes a **live A/B of HeadOn vs Pattern at long range the highest-value cheap
experiment in the campaign** (see §0.6). No offline claim that "HeadOn
beats Pattern" is permitted — only the live A/B decides.
**0.3.5 BitBrain, as shipped, is Pattern [MEASURED].** BitBrain's base is
Pattern and its ADE/SBC corrector changes almost nothing: 450+ `mean|err|`
16.200° vs Pattern 16.193°, `hitProxy` 0.0767 vs 0.0767. TMHorizon likewise
(16.199° / 0.0757). The corrector is currently **adding no measurable aim
information** on this corpus. That is the thing Phase 1 must change.
**0.3.6 Ruler resolution is NOT the limiter [MEASURED].** Aiming at the
integer-tick solve instead of the exact intercept costs only 0.29–1.10° of mean
error (`OracleQuant` column). So the "maybe the oracle only reaches 35 % because
the solve is coarse" worry is dead: the coarse/fine difference is ≈0.3° at long
range, far below the target tolerance (1.93° at 450+). Whatever caps the score,
it is the enemy's unpredictability, not the ruler.
### 0.4 Speed **[MEASURED]**
Full 70-run / 899 607-tick / 3 598 428 tick-bin sweep, 10 arms:
**406.9 s wall**, i.e. `0.1131 ms per tick-bin` over 10 arms,
**≈ 0.045 s per gun per 1000 ticks** (1000 ticks × 4 power bins).
BitBrain is the dominant cost (its ADE pass runs on every `predict` call);
Pattern/TMHorizon cache their per-tick work. A single-arm Pattern-only sweep is
several times cheaper. The binary cache (§0.1) is what makes repeat sweeps
affordable: without it every run re-parses 149 MB of JSONL.
### 0.5 How to reproduce **[MEASURED]**
```
nim c -d:release --nimcache:/tmp/nc_j98 -o:/tmp/bbq_run \
common_libs/tests/run_prediction_quality.nim
/tmp/bbq_run --corpus /tmp/tfil_ab2/out # full bar, ~7 min
/tmp/bbq_run --corpus /tmp/tfil_ab2/out --limit 10 # fast subset
/tmp/bbq_run --corpus /tmp/tfil_ab2/out --ruler quant # integer-tick solve
```
Verbatim full output: `common_libs/tests/prediction_quality_results.txt`.
Determinism: two consecutive full runs are byte-identical except the two
wall-time lines.
**Clean-checkout proof [MEASURED]:** `git archive HEAD | tar -x -C /tmp/bbq_clean`
then, from `/tmp/bbq_clean`,
`nim c -d:release --nimcache:/tmp/nc_j98 -o:bbq_run common_libs/tests/run_prediction_quality.nim`
builds, and `./bbq_run --corpus /tmp/tfil_ab2/out --limit 3` runs and prints the
same tables (separation 13.68× px on the 3-run subset). The committed harness is
self-contained; only the corpus is external.
### 0.6 Designs still to try (seed for later phases)
Ordered by expected value per unit of effort. Phase 0 has already killed one.
| # | design | why it is worth trying | status |
|---|---|---|---|
| D1 | **Live A/B: HeadOn at 450+ vs Pattern** (distance-gated switch, or HeadOn-only control) | Offline says a static gun beats Pattern at long range; the cheapest possible test of the campaign's central premise | **TODO (highest value, live)** |
| D2 | **Pattern variants that raise lead CORRELATION at long range**: longer keys / multi-length keys, per-distance learned pattern tables, different match weighting, k-NN over movement signatures | The lever is corr (0.165 at 450+), not gain; the ruler measures exactly this | TODO (offline-searchable) |
| D3 | **Supervise BitBrain with the ruler's own labels** — per-tick bearing error to the true intercept, trained on N−1 runs, evaluated on a held-out run | BitBrain's corrector currently adds nothing (0.3.5); the ruler gives it a real target. Must hold out runs or it is overfitting | TODO (offline-searchable) |
| D4 | **A causal "predictability" gate**: at each tick estimate whether the future is predictable (e.g. recent pattern-match score, reversal entropy) and fall back to HeadOn/low-variance aim when it is not | Directly attacks the 0.3.4 failure mode without needing a better long-range predictor | TODO |
| D5 | Power policy at long range (already partly done live): lower power = faster bullet = less lead error | Shortens the horizon the predictor must extrapolate; affects hit rate, live-only verdict | TODO (offline proxy only) |
| D6 | Lead-gain sweep 1.0/1.5/2.0/3.0 | **DEAD — measured.** Gain 1.0 wins at every band (§0.3.2) | **KILLED** |
Every D-item must end in a live A/B before any phase verdict.
### 0.7 What would make us quit
> If (a) no causal design raises the 450+ `hitProxy` above the static-gun
> reference (~0.10) on held-out runs by a margin larger than the run-to-run
> spread, **and** (b) the live A/B of the best such design shows no hit-rate or
> damage gain over Pattern with the left-running liveness check satisfied, then
> the campaign stops and we ship the simpler gun. We do not keep tuning an
> offline proxy that has stopped predicting live outcomes.
---
## Phase 1 — *(unclaimed; append below)*
## Phase 2 — *(unclaimed; append below)*