bitbrain campaign phase 0: offline prediction-quality ruler and the bar
New harness (common_libs/gun_harness/prediction_quality.nim + common_libs/tests/run_prediction_quality.nim): per-gun single-tick aim error in degrees against the true continuous interception point on the recorded live-vs-real-DrussGT corpus (/tmp/tfil_ab2/out, 70 runs, 899607 ticks), per range band, with the hit-probability proxy mean(|err|<=atan(18/range)). Validated: recorded hits separate from misses 13.34x px (reference 11.59x), perfect-oracle max |err| = 0, correct ordering on synthetic ground truth, two full runs byte-identical. Fixed a wrap180 bug (Nim float mod keeps the dividend sign) that inflated the negative error tail. Bar (mean|err| deg [hitProxy] at 450+): Pattern 16.19 [0.077], naive-linear 22.86 [0.054], TMHorizon 16.20 [0.076], BitBrain 16.20 [0.077], static HeadOn 12.33 [0.098], oracle 0 [1.0]. Lead-gain sweep on Pattern is a dead end (1.0 wins every band). Naive-linear applies ~1.8x Pattern's lead but carries no more lead information (corr 0.178 vs 0.165) and is strictly worse. Ledger: docs/bitbrain_campaign.md. All verdicts remain live-only.
This commit is contained in:
@@ -0,0 +1,258 @@
|
||||
# BitBrain campaign ledger
|
||||
|
||||
**Goal:** make ModularBot's gun **beat Pattern live** against the real DrussGT.
|
||||
The user has granted full freedom over the gun ("change input, output, every
|
||||
knob of it") and accepts it may fail — the deliverable is that the attempt is
|
||||
visible and evidence-backed.
|
||||
|
||||
**THE FINAL VERDICT IS ALWAYS LIVE.** Everything in this file except the
|
||||
`## Phase N` verdict lines is offline, open-loop, on a *fixed recorded enemy
|
||||
trajectory*. Per `docs/offline_harness_trust.md` (commit `e40c849`) the offline
|
||||
harness is trustworthy for exactly one thing: **per-gun single-tick prediction
|
||||
quality on a fixed enemy trajectory** — and it is *never* trustworthy for
|
||||
closed-loop questions (movement, range, round length, adaptation, gun
|
||||
selection, damage, wins, survival). No offline number here is a win/damage
|
||||
claim, and no phase may be called a success without a live A/B
|
||||
(`tools/ab/ab_run.sh`, server-side event hit rate, left-running).
|
||||
|
||||
Every claim below is tagged **[MEASURED]** (a command in §0 reproduces it) or
|
||||
**[INFERRED]** (reasoning from measured facts).
|
||||
|
||||
---
|
||||
|
||||
## Phase 0 — BUILD THE RULER AND ESTABLISH THE BAR *(owner: overnight job, committed)*
|
||||
|
||||
### 0.1 The ruler
|
||||
|
||||
`common_libs/gun_harness/prediction_quality.nim` + `common_libs/tests/run_prediction_quality.nim`.
|
||||
|
||||
At each recorded tick the shooter sits at `O = (selfX, selfY)`. For a bullet of
|
||||
speed `v` the **true interception point** is the first fractional time `t > 0`
|
||||
at which the enemy's ACTUAL recorded track reaches distance `v*t` from `O`
|
||||
(linear interpolation between recorded ticks). A bullet fired along the bearing
|
||||
to `E(t)` coincides with the enemy at `t`. Angular error is
|
||||
`wrap180(bearing(O→pred) − bearing(O→E(t)))` in **degrees**; every tick is
|
||||
scored for the four power bins (speeds 17/15.5/14/11), all bands share that
|
||||
horizon set. Per range band we report `mean|err|`, RMSE, mean signed err and the
|
||||
hit-probability proxy `mean(|err| ≤ atan(18/range))`.
|
||||
|
||||
The integer-tick solve from `analyze_lead_capture_by_range.py` (commit
|
||||
`f91e121`) is kept as `interceptBearingQuant` and reported as `OracleQuant`; the
|
||||
ruler ships the **continuous** solve because it separates recorded hits from
|
||||
misses slightly better and removes the coarse solve's own overshoot
|
||||
(§0.3.6). `--ruler quant` selects the integer solve.
|
||||
|
||||
**Data:** the recorded live-vs-real-DrussGT corpus `/tmp/tfil_ab2/out`
|
||||
(70 battles / 490 rounds / 899 607 ticks + `.events.jsonl` + `.rounds.json`),
|
||||
**verified present before use**. It lives in `/tmp` and is therefore ephemeral;
|
||||
if a later job finds it gone, regenerate it with the A/B harness
|
||||
(`tools/ab/ab_run.sh`, which sets `TR_RECORD_WORLDSTATE` so ModularBot appends
|
||||
per-tick world state) and point `--corpus` at the new output root. Layout:
|
||||
`<root>/<arm>/runN.jsonl` + `runN.events.jsonl` + `runN.jsonl.rounds.json`.
|
||||
149 MB of JSONL is converted once per run into a compact float32 `.qcache`
|
||||
(keyed on source mtime+size) and ALL measurement is taken from the cache, so two
|
||||
runs are byte-identical. See §0.4 for speed.
|
||||
|
||||
### 0.2 Validation — the ruler must pass ALL of these **[MEASURED]**
|
||||
|
||||
Run: `nim c -d:release --nimcache:/tmp/nc_j98 -r common_libs/tests/run_prediction_quality.nim`
|
||||
|
||||
**1. Recorded HITS separate from recorded MISSES** (our ACTUAL server-fired
|
||||
bearings, scored against the SAME interception solve):
|
||||
|
||||
| ruler | hits n | hits mean\|err\| | misses n | misses mean\|err\| | separation |
|
||||
|---|---|---|---|---|---|
|
||||
| continuous | 5480 | **1.360° / 10.5 px** | 48304 | **16.597° / 140.3 px** | **12.20× deg / 13.34× px** |
|
||||
| integer | 5480 | 1.478° / 11.4 px | 48304 | 16.724° / 141.3 px | 11.32× / 12.43× |
|
||||
|
||||
(The earlier validated run quoted 11.59× / 11.6 px on hits; reproduced and
|
||||
improved.) The continuous ruler is shipped because it separates better.
|
||||
|
||||
**2. Perfect oracle scores 0.** Max `|err|` over all 3 598 428 tick-bins =
|
||||
**0.000000°**. OK.
|
||||
|
||||
**3. A static line-of-sight gun is far from the predictor on learnable motion.**
|
||||
On a synthetic constant-velocity and a seeded random-walk trajectory the
|
||||
ordering is exactly as physics demands: HeadOn (zero lead) is the worst, Pattern
|
||||
and naive-linear are near-zero, and the lead-gain arms overshoot monotonically.
|
||||
On the real DrussGT corpus the static gun is *not* worst — see §0.3.4, this is a
|
||||
genuine property of the corpus, not a harness defect.
|
||||
|
||||
**4. Determinism.** Two full 70-run sweeps, stdout diffed with the two wall-time
|
||||
lines excluded: **byte-identical**. (The only difference between the two raw
|
||||
outputs is `wall time 406.88s` vs `402.41s` and the derived ms-per-tick-bin.)
|
||||
**[MEASURED]**
|
||||
|
||||
**5. A real bug was found and fixed by this validation.** The ruler's
|
||||
`wrap180` used Nim's float `mod`, which keeps the dividend's sign (C `fmod`), so
|
||||
`(x+180) mod 360 − 180` returned `x−360` instead of the wrapped equivalent for
|
||||
`x < −180`. This inflated the negative tail of every error (maxAbs read ~360°
|
||||
instead of ~180°) and made HeadOn's mean error disagree with `mean|required|`.
|
||||
After the fix HeadOn's `mean|err|` equals `mean|required lead|` to the last
|
||||
digit at every band (see the `HO |err| / HO |req| / Pat|req|` columns in the
|
||||
fixture). **A wrong ruler is worse than no ruler; this was the most important
|
||||
10 minutes of the phase.**
|
||||
|
||||
### 0.3 THE BAR — per-band numbers (70 runs, 3 598 428 tick-bins) **[MEASURED]**
|
||||
|
||||
Format: `mean|err| deg` and, in brackets, `hitProxy`. `hitProxy` is the fraction
|
||||
of tick-bins aimed within `atan(18/range)` of the true interception point.
|
||||
|
||||
| band | Pattern | naive-linear | TMHorizon | BitBrain | HeadOn (static) | Oracle |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 0–100 | **10.56** [0.699] | 17.92 [0.651] | 10.68 [0.696] | 11.12 [0.681] | 19.62 [0.342] | 0.00 [1.000] |
|
||||
| 100–200 | **14.75** [0.342] | 14.63 [0.388] | 14.84 [0.342] | 15.11 [0.324] | 19.98 [0.172] | 0.00 [1.000] |
|
||||
| 200–300 | **16.61** [0.185] | 17.57 [0.192] | 16.64 [0.179] | 16.84 [0.175] | 17.34 [0.133] | 0.00 [1.000] |
|
||||
| 300–450 | 17.53 [0.104] | 20.98 [0.100] | 17.57 [0.100] | 17.58 [0.103] | **14.61** [0.105] | 0.00 [1.000] |
|
||||
| 450+ | 16.19 [0.077] | 22.86 [0.054] | 16.20 [0.076] | 16.20 [0.077] | **12.33** [0.098] | 0.00 [1.000] |
|
||||
|
||||
`n`: 4 423 / 24 908 / 74 215 / 1 119 777 / 2 311 323. Skipped (no valid
|
||||
interception): 63 782 tick-bins (≈1.7 %).
|
||||
|
||||
**0.3.1 DIRECT ANSWER — the gap between Pattern and the oracle ceiling:**
|
||||
|
||||
| band | Pattern hitProxy | Oracle hitProxy | headroom (pp) |
|
||||
|---|---|---|---|
|
||||
| 0–100 | 0.6993 | 1.0000 | **+30.07** |
|
||||
| 100–200 | 0.3418 | 1.0000 | **+65.82** |
|
||||
| 200–300 | 0.1850 | 1.0000 | **+81.50** |
|
||||
| 300–450 | 0.1036 | 1.0000 | **+89.64** |
|
||||
| 450+ | 0.0767 | 1.0000 | **+92.33** |
|
||||
|
||||
**[INFERRED, important]** The oracle is *non-causal*: it aims with perfect
|
||||
knowledge of the enemy's future, so its 100 % is a definition, not an
|
||||
achievement, and the 92 pp at 450+ is an UPPER bound that contains both "a
|
||||
better predictor could get this" and "this is physically unknowable". The
|
||||
**realistic** causal bound measured today is the best arm at 450+: **HeadOn at
|
||||
9.8 %**, barely above Pattern's 7.7 %. So the campaign is playing for a few
|
||||
percentage points at long range, not for 92 pp. The honest target statement is
|
||||
"raise the 450+ proxy from 7.7 % toward the ~10 % causal band", not "toward
|
||||
100 %".
|
||||
|
||||
**0.3.2 The lead-gain sweep on Pattern is a DEAD END [MEASURED].** Multiply
|
||||
Pattern's angular lead over LOS by a constant, per band:
|
||||
|
||||
| band | gain 1.0 | gain 1.5 | gain 2.0 | gain 3.0 |
|
||||
|---|---|---|---|---|
|
||||
| 0–100 | **10.56** | 13.76 | 20.11 | 34.77 |
|
||||
| 100–200 | **14.75** | 19.64 | 26.78 | 42.97 |
|
||||
| 200–300 | **16.61** | 22.26 | 29.47 | 45.23 |
|
||||
| 300–450 | **17.53** | 23.28 | 29.96 | 44.27 |
|
||||
| 450+ | **16.19** | 21.25 | 27.00 | 39.27 |
|
||||
|
||||
Gain 1.0 wins at EVERY band. Scaling Pattern's lead up makes it strictly worse.
|
||||
This is the single most important negative result of Phase 0 and it should stop
|
||||
any later job from "just adding more lead".
|
||||
|
||||
**0.3.3 The naive-linear / capture tension, resolved [MEASURED].** Capture slope
|
||||
= regression of the arm's own lead on the required lead (job-95's statistic);
|
||||
corr = Pearson correlation of the arm's lead with the required lead. **corr is
|
||||
the informative number; a large slope on an uncorrelated lead is just amplified
|
||||
noise.**
|
||||
|
||||
| band | mean\|req\| | Pattern cap / corr | naive-linear cap / corr | TMHorizon | BitBrain |
|
||||
|---|---|---|---|---|---|
|
||||
| 100–200 | 19.98 | 0.553 / 0.612 | 0.592 / 0.523 | 0.560 / 0.614 | 0.570 / 0.611 |
|
||||
| 300–450 | 14.61 | 0.278 / 0.266 | 0.476 / 0.324 | 0.273 / 0.262 | 0.280 / 0.266 |
|
||||
| 450+ | 12.33 | 0.175 / 0.165 | 0.310 / 0.178 | 0.174 / 0.164 | 0.175 / 0.165 |
|
||||
|
||||
Yes — on this corpus the naive-linear predictor applies **~1.8× more lead** than
|
||||
Pattern at 450+ (0.310 vs 0.175; job-95 measured ~2×). Job-95's "we under-lead"
|
||||
reading is confirmed. **But** the two arms carry almost the same lead
|
||||
*information* (corr 0.178 vs 0.165), so the extra amplitude buys nothing and
|
||||
costs angular accuracy: naive-linear's `mean|err|` is 22.86° vs Pattern's
|
||||
16.19° at 450+. **Conclusion: the campaign's lever is lead INFORMATION
|
||||
(correlation), not lead RESPONSE (capture slope).** Capturing more of an
|
||||
uninformative lead is worse than capturing little of it — which is also exactly
|
||||
why the gain sweep fails.
|
||||
|
||||
**0.3.4 The surprise: at long range, static line-of-sight beats Pattern.**
|
||||
HeadOn (aim at the enemy's current position) has `mean|err|` 14.61°/12.33° and
|
||||
`hitProxy` 0.105/0.098 at 300–450/450+, both better than Pattern's
|
||||
17.53°/16.19° and 0.104/0.077. **[INFERRED]** At 450+ the required lead
|
||||
(`mean|req|` = 12.3°) is essentially unpredictable from the past (Pattern
|
||||
corr 0.165), so Pattern's predicted lead is mostly variance added to a nearly
|
||||
uninformative signal; a zero-lead aim has error = `|required lead|`, which is
|
||||
smaller. Consistent with the live record: the live bot's own applied lead
|
||||
capture was 0.135 at 450+ (job-95), i.e. the live bot was already nearly
|
||||
zero-lead and hit 9.14 % there; Pattern's offline proxy is 7.7 %.
|
||||
|
||||
**[INFERRED / CAVEAT]** The corpus is open-loop: DrussGT's recorded dodge was a
|
||||
reaction to the LIVE bot's (near-zero-lead) bullets. Replaying Pattern on that
|
||||
trajectory cannot show what DrussGT would do against Pattern's bullets. This
|
||||
makes a **live A/B of HeadOn vs Pattern at long range the highest-value cheap
|
||||
experiment in the campaign** (see §0.6). No offline claim that "HeadOn
|
||||
beats Pattern" is permitted — only the live A/B decides.
|
||||
|
||||
**0.3.5 BitBrain, as shipped, is Pattern [MEASURED].** BitBrain's base is
|
||||
Pattern and its ADE/SBC corrector changes almost nothing: 450+ `mean|err|`
|
||||
16.200° vs Pattern 16.193°, `hitProxy` 0.0767 vs 0.0767. TMHorizon likewise
|
||||
(16.199° / 0.0757). The corrector is currently **adding no measurable aim
|
||||
information** on this corpus. That is the thing Phase 1 must change.
|
||||
|
||||
**0.3.6 Ruler resolution is NOT the limiter [MEASURED].** Aiming at the
|
||||
integer-tick solve instead of the exact intercept costs only 0.29–1.10° of mean
|
||||
error (`OracleQuant` column). So the "maybe the oracle only reaches 35 % because
|
||||
the solve is coarse" worry is dead: the coarse/fine difference is ≈0.3° at long
|
||||
range, far below the target tolerance (1.93° at 450+). Whatever caps the score,
|
||||
it is the enemy's unpredictability, not the ruler.
|
||||
|
||||
### 0.4 Speed **[MEASURED]**
|
||||
|
||||
Full 70-run / 899 607-tick / 3 598 428 tick-bin sweep, 10 arms:
|
||||
**406.9 s wall**, i.e. `0.1131 ms per tick-bin` over 10 arms,
|
||||
**≈ 0.045 s per gun per 1000 ticks** (1000 ticks × 4 power bins).
|
||||
BitBrain is the dominant cost (its ADE pass runs on every `predict` call);
|
||||
Pattern/TMHorizon cache their per-tick work. A single-arm Pattern-only sweep is
|
||||
several times cheaper. The binary cache (§0.1) is what makes repeat sweeps
|
||||
affordable: without it every run re-parses 149 MB of JSONL.
|
||||
|
||||
### 0.5 How to reproduce **[MEASURED]**
|
||||
|
||||
```
|
||||
nim c -d:release --nimcache:/tmp/nc_j98 -o:/tmp/bbq_run \
|
||||
common_libs/tests/run_prediction_quality.nim
|
||||
/tmp/bbq_run --corpus /tmp/tfil_ab2/out # full bar, ~7 min
|
||||
/tmp/bbq_run --corpus /tmp/tfil_ab2/out --limit 10 # fast subset
|
||||
/tmp/bbq_run --corpus /tmp/tfil_ab2/out --ruler quant # integer-tick solve
|
||||
```
|
||||
Verbatim full output: `common_libs/tests/prediction_quality_results.txt`.
|
||||
Determinism: two consecutive full runs are byte-identical except the two
|
||||
wall-time lines.
|
||||
|
||||
**Clean-checkout proof [MEASURED]:** `git archive HEAD | tar -x -C /tmp/bbq_clean`
|
||||
then, from `/tmp/bbq_clean`,
|
||||
`nim c -d:release --nimcache:/tmp/nc_j98 -o:bbq_run common_libs/tests/run_prediction_quality.nim`
|
||||
builds, and `./bbq_run --corpus /tmp/tfil_ab2/out --limit 3` runs and prints the
|
||||
same tables (separation 13.68× px on the 3-run subset). The committed harness is
|
||||
self-contained; only the corpus is external.
|
||||
|
||||
### 0.6 Designs still to try (seed for later phases)
|
||||
|
||||
Ordered by expected value per unit of effort. Phase 0 has already killed one.
|
||||
|
||||
| # | design | why it is worth trying | status |
|
||||
|---|---|---|---|
|
||||
| D1 | **Live A/B: HeadOn at 450+ vs Pattern** (distance-gated switch, or HeadOn-only control) | Offline says a static gun beats Pattern at long range; the cheapest possible test of the campaign's central premise | **TODO (highest value, live)** |
|
||||
| D2 | **Pattern variants that raise lead CORRELATION at long range**: longer keys / multi-length keys, per-distance learned pattern tables, different match weighting, k-NN over movement signatures | The lever is corr (0.165 at 450+), not gain; the ruler measures exactly this | TODO (offline-searchable) |
|
||||
| D3 | **Supervise BitBrain with the ruler's own labels** — per-tick bearing error to the true intercept, trained on N−1 runs, evaluated on a held-out run | BitBrain's corrector currently adds nothing (0.3.5); the ruler gives it a real target. Must hold out runs or it is overfitting | TODO (offline-searchable) |
|
||||
| D4 | **A causal "predictability" gate**: at each tick estimate whether the future is predictable (e.g. recent pattern-match score, reversal entropy) and fall back to HeadOn/low-variance aim when it is not | Directly attacks the 0.3.4 failure mode without needing a better long-range predictor | TODO |
|
||||
| D5 | Power policy at long range (already partly done live): lower power = faster bullet = less lead error | Shortens the horizon the predictor must extrapolate; affects hit rate, live-only verdict | TODO (offline proxy only) |
|
||||
| D6 | Lead-gain sweep 1.0/1.5/2.0/3.0 | **DEAD — measured.** Gain 1.0 wins at every band (§0.3.2) | **KILLED** |
|
||||
|
||||
Every D-item must end in a live A/B before any phase verdict.
|
||||
|
||||
### 0.7 What would make us quit
|
||||
|
||||
> If (a) no causal design raises the 450+ `hitProxy` above the static-gun
|
||||
> reference (~0.10) on held-out runs by a margin larger than the run-to-run
|
||||
> spread, **and** (b) the live A/B of the best such design shows no hit-rate or
|
||||
> damage gain over Pattern with the left-running liveness check satisfied, then
|
||||
> the campaign stops and we ship the simpler gun. We do not keep tuning an
|
||||
> offline proxy that has stopped predicting live outcomes.
|
||||
|
||||
---
|
||||
|
||||
## Phase 1 — *(unclaimed; append below)*
|
||||
|
||||
## Phase 2 — *(unclaimed; append below)*
|
||||
Reference in New Issue
Block a user