15 KiB
BitBrain gate test — can ADE+SBC predict a fine-grained aim correction?
Question. Given the already-built generic BitBrain library
(common_libs/bitbrain/, commit 77e6dac, which reproduced the reference C on
MNIST to the digit), is there signal in a fine-grained angular aim
correction that neither a naive predictor nor the shipped Pattern gun already
has? If not, we stop before writing a gun.
Scope. Offline only. No gun wiring, no battle, no server, no rack registration, no changed defaults. Fixtures are read-only.
Tooling (the evidence):
common_libs/tests/measure_bitbrain_gate.nim— the analyzer.common_libs/tests/measure_bitbrain_gate_results.txt— its full deterministic output.common_libs/bitbrain/— the library under test (untouched).
Reproduce:
nim c -r -d:release --nimcache:/tmp/nc_j92 --path:common_libs \
common_libs/tests/measure_bitbrain_gate.nim
# optional knobs: BB_POWER, BB_PMAX, BB_TARGET, BB_PASSES, BB_STEP, BB_STRIDE,
# BB_NADE, BB_NLIST, BB_FILES
Direct answer
There is signal, but it does not clear the bar as a shipping gun.
- vs straight-line naive: BitBrain wins decisively on every readout and every configuration (pooled mean arrival error 124.7 px vs 160.2 px).
- vs the shipped Pattern gun: pooled, the best BitBrain configuration (per-round reset, argmax readout, N=32, nAde=256) beats Pattern on mean arrival error (124.6 px vs 139.8 px, −10.8 %; 13.3° vs 15.6°, −14.8 %) and on both hit rates (18 px hit 8.3 % vs 7.0 %; angular hit 26.1 % vs 19.1 %). However the hit-rate gain is not robust across fixtures: it is concentrated in the two fixtures where Pattern is weak (modularbot, modularbot_shield) and BitBrain loses hit rate on the two fixtures where Pattern is strongest (spinbot, crazy). It also only works in the per-round-reset regime; the retained-across-rounds regime the user actually wants is the weakest (it improves average error slightly but lowers the hit rate).
- Bottleneck (MEASURED): sample starvation / SBC memory saturation, not the AD synthesis and not an absence of signal.
A sword that helps exactly where the incumbent is already weak, and hurts where it is strong, is not a gun improvement. The honest verdict is signal yes, shippable improvement no (yet) — see the bottleneck section.
What was measured (MEASURED unless tagged INFERRED)
Data
The 5 committed Tank-Royale bridge fixtures tools/fixtures/tr_drussgt_vs_*
(open-loop replay), read-only:
| fixture | ticks | resolved samples |
|---|---|---|
tr_drussgt_vs_corners.jsonl |
2575 | 2173 |
tr_drussgt_vs_crazy.jsonl |
11507 | 11209 |
tr_drussgt_vs_modularbot.jsonl |
20026 | 19548 |
tr_drussgt_vs_modularbot_shield.jsonl |
12629 | 12308 |
tr_drussgt_vs_spinbot.jsonl |
10824 | 10494 |
| total | 55732 (55 rounds, ~1013 samples/round) |
The fixtures are open-loop: the recorded enemy does not react to us. That is acceptable here because this test measures single-tick prediction quality, the one category where fixture replay reproduces live behaviour faithfully. It would not be acceptable evidence for a movement or adaptation claim. Do not over-read the hit numbers as live hit rates.
Input — the TMHorizon 53 bits, reused, not re-derived
The analyzer drives the live TmHorizonGun one tick at a time and reads its
own exported builder tmhBaseBits (49 draft bits) plus the 4-bit horizon
one-hot via tmhLits — the exact cachedBits per-tick path, with
g.resetRoundState() called at each round boundary to mirror the live
onRoundStarted. (The brief calls this tmhBuildBits; the actual symbol is
tmhBaseBits.) This keeps the comparison apples-to-apples with the TM gun
already measured.
Output — fine-grained angular correction class
The correction is the bearing offset added to Pattern's prediction. N class bins are laid over a fixed ±40° range (class width 80°/N). Two readouts:
- wm = the count-weighted mean of the class centres, weighted by the per-class set-bit counts summed over the 6 cross-AD SBCs. No evidence → 0 correction (i.e. Pattern).
- arg = the argmax class centre; no evidence → 0.
Label — the +h-tick fact (never across a round)
h = tmhHorizonFor(dist, speed) = clamp(round(dist/(20−3·power)), 10, 50).
At fire tick t, the label is the actual angular offset of the enemy at
t+h (from the same fixture, which under perfect-info replay equals the bot's
own observation ring) relative to Pattern's base bearing. Samples with
t+h past the round end are dropped (never a cross-boundary label). Pooled
|label err|: mean 15.9°, p50 11.6°, p90 37.2°, p99 52.0°.
Metric — arrival aim error, not accuracy
Harness bmPoint geometry, the relation measure_aim_vs_power.nim validated
against the harness resolver: fireDist = |Pattern − self|,
arrivalTick = t + ceil(fireDist/v) − 1. A rotation preserves fireDist, so
the corrected aim point is rotated around the shooter. miss = |aim − actual enemy pos at arrivalTick|; hit = miss < 18 px. Angular error = |aim bearing − actual bearing|; angular hit = |angErr| < atan(18/range). Timed
resolution happens on arrivalTick, which can differ from t+h by ≤ half a
tick; the label uses h, the metric uses arrivalTick, exactly as the brief
specifies.
The px metric includes range error, and because the head is a pure rotation it cannot fix range; this is why the 18 px hit is dominated by range error and the angular metric is the cleaner measure of a rotation head. Both are reported.
Protocol — prequential (predict-then-learn, streaming)
For every sample the model predicts before it is updated with the label. Two regimes:
- retained — learning accumulates across all rounds of one battle (fixture); reset only when the battle/enemy changes. This is what the user asked for.
- perRound — reset at every round boundary (the worst case).
Baselines
- straight-line naive — enemy keeps its fire-tick velocity over the same
arrivalTickwindow. - always-the-same-answer — a fixed correction equal to the global mean label (+0.20°).
- Pattern — the shipped gun's own prediction (zero correction).
Pooled over all 55732 samples:
| predictor | meanPx | medPx | p90Px | meanDeg | medDeg | pxHit% | angHit% |
|---|---|---|---|---|---|---|---|
| Pattern (zero corr) | 139.78 | 112.71 | 299.10 | 15.56 | 11.73 | 7.0 | 19.1 |
| straight-line naive | 160.17 | 131.57 | 336.48 | 16.01 | 12.18 | 5.0 | 18.8 |
| fixed (+0.20°) | 139.76 | 112.66 | 298.85 | 15.56 | 11.74 | 7.0 | 19.0 |
The AD layer (synthesised for our data)
Random ADs (initRandomAddressDecoder, widths {6, 8, 10, 12}, center = 0),
then the deterministic homeostatic controller (accumulateFiring +
adaptThresholds, target 1 %). center = 0 is forced by our data: the
reference's 127 is the midpoint of 0..255; centring binary 0/1 inputs at 127
makes every synapse contribute ≈ −127 and collapses the ADE code to a mere
polarity count, destroying the signal.
The paper's step = 1 controller would need thousands of intervals to find the
1 % operating point — far more than a battle (500–2000 ticks) provides.
This is itself the first measured symptom of sample starvation. The analyzer
therefore initialises each ADE's threshold at the score that puts it closest to
the 1 % firing count (a fast, unsupervised percentile), then runs the
deterministic controller (step = 1, 2 passes) to refine it. Achieved firing
rates (MEASURED, on the fit stride):
| nAde | w6 | w8 | w10 | w12 | mean |
|---|---|---|---|---|---|
| 128, pct-init | 0.47 % | 0.56 % | 0.66 % | 0.70 % | 0.59 % |
| 128, +homeostasis | 1.20 % | 1.50 % | 1.34 % | 1.32 % | 1.34 % |
| 256, pct-init | 0.37 % | 0.50 % | 0.62 % | 0.69 % | 0.54 % |
| 256, +homeostasis | 1.16 % | 1.30 % | 1.31 % | 1.47 % | 1.31 % |
So the paper's ~1 % operating point is reached. (The controller alone
overshot in an earlier pass at step = 2; step = 1 lands it.)
Sweep — N × AD size × regime (pooled, mean px error)
wm = count-weighted mean, arg = argmax. hit% is the 18 px arrival hit.
| config | wm meanPx | arg meanPx | wm pxHit% | arg pxHit% |
|---|---|---|---|---|
| Pattern | 139.78 | — | 7.0 | — |
| straight-line | 160.17 | — | 5.0 | — |
| fixed | 139.76 | — | 7.0 | — |
| nAde128 / N4 / retained | 136.93 | 166.79 | 4.8 | 2.4 |
| nAde128 / N4 / perRound | 130.27 | 140.28 | 4.3 | 3.0 |
| nAde128 / N8 / perRound | 128.87 | 134.81 | 4.5 | 4.7 |
| nAde128 / N16 / perRound | 128.78 | 133.72 | 5.2 | 6.7 |
| nAde128 / N32 / perRound | 128.83 | 133.97 | 5.2 | 7.5 |
| nAde128 / N64 / perRound | 128.97 | 134.67 | 5.2 | 7.6 |
| nAde256 / N8 / perRound | 124.78 | 125.60 | 4.3 | 4.7 |
| nAde256 / N16 / perRound | 124.72 | 124.46 | 4.8 | 7.3 |
| nAde256 / N32 / perRound | 124.95 | 124.62 | 4.9 | 8.3 |
| nAde256 / N64 / perRound | 125.13 | 125.34 | 5.0 | 8.5 |
| nAde256 / N16 / retained | 134.31 | 149.00 | 5.4 | 5.6 |
| nAde256 / N32 / retained | 134.23 | 149.43 | 5.3 | 6.5 |
| nAde256 / N64 / retained | 134.30 | 151.60 | 5.3 | 6.5 |
Full per-config degrees/p90/angular-hit rows are in
measure_bitbrain_gate_results.txt.
Where the error stops falling, and why
- N: the
wmerror is flat from N=8 to N=64 (≈124.7–125.1 px); theargerror falls to N=16–32 then flattens. Optimum N ≈ 16–32. Beyond it the correction classes get finer than the loop can resolve, and the SBC coincidence cells are already too few to constrain their class bits — more classes only split the same evidence. - AD size: nAde=256 beats 128 by a modest ~3 % in
wm; both are far from saturating, but the classes saturate first. Doubling the ADE count does not double the information. - Why it stops: MNIST needed ~60 000 examples for 10 mutually exclusive classes. Here we have ~55 000 samples for 16–64 correction classes whose evidence must separate by 1–2° — i.e. ~2 orders of magnitude less evidence per class. The SBC is idempotent (a cell accumulates every class that ever co-occurred, with no decay), so with too few examples per cell the per-class counts blur toward uniform and the readout regresses toward the mean. That is sample starvation / memory saturation, and it is consistent with every other observation (flat N tail, weak nAde scaling, retained < perRound).
Count-weighted mean vs argmax (the brief's hypothesis)
The brief expected the count-weighted mean to be the key readout because the TM's discarded magnitude. MEASURED, that is only half right:
- The
wmis the shrinkage readout: it reduces mean error (and extreme misses) but lowers the hit rate (pooled pxHit 4.8 % vs Pattern 7.0 %, angular hit 14.9 % vs 19.1 %). It never makes a confident, sharp correction. - The
argmaxis the decision readout: it keeps the same mean-error reduction and improves the hit rate (pxHit 8.3 %, angular hit 26.1 %). It is the readout that beats Pattern on all four metrics.
So the fine-grained head works, but as a classifier (argmax), not as a soft regression (weighted mean). The weighted mean is a useful control: it is the readout whose shuffled-label null collapses to the baseline.
Retained vs per-round reset
Per-round reset beats retained across rounds on every readout and every N
(retained wm ≈ 134.2 px, retained arg ≈ 149–162 px; per-round wm ≈ 124.7,
per-round arg ≈ 124.5). The user wants retention across the battle; the
measurement says the idempotent SBC accumulates stale, conflicting class bits
across rounds and the extra evidence hurts. This mirrors the project's earlier
TM finding ("forgetting is stronger than accumulation"). A viable gun would
need a bounded/decaying SBC, which the library does not have.
Per-fixture breakdown (perRound, N=32, argmax)
| fixture (n) | Pattern meanPx / pxHit% / angHit% | BitBrain arg meanPx / pxHit% / angHit% |
|---|---|---|
| corners (2173) | 164.76 / 3.9 / 18.5 | 132.64 / 3.6 / 23.8 |
| crazy (11209) | 126.34 / 7.3 / 25.8 | 111.95 / 5.9 / 24.6 |
| modularbot (19548) | 151.80 / 3.3 / 10.5 | 134.46 / 7.5 / 25.1 |
| shield (12308) | 144.28 / 6.6 / 14.1 | 124.98 / 10.5 / 28.5 |
| spinbot (10494) | 121.29 / 15.0 / 33.7 | 117.72 / 10.6 / 27.3 |
Mean error improves on all five. Hit rate improves on modularbot and shield (where Pattern is weak) and regresses on spinbot and crazy (where Pattern is strong), with corners a wash. That is the whole verdict in one table.
Shuffled-label control (must collapse)
Labels permuted across all samples (3 seeds), same inputs:
| regime | wm shuffled meanPx / pxHit% | arg shuffled meanPx / pxHit% |
|---|---|---|
| retained | 141.51 / 5.8 | 200.10 / 2.4 |
| perRound | 146.74 / 4.6 | 193.22 / 2.4 |
The weighted mean collapses toward the baseline (141.5 vs Pattern 139.8) — expected, because it shrinks to the (near-zero) label mean. The argmax does not collapse to the baseline: its null is worse than the baseline — with no signal it still makes a confident, essentially random rotation, which is worse than no correction. That is the correct null behaviour for a non-shrinking readout, and it is why the honest control is real vs shuffled within the same readout: BitBrain argmax is ~124.6 px on real labels vs ~193 px on shuffled labels. The learning is real; the signal is not an artifact.
Bottleneck and recommendation (MEASURED)
Ranked by how much each could plausibly close the gap:
- Sample starvation / SBC memory saturation — the dominant one. N saturates at ~16–32, nAde barely scales, and per-round reset beats retention. The library has no bounded/decaying SBC, so a long battle only blurs.
- Fixture-dependent gain. The pooled hit win is carried by the weak-Pattern fixtures. Without an online per-fixture selector, a blanket substitution would lose on spinbot/crazy.
- AD synthesis is not the bottleneck. The ~1 % operating point is
reached and the shuffle control shows the ADs are informative. The forced
center = 0for binary inputs is a correctness requirement, not a defect.
Do not build the gun yet. The cheap decisive next step, if pursued, is a bounded/decaying SBC (a per-round or recency-weighted memory) plus an online selection gate that keeps Pattern where BitBrain is worse — the only shape the data supports. A wider class range or a larger nAde will not fix the starvation.
All numbers MEASURED by common_libs/tests/measure_bitbrain_gate.nim on this
machine, deterministic (fixed seeds). Arrival geometry is the harness bmPoint
relation validated in measure_aim_vs_power.nim. The fixtures are open-loop;
treat the hit rates as prediction-quality evidence only.