Files
SirRoboGarage/docs/bitbrain_gate_test.md

15 KiB
Raw Permalink Blame History

BitBrain gate test — can ADE+SBC predict a fine-grained aim correction?

Question. Given the already-built generic BitBrain library (common_libs/bitbrain/, commit 77e6dac, which reproduced the reference C on MNIST to the digit), is there signal in a fine-grained angular aim correction that neither a naive predictor nor the shipped Pattern gun already has? If not, we stop before writing a gun.

Scope. Offline only. No gun wiring, no battle, no server, no rack registration, no changed defaults. Fixtures are read-only.

Tooling (the evidence):

  • common_libs/tests/measure_bitbrain_gate.nim — the analyzer.
  • common_libs/tests/measure_bitbrain_gate_results.txt — its full deterministic output.
  • common_libs/bitbrain/ — the library under test (untouched).

Reproduce:

nim c -r -d:release --nimcache:/tmp/nc_j92 --path:common_libs \
  common_libs/tests/measure_bitbrain_gate.nim
# optional knobs: BB_POWER, BB_PMAX, BB_TARGET, BB_PASSES, BB_STEP, BB_STRIDE,
#                 BB_NADE, BB_NLIST, BB_FILES

Direct answer

There is signal, but it does not clear the bar as a shipping gun.

  • vs straight-line naive: BitBrain wins decisively on every readout and every configuration (pooled mean arrival error 124.7 px vs 160.2 px).
  • vs the shipped Pattern gun: pooled, the best BitBrain configuration (per-round reset, argmax readout, N=32, nAde=256) beats Pattern on mean arrival error (124.6 px vs 139.8 px, −10.8 %; 13.3° vs 15.6°, −14.8 %) and on both hit rates (18 px hit 8.3 % vs 7.0 %; angular hit 26.1 % vs 19.1 %). However the hit-rate gain is not robust across fixtures: it is concentrated in the two fixtures where Pattern is weak (modularbot, modularbot_shield) and BitBrain loses hit rate on the two fixtures where Pattern is strongest (spinbot, crazy). It also only works in the per-round-reset regime; the retained-across-rounds regime the user actually wants is the weakest (it improves average error slightly but lowers the hit rate).
  • Bottleneck (MEASURED): sample starvation / SBC memory saturation, not the AD synthesis and not an absence of signal.

A sword that helps exactly where the incumbent is already weak, and hurts where it is strong, is not a gun improvement. The honest verdict is signal yes, shippable improvement no (yet) — see the bottleneck section.


What was measured (MEASURED unless tagged INFERRED)

Data

The 5 committed Tank-Royale bridge fixtures tools/fixtures/tr_drussgt_vs_* (open-loop replay), read-only:

fixture ticks resolved samples
tr_drussgt_vs_corners.jsonl 2575 2173
tr_drussgt_vs_crazy.jsonl 11507 11209
tr_drussgt_vs_modularbot.jsonl 20026 19548
tr_drussgt_vs_modularbot_shield.jsonl 12629 12308
tr_drussgt_vs_spinbot.jsonl 10824 10494
total 55732 (55 rounds, ~1013 samples/round)

The fixtures are open-loop: the recorded enemy does not react to us. That is acceptable here because this test measures single-tick prediction quality, the one category where fixture replay reproduces live behaviour faithfully. It would not be acceptable evidence for a movement or adaptation claim. Do not over-read the hit numbers as live hit rates.

Input — the TMHorizon 53 bits, reused, not re-derived

The analyzer drives the live TmHorizonGun one tick at a time and reads its own exported builder tmhBaseBits (49 draft bits) plus the 4-bit horizon one-hot via tmhLits — the exact cachedBits per-tick path, with g.resetRoundState() called at each round boundary to mirror the live onRoundStarted. (The brief calls this tmhBuildBits; the actual symbol is tmhBaseBits.) This keeps the comparison apples-to-apples with the TM gun already measured.

Output — fine-grained angular correction class

The correction is the bearing offset added to Pattern's prediction. N class bins are laid over a fixed ±40° range (class width 80°/N). Two readouts:

  • wm = the count-weighted mean of the class centres, weighted by the per-class set-bit counts summed over the 6 cross-AD SBCs. No evidence → 0 correction (i.e. Pattern).
  • arg = the argmax class centre; no evidence → 0.

Label — the +h-tick fact (never across a round)

h = tmhHorizonFor(dist, speed) = clamp(round(dist/(20−3·power)), 10, 50). At fire tick t, the label is the actual angular offset of the enemy at t+h (from the same fixture, which under perfect-info replay equals the bot's own observation ring) relative to Pattern's base bearing. Samples with t+h past the round end are dropped (never a cross-boundary label). Pooled |label err|: mean 15.9°, p50 11.6°, p90 37.2°, p99 52.0°.

Metric — arrival aim error, not accuracy

Harness bmPoint geometry, the relation measure_aim_vs_power.nim validated against the harness resolver: fireDist = |Pattern − self|, arrivalTick = t + ceil(fireDist/v) − 1. A rotation preserves fireDist, so the corrected aim point is rotated around the shooter. miss = |aim − actual enemy pos at arrivalTick|; hit = miss < 18 px. Angular error = |aim bearing − actual bearing|; angular hit = |angErr| < atan(18/range). Timed resolution happens on arrivalTick, which can differ from t+h by ≤ half a tick; the label uses h, the metric uses arrivalTick, exactly as the brief specifies.

The px metric includes range error, and because the head is a pure rotation it cannot fix range; this is why the 18 px hit is dominated by range error and the angular metric is the cleaner measure of a rotation head. Both are reported.

Protocol — prequential (predict-then-learn, streaming)

For every sample the model predicts before it is updated with the label. Two regimes:

  • retained — learning accumulates across all rounds of one battle (fixture); reset only when the battle/enemy changes. This is what the user asked for.
  • perRound — reset at every round boundary (the worst case).

Baselines

  1. straight-line naive — enemy keeps its fire-tick velocity over the same arrivalTick window.
  2. always-the-same-answer — a fixed correction equal to the global mean label (+0.20°).
  3. Pattern — the shipped gun's own prediction (zero correction).

Pooled over all 55732 samples:

predictor meanPx medPx p90Px meanDeg medDeg pxHit% angHit%
Pattern (zero corr) 139.78 112.71 299.10 15.56 11.73 7.0 19.1
straight-line naive 160.17 131.57 336.48 16.01 12.18 5.0 18.8
fixed (+0.20°) 139.76 112.66 298.85 15.56 11.74 7.0 19.0

The AD layer (synthesised for our data)

Random ADs (initRandomAddressDecoder, widths {6, 8, 10, 12}, center = 0), then the deterministic homeostatic controller (accumulateFiring + adaptThresholds, target 1 %). center = 0 is forced by our data: the reference's 127 is the midpoint of 0..255; centring binary 0/1 inputs at 127 makes every synapse contribute ≈ −127 and collapses the ADE code to a mere polarity count, destroying the signal.

The paper's step = 1 controller would need thousands of intervals to find the 1 % operating point — far more than a battle (500–2000 ticks) provides. This is itself the first measured symptom of sample starvation. The analyzer therefore initialises each ADE's threshold at the score that puts it closest to the 1 % firing count (a fast, unsupervised percentile), then runs the deterministic controller (step = 1, 2 passes) to refine it. Achieved firing rates (MEASURED, on the fit stride):

nAde w6 w8 w10 w12 mean
128, pct-init 0.47 % 0.56 % 0.66 % 0.70 % 0.59 %
128, +homeostasis 1.20 % 1.50 % 1.34 % 1.32 % 1.34 %
256, pct-init 0.37 % 0.50 % 0.62 % 0.69 % 0.54 %
256, +homeostasis 1.16 % 1.30 % 1.31 % 1.47 % 1.31 %

So the paper's ~1 % operating point is reached. (The controller alone overshot in an earlier pass at step = 2; step = 1 lands it.)


Sweep — N × AD size × regime (pooled, mean px error)

wm = count-weighted mean, arg = argmax. hit% is the 18 px arrival hit.

config wm meanPx arg meanPx wm pxHit% arg pxHit%
Pattern 139.78 — 7.0 —
straight-line 160.17 — 5.0 —
fixed 139.76 — 7.0 —
nAde128 / N4 / retained 136.93 166.79 4.8 2.4
nAde128 / N4 / perRound 130.27 140.28 4.3 3.0
nAde128 / N8 / perRound 128.87 134.81 4.5 4.7
nAde128 / N16 / perRound 128.78 133.72 5.2 6.7
nAde128 / N32 / perRound 128.83 133.97 5.2 7.5
nAde128 / N64 / perRound 128.97 134.67 5.2 7.6
nAde256 / N8 / perRound 124.78 125.60 4.3 4.7
nAde256 / N16 / perRound 124.72 124.46 4.8 7.3
nAde256 / N32 / perRound 124.95 124.62 4.9 8.3
nAde256 / N64 / perRound 125.13 125.34 5.0 8.5
nAde256 / N16 / retained 134.31 149.00 5.4 5.6
nAde256 / N32 / retained 134.23 149.43 5.3 6.5
nAde256 / N64 / retained 134.30 151.60 5.3 6.5

Full per-config degrees/p90/angular-hit rows are in measure_bitbrain_gate_results.txt.

Where the error stops falling, and why

  • N: the wm error is flat from N=8 to N=64 (≈124.7–125.1 px); the arg error falls to N=16–32 then flattens. Optimum N ≈ 16–32. Beyond it the correction classes get finer than the loop can resolve, and the SBC coincidence cells are already too few to constrain their class bits — more classes only split the same evidence.
  • AD size: nAde=256 beats 128 by a modest ~3 % in wm; both are far from saturating, but the classes saturate first. Doubling the ADE count does not double the information.
  • Why it stops: MNIST needed ~60 000 examples for 10 mutually exclusive classes. Here we have ~55 000 samples for 16–64 correction classes whose evidence must separate by 1–2° — i.e. ~2 orders of magnitude less evidence per class. The SBC is idempotent (a cell accumulates every class that ever co-occurred, with no decay), so with too few examples per cell the per-class counts blur toward uniform and the readout regresses toward the mean. That is sample starvation / memory saturation, and it is consistent with every other observation (flat N tail, weak nAde scaling, retained < perRound).

Count-weighted mean vs argmax (the brief's hypothesis)

The brief expected the count-weighted mean to be the key readout because the TM's discarded magnitude. MEASURED, that is only half right:

  • The wm is the shrinkage readout: it reduces mean error (and extreme misses) but lowers the hit rate (pooled pxHit 4.8 % vs Pattern 7.0 %, angular hit 14.9 % vs 19.1 %). It never makes a confident, sharp correction.
  • The argmax is the decision readout: it keeps the same mean-error reduction and improves the hit rate (pxHit 8.3 %, angular hit 26.1 %). It is the readout that beats Pattern on all four metrics.

So the fine-grained head works, but as a classifier (argmax), not as a soft regression (weighted mean). The weighted mean is a useful control: it is the readout whose shuffled-label null collapses to the baseline.

Retained vs per-round reset

Per-round reset beats retained across rounds on every readout and every N (retained wm ≈ 134.2 px, retained arg ≈ 149–162 px; per-round wm ≈ 124.7, per-round arg ≈ 124.5). The user wants retention across the battle; the measurement says the idempotent SBC accumulates stale, conflicting class bits across rounds and the extra evidence hurts. This mirrors the project's earlier TM finding ("forgetting is stronger than accumulation"). A viable gun would need a bounded/decaying SBC, which the library does not have.

Per-fixture breakdown (perRound, N=32, argmax)

fixture (n) Pattern meanPx / pxHit% / angHit% BitBrain arg meanPx / pxHit% / angHit%
corners (2173) 164.76 / 3.9 / 18.5 132.64 / 3.6 / 23.8
crazy (11209) 126.34 / 7.3 / 25.8 111.95 / 5.9 / 24.6
modularbot (19548) 151.80 / 3.3 / 10.5 134.46 / 7.5 / 25.1
shield (12308) 144.28 / 6.6 / 14.1 124.98 / 10.5 / 28.5
spinbot (10494) 121.29 / 15.0 / 33.7 117.72 / 10.6 / 27.3

Mean error improves on all five. Hit rate improves on modularbot and shield (where Pattern is weak) and regresses on spinbot and crazy (where Pattern is strong), with corners a wash. That is the whole verdict in one table.

Shuffled-label control (must collapse)

Labels permuted across all samples (3 seeds), same inputs:

regime wm shuffled meanPx / pxHit% arg shuffled meanPx / pxHit%
retained 141.51 / 5.8 200.10 / 2.4
perRound 146.74 / 4.6 193.22 / 2.4

The weighted mean collapses toward the baseline (141.5 vs Pattern 139.8) — expected, because it shrinks to the (near-zero) label mean. The argmax does not collapse to the baseline: its null is worse than the baseline — with no signal it still makes a confident, essentially random rotation, which is worse than no correction. That is the correct null behaviour for a non-shrinking readout, and it is why the honest control is real vs shuffled within the same readout: BitBrain argmax is ~124.6 px on real labels vs ~193 px on shuffled labels. The learning is real; the signal is not an artifact.


Bottleneck and recommendation (MEASURED)

Ranked by how much each could plausibly close the gap:

  1. Sample starvation / SBC memory saturation — the dominant one. N saturates at ~16–32, nAde barely scales, and per-round reset beats retention. The library has no bounded/decaying SBC, so a long battle only blurs.
  2. Fixture-dependent gain. The pooled hit win is carried by the weak-Pattern fixtures. Without an online per-fixture selector, a blanket substitution would lose on spinbot/crazy.
  3. AD synthesis is not the bottleneck. The ~1 % operating point is reached and the shuffle control shows the ADs are informative. The forced center = 0 for binary inputs is a correctness requirement, not a defect.

Do not build the gun yet. The cheap decisive next step, if pursued, is a bounded/decaying SBC (a per-round or recency-weighted memory) plus an online selection gate that keeps Pattern where BitBrain is worse — the only shape the data supports. A wider class range or a larger nAde will not fix the starvation.


All numbers MEASURED by common_libs/tests/measure_bitbrain_gate.nim on this machine, deterministic (fixed seeds). Arrival geometry is the harness bmPoint relation validated in measure_aim_vs_power.nim. The fixtures are open-loop; treat the hit rates as prediction-quality evidence only.