Files
SirRoboGarage/docs/bitbrain_gate_test.md

312 lines
15 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# BitBrain gate test — can ADE+SBC predict a fine-grained aim correction?
**Question.** Given the already-built generic BitBrain library
(`common_libs/bitbrain/`, commit `77e6dac`, which reproduced the reference C on
MNIST to the digit), is there signal in a **fine-grained angular aim
correction** that neither a naive predictor nor the shipped Pattern gun already
has? If not, we stop before writing a gun.
**Scope.** Offline only. No gun wiring, no battle, no server, no rack
registration, no changed defaults. Fixtures are read-only.
**Tooling (the evidence):**
- `common_libs/tests/measure_bitbrain_gate.nim` — the analyzer.
- `common_libs/tests/measure_bitbrain_gate_results.txt` — its full deterministic
output.
- `common_libs/bitbrain/` — the library under test (untouched).
Reproduce:
```bash
nim c -r -d:release --nimcache:/tmp/nc_j92 --path:common_libs \
common_libs/tests/measure_bitbrain_gate.nim
# optional knobs: BB_POWER, BB_PMAX, BB_TARGET, BB_PASSES, BB_STEP, BB_STRIDE,
# BB_NADE, BB_NLIST, BB_FILES
```
---
## Direct answer
**There is signal, but it does not clear the bar as a shipping gun.**
- **vs straight-line naive:** BitBrain wins decisively on **every** readout and
every configuration (pooled mean arrival error **124.7 px vs 160.2 px**).
- **vs the shipped Pattern gun:** pooled, the best BitBrain configuration
(per-round reset, argmax readout, N=32, nAde=256) beats Pattern on **mean
arrival error** (124.6 px vs 139.8 px, −10.8 %; 13.3° vs 15.6°, −14.8 %) and
on **both hit rates** (18 px hit 8.3 % vs 7.0 %; angular hit 26.1 % vs
19.1 %). **However** the hit-rate gain is **not robust across fixtures**: it
is concentrated in the two fixtures where Pattern is weak (modularbot,
modularbot_shield) and **BitBrain loses hit rate on the two fixtures where
Pattern is strongest** (spinbot, crazy). It also only works in the
**per-round-reset** regime; the **retained-across-rounds** regime the user
actually wants is the weakest (it improves average error slightly but lowers
the hit rate).
- **Bottleneck (MEASURED):** sample starvation / SBC memory saturation, not the
AD synthesis and not an absence of signal.
A sword that helps exactly where the incumbent is already weak, and hurts where
it is strong, is not a gun improvement. The honest verdict is **signal yes,
shippable improvement no (yet)** — see the bottleneck section.
---
## What was measured (MEASURED unless tagged INFERRED)
### Data
The 5 committed Tank-Royale bridge fixtures `tools/fixtures/tr_drussgt_vs_*`
(open-loop replay), read-only:
| fixture | ticks | resolved samples |
|---|---:|---:|
| `tr_drussgt_vs_corners.jsonl` | 2575 | 2173 |
| `tr_drussgt_vs_crazy.jsonl` | 11507 | 11209 |
| `tr_drussgt_vs_modularbot.jsonl` | 20026 | 19548 |
| `tr_drussgt_vs_modularbot_shield.jsonl` | 12629 | 12308 |
| `tr_drussgt_vs_spinbot.jsonl` | 10824 | 10494 |
| **total** | | **55732** (55 rounds, ~1013 samples/round) |
The fixtures are **open-loop**: the recorded enemy does not react to us. That is
acceptable *here* because this test measures **single-tick prediction quality**,
the one category where fixture replay reproduces live behaviour faithfully. It
would **not** be acceptable evidence for a movement or adaptation claim. Do not
over-read the hit numbers as live hit rates.
### Input — the TMHorizon 53 bits, reused, not re-derived
The analyzer drives the **live `TmHorizonGun`** one tick at a time and reads its
own exported builder `tmhBaseBits` (49 draft bits) plus the 4-bit horizon
one-hot via `tmhLits` — the exact `cachedBits` per-tick path, with
`g.resetRoundState()` called at each round boundary to mirror the live
`onRoundStarted`. (The brief calls this `tmhBuildBits`; the actual symbol is
`tmhBaseBits`.) This keeps the comparison apples-to-apples with the TM gun
already measured.
### Output — fine-grained angular correction class
The correction is the bearing offset added to Pattern's prediction. N class
bins are laid over a fixed **±40°** range (class width 80°/N). Two readouts:
- **wm** = the **count-weighted mean** of the class centres, weighted by the
per-class set-bit counts summed over the 6 cross-AD SBCs. No evidence → 0
correction (i.e. Pattern).
- **arg** = the argmax class centre; no evidence → 0.
### Label — the +h-tick fact (never across a round)
`h = tmhHorizonFor(dist, speed) = clamp(round(dist/(20−3·power)), 10, 50)`.
At fire tick `t`, the label is the actual angular offset of the enemy at
`t+h` (from the same fixture, which under perfect-info replay equals the bot's
own observation ring) relative to Pattern's base bearing. Samples with
`t+h` past the round end are dropped (never a cross-boundary label). Pooled
`|label err|`: mean 15.9°, p50 11.6°, p90 37.2°, p99 52.0°.
### Metric — arrival aim error, not accuracy
Harness `bmPoint` geometry, the relation `measure_aim_vs_power.nim` validated
against the harness resolver: `fireDist = |Pattern − self|`,
`arrivalTick = t + ceil(fireDist/v) − 1`. A rotation preserves `fireDist`, so
the corrected aim point is rotated around the shooter. `miss = |aim − actual
enemy pos at arrivalTick|`; hit = `miss < 18 px`. Angular error = `|aim
bearing − actual bearing|`; angular hit = `|angErr| < atan(18/range)`. Timed
resolution happens on `arrivalTick`, which can differ from `t+h` by ≤ half a
tick; the label uses `h`, the metric uses `arrivalTick`, exactly as the brief
specifies.
**The px metric includes range error**, and because the head is a pure
rotation it cannot fix range; this is why the 18 px hit is dominated by range
error and the angular metric is the cleaner measure of a rotation head. Both
are reported.
### Protocol — prequential (predict-then-learn, streaming)
For every sample the model predicts **before** it is updated with the label.
Two regimes:
- **retained** — learning accumulates across all rounds of one battle
(fixture); reset only when the battle/enemy changes. This is what the user
asked for.
- **perRound** — reset at every round boundary (the worst case).
### Baselines
1. **straight-line naive** — enemy keeps its fire-tick velocity over the same
`arrivalTick` window.
2. **always-the-same-answer** — a fixed correction equal to the global mean
label (+0.20°).
3. **Pattern** — the shipped gun's own prediction (zero correction).
Pooled over all 55732 samples:
| predictor | meanPx | medPx | p90Px | meanDeg | medDeg | pxHit% | angHit% |
|---|---:|---:|---:|---:|---:|---:|---:|
| Pattern (zero corr) | 139.78 | 112.71 | 299.10 | 15.56 | 11.73 | 7.0 | 19.1 |
| straight-line naive | 160.17 | 131.57 | 336.48 | 16.01 | 12.18 | 5.0 | 18.8 |
| fixed (+0.20°) | 139.76 | 112.66 | 298.85 | 15.56 | 11.74 | 7.0 | 19.0 |
---
## The AD layer (synthesised for our data)
Random ADs (`initRandomAddressDecoder`, widths {6, 8, 10, 12}, **center = 0**),
then the deterministic homeostatic controller (`accumulateFiring` +
`adaptThresholds`, target 1 %). **center = 0 is forced by our data:** the
reference's 127 is the midpoint of 0..255; centring **binary** 0/1 inputs at 127
makes every synapse contribute ≈ −127 and collapses the ADE code to a mere
polarity count, destroying the signal.
The paper's `step = 1` controller would need thousands of intervals to find the
1 % operating point — **far more than a battle (500–2000 ticks) provides**.
This is itself the first measured symptom of sample starvation. The analyzer
therefore initialises each ADE's threshold at the score that puts it closest to
the 1 % firing count (a fast, unsupervised percentile), then runs the
deterministic controller (`step = 1`, 2 passes) to refine it. Achieved firing
rates (MEASURED, on the fit stride):
| nAde | w6 | w8 | w10 | w12 | mean |
|---:|---:|---:|---:|---:|---:|
| 128, pct-init | 0.47 % | 0.56 % | 0.66 % | 0.70 % | 0.59 % |
| 128, +homeostasis | 1.20 % | 1.50 % | 1.34 % | 1.32 % | **1.34 %** |
| 256, pct-init | 0.37 % | 0.50 % | 0.62 % | 0.69 % | 0.54 % |
| 256, +homeostasis | 1.16 % | 1.30 % | 1.31 % | 1.47 % | **1.31 %** |
So the paper's ~1 % operating point **is** reached. (The controller alone
overshot in an earlier pass at `step = 2`; `step = 1` lands it.)
---
## Sweep — N × AD size × regime (pooled, mean px error)
`wm` = count-weighted mean, `arg` = argmax. `hit%` is the 18 px arrival hit.
| config | wm meanPx | arg meanPx | wm pxHit% | arg pxHit% |
|---|---:|---:|---:|---:|
| Pattern | 139.78 | — | 7.0 | — |
| straight-line | 160.17 | — | 5.0 | — |
| fixed | 139.76 | — | 7.0 | — |
| nAde128 / N4 / retained | 136.93 | 166.79 | 4.8 | 2.4 |
| nAde128 / N4 / perRound | 130.27 | 140.28 | 4.3 | 3.0 |
| nAde128 / N8 / perRound | 128.87 | 134.81 | 4.5 | 4.7 |
| nAde128 / N16 / perRound | 128.78 | 133.72 | 5.2 | 6.7 |
| nAde128 / N32 / perRound | 128.83 | 133.97 | 5.2 | 7.5 |
| nAde128 / N64 / perRound | 128.97 | 134.67 | 5.2 | 7.6 |
| nAde256 / N8 / perRound | 124.78 | 125.60 | 4.3 | 4.7 |
| nAde256 / N16 / perRound | 124.72 | 124.46 | 4.8 | 7.3 |
| **nAde256 / N32 / perRound** | 124.95 | **124.62** | 4.9 | **8.3** |
| nAde256 / N64 / perRound | 125.13 | 125.34 | 5.0 | 8.5 |
| nAde256 / N16 / retained | 134.31 | 149.00 | 5.4 | 5.6 |
| nAde256 / N32 / retained | 134.23 | 149.43 | 5.3 | 6.5 |
| nAde256 / N64 / retained | 134.30 | 151.60 | 5.3 | 6.5 |
Full per-config degrees/p90/angular-hit rows are in
`measure_bitbrain_gate_results.txt`.
### Where the error stops falling, and why
- **N:** the `wm` error is flat from N=8 to N=64 (≈124.7–125.1 px); the `arg`
error falls to N=16–32 then flattens. **Optimum N ≈ 16–32.** Beyond it the
correction classes get finer than the loop can resolve, and the SBC
coincidence cells are already too few to constrain their class bits — more
classes only split the same evidence.
- **AD size:** nAde=256 beats 128 by a modest ~3 % in `wm`; both are far from
saturating, but the classes saturate first. Doubling the ADE count does not
double the information.
- **Why it stops:** MNIST needed ~60 000 examples for **10** mutually exclusive
classes. Here we have ~55 000 samples for **16–64** correction classes whose
evidence must separate by 1–2° — i.e. ~2 orders of magnitude less evidence per
class. The SBC is idempotent (a cell accumulates *every* class that ever
co-occurred, with no decay), so with too few examples per cell the per-class
counts blur toward uniform and the readout regresses toward the mean. That is
**sample starvation / memory saturation**, and it is consistent with every
other observation (flat N tail, weak nAde scaling, retained < perRound).
### Count-weighted mean vs argmax (the brief's hypothesis)
The brief expected the **count-weighted mean** to be the key readout because
the TM's discarded magnitude. **MEASURED, that is only half right:**
- The `wm` is the **shrinkage** readout: it reduces *mean* error (and extreme
misses) but **lowers the hit rate** (pooled pxHit 4.8 % vs Pattern 7.0 %,
angular hit 14.9 % vs 19.1 %). It never makes a confident, sharp correction.
- The `argmax` is the **decision** readout: it keeps the same mean-error
reduction *and* improves the hit rate (pxHit 8.3 %, angular hit 26.1 %). It is
the readout that beats Pattern on all four metrics.
So the fine-grained head works, but as a **classifier** (argmax), not as a soft
regression (weighted mean). The weighted mean is a useful control: it is the
readout whose shuffled-label null collapses to the baseline.
### Retained vs per-round reset
**Per-round reset beats retained across rounds on every readout and every N**
(retained `wm` ≈ 134.2 px, retained `arg` ≈ 149–162 px; per-round `wm` ≈ 124.7,
per-round `arg` ≈ 124.5). The user wants retention across the battle; the
measurement says the idempotent SBC **accumulates stale, conflicting class bits
across rounds** and the extra evidence hurts. This mirrors the project's earlier
TM finding ("forgetting is stronger than accumulation"). A viable gun would
need a bounded/decaying SBC, which the library does not have.
### Per-fixture breakdown (perRound, N=32, argmax)
| fixture (n) | Pattern meanPx / pxHit% / angHit% | BitBrain arg meanPx / pxHit% / angHit% |
|---|---|---|
| corners (2173) | 164.76 / 3.9 / 18.5 | 132.64 / 3.6 / 23.8 |
| crazy (11209) | 126.34 / 7.3 / 25.8 | 111.95 / **5.9** / **24.6** |
| modularbot (19548) | 151.80 / 3.3 / 10.5 | 134.46 / **7.5** / **25.1** |
| shield (12308) | 144.28 / 6.6 / 14.1 | 124.98 / **10.5** / **28.5** |
| spinbot (10494) | 121.29 / 15.0 / 33.7 | 117.72 / **10.6** / **27.3** |
Mean error improves on **all five**. Hit rate improves on modularbot and shield
(where Pattern is weak) and **regresses on spinbot and crazy** (where Pattern is
strong), with corners a wash. That is the whole verdict in one table.
### Shuffled-label control (must collapse)
Labels permuted across all samples (3 seeds), same inputs:
| regime | wm shuffled meanPx / pxHit% | arg shuffled meanPx / pxHit% |
|---|---|---|
| retained | 141.51 / 5.8 | 200.10 / 2.4 |
| perRound | 146.74 / 4.6 | 193.22 / 2.4 |
The **weighted mean collapses toward the baseline** (141.5 vs Pattern 139.8) —
expected, because it shrinks to the (near-zero) label mean. The **argmax does
not collapse to the baseline: its null is worse than the baseline** — with no
signal it still makes a confident, essentially random rotation, which is worse
than no correction. That is the correct null behaviour for a non-shrinking
readout, and it is why the honest control is **real vs shuffled within the same
readout**: BitBrain argmax is ~124.6 px on real labels vs ~193 px on shuffled
labels. The learning is real; the signal is not an artifact.
---
## Bottleneck and recommendation (MEASURED)
Ranked by how much each could plausibly close the gap:
1. **Sample starvation / SBC memory saturation — the dominant one.** N
saturates at ~16–32, nAde barely scales, and per-round reset beats retention.
The library has no bounded/decaying SBC, so a long battle only blurs.
2. **Fixture-dependent gain.** The pooled hit win is carried by the
weak-Pattern fixtures. Without an online per-fixture selector, a blanket
substitution would lose on spinbot/crazy.
3. **AD synthesis is *not* the bottleneck.** The ~1 % operating point is
reached and the shuffle control shows the ADs are informative. The forced
`center = 0` for binary inputs is a correctness requirement, not a defect.
**Do not build the gun yet.** The cheap decisive next step, if pursued, is a
**bounded/decaying SBC** (a per-round or recency-weighted memory) plus an
**online selection gate** that keeps Pattern where BitBrain is worse — the only
shape the data supports. A wider class range or a larger nAde will not fix the
starvation.
---
*All numbers MEASURED by `common_libs/tests/measure_bitbrain_gate.nim` on this
machine, deterministic (fixed seeds). Arrival geometry is the harness `bmPoint`
relation validated in `measure_aim_vs_power.nim`. The fixtures are open-loop;
treat the hit rates as prediction-quality evidence only.*