Files
SirRoboGarage/docs/bitbrain_counted_sbc.md
T

9.7 KiB
Raw Blame History

Counted SBC with forgetting — design and measured evidence

Step 1 of 2 of the "give the SBC a counter and a forgetting mechanism" change. Library: common_libs/bitbrain/sbc.nim, common_libs/bitbrain/bitbrain.nim. Harness: common_libs/tests/measure_counted_sbc.nim. Unit tests: common_libs/tests/test_bitbrain.nim (56 checks, up from 32).

All numbers below are MEASURED on this machine with -d:release, single-threaded, deterministic (fixed seed / fixed permutation), unless tagged INFERRED.


Why (recap of the diagnosed defects)

The default SBC is a set-union: bits: seq[uint32], learn only sets bits. Two consequences:

  1. A bit records that a coincidence went with class k, never how often. The target is stochastic, so the honest object is a probability.
  2. A bit cannot be cleared. docs/bitbrain_gate_test.md (f41cd08) measured that the retained-across-rounds regime was the weakest precisely because the idempotent memory only grows and saturates with noise.

Counted mode replaces the bit with a small counter and adds decay.


Design

Counter. One saturating uint8 per (i, j, class), 0..255. Chosen over packed 4-bit for clarity and speed; the 4-bit cost is reported below as INFERRED. learn increments the observed class's counter (saturating) and returns the number of coincidence cells touched. Repetition is now evidence: learning the same sample ten times gives a count of 10, where the bitset is a no-op after the first.

Decay — global fractional decay every decayEvery learns. On a schedule, every counter is aged once: c -= c shr decayShift. decayShift = 1 is a halving; higher values forget more slowly; 0 disables decay. Chosen over the alternatives for three reasons:

  • Cost. The decay pass is O(cells) but runs once per decayEvery learns, so the per-learn amortised cost is O(cells / decayEvery) and the hot per-tick learn path stays as cheap as a bit-set. A per-cell EMA decays every touched cell on every learn — ~nClasses× more work per coincidence (at MNIST scale that is ~10× the learn cost).
  • True forgetting. It ages cells that are never visited again, which a per-cell EMA cannot (an EMA only decays cells it touches).
  • Simplicity/determinism. No per-cell timestamps, no extra state beyond a learn counter, and the same input stream always produces the same memory.

Readouts.

  • infer — literal "sum the counters per class" (raw frequency sum). Kept for the requested semantics and as the baseline.
  • inferProb — for each observed coincidence, form the per-cell posterior P(class | cell) = count[class] / Σ_k count[k] and sum it per class. This is the recommended counted readout: it is scale-free in the class marginals, so a single high-count cell cannot dominate a majority of low-count cells. The bitset infer is unchanged.

Config (which is which). The default is and remains smBitset; the bitset path's behaviour is exactly unchanged. Counted mode is selected explicitly or at runtime:

knob kind values
TR_BITBRAIN_MODE runtime env bitset (default) / counted
TR_BITBRAIN_DECAY_EVERY runtime env learns between decay passes
TR_BITBRAIN_DECAY_SHIFT runtime env decay strength (0 = off)
-d:bitbrainDecayEvery=N compile-time overrides DefaultDecayEvery (1024)
-d:bitbrainDecayShift=N compile-time overrides DefaultDecayShift (1)

Env is read by envSbcMode / envDecayEvery / envDecayShift; unknown values fall back to the shipped bitset defaults. The -d: defines use {.intdefine.} and were verified to change DefaultDecayShift at compile time.


Memory cost

One uint8 per (i, j, class), so 8× the packed bit tensor, plus the (small) AD term.

configuration bitset SBC counted SBC ADs counted total packed 4-bit (INFERRED)
Reference MNIST: 6 × 2048² × 10 30.0 MiB (31,457,280 B) 240.0 MiB (251,658,240 B) 0.34 MiB 240.3 MiB 120 MiB
Gun-sized: 6 × 512² × 8 1.5 MiB (1,572,864 B) 12.0 MiB (12,582,912 B) 90,112 B 12.1 MiB (12,673,024 B) 6 MiB

The gun-sized byte-per-cell cost is 12 MiB — acceptable for a gun (and the whole 12.1 MiB model fits comfortably next to the rest of a bot). The MNIST 240 MiB is fine for an offline measurement machine but is 8× the bit tensor; packed 4-bit (INFERRED, not implemented) would halve both figures. At MNIST scale, decayEvery = 2000 costs ≈ 0.07–0.08 ms/learn amortised (measured train 28–30 s vs 24–25 s over 60k for the decay arm).


MNIST regression (the proof the library still matches the reference)

test_bitbrain_mnist.nim was run unchanged against the pretrained ADs and full MNIST:

reader this library reference
Corrected (clean-room) 97.210 % 97.210
Bug-compatible (i % 32 < 8) 96.540 % 96.540

Both anchors reproduce to the digit, and the fixed read_from_sbc truncation bug is not regressed. test_bitbrain is 56 checks, 0 failures (the original 32 checks still pass unchanged, plus 24 new counted-mode checks).


Experiment A — non-stationary adaptation (the forgetting proof)

The same 20,000 MNIST training images are streamed twice: pass 1 with the true labels, pass 2 with the labels permuted by the fixed π = [7,2,9,0,4,6,1,8,3,5]. Evaluation is on the first 2,000 test images, under the "old" (true) labels and the "new" (π) labels. Because the same images recur, the coincidence cells are revisited under the new mapping — exactly the case where forgetting must overwrite stale associations. All arms use infer (raw argmax).

arm old@pass1 new@pass1 old@pass2 new@pass2
bitset (default) 94.05 11.10 55.75 48.90
counted, no decay 85.40 11.55 49.35 43.35
counted + decay (every=2000, shift=1) 88.45 11.10 12.60 86.20

New-task accuracy as pass 2 proceeds (samples seen in pass 2):

arm 2500 5000 7500 10000 12500 15000 17500 20000
bitset 11.9 12.8 13.8 15.7 17.4 22.6 31.4 48.9
counted, no decay 11.9 12.4 13.1 13.8 14.7 17.2 21.3 43.4
counted + decay 41.7 76.2 82.8 82.8 83.7 83.5 84.9 86.2

The bitset climbs only to ~49 % and the counters without decay do not forget at all (~43 %): both are stuck with the pass-1 associations. Counted+decay tracks the change and reaches 86.2 % on the new task (and its old-task accuracy falls to 12.6 %, i.e. it genuinely abandoned the old mapping). This is the measured proof of the mechanism diagnosed in the gate test.


Experiment B — probabilities (rare but predictable vs common but noisy)

A synthetic SBC-level stream (256 columns, row = [0], so cells are the active columns). Class 1 is rare (20 %) but predictable: its 16-cell signature is always active. Class 0 is common (80 %) but noisy: every sample activates a random 2 % of all columns, so class 0 slowly sets a bit / lays a small count on almost every cell. Tested on pure class-1 and pure class-0 inputs.

readout recall(class 0) recall(class 1) balanced
bitset / vote (infer) 1.000 0.000 0.500
counted / raw sum (infer) 0.911 1.000 0.956
counted / per-cell posterior (inferProb) 0.996 1.000 0.998
bitset / posterior (inferProb) 1.000 0.000 0.500

The set-bit vote gives the common class a full vote at every cell it ever touched, so it never finds the rare class (balanced 0.500). Summing counters finds it, and the per-cell posterior beats the raw sum (0.998 vs 0.956): the raw sum lets one high-count cell dominate, the posterior normalises per cell.


Experiment C — does counting cost anything when the data is stationary?

Reference setup, one online pass over the full 60k MNIST train set, evaluated on the first 2,000 test images (the bitset arm is run through the same harness so the comparison is apples-to-apples):

arm vote-argmax % prob-argmax % SBC bytes
bitset 95.700 95.850 31,457,280
counted, no decay 87.700 93.450 251,658,240
counted + decay 88.150 94.900 251,658,240

The full-10k bitset anchor is 97.210 % (the 2,000-image subset is simply harder). Counting hurts on stationary MNIST. The raw counter sum loses ~8 points (87.7 vs 95.7); the per-cell posterior recovers most of it (93.5 / 94.9) but still trails the bitset by ~1–2 points. This is not a saturation artifact: the same gap appears at 2,000 training samples (bitset 89.65 vs counted-no-decay 80.35 vote / 88.60 prob) and 8,000 samples (92.90 vs 83.35 / 91.70), where per-cell counts are far from 255.

So the honest reading is: counting+decay is a trade, not a free win — it buys forgetting and true probabilities at the cost of roughly a point of stationary accuracy (with the recommended posterior readout) and 8× the memory.


Direct answer

  • Forgetting: YES, MEASURED. After a label permutation the bitset is stuck at 48.9 % on the new task and no-decay counters at 43.4 %, while counters+decay reaches 86.2 %. The global decay bound is what makes the memory adaptive.
  • Probabilities: YES, MEASURED. Summing counters resolves a rare but predictable class that the set-bit vote cannot (balanced 0.956–0.998 vs 0.500). The per-cell posterior (inferProb) is the readout to use, and it beats the raw counter sum.
  • Stationary cost: real, MEASURED. Counting does not help MNIST; the raw sum loses ~8 points and the posterior readout ~1–2 points versus the bitset. Bitset remains the default for exactly this reason.