9.7 KiB
Counted SBC with forgetting — design and measured evidence
Step 1 of 2 of the "give the SBC a counter and a forgetting mechanism" change.
Library: common_libs/bitbrain/sbc.nim, common_libs/bitbrain/bitbrain.nim.
Harness: common_libs/tests/measure_counted_sbc.nim.
Unit tests: common_libs/tests/test_bitbrain.nim (56 checks, up from 32).
All numbers below are MEASURED on this machine with -d:release,
single-threaded, deterministic (fixed seed / fixed permutation), unless tagged
INFERRED.
Why (recap of the diagnosed defects)
The default SBC is a set-union: bits: seq[uint32], learn only sets bits.
Two consequences:
- A bit records that a coincidence went with class
k, never how often. The target is stochastic, so the honest object is a probability. - A bit cannot be cleared.
docs/bitbrain_gate_test.md(f41cd08) measured that the retained-across-rounds regime was the weakest precisely because the idempotent memory only grows and saturates with noise.
Counted mode replaces the bit with a small counter and adds decay.
Design
Counter. One saturating uint8 per (i, j, class), 0..255. Chosen over
packed 4-bit for clarity and speed; the 4-bit cost is reported below as
INFERRED. learn increments the observed class's counter (saturating) and
returns the number of coincidence cells touched. Repetition is now evidence:
learning the same sample ten times gives a count of 10, where the bitset is a
no-op after the first.
Decay — global fractional decay every decayEvery learns. On a schedule,
every counter is aged once: c -= c shr decayShift. decayShift = 1 is a
halving; higher values forget more slowly; 0 disables decay. Chosen over the
alternatives for three reasons:
- Cost. The decay pass is O(cells) but runs once per
decayEverylearns, so the per-learn amortised cost is O(cells / decayEvery) and the hot per-ticklearnpath stays as cheap as a bit-set. A per-cell EMA decays every touched cell on every learn — ~nClasses× more work per coincidence (at MNIST scale that is ~10× the learn cost). - True forgetting. It ages cells that are never visited again, which a per-cell EMA cannot (an EMA only decays cells it touches).
- Simplicity/determinism. No per-cell timestamps, no extra state beyond a learn counter, and the same input stream always produces the same memory.
Readouts.
infer— literal "sum the counters per class" (raw frequency sum). Kept for the requested semantics and as the baseline.inferProb— for each observed coincidence, form the per-cell posteriorP(class | cell) = count[class] / Σ_k count[k]and sum it per class. This is the recommended counted readout: it is scale-free in the class marginals, so a single high-count cell cannot dominate a majority of low-count cells. The bitsetinferis unchanged.
Config (which is which). The default is and remains smBitset; the bitset
path's behaviour is exactly unchanged. Counted mode is selected explicitly or at
runtime:
| knob | kind | values |
|---|---|---|
TR_BITBRAIN_MODE |
runtime env | bitset (default) / counted |
TR_BITBRAIN_DECAY_EVERY |
runtime env | learns between decay passes |
TR_BITBRAIN_DECAY_SHIFT |
runtime env | decay strength (0 = off) |
-d:bitbrainDecayEvery=N |
compile-time | overrides DefaultDecayEvery (1024) |
-d:bitbrainDecayShift=N |
compile-time | overrides DefaultDecayShift (1) |
Env is read by envSbcMode / envDecayEvery / envDecayShift; unknown values
fall back to the shipped bitset defaults. The -d: defines use {.intdefine.}
and were verified to change DefaultDecayShift at compile time.
Memory cost
One uint8 per (i, j, class), so 8× the packed bit tensor, plus the (small)
AD term.
| configuration | bitset SBC | counted SBC | ADs | counted total | packed 4-bit (INFERRED) |
|---|---|---|---|---|---|
| Reference MNIST: 6 × 2048² × 10 | 30.0 MiB (31,457,280 B) | 240.0 MiB (251,658,240 B) | 0.34 MiB | 240.3 MiB | 120 MiB |
| Gun-sized: 6 × 512² × 8 | 1.5 MiB (1,572,864 B) | 12.0 MiB (12,582,912 B) | 90,112 B | 12.1 MiB (12,673,024 B) | 6 MiB |
The gun-sized byte-per-cell cost is 12 MiB — acceptable for a gun (and the
whole 12.1 MiB model fits comfortably next to the rest of a bot). The MNIST
240 MiB is fine for an offline measurement machine but is 8× the bit tensor;
packed 4-bit (INFERRED, not implemented) would halve both figures. At
MNIST scale, decayEvery = 2000 costs ≈ 0.07–0.08 ms/learn amortised (measured
train 28–30 s vs 24–25 s over 60k for the decay arm).
MNIST regression (the proof the library still matches the reference)
test_bitbrain_mnist.nim was run unchanged against the pretrained ADs and
full MNIST:
| reader | this library | reference |
|---|---|---|
| Corrected (clean-room) | 97.210 % | 97.210 |
Bug-compatible (i % 32 < 8) |
96.540 % | 96.540 |
Both anchors reproduce to the digit, and the fixed read_from_sbc truncation
bug is not regressed. test_bitbrain is 56 checks, 0 failures (the original
32 checks still pass unchanged, plus 24 new counted-mode checks).
Experiment A — non-stationary adaptation (the forgetting proof)
The same 20,000 MNIST training images are streamed twice: pass 1 with the true
labels, pass 2 with the labels permuted by the fixed π = [7,2,9,0,4,6,1,8,3,5].
Evaluation is on the first 2,000 test images, under the "old" (true) labels and
the "new" (π) labels. Because the same images recur, the coincidence cells
are revisited under the new mapping — exactly the case where forgetting must
overwrite stale associations. All arms use infer (raw argmax).
| arm | old@pass1 | new@pass1 | old@pass2 | new@pass2 |
|---|---|---|---|---|
| bitset (default) | 94.05 | 11.10 | 55.75 | 48.90 |
| counted, no decay | 85.40 | 11.55 | 49.35 | 43.35 |
counted + decay (every=2000, shift=1) |
88.45 | 11.10 | 12.60 | 86.20 |
New-task accuracy as pass 2 proceeds (samples seen in pass 2):
| arm | 2500 | 5000 | 7500 | 10000 | 12500 | 15000 | 17500 | 20000 |
|---|---|---|---|---|---|---|---|---|
| bitset | 11.9 | 12.8 | 13.8 | 15.7 | 17.4 | 22.6 | 31.4 | 48.9 |
| counted, no decay | 11.9 | 12.4 | 13.1 | 13.8 | 14.7 | 17.2 | 21.3 | 43.4 |
| counted + decay | 41.7 | 76.2 | 82.8 | 82.8 | 83.7 | 83.5 | 84.9 | 86.2 |
The bitset climbs only to ~49 % and the counters without decay do not forget at all (~43 %): both are stuck with the pass-1 associations. Counted+decay tracks the change and reaches 86.2 % on the new task (and its old-task accuracy falls to 12.6 %, i.e. it genuinely abandoned the old mapping). This is the measured proof of the mechanism diagnosed in the gate test.
Experiment B — probabilities (rare but predictable vs common but noisy)
A synthetic SBC-level stream (256 columns, row = [0], so cells are the active
columns). Class 1 is rare (20 %) but predictable: its 16-cell signature is
always active. Class 0 is common (80 %) but noisy: every sample activates a
random 2 % of all columns, so class 0 slowly sets a bit / lays a small count on
almost every cell. Tested on pure class-1 and pure class-0 inputs.
| readout | recall(class 0) | recall(class 1) | balanced |
|---|---|---|---|
bitset / vote (infer) |
1.000 | 0.000 | 0.500 |
counted / raw sum (infer) |
0.911 | 1.000 | 0.956 |
counted / per-cell posterior (inferProb) |
0.996 | 1.000 | 0.998 |
bitset / posterior (inferProb) |
1.000 | 0.000 | 0.500 |
The set-bit vote gives the common class a full vote at every cell it ever touched, so it never finds the rare class (balanced 0.500). Summing counters finds it, and the per-cell posterior beats the raw sum (0.998 vs 0.956): the raw sum lets one high-count cell dominate, the posterior normalises per cell.
Experiment C — does counting cost anything when the data is stationary?
Reference setup, one online pass over the full 60k MNIST train set, evaluated on the first 2,000 test images (the bitset arm is run through the same harness so the comparison is apples-to-apples):
| arm | vote-argmax % | prob-argmax % | SBC bytes |
|---|---|---|---|
| bitset | 95.700 | 95.850 | 31,457,280 |
| counted, no decay | 87.700 | 93.450 | 251,658,240 |
| counted + decay | 88.150 | 94.900 | 251,658,240 |
The full-10k bitset anchor is 97.210 % (the 2,000-image subset is simply harder). Counting hurts on stationary MNIST. The raw counter sum loses ~8 points (87.7 vs 95.7); the per-cell posterior recovers most of it (93.5 / 94.9) but still trails the bitset by ~1–2 points. This is not a saturation artifact: the same gap appears at 2,000 training samples (bitset 89.65 vs counted-no-decay 80.35 vote / 88.60 prob) and 8,000 samples (92.90 vs 83.35 / 91.70), where per-cell counts are far from 255.
So the honest reading is: counting+decay is a trade, not a free win — it buys forgetting and true probabilities at the cost of roughly a point of stationary accuracy (with the recommended posterior readout) and 8× the memory.
Direct answer
- Forgetting: YES, MEASURED. After a label permutation the bitset is stuck at 48.9 % on the new task and no-decay counters at 43.4 %, while counters+decay reaches 86.2 %. The global decay bound is what makes the memory adaptive.
- Probabilities: YES, MEASURED. Summing counters resolves a rare but
predictable class that the set-bit vote cannot (balanced 0.956–0.998 vs
0.500). The per-cell posterior (
inferProb) is the readout to use, and it beats the raw counter sum. - Stationary cost: real, MEASURED. Counting does not help MNIST; the raw sum loses ~8 points and the posterior readout ~1–2 points versus the bitset. Bitset remains the default for exactly this reason.