Files
SirStone 40ba96f649 BitBrain SBC: counted mode + global decay (forgetting, probabilities)
Adds an smCounted storage mode alongside the default smBitset. Each
(i,j,class) cell becomes a saturating uint8 counter; learn increments it and
a global fractional decay (c -= c shr decayShift every decayEvery learns)
makes forgetting possible. infer sums raw counters; new inferProb sums the
per-cell posterior P(class|cell) (scale-free, recommended readout).

Bitset path is the default and byte-for-byte unchanged: test_bitbrain 56/56
(was 32), and test_bitbrain_mnist reproduces 97.210% corrected / 96.540%
bug-compatible exactly.

Counted mode configurable at runtime (TR_BITBRAIN_MODE / TR_BITBRAIN_DECAY_*)
and compile time (-d:bitbrainDecay*). Measured: forgetting (86.2% vs 48.9% on
a permuted-label stream), probabilities (rare-class balanced 0.998 vs 0.500),
and the stationary cost (counted hurts MNIST; see docs/bitbrain_counted_sbc.md).

Harness: common_libs/tests/measure_counted_sbc.nim
2026-09-25 08:39:10 +02:00
..

BitBrain (ADE + SBC) — generic Nim library

A clean-room Nim implementation of the classifier in

BitBrain and Sparse Binary Coincidence (SBC) memories, Frontiers in Neuroinformatics 17:1125844, 2023.

This is a library, not a gun. There is no I/O design, no battle wiring and no environment knobs here yet — input and output formats are to be agreed separately.

Clean-room note

The implementation was written from the published algorithm description only. No code was copied from the reference C program (full_mnist_2048.c) shipped with the paper; that file is GPL-3.0-or-later, © 2022 The University of Manchester, and copying it would impose that licence on this repository. The published defaults (scale = 64, centre = 127) make the scoring rule numerically identical to the reference, which the MNIST acceptance test confirms to the digit.

Mechanism

  1. ADE (Address Decoder Element) — a sparse, signed, thresholded random projection. Each ADE has width synapses (inputIndex, polarity) and scores an input as

    raw   = Σ_j polarity_j * (input[inputIndex_j] - center)
    score = scale * raw
    

    firing iff score >= threshold. Defaults scale = 64, center = 127 reproduce the reference exactly. Multi-width ADs (the paper's best setup uses widths {6, 8, 10, 12}) detect features at different scales.

  2. Homeostatic threshold learning (optional, unsupervised) — accumulate each ADE's firing count over inputs, then nudge its threshold toward a target firing rate (~1% in the paper). accumulateFiring + adaptThresholds implement the paper's deterministic controller. The paper's additional Hebbian longevity step (retire the weakest synapse, resample a new input index) and its Metropolis–Hastings input-position sampling are described in ade.nim but not implemented — initRandomAddressDecoder uses the paper's uniform random initialisation.

  3. SBC memory (supervised) — two ADs on the axes; a pair of simultaneously firing ADEs (i, j) is a coincidence indexing one cell holding a class bitmask. learn sets the class bit and is idempotent: setting it again is a no-op. No clearing, no learning rate, no decay, no epochs. infer accesses the same locations but counts set bits per class.

  4. BitBrain container — several ADs (possibly different widths) plus several SBCs built from pairs of them. Inference aggregates class counts across SBCs; argmax wins (ties to the lowest class index).

API

import bitbrain/bitbrain   # re-exports ade + sbc

# --- construction ---
var bb = buildRandomBitBrain(
  widths    = @[6, 8, 10, 12],  # one AD per width
  nAde      = 512,              # ADEs per AD
  inputWidth = 256,             # input vector length
  nClasses  = 8,
  seed      = 1234'i64)         # deterministic

# or build ADs yourself and assemble:
# var ad = initAddressDecoder(nAde = 2048, width = 6)
# ad.codes = ...                  # signed 1-based input codes (pretrained)
# ad.thresholds = ...
# var bb = initBitBrain(@[ad, ...], crossPairs(4) & withinPairs(4), nClasses = 10)

# --- online by construction (order-free, repeatable, interleaved) ---
bb.learn(input, class)                       # sets class bits, idempotent
let (label, counts) = bb.infer(input)        # argmax + per-class counts
bb.resetLearning()                           # wipe SBCs (ADs unchanged)

# --- optional unsupervised homeostasis ---
for input in trainingStream:
  ad.accumulateFiring(input)
ad.adaptThresholds(interval = 2000, targetRate = 0.01, step = 1)

# --- accounting ---
bb.memoryBytes        # ADs + SBC tensors
bb.sbcMemoryBytes     # SBC tensors only (the dominant term)

Generic over the input element type (openArray[SomeInteger]) and over the input width, ADE count and class count — nothing is hardcoded to 784/10. Deterministic given the seed. Dependencies: std/ only.

Counted SBC mode (saturating counters + forgetting)

An optional smCounted mode replaces each bit with a saturating uint8 counter and adds global fractional decay (c -= c shr decayShift every decayEvery learns). initBitBrain(..., mode, decayEvery, decayShift) selects it; initCountedSbc builds a single counted memory. infer sums raw counters; inferProb sums the per-cell posterior P(class | cell) (the recommended counted readout). Runtime knobs: TR_BITBRAIN_MODE (bitset default, counted), TR_BITBRAIN_DECAY_EVERY, TR_BITBRAIN_DECAY_SHIFT; compile-time defaults: -d:bitbrainDecayEvery=N, -d:bitbrainDecayShift=N. The default bitset path is unchanged and remains the reference-compatible one. Design, memory cost and the measured forgetting/probability/stationary evidence are in docs/bitbrain_counted_sbc.md.

Reference SBC wiring: crossPairs(4) gives the 6 cross-AD SBCs used by the reference C. withinPairs(n) adds the paper's 4 within-AD ("half-size") SBCs; this implementation stores them full-size (half-size packing is a separate memory optimisation).

Measured results

All figures measured on this machine with -d:release, single-threaded, on the reference fixtures (pretrained ADs + MNIST). Nothing is committed: fixtures live under /tmp/bitbrain/BitBrain_C_code (override with $BITBRAIN_FIXTURES).

MNIST acceptance — exact reproduction of the reference

4 ADs × 2048 ADEs, widths {6, 8, 10, 12}, 6 cross-AD SBCs, 10 classes, one online pass over 60,000 training samples, evaluated on all 10,000 test images:

Reader This library Reference C
Corrected (clean-room, all row ADEs counted) 97.210% 97.210%
Bug-compatible (uint8_t bit_test: only i % 32 < 8) 96.540% 96.540%

The bug-compatible mode reproduces the shipped reference's truncation bug exactly, which proves the ADE scoring, the coincidence indexing and the idempotent SBC rule are all correct. The default corrected reader is the one to use.

Per-sample cost (reference MNIST configuration)

Operation Measured
learn (4 ADs + 6 SBCs, dense active lists) 0.384 ms/sample
infer (corrected reader) 0.630 ms/sample
infer (bug-compatible reader) 0.454 ms/sample

All well inside this project's live budget of ~13.16 ms/tick. The reference C measured ~0.2–0.56 ms/sample for the same geometry.

Memory (measured via sbcMemoryBytes / memoryBytes)

The SBC tensor dominates: nAde × nAde × nClasses bits per SBC, rounded up to uint32 slots.

Configuration SBC tensors ADs Total
Reference: 6 SBCs × 2048² × 10 31,457,280 B (30.0 MiB) 360,448 B 31,817,728 B (30.3 MiB)
Gun-sized: 6 SBCs × 512² × 8 1,572,864 B (1.5 MiB) 90,112 B 1,662,976 B (1.59 MiB)
10 SBCs × 2048² × 10 (paper variant) 52,428,800 B (50.0 MiB) 360,448 B 52,789,248 B (50.3 MiB)

The reference configuration is L2/L3-hostile; a battle-sized configuration should size nAde to the feature count, not copy 2048².

Tests

# Unit / sanity suite (no fixtures required, ~3 s)
nim c -r --nimcache:/tmp/nc_j90 -d:release \
  --path:common_libs common_libs/tests/test_bitbrain.nim

# MNIST acceptance harness (needs fixtures in /tmp; skips cleanly if absent)
nim c -r --nimcache:/tmp/nc_j90 -d:release --path:common_libs \
  -o:/tmp/acceptance_bitbrain_mnist common_libs/tests/test_bitbrain_mnist.nim

The unit suite covers: hand-computed ADE scoring and the >= threshold, idempotent SBC learning (learn one sample 1000× → bit-identical memory), a planted rule learned near-perfectly with a monotone online accuracy curve, a shuffled-label control that degrades to chance, unseen-input behaviour, homeostatic threshold adaptation, and memory accounting.