Files
SirRoboGarage/docs/bitbrain_algorithm.md
SirStone c8b2a8c6c5 docs: reverse-engineer the BitBrain algorithm from its C source; assess as a gun
Reads docs/BitBrain_C_code.zip end to end (full_mnist_2048.c 568 lines +
bitarray.h), reconciles every weight/data file's byte size against its
loader, and answers two questions:

1. What the algorithm is: signed thresholded address decoders (random
   projections) -> write-once sparse binary coincidence bit tensors ->
   counting/argmax readout. Thresholds and ADs come from an unsupervised
   program that is NOT in the zip; only the SBC tensors are learned here.
   The supervised rule is genuinely online, single-pass, order-free, and
   has no learning rate - so it satisfies ModularBot's reset-on-enemy-change
   constraint. Found and isolated a real bug: read_from_sbc's uint8_t
   bit_test truncates the 32-bit bit test, so only ADEs with i%%32 < 8 are
   ever counted. Shipped reader: 96.540%% (reproduced exactly); with the bug
   fixed: 97.210%%.

2. Whether it can become a gun: not on this evidence. Measured full
   inference at ~0.56 ms/sample (cost is not the blocker), but the AD layer
   is unobtainable, and this repo has already shown the binding constraint
   is the target signal, not the learner - the TM head scored 2.08pp below
   its own majority class and the AD/SBC primitives are already in the
   16-family shootout as WiSARD/Bloom/SDM. Recommends a readout swap on the
   existing WiSARD feature extractor and a Gate-2b-style >=80%% side-accuracy
   gate before any port.

Also: recommends committing the 11 MB zip under its explicit name.
2026-09-24 20:17:08 +02:00

39 KiB
Raw Permalink Blame History

BitBrain — what the algorithm actually is, and whether it can become a gun

Issue: N/A (read-only research requested by orchestrator) Date: 2026-09-24 Primary source: docs/BitBrain_C_code.zip (11,656,362 B, untracked), extracted to /tmp/bitbrain/BitBrain_C_code/. Everything below cites that tree:

  • full_mnist_2048.c — 568 lines, the whole implementation
  • bitarray.h — 41 lines, bit-array macros (uint32_t slots; BITSET 35, BITCLEAR 36, BITTEST 37, BITNSLOTS 39)
  • pretrained weights AD1_2048 … AD4_2048, thresh1_2048 … thresh4_2048
  • MNIST data train_data … test_label (the classic uint8 MNIST dump)

The C file carries a GNU GPL-3.0-or-later header, © 2022 The University of Manchester (full_mnist_2048.c:1-19). The zip contains no build files, no header, no README, and no unsupervised-learning program.

Verification legend:

Tag Meaning
[MEASURED] I compiled and/or ran the code (or an exact verbatim copy of it) here and read the number off the output.
[VERIFIED] I read the exact line, and/or reconciled it against byte sizes / arithmetic I recomputed.
[INFERRED] My reasoning from code + measurements; not directly observed.
[UNKNOWN] I could not determine it from the primary source.

Unless stated otherwise, every measurement was taken with gcc 15.3.0, -O3, single threaded (OpenMP is off by default — see §4.1), on this machine (16 cores).


0. One-paragraph answer

BitBrain (as shipped here) is a three-stage, purely feed-forward pipeline: (1) a random-projection threshold layer of 2048 "address decoders" (ADs) per layer, across 4 layers with 6/8/10/12 synapses each, whose ~74000 integer (position, sign) weights and 2048×4 integer firing thresholds arrive as two files from a program we do not have (full_mnist_2048.c:376-384, 409-417); (2) a sparse binary coincidence (SBC) memory — six 2048×2048×10 bit tensors (31.5 MB) that are written once per labelled sample, a pure BITSET, never cleared (full_mnist_2048.c:129-152); (3) a counting readout that sums, over the 6 memories, how many of the observed AD-pair coincidences are "known" for each of the 10 labels and takes the argmax (full_mnist_2048.c:157-181, 527-540). The program prints 346 wrong = 96.540 pct correct on all 10,000 MNIST test images [MEASURED], but it does so with a real bug: the readout's uint8_t bit_test truncates the 32-bit bit-test result, so only ADs with i % 32 < 8 are ever counted (full_mnist_2048.c:161,173-175). Removing that truncation gives 97.210% [MEASURED]. Stages 2 and 3 are genuinely online, single-pass, per-sample, order-free and therefore satisfy the ModularBot constraint; stage 1 is the part we cannot rebuild from this source. On the gun question the honest answer is: no, not on this evidence. The learner is cheap (0.56 ms/sample for one full forward pass here [MEASURED]), but this repo has already evaluated every structurally equivalent learner on this target — WiSARD/n-tuple LUTs (which are the AD primitive), Tsetlin, Kanerva SDM, Bloom filters, HDC/VSA — and the binding constraint was the target signal, not the learner (BNNBot_garage/analysis/shootout_results.md, common_libs/tests/tm_pattern_sweep_results.md:611-631, common_libs/guns/tm_horizon.nim:49-53). BitBrain would be the third learner swapped onto that same signal-poor target.


1. What the algorithm is

1.1 The core primitive: one ADE = one signed, thresholded random projection

find_AD_firing_pattern (full_mnist_2048.c:98-124) is the whole of stage 1. For each of the w ADEs in a layer, for each of the n synapses of that ADE:

if( address_decoder[i][j] > 0 ) { yang =  64; position =  address_decoder[i][j]; }   // line 112
else                           { yang = -64; position = -address_decoder[i][j]; }    // line 114
pixel = sensory_input[ loop ][ position-1 ];                                        // line 116
count += ( pixel - 127 ) * yang;                                                    // line 118
...
ad_row_fired[ i ] = ( count >= ad_row_thresh[ i ] ? TRUE : FALSE );                 // line 122

So, precisely [VERIFIED]:

  • An address decoder element (ADE) is described by an n-long vector of non-zero int32_t. The sign is the synapse sign — positive = excitatory, negative = inhibitory (full_mnist_2048.c:90-91, doc comment). The magnitude is the input position, 1-based into the flattened input row (position-1, line 116).
  • count is a signed, 127-centred, x64-scaled sum of the sampled input bytes. yang is only ever ±64, a pure scale factor; it is not a learned weight. So count/64 = Σ_j ±(pixel_j − 127), i.e. each "pixel" is treated as a real in [−127, 128] with 127 as the decision centre. [VERIFIED]
  • The ADE fires iff count >= ad_row_thresh[i] — a ≥ comparison against a per-ADE integer threshold, i.e. the ADE is a thresholded linear unit on a random n-dimensional subspace. [VERIFIED]
  • The output is a dense uint8_t[w] 0/1 vector, ad_row_fired (line 122).

The doc comment calls the input "0.255 for E/MNIST" (full_mnist_2048.c:84) — a typo for "0..255"; the values are raw MNIST uint8 [VERIFIED].

Layer geometry (full_mnist_2048.c:345-352,358): W = 2048 ADEs per layer; four layers with N1..N4 = 6, 8, 10, 12 synapses; SBC_CT = 6 coincidence memories, one for each of the 6 unordered layer pairs: (1,2) (1,3) (1,4) (2,3) (2,4) (3,4) — see the six write_to_sbc calls at 470-480 and the six read_from_sbc calls at 516-525. D = 10 label classes (line 347). Input width INPUT_SZ = 784 (line 356).

This is an n-tuple / RAM-node / WiSARD-style address decoder with a threshold [INFERRED] — i.e. the classic "random receptive field + threshold" feature layer. It is not novel machinery; this repo already has one (BNNBot_garage/src/wisard_predictor.nim, K=14, 50 LUT nodes, ~16384 entries each, averaged readout).

1.2 Sparsity (measured)

Layer n mean ADEs fired / 2048 mean fire rate
1 6 25.4 1.24%
2 8 26.7 1.30%
3 10 28.7 1.40%
4 12 29.8 1.45%

[MEASURED] — 2,000 test images, verbatim find_AD_firing_pattern. Over 20,000 training images the rates are 1.269 / 1.355 / 1.461 / 1.504% and every image fires ≥1 ADE in every layer (so no image is ever "empty") [MEASURED].

The per-ADE rate distribution is very heterogeneous [MEASURED] (20,000 train images):

Layer never fires 1–2% 2–3% ≥10% hottest single ADE
1 778 (38%) 1187 67 7 51.90%
2 656 (32%) 1272 97 9 23.82%
3 592 (29%) 1238 163 13 31.37%
4 522 (25%) 1293 175 17 12.75%

This is not a uniform quantile calibration: a third of the ADEs are dead and a handful are heavy hitters. That is the fingerprint of a selection/pruning process, not a global rule.

1.3 What is trainable HERE vs PRETRAINED ELSEWHERE — the decisive split

[VERIFIED] — reading the file end to end, the only learnable things updated by code in this tree are the bits of the six SBC tensors.

Component Where it comes from Learned in this file?
4 × 2048 firing thresholds (thresh*_2048) file, loaded full_mnist_2048.c:414-417; documented "thresholds from unsupervised learning" (376-379) No
4 × 2048 ADs (AD*_2048) file, loaded full_mnist_2048.c:409-412; documented "ADs from unsupervised learning" (381-384) No
6 × 2048×2048×10 SBC bit tensors calloc'd zero at full_mnist_2048.c:386-391 Yes — write_to_sbc, 129-152, called in the loop at 435-482

There is no code anywhere in the 568 lines that modifies an AD or a threshold. I listed the complete function inventory: find_AD_firing_pattern (98), write_to_sbc (129), read_from_sbc (157), four file loaders (184, 209, 234, 259), three display helpers (284, 296, 310), and main (329) [VERIFIED]. show_thresholds/show_AD are printf-only and are behind a commented-out #define SHOW_LOADED_FILES (420-431). [VERIFIED]

Say it loudly: we cannot regenerate the AD layer from this zip. The unsupervised program that produced AD*_2048 and thresh*_2048 is not present, not referenced, and not described beyond the two comments above. Any BitBrain gun must therefore either ship those exact files or synthesise a substitute AD layer of its own — and the original was trained on 784-pixel MNIST, so the files are useless for any other input domain. There is no third option.

I did, however, characterise the files well enough to say what a substitute would have to look like, and how much of the original's structure is recoverable from the unlabelled input distribution alone.

1.3.1 Structure of the AD files (this is recoverable — partly)

[MEASURED], from the four AD*_2048 files:

  • Magnitudes (|AD|, which is the input position) lie in [37, 776] for every layer, with quasi-linear deciles (L4: min 37, then 206/261/301/355/409/462/514/556/610, max 776). So the magnitudes look like near-uniform random draws over the usable index range.

  • The range itself is explained by dead pixels: pixels 1..36 (top row + part of the second) and 777..784 (bottom-right corner) are inked in 0 of 3,000 training images, and those exact positions are sampled 0 times by every layer [MEASURED]. That is why the minimum |AD| is 37 and the maximum is 776 rather than 1 and 784.

  • Of the 134 pixels that are constant across 3,000 images, 133 / 134 / 127 / 132 are never used by layers 1..4 respectively — i.e. the position pool excludes constant pixels.

  • Position usage tracks input activity almost perfectly: Spearman(ink frequency, AD-slot usage) = 0.994 over the 784 pixels; 75.9% of all AD synapse slots sit on the 219 pixels that are inked in more than a third of images, and the 134 dead pixels get exactly 0.0% [MEASURED].

  • Synapse signs are not 50/50: excitatory fractions are 44.7% / 42.9% / 40.3% / 39.2%

    [MEASURED].

  • Still unexplained: ~93–114 inked pixels are never sampled by each layer, and the 220–247 unused positions per layer have a median ink-count of 0 with max 33–89 (vs a median of 619–698 for used positions) [MEASURED]. There is a real selection rule beyond "drop dead pixels" that I could not reverse-engineer [UNKNOWN].

  • Thresholds are not per-ADE fire-rate quantiles. For each layer, 13–25 ADEs have a threshold above the maximum count that ADE can ever produce, so they can never fire (1.22% / 0.68% / 0.98% / 0.63% of ADEs) [MEASURED]. A quantile calibration on the same data can never produce that. A handful have negative thresholds (thresh1 min −10871, thresh4 min −15496) but no ADE always fires [MEASURED].

Consequence for a port: a substitute AD layer could be synthesised cheaply from unlabelled samples of the new input domain — (a) drop constant positions, (b) sample positions with probability monotone in activity, (c) uniform random ± signs, (d) set each threshold to a quantile of that ADE's own count distribution to hit a target fire rate (~1.3–1.5%) — but it would be a statistically similar, not the same, feature layer, and the only way to validate it is end-to-end accuracy on the target task. There is no apples-to-apples baseline to compare against because the original ADs are MNIST-only. That is the single largest technical risk in the whole idea [INFERRED].

1.4 Binarisation and the count↔bit-width relationship

There is no binarisation before the ADs: sensory_input is the raw uint8 MNIST matrix, 784 bytes per row (full_mnist_2048.c:234-256 loader, 346-356 sizes) [VERIFIED]. The ADE does the analogue-to-threshold conversion itself, via (pixel − 127) * yang (line 118).

The bit-width note: count is int32_t (line 101), and its per-synapse term is int8_t(±64) × (pixel − 127), so one synapse contributes at most 64 × 128 = 8192. With n = 12 the maximum |count| is 98304, which overflows int16_t — hence the int32_t for both count and ad_row_thresh. pixel is int16_t (line 103), which is sufficient for 0..255. [VERIFIED] And indeed the largest threshold in the data, 99053 (thresh4_2048), sits just above that 98304 ceiling [MEASURED].

1.5 The supervised head: write_to_sbc / read_from_sbc

Write (learning) — full_mnist_2048.c:129-152, comment at 128:

FOR_LOOP(i, w) if (first[i])
  FOR_LOOP(j, w) if (second[j]) {
    tens_index = D3(i, j, label);                      // 138: (i, j, class) -> one bit
    if (!BITTEST(mem, tens_index)) { internal_count++; BITSET(mem, tens_index); }   // 142-144
  }

D3(i,j,k) = i + j*wd + k*ws with wd = W*D and ws = W (full_mnist_2048.c:68, globals set at 360), so the bit layout is [j][label][i] with i fastest. [VERIFIED]

Thus the learnable rule is exactly: set the bit at (coincidence of an ADE from layer A firing and an ADE from layer B firing, current label). Nothing is ever decremented or cleared; there is no learning rate, no counter, no decay, no negative evidence [VERIFIED]. The six write_to_sbc calls per sample (train loop 435-482, calls at 470-480) each receive the label as uint8_t (464).

Read (inference) — full_mnist_2048.c:157-181:

FOR_LOOP(i, w) if (first[i])
  FOR_LOOP(j, w) if (second[j])
    FOR_LOOP(k, ds) { tens_index = D3(i, j, k);          // 171
                      bit_test = BITTEST(mem, tens_index);   // 173  <-- uint8_t!
                      if (bit_test) (store[k])++; }         // 175

then the six per-memory count vectors are summed and argmaxed (527-540). So the score for class k is the number of observed AD-pair coincidences whose bit for k is set — an overlap count, not a likelihood and not a probability. [VERIFIED]

Two consequences worth stating plainly:

  1. No normalisation. There is no division by the number of stored features per class, so the score is biased toward whichever class has more bits stored — a class seen more often, or whose features are more diverse, wins ties for free. On balanced MNIST this is mild; on a gun with a few hundred labels it is a real hazard. [INFERRED] from lines 175 + 527-531.
  2. No negative evidence. A coincidence that is never seen for class k contributes 0, identical to a coincidence seen for every class. Frequent coincidences saturate to 10/10 and contribute a constant to all classes (harmless for argmax); rare coincidences carry the entire discrimination. [INFERRED]

1.6 The read_from_sbc bug — found, isolated, and reproduced [MEASURED]

bit_test is declared uint8_t (line 161) but BITTEST yields the full masked 32-bit word (bitarray.h:37). Whenever the hit bit sits at word-position ≥ 8, bit_test truncates to 0 and the hit is silently dropped. Since D3 strides by ws = W = 2048 (a multiple of 32) and wd = W*D = 20480 (also a multiple of 32), the bit's position inside its uint32 word is exactly i % 32. So the shipped reader only counts ADEs with i % 32 < 8 — three quarters of the AD range is invisible at inference time. [VERIFIED] by arithmetic and [MEASURED] by reproduction.

Verification chain (all on the same trained memory, sample 0, layer pair (1,2)):

Counter Value
Python, independent byte-level bit test, full i range 319
Python, same but restricted to i % 32 < 8 50
C, verbatim read_from_sbc 50

The write path is unaffected: write_to_sbc uses !BITTEST(...) in a boolean context (line 142), so no truncation occurs there. [VERIFIED] The bug is read-only and one-sided: 25% of the evidence is used.

End-to-end, on all 10,000 test images after the full 60,000-sample training pass [MEASURED]:

Reader Accuracy ms/sample
read_from_sbc as shipped (with the truncation bug) 96.540% 0.412
read_from_sbc with the truncation removed 97.210% 0.531
sparse active-list reader, no truncation 97.210% 0.202

The 96.540% row exactly reproduces the number the unmodified program prints [MEASURED] (./bb → 346 wrong = 96.540 pct correct, 66 s wall, single-threaded). So the bug is real, it costs 0.67 pp, and it does not invalidate the paper's claim — but anyone porting this should fix it, and anyone benchmarking against "96.5%" should know it is a lower bound.

1.7 Is the supervised rule online / usable incrementally?

Yes, unambiguously [VERIFIED] from 129-152: the update is per-sample, order-independent (it is an idempotent BITSET), single-pass, and needs no batch, no epoch, no learning rate, and no revisit of older samples. Inference is available at any point, including before any training. The only thing that is not cheap is resetting: the memory is monotone, so "forget the old enemy" means clearing 31.5 MB (calloc/memset) and starting over.

The learning curve is monotone and does not collapse [MEASURED] (cumulative single pass over the training set, evaluated on the first 2,000 test images):

training samples per class acc, shipped (buggy) reader acc, correct reader
600 60 79.05% 82.50%
1,200 120 85.05% 86.85%
2,400 240 87.80% 89.75%
4,800 480 90.80% 92.20%
12,000 1,200 93.05% 93.90%
30,000 3,000 94.20% 95.30%
60,000 6,000 95.20% 95.70%

So it needs ~a few hundred labelled samples per class before it is useful and ~2,000+ to approach its ceiling. There is no saturation collapse: at 60,000 samples only 18.9–20.8% of each SBC tensor's bits are set, 70.8–76.3% of (i,j) pairs are non-empty, and only 0.03–0.05% of non-empty pairs have all 10 labels set [MEASURED]. I initially expected monotone saturation to destroy discrimination; it does not, on this data volume.

1.8 Exact dimensions and on-disk format — reconciled against byte sizes

All loaders read one raw contiguous blob, no header, no magic, no dimension fields, native-endian int32_t/uint8_t [VERIFIED]. Nothing in a file records or checks the geometry; rows/cols come from the caller's #defines in main (345-358).

  • load_int32_mat_from_file(file, into, rows, cols) (184-207) does fread(&into[0][0], sizeof(int32_t), rows*cols, infile_ptr) — a single blob into the row-major backing store created by HEAP_MAT (52, name[i] = name[i-1] + szc). So the file is row-major int32[rows][cols].
  • load_int32_vec_from_file (209-232) — int32[size].
  • load_uint8_mat_from_file (234-257) — row-major uint8[rows][cols].
  • load_uint8_vec_from_file (259-282) — uint8[size].
  • All four printf a warning and continue if the read is short (197-198 etc.); only a missing file returns FALSE, and no caller checks the return value (409-417) [VERIFIED].

Reconciling declared dims against ls -l [MEASURED]:

File bytes declaration arithmetic OK
AD1_2048 49,152 load_int32_mat_from_file("AD1_2048", AD1, W, N1), 409, AD1 = W×N1 381 2048 × 6 × 4 ✓
AD2_2048 65,536 410, 382 2048 × 8 × 4 ✓
AD3_2048 81,920 411, 383 2048 × 10 × 4 ✓
AD4_2048 98,304 412, 384 2048 × 12 × 4 ✓
thresh1..4_2048 8,192 each 414-417 2048 × 4 ✓
train_data 47,040,000 234, TRAIN_SZ × INPUT_SZ 354,356 60000 × 784 × 1 ✓
train_label 60,000 259 60000 × 1 ✓
test_data 7,840,000 234, 355 10000 × 784 × 1 ✓
test_label 10,000 259 10000 × 1 ✓

Note the task brief's hint "49152 B = 2048 × 24 B" is right in bytes and equals 2048 × 6 × int32; the file is not packed — it is one int32 per synapse. Labels are single bytes, not one-hot.

Derived memory (full_mnist_2048.c:386-391, 360): each SBC tensor is BITNSLOTS(2048·2048·10) = 1,310,720 uint32 = 5,242,880 B; six of them = 31.46 MB, zero-initialised by calloc. Plus 295 KB of ADs+thresholds, plus 47 MB of MNIST train data resident in RAM (370-373). Peak RSS is dominated by the data-file choice, not the model.

1.9 Hyperparameters that matter

Knob Value Cite
ADEs per layer (W) 2048 346
Layers 4 349-352
Synapses per layer 6 / 8 / 10 / 12 349-352
Classes (D) 10 347
Input width 784 356
SBC memories (SBC_CT) 6, one per layer pair 358, 470-480
Label classes per memory 10 (D) 347
Synapse scale yang ±64 112,114
Pixel centre 127 118
Learning rate none (idempotent bit-set) 142-144
Firing threshold per-ADE int32, from file 122, 414-417
Fire comparison count >= thresh 122
Epochs 1 (single pass) 436-482
Test protocol all 10,000 images, argmax of summed counts 488-548

1.10 The reported accuracy and how it is measured

full_mnist_2048.c:550-553:

printf("\n %u wrong = %6.3f pct correct \n", wrong, 100.0 - wrong / ( TEST_SZ / 100.0 ) );

wrong accumulates only when inf_label != label (543-548), and a 10×10 confusion matrix is printed afterwards (555-565). [VERIFIED] This is honest top-1 accuracy on the full 10,000- image MNIST test set after a single pass over the full 60,000-image training set — but with the §1.6 truncation bug active. [MEASURED] ./bb prints 346 wrong = 96.540 pct correct in 66 s wall (single-threaded). The confusion matrix shows the usual MNIST structure: 9↔4, 9↔3, and 8↔3 are the biggest confusions; 1s and 0s are nearly perfect [MEASURED].

Nothing here is MNIST-independent. [VERIFIED] The AD files' positions index 784 pixels; the thresholds are calibrated to the MNIST count distribution; train_data/test_data are hardcoded filenames inside main (402-406). This program is a demo, not a library.


2. Can this become a gun?

2.1 Can the C be linked as a library?

Not in the shape it is in. [VERIFIED] — here is the concrete evidence, item by item:

Item Finding Cite
Clean callable API? Partially. find_AD_firing_pattern, write_to_sbc, read_from_sbc are non-static with plain C signatures, so they can be called. But there is no header; the only header present is bitarray.h. 98, 129, 157; zip listing
main does all the work? Yes. Allocation (363-400), file loading (402-417), training (435-482), inference (486-548), reporting (550-565) are all inside main (329). 329-567
Hardwired dimensions? In main only. W, D, N1..N4, TRAIN_SZ, TEST_SZ, INPUT_SZ, SBC_CT are #defines in main's body (345-358). The three core functions take w and n as parameters and are dimension-generic — except that they read the globals ws, wd, ds through the D3 macro. Any caller must assign those first. 68, 360, 62
File I/O inside the core? No. All I/O is in main and the four loaders. The core is pure computation. 184-282
OpenMP? Off by default — // #define USE_OPENMP is commented out (37), with #ifdef USE_OPENMP at 38. The #pragma omp ... directives at 440, 466, 501, 513 are therefore inert and there is no libgomp dependency. Turning it on would also introduce shared mutable state across the parallel sections. 37-38, 440-480, 513-525
Statics/globals? Five non-static globals ws, ds, wd, ww, wwd (62) — linkable but externally writable, and ws/wd/ds are load-bearing. class_count/confusion are locals in main. 62, 68
Call-signature ergonomics Poor for per-tick use. find_AD_firing_pattern(loop, w, n, uint8_t** sensory_input, ...) requires the whole input matrix to be resident and indexes it by loop (116). A single observation means passing a 1-row matrix with loop = 0. 98, 116
Licence GPL-3.0-or-later, © University of Manchester (1-19). This repo has no LICENSE file and no licence statement in README.md/CONTEXT.md [VERIFIED]. Linking this C into the bot makes the bot GPL; that is an unmade decision, not a detail. 1-19

Concrete obstacle list per interop route:

  1. {.compile: "full_mnist_2048.c".} + {.importc.} — the fastest to try, but:
    • Nim's own generated C defines main, so the C file's main (329) is a duplicate symbol at link time. {.compile.} gives no way to inject -Dmain=... for that one file; you need a one-line shim, bitbrain_shim.c, that does #define main bitbrain_unused_main then #include "full_mnist_2048.c", and compile that.
    • You must also importc the five globals (62) and assign ws/wd/ds (360) before the first call, or D3 computes into unreachable memory. Silent wrong answers, not a crash.
    • -O3 -march=native are the file's own recommended flags (full_mnist_2048.c:22-23); via {.compile.} the whole Nim build's flags apply, not per-file ones. Workable, but it means the tuning in the header comment is lost.
    • bitarray.h must ship next to it (35).
  2. Static archive (gcc -c the shim → ar rcs libbitbrain.a → {.passL: "-lbitbrain".}) — cleanest link, same shim requirement, plus build-system plumbing and an architecture-specific binary artifact. The repo's .gitignore convention is explicit binary names (see §3), so a committed .a would be a new tracked binary that only works on one platform.
  3. {.header.} / importc shim only — impossible as-is: there is no header to point at. You would have to write bitbrain_wrapper.h + bitbrain_wrapper.c anyway. Once you are writing a wrapper, you may as well have the wrapper do the useful work: set the globals, expose bb_set_dims(w, d), bb_forward(const uint8_t* row, uint8_t** fired), bb_train(const uint8_t* fired, uint8_t label), bb_read(...), bb_reset(), and fix the §1.6 truncation bug in the wrapper's own read loop.

Bottom line: the C is 40 lines of trivial integer arithmetic plus two file loaders. What it is not is a library: it is a main with a hidden global dimension state and a broken reader. Either way you must write a wrapper. Given that, a Nim port of §1.1/§1.5 is cheaper and safer than any interop route — [INFERRED], but note that a port is a derivative work, so it inherits the GPL question exactly; the licence is not dodged by porting.

2.2 The wall this project has already hit — and why BitBrain does not get around it

The brief's premise is right and the repo's own numbers back it:

  • The TM head scored below its own majority class. common_libs/tests/tm_pattern_sweep_results.md:626-631: majority class 2 at 38.77%, head accuracy 36.69% (642473/1751067), −2.08 pp. The shuffled-label control sits at chance (20.04% vs 20.12%).
  • Live side accuracy sat at chance. The same file records online radial-head accuracy at 48.8% vs 19.9% shuffled (:273) and the gated side head at 40.37% vs a 38.75% majority (:649) — i.e. +1.6 pp for a learner with a gate.
  • Every TM arm tied or lost to the shipped Pattern gun. docs/gun_rack_summary.md: rack overall 6.95% real hit rate, Tsetlin 5.8%; disabling Tsetlin gave 6.18% → 5.76% (p = 0.57) [VERIFIED].
  • The break-even bar is already quantified. common_libs/guns/tm_horizon.nim:49-53: gate 2b measured that a correcting head "needs ~80% side accuracy to break even on hits while the achievable signal is ~60%".
  • And this repo already ran a 16-family / ~260-config shootout on the same aim target: BNNBot_garage/analysis/shootout_results.md — WiSARD (best generalist, 16.20 px MAE), Tsetlin (20.69 px, loses on all 4 enemies to cold start), Kanerva SDM (10.16 px, "no improvement on baseline"), Bloom filter (17.08 px, "worse than baseline"), HDC/VSA (broken), against a linear+Hebbian baseline at 17.61 px.

That last point is the one to sit with. BitBrain's two stages are already represented in that shootout:

BitBrain stage Already evaluated as Result there
AD layer (random address decoders + threshold) WiSARD / n-tuple LUT nodes (BNNBot_garage/src/wisard_predictor.nim, K=14, 50 nodes) Best of 16 families, 16.20 px — and the repo still ships Pattern
SBC coincidence memory + counting readout Bloom filter (bit array, write-once) and Kanerva SDM (address→content memory) 17.08 px and 10.16 px, neither beat the baseline on the metric that mattered
(nearest TM analogue) Tsetlin gun, TMHorizon 5.8% real vs 6.95% rack; TMHorizon predicted to lose

[INFERRED] So BitBrain is not a new kind of learner on this problem — it is a sharper addressing scheme on top of the same primitive the repo already tried, and its readout (binary "have I seen this coincidence with this label") is strictly weaker than WiSARD's running-mean readout, because it throws away the co-occurrence counts. Its one genuine advantage over the Tsetlin machinery is that it is non-stochastic and order-free (§1.7) — no learning rate to tune, no cold-start pathology, no clause-length collapse.

What target signal would make BitBrain worth building? Stated as a gate, mirroring the repo's own Gate-2b language: a target whose offline learnability, measured on held-out resolved shots, clears ≥80% side accuracy (or equivalent break-even on the hit metric) and clears a shuffled-label control at chance. The published evidence is that the available side signal is ~60% [VERIFIED] (tm_horizon.nim:49-53). No learner change moves that number, and BitBrain is a learner change. Porting BitBrain before finding such a target would be the third repetition of the same mistake — a learner swapped onto a signal-poor target — and it would cost the AD-layer problem of §1.3 on top.

2.3 The runtime budget — measured, not extrapolated

Per the repo's own cost harness, the live budget is ~13.16 ms/tick (76 ticks/s) (common_libs/tests/measure_tm_pattern_cost.nim:6-9); the TM Pattern gun measures 0.36 ms/tick and Tsetlin ~5.3 ms/tick (common_libs/guns/tm_horizon.nim:55-57). The brief's "1 ms/tick, TMHorizon 0.19 ms, Tsetlin 1.94 ms" numbers do not match what the repo records — I am reporting the repo's numbers and flagging the discrepancy rather than guessing which is current. Either way the conclusion below does not change.

I compiled and ran the C. MEASURED (gcc 15.3.0 -O3, single-threaded, MNIST geometry 784 → 2048 ADEs × 4 layers → 6 × 5 MB SBC tensors):

Operation Measured cost
Stage 1: 4-layer find_AD_firing_pattern forward pass 0.3599 ms/sample
Stage 3: shipped read_from_sbc (buggy, 6 memories) 0.412 ms/sample
Stage 3: truncation removed (6 memories) 0.531 ms/sample
Stage 3: sparse active-list read, no truncation 0.202 ms/sample
Full inference = stage 1 + sparse stage 3 ≈ 0.56 ms/sample
Stage 2: write_to_sbc training (dense scan, 6 memories) 0.600 ms/sample
Integration harness overhead, all 6 readers 0.736 → 5.330 ms/sample as memories fill
Full original program: 60,000 train + 10,000 test 66 s wall, [MEASURED] ./bb

Two things to read off this:

  • Stage 1 is 0.36 ms; the shipped stage-3 reader is more expensive than the whole feature layer (0.41 ms), because read_from_sbc scans all 2048×2048 j slots per firing i irrespective of sparsity. The sparse reader — iterate the ~25-element active list per layer instead of a 2048-slot scan — is 2.6× faster than the fixed reader and 2.0× faster than the shipped one, and produces bit-identical accuracy (97.210% both) [MEASURED]. This is a free 2× win for any port.
  • 0.56 ms/sample is affordable. It is in the same class as the TM Pattern gun's 0.36 ms/tick, and it is a hard, non-iterative bound — no epochs, no retraining pass. Note however that the 0.202 ms sparse read is dominated by cache misses (~5,400 scattered probes into six 5 MB arrays per sample), not by arithmetic; a shooting-configuration SBC memory would be far smaller and proportionally faster. [INFERRED]
  • Do NOT copy the tensor size. 2048×2048×10 bits × 6 is 31.5 MB of L2/L3-hostile memory. A gun-sized version (say 512 ADEs × 2 layers × 6 classes) is ~0.8 MB and would not have this behaviour. [INFERRED]

2.4 The online-learning constraint — satisfied

ModularBot must learn inside one battle and must not persist across battles (state is wiped when the target changes — see the resetLearning contract documented at common_libs/guns/tm_horizon.nim:37-40, which is the repo's statement of that rule).

BitBrain's trainable part, judged against that [VERIFIED] from 129-152:

Requirement BitBrain Verdict
Per-sample update, no batch write_to_sbc is called per sample ✅
Order-independent idempotent BITSET ✅
No learning rate / hyperparameter to tune none exists ✅
Usable at inference before/while training read_from_sbc reads whatever is set ✅
Survives across rounds of the same battle monotone accumulation ✅
Wipes cleanly on enemy change clear 31.5 MB (memset/re-calloc) ✅ (cost, not correctness)
Does not require old samples to be replayed yes ✅
Needs pretrained features from outside the battle yes — the ADs ⚠️ This is the failure

So the learner satisfies the constraint perfectly; the feature extractor is a pre-battle artifact, which is exactly the thing resetLearning cannot regenerate and exactly the thing we do not have the program for.

2.5 Recommendation

Do not port BitBrain as a gun now. Reasoning, in priority order:

  1. The AD layer is unobtainable (§1.3). This is not a nice-to-have; without it BitBrain is a counting head on top of a feature layer we design ourselves — at which point BitBrain's contribution is the counting head alone, and §2.2 already tells us what happens to learner-only changes on this target.
  2. The signal is the wall, and it is already quantified (§2.2): −2.08 pp vs majority for the TM head, live side accuracy at 47–49% (chance), break-even needing ~80% where the achievable signal is ~60%. BitBrain does not touch that.
  3. The primitive is already in the repo's evaluated set (§2.2): WiSARD is the AD layer, Bloom/SDM are the SBC memory, and none of them beat the incumbent on the metric that mattered.
  4. Cost is not the blocker (§2.3): ≈0.56 ms/sample is affordable. So the decision rests entirely on 1–3, and it comes out negative.

What would change the verdict — the cheap, honest next step, if the user wants one:

  • Reuse the AD primitive, not the learner. BNNBot_garage/src/wisard_predictor.nim already implements an n-tuple address-decoder layer with a running-mean readout and eligibility traces, and it already won a 16-family shootout. If BitBrain's counting readout is the interesting part, the experiment worth running is a readout swap on that existing feature extractor, measured against the same fixture the shootout used — an offline, no-battle-change experiment, not a port.
  • Gate it exactly like Gate 2b. Do not build a gun until an offline corpus of resolved shots shows a target signal that clears the ≥80% side-accuracy break-even and a shuffled-label control at chance. If it does not, the answer is "no" regardless of learner.
  • If a port is ever justified: fix the §1.6 truncation bug, use the sparse active-list reader (§2.3, 2.6× faster for free), size the SBC tensors to the gun's feature count (not 2048²), normalise the per-class score by stored-feature count to kill the class-frequency bias (§1.5), and settle the GPL-3.0 question before writing the wrapper.

How this could still fail even if we do everything right: (a) a synthesised AD layer is statistically similar but not the same, and there is no MNIST-comparable baseline to validate it against (§1.3.1); (b) the class-frequency bias with a few hundred labels per class will be much worse than on balanced MNIST (§1.5); (c) find_AD_firing_pattern needs the whole input matrix resident and indexes by row (98,116), so a per-tick port must serialise each observation into a matrix row — cheap, but it is the kind of plumbing that hides latency; and (d) the 31.5 MB tensor if copied verbatim is a cache disaster (§2.3).


3. Should the 11 MB zip be committed or gitignored?

Recommendation: commit it, under its explicit path docs/BitBrain_C_code.zip. Facts behind that call [MEASURED]/[VERIFIED]:

  • git status shows docs/BitBrain_C_code.zip as the only untracked file besides common_libs/tests/measure_aim_vs_power.nim, and there is no .gitignore rule covering it — consistent with the repo's stated convention of listing explicit binary names (.gitignore names common_libs/tests/tm_measure, tr_bots/, *_garage/out/, etc.) rather than broad *.zip patterns.
  • The repo's norms comfortably accommodate this size: it already tracks eleven ~11 MB zips (SAC_LSTM_Bot_garage/weights*/sac_*.zip) and PDFs up to 9.3 MB (docs/papers/neuroevolution/2023_neuroevolution_recurrent_architectures.pdf). .git is already 117 MB.
  • It is a third-party primary source that cannot be re-derived from anything in the repo — the weights, the thresholds, and the MNIST dump exist nowhere else here. If it is gitignored, every citation in this document becomes unverifiable for the next reader. Tracked docs/papers/*.pdf show the repo already treats third-party primary sources as trackable content.
  • Caveat, if the user would rather not carry it: the explicit rule to add is docs/BitBrain_C_code.zip (matching the existing style), and also docs/BitBrain_C_code/ for the 55 MB extracted tree, which should never be committed in either case. Do not add *.zip — that would silently swallow the SAC weight archives.

4. Appendix — reproducing every measurement

mkdir -p /tmp/bitbrain && cd /tmp/bitbrain
unzip /home/davide/Projects/SirRoboGarage/docs/BitBrain_C_code.zip
cd BitBrain_C_code
gcc full_mnist_2048.c -O3 -lm -o bb && time ./bb      # -> 346 wrong = 96.540 pct correct, 66 s

Harnesses I wrote (throwaway, in /tmp/bitbrain/measure/, not in the repo):

File What it measures
stats.c AD fire rates, 4-layer forward-pass timing
thresh.c per-ADE fire-rate distributions, sign/position statistics, dead-ADEs
exp.c online learning curve, SBC occupancy, dense vs sparse read timing
final.c shipped vs fixed vs sparse reader accuracy + per-sample cost on all 10,000 tests
dbg*.c the §1.6 truncation-bug isolation chain

All three readers in final.c contain verbatim copies of find_AD_firing_pattern (98-124), write_to_sbc (129-152) and read_from_sbc (157-181); the only differences are the bug fix and the sparse traversal. The as-shipped reader reproduces the original program's 96.540% exactly, which is the check that the copies are faithful.