# BitBrain — what the algorithm actually is, and whether it can become a gun **Issue:** N/A (read-only research requested by orchestrator) **Date:** 2026-09-24 **Primary source:** `docs/BitBrain_C_code.zip` (11,656,362 B, untracked), extracted to `/tmp/bitbrain/BitBrain_C_code/`. Everything below cites that tree: - `full_mnist_2048.c` — 568 lines, the whole implementation - `bitarray.h` — 41 lines, bit-array macros (`uint32_t` slots; `BITSET` 35, `BITCLEAR` 36, `BITTEST` 37, `BITNSLOTS` 39) - pretrained weights `AD1_2048 … AD4_2048`, `thresh1_2048 … thresh4_2048` - MNIST data `train_data … test_label` (the classic uint8 MNIST dump) The C file carries a GNU GPL-3.0-or-later header, `© 2022 The University of Manchester` (`full_mnist_2048.c:1-19`). The zip contains **no build files, no header, no README, and no unsupervised-learning program**. **Verification legend:** | Tag | Meaning | |---|---| | `[MEASURED]` | I compiled and/or ran the code (or an exact verbatim copy of it) here and read the number off the output. | | `[VERIFIED]` | I read the exact line, and/or reconciled it against byte sizes / arithmetic I recomputed. | | `[INFERRED]` | My reasoning from code + measurements; not directly observed. | | `[UNKNOWN]` | I could not determine it from the primary source. | Unless stated otherwise, every measurement was taken with `gcc 15.3.0`, `-O3`, **single threaded** (OpenMP is off by default — see §4.1), on this machine (16 cores). --- ## 0. One-paragraph answer BitBrain (as shipped here) is a **three-stage, purely feed-forward pipeline**: (1) a **random-projection threshold layer** of 2048 "address decoders" (ADs) per layer, across 4 layers with 6/8/10/12 synapses each, whose ~74000 integer `(position, sign)` weights and 2048×4 integer firing thresholds arrive **as two files from a program we do not have** (`full_mnist_2048.c:376-384`, `409-417`); (2) a **sparse binary coincidence (SBC) memory** — six `2048×2048×10` bit tensors (31.5 MB) that are written once per labelled sample, a *pure `BITSET`, never cleared* (`full_mnist_2048.c:129-152`); (3) a **counting readout** that sums, over the 6 memories, how many of the observed AD-pair coincidences are "known" for each of the 10 labels and takes the argmax (`full_mnist_2048.c:157-181`, `527-540`). The program prints `346 wrong = 96.540 pct correct` on all 10,000 MNIST test images `[MEASURED]`, but it does so **with a real bug**: the readout's `uint8_t bit_test` truncates the 32-bit bit-test result, so only ADs with `i % 32 < 8` are ever counted (`full_mnist_2048.c:161,173-175`). Removing that truncation gives **97.210%** `[MEASURED]`. Stages 2 and 3 are genuinely **online, single-pass, per-sample, order-free** and therefore satisfy the `ModularBot` constraint; stage 1 is the part we cannot rebuild from this source. On the gun question the honest answer is: **no, not on this evidence.** The learner is cheap (0.56 ms/sample for one full forward pass here `[MEASURED]`), but this repo has already evaluated every structurally equivalent learner on this target — WiSARD/n-tuple LUTs (which *are* the AD primitive), Tsetlin, Kanerva SDM, Bloom filters, HDC/VSA — and the binding constraint was the **target signal**, not the learner (`BNNBot_garage/analysis/shootout_results.md`, `common_libs/tests/tm_pattern_sweep_results.md:611-631`, `common_libs/guns/tm_horizon.nim:49-53`). BitBrain would be the third learner swapped onto that same signal-poor target. --- ## 1. What the algorithm is ### 1.1 The core primitive: one ADE = one signed, thresholded random projection `find_AD_firing_pattern` (`full_mnist_2048.c:98-124`) is the whole of stage 1. For each of the `w` ADEs in a layer, for each of the `n` synapses of that ADE: ```c if( address_decoder[i][j] > 0 ) { yang = 64; position = address_decoder[i][j]; } // line 112 else { yang = -64; position = -address_decoder[i][j]; } // line 114 pixel = sensory_input[ loop ][ position-1 ]; // line 116 count += ( pixel - 127 ) * yang; // line 118 ... ad_row_fired[ i ] = ( count >= ad_row_thresh[ i ] ? TRUE : FALSE ); // line 122 ``` So, precisely `[VERIFIED]`: - An **address decoder element (ADE)** is described by an `n`-long vector of non-zero `int32_t`. **The sign is the synapse sign** — positive = excitatory, negative = inhibitory (`full_mnist_2048.c:90-91`, doc comment). **The magnitude is the input position**, 1-based into the flattened input row (`position-1`, line 116). - `count` is a **signed, 127-centred, x64-scaled sum of the sampled input bytes**. `yang` is only ever `±64`, a pure scale factor; it is not a learned weight. So `count/64 = Σ_j ±(pixel_j − 127)`, i.e. each "pixel" is treated as a real in `[−127, 128]` with 127 as the decision centre. `[VERIFIED]` - The ADE fires iff `count >= ad_row_thresh[i]` — a **≥** comparison against a per-ADE integer threshold, i.e. the ADE is a thresholded linear unit on a random `n`-dimensional subspace. `[VERIFIED]` - The output is a dense `uint8_t[w]` 0/1 vector, `ad_row_fired` (line 122). The doc comment calls the input "`0.255` for E/MNIST" (`full_mnist_2048.c:84`) — a typo for "0..255"; the values are raw MNIST `uint8` `[VERIFIED]`. **Layer geometry** (`full_mnist_2048.c:345-352,358`): `W = 2048` ADEs per layer; four layers with `N1..N4 = 6, 8, 10, 12` synapses; `SBC_CT = 6` coincidence memories, one for each of the 6 unordered layer pairs: (1,2) (1,3) (1,4) (2,3) (2,4) (3,4) — see the six `write_to_sbc` calls at `470-480` and the six `read_from_sbc` calls at `516-525`. `D = 10` label classes (line 347). Input width `INPUT_SZ = 784` (line 356). **This is an n-tuple / RAM-node / WiSARD-style address decoder with a threshold** `[INFERRED]` — i.e. the classic "random receptive field + threshold" feature layer. It is not novel machinery; this repo already has one (`BNNBot_garage/src/wisard_predictor.nim`, K=14, 50 LUT nodes, ~16384 entries each, averaged readout). ### 1.2 Sparsity (measured) | Layer | n | mean ADEs fired / 2048 | mean fire rate | |---|---:|---:|---:| | 1 | 6 | 25.4 | 1.24% | | 2 | 8 | 26.7 | 1.30% | | 3 | 10 | 28.7 | 1.40% | | 4 | 12 | 29.8 | 1.45% | `[MEASURED]` — 2,000 test images, verbatim `find_AD_firing_pattern`. Over 20,000 training images the rates are 1.269 / 1.355 / 1.461 / 1.504% and **every** image fires ≥1 ADE in every layer (so no image is ever "empty") `[MEASURED]`. The per-ADE rate distribution is **very heterogeneous** `[MEASURED]` (20,000 train images): | Layer | never fires | 1–2% | 2–3% | ≥10% | hottest single ADE | |---|---:|---:|---:|---:|---:| | 1 | 778 (38%) | 1187 | 67 | 7 | 51.90% | | 2 | 656 (32%) | 1272 | 97 | 9 | 23.82% | | 3 | 592 (29%) | 1238 | 163 | 13 | 31.37% | | 4 | 522 (25%) | 1293 | 175 | 17 | 12.75% | This is *not* a uniform quantile calibration: a third of the ADEs are dead and a handful are heavy hitters. That is the fingerprint of a selection/pruning process, not a global rule. ### 1.3 What is trainable HERE vs PRETRAINED ELSEWHERE — **the decisive split** `[VERIFIED]` — reading the file end to end, the only learnable things updated by code in this tree are the bits of the six SBC tensors. | Component | Where it comes from | Learned in this file? | |---|---|---| | 4 × 2048 firing thresholds (`thresh*_2048`) | **file**, loaded `full_mnist_2048.c:414-417`; documented "thresholds from unsupervised learning" (`376-379`) | **No** | | 4 × 2048 ADs (`AD*_2048`) | **file**, loaded `full_mnist_2048.c:409-412`; documented "ADs from unsupervised learning" (`381-384`) | **No** | | 6 × `2048×2048×10` SBC bit tensors | `calloc`'d zero at `full_mnist_2048.c:386-391` | **Yes** — `write_to_sbc`, `129-152`, called in the loop at `435-482` | There is **no** code anywhere in the 568 lines that modifies an AD or a threshold. I listed the complete function inventory: `find_AD_firing_pattern` (98), `write_to_sbc` (129), `read_from_sbc` (157), four file loaders (184, 209, 234, 259), three display helpers (284, 296, 310), and `main` (329) `[VERIFIED]`. `show_thresholds`/`show_AD` are `printf`-only and are behind a commented-out `#define SHOW_LOADED_FILES` (`420-431`). `[VERIFIED]` **Say it loudly: we cannot regenerate the AD layer from this zip.** The unsupervised program that produced `AD*_2048` and `thresh*_2048` is not present, not referenced, and not described beyond the two comments above. Any BitBrain gun must therefore **either** ship those exact files **or** synthesise a substitute AD layer of its own — and the original was trained on 784-pixel MNIST, so the files are useless for any other input domain. There is no third option. I did, however, characterise the files well enough to say what a substitute would have to look like, and how much of the original's structure is *recoverable* from the unlabelled input distribution alone. #### 1.3.1 Structure of the AD files (this is recoverable — partly) `[MEASURED]`, from the four `AD*_2048` files: - **Magnitudes** (`|AD|`, which *is* the input position) lie in `[37, 776]` for every layer, with **quasi-linear deciles** (`L4`: min 37, then 206/261/301/355/409/462/514/556/610, max 776). So the magnitudes look like near-uniform random draws over the usable index range. - The **range itself is explained by dead pixels**: pixels `1..36` (top row + part of the second) and `777..784` (bottom-right corner) are inked in **0 of 3,000** training images, and those exact positions are sampled **0 times** by every layer `[MEASURED]`. That is why the minimum `|AD|` is 37 and the maximum is 776 rather than 1 and 784. - Of the 134 pixels that are constant across 3,000 images, 133 / 134 / 127 / 132 are **never** used by layers 1..4 respectively — i.e. the position pool excludes constant pixels. - Position usage tracks input activity almost perfectly: **Spearman(ink frequency, AD-slot usage) = 0.994** over the 784 pixels; 75.9% of all AD synapse slots sit on the 219 pixels that are inked in more than a third of images, and the 134 dead pixels get exactly **0.0%** `[MEASURED]`. - Synapse signs are **not** 50/50: excitatory fractions are 44.7% / 42.9% / 40.3% / 39.2% `[MEASURED]`. - Still unexplained: ~93–114 *inked* pixels are never sampled by each layer, and the 220–247 unused positions per layer have a median ink-count of **0** with max 33–89 (vs a median of 619–698 for used positions) `[MEASURED]`. There is a real selection rule beyond "drop dead pixels" that I could not reverse-engineer `[UNKNOWN]`. - **Thresholds are not per-ADE fire-rate quantiles.** For each layer, 13–25 ADEs have a threshold **above the maximum `count` that ADE can ever produce**, so they can never fire (1.22% / 0.68% / 0.98% / 0.63% of ADEs) `[MEASURED]`. A quantile calibration on the same data can never produce that. A handful have negative thresholds (`thresh1` min −10871, `thresh4` min −15496) but **no** ADE always fires `[MEASURED]`. **Consequence for a port:** a substitute AD layer could be synthesised cheaply from *unlabelled* samples of the new input domain — (a) drop constant positions, (b) sample positions with probability monotone in activity, (c) uniform random `±` signs, (d) set each threshold to a quantile of that ADE's own `count` distribution to hit a target fire rate (~1.3–1.5%) — but it would be a **statistically similar, not the same, feature layer**, and the only way to validate it is end-to-end accuracy on the target task. There is no apples-to-apples baseline to compare against because the original ADs are MNIST-only. That is the single largest technical risk in the whole idea `[INFERRED]`. ### 1.4 Binarisation and the `count`↔bit-width relationship There is **no binarisation** before the ADs: `sensory_input` is the raw `uint8` MNIST matrix, 784 bytes per row (`full_mnist_2048.c:234-256` loader, `346-356` sizes) `[VERIFIED]`. The ADE does the analogue-to-threshold conversion itself, via `(pixel − 127) * yang` (line 118). The bit-width note: `count` is `int32_t` (line 101), and its per-synapse term is `int8_t(±64) × (pixel − 127)`, so one synapse contributes at most `64 × 128 = 8192`. With `n = 12` the maximum `|count|` is `98304`, which **overflows `int16_t`** — hence the `int32_t` for both `count` and `ad_row_thresh`. `pixel` is `int16_t` (line 103), which is sufficient for 0..255. `[VERIFIED]` And indeed the largest threshold in the data, 99053 (`thresh4_2048`), sits just above that 98304 ceiling `[MEASURED]`. ### 1.5 The supervised head: `write_to_sbc` / `read_from_sbc` **Write (learning)** — `full_mnist_2048.c:129-152`, comment at `128`: ```c FOR_LOOP(i, w) if (first[i]) FOR_LOOP(j, w) if (second[j]) { tens_index = D3(i, j, label); // 138: (i, j, class) -> one bit if (!BITTEST(mem, tens_index)) { internal_count++; BITSET(mem, tens_index); } // 142-144 } ``` `D3(i,j,k) = i + j*wd + k*ws` with `wd = W*D` and `ws = W` (`full_mnist_2048.c:68`, globals set at `360`), so the bit layout is `[j][label][i]` with `i` fastest. `[VERIFIED]` Thus the learnable rule is exactly: **set the bit at (coincidence of an ADE from layer A firing and an ADE from layer B firing, current label)**. Nothing is ever decremented or cleared; there is **no learning rate, no counter, no decay, no negative evidence** `[VERIFIED]`. The six `write_to_sbc` calls per sample (train loop `435-482`, calls at `470-480`) each receive the label as `uint8_t` (`464`). **Read (inference)** — `full_mnist_2048.c:157-181`: ```c FOR_LOOP(i, w) if (first[i]) FOR_LOOP(j, w) if (second[j]) FOR_LOOP(k, ds) { tens_index = D3(i, j, k); // 171 bit_test = BITTEST(mem, tens_index); // 173 <-- uint8_t! if (bit_test) (store[k])++; } // 175 ``` then the six per-memory count vectors are summed and argmaxed (`527-540`). So the score for class `k` is *the number of observed AD-pair coincidences whose bit for `k` is set* — an **overlap count**, not a likelihood and not a probability. `[VERIFIED]` Two consequences worth stating plainly: 1. **No normalisation.** There is no division by the number of stored features per class, so the score is biased toward whichever class has *more bits stored* — a class seen more often, or whose features are more diverse, wins ties for free. On balanced MNIST this is mild; on a gun with a few hundred labels it is a real hazard. `[INFERRED]` from lines 175 + 527-531. 2. **No negative evidence.** A coincidence that is *never* seen for class `k` contributes 0, identical to a coincidence seen for *every* class. Frequent coincidences saturate to 10/10 and contribute a constant to all classes (harmless for argmax); rare coincidences carry the entire discrimination. `[INFERRED]` ### 1.6 The `read_from_sbc` bug — found, isolated, and reproduced `[MEASURED]` `bit_test` is declared `uint8_t` (line 161) but `BITTEST` yields the full masked 32-bit word (`bitarray.h:37`). Whenever the hit bit sits at word-position ≥ 8, `bit_test` truncates to 0 and the hit is **silently dropped**. Since `D3` strides by `ws = W = 2048` (a multiple of 32) and `wd = W*D = 20480` (also a multiple of 32), the bit's position inside its `uint32` word is exactly `i % 32`. So **the shipped reader only counts ADEs with `i % 32 < 8` — three quarters of the AD range is invisible at inference time.** `[VERIFIED]` by arithmetic and `[MEASURED]` by reproduction. Verification chain (all on the same trained memory, sample 0, layer pair (1,2)): | Counter | Value | |---|---:| | Python, independent byte-level bit test, full `i` range | **319** | | Python, same but restricted to `i % 32 < 8` | **50** | | C, verbatim `read_from_sbc` | **50** | The write path is unaffected: `write_to_sbc` uses `!BITTEST(...)` in a boolean context (line 142), so no truncation occurs there. `[VERIFIED]` The bug is read-only and one-sided: 25% of the evidence is used. End-to-end, on all 10,000 test images after the full 60,000-sample training pass `[MEASURED]`: | Reader | Accuracy | ms/sample | |---|---:|---:| | `read_from_sbc` **as shipped** (with the truncation bug) | **96.540%** | 0.412 | | `read_from_sbc` with the truncation removed | **97.210%** | 0.531 | | sparse active-list reader, no truncation | **97.210%** | **0.202** | The 96.540% row **exactly reproduces the number the unmodified program prints** `[MEASURED]` (`./bb` → `346 wrong = 96.540 pct correct`, 66 s wall, single-threaded). So the bug is real, it costs 0.67 pp, and it does not invalidate the paper's claim — but anyone porting this should fix it, and anyone benchmarking against "96.5%" should know it is a lower bound. ### 1.7 Is the supervised rule online / usable incrementally? **Yes, unambiguously** `[VERIFIED]` from `129-152`: the update is per-sample, order-independent (it is an idempotent `BITSET`), single-pass, and needs no batch, no epoch, no learning rate, and no revisit of older samples. Inference is available at any point, including before any training. The only thing that is *not* cheap is resetting: the memory is monotone, so "forget the old enemy" means clearing 31.5 MB (`calloc`/`memset`) and starting over. The learning curve is monotone and does **not** collapse `[MEASURED]` (cumulative single pass over the training set, evaluated on the first 2,000 test images): | training samples | per class | acc, shipped (buggy) reader | acc, correct reader | |---:|---:|---:|---:| | 600 | 60 | 79.05% | 82.50% | | 1,200 | 120 | 85.05% | 86.85% | | 2,400 | 240 | 87.80% | 89.75% | | 4,800 | 480 | 90.80% | 92.20% | | 12,000 | 1,200 | 93.05% | 93.90% | | 30,000 | 3,000 | 94.20% | 95.30% | | 60,000 | 6,000 | 95.20% | 95.70% | So it needs **~a few hundred labelled samples per class** before it is useful and ~2,000+ to approach its ceiling. There is no saturation collapse: at 60,000 samples only 18.9–20.8% of each SBC tensor's bits are set, 70.8–76.3% of `(i,j)` pairs are non-empty, and only **0.03–0.05%** of non-empty pairs have all 10 labels set `[MEASURED]`. I initially expected monotone saturation to destroy discrimination; it does not, on this data volume. ### 1.8 Exact dimensions and on-disk format — reconciled against byte sizes All loaders read **one raw contiguous blob, no header, no magic, no dimension fields, native-endian `int32_t`/`uint8_t`** `[VERIFIED]`. Nothing in a file records or checks the geometry; `rows`/`cols` come from the caller's `#define`s in `main` (`345-358`). - `load_int32_mat_from_file(file, into, rows, cols)` (`184-207`) does `fread(&into[0][0], sizeof(int32_t), rows*cols, infile_ptr)` — a single blob into the row-major backing store created by `HEAP_MAT` (`52`, `name[i] = name[i-1] + szc`). So the file is **row-major `int32[rows][cols]`**. - `load_int32_vec_from_file` (`209-232`) — `int32[size]`. - `load_uint8_mat_from_file` (`234-257`) — row-major `uint8[rows][cols]`. - `load_uint8_vec_from_file` (`259-282`) — `uint8[size]`. - All four `printf` a warning and continue if the read is short (`197-198` etc.); only a missing file returns `FALSE`, and **no caller checks the return value** (`409-417`) `[VERIFIED]`. Reconciling declared dims against `ls -l` `[MEASURED]`: | File | bytes | declaration | arithmetic | OK | |---|---:|---|---|:--:| | `AD1_2048` | 49,152 | `load_int32_mat_from_file("AD1_2048", AD1, W, N1)`, `409`, `AD1 = W×N1` `381` | 2048 × 6 × 4 | ✓ | | `AD2_2048` | 65,536 | `410`, `382` | 2048 × 8 × 4 | ✓ | | `AD3_2048` | 81,920 | `411`, `383` | 2048 × 10 × 4 | ✓ | | `AD4_2048` | 98,304 | `412`, `384` | 2048 × 12 × 4 | ✓ | | `thresh1..4_2048` | 8,192 each | `414-417` | 2048 × 4 | ✓ | | `train_data` | 47,040,000 | `234`, `TRAIN_SZ × INPUT_SZ` `354,356` | 60000 × 784 × 1 | ✓ | | `train_label` | 60,000 | `259` | 60000 × 1 | ✓ | | `test_data` | 7,840,000 | `234`, `355` | 10000 × 784 × 1 | ✓ | | `test_label` | 10,000 | `259` | 10000 × 1 | ✓ | Note the task brief's hint "49152 B = 2048 × 24 B" is right in bytes and equals `2048 × 6 × int32`; the file is **not** packed — it is one `int32` per synapse. Labels are single bytes, not one-hot. Derived memory (`full_mnist_2048.c:386-391`, `360`): each SBC tensor is `BITNSLOTS(2048·2048·10) = 1,310,720 uint32 = 5,242,880 B`; six of them = **31.46 MB**, zero-initialised by `calloc`. Plus 295 KB of ADs+thresholds, plus 47 MB of MNIST train data resident in RAM (`370-373`). Peak RSS is dominated by the data-file choice, not the model. ### 1.9 Hyperparameters that matter | Knob | Value | Cite | |---|---|---| | ADEs per layer (`W`) | 2048 | `346` | | Layers | 4 | `349-352` | | Synapses per layer | 6 / 8 / 10 / 12 | `349-352` | | Classes (`D`) | 10 | `347` | | Input width | 784 | `356` | | SBC memories (`SBC_CT`) | 6, one per layer pair | `358`, `470-480` | | Label classes per memory | 10 (`D`) | `347` | | Synapse scale `yang` | ±64 | `112,114` | | Pixel centre | 127 | `118` | | Learning rate | **none** (idempotent bit-set) | `142-144` | | Firing threshold | per-ADE `int32`, from file | `122`, `414-417` | | Fire comparison | `count >= thresh` | `122` | | Epochs | 1 (single pass) | `436-482` | | Test protocol | all 10,000 images, argmax of summed counts | `488-548` | ### 1.10 The reported accuracy and how it is measured `full_mnist_2048.c:550-553`: ```c printf("\n %u wrong = %6.3f pct correct \n", wrong, 100.0 - wrong / ( TEST_SZ / 100.0 ) ); ``` `wrong` accumulates only when `inf_label != label` (`543-548`), and a 10×10 confusion matrix is printed afterwards (`555-565`). `[VERIFIED]` This is honest top-1 accuracy on the full 10,000- image MNIST test set after a single pass over the full 60,000-image training set — but with the §1.6 truncation bug active. `[MEASURED] ./bb` prints `346 wrong = 96.540 pct correct` in 66 s wall (single-threaded). The confusion matrix shows the usual MNIST structure: 9↔4, 9↔3, and 8↔3 are the biggest confusions; 1s and 0s are nearly perfect `[MEASURED]`. **Nothing here is MNIST-independent.** `[VERIFIED]` The AD files' positions index 784 pixels; the thresholds are calibrated to the MNIST `count` distribution; `train_data`/`test_data` are hardcoded filenames inside `main` (`402-406`). This program is a demo, not a library. --- ## 2. Can this become a gun? ### 2.1 Can the C be linked as a library? **Not in the shape it is in.** `[VERIFIED]` — here is the concrete evidence, item by item: | Item | Finding | Cite | |---|---|---| | Clean callable API? | **Partially.** `find_AD_firing_pattern`, `write_to_sbc`, `read_from_sbc` are non-`static` with plain C signatures, so they *can* be called. But there is **no header**; the only header present is `bitarray.h`. | `98`, `129`, `157`; zip listing | | `main` does all the work? | **Yes.** Allocation (`363-400`), file loading (`402-417`), training (`435-482`), inference (`486-548`), reporting (`550-565`) are all inside `main` (`329`). | `329-567` | | Hardwired dimensions? | **In `main` only.** `W`, `D`, `N1..N4`, `TRAIN_SZ`, `TEST_SZ`, `INPUT_SZ`, `SBC_CT` are `#define`s in `main`'s body (`345-358`). The three core functions take `w` and `n` as parameters and are dimension-generic — **except** that they read the **globals** `ws`, `wd`, `ds` through the `D3` macro. Any caller must assign those first. | `68`, `360`, `62` | | File I/O inside the core? | **No.** All I/O is in `main` and the four loaders. The core is pure computation. | `184-282` | | OpenMP? | **Off by default** — `// #define USE_OPENMP` is commented out (`37`), with `#ifdef USE_OPENMP` at `38`. The `#pragma omp ...` directives at `440`, `466`, `501`, `513` are therefore inert and there is **no libgomp dependency**. Turning it on would also introduce shared mutable state across the parallel sections. | `37-38`, `440-480`, `513-525` | | Statics/globals? | Five **non-static globals** `ws, ds, wd, ww, wwd` (`62`) — linkable but externally writable, and `ws`/`wd`/`ds` are load-bearing. `class_count`/`confusion` are locals in `main`. | `62`, `68` | | Call-signature ergonomics | **Poor for per-tick use.** `find_AD_firing_pattern(loop, w, n, uint8_t** sensory_input, ...)` requires the **whole input matrix** to be resident and indexes it by `loop` (`116`). A single observation means passing a 1-row matrix with `loop = 0`. | `98`, `116` | | Licence | **GPL-3.0-or-later, © University of Manchester** (`1-19`). This repo has **no `LICENSE` file** and no licence statement in `README.md`/`CONTEXT.md` `[VERIFIED]`. Linking this C into the bot makes the bot GPL; that is an unmade decision, not a detail. | `1-19` | **Concrete obstacle list per interop route:** 1. **`{.compile: "full_mnist_2048.c".}` + `{.importc.}`** — the fastest to try, but: - Nim's own generated C defines `main`, so the C file's `main` (`329`) is a **duplicate symbol at link time**. `{.compile.}` gives no way to inject `-Dmain=...` for that one file; you need a one-line shim, `bitbrain_shim.c`, that does `#define main bitbrain_unused_main` then `#include "full_mnist_2048.c"`, and compile *that*. - You must also `importc` the five globals (`62`) and assign `ws`/`wd`/`ds` (`360`) before the first call, or `D3` computes into unreachable memory. Silent wrong answers, not a crash. - `-O3 -march=native` are the file's own recommended flags (`full_mnist_2048.c:22-23`); via `{.compile.}` the whole Nim build's flags apply, not per-file ones. Workable, but it means the tuning in the header comment is lost. - `bitarray.h` must ship next to it (`35`). 2. **Static archive** (`gcc -c` the shim → `ar rcs libbitbrain.a` → `{.passL: "-lbitbrain".}`) — cleanest link, same shim requirement, plus build-system plumbing and an architecture-specific binary artifact. The repo's `.gitignore` convention is **explicit binary names** (see §3), so a committed `.a` would be a new tracked binary that only works on one platform. 3. **`{.header.}` / `importc` shim only** — impossible as-is: there is **no header** to point at. You would have to write `bitbrain_wrapper.h` + `bitbrain_wrapper.c` anyway. Once you are writing a wrapper, you may as well have the wrapper do the useful work: set the globals, expose `bb_set_dims(w, d)`, `bb_forward(const uint8_t* row, uint8_t** fired)`, `bb_train(const uint8_t* fired, uint8_t label)`, `bb_read(...)`, `bb_reset()`, and **fix the §1.6 truncation bug in the wrapper's own read loop**. **Bottom line:** the C is 40 lines of trivial integer arithmetic plus two file loaders. What it is *not* is a library: it is a `main` with a hidden global dimension state and a broken reader. Either way you must write a wrapper. Given that, a **Nim port of §1.1/§1.5 is cheaper and safer than any interop route** — `[INFERRED]`, but note that a port is a derivative work, so it inherits the GPL question exactly; the licence is not dodged by porting. ### 2.2 The wall this project has already hit — and why BitBrain does not get around it The brief's premise is right and the repo's own numbers back it: - **The TM head scored below its own majority class.** `common_libs/tests/tm_pattern_sweep_results.md:626-631`: majority class 2 at **38.77%**, head accuracy **36.69%** (642473/1751067), **−2.08 pp**. The shuffled-label control sits at chance (20.04% vs 20.12%). - **Live side accuracy sat at chance.** The same file records online radial-head accuracy at **48.8% vs 19.9% shuffled** (`:273`) and the gated side head at **40.37%** vs a 38.75% majority (`:649`) — i.e. +1.6 pp for a learner with a gate. - **Every TM arm tied or lost to the shipped Pattern gun.** `docs/gun_rack_summary.md`: rack overall **6.95%** real hit rate, Tsetlin **5.8%**; disabling Tsetlin gave 6.18% → 5.76% (p = 0.57) `[VERIFIED]`. - **The break-even bar is already quantified.** `common_libs/guns/tm_horizon.nim:49-53`: gate 2b measured that a correcting head "needs ~80% side accuracy to break even on hits while the achievable signal is ~60%". - And this repo already ran a **16-family / ~260-config shootout** on the *same* aim target: `BNNBot_garage/analysis/shootout_results.md` — WiSARD (best generalist, 16.20 px MAE), Tsetlin (20.69 px, loses on all 4 enemies to cold start), Kanerva SDM (10.16 px, "no improvement on baseline"), Bloom filter (17.08 px, "worse than baseline"), HDC/VSA (broken), against a linear+Hebbian baseline at 17.61 px. That last point is the one to sit with. **BitBrain's two stages are already represented in that shootout:** | BitBrain stage | Already evaluated as | Result there | |---|---|---| | AD layer (random address decoders + threshold) | **WiSARD / n-tuple LUT nodes** (`BNNBot_garage/src/wisard_predictor.nim`, K=14, 50 nodes) | Best of 16 families, 16.20 px — and the repo still ships Pattern | | SBC coincidence memory + counting readout | **Bloom filter** (bit array, write-once) and **Kanerva SDM** (address→content memory) | 17.08 px and 10.16 px, neither beat the baseline on the metric that mattered | | (nearest TM analogue) | Tsetlin gun, TMHorizon | 5.8% real vs 6.95% rack; TMHorizon predicted to lose | `[INFERRED]` So BitBrain is not a new kind of learner on this problem — it is a *sharper addressing scheme* on top of the same primitive the repo already tried, and its readout (binary "have I seen this coincidence with this label") is strictly weaker than WiSARD's running-mean readout, because it throws away the co-occurrence counts. Its one genuine advantage over the Tsetlin machinery is that it is **non-stochastic and order-free** (§1.7) — no learning rate to tune, no cold-start pathology, no clause-length collapse. **What target signal would make BitBrain worth building?** Stated as a gate, mirroring the repo's own Gate-2b language: a target whose *offline* learnability, measured on held-out resolved shots, clears **≥80% side accuracy (or equivalent break-even on the hit metric)** and clears a **shuffled-label control at chance**. The published evidence is that the available side signal is ~60% `[VERIFIED]` (`tm_horizon.nim:49-53`). No learner change moves that number, and BitBrain is a learner change. **Porting BitBrain before finding such a target would be the third repetition of the same mistake** — a learner swapped onto a signal-poor target — and it would cost the AD-layer problem of §1.3 on top. ### 2.3 The runtime budget — measured, not extrapolated Per the repo's own cost harness, the live budget is **~13.16 ms/tick (76 ticks/s)** (`common_libs/tests/measure_tm_pattern_cost.nim:6-9`); the TM Pattern gun measures **0.36 ms/tick** and Tsetlin ~**5.3 ms/tick** (`common_libs/guns/tm_horizon.nim:55-57`). **The brief's "1 ms/tick, TMHorizon 0.19 ms, Tsetlin 1.94 ms" numbers do not match what the repo records** — I am reporting the repo's numbers and flagging the discrepancy rather than guessing which is current. Either way the conclusion below does not change. I compiled and ran the C. **MEASURED** (gcc 15.3.0 `-O3`, single-threaded, MNIST geometry 784 → 2048 ADEs × 4 layers → 6 × 5 MB SBC tensors): | Operation | Measured cost | |---|---:| | Stage 1: 4-layer `find_AD_firing_pattern` forward pass | **0.3599 ms/sample** | | Stage 3: shipped `read_from_sbc` (buggy, 6 memories) | **0.412 ms/sample** | | Stage 3: truncation removed (6 memories) | 0.531 ms/sample | | Stage 3: **sparse active-list read**, no truncation | **0.202 ms/sample** | | **Full inference = stage 1 + sparse stage 3** | **≈ 0.56 ms/sample** | | Stage 2: `write_to_sbc` training (dense scan, 6 memories) | 0.600 ms/sample | | Integration harness overhead, all 6 readers | 0.736 → 5.330 ms/sample as memories fill | | Full original program: 60,000 train + 10,000 test | **66 s wall**, `[MEASURED] ./bb` | Two things to read off this: - **Stage 1 is 0.36 ms; the shipped stage-3 reader is *more expensive* than the whole feature layer (0.41 ms), because `read_from_sbc` scans all 2048×2048 `j` slots per firing `i` irrespective of sparsity.** The sparse reader — iterate the ~25-element active list per layer instead of a 2048-slot scan — is **2.6× faster than the fixed reader and 2.0× faster than the shipped one**, and produces *bit-identical accuracy* (97.210% both) `[MEASURED]`. This is a free 2× win for any port. - **0.56 ms/sample is affordable.** It is in the same class as the TM Pattern gun's 0.36 ms/tick, and it is a hard, non-iterative bound — no epochs, no retraining pass. Note however that the 0.202 ms sparse read is dominated by *cache misses* (~5,400 scattered probes into six 5 MB arrays per sample), not by arithmetic; a shooting-configuration SBC memory would be far smaller and proportionally faster. `[INFERRED]` - **Do NOT copy the tensor size.** 2048×2048×10 bits × 6 is 31.5 MB of L2/L3-hostile memory. A gun-sized version (say 512 ADEs × 2 layers × 6 classes) is ~0.8 MB and would not have this behaviour. `[INFERRED]` ### 2.4 The online-learning constraint — satisfied `ModularBot` must learn inside one battle and must **not** persist across battles (state is wiped when the target changes — see the `resetLearning` contract documented at `common_libs/guns/tm_horizon.nim:37-40`, which is the repo's statement of that rule). BitBrain's trainable part, judged against that `[VERIFIED]` from `129-152`: | Requirement | BitBrain | Verdict | |---|---|---| | Per-sample update, no batch | `write_to_sbc` is called per sample | ✅ | | Order-independent | idempotent `BITSET` | ✅ | | No learning rate / hyperparameter to tune | none exists | ✅ | | Usable at inference before/while training | `read_from_sbc` reads whatever is set | ✅ | | Survives across rounds of the same battle | monotone accumulation | ✅ | | Wipes cleanly on enemy change | clear 31.5 MB (`memset`/re-`calloc`) | ✅ (cost, not correctness) | | Does not require old samples to be replayed | yes | ✅ | | Needs pretrained features from outside the battle | **yes — the ADs** | ⚠️ **This is the failure** | So the *learner* satisfies the constraint perfectly; the *feature extractor* is a pre-battle artifact, which is exactly the thing `resetLearning` cannot regenerate and exactly the thing we do not have the program for. ### 2.5 Recommendation **Do not port BitBrain as a gun now.** Reasoning, in priority order: 1. **The AD layer is unobtainable** (§1.3). This is not a nice-to-have; without it BitBrain is a counting head on top of a feature layer we design ourselves — at which point BitBrain's contribution is the counting head alone, and §2.2 already tells us what happens to learner-only changes on this target. 2. **The signal is the wall, and it is already quantified** (§2.2): −2.08 pp vs majority for the TM head, live side accuracy at 47–49% (chance), break-even needing ~80% where the achievable signal is ~60%. BitBrain does not touch that. 3. **The primitive is already in the repo's evaluated set** (§2.2): WiSARD *is* the AD layer, Bloom/SDM are the SBC memory, and none of them beat the incumbent on the metric that mattered. 4. Cost is **not** the blocker (§2.3): ≈0.56 ms/sample is affordable. So the decision rests entirely on 1–3, and it comes out negative. **What would change the verdict** — the cheap, honest next step, if the user wants one: - **Reuse the AD primitive, not the learner.** `BNNBot_garage/src/wisard_predictor.nim` already implements an n-tuple address-decoder layer with a running-mean readout and eligibility traces, and it already won a 16-family shootout. If BitBrain's *counting* readout is the interesting part, the experiment worth running is a **readout swap on that existing feature extractor**, measured against the same fixture the shootout used — an offline, no-battle-change experiment, not a port. - **Gate it exactly like Gate 2b.** Do not build a gun until an offline corpus of resolved shots shows a target signal that clears the ≥80% side-accuracy break-even *and* a shuffled-label control at chance. If it does not, the answer is "no" regardless of learner. - If a port is ever justified: fix the §1.6 truncation bug, use the sparse active-list reader (§2.3, 2.6× faster for free), size the SBC tensors to the gun's feature count (not 2048²), normalise the per-class score by stored-feature count to kill the class-frequency bias (§1.5), and settle the GPL-3.0 question **before** writing the wrapper. **How this could still fail even if we do everything right:** (a) a synthesised AD layer is statistically similar but not the same, and there is no MNIST-comparable baseline to validate it against (§1.3.1); (b) the class-frequency bias with a few hundred labels per class will be much worse than on balanced MNIST (§1.5); (c) `find_AD_firing_pattern` needs the whole input matrix resident and indexes by row (`98,116`), so a per-tick port must serialise each observation into a matrix row — cheap, but it is the kind of plumbing that hides latency; and (d) the 31.5 MB tensor if copied verbatim is a cache disaster (§2.3). --- ## 3. Should the 11 MB zip be committed or gitignored? **Recommendation: commit it, under its explicit path `docs/BitBrain_C_code.zip`.** Facts behind that call `[MEASURED]`/`[VERIFIED]`: - `git status` shows `docs/BitBrain_C_code.zip` as the only untracked file besides `common_libs/tests/measure_aim_vs_power.nim`, and there is no `.gitignore` rule covering it — consistent with the repo's stated convention of listing **explicit binary names** (`.gitignore` names `common_libs/tests/tm_measure`, `tr_bots/`, `*_garage/out/`, etc.) rather than broad `*.zip` patterns. - The repo's norms comfortably accommodate this size: it already tracks eleven ~11 MB zips (`SAC_LSTM_Bot_garage/weights*/sac_*.zip`) and PDFs up to 9.3 MB (`docs/papers/neuroevolution/2023_neuroevolution_recurrent_architectures.pdf`). `.git` is already 117 MB. - It is a **third-party primary source that cannot be re-derived** from anything in the repo — the weights, the thresholds, and the MNIST dump exist nowhere else here. If it is gitignored, every citation in this document becomes unverifiable for the next reader. Tracked `docs/papers/*.pdf` show the repo already treats third-party primary sources as trackable content. - Caveat, if the user would rather not carry it: the **explicit** rule to add is `docs/BitBrain_C_code.zip` (matching the existing style), and *also* `docs/BitBrain_C_code/` for the 55 MB extracted tree, which should never be committed in either case. Do not add `*.zip` — that would silently swallow the SAC weight archives. --- ## 4. Appendix — reproducing every measurement ```bash mkdir -p /tmp/bitbrain && cd /tmp/bitbrain unzip /home/davide/Projects/SirRoboGarage/docs/BitBrain_C_code.zip cd BitBrain_C_code gcc full_mnist_2048.c -O3 -lm -o bb && time ./bb # -> 346 wrong = 96.540 pct correct, 66 s ``` Harnesses I wrote (throwaway, in `/tmp/bitbrain/measure/`, not in the repo): | File | What it measures | |---|---| | `stats.c` | AD fire rates, 4-layer forward-pass timing | | `thresh.c` | per-ADE fire-rate distributions, sign/position statistics, dead-ADEs | | `exp.c` | online learning curve, SBC occupancy, dense vs sparse read timing | | `final.c` | shipped vs fixed vs sparse reader accuracy + per-sample cost on all 10,000 tests | | `dbg*.c` | the §1.6 truncation-bug isolation chain | All three readers in `final.c` contain **verbatim copies** of `find_AD_firing_pattern` (98-124), `write_to_sbc` (129-152) and `read_from_sbc` (157-181); the only differences are the bug fix and the sparse traversal. The as-shipped reader reproduces the original program's 96.540% exactly, which is the check that the copies are faithful.