# Unconventional CS — Domain Survey for BNNBot Context: 690-bit binary input (10-frame temporal window, Gray-coded), reward signal from wave hit system (miss distance), ~1ms/tick budget, no gradients, no pre-training, online learning only. Task: predict enemy position (continuous x,y output). --- ## 1. Reservoir Computing / Echo State Networks ### How it works Fixed random recurrent layer (reservoir) transforms temporal input into a high-dimensional nonlinear state. Only the linear readout is trained (ridge regression or online RLS). No backpropagation through the reservoir. ### Core principle Untrained chaos is still useful: the reservoir expands low-dimensional input into a rich trajectory through state-space. The readout just needs to find a linear slice. ### Mapping to our problem **We already ARE doing this.** The 690-bit encoding is a handcrafted reservoir: 10-frame temporal window, Gray coding, sin/cos projections. The Hebbian residual table is the (very shallow) readout. The architecture philosophy section of RESEARCH.md explicitly names this. ### Can the reservoir adapt with reward? Yes — two mechanisms: - **Intrinsic Plasticity (IP):** Local unsupervised rule that tunes each neuron's gain/bias so its output distribution matches a target exponential. Maximizes information throughput without reward. Updates: `a += eta*(1/a - x*tanh(b+a*x))`, `b += eta*(-tanh(b+a*x))`. Purely local, O(N) per step. - **Hebbian Architecture Generation (HAG, 2025):** Grows connections between frequently co-activating neurons, sculpting task-specific wiring from a sparse seed. Nature Comms 2025 paper shows HAG beats IP and Anti-Oja across classification and forecasting tasks. - **Reward-modulated STDP:** Neuromodulatory signal (reward) gates whether recent correlational changes are committed. Well-studied in spiking ESNs. ### Update rule sketch For reward-modulated reservoir adaptation: ``` # Per tick, after observing reward r: for each edge (i,j) in reservoir: eligibility_ij += pre_i * post_j # accumulate Hebbian trace w_ij += alpha * r * eligibility_ij eligibility_ij *= decay # exponential trace decay ``` Readout (ridge): `w_out = (X^T X + lambda I)^{-1} X^T y`, or online RLS with O(n^2) update. For our 690-dim input, n=690, so RLS matrix is 690x690 = ~380K floats — fine for 1ms budget. ### Assessment for BNNBot - **Low risk, incremental gain.** We already have the structure; adding IP or reward- modulated reservoir edges to the binary encoding could unlock better feature representations without breaking the readout. - Readout upgrade from Hebbian table to online RLS/LMS is the lowest-hanging fruit. - Depth: shallow (reservoir + linear readout). Getting deeper is the challenge this whole survey is about. --- ## 2. Random Boolean Networks (RBNs / Kauffman Networks) ### How it works N binary nodes, each receiving K random inputs and assigned a random Boolean function (truth table of size 2^K). Iterated synchronously. No weights — each node has a lookup table of 2^K bits. ### Critical regime At K=2, p=0.5: the network sits at the "edge of chaos". Small perturbations neither die out (ordered, K<2) nor explode (chaotic, K>2). Adaptive robots using K=2 RBNs outperform K<2 and K>2 variants (Entropy 2022 paper). Crucially: reward-driven training via genetic algorithm *naturally converges to K≈2*, not by design but because K=2 is the adaptive optimum. ### Core principle Computation via attractor dynamics. Inputs push the network into different basins of attraction; the fixed point or limit cycle encodes the "answer". ### Mapping to our problem 690-bit input → seed the RBN state. Let it run T steps → read out N bit aggregate as prediction. The 690 truth tables (2^K bits each) are the parameters to learn. Reward signal: mutate truth tables of poorly-performing nodes (those whose contribution correlates with miss) and keep mutations that improve hit rate. ### Concrete update rule ``` # Evolutionary strategy on Boolean functions: for each node i where contribution_score[i] < threshold: flip one random bit in truth_table[i] # mutation evaluate on recent history if miss_distance worse: revert ``` Or stochastic: with prob proportional to miss distance, flip bits in random node truth tables. ### Computational cost Forward pass: N XOR/table lookups per step × T steps. For N=690, T=10: 6900 lookups per tick. Trivially fast. Training: O(N) per reward signal. ### Depth and credit assignment Depth = T (number of synchronous update steps). Credit assignment is the hard part: which node's truth table caused the miss? No natural gradient. Options: - Perturbation-based: change one node's table, observe reward change. O(N) samples needed per gradient estimate — too slow online. - Structural: nodes that are "downstream" of the input in the network topology get blamed first (topological credit assignment). - Caveat: RBNs are primarily studied as models of gene regulatory networks, not as general function approximators. Convergence to a target function is not guaranteed. ### Assessment for BNNBot - **Exotic and uncertain.** The attractor dynamics are not well-suited to continuous regression (they produce binary outputs, need majority-vote or thermometer readout). - The critical-regime insight is philosophically interesting — it suggests that binary networks naturally self-organize to K≈2 with adaptive pressure, which might inform how we design the connectivity of a BNN. - Not recommended as primary approach but interesting structural inspiration. --- ## 3. Tsetlin Machines (TMs) — DEEP DIVE ### How it works A TM is a team of Tsetlin Automata (TAs) that learns propositional logic clauses from binary inputs. Each clause is a conjunction (AND) of literals (features or their negations): e.g., `x3 AND NOT x7 AND x12`. Each TA controls whether its literal is Included or Excluded in its clause. TAs use a state machine: states 1..2N, midpoint divides Exclude (states 1..N) from Include (N+1..2N). Moving right → more committed to Include; moving left → more committed to Exclude. The output is a vote: sum of (positive clauses - negative clauses). Classification: sign of vote. Regression: the raw vote divided by the number of clauses. ### Exact update rules (Type I and II feedback) Let `c` be a clause, `o` its output (0/1), `y` the label (0/1 for classification), `s` a specificity parameter (typically 2-10), and `x_i` the literal value: **Type I Feedback** (given to positive-polarity clauses when `y=1`, prob 1/max(1,v)): - **Type Ia** (when `o=1`): with prob `(s-1)/s`, if `x_i=1`, Reward Include TA (move right) - **Type Ib** (when `o=0` or `x_i=0`): with prob `1/s`, Penalize Include TA (move left), Reward Exclude TA (move right) **Type II Feedback** (given to positive-polarity clauses when `y=0`, prob 1/max(1,v)): - When `o=1` and `x_i=0`: Penalize Exclude TA (move left) — force inclusion of distinguishing features to fire only when correct Where `v` is the clamped vote sum: `v = clip(sum_clauses, -T, T)` — the threshold T controls the effective voting range. As `|v|` grows, the probability of feedback decreases, creating a homeostatic balance that prevents over-fitting. **Regression TM (RTM):** No sign — raw vote is the output. Loss is `(y_hat - y)`. Feedback probabilities become functions of the error magnitude rather than binary correct/wrong. Specifically: - If `y_hat > y` (over-prediction): Type II feedback to positive clauses (shrink them) - If `y_hat < y` (under-prediction): Type I feedback to positive clauses (grow them) The exact probability for RTM: `p_feedback = clip(|y_hat - y| / y_max, 0, 1)` ### Multi-layer / Deep TMs July 2025 paper "The Tsetlin Machine Goes Deep: Logical Learning and Reasoning With Graphs" (arxiv 2507.14874) introduces hierarchical TM layers where clause outputs from one layer become binary inputs to the next. This creates hierarchical logical expressions — exactly what we need for multi-layer binary networks without gradients. Key mechanism: the output of layer L (a binary vector of clause activations) feeds directly as input bits to layer L+1. Each layer still uses its own Type I/II feedback. Credit assignment flows through the logical structure, not gradients. ### Coalesced Multi-Output TM For predicting (x, y) position simultaneously: the Coalesced TM shares clauses across multiple outputs, reducing parameter count. Each clause contributes to multiple outputs with different polarity, saving memory and improving generalization. ### Mapping to our problem - Input: 690 binary bits (our existing encoding) — **native input format** - Output: continuous position (x, y) — use Regression TM with two outputs - Learning: wave hit reward gives `(hit_x - pred_x, hit_y - pred_y)` error signal directly usable as RTM feedback - No gradients, no backprop, pure reinforcement-like TA state updates - Online: each wave hit = one training example, update TAs immediately ### Architecture sketch ``` 690 bits → [TM Layer 1: M1 clauses, each max K1 literals] → binary clause activations (M1 bits) → [TM Layer 2: M2 clauses] (optional depth) → RTM readout: vote → predicted x, y ``` ### Computational cost - Clause evaluation: for each clause, check K literals. Bitwise AND on 64-bit words. For 690 inputs: ceil(690/64)=11 words. M=100 clauses: 11×100 = 1100 AND ops/tick. Trivially within 1ms. - TA updates: one update per TA per training example = M×690 state increments. M=100 clauses: 69000 integer ops per wave hit. Fast. - Memory: M × 690 TA states, each 1 byte = 69KB for 1000 clauses. Fine. ### Convergence Mathematically proven to converge for IDENTITY and NOT operators (arxiv 2007.14268). Regression convergence: empirically shown on benchmark datasets, no formal proof yet. Online convergence: slow vs batch but works — the stochastic nature averages out over many examples. Typical battle has ~1000 wave closures = 1000 training examples. ### Why TM is purpose-built for this problem 1. **Binary input native**: 690 bits processed as-is, no float conversion 2. **No gradients**: reinforcement-style TA updates only 3. **Online**: each hit event updates TAs in place 4. **Interpretable**: resulting clauses are readable Boolean rules 5. **Regression extension**: continuous x,y output is directly supported 6. **Depth available**: multi-layer version published mid-2025 ### The catch - TMs learn propositional logic — they find which binary features co-occur with good predictions. Our input is already heavily engineered so this is appropriate. - The `s` parameter controls generalization vs specificity — requires tuning. - Clause count M is a capacity knob. Too few: underfitting. Too many: slow convergence. - For regression, the voting mechanism needs the output range to be known (or clipped). Miss distance is bounded by arena diagonal (~1131px) — manageable. **Verdict: HIGHEST PRIORITY candidate. TM is essentially designed for this exact problem: binary input, reinforcement reward, online, no gradients, continuous output available.** --- ## 4. Learning Classifier Systems (LCS / XCS) ### How it works A population of if-then rules: each rule is a ternary string `{0, 1, #}^690` (# = don't care) matched against input. Rules that match vote; vote is aggregated; reward distributed back via Q-learning (XCS) or bucket brigade (original Holland). ### Core principle Genetic algorithm discovers useful rules; RL credit-assigns reward through chains of rules. Population pressure keeps only accurate, general rules. ### Mapping to our problem - Match condition: 690-bit ternary string. Each # reduces specificity (don't care = any). - Prediction: each rule has a prediction value (learned float). Matching rules' weighted average = final prediction. - Reward: wave miss distance → penalize recently activated rules; wave hit → reward them. ### Concrete update rule (XCS Q-learning variant) ``` # On wave closure with miss d at power p: matched = [r for r in population if r.condition matches current_input] reward = max_d - d # inverted miss distance for r in matched: r.prediction += beta * (reward - r.prediction) r.error += beta * (|reward - r.prediction| - r.error) r.fitness = 1 / r.error # accuracy-based # Periodically: GA on matched set to generate new rules ``` ### Computational cost - Matching: 690-bit pattern match per rule × population size. Population = 1000 rules, 690 bits → 11 words per match → 11000 AND+XOR ops per tick. Fast. - GA: triggers infrequently. Population replacement amortizes cost. ### Depth None natively. Rules fire independently, no composition. XCSR (real-valued XCS) and XCSF extend to function approximation but add complexity. ### What the 1990s knew Holland's bucket brigade was THE solution to credit assignment before Q-learning formalized it. The insight: rules form chains (rule A enables condition for rule B), and credit flows backward through the chain like tokens in a market. Deep learning rediscovered this as temporal credit assignment. LCS communities were doing it first, with interpretable symbolic rules. ### Assessment for BNNBot - Competitive approach for moderate population sizes and simple rules. - Weaker on continuous output than TM (needs XCSF extension). - GA adds noise during learning — convergence in a single battle (few hundred updates) may be too slow. - Interesting for its interpretability: resulting rules are human-readable. - **Medium priority.** More complex than TM, less theoretically grounded for this task. --- ## 5. Swarm Intelligence / Ant Colony Optimization (ACO) on Binary Weights ### How it works Each binary weight is a choice between 0/1. Maintain a pheromone table `tau[i][b]` (probability that weight i = b). Each "ant" samples a weight vector, runs a forward pass, gets reward, deposits pheromone proportional to reward. ### Core principle Collective memory of good weight configurations, without storing weights explicitly — only their probability distribution. Biased random search that concentrates where previous successes occurred. ### Mapping to our problem ``` # Pheromone matrix: tau[i] in (0,1) = probability weight_i = 1 # Each tick: sample weights w_i ~ Bernoulli(tau[i]) # Run forward pass, get prediction, wait for wave closure for reward # On reward r: for i in range(n_weights): if w_i == 1: tau[i] += rho * r * (1 - tau[i]) else: tau[i] -= rho * r * tau[i] # Evaporation: tau *= (1 - evaporation_rate) ``` ### Computational cost - N pheromone values, one Bernoulli sample per weight = N random calls per tick. - For 690 inputs × H hidden = 690H float ops. For H=100: 69K ops/tick. Fine. - Credit assignment: problem. We sample weights at tick T, wave closes at tick T+k (variable latency). We must correlate which weight sample produced which prediction. Need to store (weight_sample, prediction) pairs per wave. ### Assessment for BNNBot - Natural for binary weights. - The delayed reward (wave hits arrive T+latency ticks later) requires careful bookkeeping — exactly what the wave system already does. - Convergence is slow for high-dimensional binary spaces; pheromone evaporation fights stagnation but also fights convergence. - ACO on continuous regression outputs is non-standard; closest is Estimation of Distribution Algorithms (EDAs) like PBIL. - **Low-medium priority.** Works but likely slower convergence than TM per battle. --- ## 6. Hyperdimensional Computing (HDC) ### How it works Represent everything as D-dimensional binary (or bipolar {-1,+1}) vectors, D=1000-10000. Operations: - **Bind**: XOR (or element-wise multiply for bipolar) — creates unique vector for combination, dissimilar to components - **Bundle**: majority vote — creates vector similar to all inputs - **Permute**: circular shift — encodes position/order Learning: accumulate positive examples into a "class prototype" vector by bundling; subtract negative examples. ### Regression via RegHD RegHD (DAC 2021) clusters similar inputs into groups, learns a linear regression model per group. Prediction = weighted sum across group models by similarity. Online update: when new (input, target) arrives, find most similar group, update its model. KalmanHD (ASP-DAC 2024) adds Kalman filtering to the readout for time-series forecasting, handling non-stationarity. ### Mapping to our problem - Our 690-bit input IS already a hypervector (nearly the right dimension). - Encode each temporal frame as a hypervector; bind across time positions (permute frame i by i); bundle all 10 frames → single D-bit context vector. - Learn an associative memory: context vector → (x_pred, y_pred). - Online update: when wave closes, update the associative memory entry. ### Update rule (bipolar) ``` # Encode input: H = majority(permute(frame_i, i) for i in 1..10) # Query: find stored vector V* most similar to H (Hamming distance) # Predict: y_pred = V*.regression_weights @ H # On wave close with actual y: err = y - y_pred V*.regression_weights += alpha * err * H ``` ### Computational cost - Encoding: 10 rotations × 690 bits = trivial. - Query: Hamming distance between H and each stored prototype. For K=50 prototypes: 50 × 690-bit XOR + popcount = 50 × 11 SIMD ops. Sub-microsecond. - Update: vector addition, O(D). Fast. ### Depth None natively. HDC is a single-layer associative architecture. Composition via binding enables some structure but not deep hierarchical computation. ### What's compelling - **Completely gradient-free** by design. - The existing 690-bit encoding is already "HDC-ready." - Extremely fast inference (bitwise ops). - Online update is exactly what we need: each wave = one update. - Robust to noise and bit errors — important since binary encoding has quantization. ### The limitation - Regression accuracy degrades vs neural approaches on complex nonlinear functions. - The codebook (stored prototypes) can fragment if too many distinct input regions. - No proven depth mechanism. **Assessment: MEDIUM-HIGH priority.** Low implementation cost (our encoding is already HDC-compatible), gradient-free, online, fast. Less powerful than TM for complex logic patterns but simpler to implement correctly. Worth a quick prototype. --- ## 7. Genetic Programming / Cartesian Genetic Programming (CGP) ### How it works CGP: a grid of nodes, each computing a function (AND, OR, XOR, NAND, etc.) of two inputs from earlier in the grid. The "chromosome" encodes which function each node uses and which earlier nodes it connects to. Evolution (mutation + selection) improves the circuit. Self-Modifying CGP (SMCGP): the evolved program can modify its own structure during execution — learns a learning algorithm, not just a function. ### Mapping to our problem - Evolve a Boolean circuit that maps 690 bits → prediction encoding. - Chromosome: node functions + connections. Mutations: change one function or reconnect one edge. - Fitness: wave hit reward (miss distance). ### Update rule ``` # Online evolution variant (1+1 ES on chromosome): mutation = mutate_one_node(current_chromosome) y_mut = evaluate(mutation, input) y_curr = evaluate(current_chromosome, input) if reward(y_mut) >= reward(y_curr): current_chromosome = mutation ``` ### Computational cost - Circuit evaluation: traversal of DAG, O(nodes). For 100 nodes: fast. - Fitness evaluation requires waiting for wave closure — same latency as other methods. - Selection pressure is very low online (1 wave = 1 fitness evaluation). ### Assessment for BNNBot - CGP is powerful for Boolean circuit discovery but **requires many fitness evaluations to converge.** A 690-input circuit needs hundreds of good-quality examples before the EA finds a useful structure. One battle (~200 wave hits) is probably insufficient. - The 1+1 ES variant is too slow for credit assignment through depth. - **Low priority for this problem.** Might be interesting for evolving the Boolean function form of individual "neurons" in a fixed-topology network. --- ## 8. Amorphous Computing ### How it works Large numbers of identical, simple agents (cells), each knowing only local state and local neighborhood. No central controller. Emergent behavior from local rules. MIT "Amorphous Computing Manifesto" (Abelson, Knight, Sussman, 1996). ### Core principle Robustness through redundancy. No single point of failure. Computation arises from the aggregate, not any individual. ### Mapping to our problem Interpret each of the 690 input bits as an "agent" that has a local rule: based on my bit value and my neighbors' bit values, output 0 or 1. The aggregate output of all agents = prediction. Reward modulates the rules: bits whose recent activations correlate with reward keep their rules; others randomize. ### Assessment for BNNBot - Beautiful concept, impractical for a function approximation problem with a continuous output target. Amorphous computing is good for pattern formation, self-assembly, robust sensing — not regression. - The "agents as bits" mapping loses the distinction between input features: all bits are equivalent, but in our encoding they represent very different things (distance vs heading vs velocity). - **Skip.** Not a good fit for the problem structure. --- ## 9. Thermodynamic Computing / Boltzmann Machines ### How it works Energy-based model: joint distribution over visible (input) and hidden units defined by `P(v,h) ∝ exp(-E(v,h))` where `E = -v^T W h - b^T v - c^T h`. Training via Contrastive Divergence (CD): approximate the gradient of log-likelihood using short Gibbs chains. **CD IS gradient descent** — it approximates `∂log P / ∂W`. This violates the no- gradient constraint if we mean parameter gradients. However, CD can be viewed as: 1. Run Gibbs sampler from data (positive phase) 2. Run Gibbs sampler freely (negative phase) 3. `ΔW = eta * (E[v h^T]_data - E[v h^T]_model)` The positive/negative phase update is not a gradient in the backprop sense — it uses only local Hebbian correlations. No chain rule, no derivative computation. ### Binary Boltzmann Machine With binary units: `h_j = sigmoid(W_j * v + c_j) > random`. All operations are binary samples. The weight update `ΔW = v_data * h_data - v_model * h_model` is pure Hebbian multiplication — local, no chain rule. ### Mapping to our problem - Use as a generative model of (input, position) pairs. - Train unsupervised on observed (input, outcome) pairs from wave hits. - Query: clamp input bits, sample hidden and output units, read prediction. ### Assessment for BNNBot - The Gibbs sampling for query (inference) is iterative and slow — multiple passes needed per prediction. Bad for 1ms budget. - The model is generative, not discriminative — it models P(input, output) not P(output | input). Conditioning is approximate. - CD is technically a gradient method (gradient of log-likelihood approximated by Gibbs sampling). Borderline against our constraints. - **Low priority.** Conceptually interesting but inference cost and gradient-adjacent training make it a poor fit. --- ## 10. Program Synthesis / Inductive Logic Programming (ILP) / Version Spaces ### How it works **Version spaces (Mitchell 1982):** Maintain the set of all hypotheses consistent with observed examples. Represented by most-specific (S) and most-general (G) boundary sets. Each new example eliminates inconsistent hypotheses. At convergence, S = G = unique correct hypothesis. **ILP:** Learn logic programs (Prolog-style rules) from positive and negative examples. Hypothesis is a set of Horn clauses. Operators: generalization (relax conditions), specialization (add conditions). ### Mapping to our problem - Each wave hit is a (binary_input, true_position) example. - Learn a logic program: `predict_x(Input, X) :- feature_a(Input), feature_b(Input), X is some_function`. - Version space: maintain set of consistent Boolean formulas over 690 bits predicting position within tolerance. ### The fundamental problem Version spaces require consistent (noise-free) examples. Our wave data has noise (enemy jitters, quantization, Gray coding). ILP hypothesis space is exponential in the number of features. For 690 binary features, the hypothesis space is 2^690. Version space collapse (from noise) and exponential search make this intractable at our scale. ### What the 1990s ILP community knew The key insight: **fewer features = tractable learning.** ILP works beautifully when the representation is already close to the logical structure of the problem. Our 690-bit encoding is over-specified for ILP — it's good for numeric approximation, not symbolic rule learning. However: the underlying insight that learning = hypothesis elimination is powerful. The TM can be seen as doing approximate ILP via stochastic clause learning. ### Assessment for BNNBot - **Skip in raw form.** Intractable at 690-feature scale. - The ILP insight informs the TM approach: learn propositional clauses online, which is tractable ILP restricted to propositional logic. --- ## Synthesis: What the 1990s Knew That Deep Learning Made Us Forget 1. **Credit assignment without gradients is solved.** Bucket brigade (Holland 1986), Q-learning (Watkins 1989), TA reinforcement (Tsetlin 1961) — all predate backpropagation's dominance. Deep learning won because it scales; these algorithms are often better when the input is already binary/symbolic. 2. **The representation IS the algorithm.** ILP, LCS, and version spaces all force you to think hard about the input language before learning. Deep learning outsources this to gradient descent. Our handcrafted 690-bit encoding is more 1990s than 2020s — and that's appropriate for the constraints. 3. **Population-based search finds structure without local minima.** GA, GP, ACO avoid the dead-end attractors of gradient descent. But they require many evaluations — the trade-off is evaluation efficiency vs search freedom. 4. **Reservoir computing predates deep learning.** ESN/LSM (Jaeger 2001, Maass 2002) showed that untrained recurrence + linear readout beats fully trained RNNs in many online settings. We're doing this implicitly already. 5. **Boolean logic is a valid computation substrate.** The TM rediscovers that conjunctive rules + voting is a universal approximator when the input is binary. Deep learning's obsession with continuous weights was never mandatory. --- ## Priority Ranking for BNNBot Implementation | # | Approach | Fit | Cost | Risk | Notes | |---|----------|-----|------|------|-------| | 1 | **Regression TM (RTM)** | Excellent | Medium | Low | Purpose-built for binary→continuous, online, no gradients | | 2 | **Deep TM (multi-layer)** | Very good | Medium | Medium | 2025 paper; hierarchical logic; credit through logical structure | | 3 | **HDC + online regression** | Good | Low | Low | 690-bit already HDC-ready; gradient-free; fast | | 4 | **Adaptive reservoir (IP + reward-modulated STDP)** | Good | Low | Low | Incremental upgrade to current architecture | | 5 | **XCS/XCSF** | Moderate | High | Medium | Works but GA convergence slow in single-battle | | 6 | **ACO on binary weights** | Moderate | Medium | High | Delayed reward bookkeeping complex; slow convergence | | 7 | **RBN** | Low | Low | High | No continuous output natively; credit assignment unsolved | | 8 | **CGP** | Low | Low | High | Needs too many evaluations per battle | | 9 | **Boltzmann Machine** | Low | High | High | Inference too slow; CD is gradient-adjacent | | 10 | **Amorphous / ILP / Version Space** | Very low | — | — | Mismatched to continuous regression task | --- ## Actionable Next Steps 1. **Implement Regression TM in Nim.** Binary input is native. Use `s=3..5`, `T=500`, `M=200` clauses as starting point. Two independent RTMs for x and y prediction. Each wave hit = one online update. Replace the Hebbian residual table. 2. **Test HDC as a cheaper baseline.** The 690-bit encoding already works as a hypervector. Add a similarity-based lookup table (K=20 prototypes) with online linear regression weights per prototype. ~50 lines of code. 3. **Add intrinsic plasticity to the binary encoding layer** (optional). Tune the gain/threshold of each bit position so its activation rate targets a target distribution. No reward signal needed — purely unsupervised entropy maximization. 4. **Consider multi-layer TM** only after single-layer RTM baseline is established. The 2025 "Goes Deep" paper is the reference. Credit assignment between layers uses the binary clause output as the inter-layer information carrier — no gradient. --- ## Sources - [Frontiers: Stochastic and Deterministic Tsetlin Machine](https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2025.1377944/full) - [Regression Tsetlin Machine (arxiv 1905.04206)](https://arxiv.org/abs/1905.04206) - [Tsetlin Machine Goes Deep (arxiv 2507.14874)](https://arxiv.org/pdf/2507.14874) - [Coalesced Multi-Output TM (arxiv 2108.07594)](https://arxiv.org/pdf/2108.07594) - [Self-timed RL with Tsetlin Machine (arxiv 2109.00846)](https://arxiv.org/pdf/2109.00846) - [Reshaping Reservoirs with Hebbian Adaptation — Nature Comms 2025](https://www.nature.com/articles/s41467-025-67137-1) - [Online Reservoir Adaptation by Intrinsic Plasticity — ScienceDirect](https://www.sciencedirect.com/science/article/abs/pii/S0893608007000317) - [On the Criticality of Adaptive Boolean Network Robots (Entropy 2022)](https://doi.org/10.3390/e24101368) - [RegHD: Regression in Hyperdimensional Computing (DAC 2021)](https://dl.acm.org/doi/10.1109/DAC18074.2021.9586284) - [KalmanHD: Time Series with HDC (ASP-DAC 2024)](https://github.com/DarthIV02/KalmanHD) - [Learning Classifier Systems Complete Intro (Urbanowicz 2009)](https://onlinelibrary.wiley.com/doi/10.1155/2009/736398) - [A Brief History of LCS (arxiv 1401.3607)](https://arxiv.org/pdf/1401.3607) - [Boosting Reservoir with Brain-inspired Adaptive Dynamics (arxiv 2504.12480)](https://arxiv.org/pdf/2504.12480) - [Amorphous Computing — CACM](https://cacm.acm.org/research/amorphous-computing/) - [Version Space Learning — Wikipedia](https://en.wikipedia.org/wiki/Version_space_learning)