feat(ModularBot): 6 guns, pattern matcher, melee modules, adversarial bots
- New guns: guess-factor (GF histogram), pattern-matcher (movement tape replay) - New modules: minimum-risk melee movement, spinning melee radar - New test bots: PatternMover, RandomMover, WaveSurfer - Fixed: FeedbackEvent now carries actualX/actualY for proper GF learning - Fixed: TM gun warmup gating + directional residuals - Fixed: circular gun integrated formula + multi-bin omega cache - Fixed: oscillator wall-bounce lockout - Fixed: phantom meteor perpendicular body orientation - 6/6 battle wins across all enemy types
This commit is contained in:
@@ -0,0 +1,625 @@
|
||||
# Unconventional CS — Domain Survey for BNNBot
|
||||
|
||||
Context: 690-bit binary input (10-frame temporal window, Gray-coded), reward signal
|
||||
from wave hit system (miss distance), ~1ms/tick budget, no gradients, no pre-training,
|
||||
online learning only. Task: predict enemy position (continuous x,y output).
|
||||
|
||||
---
|
||||
|
||||
## 1. Reservoir Computing / Echo State Networks
|
||||
|
||||
### How it works
|
||||
Fixed random recurrent layer (reservoir) transforms temporal input into a
|
||||
high-dimensional nonlinear state. Only the linear readout is trained (ridge regression
|
||||
or online RLS). No backpropagation through the reservoir.
|
||||
|
||||
### Core principle
|
||||
Untrained chaos is still useful: the reservoir expands low-dimensional input into a
|
||||
rich trajectory through state-space. The readout just needs to find a linear slice.
|
||||
|
||||
### Mapping to our problem
|
||||
**We already ARE doing this.** The 690-bit encoding is a handcrafted reservoir:
|
||||
10-frame temporal window, Gray coding, sin/cos projections. The Hebbian residual
|
||||
table is the (very shallow) readout. The architecture philosophy section of RESEARCH.md
|
||||
explicitly names this.
|
||||
|
||||
### Can the reservoir adapt with reward?
|
||||
Yes — two mechanisms:
|
||||
- **Intrinsic Plasticity (IP):** Local unsupervised rule that tunes each neuron's
|
||||
gain/bias so its output distribution matches a target exponential. Maximizes
|
||||
information throughput without reward. Updates: `a += eta*(1/a - x*tanh(b+a*x))`,
|
||||
`b += eta*(-tanh(b+a*x))`. Purely local, O(N) per step.
|
||||
- **Hebbian Architecture Generation (HAG, 2025):** Grows connections between
|
||||
frequently co-activating neurons, sculpting task-specific wiring from a sparse seed.
|
||||
Nature Comms 2025 paper shows HAG beats IP and Anti-Oja across classification and
|
||||
forecasting tasks.
|
||||
- **Reward-modulated STDP:** Neuromodulatory signal (reward) gates whether recent
|
||||
correlational changes are committed. Well-studied in spiking ESNs.
|
||||
|
||||
### Update rule sketch
|
||||
For reward-modulated reservoir adaptation:
|
||||
```
|
||||
# Per tick, after observing reward r:
|
||||
for each edge (i,j) in reservoir:
|
||||
eligibility_ij += pre_i * post_j # accumulate Hebbian trace
|
||||
w_ij += alpha * r * eligibility_ij
|
||||
eligibility_ij *= decay # exponential trace decay
|
||||
```
|
||||
Readout (ridge): `w_out = (X^T X + lambda I)^{-1} X^T y`, or online RLS with O(n^2)
|
||||
update. For our 690-dim input, n=690, so RLS matrix is 690x690 = ~380K floats — fine
|
||||
for 1ms budget.
|
||||
|
||||
### Assessment for BNNBot
|
||||
- **Low risk, incremental gain.** We already have the structure; adding IP or reward-
|
||||
modulated reservoir edges to the binary encoding could unlock better feature
|
||||
representations without breaking the readout.
|
||||
- Readout upgrade from Hebbian table to online RLS/LMS is the lowest-hanging fruit.
|
||||
- Depth: shallow (reservoir + linear readout). Getting deeper is the challenge this
|
||||
whole survey is about.
|
||||
|
||||
---
|
||||
|
||||
## 2. Random Boolean Networks (RBNs / Kauffman Networks)
|
||||
|
||||
### How it works
|
||||
N binary nodes, each receiving K random inputs and assigned a random Boolean function
|
||||
(truth table of size 2^K). Iterated synchronously. No weights — each node has a
|
||||
lookup table of 2^K bits.
|
||||
|
||||
### Critical regime
|
||||
At K=2, p=0.5: the network sits at the "edge of chaos". Small perturbations neither
|
||||
die out (ordered, K<2) nor explode (chaotic, K>2). Adaptive robots using K=2 RBNs
|
||||
outperform K<2 and K>2 variants (Entropy 2022 paper). Crucially: reward-driven
|
||||
training via genetic algorithm *naturally converges to K≈2*, not by design but because
|
||||
K=2 is the adaptive optimum.
|
||||
|
||||
### Core principle
|
||||
Computation via attractor dynamics. Inputs push the network into different basins of
|
||||
attraction; the fixed point or limit cycle encodes the "answer".
|
||||
|
||||
### Mapping to our problem
|
||||
690-bit input → seed the RBN state. Let it run T steps → read out N bit aggregate as
|
||||
prediction. The 690 truth tables (2^K bits each) are the parameters to learn.
|
||||
|
||||
Reward signal: mutate truth tables of poorly-performing nodes (those whose contribution
|
||||
correlates with miss) and keep mutations that improve hit rate.
|
||||
|
||||
### Concrete update rule
|
||||
```
|
||||
# Evolutionary strategy on Boolean functions:
|
||||
for each node i where contribution_score[i] < threshold:
|
||||
flip one random bit in truth_table[i] # mutation
|
||||
evaluate on recent history
|
||||
if miss_distance worse: revert
|
||||
```
|
||||
Or stochastic: with prob proportional to miss distance, flip bits in random node
|
||||
truth tables.
|
||||
|
||||
### Computational cost
|
||||
Forward pass: N XOR/table lookups per step × T steps. For N=690, T=10: 6900 lookups
|
||||
per tick. Trivially fast.
|
||||
Training: O(N) per reward signal.
|
||||
|
||||
### Depth and credit assignment
|
||||
Depth = T (number of synchronous update steps). Credit assignment is the hard part:
|
||||
which node's truth table caused the miss? No natural gradient. Options:
|
||||
- Perturbation-based: change one node's table, observe reward change. O(N) samples
|
||||
needed per gradient estimate — too slow online.
|
||||
- Structural: nodes that are "downstream" of the input in the network topology get
|
||||
blamed first (topological credit assignment).
|
||||
- Caveat: RBNs are primarily studied as models of gene regulatory networks, not as
|
||||
general function approximators. Convergence to a target function is not guaranteed.
|
||||
|
||||
### Assessment for BNNBot
|
||||
- **Exotic and uncertain.** The attractor dynamics are not well-suited to continuous
|
||||
regression (they produce binary outputs, need majority-vote or thermometer readout).
|
||||
- The critical-regime insight is philosophically interesting — it suggests that binary
|
||||
networks naturally self-organize to K≈2 with adaptive pressure, which might inform
|
||||
how we design the connectivity of a BNN.
|
||||
- Not recommended as primary approach but interesting structural inspiration.
|
||||
|
||||
---
|
||||
|
||||
## 3. Tsetlin Machines (TMs) — DEEP DIVE
|
||||
|
||||
### How it works
|
||||
A TM is a team of Tsetlin Automata (TAs) that learns propositional logic clauses from
|
||||
binary inputs. Each clause is a conjunction (AND) of literals (features or their
|
||||
negations): e.g., `x3 AND NOT x7 AND x12`. Each TA controls whether its literal is
|
||||
Included or Excluded in its clause. TAs use a state machine: states 1..2N, midpoint
|
||||
divides Exclude (states 1..N) from Include (N+1..2N). Moving right → more committed to
|
||||
Include; moving left → more committed to Exclude.
|
||||
|
||||
The output is a vote: sum of (positive clauses - negative clauses). Classification:
|
||||
sign of vote. Regression: the raw vote divided by the number of clauses.
|
||||
|
||||
### Exact update rules (Type I and II feedback)
|
||||
|
||||
Let `c` be a clause, `o` its output (0/1), `y` the label (0/1 for classification),
|
||||
`s` a specificity parameter (typically 2-10), and `x_i` the literal value:
|
||||
|
||||
**Type I Feedback** (given to positive-polarity clauses when `y=1`, prob 1/max(1,v)):
|
||||
- **Type Ia** (when `o=1`): with prob `(s-1)/s`, if `x_i=1`, Reward Include TA (move right)
|
||||
- **Type Ib** (when `o=0` or `x_i=0`): with prob `1/s`, Penalize Include TA (move left),
|
||||
Reward Exclude TA (move right)
|
||||
|
||||
**Type II Feedback** (given to positive-polarity clauses when `y=0`, prob 1/max(1,v)):
|
||||
- When `o=1` and `x_i=0`: Penalize Exclude TA (move left) — force inclusion of
|
||||
distinguishing features to fire only when correct
|
||||
|
||||
Where `v` is the clamped vote sum: `v = clip(sum_clauses, -T, T)` — the threshold T
|
||||
controls the effective voting range. As `|v|` grows, the probability of feedback
|
||||
decreases, creating a homeostatic balance that prevents over-fitting.
|
||||
|
||||
**Regression TM (RTM):** No sign — raw vote is the output. Loss is `(y_hat - y)`.
|
||||
Feedback probabilities become functions of the error magnitude rather than binary
|
||||
correct/wrong. Specifically:
|
||||
- If `y_hat > y` (over-prediction): Type II feedback to positive clauses (shrink them)
|
||||
- If `y_hat < y` (under-prediction): Type I feedback to positive clauses (grow them)
|
||||
The exact probability for RTM: `p_feedback = clip(|y_hat - y| / y_max, 0, 1)`
|
||||
|
||||
### Multi-layer / Deep TMs
|
||||
July 2025 paper "The Tsetlin Machine Goes Deep: Logical Learning and Reasoning With
|
||||
Graphs" (arxiv 2507.14874) introduces hierarchical TM layers where clause outputs from
|
||||
one layer become binary inputs to the next. This creates hierarchical logical
|
||||
expressions — exactly what we need for multi-layer binary networks without gradients.
|
||||
|
||||
Key mechanism: the output of layer L (a binary vector of clause activations) feeds
|
||||
directly as input bits to layer L+1. Each layer still uses its own Type I/II feedback.
|
||||
Credit assignment flows through the logical structure, not gradients.
|
||||
|
||||
### Coalesced Multi-Output TM
|
||||
For predicting (x, y) position simultaneously: the Coalesced TM shares clauses across
|
||||
multiple outputs, reducing parameter count. Each clause contributes to multiple outputs
|
||||
with different polarity, saving memory and improving generalization.
|
||||
|
||||
### Mapping to our problem
|
||||
- Input: 690 binary bits (our existing encoding) — **native input format**
|
||||
- Output: continuous position (x, y) — use Regression TM with two outputs
|
||||
- Learning: wave hit reward gives `(hit_x - pred_x, hit_y - pred_y)` error signal
|
||||
directly usable as RTM feedback
|
||||
- No gradients, no backprop, pure reinforcement-like TA state updates
|
||||
- Online: each wave hit = one training example, update TAs immediately
|
||||
|
||||
### Architecture sketch
|
||||
```
|
||||
690 bits → [TM Layer 1: M1 clauses, each max K1 literals]
|
||||
→ binary clause activations (M1 bits)
|
||||
→ [TM Layer 2: M2 clauses] (optional depth)
|
||||
→ RTM readout: vote → predicted x, y
|
||||
```
|
||||
|
||||
### Computational cost
|
||||
- Clause evaluation: for each clause, check K literals. Bitwise AND on 64-bit words.
|
||||
For 690 inputs: ceil(690/64)=11 words. M=100 clauses: 11×100 = 1100 AND ops/tick.
|
||||
Trivially within 1ms.
|
||||
- TA updates: one update per TA per training example = M×690 state increments.
|
||||
M=100 clauses: 69000 integer ops per wave hit. Fast.
|
||||
- Memory: M × 690 TA states, each 1 byte = 69KB for 1000 clauses. Fine.
|
||||
|
||||
### Convergence
|
||||
Mathematically proven to converge for IDENTITY and NOT operators (arxiv 2007.14268).
|
||||
Regression convergence: empirically shown on benchmark datasets, no formal proof yet.
|
||||
Online convergence: slow vs batch but works — the stochastic nature averages out over
|
||||
many examples. Typical battle has ~1000 wave closures = 1000 training examples.
|
||||
|
||||
### Why TM is purpose-built for this problem
|
||||
1. **Binary input native**: 690 bits processed as-is, no float conversion
|
||||
2. **No gradients**: reinforcement-style TA updates only
|
||||
3. **Online**: each hit event updates TAs in place
|
||||
4. **Interpretable**: resulting clauses are readable Boolean rules
|
||||
5. **Regression extension**: continuous x,y output is directly supported
|
||||
6. **Depth available**: multi-layer version published mid-2025
|
||||
|
||||
### The catch
|
||||
- TMs learn propositional logic — they find which binary features co-occur with good
|
||||
predictions. Our input is already heavily engineered so this is appropriate.
|
||||
- The `s` parameter controls generalization vs specificity — requires tuning.
|
||||
- Clause count M is a capacity knob. Too few: underfitting. Too many: slow convergence.
|
||||
- For regression, the voting mechanism needs the output range to be known (or clipped).
|
||||
Miss distance is bounded by arena diagonal (~1131px) — manageable.
|
||||
|
||||
**Verdict: HIGHEST PRIORITY candidate. TM is essentially designed for this exact
|
||||
problem: binary input, reinforcement reward, online, no gradients, continuous output
|
||||
available.**
|
||||
|
||||
---
|
||||
|
||||
## 4. Learning Classifier Systems (LCS / XCS)
|
||||
|
||||
### How it works
|
||||
A population of if-then rules: each rule is a ternary string `{0, 1, #}^690` (# = don't
|
||||
care) matched against input. Rules that match vote; vote is aggregated; reward
|
||||
distributed back via Q-learning (XCS) or bucket brigade (original Holland).
|
||||
|
||||
### Core principle
|
||||
Genetic algorithm discovers useful rules; RL credit-assigns reward through chains of
|
||||
rules. Population pressure keeps only accurate, general rules.
|
||||
|
||||
### Mapping to our problem
|
||||
- Match condition: 690-bit ternary string. Each # reduces specificity (don't care = any).
|
||||
- Prediction: each rule has a prediction value (learned float). Matching rules' weighted
|
||||
average = final prediction.
|
||||
- Reward: wave miss distance → penalize recently activated rules; wave hit → reward them.
|
||||
|
||||
### Concrete update rule (XCS Q-learning variant)
|
||||
```
|
||||
# On wave closure with miss d at power p:
|
||||
matched = [r for r in population if r.condition matches current_input]
|
||||
reward = max_d - d # inverted miss distance
|
||||
for r in matched:
|
||||
r.prediction += beta * (reward - r.prediction)
|
||||
r.error += beta * (|reward - r.prediction| - r.error)
|
||||
r.fitness = 1 / r.error # accuracy-based
|
||||
# Periodically: GA on matched set to generate new rules
|
||||
```
|
||||
|
||||
### Computational cost
|
||||
- Matching: 690-bit pattern match per rule × population size. Population = 1000 rules,
|
||||
690 bits → 11 words per match → 11000 AND+XOR ops per tick. Fast.
|
||||
- GA: triggers infrequently. Population replacement amortizes cost.
|
||||
|
||||
### Depth
|
||||
None natively. Rules fire independently, no composition. XCSR (real-valued XCS) and
|
||||
XCSF extend to function approximation but add complexity.
|
||||
|
||||
### What the 1990s knew
|
||||
Holland's bucket brigade was THE solution to credit assignment before Q-learning
|
||||
formalized it. The insight: rules form chains (rule A enables condition for rule B),
|
||||
and credit flows backward through the chain like tokens in a market. Deep learning
|
||||
rediscovered this as temporal credit assignment. LCS communities were doing it first,
|
||||
with interpretable symbolic rules.
|
||||
|
||||
### Assessment for BNNBot
|
||||
- Competitive approach for moderate population sizes and simple rules.
|
||||
- Weaker on continuous output than TM (needs XCSF extension).
|
||||
- GA adds noise during learning — convergence in a single battle (few hundred updates)
|
||||
may be too slow.
|
||||
- Interesting for its interpretability: resulting rules are human-readable.
|
||||
- **Medium priority.** More complex than TM, less theoretically grounded for this task.
|
||||
|
||||
---
|
||||
|
||||
## 5. Swarm Intelligence / Ant Colony Optimization (ACO) on Binary Weights
|
||||
|
||||
### How it works
|
||||
Each binary weight is a choice between 0/1. Maintain a pheromone table `tau[i][b]`
|
||||
(probability that weight i = b). Each "ant" samples a weight vector, runs a forward
|
||||
pass, gets reward, deposits pheromone proportional to reward.
|
||||
|
||||
### Core principle
|
||||
Collective memory of good weight configurations, without storing weights explicitly —
|
||||
only their probability distribution. Biased random search that concentrates where
|
||||
previous successes occurred.
|
||||
|
||||
### Mapping to our problem
|
||||
```
|
||||
# Pheromone matrix: tau[i] in (0,1) = probability weight_i = 1
|
||||
# Each tick: sample weights w_i ~ Bernoulli(tau[i])
|
||||
# Run forward pass, get prediction, wait for wave closure for reward
|
||||
# On reward r:
|
||||
for i in range(n_weights):
|
||||
if w_i == 1: tau[i] += rho * r * (1 - tau[i])
|
||||
else: tau[i] -= rho * r * tau[i]
|
||||
# Evaporation:
|
||||
tau *= (1 - evaporation_rate)
|
||||
```
|
||||
|
||||
### Computational cost
|
||||
- N pheromone values, one Bernoulli sample per weight = N random calls per tick.
|
||||
- For 690 inputs × H hidden = 690H float ops. For H=100: 69K ops/tick. Fine.
|
||||
- Credit assignment: problem. We sample weights at tick T, wave closes at tick T+k
|
||||
(variable latency). We must correlate which weight sample produced which prediction.
|
||||
Need to store (weight_sample, prediction) pairs per wave.
|
||||
|
||||
### Assessment for BNNBot
|
||||
- Natural for binary weights.
|
||||
- The delayed reward (wave hits arrive T+latency ticks later) requires careful
|
||||
bookkeeping — exactly what the wave system already does.
|
||||
- Convergence is slow for high-dimensional binary spaces; pheromone evaporation fights
|
||||
stagnation but also fights convergence.
|
||||
- ACO on continuous regression outputs is non-standard; closest is Estimation of
|
||||
Distribution Algorithms (EDAs) like PBIL.
|
||||
- **Low-medium priority.** Works but likely slower convergence than TM per battle.
|
||||
|
||||
---
|
||||
|
||||
## 6. Hyperdimensional Computing (HDC)
|
||||
|
||||
### How it works
|
||||
Represent everything as D-dimensional binary (or bipolar {-1,+1}) vectors, D=1000-10000.
|
||||
Operations:
|
||||
- **Bind**: XOR (or element-wise multiply for bipolar) — creates unique vector for
|
||||
combination, dissimilar to components
|
||||
- **Bundle**: majority vote — creates vector similar to all inputs
|
||||
- **Permute**: circular shift — encodes position/order
|
||||
|
||||
Learning: accumulate positive examples into a "class prototype" vector by bundling;
|
||||
subtract negative examples.
|
||||
|
||||
### Regression via RegHD
|
||||
RegHD (DAC 2021) clusters similar inputs into groups, learns a linear regression model
|
||||
per group. Prediction = weighted sum across group models by similarity. Online update:
|
||||
when new (input, target) arrives, find most similar group, update its model.
|
||||
|
||||
KalmanHD (ASP-DAC 2024) adds Kalman filtering to the readout for time-series
|
||||
forecasting, handling non-stationarity.
|
||||
|
||||
### Mapping to our problem
|
||||
- Our 690-bit input IS already a hypervector (nearly the right dimension).
|
||||
- Encode each temporal frame as a hypervector; bind across time positions (permute
|
||||
frame i by i); bundle all 10 frames → single D-bit context vector.
|
||||
- Learn an associative memory: context vector → (x_pred, y_pred).
|
||||
- Online update: when wave closes, update the associative memory entry.
|
||||
|
||||
### Update rule (bipolar)
|
||||
```
|
||||
# Encode input: H = majority(permute(frame_i, i) for i in 1..10)
|
||||
# Query: find stored vector V* most similar to H (Hamming distance)
|
||||
# Predict: y_pred = V*.regression_weights @ H
|
||||
# On wave close with actual y:
|
||||
err = y - y_pred
|
||||
V*.regression_weights += alpha * err * H
|
||||
```
|
||||
|
||||
### Computational cost
|
||||
- Encoding: 10 rotations × 690 bits = trivial.
|
||||
- Query: Hamming distance between H and each stored prototype. For K=50 prototypes:
|
||||
50 × 690-bit XOR + popcount = 50 × 11 SIMD ops. Sub-microsecond.
|
||||
- Update: vector addition, O(D). Fast.
|
||||
|
||||
### Depth
|
||||
None natively. HDC is a single-layer associative architecture. Composition via binding
|
||||
enables some structure but not deep hierarchical computation.
|
||||
|
||||
### What's compelling
|
||||
- **Completely gradient-free** by design.
|
||||
- The existing 690-bit encoding is already "HDC-ready."
|
||||
- Extremely fast inference (bitwise ops).
|
||||
- Online update is exactly what we need: each wave = one update.
|
||||
- Robust to noise and bit errors — important since binary encoding has quantization.
|
||||
|
||||
### The limitation
|
||||
- Regression accuracy degrades vs neural approaches on complex nonlinear functions.
|
||||
- The codebook (stored prototypes) can fragment if too many distinct input regions.
|
||||
- No proven depth mechanism.
|
||||
|
||||
**Assessment: MEDIUM-HIGH priority.** Low implementation cost (our encoding is already
|
||||
HDC-compatible), gradient-free, online, fast. Less powerful than TM for complex logic
|
||||
patterns but simpler to implement correctly. Worth a quick prototype.
|
||||
|
||||
---
|
||||
|
||||
## 7. Genetic Programming / Cartesian Genetic Programming (CGP)
|
||||
|
||||
### How it works
|
||||
CGP: a grid of nodes, each computing a function (AND, OR, XOR, NAND, etc.) of two
|
||||
inputs from earlier in the grid. The "chromosome" encodes which function each node
|
||||
uses and which earlier nodes it connects to. Evolution (mutation + selection) improves
|
||||
the circuit.
|
||||
|
||||
Self-Modifying CGP (SMCGP): the evolved program can modify its own structure during
|
||||
execution — learns a learning algorithm, not just a function.
|
||||
|
||||
### Mapping to our problem
|
||||
- Evolve a Boolean circuit that maps 690 bits → prediction encoding.
|
||||
- Chromosome: node functions + connections. Mutations: change one function or
|
||||
reconnect one edge.
|
||||
- Fitness: wave hit reward (miss distance).
|
||||
|
||||
### Update rule
|
||||
```
|
||||
# Online evolution variant (1+1 ES on chromosome):
|
||||
mutation = mutate_one_node(current_chromosome)
|
||||
y_mut = evaluate(mutation, input)
|
||||
y_curr = evaluate(current_chromosome, input)
|
||||
if reward(y_mut) >= reward(y_curr):
|
||||
current_chromosome = mutation
|
||||
```
|
||||
|
||||
### Computational cost
|
||||
- Circuit evaluation: traversal of DAG, O(nodes). For 100 nodes: fast.
|
||||
- Fitness evaluation requires waiting for wave closure — same latency as other methods.
|
||||
- Selection pressure is very low online (1 wave = 1 fitness evaluation).
|
||||
|
||||
### Assessment for BNNBot
|
||||
- CGP is powerful for Boolean circuit discovery but **requires many fitness evaluations
|
||||
to converge.** A 690-input circuit needs hundreds of good-quality examples before
|
||||
the EA finds a useful structure. One battle (~200 wave hits) is probably insufficient.
|
||||
- The 1+1 ES variant is too slow for credit assignment through depth.
|
||||
- **Low priority for this problem.** Might be interesting for evolving the Boolean
|
||||
function form of individual "neurons" in a fixed-topology network.
|
||||
|
||||
---
|
||||
|
||||
## 8. Amorphous Computing
|
||||
|
||||
### How it works
|
||||
Large numbers of identical, simple agents (cells), each knowing only local state and
|
||||
local neighborhood. No central controller. Emergent behavior from local rules.
|
||||
MIT "Amorphous Computing Manifesto" (Abelson, Knight, Sussman, 1996).
|
||||
|
||||
### Core principle
|
||||
Robustness through redundancy. No single point of failure. Computation arises from
|
||||
the aggregate, not any individual.
|
||||
|
||||
### Mapping to our problem
|
||||
Interpret each of the 690 input bits as an "agent" that has a local rule: based on my
|
||||
bit value and my neighbors' bit values, output 0 or 1. The aggregate output of all
|
||||
agents = prediction.
|
||||
|
||||
Reward modulates the rules: bits whose recent activations correlate with reward keep
|
||||
their rules; others randomize.
|
||||
|
||||
### Assessment for BNNBot
|
||||
- Beautiful concept, impractical for a function approximation problem with a continuous
|
||||
output target. Amorphous computing is good for pattern formation, self-assembly,
|
||||
robust sensing — not regression.
|
||||
- The "agents as bits" mapping loses the distinction between input features: all bits
|
||||
are equivalent, but in our encoding they represent very different things (distance
|
||||
vs heading vs velocity).
|
||||
- **Skip.** Not a good fit for the problem structure.
|
||||
|
||||
---
|
||||
|
||||
## 9. Thermodynamic Computing / Boltzmann Machines
|
||||
|
||||
### How it works
|
||||
Energy-based model: joint distribution over visible (input) and hidden units defined
|
||||
by `P(v,h) ∝ exp(-E(v,h))` where `E = -v^T W h - b^T v - c^T h`. Training via
|
||||
Contrastive Divergence (CD): approximate the gradient of log-likelihood using short
|
||||
Gibbs chains.
|
||||
|
||||
**CD IS gradient descent** — it approximates `∂log P / ∂W`. This violates the no-
|
||||
gradient constraint if we mean parameter gradients. However, CD can be viewed as:
|
||||
1. Run Gibbs sampler from data (positive phase)
|
||||
2. Run Gibbs sampler freely (negative phase)
|
||||
3. `ΔW = eta * (E[v h^T]_data - E[v h^T]_model)`
|
||||
|
||||
The positive/negative phase update is not a gradient in the backprop sense — it uses
|
||||
only local Hebbian correlations. No chain rule, no derivative computation.
|
||||
|
||||
### Binary Boltzmann Machine
|
||||
With binary units: `h_j = sigmoid(W_j * v + c_j) > random`. All operations are
|
||||
binary samples. The weight update `ΔW = v_data * h_data - v_model * h_model` is
|
||||
pure Hebbian multiplication — local, no chain rule.
|
||||
|
||||
### Mapping to our problem
|
||||
- Use as a generative model of (input, position) pairs.
|
||||
- Train unsupervised on observed (input, outcome) pairs from wave hits.
|
||||
- Query: clamp input bits, sample hidden and output units, read prediction.
|
||||
|
||||
### Assessment for BNNBot
|
||||
- The Gibbs sampling for query (inference) is iterative and slow — multiple passes
|
||||
needed per prediction. Bad for 1ms budget.
|
||||
- The model is generative, not discriminative — it models P(input, output) not
|
||||
P(output | input). Conditioning is approximate.
|
||||
- CD is technically a gradient method (gradient of log-likelihood approximated by
|
||||
Gibbs sampling). Borderline against our constraints.
|
||||
- **Low priority.** Conceptually interesting but inference cost and gradient-adjacent
|
||||
training make it a poor fit.
|
||||
|
||||
---
|
||||
|
||||
## 10. Program Synthesis / Inductive Logic Programming (ILP) / Version Spaces
|
||||
|
||||
### How it works
|
||||
**Version spaces (Mitchell 1982):** Maintain the set of all hypotheses consistent with
|
||||
observed examples. Represented by most-specific (S) and most-general (G) boundary sets.
|
||||
Each new example eliminates inconsistent hypotheses. At convergence, S = G = unique
|
||||
correct hypothesis.
|
||||
|
||||
**ILP:** Learn logic programs (Prolog-style rules) from positive and negative examples.
|
||||
Hypothesis is a set of Horn clauses. Operators: generalization (relax conditions),
|
||||
specialization (add conditions).
|
||||
|
||||
### Mapping to our problem
|
||||
- Each wave hit is a (binary_input, true_position) example.
|
||||
- Learn a logic program: `predict_x(Input, X) :- feature_a(Input), feature_b(Input), X is some_function`.
|
||||
- Version space: maintain set of consistent Boolean formulas over 690 bits predicting
|
||||
position within tolerance.
|
||||
|
||||
### The fundamental problem
|
||||
Version spaces require consistent (noise-free) examples. Our wave data has noise
|
||||
(enemy jitters, quantization, Gray coding). ILP hypothesis space is exponential in
|
||||
the number of features. For 690 binary features, the hypothesis space is 2^690.
|
||||
|
||||
Version space collapse (from noise) and exponential search make this intractable at
|
||||
our scale.
|
||||
|
||||
### What the 1990s ILP community knew
|
||||
The key insight: **fewer features = tractable learning.** ILP works beautifully when
|
||||
the representation is already close to the logical structure of the problem.
|
||||
Our 690-bit encoding is over-specified for ILP — it's good for numeric approximation,
|
||||
not symbolic rule learning.
|
||||
|
||||
However: the underlying insight that learning = hypothesis elimination is powerful.
|
||||
The TM can be seen as doing approximate ILP via stochastic clause learning.
|
||||
|
||||
### Assessment for BNNBot
|
||||
- **Skip in raw form.** Intractable at 690-feature scale.
|
||||
- The ILP insight informs the TM approach: learn propositional clauses online, which
|
||||
is tractable ILP restricted to propositional logic.
|
||||
|
||||
---
|
||||
|
||||
## Synthesis: What the 1990s Knew That Deep Learning Made Us Forget
|
||||
|
||||
1. **Credit assignment without gradients is solved.** Bucket brigade (Holland 1986),
|
||||
Q-learning (Watkins 1989), TA reinforcement (Tsetlin 1961) — all predate
|
||||
backpropagation's dominance. Deep learning won because it scales; these algorithms
|
||||
are often better when the input is already binary/symbolic.
|
||||
|
||||
2. **The representation IS the algorithm.** ILP, LCS, and version spaces all force you
|
||||
to think hard about the input language before learning. Deep learning outsources
|
||||
this to gradient descent. Our handcrafted 690-bit encoding is more 1990s than 2020s
|
||||
— and that's appropriate for the constraints.
|
||||
|
||||
3. **Population-based search finds structure without local minima.** GA, GP, ACO avoid
|
||||
the dead-end attractors of gradient descent. But they require many evaluations —
|
||||
the trade-off is evaluation efficiency vs search freedom.
|
||||
|
||||
4. **Reservoir computing predates deep learning.** ESN/LSM (Jaeger 2001, Maass 2002)
|
||||
showed that untrained recurrence + linear readout beats fully trained RNNs in many
|
||||
online settings. We're doing this implicitly already.
|
||||
|
||||
5. **Boolean logic is a valid computation substrate.** The TM rediscovers that
|
||||
conjunctive rules + voting is a universal approximator when the input is binary.
|
||||
Deep learning's obsession with continuous weights was never mandatory.
|
||||
|
||||
---
|
||||
|
||||
## Priority Ranking for BNNBot Implementation
|
||||
|
||||
| # | Approach | Fit | Cost | Risk | Notes |
|
||||
|---|----------|-----|------|------|-------|
|
||||
| 1 | **Regression TM (RTM)** | Excellent | Medium | Low | Purpose-built for binary→continuous, online, no gradients |
|
||||
| 2 | **Deep TM (multi-layer)** | Very good | Medium | Medium | 2025 paper; hierarchical logic; credit through logical structure |
|
||||
| 3 | **HDC + online regression** | Good | Low | Low | 690-bit already HDC-ready; gradient-free; fast |
|
||||
| 4 | **Adaptive reservoir (IP + reward-modulated STDP)** | Good | Low | Low | Incremental upgrade to current architecture |
|
||||
| 5 | **XCS/XCSF** | Moderate | High | Medium | Works but GA convergence slow in single-battle |
|
||||
| 6 | **ACO on binary weights** | Moderate | Medium | High | Delayed reward bookkeeping complex; slow convergence |
|
||||
| 7 | **RBN** | Low | Low | High | No continuous output natively; credit assignment unsolved |
|
||||
| 8 | **CGP** | Low | Low | High | Needs too many evaluations per battle |
|
||||
| 9 | **Boltzmann Machine** | Low | High | High | Inference too slow; CD is gradient-adjacent |
|
||||
| 10 | **Amorphous / ILP / Version Space** | Very low | — | — | Mismatched to continuous regression task |
|
||||
|
||||
---
|
||||
|
||||
## Actionable Next Steps
|
||||
|
||||
1. **Implement Regression TM in Nim.** Binary input is native. Use `s=3..5`, `T=500`,
|
||||
`M=200` clauses as starting point. Two independent RTMs for x and y prediction.
|
||||
Each wave hit = one online update. Replace the Hebbian residual table.
|
||||
|
||||
2. **Test HDC as a cheaper baseline.** The 690-bit encoding already works as a
|
||||
hypervector. Add a similarity-based lookup table (K=20 prototypes) with online
|
||||
linear regression weights per prototype. ~50 lines of code.
|
||||
|
||||
3. **Add intrinsic plasticity to the binary encoding layer** (optional). Tune the
|
||||
gain/threshold of each bit position so its activation rate targets a target
|
||||
distribution. No reward signal needed — purely unsupervised entropy maximization.
|
||||
|
||||
4. **Consider multi-layer TM** only after single-layer RTM baseline is established.
|
||||
The 2025 "Goes Deep" paper is the reference. Credit assignment between layers uses
|
||||
the binary clause output as the inter-layer information carrier — no gradient.
|
||||
|
||||
---
|
||||
|
||||
## Sources
|
||||
|
||||
- [Frontiers: Stochastic and Deterministic Tsetlin Machine](https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2025.1377944/full)
|
||||
- [Regression Tsetlin Machine (arxiv 1905.04206)](https://arxiv.org/abs/1905.04206)
|
||||
- [Tsetlin Machine Goes Deep (arxiv 2507.14874)](https://arxiv.org/pdf/2507.14874)
|
||||
- [Coalesced Multi-Output TM (arxiv 2108.07594)](https://arxiv.org/pdf/2108.07594)
|
||||
- [Self-timed RL with Tsetlin Machine (arxiv 2109.00846)](https://arxiv.org/pdf/2109.00846)
|
||||
- [Reshaping Reservoirs with Hebbian Adaptation — Nature Comms 2025](https://www.nature.com/articles/s41467-025-67137-1)
|
||||
- [Online Reservoir Adaptation by Intrinsic Plasticity — ScienceDirect](https://www.sciencedirect.com/science/article/abs/pii/S0893608007000317)
|
||||
- [On the Criticality of Adaptive Boolean Network Robots (Entropy 2022)](https://doi.org/10.3390/e24101368)
|
||||
- [RegHD: Regression in Hyperdimensional Computing (DAC 2021)](https://dl.acm.org/doi/10.1109/DAC18074.2021.9586284)
|
||||
- [KalmanHD: Time Series with HDC (ASP-DAC 2024)](https://github.com/DarthIV02/KalmanHD)
|
||||
- [Learning Classifier Systems Complete Intro (Urbanowicz 2009)](https://onlinelibrary.wiley.com/doi/10.1155/2009/736398)
|
||||
- [A Brief History of LCS (arxiv 1401.3607)](https://arxiv.org/pdf/1401.3607)
|
||||
- [Boosting Reservoir with Brain-inspired Adaptive Dynamics (arxiv 2504.12480)](https://arxiv.org/pdf/2504.12480)
|
||||
- [Amorphous Computing — CACM](https://cacm.acm.org/research/amorphous-computing/)
|
||||
- [Version Space Learning — Wikipedia](https://en.wikipedia.org/wiki/Version_space_learning)
|
||||
Reference in New Issue
Block a user