1ed7797cb6
- New guns: guess-factor (GF histogram), pattern-matcher (movement tape replay) - New modules: minimum-risk melee movement, spinning melee radar - New test bots: PatternMover, RandomMover, WaveSurfer - Fixed: FeedbackEvent now carries actualX/actualY for proper GF learning - Fixed: TM gun warmup gating + directional residuals - Fixed: circular gun integrated formula + multi-bin omega cache - Fixed: oscillator wall-bounce lockout - Fixed: phantom meteor perpendicular body orientation - 6/6 battle wins across all enemy types
561 lines
24 KiB
Markdown
561 lines
24 KiB
Markdown
# Electronics & Hardware Domain Analysis
|
||
## BNNBot Research — Learning Without Gradient Descent
|
||
|
||
**Context:** Binary network predicting enemy position in Robocode Tank Royale.
|
||
Constraints: no gradient descent, online learning only, ~1ms/tick, binary-friendly,
|
||
reward signal from wave-hit system (miss distance known per shot).
|
||
|
||
---
|
||
|
||
## 1. Memristive Systems
|
||
|
||
### How it works in hardware
|
||
A memristor's resistance is a function of its charge history. Pass current one direction:
|
||
resistance drops (LRS, low-resistance state = "1"). Reverse: resistance rises (HRS = "0").
|
||
In a crossbar array, each crosspoint is one synapse. Reading = multiply (V × G = I),
|
||
summing wires implement dot product in Ohm's law. Writing = apply voltage pulse.
|
||
|
||
Binary memristors (two-state, SET/RESET) switch between exactly LRS and HRS — no analog
|
||
level needed. Programming is by voltage threshold: exceed it → switch state.
|
||
|
||
### Core algorithmic principle
|
||
The natural learning rule is **STDP (Spike-Timing Dependent Plasticity)** — a Hebbian
|
||
variant gated by spike timing. With three-factor extensions, a third neuromodulatory
|
||
signal (reward / dopamine analog) gates the synaptic update:
|
||
|
||
```
|
||
ΔW = eligibility_trace × reward_signal
|
||
```
|
||
|
||
The eligibility trace is a leaky counter: it captures recent pre/post coincidence and
|
||
decays. When reward arrives (delayed), it multiplies the trace. This lets delayed rewards
|
||
credit the right synapses — the hardware equivalent of credit assignment without backprop.
|
||
|
||
For binary memristors specifically:
|
||
- Hebbian: if pre AND post both fired recently → SET (LRS)
|
||
- Anti-Hebbian: if pre fired but post did not (or vice versa) → RESET (HRS)
|
||
- Three-factor: gate SET/RESET on reward sign
|
||
|
||
### Mapping to BNNBot
|
||
- Our neuron is: `popcount(XOR(weights, input)) > threshold` — an n-input binary gate
|
||
- Weights are bits; SET/RESET maps to w[i] = 1 / w[i] = 0
|
||
- Eligibility trace = which weight bits were active when the shot was fired
|
||
- Wave hit reward = delayed reward signal
|
||
- Update: for active neurons on correct output → reinforce their weight bits that matched
|
||
|
||
### Concrete update rule (software sketch)
|
||
```nim
|
||
# On wave hit:
|
||
# active_bits[i] = which input bits were 1 when this wave was fired
|
||
# reward = 1.0 - (miss_distance / max_miss) # from wave system
|
||
|
||
for i in 0..<WEIGHT_BITS:
|
||
if active_bits[i] and reward > 0.0:
|
||
weights[i] = 1 # SET: pre and post active, positive reward
|
||
elif active_bits[i] and reward < 0.0:
|
||
weights[i] = 0 # RESET: active but wrong prediction
|
||
# inactive bits: no change (Hebbian locality)
|
||
```
|
||
|
||
### Strengths
|
||
- Naturally online, one update per wave hit
|
||
- Locality: each weight updates from local pre/post activity only
|
||
- Delayed reward handled by eligibility trace — no need to hold full state
|
||
|
||
### Weaknesses
|
||
- Binary weights = coarse; capacity scales with weight count not precision
|
||
- No credit through layers — only the output neuron's weights get updated
|
||
- Eligibility trace must be stored per weight per in-flight wave (memory cost)
|
||
|
||
---
|
||
|
||
## 2. FPGA / Evolvable Hardware
|
||
|
||
### How it works in hardware
|
||
An FPGA is a fabric of LUTs (Look-Up Tables) connected by a programmable routing
|
||
network. A K-LUT implements any K-input Boolean function by storing 2^K bits.
|
||
Intrinsic evolution: the actual FPGA bitstream (which LUTs do what, how they connect)
|
||
is treated as a genome. A genetic algorithm mutates it, measures fitness on real
|
||
hardware (actual timing, actual noise), and selects survivors.
|
||
|
||
The key insight from Thompson (1996): intrinsic evolution finds solutions that exploit
|
||
physical properties of the silicon that no designer would think to encode — clock
|
||
leakage, capacitive coupling across supposedly-disconnected wires. The circuit
|
||
"knows" things about its substrate that the designer doesn't model.
|
||
|
||
### Core algorithmic principle
|
||
**Genetic search over discrete configuration space.** No gradient — just:
|
||
1. Mutate bitstream (flip random bits)
|
||
2. Measure fitness on hardware
|
||
3. Keep better configurations, discard worse
|
||
4. Repeat
|
||
|
||
For LUT-based learning (LUTNet, WNN/WiSARD): treat each LUT's truth table as
|
||
writable memory. Training = writing new entries to LUTs. A K-input LUT can learn
|
||
any K-variable Boolean function in a single write — no iterative optimization needed
|
||
for patterns it has seen.
|
||
|
||
### Mapping to BNNBot
|
||
Each "neuron" in our network is a boolean function of its binary inputs.
|
||
An LUT IS the truth table for that function. Instead of storing a weight vector
|
||
and doing popcount, we store the actual input→output mapping.
|
||
|
||
For n inputs: LUT has 2^n entries. But n=690 is way too large.
|
||
Solution: each LUT sees only k<<690 bits (chosen by connectivity pattern).
|
||
The 690-bit input is partitioned into chunks; each LUT learns its local chunk.
|
||
|
||
For online reward-based LUT update:
|
||
- Record input index (the k-bit address into the LUT) when shot was fired
|
||
- On wave hit: write output[address] = 1 if reward > threshold, else 0
|
||
- This is exactly one-shot learning — the LUT "memorizes" each input pattern
|
||
|
||
### Concrete update rule
|
||
```nim
|
||
const K = 16 # LUT arity — covers 2^16 = 65536 patterns
|
||
# lut: array[2^K, bool]
|
||
|
||
# On shot fired:
|
||
let addr = extract_k_bits(input_690bit, chosen_indices, K)
|
||
pending_updates.add((addr, wave_id))
|
||
|
||
# On wave hit:
|
||
let (addr, _) = pending_updates[wave_id]
|
||
lut[addr] = (miss_distance < HIT_THRESHOLD)
|
||
```
|
||
|
||
### Strengths
|
||
- One-shot learning: each (input pattern → outcome) pair stored directly
|
||
- Zero arithmetic: inference is a table lookup, O(1), nanoseconds
|
||
- LUT generalizes to unseen inputs via nearest neighbor in address space
|
||
|
||
### Weaknesses
|
||
- 2^K memory per LUT — exponential in K. K=16 → 8KB per neuron.
|
||
- With K=16 and 690 inputs, coverage per LUT is sparse (690/16 = ~43 non-overlapping LUTs)
|
||
- Cold start: empty LUTs output 0; need warmup before predictions are useful
|
||
- No generalization across different address patterns (each address independent)
|
||
|
||
---
|
||
|
||
## 3. Hopfield Networks — Energy Minimization
|
||
|
||
### How it works in hardware
|
||
A Hopfield network is a fully connected binary recurrent network. Each node is +1/-1.
|
||
Update rule: `s_i = sign(Σ_j W_ij × s_j)`. Weights are symmetric (W_ij = W_ji).
|
||
The network has a Lyapunov energy function `E = -½ Σ W_ij s_i s_j`.
|
||
Asynchronous updates monotonically decrease E → network converges to a local minimum.
|
||
No gradient needed — local update decreases global energy by construction.
|
||
|
||
### Core algorithmic principle
|
||
**Pattern as energy minimum.** Training stores patterns p by:
|
||
`W_ij = (1/N) Σ_patterns p_i × p_j` (Hebbian outer product, no gradient).
|
||
Retrieval: start from noisy/partial pattern, run updates → snaps to nearest stored pattern.
|
||
The network "completes" corrupted inputs.
|
||
|
||
### Mapping to BNNBot
|
||
Treat the prediction problem as pattern completion:
|
||
- Input: partial state (last 10 radar scans = 690 bits)
|
||
- Stored patterns: historical (input_context, correct_output) pairs
|
||
- Query: clamp input bits, let output bits relax to minimum energy
|
||
|
||
This IS associative memory — we've seen this state-like context before, what happened?
|
||
|
||
Storing a (context → outcome) association:
|
||
```nim
|
||
# After wave hit with known outcome:
|
||
let pattern = concat(input_690bit, outcome_69bit) # 759-bit pattern
|
||
# Hebbian weight update (outer product, only for active bits):
|
||
for i in active_bits(pattern):
|
||
for j in active_bits(pattern):
|
||
W[i][j] += 1.0 / N # symmetric
|
||
```
|
||
|
||
Inference: clamp input 690 bits, iterate output 69 bits until convergence.
|
||
|
||
### Strengths
|
||
- Convergence guaranteed (for symmetric W, no self-loops)
|
||
- Naturally handles noisy/incomplete inputs — good for sensor noise
|
||
- No gradient, no explicit credit assignment needed
|
||
- OscNet v1.5 showed Hopfield runs forward-pass-only on hardware
|
||
|
||
### Weaknesses
|
||
- Capacity: standard Hopfield stores ≈0.14N patterns for N nodes
|
||
With N=759: only ~106 distinct (context, outcome) pairs — very limited
|
||
- O(N²) weight matrix: 759² = 576K entries for float weights
|
||
- Spurious minima: can settle to wrong pattern if too many stored
|
||
- Modern dense Hopfield (transformer attention) has exponential capacity but requires softmax
|
||
|
||
---
|
||
|
||
## 4. Weightless Neural Networks (WNN / WiSARD)
|
||
|
||
### How it works in hardware
|
||
A WNN neuron (RAM node) is literally an n-input lookup table with learned 1-bit entries.
|
||
No multiplication, no addition — just memory address lookup. Training: present input,
|
||
write 1 at that address. Inference: read address, return stored bit.
|
||
|
||
WiSARD (1981, commercial 1984): the original binary pattern recognizer.
|
||
Each RAM node sees k randomly-chosen bits from the input. Multiple RAM nodes
|
||
vote (sum of outputs) to produce a confidence score per class.
|
||
|
||
### Core algorithmic principle
|
||
**Address-based memorization.** The n-input binary vector IS the memory address.
|
||
Learning = `memory[address] = label`. No parameters to optimize, no loss surface.
|
||
Generalization comes from the probabilistic overlap of binary addresses.
|
||
|
||
For online RL adaptation: instead of one-shot supervised write, count-based update:
|
||
`memory[address] += reward_signal` then threshold. This is exactly a frequency table
|
||
— "how often did this input pattern lead to a hit?"
|
||
|
||
### Mapping to BNNBot
|
||
This is the most direct hardware→software translation:
|
||
|
||
```nim
|
||
const K = 20 # bits per RAM node
|
||
const N_NODES = 35 # 35 × 20 = 700 > 690
|
||
# Each node: table[2^20] of int8, initialized to 0
|
||
|
||
# Index mapping: for each node, which 20 of 690 bits does it observe?
|
||
# Fixed random permutation, same across all shots.
|
||
|
||
# On wave hit:
|
||
for node in 0..<N_NODES:
|
||
let addr = read_k_bits(input_690, node_mapping[node], K)
|
||
let delta = if hit: +1 else: -1
|
||
tables[node][addr] = clamp(tables[node][addr] + delta, -127, 127)
|
||
|
||
# Inference: which output direction gets highest vote?
|
||
for each candidate_direction:
|
||
score = 0
|
||
for node in 0..<N_NODES:
|
||
let addr = read_k_bits(input_690, node_mapping[node], K)
|
||
score += tables[node][addr]
|
||
if score > best_score: best_direction = candidate_direction
|
||
```
|
||
|
||
### Strengths
|
||
- Inference: pure memory lookup, O(N_NODES) — easily under 1ms
|
||
- Online learning: one table write per wave hit
|
||
- No cold-start paralysis — zero score = no preference, falls back to baseline
|
||
- K=20 handles 2^20 = 1M patterns per node — generous capacity
|
||
- Naturally compositional: each node learns a different k-bit "feature"
|
||
|
||
### Weaknesses
|
||
- Memory: 35 nodes × 2^20 × 1 byte = 35 MB. Large but feasible.
|
||
With K=16: 35 × 64KB = 2.2MB. More practical.
|
||
- Random k-bit projections may not capture the most informative bit combinations
|
||
(vs. learned projections in LUTNet-style approaches)
|
||
- No cross-node generalization: two patterns differing in one bit observed by the same
|
||
node = completely different addresses
|
||
|
||
**Note:** WNN with reward-based counting is the closest software analog to what
|
||
memristor crossbars do in hardware — the table IS the synapse.
|
||
|
||
---
|
||
|
||
## 5. Belief Propagation on Factor Graphs (LDPC / Turbo codes)
|
||
|
||
### How it works in hardware
|
||
LDPC decoding: received bits are noisy versions of codeword bits. The decoder has
|
||
a factor graph: variable nodes (bits) and check nodes (parity constraints).
|
||
Messages pass back and forth: each node sends to neighbors "given what I know, here's
|
||
my probability estimate of your value." After ~50 iterations, converge to most likely codeword.
|
||
|
||
Hardware decoders run BP in a pipelined dataflow — no global state, purely local message updates.
|
||
|
||
### Core algorithmic principle
|
||
**Iterative marginal inference.** Each variable node holds a log-likelihood ratio (LLR).
|
||
Update rule for variable node x_i receiving messages m_{j→i} from check nodes j:
|
||
```
|
||
LLR_i = channel_LLR_i + Σ_j m_{j→i}
|
||
m_{i→k} = LLR_i - m_{k→i} # exclude message from k before sending back
|
||
```
|
||
No gradient. Convergence is guaranteed on tree-structured graphs; approximate on loopy graphs.
|
||
|
||
### Mapping to BNNBot
|
||
The 690-bit input has structure — bits are correlated (sin/cos pairs, position from velocity,
|
||
etc.). BP could exploit this structure for inference without learning.
|
||
|
||
Concrete use: model the 69-bit output (next position frame) as hidden variables.
|
||
Observed: 690-bit input. Factor graph: local constraints from physics (velocity ×
|
||
ticks = displacement, heading consistency). Run BP to infer most-likely output.
|
||
|
||
This is not a learning mechanism — it's inference. But it applies to our problem:
|
||
the physics of enemy movement IS the factor graph. No training needed if the
|
||
constraints are correct.
|
||
|
||
### Concrete sketch
|
||
```nim
|
||
# Factor: expected_x[t+1] = x[t] + vx[t] * speed
|
||
# Factor: heading changes slowly (soft constraint, variance from data)
|
||
# Observed: x[t], y[t], vx[t], vy[t] from radar scans
|
||
|
||
# For each output bit o_i:
|
||
# LLR[i] = initial estimate from linear extrapolation
|
||
# Messages from physics factors update LLR iteratively
|
||
# Final: sign(LLR[i]) = predicted bit
|
||
```
|
||
|
||
### Strengths
|
||
- Exploits problem structure (physics) explicitly — not black-box learning
|
||
- Iterative refinement: run more iterations when time allows
|
||
- Error-correcting codes achieve Shannon capacity with BP — proven near-optimal
|
||
|
||
### Weaknesses
|
||
- Requires knowing the factor structure (constraint graph) — domain engineering needed
|
||
- Our enemy's movement IS physics, but also strategy (evasion) — hard to model as factors
|
||
- Less useful as a "learning" mechanism; more useful as a structured inference method
|
||
- Loopy BP on dense graphs may not converge
|
||
|
||
---
|
||
|
||
## 6. Sigma-Delta Modulation — Feedback Tracking
|
||
|
||
### How it works in hardware
|
||
A sigma-delta (ΣΔ) modulator converts analog input to a binary bit stream.
|
||
Architecture: integrator → 1-bit comparator → feedback DAC.
|
||
The comparator outputs 1 if accumulator > 0, else 0. The 1-bit output is fed back
|
||
and subtracted from the input. The accumulator tracks the error (sigma = cumulative,
|
||
delta = difference). Output bit stream: density encodes analog value.
|
||
|
||
### Core algorithmic principle
|
||
**Error-accumulating feedback.** The system never "knows" the true analog value —
|
||
it only knows the sign of accumulated error. It corrects continuously:
|
||
```
|
||
accumulator += (input - feedback)
|
||
output_bit = (accumulator > 0) ? 1 : 0
|
||
feedback = output_bit ? +1 : -1
|
||
```
|
||
This is a 1-bit quantizer with infinite-precision error memory. The feedback loop
|
||
drives the accumulated error to zero over time — the bit stream density equals the
|
||
input value.
|
||
|
||
### Mapping to BNNBot
|
||
The analogy: our prediction is a "bit stream" (sequence of predictions).
|
||
The wave hit system gives us the sign of the error (too far left? too far right?).
|
||
We accumulate signed errors and correct the prediction offset continuously.
|
||
|
||
This is closer to what our existing Hebbian residual table does — but the ΣΔ
|
||
formulation is cleaner:
|
||
|
||
```nim
|
||
# Per directional axis (x and y):
|
||
var accumulator_x: float = 0.0
|
||
var accumulator_y: float = 0.0
|
||
|
||
# On wave hit:
|
||
let error_x = actual_x - predicted_x
|
||
let error_y = actual_y - predicted_y
|
||
|
||
accumulator_x += error_x
|
||
accumulator_y += error_y
|
||
|
||
# Correction for next prediction:
|
||
correction_x = sign(accumulator_x) * correction_step # 1-bit correction
|
||
# Or: correction_x = clamp(accumulator_x * gain, -max_corr, max_corr) # proportional
|
||
```
|
||
|
||
The ΣΔ insight: **you don't need the full error magnitude, just its sign.**
|
||
The accumulation of signs over time encodes the magnitude.
|
||
|
||
### Strengths
|
||
- Minimal state: just one accumulator per axis
|
||
- Noise-shaping: high-frequency jitter averages out; low-frequency bias accumulates correctly
|
||
- Robust to measurement noise (sign of error is reliable even when magnitude is noisy)
|
||
|
||
### Weaknesses
|
||
- Slow convergence if correction step is too small
|
||
- Tracks only a DC offset (systematic bias), not complex patterns
|
||
- Best as a correction layer on top of a better predictor, not standalone
|
||
|
||
---
|
||
|
||
## 7. Phase-Locked Loop (PLL) — Trajectory Tracking
|
||
|
||
### How it works in hardware
|
||
A PLL: phase detector compares incoming signal phase to VCO output phase.
|
||
Error signal (phase difference) feeds through a loop filter to the VCO control input.
|
||
VCO adjusts frequency until phase error → 0. Locked: VCO tracks input frequency and phase.
|
||
|
||
Components: Phase Detector (XOR or mixer), Loop Filter (low-pass), VCO.
|
||
No gradient — just: am I ahead or behind? Correct proportionally.
|
||
|
||
### Core algorithmic principle
|
||
**Proportional-integral control on phase error.** PI loop filter:
|
||
```
|
||
phase_error = input_phase - vco_phase
|
||
control_voltage += Kp * phase_error + Ki * integral(phase_error)
|
||
vco_frequency = f0 + K_vco * control_voltage
|
||
```
|
||
This IS a gradient-free optimizer for tracking. It converges when phase_error = 0.
|
||
Equivalent to online least-mean-squares for frequency estimation.
|
||
|
||
### Mapping to BNNBot
|
||
Enemy position is a trajectory in 2D. Our aiming angle is a phase relative to that trajectory.
|
||
A PLL-like system: track the rate of change of enemy heading (angular velocity).
|
||
|
||
```nim
|
||
# "Phase" = enemy heading direction
|
||
# "VCO" = our predicted heading trend
|
||
# "Lock" = our trend matches enemy trend
|
||
|
||
var predicted_angular_velocity: float = 0.0
|
||
var phase_error_integral: float = 0.0
|
||
const Kp = 0.3
|
||
const Ki = 0.1
|
||
|
||
# On each radar scan:
|
||
let actual_angular_velocity = delta_heading / delta_time
|
||
let phase_error = actual_angular_velocity - predicted_angular_velocity
|
||
|
||
phase_error_integral += phase_error
|
||
predicted_angular_velocity += Kp * phase_error + Ki * phase_error_integral
|
||
```
|
||
|
||
When the enemy moves at constant angular velocity (orbit, spiral), this "locks on"
|
||
and predicts ahead by extrapolating the locked phase.
|
||
|
||
### Strengths
|
||
- Handles periodic/oscillatory motion natively (enemy orbiting = pure sine wave)
|
||
- Acquisition (from cold) + tracking (once locked) are separate phases — known behavior
|
||
- Loop filter order can be tuned: 1st order = tracks constant heading, 2nd = tracks constant angular acceleration
|
||
|
||
### Weaknesses
|
||
- Assumes quasi-periodic or smooth trajectory; breaks on jerky evasion
|
||
- Loop bandwidth tradeoff: wide bandwidth tracks fast changes but amplifies noise
|
||
- Enemy headings in Robocode are piecewise linear, not sinusoidal — PLL is mismatched model
|
||
unless enemy orbits, which some do
|
||
|
||
---
|
||
|
||
## 8. Stochastic Computing
|
||
|
||
### How it works in hardware
|
||
Represent numbers as the probability that a bit in a random stream is 1.
|
||
x = 0.7 → 70% of bits are 1 in a random stream. Then:
|
||
- Multiplication: AND(stream_x, stream_y) → probability = x × y
|
||
- Addition (scaled): MUX(stream_x, stream_y, select) → (x + y)/2
|
||
- Dot product: AND all streams, OR results
|
||
|
||
All operations are single-gate logic. No carry chains, no multipliers.
|
||
Inference hardware is tiny. Accuracy scales with stream length (more bits = more precise).
|
||
|
||
### Core algorithmic principle
|
||
**Probability represented as bit density.** The number IS the stream.
|
||
Computation IS gate operations on streams.
|
||
|
||
Weight update in stochastic computing:
|
||
```
|
||
Δw_ij = AND(activation_i_stream, error_j_stream)
|
||
```
|
||
This computes the product of two probabilities using an AND gate.
|
||
The learning rule IS Hebbian — coincidence detection between pre and post streams.
|
||
|
||
### Mapping to BNNBot
|
||
Our 690-bit input is already binary — but not stochastic (each bit is deterministic).
|
||
However, stochastic computing suggests a different representation:
|
||
instead of Gray-coded positions, represent uncertainty as bit-stream density.
|
||
|
||
Alternatively: treat each weight as a stochastic bit (1 with probability p).
|
||
Inference = sample weights → compute output → measure hit → update p.
|
||
This is essentially WNN with probabilistic weights.
|
||
|
||
Concrete weight update:
|
||
```nim
|
||
# Each weight w[i] is a probability stored as float in [0,1]
|
||
# On shot fired: sample binary weights: sampled[i] = rand() < w[i]
|
||
# On wave hit:
|
||
for i in active_neurons:
|
||
w[i] += learning_rate * reward * (sampled[i] - w[i])
|
||
# REINFORCE-like: push probability toward 1 if it fired and reward was positive
|
||
```
|
||
|
||
### Strengths
|
||
- Noise is a feature, not a bug — exploration built in
|
||
- Multiplication = AND, the cheapest gate — scales to large networks
|
||
- Graceful degradation: shorter streams = noisier but not broken
|
||
|
||
### Weaknesses
|
||
- Long streams needed for precision: 8-bit precision needs 2^8=256 clock cycles per number
|
||
- Correlated streams corrupt statistics (must use truly independent random sources)
|
||
- Slower than binary networks for same precision; only wins on hardware area
|
||
|
||
---
|
||
|
||
## Synthesis: What Hardware Knows That Software Has Forgotten
|
||
|
||
### 1. Locality is not a limitation — it's the design
|
||
Every hardware learning rule (STDP, Hebbian, LUT write, ΣΔ correction) updates
|
||
using only information at the site of the computation. No global loss function.
|
||
Software tries to compensate for locality (backprop broadcasts global gradient).
|
||
Hardware makes locality a feature: each synapse computes its own update.
|
||
|
||
**Implication for BNNBot:** the wave hit system IS the non-local signal.
|
||
Use it sparingly, like a neuromodulator: it modulates the sign/scale of local updates,
|
||
but the local activity traces must already exist at the synapse.
|
||
|
||
### 2. Eligibility traces solve the temporal credit assignment problem without backprop
|
||
The hardware solution to "which synapse caused the reward 10 ticks later":
|
||
every recently-active synapse maintains a decaying trace. When reward arrives,
|
||
it multiplies the trace. Traces decay in hardware via RC circuits.
|
||
|
||
**Implication:** for BNNBot, each in-flight wave needs to carry the eligibility trace
|
||
of which neurons fired when it was spawned. The wave system already stores the shot
|
||
context — add which neurons/weights were active.
|
||
|
||
### 3. The LUT IS the function approximator
|
||
A K-LUT can represent any K-variable Boolean function. There's nothing to train
|
||
in the gradient sense — you just write the truth table entry for each observed pattern.
|
||
This is the most efficient possible function approximator for discrete inputs.
|
||
|
||
**Implication:** WNN/WiSARD with reward-based table updates is the direct software
|
||
implementation of what hardware learning does. It's not an approximation of gradient
|
||
descent — it's a different algorithm entirely, and it's O(1) per update.
|
||
|
||
### 4. Energy minimization is free in recurrent hardware
|
||
Hopfield/Ising machines settle to energy minima by running physics.
|
||
Software has to simulate this expensively. But for our problem, the "energy landscape"
|
||
is implicit in our wave data — we just need the right representation.
|
||
|
||
**Implication:** the Hopfield associative memory is most useful as a pattern completion
|
||
engine for "I've seen this state context before" — i.e., case-based reasoning over
|
||
battle history.
|
||
|
||
### 5. Sigma-Delta: you only need the sign of accumulated error
|
||
The entire ΣΔ insight is that a 1-bit quantizer + integrator converges to the
|
||
correct value. We are throwing away information by keeping full-precision errors when
|
||
a signed accumulator does the same job.
|
||
|
||
**Implication:** the existing Hebbian residual table could be replaced by a ΣΔ
|
||
correction accumulator — simpler, same convergence, more noise-robust.
|
||
|
||
---
|
||
|
||
## Priority Ranking for BNNBot
|
||
|
||
| Mechanism | Fit | Cost | Priority |
|
||
|-----------|-----|------|----------|
|
||
| WNN/WiSARD (LUT-based) | High — directly maps to binary input + wave reward | Medium (memory) | **1st** |
|
||
| Three-factor / eligibility traces | High — solves delayed credit assignment | Low (add trace to wave struct) | **2nd** |
|
||
| ΣΔ correction accumulator | High — replaces/improves Hebbian residual table | Very low | **3rd** |
|
||
| PLL trajectory tracking | Medium — works for orbiting enemies | Low | **4th** |
|
||
| Hopfield associative memory | Medium — limited capacity, O(N²) weights | High | **5th** |
|
||
| Belief propagation | Low — requires manual factor graph design | Medium | **6th** |
|
||
| Stochastic computing | Low — representation mismatch with deterministic bits | Medium | **7th** |
|
||
| Evolvable hardware / GA | Very Low — too slow for online single-battle learning | Very High | skip |
|
||
|
||
---
|
||
|
||
## Most Actionable Finding
|
||
|
||
**WNN (Weightless Neural Network) with reward-counted tables** is the closest
|
||
direct translation of hardware learning to software. It:
|
||
- Requires zero multiplication (lookup only)
|
||
- Learns online from wave hits in one write per hit
|
||
- Naturally handles the 690-bit binary input
|
||
- Has proven capacity for pattern recognition (WiSARD commercial since 1984)
|
||
- The "neurons" are literally memory addresses — no parameters to tune
|
||
|
||
With K=16 bits per RAM node and 45 nodes: 45 × 64KB = 2.9MB RAM, sub-microsecond
|
||
inference. This is the most direct hardware-inspired solution that respects all
|
||
BNNBot constraints.
|