feat(ModularBot): 6 guns, pattern matcher, melee modules, adversarial bots
- New guns: guess-factor (GF histogram), pattern-matcher (movement tape replay) - New modules: minimum-risk melee movement, spinning melee radar - New test bots: PatternMover, RandomMover, WaveSurfer - Fixed: FeedbackEvent now carries actualX/actualY for proper GF learning - Fixed: TM gun warmup gating + directional residuals - Fixed: circular gun integrated formula + multi-bin omega cache - Fixed: oscillator wall-bounce lockout - Fixed: phantom meteor perpendicular body orientation - 6/6 battle wins across all enemy types
This commit is contained in:
@@ -0,0 +1,560 @@
|
||||
# Electronics & Hardware Domain Analysis
|
||||
## BNNBot Research — Learning Without Gradient Descent
|
||||
|
||||
**Context:** Binary network predicting enemy position in Robocode Tank Royale.
|
||||
Constraints: no gradient descent, online learning only, ~1ms/tick, binary-friendly,
|
||||
reward signal from wave-hit system (miss distance known per shot).
|
||||
|
||||
---
|
||||
|
||||
## 1. Memristive Systems
|
||||
|
||||
### How it works in hardware
|
||||
A memristor's resistance is a function of its charge history. Pass current one direction:
|
||||
resistance drops (LRS, low-resistance state = "1"). Reverse: resistance rises (HRS = "0").
|
||||
In a crossbar array, each crosspoint is one synapse. Reading = multiply (V × G = I),
|
||||
summing wires implement dot product in Ohm's law. Writing = apply voltage pulse.
|
||||
|
||||
Binary memristors (two-state, SET/RESET) switch between exactly LRS and HRS — no analog
|
||||
level needed. Programming is by voltage threshold: exceed it → switch state.
|
||||
|
||||
### Core algorithmic principle
|
||||
The natural learning rule is **STDP (Spike-Timing Dependent Plasticity)** — a Hebbian
|
||||
variant gated by spike timing. With three-factor extensions, a third neuromodulatory
|
||||
signal (reward / dopamine analog) gates the synaptic update:
|
||||
|
||||
```
|
||||
ΔW = eligibility_trace × reward_signal
|
||||
```
|
||||
|
||||
The eligibility trace is a leaky counter: it captures recent pre/post coincidence and
|
||||
decays. When reward arrives (delayed), it multiplies the trace. This lets delayed rewards
|
||||
credit the right synapses — the hardware equivalent of credit assignment without backprop.
|
||||
|
||||
For binary memristors specifically:
|
||||
- Hebbian: if pre AND post both fired recently → SET (LRS)
|
||||
- Anti-Hebbian: if pre fired but post did not (or vice versa) → RESET (HRS)
|
||||
- Three-factor: gate SET/RESET on reward sign
|
||||
|
||||
### Mapping to BNNBot
|
||||
- Our neuron is: `popcount(XOR(weights, input)) > threshold` — an n-input binary gate
|
||||
- Weights are bits; SET/RESET maps to w[i] = 1 / w[i] = 0
|
||||
- Eligibility trace = which weight bits were active when the shot was fired
|
||||
- Wave hit reward = delayed reward signal
|
||||
- Update: for active neurons on correct output → reinforce their weight bits that matched
|
||||
|
||||
### Concrete update rule (software sketch)
|
||||
```nim
|
||||
# On wave hit:
|
||||
# active_bits[i] = which input bits were 1 when this wave was fired
|
||||
# reward = 1.0 - (miss_distance / max_miss) # from wave system
|
||||
|
||||
for i in 0..<WEIGHT_BITS:
|
||||
if active_bits[i] and reward > 0.0:
|
||||
weights[i] = 1 # SET: pre and post active, positive reward
|
||||
elif active_bits[i] and reward < 0.0:
|
||||
weights[i] = 0 # RESET: active but wrong prediction
|
||||
# inactive bits: no change (Hebbian locality)
|
||||
```
|
||||
|
||||
### Strengths
|
||||
- Naturally online, one update per wave hit
|
||||
- Locality: each weight updates from local pre/post activity only
|
||||
- Delayed reward handled by eligibility trace — no need to hold full state
|
||||
|
||||
### Weaknesses
|
||||
- Binary weights = coarse; capacity scales with weight count not precision
|
||||
- No credit through layers — only the output neuron's weights get updated
|
||||
- Eligibility trace must be stored per weight per in-flight wave (memory cost)
|
||||
|
||||
---
|
||||
|
||||
## 2. FPGA / Evolvable Hardware
|
||||
|
||||
### How it works in hardware
|
||||
An FPGA is a fabric of LUTs (Look-Up Tables) connected by a programmable routing
|
||||
network. A K-LUT implements any K-input Boolean function by storing 2^K bits.
|
||||
Intrinsic evolution: the actual FPGA bitstream (which LUTs do what, how they connect)
|
||||
is treated as a genome. A genetic algorithm mutates it, measures fitness on real
|
||||
hardware (actual timing, actual noise), and selects survivors.
|
||||
|
||||
The key insight from Thompson (1996): intrinsic evolution finds solutions that exploit
|
||||
physical properties of the silicon that no designer would think to encode — clock
|
||||
leakage, capacitive coupling across supposedly-disconnected wires. The circuit
|
||||
"knows" things about its substrate that the designer doesn't model.
|
||||
|
||||
### Core algorithmic principle
|
||||
**Genetic search over discrete configuration space.** No gradient — just:
|
||||
1. Mutate bitstream (flip random bits)
|
||||
2. Measure fitness on hardware
|
||||
3. Keep better configurations, discard worse
|
||||
4. Repeat
|
||||
|
||||
For LUT-based learning (LUTNet, WNN/WiSARD): treat each LUT's truth table as
|
||||
writable memory. Training = writing new entries to LUTs. A K-input LUT can learn
|
||||
any K-variable Boolean function in a single write — no iterative optimization needed
|
||||
for patterns it has seen.
|
||||
|
||||
### Mapping to BNNBot
|
||||
Each "neuron" in our network is a boolean function of its binary inputs.
|
||||
An LUT IS the truth table for that function. Instead of storing a weight vector
|
||||
and doing popcount, we store the actual input→output mapping.
|
||||
|
||||
For n inputs: LUT has 2^n entries. But n=690 is way too large.
|
||||
Solution: each LUT sees only k<<690 bits (chosen by connectivity pattern).
|
||||
The 690-bit input is partitioned into chunks; each LUT learns its local chunk.
|
||||
|
||||
For online reward-based LUT update:
|
||||
- Record input index (the k-bit address into the LUT) when shot was fired
|
||||
- On wave hit: write output[address] = 1 if reward > threshold, else 0
|
||||
- This is exactly one-shot learning — the LUT "memorizes" each input pattern
|
||||
|
||||
### Concrete update rule
|
||||
```nim
|
||||
const K = 16 # LUT arity — covers 2^16 = 65536 patterns
|
||||
# lut: array[2^K, bool]
|
||||
|
||||
# On shot fired:
|
||||
let addr = extract_k_bits(input_690bit, chosen_indices, K)
|
||||
pending_updates.add((addr, wave_id))
|
||||
|
||||
# On wave hit:
|
||||
let (addr, _) = pending_updates[wave_id]
|
||||
lut[addr] = (miss_distance < HIT_THRESHOLD)
|
||||
```
|
||||
|
||||
### Strengths
|
||||
- One-shot learning: each (input pattern → outcome) pair stored directly
|
||||
- Zero arithmetic: inference is a table lookup, O(1), nanoseconds
|
||||
- LUT generalizes to unseen inputs via nearest neighbor in address space
|
||||
|
||||
### Weaknesses
|
||||
- 2^K memory per LUT — exponential in K. K=16 → 8KB per neuron.
|
||||
- With K=16 and 690 inputs, coverage per LUT is sparse (690/16 = ~43 non-overlapping LUTs)
|
||||
- Cold start: empty LUTs output 0; need warmup before predictions are useful
|
||||
- No generalization across different address patterns (each address independent)
|
||||
|
||||
---
|
||||
|
||||
## 3. Hopfield Networks — Energy Minimization
|
||||
|
||||
### How it works in hardware
|
||||
A Hopfield network is a fully connected binary recurrent network. Each node is +1/-1.
|
||||
Update rule: `s_i = sign(Σ_j W_ij × s_j)`. Weights are symmetric (W_ij = W_ji).
|
||||
The network has a Lyapunov energy function `E = -½ Σ W_ij s_i s_j`.
|
||||
Asynchronous updates monotonically decrease E → network converges to a local minimum.
|
||||
No gradient needed — local update decreases global energy by construction.
|
||||
|
||||
### Core algorithmic principle
|
||||
**Pattern as energy minimum.** Training stores patterns p by:
|
||||
`W_ij = (1/N) Σ_patterns p_i × p_j` (Hebbian outer product, no gradient).
|
||||
Retrieval: start from noisy/partial pattern, run updates → snaps to nearest stored pattern.
|
||||
The network "completes" corrupted inputs.
|
||||
|
||||
### Mapping to BNNBot
|
||||
Treat the prediction problem as pattern completion:
|
||||
- Input: partial state (last 10 radar scans = 690 bits)
|
||||
- Stored patterns: historical (input_context, correct_output) pairs
|
||||
- Query: clamp input bits, let output bits relax to minimum energy
|
||||
|
||||
This IS associative memory — we've seen this state-like context before, what happened?
|
||||
|
||||
Storing a (context → outcome) association:
|
||||
```nim
|
||||
# After wave hit with known outcome:
|
||||
let pattern = concat(input_690bit, outcome_69bit) # 759-bit pattern
|
||||
# Hebbian weight update (outer product, only for active bits):
|
||||
for i in active_bits(pattern):
|
||||
for j in active_bits(pattern):
|
||||
W[i][j] += 1.0 / N # symmetric
|
||||
```
|
||||
|
||||
Inference: clamp input 690 bits, iterate output 69 bits until convergence.
|
||||
|
||||
### Strengths
|
||||
- Convergence guaranteed (for symmetric W, no self-loops)
|
||||
- Naturally handles noisy/incomplete inputs — good for sensor noise
|
||||
- No gradient, no explicit credit assignment needed
|
||||
- OscNet v1.5 showed Hopfield runs forward-pass-only on hardware
|
||||
|
||||
### Weaknesses
|
||||
- Capacity: standard Hopfield stores ≈0.14N patterns for N nodes
|
||||
With N=759: only ~106 distinct (context, outcome) pairs — very limited
|
||||
- O(N²) weight matrix: 759² = 576K entries for float weights
|
||||
- Spurious minima: can settle to wrong pattern if too many stored
|
||||
- Modern dense Hopfield (transformer attention) has exponential capacity but requires softmax
|
||||
|
||||
---
|
||||
|
||||
## 4. Weightless Neural Networks (WNN / WiSARD)
|
||||
|
||||
### How it works in hardware
|
||||
A WNN neuron (RAM node) is literally an n-input lookup table with learned 1-bit entries.
|
||||
No multiplication, no addition — just memory address lookup. Training: present input,
|
||||
write 1 at that address. Inference: read address, return stored bit.
|
||||
|
||||
WiSARD (1981, commercial 1984): the original binary pattern recognizer.
|
||||
Each RAM node sees k randomly-chosen bits from the input. Multiple RAM nodes
|
||||
vote (sum of outputs) to produce a confidence score per class.
|
||||
|
||||
### Core algorithmic principle
|
||||
**Address-based memorization.** The n-input binary vector IS the memory address.
|
||||
Learning = `memory[address] = label`. No parameters to optimize, no loss surface.
|
||||
Generalization comes from the probabilistic overlap of binary addresses.
|
||||
|
||||
For online RL adaptation: instead of one-shot supervised write, count-based update:
|
||||
`memory[address] += reward_signal` then threshold. This is exactly a frequency table
|
||||
— "how often did this input pattern lead to a hit?"
|
||||
|
||||
### Mapping to BNNBot
|
||||
This is the most direct hardware→software translation:
|
||||
|
||||
```nim
|
||||
const K = 20 # bits per RAM node
|
||||
const N_NODES = 35 # 35 × 20 = 700 > 690
|
||||
# Each node: table[2^20] of int8, initialized to 0
|
||||
|
||||
# Index mapping: for each node, which 20 of 690 bits does it observe?
|
||||
# Fixed random permutation, same across all shots.
|
||||
|
||||
# On wave hit:
|
||||
for node in 0..<N_NODES:
|
||||
let addr = read_k_bits(input_690, node_mapping[node], K)
|
||||
let delta = if hit: +1 else: -1
|
||||
tables[node][addr] = clamp(tables[node][addr] + delta, -127, 127)
|
||||
|
||||
# Inference: which output direction gets highest vote?
|
||||
for each candidate_direction:
|
||||
score = 0
|
||||
for node in 0..<N_NODES:
|
||||
let addr = read_k_bits(input_690, node_mapping[node], K)
|
||||
score += tables[node][addr]
|
||||
if score > best_score: best_direction = candidate_direction
|
||||
```
|
||||
|
||||
### Strengths
|
||||
- Inference: pure memory lookup, O(N_NODES) — easily under 1ms
|
||||
- Online learning: one table write per wave hit
|
||||
- No cold-start paralysis — zero score = no preference, falls back to baseline
|
||||
- K=20 handles 2^20 = 1M patterns per node — generous capacity
|
||||
- Naturally compositional: each node learns a different k-bit "feature"
|
||||
|
||||
### Weaknesses
|
||||
- Memory: 35 nodes × 2^20 × 1 byte = 35 MB. Large but feasible.
|
||||
With K=16: 35 × 64KB = 2.2MB. More practical.
|
||||
- Random k-bit projections may not capture the most informative bit combinations
|
||||
(vs. learned projections in LUTNet-style approaches)
|
||||
- No cross-node generalization: two patterns differing in one bit observed by the same
|
||||
node = completely different addresses
|
||||
|
||||
**Note:** WNN with reward-based counting is the closest software analog to what
|
||||
memristor crossbars do in hardware — the table IS the synapse.
|
||||
|
||||
---
|
||||
|
||||
## 5. Belief Propagation on Factor Graphs (LDPC / Turbo codes)
|
||||
|
||||
### How it works in hardware
|
||||
LDPC decoding: received bits are noisy versions of codeword bits. The decoder has
|
||||
a factor graph: variable nodes (bits) and check nodes (parity constraints).
|
||||
Messages pass back and forth: each node sends to neighbors "given what I know, here's
|
||||
my probability estimate of your value." After ~50 iterations, converge to most likely codeword.
|
||||
|
||||
Hardware decoders run BP in a pipelined dataflow — no global state, purely local message updates.
|
||||
|
||||
### Core algorithmic principle
|
||||
**Iterative marginal inference.** Each variable node holds a log-likelihood ratio (LLR).
|
||||
Update rule for variable node x_i receiving messages m_{j→i} from check nodes j:
|
||||
```
|
||||
LLR_i = channel_LLR_i + Σ_j m_{j→i}
|
||||
m_{i→k} = LLR_i - m_{k→i} # exclude message from k before sending back
|
||||
```
|
||||
No gradient. Convergence is guaranteed on tree-structured graphs; approximate on loopy graphs.
|
||||
|
||||
### Mapping to BNNBot
|
||||
The 690-bit input has structure — bits are correlated (sin/cos pairs, position from velocity,
|
||||
etc.). BP could exploit this structure for inference without learning.
|
||||
|
||||
Concrete use: model the 69-bit output (next position frame) as hidden variables.
|
||||
Observed: 690-bit input. Factor graph: local constraints from physics (velocity ×
|
||||
ticks = displacement, heading consistency). Run BP to infer most-likely output.
|
||||
|
||||
This is not a learning mechanism — it's inference. But it applies to our problem:
|
||||
the physics of enemy movement IS the factor graph. No training needed if the
|
||||
constraints are correct.
|
||||
|
||||
### Concrete sketch
|
||||
```nim
|
||||
# Factor: expected_x[t+1] = x[t] + vx[t] * speed
|
||||
# Factor: heading changes slowly (soft constraint, variance from data)
|
||||
# Observed: x[t], y[t], vx[t], vy[t] from radar scans
|
||||
|
||||
# For each output bit o_i:
|
||||
# LLR[i] = initial estimate from linear extrapolation
|
||||
# Messages from physics factors update LLR iteratively
|
||||
# Final: sign(LLR[i]) = predicted bit
|
||||
```
|
||||
|
||||
### Strengths
|
||||
- Exploits problem structure (physics) explicitly — not black-box learning
|
||||
- Iterative refinement: run more iterations when time allows
|
||||
- Error-correcting codes achieve Shannon capacity with BP — proven near-optimal
|
||||
|
||||
### Weaknesses
|
||||
- Requires knowing the factor structure (constraint graph) — domain engineering needed
|
||||
- Our enemy's movement IS physics, but also strategy (evasion) — hard to model as factors
|
||||
- Less useful as a "learning" mechanism; more useful as a structured inference method
|
||||
- Loopy BP on dense graphs may not converge
|
||||
|
||||
---
|
||||
|
||||
## 6. Sigma-Delta Modulation — Feedback Tracking
|
||||
|
||||
### How it works in hardware
|
||||
A sigma-delta (ΣΔ) modulator converts analog input to a binary bit stream.
|
||||
Architecture: integrator → 1-bit comparator → feedback DAC.
|
||||
The comparator outputs 1 if accumulator > 0, else 0. The 1-bit output is fed back
|
||||
and subtracted from the input. The accumulator tracks the error (sigma = cumulative,
|
||||
delta = difference). Output bit stream: density encodes analog value.
|
||||
|
||||
### Core algorithmic principle
|
||||
**Error-accumulating feedback.** The system never "knows" the true analog value —
|
||||
it only knows the sign of accumulated error. It corrects continuously:
|
||||
```
|
||||
accumulator += (input - feedback)
|
||||
output_bit = (accumulator > 0) ? 1 : 0
|
||||
feedback = output_bit ? +1 : -1
|
||||
```
|
||||
This is a 1-bit quantizer with infinite-precision error memory. The feedback loop
|
||||
drives the accumulated error to zero over time — the bit stream density equals the
|
||||
input value.
|
||||
|
||||
### Mapping to BNNBot
|
||||
The analogy: our prediction is a "bit stream" (sequence of predictions).
|
||||
The wave hit system gives us the sign of the error (too far left? too far right?).
|
||||
We accumulate signed errors and correct the prediction offset continuously.
|
||||
|
||||
This is closer to what our existing Hebbian residual table does — but the ΣΔ
|
||||
formulation is cleaner:
|
||||
|
||||
```nim
|
||||
# Per directional axis (x and y):
|
||||
var accumulator_x: float = 0.0
|
||||
var accumulator_y: float = 0.0
|
||||
|
||||
# On wave hit:
|
||||
let error_x = actual_x - predicted_x
|
||||
let error_y = actual_y - predicted_y
|
||||
|
||||
accumulator_x += error_x
|
||||
accumulator_y += error_y
|
||||
|
||||
# Correction for next prediction:
|
||||
correction_x = sign(accumulator_x) * correction_step # 1-bit correction
|
||||
# Or: correction_x = clamp(accumulator_x * gain, -max_corr, max_corr) # proportional
|
||||
```
|
||||
|
||||
The ΣΔ insight: **you don't need the full error magnitude, just its sign.**
|
||||
The accumulation of signs over time encodes the magnitude.
|
||||
|
||||
### Strengths
|
||||
- Minimal state: just one accumulator per axis
|
||||
- Noise-shaping: high-frequency jitter averages out; low-frequency bias accumulates correctly
|
||||
- Robust to measurement noise (sign of error is reliable even when magnitude is noisy)
|
||||
|
||||
### Weaknesses
|
||||
- Slow convergence if correction step is too small
|
||||
- Tracks only a DC offset (systematic bias), not complex patterns
|
||||
- Best as a correction layer on top of a better predictor, not standalone
|
||||
|
||||
---
|
||||
|
||||
## 7. Phase-Locked Loop (PLL) — Trajectory Tracking
|
||||
|
||||
### How it works in hardware
|
||||
A PLL: phase detector compares incoming signal phase to VCO output phase.
|
||||
Error signal (phase difference) feeds through a loop filter to the VCO control input.
|
||||
VCO adjusts frequency until phase error → 0. Locked: VCO tracks input frequency and phase.
|
||||
|
||||
Components: Phase Detector (XOR or mixer), Loop Filter (low-pass), VCO.
|
||||
No gradient — just: am I ahead or behind? Correct proportionally.
|
||||
|
||||
### Core algorithmic principle
|
||||
**Proportional-integral control on phase error.** PI loop filter:
|
||||
```
|
||||
phase_error = input_phase - vco_phase
|
||||
control_voltage += Kp * phase_error + Ki * integral(phase_error)
|
||||
vco_frequency = f0 + K_vco * control_voltage
|
||||
```
|
||||
This IS a gradient-free optimizer for tracking. It converges when phase_error = 0.
|
||||
Equivalent to online least-mean-squares for frequency estimation.
|
||||
|
||||
### Mapping to BNNBot
|
||||
Enemy position is a trajectory in 2D. Our aiming angle is a phase relative to that trajectory.
|
||||
A PLL-like system: track the rate of change of enemy heading (angular velocity).
|
||||
|
||||
```nim
|
||||
# "Phase" = enemy heading direction
|
||||
# "VCO" = our predicted heading trend
|
||||
# "Lock" = our trend matches enemy trend
|
||||
|
||||
var predicted_angular_velocity: float = 0.0
|
||||
var phase_error_integral: float = 0.0
|
||||
const Kp = 0.3
|
||||
const Ki = 0.1
|
||||
|
||||
# On each radar scan:
|
||||
let actual_angular_velocity = delta_heading / delta_time
|
||||
let phase_error = actual_angular_velocity - predicted_angular_velocity
|
||||
|
||||
phase_error_integral += phase_error
|
||||
predicted_angular_velocity += Kp * phase_error + Ki * phase_error_integral
|
||||
```
|
||||
|
||||
When the enemy moves at constant angular velocity (orbit, spiral), this "locks on"
|
||||
and predicts ahead by extrapolating the locked phase.
|
||||
|
||||
### Strengths
|
||||
- Handles periodic/oscillatory motion natively (enemy orbiting = pure sine wave)
|
||||
- Acquisition (from cold) + tracking (once locked) are separate phases — known behavior
|
||||
- Loop filter order can be tuned: 1st order = tracks constant heading, 2nd = tracks constant angular acceleration
|
||||
|
||||
### Weaknesses
|
||||
- Assumes quasi-periodic or smooth trajectory; breaks on jerky evasion
|
||||
- Loop bandwidth tradeoff: wide bandwidth tracks fast changes but amplifies noise
|
||||
- Enemy headings in Robocode are piecewise linear, not sinusoidal — PLL is mismatched model
|
||||
unless enemy orbits, which some do
|
||||
|
||||
---
|
||||
|
||||
## 8. Stochastic Computing
|
||||
|
||||
### How it works in hardware
|
||||
Represent numbers as the probability that a bit in a random stream is 1.
|
||||
x = 0.7 → 70% of bits are 1 in a random stream. Then:
|
||||
- Multiplication: AND(stream_x, stream_y) → probability = x × y
|
||||
- Addition (scaled): MUX(stream_x, stream_y, select) → (x + y)/2
|
||||
- Dot product: AND all streams, OR results
|
||||
|
||||
All operations are single-gate logic. No carry chains, no multipliers.
|
||||
Inference hardware is tiny. Accuracy scales with stream length (more bits = more precise).
|
||||
|
||||
### Core algorithmic principle
|
||||
**Probability represented as bit density.** The number IS the stream.
|
||||
Computation IS gate operations on streams.
|
||||
|
||||
Weight update in stochastic computing:
|
||||
```
|
||||
Δw_ij = AND(activation_i_stream, error_j_stream)
|
||||
```
|
||||
This computes the product of two probabilities using an AND gate.
|
||||
The learning rule IS Hebbian — coincidence detection between pre and post streams.
|
||||
|
||||
### Mapping to BNNBot
|
||||
Our 690-bit input is already binary — but not stochastic (each bit is deterministic).
|
||||
However, stochastic computing suggests a different representation:
|
||||
instead of Gray-coded positions, represent uncertainty as bit-stream density.
|
||||
|
||||
Alternatively: treat each weight as a stochastic bit (1 with probability p).
|
||||
Inference = sample weights → compute output → measure hit → update p.
|
||||
This is essentially WNN with probabilistic weights.
|
||||
|
||||
Concrete weight update:
|
||||
```nim
|
||||
# Each weight w[i] is a probability stored as float in [0,1]
|
||||
# On shot fired: sample binary weights: sampled[i] = rand() < w[i]
|
||||
# On wave hit:
|
||||
for i in active_neurons:
|
||||
w[i] += learning_rate * reward * (sampled[i] - w[i])
|
||||
# REINFORCE-like: push probability toward 1 if it fired and reward was positive
|
||||
```
|
||||
|
||||
### Strengths
|
||||
- Noise is a feature, not a bug — exploration built in
|
||||
- Multiplication = AND, the cheapest gate — scales to large networks
|
||||
- Graceful degradation: shorter streams = noisier but not broken
|
||||
|
||||
### Weaknesses
|
||||
- Long streams needed for precision: 8-bit precision needs 2^8=256 clock cycles per number
|
||||
- Correlated streams corrupt statistics (must use truly independent random sources)
|
||||
- Slower than binary networks for same precision; only wins on hardware area
|
||||
|
||||
---
|
||||
|
||||
## Synthesis: What Hardware Knows That Software Has Forgotten
|
||||
|
||||
### 1. Locality is not a limitation — it's the design
|
||||
Every hardware learning rule (STDP, Hebbian, LUT write, ΣΔ correction) updates
|
||||
using only information at the site of the computation. No global loss function.
|
||||
Software tries to compensate for locality (backprop broadcasts global gradient).
|
||||
Hardware makes locality a feature: each synapse computes its own update.
|
||||
|
||||
**Implication for BNNBot:** the wave hit system IS the non-local signal.
|
||||
Use it sparingly, like a neuromodulator: it modulates the sign/scale of local updates,
|
||||
but the local activity traces must already exist at the synapse.
|
||||
|
||||
### 2. Eligibility traces solve the temporal credit assignment problem without backprop
|
||||
The hardware solution to "which synapse caused the reward 10 ticks later":
|
||||
every recently-active synapse maintains a decaying trace. When reward arrives,
|
||||
it multiplies the trace. Traces decay in hardware via RC circuits.
|
||||
|
||||
**Implication:** for BNNBot, each in-flight wave needs to carry the eligibility trace
|
||||
of which neurons fired when it was spawned. The wave system already stores the shot
|
||||
context — add which neurons/weights were active.
|
||||
|
||||
### 3. The LUT IS the function approximator
|
||||
A K-LUT can represent any K-variable Boolean function. There's nothing to train
|
||||
in the gradient sense — you just write the truth table entry for each observed pattern.
|
||||
This is the most efficient possible function approximator for discrete inputs.
|
||||
|
||||
**Implication:** WNN/WiSARD with reward-based table updates is the direct software
|
||||
implementation of what hardware learning does. It's not an approximation of gradient
|
||||
descent — it's a different algorithm entirely, and it's O(1) per update.
|
||||
|
||||
### 4. Energy minimization is free in recurrent hardware
|
||||
Hopfield/Ising machines settle to energy minima by running physics.
|
||||
Software has to simulate this expensively. But for our problem, the "energy landscape"
|
||||
is implicit in our wave data — we just need the right representation.
|
||||
|
||||
**Implication:** the Hopfield associative memory is most useful as a pattern completion
|
||||
engine for "I've seen this state context before" — i.e., case-based reasoning over
|
||||
battle history.
|
||||
|
||||
### 5. Sigma-Delta: you only need the sign of accumulated error
|
||||
The entire ΣΔ insight is that a 1-bit quantizer + integrator converges to the
|
||||
correct value. We are throwing away information by keeping full-precision errors when
|
||||
a signed accumulator does the same job.
|
||||
|
||||
**Implication:** the existing Hebbian residual table could be replaced by a ΣΔ
|
||||
correction accumulator — simpler, same convergence, more noise-robust.
|
||||
|
||||
---
|
||||
|
||||
## Priority Ranking for BNNBot
|
||||
|
||||
| Mechanism | Fit | Cost | Priority |
|
||||
|-----------|-----|------|----------|
|
||||
| WNN/WiSARD (LUT-based) | High — directly maps to binary input + wave reward | Medium (memory) | **1st** |
|
||||
| Three-factor / eligibility traces | High — solves delayed credit assignment | Low (add trace to wave struct) | **2nd** |
|
||||
| ΣΔ correction accumulator | High — replaces/improves Hebbian residual table | Very low | **3rd** |
|
||||
| PLL trajectory tracking | Medium — works for orbiting enemies | Low | **4th** |
|
||||
| Hopfield associative memory | Medium — limited capacity, O(N²) weights | High | **5th** |
|
||||
| Belief propagation | Low — requires manual factor graph design | Medium | **6th** |
|
||||
| Stochastic computing | Low — representation mismatch with deterministic bits | Medium | **7th** |
|
||||
| Evolvable hardware / GA | Very Low — too slow for online single-battle learning | Very High | skip |
|
||||
|
||||
---
|
||||
|
||||
## Most Actionable Finding
|
||||
|
||||
**WNN (Weightless Neural Network) with reward-counted tables** is the closest
|
||||
direct translation of hardware learning to software. It:
|
||||
- Requires zero multiplication (lookup only)
|
||||
- Learns online from wave hits in one write per hit
|
||||
- Naturally handles the 690-bit binary input
|
||||
- Has proven capacity for pattern recognition (WiSARD commercial since 1984)
|
||||
- The "neurons" are literally memory addresses — no parameters to tune
|
||||
|
||||
With K=16 bits per RAM node and 45 nodes: 45 × 64KB = 2.9MB RAM, sub-microsecond
|
||||
inference. This is the most direct hardware-inspired solution that respects all
|
||||
BNNBot constraints.
|
||||
Reference in New Issue
Block a user