Files
SirRoboGarage/BNNBot_garage/analysis/domain_electronics.md
T
SirStone 1ed7797cb6 feat(ModularBot): 6 guns, pattern matcher, melee modules, adversarial bots
- New guns: guess-factor (GF histogram), pattern-matcher (movement tape replay)
- New modules: minimum-risk melee movement, spinning melee radar
- New test bots: PatternMover, RandomMover, WaveSurfer
- Fixed: FeedbackEvent now carries actualX/actualY for proper GF learning
- Fixed: TM gun warmup gating + directional residuals
- Fixed: circular gun integrated formula + multi-bin omega cache
- Fixed: oscillator wall-bounce lockout
- Fixed: phantom meteor perpendicular body orientation
- 6/6 battle wins across all enemy types
2026-09-20 00:59:53 +02:00

561 lines
24 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Electronics & Hardware Domain Analysis
## BNNBot Research — Learning Without Gradient Descent
**Context:** Binary network predicting enemy position in Robocode Tank Royale.
Constraints: no gradient descent, online learning only, ~1ms/tick, binary-friendly,
reward signal from wave-hit system (miss distance known per shot).
---
## 1. Memristive Systems
### How it works in hardware
A memristor's resistance is a function of its charge history. Pass current one direction:
resistance drops (LRS, low-resistance state = "1"). Reverse: resistance rises (HRS = "0").
In a crossbar array, each crosspoint is one synapse. Reading = multiply (V × G = I),
summing wires implement dot product in Ohm's law. Writing = apply voltage pulse.
Binary memristors (two-state, SET/RESET) switch between exactly LRS and HRS — no analog
level needed. Programming is by voltage threshold: exceed it → switch state.
### Core algorithmic principle
The natural learning rule is **STDP (Spike-Timing Dependent Plasticity)** — a Hebbian
variant gated by spike timing. With three-factor extensions, a third neuromodulatory
signal (reward / dopamine analog) gates the synaptic update:
```
ΔW = eligibility_trace × reward_signal
```
The eligibility trace is a leaky counter: it captures recent pre/post coincidence and
decays. When reward arrives (delayed), it multiplies the trace. This lets delayed rewards
credit the right synapses — the hardware equivalent of credit assignment without backprop.
For binary memristors specifically:
- Hebbian: if pre AND post both fired recently → SET (LRS)
- Anti-Hebbian: if pre fired but post did not (or vice versa) → RESET (HRS)
- Three-factor: gate SET/RESET on reward sign
### Mapping to BNNBot
- Our neuron is: `popcount(XOR(weights, input)) > threshold` — an n-input binary gate
- Weights are bits; SET/RESET maps to w[i] = 1 / w[i] = 0
- Eligibility trace = which weight bits were active when the shot was fired
- Wave hit reward = delayed reward signal
- Update: for active neurons on correct output → reinforce their weight bits that matched
### Concrete update rule (software sketch)
```nim
# On wave hit:
# active_bits[i] = which input bits were 1 when this wave was fired
# reward = 1.0 - (miss_distance / max_miss) # from wave system
for i in 0..<WEIGHT_BITS:
if active_bits[i] and reward > 0.0:
weights[i] = 1 # SET: pre and post active, positive reward
elif active_bits[i] and reward < 0.0:
weights[i] = 0 # RESET: active but wrong prediction
# inactive bits: no change (Hebbian locality)
```
### Strengths
- Naturally online, one update per wave hit
- Locality: each weight updates from local pre/post activity only
- Delayed reward handled by eligibility trace — no need to hold full state
### Weaknesses
- Binary weights = coarse; capacity scales with weight count not precision
- No credit through layers — only the output neuron's weights get updated
- Eligibility trace must be stored per weight per in-flight wave (memory cost)
---
## 2. FPGA / Evolvable Hardware
### How it works in hardware
An FPGA is a fabric of LUTs (Look-Up Tables) connected by a programmable routing
network. A K-LUT implements any K-input Boolean function by storing 2^K bits.
Intrinsic evolution: the actual FPGA bitstream (which LUTs do what, how they connect)
is treated as a genome. A genetic algorithm mutates it, measures fitness on real
hardware (actual timing, actual noise), and selects survivors.
The key insight from Thompson (1996): intrinsic evolution finds solutions that exploit
physical properties of the silicon that no designer would think to encode — clock
leakage, capacitive coupling across supposedly-disconnected wires. The circuit
"knows" things about its substrate that the designer doesn't model.
### Core algorithmic principle
**Genetic search over discrete configuration space.** No gradient — just:
1. Mutate bitstream (flip random bits)
2. Measure fitness on hardware
3. Keep better configurations, discard worse
4. Repeat
For LUT-based learning (LUTNet, WNN/WiSARD): treat each LUT's truth table as
writable memory. Training = writing new entries to LUTs. A K-input LUT can learn
any K-variable Boolean function in a single write — no iterative optimization needed
for patterns it has seen.
### Mapping to BNNBot
Each "neuron" in our network is a boolean function of its binary inputs.
An LUT IS the truth table for that function. Instead of storing a weight vector
and doing popcount, we store the actual input→output mapping.
For n inputs: LUT has 2^n entries. But n=690 is way too large.
Solution: each LUT sees only k<<690 bits (chosen by connectivity pattern).
The 690-bit input is partitioned into chunks; each LUT learns its local chunk.
For online reward-based LUT update:
- Record input index (the k-bit address into the LUT) when shot was fired
- On wave hit: write output[address] = 1 if reward > threshold, else 0
- This is exactly one-shot learning — the LUT "memorizes" each input pattern
### Concrete update rule
```nim
const K = 16 # LUT arity — covers 2^16 = 65536 patterns
# lut: array[2^K, bool]
# On shot fired:
let addr = extract_k_bits(input_690bit, chosen_indices, K)
pending_updates.add((addr, wave_id))
# On wave hit:
let (addr, _) = pending_updates[wave_id]
lut[addr] = (miss_distance < HIT_THRESHOLD)
```
### Strengths
- One-shot learning: each (input pattern → outcome) pair stored directly
- Zero arithmetic: inference is a table lookup, O(1), nanoseconds
- LUT generalizes to unseen inputs via nearest neighbor in address space
### Weaknesses
- 2^K memory per LUT — exponential in K. K=16 → 8KB per neuron.
- With K=16 and 690 inputs, coverage per LUT is sparse (690/16 = ~43 non-overlapping LUTs)
- Cold start: empty LUTs output 0; need warmup before predictions are useful
- No generalization across different address patterns (each address independent)
---
## 3. Hopfield Networks — Energy Minimization
### How it works in hardware
A Hopfield network is a fully connected binary recurrent network. Each node is +1/-1.
Update rule: `s_i = sign(Σ_j W_ij × s_j)`. Weights are symmetric (W_ij = W_ji).
The network has a Lyapunov energy function `E = -½ Σ W_ij s_i s_j`.
Asynchronous updates monotonically decrease E → network converges to a local minimum.
No gradient needed — local update decreases global energy by construction.
### Core algorithmic principle
**Pattern as energy minimum.** Training stores patterns p by:
`W_ij = (1/N) Σ_patterns p_i × p_j` (Hebbian outer product, no gradient).
Retrieval: start from noisy/partial pattern, run updates → snaps to nearest stored pattern.
The network "completes" corrupted inputs.
### Mapping to BNNBot
Treat the prediction problem as pattern completion:
- Input: partial state (last 10 radar scans = 690 bits)
- Stored patterns: historical (input_context, correct_output) pairs
- Query: clamp input bits, let output bits relax to minimum energy
This IS associative memory — we've seen this state-like context before, what happened?
Storing a (context → outcome) association:
```nim
# After wave hit with known outcome:
let pattern = concat(input_690bit, outcome_69bit) # 759-bit pattern
# Hebbian weight update (outer product, only for active bits):
for i in active_bits(pattern):
for j in active_bits(pattern):
W[i][j] += 1.0 / N # symmetric
```
Inference: clamp input 690 bits, iterate output 69 bits until convergence.
### Strengths
- Convergence guaranteed (for symmetric W, no self-loops)
- Naturally handles noisy/incomplete inputs — good for sensor noise
- No gradient, no explicit credit assignment needed
- OscNet v1.5 showed Hopfield runs forward-pass-only on hardware
### Weaknesses
- Capacity: standard Hopfield stores ≈0.14N patterns for N nodes
With N=759: only ~106 distinct (context, outcome) pairs — very limited
- O(N²) weight matrix: 759² = 576K entries for float weights
- Spurious minima: can settle to wrong pattern if too many stored
- Modern dense Hopfield (transformer attention) has exponential capacity but requires softmax
---
## 4. Weightless Neural Networks (WNN / WiSARD)
### How it works in hardware
A WNN neuron (RAM node) is literally an n-input lookup table with learned 1-bit entries.
No multiplication, no addition — just memory address lookup. Training: present input,
write 1 at that address. Inference: read address, return stored bit.
WiSARD (1981, commercial 1984): the original binary pattern recognizer.
Each RAM node sees k randomly-chosen bits from the input. Multiple RAM nodes
vote (sum of outputs) to produce a confidence score per class.
### Core algorithmic principle
**Address-based memorization.** The n-input binary vector IS the memory address.
Learning = `memory[address] = label`. No parameters to optimize, no loss surface.
Generalization comes from the probabilistic overlap of binary addresses.
For online RL adaptation: instead of one-shot supervised write, count-based update:
`memory[address] += reward_signal` then threshold. This is exactly a frequency table
— "how often did this input pattern lead to a hit?"
### Mapping to BNNBot
This is the most direct hardware→software translation:
```nim
const K = 20 # bits per RAM node
const N_NODES = 35 # 35 × 20 = 700 > 690
# Each node: table[2^20] of int8, initialized to 0
# Index mapping: for each node, which 20 of 690 bits does it observe?
# Fixed random permutation, same across all shots.
# On wave hit:
for node in 0..<N_NODES:
let addr = read_k_bits(input_690, node_mapping[node], K)
let delta = if hit: +1 else: -1
tables[node][addr] = clamp(tables[node][addr] + delta, -127, 127)
# Inference: which output direction gets highest vote?
for each candidate_direction:
score = 0
for node in 0..<N_NODES:
let addr = read_k_bits(input_690, node_mapping[node], K)
score += tables[node][addr]
if score > best_score: best_direction = candidate_direction
```
### Strengths
- Inference: pure memory lookup, O(N_NODES) — easily under 1ms
- Online learning: one table write per wave hit
- No cold-start paralysis — zero score = no preference, falls back to baseline
- K=20 handles 2^20 = 1M patterns per node — generous capacity
- Naturally compositional: each node learns a different k-bit "feature"
### Weaknesses
- Memory: 35 nodes × 2^20 × 1 byte = 35 MB. Large but feasible.
With K=16: 35 × 64KB = 2.2MB. More practical.
- Random k-bit projections may not capture the most informative bit combinations
(vs. learned projections in LUTNet-style approaches)
- No cross-node generalization: two patterns differing in one bit observed by the same
node = completely different addresses
**Note:** WNN with reward-based counting is the closest software analog to what
memristor crossbars do in hardware — the table IS the synapse.
---
## 5. Belief Propagation on Factor Graphs (LDPC / Turbo codes)
### How it works in hardware
LDPC decoding: received bits are noisy versions of codeword bits. The decoder has
a factor graph: variable nodes (bits) and check nodes (parity constraints).
Messages pass back and forth: each node sends to neighbors "given what I know, here's
my probability estimate of your value." After ~50 iterations, converge to most likely codeword.
Hardware decoders run BP in a pipelined dataflow — no global state, purely local message updates.
### Core algorithmic principle
**Iterative marginal inference.** Each variable node holds a log-likelihood ratio (LLR).
Update rule for variable node x_i receiving messages m_{j→i} from check nodes j:
```
LLR_i = channel_LLR_i + Σ_j m_{j→i}
m_{i→k} = LLR_i - m_{k→i} # exclude message from k before sending back
```
No gradient. Convergence is guaranteed on tree-structured graphs; approximate on loopy graphs.
### Mapping to BNNBot
The 690-bit input has structure — bits are correlated (sin/cos pairs, position from velocity,
etc.). BP could exploit this structure for inference without learning.
Concrete use: model the 69-bit output (next position frame) as hidden variables.
Observed: 690-bit input. Factor graph: local constraints from physics (velocity ×
ticks = displacement, heading consistency). Run BP to infer most-likely output.
This is not a learning mechanism — it's inference. But it applies to our problem:
the physics of enemy movement IS the factor graph. No training needed if the
constraints are correct.
### Concrete sketch
```nim
# Factor: expected_x[t+1] = x[t] + vx[t] * speed
# Factor: heading changes slowly (soft constraint, variance from data)
# Observed: x[t], y[t], vx[t], vy[t] from radar scans
# For each output bit o_i:
# LLR[i] = initial estimate from linear extrapolation
# Messages from physics factors update LLR iteratively
# Final: sign(LLR[i]) = predicted bit
```
### Strengths
- Exploits problem structure (physics) explicitly — not black-box learning
- Iterative refinement: run more iterations when time allows
- Error-correcting codes achieve Shannon capacity with BP — proven near-optimal
### Weaknesses
- Requires knowing the factor structure (constraint graph) — domain engineering needed
- Our enemy's movement IS physics, but also strategy (evasion) — hard to model as factors
- Less useful as a "learning" mechanism; more useful as a structured inference method
- Loopy BP on dense graphs may not converge
---
## 6. Sigma-Delta Modulation — Feedback Tracking
### How it works in hardware
A sigma-delta (ΣΔ) modulator converts analog input to a binary bit stream.
Architecture: integrator → 1-bit comparator → feedback DAC.
The comparator outputs 1 if accumulator > 0, else 0. The 1-bit output is fed back
and subtracted from the input. The accumulator tracks the error (sigma = cumulative,
delta = difference). Output bit stream: density encodes analog value.
### Core algorithmic principle
**Error-accumulating feedback.** The system never "knows" the true analog value —
it only knows the sign of accumulated error. It corrects continuously:
```
accumulator += (input - feedback)
output_bit = (accumulator > 0) ? 1 : 0
feedback = output_bit ? +1 : -1
```
This is a 1-bit quantizer with infinite-precision error memory. The feedback loop
drives the accumulated error to zero over time — the bit stream density equals the
input value.
### Mapping to BNNBot
The analogy: our prediction is a "bit stream" (sequence of predictions).
The wave hit system gives us the sign of the error (too far left? too far right?).
We accumulate signed errors and correct the prediction offset continuously.
This is closer to what our existing Hebbian residual table does — but the ΣΔ
formulation is cleaner:
```nim
# Per directional axis (x and y):
var accumulator_x: float = 0.0
var accumulator_y: float = 0.0
# On wave hit:
let error_x = actual_x - predicted_x
let error_y = actual_y - predicted_y
accumulator_x += error_x
accumulator_y += error_y
# Correction for next prediction:
correction_x = sign(accumulator_x) * correction_step # 1-bit correction
# Or: correction_x = clamp(accumulator_x * gain, -max_corr, max_corr) # proportional
```
The ΣΔ insight: **you don't need the full error magnitude, just its sign.**
The accumulation of signs over time encodes the magnitude.
### Strengths
- Minimal state: just one accumulator per axis
- Noise-shaping: high-frequency jitter averages out; low-frequency bias accumulates correctly
- Robust to measurement noise (sign of error is reliable even when magnitude is noisy)
### Weaknesses
- Slow convergence if correction step is too small
- Tracks only a DC offset (systematic bias), not complex patterns
- Best as a correction layer on top of a better predictor, not standalone
---
## 7. Phase-Locked Loop (PLL) — Trajectory Tracking
### How it works in hardware
A PLL: phase detector compares incoming signal phase to VCO output phase.
Error signal (phase difference) feeds through a loop filter to the VCO control input.
VCO adjusts frequency until phase error → 0. Locked: VCO tracks input frequency and phase.
Components: Phase Detector (XOR or mixer), Loop Filter (low-pass), VCO.
No gradient — just: am I ahead or behind? Correct proportionally.
### Core algorithmic principle
**Proportional-integral control on phase error.** PI loop filter:
```
phase_error = input_phase - vco_phase
control_voltage += Kp * phase_error + Ki * integral(phase_error)
vco_frequency = f0 + K_vco * control_voltage
```
This IS a gradient-free optimizer for tracking. It converges when phase_error = 0.
Equivalent to online least-mean-squares for frequency estimation.
### Mapping to BNNBot
Enemy position is a trajectory in 2D. Our aiming angle is a phase relative to that trajectory.
A PLL-like system: track the rate of change of enemy heading (angular velocity).
```nim
# "Phase" = enemy heading direction
# "VCO" = our predicted heading trend
# "Lock" = our trend matches enemy trend
var predicted_angular_velocity: float = 0.0
var phase_error_integral: float = 0.0
const Kp = 0.3
const Ki = 0.1
# On each radar scan:
let actual_angular_velocity = delta_heading / delta_time
let phase_error = actual_angular_velocity - predicted_angular_velocity
phase_error_integral += phase_error
predicted_angular_velocity += Kp * phase_error + Ki * phase_error_integral
```
When the enemy moves at constant angular velocity (orbit, spiral), this "locks on"
and predicts ahead by extrapolating the locked phase.
### Strengths
- Handles periodic/oscillatory motion natively (enemy orbiting = pure sine wave)
- Acquisition (from cold) + tracking (once locked) are separate phases — known behavior
- Loop filter order can be tuned: 1st order = tracks constant heading, 2nd = tracks constant angular acceleration
### Weaknesses
- Assumes quasi-periodic or smooth trajectory; breaks on jerky evasion
- Loop bandwidth tradeoff: wide bandwidth tracks fast changes but amplifies noise
- Enemy headings in Robocode are piecewise linear, not sinusoidal — PLL is mismatched model
unless enemy orbits, which some do
---
## 8. Stochastic Computing
### How it works in hardware
Represent numbers as the probability that a bit in a random stream is 1.
x = 0.7 → 70% of bits are 1 in a random stream. Then:
- Multiplication: AND(stream_x, stream_y) → probability = x × y
- Addition (scaled): MUX(stream_x, stream_y, select) → (x + y)/2
- Dot product: AND all streams, OR results
All operations are single-gate logic. No carry chains, no multipliers.
Inference hardware is tiny. Accuracy scales with stream length (more bits = more precise).
### Core algorithmic principle
**Probability represented as bit density.** The number IS the stream.
Computation IS gate operations on streams.
Weight update in stochastic computing:
```
Δw_ij = AND(activation_i_stream, error_j_stream)
```
This computes the product of two probabilities using an AND gate.
The learning rule IS Hebbian — coincidence detection between pre and post streams.
### Mapping to BNNBot
Our 690-bit input is already binary — but not stochastic (each bit is deterministic).
However, stochastic computing suggests a different representation:
instead of Gray-coded positions, represent uncertainty as bit-stream density.
Alternatively: treat each weight as a stochastic bit (1 with probability p).
Inference = sample weights → compute output → measure hit → update p.
This is essentially WNN with probabilistic weights.
Concrete weight update:
```nim
# Each weight w[i] is a probability stored as float in [0,1]
# On shot fired: sample binary weights: sampled[i] = rand() < w[i]
# On wave hit:
for i in active_neurons:
w[i] += learning_rate * reward * (sampled[i] - w[i])
# REINFORCE-like: push probability toward 1 if it fired and reward was positive
```
### Strengths
- Noise is a feature, not a bug — exploration built in
- Multiplication = AND, the cheapest gate — scales to large networks
- Graceful degradation: shorter streams = noisier but not broken
### Weaknesses
- Long streams needed for precision: 8-bit precision needs 2^8=256 clock cycles per number
- Correlated streams corrupt statistics (must use truly independent random sources)
- Slower than binary networks for same precision; only wins on hardware area
---
## Synthesis: What Hardware Knows That Software Has Forgotten
### 1. Locality is not a limitation — it's the design
Every hardware learning rule (STDP, Hebbian, LUT write, ΣΔ correction) updates
using only information at the site of the computation. No global loss function.
Software tries to compensate for locality (backprop broadcasts global gradient).
Hardware makes locality a feature: each synapse computes its own update.
**Implication for BNNBot:** the wave hit system IS the non-local signal.
Use it sparingly, like a neuromodulator: it modulates the sign/scale of local updates,
but the local activity traces must already exist at the synapse.
### 2. Eligibility traces solve the temporal credit assignment problem without backprop
The hardware solution to "which synapse caused the reward 10 ticks later":
every recently-active synapse maintains a decaying trace. When reward arrives,
it multiplies the trace. Traces decay in hardware via RC circuits.
**Implication:** for BNNBot, each in-flight wave needs to carry the eligibility trace
of which neurons fired when it was spawned. The wave system already stores the shot
context — add which neurons/weights were active.
### 3. The LUT IS the function approximator
A K-LUT can represent any K-variable Boolean function. There's nothing to train
in the gradient sense — you just write the truth table entry for each observed pattern.
This is the most efficient possible function approximator for discrete inputs.
**Implication:** WNN/WiSARD with reward-based table updates is the direct software
implementation of what hardware learning does. It's not an approximation of gradient
descent — it's a different algorithm entirely, and it's O(1) per update.
### 4. Energy minimization is free in recurrent hardware
Hopfield/Ising machines settle to energy minima by running physics.
Software has to simulate this expensively. But for our problem, the "energy landscape"
is implicit in our wave data — we just need the right representation.
**Implication:** the Hopfield associative memory is most useful as a pattern completion
engine for "I've seen this state context before" — i.e., case-based reasoning over
battle history.
### 5. Sigma-Delta: you only need the sign of accumulated error
The entire ΣΔ insight is that a 1-bit quantizer + integrator converges to the
correct value. We are throwing away information by keeping full-precision errors when
a signed accumulator does the same job.
**Implication:** the existing Hebbian residual table could be replaced by a ΣΔ
correction accumulator — simpler, same convergence, more noise-robust.
---
## Priority Ranking for BNNBot
| Mechanism | Fit | Cost | Priority |
|-----------|-----|------|----------|
| WNN/WiSARD (LUT-based) | High — directly maps to binary input + wave reward | Medium (memory) | **1st** |
| Three-factor / eligibility traces | High — solves delayed credit assignment | Low (add trace to wave struct) | **2nd** |
| ΣΔ correction accumulator | High — replaces/improves Hebbian residual table | Very low | **3rd** |
| PLL trajectory tracking | Medium — works for orbiting enemies | Low | **4th** |
| Hopfield associative memory | Medium — limited capacity, O(N²) weights | High | **5th** |
| Belief propagation | Low — requires manual factor graph design | Medium | **6th** |
| Stochastic computing | Low — representation mismatch with deterministic bits | Medium | **7th** |
| Evolvable hardware / GA | Very Low — too slow for online single-battle learning | Very High | skip |
---
## Most Actionable Finding
**WNN (Weightless Neural Network) with reward-counted tables** is the closest
direct translation of hardware learning to software. It:
- Requires zero multiplication (lookup only)
- Learns online from wave hits in one write per hit
- Naturally handles the 690-bit binary input
- Has proven capacity for pattern recognition (WiSARD commercial since 1984)
- The "neurons" are literally memory addresses — no parameters to tune
With K=16 bits per RAM node and 45 nodes: 45 × 64KB = 2.9MB RAM, sub-microsecond
inference. This is the most direct hardware-inspired solution that respects all
BNNBot constraints.