diff --git a/MISSION.md b/MISSION.md new file mode 100644 index 0000000..42ba9e9 --- /dev/null +++ b/MISSION.md @@ -0,0 +1,18 @@ +# Mission + +Build a spiking neural network (SNN) bot for Tank Royale that learns to aim at moving targets using binary operations and biologically-inspired learning rules — not analytical formulas. The system must generalize from predictable movers to evasive opponents. + +## Why + +- Classical aiming (analytical lead formulas) can't handle unpredictable movement +- SNNs offer energy-efficient, event-driven learning suited to real-time control +- Binary operations (Hamming distance, popcount) are fast and hardware-friendly +- The long-term goal is a bot that improves through experience, not programming + +## Current State + +Two implementations exist in `SNNBot_garage/src/`: +- **Grid accumulator** (`reservoir.nim`, USE_RESERVOIR=true): Lookup table mapping (vPerp, distance) → lead offset. Works (~48% hit rate) but can't scale to more inputs. +- **SNN** (`SNNBot.nim`, USE_RESERVOIR=false): LIF spiking network with SuperSpike-inspired learning. Dormant — not currently active. + +The grid hit its ceiling. The SNN path is the intended future. diff --git a/NOTES.md b/NOTES.md new file mode 100644 index 0000000..e73c0e7 --- /dev/null +++ b/NOTES.md @@ -0,0 +1,8 @@ +# Teaching Notes + +- User prefers learning by understanding their own code, not abstract theory +- User wants binary operations / binary vectors as the core representation +- User values deep root-cause thinking over quick patches +- User reads Italian natively, English fluently +- Previous session history: tried WTA bins, exemplar ring buffer, grid accumulator — each hit different ceilings +- User explicitly rejected analytical formulas: "we need to have a mind projected to harder movements to predict" diff --git a/RESOURCES.md b/RESOURCES.md new file mode 100644 index 0000000..428be72 --- /dev/null +++ b/RESOURCES.md @@ -0,0 +1,19 @@ +# Resources + +## Primary Sources + +- [SuperSpike paper (Zenke & Ganguli 2018)](https://doi.org/10.1162/neco_a_01086) — The learning rule implemented in the SNN path. Three-factor rule: eligibility trace × error signal. +- [Tank Royale API docs](https://robocode-dev.github.io/tank-royale/) — Bot API, game physics, event model. +- [Leaky Integrate-and-Fire model](https://neuronaldynamics.epfl.ch/online/Ch1.S3.html) — The neuron model used (EPFL textbook, free online). + +## Codebase + +- `SNNBot_garage/src/SNNBot.nim` — Main bot with SNN and reservoir paths +- `SNNBot_garage/src/reservoir.nim` — Grid accumulator (current active path) +- `SNNBot_garage/tests/test_bullet_economy.nim` — Automated battle test + +## To Explore + +- Hebbian learning / STDP for binary spikes +- Hyperdimensional computing for control tasks +- `SNNBot_garage/research/binary-snn-learning.md` — Research notes (if exists) diff --git a/SNNBot_garage/SNNBot b/SNNBot_garage/SNNBot index a2bfc31..04d649a 100755 Binary files a/SNNBot_garage/SNNBot and b/SNNBot_garage/SNNBot differ diff --git a/SNNBot_garage/tests/out/test_bullet_economy b/SNNBot_garage/tests/out/test_bullet_economy new file mode 100755 index 0000000..2e61c82 Binary files /dev/null and b/SNNBot_garage/tests/out/test_bullet_economy differ diff --git a/assets/quiz.js b/assets/quiz.js new file mode 100644 index 0000000..2ee92c1 --- /dev/null +++ b/assets/quiz.js @@ -0,0 +1,21 @@ +// Minimal quiz widget — shared across lessons +document.addEventListener('DOMContentLoaded', () => { + document.querySelectorAll('.quiz').forEach(quiz => { + const radios = quiz.querySelectorAll('input[type="radio"]'); + const feedbackCorrect = quiz.querySelector('.feedback.correct'); + const feedbackWrong = quiz.querySelector('.feedback.wrong'); + + radios.forEach(radio => { + radio.addEventListener('change', () => { + if (feedbackCorrect) feedbackCorrect.style.display = 'none'; + if (feedbackWrong) feedbackWrong.style.display = 'none'; + + if (radio.dataset.correct === 'true') { + if (feedbackCorrect) feedbackCorrect.style.display = 'block'; + } else { + if (feedbackWrong) feedbackWrong.style.display = 'block'; + } + }); + }); + }); +}); diff --git a/assets/style.css b/assets/style.css new file mode 100644 index 0000000..c0000c9 --- /dev/null +++ b/assets/style.css @@ -0,0 +1,158 @@ +/* Teaching workspace — shared lesson stylesheet */ +:root { + --fg: #2d2d2d; + --bg: #fffff8; + --accent: #e63946; + --code-bg: #f5f5f0; + --border: #ddd; + --link: #457b9d; + --success: #2a9d8f; + --warning: #e9c46a; + --font-body: 'Palatino Linotype', 'Book Antiqua', Palatino, serif; + --font-code: 'Fira Code', 'Source Code Pro', 'Consolas', monospace; + --font-sans: -apple-system, BlinkMacSystemFont, 'Segoe UI', sans-serif; +} + +* { box-sizing: border-box; margin: 0; padding: 0; } + +body { + font-family: var(--font-body); + color: var(--fg); + background: var(--bg); + max-width: 740px; + margin: 0 auto; + padding: 2rem 1.5rem 4rem; + line-height: 1.7; + font-size: 18px; +} + +h1 { font-size: 2rem; margin: 2rem 0 1rem; font-weight: 400; letter-spacing: -0.02em; } +h2 { font-size: 1.4rem; margin: 2rem 0 0.8rem; font-weight: 600; color: var(--accent); border-bottom: 1px solid var(--border); padding-bottom: 0.3rem; } +h3 { font-size: 1.1rem; margin: 1.5rem 0 0.5rem; font-weight: 600; } + +p { margin: 0.8rem 0; } +a { color: var(--link); text-decoration: none; border-bottom: 1px solid transparent; } +a:hover { border-bottom-color: var(--link); } + +code { + font-family: var(--font-code); + font-size: 0.85em; + background: var(--code-bg); + padding: 0.15em 0.4em; + border-radius: 3px; +} + +pre { + background: var(--code-bg); + border-left: 3px solid var(--accent); + padding: 1rem 1.2rem; + margin: 1rem 0; + overflow-x: auto; + line-height: 1.5; + font-size: 0.85rem; +} +pre code { background: none; padding: 0; } + +.note { + background: #f0f7ff; + border-left: 3px solid var(--link); + padding: 0.8rem 1rem; + margin: 1rem 0; + font-size: 0.95rem; +} + +.warning { + background: #fff8e1; + border-left: 3px solid var(--warning); + padding: 0.8rem 1rem; + margin: 1rem 0; + font-size: 0.95rem; +} + +.key-insight { + background: #f0fff4; + border-left: 3px solid var(--success); + padding: 0.8rem 1rem; + margin: 1rem 0; + font-size: 0.95rem; +} + +figure { + margin: 1.5rem 0; + text-align: center; +} +figcaption { + font-size: 0.85rem; + color: #666; + margin-top: 0.5rem; + font-style: italic; +} + +table { + border-collapse: collapse; + width: 100%; + margin: 1rem 0; + font-size: 0.9rem; +} +th, td { + border: 1px solid var(--border); + padding: 0.5rem 0.8rem; + text-align: left; +} +th { background: var(--code-bg); font-weight: 600; } + +.diagram { + font-family: var(--font-code); + font-size: 0.8rem; + line-height: 1.4; + white-space: pre; + background: var(--code-bg); + padding: 1.2rem; + margin: 1rem 0; + border-radius: 4px; + overflow-x: auto; +} + +.quiz { + background: var(--code-bg); + border: 1px solid var(--border); + border-radius: 6px; + padding: 1.2rem; + margin: 1.5rem 0; +} +.quiz h3 { margin-top: 0; color: var(--accent); } +.quiz label { + display: block; + padding: 0.4rem 0.6rem; + margin: 0.3rem 0; + border-radius: 4px; + cursor: pointer; + font-family: var(--font-sans); + font-size: 0.9rem; +} +.quiz label:hover { background: #e8e8e0; } +.quiz input[type="radio"] { margin-right: 0.5rem; } +.quiz .feedback { + display: none; + margin-top: 0.8rem; + padding: 0.6rem; + border-radius: 4px; + font-size: 0.9rem; +} +.quiz .feedback.correct { background: #d4edda; display: block; } +.quiz .feedback.wrong { background: #f8d7da; display: block; } + +.nav-footer { + margin-top: 3rem; + padding-top: 1rem; + border-top: 1px solid var(--border); + font-family: var(--font-sans); + font-size: 0.85rem; + color: #666; +} + +@media print { + body { max-width: none; padding: 1cm; font-size: 11pt; } + .quiz { page-break-inside: avoid; } + pre { font-size: 9pt; } +} diff --git a/lessons/0001-snnbot-architecture-overview.html b/lessons/0001-snnbot-architecture-overview.html new file mode 100644 index 0000000..149ea41 --- /dev/null +++ b/lessons/0001-snnbot-architecture-overview.html @@ -0,0 +1,267 @@ + + + + + + Lesson 1: Your SNN Bot — Architecture Overview + + + + +

Your SNN Bot: Architecture Overview

+

Lesson 1 of N — SNNBot_garage/src/SNNBot.nim

+ +

This lesson walks through your own code. No abstract theory first — we start +from what you already built and explain why it is shaped the way it is.

+ + + +

1. The State Machine

+ +

Your bot does not fire every tick. It cycles through three phases, each with a +distinct job. The cycle is encoded in the Phase enum and the +case bot.phase block inside run().

+ +
┌─────────┐ gun aimed ┌─────────┐ bullet fired ┌──────────┐ + │ DECIDE │ ──────────────→ │ WAITING │ ──────────────→ │ EVALUATE │ + │ │ │ │ │ │ + │ predict │ │ aim+fire│ │ learn │ + └─────────┘ └─────────┘ └──────────┘ + ↑ │ + └───────────────────────────────────────────────────────────┘
+ +

DECIDE

+

Takes a snapshot of the enemy state right now — position, velocity direction, +speed, distance — and asks the model: what lead angle should I aim at? +The answer is stored in bot.targetAngle. Crucially, all input values +are frozen into decide* fields (decideEnemyX, +decideVelDirDeg, decideDist, etc.) so that the learning +step later has a consistent ground truth. Phase advances to WAITING.

+ +

WAITING

+

Calls aimTo(bot.targetAngle, gunDir) every tick, which sets the gun +turn rate toward the target. When the aiming error drops below 2° and +the bot has enough energy, it fires and advances to EVALUATE. No learning happens +here — this phase is purely mechanical.

+ +

EVALUATE

+

Uses the frozen DECIDE-time snapshot to compute where the enemy will be +when the bullet arrives, derives the correct lead angle, and feeds the error into +the model (grid or SNN). Then immediately jumps back to DECIDE for the next shot. +This is the only phase where weights change.

+ +
+ Why freeze the snapshot? Between DECIDE and EVALUATE the bot + scans the arena and accumulates new sensor data. If EVALUATE read live values, + the "correct answer" would be computed from different inputs than the prediction + was made from — like grading a test with a different question than the one + asked. The decide* fields fix this. +
+ + + +

2. Two Brains — Grid vs SNN

+ +

The compile-time constant USE_RESERVOIR (line 46 of +SNNBot.nim) selects which brain runs. Both brains expose the same +interface to the state machine: given inputs, return a lead angle offset; +given error, update yourself.

+ +

Grid accumulator (reservoir.nim, currently active)

+ +

A 17×8 = 136-cell lookup table. Rows are perpendicular velocity bins +(integer values −8 to +8), columns are distance bands (125 px each). Each +cell stores a circular mean as three floats: sumSin, +sumCos, count.

+ + + + + + +
ProcWhat it does
initLeadGrid()Warm-starts every cell with the analytical value arcsin(vPerp / 14.0) so the bot is not blind on round 1.
forward(vPerp, dist)Interpolates over the 3×3 neighborhood of cells (weights 1 / 0.5 / 0.25 for direct / cardinal / diagonal). Returns -999.0 when no cell in the neighborhood has data.
learn(vPerp, dist, correctOffset)Adds sin(correctOffset) and cos(correctOffset) to the matching cell. The circular mean is implicit: atan2(sumSin/count, sumCos/count).
+ +

Current performance: ~48% hit rate against WallsBot. The ceiling is the input +space — only two dimensions (vPerp, distance). Adding heading change, evasion +pattern, etc. would require exponentially more cells.

+ +

SNN path (inline in SNNBot.nim, dormant)

+ +

A three-layer spiking network: 80 input neurons → 12 hidden LIF neurons +→ 2 output channels (sin, cos of lead angle).

+ + + + + + +
LayerSizeRole
Input80Population code: bearing (36 neurons, offset 0), velocity direction (36 neurons, offset 36), speed thermometer (8 neurons, offset 72)
Hidden12 LIFEach neuron integrates weighted input, leaks by factor 0.9/tick, spikes when membrane potential ≥ 0.08, resets to 0
Output2 linearWeighted sum of hidden spikes → sin and cos channels; decoded via atan2 to get angle
+ +

The LIF update per hidden neuron h each tick:

+
V[h] = LEAK * V[h] + Σ_i (input[i] * w_ih[i,h])
+if V[h] >= THRESH: spike[h] = 1.0; V[h] = 0.0
+else:              spike[h] = 0.0
+ +

Learning uses the SuperSpike three-factor rule (Zenke & Ganguli 2018):

+
Δw = η × pre_trace × σ'(V) × error
+

where σ'(V) is a surrogate derivative (peaks at threshold, giving +gradient direction through the non-differentiable spike). Hidden neurons receive +their error signal via fixed random feedback weights bFb — this +avoids the weight-transport problem of backpropagation.

+ +
+ Why is the SNN dormant? USE_RESERVOIR = true in the + source. The SNN path exists and compiles, but the grid is better-tuned right now. + The SNN is the intended long-term path because it can handle arbitrary input + dimensionality — you just add more input neurons. +
+ + + +

3. The Lead Prediction Problem

+ +

What is the bot actually trying to learn?

+ +

At DECIDE time the enemy is at absolute bearing B (world-frame +degrees), moving with some velocity, at distance D pixels. A +bullet fired at power 2 travels at 14 px/tick. By the time the +bullet covers distance D, the enemy has moved.

+ +
enemy now + ★ ──────────────────→ ★ enemy later + | / + | bullet path / enemy moves laterally + | / + ●─────────────────/ + your bot aim here (B + offset)
+ +

The perpendicular component of enemy velocity relative to the bullet line is +called vPerp. It is computed in the DECIDE branch:

+ +
let relVelDir = velDirDeg - absBearing   # velocity angle relative to bullet line
+let vPerp = velSpeed * sin(relVelDir)    # lateral component only
+ +

For a constant-velocity target the perfect offset is:

+
offset = arcsin(vPerp / bulletSpeed)     # ≈ 35° at full speed, perpendicular
+ +

The grid is warm-started with exactly this formula via +initLeadGrid(). The SNN must learn this same +relationship from experience, without being told the formula.

+ +
+ Why not just use the formula? The formula only works for + constant linear movement. A skilled opponent changes direction, jinks, circles. + The grid/SNN builds an empirical model from what actually happens, not what + physics predicts for an idealized case. +
+ + + +

4. The Learning Signal

+ +

EVALUATE is where the bot discovers how wrong its prediction was. The logic +(both grid and SNN paths share this upstream computation):

+ +
# 1. How long will the bullet travel?
+travelTime = decideDist / BULLET_SPEED
+
+# 2. Where will the enemy be at impact?
+futureX = decideEnemyX + cos(decideVelDirDeg) * decideVelSpeed * travelTime
+futureY = decideEnemyY + sin(decideVelDirDeg) * decideVelSpeed * travelTime
+
+# 3. What bearing should we have aimed at?
+correctAngle = directionTo(myX, myY, futureX, futureY)
+
+# 4. Error = what we should have aimed - what the bare enemy bearing was
+correctOffset = normalizeRelativeAngle(correctAngle - lastAbsBearing)
+ +

Then the two paths diverge:

+ + + + + + + + + + + +
PathWhat happens with correctOffset
GridCalls res.learn(lastVPerp, decideDist, correctOffset) which + adds sin/cos of the correct offset to the matching cell. Only learns if the + error exceeds the adaptive dead zone (0.3–0.5° depending on + enemy speed).
SNNCalls superSpikeUpdate() with the frozen spike counts and + voltages from DECIDE. The three-factor rule adjusts all weights in proportion + to pre-synaptic trace × surrogate derivative × output error.
+ +
+ The snapshot fix in practice. Before the fix, EVALUATE used + live enemy position. The grid learned from a target that had moved since the + prediction was made — corrupted signal. Now all EVALUATE inputs come from + the decide* fields frozen at DECIDE time. This is why the fields + exist despite looking redundant with lastEnemyX/Y. +
+ + + +

5. Check Your Understanding

+ +
+

Q1: Which phase updates the model weights?

+ + + + +
Correct. DECIDE predicts, WAITING aims, EVALUATE measures the error and updates. The separation is intentional: you need a complete prediction-fire-outcome cycle before you have a learning signal.
+
Not quite. Only EVALUATE updates weights — it is the only phase that knows the outcome (where the enemy ended up relative to where you aimed).
+
+ +
+

Q2: Why does the grid use circular mean (sumSin/sumCos) instead of a simple average of the offset values?

+ + + + +
Correct. This is the classic circular statistics problem. By storing sin and cos components separately, the atan2(sumSin, sumCos) reconstruction always gives the correct angular mean regardless of wrap-around. A simple numeric average of degree values would give nonsense near 0°/360°.
+
Not quite. The key issue is angle wrap-around. A simple average of 1 and 359 gives 180 — the exact opposite direction. sumSin + sumCos encodes direction as a vector, and atan2 decodes it correctly.
+
+ +
+

Q3: What does forward() return when a grid cell (and all its neighbors) has no data?

+ + + + +
Correct. -999.0 is the sentinel. When DECIDE receives it, bot.gridHasData is set false and the bot aims at the bare enemy bearing with no lead (cold-start behavior). It is not the warm-start value — warm-start is already baked into cells with count = 0.1, so a cell with only warm-start data will return a value. The sentinel fires only when totalW == 0.0 in the neighborhood loop — meaning all nearby cells are completely empty.
+
Look at the last line of the forward() proc in reservoir.nim: if totalW == 0.0: return -999.0. That sentinel is what the DECIDE branch checks with if gridOffset <= -999.0.
+
+ + + +

6. Next Steps

+ +

You now understand the full control loop: DECIDE → WAITING → +EVALUATE, how the grid accumulator stores and retrieves lead angles, and how the +SNN path is structured but dormant.

+ +

In the next lesson we will dive into the SNN path in detail: LIF neuron dynamics +tick by tick, why the surrogate derivative is needed (the spike is +non-differentiable), what SuperSpike's three factors are and why each one is +there, and what it would take to activate USE_RESERVOIR = false +without the hit rate collapsing.

+ +
+ Questions? Ask your agent — it is your teacher and can clarify anything + above, show you the exact line in the source, or run a test to check a + hypothesis. +
+ + + + + + diff --git a/lessons/0002-snn-internals-lif-and-superspike.html b/lessons/0002-snn-internals-lif-and-superspike.html new file mode 100644 index 0000000..c3fd090 --- /dev/null +++ b/lessons/0002-snn-internals-lif-and-superspike.html @@ -0,0 +1,263 @@ + + + + + + Lesson 2: Inside the SNN — LIF Neurons and SuperSpike Learning + + + + +

Inside the SNN: LIF Neurons and SuperSpike Learning

+

Lesson 2 of N — SNNBot_garage/src/SNNBot.nim

+ +

In Lesson 1, you saw the architecture: 80 input neurons → 12 hidden LIF neurons → 2 output channels (sin/cos). This lesson breaks open each piece — how neurons fire, how inputs are encoded, how the network learns. Every code snippet is from YOUR implementation in SNNBot.nim.

+ + + +

1. The Leaky Integrate-and-Fire Neuron

+ +

The LIF model as implemented:

+ +
V[t] = LEAK × V[t-1] + weighted_input
+if V >= THRESH: spike! V = 0
+ +

Key constants from the code:

+ + +
+ With 80 inputs and weights initialized at ±0.1, the expected input sum is ~0. The low threshold (0.08) means even small positive fluctuations cause spikes. This makes the hidden layer very active early on — most neurons fire most ticks. +
+ +

The exact forward pass code for the hidden layer:

+ +
for h in 0 ..< N_HID:
+    var wsum = 0.0
+    for i in 0 ..< N_IN:
+      wsum += inputs[i] * snn.wih[i * N_HID + h]
+    snn.vHid[h] = LEAK * snn.vHid[h] + wsum
+    if snn.vHid[h] >= THRESH:
+      spikesOut[h] = 1.0
+      snn.vHid[h] = 0.0  # reset
+    else:
+      spikesOut[h] = 0.0
+ +
+ The LIF neuron is just a leaky accumulator with a threshold. It is the simplest spiking neuron — one step above a perceptron. The "leak" gives it temporal memory: recent inputs matter more than old ones. +
+ + + +

2. Input Encoding — Triangular Interpolation

+ +

Continuous values (bearing, velocity direction, speed) must become spike patterns. The bearing encoder uses 36 neurons covering 360°, each neuron representing a 10° band:

+ +
proc encodeBearing(inputs: var array[N_IN, float], bearing: float, offset: int) =
+  let norm = ((bearing + 180.0) / BAND_DEG)  # 0..36
+  let lo = int(norm) mod 36
+  let hi = (lo + 1) mod 36
+  let frac = norm - float(int(norm))
+  inputs[offset + lo] = 1.0 - frac  # stronger for closer band
+  inputs[offset + hi] = frac         # weaker for farther band
+ +

Triangular interpolation in action:

+ +
bearing = 25°
+band 2 (20°): activation = 0.5  ▓▓▓▓▓░░░░░
+band 3 (30°): activation = 0.5  ▓▓▓▓▓░░░░░
+all others:   activation = 0.0  ░░░░░░░░░░
+ +
+ This is population coding — the same trick the brain uses for direction. No single neuron says "25 degrees"; the ratio between two adjacent neurons encodes it. Smooth interpolation means similar angles activate similar patterns. +
+ +

The 80-neuron layout:

+ + +
+ Notice: bearing and velocity direction both use ABSOLUTE angles. The previous session discovered that using RELATIVE bearing caused a feedback loop — aiming changed the input, which changed the aim. Absolute angles break this loop. +
+ + + +

3. The Output — Polar Coding

+ +

The output layer sums hidden spikes weighted by learned coefficients:

+ +
sinOut = 0.0; cosOut = 0.0
+for h in 0 ..< N_HID:
+    sinOut += spikesOut[h] * snn.wSin[h]
+    cosOut += spikesOut[h] * snn.wCos[h]
+ +

Then decoded: angle = atan2(sinOut, cosOut) × 180/π

+ +
+ Why sin/cos instead of outputting an angle directly? Because angles wrap — 359° and 1° are close, but numerically far apart. Sin/cos is the standard trick: the network outputs a point on the unit circle, and atan2 recovers the angle. No wrapping discontinuity. +
+ +

N_INFER = 10: the network runs 10 ticks on the SAME input, accumulating sin/cos outputs. This averaging stabilizes the output — a single tick's spikes are noisy (binary), but the average over 10 ticks is smooth.

+ + + +

4. The SuperSpike Learning Rule

+ +

4a. The Problem: Spikes Aren't Differentiable

+ +

A spike is binary: 0 or 1. You can't take the gradient of a step function — it's zero everywhere except at the threshold, where it's infinity. So backpropagation doesn't work directly.

+ +

4b. The Surrogate Gradient

+ +

SuperSpike replaces the true derivative with a smooth surrogate:

+ +
proc surrogateDerivative(v: float): float =
+  let x = BETA * (v - THRESH)
+  result = 1.0 / ((1.0 + abs(x)) * (1.0 + abs(x)))
+ +

This is a bell curve centered at v = THRESH. It is large when the voltage is NEAR the threshold (the neuron almost spiked or just barely spiked), and small when far from threshold (irrelevant neurons don't learn).

+ +
σ'(v) + 1.0 | ∧ + | / \ + 0.5 | / \ + | / \ + 0.0 |----/---------\---- + 0 THRESH 2×THRESH v
+ +

4c. The Three-Factor Rule

+ +

Each weight update is the product of THREE factors.

+ +

For hidden→output weights (wSin, wCos):

+ +
Δw = η × rate_h × σ'(V_h) × error
+ + + +

All three must be non-zero for learning to happen. A neuron that didn't fire (rate=0) doesn't learn. A neuron far from threshold (σ'≈0) doesn't learn. If the output is correct (error=0), nothing learns.

+ +

For input→hidden weights (wih):

+ +
Δw = η_ih × preTrace_i × σ'(V_h) × error_h
+ + + +

4d. Random Feedback Alignment

+ +

How does the hidden layer know its error? In backprop, you'd use the transpose of the output weights. SuperSpike uses RANDOM FIXED weights instead:

+ +
let errHid = snn.bFb[h * 2 + 0] * errSin + snn.bFb[h * 2 + 1] * errCos
+ +

These bFb weights are initialized randomly and NEVER updated. This is called feedback alignment — a controversial but effective shortcut. The hidden layer learns to align its representation with these random projections.

+ +
+ This is the key advantage over backprop for spiking networks: no need to propagate gradients through the spike function. The random feedback matrix B replaces WT. It works because the forward weights W gradually align with B during training. +
+ + + +

5. Weight Initialization

+ +
proc initSNN(snn: var SNN) =
+  for w in snn.wih.mitems:  w = rand(0.2) - 0.1   # ±0.1
+  for w in snn.wSin.mitems: w = rand(0.2) - 0.1
+  for w in snn.wCos.mitems: w = rand(0.2) - 0.1
+  for b in snn.bFb.mitems:  b = rand(2.0) - 1.0   # ±1.0, fixed forever
+ +
+ Weights start small (±0.1) with room to grow to ±1.0 (W_CLAMP). Feedback weights are larger (±1.0) because they need to project meaningful error signals. +
+ + + +

6. Why Is the SNN Dormant?

+ +

The honest answer:

+ + +
+ Active bug: The SNN path currently sets targetAngle = gunDir + snnAngle (RELATIVE to gun), not absolute bearing. This is the feedback loop bug that was fixed for the grid path. The SNN path still has this bug. +
+ + + +

7. Check Your Understanding

+ +
+

Q1: What does the surrogate derivative do?

+ + + + +
Correct. The spike function is a step — zero gradient everywhere except the threshold. The surrogate replaces it with a smooth bell curve so that gradient-based updates can flow through. It is a deliberate approximation, not a description of what the neuron physically does.
+
Not quite. The surrogate derivative exists specifically to give a usable gradient signal through the non-differentiable spike threshold, enabling weight updates that would otherwise be impossible.
+
+ +
+

Q2: In the three-factor rule Δw = η × rate × σ'(V) × error, what happens when a neuron's voltage is far from threshold?

+ + + + +
Correct. The surrogate derivative σ'(V) peaks at threshold and drops toward zero for voltages far from it. A neuron that is either deeply sub-threshold or has just reset contributes almost nothing to the weight update — only neurons near the decision boundary learn.
+
Not quite. σ'(V) is the bell curve centered at THRESH. Far from threshold it is near zero, which multiplies the whole update to near zero. The neuron is effectively excluded from learning that tick.
+
+ +
+

Q3: Why does the SNN use random feedback weights (bFb) instead of transposing the output weights?

+ + + + +
Correct. Transposing W for backprop requires passing gradients through the spike function — which has no useful gradient. Random fixed feedback weights sidestep this entirely. The forward weights gradually align with B during training (feedback alignment), so the signal is noisy but directionally correct.
+
Not quite. The fundamental barrier is the spike function's non-differentiability. Transposing W doesn't help if you can't propagate a gradient through the spike. Random feedback avoids this by not needing gradients through spikes at all.
+
+ +
+

Q4: What is the feedback loop bug in the SNN path?

+ + + + +
Correct. When targetAngle is computed as gunDir + snnAngle, the gun turns toward that target. But the input encoding includes the current gun direction, so as the gun moves, the inputs change, which changes snnAngle, which changes where the gun moves. The system chases its own tail. The fix: use absolute angles in both input and output.
+
Not quite. The bug is the relative output: targetAngle = gunDir + snnAngle. As the gun turns, gunDir changes, which changes the bearing inputs, which changes snnAngle. The output feeds back into the input, creating instability.
+
+ + + +

8. Next Steps

+ +

You now understand both paths in your bot. The grid works but can't scale. The SNN can learn but hasn't been tested with the pipeline fixes. In the next lesson, we'll activate the SNN, fix the feedback loop bug, and run it against Walls — your first live SNN training run.

+ +
+ Questions? Ask your agent. This is complex material — re-read sections 4a–4d until the three-factor rule clicks. +
+ + + + + +