milestone(SNNBot): 48% hit rate, grid accumulator with warm-start + frozen DECIDE snapshot — pausing for BNNBot exploration
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
+18
@@ -0,0 +1,18 @@
|
||||
# Mission
|
||||
|
||||
Build a spiking neural network (SNN) bot for Tank Royale that learns to aim at moving targets using binary operations and biologically-inspired learning rules — not analytical formulas. The system must generalize from predictable movers to evasive opponents.
|
||||
|
||||
## Why
|
||||
|
||||
- Classical aiming (analytical lead formulas) can't handle unpredictable movement
|
||||
- SNNs offer energy-efficient, event-driven learning suited to real-time control
|
||||
- Binary operations (Hamming distance, popcount) are fast and hardware-friendly
|
||||
- The long-term goal is a bot that improves through experience, not programming
|
||||
|
||||
## Current State
|
||||
|
||||
Two implementations exist in `SNNBot_garage/src/`:
|
||||
- **Grid accumulator** (`reservoir.nim`, USE_RESERVOIR=true): Lookup table mapping (vPerp, distance) → lead offset. Works (~48% hit rate) but can't scale to more inputs.
|
||||
- **SNN** (`SNNBot.nim`, USE_RESERVOIR=false): LIF spiking network with SuperSpike-inspired learning. Dormant — not currently active.
|
||||
|
||||
The grid hit its ceiling. The SNN path is the intended future.
|
||||
@@ -0,0 +1,8 @@
|
||||
# Teaching Notes
|
||||
|
||||
- User prefers learning by understanding their own code, not abstract theory
|
||||
- User wants binary operations / binary vectors as the core representation
|
||||
- User values deep root-cause thinking over quick patches
|
||||
- User reads Italian natively, English fluently
|
||||
- Previous session history: tried WTA bins, exemplar ring buffer, grid accumulator — each hit different ceilings
|
||||
- User explicitly rejected analytical formulas: "we need to have a mind projected to harder movements to predict"
|
||||
@@ -0,0 +1,19 @@
|
||||
# Resources
|
||||
|
||||
## Primary Sources
|
||||
|
||||
- [SuperSpike paper (Zenke & Ganguli 2018)](https://doi.org/10.1162/neco_a_01086) — The learning rule implemented in the SNN path. Three-factor rule: eligibility trace × error signal.
|
||||
- [Tank Royale API docs](https://robocode-dev.github.io/tank-royale/) — Bot API, game physics, event model.
|
||||
- [Leaky Integrate-and-Fire model](https://neuronaldynamics.epfl.ch/online/Ch1.S3.html) — The neuron model used (EPFL textbook, free online).
|
||||
|
||||
## Codebase
|
||||
|
||||
- `SNNBot_garage/src/SNNBot.nim` — Main bot with SNN and reservoir paths
|
||||
- `SNNBot_garage/src/reservoir.nim` — Grid accumulator (current active path)
|
||||
- `SNNBot_garage/tests/test_bullet_economy.nim` — Automated battle test
|
||||
|
||||
## To Explore
|
||||
|
||||
- Hebbian learning / STDP for binary spikes
|
||||
- Hyperdimensional computing for control tasks
|
||||
- `SNNBot_garage/research/binary-snn-learning.md` — Research notes (if exists)
|
||||
Binary file not shown.
Executable
BIN
Binary file not shown.
@@ -0,0 +1,21 @@
|
||||
// Minimal quiz widget — shared across lessons
|
||||
document.addEventListener('DOMContentLoaded', () => {
|
||||
document.querySelectorAll('.quiz').forEach(quiz => {
|
||||
const radios = quiz.querySelectorAll('input[type="radio"]');
|
||||
const feedbackCorrect = quiz.querySelector('.feedback.correct');
|
||||
const feedbackWrong = quiz.querySelector('.feedback.wrong');
|
||||
|
||||
radios.forEach(radio => {
|
||||
radio.addEventListener('change', () => {
|
||||
if (feedbackCorrect) feedbackCorrect.style.display = 'none';
|
||||
if (feedbackWrong) feedbackWrong.style.display = 'none';
|
||||
|
||||
if (radio.dataset.correct === 'true') {
|
||||
if (feedbackCorrect) feedbackCorrect.style.display = 'block';
|
||||
} else {
|
||||
if (feedbackWrong) feedbackWrong.style.display = 'block';
|
||||
}
|
||||
});
|
||||
});
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,158 @@
|
||||
/* Teaching workspace — shared lesson stylesheet */
|
||||
:root {
|
||||
--fg: #2d2d2d;
|
||||
--bg: #fffff8;
|
||||
--accent: #e63946;
|
||||
--code-bg: #f5f5f0;
|
||||
--border: #ddd;
|
||||
--link: #457b9d;
|
||||
--success: #2a9d8f;
|
||||
--warning: #e9c46a;
|
||||
--font-body: 'Palatino Linotype', 'Book Antiqua', Palatino, serif;
|
||||
--font-code: 'Fira Code', 'Source Code Pro', 'Consolas', monospace;
|
||||
--font-sans: -apple-system, BlinkMacSystemFont, 'Segoe UI', sans-serif;
|
||||
}
|
||||
|
||||
* { box-sizing: border-box; margin: 0; padding: 0; }
|
||||
|
||||
body {
|
||||
font-family: var(--font-body);
|
||||
color: var(--fg);
|
||||
background: var(--bg);
|
||||
max-width: 740px;
|
||||
margin: 0 auto;
|
||||
padding: 2rem 1.5rem 4rem;
|
||||
line-height: 1.7;
|
||||
font-size: 18px;
|
||||
}
|
||||
|
||||
h1 { font-size: 2rem; margin: 2rem 0 1rem; font-weight: 400; letter-spacing: -0.02em; }
|
||||
h2 { font-size: 1.4rem; margin: 2rem 0 0.8rem; font-weight: 600; color: var(--accent); border-bottom: 1px solid var(--border); padding-bottom: 0.3rem; }
|
||||
h3 { font-size: 1.1rem; margin: 1.5rem 0 0.5rem; font-weight: 600; }
|
||||
|
||||
p { margin: 0.8rem 0; }
|
||||
a { color: var(--link); text-decoration: none; border-bottom: 1px solid transparent; }
|
||||
a:hover { border-bottom-color: var(--link); }
|
||||
|
||||
code {
|
||||
font-family: var(--font-code);
|
||||
font-size: 0.85em;
|
||||
background: var(--code-bg);
|
||||
padding: 0.15em 0.4em;
|
||||
border-radius: 3px;
|
||||
}
|
||||
|
||||
pre {
|
||||
background: var(--code-bg);
|
||||
border-left: 3px solid var(--accent);
|
||||
padding: 1rem 1.2rem;
|
||||
margin: 1rem 0;
|
||||
overflow-x: auto;
|
||||
line-height: 1.5;
|
||||
font-size: 0.85rem;
|
||||
}
|
||||
pre code { background: none; padding: 0; }
|
||||
|
||||
.note {
|
||||
background: #f0f7ff;
|
||||
border-left: 3px solid var(--link);
|
||||
padding: 0.8rem 1rem;
|
||||
margin: 1rem 0;
|
||||
font-size: 0.95rem;
|
||||
}
|
||||
|
||||
.warning {
|
||||
background: #fff8e1;
|
||||
border-left: 3px solid var(--warning);
|
||||
padding: 0.8rem 1rem;
|
||||
margin: 1rem 0;
|
||||
font-size: 0.95rem;
|
||||
}
|
||||
|
||||
.key-insight {
|
||||
background: #f0fff4;
|
||||
border-left: 3px solid var(--success);
|
||||
padding: 0.8rem 1rem;
|
||||
margin: 1rem 0;
|
||||
font-size: 0.95rem;
|
||||
}
|
||||
|
||||
figure {
|
||||
margin: 1.5rem 0;
|
||||
text-align: center;
|
||||
}
|
||||
figcaption {
|
||||
font-size: 0.85rem;
|
||||
color: #666;
|
||||
margin-top: 0.5rem;
|
||||
font-style: italic;
|
||||
}
|
||||
|
||||
table {
|
||||
border-collapse: collapse;
|
||||
width: 100%;
|
||||
margin: 1rem 0;
|
||||
font-size: 0.9rem;
|
||||
}
|
||||
th, td {
|
||||
border: 1px solid var(--border);
|
||||
padding: 0.5rem 0.8rem;
|
||||
text-align: left;
|
||||
}
|
||||
th { background: var(--code-bg); font-weight: 600; }
|
||||
|
||||
.diagram {
|
||||
font-family: var(--font-code);
|
||||
font-size: 0.8rem;
|
||||
line-height: 1.4;
|
||||
white-space: pre;
|
||||
background: var(--code-bg);
|
||||
padding: 1.2rem;
|
||||
margin: 1rem 0;
|
||||
border-radius: 4px;
|
||||
overflow-x: auto;
|
||||
}
|
||||
|
||||
.quiz {
|
||||
background: var(--code-bg);
|
||||
border: 1px solid var(--border);
|
||||
border-radius: 6px;
|
||||
padding: 1.2rem;
|
||||
margin: 1.5rem 0;
|
||||
}
|
||||
.quiz h3 { margin-top: 0; color: var(--accent); }
|
||||
.quiz label {
|
||||
display: block;
|
||||
padding: 0.4rem 0.6rem;
|
||||
margin: 0.3rem 0;
|
||||
border-radius: 4px;
|
||||
cursor: pointer;
|
||||
font-family: var(--font-sans);
|
||||
font-size: 0.9rem;
|
||||
}
|
||||
.quiz label:hover { background: #e8e8e0; }
|
||||
.quiz input[type="radio"] { margin-right: 0.5rem; }
|
||||
.quiz .feedback {
|
||||
display: none;
|
||||
margin-top: 0.8rem;
|
||||
padding: 0.6rem;
|
||||
border-radius: 4px;
|
||||
font-size: 0.9rem;
|
||||
}
|
||||
.quiz .feedback.correct { background: #d4edda; display: block; }
|
||||
.quiz .feedback.wrong { background: #f8d7da; display: block; }
|
||||
|
||||
.nav-footer {
|
||||
margin-top: 3rem;
|
||||
padding-top: 1rem;
|
||||
border-top: 1px solid var(--border);
|
||||
font-family: var(--font-sans);
|
||||
font-size: 0.85rem;
|
||||
color: #666;
|
||||
}
|
||||
|
||||
@media print {
|
||||
body { max-width: none; padding: 1cm; font-size: 11pt; }
|
||||
.quiz { page-break-inside: avoid; }
|
||||
pre { font-size: 9pt; }
|
||||
}
|
||||
@@ -0,0 +1,267 @@
|
||||
<!DOCTYPE html>
|
||||
<html lang="en">
|
||||
<head>
|
||||
<meta charset="UTF-8">
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1.0">
|
||||
<title>Lesson 1: Your SNN Bot — Architecture Overview</title>
|
||||
<link rel="stylesheet" href="../assets/style.css">
|
||||
</head>
|
||||
<body>
|
||||
|
||||
<h1>Your SNN Bot: Architecture Overview</h1>
|
||||
<p style="font-family: var(--font-sans); font-size: 0.85rem; color: #888;">Lesson 1 of N — <code>SNNBot_garage/src/SNNBot.nim</code></p>
|
||||
|
||||
<p>This lesson walks through your own code. No abstract theory first — we start
|
||||
from what you already built and explain <em>why</em> it is shaped the way it is.</p>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>1. The State Machine</h2>
|
||||
|
||||
<p>Your bot does not fire every tick. It cycles through three phases, each with a
|
||||
distinct job. The cycle is encoded in the <code>Phase</code> enum and the
|
||||
<code>case bot.phase</code> block inside <code>run()</code>.</p>
|
||||
|
||||
<div class="diagram"> ┌─────────┐ gun aimed ┌─────────┐ bullet fired ┌──────────┐
|
||||
│ DECIDE │ ──────────────→ │ WAITING │ ──────────────→ │ EVALUATE │
|
||||
│ │ │ │ │ │
|
||||
│ predict │ │ aim+fire│ │ learn │
|
||||
└─────────┘ └─────────┘ └──────────┘
|
||||
↑ │
|
||||
└───────────────────────────────────────────────────────────┘</div>
|
||||
|
||||
<h3>DECIDE</h3>
|
||||
<p>Takes a snapshot of the enemy state right now — position, velocity direction,
|
||||
speed, distance — and asks the model: <em>what lead angle should I aim at?</em>
|
||||
The answer is stored in <code>bot.targetAngle</code>. Crucially, all input values
|
||||
are frozen into <code>decide*</code> fields (<code>decideEnemyX</code>,
|
||||
<code>decideVelDirDeg</code>, <code>decideDist</code>, etc.) so that the learning
|
||||
step later has a consistent ground truth. Phase advances to WAITING.</p>
|
||||
|
||||
<h3>WAITING</h3>
|
||||
<p>Calls <code>aimTo(bot.targetAngle, gunDir)</code> every tick, which sets the gun
|
||||
turn rate toward the target. When the aiming error drops below 2° <em>and</em>
|
||||
the bot has enough energy, it fires and advances to EVALUATE. No learning happens
|
||||
here — this phase is purely mechanical.</p>
|
||||
|
||||
<h3>EVALUATE</h3>
|
||||
<p>Uses the frozen DECIDE-time snapshot to compute where the enemy <em>will</em> be
|
||||
when the bullet arrives, derives the correct lead angle, and feeds the error into
|
||||
the model (grid or SNN). Then immediately jumps back to DECIDE for the next shot.
|
||||
This is the only phase where weights change.</p>
|
||||
|
||||
<div class="key-insight">
|
||||
<strong>Why freeze the snapshot?</strong> Between DECIDE and EVALUATE the bot
|
||||
scans the arena and accumulates new sensor data. If EVALUATE read live values,
|
||||
the "correct answer" would be computed from different inputs than the prediction
|
||||
was made from — like grading a test with a different question than the one
|
||||
asked. The <code>decide*</code> fields fix this.
|
||||
</div>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>2. Two Brains — Grid vs SNN</h2>
|
||||
|
||||
<p>The compile-time constant <code>USE_RESERVOIR</code> (line 46 of
|
||||
<code>SNNBot.nim</code>) selects which brain runs. Both brains expose the same
|
||||
interface to the state machine: <em>given inputs, return a lead angle offset</em>;
|
||||
<em>given error, update yourself</em>.</p>
|
||||
|
||||
<h3>Grid accumulator (<code>reservoir.nim</code>, currently active)</h3>
|
||||
|
||||
<p>A 17×8 = 136-cell lookup table. Rows are perpendicular velocity bins
|
||||
(integer values −8 to +8), columns are distance bands (125 px each). Each
|
||||
cell stores a <strong>circular mean</strong> as three floats: <code>sumSin</code>,
|
||||
<code>sumCos</code>, <code>count</code>.</p>
|
||||
|
||||
<table>
|
||||
<tr><th>Proc</th><th>What it does</th></tr>
|
||||
<tr><td><code>initLeadGrid()</code></td><td>Warm-starts every cell with the analytical value <code>arcsin(vPerp / 14.0)</code> so the bot is not blind on round 1.</td></tr>
|
||||
<tr><td><code>forward(vPerp, dist)</code></td><td>Interpolates over the 3×3 neighborhood of cells (weights 1 / 0.5 / 0.25 for direct / cardinal / diagonal). Returns <code>-999.0</code> when no cell in the neighborhood has data.</td></tr>
|
||||
<tr><td><code>learn(vPerp, dist, correctOffset)</code></td><td>Adds <code>sin(correctOffset)</code> and <code>cos(correctOffset)</code> to the matching cell. The circular mean is implicit: <code>atan2(sumSin/count, sumCos/count)</code>.</td></tr>
|
||||
</table>
|
||||
|
||||
<p>Current performance: ~48% hit rate against WallsBot. The ceiling is the input
|
||||
space — only two dimensions (vPerp, distance). Adding heading change, evasion
|
||||
pattern, etc. would require exponentially more cells.</p>
|
||||
|
||||
<h3>SNN path (inline in <code>SNNBot.nim</code>, dormant)</h3>
|
||||
|
||||
<p>A three-layer spiking network: 80 input neurons → 12 hidden LIF neurons
|
||||
→ 2 output channels (sin, cos of lead angle).</p>
|
||||
|
||||
<table>
|
||||
<tr><th>Layer</th><th>Size</th><th>Role</th></tr>
|
||||
<tr><td>Input</td><td>80</td><td>Population code: bearing (36 neurons, offset 0), velocity direction (36 neurons, offset 36), speed thermometer (8 neurons, offset 72)</td></tr>
|
||||
<tr><td>Hidden</td><td>12 LIF</td><td>Each neuron integrates weighted input, leaks by factor 0.9/tick, spikes when membrane potential ≥ 0.08, resets to 0</td></tr>
|
||||
<tr><td>Output</td><td>2 linear</td><td>Weighted sum of hidden spikes → sin and cos channels; decoded via <code>atan2</code> to get angle</td></tr>
|
||||
</table>
|
||||
|
||||
<p>The LIF update per hidden neuron <code>h</code> each tick:</p>
|
||||
<pre><code>V[h] = LEAK * V[h] + Σ_i (input[i] * w_ih[i,h])
|
||||
if V[h] >= THRESH: spike[h] = 1.0; V[h] = 0.0
|
||||
else: spike[h] = 0.0</code></pre>
|
||||
|
||||
<p>Learning uses the SuperSpike three-factor rule (Zenke & Ganguli 2018):</p>
|
||||
<pre><code>Δw = η × pre_trace × σ'(V) × error</code></pre>
|
||||
<p>where <code>σ'(V)</code> is a surrogate derivative (peaks at threshold, giving
|
||||
gradient direction through the non-differentiable spike). Hidden neurons receive
|
||||
their error signal via fixed random feedback weights <code>bFb</code> — this
|
||||
avoids the weight-transport problem of backpropagation.</p>
|
||||
|
||||
<div class="warning">
|
||||
<strong>Why is the SNN dormant?</strong> <code>USE_RESERVOIR = true</code> in the
|
||||
source. The SNN path exists and compiles, but the grid is better-tuned right now.
|
||||
The SNN is the intended long-term path because it can handle arbitrary input
|
||||
dimensionality — you just add more input neurons.
|
||||
</div>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>3. The Lead Prediction Problem</h2>
|
||||
|
||||
<p>What is the bot actually trying to learn?</p>
|
||||
|
||||
<p>At DECIDE time the enemy is at absolute bearing <strong>B</strong> (world-frame
|
||||
degrees), moving with some velocity, at distance <strong>D</strong> pixels. A
|
||||
bullet fired at power 2 travels at <strong>14 px/tick</strong>. By the time the
|
||||
bullet covers distance D, the enemy has moved.</p>
|
||||
|
||||
<div class="diagram"> enemy now
|
||||
★ ──────────────────→ ★ enemy later
|
||||
| /
|
||||
| bullet path / enemy moves laterally
|
||||
| /
|
||||
●─────────────────/
|
||||
your bot aim here (B + offset)</div>
|
||||
|
||||
<p>The perpendicular component of enemy velocity relative to the bullet line is
|
||||
called <code>vPerp</code>. It is computed in the DECIDE branch:</p>
|
||||
|
||||
<pre><code>let relVelDir = velDirDeg - absBearing # velocity angle relative to bullet line
|
||||
let vPerp = velSpeed * sin(relVelDir) # lateral component only</code></pre>
|
||||
|
||||
<p>For a constant-velocity target the perfect offset is:</p>
|
||||
<pre><code>offset = arcsin(vPerp / bulletSpeed) # ≈ 35° at full speed, perpendicular</code></pre>
|
||||
|
||||
<p>The grid is <strong>warm-started</strong> with exactly this formula via
|
||||
<code>initLeadGrid()</code>. The SNN must <strong>learn</strong> this same
|
||||
relationship from experience, without being told the formula.</p>
|
||||
|
||||
<div class="key-insight">
|
||||
<strong>Why not just use the formula?</strong> The formula only works for
|
||||
constant linear movement. A skilled opponent changes direction, jinks, circles.
|
||||
The grid/SNN builds an empirical model from what actually happens, not what
|
||||
physics predicts for an idealized case.
|
||||
</div>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>4. The Learning Signal</h2>
|
||||
|
||||
<p>EVALUATE is where the bot discovers how wrong its prediction was. The logic
|
||||
(both grid and SNN paths share this upstream computation):</p>
|
||||
|
||||
<pre><code># 1. How long will the bullet travel?
|
||||
travelTime = decideDist / BULLET_SPEED
|
||||
|
||||
# 2. Where will the enemy be at impact?
|
||||
futureX = decideEnemyX + cos(decideVelDirDeg) * decideVelSpeed * travelTime
|
||||
futureY = decideEnemyY + sin(decideVelDirDeg) * decideVelSpeed * travelTime
|
||||
|
||||
# 3. What bearing should we have aimed at?
|
||||
correctAngle = directionTo(myX, myY, futureX, futureY)
|
||||
|
||||
# 4. Error = what we should have aimed - what the bare enemy bearing was
|
||||
correctOffset = normalizeRelativeAngle(correctAngle - lastAbsBearing)</code></pre>
|
||||
|
||||
<p>Then the two paths diverge:</p>
|
||||
|
||||
<table>
|
||||
<tr><th>Path</th><th>What happens with <code>correctOffset</code></th></tr>
|
||||
<tr>
|
||||
<td>Grid</td>
|
||||
<td>Calls <code>res.learn(lastVPerp, decideDist, correctOffset)</code> which
|
||||
adds sin/cos of the correct offset to the matching cell. Only learns if the
|
||||
error exceeds the adaptive dead zone (0.3–0.5° depending on
|
||||
enemy speed).</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>SNN</td>
|
||||
<td>Calls <code>superSpikeUpdate()</code> with the frozen spike counts and
|
||||
voltages from DECIDE. The three-factor rule adjusts all weights in proportion
|
||||
to pre-synaptic trace × surrogate derivative × output error.</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
<div class="note">
|
||||
<strong>The snapshot fix in practice.</strong> Before the fix, EVALUATE used
|
||||
live enemy position. The grid learned from a target that had moved since the
|
||||
prediction was made — corrupted signal. Now all EVALUATE inputs come from
|
||||
the <code>decide*</code> fields frozen at DECIDE time. This is why the fields
|
||||
exist despite looking redundant with <code>lastEnemyX/Y</code>.
|
||||
</div>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>5. Check Your Understanding</h2>
|
||||
|
||||
<div class="quiz" id="q1">
|
||||
<h3>Q1: Which phase updates the model weights?</h3>
|
||||
<label><input type="radio" name="q1" data-correct="false"> DECIDE</label>
|
||||
<label><input type="radio" name="q1" data-correct="false"> WAITING</label>
|
||||
<label><input type="radio" name="q1" data-correct="true"> EVALUATE</label>
|
||||
<label><input type="radio" name="q1" data-correct="false"> All three phases contribute</label>
|
||||
<div class="feedback correct">Correct. DECIDE predicts, WAITING aims, EVALUATE measures the error and updates. The separation is intentional: you need a complete prediction-fire-outcome cycle before you have a learning signal.</div>
|
||||
<div class="feedback wrong">Not quite. Only EVALUATE updates weights — it is the only phase that knows the outcome (where the enemy ended up relative to where you aimed).</div>
|
||||
</div>
|
||||
|
||||
<div class="quiz" id="q2">
|
||||
<h3>Q2: Why does the grid use circular mean (sumSin/sumCos) instead of a simple average of the offset values?</h3>
|
||||
<label><input type="radio" name="q2" data-correct="false"> It is faster to compute</label>
|
||||
<label><input type="radio" name="q2" data-correct="true"> Angles wrap around 360° — averaging 1° and 359° with simple arithmetic gives 180°, but the correct mean is 0°</label>
|
||||
<label><input type="radio" name="q2" data-correct="false"> It uses less memory per cell</label>
|
||||
<label><input type="radio" name="q2" data-correct="false"> It handles negative angles better</label>
|
||||
<div class="feedback correct">Correct. This is the classic circular statistics problem. By storing sin and cos components separately, the atan2(sumSin, sumCos) reconstruction always gives the correct angular mean regardless of wrap-around. A simple numeric average of degree values would give nonsense near 0°/360°.</div>
|
||||
<div class="feedback wrong">Not quite. The key issue is angle wrap-around. A simple average of 1 and 359 gives 180 — the exact opposite direction. sumSin + sumCos encodes direction as a vector, and atan2 decodes it correctly.</div>
|
||||
</div>
|
||||
|
||||
<div class="quiz" id="q3">
|
||||
<h3>Q3: What does <code>forward()</code> return when a grid cell (and all its neighbors) has no data?</h3>
|
||||
<label><input type="radio" name="q3" data-correct="false"> 0.0 degrees (aim straight at the enemy)</label>
|
||||
<label><input type="radio" name="q3" data-correct="false"> The analytical warm-start value</label>
|
||||
<label><input type="radio" name="q3" data-correct="true"> -999.0 (sentinel value)</label>
|
||||
<label><input type="radio" name="q3" data-correct="false"> NaN</label>
|
||||
<div class="feedback correct">Correct. <code>-999.0</code> is the sentinel. When DECIDE receives it, <code>bot.gridHasData</code> is set false and the bot aims at the bare enemy bearing with no lead (cold-start behavior). It is not the warm-start value — warm-start is already baked into cells with <code>count = 0.1</code>, so a cell with only warm-start data <em>will</em> return a value. The sentinel fires only when <code>totalW == 0.0</code> in the neighborhood loop — meaning all nearby cells are completely empty.</div>
|
||||
<div class="feedback wrong">Look at the last line of the <code>forward()</code> proc in <code>reservoir.nim</code>: <code>if totalW == 0.0: return -999.0</code>. That sentinel is what the DECIDE branch checks with <code>if gridOffset <= -999.0</code>.</div>
|
||||
</div>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>6. Next Steps</h2>
|
||||
|
||||
<p>You now understand the full control loop: DECIDE → WAITING →
|
||||
EVALUATE, how the grid accumulator stores and retrieves lead angles, and how the
|
||||
SNN path is structured but dormant.</p>
|
||||
|
||||
<p>In the next lesson we will dive into the SNN path in detail: LIF neuron dynamics
|
||||
tick by tick, why the surrogate derivative is needed (the spike is
|
||||
non-differentiable), what SuperSpike's three factors are and why each one is
|
||||
there, and what it would take to activate <code>USE_RESERVOIR = false</code>
|
||||
without the hit rate collapsing.</p>
|
||||
|
||||
<div class="note">
|
||||
Questions? Ask your agent — it is your teacher and can clarify anything
|
||||
above, show you the exact line in the source, or run a test to check a
|
||||
hypothesis.
|
||||
</div>
|
||||
|
||||
<div class="nav-footer">
|
||||
Lesson 1 / Architecture Overview —
|
||||
source: <code>SNNBot_garage/src/SNNBot.nim</code>,
|
||||
<code>SNNBot_garage/src/reservoir.nim</code>
|
||||
</div>
|
||||
|
||||
<script src="../assets/quiz.js"></script>
|
||||
</body>
|
||||
</html>
|
||||
@@ -0,0 +1,263 @@
|
||||
<!DOCTYPE html>
|
||||
<html lang="en">
|
||||
<head>
|
||||
<meta charset="UTF-8">
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1.0">
|
||||
<title>Lesson 2: Inside the SNN — LIF Neurons and SuperSpike Learning</title>
|
||||
<link rel="stylesheet" href="../assets/style.css">
|
||||
</head>
|
||||
<body>
|
||||
|
||||
<h1>Inside the SNN: LIF Neurons and SuperSpike Learning</h1>
|
||||
<p style="font-family: var(--font-sans); font-size: 0.85rem; color: #888;">Lesson 2 of N — <code>SNNBot_garage/src/SNNBot.nim</code></p>
|
||||
|
||||
<p>In Lesson 1, you saw the architecture: 80 input neurons → 12 hidden LIF neurons → 2 output channels (sin/cos). This lesson breaks open each piece — how neurons fire, how inputs are encoded, how the network learns. Every code snippet is from YOUR implementation in SNNBot.nim.</p>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>1. The Leaky Integrate-and-Fire Neuron</h2>
|
||||
|
||||
<p>The LIF model as implemented:</p>
|
||||
|
||||
<pre><code>V[t] = LEAK × V[t-1] + weighted_input
|
||||
if V >= THRESH: spike! V = 0</code></pre>
|
||||
|
||||
<p>Key constants from the code:</p>
|
||||
<ul>
|
||||
<li><code>LEAK = 0.9</code> — membrane leaks 10% per tick — neuron "forgets" over ~10 ticks</li>
|
||||
<li><code>THRESH = 0.08</code> — very low threshold — easy to fire</li>
|
||||
</ul>
|
||||
|
||||
<div class="note">
|
||||
With 80 inputs and weights initialized at ±0.1, the expected input sum is ~0. The low threshold (0.08) means even small positive fluctuations cause spikes. This makes the hidden layer very active early on — most neurons fire most ticks.
|
||||
</div>
|
||||
|
||||
<p>The exact forward pass code for the hidden layer:</p>
|
||||
|
||||
<pre><code>for h in 0 ..< N_HID:
|
||||
var wsum = 0.0
|
||||
for i in 0 ..< N_IN:
|
||||
wsum += inputs[i] * snn.wih[i * N_HID + h]
|
||||
snn.vHid[h] = LEAK * snn.vHid[h] + wsum
|
||||
if snn.vHid[h] >= THRESH:
|
||||
spikesOut[h] = 1.0
|
||||
snn.vHid[h] = 0.0 # reset
|
||||
else:
|
||||
spikesOut[h] = 0.0</code></pre>
|
||||
|
||||
<div class="key-insight">
|
||||
The LIF neuron is just a leaky accumulator with a threshold. It is the simplest spiking neuron — one step above a perceptron. The "leak" gives it temporal memory: recent inputs matter more than old ones.
|
||||
</div>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>2. Input Encoding — Triangular Interpolation</h2>
|
||||
|
||||
<p>Continuous values (bearing, velocity direction, speed) must become spike patterns. The bearing encoder uses 36 neurons covering 360°, each neuron representing a 10° band:</p>
|
||||
|
||||
<pre><code>proc encodeBearing(inputs: var array[N_IN, float], bearing: float, offset: int) =
|
||||
let norm = ((bearing + 180.0) / BAND_DEG) # 0..36
|
||||
let lo = int(norm) mod 36
|
||||
let hi = (lo + 1) mod 36
|
||||
let frac = norm - float(int(norm))
|
||||
inputs[offset + lo] = 1.0 - frac # stronger for closer band
|
||||
inputs[offset + hi] = frac # weaker for farther band</code></pre>
|
||||
|
||||
<p>Triangular interpolation in action:</p>
|
||||
|
||||
<pre><code>bearing = 25°
|
||||
band 2 (20°): activation = 0.5 ▓▓▓▓▓░░░░░
|
||||
band 3 (30°): activation = 0.5 ▓▓▓▓▓░░░░░
|
||||
all others: activation = 0.0 ░░░░░░░░░░</code></pre>
|
||||
|
||||
<div class="note">
|
||||
This is population coding — the same trick the brain uses for direction. No single neuron says "25 degrees"; the ratio between two adjacent neurons encodes it. Smooth interpolation means similar angles activate similar patterns.
|
||||
</div>
|
||||
|
||||
<p>The 80-neuron layout:</p>
|
||||
<ul>
|
||||
<li>Neurons 0–35: relative bearing (36 neurons, 10°/band, wraps circularly)</li>
|
||||
<li>Neurons 36–71: velocity direction (36 neurons, 10°/band, wraps circularly)</li>
|
||||
<li>Neurons 72–79: speed (8 neurons, 1 unit/tick per band, clamped at 8)</li>
|
||||
</ul>
|
||||
|
||||
<div class="warning">
|
||||
<strong>Notice:</strong> bearing and velocity direction both use ABSOLUTE angles. The previous session discovered that using RELATIVE bearing caused a feedback loop — aiming changed the input, which changed the aim. Absolute angles break this loop.
|
||||
</div>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>3. The Output — Polar Coding</h2>
|
||||
|
||||
<p>The output layer sums hidden spikes weighted by learned coefficients:</p>
|
||||
|
||||
<pre><code>sinOut = 0.0; cosOut = 0.0
|
||||
for h in 0 ..< N_HID:
|
||||
sinOut += spikesOut[h] * snn.wSin[h]
|
||||
cosOut += spikesOut[h] * snn.wCos[h]</code></pre>
|
||||
|
||||
<p>Then decoded: <code>angle = atan2(sinOut, cosOut) × 180/π</code></p>
|
||||
|
||||
<div class="key-insight">
|
||||
Why sin/cos instead of outputting an angle directly? Because angles wrap — 359° and 1° are close, but numerically far apart. Sin/cos is the standard trick: the network outputs a point on the unit circle, and atan2 recovers the angle. No wrapping discontinuity.
|
||||
</div>
|
||||
|
||||
<p><code>N_INFER = 10</code>: the network runs 10 ticks on the SAME input, accumulating sin/cos outputs. This averaging stabilizes the output — a single tick's spikes are noisy (binary), but the average over 10 ticks is smooth.</p>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>4. The SuperSpike Learning Rule</h2>
|
||||
|
||||
<h3>4a. The Problem: Spikes Aren't Differentiable</h3>
|
||||
|
||||
<p>A spike is binary: 0 or 1. You can't take the gradient of a step function — it's zero everywhere except at the threshold, where it's infinity. So backpropagation doesn't work directly.</p>
|
||||
|
||||
<h3>4b. The Surrogate Gradient</h3>
|
||||
|
||||
<p>SuperSpike replaces the true derivative with a smooth surrogate:</p>
|
||||
|
||||
<pre><code>proc surrogateDerivative(v: float): float =
|
||||
let x = BETA * (v - THRESH)
|
||||
result = 1.0 / ((1.0 + abs(x)) * (1.0 + abs(x)))</code></pre>
|
||||
|
||||
<p>This is a bell curve centered at <code>v = THRESH</code>. It is large when the voltage is NEAR the threshold (the neuron almost spiked or just barely spiked), and small when far from threshold (irrelevant neurons don't learn).</p>
|
||||
|
||||
<div class="diagram">σ'(v)
|
||||
1.0 | ∧
|
||||
| / \
|
||||
0.5 | / \
|
||||
| / \
|
||||
0.0 |----/---------\----
|
||||
0 THRESH 2×THRESH v</div>
|
||||
|
||||
<h3>4c. The Three-Factor Rule</h3>
|
||||
|
||||
<p>Each weight update is the product of THREE factors.</p>
|
||||
|
||||
<p>For hidden→output weights (<code>wSin</code>, <code>wCos</code>):</p>
|
||||
|
||||
<pre><code>Δw = η × rate_h × σ'(V_h) × error</code></pre>
|
||||
|
||||
<ul>
|
||||
<li><strong>η = 0.1</strong>: learning rate</li>
|
||||
<li><strong>rate_h</strong>: spike rate of hidden neuron h (spikes/N_INFER) — "was this neuron active?"</li>
|
||||
<li><strong>σ'(V_h)</strong>: surrogate derivative — "was this neuron near threshold?"</li>
|
||||
<li><strong>error</strong>: target_rate − actual_rate — "how wrong was the output?"</li>
|
||||
</ul>
|
||||
|
||||
<p>All three must be non-zero for learning to happen. A neuron that didn't fire (rate=0) doesn't learn. A neuron far from threshold (σ'≈0) doesn't learn. If the output is correct (error=0), nothing learns.</p>
|
||||
|
||||
<p>For input→hidden weights (<code>wih</code>):</p>
|
||||
|
||||
<pre><code>Δw = η_ih × preTrace_i × σ'(V_h) × error_h</code></pre>
|
||||
|
||||
<ul>
|
||||
<li><strong>preTrace_i</strong>: low-pass filtered input (TRACE_DECAY=0.9) — "was this input recently active?"</li>
|
||||
<li><strong>error_h</strong>: hidden error, projected via random feedback weights</li>
|
||||
</ul>
|
||||
|
||||
<h3>4d. Random Feedback Alignment</h3>
|
||||
|
||||
<p>How does the hidden layer know its error? In backprop, you'd use the transpose of the output weights. SuperSpike uses RANDOM FIXED weights instead:</p>
|
||||
|
||||
<pre><code>let errHid = snn.bFb[h * 2 + 0] * errSin + snn.bFb[h * 2 + 1] * errCos</code></pre>
|
||||
|
||||
<p>These <code>bFb</code> weights are initialized randomly and NEVER updated. This is called feedback alignment — a controversial but effective shortcut. The hidden layer learns to align its representation with these random projections.</p>
|
||||
|
||||
<div class="key-insight">
|
||||
This is the key advantage over backprop for spiking networks: no need to propagate gradients through the spike function. The random feedback matrix B replaces W<sup>T</sup>. It works because the forward weights W gradually align with B during training.
|
||||
</div>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>5. Weight Initialization</h2>
|
||||
|
||||
<pre><code>proc initSNN(snn: var SNN) =
|
||||
for w in snn.wih.mitems: w = rand(0.2) - 0.1 # ±0.1
|
||||
for w in snn.wSin.mitems: w = rand(0.2) - 0.1
|
||||
for w in snn.wCos.mitems: w = rand(0.2) - 0.1
|
||||
for b in snn.bFb.mitems: b = rand(2.0) - 1.0 # ±1.0, fixed forever</code></pre>
|
||||
|
||||
<div class="note">
|
||||
Weights start small (±0.1) with room to grow to ±1.0 (W_CLAMP). Feedback weights are larger (±1.0) because they need to project meaningful error signals.
|
||||
</div>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>6. Why Is the SNN Dormant?</h2>
|
||||
|
||||
<p>The honest answer:</p>
|
||||
<ul>
|
||||
<li>The grid accumulator was introduced as a simpler baseline to debug the aiming pipeline</li>
|
||||
<li>The grid reached 48% hit rate — proving the pipeline works (state machine, bearing calc, fire gate)</li>
|
||||
<li>The SNN was never tuned against the corrected pipeline (frozen DECIDE snapshot, offset-based learning)</li>
|
||||
<li>Key open questions: Does the SNN converge? How many rounds to learn? Is 12 hidden neurons enough?</li>
|
||||
</ul>
|
||||
|
||||
<div class="warning">
|
||||
<strong>Active bug:</strong> The SNN path currently sets <code>targetAngle = gunDir + snnAngle</code> (RELATIVE to gun), not absolute bearing. This is the feedback loop bug that was fixed for the grid path. The SNN path still has this bug.
|
||||
</div>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>7. Check Your Understanding</h2>
|
||||
|
||||
<div class="quiz" id="q1">
|
||||
<h3>Q1: What does the surrogate derivative do?</h3>
|
||||
<label><input type="radio" name="q1" data-correct="true"> Approximates the gradient of the spike function for learning</label>
|
||||
<label><input type="radio" name="q1" data-correct="false"> Smooths the membrane voltage to prevent oscillation</label>
|
||||
<label><input type="radio" name="q1" data-correct="false"> Decays the membrane potential over time</label>
|
||||
<label><input type="radio" name="q1" data-correct="false"> Normalizes the spike rate across neurons</label>
|
||||
<div class="feedback correct">Correct. The spike function is a step — zero gradient everywhere except the threshold. The surrogate replaces it with a smooth bell curve so that gradient-based updates can flow through. It is a deliberate approximation, not a description of what the neuron physically does.</div>
|
||||
<div class="feedback wrong">Not quite. The surrogate derivative exists specifically to give a usable gradient signal through the non-differentiable spike threshold, enabling weight updates that would otherwise be impossible.</div>
|
||||
</div>
|
||||
|
||||
<div class="quiz" id="q2">
|
||||
<h3>Q2: In the three-factor rule Δw = η × rate × σ'(V) × error, what happens when a neuron's voltage is far from threshold?</h3>
|
||||
<label><input type="radio" name="q2" data-correct="false"> The weight update is large because the neuron needs to change</label>
|
||||
<label><input type="radio" name="q2" data-correct="true"> The weight update is zero because σ'(V) ≈ 0</label>
|
||||
<label><input type="radio" name="q2" data-correct="false"> The neuron fires more frequently</label>
|
||||
<label><input type="radio" name="q2" data-correct="false"> The learning rate η is adjusted automatically</label>
|
||||
<div class="feedback correct">Correct. The surrogate derivative σ'(V) peaks at threshold and drops toward zero for voltages far from it. A neuron that is either deeply sub-threshold or has just reset contributes almost nothing to the weight update — only neurons near the decision boundary learn.</div>
|
||||
<div class="feedback wrong">Not quite. σ'(V) is the bell curve centered at THRESH. Far from threshold it is near zero, which multiplies the whole update to near zero. The neuron is effectively excluded from learning that tick.</div>
|
||||
</div>
|
||||
|
||||
<div class="quiz" id="q3">
|
||||
<h3>Q3: Why does the SNN use random feedback weights (<code>bFb</code>) instead of transposing the output weights?</h3>
|
||||
<label><input type="radio" name="q3" data-correct="false"> Random weights are faster to compute</label>
|
||||
<label><input type="radio" name="q3" data-correct="true"> It avoids propagating gradients through non-differentiable spikes</label>
|
||||
<label><input type="radio" name="q3" data-correct="false"> Random weights provide better generalization</label>
|
||||
<label><input type="radio" name="q3" data-correct="false"> The output weights are too small to transpose</label>
|
||||
<div class="feedback correct">Correct. Transposing W for backprop requires passing gradients through the spike function — which has no useful gradient. Random fixed feedback weights sidestep this entirely. The forward weights gradually align with B during training (feedback alignment), so the signal is noisy but directionally correct.</div>
|
||||
<div class="feedback wrong">Not quite. The fundamental barrier is the spike function's non-differentiability. Transposing W doesn't help if you can't propagate a gradient through the spike. Random feedback avoids this by not needing gradients through spikes at all.</div>
|
||||
</div>
|
||||
|
||||
<div class="quiz" id="q4">
|
||||
<h3>Q4: What is the feedback loop bug in the SNN path?</h3>
|
||||
<label><input type="radio" name="q4" data-correct="false"> The SNN uses too few hidden neurons</label>
|
||||
<label><input type="radio" name="q4" data-correct="true"> Output is relative to gun direction, so aiming changes the input which changes the aim</label>
|
||||
<label><input type="radio" name="q4" data-correct="false"> The learning rate is too high</label>
|
||||
<label><input type="radio" name="q4" data-correct="false"> The surrogate derivative is centered at the wrong threshold</label>
|
||||
<div class="feedback correct">Correct. When targetAngle is computed as gunDir + snnAngle, the gun turns toward that target. But the input encoding includes the current gun direction, so as the gun moves, the inputs change, which changes snnAngle, which changes where the gun moves. The system chases its own tail. The fix: use absolute angles in both input and output.</div>
|
||||
<div class="feedback wrong">Not quite. The bug is the relative output: <code>targetAngle = gunDir + snnAngle</code>. As the gun turns, gunDir changes, which changes the bearing inputs, which changes snnAngle. The output feeds back into the input, creating instability.</div>
|
||||
</div>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>8. Next Steps</h2>
|
||||
|
||||
<p>You now understand both paths in your bot. The grid works but can't scale. The SNN can learn but hasn't been tested with the pipeline fixes. In the next lesson, we'll activate the SNN, fix the feedback loop bug, and run it against Walls — your first live SNN training run.</p>
|
||||
|
||||
<div class="note">
|
||||
Questions? Ask your agent. This is complex material — re-read sections 4a–4d until the three-factor rule clicks.
|
||||
</div>
|
||||
|
||||
<div class="nav-footer">
|
||||
Lesson 2 / LIF Neurons and SuperSpike —
|
||||
source: <code>SNNBot_garage/src/SNNBot.nim</code> —
|
||||
<a href="0001-snnbot-architecture-overview.html">← Lesson 1: Architecture Overview</a>
|
||||
</div>
|
||||
|
||||
<script src="../assets/quiz.js"></script>
|
||||
</body>
|
||||
</html>
|
||||
Reference in New Issue
Block a user