Files
SirRoboGarage/lessons/0001-snnbot-architecture-overview.html
T

268 lines
16 KiB
HTML
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Lesson 1: Your SNN Bot — Architecture Overview</title>
<link rel="stylesheet" href="../assets/style.css">
</head>
<body>
<h1>Your SNN Bot: Architecture Overview</h1>
<p style="font-family: var(--font-sans); font-size: 0.85rem; color: #888;">Lesson 1 of N &mdash; <code>SNNBot_garage/src/SNNBot.nim</code></p>
<p>This lesson walks through your own code. No abstract theory first &mdash; we start
from what you already built and explain <em>why</em> it is shaped the way it is.</p>
<!-- ═══════════════════════════════════════════════════════════════════════ -->
<h2>1. The State Machine</h2>
<p>Your bot does not fire every tick. It cycles through three phases, each with a
distinct job. The cycle is encoded in the <code>Phase</code> enum and the
<code>case bot.phase</code> block inside <code>run()</code>.</p>
<div class="diagram"> ┌─────────┐ gun aimed ┌─────────┐ bullet fired ┌──────────┐
│ DECIDE │ ──────────────→ │ WAITING │ ──────────────→ │ EVALUATE │
│ │ │ │ │ │
│ predict │ │ aim+fire│ │ learn │
└─────────┘ └─────────┘ └──────────┘
↑ │
└───────────────────────────────────────────────────────────┘</div>
<h3>DECIDE</h3>
<p>Takes a snapshot of the enemy state right now &mdash; position, velocity direction,
speed, distance &mdash; and asks the model: <em>what lead angle should I aim at?</em>
The answer is stored in <code>bot.targetAngle</code>. Crucially, all input values
are frozen into <code>decide*</code> fields (<code>decideEnemyX</code>,
<code>decideVelDirDeg</code>, <code>decideDist</code>, etc.) so that the learning
step later has a consistent ground truth. Phase advances to WAITING.</p>
<h3>WAITING</h3>
<p>Calls <code>aimTo(bot.targetAngle, gunDir)</code> every tick, which sets the gun
turn rate toward the target. When the aiming error drops below 2&deg; <em>and</em>
the bot has enough energy, it fires and advances to EVALUATE. No learning happens
here &mdash; this phase is purely mechanical.</p>
<h3>EVALUATE</h3>
<p>Uses the frozen DECIDE-time snapshot to compute where the enemy <em>will</em> be
when the bullet arrives, derives the correct lead angle, and feeds the error into
the model (grid or SNN). Then immediately jumps back to DECIDE for the next shot.
This is the only phase where weights change.</p>
<div class="key-insight">
<strong>Why freeze the snapshot?</strong> Between DECIDE and EVALUATE the bot
scans the arena and accumulates new sensor data. If EVALUATE read live values,
the "correct answer" would be computed from different inputs than the prediction
was made from &mdash; like grading a test with a different question than the one
asked. The <code>decide*</code> fields fix this.
</div>
<!-- ═══════════════════════════════════════════════════════════════════════ -->
<h2>2. Two Brains &mdash; Grid vs SNN</h2>
<p>The compile-time constant <code>USE_RESERVOIR</code> (line 46 of
<code>SNNBot.nim</code>) selects which brain runs. Both brains expose the same
interface to the state machine: <em>given inputs, return a lead angle offset</em>;
<em>given error, update yourself</em>.</p>
<h3>Grid accumulator (<code>reservoir.nim</code>, currently active)</h3>
<p>A 17&times;8 = 136-cell lookup table. Rows are perpendicular velocity bins
(integer values &minus;8 to +8), columns are distance bands (125 px each). Each
cell stores a <strong>circular mean</strong> as three floats: <code>sumSin</code>,
<code>sumCos</code>, <code>count</code>.</p>
<table>
<tr><th>Proc</th><th>What it does</th></tr>
<tr><td><code>initLeadGrid()</code></td><td>Warm-starts every cell with the analytical value <code>arcsin(vPerp / 14.0)</code> so the bot is not blind on round 1.</td></tr>
<tr><td><code>forward(vPerp, dist)</code></td><td>Interpolates over the 3&times;3 neighborhood of cells (weights 1 / 0.5 / 0.25 for direct / cardinal / diagonal). Returns <code>-999.0</code> when no cell in the neighborhood has data.</td></tr>
<tr><td><code>learn(vPerp, dist, correctOffset)</code></td><td>Adds <code>sin(correctOffset)</code> and <code>cos(correctOffset)</code> to the matching cell. The circular mean is implicit: <code>atan2(sumSin/count, sumCos/count)</code>.</td></tr>
</table>
<p>Current performance: ~48% hit rate against WallsBot. The ceiling is the input
space &mdash; only two dimensions (vPerp, distance). Adding heading change, evasion
pattern, etc. would require exponentially more cells.</p>
<h3>SNN path (inline in <code>SNNBot.nim</code>, dormant)</h3>
<p>A three-layer spiking network: 80 input neurons &rarr; 12 hidden LIF neurons
&rarr; 2 output channels (sin, cos of lead angle).</p>
<table>
<tr><th>Layer</th><th>Size</th><th>Role</th></tr>
<tr><td>Input</td><td>80</td><td>Population code: bearing (36 neurons, offset 0), velocity direction (36 neurons, offset 36), speed thermometer (8 neurons, offset 72)</td></tr>
<tr><td>Hidden</td><td>12 LIF</td><td>Each neuron integrates weighted input, leaks by factor 0.9/tick, spikes when membrane potential &ge; 0.08, resets to 0</td></tr>
<tr><td>Output</td><td>2 linear</td><td>Weighted sum of hidden spikes &rarr; sin and cos channels; decoded via <code>atan2</code> to get angle</td></tr>
</table>
<p>The LIF update per hidden neuron <code>h</code> each tick:</p>
<pre><code>V[h] = LEAK * V[h] + Σ_i (input[i] * w_ih[i,h])
if V[h] >= THRESH: spike[h] = 1.0; V[h] = 0.0
else: spike[h] = 0.0</code></pre>
<p>Learning uses the SuperSpike three-factor rule (Zenke &amp; Ganguli 2018):</p>
<pre><code>Δw = η × pre_trace × σ'(V) × error</code></pre>
<p>where <code>σ'(V)</code> is a surrogate derivative (peaks at threshold, giving
gradient direction through the non-differentiable spike). Hidden neurons receive
their error signal via fixed random feedback weights <code>bFb</code> &mdash; this
avoids the weight-transport problem of backpropagation.</p>
<div class="warning">
<strong>Why is the SNN dormant?</strong> <code>USE_RESERVOIR = true</code> in the
source. The SNN path exists and compiles, but the grid is better-tuned right now.
The SNN is the intended long-term path because it can handle arbitrary input
dimensionality &mdash; you just add more input neurons.
</div>
<!-- ═══════════════════════════════════════════════════════════════════════ -->
<h2>3. The Lead Prediction Problem</h2>
<p>What is the bot actually trying to learn?</p>
<p>At DECIDE time the enemy is at absolute bearing <strong>B</strong> (world-frame
degrees), moving with some velocity, at distance <strong>D</strong> pixels. A
bullet fired at power 2 travels at <strong>14 px/tick</strong>. By the time the
bullet covers distance D, the enemy has moved.</p>
<div class="diagram"> enemy now
★ ──────────────────→ ★ enemy later
| /
| bullet path / enemy moves laterally
| /
●─────────────────/
your bot aim here (B + offset)</div>
<p>The perpendicular component of enemy velocity relative to the bullet line is
called <code>vPerp</code>. It is computed in the DECIDE branch:</p>
<pre><code>let relVelDir = velDirDeg - absBearing # velocity angle relative to bullet line
let vPerp = velSpeed * sin(relVelDir) # lateral component only</code></pre>
<p>For a constant-velocity target the perfect offset is:</p>
<pre><code>offset = arcsin(vPerp / bulletSpeed) # ≈ 35° at full speed, perpendicular</code></pre>
<p>The grid is <strong>warm-started</strong> with exactly this formula via
<code>initLeadGrid()</code>. The SNN must <strong>learn</strong> this same
relationship from experience, without being told the formula.</p>
<div class="key-insight">
<strong>Why not just use the formula?</strong> The formula only works for
constant linear movement. A skilled opponent changes direction, jinks, circles.
The grid/SNN builds an empirical model from what actually happens, not what
physics predicts for an idealized case.
</div>
<!-- ═══════════════════════════════════════════════════════════════════════ -->
<h2>4. The Learning Signal</h2>
<p>EVALUATE is where the bot discovers how wrong its prediction was. The logic
(both grid and SNN paths share this upstream computation):</p>
<pre><code># 1. How long will the bullet travel?
travelTime = decideDist / BULLET_SPEED
# 2. Where will the enemy be at impact?
futureX = decideEnemyX + cos(decideVelDirDeg) * decideVelSpeed * travelTime
futureY = decideEnemyY + sin(decideVelDirDeg) * decideVelSpeed * travelTime
# 3. What bearing should we have aimed at?
correctAngle = directionTo(myX, myY, futureX, futureY)
# 4. Error = what we should have aimed - what the bare enemy bearing was
correctOffset = normalizeRelativeAngle(correctAngle - lastAbsBearing)</code></pre>
<p>Then the two paths diverge:</p>
<table>
<tr><th>Path</th><th>What happens with <code>correctOffset</code></th></tr>
<tr>
<td>Grid</td>
<td>Calls <code>res.learn(lastVPerp, decideDist, correctOffset)</code> which
adds sin/cos of the correct offset to the matching cell. Only learns if the
error exceeds the adaptive dead zone (0.3&ndash;0.5&deg; depending on
enemy speed).</td>
</tr>
<tr>
<td>SNN</td>
<td>Calls <code>superSpikeUpdate()</code> with the frozen spike counts and
voltages from DECIDE. The three-factor rule adjusts all weights in proportion
to pre-synaptic trace &times; surrogate derivative &times; output error.</td>
</tr>
</table>
<div class="note">
<strong>The snapshot fix in practice.</strong> Before the fix, EVALUATE used
live enemy position. The grid learned from a target that had moved since the
prediction was made &mdash; corrupted signal. Now all EVALUATE inputs come from
the <code>decide*</code> fields frozen at DECIDE time. This is why the fields
exist despite looking redundant with <code>lastEnemyX/Y</code>.
</div>
<!-- ═══════════════════════════════════════════════════════════════════════ -->
<h2>5. Check Your Understanding</h2>
<div class="quiz" id="q1">
<h3>Q1: Which phase updates the model weights?</h3>
<label><input type="radio" name="q1" data-correct="false"> DECIDE</label>
<label><input type="radio" name="q1" data-correct="false"> WAITING</label>
<label><input type="radio" name="q1" data-correct="true"> EVALUATE</label>
<label><input type="radio" name="q1" data-correct="false"> All three phases contribute</label>
<div class="feedback correct">Correct. DECIDE predicts, WAITING aims, EVALUATE measures the error and updates. The separation is intentional: you need a complete prediction-fire-outcome cycle before you have a learning signal.</div>
<div class="feedback wrong">Not quite. Only EVALUATE updates weights &mdash; it is the only phase that knows the outcome (where the enemy ended up relative to where you aimed).</div>
</div>
<div class="quiz" id="q2">
<h3>Q2: Why does the grid use circular mean (sumSin/sumCos) instead of a simple average of the offset values?</h3>
<label><input type="radio" name="q2" data-correct="false"> It is faster to compute</label>
<label><input type="radio" name="q2" data-correct="true"> Angles wrap around 360&deg; &mdash; averaging 1&deg; and 359&deg; with simple arithmetic gives 180&deg;, but the correct mean is 0&deg;</label>
<label><input type="radio" name="q2" data-correct="false"> It uses less memory per cell</label>
<label><input type="radio" name="q2" data-correct="false"> It handles negative angles better</label>
<div class="feedback correct">Correct. This is the classic circular statistics problem. By storing sin and cos components separately, the atan2(sumSin, sumCos) reconstruction always gives the correct angular mean regardless of wrap-around. A simple numeric average of degree values would give nonsense near 0&deg;/360&deg;.</div>
<div class="feedback wrong">Not quite. The key issue is angle wrap-around. A simple average of 1 and 359 gives 180 &mdash; the exact opposite direction. sumSin + sumCos encodes direction as a vector, and atan2 decodes it correctly.</div>
</div>
<div class="quiz" id="q3">
<h3>Q3: What does <code>forward()</code> return when a grid cell (and all its neighbors) has no data?</h3>
<label><input type="radio" name="q3" data-correct="false"> 0.0 degrees (aim straight at the enemy)</label>
<label><input type="radio" name="q3" data-correct="false"> The analytical warm-start value</label>
<label><input type="radio" name="q3" data-correct="true"> -999.0 (sentinel value)</label>
<label><input type="radio" name="q3" data-correct="false"> NaN</label>
<div class="feedback correct">Correct. <code>-999.0</code> is the sentinel. When DECIDE receives it, <code>bot.gridHasData</code> is set false and the bot aims at the bare enemy bearing with no lead (cold-start behavior). It is not the warm-start value &mdash; warm-start is already baked into cells with <code>count = 0.1</code>, so a cell with only warm-start data <em>will</em> return a value. The sentinel fires only when <code>totalW == 0.0</code> in the neighborhood loop &mdash; meaning all nearby cells are completely empty.</div>
<div class="feedback wrong">Look at the last line of the <code>forward()</code> proc in <code>reservoir.nim</code>: <code>if totalW == 0.0: return -999.0</code>. That sentinel is what the DECIDE branch checks with <code>if gridOffset &lt;= -999.0</code>.</div>
</div>
<!-- ═══════════════════════════════════════════════════════════════════════ -->
<h2>6. Next Steps</h2>
<p>You now understand the full control loop: DECIDE &rarr; WAITING &rarr;
EVALUATE, how the grid accumulator stores and retrieves lead angles, and how the
SNN path is structured but dormant.</p>
<p>In the next lesson we will dive into the SNN path in detail: LIF neuron dynamics
tick by tick, why the surrogate derivative is needed (the spike is
non-differentiable), what SuperSpike's three factors are and why each one is
there, and what it would take to activate <code>USE_RESERVOIR = false</code>
without the hit rate collapsing.</p>
<div class="note">
Questions? Ask your agent &mdash; it is your teacher and can clarify anything
above, show you the exact line in the source, or run a test to check a
hypothesis.
</div>
<div class="nav-footer">
Lesson 1 / Architecture Overview &mdash;
source: <code>SNNBot_garage/src/SNNBot.nim</code>,
<code>SNNBot_garage/src/reservoir.nim</code>
</div>
<script src="../assets/quiz.js"></script>
</body>
</html>