milestone(SNNBot): 48% hit rate, grid accumulator with warm-start + frozen DECIDE snapshot — pausing for BNNBot exploration
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,267 @@
|
||||
<!DOCTYPE html>
|
||||
<html lang="en">
|
||||
<head>
|
||||
<meta charset="UTF-8">
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1.0">
|
||||
<title>Lesson 1: Your SNN Bot — Architecture Overview</title>
|
||||
<link rel="stylesheet" href="../assets/style.css">
|
||||
</head>
|
||||
<body>
|
||||
|
||||
<h1>Your SNN Bot: Architecture Overview</h1>
|
||||
<p style="font-family: var(--font-sans); font-size: 0.85rem; color: #888;">Lesson 1 of N — <code>SNNBot_garage/src/SNNBot.nim</code></p>
|
||||
|
||||
<p>This lesson walks through your own code. No abstract theory first — we start
|
||||
from what you already built and explain <em>why</em> it is shaped the way it is.</p>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>1. The State Machine</h2>
|
||||
|
||||
<p>Your bot does not fire every tick. It cycles through three phases, each with a
|
||||
distinct job. The cycle is encoded in the <code>Phase</code> enum and the
|
||||
<code>case bot.phase</code> block inside <code>run()</code>.</p>
|
||||
|
||||
<div class="diagram"> ┌─────────┐ gun aimed ┌─────────┐ bullet fired ┌──────────┐
|
||||
│ DECIDE │ ──────────────→ │ WAITING │ ──────────────→ │ EVALUATE │
|
||||
│ │ │ │ │ │
|
||||
│ predict │ │ aim+fire│ │ learn │
|
||||
└─────────┘ └─────────┘ └──────────┘
|
||||
↑ │
|
||||
└───────────────────────────────────────────────────────────┘</div>
|
||||
|
||||
<h3>DECIDE</h3>
|
||||
<p>Takes a snapshot of the enemy state right now — position, velocity direction,
|
||||
speed, distance — and asks the model: <em>what lead angle should I aim at?</em>
|
||||
The answer is stored in <code>bot.targetAngle</code>. Crucially, all input values
|
||||
are frozen into <code>decide*</code> fields (<code>decideEnemyX</code>,
|
||||
<code>decideVelDirDeg</code>, <code>decideDist</code>, etc.) so that the learning
|
||||
step later has a consistent ground truth. Phase advances to WAITING.</p>
|
||||
|
||||
<h3>WAITING</h3>
|
||||
<p>Calls <code>aimTo(bot.targetAngle, gunDir)</code> every tick, which sets the gun
|
||||
turn rate toward the target. When the aiming error drops below 2° <em>and</em>
|
||||
the bot has enough energy, it fires and advances to EVALUATE. No learning happens
|
||||
here — this phase is purely mechanical.</p>
|
||||
|
||||
<h3>EVALUATE</h3>
|
||||
<p>Uses the frozen DECIDE-time snapshot to compute where the enemy <em>will</em> be
|
||||
when the bullet arrives, derives the correct lead angle, and feeds the error into
|
||||
the model (grid or SNN). Then immediately jumps back to DECIDE for the next shot.
|
||||
This is the only phase where weights change.</p>
|
||||
|
||||
<div class="key-insight">
|
||||
<strong>Why freeze the snapshot?</strong> Between DECIDE and EVALUATE the bot
|
||||
scans the arena and accumulates new sensor data. If EVALUATE read live values,
|
||||
the "correct answer" would be computed from different inputs than the prediction
|
||||
was made from — like grading a test with a different question than the one
|
||||
asked. The <code>decide*</code> fields fix this.
|
||||
</div>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>2. Two Brains — Grid vs SNN</h2>
|
||||
|
||||
<p>The compile-time constant <code>USE_RESERVOIR</code> (line 46 of
|
||||
<code>SNNBot.nim</code>) selects which brain runs. Both brains expose the same
|
||||
interface to the state machine: <em>given inputs, return a lead angle offset</em>;
|
||||
<em>given error, update yourself</em>.</p>
|
||||
|
||||
<h3>Grid accumulator (<code>reservoir.nim</code>, currently active)</h3>
|
||||
|
||||
<p>A 17×8 = 136-cell lookup table. Rows are perpendicular velocity bins
|
||||
(integer values −8 to +8), columns are distance bands (125 px each). Each
|
||||
cell stores a <strong>circular mean</strong> as three floats: <code>sumSin</code>,
|
||||
<code>sumCos</code>, <code>count</code>.</p>
|
||||
|
||||
<table>
|
||||
<tr><th>Proc</th><th>What it does</th></tr>
|
||||
<tr><td><code>initLeadGrid()</code></td><td>Warm-starts every cell with the analytical value <code>arcsin(vPerp / 14.0)</code> so the bot is not blind on round 1.</td></tr>
|
||||
<tr><td><code>forward(vPerp, dist)</code></td><td>Interpolates over the 3×3 neighborhood of cells (weights 1 / 0.5 / 0.25 for direct / cardinal / diagonal). Returns <code>-999.0</code> when no cell in the neighborhood has data.</td></tr>
|
||||
<tr><td><code>learn(vPerp, dist, correctOffset)</code></td><td>Adds <code>sin(correctOffset)</code> and <code>cos(correctOffset)</code> to the matching cell. The circular mean is implicit: <code>atan2(sumSin/count, sumCos/count)</code>.</td></tr>
|
||||
</table>
|
||||
|
||||
<p>Current performance: ~48% hit rate against WallsBot. The ceiling is the input
|
||||
space — only two dimensions (vPerp, distance). Adding heading change, evasion
|
||||
pattern, etc. would require exponentially more cells.</p>
|
||||
|
||||
<h3>SNN path (inline in <code>SNNBot.nim</code>, dormant)</h3>
|
||||
|
||||
<p>A three-layer spiking network: 80 input neurons → 12 hidden LIF neurons
|
||||
→ 2 output channels (sin, cos of lead angle).</p>
|
||||
|
||||
<table>
|
||||
<tr><th>Layer</th><th>Size</th><th>Role</th></tr>
|
||||
<tr><td>Input</td><td>80</td><td>Population code: bearing (36 neurons, offset 0), velocity direction (36 neurons, offset 36), speed thermometer (8 neurons, offset 72)</td></tr>
|
||||
<tr><td>Hidden</td><td>12 LIF</td><td>Each neuron integrates weighted input, leaks by factor 0.9/tick, spikes when membrane potential ≥ 0.08, resets to 0</td></tr>
|
||||
<tr><td>Output</td><td>2 linear</td><td>Weighted sum of hidden spikes → sin and cos channels; decoded via <code>atan2</code> to get angle</td></tr>
|
||||
</table>
|
||||
|
||||
<p>The LIF update per hidden neuron <code>h</code> each tick:</p>
|
||||
<pre><code>V[h] = LEAK * V[h] + Σ_i (input[i] * w_ih[i,h])
|
||||
if V[h] >= THRESH: spike[h] = 1.0; V[h] = 0.0
|
||||
else: spike[h] = 0.0</code></pre>
|
||||
|
||||
<p>Learning uses the SuperSpike three-factor rule (Zenke & Ganguli 2018):</p>
|
||||
<pre><code>Δw = η × pre_trace × σ'(V) × error</code></pre>
|
||||
<p>where <code>σ'(V)</code> is a surrogate derivative (peaks at threshold, giving
|
||||
gradient direction through the non-differentiable spike). Hidden neurons receive
|
||||
their error signal via fixed random feedback weights <code>bFb</code> — this
|
||||
avoids the weight-transport problem of backpropagation.</p>
|
||||
|
||||
<div class="warning">
|
||||
<strong>Why is the SNN dormant?</strong> <code>USE_RESERVOIR = true</code> in the
|
||||
source. The SNN path exists and compiles, but the grid is better-tuned right now.
|
||||
The SNN is the intended long-term path because it can handle arbitrary input
|
||||
dimensionality — you just add more input neurons.
|
||||
</div>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>3. The Lead Prediction Problem</h2>
|
||||
|
||||
<p>What is the bot actually trying to learn?</p>
|
||||
|
||||
<p>At DECIDE time the enemy is at absolute bearing <strong>B</strong> (world-frame
|
||||
degrees), moving with some velocity, at distance <strong>D</strong> pixels. A
|
||||
bullet fired at power 2 travels at <strong>14 px/tick</strong>. By the time the
|
||||
bullet covers distance D, the enemy has moved.</p>
|
||||
|
||||
<div class="diagram"> enemy now
|
||||
★ ──────────────────→ ★ enemy later
|
||||
| /
|
||||
| bullet path / enemy moves laterally
|
||||
| /
|
||||
●─────────────────/
|
||||
your bot aim here (B + offset)</div>
|
||||
|
||||
<p>The perpendicular component of enemy velocity relative to the bullet line is
|
||||
called <code>vPerp</code>. It is computed in the DECIDE branch:</p>
|
||||
|
||||
<pre><code>let relVelDir = velDirDeg - absBearing # velocity angle relative to bullet line
|
||||
let vPerp = velSpeed * sin(relVelDir) # lateral component only</code></pre>
|
||||
|
||||
<p>For a constant-velocity target the perfect offset is:</p>
|
||||
<pre><code>offset = arcsin(vPerp / bulletSpeed) # ≈ 35° at full speed, perpendicular</code></pre>
|
||||
|
||||
<p>The grid is <strong>warm-started</strong> with exactly this formula via
|
||||
<code>initLeadGrid()</code>. The SNN must <strong>learn</strong> this same
|
||||
relationship from experience, without being told the formula.</p>
|
||||
|
||||
<div class="key-insight">
|
||||
<strong>Why not just use the formula?</strong> The formula only works for
|
||||
constant linear movement. A skilled opponent changes direction, jinks, circles.
|
||||
The grid/SNN builds an empirical model from what actually happens, not what
|
||||
physics predicts for an idealized case.
|
||||
</div>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>4. The Learning Signal</h2>
|
||||
|
||||
<p>EVALUATE is where the bot discovers how wrong its prediction was. The logic
|
||||
(both grid and SNN paths share this upstream computation):</p>
|
||||
|
||||
<pre><code># 1. How long will the bullet travel?
|
||||
travelTime = decideDist / BULLET_SPEED
|
||||
|
||||
# 2. Where will the enemy be at impact?
|
||||
futureX = decideEnemyX + cos(decideVelDirDeg) * decideVelSpeed * travelTime
|
||||
futureY = decideEnemyY + sin(decideVelDirDeg) * decideVelSpeed * travelTime
|
||||
|
||||
# 3. What bearing should we have aimed at?
|
||||
correctAngle = directionTo(myX, myY, futureX, futureY)
|
||||
|
||||
# 4. Error = what we should have aimed - what the bare enemy bearing was
|
||||
correctOffset = normalizeRelativeAngle(correctAngle - lastAbsBearing)</code></pre>
|
||||
|
||||
<p>Then the two paths diverge:</p>
|
||||
|
||||
<table>
|
||||
<tr><th>Path</th><th>What happens with <code>correctOffset</code></th></tr>
|
||||
<tr>
|
||||
<td>Grid</td>
|
||||
<td>Calls <code>res.learn(lastVPerp, decideDist, correctOffset)</code> which
|
||||
adds sin/cos of the correct offset to the matching cell. Only learns if the
|
||||
error exceeds the adaptive dead zone (0.3–0.5° depending on
|
||||
enemy speed).</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>SNN</td>
|
||||
<td>Calls <code>superSpikeUpdate()</code> with the frozen spike counts and
|
||||
voltages from DECIDE. The three-factor rule adjusts all weights in proportion
|
||||
to pre-synaptic trace × surrogate derivative × output error.</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
<div class="note">
|
||||
<strong>The snapshot fix in practice.</strong> Before the fix, EVALUATE used
|
||||
live enemy position. The grid learned from a target that had moved since the
|
||||
prediction was made — corrupted signal. Now all EVALUATE inputs come from
|
||||
the <code>decide*</code> fields frozen at DECIDE time. This is why the fields
|
||||
exist despite looking redundant with <code>lastEnemyX/Y</code>.
|
||||
</div>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>5. Check Your Understanding</h2>
|
||||
|
||||
<div class="quiz" id="q1">
|
||||
<h3>Q1: Which phase updates the model weights?</h3>
|
||||
<label><input type="radio" name="q1" data-correct="false"> DECIDE</label>
|
||||
<label><input type="radio" name="q1" data-correct="false"> WAITING</label>
|
||||
<label><input type="radio" name="q1" data-correct="true"> EVALUATE</label>
|
||||
<label><input type="radio" name="q1" data-correct="false"> All three phases contribute</label>
|
||||
<div class="feedback correct">Correct. DECIDE predicts, WAITING aims, EVALUATE measures the error and updates. The separation is intentional: you need a complete prediction-fire-outcome cycle before you have a learning signal.</div>
|
||||
<div class="feedback wrong">Not quite. Only EVALUATE updates weights — it is the only phase that knows the outcome (where the enemy ended up relative to where you aimed).</div>
|
||||
</div>
|
||||
|
||||
<div class="quiz" id="q2">
|
||||
<h3>Q2: Why does the grid use circular mean (sumSin/sumCos) instead of a simple average of the offset values?</h3>
|
||||
<label><input type="radio" name="q2" data-correct="false"> It is faster to compute</label>
|
||||
<label><input type="radio" name="q2" data-correct="true"> Angles wrap around 360° — averaging 1° and 359° with simple arithmetic gives 180°, but the correct mean is 0°</label>
|
||||
<label><input type="radio" name="q2" data-correct="false"> It uses less memory per cell</label>
|
||||
<label><input type="radio" name="q2" data-correct="false"> It handles negative angles better</label>
|
||||
<div class="feedback correct">Correct. This is the classic circular statistics problem. By storing sin and cos components separately, the atan2(sumSin, sumCos) reconstruction always gives the correct angular mean regardless of wrap-around. A simple numeric average of degree values would give nonsense near 0°/360°.</div>
|
||||
<div class="feedback wrong">Not quite. The key issue is angle wrap-around. A simple average of 1 and 359 gives 180 — the exact opposite direction. sumSin + sumCos encodes direction as a vector, and atan2 decodes it correctly.</div>
|
||||
</div>
|
||||
|
||||
<div class="quiz" id="q3">
|
||||
<h3>Q3: What does <code>forward()</code> return when a grid cell (and all its neighbors) has no data?</h3>
|
||||
<label><input type="radio" name="q3" data-correct="false"> 0.0 degrees (aim straight at the enemy)</label>
|
||||
<label><input type="radio" name="q3" data-correct="false"> The analytical warm-start value</label>
|
||||
<label><input type="radio" name="q3" data-correct="true"> -999.0 (sentinel value)</label>
|
||||
<label><input type="radio" name="q3" data-correct="false"> NaN</label>
|
||||
<div class="feedback correct">Correct. <code>-999.0</code> is the sentinel. When DECIDE receives it, <code>bot.gridHasData</code> is set false and the bot aims at the bare enemy bearing with no lead (cold-start behavior). It is not the warm-start value — warm-start is already baked into cells with <code>count = 0.1</code>, so a cell with only warm-start data <em>will</em> return a value. The sentinel fires only when <code>totalW == 0.0</code> in the neighborhood loop — meaning all nearby cells are completely empty.</div>
|
||||
<div class="feedback wrong">Look at the last line of the <code>forward()</code> proc in <code>reservoir.nim</code>: <code>if totalW == 0.0: return -999.0</code>. That sentinel is what the DECIDE branch checks with <code>if gridOffset <= -999.0</code>.</div>
|
||||
</div>
|
||||
|
||||
|
||||
<!-- ═══════════════════════════════════════════════════════════════════════ -->
|
||||
<h2>6. Next Steps</h2>
|
||||
|
||||
<p>You now understand the full control loop: DECIDE → WAITING →
|
||||
EVALUATE, how the grid accumulator stores and retrieves lead angles, and how the
|
||||
SNN path is structured but dormant.</p>
|
||||
|
||||
<p>In the next lesson we will dive into the SNN path in detail: LIF neuron dynamics
|
||||
tick by tick, why the surrogate derivative is needed (the spike is
|
||||
non-differentiable), what SuperSpike's three factors are and why each one is
|
||||
there, and what it would take to activate <code>USE_RESERVOIR = false</code>
|
||||
without the hit rate collapsing.</p>
|
||||
|
||||
<div class="note">
|
||||
Questions? Ask your agent — it is your teacher and can clarify anything
|
||||
above, show you the exact line in the source, or run a test to check a
|
||||
hypothesis.
|
||||
</div>
|
||||
|
||||
<div class="nav-footer">
|
||||
Lesson 1 / Architecture Overview —
|
||||
source: <code>SNNBot_garage/src/SNNBot.nim</code>,
|
||||
<code>SNNBot_garage/src/reservoir.nim</code>
|
||||
</div>
|
||||
|
||||
<script src="../assets/quiz.js"></script>
|
||||
</body>
|
||||
</html>
|
||||
Reference in New Issue
Block a user