Your SNN Bot: Architecture Overview

Lesson 1 of N — SNNBot_garage/src/SNNBot.nim

This lesson walks through your own code. No abstract theory first — we start from what you already built and explain why it is shaped the way it is.

1. The State Machine

Your bot does not fire every tick. It cycles through three phases, each with a distinct job. The cycle is encoded in the Phase enum and the case bot.phase block inside run().

┌─────────┐ gun aimed ┌─────────┐ bullet fired ┌──────────┐ │ DECIDE │ ──────────────→ │ WAITING │ ──────────────→ │ EVALUATE │ │ │ │ │ │ │ │ predict │ │ aim+fire│ │ learn │ └─────────┘ └─────────┘ └──────────┘ ↑ │ └───────────────────────────────────────────────────────────┘

DECIDE

Takes a snapshot of the enemy state right now — position, velocity direction, speed, distance — and asks the model: what lead angle should I aim at? The answer is stored in bot.targetAngle. Crucially, all input values are frozen into decide* fields (decideEnemyX, decideVelDirDeg, decideDist, etc.) so that the learning step later has a consistent ground truth. Phase advances to WAITING.

WAITING

Calls aimTo(bot.targetAngle, gunDir) every tick, which sets the gun turn rate toward the target. When the aiming error drops below 2° and the bot has enough energy, it fires and advances to EVALUATE. No learning happens here — this phase is purely mechanical.

EVALUATE

Uses the frozen DECIDE-time snapshot to compute where the enemy will be when the bullet arrives, derives the correct lead angle, and feeds the error into the model (grid or SNN). Then immediately jumps back to DECIDE for the next shot. This is the only phase where weights change.

Why freeze the snapshot? Between DECIDE and EVALUATE the bot scans the arena and accumulates new sensor data. If EVALUATE read live values, the "correct answer" would be computed from different inputs than the prediction was made from — like grading a test with a different question than the one asked. The decide* fields fix this.

2. Two Brains — Grid vs SNN

The compile-time constant USE_RESERVOIR (line 46 of SNNBot.nim) selects which brain runs. Both brains expose the same interface to the state machine: given inputs, return a lead angle offset; given error, update yourself.

Grid accumulator (reservoir.nim, currently active)

A 17×8 = 136-cell lookup table. Rows are perpendicular velocity bins (integer values −8 to +8), columns are distance bands (125 px each). Each cell stores a circular mean as three floats: sumSin, sumCos, count.

ProcWhat it does
initLeadGrid()Warm-starts every cell with the analytical value arcsin(vPerp / 14.0) so the bot is not blind on round 1.
forward(vPerp, dist)Interpolates over the 3×3 neighborhood of cells (weights 1 / 0.5 / 0.25 for direct / cardinal / diagonal). Returns -999.0 when no cell in the neighborhood has data.
learn(vPerp, dist, correctOffset)Adds sin(correctOffset) and cos(correctOffset) to the matching cell. The circular mean is implicit: atan2(sumSin/count, sumCos/count).

Current performance: ~48% hit rate against WallsBot. The ceiling is the input space — only two dimensions (vPerp, distance). Adding heading change, evasion pattern, etc. would require exponentially more cells.

SNN path (inline in SNNBot.nim, dormant)

A three-layer spiking network: 80 input neurons → 12 hidden LIF neurons → 2 output channels (sin, cos of lead angle).

LayerSizeRole
Input80Population code: bearing (36 neurons, offset 0), velocity direction (36 neurons, offset 36), speed thermometer (8 neurons, offset 72)
Hidden12 LIFEach neuron integrates weighted input, leaks by factor 0.9/tick, spikes when membrane potential ≥ 0.08, resets to 0
Output2 linearWeighted sum of hidden spikes → sin and cos channels; decoded via atan2 to get angle

The LIF update per hidden neuron h each tick:

V[h] = LEAK * V[h] + Σ_i (input[i] * w_ih[i,h])
if V[h] >= THRESH: spike[h] = 1.0; V[h] = 0.0
else:              spike[h] = 0.0

Learning uses the SuperSpike three-factor rule (Zenke & Ganguli 2018):

Δw = η × pre_trace × σ'(V) × error

where σ'(V) is a surrogate derivative (peaks at threshold, giving gradient direction through the non-differentiable spike). Hidden neurons receive their error signal via fixed random feedback weights bFb — this avoids the weight-transport problem of backpropagation.

Why is the SNN dormant? USE_RESERVOIR = true in the source. The SNN path exists and compiles, but the grid is better-tuned right now. The SNN is the intended long-term path because it can handle arbitrary input dimensionality — you just add more input neurons.

3. The Lead Prediction Problem

What is the bot actually trying to learn?

At DECIDE time the enemy is at absolute bearing B (world-frame degrees), moving with some velocity, at distance D pixels. A bullet fired at power 2 travels at 14 px/tick. By the time the bullet covers distance D, the enemy has moved.

enemy now ★ ──────────────────→ ★ enemy later | / | bullet path / enemy moves laterally | / ●─────────────────/ your bot aim here (B + offset)

The perpendicular component of enemy velocity relative to the bullet line is called vPerp. It is computed in the DECIDE branch:

let relVelDir = velDirDeg - absBearing   # velocity angle relative to bullet line
let vPerp = velSpeed * sin(relVelDir)    # lateral component only

For a constant-velocity target the perfect offset is:

offset = arcsin(vPerp / bulletSpeed)     # ≈ 35° at full speed, perpendicular

The grid is warm-started with exactly this formula via initLeadGrid(). The SNN must learn this same relationship from experience, without being told the formula.

Why not just use the formula? The formula only works for constant linear movement. A skilled opponent changes direction, jinks, circles. The grid/SNN builds an empirical model from what actually happens, not what physics predicts for an idealized case.

4. The Learning Signal

EVALUATE is where the bot discovers how wrong its prediction was. The logic (both grid and SNN paths share this upstream computation):

# 1. How long will the bullet travel?
travelTime = decideDist / BULLET_SPEED

# 2. Where will the enemy be at impact?
futureX = decideEnemyX + cos(decideVelDirDeg) * decideVelSpeed * travelTime
futureY = decideEnemyY + sin(decideVelDirDeg) * decideVelSpeed * travelTime

# 3. What bearing should we have aimed at?
correctAngle = directionTo(myX, myY, futureX, futureY)

# 4. Error = what we should have aimed - what the bare enemy bearing was
correctOffset = normalizeRelativeAngle(correctAngle - lastAbsBearing)

Then the two paths diverge:

PathWhat happens with correctOffset
Grid Calls res.learn(lastVPerp, decideDist, correctOffset) which adds sin/cos of the correct offset to the matching cell. Only learns if the error exceeds the adaptive dead zone (0.3–0.5° depending on enemy speed).
SNN Calls superSpikeUpdate() with the frozen spike counts and voltages from DECIDE. The three-factor rule adjusts all weights in proportion to pre-synaptic trace × surrogate derivative × output error.
The snapshot fix in practice. Before the fix, EVALUATE used live enemy position. The grid learned from a target that had moved since the prediction was made — corrupted signal. Now all EVALUATE inputs come from the decide* fields frozen at DECIDE time. This is why the fields exist despite looking redundant with lastEnemyX/Y.

5. Check Your Understanding

Q1: Which phase updates the model weights?

Correct. DECIDE predicts, WAITING aims, EVALUATE measures the error and updates. The separation is intentional: you need a complete prediction-fire-outcome cycle before you have a learning signal.
Not quite. Only EVALUATE updates weights — it is the only phase that knows the outcome (where the enemy ended up relative to where you aimed).

Q2: Why does the grid use circular mean (sumSin/sumCos) instead of a simple average of the offset values?

Correct. This is the classic circular statistics problem. By storing sin and cos components separately, the atan2(sumSin, sumCos) reconstruction always gives the correct angular mean regardless of wrap-around. A simple numeric average of degree values would give nonsense near 0°/360°.
Not quite. The key issue is angle wrap-around. A simple average of 1 and 359 gives 180 — the exact opposite direction. sumSin + sumCos encodes direction as a vector, and atan2 decodes it correctly.

Q3: What does forward() return when a grid cell (and all its neighbors) has no data?

Correct. -999.0 is the sentinel. When DECIDE receives it, bot.gridHasData is set false and the bot aims at the bare enemy bearing with no lead (cold-start behavior). It is not the warm-start value — warm-start is already baked into cells with count = 0.1, so a cell with only warm-start data will return a value. The sentinel fires only when totalW == 0.0 in the neighborhood loop — meaning all nearby cells are completely empty.
Look at the last line of the forward() proc in reservoir.nim: if totalW == 0.0: return -999.0. That sentinel is what the DECIDE branch checks with if gridOffset <= -999.0.

6. Next Steps

You now understand the full control loop: DECIDE → WAITING → EVALUATE, how the grid accumulator stores and retrieves lead angles, and how the SNN path is structured but dormant.

In the next lesson we will dive into the SNN path in detail: LIF neuron dynamics tick by tick, why the surrogate derivative is needed (the spike is non-differentiable), what SuperSpike's three factors are and why each one is there, and what it would take to activate USE_RESERVOIR = false without the hit rate collapsing.

Questions? Ask your agent — it is your teacher and can clarify anything above, show you the exact line in the source, or run a test to check a hypothesis.