diff --git a/MISSION.md b/MISSION.md new file mode 100644 index 0000000..42ba9e9 --- /dev/null +++ b/MISSION.md @@ -0,0 +1,18 @@ +# Mission + +Build a spiking neural network (SNN) bot for Tank Royale that learns to aim at moving targets using binary operations and biologically-inspired learning rules — not analytical formulas. The system must generalize from predictable movers to evasive opponents. + +## Why + +- Classical aiming (analytical lead formulas) can't handle unpredictable movement +- SNNs offer energy-efficient, event-driven learning suited to real-time control +- Binary operations (Hamming distance, popcount) are fast and hardware-friendly +- The long-term goal is a bot that improves through experience, not programming + +## Current State + +Two implementations exist in `SNNBot_garage/src/`: +- **Grid accumulator** (`reservoir.nim`, USE_RESERVOIR=true): Lookup table mapping (vPerp, distance) → lead offset. Works (~48% hit rate) but can't scale to more inputs. +- **SNN** (`SNNBot.nim`, USE_RESERVOIR=false): LIF spiking network with SuperSpike-inspired learning. Dormant — not currently active. + +The grid hit its ceiling. The SNN path is the intended future. diff --git a/NOTES.md b/NOTES.md new file mode 100644 index 0000000..e73c0e7 --- /dev/null +++ b/NOTES.md @@ -0,0 +1,8 @@ +# Teaching Notes + +- User prefers learning by understanding their own code, not abstract theory +- User wants binary operations / binary vectors as the core representation +- User values deep root-cause thinking over quick patches +- User reads Italian natively, English fluently +- Previous session history: tried WTA bins, exemplar ring buffer, grid accumulator — each hit different ceilings +- User explicitly rejected analytical formulas: "we need to have a mind projected to harder movements to predict" diff --git a/RESOURCES.md b/RESOURCES.md new file mode 100644 index 0000000..428be72 --- /dev/null +++ b/RESOURCES.md @@ -0,0 +1,19 @@ +# Resources + +## Primary Sources + +- [SuperSpike paper (Zenke & Ganguli 2018)](https://doi.org/10.1162/neco_a_01086) — The learning rule implemented in the SNN path. Three-factor rule: eligibility trace × error signal. +- [Tank Royale API docs](https://robocode-dev.github.io/tank-royale/) — Bot API, game physics, event model. +- [Leaky Integrate-and-Fire model](https://neuronaldynamics.epfl.ch/online/Ch1.S3.html) — The neuron model used (EPFL textbook, free online). + +## Codebase + +- `SNNBot_garage/src/SNNBot.nim` — Main bot with SNN and reservoir paths +- `SNNBot_garage/src/reservoir.nim` — Grid accumulator (current active path) +- `SNNBot_garage/tests/test_bullet_economy.nim` — Automated battle test + +## To Explore + +- Hebbian learning / STDP for binary spikes +- Hyperdimensional computing for control tasks +- `SNNBot_garage/research/binary-snn-learning.md` — Research notes (if exists) diff --git a/SNNBot_garage/SNNBot b/SNNBot_garage/SNNBot index a2bfc31..04d649a 100755 Binary files a/SNNBot_garage/SNNBot and b/SNNBot_garage/SNNBot differ diff --git a/SNNBot_garage/tests/out/test_bullet_economy b/SNNBot_garage/tests/out/test_bullet_economy new file mode 100755 index 0000000..2e61c82 Binary files /dev/null and b/SNNBot_garage/tests/out/test_bullet_economy differ diff --git a/assets/quiz.js b/assets/quiz.js new file mode 100644 index 0000000..2ee92c1 --- /dev/null +++ b/assets/quiz.js @@ -0,0 +1,21 @@ +// Minimal quiz widget — shared across lessons +document.addEventListener('DOMContentLoaded', () => { + document.querySelectorAll('.quiz').forEach(quiz => { + const radios = quiz.querySelectorAll('input[type="radio"]'); + const feedbackCorrect = quiz.querySelector('.feedback.correct'); + const feedbackWrong = quiz.querySelector('.feedback.wrong'); + + radios.forEach(radio => { + radio.addEventListener('change', () => { + if (feedbackCorrect) feedbackCorrect.style.display = 'none'; + if (feedbackWrong) feedbackWrong.style.display = 'none'; + + if (radio.dataset.correct === 'true') { + if (feedbackCorrect) feedbackCorrect.style.display = 'block'; + } else { + if (feedbackWrong) feedbackWrong.style.display = 'block'; + } + }); + }); + }); +}); diff --git a/assets/style.css b/assets/style.css new file mode 100644 index 0000000..c0000c9 --- /dev/null +++ b/assets/style.css @@ -0,0 +1,158 @@ +/* Teaching workspace — shared lesson stylesheet */ +:root { + --fg: #2d2d2d; + --bg: #fffff8; + --accent: #e63946; + --code-bg: #f5f5f0; + --border: #ddd; + --link: #457b9d; + --success: #2a9d8f; + --warning: #e9c46a; + --font-body: 'Palatino Linotype', 'Book Antiqua', Palatino, serif; + --font-code: 'Fira Code', 'Source Code Pro', 'Consolas', monospace; + --font-sans: -apple-system, BlinkMacSystemFont, 'Segoe UI', sans-serif; +} + +* { box-sizing: border-box; margin: 0; padding: 0; } + +body { + font-family: var(--font-body); + color: var(--fg); + background: var(--bg); + max-width: 740px; + margin: 0 auto; + padding: 2rem 1.5rem 4rem; + line-height: 1.7; + font-size: 18px; +} + +h1 { font-size: 2rem; margin: 2rem 0 1rem; font-weight: 400; letter-spacing: -0.02em; } +h2 { font-size: 1.4rem; margin: 2rem 0 0.8rem; font-weight: 600; color: var(--accent); border-bottom: 1px solid var(--border); padding-bottom: 0.3rem; } +h3 { font-size: 1.1rem; margin: 1.5rem 0 0.5rem; font-weight: 600; } + +p { margin: 0.8rem 0; } +a { color: var(--link); text-decoration: none; border-bottom: 1px solid transparent; } +a:hover { border-bottom-color: var(--link); } + +code { + font-family: var(--font-code); + font-size: 0.85em; + background: var(--code-bg); + padding: 0.15em 0.4em; + border-radius: 3px; +} + +pre { + background: var(--code-bg); + border-left: 3px solid var(--accent); + padding: 1rem 1.2rem; + margin: 1rem 0; + overflow-x: auto; + line-height: 1.5; + font-size: 0.85rem; +} +pre code { background: none; padding: 0; } + +.note { + background: #f0f7ff; + border-left: 3px solid var(--link); + padding: 0.8rem 1rem; + margin: 1rem 0; + font-size: 0.95rem; +} + +.warning { + background: #fff8e1; + border-left: 3px solid var(--warning); + padding: 0.8rem 1rem; + margin: 1rem 0; + font-size: 0.95rem; +} + +.key-insight { + background: #f0fff4; + border-left: 3px solid var(--success); + padding: 0.8rem 1rem; + margin: 1rem 0; + font-size: 0.95rem; +} + +figure { + margin: 1.5rem 0; + text-align: center; +} +figcaption { + font-size: 0.85rem; + color: #666; + margin-top: 0.5rem; + font-style: italic; +} + +table { + border-collapse: collapse; + width: 100%; + margin: 1rem 0; + font-size: 0.9rem; +} +th, td { + border: 1px solid var(--border); + padding: 0.5rem 0.8rem; + text-align: left; +} +th { background: var(--code-bg); font-weight: 600; } + +.diagram { + font-family: var(--font-code); + font-size: 0.8rem; + line-height: 1.4; + white-space: pre; + background: var(--code-bg); + padding: 1.2rem; + margin: 1rem 0; + border-radius: 4px; + overflow-x: auto; +} + +.quiz { + background: var(--code-bg); + border: 1px solid var(--border); + border-radius: 6px; + padding: 1.2rem; + margin: 1.5rem 0; +} +.quiz h3 { margin-top: 0; color: var(--accent); } +.quiz label { + display: block; + padding: 0.4rem 0.6rem; + margin: 0.3rem 0; + border-radius: 4px; + cursor: pointer; + font-family: var(--font-sans); + font-size: 0.9rem; +} +.quiz label:hover { background: #e8e8e0; } +.quiz input[type="radio"] { margin-right: 0.5rem; } +.quiz .feedback { + display: none; + margin-top: 0.8rem; + padding: 0.6rem; + border-radius: 4px; + font-size: 0.9rem; +} +.quiz .feedback.correct { background: #d4edda; display: block; } +.quiz .feedback.wrong { background: #f8d7da; display: block; } + +.nav-footer { + margin-top: 3rem; + padding-top: 1rem; + border-top: 1px solid var(--border); + font-family: var(--font-sans); + font-size: 0.85rem; + color: #666; +} + +@media print { + body { max-width: none; padding: 1cm; font-size: 11pt; } + .quiz { page-break-inside: avoid; } + pre { font-size: 9pt; } +} diff --git a/lessons/0001-snnbot-architecture-overview.html b/lessons/0001-snnbot-architecture-overview.html new file mode 100644 index 0000000..149ea41 --- /dev/null +++ b/lessons/0001-snnbot-architecture-overview.html @@ -0,0 +1,267 @@ + + +
+ + +Lesson 1 of N — SNNBot_garage/src/SNNBot.nim
This lesson walks through your own code. No abstract theory first — we start +from what you already built and explain why it is shaped the way it is.
+ + + +Your bot does not fire every tick. It cycles through three phases, each with a
+distinct job. The cycle is encoded in the Phase enum and the
+case bot.phase block inside run().
Takes a snapshot of the enemy state right now — position, velocity direction,
+speed, distance — and asks the model: what lead angle should I aim at?
+The answer is stored in bot.targetAngle. Crucially, all input values
+are frozen into decide* fields (decideEnemyX,
+decideVelDirDeg, decideDist, etc.) so that the learning
+step later has a consistent ground truth. Phase advances to WAITING.
Calls aimTo(bot.targetAngle, gunDir) every tick, which sets the gun
+turn rate toward the target. When the aiming error drops below 2° and
+the bot has enough energy, it fires and advances to EVALUATE. No learning happens
+here — this phase is purely mechanical.
Uses the frozen DECIDE-time snapshot to compute where the enemy will be +when the bullet arrives, derives the correct lead angle, and feeds the error into +the model (grid or SNN). Then immediately jumps back to DECIDE for the next shot. +This is the only phase where weights change.
+ +decide* fields fix this.
+The compile-time constant USE_RESERVOIR (line 46 of
+SNNBot.nim) selects which brain runs. Both brains expose the same
+interface to the state machine: given inputs, return a lead angle offset;
+given error, update yourself.
reservoir.nim, currently active)A 17×8 = 136-cell lookup table. Rows are perpendicular velocity bins
+(integer values −8 to +8), columns are distance bands (125 px each). Each
+cell stores a circular mean as three floats: sumSin,
+sumCos, count.
| Proc | What it does |
|---|---|
initLeadGrid() | Warm-starts every cell with the analytical value arcsin(vPerp / 14.0) so the bot is not blind on round 1. |
forward(vPerp, dist) | Interpolates over the 3×3 neighborhood of cells (weights 1 / 0.5 / 0.25 for direct / cardinal / diagonal). Returns -999.0 when no cell in the neighborhood has data. |
learn(vPerp, dist, correctOffset) | Adds sin(correctOffset) and cos(correctOffset) to the matching cell. The circular mean is implicit: atan2(sumSin/count, sumCos/count). |
Current performance: ~48% hit rate against WallsBot. The ceiling is the input +space — only two dimensions (vPerp, distance). Adding heading change, evasion +pattern, etc. would require exponentially more cells.
+ +SNNBot.nim, dormant)A three-layer spiking network: 80 input neurons → 12 hidden LIF neurons +→ 2 output channels (sin, cos of lead angle).
+ +| Layer | Size | Role |
|---|---|---|
| Input | 80 | Population code: bearing (36 neurons, offset 0), velocity direction (36 neurons, offset 36), speed thermometer (8 neurons, offset 72) |
| Hidden | 12 LIF | Each neuron integrates weighted input, leaks by factor 0.9/tick, spikes when membrane potential ≥ 0.08, resets to 0 |
| Output | 2 linear | Weighted sum of hidden spikes → sin and cos channels; decoded via atan2 to get angle |
The LIF update per hidden neuron h each tick:
V[h] = LEAK * V[h] + Σ_i (input[i] * w_ih[i,h])
+if V[h] >= THRESH: spike[h] = 1.0; V[h] = 0.0
+else: spike[h] = 0.0
+
+Learning uses the SuperSpike three-factor rule (Zenke & Ganguli 2018):
+Δw = η × pre_trace × σ'(V) × error
+where σ'(V) is a surrogate derivative (peaks at threshold, giving
+gradient direction through the non-differentiable spike). Hidden neurons receive
+their error signal via fixed random feedback weights bFb — this
+avoids the weight-transport problem of backpropagation.
USE_RESERVOIR = true in the
+ source. The SNN path exists and compiles, but the grid is better-tuned right now.
+ The SNN is the intended long-term path because it can handle arbitrary input
+ dimensionality — you just add more input neurons.
+What is the bot actually trying to learn?
+ +At DECIDE time the enemy is at absolute bearing B (world-frame +degrees), moving with some velocity, at distance D pixels. A +bullet fired at power 2 travels at 14 px/tick. By the time the +bullet covers distance D, the enemy has moved.
+ +The perpendicular component of enemy velocity relative to the bullet line is
+called vPerp. It is computed in the DECIDE branch:
let relVelDir = velDirDeg - absBearing # velocity angle relative to bullet line
+let vPerp = velSpeed * sin(relVelDir) # lateral component only
+
+For a constant-velocity target the perfect offset is:
+offset = arcsin(vPerp / bulletSpeed) # ≈ 35° at full speed, perpendicular
+
+The grid is warm-started with exactly this formula via
+initLeadGrid(). The SNN must learn this same
+relationship from experience, without being told the formula.
EVALUATE is where the bot discovers how wrong its prediction was. The logic +(both grid and SNN paths share this upstream computation):
+ +# 1. How long will the bullet travel?
+travelTime = decideDist / BULLET_SPEED
+
+# 2. Where will the enemy be at impact?
+futureX = decideEnemyX + cos(decideVelDirDeg) * decideVelSpeed * travelTime
+futureY = decideEnemyY + sin(decideVelDirDeg) * decideVelSpeed * travelTime
+
+# 3. What bearing should we have aimed at?
+correctAngle = directionTo(myX, myY, futureX, futureY)
+
+# 4. Error = what we should have aimed - what the bare enemy bearing was
+correctOffset = normalizeRelativeAngle(correctAngle - lastAbsBearing)
+
+Then the two paths diverge:
+ +| Path | What happens with correctOffset |
|---|---|
| Grid | +Calls res.learn(lastVPerp, decideDist, correctOffset) which
+ adds sin/cos of the correct offset to the matching cell. Only learns if the
+ error exceeds the adaptive dead zone (0.3–0.5° depending on
+ enemy speed). |
+
| SNN | +Calls superSpikeUpdate() with the frozen spike counts and
+ voltages from DECIDE. The three-factor rule adjusts all weights in proportion
+ to pre-synaptic trace × surrogate derivative × output error. |
+
decide* fields frozen at DECIDE time. This is why the fields
+ exist despite looking redundant with lastEnemyX/Y.
+forward() return when a grid cell (and all its neighbors) has no data?-999.0 is the sentinel. When DECIDE receives it, bot.gridHasData is set false and the bot aims at the bare enemy bearing with no lead (cold-start behavior). It is not the warm-start value — warm-start is already baked into cells with count = 0.1, so a cell with only warm-start data will return a value. The sentinel fires only when totalW == 0.0 in the neighborhood loop — meaning all nearby cells are completely empty.forward() proc in reservoir.nim: if totalW == 0.0: return -999.0. That sentinel is what the DECIDE branch checks with if gridOffset <= -999.0.You now understand the full control loop: DECIDE → WAITING → +EVALUATE, how the grid accumulator stores and retrieves lead angles, and how the +SNN path is structured but dormant.
+ +In the next lesson we will dive into the SNN path in detail: LIF neuron dynamics
+tick by tick, why the surrogate derivative is needed (the spike is
+non-differentiable), what SuperSpike's three factors are and why each one is
+there, and what it would take to activate USE_RESERVOIR = false
+without the hit rate collapsing.
Lesson 2 of N — SNNBot_garage/src/SNNBot.nim
In Lesson 1, you saw the architecture: 80 input neurons → 12 hidden LIF neurons → 2 output channels (sin/cos). This lesson breaks open each piece — how neurons fire, how inputs are encoded, how the network learns. Every code snippet is from YOUR implementation in SNNBot.nim.
+ + + +The LIF model as implemented:
+ +V[t] = LEAK × V[t-1] + weighted_input
+if V >= THRESH: spike! V = 0
+
+Key constants from the code:
+LEAK = 0.9 — membrane leaks 10% per tick — neuron "forgets" over ~10 ticksTHRESH = 0.08 — very low threshold — easy to fireThe exact forward pass code for the hidden layer:
+ +for h in 0 ..< N_HID:
+ var wsum = 0.0
+ for i in 0 ..< N_IN:
+ wsum += inputs[i] * snn.wih[i * N_HID + h]
+ snn.vHid[h] = LEAK * snn.vHid[h] + wsum
+ if snn.vHid[h] >= THRESH:
+ spikesOut[h] = 1.0
+ snn.vHid[h] = 0.0 # reset
+ else:
+ spikesOut[h] = 0.0
+
+Continuous values (bearing, velocity direction, speed) must become spike patterns. The bearing encoder uses 36 neurons covering 360°, each neuron representing a 10° band:
+ +proc encodeBearing(inputs: var array[N_IN, float], bearing: float, offset: int) =
+ let norm = ((bearing + 180.0) / BAND_DEG) # 0..36
+ let lo = int(norm) mod 36
+ let hi = (lo + 1) mod 36
+ let frac = norm - float(int(norm))
+ inputs[offset + lo] = 1.0 - frac # stronger for closer band
+ inputs[offset + hi] = frac # weaker for farther band
+
+Triangular interpolation in action:
+ +bearing = 25°
+band 2 (20°): activation = 0.5 ▓▓▓▓▓░░░░░
+band 3 (30°): activation = 0.5 ▓▓▓▓▓░░░░░
+all others: activation = 0.0 ░░░░░░░░░░
+
+The 80-neuron layout:
+The output layer sums hidden spikes weighted by learned coefficients:
+ +sinOut = 0.0; cosOut = 0.0
+for h in 0 ..< N_HID:
+ sinOut += spikesOut[h] * snn.wSin[h]
+ cosOut += spikesOut[h] * snn.wCos[h]
+
+Then decoded: angle = atan2(sinOut, cosOut) × 180/π
N_INFER = 10: the network runs 10 ticks on the SAME input, accumulating sin/cos outputs. This averaging stabilizes the output — a single tick's spikes are noisy (binary), but the average over 10 ticks is smooth.
A spike is binary: 0 or 1. You can't take the gradient of a step function — it's zero everywhere except at the threshold, where it's infinity. So backpropagation doesn't work directly.
+ +SuperSpike replaces the true derivative with a smooth surrogate:
+ +proc surrogateDerivative(v: float): float =
+ let x = BETA * (v - THRESH)
+ result = 1.0 / ((1.0 + abs(x)) * (1.0 + abs(x)))
+
+This is a bell curve centered at v = THRESH. It is large when the voltage is NEAR the threshold (the neuron almost spiked or just barely spiked), and small when far from threshold (irrelevant neurons don't learn).
Each weight update is the product of THREE factors.
+ +For hidden→output weights (wSin, wCos):
Δw = η × rate_h × σ'(V_h) × error
+
+All three must be non-zero for learning to happen. A neuron that didn't fire (rate=0) doesn't learn. A neuron far from threshold (σ'≈0) doesn't learn. If the output is correct (error=0), nothing learns.
+ +For input→hidden weights (wih):
Δw = η_ih × preTrace_i × σ'(V_h) × error_h
+
+How does the hidden layer know its error? In backprop, you'd use the transpose of the output weights. SuperSpike uses RANDOM FIXED weights instead:
+ +let errHid = snn.bFb[h * 2 + 0] * errSin + snn.bFb[h * 2 + 1] * errCos
+
+These bFb weights are initialized randomly and NEVER updated. This is called feedback alignment — a controversial but effective shortcut. The hidden layer learns to align its representation with these random projections.
proc initSNN(snn: var SNN) =
+ for w in snn.wih.mitems: w = rand(0.2) - 0.1 # ±0.1
+ for w in snn.wSin.mitems: w = rand(0.2) - 0.1
+ for w in snn.wCos.mitems: w = rand(0.2) - 0.1
+ for b in snn.bFb.mitems: b = rand(2.0) - 1.0 # ±1.0, fixed forever
+
+The honest answer:
+targetAngle = gunDir + snnAngle (RELATIVE to gun), not absolute bearing. This is the feedback loop bug that was fixed for the grid path. The SNN path still has this bug.
+bFb) instead of transposing the output weights?targetAngle = gunDir + snnAngle. As the gun turns, gunDir changes, which changes the bearing inputs, which changes snnAngle. The output feeds back into the input, creating instability.You now understand both paths in your bot. The grid works but can't scale. The SNN can learn but hasn't been tested with the pipeline fixes. In the next lesson, we'll activate the SNN, fix the feedback loop bug, and run it against Walls — your first live SNN training run.
+ +