Decide learning loop structure #150

Closed
opened 2026-09-12 11:37:14 +02:00 by SirStone · 1 comment
Owner

Question

How should the learning loop be structured? Options to decide:

  1. Granularity: Does learning update weights every tick, every round, or every battle? Per-tick gives fastest feedback but may be noisy. Per-round is a natural episode boundary.
  2. Episode definition: What constitutes one learning episode? One round? One battle of N rounds? A fixed number of ticks?
  3. Reward signal: Error = abs(snn_output - enemy_bearing). Should reward be the raw error, the delta (improvement), or a shaped reward (e.g., exponential decay with distance)?
  4. Reset policy: Do weights reset between episodes, or does learning accumulate across battles?

These apply to both hill-climbing and STDP — the structure should be shared so results are comparable.

Parent map: #147

## Question How should the learning loop be structured? Options to decide: 1. **Granularity**: Does learning update weights every tick, every round, or every battle? Per-tick gives fastest feedback but may be noisy. Per-round is a natural episode boundary. 2. **Episode definition**: What constitutes one learning episode? One round? One battle of N rounds? A fixed number of ticks? 3. **Reward signal**: Error = `abs(snn_output - enemy_bearing)`. Should reward be the raw error, the delta (improvement), or a shaped reward (e.g., exponential decay with distance)? 4. **Reset policy**: Do weights reset between episodes, or does learning accumulate across battles? These apply to both hill-climbing and STDP — the structure should be shared so results are comparable. Parent map: #147
Author
Owner

Resolution

Update granularity: Wait-then-evaluate. SNN outputs a target angle, gun travels there (max 20°/tick), new outputs suppressed until gun arrives within tolerance. Error measured only after gun settles. No credit buffer needed — stationary target means no cost to waiting.

Episodes: None. Continuous learning loop: output → wait for gun → measure error → update weights → repeat. Round boundaries are just interruptions, not episode boundaries.

Reward signal: Shaped inverse — reward = 1.0 / (1.0 + error). Bounded 0–1, continuous gradient, rewards both improving and staying aimed.

Weight persistence: Accumulate across rounds within a battle. No reset between rounds. Persist to disk is out of scope for now.

Learning algorithms: Two, in order:

  1. Reward-modulated STDP — proves topology AND teaches spiking mechanics (replaces hill-climbing which was redundant)
  2. Surrogate gradient descent — the scalable method, deferred to fog (not yet specified)

Hill-climbing dropped — STDP does its job plus more.

## Resolution **Update granularity:** Wait-then-evaluate. SNN outputs a target angle, gun travels there (max 20°/tick), new outputs suppressed until gun arrives within tolerance. Error measured only after gun settles. No credit buffer needed — stationary target means no cost to waiting. **Episodes:** None. Continuous learning loop: output → wait for gun → measure error → update weights → repeat. Round boundaries are just interruptions, not episode boundaries. **Reward signal:** Shaped inverse — `reward = 1.0 / (1.0 + error)`. Bounded 0–1, continuous gradient, rewards both improving and staying aimed. **Weight persistence:** Accumulate across rounds within a battle. No reset between rounds. Persist to disk is out of scope for now. **Learning algorithms:** Two, in order: 1. Reward-modulated STDP — proves topology AND teaches spiking mechanics (replaces hill-climbing which was redundant) 2. Surrogate gradient descent — the scalable method, deferred to fog (not yet specified) Hill-climbing dropped — STDP does its job plus more.
SirStone added the wayfinder:grilling label 2026-09-12 12:22:03 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#150