Decide learning loop structure #150
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Question
How should the learning loop be structured? Options to decide:
abs(snn_output - enemy_bearing). Should reward be the raw error, the delta (improvement), or a shaped reward (e.g., exponential decay with distance)?These apply to both hill-climbing and STDP — the structure should be shared so results are comparable.
Parent map: #147
SirStone referenced this issue2026-09-12 11:37:39 +02:00
Resolution
Update granularity: Wait-then-evaluate. SNN outputs a target angle, gun travels there (max 20°/tick), new outputs suppressed until gun arrives within tolerance. Error measured only after gun settles. No credit buffer needed — stationary target means no cost to waiting.
Episodes: None. Continuous learning loop: output → wait for gun → measure error → update weights → repeat. Round boundaries are just interruptions, not episode boundaries.
Reward signal: Shaped inverse —
reward = 1.0 / (1.0 + error). Bounded 0–1, continuous gradient, rewards both improving and staying aimed.Weight persistence: Accumulate across rounds within a battle. No reset between rounds. Persist to disk is out of scope for now.
Learning algorithms: Two, in order:
Hill-climbing dropped — STDP does its job plus more.