Reverts reservoir.nim to the neighbor-blending forward() + single-cell
learn() version (pre-bilinear-interpolation). Adds /tmp/snnbot_round_stats.log
and /tmp/snnbot_aim_debug.log for diagnostics independent of test framework stdout.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Each learn() call now updates a 3×3 neighborhood (center w=1.0, cardinal w=0.3,
diagonal w=0.1). count field changed to float64 to support fractional weights.
Fills grid ~5× faster and eliminates one-sided interpolation at bin boundaries.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Ring buffer had amnesia — cycled out all data every 128 ticks,
preventing convergence. Grid accumulator permanently stores average
lead offsets indexed by (v_perp, distance). 136 cells, 1.5 KB.
Knowledge accumulates across rounds → convergence guaranteed for
stationary velocity patterns.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
3-bit speed encoding couldn't distinguish speed 3 from speed 5 —
9° of lead error baked into the input. Thermometer coding with 8 bits
gives 1-bit Hamming distance between adjacent speeds. MAX_K back to 128.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Replace (bearing, velDir, speed) with (relVelDir, speed, distance).
Raw bearing is irrelevant to lead offset — the correction depends on
how the target crosses the line of fire, not where it is. This lets
exemplars from one position generalize to all positions with similar
geometry.
Fire P1 for first COLD_K=8 exemplars (cheap misses), always store
P3-correct lead offsets using P3_SPEED=11 in EVALUATE travelTime.
Switches to P3 once exemplar buffer has enough data. Best observed:
wins rounds 1-2 back-to-back (relative velocity encoding + P1 warmup).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Exemplars now store lead correction offsets instead of absolute angles.
The system learns 'how much lead to apply' independently of bearing,
so one learned lead pattern generalizes to all positions on the
battlefield. Converges in a few ticks for constant-velocity targets.
Exponential decay (0.95^age) makes newest exemplars dominate the
weighted mean. Old stale exemplars fade naturally, so re-learning
after target movement is always fast regardless of buffer history.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Eliminates aim jitter by switching from 360-bin WTA to weighted circular
mean over stored exemplars. Continuous float output, no bins, no
hysteresis, no saturation. K=128 ring buffer, Hamming similarity with
quadratic weighting.
OR-only reinforcement caused all bins to fill with set bits over time,
making all scores similar and the winner random. Now the correct bin
is REPLACED with the exact input pattern (no accumulation). Neighbors
still OR for generalization.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Delete 1024-neuron reservoir — it added noise, not features.
Direct 80-bit input → 72 output bins via popcount + WTA Hebbian.
Same learning rule (OR reinforce, AND NOT punish), zero indirection.
Reservoir adds value for temporal features (step 3), not now.
Dense firing (~50%) caused all readout bins to saturate identically.
Replace threshold-based firing with k-WTA: only top 50 neurons fire
per tick (~5% sparsity). Sparse patterns give low inter-angle overlap
-> readout bins can discriminate between different bearings.
- Threshold 4→1: neurons now fire (~50% rate) instead of never firing
- Remove aggressive decay that erased learning signal immediately
- Remove EVALUATE re-forward that corrupted reservoir state before learn()
New architecture: 1024 binary neurons in fixed random reservoir,
72-bin population-coded output, WTA Hebbian learning with binary ops.
Forward pass: AND + popcount. Learning: OR (reinforce) / AND NOT (punish).
No backprop, no floats in hot path. Toggle via USE_RESERVOIR const.
Forecast: ~200-400 ticks to learn stationary target aiming.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>