Three bugs caused the radar to sweep continuously instead of locking:
1. run() loop set radar to Inf every tick, overwriting any lock
→ replaced with enemy_tracker.getRadarTurnRate()
2. onScannedBot used radarBearingTo() (math convention, east=0 CCW)
→ removed; run loop now handles radar via enemy_tracker
3. enemy_tracker.getRadarTurnRate() had arctan2(dx,dy) instead of
arctan2(dy,dx) — introduced by fd22535; bearing was off by ~90°
Also relaxed stale-lock threshold from 2 to 8 ticks to survive
brief scan gaps without falling back to full sweep.
Added tools/battle_runner for automated 1v1 testing.
Result: 1303/1308 ticks with successful scan (was ~1 in 4).
Per-tick SVG drawText overlay above the bot showing round number and
running average reward (e.g. "R:42 avg:3.50"). Per-round summary also
printed to the UI console via printToStdOut with tick count and score.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- enemy_tracker: toggle lastOvershootDir each tick; make getRadarTurnRate take var tracker
- training: remove threadvar Adam globals; pass adamStates as var param to ppoUpdate; export ACAdamStates
- PPO_Bot: carry ACAdamStates through TrainingArgs/TrainingResult; drop trainingDone bool and Lock — use resultChan.tryRecv() directly as synchronisation
- weights: sort checkpoint dirs newest-first by mtime instead of hardcoded order
- tests/test_training: pass explicit ACAdamStates to ppoUpdate
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Manual-backprop PPO with Adam: TrajectoryBuffer, computeGAE, ppoUpdate
(4 epochs, minibatch 64, clip 0.2, grad norm 0.5). Reward helpers
computeTickReward/computeRoundReward. Bot wired: tick transitions
collected in run loop, ppoUpdate called on onRoundEnded. Fix: add
arraymancer import to PPO_Bot.nim so Tensor resolves at top level.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Two-hidden-layer MLP actor-critic (42→64→64→5/1) with stochastic
actorForward, logStd floor at -3, and BotAction mapper wired into
the run() loop. Assert-based test suite covers shapes, finiteness,
logStd collapse, and all action range bounds.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>