RunTraining exits 0 after each chunk; && break ended the whole run after
the first battle (counter 3219, not 9000). Loop now falls through the
success path and re-checks the persisted counter each iteration.
Round-end reward = cumulative totalScore/50 grows unboundedly with battle
length; long battles (5841 rounds) blew the critic's value scale: vLoss
10-30 during the 60-round cert, 353 at battle-1 round 1, 25745 by round 3871,
policy drift to 0/6 wins. 60-round battles reproduce the certified regime:
bounded value targets, fresh bot process per battle (clears thread state).
The event queue's heap seq was the last GC'd block surviving across
rounds: each round runs on a freshly spawned bot thread, so the N+1
thread realloc'd a block grown by dead thread N's allocator mid-round
(at the next capacity doubling, ~turn 104) -> rawDealloc SIGSEGV in
addEvent (7 gdb-confirmed coredumps). Replace with a static
array[MAX_QUEUE_SIZE, BotEvent] + eventsLen: no heap block crosses
threads, realloc can never happen.
Also fix the harness aborting the final round mid-train: PPO_Bot's
onRoundEnded trains synchronously after the runner's RoundEndedEvent,
so the counter read right after awaitResults() is the stale pre-train
value and System.exit killed the bot inside ppoUpdate. Poll up to 60s
for the counter to catch up before declaring the battle incomplete.
Verified: 72 consecutive rounds vs Fire, 100% wins, all rounds trained
(counter advanced 1:1), zero coredumps since the fix.
Save ACAdamStates (m/v tensors + t counters) as .npy files alongside
network weights in latest/ and checkpoint dirs; save round counter to
round_counter.txt. loadBestAvailable restores both on startup; fresh
start works unchanged when files are absent.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Covers forward/reverse decision, proportional steering with speed-dependent
turn rate clamping, and deceleration using the existing getNewTargetSpeed util.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Three bugs caused the radar to sweep continuously instead of locking:
1. run() loop set radar to Inf every tick, overwriting any lock
→ replaced with enemy_tracker.getRadarTurnRate()
2. onScannedBot used radarBearingTo() (math convention, east=0 CCW)
→ removed; run loop now handles radar via enemy_tracker
3. enemy_tracker.getRadarTurnRate() had arctan2(dx,dy) instead of
arctan2(dy,dx) — introduced by fd22535; bearing was off by ~90°
Also relaxed stale-lock threshold from 2 to 8 ticks to survive
brief scan gaps without falling back to full sweep.
Added tools/battle_runner for automated 1v1 testing.
Result: 1303/1308 ticks with successful scan (was ~1 in 4).
Per-tick SVG drawText overlay above the bot showing round number and
running average reward (e.g. "R:42 avg:3.50"). Per-round summary also
printed to the UI console via printToStdOut with tick count and score.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
arctan2(dy, dx) gives east-based math bearing; Tank Royale uses north=0°, CW+.
Swapping to arctan2(dx, dy) gives the correct game-space bearing.
Symptom: radar commanded 45°/tick away from a target directly ahead.
Adds test_radar_lock.nim as regression test (20-tick lock, ±15° tolerance).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- enemy_tracker: toggle lastOvershootDir each tick; make getRadarTurnRate take var tracker
- training: remove threadvar Adam globals; pass adamStates as var param to ppoUpdate; export ACAdamStates
- PPO_Bot: carry ACAdamStates through TrainingArgs/TrainingResult; drop trainingDone bool and Lock — use resultChan.tryRecv() directly as synchronisation
- weights: sort checkpoint dirs newest-first by mtime instead of hardcoded order
- tests/test_training: pass explicit ACAdamStates to ppoUpdate
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Manual-backprop PPO with Adam: TrajectoryBuffer, computeGAE, ppoUpdate
(4 epochs, minibatch 64, clip 0.2, grad norm 0.5). Reward helpers
computeTickReward/computeRoundReward. Bot wired: tick transitions
collected in run loop, ppoUpdate called on onRoundEnded. Fix: add
arraymancer import to PPO_Bot.nim so Tensor resolves at top level.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Two-hidden-layer MLP actor-critic (42→64→64→5/1) with stochastic
actorForward, logStd floor at -3, and BotAction mapper wired into
the run() loop. Assert-based test suite covers shapes, finiteness,
logStd collapse, and all action range bounds.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Evaluates A2C, PPO, TD3, SAC, DDPG against the constraints: short
on-policy episodes, no RL library, few-hundred-ms training window.
PPO wins on implementation simplicity and stability at this scale.
Closes#3
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>