Commit Graph

393 Commits

Author SHA1 Message Date
SirStone 717ef3ead8 Merge branch 'worktree-agent-a8622248' (ticket #46 weight persistence) 2026-08-20 23:57:53 +02:00
SirStone 54b8139b11 feat(SAC_LSTM_Bot): weight persistence module (#46)
Save/load all SAC-LSTM tensors (actor, 2 critics, 2 target critics,
alpha, Adam states) into a single .zip of .npy files. Atomic write
via temp path + rename. Adam types (AdamVar, SACAdamStates) defined
here for training.nim to use.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:57:13 +02:00
SirStone 55ef22ff8b Merge branch 'worktree-agent-a3ca3066' (ticket #45 replay buffer) 2026-08-20 23:41:23 +02:00
SirStone 5ea57bcae3 Merge branch 'worktree-agent-a393a8b9' (ticket #44 reward module) 2026-08-20 23:41:23 +02:00
SirStone cb33551621 feat(SAC_LSTM_Bot): replay buffer module (#45)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:40:17 +02:00
SirStone 23c65c9ac6 feat(SAC_LSTM_Bot): reward module (#44)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:39:45 +02:00
SirStone 4ee0d8272c feat(SAC_LSTM_Bot): action mapping module (#43)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:38:45 +02:00
SirStone add3e34926 feat(SAC_LSTM_Bot): LSTM network module (#41)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:38:17 +02:00
SirStone f130bf1254 feat(SAC_LSTM_Bot): state vector module (#42)
35-dim normalized tensor (GameState → buildState). No history window —
LSTM handles temporal context. Covers own-bot (7), enemy (7), derived (4),
walls (4), bullets (12), scan staleness (1). All tests pass.
2026-08-20 23:37:56 +02:00
SirStone a0a3840980 feat(SAC_LSTM_Bot): skeleton bot with radar lock and colors (#40)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:37:50 +02:00
SirStone 7d73d32c85 Merge branch 'worktree-agent-ac811b59' (ticket #39 radar lock) 2026-08-20 23:34:24 +02:00
SirStone df3bbbd14e Merge branch 'worktree-agent-a1c1f549' (ticket #38 SAC_LSTM_Bot scaffold) 2026-08-20 23:34:21 +02:00
SirStone 4258d364b9 feat(SAC_LSTM_Bot): project scaffold (#38)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:33:43 +02:00
SirStone f88580b157 feat(radar_lock): standalone reusable module (#39)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:33:40 +02:00
SirStone ca3e3d2272 tune(PPO_Bot): logStd=-2.0 (std≈0.135), entropy=0, ceiling=-1.0
Stochastic eval at std≈0.37 was 0/10 vs Corners (deterministic: 10/10).
Warm-start policy is correct but brittle — any noise breaks it.
- log_std initialized to -2.0 (std≈0.135) for moderate exploration
- entropy_coeff=0.0 (no push toward exploration during fine-tuning)
- logStd ceiling=-1.0 (cap at std≈0.37)
2026-08-20 15:28:15 +02:00
SirStone 0d35646dc9 feat(PPO_Bot): deterministic eval + fix logStd warm-start
- actorForward: deterministic param, uses mean-only when PPOB_EVAL_ONLY=1
  (eval was adding unit Gaussian noise to every action — unreliable scores)
- warm_start.py: log_std initialized to -1.0 (std≈0.37) instead of copying
  snapshot values (were 2.27-4.68 → std 9-108, completely drowning signal)
- training.env: LOG_STD_CEILING 0.0→-0.5 (cap exploration at std≈0.6)
2026-08-20 15:18:07 +02:00
SirStone c834d2cbee fix(PPO_Bot): round_counter always written after increment
Early-return guard on empty buffer was skipping round_counter.txt write,
causing training script to think bot crashed (counter stuck at 0).
2026-08-20 15:11:12 +02:00
SirStone fedab54bc0 feat(PPO_Bot): multi-round transition accumulation (UPDATE_INTERVAL=10)
- Accumulate transitions across 10 rounds (~3000) before PPO update
  (was per-round ~300 — gradient estimates were far too noisy)
- training.nim: MAX_TRANSITIONS 4096→8192, done flag on transitions,
  GAE handles episode boundaries correctly
- PPO_Bot.nim: buffer persists across rounds, update every N rounds
- training.env: lr 5e-5→1e-4, entropy 0.001, UPDATE_INTERVAL=10
2026-08-20 15:06:00 +02:00
SirStone 82eeb53e5c tune(training): 20-round chunks, 50 eval rounds, 30k total rounds
- generalist_train.sh: CHUNK_SIZE 60→20 for faster opponent cycling
- generalist_train.sh: EVAL_ROUNDS 30→50 for more reliable eval
- training.env: TRAINING_ROUNDS→30000, removed hardcoded opponent
- warm_start.py: TARGET_DIM=57 (already committed, ensure latest)
2026-08-20 14:38:09 +02:00
SirStone 12624d3069 feat(PPO_Bot): bot-relative bullets + scan staleness (STATE_DIM=57)
- Bullet state (indices 44-55): enemy-relative → bot-relative frame
  (bot needs threat vectors to itself for dodging, not to enemy)
- New index 56: scan staleness = min(ticksSinceLastScan / 30, 1.0)
  (gives policy a confidence signal for enemy data freshness)
- warm_start.py updated: 44→57 dim expansion, TARGET_DIM variable
- Tests updated for new state layout

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 14:27:28 +02:00
SirStone 6ad51148f4 fix(PPO_Bot): SIGSEGV crash fixes + static buffers for thread safety
- bullets: seq[InFlightBullet] → array[4, InFlightBullet] + bulletCount
  (eliminates cross-thread heap realloc under ORC)
- hasFired: edge-triggered (cleared after state build, not level-triggered)
- round_counter parseInt: wrapped for empty/torn file → 0
- Static SVG + intent buffers to kill cross-thread heap realloc
- Tick-local alive/bulletData also fixed arrays

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 14:23:08 +02:00
SirStone 75e32e3315 PPO_Bot Fire campaign: 3787/3789 wins (99.95%); frozen eval 500/500
Single 3789-round battle (5212-9000), only rounds 1-2 lost (cold start);
last 100 rounds 100%. Frozen-policy eval (PPOB_EVAL_ONLY=1, 500 rounds vs
Fire, counter 9000->9500): 500/500 = 100%, zero train lines, weights mtime
and content untouched. vLoss avg 39.7/max 361 stable throughout — bounded
terminal reward (db99153) holds at cumulative score ~462k.
2026-08-19 04:00:47 +02:00
SirStone db99153f65 cap terminal reward scale (bounded score bonus) and restore single-battle campaigns
computeRoundReward used cumulative totalScore/50 — unbounded in long battles
(vLoss 353 at round 3160 → 25745 by 3871 in the 5841-round attempt). Cap the
score term at 400 before /50: bonus ∈ [0,8], so the critic's value scale stays
stable regardless of battle length and across battle boundaries.

Reverts the 60-round battle chunking (186e005/da2f825): one battle per
campaign for the whole remaining budget; keeps the crash-restart loop, the
mid-battle freeze guard and the end-of-battle counter completeness check.

Cert (5211, single 60-round battle): 59/60 wins (sole loss = cold-start round
1, score 61), vLoss avg 36.1 / max 148.5, gNorm max 596, zero NaN, zero
restarts, counter check passed. Weights persist to round 5211.
2026-08-19 03:50:14 +02:00
SirStone da2f825ad8 fix(training): keep run.sh looping across 60-round battle chunks
RunTraining exits 0 after each chunk; && break ended the whole run after
the first battle (counter 3219, not 9000). Loop now falls through the
success path and re-checks the persisted counter each iteration.
2026-08-19 03:36:20 +02:00
SirStone 186e005a96 fix(training): cap battles at 60 rounds to bound round-end reward scale
Round-end reward = cumulative totalScore/50 grows unboundedly with battle
length; long battles (5841 rounds) blew the critic's value scale: vLoss
10-30 during the 60-round cert, 353 at battle-1 round 1, 25745 by round 3871,
policy drift to 0/6 wins. 60-round battles reproduce the certified regime:
bounded value targets, fresh bot process per battle (clears thread state).
2026-08-19 03:32:31 +02:00
SirStone a4e830531b fix(botapi): static SVG + intent buffers to kill cross-thread heap realloc
Round N+1's fresh bot thread realloc'd module-level strings/seqs (SVG buffer,
intent stdout/stderr, team messages) left behind by dead round N's thread —
same rawDealloc SIGSEGV class as the event queue, seen at graphics.nim:274
(drawText->prepareAdd, core 2490478 @ 03:13:40, battle round 257).

- graphics.nim: gSvgBuffer -> array[16384, char] + gSvgLen, appendSvg
- bot.nim: intent stdout/stderr -> static char arrays; team messages ->
  array[16, TeamMessage] + len; buildIntentJson/printToStdOut/Err/
  broadcastTeamMessage bounded appends
- botThreadEntry: reset graphics+intent buffers on the owning thread

Also fixes stale mapActions call sites in tests/ (missing enemyX/enemyY).
2026-08-19 03:32:28 +02:00
SirStone 64697f917e fix(botapi): static event queue storage + end-of-battle train wait
The event queue's heap seq was the last GC'd block surviving across
rounds: each round runs on a freshly spawned bot thread, so the N+1
thread realloc'd a block grown by dead thread N's allocator mid-round
(at the next capacity doubling, ~turn 104) -> rawDealloc SIGSEGV in
addEvent (7 gdb-confirmed coredumps). Replace with a static
array[MAX_QUEUE_SIZE, BotEvent] + eventsLen: no heap block crosses
threads, realloc can never happen.

Also fix the harness aborting the final round mid-train: PPO_Bot's
onRoundEnded trains synchronously after the runner's RoundEndedEvent,
so the counter read right after awaitResults() is the stale pre-train
value and System.exit killed the bot inside ppoUpdate. Poll up to 60s
for the counter to catch up before declaring the battle incomplete.

Verified: 72 consecutive rounds vs Fire, 100% wins, all rounds trained
(counter advanced 1:1), zero coredumps since the fix.
2026-08-19 03:12:22 +02:00
SirStone 766b9e03ee feat(PPO_Bot): enemy-centered action space + reward shaping for 100% vs Target
- actions.nim: goto/aimTo coordinates now offset from enemy position
  (enemyX + tanh(raw) * scale) instead of absolute arena coords
  (sigmoid(raw) * arenaSize). Initial random policy defaults to
  approaching and aiming at enemy.
- training.nim: added dense reward shaping (distance closeness +
  gun bearing) to computeTickReward, doubled round reward scaling.
- PPO_Bot.nim: passes enemy position to mapActions, computes
  gun-to-enemy bearing for reward shaping.

Result: 100/100 win rate vs Target with frozen weights.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-18 17:40:59 +02:00
SirStone 0b17430735 feat(PPO_Bot): persist Adam optimizer state and round counter across restarts (#35)
Save ACAdamStates (m/v tensors + t counters) as .npy files alongside
network weights in latest/ and checkpoint dirs; save round counter to
round_counter.txt. loadBestAvailable restores both on startup; fresh
start works unchanged when files are absent.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-18 15:40:54 +02:00
SirStone cdde60d79f feat(PPO_Bot): command abstraction layer — goto/aimTo controllers (#24)
- Add gotoTick/aimToTick controller functions (#25)
- Update network dims: actor 5→6, state 42→44 (#26)
- Rewrite mapActions for 6-dim command space (#27)
- Delete stale weight files (shape mismatch)
- Fix existing tests for new signatures

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-17 19:20:40 +02:00
SirStone 56e0b306c9 docs(research): goto controller algorithm for issue #20
Covers forward/reverse decision, proportional steering with speed-dependent
turn rate clamping, and deceleration using the existing getNewTargetSpeed util.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-17 17:04:57 +02:00
SirStone e5609a7d9b fix(PPO_Bot): radar lock — use enemy_tracker width-lock, fix arctan2 arg order
Three bugs caused the radar to sweep continuously instead of locking:

1. run() loop set radar to Inf every tick, overwriting any lock
   → replaced with enemy_tracker.getRadarTurnRate()

2. onScannedBot used radarBearingTo() (math convention, east=0 CCW)
   → removed; run loop now handles radar via enemy_tracker

3. enemy_tracker.getRadarTurnRate() had arctan2(dx,dy) instead of
   arctan2(dy,dx) — introduced by fd22535; bearing was off by ~90°

Also relaxed stale-lock threshold from 2 to 8 ticks to survive
brief scan gaps without falling back to full sweep.

Added tools/battle_runner for automated 1v1 testing.

Result: 1303/1308 ticks with successful scan (was ~1 in 4).
2026-08-17 11:52:36 +02:00
SirStone a8ee2a86e3 feat(PPO_Bot): show training progress in game UI
Per-tick SVG drawText overlay above the bot showing round number and
running average reward (e.g. "R:42 avg:3.50"). Per-round summary also
printed to the UI console via printToStdOut with tick count and score.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-16 16:48:08 +02:00
SirStone bbc9e51166 fix(PPO_Bot): state vector bearing uses game coords (north=0° CW)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-16 16:40:28 +02:00
SirStone fd22535f5b fix(PPO_Bot): radar lock oscillation bug — arctan2 arg order wrong for Tank Royale coords
arctan2(dy, dx) gives east-based math bearing; Tank Royale uses north=0°, CW+.
Swapping to arctan2(dx, dy) gives the correct game-space bearing.
Symptom: radar commanded 45°/tick away from a target directly ahead.
Adds test_radar_lock.nim as regression test (20-tick lock, ±15° tolerance).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-16 16:38:15 +02:00
SirStone f27b0238f0 feat(PPO_Bot): full PPO RL implementation (#14, #15, #16, #17) 2026-08-16 16:34:37 +02:00
SirStone aea0724d3a fix(PPO_Bot): radar oscillation, Adam persistence, checkpoint order, channel race
- enemy_tracker: toggle lastOvershootDir each tick; make getRadarTurnRate take var tracker
- training: remove threadvar Adam globals; pass adamStates as var param to ppoUpdate; export ACAdamStates
- PPO_Bot: carry ACAdamStates through TrainingArgs/TrainingResult; drop trainingDone bool and Lock — use resultChan.tryRecv() directly as synchronisation
- weights: sort checkpoint dirs newest-first by mtime instead of hardcoded order
- tests/test_training: pass explicit ACAdamStates to ppoUpdate

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-16 15:42:06 +02:00
SirStone 473d67f644 feat(PPO_Bot): weight persistence + background training (#17)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-16 15:35:42 +02:00
SirStone eadd177d3b feat(PPO_Bot): reward + trajectory + GAE + PPO training (#16)
Manual-backprop PPO with Adam: TrajectoryBuffer, computeGAE, ppoUpdate
(4 epochs, minibatch 64, clip 0.2, grad norm 0.5). Reward helpers
computeTickReward/computeRoundReward. Bot wired: tick transitions
collected in run loop, ppoUpdate called on onRoundEnded. Fix: add
arraymancer import to PPO_Bot.nim so Tensor resolves at top level.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-16 15:27:15 +02:00
SirStone 588c9ebc2f feat(PPO_Bot): enemy tracker + 42-float state vector (#15)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-16 15:12:13 +02:00
SirStone aa4bc77068 feat(PPO_Bot): network forward pass + action mapping (#14)
Two-hidden-layer MLP actor-critic (42→64→64→5/1) with stochastic
actorForward, logStd floor at -3, and BotAction mapper wired into
the run() loop. Assert-based test suite covers shapes, finiteness,
logStd collapse, and all action range bounds.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-16 15:10:46 +02:00
SirStone 30cda871cc research: RL algorithm choice — recommend PPO for Tank Royale bot
Evaluates A2C, PPO, TD3, SAC, DDPG against the constraints: short
on-policy episodes, no RL library, few-hundred-ms training window.
PPO wins on implementation simplicity and stability at this scale.

Closes #3

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-15 22:07:47 +02:00
SirStone 64f73dd413 chore: add agent skills configuration
Add CLAUDE.md and docs/agents/ with issue tracker (Gitea), triage labels, and domain doc consumer rules.
2026-08-15 19:26:43 +02:00