Commit Graph

213 Commits

Author SHA1 Message Date
SirStone 1619b86f25 docs(SAC_LSTM_Bot): v2 restart saga, twin-freeze correction, pacing decisions 2026-08-22 21:26:24 +02:00
SirStone 19f34abf0c fix(SAC_LSTM_Bot): freeze mirror-twin via eval-mode gate in twin launcher
SACLSTM_EVAL_MODE=1 in SacTwin.sh suppresses all sendTrainingMsg traffic
(lever-4 gate), so the twin never trains — not even in-RAM within a battle.
Required now that the main bot's SACLSTM_SAVE_INTERVAL drops to 1 (v2 relaunch
after checkpoint-cadence diagnosis): without the gate the twin would persist
per-battle drift and stop being the frozen reproducible opponent #54 specifies.
2026-08-22 21:09:07 +02:00
SirStone bd58794b4c docs(SAC_LSTM_Bot): campaign v2 chapter — locked config, fresh-start archive, twin reseed, ceiling net, launch health evidence (#59) 2026-08-22 20:31:33 +02:00
SirStone 6fc01eb4e5 feat(SAC_LSTM_Bot): campaign v2 levers — aggression/anti-ram reward shaping + stability knob overrides (part 2)
Lever 2 (#59): x1.25 aggression mult on damage dealt, flat +0.5 hit bonus,
-3.0 per bot-bot collision (server deals RAM_DAMAGE=0.6 to both parties but
only notifies the hitter), escalating proximity deterrent below 12% arena
diagonal suppressed while dealing damage. Win/loss terminals unchanged and
dominant. All weights TUNABLE consts marked ponytail. SACLSTM_REWARD_DEBUG=1
env-gated reward_debug.log for calibration greps.

Lever 5 (#59): no code needed — SACLSTM_LR_ACTOR/LR_CRITIC/LR_ALPHA (3e-4)
and SACLSTM_TARGET_ENTROPY (-4.0) were already env-overridable in training.nim.

Smoke vs RamFire+Crazy (hidden=32, random init, isolated weights): 75 ram
penalties, 381 charge events, hit bonuses firing, 0 crashes, metrics JSONL
flowing. Tests: 8/8 suites green incl. new assert-level term math.
2026-08-22 19:44:58 +02:00
SirStone a07e5305f5 feat(SAC_LSTM_Bot): campaign v2 levers — loss metrics, eval-mode gate, eval rotation + MA gating (part 1)
Levers 3, 4, 1 of the #57 sign-off (execution order 3->4->1), tracked in #59.

- Lever 3 (#59): one JSONL line per trainPass in training_metrics.jsonl with
  exactly the scalars sacUpdate already exposes (SACMetrics: critic/actor/alpha
  losses + alpha, averaged per pass) plus epoch, buffer size (replay_buffer.len),
  cumulative steps and drained count. No trainer change needed.
- Lever 4 (#59): sendTrainingMsg drops all training input while SACLSTM_EVAL_MODE=1
  (existing #49 harness mechanism) — eval battles can neither pollute the replay
  buffer nor trigger gradient updates; one-time stderr notice at bot init.
- Lever 1 (#59): sac_train.sh evaluates every SAC_EVAL_OPPONENTS entry per cycle
  (results carry opponent name in eval_log.jsonl); best-gating now uses a
  composite = mean over opponents of the last-5-evals moving average per
  opponent. best_score.txt format change: float composite replaces the
  single-opponent integer win rate semantics (retired).
- Tests: metricsLine JSONL scalars + eval-mode suppression asserts.

Refs: #59, #57
2026-08-22 18:58:47 +02:00
SirStone 2619ba06fc docs(SAC_LSTM_Bot): campaign ending, results table, verdict, hygiene (#57) 2026-08-22 14:40:57 +02:00
SirStone f45e8f2717 fix(SAC_LSTM_Bot): startup sweep of stale .part checkpoint corpses (#57) 2026-08-22 14:40:57 +02:00
SirStone b7492f1080 docs(SAC_LSTM_Bot): notebook — score:60 anatomy, .part forensics+sweep, ceiling defused 2026-08-22 08:08:31 +02:00
SirStone f1962c7506 docs(SAC_LSTM_Bot): backfill shift-1 milestone entry + morning audit 2026-08-22 07:27:37 +02:00
SirStone 05929d2dbd docs(SAC_LSTM_Bot): notebook — shift 1, deferred-intervention milestone 2026-08-22 07:14:35 +02:00
SirStone 4b64bf18ac docs(SAC_LSTM_Bot): campaign-v1 launch record — save-check incident, health check, check-in procedure (#56) 2026-08-22 00:31:47 +02:00
SirStone 26536713ba fix(SAC_LSTM_Bot): checkpoint save check inside gradient-step loop (#56)
Launch finding during campaign-v1 verification: at production sizes
(hidden 256, ~1s/step, ~13s trainer CPU per ~40s chunk process) the
save check ran only between drain-burst passes, so stepCount never
crossed nextSave before the process died — zero checkpoints persisted
across entire runs (masked at #49/#54 smoke sizes where steps were
sub-millisecond). Check now fires mid-loop; with SAVE_INTERVAL<=5
(within the per-process step budget) every chunk persists its chain.
2026-08-22 00:21:43 +02:00
SirStone 2f49cb243f docs(SAC_LSTM_Bot): campaign-v1 notebook — locked config + story so far (#56) 2026-08-21 23:26:31 +02:00
SirStone 6a294ad7ad feat(SAC_LSTM_Bot): mirror-twin sparring partner + readiness check (#54)
- make_twin.sh: generates self-contained SacTwin dir in the sample-bots
  archive (own json/sh identity, own weights dir seeded from a frozen
  sac_best.zip copy, own round_counter) so RunTraining.java resolves it
  like any sample bot; re-running resets the twin to the frozen baseline.
- src/SAC_LSTM_Bot.nim: SACLSTM_BOT_JSON env overrides the baked-in bot
  json (loadBotInfo gives json total precedence, #49) so the same binary
  boots under the twin's name.
- sac_train.sh: chunk loop is a while, not for-over-seq — a crash on the
  FINAL chunk previously fell through ((chunk--);continue on an exhausted
  seq list) and exited 0 with budget incomplete; observed live vs SacTwin.

Readiness dry-run (#54): weighted pool Corners:1,SacTwin:3 picked the twin
in 3/4 chunks; all battles counter-checked; deterministic eval parsed;
main sac_best.zip/counter untouched by twin (twin counter advanced
independently); crash-restart proven end-to-end incl. final-chunk retry.
2026-08-21 23:19:27 +02:00
SirStone edf26aa45d fix(SAC_LSTM_Bot): enforce MaxHidden cap on SACLSTM_HIDDEN_SIZE (review of #48/#49) 2026-08-21 22:17:55 +02:00
SirStone df256b4d3e feat(SAC_LSTM_Bot): training harness (#49)
sac_train.sh orchestrates chunked self-play via tools/training_runner/
RunTraining.java: weighted opponent sampling per chunk, deterministic
eval (SACLSTM_EVAL_MODE=1) every N chunks with win-rate tracking, best
checkpoint (weights/sac_best.zip) by eval score, crash-restart loop on
the runner's liveness detection.

Supporting changes:
- integration.nim: opponentKey() keys the NewBattle buffer-clear rule on
  getBotName(id) with numeric-id fallback (#49 Q14 follow-up);
  bumpRoundCounter() emits the per-round liveness signal.
- SAC_LSTM_Bot.nim: onRoundEnded -> bumpRoundCounter().
- RunTraining.java: BOT_NAME env parameterizes result matching
  (default PPO_Bot, unchanged behavior for PPO).
- Launch packaging: root SAC_LSTM_Bot.json + .sh for the booter;
  src json name aligned to 'SAC_LSTM_Bot' so self-reported identity
  matches the booted identity (mismatch = runner connect timeout).
2026-08-21 21:51:27 +02:00
SirStone 7104645f5d chore(deps): update vendored tankroyale botapi to v1.0.1
Syncs libs/tankroyale_botapi with SirStone/robocode_tankroyale_botapi
v1.0.1 (extracted from tank-royale nim branch @ 03195a814). The local
SIGSEGV fixes (static event queue/SVG/intent buffers) were already
ported upstream in issue #24 — content is otherwise identical.

New capability (upstream #23): opponent name exposure for ticket #49.
- bot.nim: gBotNames id→name table + getBotName(id) / updateBotNames()
- umbrella module: dispatch BotListUpdate messages to updateBotNames()

Vendored layout and wiring unchanged (--path via config.nims, module
name stays tankroyale_botapi).
2026-08-21 21:03:17 +02:00
SirStone 32b71d9fc8 feat(SAC_LSTM_Bot): main bot integration (#48) 2026-08-21 20:19:05 +02:00
SirStone 62a6cc8ccf fix(SAC_LSTM_Bot): training review fixes — hidden state ordering, redundant forwards, actor grad clip (#47)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-21 00:16:47 +02:00
SirStone 415d4e3738 feat(SAC_LSTM_Bot): SAC training module (#47)
Implements sacUpdate with burn-in LSTM warm-up, twin-critic TD update,
actor reparameterization gradient, auto-alpha, and soft target update.
Manual backprop (linear + LSTM single-step, truncated BPTT). 11 new tests
all green; full regression suite (57+ tests) unaffected.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-21 00:11:35 +02:00
SirStone 717ef3ead8 Merge branch 'worktree-agent-a8622248' (ticket #46 weight persistence) 2026-08-20 23:57:53 +02:00
SirStone 54b8139b11 feat(SAC_LSTM_Bot): weight persistence module (#46)
Save/load all SAC-LSTM tensors (actor, 2 critics, 2 target critics,
alpha, Adam states) into a single .zip of .npy files. Atomic write
via temp path + rename. Adam types (AdamVar, SACAdamStates) defined
here for training.nim to use.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:57:13 +02:00
SirStone 55ef22ff8b Merge branch 'worktree-agent-a3ca3066' (ticket #45 replay buffer) 2026-08-20 23:41:23 +02:00
SirStone 5ea57bcae3 Merge branch 'worktree-agent-a393a8b9' (ticket #44 reward module) 2026-08-20 23:41:23 +02:00
SirStone cb33551621 feat(SAC_LSTM_Bot): replay buffer module (#45)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:40:17 +02:00
SirStone 23c65c9ac6 feat(SAC_LSTM_Bot): reward module (#44)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:39:45 +02:00
SirStone 4ee0d8272c feat(SAC_LSTM_Bot): action mapping module (#43)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:38:45 +02:00
SirStone add3e34926 feat(SAC_LSTM_Bot): LSTM network module (#41)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:38:17 +02:00
SirStone f130bf1254 feat(SAC_LSTM_Bot): state vector module (#42)
35-dim normalized tensor (GameState → buildState). No history window —
LSTM handles temporal context. Covers own-bot (7), enemy (7), derived (4),
walls (4), bullets (12), scan staleness (1). All tests pass.
2026-08-20 23:37:56 +02:00
SirStone a0a3840980 feat(SAC_LSTM_Bot): skeleton bot with radar lock and colors (#40)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:37:50 +02:00
SirStone 7d73d32c85 Merge branch 'worktree-agent-ac811b59' (ticket #39 radar lock) 2026-08-20 23:34:24 +02:00
SirStone df3bbbd14e Merge branch 'worktree-agent-a1c1f549' (ticket #38 SAC_LSTM_Bot scaffold) 2026-08-20 23:34:21 +02:00
SirStone 4258d364b9 feat(SAC_LSTM_Bot): project scaffold (#38)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:33:43 +02:00
SirStone f88580b157 feat(radar_lock): standalone reusable module (#39)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:33:40 +02:00
SirStone ca3e3d2272 tune(PPO_Bot): logStd=-2.0 (std≈0.135), entropy=0, ceiling=-1.0
Stochastic eval at std≈0.37 was 0/10 vs Corners (deterministic: 10/10).
Warm-start policy is correct but brittle — any noise breaks it.
- log_std initialized to -2.0 (std≈0.135) for moderate exploration
- entropy_coeff=0.0 (no push toward exploration during fine-tuning)
- logStd ceiling=-1.0 (cap at std≈0.37)
2026-08-20 15:28:15 +02:00
SirStone 0d35646dc9 feat(PPO_Bot): deterministic eval + fix logStd warm-start
- actorForward: deterministic param, uses mean-only when PPOB_EVAL_ONLY=1
  (eval was adding unit Gaussian noise to every action — unreliable scores)
- warm_start.py: log_std initialized to -1.0 (std≈0.37) instead of copying
  snapshot values (were 2.27-4.68 → std 9-108, completely drowning signal)
- training.env: LOG_STD_CEILING 0.0→-0.5 (cap exploration at std≈0.6)
2026-08-20 15:18:07 +02:00
SirStone c834d2cbee fix(PPO_Bot): round_counter always written after increment
Early-return guard on empty buffer was skipping round_counter.txt write,
causing training script to think bot crashed (counter stuck at 0).
2026-08-20 15:11:12 +02:00
SirStone fedab54bc0 feat(PPO_Bot): multi-round transition accumulation (UPDATE_INTERVAL=10)
- Accumulate transitions across 10 rounds (~3000) before PPO update
  (was per-round ~300 — gradient estimates were far too noisy)
- training.nim: MAX_TRANSITIONS 4096→8192, done flag on transitions,
  GAE handles episode boundaries correctly
- PPO_Bot.nim: buffer persists across rounds, update every N rounds
- training.env: lr 5e-5→1e-4, entropy 0.001, UPDATE_INTERVAL=10
2026-08-20 15:06:00 +02:00
SirStone 82eeb53e5c tune(training): 20-round chunks, 50 eval rounds, 30k total rounds
- generalist_train.sh: CHUNK_SIZE 60→20 for faster opponent cycling
- generalist_train.sh: EVAL_ROUNDS 30→50 for more reliable eval
- training.env: TRAINING_ROUNDS→30000, removed hardcoded opponent
- warm_start.py: TARGET_DIM=57 (already committed, ensure latest)
2026-08-20 14:38:09 +02:00
SirStone 12624d3069 feat(PPO_Bot): bot-relative bullets + scan staleness (STATE_DIM=57)
- Bullet state (indices 44-55): enemy-relative → bot-relative frame
  (bot needs threat vectors to itself for dodging, not to enemy)
- New index 56: scan staleness = min(ticksSinceLastScan / 30, 1.0)
  (gives policy a confidence signal for enemy data freshness)
- warm_start.py updated: 44→57 dim expansion, TARGET_DIM variable
- Tests updated for new state layout

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 14:27:28 +02:00
SirStone 6ad51148f4 fix(PPO_Bot): SIGSEGV crash fixes + static buffers for thread safety
- bullets: seq[InFlightBullet] → array[4, InFlightBullet] + bulletCount
  (eliminates cross-thread heap realloc under ORC)
- hasFired: edge-triggered (cleared after state build, not level-triggered)
- round_counter parseInt: wrapped for empty/torn file → 0
- Static SVG + intent buffers to kill cross-thread heap realloc
- Tick-local alive/bulletData also fixed arrays

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 14:23:08 +02:00
SirStone 75e32e3315 PPO_Bot Fire campaign: 3787/3789 wins (99.95%); frozen eval 500/500
Single 3789-round battle (5212-9000), only rounds 1-2 lost (cold start);
last 100 rounds 100%. Frozen-policy eval (PPOB_EVAL_ONLY=1, 500 rounds vs
Fire, counter 9000->9500): 500/500 = 100%, zero train lines, weights mtime
and content untouched. vLoss avg 39.7/max 361 stable throughout — bounded
terminal reward (db99153) holds at cumulative score ~462k.
2026-08-19 04:00:47 +02:00
SirStone db99153f65 cap terminal reward scale (bounded score bonus) and restore single-battle campaigns
computeRoundReward used cumulative totalScore/50 — unbounded in long battles
(vLoss 353 at round 3160 → 25745 by 3871 in the 5841-round attempt). Cap the
score term at 400 before /50: bonus ∈ [0,8], so the critic's value scale stays
stable regardless of battle length and across battle boundaries.

Reverts the 60-round battle chunking (186e005/da2f825): one battle per
campaign for the whole remaining budget; keeps the crash-restart loop, the
mid-battle freeze guard and the end-of-battle counter completeness check.

Cert (5211, single 60-round battle): 59/60 wins (sole loss = cold-start round
1, score 61), vLoss avg 36.1 / max 148.5, gNorm max 596, zero NaN, zero
restarts, counter check passed. Weights persist to round 5211.
2026-08-19 03:50:14 +02:00
SirStone da2f825ad8 fix(training): keep run.sh looping across 60-round battle chunks
RunTraining exits 0 after each chunk; && break ended the whole run after
the first battle (counter 3219, not 9000). Loop now falls through the
success path and re-checks the persisted counter each iteration.
2026-08-19 03:36:20 +02:00
SirStone 186e005a96 fix(training): cap battles at 60 rounds to bound round-end reward scale
Round-end reward = cumulative totalScore/50 grows unboundedly with battle
length; long battles (5841 rounds) blew the critic's value scale: vLoss
10-30 during the 60-round cert, 353 at battle-1 round 1, 25745 by round 3871,
policy drift to 0/6 wins. 60-round battles reproduce the certified regime:
bounded value targets, fresh bot process per battle (clears thread state).
2026-08-19 03:32:31 +02:00
SirStone a4e830531b fix(botapi): static SVG + intent buffers to kill cross-thread heap realloc
Round N+1's fresh bot thread realloc'd module-level strings/seqs (SVG buffer,
intent stdout/stderr, team messages) left behind by dead round N's thread —
same rawDealloc SIGSEGV class as the event queue, seen at graphics.nim:274
(drawText->prepareAdd, core 2490478 @ 03:13:40, battle round 257).

- graphics.nim: gSvgBuffer -> array[16384, char] + gSvgLen, appendSvg
- bot.nim: intent stdout/stderr -> static char arrays; team messages ->
  array[16, TeamMessage] + len; buildIntentJson/printToStdOut/Err/
  broadcastTeamMessage bounded appends
- botThreadEntry: reset graphics+intent buffers on the owning thread

Also fixes stale mapActions call sites in tests/ (missing enemyX/enemyY).
2026-08-19 03:32:28 +02:00
SirStone 64697f917e fix(botapi): static event queue storage + end-of-battle train wait
The event queue's heap seq was the last GC'd block surviving across
rounds: each round runs on a freshly spawned bot thread, so the N+1
thread realloc'd a block grown by dead thread N's allocator mid-round
(at the next capacity doubling, ~turn 104) -> rawDealloc SIGSEGV in
addEvent (7 gdb-confirmed coredumps). Replace with a static
array[MAX_QUEUE_SIZE, BotEvent] + eventsLen: no heap block crosses
threads, realloc can never happen.

Also fix the harness aborting the final round mid-train: PPO_Bot's
onRoundEnded trains synchronously after the runner's RoundEndedEvent,
so the counter read right after awaitResults() is the stale pre-train
value and System.exit killed the bot inside ppoUpdate. Poll up to 60s
for the counter to catch up before declaring the battle incomplete.

Verified: 72 consecutive rounds vs Fire, 100% wins, all rounds trained
(counter advanced 1:1), zero coredumps since the fix.
2026-08-19 03:12:22 +02:00
SirStone 766b9e03ee feat(PPO_Bot): enemy-centered action space + reward shaping for 100% vs Target
- actions.nim: goto/aimTo coordinates now offset from enemy position
  (enemyX + tanh(raw) * scale) instead of absolute arena coords
  (sigmoid(raw) * arenaSize). Initial random policy defaults to
  approaching and aiming at enemy.
- training.nim: added dense reward shaping (distance closeness +
  gun bearing) to computeTickReward, doubled round reward scaling.
- PPO_Bot.nim: passes enemy position to mapActions, computes
  gun-to-enemy bearing for reward shaping.

Result: 100/100 win rate vs Target with frozen weights.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-18 17:40:59 +02:00
SirStone 0b17430735 feat(PPO_Bot): persist Adam optimizer state and round counter across restarts (#35)
Save ACAdamStates (m/v tensors + t counters) as .npy files alongside
network weights in latest/ and checkpoint dirs; save round counter to
round_counter.txt. loadBestAvailable restores both on startup; fresh
start works unchanged when files are absent.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-18 15:40:54 +02:00
SirStone cdde60d79f feat(PPO_Bot): command abstraction layer — goto/aimTo controllers (#24)
- Add gotoTick/aimToTick controller functions (#25)
- Update network dims: actor 5→6, state 42→44 (#26)
- Rewrite mapActions for 6-dim command space (#27)
- Delete stale weight files (shape mismatch)
- Fix existing tests for new signatures

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-17 19:20:40 +02:00