Commit Graph

14 Commits

Author SHA1 Message Date
SirStone 2097c2e6fa chore: gitignore binaries/artifacts, remove drafts and stale files 2026-08-27 18:28:21 +02:00
SirStone b509195ee9 chore: rename libs→common_libs, all bot dirs to _garage suffix, fix all path refs
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-27 18:18:41 +02:00
SirStone df256b4d3e feat(SAC_LSTM_Bot): training harness (#49)
sac_train.sh orchestrates chunked self-play via tools/training_runner/
RunTraining.java: weighted opponent sampling per chunk, deterministic
eval (SACLSTM_EVAL_MODE=1) every N chunks with win-rate tracking, best
checkpoint (weights/sac_best.zip) by eval score, crash-restart loop on
the runner's liveness detection.

Supporting changes:
- integration.nim: opponentKey() keys the NewBattle buffer-clear rule on
  getBotName(id) with numeric-id fallback (#49 Q14 follow-up);
  bumpRoundCounter() emits the per-round liveness signal.
- SAC_LSTM_Bot.nim: onRoundEnded -> bumpRoundCounter().
- RunTraining.java: BOT_NAME env parameterizes result matching
  (default PPO_Bot, unchanged behavior for PPO).
- Launch packaging: root SAC_LSTM_Bot.json + .sh for the booter;
  src json name aligned to 'SAC_LSTM_Bot' so self-reported identity
  matches the booted identity (mismatch = runner connect timeout).
2026-08-21 21:51:27 +02:00
SirStone ca3e3d2272 tune(PPO_Bot): logStd=-2.0 (std≈0.135), entropy=0, ceiling=-1.0
Stochastic eval at std≈0.37 was 0/10 vs Corners (deterministic: 10/10).
Warm-start policy is correct but brittle — any noise breaks it.
- log_std initialized to -2.0 (std≈0.135) for moderate exploration
- entropy_coeff=0.0 (no push toward exploration during fine-tuning)
- logStd ceiling=-1.0 (cap at std≈0.37)
2026-08-20 15:28:15 +02:00
SirStone 0d35646dc9 feat(PPO_Bot): deterministic eval + fix logStd warm-start
- actorForward: deterministic param, uses mean-only when PPOB_EVAL_ONLY=1
  (eval was adding unit Gaussian noise to every action — unreliable scores)
- warm_start.py: log_std initialized to -1.0 (std≈0.37) instead of copying
  snapshot values (were 2.27-4.68 → std 9-108, completely drowning signal)
- training.env: LOG_STD_CEILING 0.0→-0.5 (cap exploration at std≈0.6)
2026-08-20 15:18:07 +02:00
SirStone fedab54bc0 feat(PPO_Bot): multi-round transition accumulation (UPDATE_INTERVAL=10)
- Accumulate transitions across 10 rounds (~3000) before PPO update
  (was per-round ~300 — gradient estimates were far too noisy)
- training.nim: MAX_TRANSITIONS 4096→8192, done flag on transitions,
  GAE handles episode boundaries correctly
- PPO_Bot.nim: buffer persists across rounds, update every N rounds
- training.env: lr 5e-5→1e-4, entropy 0.001, UPDATE_INTERVAL=10
2026-08-20 15:06:00 +02:00
SirStone 82eeb53e5c tune(training): 20-round chunks, 50 eval rounds, 30k total rounds
- generalist_train.sh: CHUNK_SIZE 60→20 for faster opponent cycling
- generalist_train.sh: EVAL_ROUNDS 30→50 for more reliable eval
- training.env: TRAINING_ROUNDS→30000, removed hardcoded opponent
- warm_start.py: TARGET_DIM=57 (already committed, ensure latest)
2026-08-20 14:38:09 +02:00
SirStone 12624d3069 feat(PPO_Bot): bot-relative bullets + scan staleness (STATE_DIM=57)
- Bullet state (indices 44-55): enemy-relative → bot-relative frame
  (bot needs threat vectors to itself for dodging, not to enemy)
- New index 56: scan staleness = min(ticksSinceLastScan / 30, 1.0)
  (gives policy a confidence signal for enemy data freshness)
- warm_start.py updated: 44→57 dim expansion, TARGET_DIM variable
- Tests updated for new state layout

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 14:27:28 +02:00
SirStone 6ad51148f4 fix(PPO_Bot): SIGSEGV crash fixes + static buffers for thread safety
- bullets: seq[InFlightBullet] → array[4, InFlightBullet] + bulletCount
  (eliminates cross-thread heap realloc under ORC)
- hasFired: edge-triggered (cleared after state build, not level-triggered)
- round_counter parseInt: wrapped for empty/torn file → 0
- Static SVG + intent buffers to kill cross-thread heap realloc
- Tick-local alive/bulletData also fixed arrays

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 14:23:08 +02:00
SirStone 75e32e3315 PPO_Bot Fire campaign: 3787/3789 wins (99.95%); frozen eval 500/500
Single 3789-round battle (5212-9000), only rounds 1-2 lost (cold start);
last 100 rounds 100%. Frozen-policy eval (PPOB_EVAL_ONLY=1, 500 rounds vs
Fire, counter 9000->9500): 500/500 = 100%, zero train lines, weights mtime
and content untouched. vLoss avg 39.7/max 361 stable throughout — bounded
terminal reward (db99153) holds at cumulative score ~462k.
2026-08-19 04:00:47 +02:00
SirStone db99153f65 cap terminal reward scale (bounded score bonus) and restore single-battle campaigns
computeRoundReward used cumulative totalScore/50 — unbounded in long battles
(vLoss 353 at round 3160 → 25745 by 3871 in the 5841-round attempt). Cap the
score term at 400 before /50: bonus ∈ [0,8], so the critic's value scale stays
stable regardless of battle length and across battle boundaries.

Reverts the 60-round battle chunking (186e005/da2f825): one battle per
campaign for the whole remaining budget; keeps the crash-restart loop, the
mid-battle freeze guard and the end-of-battle counter completeness check.

Cert (5211, single 60-round battle): 59/60 wins (sole loss = cold-start round
1, score 61), vLoss avg 36.1 / max 148.5, gNorm max 596, zero NaN, zero
restarts, counter check passed. Weights persist to round 5211.
2026-08-19 03:50:14 +02:00
SirStone da2f825ad8 fix(training): keep run.sh looping across 60-round battle chunks
RunTraining exits 0 after each chunk; && break ended the whole run after
the first battle (counter 3219, not 9000). Loop now falls through the
success path and re-checks the persisted counter each iteration.
2026-08-19 03:36:20 +02:00
SirStone 186e005a96 fix(training): cap battles at 60 rounds to bound round-end reward scale
Round-end reward = cumulative totalScore/50 grows unboundedly with battle
length; long battles (5841 rounds) blew the critic's value scale: vLoss
10-30 during the 60-round cert, 353 at battle-1 round 1, 25745 by round 3871,
policy drift to 0/6 wins. 60-round battles reproduce the certified regime:
bounded value targets, fresh bot process per battle (clears thread state).
2026-08-19 03:32:31 +02:00
SirStone 64697f917e fix(botapi): static event queue storage + end-of-battle train wait
The event queue's heap seq was the last GC'd block surviving across
rounds: each round runs on a freshly spawned bot thread, so the N+1
thread realloc'd a block grown by dead thread N's allocator mid-round
(at the next capacity doubling, ~turn 104) -> rawDealloc SIGSEGV in
addEvent (7 gdb-confirmed coredumps). Replace with a static
array[MAX_QUEUE_SIZE, BotEvent] + eventsLen: no heap block crosses
threads, realloc can never happen.

Also fix the harness aborting the final round mid-train: PPO_Bot's
onRoundEnded trains synchronously after the runner's RoundEndedEvent,
so the counter read right after awaitResults() is the stale pre-train
value and System.exit killed the bot inside ppoUpdate. Poll up to 60s
for the counter to catch up before declaring the battle incomplete.

Verified: 72 consecutive rounds vs Fire, 100% wins, all rounds trained
(counter advanced 1:1), zero coredumps since the fix.
2026-08-19 03:12:22 +02:00