39 Commits

Author SHA1 Message Date
SirStone 96065597c0 fix(divergence): guard reward-normalizer cold start; drop arctanh recovery
Welford variance-collapse divided by 1e-8 producing bit-exact +/−5e8 /
+−1.25e8 poisoned rewards into TD targets; guard skips normalization
until stats meaningful; tanh-inversion removal bounds log-prob path.

Fixes #60.
2026-08-24 01:03:04 +02:00
SirStone eae6fc15a2 research(Evo_Bot): GA/ES parameter recommendations for ~750-weight neuroevolution (#65)
Extracted concrete numbers from 13 papers in docs/papers/neuroevolution/.
Key findings: use CMA-ES or mutation-only truncation GA, mutate ALL weights
(not 5%), sigma=0.005-0.01, pop=64-200, no crossover, single elite.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-08-23 23:32:10 +02:00
SirStone 81718e3a4c fix(SAC_LSTM_Bot): un-invert dashboard axes — orientation selftest added 2026-08-23 10:47:48 +02:00
SirStone 04c149ea28 refactor(SAC_LSTM_Bot): reconcile graph tooling — dashboard-only output 2026-08-23 10:30:16 +02:00
SirStone 533a146342 feat(SAC_LSTM_Bot): live campaign dashboard — five panels, auto-reload, run-1 comparison dropped 2026-08-23 10:25:51 +02:00
SirStone 03a1853e57 feat(SAC_LSTM_Bot): embedded plain-English reading guides in graphs 2026-08-23 10:11:44 +02:00
SirStone 002a568c7e fix(SAC_LSTM_Bot): atomic+validated SVG output for progress graphs 2026-08-23 10:06:05 +02:00
SirStone f3ac0888bb feat(SAC_LSTM_Bot): readable eval chart — trends primary, raw dots secondary, v1 comparison separated 2026-08-23 09:59:30 +02:00
SirStone 0b0933294a feat(SAC_LSTM_Bot): progress graph tooling + first graphs 2026-08-23 09:08:53 +02:00
SirStone 781e41595e docs(SAC_LSTM_Bot): simple-words story + dictionary for readability 2026-08-23 08:48:37 +02:00
SirStone 40e074e5a3 docs(SAC_LSTM_Bot): lever-5 trigger + attempt-3 chapter — MA-wipe root cause, LR_CRITIC=1e-4 decision, launch health (#59 #60) 2026-08-23 07:11:09 +02:00
SirStone 167bcc4ce5 fix(SAC_LSTM_Bot): MA history self-truncation — read before > redirect; campaign v2 attempt-3 prep (#60)
$(cat f) inside a command redirected to f saw the already-truncated file,
so every eval cycle wiped ma_history_*.txt back to one leading-space value
and degraded the composite best-gate to last-cycle mean. Read is hoisted
into its own statement; unquoted expansion + tail -n keeps exactly the last
MA_WINDOW values. Verified: fresh/5+/6+ cycle edges reproduce sliding window.
2026-08-23 07:02:36 +02:00
SirStone 1619b86f25 docs(SAC_LSTM_Bot): v2 restart saga, twin-freeze correction, pacing decisions 2026-08-22 21:26:24 +02:00
SirStone 19f34abf0c fix(SAC_LSTM_Bot): freeze mirror-twin via eval-mode gate in twin launcher
SACLSTM_EVAL_MODE=1 in SacTwin.sh suppresses all sendTrainingMsg traffic
(lever-4 gate), so the twin never trains — not even in-RAM within a battle.
Required now that the main bot's SACLSTM_SAVE_INTERVAL drops to 1 (v2 relaunch
after checkpoint-cadence diagnosis): without the gate the twin would persist
per-battle drift and stop being the frozen reproducible opponent #54 specifies.
2026-08-22 21:09:07 +02:00
SirStone bd58794b4c docs(SAC_LSTM_Bot): campaign v2 chapter — locked config, fresh-start archive, twin reseed, ceiling net, launch health evidence (#59) 2026-08-22 20:31:33 +02:00
SirStone 6fc01eb4e5 feat(SAC_LSTM_Bot): campaign v2 levers — aggression/anti-ram reward shaping + stability knob overrides (part 2)
Lever 2 (#59): x1.25 aggression mult on damage dealt, flat +0.5 hit bonus,
-3.0 per bot-bot collision (server deals RAM_DAMAGE=0.6 to both parties but
only notifies the hitter), escalating proximity deterrent below 12% arena
diagonal suppressed while dealing damage. Win/loss terminals unchanged and
dominant. All weights TUNABLE consts marked ponytail. SACLSTM_REWARD_DEBUG=1
env-gated reward_debug.log for calibration greps.

Lever 5 (#59): no code needed — SACLSTM_LR_ACTOR/LR_CRITIC/LR_ALPHA (3e-4)
and SACLSTM_TARGET_ENTROPY (-4.0) were already env-overridable in training.nim.

Smoke vs RamFire+Crazy (hidden=32, random init, isolated weights): 75 ram
penalties, 381 charge events, hit bonuses firing, 0 crashes, metrics JSONL
flowing. Tests: 8/8 suites green incl. new assert-level term math.
2026-08-22 19:44:58 +02:00
SirStone a07e5305f5 feat(SAC_LSTM_Bot): campaign v2 levers — loss metrics, eval-mode gate, eval rotation + MA gating (part 1)
Levers 3, 4, 1 of the #57 sign-off (execution order 3->4->1), tracked in #59.

- Lever 3 (#59): one JSONL line per trainPass in training_metrics.jsonl with
  exactly the scalars sacUpdate already exposes (SACMetrics: critic/actor/alpha
  losses + alpha, averaged per pass) plus epoch, buffer size (replay_buffer.len),
  cumulative steps and drained count. No trainer change needed.
- Lever 4 (#59): sendTrainingMsg drops all training input while SACLSTM_EVAL_MODE=1
  (existing #49 harness mechanism) — eval battles can neither pollute the replay
  buffer nor trigger gradient updates; one-time stderr notice at bot init.
- Lever 1 (#59): sac_train.sh evaluates every SAC_EVAL_OPPONENTS entry per cycle
  (results carry opponent name in eval_log.jsonl); best-gating now uses a
  composite = mean over opponents of the last-5-evals moving average per
  opponent. best_score.txt format change: float composite replaces the
  single-opponent integer win rate semantics (retired).
- Tests: metricsLine JSONL scalars + eval-mode suppression asserts.

Refs: #59, #57
2026-08-22 18:58:47 +02:00
SirStone 2619ba06fc docs(SAC_LSTM_Bot): campaign ending, results table, verdict, hygiene (#57) 2026-08-22 14:40:57 +02:00
SirStone f45e8f2717 fix(SAC_LSTM_Bot): startup sweep of stale .part checkpoint corpses (#57) 2026-08-22 14:40:57 +02:00
SirStone b7492f1080 docs(SAC_LSTM_Bot): notebook — score:60 anatomy, .part forensics+sweep, ceiling defused 2026-08-22 08:08:31 +02:00
SirStone f1962c7506 docs(SAC_LSTM_Bot): backfill shift-1 milestone entry + morning audit 2026-08-22 07:27:37 +02:00
SirStone 05929d2dbd docs(SAC_LSTM_Bot): notebook — shift 1, deferred-intervention milestone 2026-08-22 07:14:35 +02:00
SirStone 4b64bf18ac docs(SAC_LSTM_Bot): campaign-v1 launch record — save-check incident, health check, check-in procedure (#56) 2026-08-22 00:31:47 +02:00
SirStone 26536713ba fix(SAC_LSTM_Bot): checkpoint save check inside gradient-step loop (#56)
Launch finding during campaign-v1 verification: at production sizes
(hidden 256, ~1s/step, ~13s trainer CPU per ~40s chunk process) the
save check ran only between drain-burst passes, so stepCount never
crossed nextSave before the process died — zero checkpoints persisted
across entire runs (masked at #49/#54 smoke sizes where steps were
sub-millisecond). Check now fires mid-loop; with SAVE_INTERVAL<=5
(within the per-process step budget) every chunk persists its chain.
2026-08-22 00:21:43 +02:00
SirStone 2f49cb243f docs(SAC_LSTM_Bot): campaign-v1 notebook — locked config + story so far (#56) 2026-08-21 23:26:31 +02:00
SirStone 6a294ad7ad feat(SAC_LSTM_Bot): mirror-twin sparring partner + readiness check (#54)
- make_twin.sh: generates self-contained SacTwin dir in the sample-bots
  archive (own json/sh identity, own weights dir seeded from a frozen
  sac_best.zip copy, own round_counter) so RunTraining.java resolves it
  like any sample bot; re-running resets the twin to the frozen baseline.
- src/SAC_LSTM_Bot.nim: SACLSTM_BOT_JSON env overrides the baked-in bot
  json (loadBotInfo gives json total precedence, #49) so the same binary
  boots under the twin's name.
- sac_train.sh: chunk loop is a while, not for-over-seq — a crash on the
  FINAL chunk previously fell through ((chunk--);continue on an exhausted
  seq list) and exited 0 with budget incomplete; observed live vs SacTwin.

Readiness dry-run (#54): weighted pool Corners:1,SacTwin:3 picked the twin
in 3/4 chunks; all battles counter-checked; deterministic eval parsed;
main sac_best.zip/counter untouched by twin (twin counter advanced
independently); crash-restart proven end-to-end incl. final-chunk retry.
2026-08-21 23:19:27 +02:00
SirStone edf26aa45d fix(SAC_LSTM_Bot): enforce MaxHidden cap on SACLSTM_HIDDEN_SIZE (review of #48/#49) 2026-08-21 22:17:55 +02:00
SirStone df256b4d3e feat(SAC_LSTM_Bot): training harness (#49)
sac_train.sh orchestrates chunked self-play via tools/training_runner/
RunTraining.java: weighted opponent sampling per chunk, deterministic
eval (SACLSTM_EVAL_MODE=1) every N chunks with win-rate tracking, best
checkpoint (weights/sac_best.zip) by eval score, crash-restart loop on
the runner's liveness detection.

Supporting changes:
- integration.nim: opponentKey() keys the NewBattle buffer-clear rule on
  getBotName(id) with numeric-id fallback (#49 Q14 follow-up);
  bumpRoundCounter() emits the per-round liveness signal.
- SAC_LSTM_Bot.nim: onRoundEnded -> bumpRoundCounter().
- RunTraining.java: BOT_NAME env parameterizes result matching
  (default PPO_Bot, unchanged behavior for PPO).
- Launch packaging: root SAC_LSTM_Bot.json + .sh for the booter;
  src json name aligned to 'SAC_LSTM_Bot' so self-reported identity
  matches the booted identity (mismatch = runner connect timeout).
2026-08-21 21:51:27 +02:00
SirStone 7104645f5d chore(deps): update vendored tankroyale botapi to v1.0.1
Syncs libs/tankroyale_botapi with SirStone/robocode_tankroyale_botapi
v1.0.1 (extracted from tank-royale nim branch @ 03195a814). The local
SIGSEGV fixes (static event queue/SVG/intent buffers) were already
ported upstream in issue #24 — content is otherwise identical.

New capability (upstream #23): opponent name exposure for ticket #49.
- bot.nim: gBotNames id→name table + getBotName(id) / updateBotNames()
- umbrella module: dispatch BotListUpdate messages to updateBotNames()

Vendored layout and wiring unchanged (--path via config.nims, module
name stays tankroyale_botapi).
2026-08-21 21:03:17 +02:00
SirStone 32b71d9fc8 feat(SAC_LSTM_Bot): main bot integration (#48) 2026-08-21 20:19:05 +02:00
SirStone 62a6cc8ccf fix(SAC_LSTM_Bot): training review fixes — hidden state ordering, redundant forwards, actor grad clip (#47)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-21 00:16:47 +02:00
SirStone 415d4e3738 feat(SAC_LSTM_Bot): SAC training module (#47)
Implements sacUpdate with burn-in LSTM warm-up, twin-critic TD update,
actor reparameterization gradient, auto-alpha, and soft target update.
Manual backprop (linear + LSTM single-step, truncated BPTT). 11 new tests
all green; full regression suite (57+ tests) unaffected.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-21 00:11:35 +02:00
SirStone 717ef3ead8 Merge branch 'worktree-agent-a8622248' (ticket #46 weight persistence) 2026-08-20 23:57:53 +02:00
SirStone 54b8139b11 feat(SAC_LSTM_Bot): weight persistence module (#46)
Save/load all SAC-LSTM tensors (actor, 2 critics, 2 target critics,
alpha, Adam states) into a single .zip of .npy files. Atomic write
via temp path + rename. Adam types (AdamVar, SACAdamStates) defined
here for training.nim to use.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:57:13 +02:00
SirStone 55ef22ff8b Merge branch 'worktree-agent-a3ca3066' (ticket #45 replay buffer) 2026-08-20 23:41:23 +02:00
SirStone 5ea57bcae3 Merge branch 'worktree-agent-a393a8b9' (ticket #44 reward module) 2026-08-20 23:41:23 +02:00
SirStone cb33551621 feat(SAC_LSTM_Bot): replay buffer module (#45)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:40:17 +02:00
SirStone 4ee0d8272c feat(SAC_LSTM_Bot): action mapping module (#43)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:38:45 +02:00
SirStone add3e34926 feat(SAC_LSTM_Bot): LSTM network module (#41)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:38:17 +02:00
29 changed files with 5075 additions and 32 deletions
+11
View File
@@ -0,0 +1,11 @@
{
"name": "SAC_LSTM_Bot",
"version": "0.1.0",
"authors": ["Davide Cappellini"],
"description": "SAC+LSTM-trained Tank Royale bot (#37) — training/eval launch config",
"homepage": "",
"countryCodes": ["IT"],
"gameTypes": ["classic", "melee", "1v1"],
"platform": "Nim",
"programmingLang": "Nim"
}
+4
View File
@@ -0,0 +1,4 @@
#!/bin/sh
# Launch config for tools/training_runner/RunTraining.java (#49): the runner
# executes <json-basename>.sh inside the bot dir (sample-bots convention).
exec "$(dirname "$0")/SAC_LSTM_Bot"
+6
View File
@@ -1,7 +1,13 @@
# Static-link OpenBLAS for portable deployment # Static-link OpenBLAS for portable deployment
# ponytail: adjust path per machine, or use pkg-config # ponytail: adjust path per machine, or use pkg-config
switch("passL", "-L/nix/store/v07svn2y92bvzjl51aj7c9ca1cwg7rw7-openblas-0.3.32/lib -lopenblas") switch("passL", "-L/nix/store/v07svn2y92bvzjl51aj7c9ca1cwg7rw7-openblas-0.3.32/lib -lopenblas")
# libzip for weight checkpoint zip files
# ponytail: nix store path; adjust per machine, or use pkg-config
switch("passL", "-L/nix/store/wqvz31s598bvj3zb747943xhl38hjc6h-libzip-1.11.4/lib -lzip")
switch("threads", "on") switch("threads", "on")
# Submodules import each other as SAC_LSTM_Bot/<mod>; make that resolvable for
# the binary build too (tests already add ../src via tests/config.nims).
switch("path", thisDir() & "/src")
# begin Nimble config (version 2) # begin Nimble config (version 2)
when withDir(thisDir(), system.fileExists("nimble.paths")): when withDir(thisDir(), system.fileExists("nimble.paths")):
include "nimble.paths" include "nimble.paths"
+743
View File
@@ -0,0 +1,743 @@
<svg xmlns="http://www.w3.org/2000/svg" width="1400" height="2000" viewBox="0 0 1400 2000" font-family="sans-serif">
<rect width="1400" height="2000" fill="white"/>
<text x="700" y="32" text-anchor="middle" font-size="21" font-weight="bold">SAC-LSTM campaign dashboard - live run (current only)</text>
<text x="700" y="56" text-anchor="middle" font-size="12" fill="#555">generated 2026-08-23 10:46:52 - auto-reloads every 60 s (open this file in Chrome)</text>
<text x="70" y="100" font-size="15" font-weight="bold">Test matches - win % vs opponents</text>
<text x="70" y="115" font-size="11" fill="#555">raw dots = single test matches, thick = rolling-mean-10</text>
<line x1="70" y1="720.0" x2="697" y2="720.0" stroke="#dddddd"/>
<line x1="70" y1="600.0" x2="697" y2="600.0" stroke="#dddddd"/>
<line x1="70" y1="480.0" x2="697" y2="480.0" stroke="#dddddd"/>
<line x1="70" y1="360.0" x2="697" y2="360.0" stroke="#dddddd"/>
<line x1="70" y1="240.0" x2="697" y2="240.0" stroke="#dddddd"/>
<line x1="70" y1="120.0" x2="697" y2="120.0" stroke="#dddddd"/>
<line x1="70" y1="720" x2="697" y2="720" stroke="black"/>
<line x1="70" y1="720" x2="70" y2="120" stroke="black"/>
<line x1="128.2" y1="720" x2="128.2" y2="724" stroke="black"/>
<text x="128.2" y="737" text-anchor="middle" font-size="11">10</text>
<line x1="192.8" y1="720" x2="192.8" y2="724" stroke="black"/>
<text x="192.8" y="737" text-anchor="middle" font-size="11">20</text>
<line x1="257.5" y1="720" x2="257.5" y2="724" stroke="black"/>
<text x="257.5" y="737" text-anchor="middle" font-size="11">30</text>
<line x1="322.1" y1="720" x2="322.1" y2="724" stroke="black"/>
<text x="322.1" y="737" text-anchor="middle" font-size="11">40</text>
<line x1="386.7" y1="720" x2="386.7" y2="724" stroke="black"/>
<text x="386.7" y="737" text-anchor="middle" font-size="11">50</text>
<line x1="451.4" y1="720" x2="451.4" y2="724" stroke="black"/>
<text x="451.4" y="737" text-anchor="middle" font-size="11">60</text>
<line x1="516.0" y1="720" x2="516.0" y2="724" stroke="black"/>
<text x="516.0" y="737" text-anchor="middle" font-size="11">70</text>
<line x1="580.6" y1="720" x2="580.6" y2="724" stroke="black"/>
<text x="580.6" y="737" text-anchor="middle" font-size="11">80</text>
<line x1="645.3" y1="720" x2="645.3" y2="724" stroke="black"/>
<text x="645.3" y="737" text-anchor="middle" font-size="11">90</text>
<line x1="66" y1="720.0" x2="70" y2="720.0" stroke="black"/>
<text x="63" y="724.0" text-anchor="end" font-size="11">0</text>
<line x1="66" y1="600.0" x2="70" y2="600.0" stroke="black"/>
<text x="63" y="604.0" text-anchor="end" font-size="11">20</text>
<line x1="66" y1="480.0" x2="70" y2="480.0" stroke="black"/>
<text x="63" y="484.0" text-anchor="end" font-size="11">40</text>
<line x1="66" y1="360.0" x2="70" y2="360.0" stroke="black"/>
<text x="63" y="364.0" text-anchor="end" font-size="11">60</text>
<line x1="66" y1="240.0" x2="70" y2="240.0" stroke="black"/>
<text x="63" y="244.0" text-anchor="end" font-size="11">80</text>
<line x1="66" y1="120.0" x2="70" y2="120.0" stroke="black"/>
<text x="63" y="124.0" text-anchor="end" font-size="11">100</text>
<text x="383" y="753" text-anchor="middle" font-size="12">test match number (each opponent)</text>
<text x="16" y="420" text-anchor="middle" font-size="12" transform="rotate(-90 16 420)">win rate (%)</text>
<circle cx="70.0" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="76.5" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="82.9" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="89.4" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="95.9" cy="540.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="102.3" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="108.8" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="115.2" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="121.7" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="128.2" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="134.6" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="141.1" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="147.6" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="154.0" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="160.5" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="167.0" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="173.4" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="179.9" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="186.4" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="192.8" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="199.3" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="205.7" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="212.2" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="218.7" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="225.1" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="231.6" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="238.1" cy="660.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="244.5" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="251.0" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="257.5" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="263.9" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="270.4" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="276.8" cy="660.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="283.3" cy="660.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="289.8" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="296.2" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="302.7" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="309.2" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="315.6" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="322.1" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="328.6" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="335.0" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="341.5" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="347.9" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="354.4" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="360.9" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="367.3" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="373.8" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="380.3" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="386.7" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="393.2" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="399.7" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="406.1" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="412.6" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="419.1" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="425.5" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="432.0" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="438.4" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="444.9" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="451.4" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="457.8" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="464.3" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="470.8" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="477.2" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="483.7" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="490.2" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="496.6" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="503.1" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="509.5" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="516.0" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="522.5" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="528.9" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="535.4" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="541.9" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="548.3" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="554.8" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="561.3" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="567.7" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="574.2" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="580.6" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="587.1" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="593.6" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="600.0" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="606.5" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="613.0" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="619.4" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="625.9" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="632.4" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="638.8" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="645.3" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="651.8" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="658.2" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="664.7" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="671.1" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="677.6" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="684.1" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="690.5" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<circle cx="697.0" cy="720.0" r="2" fill="#d62728" opacity="0.25"/>
<polyline points="70.0,720.0 76.5,720.0 82.9,720.0 89.4,720.0 95.9,684.0 102.3,690.0 108.8,694.3 115.2,697.5 121.7,700.0 128.2,702.0 134.6,702.0 141.1,702.0 147.6,702.0 154.0,702.0 160.5,720.0 167.0,720.0 173.4,720.0 179.9,720.0 186.4,720.0 192.8,720.0 199.3,720.0 205.7,720.0 212.2,720.0 218.7,720.0 225.1,720.0 231.6,720.0 238.1,714.0 244.5,714.0 251.0,714.0 257.5,714.0 263.9,714.0 270.4,714.0 276.8,708.0 283.3,702.0 289.8,702.0 296.2,702.0 302.7,708.0 309.2,708.0 315.6,708.0 322.1,708.0 328.6,708.0 335.0,708.0 341.5,714.0 347.9,720.0 354.4,720.0 360.9,720.0 367.3,720.0 373.8,720.0 380.3,720.0 386.7,720.0 393.2,720.0 399.7,720.0 406.1,720.0 412.6,720.0 419.1,720.0 425.5,720.0 432.0,720.0 438.4,720.0 444.9,720.0 451.4,720.0 457.8,720.0 464.3,720.0 470.8,720.0 477.2,720.0 483.7,720.0 490.2,720.0 496.6,720.0 503.1,720.0 509.5,720.0 516.0,720.0 522.5,720.0 528.9,720.0 535.4,720.0 541.9,720.0 548.3,720.0 554.8,720.0 561.3,720.0 567.7,720.0 574.2,720.0 580.6,720.0 587.1,720.0 593.6,720.0 600.0,720.0 606.5,720.0 613.0,720.0 619.4,720.0 625.9,720.0 632.4,720.0 638.8,720.0 645.3,720.0 651.8,720.0 658.2,720.0 664.7,720.0 671.1,720.0 677.6,720.0 684.1,720.0 690.5,720.0 697.0,720.0" fill="none" stroke="#d62728" stroke-width="3.5" opacity="1.0"/>
<circle cx="70.0" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="76.5" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="82.9" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="89.4" cy="180.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="95.9" cy="480.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="102.3" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="108.8" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="115.2" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="121.7" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="128.2" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="134.6" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="141.1" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="147.6" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="154.0" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="160.5" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="167.0" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="173.4" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="179.9" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="186.4" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="192.8" cy="360.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="199.3" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="205.7" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="212.2" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="218.7" cy="120.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="225.1" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="231.6" cy="600.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="238.1" cy="600.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="244.5" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="251.0" cy="120.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="257.5" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="263.9" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="270.4" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="276.8" cy="540.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="283.3" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="289.8" cy="600.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="296.2" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="302.7" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="309.2" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="315.6" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="322.1" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="328.6" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="335.0" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="341.5" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="347.9" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="354.4" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="360.9" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="367.3" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="373.8" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="380.3" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="386.7" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="393.2" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="399.7" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="406.1" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="412.6" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="419.1" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="425.5" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="432.0" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="438.4" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="444.9" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="451.4" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="457.8" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="464.3" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="470.8" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="477.2" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="483.7" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="490.2" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="496.6" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="503.1" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="509.5" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="516.0" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="522.5" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="528.9" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="535.4" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="541.9" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="548.3" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="554.8" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="561.3" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="567.7" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="574.2" cy="540.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="580.6" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="587.1" cy="600.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="593.6" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="600.0" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="606.5" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="613.0" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="619.4" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="625.9" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="632.4" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="638.8" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="645.3" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="651.8" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="658.2" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="664.7" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="671.1" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="677.6" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="684.1" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="690.5" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<circle cx="697.0" cy="720.0" r="2" fill="#1f77b4" opacity="0.25"/>
<polyline points="70.0,720.0 76.5,720.0 82.9,720.0 89.4,585.0 95.9,564.0 102.3,590.0 108.8,608.6 115.2,622.5 121.7,633.3 128.2,642.0 134.6,642.0 141.1,642.0 147.6,642.0 154.0,696.0 160.5,720.0 167.0,720.0 173.4,720.0 179.9,720.0 186.4,720.0 192.8,684.0 199.3,684.0 205.7,684.0 212.2,684.0 218.7,624.0 225.1,624.0 231.6,612.0 238.1,600.0 244.5,600.0 251.0,540.0 257.5,576.0 263.9,576.0 270.4,576.0 276.8,558.0 283.3,618.0 289.8,606.0 296.2,618.0 302.7,630.0 309.2,630.0 315.6,690.0 322.1,690.0 328.6,690.0 335.0,690.0 341.5,708.0 347.9,708.0 354.4,720.0 360.9,720.0 367.3,720.0 373.8,720.0 380.3,720.0 386.7,720.0 393.2,720.0 399.7,720.0 406.1,720.0 412.6,720.0 419.1,720.0 425.5,720.0 432.0,720.0 438.4,720.0 444.9,720.0 451.4,720.0 457.8,720.0 464.3,720.0 470.8,720.0 477.2,720.0 483.7,720.0 490.2,720.0 496.6,720.0 503.1,720.0 509.5,720.0 516.0,720.0 522.5,720.0 528.9,720.0 535.4,720.0 541.9,720.0 548.3,720.0 554.8,720.0 561.3,720.0 567.7,720.0 574.2,702.0 580.6,702.0 587.1,690.0 593.6,690.0 600.0,690.0 606.5,690.0 613.0,690.0 619.4,690.0 625.9,690.0 632.4,690.0 638.8,708.0 645.3,708.0 651.8,720.0 658.2,720.0 664.7,720.0 671.1,720.0 677.6,720.0 684.1,720.0 690.5,720.0 697.0,720.0" fill="none" stroke="#1f77b4" stroke-width="3.5" opacity="1.0"/>
<circle cx="70.0" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="76.5" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="82.9" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="89.4" cy="600.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="95.9" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="102.3" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="108.8" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="115.2" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="121.7" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="128.2" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="134.6" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="141.1" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="147.6" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="154.0" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="160.5" cy="660.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="167.0" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="173.4" cy="240.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="179.9" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="186.4" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="192.8" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="199.3" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="205.7" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="212.2" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="218.7" cy="660.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="225.1" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="231.6" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="238.1" cy="360.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="244.5" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="251.0" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="257.5" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="263.9" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="270.4" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="276.8" cy="660.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="283.3" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="289.8" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="296.2" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="302.7" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="309.2" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="315.6" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="322.1" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="328.6" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="335.0" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="341.5" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="347.9" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="354.4" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="360.9" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="367.3" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="373.8" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="380.3" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="386.7" cy="660.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="393.2" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="399.7" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="406.1" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="412.6" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="419.1" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="425.5" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="432.0" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="438.4" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="444.9" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="451.4" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="457.8" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="464.3" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="470.8" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="477.2" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="483.7" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="490.2" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="496.6" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="503.1" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="509.5" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="516.0" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="522.5" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="528.9" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="535.4" cy="660.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="541.9" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="548.3" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="554.8" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="561.3" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="567.7" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="574.2" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="580.6" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="587.1" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="593.6" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="600.0" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="606.5" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="613.0" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="619.4" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="625.9" cy="660.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="632.4" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="638.8" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="645.3" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="651.8" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="658.2" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="664.7" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="671.1" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="677.6" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="684.1" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="690.5" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<circle cx="697.0" cy="720.0" r="2" fill="#2ca02c" opacity="0.25"/>
<polyline points="70.0,720.0 76.5,720.0 82.9,720.0 89.4,690.0 95.9,696.0 102.3,700.0 108.8,702.9 115.2,705.0 121.7,706.7 128.2,708.0 134.6,708.0 141.1,708.0 147.6,708.0 154.0,720.0 160.5,714.0 167.0,714.0 173.4,666.0 179.9,666.0 186.4,666.0 192.8,666.0 199.3,666.0 205.7,666.0 212.2,666.0 218.7,660.0 225.1,666.0 231.6,666.0 238.1,678.0 244.5,678.0 251.0,678.0 257.5,678.0 263.9,678.0 270.4,678.0 276.8,672.0 283.3,678.0 289.8,678.0 296.2,678.0 302.7,714.0 309.2,714.0 315.6,714.0 322.1,714.0 328.6,714.0 335.0,714.0 341.5,720.0 347.9,720.0 354.4,720.0 360.9,720.0 367.3,720.0 373.8,720.0 380.3,720.0 386.7,714.0 393.2,714.0 399.7,714.0 406.1,714.0 412.6,714.0 419.1,714.0 425.5,714.0 432.0,714.0 438.4,714.0 444.9,714.0 451.4,720.0 457.8,720.0 464.3,720.0 470.8,720.0 477.2,720.0 483.7,720.0 490.2,720.0 496.6,720.0 503.1,720.0 509.5,720.0 516.0,720.0 522.5,720.0 528.9,720.0 535.4,714.0 541.9,714.0 548.3,714.0 554.8,714.0 561.3,714.0 567.7,714.0 574.2,714.0 580.6,714.0 587.1,714.0 593.6,714.0 600.0,720.0 606.5,720.0 613.0,720.0 619.4,720.0 625.9,714.0 632.4,714.0 638.8,714.0 645.3,714.0 651.8,714.0 658.2,714.0 664.7,714.0 671.1,714.0 677.6,714.0 684.1,714.0 690.5,720.0 697.0,720.0" fill="none" stroke="#2ca02c" stroke-width="3.5" opacity="1.0"/>
<line x1="82" y1="772" x2="110" y2="772" stroke="#d62728" stroke-width="3"/>
<text x="116" y="776" font-size="12">Corners - 98 evals</text>
<line x1="82" y1="790" x2="110" y2="790" stroke="#1f77b4" stroke-width="3"/>
<text x="116" y="794" font-size="12">Crazy - 98 evals</text>
<line x1="82" y1="808" x2="110" y2="808" stroke="#2ca02c" stroke-width="3"/>
<text x="116" y="812" font-size="12">Target - 98 evals</text>
<text x="747" y="100" font-size="15" font-weight="bold">Real fights - win % per opponent</text>
<text x="747" y="115" font-size="11" fill="#555">training_log.jsonl only - learning in REAL battles, not tests</text>
<line x1="747" y1="720.0" x2="1375" y2="720.0" stroke="#dddddd"/>
<line x1="747" y1="600.0" x2="1375" y2="600.0" stroke="#dddddd"/>
<line x1="747" y1="480.0" x2="1375" y2="480.0" stroke="#dddddd"/>
<line x1="747" y1="360.0" x2="1375" y2="360.0" stroke="#dddddd"/>
<line x1="747" y1="240.0" x2="1375" y2="240.0" stroke="#dddddd"/>
<line x1="747" y1="120.0" x2="1375" y2="120.0" stroke="#dddddd"/>
<line x1="747" y1="720" x2="1375" y2="720" stroke="black"/>
<line x1="747" y1="720" x2="747" y2="120" stroke="black"/>
<line x1="840.9" y1="720" x2="840.9" y2="724" stroke="black"/>
<text x="840.9" y="737" text-anchor="middle" font-size="11">300</text>
<line x1="935.2" y1="720" x2="935.2" y2="724" stroke="black"/>
<text x="935.2" y="737" text-anchor="middle" font-size="11">600</text>
<line x1="1029.4" y1="720" x2="1029.4" y2="724" stroke="black"/>
<text x="1029.4" y="737" text-anchor="middle" font-size="11">900</text>
<line x1="1123.7" y1="720" x2="1123.7" y2="724" stroke="black"/>
<text x="1123.7" y="737" text-anchor="middle" font-size="11">1200</text>
<line x1="1217.9" y1="720" x2="1217.9" y2="724" stroke="black"/>
<text x="1217.9" y="737" text-anchor="middle" font-size="11">1500</text>
<line x1="1312.2" y1="720" x2="1312.2" y2="724" stroke="black"/>
<text x="1312.2" y="737" text-anchor="middle" font-size="11">1800</text>
<line x1="743" y1="720.0" x2="747" y2="720.0" stroke="black"/>
<text x="740" y="724.0" text-anchor="end" font-size="11">0</text>
<line x1="743" y1="600.0" x2="747" y2="600.0" stroke="black"/>
<text x="740" y="604.0" text-anchor="end" font-size="11">20</text>
<line x1="743" y1="480.0" x2="747" y2="480.0" stroke="black"/>
<text x="740" y="484.0" text-anchor="end" font-size="11">40</text>
<line x1="743" y1="360.0" x2="747" y2="360.0" stroke="black"/>
<text x="740" y="364.0" text-anchor="end" font-size="11">60</text>
<line x1="743" y1="240.0" x2="747" y2="240.0" stroke="black"/>
<text x="740" y="244.0" text-anchor="end" font-size="11">80</text>
<line x1="743" y1="120.0" x2="747" y2="120.0" stroke="black"/>
<text x="740" y="124.0" text-anchor="end" font-size="11">100</text>
<text x="1061" y="753" text-anchor="middle" font-size="12">game number (100-game buckets)</text>
<text x="16" y="420" text-anchor="middle" font-size="12" transform="rotate(-90 16 420)">win %</text>
<polyline points="762.4,720.0 793.8,720.0 825.2,720.0 856.6,720.0 888.1,720.0 919.5,720.0 950.9,720.0 982.3,720.0 1013.7,720.0 1045.1,720.0 1076.6,720.0 1108.0,720.0 1139.4,720.0 1170.8,720.0 1202.2,720.0 1233.6,720.0 1265.0,720.0 1296.5,720.0 1327.9,700.0 1359.3,720.0" fill="none" stroke="#1f77b4" stroke-width="3.5" opacity="1.0"/>
<polyline points="825.2,720.0 856.6,720.0 888.1,720.0 919.5,720.0 950.9,720.0 982.3,720.0 1013.7,720.0 1045.1,720.0" fill="none" stroke="#ff7f0e" stroke-width="3.5" opacity="1.0"/>
<polyline points="1108.0,720.0 1139.4,720.0 1170.8,720.0" fill="none" stroke="#ff7f0e" stroke-width="3.5" opacity="1.0"/>
<polyline points="1233.6,720.0 1265.0,720.0 1296.5,720.0 1327.9,720.0 1359.3,720.0" fill="none" stroke="#ff7f0e" stroke-width="3.5" opacity="1.0"/>
<polyline points="762.4,720.0 793.8,720.0" fill="none" stroke="#2ca02c" stroke-width="3.5" opacity="1.0"/>
<polyline points="888.1,720.0 919.5,720.0 950.9,720.0 982.3,720.0 1013.7,720.0 1045.1,720.0" fill="none" stroke="#2ca02c" stroke-width="3.5" opacity="1.0"/>
<polyline points="1202.2,690.0 1233.6,720.0" fill="none" stroke="#2ca02c" stroke-width="3.5" opacity="1.0"/>
<polyline points="1296.5,720.0 1327.9,720.0 1359.3,720.0" fill="none" stroke="#2ca02c" stroke-width="3.5" opacity="1.0"/>
<polyline points="762.4,720.0 793.8,720.0 825.2,720.0 856.6,720.0 888.1,720.0 919.5,720.0 950.9,660.0 982.3,720.0 1013.7,720.0" fill="none" stroke="#9467bd" stroke-width="3.5" opacity="1.0"/>
<polyline points="1139.4,720.0 1170.8,720.0" fill="none" stroke="#9467bd" stroke-width="3.5" opacity="1.0"/>
<polyline points="762.4,720.0 793.8,720.0 825.2,720.0 856.6,720.0 888.1,720.0 919.5,720.0 950.9,720.0 982.3,720.0 1013.7,720.0 1045.1,720.0 1076.6,720.0 1108.0,720.0 1139.4,720.0 1170.8,720.0 1202.2,720.0 1233.6,720.0 1265.0,720.0 1296.5,720.0 1327.9,720.0 1359.3,720.0" fill="none" stroke="#d62728" stroke-width="3.5" opacity="1.0"/>
<line x1="759" y1="772" x2="787" y2="772" stroke="#1f77b4" stroke-width="3"/>
<text x="793" y="776" font-size="12">Crazy (470 games)</text>
<line x1="759" y1="790" x2="787" y2="790" stroke="#ff7f0e" stroke-width="3"/>
<text x="793" y="794" font-size="12">RamFire (450 games)</text>
<line x1="759" y1="808" x2="787" y2="808" stroke="#2ca02c" stroke-width="3"/>
<text x="793" y="812" font-size="12">Target (220 games)</text>
<line x1="759" y1="826" x2="787" y2="826" stroke="#9467bd" stroke-width="3"/>
<text x="793" y="830" font-size="12">SacTwin (190 games)</text>
<line x1="759" y1="844" x2="787" y2="844" stroke="#d62728" stroke-width="3"/>
<text x="793" y="848" font-size="12">Corners (670 games)</text>
<text x="70" y="940" font-size="15" font-weight="bold">Training losses (log scale)</text>
<text x="70" y="955" font-size="11" fill="#555">training_metrics.jsonl - big early spikes are normal</text>
<line x1="70" y1="1428.0" x2="697" y2="1428.0" stroke="#dddddd"/>
<line x1="70" y1="1401.9" x2="697" y2="1401.9" stroke="#dddddd"/>
<line x1="70" y1="1375.8" x2="697" y2="1375.8" stroke="#dddddd"/>
<line x1="70" y1="1349.7" x2="697" y2="1349.7" stroke="#dddddd"/>
<line x1="70" y1="1323.6" x2="697" y2="1323.6" stroke="#dddddd"/>
<line x1="70" y1="1297.4" x2="697" y2="1297.4" stroke="#dddddd"/>
<line x1="70" y1="1271.3" x2="697" y2="1271.3" stroke="#dddddd"/>
<line x1="70" y1="1245.2" x2="697" y2="1245.2" stroke="#dddddd"/>
<line x1="70" y1="1219.1" x2="697" y2="1219.1" stroke="#dddddd"/>
<line x1="70" y1="1193.0" x2="697" y2="1193.0" stroke="#dddddd"/>
<line x1="70" y1="1166.9" x2="697" y2="1166.9" stroke="#dddddd"/>
<line x1="70" y1="1140.8" x2="697" y2="1140.8" stroke="#dddddd"/>
<line x1="70" y1="1114.7" x2="697" y2="1114.7" stroke="#dddddd"/>
<line x1="70" y1="1088.6" x2="697" y2="1088.6" stroke="#dddddd"/>
<line x1="70" y1="1062.4" x2="697" y2="1062.4" stroke="#dddddd"/>
<line x1="70" y1="1036.3" x2="697" y2="1036.3" stroke="#dddddd"/>
<line x1="70" y1="1010.2" x2="697" y2="1010.2" stroke="#dddddd"/>
<line x1="70" y1="984.1" x2="697" y2="984.1" stroke="#dddddd"/>
<line x1="70" y1="958.0" x2="697" y2="958.0" stroke="#dddddd"/>
<line x1="70" y1="1428" x2="697" y2="1428" stroke="black"/>
<line x1="70" y1="1428" x2="70" y2="958" stroke="black"/>
<line x1="70.0" y1="1428" x2="70.0" y2="1432" stroke="black"/>
<text x="70.0" y="1445" text-anchor="middle" font-size="11">1</text>
<line x1="226.8" y1="1428" x2="226.8" y2="1432" stroke="black"/>
<text x="226.8" y="1445" text-anchor="middle" font-size="11">51</text>
<line x1="383.5" y1="1428" x2="383.5" y2="1432" stroke="black"/>
<text x="383.5" y="1445" text-anchor="middle" font-size="11">101</text>
<line x1="540.2" y1="1428" x2="540.2" y2="1432" stroke="black"/>
<text x="540.2" y="1445" text-anchor="middle" font-size="11">151</text>
<line x1="697.0" y1="1428" x2="697.0" y2="1432" stroke="black"/>
<text x="697.0" y="1445" text-anchor="middle" font-size="11">201</text>
<line x1="66" y1="1428.0" x2="70" y2="1428.0" stroke="black"/>
<text x="63" y="1432.0" text-anchor="end" font-size="11">0.1</text>
<line x1="66" y1="1310.5" x2="70" y2="1310.5" stroke="black"/>
<text x="63" y="1314.5" text-anchor="end" font-size="11">3.16e+03</text>
<line x1="66" y1="1193.0" x2="70" y2="1193.0" stroke="black"/>
<text x="63" y="1197.0" text-anchor="end" font-size="11">1e+08</text>
<line x1="66" y1="1075.5" x2="70" y2="1075.5" stroke="black"/>
<text x="63" y="1079.5" text-anchor="end" font-size="11">3.16e+12</text>
<line x1="66" y1="958.0" x2="70" y2="958.0" stroke="black"/>
<text x="63" y="962.0" text-anchor="end" font-size="11">1e+17</text>
<text x="383" y="1461" text-anchor="middle" font-size="12">metric line number</text>
<text x="16" y="1193" text-anchor="middle" font-size="12" transform="rotate(-90 16 1193)">loss (log)</text>
<polyline points="70.0,1366.3 73.1,1368.7 76.3,1365.5 79.4,1360.9 82.5,1355.0 85.7,1351.6 88.8,1353.9 91.9,1357.7 95.1,979.1 98.2,1370.9 101.3,1393.2 104.5,1379.4 107.6,1366.7 110.8,1357.6 113.9,1351.5 117.0,1349.9 120.2,1348.1 123.3,1347.8 126.4,1345.5 129.6,1346.3 132.7,1348.6 135.8,1350.0 139.0,1103.1 142.1,1359.7 145.2,1364.0 148.4,1366.7 151.5,1064.5 154.6,1382.5 157.8,1382.4 160.9,1385.5 164.1,1383.9 167.2,1386.3 170.3,1390.4 173.5,1392.9 176.6,1392.1 179.7,1392.5 182.9,1404.7 186.0,1403.6 189.1,1397.8 192.3,1409.3 195.4,1415.2 198.5,979.1 201.7,979.1 204.8,1391.7 207.9,1385.3 211.1,1375.4 214.2,1372.1 217.3,979.1 220.5,1364.5 223.6,1359.7 226.8,1101.7 229.9,1357.2 233.0,1354.6 236.2,1363.0 239.3,1377.5 242.4,1390.2 245.6,1375.3 248.7,979.1 251.8,979.1 255.0,1353.1 258.1,1350.3 261.2,1352.5 264.4,1352.6 267.5,979.1 270.6,1352.6 273.8,1350.5 276.9,1354.6 280.0,1359.8 283.2,1355.3 286.3,1359.3 289.4,1354.0 292.6,1362.1 295.7,1365.6 298.9,1367.3 302.0,1374.4 305.1,1397.8 308.3,1104.4 311.4,1375.4 314.5,1261.1 317.7,1156.2 320.8,1076.6 323.9,1059.3 327.1,1058.7 330.2,1058.0 333.3,1057.7 336.5,1057.4 339.6,1057.2 342.7,1057.4 345.9,1056.5 349.0,1056.0 352.1,1055.7 355.3,1055.3 358.4,1054.6 361.6,1054.1 364.7,979.2 367.8,979.2 371.0,979.2 374.1,979.3 377.2,1052.4 380.4,1051.9 383.5,1051.6 386.6,1051.2 389.8,1050.7 392.9,1050.2 396.0,1049.7 399.2,1049.3 402.3,1048.8 405.4,1048.3 408.6,1047.8 411.7,1046.9 414.9,1046.4 418.0,1045.9 421.1,1045.5 424.3,1045.0 427.4,1044.6 430.5,1044.1 433.7,1043.6 436.8,1043.1 439.9,979.4 443.1,1042.3 446.2,1042.0 449.3,1041.6 452.5,1041.3 455.6,1040.9 458.7,1040.2 461.9,1039.9 465.0,1039.6 468.1,1039.2 471.3,1038.9 474.4,1038.5 477.6,1038.9 480.7,1037.9 483.8,1037.4 487.0,979.4 490.1,1036.9 493.2,1036.5 496.4,1036.0 499.5,979.4 502.6,1035.4 505.8,1035.1 508.9,1034.7 512.0,1034.4 515.2,1034.0 518.3,1033.7 521.4,1033.2 524.6,1032.8 527.7,979.5 530.8,1032.0 534.0,979.5 537.1,1031.3 540.2,1030.9 543.4,1030.5 546.5,1030.2 549.7,1029.9 552.8,1029.6 555.9,1029.4 559.1,1029.1 562.2,1028.7 565.3,1028.3 568.5,1028.0 571.6,1028.1 574.7,1027.3 577.9,1026.9 581.0,979.6 584.1,1026.2 587.3,1025.6 590.4,1025.3 593.5,1025.0 596.7,979.6 599.8,979.6 603.0,1023.8 606.1,1023.2 609.2,1022.9 612.4,1022.6 615.5,1022.3 618.6,1022.0 621.8,1021.7 624.9,1021.4 628.0,1021.1 631.2,1020.8 634.3,1020.5 637.4,1020.2 640.6,1019.9 643.7,1019.3 646.8,1019.1 650.0,1018.8 653.1,1018.6 656.2,1018.5 659.4,1018.0 662.5,1017.5 665.6,1017.3 668.8,1017.0 671.9,1016.7 675.1,1016.5 678.2,1016.2 681.3,1016.0 684.5,1015.8 687.6,1015.5 690.7,1015.3 693.9,979.8 697.0,1014.8" fill="none" stroke="#1f77b4" stroke-width="1.8" opacity="1.0"/>
<polyline points="70.0,1388.6 73.1,1384.3 76.3,1382.4 79.4,1380.1 82.5,1377.2 85.7,1375.8 88.8,1376.7 91.9,1378.2 95.1,1381.7 98.2,1384.2 101.3,1390.4 104.5,1394.2 107.6,1384.3 110.8,1379.2 113.9,1376.0 117.0,1373.9 120.2,1373.1 123.3,1372.6 126.4,1371.7 129.6,1371.0 132.7,1371.2 135.8,1371.3 139.0,1371.8 142.1,1372.0 145.2,1373.0 148.4,1373.4 151.5,1375.0 154.6,1375.2 157.8,1377.9 160.9,1379.3 164.1,1380.5 167.2,1381.0 170.3,1380.8 173.5,1380.4 176.6,1382.1 179.7,1382.3 182.9,1380.9 186.0,1380.5 189.1,1382.1 192.3,1379.9 195.4,1379.2 198.5,1378.0 201.7,1375.1 204.8,1372.7 207.9,1371.7 211.1,1370.2 214.2,1369.3 217.3,1367.8 220.5,1366.8 223.6,1365.7 226.8,1364.8 229.9,1364.0 233.0,1363.3 236.2,1363.7 239.3,1363.9 242.4,1364.5 245.6,1365.6 248.7,1367.1 251.8,1369.1 255.0,1370.0 258.1,1371.0 261.2,1371.8 264.4,1371.9 267.5,1371.2 270.6,1370.3 273.8,1371.1 276.9,1370.3 280.0,1369.3 283.2,1370.3 286.3,1370.0 289.4,1371.0 292.6,1369.7 295.7,1369.3 298.9,1369.5 302.0,1369.1 305.1,1367.5 308.3,1365.0 311.4,1364.7 314.5,1332.0 317.7,1278.9 320.8,1239.1 323.9,1230.5 327.1,1230.2 330.2,1229.8 333.3,1229.7 336.5,1229.5 339.6,1229.4 342.7,1229.2 345.9,1229.1 349.0,1228.8 352.1,1228.7 355.3,1228.5 358.4,1228.1 361.6,1227.9 364.7,1227.8 367.8,1227.6 371.0,1227.5 374.1,1227.3 377.2,1227.0 380.4,1226.8 383.5,1226.6 386.6,1226.4 389.8,1226.2 392.9,1225.9 396.0,1225.7 399.2,1225.5 402.3,1225.2 405.4,1225.0 408.6,1224.7 411.7,1224.3 414.9,1224.0 418.0,1223.8 421.1,1223.6 424.3,1223.3 427.4,1223.1 430.5,1222.9 433.7,1222.6 436.8,1222.4 439.9,1222.2 443.1,1222.0 446.2,1221.8 449.3,1221.7 452.5,1221.5 455.6,1221.3 458.7,1220.9 461.9,1220.8 465.0,1220.6 468.1,1220.5 471.3,1220.3 474.4,1220.1 477.6,1219.9 480.7,1219.8 483.8,1219.5 487.0,1219.4 490.1,1219.3 493.2,1219.1 496.4,1218.8 499.5,1218.6 502.6,1218.5 505.8,1218.4 508.9,1218.2 512.0,1218.0 515.2,1217.8 518.3,1217.7 521.4,1217.5 524.6,1217.3 527.7,1217.0 530.8,1216.8 534.0,1216.7 537.1,1216.5 540.2,1216.3 543.4,1216.1 546.5,1215.9 549.7,1215.8 552.8,1215.6 555.9,1215.5 559.1,1215.4 562.2,1215.2 565.3,1215.0 568.5,1214.8 571.6,1214.6 574.7,1214.5 577.9,1214.3 581.0,1214.1 584.1,1213.8 587.3,1213.6 590.4,1213.5 593.5,1213.3 596.7,1213.1 599.8,1213.0 603.0,1212.8 606.1,1212.4 609.2,1212.3 612.4,1212.1 615.5,1212.0 618.6,1211.8 621.8,1211.7 624.9,1211.5 628.0,1211.4 631.2,1211.2 634.3,1211.1 637.4,1210.9 640.6,1210.8 643.7,1210.5 646.8,1210.4 650.0,1210.2 653.1,1210.1 656.2,1210.0 659.4,1209.8 662.5,1209.6 665.6,1209.5 668.8,1209.3 671.9,1209.2 675.1,1209.1 678.2,1208.9 681.3,1208.8 684.5,1208.7 687.6,1208.6 690.7,1208.5 693.9,1208.3 697.0,1208.2" fill="none" stroke="#ff7f0e" stroke-width="1.8" opacity="1.0"/>
<line x1="82" y1="1480" x2="110" y2="1480" stroke="#1f77b4" stroke-width="3"/>
<text x="116" y="1484" font-size="12">critic_loss</text>
<line x1="82" y1="1498" x2="110" y2="1498" stroke="#ff7f0e" stroke-width="3"/>
<text x="116" y="1502" font-size="12">|actor_loss|</text>
<text x="747" y="940" font-size="15" font-weight="bold">Alpha temperature</text>
<text x="747" y="955" font-size="11" fill="#555">training_metrics.jsonl - high = exploring, low = exploiting</text>
<line x1="747" y1="1428.0" x2="1375" y2="1428.0" stroke="#dddddd"/>
<line x1="747" y1="1310.5" x2="1375" y2="1310.5" stroke="#dddddd"/>
<line x1="747" y1="1193.0" x2="1375" y2="1193.0" stroke="#dddddd"/>
<line x1="747" y1="1075.5" x2="1375" y2="1075.5" stroke="#dddddd"/>
<line x1="747" y1="958.0" x2="1375" y2="958.0" stroke="#dddddd"/>
<line x1="747" y1="1428" x2="1375" y2="1428" stroke="black"/>
<line x1="747" y1="1428" x2="747" y2="958" stroke="black"/>
<line x1="747.0" y1="1428" x2="747.0" y2="1432" stroke="black"/>
<text x="747.0" y="1445" text-anchor="middle" font-size="11">1</text>
<line x1="904.0" y1="1428" x2="904.0" y2="1432" stroke="black"/>
<text x="904.0" y="1445" text-anchor="middle" font-size="11">51</text>
<line x1="1061.0" y1="1428" x2="1061.0" y2="1432" stroke="black"/>
<text x="1061.0" y="1445" text-anchor="middle" font-size="11">101</text>
<line x1="1218.0" y1="1428" x2="1218.0" y2="1432" stroke="black"/>
<text x="1218.0" y="1445" text-anchor="middle" font-size="11">151</text>
<line x1="1375.0" y1="1428" x2="1375.0" y2="1432" stroke="black"/>
<text x="1375.0" y="1445" text-anchor="middle" font-size="11">201</text>
<line x1="743" y1="1428.0" x2="747" y2="1428.0" stroke="black"/>
<text x="740" y="1432.0" text-anchor="end" font-size="11">0</text>
<line x1="743" y1="1310.5" x2="747" y2="1310.5" stroke="black"/>
<text x="740" y="1314.5" text-anchor="end" font-size="11">0.252</text>
<line x1="743" y1="1193.0" x2="747" y2="1193.0" stroke="black"/>
<text x="740" y="1197.0" text-anchor="end" font-size="11">0.504</text>
<line x1="743" y1="1075.5" x2="747" y2="1075.5" stroke="black"/>
<text x="740" y="1079.5" text-anchor="end" font-size="11">0.756</text>
<line x1="743" y1="958.0" x2="747" y2="958.0" stroke="black"/>
<text x="740" y="962.0" text-anchor="end" font-size="11">1.01</text>
<text x="1061" y="1461" text-anchor="middle" font-size="12">metric line number</text>
<text x="16" y="1193" text-anchor="middle" font-size="12" transform="rotate(-90 16 1193)">alpha</text>
<polyline points="747.0,961.7 750.1,962.3 753.3,962.6 756.4,963.0 759.6,963.6 762.7,964.5 765.8,964.8 769.0,965.1 772.1,965.6 775.3,965.9 778.4,966.5 781.5,967.4 784.7,967.6 787.8,967.0 791.0,966.5 794.1,965.9 797.2,965.4 800.4,964.8 803.5,964.1 806.7,963.6 809.8,963.0 812.9,962.5 816.1,961.9 819.2,961.3 822.4,960.8 825.5,960.5 828.6,959.6 831.8,959.3 834.9,959.8 838.1,960.1 841.2,960.5 844.3,960.8 847.5,961.1 850.6,961.5 853.8,962.0 856.9,962.5 860.0,963.3 863.2,963.9 866.3,964.4 869.5,965.0 872.6,965.5 875.7,966.0 878.9,966.3 882.0,965.9 885.2,965.5 888.3,965.2 891.4,964.9 894.6,964.4 897.7,964.0 900.9,963.5 904.0,963.0 907.1,962.4 910.3,962.0 913.4,961.6 916.6,960.6 919.7,960.1 922.8,959.5 926.0,958.5 929.1,958.0 932.3,958.4 935.4,958.8 938.5,959.4 941.7,960.0 944.8,960.5 948.0,961.5 951.1,962.1 954.2,962.6 957.4,963.2 960.5,963.7 963.7,964.0 966.8,964.6 969.9,965.1 973.1,965.7 976.2,965.9 979.4,965.9 982.5,965.4 985.6,964.3 988.8,963.7 991.9,963.5 995.1,964.1 998.2,964.6 1001.3,964.9 1004.5,965.6 1007.6,966.4 1010.8,966.7 1013.9,967.0 1017.0,967.2 1020.2,967.8 1023.3,968.1 1026.5,968.6 1029.6,969.2 1032.7,969.7 1035.9,970.4 1039.0,971.0 1042.2,971.2 1045.3,971.5 1048.4,971.8 1051.6,972.3 1054.7,972.9 1057.9,973.4 1061.0,973.7 1064.1,974.2 1067.3,974.8 1070.4,975.3 1073.6,975.9 1076.7,976.4 1079.8,977.0 1083.0,977.5 1086.1,978.0 1089.3,979.0 1092.4,979.5 1095.5,980.1 1098.7,980.5 1101.8,981.0 1105.0,981.4 1108.1,981.9 1111.2,982.5 1114.4,983.0 1117.5,983.5 1120.7,983.9 1123.8,984.3 1126.9,984.7 1130.1,985.3 1133.2,985.8 1136.4,986.7 1139.5,987.3 1142.6,987.7 1145.8,988.2 1148.9,988.7 1152.1,989.2 1155.2,989.8 1158.3,990.2 1161.5,990.9 1164.6,991.3 1167.8,991.7 1170.9,992.3 1174.0,993.0 1177.2,993.6 1180.3,994.0 1183.5,994.3 1186.6,994.9 1189.7,995.4 1192.9,995.9 1196.0,996.4 1199.2,996.9 1202.3,997.5 1205.4,998.0 1208.6,998.5 1211.7,999.0 1214.9,999.5 1218.0,1000.0 1221.1,1000.5 1224.3,1000.9 1227.4,1001.3 1230.6,1001.7 1233.7,1002.1 1236.8,1002.5 1240.0,1003.0 1243.1,1003.5 1246.3,1004.0 1249.4,1004.5 1252.5,1005.0 1255.7,1005.5 1258.8,1006.0 1262.0,1006.9 1265.1,1007.4 1268.2,1007.9 1271.4,1008.4 1274.5,1008.9 1277.7,1009.4 1280.8,1010.1 1283.9,1011.1 1287.1,1011.6 1290.2,1012.1 1293.4,1012.6 1296.5,1013.1 1299.6,1013.6 1302.8,1014.1 1305.9,1014.5 1309.1,1015.0 1312.2,1015.7 1315.3,1016.2 1318.5,1016.7 1321.6,1017.6 1324.8,1018.0 1327.9,1018.5 1331.0,1019.0 1334.2,1019.5 1337.3,1020.0 1340.5,1020.8 1343.6,1021.3 1346.7,1021.8 1349.9,1022.3 1353.0,1022.8 1356.2,1023.3 1359.3,1023.6 1362.4,1024.1 1365.6,1024.6 1368.7,1025.1 1371.9,1025.6 1375.0,1026.0" fill="none" stroke="#9467bd" stroke-width="1.8" opacity="1.0"/>
<text x="70" y="1600" font-size="15" font-weight="bold">Throughput - games per hour</text>
<text x="70" y="1615" font-size="11" fill="#555">method: training_metrics.jsonl 'epoch' deltas (training_log.jsonl has no timestamps); 1 row = one 10-game chunk</text>
<line x1="70" y1="1780.0" x2="1375" y2="1780.0" stroke="#dddddd"/>
<line x1="70" y1="1739.5" x2="1375" y2="1739.5" stroke="#dddddd"/>
<line x1="70" y1="1699.0" x2="1375" y2="1699.0" stroke="#dddddd"/>
<line x1="70" y1="1658.5" x2="1375" y2="1658.5" stroke="#dddddd"/>
<line x1="70" y1="1618.0" x2="1375" y2="1618.0" stroke="#dddddd"/>
<line x1="70" y1="1780" x2="1375" y2="1780" stroke="black"/>
<line x1="70" y1="1780" x2="70" y2="1618" stroke="black"/>
<line x1="194.6" y1="1780" x2="194.6" y2="1784" stroke="black"/>
<text x="194.6" y="1797" text-anchor="middle" font-size="11">20</text>
<line x1="325.8" y1="1780" x2="325.8" y2="1784" stroke="black"/>
<text x="325.8" y="1797" text-anchor="middle" font-size="11">40</text>
<line x1="456.9" y1="1780" x2="456.9" y2="1784" stroke="black"/>
<text x="456.9" y="1797" text-anchor="middle" font-size="11">60</text>
<line x1="588.1" y1="1780" x2="588.1" y2="1784" stroke="black"/>
<text x="588.1" y="1797" text-anchor="middle" font-size="11">80</text>
<line x1="719.2" y1="1780" x2="719.2" y2="1784" stroke="black"/>
<text x="719.2" y="1797" text-anchor="middle" font-size="11">100</text>
<line x1="850.4" y1="1780" x2="850.4" y2="1784" stroke="black"/>
<text x="850.4" y="1797" text-anchor="middle" font-size="11">120</text>
<line x1="981.5" y1="1780" x2="981.5" y2="1784" stroke="black"/>
<text x="981.5" y="1797" text-anchor="middle" font-size="11">140</text>
<line x1="1112.7" y1="1780" x2="1112.7" y2="1784" stroke="black"/>
<text x="1112.7" y="1797" text-anchor="middle" font-size="11">160</text>
<line x1="1243.8" y1="1780" x2="1243.8" y2="1784" stroke="black"/>
<text x="1243.8" y="1797" text-anchor="middle" font-size="11">180</text>
<line x1="1375.0" y1="1780" x2="1375.0" y2="1784" stroke="black"/>
<text x="1375.0" y="1797" text-anchor="middle" font-size="11">200</text>
<line x1="66" y1="1780.0" x2="70" y2="1780.0" stroke="black"/>
<text x="63" y="1784.0" text-anchor="end" font-size="11">0</text>
<line x1="66" y1="1739.5" x2="70" y2="1739.5" stroke="black"/>
<text x="63" y="1743.5" text-anchor="end" font-size="11">1688</text>
<line x1="66" y1="1699.0" x2="70" y2="1699.0" stroke="black"/>
<text x="63" y="1703.0" text-anchor="end" font-size="11">3375</text>
<line x1="66" y1="1658.5" x2="70" y2="1658.5" stroke="black"/>
<text x="63" y="1662.5" text-anchor="end" font-size="11">5063</text>
<line x1="66" y1="1618.0" x2="70" y2="1618.0" stroke="black"/>
<text x="63" y="1622.0" text-anchor="end" font-size="11">6750</text>
<text x="722" y="1813" text-anchor="middle" font-size="12">chunk interval (10-game chunks)</text>
<text x="16" y="1699" text-anchor="middle" font-size="12" transform="rotate(-90 16 1699)">games / hour</text>
<circle cx="70.0" cy="1698.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="76.6" cy="1774.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="83.1" cy="1658.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="89.7" cy="1776.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="96.2" cy="1762.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="102.8" cy="1776.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="109.3" cy="1655.7" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="115.9" cy="1775.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="122.5" cy="1654.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="129.0" cy="1776.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="135.6" cy="1762.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="142.1" cy="1773.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="148.7" cy="1689.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="155.3" cy="1772.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="161.8" cy="1698.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="168.4" cy="1771.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="174.9" cy="1693.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="181.5" cy="1772.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="188.0" cy="1696.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="194.6" cy="1771.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="201.2" cy="1692.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="207.7" cy="1771.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="214.3" cy="1689.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="220.8" cy="1771.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="227.4" cy="1653.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="233.9" cy="1773.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="240.5" cy="1687.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="247.1" cy="1772.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="253.6" cy="1653.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="260.2" cy="1775.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="266.7" cy="1650.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="273.3" cy="1773.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="279.8" cy="1682.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="286.4" cy="1774.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="293.0" cy="1687.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="299.5" cy="1774.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="306.1" cy="1691.6" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="312.6" cy="1774.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="319.2" cy="1690.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="325.8" cy="1775.7" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="332.3" cy="1690.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="338.9" cy="1776.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="345.4" cy="1659.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="352.0" cy="1772.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="358.5" cy="1645.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="365.1" cy="1773.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="371.7" cy="1680.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="378.2" cy="1772.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="384.8" cy="1666.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="391.3" cy="1775.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="397.9" cy="1690.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="404.4" cy="1772.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="411.0" cy="1679.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="417.6" cy="1775.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="424.1" cy="1686.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="430.7" cy="1775.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="437.2" cy="1762.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="443.8" cy="1775.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="450.4" cy="1694.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="456.9" cy="1775.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="463.5" cy="1680.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="470.0" cy="1775.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="476.6" cy="1681.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="483.1" cy="1775.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="489.7" cy="1691.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="496.3" cy="1772.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="502.8" cy="1683.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="509.4" cy="1774.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="515.9" cy="1642.7" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="522.5" cy="1775.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="529.0" cy="1683.7" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="535.6" cy="1774.6" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="542.2" cy="1639.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="548.7" cy="1775.6" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="555.3" cy="1696.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="561.8" cy="1775.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="568.4" cy="1691.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="574.9" cy="1773.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="581.5" cy="1696.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="588.1" cy="1771.7" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="594.6" cy="1632.7" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="601.2" cy="1757.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="607.7" cy="1771.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="614.3" cy="1641.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="620.9" cy="1770.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="627.4" cy="1650.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="634.0" cy="1771.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="640.5" cy="1639.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="647.1" cy="1770.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="653.6" cy="1689.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="660.2" cy="1770.7" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="666.8" cy="1705.7" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="673.3" cy="1771.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="679.9" cy="1652.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="686.4" cy="1771.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="693.0" cy="1660.7" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="699.5" cy="1771.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="706.1" cy="1681.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="712.7" cy="1770.7" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="719.2" cy="1643.7" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="725.8" cy="1770.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="732.3" cy="1690.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="738.9" cy="1771.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="745.5" cy="1684.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="752.0" cy="1769.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="758.6" cy="1688.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="765.1" cy="1769.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="771.7" cy="1695.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="778.2" cy="1773.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="784.8" cy="1694.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="791.4" cy="1770.6" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="797.9" cy="1673.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="804.5" cy="1770.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="811.0" cy="1676.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="817.6" cy="1771.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="824.1" cy="1691.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="830.7" cy="1769.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="837.3" cy="1692.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="843.8" cy="1770.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="850.4" cy="1670.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="856.9" cy="1770.6" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="863.5" cy="1698.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="870.1" cy="1770.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="876.6" cy="1762.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="883.2" cy="1770.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="889.7" cy="1688.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="896.3" cy="1770.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="902.8" cy="1690.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="909.4" cy="1771.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="916.0" cy="1693.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="922.5" cy="1769.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="929.1" cy="1760.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="935.6" cy="1769.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="942.2" cy="1682.6" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="948.7" cy="1770.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="955.3" cy="1761.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="961.9" cy="1770.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="968.4" cy="1680.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="975.0" cy="1769.6" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="981.5" cy="1681.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="988.1" cy="1769.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="994.6" cy="1692.7" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1001.2" cy="1771.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1007.8" cy="1685.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1014.3" cy="1771.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1020.9" cy="1690.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1027.4" cy="1772.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1034.0" cy="1696.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1040.6" cy="1770.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1047.1" cy="1682.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1053.7" cy="1771.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1060.2" cy="1673.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1066.8" cy="1770.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1073.3" cy="1676.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1079.9" cy="1769.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1086.5" cy="1668.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1093.0" cy="1771.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1099.6" cy="1694.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1106.1" cy="1771.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1112.7" cy="1695.6" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1119.2" cy="1771.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1125.8" cy="1689.7" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1132.4" cy="1771.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1138.9" cy="1761.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1145.5" cy="1771.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1152.0" cy="1688.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1158.6" cy="1771.7" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1165.2" cy="1684.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1171.7" cy="1772.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1178.3" cy="1702.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1184.8" cy="1774.6" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1191.4" cy="1702.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1197.9" cy="1772.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1204.5" cy="1688.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1211.1" cy="1772.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1217.6" cy="1696.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1224.2" cy="1772.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1230.7" cy="1695.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1237.3" cy="1771.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1243.8" cy="1707.2" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1250.4" cy="1772.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1257.0" cy="1699.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1263.5" cy="1774.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1270.1" cy="1672.7" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1276.6" cy="1771.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1283.2" cy="1689.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1289.7" cy="1771.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1296.3" cy="1696.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1302.9" cy="1774.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1309.4" cy="1689.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1316.0" cy="1772.3" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1322.5" cy="1687.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1329.1" cy="1771.8" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1335.7" cy="1690.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1342.2" cy="1772.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1348.8" cy="1691.0" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1355.3" cy="1772.5" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1361.9" cy="1701.1" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1368.4" cy="1770.4" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<circle cx="1375.0" cy="1682.9" r="1.6" fill="#7f7f7f" opacity="0.25"/>
<polyline points="96.2,1730.7 161.8,1740.2 227.4,1724.1 293.0,1727.4 358.5,1721.3 424.1,1738.6 489.7,1725.4 555.3,1727.8 620.9,1709.3 686.4,1720.0 752.0,1730.7 817.6,1725.8 883.2,1738.6 948.7,1741.6 1014.3,1730.2 1079.9,1726.2 1145.5,1738.5 1211.1,1735.3 1276.6,1731.1 1342.2,1731.1" fill="none" stroke="#2ca02c" stroke-width="3.5" opacity="1.0"/>
<rect x="0" y="1820" width="1400" height="180" fill="#f2f2f2"/>
<text x="16" y="1841" font-size="15" font-weight="bold">How to read</text>
<text x="1384" y="1840" text-anchor="end" font-size="11" fill="#666">Regenerate anytime: python3 tools/plot_progress.py</text>
<text x="16" y="1863" font-size="13">Test matches: dots are single fights, thick line shows trend.</text>
<text x="16" y="1880" font-size="13">Real battles only. Rising lines mean the bot improves.</text>
<text x="16" y="1897" font-size="13">Loss spikes are normal early; endless growth is bad.</text>
<text x="16" y="1914" font-size="13">Alpha high means experimenting; falling too fast freezes habits.</text>
<text x="16" y="1931" font-size="13">Throughput flat is healthy; dips mean something slowed.</text>
<text x="16" y="1948" font-size="13">This file reloads itself in Chrome every sixty seconds.</text>
<text x="16" y="1965" font-size="13">Regenerate anytime with tools/watch_dashboard.sh or the python command.</text>
<script type="text/javascript"><![CDATA[ setTimeout(function(){ location.reload(); }, 60000); ]]></script>
</svg>

After

Width:  |  Height:  |  Size: 63 KiB

+332
View File
@@ -0,0 +1,332 @@
## The story so far, in simple words
This project trains a robot tank. It plays many fights against other tanks.
After each fight it changes itself a little. It keeps the changes that helped it win.
Night 1 (run 1) finished without problems. It ran for 14 hours alone. It never crashed.
It beat an old copy of itself most of the time. It beat Crazy about half the time.
Three things went badly. First, learning was not stable. Good skill appeared, then disappeared again.
Second, the saved "best" version came from one lucky perfect score. It was not really its best.
Third, the bot learned to hide and survive. It almost never shot back.
The human approved five fixes. All five were put into the code.
Run 2 used these fixes. Its error numbers grew far too big. Learning broke.
We made one speed number smaller. This number sets how fast one part learns.
Then we dropped the broken progress and started clean. This is run 3. It is running now.
Next we watch run 3. One of three doors will open.
Door 1: it stays steady. We let it run to the end.
Door 2: the numbers grow too big again. We turn the next speed number down.
Door 3: it stays steady but still fights badly. We teach aiming as a separate, direct lesson.
Updated: 2026-08-23 — this section is refreshed at every major step.
## Small dictionary
- **training**: the time when the bot plays fights and changes itself to improve. It learns only during training.
- **battle**: one group of fights against one opponent. The bot restarts between groups.
- **round**: one single fight. Win it by destroying the enemy tank or outliving it.
- **chunk**: one work block: a battle of up to 10 rounds, then some learning from it.
- **eval (test match)**: a test match. The bot does not learn during these. We use them only to measure.
- **win rate**: how many test matches were won, as a percent. 8 wins in 10 matches = 80%.
- **checkpoint**: a saved copy of the bot's brain (a zip file). Written every few learning steps.
- **"best" checkpoint**: the saved copy we currently call best. Run 1 picked one from a lucky score, hence the quotes.
- **replay buffer**: the bot's memory of past moments: what it saw, did, and received. Learning picks old moments from it.
- **loss (critic/actor)**: a number saying how wrong the bot's inner guesses are. Lower usually means better. Losses growing huge mean trouble.
- **alpha**: a dial setting how much the bot tries new moves instead of repeating known good ones.
- **MA / composite score**: MA is the average of the last few win rates; it smooths luck. Composite is the average of MAs across all test opponents.
- **twin (SacTwin)**: a frozen copy of our own bot, used as a practice partner. Beating it proves real improvement.
- **lever**: one numbered change we prepared, waiting for approval. There are levers 1 to 5.
- **watchman**: a helper who checks the running training at set times and stops it if something breaks.
---
# Campaign Notebook — campaign-v1 (SAC_LSTM_Bot)
Overnight training campaign on branch `research/goto-controller`.
Companion tickets: config+launch = **#56**, morning verdict = **#57**, map = **#53**.
Monitoring contract: observability inventory in **#55 comment 461** (9 signals, thresholds, four-way discrimination).
## Locked config (campaign-v1)
Launch command (tmux session `sac_campaign`, stdout teed to `campaign_stdout.log`):
```bash
SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1' \
SAC_EVAL_OPPONENT=Corners \
SAC_TOTAL_ROUNDS=25000 \
SAC_CHUNK_SIZE=10 \
SAC_EVAL_INTERVAL=2 \
SAC_EVAL_ROUNDS=10 \
SAC_MAX_CRASHES=5 \
SACLSTM_HIDDEN_SIZE=256 \
SACLSTM_BATCH_SIZE=16 \
SACLSTM_UTD_RATIO=1 \
SACLSTM_SAVE_INTERVAL=5 \
./sac_train.sh 2>&1 | tee -a campaign_stdout.log
```
| Knob | Value | Source / rationale |
|------|-------|--------------------|
| `SAC_OPPONENTS` | `Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1` | Working recommendation kept. Twin pinned at **1 not 2**: #54 showed mirror battles end early ⇒ fewer transitions per chunk; weight 2 would starve the replay buffer. |
| `SAC_EVAL_OPPONENT` | `Corners` | Pinned explicitly (= first pool entry default, #54 note) — removes reorder footgun. |
| `SAC_TOTAL_ROUNDS` | `25000` | Sized so the **wall-clock ceiling binds first**: measured throughput ~3400 rounds/h early (drops as battles lengthen) ⇒ 2000 would have exhausted in ~1 h. |
| `SAC_CHUNK_SIZE` | `10` | Harness default. |
| `SAC_EVAL_INTERVAL` | `2` | Harness default — eval every ~20 rounds. |
| `SAC_EVAL_ROUNDS` | `10` | Harness default. |
| `SAC_MAX_CRASHES` | `5` | Harness self-abort; monitor intervenes earlier at ≥3 consecutive crashes (#55). |
| `SACLSTM_HIDDEN_SIZE` | `256` | Module default (`network.nim`); real capacity vs #49 smoke's 32; under `MaxHidden`=512 cap. |
| `SACLSTM_BATCH_SIZE` | `16` | Module default (`integration.nim`). |
| `SACLSTM_UTD_RATIO` | `1` | Module default. |
| `SACLSTM_SAVE_INTERVAL` | `5` | **Deviation** from default 500 and from #54's "10–20": at hidden 256 a gradient step takes ~1 s and a chunk process fits only ~10–18 steps (see incident below) — interval must sit **inside the per-process step budget**. 5 ⇒ checkpoint every ~5–10 s of active training; IO trivial (10.5 MB zip, atomic replace). |
| *(not pinned)* | module defaults | `LR_ACTOR/LR_CRITIC/LR_ALPHA=3e-4`, `GAMMA=0.99`, `TAU=0.005`, `TARGET_ENTROPY=-4.0`, `BUFFER_CAPACITY=500000`, `BURN_IN=8`, `TRAIN_WINDOW=16`. |
**Budget**: generous wall-clock **ceiling, not a deadline** — originally **T+12 h** from launch 00:21:55 CEST 2026-08-22 ⇒ 12:22 CEST (epoch 1787394115); **extended 2026-08-22 ~07:35 by orchestrator decision on human mandate** ("no deadlines — let the 25000-round budget complete", ~14:35 projected) ⇒ ceiling now **16:30 CEST 2026-08-22 (epoch 1787409000)**, enforced by a hard user-systemd net unit `sac-ceiling-net` (sleeps to the epoch, then kills the tmux session and any straggler harness processes). While HEALTHY per #55 discrimination rules the run continues; a monitor kills the tmux session at the ceiling or on an intervene threshold.
**Fresh start**: pre-campaign `weights/` held #49-smoke 32-hidden checkpoints, incompatible with hidden=256. Archived to `weights_smoke49_backup/`; campaign baseline re-established by probe battles vs Corners (random-init hidden-256 checkpoint, `best_score.txt` reset then re-raised to 20 by a genuine eval). Twin regenerated via `./make_twin.sh` from that baseline (md5 `61521cff…` verified seed).
## Phase log
Every entry below starts with a plain-language first sentence. Technical detail follows for those who want it.
- **2026-08-21 23:22** — Claim posted on #56 (comment 471). Config locked, notebook committed (`2f49cb2`).
- **2026-08-21 23:27** — Smoke weights archived; release build; bootstrap + probe battles vs Corners established a hidden-256 baseline checkpoint (`sac_latest.zip`, 10.5 MB) and twin seed.
- **2026-08-21 23:32** — Twin regenerated (md5-verified). **Launch attempt 1** (SAVE_INTERVAL=20, TOTAL_ROUNDS=2000): ran 16+ chunks, evals every 2 chunks — but **zero checkpoints persisted** (see incident). Killed 23:48.
- **2026-08-21 23:52–00:10** — Diagnosis (see incident): interval=1 fired, interval=2/20 never; instrumentation + /proc thread forensics ⇒ per-process step budget ~10–18 at ~1 s/step; save check ran only between drain-burst passes.
- **2026-08-22 00:12** — Fix: save check moved inside the gradient-step loop, committed `2653671`. Validated: interval=5 save fired ~12 s into a battle.
- **2026-08-22 00:21:55** — **Launch (final)**: tmux `sac_campaign`, config above. First campaign save on disk at t+54 s; eval #1 on cadence.
- **2026-08-22 00:30** — HEALTHY checklist passed (see below).
- **2026-08-22 03:45** — Watch shift 1 (00:34–03:35): liveness flawless (10/10 HEALTHY, zero banners, zip ≤30 s). Learning signal: steady-state eval vs Corners 0–5% with two isolated 10/10 spikes (~01:15) → capability emerged, then lost. Eval-regression intervene threshold fired per #55; intervention DEFERRED to Campaign verdict (#57) — rationale: n=2 evidence, no loss metrics, buffer-loss on restart, run completes ~08:05 anyway. Milestone issue: #58 "Campaign-v1 watch: eval-regression threshold fired — intervention deferred to verdict".
- **2026-08-22 ~07:15** — Morning audit: policy demonstrably learning off-benchmark (SacTwin 73→100%, Crazy 26→56%) while Corners eval stays ~0–9% with 3 transient 10/10s; wall-clock ceiling extended to let the 25k complete (~14:35 projected); instability-vs-plateau question left to the curve.
- **2026-08-22 ~07:50 — score:60 anatomy**: Tank Royale survival(50)+last-survivor(10) awarded when opponent dies while we survive; exactly-60 ⇒ zero damage dealt by us that round (opponent self-destructed via wasted shots + 0.1/turn inactivity drain). 3,878 rounds (19%); modal vs SacTwin; vs Corners 986 damageless outlives vs 242 true wins; combined with 43% of rounds being score:0, texture = survivor-not-fighter against walls. RL rewards are event-driven (rewards module), so behavioral evidence, not reward poisoning. Feeds #57 levers: aggression shaping / specialist-vs-generalist.
- **2026-08-22 ~08:05 — `.part` debris forensics + sweep**: 99 `*.zip.tmp.*.part` (483 MB) are NOT weights.nim debris (that proc uses fixed `.tmp` + finally-cleanup, working); naming matches an external write-temp→rename copier killed mid-write, bursts correlating with kill events; possible culprit: a folder-sync client fighting a file that changes every ~20 s (**human asked to confirm**). Swept with `-mmin +10` age guard (protects in-flight writes); 483 MB freed, real zips untouched — note 2 fresh `.part` reappeared minutes later, copier still active.
- **2026-08-22 ~08:00 — ceiling defused**: the 12:22 'self-kill' was notebook prose instructing watchmen — never an OS mechanism; rewritten to 16:30 CEST AND armed a real systemd --user net `sac-ceiling-net` firing 16:30:00 (epoch 1787409000); training uninterrupted (counter +99/147 s verified); annotated #56 comment 487.
- **2026-08-22 ~14:28 — CAMPAIGN ENDED NATURALLY**: banner `>>> training complete: 2500 chunks` after **14 h 07 m** (00:21:55 → ~14:28 CEST); round_counter **38912**; **zero crash banners across the whole run**; the `sac-ceiling-net` backstop never fired — cancelled unneeded. Budget note: the harness loop is **chunk-based** — `SAC_TOTAL_ROUNDS=25000` ÷ `CHUNK_SIZE=10` ⇒ **2500 chunk battles** of ≤10 rounds each; the "25k-rounds" label was a misnomer (the counter also accrues 1287×10 eval rounds and rerun chunks). Final eval vs Corners: 10%.
- **2026-08-22 ~14:50 — VERDICT posted (→ #57)**: **RETUNE BEFORE SCALING.** Ops layer PROVEN (14 h autonomous, zero crashes, self-healing restarts, natural completion — the harness scales); learning REAL BUT NARROW (within-opponent gains genuine — SacTwin 90.7%, Crazy 26→48% — but specialist-not-generalist, walls untouched; 12 eval spikes ≥8/10 incl. 5×10/10, none retained); benchmark pathology: Corners-only deterministic eval + single-max best gating froze `sac_best.zip` at 01:01:57 on a fluke 10/10. Five code-level levers staged **awaiting human sign-off**; v2 NOT launched. Full rationale: Results below + #57 resolution comment.
- **2026-08-22 ~14:50 — HYGIENE (campaign over, no live writers — safe)**: `sac-ceiling-net` stopped + `reset-failed` (backstop obsolete); `.part` corpse sweep **48 → 0** (no age guard needed — nothing writes anymore); future-run guard added to `sac_train.sh`: startup `rm -f "$WEIGHTS_DIR"/sac_latest.zip.tmp.*.part "$WEIGHTS_DIR"/sac_latest.zip.tmp` so a SIGKILLed run's libzip modify-path corpses can't accumulate again.
## Decision-issue index
| Issue | What it decided |
|-------|-----------------|
| #37–#48 | Bot built: skeleton, state, actions, rewards, LSTM network, weights, SAC+LSTM training, integration. |
| #49 | Training harness + smoke run (toy hyperparams: hidden 32). |
| #54 | Mirror-twin sparring partner; SAVE_INTERVAL persistence rule; eval opponent = first pool entry. |
| #55 | 9-signal observability inventory; CRASHED/STALLED/SLOW-LEARNER/HEALTHY discriminators; monitor thresholds. |
| #56 | This campaign: locked config + launch + the save-check fix (`2653671`). |
| #57 | Morning verdict — consumes this notebook + logs. |
## Incidents & checks
### Incident 1 — zero checkpoint persistence at production sizes (launch blockers, fixed)
**Symptom**: campaign ran 26+ chunk processes across two attempts without a single `sac_latest.zip` update, while rounds/evals flowed normally. #54's rule ("keep `SACLSTM_SAVE_INTERVAL` well below per-chunk gradient-step counts, 10–20 fired in smokes") silently broke at hidden 256.
**Diagnosis chain** (all reproducible):
1. Interval=1 saved within seconds; interval=2 and 20 never saved — through the *same* harness ⇒ not env propagation.
2. Temporary step instrumentation (bot stderr via a one-line `SAC_LSTM_Bot.sh` redirect — the vendored runner swallows bot stderr, #55 gap S9-adjacent): steps cost **~1.06 s each**; a drain burst queued 53 steps; logging stopped mid-pass while rounds kept completing.
3. `/proc/<pid>/task` sampling: training thread alive and RUNNING (~13 s CPU per ~40 s process) — not deadlocked, just slow ⇒ **per-process step budget ≈ 10–18 steps**.
4. The save check lived *between* drain-burst passes; with bursts queueing minutes of steps, `stepCount` never reached `nextSave` before process teardown. Smoke runs masked this: hidden 32 steps were sub-millisecond, so hundreds of steps fit per chunk.
**Fix** (commit `2653671`): save check relocated **inside** the step loop (checked every gradient step; `packFull`+`trySend` unchanged). Validated: interval=5 save fires ~12 s into a battle; campaign save fired 54 s after launch.
**Config consequences**: `SACLSTM_SAVE_INTERVAL=5` (inside the per-process budget; #54's 10–20 was derived at smoke speeds). `SAC_TOTAL_ROUNDS=25000` (throughput measured ~3400 rounds/h, so 2000 was a 1-hour budget, not an overnight one). OMP_NUM_THREADS=1 tested and **not** needed (hang was step-budget exhaustion, not OpenMP).
### Launch health check (t+9 min, 00:30:16) — **HEALTHY** per #55 checklist
| Signal | Reading | Verdict |
|--------|---------|---------|
| S1 harness stdout | teed to `campaign_stdout.log`; **0** `crash`/`aborted` banners | ✓ |
| S4 round_counter | 1890, +460 in 9 min (~51 rounds/min) | ✓ advancing |
| S6 sac_latest.zip mtime | **6 s old**; first save at t+54 s | ✓ fresh |
| S2 training_log.jsonl | 1492 lines, growing; last ticks=668, plausible | ✓ |
| S3 eval_log.jsonl | age 3 s (atomic replace); eval every 2 chunks | ✓ on cadence |
| S5 best_score | 20 (from a genuine campaign-1 eval; non-decreasing) | ✓ |
| Sampling | 112 chunks: Corners 41 / Crazy 25 / RamFire 15 / SacTwin 16 / Target 15 ≈ weights 3:2:1:1:1 | ✓ plausible |
| Disk | 418 GB free | ✓ |
Early evals 0% vs Corners — expected for a near-random policy minutes in; SLOW-LEARNER watch rule (flat ≥5 evals = watch) applies, never intervene.
## Check-in procedure (for monitor sessions)
```bash
tmux capture-pane -p -t sac_campaign | tail -5 # S1: banners, crashes
cat ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/round_counter.txt
stat -c '%Y' ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/sac_latest.zip # age <~600s = training alive
tail -1 ~/Projects/SirRoboGarage/SAC_LSTM_Bot/training_log.jsonl
tail -3 ~/Projects/SirRoboGarage/SAC_LSTM_Bot/campaign_stdout.log # eval results / new best
cat ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/best_score.txt
```
Intervene per #55 thresholds: ≥3 consecutive `crash #N` banners; ΔS4=0 over ≥15 min; zip mtime >10 min stale while S4 advances (STALLED); disk <1 GB. At the **ceiling (16:30 CEST Aug 22, epoch 1787409000 — extended from 12:22 per human mandate)**: `tmux kill-session -t sac_campaign` if still running (the hard net unit `sac-ceiling-net` fires at the same epoch regardless of monitors) — final state is in `weights/`, logs, and this notebook.
## Results (campaign-v1 — filled by #57)
**Run**: 2500/2500 chunks · **14 h 07 m** autonomous (00:21:55 → ~14:28 CEST 2026-08-22) · round_counter 38912 · **zero crashes** · graceful banner `>>> training complete: 2500 chunks` · systemd net never fired. Budget was chunk-based (see phase log) — the "25k-rounds" label was a misnomer.
**Per-opponent training win rates** (run-3 slice of `training_log.jsonl`):
| Opponent | Win rate | Record | Note |
|----------|----------|--------|------|
| SacTwin | **90.7%** | 2693/2970 | vs frozen past-self — genuine self-play gain |
| Crazy | 48.3% | 2965/6140 | doubled from 26% early-run |
| Target | 9.6% | — | static, barely moved |
| Corners | 7.2% | — | walls untouched |
| RamFire | **0%** | 0/2920 | mirrors the PPO-era ladder — ram-class needs dedicated pressure |
**Eval vs Corners (pinned benchmark)**: 1287 evals · overall mean **7.7%** · histogram headline: `0/10 = 867 (67%)`, spikes ≥8/10 = **12** (incl. **5× perfect 10/10**) · final eval 10%. Stdout log carries no timestamps; timing reconstructed from file mtimes.
**Best-zip paradox**: `sac_best.zip` frozen since **01:01:57** — a single lucky 10/10 at ~round 3.5k wrote `best_score=100`, and no later eval could outrank a perfect score (even genuine ~50%-winrate stretches elsewhere). Best checkpoint = lottery ticket, decoupled from the steady-state policy (which sat at 0–10% vs Corners).
**Verdict: RETUNE BEFORE SCALING** — full rationale in #57 resolution comment. Five code-level levers staged for human sign-off: (1) eval rotation across pool + moving-average best gating; (2) reward shaping toward damage/aggression incl. anti-ram signal; (3) training-loss/step metrics logged from the training thread (#55 gap #1); (4) gate `sendTrainingMsg` off in eval mode (#55 gap #3); (5) optional stability knobs (lower LR / entropy coeff) once loss curves exist.
---
# Campaign Notebook — campaign-v2 (SAC_LSTM_Bot)
## Locked config (campaign-v2)
Launched verbatim from orchestrator mandate (umbrella ticket #59, all five levers approved & implemented in `a07e530` + `6fc01eb`):
```bash
cd /home/davide/Projects/SirRoboGarage/SAC_LSTM_Bot && \
SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:2,Target:1,SacTwin:1' \
SAC_EVAL_OPPONENTS='Corners,Crazy,Target' \
SAC_EVAL_INTERVAL=2 SAC_EVAL_ROUNDS=10 \
SAC_TOTAL_ROUNDS=25000 SAC_CHUNK_SIZE=10 SAC_MAX_CRASHES=5 \
SACLSTM_HIDDEN_SIZE=256 SACLSTM_BATCH_SIZE=16 SACLSTM_SAVE_INTERVAL=5 \
./sac_train.sh 2>&1 | tee -a campaign_v2_stdout.log
```
| Knob | Value | Rationale |
|------|-------|-----------|
| `SAC_OPPONENTS` | Corners:3, Crazy:2, **RamFire:2**, Target:1, SacTwin:1 | RamFire bumped 1→2 vs v1: anti-ram shaping (lever 2, `6fc01eb`) needs exposure to fire; without samples there is no gradient signal against ram-class |
| `SAC_EVAL_OPPONENTS` | Corners,Crazy,Target | Lever-1 rotation set (default); composite = mean of per-opponent MA-5 win rates |
| `SAC_EVAL_INTERVAL/ROUNDS` | 2 / 10 | Unchanged from v1 cadence |
| `SAC_TOTAL_ROUNDS/CHUNK_SIZE/MAX_CRASHES` | 25000 / 10 / 5 | Same budget semantics as v1 (chunk-based) |
| `SACLSTM_HIDDEN_SIZE/BATCH_SIZE/SAVE_INTERVAL` | 256 / 16 / 5 | Architecture + throughput knobs carried over; LR/entropy defaults untouched = **lever-5 conditional posture** (activation decided by loss-curve evidence, not upfront) |
## Fresh start & archive
v1 state archived intact (notebook references preserved) into `SAC_LSTM_Bot/weights_v1_archive/`: `sac_latest.zip`, `sac_best.zip`, `best_score.txt`, `round_counter.txt`, `training_log.jsonl`, `eval_log.jsonl`, `campaign_stdout.log`, plus the lever-3 smoke leftover `training_metrics.jsonl` (from `src/SAC_LSTM_Bot/`, moved so v2 loss curves start clean for lever-5 reading). Main `weights/` verified empty afterwards ⇒ main bot takes the genuine random-init path (`loadOrInitFull` → `randomFull()`).
**Twin reseed**: `make_twin.sh` requires a seed zip, but the fresh-start baseline has none. Generated a fresh random-init checkpoint (hidden=256, alpha=1.0 matching `logAlpha=0`) via a throwaway Nim script against `network.nim`/`weights.nim`, seeded it as `sac_best.zip` transiently, ran `./make_twin.sh`, removed the transient copy. Verified twin dir got byte-identical fresh zips (`cmp` OK; NOT v1 zips) + `round_counter.txt=0`. No script changes needed.
## Safety net
`systemd-run --user --unit=sac-ceiling-net-v2` armed at launch: sleeps 72000 s then `tmux kill-session -t sac_campaign_v2; sleep 5; pkill -f sac_train.sh`. Unit active at 19:59:40 CEST 2026-08-22, fires **15:59:40 CEST 2026-08-23** (epoch 1787493580).
## Launch & health evidence (first ~45 min)
Launched 20:00:19 CEST 2026-08-22 (epoch 1787421619), tmux session `sac_campaign_v2`.
| Check | Evidence |
|-------|----------|
| Round counter advances | `round_counter.txt` 95→100→465 across polls; RunTraining `Counter check passed: N == N` every chunk |
| Metrics JSONL with scalars | `training_metrics.jsonl` growing (20 lines @ t+45m): full `{epoch, steps, buffer_size, drained, grad_steps, critic_loss, actor_loss, alpha_loss, alpha}` per line |
| Eval rotation cycles ≥2 | All 3 opponents EVERY cycle: Corners→Crazy→Target ×4+ cycles in `eval_log.jsonl` (10 games each per cycle) |
| MA files written | `weights/ma_history_{Corners,Crazy,Target}.txt` created at first cycle, appended since |
| Best-gate on composite only | First write exactly when composite 0.0000 > −1 (missing-file default); later 0% cycles correctly did NOT rewrite (strict improvement enforced) |
| No transitions during eval windows | Lever-4 gate active by construction (`sendTrainingMsg` drops all msgs under `SACLSTM_EVAL_MODE=1`, unit-tested); metrics epochs cluster at chunk boundaries |
| Zero crash banners | `grep -c 'crash #'` = 0 through 20 chunks |
| Sampling distribution plausible | 20 chunks: Corners 10, RamFire 5, Crazy 3, SacTwin 2, Target 0 — within small-n noise of weights (3/2/2/1/1)/9; RamFire already sampled (exposure goal met) |
Twin liveness: own `round_counter.txt` advancing, own `sac_latest.zip` updating during SacTwin chunks, own metrics file separate from the main bot's.
### Watch items (not blockers)
1. **Training-throughput signature**: metrics lines consistently show `steps=1, buffer_size=24 (= burnIn 8 + trainWindow 16, i.e. exact canSample threshold), drained=1` — one gradient step per pass at threshold-crossing moments rather than large drain bursts. Mechanism unexplained by static code read (per-tick sends should yield bigger bursts); v1 learned to its score-60 state under the same integration code without instrumentation, so learning is not obviously broken — but effective grad-steps/hour is THE number to check at first review. This is precisely what lever-3 instrumentation exists to surface.
2. **Early critic-loss spikes**: two `1.56e16` outliers (t+6:16, t+12:04) amid otherwise sane values (~8–35) — same class as the random-init artifact flagged in #59 phase-2 notes (~1e12 there); expect decay. If persistent past early chunks, feeds the lever-5 decision.
3. **`.part` corpses**: three `sac_latest.zip.tmp.*.part` files accumulated mid-run (libzip interrupted-write artifact, #57 forensics); harmless — atomic renames keep the main zips valid, startup sweep clears them next restart.
## ~21:30 — v2 attempt-1 checkpoint finding
Attempt-1's rolling zip was dead ~40 min at birth: `SACLSTM_SAVE_INTERVAL` counts **gradient steps**, and each training pass inside the short battle processes carries exactly **1 step** (the `steps=1, drained=1` signature of watch item 1) ⇒ interval-5 never reached its trigger within a process lifetime — no `sac_latest.zip` roll ever fired despite healthy training. Same class as v1's 23:32 incident, resurfacing through a different seam (per-process step budget vs per-pass step count).
**Fix (env-level, no rebuild)**: `SACLSTM_SAVE_INTERVAL=1`, restart 21:09:30 CEST. Saves verified flowing: 21:09 / 21:13 / 21:16 (and still flowing at audit time — `sac_latest.zip` mtime 21:22:28, `sac_best.zip` 21:12:50). The locked-config table above keeps the original attempt-1 launch line for the record; live relaunch differs only in this knob.
## ~21:30 — twin-freeze contract corrected
`make_twin.sh` twin launcher now exports `SACLSTM_EVAL_MODE=1` (commit `19f34ab`). Without lever-4's gate on the twin side, the **v1 SacTwin had been TRAINING throughout**, not frozen — every mirror battle updated the twin's own weights. Consequence: v1's headline "**90.7% vs twin**" is retroactively an **arms-race win rate** (both policies co-evolving), not a fixed-benchmark score, and the #57 verdict phrasing implying a frozen sparring partner is corrected by this entry. No numbers change; the story does.
## ~21:30 — net extended + pacing decision
Slowdown attribution complete: **89% of wall clock = eval rotation by design** (`SAC_EVAL_INTERVAL=2` × 3 opponents × 10 rounds each); trainer exonerated at **~500 ticks/s**. Decision recorded in issue **#60**.
Ceiling net re-armed to match the extended budget: fires **Thu 2026-08-27 07:59:40 CEST (epoch 1787810380)** — supersedes the 2026-08-23 15:59:40 fire noted under Safety net. Attempt-1 artifacts preserved in `/tmp/v2_attempt1_backup/`.
# Campaign v2 — attempt-3 (2026-08-23)
## ~06:45 — LEVER-5 TRIGGERED (watchman shift 2, 02:11–06:22 CEST)
Evidence-gated activation of the lever-5 conditional posture (`SACLSTM_LR_CRITIC`), human signed off. Trigger evidence:
| Signal | Observation |
|--------|-------------|
| actor_loss | \|34M\| → \|106M\| monotone (+15M/h), **no plateau** |
| critic_loss | spikes >1e12 in **100%** of NEW metric records; campaign share 80.2% |
| Signature | exact `1.5625e16` recurring (= float32 saturation neighborhood) |
| alpha | decayed 0.951 → 0.711 (entropy collapse under runaway Q scale) |
| best_score.txt | frozen 00:32 (=83.3333) through 5h50m of composite oscillation 3.3–43.3 |
| Per-opponent MAs | whipsawing (Crazy 30→100→90→0→60→90→10→0) |
## DECISION — LR_CRITIC=1e-4, single knob, fresh start
- **Adjudication**: primary suspect is critic/Q value-scale growth. Actor loss inherits Q magnitude through the policy gradient, so the actor explosion is downstream; alpha decay is a *symptom* (entropy temperature chasing a blown value scale), not a cause. Therefore ONE knob moves: `SACLSTM_LR_CRITIC=1e-4` (critic learns slower → Q estimates stay estimable); LR_ACTOR/LR_ALPHA stay at default 3e-4, TARGET_ENTROPY −4. Multi-knob changes would confound attempt-3's attribution.
- **Fresh start over warm start**: attempt-2's weights carry saturated Q representations; warm-starting them under a new LR would inherit the pathology we're trying to escape (contamination risk). Attempt-2 archived intact, nothing discarded.
## MA-wipe root cause FOUND & FIXED (`167bcc4`)
Watchman anomaly explained: `ma_history_*.txt` held exactly one leading-space value per file (e.g. `[ 10 ]`, `[ 20 ]`, `[ 0 ]` in the attempt-2 backup) instead of the designed 5-value history, degrading the composite best-gate to last-cycle mean — which fully accounts for the composite whipsaw 3.3–43.3.
Root cause: in `eval_checkpoint`, the append line was
```bash
printf '%s\n' "$(cat "$(ma_hist_file "$opp")")" "$wr" | tail -n 5 | tr '\n' ' ' > "$(ma_hist_file "$opp")"
```
Bash sets up the `>` redirect **before** running the command substitution, so `$(cat f)` always read the freshly truncated file ⇒ every cycle wiped history to `" $wr "`. Reproduced standalone on bash 5.3 before touching the script; no other writer exists (grep). Fix hoists the read into its own statement and word-splits it so `tail -n 5` keeps exactly the last MA_WINDOW values. Verified live in attempt-3: two eval cycles produced `[0 0 ]` per opponent (previously impossible).
## Attempt-3 launch (2026-08-23 07:01:28 CEST, epoch 1787461288)
- Attempt-2 stopped cleanly (tmux kill; no stray procs); artifacts (76 items, 336 MB incl. `.part` corpses + MA/best/latest/counter) → `/tmp/v2_attempt2_backup/`; final round_counter **8452**; `weights/` verified empty.
- Twin regenerated via documented transient-seed method: throwaway Nim script (`initSACTrainer(35,4)` + `saveWeights`, hidden=256 env, alpha=1.0) seeded as transient `weights/sac_best.zip` → `./make_twin.sh` → twin zips byte-identical to seed (`cmp` OK), twin counter=0 → transient removed.
- Launch line = locked config verbatim plus the two live deltas:
```bash
SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:2,Target:1,SacTwin:1' \
SAC_EVAL_OPPONENTS='Corners,Crazy,Target' \
SAC_EVAL_INTERVAL=2 SAC_EVAL_ROUNDS=10 \
SAC_TOTAL_ROUNDS=25000 SAC_CHUNK_SIZE=10 SAC_MAX_CRASHES=5 \
SACLSTM_HIDDEN_SIZE=256 SACLSTM_BATCH_SIZE=16 SACLSTM_SAVE_INTERVAL=1 \
SACLSTM_LR_CRITIC=1e-4 \
./sac_train.sh 2>&1 | tee -a campaign_v2_stdout.log
```
- Safety net: existing transient unit `sac-ceiling-net-v2.service` re-verified active; fire epoch start+384587 s = **1787810380 = Thu 2026-08-27 07:59:40 CEST exactly** (delta 0 vs mandate) — kept, not duplicated.
### Health baseline (t+15 m)
| Check | Evidence |
|-------|----------|
| Counter | 33 (t+8m) → **125** (t+15m) |
| Metrics JSONL | flowing (6 lines); baseline line: `critic_loss=23.09, actor_loss=-3.228, alpha_loss=0.0, alpha=0.99970` (random-init magnitudes — attempt-3's divergence curve starts here); t+15m line: critic 84.5, actor −9.99, alpha 0.9937 |
| Saves | `SAVE_INTERVAL=1`: `sac_latest.zip` written from first grad-step onward |
| Eval rotation | full cycle landed: Corners/Crazy/Target ×10 games each in `eval_log.jsonl` |
| MA files | `[0 0 ]` per opponent after two cycles — fix confirmed in vivo |
| Best gate | `best_score.txt=0.0000` written on first composite (0 > −1 default) |
| Crash banners | 0 |
### Progress graphs
One live dashboard: `docs/campaign_dashboard.svg` (current run only — test wins, real-fight wins, losses, alpha, throughput; auto-reloads every 60 s when open in Chrome). Keep it fresh with `tools/watch_dashboard.sh` (regenerates every 60 s), or one-shot `python3 tools/plot_progress.py` (pure stdlib; paths overridable via argv, `--selftest` for sanity check).
## ~10:47 — the dashboard's axes were upside-down since creation
The progress graphs have been lying since they were made: 0% was drawn at the TOP of every panel and the newest games appeared on the LEFT. The cause is a one-line formula bug in `tools/plot_progress.py`: `map_fn` interpolated as `p1 - t*(p1-p0)` instead of `p0 + t*(p1-p0)`, so all five panels plotted `100 - value` on y and reversed time on x. Tick labels were computed by separate (correct) code, which is why the numbers on the axes never matched the ink.
Fix + guard: formula corrected; every call site audited (panels 1/2/5 use `map_fn` for both axes and are fixed by the same line; panels 3/4 already used a correct local x-lambda; no other consumer of `map_fn` exists in the repo). `--selftest` now renders a known rising series through the full build path and fails loudly unless higher value = smaller SVG y and newer data = further right — proven to catch this exact bug when the old formula is re-injected. Dashboard regenerated from live logs.
Honest status while reading the now-correct charts: run 3 is only hours old and winning ~0% — recent evals are 0/10 vs Corners, Crazy and Target alike, and real-fight buckets sit at 0–1 wins per 100 games. Expected for a fresh brain. Night-1's gains were real but were intentionally reset by the stability restart that began attempt-3; the curve starts from zero again here.
+61
View File
@@ -0,0 +1,61 @@
#!/usr/bin/env bash
# make_twin.sh — #54 mirror-twin sparring partner generator.
#
# Builds a self-contained `SacTwin` bot dir inside the sample-bots archive so
# RunTraining.java resolves it like any sample bot ($SAMPLE_BOTS_DIR/<name>).
# The twin is the SAME binary as SAC_LSTM_Bot but with:
# - distinct identity (SacTwin.json; SACLSTM_BOT_JSON overrides the baked-in
# src json — loadBotInfo gives json total precedence over env, #49 lesson)
# - its OWN weights dir, seeded from a FROZEN copy of weights/sac_best.zip
# (no checkpoint write races with the main bot)
# - its OWN round_counter.txt (main liveness guard untouched)
#
# Re-running resets the twin to the frozen baseline (reproducible opponent).
#
# Usage: ./make_twin.sh [target-archive-dir]
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
TARGET="${1:-${SAMPLE_BOTS_DIR:-/home/davide/Projects/tank-royale/sample-bots/java/build/archive}}/SacTwin"
BIN="$SCRIPT_DIR/SAC_LSTM_Bot" # nimble build -d:release output
SEED="$SCRIPT_DIR/weights/sac_best.zip" # frozen baseline
[ -x "$BIN" ] || { echo ">>> $BIN missing — run 'nimble build -d:release' first"; exit 1; }
[ -f "$SEED" ] || { echo ">>> $SEED missing — need at least one eval'd checkpoint"; exit 1; }
mkdir -p "$TARGET/weights"
cp "$BIN" "$TARGET/SAC_LSTM_Bot"
cat > "$TARGET/SacTwin.json" <<EOF
{
"name": "SacTwin",
"version": "0.1.0",
"authors": ["Davide Cappellini"],
"description": "Mirror-twin sparring partner of SAC_LSTM_Bot (#54), generated by make_twin.sh — own weights dir, frozen seed",
"homepage": "",
"countryCodes": ["IT"],
"gameTypes": ["classic", "melee", "1v1"],
"platform": "Nim",
"programmingLang": "Nim"
}
EOF
cat > "$TARGET/SacTwin.sh" <<'EOF'
#!/bin/sh
# Twin launcher (#54): identity + weights fully decoupled from the main bot.
DIR="$(cd "$(dirname "$0")" && pwd)"
export SACLSTM_BOT_JSON="$DIR/SacTwin.json"
export SACLSTM_WEIGHTS_PATH="$DIR/weights/sac_latest.zip"
# Freeze contract (#54): the twin is a FROZEN sparring partner, not a
# co-learner. Eval-mode gate (lever 4) suppresses all sendTrainingMsg traffic,
# so the twin never trains — not even in-RAM within a battle. Without this,
# any main-bot checkpoint-interval change would let twin drift accumulate.
export SACLSTM_EVAL_MODE=1
exec "$DIR/SAC_LSTM_Bot"
EOF
chmod +x "$TARGET/SacTwin.sh" "$TARGET/SAC_LSTM_Bot"
cp "$SEED" "$TARGET/weights/sac_latest.zip"
cp "$SEED" "$TARGET/weights/sac_best.zip"
echo 0 > "$TARGET/weights/round_counter.txt"
echo ">>> twin ready: $TARGET (seeded from $(stat -c %y "$SEED" | cut -d. -f1) snapshot of sac_best.zip)"
+181
View File
@@ -0,0 +1,181 @@
#!/usr/bin/env bash
# sac_train.sh — #49 training orchestration for SAC_LSTM_Bot.
#
# Drives chunked self-play via tools/training_runner/RunTraining.java (which
# owns server lifecycle, opponent connection and dead-bot liveness detection
# through weights/round_counter.txt), samples opponents by weight per chunk,
# runs deterministic evaluation (SACLSTM_EVAL_MODE=1) every N chunks, and keeps
# the best checkpoint (weights/sac_best.zip) by a moving-average composite over
# the eval opponent set (campaign v2 lever 1, #59).
#
# Config (env vars):
# SAC_OPPONENTS "Name:weight,Name:weight,..." (default below)
# SAC_TOTAL_ROUNDS total training-round budget (default 100)
# SAC_CHUNK_SIZE rounds per RunTraining battle (default 10)
# SAC_EVAL_INTERVAL eval every N chunks (default 2)
# SAC_EVAL_ROUNDS rounds per evaluation battle (default 10)
# SAC_EVAL_OPPONENTS comma-separated eval set (default Corners,Crazy,Target)
# — each cycle evaluates EVERY one; results all land in
# eval_log.jsonl (lines carry "opponent":"Name")
# SAC_MAX_CRASHES consecutive crashes before abort (default 5)
# SAC_LOG_FILE / SAC_EVAL_LOG_FILE (JSON-lines logs)
# SACLSTM_* passed through to the bot (UTD_RATIO, BATCH_SIZE, ...)
set -uo pipefail
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
REPO_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
RUNNER_DIR="$REPO_ROOT/tools/training_runner"
JAR="${TANK_ROYALE_JAR:-/home/davide/Projects/tank-royale/runner/examples/lib/robocode-tankroyale-runner.jar}"
export PPO_BOT_DIR="$SCRIPT_DIR" # runner launches THIS bot dir
export BOT_NAME="${BOT_NAME:-SAC_LSTM_Bot}" # RunTraining result matching
export SAMPLE_BOTS_DIR="${SAMPLE_BOTS_DIR:-/home/davide/Projects/tank-royale/sample-bots/java/build/archive}"
# Liveness contract (#49): RunTraining.java watches $BOT_DIR/weights/round_counter.txt
# and integration.bumpRoundCounter() writes it next to the weights — so the bot's
# weights path is pinned here, NOT env-overridable.
export SACLSTM_WEIGHTS_PATH="$SCRIPT_DIR/weights/sac_latest.zip"
WEIGHTS_DIR="$(dirname "$SACLSTM_WEIGHTS_PATH")"
OPPONENTS="${SAC_OPPONENTS:-Corners:3,Crazy:2,RamFire:1,Target:1}"
TOTAL_ROUNDS="${SAC_TOTAL_ROUNDS:-100}"
CHUNK_SIZE="${SAC_CHUNK_SIZE:-10}"
EVAL_INTERVAL="${SAC_EVAL_INTERVAL:-2}"
EVAL_ROUNDS="${SAC_EVAL_ROUNDS:-10}"
EVAL_OPPONENTS="${SAC_EVAL_OPPONENTS:-Corners,Crazy,Target}"
MA_WINDOW=5 # lever 1 (#59): per-opponent moving average over last N evals
MAX_CRASHES="${SAC_MAX_CRASHES:-5}"
LOG_FILE="${SAC_LOG_FILE:-$SCRIPT_DIR/training_log.jsonl}"
EVAL_LOG_FILE="${SAC_EVAL_LOG_FILE:-$SCRIPT_DIR/eval_log.jsonl}"
CLASSES_DIR="/tmp/opencode/sac_train_classes"
echo "=== SAC_LSTM_Bot training harness ==="
echo "Opponents: $OPPONENTS | budget: $TOTAL_ROUNDS rounds in chunks of $CHUNK_SIZE"
echo "Eval: every $EVAL_INTERVAL chunks, $EVAL_ROUNDS rounds vs [$EVAL_OPPONENTS], MA-$MA_WINDOW composite best-gating"
echo "Weights: $SACLSTM_WEIGHTS_PATH"
# ── compile bot + java runner ─────────────────────────────────────────────────
(cd "$SCRIPT_DIR" && nimble build -d:release) || { echo ">>> bot build failed"; exit 1; }
mkdir -p "$WEIGHTS_DIR" "$CLASSES_DIR"
# Startup sweep: atomic saves leave sac_latest.zip.tmp.*.part corpses behind if
# a run is SIGKILLed (libzip modify-path, see notebook forensics) — clear them.
rm -f "$WEIGHTS_DIR"/sac_latest.zip.tmp.*.part "$WEIGHTS_DIR"/sac_latest.zip.tmp
javac -cp "$JAR" -d "$CLASSES_DIR" "$RUNNER_DIR/RunTraining.java" || { echo ">>> javac failed"; exit 1; }
# ── weighted opponent pick over "Name:w,Name:w" ───────────────────────────────
pick_opponent() {
local total=0 p name w r
local pairs
IFS=',' read -ra pairs <<< "$OPPONENTS"
for p in "${pairs[@]}"; do total=$(( total + ${p##*:} )); done
r=$(( RANDOM % total ))
for p in "${pairs[@]}"; do
name="${p%%:*}"; w="${p##*:}"
if (( r < w )); then echo "$name"; return; fi
r=$(( r - w ))
done
echo "${pairs[0]%%:*}"
}
run_battle() { # $1=opponent $2=rounds $3=log file
PPOB_LOG_FILE="$3" java -cp "$CLASSES_DIR:$JAR" RunTraining "$1" "$2"
}
# ── Lever 1 (#59): eval rotation + MA best-gating ─────────────────────────────
# best_score.txt FORMAT CHANGE: it used to store the single-opponent integer
# win rate (%); that semantics is retired. It now stores the COMPOSITE score —
# the mean over SAC_EVAL_OPPONENTS of each opponent's moving average (last
# MA_WINDOW eval win rates, %). sac_best.zip is rewritten only when the
# composite strictly improves.
ma_hist_file() { echo "$WEIGHTS_DIR/ma_history_$1.txt"; }
composite_of() { # reads one "w w w ..." history line per opponent on stdin
awk -v W="$MA_WINDOW" '
NF > 0 { n=NF; k=(n>W)?W:n; s=0; for(j=n-k+1;j<=n;j++) s+=$j; tot+=s/k; c++ }
END { if (c>0) printf "%.4f", tot/c; else print "-1" }'
}
eval_checkpoint() {
# ponytail: opponent names are split by whitespace — fine for Tank Royale bot
# names (no spaces); switch to a mapfile IFS=',\n' read if that ever changes.
local opps=(${EVAL_OPPONENTS//,/ })
local tmp="$EVAL_LOG_FILE.tmp" otmp opp wins rounds wr composite best
: > "$tmp"
for opp in "${opps[@]}"; do
otmp="$EVAL_LOG_FILE.$opp.tmp"
: > "$otmp"
echo ">>> [eval] $EVAL_ROUNDS deterministic rounds vs $opp"
if ! SACLSTM_EVAL_MODE=1 run_battle "$opp" "$EVAL_ROUNDS" "$otmp"; then
rm -f "$otmp" "$tmp"
echo ">>> [eval] crashed vs $opp — keeping previous best"
return 0
fi
wins=$(grep -c '"win":true' "$otmp" || true)
rounds=$(grep -c '"type":"game"' "$otmp" || true)
if (( rounds == 0 )); then
rm -f "$otmp" "$tmp"
echo ">>> [eval] no results vs $opp — keeping previous best"
return 0
fi
wr=$(( 100 * wins / rounds ))
echo ">>> [eval] win rate: $wins/$rounds ($wr%) vs $opp"
cat "$otmp" >> "$tmp"; rm -f "$otmp"
# Per-opponent history: append this cycle's win rate, keep last MA_WINDOW
# values on one line. Read FIRST, separately: `$(cat f)` inside a command
# redirected `> f` executes against the already-truncated file (bash sets
# up the redirect before running the substitution) — every cycle wiped the
# history back to a single leading-space value (#60). $hist is UNQUOTED on
# purpose: word-splitting turns the stored line into one value per line so
# tail keeps the last MA_WINDOW values.
local hist
hist="$(cat "$(ma_hist_file "$opp")" 2>/dev/null)"
printf '%s\n' $hist "$wr" \
| tail -n "$MA_WINDOW" | tr '\n' ' ' > "$(ma_hist_file "$opp")"
done
mv "$tmp" "$EVAL_LOG_FILE"
composite=$(for opp in "${opps[@]}"; do cat "$(ma_hist_file "$opp")"; echo; done | composite_of)
# ponytail: best-score state is a plain file next to the checkpoint; survives
# harness restarts, no lock needed (single harness instance assumed).
best=$(cat "$WEIGHTS_DIR/best_score.txt" 2>/dev/null)
[ -z "$best" ] && best=-1
if awk -v a="$composite" -v b="$best" 'BEGIN{exit !(a+0 > b+0)}' \
&& [ -f "$SACLSTM_WEIGHTS_PATH" ]; then
echo "$composite" > "$WEIGHTS_DIR/best_score.txt"
cp "$SACLSTM_WEIGHTS_PATH" "$WEIGHTS_DIR/sac_best.zip"
echo ">>> [eval] new best composite ($composite) -> sac_best.zip"
fi
}
NUM_CHUNKS=$(( (TOTAL_ROUNDS + CHUNK_SIZE - 1) / CHUNK_SIZE ))
fails=0
chunk=1
# while, not `for chunk in $(seq ...)`: a crash on the FINAL chunk must rerun
# it (#54 — seq list is exhausted by then, so ((chunk--));continue fell through
# and the harness exited 0 with the budget incomplete).
while (( chunk <= NUM_CHUNKS )); do
ROUNDS=$CHUNK_SIZE
(( TOTAL_ROUNDS - (chunk - 1) * CHUNK_SIZE < CHUNK_SIZE )) && \
ROUNDS=$(( TOTAL_ROUNDS - (chunk - 1) * CHUNK_SIZE ))
OPP=$(pick_opponent)
echo "=== Chunk $chunk/$NUM_CHUNKS: $ROUNDS rounds vs $OPP ==="
if ! run_battle "$OPP" "$ROUNDS" "$LOG_FILE"; then
fails=$(( fails + 1 ))
if (( fails >= MAX_CRASHES )); then
echo ">>> aborted: $fails consecutive crashes (bot process dying?)"
exit 1
fi
# Crash recovery: RunTraining's liveness detection exited; the bot reloads
# its latest checkpoint on restart, so just rerun this chunk.
echo ">>> crash #$fails — restarting chunk from latest checkpoint"
(( chunk-- )); continue
fi
fails=0
(( chunk % EVAL_INTERVAL == 0 )) && eval_checkpoint
((chunk += 1))
done
echo ">>> training complete: $NUM_CHUNKS chunks. Logs:"
echo " training: $LOG_FILE"
echo " eval: $EVAL_LOG_FILE"
[ -f "$WEIGHTS_DIR/sac_best.zip" ] && \
echo " best: $WEIGHTS_DIR/sac_best.zip (composite $(cat "$WEIGHTS_DIR/best_score.txt"))"
exit 0
+2 -2
View File
@@ -1,8 +1,8 @@
{ {
"name": "Recurrent Royalty", "name": "SAC_LSTM_Bot",
"version": "0.1.0", "version": "0.1.0",
"authors": ["Davide Cappellini"], "authors": ["Davide Cappellini"],
"description": "SAC+LSTM Tank Royale bot — skeleton with radar lock", "description": "SAC+LSTM Tank Royale bot — self-reported identity; MUST match the name in ../SAC_LSTM_Bot.json (booter identity) or the training runner never sees this bot join",
"homepage": "", "homepage": "",
"countryCodes": ["IT"], "countryCodes": ["IT"],
"gameTypes": ["classic", "melee", "1v1"], "gameTypes": ["classic", "melee", "1v1"],
+299 -10
View File
@@ -1,13 +1,30 @@
## SAC_LSTM_Bot — skeleton: radar lock + "Recurrent Royalty" color scheme. ## SAC_LSTM_Bot — Recurrent SAC-v2 bot (Gitea #48).
## No RL yet. Connects, sets colors, locks radar onto enemy. ##
## Thread layout (decisions Q1–Q14, see integration.nim for plumbing):
## bot thread — this file's run(): inference only, <2ms/tick.
## training thread — permanent background SAC updates (integration.nim).
## I/O thread — atomic weight saves (integration.nim).
## No Arraymancer tensor ever crosses a thread boundary: the bot object and
## channels carry plain scalars / fixed arrays / plain seqs only.
import std/os import std/[os, math, random, algorithm]
import arraymancer except Linear
import tankroyale_botapi import tankroyale_botapi
import radar_lock import radar_lock
import SAC_LSTM_Bot/state
import SAC_LSTM_Bot/network
import SAC_LSTM_Bot/actions
import SAC_LSTM_Bot/rewards
import SAC_LSTM_Bot/integration
const botJsonPath = currentSourcePath().parentDir / "SAC_LSTM_Bot.json" # Identity json: baked-in src json by default; SACLSTM_BOT_JSON lets a mirror
# twin (#54) boot the same binary under its own name (loadBotInfo gives the
# json total precedence over env, so the twin must point at its own file).
let botJsonPath = getEnv("SACLSTM_BOT_JSON",
currentSourcePath().parentDir / "SAC_LSTM_Bot.json")
# ── Colors (Recurrent Royalty palette) ─────────────────────────────────────── # ── Colors (Recurrent Royalty palette) ───────────────────────────────────────
const const
ColBody = fromHex("#7B2FBE") ColBody = fromHex("#7B2FBE")
ColTurret = fromHex("#FFD700") ColTurret = fromHex("#FFD700")
@@ -26,35 +43,307 @@ proc applyColors() =
setBulletColor(ColBullet) setBulletColor(ColBullet)
setTracksColor(ColTracks) setTracksColor(ColTracks)
# ── Bot type ────────────────────────────────────────────────────────────────── # ── Enemy bullet tracking ────────────────────────────────────────────────────
# PPO_Bot-proven pattern: dead-reckoned fixed buffer. getBulletStates() is NOT
# used from the bot thread — its seq refcount is shared with the main thread.
type InFlightBullet = object
x, y, vx, vy, power: float64
const MaxBotBullets = 4
const MaxHidden = 512 # ponytail: cap for the plain-array LSTM persistence; raise if SACLSTM_HIDDEN_SIZE > 512
# ── Bot type — PLAIN DATA ONLY on the shared object (no tensors/heap seqs:
# each round runs a fresh bot thread; heap blocks owned by the previous
# round's thread must not be freed from another thread) ─────────────────────
type SacBot = ref object of Bot type SacBot = ref object of Bot
enemyBearing: float # last known absolute bearing to enemy enemyBearing: float # last known absolute bearing to enemy
battleId: int # main thread bumps in onGameStarted; bot thread compares
seenBattle: int # bot-thread copy for battle-change detection
newBattleSent: bool # first scan of THIS battle emits NewBattle
hasContact: bool
enemy: EnemyData
ticksSinceScan: int
# per-step reward accumulators (consumed by the next tick's transition)
dmgDealt, dmgTaken, wastedPower: float64
wallHits, hits, ramTaken: int
# pending transition (episode spans the whole battle; round end is NOT a boundary)
hasLastTrans: bool
lastState: array[STATE_DIM, float32]
lastAction: array[ACTION_DIM, float32]
rn: RewardNormalizer # Welford running stats, persists across battles
bullets: array[MaxBotBullets, InFlightBullet]
bulletCount: int
hArr, cArr: array[MaxHidden, float32] # LSTM state across rounds; zeros at battle start
# ── Plain-array <-> tensor helpers (bot thread only) ──────────────────────────
proc stateToArr(t: Tensor[float32]): array[STATE_DIM, float32] =
for i in 0 ..< STATE_DIM: result[i] = t[i]
proc actionToArr(t: Tensor[float32]): array[ACTION_DIM, float32] =
for i in 0 ..< ACTION_DIM: result[i] = t[i]
proc hiddenToTensor(arr: array[MaxHidden, float32]; n: int): Tensor[float32] =
result = newTensor[float32](n)
for i in 0 ..< n: result[i] = arr[i]
proc tensorToHidden(t: Tensor[float32]; arr: var array[MaxHidden, float32]) =
for i in 0 ..< t.shape[0]: arr[i] = t[i]
# ── Reward ────────────────────────────────────────────────────────────────────
proc takeReward(bot: SacBot; win = false, loss = false): float32 =
## Consume accumulated step events -> Welford-normalized reward (#44).
## Lever 2 (#59): pass enemy distance (frac of arena diagonal) for the
## anti-charge term; sentinel 2.0 (> ChargeDistFrac) when no contact.
var distFrac = 2.0
if bot.hasContact:
let diag = hypot(getArenaWidth().float64, getArenaHeight().float64)
distFrac = hypot(bot.enemy.x - getX(), bot.enemy.y - getY()) / diag
let raw = computeReward(
damageInflicted = bot.dmgDealt,
damageReceived = bot.dmgTaken,
wallHitTicks = bot.wallHits,
wastedShotPower = bot.wastedPower,
hitCount = bot.hits,
ramTakenCount = bot.ramTaken,
enemyDistFrac = distFrac,
win = win, loss = loss)
# Lever-2 (#59) observability: env-gated one-liner for smoke/calibration
# greps — proves hit/ram/charge terms fire and shows raw magnitudes. File
# (not stderr): the battle runner swallows bot process streams.
# ponytail: grows unbounded if left on; keep off outside smokes.
if getEnv("SACLSTM_REWARD_DEBUG") == "1" and
(bot.hits > 0 or bot.ramTaken > 0 or (distFrac < ChargeDistFrac and bot.dmgDealt <= 0.0)):
try:
let f = open(getWeightsPath().parentDir.parentDir / "reward_debug.log", fmAppend)
f.writeLine("raw=" & $raw & " hits=" & $bot.hits & " ram=" & $bot.ramTaken &
" distFrac=" & $distFrac)
f.close()
except CatchableError:
discard
bot.dmgDealt = 0; bot.dmgTaken = 0; bot.wastedPower = 0; bot.wallHits = 0
bot.hits = 0; bot.ramTaken = 0
let norm = rewards.normalize(bot.rn, raw)
rewards.update(bot.rn, raw)
norm.float32
# ── Event handlers ──────────────────────────────────────────────────────────── # ── Event handlers ────────────────────────────────────────────────────────────
# onGameStarted/onRoundStarted fire on the MAIN thread (bot thread not yet
# started or already joined) — plain-field writes only, no tensors here.
method onGameStarted*(bot: SacBot, e: GameStartedEventForBot) =
inc bot.battleId # bot thread zeroes LSTM + per-battle flags at next tick (Q4/Q12)
method onRoundStarted*(bot: SacBot, e: RoundStartedEvent) = method onRoundStarted*(bot: SacBot, e: RoundStartedEvent) =
setAdjustRadarForBodyTurn(true) setAdjustRadarForBodyTurn(true)
setAdjustRadarForGunTurn(true) setAdjustRadarForGunTurn(true)
radar_lock.init() radar_lock.init()
applyColors() applyColors()
# Per-round reset ONLY. NOT the LSTM hidden state (persists across rounds, Q4);
# NOT hasLastTrans (the pending transition spans the round boundary — episode
# ends at battle end only).
bot.hasContact = false
bot.enemy = EnemyData()
bot.ticksSinceScan = 0
bot.bulletCount = 0
method onScannedBot*(bot: SacBot, e: ScannedBotEvent) = method onScannedBot*(bot: SacBot, e: ScannedBotEvent) =
bot.enemyBearing = directionTo(getX(), getY(), e.x, e.y) bot.enemyBearing = directionTo(getX(), getY(), e.x, e.y)
# Same-tick radar lock: apply turn rate immediately so it takes effect this tick. # Same-tick radar lock: apply turn rate immediately so it takes effect this tick.
setRadarTurnRate(radar_lock.doRadar(getRadarDirection(), bot.enemyBearing)) setRadarTurnRate(radar_lock.doRadar(getRadarDirection(), bot.enemyBearing))
# Fire detection: energy drop in [0.1, 3.0] between scans (PPO_Bot heuristic).
let prevE = if bot.hasContact: bot.enemy.energy else: e.energy
let drop = prevE - e.energy
bot.enemy.hasFired = bot.hasContact and drop >= 0.1 and drop <= 3.0
if bot.enemy.hasFired:
bot.enemy.lastFirePower = drop
# Keep previous-scan deltas before overwriting (state.nim derives accel/turn rate).
bot.enemy.prevSpeed = bot.enemy.speed
bot.enemy.prevDirection = bot.enemy.direction
bot.enemy.hasPrevScan = bot.hasContact
bot.enemy.x = e.x
bot.enemy.y = e.y
bot.enemy.direction = e.direction
bot.enemy.speed = e.speed
bot.enemy.energy = e.energy
bot.hasContact = true
bot.ticksSinceScan = 0
# Q12b/Q14+#49: one NewBattle per battle, numeric scannedBotId. Identity is
# keyed on getBotName(id) training-side (integration.opponentKey), numeric
# fallback in the pre-BotListUpdate window.
if not bot.newBattleSent:
bot.newBattleSent = true
discard sendTrainingMsg(TrainingMsg(kind: tmkNewBattle, enemyId: e.scannedBotId))
# ── Run loop ────────────────────────────────────────────────────────────────── method onBulletHit*(bot: SacBot, e: BulletHitBotEvent) =
bot.dmgDealt += e.damage
inc bot.hits
method onHitByBullet*(bot: SacBot, e: HitByBulletEvent) =
bot.dmgTaken += e.damage
# Lever 2 (#59): anti-ram — every BotHitBotEvent receipt means a bot-bot
# collision happened and we ate RAM_DAMAGE (server deals 0.6 to both parties;
# only the hitter gets notified). Flat per-event penalty; being rammed without
# hitting back stays event-invisible.
# ponytail: enemy-initiated rams undetected — add energy-residual detection if
# v2 battle data shows ram-heavy losses.
method onHitBot*(bot: SacBot, e: BotHitBotEvent) =
inc bot.ramTaken
method onHitWall*(bot: SacBot, e: BotHitWallEvent) =
inc bot.wallHits
method onBulletHitWall*(bot: SacBot, e: BulletHitWallEvent) =
if e.bullet.ownerId == getMyId():
bot.wastedPower += e.bullet.power
method onGameAborted*(bot: SacBot) =
# Mid-round abort: drop the pending transition rather than leak it into the
# next battle's data.
bot.hasLastTrans = false
bot.dmgDealt = 0; bot.dmgTaken = 0; bot.wastedPower = 0; bot.wallHits = 0
bot.hits = 0; bot.ramTaken = 0
method onRoundEnded*(bot: SacBot, e: RoundEndedEventForBot) =
## Harness liveness signal (#49): RunTraining.java watches round_counter.txt
## and aborts the battle if it freezes (dead bot process). Main thread.
bumpRoundCounter()
method onGameEnded*(bot: SacBot, e: GameEndedEventForBot) =
## Battle end -> terminal transition with done=true. Main thread; the API has
## joined the bot thread before this fires, so these plain fields are quiescent.
if bot.hasLastTrans:
let r = bot.takeReward(win = e.results.rank == 1, loss = e.results.rank != 1)
discard sendTrainingMsg(TrainingMsg(kind: tmkTransition,
state: bot.lastState, action: bot.lastAction,
reward: r, nextState: bot.lastState, done: true))
bot.hasLastTrans = false
# ── Run loop (bot thread) ─────────────────────────────────────────────────────
method run(bot: SacBot) = method run(bot: SacBot) =
randomize()
# Per-thread locals: born and freed on THIS thread, every round. Nothing
# heap-owned survives the round boundary except the bot object's plain fields.
var actor: ActorNet
var actorReady = false
var myVersion = 0
var myHidden = 0
var flat: seq[float32]
while isRunning(): while isRunning():
# Spin radar when no enemy is visible (full sweep). # Battle boundary (Q4/Q12): zero LSTM persistence + per-battle flags.
if bot.enemyBearing == 0.0: if bot.battleId != bot.seenBattle:
setRadarTurnRate(45.0) bot.seenBattle = bot.battleId
bot.newBattleSent = false
zeroMem(addr bot.hArr, sizeof(bot.hArr))
zeroMem(addr bot.cArr, sizeof(bot.cArr))
# Weight sync (Q7/Q11): always-latest; rebuild this thread's tensors on change.
if pullWeights(myVersion, myHidden, flat):
if myHidden > MaxHidden:
# hArr/cArr are fixed-capacity; a bigger SACLSTM_HIDDEN_SIZE would
# heap-overflow them in tensorToHidden. Loud misconfig beats corruption.
raise newException(ValueError, "SACLSTM_HIDDEN_SIZE=" & $myHidden &
" exceeds MaxHidden=" & $MaxHidden & " (bot-side LSTM persistence cap)")
var cur = 0
actor = actorFromFlat(flat, cur, myHidden)
actorReady = true
# Spawn an enemy bullet when a fresh scan shows they fired; then advance and
# prune the tracked bullets (positions feed state slots 22–33).
if bot.hasContact and bot.enemy.hasFired and bot.bulletCount < MaxBotBullets:
let p = bot.enemy.lastFirePower
let spd = 20.0 - 3.0 * p
let ang = arctan2(getY() - bot.enemy.y, getX() - bot.enemy.x)
bot.bullets[bot.bulletCount] = InFlightBullet(x: bot.enemy.x, y: bot.enemy.y,
vx: spd * cos(ang), vy: spd * sin(ang), power: p)
inc bot.bulletCount
let aW = float64(getArenaWidth())
let aH = float64(getArenaHeight())
var alive = 0
for i in 0 ..< bot.bulletCount:
let b = bot.bullets[i]
let nx = b.x + b.vx
let ny = b.y + b.vy
if nx >= 0.0 and nx <= aW and ny >= 0.0 and ny <= aH:
bot.bullets[alive] = InFlightBullet(x: nx, y: ny, vx: b.vx, vy: b.vy, power: b.power)
inc alive
bot.bulletCount = alive
# State build (35-dim, this thread's tensor).
inc bot.ticksSinceScan
var gs: GameState
gs.x = getX()
gs.y = getY()
gs.direction = getDirection()
gs.speed = getSpeed()
gs.energy = getEnergy()
gs.gunDirection = getGunDirection()
gs.gunHeat = getGunHeat()
gs.arenaWidth = aW
gs.arenaHeight = aH
gs.hasContact = bot.hasContact
gs.enemy = bot.enemy
gs.ticksSinceLastScan = bot.ticksSinceScan
var bd: array[MaxBotBullets, BulletData]
for i in 0 ..< bot.bulletCount:
bd[i] = BulletData(x: bot.bullets[i].x, y: bot.bullets[i].y, power: bot.bullets[i].power)
if bot.bulletCount > 1: # closest threats fill slots 0-2
bd.toOpenArray(0, bot.bulletCount - 1).sort(proc(a, b: BulletData): int =
cmp(hypot(a.x - gs.x, a.y - gs.y), hypot(b.x - gs.x, b.y - gs.y)))
gs.bulletCount = min(bot.bulletCount, 3)
for i in 0 ..< gs.bulletCount:
gs.bullets[i] = bd[i]
# Consume the fired pulse AFTER the state saw it (one shot -> one bullet).
bot.enemy.hasFired = false
let stateT = buildState(gs)
# Finalize the PREVIOUS transition: reward from events since the last tick,
# nextState is this tick's observation (PPO_Bot alignment).
if actorReady and bot.hasLastTrans:
discard sendTrainingMsg(TrainingMsg(kind: tmkTransition,
state: bot.lastState, action: bot.lastAction,
reward: bot.takeReward(), nextState: stateToArr(stateT), done: false))
bot.lastState = stateToArr(stateT)
if not actorReady:
setRadarTurnRate(45.0) # no weights yet (defensive; main pre-inits) — just sweep
go()
continue
# Inference: hidden state lives as plain arrays on the bot object (persists
# across rounds); tensors are rebuilt per tick on this thread.
let h = hiddenToTensor(bot.hArr, myHidden)
let c = hiddenToTensor(bot.cArr, myHidden)
let fwd = actor.actorForward(stateT, (h: h, c: c))
tensorToHidden(fwd.lstm.h, bot.hArr)
tensorToHidden(fwd.lstm.c, bot.cArr)
bot.lastAction = actionToArr(fwd.actions)
bot.hasLastTrans = true
# Actions -> intents (go() snapshots them at send time).
let mapped = mapActions(fwd.actions, getSpeed(), getGunHeat())
setTurnRate(mapped.turnRate)
setTargetSpeed(getSpeed() + mapped.acceleration) # actions.nim contract
setGunTurnRate(mapped.gunTurnRate)
if mapped.firePower > 0.0:
discard setFire(mapped.firePower)
if not bot.hasContact:
setRadarTurnRate(45.0) # sweep until first lock (onScannedBot overrides same-tick)
go() go()
# ── Entry point ─────────────────────────────────────────────────────────────── # ── Entry point ───────────────────────────────────────────────────────────────
when isMainModule: when isMainModule:
initIntegration() # spawn training + I/O threads, seed weight snapshot
var bot = SacBot() var bot = SacBot()
start(bot, botJsonPath) start(bot, botJsonPath) # blocks until server disconnect
shutdownIntegration() # Shutdown msg -> final save -> joins
+36
View File
@@ -0,0 +1,36 @@
## actions.nim — map raw network output (4 tanh values) to bot intent fields.
##
## Note on acceleration vs targetSpeed:
## TankRoyale uses setTargetSpeed(), not setAcceleration().
## The mapped `acceleration` field is a delta; callers must compute:
## newTargetSpeed = clamp(currentSpeed + acceleration, -8.0, 8.0)
## and call setTargetSpeed(newTargetSpeed).
import arraymancer
const ACTION_DIM* = 4
type
MappedActions* = object
turnRate*: float ## degrees/tick, speed-aware; [-10, 10] at speed 0
acceleration*: float ## delta speed in [-2, +1]; caller adds to currentSpeed
gunTurnRate*: float ## degrees/tick in [-20, 20]
firePower*: float ## 0 = don't fire; (0.1, 3.0] = fire with this power
proc mapActions*(networkOutput: Tensor[float32],
currentSpeed: float,
gunHeat: float): MappedActions =
## networkOutput: [4] tensor of tanh values in [-1, 1].
let a0 = networkOutput[0].float
let a1 = networkOutput[1].float
let a2 = networkOutput[2].float
let a3 = networkOutput[3].float
result.turnRate = a0 * (10.0 - 0.75 * abs(currentSpeed))
# asymmetric accel: [-1,1] -> [-2, +1] via (value * 1.5 - 0.5)
result.acceleration = a1 * 1.5 - 0.5
result.gunTurnRate = a2 * 20.0
if a3 > 0.0 and gunHeat <= 0.0:
result.firePower = a3 * 2.9 + 0.1
else:
result.firePower = 0.0
@@ -0,0 +1,463 @@
## integration.nim — #48 thread plumbing for SAC_LSTM_Bot.
##
## Threads added by the bot (on top of the bot API's main/bot/sender):
## training thread — permanent, drain-then-train loop, owns SACTrainer +
## ReplayBuffer. Q1/Q2/Q10/Q12.
## I/O thread — cap-1 channel of weight snapshots, atomic zip saves. Q5.
##
## Cross-thread payloads are plain arrays/seqs ONLY. No Arraymancer tensor ever
## crosses a thread boundary: Tensor is a ref type and ORC refcounts are
## non-atomic — sharing them across threads is SIGSEGV territory (PPO_Bot,
## empirically confirmed). Each thread builds its own tensors from plain data.
## Decisions Q1–Q14: Gitea #48.
import arraymancer except Linear
import std/[locks, os, math, random, strutils, times]
import tankroyale_botapi # getBotName (#49 name-based opponent identity)
import SAC_LSTM_Bot/network
import SAC_LSTM_Bot/state # STATE_DIM
import SAC_LSTM_Bot/actions # ACTION_DIM
import SAC_LSTM_Bot/replay_buffer
import SAC_LSTM_Bot/training
import SAC_LSTM_Bot/weights
# ── Config ────────────────────────────────────────────────────────────────────
proc getUtdRatio*(): int =
## Gradient steps per drained transition (Q2). 200 ticks -> 200 steps at 1.
parseInt(getEnv("SACLSTM_UTD_RATIO", "1"))
proc getBatchSize*(): int =
# ponytail: 16 is an unprofiled guess sized so UTD=1 keeps up with 30tps;
# env knob is the calibration point if steps/sec falls short.
parseInt(getEnv("SACLSTM_BATCH_SIZE", "16"))
proc getSaveInterval*(): int =
parseInt(getEnv("SACLSTM_SAVE_INTERVAL", "500"))
proc getWeightsPath*(): string =
getEnv("SACLSTM_WEIGHTS_PATH",
currentSourcePath().parentDir / "weights" / "sac_latest.zip")
proc opponentKey*(enemyId: int): string =
## Q14 follow-up (#49): name-based opponent identity from the v1.0.1
## BotListUpdate table; numeric-id fallback for the window before the first
## update arrives (getBotName still returns "" then). Called on the training
## thread — the lookup is lock-guarded in the API, no cross-thread refs.
let name = getBotName(enemyId)
if name.len > 0: name else: $enemyId
proc bumpRoundCounter*() =
## Liveness signal for tools/training_runner/RunTraining.java (#49): one
## increment per round end; the runner aborts when it freezes (dead bot).
# ponytail: non-atomic read-modify-write; single writer (main thread) and the
# runner re-polls every 500ms with multi-round tolerance, torn reads self-heal.
let p = getWeightsPath().parentDir / "round_counter.txt"
var n = 0
try:
n = parseInt(readFile(p).strip())
except CatchableError:
discard # absent/garbage -> start at 1
try:
createDir(p.parentDir)
writeFile(p, $(n + 1))
except CatchableError:
discard # counter is best-effort liveness; never kill an event handler
# ── Channel message (plain data only) ────────────────────────────────────────
type
TrainingMsgKind* = enum tmkTransition, tmkNewBattle, tmkShutdown
TrainingMsg* = object
case kind*: TrainingMsgKind
of tmkTransition:
state*: array[STATE_DIM, float32] # Q13 plain arrays, no tensors
action*: array[ACTION_DIM, float32]
reward*: float32
nextState*: array[STATE_DIM, float32]
done*: bool # true only at battle end
of tmkNewBattle:
enemyId*: int # Q14 numeric scannedBotId
of tmkShutdown:
discard
proc arrToTensor*[N: static int](arr: array[N, float32]): Tensor[float32] =
result = newTensor[float32](N)
for i in 0 ..< N: result[i] = arr[i]
# ── Flat weight snapshots ─────────────────────────────────────────────────────
# Layout mirrors network.nim's fixed architecture (fc1 -> hiddenDim-wide LSTM
# with [4h, 2h] combined weights, 128-wide fc2, 4-out heads). The asserts catch
# layout drift if network.nim shapes ever change.
proc actorSize*(h: int): int = 8*h*h + 168*h + 1160
proc criticSize*(h: int): int = 8*h*h + 172*h + 257
proc putT(t: Tensor[float32]; dst: var seq[float32]; c: var int) =
for v in t:
dst[c] = v
inc c
proc takeT(src: seq[float32]; c: var int; rows, cols: int): Tensor[float32] =
# seq slice copies, then toTensor copies again: result owns its memory —
# never a view into src (src may be a cross-thread buffer).
let n = rows * cols
result = src[c ..< c + n].toTensor().reshape(rows, cols)
c += n
proc takeV(src: seq[float32]; c: var int; n: int): Tensor[float32] =
## Rank-1 vector (biases) — reshape(n) keeps rank 1.
result = src[c ..< c + n].toTensor().reshape(n)
c += n
proc packActor*(a: ActorNet; dst: var seq[float32]; c: var int) =
putT(a.fc1.w, dst, c); putT(a.fc1.b, dst, c)
putT(a.lstm.wCombined, dst, c); putT(a.lstm.bCombined, dst, c)
putT(a.fc2.w, dst, c); putT(a.fc2.b, dst, c)
putT(a.muHead.w, dst, c); putT(a.muHead.b, dst, c)
putT(a.logStdHead.w, dst, c); putT(a.logStdHead.b, dst, c)
proc packCritic*(net: CriticNet; dst: var seq[float32]; c: var int) =
putT(net.fc1.w, dst, c); putT(net.fc1.b, dst, c)
putT(net.lstm.wCombined, dst, c); putT(net.lstm.bCombined, dst, c)
putT(net.fc2.w, dst, c); putT(net.fc2.b, dst, c)
putT(net.fc3.w, dst, c); putT(net.fc3.b, dst, c)
proc actorFromFlat*(src: seq[float32]; c: var int; h: int): ActorNet =
result.fc1.w = takeT(src, c, h, 35)
result.fc1.b = takeV(src, c, h)
result.lstm.wCombined = takeT(src, c, 4*h, 2*h)
result.lstm.bCombined = takeV(src, c, 4*h)
result.fc2.w = takeT(src, c, 128, h)
result.fc2.b = takeV(src, c, 128)
result.muHead.w = takeT(src, c, 4, 128)
result.muHead.b = takeV(src, c, 4)
result.logStdHead.w = takeT(src, c, 4, 128)
result.logStdHead.b = takeV(src, c, 4)
result.lstm.hiddenDim = result.lstm.bCombined.size div 4 # same as weights.nim loadLSTMCell
result.hiddenDim = h
assert c == actorSize(h), "actor flat layout drift"
proc criticFromFlat*(src: seq[float32]; c: var int; h: int): CriticNet =
result.fc1.w = takeT(src, c, h, 39)
result.fc1.b = takeV(src, c, h)
result.lstm.wCombined = takeT(src, c, 4*h, 2*h)
result.lstm.bCombined = takeV(src, c, 4*h)
result.fc2.w = takeT(src, c, 128, h)
result.fc2.b = takeV(src, c, 128)
result.fc3.w = takeT(src, c, 1, 128)
result.fc3.b = takeV(src, c, 1)
result.lstm.hiddenDim = result.lstm.bCombined.size div 4
result.hiddenDim = h
type
FullSnap* = object
hiddenDim*: int
data*: seq[float32] # actor | critic1 | critic2 | targetCritic1 | targetCritic2 | alpha
proc packFull*(t: SACTrainer): FullSnap =
let h = t.actor.hiddenDim
result.hiddenDim = h
result.data = newSeq[float32](actorSize(h) + 4 * criticSize(h) + 1)
var c = 0
packActor(t.actor, result.data, c)
packCritic(t.critic1, result.data, c)
packCritic(t.critic2, result.data, c)
packCritic(t.targetCritic1, result.data, c)
packCritic(t.targetCritic2, result.data, c)
assert abs(t.alpha().float64 - exp(t.logAlpha.float64)) < 1e-6
result.data[c] = t.alpha()
inc c
assert c == result.data.len, "full snapshot layout drift"
proc unpackFull*(fs: FullSnap):
tuple[a: ActorNet, c1, c2, t1, t2: CriticNet, alpha: float32] =
var c = 0
result.a = actorFromFlat(fs.data, c, fs.hiddenDim)
result.c1 = criticFromFlat(fs.data, c, fs.hiddenDim)
result.c2 = criticFromFlat(fs.data, c, fs.hiddenDim)
result.t1 = criticFromFlat(fs.data, c, fs.hiddenDim)
result.t2 = criticFromFlat(fs.data, c, fs.hiddenDim)
result.alpha = fs.data[c]
# ── Shared weight snapshot (training thread writes, bot thread copies out) ────
# Q7/Q11: Lock + always-latest semantics. `data` is allocated ONCE and written
# IN PLACE under gWeightLock — it is never reassigned, so the shared heap block
# never sees cross-thread refcount traffic. Readers copy element-wise out.
type
WeightSnapshot* = object
hiddenDim*: int
version*: int
data*: seq[float32]
var gWeightLock: Lock
var gSharedSnap: WeightSnapshot
var gTrainChan: Channel[TrainingMsg]
var gSaveChan: Channel[FullSnap]
var gTrainingThread: Thread[void]
var gIoThread: Thread[void]
var gInitialFull: FullSnap # built on main before spawn; read-once after (happens-before)
var gWeightsPath: string
proc pullWeights*(myVersion: var int; hidden: var int;
flat: var seq[float32]): bool =
## Copy the latest actor snapshot out under the lock. Returns true when a new
## version arrived (caller rebuilds its tensors on ITS OWN thread).
withLock(gWeightLock):
if gSharedSnap.version == myVersion:
return false
if flat.len != gSharedSnap.data.len:
flat = newSeq[float32](gSharedSnap.data.len) # caller-thread-owned buffer
for i in 0 ..< flat.len:
flat[i] = gSharedSnap.data[i]
myVersion = gSharedSnap.version
hidden = gSharedSnap.hiddenDim
true
proc evalModeActive*(): bool {.inline.} =
## Lever 4 (#59): the harness's deterministic eval battles already run the bot
## with SACLSTM_EVAL_MODE=1 (sac_train.sh eval_checkpoint, mechanism from #49).
## While set, eval ticks must NOT feed the trainer — transitions would pollute
## the replay buffer with eval-only data and trigger gradient updates.
getEnv("SACLSTM_EVAL_MODE") == "1"
proc sendTrainingMsg*(msg: TrainingMsg): bool {.inline.} =
## Bot-side enqueue (cap-256, drops on overflow per Q10). Thread-safe.
## Lever 4 (#59): fully suppressed in eval mode — NewBattle drops too, so an
## eval battle can neither add transitions nor clear/retarget the buffer.
if evalModeActive(): return false
gTrainChan.trySend(msg)
# ── Training state (testable without threads) ─────────────────────────────────
type
TrainState* = object
trainer*: SACTrainer
buf*: ReplayBuffer
lastEnemyKey*: string # opponent identity key (#49 name-based, Q14)
stepCount*: int
nextSave*: int
proc trainerFromFull*(initial: FullSnap): SACTrainer =
let (a, c1, c2, t1, t2, alpha) = unpackFull(initial)
result.actor = a
result.critic1 = c1
result.critic2 = c2
result.targetCritic1 = t1
result.targetCritic2 = t2
assert alpha > 0.0'f32, "checkpoint alpha must be positive"
result.logAlpha = ln(alpha)
result.targetEntropy = getTargetEntropy()
result.tau = getTau()
result.lrActor = getLrActor()
result.lrCritic = getLrCritic()
result.lrAlpha = getLrAlpha()
result.gamma = getGamma()
result.adam = initSACAdamStates(result.actor, result.critic1, result.critic2)
# ponytail: adam momentum not carried in the flat snapshot — optimizer restarts
# fresh each process; switch to saveCheckpoint/loadCheckpoint end-to-end when
# resume quality matters (#49 harness owns checkpoint management).
proc initTrainState*(initial: FullSnap): TrainState =
result.trainer = trainerFromFull(initial)
result.buf = newReplayBuffer(getBufferCapacity(), STATE_DIM, ACTION_DIM)
result.lastEnemyKey = ""
result.nextSave = getSaveInterval()
# ── Training-loss metrics (campaign v2 lever 3, #59) ──────────────────────────
proc metricsFilePath*(): string =
## Sits next to the weights dir's parent: SAC_LSTM_Bot/training_metrics.jsonl
## under the #49 harness (weights live in SAC_LSTM_Bot/weights/).
getWeightsPath().parentDir.parentDir / "training_metrics.jsonl"
proc metricsLine*(epoch: float64; stepCount, bufferLen, drained, gradSteps: int;
m: SACMetrics): string =
## One JSONL line with exactly the scalars SACTrainer.sacUpdate exposes
## (#59 lever 3 — SACMetrics was already returned, no trainer change needed):
## losses/alpha averaged over this pass's gradient steps, buffer size from
## replay_buffer.len, cumulative step count and drained transition count.
"{\"epoch\":" & $epoch &
",\"steps\":" & $stepCount &
",\"buffer_size\":" & $bufferLen &
",\"drained\":" & $drained &
",\"grad_steps\":" & $gradSteps &
",\"critic_loss\":" & $m.criticLoss &
",\"actor_loss\":" & $m.actorLoss &
",\"alpha_loss\":" & $m.alphaLoss &
",\"alpha\":" & $m.alpha & "}"
proc appendMetricsLine(st: TrainState; drained, gradSteps: int; m: SACMetrics) =
## Lever 3 (#59): one append per trainPass (never per gradient step). Open,
## write, close — cheap and crash-tolerant; a metrics failure never kills
## training.
try:
let f = open(metricsFilePath(), fmAppend)
f.writeLine(metricsLine(epochTime(), st.stepCount, st.buf.len,
drained, gradSteps, m))
f.close()
except CatchableError:
discard
proc handleTrainingMsg*(st: var TrainState; msg: TrainingMsg): bool =
## Process one message. Returns false for Shutdown (caller stops).
## Tensors are born HERE from the message's plain arrays — training thread only.
case msg.kind
of tmkTransition:
st.buf.add(Transition(
state: msg.state.arrToTensor,
action: msg.action.arrToTensor,
reward: msg.reward,
nextState: msg.nextState.arrToTensor,
done: msg.done))
of tmkNewBattle:
# Q12a: keep the buffer if the opponent is unchanged, clear otherwise.
# Identity keyed on NAME (#49); numeric-id fallback pre-BotListUpdate.
let key = opponentKey(msg.enemyId)
if key != st.lastEnemyKey:
st.buf.clear()
st.lastEnemyKey = key
of tmkShutdown:
return false
true
proc trainPass*(st: var TrainState; drained: int) =
## UTD gradient steps for the transitions drained this pass, then publish the
## latest actor to the shared snapshot and request periodic disk saves.
if drained <= 0 or not st.buf.canSample:
return
let steps = drained * getUtdRatio() # Q2
var gradSteps = 0
var sumCritic, sumActor, sumAlphaLoss, sumAlpha = 0.0'f32
for i in 1 .. steps:
let seqs = st.buf.sampleSequences(getBatchSize())
if seqs.len == 0:
break
let m = sacUpdate(st.trainer, seqs)
sumCritic += m.criticLoss; sumActor += m.actorLoss
sumAlphaLoss += m.alphaLoss; sumAlpha += m.alpha
inc gradSteps
inc st.stepCount
# Save check INSIDE the step loop (#56 launch finding): at production sizes
# (hidden 256 ⇒ ~1 s/step) a drain burst queues minutes of steps; checking
# only between passes meant the process died mid-loop before stepCount ever
# reached nextSave — zero checkpoints persisted for the whole campaign.
# Mid-loop checks + SAVE_INTERVAL≤20 (#54) keep saves ~20 s apart.
if st.stepCount >= st.nextSave:
st.nextSave += getSaveInterval()
var full = packFull(st.trainer)
discard gSaveChan.trySend(move(full)) # cap-1: drop if I/O thread is busy (Q5)
if gradSteps > 0:
# Lever 3 (#59): one metrics line per pass, losses averaged over its steps.
appendMetricsLine(st, drained, gradSteps, SACMetrics(
criticLoss: sumCritic / gradSteps.float32,
actorLoss: sumActor / gradSteps.float32,
alphaLoss: sumAlphaLoss / gradSteps.float32,
alpha: sumAlpha / gradSteps.float32))
# Publish latest actor (Q7): in-place write under the lock, bump version.
withLock(gWeightLock):
assert gSharedSnap.hiddenDim == st.trainer.actor.hiddenDim,
"snapshot/trainer hidden size mismatch"
var c = 0
packActor(st.trainer.actor, gSharedSnap.data, c)
inc gSharedSnap.version
# ── Threads ───────────────────────────────────────────────────────────────────
proc trainingThreadEntry() {.thread.} =
{.cast(gcsafe).}:
randomize()
var st = initTrainState(gInitialFull)
var running = true
while running:
let first = gTrainChan.recv() # block until traffic (no busy spin)
var drained = 0
var msg = first
while running:
if msg.kind == tmkTransition:
inc drained
if not handleTrainingMsg(st, msg):
running = false # Shutdown
break
let (more, nxt) = gTrainChan.tryRecv()
if not more:
break # drained — now train (Q10)
msg = nxt
if not running:
var full = packFull(st.trainer) # final save request, then exit
discard gSaveChan.trySend(move(full))
break
trainPass(st, drained)
proc ioThreadEntry() {.thread.} =
{.cast(gcsafe).}:
while true:
let fs = gSaveChan.recv() # blocks; exits via empty-data sentinel
if fs.data.len == 0:
break
let (a, c1, c2, t1, t2, alpha) = unpackFull(fs)
saveWeights(gWeightsPath, a, c1, c2, t1, t2, alpha) # atomic zip (weights.nim)
# ── Lifecycle ─────────────────────────────────────────────────────────────────
proc randomFull(): FullSnap =
let t = initSACTrainer(STATE_DIM, ACTION_DIM) # random nets, born + freed here
packFull(t)
proc loadOrInitFull(): FullSnap =
let path = getWeightsPath()
if fileExists(path):
try:
let cp = loadCheckpoint(path)
var t = initSACTrainer(STATE_DIM, ACTION_DIM) # env hyperparams
t.actor = cp.actor
t.critic1 = cp.critic1
t.critic2 = cp.critic2
t.targetCritic1 = cp.targetCritic1
t.targetCritic2 = cp.targetCritic2
assert cp.alpha > 0.0'f32, "checkpoint alpha must be positive"
t.logAlpha = ln(cp.alpha)
result = packFull(t)
except Exception as e:
stderr.writeLine "[sac] checkpoint load failed (" & e.msg & ") — random init"
result = randomFull()
else:
result = randomFull()
proc initIntegration*() =
## Open channels, build the initial weight snapshot, spawn both threads.
## Call once from the main module before start().
gWeightsPath = getWeightsPath()
gInitialFull = loadOrInitFull()
gSharedSnap.hiddenDim = gInitialFull.hiddenDim
gSharedSnap.data = newSeq[float32](actorSize(gInitialFull.hiddenDim))
for i in 0 ..< gSharedSnap.data.len: # element-wise: no refcount traffic
gSharedSnap.data[i] = gInitialFull.data[i]
gSharedSnap.version = 1
initLock(gWeightLock)
gTrainChan.open(256) # Q10 cap-256
gSaveChan.open(1) # Q5 cap-1
# Lever 4 (#59): one-time visibility for the suppression gate (see
# sendTrainingMsg) — the eval bot trains nothing by design.
if evalModeActive():
stderr.writeLine "[sac] SACLSTM_EVAL_MODE=1 — training input suppressed (lever 4, #59)"
createThread(gTrainingThread, trainingThreadEntry)
createThread(gIoThread, ioThreadEntry)
proc shutdownIntegration*() =
## Stop both threads cleanly. Called after the bot disconnects (start returned).
while gTrainChan.tryRecv().dataAvailable:
discard # drop pending transitions — process exiting
discard gTrainChan.trySend(TrainingMsg(kind: tmkShutdown))
joinThread(gTrainingThread) # training requests one final save
# Nim's recv() blocks even on closed channels, so the I/O thread exits via an
# empty-data sentinel. Retry while it is still busy saving: every failed
# trySend means a save is in flight and will be consumed, so this terminates.
var stop = FullSnap(hiddenDim: -1)
while not gSaveChan.trySend(move(stop)):
sleep(50)
stop = FullSnap(hiddenDim: -1)
joinThread(gIoThread)
gTrainChan.close()
+161
View File
@@ -0,0 +1,161 @@
## network.nim — LSTM-based Actor and dual Critic for SAC-v2.
## No autograd; inference only. Manual LSTM cell from scratch.
import arraymancer
import std/[math, random, os, strutils]
# ── Configuration ─────────────────────────────────────────────────────────────
proc getHiddenSize*(): int =
let s = getEnv("SACLSTM_HIDDEN_SIZE", "256")
result = parseInt(s)
proc isEvalMode*(): bool =
getEnv("SACLSTM_EVAL_MODE", "0") == "1"
# ── Types ─────────────────────────────────────────────────────────────────────
type
Linear* = object
w*, b*: Tensor[float32] # w: [out, in], b: [out]
LSTMCell* = object
## Combined weight matrix Wi|Wf|Wg|Wo stacked: [4*hidden, input+hidden]
## Combined bias stacked: [4*hidden]
wCombined*: Tensor[float32]
bCombined*: Tensor[float32]
hiddenDim*: int
ActorNet* = object
fc1*: Linear
lstm*: LSTMCell
fc2*: Linear
muHead*: Linear
logStdHead*: Linear
hiddenDim*: int
CriticNet* = object
fc1*: Linear
lstm*: LSTMCell
fc2*: Linear
fc3*: Linear
hiddenDim*: int
LSTMState* = tuple[h, c: Tensor[float32]] # each [hiddenDim]
# ── Init helpers ──────────────────────────────────────────────────────────────
proc initLinear*(inDim, outDim: int; scale: float32): Linear =
result.w = randomNormalTensor[float32]([outDim, inDim]) *. scale
result.b = zeros[float32](outDim)
proc initLinearHe*(inDim, outDim: int): Linear =
initLinear(inDim, outDim, sqrt(2.0'f32 / inDim.float32))
proc initLinearOut*(inDim, outDim: int): Linear =
initLinear(inDim, outDim, sqrt(1.0'f32 / inDim.float32))
proc initLSTMCell*(inputDim, hiddenDim: int): LSTMCell =
result.hiddenDim = hiddenDim
let fanIn = (inputDim + hiddenDim).float32
let scale = sqrt(1.0'f32 / fanIn)
result.wCombined = randomNormalTensor[float32]([4 * hiddenDim, inputDim + hiddenDim]) *. scale
result.bCombined = zeros[float32](4 * hiddenDim)
proc zeroState*(hiddenDim: int): LSTMState =
result = (h: zeros[float32](hiddenDim), c: zeros[float32](hiddenDim))
proc initActorNet*(stateDim: int): ActorNet =
let h = getHiddenSize()
result.hiddenDim = h
result.fc1 = initLinearHe(stateDim, h)
result.lstm = initLSTMCell(h, h)
result.fc2 = initLinearHe(h, 128)
result.muHead = initLinearOut(128, 4)
result.logStdHead = initLinearOut(128, 4)
proc initCriticNet*(stateDim, actionDim: int): CriticNet =
let h = getHiddenSize()
result.hiddenDim = h
result.fc1 = initLinearHe(stateDim + actionDim, h)
result.lstm = initLSTMCell(h, h)
result.fc2 = initLinearHe(h, 128)
result.fc3 = initLinearOut(128, 1)
# ── Forward helpers ───────────────────────────────────────────────────────────
proc linear*(l: Linear; x: Tensor[float32]): Tensor[float32] =
l.w * x + l.b
proc relu*(x: Tensor[float32]): Tensor[float32] =
x.map(proc(v: float32): float32 = max(0.0'f32, v))
proc sigmoid*(x: Tensor[float32]): Tensor[float32] =
x.map(proc(v: float32): float32 = 1.0'f32 / (1.0'f32 + exp(-v)))
proc tanhT*(x: Tensor[float32]): Tensor[float32] =
x.map(proc(v: float32): float32 = tanh(v))
proc lstmStep*(cell: LSTMCell; x, h, c: Tensor[float32]): LSTMState =
## x: [inputDim], h/c: [hiddenDim] → h', c': [hiddenDim]
let xh = concat(x, h, axis = 0) # [inputDim + hiddenDim]
let gates = cell.wCombined * xh + cell.bCombined # [4*hidden]
let hd = cell.hiddenDim
let iGate = sigmoid(gates[0 ..< hd])
let fGate = sigmoid(gates[hd ..< 2*hd])
let gGate = tanhT(gates[2*hd ..< 3*hd])
let oGate = sigmoid(gates[3*hd ..< 4*hd])
let cPrime = fGate *. c + iGate *. gGate
let hPrime = oGate *. tanhT(cPrime)
result = (h: hPrime, c: cPrime)
# ── Actor forward ─────────────────────────────────────────────────────────────
const
LOG_STD_MIN = -5.0'f32
LOG_STD_MAX = 2.0'f32
LOG_PROB_EPS = 1e-6'f32
proc actorForward*(net: ActorNet; state: Tensor[float32]; lstm: LSTMState;
deterministic = false):
tuple[actions: Tensor[float32]; logProb: float32; lstm: LSTMState] =
## state: [stateDim], lstm: (h,c) each [hiddenDim]
## Returns actions [4], scalar logProb, updated (h',c').
let h1 = relu(net.fc1.linear(state))
let lstmOut = lstmStep(net.lstm, h1, lstm.h, lstm.c)
let h2 = relu(net.fc2.linear(lstmOut.h))
let mu = net.muHead.linear(h2)
let logStdRaw = net.logStdHead.linear(h2)
let logStd = logStdRaw.map(proc(v: float32): float32 = clamp(v, LOG_STD_MIN, LOG_STD_MAX))
if deterministic or isEvalMode():
let actions = tanhT(mu)
return (actions: actions, logProb: 0.0'f32, lstm: lstmOut)
# Reparameterization: z = mu + std * eps, action = tanh(z)
let std = logStd.map(proc(v: float32): float32 = exp(v))
var actions = newTensor[float32](4)
var logProb = 0.0'f32
let twoPiLog = 0.5'f32 * ln(2.0'f32 * PI.float32)
for i in 0 ..< 4:
let eps = gauss(0.0'f64, 1.0'f64).float32
let z = mu[i] + std[i] * eps
actions[i] = tanh(z)
# log N(z | mu, std) - log(1 - tanh²(z) + eps)
let diff = (z - mu[i]) / std[i]
let logNorm = -0.5'f32 * diff * diff - ln(std[i]) - twoPiLog
let tanhCorr = ln(1.0'f32 - actions[i] * actions[i] + LOG_PROB_EPS)
logProb += logNorm - tanhCorr
result = (actions: actions, logProb: logProb, lstm: lstmOut)
# ── Critic forward ────────────────────────────────────────────────────────────
proc criticForward*(net: CriticNet; stateAction: Tensor[float32]; lstm: LSTMState):
tuple[q: float32; lstm: LSTMState] =
## stateAction: [stateDim + actionDim], lstm: (h,c) each [hiddenDim]
let h1 = relu(net.fc1.linear(stateAction))
let lstmOut = lstmStep(net.lstm, h1, lstm.h, lstm.c)
let h2 = relu(net.fc2.linear(lstmOut.h))
let q = net.fc3.linear(h2)
result = (q: q[0], lstm: lstmOut)
@@ -0,0 +1,136 @@
## replay_buffer.nim — sequential ring buffer for off-policy SAC+LSTM training.
##
## Stores transitions and samples contiguous sequences for recurrent training.
## Sequences NEVER cross battle boundaries (done=true).
##
## Config env vars:
## SACLSTM_BUFFER_CAPACITY (default: 500_000)
## SACLSTM_BURN_IN (default: 8)
## SACLSTM_TRAIN_WINDOW (default: 16)
import arraymancer
import std/[os, strutils, random]
# ── Config ────────────────────────────────────────────────────────────────────
proc getBufferCapacity*(): int =
parseInt(getEnv("SACLSTM_BUFFER_CAPACITY", "500000"))
proc getBurnIn*(): int =
parseInt(getEnv("SACLSTM_BURN_IN", "8"))
proc getTrainWindow*(): int =
parseInt(getEnv("SACLSTM_TRAIN_WINDOW", "16"))
# ── Types ─────────────────────────────────────────────────────────────────────
type
Transition* = object
state*: Tensor[float32] # [stateDim]
action*: Tensor[float32] # [actionDim]
reward*: float32
nextState*: Tensor[float32] # [stateDim]
done*: bool # true = battle end
Sequence* = object
burnIn*: seq[Transition] # first burnIn steps (for LSTM warm-up)
train*: seq[Transition] # next trainWindow steps (for gradient computation)
ReplayBuffer* = object
## Ring buffer. `head` is the next write position. `count` tracks fill level.
transitions: seq[Transition]
capacity: int
stateDim: int
actionDim: int
head: int # next write index
count: int # number of valid transitions stored
burnIn: int
trainWindow: int
# ── Construction ──────────────────────────────────────────────────────────────
proc newReplayBuffer*(capacity, stateDim, actionDim: int;
burnIn = getBurnIn();
trainWindow = getTrainWindow()): ReplayBuffer =
result.capacity = capacity
result.stateDim = stateDim
result.actionDim = actionDim
result.burnIn = burnIn
result.trainWindow = trainWindow
result.head = 0
result.count = 0
result.transitions = newSeq[Transition](capacity)
# ── Core operations ───────────────────────────────────────────────────────────
proc add*(buf: var ReplayBuffer; t: Transition) =
buf.transitions[buf.head] = t
buf.head = (buf.head + 1) mod buf.capacity
if buf.count < buf.capacity:
inc buf.count
proc len*(buf: ReplayBuffer): int = buf.count
proc clear*(buf: var ReplayBuffer) =
## Drop all transitions (#48: opponent changed across battles).
## Old slots keep stale tensors until the ring overwrites them.
buf.head = 0
buf.count = 0
proc canSample*(buf: ReplayBuffer): bool =
buf.count >= buf.burnIn + buf.trainWindow
# ── Sampling ──────────────────────────────────────────────────────────────────
proc sampleSequences*(buf: ReplayBuffer; batchSize: int): seq[Sequence] =
## Sample `batchSize` contiguous sequences of length burnIn+trainWindow.
## Sequences never cross a done=true boundary and never wrap the ring buffer.
##
## Returns fewer than batchSize sequences if not enough valid starts exist.
## Returns empty seq if canSample is false.
if not buf.canSample: return @[]
let seqLen = buf.burnIn + buf.trainWindow
let oldest = if buf.count < buf.capacity: 0
else: buf.head # oldest valid index when full
# Build valid starting indices.
# ponytail: O(count) scan per sample call; upgrade to an indexed set of
# boundary positions if count reaches hundreds of thousands and profiling shows
# this is a bottleneck.
var validStarts: seq[int]
for i in 0 ..< buf.count - seqLen + 1:
# Absolute ring-buffer index for the i-th oldest transition
let startIdx = (oldest + i) mod buf.capacity
# Check: the sequence [startIdx .. startIdx+seqLen-2] must not contain done=true
# (a done at position k means the battle ended there; the next transition is
# from a new battle, so the sequence would cross a boundary).
# Also, the sequence must not wrap around the ring buffer.
let endIdx = startIdx + seqLen - 1 # exclusive of wrap check
if endIdx >= buf.capacity:
# Sequence wraps the ring buffer — invalid starting point.
continue
var crosses = false
for j in 0 ..< seqLen - 1:
if buf.transitions[startIdx + j].done:
crosses = true
break
if not crosses:
validStarts.add(startIdx)
if validStarts.len == 0: return @[]
result = newSeq[Sequence](min(batchSize, validStarts.len))
# Sample with replacement if batchSize > validStarts.len, else sample without.
# ponytail: sampling with replacement for simplicity; shuffle+take for
# without-replacement if the caller needs it.
for i in 0 ..< result.len:
let startIdx = validStarts[rand(validStarts.len - 1)]
var s: Sequence
s.burnIn = newSeq[Transition](buf.burnIn)
s.train = newSeq[Transition](buf.trainWindow)
for j in 0 ..< buf.burnIn:
s.burnIn[j] = buf.transitions[startIdx + j]
for j in 0 ..< buf.trainWindow:
s.train[j] = buf.transitions[startIdx + buf.burnIn + j]
result[i] = s
+40 -5
View File
@@ -3,6 +3,25 @@
import std/math import std/math
# ── Lever-2 shaping constants (#59, campaign v2) — TUNABLE ────────────────────
# Scale discipline: commensurate with existing magnitudes (dealt p=1 was +4,
# wall tick -5/tick, win +20). Death/loss and win terms stay dominant; these
# only re-rank mid-band behaviors (fight vs outlive vs get-rammed).
const
# Multiplier on the bullet-damage-dealt term: p=1 hit +4 -> +5. Low-power
# spam stays unprofitable (6*0.1-2 = -1.4 < 0 even after x1.25).
AggressionMult* = 1.25 # ponytail: TUNABLE — raise toward 1.5 if v2 bot still passivity-leaning
# Flat per landed shot on top of damage: discrete accuracy signal.
HitBonus* = 0.5 # ponytail: TUNABLE — keep < 6p-2 at min viable power (~0.34)
# Per bot-bot collision (BotHitBotEvent): server deals RAM_DAMAGE=0.6 to
# both parties but only notifies the hitter — each receipt = damage taken.
RamTakenPenalty* = 3.0 # ponytail: TUNABLE — vs p=0.8 bullet received (-6.8)
# Enemy-charging deterrent: distance/diagonal below this => escalating
# negative (max at zero distance), suppressed while we deal damage that step.
ChargeDistFrac* = 0.12 # ponytail: TUNABLE — ~120u of 800x600 diag (1000)
ChargePenalty* = 2.0 # ponytail: TUNABLE — per-tick ceiling, milder than wall (-5/tick)
# ── Raw reward ──────────────────────────────────────────────────────────────── # ── Raw reward ────────────────────────────────────────────────────────────────
proc computeReward*( proc computeReward*(
@@ -10,6 +29,9 @@ proc computeReward*(
damageReceived: float64 = 0.0, # fire power p_e of enemy shot that hit damageReceived: float64 = 0.0, # fire power p_e of enemy shot that hit
wallHitTicks: int = 0, # ticks in wall contact this step wallHitTicks: int = 0, # ticks in wall contact this step
wastedShotPower: float64 = 0.0, # fire power of shot that missed/hit wall wastedShotPower: float64 = 0.0, # fire power of shot that missed/hit wall
hitCount: int = 0, # own bullets that hit the enemy this step (#59)
ramTakenCount: int = 0, # collisions where we were the victim (#59)
enemyDistFrac: float64 = 2.0, # enemy dist / arena diag; >ChargeDistFrac when no contact (#59)
win: bool = false, win: bool = false,
loss: bool = false loss: bool = false
): float64 = ): float64 =
@@ -17,8 +39,13 @@ proc computeReward*(
## Damage formula: 4p + 2(p-1) = 6p - 2 (matches Tank Royale bullet rules). ## Damage formula: 4p + 2(p-1) = 6p - 2 (matches Tank Royale bullet rules).
let p = damageInflicted let p = damageInflicted
let pe = damageReceived let pe = damageReceived
if p > 0.0: result += 6.0 * p - 2.0 if p > 0.0:
result += AggressionMult * (6.0 * p - 2.0)
if hitCount > 0: result += HitBonus * hitCount.float64
if pe > 0.0: result -= 6.0 * pe - 2.0 if pe > 0.0: result -= 6.0 * pe - 2.0
result -= RamTakenPenalty * ramTakenCount.float64
if enemyDistFrac < ChargeDistFrac and p <= 0.0:
result -= ChargePenalty * (1.0 - enemyDistFrac / ChargeDistFrac)
result -= 5.0 * wallHitTicks.float64 result -= 5.0 * wallHitTicks.float64
if wastedShotPower > 0.0: result -= 0.1 * wastedShotPower if wastedShotPower > 0.0: result -= 0.1 * wastedShotPower
if win: result += 20.0 if win: result += 20.0
@@ -42,8 +69,16 @@ proc update*(rn: var RewardNormalizer; r: float64) =
rn.m2 += delta * delta2 rn.m2 += delta * delta2
proc normalize*(rn: RewardNormalizer; r: float64): float64 = proc normalize*(rn: RewardNormalizer; r: float64): float64 =
## Returns (r - mean) / (std + eps). ## Returns (r - mean) / (std + eps) once statistics are meaningful
## Cold start (n < 2): returns 0.0 to avoid NaN/inf. ## (n >= 4 and spread well above zero). Before that, returns the RAW
if rn.n < 2: return 0.0 ## reward unchanged — Welford M2 collapses to exactly 0 when early raw
## rewards are identical, and dividing by the 1e-8 floor then z-scores
## the first differing reward to ~1e8, poisoning TD targets.
if rn.n < 4: return r
let variance = rn.m2 / rn.n.float64 # ponytail: population var; switch to n-1 if bias matters let variance = rn.m2 / rn.n.float64 # ponytail: population var; switch to n-1 if bias matters
result = (r - rn.mean) / (sqrt(variance) + NormEps) let stddev = sqrt(variance)
if stddev <= 1e-3 * (abs(rn.mean) + 1.0): return r
# ponytail: warm-up pass-through ceiling — raw rewards bypass normalization
# until stats are meaningful; upgrade = persist Welford state in checkpoint
# if warm-up noise ever hurts learning.
result = (r - rn.mean) / (stddev + NormEps)
+611
View File
@@ -0,0 +1,611 @@
## training.nim — SAC-v2 update for the LSTM Actor + twin Critic.
## Manual backprop; no autograd. Uses Arraymancer tensors throughout.
##
## Config env vars:
## SACLSTM_LR_ACTOR (default: 3e-4)
## SACLSTM_LR_CRITIC (default: 3e-4)
## SACLSTM_LR_ALPHA (default: 3e-4)
## SACLSTM_GAMMA (default: 0.99)
## SACLSTM_TAU (default: 0.005)
## SACLSTM_TARGET_ENTROPY (default: -4.0)
import arraymancer except Linear
import std/[math, os, strutils]
import SAC_LSTM_Bot/network
import SAC_LSTM_Bot/weights
import SAC_LSTM_Bot/replay_buffer
# ── Config ────────────────────────────────────────────────────────────────────
proc getLrActor*(): float32 = parseFloat(getEnv("SACLSTM_LR_ACTOR", "3e-4")).float32
proc getLrCritic*(): float32 = parseFloat(getEnv("SACLSTM_LR_CRITIC", "3e-4")).float32
proc getLrAlpha*(): float32 = parseFloat(getEnv("SACLSTM_LR_ALPHA", "3e-4")).float32
proc getGamma*(): float32 = parseFloat(getEnv("SACLSTM_GAMMA", "0.99")).float32
proc getTau*(): float32 = parseFloat(getEnv("SACLSTM_TAU", "0.005")).float32
proc getTargetEntropy*(): float32 =
parseFloat(getEnv("SACLSTM_TARGET_ENTROPY", "-4.0")).float32
# ── SACTrainer ────────────────────────────────────────────────────────────────
type
SACTrainer* = object
actor*: ActorNet
critic1*: CriticNet
critic2*: CriticNet
targetCritic1*: CriticNet
targetCritic2*: CriticNet
logAlpha*: float32 ## log of entropy temperature; alpha = exp(logAlpha)
targetEntropy*: float32
tau*: float32
lrActor*: float32
lrCritic*: float32
lrAlpha*: float32
gamma*: float32
adam*: SACAdamStates
SACMetrics* = object
criticLoss*: float32
actorLoss*: float32
alphaLoss*: float32
alpha*: float32
proc initSACTrainer*(stateDim, actionDim: int): SACTrainer =
result.actor = initActorNet(stateDim)
result.critic1 = initCriticNet(stateDim, actionDim)
result.critic2 = initCriticNet(stateDim, actionDim)
result.targetCritic1 = result.critic1
result.targetCritic2 = result.critic2
result.logAlpha = 0.0'f32
result.targetEntropy = getTargetEntropy()
result.tau = getTau()
result.lrActor = getLrActor()
result.lrCritic = getLrCritic()
result.lrAlpha = getLrAlpha()
result.gamma = getGamma()
result.adam = initSACAdamStates(result.actor, result.critic1, result.critic2)
proc alpha*(t: SACTrainer): float32 = exp(t.logAlpha)
# ── Adam steps ────────────────────────────────────────────────────────────────
proc adamStepScalar(param: var float32; grad: float32;
state: var AdamVar; lr: float32) =
## Scalar Adam for logAlpha (state.m/v are shape [1] tensors).
inc state.t
let b1 = 0.9'f32; let b2 = 0.999'f32; let eps = 1e-8'f32
state.m[0] = b1 * state.m[0] + (1.0'f32 - b1) * grad
state.v[0] = b2 * state.v[0] + (1.0'f32 - b2) * grad * grad
let mHat = state.m[0] / (1.0'f32 - b1 ^ state.t.float32)
let vHat = state.v[0] / (1.0'f32 - b2 ^ state.t.float32)
param -= lr * mHat / (sqrt(vHat) + eps)
proc adamStepTensor(param: var Tensor[float32]; grad: Tensor[float32];
state: var AdamVar; lr: float32) =
## Tensor Adam (same pattern as PPO_Bot/training.nim adamStep).
inc state.t
let b1 = 0.9'f32; let b2 = 0.999'f32; let eps = 1e-8'f32
state.m = b1 *. state.m + (1.0'f32 - b1) *. grad
state.v = b2 *. state.v + (1.0'f32 - b2) *. (grad *. grad)
let mHat = state.m /. (1.0'f32 - b1 ^ state.t.float32)
let vHat = state.v /. (1.0'f32 - b2 ^ state.t.float32)
param -= lr *. mHat /. vHat.map(proc(x: float32): float32 = sqrt(x) + eps)
# ── Gradient clipping ─────────────────────────────────────────────────────────
proc globalNorm(grads: seq[Tensor[float32]]): float32 =
var sumSq = 0.0'f32
for g in grads:
for v in g: sumSq += v * v
sqrt(sumSq)
proc clipGrads(grads: var seq[Tensor[float32]]; maxNorm: float32) =
let norm = globalNorm(grads)
if norm > maxNorm and norm == norm:
let scale = maxNorm / norm
for g in grads.mitems: g = g *. scale
# ── Forward caches (for backprop) ─────────────────────────────────────────────
type
LinearFwd = object
inp, pre, act: Tensor[float32] # input, pre-relu, post-relu (or linear)
proc linearReluFwd(l: Linear; x: Tensor[float32]): LinearFwd =
result.inp = x
result.pre = l.w * x + l.b
result.act = relu(result.pre)
proc linearFwd(l: Linear; x: Tensor[float32]): LinearFwd =
result.inp = x
result.pre = l.w * x + l.b
result.act = result.pre # no nonlinearity
type
LSTMFwdCache = object
xh, gatesPre: Tensor[float32] # [inputDim+hd], [4*hd]
iGate, fGate, gGate, oGate: Tensor[float32] # [hd] each
cPrev, cPrime, hPrime: Tensor[float32] # [hd] each
proc lstmStepCached(cell: LSTMCell; x, h, c: Tensor[float32]): LSTMFwdCache =
result.cPrev = c
result.xh = concat(x, h, axis = 0)
result.gatesPre = cell.wCombined * result.xh + cell.bCombined
let hd = cell.hiddenDim
result.iGate = sigmoid(result.gatesPre[0 ..< hd])
result.fGate = sigmoid(result.gatesPre[hd ..< 2*hd])
result.gGate = tanhT(result.gatesPre[2*hd ..< 3*hd])
result.oGate = sigmoid(result.gatesPre[3*hd ..< 4*hd])
result.cPrime = result.fGate *. c + result.iGate *. result.gGate
result.hPrime = result.oGate *. tanhT(result.cPrime)
# ── Backward helpers ──────────────────────────────────────────────────────────
proc reluGrad(pre, dAct: Tensor[float32]): Tensor[float32] =
result = newTensor[float32](dAct.shape)
for i in 0 ..< dAct.shape[0]:
result[i] = if pre[i] > 0.0'f32: dAct[i] else: 0.0'f32
## Linear layer backward: returns (dx, dw, db) given upstream grad dAct.
## If hasRelu, applies relu' gate before computing gradients.
proc linearBack(w: Tensor[float32]; fwd: LinearFwd;
dAct: Tensor[float32]; hasRelu: bool):
tuple[dx, dw, db: Tensor[float32]] =
let dPre = if hasRelu: reluGrad(fwd.pre, dAct) else: dAct
result.dw = dPre.unsqueeze(1) * fwd.inp.unsqueeze(0) # [out, in]
result.db = dPre
result.dx = w.transpose * dPre # [in]
## LSTM single-step backward. dHPrime: [hd], dCPrime: [hd] (use zeros for truncated BPTT).
## Returns (dwCombined, dbCombined, dxh).
proc lstmBack(cell: LSTMCell; cache: LSTMFwdCache;
dHPrime, dCPrime: Tensor[float32]):
tuple[dwCombined, dbCombined, dxh: Tensor[float32]] =
let hd = cell.hiddenDim
let tanhCPrime = tanhT(cache.cPrime)
# Output gate
let dOGate_post = dHPrime *. tanhCPrime
# Cell state: gradient from h' and from downstream dCPrime
let dCPrimeTotal = dHPrime *. cache.oGate *.
(ones[float32](hd) - tanhCPrime *. tanhCPrime) + dCPrime
# Gate post-activation gradients
let dFGate_post = dCPrimeTotal *. cache.cPrev
let dIGate_post = dCPrimeTotal *. cache.gGate
let dGGate_post = dCPrimeTotal *. cache.iGate
# Gate pre-activation gradients (sigmoid', tanh')
let dIPre = dIGate_post *. cache.iGate *. (ones[float32](hd) - cache.iGate)
let dFPre = dFGate_post *. cache.fGate *. (ones[float32](hd) - cache.fGate)
let dGPre = dGGate_post *. (ones[float32](hd) - cache.gGate *. cache.gGate)
let dOPre = dOGate_post *. cache.oGate *. (ones[float32](hd) - cache.oGate)
# Concatenated gate gradient [4*hd]
let dGatesPre = concat(dIPre, dFPre, dGPre, dOPre, axis = 0)
result.dwCombined = dGatesPre.unsqueeze(1) * cache.xh.unsqueeze(0)
result.dbCombined = dGatesPre
result.dxh = cell.wCombined.transpose * dGatesPre
# ── Squashed-Gaussian log-prob and its gradients ──────────────────────────────
const
LOG_PROB_EPS = 1e-6'f32
LOG_STD_MIN = -5.0'f32
LOG_STD_MAX = 2.0'f32
## Given stored mu, clamped logStd, and sampled action = tanh(z), recover
## log π(a|s) and gradients w.r.t. mu and logStd.
proc squashedLogProb(mu, logStd, action: Tensor[float32]):
tuple[logProb: float32;
dLogProbDMu, dLogProbDLogStd: Tensor[float32]] =
let std = logStd.map(proc(v: float32): float32 = exp(v))
let z = mu # deterministic reparam: action = tanh(mu), so z ≡ mu, diff ≡ 0
result.dLogProbDMu = newTensor[float32](4)
result.dLogProbDLogStd = newTensor[float32](4)
let twoPiLog = 0.5'f32 * ln(2.0'f32 * PI.float32)
var lp = 0.0'f32
for i in 0 ..< 4:
let diff = (z[i] - mu[i]) / std[i]
let logNorm = -0.5'f32 * diff * diff - ln(std[i]) - twoPiLog
let tanhCorr = ln(1.0'f32 - action[i] * action[i] + LOG_PROB_EPS)
lp += logNorm - tanhCorr
result.dLogProbDMu[i] = diff / std[i] # (z-mu)/std²
result.dLogProbDLogStd[i] = diff * diff - 1.0'f32 # d logN / d logStd
result.logProb = lp
# ── Critic forward with activation cache ──────────────────────────────────────
type
CriticFwdCache = object
fc1: LinearFwd
lstm: LSTMFwdCache
fc2: LinearFwd
fc3: LinearFwd
q: float32
proc criticFwdCached(net: CriticNet; stateAction: Tensor[float32];
h, c: Tensor[float32]): CriticFwdCache =
result.fc1 = linearReluFwd(net.fc1, stateAction)
result.lstm = lstmStepCached(net.lstm, result.fc1.act, h, c)
result.fc2 = linearReluFwd(net.fc2, result.lstm.hPrime)
result.fc3 = linearFwd(net.fc3, result.fc2.act)
result.q = result.fc3.act[0]
# ── Critic backward ───────────────────────────────────────────────────────────
type
CriticGrads = object
dFc1W, dFc1B: Tensor[float32]
dFc2W, dFc2B: Tensor[float32]
dFc3W, dFc3B: Tensor[float32]
dLstmW, dLstmB: Tensor[float32]
dInput: Tensor[float32] ## grad w.r.t. stateAction input
proc criticBack(net: CriticNet; cache: CriticFwdCache; dQ: float32): CriticGrads =
let dFc3Act = [dQ].toTensor()
let fc3b = linearBack(net.fc3.w, cache.fc3, dFc3Act, hasRelu = false)
result.dFc3W = fc3b.dw; result.dFc3B = fc3b.db
let fc2b = linearBack(net.fc2.w, cache.fc2, fc3b.dx, hasRelu = true)
result.dFc2W = fc2b.dw; result.dFc2B = fc2b.db
let zeros_hd = zeros[float32](net.hiddenDim)
let lstmb = lstmBack(net.lstm, cache.lstm, fc2b.dx, zeros_hd)
result.dLstmW = lstmb.dwCombined; result.dLstmB = lstmb.dbCombined
# xh = [fc1.act | h_prev], dx is the x-part (fc1 output dim = hiddenDim)
let dLstmX = lstmb.dxh[0 ..< cache.fc1.act.shape[0]]
let fc1b = linearBack(net.fc1.w, cache.fc1, dLstmX, hasRelu = true)
result.dFc1W = fc1b.dw; result.dFc1B = fc1b.db
result.dInput = fc1b.dx # [stateDim + actionDim]
# ── Actor forward with activation cache ───────────────────────────────────────
type
ActorFwdCache = object
fc1: LinearFwd
lstm: LSTMFwdCache
fc2: LinearFwd
muHead: LinearFwd
lsHead: LinearFwd ## logStd head
mu: Tensor[float32] ## [4]
logStd: Tensor[float32] ## [4] clamped
action: Tensor[float32] ## [4] tanh(mu) — deterministic for gradient
proc actorFwdCached(net: ActorNet; state: Tensor[float32];
h, c: Tensor[float32]): ActorFwdCache =
result.fc1 = linearReluFwd(net.fc1, state)
result.lstm = lstmStepCached(net.lstm, result.fc1.act, h, c)
result.fc2 = linearReluFwd(net.fc2, result.lstm.hPrime)
result.muHead = linearFwd(net.muHead, result.fc2.act)
result.lsHead = linearFwd(net.logStdHead, result.fc2.act)
result.mu = result.muHead.act
result.logStd = result.lsHead.act.map(
proc(v: float32): float32 = clamp(v, LOG_STD_MIN, LOG_STD_MAX))
# Use tanh(mu) as the action for gradient computation (reparameterization).
# ponytail: deterministic here; add stochastic sample if off-policy bias matters.
result.action = tanhT(result.mu)
# ── Actor backward ────────────────────────────────────────────────────────────
type
ActorGrads = object
dFc1W, dFc1B: Tensor[float32]
dFc2W, dFc2B: Tensor[float32]
dMuW, dMuB: Tensor[float32]
dLogStdW, dLogStdB: Tensor[float32]
dLstmW, dLstmB: Tensor[float32]
proc actorBack(net: ActorNet; cache: ActorFwdCache;
dMu, dLogStd: Tensor[float32]): ActorGrads =
let muBack = linearBack(net.muHead.w, cache.muHead, dMu, hasRelu = false)
result.dMuW = muBack.dw; result.dMuB = muBack.db
let lsBack = linearBack(net.logStdHead.w, cache.lsHead, dLogStd, hasRelu = false)
result.dLogStdW = lsBack.dw; result.dLogStdB = lsBack.db
# fc2 gets grads from both output heads
let fc2b = linearBack(net.fc2.w, cache.fc2, muBack.dx + lsBack.dx, hasRelu = true)
result.dFc2W = fc2b.dw; result.dFc2B = fc2b.db
let zeros_hd = zeros[float32](net.hiddenDim)
let lstmb = lstmBack(net.lstm, cache.lstm, fc2b.dx, zeros_hd)
result.dLstmW = lstmb.dwCombined; result.dLstmB = lstmb.dbCombined
let dLstmX = lstmb.dxh[0 ..< cache.fc1.act.shape[0]]
let fc1b = linearBack(net.fc1.w, cache.fc1, dLstmX, hasRelu = true)
result.dFc1W = fc1b.dw; result.dFc1B = fc1b.db
# ── Adam application ──────────────────────────────────────────────────────────
proc applyActorAdam(net: var ActorNet; g: ActorGrads;
adam: var ActorAdam; lr: float32) =
adamStepTensor(net.fc1.w, g.dFc1W, adam.fc1.w, lr)
adamStepTensor(net.fc1.b, g.dFc1B, adam.fc1.b, lr)
adamStepTensor(net.fc2.w, g.dFc2W, adam.fc2.w, lr)
adamStepTensor(net.fc2.b, g.dFc2B, adam.fc2.b, lr)
adamStepTensor(net.muHead.w, g.dMuW, adam.muHead.w, lr)
adamStepTensor(net.muHead.b, g.dMuB, adam.muHead.b, lr)
adamStepTensor(net.logStdHead.w, g.dLogStdW, adam.logStdHead.w, lr)
adamStepTensor(net.logStdHead.b, g.dLogStdB, adam.logStdHead.b, lr)
adamStepTensor(net.lstm.wCombined, g.dLstmW, adam.lstm.wCombined, lr)
adamStepTensor(net.lstm.bCombined, g.dLstmB, adam.lstm.bCombined, lr)
proc applyCriticAdam(net: var CriticNet; g: CriticGrads;
adam: var CriticAdam; lr: float32) =
adamStepTensor(net.fc1.w, g.dFc1W, adam.fc1.w, lr)
adamStepTensor(net.fc1.b, g.dFc1B, adam.fc1.b, lr)
adamStepTensor(net.fc2.w, g.dFc2W, adam.fc2.w, lr)
adamStepTensor(net.fc2.b, g.dFc2B, adam.fc2.b, lr)
adamStepTensor(net.fc3.w, g.dFc3W, adam.fc3.w, lr)
adamStepTensor(net.fc3.b, g.dFc3B, adam.fc3.b, lr)
adamStepTensor(net.lstm.wCombined, g.dLstmW, adam.lstm.wCombined, lr)
adamStepTensor(net.lstm.bCombined, g.dLstmB, adam.lstm.bCombined, lr)
# ── Soft target update ────────────────────────────────────────────────────────
proc softUpdateLinear(target: var Linear; src: Linear; tau: float32) =
target.w = tau *. src.w + (1.0'f32 - tau) *. target.w
target.b = tau *. src.b + (1.0'f32 - tau) *. target.b
proc softUpdateLSTM(target: var LSTMCell; src: LSTMCell; tau: float32) =
target.wCombined = tau *. src.wCombined + (1.0'f32 - tau) *. target.wCombined
target.bCombined = tau *. src.bCombined + (1.0'f32 - tau) *. target.bCombined
proc softUpdateCritic(target: var CriticNet; src: CriticNet; tau: float32) =
softUpdateLinear(target.fc1, src.fc1, tau)
softUpdateLSTM(target.lstm, src.lstm, tau)
softUpdateLinear(target.fc2, src.fc2, tau)
softUpdateLinear(target.fc3, src.fc3, tau)
# ── Gradient accumulators ─────────────────────────────────────────────────────
proc zeroCriticGrads(net: CriticNet): CriticGrads =
result.dFc1W = zeros[float32](net.fc1.w.shape)
result.dFc1B = zeros[float32](net.fc1.b.shape)
result.dFc2W = zeros[float32](net.fc2.w.shape)
result.dFc2B = zeros[float32](net.fc2.b.shape)
result.dFc3W = zeros[float32](net.fc3.w.shape)
result.dFc3B = zeros[float32](net.fc3.b.shape)
result.dLstmW = zeros[float32](net.lstm.wCombined.shape)
result.dLstmB = zeros[float32](net.lstm.bCombined.shape)
result.dInput = zeros[float32](net.fc1.w.shape[1]) # [stateDim+actionDim]
proc zeroActorGrads(net: ActorNet): ActorGrads =
result.dFc1W = zeros[float32](net.fc1.w.shape)
result.dFc1B = zeros[float32](net.fc1.b.shape)
result.dFc2W = zeros[float32](net.fc2.w.shape)
result.dFc2B = zeros[float32](net.fc2.b.shape)
result.dMuW = zeros[float32](net.muHead.w.shape)
result.dMuB = zeros[float32](net.muHead.b.shape)
result.dLogStdW = zeros[float32](net.logStdHead.w.shape)
result.dLogStdB = zeros[float32](net.logStdHead.b.shape)
result.dLstmW = zeros[float32](net.lstm.wCombined.shape)
result.dLstmB = zeros[float32](net.lstm.bCombined.shape)
proc addCriticGrads(a: var CriticGrads; b: CriticGrads) =
a.dFc1W += b.dFc1W; a.dFc1B += b.dFc1B
a.dFc2W += b.dFc2W; a.dFc2B += b.dFc2B
a.dFc3W += b.dFc3W; a.dFc3B += b.dFc3B
a.dLstmW += b.dLstmW; a.dLstmB += b.dLstmB
# dInput not accumulated (not used for parameter update)
proc addActorGrads(a: var ActorGrads; b: ActorGrads) =
a.dFc1W += b.dFc1W; a.dFc1B += b.dFc1B
a.dFc2W += b.dFc2W; a.dFc2B += b.dFc2B
a.dMuW += b.dMuW; a.dMuB += b.dMuB
a.dLogStdW += b.dLogStdW; a.dLogStdB += b.dLogStdB
a.dLstmW += b.dLstmW; a.dLstmB += b.dLstmB
proc scaleCriticGrads(g: var CriticGrads; s: float32) =
g.dFc1W = g.dFc1W *. s; g.dFc1B = g.dFc1B *. s
g.dFc2W = g.dFc2W *. s; g.dFc2B = g.dFc2B *. s
g.dFc3W = g.dFc3W *. s; g.dFc3B = g.dFc3B *. s
g.dLstmW = g.dLstmW *. s; g.dLstmB = g.dLstmB *. s
proc scaleActorGrads(g: var ActorGrads; s: float32) =
g.dFc1W = g.dFc1W *. s; g.dFc1B = g.dFc1B *. s
g.dFc2W = g.dFc2W *. s; g.dFc2B = g.dFc2B *. s
g.dMuW = g.dMuW *. s; g.dMuB = g.dMuB *. s
g.dLogStdW = g.dLogStdW *. s; g.dLogStdB = g.dLogStdB *. s
g.dLstmW = g.dLstmW *. s; g.dLstmB = g.dLstmB *. s
proc criticGradsAsSeq(g: CriticGrads): seq[Tensor[float32]] =
@[g.dFc1W, g.dFc1B, g.dFc2W, g.dFc2B, g.dFc3W, g.dFc3B, g.dLstmW, g.dLstmB]
proc applyClipToCritic(g: var CriticGrads; maxNorm: float32) =
var gs = criticGradsAsSeq(g)
clipGrads(gs, maxNorm)
g.dFc1W = gs[0]; g.dFc1B = gs[1]
g.dFc2W = gs[2]; g.dFc2B = gs[3]
g.dFc3W = gs[4]; g.dFc3B = gs[5]
g.dLstmW = gs[6]; g.dLstmB = gs[7]
proc applyClipToActor(g: var ActorGrads; maxNorm: float32) =
var gs = @[g.dFc1W, g.dFc1B, g.dFc2W, g.dFc2B,
g.dMuW, g.dMuB, g.dLogStdW, g.dLogStdB, g.dLstmW, g.dLstmB]
clipGrads(gs, maxNorm)
g.dFc1W = gs[0]; g.dFc1B = gs[1]
g.dFc2W = gs[2]; g.dFc2B = gs[3]
g.dMuW = gs[4]; g.dMuB = gs[5]
g.dLogStdW = gs[6]; g.dLogStdB = gs[7]
g.dLstmW = gs[8]; g.dLstmB = gs[9]
# ── SAC update ────────────────────────────────────────────────────────────────
proc sacUpdate*(trainer: var SACTrainer; sequences: seq[Sequence]): SACMetrics =
## One SAC-v2 update given a batch of sequences. No-op if empty.
if sequences.len == 0: return
let N = sequences.len.float32
let alph = trainer.alpha()
let gamma = trainer.gamma
var totalCriticLoss = 0.0'f32
var totalActorLoss = 0.0'f32
var totalAlphaLoss = 0.0'f32
var accC1Grads = zeroCriticGrads(trainer.critic1)
var accC2Grads = zeroCriticGrads(trainer.critic2)
var accAGrads = zeroActorGrads(trainer.actor)
var dLogAlpha = 0.0'f32
for sq in sequences:
# ── 1. Burn-in: warm up hidden states, no gradient ──────────────────────
var actorH = zeros[float32](trainer.actor.hiddenDim)
var actorC = zeros[float32](trainer.actor.hiddenDim)
var c1H = zeros[float32](trainer.critic1.hiddenDim)
var c1C = zeros[float32](trainer.critic1.hiddenDim)
var c2H = zeros[float32](trainer.critic2.hiddenDim)
var c2C = zeros[float32](trainer.critic2.hiddenDim)
var tc1H = zeros[float32](trainer.targetCritic1.hiddenDim)
var tc1C = zeros[float32](trainer.targetCritic1.hiddenDim)
var tc2H = zeros[float32](trainer.targetCritic2.hiddenDim)
var tc2C = zeros[float32](trainer.targetCritic2.hiddenDim)
for tr in sq.burnIn:
let sa = concat(tr.state, tr.action, axis = 0)
let af = lstmStepCached(trainer.actor.lstm,
relu(trainer.actor.fc1.linear(tr.state)), actorH, actorC)
actorH = af.hPrime; actorC = af.cPrime
let c1f = lstmStepCached(trainer.critic1.lstm,
relu(trainer.critic1.fc1.linear(sa)), c1H, c1C)
c1H = c1f.hPrime; c1C = c1f.cPrime
let c2f = lstmStepCached(trainer.critic2.lstm,
relu(trainer.critic2.fc1.linear(sa)), c2H, c2C)
c2H = c2f.hPrime; c2C = c2f.cPrime
let tc1f = lstmStepCached(trainer.targetCritic1.lstm,
relu(trainer.targetCritic1.fc1.linear(sa)), tc1H, tc1C)
tc1H = tc1f.hPrime; tc1C = tc1f.cPrime
let tc2f = lstmStepCached(trainer.targetCritic2.lstm,
relu(trainer.targetCritic2.fc1.linear(sa)), tc2H, tc2C)
tc2H = tc2f.hPrime; tc2C = tc2f.cPrime
# ── 2–4. Training window ─────────────────────────────────────────────────
let T = sq.train.len.float32
var seqC1Grads = zeroCriticGrads(trainer.critic1)
var seqC2Grads = zeroCriticGrads(trainer.critic2)
var seqAGrads = zeroActorGrads(trainer.actor)
var seqDLogAlpha = 0.0'f32
for tr in sq.train:
let s = tr.state
let a = tr.action
let r = tr.reward
let sn = tr.nextState
let d = if tr.done: 0.0'f32 else: 1.0'f32
let sa = concat(s, a, axis = 0)
# ── 2. Critic update ─────────────────────────────────────────────────
let c1Cache = criticFwdCached(trainer.critic1, sa, c1H, c1C)
let c2Cache = criticFwdCached(trainer.critic2, sa, c2H, c2C)
# ── 3. Actor update (run first to get actorFwd on s before advancing h/c) ─
let actorFwd = actorFwdCached(trainer.actor, s, actorH, actorC)
let aCurr = actorFwd.action
let lpResult = squashedLogProb(actorFwd.mu, actorFwd.logStd, aCurr)
let logProbA = lpResult.logProb
# Advance actor hidden state from s → sn before computing actorNxt
actorH = actorFwd.lstm.hPrime; actorC = actorFwd.lstm.cPrime
# Next-state action from current actor (uses h/c advanced through s)
let actorNxt = actorFwdCached(trainer.actor, sn, actorH, actorC)
let aN = actorNxt.action
let lpN = squashedLogProb(actorNxt.mu, actorNxt.logStd, aN).logProb
let saN = concat(sn, aN, axis = 0)
# Target Q
let tc1Cache = criticFwdCached(trainer.targetCritic1, saN, tc1H, tc1C)
let tc2Cache = criticFwdCached(trainer.targetCritic2, saN, tc2H, tc2C)
let minQTarg = min(tc1Cache.q, tc2Cache.q)
# Bellman target
let y = r + gamma * d * (minQTarg - alph * lpN)
let errQ1 = c1Cache.q - y
let errQ2 = c2Cache.q - y
totalCriticLoss += 0.5'f32 * (errQ1 * errQ1 + errQ2 * errQ2)
# MSE gradient: d_loss/d_q = (q - y) [scaling applied at accumulation]
addCriticGrads(seqC1Grads, criticBack(trainer.critic1, c1Cache, errQ1))
addCriticGrads(seqC2Grads, criticBack(trainer.critic2, c2Cache, errQ2))
# Advance critic hidden states
c1H = c1Cache.lstm.hPrime; c1C = c1Cache.lstm.cPrime
c2H = c2Cache.lstm.hPrime; c2C = c2Cache.lstm.cPrime
tc1H = tc1Cache.lstm.hPrime; tc1C = tc1Cache.lstm.cPrime
tc2H = tc2Cache.lstm.hPrime; tc2C = tc2Cache.lstm.cPrime
# ── 3 (cont). Actor gradient via critic ─────────────────────────────────
# Q-values for current policy action (critics used as frozen estimators)
let saCurr = concat(s, aCurr, axis = 0)
let qA1Cache = criticFwdCached(trainer.critic1, saCurr, c1H, c1C)
let qA2Cache = criticFwdCached(trainer.critic2, saCurr, c2H, c2C)
let q1Val = qA1Cache.q
let q2Val = qA2Cache.q
let qA1Back = criticBack(trainer.critic1, qA1Cache, -1.0'f32)
let qA2Back = criticBack(trainer.critic2, qA2Cache, -1.0'f32)
totalActorLoss += alph * logProbA - min(q1Val, q2Val)
# Gradient of -minQ w.r.t. action = dInput[stateDim ..< stateDim+actionDim]
# from the critic whose Q was smaller.
let minQBack = if q1Val <= q2Val: qA1Back else: qA2Back
let stateDim = s.shape[0]
let actionDim = aCurr.shape[0]
let dQdA = minQBack.dInput[stateDim ..< stateDim + actionDim]
# Chain through tanh: d(tanh(mu))/d(mu) = 1 - action²
let dTanh = aCurr.map(proc(a: float32): float32 = 1.0'f32 - a * a)
# Total gradient w.r.t. mu: (alpha * dLogP/dMu + dQ/dA) * dTanh/dMu
let dMu = (alph *. lpResult.dLogProbDMu + dQdA) *. dTanh
let dLogStd = alph *. lpResult.dLogProbDLogStd
addActorGrads(seqAGrads, actorBack(trainer.actor, actorFwd, dMu, dLogStd))
# ── 4. Alpha update ──────────────────────────────────────────────────
# Loss = -log_alpha * stop_grad(logProb + targetEntropy)
# d_loss/d_log_alpha = -(logProb + targetEntropy)
totalAlphaLoss += -trainer.logAlpha * (logProbA + trainer.targetEntropy)
seqDLogAlpha += -(logProbA + trainer.targetEntropy)
# Average sequence grads over T steps, accumulate over batch
scaleCriticGrads(seqC1Grads, 1.0'f32 / T)
scaleCriticGrads(seqC2Grads, 1.0'f32 / T)
scaleActorGrads(seqAGrads, 1.0'f32 / T)
addCriticGrads(accC1Grads, seqC1Grads)
addCriticGrads(accC2Grads, seqC2Grads)
addActorGrads(accAGrads, seqAGrads)
dLogAlpha += seqDLogAlpha / T
# Average over batch
scaleCriticGrads(accC1Grads, 1.0'f32 / N)
scaleCriticGrads(accC2Grads, 1.0'f32 / N)
scaleActorGrads(accAGrads, 1.0'f32 / N)
dLogAlpha /= N
# Gradient clipping (max_norm = 1.0)
applyClipToCritic(accC1Grads, 1.0'f32)
applyClipToCritic(accC2Grads, 1.0'f32)
applyClipToActor(accAGrads, 1.0'f32)
# Apply Adam updates
applyCriticAdam(trainer.critic1, accC1Grads, trainer.adam.critic1, trainer.lrCritic)
applyCriticAdam(trainer.critic2, accC2Grads, trainer.adam.critic2, trainer.lrCritic)
applyActorAdam(trainer.actor, accAGrads, trainer.adam.actor, trainer.lrActor)
adamStepScalar(trainer.logAlpha, dLogAlpha, trainer.adam.alpha, trainer.lrAlpha)
# ── 5. Soft target update ──────────────────────────────────────────────────
softUpdateCritic(trainer.targetCritic1, trainer.critic1, trainer.tau)
softUpdateCritic(trainer.targetCritic2, trainer.critic2, trainer.tau)
let totalSteps = N * sequences[0].train.len.float32
result.criticLoss = totalCriticLoss / totalSteps
result.actorLoss = totalActorLoss / totalSteps
result.alphaLoss = totalAlphaLoss / totalSteps
result.alpha = trainer.alpha()
+352
View File
@@ -0,0 +1,352 @@
## weights.nim — save/load all SAC-LSTM network tensors as .npy inside a .zip.
##
## Strategy: write_npy writes to paths; zip/zipfiles.addFile reads from paths.
## So we write each tensor to a temp .npy, add it to the zip, then delete temps.
## Load reverses: extract each entry to a temp .npy, read_npy, delete.
## Atomic save: build the zip in a temp path, then rename over the target.
import arraymancer except Linear
import zip/zipfiles
import std/[os, times, strutils]
import SAC_LSTM_Bot/network
# ── Adam state types (used by training.nim) ───────────────────────────────────
type
AdamVar* = object
m*, v*: Tensor[float32]
t*: int
## Adam states for one Linear layer (w and b).
LinearAdam* = object
w*, b*: AdamVar
## Adam states for one LSTMCell (wCombined and bCombined).
LSTMCellAdam* = object
wCombined*, bCombined*: AdamVar
## Adam states for one ActorNet.
ActorAdam* = object
fc1*, fc2*, muHead*, logStdHead*: LinearAdam
lstm*: LSTMCellAdam
## Adam states for one CriticNet.
CriticAdam* = object
fc1*, fc2*, fc3*: LinearAdam
lstm*: LSTMCellAdam
SACAdamStates* = object
actor*: ActorAdam
critic1*: CriticAdam
critic2*: CriticAdam
alpha*: AdamVar # scalar, shape [1]
initialized*: bool
# ── Init helpers ──────────────────────────────────────────────────────────────
proc initAdamVar(t: Tensor[float32]): AdamVar =
AdamVar(m: zeros[float32](t.shape), v: zeros[float32](t.shape), t: 0)
proc initLinearAdam*(l: Linear): LinearAdam =
LinearAdam(w: initAdamVar(l.w), b: initAdamVar(l.b))
proc initLSTMCellAdam*(c: LSTMCell): LSTMCellAdam =
LSTMCellAdam(
wCombined: initAdamVar(c.wCombined),
bCombined: initAdamVar(c.bCombined))
proc initActorAdam*(a: ActorNet): ActorAdam =
ActorAdam(
fc1: initLinearAdam(a.fc1),
fc2: initLinearAdam(a.fc2),
muHead: initLinearAdam(a.muHead),
logStdHead: initLinearAdam(a.logStdHead),
lstm: initLSTMCellAdam(a.lstm))
proc initCriticAdam*(c: CriticNet): CriticAdam =
CriticAdam(
fc1: initLinearAdam(c.fc1),
fc2: initLinearAdam(c.fc2),
fc3: initLinearAdam(c.fc3),
lstm: initLSTMCellAdam(c.lstm))
proc initSACAdamStates*(actor: ActorNet; critic1, critic2: CriticNet): SACAdamStates =
result.actor = initActorAdam(actor)
result.critic1 = initCriticAdam(critic1)
result.critic2 = initCriticAdam(critic2)
result.alpha = initAdamVar(ones[float32](1))
result.initialized = true
# ── Internal: temp dir per save ───────────────────────────────────────────────
proc tmpDir(): string =
getTempDir() / ("sacw_" & $int(epochTime() * 1000))
# ── Save helpers ──────────────────────────────────────────────────────────────
template addT(z: var ZipArchive; name: string; t: Tensor[float32]; tmp: string) =
## Write tensor to a temp file, add to zip, delete temp file.
let p = tmp / name
t.write_npy(p)
z.addFile(name, p)
proc addLinear(z: var ZipArchive; prefix: string; l: Linear; tmp: string) =
addT(z, prefix & "_w.npy", l.w, tmp)
addT(z, prefix & "_b.npy", l.b, tmp)
proc addLSTMCell(z: var ZipArchive; prefix: string; c: LSTMCell; tmp: string) =
addT(z, prefix & "_wc.npy", c.wCombined, tmp)
addT(z, prefix & "_bc.npy", c.bCombined, tmp)
proc addActorNet(z: var ZipArchive; prefix: string; a: ActorNet; tmp: string) =
addLinear(z, prefix & "_fc1", a.fc1, tmp)
addLSTMCell(z, prefix & "_lstm", a.lstm, tmp)
addLinear(z, prefix & "_fc2", a.fc2, tmp)
addLinear(z, prefix & "_mu", a.muHead, tmp)
addLinear(z, prefix & "_logstd", a.logStdHead, tmp)
proc addCriticNet(z: var ZipArchive; prefix: string; c: CriticNet; tmp: string) =
addLinear(z, prefix & "_fc1", c.fc1, tmp)
addLSTMCell(z, prefix & "_lstm", c.lstm, tmp)
addLinear(z, prefix & "_fc2", c.fc2, tmp)
addLinear(z, prefix & "_fc3", c.fc3, tmp)
proc addAdamVar(z: var ZipArchive; prefix: string; v: AdamVar; tmp: string) =
addT(z, prefix & "_m.npy", v.m, tmp)
addT(z, prefix & "_v.npy", v.v, tmp)
proc addLinearAdam(z: var ZipArchive; prefix: string; la: LinearAdam; tmp: string) =
addAdamVar(z, prefix & "_w", la.w, tmp)
addAdamVar(z, prefix & "_b", la.b, tmp)
proc addLSTMCellAdam(z: var ZipArchive; prefix: string; la: LSTMCellAdam; tmp: string) =
addAdamVar(z, prefix & "_wc", la.wCombined, tmp)
addAdamVar(z, prefix & "_bc", la.bCombined, tmp)
proc addActorAdam(z: var ZipArchive; prefix: string; a: ActorAdam; tmp: string) =
addLinearAdam(z, prefix & "_fc1", a.fc1, tmp)
addLSTMCellAdam(z, prefix & "_lstm", a.lstm, tmp)
addLinearAdam(z, prefix & "_fc2", a.fc2, tmp)
addLinearAdam(z, prefix & "_mu", a.muHead, tmp)
addLinearAdam(z, prefix & "_logstd", a.logStdHead, tmp)
proc addCriticAdam(z: var ZipArchive; prefix: string; c: CriticAdam; tmp: string) =
addLinearAdam(z, prefix & "_fc1", c.fc1, tmp)
addLSTMCellAdam(z, prefix & "_lstm", c.lstm, tmp)
addLinearAdam(z, prefix & "_fc2", c.fc2, tmp)
addLinearAdam(z, prefix & "_fc3", c.fc3, tmp)
# ── Public API ────────────────────────────────────────────────────────────────
proc saveWeights*(path: string;
actor: ActorNet;
critic1, critic2: CriticNet;
targetCritic1, targetCritic2: CriticNet;
alpha: float32) =
## Save network tensors (no Adam states) to `path` (.zip).
## Atomic: writes to a temp path first, then renames.
let tmp = tmpDir()
createDir(tmp)
let tmpZip = path & ".tmp"
try:
var z: ZipArchive
if not z.open(tmpZip, fmWrite):
raise newException(IOError, "cannot create zip: " & tmpZip)
addActorNet(z, "actor", actor, tmp)
addCriticNet(z, "c1", critic1, tmp)
addCriticNet(z, "c2", critic2, tmp)
addCriticNet(z, "tc1", targetCritic1,tmp)
addCriticNet(z, "tc2", targetCritic2,tmp)
# alpha: store as a 1-element tensor
addT(z, "alpha.npy", [alpha].toTensor.asType(float32), tmp)
z.close()
createDir(path.parentDir)
moveFile(tmpZip, path)
finally:
removeDir(tmp)
if fileExists(tmpZip): removeFile(tmpZip)
proc saveCheckpoint*(path: string;
actor: ActorNet;
critic1, critic2: CriticNet;
targetCritic1, targetCritic2: CriticNet;
alpha: float32;
adam: SACAdamStates) =
## Save networks + Adam states to `path` (.zip). Atomic.
let tmp = tmpDir()
createDir(tmp)
let tmpZip = path & ".tmp"
try:
var z: ZipArchive
if not z.open(tmpZip, fmWrite):
raise newException(IOError, "cannot create zip: " & tmpZip)
addActorNet(z, "actor", actor, tmp)
addCriticNet(z, "c1", critic1, tmp)
addCriticNet(z, "c2", critic2, tmp)
addCriticNet(z, "tc1", targetCritic1,tmp)
addCriticNet(z, "tc2", targetCritic2,tmp)
addT(z, "alpha.npy", [alpha].toTensor.asType(float32), tmp)
if adam.initialized:
addActorAdam(z, "adam_actor", adam.actor, tmp)
addCriticAdam(z, "adam_c1", adam.critic1, tmp)
addCriticAdam(z, "adam_c2", adam.critic2, tmp)
addAdamVar(z, "adam_alpha", adam.alpha, tmp)
# t counters (all in lockstep; store as text)
writeFile(tmp / "adam_t.txt",
$adam.actor.fc1.w.t & "\n" &
$adam.critic1.fc1.w.t & "\n" &
$adam.critic2.fc1.w.t & "\n" &
$adam.alpha.t)
z.addFile("adam_t.txt", tmp / "adam_t.txt")
z.close()
createDir(path.parentDir)
moveFile(tmpZip, path)
finally:
removeDir(tmp)
if fileExists(tmpZip): removeFile(tmpZip)
# ── Load helpers ──────────────────────────────────────────────────────────────
template loadT(name: string; tmp: string): Tensor[float32] =
read_npy[float32](tmp / name)
proc loadLinear(z: var ZipArchive; prefix, tmp: string): Linear =
z.extractFile(prefix & "_w.npy", tmp / (prefix & "_w.npy"))
z.extractFile(prefix & "_b.npy", tmp / (prefix & "_b.npy"))
result.w = read_npy[float32](tmp / (prefix & "_w.npy"))
result.b = read_npy[float32](tmp / (prefix & "_b.npy"))
proc loadLSTMCell(z: var ZipArchive; prefix, tmp: string): LSTMCell =
z.extractFile(prefix & "_wc.npy", tmp / (prefix & "_wc.npy"))
z.extractFile(prefix & "_bc.npy", tmp / (prefix & "_bc.npy"))
result.wCombined = read_npy[float32](tmp / (prefix & "_wc.npy"))
result.bCombined = read_npy[float32](tmp / (prefix & "_bc.npy"))
result.hiddenDim = result.bCombined.shape[0] div 4
proc loadActorNet(z: var ZipArchive; prefix, tmp: string): ActorNet =
result.fc1 = loadLinear(z, prefix & "_fc1", tmp)
result.lstm = loadLSTMCell(z, prefix & "_lstm", tmp)
result.fc2 = loadLinear(z, prefix & "_fc2", tmp)
result.muHead = loadLinear(z, prefix & "_mu", tmp)
result.logStdHead = loadLinear(z, prefix & "_logstd", tmp)
result.hiddenDim = result.lstm.hiddenDim
proc loadCriticNet(z: var ZipArchive; prefix, tmp: string): CriticNet =
result.fc1 = loadLinear(z, prefix & "_fc1", tmp)
result.lstm = loadLSTMCell(z, prefix & "_lstm", tmp)
result.fc2 = loadLinear(z, prefix & "_fc2", tmp)
result.fc3 = loadLinear(z, prefix & "_fc3", tmp)
result.hiddenDim = result.lstm.hiddenDim
proc loadAdamVarFromZip(z: var ZipArchive; prefix, tmp: string): AdamVar =
z.extractFile(prefix & "_m.npy", tmp / (prefix & "_m.npy"))
z.extractFile(prefix & "_v.npy", tmp / (prefix & "_v.npy"))
result.m = read_npy[float32](tmp / (prefix & "_m.npy"))
result.v = read_npy[float32](tmp / (prefix & "_v.npy"))
proc loadLinearAdam(z: var ZipArchive; prefix, tmp: string): LinearAdam =
result.w = loadAdamVarFromZip(z, prefix & "_w", tmp)
result.b = loadAdamVarFromZip(z, prefix & "_b", tmp)
proc loadLSTMCellAdam(z: var ZipArchive; prefix, tmp: string): LSTMCellAdam =
result.wCombined = loadAdamVarFromZip(z, prefix & "_wc", tmp)
result.bCombined = loadAdamVarFromZip(z, prefix & "_bc", tmp)
proc loadActorAdam(z: var ZipArchive; prefix, tmp: string): ActorAdam =
result.fc1 = loadLinearAdam(z, prefix & "_fc1", tmp)
result.lstm = loadLSTMCellAdam(z, prefix & "_lstm", tmp)
result.fc2 = loadLinearAdam(z, prefix & "_fc2", tmp)
result.muHead = loadLinearAdam(z, prefix & "_mu", tmp)
result.logStdHead = loadLinearAdam(z, prefix & "_logstd", tmp)
proc loadCriticAdam(z: var ZipArchive; prefix, tmp: string): CriticAdam =
result.fc1 = loadLinearAdam(z, prefix & "_fc1", tmp)
result.lstm = loadLSTMCellAdam(z, prefix & "_lstm", tmp)
result.fc2 = loadLinearAdam(z, prefix & "_fc2", tmp)
result.fc3 = loadLinearAdam(z, prefix & "_fc3", tmp)
type
WeightCheckpoint* = object
actor*: ActorNet
critic1*: CriticNet
critic2*: CriticNet
targetCritic1*: CriticNet
targetCritic2*: CriticNet
alpha*: float32
adam*: SACAdamStates ## initialized=false if not present in zip
proc loadCheckpoint*(path: string): WeightCheckpoint =
## Load all tensors from `path` (.zip). Raises IOError if file not found.
## Adam states loaded only if present; result.adam.initialized reflects this.
if not fileExists(path):
raise newException(IOError, "checkpoint not found: " & path)
let tmp = tmpDir()
createDir(tmp)
try:
var z: ZipArchive
if not z.open(path, fmRead):
raise newException(IOError, "cannot open zip: " & path)
result.actor = loadActorNet(z, "actor", tmp)
result.critic1 = loadCriticNet(z, "c1", tmp)
result.critic2 = loadCriticNet(z, "c2", tmp)
result.targetCritic1 = loadCriticNet(z, "tc1", tmp)
result.targetCritic2 = loadCriticNet(z, "tc2", tmp)
z.extractFile("alpha.npy", tmp / "alpha.npy")
let alphaTensor = read_npy[float32](tmp / "alpha.npy")
result.alpha = alphaTensor[0]
# Adam states — optional
var hasAdam = false
for f in z.walkFiles:
if f.startsWith("adam_"):
hasAdam = true
break
if hasAdam:
result.adam.actor = loadActorAdam(z, "adam_actor", tmp)
result.adam.critic1 = loadCriticAdam(z, "adam_c1", tmp)
result.adam.critic2 = loadCriticAdam(z, "adam_c2", tmp)
result.adam.alpha = loadAdamVarFromZip(z, "adam_alpha", tmp)
# t counters
z.extractFile("adam_t.txt", tmp / "adam_t.txt")
let ts = readFile(tmp / "adam_t.txt").strip().splitLines()
if ts.len >= 4:
let tActor = parseInt(ts[0])
let tCritic1 = parseInt(ts[1])
let tCritic2 = parseInt(ts[2])
let tAlpha = parseInt(ts[3])
# propagate t to all Adam vars
template setT(v: var AdamVar; tval: int) = v.t = tval
setT(result.adam.actor.fc1.w, tActor)
setT(result.adam.actor.fc1.b, tActor)
setT(result.adam.actor.lstm.wCombined,tActor)
setT(result.adam.actor.lstm.bCombined,tActor)
setT(result.adam.actor.fc2.w, tActor)
setT(result.adam.actor.fc2.b, tActor)
setT(result.adam.actor.muHead.w, tActor)
setT(result.adam.actor.muHead.b, tActor)
setT(result.adam.actor.logStdHead.w, tActor)
setT(result.adam.actor.logStdHead.b, tActor)
setT(result.adam.critic1.fc1.w, tCritic1)
setT(result.adam.critic1.fc1.b, tCritic1)
setT(result.adam.critic1.lstm.wCombined,tCritic1)
setT(result.adam.critic1.lstm.bCombined,tCritic1)
setT(result.adam.critic1.fc2.w, tCritic1)
setT(result.adam.critic1.fc2.b, tCritic1)
setT(result.adam.critic1.fc3.w, tCritic1)
setT(result.adam.critic1.fc3.b, tCritic1)
setT(result.adam.critic2.fc1.w, tCritic2)
setT(result.adam.critic2.fc1.b, tCritic2)
setT(result.adam.critic2.lstm.wCombined,tCritic2)
setT(result.adam.critic2.lstm.bCombined,tCritic2)
setT(result.adam.critic2.fc2.w, tCritic2)
setT(result.adam.critic2.fc2.b, tCritic2)
setT(result.adam.critic2.fc3.w, tCritic2)
setT(result.adam.critic2.fc3.b, tCritic2)
setT(result.adam.alpha, tAlpha)
result.adam.initialized = true
z.close()
finally:
removeDir(tmp)
+97
View File
@@ -0,0 +1,97 @@
import unittest
import arraymancer
import std/math
import SAC_LSTM_Bot/actions
proc makeOutput(a0, a1, a2, a3: float): Tensor[float32] =
result = newTensor[float32](4)
result[0] = a0.float32
result[1] = a1.float32
result[2] = a2.float32
result[3] = a3.float32
suite "mapActions":
test "ACTION_DIM is 4":
check ACTION_DIM == 4
# Speed-aware turn rate
test "turn: output +1 at speed 0 -> +10 degrees":
let m = mapActions(makeOutput(1.0, 0.0, 0.0, -1.0), 0.0, 0.0)
check abs(m.turnRate - 10.0) < 1e-6
test "turn: output -1 at speed 0 -> -10 degrees":
let m = mapActions(makeOutput(-1.0, 0.0, 0.0, -1.0), 0.0, 0.0)
check abs(m.turnRate - (-10.0)) < 1e-6
test "turn: output +1 at speed 8 -> +4 degrees":
let m = mapActions(makeOutput(1.0, 0.0, 0.0, -1.0), 8.0, 0.0)
check abs(m.turnRate - 4.0) < 1e-6
test "turn: output -1 at speed 8 -> -4 degrees":
let m = mapActions(makeOutput(-1.0, 0.0, 0.0, -1.0), 8.0, 0.0)
check abs(m.turnRate - (-4.0)) < 1e-6
test "turn: output 0 -> 0 regardless of speed":
let m = mapActions(makeOutput(0.0, 0.0, 0.0, -1.0), 5.0, 0.0)
check abs(m.turnRate) < 1e-6
# Acceleration range
test "accel: output -1 -> -2.0":
let m = mapActions(makeOutput(0.0, -1.0, 0.0, -1.0), 0.0, 0.0)
check abs(m.acceleration - (-2.0)) < 1e-6
test "accel: output +1 -> +1.0":
let m = mapActions(makeOutput(0.0, 1.0, 0.0, -1.0), 0.0, 0.0)
check abs(m.acceleration - 1.0) < 1e-6
test "accel: output 0 -> midpoint -0.5":
# value*1.5 - 0.5 at value=0 -> -0.5 (correct midpoint between -2 and +1)
let m = mapActions(makeOutput(0.0, 0.0, 0.0, -1.0), 0.0, 0.0)
check abs(m.acceleration - (-0.5)) < 1e-6
# Gun turn rate
test "gun turn: output +1 -> +20 degrees":
let m = mapActions(makeOutput(0.0, 0.0, 1.0, -1.0), 0.0, 0.0)
check abs(m.gunTurnRate - 20.0) < 1e-6
test "gun turn: output -1 -> -20 degrees":
let m = mapActions(makeOutput(0.0, 0.0, -1.0, -1.0), 0.0, 0.0)
check abs(m.gunTurnRate - (-20.0)) < 1e-6
test "gun turn: output 0 -> 0":
let m = mapActions(makeOutput(0.0, 0.0, 0.0, -1.0), 0.0, 0.0)
check abs(m.gunTurnRate) < 1e-6
# Fire threshold
test "fire: output -0.5 -> no fire (firePower == 0)":
let m = mapActions(makeOutput(0.0, 0.0, 0.0, -0.5), 0.0, 0.0)
check m.firePower == 0.0
test "fire: output 0.0 -> no fire (boundary, not positive)":
let m = mapActions(makeOutput(0.0, 0.0, 0.0, 0.0), 0.0, 0.0)
check m.firePower == 0.0
test "fire: output +0.5 -> fire with correct power":
let m = mapActions(makeOutput(0.0, 0.0, 0.0, 0.5), 0.0, 0.0)
# 0.5 * 2.9 + 0.1 = 1.55
check abs(m.firePower - 1.55) < 1e-5
test "fire: output +1.0 -> fire power near 3.0":
let m = mapActions(makeOutput(0.0, 0.0, 0.0, 1.0), 0.0, 0.0)
check abs(m.firePower - 3.0) < 1e-5
test "fire: output +1.0 but gunHeat > 0 -> no fire":
let m = mapActions(makeOutput(0.0, 0.0, 0.0, 1.0), 0.0, 1.5)
check m.firePower == 0.0
test "fire: output +0.001 (just above 0) -> fires with power near 0.1":
let m = mapActions(makeOutput(0.0, 0.0, 0.0, 0.001), 0.0, 0.0)
check m.firePower > 0.0
check m.firePower < 0.2
# Speed-aware turn with negative speed (reverse)
test "turn: speed -8 (reversing) -> same magnitude as speed +8":
let fwd = mapActions(makeOutput(1.0, 0.0, 0.0, -1.0), 8.0, 0.0)
let rev = mapActions(makeOutput(1.0, 0.0, 0.0, -1.0), -8.0, 0.0)
check abs(fwd.turnRate - rev.turnRate) < 1e-6
+150
View File
@@ -0,0 +1,150 @@
## Tests for integration.nim (#48) — assert-based, no framework.
## Covers: TrainingMsg channel round-trip (plain arrays through a channel),
## drain-then-train NewBattle/Shutdown handling, trainPass safety below canSample.
import arraymancer except Linear
import std/[locks, json, os]
import tankroyale_botapi # updateBotNames: seed the vendored id->name table
import SAC_LSTM_Bot/integration
import SAC_LSTM_Bot/state # STATE_DIM
import SAC_LSTM_Bot/actions # ACTION_DIM
import SAC_LSTM_Bot/training # initSACTrainer
import SAC_LSTM_Bot/replay_buffer
# ── 1. TrainingMsg round-trips through a channel with arrays intact ──────────
block:
var ch: Channel[TrainingMsg]
ch.open(4)
var msg = TrainingMsg(kind: tmkTransition)
for i in 0 ..< STATE_DIM:
msg.state[i] = float32(i) * 0.5'f32
msg.nextState[i] = float32(i) * 2.0'f32
for i in 0 ..< ACTION_DIM:
msg.action[i] = float32(i) - 2.0'f32
msg.reward = -1.25'f32
msg.done = true
assert ch.trySend(msg)
assert ch.trySend(TrainingMsg(kind: tmkNewBattle, enemyId: 4242))
assert ch.trySend(TrainingMsg(kind: tmkShutdown))
let r1 = ch.recv()
assert r1.kind == tmkTransition, "first msg is a transition"
for i in 0 ..< STATE_DIM:
assert r1.state[i] == float32(i) * 0.5'f32, "state round-trip at " & $i
assert r1.nextState[i] == float32(i) * 2.0'f32, "nextState round-trip at " & $i
for i in 0 ..< ACTION_DIM:
assert r1.action[i] == float32(i) - 2.0'f32, "action round-trip at " & $i
assert r1.reward == -1.25'f32 and r1.done
let r2 = ch.recv()
assert r2.kind == tmkNewBattle and r2.enemyId == 4242
let r3 = ch.recv()
assert r3.kind == tmkShutdown
ch.close()
# closed + empty -> tryRecv reports no data (Nim 2.2: recv would block forever)
let (ok4, _) = ch.tryRecv()
assert not ok4, "closed channel must report dataAvailable=false"
echo "PASS TrainingMsg channel round-trip"
# ── 2. NewBattle clears only when the opponent NAME changes; Shutdown stops ───
block:
# Seed the API's id->name table (v1.0.1) for name-based identity (#49).
updateBotNames(parseJson(
"""{"bots":[{"id":7,"name":"Corners"},{"id":8,"name":"Crazy"}]}"""))
assert opponentKey(7) == "Corners", "known id resolves to name"
assert opponentKey(99) == "99", "unknown id falls back to numeric string"
var st: TrainState
st.trainer = initSACTrainer(STATE_DIM, ACTION_DIM)
st.buf = newReplayBuffer(64, STATE_DIM, ACTION_DIM, burnIn = 2, trainWindow = 3)
assert handleTrainingMsg(st, TrainingMsg(kind: tmkNewBattle, enemyId: 7))
assert st.lastEnemyKey == "Corners"
for i in 0 ..< 5:
var m = TrainingMsg(kind: tmkTransition)
m.reward = float32(i)
assert handleTrainingMsg(st, m)
assert st.buf.len == 5, "transitions stored"
# Same opponent -> buffer kept (Q12a).
assert handleTrainingMsg(st, TrainingMsg(kind: tmkNewBattle, enemyId: 7))
assert st.buf.len == 5, "same opponent must NOT clear"
# Opponent changed -> clear.
assert handleTrainingMsg(st, TrainingMsg(kind: tmkNewBattle, enemyId: 8))
assert st.buf.len == 0, "opponent change must clear"
assert st.lastEnemyKey == "Crazy"
# Nameless window (pre-BotListUpdate) is a distinct key.
assert handleTrainingMsg(st, TrainingMsg(kind: tmkNewBattle, enemyId: 99))
assert st.lastEnemyKey == "99" and st.buf.len == 0
# Shutdown stops the caller's loop.
assert not handleTrainingMsg(st, TrainingMsg(kind: tmkShutdown))
echo "PASS name-keyed NewBattle clear / Shutdown"
# ── 2b. bumpRoundCounter increments per call, cold-starts at 1 ────────────────
block:
let tmp = getTempDir() / "sac_test_weights_" & $getCurrentProcessId()
putEnv("SACLSTM_WEIGHTS_PATH", tmp / "sac_latest.zip")
bumpRoundCounter()
bumpRoundCounter()
assert readFile(tmp / "round_counter.txt") == "2", "counter increments per round"
echo "PASS bumpRoundCounter"
# ── 3. trainPass is a safe no-op below canSample (no steps, no publish) ───────
block:
var st: TrainState
st.trainer = initSACTrainer(STATE_DIM, ACTION_DIM)
st.buf = newReplayBuffer(64, STATE_DIM, ACTION_DIM, burnIn = 2, trainWindow = 3)
for i in 0 ..< 4:
assert handleTrainingMsg(st, TrainingMsg(kind: tmkTransition))
trainPass(st, 4) # 4 < burnIn+trainWindow = 5
assert st.stepCount == 0, "no gradient steps below canSample"
echo "PASS trainPass no-op below canSample"
# ── 4. Flat snapshot layout: pack/unpack round-trips weights exactly ──────────
block:
let t0 = initSACTrainer(STATE_DIM, ACTION_DIM)
let fs = packFull(t0)
assert fs.data.len == actorSize(t0.actor.hiddenDim) + 4 * criticSize(t0.actor.hiddenDim) + 1
let (a, c1, c2, tc1, tc2, alpha) = unpackFull(fs)
assert a.hiddenDim == t0.actor.hiddenDim
assert c1.fc3.b.shape[0] == 1
let fw = a.muHead.w.flatten()
let fw0 = t0.actor.muHead.w.flatten()
for i in 0 ..< fw.size:
assert fw[i] == fw0[i], "actor mu weights round-trip"
for i in 0 ..< c2.lstm.bCombined.size:
assert c2.lstm.bCombined[i] == t0.critic2.lstm.bCombined[i], "critic lstm bias round-trip"
assert alpha == t0.alpha()
discard tc1
echo "PASS flat snapshot pack/unpack round-trip"
# ── 5. Lever 3 (#59): metricsLine emits exactly the exposed trainer scalars ───
block:
let line = metricsLine(1787394115.123, 42, 500, 20, 20,
SACMetrics(criticLoss: 0.5'f32, actorLoss: -1.5'f32,
alphaLoss: 0.25'f32, alpha: 2.0'f32))
let j = parseJson(line) # throws on malformed JSONL
assert j["steps"].getInt() == 42 and j["buffer_size"].getInt() == 500
assert j["drained"].getInt() == 20 and j["grad_steps"].getInt() == 20
assert abs(j["critic_loss"].getFloat() - 0.5) < 1e-3
assert abs(j["actor_loss"].getFloat() + 1.5) < 1e-3
assert abs(j["alpha_loss"].getFloat() - 0.25) < 1e-3
assert abs(j["alpha"].getFloat() - 2.0) < 1e-3
assert j["epoch"].getFloat() > 1e9
echo "PASS metricsLine JSONL scalars"
# ── 6. Lever 4 (#59): SACLSTM_EVAL_MODE=1 suppresses training input ───────────
block:
putEnv("SACLSTM_EVAL_MODE", "1")
# Gate fires before any channel traffic: false = dropped, nothing enqueued.
assert not sendTrainingMsg(TrainingMsg(kind: tmkTransition)),
"eval mode must drop transitions"
assert not sendTrainingMsg(TrainingMsg(kind: tmkNewBattle, enemyId: 7)),
"eval mode must drop NewBattle (no buffer clears from eval)"
delEnv("SACLSTM_EVAL_MODE")
echo "PASS eval-mode training-input suppression"
+99
View File
@@ -0,0 +1,99 @@
import unittest
import arraymancer
import std/[math, os]
import SAC_LSTM_Bot/network
const
STATE_DIM = 20
ACTION_DIM = 4
suite "ActorNet":
setup:
let actor = initActorNet(STATE_DIM)
let state = randomNormalTensor[float32](STATE_DIM)
let ls0 = zeroState(actor.hiddenDim)
test "forward output shapes":
let (actions, lp, ls1) = actorForward(actor, state, ls0)
check actions.shape == [4]
check ls1.h.shape == [actor.hiddenDim]
check ls1.c.shape == [actor.hiddenDim]
check not lp.isNaN
test "actions clamped in (-1, 1) — tanh output":
let (actions, _, _) = actorForward(actor, state, ls0)
for i in 0 ..< 4:
check actions[i] > -1.0'f32
check actions[i] < 1.0'f32
test "hidden state propagates (h/c change after step)":
let (_, _, ls1) = actorForward(actor, state, ls0)
# h' should differ from zero init for non-trivial input
var hChanged = false
for i in 0 ..< actor.hiddenDim:
if abs(ls1.h[i] - ls0.h[i]) > 1e-7'f32:
hChanged = true
break
check hChanged
test "deterministic mode: same input → same output":
let (a1, _, _) = actorForward(actor, state, ls0, deterministic = true)
let (a2, _, _) = actorForward(actor, state, ls0, deterministic = true)
for i in 0 ..< 4:
check abs(a1[i] - a2[i]) < 1e-7'f32
test "stochastic mode: outputs may differ (sampling)":
# Run many times; at least one pair should differ
let (a1, _, _) = actorForward(actor, state, ls0)
let (a2, _, _) = actorForward(actor, state, ls0)
var anyDiff = false
for i in 0 ..< 4:
if abs(a1[i] - a2[i]) > 1e-7'f32:
anyDiff = true
break
check anyDiff
suite "CriticNet":
setup:
let critic = initCriticNet(STATE_DIM, ACTION_DIM)
let state = randomNormalTensor[float32](STATE_DIM)
let actions = randomNormalTensor[float32](ACTION_DIM)
let stateAct = concat(state, actions, axis = 0)
let ls0 = zeroState(critic.hiddenDim)
test "forward output shape":
let (q, ls1) = criticForward(critic, stateAct, ls0)
check ls1.h.shape == [critic.hiddenDim]
check ls1.c.shape == [critic.hiddenDim]
# q is scalar float — just ensure it doesn't NaN
check not q.isNaN
test "hidden state propagates":
let (_, ls1) = criticForward(critic, stateAct, ls0)
var hChanged = false
for i in 0 ..< critic.hiddenDim:
if abs(ls1.h[i] - ls0.h[i]) > 1e-7'f32:
hChanged = true
break
check hChanged
suite "Dual critics":
test "two independent critics produce different Q values":
let c1 = initCriticNet(STATE_DIM, ACTION_DIM)
let c2 = initCriticNet(STATE_DIM, ACTION_DIM)
let sa = randomNormalTensor[float32](STATE_DIM + ACTION_DIM)
let ls = zeroState(c1.hiddenDim)
let (q1, _) = criticForward(c1, sa, ls)
let (q2, _) = criticForward(c2, sa, ls)
check abs(q1 - q2) > 1e-7'f32
suite "Hidden size config":
test "hidden size 128 works":
putEnv("SACLSTM_HIDDEN_SIZE", "128")
let actor = initActorNet(STATE_DIM)
let state = randomNormalTensor[float32](STATE_DIM)
let ls0 = zeroState(128)
let (actions, _, ls1) = actorForward(actor, state, ls0)
check actions.shape == [4]
check ls1.h.shape == [128]
putEnv("SACLSTM_HIDDEN_SIZE", "256")
+156
View File
@@ -0,0 +1,156 @@
## Tests for replay_buffer.nim — assert-based, no framework.
import arraymancer
import std/sequtils
import SAC_LSTM_Bot/replay_buffer
import SAC_LSTM_Bot/state # STATE_DIM
const
S_DIM = STATE_DIM # 35
A_DIM = 4
proc makeTrans(reward: float32; done: bool): Transition =
Transition(
state: zeros[float32](S_DIM),
action: zeros[float32](A_DIM),
reward: reward,
nextState: zeros[float32](S_DIM),
done: done
)
# ── 1. len and canSample on empty buffer ──────────────────────────────────────
block:
var buf = newReplayBuffer(100, S_DIM, A_DIM, burnIn = 8, trainWindow = 16)
assert buf.len == 0, "empty len"
assert not buf.canSample, "empty canSample"
echo "PASS empty buffer"
# ── 2. len grows, canSample becomes true ──────────────────────────────────────
block:
var buf = newReplayBuffer(100, S_DIM, A_DIM, burnIn = 2, trainWindow = 3)
for i in 0 ..< 4:
buf.add(makeTrans(float32(i), false))
assert buf.len == 4
assert not buf.canSample, "needs 5 to sample"
buf.add(makeTrans(99, false))
assert buf.len == 5
assert buf.canSample, "5 transitions, seqLen=5 → canSample"
echo "PASS len + canSample"
# ── 3. Ring buffer wraps at capacity ─────────────────────────────────────────
block:
var buf = newReplayBuffer(10, S_DIM, A_DIM, burnIn = 2, trainWindow = 3)
for i in 0 ..< 15:
buf.add(makeTrans(float32(i), false))
assert buf.len == 10, "wraps at capacity, len stays 10"
echo "PASS wrap at capacity"
# ── 4. Sampled sequences have correct length ──────────────────────────────────
block:
var buf = newReplayBuffer(200, S_DIM, A_DIM, burnIn = 4, trainWindow = 6)
for i in 0 ..< 50:
buf.add(makeTrans(float32(i), false))
let seqs = buf.sampleSequences(8)
assert seqs.len > 0, "should have valid starts"
for s in seqs:
assert s.burnIn.len == 4, "burnIn len"
assert s.train.len == 6, "train len"
echo "PASS sequence length"
# ── 5. Burn-in / train split is correct ───────────────────────────────────────
block:
# Fill with distinct rewards so we can identify positions
var buf = newReplayBuffer(200, S_DIM, A_DIM, burnIn = 3, trainWindow = 4)
for i in 0 ..< 30:
buf.add(makeTrans(float32(i), false))
let seqs = buf.sampleSequences(1)
assert seqs.len == 1
let s = seqs[0]
# The 4th reward of the sequence (index 3) must match s.train[0].reward
# and s.burnIn[2].reward must be s.burnIn[2].reward (just check no overlap)
let allRewards = s.burnIn.mapIt(it.reward) & s.train.mapIt(it.reward)
# consecutive integer rewards → each must be strictly increasing by 1
var ok = true
for i in 1 ..< allRewards.len:
if allRewards[i] != allRewards[i-1] + 1.0f32:
ok = false
break
assert ok, "burn-in and train must form a contiguous sequence"
echo "PASS burn-in/train split"
# ── 6. Sequences never cross a done=true boundary ─────────────────────────────
block:
var buf = newReplayBuffer(200, S_DIM, A_DIM, burnIn = 2, trainWindow = 3)
# Episode 1: transitions 0..4 (done at index 4)
for i in 0 ..< 4:
buf.add(makeTrans(float32(i), false))
buf.add(makeTrans(99, true)) # battle end at index 4
# Episode 2: transitions 5..14
for i in 5 ..< 15:
buf.add(makeTrans(float32(i), false))
let seqs = buf.sampleSequences(20)
for s in seqs:
# No transition in burnIn (except the last) or train (except the last)
# may have done=true, since that would mean the next step crosses a boundary.
for i in 0 ..< s.burnIn.len - 1:
assert not s.burnIn[i].done, "done in middle of burnIn"
for i in 0 ..< s.train.len - 1:
assert not s.train[i].done, "done in middle of train"
# The join between burnIn and train must not cross a done=true
if s.burnIn.len > 0:
assert not s.burnIn[^1].done, "done at end of burnIn crosses boundary to train"
echo "PASS no cross-boundary sequences"
# ── 7. canSample false when buffer too small ───────────────────────────────────
block:
var buf = newReplayBuffer(100, S_DIM, A_DIM, burnIn = 8, trainWindow = 16)
for i in 0 ..< 23: # seqLen = 24; 23 < 24
buf.add(makeTrans(float32(i), false))
assert not buf.canSample, "23 < 24 seqLen"
buf.add(makeTrans(23, false))
assert buf.canSample, "24 == seqLen"
echo "PASS canSample threshold"
# ── 8. Empty buffer doesn't crash on sample ────────────────────────────────────
block:
var buf = newReplayBuffer(100, S_DIM, A_DIM, burnIn = 8, trainWindow = 16)
let seqs = buf.sampleSequences(4)
assert seqs.len == 0, "empty buffer → empty result"
echo "PASS empty sample"
# ── 9. Single episode (no done except at very end) ────────────────────────────
block:
var buf = newReplayBuffer(200, S_DIM, A_DIM, burnIn = 3, trainWindow = 5)
for i in 0 ..< 19:
buf.add(makeTrans(float32(i), false))
buf.add(makeTrans(19, true)) # last
let seqs = buf.sampleSequences(5)
assert seqs.len > 0, "should find valid starts"
for s in seqs:
assert s.burnIn.len == 3
assert s.train.len == 5
echo "PASS single episode"
# ── 10. Multiple short episodes: all boundaries respected ─────────────────────
block:
var buf = newReplayBuffer(300, S_DIM, A_DIM, burnIn = 2, trainWindow = 3)
# 10 episodes of 5 transitions each (done at end of each episode)
var reward = 0'f32
for ep in 0 ..< 10:
for i in 0 ..< 4:
buf.add(makeTrans(reward, false))
reward += 1
buf.add(makeTrans(reward, true)) # battle end
reward += 1
let seqs = buf.sampleSequences(30)
assert seqs.len > 0
for s in seqs:
# No done in the middle of any sequence
let all = s.burnIn & s.train
for i in 0 ..< all.len - 1:
assert not all[i].done, "boundary crossed in multi-episode test"
echo "PASS multiple episodes"
echo "ALL TESTS PASSED"
+35 -5
View File
@@ -11,17 +11,46 @@ template check(cond: bool, msg: string) =
# ── computeReward ───────────────────────────────────────────────────────────── # ── computeReward ─────────────────────────────────────────────────────────────
block damageInflicted: block damageInflicted:
# p=1: 6*1 - 2 = 4 # p=1: 1.25 * (6*1 - 2) = 5 (lever 2 aggression mult, #59)
check abs(computeReward(damageInflicted = 1.0) - 4.0) < 1e-9, "p=1 damage = +4" check abs(computeReward(damageInflicted = 1.0) - 5.0) < 1e-9, "p=1 damage = +5"
# p=3: 6*3 - 2 = 16 # p=3: 1.25 * (6*3 - 2) = 20
check abs(computeReward(damageInflicted = 3.0) - 16.0) < 1e-9, "p=3 damage = +16" check abs(computeReward(damageInflicted = 3.0) - 20.0) < 1e-9, "p=3 damage = +20"
# low-power spam stays unprofitable: 1.25*(6*0.1-2) < 0
check computeReward(damageInflicted = 0.1) < 0.0, "p=0.1 spam still negative"
block damageReceived: block damageReceived:
# p_e=1: -(6*1 - 2) = -4 # p_e=1: -(6*1 - 2) = -4 (unchanged by lever 2)
check abs(computeReward(damageReceived = 1.0) - (-4.0)) < 1e-9, "p_e=1 received = -4" check abs(computeReward(damageReceived = 1.0) - (-4.0)) < 1e-9, "p_e=1 received = -4"
# p_e=3: -(6*3 - 2) = -16 # p_e=3: -(6*3 - 2) = -16
check abs(computeReward(damageReceived = 3.0) - (-16.0)) < 1e-9, "p_e=3 received = -16" check abs(computeReward(damageReceived = 3.0) - (-16.0)) < 1e-9, "p_e=3 received = -16"
block hitBonus:
# flat +0.5 per landed shot: p=1 hit -> 5.0 + 0.5
check abs(computeReward(damageInflicted = 1.0, hitCount = 1) - 5.5) < 1e-9,
"p=1 hit = +5.5"
# two hits in one step: 1.25*(6*2-2) + 2*0.5 = 12.5 + 1.0 = 13.5
check abs(computeReward(damageInflicted = 2.0, hitCount = 2) - 13.5) < 1e-9,
"two hits = +13.5"
block ramTaken:
# flat per victim collision (#59)
check abs(computeReward(ramTakenCount = 1) - (-3.0)) < 1e-9, "ram taken x1 = -3"
check abs(computeReward(ramTakenCount = 2) - (-6.0)) < 1e-9, "ram taken x2 = -6"
block chargeDeterrent:
# zero-damage case at half threshold depth: -2 * (1 - 0.06/0.12) = -1
let rHalf = computeReward(enemyDistFrac = 0.06)
check abs(rHalf - (-1.0)) < 1e-9, "charge at frac 0.06 = -1"
# at zero distance: full ceiling
check abs(computeReward(enemyDistFrac = 0.0) - (-2.0)) < 1e-9, "charge at frac 0 = -2"
# at/beyond threshold and no-contact sentinel: no penalty
check abs(computeReward(enemyDistFrac = 0.12)) < 1e-9, "at threshold = 0"
check abs(computeReward(enemyDistFrac = 0.5)) < 1e-9, "beyond threshold = 0"
check abs(computeReward(enemyDistFrac = 2.0)) < 1e-9, "no-contact sentinel = 0"
# suppressed while dealing damage that step (fighting back at close range is fine)
let rFight = computeReward(damageInflicted = 1.0, enemyDistFrac = 0.06)
check abs(rFight - 5.0) < 1e-9, "dealing damage cancels charge penalty"
block wallHit: block wallHit:
check abs(computeReward(wallHitTicks = 1) - (-5.0)) < 1e-9, "1 wall tick = -5" check abs(computeReward(wallHitTicks = 1) - (-5.0)) < 1e-9, "1 wall tick = -5"
@@ -30,6 +59,7 @@ block wastedShot:
check abs(computeReward(wastedShotPower = 2.0) - (-0.2)) < 1e-9, "missed p=2 = -0.2" check abs(computeReward(wastedShotPower = 2.0) - (-0.2)) < 1e-9, "missed p=2 = -0.2"
block winLoss: block winLoss:
# terminal terms stay dominant over shaping (#59 scale discipline)
check abs(computeReward(win = true) - 20.0) < 1e-9, "win = +20" check abs(computeReward(win = true) - 20.0) < 1e-9, "win = +20"
check abs(computeReward(loss = true) - (-10.0)) < 1e-9, "loss = -10" check abs(computeReward(loss = true) - (-10.0)) < 1e-9, "loss = -10"
+159
View File
@@ -0,0 +1,159 @@
## test_training.nim — stdlib unittest for SAC training module.
import unittest
import arraymancer
import std/[math, random, os]
import SAC_LSTM_Bot/network
import SAC_LSTM_Bot/weights
import SAC_LSTM_Bot/replay_buffer
import SAC_LSTM_Bot/training
const
STATE_DIM = 10
ACTION_DIM = 4
HIDDEN = 16 # small for speed; override via env not needed in tests
proc makeTrainer(): SACTrainer =
## Small trainer for tests — override hidden size via env before calling.
putEnv("SACLSTM_HIDDEN_SIZE", "16")
initSACTrainer(STATE_DIM, ACTION_DIM)
proc makeBuffer(): ReplayBuffer =
newReplayBuffer(capacity = 2000, stateDim = STATE_DIM, actionDim = ACTION_DIM,
burnIn = 4, trainWindow = 8)
proc randState(): Tensor[float32] =
randomNormalTensor[float32](STATE_DIM)
proc randAction(): Tensor[float32] =
randomTensor[float32](ACTION_DIM, 1.0'f32) *. 2.0'f32 -. 1.0'f32 # uniform [-1,1]
proc fillBuffer(buf: var ReplayBuffer; n: int; donePeriod = 20) =
for i in 0 ..< n:
let t = Transition(
state: randState(),
action: randAction(),
reward: rand(-1.0'f32 .. 1.0'f32),
nextState: randState(),
done: (i mod donePeriod == donePeriod - 1))
buf.add(t)
proc isFinite(x: float32): bool =
not (x != x) and x < Inf and x > -Inf # not NaN and not Inf
suite "SACTrainer — basic update":
setup:
randomize(42)
var trainer = makeTrainer()
var buf = makeBuffer()
fillBuffer(buf, 1000)
let seqs = buf.sampleSequences(4)
test "sampleSequences returns non-empty batch":
check seqs.len > 0
test "sacUpdate returns finite losses":
let m = sacUpdate(trainer, seqs)
check isFinite(m.criticLoss)
check isFinite(m.actorLoss)
check isFinite(m.alphaLoss)
test "alpha stays positive after update":
var t2 = trainer
discard sacUpdate(t2, seqs)
check t2.alpha() > 0.0'f32
test "critic1 weights change after update":
let w_before = trainer.critic1.fc3.w.clone()
discard sacUpdate(trainer, seqs)
let w_after = trainer.critic1.fc3.w
var changed = false
for i in 0 ..< w_before.size:
if abs(w_before.unsafe_raw_offset[i] - w_after.unsafe_raw_offset[i]) > 1e-10'f32:
changed = true
break
check changed
test "critic2 weights change after update":
let w_before = trainer.critic2.fc3.w.clone()
discard sacUpdate(trainer, seqs)
let w_after = trainer.critic2.fc3.w
var changed = false
for i in 0 ..< w_before.size:
if abs(w_before.unsafe_raw_offset[i] - w_after.unsafe_raw_offset[i]) > 1e-10'f32:
changed = true
break
check changed
test "actor weights change after update":
let w_before = trainer.actor.fc1.w.clone()
discard sacUpdate(trainer, seqs)
let w_after = trainer.actor.fc1.w
var changed = false
for i in 0 ..< w_before.size:
if abs(w_before.unsafe_raw_offset[i] - w_after.unsafe_raw_offset[i]) > 1e-10'f32:
changed = true
break
check changed
test "soft target update: target moves toward critic":
## After update, target fc3.w should be closer to critic1.fc3.w than before.
let targetBefore = trainer.targetCritic1.fc3.w.clone()
let criticW = trainer.critic1.fc3.w.clone()
discard sacUpdate(trainer, seqs)
let targetAfter = trainer.targetCritic1.fc3.w
# Distance before: ||targetBefore - criticW||
var distBefore = 0.0'f32
for i in 0 ..< targetBefore.size:
let d = targetBefore.unsafe_raw_offset[i] - criticW.unsafe_raw_offset[i]
distBefore += d * d
# Distance after: ||targetAfter - criticW_new|| (critic changed too, use original for ref)
var distAfter = 0.0'f32
for i in 0 ..< targetAfter.size:
let d = targetAfter.unsafe_raw_offset[i] - criticW.unsafe_raw_offset[i]
distAfter += d * d
# Target moved toward critic (distance decreased).
# With tau=0.005, it moves a tiny bit — just check direction.
check distAfter <= distBefore + 1e-3'f32 # soft bound; critic also moves
test "empty sequences is a no-op":
let m = sacUpdate(trainer, @[])
check m.criticLoss == 0.0'f32
check m.actorLoss == 0.0'f32
check m.alpha == 0.0'f32
suite "SACTrainer — done=true terminal transitions":
test "update with done=true transitions produces finite losses":
randomize(7)
putEnv("SACLSTM_HIDDEN_SIZE", "16")
var trainer = makeTrainer()
var buf = makeBuffer()
# Fill with short episodes: done every 12 steps (burnIn=4, trainWindow=8 → seqLen=12)
fillBuffer(buf, 800, donePeriod = 12)
let seqs = buf.sampleSequences(2)
if seqs.len > 0:
let m = sacUpdate(trainer, seqs)
check isFinite(m.criticLoss)
check isFinite(m.actorLoss)
check isFinite(m.alphaLoss)
check m.alpha > 0.0'f32
suite "SACTrainer — multiple updates":
test "three sequential updates all stay finite":
randomize(99)
putEnv("SACLSTM_HIDDEN_SIZE", "16")
var trainer = makeTrainer()
var buf = makeBuffer()
fillBuffer(buf, 1000)
for _ in 1..3:
let seqs = buf.sampleSequences(4)
if seqs.len > 0:
let m = sacUpdate(trainer, seqs)
check isFinite(m.criticLoss)
check isFinite(m.actorLoss)
check m.alpha > 0.0'f32
+141
View File
@@ -0,0 +1,141 @@
import unittest
import arraymancer
import zip/zipfiles
import std/[os, math, strutils]
import SAC_LSTM_Bot/network
import SAC_LSTM_Bot/weights
const
STATE_DIM = 20
ACTION_DIM = 4
proc tensorsEqual(a, b: Tensor[float32]; tol: float32 = 1e-6'f32): bool =
if a.shape != b.shape: return false
for i in 0 ..< a.size:
if abs(a.unsafe_raw_offset[i] - b.unsafe_raw_offset[i]) > tol: return false
true
suite "saveWeights / loadCheckpoint":
setup:
let actor = initActorNet(STATE_DIM)
let c1 = initCriticNet(STATE_DIM, ACTION_DIM)
let c2 = initCriticNet(STATE_DIM, ACTION_DIM)
let tc1 = initCriticNet(STATE_DIM, ACTION_DIM)
let tc2 = initCriticNet(STATE_DIM, ACTION_DIM)
let alpha = 0.2'f32
let zipPath = getTempDir() / "test_weights_latest.zip"
teardown:
if fileExists(zipPath): removeFile(zipPath)
test "creates a valid zip file":
saveWeights(zipPath, actor, c1, c2, tc1, tc2, alpha)
check fileExists(zipPath)
var z: ZipArchive
check z.open(zipPath, fmRead)
var count = 0
for f in z.walkFiles: inc count
z.close()
check count > 0
test "all entries are .npy files (+ alpha.npy)":
saveWeights(zipPath, actor, c1, c2, tc1, tc2, alpha)
var z: ZipArchive
discard z.open(zipPath, fmRead)
var allNpy = true
for f in z.walkFiles:
if not f.endsWith(".npy"): allNpy = false
z.close()
check allNpy
test "round-trip: actor weights preserved":
saveWeights(zipPath, actor, c1, c2, tc1, tc2, alpha)
let ck = loadCheckpoint(zipPath)
check tensorsEqual(actor.fc1.w, ck.actor.fc1.w)
check tensorsEqual(actor.fc1.b, ck.actor.fc1.b)
check tensorsEqual(actor.lstm.wCombined, ck.actor.lstm.wCombined)
check tensorsEqual(actor.lstm.bCombined, ck.actor.lstm.bCombined)
check tensorsEqual(actor.fc2.w, ck.actor.fc2.w)
check tensorsEqual(actor.muHead.w, ck.actor.muHead.w)
check tensorsEqual(actor.logStdHead.w, ck.actor.logStdHead.w)
test "round-trip: critic weights preserved":
saveWeights(zipPath, actor, c1, c2, tc1, tc2, alpha)
let ck = loadCheckpoint(zipPath)
check tensorsEqual(c1.fc1.w, ck.critic1.fc1.w)
check tensorsEqual(c1.lstm.wCombined, ck.critic1.lstm.wCombined)
check tensorsEqual(c1.fc3.w, ck.critic1.fc3.w)
check tensorsEqual(tc1.fc1.w, ck.targetCritic1.fc1.w)
check tensorsEqual(tc2.fc1.w, ck.targetCritic2.fc1.w)
test "round-trip: alpha preserved":
saveWeights(zipPath, actor, c1, c2, tc1, tc2, alpha)
let ck = loadCheckpoint(zipPath)
check abs(ck.alpha - alpha) < 1e-6'f32
test "round-trip: hiddenDim reconstructed":
saveWeights(zipPath, actor, c1, c2, tc1, tc2, alpha)
let ck = loadCheckpoint(zipPath)
check ck.actor.hiddenDim == actor.hiddenDim
check ck.critic1.hiddenDim == c1.hiddenDim
test "adam not present → initialized=false":
saveWeights(zipPath, actor, c1, c2, tc1, tc2, alpha)
let ck = loadCheckpoint(zipPath)
check not ck.adam.initialized
suite "saveCheckpoint (with Adam)":
setup:
let actor = initActorNet(STATE_DIM)
let c1 = initCriticNet(STATE_DIM, ACTION_DIM)
let c2 = initCriticNet(STATE_DIM, ACTION_DIM)
let tc1 = initCriticNet(STATE_DIM, ACTION_DIM)
let tc2 = initCriticNet(STATE_DIM, ACTION_DIM)
let alpha = 0.1'f32
var adam = initSACAdamStates(actor, c1, c2)
# put some non-zero values in Adam state
adam.actor.fc1.w.m[0, 0] = 0.5'f32
adam.actor.fc1.w.t = 42
adam.critic1.fc1.w.t = 7
adam.alpha.t = 99
let zipPath = getTempDir() / "test_checkpoint.zip"
teardown:
if fileExists(zipPath): removeFile(zipPath)
test "round-trip: Adam m tensor":
saveCheckpoint(zipPath, actor, c1, c2, tc1, tc2, alpha, adam)
let ck = loadCheckpoint(zipPath)
check ck.adam.initialized
check abs(ck.adam.actor.fc1.w.m[0, 0] - 0.5'f32) < 1e-6'f32
test "round-trip: Adam t counters":
saveCheckpoint(zipPath, actor, c1, c2, tc1, tc2, alpha, adam)
let ck = loadCheckpoint(zipPath)
check ck.adam.actor.fc1.w.t == 42
check ck.adam.critic1.fc1.w.t == 7
check ck.adam.alpha.t == 99
suite "Atomic save":
test "temp file is cleaned up after successful save":
let zipPath = getTempDir() / "test_atomic.zip"
let tmpZip = zipPath & ".tmp"
let actor = initActorNet(STATE_DIM)
let c = initCriticNet(STATE_DIM, ACTION_DIM)
saveWeights(zipPath, actor, c, c, c, c, 0.2'f32)
check fileExists(zipPath)
check not fileExists(tmpZip)
removeFile(zipPath)
suite "Error handling":
test "loadCheckpoint missing file → IOError":
var raised = false
try:
discard loadCheckpoint(getTempDir() / "nonexistent_xxxxxx.zip")
except IOError:
raised = true
check raised
+578
View File
@@ -0,0 +1,578 @@
#!/usr/bin/env python3
"""Build the single self-updating SAC training dashboard.
Pure-stdlib SVG output (matplotlib not available on this box).
Generates ONE file:
docs/campaign_dashboard.svg - five panels, current (v2) run only:
1. test-match win % vs opponents (campaign_v2_stdout.log eval lines)
2. real-fight win % per opponent (training_log.jsonl, ~100-game buckets)
3. critic_loss / |actor_loss| (training_metrics.jsonl, shared log-y)
4. alpha temperature (training_metrics.jsonl, linear)
5. throughput, games/hour buckets (training_metrics.jsonl 'epoch' deltas;
training_log.jsonl has NO timestamps -
verified field names - and one metrics
row == one 10-game chunk, counts match
the stdout "=== Chunk N/N ===" markers)
plus an embedded JS snippet that reloads the page every 60 s when the SVG is
opened as a top-level document in Chrome.
Usage:
python3 tools/plot_progress.py [campaign_log] [metrics_jsonl] [games_jsonl] [outdir]
All args optional; defaults relative to the SAC_LSTM_Bot/ root (parent of tools/).
python3 tools/plot_progress.py --selftest # tiny built-in sanity check
"""
import json
import math
import os
import re
import sys
import tempfile
from datetime import datetime
from pathlib import Path
import xml.etree.ElementTree as ET
ROOT = Path(__file__).resolve().parent.parent
EVAL_RE = re.compile(r">>> \[eval\] win rate: (\d+)/(\d+) \(([\d.]+)%\) vs (\S+)")
TREND_WINDOW = 10 # rolling mean shown as the thick trend line (panel 1)
GAME_BUCKET = 100 # games per bucket, real-fight panel
RATE_BUCKET = 20 # metric intervals per throughput bucket (~200 games)
GAMES_PER_ROW = 10 # one training_metrics.jsonl row per 10-round chunk
COLORS = {"Corners": "#d62728", "Crazy": "#1f77b4", "Target": "#2ca02c",
"RamFire": "#ff7f0e", "SacTwin": "#9467bd"}
CRITIC_C, ACTOR_C = "#1f77b4", "#ff7f0e"
RELOAD_JS = ('<script type="text/javascript"><![CDATA[ '
'setTimeout(function(){ location.reload(); }, 60000); ]]></script>')
W, H = 1400, 2000
TITLE_H, GUIDE_H = 80, 180
# rows: (header_y, panel_top_y, panel_bottom_y, x_left, x_right)
C1_L, C1_R = 70, 697
C2_L, C2_R = 747, 1375
ROWS = {
1: (100, 120, 720, C1_L, C1_R),
2: (100, 120, 720, C2_L, C2_R),
3: (940, 958, 1428, C1_L, C1_R),
4: (940, 958, 1428, C2_L, C2_R),
5: (1600, 1618, 1780, C1_L, C2_R),
}
PANEL_TITLES = [
"Test matches - win % vs opponents",
"Real fights - win % per opponent",
"Training losses (log scale)",
"Alpha temperature",
"Throughput - games per hour",
]
GUIDE_LINES = [
"Test matches: dots are single fights, thick line shows trend.",
"Real battles only. Rising lines mean the bot improves.",
"Loss spikes are normal early; endless growth is bad.",
"Alpha high means experimenting; falling too fast freezes habits.",
"Throughput flat is healthy; dips mean something slowed.",
"This file reloads itself in Chrome every sixty seconds.",
"Regenerate anytime with tools/watch_dashboard.sh or the python command.",
]
# ---------- parsing ----------
def parse_eval_series(path):
"""Return {opponent: [win% per eval, in file order]}."""
series = {}
if not path.is_file():
print(f"[skip] eval log not found: {path}")
return series
for line in path.read_text(errors="replace").splitlines():
m = EVAL_RE.search(line)
if not m:
continue
try:
pct = float(m.group(3))
except ValueError:
continue
series.setdefault(m.group(4), []).append(pct)
return series
def parse_metrics(path):
"""Return list of metric dicts, skipping malformed lines."""
rows = []
if not path.is_file():
print(f"[skip] metrics file not found: {path}")
return rows
for line in path.read_text(errors="replace").splitlines():
line = line.strip()
if not line:
continue
try:
rows.append(json.loads(line))
except json.JSONDecodeError:
continue
return rows
def parse_games(path):
"""Return [(opponent, won_bool)] for type=='game' rows, skipping junk."""
out = []
if not path.is_file():
print(f"[skip] games log not found: {path}")
return out
for line in path.read_text(errors="replace").splitlines():
try:
r = json.loads(line)
except json.JSONDecodeError:
continue
if r.get("type") != "game":
continue
opp, win = r.get("opponent"), r.get("win")
if isinstance(opp, str) and isinstance(win, bool):
out.append((opp, win))
return out
def metric_col(rows, key, positive=False):
"""[(index, value)] for float-parseable rows; abs() applied; optional >0 filter."""
out = []
for i, r in enumerate(rows):
try:
v = abs(float(r[key]))
except (KeyError, TypeError, ValueError):
continue
if positive and v <= 0:
continue
out.append((i, v))
return out
def rolling(vals, w=TREND_WINDOW):
out, s = [], 0.0
for i, v in enumerate(vals):
s += v
if i >= w:
s -= vals[i - w]
out.append(s / min(i + 1, w))
return out
def bucket_means(vals, n):
"""Split vals into <=n contiguous buckets of near-equal size; per-bucket mean."""
if not vals:
return []
n = min(n, len(vals))
k, rem = divmod(len(vals), n)
out, i = [], 0
for b in range(n):
size = k + (1 if b < rem else 0)
out.append(sum(vals[i:i + size]) / size)
i += size
return out
# ---------- tiny SVG helpers ----------
def esc(s):
return str(s).replace("&", "&amp;").replace("<", "&lt;").replace(">", "&gt;")
def write_svg(path, text):
"""Validate the finished SVG, then atomically swap it into place.
Readers never see partial output; an invalid render aborts without
touching the previous good file."""
try:
ET.fromstring(text)
except ET.ParseError as e:
print(f"[error] {path.name}: generated SVG invalid, keeping old file ({e})")
return False
tmp = path.with_name(path.name + ".tmp")
tmp.write_text(text)
os.replace(tmp, path)
return True
def polyline(pts, color, width=1.5, dash=None, opacity=1.0):
if len(pts) < 2:
return ""
d = f' stroke-dasharray="{dash}"' if dash else ""
p = " ".join(f"{x:.1f},{y:.1f}" for x, y in pts)
return (f'<polyline points="{p}" fill="none" stroke="{color}" '
f'stroke-width="{width}" opacity="{opacity}"{d}/>\n')
def dots(pts, color, r=2, opacity=0.25):
return "".join(f'<circle cx="{x:.1f}" cy="{y:.1f}" r="{r}" '
f'fill="{color}" opacity="{opacity}"/>\n' for x, y in pts)
def hgrid(x0, x1, ys):
return "".join(f'<line x1="{x0}" y1="{y:.1f}" x2="{x1}" y2="{y:.1f}" '
f'stroke="#dddddd"/>\n' for y in ys)
def axis(x0, y0, x1, y1, xt, yt, xlabel, ylabel, ylog=False):
"""Draw axes + ticks + labels. xt/yt are (value, px) tick lists."""
s = (f'<line x1="{x0}" y1="{y0}" x2="{x1}" y2="{y0}" stroke="black"/>\n'
f'<line x1="{x0}" y1="{y0}" x2="{x0}" y2="{y1}" stroke="black"/>\n')
for v, px in xt:
s += (f'<line x1="{px:.1f}" y1="{y0}" x2="{px:.1f}" y2="{y0 + 4}" stroke="black"/>\n'
f'<text x="{px:.1f}" y="{y0 + 17}" text-anchor="middle" font-size="11">'
f"{esc(v)}</text>\n")
for v, py in yt:
s += (f'<line x1="{x0 - 4}" y1="{py:.1f}" x2="{x0}" y2="{py:.1f}" stroke="black"/>\n'
f'<text x="{x0 - 7}" y="{py + 4:.1f}" text-anchor="end" font-size="11">'
f"{esc(v)}</text>\n")
s += (f'<text x="{(x0 + x1) // 2}" y="{y0 + 33}" text-anchor="middle" font-size="12">'
f"{esc(xlabel)}</text>\n"
f'<text x="16" y="{(y0 + y1) // 2}" text-anchor="middle" font-size="12" '
f'transform="rotate(-90 16 {(y0 + y1) // 2})">{esc(ylabel)}</text>\n')
return s
def ticks_linear(vmin, vmax, p0, p1, n=5, fmt="{:.0f}"):
return [(fmt.format(vmin + (vmax - vmin) * i / (n - 1)),
p0 + (p1 - p0) * i / (n - 1)) for i in range(n)]
def ticks_log(vmin, vmax, p0, p1, n=5):
vals = [10 ** (vmin + (vmax - vmin) * i / (n - 1)) for i in range(n)]
return [("{:.3g}".format(v), p0 + (p1 - p0) * i / (n - 1)) for i, v in enumerate(vals)]
def legend(items, x, y):
"""items: [(color, label)]"""
s = ""
for i, (color, label) in enumerate(items):
yy = y + i * 18
s += (f'<line x1="{x}" y1="{yy}" x2="{x + 28}" y2="{yy}" stroke="{color}" '
f'stroke-width="3"/>\n'
f'<text x="{x + 34}" y="{yy + 4}" font-size="12">{esc(label)}</text>\n')
return s
def map_fn(p0, p1, vmin, vmax, log=False):
def f(v):
t = (math.log10(v) - vmin) / (vmax - vmin) if log else (v - vmin) / (vmax - vmin)
return p0 + max(0.0, min(1.0, t)) * (p1 - p0)
return f
def guide_block(w, h, lines):
"""Plain-English "How to read" band at the bottom of the canvas."""
y = H - GUIDE_H
s = f'<rect x="0" y="{y}" width="{w}" height="{GUIDE_H}" fill="#f2f2f2"/>\n'
s += (f'<text x="16" y="{y + 21}" font-size="15" font-weight="bold">'
f"How to read</text>\n"
f'<text x="{w - 16}" y="{y + 20}" text-anchor="end" font-size="11" '
f'fill="#666">Regenerate anytime: python3 tools/plot_progress.py</text>\n')
for i, ln in enumerate(lines):
s += f'<text x="16" y="{y + 43 + i * 17}" font-size="13">{esc(ln)}</text>\n'
return s
def header(x, y, title, sub=None):
s = (f'<text x="{x}" y="{y}" font-size="15" font-weight="bold">'
f"{esc(title)}</text>\n")
if sub:
s += (f'<text x="{x}" y="{y + 15}" font-size="11" fill="#555">'
f"{esc(sub)}</text>\n")
return s
# ---------- panels ----------
def panel_test_matches(series, geo):
_, pt, pb, x0, x1 = geo
s = ""
if not series:
s += f'<text x="{x0 + 10}" y="{pt + 40}" font-size="12" fill="#a00">' \
"no eval lines found</text>\n"
return
xmax = max(max(len(v) for v in series.values()), 2)
xm, ym = map_fn(x0, x1, 1, xmax), map_fn(pb, pt, 0, 100)
s += hgrid(x0, x1, [ym(v) for v in range(0, 101, 20)])
step = max(1, (xmax // 8 // 10) * 10)
xt = [(str(v), xm(v)) for v in range(step, xmax + 1, step)] or [("1", xm(1))]
s += axis(x0, pb, x1, pt, xt, ticks_linear(0, 100, pb, pt, n=6),
"test match number (each opponent)", "win rate (%)")
items = []
for name in ("Corners", "Crazy", "Target"):
vals = series.get(name, [])
if not vals:
continue
c = COLORS[name]
s += dots([(xm(i + 1), ym(v)) for i, v in enumerate(vals)], c)
s += polyline([(xm(i + 1), ym(v))
for i, v in enumerate(rolling(vals))], c, 3.5)
items.append((c, f"{name} - {len(vals)} evals"))
s += legend(items, x0 + 12, pb + 52)
return s
def panel_real_fights(games, geo):
_, pt, pb, x0, x1 = geo
s = ""
if not games:
s += f'<text x="{x0 + 10}" y="{pt + 40}" font-size="12" fill="#a00">' \
"no game rows found</text>\n"
return
edges = list(range(0, len(games) + 1, GAME_BUCKET))
if edges[-1] != len(games):
edges.append(len(games))
buckets = list(zip(edges[:-1], edges[1:]))
xm, ym = map_fn(x0, x1, 1, max(len(games), 2)), map_fn(pb, pt, 0, 100)
s += hgrid(x0, x1, [ym(v) for v in range(0, 101, 20)])
step = max(GAME_BUCKET, GAME_BUCKET * (len(games) // GAME_BUCKET // 8 + 1))
xt = [(str(v), xm(v)) for v in range(step, len(games) + 1, step)]
s += axis(x0, pb, x1, pt, xt, ticks_linear(0, 100, pb, pt, n=6),
f"game number ({GAME_BUCKET}-game buckets)", "win %")
opponents = []
for opp, _ in games:
if opp not in opponents:
opponents.append(opp)
items = []
for name in opponents:
c = COLORS.get(name, "#7f7f7f")
by_b = {}
for bi, (lo, hi) in enumerate(buckets):
sub = [w for o, w in games[lo:hi] if o == name]
if sub:
by_b[bi] = 100.0 * sum(sub) / len(sub)
pts = []
segs, prev = [], None
for bi in sorted(by_b):
if prev is not None and bi != prev + 1:
segs.append(pts)
pts = []
center = (buckets[bi][0] + buckets[bi][1]) / 2
pts.append((xm(center), ym(by_b[bi])))
prev = bi
if len(pts) >= 2:
segs.append(pts)
for seg in segs:
s += polyline(seg, c, 3.5)
n_played = sum(1 for o, _ in games if o == name)
items.append((c, f"{name} ({n_played} games)"))
s += legend(items, x0 + 12, pb + 52)
return s
def panel_losses(rows, geo):
_, pt, pb, x0, x1 = geo
s = ""
critic = metric_col(rows, "critic_loss", positive=True)
actor = metric_col(rows, "actor_loss") # abs() applied; sign dropped
actor = [(i, v) for i, v in actor if v > 0]
if not (critic or actor):
s += f'<text x="{x0 + 10}" y="{pt + 40}" font-size="12" fill="#a00">' \
"no valid loss points</text>\n"
return
allv = [v for _, v in critic + actor]
lo, hi = math.floor(math.log10(min(allv))), math.ceil(math.log10(max(allv)))
if lo == hi:
hi = lo + 1
n = len(rows)
xm = lambda i: x0 + (x1 - x0) * i / max(n - 1, 1)
ym = map_fn(pb, pt, lo, hi, log=True)
s += hgrid(x0, x1, [ym(10 ** e) for e in range(lo, hi + 1)])
s += axis(x0, pb, x1, pt, ticks_linear(1, n, x0, x1, n=5),
ticks_log(lo, hi, pb, pt), "metric line number", "loss (log)")
s += polyline([(xm(i), ym(v)) for i, v in critic], CRITIC_C, 1.8)
s += polyline([(xm(i), ym(v)) for i, v in actor], ACTOR_C, 1.8)
s += legend([(CRITIC_C, "critic_loss"), (ACTOR_C, "|actor_loss|")],
x0 + 12, pb + 52)
return s
def panel_alpha(rows, geo):
_, pt, pb, x0, x1 = geo
s = ""
alpha = metric_col(rows, "alpha")
if not alpha:
s += f'<text x="{x0 + 10}" y="{pt + 40}" font-size="12" fill="#a00">' \
"no alpha points</text>\n"
return
n = len(rows)
hi = max(1.0, max(v for _, v in alpha))
xm = lambda i: x0 + (x1 - x0) * i / max(n - 1, 1)
ym = map_fn(pb, pt, 0, hi)
s += hgrid(x0, x1, [ym(v) for v in
[hi * k / 4 for k in range(5)]])
s += axis(x0, pb, x1, pt, ticks_linear(1, n, x0, x1, n=5),
ticks_linear(0, hi, pb, pt, n=5, fmt="{:.3g}"),
"metric line number", "alpha")
s += polyline([(xm(i), ym(v)) for i, v in alpha], "#9467bd", 1.8)
return s
def panel_throughput(rows, geo):
_, pt, pb, x0, x1 = geo
s = ""
eps = []
for r in rows:
try:
eps.append(float(r["epoch"]))
except (KeyError, TypeError, ValueError):
continue
rates = [] # games/hour per inter-row interval
for a, b in zip(eps, eps[1:]):
dt = b - a
if dt > 0:
rates.append(3600.0 * GAMES_PER_ROW / dt)
if not rates:
s += f'<text x="{x0 + 10}" y="{pt + 40}" font-size="12" fill="#a00">' \
"no usable epoch timestamps</text>\n"
return
bm = bucket_means(rates, RATE_BUCKET)
xm = map_fn(x0, x1, 1, len(rates))
ymax = max(max(rates), max(bm)) * 1.1
ym = map_fn(pb, pt, 0, ymax)
s += hgrid(x0, x1, [ym(ymax * k / 4) for k in range(5)])
step = max(1, len(rates) // 10)
xt = [(str(v), xm(v)) for v in range(step, len(rates) + 1, step)]
s += axis(x0, pb, x1, pt, xt, ticks_linear(0, ymax, pb, pt, n=5, fmt="{:.0f}"),
f"chunk interval ({GAMES_PER_ROW}-game chunks)", "games / hour")
s += dots([(xm(i + 1), ym(v)) for i, v in enumerate(rates)], "#7f7f7f", r=1.6)
if len(bm) >= 2:
ctr = [xm(round((i + 0.5) * len(rates) / len(bm))) for i in range(len(bm))]
s += polyline(list(zip(ctr, map(ym, bm))), "#2ca02c", 3.5)
return s
# ---------- assembly ----------
def build_dashboard(campaign, metrics_f, games_f, out):
series = parse_eval_series(campaign)
print("[info] evals parsed: " +
(", ".join(f"{k}={len(v)}" for k, v in sorted(series.items())) or "(none)"))
rows = parse_metrics(metrics_f)
print(f"[info] metric rows parsed: {len(rows)}")
games = parse_games(games_f)
print(f"[info] game rows parsed: {len(games)}")
s = (f'<svg xmlns="http://www.w3.org/2000/svg" width="{W}" height="{H}" '
f'viewBox="0 0 {W} {H}" font-family="sans-serif">\n'
f'<rect width="{W}" height="{H}" fill="white"/>\n'
f'<text x="{W // 2}" y="32" text-anchor="middle" font-size="21" '
f'font-weight="bold">SAC-LSTM campaign dashboard - live run (current only)</text>\n'
f'<text x="{W // 2}" y="56" text-anchor="middle" font-size="12" fill="#555">'
f'generated {datetime.now():%Y-%m-%d %H:%M:%S} - auto-reloads every 60 s '
f'(open this file in Chrome)</text>\n')
drawers = [
(ROWS[1], PANEL_TITLES[0],
"raw dots = single test matches, thick = rolling-mean-%d" % TREND_WINDOW,
lambda: panel_test_matches(series, ROWS[1])),
(ROWS[2], PANEL_TITLES[1],
"training_log.jsonl only - learning in REAL battles, not tests",
lambda: panel_real_fights(games, ROWS[2])),
(ROWS[3], PANEL_TITLES[2],
"training_metrics.jsonl - big early spikes are normal",
lambda: panel_losses(rows, ROWS[3])),
(ROWS[4], PANEL_TITLES[3],
"training_metrics.jsonl - high = exploring, low = exploiting",
lambda: panel_alpha(rows, ROWS[4])),
(ROWS[5], PANEL_TITLES[4],
"method: training_metrics.jsonl 'epoch' deltas (training_log.jsonl has "
"no timestamps); 1 row = one 10-game chunk",
lambda: panel_throughput(rows, ROWS[5])),
]
for geo, title, sub, drawer in drawers:
s += header(geo[3], geo[0], title, sub)
s += drawer()
s += guide_block(W, H, GUIDE_LINES)
s += RELOAD_JS + "\n"
s += "</svg>\n"
return write_svg(out, s)
# ---------- selftest ----------
def selftest():
with tempfile.TemporaryDirectory() as td:
td = Path(td)
(td / "log").write_text(
">>> [eval] win rate: 3/10 (30%) vs Corners\n"
"garbage line\n"
">>> [eval] win rate: 7/10 (70%) vs Crazy\n"
">>> [eval] win rate: broken\n"
">>> [eval] win rate: 5/10 (50%) vs Corners\n"
">>> [eval] win rate: 4/10 (40%) vs Crazy\n")
ser = parse_eval_series(td / "log")
assert ser == {"Corners": [30.0, 50.0], "Crazy": [70.0, 40.0]}, ser
assert rolling([10] * 25, 20)[-1] == 10.0
assert rolling([1, 2, 3], 20) == [1.0, 1.5, 2.0]
assert len(bucket_means(list(range(1287)), RATE_BUCKET)) == RATE_BUCKET
bm = bucket_means([0, 10], RATE_BUCKET)
assert bm == [0.0, 10.0], bm # fewer points than buckets -> no empty buckets
(td / "m.jsonl").write_text(
'{"epoch": 1000.0, "critic_loss": 10, "actor_loss": -2, "alpha": 0.5}\n'
"not json\n"
'{"epoch": 1060.0, "critic_loss": 100, "actor_loss": -4, "alpha": 0.25}\n'
'{"epoch": 1090.0, "critic_loss": 50, "actor_loss": 3, "alpha": 0.2}\n')
rows = parse_metrics(td / "m.jsonl")
assert len(rows) == 3 and rows[1]["critic_loss"] == 100
glines = []
for i in range(150): # 2 full GAME_BUCKETs, both opponents in both
glines.append(json.dumps(
{"type": "game", "round": i % 10 + 1, "ticks": 100,
"score": i % 3, "total_score": i, "win": i % 3 == 0,
"opponent": ("Corners", "Crazy")[i % 2]}))
(td / "g.jsonl").write_text("\n".join(glines) + "\n")
games = parse_games(td / "g.jsonl")
assert len(games) == 150 and games[0] == ("Corners", True)
assert games[-1] == ("Crazy", False) # i=149: odd -> Crazy; 149%3!=0 -> loss
assert metric_col(rows, "actor_loss") == [(0, 2.0), (1, 4.0), (2, 3.0)]
dash = td / "dash.svg"
assert build_dashboard(td / "log", td / "m.jsonl", td / "g.jsonl", dash)
text = dash.read_text()
ET.fromstring(text) # whole doc must parse -> closing tag present
assert RELOAD_JS in text, "auto-reload script missing"
for t in PANEL_TITLES:
assert t in text, f"panel title missing: {t}"
assert text.count(PANEL_TITLES[0]) == 1
assert "How to read" in text, "reading guide missing"
assert 'width="1400"' in text and 'height="2000"' in text
# circles: 4 eval dots (panel 1) + 2 throughput rate dots (panel 5)
assert text.count("<circle") == 6, text.count("<circle")
# 2 trends + 2 real-fight opp lines + 2 losses + 1 alpha + 1 throughput
assert text.count("<polyline") == 8, text.count("<polyline")
# orientation guard: a known rising series (10% -> 90%) rendered through
# the FULL build path must plot upward (smaller SVG y) and forward in
# time (larger x). Fails loudly if axis mapping is ever inverted again.
ori_log = td / "ori.log"
ori_log.write_text(
">>> [eval] win rate: 1/10 (10%) vs Corners\n"
">>> [eval] win rate: 9/10 (90%) vs Corners\n")
ori_dash = td / "ori.svg"
assert build_dashboard(ori_log, td / "m.jsonl", td / "g.jsonl", ori_dash)
m = re.search(r'<polyline points="([^"]+)"', ori_dash.read_text())
pts = [tuple(map(float, p.split(","))) for p in m.group(1).split()]
assert len(pts) == 2, pts
(xa, ya), (xb, yb) = pts
assert yb < ya, f"y-axis inverted: win % rose 10->90 but ink moved down ({ya} -> {yb})"
assert xb > xa, f"x-axis reversed: newer eval plotted left ({xa} -> {xb})"
print("selftest OK")
def main():
if "--selftest" in sys.argv:
selftest()
return
args = [a for a in sys.argv[1:] if not a.startswith("-")]
campaign = Path(args[0]) if len(args) > 0 else ROOT / "campaign_v2_stdout.log"
metrics = Path(args[1]) if len(args) > 1 else ROOT / "training_metrics.jsonl"
games = Path(args[2]) if len(args) > 2 else ROOT / "training_log.jsonl"
outdir = Path(args[3]) if len(args) > 3 else ROOT / "docs"
outdir.mkdir(parents=True, exist_ok=True)
try:
ok = build_dashboard(campaign, metrics, games, outdir / "campaign_dashboard.svg")
except Exception as e:
print(f"[error] dashboard build failed: {e}")
ok = False
if ok:
print(f"[done] dashboard written to {outdir / 'campaign_dashboard.svg'}")
else:
print("[error] dashboard NOT updated - check paths above")
sys.exit(1)
if __name__ == "__main__":
main()
+10
View File
@@ -0,0 +1,10 @@
#!/bin/sh
# Keep docs/campaign_dashboard.svg fresh: regenerate every 60 s.
# Errors go to stderr and never exit the loop silently.
dir=$(dirname "$0")
while :; do
if ! python3 "$dir/plot_progress.py"; then
echo "[watch_dashboard] $(date '+%F %T') regeneration failed (see error above)" >&2
fi
sleep 60
done
@@ -0,0 +1,173 @@
# GA/ES Parameters for ~750-Weight Neuroevolution
Research for issue #65. Concrete parameter recommendations extracted from 13 papers in `docs/papers/neuroevolution/`.
## Context
Evo_Bot's TOPO_Gun: fixed-topology feedforward ANN, 745-1489 weights depending on layer sizes. Task is predicting enemy dodge behavior from a 30-tick sliding window, outputting a guess factor. Online evolution during Robocode matches.
---
## 1. Population Size
**Recommendation: 64-256, not 300.**
| Source | Network size | Population | Notes |
|--------|-------------|------------|-------|
| Uber Deep GA (Such et al. 2017) | 4M params (Atari), 167k (Humanoid) | 1,000 | Massive networks, distributed; overkill for ~750 weights |
| OpenAI ES (Salimans et al. 2017) | 1.7M params | 720-1,440 workers | NES-style, not a population GA |
| Canonical ES (Chrabaszcz et al. 2018) | 1.7M params | 798 (lambda), mu=50 | mu=50 selected as best across games |
| World Models CMA-ES (Ha & Schmidhuber 2018) | 867-1,088 params (controller) | 64 | CMA-ES with 16 evals per individual |
| Challenges paper (Muller & Glasmachers 2018) | 1,352-2,349 weights | default CMA-ES lambda | LM-MA-ES for ~1k-2k weights |
| Evolving Generalists (Triebold & Yaman 2023) | 246-728 weights | xNES default: 4+floor(3*ln(d)) | For d=745 -> ~24; for d=1489 -> ~26 |
| NRA (Le Clei & Bellec 2022) | 3-322 params (dynamic) | 8-512 | 256+elitism best for ~300-param tasks; 16 enough for <100 params |
| Playing Atari 6 Neurons (Cuccu et al. 2018) | ~3k connections, 6-18 neurons | xNES default | Small networks, 100 generations sufficient |
**Key finding:** For ~750 weights, CMA-ES default lambda = 4+floor(3*ln(745)) = ~24 is a starting floor. The World Models paper (867-1,088 params, closest to our size) used 64 with CMA-ES and solved CarRacing. NRA used 256+elitism for ~300-param Pendulum networks. For a simple truncation-selection GA (not CMA-ES), 100-200 is a reasonable population; 300 is slightly wasteful but not harmful if evaluation is cheap.
**Verdict: Start with 100-200 for a truncation GA. If using CMA-ES or xNES, use their defaults (~24-26). 300 is too large for the weight count but acceptable if per-evaluation cost is low (Robocode rounds are fast).**
---
## 2. Mutation Rate and Distribution
**Recommendation: Additive Gaussian on ALL weights, sigma=0.002-0.02, not 5% at sigma=0.1.**
| Source | Mutation scheme | Notes |
|--------|----------------|-------|
| Uber Deep GA (Such et al. 2017) | theta' = theta + sigma * epsilon, epsilon ~ N(0,I). Sigma determined empirically per task. | Mutates ALL weights every generation, not a fraction. No "mutation rate" — every weight gets noise. |
| OpenAI ES (Salimans et al. 2017) | sigma = fixed hyperparameter (not adapted). Perturbation on full parameter vector. | Full-vector Gaussian perturbation, sigma tuned. |
| Canonical ES (Chrabaszcz et al. 2018) | N(0, sigma^2) added to all params. Network init from N(0, 0.05). | sigma is the step-size, adapted or fixed. |
| World Models (Ha & Schmidhuber 2018) | CMA-ES adapts sigma and full covariance matrix. | Self-adapting sigma — no manual sigma needed. |
| Challenges paper (Muller & Glasmachers 2018) | CSA (cumulative step-size adaptation) essential. Fixed sigma converges as slowly as random search. | Step-size adaptation is critical; fixed sigma is a known failure mode. |
| NRA (Le Clei & Bellec 2022) | N(0, 0.01) perturbation to all weights and biases. Top 50% selection. | sigma=0.01 for small networks. |
**Key finding:** No paper uses a "5% mutation rate" (mutating only 5% of weights). ALL papers mutate ALL weights simultaneously with small additive Gaussian noise. The "mutation rate" concept from traditional GAs (flip probability per gene) does not apply to real-valued neuroevolution. Instead, the noise magnitude (sigma) controls exploration.
For ~750 weights:
- NRA uses sigma=0.01 for networks up to ~300 params
- Uber GA uses sigma empirically per task (typical range 0.002-0.02 for Atari)
- CMA-ES/xNES adapt sigma automatically
**Verdict: Mutate ALL weights every generation. Use sigma=0.005-0.01 as starting point. If using CMA-ES, sigma self-adapts. The "5% of weights mutated" approach is non-standard and likely harmful — it under-explores the search space.**
---
## 3. Selection Pressure
**Recommendation: Top 10-50% (truncation) or top mu out of lambda. 20% is reasonable.**
| Source | Selection | Notes |
|--------|-----------|-------|
| Uber Deep GA (Such et al. 2017) | Truncation selection, top T individuals become parents. T not specified as percentage — varies. | Parents chosen uniformly at random from top T. |
| Canonical ES (Chrabaszcz et al. 2018) | Top mu=50 out of lambda=798 (~6%). Weighted mean of top mu. | mu=50 found optimal across games; tested mu in {10,20,50,100,200,400}. |
| NRA (Le Clei & Bellec 2022) | Top 50% duplicated, bottom 50% replaced. | Simple and effective for small populations. |
| Evolving Generalists (Triebold & Yaman 2023) | xNES default selection. | NES uses weighted rank-based update. |
**Key finding:** Selection pressure varies widely. Canonical ES uses ~6% (mu=50 out of 798). NRA uses 50%. Standard CMA-ES uses mu = lambda/2 (50%). Top 20% is in the middle range and is fine for a truncation GA.
**Verdict: 20% is reasonable. For small populations (64-100), 50% (top half) may work better. For larger populations (200+), stricter selection (10-20%) is appropriate. The Canonical ES result suggests mu=50 works well regardless of lambda for Atari-scale problems.**
---
## 4. Crossover
**Recommendation: No crossover. Mutation-only.**
| Source | Crossover? | Notes |
|--------|-----------|-------|
| Uber Deep GA (Such et al. 2017) | **No crossover.** "Historically, GAs often involve crossover, but for simplicity we did not include it." | Explicitly dropped crossover for DNN weights. |
| OpenAI ES (Salimans et al. 2017) | No crossover. | ES-style: mean update, not recombination of individuals. |
| NRA (Le Clei & Bellec 2022) | **No crossover.** "stripping down many mechanisms popular in traditional evolutionary methods, like agent crossover and speciation" | Crossover explicitly excluded. |
| CMA-ES/xNES | No crossover in the traditional sense. | Weighted recombination of top individuals into distribution mean — not pairwise crossover. |
| NEAT (Stanley & Miikkulainen 2011) | Has crossover via innovation numbers. | But NEAT is for topology evolution, not fixed-topology weight-only GA. |
**Key finding:** Every modern neuroevolution paper that works with fixed-topology networks drops crossover. Fogel & Stayton (1994, cited by Such et al.) showed crossover is often ineffective for simulated evolutionary optimization. For ANN weight vectors, crossover tends to be destructive because individual weights are not independent genes — they form functional units (layers, pathways) where mixing two different solutions creates non-functional hybrids.
**Verdict: No crossover. Mutation-only. Uniform crossover of ANN weights is harmful — it breaks co-adapted weight configurations. If recombination is desired, use CMA-ES/xNES-style weighted mean of top solutions, which is mathematically sound.**
---
## 5. Elitism
**Recommendation: Yes, keep top 1 unchanged (single elite).**
| Source | Elitism? | Notes |
|--------|---------|-------|
| Uber Deep GA (Such et al. 2017) | **Yes, 1 elite.** "The Nth individual is an unmodified copy of the best individual from the previous generation." Additionally, top 10 re-evaluated 30 times to find the true elite. | Single elite with robust re-evaluation. |
| NRA (Le Clei & Bellec 2022) | **Yes, elitism tested and beneficial.** Population sizes labeled "(elite)" consistently outperform non-elite variants in all figures. | Elitism was the single most impactful improvement for small populations. |
| CMA-ES | Elitist variants exist (mu+lambda). Standard CMA-ES is (mu,lambda) — non-elitist. | Non-elitist CMA-ES relies on distribution adaptation, not individual survival. |
**Key finding:** For simple truncation GAs, elitism (keeping top 1) prevents regression and is universally recommended. The NRA paper shows that adding elitism to even a population of 16 dramatically improves results. Uber's Deep GA uses elitism with robust re-evaluation (30 episodes to confirm the elite).
**Verdict: Keep top 1 elite unchanged. In noisy evaluation environments (Robocode), re-evaluate the top few candidates multiple times to find the true elite, following Uber's approach.**
---
## 6. Generations to Convergence
**Recommendation: 100-1,500 generations for ~750 weights, depending on the algorithm.**
| Source | Network size | Generations | Notes |
|--------|-------------|-------------|-------|
| Uber Deep GA (Such et al. 2017) | 4M params | 348-1,834 gens (at 1k pop) | Many games: best-in-run found in 1-29 gens |
| World Models CMA-ES (Ha & Schmidhuber 2018) | 867 params | ~1,800 gens | CMA-ES, pop=64, solved CarRacing |
| NRA (Le Clei & Bellec 2022) | 3-322 params (dynamic) | 100-5,000 gens | Simple tasks: <300 gens. Complex (Ant/Humanoid): 5,000+ |
| Playing Atari 6 Neurons (Cuccu et al. 2018) | ~3k connections | 100 gens | Extremely tight budget, still achieved competitive results |
| Challenges paper (Muller & Glasmachers 2018) | 1,352-2,349 weights | 100k-300k evals | LM-MA-ES, ~150k evals for bipedal walker convergence |
| Evolving Generalists (Triebold & Yaman 2023) | 728 weights (Ant) | 5,000 gens max | xNES, some tasks solved in <100 gens |
**Key finding:** For ~750 weights with a simple truncation GA (pop=100), expect 200-500 generations for a well-tuned sigma. CMA-ES/xNES may converge faster in generations but each generation is more expensive. The Challenges paper warns that halving the distance to the optimum requires O(d) samples, so for d=750, expect ~750 evaluations per halving step.
For Evo_Bot's online evolution during matches: each Robocode round can evaluate one individual. With 35-round matches (typical), ~5 generations of pop=7 per match, or ~2 generations of pop=15. Convergence within a single match is unlikely; evolution must persist across matches via weight persistence.
**Verdict: Budget 500-2,000 generations. With pop=100, that's 50k-200k evaluations. Online evolution will need many matches to converge — weight persistence is essential.**
---
## 7. CMA-ES vs Simple GA vs Tournament Selection
**Recommendation: CMA-ES or xNES for ~750 weights. Simple GA as a simpler fallback.**
| Algorithm | Sweet spot | Pros | Cons | Source |
|-----------|-----------|------|------|--------|
| **CMA-ES** | d <= 1,000 (ideal), up to ~2,000 (practical) | Self-adapts sigma and covariance, best convergence rate, handles ill-conditioned landscapes | O(d^2) memory/time per generation, needs O(d^2) evals for full covariance learning | Muller & Glasmachers 2018, Ha & Schmidhuber 2018 |
| **LM-MA-ES** | d = 1,000-10,000 | O(d) complexity, adapts fastest-evolving subspace, strong on ~2k weights | More complex to implement | Muller & Glasmachers 2018 |
| **xNES** | d <= 1,000 | Natural gradient, self-adapting, elegant | Similar scaling limits to CMA-ES | Cuccu et al. 2018, Triebold & Yaman 2023 |
| **Simple truncation GA** | Any d | Dead simple, trivially parallel, no internal state beyond population | Needs manual sigma tuning, no adaptation, converges slowly | Such et al. 2017, Le Clei & Bellec 2022 |
| **Canonical (mu,lambda)-ES** | Any d | Step-size adaptation via CSA, simple | mu tuning matters; mu=50 worked well in Chrabaszcz 2018 | Chrabaszcz et al. 2018 |
| **Tournament selection** | Traditional GA context | Tunable selection pressure | No advantage over truncation for ANN weights | Not specifically tested in any of the 13 papers |
**Key finding at d=750:** CMA-ES is in its sweet spot. The World Models paper (Ha & Schmidhuber 2018) used CMA-ES with pop=64 on 867-1,088 params and solved complex control tasks. The Evolving Generalists paper (Triebold & Yaman 2023) used xNES on 728 weights (Ant controller) with default population sizes. The Challenges paper (Muller & Glasmachers 2018) explicitly shows CMA-ES and LM-MA-ES outperforming simple ES on problems with 769-2,738 weights.
However, CMA-ES requires O(d^2) = O(560k) memory for the covariance matrix at d=750. This is manageable but not trivial for an online Robocode bot. A simpler option is a (mu,lambda)-ES with CSA for step-size adaptation.
**Verdict: CMA-ES or xNES is the best fit for 750 weights. If implementation complexity is a concern, a truncation GA with adaptive sigma (or even fixed sigma=0.005) is the pragmatic choice. Tournament selection offers no advantage.**
---
## Summary: Recommended Parameters for Evo_Bot TOPO_Gun
| Parameter | Current assumption | Recommendation | Rationale |
|-----------|-------------------|----------------|-----------|
| Population | 300 | 64-200 | 300 is oversized for ~750 weights; 64 (CMA-ES) to 200 (truncation GA) |
| Mutation | 5% of weights, gaussian sigma=0.1 | ALL weights, sigma=0.005-0.01 | Every paper mutates all weights. sigma=0.1 is too large. |
| Selection | Top 20% | Top 20-50% | 20% is fine; 50% if pop is small |
| Crossover | Uniform | None | Uniformly dropped in all modern neuroevolution papers |
| Elitism | Not specified | Top 1, re-evaluated | Single elite prevents regression; re-evaluate to handle noise |
| Algorithm | Simple GA | CMA-ES or truncation GA + CSA | CMA-ES is in its sweet spot at d=750; simple GA works but converges slower |
| Generations | Not specified | 500-2,000 | Online evolution needs many matches for convergence |
---
## Sources
1. Such et al. 2017 — "Deep Neuroevolution: Genetic Algorithms are a Competitive Alternative" (Uber AI Labs)
2. Salimans et al. 2017 — "Evolution Strategies as a Scalable Alternative to Reinforcement Learning" (OpenAI)
3. Chrabaszcz et al. 2018 — "Back to Basics: Benchmarking Canonical Evolution Strategies for Playing Atari"
4. Muller & Glasmachers 2018 — "Challenges in High-dimensional Reinforcement Learning with Evolution Strategies"
5. Ha & Schmidhuber 2018 — "Recurrent World Models Facilitate Policy Evolution" (World Models)
6. Cuccu et al. 2018 — "Playing Atari with Six Neurons"
7. Le Clei & Bellec 2022 — "Neuroevolution of Recurrent Architectures on Control Tasks"
8. Triebold & Yaman 2023 — "Evolving Generalist Controllers to Handle a Wide Range of Morphological Variations"
9. Stanley & Miikkulainen 2011 — "Competitive Coevolution through Evolutionary Complexification" (NEAT)
@@ -209,6 +209,8 @@ proc runReceiveLoop*(ws: SyncWebSocket; info: BotInfo; secret: string; serverUrl
of "SkippedTurnEvent": of "SkippedTurnEvent":
let e = node.to(SkippedTurnEvent) let e = node.to(SkippedTurnEvent)
gBot.onSkippedTurn(e) gBot.onSkippedTurn(e)
of "BotListUpdate":
updateBotNames(node)
else: else:
discard # unknown message type — ignore discard # unknown message type — ignore
except Exception as e: except Exception as e:
@@ -13,7 +13,7 @@
## tickChan main → bot (true = new tick ready; false = stop) ## tickChan main → bot (true = new tick ready; false = stop)
## intentChan bot → sender (JSON string to send to server; "" = stop) ## intentChan bot → sender (JSON string to send to server; "" = stop)
import std/[json, locks, math, os, posix, syncio] import std/[json, locks, math, os, posix, syncio, tables]
import ./constants import ./constants
import ./schemas import ./schemas
import ./color import ./color
@@ -68,6 +68,7 @@ var gGameSetup {.guard: gLock.}: GameSetup
var gTeammateIds{.guard: gLock.}: seq[int] var gTeammateIds{.guard: gLock.}: seq[int]
var gVariant {.guard: gLock.}: string var gVariant {.guard: gLock.}: string
var gServerVersion {.guard: gLock.}: string var gServerVersion {.guard: gLock.}: string
var gBotNames {.guard: gLock.}: Table[int, string]
# Events handed main -> bot thread. Channel move, NOT a locked shared seq: # Events handed main -> bot thread. Channel move, NOT a locked shared seq:
# the old locked seq[BotEvent] copied GC'd payloads (strings/teamMessages) # the old locked seq[BotEvent] copied GC'd payloads (strings/teamMessages)
# across threads on every tick -> refcount churn under ORC + --threads:on -> # across threads on every tick -> refcount churn under ORC + --threads:on ->
@@ -166,6 +167,28 @@ proc getTurnRemaining*(): float = gTurnRemaining
proc getGunTurnRemaining*(): float = gGunTurnRemaining proc getGunTurnRemaining*(): float = gGunTurnRemaining
proc getRadarTurnRemaining*(): float = gRadarTurnRemaining proc getRadarTurnRemaining*(): float = gRadarTurnRemaining
proc getBotName*(id: int): string =
## Lookup bot name by id from the last BotListUpdate. Returns "" if unknown.
withLock(gLock): result = gBotNames.getOrDefault(id, "")
proc updateBotNames*(node: JsonNode) =
## Update the id → name table from a BotListUpdate message (full replacement).
withLock(gLock):
gBotNames.clear()
if node.isNil: return
let botsNode = node{"bots"}
if botsNode.isNil or botsNode.kind != JArray: return
for b in botsNode:
let name = b{"name"}.getStr("")
if name.len == 0: continue
var id = -1
if not b{"id"}.isNil:
id = b{"id"}.getInt(-1)
if id == -1 and not b{"botId"}.isNil:
id = b{"botId"}.getInt(-1)
if id == -1: continue
gBotNames[id] = name
proc getMaxSpeed*(): float = gMaxSpeed proc getMaxSpeed*(): float = gMaxSpeed
proc getMaxTurnRate*(): float = gMaxTurnRate proc getMaxTurnRate*(): float = gMaxTurnRate
proc getMaxGunTurnRate*(): float = gMaxGunTurnRate proc getMaxGunTurnRate*(): float = gMaxGunTurnRate
@@ -862,6 +885,8 @@ proc initGlobals*() =
gEventChan.open(8) gEventChan.open(8)
initLock(gLock) initLock(gLock)
gEventQueue = initEventQueue() gEventQueue = initEventQueue()
withLock(gLock):
gBotNames = initTable[int, string]()
# Debug log is opt-in: it is written every tick from two threads, so leaving # Debug log is opt-in: it is written every tick from two threads, so leaving
# it on by default is a disk hog and an I/O stall source. Enable with # it on by default is a disk hog and an I/O stall source. Enable with
# PPOB_DEBUG_LOG=1 to debug; starts fresh (truncated) each run. # PPOB_DEBUG_LOG=1 to debug; starts fresh (truncated) each run.
+10 -8
View File
@@ -18,9 +18,10 @@ import java.util.logging.Logger;
* to the same log file via PPOB_LOG_FILE — the shell wrapper stitches them. * to the same log file via PPOB_LOG_FILE — the shell wrapper stitches them.
* *
* Usage (env vars): * Usage (env vars):
* PPO_BOT_DIR — path to PPO_Bot dir * PPO_BOT_DIR — path to the bot dir (any Tank Royale bot)
* SAMPLE_BOTS_DIR — path to sample bots archive * SAMPLE_BOTS_DIR — path to sample bots archive
* PPOB_LOG_FILE — path to training_log.jsonl (appended) * PPOB_LOG_FILE — path to training_log.jsonl (appended)
* BOT_NAME — bot name to match in round results (default: PPO_Bot)
* TRAINING_OPPONENT — opponent bot name (default: Target) * TRAINING_OPPONENT — opponent bot name (default: Target)
* TRAINING_ROUNDS — number of rounds to run (CLI arg or env var) * TRAINING_ROUNDS — number of rounds to run (CLI arg or env var)
* *
@@ -39,8 +40,9 @@ public class RunTraining {
: System.getenv().getOrDefault("TRAINING_OPPONENT", "Target"); : System.getenv().getOrDefault("TRAINING_OPPONENT", "Target");
int totalRounds = args.length > 1 ? Integer.parseInt(args[1]) int totalRounds = args.length > 1 ? Integer.parseInt(args[1])
: Integer.parseInt(System.getenv().getOrDefault("TRAINING_ROUNDS", "100")); : Integer.parseInt(System.getenv().getOrDefault("TRAINING_ROUNDS", "100"));
String botName = System.getenv().getOrDefault("BOT_NAME", "PPO_Bot");
System.out.printf("Training: PPO_Bot vs %s for %d rounds%n", opponent, totalRounds); System.out.printf("Training: %s vs %s for %d rounds%n", botName, opponent, totalRounds);
System.out.printf("Log: %s%n", logFile); System.out.printf("Log: %s%n", logFile);
// Dead-bot guard: the runner keeps listing a crashed PPO_Bot in the // Dead-bot guard: the runner keeps listing a crashed PPO_Bot in the
@@ -81,7 +83,7 @@ public class RunTraining {
boolean win = false; boolean win = false;
boolean found = false; boolean found = false;
for (var r : event.getResults()) { for (var r : event.getResults()) {
if (r.getName().equals("PPO_Bot")) { if (r.getName().equals(botName)) {
found = true; found = true;
totalScore = r.getTotalScore(); totalScore = r.getTotalScore();
win = r.getRank() == 1; win = r.getRank() == 1;
@@ -99,9 +101,9 @@ public class RunTraining {
frozenRounds[0]++; frozenRounds[0]++;
long frozenMs = System.currentTimeMillis() - lastAdvanceMs[0]; long frozenMs = System.currentTimeMillis() - lastAdvanceMs[0];
if (frozenRounds[0] >= 10 && frozenMs >= 10_000) { if (frozenRounds[0] >= 10 && frozenMs >= 10_000) {
System.err.printf("PPO_Bot round_counter frozen at %d for %d " System.err.printf("%s round_counter frozen at %d for %d "
+ "harness rounds / %.0fs — process dead, aborting for restart%n", + "harness rounds / %.0fs — process dead, aborting for restart%n",
ctr, frozenRounds[0], frozenMs / 1000.0); botName, ctr, frozenRounds[0], frozenMs / 1000.0);
System.exit(1); System.exit(1);
} }
} else { } else {
@@ -110,7 +112,7 @@ public class RunTraining {
lastAdvanceMs[0] = System.currentTimeMillis(); lastAdvanceMs[0] = System.currentTimeMillis();
} }
if (!found) { if (!found) {
System.err.println("PPO_Bot missing from round " + round System.err.println(botName + " missing from round " + round
+ " results — process died, aborting battle for restart"); + " results — process died, aborting battle for restart");
System.exit(1); System.exit(1);
} }
@@ -147,10 +149,10 @@ public class RunTraining {
endCounter = readCounter(counterPath); endCounter = readCounter(counterPath);
} }
if (endCounter < expectedEnd) { if (endCounter < expectedEnd) {
System.err.printf("PPO_Bot round_counter %d < expected %d (start+%d) at battle " System.err.printf("%s round_counter %d < expected %d (start+%d) at battle "
+ "end (waited 60s) — %d rounds never trained (corpse?), aborting for " + "end (waited 60s) — %d rounds never trained (corpse?), aborting for "
+ "restart%n", + "restart%n",
endCounter, expectedEnd, totalRounds, expectedEnd - endCounter); botName, endCounter, expectedEnd, totalRounds, expectedEnd - endCounter);
System.exit(1); System.exit(1);
} }
System.out.printf("Counter check passed: %d == expected %d%n", endCounter, expectedEnd); System.out.printf("Counter check passed: %d == expected %d%n", endCounter, expectedEnd);