Commit Graph

332 Commits

Author SHA1 Message Date
SirStone 6241c41e52 Merge branch 'worktree-agent-a4c06ae3' into research/goto-controller 2026-08-24 08:23:19 +02:00
SirStone 156b4ae8db Merge branch 'worktree-agent-af671e69' into research/goto-controller 2026-08-24 08:23:12 +02:00
SirStone eee48dec38 prototype: GA evolution spike -- sin(x) prediction validates pipeline (#66)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-08-24 08:14:36 +02:00
SirStone eb8faa7c99 prototype: GA evolution spike -- sin(x) prediction validates pipeline (#66)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-08-24 08:14:32 +02:00
SirStone 96065597c0 fix(divergence): guard reward-normalizer cold start; drop arctanh recovery
Welford variance-collapse divided by 1e-8 producing bit-exact +/−5e8 /
+−1.25e8 poisoned rewards into TD targets; guard skips normalization
until stats meaningful; tanh-inversion removal bounds log-prob path.

Fixes #60.
2026-08-24 01:03:04 +02:00
SirStone cb1bbc35dc docs(adr): Evo_Bot neuroevolution gun architecture (#63)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-08-24 00:02:47 +02:00
SirStone b7b10af811 feat: OscillatorBot sparring partner for GA gun testing (#64)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-08-24 00:01:38 +02:00
SirStone e455566c8e docs: CONTEXT.md — Evo_Bot domain vocabulary (#62)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-08-24 00:00:21 +02:00
SirStone eae6fc15a2 research(Evo_Bot): GA/ES parameter recommendations for ~750-weight neuroevolution (#65)
Extracted concrete numbers from 13 papers in docs/papers/neuroevolution/.
Key findings: use CMA-ES or mutation-only truncation GA, mutate ALL weights
(not 5%), sigma=0.005-0.01, pop=64-200, no crossover, single elite.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-08-23 23:32:10 +02:00
SirStone 81718e3a4c fix(SAC_LSTM_Bot): un-invert dashboard axes — orientation selftest added 2026-08-23 10:47:48 +02:00
SirStone 04c149ea28 refactor(SAC_LSTM_Bot): reconcile graph tooling — dashboard-only output 2026-08-23 10:30:16 +02:00
SirStone 533a146342 feat(SAC_LSTM_Bot): live campaign dashboard — five panels, auto-reload, run-1 comparison dropped 2026-08-23 10:25:51 +02:00
SirStone 03a1853e57 feat(SAC_LSTM_Bot): embedded plain-English reading guides in graphs 2026-08-23 10:11:44 +02:00
SirStone 002a568c7e fix(SAC_LSTM_Bot): atomic+validated SVG output for progress graphs 2026-08-23 10:06:05 +02:00
SirStone f3ac0888bb feat(SAC_LSTM_Bot): readable eval chart — trends primary, raw dots secondary, v1 comparison separated 2026-08-23 09:59:30 +02:00
SirStone 0b0933294a feat(SAC_LSTM_Bot): progress graph tooling + first graphs 2026-08-23 09:08:53 +02:00
SirStone 781e41595e docs(SAC_LSTM_Bot): simple-words story + dictionary for readability 2026-08-23 08:48:37 +02:00
SirStone 40e074e5a3 docs(SAC_LSTM_Bot): lever-5 trigger + attempt-3 chapter — MA-wipe root cause, LR_CRITIC=1e-4 decision, launch health (#59 #60) 2026-08-23 07:11:09 +02:00
SirStone 167bcc4ce5 fix(SAC_LSTM_Bot): MA history self-truncation — read before > redirect; campaign v2 attempt-3 prep (#60)
$(cat f) inside a command redirected to f saw the already-truncated file,
so every eval cycle wiped ma_history_*.txt back to one leading-space value
and degraded the composite best-gate to last-cycle mean. Read is hoisted
into its own statement; unquoted expansion + tail -n keeps exactly the last
MA_WINDOW values. Verified: fresh/5+/6+ cycle edges reproduce sliding window.
2026-08-23 07:02:36 +02:00
SirStone 1619b86f25 docs(SAC_LSTM_Bot): v2 restart saga, twin-freeze correction, pacing decisions 2026-08-22 21:26:24 +02:00
SirStone 19f34abf0c fix(SAC_LSTM_Bot): freeze mirror-twin via eval-mode gate in twin launcher
SACLSTM_EVAL_MODE=1 in SacTwin.sh suppresses all sendTrainingMsg traffic
(lever-4 gate), so the twin never trains — not even in-RAM within a battle.
Required now that the main bot's SACLSTM_SAVE_INTERVAL drops to 1 (v2 relaunch
after checkpoint-cadence diagnosis): without the gate the twin would persist
per-battle drift and stop being the frozen reproducible opponent #54 specifies.
2026-08-22 21:09:07 +02:00
SirStone bd58794b4c docs(SAC_LSTM_Bot): campaign v2 chapter — locked config, fresh-start archive, twin reseed, ceiling net, launch health evidence (#59) 2026-08-22 20:31:33 +02:00
SirStone 6fc01eb4e5 feat(SAC_LSTM_Bot): campaign v2 levers — aggression/anti-ram reward shaping + stability knob overrides (part 2)
Lever 2 (#59): x1.25 aggression mult on damage dealt, flat +0.5 hit bonus,
-3.0 per bot-bot collision (server deals RAM_DAMAGE=0.6 to both parties but
only notifies the hitter), escalating proximity deterrent below 12% arena
diagonal suppressed while dealing damage. Win/loss terminals unchanged and
dominant. All weights TUNABLE consts marked ponytail. SACLSTM_REWARD_DEBUG=1
env-gated reward_debug.log for calibration greps.

Lever 5 (#59): no code needed — SACLSTM_LR_ACTOR/LR_CRITIC/LR_ALPHA (3e-4)
and SACLSTM_TARGET_ENTROPY (-4.0) were already env-overridable in training.nim.

Smoke vs RamFire+Crazy (hidden=32, random init, isolated weights): 75 ram
penalties, 381 charge events, hit bonuses firing, 0 crashes, metrics JSONL
flowing. Tests: 8/8 suites green incl. new assert-level term math.
2026-08-22 19:44:58 +02:00
SirStone a07e5305f5 feat(SAC_LSTM_Bot): campaign v2 levers — loss metrics, eval-mode gate, eval rotation + MA gating (part 1)
Levers 3, 4, 1 of the #57 sign-off (execution order 3->4->1), tracked in #59.

- Lever 3 (#59): one JSONL line per trainPass in training_metrics.jsonl with
  exactly the scalars sacUpdate already exposes (SACMetrics: critic/actor/alpha
  losses + alpha, averaged per pass) plus epoch, buffer size (replay_buffer.len),
  cumulative steps and drained count. No trainer change needed.
- Lever 4 (#59): sendTrainingMsg drops all training input while SACLSTM_EVAL_MODE=1
  (existing #49 harness mechanism) — eval battles can neither pollute the replay
  buffer nor trigger gradient updates; one-time stderr notice at bot init.
- Lever 1 (#59): sac_train.sh evaluates every SAC_EVAL_OPPONENTS entry per cycle
  (results carry opponent name in eval_log.jsonl); best-gating now uses a
  composite = mean over opponents of the last-5-evals moving average per
  opponent. best_score.txt format change: float composite replaces the
  single-opponent integer win rate semantics (retired).
- Tests: metricsLine JSONL scalars + eval-mode suppression asserts.

Refs: #59, #57
2026-08-22 18:58:47 +02:00
SirStone 2619ba06fc docs(SAC_LSTM_Bot): campaign ending, results table, verdict, hygiene (#57) 2026-08-22 14:40:57 +02:00
SirStone f45e8f2717 fix(SAC_LSTM_Bot): startup sweep of stale .part checkpoint corpses (#57) 2026-08-22 14:40:57 +02:00
SirStone b7492f1080 docs(SAC_LSTM_Bot): notebook — score:60 anatomy, .part forensics+sweep, ceiling defused 2026-08-22 08:08:31 +02:00
SirStone f1962c7506 docs(SAC_LSTM_Bot): backfill shift-1 milestone entry + morning audit 2026-08-22 07:27:37 +02:00
SirStone 05929d2dbd docs(SAC_LSTM_Bot): notebook — shift 1, deferred-intervention milestone 2026-08-22 07:14:35 +02:00
SirStone 4b64bf18ac docs(SAC_LSTM_Bot): campaign-v1 launch record — save-check incident, health check, check-in procedure (#56) 2026-08-22 00:31:47 +02:00
SirStone 26536713ba fix(SAC_LSTM_Bot): checkpoint save check inside gradient-step loop (#56)
Launch finding during campaign-v1 verification: at production sizes
(hidden 256, ~1s/step, ~13s trainer CPU per ~40s chunk process) the
save check ran only between drain-burst passes, so stepCount never
crossed nextSave before the process died — zero checkpoints persisted
across entire runs (masked at #49/#54 smoke sizes where steps were
sub-millisecond). Check now fires mid-loop; with SAVE_INTERVAL<=5
(within the per-process step budget) every chunk persists its chain.
2026-08-22 00:21:43 +02:00
SirStone 2f49cb243f docs(SAC_LSTM_Bot): campaign-v1 notebook — locked config + story so far (#56) 2026-08-21 23:26:31 +02:00
SirStone 6a294ad7ad feat(SAC_LSTM_Bot): mirror-twin sparring partner + readiness check (#54)
- make_twin.sh: generates self-contained SacTwin dir in the sample-bots
  archive (own json/sh identity, own weights dir seeded from a frozen
  sac_best.zip copy, own round_counter) so RunTraining.java resolves it
  like any sample bot; re-running resets the twin to the frozen baseline.
- src/SAC_LSTM_Bot.nim: SACLSTM_BOT_JSON env overrides the baked-in bot
  json (loadBotInfo gives json total precedence, #49) so the same binary
  boots under the twin's name.
- sac_train.sh: chunk loop is a while, not for-over-seq — a crash on the
  FINAL chunk previously fell through ((chunk--);continue on an exhausted
  seq list) and exited 0 with budget incomplete; observed live vs SacTwin.

Readiness dry-run (#54): weighted pool Corners:1,SacTwin:3 picked the twin
in 3/4 chunks; all battles counter-checked; deterministic eval parsed;
main sac_best.zip/counter untouched by twin (twin counter advanced
independently); crash-restart proven end-to-end incl. final-chunk retry.
2026-08-21 23:19:27 +02:00
SirStone edf26aa45d fix(SAC_LSTM_Bot): enforce MaxHidden cap on SACLSTM_HIDDEN_SIZE (review of #48/#49) 2026-08-21 22:17:55 +02:00
SirStone df256b4d3e feat(SAC_LSTM_Bot): training harness (#49)
sac_train.sh orchestrates chunked self-play via tools/training_runner/
RunTraining.java: weighted opponent sampling per chunk, deterministic
eval (SACLSTM_EVAL_MODE=1) every N chunks with win-rate tracking, best
checkpoint (weights/sac_best.zip) by eval score, crash-restart loop on
the runner's liveness detection.

Supporting changes:
- integration.nim: opponentKey() keys the NewBattle buffer-clear rule on
  getBotName(id) with numeric-id fallback (#49 Q14 follow-up);
  bumpRoundCounter() emits the per-round liveness signal.
- SAC_LSTM_Bot.nim: onRoundEnded -> bumpRoundCounter().
- RunTraining.java: BOT_NAME env parameterizes result matching
  (default PPO_Bot, unchanged behavior for PPO).
- Launch packaging: root SAC_LSTM_Bot.json + .sh for the booter;
  src json name aligned to 'SAC_LSTM_Bot' so self-reported identity
  matches the booted identity (mismatch = runner connect timeout).
2026-08-21 21:51:27 +02:00
SirStone 7104645f5d chore(deps): update vendored tankroyale botapi to v1.0.1
Syncs libs/tankroyale_botapi with SirStone/robocode_tankroyale_botapi
v1.0.1 (extracted from tank-royale nim branch @ 03195a814). The local
SIGSEGV fixes (static event queue/SVG/intent buffers) were already
ported upstream in issue #24 — content is otherwise identical.

New capability (upstream #23): opponent name exposure for ticket #49.
- bot.nim: gBotNames id→name table + getBotName(id) / updateBotNames()
- umbrella module: dispatch BotListUpdate messages to updateBotNames()

Vendored layout and wiring unchanged (--path via config.nims, module
name stays tankroyale_botapi).
2026-08-21 21:03:17 +02:00
SirStone 32b71d9fc8 feat(SAC_LSTM_Bot): main bot integration (#48) 2026-08-21 20:19:05 +02:00
SirStone 62a6cc8ccf fix(SAC_LSTM_Bot): training review fixes — hidden state ordering, redundant forwards, actor grad clip (#47)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-21 00:16:47 +02:00
SirStone 415d4e3738 feat(SAC_LSTM_Bot): SAC training module (#47)
Implements sacUpdate with burn-in LSTM warm-up, twin-critic TD update,
actor reparameterization gradient, auto-alpha, and soft target update.
Manual backprop (linear + LSTM single-step, truncated BPTT). 11 new tests
all green; full regression suite (57+ tests) unaffected.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-21 00:11:35 +02:00
SirStone 717ef3ead8 Merge branch 'worktree-agent-a8622248' (ticket #46 weight persistence) 2026-08-20 23:57:53 +02:00
SirStone 54b8139b11 feat(SAC_LSTM_Bot): weight persistence module (#46)
Save/load all SAC-LSTM tensors (actor, 2 critics, 2 target critics,
alpha, Adam states) into a single .zip of .npy files. Atomic write
via temp path + rename. Adam types (AdamVar, SACAdamStates) defined
here for training.nim to use.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:57:13 +02:00
SirStone 55ef22ff8b Merge branch 'worktree-agent-a3ca3066' (ticket #45 replay buffer) 2026-08-20 23:41:23 +02:00
SirStone 5ea57bcae3 Merge branch 'worktree-agent-a393a8b9' (ticket #44 reward module) 2026-08-20 23:41:23 +02:00
SirStone cb33551621 feat(SAC_LSTM_Bot): replay buffer module (#45)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:40:17 +02:00
SirStone 23c65c9ac6 feat(SAC_LSTM_Bot): reward module (#44)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:39:45 +02:00
SirStone 4ee0d8272c feat(SAC_LSTM_Bot): action mapping module (#43)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:38:45 +02:00
SirStone add3e34926 feat(SAC_LSTM_Bot): LSTM network module (#41)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:38:17 +02:00
SirStone f130bf1254 feat(SAC_LSTM_Bot): state vector module (#42)
35-dim normalized tensor (GameState → buildState). No history window —
LSTM handles temporal context. Covers own-bot (7), enemy (7), derived (4),
walls (4), bullets (12), scan staleness (1). All tests pass.
2026-08-20 23:37:56 +02:00
SirStone a0a3840980 feat(SAC_LSTM_Bot): skeleton bot with radar lock and colors (#40)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-20 23:37:50 +02:00
SirStone 7d73d32c85 Merge branch 'worktree-agent-ac811b59' (ticket #39 radar lock) 2026-08-20 23:34:24 +02:00