[Map] Autonomous LLM-supervised PPO_Bot training pipeline #29

Closed
opened 2026-08-17 19:48:13 +02:00 by SirStone · 1 comment
Owner

Destination

A standalone CLI training tool (tools/training_runner/) that runs PPO_Bot against a chosen opponent for unlimited rounds with crash recovery, outputting structured training stats to files. Runnable, stoppable, and re-runnable from the command line — designed so an external LLM agent (or human) can monitor output, edit hyperparameters, and restart training without the tool knowing or caring who's driving it.

Notes

  • Domain: reinforcement learning (PPO) for Tank Royale bot combat
  • PPO_Bot is a Nim bot that trains itself between rounds; the runner just orchestrates battles
  • Tank Royale battles run via Java's embedded BattleRunner API
  • Existing 1-round runner at tools/battle_runner/ — keep it untouched for debugging
  • PPO_Bot auto-loads weights from weights/latest/ on startup (crash recovery baseline works)
  • Full hyperparameter inventory documented in map ticket discussions
  • Skills to consult: /grilling, /domain-modeling

Decisions so far

  • LLM supervisor agent design — out of scope; the runner is a dumb CLI tool, the LLM supervisor is a separate effort
  • Stall detection: hardcoded heuristics vs LLM judgment — no heuristics in the runner; it outputs rich stats, external observer judges; merged into #32
  • Structured output format for LLM consumption — JSON lines to training_log.jsonl, one object per round with game outcome + training health + hyperparam snapshot; no derived stats, reader computes those
  • [#35] Adam optimizer state (m/v tensors + t counters) and round counter persisted to disk alongside weights; restored on startup, cold-start if absent
  • Hyperparam delivery — flat env vars with PPOB_ prefix, stdlib getEnv/parseFloat, sensible defaults; no recompile risk
  • Build the training runner — tools/training_runner/ with shell wrapper (compile + crash-restart loop) and Java runner; JSONL stats output
  • Tunable parameter set — all 11 non-structural hyperparams exposed via PPOB_* env vars; logStd floor consolidated from 4 sites to 1

Not yet specified

  • Reward shaping knobs: whether the LLM should be able to tune reward function coefficients, or if that's too structural/dangerous to expose as a runtime knob
  • Multi-opponent curriculum: running against a sequence of increasingly difficult opponents automatically — out of scope for now (single opponent) but likely next after this map
  • Distributed training: running multiple battle instances in parallel for faster data collection
  • Metrics visualization: live dashboards, TensorBoard integration, or similar — for now, structured logs that the LLM reads directly

Out of scope

  • Multi-opponent curriculum training (future effort, single opponent first)
  • GUI/dashboard for human monitoring (LLM is the monitor)
  • Changes to network architecture (STATE_DIM, ACTION_DIM, layer count)
  • LLM supervisor agent (#36) — separate effort; this tool just needs to be CLI-friendly and output structured stats
  • Stall detection heuristics — external concern, not baked into the runner
## Destination A standalone CLI training tool (`tools/training_runner/`) that runs PPO_Bot against a chosen opponent for unlimited rounds with crash recovery, outputting structured training stats to files. Runnable, stoppable, and re-runnable from the command line — designed so an external LLM agent (or human) can monitor output, edit hyperparameters, and restart training without the tool knowing or caring who's driving it. ## Notes - Domain: reinforcement learning (PPO) for Tank Royale bot combat - PPO_Bot is a Nim bot that trains itself between rounds; the runner just orchestrates battles - Tank Royale battles run via Java's embedded `BattleRunner` API - Existing 1-round runner at `tools/battle_runner/` — keep it untouched for debugging - PPO_Bot auto-loads weights from `weights/latest/` on startup (crash recovery baseline works) - Full hyperparameter inventory documented in map ticket discussions - Skills to consult: `/grilling`, `/domain-modeling` ## Decisions so far - [LLM supervisor agent design](https://git.fossellini.top/SirStone/SirRoboGarage/issues/36) — out of scope; the runner is a dumb CLI tool, the LLM supervisor is a separate effort - [Stall detection: hardcoded heuristics vs LLM judgment](https://git.fossellini.top/SirStone/SirRoboGarage/issues/33) — no heuristics in the runner; it outputs rich stats, external observer judges; merged into #32 - [Structured output format for LLM consumption](https://git.fossellini.top/SirStone/SirRoboGarage/issues/32) — JSON lines to `training_log.jsonl`, one object per round with game outcome + training health + hyperparam snapshot; no derived stats, reader computes those - [#35] Adam optimizer state (m/v tensors + t counters) and round counter persisted to disk alongside weights; restored on startup, cold-start if absent - [Hyperparam delivery](https://git.fossellini.top/SirStone/SirRoboGarage/issues/30) — flat env vars with PPOB_ prefix, stdlib getEnv/parseFloat, sensible defaults; no recompile risk - [Build the training runner](https://git.fossellini.top/SirStone/SirRoboGarage/issues/34) — tools/training_runner/ with shell wrapper (compile + crash-restart loop) and Java runner; JSONL stats output - [Tunable parameter set](https://git.fossellini.top/SirStone/SirRoboGarage/issues/31) — all 11 non-structural hyperparams exposed via PPOB_* env vars; logStd floor consolidated from 4 sites to 1 ## Not yet specified - **Reward shaping knobs**: whether the LLM should be able to tune reward function coefficients, or if that's too structural/dangerous to expose as a runtime knob - **Multi-opponent curriculum**: running against a sequence of increasingly difficult opponents automatically — out of scope for now (single opponent) but likely next after this map - **Distributed training**: running multiple battle instances in parallel for faster data collection - **Metrics visualization**: live dashboards, TensorBoard integration, or similar — for now, structured logs that the LLM reads directly ## Out of scope - Multi-opponent curriculum training (future effort, single opponent first) - GUI/dashboard for human monitoring (LLM is the monitor) - Changes to network architecture (STATE_DIM, ACTION_DIM, layer count) - LLM supervisor agent (#36) — separate effort; this tool just needs to be CLI-friendly and output structured stats - Stall detection heuristics — external concern, not baked into the runner
SirStone added the wayfinder:map label 2026-08-17 19:48:13 +02:00
Author
Owner

Map complete

All tickets resolved. The destination is reached: tools/training_runner/ is a dumb CLI tool that runs PPO_Bot against a chosen opponent for unlimited rounds with crash recovery, outputting structured JSONL training stats. Designed to be driven by an external LLM agent or human.

Summary of decisions:

  • Hyperparam delivery (#30): flat env vars, PPOB_* prefix, stdlib only
  • Tunable params (#31): all 11 non-structural hyperparams exposed via env vars
  • Output format (#32): JSON lines to training_log.jsonl
  • Stall detection (#33): no heuristics in runner; rich stats for external observer
  • Training runner (#34): tools/training_runner/ — Java runner + shell wrapper with crash-restart loop
  • Adam persistence (#35): optimizer state + round counter survive restarts
  • LLM supervisor (#36): out of scope — separate effort
## Map complete All tickets resolved. The destination is reached: `tools/training_runner/` is a dumb CLI tool that runs PPO_Bot against a chosen opponent for unlimited rounds with crash recovery, outputting structured JSONL training stats. Designed to be driven by an external LLM agent or human. Summary of decisions: - **Hyperparam delivery (#30):** flat env vars, `PPOB_*` prefix, stdlib only - **Tunable params (#31):** all 11 non-structural hyperparams exposed via env vars - **Output format (#32):** JSON lines to `training_log.jsonl` - **Stall detection (#33):** no heuristics in runner; rich stats for external observer - **Training runner (#34):** `tools/training_runner/` — Java runner + shell wrapper with crash-restart loop - **Adam persistence (#35):** optimizer state + round counter survive restarts - **LLM supervisor (#36):** out of scope — separate effort
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#29