Build the training runner tool #34

Closed
opened 2026-08-17 19:48:57 +02:00 by SirStone · 1 comment
Owner

Parent: #29
Blocked by: #30, #32

Question

Build tools/training_runner/ — a Java + shell tool that runs PPO_Bot against a chosen opponent for a large number of rounds with crash recovery.

Requirements:

  • Java runner (based on RunBattle.java pattern): accept opponent name + round count as CLI args
  • Shell wrapper: compile PPO_Bot, launch Java runner, detect crashes, restart automatically
  • Structured stats output (format per the output format decision)
  • Track total rounds across restarts
  • Separate from existing tools/battle_runner/ (keep that for debugging)

Depends on: structured output format decision, hyperparam delivery decision.

Parent: #29 Blocked by: #30, #32 ## Question Build `tools/training_runner/` — a Java + shell tool that runs PPO_Bot against a chosen opponent for a large number of rounds with crash recovery. **Requirements:** - Java runner (based on RunBattle.java pattern): accept opponent name + round count as CLI args - Shell wrapper: compile PPO_Bot, launch Java runner, detect crashes, restart automatically - Structured stats output (format per the output format decision) - Track total rounds across restarts - Separate from existing `tools/battle_runner/` (keep that for debugging) **Depends on:** structured output format decision, hyperparam delivery decision.
SirStone added the wayfinder:task label 2026-08-17 19:48:57 +02:00
Author
Owner

Built: tools/training_runner/

What was built

tools/training_runner/RunTraining.java — Java battle runner that accepts [opponent] [rounds] CLI args (also via env vars TRAINING_OPPONENT/TRAINING_ROUNDS). Runs one round at a time in a loop, appending a {"type":"game",...} JSON line to PPOB_LOG_FILE after each round with: round number, ticks, score, win/loss, opponent name.

tools/training_runner/run.sh — Shell wrapper that:

  1. Compiles PPO_Bot (nimble build -d:release)
  2. Compiles RunTraining.java
  3. Loops: reads persisted weights/round_counter.txt to compute remaining rounds, restarts Java runner on crash, exits when all rounds done

PPO_Bot/PPO_Bot.nim — Extended with:

  • PPOB_* env var reading at startup (decision #30): PPOB_LR, PPOB_CLIP_EPSILON, PPOB_ENTROPY_COEFF, PPOB_VALUE_LOSS_COEFF, PPOB_MAX_GRAD_NORM, PPOB_GAMMA, PPOB_LAM, PPOB_EPOCHS, PPOB_MINI_BATCH_SIZE — all optional, defaults match existing code
  • PPOB_LOG_FILE env var: if set, PPO_Bot appends two JSON line types per round:
    • {"type":"round", "round":N, "ticks":N, "avgReward":f, "score":N, "ts":N} — game outcome at round end
    • {"type":"train", "round":N, "actorLoss":f, "valueLoss":f, "gradNorm":f, "ts":N, "lr":f, ...hyperparams} — training health when background thread finishes

PPO_Bot/training.nim — ppoUpdate now accepts gamma and lam params (forwarded from env vars via TrainingArgs).

Log format

Three JSON line types in training_log.jsonl (also written by Java runner):

  • {"type":"game", ...} — from Java runner (game-engine outcome)
  • {"type":"round", ...} — from PPO_Bot (reward stats)
  • {"type":"train", ...} — from PPO_Bot (gradient/loss health + hyperparam snapshot)

Reader correlates by round field. No derived stats — reader computes those (decision #32).

Crash recovery

PPO_Bot persists round counter to weights/round_counter.txt (decision #35). Shell wrapper reads it on each restart to compute remaining rounds. No work is lost.

Usage

cd tools/training_runner
./run.sh Target 500          # 500 rounds vs Target
PPOB_LR=1e-4 ./run.sh Corners 200   # tune lr on the fly
## Built: `tools/training_runner/` ### What was built **`tools/training_runner/RunTraining.java`** — Java battle runner that accepts `[opponent] [rounds]` CLI args (also via env vars `TRAINING_OPPONENT`/`TRAINING_ROUNDS`). Runs one round at a time in a loop, appending a `{"type":"game",...}` JSON line to `PPOB_LOG_FILE` after each round with: round number, ticks, score, win/loss, opponent name. **`tools/training_runner/run.sh`** — Shell wrapper that: 1. Compiles PPO_Bot (`nimble build -d:release`) 2. Compiles `RunTraining.java` 3. Loops: reads persisted `weights/round_counter.txt` to compute remaining rounds, restarts Java runner on crash, exits when all rounds done **`PPO_Bot/PPO_Bot.nim`** — Extended with: - `PPOB_*` env var reading at startup (decision #30): `PPOB_LR`, `PPOB_CLIP_EPSILON`, `PPOB_ENTROPY_COEFF`, `PPOB_VALUE_LOSS_COEFF`, `PPOB_MAX_GRAD_NORM`, `PPOB_GAMMA`, `PPOB_LAM`, `PPOB_EPOCHS`, `PPOB_MINI_BATCH_SIZE` — all optional, defaults match existing code - `PPOB_LOG_FILE` env var: if set, PPO_Bot appends two JSON line types per round: - `{"type":"round", "round":N, "ticks":N, "avgReward":f, "score":N, "ts":N}` — game outcome at round end - `{"type":"train", "round":N, "actorLoss":f, "valueLoss":f, "gradNorm":f, "ts":N, "lr":f, ...hyperparams}` — training health when background thread finishes **`PPO_Bot/training.nim`** — `ppoUpdate` now accepts `gamma` and `lam` params (forwarded from env vars via `TrainingArgs`). ### Log format Three JSON line types in `training_log.jsonl` (also written by Java runner): - `{"type":"game", ...}` — from Java runner (game-engine outcome) - `{"type":"round", ...}` — from PPO_Bot (reward stats) - `{"type":"train", ...}` — from PPO_Bot (gradient/loss health + hyperparam snapshot) Reader correlates by `round` field. No derived stats — reader computes those (decision #32). ### Crash recovery PPO_Bot persists round counter to `weights/round_counter.txt` (decision #35). Shell wrapper reads it on each restart to compute remaining rounds. No work is lost. ### Usage ```bash cd tools/training_runner ./run.sh Target 500 # 500 rounds vs Target PPOB_LR=1e-4 ./run.sh Corners 200 # tune lr on the fly ```
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#34