[Map] Autonomous LLM-supervised PPO_Bot training pipeline #29
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Destination
A standalone CLI training tool (
tools/training_runner/) that runs PPO_Bot against a chosen opponent for unlimited rounds with crash recovery, outputting structured training stats to files. Runnable, stoppable, and re-runnable from the command line — designed so an external LLM agent (or human) can monitor output, edit hyperparameters, and restart training without the tool knowing or caring who's driving it.Notes
BattleRunnerAPItools/battle_runner/— keep it untouched for debuggingweights/latest/on startup (crash recovery baseline works)/grilling,/domain-modelingDecisions so far
training_log.jsonl, one object per round with game outcome + training health + hyperparam snapshot; no derived stats, reader computes thoseNot yet specified
Out of scope
Map complete
All tickets resolved. The destination is reached:
tools/training_runner/is a dumb CLI tool that runs PPO_Bot against a chosen opponent for unlimited rounds with crash recovery, outputting structured JSONL training stats. Designed to be driven by an external LLM agent or human.Summary of decisions:
PPOB_*prefix, stdlib onlytraining_log.jsonltools/training_runner/— Java runner + shell wrapper with crash-restart loop