Structured output format for LLM consumption #32

Closed
opened 2026-08-17 19:48:44 +02:00 by SirStone · 1 comment
Owner

Parent: #29

Question

What machine-readable format should PPO_Bot / the training runner emit so the LLM supervisor can parse training stats reliably?

Options:

  1. JSON lines (one JSON object per round to stdout/file) — self-describing, easy to parse, LLM-native. E.g.: {"round":42,"ticks":300,"reward":1.2,"avgReward":0.8,"score":150,"actorLoss":0.03,"valueLoss":0.5,"gradNorm":0.3,"win":true}
  2. CSV — compact, appendable, but needs header convention and is fragile to schema changes
  3. Both — JSON to a log file, human-readable summary to console

Current state: PPO_Bot prints unstructured text: R:{n} ticks:{n} avgR:{f} score:{n} and trained R:{n} aLoss:{f} vLoss:{f} gNorm:{f}. This would need to change or be supplemented.

Constraints:

  • The LLM will read this output to decide if training is stuck
  • Must include enough signal: reward trend, loss magnitudes, win/loss, grad norm health
  • Must survive process restarts (append, not overwrite)
Parent: #29 ## Question What machine-readable format should PPO_Bot / the training runner emit so the LLM supervisor can parse training stats reliably? **Options:** 1. **JSON lines** (one JSON object per round to stdout/file) — self-describing, easy to parse, LLM-native. E.g.: `{"round":42,"ticks":300,"reward":1.2,"avgReward":0.8,"score":150,"actorLoss":0.03,"valueLoss":0.5,"gradNorm":0.3,"win":true}` 2. **CSV** — compact, appendable, but needs header convention and is fragile to schema changes 3. **Both** — JSON to a log file, human-readable summary to console **Current state:** PPO_Bot prints unstructured text: `R:{n} ticks:{n} avgR:{f} score:{n}` and `trained R:{n} aLoss:{f} vLoss:{f} gNorm:{f}`. This would need to change or be supplemented. **Constraints:** - The LLM will read this output to decide if training is stuck - Must include enough signal: reward trend, loss magnitudes, win/loss, grad norm health - Must survive process restarts (append, not overwrite)
SirStone added the wayfinder:grilling label 2026-08-17 19:48:44 +02:00
Author
Owner

Resolution

Format: JSON lines, appended to training_log.jsonl, one object per round.

Fields per round:

  • round — cumulative across restarts
  • ticks — round length
  • reward — total reward this round
  • avgReward — average reward per tick
  • score — Tank Royale score
  • win — boolean, did PPO_Bot win
  • actorLoss — PPO actor loss
  • valueLoss — PPO value/critic loss
  • gradNorm — gradient norm after clipping
  • enemyName — opponent bot name
  • timestamp — ISO 8601 wall clock
  • hyperparams — object snapshot of current tunable params (lr, clipEpsilon, entropyCoeff, etc.)

No rolling averages or trend indicators — reader computes those from raw rows.

No console-only stats — everything goes to the file. Console can echo a human-readable summary but the file is the source of truth.

## Resolution **Format:** JSON lines, appended to `training_log.jsonl`, one object per round. **Fields per round:** - `round` — cumulative across restarts - `ticks` — round length - `reward` — total reward this round - `avgReward` — average reward per tick - `score` — Tank Royale score - `win` — boolean, did PPO_Bot win - `actorLoss` — PPO actor loss - `valueLoss` — PPO value/critic loss - `gradNorm` — gradient norm after clipping - `enemyName` — opponent bot name - `timestamp` — ISO 8601 wall clock - `hyperparams` — object snapshot of current tunable params (lr, clipEpsilon, entropyCoeff, etc.) **No rolling averages or trend indicators** — reader computes those from raw rows. **No console-only stats** — everything goes to the file. Console can echo a human-readable summary but the file is the source of truth.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#32