Structured output format for LLM consumption #32
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Parent: #29
Question
What machine-readable format should PPO_Bot / the training runner emit so the LLM supervisor can parse training stats reliably?
Options:
{"round":42,"ticks":300,"reward":1.2,"avgReward":0.8,"score":150,"actorLoss":0.03,"valueLoss":0.5,"gradNorm":0.3,"win":true}Current state: PPO_Bot prints unstructured text:
R:{n} ticks:{n} avgR:{f} score:{n}andtrained R:{n} aLoss:{f} vLoss:{f} gNorm:{f}. This would need to change or be supplemented.Constraints:
Resolution
Format: JSON lines, appended to
training_log.jsonl, one object per round.Fields per round:
round— cumulative across restartsticks— round lengthreward— total reward this roundavgReward— average reward per tickscore— Tank Royale scorewin— boolean, did PPO_Bot winactorLoss— PPO actor lossvalueLoss— PPO value/critic lossgradNorm— gradient norm after clippingenemyName— opponent bot nametimestamp— ISO 8601 wall clockhyperparams— object snapshot of current tunable params (lr, clipEpsilon, entropyCoeff, etc.)No rolling averages or trend indicators — reader computes those from raw rows.
No console-only stats — everything goes to the file. Console can echo a human-readable summary but the file is the source of truth.