feat(PPO_Bot): multi-round transition accumulation (UPDATE_INTERVAL=10)

- Accumulate transitions across 10 rounds (~3000) before PPO update
  (was per-round ~300 — gradient estimates were far too noisy)
- training.nim: MAX_TRANSITIONS 4096→8192, done flag on transitions,
  GAE handles episode boundaries correctly
- PPO_Bot.nim: buffer persists across rounds, update every N rounds
- training.env: lr 5e-5→1e-4, entropy 0.001, UPDATE_INTERVAL=10
This commit is contained in:
2026-08-20 15:06:00 +02:00
parent 82eeb53e5c
commit fedab54bc0
3 changed files with 71 additions and 50 deletions
+5 -3
View File
@@ -1,9 +1,10 @@
# TRAINING_OPPONENT=Fire
TRAINING_ROUNDS=30000
PPOB_LOG_FILE=/home/davide/Projects/SirRoboGarage/tools/training_runner/logs/fire_training.jsonl
PPOB_LR=0.0003
PPOB_LR=1e-4
PPOB_UPDATE_INTERVAL=10
PPOB_CLIP_EPSILON=0.2
PPOB_ENTROPY_COEFF=0.01
PPOB_ENTROPY_COEFF=0.001
PPOB_VALUE_LOSS_COEFF=0.5
PPOB_MAX_GRAD_NORM=0.5
PPOB_GAMMA=0.99
@@ -11,6 +12,7 @@ PPOB_LAM=0.95
PPOB_EPOCHS=4
PPOB_MINI_BATCH_SIZE=64
PPOB_LOG_STD_FLOOR=-3.0
PPOB_INITIAL_LOG_STD=0.0
PPOB_LOG_STD_CEILING=0.0
PPOB_INITIAL_LOG_STD=-0.5
# PPOB_EVAL_ONLY=1 → freeze training (pure evaluation)