Files
SirRoboGarage/tools/training_runner/training.env
T
SirStone 82eeb53e5c tune(training): 20-round chunks, 50 eval rounds, 30k total rounds
- generalist_train.sh: CHUNK_SIZE 60→20 for faster opponent cycling
- generalist_train.sh: EVAL_ROUNDS 30→50 for more reliable eval
- training.env: TRAINING_ROUNDS→30000, removed hardcoded opponent
- warm_start.py: TARGET_DIM=57 (already committed, ensure latest)
2026-08-20 14:38:09 +02:00

17 lines
429 B
Bash

# TRAINING_OPPONENT=Fire
TRAINING_ROUNDS=30000
PPOB_LOG_FILE=/home/davide/Projects/SirRoboGarage/tools/training_runner/logs/fire_training.jsonl
PPOB_LR=0.0003
PPOB_CLIP_EPSILON=0.2
PPOB_ENTROPY_COEFF=0.01
PPOB_VALUE_LOSS_COEFF=0.5
PPOB_MAX_GRAD_NORM=0.5
PPOB_GAMMA=0.99
PPOB_LAM=0.95
PPOB_EPOCHS=4
PPOB_MINI_BATCH_SIZE=64
PPOB_LOG_STD_FLOOR=-3.0
PPOB_INITIAL_LOG_STD=0.0
# PPOB_EVAL_ONLY=1 → freeze training (pure evaluation)