PPO_Bot Fire campaign: 3787/3789 wins (99.95%); frozen eval 500/500
Single 3789-round battle (5212-9000), only rounds 1-2 lost (cold start);
last 100 rounds 100%. Frozen-policy eval (PPOB_EVAL_ONLY=1, 500 rounds vs
Fire, counter 9000->9500): 500/500 = 100%, zero train lines, weights mtime
and content untouched. vLoss avg 39.7/max 361 stable throughout — bounded
terminal reward (db99153) holds at cumulative score ~462k.
This commit is contained in:
@@ -0,0 +1,24 @@
|
||||
# Fire campaign — final results (2026-08-19)
|
||||
|
||||
Policy: PPO_Bot, single 3789-round battle (global rounds 5212–9000), after the
|
||||
bounded terminal-reward fix (db99153: computeRoundReward caps cumulative score
|
||||
at 400/50 → bonus ∈ [0,8]).
|
||||
|
||||
## Training battle (run.sh Fire 9000, counter 5211 → 9000)
|
||||
- 3789 rounds: **3787 wins / 3789 = 99.95%**
|
||||
- Only losses: round 1 (score 48, ticks 645) and round 2 (score 128, ticks 449) — cold start
|
||||
- 100/100 wins in the last 100 rounds
|
||||
- train:game:round = 3789:3789:3789 (1:1:1), zero NaN
|
||||
- valueLoss avg 39.7, max 361.1, last (r9000) 32.5; gradNorm max 1151.3 — stable, no divergence
|
||||
- Battle wall time ≈ 6.4 min (~0.1 s/round; round ticks ~440–510, no learning speedup)
|
||||
|
||||
## Frozen-policy eval (PPOB_EVAL_ONLY=1, run.sh Fire 9500, counter 9000 → 9500)
|
||||
- 500 rounds vs Fire: **500/500 = 100.0%**
|
||||
- **0 train lines** during eval (no learning); weights/latest .npy mtimes and
|
||||
content unchanged before/after eval — policy frozen
|
||||
- Final weights committed at counter 9500
|
||||
|
||||
## Reproduce a 100% run
|
||||
1. Weights: `PPO_Bot/weights/latest` at counter 9500 (this commit).
|
||||
2. `PPOB_EVAL_ONLY=1 ./tools/training_runner/run.sh Fire <counter+N>`
|
||||
— one battle of N rounds vs Fire, frozen policy, no training.
|
||||
Reference in New Issue
Block a user