PPO_Bot Fire campaign: 3787/3789 wins (99.95%); frozen eval 500/500

Single 3789-round battle (5212-9000), only rounds 1-2 lost (cold start);
last 100 rounds 100%. Frozen-policy eval (PPOB_EVAL_ONLY=1, 500 rounds vs
Fire, counter 9000->9500): 500/500 = 100%, zero train lines, weights mtime
and content untouched. vLoss avg 39.7/max 361 stable throughout — bounded
terminal reward (db99153) holds at cumulative score ~462k.
This commit is contained in:
2026-08-19 04:00:47 +02:00
parent db99153f65
commit 75e32e3315
42 changed files with 38 additions and 14 deletions
@@ -0,0 +1,24 @@
# Fire campaign — final results (2026-08-19)
Policy: PPO_Bot, single 3789-round battle (global rounds 5212–9000), after the
bounded terminal-reward fix (db99153: computeRoundReward caps cumulative score
at 400/50 → bonus ∈ [0,8]).
## Training battle (run.sh Fire 9000, counter 5211 → 9000)
- 3789 rounds: **3787 wins / 3789 = 99.95%**
- Only losses: round 1 (score 48, ticks 645) and round 2 (score 128, ticks 449) — cold start
- 100/100 wins in the last 100 rounds
- train:game:round = 3789:3789:3789 (1:1:1), zero NaN
- valueLoss avg 39.7, max 361.1, last (r9000) 32.5; gradNorm max 1151.3 — stable, no divergence
- Battle wall time ≈ 6.4 min (~0.1 s/round; round ticks ~440–510, no learning speedup)
## Frozen-policy eval (PPOB_EVAL_ONLY=1, run.sh Fire 9500, counter 9000 → 9500)
- 500 rounds vs Fire: **500/500 = 100.0%**
- **0 train lines** during eval (no learning); weights/latest .npy mtimes and
content unchanged before/after eval — policy frozen
- Final weights committed at counter 9500
## Reproduce a 100% run
1. Weights: `PPO_Bot/weights/latest` at counter 9500 (this commit).
2. `PPOB_EVAL_ONLY=1 ./tools/training_runner/run.sh Fire <counter+N>`
— one battle of N rounds vs Fire, frozen policy, no training.