75e32e3315
Single 3789-round battle (5212-9000), only rounds 1-2 lost (cold start);
last 100 rounds 100%. Frozen-policy eval (PPOB_EVAL_ONLY=1, 500 rounds vs
Fire, counter 9000->9500): 500/500 = 100%, zero train lines, weights mtime
and content untouched. vLoss avg 39.7/max 361 stable throughout — bounded
terminal reward (db99153) holds at cumulative score ~462k.
24 lines
1.2 KiB
Markdown
24 lines
1.2 KiB
Markdown
# Fire campaign — final results (2026-08-19)
|
||
|
||
Policy: PPO_Bot, single 3789-round battle (global rounds 5212–9000), after the
|
||
bounded terminal-reward fix (db99153: computeRoundReward caps cumulative score
|
||
at 400/50 → bonus ∈ [0,8]).
|
||
|
||
## Training battle (run.sh Fire 9000, counter 5211 → 9000)
|
||
- 3789 rounds: **3787 wins / 3789 = 99.95%**
|
||
- Only losses: round 1 (score 48, ticks 645) and round 2 (score 128, ticks 449) — cold start
|
||
- 100/100 wins in the last 100 rounds
|
||
- train:game:round = 3789:3789:3789 (1:1:1), zero NaN
|
||
- valueLoss avg 39.7, max 361.1, last (r9000) 32.5; gradNorm max 1151.3 — stable, no divergence
|
||
- Battle wall time ≈ 6.4 min (~0.1 s/round; round ticks ~440–510, no learning speedup)
|
||
|
||
## Frozen-policy eval (PPOB_EVAL_ONLY=1, run.sh Fire 9500, counter 9000 → 9500)
|
||
- 500 rounds vs Fire: **500/500 = 100.0%**
|
||
- **0 train lines** during eval (no learning); weights/latest .npy mtimes and
|
||
content unchanged before/after eval — policy frozen
|
||
- Final weights committed at counter 9500
|
||
|
||
## Reproduce a 100% run
|
||
1. Weights: `PPO_Bot/weights/latest` at counter 9500 (this commit).
|
||
2. `PPOB_EVAL_ONLY=1 ./tools/training_runner/run.sh Fire <counter+N>`
|
||
— one battle of N rounds vs Fire, frozen policy, no training. |