cap terminal reward scale (bounded score bonus) and restore single-battle campaigns

computeRoundReward used cumulative totalScore/50 — unbounded in long battles
(vLoss 353 at round 3160 → 25745 by 3871 in the 5841-round attempt). Cap the
score term at 400 before /50: bonus ∈ [0,8], so the critic's value scale stays
stable regardless of battle length and across battle boundaries.

Reverts the 60-round battle chunking (186e005/da2f825): one battle per
campaign for the whole remaining budget; keeps the crash-restart loop, the
mid-battle freeze guard and the end-of-battle counter completeness check.

Cert (5211, single 60-round battle): 59/60 wins (sole loss = cold-start round
1, score 61), vLoss avg 36.1 / max 148.5, gNorm max 596, zero NaN, zero
restarts, counter check passed. Weights persist to round 5211.
This commit is contained in:
2026-08-19 03:50:14 +02:00
parent da2f825ad8
commit db99153f65
46 changed files with 25 additions and 22 deletions
BIN
View File
Binary file not shown.
Binary file not shown.
+3
View File
@@ -23,6 +23,9 @@ block testTickReward:
block testRoundReward: block testRoundReward:
let r = computeRoundReward(350.0'f32) let r = computeRoundReward(350.0'f32)
check abs(r - 7.0'f32) < 1e-6'f32, "computeRoundReward(350) == 7.0, got " & $r check abs(r - 7.0'f32) < 1e-6'f32, "computeRoundReward(350) == 7.0, got " & $r
# bounded: long-battle cumulative scores must saturate, not blow the value scale
check abs(computeRoundReward(89299.0'f32) - 8.0'f32) < 1e-6'f32,
"computeRoundReward(89299) == 8.0 (capped), got " & $computeRoundReward(89299.0'f32)
# ── TrajectoryBuffer ────────────────────────────────────────────────────────── # ── TrajectoryBuffer ──────────────────────────────────────────────────────────
+3 -1
View File
@@ -69,7 +69,9 @@ proc computeTickReward*(myEnergyDelta, enemyEnergyDelta: float32;
proc computeRoundReward*(roundScore: float32): float32 = proc computeRoundReward*(roundScore: float32): float32 =
## Normalise round-end score to a rough ±6 range (doubled win signal). ## Normalise round-end score to a rough ±6 range (doubled win signal).
result = roundScore / 50.0'f32 ## ponytail: cap cumulative score at 400 before /50 — bounded terminal bonus
## keeps critic value scale stable across battle boundaries and long battles.
result = min(roundScore, 400.0'f32) / 50.0'f32
# ── GAE ─────────────────────────────────────────────────────────────────────── # ── GAE ───────────────────────────────────────────────────────────────────────
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+13 -13
View File
@@ -1,13 +1,13 @@
134308 200200
134308 200200
134308 200200
134308 200200
134308 200200
134308 200200
134308 200200
134308 200200
134308 200200
134308 200200
134308 200200
134308 200200
134308 200200
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+1 -1
View File
@@ -1 +1 @@
3159 5211
+5 -7
View File
@@ -79,13 +79,11 @@ while true; do
sleep 2 sleep 2
fi fi
# ponytail: cap each battle at 60 rounds — matches the certified config. # One battle for the whole remaining budget: round-end reward is a bounded
# Round-end reward is cumulative totalScore/50; long battles grow it unboundedly # terminal bonus (computeRoundReward caps totalScore at 400), so long battles
# and blow up the critic's value scale (seen at round 3160, vLoss 25745 by 3871). # keep the critic's value scale stable — no per-battle chunking needed.
# Short battles bound the reward, and fresh-process restarts clear thread state. echo ">>> Running $remaining rounds (opponent: $OPPONENT)..."
battle_rounds=$(( remaining > 60 ? 60 : remaining )) java -cp ".:$JAR" RunTraining "$OPPONENT" "$remaining" || {
echo ">>> Running $battle_rounds rounds (opponent: $OPPONENT)..."
java -cp ".:$JAR" RunTraining "$OPPONENT" "$battle_rounds" || {
exit_code=$? exit_code=$?
echo ">>> Java runner exited with code $exit_code — checking if done..." echo ">>> Java runner exited with code $exit_code — checking if done..."
remaining=$(remaining_rounds) remaining=$(remaining_rounds)