cap terminal reward scale (bounded score bonus) and restore single-battle campaigns
computeRoundReward used cumulative totalScore/50 — unbounded in long battles (vLoss 353 at round 3160 → 25745 by 3871 in the 5841-round attempt). Cap the score term at 400 before /50: bonus ∈ [0,8], so the critic's value scale stays stable regardless of battle length and across battle boundaries. Reverts the 60-round battle chunking (186e005/da2f825): one battle per campaign for the whole remaining budget; keeps the crash-restart loop, the mid-battle freeze guard and the end-of-battle counter completeness check. Cert (5211, single 60-round battle): 59/60 wins (sole loss = cold-start round 1, score 61), vLoss avg 36.1 / max 148.5, gNorm max 596, zero NaN, zero restarts, counter check passed. Weights persist to round 5211.
This commit is contained in:
Binary file not shown.
@@ -23,6 +23,9 @@ block testTickReward:
|
||||
block testRoundReward:
|
||||
let r = computeRoundReward(350.0'f32)
|
||||
check abs(r - 7.0'f32) < 1e-6'f32, "computeRoundReward(350) == 7.0, got " & $r
|
||||
# bounded: long-battle cumulative scores must saturate, not blow the value scale
|
||||
check abs(computeRoundReward(89299.0'f32) - 8.0'f32) < 1e-6'f32,
|
||||
"computeRoundReward(89299) == 8.0 (capped), got " & $computeRoundReward(89299.0'f32)
|
||||
|
||||
# ── TrajectoryBuffer ──────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
Reference in New Issue
Block a user