cap terminal reward scale (bounded score bonus) and restore single-battle campaigns
computeRoundReward used cumulative totalScore/50 — unbounded in long battles (vLoss 353 at round 3160 → 25745 by 3871 in the 5841-round attempt). Cap the score term at 400 before /50: bonus ∈ [0,8], so the critic's value scale stays stable regardless of battle length and across battle boundaries. Reverts the 60-round battle chunking (186e005/da2f825): one battle per campaign for the whole remaining budget; keeps the crash-restart loop, the mid-battle freeze guard and the end-of-battle counter completeness check. Cert (5211, single 60-round battle): 59/60 wins (sole loss = cold-start round 1, score 61), vLoss avg 36.1 / max 148.5, gNorm max 596, zero NaN, zero restarts, counter check passed. Weights persist to round 5211.
This commit is contained in:
@@ -79,13 +79,11 @@ while true; do
|
||||
sleep 2
|
||||
fi
|
||||
|
||||
# ponytail: cap each battle at 60 rounds — matches the certified config.
|
||||
# Round-end reward is cumulative totalScore/50; long battles grow it unboundedly
|
||||
# and blow up the critic's value scale (seen at round 3160, vLoss 25745 by 3871).
|
||||
# Short battles bound the reward, and fresh-process restarts clear thread state.
|
||||
battle_rounds=$(( remaining > 60 ? 60 : remaining ))
|
||||
echo ">>> Running $battle_rounds rounds (opponent: $OPPONENT)..."
|
||||
java -cp ".:$JAR" RunTraining "$OPPONENT" "$battle_rounds" || {
|
||||
# One battle for the whole remaining budget: round-end reward is a bounded
|
||||
# terminal bonus (computeRoundReward caps totalScore at 400), so long battles
|
||||
# keep the critic's value scale stable — no per-battle chunking needed.
|
||||
echo ">>> Running $remaining rounds (opponent: $OPPONENT)..."
|
||||
java -cp ".:$JAR" RunTraining "$OPPONENT" "$remaining" || {
|
||||
exit_code=$?
|
||||
echo ">>> Java runner exited with code $exit_code — checking if done..."
|
||||
remaining=$(remaining_rounds)
|
||||
|
||||
Reference in New Issue
Block a user