fix(training): cap battles at 60 rounds to bound round-end reward scale

Round-end reward = cumulative totalScore/50 grows unboundedly with battle
length; long battles (5841 rounds) blew the critic's value scale: vLoss
10-30 during the 60-round cert, 353 at battle-1 round 1, 25745 by round 3871,
policy drift to 0/6 wins. 60-round battles reproduce the certified regime:
bounded value targets, fresh bot process per battle (clears thread state).
This commit is contained in:
2026-08-19 03:32:31 +02:00
parent a4e830531b
commit 186e005a96
+7 -2
View File
@@ -79,8 +79,13 @@ while true; do
sleep 2
fi
echo ">>> Running $remaining rounds (opponent: $OPPONENT)..."
java -cp ".:$JAR" RunTraining "$OPPONENT" "$remaining" && break || {
# ponytail: cap each battle at 60 rounds — matches the certified config.
# Round-end reward is cumulative totalScore/50; long battles grow it unboundedly
# and blow up the critic's value scale (seen at round 3160, vLoss 25745 by 3871).
# Short battles bound the reward, and fresh-process restarts clear thread state.
battle_rounds=$(( remaining > 60 ? 60 : remaining ))
echo ">>> Running $battle_rounds rounds (opponent: $OPPONENT)..."
java -cp ".:$JAR" RunTraining "$OPPONENT" "$battle_rounds" && break || {
exit_code=$?
echo ">>> Java runner exited with code $exit_code — checking if done..."
remaining=$(remaining_rounds)