There are no genuinely competitive adversaries for Tank Royale, and the
in-repo ones were broken until recently. Classic Robocode 1.9.5.5 is
obtainable (SourceForge, 20.4 MB) and its programmatic control API
(robocode.control.RobocodeEngine + BattleAdaptor.onTurnEnded) can run real
battles headless and expose per-turn robot state. So a legacy leader bot's
MOVEMENT can be captured and used as a gun-testing fixture with no port.
Captured unmodified DrussGT 3.1.4159 vs spinbot/ramfire/crazy/corners and a
mirror match: 28,797 ticks, plus two trivial-bot contrasts. Conversion to
the Tank Royale convention is validated to 0.000-0.001 deg by recomputing
the direction implied by (heading, speed) and comparing it against the
recorded per-tick displacement -- i.e. the data is proven to be genuine
recorded motion rather than a mangled export. (A first attempt treated the
snapshot API's headings as degrees; they are radians, ~95 deg off.)
The statistics confirm it is really a wave surfer: perpendicular to the
opponent 65-96% of ticks, radial ~0.001, 41-72% of ticks at full speed,
reversing on 42-46% of ticks, holding range at a 283-526 px median. The
straight-line contrast is radial-dominant (0.75) with ZERO reversals.
Discovery: DrussGT detects predictable guns and switches to a bullet-shield
stand-still mode, so captures against sample.Walls/TrackFire had to be
rejected as non-movement.
CAVEATS, recorded in DRUSSGT_FIXTURES.md: these are open-loop (replayed
DrussGT never dodges OUR bullets) and perfect-information (the observer
gives true positions every tick, unlike our stale live WorldState). Both
make our guns look better than in live play, so use them for RELATIVE gun
ranking, not absolute hit rates.
Jars stay out of git; capture tooling is reproducible via capture.sh.
Gun evaluation previously required a full end-to-end battle (Java server +
battle runner + websocket IPC to 2 bot processes, 50 rounds, ~3.4 min) and
yielded only ~300-900 REAL shots across 13 guns -- far too few to rank
guns, which is why tuning needed many repetitions.
VirtualTracker is already a pure function of (WorldState stream, gun list);
the only reason it needed Java was where WorldState came from. So the range
replays a seq[WorldState] through the SAME tracker: offline and online
scores are the same metric by construction, not an approximation.
ACCEPTANCE TEST (the point of the whole thing): record one live round, replay
it offline, compare per-gun virtual hit rates. 12/12 deterministic guns match
EXACTLY, reproduced twice. Tsetlin is compared separately because tmLearnOne
calls rand(). Getting to 12/12 exposed two real ordering quirks in the live
loop: run() calls go() before the aim/fire block, so tickBullets resolves
against the NEXT tick's scan while the prediction used the previous one; and
if the target dies during that go() the final tick's spawn+resolution is
skipped entirely. The recorder emits an end marker for the second case.
The 5th (selected-gun) predict call was verified to be a no-op.
Measured cost: 8 fixtures (1770 ticks, ~92k virtual bullets, 13 guns) replay
in 2.9 s, ~32k virtual bullets/s -- roughly 70x faster and 100x more samples
than a live gauntlet.
Also adds a per-tick WorldState recorder behind const RecordWorldState
(default off, mirrors the ShotLog idiom) which records the state the bot
ACTUALLY builds, staleness included, rather than true positions -- recording
the latter would hand the guns perfect information and produce flattering
scores.
9 new guard checks (33 total, all passing), including fixture round-trip,
replay determinism, stationary->HeadOn 100%, constant-velocity->Linear>HeadOn,
and the energy-threshold turner crossing at t=41.
Add --max-speed flag to TestBattleRunner (sets defaultTurnsPerSecond=-1 for unlimited TPS).
Add SNNBot_garage/tests/test_bullet_economy.nim: 10-round vs WallsBot, prints per-round and summary stats for tuning.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Document exception handling, zero-value BotResult trap, shared adversary bots
- Add offline parsing example using parseServerOutput
- Skip tests gracefully when JARs missing (guard before suite blocks)
- Fix blocking readLine in runner_process.nim: poll with 50ms sleep + atEnd check
(was preventing timeout enforcement, now blocks correctly during battle)
- Add test task to QBot.nimble and config.nims setup docs to AGENTS.md
- Add debug logging to TestBattleRunner for bot identity tracking
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Implements:
- BattleResult type and JSON-lines parser (#127)
- TR server lifecycle manager (#128)
- Bot compiler using nim c (#129)
- runBattle() orchestrator (#130)
- Example test in OscillatorBot_garage (#131)
- Framework usage guide (#132)
- TestBattleRunner.java for external server (#133)
- BattleRunner process lifecycle (#134)
Fix: runner_process.nim was redefining TimeoutError locally; now
uses std/net.TimeoutError consistently with server_manager.nim.
sac_train.sh orchestrates chunked self-play via tools/training_runner/
RunTraining.java: weighted opponent sampling per chunk, deterministic
eval (SACLSTM_EVAL_MODE=1) every N chunks with win-rate tracking, best
checkpoint (weights/sac_best.zip) by eval score, crash-restart loop on
the runner's liveness detection.
Supporting changes:
- integration.nim: opponentKey() keys the NewBattle buffer-clear rule on
getBotName(id) with numeric-id fallback (#49 Q14 follow-up);
bumpRoundCounter() emits the per-round liveness signal.
- SAC_LSTM_Bot.nim: onRoundEnded -> bumpRoundCounter().
- RunTraining.java: BOT_NAME env parameterizes result matching
(default PPO_Bot, unchanged behavior for PPO).
- Launch packaging: root SAC_LSTM_Bot.json + .sh for the booter;
src json name aligned to 'SAC_LSTM_Bot' so self-reported identity
matches the booted identity (mismatch = runner connect timeout).
Stochastic eval at std≈0.37 was 0/10 vs Corners (deterministic: 10/10).
Warm-start policy is correct but brittle — any noise breaks it.
- log_std initialized to -2.0 (std≈0.135) for moderate exploration
- entropy_coeff=0.0 (no push toward exploration during fine-tuning)
- logStd ceiling=-1.0 (cap at std≈0.37)
- Accumulate transitions across 10 rounds (~3000) before PPO update
(was per-round ~300 — gradient estimates were far too noisy)
- training.nim: MAX_TRANSITIONS 4096→8192, done flag on transitions,
GAE handles episode boundaries correctly
- PPO_Bot.nim: buffer persists across rounds, update every N rounds
- training.env: lr 5e-5→1e-4, entropy 0.001, UPDATE_INTERVAL=10
- Bullet state (indices 44-55): enemy-relative → bot-relative frame
(bot needs threat vectors to itself for dodging, not to enemy)
- New index 56: scan staleness = min(ticksSinceLastScan / 30, 1.0)
(gives policy a confidence signal for enemy data freshness)
- warm_start.py updated: 44→57 dim expansion, TARGET_DIM variable
- Tests updated for new state layout
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
computeRoundReward used cumulative totalScore/50 — unbounded in long battles
(vLoss 353 at round 3160 → 25745 by 3871 in the 5841-round attempt). Cap the
score term at 400 before /50: bonus ∈ [0,8], so the critic's value scale stays
stable regardless of battle length and across battle boundaries.
Reverts the 60-round battle chunking (186e005/da2f825): one battle per
campaign for the whole remaining budget; keeps the crash-restart loop, the
mid-battle freeze guard and the end-of-battle counter completeness check.
Cert (5211, single 60-round battle): 59/60 wins (sole loss = cold-start round
1, score 61), vLoss avg 36.1 / max 148.5, gNorm max 596, zero NaN, zero
restarts, counter check passed. Weights persist to round 5211.
RunTraining exits 0 after each chunk; && break ended the whole run after
the first battle (counter 3219, not 9000). Loop now falls through the
success path and re-checks the persisted counter each iteration.
Round-end reward = cumulative totalScore/50 grows unboundedly with battle
length; long battles (5841 rounds) blew the critic's value scale: vLoss
10-30 during the 60-round cert, 353 at battle-1 round 1, 25745 by round 3871,
policy drift to 0/6 wins. 60-round battles reproduce the certified regime:
bounded value targets, fresh bot process per battle (clears thread state).
The event queue's heap seq was the last GC'd block surviving across
rounds: each round runs on a freshly spawned bot thread, so the N+1
thread realloc'd a block grown by dead thread N's allocator mid-round
(at the next capacity doubling, ~turn 104) -> rawDealloc SIGSEGV in
addEvent (7 gdb-confirmed coredumps). Replace with a static
array[MAX_QUEUE_SIZE, BotEvent] + eventsLen: no heap block crosses
threads, realloc can never happen.
Also fix the harness aborting the final round mid-train: PPO_Bot's
onRoundEnded trains synchronously after the runner's RoundEndedEvent,
so the counter read right after awaitResults() is the stale pre-train
value and System.exit killed the bot inside ppoUpdate. Poll up to 60s
for the counter to catch up before declaring the battle incomplete.
Verified: 72 consecutive rounds vs Fire, 100% wins, all rounds trained
(counter advanced 1:1), zero coredumps since the fix.
Three bugs caused the radar to sweep continuously instead of locking:
1. run() loop set radar to Inf every tick, overwriting any lock
→ replaced with enemy_tracker.getRadarTurnRate()
2. onScannedBot used radarBearingTo() (math convention, east=0 CCW)
→ removed; run loop now handles radar via enemy_tracker
3. enemy_tracker.getRadarTurnRate() had arctan2(dx,dy) instead of
arctan2(dy,dx) — introduced by fd22535; bearing was off by ~90°
Also relaxed stale-lock threshold from 2 to 8 ticks to survive
brief scan gaps without falling back to full sweep.
Added tools/battle_runner for automated 1v1 testing.
Result: 1303/1308 ticks with successful scan (was ~1 in 4).