The user pushed back on "the TM can't be your best 1v1 gun", correctly, because two
decisive tests had never been run. Both are now run and they agree.
TASK 1 - THE GF HEAD vs ITS MAJORITY-CLASS BASELINE (offline, n=1,751,067):
label histogram [254286, 284578, 678879, 297055, 236269]
majority class = 2 (the CENTRE bucket) = 38.77%
RAW head accuracy = 36.69% -> margin **-2.08 pp, BELOW majority**
GATED head accuracy = 40.37% vs 38.75% majority -> +1.62 pp, BUT it predicts the
majority class on 62.4% of ticks and its minority recall is 13.6% / 12.9% - a
base-rate predictor wearing a classifier's clothes.
Shuffled control sits at its own majority (20.04% vs 20.12%), confirming chance.
**THE OLD "46% vs 20% CHANCE" FIGURE I QUOTED WAS WRONG ON TWO COUNTS:** the
baseline is 38.8%, not 20%, and the 46% predated the deferred-label fix. Against
the correct baseline the head is BELOW it.
TASK 2 - THE FIRST-EVER LIVE A/B OF THE TM GUN (7 runs x 7 rounds per arm, one
frozen binary from git archive HEAD = eb74f9b2, sha256 cb66d66b..., real DrussGT,
every arm forced alone with TR_RACK_<GUN>=both and all 14 others off, liveness
confirmed per run):
arm shots real % dmg/run round wins
onlyPattern 4610 10.74% 285 25/49
onlyTMPATTERN (radial) 3374 3.50% 71 0/49
onlyLinear 3218 3.23% 61 0/49
Pattern vs TM: +7.22 pp / +213.7 dmg, exact p=0.0006
TM vs Linear: +0.30 pp, p=0.659 (dmg p=0.438)
**The TM is statistically INDISTINGUISHABLE from its own Linear base live.** So it
is not "the TM works and we are aiming it wrong".
DIRECT ANSWER: **(c) It loses live AND sits at/below majority - the target carries
no learnable signal beyond the base rate, and that is the reason.** The reason is
not the machine, not the knobs, and not the application alone: the thing it was
asked to predict is dominated by the modal answer.
This closes the TM-as-gun thread. If a TM is wanted in the bot, a firing gate or a
movement decision is a better fit for a boolean-rule classifier than an aim point -
that is untested and is a different project.
A LIVE GF-MODE ARM WAS NOT RUN (stated as unmeasured): the task pinned one frozen
HEAD binary and HEAD registers the TM gun as radial only; Task 1 already makes GF
the unpromising candidate.
HARNESS FIX WORTH KEEPING: `tools/ab/which_gun_arm_env.sh` left the TARGET gun
unset, so with the now-Pattern-only default it silently fell back to the FULL rack
- an arm could appear to test a single gun while actually running the whole rack.
It now emits `TR_RACK_<GUN>=both` for the target and `=off` for all 14 others.
(Earlier which-gun results are unaffected: they ran before the Pattern-only default,
or - as in the melee/1v1 campaign - set the explicit `=both` themselves.)
tm_pattern.nim gains a per-class confusion matrix (warm samples only) to support the
majority baseline; no behaviour change. Adds Round 4 to
tm_pattern_sweep_results.md with both tasks and the interpretation rule.
Follow-up to e0666a5, which showed Pattern alone (10.78%) beats the full rack
(6.93%). That left two open questions: is a SMALL rack of good guns better than
Pattern alone, and does the selector add value on a good rack (rather than only
on the bloated one)? Both are now answered: NO and NO.
6 arms x 7 runs x 7 rounds, one frozen binary from CLEAN HEAD e0666a5 (built via
`git archive`, source verified byte-identical to the clean tree), rack knobs
only, 8 concurrent battles, real server-side hit rate vs the real DrussGT, exact
two-sided permutation test on per-run rates.
arm guns (selector active?) real % dmg/run p vs onlyPattern
onlyPattern Pattern, NO selection 10.36 264 --
lean8 HeadOn,Linear,Circular,Accel,Pattern,GF,KNN,WallBounce 6.31 146 0.0169
lean6 lean8 - HeadOn 8.83 212 0.0262
pairPC Pattern + Circular 8.23 185 0.0460
pairPK Pattern + KNN 9.80 264 0.3998
pairPL Pattern + Linear 8.23 200 0.0035
The control replicates the prior run (10.36% vs 10.78% before; same binary tree,
different build path).
THE MECHANISM, from the per-arm selected-gun mix - the virtual signal keeps
ranking the WRONG guns first, even on a two-gun rack:
lean8: HeadOn 46.2% of ticks at 2.0% REAL; Pattern only 14.4% (12.6% real)
lean6: Pattern 29.2% (10.4% real) vs KNN 25.3% (8.1%) and Linear 15.2% (8.0%)
pairPC: Circular 66.8% (6.9% real) vs Pattern 33.2% (11.2% real) - over-picks Circular
pairPL: Linear 57.7% (6.0% real) vs Pattern 42.3% (11.1% real) - over-picks Linear
pairPK: Pattern 86.4% - ties ONLY because the selector happens to pick Pattern
most of the time; it is numerically lower with identical dmg/run
So the failure is NOT rack size. Pruning does not fix it; the ranking is wrong.
VERDICT: ship `onlyPattern` - Pattern alone with selection bypassed - at 10.36%
real and 264 dmg/run, vs lean8 6.31%/146 and the prior full rack 6.93%/159.
This DIRECTLY CONTRADICTS the standing user directive to keep virtual-fitness
selection, so it is recorded here plainly rather than quietly acted on: disable
the selector (`TR_RACK_<every gun but PATTERN>=off`) pending a better fitness
signal. The mechanism itself is left intact and functional so it can be re-enabled
with one env var, and so it can be fixed rather than discarded.
REMAINING CAVEAT: ONE ADVERSARY. All of this is vs DrussGT. Pattern as the default
must be re-checked against other bots first - that is the next job.
Extends the reusable harness (tools/ab/which_gun_arm_env.sh now has lean8/lean6/
pairPC/pairPK/pairPL; which_gun_analyze.py is parameterised by WHICHGUN_OUT and
compares against both `full` and `onlyPattern`).
MEASURED against the real DrussGT, one frozen binary built from clean HEAD, rack
knobs only (no source edits), 5 arms x 7 runs x 7 rounds, 8 concurrent battles,
judged ONLY on server-side real hit rate from the events sidecar, exact
two-sided permutation test on per-run rates.
arm runs shots hits real % dmg/run p vs full
full (shipped) 7 3898 270 6.93 159 --
onlyPattern 7 4582 494 10.78 287 0.0012 <- BETTER
onlyKNN 7 4033 207 5.13 119 0.1340
onlyLinear 7 3215 105 3.27 65 0.0082
onlyGF 7 3193 72 2.25 45 0.0012
Firing Pattern ALONE gives +3.85pp pooled hit rate and +80% damage per run, and
it fires MORE shots (4582 vs 3898) - it dominates on rate and volume. This is not
"any single gun wins" (full beats Linear, GF and KNN); it is specifically
"Pattern alone beats the rack".
WHY - the virtual fitness signal mis-ranks guns against real outcomes:
- HeadOn is massively over-selected: 31.4% of ticks, the most real shots (1070),
but only 4.5% REAL. It alone drags the rack down.
- Pattern has the best virtual rank and near-best real rate (11.9%, rank 2), yet
is selected only 22.6% of the time.
- Linear's apparent strength was SELECTION BIAS: conditional on being selected it
looked like 15.2% (n=33), but its UNCONDITIONAL rate (onlyLinear) is 3.27%.
Every earlier per-gun "real rate" in this repo is conditional on selection and
is therefore confounded. This experiment is the clean measurement.
NOT YET SETTLED (do not overclaim):
- ONE ADVERSARY. All of this is vs DrussGT. Pattern must be re-checked against
other bots before it becomes the default on this evidence alone.
- Whether a SMALL rack of good guns beats Pattern alone. The selector is negative
value on the CURRENT bloated rack; that does not prove it is negative value on
a rack of only good guns. That is the next experiment and it decides whether
the selection apparatus is fixed or disabled.
- The user's standing directive is to KEEP virtual-fitness selection. This
measurement conflicts with it, so the next step tests the selector on a small
good rack rather than assuming either answer.
Context - three prior selection-side attempts all failed: hysteresis (7.02% ->
5.10%, p=0.002), commitment (7.17% -> 4.44%, p=0.0012), arrival-accuracy
tie-break (7.08%, p=0.88 null). The per-tick random draw is load-bearing on
three independent measurements. This experiment locates the real problem one
level up: which guns are in the rack, and that the virtual signal ranks them
wrongly.
Preserves the reusable harness (tools/ab/which_gun_run_one.sh,
which_gun_arm_env.sh, which_gun_analyze.py) and the full writeup
(docs/selector_negative_value.md).
Running /tmp/tr_bots/DrussGT/DrussGT.sh manually failed with:
BotException: Required bot property 'name' is missing.
The Tank Royale Java bot API reads bot identity from BOT_* environment
variables (EnvVars: BOT_NAME, BOT_VERSION, BOT_AUTHORS, BOT_DESCRIPTION,
BOT_HOMEPAGE, BOT_COUNTRY_CODES, BOT_GAME_TYPES, BOT_PLATFORM, BOT_PROG_LANG,
BOT_INITIAL_POS). When the TR booter launches the directory it supplies those,
derived from the .json - which is why every headless battle worked - but
launching the script directly supplies nothing, so the API rejects the
handshake.
make_botdir.sh now exports them in the generated <dir>.sh (kept in sync with
the .json it writes), so the script works standalone as well as under the
booter.
Verified against a server started for the test: before, the exact error above;
after, 'Connected to: ws://localhost:4599' and the server logs
'Bot joined: DrussGT 3.1.4159'.
The classic captures were OPEN-LOOP: replayed DrussGT never dodged OUR
bullets. These come from real TR battles through the working Java bridge, so
the recording contains genuine reactions to ModularBot's live fire. The
open-loop caveat is gone (perfect-information remains).
PRIMARY RESULT - the boss beats us badly. DrussGT 1447 - ModularBot 300 over
15 rounds, ModularBot winning only round 5 (DrussGT died at tick 1893). Rounds
are long, not truncated: mean 1335 ticks, ModularBot got off 1134 shots.
ModularBot 1134 shots / 60 hits = 5.3% real hit rate
DrussGT 1400 shots / 169 hits = 12.1% real hit rate
So DrussGT's gun is ~2.3x more accurate than our entire rack, on top of far
better movement. That is the number to move.
Also captured: shield-on variant (DrussGT 939-287, 9/10 - ModularBot takes
round 1 to the known shield warm-up), and vs SpinBot 1175-0, Crazy 1080-1,
Corners 1659-0. 20,026 + 12,629 + 10,824 + 11,507 + 2,575 ticks.
Movement statistics match the classic set within ~0.04 on the perpendicular
and radial fractions, so this is the same wave surfer in TR physics:
TR vs modularbot: perp 0.967, radial 0.001, 52.8% at full speed,
reversing 46.3%, median range 464 px.
CLOSED LOOP PROVEN, not asserted. ModularBot's fire is a heat-limited near
metronome (median interval 14 ticks), which gives a usable exogenous clock:
- event-locked |delta heading| oscillates 0.96 -> 2.69 deg about a 1.47 deg
mean with the fire period, almost every lag outside the 95% band of a
400-iteration phase-shuffled null;
- cross-correlation of |delta heading| against the fire impulse peaks at
r = +0.111, lag 12 ticks, permutation p = 0.005 (null peak mean +0.016);
- OWN-FIRE CONTROL is flat, so the oscillation is enemy-driven rather than
an internal cadence;
- range response is weak (~3 px over 30 ticks, near noise) and is therefore
NOT claimed, and per-bullet dodging is not claimed either because the
bullet detector (shield) is off.
CAVEATS: still perfect-information (observer gives true positions every tick,
unlike the live bot's stale between-scan WorldState) so these remain optimistic
vs live play; and they are open-loop AT REPLAY TIME - 'closed_loop' describes
the capture, not a later replay. TR conversion residual is ~1.5 deg mean
because the TR server moves along the pre-turn heading, vs 0.000 deg for the
classic captures.
Adds analyze_closed_loop.py (PSTH event-locking, phase-shuffle permutation
null, cross-correlation, own-fire control) and per-round result sidecars.
The unmodified DrussGT.jar connects, wave-surfs, fires and beats every
adversary we have. 5 rounds each, all rounds won:
SpinBot 542-16, Corners 823-4, Crazy 604-0, RamFire 900-0,
ModularBot 506-72.
It is genuinely surfing, not drifting or stalling. Movement statistics
against the classic captures, same metrics, same analyzer:
opp perp TR/classic reversing TR/cl median range TR/cl
SpinBot 0.967 / 0.962 0.466 / 0.464 462 / 402
Corners 0.894 / 0.920 0.462 / 0.419 520 / 504
Crazy 0.841 / 0.829 0.527 / 0.439 370 / 378
RamFire 0.762 / 0.651 0.533 / 0.417 326 / 283
ModularBot 0.956 / - 0.466 / - 453 / -
Every round starts moving within 3-11 ticks and runs at 75-82% full speed.
Implemented: the classic remaining-quantity motion model (delegated to the
TR Bot's own Nat-Pavasant model - identical constants: accel 1, decel -2,
max 8, body turn 10-0.75|v|, gun 20, radar 45 - so getDistanceRemaining and
getTurnRemaining are exactly self-consistent with what is emulated); event
synthesis with classic ordering; bullet identity via object identity; gun
heat; rounds; radar cadence; firing translation; and the ThreadManager
landmine is killed by installing a no-op IThreadManagerBase in
ContainerBase.instance (verified: without it 'RobotException: ThreadManager
cannot be null!' kills the bot thread; with it the write succeeds).
Also adds TrBattleCapture, an observer that dumps per-tick state in the SAME
JSONL fixture format as the classic capture, so legacy bots can be captured
from Tank Royale battles too.
HONEST DIVERGENCES (README section 5.9): the TR server moves along the
PRE-turn heading and then turns, while classic aligns displacement with the
POST-turn heading, so the analyzer's conversion error is ~1.5 deg rather than
0.000 deg; distanceRemaining decrements by target speed rather than actual
distance; collision clamping differs; BulletMissedEvent can fire less often
because age-expiry has no TR event; bulletId is a local temp id;
StatusEvent/onPaint/SkippedTurnEvent are never delivered. Physics fidelity
diverges by construction - expect to retune.
EnergyDomeWorker (the bullet shield) is off by default: its precise
bullet-detection warm-up makes round 1 up to 79% stationary vs 35% in
classic. Rounds 2+ match classic closely with it on, but the pure surfer is
the consistent path. DRUSSGT_SHIELD=1 re-enables it.
Jars remain out of git.
The question was whether we can fight a genuine legacy leader bot in Tank
Royale. Answer: the API side is now PROVEN, not estimated.
KEY FINDING: the classic robocode.* API is a thin delegation layer over a
public seam. javap -c shows AdvancedRobot forwarding every call to
_RobotBase.peer (IBasicRobotPeer/IAdvancedRobotPeer), and _RobotBase.setPeer
is public final. So we do NOT need to reimplement the API: we reuse the
genuine robocode.jar and implement only the 75-method peer interface.
Consequences:
- DrussGT's 22 sources compile against the real API with ZERO unresolved
symbols. (The literal 22-file javac fails only on two PRE-EXISTING
duplicate classes - GFRange and Indice are declared both inline in
DrussGunDC.java and as standalone files - and the 6 classes that ship
without source. The jar supplies all of them, so a shim never cares.)
- Runtime smoke test PASSES: the unmodified DrussGT.jar is loaded through a
child URLClassLoader and driven for 200 synthetic ticks, emitting movement
intents every tick and 169 fire requests, with its thread surviving.
This is the cheap path to the 'final boss', and it also unlocks the whole
roborumble archive rather than one bot.
Traps found by measurement:
- ScannedRobotEvent's constructor order is (name, energy, bearing, distance,
heading, velocity) - NOT heading-before-bearing. The wrong order silently
yields distance=0, an immediate KD-tree insert and an NPE in
EnemyMoves.predict; it hung the first smoke run.
- RobocodeFileOutputStream has a hard engine dependency (resolves
IThreadManagerBase via ContainerBase) and throws 'ThreadManager cannot be
null!' outside the engine, killing the bot thread. Reached from DrussGT's
own contain() error logging, so it must be stubbed.
- robocode.RobotDeathEvent is required and was NOT in the predicted API list.
- Bullet.equals() is genuinely called for bullet identity, not just getters.
REMAINING WORK (README section 5): coordinate rotation DONE, execute()->tick
bridge DONE and proven, ThreadManager fix scoped. The main open item is the
classic motion model (setAhead distance semantics vs TR speed), estimated
1-3 days, plus event synthesis/ordering, gun heat, bullet identity, round
and radar cadence, and firing translation. NO API UNKNOWNS REMAIN.
Physics fidelity will still diverge from classic - expect to retune.
Jars stay out of git (blocked by tools/robocode_shim/.gitignore); the
genuine robocode.jar and DrussGT.jar are referenced from /tmp via
ROBOCODE_JAR / DRUSSGT_JAR.
There are no genuinely competitive adversaries for Tank Royale, and the
in-repo ones were broken until recently. Classic Robocode 1.9.5.5 is
obtainable (SourceForge, 20.4 MB) and its programmatic control API
(robocode.control.RobocodeEngine + BattleAdaptor.onTurnEnded) can run real
battles headless and expose per-turn robot state. So a legacy leader bot's
MOVEMENT can be captured and used as a gun-testing fixture with no port.
Captured unmodified DrussGT 3.1.4159 vs spinbot/ramfire/crazy/corners and a
mirror match: 28,797 ticks, plus two trivial-bot contrasts. Conversion to
the Tank Royale convention is validated to 0.000-0.001 deg by recomputing
the direction implied by (heading, speed) and comparing it against the
recorded per-tick displacement -- i.e. the data is proven to be genuine
recorded motion rather than a mangled export. (A first attempt treated the
snapshot API's headings as degrees; they are radians, ~95 deg off.)
The statistics confirm it is really a wave surfer: perpendicular to the
opponent 65-96% of ticks, radial ~0.001, 41-72% of ticks at full speed,
reversing on 42-46% of ticks, holding range at a 283-526 px median. The
straight-line contrast is radial-dominant (0.75) with ZERO reversals.
Discovery: DrussGT detects predictable guns and switches to a bullet-shield
stand-still mode, so captures against sample.Walls/TrackFire had to be
rejected as non-movement.
CAVEATS, recorded in DRUSSGT_FIXTURES.md: these are open-loop (replayed
DrussGT never dodges OUR bullets) and perfect-information (the observer
gives true positions every tick, unlike our stale live WorldState). Both
make our guns look better than in live play, so use them for RELATIVE gun
ranking, not absolute hit rates.
Jars stay out of git; capture tooling is reproducible via capture.sh.
Gun evaluation previously required a full end-to-end battle (Java server +
battle runner + websocket IPC to 2 bot processes, 50 rounds, ~3.4 min) and
yielded only ~300-900 REAL shots across 13 guns -- far too few to rank
guns, which is why tuning needed many repetitions.
VirtualTracker is already a pure function of (WorldState stream, gun list);
the only reason it needed Java was where WorldState came from. So the range
replays a seq[WorldState] through the SAME tracker: offline and online
scores are the same metric by construction, not an approximation.
ACCEPTANCE TEST (the point of the whole thing): record one live round, replay
it offline, compare per-gun virtual hit rates. 12/12 deterministic guns match
EXACTLY, reproduced twice. Tsetlin is compared separately because tmLearnOne
calls rand(). Getting to 12/12 exposed two real ordering quirks in the live
loop: run() calls go() before the aim/fire block, so tickBullets resolves
against the NEXT tick's scan while the prediction used the previous one; and
if the target dies during that go() the final tick's spawn+resolution is
skipped entirely. The recorder emits an end marker for the second case.
The 5th (selected-gun) predict call was verified to be a no-op.
Measured cost: 8 fixtures (1770 ticks, ~92k virtual bullets, 13 guns) replay
in 2.9 s, ~32k virtual bullets/s -- roughly 70x faster and 100x more samples
than a live gauntlet.
Also adds a per-tick WorldState recorder behind const RecordWorldState
(default off, mirrors the ShotLog idiom) which records the state the bot
ACTUALLY builds, staleness included, rather than true positions -- recording
the latter would hand the guns perfect information and produce flattering
scores.
9 new guard checks (33 total, all passing), including fixture round-trip,
replay determinism, stationary->HeadOn 100%, constant-velocity->Linear>HeadOn,
and the energy-threshold turner crossing at t=41.
Add --max-speed flag to TestBattleRunner (sets defaultTurnsPerSecond=-1 for unlimited TPS).
Add SNNBot_garage/tests/test_bullet_economy.nim: 10-round vs WallsBot, prints per-round and summary stats for tuning.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Document exception handling, zero-value BotResult trap, shared adversary bots
- Add offline parsing example using parseServerOutput
- Skip tests gracefully when JARs missing (guard before suite blocks)
- Fix blocking readLine in runner_process.nim: poll with 50ms sleep + atEnd check
(was preventing timeout enforcement, now blocks correctly during battle)
- Add test task to QBot.nimble and config.nims setup docs to AGENTS.md
- Add debug logging to TestBattleRunner for bot identity tracking
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Implements:
- BattleResult type and JSON-lines parser (#127)
- TR server lifecycle manager (#128)
- Bot compiler using nim c (#129)
- runBattle() orchestrator (#130)
- Example test in OscillatorBot_garage (#131)
- Framework usage guide (#132)
- TestBattleRunner.java for external server (#133)
- BattleRunner process lifecycle (#134)
Fix: runner_process.nim was redefining TimeoutError locally; now
uses std/net.TimeoutError consistently with server_manager.nim.
sac_train.sh orchestrates chunked self-play via tools/training_runner/
RunTraining.java: weighted opponent sampling per chunk, deterministic
eval (SACLSTM_EVAL_MODE=1) every N chunks with win-rate tracking, best
checkpoint (weights/sac_best.zip) by eval score, crash-restart loop on
the runner's liveness detection.
Supporting changes:
- integration.nim: opponentKey() keys the NewBattle buffer-clear rule on
getBotName(id) with numeric-id fallback (#49 Q14 follow-up);
bumpRoundCounter() emits the per-round liveness signal.
- SAC_LSTM_Bot.nim: onRoundEnded -> bumpRoundCounter().
- RunTraining.java: BOT_NAME env parameterizes result matching
(default PPO_Bot, unchanged behavior for PPO).
- Launch packaging: root SAC_LSTM_Bot.json + .sh for the booter;
src json name aligned to 'SAC_LSTM_Bot' so self-reported identity
matches the booted identity (mismatch = runner connect timeout).
Stochastic eval at std≈0.37 was 0/10 vs Corners (deterministic: 10/10).
Warm-start policy is correct but brittle — any noise breaks it.
- log_std initialized to -2.0 (std≈0.135) for moderate exploration
- entropy_coeff=0.0 (no push toward exploration during fine-tuning)
- logStd ceiling=-1.0 (cap at std≈0.37)
- Accumulate transitions across 10 rounds (~3000) before PPO update
(was per-round ~300 — gradient estimates were far too noisy)
- training.nim: MAX_TRANSITIONS 4096→8192, done flag on transitions,
GAE handles episode boundaries correctly
- PPO_Bot.nim: buffer persists across rounds, update every N rounds
- training.env: lr 5e-5→1e-4, entropy 0.001, UPDATE_INTERVAL=10
- Bullet state (indices 44-55): enemy-relative → bot-relative frame
(bot needs threat vectors to itself for dodging, not to enemy)
- New index 56: scan staleness = min(ticksSinceLastScan / 30, 1.0)
(gives policy a confidence signal for enemy data freshness)
- warm_start.py updated: 44→57 dim expansion, TARGET_DIM variable
- Tests updated for new state layout
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
computeRoundReward used cumulative totalScore/50 — unbounded in long battles
(vLoss 353 at round 3160 → 25745 by 3871 in the 5841-round attempt). Cap the
score term at 400 before /50: bonus ∈ [0,8], so the critic's value scale stays
stable regardless of battle length and across battle boundaries.
Reverts the 60-round battle chunking (186e005/da2f825): one battle per
campaign for the whole remaining budget; keeps the crash-restart loop, the
mid-battle freeze guard and the end-of-battle counter completeness check.
Cert (5211, single 60-round battle): 59/60 wins (sole loss = cold-start round
1, score 61), vLoss avg 36.1 / max 148.5, gNorm max 596, zero NaN, zero
restarts, counter check passed. Weights persist to round 5211.
RunTraining exits 0 after each chunk; && break ended the whole run after
the first battle (counter 3219, not 9000). Loop now falls through the
success path and re-checks the persisted counter each iteration.
Round-end reward = cumulative totalScore/50 grows unboundedly with battle
length; long battles (5841 rounds) blew the critic's value scale: vLoss
10-30 during the 60-round cert, 353 at battle-1 round 1, 25745 by round 3871,
policy drift to 0/6 wins. 60-round battles reproduce the certified regime:
bounded value targets, fresh bot process per battle (clears thread state).
The event queue's heap seq was the last GC'd block surviving across
rounds: each round runs on a freshly spawned bot thread, so the N+1
thread realloc'd a block grown by dead thread N's allocator mid-round
(at the next capacity doubling, ~turn 104) -> rawDealloc SIGSEGV in
addEvent (7 gdb-confirmed coredumps). Replace with a static
array[MAX_QUEUE_SIZE, BotEvent] + eventsLen: no heap block crosses
threads, realloc can never happen.
Also fix the harness aborting the final round mid-train: PPO_Bot's
onRoundEnded trains synchronously after the runner's RoundEndedEvent,
so the counter read right after awaitResults() is the stale pre-train
value and System.exit killed the bot inside ppoUpdate. Poll up to 60s
for the counter to catch up before declaring the battle incomplete.
Verified: 72 consecutive rounds vs Fire, 100% wins, all rounds trained
(counter advanced 1:1), zero coredumps since the fix.
Three bugs caused the radar to sweep continuously instead of locking:
1. run() loop set radar to Inf every tick, overwriting any lock
→ replaced with enemy_tracker.getRadarTurnRate()
2. onScannedBot used radarBearingTo() (math convention, east=0 CCW)
→ removed; run loop now handles radar via enemy_tracker
3. enemy_tracker.getRadarTurnRate() had arctan2(dx,dy) instead of
arctan2(dy,dx) — introduced by fd22535; bearing was off by ~90°
Also relaxed stale-lock threshold from 2 to 8 ticks to survive
brief scan gaps without falling back to full sweep.
Added tools/battle_runner for automated 1v1 testing.
Result: 1303/1308 ticks with successful scan (was ~1 in 4).