3 Commits

Author SHA1 Message Date
SirStone 17c99f542f feat(fixtures): closed-loop DrussGT captures from real Tank Royale battles
The classic captures were OPEN-LOOP: replayed DrussGT never dodged OUR
bullets. These come from real TR battles through the working Java bridge, so
the recording contains genuine reactions to ModularBot's live fire. The
open-loop caveat is gone (perfect-information remains).

PRIMARY RESULT - the boss beats us badly. DrussGT 1447 - ModularBot 300 over
15 rounds, ModularBot winning only round 5 (DrussGT died at tick 1893). Rounds
are long, not truncated: mean 1335 ticks, ModularBot got off 1134 shots.
  ModularBot  1134 shots /  60 hits =  5.3% real hit rate
  DrussGT     1400 shots / 169 hits = 12.1% real hit rate
So DrussGT's gun is ~2.3x more accurate than our entire rack, on top of far
better movement. That is the number to move.

Also captured: shield-on variant (DrussGT 939-287, 9/10 - ModularBot takes
round 1 to the known shield warm-up), and vs SpinBot 1175-0, Crazy 1080-1,
Corners 1659-0. 20,026 + 12,629 + 10,824 + 11,507 + 2,575 ticks.

Movement statistics match the classic set within ~0.04 on the perpendicular
and radial fractions, so this is the same wave surfer in TR physics:
  TR vs modularbot: perp 0.967, radial 0.001, 52.8% at full speed,
                    reversing 46.3%, median range 464 px.

CLOSED LOOP PROVEN, not asserted. ModularBot's fire is a heat-limited near
metronome (median interval 14 ticks), which gives a usable exogenous clock:
  - event-locked |delta heading| oscillates 0.96 -> 2.69 deg about a 1.47 deg
    mean with the fire period, almost every lag outside the 95% band of a
    400-iteration phase-shuffled null;
  - cross-correlation of |delta heading| against the fire impulse peaks at
    r = +0.111, lag 12 ticks, permutation p = 0.005 (null peak mean +0.016);
  - OWN-FIRE CONTROL is flat, so the oscillation is enemy-driven rather than
    an internal cadence;
  - range response is weak (~3 px over 30 ticks, near noise) and is therefore
    NOT claimed, and per-bullet dodging is not claimed either because the
    bullet detector (shield) is off.

CAVEATS: still perfect-information (observer gives true positions every tick,
unlike the live bot's stale between-scan WorldState) so these remain optimistic
vs live play; and they are open-loop AT REPLAY TIME - 'closed_loop' describes
the capture, not a later replay. TR conversion residual is ~1.5 deg mean
because the TR server moves along the pre-turn heading, vs 0.000 deg for the
classic captures.

Adds analyze_closed_loop.py (PSTH event-locking, phase-shuffle permutation
null, cross-correlation, own-fire control) and per-round result sidecars.
2026-09-21 01:18:13 +02:00
SirStone 4f18c8ce07 feat(tools): capture real DrussGT movement from classic Robocode as fixtures
There are no genuinely competitive adversaries for Tank Royale, and the
in-repo ones were broken until recently. Classic Robocode 1.9.5.5 is
obtainable (SourceForge, 20.4 MB) and its programmatic control API
(robocode.control.RobocodeEngine + BattleAdaptor.onTurnEnded) can run real
battles headless and expose per-turn robot state. So a legacy leader bot's
MOVEMENT can be captured and used as a gun-testing fixture with no port.

Captured unmodified DrussGT 3.1.4159 vs spinbot/ramfire/crazy/corners and a
mirror match: 28,797 ticks, plus two trivial-bot contrasts. Conversion to
the Tank Royale convention is validated to 0.000-0.001 deg by recomputing
the direction implied by (heading, speed) and comparing it against the
recorded per-tick displacement -- i.e. the data is proven to be genuine
recorded motion rather than a mangled export. (A first attempt treated the
snapshot API's headings as degrees; they are radians, ~95 deg off.)

The statistics confirm it is really a wave surfer: perpendicular to the
opponent 65-96% of ticks, radial ~0.001, 41-72% of ticks at full speed,
reversing on 42-46% of ticks, holding range at a 283-526 px median. The
straight-line contrast is radial-dominant (0.75) with ZERO reversals.

Discovery: DrussGT detects predictable guns and switches to a bullet-shield
stand-still mode, so captures against sample.Walls/TrackFire had to be
rejected as non-movement.

CAVEATS, recorded in DRUSSGT_FIXTURES.md: these are open-loop (replayed
DrussGT never dodges OUR bullets) and perfect-information (the observer
gives true positions every tick, unlike our stale live WorldState). Both
make our guns look better than in live play, so use them for RELATIVE gun
ranking, not absolute hit rates.

Jars stay out of git; capture tooling is reproducible via capture.sh.
2026-09-20 23:44:42 +02:00
SirStone 974528d5cf feat(gun_harness): offline gun range, proven equivalent to live play
Gun evaluation previously required a full end-to-end battle (Java server +
battle runner + websocket IPC to 2 bot processes, 50 rounds, ~3.4 min) and
yielded only ~300-900 REAL shots across 13 guns -- far too few to rank
guns, which is why tuning needed many repetitions.

VirtualTracker is already a pure function of (WorldState stream, gun list);
the only reason it needed Java was where WorldState came from. So the range
replays a seq[WorldState] through the SAME tracker: offline and online
scores are the same metric by construction, not an approximation.

ACCEPTANCE TEST (the point of the whole thing): record one live round, replay
it offline, compare per-gun virtual hit rates. 12/12 deterministic guns match
EXACTLY, reproduced twice. Tsetlin is compared separately because tmLearnOne
calls rand(). Getting to 12/12 exposed two real ordering quirks in the live
loop: run() calls go() before the aim/fire block, so tickBullets resolves
against the NEXT tick's scan while the prediction used the previous one; and
if the target dies during that go() the final tick's spawn+resolution is
skipped entirely. The recorder emits an end marker for the second case.
The 5th (selected-gun) predict call was verified to be a no-op.

Measured cost: 8 fixtures (1770 ticks, ~92k virtual bullets, 13 guns) replay
in 2.9 s, ~32k virtual bullets/s -- roughly 70x faster and 100x more samples
than a live gauntlet.

Also adds a per-tick WorldState recorder behind const RecordWorldState
(default off, mirrors the ShotLog idiom) which records the state the bot
ACTUALLY builds, staleness included, rather than true positions -- recording
the latter would hand the guns perfect information and produce flattering
scores.

9 new guard checks (33 total, all passing), including fixture round-trip,
replay determinism, stationary->HeadOn 100%, constant-velocity->Linear>HeadOn,
and the energy-threshold turner crossing at t=41.
2026-09-20 23:44:42 +02:00