Commit Graph

2 Commits

Author SHA1 Message Date
SirStone 3b5d70b7c3 feat(gun_harness): runtime metric switch + A/B proving the point metric mis-selects
Adds GUN_VBULLET_METRIC (point|path, default point = unchanged behaviour) so
the virtual-bullet hit model can be selected at runtime with no rebuild. Both
the live tracker and the offline replay read the same value, so the 12/12
offline==online acceptance holds under EITHER setting (verified for both).

A/B AGAINST THE LIVE BOSS, real server-side hit rate as ground truth, 5
battles x 12 rounds per metric on one frozen binary:
  point  4660 shots / 219 hits = 4.70%   (per-run 3.16-5.53)
  path   4834 shots / 359 hits = 7.43%   (per-run 6.55-8.24)
The distributions DO NOT OVERLAP: path's worst run beats point's best run.
+2.73pp, +58% relative, z = 5.56, p < 0.0001. Range distributions were
identical (~460-478 px), so this is not a range confound.

MECHANISM - and this is the important part. The gain is SELECTION, not better
gun learning. Under the point model every gun's virtual rate is compressed
into 0.6-4.4%, so HeadOn sits inside the 2pp tie margin and takes 72.6% of
selection ticks / 76.9% of shots - while HeadOn is 11th of 13 by REAL hit rate
(2.3%). The path model widens the band to 4.7-13.7% and ranks HeadOn 10th, so
its shot share falls to 35.9% and Pattern/Accel/WallBounce get picked instead.
Counterfactual: applying the point model's per-gun real rates to the path
model's shot mix yields 7.65%, i.e. essentially the whole observed gain.
So the selector, not the guns, is where the win lives.

PER-GUN REAL HIT RATE vs DrussGT (path mix, the answer to 'which guns are
worth keeping'): WallBounce 10.8, Pattern 10.5, Accel 10.0, Displace 9.3,
Circular 9.2, AvgLead 8.5, KNN 5.7, StopShot 5.2, GuessFactor 3.7,
Tsetlin 2.9. Per-gun N is small (hundreds of shots) so single-gun ordering is
indicative, not definitive.

TWO CAVEATS, recorded because they undercut a naive reading:
1. One adversary. DrussGT is a wave surfer and HeadOn is genuinely bad against
   surfers, so part of this may be matchup-specific.
2. The path model is NOT a better general ranker. Spearman(virtual rank, real
   rank) is 0.52 under point vs -0.04 under path. It wins by accidentally
   fixing HeadOn's mis-rank, not by ranking guns better. A more durable fix is
   to address the selection logic directly - which is the next job.

Also adds a focused guard test (test_vbullet_metric) covering parsing/default,
a receding-target point-miss/path-hit, a perpendicular-target path-miss, and
replay determinism.

Verified: 33 guard checks, 12/12 acceptance under both metrics, tsetlin tests
green, range 34.3% (point, unchanged) / 50.8% (path).
2026-09-21 03:58:27 +02:00
SirStone 974528d5cf feat(gun_harness): offline gun range, proven equivalent to live play
Gun evaluation previously required a full end-to-end battle (Java server +
battle runner + websocket IPC to 2 bot processes, 50 rounds, ~3.4 min) and
yielded only ~300-900 REAL shots across 13 guns -- far too few to rank
guns, which is why tuning needed many repetitions.

VirtualTracker is already a pure function of (WorldState stream, gun list);
the only reason it needed Java was where WorldState came from. So the range
replays a seq[WorldState] through the SAME tracker: offline and online
scores are the same metric by construction, not an approximation.

ACCEPTANCE TEST (the point of the whole thing): record one live round, replay
it offline, compare per-gun virtual hit rates. 12/12 deterministic guns match
EXACTLY, reproduced twice. Tsetlin is compared separately because tmLearnOne
calls rand(). Getting to 12/12 exposed two real ordering quirks in the live
loop: run() calls go() before the aim/fire block, so tickBullets resolves
against the NEXT tick's scan while the prediction used the previous one; and
if the target dies during that go() the final tick's spawn+resolution is
skipped entirely. The recorder emits an end marker for the second case.
The 5th (selected-gun) predict call was verified to be a no-op.

Measured cost: 8 fixtures (1770 ticks, ~92k virtual bullets, 13 guns) replay
in 2.9 s, ~32k virtual bullets/s -- roughly 70x faster and 100x more samples
than a live gauntlet.

Also adds a per-tick WorldState recorder behind const RecordWorldState
(default off, mirrors the ShotLog idiom) which records the state the bot
ACTUALLY builds, staleness included, rather than true positions -- recording
the latter would hand the guns perfect information and produce flattering
scores.

9 new guard checks (33 total, all passing), including fixture round-trip,
replay determinism, stationary->HeadOn 100%, constant-velocity->Linear>HeadOn,
and the energy-threshold turner crossing at t=41.
2026-09-20 23:44:42 +02:00