Commit Graph

1 Commits

Author SHA1 Message Date
SirStone ab86c0481f gun audit: the virtual system is sound; the rack is redundant, not broken
Audited all 14 guns offline over the committed DrussGT fixtures (~150k resolved
bullets/gun) plus 123 rounds of live gun_stats. Prompted by a GUI observation
that selected guns "fire dozens of pixels away" and a suspicion of reverse
selection.

MY HYPOTHESIS WAS WRONG. I expected guns to be ignoring `bulletSpeed`, which
would make their 4 power bins identical and the per-bin fitness pure noise.
MEASURED: only `HeadOn` is speed-blind (100% identical bins) and that is its
correct definition. Every other gun emits 91-94% DISTINCT per-bin predictions
(mean intra-tick bin spread 62-103px). The earlier "tick-only cache collapsed
all bins onto bin 0" fix is complete across the whole rack.

LEAD/SIGN/UNITS ARE CORRECT: replaying synthetic ground truth, all 14 guns score
100% on a stationary target (which also proves predictions are ABSOLUTE - a
relative or angle return would score 0), ~100% on constant-velocity for every
leaded gun, Circular 100% / Accel 99.8% on a 3deg/tick circle, WallBounce 98.2%
on a bounce. No missing lead, no sign inversion. Resolution is right (BotRadius
18, hit credited to the owning gun).

FEEDBACK IS INTACT: offline pushes=151260/starved=0, Tsetlin trained=149205/
traceMisses=0; live `vStarved=0` and `vDropped=0` across all 123 rounds.

THE "43% FLAT / 4x OPTIMISTIC" EVIDENCE I CITED IS NOT REPRODUCIBLE on current
code/data. Live virtual/real ratios against DrussGT are 0.8-2.0 for most guns
(Linear 10.9 virt / 13.4 real; KNN 8.1/7.1; DecayGF 10.5/9.8). The 43%-flat
session matches an older config or a weak opponent (SittingDuck), not DrussGT.
`bmPath` IS 2.3-3.6x `bmPoint` - but by design and documented: it asks "does the
ray eventually sweep the target's path", a deliberately generous relative
signal. So the flat tie is a RANKING artefact: many guns share the same base
forecast and, with a near-zero learned correction, collapse onto the same ray;
RelTieMargin=0.20 then treats the top ~half of the rack as tied.

DUTY AND OVERLAP (>=50% of ticks within 20px = redundant):
  Tsetlin   ~ StopShot 87%            -> duplicate pair
  DecayGF   ~ GuessFactor 92%         -> duplicate pair
  Accel     ~ Circular 65%            -> partial duplicate
  WallBounce~ Linear 58%
  AvgLead   = the MEAN of Linear+Circular+WallBounce (constructed redundancy)
  Displace  worst point% (6.6) AND worst real% (2.4); wins no bucket
  TMSelect  DEAD - never spawned (EnableTmSelector=false), 0 shots in every log
  Pattern   the ONLY gun competitive in every distance/speed bucket
  HeadOn/Linear/Tsetlin/StopShot are identical copies of each other on a real
  surfer (v<1 ~50.8%, everything else ~2%)

RECOMMENDED LEAN RACK (8): HeadOn, Linear, Circular, Accel, Pattern,
GuessFactor, KNN, WallBounce.
DROP (6): TMSelect (dead), AvgLead (constructed mean), Displace (worst), DecayGF
(92% GF), StopShot (87% Tsetlin), Tsetlin (the repo's own sweep already showed
it learns nothing on DrussGT).

HONEST HEADLINE: pruning is NOT expected to raise hit rate - an earlier
15-paired-run experiment found it neutral-to-negative (p=0.57/0.21). The
mechanism by which it could help is a SELECTOR effect (shrinking the tied band),
not a gun effect, and that is UNVERIFIED until A/B'd. The virtual system and the
rack are basically sound; the lever that matters most is the selector's
metric/tie-band, not deleting guns.

DESIGN SMELL FOUND (INFERRED, not measured): GF/DecayGF/KNN `onResult` pops the
OLDEST wave, but under bmPath bullets leave the arena in non-FIFO order, so a
resolution can be paired with a neighbouring tick's wave. starved=0 does not
rule this out. Candidate fix: key waves by fireTick, as Tsetlin/TMSelect do.

Adds common_libs/tests/audit_virtual_guns.nim (offline, no shipped file touched).
2026-09-22 00:44:19 +02:00