Commit Graph

3 Commits

Author SHA1 Message Date
SirStone 013b9fe01e docs: correct the overstated 'inverted metric' claim; record tonight's fixes
Three corrections, all prompted by later measurements:

1. The virtual-vs-real rank correlation is NOT robustly negative. Six
   independent Spearman measurements now exist (-0.374, +0.335, +0.522,
   -0.371, -0.073, -0.037) and the sign flips on large samples, so it is near
   zero on average. The honest headline is that virtual hit rate is a POOR
   RANKER, not an inverted one. The report said 'not weak - it is inverted' in
   six places; it now says so in none. The practical conclusion (do not trust
   it for ranking) is unchanged; the mechanism claimed was wrong.
2. The offline==online acceptance is FIXED, not flaky. Root cause was that the
   replay spawned gun 13 (TMSelect) while the live rack has it disabled, and
   the shared VirtualTracker ring is ORDER-SENSITIVE, so gun 13's extra 4
   bullets/tick permuted the per-tick resolution order for every other gun and
   shifted the learning guns' observations. After closing gun 13's ready gate
   offline the live and offline KNN traces are byte-identical (904/904 lines,
   empty diff). 5/5 consecutive runs now report 12/12 exact with the death
   boundary included. Recorded with the lesson: a flaky proof was hiding a real
   bug. Also records the general A/B confound - disabling a gun removes its 4
   spawns/tick from the shared ring, perturbing resolution order for the rest.
3. Pruning was tested and does NOT help, so the verdict for Tsetlin and
   Displace changes from an implied drop to BELOW OVERALL - KEEP. 15 paired
   runs: baseline 6.18%, Tsetlin-off 5.76%, Tsetlin+Displace-off 5.46%;
   paired permutation p=0.57 and p=0.21; distributions completely overlap; a
   non-surfer control showed no separation. Being below average does not
   justify removal.

Also records the tie-break randomness fix, and quotes run counts with every
rate (6.95% over 13 runs vs 6.18% over 15 runs, same binary) rather than
presenting a single figure as definitive.
2026-09-21 06:59:00 +02:00
SirStone 19410164f1 docs: definitive gun-rack report on real measured numbers
Replaces the stale 2026-09-20 docs, which predated per-gun real attribution,
the offline gun range and the DrussGT boss, and whose verdicts were built on
virtual hit rates that turned out to be ANTI-correlated with reality.

docs/gun_rack_analysis.md (751 lines) covers: the test infrastructure
described honestly (offline range with its flaky-acceptance caveat, the 20
fixtures and what each set is good for, the live boss, and the A/B methodology
of per-run server-side real hit rate with an explicit overlap test); the
virtual-vs-real metric lesson with Spearman -0.374 and the point-vs-path A/B;
per-gun real performance and the 16-rule ranking A/B; the offline per-fixture
gun matrix; KEEP/MARGINAL/BELOW verdicts; and the five root-cause bugs with
before/after numbers.
docs/gun_rack_summary.md (58 lines) is the verdict table plus top actions.

The '~230 point' score-noise band that has been steering methodology all night
was re-derived from the artifacts rather than asserted: the 13 shipped-config
run scores span 175-526, s.d. ~105, i.e. a ~210-point 2-s.d. band.

Caveats recorded verbatim rather than softened: the offline==online acceptance
is flaky (typically 11/12 on unmodified HEAD), fixtures are perfect-information
and therefore optimistic vs live play, per-gun real N is small so single-gun
ordering is indicative, the headline numbers come from ONE wave-surfer
adversary, and HeadOn must stay despite being lowest because it is the floor
fallback (disabling it: 5.08% / 175 dmg vs 6.95% / 251 dmg).
2026-09-21 06:37:20 +02:00
SirStone c034eb9d25 feat(testing): gun rack gauntlet + analysis reports
- fix(ModularBot): onBulletHitBot → onBulletHit (real hits were never tracked)
- feat(ModularBot): per-round gun stats dump to /tmp/gun_stats.jsonl
- feat(ModularBot): gun selection counter per round
- fix(tests): adversary paths _garage suffix removed from 7 test files
- feat(tests): test_gauntlet_5bots.nim — 10-round gauntlet vs all 5 adversaries
- feat(tests): analyze_gun_stats.nim — JSONL parser for gun performance tables
- docs: gun_rack_analysis.md — full per-gun performance report
- docs: gun_rack_summary.md — TL;DR verdict table (keep/drop/tune)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-09-20 21:14:12 +02:00