Files
SirRoboGarage/docs/gun_rack_summary.md
T
SirStone 19410164f1 docs: definitive gun-rack report on real measured numbers
Replaces the stale 2026-09-20 docs, which predated per-gun real attribution,
the offline gun range and the DrussGT boss, and whose verdicts were built on
virtual hit rates that turned out to be ANTI-correlated with reality.

docs/gun_rack_analysis.md (751 lines) covers: the test infrastructure
described honestly (offline range with its flaky-acceptance caveat, the 20
fixtures and what each set is good for, the live boss, and the A/B methodology
of per-run server-side real hit rate with an explicit overlap test); the
virtual-vs-real metric lesson with Spearman -0.374 and the point-vs-path A/B;
per-gun real performance and the 16-rule ranking A/B; the offline per-fixture
gun matrix; KEEP/MARGINAL/BELOW verdicts; and the five root-cause bugs with
before/after numbers.
docs/gun_rack_summary.md (58 lines) is the verdict table plus top actions.

The '~230 point' score-noise band that has been steering methodology all night
was re-derived from the artifacts rather than asserted: the 13 shipped-config
run scores span 175-526, s.d. ~105, i.e. a ~210-point 2-s.d. band.

Caveats recorded verbatim rather than softened: the offline==online acceptance
is flaky (typically 11/12 on unmodified HEAD), fixtures are perfect-information
and therefore optimistic vs live play, per-gun real N is small so single-gun
ordering is indicative, the headline numbers come from ONE wave-surfer
adversary, and HeadOn must stay despite being lowest because it is the floor
fallback (disabling it: 5.08% / 175 dmg vs 6.95% / 251 dmg).
2026-09-21 06:37:20 +02:00

2.6 KiB
Raw Blame History

Gun Rack Summary — TL;DR

Date: 2026-09-21 · Bot: ModularBot (13 active guns + 1 disabled TM gate) Shipped config: metric path, selector relative. Ground truth: server-side real hit rate (events sidecar), per-run, with an explicit overlap test. Scores are not used — they swing by a couple of hundred points per run.

Boss: the unmodified DrussGT jar through tools/robocode_shim/. Headline (13 runs vs DrussGT, 3,612 server-side shots): 6.95% real hit rate, 251 dmg/run.

The one thing to know: virtual hit rate is not a proxy for real hit rate. For the shipped config Spearman(virtual rank, real rank) = −0.374 — anti-correlated. 16 candidate ranking rules all failed to beat the shipped config; what works is the selector's floor/tie hedging (removing the floor: 5.08% / 175 dmg vs 6.95% / 251 dmg). Detail: gun_rack_analysis.md.

Verdict table

Verdict Gun Real % Virtual %
KEEP Linear 10.7 10.2
KEEP Circular 9.9 11.9
KEEP KNN 9.0 7.5
KEEP Pattern 8.6 12.0
KEEP Accel 7.3 12.1
KEEP AvgLead 7.0 12.3
MARGINAL GuessFactor 6.9 10.3
MARGINAL DecayGF 6.4 9.2
MARGINAL WallBounce 6.2 12.9
MARGINAL StopShot 6.1 12.6
BELOW Tsetlin 5.8 12.9
BELOW Displace 5.3 12.3
FLOOR — KEEP HeadOn 5.2 8.6
DISABLED TMSelect — —

HeadOn is lowest but must stay: it is the floor fallback, and disabling the floor measurably hurt (5.08% / 175 dmg).

Top actions

  1. Stop trusting the virtual metric as a ranker. It is anti-correlated with real hit rate and unstable across run sets. The selector survives on its floor/tie hedge, not on its ordering.
  2. Fix the offline==online acceptance race. It is currently flaky (typically 11/12), so offline numbers are strong-but-not-exact.
  3. Get a second adversary. Every per-gun verdict rests on one wave surfer; the KEEP/BELOW boundaries are matchup-specific and per-gun N is small (47–898 shots).
  4. Decide Tsetlin and Displace. Both are below overall on small N; Tsetlin now learns (clauses 714→13.8 literals) but is not competitive — tune the regression head or drop.
  5. Do not re-enable TMSelect until the gate-margin/label problem is fixed; it cost 7.47% → 5.59% despite showing real energy structure in its clauses.
  6. Real-hit-rate-driven selection is not viable yet — unselected guns get near-zero shots, so it needs forced exploration + shrinkage + thousands of shots per gun.