Replaces the stale 2026-09-20 docs, which predated per-gun real attribution, the offline gun range and the DrussGT boss, and whose verdicts were built on virtual hit rates that turned out to be ANTI-correlated with reality. docs/gun_rack_analysis.md (751 lines) covers: the test infrastructure described honestly (offline range with its flaky-acceptance caveat, the 20 fixtures and what each set is good for, the live boss, and the A/B methodology of per-run server-side real hit rate with an explicit overlap test); the virtual-vs-real metric lesson with Spearman -0.374 and the point-vs-path A/B; per-gun real performance and the 16-rule ranking A/B; the offline per-fixture gun matrix; KEEP/MARGINAL/BELOW verdicts; and the five root-cause bugs with before/after numbers. docs/gun_rack_summary.md (58 lines) is the verdict table plus top actions. The '~230 point' score-noise band that has been steering methodology all night was re-derived from the artifacts rather than asserted: the 13 shipped-config run scores span 175-526, s.d. ~105, i.e. a ~210-point 2-s.d. band. Caveats recorded verbatim rather than softened: the offline==online acceptance is flaky (typically 11/12 on unmodified HEAD), fixtures are perfect-information and therefore optimistic vs live play, per-gun real N is small so single-gun ordering is indicative, the headline numbers come from ONE wave-surfer adversary, and HeadOn must stay despite being lowest because it is the floor fallback (disabling it: 5.08% / 175 dmg vs 6.95% / 251 dmg).
2.6 KiB
Gun Rack Summary — TL;DR
Date: 2026-09-21 · Bot: ModularBot (13 active guns + 1 disabled TM gate)
Shipped config: metric path, selector relative.
Ground truth: server-side real hit rate (events sidecar), per-run, with an
explicit overlap test. Scores are not used — they swing by a couple of hundred
points per run.
Boss: the unmodified DrussGT jar through tools/robocode_shim/.
Headline (13 runs vs DrussGT, 3,612 server-side shots): 6.95% real hit rate,
251 dmg/run.
The one thing to know: virtual hit rate is not a proxy for real hit rate.
For the shipped config Spearman(virtual rank, real rank) = −0.374 —
anti-correlated. 16 candidate ranking rules all failed to beat the shipped
config; what works is the selector's floor/tie hedging (removing the floor:
5.08% / 175 dmg vs 6.95% / 251 dmg). Detail: gun_rack_analysis.md.
Verdict table
| Verdict | Gun | Real % | Virtual % |
|---|---|---|---|
| KEEP | Linear | 10.7 | 10.2 |
| KEEP | Circular | 9.9 | 11.9 |
| KEEP | KNN | 9.0 | 7.5 |
| KEEP | Pattern | 8.6 | 12.0 |
| KEEP | Accel | 7.3 | 12.1 |
| KEEP | AvgLead | 7.0 | 12.3 |
| MARGINAL | GuessFactor | 6.9 | 10.3 |
| MARGINAL | DecayGF | 6.4 | 9.2 |
| MARGINAL | WallBounce | 6.2 | 12.9 |
| MARGINAL | StopShot | 6.1 | 12.6 |
| BELOW | Tsetlin | 5.8 | 12.9 |
| BELOW | Displace | 5.3 | 12.3 |
| FLOOR — KEEP | HeadOn | 5.2 | 8.6 |
| DISABLED | TMSelect | — | — |
HeadOn is lowest but must stay: it is the floor fallback, and disabling the floor measurably hurt (5.08% / 175 dmg).
Top actions
- Stop trusting the virtual metric as a ranker. It is anti-correlated with real hit rate and unstable across run sets. The selector survives on its floor/tie hedge, not on its ordering.
- Fix the offline==online acceptance race. It is currently flaky (typically 11/12), so offline numbers are strong-but-not-exact.
- Get a second adversary. Every per-gun verdict rests on one wave surfer; the KEEP/BELOW boundaries are matchup-specific and per-gun N is small (47–898 shots).
- Decide Tsetlin and Displace. Both are below overall on small N; Tsetlin now learns (clauses 714→13.8 literals) but is not competitive — tune the regression head or drop.
- Do not re-enable TMSelect until the gate-margin/label problem is fixed; it cost 7.47% → 5.59% despite showing real energy structure in its clauses.
- Real-hit-rate-driven selection is not viable yet — unselected guns get near-zero shots, so it needs forced exploration + shrinkage + thousands of shots per gun.