19410164f1
Replaces the stale 2026-09-20 docs, which predated per-gun real attribution, the offline gun range and the DrussGT boss, and whose verdicts were built on virtual hit rates that turned out to be ANTI-correlated with reality. docs/gun_rack_analysis.md (751 lines) covers: the test infrastructure described honestly (offline range with its flaky-acceptance caveat, the 20 fixtures and what each set is good for, the live boss, and the A/B methodology of per-run server-side real hit rate with an explicit overlap test); the virtual-vs-real metric lesson with Spearman -0.374 and the point-vs-path A/B; per-gun real performance and the 16-rule ranking A/B; the offline per-fixture gun matrix; KEEP/MARGINAL/BELOW verdicts; and the five root-cause bugs with before/after numbers. docs/gun_rack_summary.md (58 lines) is the verdict table plus top actions. The '~230 point' score-noise band that has been steering methodology all night was re-derived from the artifacts rather than asserted: the 13 shipped-config run scores span 175-526, s.d. ~105, i.e. a ~210-point 2-s.d. band. Caveats recorded verbatim rather than softened: the offline==online acceptance is flaky (typically 11/12 on unmodified HEAD), fixtures are perfect-information and therefore optimistic vs live play, per-gun real N is small so single-gun ordering is indicative, the headline numbers come from ONE wave-surfer adversary, and HeadOn must stay despite being lowest because it is the floor fallback (disabling it: 5.08% / 175 dmg vs 6.95% / 251 dmg).
59 lines
2.6 KiB
Markdown
59 lines
2.6 KiB
Markdown
# Gun Rack Summary — TL;DR
|
||
|
||
**Date:** 2026-09-21 · **Bot:** ModularBot (13 active guns + 1 disabled TM gate)
|
||
**Shipped config:** metric `path`, selector `relative`.
|
||
**Ground truth:** server-side real hit rate (events sidecar), per-run, with an
|
||
explicit overlap test. Scores are not used — they swing by a couple of hundred
|
||
points per run.
|
||
|
||
**Boss:** the unmodified DrussGT jar through `tools/robocode_shim/`.
|
||
**Headline (13 runs vs DrussGT, 3,612 server-side shots): 6.95% real hit rate,
|
||
251 dmg/run.**
|
||
|
||
**The one thing to know:** *virtual hit rate is not a proxy for real hit rate.*
|
||
For the shipped config Spearman(virtual rank, real rank) = **−0.374** —
|
||
anti-correlated. 16 candidate ranking rules all failed to beat the shipped
|
||
config; what works is the selector's floor/tie hedging (removing the floor:
|
||
5.08% / 175 dmg vs 6.95% / 251 dmg). Detail: [`gun_rack_analysis.md`](gun_rack_analysis.md).
|
||
|
||
## Verdict table
|
||
|
||
| Verdict | Gun | Real % | Virtual % |
|
||
|---|---|---:|---:|
|
||
| **KEEP** | Linear | 10.7 | 10.2 |
|
||
| **KEEP** | Circular | 9.9 | 11.9 |
|
||
| **KEEP** | KNN | 9.0 | 7.5 |
|
||
| **KEEP** | Pattern | 8.6 | 12.0 |
|
||
| **KEEP** | Accel | 7.3 | 12.1 |
|
||
| **KEEP** | AvgLead | 7.0 | 12.3 |
|
||
| MARGINAL | GuessFactor | 6.9 | 10.3 |
|
||
| MARGINAL | DecayGF | 6.4 | 9.2 |
|
||
| MARGINAL | WallBounce | 6.2 | 12.9 |
|
||
| MARGINAL | StopShot | 6.1 | 12.6 |
|
||
| BELOW | Tsetlin | 5.8 | 12.9 |
|
||
| BELOW | Displace | 5.3 | 12.3 |
|
||
| **FLOOR — KEEP** | HeadOn | 5.2 | 8.6 |
|
||
| DISABLED | TMSelect | — | — |
|
||
|
||
HeadOn is lowest but **must stay**: it is the floor fallback, and disabling the
|
||
floor measurably hurt (5.08% / 175 dmg).
|
||
|
||
## Top actions
|
||
|
||
1. **Stop trusting the virtual metric as a ranker.** It is anti-correlated with
|
||
real hit rate and unstable across run sets. The selector survives on its
|
||
floor/tie hedge, not on its ordering.
|
||
2. **Fix the offline==online acceptance race.** It is currently flaky
|
||
(typically 11/12), so offline numbers are strong-but-not-exact.
|
||
3. **Get a second adversary.** Every per-gun verdict rests on one wave surfer;
|
||
the KEEP/BELOW boundaries are matchup-specific and per-gun N is small
|
||
(47–898 shots).
|
||
4. **Decide Tsetlin and Displace.** Both are below overall on small N; Tsetlin
|
||
now learns (clauses 714→13.8 literals) but is not competitive — tune the
|
||
regression head or drop.
|
||
5. **Do not re-enable TMSelect** until the gate-margin/label problem is fixed;
|
||
it cost 7.47% → 5.59% despite showing real energy structure in its clauses.
|
||
6. **Real-hit-rate-driven selection is not viable yet** — unselected guns get
|
||
near-zero shots, so it needs forced exploration + shrinkage + thousands of
|
||
shots per gun.
|