Files
SirRoboGarage/docs/gun_rack_summary.md
T
SirStone 19410164f1 docs: definitive gun-rack report on real measured numbers
Replaces the stale 2026-09-20 docs, which predated per-gun real attribution,
the offline gun range and the DrussGT boss, and whose verdicts were built on
virtual hit rates that turned out to be ANTI-correlated with reality.

docs/gun_rack_analysis.md (751 lines) covers: the test infrastructure
described honestly (offline range with its flaky-acceptance caveat, the 20
fixtures and what each set is good for, the live boss, and the A/B methodology
of per-run server-side real hit rate with an explicit overlap test); the
virtual-vs-real metric lesson with Spearman -0.374 and the point-vs-path A/B;
per-gun real performance and the 16-rule ranking A/B; the offline per-fixture
gun matrix; KEEP/MARGINAL/BELOW verdicts; and the five root-cause bugs with
before/after numbers.
docs/gun_rack_summary.md (58 lines) is the verdict table plus top actions.

The '~230 point' score-noise band that has been steering methodology all night
was re-derived from the artifacts rather than asserted: the 13 shipped-config
run scores span 175-526, s.d. ~105, i.e. a ~210-point 2-s.d. band.

Caveats recorded verbatim rather than softened: the offline==online acceptance
is flaky (typically 11/12 on unmodified HEAD), fixtures are perfect-information
and therefore optimistic vs live play, per-gun real N is small so single-gun
ordering is indicative, the headline numbers come from ONE wave-surfer
adversary, and HeadOn must stay despite being lowest because it is the floor
fallback (disabling it: 5.08% / 175 dmg vs 6.95% / 251 dmg).
2026-09-21 06:37:20 +02:00

59 lines
2.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Gun Rack Summary — TL;DR
**Date:** 2026-09-21 · **Bot:** ModularBot (13 active guns + 1 disabled TM gate)
**Shipped config:** metric `path`, selector `relative`.
**Ground truth:** server-side real hit rate (events sidecar), per-run, with an
explicit overlap test. Scores are not used — they swing by a couple of hundred
points per run.
**Boss:** the unmodified DrussGT jar through `tools/robocode_shim/`.
**Headline (13 runs vs DrussGT, 3,612 server-side shots): 6.95% real hit rate,
251 dmg/run.**
**The one thing to know:** *virtual hit rate is not a proxy for real hit rate.*
For the shipped config Spearman(virtual rank, real rank) = **−0.374** —
anti-correlated. 16 candidate ranking rules all failed to beat the shipped
config; what works is the selector's floor/tie hedging (removing the floor:
5.08% / 175 dmg vs 6.95% / 251 dmg). Detail: [`gun_rack_analysis.md`](gun_rack_analysis.md).
## Verdict table
| Verdict | Gun | Real % | Virtual % |
|---|---|---:|---:|
| **KEEP** | Linear | 10.7 | 10.2 |
| **KEEP** | Circular | 9.9 | 11.9 |
| **KEEP** | KNN | 9.0 | 7.5 |
| **KEEP** | Pattern | 8.6 | 12.0 |
| **KEEP** | Accel | 7.3 | 12.1 |
| **KEEP** | AvgLead | 7.0 | 12.3 |
| MARGINAL | GuessFactor | 6.9 | 10.3 |
| MARGINAL | DecayGF | 6.4 | 9.2 |
| MARGINAL | WallBounce | 6.2 | 12.9 |
| MARGINAL | StopShot | 6.1 | 12.6 |
| BELOW | Tsetlin | 5.8 | 12.9 |
| BELOW | Displace | 5.3 | 12.3 |
| **FLOOR — KEEP** | HeadOn | 5.2 | 8.6 |
| DISABLED | TMSelect | — | — |
HeadOn is lowest but **must stay**: it is the floor fallback, and disabling the
floor measurably hurt (5.08% / 175 dmg).
## Top actions
1. **Stop trusting the virtual metric as a ranker.** It is anti-correlated with
real hit rate and unstable across run sets. The selector survives on its
floor/tie hedge, not on its ordering.
2. **Fix the offline==online acceptance race.** It is currently flaky
(typically 11/12), so offline numbers are strong-but-not-exact.
3. **Get a second adversary.** Every per-gun verdict rests on one wave surfer;
the KEEP/BELOW boundaries are matchup-specific and per-gun N is small
(47–898 shots).
4. **Decide Tsetlin and Displace.** Both are below overall on small N; Tsetlin
now learns (clauses 714→13.8 literals) but is not competitive — tune the
regression head or drop.
5. **Do not re-enable TMSelect** until the gate-margin/label problem is fixed;
it cost 7.47% → 5.59% despite showing real energy structure in its clauses.
6. **Real-hit-rate-driven selection is not viable yet** — unselected guns get
near-zero shots, so it needs forced exploration + shrinkage + thousands of
shots per gun.