docs: definitive gun-rack report on real measured numbers

Replaces the stale 2026-09-20 docs, which predated per-gun real attribution,
the offline gun range and the DrussGT boss, and whose verdicts were built on
virtual hit rates that turned out to be ANTI-correlated with reality.

docs/gun_rack_analysis.md (751 lines) covers: the test infrastructure
described honestly (offline range with its flaky-acceptance caveat, the 20
fixtures and what each set is good for, the live boss, and the A/B methodology
of per-run server-side real hit rate with an explicit overlap test); the
virtual-vs-real metric lesson with Spearman -0.374 and the point-vs-path A/B;
per-gun real performance and the 16-rule ranking A/B; the offline per-fixture
gun matrix; KEEP/MARGINAL/BELOW verdicts; and the five root-cause bugs with
before/after numbers.
docs/gun_rack_summary.md (58 lines) is the verdict table plus top actions.

The '~230 point' score-noise band that has been steering methodology all night
was re-derived from the artifacts rather than asserted: the 13 shipped-config
run scores span 175-526, s.d. ~105, i.e. a ~210-point 2-s.d. band.

Caveats recorded verbatim rather than softened: the offline==online acceptance
is flaky (typically 11/12 on unmodified HEAD), fixtures are perfect-information
and therefore optimistic vs live play, per-gun real N is small so single-gun
ordering is indicative, the headline numbers come from ONE wave-surfer
adversary, and HeadOn must stay despite being lowest because it is the floor
fallback (disabling it: 5.08% / 175 dmg vs 6.95% / 251 dmg).
This commit is contained in:
2026-09-21 06:37:20 +02:00
parent 2c94dc221a
commit 19410164f1
2 changed files with 763 additions and 241 deletions
+50 -32
View File
@@ -1,40 +1,58 @@
# Gun Rack Summary — TL;DR
**Gauntlet:** 5 adversaries × 10 rounds, 14 guns. Data: `/tmp/gun_stats.jsonl`
**Date:** 2026-09-21 · **Bot:** ModularBot (13 active guns + 1 disabled TM gate)
**Shipped config:** metric `path`, selector `relative`.
**Ground truth:** server-side real hit rate (events sidecar), per-run, with an
explicit overlap test. Scores are not used — they swing by a couple of hundred
points per run.
## Scoreboard
**Boss:** the unmodified DrussGT jar through `tools/robocode_shim/`.
**Headline (13 runs vs DrussGT, 3,612 server-side shots): 6.95% real hit rate,
251 dmg/run.**
| Adversary | Score | Real Hit% | Best Gun |
|-----------|-------|-----------|----------|
| SittingDuck | 1944 | 85% | AvgLead / WallBounce (99% vhit) |
| OscillatorBot | 1645 | 59% | AvgLead (65% vhit) |
| RandomMover | 1916 | 61% | Linear / WallBounce / AvgLead (82%) |
| PatternMover | 1914 | 73% | AvgLead / Displace / Accel (96%) |
| WaveSurfer | 1873 | 72% | AvgLead / StopShot / Accel (86%) |
**The one thing to know:** *virtual hit rate is not a proxy for real hit rate.*
For the shipped config Spearman(virtual rank, real rank) = **−0.374** —
anti-correlated. 16 candidate ranking rules all failed to beat the shipped
config; what works is the selector's floor/tie hedging (removing the floor:
5.08% / 175 dmg vs 6.95% / 251 dmg). Detail: [`gun_rack_analysis.md`](gun_rack_analysis.md).
## Verdict Table
## Verdict table
| Gun | Best Use | Verdict |
|-----|----------|---------|
| **HeadOn** | Stationary / pattern bots | KEEP — reliable floor, over-selected |
| **Linear** | Random movers | KEEP — under-selected despite 82% |
| **Tsetlin** | Unknown | TUNE — 0 selection ticks ever; investigate |
| **Circular** | Orbit-heavy bots | KEEP (marginal) |
| **GuessFactor** | General | KEEP — underperforms vs expectation; tune bins |
| **Pattern** | Pattern movers | TUNE — catastrophic vs OscillatorBot (9%, 1081 ticks); raise MinObs gate |
| **AntiSurf** | — | DROP — 0% vhit every adversary including stationary; broken |
| **WallBounce** | All types | KEEP — most consistent overall |
| **Accel** | Pattern / wave bots | KEEP — avoid vs random (19% cliff) |
| **StopShot** | All types | KEEP — top rates, chronically under-selected |
| **Displace** | Pattern / wave bots | KEEP |
| **AvgLead** | All types | KEEP — best all-rounder in the rack |
| **DecayGF** | Wave / pattern bots | KEEP |
| **KNN** | Pattern / wave bots | KEEP — improves with history |
| Verdict | Gun | Real % | Virtual % |
|---|---|---:|---:|
| **KEEP** | Linear | 10.7 | 10.2 |
| **KEEP** | Circular | 9.9 | 11.9 |
| **KEEP** | KNN | 9.0 | 7.5 |
| **KEEP** | Pattern | 8.6 | 12.0 |
| **KEEP** | Accel | 7.3 | 12.1 |
| **KEEP** | AvgLead | 7.0 | 12.3 |
| MARGINAL | GuessFactor | 6.9 | 10.3 |
| MARGINAL | DecayGF | 6.4 | 9.2 |
| MARGINAL | WallBounce | 6.2 | 12.9 |
| MARGINAL | StopShot | 6.1 | 12.6 |
| BELOW | Tsetlin | 5.8 | 12.9 |
| BELOW | Displace | 5.3 | 12.3 |
| **FLOOR — KEEP** | HeadOn | 5.2 | 8.6 |
| DISABLED | TMSelect | — | — |
## Top Actions
HeadOn is lowest but **must stay**: it is the floor fallback, and disabling the
floor measurably hurt (5.08% / 175 dmg).
1. **Fix AntiSurf** — 0% vhit against SittingDuck means wrong angle computation, not just weak targeting.
2. **Raise Pattern's MinObsBeforeCompete** — 15 is too low; 1081 wasted ticks at 9% hit rate vs OscillatorBot.
3. **Re-run gauntlet** with `onBulletHit` fix to get clean per-gun real hit attribution.
4. **AvgLead** is the star gun — confirm it stays in the top selection tier.
5. **HeadOn selection dominance** is a selector bias problem, not a gun problem — add ±2% tiebreak randomisation.
## Top actions
1. **Stop trusting the virtual metric as a ranker.** It is anti-correlated with
real hit rate and unstable across run sets. The selector survives on its
floor/tie hedge, not on its ordering.
2. **Fix the offline==online acceptance race.** It is currently flaky
(typically 11/12), so offline numbers are strong-but-not-exact.
3. **Get a second adversary.** Every per-gun verdict rests on one wave surfer;
the KEEP/BELOW boundaries are matchup-specific and per-gun N is small
(47–898 shots).
4. **Decide Tsetlin and Displace.** Both are below overall on small N; Tsetlin
now learns (clauses 714→13.8 literals) but is not competitive — tune the
regression head or drop.
5. **Do not re-enable TMSelect** until the gate-margin/label problem is fixed;
it cost 7.47% → 5.59% despite showing real energy structure in its clauses.
6. **Real-hit-rate-driven selection is not viable yet** — unselected guns get
near-zero shots, so it needs forced exploration + shrinkage + thousands of
shots per gun.