docs: definitive gun-rack report on real measured numbers
Replaces the stale 2026-09-20 docs, which predated per-gun real attribution, the offline gun range and the DrussGT boss, and whose verdicts were built on virtual hit rates that turned out to be ANTI-correlated with reality. docs/gun_rack_analysis.md (751 lines) covers: the test infrastructure described honestly (offline range with its flaky-acceptance caveat, the 20 fixtures and what each set is good for, the live boss, and the A/B methodology of per-run server-side real hit rate with an explicit overlap test); the virtual-vs-real metric lesson with Spearman -0.374 and the point-vs-path A/B; per-gun real performance and the 16-rule ranking A/B; the offline per-fixture gun matrix; KEEP/MARGINAL/BELOW verdicts; and the five root-cause bugs with before/after numbers. docs/gun_rack_summary.md (58 lines) is the verdict table plus top actions. The '~230 point' score-noise band that has been steering methodology all night was re-derived from the artifacts rather than asserted: the 13 shipped-config run scores span 175-526, s.d. ~105, i.e. a ~210-point 2-s.d. band. Caveats recorded verbatim rather than softened: the offline==online acceptance is flaky (typically 11/12 on unmodified HEAD), fixtures are perfect-information and therefore optimistic vs live play, per-gun real N is small so single-gun ordering is indicative, the headline numbers come from ONE wave-surfer adversary, and HeadOn must stay despite being lowest because it is the floor fallback (disabling it: 5.08% / 175 dmg vs 6.95% / 251 dmg).
This commit is contained in:
+50
-32
@@ -1,40 +1,58 @@
|
||||
# Gun Rack Summary — TL;DR
|
||||
|
||||
**Gauntlet:** 5 adversaries × 10 rounds, 14 guns. Data: `/tmp/gun_stats.jsonl`
|
||||
**Date:** 2026-09-21 · **Bot:** ModularBot (13 active guns + 1 disabled TM gate)
|
||||
**Shipped config:** metric `path`, selector `relative`.
|
||||
**Ground truth:** server-side real hit rate (events sidecar), per-run, with an
|
||||
explicit overlap test. Scores are not used — they swing by a couple of hundred
|
||||
points per run.
|
||||
|
||||
## Scoreboard
|
||||
**Boss:** the unmodified DrussGT jar through `tools/robocode_shim/`.
|
||||
**Headline (13 runs vs DrussGT, 3,612 server-side shots): 6.95% real hit rate,
|
||||
251 dmg/run.**
|
||||
|
||||
| Adversary | Score | Real Hit% | Best Gun |
|
||||
|-----------|-------|-----------|----------|
|
||||
| SittingDuck | 1944 | 85% | AvgLead / WallBounce (99% vhit) |
|
||||
| OscillatorBot | 1645 | 59% | AvgLead (65% vhit) |
|
||||
| RandomMover | 1916 | 61% | Linear / WallBounce / AvgLead (82%) |
|
||||
| PatternMover | 1914 | 73% | AvgLead / Displace / Accel (96%) |
|
||||
| WaveSurfer | 1873 | 72% | AvgLead / StopShot / Accel (86%) |
|
||||
**The one thing to know:** *virtual hit rate is not a proxy for real hit rate.*
|
||||
For the shipped config Spearman(virtual rank, real rank) = **−0.374** —
|
||||
anti-correlated. 16 candidate ranking rules all failed to beat the shipped
|
||||
config; what works is the selector's floor/tie hedging (removing the floor:
|
||||
5.08% / 175 dmg vs 6.95% / 251 dmg). Detail: [`gun_rack_analysis.md`](gun_rack_analysis.md).
|
||||
|
||||
## Verdict Table
|
||||
## Verdict table
|
||||
|
||||
| Gun | Best Use | Verdict |
|
||||
|-----|----------|---------|
|
||||
| **HeadOn** | Stationary / pattern bots | KEEP — reliable floor, over-selected |
|
||||
| **Linear** | Random movers | KEEP — under-selected despite 82% |
|
||||
| **Tsetlin** | Unknown | TUNE — 0 selection ticks ever; investigate |
|
||||
| **Circular** | Orbit-heavy bots | KEEP (marginal) |
|
||||
| **GuessFactor** | General | KEEP — underperforms vs expectation; tune bins |
|
||||
| **Pattern** | Pattern movers | TUNE — catastrophic vs OscillatorBot (9%, 1081 ticks); raise MinObs gate |
|
||||
| **AntiSurf** | — | DROP — 0% vhit every adversary including stationary; broken |
|
||||
| **WallBounce** | All types | KEEP — most consistent overall |
|
||||
| **Accel** | Pattern / wave bots | KEEP — avoid vs random (19% cliff) |
|
||||
| **StopShot** | All types | KEEP — top rates, chronically under-selected |
|
||||
| **Displace** | Pattern / wave bots | KEEP |
|
||||
| **AvgLead** | All types | KEEP — best all-rounder in the rack |
|
||||
| **DecayGF** | Wave / pattern bots | KEEP |
|
||||
| **KNN** | Pattern / wave bots | KEEP — improves with history |
|
||||
| Verdict | Gun | Real % | Virtual % |
|
||||
|---|---|---:|---:|
|
||||
| **KEEP** | Linear | 10.7 | 10.2 |
|
||||
| **KEEP** | Circular | 9.9 | 11.9 |
|
||||
| **KEEP** | KNN | 9.0 | 7.5 |
|
||||
| **KEEP** | Pattern | 8.6 | 12.0 |
|
||||
| **KEEP** | Accel | 7.3 | 12.1 |
|
||||
| **KEEP** | AvgLead | 7.0 | 12.3 |
|
||||
| MARGINAL | GuessFactor | 6.9 | 10.3 |
|
||||
| MARGINAL | DecayGF | 6.4 | 9.2 |
|
||||
| MARGINAL | WallBounce | 6.2 | 12.9 |
|
||||
| MARGINAL | StopShot | 6.1 | 12.6 |
|
||||
| BELOW | Tsetlin | 5.8 | 12.9 |
|
||||
| BELOW | Displace | 5.3 | 12.3 |
|
||||
| **FLOOR — KEEP** | HeadOn | 5.2 | 8.6 |
|
||||
| DISABLED | TMSelect | — | — |
|
||||
|
||||
## Top Actions
|
||||
HeadOn is lowest but **must stay**: it is the floor fallback, and disabling the
|
||||
floor measurably hurt (5.08% / 175 dmg).
|
||||
|
||||
1. **Fix AntiSurf** — 0% vhit against SittingDuck means wrong angle computation, not just weak targeting.
|
||||
2. **Raise Pattern's MinObsBeforeCompete** — 15 is too low; 1081 wasted ticks at 9% hit rate vs OscillatorBot.
|
||||
3. **Re-run gauntlet** with `onBulletHit` fix to get clean per-gun real hit attribution.
|
||||
4. **AvgLead** is the star gun — confirm it stays in the top selection tier.
|
||||
5. **HeadOn selection dominance** is a selector bias problem, not a gun problem — add ±2% tiebreak randomisation.
|
||||
## Top actions
|
||||
|
||||
1. **Stop trusting the virtual metric as a ranker.** It is anti-correlated with
|
||||
real hit rate and unstable across run sets. The selector survives on its
|
||||
floor/tie hedge, not on its ordering.
|
||||
2. **Fix the offline==online acceptance race.** It is currently flaky
|
||||
(typically 11/12), so offline numbers are strong-but-not-exact.
|
||||
3. **Get a second adversary.** Every per-gun verdict rests on one wave surfer;
|
||||
the KEEP/BELOW boundaries are matchup-specific and per-gun N is small
|
||||
(47–898 shots).
|
||||
4. **Decide Tsetlin and Displace.** Both are below overall on small N; Tsetlin
|
||||
now learns (clauses 714→13.8 literals) but is not competitive — tune the
|
||||
regression head or drop.
|
||||
5. **Do not re-enable TMSelect** until the gate-margin/label problem is fixed;
|
||||
it cost 7.47% → 5.59% despite showing real energy structure in its clauses.
|
||||
6. **Real-hit-rate-driven selection is not viable yet** — unselected guns get
|
||||
near-zero shots, so it needs forced exploration + shrinkage + thousands of
|
||||
shots per gun.
|
||||
|
||||
Reference in New Issue
Block a user