013b9fe01e
Three corrections, all prompted by later measurements: 1. The virtual-vs-real rank correlation is NOT robustly negative. Six independent Spearman measurements now exist (-0.374, +0.335, +0.522, -0.371, -0.073, -0.037) and the sign flips on large samples, so it is near zero on average. The honest headline is that virtual hit rate is a POOR RANKER, not an inverted one. The report said 'not weak - it is inverted' in six places; it now says so in none. The practical conclusion (do not trust it for ranking) is unchanged; the mechanism claimed was wrong. 2. The offline==online acceptance is FIXED, not flaky. Root cause was that the replay spawned gun 13 (TMSelect) while the live rack has it disabled, and the shared VirtualTracker ring is ORDER-SENSITIVE, so gun 13's extra 4 bullets/tick permuted the per-tick resolution order for every other gun and shifted the learning guns' observations. After closing gun 13's ready gate offline the live and offline KNN traces are byte-identical (904/904 lines, empty diff). 5/5 consecutive runs now report 12/12 exact with the death boundary included. Recorded with the lesson: a flaky proof was hiding a real bug. Also records the general A/B confound - disabling a gun removes its 4 spawns/tick from the shared ring, perturbing resolution order for the rest. 3. Pruning was tested and does NOT help, so the verdict for Tsetlin and Displace changes from an implied drop to BELOW OVERALL - KEEP. 15 paired runs: baseline 6.18%, Tsetlin-off 5.76%, Tsetlin+Displace-off 5.46%; paired permutation p=0.57 and p=0.21; distributions completely overlap; a non-surfer control showed no separation. Being below average does not justify removal. Also records the tie-break randomness fix, and quotes run counts with every rate (6.95% over 13 runs vs 6.18% over 15 runs, same binary) rather than presenting a single figure as definitive.
81 lines
4.2 KiB
Markdown
81 lines
4.2 KiB
Markdown
# Gun Rack Summary — TL;DR
|
||
|
||
**Date:** 2026-09-21 · **Bot:** ModularBot (13 active guns + 1 disabled TM gate)
|
||
**Shipped config:** metric `path`, selector `relative`.
|
||
**Ground truth:** server-side real hit rate (events sidecar), per-run, with an
|
||
explicit overlap test. Scores are not used — they swing by a couple of hundred
|
||
points per run.
|
||
|
||
**Boss:** the unmodified DrussGT jar through `tools/robocode_shim/`.
|
||
**Headline (13 runs vs DrussGT, 3,612 server-side shots): 6.95% real hit rate,
|
||
251 dmg/run.** The **same binary** on a 15-run set measures **6.18%** (events
|
||
6.16%), so always read a rate with its run count.
|
||
|
||
**The one thing to know:** *virtual hit rate is a poor ranker, not a reliable
|
||
proxy for real hit rate.* Its Spearman correlation with real hit rate is weak
|
||
and **sign-unstable** across run sets — **−0.374** over 13 runs, but **+0.335**
|
||
over a 15-run paired baseline (same 13 guns, different but equally defensible
|
||
aggregation) and +0.522 on the 5-run `relative+path` set — i.e. near zero on
|
||
average, **not** reliably anti-correlated. An earlier version of this report
|
||
overstated it as an inversion; the later 15-run measurement refuted that. 16
|
||
candidate ranking rules all failed to beat the shipped config; what works is the
|
||
selector's floor/tie hedging (removing the floor: 5.08% / 175 dmg vs 6.95% /
|
||
251 dmg). Detail: [`gun_rack_analysis.md`](gun_rack_analysis.md).
|
||
|
||
## Verdict table
|
||
|
||
| Verdict | Gun | Real % | Virtual % |
|
||
|---|---|---:|---:|
|
||
| **KEEP** | Linear | 10.7 | 10.2 |
|
||
| **KEEP** | Circular | 9.9 | 11.9 |
|
||
| **KEEP** | KNN | 9.0 | 7.5 |
|
||
| **KEEP** | Pattern | 8.6 | 12.0 |
|
||
| **KEEP** | Accel | 7.3 | 12.1 |
|
||
| **KEEP** | AvgLead | 7.0 | 12.3 |
|
||
| MARGINAL | GuessFactor | 6.9 | 10.3 |
|
||
| MARGINAL | DecayGF | 6.4 | 9.2 |
|
||
| MARGINAL | WallBounce | 6.2 | 12.9 |
|
||
| MARGINAL | StopShot | 6.1 | 12.6 |
|
||
| **KEEP** (below overall; pruning no help) | Tsetlin | 5.8 | 12.9 |
|
||
| **KEEP** (below overall; pruning no help) | Displace | 5.3 | 12.3 |
|
||
| **FLOOR — KEEP** | HeadOn | 5.2 | 8.6 |
|
||
| DISABLED | TMSelect | — | — |
|
||
|
||
HeadOn is lowest but **must stay**: it is the floor fallback, and disabling the
|
||
floor measurably hurt (5.08% / 175 dmg).
|
||
|
||
**Keep the full rack.** Tsetlin and Displace are below overall, but removing
|
||
them was **tested**: 15 paired runs per variant gave 6.18% → 5.76% (Tsetlin
|
||
disabled) and 5.46% (Tsetlin+Displace disabled), with fully overlapping per-run
|
||
distributions and paired permutation p = 0.57 / 0.21. Being below overall does
|
||
**not** justify removal.
|
||
|
||
## Top actions
|
||
|
||
1. **Stop trusting the virtual metric as a ranker.** It is a poor ranker —
|
||
weak, sign-unstable across run sets (−0.374 over 13 runs vs +0.335 over 15),
|
||
and near zero on average, **not** reliably anti-correlated. The selector
|
||
survives on its floor/tie hedge, not on its ordering.
|
||
2. **Acceptance test is fixed, not flaky.** The old 11/12 was a real replay
|
||
bug (the replay spawned disabled gun 13, and the shared ring is
|
||
order-sensitive, so it permuted every other gun's resolution order). It now
|
||
mirrors the live rack: 5/5 runs byte-identical 12/12 with the death boundary
|
||
included.
|
||
3. **Get a second adversary.** Every per-gun verdict rests on one wave surfer;
|
||
the KEEP/BELOW boundaries are matchup-specific and per-gun N is small
|
||
(47–898 shots).
|
||
4. **Keep Tsetlin and Displace (pruning tested).** Both are below overall on
|
||
small N, but disabling Tsetlin (15 paired runs: 6.18% → 5.76%) and
|
||
Tsetlin+Displace (→ 5.46%) was neutral-to-slightly-negative; keep the full
|
||
rack. Tsetlin now learns (clauses 714→13.8 literals) but is not competitive —
|
||
the regression head is a follow-up, not grounds for removal.
|
||
5. **Do not re-enable TMSelect** until the gate-margin/label problem is fixed;
|
||
it cost 7.47% → 5.59% despite showing real energy structure in its clauses.
|
||
6. **Real-hit-rate-driven selection is not viable yet** — unselected guns get
|
||
near-zero shots, so it needs forced exploration + shrinkage + thousands of
|
||
shots per gun.
|
||
7. **Selector tie-break is now seeded.** It had been effectively deterministic
|
||
(`randomize()` reached only incidentally via the Tsetlin constructor); now an
|
||
explicit startup seed plus a `GUN_SELECTOR_SEED` override makes seeded runs
|
||
reproducible and unseeded runs vary.
|