0ede6d12ec
Hypothesis under test (from the gun audit, which named the tie-band as "the lever that matters most"): `bmPath` is deliberately generous (2.3-3.6x `bmPoint`), so a gun can sit in the tied band on a ray that sweeps the target's path while its bullets ARRIVE badly. So: keep the `path`-ranked band (path beat point on real hit rate 7.43% vs 4.70%, z=5.56), but narrow the random draw inside it using a parallel `point` (arrival-accuracy) window. RESULT: NO EFFECT. Real DrussGT, ONE frozen binary (/tmp/ModularBot_tieband, md5 2c0c56e6...), env knobs only, 7 runs x 7 rounds per arm, server-side events sidecar, exact two-sided permutation test on per-run rates. arm runs shots real % dmg/run d p tbbase (shipped) 7 4128 7.17 175 -- -- tbpt path-rank + point-narrow 7 3938 7.08 165 +0.14 0.88 tbpc =commit control 7 3759 4.44 98 +2.74 0.0012 tbpt25 point margin 0.25 7 3683 5.59 119 +1.65 0.20 tbtie05 / tbtie40 (band width) 7 3937/3917 5.84/6.28 133/144 1.49/1.00 0.11/0.25 tbwin50 (SelectorWindow=50) 7 3983 6.05 139 +1.20 0.11 tbfloor10 (FloorPeakFrac=0.10) 7 3829 5.33 118 +2.12 0.11 tbpt vs base: fully overlapping ranges, p=0.88. This is a REAL null, not a dead arm - the mechanism was live, and it visibly changed the selected-gun mix (Pattern 24%->16%, Accel 6%->16%, Tsetlin ~0%->13%). CONTROL VALIDATED, AND THIS IS THE THIRD TIME: removing the random draw inside the band is SIGNIFICANTLY WORSE (4.44%, p=0.0012). Combined with the earlier hysteresis A/B (7.02% -> 5.10% for commitment) and the light-hysteresis result, the selector's per-tick randomness is now load-bearing on three independent measurements. Narrowing the band on ANY second virtual statistic has not helped. Every knob swept (band width, floor, window) is nominally worse than shipped at n=7; that is "no credible win" rather than "proven harm" (sd ~1.8pp, ~1pp resolution, underpowered). Shipped default stays `GUN_SELECTOR_TIEBREAK=off`; the feature is opt-in, fully guarded, and costs zero extra work on the default path (point windows are scored only when the mode is on). Guards: test_selector_tiebreak 19 (new, pure), test_gun_harness 39, test_vbullet_metric 11, test_adaptive_radar 41, test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28, test_rack_membership 38, acceptance_offline_vs_online 12/12 PASS (offline path calls neither chooseFromFit nor the tie-break). STRATEGIC CONCLUSION: three selection-side attempts have now failed (hysteresis, commitment, point tie-break). The selector is at a local optimum and the remaining lever is the QUALITY OF THE GUNS, not the selection among them.
88 lines
4.7 KiB
Markdown
88 lines
4.7 KiB
Markdown
# Gun Rack Summary — TL;DR
|
||
|
||
**Date:** 2026-09-21 · **Bot:** ModularBot (13 active guns + 1 disabled TM gate)
|
||
**Shipped config:** metric `path`, selector `relative`.
|
||
**Ground truth:** server-side real hit rate (events sidecar), per-run, with an
|
||
explicit overlap test. Scores are not used — they swing by a couple of hundred
|
||
points per run.
|
||
|
||
**Boss:** the unmodified DrussGT jar through `tools/robocode_shim/`.
|
||
**Headline (13 runs vs DrussGT, 3,612 server-side shots): 6.95% real hit rate,
|
||
251 dmg/run.** The **same binary** on a 15-run set measures **6.18%** (events
|
||
6.16%), so always read a rate with its run count.
|
||
|
||
**The one thing to know:** *virtual hit rate is a poor ranker, not a reliable
|
||
proxy for real hit rate.* Its Spearman correlation with real hit rate is weak
|
||
and **sign-unstable** across run sets — **−0.374** over 13 runs, but **+0.335**
|
||
over a 15-run paired baseline (same 13 guns, different but equally defensible
|
||
aggregation) and +0.522 on the 5-run `relative+path` set — i.e. near zero on
|
||
average, **not** reliably anti-correlated. An earlier version of this report
|
||
overstated it as an inversion; the later 15-run measurement refuted that. 16
|
||
candidate ranking rules all failed to beat the shipped config; what works is the
|
||
selector's floor/tie hedging (removing the floor: 5.08% / 175 dmg vs 6.95% /
|
||
251 dmg). Detail: [`gun_rack_analysis.md`](gun_rack_analysis.md).
|
||
|
||
## Verdict table
|
||
|
||
| Verdict | Gun | Real % | Virtual % |
|
||
|---|---|---:|---:|
|
||
| **KEEP** | Linear | 10.7 | 10.2 |
|
||
| **KEEP** | Circular | 9.9 | 11.9 |
|
||
| **KEEP** | KNN | 9.0 | 7.5 |
|
||
| **KEEP** | Pattern | 8.6 | 12.0 |
|
||
| **KEEP** | Accel | 7.3 | 12.1 |
|
||
| **KEEP** | AvgLead | 7.0 | 12.3 |
|
||
| MARGINAL | GuessFactor | 6.9 | 10.3 |
|
||
| MARGINAL | DecayGF | 6.4 | 9.2 |
|
||
| MARGINAL | WallBounce | 6.2 | 12.9 |
|
||
| MARGINAL | StopShot | 6.1 | 12.6 |
|
||
| **KEEP** (below overall; pruning no help) | Tsetlin | 5.8 | 12.9 |
|
||
| **KEEP** (below overall; pruning no help) | Displace | 5.3 | 12.3 |
|
||
| **FLOOR — KEEP** | HeadOn | 5.2 | 8.6 |
|
||
| DISABLED | TMSelect | — | — |
|
||
|
||
HeadOn is lowest but **must stay**: it is the floor fallback, and disabling the
|
||
floor measurably hurt (5.08% / 175 dmg).
|
||
|
||
**Keep the full rack.** Tsetlin and Displace are below overall, but removing
|
||
them was **tested**: 15 paired runs per variant gave 6.18% → 5.76% (Tsetlin
|
||
disabled) and 5.46% (Tsetlin+Displace disabled), with fully overlapping per-run
|
||
distributions and paired permutation p = 0.57 / 0.21. Being below overall does
|
||
**not** justify removal.
|
||
|
||
## Top actions
|
||
|
||
1. **Stop trusting the virtual metric as a ranker.** It is a poor ranker —
|
||
weak, sign-unstable across run sets (−0.374 over 13 runs vs +0.335 over 15),
|
||
and near zero on average, **not** reliably anti-correlated. The selector
|
||
survives on its floor/tie hedge, not on its ordering.
|
||
2. **Acceptance test is fixed, not flaky.** The old 11/12 was a real replay
|
||
bug (the replay spawned disabled gun 13, and the shared ring is
|
||
order-sensitive, so it permuted every other gun's resolution order). It now
|
||
mirrors the live rack: 5/5 runs byte-identical 12/12 with the death boundary
|
||
included.
|
||
3. **Get a second adversary.** Every per-gun verdict rests on one wave surfer;
|
||
the KEEP/BELOW boundaries are matchup-specific and per-gun N is small
|
||
(47–898 shots).
|
||
4. **Keep Tsetlin and Displace (pruning tested).** Both are below overall on
|
||
small N, but disabling Tsetlin (15 paired runs: 6.18% → 5.76%) and
|
||
Tsetlin+Displace (→ 5.46%) was neutral-to-slightly-negative; keep the full
|
||
rack. Tsetlin now learns (clauses 714→13.8 literals) but is not competitive —
|
||
the regression head is a follow-up, not grounds for removal.
|
||
5. **Do not re-enable TMSelect** until the gate-margin/label problem is fixed;
|
||
it cost 7.47% → 5.59% despite showing real energy structure in its clauses.
|
||
6. **Real-hit-rate-driven selection is not viable yet** — unselected guns get
|
||
near-zero shots, so it needs forced exploration + shrinkage + thousands of
|
||
shots per gun.
|
||
7. **Selector tie-break is now seeded.** It had been effectively deterministic
|
||
(`randomize()` reached only incidentally via the Tsetlin constructor); now an
|
||
explicit startup seed plus a `GUN_SELECTOR_SEED` override makes seeded runs
|
||
reproducible and unseeded runs vary.
|
||
8. **The arrival-accuracy tie-band does not beat the shipped band** (7 runs x 7
|
||
rounds/arm vs DrussGT, one frozen binary): path-ranking + point-narrowing
|
||
7.08% vs shipped 7.17% (overlapping, p = 0.88), while the no-randomness
|
||
control (`GUN_SELECTOR_TIEBREAK=commit`) is significantly worse at 4.44%
|
||
(p = 0.0012). The `TIE`/`FLOOR`/`WINDOW` sweep is also nominally worse at
|
||
every setting. **Shipped selector unchanged**; `GUN_SELECTOR_TIEBREAK`
|
||
defaults to `off`. Detail: [`gun_rack_analysis.md`](gun_rack_analysis.md) §6.8.
|