Hypothesis under test (from the gun audit, which named the tie-band as "the lever that matters most"): `bmPath` is deliberately generous (2.3-3.6x `bmPoint`), so a gun can sit in the tied band on a ray that sweeps the target's path while its bullets ARRIVE badly. So: keep the `path`-ranked band (path beat point on real hit rate 7.43% vs 4.70%, z=5.56), but narrow the random draw inside it using a parallel `point` (arrival-accuracy) window. RESULT: NO EFFECT. Real DrussGT, ONE frozen binary (/tmp/ModularBot_tieband, md5 2c0c56e6...), env knobs only, 7 runs x 7 rounds per arm, server-side events sidecar, exact two-sided permutation test on per-run rates. arm runs shots real % dmg/run d p tbbase (shipped) 7 4128 7.17 175 -- -- tbpt path-rank + point-narrow 7 3938 7.08 165 +0.14 0.88 tbpc =commit control 7 3759 4.44 98 +2.74 0.0012 tbpt25 point margin 0.25 7 3683 5.59 119 +1.65 0.20 tbtie05 / tbtie40 (band width) 7 3937/3917 5.84/6.28 133/144 1.49/1.00 0.11/0.25 tbwin50 (SelectorWindow=50) 7 3983 6.05 139 +1.20 0.11 tbfloor10 (FloorPeakFrac=0.10) 7 3829 5.33 118 +2.12 0.11 tbpt vs base: fully overlapping ranges, p=0.88. This is a REAL null, not a dead arm - the mechanism was live, and it visibly changed the selected-gun mix (Pattern 24%->16%, Accel 6%->16%, Tsetlin ~0%->13%). CONTROL VALIDATED, AND THIS IS THE THIRD TIME: removing the random draw inside the band is SIGNIFICANTLY WORSE (4.44%, p=0.0012). Combined with the earlier hysteresis A/B (7.02% -> 5.10% for commitment) and the light-hysteresis result, the selector's per-tick randomness is now load-bearing on three independent measurements. Narrowing the band on ANY second virtual statistic has not helped. Every knob swept (band width, floor, window) is nominally worse than shipped at n=7; that is "no credible win" rather than "proven harm" (sd ~1.8pp, ~1pp resolution, underpowered). Shipped default stays `GUN_SELECTOR_TIEBREAK=off`; the feature is opt-in, fully guarded, and costs zero extra work on the default path (point windows are scored only when the mode is on). Guards: test_selector_tiebreak 19 (new, pure), test_gun_harness 39, test_vbullet_metric 11, test_adaptive_radar 41, test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28, test_rack_membership 38, acceptance_offline_vs_online 12/12 PASS (offline path calls neither chooseFromFit nor the tie-break). STRATEGIC CONCLUSION: three selection-side attempts have now failed (hysteresis, commitment, point tie-break). The selector is at a local optimum and the remaining lever is the QUALITY OF THE GUNS, not the selection among them.
4.7 KiB
Gun Rack Summary — TL;DR
Date: 2026-09-21 · Bot: ModularBot (13 active guns + 1 disabled TM gate)
Shipped config: metric path, selector relative.
Ground truth: server-side real hit rate (events sidecar), per-run, with an
explicit overlap test. Scores are not used — they swing by a couple of hundred
points per run.
Boss: the unmodified DrussGT jar through tools/robocode_shim/.
Headline (13 runs vs DrussGT, 3,612 server-side shots): 6.95% real hit rate,
251 dmg/run. The same binary on a 15-run set measures 6.18% (events
6.16%), so always read a rate with its run count.
The one thing to know: virtual hit rate is a poor ranker, not a reliable
proxy for real hit rate. Its Spearman correlation with real hit rate is weak
and sign-unstable across run sets — −0.374 over 13 runs, but +0.335
over a 15-run paired baseline (same 13 guns, different but equally defensible
aggregation) and +0.522 on the 5-run relative+path set — i.e. near zero on
average, not reliably anti-correlated. An earlier version of this report
overstated it as an inversion; the later 15-run measurement refuted that. 16
candidate ranking rules all failed to beat the shipped config; what works is the
selector's floor/tie hedging (removing the floor: 5.08% / 175 dmg vs 6.95% /
251 dmg). Detail: gun_rack_analysis.md.
Verdict table
| Verdict | Gun | Real % | Virtual % |
|---|---|---|---|
| KEEP | Linear | 10.7 | 10.2 |
| KEEP | Circular | 9.9 | 11.9 |
| KEEP | KNN | 9.0 | 7.5 |
| KEEP | Pattern | 8.6 | 12.0 |
| KEEP | Accel | 7.3 | 12.1 |
| KEEP | AvgLead | 7.0 | 12.3 |
| MARGINAL | GuessFactor | 6.9 | 10.3 |
| MARGINAL | DecayGF | 6.4 | 9.2 |
| MARGINAL | WallBounce | 6.2 | 12.9 |
| MARGINAL | StopShot | 6.1 | 12.6 |
| KEEP (below overall; pruning no help) | Tsetlin | 5.8 | 12.9 |
| KEEP (below overall; pruning no help) | Displace | 5.3 | 12.3 |
| FLOOR — KEEP | HeadOn | 5.2 | 8.6 |
| DISABLED | TMSelect | — | — |
HeadOn is lowest but must stay: it is the floor fallback, and disabling the floor measurably hurt (5.08% / 175 dmg).
Keep the full rack. Tsetlin and Displace are below overall, but removing them was tested: 15 paired runs per variant gave 6.18% → 5.76% (Tsetlin disabled) and 5.46% (Tsetlin+Displace disabled), with fully overlapping per-run distributions and paired permutation p = 0.57 / 0.21. Being below overall does not justify removal.
Top actions
- Stop trusting the virtual metric as a ranker. It is a poor ranker — weak, sign-unstable across run sets (−0.374 over 13 runs vs +0.335 over 15), and near zero on average, not reliably anti-correlated. The selector survives on its floor/tie hedge, not on its ordering.
- Acceptance test is fixed, not flaky. The old 11/12 was a real replay bug (the replay spawned disabled gun 13, and the shared ring is order-sensitive, so it permuted every other gun's resolution order). It now mirrors the live rack: 5/5 runs byte-identical 12/12 with the death boundary included.
- Get a second adversary. Every per-gun verdict rests on one wave surfer; the KEEP/BELOW boundaries are matchup-specific and per-gun N is small (47–898 shots).
- Keep Tsetlin and Displace (pruning tested). Both are below overall on small N, but disabling Tsetlin (15 paired runs: 6.18% → 5.76%) and Tsetlin+Displace (→ 5.46%) was neutral-to-slightly-negative; keep the full rack. Tsetlin now learns (clauses 714→13.8 literals) but is not competitive — the regression head is a follow-up, not grounds for removal.
- Do not re-enable TMSelect until the gate-margin/label problem is fixed; it cost 7.47% → 5.59% despite showing real energy structure in its clauses.
- Real-hit-rate-driven selection is not viable yet — unselected guns get near-zero shots, so it needs forced exploration + shrinkage + thousands of shots per gun.
- Selector tie-break is now seeded. It had been effectively deterministic
(
randomize()reached only incidentally via the Tsetlin constructor); now an explicit startup seed plus aGUN_SELECTOR_SEEDoverride makes seeded runs reproducible and unseeded runs vary. - The arrival-accuracy tie-band does not beat the shipped band (7 runs x 7
rounds/arm vs DrussGT, one frozen binary): path-ranking + point-narrowing
7.08% vs shipped 7.17% (overlapping, p = 0.88), while the no-randomness
control (
GUN_SELECTOR_TIEBREAK=commit) is significantly worse at 4.44% (p = 0.0012). TheTIE/FLOOR/WINDOWsweep is also nominally worse at every setting. Shipped selector unchanged;GUN_SELECTOR_TIEBREAKdefaults tooff. Detail:gun_rack_analysis.md§6.8.