Files
SirRoboGarage/docs/gun_rack_summary.md
SirStone 0ede6d12ec selector: arrival-accuracy tie-break measured NEGATIVE; randomness is load-bearing
Hypothesis under test (from the gun audit, which named the tie-band as "the
lever that matters most"): `bmPath` is deliberately generous (2.3-3.6x
`bmPoint`), so a gun can sit in the tied band on a ray that sweeps the target's
path while its bullets ARRIVE badly. So: keep the `path`-ranked band (path beat
point on real hit rate 7.43% vs 4.70%, z=5.56), but narrow the random draw
inside it using a parallel `point` (arrival-accuracy) window.

RESULT: NO EFFECT. Real DrussGT, ONE frozen binary (/tmp/ModularBot_tieband,
md5 2c0c56e6...), env knobs only, 7 runs x 7 rounds per arm, server-side events
sidecar, exact two-sided permutation test on per-run rates.

  arm                              runs  shots  real %  dmg/run   d      p
  tbbase (shipped)                    7   4128   7.17     175      --     --
  tbpt  path-rank + point-narrow      7   3938   7.08     165    +0.14  0.88
  tbpc  =commit control               7   3759   4.44      98    +2.74  0.0012
  tbpt25 point margin 0.25            7   3683   5.59     119    +1.65  0.20
  tbtie05 / tbtie40 (band width)      7   3937/3917  5.84/6.28  133/144  1.49/1.00  0.11/0.25
  tbwin50 (SelectorWindow=50)         7   3983   6.05     139    +1.20  0.11
  tbfloor10 (FloorPeakFrac=0.10)      7   3829   5.33     118    +2.12  0.11

tbpt vs base: fully overlapping ranges, p=0.88. This is a REAL null, not a dead
arm - the mechanism was live, and it visibly changed the selected-gun mix
(Pattern 24%->16%, Accel 6%->16%, Tsetlin ~0%->13%).

CONTROL VALIDATED, AND THIS IS THE THIRD TIME: removing the random draw inside
the band is SIGNIFICANTLY WORSE (4.44%, p=0.0012). Combined with the earlier
hysteresis A/B (7.02% -> 5.10% for commitment) and the light-hysteresis result,
the selector's per-tick randomness is now load-bearing on three independent
measurements. Narrowing the band on ANY second virtual statistic has not helped.

Every knob swept (band width, floor, window) is nominally worse than shipped at
n=7; that is "no credible win" rather than "proven harm" (sd ~1.8pp, ~1pp
resolution, underpowered).

Shipped default stays `GUN_SELECTOR_TIEBREAK=off`; the feature is opt-in, fully
guarded, and costs zero extra work on the default path (point windows are scored
only when the mode is on).

Guards: test_selector_tiebreak 19 (new, pure), test_gun_harness 39,
test_vbullet_metric 11, test_adaptive_radar 41, test_tfil_ring_weights 24,
test_power_policy 26, test_ram_decision 28, test_rack_membership 38,
acceptance_offline_vs_online 12/12 PASS (offline path calls neither
chooseFromFit nor the tie-break).

STRATEGIC CONCLUSION: three selection-side attempts have now failed (hysteresis,
commitment, point tie-break). The selector is at a local optimum and the
remaining lever is the QUALITY OF THE GUNS, not the selection among them.
2026-09-22 01:07:12 +02:00

88 lines
4.7 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Gun Rack Summary — TL;DR
**Date:** 2026-09-21 · **Bot:** ModularBot (13 active guns + 1 disabled TM gate)
**Shipped config:** metric `path`, selector `relative`.
**Ground truth:** server-side real hit rate (events sidecar), per-run, with an
explicit overlap test. Scores are not used — they swing by a couple of hundred
points per run.
**Boss:** the unmodified DrussGT jar through `tools/robocode_shim/`.
**Headline (13 runs vs DrussGT, 3,612 server-side shots): 6.95% real hit rate,
251 dmg/run.** The **same binary** on a 15-run set measures **6.18%** (events
6.16%), so always read a rate with its run count.
**The one thing to know:** *virtual hit rate is a poor ranker, not a reliable
proxy for real hit rate.* Its Spearman correlation with real hit rate is weak
and **sign-unstable** across run sets — **−0.374** over 13 runs, but **+0.335**
over a 15-run paired baseline (same 13 guns, different but equally defensible
aggregation) and +0.522 on the 5-run `relative+path` set — i.e. near zero on
average, **not** reliably anti-correlated. An earlier version of this report
overstated it as an inversion; the later 15-run measurement refuted that. 16
candidate ranking rules all failed to beat the shipped config; what works is the
selector's floor/tie hedging (removing the floor: 5.08% / 175 dmg vs 6.95% /
251 dmg). Detail: [`gun_rack_analysis.md`](gun_rack_analysis.md).
## Verdict table
| Verdict | Gun | Real % | Virtual % |
|---|---|---:|---:|
| **KEEP** | Linear | 10.7 | 10.2 |
| **KEEP** | Circular | 9.9 | 11.9 |
| **KEEP** | KNN | 9.0 | 7.5 |
| **KEEP** | Pattern | 8.6 | 12.0 |
| **KEEP** | Accel | 7.3 | 12.1 |
| **KEEP** | AvgLead | 7.0 | 12.3 |
| MARGINAL | GuessFactor | 6.9 | 10.3 |
| MARGINAL | DecayGF | 6.4 | 9.2 |
| MARGINAL | WallBounce | 6.2 | 12.9 |
| MARGINAL | StopShot | 6.1 | 12.6 |
| **KEEP** (below overall; pruning no help) | Tsetlin | 5.8 | 12.9 |
| **KEEP** (below overall; pruning no help) | Displace | 5.3 | 12.3 |
| **FLOOR — KEEP** | HeadOn | 5.2 | 8.6 |
| DISABLED | TMSelect | — | — |
HeadOn is lowest but **must stay**: it is the floor fallback, and disabling the
floor measurably hurt (5.08% / 175 dmg).
**Keep the full rack.** Tsetlin and Displace are below overall, but removing
them was **tested**: 15 paired runs per variant gave 6.18% → 5.76% (Tsetlin
disabled) and 5.46% (Tsetlin+Displace disabled), with fully overlapping per-run
distributions and paired permutation p = 0.57 / 0.21. Being below overall does
**not** justify removal.
## Top actions
1. **Stop trusting the virtual metric as a ranker.** It is a poor ranker —
weak, sign-unstable across run sets (−0.374 over 13 runs vs +0.335 over 15),
and near zero on average, **not** reliably anti-correlated. The selector
survives on its floor/tie hedge, not on its ordering.
2. **Acceptance test is fixed, not flaky.** The old 11/12 was a real replay
bug (the replay spawned disabled gun 13, and the shared ring is
order-sensitive, so it permuted every other gun's resolution order). It now
mirrors the live rack: 5/5 runs byte-identical 12/12 with the death boundary
included.
3. **Get a second adversary.** Every per-gun verdict rests on one wave surfer;
the KEEP/BELOW boundaries are matchup-specific and per-gun N is small
(47–898 shots).
4. **Keep Tsetlin and Displace (pruning tested).** Both are below overall on
small N, but disabling Tsetlin (15 paired runs: 6.18% → 5.76%) and
Tsetlin+Displace (→ 5.46%) was neutral-to-slightly-negative; keep the full
rack. Tsetlin now learns (clauses 714→13.8 literals) but is not competitive —
the regression head is a follow-up, not grounds for removal.
5. **Do not re-enable TMSelect** until the gate-margin/label problem is fixed;
it cost 7.47% → 5.59% despite showing real energy structure in its clauses.
6. **Real-hit-rate-driven selection is not viable yet** — unselected guns get
near-zero shots, so it needs forced exploration + shrinkage + thousands of
shots per gun.
7. **Selector tie-break is now seeded.** It had been effectively deterministic
(`randomize()` reached only incidentally via the Tsetlin constructor); now an
explicit startup seed plus a `GUN_SELECTOR_SEED` override makes seeded runs
reproducible and unseeded runs vary.
8. **The arrival-accuracy tie-band does not beat the shipped band** (7 runs x 7
rounds/arm vs DrussGT, one frozen binary): path-ranking + point-narrowing
7.08% vs shipped 7.17% (overlapping, p = 0.88), while the no-randomness
control (`GUN_SELECTOR_TIEBREAK=commit`) is significantly worse at 4.44%
(p = 0.0012). The `TIE`/`FLOOR`/`WINDOW` sweep is also nominally worse at
every setting. **Shipped selector unchanged**; `GUN_SELECTOR_TIEBREAK`
defaults to `off`. Detail: [`gun_rack_analysis.md`](gun_rack_analysis.md) §6.8.