Hypothesis under test (from the gun audit, which named the tie-band as "the lever that matters most"): `bmPath` is deliberately generous (2.3-3.6x `bmPoint`), so a gun can sit in the tied band on a ray that sweeps the target's path while its bullets ARRIVE badly. So: keep the `path`-ranked band (path beat point on real hit rate 7.43% vs 4.70%, z=5.56), but narrow the random draw inside it using a parallel `point` (arrival-accuracy) window. RESULT: NO EFFECT. Real DrussGT, ONE frozen binary (/tmp/ModularBot_tieband, md5 2c0c56e6...), env knobs only, 7 runs x 7 rounds per arm, server-side events sidecar, exact two-sided permutation test on per-run rates. arm runs shots real % dmg/run d p tbbase (shipped) 7 4128 7.17 175 -- -- tbpt path-rank + point-narrow 7 3938 7.08 165 +0.14 0.88 tbpc =commit control 7 3759 4.44 98 +2.74 0.0012 tbpt25 point margin 0.25 7 3683 5.59 119 +1.65 0.20 tbtie05 / tbtie40 (band width) 7 3937/3917 5.84/6.28 133/144 1.49/1.00 0.11/0.25 tbwin50 (SelectorWindow=50) 7 3983 6.05 139 +1.20 0.11 tbfloor10 (FloorPeakFrac=0.10) 7 3829 5.33 118 +2.12 0.11 tbpt vs base: fully overlapping ranges, p=0.88. This is a REAL null, not a dead arm - the mechanism was live, and it visibly changed the selected-gun mix (Pattern 24%->16%, Accel 6%->16%, Tsetlin ~0%->13%). CONTROL VALIDATED, AND THIS IS THE THIRD TIME: removing the random draw inside the band is SIGNIFICANTLY WORSE (4.44%, p=0.0012). Combined with the earlier hysteresis A/B (7.02% -> 5.10% for commitment) and the light-hysteresis result, the selector's per-tick randomness is now load-bearing on three independent measurements. Narrowing the band on ANY second virtual statistic has not helped. Every knob swept (band width, floor, window) is nominally worse than shipped at n=7; that is "no credible win" rather than "proven harm" (sd ~1.8pp, ~1pp resolution, underpowered). Shipped default stays `GUN_SELECTOR_TIEBREAK=off`; the feature is opt-in, fully guarded, and costs zero extra work on the default path (point windows are scored only when the mode is on). Guards: test_selector_tiebreak 19 (new, pure), test_gun_harness 39, test_vbullet_metric 11, test_adaptive_radar 41, test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28, test_rack_membership 38, acceptance_offline_vs_online 12/12 PASS (offline path calls neither chooseFromFit nor the tie-break). STRATEGIC CONCLUSION: three selection-side attempts have now failed (hysteresis, commitment, point tie-break). The selector is at a local optimum and the remaining lever is the QUALITY OF THE GUNS, not the selection among them.
50 KiB
Gun Rack Analysis — ModularBot guns vs the real DrussGT boss
Date: 2026-09-21
Bot: ModularBot (13 active guns + 1 disabled TM classifier gun)
Shipped config: virtual-bullet metric path, selector thresholds relative
(GUN_VBULLET_METRIC default path, GUN_SELECTOR_MODE default relative)
Rewritten: this file supersedes the 2026-09-20 version, whose numbers
predated per-gun real-hit attribution, the offline gun range, and the live
DrussGT boss. No number from the old version is retained unless it was
re-measured below.
Corrected after later runs (same day). Three claims in the first version of this rewrite were revised by follow-up measurements: the virtual-vs-real correlation is a poor, sign-unstable ranker rather than an "inversion" (§2); the offline==online acceptance test is fixed, not flaky (§1.1); and pruning the below-overall guns was tested and does not help (§3). Each correction is restated plainly at the point of the old claim.
Evidence tags used throughout:
- [MEASURED] — I read it from a recorded artifact (commit message,
/tmpoutput, a log, or a source file). The provenance is named every time. - [INFERRED] — reasoning from measured facts; explicitly not measured.
A note on provenance: the measurements below were produced overnight in
commits 343e631 … 2c94dc2 on branch research/lead-targeting. Each commit
message records the numbers, the hypotheses it refuted, and its caveats. Raw
artifacts are in /tmp (/tmp/gun_stats_base_r*.jsonl,
/tmp/events_base_r*.json, /tmp/range_path.txt,
/tmp/ab_logs/FINAL_ANALYSIS.txt, /tmp/ab_logs3/FINAL_AB.txt,
/tmp/ab_logs3/selector_diag.txt). Where a commit message and a raw artifact
disagree slightly, both are stated.
1. The test infrastructure
The numbers are only as good as the rig that produced them, so the rig is described first.
1.1 The offline gun range
common_libs/gun_harness/offline_range.nim replays a recorded
seq[WorldState] through the same VirtualTracker
(common_libs/gun_harness/virtual_bullets.nim) that the live bot drives. The
claim is not "an approximation of the live metric", it is "the same metric":
virtual-bullet fitness is already a pure function of (a stream of
WorldState, a list of guns), and the Java battle only supplies where the
states come from. Guns keep their own internal history, so a replayed stream in
order is a complete movement history. [MEASURED] (offline_range.nim,
header comment; commit 974528d).
- Speed and sample count. 8 fixtures / 1,770 ticks / ~92 k virtual bullets /
13 guns replay in 2.9 s at ~32 k virtual bullets/s — roughly 70× faster and
100× more samples than a live gauntlet. [MEASURED] (commit
974528d). - Acceptance test.
common_libs/tests/acceptance_offline_vs_online.nimrecords one live round, replays it offline, and compares per-gun virtual hit counts. It passed 12/12 deterministic guns (Tsetlin is separately labelled stochastic becausetmLearnOnecallsrand()). Two real live-loop ordering quirks had to be modelled to reach that:run()callsgo()before the aim/fire block, sotickBulletsresolves against the next tick's scan; and if the target dies during thatgo(), the final tick's spawn+resolution is skipped. [MEASURED] (commit974528d; a passing run is preserved in/tmp/ab_logs3/test_acceptance.log: 12/12, 534-tick round). - Acceptance test: the equivalence is now proven and stable. An earlier
version of this report recorded the test as flaky — typically 11/12 on
unmodified HEAD, with the mismatching gun moving between runs (KNN, then
WallBounce) — and guessed it was a live/offline boundary race. That guess was
wrong. The root cause was a real replay bug: the offline replay
spawned gun 13 (TMSelect) while the live rack has
EnableTmSelector = falseand never does. The sharedVirtualTrackerring is order-sensitive, so gun 13's extra 4 bullets/tick permuted the per-tick resolution order of every other gun, shifting the learning guns' observations. Closing gun 13's ready gate offline made the live and offline KNN traces byte-identical (904/904 lines, empty diff). The fix mirrors the live rack in the replay — no tick exclusion, no tolerance loosening. Stability: 5/5 consecutive runs report 12/12 exact, each with the death boundary included. So "offline == online" is exact on these runs. A flaky proof had hidden a real bug. [MEASURED]. - General lesson for rack A/B. Because the ring is order-sensitive, any rack A/B that disables a gun also removes that gun's 4 spawns/tick from the shared ring, which perturbs the resolution order — and therefore the learning observations — of every other gun. That is a confound to record for anyone repeating these experiments. [INFERRED].
1.2 The fixture sets and what each is good for
There are 20 JSONL fixtures under tools/fixtures/. They fall into four
groups. [MEASURED] (tools/fixtures/, DRUSSGT_FIXTURES.md,
TR_BRIDGE_FIXTURES.md).
(a) 8 synthetic fixtures — ground truth by construction.
common_libs/gun_harness/offline_range.nim generates them; each has a known
rule:
| Fixture | Generator / rule | What it is good for |
|---|---|---|
stationary |
fixed enemy | ceiling: any correct gun scores ~100% |
constant-velocity |
straight line, no walls | does the gun lead a moving target at all |
circular |
constant turn 3°/tick | circular/accel models |
wall-bounce |
specular reflection off all 4 walls | wall-aware prediction |
oscillator |
east 30 ticks, west 30 ticks | phase/timing |
random-walk |
seeded ±15°/tick jitter | generality under noise |
decel-before-turn |
cruise → full stop → pivot 3×45° → accelerate | stop-shot detection |
energy-threshold-turner |
KNOWN RULE: straight while energy ≥ 30, hard 20°/tick turn while energy < 30, energy = max(5, 50 − 0.5·t) |
can a learner find a readable high-level rule; threshold crosses at t=41 |
The energy-threshold turner is the falsifiable one: the label is literally a predicate over the 11-bit Gray-coded energy field, so a learner that reads energy can be shown to have read the right variable (see §6.3).
(b) 2 classic-Robocode contrast fixtures. contrast_stationary_sittingduck
and contrast_straightline — trivial motion, used as a sanity/ceiling check.
[MEASURED].
(c) 5 classic-Robocode DrussGT captures. Real, unmodified DrussGT
3.1.4159 movement captured from Robocode 1.9.5.5 via
tools/robocode_fixture_capture/ (file DRUSSGT_FIXTURES.md). Opponents:
SpinBot, RamFire, Crazy, Corners, and a DrussGT mirror; 28,797 ticks. The
coordinate conversion is validated to 0.000–0.001° on every fixture by
recomputing the direction implied by (heading, speed) against the recorded
per-tick displacement. These are OPEN-LOOP and perfect-information: the
replayed DrussGT never dodges our bullets, and the observer reports true
positions every tick (unlike the live bot's stale between-scan WorldState).
They are therefore optimistic and good for relative gun ranking, not absolute
hit rates. [MEASURED] (DRUSSGT_FIXTURES.md).
(d) 5 closed-loop Tank Royale DrussGT captures. The real DrussGT jar playing
Tank Royale through tools/robocode_shim/, captured by
tools/robocode_shim/src/robocode_shim/TrBattleCapture.java (file
TR_BRIDGE_FIXTURES.md). Primary file tr_drussgt_vs_modularbot.jsonl:
15 rounds, 20,026 ticks, ModularBot fired 1,134 shots. At capture time DrussGT
was reacting to our real bullets. Closed-loop proven, not asserted:
tools/robocode_shim/analyze_closed_loop.py event-locks |Δheading| to
ModularBot's fire times (a heat-limited near-metronome, median interval
14 ticks) and gets an oscillating response with the fire period; the
cross-correlation peaks at r = +0.111, lag 12, permutation p = 0.005 (null
peak mean +0.016), and the own-fire control is flat, so the lock is
enemy-driven, not internal cadence. Still perfect-information, and
open-loop at replay time — "closed_loop" describes the capture, not a later
replay. TR angle conversion residual is ~1.5° mean (vs 0.000° classic) because
the TR server moves along the pre-turn heading. [MEASURED]
(TR_BRIDGE_FIXTURES.md).
| Fixture | Source | rounds | ticks | adversary / note |
|---|---|---|---|---|
circular … energy-threshold-turner |
synthetic | — | 150–260 | 8 known-rule trajectories |
contrast_stationary_sittingduck, contrast_straightline |
classic-robocode | 1 each | 1,451 | sanity contrasts |
drussgt_vs_spinbot |
classic-robocode | 20 | 5,002 | open-loop, perfect-info |
drussgt_vs_ramfire |
classic-robocode | 20 | 3,240 | open-loop, perfect-info |
drussgt_vs_crazy |
classic-robocode | 20 | 9,025 | open-loop, perfect-info |
drussgt_vs_corners |
classic-robocode | 20 | 4,975 | open-loop, perfect-info |
drussgt_vs_drussgt |
classic-robocode | 2 | 6,555 | open-loop, perfect-info (mirror) |
tr_drussgt_vs_modularbot |
tr-bridge | 15 | 20,026 | closed-loop at capture, perfect-info |
tr_drussgt_vs_modularbot_shield |
tr-bridge | 10 | 12,629 | shield on |
tr_drussgt_vs_spinbot / _crazy / _corners |
tr-bridge | 10 each | 10,824 / 11,507 / 2,575 | closed-loop at capture |
1.3 The live boss — the real DrussGT jar
tools/robocode_shim/ runs the unmodified DrussGT.jar (159,289 bytes,
md5 5cd6015dcc6d6da8a7e6aeecb1fec211) as a Tank Royale bot. The classic
robocode.* API is a thin delegation layer over the public
IBasicRobotPeer/IAdvancedRobotPeer seam, so the shim reuses the genuine
robocode.jar and implements only the 75-method peer interface
(ClassicPeer), plus BotHost/ThreadManagerFix. DrussGT compiles with zero
shim API symbols and runs real battles; the EnergyDome shield is disabled by
default (pure wave surfer). Known physics divergences (move/turn ordering,
distance bookkeeping, etc.) are enumerated in section 5.9 of
tools/robocode_shim/README.md. [MEASURED].
The boss is far stronger than us: in the capture battle DrussGT beat ModularBot
1447–300 over 15 rounds (ModularBot won round 5 only), firing 1,400 bullets
at a 12.1% hit rate against ModularBot's 1,134 bullets at 5.3%. That is
the number the whole gun rack is trying to move. [MEASURED] (commit
17c99f5, TR_BRIDGE_FIXTURES.md).
1.4 A/B methodology — server-side per-run hit rate, never scores
The A/B that decides configs uses server-side ground truth, not the bot's
own counters and not scores. [MEASURED] (/tmp/analyze.py,
/tmp/compare.py, /tmp/ab_logs3/FINAL_AB.txt):
- The battle runner writes a per-shot events sidecar (
/tmp/events_*.json) withfire/hit/ damage events stamped with the server's per-round bullet id (GunEngine.nextBulletId).26b66cbproved per-gun attribution: the server assigns the id once and reuses it onBulletFired,BulletHitBot,BulletHitWall,BulletHitBullet; hits that arrive before the fire event (client priority 70 > 60) are deferred. 99.9% of shots and 99.8% of hits were attributed in that session. - A config is judged on its per-run real hit rate (
hits/shotsfor one battle), and two configs are called different only if their per-run ranges do not overlap. This is why the overnight A/B reports "SEPARATED" or "OVERLAP" for every pair rather than a single pooled p-value. - Scores are not used to judge configs. Single-run scores swing by a
couple of hundred points: the 13 shipped-config runs span 175–526
(s.d. ≈ 105, range 351;
/tmp/battle_base_r*.log), so a ~210-point 2-s.d. band swamps any plausible config effect. The gate study (3c90a59) made the same call: no per-adversary score delta exceeded the ~300-point run-to-run noise band.
The bot-side per-gun attribution (realShots/realHits in
/tmp/gun_stats_base_r*.jsonl) covers ~87% of the server's total shots
uniformly (13 runs: 220/3,157 = 6.97% attributed vs 251/3,612 = 6.95%
server-side), so it is used for per-gun ranking but the server sidecar is the
ground truth for config decisions. [MEASURED] (re-aggregated from
/tmp/events_base_r*.json and /tmp/gun_stats_base_r*.jsonl).
2. The metric lesson: virtual hit rate is a poor ranker, not a proxy for real hit rate
This is the most important conceptual result of the night and it invalidates a naive reading of every offline table in this report.
[MEASURED] Aggregating the 13 shipped-config runs against the live DrussGT
boss (/tmp/gun_stats_base_r{1..13}.jsonl; python3 /tmp/agg2.py base), the
Spearman rank correlation between a gun's virtual hit rate and its real
hit rate is
Spearman(virtual rank, real rank) = -0.374 (n = 13 guns, all with ≥10 real shots)
The correlation is weak and sign-unstable — the honest headline is that virtual hit rate is a poor ranker, not a reliable inverse. The table below is what a poor ranker looks like: on this run set the ordering it produces tracks the opposite of the real ordering, but that does not hold on other run sets (see the full measurement set after the table). An earlier version of this report stated this as a clean inversion; a later 15-run measurement refuted that.
| Gun | Selected (ticks) | Real hits/shots | Real % | Virtual % |
|---|---|---|---|---|
| Linear | 2,707 | 9/84 | 10.7 | 10.2 |
| Circular | 4,051 | 17/172 | 9.9 | 11.9 |
| KNN | 4,886 | 11/122 | 9.0 | 7.5 |
| Pattern | 14,786 | 50/582 | 8.6 | 12.0 |
| Accel | 9,205 | 28/382 | 7.3 | 12.1 |
| AvgLead | 5,119 | 14/200 | 7.0 | 12.3 |
| GuessFactor | 2,345 | 5/72 | 6.9 | 10.3 |
| DecayGF | 1,360 | 3/47 | 6.4 | 9.2 |
| WallBounce | 7,618 | 18/288 | 6.2 | 12.9 |
| StopShot | 3,348 | 8/132 | 6.1 | 12.6 |
| Tsetlin | 3,284 | 6/103 | 5.8 | 12.9 |
| Displace | 2,664 | 4/75 | 5.3 | 12.3 |
| HeadOn | 16,975 | 47/898 | 5.2 | 8.6 |
Read the top and bottom: Tsetlin, WallBounce and StopShot have the highest virtual rates (12.6–12.9%) and near-bottom real rates (5.8–6.2%); Linear and KNN sit at 10.2% / 7.5% virtual but 10.7% / 9.0% real. On this run set the virtual ordering inverts the real one — but because the sign flips on other run sets (next paragraph), the safe reading is that the virtual ranking is uninformative about the real ranking, not that it is reliably inverted.
Why this matters for selection. What has kept the rack alive is the
selector's floor/tie hedging, not its ranking: removing the floor
(GUN_SELECTOR_FLOOR=0.0, config floor00) drops the rack from 6.95% to
5.08% at 175 dmg/run (vs 251) over 788 shots. [MEASURED]
(/tmp/compare.py). So the selector is useful because it refuses to commit to
a bad field, not because its virtual-rate ordering is good.
The headline: across run sets the correlation is sign-unstable, so it is near
zero on average — not robustly negative. The full set of independent
Spearman measurements (13-run source /tmp/agg2.py base; 5-run configs
/tmp/ab_logs3/FINAL_AB.txt; 15-run paired baseline §3) is:
| Run set / aggregation | Spearman |
|---|---|
13-run base, shipped relative+path |
−0.374 |
| 15-run paired baseline (different but equally defensible aggregation) | +0.335 |
5-run 12-round A/B, relative+path |
+0.522 |
5-run A/B, absolute+point |
−0.371 |
5-run A/B, absolute+path |
−0.073 |
5-run A/B, relative+point |
−0.037 |
Two opposite signs on large samples (−0.374 over 13 runs, +0.335 over 15 runs, all over the same 13 guns) mean virtual hit rate is not a reliable inverse of real hit rate. It is a poor ranker: weak correlation, sign flipping between run sets, near zero on average. The practical conclusion is unchanged — do not build a ranking rule on it — but an earlier version of this report overstated the mechanism as an inversion; the later 15-run measurement refuted that. [MEASURED] + [INFERRED] (the six numbers are measured; "poor ranker / near zero on average" is the reasoning).
2.1 The metric A/B: point vs path (this one is real, and it is selection)
GUN_VBULLET_METRIC picks how a virtual bullet is scored
(common_libs/gun_harness/virtual_bullets.nim):
bmPoint— resolve at the fire-time aim distance and score that single point. Measures prediction accuracy.bmPath(shipped) — fly the ray to the wall and test each swept segment against the target radius. Measures hypothetical hit chance.
[MEASURED] Live A/B against the boss, 5 battles × 12 rounds, one frozen
binary (commit 3b5d70b; per-run detail in /tmp/ab_logs/FINAL_ANALYSIS.txt):
| Metric | Shots | Hits | Real hit rate | Per-run rates | Spearman |
|---|---|---|---|---|---|
| point | 4,660 | 219 | 4.70% | 5.53 / 5.30 / 4.92 / 3.16 / 4.57 | −0.04 |
| path | 4,834 | 359 | 7.43% | 6.76 / 8.20 / 8.24 / 6.55 / 7.30 | +0.52 |
The distributions do not overlap: path's worst run (6.55%) beats point's best (5.53%). +2.73 pp, +58% relative, z = 5.56, p < 0.0001. Range distributions were identical (~460–478 px), so this is not a range confound.
The gain is selection, not better gun learning. Under point every gun's
virtual rate is compressed into 0.6–4.4%, so HeadOn sits inside the 2 pp tie
margin and takes 72.6% of selection ticks / 76.9% of shots while ranking
11th of 13 by real hit rate (2.3%). Under path the band widens to 4.7–13.7%
and HeadOn's shot share falls to 35.9%, so Pattern/Accel/WallBounce get picked.
The counterfactual confirms it: applying the point model's per-gun real rates
to the path model's shot mix yields 7.65%, i.e. essentially the whole observed
gain. [MEASURED] (commit 3b5d70b).
Offline range total moves the same way: 34.3% under point
(/tmp/final_range.txt, 35,636/104,000) vs 50.8% under path
(/tmp/range_path.txt, 52,770/103,938). The offline totals are inflated by the
perfect-information synthetic fixtures (three of them score 100% under path for
every gun), so the offline totals are not comparable to live rates — only to
each other.
2.2 The selector-threshold A/B: absolute vs relative
The legacy thresholds were calibrated for a rate scale that does not exist.
[MEASURED] offline replay of a fogged live WorldState vs DrussGT
(1,397 selection ticks, /tmp/ab_logs3/selector_diag.txt):
| Config | Floor fires | HeadOn selection share | bestRate med |
|---|---|---|---|
| absolute + point | 53.0% | 69.1% | 8.0% |
| relative + point | 21.2% | 43.5% | 5.25% |
| absolute + path | 3.0% | 23.1% | 24.0% |
| relative + path (shipped) | 8.4% | 24.2% | 16.75% |
The 0.10 absolute floor fires on 53.0% of point-metric ticks and forces
HeadOn, whose real rate was 2.0–4.4%. (An earlier claim that the floor fires
always is refuted: it is 53%, because bestRate is a max over
gun×power-bin and an occasional ≥50-sample bin clears 10%.)
The scale-aware replacement (commit dea4dcb):
RelTieMargin = 0.20 (tie band is a fraction of bestRate), FloorPeakFrac = 0.25 (floor fires only if the field collapsed vs its own recent peak over a
256-tick window, counting only guns with ≥ MinObsBeforeCompete = 50 samples),
and pooled-over-bins ranking instead of max-over-bins.
Live A/B, 3 runs × 10 rounds (commit dea4dcb, /tmp/ab_logs3/FINAL_AB.txt):
| Config | Per-run rates | Pooled | vs absolute+point |
|---|---|---|---|
| absolute + point | 3.66 / 2.45 / 5.01 | 3.76% | — |
| absolute + path | 7.55 / 8.21 / 6.83 | 7.57% | SEPARATED (p<0.0001) |
| relative + point | 7.66 / 6.18 / 5.79 | 6.59% | SEPARATED |
| relative + path (shipped) | 7.15 / 7.55 / 6.90 | 7.21% | SEPARATED |
absolute+path is nominally 0.35 pp above relative+path, but they overlap
(p = 0.64); so do relative+point and both path configs. The metric is the
dominant lever; under path the two threshold models are statistically tied.
relative was shipped because it is the principled scale-aware fix, works
under both metrics, and prevents the point-metric catastrophe if anyone
switches back. [MEASURED].
3. Per-gun performance on real numbers
The final per-gun table, shipped config, 13 runs vs the live DrussGT boss,
3,612 server-side shots, 6.95% overall (server sidecar; per-run rates 6.77 /
5.90 / 6.57 / 6.34 / 6.10 / 7.59 / 2.90 / 7.95 / 7.49 / 9.18 / 8.44 / 7.48 /
6.42%; 251 dmg/run). Per-gun rows are the bot-side attribution over the same
runs. The same binary on a different 15-run set gives 6.18% (events
6.16%, 200 dmg/run, §3 pruning baseline), so every rate here is quoted with its
run count — a single figure is not definitive. [MEASURED] (commit 2c94dc2;
/tmp/gun_stats_base_r*.jsonl, re-aggregated with /tmp/agg2.py base;
/tmp/events_base_r*.json).
| Verdict | Gun | Real hits/shots | Real % | Virtual % | Selected |
|---|---|---|---|---|---|
| KEEP | Linear | 9/84 | 10.7 | 10.2 | 2,707 |
| KEEP | Circular | 17/172 | 9.9 | 11.9 | 4,051 |
| KEEP | KNN | 11/122 | 9.0 | 7.5 | 4,886 |
| KEEP | Pattern | 50/582 | 8.6 | 12.0 | 14,786 |
| KEEP | Accel | 28/382 | 7.3 | 12.1 | 9,205 |
| KEEP | AvgLead | 14/200 | 7.0 | 12.3 | 5,119 |
| MARGINAL | GuessFactor | 5/72 | 6.9 | 10.3 | 2,345 |
| MARGINAL | DecayGF | 3/47 | 6.4 | 9.2 | 1,360 |
| MARGINAL | WallBounce | 18/288 | 6.2 | 12.9 | 7,618 |
| MARGINAL | StopShot | 8/132 | 6.1 | 12.6 | 3,348 |
| BELOW — KEEP | Tsetlin | 6/103 | 5.8 | 12.9 | 3,284 |
| BELOW — KEEP | Displace | 4/75 | 5.3 | 12.3 | 2,664 |
| FLOOR — STAYS | HeadOn | 47/898 | 5.2 | 8.6 | 16,975 |
Verdicts.
- KEEP: Linear, Circular, KNN, Pattern, Accel, AvgLead. These six are at or above the 6.95% overall, yet their virtual rates are mid-pack to low: the metric's three favourites (Tsetlin 12.9%, WallBounce 12.9%, StopShot 12.6%) are near the bottom of the real ranking, while the real leader (Linear) sits at 10.2% virtual. Further evidence the virtual ranking is uninformative about the real ranking (and, on this run set, roughly its opposite).
- MARGINAL: GuessFactor, DecayGF, WallBounce, StopShot. Within ~1 pp of overall on small N (47–288 shots). They are not obviously worth deleting, but they have not earned a larger share.
- BELOW OVERALL — but KEEP: Tsetlin, Displace. Both sit below the 6.95% overall on small N (75–103 shots), which an earlier version of this report read as an implied recommendation to drop. That was tested and refuted: 15 paired runs per variant against DrussGT (identical seeds, 8 rounds, same binary) gave baseline 3,238 shots / 6.18% (events 6.16%) / 200 dmg/run; Tsetlin disabled 3,522 shots / 5.76% (events 5.71%) / 197 dmg/run; and Tsetlin+Displace disabled 3,478 shots / 5.46% (events 5.37%) / 183 dmg/run. Paired permutation tests: −0.34 pp (p = 0.57) and −0.70 pp (p = 0.21); the per-run distributions completely overlap, and a Crazy (non-surfer) control showed no separation either. Removing the measured-worst real performers is therefore neutral-to-slightly-negative on both hit rate and damage. With sd ≈ 1.8 pp a definitive claim would need far more runs, so keep the full rack — being below overall does not justify removal. [MEASURED].
- HeadOn MUST STAY despite being lowest (5.2%). It is the floor fallback:
when the field collapses the selector returns gun 0. Disabling the floor
measurably hurt — 5.08% / 175 dmg vs 6.95% / 251 dmg (config
floor00, 788 shots). Do not delete HeadOn to improve the per-gun average; that average is computed over shots it only gets because nothing better was available. [MEASURED] (/tmp/compare.py).
The 16-candidate ranking A/B found no winner. Runtime knobs were added to
the selector (GUN_SELECTOR_WINDOW, MINOBS, TIE, FLOOR, POOL, RANK,
SHRINK, SEED; rankScore supports mean/Wilson/UCB/Thompson/shrinkage), all
defaulting to the shipped values. 16 candidates were A/B'd against the boss.
None credibly beat the shipped config; every candidate's per-run interval
overlaps base, and the nominal "winners" are ≤0.6 SE apart on far fewer shots.
[MEASURED] (commit 2c94dc2; /tmp/compare.py):
| Config | Runs | Shots | Rate % | dmg/run | Spearman |
|---|---|---|---|---|---|
| base (shipped) | 14* | 3,612 | 6.95 | 251 | −0.374 |
| tie00 (TIE=0.0) | 2 | 540 | 7.04 | 242 | +0.018 |
| win50 (WINDOW=50) | 2 | 559 | 6.08 | 216 | +0.588 |
| wilson (RANK=wilson) | 13 | 3,045 | 6.67 | 196 | +0.088 |
| thompson (RANK=thompson) | 2 | 531 | 4.90 | 166 | +0.083 |
| maxbin (POOL=0) | 2 | 574 | 5.23 | 198 | −0.264 |
| minobs20 (MINOBS=20) | 2 | 535 | 5.98 | 212 | +0.144 |
| tie05 (TIE=0.05) | 12 | 3,357 | 6.20 | 222 | −0.060 |
| tie10 (TIE=0.10) | 2 | 556 | 6.65 | 233 | −0.150 |
| tie40 (TIE=0.40) | 2 | 583 | 6.35 | 216 | −0.160 |
| floor10 (FLOOR=0.10) | 6 | 1,694 | 6.49 | 232 | −0.578 |
| t05f10 (TIE=0.05, FLOOR=0.10) | 4 | 1,073 | 5.50 | 190 | −0.041 |
| wilf10 (RANK=wilson, FLOOR=0.10) | 4 | 1,252 | 6.71 | 271 | −0.410 |
| floor00 (FLOOR=0.0) | 3 | 788 | 5.08 | 175 | +0.055 |
| f00t05 (FLOOR=0, TIE=0.05) | 3 | 540 | 6.48 | 153 | +0.226 |
| f00wil (FLOOR=0, RANK=wilson) | 2 | 595 | 6.72 | 221 | −0.116 |
| f00w50 (FLOOR=0, WINDOW=50) | 2 | 593 | 6.58 | 210 | +0.178 |
* compare.py counts 14 events files, but run 14 has no fire events; the 13 runs with data carry all 3,612 shots.
No ranking rule produced a stable, useful correlation. The best Spearman in the table (win50, +0.588) is on 2 runs / 559 shots; none of the 16 tested rules recovered a sign-stable signal. The shipped config's −0.374 is the largest single-run estimate but is contradicted in sign by the +0.335 over 15 runs, so a single Spearman value on one run set is not a reliable estimate. [MEASURED] + [INFERRED].
3.1 Offline range: which gun wins which trajectory family
Shipped path metric, /tmp/range_path.txt (52,770/103,938 = 50.8%). Cells
are hit-% per gun per fixture; bold = best gun for that fixture. 400 shots
per gun per fixture.
| Fixture | HeadOn | Linear | Tsetlin | Circular | GuessF | Pattern | WallBn | Accel | StopSh | Displ | AvgLead | DecayG | KNN |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| circular | 12 | 29 | 29 | 100 | 15 | 66 | 34 | 100 | 27 | 18 | 54 | 13 | 73 |
| constant-velocity | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| contr-SittingDuck | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| contr-StraightLine | 43 | 100 | 72 | 100 | 100 | 100 | 94 | 98 | 84 | 98 | 100 | 100 | 77 |
| decel-before-turn | 79 | 100 | 90 | 100 | 100 | 100 | 100 | 100 | 98 | 100 | 100 | 100 | 78 |
| classC-corners | 6 | 9 | 23 | 17 | 9 | 17 | 12 | 19 | 19 | 22 | 13 | 10 | 4 |
| classC-crazy | 4 | 36 | 24 | 34 | 35 | 32 | 54 | 35 | 22 | 26 | 39 | 27 | 22 |
| classC-mirror | 12 | 6 | 7 | 6 | 6 | 9 | 6 | 7 | 8 | 5 | 8 | 7 | 8 |
| classC-ramfire | 35 | 50 | 32 | 54 | 50 | 54 | 49 | 49 | 35 | 30 | 56 | 52 | 46 |
| classC-spinbot | 3 | 58 | 10 | 50 | 58 | 37 | 48 | 50 | 13 | 46 | 52 | 58 | 42 |
| energy-threshold-turner | 58 | 34 | 39 | 100 | 34 | 88 | 38 | 100 | 40 | 44 | 63 | 34 | 42 |
| oscillator | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| random-walk | 26 | 67 | 64 | 28 | 63 | 41 | 68 | 26 | 61 | 67 | 55 | 62 | 47 |
| stationary | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| TR-corners | 3 | 30 | 28 | 33 | 39 | 26 | 33 | 34 | 30 | 34 | 56 | 29 | 32 |
| TR-crazy | 5 | 7 | 12 | 16 | 2 | 14 | 19 | 28 | 11 | 20 | 27 | 2 | 3 |
| TR-ModularBot | 15 | 12 | 18 | 13 | 12 | 12 | 14 | 15 | 21 | 12 | 13 | 11 | 7 |
| TR-MB-shield | 4 | 10 | 32 | 27 | 10 | 28 | 32 | 28 | 24 | 37 | 21 | 27 | 5 |
| TR-spinbot | 3 | 21 | 20 | 21 | 37 | 37 | 25 | 30 | 27 | 14 | 24 | 15 | 11 |
| wall-bounce | 0 | 49 | 44 | 48 | 37 | 41 | 100 | 60 | 41 | 34 | 52 | 34 | 28 |
Aggregated hit-% by family (400 shots/gun/fixture):
| Family | HeadOn | Linear | Tsetlin | Circular | GuessF | Pattern | WallBn | Accel | StopSh | Displ | AvgLead | DecayG | KNN |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| synthetic (8) | 59 | 72 | 71 | 84 | 69 | 79 | 80 | 86 | 71 | 70 | 78 | 68 | 71 |
| classic DrussGT (5) | 12 | 32 | 19 | 32 | 32 | 30 | 34 | 32 | 20 | 26 | 34 | 31 | 25 |
| TR bridge (5) | 6 | 16 | 22 | 22 | 15 | 23 | 24 | 27 | 23 | 23 | 28 | 17 | 12 |
| all 20 | 35 | 51 | 47 | 57 | 49 | 55 | 56 | 59 | 48 | 50 | 57 | 49 | 46 |
Which gun wins which trajectory (offline, path metric):
- Stationary / constant-velocity / oscillator / decel-before-turn / straight-line: every useful gun ≥ 94% (perfect-information + path metric makes these uninformative). Only HeadOn (43–79%) and KNN (77–78%) stand out as weak.
- Circular: Circular and Accel 100% by construction; KNN 73, Pattern 66.
- Wall-bounce: WallBounce 100, Accel 60, AvgLead 52 — the only fixture where WallBounce is dominant.
- Random-walk: WallBounce 68, Displace 67, Linear 67, Tsetlin 64; Accel collapses to 26.
- Energy-threshold rule: Circular and Accel 100, Pattern 88, AvgLead 63. Tsetlin 39 — it beat Linear (34) but is far from reading the rule.
- Classic DrussGT (a real surfer): WallBounce 34 and AvgLead 34 at the top, HeadOn 12 at the bottom. Cornered surfing is the one case where Tsetlin (23) leads.
- TR bridge DrussGT: AvgLead 28 overall; Accel 28 on TR-crazy, AvgLead 56 on TR-corners, StopShot 21 on the ModularBot mirror, Pattern 37 on TR-spinbot.
Do not read the offline winner as the rack verdict. The offline range's classic DrussGT order (WallBounce 34 top, KNN 25 low) is close to the inverse of the real order (KNN 9.0% third, WallBounce 6.2% ninth). The offline range is excellent for catching structural bugs (§6) and for per-trajectory sanity, and poor as a selector signal. [MEASURED] + [INFERRED].
4. Verdict table
| Gun | Family / model | Real % (13 runs) | Offline all-20 % | Verdict |
|---|---|---|---|---|
| Linear | constant-velocity lead | 10.7 | 51 | KEEP — this is what a surfer cannot defeat; keep warm |
| Circular | constant-turn lead | 9.9 | 57 | KEEP |
| KNN | k-NN on motion history | 9.0 | 46 | KEEP — offline under-rates it badly |
| Pattern | pattern replay | 8.6 | 55 | KEEP |
| Accel | acceleration-aware lead | 7.3 | 59 | KEEP |
| AvgLead | windowed average lead | 7.0 | 57 | KEEP |
| GuessFactor | GF histogram | 6.9 | 49 | MARGINAL — small N, no clear edge |
| DecayGF | recency-weighted GF | 6.4 | 49 | MARGINAL |
| WallBounce | wall-reflection model | 6.2 | 56 | MARGINAL — offline favourite, real underperformer |
| StopShot | deceleration/stop point | 6.1 | 48 | MARGINAL |
| Tsetlin | Tsetlin-Machine correction | 5.8 | 47 | KEEP — below overall; pruning tested neutral-to-negative (§3) |
| Displace | displacement vector | 5.3 | 50 | KEEP — below overall; pruning tested neutral-to-negative (§3) |
| HeadOn | aim at current position | 5.2 | 35 | KEEP — mandatory floor fallback |
| TMSelect | TM mixture-of-experts gate | — | — | DISABLED (EnableTmSelector = false) — see §6.4 |
5. What "worth keeping" means, and what is not proven
The KEEP/MARGINAL/BELOW split is a statement about a single adversary (a wave surfer), judged on the shipped config, on a few hundred real shots per gun. It is a starting point, not a final ranking. In particular, BELOW does not mean "drop": pruning the measured-worst real performers (Tsetlin, then Tsetlin+Displace) was tested in 15 paired runs each and was neutral-to-slightly-negative on both hit rate and damage (§3), so the verdict is keep the full rack. The concrete caveats are in §7.
6. Bugs found and fixed tonight (why earlier rack verdicts were wrong)
Five of these changed the rack ordering; all are [MEASURED] from the commit messages and the offline range.
6.1 The GF family aimed at the fire-time RADIUS, not the angle
guess_factor, decay_gf and knn_gun aimed at the fire-time distance. But
the virtual-bullet metric resolves a bullet at its aim-point distance and
scores that single point against the enemy's position on that tick, so with any
radial target motion the bullet stopped at the wrong radius and missed even
with a perfect angle. Angle-only prediction is structurally unscoreable
under the point metric. Two competing hypotheses were tested and both
refuted: (a) MEA range too narrow — 0 clamped shots out of 837/849/957, with
required offsets peaking at ~33° against MEA 28.1–46.7°, and arcsin(8/bulletSpeed)
correctly uses max robot speed, not the hit radius; (b) wrong GF peak — a
sweep of every constant GF value showed the oracle-best constant offset was
only 6% on circular, 4% on wall-bounce, 7.5% on random-walk. Learning was fine
too (~850–960 observations per fixture, 0 starved waves). [MEASURED]
(commit 7f706e5).
Fix: a self-consistent constant-velocity lead_forecast.nim base, so the
histogram learns the residual and the aim point lands at the right radius;
also fixed linear.nim (it did a one-shot extrapolation and never iterated its
flight time). Before/after, offline: circular GF 6→23, DecayGF 6→21;
wall-bounce GF 0→60.2, DecayGF 0→60.2; constant-velocity GF/DecayGF/KNN
26→100; random-walk GF 0→53; StraightLine GF 8→77. The oracle-best constant GF
moved 6%→20% (circular), 4%→57% (wall-bounce), 7.5%→49% (random-walk),
proving the structural fix independently of tuning. Honest trade-off: on
the 5 real DrussGT surfer captures the GF family regressed (GuessFactor
108→55, DecayGF 108→76, KNN 101→74 hits/2000) because a linear base is a poor
model for a surfer and the residual histogram is noisier than the old
total-lead histogram.
That regression was then recovered by blending the range between a
radial-only forecast and the geometric one by the measured radial fraction
(radialFrac), keeping the constant-velocity bearing. Nine candidate bases
were measured and rejected with numbers (velocity scaling 0.8 recovered DrussGT
but destroyed wall-bounce 241→20; radial-only range wall-bounce 241→140;
short-window average worse than both; reversal/speed gates weaker than the
blend). Result (hits/2000): classic-5 GF 55→171, DecayGF 76→100; TR-5 GF
9→86, DecayGF 4→87; synthetic-10 GF 2702→2717. The only figure below the old
base is classic-5 DecayGF (108→100, within noise). [MEASURED] (commit
e2ca2fc).
6.2 The wave queues were starved 1-push-vs-4-pops
predict() stored one wave per tick while onResult() popped one per
resolved bullet (~4/tick), so the queue drained within a few dozen ticks and
~3 of every 4 resolutions returned without learning; the survivor paired with
a same-tick wave (bearingDelta ≈ 0), pinning the histogram at centre.
Proof: GF.vHits == HeadOn.vHits and DecayGF.vHits == HeadOn.vHits
byte-for-byte in every one of 50 rounds — GF, DecayGF and KNN had
degenerated into HeadOn clones. Fix: per-bin FIFO with an O(1) head cursor, at
most one push per (tick, bin). Also maxBullets 2,048→8,192: the rack spawns
52 bullets/tick so the ring wrapped every ~39 ticks while a long power-3 shot
needs ~90, silently discarding unresolved bullets and biasing every measured
hit rate by range; a droppedBullets counter was added. After the fix
vDropped = 0 and vStarved = 0 across all 48 recorded rounds. [MEASURED]
(commit 0cc6821).
6.3 Tsetlin's clauses saturated at ~714 included literals each
Tsetlin.vHits was byte-for-byte equal to Linear.vHits in every measured
round of every run because its learned correction was always exactly 0.
Root cause: tmLearnOne rewarded included true literals unconditionally,
omitting Granmo's (c=0, lk=1) → toward Exclude counter-force, so true
literals ratcheted toward Include forever; Type II was unreachable dead code
with the wrong direction; resource allocation was an |error| heuristic
instead of Granmo's (T − clip(v,−T,T))/(2T); the label baseline had a
factor-2 shrink (error = δ − 2c, fixed point c = δ/2); hits zeroed their
residual; the enemy-energy feature was duplicated (state.selfEnergy fed
where WorldState.enemyEnergy exists, so energy rules were literally
unrepresentable); and tmEvalClause needed Granmo Eq. 6 (all-Exclude clause
outputs 1 during learning, 0 during classification) or fix #1 deadlocks every
clause at empty.
Measured effect (energy-threshold-turner, seed 1): mean included
literals/clause 714.0 → 13.8; active clauses 100/100 → 53/100; nonzero
corrections 8/764 → 708/764; Tsetlin virtual hits 27/400 → 69/400 (Linear
43/400). Tsetlin now learns but is not yet competitive with Linear —
the regression head is untuned, flagged as follow-up rather than claimed as a
win. [MEASURED] (commit 8937000; /tmp/ab_logs3/final_test_tsetlin_gun.log).
6.4 The TM classifier gun did not earn its slot (but its clauses are real)
A Tsetlin-Machine mixture-of-experts gate over HeadOn/Linear/Circular/WallBounce/Accel was built with the corrected feedback and labelled by which expert's prediction was closest to the actual enemy position (an exact, supervised, per-shot label — no delayed credit). It loses to the best of its own experts offline on nearly every fixture, and against DrussGT it cost real performance:
baseline (path + relative) 7.56% real hit rate, 157 dmg
+ power fix 7.47%, 239 dmg
+ power fix + TM selector 5.59%, 133 dmg
It was selected on 806 ticks and fired 24 real shots at 4.2%. It ships disabled
(EnableTmSelector = false; code and wiring kept intact). However, the gate
latched onto meaningful structure: on the energy-threshold turner, HeadOn's
clauses key on the energy bits (the rule's own driving variable) while
Circular keys on distance/velocity. So the TM learned something real and
interpretable; it simply could not beat "always pick the best expert".
[INFERRED] root cause: the closest-expert label is noisy because several
experts are near-tied, and under the path metric the winner varies by power bin
while the gate sees one shared per-tick input, so a one-vs-rest gate over a
saturated 870-bit clause space has no margin to exploit. A standalone Granmo
classifier on the same encoding reaches ~99% on the rule but, per the
counterfactual probe, does not read energy (follow rate 24% high / 62% mean
— statistically identical at 1, 2 and 10 frames), so even the "it learned the
rule" claim is limited to ~99% accuracy, not to a readable energy threshold.
The best recovered proposition was !g9 ∧ !g8 (energy < 25.6, not the labelled
30) — a genuine simple threshold, but not the ensemble's decision mechanism.
[MEASURED] (commits 57b2ac3, d5061ee; test_tm_pattern_learning.nim).
6.5 The selector thresholds were absolute on a rescaled metric
Covered in §2.2: the 0.10 absolute floor fired on 53.0% of point-metric ticks
and forced HeadOn (real 2.0–4.4%, 11th of 13); HeadOn selection share fell
69.1% → 43.5% under relative thresholds (and 23.1% → 24.2% under path). Also:
bestGun was first-index-wins argmax, so HeadOn at index 0 silently won every
tie until the random tie-break landed (343e631); bestPower had the same
absolute-40% defect (below).
6.6 Power selection was stuck at power 1.0 (MinHitRate = 0.40)
bestPower used an absolute MinHitRate = 0.40 bar. Measured per-bin
virtual rates show no bin ever clears 40%, so 11 of 14 guns were stuck at bin 0
(power 1.0) even where higher bins were comparable:
Linear p1.0 44% p1.5 39% p2.0 30% p3.0 29% old bin 0 -> new bin 3
Accel p1.0 44% p1.5 40% p2.0 26% p3.0 29% old bin 1 -> new bin 3
Pattern p1.0 50% p1.5 40% p2.0 27% p3.0 12% old bin 1 -> new bin 2
Replaced with a scale-aware PowerBarFrac = 0.50 (a dimensionless fraction of
the gun's own best-bin rate); 13 of 14 selections now pick heavier bullets.
Real effect vs DrussGT (8 rounds × 3 runs): hit rate unchanged (7.56% → 7.47%),
damage +52% (157 → 239 per run) and rounds end faster. [MEASURED]
(commit 57b2ac3).
Smaller fixes in the same family: bestPower on a cold gun returned the
highest bin (empty bin satisfied the count == 0 clause); fitnessFor
aggregated enemies in nondeterministic hash order; stop_shot had an
unreachable deceleration branch and several guns had tick-only caches that made
all four power bins return bin 0's lead (e536900). [MEASURED].
6.7 The selector's tie-break was not actually random
randomize() was reached only incidentally, through the Tsetlin gun's
constructor, so ties resolved identically across process restarts — the
"random" tie-break was effectively deterministic. Now fixed with an explicit
startup seed plus a GUN_SELECTOR_SEED override. Evidence: unseeded runs vary
across processes, seeded runs are identical. [MEASURED].
6.8 The arrival-accuracy tie-band does not beat the shipped band
Hypothesis. Rank by path (robust, keeps its measured advantage) but narrow
the tied random draw by arrival accuracy (point): a gun whose ray sweeps
the target generously can sit in the band while its bullets arrive badly, so
making the band informative should improve the real hit rate without removing
the load-bearing randomness.
Implementation (GUN_SELECTOR_TIEBREAK, common_libs/gun_harness/,
default off): each virtual bullet is additionally scored with the point
model at the exact tick it reaches its aim distance into a parallel
GunFitness.pointBins window, and the path tie band is narrowed to the guns
within GUN_SELECTOR_POINT_TIE (default 0.5) of the best in-band point rate. The
uniform random draw over the narrowed band is kept. GUN_SELECTOR_TIEBREAK=point
selects it; =commit is the no-randomness control. off performs no parallel
scoring at all, so the shipped path is byte-identical. The recording rule and
the ranking rule are pinned by common_libs/tests/test_selector_tiebreak.nim
(19 checks, no battle).
Live A/B vs the real DrussGT, one frozen binary, 7 runs x 7 rounds per arm,
server-side events sidecar (~200-260 shots/run). d = base - arm (positive =
arm worse); p is the exact two-sided permutation test on per-run rates.
[MEASURED] (/tmp/battle_tb*_r*.log, /tmp/events_tb*_r*.json):
| Arm | Runs | Shots | Real % | dmg/run | d | p |
|---|---|---|---|---|---|---|
tbbase (shipped) |
7 | 4,128 | 7.17 | 175 | — | — |
tbpt (path + point narrow) |
7 | 3,938 | 7.08 | 165 | +0.14 | 0.88 |
tbpc (=commit control) |
7 | 3,759 | 4.44 | 98 | +2.74 | 0.0012 |
tbpt25 (point margin 0.25) |
7 | 3,683 | 5.59 | 119 | +1.65 | 0.20 |
tbtie05 (TIE=0.05) |
7 | 3,937 | 5.84 | 133 | +1.49 | 0.11 |
tbtie40 (TIE=0.40) |
7 | 3,917 | 6.28 | 144 | +1.00 | 0.25 |
tbwin50 (WINDOW=50) |
7 | 3,983 | 6.05 | 139 | +1.20 | 0.11 |
tbfloor10 (FLOOR=0.10) |
7 | 3,829 | 5.33 | 118 | +2.12 | 0.11 |
The mechanism DID fire — the tie-break re-shaped the selection mix (over all 7
runs: Pattern 24%→16%, Accel 6%→16%, Tsetlin ~0%→13% under tbpt; the
no-randomness control tbpc collapses to HeadOn 43% vs 27%) — but it did not
improve the real hit rate: 7.08% vs 7.17%, fully
overlapping per-run ranges (base 5.29-8.73, arm 3.71-10.39), p = 0.88. The
no-randomness control tbpc is significantly worse (4.44%, p = 0.0012),
which independently replicates the earlier "commitment to the virtual best
costs real hit rate" result and validates that the arm was live. Every knob
variant (TIE, FLOOR, WINDOW) is also nominally worse than the shipped
values, none credibly better. Verdict: clean negative — the shipped selector
is unchanged (GUN_SELECTOR_TIEBREAK defaults to off).
This is consistent with §2: the virtual rate is a poor ranker, and the selector's value is its floor/tie hedging, not the ordering it computes. Narrowing the band with a second virtual statistic changes which guns are drawn without making that draw any better.
7. Known caveats and open problems
Stated without hedging.
-
The headline per-gun numbers come from ONE adversary, a wave surfer. HeadOn is genuinely bad against surfers, so part of the rack ordering may be matchup-specific. A SpinBot guard was inconclusive: ModularBot fires only 17–31 real shots/run against a fast bot because the range-aware firing gate is strict at long range, so the guard had little power (Wilson looked better, 18.5% vs 8.6%, but on 70–92 shots with a 5–33% spread). [MEASURED] (commit
2c94dc2). A second, independent adversary at scale is missing. -
Per-gun real N is small. 47–898 shots per gun; n < 200 gives roughly ±5 pp across a 3–15% spread. Single-gun ordering is indicative, not definitive. The KEEP/MARGINAL/BELOW boundaries should be treated as soft.
-
The fixtures are perfect-information and therefore optimistic. Every fixture is an observer capture with true positions every tick; the classic set is additionally open-loop (replayed DrussGT never dodges our bullets). Absolute offline hit rates are inflated by an unknown amount; only relative comparisons are safe.
-
The virtual metric is a poor ranker, not a reliable inverse. The correlation with real hit rate is weak and sign-unstable across run sets: −0.374 over the 13-run base, +0.335 over a 15-run paired baseline (different aggregation), +0.522 on the 5-run
relative+pathset, and −0.371 / −0.073 / −0.037 on the other 5-run configs. Two opposite signs on large samples mean it is near zero on average, not reliably anti-correlated; an earlier version of this report overstated it as an inversion and a later measurement refuted that. 16 candidate ranking rules all overlapped the shipped config, so none produced a stable, useful correlation. The selector's value lives in its floor/tie hedging (5.08% without the floor vs 6.95% with it), not in its ranking. [MEASURED] + [INFERRED]. -
The TM classifier gun did not earn its slot. It cost real performance (7.47% → 5.59%, 133 dmg) despite showing interpretable energy structure in its clauses (§6.4). It is disabled; re-enabling requires a fix to the gate margin/label problem, not more training.
-
Real-hit-rate-driven selection is not viable yet. Only the selected gun fires, so unselected guns get near-zero real shots (GuessFactor 20, Linear 24 vs HeadOn 733 in the point A/B); noise is fatal (n = 470 at p = 10% gives ±2.8 pp, most guns n < 200 gives ±5 pp+); and real rate is conditional on when the gun was selected. A blended signal with forced exploration and shrinkage is defensible in principle but needs thousands of shots per gun across many battles. Real rate is currently best used offline as the evaluation metric — which is exactly what the A/B does. [MEASURED] (commit
dea4dcb). -
The offline==online acceptance test is fixed and stable (§1.1): the old 11/12 flakiness was a real replay bug (the replay spawned disabled gun 13, and the shared order-sensitive ring then permuted every other gun's resolution order), now fixed by mirroring the live rack. 5/5 consecutive runs give a byte-identical 12/12 with the death boundary included.
-
The selector's tie-break is now explicitly seeded (§6.7). It had been effectively non-random —
randomize()was reached only incidentally through the Tsetlin gun's constructor — so ties resolved identically across process restarts. Fixed with an explicit startup seed plus aGUN_SELECTOR_SEEDoverride; seeded runs are reproducible, unseeded runs vary. -
The firing gate is not the bottleneck. The shipped range-aware gate does not beat a fixed 2.0° gate on hit rate (55.8% vs 57.9%, ~1.5 σ), though it fires 22–28% more shots. No per-adversary score delta exceeded the ~300-point run-to-run noise band. [MEASURED] (commit
3c90a59). -
The boss is ~2.3× more accurate than the whole rack (12.1% vs 5.3% in the capture). Closing that gap is the point of the rack; the current best single gun is 10.7%.
8. Reproduction
Commands recorded in the commits and tool READMEs. (I was instructed not to run builds/tests while writing this report; these are the documented invocations, not a fresh verification by me.)
# Offline gun range over all 20 fixtures, shipped path metric (default):
nim c -r common_libs/tests/run_range.nim
# Point metric for comparison:
GUN_VBULLET_METRIC=point nim c -r common_libs/tests/run_range.nim
# Add timing:
nim c -r common_libs/tests/run_range.nim --timing
# Selector diagnostics (floor/tie/bestRate/HeadOn-share) on a fixture:
GUN_SELECTOR_MODE=relative nim c -d:release -r \
common_libs/tests/analyze_selector.nim tools/fixtures/drussgt_vs_spinbot.jsonl
# Offline == online acceptance (fixed; stable exact 12/12 — see §1.1):
nim c -r common_libs/tests/acceptance_offline_vs_online.nim
# Tsetlin gun clause sparsity / divergence:
nim c -r common_libs/tests/test_tsetlin_gun.nim
# TM readability (standalone Granmo classifier on the energy-threshold rule):
nim c -r common_libs/tests/test_tm_pattern_learning.nim
# Live boss (real DrussGT jar; jars stay out of git, see the README):
# tools/robocode_shim/run_bridge_battle.sh <bot_dir> <rounds> <capture.jsonl>
# Closed-loop evidence for the TR captures:
python3 tools/robocode_shim/analyze_closed_loop.py \
tools/fixtures/tr_drussgt_vs_modularbot.jsonl \
tools/robocode_shim/evidence/tr_drussgt_vs_modularbot.events.json
Per-gun aggregation scripts used for the tables above:
python3 /tmp/agg2.py base (virtual-vs-real + Spearman),
python3 /tmp/compare.py (server-side per-run A/B + overlap).
Raw range output: /tmp/range_path.txt (path), /tmp/final_range.txt (point).