# Gun Rack Analysis — ModularBot guns vs the real DrussGT boss **Date:** 2026-09-21 **Bot:** ModularBot (13 active guns + 1 disabled TM classifier gun) **Shipped config:** virtual-bullet metric `path`, selector thresholds `relative` (`GUN_VBULLET_METRIC` default `path`, `GUN_SELECTOR_MODE` default `relative`) **Rewritten:** this file supersedes the 2026-09-20 version, whose numbers predated per-gun real-hit attribution, the offline gun range, and the live DrussGT boss. No number from the old version is retained unless it was re-measured below. Evidence tags used throughout: - **[MEASURED]** — I read it from a recorded artifact (commit message, `/tmp` output, a log, or a source file). The provenance is named every time. - **[INFERRED]** — reasoning from measured facts; explicitly not measured. A note on provenance: the measurements below were produced overnight in commits `343e631` … `2c94dc2` on branch `research/lead-targeting`. Each commit message records the numbers, the hypotheses it refuted, and its caveats. Raw artifacts are in `/tmp` (`/tmp/gun_stats_base_r*.jsonl`, `/tmp/events_base_r*.json`, `/tmp/range_path.txt`, `/tmp/ab_logs/FINAL_ANALYSIS.txt`, `/tmp/ab_logs3/FINAL_AB.txt`, `/tmp/ab_logs3/selector_diag.txt`). Where a commit message and a raw artifact disagree slightly, both are stated. --- ## 1. The test infrastructure The numbers are only as good as the rig that produced them, so the rig is described first. ### 1.1 The offline gun range `common_libs/gun_harness/offline_range.nim` replays a recorded `seq[WorldState]` through the **same** `VirtualTracker` (`common_libs/gun_harness/virtual_bullets.nim`) that the live bot drives. The claim is not "an approximation of the live metric", it is "the same metric": virtual-bullet fitness is already a pure function of (a stream of `WorldState`, a list of guns), and the Java battle only supplies where the states come from. Guns keep their own internal history, so a replayed stream in order is a complete movement history. **[MEASURED]** (`offline_range.nim`, header comment; commit `974528d`). - **Speed and sample count.** 8 fixtures / 1,770 ticks / ~92 k virtual bullets / 13 guns replay in 2.9 s at ~32 k virtual bullets/s — roughly **70× faster and 100× more samples** than a live gauntlet. **[MEASURED]** (commit `974528d`). - **Acceptance test.** `common_libs/tests/acceptance_offline_vs_online.nim` records one live round, replays it offline, and compares per-gun virtual hit counts. It passed **12/12 deterministic guns** (Tsetlin is separately labelled stochastic because `tmLearnOne` calls `rand()`). Two real live-loop ordering quirks had to be modelled to reach that: `run()` calls `go()` before the aim/fire block, so `tickBullets` resolves against the *next* tick's scan; and if the target dies during that `go()`, the final tick's spawn+resolution is skipped. **[MEASURED]** (commit `974528d`; a passing run is preserved in `/tmp/ab_logs3/test_acceptance.log`: 12/12, 534-tick round). - **Honest caveat — the acceptance test is currently FLAKY.** On the *unmodified HEAD* source it normally reaches only **11/12**, e.g. KNN 81 online vs 71 offline, and the mismatching gun moves between runs (KNN, then WallBounce). It is a live/offline boundary race, pre-existing, and not caused by the selector work (the replay never calls the selector). Treat "offline == online" as **strong but not exact until the race is fixed**. **[MEASURED]** (commit `2c94dc2`). The 12/12 runs above are real; they were lucky runs. ### 1.2 The fixture sets and what each is good for There are **20 JSONL fixtures** under `tools/fixtures/`. They fall into four groups. **[MEASURED]** (`tools/fixtures/`, `DRUSSGT_FIXTURES.md`, `TR_BRIDGE_FIXTURES.md`). **(a) 8 synthetic fixtures — ground truth by construction.** `common_libs/gun_harness/offline_range.nim` generates them; each has a known rule: | Fixture | Generator / rule | What it is good for | |---|---|---| | `stationary` | fixed enemy | ceiling: any correct gun scores ~100% | | `constant-velocity` | straight line, no walls | does the gun lead a moving target at all | | `circular` | constant turn 3°/tick | circular/accel models | | `wall-bounce` | specular reflection off all 4 walls | wall-aware prediction | | `oscillator` | east 30 ticks, west 30 ticks | phase/timing | | `random-walk` | seeded ±15°/tick jitter | generality under noise | | `decel-before-turn` | cruise → full stop → pivot 3×45° → accelerate | stop-shot detection | | `energy-threshold-turner` | **KNOWN RULE**: straight while `energy ≥ 30`, hard 20°/tick turn while `energy < 30`, `energy = max(5, 50 − 0.5·t)` | can a learner find a readable high-level rule; threshold crosses at t=41 | The energy-threshold turner is the falsifiable one: the label is literally a predicate over the 11-bit Gray-coded energy field, so a learner that reads energy can be shown to have read the *right* variable (see §6.3). **(b) 2 classic-Robocode contrast fixtures.** `contrast_stationary_sittingduck` and `contrast_straightline` — trivial motion, used as a sanity/ceiling check. **[MEASURED]**. **(c) 5 classic-Robocode DrussGT captures.** Real, unmodified DrussGT 3.1.4159 movement captured from Robocode 1.9.5.5 via `tools/robocode_fixture_capture/` (file `DRUSSGT_FIXTURES.md`). Opponents: SpinBot, RamFire, Crazy, Corners, and a DrussGT mirror; 28,797 ticks. The coordinate conversion is validated to **0.000–0.001°** on every fixture by recomputing the direction implied by (heading, speed) against the recorded per-tick displacement. **These are OPEN-LOOP and perfect-information:** the replayed DrussGT never dodges *our* bullets, and the observer reports true positions every tick (unlike the live bot's stale between-scan `WorldState`). They are therefore optimistic and good for *relative* gun ranking, not absolute hit rates. **[MEASURED]** (`DRUSSGT_FIXTURES.md`). **(d) 5 closed-loop Tank Royale DrussGT captures.** The real DrussGT jar playing Tank Royale through `tools/robocode_shim/`, captured by `tools/robocode_shim/src/robocode_shim/TrBattleCapture.java` (file `TR_BRIDGE_FIXTURES.md`). Primary file `tr_drussgt_vs_modularbot.jsonl`: 15 rounds, 20,026 ticks, ModularBot fired 1,134 shots. At capture time DrussGT was reacting to *our* real bullets. **Closed-loop proven, not asserted:** `tools/robocode_shim/analyze_closed_loop.py` event-locks `|Δheading|` to ModularBot's fire times (a heat-limited near-metronome, median interval 14 ticks) and gets an oscillating response with the fire period; the cross-correlation peaks at **r = +0.111, lag 12, permutation p = 0.005** (null peak mean +0.016), and the **own-fire control is flat**, so the lock is enemy-driven, not internal cadence. **Still perfect-information**, and open-loop *at replay time* — "closed_loop" describes the capture, not a later replay. TR angle conversion residual is ~1.5° mean (vs 0.000° classic) because the TR server moves along the pre-turn heading. **[MEASURED]** (`TR_BRIDGE_FIXTURES.md`). | Fixture | Source | rounds | ticks | adversary / note | |---|---|---:|---:|---| | `circular` … `energy-threshold-turner` | synthetic | — | 150–260 | 8 known-rule trajectories | | `contrast_stationary_sittingduck`, `contrast_straightline` | classic-robocode | 1 each | 1,451 | sanity contrasts | | `drussgt_vs_spinbot` | classic-robocode | 20 | 5,002 | open-loop, perfect-info | | `drussgt_vs_ramfire` | classic-robocode | 20 | 3,240 | open-loop, perfect-info | | `drussgt_vs_crazy` | classic-robocode | 20 | 9,025 | open-loop, perfect-info | | `drussgt_vs_corners` | classic-robocode | 20 | 4,975 | open-loop, perfect-info | | `drussgt_vs_drussgt` | classic-robocode | 2 | 6,555 | open-loop, perfect-info (mirror) | | `tr_drussgt_vs_modularbot` | tr-bridge | 15 | 20,026 | **closed-loop at capture**, perfect-info | | `tr_drussgt_vs_modularbot_shield` | tr-bridge | 10 | 12,629 | shield on | | `tr_drussgt_vs_spinbot` / `_crazy` / `_corners` | tr-bridge | 10 each | 10,824 / 11,507 / 2,575 | closed-loop at capture | ### 1.3 The live boss — the real DrussGT jar `tools/robocode_shim/` runs the **unmodified** `DrussGT.jar` (159,289 bytes, md5 `5cd6015dcc6d6da8a7e6aeecb1fec211`) as a Tank Royale bot. The classic `robocode.*` API is a thin delegation layer over the public `IBasicRobotPeer`/`IAdvancedRobotPeer` seam, so the shim reuses the genuine `robocode.jar` and implements only the 75-method peer interface (`ClassicPeer`), plus `BotHost`/`ThreadManagerFix`. DrussGT compiles with zero shim API symbols and runs real battles; the EnergyDome shield is disabled by default (pure wave surfer). Known physics divergences (move/turn ordering, distance bookkeeping, etc.) are enumerated in section 5.9 of `tools/robocode_shim/README.md`. **[MEASURED]**. The boss is far stronger than us: in the capture battle DrussGT beat ModularBot **1447–300 over 15 rounds** (ModularBot won round 5 only), firing 1,400 bullets at a **12.1%** hit rate against ModularBot's 1,134 bullets at **5.3%**. That is the number the whole gun rack is trying to move. **[MEASURED]** (commit `17c99f5`, `TR_BRIDGE_FIXTURES.md`). ### 1.4 A/B methodology — server-side per-run hit rate, never scores The A/B that decides configs uses **server-side ground truth**, not the bot's own counters and not scores. **[MEASURED]** (`/tmp/analyze.py`, `/tmp/compare.py`, `/tmp/ab_logs3/FINAL_AB.txt`): 1. The battle runner writes a per-shot **events sidecar** (`/tmp/events_*.json`) with `fire` / `hit` / damage events stamped with the server's per-round bullet id (`GunEngine.nextBulletId`). `26b66cb` proved per-gun attribution: the server assigns the id once and reuses it on `BulletFired`, `BulletHitBot`, `BulletHitWall`, `BulletHitBullet`; hits that arrive before the fire event (client priority 70 > 60) are deferred. 99.9% of shots and 99.8% of hits were attributed in that session. 2. A config is judged on its **per-run real hit rate** (`hits/shots` for one battle), and two configs are called different only if their per-run ranges **do not overlap**. This is why the overnight A/B reports "SEPARATED" or "OVERLAP" for every pair rather than a single pooled p-value. 3. **Scores are not used to judge configs.** Single-run scores swing by a couple of hundred points: the 13 shipped-config runs span **175–526** (s.d. ≈ 105, range 351; `/tmp/battle_base_r*.log`), so a ~210-point 2-s.d. band swamps any plausible config effect. The gate study (`3c90a59`) made the same call: no per-adversary score delta exceeded the ~300-point run-to-run noise band. The bot-side per-gun attribution (`realShots`/`realHits` in `/tmp/gun_stats_base_r*.jsonl`) covers ~87% of the server's total shots uniformly (13 runs: 220/3,157 = 6.97% attributed vs 251/3,612 = 6.95% server-side), so it is used for per-gun *ranking* but the server sidecar is the ground truth for config decisions. **[MEASURED]** (re-aggregated from `/tmp/events_base_r*.json` and `/tmp/gun_stats_base_r*.jsonl`). --- ## 2. The metric lesson: virtual hit rate is NOT a proxy for real hit rate This is the most important conceptual result of the night and it invalidates a naive reading of every offline table in this report. **[MEASURED]** Aggregating the 13 shipped-config runs against the live DrussGT boss (`/tmp/gun_stats_base_r{1..13}.jsonl`; `python3 /tmp/agg2.py base`), the Spearman rank correlation between a gun's **virtual** hit rate and its **real** hit rate is ``` Spearman(virtual rank, real rank) = -0.374 (n = 13 guns, all with ≥10 real shots) ``` It is not weak — it is **inverted**. The guns with the highest virtual rates have among the lowest real rates, and vice versa: | Gun | Selected (ticks) | Real hits/shots | Real % | Virtual % | |---|---:|---:|---:|---:| | Linear | 2,707 | 9/84 | **10.7** | 10.2 | | Circular | 4,051 | 17/172 | **9.9** | 11.9 | | KNN | 4,886 | 11/122 | **9.0** | 7.5 | | Pattern | 14,786 | 50/582 | **8.6** | 12.0 | | Accel | 9,205 | 28/382 | **7.3** | 12.1 | | AvgLead | 5,119 | 14/200 | **7.0** | 12.3 | | GuessFactor | 2,345 | 5/72 | **6.9** | 10.3 | | DecayGF | 1,360 | 3/47 | **6.4** | 9.2 | | WallBounce | 7,618 | 18/288 | **6.2** | 12.9 | | StopShot | 3,348 | 8/132 | **6.1** | 12.6 | | Tsetlin | 3,284 | 6/103 | **5.8** | 12.9 | | Displace | 2,664 | 4/75 | **5.3** | 12.3 | | HeadOn | 16,975 | 47/898 | **5.2** | 8.6 | Read the top and bottom: **Tsetlin, WallBounce and StopShot have the highest virtual rates (12.6–12.9%) and near-bottom real rates (5.8–6.2%); Linear and KNN sit at 10.2% / 7.5% virtual but 10.7% / 9.0% real.** The ranking the virtual metric produces is not merely uninformative, it points the wrong way. **Why this matters for selection.** What has kept the rack alive is the selector's **floor/tie hedging**, not its ranking: removing the floor (`GUN_SELECTOR_FLOOR=0.0`, config `floor00`) drops the rack from 6.95% to **5.08%** at 175 dmg/run (vs 251) over 788 shots. **[MEASURED]** (`/tmp/compare.py`). So the selector is useful because it refuses to commit to a bad field, not because its virtual-rate ordering is good. **The virtual metric appears anti-correlated no matter which config you pick.** Measured Spearman per config (5-run 12-round A/B, `/tmp/ab_logs3/FINAL_AB.txt`): `absolute+point` −0.371, `absolute+path` −0.073, `relative+point` −0.037, `relative+path` +0.522. But on the large 13-run base set the shipped config (`relative+path`) is **−0.374**. The sign **flips between run sets**, which is itself the finding: the correlation is unstable, so no ranking rule built on it can be trusted. **[MEASURED]** + **[INFERRED]** (the flip is measured; the conclusion is reasoning). ### 2.1 The metric A/B: point vs path (this one is real, and it is selection) `GUN_VBULLET_METRIC` picks how a virtual bullet is scored (`common_libs/gun_harness/virtual_bullets.nim`): - `bmPoint` — resolve at the fire-time aim distance and score that single point. Measures prediction accuracy. - `bmPath` (**shipped**) — fly the ray to the wall and test each swept segment against the target radius. Measures hypothetical hit chance. **[MEASURED]** Live A/B against the boss, 5 battles × 12 rounds, one frozen binary (commit `3b5d70b`; per-run detail in `/tmp/ab_logs/FINAL_ANALYSIS.txt`): | Metric | Shots | Hits | Real hit rate | Per-run rates | Spearman | |---|---:|---:|---:|---|---:| | point | 4,660 | 219 | 4.70% | 5.53 / 5.30 / 4.92 / 3.16 / 4.57 | −0.04 | | path | 4,834 | 359 | **7.43%** | 6.76 / 8.20 / 8.24 / 6.55 / 7.30 | +0.52 | The distributions **do not overlap**: path's worst run (6.55%) beats point's best (5.53%). +2.73 pp, +58% relative, z = 5.56, p < 0.0001. Range distributions were identical (~460–478 px), so this is not a range confound. **The gain is selection, not better gun learning.** Under `point` every gun's virtual rate is compressed into 0.6–4.4%, so HeadOn sits inside the 2 pp tie margin and takes **72.6% of selection ticks / 76.9% of shots** while ranking 11th of 13 by real hit rate (2.3%). Under `path` the band widens to 4.7–13.7% and HeadOn's shot share falls to 35.9%, so Pattern/Accel/WallBounce get picked. The counterfactual confirms it: applying the point model's per-gun real rates to the path model's shot mix yields 7.65%, i.e. essentially the whole observed gain. **[MEASURED]** (commit `3b5d70b`). Offline range total moves the same way: **34.3%** under point (`/tmp/final_range.txt`, 35,636/104,000) vs **50.8%** under path (`/tmp/range_path.txt`, 52,770/103,938). The offline totals are inflated by the perfect-information synthetic fixtures (three of them score 100% under path for every gun), so the offline totals are *not* comparable to live rates — only to each other. ### 2.2 The selector-threshold A/B: absolute vs relative The legacy thresholds were calibrated for a rate scale that does not exist. **[MEASURED]** offline replay of a *fogged live* `WorldState` vs DrussGT (1,397 selection ticks, `/tmp/ab_logs3/selector_diag.txt`): | Config | Floor fires | HeadOn selection share | bestRate med | |---|---:|---:|---:| | absolute + point | 53.0% | 69.1% | 8.0% | | relative + point | 21.2% | 43.5% | 5.25% | | absolute + path | 3.0% | 23.1% | 24.0% | | relative + path (shipped) | 8.4% | 24.2% | 16.75% | The `0.10` absolute floor fires on **53.0%** of point-metric ticks and forces HeadOn, whose real rate was 2.0–4.4%. (An earlier claim that the floor fires *always* is **refuted**: it is 53%, because `bestRate` is a max over gun×power-bin and an occasional ≥50-sample bin clears 10%.) The scale-aware replacement (commit `dea4dcb`): `RelTieMargin = 0.20` (tie band is a fraction of `bestRate`), `FloorPeakFrac = 0.25` (floor fires only if the field collapsed vs its own recent peak over a 256-tick window, counting only guns with ≥ MinObsBeforeCompete = 50 samples), and pooled-over-bins ranking instead of max-over-bins. Live A/B, 3 runs × 10 rounds (commit `dea4dcb`, `/tmp/ab_logs3/FINAL_AB.txt`): | Config | Per-run rates | Pooled | vs `absolute+point` | |---|---|---:|---| | absolute + point | 3.66 / 2.45 / 5.01 | 3.76% | — | | absolute + path | 7.55 / 8.21 / 6.83 | 7.57% | SEPARATED (p<0.0001) | | relative + point | 7.66 / 6.18 / 5.79 | 6.59% | SEPARATED | | relative + path (**shipped**) | 7.15 / 7.55 / 6.90 | 7.21% | SEPARATED | `absolute+path` is nominally 0.35 pp above `relative+path`, but they **overlap** (p = 0.64); so do `relative+point` and both path configs. The **metric** is the dominant lever; under `path` the two threshold models are statistically tied. `relative` was shipped because it is the principled scale-aware fix, works under both metrics, and prevents the point-metric catastrophe if anyone switches back. **[MEASURED]**. --- ## 3. Per-gun performance on real numbers The final per-gun table, shipped config, **13 runs vs the live DrussGT boss, 3,612 server-side shots, 6.95% overall** (server sidecar; per-run rates 6.77 / 5.90 / 6.57 / 6.34 / 6.10 / 7.59 / 2.90 / 7.95 / 7.49 / 9.18 / 8.44 / 7.48 / 6.42%; 251 dmg/run). Per-gun rows are the bot-side attribution over the same runs. **[MEASURED]** (commit `2c94dc2`; `/tmp/gun_stats_base_r*.jsonl`, re-aggregated with `/tmp/agg2.py base`; `/tmp/events_base_r*.json`). | Verdict | Gun | Real hits/shots | Real % | Virtual % | Selected | |---|---|---:|---:|---:|---:| | **KEEP** | Linear | 9/84 | 10.7 | 10.2 | 2,707 | | **KEEP** | Circular | 17/172 | 9.9 | 11.9 | 4,051 | | **KEEP** | KNN | 11/122 | 9.0 | 7.5 | 4,886 | | **KEEP** | Pattern | 50/582 | 8.6 | 12.0 | 14,786 | | **KEEP** | Accel | 28/382 | 7.3 | 12.1 | 9,205 | | **KEEP** | AvgLead | 14/200 | 7.0 | 12.3 | 5,119 | | MARGINAL | GuessFactor | 5/72 | 6.9 | 10.3 | 2,345 | | MARGINAL | DecayGF | 3/47 | 6.4 | 9.2 | 1,360 | | MARGINAL | WallBounce | 18/288 | 6.2 | 12.9 | 7,618 | | MARGINAL | StopShot | 8/132 | 6.1 | 12.6 | 3,348 | | BELOW | Tsetlin | 6/103 | 5.8 | 12.9 | 3,284 | | BELOW | Displace | 4/75 | 5.3 | 12.3 | 2,664 | | **FLOOR — STAYS** | HeadOn | 47/898 | 5.2 | 8.6 | 16,975 | **Verdicts.** - **KEEP: Linear, Circular, KNN, Pattern, Accel, AvgLead.** These six are at or above the 6.95% overall, yet their virtual rates are mid-pack to low: the metric's three favourites (Tsetlin 12.9%, WallBounce 12.9%, StopShot 12.6%) are near the *bottom* of the real ranking, while the real leader (Linear) sits at 10.2% virtual. Further evidence the virtual ranking is inverted. - **MARGINAL: GuessFactor, DecayGF, WallBounce, StopShot.** Within ~1 pp of overall on small N (47–288 shots). They are not obviously worth deleting, but they have not earned a larger share. - **BELOW OVERALL: Tsetlin, Displace.** Below 6% on 75–103 shots. Candidates to drop or re-tune, but the N is small. - **HeadOn MUST STAY** despite being lowest (5.2%). It is the floor fallback: when the field collapses the selector returns gun 0. Disabling the floor measurably hurt — 5.08% / 175 dmg vs 6.95% / 251 dmg (config `floor00`, 788 shots). Do not delete HeadOn to improve the per-gun average; that average is computed over shots it only gets because nothing better was available. **[MEASURED]** (`/tmp/compare.py`). **The 16-candidate ranking A/B found no winner.** Runtime knobs were added to the selector (`GUN_SELECTOR_WINDOW`, `MINOBS`, `TIE`, `FLOOR`, `POOL`, `RANK`, `SHRINK`, `SEED`; `rankScore` supports mean/Wilson/UCB/Thompson/shrinkage), all defaulting to the shipped values. 16 candidates were A/B'd against the boss. None credibly beat the shipped config; every candidate's per-run interval overlaps base, and the nominal "winners" are ≤0.6 SE apart on far fewer shots. **[MEASURED]** (commit `2c94dc2`; `/tmp/compare.py`): | Config | Runs | Shots | Rate % | dmg/run | Spearman | |---|---:|---:|---:|---:|---:| | **base (shipped)** | 14* | 3,612 | **6.95** | 251 | −0.374 | | tie00 (TIE=0.0) | 2 | 540 | 7.04 | 242 | +0.018 | | win50 (WINDOW=50) | 2 | 559 | 6.08 | 216 | +0.588 | | wilson (RANK=wilson) | 13 | 3,045 | 6.67 | 196 | +0.088 | | thompson (RANK=thompson) | 2 | 531 | 4.90 | 166 | +0.083 | | maxbin (POOL=0) | 2 | 574 | 5.23 | 198 | −0.264 | | minobs20 (MINOBS=20) | 2 | 535 | 5.98 | 212 | +0.144 | | tie05 (TIE=0.05) | 12 | 3,357 | 6.20 | 222 | −0.060 | | tie10 (TIE=0.10) | 2 | 556 | 6.65 | 233 | −0.150 | | tie40 (TIE=0.40) | 2 | 583 | 6.35 | 216 | −0.160 | | floor10 (FLOOR=0.10) | 6 | 1,694 | 6.49 | 232 | −0.578 | | t05f10 (TIE=0.05, FLOOR=0.10) | 4 | 1,073 | 5.50 | 190 | −0.041 | | wilf10 (RANK=wilson, FLOOR=0.10) | 4 | 1,252 | 6.71 | 271 | −0.410 | | floor00 (FLOOR=0.0) | 3 | 788 | **5.08** | 175 | +0.055 | | f00t05 (FLOOR=0, TIE=0.05) | 3 | 540 | 6.48 | 153 | +0.226 | | f00wil (FLOOR=0, RANK=wilson) | 2 | 595 | 6.72 | 221 | −0.116 | | f00w50 (FLOOR=0, WINDOW=50) | 2 | 593 | 6.58 | 210 | +0.178 | \* compare.py counts 14 events files, but run 14 has no fire events; the 13 runs with data carry all 3,612 shots. **No ranking rule fixed the anti-correlation.** The best Spearman in the table (win50, +0.588) is on 2 runs / 559 shots. The shipped config's −0.374 over 13 runs is the most reliable estimate. **[MEASURED]**. ### 3.1 Offline range: which gun wins which trajectory family Shipped `path` metric, `/tmp/range_path.txt` (52,770/103,938 = 50.8%). Cells are hit-% per gun per fixture; **bold** = best gun for that fixture. 400 shots per gun per fixture. | Fixture | HeadOn | Linear | Tsetlin | Circular | GuessF | Pattern | WallBn | Accel | StopSh | Displ | AvgLead | DecayG | KNN | |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:| | circular | 12 | 29 | 29 | **100** | 15 | 66 | 34 | 100 | 27 | 18 | 54 | 13 | 73 | | constant-velocity | **100** | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | | contr-SittingDuck | **100** | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | | contr-StraightLine | 43 | **100** | 72 | 100 | 100 | 100 | 94 | 98 | 84 | 98 | 100 | 100 | 77 | | decel-before-turn | 79 | **100** | 90 | 100 | 100 | 100 | 100 | 100 | 98 | 100 | 100 | 100 | 78 | | classC-corners | 6 | 9 | **23** | 17 | 9 | 17 | 12 | 19 | 19 | 22 | 13 | 10 | 4 | | classC-crazy | 4 | 36 | 24 | 34 | 35 | 32 | **54** | 35 | 22 | 26 | 39 | 27 | 22 | | classC-mirror | **12** | 6 | 7 | 6 | 6 | 9 | 6 | 7 | 8 | 5 | 8 | 7 | 8 | | classC-ramfire | 35 | 50 | 32 | 54 | 50 | 54 | 49 | 49 | 35 | 30 | **56** | 52 | 46 | | classC-spinbot | 3 | **58** | 10 | 50 | 58 | 37 | 48 | 50 | 13 | 46 | 52 | 58 | 42 | | energy-threshold-turner | 58 | 34 | 39 | **100** | 34 | 88 | 38 | 100 | 40 | 44 | 63 | 34 | 42 | | oscillator | **100** | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | | random-walk | 26 | 67 | 64 | 28 | 63 | 41 | **68** | 26 | 61 | 67 | 55 | 62 | 47 | | stationary | **100** | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | | TR-corners | 3 | 30 | 28 | 33 | 39 | 26 | 33 | 34 | 30 | 34 | **56** | 29 | 32 | | TR-crazy | 5 | 7 | 12 | 16 | 2 | 14 | 19 | **28** | 11 | 20 | 27 | 2 | 3 | | TR-ModularBot | 15 | 12 | 18 | 13 | 12 | 12 | 14 | 15 | **21** | 12 | 13 | 11 | 7 | | TR-MB-shield | 4 | 10 | 32 | 27 | 10 | 28 | 32 | 28 | 24 | **37** | 21 | 27 | 5 | | TR-spinbot | 3 | 21 | 20 | 21 | **37** | 37 | 25 | 30 | 27 | 14 | 24 | 15 | 11 | | wall-bounce | 0 | 49 | 44 | 48 | 37 | 41 | **100** | 60 | 41 | 34 | 52 | 34 | 28 | Aggregated hit-% by family (400 shots/gun/fixture): | Family | HeadOn | Linear | Tsetlin | Circular | GuessF | Pattern | WallBn | Accel | StopSh | Displ | AvgLead | DecayG | KNN | |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:| | synthetic (8) | 59 | 72 | 71 | 84 | 69 | 79 | 80 | **86** | 71 | 70 | 78 | 68 | 71 | | classic DrussGT (5) | 12 | 32 | 19 | 32 | 32 | 30 | **34** | 32 | 20 | 26 | 34 | 31 | 25 | | TR bridge (5) | 6 | 16 | 22 | 22 | 15 | 23 | 24 | 27 | 23 | 23 | **28** | 17 | 12 | | all 20 | 35 | 51 | 47 | 57 | 49 | 55 | 56 | **59** | 48 | 50 | 57 | 49 | 46 | **Which gun wins which trajectory (offline, path metric):** - **Stationary / constant-velocity / oscillator / decel-before-turn / straight-line:** every useful gun ≥ 94% (perfect-information + path metric makes these uninformative). Only HeadOn (43–79%) and KNN (77–78%) stand out as weak. - **Circular:** Circular and Accel 100% by construction; KNN 73, Pattern 66. - **Wall-bounce:** WallBounce 100, Accel 60, AvgLead 52 — the only fixture where WallBounce is dominant. - **Random-walk:** WallBounce 68, Displace 67, Linear 67, Tsetlin 64; Accel collapses to 26. - **Energy-threshold rule:** Circular and Accel 100, Pattern 88, AvgLead 63. Tsetlin 39 — it beat Linear (34) but is far from reading the rule. - **Classic DrussGT (a real surfer):** WallBounce 34 and AvgLead 34 at the top, HeadOn 12 at the bottom. Cornered surfing is the one case where Tsetlin (23) leads. - **TR bridge DrussGT:** AvgLead 28 overall; Accel 28 on TR-crazy, AvgLead 56 on TR-corners, StopShot 21 on the ModularBot mirror, Pattern 37 on TR-spinbot. **Do not read the offline winner as the rack verdict.** The offline range's *classic DrussGT* order (WallBounce 34 top, KNN 25 low) is close to the **inverse** of the real order (KNN 9.0% third, WallBounce 6.2% ninth). The offline range is excellent for catching structural bugs (§6) and for per-trajectory sanity, and poor as a selector signal. **[MEASURED]** + **[INFERRED]**. --- ## 4. Verdict table | Gun | Family / model | Real % (13 runs) | Offline all-20 % | Verdict | |---|---|---:|---:|---| | Linear | constant-velocity lead | 10.7 | 51 | **KEEP — this is what a surfer cannot defeat; keep warm** | | Circular | constant-turn lead | 9.9 | 57 | **KEEP** | | KNN | k-NN on motion history | 9.0 | 46 | **KEEP — offline under-rates it badly** | | Pattern | pattern replay | 8.6 | 55 | **KEEP** | | Accel | acceleration-aware lead | 7.3 | 59 | **KEEP** | | AvgLead | windowed average lead | 7.0 | 57 | **KEEP** | | GuessFactor | GF histogram | 6.9 | 49 | MARGINAL — small N, no clear edge | | DecayGF | recency-weighted GF | 6.4 | 49 | MARGINAL | | WallBounce | wall-reflection model | 6.2 | 56 | MARGINAL — offline favourite, real underperformer | | StopShot | deceleration/stop point | 6.1 | 48 | MARGINAL | | Tsetlin | Tsetlin-Machine correction | 5.8 | 47 | BELOW — learns, not yet competitive | | Displace | displacement vector | 5.3 | 50 | BELOW | | HeadOn | aim at current position | 5.2 | 35 | **KEEP — mandatory floor fallback** | | TMSelect | TM mixture-of-experts gate | — | — | **DISABLED (`EnableTmSelector = false`)** — see §6.4 | --- ## 5. What "worth keeping" means, and what is not proven The KEEP/MARGINAL/BELOW split is a statement about a **single adversary (a wave surfer)**, judged on the shipped config, on a few hundred real shots per gun. It is a starting point, not a final ranking. The concrete caveats are in §7. --- ## 6. Bugs found and fixed tonight (why earlier rack verdicts were wrong) Five of these changed the rack ordering; all are [MEASURED] from the commit messages and the offline range. ### 6.1 The GF family aimed at the fire-time RADIUS, not the angle `guess_factor`, `decay_gf` and `knn_gun` aimed at the fire-time distance. But the virtual-bullet metric resolves a bullet at its **aim-point distance** and scores that single point against the enemy's position on that tick, so with any radial target motion the bullet stopped at the wrong radius and missed even with a perfect angle. **Angle-only prediction is structurally unscoreable under the point metric.** Two competing hypotheses were tested and **both refuted**: (a) MEA range too narrow — 0 clamped shots out of 837/849/957, with required offsets peaking at ~33° against MEA 28.1–46.7°, and `arcsin(8/bulletSpeed)` correctly uses max robot *speed*, not the hit radius; (b) wrong GF peak — a sweep of every constant GF value showed the oracle-best constant offset was only 6% on circular, 4% on wall-bounce, 7.5% on random-walk. Learning was fine too (~850–960 observations per fixture, 0 starved waves). **[MEASURED]** (commit `7f706e5`). Fix: a self-consistent constant-velocity `lead_forecast.nim` base, so the histogram learns the **residual** and the aim point lands at the right radius; also fixed `linear.nim` (it did a one-shot extrapolation and never iterated its flight time). Before/after, offline: circular GF 6→23, DecayGF 6→21; wall-bounce GF 0→60.2, DecayGF 0→60.2; constant-velocity GF/DecayGF/KNN 26→100; random-walk GF 0→53; StraightLine GF 8→77. The oracle-best constant GF moved 6%→20% (circular), 4%→57% (wall-bounce), 7.5%→49% (random-walk), proving the structural fix independently of tuning. **Honest trade-off:** on the 5 real DrussGT surfer captures the GF family regressed (GuessFactor 108→55, DecayGF 108→76, KNN 101→74 hits/2000) because a linear base is a poor model for a surfer and the residual histogram is noisier than the old total-lead histogram. That regression was then recovered by **blending the range** between a radial-only forecast and the geometric one by the measured radial fraction (`radialFrac`), keeping the constant-velocity bearing. Nine candidate bases were measured and rejected with numbers (velocity scaling 0.8 recovered DrussGT but destroyed wall-bounce 241→20; radial-only range wall-bounce 241→140; short-window average worse than both; reversal/speed gates weaker than the blend). Result (hits/2000): classic-5 GF 55→**171**, DecayGF 76→100; TR-5 GF 9→86, DecayGF 4→87; synthetic-10 GF 2702→2717. The only figure below the old base is classic-5 DecayGF (108→100, within noise). **[MEASURED]** (commit `e2ca2fc`). ### 6.2 The wave queues were starved 1-push-vs-4-pops `predict()` stored **one** wave per tick while `onResult()` popped one per resolved bullet (~4/tick), so the queue drained within a few dozen ticks and ~3 of every 4 resolutions returned without learning; the survivor paired with a same-tick wave (`bearingDelta ≈ 0`), pinning the histogram at centre. **Proof:** `GF.vHits == HeadOn.vHits` and `DecayGF.vHits == HeadOn.vHits` byte-for-byte in **every one of 50 rounds** — GF, DecayGF and KNN had degenerated into HeadOn clones. Fix: per-bin FIFO with an O(1) head cursor, at most one push per (tick, bin). Also `maxBullets` 2,048→8,192: the rack spawns 52 bullets/tick so the ring wrapped every ~39 ticks while a long power-3 shot needs ~90, silently discarding unresolved bullets and biasing every measured hit rate by range; a `droppedBullets` counter was added. After the fix `vDropped = 0` and `vStarved = 0` across all 48 recorded rounds. **[MEASURED]** (commit `0cc6821`). ### 6.3 Tsetlin's clauses saturated at ~714 included literals each `Tsetlin.vHits` was byte-for-byte equal to `Linear.vHits` in every measured round of every run because its learned correction was always exactly 0. Root cause: `tmLearnOne` rewarded included true literals unconditionally, omitting Granmo's `(c=0, lk=1) → toward Exclude` counter-force, so true literals ratcheted toward Include forever; Type II was unreachable dead code with the wrong direction; resource allocation was an `|error|` heuristic instead of Granmo's `(T − clip(v,−T,T))/(2T)`; the label baseline had a factor-2 shrink (`error = δ − 2c`, fixed point `c = δ/2`); hits zeroed their residual; the enemy-energy feature was duplicated (`state.selfEnergy` fed where `WorldState.enemyEnergy` exists, so energy rules were literally unrepresentable); and `tmEvalClause` needed Granmo Eq. 6 (all-Exclude clause outputs 1 during learning, 0 during classification) or fix #1 deadlocks every clause at empty. Measured effect (energy-threshold-turner, seed 1): mean included literals/clause **714.0 → 13.8**; active clauses 100/100 → 53/100; nonzero corrections 8/764 → 708/764; Tsetlin virtual hits **27/400 → 69/400** (Linear 43/400). Tsetlin now **learns** but is **not yet competitive with Linear** — the regression head is untuned, flagged as follow-up rather than claimed as a win. **[MEASURED]** (commit `8937000`; `/tmp/ab_logs3/final_test_tsetlin_gun.log`). ### 6.4 The TM classifier gun did not earn its slot (but its clauses are real) A Tsetlin-Machine mixture-of-experts gate over HeadOn/Linear/Circular/WallBounce/Accel was built with the corrected feedback and labelled by which expert's prediction was closest to the actual enemy position (an exact, supervised, per-shot label — no delayed credit). It loses to the best of its own experts offline on nearly every fixture, and against DrussGT it cost real performance: ``` baseline (path + relative) 7.56% real hit rate, 157 dmg + power fix 7.47%, 239 dmg + power fix + TM selector 5.59%, 133 dmg ``` It was selected on 806 ticks and fired 24 real shots at 4.2%. It ships disabled (`EnableTmSelector = false`; code and wiring kept intact). **However, the gate latched onto meaningful structure:** on the energy-threshold turner, HeadOn's clauses key on the **energy bits** (the rule's own driving variable) while Circular keys on distance/velocity. So the TM learned something real and interpretable; it simply could not beat "always pick the best expert". **[INFERRED]** root cause: the closest-expert label is noisy because several experts are near-tied, and under the path metric the winner varies by power bin while the gate sees one shared per-tick input, so a one-vs-rest gate over a saturated 870-bit clause space has no margin to exploit. A standalone Granmo classifier on the same encoding reaches ~99% on the rule but, per the counterfactual probe, does **not** read energy (follow rate 24% high / 62% mean — statistically identical at 1, 2 and 10 frames), so even the "it learned the rule" claim is limited to ~99% accuracy, not to a readable energy threshold. The best recovered proposition was `!g9 ∧ !g8` (energy < 25.6, not the labelled 30) — a genuine simple threshold, but not the ensemble's decision mechanism. **[MEASURED]** (commits `57b2ac3`, `d5061ee`; `test_tm_pattern_learning.nim`). ### 6.5 The selector thresholds were absolute on a rescaled metric Covered in §2.2: the `0.10` absolute floor fired on 53.0% of point-metric ticks and forced HeadOn (real 2.0–4.4%, 11th of 13); HeadOn selection share fell 69.1% → 43.5% under relative thresholds (and 23.1% → 24.2% under path). Also: `bestGun` was first-index-wins argmax, so HeadOn at index 0 silently won every tie until the random tie-break landed (`343e631`); `bestPower` had the same absolute-40% defect (below). ### 6.6 Power selection was stuck at power 1.0 (`MinHitRate = 0.40`) `bestPower` used an **absolute** `MinHitRate = 0.40` bar. Measured per-bin virtual rates show no bin ever clears 40%, so 11 of 14 guns were stuck at bin 0 (power 1.0) even where higher bins were comparable: ``` Linear p1.0 44% p1.5 39% p2.0 30% p3.0 29% old bin 0 -> new bin 3 Accel p1.0 44% p1.5 40% p2.0 26% p3.0 29% old bin 1 -> new bin 3 Pattern p1.0 50% p1.5 40% p2.0 27% p3.0 12% old bin 1 -> new bin 2 ``` Replaced with a scale-aware `PowerBarFrac = 0.50` (a dimensionless fraction of the gun's own best-bin rate); 13 of 14 selections now pick heavier bullets. Real effect vs DrussGT (8 rounds × 3 runs): hit rate unchanged (7.56% → 7.47%), **damage +52% (157 → 239 per run)** and rounds end faster. **[MEASURED]** (commit `57b2ac3`). Smaller fixes in the same family: `bestPower` on a cold gun returned the *highest* bin (empty bin satisfied the `count == 0` clause); `fitnessFor` aggregated enemies in nondeterministic hash order; `stop_shot` had an unreachable deceleration branch and several guns had tick-only caches that made all four power bins return bin 0's lead (`e536900`). **[MEASURED]**. --- ## 7. Known caveats and open problems Stated without hedging. 1. **The headline per-gun numbers come from ONE adversary, a wave surfer.** HeadOn is genuinely bad against surfers, so part of the rack ordering may be matchup-specific. A SpinBot guard was inconclusive: ModularBot fires only 17–31 real shots/run against a fast bot because the range-aware firing gate is strict at long range, so the guard had little power (Wilson looked better, 18.5% vs 8.6%, but on 70–92 shots with a 5–33% spread). **[MEASURED]** (commit `2c94dc2`). A second, independent adversary at scale is missing. 2. **Per-gun real N is small.** 47–898 shots per gun; n < 200 gives roughly ±5 pp across a 3–15% spread. Single-gun ordering is **indicative, not definitive**. The KEEP/MARGINAL/BELOW boundaries should be treated as soft. 3. **The fixtures are perfect-information and therefore optimistic.** Every fixture is an observer capture with true positions every tick; the classic set is additionally **open-loop** (replayed DrussGT never dodges our bullets). Absolute offline hit rates are inflated by an unknown amount; only relative comparisons are safe. 4. **The virtual metric is anti-correlated with real hit rate and no tested ranking rule fixed it.** Shipped config Spearman ≈ **−0.374** over 13 runs (the sign flips to +0.52 on the smaller 5-run set, so it is unstable). 16 candidate ranking rules all overlapped the shipped config. The selector's value lives in its **floor/tie hedging** (5.08% without the floor vs 6.95% with it), not in its ranking. **[MEASURED]**. 5. **The TM classifier gun did not earn its slot.** It cost real performance (7.47% → 5.59%, 133 dmg) despite showing interpretable energy structure in its clauses (§6.4). It is disabled; re-enabling requires a fix to the gate margin/label problem, not more training. 6. **Real-hit-rate-driven selection is not viable yet.** Only the selected gun fires, so unselected guns get near-zero real shots (GuessFactor 20, Linear 24 vs HeadOn 733 in the point A/B); noise is fatal (n = 470 at p = 10% gives ±2.8 pp, most guns n < 200 gives ±5 pp+); and real rate is conditional on when the gun was selected. A blended signal with forced exploration and shrinkage is defensible in principle but needs thousands of shots per gun across many battles. Real rate is currently best used **offline** as the evaluation metric — which is exactly what the A/B does. **[MEASURED]** (commit `dea4dcb`). 7. **The offline==online acceptance test is flaky** (§1.1): typically 11/12 on unmodified HEAD, with the mismatching gun varying run to run. The equivalence claim is strong-but-not-exact until the boundary race is fixed. 8. **The selector's random tie-break is not randomised in the live bot.** The shipped bot never calls `randomize()`, so the "random" sequence is fixed across process restarts (a side finding of `2c94dc2`, not fixed). 9. **The firing gate is not the bottleneck.** The shipped range-aware gate does not beat a fixed 2.0° gate on hit rate (55.8% vs 57.9%, ~1.5 σ), though it fires 22–28% more shots. No per-adversary score delta exceeded the ~300-point run-to-run noise band. **[MEASURED]** (commit `3c90a59`). 10. **The boss is ~2.3× more accurate than the whole rack** (12.1% vs 5.3% in the capture). Closing that gap is the point of the rack; the current best single gun is 10.7%. --- ## 8. Reproduction Commands recorded in the commits and tool READMEs. (I was instructed not to run builds/tests while writing this report; these are the documented invocations, not a fresh verification by me.) ```bash # Offline gun range over all 20 fixtures, shipped path metric (default): nim c -r common_libs/tests/run_range.nim # Point metric for comparison: GUN_VBULLET_METRIC=point nim c -r common_libs/tests/run_range.nim # Add timing: nim c -r common_libs/tests/run_range.nim --timing # Selector diagnostics (floor/tie/bestRate/HeadOn-share) on a fixture: GUN_SELECTOR_MODE=relative nim c -d:release -r \ common_libs/tests/analyze_selector.nim tools/fixtures/drussgt_vs_spinbot.jsonl # Offline == online acceptance (currently flaky): nim c -r common_libs/tests/acceptance_offline_vs_online.nim # Tsetlin gun clause sparsity / divergence: nim c -r common_libs/tests/test_tsetlin_gun.nim # TM readability (standalone Granmo classifier on the energy-threshold rule): nim c -r common_libs/tests/test_tm_pattern_learning.nim # Live boss (real DrussGT jar; jars stay out of git, see the README): # tools/robocode_shim/run_bridge_battle.sh # Closed-loop evidence for the TR captures: python3 tools/robocode_shim/analyze_closed_loop.py \ tools/fixtures/tr_drussgt_vs_modularbot.jsonl \ tools/robocode_shim/evidence/tr_drussgt_vs_modularbot.events.json ``` Per-gun aggregation scripts used for the tables above: `python3 /tmp/agg2.py base` (virtual-vs-real + Spearman), `python3 /tmp/compare.py` (server-side per-run A/B + overlap). Raw range output: `/tmp/range_path.txt` (path), `/tmp/final_range.txt` (point).