Files
SirRoboGarage/docs/gun_rack_analysis.md
T
SirStone 19410164f1 docs: definitive gun-rack report on real measured numbers
Replaces the stale 2026-09-20 docs, which predated per-gun real attribution,
the offline gun range and the DrussGT boss, and whose verdicts were built on
virtual hit rates that turned out to be ANTI-correlated with reality.

docs/gun_rack_analysis.md (751 lines) covers: the test infrastructure
described honestly (offline range with its flaky-acceptance caveat, the 20
fixtures and what each set is good for, the live boss, and the A/B methodology
of per-run server-side real hit rate with an explicit overlap test); the
virtual-vs-real metric lesson with Spearman -0.374 and the point-vs-path A/B;
per-gun real performance and the 16-rule ranking A/B; the offline per-fixture
gun matrix; KEEP/MARGINAL/BELOW verdicts; and the five root-cause bugs with
before/after numbers.
docs/gun_rack_summary.md (58 lines) is the verdict table plus top actions.

The '~230 point' score-noise band that has been steering methodology all night
was re-derived from the artifacts rather than asserted: the 13 shipped-config
run scores span 175-526, s.d. ~105, i.e. a ~210-point 2-s.d. band.

Caveats recorded verbatim rather than softened: the offline==online acceptance
is flaky (typically 11/12 on unmodified HEAD), fixtures are perfect-information
and therefore optimistic vs live play, per-gun real N is small so single-gun
ordering is indicative, the headline numbers come from ONE wave-surfer
adversary, and HeadOn must stay despite being lowest because it is the floor
fallback (disabling it: 5.08% / 175 dmg vs 6.95% / 251 dmg).
2026-09-21 06:37:20 +02:00

41 KiB
Raw Blame History

Gun Rack Analysis — ModularBot guns vs the real DrussGT boss

Date: 2026-09-21 Bot: ModularBot (13 active guns + 1 disabled TM classifier gun) Shipped config: virtual-bullet metric path, selector thresholds relative (GUN_VBULLET_METRIC default path, GUN_SELECTOR_MODE default relative) Rewritten: this file supersedes the 2026-09-20 version, whose numbers predated per-gun real-hit attribution, the offline gun range, and the live DrussGT boss. No number from the old version is retained unless it was re-measured below.

Evidence tags used throughout:

  • [MEASURED] — I read it from a recorded artifact (commit message, /tmp output, a log, or a source file). The provenance is named every time.
  • [INFERRED] — reasoning from measured facts; explicitly not measured.

A note on provenance: the measurements below were produced overnight in commits 343e631 … 2c94dc2 on branch research/lead-targeting. Each commit message records the numbers, the hypotheses it refuted, and its caveats. Raw artifacts are in /tmp (/tmp/gun_stats_base_r*.jsonl, /tmp/events_base_r*.json, /tmp/range_path.txt, /tmp/ab_logs/FINAL_ANALYSIS.txt, /tmp/ab_logs3/FINAL_AB.txt, /tmp/ab_logs3/selector_diag.txt). Where a commit message and a raw artifact disagree slightly, both are stated.


1. The test infrastructure

The numbers are only as good as the rig that produced them, so the rig is described first.

1.1 The offline gun range

common_libs/gun_harness/offline_range.nim replays a recorded seq[WorldState] through the same VirtualTracker (common_libs/gun_harness/virtual_bullets.nim) that the live bot drives. The claim is not "an approximation of the live metric", it is "the same metric": virtual-bullet fitness is already a pure function of (a stream of WorldState, a list of guns), and the Java battle only supplies where the states come from. Guns keep their own internal history, so a replayed stream in order is a complete movement history. [MEASURED] (offline_range.nim, header comment; commit 974528d).

  • Speed and sample count. 8 fixtures / 1,770 ticks / ~92 k virtual bullets / 13 guns replay in 2.9 s at ~32 k virtual bullets/s — roughly 70× faster and 100× more samples than a live gauntlet. [MEASURED] (commit 974528d).
  • Acceptance test. common_libs/tests/acceptance_offline_vs_online.nim records one live round, replays it offline, and compares per-gun virtual hit counts. It passed 12/12 deterministic guns (Tsetlin is separately labelled stochastic because tmLearnOne calls rand()). Two real live-loop ordering quirks had to be modelled to reach that: run() calls go() before the aim/fire block, so tickBullets resolves against the next tick's scan; and if the target dies during that go(), the final tick's spawn+resolution is skipped. [MEASURED] (commit 974528d; a passing run is preserved in /tmp/ab_logs3/test_acceptance.log: 12/12, 534-tick round).
  • Honest caveat — the acceptance test is currently FLAKY. On the unmodified HEAD source it normally reaches only 11/12, e.g. KNN 81 online vs 71 offline, and the mismatching gun moves between runs (KNN, then WallBounce). It is a live/offline boundary race, pre-existing, and not caused by the selector work (the replay never calls the selector). Treat "offline == online" as strong but not exact until the race is fixed. [MEASURED] (commit 2c94dc2). The 12/12 runs above are real; they were lucky runs.

1.2 The fixture sets and what each is good for

There are 20 JSONL fixtures under tools/fixtures/. They fall into four groups. [MEASURED] (tools/fixtures/, DRUSSGT_FIXTURES.md, TR_BRIDGE_FIXTURES.md).

(a) 8 synthetic fixtures — ground truth by construction. common_libs/gun_harness/offline_range.nim generates them; each has a known rule:

Fixture Generator / rule What it is good for
stationary fixed enemy ceiling: any correct gun scores ~100%
constant-velocity straight line, no walls does the gun lead a moving target at all
circular constant turn 3°/tick circular/accel models
wall-bounce specular reflection off all 4 walls wall-aware prediction
oscillator east 30 ticks, west 30 ticks phase/timing
random-walk seeded ±15°/tick jitter generality under noise
decel-before-turn cruise → full stop → pivot 3×45° → accelerate stop-shot detection
energy-threshold-turner KNOWN RULE: straight while energy ≥ 30, hard 20°/tick turn while energy < 30, energy = max(5, 50 − 0.5·t) can a learner find a readable high-level rule; threshold crosses at t=41

The energy-threshold turner is the falsifiable one: the label is literally a predicate over the 11-bit Gray-coded energy field, so a learner that reads energy can be shown to have read the right variable (see §6.3).

(b) 2 classic-Robocode contrast fixtures. contrast_stationary_sittingduck and contrast_straightline — trivial motion, used as a sanity/ceiling check. [MEASURED].

(c) 5 classic-Robocode DrussGT captures. Real, unmodified DrussGT 3.1.4159 movement captured from Robocode 1.9.5.5 via tools/robocode_fixture_capture/ (file DRUSSGT_FIXTURES.md). Opponents: SpinBot, RamFire, Crazy, Corners, and a DrussGT mirror; 28,797 ticks. The coordinate conversion is validated to 0.000–0.001° on every fixture by recomputing the direction implied by (heading, speed) against the recorded per-tick displacement. These are OPEN-LOOP and perfect-information: the replayed DrussGT never dodges our bullets, and the observer reports true positions every tick (unlike the live bot's stale between-scan WorldState). They are therefore optimistic and good for relative gun ranking, not absolute hit rates. [MEASURED] (DRUSSGT_FIXTURES.md).

(d) 5 closed-loop Tank Royale DrussGT captures. The real DrussGT jar playing Tank Royale through tools/robocode_shim/, captured by tools/robocode_shim/src/robocode_shim/TrBattleCapture.java (file TR_BRIDGE_FIXTURES.md). Primary file tr_drussgt_vs_modularbot.jsonl: 15 rounds, 20,026 ticks, ModularBot fired 1,134 shots. At capture time DrussGT was reacting to our real bullets. Closed-loop proven, not asserted: tools/robocode_shim/analyze_closed_loop.py event-locks |Δheading| to ModularBot's fire times (a heat-limited near-metronome, median interval 14 ticks) and gets an oscillating response with the fire period; the cross-correlation peaks at r = +0.111, lag 12, permutation p = 0.005 (null peak mean +0.016), and the own-fire control is flat, so the lock is enemy-driven, not internal cadence. Still perfect-information, and open-loop at replay time — "closed_loop" describes the capture, not a later replay. TR angle conversion residual is ~1.5° mean (vs 0.000° classic) because the TR server moves along the pre-turn heading. [MEASURED] (TR_BRIDGE_FIXTURES.md).

Fixture Source rounds ticks adversary / note
circular … energy-threshold-turner synthetic — 150–260 8 known-rule trajectories
contrast_stationary_sittingduck, contrast_straightline classic-robocode 1 each 1,451 sanity contrasts
drussgt_vs_spinbot classic-robocode 20 5,002 open-loop, perfect-info
drussgt_vs_ramfire classic-robocode 20 3,240 open-loop, perfect-info
drussgt_vs_crazy classic-robocode 20 9,025 open-loop, perfect-info
drussgt_vs_corners classic-robocode 20 4,975 open-loop, perfect-info
drussgt_vs_drussgt classic-robocode 2 6,555 open-loop, perfect-info (mirror)
tr_drussgt_vs_modularbot tr-bridge 15 20,026 closed-loop at capture, perfect-info
tr_drussgt_vs_modularbot_shield tr-bridge 10 12,629 shield on
tr_drussgt_vs_spinbot / _crazy / _corners tr-bridge 10 each 10,824 / 11,507 / 2,575 closed-loop at capture

1.3 The live boss — the real DrussGT jar

tools/robocode_shim/ runs the unmodified DrussGT.jar (159,289 bytes, md5 5cd6015dcc6d6da8a7e6aeecb1fec211) as a Tank Royale bot. The classic robocode.* API is a thin delegation layer over the public IBasicRobotPeer/IAdvancedRobotPeer seam, so the shim reuses the genuine robocode.jar and implements only the 75-method peer interface (ClassicPeer), plus BotHost/ThreadManagerFix. DrussGT compiles with zero shim API symbols and runs real battles; the EnergyDome shield is disabled by default (pure wave surfer). Known physics divergences (move/turn ordering, distance bookkeeping, etc.) are enumerated in section 5.9 of tools/robocode_shim/README.md. [MEASURED].

The boss is far stronger than us: in the capture battle DrussGT beat ModularBot 1447–300 over 15 rounds (ModularBot won round 5 only), firing 1,400 bullets at a 12.1% hit rate against ModularBot's 1,134 bullets at 5.3%. That is the number the whole gun rack is trying to move. [MEASURED] (commit 17c99f5, TR_BRIDGE_FIXTURES.md).

1.4 A/B methodology — server-side per-run hit rate, never scores

The A/B that decides configs uses server-side ground truth, not the bot's own counters and not scores. [MEASURED] (/tmp/analyze.py, /tmp/compare.py, /tmp/ab_logs3/FINAL_AB.txt):

  1. The battle runner writes a per-shot events sidecar (/tmp/events_*.json) with fire / hit / damage events stamped with the server's per-round bullet id (GunEngine.nextBulletId). 26b66cb proved per-gun attribution: the server assigns the id once and reuses it on BulletFired, BulletHitBot, BulletHitWall, BulletHitBullet; hits that arrive before the fire event (client priority 70 > 60) are deferred. 99.9% of shots and 99.8% of hits were attributed in that session.
  2. A config is judged on its per-run real hit rate (hits/shots for one battle), and two configs are called different only if their per-run ranges do not overlap. This is why the overnight A/B reports "SEPARATED" or "OVERLAP" for every pair rather than a single pooled p-value.
  3. Scores are not used to judge configs. Single-run scores swing by a couple of hundred points: the 13 shipped-config runs span 175–526 (s.d. ≈ 105, range 351; /tmp/battle_base_r*.log), so a ~210-point 2-s.d. band swamps any plausible config effect. The gate study (3c90a59) made the same call: no per-adversary score delta exceeded the ~300-point run-to-run noise band.

The bot-side per-gun attribution (realShots/realHits in /tmp/gun_stats_base_r*.jsonl) covers ~87% of the server's total shots uniformly (13 runs: 220/3,157 = 6.97% attributed vs 251/3,612 = 6.95% server-side), so it is used for per-gun ranking but the server sidecar is the ground truth for config decisions. [MEASURED] (re-aggregated from /tmp/events_base_r*.json and /tmp/gun_stats_base_r*.jsonl).


2. The metric lesson: virtual hit rate is NOT a proxy for real hit rate

This is the most important conceptual result of the night and it invalidates a naive reading of every offline table in this report.

[MEASURED] Aggregating the 13 shipped-config runs against the live DrussGT boss (/tmp/gun_stats_base_r{1..13}.jsonl; python3 /tmp/agg2.py base), the Spearman rank correlation between a gun's virtual hit rate and its real hit rate is

Spearman(virtual rank, real rank) = -0.374   (n = 13 guns, all with ≥10 real shots)

It is not weak — it is inverted. The guns with the highest virtual rates have among the lowest real rates, and vice versa:

Gun Selected (ticks) Real hits/shots Real % Virtual %
Linear 2,707 9/84 10.7 10.2
Circular 4,051 17/172 9.9 11.9
KNN 4,886 11/122 9.0 7.5
Pattern 14,786 50/582 8.6 12.0
Accel 9,205 28/382 7.3 12.1
AvgLead 5,119 14/200 7.0 12.3
GuessFactor 2,345 5/72 6.9 10.3
DecayGF 1,360 3/47 6.4 9.2
WallBounce 7,618 18/288 6.2 12.9
StopShot 3,348 8/132 6.1 12.6
Tsetlin 3,284 6/103 5.8 12.9
Displace 2,664 4/75 5.3 12.3
HeadOn 16,975 47/898 5.2 8.6

Read the top and bottom: Tsetlin, WallBounce and StopShot have the highest virtual rates (12.6–12.9%) and near-bottom real rates (5.8–6.2%); Linear and KNN sit at 10.2% / 7.5% virtual but 10.7% / 9.0% real. The ranking the virtual metric produces is not merely uninformative, it points the wrong way.

Why this matters for selection. What has kept the rack alive is the selector's floor/tie hedging, not its ranking: removing the floor (GUN_SELECTOR_FLOOR=0.0, config floor00) drops the rack from 6.95% to 5.08% at 175 dmg/run (vs 251) over 788 shots. [MEASURED] (/tmp/compare.py). So the selector is useful because it refuses to commit to a bad field, not because its virtual-rate ordering is good.

The virtual metric appears anti-correlated no matter which config you pick. Measured Spearman per config (5-run 12-round A/B, /tmp/ab_logs3/FINAL_AB.txt): absolute+point −0.371, absolute+path −0.073, relative+point −0.037, relative+path +0.522. But on the large 13-run base set the shipped config (relative+path) is −0.374. The sign flips between run sets, which is itself the finding: the correlation is unstable, so no ranking rule built on it can be trusted. [MEASURED] + [INFERRED] (the flip is measured; the conclusion is reasoning).

2.1 The metric A/B: point vs path (this one is real, and it is selection)

GUN_VBULLET_METRIC picks how a virtual bullet is scored (common_libs/gun_harness/virtual_bullets.nim):

  • bmPoint — resolve at the fire-time aim distance and score that single point. Measures prediction accuracy.
  • bmPath (shipped) — fly the ray to the wall and test each swept segment against the target radius. Measures hypothetical hit chance.

[MEASURED] Live A/B against the boss, 5 battles × 12 rounds, one frozen binary (commit 3b5d70b; per-run detail in /tmp/ab_logs/FINAL_ANALYSIS.txt):

Metric Shots Hits Real hit rate Per-run rates Spearman
point 4,660 219 4.70% 5.53 / 5.30 / 4.92 / 3.16 / 4.57 −0.04
path 4,834 359 7.43% 6.76 / 8.20 / 8.24 / 6.55 / 7.30 +0.52

The distributions do not overlap: path's worst run (6.55%) beats point's best (5.53%). +2.73 pp, +58% relative, z = 5.56, p < 0.0001. Range distributions were identical (~460–478 px), so this is not a range confound.

The gain is selection, not better gun learning. Under point every gun's virtual rate is compressed into 0.6–4.4%, so HeadOn sits inside the 2 pp tie margin and takes 72.6% of selection ticks / 76.9% of shots while ranking 11th of 13 by real hit rate (2.3%). Under path the band widens to 4.7–13.7% and HeadOn's shot share falls to 35.9%, so Pattern/Accel/WallBounce get picked. The counterfactual confirms it: applying the point model's per-gun real rates to the path model's shot mix yields 7.65%, i.e. essentially the whole observed gain. [MEASURED] (commit 3b5d70b).

Offline range total moves the same way: 34.3% under point (/tmp/final_range.txt, 35,636/104,000) vs 50.8% under path (/tmp/range_path.txt, 52,770/103,938). The offline totals are inflated by the perfect-information synthetic fixtures (three of them score 100% under path for every gun), so the offline totals are not comparable to live rates — only to each other.

2.2 The selector-threshold A/B: absolute vs relative

The legacy thresholds were calibrated for a rate scale that does not exist. [MEASURED] offline replay of a fogged live WorldState vs DrussGT (1,397 selection ticks, /tmp/ab_logs3/selector_diag.txt):

Config Floor fires HeadOn selection share bestRate med
absolute + point 53.0% 69.1% 8.0%
relative + point 21.2% 43.5% 5.25%
absolute + path 3.0% 23.1% 24.0%
relative + path (shipped) 8.4% 24.2% 16.75%

The 0.10 absolute floor fires on 53.0% of point-metric ticks and forces HeadOn, whose real rate was 2.0–4.4%. (An earlier claim that the floor fires always is refuted: it is 53%, because bestRate is a max over gun×power-bin and an occasional ≥50-sample bin clears 10%.)

The scale-aware replacement (commit dea4dcb): RelTieMargin = 0.20 (tie band is a fraction of bestRate), FloorPeakFrac = 0.25 (floor fires only if the field collapsed vs its own recent peak over a 256-tick window, counting only guns with ≥ MinObsBeforeCompete = 50 samples), and pooled-over-bins ranking instead of max-over-bins.

Live A/B, 3 runs × 10 rounds (commit dea4dcb, /tmp/ab_logs3/FINAL_AB.txt):

Config Per-run rates Pooled vs absolute+point
absolute + point 3.66 / 2.45 / 5.01 3.76% —
absolute + path 7.55 / 8.21 / 6.83 7.57% SEPARATED (p<0.0001)
relative + point 7.66 / 6.18 / 5.79 6.59% SEPARATED
relative + path (shipped) 7.15 / 7.55 / 6.90 7.21% SEPARATED

absolute+path is nominally 0.35 pp above relative+path, but they overlap (p = 0.64); so do relative+point and both path configs. The metric is the dominant lever; under path the two threshold models are statistically tied. relative was shipped because it is the principled scale-aware fix, works under both metrics, and prevents the point-metric catastrophe if anyone switches back. [MEASURED].


3. Per-gun performance on real numbers

The final per-gun table, shipped config, 13 runs vs the live DrussGT boss, 3,612 server-side shots, 6.95% overall (server sidecar; per-run rates 6.77 / 5.90 / 6.57 / 6.34 / 6.10 / 7.59 / 2.90 / 7.95 / 7.49 / 9.18 / 8.44 / 7.48 / 6.42%; 251 dmg/run). Per-gun rows are the bot-side attribution over the same runs. [MEASURED] (commit 2c94dc2; /tmp/gun_stats_base_r*.jsonl, re-aggregated with /tmp/agg2.py base; /tmp/events_base_r*.json).

Verdict Gun Real hits/shots Real % Virtual % Selected
KEEP Linear 9/84 10.7 10.2 2,707
KEEP Circular 17/172 9.9 11.9 4,051
KEEP KNN 11/122 9.0 7.5 4,886
KEEP Pattern 50/582 8.6 12.0 14,786
KEEP Accel 28/382 7.3 12.1 9,205
KEEP AvgLead 14/200 7.0 12.3 5,119
MARGINAL GuessFactor 5/72 6.9 10.3 2,345
MARGINAL DecayGF 3/47 6.4 9.2 1,360
MARGINAL WallBounce 18/288 6.2 12.9 7,618
MARGINAL StopShot 8/132 6.1 12.6 3,348
BELOW Tsetlin 6/103 5.8 12.9 3,284
BELOW Displace 4/75 5.3 12.3 2,664
FLOOR — STAYS HeadOn 47/898 5.2 8.6 16,975

Verdicts.

  • KEEP: Linear, Circular, KNN, Pattern, Accel, AvgLead. These six are at or above the 6.95% overall, yet their virtual rates are mid-pack to low: the metric's three favourites (Tsetlin 12.9%, WallBounce 12.9%, StopShot 12.6%) are near the bottom of the real ranking, while the real leader (Linear) sits at 10.2% virtual. Further evidence the virtual ranking is inverted.
  • MARGINAL: GuessFactor, DecayGF, WallBounce, StopShot. Within ~1 pp of overall on small N (47–288 shots). They are not obviously worth deleting, but they have not earned a larger share.
  • BELOW OVERALL: Tsetlin, Displace. Below 6% on 75–103 shots. Candidates to drop or re-tune, but the N is small.
  • HeadOn MUST STAY despite being lowest (5.2%). It is the floor fallback: when the field collapses the selector returns gun 0. Disabling the floor measurably hurt — 5.08% / 175 dmg vs 6.95% / 251 dmg (config floor00, 788 shots). Do not delete HeadOn to improve the per-gun average; that average is computed over shots it only gets because nothing better was available. [MEASURED] (/tmp/compare.py).

The 16-candidate ranking A/B found no winner. Runtime knobs were added to the selector (GUN_SELECTOR_WINDOW, MINOBS, TIE, FLOOR, POOL, RANK, SHRINK, SEED; rankScore supports mean/Wilson/UCB/Thompson/shrinkage), all defaulting to the shipped values. 16 candidates were A/B'd against the boss. None credibly beat the shipped config; every candidate's per-run interval overlaps base, and the nominal "winners" are ≤0.6 SE apart on far fewer shots. [MEASURED] (commit 2c94dc2; /tmp/compare.py):

Config Runs Shots Rate % dmg/run Spearman
base (shipped) 14* 3,612 6.95 251 −0.374
tie00 (TIE=0.0) 2 540 7.04 242 +0.018
win50 (WINDOW=50) 2 559 6.08 216 +0.588
wilson (RANK=wilson) 13 3,045 6.67 196 +0.088
thompson (RANK=thompson) 2 531 4.90 166 +0.083
maxbin (POOL=0) 2 574 5.23 198 −0.264
minobs20 (MINOBS=20) 2 535 5.98 212 +0.144
tie05 (TIE=0.05) 12 3,357 6.20 222 −0.060
tie10 (TIE=0.10) 2 556 6.65 233 −0.150
tie40 (TIE=0.40) 2 583 6.35 216 −0.160
floor10 (FLOOR=0.10) 6 1,694 6.49 232 −0.578
t05f10 (TIE=0.05, FLOOR=0.10) 4 1,073 5.50 190 −0.041
wilf10 (RANK=wilson, FLOOR=0.10) 4 1,252 6.71 271 −0.410
floor00 (FLOOR=0.0) 3 788 5.08 175 +0.055
f00t05 (FLOOR=0, TIE=0.05) 3 540 6.48 153 +0.226
f00wil (FLOOR=0, RANK=wilson) 2 595 6.72 221 −0.116
f00w50 (FLOOR=0, WINDOW=50) 2 593 6.58 210 +0.178

* compare.py counts 14 events files, but run 14 has no fire events; the 13 runs with data carry all 3,612 shots.

No ranking rule fixed the anti-correlation. The best Spearman in the table (win50, +0.588) is on 2 runs / 559 shots. The shipped config's −0.374 over 13 runs is the most reliable estimate. [MEASURED].

3.1 Offline range: which gun wins which trajectory family

Shipped path metric, /tmp/range_path.txt (52,770/103,938 = 50.8%). Cells are hit-% per gun per fixture; bold = best gun for that fixture. 400 shots per gun per fixture.

Fixture HeadOn Linear Tsetlin Circular GuessF Pattern WallBn Accel StopSh Displ AvgLead DecayG KNN
circular 12 29 29 100 15 66 34 100 27 18 54 13 73
constant-velocity 100 100 100 100 100 100 100 100 100 100 100 100 100
contr-SittingDuck 100 100 100 100 100 100 100 100 100 100 100 100 100
contr-StraightLine 43 100 72 100 100 100 94 98 84 98 100 100 77
decel-before-turn 79 100 90 100 100 100 100 100 98 100 100 100 78
classC-corners 6 9 23 17 9 17 12 19 19 22 13 10 4
classC-crazy 4 36 24 34 35 32 54 35 22 26 39 27 22
classC-mirror 12 6 7 6 6 9 6 7 8 5 8 7 8
classC-ramfire 35 50 32 54 50 54 49 49 35 30 56 52 46
classC-spinbot 3 58 10 50 58 37 48 50 13 46 52 58 42
energy-threshold-turner 58 34 39 100 34 88 38 100 40 44 63 34 42
oscillator 100 100 100 100 100 100 100 100 100 100 100 100 100
random-walk 26 67 64 28 63 41 68 26 61 67 55 62 47
stationary 100 100 100 100 100 100 100 100 100 100 100 100 100
TR-corners 3 30 28 33 39 26 33 34 30 34 56 29 32
TR-crazy 5 7 12 16 2 14 19 28 11 20 27 2 3
TR-ModularBot 15 12 18 13 12 12 14 15 21 12 13 11 7
TR-MB-shield 4 10 32 27 10 28 32 28 24 37 21 27 5
TR-spinbot 3 21 20 21 37 37 25 30 27 14 24 15 11
wall-bounce 0 49 44 48 37 41 100 60 41 34 52 34 28

Aggregated hit-% by family (400 shots/gun/fixture):

Family HeadOn Linear Tsetlin Circular GuessF Pattern WallBn Accel StopSh Displ AvgLead DecayG KNN
synthetic (8) 59 72 71 84 69 79 80 86 71 70 78 68 71
classic DrussGT (5) 12 32 19 32 32 30 34 32 20 26 34 31 25
TR bridge (5) 6 16 22 22 15 23 24 27 23 23 28 17 12
all 20 35 51 47 57 49 55 56 59 48 50 57 49 46

Which gun wins which trajectory (offline, path metric):

  • Stationary / constant-velocity / oscillator / decel-before-turn / straight-line: every useful gun ≥ 94% (perfect-information + path metric makes these uninformative). Only HeadOn (43–79%) and KNN (77–78%) stand out as weak.
  • Circular: Circular and Accel 100% by construction; KNN 73, Pattern 66.
  • Wall-bounce: WallBounce 100, Accel 60, AvgLead 52 — the only fixture where WallBounce is dominant.
  • Random-walk: WallBounce 68, Displace 67, Linear 67, Tsetlin 64; Accel collapses to 26.
  • Energy-threshold rule: Circular and Accel 100, Pattern 88, AvgLead 63. Tsetlin 39 — it beat Linear (34) but is far from reading the rule.
  • Classic DrussGT (a real surfer): WallBounce 34 and AvgLead 34 at the top, HeadOn 12 at the bottom. Cornered surfing is the one case where Tsetlin (23) leads.
  • TR bridge DrussGT: AvgLead 28 overall; Accel 28 on TR-crazy, AvgLead 56 on TR-corners, StopShot 21 on the ModularBot mirror, Pattern 37 on TR-spinbot.

Do not read the offline winner as the rack verdict. The offline range's classic DrussGT order (WallBounce 34 top, KNN 25 low) is close to the inverse of the real order (KNN 9.0% third, WallBounce 6.2% ninth). The offline range is excellent for catching structural bugs (§6) and for per-trajectory sanity, and poor as a selector signal. [MEASURED] + [INFERRED].


4. Verdict table

Gun Family / model Real % (13 runs) Offline all-20 % Verdict
Linear constant-velocity lead 10.7 51 KEEP — this is what a surfer cannot defeat; keep warm
Circular constant-turn lead 9.9 57 KEEP
KNN k-NN on motion history 9.0 46 KEEP — offline under-rates it badly
Pattern pattern replay 8.6 55 KEEP
Accel acceleration-aware lead 7.3 59 KEEP
AvgLead windowed average lead 7.0 57 KEEP
GuessFactor GF histogram 6.9 49 MARGINAL — small N, no clear edge
DecayGF recency-weighted GF 6.4 49 MARGINAL
WallBounce wall-reflection model 6.2 56 MARGINAL — offline favourite, real underperformer
StopShot deceleration/stop point 6.1 48 MARGINAL
Tsetlin Tsetlin-Machine correction 5.8 47 BELOW — learns, not yet competitive
Displace displacement vector 5.3 50 BELOW
HeadOn aim at current position 5.2 35 KEEP — mandatory floor fallback
TMSelect TM mixture-of-experts gate — — DISABLED (EnableTmSelector = false) — see §6.4

5. What "worth keeping" means, and what is not proven

The KEEP/MARGINAL/BELOW split is a statement about a single adversary (a wave surfer), judged on the shipped config, on a few hundred real shots per gun. It is a starting point, not a final ranking. The concrete caveats are in §7.


6. Bugs found and fixed tonight (why earlier rack verdicts were wrong)

Five of these changed the rack ordering; all are [MEASURED] from the commit messages and the offline range.

6.1 The GF family aimed at the fire-time RADIUS, not the angle

guess_factor, decay_gf and knn_gun aimed at the fire-time distance. But the virtual-bullet metric resolves a bullet at its aim-point distance and scores that single point against the enemy's position on that tick, so with any radial target motion the bullet stopped at the wrong radius and missed even with a perfect angle. Angle-only prediction is structurally unscoreable under the point metric. Two competing hypotheses were tested and both refuted: (a) MEA range too narrow — 0 clamped shots out of 837/849/957, with required offsets peaking at ~33° against MEA 28.1–46.7°, and arcsin(8/bulletSpeed) correctly uses max robot speed, not the hit radius; (b) wrong GF peak — a sweep of every constant GF value showed the oracle-best constant offset was only 6% on circular, 4% on wall-bounce, 7.5% on random-walk. Learning was fine too (~850–960 observations per fixture, 0 starved waves). [MEASURED] (commit 7f706e5).

Fix: a self-consistent constant-velocity lead_forecast.nim base, so the histogram learns the residual and the aim point lands at the right radius; also fixed linear.nim (it did a one-shot extrapolation and never iterated its flight time). Before/after, offline: circular GF 6→23, DecayGF 6→21; wall-bounce GF 0→60.2, DecayGF 0→60.2; constant-velocity GF/DecayGF/KNN 26→100; random-walk GF 0→53; StraightLine GF 8→77. The oracle-best constant GF moved 6%→20% (circular), 4%→57% (wall-bounce), 7.5%→49% (random-walk), proving the structural fix independently of tuning. Honest trade-off: on the 5 real DrussGT surfer captures the GF family regressed (GuessFactor 108→55, DecayGF 108→76, KNN 101→74 hits/2000) because a linear base is a poor model for a surfer and the residual histogram is noisier than the old total-lead histogram.

That regression was then recovered by blending the range between a radial-only forecast and the geometric one by the measured radial fraction (radialFrac), keeping the constant-velocity bearing. Nine candidate bases were measured and rejected with numbers (velocity scaling 0.8 recovered DrussGT but destroyed wall-bounce 241→20; radial-only range wall-bounce 241→140; short-window average worse than both; reversal/speed gates weaker than the blend). Result (hits/2000): classic-5 GF 55→171, DecayGF 76→100; TR-5 GF 9→86, DecayGF 4→87; synthetic-10 GF 2702→2717. The only figure below the old base is classic-5 DecayGF (108→100, within noise). [MEASURED] (commit e2ca2fc).

6.2 The wave queues were starved 1-push-vs-4-pops

predict() stored one wave per tick while onResult() popped one per resolved bullet (~4/tick), so the queue drained within a few dozen ticks and ~3 of every 4 resolutions returned without learning; the survivor paired with a same-tick wave (bearingDelta ≈ 0), pinning the histogram at centre. Proof: GF.vHits == HeadOn.vHits and DecayGF.vHits == HeadOn.vHits byte-for-byte in every one of 50 rounds — GF, DecayGF and KNN had degenerated into HeadOn clones. Fix: per-bin FIFO with an O(1) head cursor, at most one push per (tick, bin). Also maxBullets 2,048→8,192: the rack spawns 52 bullets/tick so the ring wrapped every ~39 ticks while a long power-3 shot needs ~90, silently discarding unresolved bullets and biasing every measured hit rate by range; a droppedBullets counter was added. After the fix vDropped = 0 and vStarved = 0 across all 48 recorded rounds. [MEASURED] (commit 0cc6821).

6.3 Tsetlin's clauses saturated at ~714 included literals each

Tsetlin.vHits was byte-for-byte equal to Linear.vHits in every measured round of every run because its learned correction was always exactly 0. Root cause: tmLearnOne rewarded included true literals unconditionally, omitting Granmo's (c=0, lk=1) → toward Exclude counter-force, so true literals ratcheted toward Include forever; Type II was unreachable dead code with the wrong direction; resource allocation was an |error| heuristic instead of Granmo's (T − clip(v,−T,T))/(2T); the label baseline had a factor-2 shrink (error = δ − 2c, fixed point c = δ/2); hits zeroed their residual; the enemy-energy feature was duplicated (state.selfEnergy fed where WorldState.enemyEnergy exists, so energy rules were literally unrepresentable); and tmEvalClause needed Granmo Eq. 6 (all-Exclude clause outputs 1 during learning, 0 during classification) or fix #1 deadlocks every clause at empty.

Measured effect (energy-threshold-turner, seed 1): mean included literals/clause 714.0 → 13.8; active clauses 100/100 → 53/100; nonzero corrections 8/764 → 708/764; Tsetlin virtual hits 27/400 → 69/400 (Linear 43/400). Tsetlin now learns but is not yet competitive with Linear — the regression head is untuned, flagged as follow-up rather than claimed as a win. [MEASURED] (commit 8937000; /tmp/ab_logs3/final_test_tsetlin_gun.log).

6.4 The TM classifier gun did not earn its slot (but its clauses are real)

A Tsetlin-Machine mixture-of-experts gate over HeadOn/Linear/Circular/WallBounce/Accel was built with the corrected feedback and labelled by which expert's prediction was closest to the actual enemy position (an exact, supervised, per-shot label — no delayed credit). It loses to the best of its own experts offline on nearly every fixture, and against DrussGT it cost real performance:

baseline (path + relative)         7.56% real hit rate, 157 dmg
+ power fix                        7.47%,               239 dmg
+ power fix + TM selector          5.59%,               133 dmg

It was selected on 806 ticks and fired 24 real shots at 4.2%. It ships disabled (EnableTmSelector = false; code and wiring kept intact). However, the gate latched onto meaningful structure: on the energy-threshold turner, HeadOn's clauses key on the energy bits (the rule's own driving variable) while Circular keys on distance/velocity. So the TM learned something real and interpretable; it simply could not beat "always pick the best expert". [INFERRED] root cause: the closest-expert label is noisy because several experts are near-tied, and under the path metric the winner varies by power bin while the gate sees one shared per-tick input, so a one-vs-rest gate over a saturated 870-bit clause space has no margin to exploit. A standalone Granmo classifier on the same encoding reaches ~99% on the rule but, per the counterfactual probe, does not read energy (follow rate 24% high / 62% mean — statistically identical at 1, 2 and 10 frames), so even the "it learned the rule" claim is limited to ~99% accuracy, not to a readable energy threshold. The best recovered proposition was !g9 ∧ !g8 (energy < 25.6, not the labelled 30) — a genuine simple threshold, but not the ensemble's decision mechanism. [MEASURED] (commits 57b2ac3, d5061ee; test_tm_pattern_learning.nim).

6.5 The selector thresholds were absolute on a rescaled metric

Covered in §2.2: the 0.10 absolute floor fired on 53.0% of point-metric ticks and forced HeadOn (real 2.0–4.4%, 11th of 13); HeadOn selection share fell 69.1% → 43.5% under relative thresholds (and 23.1% → 24.2% under path). Also: bestGun was first-index-wins argmax, so HeadOn at index 0 silently won every tie until the random tie-break landed (343e631); bestPower had the same absolute-40% defect (below).

6.6 Power selection was stuck at power 1.0 (MinHitRate = 0.40)

bestPower used an absolute MinHitRate = 0.40 bar. Measured per-bin virtual rates show no bin ever clears 40%, so 11 of 14 guns were stuck at bin 0 (power 1.0) even where higher bins were comparable:

Linear  p1.0 44%  p1.5 39%  p2.0 30%  p3.0 29%   old bin 0 -> new bin 3
Accel   p1.0 44%  p1.5 40%  p2.0 26%  p3.0 29%   old bin 1 -> new bin 3
Pattern p1.0 50%  p1.5 40%  p2.0 27%  p3.0 12%   old bin 1 -> new bin 2

Replaced with a scale-aware PowerBarFrac = 0.50 (a dimensionless fraction of the gun's own best-bin rate); 13 of 14 selections now pick heavier bullets. Real effect vs DrussGT (8 rounds × 3 runs): hit rate unchanged (7.56% → 7.47%), damage +52% (157 → 239 per run) and rounds end faster. [MEASURED] (commit 57b2ac3).

Smaller fixes in the same family: bestPower on a cold gun returned the highest bin (empty bin satisfied the count == 0 clause); fitnessFor aggregated enemies in nondeterministic hash order; stop_shot had an unreachable deceleration branch and several guns had tick-only caches that made all four power bins return bin 0's lead (e536900). [MEASURED].


7. Known caveats and open problems

Stated without hedging.

  1. The headline per-gun numbers come from ONE adversary, a wave surfer. HeadOn is genuinely bad against surfers, so part of the rack ordering may be matchup-specific. A SpinBot guard was inconclusive: ModularBot fires only 17–31 real shots/run against a fast bot because the range-aware firing gate is strict at long range, so the guard had little power (Wilson looked better, 18.5% vs 8.6%, but on 70–92 shots with a 5–33% spread). [MEASURED] (commit 2c94dc2). A second, independent adversary at scale is missing.

  2. Per-gun real N is small. 47–898 shots per gun; n < 200 gives roughly ±5 pp across a 3–15% spread. Single-gun ordering is indicative, not definitive. The KEEP/MARGINAL/BELOW boundaries should be treated as soft.

  3. The fixtures are perfect-information and therefore optimistic. Every fixture is an observer capture with true positions every tick; the classic set is additionally open-loop (replayed DrussGT never dodges our bullets). Absolute offline hit rates are inflated by an unknown amount; only relative comparisons are safe.

  4. The virtual metric is anti-correlated with real hit rate and no tested ranking rule fixed it. Shipped config Spearman ≈ −0.374 over 13 runs (the sign flips to +0.52 on the smaller 5-run set, so it is unstable). 16 candidate ranking rules all overlapped the shipped config. The selector's value lives in its floor/tie hedging (5.08% without the floor vs 6.95% with it), not in its ranking. [MEASURED].

  5. The TM classifier gun did not earn its slot. It cost real performance (7.47% → 5.59%, 133 dmg) despite showing interpretable energy structure in its clauses (§6.4). It is disabled; re-enabling requires a fix to the gate margin/label problem, not more training.

  6. Real-hit-rate-driven selection is not viable yet. Only the selected gun fires, so unselected guns get near-zero real shots (GuessFactor 20, Linear 24 vs HeadOn 733 in the point A/B); noise is fatal (n = 470 at p = 10% gives ±2.8 pp, most guns n < 200 gives ±5 pp+); and real rate is conditional on when the gun was selected. A blended signal with forced exploration and shrinkage is defensible in principle but needs thousands of shots per gun across many battles. Real rate is currently best used offline as the evaluation metric — which is exactly what the A/B does. [MEASURED] (commit dea4dcb).

  7. The offline==online acceptance test is flaky (§1.1): typically 11/12 on unmodified HEAD, with the mismatching gun varying run to run. The equivalence claim is strong-but-not-exact until the boundary race is fixed.

  8. The selector's random tie-break is not randomised in the live bot. The shipped bot never calls randomize(), so the "random" sequence is fixed across process restarts (a side finding of 2c94dc2, not fixed).

  9. The firing gate is not the bottleneck. The shipped range-aware gate does not beat a fixed 2.0° gate on hit rate (55.8% vs 57.9%, ~1.5 σ), though it fires 22–28% more shots. No per-adversary score delta exceeded the ~300-point run-to-run noise band. [MEASURED] (commit 3c90a59).

  10. The boss is ~2.3× more accurate than the whole rack (12.1% vs 5.3% in the capture). Closing that gap is the point of the rack; the current best single gun is 10.7%.


8. Reproduction

Commands recorded in the commits and tool READMEs. (I was instructed not to run builds/tests while writing this report; these are the documented invocations, not a fresh verification by me.)

# Offline gun range over all 20 fixtures, shipped path metric (default):
nim c -r common_libs/tests/run_range.nim
# Point metric for comparison:
GUN_VBULLET_METRIC=point nim c -r common_libs/tests/run_range.nim
# Add timing:
nim c -r common_libs/tests/run_range.nim --timing

# Selector diagnostics (floor/tie/bestRate/HeadOn-share) on a fixture:
GUN_SELECTOR_MODE=relative nim c -d:release -r \
    common_libs/tests/analyze_selector.nim tools/fixtures/drussgt_vs_spinbot.jsonl

# Offline == online acceptance (currently flaky):
nim c -r common_libs/tests/acceptance_offline_vs_online.nim

# Tsetlin gun clause sparsity / divergence:
nim c -r common_libs/tests/test_tsetlin_gun.nim
# TM readability (standalone Granmo classifier on the energy-threshold rule):
nim c -r common_libs/tests/test_tm_pattern_learning.nim

# Live boss (real DrussGT jar; jars stay out of git, see the README):
#   tools/robocode_shim/run_bridge_battle.sh <bot_dir> <rounds> <capture.jsonl>
# Closed-loop evidence for the TR captures:
python3 tools/robocode_shim/analyze_closed_loop.py \
  tools/fixtures/tr_drussgt_vs_modularbot.jsonl \
  tools/robocode_shim/evidence/tr_drussgt_vs_modularbot.events.json

Per-gun aggregation scripts used for the tables above: python3 /tmp/agg2.py base (virtual-vs-real + Spearman), python3 /tmp/compare.py (server-side per-run A/B + overlap). Raw range output: /tmp/range_path.txt (path), /tmp/final_range.txt (point).