docs: definitive gun-rack report on real measured numbers

Replaces the stale 2026-09-20 docs, which predated per-gun real attribution,
the offline gun range and the DrussGT boss, and whose verdicts were built on
virtual hit rates that turned out to be ANTI-correlated with reality.

docs/gun_rack_analysis.md (751 lines) covers: the test infrastructure
described honestly (offline range with its flaky-acceptance caveat, the 20
fixtures and what each set is good for, the live boss, and the A/B methodology
of per-run server-side real hit rate with an explicit overlap test); the
virtual-vs-real metric lesson with Spearman -0.374 and the point-vs-path A/B;
per-gun real performance and the 16-rule ranking A/B; the offline per-fixture
gun matrix; KEEP/MARGINAL/BELOW verdicts; and the five root-cause bugs with
before/after numbers.
docs/gun_rack_summary.md (58 lines) is the verdict table plus top actions.

The '~230 point' score-noise band that has been steering methodology all night
was re-derived from the artifacts rather than asserted: the 13 shipped-config
run scores span 175-526, s.d. ~105, i.e. a ~210-point 2-s.d. band.

Caveats recorded verbatim rather than softened: the offline==online acceptance
is flaky (typically 11/12 on unmodified HEAD), fixtures are perfect-information
and therefore optimistic vs live play, per-gun real N is small so single-gun
ordering is indicative, the headline numbers come from ONE wave-surfer
adversary, and HeadOn must stay despite being lowest because it is the floor
fallback (disabling it: 5.08% / 175 dmg vs 6.95% / 251 dmg).
This commit is contained in:
2026-09-21 06:37:20 +02:00
parent 2c94dc221a
commit 19410164f1
2 changed files with 763 additions and 241 deletions
+713 -209
View File
@@ -1,247 +1,751 @@
# Gun Rack Analysis Report
# Gun Rack Analysis — ModularBot guns vs the real DrussGT boss
**Date:** 2026-09-20
**Bot:** ModularBot (14 guns)
**Data source:** `/tmp/gun_stats.jsonl`
**Date:** 2026-09-21
**Bot:** ModularBot (13 active guns + 1 disabled TM classifier gun)
**Shipped config:** virtual-bullet metric `path`, selector thresholds `relative`
(`GUN_VBULLET_METRIC` default `path`, `GUN_SELECTOR_MODE` default `relative`)
**Rewritten:** this file supersedes the 2026-09-20 version, whose numbers
predated per-gun real-hit attribution, the offline gun range, and the live
DrussGT boss. No number from the old version is retained unless it was
re-measured below.
Evidence tags used throughout:
- **[MEASURED]** — I read it from a recorded artifact (commit message, `/tmp`
output, a log, or a source file). The provenance is named every time.
- **[INFERRED]** — reasoning from measured facts; explicitly not measured.
A note on provenance: the measurements below were produced overnight in
commits `343e631` … `2c94dc2` on branch `research/lead-targeting`. Each commit
message records the numbers, the hypotheses it refuted, and its caveats. Raw
artifacts are in `/tmp` (`/tmp/gun_stats_base_r*.jsonl`,
`/tmp/events_base_r*.json`, `/tmp/range_path.txt`,
`/tmp/ab_logs/FINAL_ANALYSIS.txt`, `/tmp/ab_logs3/FINAL_AB.txt`,
`/tmp/ab_logs3/selector_diag.txt`). Where a commit message and a raw artifact
disagree slightly, both are stated.
---
## Test Methodology
## 1. The test infrastructure
- **Adversaries:** 5 bots, 10 rounds each (OscillatorBot: 9 rounds — one round lost to disconnect)
- **Format:** 1v1, score accumulated across rounds
- **Virtual bullet system:** 4 power bins (1.0, 1.5, 2.0, 3.0), 100-tick rolling window per gun per bin
- **Gun selection rule:** highest virtual hit rate, minimum 15 observations before a gun is eligible to compete
- **Stats written to:** `/tmp/gun_stats.jsonl` (one JSON line per round, prefixed by `session_start` markers)
The numbers are only as good as the rig that produced them, so the rig is
described first.
### 1.1 The offline gun range
`common_libs/gun_harness/offline_range.nim` replays a recorded
`seq[WorldState]` through the **same** `VirtualTracker`
(`common_libs/gun_harness/virtual_bullets.nim`) that the live bot drives. The
claim is not "an approximation of the live metric", it is "the same metric":
virtual-bullet fitness is already a pure function of (a stream of
`WorldState`, a list of guns), and the Java battle only supplies where the
states come from. Guns keep their own internal history, so a replayed stream in
order is a complete movement history. **[MEASURED]** (`offline_range.nim`,
header comment; commit `974528d`).
- **Speed and sample count.** 8 fixtures / 1,770 ticks / ~92 k virtual bullets /
13 guns replay in 2.9 s at ~32 k virtual bullets/s — roughly **70× faster and
100× more samples** than a live gauntlet. **[MEASURED]** (commit `974528d`).
- **Acceptance test.** `common_libs/tests/acceptance_offline_vs_online.nim`
records one live round, replays it offline, and compares per-gun virtual hit
counts. It passed **12/12 deterministic guns** (Tsetlin is separately labelled
stochastic because `tmLearnOne` calls `rand()`). Two real live-loop ordering
quirks had to be modelled to reach that: `run()` calls `go()` before the
aim/fire block, so `tickBullets` resolves against the *next* tick's scan; and
if the target dies during that `go()`, the final tick's spawn+resolution is
skipped. **[MEASURED]** (commit `974528d`; a passing run is preserved in
`/tmp/ab_logs3/test_acceptance.log`: 12/12, 534-tick round).
- **Honest caveat — the acceptance test is currently FLAKY.** On the
*unmodified HEAD* source it normally reaches only **11/12**, e.g. KNN 81
online vs 71 offline, and the mismatching gun moves between runs (KNN, then
WallBounce). It is a live/offline boundary race, pre-existing, and not caused
by the selector work (the replay never calls the selector). Treat
"offline == online" as **strong but not exact until the race is fixed**.
**[MEASURED]** (commit `2c94dc2`). The 12/12 runs above are real; they were
lucky runs.
### 1.2 The fixture sets and what each is good for
There are **20 JSONL fixtures** under `tools/fixtures/`. They fall into four
groups. **[MEASURED]** (`tools/fixtures/`, `DRUSSGT_FIXTURES.md`,
`TR_BRIDGE_FIXTURES.md`).
**(a) 8 synthetic fixtures — ground truth by construction.**
`common_libs/gun_harness/offline_range.nim` generates them; each has a known
rule:
| Fixture | Generator / rule | What it is good for |
|---|---|---|
| `stationary` | fixed enemy | ceiling: any correct gun scores ~100% |
| `constant-velocity` | straight line, no walls | does the gun lead a moving target at all |
| `circular` | constant turn 3°/tick | circular/accel models |
| `wall-bounce` | specular reflection off all 4 walls | wall-aware prediction |
| `oscillator` | east 30 ticks, west 30 ticks | phase/timing |
| `random-walk` | seeded ±15°/tick jitter | generality under noise |
| `decel-before-turn` | cruise → full stop → pivot 3×45° → accelerate | stop-shot detection |
| `energy-threshold-turner` | **KNOWN RULE**: straight while `energy ≥ 30`, hard 20°/tick turn while `energy < 30`, `energy = max(5, 50 − 0.5·t)` | can a learner find a readable high-level rule; threshold crosses at t=41 |
The energy-threshold turner is the falsifiable one: the label is literally a
predicate over the 11-bit Gray-coded energy field, so a learner that reads
energy can be shown to have read the *right* variable (see §6.3).
**(b) 2 classic-Robocode contrast fixtures.** `contrast_stationary_sittingduck`
and `contrast_straightline` — trivial motion, used as a sanity/ceiling check.
**[MEASURED]**.
**(c) 5 classic-Robocode DrussGT captures.** Real, unmodified DrussGT
3.1.4159 movement captured from Robocode 1.9.5.5 via
`tools/robocode_fixture_capture/` (file `DRUSSGT_FIXTURES.md`). Opponents:
SpinBot, RamFire, Crazy, Corners, and a DrussGT mirror; 28,797 ticks. The
coordinate conversion is validated to **0.000–0.001°** on every fixture by
recomputing the direction implied by (heading, speed) against the recorded
per-tick displacement. **These are OPEN-LOOP and perfect-information:** the
replayed DrussGT never dodges *our* bullets, and the observer reports true
positions every tick (unlike the live bot's stale between-scan `WorldState`).
They are therefore optimistic and good for *relative* gun ranking, not absolute
hit rates. **[MEASURED]** (`DRUSSGT_FIXTURES.md`).
**(d) 5 closed-loop Tank Royale DrussGT captures.** The real DrussGT jar playing
Tank Royale through `tools/robocode_shim/`, captured by
`tools/robocode_shim/src/robocode_shim/TrBattleCapture.java` (file
`TR_BRIDGE_FIXTURES.md`). Primary file `tr_drussgt_vs_modularbot.jsonl`:
15 rounds, 20,026 ticks, ModularBot fired 1,134 shots. At capture time DrussGT
was reacting to *our* real bullets. **Closed-loop proven, not asserted:**
`tools/robocode_shim/analyze_closed_loop.py` event-locks `|Δheading|` to
ModularBot's fire times (a heat-limited near-metronome, median interval
14 ticks) and gets an oscillating response with the fire period; the
cross-correlation peaks at **r = +0.111, lag 12, permutation p = 0.005** (null
peak mean +0.016), and the **own-fire control is flat**, so the lock is
enemy-driven, not internal cadence. **Still perfect-information**, and
open-loop *at replay time* — "closed_loop" describes the capture, not a later
replay. TR angle conversion residual is ~1.5° mean (vs 0.000° classic) because
the TR server moves along the pre-turn heading. **[MEASURED]**
(`TR_BRIDGE_FIXTURES.md`).
| Fixture | Source | rounds | ticks | adversary / note |
|---|---|---:|---:|---|
| `circular` … `energy-threshold-turner` | synthetic | — | 150–260 | 8 known-rule trajectories |
| `contrast_stationary_sittingduck`, `contrast_straightline` | classic-robocode | 1 each | 1,451 | sanity contrasts |
| `drussgt_vs_spinbot` | classic-robocode | 20 | 5,002 | open-loop, perfect-info |
| `drussgt_vs_ramfire` | classic-robocode | 20 | 3,240 | open-loop, perfect-info |
| `drussgt_vs_crazy` | classic-robocode | 20 | 9,025 | open-loop, perfect-info |
| `drussgt_vs_corners` | classic-robocode | 20 | 4,975 | open-loop, perfect-info |
| `drussgt_vs_drussgt` | classic-robocode | 2 | 6,555 | open-loop, perfect-info (mirror) |
| `tr_drussgt_vs_modularbot` | tr-bridge | 15 | 20,026 | **closed-loop at capture**, perfect-info |
| `tr_drussgt_vs_modularbot_shield` | tr-bridge | 10 | 12,629 | shield on |
| `tr_drussgt_vs_spinbot` / `_crazy` / `_corners` | tr-bridge | 10 each | 10,824 / 11,507 / 2,575 | closed-loop at capture |
### 1.3 The live boss — the real DrussGT jar
`tools/robocode_shim/` runs the **unmodified** `DrussGT.jar` (159,289 bytes,
md5 `5cd6015dcc6d6da8a7e6aeecb1fec211`) as a Tank Royale bot. The classic
`robocode.*` API is a thin delegation layer over the public
`IBasicRobotPeer`/`IAdvancedRobotPeer` seam, so the shim reuses the genuine
`robocode.jar` and implements only the 75-method peer interface
(`ClassicPeer`), plus `BotHost`/`ThreadManagerFix`. DrussGT compiles with zero
shim API symbols and runs real battles; the EnergyDome shield is disabled by
default (pure wave surfer). Known physics divergences (move/turn ordering,
distance bookkeeping, etc.) are enumerated in section 5.9 of
`tools/robocode_shim/README.md`. **[MEASURED]**.
The boss is far stronger than us: in the capture battle DrussGT beat ModularBot
**1447–300 over 15 rounds** (ModularBot won round 5 only), firing 1,400 bullets
at a **12.1%** hit rate against ModularBot's 1,134 bullets at **5.3%**. That is
the number the whole gun rack is trying to move. **[MEASURED]** (commit
`17c99f5`, `TR_BRIDGE_FIXTURES.md`).
### 1.4 A/B methodology — server-side per-run hit rate, never scores
The A/B that decides configs uses **server-side ground truth**, not the bot's
own counters and not scores. **[MEASURED]** (`/tmp/analyze.py`,
`/tmp/compare.py`, `/tmp/ab_logs3/FINAL_AB.txt`):
1. The battle runner writes a per-shot **events sidecar** (`/tmp/events_*.json`)
with `fire` / `hit` / damage events stamped with the server's per-round
bullet id (`GunEngine.nextBulletId`). `26b66cb` proved per-gun attribution:
the server assigns the id once and reuses it on `BulletFired`,
`BulletHitBot`, `BulletHitWall`, `BulletHitBullet`; hits that arrive before
the fire event (client priority 70 > 60) are deferred. 99.9% of shots and
99.8% of hits were attributed in that session.
2. A config is judged on its **per-run real hit rate** (`hits/shots` for one
battle), and two configs are called different only if their per-run ranges
**do not overlap**. This is why the overnight A/B reports "SEPARATED" or
"OVERLAP" for every pair rather than a single pooled p-value.
3. **Scores are not used to judge configs.** Single-run scores swing by a
couple of hundred points: the 13 shipped-config runs span **175–526**
(s.d. ≈ 105, range 351; `/tmp/battle_base_r*.log`), so a ~210-point
2-s.d. band swamps any plausible config effect. The gate study (`3c90a59`)
made the same call: no per-adversary score delta exceeded the ~300-point
run-to-run noise band.
The bot-side per-gun attribution (`realShots`/`realHits` in
`/tmp/gun_stats_base_r*.jsonl`) covers ~87% of the server's total shots
uniformly (13 runs: 220/3,157 = 6.97% attributed vs 251/3,612 = 6.95%
server-side), so it is used for per-gun *ranking* but the server sidecar is the
ground truth for config decisions. **[MEASURED]** (re-aggregated from
`/tmp/events_base_r*.json` and `/tmp/gun_stats_base_r*.jsonl`).
---
## Adversary Profiles
## 2. The metric lesson: virtual hit rate is NOT a proxy for real hit rate
| Adversary | Movement | What it tests |
|-----------|----------|---------------|
| **SittingDuck** | Stationary | Baseline ceiling — any gun should near-100%; exposes broken guns |
| **OscillatorBot** | Periodic side-to-side oscillation | Predictable but timed — rewards guns that track phase, punishes naive straight-line |
| **RandomMover** | Uniform random heading changes | Stochastic coverage — tests gun generality under noise |
| **PatternMover** | Repeating movement macro | Tests pattern/memory guns; predictable once pattern is locked |
| **WaveSurfer** | Surfing (dodges incoming bullets reactively) | Most realistic; tests guns that account for dodge response |
This is the most important conceptual result of the night and it invalidates a
naive reading of every offline table in this report.
**[MEASURED]** Aggregating the 13 shipped-config runs against the live DrussGT
boss (`/tmp/gun_stats_base_r{1..13}.jsonl`; `python3 /tmp/agg2.py base`), the
Spearman rank correlation between a gun's **virtual** hit rate and its **real**
hit rate is
```
Spearman(virtual rank, real rank) = -0.374 (n = 13 guns, all with ≥10 real shots)
```
It is not weak — it is **inverted**. The guns with the highest virtual rates
have among the lowest real rates, and vice versa:
| Gun | Selected (ticks) | Real hits/shots | Real % | Virtual % |
|---|---:|---:|---:|---:|
| Linear | 2,707 | 9/84 | **10.7** | 10.2 |
| Circular | 4,051 | 17/172 | **9.9** | 11.9 |
| KNN | 4,886 | 11/122 | **9.0** | 7.5 |
| Pattern | 14,786 | 50/582 | **8.6** | 12.0 |
| Accel | 9,205 | 28/382 | **7.3** | 12.1 |
| AvgLead | 5,119 | 14/200 | **7.0** | 12.3 |
| GuessFactor | 2,345 | 5/72 | **6.9** | 10.3 |
| DecayGF | 1,360 | 3/47 | **6.4** | 9.2 |
| WallBounce | 7,618 | 18/288 | **6.2** | 12.9 |
| StopShot | 3,348 | 8/132 | **6.1** | 12.6 |
| Tsetlin | 3,284 | 6/103 | **5.8** | 12.9 |
| Displace | 2,664 | 4/75 | **5.3** | 12.3 |
| HeadOn | 16,975 | 47/898 | **5.2** | 8.6 |
Read the top and bottom: **Tsetlin, WallBounce and StopShot have the highest
virtual rates (12.6–12.9%) and near-bottom real rates (5.8–6.2%); Linear and
KNN sit at 10.2% / 7.5% virtual but 10.7% / 9.0% real.** The ranking the
virtual metric produces is not merely uninformative, it points the wrong way.
**Why this matters for selection.** What has kept the rack alive is the
selector's **floor/tie hedging**, not its ranking: removing the floor
(`GUN_SELECTOR_FLOOR=0.0`, config `floor00`) drops the rack from 6.95% to
**5.08%** at 175 dmg/run (vs 251) over 788 shots. **[MEASURED]**
(`/tmp/compare.py`). So the selector is useful because it refuses to commit to
a bad field, not because its virtual-rate ordering is good.
**The virtual metric appears anti-correlated no matter which config you pick.**
Measured Spearman per config (5-run 12-round A/B, `/tmp/ab_logs3/FINAL_AB.txt`):
`absolute+point` −0.371, `absolute+path` −0.073, `relative+point` −0.037,
`relative+path` +0.522. But on the large 13-run base set the shipped config
(`relative+path`) is **−0.374**. The sign **flips between run sets**, which is
itself the finding: the correlation is unstable, so no ranking rule built on
it can be trusted. **[MEASURED]** + **[INFERRED]** (the flip is measured; the
conclusion is reasoning).
### 2.1 The metric A/B: point vs path (this one is real, and it is selection)
`GUN_VBULLET_METRIC` picks how a virtual bullet is scored
(`common_libs/gun_harness/virtual_bullets.nim`):
- `bmPoint` — resolve at the fire-time aim distance and score that single
point. Measures prediction accuracy.
- `bmPath` (**shipped**) — fly the ray to the wall and test each swept segment
against the target radius. Measures hypothetical hit chance.
**[MEASURED]** Live A/B against the boss, 5 battles × 12 rounds, one frozen
binary (commit `3b5d70b`; per-run detail in `/tmp/ab_logs/FINAL_ANALYSIS.txt`):
| Metric | Shots | Hits | Real hit rate | Per-run rates | Spearman |
|---|---:|---:|---:|---|---:|
| point | 4,660 | 219 | 4.70% | 5.53 / 5.30 / 4.92 / 3.16 / 4.57 | −0.04 |
| path | 4,834 | 359 | **7.43%** | 6.76 / 8.20 / 8.24 / 6.55 / 7.30 | +0.52 |
The distributions **do not overlap**: path's worst run (6.55%) beats point's
best (5.53%). +2.73 pp, +58% relative, z = 5.56, p < 0.0001. Range
distributions were identical (~460–478 px), so this is not a range confound.
**The gain is selection, not better gun learning.** Under `point` every gun's
virtual rate is compressed into 0.6–4.4%, so HeadOn sits inside the 2 pp tie
margin and takes **72.6% of selection ticks / 76.9% of shots** while ranking
11th of 13 by real hit rate (2.3%). Under `path` the band widens to 4.7–13.7%
and HeadOn's shot share falls to 35.9%, so Pattern/Accel/WallBounce get picked.
The counterfactual confirms it: applying the point model's per-gun real rates
to the path model's shot mix yields 7.65%, i.e. essentially the whole observed
gain. **[MEASURED]** (commit `3b5d70b`).
Offline range total moves the same way: **34.3%** under point
(`/tmp/final_range.txt`, 35,636/104,000) vs **50.8%** under path
(`/tmp/range_path.txt`, 52,770/103,938). The offline totals are inflated by the
perfect-information synthetic fixtures (three of them score 100% under path for
every gun), so the offline totals are *not* comparable to live rates — only to
each other.
### 2.2 The selector-threshold A/B: absolute vs relative
The legacy thresholds were calibrated for a rate scale that does not exist.
**[MEASURED]** offline replay of a *fogged live* `WorldState` vs DrussGT
(1,397 selection ticks, `/tmp/ab_logs3/selector_diag.txt`):
| Config | Floor fires | HeadOn selection share | bestRate med |
|---|---:|---:|---:|
| absolute + point | 53.0% | 69.1% | 8.0% |
| relative + point | 21.2% | 43.5% | 5.25% |
| absolute + path | 3.0% | 23.1% | 24.0% |
| relative + path (shipped) | 8.4% | 24.2% | 16.75% |
The `0.10` absolute floor fires on **53.0%** of point-metric ticks and forces
HeadOn, whose real rate was 2.0–4.4%. (An earlier claim that the floor fires
*always* is **refuted**: it is 53%, because `bestRate` is a max over
gun×power-bin and an occasional ≥50-sample bin clears 10%.)
The scale-aware replacement (commit `dea4dcb`):
`RelTieMargin = 0.20` (tie band is a fraction of `bestRate`), `FloorPeakFrac =
0.25` (floor fires only if the field collapsed vs its own recent peak over a
256-tick window, counting only guns with ≥ MinObsBeforeCompete = 50 samples),
and pooled-over-bins ranking instead of max-over-bins.
Live A/B, 3 runs × 10 rounds (commit `dea4dcb`, `/tmp/ab_logs3/FINAL_AB.txt`):
| Config | Per-run rates | Pooled | vs `absolute+point` |
|---|---|---:|---|
| absolute + point | 3.66 / 2.45 / 5.01 | 3.76% | — |
| absolute + path | 7.55 / 8.21 / 6.83 | 7.57% | SEPARATED (p<0.0001) |
| relative + point | 7.66 / 6.18 / 5.79 | 6.59% | SEPARATED |
| relative + path (**shipped**) | 7.15 / 7.55 / 6.90 | 7.21% | SEPARATED |
`absolute+path` is nominally 0.35 pp above `relative+path`, but they **overlap**
(p = 0.64); so do `relative+point` and both path configs. The **metric** is the
dominant lever; under `path` the two threshold models are statistically tied.
`relative` was shipped because it is the principled scale-aware fix, works
under both metrics, and prevents the point-metric catastrophe if anyone
switches back. **[MEASURED]**.
---
## Results Summary
## 3. Per-gun performance on real numbers
| Adversary | ModularBot Score | Real Hits / Shots | Real Hit% |
|-----------|-----------------|-------------------|-----------|
| SittingDuck | 1944 | 69 / 81 | **85%** |
| OscillatorBot | 1645 | 100 / 170 | **59%** |
| RandomMover | 1916 | 86 / 142 | **61%** |
| PatternMover | 1914 | 80 / 109 | **73%** |
| WaveSurfer | 1873 | 76 / 105 | **72%** |
The final per-gun table, shipped config, **13 runs vs the live DrussGT boss,
3,612 server-side shots, 6.95% overall** (server sidecar; per-run rates 6.77 /
5.90 / 6.57 / 6.34 / 6.10 / 7.59 / 2.90 / 7.95 / 7.49 / 9.18 / 8.44 / 7.48 /
6.42%; 251 dmg/run). Per-gun rows are the bot-side attribution over the same
runs. **[MEASURED]** (commit `2c94dc2`; `/tmp/gun_stats_base_r*.jsonl`,
re-aggregated with `/tmp/agg2.py base`; `/tmp/events_base_r*.json`).
> **Note:** Real hit% was dead during these tests due to the `onBulletHitBot` bug (see Bugs section). Numbers reflect post-fix tracking where the callback fired correctly, but per-gun real-hit attribution was not collected — only aggregate real hits per round are reliable.
| Verdict | Gun | Real hits/shots | Real % | Virtual % | Selected |
|---|---|---:|---:|---:|---:|
| **KEEP** | Linear | 9/84 | 10.7 | 10.2 | 2,707 |
| **KEEP** | Circular | 17/172 | 9.9 | 11.9 | 4,051 |
| **KEEP** | KNN | 11/122 | 9.0 | 7.5 | 4,886 |
| **KEEP** | Pattern | 50/582 | 8.6 | 12.0 | 14,786 |
| **KEEP** | Accel | 28/382 | 7.3 | 12.1 | 9,205 |
| **KEEP** | AvgLead | 14/200 | 7.0 | 12.3 | 5,119 |
| MARGINAL | GuessFactor | 5/72 | 6.9 | 10.3 | 2,345 |
| MARGINAL | DecayGF | 3/47 | 6.4 | 9.2 | 1,360 |
| MARGINAL | WallBounce | 18/288 | 6.2 | 12.9 | 7,618 |
| MARGINAL | StopShot | 8/132 | 6.1 | 12.6 | 3,348 |
| BELOW | Tsetlin | 6/103 | 5.8 | 12.9 | 3,284 |
| BELOW | Displace | 4/75 | 5.3 | 12.3 | 2,664 |
| **FLOOR — STAYS** | HeadOn | 47/898 | 5.2 | 8.6 | 16,975 |
**Verdicts.**
- **KEEP: Linear, Circular, KNN, Pattern, Accel, AvgLead.** These six are at or
above the 6.95% overall, yet their virtual rates are mid-pack to low: the
metric's three favourites (Tsetlin 12.9%, WallBounce 12.9%, StopShot 12.6%)
are near the *bottom* of the real ranking, while the real leader (Linear) sits
at 10.2% virtual. Further evidence the virtual ranking is inverted.
- **MARGINAL: GuessFactor, DecayGF, WallBounce, StopShot.** Within ~1 pp of
overall on small N (47–288 shots). They are not obviously worth deleting, but
they have not earned a larger share.
- **BELOW OVERALL: Tsetlin, Displace.** Below 6% on 75–103 shots. Candidates to
drop or re-tune, but the N is small.
- **HeadOn MUST STAY** despite being lowest (5.2%). It is the floor fallback:
when the field collapses the selector returns gun 0. Disabling the floor
measurably hurt — 5.08% / 175 dmg vs 6.95% / 251 dmg (config `floor00`,
788 shots). Do not delete HeadOn to improve the per-gun average; that
average is computed over shots it only gets because nothing better was
available. **[MEASURED]** (`/tmp/compare.py`).
**The 16-candidate ranking A/B found no winner.** Runtime knobs were added to
the selector (`GUN_SELECTOR_WINDOW`, `MINOBS`, `TIE`, `FLOOR`, `POOL`, `RANK`,
`SHRINK`, `SEED`; `rankScore` supports mean/Wilson/UCB/Thompson/shrinkage), all
defaulting to the shipped values. 16 candidates were A/B'd against the boss.
None credibly beat the shipped config; every candidate's per-run interval
overlaps base, and the nominal "winners" are ≤0.6 SE apart on far fewer shots.
**[MEASURED]** (commit `2c94dc2`; `/tmp/compare.py`):
| Config | Runs | Shots | Rate % | dmg/run | Spearman |
|---|---:|---:|---:|---:|---:|
| **base (shipped)** | 14* | 3,612 | **6.95** | 251 | −0.374 |
| tie00 (TIE=0.0) | 2 | 540 | 7.04 | 242 | +0.018 |
| win50 (WINDOW=50) | 2 | 559 | 6.08 | 216 | +0.588 |
| wilson (RANK=wilson) | 13 | 3,045 | 6.67 | 196 | +0.088 |
| thompson (RANK=thompson) | 2 | 531 | 4.90 | 166 | +0.083 |
| maxbin (POOL=0) | 2 | 574 | 5.23 | 198 | −0.264 |
| minobs20 (MINOBS=20) | 2 | 535 | 5.98 | 212 | +0.144 |
| tie05 (TIE=0.05) | 12 | 3,357 | 6.20 | 222 | −0.060 |
| tie10 (TIE=0.10) | 2 | 556 | 6.65 | 233 | −0.150 |
| tie40 (TIE=0.40) | 2 | 583 | 6.35 | 216 | −0.160 |
| floor10 (FLOOR=0.10) | 6 | 1,694 | 6.49 | 232 | −0.578 |
| t05f10 (TIE=0.05, FLOOR=0.10) | 4 | 1,073 | 5.50 | 190 | −0.041 |
| wilf10 (RANK=wilson, FLOOR=0.10) | 4 | 1,252 | 6.71 | 271 | −0.410 |
| floor00 (FLOOR=0.0) | 3 | 788 | **5.08** | 175 | +0.055 |
| f00t05 (FLOOR=0, TIE=0.05) | 3 | 540 | 6.48 | 153 | +0.226 |
| f00wil (FLOOR=0, RANK=wilson) | 2 | 595 | 6.72 | 221 | −0.116 |
| f00w50 (FLOOR=0, WINDOW=50) | 2 | 593 | 6.58 | 210 | +0.178 |
\* compare.py counts 14 events files, but run 14 has no fire events; the 13
runs with data carry all 3,612 shots.
**No ranking rule fixed the anti-correlation.** The best Spearman in the table
(win50, +0.588) is on 2 runs / 559 shots. The shipped config's −0.374 over 13
runs is the most reliable estimate. **[MEASURED]**.
### 3.1 Offline range: which gun wins which trajectory family
Shipped `path` metric, `/tmp/range_path.txt` (52,770/103,938 = 50.8%). Cells
are hit-% per gun per fixture; **bold** = best gun for that fixture. 400 shots
per gun per fixture.
| Fixture | HeadOn | Linear | Tsetlin | Circular | GuessF | Pattern | WallBn | Accel | StopSh | Displ | AvgLead | DecayG | KNN |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| circular | 12 | 29 | 29 | **100** | 15 | 66 | 34 | 100 | 27 | 18 | 54 | 13 | 73 |
| constant-velocity | **100** | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| contr-SittingDuck | **100** | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| contr-StraightLine | 43 | **100** | 72 | 100 | 100 | 100 | 94 | 98 | 84 | 98 | 100 | 100 | 77 |
| decel-before-turn | 79 | **100** | 90 | 100 | 100 | 100 | 100 | 100 | 98 | 100 | 100 | 100 | 78 |
| classC-corners | 6 | 9 | **23** | 17 | 9 | 17 | 12 | 19 | 19 | 22 | 13 | 10 | 4 |
| classC-crazy | 4 | 36 | 24 | 34 | 35 | 32 | **54** | 35 | 22 | 26 | 39 | 27 | 22 |
| classC-mirror | **12** | 6 | 7 | 6 | 6 | 9 | 6 | 7 | 8 | 5 | 8 | 7 | 8 |
| classC-ramfire | 35 | 50 | 32 | 54 | 50 | 54 | 49 | 49 | 35 | 30 | **56** | 52 | 46 |
| classC-spinbot | 3 | **58** | 10 | 50 | 58 | 37 | 48 | 50 | 13 | 46 | 52 | 58 | 42 |
| energy-threshold-turner | 58 | 34 | 39 | **100** | 34 | 88 | 38 | 100 | 40 | 44 | 63 | 34 | 42 |
| oscillator | **100** | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| random-walk | 26 | 67 | 64 | 28 | 63 | 41 | **68** | 26 | 61 | 67 | 55 | 62 | 47 |
| stationary | **100** | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| TR-corners | 3 | 30 | 28 | 33 | 39 | 26 | 33 | 34 | 30 | 34 | **56** | 29 | 32 |
| TR-crazy | 5 | 7 | 12 | 16 | 2 | 14 | 19 | **28** | 11 | 20 | 27 | 2 | 3 |
| TR-ModularBot | 15 | 12 | 18 | 13 | 12 | 12 | 14 | 15 | **21** | 12 | 13 | 11 | 7 |
| TR-MB-shield | 4 | 10 | 32 | 27 | 10 | 28 | 32 | 28 | 24 | **37** | 21 | 27 | 5 |
| TR-spinbot | 3 | 21 | 20 | 21 | **37** | 37 | 25 | 30 | 27 | 14 | 24 | 15 | 11 |
| wall-bounce | 0 | 49 | 44 | 48 | 37 | 41 | **100** | 60 | 41 | 34 | 52 | 34 | 28 |
Aggregated hit-% by family (400 shots/gun/fixture):
| Family | HeadOn | Linear | Tsetlin | Circular | GuessF | Pattern | WallBn | Accel | StopSh | Displ | AvgLead | DecayG | KNN |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| synthetic (8) | 59 | 72 | 71 | 84 | 69 | 79 | 80 | **86** | 71 | 70 | 78 | 68 | 71 |
| classic DrussGT (5) | 12 | 32 | 19 | 32 | 32 | 30 | **34** | 32 | 20 | 26 | 34 | 31 | 25 |
| TR bridge (5) | 6 | 16 | 22 | 22 | 15 | 23 | 24 | 27 | 23 | 23 | **28** | 17 | 12 |
| all 20 | 35 | 51 | 47 | 57 | 49 | 55 | 56 | **59** | 48 | 50 | 57 | 49 | 46 |
**Which gun wins which trajectory (offline, path metric):**
- **Stationary / constant-velocity / oscillator / decel-before-turn /
straight-line:** every useful gun ≥ 94% (perfect-information + path metric
makes these uninformative). Only HeadOn (43–79%) and KNN (77–78%) stand out
as weak.
- **Circular:** Circular and Accel 100% by construction; KNN 73, Pattern 66.
- **Wall-bounce:** WallBounce 100, Accel 60, AvgLead 52 — the only fixture
where WallBounce is dominant.
- **Random-walk:** WallBounce 68, Displace 67, Linear 67, Tsetlin 64; Accel
collapses to 26.
- **Energy-threshold rule:** Circular and Accel 100, Pattern 88, AvgLead 63.
Tsetlin 39 — it beat Linear (34) but is far from reading the rule.
- **Classic DrussGT (a real surfer):** WallBounce 34 and AvgLead 34 at the top,
HeadOn 12 at the bottom. Cornered surfing is the one case where Tsetlin (23)
leads.
- **TR bridge DrussGT:** AvgLead 28 overall; Accel 28 on TR-crazy, AvgLead 56
on TR-corners, StopShot 21 on the ModularBot mirror, Pattern 37 on TR-spinbot.
**Do not read the offline winner as the rack verdict.** The offline range's
*classic DrussGT* order (WallBounce 34 top, KNN 25 low) is close to the
**inverse** of the real order (KNN 9.0% third, WallBounce 6.2% ninth). The
offline range is excellent for catching structural bugs (§6) and for
per-trajectory sanity, and poor as a selector signal. **[MEASURED]** +
**[INFERRED]**.
---
## Per-Gun Virtual Hit Rate Matrix
## 4. Verdict table
Average virtual hit rate (%) across all rounds per adversary. **Bold** = best gun for that adversary. Selection count (total ticks gun was active) shown in parentheses.
| Gun | SittingDuck | OscillatorBot | RandomMover | PatternMover | WaveSurfer |
|-----|-------------|---------------|-------------|--------------|------------|
| **HeadOn** | 98% (1470) | 63% (664) | 79% (1010) | **95%** (1188) | 84% (1215) |
| **Linear** | 98% (0) | 63% (552) | **82%** (251) | 95% (27) | 84% (13) |
| **Tsetlin** | 98% (0) | 63% (0) | **82%** (0) | 95% (0) | 84% (0) |
| **Circular** | 98% (0) | 63% (3) | 80% (350) | 95% (70) | 84% (137) |
| **GuessFactor** | 98% (0) | 63% (0) | 79% (100) | 94% (0) | 84% (0) |
| **Pattern** | 98% (0) | 9% **(1081)** | 17% (254) | 95% (0) | 84% (77) |
| **AntiSurf** | 0% (0) | 0% (147) | 2% (0) | 0% (19) | 0% (0) |
| **WallBounce** | **99%** (68) | 64% (82) | **82%** (215) | 95% (211) | 85% (51) |
| **Accel** | **99%** (0) | 63% (281) | 19% (152) | 96% (24) | **86%** (255) |
| **StopShot** | **99%** (0) | 64% (1) | **82%** (1) | 96% (6) | **86%** (1) |
| **Displace** | **99%** (0) | 63% (5) | 81% (55) | **96%** (243) | **86%** (10) |
| **AvgLead** | **99%** (0) | **65%** (105) | **82%** (110) | **96%** (107) | **86%** (97) |
| **DecayGF** | **99%** (0) | 63% (0) | 79% (1) | **96%** (20) | **86%** (30) |
| **KNN** | **99%** (0) | 63% (0) | 79% (1) | **96%** (50) | 85% (24) |
**Best gun per adversary:**
- SittingDuck → WallBounce / Accel / StopShot / Displace / AvgLead / DecayGF / KNN (all tied at 99%)
- OscillatorBot → **AvgLead** (65%)
- RandomMover → **Linear / Tsetlin / WallBounce / StopShot / AvgLead** (tied at 82%)
- PatternMover → **Displace / Accel / StopShot / DecayGF / AvgLead / KNN** (tied at 96%)
- WaveSurfer → **StopShot / DecayGF / AvgLead / Accel / Displace** (tied at 86%)
| Gun | Family / model | Real % (13 runs) | Offline all-20 % | Verdict |
|---|---|---:|---:|---|
| Linear | constant-velocity lead | 10.7 | 51 | **KEEP — this is what a surfer cannot defeat; keep warm** |
| Circular | constant-turn lead | 9.9 | 57 | **KEEP** |
| KNN | k-NN on motion history | 9.0 | 46 | **KEEP — offline under-rates it badly** |
| Pattern | pattern replay | 8.6 | 55 | **KEEP** |
| Accel | acceleration-aware lead | 7.3 | 59 | **KEEP** |
| AvgLead | windowed average lead | 7.0 | 57 | **KEEP** |
| GuessFactor | GF histogram | 6.9 | 49 | MARGINAL — small N, no clear edge |
| DecayGF | recency-weighted GF | 6.4 | 49 | MARGINAL |
| WallBounce | wall-reflection model | 6.2 | 56 | MARGINAL — offline favourite, real underperformer |
| StopShot | deceleration/stop point | 6.1 | 48 | MARGINAL |
| Tsetlin | Tsetlin-Machine correction | 5.8 | 47 | BELOW — learns, not yet competitive |
| Displace | displacement vector | 5.3 | 50 | BELOW |
| HeadOn | aim at current position | 5.2 | 35 | **KEEP — mandatory floor fallback** |
| TMSelect | TM mixture-of-experts gate | — | — | **DISABLED (`EnableTmSelector = false`)** — see §6.4 |
---
## Gun-by-Gun Analysis
## 5. What "worth keeping" means, and what is not proven
### HeadOn
**What it does:** Fires directly at current enemy position — no lead, no prediction.
- **Strengths:** SittingDuck (98%), PatternMover (95%), WaveSurfer (84%). Competent baseline everywhere.
- **Weaknesses:** No lead — loses to faster movers once Random/Oscillator velocity matters. Matched or beaten by 8+ other guns on every adversary except SittingDuck.
- **Selection bias:** Receives the most ticks in every matchup (664–1470), consistently crowding out marginally better guns. This is a selector problem, not a gun problem.
- **Verdict:** KEEP — but its selection dominance needs addressing. It is a reliable floor, not the ceiling.
The KEEP/MARGINAL/BELOW split is a statement about a **single adversary (a
wave surfer)**, judged on the shipped config, on a few hundred real shots per
gun. It is a starting point, not a final ranking. The concrete caveats are in
§7.
---
### Linear
**What it does:** Straight-line lead — assumes constant velocity from last observed heading.
- **Strengths:** RandomMover (82%, tied best), WaveSurfer (84%), PatternMover (95%).
- **Weaknesses:** OscillatorBot (63%, same as HeadOn) — oscillation breaks the constant-velocity assumption.
- **Verdict:** KEEP — genuinely better than HeadOn on random movers. Gets 0 selection ticks on most matchups despite competitive rates; underutilised.
## 6. Bugs found and fixed tonight (why earlier rack verdicts were wrong)
Five of these changed the rack ordering; all are [MEASURED] from the commit
messages and the offline range.
### 6.1 The GF family aimed at the fire-time RADIUS, not the angle
`guess_factor`, `decay_gf` and `knn_gun` aimed at the fire-time distance. But
the virtual-bullet metric resolves a bullet at its **aim-point distance** and
scores that single point against the enemy's position on that tick, so with any
radial target motion the bullet stopped at the wrong radius and missed even
with a perfect angle. **Angle-only prediction is structurally unscoreable
under the point metric.** Two competing hypotheses were tested and **both
refuted**: (a) MEA range too narrow — 0 clamped shots out of 837/849/957, with
required offsets peaking at ~33° against MEA 28.1–46.7°, and `arcsin(8/bulletSpeed)`
correctly uses max robot *speed*, not the hit radius; (b) wrong GF peak — a
sweep of every constant GF value showed the oracle-best constant offset was
only 6% on circular, 4% on wall-bounce, 7.5% on random-walk. Learning was fine
too (~850–960 observations per fixture, 0 starved waves). **[MEASURED]**
(commit `7f706e5`).
Fix: a self-consistent constant-velocity `lead_forecast.nim` base, so the
histogram learns the **residual** and the aim point lands at the right radius;
also fixed `linear.nim` (it did a one-shot extrapolation and never iterated its
flight time). Before/after, offline: circular GF 6→23, DecayGF 6→21;
wall-bounce GF 0→60.2, DecayGF 0→60.2; constant-velocity GF/DecayGF/KNN
26→100; random-walk GF 0→53; StraightLine GF 8→77. The oracle-best constant GF
moved 6%→20% (circular), 4%→57% (wall-bounce), 7.5%→49% (random-walk),
proving the structural fix independently of tuning. **Honest trade-off:** on
the 5 real DrussGT surfer captures the GF family regressed (GuessFactor
108→55, DecayGF 108→76, KNN 101→74 hits/2000) because a linear base is a poor
model for a surfer and the residual histogram is noisier than the old
total-lead histogram.
That regression was then recovered by **blending the range** between a
radial-only forecast and the geometric one by the measured radial fraction
(`radialFrac`), keeping the constant-velocity bearing. Nine candidate bases
were measured and rejected with numbers (velocity scaling 0.8 recovered DrussGT
but destroyed wall-bounce 241→20; radial-only range wall-bounce 241→140;
short-window average worse than both; reversal/speed gates weaker than the
blend). Result (hits/2000): classic-5 GF 55→**171**, DecayGF 76→100; TR-5 GF
9→86, DecayGF 4→87; synthetic-10 GF 2702→2717. The only figure below the old
base is classic-5 DecayGF (108→100, within noise). **[MEASURED]** (commit
`e2ca2fc`).
### 6.2 The wave queues were starved 1-push-vs-4-pops
`predict()` stored **one** wave per tick while `onResult()` popped one per
resolved bullet (~4/tick), so the queue drained within a few dozen ticks and
~3 of every 4 resolutions returned without learning; the survivor paired with
a same-tick wave (`bearingDelta ≈ 0`), pinning the histogram at centre.
**Proof:** `GF.vHits == HeadOn.vHits` and `DecayGF.vHits == HeadOn.vHits`
byte-for-byte in **every one of 50 rounds** — GF, DecayGF and KNN had
degenerated into HeadOn clones. Fix: per-bin FIFO with an O(1) head cursor, at
most one push per (tick, bin). Also `maxBullets` 2,048→8,192: the rack spawns
52 bullets/tick so the ring wrapped every ~39 ticks while a long power-3 shot
needs ~90, silently discarding unresolved bullets and biasing every measured
hit rate by range; a `droppedBullets` counter was added. After the fix
`vDropped = 0` and `vStarved = 0` across all 48 recorded rounds. **[MEASURED]**
(commit `0cc6821`).
### 6.3 Tsetlin's clauses saturated at ~714 included literals each
`Tsetlin.vHits` was byte-for-byte equal to `Linear.vHits` in every measured
round of every run because its learned correction was always exactly 0.
Root cause: `tmLearnOne` rewarded included true literals unconditionally,
omitting Granmo's `(c=0, lk=1) → toward Exclude` counter-force, so true
literals ratcheted toward Include forever; Type II was unreachable dead code
with the wrong direction; resource allocation was an `|error|` heuristic
instead of Granmo's `(T − clip(v,−T,T))/(2T)`; the label baseline had a
factor-2 shrink (`error = δ − 2c`, fixed point `c = δ/2`); hits zeroed their
residual; the enemy-energy feature was duplicated (`state.selfEnergy` fed
where `WorldState.enemyEnergy` exists, so energy rules were literally
unrepresentable); and `tmEvalClause` needed Granmo Eq. 6 (all-Exclude clause
outputs 1 during learning, 0 during classification) or fix #1 deadlocks every
clause at empty.
Measured effect (energy-threshold-turner, seed 1): mean included
literals/clause **714.0 → 13.8**; active clauses 100/100 → 53/100; nonzero
corrections 8/764 → 708/764; Tsetlin virtual hits **27/400 → 69/400** (Linear
43/400). Tsetlin now **learns** but is **not yet competitive with Linear** —
the regression head is untuned, flagged as follow-up rather than claimed as a
win. **[MEASURED]** (commit `8937000`; `/tmp/ab_logs3/final_test_tsetlin_gun.log`).
### 6.4 The TM classifier gun did not earn its slot (but its clauses are real)
A Tsetlin-Machine mixture-of-experts gate over
HeadOn/Linear/Circular/WallBounce/Accel was built with the corrected feedback
and labelled by which expert's prediction was closest to the actual enemy
position (an exact, supervised, per-shot label — no delayed credit). It loses
to the best of its own experts offline on nearly every fixture, and against
DrussGT it cost real performance:
```
baseline (path + relative) 7.56% real hit rate, 157 dmg
+ power fix 7.47%, 239 dmg
+ power fix + TM selector 5.59%, 133 dmg
```
It was selected on 806 ticks and fired 24 real shots at 4.2%. It ships disabled
(`EnableTmSelector = false`; code and wiring kept intact). **However, the gate
latched onto meaningful structure:** on the energy-threshold turner, HeadOn's
clauses key on the **energy bits** (the rule's own driving variable) while
Circular keys on distance/velocity. So the TM learned something real and
interpretable; it simply could not beat "always pick the best expert".
**[INFERRED]** root cause: the closest-expert label is noisy because several
experts are near-tied, and under the path metric the winner varies by power bin
while the gate sees one shared per-tick input, so a one-vs-rest gate over a
saturated 870-bit clause space has no margin to exploit. A standalone Granmo
classifier on the same encoding reaches ~99% on the rule but, per the
counterfactual probe, does **not** read energy (follow rate 24% high / 62% mean
— statistically identical at 1, 2 and 10 frames), so even the "it learned the
rule" claim is limited to ~99% accuracy, not to a readable energy threshold.
The best recovered proposition was `!g9 ∧ !g8` (energy < 25.6, not the labelled
30) — a genuine simple threshold, but not the ensemble's decision mechanism.
**[MEASURED]** (commits `57b2ac3`, `d5061ee`; `test_tm_pattern_learning.nim`).
### 6.5 The selector thresholds were absolute on a rescaled metric
Covered in §2.2: the `0.10` absolute floor fired on 53.0% of point-metric ticks
and forced HeadOn (real 2.0–4.4%, 11th of 13); HeadOn selection share fell
69.1% → 43.5% under relative thresholds (and 23.1% → 24.2% under path). Also:
`bestGun` was first-index-wins argmax, so HeadOn at index 0 silently won every
tie until the random tie-break landed (`343e631`); `bestPower` had the same
absolute-40% defect (below).
### 6.6 Power selection was stuck at power 1.0 (`MinHitRate = 0.40`)
`bestPower` used an **absolute** `MinHitRate = 0.40` bar. Measured per-bin
virtual rates show no bin ever clears 40%, so 11 of 14 guns were stuck at bin 0
(power 1.0) even where higher bins were comparable:
```
Linear p1.0 44% p1.5 39% p2.0 30% p3.0 29% old bin 0 -> new bin 3
Accel p1.0 44% p1.5 40% p2.0 26% p3.0 29% old bin 1 -> new bin 3
Pattern p1.0 50% p1.5 40% p2.0 27% p3.0 12% old bin 1 -> new bin 2
```
Replaced with a scale-aware `PowerBarFrac = 0.50` (a dimensionless fraction of
the gun's own best-bin rate); 13 of 14 selections now pick heavier bullets.
Real effect vs DrussGT (8 rounds × 3 runs): hit rate unchanged (7.56% → 7.47%),
**damage +52% (157 → 239 per run)** and rounds end faster. **[MEASURED]**
(commit `57b2ac3`).
Smaller fixes in the same family: `bestPower` on a cold gun returned the
*highest* bin (empty bin satisfied the `count == 0` clause); `fitnessFor`
aggregated enemies in nondeterministic hash order; `stop_shot` had an
unreachable deceleration branch and several guns had tick-only caches that made
all four power bins return bin 0's lead (`e536900`). **[MEASURED]**.
---
### Tsetlin
**What it does:** Tsetlin Machine classifier for aim prediction (binary learning automaton).
- **Strengths:** Identical rates to Linear across all adversaries (82% Random, 63% Oscillator, 95% Pattern, 84% Wave, 98% Sitting).
- **Weaknesses:** Gets 0 selection ticks in every single matchup — never selected once. Either never crosses the 15-observation threshold, or its virtual bullets are being computed identically to Linear.
- **Verdict:** TUNE — investigate why selection count is always 0. If the virtual hit computation is a clone of another gun's trajectory, the gun is effectively dead weight. Confirm it fires unique virtual bullets.
## 7. Known caveats and open problems
Stated without hedging.
1. **The headline per-gun numbers come from ONE adversary, a wave surfer.**
HeadOn is genuinely bad against surfers, so part of the rack ordering may be
matchup-specific. A SpinBot guard was inconclusive: ModularBot fires only
17–31 real shots/run against a fast bot because the range-aware firing gate
is strict at long range, so the guard had little power (Wilson looked better,
18.5% vs 8.6%, but on 70–92 shots with a 5–33% spread). **[MEASURED]**
(commit `2c94dc2`). A second, independent adversary at scale is missing.
2. **Per-gun real N is small.** 47–898 shots per gun; n < 200 gives roughly
±5 pp across a 3–15% spread. Single-gun ordering is **indicative, not
definitive**. The KEEP/MARGINAL/BELOW boundaries should be treated as soft.
3. **The fixtures are perfect-information and therefore optimistic.** Every
fixture is an observer capture with true positions every tick; the classic
set is additionally **open-loop** (replayed DrussGT never dodges our
bullets). Absolute offline hit rates are inflated by an unknown amount; only
relative comparisons are safe.
4. **The virtual metric is anti-correlated with real hit rate and no tested
ranking rule fixed it.** Shipped config Spearman ≈ **−0.374** over 13 runs
(the sign flips to +0.52 on the smaller 5-run set, so it is unstable). 16
candidate ranking rules all overlapped the shipped config. The selector's
value lives in its **floor/tie hedging** (5.08% without the floor vs 6.95%
with it), not in its ranking. **[MEASURED]**.
5. **The TM classifier gun did not earn its slot.** It cost real performance
(7.47% → 5.59%, 133 dmg) despite showing interpretable energy structure in
its clauses (§6.4). It is disabled; re-enabling requires a fix to the gate
margin/label problem, not more training.
6. **Real-hit-rate-driven selection is not viable yet.** Only the selected gun
fires, so unselected guns get near-zero real shots (GuessFactor 20, Linear
24 vs HeadOn 733 in the point A/B); noise is fatal (n = 470 at p = 10% gives
±2.8 pp, most guns n < 200 gives ±5 pp+); and real rate is conditional on
when the gun was selected. A blended signal with forced exploration and
shrinkage is defensible in principle but needs thousands of shots per gun
across many battles. Real rate is currently best used **offline** as the
evaluation metric — which is exactly what the A/B does. **[MEASURED]**
(commit `dea4dcb`).
7. **The offline==online acceptance test is flaky** (§1.1): typically 11/12 on
unmodified HEAD, with the mismatching gun varying run to run. The
equivalence claim is strong-but-not-exact until the boundary race is fixed.
8. **The selector's random tie-break is not randomised in the live bot.** The
shipped bot never calls `randomize()`, so the "random" sequence is fixed
across process restarts (a side finding of `2c94dc2`, not fixed).
9. **The firing gate is not the bottleneck.** The shipped range-aware gate does
not beat a fixed 2.0° gate on hit rate (55.8% vs 57.9%, ~1.5 σ), though it
fires 22–28% more shots. No per-adversary score delta exceeded the
~300-point run-to-run noise band. **[MEASURED]** (commit `3c90a59`).
10. **The boss is ~2.3× more accurate than the whole rack** (12.1% vs 5.3% in
the capture). Closing that gap is the point of the rack; the current
best single gun is 10.7%.
---
### Circular
**What it does:** Circular lead — assumes constant angular velocity (orbit).
- **Strengths:** RandomMover (80%), decent on most adversaries.
- **Weaknesses:** Lags behind Linear/WallBounce on Random; barely competitive elsewhere.
- **Verdict:** KEEP (marginal) — distinct from Linear (angular vs linear velocity assumption), worth keeping for orbit-heavy bots. Gets real selection only on RandomMover (350 ticks) and WaveSurfer (137).
## 8. Reproduction
---
Commands recorded in the commits and tool READMEs. (I was instructed not to
run builds/tests while writing this report; these are the documented
invocations, not a fresh verification by me.)
### GuessFactor
**What it does:** Guess-factor segmentation gun — divides lateral movement into bins and tracks historical hit density.
- **Strengths:** SittingDuck 98%, generally solid floor across all enemies.
- **Weaknesses:** Never edges above HeadOn on any adversary; OscillatorBot 63% (same floor). Needs history to warm up — may lag early rounds.
- **Selection:** Modest (100 RandomMover, 0 elsewhere). Learning gun that never became dominant.
- **Verdict:** KEEP — GF guns are the backbone of competitive targeting. Gets outperformed by AvgLead/DecayGF variants — investigate whether GF segmentation resolution needs tuning.
```bash
# Offline gun range over all 20 fixtures, shipped path metric (default):
nim c -r common_libs/tests/run_range.nim
# Point metric for comparison:
GUN_VBULLET_METRIC=point nim c -r common_libs/tests/run_range.nim
# Add timing:
nim c -r common_libs/tests/run_range.nim --timing
---
# Selector diagnostics (floor/tie/bestRate/HeadOn-share) on a fixture:
GUN_SELECTOR_MODE=relative nim c -d:release -r \
common_libs/tests/analyze_selector.nim tools/fixtures/drussgt_vs_spinbot.jsonl
### Pattern
**What it does:** Movement pattern matching — records enemy movement sequences and replays prediction.
- **Strengths:** SittingDuck 98%, PatternMover 95%, WaveSurfer 84% (when selected).
- **Weaknesses:** **OscillatorBot 9%** — catastrophically wrong; 1081 selection ticks wasted there. RandomMover 17% — random movement breaks pattern replay entirely.
- **Critical bug:** The selector gave Pattern 1081 ticks vs OscillatorBot despite a 9% rate. This means `MinObsBeforeCompete` let it compete before it had enough data to show a representative rate, or the rolling window allowed a lucky early patch to lock in selection.
- **Verdict:** TUNE — the gun itself is sound against true pattern movers; the problem is premature selection. Raise `MinObsBeforeCompete` or add a "minimum rounds active" gate before Pattern can dominate the selector.
# Offline == online acceptance (currently flaky):
nim c -r common_libs/tests/acceptance_offline_vs_online.nim
---
# Tsetlin gun clause sparsity / divergence:
nim c -r common_libs/tests/test_tsetlin_gun.nim
# TM readability (standalone Granmo classifier on the energy-threshold rule):
nim c -r common_libs/tests/test_tm_pattern_learning.nim
### AntiSurf
**What it does:** Anti-surfing gun — attempts to predict dodge direction by modelling bullet-wave response.
- **Strengths:** None observed. 0% virtual hit rate against every adversary in every round.
- **Weaknesses:** Broken or misconfigured — 0% is not a statistical artifact; it is a systematic failure. Scored 147 selection ticks vs OscillatorBot and 19 vs PatternMover despite 0% rate.
- **Verdict:** DROP or fix before retest. 0% vhit means virtual bullets are being fired in the wrong direction entirely. The selection ticks it received represent rounds where it was preferred over functional guns — a liability.
# Live boss (real DrussGT jar; jars stay out of git, see the README):
# tools/robocode_shim/run_bridge_battle.sh <bot_dir> <rounds> <capture.jsonl>
# Closed-loop evidence for the TR captures:
python3 tools/robocode_shim/analyze_closed_loop.py \
tools/fixtures/tr_drussgt_vs_modularbot.jsonl \
tools/robocode_shim/evidence/tr_drussgt_vs_modularbot.events.json
```
---
### WallBounce
**What it does:** Predicts enemy movement reflecting off walls.
- **Strengths:** SittingDuck 99% (joint best), RandomMover 82% (tied best), OscillatorBot 64%, PatternMover 95%, WaveSurfer 85%.
- **Weaknesses:** None significant — consistently top-tier across all adversary types.
- **Selection:** Gets real ticks on SittingDuck (68), OscillatorBot (82), RandomMover (215), PatternMover (211), WaveSurfer (51) — the selector does pick it when scores differentiate.
- **Verdict:** KEEP — one of the most consistent guns in the rack. Wall-aware prediction generalises well.
---
### Accel
**What it does:** Acceleration-based lead — accounts for changing velocity.
- **Strengths:** SittingDuck 99%, PatternMover 96%, WaveSurfer 86% (joint best). Strong against momentum-based movers.
- **Weaknesses:** **RandomMover 19%** — severe drop. Random heading changes invalidate acceleration extrapolation completely.
- **Selection:** 281 ticks OscillatorBot, 255 WaveSurfer, 152 RandomMover. Gets selection time on Oscillator despite 63% (same as HeadOn) — selector noise.
- **Verdict:** KEEP but note the RandomMover cliff. The 19% rate while receiving 152 selection ticks vs Random is wasteful — the selector should have dropped it faster.
---
### StopShot
**What it does:** Fires at the predicted stop point — targets deceleration patterns.
- **Strengths:** WaveSurfer 86% (joint best), PatternMover 96%, SittingDuck 99%, RandomMover 82%. Broad competence.
- **Weaknesses:** Gets almost zero selection despite top-tier rates (1 tick on most matchups). Under-selected almost everywhere.
- **Verdict:** KEEP — likely under-selected because it ties with AvgLead but has fewer early observations. Its consistent top performance suggests it should get more selection time.
---
### Displace
**What it does:** Displacement-based prediction — fires at position + expected displacement vector.
- **Strengths:** PatternMover 96% (joint best), WaveSurfer 86%, SittingDuck 99%.
- **Weaknesses:** RandomMover 81% (slightly below Linear/WallBounce).
- **Selection:** Gets decent ticks on PatternMover (243) — the selector finds it there.
- **Verdict:** KEEP — solid and well-selected where it's strong.
---
### AvgLead
**What it does:** Averaged lead angle over a window — smoothed velocity prediction.
- **Strengths:** **OscillatorBot 65% (best gun)**, WaveSurfer 86% (joint best), PatternMover 96%, RandomMover 82%, SittingDuck 99%.
- **Weaknesses:** None — top or joint-top across all five adversaries.
- **Selection:** Gets consistent moderate ticks everywhere (97–110), indicating the selector finds it reliable.
- **Verdict:** KEEP — the most consistently high-performing gun in the rack. If only one gun were kept, this would be it.
---
### DecayGF
**What it does:** Guess-factor with exponential decay — recent history weighted more heavily.
- **Strengths:** WaveSurfer 86% (joint best), PatternMover 96%, SittingDuck 99%.
- **Weaknesses:** OscillatorBot 63% (same as HeadOn floor). Needs time to build a useful GF profile.
- **Selection:** Gets 30 ticks on WaveSurfer, 20 on PatternMover — modest but real.
- **Verdict:** KEEP — the decay weighting should help against adapting movers; verify the GF profile is being updated correctly.
---
### KNN
**What it does:** K-nearest neighbours targeting — finds historical situations most similar to current and fires the historically best angle.
- **Strengths:** PatternMover 96% (joint best), WaveSurfer 85%, SittingDuck 99%.
- **Weaknesses:** OscillatorBot 63% (floor), needs history to warm up.
- **Selection:** 50 PatternMover, 24 WaveSurfer — gets some selection where it performs.
- **Verdict:** KEEP — KNN is a sound approach that should improve as history accumulates across rounds. Cold-start performance is expected to be weak.
---
## Recommendations
### 1. Guns to DROP
| Gun | Reason |
|-----|--------|
| **AntiSurf** | 0% virtual hit rate against all adversaries. Not statistically bad — systematically broken. Received selection ticks on two matchups despite never hitting anything. Fix or remove; as-is it is pure negative value. |
### 2. Guns to TUNE
| Gun | What to change |
|-----|---------------|
| **Pattern** | Raise `MinObsBeforeCompete` significantly (suggest ≥ 50 obs). Add a round-count gate: Pattern should not be eligible to dominate the selector until it has fired virtual bullets for at least 2 full rounds. The 1081-tick / 9% disaster vs OscillatorBot is entirely a selector gate failure. |
| **Tsetlin** | Investigate zero selection count. Check that virtual bullets are computed with a unique angle (distinct from Linear). If it is accidentally sharing a firing angle with another gun, it will never differentiate. |
| **Accel** | Its 19% vs RandomMover while receiving 152 selection ticks is wasteful. The rolling window should surface this faster — check if the window length (100 ticks) is too long for Accel's decay signal to flush bad data quickly. |
| **GuessFactor** | Investigate why it never edges above HeadOn. GF is a proven algorithm — if it is not beating HeadOn after 10 rounds vs PatternMover or WaveSurfer, the bin resolution or lateral-velocity calculation may be wrong. |
### 3. Selector Improvements
- **Pattern over-selection:** `MinObsBeforeCompete = 15` is too low for a memory gun. Pattern needs hundreds of ticks to build a meaningful database. Gate it at 100+ observations or require a minimum database size before competing.
- **HeadOn selection dominance:** HeadOn gets 1010–1470 ticks per matchup while marginally better guns (AvgLead, WallBounce) get 50–250. The selector may be hitting HeadOn early when all guns are equal, then staying there due to recency bias. Consider a uniform-random tiebreak when rates are within 1–2%.
- **Accel vs Random:** 152 wasted ticks at 19% vhit. The 100-tick rolling window should expose this; check if the window resets between rounds or carries over (if it carries over from a good round, bad-round data is diluted).
### 4. Missing Coverage
- **True surfer with timing randomness:** WaveSurfer is the closest to real tournament surfing bots, but AvgLead/StopShot/Accel all hit 86%. Consider testing against a stronger surfer variant to stress-test wave-aware guns.
- **Fast random with perpendicular escape:** Accel's 19% on RandomMover suggests fast lateral movers are poorly handled. No gun clearly dominates RandomMover — the 82% ceiling is shared by four guns, none clearly specialised.
- **Anti-mirror bots:** No adversary tests symmetric evasion. If AntiSurf is ever fixed, it needs a proper test target.
---
## Bugs Found During Testing
### 1. `onBulletHitBot` → `onBulletHit` (critical)
**Location:** ModularBot event handler registration.
**Problem:** The callback method was named `onBulletHitBot` but the Robocode API fires `onBulletHit`. Real hit tracking was silently dead — the callback never fired, so `realHits` in the gun stats relied on whatever fallback was in place.
**Impact:** Per-gun real hit rates could not be attributed correctly. The `realHits` field in the JSONL reflects aggregate post-fix behaviour but per-gun real-hit breakdown was unavailable for this gauntlet. Virtual hit rates remain valid (they are computed independently of the event callback).
**Fix:** Rename to `onBulletHit`. Already corrected; re-run the gauntlet to get clean real-hit attribution per gun.
### 2. AntiSurf 0% vhit (potential implementation bug)
**Observed:** AntiSurf returns 0 virtual hits against every adversary including SittingDuck. SittingDuck does not move — any gun aimed at the current position should achieve near-100%.
**Implication:** AntiSurf is not firing virtual bullets at the target position at all, or its angle computation produces NaN/wrong quadrant. This is a code bug, not a tuning issue.
**Suggested check:** Print the computed firing angle for AntiSurf virtual bullets and verify it is within ±π/2 of the actual bearing to target.
### 3. Pattern gun gets selected during initial warm-up
**Observed:** Pattern accumulates 1081 selection ticks vs OscillatorBot at 9% vhit. With `MinObsBeforeCompete = 15`, it became eligible before its pattern database was meaningful, won a random tiebreak at 100% during early ticks (before any miss was recorded), then the rolling window retained that optimistic early rate too long.
**Fix:** See Recommendations §3.
Per-gun aggregation scripts used for the tables above:
`python3 /tmp/agg2.py base` (virtual-vs-real + Spearman),
`python3 /tmp/compare.py` (server-side per-run A/B + overlap).
Raw range output: `/tmp/range_path.txt` (path), `/tmp/final_range.txt` (point).
+50 -32
View File
@@ -1,40 +1,58 @@
# Gun Rack Summary — TL;DR
**Gauntlet:** 5 adversaries × 10 rounds, 14 guns. Data: `/tmp/gun_stats.jsonl`
**Date:** 2026-09-21 · **Bot:** ModularBot (13 active guns + 1 disabled TM gate)
**Shipped config:** metric `path`, selector `relative`.
**Ground truth:** server-side real hit rate (events sidecar), per-run, with an
explicit overlap test. Scores are not used — they swing by a couple of hundred
points per run.
## Scoreboard
**Boss:** the unmodified DrussGT jar through `tools/robocode_shim/`.
**Headline (13 runs vs DrussGT, 3,612 server-side shots): 6.95% real hit rate,
251 dmg/run.**
| Adversary | Score | Real Hit% | Best Gun |
|-----------|-------|-----------|----------|
| SittingDuck | 1944 | 85% | AvgLead / WallBounce (99% vhit) |
| OscillatorBot | 1645 | 59% | AvgLead (65% vhit) |
| RandomMover | 1916 | 61% | Linear / WallBounce / AvgLead (82%) |
| PatternMover | 1914 | 73% | AvgLead / Displace / Accel (96%) |
| WaveSurfer | 1873 | 72% | AvgLead / StopShot / Accel (86%) |
**The one thing to know:** *virtual hit rate is not a proxy for real hit rate.*
For the shipped config Spearman(virtual rank, real rank) = **−0.374** —
anti-correlated. 16 candidate ranking rules all failed to beat the shipped
config; what works is the selector's floor/tie hedging (removing the floor:
5.08% / 175 dmg vs 6.95% / 251 dmg). Detail: [`gun_rack_analysis.md`](gun_rack_analysis.md).
## Verdict Table
## Verdict table
| Gun | Best Use | Verdict |
|-----|----------|---------|
| **HeadOn** | Stationary / pattern bots | KEEP — reliable floor, over-selected |
| **Linear** | Random movers | KEEP — under-selected despite 82% |
| **Tsetlin** | Unknown | TUNE — 0 selection ticks ever; investigate |
| **Circular** | Orbit-heavy bots | KEEP (marginal) |
| **GuessFactor** | General | KEEP — underperforms vs expectation; tune bins |
| **Pattern** | Pattern movers | TUNE — catastrophic vs OscillatorBot (9%, 1081 ticks); raise MinObs gate |
| **AntiSurf** | — | DROP — 0% vhit every adversary including stationary; broken |
| **WallBounce** | All types | KEEP — most consistent overall |
| **Accel** | Pattern / wave bots | KEEP — avoid vs random (19% cliff) |
| **StopShot** | All types | KEEP — top rates, chronically under-selected |
| **Displace** | Pattern / wave bots | KEEP |
| **AvgLead** | All types | KEEP — best all-rounder in the rack |
| **DecayGF** | Wave / pattern bots | KEEP |
| **KNN** | Pattern / wave bots | KEEP — improves with history |
| Verdict | Gun | Real % | Virtual % |
|---|---|---:|---:|
| **KEEP** | Linear | 10.7 | 10.2 |
| **KEEP** | Circular | 9.9 | 11.9 |
| **KEEP** | KNN | 9.0 | 7.5 |
| **KEEP** | Pattern | 8.6 | 12.0 |
| **KEEP** | Accel | 7.3 | 12.1 |
| **KEEP** | AvgLead | 7.0 | 12.3 |
| MARGINAL | GuessFactor | 6.9 | 10.3 |
| MARGINAL | DecayGF | 6.4 | 9.2 |
| MARGINAL | WallBounce | 6.2 | 12.9 |
| MARGINAL | StopShot | 6.1 | 12.6 |
| BELOW | Tsetlin | 5.8 | 12.9 |
| BELOW | Displace | 5.3 | 12.3 |
| **FLOOR — KEEP** | HeadOn | 5.2 | 8.6 |
| DISABLED | TMSelect | — | — |
## Top Actions
HeadOn is lowest but **must stay**: it is the floor fallback, and disabling the
floor measurably hurt (5.08% / 175 dmg).
1. **Fix AntiSurf** — 0% vhit against SittingDuck means wrong angle computation, not just weak targeting.
2. **Raise Pattern's MinObsBeforeCompete** — 15 is too low; 1081 wasted ticks at 9% hit rate vs OscillatorBot.
3. **Re-run gauntlet** with `onBulletHit` fix to get clean per-gun real hit attribution.
4. **AvgLead** is the star gun — confirm it stays in the top selection tier.
5. **HeadOn selection dominance** is a selector bias problem, not a gun problem — add ±2% tiebreak randomisation.
## Top actions
1. **Stop trusting the virtual metric as a ranker.** It is anti-correlated with
real hit rate and unstable across run sets. The selector survives on its
floor/tie hedge, not on its ordering.
2. **Fix the offline==online acceptance race.** It is currently flaky
(typically 11/12), so offline numbers are strong-but-not-exact.
3. **Get a second adversary.** Every per-gun verdict rests on one wave surfer;
the KEEP/BELOW boundaries are matchup-specific and per-gun N is small
(47–898 shots).
4. **Decide Tsetlin and Displace.** Both are below overall on small N; Tsetlin
now learns (clauses 714→13.8 literals) but is not competitive — tune the
regression head or drop.
5. **Do not re-enable TMSelect** until the gate-margin/label problem is fixed;
it cost 7.47% → 5.59% despite showing real energy structure in its clauses.
6. **Real-hit-rate-driven selection is not viable yet** — unselected guns get
near-zero shots, so it needs forced exploration + shrinkage + thousands of
shots per gun.