19410164f1
Replaces the stale 2026-09-20 docs, which predated per-gun real attribution, the offline gun range and the DrussGT boss, and whose verdicts were built on virtual hit rates that turned out to be ANTI-correlated with reality. docs/gun_rack_analysis.md (751 lines) covers: the test infrastructure described honestly (offline range with its flaky-acceptance caveat, the 20 fixtures and what each set is good for, the live boss, and the A/B methodology of per-run server-side real hit rate with an explicit overlap test); the virtual-vs-real metric lesson with Spearman -0.374 and the point-vs-path A/B; per-gun real performance and the 16-rule ranking A/B; the offline per-fixture gun matrix; KEEP/MARGINAL/BELOW verdicts; and the five root-cause bugs with before/after numbers. docs/gun_rack_summary.md (58 lines) is the verdict table plus top actions. The '~230 point' score-noise band that has been steering methodology all night was re-derived from the artifacts rather than asserted: the 13 shipped-config run scores span 175-526, s.d. ~105, i.e. a ~210-point 2-s.d. band. Caveats recorded verbatim rather than softened: the offline==online acceptance is flaky (typically 11/12 on unmodified HEAD), fixtures are perfect-information and therefore optimistic vs live play, per-gun real N is small so single-gun ordering is indicative, the headline numbers come from ONE wave-surfer adversary, and HeadOn must stay despite being lowest because it is the floor fallback (disabling it: 5.08% / 175 dmg vs 6.95% / 251 dmg).
752 lines
41 KiB
Markdown
752 lines
41 KiB
Markdown
# Gun Rack Analysis — ModularBot guns vs the real DrussGT boss
|
||
|
||
**Date:** 2026-09-21
|
||
**Bot:** ModularBot (13 active guns + 1 disabled TM classifier gun)
|
||
**Shipped config:** virtual-bullet metric `path`, selector thresholds `relative`
|
||
(`GUN_VBULLET_METRIC` default `path`, `GUN_SELECTOR_MODE` default `relative`)
|
||
**Rewritten:** this file supersedes the 2026-09-20 version, whose numbers
|
||
predated per-gun real-hit attribution, the offline gun range, and the live
|
||
DrussGT boss. No number from the old version is retained unless it was
|
||
re-measured below.
|
||
|
||
Evidence tags used throughout:
|
||
|
||
- **[MEASURED]** — I read it from a recorded artifact (commit message, `/tmp`
|
||
output, a log, or a source file). The provenance is named every time.
|
||
- **[INFERRED]** — reasoning from measured facts; explicitly not measured.
|
||
|
||
A note on provenance: the measurements below were produced overnight in
|
||
commits `343e631` … `2c94dc2` on branch `research/lead-targeting`. Each commit
|
||
message records the numbers, the hypotheses it refuted, and its caveats. Raw
|
||
artifacts are in `/tmp` (`/tmp/gun_stats_base_r*.jsonl`,
|
||
`/tmp/events_base_r*.json`, `/tmp/range_path.txt`,
|
||
`/tmp/ab_logs/FINAL_ANALYSIS.txt`, `/tmp/ab_logs3/FINAL_AB.txt`,
|
||
`/tmp/ab_logs3/selector_diag.txt`). Where a commit message and a raw artifact
|
||
disagree slightly, both are stated.
|
||
|
||
---
|
||
|
||
## 1. The test infrastructure
|
||
|
||
The numbers are only as good as the rig that produced them, so the rig is
|
||
described first.
|
||
|
||
### 1.1 The offline gun range
|
||
|
||
`common_libs/gun_harness/offline_range.nim` replays a recorded
|
||
`seq[WorldState]` through the **same** `VirtualTracker`
|
||
(`common_libs/gun_harness/virtual_bullets.nim`) that the live bot drives. The
|
||
claim is not "an approximation of the live metric", it is "the same metric":
|
||
virtual-bullet fitness is already a pure function of (a stream of
|
||
`WorldState`, a list of guns), and the Java battle only supplies where the
|
||
states come from. Guns keep their own internal history, so a replayed stream in
|
||
order is a complete movement history. **[MEASURED]** (`offline_range.nim`,
|
||
header comment; commit `974528d`).
|
||
|
||
- **Speed and sample count.** 8 fixtures / 1,770 ticks / ~92 k virtual bullets /
|
||
13 guns replay in 2.9 s at ~32 k virtual bullets/s — roughly **70× faster and
|
||
100× more samples** than a live gauntlet. **[MEASURED]** (commit `974528d`).
|
||
- **Acceptance test.** `common_libs/tests/acceptance_offline_vs_online.nim`
|
||
records one live round, replays it offline, and compares per-gun virtual hit
|
||
counts. It passed **12/12 deterministic guns** (Tsetlin is separately labelled
|
||
stochastic because `tmLearnOne` calls `rand()`). Two real live-loop ordering
|
||
quirks had to be modelled to reach that: `run()` calls `go()` before the
|
||
aim/fire block, so `tickBullets` resolves against the *next* tick's scan; and
|
||
if the target dies during that `go()`, the final tick's spawn+resolution is
|
||
skipped. **[MEASURED]** (commit `974528d`; a passing run is preserved in
|
||
`/tmp/ab_logs3/test_acceptance.log`: 12/12, 534-tick round).
|
||
- **Honest caveat — the acceptance test is currently FLAKY.** On the
|
||
*unmodified HEAD* source it normally reaches only **11/12**, e.g. KNN 81
|
||
online vs 71 offline, and the mismatching gun moves between runs (KNN, then
|
||
WallBounce). It is a live/offline boundary race, pre-existing, and not caused
|
||
by the selector work (the replay never calls the selector). Treat
|
||
"offline == online" as **strong but not exact until the race is fixed**.
|
||
**[MEASURED]** (commit `2c94dc2`). The 12/12 runs above are real; they were
|
||
lucky runs.
|
||
|
||
### 1.2 The fixture sets and what each is good for
|
||
|
||
There are **20 JSONL fixtures** under `tools/fixtures/`. They fall into four
|
||
groups. **[MEASURED]** (`tools/fixtures/`, `DRUSSGT_FIXTURES.md`,
|
||
`TR_BRIDGE_FIXTURES.md`).
|
||
|
||
**(a) 8 synthetic fixtures — ground truth by construction.**
|
||
`common_libs/gun_harness/offline_range.nim` generates them; each has a known
|
||
rule:
|
||
|
||
| Fixture | Generator / rule | What it is good for |
|
||
|---|---|---|
|
||
| `stationary` | fixed enemy | ceiling: any correct gun scores ~100% |
|
||
| `constant-velocity` | straight line, no walls | does the gun lead a moving target at all |
|
||
| `circular` | constant turn 3°/tick | circular/accel models |
|
||
| `wall-bounce` | specular reflection off all 4 walls | wall-aware prediction |
|
||
| `oscillator` | east 30 ticks, west 30 ticks | phase/timing |
|
||
| `random-walk` | seeded ±15°/tick jitter | generality under noise |
|
||
| `decel-before-turn` | cruise → full stop → pivot 3×45° → accelerate | stop-shot detection |
|
||
| `energy-threshold-turner` | **KNOWN RULE**: straight while `energy ≥ 30`, hard 20°/tick turn while `energy < 30`, `energy = max(5, 50 − 0.5·t)` | can a learner find a readable high-level rule; threshold crosses at t=41 |
|
||
|
||
The energy-threshold turner is the falsifiable one: the label is literally a
|
||
predicate over the 11-bit Gray-coded energy field, so a learner that reads
|
||
energy can be shown to have read the *right* variable (see §6.3).
|
||
|
||
**(b) 2 classic-Robocode contrast fixtures.** `contrast_stationary_sittingduck`
|
||
and `contrast_straightline` — trivial motion, used as a sanity/ceiling check.
|
||
**[MEASURED]**.
|
||
|
||
**(c) 5 classic-Robocode DrussGT captures.** Real, unmodified DrussGT
|
||
3.1.4159 movement captured from Robocode 1.9.5.5 via
|
||
`tools/robocode_fixture_capture/` (file `DRUSSGT_FIXTURES.md`). Opponents:
|
||
SpinBot, RamFire, Crazy, Corners, and a DrussGT mirror; 28,797 ticks. The
|
||
coordinate conversion is validated to **0.000–0.001°** on every fixture by
|
||
recomputing the direction implied by (heading, speed) against the recorded
|
||
per-tick displacement. **These are OPEN-LOOP and perfect-information:** the
|
||
replayed DrussGT never dodges *our* bullets, and the observer reports true
|
||
positions every tick (unlike the live bot's stale between-scan `WorldState`).
|
||
They are therefore optimistic and good for *relative* gun ranking, not absolute
|
||
hit rates. **[MEASURED]** (`DRUSSGT_FIXTURES.md`).
|
||
|
||
**(d) 5 closed-loop Tank Royale DrussGT captures.** The real DrussGT jar playing
|
||
Tank Royale through `tools/robocode_shim/`, captured by
|
||
`tools/robocode_shim/src/robocode_shim/TrBattleCapture.java` (file
|
||
`TR_BRIDGE_FIXTURES.md`). Primary file `tr_drussgt_vs_modularbot.jsonl`:
|
||
15 rounds, 20,026 ticks, ModularBot fired 1,134 shots. At capture time DrussGT
|
||
was reacting to *our* real bullets. **Closed-loop proven, not asserted:**
|
||
`tools/robocode_shim/analyze_closed_loop.py` event-locks `|Δheading|` to
|
||
ModularBot's fire times (a heat-limited near-metronome, median interval
|
||
14 ticks) and gets an oscillating response with the fire period; the
|
||
cross-correlation peaks at **r = +0.111, lag 12, permutation p = 0.005** (null
|
||
peak mean +0.016), and the **own-fire control is flat**, so the lock is
|
||
enemy-driven, not internal cadence. **Still perfect-information**, and
|
||
open-loop *at replay time* — "closed_loop" describes the capture, not a later
|
||
replay. TR angle conversion residual is ~1.5° mean (vs 0.000° classic) because
|
||
the TR server moves along the pre-turn heading. **[MEASURED]**
|
||
(`TR_BRIDGE_FIXTURES.md`).
|
||
|
||
| Fixture | Source | rounds | ticks | adversary / note |
|
||
|---|---|---:|---:|---|
|
||
| `circular` … `energy-threshold-turner` | synthetic | — | 150–260 | 8 known-rule trajectories |
|
||
| `contrast_stationary_sittingduck`, `contrast_straightline` | classic-robocode | 1 each | 1,451 | sanity contrasts |
|
||
| `drussgt_vs_spinbot` | classic-robocode | 20 | 5,002 | open-loop, perfect-info |
|
||
| `drussgt_vs_ramfire` | classic-robocode | 20 | 3,240 | open-loop, perfect-info |
|
||
| `drussgt_vs_crazy` | classic-robocode | 20 | 9,025 | open-loop, perfect-info |
|
||
| `drussgt_vs_corners` | classic-robocode | 20 | 4,975 | open-loop, perfect-info |
|
||
| `drussgt_vs_drussgt` | classic-robocode | 2 | 6,555 | open-loop, perfect-info (mirror) |
|
||
| `tr_drussgt_vs_modularbot` | tr-bridge | 15 | 20,026 | **closed-loop at capture**, perfect-info |
|
||
| `tr_drussgt_vs_modularbot_shield` | tr-bridge | 10 | 12,629 | shield on |
|
||
| `tr_drussgt_vs_spinbot` / `_crazy` / `_corners` | tr-bridge | 10 each | 10,824 / 11,507 / 2,575 | closed-loop at capture |
|
||
|
||
### 1.3 The live boss — the real DrussGT jar
|
||
|
||
`tools/robocode_shim/` runs the **unmodified** `DrussGT.jar` (159,289 bytes,
|
||
md5 `5cd6015dcc6d6da8a7e6aeecb1fec211`) as a Tank Royale bot. The classic
|
||
`robocode.*` API is a thin delegation layer over the public
|
||
`IBasicRobotPeer`/`IAdvancedRobotPeer` seam, so the shim reuses the genuine
|
||
`robocode.jar` and implements only the 75-method peer interface
|
||
(`ClassicPeer`), plus `BotHost`/`ThreadManagerFix`. DrussGT compiles with zero
|
||
shim API symbols and runs real battles; the EnergyDome shield is disabled by
|
||
default (pure wave surfer). Known physics divergences (move/turn ordering,
|
||
distance bookkeeping, etc.) are enumerated in section 5.9 of
|
||
`tools/robocode_shim/README.md`. **[MEASURED]**.
|
||
|
||
The boss is far stronger than us: in the capture battle DrussGT beat ModularBot
|
||
**1447–300 over 15 rounds** (ModularBot won round 5 only), firing 1,400 bullets
|
||
at a **12.1%** hit rate against ModularBot's 1,134 bullets at **5.3%**. That is
|
||
the number the whole gun rack is trying to move. **[MEASURED]** (commit
|
||
`17c99f5`, `TR_BRIDGE_FIXTURES.md`).
|
||
|
||
### 1.4 A/B methodology — server-side per-run hit rate, never scores
|
||
|
||
The A/B that decides configs uses **server-side ground truth**, not the bot's
|
||
own counters and not scores. **[MEASURED]** (`/tmp/analyze.py`,
|
||
`/tmp/compare.py`, `/tmp/ab_logs3/FINAL_AB.txt`):
|
||
|
||
1. The battle runner writes a per-shot **events sidecar** (`/tmp/events_*.json`)
|
||
with `fire` / `hit` / damage events stamped with the server's per-round
|
||
bullet id (`GunEngine.nextBulletId`). `26b66cb` proved per-gun attribution:
|
||
the server assigns the id once and reuses it on `BulletFired`,
|
||
`BulletHitBot`, `BulletHitWall`, `BulletHitBullet`; hits that arrive before
|
||
the fire event (client priority 70 > 60) are deferred. 99.9% of shots and
|
||
99.8% of hits were attributed in that session.
|
||
2. A config is judged on its **per-run real hit rate** (`hits/shots` for one
|
||
battle), and two configs are called different only if their per-run ranges
|
||
**do not overlap**. This is why the overnight A/B reports "SEPARATED" or
|
||
"OVERLAP" for every pair rather than a single pooled p-value.
|
||
3. **Scores are not used to judge configs.** Single-run scores swing by a
|
||
couple of hundred points: the 13 shipped-config runs span **175–526**
|
||
(s.d. ≈ 105, range 351; `/tmp/battle_base_r*.log`), so a ~210-point
|
||
2-s.d. band swamps any plausible config effect. The gate study (`3c90a59`)
|
||
made the same call: no per-adversary score delta exceeded the ~300-point
|
||
run-to-run noise band.
|
||
|
||
The bot-side per-gun attribution (`realShots`/`realHits` in
|
||
`/tmp/gun_stats_base_r*.jsonl`) covers ~87% of the server's total shots
|
||
uniformly (13 runs: 220/3,157 = 6.97% attributed vs 251/3,612 = 6.95%
|
||
server-side), so it is used for per-gun *ranking* but the server sidecar is the
|
||
ground truth for config decisions. **[MEASURED]** (re-aggregated from
|
||
`/tmp/events_base_r*.json` and `/tmp/gun_stats_base_r*.jsonl`).
|
||
|
||
---
|
||
|
||
## 2. The metric lesson: virtual hit rate is NOT a proxy for real hit rate
|
||
|
||
This is the most important conceptual result of the night and it invalidates a
|
||
naive reading of every offline table in this report.
|
||
|
||
**[MEASURED]** Aggregating the 13 shipped-config runs against the live DrussGT
|
||
boss (`/tmp/gun_stats_base_r{1..13}.jsonl`; `python3 /tmp/agg2.py base`), the
|
||
Spearman rank correlation between a gun's **virtual** hit rate and its **real**
|
||
hit rate is
|
||
|
||
```
|
||
Spearman(virtual rank, real rank) = -0.374 (n = 13 guns, all with ≥10 real shots)
|
||
```
|
||
|
||
It is not weak — it is **inverted**. The guns with the highest virtual rates
|
||
have among the lowest real rates, and vice versa:
|
||
|
||
| Gun | Selected (ticks) | Real hits/shots | Real % | Virtual % |
|
||
|---|---:|---:|---:|---:|
|
||
| Linear | 2,707 | 9/84 | **10.7** | 10.2 |
|
||
| Circular | 4,051 | 17/172 | **9.9** | 11.9 |
|
||
| KNN | 4,886 | 11/122 | **9.0** | 7.5 |
|
||
| Pattern | 14,786 | 50/582 | **8.6** | 12.0 |
|
||
| Accel | 9,205 | 28/382 | **7.3** | 12.1 |
|
||
| AvgLead | 5,119 | 14/200 | **7.0** | 12.3 |
|
||
| GuessFactor | 2,345 | 5/72 | **6.9** | 10.3 |
|
||
| DecayGF | 1,360 | 3/47 | **6.4** | 9.2 |
|
||
| WallBounce | 7,618 | 18/288 | **6.2** | 12.9 |
|
||
| StopShot | 3,348 | 8/132 | **6.1** | 12.6 |
|
||
| Tsetlin | 3,284 | 6/103 | **5.8** | 12.9 |
|
||
| Displace | 2,664 | 4/75 | **5.3** | 12.3 |
|
||
| HeadOn | 16,975 | 47/898 | **5.2** | 8.6 |
|
||
|
||
Read the top and bottom: **Tsetlin, WallBounce and StopShot have the highest
|
||
virtual rates (12.6–12.9%) and near-bottom real rates (5.8–6.2%); Linear and
|
||
KNN sit at 10.2% / 7.5% virtual but 10.7% / 9.0% real.** The ranking the
|
||
virtual metric produces is not merely uninformative, it points the wrong way.
|
||
|
||
**Why this matters for selection.** What has kept the rack alive is the
|
||
selector's **floor/tie hedging**, not its ranking: removing the floor
|
||
(`GUN_SELECTOR_FLOOR=0.0`, config `floor00`) drops the rack from 6.95% to
|
||
**5.08%** at 175 dmg/run (vs 251) over 788 shots. **[MEASURED]**
|
||
(`/tmp/compare.py`). So the selector is useful because it refuses to commit to
|
||
a bad field, not because its virtual-rate ordering is good.
|
||
|
||
**The virtual metric appears anti-correlated no matter which config you pick.**
|
||
Measured Spearman per config (5-run 12-round A/B, `/tmp/ab_logs3/FINAL_AB.txt`):
|
||
`absolute+point` −0.371, `absolute+path` −0.073, `relative+point` −0.037,
|
||
`relative+path` +0.522. But on the large 13-run base set the shipped config
|
||
(`relative+path`) is **−0.374**. The sign **flips between run sets**, which is
|
||
itself the finding: the correlation is unstable, so no ranking rule built on
|
||
it can be trusted. **[MEASURED]** + **[INFERRED]** (the flip is measured; the
|
||
conclusion is reasoning).
|
||
|
||
### 2.1 The metric A/B: point vs path (this one is real, and it is selection)
|
||
|
||
`GUN_VBULLET_METRIC` picks how a virtual bullet is scored
|
||
(`common_libs/gun_harness/virtual_bullets.nim`):
|
||
|
||
- `bmPoint` — resolve at the fire-time aim distance and score that single
|
||
point. Measures prediction accuracy.
|
||
- `bmPath` (**shipped**) — fly the ray to the wall and test each swept segment
|
||
against the target radius. Measures hypothetical hit chance.
|
||
|
||
**[MEASURED]** Live A/B against the boss, 5 battles × 12 rounds, one frozen
|
||
binary (commit `3b5d70b`; per-run detail in `/tmp/ab_logs/FINAL_ANALYSIS.txt`):
|
||
|
||
| Metric | Shots | Hits | Real hit rate | Per-run rates | Spearman |
|
||
|---|---:|---:|---:|---|---:|
|
||
| point | 4,660 | 219 | 4.70% | 5.53 / 5.30 / 4.92 / 3.16 / 4.57 | −0.04 |
|
||
| path | 4,834 | 359 | **7.43%** | 6.76 / 8.20 / 8.24 / 6.55 / 7.30 | +0.52 |
|
||
|
||
The distributions **do not overlap**: path's worst run (6.55%) beats point's
|
||
best (5.53%). +2.73 pp, +58% relative, z = 5.56, p < 0.0001. Range
|
||
distributions were identical (~460–478 px), so this is not a range confound.
|
||
|
||
**The gain is selection, not better gun learning.** Under `point` every gun's
|
||
virtual rate is compressed into 0.6–4.4%, so HeadOn sits inside the 2 pp tie
|
||
margin and takes **72.6% of selection ticks / 76.9% of shots** while ranking
|
||
11th of 13 by real hit rate (2.3%). Under `path` the band widens to 4.7–13.7%
|
||
and HeadOn's shot share falls to 35.9%, so Pattern/Accel/WallBounce get picked.
|
||
The counterfactual confirms it: applying the point model's per-gun real rates
|
||
to the path model's shot mix yields 7.65%, i.e. essentially the whole observed
|
||
gain. **[MEASURED]** (commit `3b5d70b`).
|
||
|
||
Offline range total moves the same way: **34.3%** under point
|
||
(`/tmp/final_range.txt`, 35,636/104,000) vs **50.8%** under path
|
||
(`/tmp/range_path.txt`, 52,770/103,938). The offline totals are inflated by the
|
||
perfect-information synthetic fixtures (three of them score 100% under path for
|
||
every gun), so the offline totals are *not* comparable to live rates — only to
|
||
each other.
|
||
|
||
### 2.2 The selector-threshold A/B: absolute vs relative
|
||
|
||
The legacy thresholds were calibrated for a rate scale that does not exist.
|
||
**[MEASURED]** offline replay of a *fogged live* `WorldState` vs DrussGT
|
||
(1,397 selection ticks, `/tmp/ab_logs3/selector_diag.txt`):
|
||
|
||
| Config | Floor fires | HeadOn selection share | bestRate med |
|
||
|---|---:|---:|---:|
|
||
| absolute + point | 53.0% | 69.1% | 8.0% |
|
||
| relative + point | 21.2% | 43.5% | 5.25% |
|
||
| absolute + path | 3.0% | 23.1% | 24.0% |
|
||
| relative + path (shipped) | 8.4% | 24.2% | 16.75% |
|
||
|
||
The `0.10` absolute floor fires on **53.0%** of point-metric ticks and forces
|
||
HeadOn, whose real rate was 2.0–4.4%. (An earlier claim that the floor fires
|
||
*always* is **refuted**: it is 53%, because `bestRate` is a max over
|
||
gun×power-bin and an occasional ≥50-sample bin clears 10%.)
|
||
|
||
The scale-aware replacement (commit `dea4dcb`):
|
||
`RelTieMargin = 0.20` (tie band is a fraction of `bestRate`), `FloorPeakFrac =
|
||
0.25` (floor fires only if the field collapsed vs its own recent peak over a
|
||
256-tick window, counting only guns with ≥ MinObsBeforeCompete = 50 samples),
|
||
and pooled-over-bins ranking instead of max-over-bins.
|
||
|
||
Live A/B, 3 runs × 10 rounds (commit `dea4dcb`, `/tmp/ab_logs3/FINAL_AB.txt`):
|
||
|
||
| Config | Per-run rates | Pooled | vs `absolute+point` |
|
||
|---|---|---:|---|
|
||
| absolute + point | 3.66 / 2.45 / 5.01 | 3.76% | — |
|
||
| absolute + path | 7.55 / 8.21 / 6.83 | 7.57% | SEPARATED (p<0.0001) |
|
||
| relative + point | 7.66 / 6.18 / 5.79 | 6.59% | SEPARATED |
|
||
| relative + path (**shipped**) | 7.15 / 7.55 / 6.90 | 7.21% | SEPARATED |
|
||
|
||
`absolute+path` is nominally 0.35 pp above `relative+path`, but they **overlap**
|
||
(p = 0.64); so do `relative+point` and both path configs. The **metric** is the
|
||
dominant lever; under `path` the two threshold models are statistically tied.
|
||
`relative` was shipped because it is the principled scale-aware fix, works
|
||
under both metrics, and prevents the point-metric catastrophe if anyone
|
||
switches back. **[MEASURED]**.
|
||
|
||
---
|
||
|
||
## 3. Per-gun performance on real numbers
|
||
|
||
The final per-gun table, shipped config, **13 runs vs the live DrussGT boss,
|
||
3,612 server-side shots, 6.95% overall** (server sidecar; per-run rates 6.77 /
|
||
5.90 / 6.57 / 6.34 / 6.10 / 7.59 / 2.90 / 7.95 / 7.49 / 9.18 / 8.44 / 7.48 /
|
||
6.42%; 251 dmg/run). Per-gun rows are the bot-side attribution over the same
|
||
runs. **[MEASURED]** (commit `2c94dc2`; `/tmp/gun_stats_base_r*.jsonl`,
|
||
re-aggregated with `/tmp/agg2.py base`; `/tmp/events_base_r*.json`).
|
||
|
||
| Verdict | Gun | Real hits/shots | Real % | Virtual % | Selected |
|
||
|---|---|---:|---:|---:|---:|
|
||
| **KEEP** | Linear | 9/84 | 10.7 | 10.2 | 2,707 |
|
||
| **KEEP** | Circular | 17/172 | 9.9 | 11.9 | 4,051 |
|
||
| **KEEP** | KNN | 11/122 | 9.0 | 7.5 | 4,886 |
|
||
| **KEEP** | Pattern | 50/582 | 8.6 | 12.0 | 14,786 |
|
||
| **KEEP** | Accel | 28/382 | 7.3 | 12.1 | 9,205 |
|
||
| **KEEP** | AvgLead | 14/200 | 7.0 | 12.3 | 5,119 |
|
||
| MARGINAL | GuessFactor | 5/72 | 6.9 | 10.3 | 2,345 |
|
||
| MARGINAL | DecayGF | 3/47 | 6.4 | 9.2 | 1,360 |
|
||
| MARGINAL | WallBounce | 18/288 | 6.2 | 12.9 | 7,618 |
|
||
| MARGINAL | StopShot | 8/132 | 6.1 | 12.6 | 3,348 |
|
||
| BELOW | Tsetlin | 6/103 | 5.8 | 12.9 | 3,284 |
|
||
| BELOW | Displace | 4/75 | 5.3 | 12.3 | 2,664 |
|
||
| **FLOOR — STAYS** | HeadOn | 47/898 | 5.2 | 8.6 | 16,975 |
|
||
|
||
**Verdicts.**
|
||
|
||
- **KEEP: Linear, Circular, KNN, Pattern, Accel, AvgLead.** These six are at or
|
||
above the 6.95% overall, yet their virtual rates are mid-pack to low: the
|
||
metric's three favourites (Tsetlin 12.9%, WallBounce 12.9%, StopShot 12.6%)
|
||
are near the *bottom* of the real ranking, while the real leader (Linear) sits
|
||
at 10.2% virtual. Further evidence the virtual ranking is inverted.
|
||
- **MARGINAL: GuessFactor, DecayGF, WallBounce, StopShot.** Within ~1 pp of
|
||
overall on small N (47–288 shots). They are not obviously worth deleting, but
|
||
they have not earned a larger share.
|
||
- **BELOW OVERALL: Tsetlin, Displace.** Below 6% on 75–103 shots. Candidates to
|
||
drop or re-tune, but the N is small.
|
||
- **HeadOn MUST STAY** despite being lowest (5.2%). It is the floor fallback:
|
||
when the field collapses the selector returns gun 0. Disabling the floor
|
||
measurably hurt — 5.08% / 175 dmg vs 6.95% / 251 dmg (config `floor00`,
|
||
788 shots). Do not delete HeadOn to improve the per-gun average; that
|
||
average is computed over shots it only gets because nothing better was
|
||
available. **[MEASURED]** (`/tmp/compare.py`).
|
||
|
||
**The 16-candidate ranking A/B found no winner.** Runtime knobs were added to
|
||
the selector (`GUN_SELECTOR_WINDOW`, `MINOBS`, `TIE`, `FLOOR`, `POOL`, `RANK`,
|
||
`SHRINK`, `SEED`; `rankScore` supports mean/Wilson/UCB/Thompson/shrinkage), all
|
||
defaulting to the shipped values. 16 candidates were A/B'd against the boss.
|
||
None credibly beat the shipped config; every candidate's per-run interval
|
||
overlaps base, and the nominal "winners" are ≤0.6 SE apart on far fewer shots.
|
||
**[MEASURED]** (commit `2c94dc2`; `/tmp/compare.py`):
|
||
|
||
| Config | Runs | Shots | Rate % | dmg/run | Spearman |
|
||
|---|---:|---:|---:|---:|---:|
|
||
| **base (shipped)** | 14* | 3,612 | **6.95** | 251 | −0.374 |
|
||
| tie00 (TIE=0.0) | 2 | 540 | 7.04 | 242 | +0.018 |
|
||
| win50 (WINDOW=50) | 2 | 559 | 6.08 | 216 | +0.588 |
|
||
| wilson (RANK=wilson) | 13 | 3,045 | 6.67 | 196 | +0.088 |
|
||
| thompson (RANK=thompson) | 2 | 531 | 4.90 | 166 | +0.083 |
|
||
| maxbin (POOL=0) | 2 | 574 | 5.23 | 198 | −0.264 |
|
||
| minobs20 (MINOBS=20) | 2 | 535 | 5.98 | 212 | +0.144 |
|
||
| tie05 (TIE=0.05) | 12 | 3,357 | 6.20 | 222 | −0.060 |
|
||
| tie10 (TIE=0.10) | 2 | 556 | 6.65 | 233 | −0.150 |
|
||
| tie40 (TIE=0.40) | 2 | 583 | 6.35 | 216 | −0.160 |
|
||
| floor10 (FLOOR=0.10) | 6 | 1,694 | 6.49 | 232 | −0.578 |
|
||
| t05f10 (TIE=0.05, FLOOR=0.10) | 4 | 1,073 | 5.50 | 190 | −0.041 |
|
||
| wilf10 (RANK=wilson, FLOOR=0.10) | 4 | 1,252 | 6.71 | 271 | −0.410 |
|
||
| floor00 (FLOOR=0.0) | 3 | 788 | **5.08** | 175 | +0.055 |
|
||
| f00t05 (FLOOR=0, TIE=0.05) | 3 | 540 | 6.48 | 153 | +0.226 |
|
||
| f00wil (FLOOR=0, RANK=wilson) | 2 | 595 | 6.72 | 221 | −0.116 |
|
||
| f00w50 (FLOOR=0, WINDOW=50) | 2 | 593 | 6.58 | 210 | +0.178 |
|
||
|
||
\* compare.py counts 14 events files, but run 14 has no fire events; the 13
|
||
runs with data carry all 3,612 shots.
|
||
|
||
**No ranking rule fixed the anti-correlation.** The best Spearman in the table
|
||
(win50, +0.588) is on 2 runs / 559 shots. The shipped config's −0.374 over 13
|
||
runs is the most reliable estimate. **[MEASURED]**.
|
||
|
||
### 3.1 Offline range: which gun wins which trajectory family
|
||
|
||
Shipped `path` metric, `/tmp/range_path.txt` (52,770/103,938 = 50.8%). Cells
|
||
are hit-% per gun per fixture; **bold** = best gun for that fixture. 400 shots
|
||
per gun per fixture.
|
||
|
||
| Fixture | HeadOn | Linear | Tsetlin | Circular | GuessF | Pattern | WallBn | Accel | StopSh | Displ | AvgLead | DecayG | KNN |
|
||
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
||
| circular | 12 | 29 | 29 | **100** | 15 | 66 | 34 | 100 | 27 | 18 | 54 | 13 | 73 |
|
||
| constant-velocity | **100** | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
|
||
| contr-SittingDuck | **100** | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
|
||
| contr-StraightLine | 43 | **100** | 72 | 100 | 100 | 100 | 94 | 98 | 84 | 98 | 100 | 100 | 77 |
|
||
| decel-before-turn | 79 | **100** | 90 | 100 | 100 | 100 | 100 | 100 | 98 | 100 | 100 | 100 | 78 |
|
||
| classC-corners | 6 | 9 | **23** | 17 | 9 | 17 | 12 | 19 | 19 | 22 | 13 | 10 | 4 |
|
||
| classC-crazy | 4 | 36 | 24 | 34 | 35 | 32 | **54** | 35 | 22 | 26 | 39 | 27 | 22 |
|
||
| classC-mirror | **12** | 6 | 7 | 6 | 6 | 9 | 6 | 7 | 8 | 5 | 8 | 7 | 8 |
|
||
| classC-ramfire | 35 | 50 | 32 | 54 | 50 | 54 | 49 | 49 | 35 | 30 | **56** | 52 | 46 |
|
||
| classC-spinbot | 3 | **58** | 10 | 50 | 58 | 37 | 48 | 50 | 13 | 46 | 52 | 58 | 42 |
|
||
| energy-threshold-turner | 58 | 34 | 39 | **100** | 34 | 88 | 38 | 100 | 40 | 44 | 63 | 34 | 42 |
|
||
| oscillator | **100** | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
|
||
| random-walk | 26 | 67 | 64 | 28 | 63 | 41 | **68** | 26 | 61 | 67 | 55 | 62 | 47 |
|
||
| stationary | **100** | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
|
||
| TR-corners | 3 | 30 | 28 | 33 | 39 | 26 | 33 | 34 | 30 | 34 | **56** | 29 | 32 |
|
||
| TR-crazy | 5 | 7 | 12 | 16 | 2 | 14 | 19 | **28** | 11 | 20 | 27 | 2 | 3 |
|
||
| TR-ModularBot | 15 | 12 | 18 | 13 | 12 | 12 | 14 | 15 | **21** | 12 | 13 | 11 | 7 |
|
||
| TR-MB-shield | 4 | 10 | 32 | 27 | 10 | 28 | 32 | 28 | 24 | **37** | 21 | 27 | 5 |
|
||
| TR-spinbot | 3 | 21 | 20 | 21 | **37** | 37 | 25 | 30 | 27 | 14 | 24 | 15 | 11 |
|
||
| wall-bounce | 0 | 49 | 44 | 48 | 37 | 41 | **100** | 60 | 41 | 34 | 52 | 34 | 28 |
|
||
|
||
Aggregated hit-% by family (400 shots/gun/fixture):
|
||
|
||
| Family | HeadOn | Linear | Tsetlin | Circular | GuessF | Pattern | WallBn | Accel | StopSh | Displ | AvgLead | DecayG | KNN |
|
||
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
||
| synthetic (8) | 59 | 72 | 71 | 84 | 69 | 79 | 80 | **86** | 71 | 70 | 78 | 68 | 71 |
|
||
| classic DrussGT (5) | 12 | 32 | 19 | 32 | 32 | 30 | **34** | 32 | 20 | 26 | 34 | 31 | 25 |
|
||
| TR bridge (5) | 6 | 16 | 22 | 22 | 15 | 23 | 24 | 27 | 23 | 23 | **28** | 17 | 12 |
|
||
| all 20 | 35 | 51 | 47 | 57 | 49 | 55 | 56 | **59** | 48 | 50 | 57 | 49 | 46 |
|
||
|
||
**Which gun wins which trajectory (offline, path metric):**
|
||
|
||
- **Stationary / constant-velocity / oscillator / decel-before-turn /
|
||
straight-line:** every useful gun ≥ 94% (perfect-information + path metric
|
||
makes these uninformative). Only HeadOn (43–79%) and KNN (77–78%) stand out
|
||
as weak.
|
||
- **Circular:** Circular and Accel 100% by construction; KNN 73, Pattern 66.
|
||
- **Wall-bounce:** WallBounce 100, Accel 60, AvgLead 52 — the only fixture
|
||
where WallBounce is dominant.
|
||
- **Random-walk:** WallBounce 68, Displace 67, Linear 67, Tsetlin 64; Accel
|
||
collapses to 26.
|
||
- **Energy-threshold rule:** Circular and Accel 100, Pattern 88, AvgLead 63.
|
||
Tsetlin 39 — it beat Linear (34) but is far from reading the rule.
|
||
- **Classic DrussGT (a real surfer):** WallBounce 34 and AvgLead 34 at the top,
|
||
HeadOn 12 at the bottom. Cornered surfing is the one case where Tsetlin (23)
|
||
leads.
|
||
- **TR bridge DrussGT:** AvgLead 28 overall; Accel 28 on TR-crazy, AvgLead 56
|
||
on TR-corners, StopShot 21 on the ModularBot mirror, Pattern 37 on TR-spinbot.
|
||
|
||
**Do not read the offline winner as the rack verdict.** The offline range's
|
||
*classic DrussGT* order (WallBounce 34 top, KNN 25 low) is close to the
|
||
**inverse** of the real order (KNN 9.0% third, WallBounce 6.2% ninth). The
|
||
offline range is excellent for catching structural bugs (§6) and for
|
||
per-trajectory sanity, and poor as a selector signal. **[MEASURED]** +
|
||
**[INFERRED]**.
|
||
|
||
---
|
||
|
||
## 4. Verdict table
|
||
|
||
| Gun | Family / model | Real % (13 runs) | Offline all-20 % | Verdict |
|
||
|---|---|---:|---:|---|
|
||
| Linear | constant-velocity lead | 10.7 | 51 | **KEEP — this is what a surfer cannot defeat; keep warm** |
|
||
| Circular | constant-turn lead | 9.9 | 57 | **KEEP** |
|
||
| KNN | k-NN on motion history | 9.0 | 46 | **KEEP — offline under-rates it badly** |
|
||
| Pattern | pattern replay | 8.6 | 55 | **KEEP** |
|
||
| Accel | acceleration-aware lead | 7.3 | 59 | **KEEP** |
|
||
| AvgLead | windowed average lead | 7.0 | 57 | **KEEP** |
|
||
| GuessFactor | GF histogram | 6.9 | 49 | MARGINAL — small N, no clear edge |
|
||
| DecayGF | recency-weighted GF | 6.4 | 49 | MARGINAL |
|
||
| WallBounce | wall-reflection model | 6.2 | 56 | MARGINAL — offline favourite, real underperformer |
|
||
| StopShot | deceleration/stop point | 6.1 | 48 | MARGINAL |
|
||
| Tsetlin | Tsetlin-Machine correction | 5.8 | 47 | BELOW — learns, not yet competitive |
|
||
| Displace | displacement vector | 5.3 | 50 | BELOW |
|
||
| HeadOn | aim at current position | 5.2 | 35 | **KEEP — mandatory floor fallback** |
|
||
| TMSelect | TM mixture-of-experts gate | — | — | **DISABLED (`EnableTmSelector = false`)** — see §6.4 |
|
||
|
||
---
|
||
|
||
## 5. What "worth keeping" means, and what is not proven
|
||
|
||
The KEEP/MARGINAL/BELOW split is a statement about a **single adversary (a
|
||
wave surfer)**, judged on the shipped config, on a few hundred real shots per
|
||
gun. It is a starting point, not a final ranking. The concrete caveats are in
|
||
§7.
|
||
|
||
---
|
||
|
||
## 6. Bugs found and fixed tonight (why earlier rack verdicts were wrong)
|
||
|
||
Five of these changed the rack ordering; all are [MEASURED] from the commit
|
||
messages and the offline range.
|
||
|
||
### 6.1 The GF family aimed at the fire-time RADIUS, not the angle
|
||
|
||
`guess_factor`, `decay_gf` and `knn_gun` aimed at the fire-time distance. But
|
||
the virtual-bullet metric resolves a bullet at its **aim-point distance** and
|
||
scores that single point against the enemy's position on that tick, so with any
|
||
radial target motion the bullet stopped at the wrong radius and missed even
|
||
with a perfect angle. **Angle-only prediction is structurally unscoreable
|
||
under the point metric.** Two competing hypotheses were tested and **both
|
||
refuted**: (a) MEA range too narrow — 0 clamped shots out of 837/849/957, with
|
||
required offsets peaking at ~33° against MEA 28.1–46.7°, and `arcsin(8/bulletSpeed)`
|
||
correctly uses max robot *speed*, not the hit radius; (b) wrong GF peak — a
|
||
sweep of every constant GF value showed the oracle-best constant offset was
|
||
only 6% on circular, 4% on wall-bounce, 7.5% on random-walk. Learning was fine
|
||
too (~850–960 observations per fixture, 0 starved waves). **[MEASURED]**
|
||
(commit `7f706e5`).
|
||
|
||
Fix: a self-consistent constant-velocity `lead_forecast.nim` base, so the
|
||
histogram learns the **residual** and the aim point lands at the right radius;
|
||
also fixed `linear.nim` (it did a one-shot extrapolation and never iterated its
|
||
flight time). Before/after, offline: circular GF 6→23, DecayGF 6→21;
|
||
wall-bounce GF 0→60.2, DecayGF 0→60.2; constant-velocity GF/DecayGF/KNN
|
||
26→100; random-walk GF 0→53; StraightLine GF 8→77. The oracle-best constant GF
|
||
moved 6%→20% (circular), 4%→57% (wall-bounce), 7.5%→49% (random-walk),
|
||
proving the structural fix independently of tuning. **Honest trade-off:** on
|
||
the 5 real DrussGT surfer captures the GF family regressed (GuessFactor
|
||
108→55, DecayGF 108→76, KNN 101→74 hits/2000) because a linear base is a poor
|
||
model for a surfer and the residual histogram is noisier than the old
|
||
total-lead histogram.
|
||
|
||
That regression was then recovered by **blending the range** between a
|
||
radial-only forecast and the geometric one by the measured radial fraction
|
||
(`radialFrac`), keeping the constant-velocity bearing. Nine candidate bases
|
||
were measured and rejected with numbers (velocity scaling 0.8 recovered DrussGT
|
||
but destroyed wall-bounce 241→20; radial-only range wall-bounce 241→140;
|
||
short-window average worse than both; reversal/speed gates weaker than the
|
||
blend). Result (hits/2000): classic-5 GF 55→**171**, DecayGF 76→100; TR-5 GF
|
||
9→86, DecayGF 4→87; synthetic-10 GF 2702→2717. The only figure below the old
|
||
base is classic-5 DecayGF (108→100, within noise). **[MEASURED]** (commit
|
||
`e2ca2fc`).
|
||
|
||
### 6.2 The wave queues were starved 1-push-vs-4-pops
|
||
|
||
`predict()` stored **one** wave per tick while `onResult()` popped one per
|
||
resolved bullet (~4/tick), so the queue drained within a few dozen ticks and
|
||
~3 of every 4 resolutions returned without learning; the survivor paired with
|
||
a same-tick wave (`bearingDelta ≈ 0`), pinning the histogram at centre.
|
||
**Proof:** `GF.vHits == HeadOn.vHits` and `DecayGF.vHits == HeadOn.vHits`
|
||
byte-for-byte in **every one of 50 rounds** — GF, DecayGF and KNN had
|
||
degenerated into HeadOn clones. Fix: per-bin FIFO with an O(1) head cursor, at
|
||
most one push per (tick, bin). Also `maxBullets` 2,048→8,192: the rack spawns
|
||
52 bullets/tick so the ring wrapped every ~39 ticks while a long power-3 shot
|
||
needs ~90, silently discarding unresolved bullets and biasing every measured
|
||
hit rate by range; a `droppedBullets` counter was added. After the fix
|
||
`vDropped = 0` and `vStarved = 0` across all 48 recorded rounds. **[MEASURED]**
|
||
(commit `0cc6821`).
|
||
|
||
### 6.3 Tsetlin's clauses saturated at ~714 included literals each
|
||
|
||
`Tsetlin.vHits` was byte-for-byte equal to `Linear.vHits` in every measured
|
||
round of every run because its learned correction was always exactly 0.
|
||
Root cause: `tmLearnOne` rewarded included true literals unconditionally,
|
||
omitting Granmo's `(c=0, lk=1) → toward Exclude` counter-force, so true
|
||
literals ratcheted toward Include forever; Type II was unreachable dead code
|
||
with the wrong direction; resource allocation was an `|error|` heuristic
|
||
instead of Granmo's `(T − clip(v,−T,T))/(2T)`; the label baseline had a
|
||
factor-2 shrink (`error = δ − 2c`, fixed point `c = δ/2`); hits zeroed their
|
||
residual; the enemy-energy feature was duplicated (`state.selfEnergy` fed
|
||
where `WorldState.enemyEnergy` exists, so energy rules were literally
|
||
unrepresentable); and `tmEvalClause` needed Granmo Eq. 6 (all-Exclude clause
|
||
outputs 1 during learning, 0 during classification) or fix #1 deadlocks every
|
||
clause at empty.
|
||
|
||
Measured effect (energy-threshold-turner, seed 1): mean included
|
||
literals/clause **714.0 → 13.8**; active clauses 100/100 → 53/100; nonzero
|
||
corrections 8/764 → 708/764; Tsetlin virtual hits **27/400 → 69/400** (Linear
|
||
43/400). Tsetlin now **learns** but is **not yet competitive with Linear** —
|
||
the regression head is untuned, flagged as follow-up rather than claimed as a
|
||
win. **[MEASURED]** (commit `8937000`; `/tmp/ab_logs3/final_test_tsetlin_gun.log`).
|
||
|
||
### 6.4 The TM classifier gun did not earn its slot (but its clauses are real)
|
||
|
||
A Tsetlin-Machine mixture-of-experts gate over
|
||
HeadOn/Linear/Circular/WallBounce/Accel was built with the corrected feedback
|
||
and labelled by which expert's prediction was closest to the actual enemy
|
||
position (an exact, supervised, per-shot label — no delayed credit). It loses
|
||
to the best of its own experts offline on nearly every fixture, and against
|
||
DrussGT it cost real performance:
|
||
|
||
```
|
||
baseline (path + relative) 7.56% real hit rate, 157 dmg
|
||
+ power fix 7.47%, 239 dmg
|
||
+ power fix + TM selector 5.59%, 133 dmg
|
||
```
|
||
|
||
It was selected on 806 ticks and fired 24 real shots at 4.2%. It ships disabled
|
||
(`EnableTmSelector = false`; code and wiring kept intact). **However, the gate
|
||
latched onto meaningful structure:** on the energy-threshold turner, HeadOn's
|
||
clauses key on the **energy bits** (the rule's own driving variable) while
|
||
Circular keys on distance/velocity. So the TM learned something real and
|
||
interpretable; it simply could not beat "always pick the best expert".
|
||
**[INFERRED]** root cause: the closest-expert label is noisy because several
|
||
experts are near-tied, and under the path metric the winner varies by power bin
|
||
while the gate sees one shared per-tick input, so a one-vs-rest gate over a
|
||
saturated 870-bit clause space has no margin to exploit. A standalone Granmo
|
||
classifier on the same encoding reaches ~99% on the rule but, per the
|
||
counterfactual probe, does **not** read energy (follow rate 24% high / 62% mean
|
||
— statistically identical at 1, 2 and 10 frames), so even the "it learned the
|
||
rule" claim is limited to ~99% accuracy, not to a readable energy threshold.
|
||
The best recovered proposition was `!g9 ∧ !g8` (energy < 25.6, not the labelled
|
||
30) — a genuine simple threshold, but not the ensemble's decision mechanism.
|
||
**[MEASURED]** (commits `57b2ac3`, `d5061ee`; `test_tm_pattern_learning.nim`).
|
||
|
||
### 6.5 The selector thresholds were absolute on a rescaled metric
|
||
|
||
Covered in §2.2: the `0.10` absolute floor fired on 53.0% of point-metric ticks
|
||
and forced HeadOn (real 2.0–4.4%, 11th of 13); HeadOn selection share fell
|
||
69.1% → 43.5% under relative thresholds (and 23.1% → 24.2% under path). Also:
|
||
`bestGun` was first-index-wins argmax, so HeadOn at index 0 silently won every
|
||
tie until the random tie-break landed (`343e631`); `bestPower` had the same
|
||
absolute-40% defect (below).
|
||
|
||
### 6.6 Power selection was stuck at power 1.0 (`MinHitRate = 0.40`)
|
||
|
||
`bestPower` used an **absolute** `MinHitRate = 0.40` bar. Measured per-bin
|
||
virtual rates show no bin ever clears 40%, so 11 of 14 guns were stuck at bin 0
|
||
(power 1.0) even where higher bins were comparable:
|
||
|
||
```
|
||
Linear p1.0 44% p1.5 39% p2.0 30% p3.0 29% old bin 0 -> new bin 3
|
||
Accel p1.0 44% p1.5 40% p2.0 26% p3.0 29% old bin 1 -> new bin 3
|
||
Pattern p1.0 50% p1.5 40% p2.0 27% p3.0 12% old bin 1 -> new bin 2
|
||
```
|
||
|
||
Replaced with a scale-aware `PowerBarFrac = 0.50` (a dimensionless fraction of
|
||
the gun's own best-bin rate); 13 of 14 selections now pick heavier bullets.
|
||
Real effect vs DrussGT (8 rounds × 3 runs): hit rate unchanged (7.56% → 7.47%),
|
||
**damage +52% (157 → 239 per run)** and rounds end faster. **[MEASURED]**
|
||
(commit `57b2ac3`).
|
||
|
||
Smaller fixes in the same family: `bestPower` on a cold gun returned the
|
||
*highest* bin (empty bin satisfied the `count == 0` clause); `fitnessFor`
|
||
aggregated enemies in nondeterministic hash order; `stop_shot` had an
|
||
unreachable deceleration branch and several guns had tick-only caches that made
|
||
all four power bins return bin 0's lead (`e536900`). **[MEASURED]**.
|
||
|
||
---
|
||
|
||
## 7. Known caveats and open problems
|
||
|
||
Stated without hedging.
|
||
|
||
1. **The headline per-gun numbers come from ONE adversary, a wave surfer.**
|
||
HeadOn is genuinely bad against surfers, so part of the rack ordering may be
|
||
matchup-specific. A SpinBot guard was inconclusive: ModularBot fires only
|
||
17–31 real shots/run against a fast bot because the range-aware firing gate
|
||
is strict at long range, so the guard had little power (Wilson looked better,
|
||
18.5% vs 8.6%, but on 70–92 shots with a 5–33% spread). **[MEASURED]**
|
||
(commit `2c94dc2`). A second, independent adversary at scale is missing.
|
||
|
||
2. **Per-gun real N is small.** 47–898 shots per gun; n < 200 gives roughly
|
||
±5 pp across a 3–15% spread. Single-gun ordering is **indicative, not
|
||
definitive**. The KEEP/MARGINAL/BELOW boundaries should be treated as soft.
|
||
|
||
3. **The fixtures are perfect-information and therefore optimistic.** Every
|
||
fixture is an observer capture with true positions every tick; the classic
|
||
set is additionally **open-loop** (replayed DrussGT never dodges our
|
||
bullets). Absolute offline hit rates are inflated by an unknown amount; only
|
||
relative comparisons are safe.
|
||
|
||
4. **The virtual metric is anti-correlated with real hit rate and no tested
|
||
ranking rule fixed it.** Shipped config Spearman ≈ **−0.374** over 13 runs
|
||
(the sign flips to +0.52 on the smaller 5-run set, so it is unstable). 16
|
||
candidate ranking rules all overlapped the shipped config. The selector's
|
||
value lives in its **floor/tie hedging** (5.08% without the floor vs 6.95%
|
||
with it), not in its ranking. **[MEASURED]**.
|
||
|
||
5. **The TM classifier gun did not earn its slot.** It cost real performance
|
||
(7.47% → 5.59%, 133 dmg) despite showing interpretable energy structure in
|
||
its clauses (§6.4). It is disabled; re-enabling requires a fix to the gate
|
||
margin/label problem, not more training.
|
||
|
||
6. **Real-hit-rate-driven selection is not viable yet.** Only the selected gun
|
||
fires, so unselected guns get near-zero real shots (GuessFactor 20, Linear
|
||
24 vs HeadOn 733 in the point A/B); noise is fatal (n = 470 at p = 10% gives
|
||
±2.8 pp, most guns n < 200 gives ±5 pp+); and real rate is conditional on
|
||
when the gun was selected. A blended signal with forced exploration and
|
||
shrinkage is defensible in principle but needs thousands of shots per gun
|
||
across many battles. Real rate is currently best used **offline** as the
|
||
evaluation metric — which is exactly what the A/B does. **[MEASURED]**
|
||
(commit `dea4dcb`).
|
||
|
||
7. **The offline==online acceptance test is flaky** (§1.1): typically 11/12 on
|
||
unmodified HEAD, with the mismatching gun varying run to run. The
|
||
equivalence claim is strong-but-not-exact until the boundary race is fixed.
|
||
|
||
8. **The selector's random tie-break is not randomised in the live bot.** The
|
||
shipped bot never calls `randomize()`, so the "random" sequence is fixed
|
||
across process restarts (a side finding of `2c94dc2`, not fixed).
|
||
|
||
9. **The firing gate is not the bottleneck.** The shipped range-aware gate does
|
||
not beat a fixed 2.0° gate on hit rate (55.8% vs 57.9%, ~1.5 σ), though it
|
||
fires 22–28% more shots. No per-adversary score delta exceeded the
|
||
~300-point run-to-run noise band. **[MEASURED]** (commit `3c90a59`).
|
||
|
||
10. **The boss is ~2.3× more accurate than the whole rack** (12.1% vs 5.3% in
|
||
the capture). Closing that gap is the point of the rack; the current
|
||
best single gun is 10.7%.
|
||
|
||
---
|
||
|
||
## 8. Reproduction
|
||
|
||
Commands recorded in the commits and tool READMEs. (I was instructed not to
|
||
run builds/tests while writing this report; these are the documented
|
||
invocations, not a fresh verification by me.)
|
||
|
||
```bash
|
||
# Offline gun range over all 20 fixtures, shipped path metric (default):
|
||
nim c -r common_libs/tests/run_range.nim
|
||
# Point metric for comparison:
|
||
GUN_VBULLET_METRIC=point nim c -r common_libs/tests/run_range.nim
|
||
# Add timing:
|
||
nim c -r common_libs/tests/run_range.nim --timing
|
||
|
||
# Selector diagnostics (floor/tie/bestRate/HeadOn-share) on a fixture:
|
||
GUN_SELECTOR_MODE=relative nim c -d:release -r \
|
||
common_libs/tests/analyze_selector.nim tools/fixtures/drussgt_vs_spinbot.jsonl
|
||
|
||
# Offline == online acceptance (currently flaky):
|
||
nim c -r common_libs/tests/acceptance_offline_vs_online.nim
|
||
|
||
# Tsetlin gun clause sparsity / divergence:
|
||
nim c -r common_libs/tests/test_tsetlin_gun.nim
|
||
# TM readability (standalone Granmo classifier on the energy-threshold rule):
|
||
nim c -r common_libs/tests/test_tm_pattern_learning.nim
|
||
|
||
# Live boss (real DrussGT jar; jars stay out of git, see the README):
|
||
# tools/robocode_shim/run_bridge_battle.sh <bot_dir> <rounds> <capture.jsonl>
|
||
# Closed-loop evidence for the TR captures:
|
||
python3 tools/robocode_shim/analyze_closed_loop.py \
|
||
tools/fixtures/tr_drussgt_vs_modularbot.jsonl \
|
||
tools/robocode_shim/evidence/tr_drussgt_vs_modularbot.events.json
|
||
```
|
||
|
||
Per-gun aggregation scripts used for the tables above:
|
||
`python3 /tmp/agg2.py base` (virtual-vs-real + Spearman),
|
||
`python3 /tmp/compare.py` (server-side per-run A/B + overlap).
|
||
Raw range output: `/tmp/range_path.txt` (path), `/tmp/final_range.txt` (point).
|