0ede6d12ec
Hypothesis under test (from the gun audit, which named the tie-band as "the lever that matters most"): `bmPath` is deliberately generous (2.3-3.6x `bmPoint`), so a gun can sit in the tied band on a ray that sweeps the target's path while its bullets ARRIVE badly. So: keep the `path`-ranked band (path beat point on real hit rate 7.43% vs 4.70%, z=5.56), but narrow the random draw inside it using a parallel `point` (arrival-accuracy) window. RESULT: NO EFFECT. Real DrussGT, ONE frozen binary (/tmp/ModularBot_tieband, md5 2c0c56e6...), env knobs only, 7 runs x 7 rounds per arm, server-side events sidecar, exact two-sided permutation test on per-run rates. arm runs shots real % dmg/run d p tbbase (shipped) 7 4128 7.17 175 -- -- tbpt path-rank + point-narrow 7 3938 7.08 165 +0.14 0.88 tbpc =commit control 7 3759 4.44 98 +2.74 0.0012 tbpt25 point margin 0.25 7 3683 5.59 119 +1.65 0.20 tbtie05 / tbtie40 (band width) 7 3937/3917 5.84/6.28 133/144 1.49/1.00 0.11/0.25 tbwin50 (SelectorWindow=50) 7 3983 6.05 139 +1.20 0.11 tbfloor10 (FloorPeakFrac=0.10) 7 3829 5.33 118 +2.12 0.11 tbpt vs base: fully overlapping ranges, p=0.88. This is a REAL null, not a dead arm - the mechanism was live, and it visibly changed the selected-gun mix (Pattern 24%->16%, Accel 6%->16%, Tsetlin ~0%->13%). CONTROL VALIDATED, AND THIS IS THE THIRD TIME: removing the random draw inside the band is SIGNIFICANTLY WORSE (4.44%, p=0.0012). Combined with the earlier hysteresis A/B (7.02% -> 5.10% for commitment) and the light-hysteresis result, the selector's per-tick randomness is now load-bearing on three independent measurements. Narrowing the band on ANY second virtual statistic has not helped. Every knob swept (band width, floor, window) is nominally worse than shipped at n=7; that is "no credible win" rather than "proven harm" (sd ~1.8pp, ~1pp resolution, underpowered). Shipped default stays `GUN_SELECTOR_TIEBREAK=off`; the feature is opt-in, fully guarded, and costs zero extra work on the default path (point windows are scored only when the mode is on). Guards: test_selector_tiebreak 19 (new, pure), test_gun_harness 39, test_vbullet_metric 11, test_adaptive_radar 41, test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28, test_rack_membership 38, acceptance_offline_vs_online 12/12 PASS (offline path calls neither chooseFromFit nor the tie-break). STRATEGIC CONCLUSION: three selection-side attempts have now failed (hysteresis, commitment, point tie-break). The selector is at a local optimum and the remaining lever is the QUALITY OF THE GUNS, not the selection among them.
882 lines
50 KiB
Markdown
882 lines
50 KiB
Markdown
# Gun Rack Analysis — ModularBot guns vs the real DrussGT boss
|
||
|
||
**Date:** 2026-09-21
|
||
**Bot:** ModularBot (13 active guns + 1 disabled TM classifier gun)
|
||
**Shipped config:** virtual-bullet metric `path`, selector thresholds `relative`
|
||
(`GUN_VBULLET_METRIC` default `path`, `GUN_SELECTOR_MODE` default `relative`)
|
||
**Rewritten:** this file supersedes the 2026-09-20 version, whose numbers
|
||
predated per-gun real-hit attribution, the offline gun range, and the live
|
||
DrussGT boss. No number from the old version is retained unless it was
|
||
re-measured below.
|
||
|
||
**Corrected after later runs (same day).** Three claims in the first version of
|
||
this rewrite were revised by follow-up measurements: the virtual-vs-real
|
||
correlation is a poor, sign-unstable ranker rather than an "inversion" (§2);
|
||
the offline==online acceptance test is fixed, not flaky (§1.1); and pruning the
|
||
below-overall guns was tested and does not help (§3). Each correction is
|
||
restated plainly at the point of the old claim.
|
||
|
||
Evidence tags used throughout:
|
||
|
||
- **[MEASURED]** — I read it from a recorded artifact (commit message, `/tmp`
|
||
output, a log, or a source file). The provenance is named every time.
|
||
- **[INFERRED]** — reasoning from measured facts; explicitly not measured.
|
||
|
||
A note on provenance: the measurements below were produced overnight in
|
||
commits `343e631` … `2c94dc2` on branch `research/lead-targeting`. Each commit
|
||
message records the numbers, the hypotheses it refuted, and its caveats. Raw
|
||
artifacts are in `/tmp` (`/tmp/gun_stats_base_r*.jsonl`,
|
||
`/tmp/events_base_r*.json`, `/tmp/range_path.txt`,
|
||
`/tmp/ab_logs/FINAL_ANALYSIS.txt`, `/tmp/ab_logs3/FINAL_AB.txt`,
|
||
`/tmp/ab_logs3/selector_diag.txt`). Where a commit message and a raw artifact
|
||
disagree slightly, both are stated.
|
||
|
||
---
|
||
|
||
## 1. The test infrastructure
|
||
|
||
The numbers are only as good as the rig that produced them, so the rig is
|
||
described first.
|
||
|
||
### 1.1 The offline gun range
|
||
|
||
`common_libs/gun_harness/offline_range.nim` replays a recorded
|
||
`seq[WorldState]` through the **same** `VirtualTracker`
|
||
(`common_libs/gun_harness/virtual_bullets.nim`) that the live bot drives. The
|
||
claim is not "an approximation of the live metric", it is "the same metric":
|
||
virtual-bullet fitness is already a pure function of (a stream of
|
||
`WorldState`, a list of guns), and the Java battle only supplies where the
|
||
states come from. Guns keep their own internal history, so a replayed stream in
|
||
order is a complete movement history. **[MEASURED]** (`offline_range.nim`,
|
||
header comment; commit `974528d`).
|
||
|
||
- **Speed and sample count.** 8 fixtures / 1,770 ticks / ~92 k virtual bullets /
|
||
13 guns replay in 2.9 s at ~32 k virtual bullets/s — roughly **70× faster and
|
||
100× more samples** than a live gauntlet. **[MEASURED]** (commit `974528d`).
|
||
- **Acceptance test.** `common_libs/tests/acceptance_offline_vs_online.nim`
|
||
records one live round, replays it offline, and compares per-gun virtual hit
|
||
counts. It passed **12/12 deterministic guns** (Tsetlin is separately labelled
|
||
stochastic because `tmLearnOne` calls `rand()`). Two real live-loop ordering
|
||
quirks had to be modelled to reach that: `run()` calls `go()` before the
|
||
aim/fire block, so `tickBullets` resolves against the *next* tick's scan; and
|
||
if the target dies during that `go()`, the final tick's spawn+resolution is
|
||
skipped. **[MEASURED]** (commit `974528d`; a passing run is preserved in
|
||
`/tmp/ab_logs3/test_acceptance.log`: 12/12, 534-tick round).
|
||
- **Acceptance test: the equivalence is now proven and stable.** An earlier
|
||
version of this report recorded the test as flaky — typically **11/12** on
|
||
unmodified HEAD, with the mismatching gun moving between runs (KNN, then
|
||
WallBounce) — and guessed it was a live/offline boundary race. That guess was
|
||
**wrong**. The root cause was a **real replay bug**: the offline replay
|
||
spawned gun 13 (TMSelect) while the live rack has `EnableTmSelector = false`
|
||
and never does. The shared `VirtualTracker` ring is **order-sensitive**, so
|
||
gun 13's extra 4 bullets/tick permuted the per-tick **resolution order** of
|
||
every other gun, shifting the learning guns' observations. Closing gun 13's
|
||
ready gate offline made the live and offline KNN traces **byte-identical**
|
||
(904/904 lines, empty diff). The fix mirrors the live rack in the replay — no
|
||
tick exclusion, no tolerance loosening. Stability: **5/5 consecutive runs
|
||
report 12/12 exact, each with the death boundary included.** So
|
||
"offline == online" is exact on these runs. A flaky proof had hidden a real
|
||
bug. **[MEASURED]**.
|
||
- **General lesson for rack A/B.** Because the ring is order-sensitive, any
|
||
rack A/B that disables a gun also removes that gun's **4 spawns/tick** from
|
||
the shared ring, which perturbs the resolution order — and therefore the
|
||
learning observations — of every other gun. That is a confound to record for
|
||
anyone repeating these experiments. **[INFERRED]**.
|
||
|
||
### 1.2 The fixture sets and what each is good for
|
||
|
||
There are **20 JSONL fixtures** under `tools/fixtures/`. They fall into four
|
||
groups. **[MEASURED]** (`tools/fixtures/`, `DRUSSGT_FIXTURES.md`,
|
||
`TR_BRIDGE_FIXTURES.md`).
|
||
|
||
**(a) 8 synthetic fixtures — ground truth by construction.**
|
||
`common_libs/gun_harness/offline_range.nim` generates them; each has a known
|
||
rule:
|
||
|
||
| Fixture | Generator / rule | What it is good for |
|
||
|---|---|---|
|
||
| `stationary` | fixed enemy | ceiling: any correct gun scores ~100% |
|
||
| `constant-velocity` | straight line, no walls | does the gun lead a moving target at all |
|
||
| `circular` | constant turn 3°/tick | circular/accel models |
|
||
| `wall-bounce` | specular reflection off all 4 walls | wall-aware prediction |
|
||
| `oscillator` | east 30 ticks, west 30 ticks | phase/timing |
|
||
| `random-walk` | seeded ±15°/tick jitter | generality under noise |
|
||
| `decel-before-turn` | cruise → full stop → pivot 3×45° → accelerate | stop-shot detection |
|
||
| `energy-threshold-turner` | **KNOWN RULE**: straight while `energy ≥ 30`, hard 20°/tick turn while `energy < 30`, `energy = max(5, 50 − 0.5·t)` | can a learner find a readable high-level rule; threshold crosses at t=41 |
|
||
|
||
The energy-threshold turner is the falsifiable one: the label is literally a
|
||
predicate over the 11-bit Gray-coded energy field, so a learner that reads
|
||
energy can be shown to have read the *right* variable (see §6.3).
|
||
|
||
**(b) 2 classic-Robocode contrast fixtures.** `contrast_stationary_sittingduck`
|
||
and `contrast_straightline` — trivial motion, used as a sanity/ceiling check.
|
||
**[MEASURED]**.
|
||
|
||
**(c) 5 classic-Robocode DrussGT captures.** Real, unmodified DrussGT
|
||
3.1.4159 movement captured from Robocode 1.9.5.5 via
|
||
`tools/robocode_fixture_capture/` (file `DRUSSGT_FIXTURES.md`). Opponents:
|
||
SpinBot, RamFire, Crazy, Corners, and a DrussGT mirror; 28,797 ticks. The
|
||
coordinate conversion is validated to **0.000–0.001°** on every fixture by
|
||
recomputing the direction implied by (heading, speed) against the recorded
|
||
per-tick displacement. **These are OPEN-LOOP and perfect-information:** the
|
||
replayed DrussGT never dodges *our* bullets, and the observer reports true
|
||
positions every tick (unlike the live bot's stale between-scan `WorldState`).
|
||
They are therefore optimistic and good for *relative* gun ranking, not absolute
|
||
hit rates. **[MEASURED]** (`DRUSSGT_FIXTURES.md`).
|
||
|
||
**(d) 5 closed-loop Tank Royale DrussGT captures.** The real DrussGT jar playing
|
||
Tank Royale through `tools/robocode_shim/`, captured by
|
||
`tools/robocode_shim/src/robocode_shim/TrBattleCapture.java` (file
|
||
`TR_BRIDGE_FIXTURES.md`). Primary file `tr_drussgt_vs_modularbot.jsonl`:
|
||
15 rounds, 20,026 ticks, ModularBot fired 1,134 shots. At capture time DrussGT
|
||
was reacting to *our* real bullets. **Closed-loop proven, not asserted:**
|
||
`tools/robocode_shim/analyze_closed_loop.py` event-locks `|Δheading|` to
|
||
ModularBot's fire times (a heat-limited near-metronome, median interval
|
||
14 ticks) and gets an oscillating response with the fire period; the
|
||
cross-correlation peaks at **r = +0.111, lag 12, permutation p = 0.005** (null
|
||
peak mean +0.016), and the **own-fire control is flat**, so the lock is
|
||
enemy-driven, not internal cadence. **Still perfect-information**, and
|
||
open-loop *at replay time* — "closed_loop" describes the capture, not a later
|
||
replay. TR angle conversion residual is ~1.5° mean (vs 0.000° classic) because
|
||
the TR server moves along the pre-turn heading. **[MEASURED]**
|
||
(`TR_BRIDGE_FIXTURES.md`).
|
||
|
||
| Fixture | Source | rounds | ticks | adversary / note |
|
||
|---|---|---:|---:|---|
|
||
| `circular` … `energy-threshold-turner` | synthetic | — | 150–260 | 8 known-rule trajectories |
|
||
| `contrast_stationary_sittingduck`, `contrast_straightline` | classic-robocode | 1 each | 1,451 | sanity contrasts |
|
||
| `drussgt_vs_spinbot` | classic-robocode | 20 | 5,002 | open-loop, perfect-info |
|
||
| `drussgt_vs_ramfire` | classic-robocode | 20 | 3,240 | open-loop, perfect-info |
|
||
| `drussgt_vs_crazy` | classic-robocode | 20 | 9,025 | open-loop, perfect-info |
|
||
| `drussgt_vs_corners` | classic-robocode | 20 | 4,975 | open-loop, perfect-info |
|
||
| `drussgt_vs_drussgt` | classic-robocode | 2 | 6,555 | open-loop, perfect-info (mirror) |
|
||
| `tr_drussgt_vs_modularbot` | tr-bridge | 15 | 20,026 | **closed-loop at capture**, perfect-info |
|
||
| `tr_drussgt_vs_modularbot_shield` | tr-bridge | 10 | 12,629 | shield on |
|
||
| `tr_drussgt_vs_spinbot` / `_crazy` / `_corners` | tr-bridge | 10 each | 10,824 / 11,507 / 2,575 | closed-loop at capture |
|
||
|
||
### 1.3 The live boss — the real DrussGT jar
|
||
|
||
`tools/robocode_shim/` runs the **unmodified** `DrussGT.jar` (159,289 bytes,
|
||
md5 `5cd6015dcc6d6da8a7e6aeecb1fec211`) as a Tank Royale bot. The classic
|
||
`robocode.*` API is a thin delegation layer over the public
|
||
`IBasicRobotPeer`/`IAdvancedRobotPeer` seam, so the shim reuses the genuine
|
||
`robocode.jar` and implements only the 75-method peer interface
|
||
(`ClassicPeer`), plus `BotHost`/`ThreadManagerFix`. DrussGT compiles with zero
|
||
shim API symbols and runs real battles; the EnergyDome shield is disabled by
|
||
default (pure wave surfer). Known physics divergences (move/turn ordering,
|
||
distance bookkeeping, etc.) are enumerated in section 5.9 of
|
||
`tools/robocode_shim/README.md`. **[MEASURED]**.
|
||
|
||
The boss is far stronger than us: in the capture battle DrussGT beat ModularBot
|
||
**1447–300 over 15 rounds** (ModularBot won round 5 only), firing 1,400 bullets
|
||
at a **12.1%** hit rate against ModularBot's 1,134 bullets at **5.3%**. That is
|
||
the number the whole gun rack is trying to move. **[MEASURED]** (commit
|
||
`17c99f5`, `TR_BRIDGE_FIXTURES.md`).
|
||
|
||
### 1.4 A/B methodology — server-side per-run hit rate, never scores
|
||
|
||
The A/B that decides configs uses **server-side ground truth**, not the bot's
|
||
own counters and not scores. **[MEASURED]** (`/tmp/analyze.py`,
|
||
`/tmp/compare.py`, `/tmp/ab_logs3/FINAL_AB.txt`):
|
||
|
||
1. The battle runner writes a per-shot **events sidecar** (`/tmp/events_*.json`)
|
||
with `fire` / `hit` / damage events stamped with the server's per-round
|
||
bullet id (`GunEngine.nextBulletId`). `26b66cb` proved per-gun attribution:
|
||
the server assigns the id once and reuses it on `BulletFired`,
|
||
`BulletHitBot`, `BulletHitWall`, `BulletHitBullet`; hits that arrive before
|
||
the fire event (client priority 70 > 60) are deferred. 99.9% of shots and
|
||
99.8% of hits were attributed in that session.
|
||
2. A config is judged on its **per-run real hit rate** (`hits/shots` for one
|
||
battle), and two configs are called different only if their per-run ranges
|
||
**do not overlap**. This is why the overnight A/B reports "SEPARATED" or
|
||
"OVERLAP" for every pair rather than a single pooled p-value.
|
||
3. **Scores are not used to judge configs.** Single-run scores swing by a
|
||
couple of hundred points: the 13 shipped-config runs span **175–526**
|
||
(s.d. ≈ 105, range 351; `/tmp/battle_base_r*.log`), so a ~210-point
|
||
2-s.d. band swamps any plausible config effect. The gate study (`3c90a59`)
|
||
made the same call: no per-adversary score delta exceeded the ~300-point
|
||
run-to-run noise band.
|
||
|
||
The bot-side per-gun attribution (`realShots`/`realHits` in
|
||
`/tmp/gun_stats_base_r*.jsonl`) covers ~87% of the server's total shots
|
||
uniformly (13 runs: 220/3,157 = 6.97% attributed vs 251/3,612 = 6.95%
|
||
server-side), so it is used for per-gun *ranking* but the server sidecar is the
|
||
ground truth for config decisions. **[MEASURED]** (re-aggregated from
|
||
`/tmp/events_base_r*.json` and `/tmp/gun_stats_base_r*.jsonl`).
|
||
|
||
---
|
||
|
||
## 2. The metric lesson: virtual hit rate is a poor ranker, not a proxy for real hit rate
|
||
|
||
This is the most important conceptual result of the night and it invalidates a
|
||
naive reading of every offline table in this report.
|
||
|
||
**[MEASURED]** Aggregating the 13 shipped-config runs against the live DrussGT
|
||
boss (`/tmp/gun_stats_base_r{1..13}.jsonl`; `python3 /tmp/agg2.py base`), the
|
||
Spearman rank correlation between a gun's **virtual** hit rate and its **real**
|
||
hit rate is
|
||
|
||
```
|
||
Spearman(virtual rank, real rank) = -0.374 (n = 13 guns, all with ≥10 real shots)
|
||
```
|
||
|
||
The correlation is **weak and sign-unstable** — the honest headline is that
|
||
virtual hit rate is a **poor ranker**, not a reliable inverse. The table below
|
||
is what a poor ranker looks like: on this run set the ordering it produces
|
||
tracks the opposite of the real ordering, but that does not hold on other run
|
||
sets (see the full measurement set after the table). An earlier version of this
|
||
report stated this as a clean inversion; a later 15-run measurement refuted
|
||
that.
|
||
|
||
| Gun | Selected (ticks) | Real hits/shots | Real % | Virtual % |
|
||
|---|---:|---:|---:|---:|
|
||
| Linear | 2,707 | 9/84 | **10.7** | 10.2 |
|
||
| Circular | 4,051 | 17/172 | **9.9** | 11.9 |
|
||
| KNN | 4,886 | 11/122 | **9.0** | 7.5 |
|
||
| Pattern | 14,786 | 50/582 | **8.6** | 12.0 |
|
||
| Accel | 9,205 | 28/382 | **7.3** | 12.1 |
|
||
| AvgLead | 5,119 | 14/200 | **7.0** | 12.3 |
|
||
| GuessFactor | 2,345 | 5/72 | **6.9** | 10.3 |
|
||
| DecayGF | 1,360 | 3/47 | **6.4** | 9.2 |
|
||
| WallBounce | 7,618 | 18/288 | **6.2** | 12.9 |
|
||
| StopShot | 3,348 | 8/132 | **6.1** | 12.6 |
|
||
| Tsetlin | 3,284 | 6/103 | **5.8** | 12.9 |
|
||
| Displace | 2,664 | 4/75 | **5.3** | 12.3 |
|
||
| HeadOn | 16,975 | 47/898 | **5.2** | 8.6 |
|
||
|
||
Read the top and bottom: **Tsetlin, WallBounce and StopShot have the highest
|
||
virtual rates (12.6–12.9%) and near-bottom real rates (5.8–6.2%); Linear and
|
||
KNN sit at 10.2% / 7.5% virtual but 10.7% / 9.0% real.** On this run set the
|
||
virtual ordering inverts the real one — but because the sign flips on other
|
||
run sets (next paragraph), the safe reading is that the virtual ranking is
|
||
**uninformative about the real ranking**, not that it is reliably inverted.
|
||
|
||
**Why this matters for selection.** What has kept the rack alive is the
|
||
selector's **floor/tie hedging**, not its ranking: removing the floor
|
||
(`GUN_SELECTOR_FLOOR=0.0`, config `floor00`) drops the rack from 6.95% to
|
||
**5.08%** at 175 dmg/run (vs 251) over 788 shots. **[MEASURED]**
|
||
(`/tmp/compare.py`). So the selector is useful because it refuses to commit to
|
||
a bad field, not because its virtual-rate ordering is good.
|
||
|
||
**The headline: across run sets the correlation is sign-unstable, so it is near
|
||
zero on average — not robustly negative.** The full set of independent
|
||
Spearman measurements (13-run source `/tmp/agg2.py base`; 5-run configs
|
||
`/tmp/ab_logs3/FINAL_AB.txt`; 15-run paired baseline §3) is:
|
||
|
||
| Run set / aggregation | Spearman |
|
||
|---|---:|
|
||
| 13-run base, shipped `relative+path` | **−0.374** |
|
||
| 15-run paired baseline (different but equally defensible aggregation) | **+0.335** |
|
||
| 5-run 12-round A/B, `relative+path` | +0.522 |
|
||
| 5-run A/B, `absolute+point` | −0.371 |
|
||
| 5-run A/B, `absolute+path` | −0.073 |
|
||
| 5-run A/B, `relative+point` | −0.037 |
|
||
|
||
Two **opposite signs on large samples** (−0.374 over 13 runs, +0.335 over 15
|
||
runs, all over the same 13 guns) mean virtual hit rate is **not** a reliable
|
||
inverse of real hit rate. It is a **poor ranker**: weak correlation, sign
|
||
flipping between run sets, near zero on average. The practical conclusion is
|
||
unchanged — do not build a ranking rule on it — but an **earlier version of
|
||
this report overstated the mechanism as an inversion**; the later 15-run
|
||
measurement refuted that. **[MEASURED]** + **[INFERRED]** (the six numbers are
|
||
measured; "poor ranker / near zero on average" is the reasoning).
|
||
|
||
### 2.1 The metric A/B: point vs path (this one is real, and it is selection)
|
||
|
||
`GUN_VBULLET_METRIC` picks how a virtual bullet is scored
|
||
(`common_libs/gun_harness/virtual_bullets.nim`):
|
||
|
||
- `bmPoint` — resolve at the fire-time aim distance and score that single
|
||
point. Measures prediction accuracy.
|
||
- `bmPath` (**shipped**) — fly the ray to the wall and test each swept segment
|
||
against the target radius. Measures hypothetical hit chance.
|
||
|
||
**[MEASURED]** Live A/B against the boss, 5 battles × 12 rounds, one frozen
|
||
binary (commit `3b5d70b`; per-run detail in `/tmp/ab_logs/FINAL_ANALYSIS.txt`):
|
||
|
||
| Metric | Shots | Hits | Real hit rate | Per-run rates | Spearman |
|
||
|---|---:|---:|---:|---|---:|
|
||
| point | 4,660 | 219 | 4.70% | 5.53 / 5.30 / 4.92 / 3.16 / 4.57 | −0.04 |
|
||
| path | 4,834 | 359 | **7.43%** | 6.76 / 8.20 / 8.24 / 6.55 / 7.30 | +0.52 |
|
||
|
||
The distributions **do not overlap**: path's worst run (6.55%) beats point's
|
||
best (5.53%). +2.73 pp, +58% relative, z = 5.56, p < 0.0001. Range
|
||
distributions were identical (~460–478 px), so this is not a range confound.
|
||
|
||
**The gain is selection, not better gun learning.** Under `point` every gun's
|
||
virtual rate is compressed into 0.6–4.4%, so HeadOn sits inside the 2 pp tie
|
||
margin and takes **72.6% of selection ticks / 76.9% of shots** while ranking
|
||
11th of 13 by real hit rate (2.3%). Under `path` the band widens to 4.7–13.7%
|
||
and HeadOn's shot share falls to 35.9%, so Pattern/Accel/WallBounce get picked.
|
||
The counterfactual confirms it: applying the point model's per-gun real rates
|
||
to the path model's shot mix yields 7.65%, i.e. essentially the whole observed
|
||
gain. **[MEASURED]** (commit `3b5d70b`).
|
||
|
||
Offline range total moves the same way: **34.3%** under point
|
||
(`/tmp/final_range.txt`, 35,636/104,000) vs **50.8%** under path
|
||
(`/tmp/range_path.txt`, 52,770/103,938). The offline totals are inflated by the
|
||
perfect-information synthetic fixtures (three of them score 100% under path for
|
||
every gun), so the offline totals are *not* comparable to live rates — only to
|
||
each other.
|
||
|
||
### 2.2 The selector-threshold A/B: absolute vs relative
|
||
|
||
The legacy thresholds were calibrated for a rate scale that does not exist.
|
||
**[MEASURED]** offline replay of a *fogged live* `WorldState` vs DrussGT
|
||
(1,397 selection ticks, `/tmp/ab_logs3/selector_diag.txt`):
|
||
|
||
| Config | Floor fires | HeadOn selection share | bestRate med |
|
||
|---|---:|---:|---:|
|
||
| absolute + point | 53.0% | 69.1% | 8.0% |
|
||
| relative + point | 21.2% | 43.5% | 5.25% |
|
||
| absolute + path | 3.0% | 23.1% | 24.0% |
|
||
| relative + path (shipped) | 8.4% | 24.2% | 16.75% |
|
||
|
||
The `0.10` absolute floor fires on **53.0%** of point-metric ticks and forces
|
||
HeadOn, whose real rate was 2.0–4.4%. (An earlier claim that the floor fires
|
||
*always* is **refuted**: it is 53%, because `bestRate` is a max over
|
||
gun×power-bin and an occasional ≥50-sample bin clears 10%.)
|
||
|
||
The scale-aware replacement (commit `dea4dcb`):
|
||
`RelTieMargin = 0.20` (tie band is a fraction of `bestRate`), `FloorPeakFrac =
|
||
0.25` (floor fires only if the field collapsed vs its own recent peak over a
|
||
256-tick window, counting only guns with ≥ MinObsBeforeCompete = 50 samples),
|
||
and pooled-over-bins ranking instead of max-over-bins.
|
||
|
||
Live A/B, 3 runs × 10 rounds (commit `dea4dcb`, `/tmp/ab_logs3/FINAL_AB.txt`):
|
||
|
||
| Config | Per-run rates | Pooled | vs `absolute+point` |
|
||
|---|---|---:|---|
|
||
| absolute + point | 3.66 / 2.45 / 5.01 | 3.76% | — |
|
||
| absolute + path | 7.55 / 8.21 / 6.83 | 7.57% | SEPARATED (p<0.0001) |
|
||
| relative + point | 7.66 / 6.18 / 5.79 | 6.59% | SEPARATED |
|
||
| relative + path (**shipped**) | 7.15 / 7.55 / 6.90 | 7.21% | SEPARATED |
|
||
|
||
`absolute+path` is nominally 0.35 pp above `relative+path`, but they **overlap**
|
||
(p = 0.64); so do `relative+point` and both path configs. The **metric** is the
|
||
dominant lever; under `path` the two threshold models are statistically tied.
|
||
`relative` was shipped because it is the principled scale-aware fix, works
|
||
under both metrics, and prevents the point-metric catastrophe if anyone
|
||
switches back. **[MEASURED]**.
|
||
|
||
---
|
||
|
||
## 3. Per-gun performance on real numbers
|
||
|
||
The final per-gun table, shipped config, **13 runs vs the live DrussGT boss,
|
||
3,612 server-side shots, 6.95% overall** (server sidecar; per-run rates 6.77 /
|
||
5.90 / 6.57 / 6.34 / 6.10 / 7.59 / 2.90 / 7.95 / 7.49 / 9.18 / 8.44 / 7.48 /
|
||
6.42%; 251 dmg/run). Per-gun rows are the bot-side attribution over the same
|
||
runs. The **same binary** on a different 15-run set gives **6.18%** (events
|
||
6.16%, 200 dmg/run, §3 pruning baseline), so every rate here is quoted with its
|
||
run count — a single figure is not definitive. **[MEASURED]** (commit `2c94dc2`;
|
||
`/tmp/gun_stats_base_r*.jsonl`, re-aggregated with `/tmp/agg2.py base`;
|
||
`/tmp/events_base_r*.json`).
|
||
|
||
| Verdict | Gun | Real hits/shots | Real % | Virtual % | Selected |
|
||
|---|---|---:|---:|---:|---:|
|
||
| **KEEP** | Linear | 9/84 | 10.7 | 10.2 | 2,707 |
|
||
| **KEEP** | Circular | 17/172 | 9.9 | 11.9 | 4,051 |
|
||
| **KEEP** | KNN | 11/122 | 9.0 | 7.5 | 4,886 |
|
||
| **KEEP** | Pattern | 50/582 | 8.6 | 12.0 | 14,786 |
|
||
| **KEEP** | Accel | 28/382 | 7.3 | 12.1 | 9,205 |
|
||
| **KEEP** | AvgLead | 14/200 | 7.0 | 12.3 | 5,119 |
|
||
| MARGINAL | GuessFactor | 5/72 | 6.9 | 10.3 | 2,345 |
|
||
| MARGINAL | DecayGF | 3/47 | 6.4 | 9.2 | 1,360 |
|
||
| MARGINAL | WallBounce | 18/288 | 6.2 | 12.9 | 7,618 |
|
||
| MARGINAL | StopShot | 8/132 | 6.1 | 12.6 | 3,348 |
|
||
| BELOW — KEEP | Tsetlin | 6/103 | 5.8 | 12.9 | 3,284 |
|
||
| BELOW — KEEP | Displace | 4/75 | 5.3 | 12.3 | 2,664 |
|
||
| **FLOOR — STAYS** | HeadOn | 47/898 | 5.2 | 8.6 | 16,975 |
|
||
|
||
**Verdicts.**
|
||
|
||
- **KEEP: Linear, Circular, KNN, Pattern, Accel, AvgLead.** These six are at or
|
||
above the 6.95% overall, yet their virtual rates are mid-pack to low: the
|
||
metric's three favourites (Tsetlin 12.9%, WallBounce 12.9%, StopShot 12.6%)
|
||
are near the *bottom* of the real ranking, while the real leader (Linear) sits
|
||
at 10.2% virtual. Further evidence the virtual ranking is uninformative about
|
||
the real ranking (and, on this run set, roughly its opposite).
|
||
- **MARGINAL: GuessFactor, DecayGF, WallBounce, StopShot.** Within ~1 pp of
|
||
overall on small N (47–288 shots). They are not obviously worth deleting, but
|
||
they have not earned a larger share.
|
||
- **BELOW OVERALL — but KEEP: Tsetlin, Displace.** Both sit below the 6.95%
|
||
overall on small N (75–103 shots), which an earlier version of this report
|
||
read as an implied recommendation to drop. That was **tested and refuted**:
|
||
15 **paired** runs per variant against DrussGT (identical seeds, 8 rounds,
|
||
same binary) gave baseline 3,238 shots / 6.18% (events 6.16%) / 200 dmg/run;
|
||
Tsetlin disabled 3,522 shots / 5.76% (events 5.71%) / 197 dmg/run; and
|
||
Tsetlin+Displace disabled 3,478 shots / 5.46% (events 5.37%) / 183 dmg/run.
|
||
Paired permutation tests: −0.34 pp (p = 0.57) and −0.70 pp (p = 0.21); the
|
||
per-run distributions completely overlap, and a Crazy (non-surfer) control
|
||
showed no separation either. Removing the measured-worst real performers is
|
||
therefore **neutral-to-slightly-negative** on both hit rate and damage. With
|
||
sd ≈ 1.8 pp a definitive claim would need far more runs, so **keep the full
|
||
rack** — being below overall does not justify removal. **[MEASURED]**.
|
||
- **HeadOn MUST STAY** despite being lowest (5.2%). It is the floor fallback:
|
||
when the field collapses the selector returns gun 0. Disabling the floor
|
||
measurably hurt — 5.08% / 175 dmg vs 6.95% / 251 dmg (config `floor00`,
|
||
788 shots). Do not delete HeadOn to improve the per-gun average; that
|
||
average is computed over shots it only gets because nothing better was
|
||
available. **[MEASURED]** (`/tmp/compare.py`).
|
||
|
||
**The 16-candidate ranking A/B found no winner.** Runtime knobs were added to
|
||
the selector (`GUN_SELECTOR_WINDOW`, `MINOBS`, `TIE`, `FLOOR`, `POOL`, `RANK`,
|
||
`SHRINK`, `SEED`; `rankScore` supports mean/Wilson/UCB/Thompson/shrinkage), all
|
||
defaulting to the shipped values. 16 candidates were A/B'd against the boss.
|
||
None credibly beat the shipped config; every candidate's per-run interval
|
||
overlaps base, and the nominal "winners" are ≤0.6 SE apart on far fewer shots.
|
||
**[MEASURED]** (commit `2c94dc2`; `/tmp/compare.py`):
|
||
|
||
| Config | Runs | Shots | Rate % | dmg/run | Spearman |
|
||
|---|---:|---:|---:|---:|---:|
|
||
| **base (shipped)** | 14* | 3,612 | **6.95** | 251 | −0.374 |
|
||
| tie00 (TIE=0.0) | 2 | 540 | 7.04 | 242 | +0.018 |
|
||
| win50 (WINDOW=50) | 2 | 559 | 6.08 | 216 | +0.588 |
|
||
| wilson (RANK=wilson) | 13 | 3,045 | 6.67 | 196 | +0.088 |
|
||
| thompson (RANK=thompson) | 2 | 531 | 4.90 | 166 | +0.083 |
|
||
| maxbin (POOL=0) | 2 | 574 | 5.23 | 198 | −0.264 |
|
||
| minobs20 (MINOBS=20) | 2 | 535 | 5.98 | 212 | +0.144 |
|
||
| tie05 (TIE=0.05) | 12 | 3,357 | 6.20 | 222 | −0.060 |
|
||
| tie10 (TIE=0.10) | 2 | 556 | 6.65 | 233 | −0.150 |
|
||
| tie40 (TIE=0.40) | 2 | 583 | 6.35 | 216 | −0.160 |
|
||
| floor10 (FLOOR=0.10) | 6 | 1,694 | 6.49 | 232 | −0.578 |
|
||
| t05f10 (TIE=0.05, FLOOR=0.10) | 4 | 1,073 | 5.50 | 190 | −0.041 |
|
||
| wilf10 (RANK=wilson, FLOOR=0.10) | 4 | 1,252 | 6.71 | 271 | −0.410 |
|
||
| floor00 (FLOOR=0.0) | 3 | 788 | **5.08** | 175 | +0.055 |
|
||
| f00t05 (FLOOR=0, TIE=0.05) | 3 | 540 | 6.48 | 153 | +0.226 |
|
||
| f00wil (FLOOR=0, RANK=wilson) | 2 | 595 | 6.72 | 221 | −0.116 |
|
||
| f00w50 (FLOOR=0, WINDOW=50) | 2 | 593 | 6.58 | 210 | +0.178 |
|
||
|
||
\* compare.py counts 14 events files, but run 14 has no fire events; the 13
|
||
runs with data carry all 3,612 shots.
|
||
|
||
**No ranking rule produced a stable, useful correlation.** The best Spearman in
|
||
the table (win50, +0.588) is on 2 runs / 559 shots; none of the 16 tested rules
|
||
recovered a sign-stable signal. The shipped config's −0.374 is the largest
|
||
single-run estimate but is contradicted in sign by the +0.335 over 15 runs, so
|
||
a single Spearman value on one run set is not a reliable estimate.
|
||
**[MEASURED]** + **[INFERRED]**.
|
||
|
||
### 3.1 Offline range: which gun wins which trajectory family
|
||
|
||
Shipped `path` metric, `/tmp/range_path.txt` (52,770/103,938 = 50.8%). Cells
|
||
are hit-% per gun per fixture; **bold** = best gun for that fixture. 400 shots
|
||
per gun per fixture.
|
||
|
||
| Fixture | HeadOn | Linear | Tsetlin | Circular | GuessF | Pattern | WallBn | Accel | StopSh | Displ | AvgLead | DecayG | KNN |
|
||
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
||
| circular | 12 | 29 | 29 | **100** | 15 | 66 | 34 | 100 | 27 | 18 | 54 | 13 | 73 |
|
||
| constant-velocity | **100** | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
|
||
| contr-SittingDuck | **100** | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
|
||
| contr-StraightLine | 43 | **100** | 72 | 100 | 100 | 100 | 94 | 98 | 84 | 98 | 100 | 100 | 77 |
|
||
| decel-before-turn | 79 | **100** | 90 | 100 | 100 | 100 | 100 | 100 | 98 | 100 | 100 | 100 | 78 |
|
||
| classC-corners | 6 | 9 | **23** | 17 | 9 | 17 | 12 | 19 | 19 | 22 | 13 | 10 | 4 |
|
||
| classC-crazy | 4 | 36 | 24 | 34 | 35 | 32 | **54** | 35 | 22 | 26 | 39 | 27 | 22 |
|
||
| classC-mirror | **12** | 6 | 7 | 6 | 6 | 9 | 6 | 7 | 8 | 5 | 8 | 7 | 8 |
|
||
| classC-ramfire | 35 | 50 | 32 | 54 | 50 | 54 | 49 | 49 | 35 | 30 | **56** | 52 | 46 |
|
||
| classC-spinbot | 3 | **58** | 10 | 50 | 58 | 37 | 48 | 50 | 13 | 46 | 52 | 58 | 42 |
|
||
| energy-threshold-turner | 58 | 34 | 39 | **100** | 34 | 88 | 38 | 100 | 40 | 44 | 63 | 34 | 42 |
|
||
| oscillator | **100** | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
|
||
| random-walk | 26 | 67 | 64 | 28 | 63 | 41 | **68** | 26 | 61 | 67 | 55 | 62 | 47 |
|
||
| stationary | **100** | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
|
||
| TR-corners | 3 | 30 | 28 | 33 | 39 | 26 | 33 | 34 | 30 | 34 | **56** | 29 | 32 |
|
||
| TR-crazy | 5 | 7 | 12 | 16 | 2 | 14 | 19 | **28** | 11 | 20 | 27 | 2 | 3 |
|
||
| TR-ModularBot | 15 | 12 | 18 | 13 | 12 | 12 | 14 | 15 | **21** | 12 | 13 | 11 | 7 |
|
||
| TR-MB-shield | 4 | 10 | 32 | 27 | 10 | 28 | 32 | 28 | 24 | **37** | 21 | 27 | 5 |
|
||
| TR-spinbot | 3 | 21 | 20 | 21 | **37** | 37 | 25 | 30 | 27 | 14 | 24 | 15 | 11 |
|
||
| wall-bounce | 0 | 49 | 44 | 48 | 37 | 41 | **100** | 60 | 41 | 34 | 52 | 34 | 28 |
|
||
|
||
Aggregated hit-% by family (400 shots/gun/fixture):
|
||
|
||
| Family | HeadOn | Linear | Tsetlin | Circular | GuessF | Pattern | WallBn | Accel | StopSh | Displ | AvgLead | DecayG | KNN |
|
||
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
||
| synthetic (8) | 59 | 72 | 71 | 84 | 69 | 79 | 80 | **86** | 71 | 70 | 78 | 68 | 71 |
|
||
| classic DrussGT (5) | 12 | 32 | 19 | 32 | 32 | 30 | **34** | 32 | 20 | 26 | 34 | 31 | 25 |
|
||
| TR bridge (5) | 6 | 16 | 22 | 22 | 15 | 23 | 24 | 27 | 23 | 23 | **28** | 17 | 12 |
|
||
| all 20 | 35 | 51 | 47 | 57 | 49 | 55 | 56 | **59** | 48 | 50 | 57 | 49 | 46 |
|
||
|
||
**Which gun wins which trajectory (offline, path metric):**
|
||
|
||
- **Stationary / constant-velocity / oscillator / decel-before-turn /
|
||
straight-line:** every useful gun ≥ 94% (perfect-information + path metric
|
||
makes these uninformative). Only HeadOn (43–79%) and KNN (77–78%) stand out
|
||
as weak.
|
||
- **Circular:** Circular and Accel 100% by construction; KNN 73, Pattern 66.
|
||
- **Wall-bounce:** WallBounce 100, Accel 60, AvgLead 52 — the only fixture
|
||
where WallBounce is dominant.
|
||
- **Random-walk:** WallBounce 68, Displace 67, Linear 67, Tsetlin 64; Accel
|
||
collapses to 26.
|
||
- **Energy-threshold rule:** Circular and Accel 100, Pattern 88, AvgLead 63.
|
||
Tsetlin 39 — it beat Linear (34) but is far from reading the rule.
|
||
- **Classic DrussGT (a real surfer):** WallBounce 34 and AvgLead 34 at the top,
|
||
HeadOn 12 at the bottom. Cornered surfing is the one case where Tsetlin (23)
|
||
leads.
|
||
- **TR bridge DrussGT:** AvgLead 28 overall; Accel 28 on TR-crazy, AvgLead 56
|
||
on TR-corners, StopShot 21 on the ModularBot mirror, Pattern 37 on TR-spinbot.
|
||
|
||
**Do not read the offline winner as the rack verdict.** The offline range's
|
||
*classic DrussGT* order (WallBounce 34 top, KNN 25 low) is close to the
|
||
**inverse** of the real order (KNN 9.0% third, WallBounce 6.2% ninth). The
|
||
offline range is excellent for catching structural bugs (§6) and for
|
||
per-trajectory sanity, and poor as a selector signal. **[MEASURED]** +
|
||
**[INFERRED]**.
|
||
|
||
---
|
||
|
||
## 4. Verdict table
|
||
|
||
| Gun | Family / model | Real % (13 runs) | Offline all-20 % | Verdict |
|
||
|---|---|---:|---:|---|
|
||
| Linear | constant-velocity lead | 10.7 | 51 | **KEEP — this is what a surfer cannot defeat; keep warm** |
|
||
| Circular | constant-turn lead | 9.9 | 57 | **KEEP** |
|
||
| KNN | k-NN on motion history | 9.0 | 46 | **KEEP — offline under-rates it badly** |
|
||
| Pattern | pattern replay | 8.6 | 55 | **KEEP** |
|
||
| Accel | acceleration-aware lead | 7.3 | 59 | **KEEP** |
|
||
| AvgLead | windowed average lead | 7.0 | 57 | **KEEP** |
|
||
| GuessFactor | GF histogram | 6.9 | 49 | MARGINAL — small N, no clear edge |
|
||
| DecayGF | recency-weighted GF | 6.4 | 49 | MARGINAL |
|
||
| WallBounce | wall-reflection model | 6.2 | 56 | MARGINAL — offline favourite, real underperformer |
|
||
| StopShot | deceleration/stop point | 6.1 | 48 | MARGINAL |
|
||
| Tsetlin | Tsetlin-Machine correction | 5.8 | 47 | **KEEP — below overall; pruning tested neutral-to-negative (§3)** |
|
||
| Displace | displacement vector | 5.3 | 50 | **KEEP — below overall; pruning tested neutral-to-negative (§3)** |
|
||
| HeadOn | aim at current position | 5.2 | 35 | **KEEP — mandatory floor fallback** |
|
||
| TMSelect | TM mixture-of-experts gate | — | — | **DISABLED (`EnableTmSelector = false`)** — see §6.4 |
|
||
|
||
---
|
||
|
||
## 5. What "worth keeping" means, and what is not proven
|
||
|
||
The KEEP/MARGINAL/BELOW split is a statement about a **single adversary (a
|
||
wave surfer)**, judged on the shipped config, on a few hundred real shots per
|
||
gun. It is a starting point, not a final ranking. In particular, **BELOW does
|
||
not mean "drop"**: pruning the measured-worst real performers (Tsetlin, then
|
||
Tsetlin+Displace) was tested in 15 paired runs each and was
|
||
neutral-to-slightly-negative on both hit rate and damage (§3), so the verdict
|
||
is **keep the full rack**. The concrete caveats are in §7.
|
||
|
||
---
|
||
|
||
## 6. Bugs found and fixed tonight (why earlier rack verdicts were wrong)
|
||
|
||
Five of these changed the rack ordering; all are [MEASURED] from the commit
|
||
messages and the offline range.
|
||
|
||
### 6.1 The GF family aimed at the fire-time RADIUS, not the angle
|
||
|
||
`guess_factor`, `decay_gf` and `knn_gun` aimed at the fire-time distance. But
|
||
the virtual-bullet metric resolves a bullet at its **aim-point distance** and
|
||
scores that single point against the enemy's position on that tick, so with any
|
||
radial target motion the bullet stopped at the wrong radius and missed even
|
||
with a perfect angle. **Angle-only prediction is structurally unscoreable
|
||
under the point metric.** Two competing hypotheses were tested and **both
|
||
refuted**: (a) MEA range too narrow — 0 clamped shots out of 837/849/957, with
|
||
required offsets peaking at ~33° against MEA 28.1–46.7°, and `arcsin(8/bulletSpeed)`
|
||
correctly uses max robot *speed*, not the hit radius; (b) wrong GF peak — a
|
||
sweep of every constant GF value showed the oracle-best constant offset was
|
||
only 6% on circular, 4% on wall-bounce, 7.5% on random-walk. Learning was fine
|
||
too (~850–960 observations per fixture, 0 starved waves). **[MEASURED]**
|
||
(commit `7f706e5`).
|
||
|
||
Fix: a self-consistent constant-velocity `lead_forecast.nim` base, so the
|
||
histogram learns the **residual** and the aim point lands at the right radius;
|
||
also fixed `linear.nim` (it did a one-shot extrapolation and never iterated its
|
||
flight time). Before/after, offline: circular GF 6→23, DecayGF 6→21;
|
||
wall-bounce GF 0→60.2, DecayGF 0→60.2; constant-velocity GF/DecayGF/KNN
|
||
26→100; random-walk GF 0→53; StraightLine GF 8→77. The oracle-best constant GF
|
||
moved 6%→20% (circular), 4%→57% (wall-bounce), 7.5%→49% (random-walk),
|
||
proving the structural fix independently of tuning. **Honest trade-off:** on
|
||
the 5 real DrussGT surfer captures the GF family regressed (GuessFactor
|
||
108→55, DecayGF 108→76, KNN 101→74 hits/2000) because a linear base is a poor
|
||
model for a surfer and the residual histogram is noisier than the old
|
||
total-lead histogram.
|
||
|
||
That regression was then recovered by **blending the range** between a
|
||
radial-only forecast and the geometric one by the measured radial fraction
|
||
(`radialFrac`), keeping the constant-velocity bearing. Nine candidate bases
|
||
were measured and rejected with numbers (velocity scaling 0.8 recovered DrussGT
|
||
but destroyed wall-bounce 241→20; radial-only range wall-bounce 241→140;
|
||
short-window average worse than both; reversal/speed gates weaker than the
|
||
blend). Result (hits/2000): classic-5 GF 55→**171**, DecayGF 76→100; TR-5 GF
|
||
9→86, DecayGF 4→87; synthetic-10 GF 2702→2717. The only figure below the old
|
||
base is classic-5 DecayGF (108→100, within noise). **[MEASURED]** (commit
|
||
`e2ca2fc`).
|
||
|
||
### 6.2 The wave queues were starved 1-push-vs-4-pops
|
||
|
||
`predict()` stored **one** wave per tick while `onResult()` popped one per
|
||
resolved bullet (~4/tick), so the queue drained within a few dozen ticks and
|
||
~3 of every 4 resolutions returned without learning; the survivor paired with
|
||
a same-tick wave (`bearingDelta ≈ 0`), pinning the histogram at centre.
|
||
**Proof:** `GF.vHits == HeadOn.vHits` and `DecayGF.vHits == HeadOn.vHits`
|
||
byte-for-byte in **every one of 50 rounds** — GF, DecayGF and KNN had
|
||
degenerated into HeadOn clones. Fix: per-bin FIFO with an O(1) head cursor, at
|
||
most one push per (tick, bin). Also `maxBullets` 2,048→8,192: the rack spawns
|
||
52 bullets/tick so the ring wrapped every ~39 ticks while a long power-3 shot
|
||
needs ~90, silently discarding unresolved bullets and biasing every measured
|
||
hit rate by range; a `droppedBullets` counter was added. After the fix
|
||
`vDropped = 0` and `vStarved = 0` across all 48 recorded rounds. **[MEASURED]**
|
||
(commit `0cc6821`).
|
||
|
||
### 6.3 Tsetlin's clauses saturated at ~714 included literals each
|
||
|
||
`Tsetlin.vHits` was byte-for-byte equal to `Linear.vHits` in every measured
|
||
round of every run because its learned correction was always exactly 0.
|
||
Root cause: `tmLearnOne` rewarded included true literals unconditionally,
|
||
omitting Granmo's `(c=0, lk=1) → toward Exclude` counter-force, so true
|
||
literals ratcheted toward Include forever; Type II was unreachable dead code
|
||
with the wrong direction; resource allocation was an `|error|` heuristic
|
||
instead of Granmo's `(T − clip(v,−T,T))/(2T)`; the label baseline had a
|
||
factor-2 shrink (`error = δ − 2c`, fixed point `c = δ/2`); hits zeroed their
|
||
residual; the enemy-energy feature was duplicated (`state.selfEnergy` fed
|
||
where `WorldState.enemyEnergy` exists, so energy rules were literally
|
||
unrepresentable); and `tmEvalClause` needed Granmo Eq. 6 (all-Exclude clause
|
||
outputs 1 during learning, 0 during classification) or fix #1 deadlocks every
|
||
clause at empty.
|
||
|
||
Measured effect (energy-threshold-turner, seed 1): mean included
|
||
literals/clause **714.0 → 13.8**; active clauses 100/100 → 53/100; nonzero
|
||
corrections 8/764 → 708/764; Tsetlin virtual hits **27/400 → 69/400** (Linear
|
||
43/400). Tsetlin now **learns** but is **not yet competitive with Linear** —
|
||
the regression head is untuned, flagged as follow-up rather than claimed as a
|
||
win. **[MEASURED]** (commit `8937000`; `/tmp/ab_logs3/final_test_tsetlin_gun.log`).
|
||
|
||
### 6.4 The TM classifier gun did not earn its slot (but its clauses are real)
|
||
|
||
A Tsetlin-Machine mixture-of-experts gate over
|
||
HeadOn/Linear/Circular/WallBounce/Accel was built with the corrected feedback
|
||
and labelled by which expert's prediction was closest to the actual enemy
|
||
position (an exact, supervised, per-shot label — no delayed credit). It loses
|
||
to the best of its own experts offline on nearly every fixture, and against
|
||
DrussGT it cost real performance:
|
||
|
||
```
|
||
baseline (path + relative) 7.56% real hit rate, 157 dmg
|
||
+ power fix 7.47%, 239 dmg
|
||
+ power fix + TM selector 5.59%, 133 dmg
|
||
```
|
||
|
||
It was selected on 806 ticks and fired 24 real shots at 4.2%. It ships disabled
|
||
(`EnableTmSelector = false`; code and wiring kept intact). **However, the gate
|
||
latched onto meaningful structure:** on the energy-threshold turner, HeadOn's
|
||
clauses key on the **energy bits** (the rule's own driving variable) while
|
||
Circular keys on distance/velocity. So the TM learned something real and
|
||
interpretable; it simply could not beat "always pick the best expert".
|
||
**[INFERRED]** root cause: the closest-expert label is noisy because several
|
||
experts are near-tied, and under the path metric the winner varies by power bin
|
||
while the gate sees one shared per-tick input, so a one-vs-rest gate over a
|
||
saturated 870-bit clause space has no margin to exploit. A standalone Granmo
|
||
classifier on the same encoding reaches ~99% on the rule but, per the
|
||
counterfactual probe, does **not** read energy (follow rate 24% high / 62% mean
|
||
— statistically identical at 1, 2 and 10 frames), so even the "it learned the
|
||
rule" claim is limited to ~99% accuracy, not to a readable energy threshold.
|
||
The best recovered proposition was `!g9 ∧ !g8` (energy < 25.6, not the labelled
|
||
30) — a genuine simple threshold, but not the ensemble's decision mechanism.
|
||
**[MEASURED]** (commits `57b2ac3`, `d5061ee`; `test_tm_pattern_learning.nim`).
|
||
|
||
### 6.5 The selector thresholds were absolute on a rescaled metric
|
||
|
||
Covered in §2.2: the `0.10` absolute floor fired on 53.0% of point-metric ticks
|
||
and forced HeadOn (real 2.0–4.4%, 11th of 13); HeadOn selection share fell
|
||
69.1% → 43.5% under relative thresholds (and 23.1% → 24.2% under path). Also:
|
||
`bestGun` was first-index-wins argmax, so HeadOn at index 0 silently won every
|
||
tie until the random tie-break landed (`343e631`); `bestPower` had the same
|
||
absolute-40% defect (below).
|
||
|
||
### 6.6 Power selection was stuck at power 1.0 (`MinHitRate = 0.40`)
|
||
|
||
`bestPower` used an **absolute** `MinHitRate = 0.40` bar. Measured per-bin
|
||
virtual rates show no bin ever clears 40%, so 11 of 14 guns were stuck at bin 0
|
||
(power 1.0) even where higher bins were comparable:
|
||
|
||
```
|
||
Linear p1.0 44% p1.5 39% p2.0 30% p3.0 29% old bin 0 -> new bin 3
|
||
Accel p1.0 44% p1.5 40% p2.0 26% p3.0 29% old bin 1 -> new bin 3
|
||
Pattern p1.0 50% p1.5 40% p2.0 27% p3.0 12% old bin 1 -> new bin 2
|
||
```
|
||
|
||
Replaced with a scale-aware `PowerBarFrac = 0.50` (a dimensionless fraction of
|
||
the gun's own best-bin rate); 13 of 14 selections now pick heavier bullets.
|
||
Real effect vs DrussGT (8 rounds × 3 runs): hit rate unchanged (7.56% → 7.47%),
|
||
**damage +52% (157 → 239 per run)** and rounds end faster. **[MEASURED]**
|
||
(commit `57b2ac3`).
|
||
|
||
Smaller fixes in the same family: `bestPower` on a cold gun returned the
|
||
*highest* bin (empty bin satisfied the `count == 0` clause); `fitnessFor`
|
||
aggregated enemies in nondeterministic hash order; `stop_shot` had an
|
||
unreachable deceleration branch and several guns had tick-only caches that made
|
||
all four power bins return bin 0's lead (`e536900`). **[MEASURED]**.
|
||
|
||
### 6.7 The selector's tie-break was not actually random
|
||
|
||
`randomize()` was reached only **incidentally**, through the Tsetlin gun's
|
||
constructor, so ties resolved **identically across process restarts** — the
|
||
"random" tie-break was effectively deterministic. Now fixed with an explicit
|
||
startup seed plus a `GUN_SELECTOR_SEED` override. Evidence: unseeded runs vary
|
||
across processes, seeded runs are identical. **[MEASURED]**.
|
||
|
||
### 6.8 The arrival-accuracy tie-band does not beat the shipped band
|
||
|
||
**Hypothesis.** Rank by `path` (robust, keeps its measured advantage) but narrow
|
||
the tied random draw by **arrival accuracy** (`point`): a gun whose ray sweeps
|
||
the target generously can sit in the band while its bullets arrive badly, so
|
||
making the band informative should improve the real hit rate without removing
|
||
the load-bearing randomness.
|
||
|
||
**Implementation** (`GUN_SELECTOR_TIEBREAK`, `common_libs/gun_harness/`,
|
||
default `off`): each virtual bullet is additionally scored with the `point`
|
||
model at the exact tick it reaches its aim distance into a parallel
|
||
`GunFitness.pointBins` window, and the `path` tie band is narrowed to the guns
|
||
within `GUN_SELECTOR_POINT_TIE` (default 0.5) of the best in-band point rate. The
|
||
uniform random draw over the narrowed band is kept. `GUN_SELECTOR_TIEBREAK=point`
|
||
selects it; `=commit` is the no-randomness control. `off` performs no parallel
|
||
scoring at all, so the shipped path is byte-identical. The recording rule and
|
||
the ranking rule are pinned by `common_libs/tests/test_selector_tiebreak.nim`
|
||
(19 checks, no battle).
|
||
|
||
**Live A/B** vs the real DrussGT, one frozen binary, 7 runs x 7 rounds per arm,
|
||
server-side events sidecar (~200-260 shots/run). `d = base - arm` (positive =
|
||
arm worse); p is the exact two-sided permutation test on per-run rates.
|
||
**[MEASURED]** (`/tmp/battle_tb*_r*.log`, `/tmp/events_tb*_r*.json`):
|
||
|
||
| Arm | Runs | Shots | Real % | dmg/run | d | p |
|
||
|---|---:|---:|---:|---:|---:|---:|
|
||
| `tbbase` (shipped) | 7 | 4,128 | **7.17** | **175** | — | — |
|
||
| `tbpt` (path + point narrow) | 7 | 3,938 | 7.08 | 165 | +0.14 | 0.88 |
|
||
| `tbpc` (`=commit` control) | 7 | 3,759 | 4.44 | 98 | +2.74 | **0.0012** |
|
||
| `tbpt25` (point margin 0.25) | 7 | 3,683 | 5.59 | 119 | +1.65 | 0.20 |
|
||
| `tbtie05` (`TIE=0.05`) | 7 | 3,937 | 5.84 | 133 | +1.49 | 0.11 |
|
||
| `tbtie40` (`TIE=0.40`) | 7 | 3,917 | 6.28 | 144 | +1.00 | 0.25 |
|
||
| `tbwin50` (`WINDOW=50`) | 7 | 3,983 | 6.05 | 139 | +1.20 | 0.11 |
|
||
| `tbfloor10` (`FLOOR=0.10`) | 7 | 3,829 | 5.33 | 118 | +2.12 | 0.11 |
|
||
|
||
The mechanism DID fire — the tie-break re-shaped the selection mix (over all 7
|
||
runs: Pattern 24%→16%, Accel 6%→16%, Tsetlin ~0%→13% under `tbpt`; the
|
||
no-randomness control `tbpc` collapses to HeadOn 43% vs 27%) — but it **did not
|
||
improve the real hit rate**: 7.08% vs 7.17%, fully
|
||
overlapping per-run ranges (base 5.29-8.73, arm 3.71-10.39), p = 0.88. The
|
||
no-randomness control `tbpc` is **significantly worse** (4.44%, p = 0.0012),
|
||
which independently replicates the earlier "commitment to the virtual best
|
||
costs real hit rate" result and validates that the arm was live. Every knob
|
||
variant (`TIE`, `FLOOR`, `WINDOW`) is also nominally *worse* than the shipped
|
||
values, none credibly better. **Verdict: clean negative — the shipped selector
|
||
is unchanged** (`GUN_SELECTOR_TIEBREAK` defaults to `off`).
|
||
|
||
This is consistent with §2: the virtual rate is a poor ranker, and the
|
||
selector's value is its **floor/tie hedging**, not the ordering it computes.
|
||
Narrowing the band with a second virtual statistic changes *which* guns are
|
||
drawn without making that draw any better.
|
||
|
||
---
|
||
|
||
## 7. Known caveats and open problems
|
||
|
||
Stated without hedging.
|
||
|
||
1. **The headline per-gun numbers come from ONE adversary, a wave surfer.**
|
||
HeadOn is genuinely bad against surfers, so part of the rack ordering may be
|
||
matchup-specific. A SpinBot guard was inconclusive: ModularBot fires only
|
||
17–31 real shots/run against a fast bot because the range-aware firing gate
|
||
is strict at long range, so the guard had little power (Wilson looked better,
|
||
18.5% vs 8.6%, but on 70–92 shots with a 5–33% spread). **[MEASURED]**
|
||
(commit `2c94dc2`). A second, independent adversary at scale is missing.
|
||
|
||
2. **Per-gun real N is small.** 47–898 shots per gun; n < 200 gives roughly
|
||
±5 pp across a 3–15% spread. Single-gun ordering is **indicative, not
|
||
definitive**. The KEEP/MARGINAL/BELOW boundaries should be treated as soft.
|
||
|
||
3. **The fixtures are perfect-information and therefore optimistic.** Every
|
||
fixture is an observer capture with true positions every tick; the classic
|
||
set is additionally **open-loop** (replayed DrussGT never dodges our
|
||
bullets). Absolute offline hit rates are inflated by an unknown amount; only
|
||
relative comparisons are safe.
|
||
|
||
4. **The virtual metric is a poor ranker, not a reliable inverse.** The
|
||
correlation with real hit rate is weak and **sign-unstable** across run
|
||
sets: −0.374 over the 13-run base, +0.335 over a 15-run paired baseline
|
||
(different aggregation), +0.522 on the 5-run `relative+path` set, and
|
||
−0.371 / −0.073 / −0.037 on the other 5-run configs. Two opposite signs on
|
||
large samples mean it is near zero on average, not reliably anti-correlated;
|
||
an earlier version of this report overstated it as an inversion and a later
|
||
measurement refuted that. 16 candidate ranking rules all overlapped the
|
||
shipped config, so none produced a stable, useful correlation. The
|
||
selector's value lives in its **floor/tie hedging** (5.08% without the floor
|
||
vs 6.95% with it), not in its ranking. **[MEASURED]** + **[INFERRED]**.
|
||
|
||
5. **The TM classifier gun did not earn its slot.** It cost real performance
|
||
(7.47% → 5.59%, 133 dmg) despite showing interpretable energy structure in
|
||
its clauses (§6.4). It is disabled; re-enabling requires a fix to the gate
|
||
margin/label problem, not more training.
|
||
|
||
6. **Real-hit-rate-driven selection is not viable yet.** Only the selected gun
|
||
fires, so unselected guns get near-zero real shots (GuessFactor 20, Linear
|
||
24 vs HeadOn 733 in the point A/B); noise is fatal (n = 470 at p = 10% gives
|
||
±2.8 pp, most guns n < 200 gives ±5 pp+); and real rate is conditional on
|
||
when the gun was selected. A blended signal with forced exploration and
|
||
shrinkage is defensible in principle but needs thousands of shots per gun
|
||
across many battles. Real rate is currently best used **offline** as the
|
||
evaluation metric — which is exactly what the A/B does. **[MEASURED]**
|
||
(commit `dea4dcb`).
|
||
|
||
7. **The offline==online acceptance test is fixed and stable** (§1.1): the old
|
||
11/12 flakiness was a real replay bug (the replay spawned disabled gun 13,
|
||
and the shared order-sensitive ring then permuted every other gun's
|
||
resolution order), now fixed by mirroring the live rack. 5/5 consecutive runs
|
||
give a byte-identical 12/12 with the death boundary included.
|
||
|
||
8. **The selector's tie-break is now explicitly seeded** (§6.7). It had been
|
||
effectively non-random — `randomize()` was reached only incidentally through
|
||
the Tsetlin gun's constructor — so ties resolved identically across process
|
||
restarts. Fixed with an explicit startup seed plus a `GUN_SELECTOR_SEED`
|
||
override; seeded runs are reproducible, unseeded runs vary.
|
||
|
||
9. **The firing gate is not the bottleneck.** The shipped range-aware gate does
|
||
not beat a fixed 2.0° gate on hit rate (55.8% vs 57.9%, ~1.5 σ), though it
|
||
fires 22–28% more shots. No per-adversary score delta exceeded the
|
||
~300-point run-to-run noise band. **[MEASURED]** (commit `3c90a59`).
|
||
|
||
10. **The boss is ~2.3× more accurate than the whole rack** (12.1% vs 5.3% in
|
||
the capture). Closing that gap is the point of the rack; the current
|
||
best single gun is 10.7%.
|
||
|
||
---
|
||
|
||
## 8. Reproduction
|
||
|
||
Commands recorded in the commits and tool READMEs. (I was instructed not to
|
||
run builds/tests while writing this report; these are the documented
|
||
invocations, not a fresh verification by me.)
|
||
|
||
```bash
|
||
# Offline gun range over all 20 fixtures, shipped path metric (default):
|
||
nim c -r common_libs/tests/run_range.nim
|
||
# Point metric for comparison:
|
||
GUN_VBULLET_METRIC=point nim c -r common_libs/tests/run_range.nim
|
||
# Add timing:
|
||
nim c -r common_libs/tests/run_range.nim --timing
|
||
|
||
# Selector diagnostics (floor/tie/bestRate/HeadOn-share) on a fixture:
|
||
GUN_SELECTOR_MODE=relative nim c -d:release -r \
|
||
common_libs/tests/analyze_selector.nim tools/fixtures/drussgt_vs_spinbot.jsonl
|
||
|
||
# Offline == online acceptance (fixed; stable exact 12/12 — see §1.1):
|
||
nim c -r common_libs/tests/acceptance_offline_vs_online.nim
|
||
|
||
# Tsetlin gun clause sparsity / divergence:
|
||
nim c -r common_libs/tests/test_tsetlin_gun.nim
|
||
# TM readability (standalone Granmo classifier on the energy-threshold rule):
|
||
nim c -r common_libs/tests/test_tm_pattern_learning.nim
|
||
|
||
# Live boss (real DrussGT jar; jars stay out of git, see the README):
|
||
# tools/robocode_shim/run_bridge_battle.sh <bot_dir> <rounds> <capture.jsonl>
|
||
# Closed-loop evidence for the TR captures:
|
||
python3 tools/robocode_shim/analyze_closed_loop.py \
|
||
tools/fixtures/tr_drussgt_vs_modularbot.jsonl \
|
||
tools/robocode_shim/evidence/tr_drussgt_vs_modularbot.events.json
|
||
```
|
||
|
||
Per-gun aggregation scripts used for the tables above:
|
||
`python3 /tmp/agg2.py base` (virtual-vs-real + Spearman),
|
||
`python3 /tmp/compare.py` (server-side per-run A/B + overlap).
|
||
Raw range output: `/tmp/range_path.txt` (path), `/tmp/final_range.txt` (point).
|