Ablation: the radial TM is replaceable by a CONSTANT, and its avenue is dead on the
shipped metric The radial TM beats Linear on bmPoint, but its head never beat the majority baseline after the label bias was fixed - suggesting the win is a constant lean rather than learning. So: sweep a stateless constant short-range offset (new `common_libs/guns/radial_offset.nim`, no learning at all) against the learned TM. VERDICT (measured, offline range, seeds=3, 18 paired runs, 231 TM rounds): 1. **bmPoint - REPLACE the TM with a constant.** `RO_s0.95` (aim distance x0.95) TIES it early (9/9, p=1.0) and BEATS it overall (15/3, p=0.0075; 7.47% vs 6.89% per-run mean). A fixed -20px does the same. The head never beats its majority baseline (56.2% vs 57.2%). 2. **bmPath (the SHIPPED metric) - the radial avenue is a DEAD END.** TMRadial is a systematic LOSS there (2/16, p=0.0013); every constant is within +-0.2pp; the only real bmPath effect is the BotRadius clamp. So the radial shift cannot help the shipped configuration. 3. The per-adversary optimum DOES vary (fixed -10 for crazy, -30 for tr_crazy, scale 0.95 for three others) - but ONE GLOBAL CONSTANT still beats the adaptively-trained head, so the "fragility justifies learning" argument FAILS. THE REAL FINDING UNDERNEATH, and it generalises beyond this gun: the base linear prediction systematically OVERSHOOTS. Measured raw per-tick base radial error has mean -71 to -100 px; the enemy is NEARER than the prediction in 63-81% of shots and farther in only 4-14%, CONSISTENT ACROSS ALL SIX CAPTURES. Radial label histogram [415166,126461,120325,44694,19351] = 57.2% majority class, mean label -82.3 px, mean applied shift -37.6 px. So the net-short bias is a GENUINE property of these range-holders against a constant-velocity extrapolation (they decelerate and turn, so the true position is closer than the straight-line guess) - NOT a fixture artefact. That is worth chasing for the guns that actually ship. Caveat: bmPoint is not the shipped metric (bmPath won the real-hit-rate A/B for SELECTION), so a bmPoint win is not yet evidence of a real win. That needs a live test - and the natural target is Pattern, which is now the default and best gun. Adds radial_offset.nim + sweep_radial_offset.nim; tm_pattern.nim gains additive instrumentation only (radial label mean and applied-shift mean; no behaviour change, and test_tm_pattern_registration still passes all 20 checks).
This commit is contained in:
@@ -364,3 +364,229 @@ by these numbers. There is no persistence across battles.
|
||||
* INFERRED: that the radial win comes from surfers being NEARER than the base
|
||||
prediction (range-holding), supported by the asymmetric radial label histogram
|
||||
but not separately modelled.
|
||||
|
||||
---
|
||||
|
||||
# ROUND 3 — ABLATION: does the radial win even need the Tsetlin Machine?
|
||||
|
||||
Date: 2026-09-22. Artifacts: `common_libs/guns/radial_offset.nim` (new),
|
||||
`common_libs/tests/sweep_radial_offset.nim` (new), `common_libs/guns/tm_pattern.nim`
|
||||
(label/readout instrumentation only). Raw outputs: `/tmp/ro_point_s3.txt`,
|
||||
`/tmp/ro_path_s3.txt` (`--set=real --seeds=3`), `/tmp/ro_point_s1.txt`.
|
||||
|
||||
## The question
|
||||
|
||||
Commit 589a230 retracted the radial head's "conditional learning" claim: the
|
||||
radial head's online accuracy (57.0%) is AT/BELOW the 58.2% majority baseline, and
|
||||
the bmPoint metric win was attributed to a **NET-POSITIVE AVERAGE RADIAL SHIFT**.
|
||||
If that is what it is, a FIXED radial shift should reproduce the win with no
|
||||
learning, no 0.36 ms/tick cost and no risk.
|
||||
|
||||
`radial_offset.nim` is exactly that fixed shift and nothing else: the exact
|
||||
`forecastLinear` bearing, aim distance `f.dist * scale + offsetPx`, and the SAME
|
||||
`[BotRadius, arena-BotRadius]` clamp the TM's corrective path uses. It is
|
||||
stateless, so two runs are identical by construction. The sweep runs **Linear**,
|
||||
a grid of constants, **TMRadial** and **TMRadialShuf** under BOTH metrics, using
|
||||
the identical fixture set / round splitting / per-fixture fresh gun methodology
|
||||
as `sweep_tm_pattern.nim`.
|
||||
|
||||
**Harness sanity check (MEASURED):** `TMRadial` in this report reproduces the
|
||||
committed numbers exactly — bmPoint early 9.1% / overall 5.7%, vs Linear early
|
||||
16/2 p=0.0013 and overall 18/0 p<0.0001, vs Shuf 18/0 p<0.0001. So the ablation
|
||||
runs on the same experiment, not a re-derivation.
|
||||
|
||||
0 dropped bullets in every run (ring never clobbered); `labelMiss=0` and
|
||||
`traceMiss=0` for every TM row (the deferred-label fix holds).
|
||||
|
||||
## bmPoint (seeds=3)
|
||||
|
||||
Pooled over 77 rounds for the deterministic arms (one pass; replicated across the
|
||||
3 seeds only for the paired test) and 231 rounds for the TM arms (3 seeds).
|
||||
|
||||
| variant | early | overall |
|
||||
|---|---|---|
|
||||
| Linear | 7.2% (1480/20498) | 4.7% (11277/241423) |
|
||||
| `RO_s1.00` (base + BotRadius clamp only) | 7.2% (1481/20626) | 4.7% (11388/241551) |
|
||||
| `RO_s0.98` (scale 0.98) | 8.0% (1669/20781) | 5.8% (14033/241677) |
|
||||
| **`RO_s0.95`** | **8.8% (1852/21001)** | **6.3% (15299/241890)** |
|
||||
| `RO_o-10` (fixed -10 px) | 8.4% (1756/20790) | 5.9% (14336/241689) |
|
||||
| **`RO_o-20`** | 7.9% (1658/20939) | **6.4% (15445/241836)** |
|
||||
| `RO_o-30` | 5.4% (1130/21117) | 4.9% (11740/242007) |
|
||||
| `RO_o-40` | 4.1% (883/21292) | 3.7% (8989/242166) |
|
||||
| `RO_o-60` | 4.7% (1019/21655) | 3.0% (7173/242495) |
|
||||
| `RO_o+30` (opposite-direction control) | 0.6% (131/20159) | 0.3% (650/241168) |
|
||||
| **TMRadial** | **9.1% (5776/63518)** | 5.7% (41232/726357) |
|
||||
| TMRadialShuf | 7.0% (4314/61672) | 3.8% (27317/724473) |
|
||||
|
||||
Per-run means over the 18 fixture×seed runs:
|
||||
Linear 13.61 / 5.77; `RO_s0.95` **14.86 / 7.47**; `RO_o-20` 10.35 / 7.45;
|
||||
TMRadial **15.08 / 6.89**; Shuf 12.89 / 4.54.
|
||||
|
||||
Paired sign tests (exact two-sided binomial, 18 pairs):
|
||||
|
||||
| A | B | metric | nA>B | nB>A | p |
|
||||
|---|---|---|---|---|---|
|
||||
| RO_s0.95 | Linear | early | 15 | 3 | **0.0075** |
|
||||
| RO_s0.95 | Linear | overall | 15 | 3 | **0.0075** |
|
||||
| RO_s0.98 | Linear | overall | 18 | 0 | **<0.0001** |
|
||||
| RO_o-20 | Linear | early | 15 | 3 | **0.0075** |
|
||||
| RO_o-20 | Linear | overall | 15 | 3 | **0.0075** |
|
||||
| RO_o-10 | Linear | overall | 18 | 0 | **<0.0001** |
|
||||
| RO_o+30 | Linear | both | 0 | 18 | **<0.0001** (wrong direction) |
|
||||
| TMRadial | Linear | early | 16 | 2 | **0.0013** |
|
||||
| TMRadial | Linear | overall | 18 | 0 | **<0.0001** |
|
||||
| TMRadial | Shuf | both | 18 | 0 | **<0.0001** |
|
||||
| **RO_s0.95** | **TMRadial** | **early** | **9** | **9** | **1.0000 (tie)** |
|
||||
| **RO_s0.95** | **TMRadial** | **overall** | **15** | **3** | **0.0075 (constant wins)** |
|
||||
| RO_o-20 | TMRadial | early | 6 | 12 | 0.2379 |
|
||||
| RO_o-20 | TMRadial | overall | 12 | 6 | 0.2379 |
|
||||
| RO_s0.98 | TMRadial | overall | 6 | 12 | 0.2379 |
|
||||
|
||||
**MEASURED conclusion (bmPoint):** a constant radial shift of **scale 0.95**
|
||||
(or fixed **−20 px**) *ties the learned radial head on EARLY and beats it on
|
||||
OVERALL* (7.47% vs 6.89% per-run mean, 15/3 p=0.0075). The learned head's whole
|
||||
nominal advantage is its slightly higher early rate (9.1% vs 8.8% pooled,
|
||||
15.08% vs 14.86% per-run), and that is a statistical tie (9/9, p=1.0).
|
||||
`RO_o+30` (aiming the opposite way) collapses to 0.3%, confirming the direction of
|
||||
the effect is real and not a clamp artefact.
|
||||
|
||||
## bmPath (the SHIPPED metric) — the radial avenue is a dead end
|
||||
|
||||
| variant | early | overall |
|
||||
|---|---|---|
|
||||
| Linear | 34.0% (6358/18715) | 24.3% (58297/239943) |
|
||||
| `RO_s1.00` (clamp only) | 34.2% (6350/18592) | **24.7%** (59217/239891) |
|
||||
| `RO_s0.98` | 34.2% (6360/18623) | 24.6% (58950/239906) |
|
||||
| `RO_s0.95` (best bmPoint constant) | 34.0% (6344/18675) | 24.4% (58530/239937) |
|
||||
| `RO_o-20` (best bmPoint constant) | 34.1% (6354/18649) | 24.5% (58699/239921) |
|
||||
| `RO_o+30` (best bmPath constant) | 34.4% (6357/18498) | **24.8%** (59366/239852) |
|
||||
| `RO_s0.80` | 33.6% (6346/18886) | 23.4% (56151/240035) |
|
||||
| TMRadial | 33.8% (19002/56211) | 24.0% (173001/719875) |
|
||||
| TMRadialShuf | 33.9% (19026/56076) | 24.3% (175262/719776) |
|
||||
|
||||
Per-run means (n=18): Linear 35.52 / 26.86; `RO_s1.00` 35.71 / **27.23**;
|
||||
`RO_s0.95` 35.59 / 26.97; `RO_o-20` 35.63 / 27.04; **TMRadial 35.37 / 26.63**
|
||||
(a loss).
|
||||
|
||||
Paired sign tests (bmPath):
|
||||
|
||||
* **TMRadial vs Linear: early 5/13 p=0.0963; overall 2/16 p=0.0013 — a
|
||||
systematic LOSS.** (Matches the committed Round-2 finding.)
|
||||
* `RO_s1.00` vs Linear: overall 18/0 p<0.0001 (+0.4 pp) — this is the pre-existing
|
||||
**BotRadius-clamp** effect, *not* the radial shift.
|
||||
* Every constant between 0.98 and 0.95 and every fixed offset −10..−20 is within
|
||||
±0.2 pp of Linear; `RO_s0.95` (a real +1.6 pp win on bmPoint overall) is only
|
||||
+0.1 pp here and *below* the clamp-only arm. Strong shrink (≤0.90, ≤−40 px) is a
|
||||
significant loss (e.g. `RO_s0.80` 3/15 p=0.0075 early, 0/18 overall).
|
||||
* `RO_o+30` — the OPPOSITE direction to what bmPoint wants — is the best bmPath
|
||||
constant (+0.5 pp, 18/0). So the bmPath response to the radial knob is
|
||||
**clamp-mediated and direction-insensitive**, i.e. the radial degree of freedom
|
||||
is a structural no-op, exactly as Round 2 argued.
|
||||
|
||||
**MEASURED conclusion (bmPath):** neither the learned radial head nor any
|
||||
constant reproduces a real gain. TMRadial is a systematic *loss*. **The whole
|
||||
radial avenue cannot help the shipped configuration.**
|
||||
|
||||
## The radial-label distribution (MEASURED)
|
||||
|
||||
Pooled `TMRadial` training labels (3 seeds, n=725997), class centres
|
||||
−60/−30/0/+30/+60 px:
|
||||
`radLabelHist = [415166, 126461, 120325, 44694, 19351]` → **majority class 57.2%**,
|
||||
`meanRadDelta = −82.30 px`, `meanAbsRadDelta = 90.13 px`.
|
||||
The head's online accuracy is **56.2%**, i.e. at/below that majority.
|
||||
The mean APPLIED shift is **−37.61 px** (`radChosenHist =
|
||||
[460960, 35411, 249071, 6832, 2994]`).
|
||||
|
||||
Metric-free raw base radial error (each fired bullet's enemy radius minus
|
||||
`forecastLinear`'s fire distance, at the base arrival tick):
|
||||
|
||||
| fixture | n | mean px | meanAbs px | frac nearer | frac farther |
|
||||
|---|---|---|---|---|---|
|
||||
| drussgt_vs_crazy | 35959 | −87.6 | 92.7 | 0.693 | 0.053 |
|
||||
| drussgt_vs_spinbot | 19830 | −100.4 | 107.7 | 0.769 | 0.062 |
|
||||
| drussgt_vs_drussgt | 26060 | −70.8 | 85.8 | 0.627 | 0.142 |
|
||||
| tr_drussgt_vs_crazy | 45905 | −99.0 | 104.4 | 0.805 | 0.053 |
|
||||
| tr_drussgt_vs_spinbot | 43150 | −92.4 | 96.7 | 0.793 | 0.044 |
|
||||
| tr_drussgt_vs_modularbot | 79970 | −79.9 | 90.7 | 0.705 | 0.108 |
|
||||
|
||||
The enemy is NEARER than the constant-velocity prediction in **63–81%** of fired
|
||||
bullets and FARTHER in only **4–14%**, consistently across all six captures and
|
||||
both metrics. So the net-short bias is a GENUINE property of these range-holding
|
||||
surfers against the constant-velocity base (the base lets range grow
|
||||
geometrically; they hold it) — not a one-fixture artefact.
|
||||
|
||||
**INFERRED:** the mean label (−82 px) is much larger than the OPTIMAL constant
|
||||
shift (−20 px). The mean is dominated by large radial errors that miss regardless
|
||||
of the shift; the near-miss window is served better by a small shift, and a
|
||||
larger shift also resolves the bullet at an earlier tick. The exact reason a
|
||||
−20 px shift beats −60 px is not separately modelled.
|
||||
|
||||
## Does the optimal constant vary by adversary? (MEASURED)
|
||||
|
||||
bmPoint, best constant PER FIXTURE (in-sample upper bound), with Linear and
|
||||
TMRadial for reference:
|
||||
|
||||
| fixture | best early | best overall | Linear overall | TMRadial overall |
|
||||
|---|---|---|---|---|
|
||||
| drussgt_vs_crazy | RO_s0.80 (10.1%) | RO_o-10 (16.9%) | 15.8% | 17.0% |
|
||||
| drussgt_vs_spinbot | RO_s0.95 (9.9%) | RO_s0.95 (11.5%) | 7.5% | 9.9% |
|
||||
| drussgt_vs_drussgt | RO_s1.00 (55.3%) | RO_s0.98 (4.3%) | 3.9% | 4.1% |
|
||||
| tr_drussgt_vs_crazy | RO_s0.85 (9.5%) | RO_o-30 (5.8%) | 2.9% | 4.4% |
|
||||
| tr_drussgt_vs_modularbot | RO_s0.95 (6.8%) | RO_s0.95 (2.7%) | 1.3% | 2.1% |
|
||||
| tr_drussgt_vs_spinbot | RO_s0.80 (6.8%) | RO_s0.95 (7.4%) | 3.3% | 3.7% |
|
||||
|
||||
**MEASURED:** the per-fixture optimum DOES vary (fixed −10 for `crazy`, −30 for
|
||||
`tr_crazy`, scale 0.95 for three others). **But** one GLOBAL constant (0.95)
|
||||
still beats TMRadial on pooled and per-run overall bmPoint. So the per-adversary
|
||||
variation is not enough to justify the learned head: a fixed 0.95 is already the
|
||||
best pooled point-metric arm measured here.
|
||||
|
||||
## DIRECT VERDICT
|
||||
|
||||
1. **bmPoint: REPLACE the radial TM with a constant.** The best constant
|
||||
(scale 0.95, equivalently fixed −20 px) statistically TIES the TM on early
|
||||
(9/9 p=1.0) and BEATS it on overall (15/3 p=0.0075; 7.47% vs 6.89% per-run
|
||||
mean). The head never out-classifies its majority baseline (56.2% vs 57.2%)
|
||||
and its mean applied shift (−37.6 px) is roughly twice the optimal constant.
|
||||
The TM buys nothing a constant does not, and costs 0.36 ms/tick + complexity.
|
||||
2. **bmPath (shipped): DEAD END.** No constant and no TM improves it; TMRadial
|
||||
is a systematic loss (2/16 p=0.0013), and the only real bmPath effect in the
|
||||
table is the BotRadius clamp, which is direction-insensitive. Drop the radial
|
||||
mode from any shipped configuration.
|
||||
3. **Fragility argument fails.** The optimum does vary per adversary, but a
|
||||
single global constant already matches/beats the adaptively-trained head — so
|
||||
the TM is not earning its cost even by the "per-adversary adaptation"
|
||||
argument (it is cold-every-battle and trains online within the battle, yet
|
||||
still loses to the global 0.95).
|
||||
|
||||
Net: **do not keep the radial TM.** If the arrival metric ever matters, ship the
|
||||
stateless constant; for the current shipped metric, the radial mode (and its
|
||||
registered gun id 14) is not justified.
|
||||
|
||||
## MEASURED vs INFERRED (round 3)
|
||||
|
||||
* MEASURED: every table, pooled rate, per-run mean, paired sign test, label
|
||||
histogram, mean/abs radial delta, applied-shift mean, online accuracy, and the
|
||||
exact reproduction of the committed TMRadial numbers.
|
||||
* MEASURED: the constant-offset arm is stateless (verbatim `forecastLinear`
|
||||
bearing, `f.dist*scale+offsetPx`, TM corrective clamp), so its rollout is
|
||||
deterministic and its replicated per-seed values are legitimate.
|
||||
* MEASURED: `RO_s1.00` isolates the BotRadius clamp — on bmPoint it is identical
|
||||
to Linear (7.2/4.7%), on bmPath it is the +0.4 pp arm; the radial shift itself
|
||||
adds nothing on bmPath.
|
||||
* INFERRED: the explanation of the label-mean (−82 px) vs optimal shift (−20 px)
|
||||
gap (large-error tail + earlier resolution tick); the direction claim itself is
|
||||
measured (the +30 control collapses).
|
||||
* INFERRED (not measured): whether a per-adversary constant would beat a global
|
||||
one out-of-sample — the per-fixture optima above are in-sample.
|
||||
|
||||
## How to reproduce (round 3)
|
||||
|
||||
```
|
||||
nim c --path:common_libs -d:release -o:/tmp/sweep_radial_offset \
|
||||
common_libs/tests/sweep_radial_offset.nim
|
||||
/tmp/sweep_radial_offset --set=real --metric=point --seeds=3
|
||||
/tmp/sweep_radial_offset --set=real --metric=path --seeds=3
|
||||
# smoke: --seeds=1
|
||||
```
|
||||
|
||||
|
||||
Reference in New Issue
Block a user