b0654d18eb
The user pushed back on "the TM can't be your best 1v1 gun", correctly, because two
decisive tests had never been run. Both are now run and they agree.
TASK 1 - THE GF HEAD vs ITS MAJORITY-CLASS BASELINE (offline, n=1,751,067):
label histogram [254286, 284578, 678879, 297055, 236269]
majority class = 2 (the CENTRE bucket) = 38.77%
RAW head accuracy = 36.69% -> margin **-2.08 pp, BELOW majority**
GATED head accuracy = 40.37% vs 38.75% majority -> +1.62 pp, BUT it predicts the
majority class on 62.4% of ticks and its minority recall is 13.6% / 12.9% - a
base-rate predictor wearing a classifier's clothes.
Shuffled control sits at its own majority (20.04% vs 20.12%), confirming chance.
**THE OLD "46% vs 20% CHANCE" FIGURE I QUOTED WAS WRONG ON TWO COUNTS:** the
baseline is 38.8%, not 20%, and the 46% predated the deferred-label fix. Against
the correct baseline the head is BELOW it.
TASK 2 - THE FIRST-EVER LIVE A/B OF THE TM GUN (7 runs x 7 rounds per arm, one
frozen binary from git archive HEAD = eb74f9b2, sha256 cb66d66b..., real DrussGT,
every arm forced alone with TR_RACK_<GUN>=both and all 14 others off, liveness
confirmed per run):
arm shots real % dmg/run round wins
onlyPattern 4610 10.74% 285 25/49
onlyTMPATTERN (radial) 3374 3.50% 71 0/49
onlyLinear 3218 3.23% 61 0/49
Pattern vs TM: +7.22 pp / +213.7 dmg, exact p=0.0006
TM vs Linear: +0.30 pp, p=0.659 (dmg p=0.438)
**The TM is statistically INDISTINGUISHABLE from its own Linear base live.** So it
is not "the TM works and we are aiming it wrong".
DIRECT ANSWER: **(c) It loses live AND sits at/below majority - the target carries
no learnable signal beyond the base rate, and that is the reason.** The reason is
not the machine, not the knobs, and not the application alone: the thing it was
asked to predict is dominated by the modal answer.
This closes the TM-as-gun thread. If a TM is wanted in the bot, a firing gate or a
movement decision is a better fit for a boolean-rule classifier than an aim point -
that is untested and is a different project.
A LIVE GF-MODE ARM WAS NOT RUN (stated as unmeasured): the task pinned one frozen
HEAD binary and HEAD registers the TM gun as radial only; Task 1 already makes GF
the unpromising candidate.
HARNESS FIX WORTH KEEPING: `tools/ab/which_gun_arm_env.sh` left the TARGET gun
unset, so with the now-Pattern-only default it silently fell back to the FULL rack
- an arm could appear to test a single gun while actually running the whole rack.
It now emits `TR_RACK_<GUN>=both` for the target and `=off` for all 14 others.
(Earlier which-gun results are unaffected: they ran before the Pattern-only default,
or - as in the melee/1v1 campaign - set the explicit `=both` themselves.)
tm_pattern.nim gains a per-class confusion matrix (warm samples only) to support the
majority baseline; no behaviour change. Adds Round 4 to
tm_pattern_sweep_results.md with both tasks and the interpretation rule.
798 lines
41 KiB
Markdown
798 lines
41 KiB
Markdown
# TM pattern gun — discrete-target sweep results
|
||
|
||
Date: 2026-09-21. Author: background worker (executor-heavy).
|
||
Artifacts implementing this: `common_libs/guns/tm_pattern.nim`,
|
||
`common_libs/tests/sweep_tm_pattern.nim`. Do not commit.
|
||
|
||
## What was built
|
||
|
||
`tm_pattern.nim` is a NEW gun (the old `guns/tsetlin.nim` is untouched). It
|
||
attacks both the REPRESENTATION and the TARGET as the brief asked:
|
||
|
||
* **Base**: `forecastLinear` (the exact self-consistent forecast `LinearGun`
|
||
uses). GF class 0 (centre) reproduces the Linear gun byte-for-byte, so any
|
||
measured difference is attributable to the TM.
|
||
* **Target**: a discrete multi-class GUESS-FACTOR BUCKET — which lateral escape
|
||
sector (in max-escape-angle units) the enemy occupied at the tick the bullet
|
||
would have reached the BASE fire distance. 5 or 9 classes.
|
||
* **Label**: read from a per-tick ring of our own recorded enemy positions at the
|
||
base arrival tick, NOT from `FeedbackEvent.actualXY`. Under the shipped
|
||
`bmPath` metric `actualXY` is the closest-approach point on the gun's OWN aim
|
||
ray, which biases the label toward the gun's own last output; the ring gives a
|
||
clean, metric-independent label.
|
||
* **Features**: 40 hand-built binary/bucketed motion features (lateral-velocity
|
||
sign over 3 ticks, turn-rate sign over 3 ticks, time since reversal, lateral
|
||
magnitude, speed/distance/flight-time bands, four per-wall proximity bits,
|
||
radial-fraction band, energy band, heading relative to LOS, approach sign).
|
||
* **TM core**: compact self-contained Granmo Table 2/3 with the corrected
|
||
feedback rules and Eq. 6 empty-clause bootstrap (same corrected core as
|
||
tsetlin.nim / tm_selector.nim, re-derived at 40-bit width).
|
||
* **Per-enemy / freshness**: a fresh net per gun instance; the net and history
|
||
reset if the target id changes. Each offline round is replayed with a fresh
|
||
instance (cold every battle, overfit within the battle).
|
||
|
||
Config overrides used in the final run: `-d:TM_CONF_MARGIN_DEF=0.25
|
||
-d:TM_SHRINK_DEF=0.5` (confidence gate + shrink). Defaults are 0.0 / 1.0
|
||
(= raw argmax). Compile-time knobs: `TM_CLASSES`, `TM_NCLAUSES`, `TM_NSTATES`,
|
||
`TM_S_DEF`, `TM_MIN_OBS`, `TM_CONF_MARGIN_DEF`, `TM_SHRINK_DEF`, `TM_GF_MODE`
|
||
(hard|soft), `TM_SOFT_BETA_DEF`.
|
||
|
||
## How to reproduce
|
||
|
||
```
|
||
nim c --path:common_libs -d:release \
|
||
-d:TM_CONF_MARGIN_DEF=0.25 -d:TM_SHRINK_DEF=0.5 \
|
||
-o:/tmp/sweep_tm_pattern common_libs/tests/sweep_tm_pattern.nim
|
||
/tmp/sweep_tm_pattern --set=real --seeds=3 --metric=path \
|
||
--variants=linear,tsetlin,tmpat,tmpat_shuf
|
||
/tmp/sweep_tm_pattern --set=real --seeds=3 --metric=point \
|
||
--variants=linear,tsetlin,tmpat,tmpat_shuf
|
||
```
|
||
|
||
Raw outputs: `/tmp/final_path_s3.txt`, `/tmp/final_point_s3.txt`,
|
||
`/tmp/final_ungated_path_s3.txt`, `/tmp/syn_*`.
|
||
|
||
## Metric
|
||
|
||
EARLY = resolutions in the first 100 ticks of each round (a cold TM every
|
||
round). OVERALL = whole fixture. Pooled over all rounds / fixtures / seeds.
|
||
`TMPatternShuf` = identical gun/encoding/cadence but the training label is a
|
||
uniform-random class (the mandatory shuffled-feedback control). Per-run = one
|
||
fixture × one seed (Linear is deterministic and replicated across seeds for
|
||
pairing). Significance = exact two-sided paired sign test, 18 pairs.
|
||
|
||
## The core result — the discrete target IS learnable, but does not beat the base
|
||
|
||
Online classification accuracy of the GF bucket (warm predictions only,
|
||
seeds=1, n ≈ 1.26 M for each arm):
|
||
|
||
| arm | correct/total | accuracy |
|
||
|---|---|---|
|
||
| TMPattern (real labels) | 578722/1258488 | **46.0%** |
|
||
| TMPatternShuf (random labels) | 246733/1231116 | **20.0%** (chance) |
|
||
|
||
So the Tsetlin Machine genuinely learns the discrete target (2.3× chance). The
|
||
representation mismatch was real and is fixed. The problem is that the target
|
||
is not aligned with what wins the metric.
|
||
|
||
### Real DrussGT fixtures, bmPath (shipped), seeds=3
|
||
|
||
| variant | early | overall |
|
||
|---|---|---|
|
||
| Linear | 34.0% (6358/18715) | 24.3% (58297/239943) |
|
||
| Tsetlin (default) | 22.3% (12344/55463) | 20.3% (145828/719205) |
|
||
| **TMPattern (gated)** | **27.9% (15514/55535)** | **22.0% (158658/719681)** |
|
||
| TMPatternShuf | 28.7% (16049/55969) | 19.4% (139639/719790) |
|
||
|
||
Paired sign tests (18 runs; ranges overlap, so the paired test is the test):
|
||
|
||
* Linear > TMPattern: early 15/18 p=0.0075; overall 15/18 p=0.0075. **Significantly
|
||
worse than Linear.**
|
||
* TMPattern > TMPatternShuf: early 10/8 p=0.81 (tie); overall 17/1 p=0.0001.
|
||
**Learning is real but shows up mainly in the whole-round aggregate, not early.**
|
||
* TMPattern > Tsetlin: early 17/1 p=0.0001; overall 12/6 p=0.24. **Beats the
|
||
default TM gun early, ties overall.**
|
||
|
||
Per-run distributions (mean [min,max], 18 runs):
|
||
Linear early 35.52 [26.03,50.55] / overall 26.86 [9.79,43.36];
|
||
Tsetlin 25.76 [19.79,47.68] / 21.99 [9.64,31.34];
|
||
TMPattern 31.18 [20.17,55.11] / 24.45 [11.24,37.78];
|
||
Shuf 31.42 [20.72,53.49] / 21.07 [9.55,32.65].
|
||
|
||
### Raw ungated hard argmax (margin 0.0, shrink 1.0), bmPath, seeds=3
|
||
|
||
| variant | early | overall |
|
||
|---|---|---|
|
||
| Linear | 34.0% | 24.3% |
|
||
| Tsetlin | 22.3% | 20.3% |
|
||
| TMPattern | 21.2% (11783/55664) | 18.6% (133465/719432) |
|
||
| TMPatternShuf | 15.2% (8710/57169) | 8.3% (59903/720583) |
|
||
|
||
TMPattern > Shuf 18/18 p<0.0001 on BOTH early and overall; TMPattern < Linear
|
||
3/15 p=0.0075 on both. The raw classifier is a clear, decisive learner and a
|
||
clear loser to the Linear base: applying an argmax GF bucket costs ~13 pp early.
|
||
|
||
### Real DrussGT fixtures, bmPoint, seeds=3 (gated)
|
||
|
||
| variant | early | overall |
|
||
|---|---|---|
|
||
| Linear | 7.2% (1480/20498) | 4.7% (11277/241423) |
|
||
| Tsetlin | 7.0% (4403/62617) | 4.8% (34588/724717) |
|
||
| TMPattern | 7.2% (4441/61842) | 4.6% (33341/724556) |
|
||
| TMPatternShuf | 6.6% (4112/61979) | 3.4% (24847/724655) |
|
||
|
||
Online accuracy 50.8%. TMPattern is statistically indistinguishable from Linear
|
||
here (early per-run mean 13.67 vs 13.61; overall 5.70 vs 5.77) and beats its
|
||
control on overall — i.e. on the arrival-time metric the correction is neutral,
|
||
not harmful.
|
||
|
||
### Synthetic fixtures (known rules) — the mechanism works when motion is predictable
|
||
|
||
bmPath, seeds=1, soft readout K=9: Linear early 76.3% / overall 71.4%;
|
||
TMPattern early 76.6% / overall 71.8%; Shuf early 76.6% / overall 67.8%.
|
||
Per-fixture gains vs Linear: wall-bounce 567 vs 537, energy-threshold-turner 332
|
||
vs 319; loss: constant-velocity 417 vs 431.
|
||
|
||
bmPoint, seeds=1, gated hard K=5: Linear 66.4% / 59.6%; TMPattern 66.8% / 60.6%;
|
||
Shuf 55.7% / 50.1%. Energy-threshold-turner 268 vs 212, wall-bounce 585 vs 573.
|
||
|
||
## Verdict
|
||
|
||
* **Learning**: YES, decisively. The discrete-target TM predicts the GF bucket
|
||
far above chance (46% vs 20%) and beats its shuffled control (ungated 18/18,
|
||
p<0.0001). The "regression is a TM mismatch" diagnosis was correct.
|
||
* **Beats Linear**: NO on the real surfers under bmPath (significantly worse,
|
||
p=0.0075). Neutral under bmPoint. Matches/slightly beats Linear only on
|
||
synthetic motion whose future is genuinely predictable.
|
||
* **Best configuration found**: gated hard K=5, `TM_CONF_MARGIN=0.25`,
|
||
`TM_SHRINK=0.5` → 27.9% early / 22.0% overall (bmPath, real), +5.6 pp early /
|
||
+1.7 pp overall vs the default TM gun, but 6.1 pp early / 2.3 pp overall
|
||
behind Linear.
|
||
|
||
## MEASURED vs INFERRED
|
||
|
||
MEASURED: every number in the tables above (pooled hits/shots, per-run
|
||
distributions, paired sign tests, online classification accuracies). The
|
||
position-ring label is our own recorded history at the base arrival tick; the
|
||
shuffled control replaces only the label class with a uniform random draw.
|
||
|
||
INFERRED: that the residual loss on real surfers is because the linear lead is
|
||
already the modal GF (label histogram is centred: real labels
|
||
[3.8,6.6,17.6,6.6,3.5]×10⁵ for 5 classes) and the enemy's per-tick lateral
|
||
reversal sign is not predictable enough from the 40 context bits to make a
|
||
corrective excursion net-positive. Not directly measured.
|
||
|
||
## What to try next (not done, time-boxed out)
|
||
|
||
1. **Radial target instead of angular.** `forecastRadialBlend` work showed the
|
||
dominant surfer error is range-holding (radial), not angle. A TM classifier
|
||
over a RADIAL displacement bucket applied as an aim-distance correction
|
||
targets the error the base actually has room to fix, and should matter most
|
||
under bmPoint.
|
||
2. **Binary reversal with a two-candidate aim** (brief candidate #1, unimplemented):
|
||
predict "will the enemy reverse lateral direction before arrival?" and choose
|
||
between the linear lead and a reversed lead. Same GF family, but a 2-class
|
||
target is far more data-efficient; expected neutral given the GF result.
|
||
3. **Condition a genuinely weaker base.** The measured wall says the deficit is
|
||
the baseline; the Linear base leaves the TM no headroom. Feeding the TM the
|
||
residual of `forecastRadialBlend` (a base that is worse on straight-liners but
|
||
range-correct on surfers) is where a learned correction could plausibly pay.
|
||
4. **Richer context.** 46% accuracy leaves room; the current context lacks the
|
||
enemy's own recent GF history / segmentation that KNN/DecayGF exploit.
|
||
|
||
---
|
||
|
||
# ROUND 2 — fix the base, then try a target Linear cannot predict
|
||
|
||
Date: 2026-09-22. Artifacts: `common_libs/guns/tm_pattern.nim` (extended),
|
||
`common_libs/tests/sweep_tm_pattern.nim` (extended). Raw outputs:
|
||
`/tmp/tm2_real_path_s3.txt`, `/tmp/tm2_real_point_s3.txt`,
|
||
`/tmp/tm_best_path_s3.txt`. All numbers below are MEASURED unless a line says
|
||
INFERRED.
|
||
|
||
Round 1's "best" gun predicted the LATERAL GF bucket. Round 2 adds a RADIAL
|
||
head (aim-distance correction) and a binary REVERSAL head (flip the GF sign),
|
||
both on the same 40-bit context and the same TM core, selected by a runtime
|
||
`targetMode` (`tmGF` | `tmRadial` | `tmReversal`). The shuffled-feedback control
|
||
now randomises only the head the active mode is claiming.
|
||
|
||
## Task 1 — the base is EXACTLY Linear (premise refuted)
|
||
|
||
`tm_pattern`'s base is `forecastLinear`, which already iterates the flight time
|
||
(5-iteration fixed point, same as `LinearGun`). The only deviation from
|
||
`LinearGun` was the wall clamp: the base path clamped to `[BotRadius, W-BotRadius]`
|
||
(17 px inset) instead of `LinearGun`'s `[0, W]`. Added a `forceBase` flag and a
|
||
`TMPatternBase` variant, and made the zero-correction path return `f.x, f.y`
|
||
with the exact `[0, W]` clamp.
|
||
|
||
Real DrussGT fixtures, bmPath, seeds=3, 18 fixture×seed runs, 77 rounds pooled:
|
||
|
||
| variant | early | overall |
|
||
|---|---|---|
|
||
| Linear | 34.0% (6358/18715) | 24.3% (58297/239943) |
|
||
| LinearOldClamp (pre-fix base, BotRadius inset) | 34.2% (6350/18592) | 24.7% (59217/239891) |
|
||
| **TMPatternBase (forceBase, exact Linear clamp)** | **34.0% (6358/18715)** | **24.3% (58297/239943)** |
|
||
|
||
Paired sign test Linear vs TMPatternBase: **18 ties, 0 wins each, p=1.000** on
|
||
both early and overall; every per-run row (hits, shots, per-bin) is byte-for-byte
|
||
identical. Per-run means identical: early 35.52%, overall 26.86%.
|
||
|
||
Verdict: **the base was never behind.** It IS `LinearGun` to the last floating
|
||
point. The earlier "one-shot, non-iterating baseline" finding belonged to the OLD
|
||
`guns/tsetlin.nim`, not to `tm_pattern`. The clamp fix is a wash (the old inset
|
||
was marginally BETTER on overall: 24.7% vs 24.3%), so there is **zero baseline
|
||
headroom** to recover: the entire deficit vs Linear is the TM's corrective
|
||
excursions.
|
||
|
||
## Task 2 — radial target: a structural no-op under bmPath, a real WIN under bmPoint
|
||
|
||
**Structural fact (from `virtual_bullets.nim`, INFERRED then confirmed):** under
|
||
`bmPath` a bullet flies along the aim RAY until it leaves the arena; the aim
|
||
distance only sets `fireDist` (used for the tie-break probe), it does NOT change
|
||
the ray. Moving the aim point radially along the base bearing therefore cannot
|
||
change a `bmPath` hit. Confirmed exactly: on the synthetic set, `TMRadial` vs
|
||
`Linear` scored **8/8 exact ties, p=1.000** under bmPath.
|
||
|
||
Under `bmPoint` the bullet resolves when `travelDist >= fireDist`, so the aim
|
||
distance selects the arrival tick — the radial degree of freedom is live.
|
||
|
||
### bmPath (shipped), real, default config, seeds=3
|
||
|
||
| variant | early | overall |
|
||
|---|---|---|
|
||
| Linear | 34.0% | 24.3% |
|
||
| TMRadial | 33.9% (19038/56203) | 24.1% (173577/719860) |
|
||
| TMRadialShuf | 33.8% (18976/56085) | 24.3% (175154/719781) |
|
||
|
||
Per-run means (n=18): Linear 35.52/26.86; TMRadial 35.44/26.70; Shuf 35.40/26.92.
|
||
Paired sign tests: Linear vs TMRadial early 13/5 p=0.096; overall 15/3 p=0.0075
|
||
(a tiny systematic LOSS, traceable to the `BotRadius` clamp perturbing the ray
|
||
near walls when `radOffset != 0`). TMRadial vs TMRadialShuf early 8/8 p=1.000.
|
||
|
||
**Verdict bmPath: no gain.** Radial is a structural no-op; the shipped metric
|
||
therefore cannot reward Task 2.
|
||
|
||
### bmPoint, real, default config, seeds=3
|
||
|
||
| variant | early | overall |
|
||
|---|---|---|
|
||
| Linear | 7.2% (1480/20498) | 4.7% (11277/241423) |
|
||
| Tsetlin (default gun) | 7.0% (4403/62617) | 4.8% (34588/724717) |
|
||
| **TMRadial** | **9.4% (6013/63785)** | **5.8% (42079/726652)** |
|
||
| TMRadialShuf (control) | 7.0% (4329/61691) | 3.6% (26116/724594) |
|
||
|
||
Per-run means (n=18): Linear 13.61/5.77; Tsetlin 13.13/5.82; TMRadial
|
||
15.32/7.06; Shuf 12.86/4.34. Paired sign tests:
|
||
|
||
* **TMRadial > Linear: early 14/4 p=0.0309; overall 17/1 p=0.0001.**
|
||
* **TMRadial > TMRadialShuf: early 17/1 p=0.0001; overall 18/0 p<0.0001.**
|
||
* TMRadial > Tsetlin: early 17/1 p=0.0001; overall 15/3 p=0.0075.
|
||
* TMRadialShuf vs Linear: early 10/8 p=0.81; overall 12/6 p=0.24 (control sits
|
||
at baseline).
|
||
|
||
Online accuracy of the radial head: 48.8% (511949/1049453) vs **19.9%** shuffled
|
||
chance under bmPoint (46.6% vs 20.0% under bmPath). Radial label histogram
|
||
(raw, seeds=1) = [9058, 5392, 10305, 2236, 1112]: strongly asymmetric — surfers
|
||
are often NEARER than the base constant-velocity prediction at the arrival tick
|
||
(the base overshoots range on range-holders), so class 0 (aim 45–60 px short)
|
||
dominates. That is the mechanism behind the win.
|
||
|
||
Caveat (MEASURED): `labelMiss` is much higher for radial mode (~4.3 M vs ~1.7 M
|
||
for GF) because aiming SHORT resolves the bullet before the base arrival tick,
|
||
so the arrival-tick ring sample is not yet recorded. The radial head is trained
|
||
only on resolvable samples; the win is nonetheless measured on the metric, which
|
||
is label-independent. A deferred-label fix would be the next refinement.
|
||
|
||
## Task 3 — binary reversal: not learnable, and the flip is a no-op
|
||
|
||
Label: net heading turn over the flight opposes the direction the enemy was
|
||
turning at fire time (threshold 10°). Readout: train GF as in round 1, and if
|
||
reversal is predicted, negate the GF correction (`TM_REV_GAIN=1.0`).
|
||
|
||
Label base rate (real, seeds=1, 1 round/fixture): `rev=[24772, 2673]` → the
|
||
positive class is only **9.7%**. The head scores 86.8% (23030/26517) — **below
|
||
the 90.3% majority-class base rate**, i.e. it is not detecting reversals at all,
|
||
only predicting "no reversal". (The shuffled control is 50.2% because its labels
|
||
are balanced.)
|
||
|
||
### bmPath (shipped), real, default config, seeds=3
|
||
|
||
| variant | early | overall |
|
||
|---|---|---|
|
||
| Linear | 34.0% | 24.3% |
|
||
| Tsetlin (round-1 measurement) | 22.3% | 20.3% |
|
||
| TMReversal | 19.5% (11062/56704) | 18.4% (132356/719725) |
|
||
| TMReversalShuf | 19.1% (10925/57063) | 17.8% (128450/720126) |
|
||
|
||
Per-run means: TMReversal 23.45/20.32; Shuf 22.51/20.01. Paired: TMReversal vs
|
||
Shuf early 12/6 p=0.238; overall 8/10 p=0.815 → **no learning effect on hits.**
|
||
|
||
### bmPath, best config (margin=0.25, shrink=0.5), real, seeds=3
|
||
|
||
| variant | early | overall |
|
||
|---|---|---|
|
||
| Linear | 34.0% | 24.3% |
|
||
| TMPattern (gated GF, round-1 best) | 28.3% (15815/55902) | 22.2% (159674/719742) |
|
||
| TMReversal (gated GF + flip) | 28.4% (15944/56067) | 22.3% (160163/719761) |
|
||
| TMReversalShuf | 28.7% (16146/56195) | 21.7% (155973/719841) |
|
||
|
||
Paired: TMPattern vs TMReversal early 11/7 p=0.481, overall 7/11 p=0.481 — the
|
||
flip changes nothing. TMReversal vs Shuf overall 13/5 p=0.096 (not significant).
|
||
|
||
**Verdict: clean negative.** The reversal target as defined is too rare to learn
|
||
(head below the majority baseline), and using it to flip the GF sign is neutral
|
||
to slightly negative on hits. Do not pursue this label; if revisited, balance the
|
||
positive class (per-tick reversal events, or predict the arrival turn direction
|
||
rather than "a reversal happened").
|
||
|
||
## Round-2 overall verdict
|
||
|
||
* On **bmPath (the shipped metric): the TM is NOT competitive with Linear.**
|
||
Base = Linear exactly; radial is a structural no-op; gated GF is significantly
|
||
worse (28.3%/22.2% vs 34.0%/24.3%, p=0.0075); reversal does nothing. The linear
|
||
lead is already the best aim DIRECTION on these surfers and every learned
|
||
angular excursion loses.
|
||
* On **bmPoint: the TM now BEATS Linear and the default Tsetlin gun.**
|
||
`TMRadial` (radial head, `TM_RADIAL_RANGE=60`, `TM_RAD_MARGIN=0.25`, 5 classes,
|
||
gated), 9.4%/5.8% vs Linear 7.2%/4.7% (overall 17/1, p=0.0001) and vs Tsetlin
|
||
7.0%/4.8% (overall 15/3, p=0.0075), with its shuffled control at 7.0%/3.6%.
|
||
This is the first configuration in the whole TM effort that beats both
|
||
baselines with a control-validated margin.
|
||
* **Learning vs controls:** radial head 48.8% vs 19.9% chance (bmPoint); GF head
|
||
reproduces round 1 (48.5% vs ~20%); reversal head does not beat majority.
|
||
|
||
Best configuration if the arrival-time metric is what matters: **TMRadial**. Best
|
||
configuration under the shipped bmPath: **do nothing — keep the Linear base**. The
|
||
evidence says the next step under bmPath is not another bucket target but either a
|
||
richer DIRECTION representation (segmentation / pattern matching, as KNN and
|
||
DecayGF use) or a metric that exposes the radial degree of freedom.
|
||
|
||
Per-enemy specialisation / freshness (MEASURED, unchanged from round 1): each
|
||
offline round is replayed with a FRESH gun instance and the gun calls
|
||
`resetLearning` if the target id changes mid-battle. The offline fixtures are
|
||
single-target, so the mid-battle reset never fires there; its effect is untested
|
||
by these numbers. There is no persistence across battles.
|
||
|
||
## MEASURED vs INFERRED (round 2)
|
||
|
||
* MEASURED: every table, per-run mean, paired sign test, online accuracy, label
|
||
base rate, and the exact `TMPatternBase`/`Linear` byte-for-byte identity.
|
||
* MEASURED: the bmPath radial no-op (synthetic exact ties; real bmPath tiny
|
||
clamp-induced loss).
|
||
* INFERRED: that bmPath ignores radial distance because it flies a ray — read
|
||
from `virtual_bullets.nim`, then confirmed by the synthetic tie.
|
||
* INFERRED: that the radial win comes from surfers being NEARER than the base
|
||
prediction (range-holding), supported by the asymmetric radial label histogram
|
||
but not separately modelled.
|
||
|
||
---
|
||
|
||
# ROUND 3 — ABLATION: does the radial win even need the Tsetlin Machine?
|
||
|
||
Date: 2026-09-22. Artifacts: `common_libs/guns/radial_offset.nim` (new),
|
||
`common_libs/tests/sweep_radial_offset.nim` (new), `common_libs/guns/tm_pattern.nim`
|
||
(label/readout instrumentation only). Raw outputs: `/tmp/ro_point_s3.txt`,
|
||
`/tmp/ro_path_s3.txt` (`--set=real --seeds=3`), `/tmp/ro_point_s1.txt`.
|
||
|
||
## The question
|
||
|
||
Commit 589a230 retracted the radial head's "conditional learning" claim: the
|
||
radial head's online accuracy (57.0%) is AT/BELOW the 58.2% majority baseline, and
|
||
the bmPoint metric win was attributed to a **NET-POSITIVE AVERAGE RADIAL SHIFT**.
|
||
If that is what it is, a FIXED radial shift should reproduce the win with no
|
||
learning, no 0.36 ms/tick cost and no risk.
|
||
|
||
`radial_offset.nim` is exactly that fixed shift and nothing else: the exact
|
||
`forecastLinear` bearing, aim distance `f.dist * scale + offsetPx`, and the SAME
|
||
`[BotRadius, arena-BotRadius]` clamp the TM's corrective path uses. It is
|
||
stateless, so two runs are identical by construction. The sweep runs **Linear**,
|
||
a grid of constants, **TMRadial** and **TMRadialShuf** under BOTH metrics, using
|
||
the identical fixture set / round splitting / per-fixture fresh gun methodology
|
||
as `sweep_tm_pattern.nim`.
|
||
|
||
**Harness sanity check (MEASURED):** `TMRadial` in this report reproduces the
|
||
committed numbers exactly — bmPoint early 9.1% / overall 5.7%, vs Linear early
|
||
16/2 p=0.0013 and overall 18/0 p<0.0001, vs Shuf 18/0 p<0.0001. So the ablation
|
||
runs on the same experiment, not a re-derivation.
|
||
|
||
0 dropped bullets in every run (ring never clobbered); `labelMiss=0` and
|
||
`traceMiss=0` for every TM row (the deferred-label fix holds).
|
||
|
||
## bmPoint (seeds=3)
|
||
|
||
Pooled over 77 rounds for the deterministic arms (one pass; replicated across the
|
||
3 seeds only for the paired test) and 231 rounds for the TM arms (3 seeds).
|
||
|
||
| variant | early | overall |
|
||
|---|---|---|
|
||
| Linear | 7.2% (1480/20498) | 4.7% (11277/241423) |
|
||
| `RO_s1.00` (base + BotRadius clamp only) | 7.2% (1481/20626) | 4.7% (11388/241551) |
|
||
| `RO_s0.98` (scale 0.98) | 8.0% (1669/20781) | 5.8% (14033/241677) |
|
||
| **`RO_s0.95`** | **8.8% (1852/21001)** | **6.3% (15299/241890)** |
|
||
| `RO_o-10` (fixed -10 px) | 8.4% (1756/20790) | 5.9% (14336/241689) |
|
||
| **`RO_o-20`** | 7.9% (1658/20939) | **6.4% (15445/241836)** |
|
||
| `RO_o-30` | 5.4% (1130/21117) | 4.9% (11740/242007) |
|
||
| `RO_o-40` | 4.1% (883/21292) | 3.7% (8989/242166) |
|
||
| `RO_o-60` | 4.7% (1019/21655) | 3.0% (7173/242495) |
|
||
| `RO_o+30` (opposite-direction control) | 0.6% (131/20159) | 0.3% (650/241168) |
|
||
| **TMRadial** | **9.1% (5776/63518)** | 5.7% (41232/726357) |
|
||
| TMRadialShuf | 7.0% (4314/61672) | 3.8% (27317/724473) |
|
||
|
||
Per-run means over the 18 fixture×seed runs:
|
||
Linear 13.61 / 5.77; `RO_s0.95` **14.86 / 7.47**; `RO_o-20` 10.35 / 7.45;
|
||
TMRadial **15.08 / 6.89**; Shuf 12.89 / 4.54.
|
||
|
||
Paired sign tests (exact two-sided binomial, 18 pairs):
|
||
|
||
| A | B | metric | nA>B | nB>A | p |
|
||
|---|---|---|---|---|---|
|
||
| RO_s0.95 | Linear | early | 15 | 3 | **0.0075** |
|
||
| RO_s0.95 | Linear | overall | 15 | 3 | **0.0075** |
|
||
| RO_s0.98 | Linear | overall | 18 | 0 | **<0.0001** |
|
||
| RO_o-20 | Linear | early | 15 | 3 | **0.0075** |
|
||
| RO_o-20 | Linear | overall | 15 | 3 | **0.0075** |
|
||
| RO_o-10 | Linear | overall | 18 | 0 | **<0.0001** |
|
||
| RO_o+30 | Linear | both | 0 | 18 | **<0.0001** (wrong direction) |
|
||
| TMRadial | Linear | early | 16 | 2 | **0.0013** |
|
||
| TMRadial | Linear | overall | 18 | 0 | **<0.0001** |
|
||
| TMRadial | Shuf | both | 18 | 0 | **<0.0001** |
|
||
| **RO_s0.95** | **TMRadial** | **early** | **9** | **9** | **1.0000 (tie)** |
|
||
| **RO_s0.95** | **TMRadial** | **overall** | **15** | **3** | **0.0075 (constant wins)** |
|
||
| RO_o-20 | TMRadial | early | 6 | 12 | 0.2379 |
|
||
| RO_o-20 | TMRadial | overall | 12 | 6 | 0.2379 |
|
||
| RO_s0.98 | TMRadial | overall | 6 | 12 | 0.2379 |
|
||
|
||
**MEASURED conclusion (bmPoint):** a constant radial shift of **scale 0.95**
|
||
(or fixed **−20 px**) *ties the learned radial head on EARLY and beats it on
|
||
OVERALL* (7.47% vs 6.89% per-run mean, 15/3 p=0.0075). The learned head's whole
|
||
nominal advantage is its slightly higher early rate (9.1% vs 8.8% pooled,
|
||
15.08% vs 14.86% per-run), and that is a statistical tie (9/9, p=1.0).
|
||
`RO_o+30` (aiming the opposite way) collapses to 0.3%, confirming the direction of
|
||
the effect is real and not a clamp artefact.
|
||
|
||
## bmPath (the SHIPPED metric) — the radial avenue is a dead end
|
||
|
||
| variant | early | overall |
|
||
|---|---|---|
|
||
| Linear | 34.0% (6358/18715) | 24.3% (58297/239943) |
|
||
| `RO_s1.00` (clamp only) | 34.2% (6350/18592) | **24.7%** (59217/239891) |
|
||
| `RO_s0.98` | 34.2% (6360/18623) | 24.6% (58950/239906) |
|
||
| `RO_s0.95` (best bmPoint constant) | 34.0% (6344/18675) | 24.4% (58530/239937) |
|
||
| `RO_o-20` (best bmPoint constant) | 34.1% (6354/18649) | 24.5% (58699/239921) |
|
||
| `RO_o+30` (best bmPath constant) | 34.4% (6357/18498) | **24.8%** (59366/239852) |
|
||
| `RO_s0.80` | 33.6% (6346/18886) | 23.4% (56151/240035) |
|
||
| TMRadial | 33.8% (19002/56211) | 24.0% (173001/719875) |
|
||
| TMRadialShuf | 33.9% (19026/56076) | 24.3% (175262/719776) |
|
||
|
||
Per-run means (n=18): Linear 35.52 / 26.86; `RO_s1.00` 35.71 / **27.23**;
|
||
`RO_s0.95` 35.59 / 26.97; `RO_o-20` 35.63 / 27.04; **TMRadial 35.37 / 26.63**
|
||
(a loss).
|
||
|
||
Paired sign tests (bmPath):
|
||
|
||
* **TMRadial vs Linear: early 5/13 p=0.0963; overall 2/16 p=0.0013 — a
|
||
systematic LOSS.** (Matches the committed Round-2 finding.)
|
||
* `RO_s1.00` vs Linear: overall 18/0 p<0.0001 (+0.4 pp) — this is the pre-existing
|
||
**BotRadius-clamp** effect, *not* the radial shift.
|
||
* Every constant between 0.98 and 0.95 and every fixed offset −10..−20 is within
|
||
±0.2 pp of Linear; `RO_s0.95` (a real +1.6 pp win on bmPoint overall) is only
|
||
+0.1 pp here and *below* the clamp-only arm. Strong shrink (≤0.90, ≤−40 px) is a
|
||
significant loss (e.g. `RO_s0.80` 3/15 p=0.0075 early, 0/18 overall).
|
||
* `RO_o+30` — the OPPOSITE direction to what bmPoint wants — is the best bmPath
|
||
constant (+0.5 pp, 18/0). So the bmPath response to the radial knob is
|
||
**clamp-mediated and direction-insensitive**, i.e. the radial degree of freedom
|
||
is a structural no-op, exactly as Round 2 argued.
|
||
|
||
**MEASURED conclusion (bmPath):** neither the learned radial head nor any
|
||
constant reproduces a real gain. TMRadial is a systematic *loss*. **The whole
|
||
radial avenue cannot help the shipped configuration.**
|
||
|
||
## The radial-label distribution (MEASURED)
|
||
|
||
Pooled `TMRadial` training labels (3 seeds, n=725997), class centres
|
||
−60/−30/0/+30/+60 px:
|
||
`radLabelHist = [415166, 126461, 120325, 44694, 19351]` → **majority class 57.2%**,
|
||
`meanRadDelta = −82.30 px`, `meanAbsRadDelta = 90.13 px`.
|
||
The head's online accuracy is **56.2%**, i.e. at/below that majority.
|
||
The mean APPLIED shift is **−37.61 px** (`radChosenHist =
|
||
[460960, 35411, 249071, 6832, 2994]`).
|
||
|
||
Metric-free raw base radial error (each fired bullet's enemy radius minus
|
||
`forecastLinear`'s fire distance, at the base arrival tick):
|
||
|
||
| fixture | n | mean px | meanAbs px | frac nearer | frac farther |
|
||
|---|---|---|---|---|---|
|
||
| drussgt_vs_crazy | 35959 | −87.6 | 92.7 | 0.693 | 0.053 |
|
||
| drussgt_vs_spinbot | 19830 | −100.4 | 107.7 | 0.769 | 0.062 |
|
||
| drussgt_vs_drussgt | 26060 | −70.8 | 85.8 | 0.627 | 0.142 |
|
||
| tr_drussgt_vs_crazy | 45905 | −99.0 | 104.4 | 0.805 | 0.053 |
|
||
| tr_drussgt_vs_spinbot | 43150 | −92.4 | 96.7 | 0.793 | 0.044 |
|
||
| tr_drussgt_vs_modularbot | 79970 | −79.9 | 90.7 | 0.705 | 0.108 |
|
||
|
||
The enemy is NEARER than the constant-velocity prediction in **63–81%** of fired
|
||
bullets and FARTHER in only **4–14%**, consistently across all six captures and
|
||
both metrics. So the net-short bias is a GENUINE property of these range-holding
|
||
surfers against the constant-velocity base (the base lets range grow
|
||
geometrically; they hold it) — not a one-fixture artefact.
|
||
|
||
**INFERRED:** the mean label (−82 px) is much larger than the OPTIMAL constant
|
||
shift (−20 px). The mean is dominated by large radial errors that miss regardless
|
||
of the shift; the near-miss window is served better by a small shift, and a
|
||
larger shift also resolves the bullet at an earlier tick. The exact reason a
|
||
−20 px shift beats −60 px is not separately modelled.
|
||
|
||
## Does the optimal constant vary by adversary? (MEASURED)
|
||
|
||
bmPoint, best constant PER FIXTURE (in-sample upper bound), with Linear and
|
||
TMRadial for reference:
|
||
|
||
| fixture | best early | best overall | Linear overall | TMRadial overall |
|
||
|---|---|---|---|---|
|
||
| drussgt_vs_crazy | RO_s0.80 (10.1%) | RO_o-10 (16.9%) | 15.8% | 17.0% |
|
||
| drussgt_vs_spinbot | RO_s0.95 (9.9%) | RO_s0.95 (11.5%) | 7.5% | 9.9% |
|
||
| drussgt_vs_drussgt | RO_s1.00 (55.3%) | RO_s0.98 (4.3%) | 3.9% | 4.1% |
|
||
| tr_drussgt_vs_crazy | RO_s0.85 (9.5%) | RO_o-30 (5.8%) | 2.9% | 4.4% |
|
||
| tr_drussgt_vs_modularbot | RO_s0.95 (6.8%) | RO_s0.95 (2.7%) | 1.3% | 2.1% |
|
||
| tr_drussgt_vs_spinbot | RO_s0.80 (6.8%) | RO_s0.95 (7.4%) | 3.3% | 3.7% |
|
||
|
||
**MEASURED:** the per-fixture optimum DOES vary (fixed −10 for `crazy`, −30 for
|
||
`tr_crazy`, scale 0.95 for three others). **But** one GLOBAL constant (0.95)
|
||
still beats TMRadial on pooled and per-run overall bmPoint. So the per-adversary
|
||
variation is not enough to justify the learned head: a fixed 0.95 is already the
|
||
best pooled point-metric arm measured here.
|
||
|
||
## DIRECT VERDICT
|
||
|
||
1. **bmPoint: REPLACE the radial TM with a constant.** The best constant
|
||
(scale 0.95, equivalently fixed −20 px) statistically TIES the TM on early
|
||
(9/9 p=1.0) and BEATS it on overall (15/3 p=0.0075; 7.47% vs 6.89% per-run
|
||
mean). The head never out-classifies its majority baseline (56.2% vs 57.2%)
|
||
and its mean applied shift (−37.6 px) is roughly twice the optimal constant.
|
||
The TM buys nothing a constant does not, and costs 0.36 ms/tick + complexity.
|
||
2. **bmPath (shipped): DEAD END.** No constant and no TM improves it; TMRadial
|
||
is a systematic loss (2/16 p=0.0013), and the only real bmPath effect in the
|
||
table is the BotRadius clamp, which is direction-insensitive. Drop the radial
|
||
mode from any shipped configuration.
|
||
3. **Fragility argument fails.** The optimum does vary per adversary, but a
|
||
single global constant already matches/beats the adaptively-trained head — so
|
||
the TM is not earning its cost even by the "per-adversary adaptation"
|
||
argument (it is cold-every-battle and trains online within the battle, yet
|
||
still loses to the global 0.95).
|
||
|
||
Net: **do not keep the radial TM.** If the arrival metric ever matters, ship the
|
||
stateless constant; for the current shipped metric, the radial mode (and its
|
||
registered gun id 14) is not justified.
|
||
|
||
## MEASURED vs INFERRED (round 3)
|
||
|
||
* MEASURED: every table, pooled rate, per-run mean, paired sign test, label
|
||
histogram, mean/abs radial delta, applied-shift mean, online accuracy, and the
|
||
exact reproduction of the committed TMRadial numbers.
|
||
* MEASURED: the constant-offset arm is stateless (verbatim `forecastLinear`
|
||
bearing, `f.dist*scale+offsetPx`, TM corrective clamp), so its rollout is
|
||
deterministic and its replicated per-seed values are legitimate.
|
||
* MEASURED: `RO_s1.00` isolates the BotRadius clamp — on bmPoint it is identical
|
||
to Linear (7.2/4.7%), on bmPath it is the +0.4 pp arm; the radial shift itself
|
||
adds nothing on bmPath.
|
||
* INFERRED: the explanation of the label-mean (−82 px) vs optimal shift (−20 px)
|
||
gap (large-error tail + earlier resolution tick); the direction claim itself is
|
||
measured (the +30 control collapses).
|
||
* INFERRED (not measured): whether a per-adversary constant would beat a global
|
||
one out-of-sample — the per-fixture optima above are in-sample.
|
||
|
||
## How to reproduce (round 3)
|
||
|
||
```
|
||
nim c --path:common_libs -d:release -o:/tmp/sweep_radial_offset \
|
||
common_libs/tests/sweep_radial_offset.nim
|
||
/tmp/sweep_radial_offset --set=real --metric=point --seeds=3
|
||
/tmp/sweep_radial_offset --set=real --metric=path --seeds=3
|
||
# smoke: --seeds=1
|
||
```
|
||
|
||
|
||
---
|
||
|
||
# ROUND 4 — THE TWO DECISIVE TESTS: the majority-class baseline, and the LIVE A/B
|
||
|
||
Date: 2026-09-22. Author: background worker (executor-heavy). Artifacts: the
|
||
`confusion`/`radConfusion` instrumentation in `common_libs/guns/tm_pattern.nim`
|
||
(+`GunStats.confusion` in `common_libs/tests/sweep_tm_pattern.nim`); the live A/B
|
||
table below. Raw outputs: `/tmp/tmjob/raw_path_s1.txt`, `/tmp/tmjob/gated_path_s1.txt`,
|
||
`/tmp/tmjob/gated_path_s3.txt`, `/tmp/tmjob/whichgun/analyze_out.txt`,
|
||
`/tmp/tmjob/whichgun/analyze_tm_out.txt`. Do NOT commit.
|
||
|
||
The two questions the earlier rounds never answered honestly:
|
||
1. Is the GF-target head actually beating the MAJORITY-CLASS baseline, or only
|
||
the 20% "chance" figure that was never the right comparison?
|
||
2. Has the TM gun ever been A/B'd LIVE against Pattern? (It had not — only
|
||
offline on bmPath/bmPoint, and the offline metric is a poor live predictor.)
|
||
|
||
## TASK 1 — GF head vs its majority class (MEASURED)
|
||
|
||
Method: `sweep_tm_pattern.nim`, committed DrussGT fixtures
|
||
(`--set=real`, 6 fixtures / 77 rounds), position-ring labels, seeds=1, bmPath.
|
||
The new `confusion[true][pred]` matrix is counted over the SAME warm samples the
|
||
`classCorrect/classTotal` accuracy already scores, so the baseline is on the same
|
||
samples. `warmLabelHist` = row sums; `accuracy` = diagonal / total; majority =
|
||
`max(warmLabelHist)/total`. The shuffled control randomises only the claimed
|
||
label. Configs: RAW = the default hard argmax (`TM_CONF_MARGIN=0.0`); GATED =
|
||
round-1/2 "best" (`TM_CONF_MARGIN=0.25`, `TM_SHRINK=0.5`).
|
||
|
||
### RAW head (default hard argmax) — seeds=1, n=1,751,067 warm predictions
|
||
|
||
| | |
|
||
|---|---|
|
||
| warm label histogram `[c0..c4]` | `[254286, 284578, 678879, 297055, 236269]` |
|
||
| majority class | **class 2 (centre)** |
|
||
| majority share | **38.77%** |
|
||
| head accuracy | **36.69%** (642473/1751067) |
|
||
| **margin over majority** | **−2.08 pp (AT/BELOW)** |
|
||
| shuffled control | 20.04% vs 20.12% majority → chance (margin −0.08 pp) |
|
||
|
||
Per-class precision/recall (RAW):
|
||
|
||
| class | true N | pred N | TP | recall | precision |
|
||
|---|---|---|---|---|---|
|
||
| 0 | 254286 | 387266 | 107056 | 42.1% | 27.6% |
|
||
| 1 | 284578 | 214589 | 63224 | 22.2% | 29.5% |
|
||
| 2 | 678879 | 692044 | 336818 | 49.6% | 48.7% |
|
||
| 3 | 297055 | 215960 | 57809 | 19.5% | 26.8% |
|
||
| 4 | 236269 | 241208 | 77566 | 32.8% | 32.2% |
|
||
|
||
### GATED head (margin=0.25, shrink=0.5) — seeds=1, n=1,751,844
|
||
|
||
| | |
|
||
|---|---|
|
||
| warm label histogram | `[255146, 284332, 678780, 297017, 236569]` |
|
||
| majority | class 2, **38.75%** |
|
||
| accuracy | **40.37%** (707199/1751844) |
|
||
| **margin over majority** | **+1.62 pp** |
|
||
| shuffled control | 20.24% vs 20.10% majority → chance (margin +0.13 pp) |
|
||
|
||
GATED per-class precision/recall:
|
||
|
||
| class | true N | pred N | TP | recall | precision |
|
||
|---|---|---|---|---|---|
|
||
| 0 | 255146 | 224606 | 77412 | 30.3% | 34.5% |
|
||
| 1 | 284332 | 124841 | 38752 | **13.6%** | 31.0% |
|
||
| 2 | 678780 | 1092616 | 484804 | 71.4% | 44.4% |
|
||
| 3 | 297017 | 128494 | 38298 | **12.9%** | 29.8% |
|
||
| 4 | 236569 | 181287 | 67933 | 28.7% | 37.5% |
|
||
|
||
seeds=3 GATED (n=5,254,780): accuracy **40.1%** vs **38.8%** majority
|
||
(+1.4 pp); minority recall class 1 **13.9%**, class 3 **13.0%**. The shuffled
|
||
control seeds=3 sits at 20.1% vs 20.1% majority.
|
||
|
||
**What the gate is doing (MEASURED):** the gated head predicts the majority class
|
||
on 1,092,616 / 1,751,844 = **62.4%** of warm ticks (vs 38.8% base rate); the
|
||
shuffled gated control does the same thing (69.4% class-2), which is why its
|
||
"accuracy" is also ~its majority. The real head's +1.6 pp over majority is
|
||
statistically distinguishable at this n (SE ≈ 0.037 pp) but comes almost
|
||
entirely from re-allocating predictions toward the centre while recovering a
|
||
little mass on classes 0/4; recall on classes 1 and 3 is near-trivial.
|
||
|
||
**INTERPRETATION RULE (stated explicitly, and applied):** if the head is at or
|
||
below the majority baseline, the TM is NOT extracting conditional information
|
||
beyond the base rate, and no amount of engineering fixes that. If it is
|
||
meaningfully above majority AND has non-trivial minority-class recall, the TM IS
|
||
learning and the problem is the application, not the machine.
|
||
|
||
**Task 1 verdict (MEASURED):** the RAW classifier — the head's actual output —
|
||
is **2.1 pp BELOW its majority class**. Only a deliberately gated/abstaining
|
||
config edges 1.4–1.6 pp above, by predicting the majority class 62% of the time,
|
||
with minority-class recall of 13–14%. On the rule above this is **at/below the
|
||
majority baseline** — the GF target carries essentially no learnable conditional
|
||
signal beyond the base rate. The earlier "46.0% vs 20% chance" claim was
|
||
misleading on two counts: the correct baseline is **38.8%** (not 20%), and the
|
||
46% figure was measured on the pre-deferred-label-fix sample set (n≈1.26 M); with
|
||
the deferred-label fix (n≈1.75 M, `labelMiss=0`, `traceMiss=0`) the unbiased raw
|
||
head falls below majority. (Cause of the 46→36.7/40.4 change: INFERRED — the
|
||
deferred-label fix added the previously-dropped short-aim samples; those samples
|
||
are the ones the head gets wrong.)
|
||
|
||
## TASK 2 — THE LIVE A/B (MEASURED)
|
||
|
||
This is the first time the TM gun has ever been fought live against the shipped
|
||
gun. The registered, committed candidate is `initTmRadialGun()` (TMPATTERN id 14,
|
||
**radial** target mode) — that is the arm tested; a GF-mode live arm was not run
|
||
because the task specifies ONE frozen HEAD binary and HEAD registers radial only
|
||
(and Task 1 above makes GF the unpromising candidate).
|
||
|
||
* **Frozen binary**: built from `git archive HEAD` → source commit
|
||
**eb74f9b2e39af70caa36c092e94f1ac5eca8b5e6**
|
||
("Ram: finisher-only by default ..."), sha256
|
||
**cb66d66b3bf4c0d3a358ff20133c7e8455cda37e8236109ce32b0283c2fd01be**
|
||
(`/tmp/tmjob/ModularBot_frozen`). My working-tree instrumentation edits are NOT
|
||
in it (HEAD's `tm_pattern.nim` is radial-registered and untouched).
|
||
* **Adversary**: real DrussGT via `tools/robocode_shim/run_bridge_battle.sh`.
|
||
* **Arms**: `onlyPattern` (incumbent), `onlyTMPATTERN` (TM, radial), `onlyLinear`
|
||
(reference = the TM's own base). Every arm forced with `TR_RACK_<GUN>=both`
|
||
and **all 14 other guns `off`**, including TMPATTERN in the non-TM arms (the
|
||
repo `arm_env.sh` was fixed to emit `=both` explicitly so an arm can never fall
|
||
back to the FULL rack).
|
||
* **7 runs × 7 rounds per arm, 7 concurrent**, distinct ports/output files.
|
||
* **Liveness (MEASURED)**: the `[rack] mode=1v1 active=` line is exactly the
|
||
target gun in every run (`PATTERN` / `TMPATTERN` / `LINEAR`) and the
|
||
selected-gun mix is 100% that gun (Pattern 79401 sel, TMPattern 83478 sel,
|
||
Linear 80233 sel). TMPattern fired real bullets every run.
|
||
|
||
### Live results — damage/run is PRIMARY
|
||
|
||
| arm | runs | shots | hits | real % | **dmg/run** | **round wins** | mean round len | real % range | dmg range |
|
||
|---|---|---|---|---|---|---|---|---|---|
|
||
| **onlyPattern** | 7 | 4610 | 495 | **10.74%** | **285** | **25/49** | 1625 | 9.31–12.01 | 236–344 |
|
||
| onlyTMPATTERN | 7 | 3374 | 118 | 3.50% | 71 | 0/49 | 1743 | 2.22–5.41 | 44–104 |
|
||
| onlyLinear | 7 | 3218 | 104 | 3.23% | 61 | 0/49 | 1642 | 1.43–5.14 | 28–106 |
|
||
|
||
Per-run values (`run: hits/shots rate, dmg, roundWins/7, meanTurn`):
|
||
|
||
* `onlyPattern` — r1 `86/716 12.01% 344 4/7 1772`; r2 `67/620 10.81% 274 4/7 1481`;
|
||
r3 `60/534 11.24% 246 0/7 1314`; r4 `73/703 10.38% 292 7/7 1719`;
|
||
r5 `59/634 9.31% 236 0/7 1646`; r6 `74/695 10.65% 299 7/7 1692`;
|
||
r7 `76/708 10.73% 304 3/7 1753`.
|
||
* `onlyTMPATTERN` — r1 `19/437 4.35% 76 0/7 1691`; r2 `14/489 2.86% 68 0/7 1837`;
|
||
r3 `21/490 4.29% 84 0/7 1805`; r4 `16/505 3.17% 73 0/7 1681`;
|
||
r5 `11/477 2.31% 50 0/7 1713`; r6 `11/495 2.22% 44 0/7 1723`;
|
||
r7 `26/481 5.41% 104 0/7 1747`.
|
||
* `onlyLinear` — r1 `19/465 4.09% 76 0/7 1781`; r2 `7/491 1.43% 28 0/7 1674`;
|
||
r3 `25/486 5.14% 106 0/7 1837`; r4 `7/425 1.65% 34 0/7 1447`;
|
||
r5 `14/417 3.36% 56 0/7 1456`; r6 `17/454 3.74% 68 0/7 1638`;
|
||
r7 `15/480 3.12% 60 0/7 1664`.
|
||
|
||
Exact two-sided permutation test on per-run values (7 vs 7, C(14,7)=3432 exact):
|
||
|
||
| A vs B | metric | mean diff (A−B) | exact p |
|
||
|---|---|---|---|
|
||
| onlyPattern vs onlyTMPATTERN | hit-rate pp | **+7.22** | **0.0006** |
|
||
| onlyPattern vs onlyTMPATTERN | dmg | **+213.7** | **0.0006** |
|
||
| onlyPattern vs onlyLinear | hit-rate pp | +7.51 | 0.0006 |
|
||
| onlyPattern vs onlyLinear | dmg | +223.9 | 0.0006 |
|
||
| **onlyTMPATTERN vs onlyLinear** | hit-rate pp | **+0.30** | **0.6591** |
|
||
| **onlyTMPATTERN vs onlyLinear** | dmg | **+10.1** | **0.4376** |
|
||
|
||
**Task 2 verdict (MEASURED):** Pattern **decisively** beats the TM gun live —
|
||
10.74% vs 3.50% hit rate, 285 vs 71 dmg/run, 25 vs **0** round wins of 49,
|
||
p=0.0006 on both hit rate and damage, with **non-overlapping** per-run ranges
|
||
(Pattern 9.31–12.01 vs TM 2.22–5.41). The TM gun is **statistically
|
||
indistinguishable from its own Linear base** (p=0.66 hit rate, p=0.44 damage):
|
||
the whole radial TM correction is a live no-op, exactly as Round 3 predicted for
|
||
the structural-radial dead end. The TM gun cannot be the best 1v1 gun.
|
||
|
||
## DIRECT ANSWER — can the TM gun be the best 1v1 gun?
|
||
|
||
**(c) It loses live AND it sits at/below its majority class — the target carries
|
||
no learnable conditional signal, and that is the reason.**
|
||
|
||
* LIVE: the registered TM (radial) loses decisively to Pattern (3.50% vs 10.74%,
|
||
p=0.0006; 0 vs 25 round wins) and is no better than its own Linear base
|
||
(p=0.66). It was never the best gun live.
|
||
* OFFLINE: the GF head at its honest (raw) readout is 2.1 pp BELOW the 38.8%
|
||
majority class; the gated config is only 1.4–1.6 pp above, by predicting the
|
||
majority 62% of the time and with 13–14% recall on the two minority classes the
|
||
correction would need. The radial head was already measured at/below majority
|
||
in Round 3 (56.2% vs 57.2%).
|
||
* Therefore the TM machine is not the limiting factor in the sense "it learns and
|
||
we aim it wrong" (that would be (b)); the discrete GF/radial targets simply do
|
||
not contain conditional information beyond the base rate for these surfers, and
|
||
the live result agrees with the offline majority-class analysis.
|
||
|
||
## MEASURED vs INFERRED (round 4)
|
||
|
||
* **MEASURED**: the GF head confusion matrices, warm label histograms, majority
|
||
shares, accuracies, margins, per-class precision/recall, and the shuffled
|
||
controls (seeds=1 RAW+GATED, seeds=3 GATED); the fact that RAW is below
|
||
majority and GATED is +1.4–1.6 pp above by majority-class over-prediction.
|
||
* **MEASURED**: the live table (shots, hits, real %, dmg/run, round wins, mean
|
||
round length, per-run ranges, exact permutation p vs onlyPattern and TM vs
|
||
Linear), the frozen-binary commit + sha256, and the per-arm `[rack] active=`
|
||
liveness lines.
|
||
* **MEASURED**: `onlyTMPATTERN` and `onlyLinear` are indistinguishable live
|
||
(p=0.66/0.44) — the radial correction changes nothing in real physics.
|
||
* **INFERRED**: that the drop from the old 46.0% to the current 36.7%/40.4% is
|
||
caused by the deferred-label fix enlarging the sample set (1.26 M → 1.75 M)
|
||
with the short-aim samples the head gets wrong; the current numbers are
|
||
measured, the causal attribution is not separately ablated.
|
||
* **INFERRED (not measured)**: a live GF-mode arm — not run (one frozen HEAD
|
||
binary; HEAD registers radial only). Task 1 makes it the unpromising candidate.
|