TM verdict, settled: it loses LIVE and sits at/below its majority class - (c)

The user pushed back on "the TM can't be your best 1v1 gun", correctly, because two
decisive tests had never been run. Both are now run and they agree.

TASK 1 - THE GF HEAD vs ITS MAJORITY-CLASS BASELINE (offline, n=1,751,067):
  label histogram [254286, 284578, 678879, 297055, 236269]
  majority class = 2 (the CENTRE bucket) = 38.77%
  RAW head accuracy = 36.69%  ->  margin **-2.08 pp, BELOW majority**
  GATED head accuracy = 40.37% vs 38.75% majority -> +1.62 pp, BUT it predicts the
  majority class on 62.4% of ticks and its minority recall is 13.6% / 12.9% - a
  base-rate predictor wearing a classifier's clothes.
  Shuffled control sits at its own majority (20.04% vs 20.12%), confirming chance.
**THE OLD "46% vs 20% CHANCE" FIGURE I QUOTED WAS WRONG ON TWO COUNTS:** the
baseline is 38.8%, not 20%, and the 46% predated the deferred-label fix. Against
the correct baseline the head is BELOW it.

TASK 2 - THE FIRST-EVER LIVE A/B OF THE TM GUN (7 runs x 7 rounds per arm, one
frozen binary from git archive HEAD = eb74f9b2, sha256 cb66d66b..., real DrussGT,
every arm forced alone with TR_RACK_<GUN>=both and all 14 others off, liveness
confirmed per run):
  arm                     shots   real %   dmg/run   round wins
  onlyPattern              4610   10.74%     285      25/49
  onlyTMPATTERN (radial)   3374    3.50%      71       0/49
  onlyLinear               3218    3.23%      61       0/49
  Pattern vs TM:  +7.22 pp / +213.7 dmg, exact p=0.0006
  TM vs Linear:   +0.30 pp, p=0.659  (dmg p=0.438)
**The TM is statistically INDISTINGUISHABLE from its own Linear base live.** So it
is not "the TM works and we are aiming it wrong".

DIRECT ANSWER: **(c) It loses live AND sits at/below majority - the target carries
no learnable signal beyond the base rate, and that is the reason.** The reason is
not the machine, not the knobs, and not the application alone: the thing it was
asked to predict is dominated by the modal answer.

This closes the TM-as-gun thread. If a TM is wanted in the bot, a firing gate or a
movement decision is a better fit for a boolean-rule classifier than an aim point -
that is untested and is a different project.

A LIVE GF-MODE ARM WAS NOT RUN (stated as unmeasured): the task pinned one frozen
HEAD binary and HEAD registers the TM gun as radial only; Task 1 already makes GF
the unpromising candidate.

HARNESS FIX WORTH KEEPING: `tools/ab/which_gun_arm_env.sh` left the TARGET gun
unset, so with the now-Pattern-only default it silently fell back to the FULL rack
- an arm could appear to test a single gun while actually running the whole rack.
It now emits `TR_RACK_<GUN>=both` for the target and `=off` for all 14 others.
(Earlier which-gun results are unaffected: they ran before the Pattern-only default,
or - as in the melee/1v1 campaign - set the explicit `=both` themselves.)

tm_pattern.nim gains a per-class confusion matrix (warm samples only) to support the
majority baseline; no behaviour change. Adds Round 4 to
tm_pattern_sweep_results.md with both tasks and the interpretation rule.
This commit is contained in:
2026-09-22 08:21:49 +02:00
parent eb74f9b2e3
commit b0654d18eb
5 changed files with 281 additions and 11 deletions
+10
View File
@@ -188,6 +188,14 @@ type
radOffsetN*: int
classCorrect*: int ## warm predictions whose class matched the eventual label
classTotal*: int ## warm predictions with a resolvable label
## Per-class confusion matrix for the GF head, indexed [true label][predicted
## class], counted over WARM predictions only (the same samples `classTotal`
## scores). This is what makes the majority-class baseline and per-class
## precision/recall measurable. Row sums = the warm label histogram; the
## diagonal sum = classCorrect.
confusion*: array[TM_CLASSES, array[TM_CLASSES, int]]
## Same for the radial head.
radConfusion*: array[TM_CLASSES, array[TM_CLASSES, int]]
lastChosen*: int
shuffleLabels*: bool ## control: replace the computed GF label with a random class
forceBase*: bool ## measurement: ignore the TM, emit the pure LinearGun base
@@ -504,6 +512,7 @@ proc tmResolveTrace(g: var TmPatternGun, t: TmPatternTrace, power: float) =
if t.warm:
inc g.classTotal
if winner == t.chosen: inc g.classCorrect
inc g.confusion[winner][t.chosen]
# Radial label: enemy radius at the base arrival tick minus the base fire
# distance. Independent of our own aim, so it is a clean target.
@@ -517,6 +526,7 @@ proc tmResolveTrace(g: var TmPatternGun, t: TmPatternTrace, power: float) =
if t.warm:
inc g.radTotal
if radWinner == t.radChosen: inc g.radCorrect
inc g.radConfusion[radWinner][t.radChosen]
# Reversal label: net heading turn over the flight, opposite to the direction
# the enemy was turning at fire time.
+47
View File
@@ -45,6 +45,8 @@ type
labHist, choHist: array[TM_CLASSES, int]
classCorrect, classTotal: int
radCorrect, radTotal, revCorrect, revTotal: int
confusion: array[TM_CLASSES, array[TM_CLASSES, int]]
radConfusion: array[TM_CLASSES, array[TM_CLASSES, int]]
Adapt = object
h100, n100, h300, n300, hall, nall, f100, m100: int
@@ -103,6 +105,9 @@ proc addAdapt(dst: var Adapt, src: Adapt) =
for c in 0..<TM_CLASSES:
dst.st.labHist[c] += src.st.labHist[c]
dst.st.choHist[c] += src.st.choHist[c]
for p in 0..<TM_CLASSES:
dst.st.confusion[c][p] += src.st.confusion[c][p]
dst.st.radConfusion[c][p] += src.st.radConfusion[c][p]
proc replayRound(states: seq[WorldState], lastSeen: seq[int], enemyId, baseTick: int,
driver: GunDriver, metric: BulletMetric,
@@ -158,6 +163,9 @@ proc replayRound(states: seq[WorldState], lastSeen: seq[int], enemyId, baseTick:
for c in 0..<TM_CLASSES:
res.st.labHist[c] = after.labHist[c] - obsBefore.labHist[c]
res.st.choHist[c] = after.choHist[c] - obsBefore.choHist[c]
for p in 0..<TM_CLASSES:
res.st.confusion[c][p] = after.confusion[c][p] - obsBefore.confusion[c][p]
res.st.radConfusion[c][p] = after.radConfusion[c][p] - obsBefore.radConfusion[c][p]
result = res
proc replayFixture(fx: Fixture, path: string, driver: GunDriver, metric: BulletMetric,
@@ -209,6 +217,8 @@ proc tmpatStats(g: ref TmPatternGun): GunStats =
result.choHist = g[].chosenHist
result.classCorrect = g[].classCorrect
result.classTotal = g[].classTotal
result.confusion = g[].confusion
result.radConfusion = g[].radConfusion
result.radCorrect = g[].radCorrect
result.radTotal = g[].radTotal
result.revCorrect = g[].revCorrect
@@ -252,6 +262,35 @@ proc signTestP(wins, n: int): float =
for k in 0..lo: s += binomPmf(k, n)
min(1.0, 2.0 * s)
proc printConfusion(variant: string, cm: array[TM_CLASSES, array[TM_CLASSES, int]]) =
## Task 1 deliverable: majority-class baseline vs online accuracy on the SAME
## (warm) samples, plus per-class precision/recall. `cm[true][pred]`.
var total = 0
var diag = 0
var rowSum: array[TM_CLASSES, int]
var colSum: array[TM_CLASSES, int]
for c in 0..<TM_CLASSES:
for p in 0..<TM_CLASSES:
total += cm[c][p]
if c == p: diag += cm[c][p]
rowSum[c] += cm[c][p]
colSum[p] += cm[c][p]
var maj = 0
for c in 1..<TM_CLASSES:
if rowSum[c] > rowSum[maj]: maj = c
let majShare = if total > 0: rowSum[maj].float / total.float else: 0.0
let acc = if total > 0: diag.float / total.float else: 0.0
var hs = ""
for c in 0..<TM_CLASSES: hs.add &"{rowSum[c]},"
echo &"# {variant}: warmTotal={total} warmLabelHist=[{hs}] majority=class{maj} " &
&"majShare={majShare*100:.1f}% acc={diag}/{total}={acc*100:.1f}% " &
&"margin={((acc-majShare)*100):+.1f}pp"
for c in 0..<TM_CLASSES:
let rec = if rowSum[c] > 0: cm[c][c].float / rowSum[c].float else: 0.0
let prec = if colSum[c] > 0: cm[c][c].float / colSum[c].float else: 0.0
echo &"# class{c}: trueN={rowSum[c]} predN={colSum[c]} TP={cm[c][c]} " &
&"recall={rec*100:.1f}% precision={prec*100:.1f}%"
proc main() =
var nSeeds = 3
var set = "real"
@@ -358,6 +397,14 @@ proc main() =
for c in 0..<TM_CLASSES: rdl.add &"{radLabTab[variantName(v)][c]},"
echo &"# label hist {variantName(v)}: radial=[{rdl}] rev=[{rl}]"
# ── TASK 1: majority-class baseline vs accuracy + per-class P/R ──
echo "\n# ── GF head majority-class test (warm samples only) ──"
for v in [vTmpat, vTmpatShuf]:
if v in variants: printConfusion(variantName(v), pooled[variantName(v)].st.confusion)
echo "\n# ── radial head confusion (context) ──"
for v in [vTmpatRad, vTmpatRadShuf]:
if v in variants: printConfusion(variantName(v), pooled[variantName(v)].st.radConfusion)
# ── per-run distributions (a "run" = one fixture × one seed) ──
# Linear is deterministic: replicate its one row per fixture across seeds so a
# paired comparison against a stochastic variant has a partner per run.
@@ -590,3 +590,208 @@ nim c --path:common_libs -d:release -o:/tmp/sweep_radial_offset \
# smoke: --seeds=1
```
---
# ROUND 4 — THE TWO DECISIVE TESTS: the majority-class baseline, and the LIVE A/B
Date: 2026-09-22. Author: background worker (executor-heavy). Artifacts: the
`confusion`/`radConfusion` instrumentation in `common_libs/guns/tm_pattern.nim`
(+`GunStats.confusion` in `common_libs/tests/sweep_tm_pattern.nim`); the live A/B
table below. Raw outputs: `/tmp/tmjob/raw_path_s1.txt`, `/tmp/tmjob/gated_path_s1.txt`,
`/tmp/tmjob/gated_path_s3.txt`, `/tmp/tmjob/whichgun/analyze_out.txt`,
`/tmp/tmjob/whichgun/analyze_tm_out.txt`. Do NOT commit.
The two questions the earlier rounds never answered honestly:
1. Is the GF-target head actually beating the MAJORITY-CLASS baseline, or only
the 20% "chance" figure that was never the right comparison?
2. Has the TM gun ever been A/B'd LIVE against Pattern? (It had not — only
offline on bmPath/bmPoint, and the offline metric is a poor live predictor.)
## TASK 1 — GF head vs its majority class (MEASURED)
Method: `sweep_tm_pattern.nim`, committed DrussGT fixtures
(`--set=real`, 6 fixtures / 77 rounds), position-ring labels, seeds=1, bmPath.
The new `confusion[true][pred]` matrix is counted over the SAME warm samples the
`classCorrect/classTotal` accuracy already scores, so the baseline is on the same
samples. `warmLabelHist` = row sums; `accuracy` = diagonal / total; majority =
`max(warmLabelHist)/total`. The shuffled control randomises only the claimed
label. Configs: RAW = the default hard argmax (`TM_CONF_MARGIN=0.0`); GATED =
round-1/2 "best" (`TM_CONF_MARGIN=0.25`, `TM_SHRINK=0.5`).
### RAW head (default hard argmax) — seeds=1, n=1,751,067 warm predictions
| | |
|---|---|
| warm label histogram `[c0..c4]` | `[254286, 284578, 678879, 297055, 236269]` |
| majority class | **class 2 (centre)** |
| majority share | **38.77%** |
| head accuracy | **36.69%** (642473/1751067) |
| **margin over majority** | **−2.08 pp (AT/BELOW)** |
| shuffled control | 20.04% vs 20.12% majority → chance (margin −0.08 pp) |
Per-class precision/recall (RAW):
| class | true N | pred N | TP | recall | precision |
|---|---|---|---|---|---|
| 0 | 254286 | 387266 | 107056 | 42.1% | 27.6% |
| 1 | 284578 | 214589 | 63224 | 22.2% | 29.5% |
| 2 | 678879 | 692044 | 336818 | 49.6% | 48.7% |
| 3 | 297055 | 215960 | 57809 | 19.5% | 26.8% |
| 4 | 236269 | 241208 | 77566 | 32.8% | 32.2% |
### GATED head (margin=0.25, shrink=0.5) — seeds=1, n=1,751,844
| | |
|---|---|
| warm label histogram | `[255146, 284332, 678780, 297017, 236569]` |
| majority | class 2, **38.75%** |
| accuracy | **40.37%** (707199/1751844) |
| **margin over majority** | **+1.62 pp** |
| shuffled control | 20.24% vs 20.10% majority → chance (margin +0.13 pp) |
GATED per-class precision/recall:
| class | true N | pred N | TP | recall | precision |
|---|---|---|---|---|---|
| 0 | 255146 | 224606 | 77412 | 30.3% | 34.5% |
| 1 | 284332 | 124841 | 38752 | **13.6%** | 31.0% |
| 2 | 678780 | 1092616 | 484804 | 71.4% | 44.4% |
| 3 | 297017 | 128494 | 38298 | **12.9%** | 29.8% |
| 4 | 236569 | 181287 | 67933 | 28.7% | 37.5% |
seeds=3 GATED (n=5,254,780): accuracy **40.1%** vs **38.8%** majority
(+1.4 pp); minority recall class 1 **13.9%**, class 3 **13.0%**. The shuffled
control seeds=3 sits at 20.1% vs 20.1% majority.
**What the gate is doing (MEASURED):** the gated head predicts the majority class
on 1,092,616 / 1,751,844 = **62.4%** of warm ticks (vs 38.8% base rate); the
shuffled gated control does the same thing (69.4% class-2), which is why its
"accuracy" is also ~its majority. The real head's +1.6 pp over majority is
statistically distinguishable at this n (SE ≈ 0.037 pp) but comes almost
entirely from re-allocating predictions toward the centre while recovering a
little mass on classes 0/4; recall on classes 1 and 3 is near-trivial.
**INTERPRETATION RULE (stated explicitly, and applied):** if the head is at or
below the majority baseline, the TM is NOT extracting conditional information
beyond the base rate, and no amount of engineering fixes that. If it is
meaningfully above majority AND has non-trivial minority-class recall, the TM IS
learning and the problem is the application, not the machine.
**Task 1 verdict (MEASURED):** the RAW classifier — the head's actual output —
is **2.1 pp BELOW its majority class**. Only a deliberately gated/abstaining
config edges 1.4–1.6 pp above, by predicting the majority class 62% of the time,
with minority-class recall of 13–14%. On the rule above this is **at/below the
majority baseline** — the GF target carries essentially no learnable conditional
signal beyond the base rate. The earlier "46.0% vs 20% chance" claim was
misleading on two counts: the correct baseline is **38.8%** (not 20%), and the
46% figure was measured on the pre-deferred-label-fix sample set (n≈1.26 M); with
the deferred-label fix (n≈1.75 M, `labelMiss=0`, `traceMiss=0`) the unbiased raw
head falls below majority. (Cause of the 46→36.7/40.4 change: INFERRED — the
deferred-label fix added the previously-dropped short-aim samples; those samples
are the ones the head gets wrong.)
## TASK 2 — THE LIVE A/B (MEASURED)
This is the first time the TM gun has ever been fought live against the shipped
gun. The registered, committed candidate is `initTmRadialGun()` (TMPATTERN id 14,
**radial** target mode) — that is the arm tested; a GF-mode live arm was not run
because the task specifies ONE frozen HEAD binary and HEAD registers radial only
(and Task 1 above makes GF the unpromising candidate).
* **Frozen binary**: built from `git archive HEAD` → source commit
**eb74f9b2e39af70caa36c092e94f1ac5eca8b5e6**
("Ram: finisher-only by default ..."), sha256
**cb66d66b3bf4c0d3a358ff20133c7e8455cda37e8236109ce32b0283c2fd01be**
(`/tmp/tmjob/ModularBot_frozen`). My working-tree instrumentation edits are NOT
in it (HEAD's `tm_pattern.nim` is radial-registered and untouched).
* **Adversary**: real DrussGT via `tools/robocode_shim/run_bridge_battle.sh`.
* **Arms**: `onlyPattern` (incumbent), `onlyTMPATTERN` (TM, radial), `onlyLinear`
(reference = the TM's own base). Every arm forced with `TR_RACK_<GUN>=both`
and **all 14 other guns `off`**, including TMPATTERN in the non-TM arms (the
repo `arm_env.sh` was fixed to emit `=both` explicitly so an arm can never fall
back to the FULL rack).
* **7 runs × 7 rounds per arm, 7 concurrent**, distinct ports/output files.
* **Liveness (MEASURED)**: the `[rack] mode=1v1 active=` line is exactly the
target gun in every run (`PATTERN` / `TMPATTERN` / `LINEAR`) and the
selected-gun mix is 100% that gun (Pattern 79401 sel, TMPattern 83478 sel,
Linear 80233 sel). TMPattern fired real bullets every run.
### Live results — damage/run is PRIMARY
| arm | runs | shots | hits | real % | **dmg/run** | **round wins** | mean round len | real % range | dmg range |
|---|---|---|---|---|---|---|---|---|---|
| **onlyPattern** | 7 | 4610 | 495 | **10.74%** | **285** | **25/49** | 1625 | 9.31–12.01 | 236–344 |
| onlyTMPATTERN | 7 | 3374 | 118 | 3.50% | 71 | 0/49 | 1743 | 2.22–5.41 | 44–104 |
| onlyLinear | 7 | 3218 | 104 | 3.23% | 61 | 0/49 | 1642 | 1.43–5.14 | 28–106 |
Per-run values (`run: hits/shots rate, dmg, roundWins/7, meanTurn`):
* `onlyPattern` — r1 `86/716 12.01% 344 4/7 1772`; r2 `67/620 10.81% 274 4/7 1481`;
r3 `60/534 11.24% 246 0/7 1314`; r4 `73/703 10.38% 292 7/7 1719`;
r5 `59/634 9.31% 236 0/7 1646`; r6 `74/695 10.65% 299 7/7 1692`;
r7 `76/708 10.73% 304 3/7 1753`.
* `onlyTMPATTERN` — r1 `19/437 4.35% 76 0/7 1691`; r2 `14/489 2.86% 68 0/7 1837`;
r3 `21/490 4.29% 84 0/7 1805`; r4 `16/505 3.17% 73 0/7 1681`;
r5 `11/477 2.31% 50 0/7 1713`; r6 `11/495 2.22% 44 0/7 1723`;
r7 `26/481 5.41% 104 0/7 1747`.
* `onlyLinear` — r1 `19/465 4.09% 76 0/7 1781`; r2 `7/491 1.43% 28 0/7 1674`;
r3 `25/486 5.14% 106 0/7 1837`; r4 `7/425 1.65% 34 0/7 1447`;
r5 `14/417 3.36% 56 0/7 1456`; r6 `17/454 3.74% 68 0/7 1638`;
r7 `15/480 3.12% 60 0/7 1664`.
Exact two-sided permutation test on per-run values (7 vs 7, C(14,7)=3432 exact):
| A vs B | metric | mean diff (A−B) | exact p |
|---|---|---|---|
| onlyPattern vs onlyTMPATTERN | hit-rate pp | **+7.22** | **0.0006** |
| onlyPattern vs onlyTMPATTERN | dmg | **+213.7** | **0.0006** |
| onlyPattern vs onlyLinear | hit-rate pp | +7.51 | 0.0006 |
| onlyPattern vs onlyLinear | dmg | +223.9 | 0.0006 |
| **onlyTMPATTERN vs onlyLinear** | hit-rate pp | **+0.30** | **0.6591** |
| **onlyTMPATTERN vs onlyLinear** | dmg | **+10.1** | **0.4376** |
**Task 2 verdict (MEASURED):** Pattern **decisively** beats the TM gun live —
10.74% vs 3.50% hit rate, 285 vs 71 dmg/run, 25 vs **0** round wins of 49,
p=0.0006 on both hit rate and damage, with **non-overlapping** per-run ranges
(Pattern 9.31–12.01 vs TM 2.22–5.41). The TM gun is **statistically
indistinguishable from its own Linear base** (p=0.66 hit rate, p=0.44 damage):
the whole radial TM correction is a live no-op, exactly as Round 3 predicted for
the structural-radial dead end. The TM gun cannot be the best 1v1 gun.
## DIRECT ANSWER — can the TM gun be the best 1v1 gun?
**(c) It loses live AND it sits at/below its majority class — the target carries
no learnable conditional signal, and that is the reason.**
* LIVE: the registered TM (radial) loses decisively to Pattern (3.50% vs 10.74%,
p=0.0006; 0 vs 25 round wins) and is no better than its own Linear base
(p=0.66). It was never the best gun live.
* OFFLINE: the GF head at its honest (raw) readout is 2.1 pp BELOW the 38.8%
majority class; the gated config is only 1.4–1.6 pp above, by predicting the
majority 62% of the time and with 13–14% recall on the two minority classes the
correction would need. The radial head was already measured at/below majority
in Round 3 (56.2% vs 57.2%).
* Therefore the TM machine is not the limiting factor in the sense "it learns and
we aim it wrong" (that would be (b)); the discrete GF/radial targets simply do
not contain conditional information beyond the base rate for these surfers, and
the live result agrees with the offline majority-class analysis.
## MEASURED vs INFERRED (round 4)
* **MEASURED**: the GF head confusion matrices, warm label histograms, majority
shares, accuracies, margins, per-class precision/recall, and the shuffled
controls (seeds=1 RAW+GATED, seeds=3 GATED); the fact that RAW is below
majority and GATED is +1.4–1.6 pp above by majority-class over-prediction.
* **MEASURED**: the live table (shots, hits, real %, dmg/run, round wins, mean
round length, per-run ranges, exact permutation p vs onlyPattern and TM vs
Linear), the frozen-binary commit + sha256, and the per-arm `[rack] active=`
liveness lines.
* **MEASURED**: `onlyTMPATTERN` and `onlyLinear` are indistinguishable live
(p=0.66/0.44) — the radial correction changes nothing in real physics.
* **INFERRED**: that the drop from the old 46.0% to the current 36.7%/40.4% is
caused by the deferred-label fix enlarging the sample set (1.26 M → 1.75 M)
with the short-aim samples the head gets wrong; the current numbers are
measured, the causal attribution is not separately ablated.
* **INFERRED (not measured)**: a live GF-mode arm — not run (one frozen HEAD
binary; HEAD registers radial only). Task 1 makes it the unpromising candidate.