TM verdict, settled: it loses LIVE and sits at/below its majority class - (c)
The user pushed back on "the TM can't be your best 1v1 gun", correctly, because two
decisive tests had never been run. Both are now run and they agree.
TASK 1 - THE GF HEAD vs ITS MAJORITY-CLASS BASELINE (offline, n=1,751,067):
label histogram [254286, 284578, 678879, 297055, 236269]
majority class = 2 (the CENTRE bucket) = 38.77%
RAW head accuracy = 36.69% -> margin **-2.08 pp, BELOW majority**
GATED head accuracy = 40.37% vs 38.75% majority -> +1.62 pp, BUT it predicts the
majority class on 62.4% of ticks and its minority recall is 13.6% / 12.9% - a
base-rate predictor wearing a classifier's clothes.
Shuffled control sits at its own majority (20.04% vs 20.12%), confirming chance.
**THE OLD "46% vs 20% CHANCE" FIGURE I QUOTED WAS WRONG ON TWO COUNTS:** the
baseline is 38.8%, not 20%, and the 46% predated the deferred-label fix. Against
the correct baseline the head is BELOW it.
TASK 2 - THE FIRST-EVER LIVE A/B OF THE TM GUN (7 runs x 7 rounds per arm, one
frozen binary from git archive HEAD = eb74f9b2, sha256 cb66d66b..., real DrussGT,
every arm forced alone with TR_RACK_<GUN>=both and all 14 others off, liveness
confirmed per run):
arm shots real % dmg/run round wins
onlyPattern 4610 10.74% 285 25/49
onlyTMPATTERN (radial) 3374 3.50% 71 0/49
onlyLinear 3218 3.23% 61 0/49
Pattern vs TM: +7.22 pp / +213.7 dmg, exact p=0.0006
TM vs Linear: +0.30 pp, p=0.659 (dmg p=0.438)
**The TM is statistically INDISTINGUISHABLE from its own Linear base live.** So it
is not "the TM works and we are aiming it wrong".
DIRECT ANSWER: **(c) It loses live AND sits at/below majority - the target carries
no learnable signal beyond the base rate, and that is the reason.** The reason is
not the machine, not the knobs, and not the application alone: the thing it was
asked to predict is dominated by the modal answer.
This closes the TM-as-gun thread. If a TM is wanted in the bot, a firing gate or a
movement decision is a better fit for a boolean-rule classifier than an aim point -
that is untested and is a different project.
A LIVE GF-MODE ARM WAS NOT RUN (stated as unmeasured): the task pinned one frozen
HEAD binary and HEAD registers the TM gun as radial only; Task 1 already makes GF
the unpromising candidate.
HARNESS FIX WORTH KEEPING: `tools/ab/which_gun_arm_env.sh` left the TARGET gun
unset, so with the now-Pattern-only default it silently fell back to the FULL rack
- an arm could appear to test a single gun while actually running the whole rack.
It now emits `TR_RACK_<GUN>=both` for the target and `=off` for all 14 others.
(Earlier which-gun results are unaffected: they ran before the Pattern-only default,
or - as in the melee/1v1 campaign - set the explicit `=both` themselves.)
tm_pattern.nim gains a per-class confusion matrix (warm samples only) to support the
majority baseline; no behaviour change. Adds Round 4 to
tm_pattern_sweep_results.md with both tasks and the interpretation rule.
This commit is contained in:
@@ -590,3 +590,208 @@ nim c --path:common_libs -d:release -o:/tmp/sweep_radial_offset \
|
||||
# smoke: --seeds=1
|
||||
```
|
||||
|
||||
|
||||
---
|
||||
|
||||
# ROUND 4 — THE TWO DECISIVE TESTS: the majority-class baseline, and the LIVE A/B
|
||||
|
||||
Date: 2026-09-22. Author: background worker (executor-heavy). Artifacts: the
|
||||
`confusion`/`radConfusion` instrumentation in `common_libs/guns/tm_pattern.nim`
|
||||
(+`GunStats.confusion` in `common_libs/tests/sweep_tm_pattern.nim`); the live A/B
|
||||
table below. Raw outputs: `/tmp/tmjob/raw_path_s1.txt`, `/tmp/tmjob/gated_path_s1.txt`,
|
||||
`/tmp/tmjob/gated_path_s3.txt`, `/tmp/tmjob/whichgun/analyze_out.txt`,
|
||||
`/tmp/tmjob/whichgun/analyze_tm_out.txt`. Do NOT commit.
|
||||
|
||||
The two questions the earlier rounds never answered honestly:
|
||||
1. Is the GF-target head actually beating the MAJORITY-CLASS baseline, or only
|
||||
the 20% "chance" figure that was never the right comparison?
|
||||
2. Has the TM gun ever been A/B'd LIVE against Pattern? (It had not — only
|
||||
offline on bmPath/bmPoint, and the offline metric is a poor live predictor.)
|
||||
|
||||
## TASK 1 — GF head vs its majority class (MEASURED)
|
||||
|
||||
Method: `sweep_tm_pattern.nim`, committed DrussGT fixtures
|
||||
(`--set=real`, 6 fixtures / 77 rounds), position-ring labels, seeds=1, bmPath.
|
||||
The new `confusion[true][pred]` matrix is counted over the SAME warm samples the
|
||||
`classCorrect/classTotal` accuracy already scores, so the baseline is on the same
|
||||
samples. `warmLabelHist` = row sums; `accuracy` = diagonal / total; majority =
|
||||
`max(warmLabelHist)/total`. The shuffled control randomises only the claimed
|
||||
label. Configs: RAW = the default hard argmax (`TM_CONF_MARGIN=0.0`); GATED =
|
||||
round-1/2 "best" (`TM_CONF_MARGIN=0.25`, `TM_SHRINK=0.5`).
|
||||
|
||||
### RAW head (default hard argmax) — seeds=1, n=1,751,067 warm predictions
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| warm label histogram `[c0..c4]` | `[254286, 284578, 678879, 297055, 236269]` |
|
||||
| majority class | **class 2 (centre)** |
|
||||
| majority share | **38.77%** |
|
||||
| head accuracy | **36.69%** (642473/1751067) |
|
||||
| **margin over majority** | **−2.08 pp (AT/BELOW)** |
|
||||
| shuffled control | 20.04% vs 20.12% majority → chance (margin −0.08 pp) |
|
||||
|
||||
Per-class precision/recall (RAW):
|
||||
|
||||
| class | true N | pred N | TP | recall | precision |
|
||||
|---|---|---|---|---|---|
|
||||
| 0 | 254286 | 387266 | 107056 | 42.1% | 27.6% |
|
||||
| 1 | 284578 | 214589 | 63224 | 22.2% | 29.5% |
|
||||
| 2 | 678879 | 692044 | 336818 | 49.6% | 48.7% |
|
||||
| 3 | 297055 | 215960 | 57809 | 19.5% | 26.8% |
|
||||
| 4 | 236269 | 241208 | 77566 | 32.8% | 32.2% |
|
||||
|
||||
### GATED head (margin=0.25, shrink=0.5) — seeds=1, n=1,751,844
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| warm label histogram | `[255146, 284332, 678780, 297017, 236569]` |
|
||||
| majority | class 2, **38.75%** |
|
||||
| accuracy | **40.37%** (707199/1751844) |
|
||||
| **margin over majority** | **+1.62 pp** |
|
||||
| shuffled control | 20.24% vs 20.10% majority → chance (margin +0.13 pp) |
|
||||
|
||||
GATED per-class precision/recall:
|
||||
|
||||
| class | true N | pred N | TP | recall | precision |
|
||||
|---|---|---|---|---|---|
|
||||
| 0 | 255146 | 224606 | 77412 | 30.3% | 34.5% |
|
||||
| 1 | 284332 | 124841 | 38752 | **13.6%** | 31.0% |
|
||||
| 2 | 678780 | 1092616 | 484804 | 71.4% | 44.4% |
|
||||
| 3 | 297017 | 128494 | 38298 | **12.9%** | 29.8% |
|
||||
| 4 | 236569 | 181287 | 67933 | 28.7% | 37.5% |
|
||||
|
||||
seeds=3 GATED (n=5,254,780): accuracy **40.1%** vs **38.8%** majority
|
||||
(+1.4 pp); minority recall class 1 **13.9%**, class 3 **13.0%**. The shuffled
|
||||
control seeds=3 sits at 20.1% vs 20.1% majority.
|
||||
|
||||
**What the gate is doing (MEASURED):** the gated head predicts the majority class
|
||||
on 1,092,616 / 1,751,844 = **62.4%** of warm ticks (vs 38.8% base rate); the
|
||||
shuffled gated control does the same thing (69.4% class-2), which is why its
|
||||
"accuracy" is also ~its majority. The real head's +1.6 pp over majority is
|
||||
statistically distinguishable at this n (SE ≈ 0.037 pp) but comes almost
|
||||
entirely from re-allocating predictions toward the centre while recovering a
|
||||
little mass on classes 0/4; recall on classes 1 and 3 is near-trivial.
|
||||
|
||||
**INTERPRETATION RULE (stated explicitly, and applied):** if the head is at or
|
||||
below the majority baseline, the TM is NOT extracting conditional information
|
||||
beyond the base rate, and no amount of engineering fixes that. If it is
|
||||
meaningfully above majority AND has non-trivial minority-class recall, the TM IS
|
||||
learning and the problem is the application, not the machine.
|
||||
|
||||
**Task 1 verdict (MEASURED):** the RAW classifier — the head's actual output —
|
||||
is **2.1 pp BELOW its majority class**. Only a deliberately gated/abstaining
|
||||
config edges 1.4–1.6 pp above, by predicting the majority class 62% of the time,
|
||||
with minority-class recall of 13–14%. On the rule above this is **at/below the
|
||||
majority baseline** — the GF target carries essentially no learnable conditional
|
||||
signal beyond the base rate. The earlier "46.0% vs 20% chance" claim was
|
||||
misleading on two counts: the correct baseline is **38.8%** (not 20%), and the
|
||||
46% figure was measured on the pre-deferred-label-fix sample set (n≈1.26 M); with
|
||||
the deferred-label fix (n≈1.75 M, `labelMiss=0`, `traceMiss=0`) the unbiased raw
|
||||
head falls below majority. (Cause of the 46→36.7/40.4 change: INFERRED — the
|
||||
deferred-label fix added the previously-dropped short-aim samples; those samples
|
||||
are the ones the head gets wrong.)
|
||||
|
||||
## TASK 2 — THE LIVE A/B (MEASURED)
|
||||
|
||||
This is the first time the TM gun has ever been fought live against the shipped
|
||||
gun. The registered, committed candidate is `initTmRadialGun()` (TMPATTERN id 14,
|
||||
**radial** target mode) — that is the arm tested; a GF-mode live arm was not run
|
||||
because the task specifies ONE frozen HEAD binary and HEAD registers radial only
|
||||
(and Task 1 above makes GF the unpromising candidate).
|
||||
|
||||
* **Frozen binary**: built from `git archive HEAD` → source commit
|
||||
**eb74f9b2e39af70caa36c092e94f1ac5eca8b5e6**
|
||||
("Ram: finisher-only by default ..."), sha256
|
||||
**cb66d66b3bf4c0d3a358ff20133c7e8455cda37e8236109ce32b0283c2fd01be**
|
||||
(`/tmp/tmjob/ModularBot_frozen`). My working-tree instrumentation edits are NOT
|
||||
in it (HEAD's `tm_pattern.nim` is radial-registered and untouched).
|
||||
* **Adversary**: real DrussGT via `tools/robocode_shim/run_bridge_battle.sh`.
|
||||
* **Arms**: `onlyPattern` (incumbent), `onlyTMPATTERN` (TM, radial), `onlyLinear`
|
||||
(reference = the TM's own base). Every arm forced with `TR_RACK_<GUN>=both`
|
||||
and **all 14 other guns `off`**, including TMPATTERN in the non-TM arms (the
|
||||
repo `arm_env.sh` was fixed to emit `=both` explicitly so an arm can never fall
|
||||
back to the FULL rack).
|
||||
* **7 runs × 7 rounds per arm, 7 concurrent**, distinct ports/output files.
|
||||
* **Liveness (MEASURED)**: the `[rack] mode=1v1 active=` line is exactly the
|
||||
target gun in every run (`PATTERN` / `TMPATTERN` / `LINEAR`) and the
|
||||
selected-gun mix is 100% that gun (Pattern 79401 sel, TMPattern 83478 sel,
|
||||
Linear 80233 sel). TMPattern fired real bullets every run.
|
||||
|
||||
### Live results — damage/run is PRIMARY
|
||||
|
||||
| arm | runs | shots | hits | real % | **dmg/run** | **round wins** | mean round len | real % range | dmg range |
|
||||
|---|---|---|---|---|---|---|---|---|---|
|
||||
| **onlyPattern** | 7 | 4610 | 495 | **10.74%** | **285** | **25/49** | 1625 | 9.31–12.01 | 236–344 |
|
||||
| onlyTMPATTERN | 7 | 3374 | 118 | 3.50% | 71 | 0/49 | 1743 | 2.22–5.41 | 44–104 |
|
||||
| onlyLinear | 7 | 3218 | 104 | 3.23% | 61 | 0/49 | 1642 | 1.43–5.14 | 28–106 |
|
||||
|
||||
Per-run values (`run: hits/shots rate, dmg, roundWins/7, meanTurn`):
|
||||
|
||||
* `onlyPattern` — r1 `86/716 12.01% 344 4/7 1772`; r2 `67/620 10.81% 274 4/7 1481`;
|
||||
r3 `60/534 11.24% 246 0/7 1314`; r4 `73/703 10.38% 292 7/7 1719`;
|
||||
r5 `59/634 9.31% 236 0/7 1646`; r6 `74/695 10.65% 299 7/7 1692`;
|
||||
r7 `76/708 10.73% 304 3/7 1753`.
|
||||
* `onlyTMPATTERN` — r1 `19/437 4.35% 76 0/7 1691`; r2 `14/489 2.86% 68 0/7 1837`;
|
||||
r3 `21/490 4.29% 84 0/7 1805`; r4 `16/505 3.17% 73 0/7 1681`;
|
||||
r5 `11/477 2.31% 50 0/7 1713`; r6 `11/495 2.22% 44 0/7 1723`;
|
||||
r7 `26/481 5.41% 104 0/7 1747`.
|
||||
* `onlyLinear` — r1 `19/465 4.09% 76 0/7 1781`; r2 `7/491 1.43% 28 0/7 1674`;
|
||||
r3 `25/486 5.14% 106 0/7 1837`; r4 `7/425 1.65% 34 0/7 1447`;
|
||||
r5 `14/417 3.36% 56 0/7 1456`; r6 `17/454 3.74% 68 0/7 1638`;
|
||||
r7 `15/480 3.12% 60 0/7 1664`.
|
||||
|
||||
Exact two-sided permutation test on per-run values (7 vs 7, C(14,7)=3432 exact):
|
||||
|
||||
| A vs B | metric | mean diff (A−B) | exact p |
|
||||
|---|---|---|---|
|
||||
| onlyPattern vs onlyTMPATTERN | hit-rate pp | **+7.22** | **0.0006** |
|
||||
| onlyPattern vs onlyTMPATTERN | dmg | **+213.7** | **0.0006** |
|
||||
| onlyPattern vs onlyLinear | hit-rate pp | +7.51 | 0.0006 |
|
||||
| onlyPattern vs onlyLinear | dmg | +223.9 | 0.0006 |
|
||||
| **onlyTMPATTERN vs onlyLinear** | hit-rate pp | **+0.30** | **0.6591** |
|
||||
| **onlyTMPATTERN vs onlyLinear** | dmg | **+10.1** | **0.4376** |
|
||||
|
||||
**Task 2 verdict (MEASURED):** Pattern **decisively** beats the TM gun live —
|
||||
10.74% vs 3.50% hit rate, 285 vs 71 dmg/run, 25 vs **0** round wins of 49,
|
||||
p=0.0006 on both hit rate and damage, with **non-overlapping** per-run ranges
|
||||
(Pattern 9.31–12.01 vs TM 2.22–5.41). The TM gun is **statistically
|
||||
indistinguishable from its own Linear base** (p=0.66 hit rate, p=0.44 damage):
|
||||
the whole radial TM correction is a live no-op, exactly as Round 3 predicted for
|
||||
the structural-radial dead end. The TM gun cannot be the best 1v1 gun.
|
||||
|
||||
## DIRECT ANSWER — can the TM gun be the best 1v1 gun?
|
||||
|
||||
**(c) It loses live AND it sits at/below its majority class — the target carries
|
||||
no learnable conditional signal, and that is the reason.**
|
||||
|
||||
* LIVE: the registered TM (radial) loses decisively to Pattern (3.50% vs 10.74%,
|
||||
p=0.0006; 0 vs 25 round wins) and is no better than its own Linear base
|
||||
(p=0.66). It was never the best gun live.
|
||||
* OFFLINE: the GF head at its honest (raw) readout is 2.1 pp BELOW the 38.8%
|
||||
majority class; the gated config is only 1.4–1.6 pp above, by predicting the
|
||||
majority 62% of the time and with 13–14% recall on the two minority classes the
|
||||
correction would need. The radial head was already measured at/below majority
|
||||
in Round 3 (56.2% vs 57.2%).
|
||||
* Therefore the TM machine is not the limiting factor in the sense "it learns and
|
||||
we aim it wrong" (that would be (b)); the discrete GF/radial targets simply do
|
||||
not contain conditional information beyond the base rate for these surfers, and
|
||||
the live result agrees with the offline majority-class analysis.
|
||||
|
||||
## MEASURED vs INFERRED (round 4)
|
||||
|
||||
* **MEASURED**: the GF head confusion matrices, warm label histograms, majority
|
||||
shares, accuracies, margins, per-class precision/recall, and the shuffled
|
||||
controls (seeds=1 RAW+GATED, seeds=3 GATED); the fact that RAW is below
|
||||
majority and GATED is +1.4–1.6 pp above by majority-class over-prediction.
|
||||
* **MEASURED**: the live table (shots, hits, real %, dmg/run, round wins, mean
|
||||
round length, per-run ranges, exact permutation p vs onlyPattern and TM vs
|
||||
Linear), the frozen-binary commit + sha256, and the per-arm `[rack] active=`
|
||||
liveness lines.
|
||||
* **MEASURED**: `onlyTMPATTERN` and `onlyLinear` are indistinguishable live
|
||||
(p=0.66/0.44) — the radial correction changes nothing in real physics.
|
||||
* **INFERRED**: that the drop from the old 46.0% to the current 36.7%/40.4% is
|
||||
caused by the deferred-label fix enlarging the sample set (1.26 M → 1.75 M)
|
||||
with the short-aim samples the head gets wrong; the current numbers are
|
||||
measured, the causal attribution is not separately ablated.
|
||||
* **INFERRED (not measured)**: a live GF-mode arm — not run (one frozen HEAD
|
||||
binary; HEAD registers radial only). Task 1 makes it the unpromising candidate.
|
||||
|
||||
Reference in New Issue
Block a user