TM verdict, settled: it loses LIVE and sits at/below its majority class - (c)
The user pushed back on "the TM can't be your best 1v1 gun", correctly, because two
decisive tests had never been run. Both are now run and they agree.
TASK 1 - THE GF HEAD vs ITS MAJORITY-CLASS BASELINE (offline, n=1,751,067):
label histogram [254286, 284578, 678879, 297055, 236269]
majority class = 2 (the CENTRE bucket) = 38.77%
RAW head accuracy = 36.69% -> margin **-2.08 pp, BELOW majority**
GATED head accuracy = 40.37% vs 38.75% majority -> +1.62 pp, BUT it predicts the
majority class on 62.4% of ticks and its minority recall is 13.6% / 12.9% - a
base-rate predictor wearing a classifier's clothes.
Shuffled control sits at its own majority (20.04% vs 20.12%), confirming chance.
**THE OLD "46% vs 20% CHANCE" FIGURE I QUOTED WAS WRONG ON TWO COUNTS:** the
baseline is 38.8%, not 20%, and the 46% predated the deferred-label fix. Against
the correct baseline the head is BELOW it.
TASK 2 - THE FIRST-EVER LIVE A/B OF THE TM GUN (7 runs x 7 rounds per arm, one
frozen binary from git archive HEAD = eb74f9b2, sha256 cb66d66b..., real DrussGT,
every arm forced alone with TR_RACK_<GUN>=both and all 14 others off, liveness
confirmed per run):
arm shots real % dmg/run round wins
onlyPattern 4610 10.74% 285 25/49
onlyTMPATTERN (radial) 3374 3.50% 71 0/49
onlyLinear 3218 3.23% 61 0/49
Pattern vs TM: +7.22 pp / +213.7 dmg, exact p=0.0006
TM vs Linear: +0.30 pp, p=0.659 (dmg p=0.438)
**The TM is statistically INDISTINGUISHABLE from its own Linear base live.** So it
is not "the TM works and we are aiming it wrong".
DIRECT ANSWER: **(c) It loses live AND sits at/below majority - the target carries
no learnable signal beyond the base rate, and that is the reason.** The reason is
not the machine, not the knobs, and not the application alone: the thing it was
asked to predict is dominated by the modal answer.
This closes the TM-as-gun thread. If a TM is wanted in the bot, a firing gate or a
movement decision is a better fit for a boolean-rule classifier than an aim point -
that is untested and is a different project.
A LIVE GF-MODE ARM WAS NOT RUN (stated as unmeasured): the task pinned one frozen
HEAD binary and HEAD registers the TM gun as radial only; Task 1 already makes GF
the unpromising candidate.
HARNESS FIX WORTH KEEPING: `tools/ab/which_gun_arm_env.sh` left the TARGET gun
unset, so with the now-Pattern-only default it silently fell back to the FULL rack
- an arm could appear to test a single gun while actually running the whole rack.
It now emits `TR_RACK_<GUN>=both` for the target and `=off` for all 14 others.
(Earlier which-gun results are unaffected: they ran before the Pattern-only default,
or - as in the melee/1v1 campaign - set the explicit `=both` themselves.)
tm_pattern.nim gains a per-class confusion matrix (warm samples only) to support the
majority baseline; no behaviour change. Adds Round 4 to
tm_pattern_sweep_results.md with both tasks and the interpretation rule.
This commit is contained in:
@@ -188,6 +188,14 @@ type
|
||||
radOffsetN*: int
|
||||
classCorrect*: int ## warm predictions whose class matched the eventual label
|
||||
classTotal*: int ## warm predictions with a resolvable label
|
||||
## Per-class confusion matrix for the GF head, indexed [true label][predicted
|
||||
## class], counted over WARM predictions only (the same samples `classTotal`
|
||||
## scores). This is what makes the majority-class baseline and per-class
|
||||
## precision/recall measurable. Row sums = the warm label histogram; the
|
||||
## diagonal sum = classCorrect.
|
||||
confusion*: array[TM_CLASSES, array[TM_CLASSES, int]]
|
||||
## Same for the radial head.
|
||||
radConfusion*: array[TM_CLASSES, array[TM_CLASSES, int]]
|
||||
lastChosen*: int
|
||||
shuffleLabels*: bool ## control: replace the computed GF label with a random class
|
||||
forceBase*: bool ## measurement: ignore the TM, emit the pure LinearGun base
|
||||
@@ -504,6 +512,7 @@ proc tmResolveTrace(g: var TmPatternGun, t: TmPatternTrace, power: float) =
|
||||
if t.warm:
|
||||
inc g.classTotal
|
||||
if winner == t.chosen: inc g.classCorrect
|
||||
inc g.confusion[winner][t.chosen]
|
||||
|
||||
# Radial label: enemy radius at the base arrival tick minus the base fire
|
||||
# distance. Independent of our own aim, so it is a clean target.
|
||||
@@ -517,6 +526,7 @@ proc tmResolveTrace(g: var TmPatternGun, t: TmPatternTrace, power: float) =
|
||||
if t.warm:
|
||||
inc g.radTotal
|
||||
if radWinner == t.radChosen: inc g.radCorrect
|
||||
inc g.radConfusion[radWinner][t.radChosen]
|
||||
|
||||
# Reversal label: net heading turn over the flight, opposite to the direction
|
||||
# the enemy was turning at fire time.
|
||||
|
||||
Reference in New Issue
Block a user