From 9cd6e9b8ce19a5dc03ae6f1c3f771c7bab8f0e06 Mon Sep 17 00:00:00 2001 From: Davide Cappellini Date: Tue, 22 Sep 2026 02:08:38 +0200 Subject: [PATCH] Ablation: the radial TM is replaceable by a CONSTANT, and its avenue is dead on the shipped metric The radial TM beats Linear on bmPoint, but its head never beat the majority baseline after the label bias was fixed - suggesting the win is a constant lean rather than learning. So: sweep a stateless constant short-range offset (new `common_libs/guns/radial_offset.nim`, no learning at all) against the learned TM. VERDICT (measured, offline range, seeds=3, 18 paired runs, 231 TM rounds): 1. **bmPoint - REPLACE the TM with a constant.** `RO_s0.95` (aim distance x0.95) TIES it early (9/9, p=1.0) and BEATS it overall (15/3, p=0.0075; 7.47% vs 6.89% per-run mean). A fixed -20px does the same. The head never beats its majority baseline (56.2% vs 57.2%). 2. **bmPath (the SHIPPED metric) - the radial avenue is a DEAD END.** TMRadial is a systematic LOSS there (2/16, p=0.0013); every constant is within +-0.2pp; the only real bmPath effect is the BotRadius clamp. So the radial shift cannot help the shipped configuration. 3. The per-adversary optimum DOES vary (fixed -10 for crazy, -30 for tr_crazy, scale 0.95 for three others) - but ONE GLOBAL CONSTANT still beats the adaptively-trained head, so the "fragility justifies learning" argument FAILS. THE REAL FINDING UNDERNEATH, and it generalises beyond this gun: the base linear prediction systematically OVERSHOOTS. Measured raw per-tick base radial error has mean -71 to -100 px; the enemy is NEARER than the prediction in 63-81% of shots and farther in only 4-14%, CONSISTENT ACROSS ALL SIX CAPTURES. Radial label histogram [415166,126461,120325,44694,19351] = 57.2% majority class, mean label -82.3 px, mean applied shift -37.6 px. So the net-short bias is a GENUINE property of these range-holders against a constant-velocity extrapolation (they decelerate and turn, so the true position is closer than the straight-line guess) - NOT a fixture artefact. That is worth chasing for the guns that actually ship. Caveat: bmPoint is not the shipped metric (bmPath won the real-hit-rate A/B for SELECTION), so a bmPoint win is not yet evidence of a real win. That needs a live test - and the natural target is Pattern, which is now the default and best gun. Adds radial_offset.nim + sweep_radial_offset.nim; tm_pattern.nim gains additive instrumentation only (radial label mean and applied-shift mean; no behaviour change, and test_tm_pattern_registration still passes all 20 checks). --- common_libs/guns/radial_offset.nim | 55 ++ common_libs/guns/tm_pattern.nim | 12 + common_libs/tests/sweep_radial_offset.nim | 499 ++++++++++++++++++ common_libs/tests/tm_pattern_sweep_results.md | 226 ++++++++ 4 files changed, 792 insertions(+) create mode 100644 common_libs/guns/radial_offset.nim create mode 100644 common_libs/tests/sweep_radial_offset.nim diff --git a/common_libs/guns/radial_offset.nim b/common_libs/guns/radial_offset.nim new file mode 100644 index 0000000..a8aa158 --- /dev/null +++ b/common_libs/guns/radial_offset.nim @@ -0,0 +1,55 @@ +## Constant-offset radial gun — the ABLATION CONTROL for `guns/tm_pattern.nim`'s +## radial head. +## +## Why this exists: the radial Tsetlin head was shown to win on `bmPoint` not by +## out-classifying a majority baseline (its online accuracy sits AT/BELOW the +## majority-class rate) but — per its own committed write-up — because the net +## applied radial correction is a positive average shift of the aim distance. If +## that is true, a FIXED radial shift should reproduce most or all of the win +## with no learning, no 0.36 ms/tick cost and no risk. +## +## This gun is that fixed shift, and nothing else: +## +## base = `forecastLinear` — the exact self-consistent forecast `LinearGun` +## and the TM base use. +## aim = the base BEARING unchanged; the aim DISTANCE scaled by `scale` and +## shifted by `offsetPx`: +## aimDist = f.dist * scale + offsetPx +## clamp = the SAME `[BotRadius, arena-BotRadius]` clamp the TM's corrective +## (non-base) path uses, so this is byte-comparable with TMRadial. +## +## `scale == 1.0 and offsetPx == 0.0` is the pure base through the corrective +## clamp — useful as a clamp-only diagnostic against `LinearGun`'s `[0, arena]`. +## There is NO learning, no history and no per-tick state: two runs on the same +## state stream are identical by construction. + +import std/math +import gun_harness/gun_interface +import guns/lead_forecast + +type + RadialOffsetGun* = object + scale*: float ## multiplicative factor on the base fire distance + offsetPx*: float ## fixed px added to the base fire distance (negative = short) + debugGraphics*: bool + +proc initRadialOffsetGun*(scale = 1.0, offsetPx = 0.0): RadialOffsetGun = + RadialOffsetGun(scale: scale, offsetPx: offsetPx) + +proc predict*(g: var RadialOffsetGun, state: WorldState, + bulletSpeed: float): GunPrediction = + if bulletSpeed <= 0.0: + return GunPrediction(x: state.enemyX, y: state.enemyY) + let f = forecastLinear(state, bulletSpeed) + let aimDist = f.dist * g.scale + g.offsetPx + let px = state.selfX + cos(f.bearing) * aimDist + let py = state.selfY + sin(f.bearing) * aimDist + # Mirror the TM's corrective path exactly (TMRadial when radOffset != 0): + # same bearing, adjusted distance, BotRadius-inset clamp. + GunPrediction( + x: clamp(px, BotRadius, state.arenaWidth - BotRadius), + y: clamp(py, BotRadius, state.arenaHeight - BotRadius), + ) + +proc onResult*(g: var RadialOffsetGun, e: FeedbackEvent) = + discard # analytical, stateless — the ablation has no learning by design diff --git a/common_libs/guns/tm_pattern.nim b/common_libs/guns/tm_pattern.nim index 090f5f7..e3ba0b4 100644 --- a/common_libs/guns/tm_pattern.nim +++ b/common_libs/guns/tm_pattern.nim @@ -179,6 +179,13 @@ type revLabelHist*: array[2, int] radCorrect*, radTotal*: int revCorrect*, revTotal*: int + ## Radial-target instrumentation (ablation support): raw label statistics + ## and the mean APPLIED aim-distance shift, so a constant-offset control can + ## be compared against the actual average readout the TM produces. + radDeltaSum*, radDeltaAbsSum*: float + radDeltaN*: int + radOffsetSum*: float + radOffsetN*: int classCorrect*: int ## warm predictions whose class matched the eventual label classTotal*: int ## warm predictions with a resolvable label lastChosen*: int @@ -502,6 +509,9 @@ proc tmResolveTrace(g: var TmPatternGun, t: TmPatternTrace, power: float) = # distance. Independent of our own aim, so it is a clean target. let actualRadius = hypot(g.posRing[s].x - t.fireX, g.posRing[s].y - t.fireY) let radDelta = actualRadius - t.fireDist + inc g.radDeltaN + g.radDeltaSum += radDelta + g.radDeltaAbsSum += abs(radDelta) let radWinner = if shuffleRad: rand(TM_CLASSES - 1) else: radToBucket(radDelta) inc g.radLabelHist[radWinner] if t.warm: @@ -612,6 +622,8 @@ proc predict*(g: var TmPatternGun, state: WorldState, bulletSpeed: float): else: gf = TM_SHRINK * bucketToGF(chosen) of tmRadial: radOffset = bucketToRadial(radChosen) + g.radOffsetSum += radOffset + inc g.radOffsetN of tmReversal: if TM_GF_MODE == "soft": gf = TM_SHRINK * tmSoftGF(votes) else: gf = TM_SHRINK * bucketToGF(chosen) diff --git a/common_libs/tests/sweep_radial_offset.nim b/common_libs/tests/sweep_radial_offset.nim new file mode 100644 index 0000000..c000b1e --- /dev/null +++ b/common_libs/tests/sweep_radial_offset.nim @@ -0,0 +1,499 @@ +## ABLATION: does the radial TM's `bmPoint` advantage need the Tsetlin Machine? +## +## The committed radial-TM finding (589a230) is that its `bmPoint` win does NOT +## come from out-classifying a majority baseline (the radial head sits AT/BELOW +## the 58.2% majority rate) but from a NET-POSITIVE AVERAGE radial aim-distance +## shift. If so, a FIXED radial shift should reproduce most or all of the win +## with no learning at all. +## +## This sweep builds a constant-offset variant of the base gun +## (`guns/radial_offset.nim`: exact `forecastLinear` bearing, aim distance scaled +## and shifted by a fixed constant, same BotRadius clamp as the TM corrective +## path) and sweeps the constant. It compares against: +## 1. `Linear` — the unmodified base; +## 2. `TMRadial` — the learned radial head (current best config, defaults +## TM_RADIAL_RANGE=60, TM_RAD_MARGIN=0.25, 5 classes); +## 3. `TMRadialShuf` — the mandatory shuffled-label control; +## under BOTH metrics (`bmPoint`, and the SHIPPED `bmPath`). +## +## It also reports the raw radial-label distribution and per-adversary optima, +## so the reader can judge whether the net shift is a genuine surfer property or +## an artefact, and whether a single constant is fragile. +## +## Usage: +## nim c -r -d:release --path:common_libs \ +## common_libs/tests/sweep_radial_offset.nim \ +## --set=real --metric=point --seeds=3 +## Flags: --set=real|range|synthetic, --metric=point|path, --seeds=N, +## --maxrounds=N, --summaryOnly=0|1 +## +## EARLY = resolutions whose LOCAL tick is in the first 100 ticks of a round. +## OVERALL = every resolution in the round. Methodology (fixture set, round +## splitting, fresh gun per fixture) matches `sweep_tm_pattern.nim` exactly so +## the TMRadial numbers here are directly comparable to the committed ones. + +import std/[os, strformat, strutils, json, tables, math, random, algorithm, sequtils] +import gun_harness/offline_range +import gun_harness/gun_interface +import gun_harness/virtual_bullets as vb +import guns/linear +import guns/tm_pattern +import guns/lead_forecast +import guns/radial_offset + +const + repoRoot = currentSourcePath().parentDir.parentDir.parentDir + fixturesDir = repoRoot / "tools" / "fixtures" + +type + Adapt = object + h100, n100, h300, n300, hall, nall, f100, m100: int + rounds: int + + GunStats = object + obs, labelMiss, traceMiss: int + radLabel: array[TM_CLASSES, int] + radChosen: array[TM_CLASSES, int] + radCorrect, radTotal: int + radDeltaSum, radDeltaAbsSum: float + radDeltaN: int + radOffsetSum: float + radOffsetN: int + + ArmKind = enum akLinear, akOffset, akTMRad, akTMRadShuf + + Arm = object + name: string + kind: ArmKind + scale, offsetPx: float + + RoundSpan = tuple[start, count: int] + + Row = object + variant, fixture: string + seed: int + r: Adapt + st: GunStats + +var gDropped = 0 ## unresolved bullets clobbered by the tracker ring (must be 0) + +# ── stats helpers ──────────────────────────────────────────────────────────── + +proc addAdapt(dst: var Adapt, src: Adapt) = + inc dst.rounds, src.rounds + dst.h100 += src.h100; dst.n100 += src.n100 + dst.h300 += src.h300; dst.n300 += src.n300 + dst.hall += src.hall; dst.nall += src.nall + dst.f100 += src.f100; dst.m100 += src.m100 + +proc addStats(dst: var GunStats, src: GunStats) = + dst.obs += src.obs; dst.labelMiss += src.labelMiss; dst.traceMiss += src.traceMiss + dst.radCorrect += src.radCorrect; dst.radTotal += src.radTotal + dst.radDeltaSum += src.radDeltaSum; dst.radDeltaAbsSum += src.radDeltaAbsSum + dst.radDeltaN += src.radDeltaN + dst.radOffsetSum += src.radOffsetSum; dst.radOffsetN += src.radOffsetN + for c in 0.. n: return 0.0 + var lg = 0.0 + for i in 1..k: lg += ln(float(n - k + i)) - ln(float(i)) + exp(lg - float(n) * ln(2.0)) + +proc signTestP(wins, n: int): float = + if n == 0: return 1.0 + let lo = min(wins, n - wins) + var s = 0.0 + for k in 0..lo: s += binomPmf(k, n) + min(1.0, 2.0 * s) + +# ── fixture / round I/O ────────────────────────────────────────────────────── + +proc loadRounds(path: string): seq[RoundSpan] = + let dir = path.parentDir + let base = path.extractFilename + var side = dir / "drussgt_meta" / (base & ".rounds.json") + if not fileExists(side): side = dir / (base & ".rounds.json") + if not fileExists(side): return @[] + let node = parseJson(readFile(side)) + if not node.hasKey("rounds"): return @[] + for r in node["rounds"]: + result.add (r["startTick"].getInt(), r["count"].getInt()) + +proc resolve(name: string): tuple[fx: Fixture, path: string] = + let p = if fileExists(name): name else: fixturesDir / (name & ".jsonl") + (loadFixture(p), p) + +proc fixtureSet(name: string): seq[string] = + case name + of "synthetic", "": + for n in SyntheticFixtureNames: result.add n + of "real": + for n in ["drussgt_vs_crazy", "drussgt_vs_spinbot", "drussgt_vs_drussgt", + "tr_drussgt_vs_crazy", "tr_drussgt_vs_spinbot", + "tr_drussgt_vs_modularbot"]: + result.add(fixturesDir / (n & ".jsonl")) + of "range": + for n in SyntheticFixtureNames: result.add n + for n in ["drussgt_vs_crazy", "drussgt_vs_spinbot", "tr_drussgt_vs_crazy"]: + result.add(fixturesDir / (n & ".jsonl")) + else: discard + +# ── replay ─────────────────────────────────────────────────────────────────── + +proc runRound(states: seq[WorldState], lastSeen: seq[int], enemyId, baseTick: int, + drivers: seq[GunDriver], metric: BulletMetric): seq[Adapt] = + var tracker = initTracker(drivers.len, metric) + var accum = newSeq[Adapt](drivers.len) + for ad in accum.mitems: inc ad.rounds + for si in 0..= 0: lst = lastSeen[si] + if state.enemies.len > 0: + for e in state.enemies: + enemyPositions[e.id] = (x: e.x, y: e.y, lastSeenTick: lst, alive: true) + else: + enemyPositions[enemyId] = (x: state.enemyX, y: state.enemyY, + lastSeenTick: lst, alive: true) + + let localTick = state.tick - baseTick + let dref = drivers + tracker.tickBullets(state, enemyPositions, + proc(gunId: GunId, binIdx: int, e: FeedbackEvent) = + inc accum[gunId].nall + if e.hit: inc accum[gunId].hall + if localTick < 100: + inc accum[gunId].n100 + if e.hit: inc accum[gunId].h100 + if localTick < 300: + inc accum[gunId].n300 + if e.hit: inc accum[gunId].h300 + let fireTick = e.fireTick - baseTick + if fireTick < 100: + inc accum[gunId].m100 + if e.hit: inc accum[gunId].f100 + dref[gunId].resultCb(e)) + gDropped += tracker.droppedBullets + result = accum + +proc runFixtureClean(fx: Fixture, path: string, drivers: seq[GunDriver], + metric: BulletMetric, maxRounds = 0): seq[Adapt] = + var spans = + if fx.meta.source == "synthetic": @[(start: 0, count: fx.states.len)] + else: loadRounds(path) + if maxRounds > 0 and spans.len > maxRounds: spans.setLen(maxRounds) + result = newSeq[Adapt](drivers.len) + if spans.len == 0: + result = runRound(fx.states, fx.lastSeen, fx.enemyId, 0, drivers, metric) + return + for sp in spans: + var st: seq[WorldState] + var ls: seq[int] + for i in 0..= sp.start and t < sp.start + sp.count: + st.add fx.states[i] + ls.add(if i < fx.lastSeen.len: fx.lastSeen[i] else: -1) + if st.len == 0: continue + let rr = runRound(st, ls, fx.enemyId, sp.start, drivers, metric) + for gi in 0..= 0: randomize(seed) + result.gun = g + result.driver = GunDriver( + name: name, + predictCb: proc(state: WorldState, bulletSpeed: float): GunPrediction = + g[].predict(state, bulletSpeed), + resultCb: proc(e: FeedbackEvent) = g[].onResult(e), + readyCb: proc(): bool = g[].isWarmedUp()) + +proc detDriver(a: Arm): GunDriver = + case a.kind + of akLinear: makeDriver(a.name, LinearGun()) + of akOffset: makeDriver(a.name, initRadialOffsetGun(a.scale, a.offsetPx)) + else: + raise newException(ValueError, "not a deterministic arm: " & a.name) + +# ── main ───────────────────────────────────────────────────────────────────── + +proc main() = + var set = "real" + var metricName = "point" + var nSeeds = 3 + var maxRounds = 0 + for i in 1..paramCount(): + let a = paramStr(i) + if a.startsWith("--set="): set = a[6..^1] + elif a.startsWith("--metric="): metricName = a[9..^1] + elif a.startsWith("--seeds="): nSeeds = parseInt(a[8..^1]) + elif a.startsWith("--maxrounds="): maxRounds = parseInt(a[12..^1]) + else: discard + let metric = if metricName == "point": bmPoint else: bmPath + let names = fixtureSet(set) + + # ── arms ─────────────────────────────────────────────────────────────────── + var arms: seq[Arm] + arms.add Arm(name: "Linear", kind: akLinear) + # multiplicative shrink (fraction of base fire distance) + for s in [1.00, 0.98, 0.95, 0.90, 0.85, 0.80]: + arms.add Arm(name: &"RO_s{s:.2f}", kind: akOffset, scale: s) + # fixed px shift (negative = aim short); +30 = opposite-direction control + for o in [-10.0, -20.0, -30.0, -40.0, -60.0, 30.0]: + arms.add Arm(name: &"RO_o{int(o):+d}", kind: akOffset, scale: 1.0, offsetPx: o) + arms.add Arm(name: "TMRadial", kind: akTMRad) + arms.add Arm(name: "TMRadialShuf", kind: akTMRadShuf) + + let detArms = arms.filterIt(it.kind in {akLinear, akOffset}) + let tmArms = arms.filterIt(it.kind in {akTMRad, akTMRadShuf}) + + echo "# set=", set, " metric=", metricName, " seeds=", nSeeds, + " fixtureCount=", names.len + + var rows: seq[Row] + echo "variant,fixture,seed,h100,n100,h300,n300,hall,nall,rounds,obs,labelMiss,traceMiss" + for name in names: + let (fx, path) = resolve(name) + let fxName = path.extractFilename.replace(".jsonl", "") + + # Deterministic arms: one batched pass (stateless, order-independent). + let dDrivers = detArms.mapIt(detDriver(it)) + let dRes = runFixtureClean(fx, path, dDrivers, metric, maxRounds) + for i, a in detArms: + let row = Row(variant: a.name, fixture: fxName, seed: 1, r: dRes[i], st: GunStats()) + rows.add row + echo &"{row.variant},{fxName},1,{row.r.h100},{row.r.n100},{row.r.h300},{row.r.n300}," & + &"{row.r.hall},{row.r.nall},{row.r.rounds},0,0,0" + + # TM arms: single-driver, fresh gun per fixture, seeded. + for a in tmArms: + for seed in 1..nSeeds: + let pair = makeTMRadDriver(a.name, seed, shuffle = (a.kind == akTMRadShuf)) + let rr = runFixtureClean(fx, path, @[pair.driver], metric, maxRounds)[0] + let row = Row(variant: a.name, fixture: fxName, seed: seed, r: rr, + st: tmStats(pair.gun)) + rows.add row + echo &"{row.variant},{fxName},{seed},{row.r.h100},{row.r.n100},{row.r.h300}," & + &"{row.r.n300},{row.r.hall},{row.r.nall},{row.r.rounds},{row.st.obs}," & + &"{row.st.labelMiss},{row.st.traceMiss}" + + if gDropped > 0: + echo &"\n# WARNING: droppedBullets={gDropped} (ring clobbered unresolved bullets)" + + # ── pooled summary ───────────────────────────────────────────────────────── + var pooled = initTable[string, Adapt]() + var pooledSt = initTable[string, GunStats]() + for a in arms: + pooled[a.name] = Adapt() + pooledSt[a.name] = GunStats() + for row in rows: + addAdapt(pooled[row.variant], row.r) + addStats(pooledSt[row.variant], row.st) + + echo "\n# ── pooled summary ──" + echo "variant,runs,early%,early_hits,early_n,overall%,overall_hits,overall_n" + for a in arms: + let p = pooled[a.name] + echo &"{a.name},{p.rounds},{rateStr(p.h100, p.n100)},{p.h100},{p.n100}," & + &"{rateStr(p.hall, p.nall)},{p.hall},{p.nall}" + + # ── per-run tables: deterministic arms replicated across seeds ───────────── + type Key = tuple[fixture: string, seed: int] + var byEarly = initTable[string, Table[Key, float]]() + var byAll = initTable[string, Table[Key, float]]() + var byEarlyHits = initTable[string, Table[Key, tuple[h, n: int]]]() + var byAllHits = initTable[string, Table[Key, tuple[h, n: int]]]() + for a in arms: + byEarly[a.name] = initTable[Key, float]() + byAll[a.name] = initTable[Key, float]() + byEarlyHits[a.name] = initTable[Key, tuple[h, n: int]]() + byAllHits[a.name] = initTable[Key, tuple[h, n: int]]() + for row in rows: + let k = (fixture: row.fixture, seed: row.seed) + let e = if row.r.n100 > 0: row.r.h100.float / row.r.n100.float else: 0.0 + let o = if row.r.nall > 0: row.r.hall.float / row.r.nall.float else: 0.0 + byEarly[row.variant][k] = e + byAll[row.variant][k] = o + byEarlyHits[row.variant][k] = (row.r.h100, row.r.n100) + byAllHits[row.variant][k] = (row.r.hall, row.r.nall) + # replicate deterministic arms across the seed range for paired comparison + for a in detArms: + for row in rows: + if row.variant == a.name: + for s in 2..nSeeds: + byEarly[a.name][(row.fixture, s)] = byEarly[a.name][(row.fixture, 1)] + byAll[a.name][(row.fixture, s)] = byAll[a.name][(row.fixture, 1)] + byEarlyHits[a.name][(row.fixture, s)] = byEarlyHits[a.name][(row.fixture, 1)] + byAllHits[a.name][(row.fixture, s)] = byAllHits[a.name][(row.fixture, 1)] + + echo "\n# ── per-run distribution (mean / min / max, n runs) ──" + echo "variant,earlyMean%,earlyMin%,earlyMax%,overallMean%,overallMin%,overallMax%,n" + for a in arms: + var es, osx: seq[float] + for v in byEarly[a.name].values: es.add v + for v in byAll[a.name].values: osx.add v + if es.len == 0: continue + es.sort(); osx.sort() + echo &"{a.name},{es.sum/float(es.len)*100:.2f},{es[0]*100:.2f},{es[^1]*100:.2f}," & + &"{osx.sum/float(osx.len)*100:.2f},{osx[0]*100:.2f},{osx[^1]*100:.2f},{es.len}" + + # ── sign tests ───────────────────────────────────────────────────────────── + echo "\n# ── paired sign tests (rows = fixture x seed; exact two-sided binomial) ──" + echo "A,B,metric,nA>B,nB>A,ties,p" + proc signRow(aName, bName, label: string, tab: Table[string, Table[Key, float]]) = + var winsA, winsB, ties, n = 0 + for k, va in tab[aName].pairs: + if k notin tab[bName]: continue + let v = tab[bName][k] + inc n + if va > v: inc winsA + elif v > va: inc winsB + else: inc ties + if n == 0: return + echo &"{aName},{bName},{label},{n},{winsA},{winsB},{ties},{signTestP(winsA, n - ties):.4f}" + + for a in detArms: + if a.name != "Linear": + signRow(a.name, "Linear", "early", byEarly) + signRow(a.name, "Linear", "overall", byAll) + signRow("TMRadial", "Linear", "early", byEarly) + signRow("TMRadial", "Linear", "overall", byAll) + signRow("TMRadial", "TMRadialShuf", "early", byEarly) + signRow("TMRadial", "TMRadialShuf", "overall", byAll) + for a in detArms: + if a.name != "Linear": + signRow(a.name, "TMRadial", "early", byEarly) + signRow(a.name, "TMRadial", "overall", byAll) + + # ── per-adversary optimum for the constant offset ────────────────────────── + echo "\n# ── per-fixture best constant offset (pooled over seeds) ──" + echo "fixture,bestEarlyVariant,bestEarly%,bestOverallVariant,bestOverall%,linearEarly%,linearOverall%,tmEarly%,tmOverall%" + var fxNames: seq[string] + for row in rows: + if row.fixture notin fxNames: fxNames.add row.fixture + fxNames.sort() + for fxName in fxNames: + var bestE = ("none", -1.0) + var bestO = ("none", -1.0) + for a in detArms: + if a.name == "Linear": continue + var eh, en, oh, on: int + for row in rows: + if row.variant == a.name and row.fixture == fxName: + eh += row.r.h100; en += row.r.n100 + oh += row.r.hall; on += row.r.nall + let er = if en > 0: eh.float/en.float else: -1.0 + let orr = if on > 0: oh.float/on.float else: -1.0 + if er > bestE[1]: bestE = (a.name, er) + if orr > bestO[1]: bestO = (a.name, orr) + var lh, ln, loh, lon: int + var th, tn, toh, ton: int + for row in rows: + if row.fixture != fxName: continue + if row.variant == "Linear": + lh += row.r.h100; ln += row.r.n100; loh += row.r.hall; lon += row.r.nall + elif row.variant == "TMRadial": + th += row.r.h100; tn += row.r.n100; toh += row.r.hall; ton += row.r.nall + echo &"{fxName},{bestE[0]},{pctStr(bestE[1])},{bestO[0]},{pctStr(bestO[1])}," & + &"{rateStr(lh,ln)},{rateStr(loh,lon)},{rateStr(th,tn)},{rateStr(toh,ton)}" + + # ── radial-label distribution + applied shift (from the real TM arm) ─────── + echo "\n# ── radial-label distribution and applied shift (TMRadial, delta per fixture) ──" + echo "scope,n,labelHist,labelMajority%,meanRadDeltaPx,meanAbsRadDeltaPx,chosenHist,meanAppliedShiftPx,onlineAcc" + proc radReport(scope: string, st: GunStats) = + var n = 0 + var major = 0 + var hs: string + for c in 0.. 0: st.radDeltaSum / st.radDeltaN.float else: 0.0 + let meanAbsD = if st.radDeltaN > 0: st.radDeltaAbsSum / st.radDeltaN.float else: 0.0 + let meanShift = if st.radOffsetN > 0: st.radOffsetSum / st.radOffsetN.float else: 0.0 + let acc = rateStr(st.radCorrect, st.radTotal) + let majPct = if n > 0: &"{major.float/n.float*100.0:.1f}" else: "n/a" + echo &"{scope},{n},{hs},{majPct},{meanD:.2f},{meanAbsD:.2f},{chs},{meanShift:.2f},{acc}" + var poolSt: GunStats + for row in rows: + if row.variant == "TMRadial": + addStats(poolSt, row.st) + radReport(row.fixture, row.st) + radReport("POOLED", poolSt) + + # ── raw radial DELTA distribution from the base forecast (metric-free) ───── + echo "\n# ── raw base radial-error distribution (per fired-bullet arrival, from forecastLinear) ──" + echo "scope,n,meanPx,meanAbsPx,fracNearer(short),fracFarther(long)" + for name in names: + let (fx, path) = resolve(name) + let fxName = path.extractFilename.replace(".jsonl", "") + # map tick -> enemy pose as the TM's own posRing would + var pose = initTable[int, tuple[x, y: float]]() + for s in fx.states: pose[s.tick] = (s.enemyX, s.enemyY) + var sum, asum = 0.0 + var n, near, far: int + for s in fx.states: + for b in 0.. BotRadius: inc far + let meanD = if n > 0: sum/n.float else: 0.0 + let meanA = if n > 0: asum/n.float else: 0.0 + echo &"{fxName},{n},{meanD:.2f},{meanA:.2f},{near.float/max(1,n).float:.3f},{far.float/max(1,n).float:.3f}" + +when isMainModule: + main() diff --git a/common_libs/tests/tm_pattern_sweep_results.md b/common_libs/tests/tm_pattern_sweep_results.md index 79c41ac..fa10474 100644 --- a/common_libs/tests/tm_pattern_sweep_results.md +++ b/common_libs/tests/tm_pattern_sweep_results.md @@ -364,3 +364,229 @@ by these numbers. There is no persistence across battles. * INFERRED: that the radial win comes from surfers being NEARER than the base prediction (range-holding), supported by the asymmetric radial label histogram but not separately modelled. + +--- + +# ROUND 3 — ABLATION: does the radial win even need the Tsetlin Machine? + +Date: 2026-09-22. Artifacts: `common_libs/guns/radial_offset.nim` (new), +`common_libs/tests/sweep_radial_offset.nim` (new), `common_libs/guns/tm_pattern.nim` +(label/readout instrumentation only). Raw outputs: `/tmp/ro_point_s3.txt`, +`/tmp/ro_path_s3.txt` (`--set=real --seeds=3`), `/tmp/ro_point_s1.txt`. + +## The question + +Commit 589a230 retracted the radial head's "conditional learning" claim: the +radial head's online accuracy (57.0%) is AT/BELOW the 58.2% majority baseline, and +the bmPoint metric win was attributed to a **NET-POSITIVE AVERAGE RADIAL SHIFT**. +If that is what it is, a FIXED radial shift should reproduce the win with no +learning, no 0.36 ms/tick cost and no risk. + +`radial_offset.nim` is exactly that fixed shift and nothing else: the exact +`forecastLinear` bearing, aim distance `f.dist * scale + offsetPx`, and the SAME +`[BotRadius, arena-BotRadius]` clamp the TM's corrective path uses. It is +stateless, so two runs are identical by construction. The sweep runs **Linear**, +a grid of constants, **TMRadial** and **TMRadialShuf** under BOTH metrics, using +the identical fixture set / round splitting / per-fixture fresh gun methodology +as `sweep_tm_pattern.nim`. + +**Harness sanity check (MEASURED):** `TMRadial` in this report reproduces the +committed numbers exactly — bmPoint early 9.1% / overall 5.7%, vs Linear early +16/2 p=0.0013 and overall 18/0 p<0.0001, vs Shuf 18/0 p<0.0001. So the ablation +runs on the same experiment, not a re-derivation. + +0 dropped bullets in every run (ring never clobbered); `labelMiss=0` and +`traceMiss=0` for every TM row (the deferred-label fix holds). + +## bmPoint (seeds=3) + +Pooled over 77 rounds for the deterministic arms (one pass; replicated across the +3 seeds only for the paired test) and 231 rounds for the TM arms (3 seeds). + +| variant | early | overall | +|---|---|---| +| Linear | 7.2% (1480/20498) | 4.7% (11277/241423) | +| `RO_s1.00` (base + BotRadius clamp only) | 7.2% (1481/20626) | 4.7% (11388/241551) | +| `RO_s0.98` (scale 0.98) | 8.0% (1669/20781) | 5.8% (14033/241677) | +| **`RO_s0.95`** | **8.8% (1852/21001)** | **6.3% (15299/241890)** | +| `RO_o-10` (fixed -10 px) | 8.4% (1756/20790) | 5.9% (14336/241689) | +| **`RO_o-20`** | 7.9% (1658/20939) | **6.4% (15445/241836)** | +| `RO_o-30` | 5.4% (1130/21117) | 4.9% (11740/242007) | +| `RO_o-40` | 4.1% (883/21292) | 3.7% (8989/242166) | +| `RO_o-60` | 4.7% (1019/21655) | 3.0% (7173/242495) | +| `RO_o+30` (opposite-direction control) | 0.6% (131/20159) | 0.3% (650/241168) | +| **TMRadial** | **9.1% (5776/63518)** | 5.7% (41232/726357) | +| TMRadialShuf | 7.0% (4314/61672) | 3.8% (27317/724473) | + +Per-run means over the 18 fixture×seed runs: +Linear 13.61 / 5.77; `RO_s0.95` **14.86 / 7.47**; `RO_o-20` 10.35 / 7.45; +TMRadial **15.08 / 6.89**; Shuf 12.89 / 4.54. + +Paired sign tests (exact two-sided binomial, 18 pairs): + +| A | B | metric | nA>B | nB>A | p | +|---|---|---|---|---|---| +| RO_s0.95 | Linear | early | 15 | 3 | **0.0075** | +| RO_s0.95 | Linear | overall | 15 | 3 | **0.0075** | +| RO_s0.98 | Linear | overall | 18 | 0 | **<0.0001** | +| RO_o-20 | Linear | early | 15 | 3 | **0.0075** | +| RO_o-20 | Linear | overall | 15 | 3 | **0.0075** | +| RO_o-10 | Linear | overall | 18 | 0 | **<0.0001** | +| RO_o+30 | Linear | both | 0 | 18 | **<0.0001** (wrong direction) | +| TMRadial | Linear | early | 16 | 2 | **0.0013** | +| TMRadial | Linear | overall | 18 | 0 | **<0.0001** | +| TMRadial | Shuf | both | 18 | 0 | **<0.0001** | +| **RO_s0.95** | **TMRadial** | **early** | **9** | **9** | **1.0000 (tie)** | +| **RO_s0.95** | **TMRadial** | **overall** | **15** | **3** | **0.0075 (constant wins)** | +| RO_o-20 | TMRadial | early | 6 | 12 | 0.2379 | +| RO_o-20 | TMRadial | overall | 12 | 6 | 0.2379 | +| RO_s0.98 | TMRadial | overall | 6 | 12 | 0.2379 | + +**MEASURED conclusion (bmPoint):** a constant radial shift of **scale 0.95** +(or fixed **−20 px**) *ties the learned radial head on EARLY and beats it on +OVERALL* (7.47% vs 6.89% per-run mean, 15/3 p=0.0075). The learned head's whole +nominal advantage is its slightly higher early rate (9.1% vs 8.8% pooled, +15.08% vs 14.86% per-run), and that is a statistical tie (9/9, p=1.0). +`RO_o+30` (aiming the opposite way) collapses to 0.3%, confirming the direction of +the effect is real and not a clamp artefact. + +## bmPath (the SHIPPED metric) — the radial avenue is a dead end + +| variant | early | overall | +|---|---|---| +| Linear | 34.0% (6358/18715) | 24.3% (58297/239943) | +| `RO_s1.00` (clamp only) | 34.2% (6350/18592) | **24.7%** (59217/239891) | +| `RO_s0.98` | 34.2% (6360/18623) | 24.6% (58950/239906) | +| `RO_s0.95` (best bmPoint constant) | 34.0% (6344/18675) | 24.4% (58530/239937) | +| `RO_o-20` (best bmPoint constant) | 34.1% (6354/18649) | 24.5% (58699/239921) | +| `RO_o+30` (best bmPath constant) | 34.4% (6357/18498) | **24.8%** (59366/239852) | +| `RO_s0.80` | 33.6% (6346/18886) | 23.4% (56151/240035) | +| TMRadial | 33.8% (19002/56211) | 24.0% (173001/719875) | +| TMRadialShuf | 33.9% (19026/56076) | 24.3% (175262/719776) | + +Per-run means (n=18): Linear 35.52 / 26.86; `RO_s1.00` 35.71 / **27.23**; +`RO_s0.95` 35.59 / 26.97; `RO_o-20` 35.63 / 27.04; **TMRadial 35.37 / 26.63** +(a loss). + +Paired sign tests (bmPath): + +* **TMRadial vs Linear: early 5/13 p=0.0963; overall 2/16 p=0.0013 — a + systematic LOSS.** (Matches the committed Round-2 finding.) +* `RO_s1.00` vs Linear: overall 18/0 p<0.0001 (+0.4 pp) — this is the pre-existing + **BotRadius-clamp** effect, *not* the radial shift. +* Every constant between 0.98 and 0.95 and every fixed offset −10..−20 is within + ±0.2 pp of Linear; `RO_s0.95` (a real +1.6 pp win on bmPoint overall) is only + +0.1 pp here and *below* the clamp-only arm. Strong shrink (≤0.90, ≤−40 px) is a + significant loss (e.g. `RO_s0.80` 3/15 p=0.0075 early, 0/18 overall). +* `RO_o+30` — the OPPOSITE direction to what bmPoint wants — is the best bmPath + constant (+0.5 pp, 18/0). So the bmPath response to the radial knob is + **clamp-mediated and direction-insensitive**, i.e. the radial degree of freedom + is a structural no-op, exactly as Round 2 argued. + +**MEASURED conclusion (bmPath):** neither the learned radial head nor any +constant reproduces a real gain. TMRadial is a systematic *loss*. **The whole +radial avenue cannot help the shipped configuration.** + +## The radial-label distribution (MEASURED) + +Pooled `TMRadial` training labels (3 seeds, n=725997), class centres +−60/−30/0/+30/+60 px: +`radLabelHist = [415166, 126461, 120325, 44694, 19351]` → **majority class 57.2%**, +`meanRadDelta = −82.30 px`, `meanAbsRadDelta = 90.13 px`. +The head's online accuracy is **56.2%**, i.e. at/below that majority. +The mean APPLIED shift is **−37.61 px** (`radChosenHist = +[460960, 35411, 249071, 6832, 2994]`). + +Metric-free raw base radial error (each fired bullet's enemy radius minus +`forecastLinear`'s fire distance, at the base arrival tick): + +| fixture | n | mean px | meanAbs px | frac nearer | frac farther | +|---|---|---|---|---|---| +| drussgt_vs_crazy | 35959 | −87.6 | 92.7 | 0.693 | 0.053 | +| drussgt_vs_spinbot | 19830 | −100.4 | 107.7 | 0.769 | 0.062 | +| drussgt_vs_drussgt | 26060 | −70.8 | 85.8 | 0.627 | 0.142 | +| tr_drussgt_vs_crazy | 45905 | −99.0 | 104.4 | 0.805 | 0.053 | +| tr_drussgt_vs_spinbot | 43150 | −92.4 | 96.7 | 0.793 | 0.044 | +| tr_drussgt_vs_modularbot | 79970 | −79.9 | 90.7 | 0.705 | 0.108 | + +The enemy is NEARER than the constant-velocity prediction in **63–81%** of fired +bullets and FARTHER in only **4–14%**, consistently across all six captures and +both metrics. So the net-short bias is a GENUINE property of these range-holding +surfers against the constant-velocity base (the base lets range grow +geometrically; they hold it) — not a one-fixture artefact. + +**INFERRED:** the mean label (−82 px) is much larger than the OPTIMAL constant +shift (−20 px). The mean is dominated by large radial errors that miss regardless +of the shift; the near-miss window is served better by a small shift, and a +larger shift also resolves the bullet at an earlier tick. The exact reason a +−20 px shift beats −60 px is not separately modelled. + +## Does the optimal constant vary by adversary? (MEASURED) + +bmPoint, best constant PER FIXTURE (in-sample upper bound), with Linear and +TMRadial for reference: + +| fixture | best early | best overall | Linear overall | TMRadial overall | +|---|---|---|---|---| +| drussgt_vs_crazy | RO_s0.80 (10.1%) | RO_o-10 (16.9%) | 15.8% | 17.0% | +| drussgt_vs_spinbot | RO_s0.95 (9.9%) | RO_s0.95 (11.5%) | 7.5% | 9.9% | +| drussgt_vs_drussgt | RO_s1.00 (55.3%) | RO_s0.98 (4.3%) | 3.9% | 4.1% | +| tr_drussgt_vs_crazy | RO_s0.85 (9.5%) | RO_o-30 (5.8%) | 2.9% | 4.4% | +| tr_drussgt_vs_modularbot | RO_s0.95 (6.8%) | RO_s0.95 (2.7%) | 1.3% | 2.1% | +| tr_drussgt_vs_spinbot | RO_s0.80 (6.8%) | RO_s0.95 (7.4%) | 3.3% | 3.7% | + +**MEASURED:** the per-fixture optimum DOES vary (fixed −10 for `crazy`, −30 for +`tr_crazy`, scale 0.95 for three others). **But** one GLOBAL constant (0.95) +still beats TMRadial on pooled and per-run overall bmPoint. So the per-adversary +variation is not enough to justify the learned head: a fixed 0.95 is already the +best pooled point-metric arm measured here. + +## DIRECT VERDICT + +1. **bmPoint: REPLACE the radial TM with a constant.** The best constant + (scale 0.95, equivalently fixed −20 px) statistically TIES the TM on early + (9/9 p=1.0) and BEATS it on overall (15/3 p=0.0075; 7.47% vs 6.89% per-run + mean). The head never out-classifies its majority baseline (56.2% vs 57.2%) + and its mean applied shift (−37.6 px) is roughly twice the optimal constant. + The TM buys nothing a constant does not, and costs 0.36 ms/tick + complexity. +2. **bmPath (shipped): DEAD END.** No constant and no TM improves it; TMRadial + is a systematic loss (2/16 p=0.0013), and the only real bmPath effect in the + table is the BotRadius clamp, which is direction-insensitive. Drop the radial + mode from any shipped configuration. +3. **Fragility argument fails.** The optimum does vary per adversary, but a + single global constant already matches/beats the adaptively-trained head — so + the TM is not earning its cost even by the "per-adversary adaptation" + argument (it is cold-every-battle and trains online within the battle, yet + still loses to the global 0.95). + +Net: **do not keep the radial TM.** If the arrival metric ever matters, ship the +stateless constant; for the current shipped metric, the radial mode (and its +registered gun id 14) is not justified. + +## MEASURED vs INFERRED (round 3) + +* MEASURED: every table, pooled rate, per-run mean, paired sign test, label + histogram, mean/abs radial delta, applied-shift mean, online accuracy, and the + exact reproduction of the committed TMRadial numbers. +* MEASURED: the constant-offset arm is stateless (verbatim `forecastLinear` + bearing, `f.dist*scale+offsetPx`, TM corrective clamp), so its rollout is + deterministic and its replicated per-seed values are legitimate. +* MEASURED: `RO_s1.00` isolates the BotRadius clamp — on bmPoint it is identical + to Linear (7.2/4.7%), on bmPath it is the +0.4 pp arm; the radial shift itself + adds nothing on bmPath. +* INFERRED: the explanation of the label-mean (−82 px) vs optimal shift (−20 px) + gap (large-error tail + earlier resolution tick); the direction claim itself is + measured (the +30 control collapses). +* INFERRED (not measured): whether a per-adversary constant would beat a global + one out-of-sample — the per-fixture optima above are in-sample. + +## How to reproduce (round 3) + +``` +nim c --path:common_libs -d:release -o:/tmp/sweep_radial_offset \ + common_libs/tests/sweep_radial_offset.nim +/tmp/sweep_radial_offset --set=real --metric=point --seeds=3 +/tmp/sweep_radial_offset --set=real --metric=path --seeds=3 +# smoke: --seeds=1 +``` +