From a82c864c6075d8325589681bae23b371e3c4f1b5 Mon Sep 17 00:00:00 2001 From: Davide Cappellini Date: Thu, 24 Sep 2026 23:38:45 +0200 Subject: [PATCH] bitbrain campaign phase 0: offline prediction-quality ruler and the bar New harness (common_libs/gun_harness/prediction_quality.nim + common_libs/tests/run_prediction_quality.nim): per-gun single-tick aim error in degrees against the true continuous interception point on the recorded live-vs-real-DrussGT corpus (/tmp/tfil_ab2/out, 70 runs, 899607 ticks), per range band, with the hit-probability proxy mean(|err|<=atan(18/range)). Validated: recorded hits separate from misses 13.34x px (reference 11.59x), perfect-oracle max |err| = 0, correct ordering on synthetic ground truth, two full runs byte-identical. Fixed a wrap180 bug (Nim float mod keeps the dividend sign) that inflated the negative error tail. Bar (mean|err| deg [hitProxy] at 450+): Pattern 16.19 [0.077], naive-linear 22.86 [0.054], TMHorizon 16.20 [0.076], BitBrain 16.20 [0.077], static HeadOn 12.33 [0.098], oracle 0 [1.0]. Lead-gain sweep on Pattern is a dead end (1.0 wins every band). Naive-linear applies ~1.8x Pattern's lead but carries no more lead information (corr 0.178 vs 0.165) and is strictly worse. Ledger: docs/bitbrain_campaign.md. All verdicts remain live-only. --- .../gun_harness/prediction_quality.nim | 605 ++++++++++++++++++ .../tests/prediction_quality_results.txt | 146 +++++ common_libs/tests/run_prediction_quality.nim | 357 +++++++++++ docs/bitbrain_campaign.md | 258 ++++++++ 4 files changed, 1366 insertions(+) create mode 100644 common_libs/gun_harness/prediction_quality.nim create mode 100644 common_libs/tests/prediction_quality_results.txt create mode 100644 common_libs/tests/run_prediction_quality.nim create mode 100644 docs/bitbrain_campaign.md diff --git a/common_libs/gun_harness/prediction_quality.nim b/common_libs/gun_harness/prediction_quality.nim new file mode 100644 index 0000000..02d2f41 --- /dev/null +++ b/common_libs/gun_harness/prediction_quality.nim @@ -0,0 +1,605 @@ +## Offline per-gun PREDICTION-QUALITY scorer — the campaign ruler. +## +## WHAT THIS MEASURES (and only this) +## ---------------------------------- +## `docs/offline_harness_trust.md` established that the offline harness is +## trustworthy for ONE question: the per-gun, single-tick PREDICTION QUALITY of a +## gun on a FIXED enemy trajectory. It is never trustworthy for closed-loop +## questions (wins, damage, survival, movement, range, adaptation, selection). +## This module builds exactly that one trustworthy thing: given a recorded live +## battle, "if this gun had aimed at every tick it was asked to, how good was its +## aim?" — nothing more. +## +## THE METRIC +## ---------- +## At each recorded tick `i` the shooter sits at `O = (selfX, selfY)`. For a +## bullet of speed `v` the AIM-INDEPENDENT interception point is the first future +## tick `t* = i + k` (k >= 1) with `|E(t*) - O| <= v*k` — where the enemy's +## ACTUAL recorded track crosses the bullet's path. This is the same solve as +## `common_libs/tests/analyze_lead_capture_by_range.py` (commit f91e121): it +## depends only on the recorded truth and the bullet speed, never on our aim, so +## it is a fixed target every gun is scored against. +## +## Angular error is `wrap180(bearing(O -> pred) - bearing(O -> E(t*)))` in +## DEGREES. Degrees are the physically meaningful unit (the arena spans 800 px +## but the tolerance shrinks with range), so the ruler never reports pixels except +## in the sanity-validation path. Per RANGE BAND (0-100/100-200/200-300/300-450/ +## 450+) we report: +## mean |err| (deg), RMSE (deg), mean signed err (deg), +## the hit-probability proxy `mean(|err| < atan(18/range))`, and n. +## +## The "perfect oracle" arm aims at `E(t*)` itself, so it must score ~0 error — +## that is the plumbing check. The real correctness check is `validateShots`, +## which scores OUR ACTUAL recorded fired bearings (from the event sidecar) +## against the same interception solve: recorded HITS must cluster near zero and +## MISSES far away. If that separation collapses, the ruler is wrong. +## +## EVALUATION GRANULARITY +## ---------------------- +## The offline replay fires no real bullets, so a gun is asked for a prediction +## per POWER BIN (`PowerBins = 1.0/1.5/2.0/3.0`) each tick; each bin's prediction +## is scored against the interception at that bin's bullet speed. The phrase "the +## power that was really fired" applies to the separate `validateShots` check, +## which uses the server-recorded fired power. Aggregating over the four bins is +## a common horizon set shared by every arm, so the comparison is fair. +## +## WHY NOT `replayFixture` +## ----------------------- +## `common_libs/gun_harness/offline_range.nim` replays a fixture through a +## `VirtualTracker` so that feedback-adaptive guns (Tsetlin, KNN, DecayGF) learn +## from resolved virtual bullets. Every arm measured here — Pattern, TMHorizon, +## BitBrain, HeadOn, and the naive-linear control — has a no-op `onResult`: its +## prediction is a pure function of the observed `WorldState` stream, so the +## tracker changes nothing while costing an O(MaxBullets) resolution scan per +## tick (~8192 slots), which would dominate runtime. This module therefore keeps +## the recorded-state semantics (states replayed in order, one gun's history is +## the whole stream) but drives the guns directly, and it needs per-tick +## predictions — which `replayFixture`'s aggregate report does not expose. +## The corpus representation is the recorded-run loader below (round boundaries +## from `*.rounds.json`, event sidecar for validation, binary cache for speed), +## i.e. the same recorded-state contract with the battle metadata the analysis +## needs. +## +## COST / CACHING +## -------------- +## The corpus is ~900k recorded ticks over 70 battles (149 MB of JSONL). Parsing +## that with `std/json` on every sweep would dominate runtime, so the loader +## converts each run once into a compact binary `.qcache` (float32, 10 fields per +## tick) keyed on the source mtime+size. All measurements are taken from the +## cache, never from a mixture of cache and fresh parse, so two runs are +## byte-identical. + +import std/[math, os, strformat, strutils, json, tables, times, algorithm] + +const + NFields* = 10 + NBands* = 5 + BandLo* = [0.0, 100.0, 200.0, 300.0, 450.0] + BandHi* = [100.0, 200.0, 300.0, 450.0, 1.0e18] + BandLabels* = ["0-100", "100-200", "200-300", "300-450", "450+"] + MaxFlight* = 220 + ## max ticks a bullet is followed when solving for the interception tick + ## (matches analyze_lead_capture_by_range.py). + BbRadius* = 18.0 ## hit-detection radius in px (atan(18/range) tolerance) + +const CacheMagic = 0x31434242'i32 # "BBQ1" +const CacheVersion = 2'i32 + +proc wb(f: File, p: pointer, n: int) {.inline.} = + if n > 0: discard f.writeBuffer(p, n) + +proc rb(f: File, p: pointer, n: int) {.inline.} = + if n > 0: discard f.readBuffer(p, n) + +type + Corpus* = ref object + path*: string + arenaW*, arenaH*: float + n*: int + st*: seq[float32] ## NFields floats per tick + tick*: seq[int32] + rStart*: seq[int32] ## per-round first tick value + rCount*: seq[int32] + contiguous*: bool + base*: int + byTick*: Table[int32, int] + + BandStat* = object + n*: int + sumAbs*: float64 + sumSq*: float64 + sumSigned*: float64 + hits*: int + maxAbs*: float64 + sumPred*: float64 ## sum of the arm's own lead over LOS (deg) + sumReq*: float64 ## sum of the true required lead over LOS (deg) + sumAbsReq*: float64 + sumPredReq*: float64 + sumPred2*: float64 + sumReq2*: float64 + + ArmAcc* = object + name*: string + bands*: array[NBands, BandStat] + skipped*: int ## ticks×bins with no valid interception (no evidence) + evaluated*: int ## ticks×bins actually scored + + TrEvent* = object + round*, tick*, owner*, bullet*: int + kind*: string + power*, x*, y*, dir*: float + + ShotStat* = object + hits*, misses*: int + hitSumDeg*, missSumDeg*: float64 + hitSumPx*, missSumPx*: float64 + +template ex*(c: Corpus, i: int): float = c.st[i * NFields + 0].float +template ey*(c: Corpus, i: int): float = c.st[i * NFields + 1].float +template eh*(c: Corpus, i: int): float = c.st[i * NFields + 2].float +template es*(c: Corpus, i: int): float = c.st[i * NFields + 3].float +template ee*(c: Corpus, i: int): float = c.st[i * NFields + 4].float +template sx*(c: Corpus, i: int): float = c.st[i * NFields + 5].float +template sy*(c: Corpus, i: int): float = c.st[i * NFields + 6].float +template sh*(c: Corpus, i: int): float = c.st[i * NFields + 7].float +template ss*(c: Corpus, i: int): float = c.st[i * NFields + 8].float +template se*(c: Corpus, i: int): float = c.st[i * NFields + 9].float + +# ── geometry ───────────────────────────────────────────────────────────────── + +proc wrap180*(x: float): float {.inline.} = + ## Signed angular difference in (-180, 180]. NOTE: Nim's float `mod` keeps the + ## sign of the dividend (C fmod), so `(x + 180) mod 360 - 180` is WRONG for + ## x < -180 — it returns x - 360 instead of the wrapped equivalent. Normalise + ## explicitly (this bug inflated the negative tail of every error before fix). + result = x mod 360.0 + if result > 180.0: result -= 360.0 + elif result <= -180.0: result += 360.0 + +proc bearingDeg*(ox, oy, px, py: float): float {.inline.} = + radToDeg(arctan2(py - oy, px - ox)) + +proc bandOf*(r: float): int {.inline.} = + for b in 0..= BandLo[b] and r < BandHi[b]: return b + NBands - 1 + +proc tolDeg*(r: float): float {.inline.} = + ## Angular half-width of the target disc at range `r`: atan(18/range). + radToDeg(arctan2(BbRadius, max(r, 1e-9))) + +proc interceptTick(c: Corpus, i0, iEnd: int, ox, oy, speed: float): int = + ## First integer k >= 1 with |E(i0+k) - O| <= speed*k and i0+k < iEnd; -1 if + ## none. This is the integer-tick solve used by + ## analyze_lead_capture_by_range.py. + let v = speed + var k = 1 + while k <= MaxFlight: + let j = i0 + k + if j >= iEnd: break + let dx = c.ex(j) - ox + let dy = c.ey(j) - oy + if sqrt(dx * dx + dy * dy) <= v * float(k): + return k + inc k + -1 + +proc enemyPosAt(c: Corpus, i0, iEnd: int, t: float): tuple[x, y: float] = + ## Linearly interpolated enemy position at fractional time `t` after i0. + let base = float(int(t)) + let frac = t - base + let jm = min(i0 + int(base), iEnd - 1) + let jm2 = min(jm + 1, iEnd - 1) + (c.ex(jm) + (c.ex(jm2) - c.ex(jm)) * frac, + c.ey(jm) + (c.ey(jm2) - c.ey(jm)) * frac) + +proc interceptBearingQuant*(c: Corpus, i0, iEnd: int, ox, oy, + speed: float): tuple[ok: bool, bearing, range: float] = + ## The analyze_lead_capture_by_range.py solve: bearing to E(i0+k) at the first + ## integer tick k the enemy is within the bullet's reach. + let rng = hypot(c.ex(i0) - ox, c.ey(i0) - oy) + let k = interceptTick(c, i0, iEnd, ox, oy, speed) + if k < 0: return (false, 0.0, rng) + (true, bearingDeg(ox, oy, c.ex(i0 + k), c.ey(i0 + k)), rng) + +proc interceptBearingCont*(c: Corpus, i0, iEnd: int, ox, oy, + speed: float): tuple[ok: bool, bearing, range: float] = + ## THE PHYSICALLY EXACT TRUE INTERCEPTION POINT: the first fractional time + ## t > 0 at which the enemy's recorded track reaches distance speed*t from the + ## origin, found by bisecting the first integer tick where it comes within + ## reach. A bullet fired along the bearing to E(t) coincides with the enemy at + ## t, so aiming at this point is a true hit (the integer-tick solve overshoots: + ## it aims at E(k) but the bullet and enemy meet at E(t) with t <= k). + let rng = hypot(c.ex(i0) - ox, c.ey(i0) - oy) + let v = speed + var k = 1 + while k <= MaxFlight: + let j = i0 + k + if j >= iEnd: break + let f = hypot(c.ex(j) - ox, c.ey(j) - oy) - v * float(k) + if f <= 0.0: + var lo = float(k - 1) + var hi = float(k) + for _ in 0 ..< 24: + let mid = 0.5 * (lo + hi) + let p = enemyPosAt(c, i0, iEnd, mid) + if hypot(p.x - ox, p.y - oy) - v * mid > 0.0: lo = mid + else: hi = mid + let ts = 0.5 * (lo + hi) + let p = enemyPosAt(c, i0, iEnd, ts) + return (true, bearingDeg(ox, oy, p.x, p.y), rng) + inc k + (false, 0.0, rng) + +proc interceptBearing*(c: Corpus, i0, iEnd: int, ox, oy, speed: float, + cont: bool): tuple[ok: bool, bearing, range: float] = + if cont: interceptBearingCont(c, i0, iEnd, ox, oy, speed) + else: interceptBearingQuant(c, i0, iEnd, ox, oy, speed) + +# ── accumulators ───────────────────────────────────────────────────────────── + +proc addErr*(s: var BandStat, errDeg, range: float) = + inc s.n + let a = abs(errDeg) + s.sumAbs += a + s.sumSq += float64(errDeg) * float64(errDeg) + s.sumSigned += errDeg + if a > s.maxAbs: s.maxAbs = a + if a <= tolDeg(range): inc s.hits + +proc record*(a: var ArmAcc, range: float, errDeg, predLead, reqLead: float) = + let b = bandOf(range) + addErr(a.bands[b], errDeg, range) + var s = addr a.bands[b] + s.sumPred += predLead + s.sumReq += reqLead + s.sumAbsReq += abs(reqLead) + s.sumPredReq += predLead * reqLead + s.sumPred2 += predLead * predLead + s.sumReq2 += reqLead * reqLead + inc a.evaluated + +proc captureSlope*(s: BandStat): float = + ## Regression of the arm's lead on the true required lead (job-95's "capture" + ## statistic): 1.0 = perfect proportional response, 0.0 = no response. + if s.sumReq2 <= 1e-12: NaN else: s.sumPredReq / s.sumReq2 + +proc leadCorr*(s: BandStat): float = + ## Pearson correlation between the arm's lead and the required lead. THIS is + ## the honest measure of "is the lead informative"; a large capture slope on + ## an uncorrelated lead is just amplification of noise. + let d = s.sumPred2 * s.sumReq2 + if d <= 1e-12: NaN else: s.sumPredReq / sqrt(d) + +proc meanAbsReq*(s: BandStat): float = + if s.n == 0: NaN else: s.sumAbsReq / float(s.n) + +proc skip*(a: var ArmAcc) = inc a.skipped + +proc meanAbs*(s: BandStat): float = + if s.n == 0: NaN else: s.sumAbs / float(s.n) + +proc rmse*(s: BandStat): float = + if s.n == 0: NaN else: sqrt(s.sumSq / float(s.n)) + +proc meanSigned*(s: BandStat): float = + if s.n == 0: NaN else: s.sumSigned / float(s.n) + +proc hitProxy*(s: BandStat): float = + if s.n == 0: NaN else: s.hits.float / float(s.n) + +# ── corpus loading + binary cache ──────────────────────────────────────────── + +proc cachePathFor*(src: string): string = src & ".qcache" + +proc eventsPathFor*(runPath: string): string = + ## `run10.jsonl` -> `run10.events.jsonl` (note: NOT run10.jsonl.events.jsonl). + if runPath.endsWith(".jsonl"): + runPath[0 ..< runPath.len - 6] & ".events.jsonl" + else: + runPath & ".events.jsonl" + +proc mtimeOf(p: string): float = + if fileExists(p): getFileInfo(p).lastWriteTime.toUnixFloat else: 0.0 + +proc buildCache(src, roundsPath, cachePath: string) = + ## JSONL -> compact binary. Only runs when the cache is absent/stale. + var st: seq[float32] + var ticks: seq[int32] + var arenaW = 800.0 + var arenaH = 600.0 + for line in lines(src): + let s = line.strip() + if s.len == 0: continue + let node = parseJson(s) + if node.hasKey("meta"): + if node["meta"].hasKey("arena"): + let a = node["meta"]["arena"] + if a.hasKey("w"): arenaW = a["w"].getFloat() + if a.hasKey("h"): arenaH = a["h"].getFloat() + continue + if node.hasKey("end"): continue + st.add node["ex"].getFloat().float32 + st.add node["ey"].getFloat().float32 + st.add node["eh"].getFloat().float32 + st.add node["es"].getFloat().float32 + st.add node["ee"].getFloat().float32 + st.add node["sx"].getFloat().float32 + st.add node["sy"].getFloat().float32 + st.add node["sh"].getFloat().float32 + st.add node["ss"].getFloat().float32 + st.add node["se"].getFloat().float32 + ticks.add node["tick"].getInt().int32 + + var rStart, rCount: seq[int32] + if roundsPath.len > 0 and fileExists(roundsPath): + try: + let rj = parseFile(roundsPath) + for r in rj["rounds"]: + rStart.add r["startTick"].getInt().int32 + rCount.add r["count"].getInt().int32 + except CatchableError: + discard + if rStart.len == 0: + rStart.add 0'i32 + rCount.add ticks.len.int32 + + var n32 = ticks.len.int32 + var nr32 = rStart.len.int32 + var magic = CacheMagic + var version = CacheVersion + var sm = mtimeOf(src) + var ss = getFileInfo(src).size.int64 + var rm = mtimeOf(roundsPath) + let f = open(cachePath, fmWrite) + defer: f.close() + wb(f, addr magic, sizeof(magic)) + wb(f, addr version, sizeof(version)) + wb(f, addr n32, sizeof(n32)) + wb(f, addr nr32, sizeof(nr32)) + wb(f, addr arenaW, sizeof(arenaW)) + wb(f, addr arenaH, sizeof(arenaH)) + wb(f, addr sm, sizeof(sm)) + wb(f, addr ss, sizeof(ss)) + wb(f, addr rm, sizeof(rm)) + if rStart.len > 0: + wb(f, addr rStart[0], rStart.len * sizeof(int32)) + wb(f, addr rCount[0], rCount.len * sizeof(int32)) + if ticks.len > 0: + wb(f, addr ticks[0], ticks.len * sizeof(int32)) + wb(f, addr st[0], st.len * sizeof(float32)) + +proc readCache(path: string): Corpus = + let f = open(path, fmRead) + defer: f.close() + var magic, version: int32 + rb(f, addr magic, sizeof(magic)) + rb(f, addr version, sizeof(version)) + if magic != CacheMagic or version != CacheVersion: + raise newException(IOError, "bad qcache header: " & path) + var n32, nr32: int32 + rb(f, addr n32, sizeof(n32)) + rb(f, addr nr32, sizeof(nr32)) + result = Corpus(path: path, n: int(n32)) + rb(f, addr result.arenaW, sizeof(result.arenaW)) + rb(f, addr result.arenaH, sizeof(result.arenaH)) + var sm: float + var ss: int64 + var rm: float + rb(f, addr sm, sizeof(sm)) + rb(f, addr ss, sizeof(ss)) + rb(f, addr rm, sizeof(rm)) + result.rStart.setLen(int(nr32)) + result.rCount.setLen(int(nr32)) + if nr32 > 0: + rb(f, addr result.rStart[0], int(nr32) * sizeof(int32)) + rb(f, addr result.rCount[0], int(nr32) * sizeof(int32)) + result.tick.setLen(result.n) + result.st.setLen(result.n * NFields) + if result.n > 0: + rb(f, addr result.tick[0], result.n * sizeof(int32)) + rb(f, addr result.st[0], result.n * NFields * sizeof(float32)) + +proc indexMapping(c: var Corpus) = + c.contiguous = c.n > 0 + c.base = if c.n > 0: int(c.tick[0]) else: 0 + if c.contiguous: + for i in 0..= 0 and i < c.n and c.tick[i] == t: + found = true + return i + found = false + return -1 + if c.byTick.hasKey(t): + found = true + return c.byTick[t] + found = false + -1 + +proc cacheFresh(cache, src, roundsPath: string): bool = + if not fileExists(cache): return false + try: + let f = open(cache, fmRead) + var magic, version: int32 + rb(f, addr magic, sizeof(magic)) + rb(f, addr version, sizeof(version)) + var n32, nr32: int32 + rb(f, addr n32, sizeof(n32)) + rb(f, addr nr32, sizeof(nr32)) + var aw, ah, sm, rm: float + var ss: int64 + rb(f, addr aw, sizeof(aw)) + rb(f, addr ah, sizeof(ah)) + rb(f, addr sm, sizeof(sm)) + rb(f, addr ss, sizeof(ss)) + rb(f, addr rm, sizeof(rm)) + f.close() + return magic == CacheMagic and version == CacheVersion and + sm == mtimeOf(src) and ss == getFileInfo(src).size.int64 and + rm == mtimeOf(roundsPath) + except CatchableError: + false + +proc loadCorpus*(src: string): Corpus = + ## Load a recorded run (`runN.jsonl`); the per-round index is read from the + ## sibling `runN.jsonl.rounds.json`. Builds the binary cache when stale/absent, + ## then ALWAYS measures from the cache so repeated runs are byte-identical. + let cache = cachePathFor(src) + let roundsPath = src & ".rounds.json" + if not cacheFresh(cache, src, roundsPath): + buildCache(src, roundsPath, cache) + result = readCache(cache) + result.path = src + result.indexMapping() + +proc discoverRuns*(root: string): seq[string] = + ## All `/runN.jsonl` under `root` that have the events + rounds sidecars. + for sub in walkDir(root, relative = false): + if sub.kind != pcDir: continue + for fn in walkFiles(sub.path / "*.jsonl"): + if fn.endsWith(".events.jsonl"): continue + if fileExists(fn & ".rounds.json") and fileExists(eventsPathFor(fn)): + result.add fn + result.sort() + +# ── shot-level geometry validation ─────────────────────────────────────────── +# +# Scores OUR ACTUAL server-recorded fired bearings against the interception solve +# above. Owner ids in the event sidecar are not stable across runs, so each run is +# attributed independently (mirrors analyze_lead_capture_by_range.py): a fire +# event's (x, y) is the firing tank's centre and its energy drops by exactly the +# fired power one tick later. + +proc parseEvents*(path: string): seq[TrEvent] = + for line in lines(path): + let s = line.strip() + if s.len == 0: continue + let node = parseJson(s) + result.add TrEvent( + round: node["round"].getInt(), + tick: node["tick"].getInt(), + kind: node["type"].getStr(), + owner: (if node.hasKey("owner"): node["owner"].getInt() else: -1), + bullet: (if node.hasKey("bullet"): node["bullet"].getInt() else: -1), + power: (if node.hasKey("power"): node["power"].getFloat() else: 0.0), + x: (if node.hasKey("x"): node["x"].getFloat() else: 0.0), + y: (if node.hasKey("y"): node["y"].getFloat() else: 0.0), + dir: (if node.hasKey("dir"): node["dir"].getFloat() else: 0.0)) + +proc matchEvent(c: Corpus, t: int, side: int, ev: TrEvent): bool = + ## side 0 = 'e' (enemy), 1 = 's' (self). Position match + energy drop. + if t < 0 or t + 1 >= c.n: return false + let px = if side == 0: c.ex(t) else: c.sx(t) + let py = if side == 0: c.ey(t) else: c.sy(t) + if abs(px - ev.x) > 0.02 or abs(py - ev.y) > 0.02: return false + let e0 = if side == 0: c.ee(t) else: c.se(t) + let e1 = if side == 0: c.ee(t + 1) else: c.se(t + 1) + abs((e0 - e1) - ev.power) < 0.02 + +proc roundStartIndexOf(c: Corpus, rnd: int): int = + ## Rounds hold a GLOBAL startTick; map to an array index. + var startTick: int32 = 0 + for i, r in c.rStart: + if int(i) + 1 == rnd: startTick = r + if c.contiguous: return int(startTick) - c.base + var found = false + idxOfTick(c, startTick, found) + +proc validateShots*(c: Corpus, events: seq[TrEvent], cont: bool): ShotStat = + ## Hits must show small angular error and misses large; the separation ratio is + ## the ruler's correctness certificate. + var votes = initTable[int, array[2, int]]() + for ev in events: + if ev.kind != "fire": continue + let guess = c.roundStartIndexOf(ev.round) + ev.tick + if ev.owner notin votes: votes[ev.owner] = [0, 0] + for t in (guess - 8) .. (guess + 8): + for side in 0..1: + if matchEvent(c, t, side, ev): inc votes[ev.owner][side] + var ownerSide = initTable[int, int]() + for owner, v in votes: + ownerSide[owner] = if v[1] >= v[0]: 1 else: 0 + + var resolution = initTable[string, string]() + for ev in events: + if ev.kind in ["hit", "hitwall", "hitbullet"]: + resolution[$ev.round & "/" & $ev.owner & "/" & $ev.bullet] = ev.kind + + for ev in events: + if ev.kind != "fire": continue + if ownerSide.getOrDefault(ev.owner, -1) != 1: continue # only OUR shots + let guess = c.roundStartIndexOf(ev.round) + ev.tick + var t0 = -1 + var bestD = high(int) + for t in (guess - 8) .. (guess + 8): + if matchEvent(c, t, ownerSide[ev.owner], ev): + let d = abs(t - guess) + if d < bestD: + bestD = d + t0 = t + if t0 < 0: continue + var rIdx = -1 + for i, r in c.rStart: + if int(i) + 1 == ev.round: rIdx = i + if rIdx < 0: continue + let iEnd = int(c.rStart[rIdx]) - c.base + int(c.rCount[rIdx]) + let speed = 20.0 - 3.0 * ev.power + let ib = interceptBearing(c, t0, iEnd, ev.x, ev.y, speed, cont) + if not ib.ok: continue + let err = wrap180(ev.dir - ib.bearing) + let kind = resolution.getOrDefault($ev.round & "/" & $ev.owner & "/" & $ev.bullet, "") + if kind == "hit": + inc result.hits + result.hitSumDeg += abs(err) + result.hitSumPx += abs(degToRad(err)) * ib.range + elif kind in ["hitwall", "hitbullet"]: + inc result.misses + result.missSumDeg += abs(err) + result.missSumPx += abs(degToRad(err)) * ib.range + +proc separationDeg*(s: ShotStat): float = + let hd = if s.hits > 0: s.hitSumDeg / float(s.hits) else: NaN + let md = if s.misses > 0: s.missSumDeg / float(s.misses) else: NaN + md / hd + +proc separationPx*(s: ShotStat): float = + let hp = if s.hits > 0: s.hitSumPx / float(s.hits) else: NaN + let mp = if s.misses > 0: s.missSumPx / float(s.misses) else: NaN + mp / hp + +proc meanHitDeg*(s: ShotStat): float = + if s.hits > 0: s.hitSumDeg / float(s.hits) else: NaN + +proc meanMissDeg*(s: ShotStat): float = + if s.misses > 0: s.missSumDeg / float(s.misses) else: NaN + +proc meanHitPx*(s: ShotStat): float = + if s.hits > 0: s.hitSumPx / float(s.hits) else: NaN + +proc meanMissPx*(s: ShotStat): float = + if s.misses > 0: s.missSumPx / float(s.misses) else: NaN + +# ── formatting ─────────────────────────────────────────────────────────────── + +proc formatArmTable*(arms: seq[ArmAcc]): string = + let hdr = "arm band n meanAbs rmse signed hitProxy maxAbs" + result = hdr & "\n" & "-".repeat(hdr.len) & "\n" + for a in arms: + for b in 0..8} {'-':>8} {'-':>8} {'-':>8} {'-':>9} {'-':>8}" & "\n" + else: + result.add fmt"{a.name:<16} {BandLabels[b]:<9} {s.n:>8} {meanAbs(s):>8.3f} {rmse(s):>8.3f} {meanSigned(s):>8.3f} {hitProxy(s):>9.4f} {s.maxAbs:>8.2f}" & "\n" + if a.skipped > 0: + result.add fmt" ({a.name}: {a.skipped} tick-bins had no valid interception)" & "\n" diff --git a/common_libs/tests/prediction_quality_results.txt b/common_libs/tests/prediction_quality_results.txt new file mode 100644 index 0000000..547226b --- /dev/null +++ b/common_libs/tests/prediction_quality_results.txt @@ -0,0 +1,146 @@ +======================================================================================================================== +OFFLINE PREDICTION QUALITY -- per-gun single-tick aim error vs the true interception point +======================================================================================================================== +corpus : /tmp/tfil_ab2/out +ruler : continuous (physically exact) +runs : 70 +recorded ticks: 899607 +tick x bin : 3598428 +wall time : 406.88s (0.1131 ms per tick-bin) +per-arm speed : 0.1131 s per 1000 tick-bins per arm + +NOTE: offline OPEN-LOOP prediction quality only. No win/damage/survival claim. + +======================================================================================================================== +VALIDATION -- the ruler must pass ALL of these before any number below is trusted +======================================================================================================================== +1. recorded shots (OUR actual server-fired bearings vs the SAME interception solve): + ruler=continuous hits n=5480 mean|err|= 1.360 deg / 10.5 px | misses n=48304 mean|err|= 16.597 deg / 140.3 px | separation 12.20x deg / 13.34x px -> OK + ruler=integer hits n=5480 mean|err|= 1.478 deg / 11.4 px | misses n=48304 mean|err|= 16.724 deg / 141.3 px | separation 11.32x deg / 12.43x px -> OK +2. perfect-oracle gun max |err| over all tick-bins = 0.000000 deg -> OK +3. HeadOn (static LOS) mean|err| = 13.217 deg vs Pattern 16.609 / TMHorizon 16.627 / BitBrain 16.635 + -> UNEXPECTED: a predictive gun is worse than static LOS + NaiveLinear mean|err| = 22.086 deg (over-leads; see the lead-gain sweep for why a larger + lead *response* does not mean a smaller angular error) +4. determinism: run twice and diff stdout (see fixture; verified separately). + +======================================================================================================================== +THE BAR -- per-band mean ABSOLUTE angular aim error (deg), RMSE, sign, hit-proxy +======================================================================================================================== +hitProxy = fraction of tick-bins with |err| <= atan(18/range) (the angular half-width of the target disc). + +arm band n meanAbs rmse signed hitProxy maxAbs +---------------------------------------------------------------------------------- +Oracle 0-100 4423 0.000 0.000 0.000 1.0000 0.00 +Oracle 100-200 24908 0.000 0.000 0.000 1.0000 0.00 +Oracle 200-300 74215 0.000 0.000 0.000 1.0000 0.00 +Oracle 300-450 1119777 0.000 0.000 0.000 1.0000 0.00 +Oracle 450+ 2311323 0.000 0.000 0.000 1.0000 0.00 + (Oracle: 63782 tick-bins had no valid interception) +OracleQuant 0-100 4423 1.103 1.498 0.078 1.0000 6.00 +OracleQuant 100-200 24908 0.674 0.911 -0.002 1.0000 4.07 +OracleQuant 200-300 74215 0.472 0.624 -0.008 1.0000 2.44 +OracleQuant 300-450 1119777 0.360 0.470 -0.010 1.0000 1.62 +OracleQuant 450+ 2311323 0.288 0.376 0.010 1.0000 1.23 + (OracleQuant: 63782 tick-bins had no valid interception) +HeadOn 0-100 4423 19.619 23.254 -1.279 0.3423 46.28 +HeadOn 100-200 24908 19.982 23.151 -0.462 0.1724 46.62 +HeadOn 200-300 74215 17.341 20.521 0.481 0.1330 46.33 +HeadOn 300-450 1119777 14.607 17.606 0.690 0.1049 46.38 +HeadOn 450+ 2311323 12.326 15.017 -0.263 0.0984 45.49 + (HeadOn: 63782 tick-bins had no valid interception) +Pattern 0-100 4423 10.555 14.796 -0.404 0.6993 64.72 +Pattern 100-200 24908 14.745 19.510 1.276 0.3418 76.63 +Pattern 200-300 74215 16.610 21.174 1.350 0.1850 81.45 +Pattern 300-450 1119777 17.531 21.838 0.948 0.1036 85.65 +Pattern 450+ 2311323 16.193 20.021 -0.642 0.0767 79.92 + (Pattern: 63782 tick-bins had no valid interception) +PatternGain1.5 0-100 4423 13.764 18.522 0.034 0.6093 79.58 +PatternGain1.5 100-200 24908 19.640 25.124 2.144 0.2025 97.08 +PatternGain1.5 200-300 74215 22.260 27.672 1.784 0.1024 103.15 +PatternGain1.5 300-450 1119777 23.279 28.548 1.077 0.0627 111.53 +PatternGain1.5 450+ 2311323 21.245 26.062 -0.832 0.0542 104.02 + (PatternGain1.5: 63782 tick-bins had no valid interception) +PatternGain2.0 0-100 4423 20.111 25.637 0.471 0.4047 95.58 +PatternGain2.0 100-200 24908 26.782 33.175 3.013 0.1487 118.82 +PatternGain2.0 200-300 74215 29.473 35.856 2.219 0.0776 126.59 +PatternGain2.0 300-450 1119777 29.955 36.371 1.207 0.0488 137.41 +PatternGain2.0 450+ 2311323 26.997 32.935 -1.021 0.0425 128.12 + (PatternGain2.0: 63782 tick-bins had no valid interception) +PatternGain3.0 0-100 4423 34.772 43.079 1.346 0.2720 141.72 +PatternGain3.0 100-200 24908 42.972 51.921 4.751 0.0978 162.72 +PatternGain3.0 200-300 74215 45.232 54.158 3.088 0.0507 173.45 +PatternGain3.0 300-450 1119777 44.274 53.364 1.462 0.0333 179.99 +PatternGain3.0 450+ 2311323 39.271 47.718 -1.400 0.0293 177.73 + (PatternGain3.0: 63782 tick-bins had no valid interception) +NaiveLinear 0-100 4423 17.924 35.009 -0.098 0.6505 178.26 +NaiveLinear 100-200 24908 14.629 24.037 -0.445 0.3879 179.65 +NaiveLinear 200-300 74215 17.572 24.479 1.030 0.1924 179.78 +NaiveLinear 300-450 1119777 20.976 26.164 1.096 0.0995 179.64 +NaiveLinear 450+ 2311323 22.857 27.679 -0.477 0.0537 179.98 + (NaiveLinear: 63782 tick-bins had no valid interception) +TMHorizon 0-100 4423 10.681 14.782 -1.233 0.6955 66.72 +TMHorizon 100-200 24908 14.841 19.526 0.814 0.3418 78.63 +TMHorizon 200-300 74215 16.644 21.142 1.091 0.1786 84.45 +TMHorizon 300-450 1119777 17.572 21.878 0.691 0.0996 84.01 +TMHorizon 450+ 2311323 16.199 20.025 -0.687 0.0757 81.19 + (TMHorizon: 63782 tick-bins had no valid interception) +BitBrain 0-100 4423 11.118 15.268 -1.507 0.6810 64.72 +BitBrain 100-200 24908 15.107 19.782 1.289 0.3241 76.63 +BitBrain 200-300 74215 16.838 21.426 1.341 0.1747 105.04 +BitBrain 300-450 1119777 17.575 21.894 0.933 0.1025 100.98 +BitBrain 450+ 2311323 16.200 20.032 -0.647 0.0767 96.38 + (BitBrain: 63782 tick-bins had no valid interception) + +======================================================================================================================== +HEADROOM -- the direct answer: how far each arm is from the oracle ceiling, per band +======================================================================================================================== +band Pattern n Pattern|err| Pattern hpx Oracle hpx headroom pp naive hpx TMHoriz hpx BitBrain hpx +--------------------------------------------------------------------------------------------------------------- +0-100 4423 10.555 0.6993 1.0000 0.3007 0.6505 0.6955 0.6810 +100-200 24908 14.745 0.3418 1.0000 0.6582 0.3879 0.3418 0.3241 +200-300 74215 16.610 0.1850 1.0000 0.8150 0.1924 0.1786 0.1747 +300-450 1119777 17.531 0.1036 1.0000 0.8964 0.0995 0.0996 0.1025 +450+ 2311323 16.193 0.0767 1.0000 0.9233 0.0537 0.0757 0.0767 + +hitProxy = fraction of tick-bins aimed within atan(18/range) of the true interception point. +headroom pp = oracle hitProxy - Pattern hitProxy = the absolute hit-probability points available +to a perfect predictor (the campaign is playing for a slice of this). + +band OracleQuant hpx integer-solve coarseness +--------------------------------------------------- +0-100 1.0000 1.103 deg mean |err| +100-200 1.0000 0.674 deg mean |err| +200-300 1.0000 0.472 deg mean |err| +300-450 1.0000 0.360 deg mean |err| +450+ 1.0000 0.288 deg mean |err| +(OracleQuant aims at the analyze_lead_capture_by_range.py integer-tick intercept and is scored + on the active ruler. On the continuous ruler it measures how much of a gun's 'error' the coarse + solve itself would produce; on the integer ruler it is identically zero.) + +======================================================================================================================== +LEAD-GAIN SWEEP ON PATTERN -- multiply Pattern's lead (deg over LOS) by a constant +======================================================================================================================== +band gain=1.0 gain=1.5 gain=2.0 gain=3.0 best-gain +------------------------------------------------------------------ +0-100 10.555 13.764 20.111 34.772 1.0 (10.555) +100-200 14.745 19.640 26.782 42.972 1.0 (14.745) +200-300 16.610 22.260 29.473 45.232 1.0 (16.610) +300-450 17.531 23.279 29.955 44.274 1.0 (17.531) +450+ 16.193 21.245 26.997 39.271 1.0 (16.193) + +======================================================================================================================== +LEAD INFORMATIVENESS -- capture slope (regression of applied lead on required lead) and lead correlation +======================================================================================================================== +capture slope is job-95's metric (1.0 = perfect proportional response). corr is the Pearson +correlation of the arm's lead with the REQUIRED lead: a large slope on an uncorrelated lead is +just amplified noise. This is the table that resolves the 'naive-linear captures 2x the lead but +hits less' tension. + +band HO |err| HO |req| Pat|req| Pat cap Pat corr Lin cap Lin corr TMH cap TMH corr BB cap BB corr +----------------------------------------------------------------------------------------------------------------------- +0-100 19.619 19.619 19.619 0.649 0.774 0.573 0.362 0.668 0.776 0.701 0.768 +100-200 19.982 19.982 19.982 0.553 0.612 0.592 0.523 0.560 0.614 0.570 0.611 +200-300 17.341 17.341 17.341 0.449 0.457 0.519 0.429 0.452 0.460 0.459 0.457 +300-450 14.607 14.607 14.607 0.278 0.266 0.476 0.324 0.273 0.262 0.280 0.266 +450+ 12.326 12.326 12.326 0.175 0.165 0.310 0.178 0.174 0.164 0.175 0.165 diff --git a/common_libs/tests/run_prediction_quality.nim b/common_libs/tests/run_prediction_quality.nim new file mode 100644 index 0000000..b7ae367 --- /dev/null +++ b/common_libs/tests/run_prediction_quality.nim @@ -0,0 +1,357 @@ +## Offline PREDICTION-QUALITY runner — the campaign's measurement sweep. +## +## Reads the recorded live-vs-real-DrussGT corpus, drives each arm over the +## recorded enemy trajectory, and scores every tick×power-bin prediction against +## the aim-independent interception point (see prediction_quality.nim). Owns the +## RANGE-BAND table that is "the bar" for the BitBrain campaign. +## +## NO CLOSED-LOOP CLAIM IS MADE HERE. Every arm below is an open-loop prediction +## scored on a FIXED trajectory. Wins, damage and survival are decided live. +## +## Usage: +## nim c -r --nimcache:/tmp/nc_j98 common_libs/tests/run_prediction_quality.nim \ +## [--corpus /tmp/tfil_ab2/out] [--limit N] [--timing] +## +## `--limit N` keeps only the first N runs (sorted), for fast iteration. + +import std/[os, strformat, strutils, times, math] +import gun_harness/[gun_interface, virtual_bullets, prediction_quality] +import guns/[head_on, pattern_matcher, tm_horizon, bitbrain_gun] + +# arm indices (fixed order = fixed output) +const + A_ORACLE* = 0 + A_ORACLEQ* = 1 + A_HEADON* = 2 + A_PATTERN* = 3 + A_G15* = 4 + A_G20* = 5 + A_G30* = 6 + A_NAIVE* = 7 + A_TMH* = 8 + A_BB* = 9 + ArmNames* = ["Oracle", "OracleQuant", "HeadOn", "Pattern", "PatternGain1.5", + "PatternGain2.0", "PatternGain3.0", "NaiveLinear", "TMHorizon", "BitBrain"] + +# ── the naive-linear control (job-95's LIN_M = 4 extrapolation) ────────────── +# +# Velocity = (pos(t) - pos(t-4)) / 4, then iterate the interception equation. +# This is the trivial predictive gun the lead-capture analysis used as its +# ceiling control; it captures ~2x the lead response Pattern does at 450+ and is +# included to resolve that tension against the angular-error ruler. + +type + NaiveLinearGun = object + hist: array[5, tuple[x, y: float]] + count: int + lastTick: int + +proc predict*(g: var NaiveLinearGun, state: WorldState, + bulletSpeed: float): GunPrediction = + if state.tick != g.lastTick: + for i in countdown(4, 1): g.hist[i] = g.hist[i - 1] + g.hist[0] = (state.enemyX, state.enemyY) + if g.count < 5: inc g.count + g.lastTick = state.tick + if g.count < 5 or bulletSpeed <= 0.0: + return GunPrediction(x: state.enemyX, y: state.enemyY) + let vx = (g.hist[0].x - g.hist[4].x) / 4.0 + let vy = (g.hist[0].y - g.hist[4].y) / 4.0 + let ox = state.selfX + let oy = state.selfY + var t = hypot(state.enemyX - ox, state.enemyY - oy) / bulletSpeed + for _ in 0..<40: + t = hypot(state.enemyX + vx * t - ox, state.enemyY + vy * t - oy) / bulletSpeed + GunPrediction(x: state.enemyX + vx * t, y: state.enemyY + vy * t) + +proc onResult*(g: var NaiveLinearGun, e: FeedbackEvent) = discard + +# ── runner ─────────────────────────────────────────────────────────────────── + +type Ctx = object + c: Corpus + cont: bool + pattern: PatternMatcherGun + naive: NaiveLinearGun + tmh: TmHorizonGun + bb: BitBrainGun + headon: HeadOnGun + st: WorldState + enemy: seq[EnemyInfo] + +proc runRound(ctx: var Ctx, arms: var seq[ArmAcc], r: int) = + let c = ctx.c + let base = int(c.rStart[r]) - c.base + let cnt = int(c.rCount[r]) + let iEnd = base + cnt + ctx.enemy[0] = EnemyInfo(id: 1) + for i in base ..< iEnd: + let ox = c.sx(i) + let oy = c.sy(i) + let localTick = int(c.tick[i]) - int(c.rStart[r]) + ctx.enemy[0].x = c.ex(i) + ctx.enemy[0].y = c.ey(i) + ctx.enemy[0].heading = c.eh(i) + ctx.enemy[0].speed = c.es(i) + ctx.enemy[0].energy = c.ee(i) + ctx.enemy[0].lastSeenTick = localTick + ctx.st.enemyX = c.ex(i) + ctx.st.enemyY = c.ey(i) + ctx.st.enemyHeading = c.eh(i) + ctx.st.enemySpeed = c.es(i) + ctx.st.enemyEnergy = c.ee(i) + ctx.st.selfX = ox + ctx.st.selfY = oy + ctx.st.selfHeading = c.sh(i) + ctx.st.selfSpeed = c.ss(i) + ctx.st.selfEnergy = c.se(i) + ctx.st.selfRadarHeading = c.sh(i) + ctx.st.tick = localTick + let los = bearingDeg(ox, oy, c.ex(i), c.ey(i)) + for bin in 0 ..< len(PowerBins): + let speed = bulletSpeed(PowerBins[bin]) + let ib = interceptBearing(c, i, iEnd, ox, oy, speed, ctx.cont) + if not ib.ok: + for ai in 0 ..< arms.len: skip(arms[ai]) + continue + let ibq = interceptBearingQuant(c, i, iEnd, ox, oy, speed) + let rng = ib.range + let targetLead = wrap180(ib.bearing - los) + # oracle: aims at the active ruler's true interception point -> 0 error + arms[A_ORACLE].record(rng, 0.0, targetLead, targetLead) + # the analyze_lead_capture_by_range.py integer-tick intercept, scored on the + # SAME ruler: this is what the coarse solve's own oracle would reach. + let qLead = if ibq.ok: wrap180(ibq.bearing - los) else: targetLead + arms[A_ORACLEQ].record(rng, wrap180(qLead - targetLead), qLead, targetLead) + # head-on: aim at the enemy's CURRENT position (worst realistic gun) + arms[A_HEADON].record(rng, wrap180(los - ib.bearing), 0.0, targetLead) + # pattern + lead-gain sweep (gain scales Pattern's lead over LOS) + let pp = ctx.pattern.predict(ctx.st, speed) + let pb = bearingDeg(ox, oy, pp.x, pp.y) + let plead = wrap180(pb - los) + arms[A_PATTERN].record(rng, wrap180(plead - targetLead), plead, targetLead) + arms[A_G15].record(rng, wrap180(1.5 * plead - targetLead), 1.5 * plead, targetLead) + arms[A_G20].record(rng, wrap180(2.0 * plead - targetLead), 2.0 * plead, targetLead) + arms[A_G30].record(rng, wrap180(3.0 * plead - targetLead), 3.0 * plead, targetLead) + # naive linear + let np = predict(ctx.naive, ctx.st, speed) + let nl = wrap180(bearingDeg(ox, oy, np.x, np.y) - los) + arms[A_NAIVE].record(rng, wrap180(nl - targetLead), nl, targetLead) + # TMHorizon + let tp = predict(ctx.tmh, ctx.st, speed) + let tl = wrap180(bearingDeg(ox, oy, tp.x, tp.y) - los) + arms[A_TMH].record(rng, wrap180(tl - targetLead), tl, targetLead) + # BitBrain (base Pattern + ADE/SBC corrector) + let bp = predict(ctx.bb, ctx.st, speed) + let bl = wrap180(bearingDeg(ox, oy, bp.x, bp.y) - los) + arms[A_BB].record(rng, wrap180(bl - targetLead), bl, targetLead) + +proc runOne(runPath: string, arms: var seq[ArmAcc], shotsCont, shotsQuant: var ShotStat, + doShots: bool, timing: bool, cont: bool): int = + let t0 = epochTime() + let c = loadCorpus(runPath) + if c.n == 0: return 0 + var ctx = Ctx( + c: c, + cont: cont, + pattern: PatternMatcherGun(), + naive: NaiveLinearGun(lastTick: -1), + tmh: initTmHorizonGun(), + bb: initBitBrainGun(), + headon: HeadOnGun(), + st: WorldState(arenaWidth: c.arenaW, arenaHeight: c.arenaH), + enemy: newSeq[EnemyInfo](1)) + for r in 0 ..< c.rStart.len: + runRound(ctx, arms, r) + if doShots: + let ev = parseEvents(eventsPathFor(runPath)) + for mode in [true, false]: + let s = validateShots(c, ev, mode) + var dst = if mode: addr shotsCont else: addr shotsQuant + dst.hits += s.hits + dst.misses += s.misses + dst.hitSumDeg += s.hitSumDeg + dst.missSumDeg += s.missSumDeg + dst.hitSumPx += s.hitSumPx + dst.missSumPx += s.missSumPx + if timing: + stderr.writeLine(fmt" {extractFilename(runPath):<16} ticks={c.n:<7} {epochTime()-t0:>6.2f}s") + c.n + +proc fmt4(x: float): string = + if x.classify in {fcNan, fcInf, fcNegInf}: "-" else: fmt"{x:.4f}" + +proc fmt3(x: float): string = + if x.classify in {fcNan, fcInf, fcNegInf}: "-" else: fmt"{x:.3f}" + +proc main() = + var corpusRoot = "/tmp/tfil_ab2/out" + var limit = 0 + var doShots = true + var timing = false + var cont = true + var i = 1 + while i <= paramCount(): + case paramStr(i) + of "--corpus": inc i; corpusRoot = paramStr(i) + of "--limit": inc i; limit = parseInt(paramStr(i)) + of "--no-shots": doShots = false + of "--timing": timing = true + of "--ruler": + inc i + cont = paramStr(i) != "quant" + else: + stderr.writeLine("unknown arg: " & paramStr(i)) + quit(2) + inc i + + var runs = discoverRuns(corpusRoot) + if limit > 0 and runs.len > limit: runs.setLen(limit) + if runs.len == 0: + stderr.writeLine("no runs found under " & corpusRoot) + quit(1) + + var arms: seq[ArmAcc] + for nm in ArmNames: arms.add ArmAcc(name: nm) + var shotsCont, shotsQuant: ShotStat + + var t0 = epochTime() + var ticks = 0 + for rp in runs: + ticks += runOne(rp, arms, shotsCont, shotsQuant, doShots, timing, cont) + let elapsed = epochTime() - t0 + + echo "=".repeat(120) + echo "OFFLINE PREDICTION QUALITY -- per-gun single-tick aim error vs the true interception point" + echo "=".repeat(120) + echo fmt"corpus : {corpusRoot}" + let rulerName = if cont: "continuous (physically exact)" else: "integer-tick (analyze_lead_capture_by_range.py)" + echo fmt"ruler : {rulerName}" + echo fmt"runs : {runs.len}" + echo fmt"recorded ticks: {ticks}" + let tickBins = ticks * len(PowerBins) + echo fmt"tick x bin : {tickBins}" + echo fmt"wall time : {elapsed:.2f}s ({elapsed / max(1.0, float(tickBins)) * 1000.0:.4f} ms per tick-bin)" + echo fmt"per-arm speed : {elapsed / max(1.0, float(tickBins)) * 1000.0:.4f} s per 1000 tick-bins per arm" + echo "" + echo "NOTE: offline OPEN-LOOP prediction quality only. No win/damage/survival claim." + echo "" + + # validation block + echo "=".repeat(120) + echo "VALIDATION -- the ruler must pass ALL of these before any number below is trusted" + echo "=".repeat(120) + if doShots and shotsCont.hits > 0 and shotsQuant.hits > 0: + echo fmt"1. recorded shots (OUR actual server-fired bearings vs the SAME interception solve):" + for mode in [("continuous", shotsCont), ("integer", shotsQuant)]: + let s = mode[1] + let sepOk = if separationPx(s) > 2.0: "OK" else: "WEAK" + echo fmt" ruler={mode[0]:<11} hits n={s.hits:<6} mean|err|={meanHitDeg(s):>7.3f} deg / {meanHitPx(s):>7.1f} px | " & + fmt"misses n={s.misses:<6} mean|err|={meanMissDeg(s):>7.3f} deg / {meanMissPx(s):>7.1f} px | " & + fmt"separation {separationDeg(s):>6.2f}x deg / {separationPx(s):>6.2f}x px -> {sepOk}" + else: + echo "1. recorded-shot validation: SKIPPED" + let oMax = max([arms[A_ORACLE].bands[0].maxAbs, arms[A_ORACLE].bands[1].maxAbs, + arms[A_ORACLE].bands[2].maxAbs, arms[A_ORACLE].bands[3].maxAbs, + arms[A_ORACLE].bands[4].maxAbs]) + let oOk = if oMax < 1e-6: "OK" else: "BROKEN" + echo fmt"2. perfect-oracle gun max |err| over all tick-bins = {oMax:.6f} deg -> {oOk}" + # ordering check: the static LOS gun must be worse than every predictive gun. + proc overallMean(arms: seq[ArmAcc], ai: int): float = + var sAbs = 0.0 + var n = 0 + for b in 0 ..< NBands: + sAbs += arms[ai].bands[b].sumAbs + n += arms[ai].bands[b].n + if n > 0: sAbs / float(n) else: 0.0 + let mHead = overallMean(arms, A_HEADON) + let mPat = overallMean(arms, A_PATTERN) + let mTmh = overallMean(arms, A_TMH) + let mBb = overallMean(arms, A_BB) + let mLin = overallMean(arms, A_NAIVE) + let ordOk = mHead > mPat and mHead > mTmh and mHead > mBb + echo fmt"3. HeadOn (static LOS) mean|err| = {mHead:.3f} deg vs Pattern {mPat:.3f} / TMHorizon {mTmh:.3f} / BitBrain {mBb:.3f}" + let ordMsg = if ordOk: "OK (static gun worst among real guns)" else: "UNEXPECTED: a predictive gun is worse than static LOS" + echo fmt" -> {ordMsg}" + echo fmt" NaiveLinear mean|err| = {mLin:.3f} deg (over-leads; see the lead-gain sweep for why a larger" + echo fmt" lead *response* does not mean a smaller angular error)" + echo "4. determinism: run twice and diff stdout (see fixture; verified separately)." + echo "" + + echo "=".repeat(120) + echo "THE BAR -- per-band mean ABSOLUTE angular aim error (deg), RMSE, sign, hit-proxy" + echo "=".repeat(120) + echo "hitProxy = fraction of tick-bins with |err| <= atan(18/range) (the angular half-width of the target disc)." + echo "" + stdout.write formatArmTable(arms) + + echo "" + echo "=".repeat(120) + echo "HEADROOM -- the direct answer: how far each arm is from the oracle ceiling, per band" + echo "=" .repeat(120) + let hdr = "band Pattern n Pattern|err| Pattern hpx Oracle hpx headroom pp naive hpx TMHoriz hpx BitBrain hpx" + echo hdr + echo "-".repeat(hdr.len) + for b in 0 ..< NBands: + let pat = arms[A_PATTERN].bands[b] + let orc = arms[A_ORACLE].bands[b] + let hp = pat.hitProxy + let ohp = orc.hitProxy + echo fmt"{BandLabels[b]:<9} {pat.n:>8} {fmt3(meanAbs(pat)):>12} {fmt4(hp):>12} {fmt4(ohp):>12} {ohp - hp:>13.4f} {fmt4(arms[A_NAIVE].bands[b].hitProxy):>11} {fmt4(arms[A_TMH].bands[b].hitProxy):>12} {fmt4(arms[A_BB].bands[b].hitProxy):>13}" + echo "" + echo "hitProxy = fraction of tick-bins aimed within atan(18/range) of the true interception point." + echo "headroom pp = oracle hitProxy - Pattern hitProxy = the absolute hit-probability points available" + echo "to a perfect predictor (the campaign is playing for a slice of this)." + echo "" + let hdrq = "band OracleQuant hpx integer-solve coarseness" + echo hdrq + echo "-".repeat(hdrq.len) + for b in 0 ..< NBands: + let oq = arms[A_ORACLEQ].bands[b] + echo fmt"{BandLabels[b]:<9} {fmt4(oq.hitProxy):>15} {fmt3(meanAbs(oq)):>10} deg mean |err|" + echo "(OracleQuant aims at the analyze_lead_capture_by_range.py integer-tick intercept and is scored" + echo " on the active ruler. On the continuous ruler it measures how much of a gun's 'error' the coarse" + echo " solve itself would produce; on the integer ruler it is identically zero.)" + + echo "" + echo "=".repeat(120) + echo "LEAD-GAIN SWEEP ON PATTERN -- multiply Pattern's lead (deg over LOS) by a constant" + echo "=".repeat(120) + let hdr2 = "band gain=1.0 gain=1.5 gain=2.0 gain=3.0 best-gain" + echo hdr2 + echo "-".repeat(hdr2.len) + for b in 0 ..< NBands: + let g1 = meanAbs(arms[A_PATTERN].bands[b]) + let g15 = meanAbs(arms[A_G15].bands[b]) + let g20 = meanAbs(arms[A_G20].bands[b]) + let g30 = meanAbs(arms[A_G30].bands[b]) + var best = "1.0" + var bestV = g1 + if g15 < bestV: bestV = g15; best = "1.5" + if g20 < bestV: bestV = g20; best = "2.0" + if g30 < bestV: bestV = g30; best = "3.0" + echo fmt"{BandLabels[b]:<9} {fmt3(g1):>10} {fmt3(g15):>10} {fmt3(g20):>10} {fmt3(g30):>10} {best} ({fmt3(bestV)})" + + echo "" + echo "=".repeat(120) + echo "LEAD INFORMATIVENESS -- capture slope (regression of applied lead on required lead) and lead correlation" + echo "=".repeat(120) + echo "capture slope is job-95's metric (1.0 = perfect proportional response). corr is the Pearson" + echo "correlation of the arm's lead with the REQUIRED lead: a large slope on an uncorrelated lead is" + echo "just amplified noise. This is the table that resolves the 'naive-linear captures 2x the lead but" + echo "hits less' tension." + echo "" + let hdr3 = "band HO |err| HO |req| Pat|req| Pat cap Pat corr Lin cap Lin corr TMH cap TMH corr BB cap BB corr" + echo hdr3 + echo "-".repeat(hdr3.len) + for b in 0 ..< NBands: + let sp = arms[A_PATTERN].bands[b] + let sn = arms[A_NAIVE].bands[b] + let st = arms[A_TMH].bands[b] + let sb = arms[A_BB].bands[b] + echo fmt"{BandLabels[b]:<9} {fmt3(meanAbs(arms[A_HEADON].bands[b])):>9} {fmt3(meanAbsReq(arms[A_HEADON].bands[b])):>9} {fmt3(meanAbsReq(sp)):>9} {fmt3(captureSlope(sp)):>10} {fmt3(leadCorr(sp)):>10} " & + fmt"{fmt3(captureSlope(sn)):>10} {fmt3(leadCorr(sn)):>10} {fmt3(captureSlope(st)):>10} " & + fmt"{fmt3(leadCorr(st)):>10} {fmt3(captureSlope(sb)):>10} {fmt3(leadCorr(sb)):>10}" + +when isMainModule: + main() diff --git a/docs/bitbrain_campaign.md b/docs/bitbrain_campaign.md new file mode 100644 index 0000000..876cb25 --- /dev/null +++ b/docs/bitbrain_campaign.md @@ -0,0 +1,258 @@ +# BitBrain campaign ledger + +**Goal:** make ModularBot's gun **beat Pattern live** against the real DrussGT. +The user has granted full freedom over the gun ("change input, output, every +knob of it") and accepts it may fail — the deliverable is that the attempt is +visible and evidence-backed. + +**THE FINAL VERDICT IS ALWAYS LIVE.** Everything in this file except the +`## Phase N` verdict lines is offline, open-loop, on a *fixed recorded enemy +trajectory*. Per `docs/offline_harness_trust.md` (commit `e40c849`) the offline +harness is trustworthy for exactly one thing: **per-gun single-tick prediction +quality on a fixed enemy trajectory** — and it is *never* trustworthy for +closed-loop questions (movement, range, round length, adaptation, gun +selection, damage, wins, survival). No offline number here is a win/damage +claim, and no phase may be called a success without a live A/B +(`tools/ab/ab_run.sh`, server-side event hit rate, left-running). + +Every claim below is tagged **[MEASURED]** (a command in §0 reproduces it) or +**[INFERRED]** (reasoning from measured facts). + +--- + +## Phase 0 — BUILD THE RULER AND ESTABLISH THE BAR *(owner: overnight job, committed)* + +### 0.1 The ruler + +`common_libs/gun_harness/prediction_quality.nim` + `common_libs/tests/run_prediction_quality.nim`. + +At each recorded tick the shooter sits at `O = (selfX, selfY)`. For a bullet of +speed `v` the **true interception point** is the first fractional time `t > 0` +at which the enemy's ACTUAL recorded track reaches distance `v*t` from `O` +(linear interpolation between recorded ticks). A bullet fired along the bearing +to `E(t)` coincides with the enemy at `t`. Angular error is +`wrap180(bearing(O→pred) − bearing(O→E(t)))` in **degrees**; every tick is +scored for the four power bins (speeds 17/15.5/14/11), all bands share that +horizon set. Per range band we report `mean|err|`, RMSE, mean signed err and the +hit-probability proxy `mean(|err| ≤ atan(18/range))`. + +The integer-tick solve from `analyze_lead_capture_by_range.py` (commit +`f91e121`) is kept as `interceptBearingQuant` and reported as `OracleQuant`; the +ruler ships the **continuous** solve because it separates recorded hits from +misses slightly better and removes the coarse solve's own overshoot +(§0.3.6). `--ruler quant` selects the integer solve. + +**Data:** the recorded live-vs-real-DrussGT corpus `/tmp/tfil_ab2/out` +(70 battles / 490 rounds / 899 607 ticks + `.events.jsonl` + `.rounds.json`), +**verified present before use**. It lives in `/tmp` and is therefore ephemeral; +if a later job finds it gone, regenerate it with the A/B harness +(`tools/ab/ab_run.sh`, which sets `TR_RECORD_WORLDSTATE` so ModularBot appends +per-tick world state) and point `--corpus` at the new output root. Layout: +`//runN.jsonl` + `runN.events.jsonl` + `runN.jsonl.rounds.json`. +149 MB of JSONL is converted once per run into a compact float32 `.qcache` +(keyed on source mtime+size) and ALL measurement is taken from the cache, so two +runs are byte-identical. See §0.4 for speed. + +### 0.2 Validation — the ruler must pass ALL of these **[MEASURED]** + +Run: `nim c -d:release --nimcache:/tmp/nc_j98 -r common_libs/tests/run_prediction_quality.nim` + +**1. Recorded HITS separate from recorded MISSES** (our ACTUAL server-fired +bearings, scored against the SAME interception solve): + +| ruler | hits n | hits mean\|err\| | misses n | misses mean\|err\| | separation | +|---|---|---|---|---|---| +| continuous | 5480 | **1.360° / 10.5 px** | 48304 | **16.597° / 140.3 px** | **12.20× deg / 13.34× px** | +| integer | 5480 | 1.478° / 11.4 px | 48304 | 16.724° / 141.3 px | 11.32× / 12.43× | + +(The earlier validated run quoted 11.59× / 11.6 px on hits; reproduced and +improved.) The continuous ruler is shipped because it separates better. + +**2. Perfect oracle scores 0.** Max `|err|` over all 3 598 428 tick-bins = +**0.000000°**. OK. + +**3. A static line-of-sight gun is far from the predictor on learnable motion.** +On a synthetic constant-velocity and a seeded random-walk trajectory the +ordering is exactly as physics demands: HeadOn (zero lead) is the worst, Pattern +and naive-linear are near-zero, and the lead-gain arms overshoot monotonically. +On the real DrussGT corpus the static gun is *not* worst — see §0.3.4, this is a +genuine property of the corpus, not a harness defect. + +**4. Determinism.** Two full 70-run sweeps, stdout diffed with the two wall-time +lines excluded: **byte-identical**. (The only difference between the two raw +outputs is `wall time 406.88s` vs `402.41s` and the derived ms-per-tick-bin.) +**[MEASURED]** + +**5. A real bug was found and fixed by this validation.** The ruler's +`wrap180` used Nim's float `mod`, which keeps the dividend's sign (C `fmod`), so +`(x+180) mod 360 − 180` returned `x−360` instead of the wrapped equivalent for +`x < −180`. This inflated the negative tail of every error (maxAbs read ~360° +instead of ~180°) and made HeadOn's mean error disagree with `mean|required|`. +After the fix HeadOn's `mean|err|` equals `mean|required lead|` to the last +digit at every band (see the `HO |err| / HO |req| / Pat|req|` columns in the +fixture). **A wrong ruler is worse than no ruler; this was the most important +10 minutes of the phase.** + +### 0.3 THE BAR — per-band numbers (70 runs, 3 598 428 tick-bins) **[MEASURED]** + +Format: `mean|err| deg` and, in brackets, `hitProxy`. `hitProxy` is the fraction +of tick-bins aimed within `atan(18/range)` of the true interception point. + +| band | Pattern | naive-linear | TMHorizon | BitBrain | HeadOn (static) | Oracle | +|---|---|---|---|---|---|---| +| 0–100 | **10.56** [0.699] | 17.92 [0.651] | 10.68 [0.696] | 11.12 [0.681] | 19.62 [0.342] | 0.00 [1.000] | +| 100–200 | **14.75** [0.342] | 14.63 [0.388] | 14.84 [0.342] | 15.11 [0.324] | 19.98 [0.172] | 0.00 [1.000] | +| 200–300 | **16.61** [0.185] | 17.57 [0.192] | 16.64 [0.179] | 16.84 [0.175] | 17.34 [0.133] | 0.00 [1.000] | +| 300–450 | 17.53 [0.104] | 20.98 [0.100] | 17.57 [0.100] | 17.58 [0.103] | **14.61** [0.105] | 0.00 [1.000] | +| 450+ | 16.19 [0.077] | 22.86 [0.054] | 16.20 [0.076] | 16.20 [0.077] | **12.33** [0.098] | 0.00 [1.000] | + +`n`: 4 423 / 24 908 / 74 215 / 1 119 777 / 2 311 323. Skipped (no valid +interception): 63 782 tick-bins (≈1.7 %). + +**0.3.1 DIRECT ANSWER — the gap between Pattern and the oracle ceiling:** + +| band | Pattern hitProxy | Oracle hitProxy | headroom (pp) | +|---|---|---|---| +| 0–100 | 0.6993 | 1.0000 | **+30.07** | +| 100–200 | 0.3418 | 1.0000 | **+65.82** | +| 200–300 | 0.1850 | 1.0000 | **+81.50** | +| 300–450 | 0.1036 | 1.0000 | **+89.64** | +| 450+ | 0.0767 | 1.0000 | **+92.33** | + +**[INFERRED, important]** The oracle is *non-causal*: it aims with perfect +knowledge of the enemy's future, so its 100 % is a definition, not an +achievement, and the 92 pp at 450+ is an UPPER bound that contains both "a +better predictor could get this" and "this is physically unknowable". The +**realistic** causal bound measured today is the best arm at 450+: **HeadOn at +9.8 %**, barely above Pattern's 7.7 %. So the campaign is playing for a few +percentage points at long range, not for 92 pp. The honest target statement is +"raise the 450+ proxy from 7.7 % toward the ~10 % causal band", not "toward +100 %". + +**0.3.2 The lead-gain sweep on Pattern is a DEAD END [MEASURED].** Multiply +Pattern's angular lead over LOS by a constant, per band: + +| band | gain 1.0 | gain 1.5 | gain 2.0 | gain 3.0 | +|---|---|---|---|---| +| 0–100 | **10.56** | 13.76 | 20.11 | 34.77 | +| 100–200 | **14.75** | 19.64 | 26.78 | 42.97 | +| 200–300 | **16.61** | 22.26 | 29.47 | 45.23 | +| 300–450 | **17.53** | 23.28 | 29.96 | 44.27 | +| 450+ | **16.19** | 21.25 | 27.00 | 39.27 | + +Gain 1.0 wins at EVERY band. Scaling Pattern's lead up makes it strictly worse. +This is the single most important negative result of Phase 0 and it should stop +any later job from "just adding more lead". + +**0.3.3 The naive-linear / capture tension, resolved [MEASURED].** Capture slope += regression of the arm's own lead on the required lead (job-95's statistic); +corr = Pearson correlation of the arm's lead with the required lead. **corr is +the informative number; a large slope on an uncorrelated lead is just amplified +noise.** + +| band | mean\|req\| | Pattern cap / corr | naive-linear cap / corr | TMHorizon | BitBrain | +|---|---|---|---|---|---| +| 100–200 | 19.98 | 0.553 / 0.612 | 0.592 / 0.523 | 0.560 / 0.614 | 0.570 / 0.611 | +| 300–450 | 14.61 | 0.278 / 0.266 | 0.476 / 0.324 | 0.273 / 0.262 | 0.280 / 0.266 | +| 450+ | 12.33 | 0.175 / 0.165 | 0.310 / 0.178 | 0.174 / 0.164 | 0.175 / 0.165 | + +Yes — on this corpus the naive-linear predictor applies **~1.8× more lead** than +Pattern at 450+ (0.310 vs 0.175; job-95 measured ~2×). Job-95's "we under-lead" +reading is confirmed. **But** the two arms carry almost the same lead +*information* (corr 0.178 vs 0.165), so the extra amplitude buys nothing and +costs angular accuracy: naive-linear's `mean|err|` is 22.86° vs Pattern's +16.19° at 450+. **Conclusion: the campaign's lever is lead INFORMATION +(correlation), not lead RESPONSE (capture slope).** Capturing more of an +uninformative lead is worse than capturing little of it — which is also exactly +why the gain sweep fails. + +**0.3.4 The surprise: at long range, static line-of-sight beats Pattern.** +HeadOn (aim at the enemy's current position) has `mean|err|` 14.61°/12.33° and +`hitProxy` 0.105/0.098 at 300–450/450+, both better than Pattern's +17.53°/16.19° and 0.104/0.077. **[INFERRED]** At 450+ the required lead +(`mean|req|` = 12.3°) is essentially unpredictable from the past (Pattern +corr 0.165), so Pattern's predicted lead is mostly variance added to a nearly +uninformative signal; a zero-lead aim has error = `|required lead|`, which is +smaller. Consistent with the live record: the live bot's own applied lead +capture was 0.135 at 450+ (job-95), i.e. the live bot was already nearly +zero-lead and hit 9.14 % there; Pattern's offline proxy is 7.7 %. + +**[INFERRED / CAVEAT]** The corpus is open-loop: DrussGT's recorded dodge was a +reaction to the LIVE bot's (near-zero-lead) bullets. Replaying Pattern on that +trajectory cannot show what DrussGT would do against Pattern's bullets. This +makes a **live A/B of HeadOn vs Pattern at long range the highest-value cheap +experiment in the campaign** (see §0.6). No offline claim that "HeadOn +beats Pattern" is permitted — only the live A/B decides. + +**0.3.5 BitBrain, as shipped, is Pattern [MEASURED].** BitBrain's base is +Pattern and its ADE/SBC corrector changes almost nothing: 450+ `mean|err|` +16.200° vs Pattern 16.193°, `hitProxy` 0.0767 vs 0.0767. TMHorizon likewise +(16.199° / 0.0757). The corrector is currently **adding no measurable aim +information** on this corpus. That is the thing Phase 1 must change. + +**0.3.6 Ruler resolution is NOT the limiter [MEASURED].** Aiming at the +integer-tick solve instead of the exact intercept costs only 0.29–1.10° of mean +error (`OracleQuant` column). So the "maybe the oracle only reaches 35 % because +the solve is coarse" worry is dead: the coarse/fine difference is ≈0.3° at long +range, far below the target tolerance (1.93° at 450+). Whatever caps the score, +it is the enemy's unpredictability, not the ruler. + +### 0.4 Speed **[MEASURED]** + +Full 70-run / 899 607-tick / 3 598 428 tick-bin sweep, 10 arms: +**406.9 s wall**, i.e. `0.1131 ms per tick-bin` over 10 arms, +**≈ 0.045 s per gun per 1000 ticks** (1000 ticks × 4 power bins). +BitBrain is the dominant cost (its ADE pass runs on every `predict` call); +Pattern/TMHorizon cache their per-tick work. A single-arm Pattern-only sweep is +several times cheaper. The binary cache (§0.1) is what makes repeat sweeps +affordable: without it every run re-parses 149 MB of JSONL. + +### 0.5 How to reproduce **[MEASURED]** + +``` +nim c -d:release --nimcache:/tmp/nc_j98 -o:/tmp/bbq_run \ + common_libs/tests/run_prediction_quality.nim +/tmp/bbq_run --corpus /tmp/tfil_ab2/out # full bar, ~7 min +/tmp/bbq_run --corpus /tmp/tfil_ab2/out --limit 10 # fast subset +/tmp/bbq_run --corpus /tmp/tfil_ab2/out --ruler quant # integer-tick solve +``` +Verbatim full output: `common_libs/tests/prediction_quality_results.txt`. +Determinism: two consecutive full runs are byte-identical except the two +wall-time lines. + +**Clean-checkout proof [MEASURED]:** `git archive HEAD | tar -x -C /tmp/bbq_clean` +then, from `/tmp/bbq_clean`, +`nim c -d:release --nimcache:/tmp/nc_j98 -o:bbq_run common_libs/tests/run_prediction_quality.nim` +builds, and `./bbq_run --corpus /tmp/tfil_ab2/out --limit 3` runs and prints the +same tables (separation 13.68× px on the 3-run subset). The committed harness is +self-contained; only the corpus is external. + +### 0.6 Designs still to try (seed for later phases) + +Ordered by expected value per unit of effort. Phase 0 has already killed one. + +| # | design | why it is worth trying | status | +|---|---|---|---| +| D1 | **Live A/B: HeadOn at 450+ vs Pattern** (distance-gated switch, or HeadOn-only control) | Offline says a static gun beats Pattern at long range; the cheapest possible test of the campaign's central premise | **TODO (highest value, live)** | +| D2 | **Pattern variants that raise lead CORRELATION at long range**: longer keys / multi-length keys, per-distance learned pattern tables, different match weighting, k-NN over movement signatures | The lever is corr (0.165 at 450+), not gain; the ruler measures exactly this | TODO (offline-searchable) | +| D3 | **Supervise BitBrain with the ruler's own labels** — per-tick bearing error to the true intercept, trained on N−1 runs, evaluated on a held-out run | BitBrain's corrector currently adds nothing (0.3.5); the ruler gives it a real target. Must hold out runs or it is overfitting | TODO (offline-searchable) | +| D4 | **A causal "predictability" gate**: at each tick estimate whether the future is predictable (e.g. recent pattern-match score, reversal entropy) and fall back to HeadOn/low-variance aim when it is not | Directly attacks the 0.3.4 failure mode without needing a better long-range predictor | TODO | +| D5 | Power policy at long range (already partly done live): lower power = faster bullet = less lead error | Shortens the horizon the predictor must extrapolate; affects hit rate, live-only verdict | TODO (offline proxy only) | +| D6 | Lead-gain sweep 1.0/1.5/2.0/3.0 | **DEAD — measured.** Gain 1.0 wins at every band (§0.3.2) | **KILLED** | + +Every D-item must end in a live A/B before any phase verdict. + +### 0.7 What would make us quit + +> If (a) no causal design raises the 450+ `hitProxy` above the static-gun +> reference (~0.10) on held-out runs by a margin larger than the run-to-run +> spread, **and** (b) the live A/B of the best such design shows no hit-rate or +> damage gain over Pattern with the left-running liveness check satisfied, then +> the campaign stops and we ship the simpler gun. We do not keep tuning an +> offline proxy that has stopped predicting live outcomes. + +--- + +## Phase 1 — *(unclaimed; append below)* + +## Phase 2 — *(unclaimed; append below)*