bitbrain campaign phase 0: offline prediction-quality ruler and the bar

New harness (common_libs/gun_harness/prediction_quality.nim +
common_libs/tests/run_prediction_quality.nim): per-gun single-tick aim error in
degrees against the true continuous interception point on the recorded
live-vs-real-DrussGT corpus (/tmp/tfil_ab2/out, 70 runs, 899607 ticks), per
range band, with the hit-probability proxy mean(|err|<=atan(18/range)).
Validated: recorded hits separate from misses 13.34x px (reference 11.59x),
perfect-oracle max |err| = 0, correct ordering on synthetic ground truth, two
full runs byte-identical. Fixed a wrap180 bug (Nim float mod keeps the dividend
sign) that inflated the negative error tail.

Bar (mean|err| deg [hitProxy] at 450+): Pattern 16.19 [0.077], naive-linear
22.86 [0.054], TMHorizon 16.20 [0.076], BitBrain 16.20 [0.077], static HeadOn
12.33 [0.098], oracle 0 [1.0]. Lead-gain sweep on Pattern is a dead end (1.0
wins every band). Naive-linear applies ~1.8x Pattern's lead but carries no more
lead information (corr 0.178 vs 0.165) and is strictly worse. Ledger:
docs/bitbrain_campaign.md. All verdicts remain live-only.
This commit is contained in:
2026-09-24 23:38:45 +02:00
parent 32a5e72fac
commit a82c864c60
4 changed files with 1366 additions and 0 deletions
@@ -0,0 +1,605 @@
## Offline per-gun PREDICTION-QUALITY scorer — the campaign ruler.
##
## WHAT THIS MEASURES (and only this)
## ----------------------------------
## `docs/offline_harness_trust.md` established that the offline harness is
## trustworthy for ONE question: the per-gun, single-tick PREDICTION QUALITY of a
## gun on a FIXED enemy trajectory. It is never trustworthy for closed-loop
## questions (wins, damage, survival, movement, range, adaptation, selection).
## This module builds exactly that one trustworthy thing: given a recorded live
## battle, "if this gun had aimed at every tick it was asked to, how good was its
## aim?" — nothing more.
##
## THE METRIC
## ----------
## At each recorded tick `i` the shooter sits at `O = (selfX, selfY)`. For a
## bullet of speed `v` the AIM-INDEPENDENT interception point is the first future
## tick `t* = i + k` (k >= 1) with `|E(t*) - O| <= v*k` — where the enemy's
## ACTUAL recorded track crosses the bullet's path. This is the same solve as
## `common_libs/tests/analyze_lead_capture_by_range.py` (commit f91e121): it
## depends only on the recorded truth and the bullet speed, never on our aim, so
## it is a fixed target every gun is scored against.
##
## Angular error is `wrap180(bearing(O -> pred) - bearing(O -> E(t*)))` in
## DEGREES. Degrees are the physically meaningful unit (the arena spans 800 px
## but the tolerance shrinks with range), so the ruler never reports pixels except
## in the sanity-validation path. Per RANGE BAND (0-100/100-200/200-300/300-450/
## 450+) we report:
## mean |err| (deg), RMSE (deg), mean signed err (deg),
## the hit-probability proxy `mean(|err| < atan(18/range))`, and n.
##
## The "perfect oracle" arm aims at `E(t*)` itself, so it must score ~0 error —
## that is the plumbing check. The real correctness check is `validateShots`,
## which scores OUR ACTUAL recorded fired bearings (from the event sidecar)
## against the same interception solve: recorded HITS must cluster near zero and
## MISSES far away. If that separation collapses, the ruler is wrong.
##
## EVALUATION GRANULARITY
## ----------------------
## The offline replay fires no real bullets, so a gun is asked for a prediction
## per POWER BIN (`PowerBins = 1.0/1.5/2.0/3.0`) each tick; each bin's prediction
## is scored against the interception at that bin's bullet speed. The phrase "the
## power that was really fired" applies to the separate `validateShots` check,
## which uses the server-recorded fired power. Aggregating over the four bins is
## a common horizon set shared by every arm, so the comparison is fair.
##
## WHY NOT `replayFixture`
## -----------------------
## `common_libs/gun_harness/offline_range.nim` replays a fixture through a
## `VirtualTracker` so that feedback-adaptive guns (Tsetlin, KNN, DecayGF) learn
## from resolved virtual bullets. Every arm measured here — Pattern, TMHorizon,
## BitBrain, HeadOn, and the naive-linear control — has a no-op `onResult`: its
## prediction is a pure function of the observed `WorldState` stream, so the
## tracker changes nothing while costing an O(MaxBullets) resolution scan per
## tick (~8192 slots), which would dominate runtime. This module therefore keeps
## the recorded-state semantics (states replayed in order, one gun's history is
## the whole stream) but drives the guns directly, and it needs per-tick
## predictions — which `replayFixture`'s aggregate report does not expose.
## The corpus representation is the recorded-run loader below (round boundaries
## from `*.rounds.json`, event sidecar for validation, binary cache for speed),
## i.e. the same recorded-state contract with the battle metadata the analysis
## needs.
##
## COST / CACHING
## --------------
## The corpus is ~900k recorded ticks over 70 battles (149 MB of JSONL). Parsing
## that with `std/json` on every sweep would dominate runtime, so the loader
## converts each run once into a compact binary `.qcache` (float32, 10 fields per
## tick) keyed on the source mtime+size. All measurements are taken from the
## cache, never from a mixture of cache and fresh parse, so two runs are
## byte-identical.
import std/[math, os, strformat, strutils, json, tables, times, algorithm]
const
NFields* = 10
NBands* = 5
BandLo* = [0.0, 100.0, 200.0, 300.0, 450.0]
BandHi* = [100.0, 200.0, 300.0, 450.0, 1.0e18]
BandLabels* = ["0-100", "100-200", "200-300", "300-450", "450+"]
MaxFlight* = 220
## max ticks a bullet is followed when solving for the interception tick
## (matches analyze_lead_capture_by_range.py).
BbRadius* = 18.0 ## hit-detection radius in px (atan(18/range) tolerance)
const CacheMagic = 0x31434242'i32 # "BBQ1"
const CacheVersion = 2'i32
proc wb(f: File, p: pointer, n: int) {.inline.} =
if n > 0: discard f.writeBuffer(p, n)
proc rb(f: File, p: pointer, n: int) {.inline.} =
if n > 0: discard f.readBuffer(p, n)
type
Corpus* = ref object
path*: string
arenaW*, arenaH*: float
n*: int
st*: seq[float32] ## NFields floats per tick
tick*: seq[int32]
rStart*: seq[int32] ## per-round first tick value
rCount*: seq[int32]
contiguous*: bool
base*: int
byTick*: Table[int32, int]
BandStat* = object
n*: int
sumAbs*: float64
sumSq*: float64
sumSigned*: float64
hits*: int
maxAbs*: float64
sumPred*: float64 ## sum of the arm's own lead over LOS (deg)
sumReq*: float64 ## sum of the true required lead over LOS (deg)
sumAbsReq*: float64
sumPredReq*: float64
sumPred2*: float64
sumReq2*: float64
ArmAcc* = object
name*: string
bands*: array[NBands, BandStat]
skipped*: int ## ticks×bins with no valid interception (no evidence)
evaluated*: int ## ticks×bins actually scored
TrEvent* = object
round*, tick*, owner*, bullet*: int
kind*: string
power*, x*, y*, dir*: float
ShotStat* = object
hits*, misses*: int
hitSumDeg*, missSumDeg*: float64
hitSumPx*, missSumPx*: float64
template ex*(c: Corpus, i: int): float = c.st[i * NFields + 0].float
template ey*(c: Corpus, i: int): float = c.st[i * NFields + 1].float
template eh*(c: Corpus, i: int): float = c.st[i * NFields + 2].float
template es*(c: Corpus, i: int): float = c.st[i * NFields + 3].float
template ee*(c: Corpus, i: int): float = c.st[i * NFields + 4].float
template sx*(c: Corpus, i: int): float = c.st[i * NFields + 5].float
template sy*(c: Corpus, i: int): float = c.st[i * NFields + 6].float
template sh*(c: Corpus, i: int): float = c.st[i * NFields + 7].float
template ss*(c: Corpus, i: int): float = c.st[i * NFields + 8].float
template se*(c: Corpus, i: int): float = c.st[i * NFields + 9].float
# ── geometry ─────────────────────────────────────────────────────────────────
proc wrap180*(x: float): float {.inline.} =
## Signed angular difference in (-180, 180]. NOTE: Nim's float `mod` keeps the
## sign of the dividend (C fmod), so `(x + 180) mod 360 - 180` is WRONG for
## x < -180 — it returns x - 360 instead of the wrapped equivalent. Normalise
## explicitly (this bug inflated the negative tail of every error before fix).
result = x mod 360.0
if result > 180.0: result -= 360.0
elif result <= -180.0: result += 360.0
proc bearingDeg*(ox, oy, px, py: float): float {.inline.} =
radToDeg(arctan2(py - oy, px - ox))
proc bandOf*(r: float): int {.inline.} =
for b in 0..<NBands:
if r >= BandLo[b] and r < BandHi[b]: return b
NBands - 1
proc tolDeg*(r: float): float {.inline.} =
## Angular half-width of the target disc at range `r`: atan(18/range).
radToDeg(arctan2(BbRadius, max(r, 1e-9)))
proc interceptTick(c: Corpus, i0, iEnd: int, ox, oy, speed: float): int =
## First integer k >= 1 with |E(i0+k) - O| <= speed*k and i0+k < iEnd; -1 if
## none. This is the integer-tick solve used by
## analyze_lead_capture_by_range.py.
let v = speed
var k = 1
while k <= MaxFlight:
let j = i0 + k
if j >= iEnd: break
let dx = c.ex(j) - ox
let dy = c.ey(j) - oy
if sqrt(dx * dx + dy * dy) <= v * float(k):
return k
inc k
-1
proc enemyPosAt(c: Corpus, i0, iEnd: int, t: float): tuple[x, y: float] =
## Linearly interpolated enemy position at fractional time `t` after i0.
let base = float(int(t))
let frac = t - base
let jm = min(i0 + int(base), iEnd - 1)
let jm2 = min(jm + 1, iEnd - 1)
(c.ex(jm) + (c.ex(jm2) - c.ex(jm)) * frac,
c.ey(jm) + (c.ey(jm2) - c.ey(jm)) * frac)
proc interceptBearingQuant*(c: Corpus, i0, iEnd: int, ox, oy,
speed: float): tuple[ok: bool, bearing, range: float] =
## The analyze_lead_capture_by_range.py solve: bearing to E(i0+k) at the first
## integer tick k the enemy is within the bullet's reach.
let rng = hypot(c.ex(i0) - ox, c.ey(i0) - oy)
let k = interceptTick(c, i0, iEnd, ox, oy, speed)
if k < 0: return (false, 0.0, rng)
(true, bearingDeg(ox, oy, c.ex(i0 + k), c.ey(i0 + k)), rng)
proc interceptBearingCont*(c: Corpus, i0, iEnd: int, ox, oy,
speed: float): tuple[ok: bool, bearing, range: float] =
## THE PHYSICALLY EXACT TRUE INTERCEPTION POINT: the first fractional time
## t > 0 at which the enemy's recorded track reaches distance speed*t from the
## origin, found by bisecting the first integer tick where it comes within
## reach. A bullet fired along the bearing to E(t) coincides with the enemy at
## t, so aiming at this point is a true hit (the integer-tick solve overshoots:
## it aims at E(k) but the bullet and enemy meet at E(t) with t <= k).
let rng = hypot(c.ex(i0) - ox, c.ey(i0) - oy)
let v = speed
var k = 1
while k <= MaxFlight:
let j = i0 + k
if j >= iEnd: break
let f = hypot(c.ex(j) - ox, c.ey(j) - oy) - v * float(k)
if f <= 0.0:
var lo = float(k - 1)
var hi = float(k)
for _ in 0 ..< 24:
let mid = 0.5 * (lo + hi)
let p = enemyPosAt(c, i0, iEnd, mid)
if hypot(p.x - ox, p.y - oy) - v * mid > 0.0: lo = mid
else: hi = mid
let ts = 0.5 * (lo + hi)
let p = enemyPosAt(c, i0, iEnd, ts)
return (true, bearingDeg(ox, oy, p.x, p.y), rng)
inc k
(false, 0.0, rng)
proc interceptBearing*(c: Corpus, i0, iEnd: int, ox, oy, speed: float,
cont: bool): tuple[ok: bool, bearing, range: float] =
if cont: interceptBearingCont(c, i0, iEnd, ox, oy, speed)
else: interceptBearingQuant(c, i0, iEnd, ox, oy, speed)
# ── accumulators ─────────────────────────────────────────────────────────────
proc addErr*(s: var BandStat, errDeg, range: float) =
inc s.n
let a = abs(errDeg)
s.sumAbs += a
s.sumSq += float64(errDeg) * float64(errDeg)
s.sumSigned += errDeg
if a > s.maxAbs: s.maxAbs = a
if a <= tolDeg(range): inc s.hits
proc record*(a: var ArmAcc, range: float, errDeg, predLead, reqLead: float) =
let b = bandOf(range)
addErr(a.bands[b], errDeg, range)
var s = addr a.bands[b]
s.sumPred += predLead
s.sumReq += reqLead
s.sumAbsReq += abs(reqLead)
s.sumPredReq += predLead * reqLead
s.sumPred2 += predLead * predLead
s.sumReq2 += reqLead * reqLead
inc a.evaluated
proc captureSlope*(s: BandStat): float =
## Regression of the arm's lead on the true required lead (job-95's "capture"
## statistic): 1.0 = perfect proportional response, 0.0 = no response.
if s.sumReq2 <= 1e-12: NaN else: s.sumPredReq / s.sumReq2
proc leadCorr*(s: BandStat): float =
## Pearson correlation between the arm's lead and the required lead. THIS is
## the honest measure of "is the lead informative"; a large capture slope on
## an uncorrelated lead is just amplification of noise.
let d = s.sumPred2 * s.sumReq2
if d <= 1e-12: NaN else: s.sumPredReq / sqrt(d)
proc meanAbsReq*(s: BandStat): float =
if s.n == 0: NaN else: s.sumAbsReq / float(s.n)
proc skip*(a: var ArmAcc) = inc a.skipped
proc meanAbs*(s: BandStat): float =
if s.n == 0: NaN else: s.sumAbs / float(s.n)
proc rmse*(s: BandStat): float =
if s.n == 0: NaN else: sqrt(s.sumSq / float(s.n))
proc meanSigned*(s: BandStat): float =
if s.n == 0: NaN else: s.sumSigned / float(s.n)
proc hitProxy*(s: BandStat): float =
if s.n == 0: NaN else: s.hits.float / float(s.n)
# ── corpus loading + binary cache ────────────────────────────────────────────
proc cachePathFor*(src: string): string = src & ".qcache"
proc eventsPathFor*(runPath: string): string =
## `run10.jsonl` -> `run10.events.jsonl` (note: NOT run10.jsonl.events.jsonl).
if runPath.endsWith(".jsonl"):
runPath[0 ..< runPath.len - 6] & ".events.jsonl"
else:
runPath & ".events.jsonl"
proc mtimeOf(p: string): float =
if fileExists(p): getFileInfo(p).lastWriteTime.toUnixFloat else: 0.0
proc buildCache(src, roundsPath, cachePath: string) =
## JSONL -> compact binary. Only runs when the cache is absent/stale.
var st: seq[float32]
var ticks: seq[int32]
var arenaW = 800.0
var arenaH = 600.0
for line in lines(src):
let s = line.strip()
if s.len == 0: continue
let node = parseJson(s)
if node.hasKey("meta"):
if node["meta"].hasKey("arena"):
let a = node["meta"]["arena"]
if a.hasKey("w"): arenaW = a["w"].getFloat()
if a.hasKey("h"): arenaH = a["h"].getFloat()
continue
if node.hasKey("end"): continue
st.add node["ex"].getFloat().float32
st.add node["ey"].getFloat().float32
st.add node["eh"].getFloat().float32
st.add node["es"].getFloat().float32
st.add node["ee"].getFloat().float32
st.add node["sx"].getFloat().float32
st.add node["sy"].getFloat().float32
st.add node["sh"].getFloat().float32
st.add node["ss"].getFloat().float32
st.add node["se"].getFloat().float32
ticks.add node["tick"].getInt().int32
var rStart, rCount: seq[int32]
if roundsPath.len > 0 and fileExists(roundsPath):
try:
let rj = parseFile(roundsPath)
for r in rj["rounds"]:
rStart.add r["startTick"].getInt().int32
rCount.add r["count"].getInt().int32
except CatchableError:
discard
if rStart.len == 0:
rStart.add 0'i32
rCount.add ticks.len.int32
var n32 = ticks.len.int32
var nr32 = rStart.len.int32
var magic = CacheMagic
var version = CacheVersion
var sm = mtimeOf(src)
var ss = getFileInfo(src).size.int64
var rm = mtimeOf(roundsPath)
let f = open(cachePath, fmWrite)
defer: f.close()
wb(f, addr magic, sizeof(magic))
wb(f, addr version, sizeof(version))
wb(f, addr n32, sizeof(n32))
wb(f, addr nr32, sizeof(nr32))
wb(f, addr arenaW, sizeof(arenaW))
wb(f, addr arenaH, sizeof(arenaH))
wb(f, addr sm, sizeof(sm))
wb(f, addr ss, sizeof(ss))
wb(f, addr rm, sizeof(rm))
if rStart.len > 0:
wb(f, addr rStart[0], rStart.len * sizeof(int32))
wb(f, addr rCount[0], rCount.len * sizeof(int32))
if ticks.len > 0:
wb(f, addr ticks[0], ticks.len * sizeof(int32))
wb(f, addr st[0], st.len * sizeof(float32))
proc readCache(path: string): Corpus =
let f = open(path, fmRead)
defer: f.close()
var magic, version: int32
rb(f, addr magic, sizeof(magic))
rb(f, addr version, sizeof(version))
if magic != CacheMagic or version != CacheVersion:
raise newException(IOError, "bad qcache header: " & path)
var n32, nr32: int32
rb(f, addr n32, sizeof(n32))
rb(f, addr nr32, sizeof(nr32))
result = Corpus(path: path, n: int(n32))
rb(f, addr result.arenaW, sizeof(result.arenaW))
rb(f, addr result.arenaH, sizeof(result.arenaH))
var sm: float
var ss: int64
var rm: float
rb(f, addr sm, sizeof(sm))
rb(f, addr ss, sizeof(ss))
rb(f, addr rm, sizeof(rm))
result.rStart.setLen(int(nr32))
result.rCount.setLen(int(nr32))
if nr32 > 0:
rb(f, addr result.rStart[0], int(nr32) * sizeof(int32))
rb(f, addr result.rCount[0], int(nr32) * sizeof(int32))
result.tick.setLen(result.n)
result.st.setLen(result.n * NFields)
if result.n > 0:
rb(f, addr result.tick[0], result.n * sizeof(int32))
rb(f, addr result.st[0], result.n * NFields * sizeof(float32))
proc indexMapping(c: var Corpus) =
c.contiguous = c.n > 0
c.base = if c.n > 0: int(c.tick[0]) else: 0
if c.contiguous:
for i in 0..<c.n:
if int(c.tick[i]) != c.base + i:
c.contiguous = false
break
if not c.contiguous:
c.byTick = initTable[int32, int](nextPowerOfTwo(max(16, c.n)))
for i in 0..<c.n: c.byTick[c.tick[i]] = i
proc idxOfTick*(c: Corpus, t: int32, found: var bool): int =
if c.contiguous:
let i = int(t) - c.base
if i >= 0 and i < c.n and c.tick[i] == t:
found = true
return i
found = false
return -1
if c.byTick.hasKey(t):
found = true
return c.byTick[t]
found = false
-1
proc cacheFresh(cache, src, roundsPath: string): bool =
if not fileExists(cache): return false
try:
let f = open(cache, fmRead)
var magic, version: int32
rb(f, addr magic, sizeof(magic))
rb(f, addr version, sizeof(version))
var n32, nr32: int32
rb(f, addr n32, sizeof(n32))
rb(f, addr nr32, sizeof(nr32))
var aw, ah, sm, rm: float
var ss: int64
rb(f, addr aw, sizeof(aw))
rb(f, addr ah, sizeof(ah))
rb(f, addr sm, sizeof(sm))
rb(f, addr ss, sizeof(ss))
rb(f, addr rm, sizeof(rm))
f.close()
return magic == CacheMagic and version == CacheVersion and
sm == mtimeOf(src) and ss == getFileInfo(src).size.int64 and
rm == mtimeOf(roundsPath)
except CatchableError:
false
proc loadCorpus*(src: string): Corpus =
## Load a recorded run (`runN.jsonl`); the per-round index is read from the
## sibling `runN.jsonl.rounds.json`. Builds the binary cache when stale/absent,
## then ALWAYS measures from the cache so repeated runs are byte-identical.
let cache = cachePathFor(src)
let roundsPath = src & ".rounds.json"
if not cacheFresh(cache, src, roundsPath):
buildCache(src, roundsPath, cache)
result = readCache(cache)
result.path = src
result.indexMapping()
proc discoverRuns*(root: string): seq[string] =
## All `<arm>/runN.jsonl` under `root` that have the events + rounds sidecars.
for sub in walkDir(root, relative = false):
if sub.kind != pcDir: continue
for fn in walkFiles(sub.path / "*.jsonl"):
if fn.endsWith(".events.jsonl"): continue
if fileExists(fn & ".rounds.json") and fileExists(eventsPathFor(fn)):
result.add fn
result.sort()
# ── shot-level geometry validation ───────────────────────────────────────────
#
# Scores OUR ACTUAL server-recorded fired bearings against the interception solve
# above. Owner ids in the event sidecar are not stable across runs, so each run is
# attributed independently (mirrors analyze_lead_capture_by_range.py): a fire
# event's (x, y) is the firing tank's centre and its energy drops by exactly the
# fired power one tick later.
proc parseEvents*(path: string): seq[TrEvent] =
for line in lines(path):
let s = line.strip()
if s.len == 0: continue
let node = parseJson(s)
result.add TrEvent(
round: node["round"].getInt(),
tick: node["tick"].getInt(),
kind: node["type"].getStr(),
owner: (if node.hasKey("owner"): node["owner"].getInt() else: -1),
bullet: (if node.hasKey("bullet"): node["bullet"].getInt() else: -1),
power: (if node.hasKey("power"): node["power"].getFloat() else: 0.0),
x: (if node.hasKey("x"): node["x"].getFloat() else: 0.0),
y: (if node.hasKey("y"): node["y"].getFloat() else: 0.0),
dir: (if node.hasKey("dir"): node["dir"].getFloat() else: 0.0))
proc matchEvent(c: Corpus, t: int, side: int, ev: TrEvent): bool =
## side 0 = 'e' (enemy), 1 = 's' (self). Position match + energy drop.
if t < 0 or t + 1 >= c.n: return false
let px = if side == 0: c.ex(t) else: c.sx(t)
let py = if side == 0: c.ey(t) else: c.sy(t)
if abs(px - ev.x) > 0.02 or abs(py - ev.y) > 0.02: return false
let e0 = if side == 0: c.ee(t) else: c.se(t)
let e1 = if side == 0: c.ee(t + 1) else: c.se(t + 1)
abs((e0 - e1) - ev.power) < 0.02
proc roundStartIndexOf(c: Corpus, rnd: int): int =
## Rounds hold a GLOBAL startTick; map to an array index.
var startTick: int32 = 0
for i, r in c.rStart:
if int(i) + 1 == rnd: startTick = r
if c.contiguous: return int(startTick) - c.base
var found = false
idxOfTick(c, startTick, found)
proc validateShots*(c: Corpus, events: seq[TrEvent], cont: bool): ShotStat =
## Hits must show small angular error and misses large; the separation ratio is
## the ruler's correctness certificate.
var votes = initTable[int, array[2, int]]()
for ev in events:
if ev.kind != "fire": continue
let guess = c.roundStartIndexOf(ev.round) + ev.tick
if ev.owner notin votes: votes[ev.owner] = [0, 0]
for t in (guess - 8) .. (guess + 8):
for side in 0..1:
if matchEvent(c, t, side, ev): inc votes[ev.owner][side]
var ownerSide = initTable[int, int]()
for owner, v in votes:
ownerSide[owner] = if v[1] >= v[0]: 1 else: 0
var resolution = initTable[string, string]()
for ev in events:
if ev.kind in ["hit", "hitwall", "hitbullet"]:
resolution[$ev.round & "/" & $ev.owner & "/" & $ev.bullet] = ev.kind
for ev in events:
if ev.kind != "fire": continue
if ownerSide.getOrDefault(ev.owner, -1) != 1: continue # only OUR shots
let guess = c.roundStartIndexOf(ev.round) + ev.tick
var t0 = -1
var bestD = high(int)
for t in (guess - 8) .. (guess + 8):
if matchEvent(c, t, ownerSide[ev.owner], ev):
let d = abs(t - guess)
if d < bestD:
bestD = d
t0 = t
if t0 < 0: continue
var rIdx = -1
for i, r in c.rStart:
if int(i) + 1 == ev.round: rIdx = i
if rIdx < 0: continue
let iEnd = int(c.rStart[rIdx]) - c.base + int(c.rCount[rIdx])
let speed = 20.0 - 3.0 * ev.power
let ib = interceptBearing(c, t0, iEnd, ev.x, ev.y, speed, cont)
if not ib.ok: continue
let err = wrap180(ev.dir - ib.bearing)
let kind = resolution.getOrDefault($ev.round & "/" & $ev.owner & "/" & $ev.bullet, "")
if kind == "hit":
inc result.hits
result.hitSumDeg += abs(err)
result.hitSumPx += abs(degToRad(err)) * ib.range
elif kind in ["hitwall", "hitbullet"]:
inc result.misses
result.missSumDeg += abs(err)
result.missSumPx += abs(degToRad(err)) * ib.range
proc separationDeg*(s: ShotStat): float =
let hd = if s.hits > 0: s.hitSumDeg / float(s.hits) else: NaN
let md = if s.misses > 0: s.missSumDeg / float(s.misses) else: NaN
md / hd
proc separationPx*(s: ShotStat): float =
let hp = if s.hits > 0: s.hitSumPx / float(s.hits) else: NaN
let mp = if s.misses > 0: s.missSumPx / float(s.misses) else: NaN
mp / hp
proc meanHitDeg*(s: ShotStat): float =
if s.hits > 0: s.hitSumDeg / float(s.hits) else: NaN
proc meanMissDeg*(s: ShotStat): float =
if s.misses > 0: s.missSumDeg / float(s.misses) else: NaN
proc meanHitPx*(s: ShotStat): float =
if s.hits > 0: s.hitSumPx / float(s.hits) else: NaN
proc meanMissPx*(s: ShotStat): float =
if s.misses > 0: s.missSumPx / float(s.misses) else: NaN
# ── formatting ───────────────────────────────────────────────────────────────
proc formatArmTable*(arms: seq[ArmAcc]): string =
let hdr = "arm band n meanAbs rmse signed hitProxy maxAbs"
result = hdr & "\n" & "-".repeat(hdr.len) & "\n"
for a in arms:
for b in 0..<NBands:
let s = a.bands[b]
if s.n == 0:
result.add fmt"{a.name:<16} {BandLabels[b]:<9} {0:>8} {'-':>8} {'-':>8} {'-':>8} {'-':>9} {'-':>8}" & "\n"
else:
result.add fmt"{a.name:<16} {BandLabels[b]:<9} {s.n:>8} {meanAbs(s):>8.3f} {rmse(s):>8.3f} {meanSigned(s):>8.3f} {hitProxy(s):>9.4f} {s.maxAbs:>8.2f}" & "\n"
if a.skipped > 0:
result.add fmt" ({a.name}: {a.skipped} tick-bins had no valid interception)" & "\n"
@@ -0,0 +1,146 @@
========================================================================================================================
OFFLINE PREDICTION QUALITY -- per-gun single-tick aim error vs the true interception point
========================================================================================================================
corpus : /tmp/tfil_ab2/out
ruler : continuous (physically exact)
runs : 70
recorded ticks: 899607
tick x bin : 3598428
wall time : 406.88s (0.1131 ms per tick-bin)
per-arm speed : 0.1131 s per 1000 tick-bins per arm
NOTE: offline OPEN-LOOP prediction quality only. No win/damage/survival claim.
========================================================================================================================
VALIDATION -- the ruler must pass ALL of these before any number below is trusted
========================================================================================================================
1. recorded shots (OUR actual server-fired bearings vs the SAME interception solve):
ruler=continuous hits n=5480 mean|err|= 1.360 deg / 10.5 px | misses n=48304 mean|err|= 16.597 deg / 140.3 px | separation 12.20x deg / 13.34x px -> OK
ruler=integer hits n=5480 mean|err|= 1.478 deg / 11.4 px | misses n=48304 mean|err|= 16.724 deg / 141.3 px | separation 11.32x deg / 12.43x px -> OK
2. perfect-oracle gun max |err| over all tick-bins = 0.000000 deg -> OK
3. HeadOn (static LOS) mean|err| = 13.217 deg vs Pattern 16.609 / TMHorizon 16.627 / BitBrain 16.635
-> UNEXPECTED: a predictive gun is worse than static LOS
NaiveLinear mean|err| = 22.086 deg (over-leads; see the lead-gain sweep for why a larger
lead *response* does not mean a smaller angular error)
4. determinism: run twice and diff stdout (see fixture; verified separately).
========================================================================================================================
THE BAR -- per-band mean ABSOLUTE angular aim error (deg), RMSE, sign, hit-proxy
========================================================================================================================
hitProxy = fraction of tick-bins with |err| <= atan(18/range) (the angular half-width of the target disc).
arm band n meanAbs rmse signed hitProxy maxAbs
----------------------------------------------------------------------------------
Oracle 0-100 4423 0.000 0.000 0.000 1.0000 0.00
Oracle 100-200 24908 0.000 0.000 0.000 1.0000 0.00
Oracle 200-300 74215 0.000 0.000 0.000 1.0000 0.00
Oracle 300-450 1119777 0.000 0.000 0.000 1.0000 0.00
Oracle 450+ 2311323 0.000 0.000 0.000 1.0000 0.00
(Oracle: 63782 tick-bins had no valid interception)
OracleQuant 0-100 4423 1.103 1.498 0.078 1.0000 6.00
OracleQuant 100-200 24908 0.674 0.911 -0.002 1.0000 4.07
OracleQuant 200-300 74215 0.472 0.624 -0.008 1.0000 2.44
OracleQuant 300-450 1119777 0.360 0.470 -0.010 1.0000 1.62
OracleQuant 450+ 2311323 0.288 0.376 0.010 1.0000 1.23
(OracleQuant: 63782 tick-bins had no valid interception)
HeadOn 0-100 4423 19.619 23.254 -1.279 0.3423 46.28
HeadOn 100-200 24908 19.982 23.151 -0.462 0.1724 46.62
HeadOn 200-300 74215 17.341 20.521 0.481 0.1330 46.33
HeadOn 300-450 1119777 14.607 17.606 0.690 0.1049 46.38
HeadOn 450+ 2311323 12.326 15.017 -0.263 0.0984 45.49
(HeadOn: 63782 tick-bins had no valid interception)
Pattern 0-100 4423 10.555 14.796 -0.404 0.6993 64.72
Pattern 100-200 24908 14.745 19.510 1.276 0.3418 76.63
Pattern 200-300 74215 16.610 21.174 1.350 0.1850 81.45
Pattern 300-450 1119777 17.531 21.838 0.948 0.1036 85.65
Pattern 450+ 2311323 16.193 20.021 -0.642 0.0767 79.92
(Pattern: 63782 tick-bins had no valid interception)
PatternGain1.5 0-100 4423 13.764 18.522 0.034 0.6093 79.58
PatternGain1.5 100-200 24908 19.640 25.124 2.144 0.2025 97.08
PatternGain1.5 200-300 74215 22.260 27.672 1.784 0.1024 103.15
PatternGain1.5 300-450 1119777 23.279 28.548 1.077 0.0627 111.53
PatternGain1.5 450+ 2311323 21.245 26.062 -0.832 0.0542 104.02
(PatternGain1.5: 63782 tick-bins had no valid interception)
PatternGain2.0 0-100 4423 20.111 25.637 0.471 0.4047 95.58
PatternGain2.0 100-200 24908 26.782 33.175 3.013 0.1487 118.82
PatternGain2.0 200-300 74215 29.473 35.856 2.219 0.0776 126.59
PatternGain2.0 300-450 1119777 29.955 36.371 1.207 0.0488 137.41
PatternGain2.0 450+ 2311323 26.997 32.935 -1.021 0.0425 128.12
(PatternGain2.0: 63782 tick-bins had no valid interception)
PatternGain3.0 0-100 4423 34.772 43.079 1.346 0.2720 141.72
PatternGain3.0 100-200 24908 42.972 51.921 4.751 0.0978 162.72
PatternGain3.0 200-300 74215 45.232 54.158 3.088 0.0507 173.45
PatternGain3.0 300-450 1119777 44.274 53.364 1.462 0.0333 179.99
PatternGain3.0 450+ 2311323 39.271 47.718 -1.400 0.0293 177.73
(PatternGain3.0: 63782 tick-bins had no valid interception)
NaiveLinear 0-100 4423 17.924 35.009 -0.098 0.6505 178.26
NaiveLinear 100-200 24908 14.629 24.037 -0.445 0.3879 179.65
NaiveLinear 200-300 74215 17.572 24.479 1.030 0.1924 179.78
NaiveLinear 300-450 1119777 20.976 26.164 1.096 0.0995 179.64
NaiveLinear 450+ 2311323 22.857 27.679 -0.477 0.0537 179.98
(NaiveLinear: 63782 tick-bins had no valid interception)
TMHorizon 0-100 4423 10.681 14.782 -1.233 0.6955 66.72
TMHorizon 100-200 24908 14.841 19.526 0.814 0.3418 78.63
TMHorizon 200-300 74215 16.644 21.142 1.091 0.1786 84.45
TMHorizon 300-450 1119777 17.572 21.878 0.691 0.0996 84.01
TMHorizon 450+ 2311323 16.199 20.025 -0.687 0.0757 81.19
(TMHorizon: 63782 tick-bins had no valid interception)
BitBrain 0-100 4423 11.118 15.268 -1.507 0.6810 64.72
BitBrain 100-200 24908 15.107 19.782 1.289 0.3241 76.63
BitBrain 200-300 74215 16.838 21.426 1.341 0.1747 105.04
BitBrain 300-450 1119777 17.575 21.894 0.933 0.1025 100.98
BitBrain 450+ 2311323 16.200 20.032 -0.647 0.0767 96.38
(BitBrain: 63782 tick-bins had no valid interception)
========================================================================================================================
HEADROOM -- the direct answer: how far each arm is from the oracle ceiling, per band
========================================================================================================================
band Pattern n Pattern|err| Pattern hpx Oracle hpx headroom pp naive hpx TMHoriz hpx BitBrain hpx
---------------------------------------------------------------------------------------------------------------
0-100 4423 10.555 0.6993 1.0000 0.3007 0.6505 0.6955 0.6810
100-200 24908 14.745 0.3418 1.0000 0.6582 0.3879 0.3418 0.3241
200-300 74215 16.610 0.1850 1.0000 0.8150 0.1924 0.1786 0.1747
300-450 1119777 17.531 0.1036 1.0000 0.8964 0.0995 0.0996 0.1025
450+ 2311323 16.193 0.0767 1.0000 0.9233 0.0537 0.0757 0.0767
hitProxy = fraction of tick-bins aimed within atan(18/range) of the true interception point.
headroom pp = oracle hitProxy - Pattern hitProxy = the absolute hit-probability points available
to a perfect predictor (the campaign is playing for a slice of this).
band OracleQuant hpx integer-solve coarseness
---------------------------------------------------
0-100 1.0000 1.103 deg mean |err|
100-200 1.0000 0.674 deg mean |err|
200-300 1.0000 0.472 deg mean |err|
300-450 1.0000 0.360 deg mean |err|
450+ 1.0000 0.288 deg mean |err|
(OracleQuant aims at the analyze_lead_capture_by_range.py integer-tick intercept and is scored
on the active ruler. On the continuous ruler it measures how much of a gun's 'error' the coarse
solve itself would produce; on the integer ruler it is identically zero.)
========================================================================================================================
LEAD-GAIN SWEEP ON PATTERN -- multiply Pattern's lead (deg over LOS) by a constant
========================================================================================================================
band gain=1.0 gain=1.5 gain=2.0 gain=3.0 best-gain
------------------------------------------------------------------
0-100 10.555 13.764 20.111 34.772 1.0 (10.555)
100-200 14.745 19.640 26.782 42.972 1.0 (14.745)
200-300 16.610 22.260 29.473 45.232 1.0 (16.610)
300-450 17.531 23.279 29.955 44.274 1.0 (17.531)
450+ 16.193 21.245 26.997 39.271 1.0 (16.193)
========================================================================================================================
LEAD INFORMATIVENESS -- capture slope (regression of applied lead on required lead) and lead correlation
========================================================================================================================
capture slope is job-95's metric (1.0 = perfect proportional response). corr is the Pearson
correlation of the arm's lead with the REQUIRED lead: a large slope on an uncorrelated lead is
just amplified noise. This is the table that resolves the 'naive-linear captures 2x the lead but
hits less' tension.
band HO |err| HO |req| Pat|req| Pat cap Pat corr Lin cap Lin corr TMH cap TMH corr BB cap BB corr
-----------------------------------------------------------------------------------------------------------------------
0-100 19.619 19.619 19.619 0.649 0.774 0.573 0.362 0.668 0.776 0.701 0.768
100-200 19.982 19.982 19.982 0.553 0.612 0.592 0.523 0.560 0.614 0.570 0.611
200-300 17.341 17.341 17.341 0.449 0.457 0.519 0.429 0.452 0.460 0.459 0.457
300-450 14.607 14.607 14.607 0.278 0.266 0.476 0.324 0.273 0.262 0.280 0.266
450+ 12.326 12.326 12.326 0.175 0.165 0.310 0.178 0.174 0.164 0.175 0.165
@@ -0,0 +1,357 @@
## Offline PREDICTION-QUALITY runner — the campaign's measurement sweep.
##
## Reads the recorded live-vs-real-DrussGT corpus, drives each arm over the
## recorded enemy trajectory, and scores every tick×power-bin prediction against
## the aim-independent interception point (see prediction_quality.nim). Owns the
## RANGE-BAND table that is "the bar" for the BitBrain campaign.
##
## NO CLOSED-LOOP CLAIM IS MADE HERE. Every arm below is an open-loop prediction
## scored on a FIXED trajectory. Wins, damage and survival are decided live.
##
## Usage:
## nim c -r --nimcache:/tmp/nc_j98 common_libs/tests/run_prediction_quality.nim \
## [--corpus /tmp/tfil_ab2/out] [--limit N] [--timing]
##
## `--limit N` keeps only the first N runs (sorted), for fast iteration.
import std/[os, strformat, strutils, times, math]
import gun_harness/[gun_interface, virtual_bullets, prediction_quality]
import guns/[head_on, pattern_matcher, tm_horizon, bitbrain_gun]
# arm indices (fixed order = fixed output)
const
A_ORACLE* = 0
A_ORACLEQ* = 1
A_HEADON* = 2
A_PATTERN* = 3
A_G15* = 4
A_G20* = 5
A_G30* = 6
A_NAIVE* = 7
A_TMH* = 8
A_BB* = 9
ArmNames* = ["Oracle", "OracleQuant", "HeadOn", "Pattern", "PatternGain1.5",
"PatternGain2.0", "PatternGain3.0", "NaiveLinear", "TMHorizon", "BitBrain"]
# ── the naive-linear control (job-95's LIN_M = 4 extrapolation) ──────────────
#
# Velocity = (pos(t) - pos(t-4)) / 4, then iterate the interception equation.
# This is the trivial predictive gun the lead-capture analysis used as its
# ceiling control; it captures ~2x the lead response Pattern does at 450+ and is
# included to resolve that tension against the angular-error ruler.
type
NaiveLinearGun = object
hist: array[5, tuple[x, y: float]]
count: int
lastTick: int
proc predict*(g: var NaiveLinearGun, state: WorldState,
bulletSpeed: float): GunPrediction =
if state.tick != g.lastTick:
for i in countdown(4, 1): g.hist[i] = g.hist[i - 1]
g.hist[0] = (state.enemyX, state.enemyY)
if g.count < 5: inc g.count
g.lastTick = state.tick
if g.count < 5 or bulletSpeed <= 0.0:
return GunPrediction(x: state.enemyX, y: state.enemyY)
let vx = (g.hist[0].x - g.hist[4].x) / 4.0
let vy = (g.hist[0].y - g.hist[4].y) / 4.0
let ox = state.selfX
let oy = state.selfY
var t = hypot(state.enemyX - ox, state.enemyY - oy) / bulletSpeed
for _ in 0..<40:
t = hypot(state.enemyX + vx * t - ox, state.enemyY + vy * t - oy) / bulletSpeed
GunPrediction(x: state.enemyX + vx * t, y: state.enemyY + vy * t)
proc onResult*(g: var NaiveLinearGun, e: FeedbackEvent) = discard
# ── runner ───────────────────────────────────────────────────────────────────
type Ctx = object
c: Corpus
cont: bool
pattern: PatternMatcherGun
naive: NaiveLinearGun
tmh: TmHorizonGun
bb: BitBrainGun
headon: HeadOnGun
st: WorldState
enemy: seq[EnemyInfo]
proc runRound(ctx: var Ctx, arms: var seq[ArmAcc], r: int) =
let c = ctx.c
let base = int(c.rStart[r]) - c.base
let cnt = int(c.rCount[r])
let iEnd = base + cnt
ctx.enemy[0] = EnemyInfo(id: 1)
for i in base ..< iEnd:
let ox = c.sx(i)
let oy = c.sy(i)
let localTick = int(c.tick[i]) - int(c.rStart[r])
ctx.enemy[0].x = c.ex(i)
ctx.enemy[0].y = c.ey(i)
ctx.enemy[0].heading = c.eh(i)
ctx.enemy[0].speed = c.es(i)
ctx.enemy[0].energy = c.ee(i)
ctx.enemy[0].lastSeenTick = localTick
ctx.st.enemyX = c.ex(i)
ctx.st.enemyY = c.ey(i)
ctx.st.enemyHeading = c.eh(i)
ctx.st.enemySpeed = c.es(i)
ctx.st.enemyEnergy = c.ee(i)
ctx.st.selfX = ox
ctx.st.selfY = oy
ctx.st.selfHeading = c.sh(i)
ctx.st.selfSpeed = c.ss(i)
ctx.st.selfEnergy = c.se(i)
ctx.st.selfRadarHeading = c.sh(i)
ctx.st.tick = localTick
let los = bearingDeg(ox, oy, c.ex(i), c.ey(i))
for bin in 0 ..< len(PowerBins):
let speed = bulletSpeed(PowerBins[bin])
let ib = interceptBearing(c, i, iEnd, ox, oy, speed, ctx.cont)
if not ib.ok:
for ai in 0 ..< arms.len: skip(arms[ai])
continue
let ibq = interceptBearingQuant(c, i, iEnd, ox, oy, speed)
let rng = ib.range
let targetLead = wrap180(ib.bearing - los)
# oracle: aims at the active ruler's true interception point -> 0 error
arms[A_ORACLE].record(rng, 0.0, targetLead, targetLead)
# the analyze_lead_capture_by_range.py integer-tick intercept, scored on the
# SAME ruler: this is what the coarse solve's own oracle would reach.
let qLead = if ibq.ok: wrap180(ibq.bearing - los) else: targetLead
arms[A_ORACLEQ].record(rng, wrap180(qLead - targetLead), qLead, targetLead)
# head-on: aim at the enemy's CURRENT position (worst realistic gun)
arms[A_HEADON].record(rng, wrap180(los - ib.bearing), 0.0, targetLead)
# pattern + lead-gain sweep (gain scales Pattern's lead over LOS)
let pp = ctx.pattern.predict(ctx.st, speed)
let pb = bearingDeg(ox, oy, pp.x, pp.y)
let plead = wrap180(pb - los)
arms[A_PATTERN].record(rng, wrap180(plead - targetLead), plead, targetLead)
arms[A_G15].record(rng, wrap180(1.5 * plead - targetLead), 1.5 * plead, targetLead)
arms[A_G20].record(rng, wrap180(2.0 * plead - targetLead), 2.0 * plead, targetLead)
arms[A_G30].record(rng, wrap180(3.0 * plead - targetLead), 3.0 * plead, targetLead)
# naive linear
let np = predict(ctx.naive, ctx.st, speed)
let nl = wrap180(bearingDeg(ox, oy, np.x, np.y) - los)
arms[A_NAIVE].record(rng, wrap180(nl - targetLead), nl, targetLead)
# TMHorizon
let tp = predict(ctx.tmh, ctx.st, speed)
let tl = wrap180(bearingDeg(ox, oy, tp.x, tp.y) - los)
arms[A_TMH].record(rng, wrap180(tl - targetLead), tl, targetLead)
# BitBrain (base Pattern + ADE/SBC corrector)
let bp = predict(ctx.bb, ctx.st, speed)
let bl = wrap180(bearingDeg(ox, oy, bp.x, bp.y) - los)
arms[A_BB].record(rng, wrap180(bl - targetLead), bl, targetLead)
proc runOne(runPath: string, arms: var seq[ArmAcc], shotsCont, shotsQuant: var ShotStat,
doShots: bool, timing: bool, cont: bool): int =
let t0 = epochTime()
let c = loadCorpus(runPath)
if c.n == 0: return 0
var ctx = Ctx(
c: c,
cont: cont,
pattern: PatternMatcherGun(),
naive: NaiveLinearGun(lastTick: -1),
tmh: initTmHorizonGun(),
bb: initBitBrainGun(),
headon: HeadOnGun(),
st: WorldState(arenaWidth: c.arenaW, arenaHeight: c.arenaH),
enemy: newSeq[EnemyInfo](1))
for r in 0 ..< c.rStart.len:
runRound(ctx, arms, r)
if doShots:
let ev = parseEvents(eventsPathFor(runPath))
for mode in [true, false]:
let s = validateShots(c, ev, mode)
var dst = if mode: addr shotsCont else: addr shotsQuant
dst.hits += s.hits
dst.misses += s.misses
dst.hitSumDeg += s.hitSumDeg
dst.missSumDeg += s.missSumDeg
dst.hitSumPx += s.hitSumPx
dst.missSumPx += s.missSumPx
if timing:
stderr.writeLine(fmt" {extractFilename(runPath):<16} ticks={c.n:<7} {epochTime()-t0:>6.2f}s")
c.n
proc fmt4(x: float): string =
if x.classify in {fcNan, fcInf, fcNegInf}: "-" else: fmt"{x:.4f}"
proc fmt3(x: float): string =
if x.classify in {fcNan, fcInf, fcNegInf}: "-" else: fmt"{x:.3f}"
proc main() =
var corpusRoot = "/tmp/tfil_ab2/out"
var limit = 0
var doShots = true
var timing = false
var cont = true
var i = 1
while i <= paramCount():
case paramStr(i)
of "--corpus": inc i; corpusRoot = paramStr(i)
of "--limit": inc i; limit = parseInt(paramStr(i))
of "--no-shots": doShots = false
of "--timing": timing = true
of "--ruler":
inc i
cont = paramStr(i) != "quant"
else:
stderr.writeLine("unknown arg: " & paramStr(i))
quit(2)
inc i
var runs = discoverRuns(corpusRoot)
if limit > 0 and runs.len > limit: runs.setLen(limit)
if runs.len == 0:
stderr.writeLine("no runs found under " & corpusRoot)
quit(1)
var arms: seq[ArmAcc]
for nm in ArmNames: arms.add ArmAcc(name: nm)
var shotsCont, shotsQuant: ShotStat
var t0 = epochTime()
var ticks = 0
for rp in runs:
ticks += runOne(rp, arms, shotsCont, shotsQuant, doShots, timing, cont)
let elapsed = epochTime() - t0
echo "=".repeat(120)
echo "OFFLINE PREDICTION QUALITY -- per-gun single-tick aim error vs the true interception point"
echo "=".repeat(120)
echo fmt"corpus : {corpusRoot}"
let rulerName = if cont: "continuous (physically exact)" else: "integer-tick (analyze_lead_capture_by_range.py)"
echo fmt"ruler : {rulerName}"
echo fmt"runs : {runs.len}"
echo fmt"recorded ticks: {ticks}"
let tickBins = ticks * len(PowerBins)
echo fmt"tick x bin : {tickBins}"
echo fmt"wall time : {elapsed:.2f}s ({elapsed / max(1.0, float(tickBins)) * 1000.0:.4f} ms per tick-bin)"
echo fmt"per-arm speed : {elapsed / max(1.0, float(tickBins)) * 1000.0:.4f} s per 1000 tick-bins per arm"
echo ""
echo "NOTE: offline OPEN-LOOP prediction quality only. No win/damage/survival claim."
echo ""
# validation block
echo "=".repeat(120)
echo "VALIDATION -- the ruler must pass ALL of these before any number below is trusted"
echo "=".repeat(120)
if doShots and shotsCont.hits > 0 and shotsQuant.hits > 0:
echo fmt"1. recorded shots (OUR actual server-fired bearings vs the SAME interception solve):"
for mode in [("continuous", shotsCont), ("integer", shotsQuant)]:
let s = mode[1]
let sepOk = if separationPx(s) > 2.0: "OK" else: "WEAK"
echo fmt" ruler={mode[0]:<11} hits n={s.hits:<6} mean|err|={meanHitDeg(s):>7.3f} deg / {meanHitPx(s):>7.1f} px | " &
fmt"misses n={s.misses:<6} mean|err|={meanMissDeg(s):>7.3f} deg / {meanMissPx(s):>7.1f} px | " &
fmt"separation {separationDeg(s):>6.2f}x deg / {separationPx(s):>6.2f}x px -> {sepOk}"
else:
echo "1. recorded-shot validation: SKIPPED"
let oMax = max([arms[A_ORACLE].bands[0].maxAbs, arms[A_ORACLE].bands[1].maxAbs,
arms[A_ORACLE].bands[2].maxAbs, arms[A_ORACLE].bands[3].maxAbs,
arms[A_ORACLE].bands[4].maxAbs])
let oOk = if oMax < 1e-6: "OK" else: "BROKEN"
echo fmt"2. perfect-oracle gun max |err| over all tick-bins = {oMax:.6f} deg -> {oOk}"
# ordering check: the static LOS gun must be worse than every predictive gun.
proc overallMean(arms: seq[ArmAcc], ai: int): float =
var sAbs = 0.0
var n = 0
for b in 0 ..< NBands:
sAbs += arms[ai].bands[b].sumAbs
n += arms[ai].bands[b].n
if n > 0: sAbs / float(n) else: 0.0
let mHead = overallMean(arms, A_HEADON)
let mPat = overallMean(arms, A_PATTERN)
let mTmh = overallMean(arms, A_TMH)
let mBb = overallMean(arms, A_BB)
let mLin = overallMean(arms, A_NAIVE)
let ordOk = mHead > mPat and mHead > mTmh and mHead > mBb
echo fmt"3. HeadOn (static LOS) mean|err| = {mHead:.3f} deg vs Pattern {mPat:.3f} / TMHorizon {mTmh:.3f} / BitBrain {mBb:.3f}"
let ordMsg = if ordOk: "OK (static gun worst among real guns)" else: "UNEXPECTED: a predictive gun is worse than static LOS"
echo fmt" -> {ordMsg}"
echo fmt" NaiveLinear mean|err| = {mLin:.3f} deg (over-leads; see the lead-gain sweep for why a larger"
echo fmt" lead *response* does not mean a smaller angular error)"
echo "4. determinism: run twice and diff stdout (see fixture; verified separately)."
echo ""
echo "=".repeat(120)
echo "THE BAR -- per-band mean ABSOLUTE angular aim error (deg), RMSE, sign, hit-proxy"
echo "=".repeat(120)
echo "hitProxy = fraction of tick-bins with |err| <= atan(18/range) (the angular half-width of the target disc)."
echo ""
stdout.write formatArmTable(arms)
echo ""
echo "=".repeat(120)
echo "HEADROOM -- the direct answer: how far each arm is from the oracle ceiling, per band"
echo "=" .repeat(120)
let hdr = "band Pattern n Pattern|err| Pattern hpx Oracle hpx headroom pp naive hpx TMHoriz hpx BitBrain hpx"
echo hdr
echo "-".repeat(hdr.len)
for b in 0 ..< NBands:
let pat = arms[A_PATTERN].bands[b]
let orc = arms[A_ORACLE].bands[b]
let hp = pat.hitProxy
let ohp = orc.hitProxy
echo fmt"{BandLabels[b]:<9} {pat.n:>8} {fmt3(meanAbs(pat)):>12} {fmt4(hp):>12} {fmt4(ohp):>12} {ohp - hp:>13.4f} {fmt4(arms[A_NAIVE].bands[b].hitProxy):>11} {fmt4(arms[A_TMH].bands[b].hitProxy):>12} {fmt4(arms[A_BB].bands[b].hitProxy):>13}"
echo ""
echo "hitProxy = fraction of tick-bins aimed within atan(18/range) of the true interception point."
echo "headroom pp = oracle hitProxy - Pattern hitProxy = the absolute hit-probability points available"
echo "to a perfect predictor (the campaign is playing for a slice of this)."
echo ""
let hdrq = "band OracleQuant hpx integer-solve coarseness"
echo hdrq
echo "-".repeat(hdrq.len)
for b in 0 ..< NBands:
let oq = arms[A_ORACLEQ].bands[b]
echo fmt"{BandLabels[b]:<9} {fmt4(oq.hitProxy):>15} {fmt3(meanAbs(oq)):>10} deg mean |err|"
echo "(OracleQuant aims at the analyze_lead_capture_by_range.py integer-tick intercept and is scored"
echo " on the active ruler. On the continuous ruler it measures how much of a gun's 'error' the coarse"
echo " solve itself would produce; on the integer ruler it is identically zero.)"
echo ""
echo "=".repeat(120)
echo "LEAD-GAIN SWEEP ON PATTERN -- multiply Pattern's lead (deg over LOS) by a constant"
echo "=".repeat(120)
let hdr2 = "band gain=1.0 gain=1.5 gain=2.0 gain=3.0 best-gain"
echo hdr2
echo "-".repeat(hdr2.len)
for b in 0 ..< NBands:
let g1 = meanAbs(arms[A_PATTERN].bands[b])
let g15 = meanAbs(arms[A_G15].bands[b])
let g20 = meanAbs(arms[A_G20].bands[b])
let g30 = meanAbs(arms[A_G30].bands[b])
var best = "1.0"
var bestV = g1
if g15 < bestV: bestV = g15; best = "1.5"
if g20 < bestV: bestV = g20; best = "2.0"
if g30 < bestV: bestV = g30; best = "3.0"
echo fmt"{BandLabels[b]:<9} {fmt3(g1):>10} {fmt3(g15):>10} {fmt3(g20):>10} {fmt3(g30):>10} {best} ({fmt3(bestV)})"
echo ""
echo "=".repeat(120)
echo "LEAD INFORMATIVENESS -- capture slope (regression of applied lead on required lead) and lead correlation"
echo "=".repeat(120)
echo "capture slope is job-95's metric (1.0 = perfect proportional response). corr is the Pearson"
echo "correlation of the arm's lead with the REQUIRED lead: a large slope on an uncorrelated lead is"
echo "just amplified noise. This is the table that resolves the 'naive-linear captures 2x the lead but"
echo "hits less' tension."
echo ""
let hdr3 = "band HO |err| HO |req| Pat|req| Pat cap Pat corr Lin cap Lin corr TMH cap TMH corr BB cap BB corr"
echo hdr3
echo "-".repeat(hdr3.len)
for b in 0 ..< NBands:
let sp = arms[A_PATTERN].bands[b]
let sn = arms[A_NAIVE].bands[b]
let st = arms[A_TMH].bands[b]
let sb = arms[A_BB].bands[b]
echo fmt"{BandLabels[b]:<9} {fmt3(meanAbs(arms[A_HEADON].bands[b])):>9} {fmt3(meanAbsReq(arms[A_HEADON].bands[b])):>9} {fmt3(meanAbsReq(sp)):>9} {fmt3(captureSlope(sp)):>10} {fmt3(leadCorr(sp)):>10} " &
fmt"{fmt3(captureSlope(sn)):>10} {fmt3(leadCorr(sn)):>10} {fmt3(captureSlope(st)):>10} " &
fmt"{fmt3(leadCorr(st)):>10} {fmt3(captureSlope(sb)):>10} {fmt3(leadCorr(sb)):>10}"
when isMainModule:
main()
+258
View File
@@ -0,0 +1,258 @@
# BitBrain campaign ledger
**Goal:** make ModularBot's gun **beat Pattern live** against the real DrussGT.
The user has granted full freedom over the gun ("change input, output, every
knob of it") and accepts it may fail — the deliverable is that the attempt is
visible and evidence-backed.
**THE FINAL VERDICT IS ALWAYS LIVE.** Everything in this file except the
`## Phase N` verdict lines is offline, open-loop, on a *fixed recorded enemy
trajectory*. Per `docs/offline_harness_trust.md` (commit `e40c849`) the offline
harness is trustworthy for exactly one thing: **per-gun single-tick prediction
quality on a fixed enemy trajectory** — and it is *never* trustworthy for
closed-loop questions (movement, range, round length, adaptation, gun
selection, damage, wins, survival). No offline number here is a win/damage
claim, and no phase may be called a success without a live A/B
(`tools/ab/ab_run.sh`, server-side event hit rate, left-running).
Every claim below is tagged **[MEASURED]** (a command in §0 reproduces it) or
**[INFERRED]** (reasoning from measured facts).
---
## Phase 0 — BUILD THE RULER AND ESTABLISH THE BAR *(owner: overnight job, committed)*
### 0.1 The ruler
`common_libs/gun_harness/prediction_quality.nim` + `common_libs/tests/run_prediction_quality.nim`.
At each recorded tick the shooter sits at `O = (selfX, selfY)`. For a bullet of
speed `v` the **true interception point** is the first fractional time `t > 0`
at which the enemy's ACTUAL recorded track reaches distance `v*t` from `O`
(linear interpolation between recorded ticks). A bullet fired along the bearing
to `E(t)` coincides with the enemy at `t`. Angular error is
`wrap180(bearing(O→pred) − bearing(O→E(t)))` in **degrees**; every tick is
scored for the four power bins (speeds 17/15.5/14/11), all bands share that
horizon set. Per range band we report `mean|err|`, RMSE, mean signed err and the
hit-probability proxy `mean(|err| ≤ atan(18/range))`.
The integer-tick solve from `analyze_lead_capture_by_range.py` (commit
`f91e121`) is kept as `interceptBearingQuant` and reported as `OracleQuant`; the
ruler ships the **continuous** solve because it separates recorded hits from
misses slightly better and removes the coarse solve's own overshoot
(§0.3.6). `--ruler quant` selects the integer solve.
**Data:** the recorded live-vs-real-DrussGT corpus `/tmp/tfil_ab2/out`
(70 battles / 490 rounds / 899 607 ticks + `.events.jsonl` + `.rounds.json`),
**verified present before use**. It lives in `/tmp` and is therefore ephemeral;
if a later job finds it gone, regenerate it with the A/B harness
(`tools/ab/ab_run.sh`, which sets `TR_RECORD_WORLDSTATE` so ModularBot appends
per-tick world state) and point `--corpus` at the new output root. Layout:
`<root>/<arm>/runN.jsonl` + `runN.events.jsonl` + `runN.jsonl.rounds.json`.
149 MB of JSONL is converted once per run into a compact float32 `.qcache`
(keyed on source mtime+size) and ALL measurement is taken from the cache, so two
runs are byte-identical. See §0.4 for speed.
### 0.2 Validation — the ruler must pass ALL of these **[MEASURED]**
Run: `nim c -d:release --nimcache:/tmp/nc_j98 -r common_libs/tests/run_prediction_quality.nim`
**1. Recorded HITS separate from recorded MISSES** (our ACTUAL server-fired
bearings, scored against the SAME interception solve):
| ruler | hits n | hits mean\|err\| | misses n | misses mean\|err\| | separation |
|---|---|---|---|---|---|
| continuous | 5480 | **1.360° / 10.5 px** | 48304 | **16.597° / 140.3 px** | **12.20× deg / 13.34× px** |
| integer | 5480 | 1.478° / 11.4 px | 48304 | 16.724° / 141.3 px | 11.32× / 12.43× |
(The earlier validated run quoted 11.59× / 11.6 px on hits; reproduced and
improved.) The continuous ruler is shipped because it separates better.
**2. Perfect oracle scores 0.** Max `|err|` over all 3 598 428 tick-bins =
**0.000000°**. OK.
**3. A static line-of-sight gun is far from the predictor on learnable motion.**
On a synthetic constant-velocity and a seeded random-walk trajectory the
ordering is exactly as physics demands: HeadOn (zero lead) is the worst, Pattern
and naive-linear are near-zero, and the lead-gain arms overshoot monotonically.
On the real DrussGT corpus the static gun is *not* worst — see §0.3.4, this is a
genuine property of the corpus, not a harness defect.
**4. Determinism.** Two full 70-run sweeps, stdout diffed with the two wall-time
lines excluded: **byte-identical**. (The only difference between the two raw
outputs is `wall time 406.88s` vs `402.41s` and the derived ms-per-tick-bin.)
**[MEASURED]**
**5. A real bug was found and fixed by this validation.** The ruler's
`wrap180` used Nim's float `mod`, which keeps the dividend's sign (C `fmod`), so
`(x+180) mod 360 − 180` returned `x−360` instead of the wrapped equivalent for
`x < −180`. This inflated the negative tail of every error (maxAbs read ~360°
instead of ~180°) and made HeadOn's mean error disagree with `mean|required|`.
After the fix HeadOn's `mean|err|` equals `mean|required lead|` to the last
digit at every band (see the `HO |err| / HO |req| / Pat|req|` columns in the
fixture). **A wrong ruler is worse than no ruler; this was the most important
10 minutes of the phase.**
### 0.3 THE BAR — per-band numbers (70 runs, 3 598 428 tick-bins) **[MEASURED]**
Format: `mean|err| deg` and, in brackets, `hitProxy`. `hitProxy` is the fraction
of tick-bins aimed within `atan(18/range)` of the true interception point.
| band | Pattern | naive-linear | TMHorizon | BitBrain | HeadOn (static) | Oracle |
|---|---|---|---|---|---|---|
| 0–100 | **10.56** [0.699] | 17.92 [0.651] | 10.68 [0.696] | 11.12 [0.681] | 19.62 [0.342] | 0.00 [1.000] |
| 100–200 | **14.75** [0.342] | 14.63 [0.388] | 14.84 [0.342] | 15.11 [0.324] | 19.98 [0.172] | 0.00 [1.000] |
| 200–300 | **16.61** [0.185] | 17.57 [0.192] | 16.64 [0.179] | 16.84 [0.175] | 17.34 [0.133] | 0.00 [1.000] |
| 300–450 | 17.53 [0.104] | 20.98 [0.100] | 17.57 [0.100] | 17.58 [0.103] | **14.61** [0.105] | 0.00 [1.000] |
| 450+ | 16.19 [0.077] | 22.86 [0.054] | 16.20 [0.076] | 16.20 [0.077] | **12.33** [0.098] | 0.00 [1.000] |
`n`: 4 423 / 24 908 / 74 215 / 1 119 777 / 2 311 323. Skipped (no valid
interception): 63 782 tick-bins (≈1.7 %).
**0.3.1 DIRECT ANSWER — the gap between Pattern and the oracle ceiling:**
| band | Pattern hitProxy | Oracle hitProxy | headroom (pp) |
|---|---|---|---|
| 0–100 | 0.6993 | 1.0000 | **+30.07** |
| 100–200 | 0.3418 | 1.0000 | **+65.82** |
| 200–300 | 0.1850 | 1.0000 | **+81.50** |
| 300–450 | 0.1036 | 1.0000 | **+89.64** |
| 450+ | 0.0767 | 1.0000 | **+92.33** |
**[INFERRED, important]** The oracle is *non-causal*: it aims with perfect
knowledge of the enemy's future, so its 100 % is a definition, not an
achievement, and the 92 pp at 450+ is an UPPER bound that contains both "a
better predictor could get this" and "this is physically unknowable". The
**realistic** causal bound measured today is the best arm at 450+: **HeadOn at
9.8 %**, barely above Pattern's 7.7 %. So the campaign is playing for a few
percentage points at long range, not for 92 pp. The honest target statement is
"raise the 450+ proxy from 7.7 % toward the ~10 % causal band", not "toward
100 %".
**0.3.2 The lead-gain sweep on Pattern is a DEAD END [MEASURED].** Multiply
Pattern's angular lead over LOS by a constant, per band:
| band | gain 1.0 | gain 1.5 | gain 2.0 | gain 3.0 |
|---|---|---|---|---|
| 0–100 | **10.56** | 13.76 | 20.11 | 34.77 |
| 100–200 | **14.75** | 19.64 | 26.78 | 42.97 |
| 200–300 | **16.61** | 22.26 | 29.47 | 45.23 |
| 300–450 | **17.53** | 23.28 | 29.96 | 44.27 |
| 450+ | **16.19** | 21.25 | 27.00 | 39.27 |
Gain 1.0 wins at EVERY band. Scaling Pattern's lead up makes it strictly worse.
This is the single most important negative result of Phase 0 and it should stop
any later job from "just adding more lead".
**0.3.3 The naive-linear / capture tension, resolved [MEASURED].** Capture slope
= regression of the arm's own lead on the required lead (job-95's statistic);
corr = Pearson correlation of the arm's lead with the required lead. **corr is
the informative number; a large slope on an uncorrelated lead is just amplified
noise.**
| band | mean\|req\| | Pattern cap / corr | naive-linear cap / corr | TMHorizon | BitBrain |
|---|---|---|---|---|---|
| 100–200 | 19.98 | 0.553 / 0.612 | 0.592 / 0.523 | 0.560 / 0.614 | 0.570 / 0.611 |
| 300–450 | 14.61 | 0.278 / 0.266 | 0.476 / 0.324 | 0.273 / 0.262 | 0.280 / 0.266 |
| 450+ | 12.33 | 0.175 / 0.165 | 0.310 / 0.178 | 0.174 / 0.164 | 0.175 / 0.165 |
Yes — on this corpus the naive-linear predictor applies **~1.8× more lead** than
Pattern at 450+ (0.310 vs 0.175; job-95 measured ~2×). Job-95's "we under-lead"
reading is confirmed. **But** the two arms carry almost the same lead
*information* (corr 0.178 vs 0.165), so the extra amplitude buys nothing and
costs angular accuracy: naive-linear's `mean|err|` is 22.86° vs Pattern's
16.19° at 450+. **Conclusion: the campaign's lever is lead INFORMATION
(correlation), not lead RESPONSE (capture slope).** Capturing more of an
uninformative lead is worse than capturing little of it — which is also exactly
why the gain sweep fails.
**0.3.4 The surprise: at long range, static line-of-sight beats Pattern.**
HeadOn (aim at the enemy's current position) has `mean|err|` 14.61°/12.33° and
`hitProxy` 0.105/0.098 at 300–450/450+, both better than Pattern's
17.53°/16.19° and 0.104/0.077. **[INFERRED]** At 450+ the required lead
(`mean|req|` = 12.3°) is essentially unpredictable from the past (Pattern
corr 0.165), so Pattern's predicted lead is mostly variance added to a nearly
uninformative signal; a zero-lead aim has error = `|required lead|`, which is
smaller. Consistent with the live record: the live bot's own applied lead
capture was 0.135 at 450+ (job-95), i.e. the live bot was already nearly
zero-lead and hit 9.14 % there; Pattern's offline proxy is 7.7 %.
**[INFERRED / CAVEAT]** The corpus is open-loop: DrussGT's recorded dodge was a
reaction to the LIVE bot's (near-zero-lead) bullets. Replaying Pattern on that
trajectory cannot show what DrussGT would do against Pattern's bullets. This
makes a **live A/B of HeadOn vs Pattern at long range the highest-value cheap
experiment in the campaign** (see §0.6). No offline claim that "HeadOn
beats Pattern" is permitted — only the live A/B decides.
**0.3.5 BitBrain, as shipped, is Pattern [MEASURED].** BitBrain's base is
Pattern and its ADE/SBC corrector changes almost nothing: 450+ `mean|err|`
16.200° vs Pattern 16.193°, `hitProxy` 0.0767 vs 0.0767. TMHorizon likewise
(16.199° / 0.0757). The corrector is currently **adding no measurable aim
information** on this corpus. That is the thing Phase 1 must change.
**0.3.6 Ruler resolution is NOT the limiter [MEASURED].** Aiming at the
integer-tick solve instead of the exact intercept costs only 0.29–1.10° of mean
error (`OracleQuant` column). So the "maybe the oracle only reaches 35 % because
the solve is coarse" worry is dead: the coarse/fine difference is ≈0.3° at long
range, far below the target tolerance (1.93° at 450+). Whatever caps the score,
it is the enemy's unpredictability, not the ruler.
### 0.4 Speed **[MEASURED]**
Full 70-run / 899 607-tick / 3 598 428 tick-bin sweep, 10 arms:
**406.9 s wall**, i.e. `0.1131 ms per tick-bin` over 10 arms,
**≈ 0.045 s per gun per 1000 ticks** (1000 ticks × 4 power bins).
BitBrain is the dominant cost (its ADE pass runs on every `predict` call);
Pattern/TMHorizon cache their per-tick work. A single-arm Pattern-only sweep is
several times cheaper. The binary cache (§0.1) is what makes repeat sweeps
affordable: without it every run re-parses 149 MB of JSONL.
### 0.5 How to reproduce **[MEASURED]**
```
nim c -d:release --nimcache:/tmp/nc_j98 -o:/tmp/bbq_run \
common_libs/tests/run_prediction_quality.nim
/tmp/bbq_run --corpus /tmp/tfil_ab2/out # full bar, ~7 min
/tmp/bbq_run --corpus /tmp/tfil_ab2/out --limit 10 # fast subset
/tmp/bbq_run --corpus /tmp/tfil_ab2/out --ruler quant # integer-tick solve
```
Verbatim full output: `common_libs/tests/prediction_quality_results.txt`.
Determinism: two consecutive full runs are byte-identical except the two
wall-time lines.
**Clean-checkout proof [MEASURED]:** `git archive HEAD | tar -x -C /tmp/bbq_clean`
then, from `/tmp/bbq_clean`,
`nim c -d:release --nimcache:/tmp/nc_j98 -o:bbq_run common_libs/tests/run_prediction_quality.nim`
builds, and `./bbq_run --corpus /tmp/tfil_ab2/out --limit 3` runs and prints the
same tables (separation 13.68× px on the 3-run subset). The committed harness is
self-contained; only the corpus is external.
### 0.6 Designs still to try (seed for later phases)
Ordered by expected value per unit of effort. Phase 0 has already killed one.
| # | design | why it is worth trying | status |
|---|---|---|---|
| D1 | **Live A/B: HeadOn at 450+ vs Pattern** (distance-gated switch, or HeadOn-only control) | Offline says a static gun beats Pattern at long range; the cheapest possible test of the campaign's central premise | **TODO (highest value, live)** |
| D2 | **Pattern variants that raise lead CORRELATION at long range**: longer keys / multi-length keys, per-distance learned pattern tables, different match weighting, k-NN over movement signatures | The lever is corr (0.165 at 450+), not gain; the ruler measures exactly this | TODO (offline-searchable) |
| D3 | **Supervise BitBrain with the ruler's own labels** — per-tick bearing error to the true intercept, trained on N−1 runs, evaluated on a held-out run | BitBrain's corrector currently adds nothing (0.3.5); the ruler gives it a real target. Must hold out runs or it is overfitting | TODO (offline-searchable) |
| D4 | **A causal "predictability" gate**: at each tick estimate whether the future is predictable (e.g. recent pattern-match score, reversal entropy) and fall back to HeadOn/low-variance aim when it is not | Directly attacks the 0.3.4 failure mode without needing a better long-range predictor | TODO |
| D5 | Power policy at long range (already partly done live): lower power = faster bullet = less lead error | Shortens the horizon the predictor must extrapolate; affects hit rate, live-only verdict | TODO (offline proxy only) |
| D6 | Lead-gain sweep 1.0/1.5/2.0/3.0 | **DEAD — measured.** Gain 1.0 wins at every band (§0.3.2) | **KILLED** |
Every D-item must end in a live A/B before any phase verdict.
### 0.7 What would make us quit
> If (a) no causal design raises the 450+ `hitProxy` above the static-gun
> reference (~0.10) on held-out runs by a margin larger than the run-to-run
> spread, **and** (b) the live A/B of the best such design shows no hit-rate or
> damage gain over Pattern with the left-running liveness check satisfied, then
> the campaign stops and we ship the simpler gun. We do not keep tuning an
> offline proxy that has stopped predicting live outcomes.
---
## Phase 1 — *(unclaimed; append below)*
## Phase 2 — *(unclaimed; append below)*