From c305ef4212b641a44a26eba98b41f8bd1e6ace69 Mon Sep 17 00:00:00 2001 From: Davide Cappellini Date: Fri, 25 Sep 2026 00:10:19 +0200 Subject: [PATCH] BitBrain campaign phase 1: the missing gain sweep and BitBrain as a lead-gain corrector MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Task A — the gain region Phase 0 never covered (gain < 1). Extend the prediction-quality ruler with gain 0.25/0.50/0.75 arms and a fixed causal per-band arm. Full 70-run result: the hitProxy-argmax curve is [1.00, 1.00, 1.00, 0.00, 0.00] — Pattern below 300 px, HeadOn above — worth +0.13 pp at 300-450 and +2.16 pp at 450+ (0.0767 -> 0.0984). Lead correlation is identical for every g>0 (Pearson is scale-invariant), so a shrinking gain adds no lead information; and the least-squares optimum [1,1,1,1,0.25,0.25] diverges from the hitProxy optimum because Pattern's lead errors are bimodal. Task B — rebuild guns/bitbrain_gun.nim as a lead-gain corrector: aim = LOS + gain*(patternAim - LOS), gain learned online per range band by ranking candidate gains on the hit-probability proxy (the observed lead label via tmhObservedAt), gated to range >= 300 px. ADE+SBC output removed. Offline (70 runs): Pattern below 300 px, +0.37 pp at 300-450, +1.90 pp at 450+ (hitProxy 0.0957 vs 0.0767), matching the fixed rule to within 0.27 pp at 450+ and exceeding it at 300-450. ~0.0007 ms/tick marginal (old gun ~0.114 ms/tick). Default off; rack membership, env report and guard tests (bitbrain 32, registration 13, rack 48, tm_pattern 20, env_report) unchanged and green. Ledger: docs/bitbrain_campaign.md Phase 1, including the three on-file negatives and the causal-shippability note. No live claim. --- common_libs/guns/bitbrain_gun.nim | 497 +++++++----------- .../tests/prediction_quality_results.txt | 105 +++- common_libs/tests/run_prediction_quality.nim | 99 +++- docs/bitbrain_campaign.md | 165 +++++- 4 files changed, 553 insertions(+), 313 deletions(-) diff --git a/common_libs/guns/bitbrain_gun.nim b/common_libs/guns/bitbrain_gun.nim index 701550b..8e9378d 100644 --- a/common_libs/guns/bitbrain_gun.nim +++ b/common_libs/guns/bitbrain_gun.nim @@ -1,77 +1,105 @@ -## bitbrain_gun.nim — BitBrain (ADE + SBC) FINE-GRAINED AIM CORRECTOR. +## bitbrain_gun.nim — BitBrain (id 16), REBUILT as a LEAD-GAIN CORRECTOR. ## -## THE BRIEF THIS IMPLEMENTS (from the offline gate test, docs/bitbrain_gate_test.md): -## * BASE — the shipped Pattern gun's prediction (`guns/pattern_matcher`). -## BitBrain supplies only a small ANGULAR CORRECTION on top of it, -## exactly the shape the gate test measured and the shape TMHorizon -## uses. That keeps the comparison against Pattern/TMHorizon -## apples-to-apples. -## * INPUT — the SAME 53 bits TMHorizon uses: the 49-bit draft spec -## (`tmhBaseBits`) PLUS a 4-bit horizon one-hot (`tmhLits`). These are -## reused from `guns/tm_horizon.nim`, not re-derived. -## * OUTPUT — a fine-grained angular-correction CLASS over ±`TR_BITBRAIN_RANGE` -## degrees (`TR_BITBRAIN_N` bins, default 32). The readout is the -## ARGMAX class centre (the gate test MEASURED that argmax is the -## winning readout; the count-weighted mean is a shrinkage predictor -## that lowers the hit rate). Zero correction when there is no -## evidence. -## * LABEL — the +h-tick FACT from our OWN observation ring (the ring the -## embedded TmHorizonGun maintains): `h = round(dist/speed)`, -## `speed = 20 - 3*power`, clamped to [10, 50]. Never crosses a round -## boundary (pending samples are dropped on a round reset). -## * TRAIN — ONLINE / PREQUENTIAL: predict, then learn the resolved fact when -## it becomes due `h` ticks later. -## * AD LAYER — synthesised for OUR data. The MNIST weights are useless. -## center = 0 (the inputs are BINARY; the reference 127 would collapse -## the code to a polarity count). Thresholds start from a small -## heuristic that fires ~1 % from the first ticks, are then calibrated -## from a running score histogram to the paper's ~1 % operating point -## (the gate test's percentile init, made online), and are nudged by -## the library's deterministic `adaptThresholds` homeostasis. +## ── WHY THIS FILE WAS REWRITTEN (Phase 0/1 evidence) ────────────────────────── +## The previous design was an ADDITIVE angular shift: an ADE+SBC network +## classified the +h-tick angular error over ±`TR_BITBRAIN_RANGE` degrees and +## added the argmax class centre to Pattern's bearing. Phase 0 measured it as +## statistically identical to Pattern (`docs/bitbrain_gun_verdict.md`, +## commit d93ce44) and as carrying no measurable aim information +## (450+: 16.200 deg vs Pattern's 16.193; `docs/bitbrain_campaign.md` §0.3.5). ## -## MEMORY MODES (`TR_BITBRAIN_MEM`): -## perRound (DEFAULT) — wipe the SBCs every round. The gate test measured this -## as the WINNING regime. -## retained — accumulate across the whole battle/enemy and wipe only -## on a target change / new battle. This is what the user -## asked for, and the gate test measured it as the WEAKEST -## regime: the idempotent SBC only ADDS, so it saturates. -## decay — retained PLUS a periodic partial wipe of the SBC -## tensors (TR_BITBRAIN_DECAY every N samples, a fraction -## TR_BITBRAIN_DECAY_FRAC of words zeroed). This is the one -## mechanism with a measured diagnosis behind it: the SBC -## saturates and a bounded/decaying memory should help. +## Phase 1 measured the actual lever. The gain sweep found that gains >= 1 are +## strictly worse at every band and that the optimal gain is BELOW 1.0 at long +## range (450+: ~0.25). A fractional gain leaves the Pearson lead *correlation* +## unchanged (correlation is invariant under positive scaling), so a smaller +## gain does not add information — it shrinks the magnitude of an uninformative +## Pattern lead toward the low-variance static (HeadOn) aim. The right output is +## therefore a multiplicative GAIN on Pattern's lead, not a class-based additive +## shift. See `docs/bitbrain_campaign.md` §Phase 1 for the measured curve. +## +## ── THE DESIGN ──────────────────────────────────────────────────────────────── +## * BASE — the shipped Pattern gun's prediction (`guns/pattern_matcher`), +## reached through the TmHorizonGun observation ring. +## * OUTPUT — `aim = LOS + gain * (patternAim - LOS)`, i.e. Pattern's lead over +## the line of sight is multiplied by a learned `gain` (one of +## `BB_CAND`, so it may be BELOW 1.0 — the point). +## * LABEL — the same deferred-label path the old corrector used: at fire +## time we remember the base lead and the aim tolerance; `h = +## round(dist/speed)` ticks later `tmhObservedAt` returns the +## enemy's OBSERVED bearing from the firing position. `requiredLead +## = observedBearing - LOS` and `baseLead = baseBearing - LOS`, so a +## candidate gain scores a hit on this sample when +## `|gain*baseLead - requiredLead| <= tolerance`. +## * TRAIN — ONLINE / PREQUENTIAL per range band: for each candidate gain we +## count the fraction of resolved samples that would have been +## within the target's angular half-width (`atan(18/range)`, the +## SAME tolerance the offline ruler uses). The band's gain is the +## argmax hit rate. THIS is the key lesson of Phase 1: the +## least-squares gain and the hit-probability-optimal gain DIVERGE +## (Pattern's lead errors are bimodal), so the learner optimises the +## hit-probability proxy directly instead of mean squared error. +## * STATE — the range band (the ruler's 5 bands). Range is known causally at +## fire time, so a per-band gain table is shippable with no learning +## at all; BitBrain learns that table online. The correction is +## additionally gated to bands with range >= 300 px +## (`BB_GAIN_BAND_MIN`), where Phase 1 measured Pattern's lead to be +## uninformative. That gate is causal (range is known). +## +## The gain statistics are battle-scale: a round boundary wipes the observation +## ring and deferred labels but NOT the gain counts (a new round is not a new +## enemy). `resetLearning` wipes them on a new battle / target change; with +## `TR_BITBRAIN_MEM=decay` every `TR_BITBRAIN_DECAY` resolved samples decays the +## counts by `TR_BITBRAIN_DECAY_FRAC` toward the gain-1.0 column. +## +## ── WHAT IS STILL HERE ONLY FOR THE BOOT REPORT / GUARD TESTS ───────────────── +## The ADE+SBC network is GONE from the gun. The 53-bit TMH input, the class +## geometry (`bbCenterDeg`/`bbClassOf`), `TR_BITBRAIN_N`/`NADE`/`WARMUP`/`ADAPT`/ +## `CALIB`/`SEED` and `TR_BITBRAIN_RANGE` are retained as resolved configuration +## so the boot report (`env_report.nim`) and the registration guard tests keep +## working unchanged; they no longer affect the gain learner. The generic +## `common_libs/bitbrain/` library is untouched and still tested by +## `test_bitbrain.nim`. ## ## DEFAULT OFF / PARITY: this gun is admitted ONLY when `TR_RACK_BITBRAIN` says so -## (default `off`) AND it never runs its network until `predict` is first called -## (`ensureInit`). With the shipped rack the live loop never calls `predict`, so -## no network is built, no RNG is touched and the shipped bot is unchanged. +## (default `off`). The shipped rack never calls `predict`, so `ensureInit` never +## runs and the shipped bot is byte-for-byte unchanged. -import std/[math, os, strutils, strformat, random] +import std/[math, os, strutils, strformat] import gun_harness/gun_interface import guns/tm_horizon import guns/pattern_matcher -import bitbrain/bitbrain const ## ── env knobs (all resolved once at gun construction) ───────────────────── BB_MEM_ENV* = "TR_BITBRAIN_MEM" ## perRound|retained|decay - BB_N_ENV* = "TR_BITBRAIN_N" ## correction classes - BB_NADE_ENV* = "TR_BITBRAIN_NADE" ## ADEs per address decoder - BB_RANGE_ENV* = "TR_BITBRAIN_RANGE" ## class half-range, degrees + BB_N_ENV* = "TR_BITBRAIN_N" ## (legacy geometry; inert) + BB_NADE_ENV* = "TR_BITBRAIN_NADE" ## (legacy ADE count; inert) + BB_RANGE_ENV* = "TR_BITBRAIN_RANGE" ## (legacy class half-range; inert) BB_LOG_ENV* = "TR_BITBRAIN_LOG" ## 1 = per-change [bb] log - BB_MIN_OBS_ENV* = "TR_BITBRAIN_MIN_OBS" ## resolved samples before correction - BB_WARMUP_ENV* = "TR_BITBRAIN_WARMUP" ## samples before percentile init - BB_ADAPT_ENV* = "TR_BITBRAIN_ADAPT" ## homeostasis interval (samples) - BB_CALIB_ENV* = "TR_BITBRAIN_CALIB" ## percentile recalibration interval + BB_MIN_OBS_ENV* = "TR_BITBRAIN_MIN_OBS" ## samples before a band is trusted + BB_WARMUP_ENV* = "TR_BITBRAIN_WARMUP" ## (legacy; inert) + BB_ADAPT_ENV* = "TR_BITBRAIN_ADAPT" ## (legacy; inert) + BB_CALIB_ENV* = "TR_BITBRAIN_CALIB" ## (legacy; inert) BB_DECAY_ENV* = "TR_BITBRAIN_DECAY" ## decay interval (samples) - BB_DECAY_FRAC_ENV* = "TR_BITBRAIN_DECAY_FRAC" ## fraction of words zeroed per decay - BB_SEED_ENV* = "TR_BITBRAIN_SEED" ## deterministic AD/decay seed + BB_DECAY_FRAC_ENV* = "TR_BITBRAIN_DECAY_FRAC" ## per-decay count shrink + BB_SEED_ENV* = "TR_BITBRAIN_SEED" ## (legacy; inert) BB_RESET_ON_TARGET_ENV* = "TR_BITBRAIN_RESET_ON_TARGET" ## ── fixed geometry ──────────────────────────────────────────────────────── - BB_WIDTHS* = [6, 8, 10, 12] ## the paper's multi-width ADs - BB_TARGET_RATE* = 0.01 ## the paper's ~1 % firing target BB_PENDING_CAP* = 512 ## deferred-label queue (>= 4 buckets x 50 ticks) + ## ── the gain learner ────────────────────────────────────────────────────── + BB_NBANDS* = 5 ## the ruler's range bands + BB_NHB* = 4 ## horizon buckets (for the per-tick label dedupe) + BB_BAND_LO* = [0.0, 100.0, 200.0, 300.0, 450.0] + BB_BAND_HI* = [100.0, 200.0, 300.0, 450.0, 1.0e18] + ## The candidate lead gains the band selector picks from. 0.0 == HeadOn (aim + ## at the current position) and 1.0 == Pattern (use the full lead). + BB_CAND* = [0.0, 0.25, 0.50, 0.75, 1.0] + BB_NCAND* = 5 + BB_BB_RADIUS* = 18.0 ## hit-detection radius in px (ruler tolerance) + ## Apply the correction only from this band up (range >= BB_BAND_LO[3] = 300). + ## [MEASURED] below 300 Pattern's lead is informative and shrinking it loses + ## hits; see the header note. + BB_GAIN_BAND_MIN* = 3 ## ── shipped defaults ────────────────────────────────────────────────────── BB_N_DEF = 32 BB_NADE_DEF = 256 @@ -90,19 +118,21 @@ type bmPerRound, bmRetained, bmDecay BbPending = object - ## One deferred training sample. `lits` is the exact literal vector the ADs - ## saw at fire time; the label is resolved `horizon` ticks later. + ## One deferred training sample. `lead` is Pattern's lead over LOS at fire + ## time (radians) and `tol` the target's angular half-width then; the label + ## is resolved `horizon` ticks later. fireTick: int horizon: int + band: int selfX*, selfY: float baseBearing: float - lits: array[TMH_NLITS, uint8] + lead: float + tolDeg: float BitBrainGun* = object tmh: TmHorizonGun - bb: BitBrain initialized: bool - # ── resolved config ────────────────────────────────────────────────────── + # ── resolved config (kept in the boot report) ───────────────────────────── nClasses*: int maxDeg*: float nAde*: int @@ -116,36 +146,24 @@ type decayFrac*: float seed*: int64 resetOnTarget*: bool - # ── AD calibration state ───────────────────────────────────────────────── - rng: Rand - hist: seq[seq[int32]] ## per-AD raw-score histogram (bins 2w+1) - histTotal: int - sampleCount*: int - sinceAdapt: int - sinceCalib: int + # ── gain learner: hit counts per (range band x candidate gain) ──────────── + bandHits*: array[BB_NBANDS, array[BB_NCAND, float64]] + bandN*: array[BB_NBANDS, float64] + trained*: int sinceDecay: int - decays*: int - # ── scratch (avoid per-sample allocation) ──────────────────────────────── - scratch: seq[seq[int32]] - counts: seq[int] - # ── deferred labels ────────────────────────────────────────────────────── + decays*: int + # ── readout / accounting ────────────────────────────────────────────────── + lastGain*: array[BB_NBANDS, float] + corrections*: int + lastLogKey: string + # ── deferred labels ─────────────────────────────────────────────────────── pending: array[BB_PENDING_CAP, BbPending] pendingCount*: int pendingDropped*: int - # ── per-tick caches ────────────────────────────────────────────────────── + # ── per-tick caches ─────────────────────────────────────────────────────── lastTick: int lastEnqTick: int lastEnqBucket: int - cachedBits: array[TMH_N_BASE, uint8] - cachedBitsTick: int - bitsValid: bool - # ── accounting / logging ───────────────────────────────────────────────── - trained*: int - lastBest: int - lastShift*: float - corrections*: int - lastLogKey: string - lastLogTick: int observedTargetId*: int # ── small pure helpers ─────────────────────────────────────────────────────── @@ -162,8 +180,8 @@ proc memModeName*(m: BitMemMode): string = of bmDecay: "decay" proc parseMemMode*(value: string): BitMemMode = - ## Empty / unknown values fall back to the shipped `perRound` (the measured - ## winning regime), so a typo cannot silently select another regime. + ## Empty / unknown values fall back to the shipped `perRound`, so a typo + ## cannot silently select another regime. case value.strip().toLowerAscii() of "retained", "retain", "accum", "accumulate": bmRetained of "decay", "forget", "age": bmDecay @@ -186,20 +204,32 @@ proc envBoolBB(name: string, default: bool): bool = else: default proc bbCenterDeg*(k, nClasses: int, maxDeg: float): float = - ## Centre (degrees) of correction class `k` over ±maxDeg. + ## Centre (degrees) of correction class `k` over ±maxDeg. Retained for the + ## registration guard test and the boot report; inert for the gain learner. let w = 2.0 * maxDeg / float(nClasses) -maxDeg + (float(k) + 0.5) * w proc bbClassOf*(errRad: float, nClasses: int, maxDeg: float): int = ## Bin a signed angular error (radians) into one of `nClasses` bins over - ## [−maxDeg, +maxDeg] (the gate test's `binOf`). + ## [−maxDeg, +maxDeg]. Retained for the registration guard test; inert. let x = radToDeg(errRad) var k = int((x + maxDeg) / (2.0 * maxDeg) * float(nClasses)) if k < 0: k = 0 if k >= nClasses: k = nClasses - 1 k -# ── construction / lazy network build ──────────────────────────────────────── +proc bbBandOf*(range: float): int {.inline.} = + ## Range band (the ruler's bands), known causally at fire time. + for b in 0 ..< BB_NBANDS: + if range >= BB_BAND_LO[b] and range < BB_BAND_HI[b]: return b + BB_NBANDS - 1 + +proc bbTolDeg*(range: float): float {.inline.} = + ## The target's angular half-width at `range` — atan(18/range) — i.e. the exact + ## tolerance the offline ruler uses for its hit-probability proxy. + radToDeg(arctan2(BB_BB_RADIUS, max(range, 1e-9))) + +# ── construction / lazy init ───────────────────────────────────────────────── proc initBitBrainGun*(): BitBrainGun = result.nClasses = clamp(envIntBB(BB_N_ENV, BB_N_DEF), 2, 512) @@ -219,148 +249,58 @@ proc initBitBrainGun*(): BitBrainGun = result.lastEnqTick = -1 result.lastEnqBucket = -1 result.observedTargetId = -1 - result.rng = initRand(result.seed + 991) - -proc resetThresholdsHeuristic(g: var BitBrainGun) = - ## Cold-start thresholds: a small multiple of the raw-score standard deviation - ## puts every ADE near the paper's ~1 % firing rate from the FIRST ticks, so - ## the SBCs see a useful (sparse) coincidence set immediately and inference - ## never degenerates into an O(nAde^2) dense scan. The running-histogram - ## percentile calibration replaces these once warmup has passed. - for a in 0..= t` is CLOSEST to - ## `BB_TARGET_RATE * total` — the gate test's percentile init, run online over - ## the running histogram. This is what pins the realised firing rate near 1 %. - if g.histTotal <= 0: return - let target = BB_TARGET_RATE * float(g.histTotal) - for a in 0.. g.warmupN: - inc g.sinceAdapt - inc g.sinceCalib - if g.sinceAdapt >= g.adaptEvery: - for a in 0..= g.calibEvery: - g.calibrate() - g.sinceCalib = 0 - -# ── one AD pass: firing counts + histogram + inference ─────────────────────── - -proc bbObserve(g: var BitBrainGun, lits: array[TMH_NLITS, uint8]) = - ## Drive every ADE: update its firing accumulator and the score histogram, - ## collect the active list, then infer the class counts into `g.counts`. - for a in 0.. 0'i32: int(c) - 1 else: int(-c) - 1 - let pol = if c > 0'i32: 1 else: -1 - raw += pol * int(lits[idx]) - inc g.hist[a][raw + w] - if raw * sc >= int(g.bb.ades[a].thresholds[e]): - g.scratch[a].add int32(e) - inc g.bb.ades[a].fireCounts[e] - inc g.histTotal - for k in 0.. 0'i32: int(c) - 1 else: int(-c) - 1 - let pol = if c > 0'i32: 1 else: -1 - raw += pol * int(lits[idx]) - if raw * sc >= int(g.bb.ades[a].thresholds[e]): - g.scratch[a].add int32(e) - for sl in 0..= 1.0: return + for b in 0 ..< BB_NBANDS: + for ci in 0 ..< BB_NCAND: + g.bandHits[b][ci] *= f + g.bandN[b] *= f inc g.decays +proc bbGain(g: BitBrainGun, band: int): float = + ## The band's gain is the candidate with the highest observed hit rate. + ## Ties keep the SMALLER candidate (the scan is ascending), which is the + ## conservative choice for the long-range regime this corrector targets. + ## Returns 1.0 (Pattern) below the range gate or when the band is cold. + if band < BB_GAIN_BAND_MIN: return 1.0 + if g.bandN[band] < float(g.minObs): return 1.0 + var best = 4 # gain 1.0 + var bestRate = -1.0 + for ci in 0 ..< BB_NCAND: + let rate = g.bandHits[band][ci] / g.bandN[band] + if rate > bestRate: + bestRate = rate + best = ci + BB_CAND[best] + # ── deferred-label resolution (prequential learning) ───────────────────────── proc resolvePending(g: var BitBrainGun, state: WorldState) = var w = 0 - for i in 0.. state.tick: @@ -370,12 +310,11 @@ proc resolvePending(g: var BitBrainGun, state: WorldState) = let obs = tmhObservedAt(g.tmh, state.tick, p.selfX, p.selfY) if obs.ok and (state.tick - obs.lastSeenTick) <= TMH_STALE_MAX: let err = wrapRadBB(obs.bearing - p.baseBearing) - let cls = bbClassOf(err, g.nClasses, g.maxDeg) - g.bbLearn(p.lits, cls) - inc g.trained + let reqLead = wrapRadBB(err + p.lead) + g.bbAccumulate(radToDeg(p.lead), radToDeg(reqLead), p.tolDeg, p.band) inc g.sinceDecay if g.memMode == bmDecay and g.sinceDecay >= g.decayEvery: - g.applyDecay() + g.bbApplyDecay() g.sinceDecay = 0 else: inc g.pendingDropped @@ -385,62 +324,49 @@ proc resolvePending(g: var BitBrainGun, state: WorldState) = # ── logging ────────────────────────────────────────────────────────────────── -proc bbLog(g: var BitBrainGun, state: WorldState, h, bucket, total, best: int) = - ## ONE change-gated `[bb]` line (behind TR_BITBRAIN_LOG=1) so the user tailing - ## the GUI log sees what the corrector is thinking, not one line per tick. +proc bbLog(g: var BitBrainGun, state: WorldState, band: int, gain: float) = + ## ONE change-gated `[bb]` line (behind TR_BITBRAIN_LOG=1) so a user tailing + ## the GUI log sees the gain the corrector is applying. if not g.logEnabled: return - let shift = bbCenterDeg(best, g.nClasses, g.maxDeg) - let key = fmt"{best}|{shift:.1f}" + let key = fmt"{gain:.2f}|{band}" if key == g.lastLogKey: return - if state.tick == g.lastLogTick: return g.lastLogKey = key - g.lastLogTick = state.tick - var nz = 0 - for k in 0.. 0: inc nz - echo fmt"[bb] t={state.tick} h={h} bucket={bucket} cls={best}/{g.nClasses} " & - fmt"shift={shift:+.1f}deg cnt={g.counts[best]}/{total} nz={nz} " & - fmt"trained={g.trained} samples={g.sampleCount} pend={g.pendingCount} " & - fmt"mode={memModeName(g.memMode)} warm={(if g.trained >= g.minObs: 1 else: 0)}" + var rate = 0.0 + for ci in 0 ..< BB_NCAND: + if abs(BB_CAND[ci] - gain) < 1e-9: rate = g.bandHits[band][ci] / max(1.0, g.bandN[band]) + echo fmt"[bb] t={state.tick} band={BB_BAND_LO[band]:.0f}+ gain={gain:.2f} " & + fmt"rate={rate:.3f} n={g.bandN[band]:.0f} trained={g.trained} " & + fmt"pend={g.pendingCount} dropped={g.pendingDropped} mode={memModeName(g.memMode)}" # ── reset hooks (mirroring TmHorizonGun) ───────────────────────────────────── proc resetRound(g: var BitBrainGun) = - ## PER-ROUND wipe. Always clear the observation ring, deferred labels and - ## per-tick caches (the bots teleport between rounds). In `perRound` mode the - ## SBCs are wiped too; `retained`/`decay` keep them across the round. + ## PER-ROUND reset: observation ring, deferred labels and per-tick caches (the + ## bots teleport between rounds). The gain counts are deliberately KEPT — they + ## are battle-scale and a new round is not a new enemy. g.tmh.resetRoundState() g.pendingCount = 0 g.lastTick = -1 g.lastEnqTick = -1 g.lastEnqBucket = -1 - g.bitsValid = false g.lastLogKey = "" - g.lastLogTick = -1 - if g.memMode == bmPerRound: - g.bb.resetLearning() - g.trained = 0 proc resetRoundState*(g: var BitBrainGun) = if not g.initialized: return g.resetRound() proc resetLearning*(g: var BitBrainGun, reason = "") = - ## PER-BATTLE / PER-ENEMY wipe: SBCs, AD thresholds, histograms and counters. + ## PER-BATTLE / PER-ENEMY wipe: gain counts, counters and the round state. if not g.initialized: return - g.bb.resetLearning() - g.resetThresholdsHeuristic() - for a in 0.. 0 and g.logEnabled: echo fmt"[bb-reset] reason={reason}" @@ -462,8 +388,8 @@ proc targetChanged*(g: var BitBrainGun, enemyId: int): bool = proc isWarmedUp*(g: BitBrainGun): bool {.inline.} = true proc networkBytes*(g: BitBrainGun): int = - ## Bytes held by the AD/SBC network (0 until the network is built). - if g.initialized: g.bb.memoryBytes else: 0 + ## No neural network is held any more; kept for the boot report / guard test. + 0 proc predict*(g: var BitBrainGun, state: WorldState, bulletSpeed: float): GunPrediction = @@ -477,58 +403,43 @@ proc predict*(g: var BitBrainGun, state: WorldState, tmhUpdateHistory(g.tmh, state) g.resolvePending(state) g.lastTick = state.tick - g.bitsValid = false - # The base prediction is Pattern; BitBrain only corrects its bearing. + # The base prediction is Pattern; BitBrain only scales its lead over LOS. let base = g.tmh.pattern.predict(state, bulletSpeed) if bulletSpeed <= 0.0: return base let dist = hypot(state.enemyX - state.selfX, state.enemyY - state.selfY) let h = tmhHorizonFor(dist, bulletSpeed) - let bucket = tmhHorizonBucket(h) + let hb = tmhHorizonBucket(h) + let band = bbBandOf(dist) - if not g.bitsValid or g.cachedBitsTick != state.tick: - g.cachedBits = tmhBaseBits(g.tmh, state) - g.cachedBitsTick = state.tick - g.bitsValid = true - let lits = tmhLits(g.cachedBits, bucket) + let los = arctan2(state.enemyY - state.selfY, state.enemyX - state.selfX) + let baseBearing = arctan2(base.y - state.selfY, base.x - state.selfX) + let lead = wrapRadBB(baseBearing - los) - # Observe this input (AD pass + inference) and advance the calibration clock. - g.bbObserve(lits) - g.afterSample() - - # Enqueue one deferred sample per (tick, bucket): predict runs once per power - # bin, so all four horizons contribute evidence. - if g.lastEnqTick != state.tick or g.lastEnqBucket != bucket: + # Enqueue one deferred sample per (tick, horizon bucket): `predict` runs once + # per power bin, so all four horizons contribute evidence. + if g.lastEnqTick != state.tick or g.lastEnqBucket != hb: if g.pendingCount < BB_PENDING_CAP: g.pending[g.pendingCount] = BbPending( - fireTick: state.tick, horizon: h, + fireTick: state.tick, horizon: h, band: band, selfX: state.selfX, selfY: state.selfY, - baseBearing: arctan2(base.y - state.selfY, base.x - state.selfX), - lits: lits) + baseBearing: baseBearing, lead: lead, tolDeg: bbTolDeg(dist)) inc g.pendingCount else: inc g.pendingDropped g.lastEnqTick = state.tick - g.lastEnqBucket = bucket + g.lastEnqBucket = hb - # Readout: argmax class centre, zero correction with no evidence / cold. - var shiftDeg = 0.0 - if g.trained >= g.minObs: - var total = 0 - for k in 0.. 0: - var best = 0 - for k in 1.. g.counts[best]: best = k - shiftDeg = bbCenterDeg(best, g.nClasses, g.maxDeg) - g.lastBest = best - g.lastShift = shiftDeg - inc g.corrections - g.bbLog(state, h, bucket, total, best) - - if shiftDeg == 0.0: return base - tmhApplyShift(state.selfX, state.selfY, base.x, base.y, shiftDeg) + # Readout: a fractional gain may be BELOW 1.0. When cold / gated out the + # learner returns 1.0 and the base prediction is returned unchanged. + let gain = g.bbGain(band) + g.lastGain[band] = gain + if abs(gain - 1.0) < 1e-9: return base + inc g.corrections + g.bbLog(state, band, gain) + tmhApplyShift(state.selfX, state.selfY, base.x, base.y, + radToDeg((gain - 1.0) * lead)) proc onResult*(g: var BitBrainGun, e: FeedbackEvent) = ## Labels come from our own observation ring, not from virtual-bullet diff --git a/common_libs/tests/prediction_quality_results.txt b/common_libs/tests/prediction_quality_results.txt index 547226b..aa1f13f 100644 --- a/common_libs/tests/prediction_quality_results.txt +++ b/common_libs/tests/prediction_quality_results.txt @@ -6,8 +6,8 @@ ruler : continuous (physically exact) runs : 70 recorded ticks: 899607 tick x bin : 3598428 -wall time : 406.88s (0.1131 ms per tick-bin) -per-arm speed : 0.1131 s per 1000 tick-bins per arm +wall time : 271.23s (0.0754 ms per tick-bin) +per-arm speed : 0.0754 s per 1000 tick-bins per arm NOTE: offline OPEN-LOOP prediction quality only. No win/damage/survival claim. @@ -18,7 +18,7 @@ VALIDATION -- the ruler must pass ALL of these before any number below is trus ruler=continuous hits n=5480 mean|err|= 1.360 deg / 10.5 px | misses n=48304 mean|err|= 16.597 deg / 140.3 px | separation 12.20x deg / 13.34x px -> OK ruler=integer hits n=5480 mean|err|= 1.478 deg / 11.4 px | misses n=48304 mean|err|= 16.724 deg / 141.3 px | separation 11.32x deg / 12.43x px -> OK 2. perfect-oracle gun max |err| over all tick-bins = 0.000000 deg -> OK -3. HeadOn (static LOS) mean|err| = 13.217 deg vs Pattern 16.609 / TMHorizon 16.627 / BitBrain 16.635 +3. HeadOn (static LOS) mean|err| = 13.217 deg vs Pattern 16.609 / TMHorizon 16.627 / BitBrain 13.627 -> UNEXPECTED: a predictive gun is worse than static LOS NaiveLinear mean|err| = 22.086 deg (over-leads; see the lead-gain sweep for why a larger lead *response* does not mean a smaller angular error) @@ -85,23 +85,47 @@ TMHorizon 200-300 74215 16.644 21.142 1.091 0.1786 84.4 TMHorizon 300-450 1119777 17.572 21.878 0.691 0.0996 84.01 TMHorizon 450+ 2311323 16.199 20.025 -0.687 0.0757 81.19 (TMHorizon: 63782 tick-bins had no valid interception) -BitBrain 0-100 4423 11.118 15.268 -1.507 0.6810 64.72 -BitBrain 100-200 24908 15.107 19.782 1.289 0.3241 76.63 -BitBrain 200-300 74215 16.838 21.426 1.341 0.1747 105.04 -BitBrain 300-450 1119777 17.575 21.894 0.933 0.1025 100.98 -BitBrain 450+ 2311323 16.200 20.032 -0.647 0.0767 96.38 +BitBrain 0-100 4423 10.555 14.796 -0.404 0.6993 64.72 +BitBrain 100-200 24908 14.745 19.510 1.276 0.3418 76.63 +BitBrain 200-300 74215 16.610 21.174 1.350 0.1850 81.45 +BitBrain 300-450 1119777 15.518 19.170 0.773 0.1073 85.65 +BitBrain 450+ 2311323 12.608 15.431 -0.351 0.0957 79.92 (BitBrain: 63782 tick-bins had no valid interception) +PatternGain0.25 0-100 4423 16.984 19.724 -1.060 0.3914 46.00 +PatternGain0.25 100-200 24908 17.791 20.374 -0.027 0.1661 51.44 +PatternGain0.25 200-300 74215 15.950 18.761 0.698 0.1276 53.22 +PatternGain0.25 300-450 1119777 14.131 16.972 0.755 0.1025 52.64 +PatternGain0.25 450+ 2311323 12.261 14.888 -0.358 0.0931 51.24 + (PatternGain0.25: 63782 tick-bins had no valid interception) +PatternGain0.50 0-100 4423 14.583 16.877 -0.841 0.4689 50.78 +PatternGain0.50 100-200 24908 16.103 18.679 0.407 0.1673 58.27 +PatternGain0.50 200-300 74215 15.316 18.261 0.915 0.1366 62.44 +PatternGain0.50 300-450 1119777 14.491 17.566 0.819 0.1003 61.45 +PatternGain0.50 450+ 2311323 12.938 15.799 -0.453 0.0883 59.63 + (PatternGain0.50: 63782 tick-bins had no valid interception) +PatternGain0.75 0-100 4423 12.356 15.103 -0.622 0.6396 57.28 +PatternGain0.75 100-200 24908 14.965 18.369 0.841 0.2337 66.41 +PatternGain0.75 200-300 74215 15.528 19.120 1.133 0.1389 71.95 +PatternGain0.75 300-450 1119777 15.639 19.275 0.884 0.0952 72.71 +PatternGain0.75 450+ 2311323 14.280 17.588 -0.548 0.0813 68.72 + (PatternGain0.75: 63782 tick-bins had no valid interception) +PatternBandGain 0-100 4423 10.555 14.796 -0.404 0.6993 64.72 +PatternBandGain 100-200 24908 14.745 19.510 1.276 0.3418 76.63 +PatternBandGain 200-300 74215 16.610 21.174 1.350 0.1850 81.45 +PatternBandGain 300-450 1119777 14.607 17.606 0.690 0.1049 46.38 +PatternBandGain 450+ 2311323 12.326 15.017 -0.263 0.0984 45.49 + (PatternBandGain: 63782 tick-bins had no valid interception) ======================================================================================================================== HEADROOM -- the direct answer: how far each arm is from the oracle ceiling, per band ======================================================================================================================== band Pattern n Pattern|err| Pattern hpx Oracle hpx headroom pp naive hpx TMHoriz hpx BitBrain hpx --------------------------------------------------------------------------------------------------------------- -0-100 4423 10.555 0.6993 1.0000 0.3007 0.6505 0.6955 0.6810 -100-200 24908 14.745 0.3418 1.0000 0.6582 0.3879 0.3418 0.3241 -200-300 74215 16.610 0.1850 1.0000 0.8150 0.1924 0.1786 0.1747 -300-450 1119777 17.531 0.1036 1.0000 0.8964 0.0995 0.0996 0.1025 -450+ 2311323 16.193 0.0767 1.0000 0.9233 0.0537 0.0757 0.0767 +0-100 4423 10.555 0.6993 1.0000 0.3007 0.6505 0.6955 0.6993 +100-200 24908 14.745 0.3418 1.0000 0.6582 0.3879 0.3418 0.3418 +200-300 74215 16.610 0.1850 1.0000 0.8150 0.1924 0.1786 0.1850 +300-450 1119777 17.531 0.1036 1.0000 0.8964 0.0995 0.0996 0.1073 +450+ 2311323 16.193 0.0767 1.0000 0.9233 0.0537 0.0757 0.0957 hitProxy = fraction of tick-bins aimed within atan(18/range) of the true interception point. headroom pp = oracle hitProxy - Pattern hitProxy = the absolute hit-probability points available @@ -129,6 +153,51 @@ band gain=1.0 gain=1.5 gain=2.0 gain=3.0 best-gain 300-450 17.531 23.279 29.955 44.274 1.0 (17.531) 450+ 16.193 21.245 26.997 39.271 1.0 (16.193) +======================================================================================================================== +PHASE 1 — THE MISSING GAIN SWEEP: Pattern lead x gain in [0.00, 1.00] (0 = HeadOn, 1 = Pattern) +======================================================================================================================== +Format per cell: mean|err| deg [hitProxy]. hitProxy is the objective. gain 0.0 is HeadOn, +gain 1.0 is Pattern. The per-band gain table IS causal to APPLY (range is known at fire time, +so a per-band lookup needs no learning); its ESTIMATION from these same runs is in-sample. + +band |req| deg g=0.00 [hpx] g=0.25 [hpx] g=0.50 [hpx] g=0.75 [hpx] g=1.00 [hpx] bestHpx dHpx bestErr +-------------------------------------------------------------------------------------------------------------------- +0-100 19.619 19.619 [0.342] 16.984 [0.391] 14.583 [0.469] 12.356 [0.640] 10.555 [0.699] 1.00 +0.0000 1.00 +100-200 19.982 19.982 [0.172] 17.791 [0.166] 16.103 [0.167] 14.965 [0.234] 14.745 [0.342] 1.00 +0.0000 1.00 +200-300 17.341 17.341 [0.133] 15.950 [0.128] 15.316 [0.137] 15.528 [0.139] 16.610 [0.185] 1.00 +0.0000 0.50 +300-450 14.607 14.607 [0.105] 14.131 [0.102] 14.491 [0.100] 15.639 [0.095] 17.531 [0.104] 0.00 +0.0013 0.25 +450+ 12.326 12.326 [0.098] 12.261 [0.093] 12.938 [0.088] 14.280 [0.081] 16.193 [0.077] 0.00 +0.0216 0.25 + +OPTIMAL GAIN CURVE (hitProxy-argmax per band) and its implied hit-probability gain vs Pattern: + 0-100 best gain 1.00 hitProxy 0.6993 vs Pattern 0.6993 => +0.0000 pp + 100-200 best gain 1.00 hitProxy 0.3418 vs Pattern 0.3418 => +0.0000 pp + 200-300 best gain 1.00 hitProxy 0.1850 vs Pattern 0.1850 => +0.0000 pp + 300-450 best gain 0.00 hitProxy 0.1049 vs Pattern 0.1036 => +0.0013 pp + 450+ best gain 0.00 hitProxy 0.0984 vs Pattern 0.0767 => +0.0216 pp + gain = [ 0-100->1.00 100-200->1.00 200-300->1.00 300-450->0.00 450+->0.00 ] + +DIRECT COMPARISON — Pattern vs the FIXED causal per-band rule [1,1,1,0.00,0.00] vs BitBrain (learned online): +band Pattern hpx fixed-band hpx BitBrain hpx fixed-Pat pp BB-Pat pp +-------------------------------------------------------------------------------- +0-100 0.6993 0.6993 0.6993 +0.0000 +0.0000 +100-200 0.3418 0.3418 0.3418 +0.0000 +0.0000 +200-300 0.1850 0.1850 0.1850 +0.0000 +0.0000 +300-450 0.1036 0.1049 0.1073 +0.0013 +0.0037 +450+ 0.0767 0.0984 0.0957 +0.0216 +0.0190 +fixed-band hpx = the [1,1,1,0,0] table applied causally; it was selected in-sample. +BitBrain is learned online from labels inside each run (cold start at gain 1.0). + +LEAD CORRELATION PER GAIN (Pearson of applied lead with required lead). Pearson is invariant +under positive scaling, so every g>0 column must be IDENTICAL to Pattern; g=0 has no lead and +therefore no correlation. If they match, a shrinking gain does NOT add lead information — it +only shrinks the magnitude of an uninformative signal (the gain-sweep mechanism). +band corr(g) g=0.25 g=0.50 g=0.75 g=1.00 +0-100 0.774 0.774 0.774 0.774 +100-200 0.612 0.612 0.612 0.612 +200-300 0.457 0.457 0.457 0.457 +300-450 0.266 0.266 0.266 0.266 +450+ 0.165 0.165 0.165 0.165 + ======================================================================================================================== LEAD INFORMATIVENESS -- capture slope (regression of applied lead on required lead) and lead correlation ======================================================================================================================== @@ -139,8 +208,8 @@ hits less' tension. band HO |err| HO |req| Pat|req| Pat cap Pat corr Lin cap Lin corr TMH cap TMH corr BB cap BB corr ----------------------------------------------------------------------------------------------------------------------- -0-100 19.619 19.619 19.619 0.649 0.774 0.573 0.362 0.668 0.776 0.701 0.768 -100-200 19.982 19.982 19.982 0.553 0.612 0.592 0.523 0.560 0.614 0.570 0.611 -200-300 17.341 17.341 17.341 0.449 0.457 0.519 0.429 0.452 0.460 0.459 0.457 -300-450 14.607 14.607 14.607 0.278 0.266 0.476 0.324 0.273 0.262 0.280 0.266 -450+ 12.326 12.326 12.326 0.175 0.165 0.310 0.178 0.174 0.164 0.175 0.165 +0-100 19.619 19.619 19.619 0.649 0.774 0.573 0.362 0.668 0.776 0.649 0.774 +100-200 19.982 19.982 19.982 0.553 0.612 0.592 0.523 0.560 0.614 0.553 0.612 +200-300 17.341 17.341 17.341 0.449 0.457 0.519 0.429 0.452 0.460 0.449 0.457 +300-450 14.607 14.607 14.607 0.278 0.266 0.476 0.324 0.273 0.262 0.148 0.213 +450+ 12.326 12.326 12.326 0.175 0.165 0.310 0.178 0.174 0.164 0.036 0.101 diff --git a/common_libs/tests/run_prediction_quality.nim b/common_libs/tests/run_prediction_quality.nim index b7ae367..ad927cd 100644 --- a/common_libs/tests/run_prediction_quality.nim +++ b/common_libs/tests/run_prediction_quality.nim @@ -30,8 +30,29 @@ const A_NAIVE* = 7 A_TMH* = 8 A_BB* = 9 + # ── Phase 1: the MISSING gain sweep. Gains >= 1 were measured worse at every + # band in Phase 0; the unexplored region is gain < 1. gain 0.0 is HeadOn + # (A_HEADON) and gain 1.0 is Pattern (A_PATTERN), so only 0.25/0.50/0.75 are + # new arms. Their `leadCorr` is IDENTICAL to Pattern's by construction (Pearson + # correlation is invariant under positive scaling) — printed only to prove it. + A_G025* = 10 + A_G050* = 11 + A_G075* = 12 + # Phase 1 fixed causal per-band gain rule: the hitProxy-argmax curve measured by + # the sub-unity sweep ([1,1,1,0,0] == Pattern below 300 px, HeadOn above). This + # is the rule BitBrain must match; it needs no learning (range is known at fire + # time). The table was selected in-sample from this corpus. + A_BAND* = 13 + BandGainTable* = [1.0, 1.0, 1.0, 0.0, 0.0] ArmNames* = ["Oracle", "OracleQuant", "HeadOn", "Pattern", "PatternGain1.5", - "PatternGain2.0", "PatternGain3.0", "NaiveLinear", "TMHorizon", "BitBrain"] + "PatternGain2.0", "PatternGain3.0", "NaiveLinear", "TMHorizon", "BitBrain", + "PatternGain0.25", "PatternGain0.50", "PatternGain0.75", "PatternBandGain"] + +const + ## The five sub-unity gain arms, in increasing order, resolved to arm indices. + ## gain 0.0 == HeadOn, gain 1.0 == Pattern. + GainArmIdx* = [A_HEADON, A_G025, A_G050, A_G075, A_PATTERN] + GainValues* = [0.0, 0.25, 0.50, 0.75, 1.0] # ── the naive-linear control (job-95's LIN_M = 4 extrapolation) ────────────── # @@ -133,6 +154,13 @@ proc runRound(ctx: var Ctx, arms: var seq[ArmAcc], r: int) = arms[A_G15].record(rng, wrap180(1.5 * plead - targetLead), 1.5 * plead, targetLead) arms[A_G20].record(rng, wrap180(2.0 * plead - targetLead), 2.0 * plead, targetLead) arms[A_G30].record(rng, wrap180(3.0 * plead - targetLead), 3.0 * plead, targetLead) + # sub-unity gains (Phase 1) — the region Phase 0 never covered + arms[A_G025].record(rng, wrap180(0.25 * plead - targetLead), 0.25 * plead, targetLead) + arms[A_G050].record(rng, wrap180(0.50 * plead - targetLead), 0.50 * plead, targetLead) + arms[A_G075].record(rng, wrap180(0.75 * plead - targetLead), 0.75 * plead, targetLead) + # the fixed causal per-band rule (Phase 1 hitProxy-argmax curve) + let bg = BandGainTable[bandOf(rng)] + arms[A_BAND].record(rng, wrap180(bg * plead - targetLead), bg * plead, targetLead) # naive linear let np = predict(ctx.naive, ctx.st, speed) let nl = wrap180(bearingDeg(ox, oy, np.x, np.y) - los) @@ -332,6 +360,75 @@ proc main() = if g30 < bestV: bestV = g30; best = "3.0" echo fmt"{BandLabels[b]:<9} {fmt3(g1):>10} {fmt3(g15):>10} {fmt3(g20):>10} {fmt3(g30):>10} {best} ({fmt3(bestV)})" + echo "" + echo "=".repeat(120) + echo "PHASE 1 — THE MISSING GAIN SWEEP: Pattern lead x gain in [0.00, 1.00] (0 = HeadOn, 1 = Pattern)" + echo "=".repeat(120) + echo "Format per cell: mean|err| deg [hitProxy]. hitProxy is the objective. gain 0.0 is HeadOn," + echo "gain 1.0 is Pattern. The per-band gain table IS causal to APPLY (range is known at fire time," + echo "so a per-band lookup needs no learning); its ESTIMATION from these same runs is in-sample." + echo "" + var hdrg = "band |req| deg" + for gi in 0 ..< GainValues.len: hdrg.add fmt" g={GainValues[gi]:.2f} [hpx]" + hdrg.add " bestHpx dHpx bestErr" + echo hdrg + echo "-".repeat(hdrg.len) + for b in 0 ..< NBands: + var line = fmt"{BandLabels[b]:<9} {fmt3(meanAbsReq(arms[A_PATTERN].bands[b])):>9}" + var bestHi = 0 + var bestHp = -1.0 + var bestEi = 0 + var bestEr = Inf + for gi in 0 ..< GainValues.len: + let s = arms[GainArmIdx[gi]].bands[b] + let e = meanAbs(s) + let hp = s.hitProxy + line.add fmt"{fmt3(e):>7} [{fmt3(hp)}] " + if hp > bestHp: bestHp = hp; bestHi = gi + if e < bestEr: bestEr = e; bestEi = gi + let patHp = arms[A_PATTERN].bands[b].hitProxy + line.add fmt" {GainValues[bestHi]:.2f} {bestHp-patHp:+.4f} {GainValues[bestEi]:.2f}" + echo line + echo "" + echo "OPTIMAL GAIN CURVE (hitProxy-argmax per band) and its implied hit-probability gain vs Pattern:" + var curve = " gain = [ " + for b in 0 ..< NBands: + var bestHi = 0 + var bestHp = -1.0 + for gi in 0 ..< GainValues.len: + let hp = arms[GainArmIdx[gi]].bands[b].hitProxy + if hp > bestHp: bestHp = hp; bestHi = gi + curve.add fmt"{BandLabels[b]}->{GainValues[bestHi]:.2f} " + let patHp = arms[A_PATTERN].bands[b].hitProxy + echo fmt" {BandLabels[b]:<9} best gain {GainValues[bestHi]:.2f} hitProxy {bestHp:.4f} vs Pattern {patHp:.4f} => {bestHp-patHp:+.4f} pp" + echo curve & "]" + echo "" + echo "DIRECT COMPARISON — Pattern vs the FIXED causal per-band rule [1,1,1,0.00,0.00] vs BitBrain (learned online):" + let hdrd = "band Pattern hpx fixed-band hpx BitBrain hpx fixed-Pat pp BB-Pat pp" + echo hdrd + echo "-".repeat(hdrd.len) + for b in 0 ..< NBands: + let patHp = arms[A_PATTERN].bands[b].hitProxy + let fixHp = arms[A_BAND].bands[b].hitProxy + let bbHp = arms[A_BB].bands[b].hitProxy + echo fmt"{BandLabels[b]:<9} {patHp:>11.4f} {fixHp:>16.4f} {bbHp:>14.4f} {fixHp-patHp:>+14.4f} {bbHp-patHp:>+10.4f}" + echo "fixed-band hpx = the [1,1,1,0,0] table applied causally; it was selected in-sample." + echo "BitBrain is learned online from labels inside each run (cold start at gain 1.0)." + echo "" + echo "LEAD CORRELATION PER GAIN (Pearson of applied lead with required lead). Pearson is invariant" + echo "under positive scaling, so every g>0 column must be IDENTICAL to Pattern; g=0 has no lead and" + echo "therefore no correlation. If they match, a shrinking gain does NOT add lead information — it" + echo "only shrinks the magnitude of an uninformative signal (the gain-sweep mechanism)." + let hdrc = "band " & " corr(g) " + var hdrc2 = hdrc + for gi in 1 ..< GainValues.len: hdrc2.add fmt" g={GainValues[gi]:.2f}" + echo hdrc2 + for b in 0 ..< NBands: + var line = fmt"{BandLabels[b]:<9}" + for gi in 1 ..< GainValues.len: + line.add fmt" {fmt3(leadCorr(arms[GainArmIdx[gi]].bands[b])):>8}" + echo line + echo "" echo "=".repeat(120) echo "LEAD INFORMATIVENESS -- capture slope (regression of applied lead on required lead) and lead correlation" diff --git a/docs/bitbrain_campaign.md b/docs/bitbrain_campaign.md index 876cb25..8b465d2 100644 --- a/docs/bitbrain_campaign.md +++ b/docs/bitbrain_campaign.md @@ -253,6 +253,169 @@ Every D-item must end in a live A/B before any phase verdict. --- -## Phase 1 — *(unclaimed; append below)* +## Phase 1: the missing gain sweep *(owner: overnight job j100, committed)* + +### Three negatives are on file (the morning reader must see these) + +1. **BitBrain as previously shipped was statistically identical to Pattern + live.** 30 runs/arm, MDE 24.4 dmg/run: `docs/bitbrain_gun_verdict.md`, + commit `d93ce44`. +2. **Gun-mixing does not disrupt DrussGT.** TMHorizon+BitBrain mixed into the + rack produced no dodge disruption: `docs/gun_mix_disruption.md`, commit + `32a5e72`. +3. **BitBrain added no measurable aim information offline.** 450+: 16.200 deg + vs Pattern 16.193, hitProxy 0.0767 vs 0.0767 (`a82c864`, §0.3.5 above). + +These bound the plausible upside: the previous BitBrain output — an ADDITIVE +angular shift — was information-free, so Phase 1 changes the output shape, not +the learning rate. + +### 1.1 The missing sweep: Pattern x gain in [0.00, 1.00] **[MEASURED]** + +`common_libs/tests/run_prediction_quality.nim` now carries sub-unity gain arms +(gain 0.0 is HeadOn, gain 1.0 is Pattern). Full 70-run sweep, 14 arms, +3 598 428 tick-bins, `common_libs/tests/prediction_quality_results.txt`. +Cells are `mean|err| deg [hitProxy]`; **hitProxy is the objective** (fraction of +tick-bins within `atan(18/range)`, the ruler's proxy for hit probability). + +| band | |req| deg | g=0.00 | g=0.25 | g=0.50 | g=0.75 | g=1.00 | best hpx | dHpx | best |err| | +|---|---|---|---|---|---|---|---|---| +| 0–100 | 19.62 | 19.62 [.342] | 16.98 [.391] | 14.58 [.469] | 12.36 [.640] | **10.56 [.699]** | 1.00 | +.0000 | 1.00 | +| 100–200 | 19.98 | 19.98 [.172] | 17.79 [.166] | 16.10 [.167] | 14.97 [.234] | **14.75 [.342]** | 1.00 | +.0000 | 1.00 | +| 200–300 | 17.34 | 17.34 [.133] | 15.95 [.128] | 15.32 [.137] | 15.53 [.139] | **16.61 [.185]** | 1.00 | +.0000 | 0.50 | +| 300–450 | 14.61 | **14.61 [.105]** | 14.13 [.102] | 14.49 [.100] | 15.64 [.095] | 17.53 [.104] | 0.00 | +.0013 | 0.25 | +| 450+ | 12.33 | **12.33 [.098]** | 12.26 [.093] | 12.94 [.088] | 14.28 [.081] | 16.19 [.077] | 0.00 | +.0216 | 0.25 | + +**The optimal gain curve (hitProxy-argmax per band) is** +**`[1.00, 1.00, 1.00, 0.00, 0.00]`** — use Pattern's full lead below 300 px and +**aim at the enemy's current position (zero lead, HeadOn) at 300+ px**. The +implied hit-probability gains vs Pattern are +0.00 pp (0–300), +0.13 pp +(300–450) and **+2.16 pp at 450+ (0.0767 -> 0.0984, a 28 % relative rise)**. +Weighted by tick-bin share (450+ is 64.2 %, 300–450 is 31.1 %) the whole-corpus +proxy rises ~+1.4 pp. + +**[CAUSAL-SHIPPABILITY, important]** The rule only needs the *range*, which is +known at fire time, so applying `[1,1,1,0,0]` is causal and needs **no learning +at all** — a per-band table is shippable as a constant, exactly like the +Pattern radial-offset knob. What is *not* causal is the **estimation** of the +table from the same runs (it is in-sample here); a shipped table would be fitted +offline on past battles or learned online, which is what BitBrain does. The +table's value is robust to that caveat because the winning entries are the two +extremes (full Pattern lead, or zero lead), not a knife-edge intermediate value. + +### 1.2 Shrinking the gain does NOT add lead information **[MEASURED]** + +| band | corr @ g=0.25 | g=0.50 | g=0.75 | g=1.00 | +|---|---|---|---|---| +| 0–100 | 0.774 | 0.774 | 0.774 | 0.774 | +| 100–200 | 0.612 | 0.612 | 0.612 | 0.612 | +| 200–300 | 0.457 | 0.457 | 0.457 | 0.457 | +| 300–450 | 0.266 | 0.266 | 0.266 | 0.266 | +| 450+ | 0.165 | 0.165 | 0.165 | 0.165 | + +Every `g > 0` column is **identical**: Pearson correlation is invariant under +positive scaling. A fractional gain therefore buys nothing on the +lead-information axis — it only shrinks the magnitude of an uninformative signal +toward the low-variance static aim. This confirms §0.3.3's reading and is the +mechanism behind the whole curve. + +**The MSE-optimal gain and the hitProxy-optimal gain diverge [MEASURED, key].** +The `best |err|` column is the least-squares optimum `[1.00, 1.00, 0.50, 0.25, +0.25]`. At 200–300, MSE wants ~0.5 but hitProxy wants 1.0: Pattern's lead errors +are **bimodal** (it either nails the lead or is far off), so shrinking every +sample trades many small-within-tolerance hits for a smaller tail. Any *learned* +corrector that minimises squared error will therefore under-perform at mid range +— a fact this phase measured the hard way (an MSE-gain BitBrain dropped the +200–300 proxy from 0.190 to 0.131). + +### 1.3 BitBrain rebuilt as a lead-gain corrector **[MEASURED]** + +`common_libs/guns/bitbrain_gun.nim` is rewritten. The ADE+SBC network is gone +from the gun; the output is now a multiplicative gain on Pattern's lead, +`aim = LOS + gain * (patternAim - LOS)`, with `gain` chosen from a fixed +candidate set `{0, 0.25, 0.5, 0.75, 1.0}`. + +* **Label path** (unchanged): at fire time we remember Pattern's lead and the +target's angular half-width `atan(18/range)`; `h = round(dist/speed)` ticks +later `tmhObservedAt` returns the enemy's observed bearing from the firing +position, giving `requiredLead = observedBearing - LOS`. +* **Learning rule** (changed): for every resolved sample we score *each* +candidate by whether `|gain*baseLead - requiredLead| <= atan(18/range)` and keep +the hit counts per range band; the band's gain is the **argmax hit rate** — the +hit-probability proxy itself, not squared error. This directly fixes the +bimodality failure above. +* **State / gate**: the state is the range band (causally known). The correction +is applied only for range >= 300 px (`BB_GAIN_BAND_MIN = 3`), below which +Pattern's lead is informative. +* **Stats are battle-scale**: a round boundary keeps the counts; a new battle or +target change (`resetLearning`/`targetChanged`) wipes them. + +Offline 70-run result (same ruler, same run as §1.1): + +| band | Pattern hpx | fixed rule [1,1,1,0,0] | BitBrain hpx | fixed−Pat | BB−Pat | +|---|---|---|---|---|---| +| 0–100 | 0.6993 | 0.6993 | 0.6993 | +0.0000 | +0.0000 | +| 100–200 | 0.3418 | 0.3418 | 0.3418 | +0.0000 | +0.0000 | +| 200–300 | 0.1850 | 0.1850 | 0.1850 | +0.0000 | +0.0000 | +| 300–450 | 0.1036 | 0.1049 | **0.1073** | +0.0013 | **+0.0037** | +| 450+ | 0.0767 | **0.0984** | 0.0957 | +0.0216 | **+0.0190** | + +BitBrain's effective point estimates match the fixed table to within 0.27 pp at +450+ and actually exceed it at 300–450. The internal log shows why it is not +exactly equal: at long range the candidate hit rates are near-tied, so the +argmax flips between 0.00 and 0.25 (450+) / 0.00 and 0.75 (300–450); the +aggregate still lands on the right side. A fixed table is more stable; the +learned version needs no table and adapts per battle. + +**Cost [MEASURED].** `/tmp/bench_bb` (200k ticks, 4 power bins) — Pattern +0.0224 ms/tick, BitBrain 0.0230 ms/tick, i.e. a **marginal ~0.0007 ms/tick** +(0.7 us/tick), versus the ~0.114 ms/tick of the old ADE+SBC gun and the +13.16 ms/tick budget. The whole 14-arm offline sweep now runs in 271 s +(0.0754 ms per tick-bin vs Phase 0's 0.1131 for 10 arms). + +**Guard tests stay green [MEASURED].** `test_bitbrain` 32 checks, +`test_bitbrain_registration` 13, `test_rack_membership` 48, +`test_tm_pattern_registration` 20, `test_env_report` all pass. `env_report.nim` +is unchanged: the legacy BitBrain knobs are still resolved and reported. The +ruler still validates (recorded hits 1.360 deg vs misses 16.597 deg, +separation 12.20x deg / 13.34x px; oracle 0.000000 deg). + +### 1.4 Verdict + +**[MEASURED] Yes — a per-band lead-gain rule beats Pattern offline**, but only in +the two long bands: `[1,1,1,0,0]` gives +0.13 pp at 300–450 and **+2.16 pp at +450+** (a ~28 % relative rise in the hit-proxy, against Phase 0's realistic +causal ceiling of HeadOn's ~0.10, which it reaches). Below 300 px Pattern is +already optimal and the rule is a no-op. + +**[MEASURED] Yes — BitBrain-as-gain-corrector matches that rule** (+1.90 pp at +450+, +0.37 pp at 300–450 over Pattern; within 0.27 pp of the fixed table at +450+, above it at 300–450), while being **causally learnable online** and +~163x cheaper per tick than the old gun. + +**[INFERRED / honest caveat] BitBrain is not required to capture the gain** — a +constant `[1,1,1,0,0]` table would do it. BitBrain's value is that it discovers +the table per battle without one; its cost is cold-start (it begins at gain 1.0, +explaining the 0.27 pp gap at 450+) and the near-tie jitter in its point +estimates. So the honest ship decision is: **the per-band rule is the thing to +test live, and it can be shipped either as a constant table or as this online +learner** — the learner is redundant if a table is acceptable, and preferable +only if the optimum is expected to drift per enemy. The highest-value live +experiment is unchanged and stronger: a **range-gated Pattern<->HeadOn switch at +~300 px vs pure Pattern** (Phase 0's D1), left-running. **No live win is claimed +here** — the live gate is a separate phase. + +### 1.5 Designs after Phase 1 + +| # | design | status | +|---|---|---| +| D1 | Live A/B: Pattern<->HeadOn switch at ~300 px vs Pattern | **TODO (highest value, live)** — offline now says +2.16 pp at 450+ | +| D2 | Raise Pattern lead *correlation* at long range (longer/multi-length keys, per-distance tables, k-NN) | TODO (offline-searchable) | +| D3 | Match-quality-conditional gain (keep full lead when the pattern match is good, shrink when bad) — attacks the bimodality that makes the MSE gain diverge | TODO (offline-searchable) | +| D4 | A causal predictability gate falling back to HeadOn when the future is unpredictable | TODO | +| D5 | Long-range power policy (already partly live) | TODO (offline proxy only) | +| D6 | Lead-gain sweep gains >= 1.0 | **DEAD — measured** (§0.3.2) | +| D7 | Constant sub-unity lead gain at all ranges | **DEAD — measured** (§1.1: gain < 1 hurts below 300) | + ## Phase 2 — *(unclaimed; append below)*