BitBrain campaign phase 1: the missing gain sweep and BitBrain as a lead-gain corrector

Task A — the gain region Phase 0 never covered (gain < 1). Extend the
prediction-quality ruler with gain 0.25/0.50/0.75 arms and a fixed causal
per-band arm. Full 70-run result: the hitProxy-argmax curve is
[1.00, 1.00, 1.00, 0.00, 0.00] — Pattern below 300 px, HeadOn above —
worth +0.13 pp at 300-450 and +2.16 pp at 450+ (0.0767 -> 0.0984). Lead
correlation is identical for every g>0 (Pearson is scale-invariant), so a
shrinking gain adds no lead information; and the least-squares optimum
[1,1,1,1,0.25,0.25] diverges from the hitProxy optimum because Pattern's
lead errors are bimodal.

Task B — rebuild guns/bitbrain_gun.nim as a lead-gain corrector:
aim = LOS + gain*(patternAim - LOS), gain learned online per range band by
ranking candidate gains on the hit-probability proxy (the observed lead
label via tmhObservedAt), gated to range >= 300 px. ADE+SBC output removed.
Offline (70 runs): Pattern below 300 px, +0.37 pp at 300-450, +1.90 pp at
450+ (hitProxy 0.0957 vs 0.0767), matching the fixed rule to within 0.27 pp
at 450+ and exceeding it at 300-450. ~0.0007 ms/tick marginal (old gun
~0.114 ms/tick). Default off; rack membership, env report and guard tests
(bitbrain 32, registration 13, rack 48, tm_pattern 20, env_report) unchanged
and green.

Ledger: docs/bitbrain_campaign.md Phase 1, including the three on-file
negatives and the causal-shippability note. No live claim.
This commit is contained in:
2026-09-25 00:10:19 +02:00
parent 140fe2519a
commit c305ef4212
4 changed files with 553 additions and 313 deletions
+204 -293
View File
@@ -1,77 +1,105 @@
## bitbrain_gun.nim — BitBrain (ADE + SBC) FINE-GRAINED AIM CORRECTOR.
## bitbrain_gun.nim — BitBrain (id 16), REBUILT as a LEAD-GAIN CORRECTOR.
##
## THE BRIEF THIS IMPLEMENTS (from the offline gate test, docs/bitbrain_gate_test.md):
## * BASE — the shipped Pattern gun's prediction (`guns/pattern_matcher`).
## BitBrain supplies only a small ANGULAR CORRECTION on top of it,
## exactly the shape the gate test measured and the shape TMHorizon
## uses. That keeps the comparison against Pattern/TMHorizon
## apples-to-apples.
## * INPUT — the SAME 53 bits TMHorizon uses: the 49-bit draft spec
## (`tmhBaseBits`) PLUS a 4-bit horizon one-hot (`tmhLits`). These are
## reused from `guns/tm_horizon.nim`, not re-derived.
## * OUTPUT — a fine-grained angular-correction CLASS over ±`TR_BITBRAIN_RANGE`
## degrees (`TR_BITBRAIN_N` bins, default 32). The readout is the
## ARGMAX class centre (the gate test MEASURED that argmax is the
## winning readout; the count-weighted mean is a shrinkage predictor
## that lowers the hit rate). Zero correction when there is no
## evidence.
## * LABEL — the +h-tick FACT from our OWN observation ring (the ring the
## embedded TmHorizonGun maintains): `h = round(dist/speed)`,
## `speed = 20 - 3*power`, clamped to [10, 50]. Never crosses a round
## boundary (pending samples are dropped on a round reset).
## * TRAIN — ONLINE / PREQUENTIAL: predict, then learn the resolved fact when
## it becomes due `h` ticks later.
## * AD LAYER — synthesised for OUR data. The MNIST weights are useless.
## center = 0 (the inputs are BINARY; the reference 127 would collapse
## the code to a polarity count). Thresholds start from a small
## heuristic that fires ~1 % from the first ticks, are then calibrated
## from a running score histogram to the paper's ~1 % operating point
## (the gate test's percentile init, made online), and are nudged by
## the library's deterministic `adaptThresholds` homeostasis.
## ── WHY THIS FILE WAS REWRITTEN (Phase 0/1 evidence) ──────────────────────────
## The previous design was an ADDITIVE angular shift: an ADE+SBC network
## classified the +h-tick angular error over ±`TR_BITBRAIN_RANGE` degrees and
## added the argmax class centre to Pattern's bearing. Phase 0 measured it as
## statistically identical to Pattern (`docs/bitbrain_gun_verdict.md`,
## commit d93ce44) and as carrying no measurable aim information
## (450+: 16.200 deg vs Pattern's 16.193; `docs/bitbrain_campaign.md` §0.3.5).
##
## MEMORY MODES (`TR_BITBRAIN_MEM`):
## perRound (DEFAULT) — wipe the SBCs every round. The gate test measured this
## as the WINNING regime.
## retained — accumulate across the whole battle/enemy and wipe only
## on a target change / new battle. This is what the user
## asked for, and the gate test measured it as the WEAKEST
## regime: the idempotent SBC only ADDS, so it saturates.
## decay — retained PLUS a periodic partial wipe of the SBC
## tensors (TR_BITBRAIN_DECAY every N samples, a fraction
## TR_BITBRAIN_DECAY_FRAC of words zeroed). This is the one
## mechanism with a measured diagnosis behind it: the SBC
## saturates and a bounded/decaying memory should help.
## Phase 1 measured the actual lever. The gain sweep found that gains >= 1 are
## strictly worse at every band and that the optimal gain is BELOW 1.0 at long
## range (450+: ~0.25). A fractional gain leaves the Pearson lead *correlation*
## unchanged (correlation is invariant under positive scaling), so a smaller
## gain does not add information — it shrinks the magnitude of an uninformative
## Pattern lead toward the low-variance static (HeadOn) aim. The right output is
## therefore a multiplicative GAIN on Pattern's lead, not a class-based additive
## shift. See `docs/bitbrain_campaign.md` §Phase 1 for the measured curve.
##
## ── THE DESIGN ────────────────────────────────────────────────────────────────
## * BASE — the shipped Pattern gun's prediction (`guns/pattern_matcher`),
## reached through the TmHorizonGun observation ring.
## * OUTPUT — `aim = LOS + gain * (patternAim - LOS)`, i.e. Pattern's lead over
## the line of sight is multiplied by a learned `gain` (one of
## `BB_CAND`, so it may be BELOW 1.0 — the point).
## * LABEL — the same deferred-label path the old corrector used: at fire
## time we remember the base lead and the aim tolerance; `h =
## round(dist/speed)` ticks later `tmhObservedAt` returns the
## enemy's OBSERVED bearing from the firing position. `requiredLead
## = observedBearing - LOS` and `baseLead = baseBearing - LOS`, so a
## candidate gain scores a hit on this sample when
## `|gain*baseLead - requiredLead| <= tolerance`.
## * TRAIN — ONLINE / PREQUENTIAL per range band: for each candidate gain we
## count the fraction of resolved samples that would have been
## within the target's angular half-width (`atan(18/range)`, the
## SAME tolerance the offline ruler uses). The band's gain is the
## argmax hit rate. THIS is the key lesson of Phase 1: the
## least-squares gain and the hit-probability-optimal gain DIVERGE
## (Pattern's lead errors are bimodal), so the learner optimises the
## hit-probability proxy directly instead of mean squared error.
## * STATE — the range band (the ruler's 5 bands). Range is known causally at
## fire time, so a per-band gain table is shippable with no learning
## at all; BitBrain learns that table online. The correction is
## additionally gated to bands with range >= 300 px
## (`BB_GAIN_BAND_MIN`), where Phase 1 measured Pattern's lead to be
## uninformative. That gate is causal (range is known).
##
## The gain statistics are battle-scale: a round boundary wipes the observation
## ring and deferred labels but NOT the gain counts (a new round is not a new
## enemy). `resetLearning` wipes them on a new battle / target change; with
## `TR_BITBRAIN_MEM=decay` every `TR_BITBRAIN_DECAY` resolved samples decays the
## counts by `TR_BITBRAIN_DECAY_FRAC` toward the gain-1.0 column.
##
## ── WHAT IS STILL HERE ONLY FOR THE BOOT REPORT / GUARD TESTS ─────────────────
## The ADE+SBC network is GONE from the gun. The 53-bit TMH input, the class
## geometry (`bbCenterDeg`/`bbClassOf`), `TR_BITBRAIN_N`/`NADE`/`WARMUP`/`ADAPT`/
## `CALIB`/`SEED` and `TR_BITBRAIN_RANGE` are retained as resolved configuration
## so the boot report (`env_report.nim`) and the registration guard tests keep
## working unchanged; they no longer affect the gain learner. The generic
## `common_libs/bitbrain/` library is untouched and still tested by
## `test_bitbrain.nim`.
##
## DEFAULT OFF / PARITY: this gun is admitted ONLY when `TR_RACK_BITBRAIN` says so
## (default `off`) AND it never runs its network until `predict` is first called
## (`ensureInit`). With the shipped rack the live loop never calls `predict`, so
## no network is built, no RNG is touched and the shipped bot is unchanged.
## (default `off`). The shipped rack never calls `predict`, so `ensureInit` never
## runs and the shipped bot is byte-for-byte unchanged.
import std/[math, os, strutils, strformat, random]
import std/[math, os, strutils, strformat]
import gun_harness/gun_interface
import guns/tm_horizon
import guns/pattern_matcher
import bitbrain/bitbrain
const
## ── env knobs (all resolved once at gun construction) ─────────────────────
BB_MEM_ENV* = "TR_BITBRAIN_MEM" ## perRound|retained|decay
BB_N_ENV* = "TR_BITBRAIN_N" ## correction classes
BB_NADE_ENV* = "TR_BITBRAIN_NADE" ## ADEs per address decoder
BB_RANGE_ENV* = "TR_BITBRAIN_RANGE" ## class half-range, degrees
BB_N_ENV* = "TR_BITBRAIN_N" ## (legacy geometry; inert)
BB_NADE_ENV* = "TR_BITBRAIN_NADE" ## (legacy ADE count; inert)
BB_RANGE_ENV* = "TR_BITBRAIN_RANGE" ## (legacy class half-range; inert)
BB_LOG_ENV* = "TR_BITBRAIN_LOG" ## 1 = per-change [bb] log
BB_MIN_OBS_ENV* = "TR_BITBRAIN_MIN_OBS" ## resolved samples before correction
BB_WARMUP_ENV* = "TR_BITBRAIN_WARMUP" ## samples before percentile init
BB_ADAPT_ENV* = "TR_BITBRAIN_ADAPT" ## homeostasis interval (samples)
BB_CALIB_ENV* = "TR_BITBRAIN_CALIB" ## percentile recalibration interval
BB_MIN_OBS_ENV* = "TR_BITBRAIN_MIN_OBS" ## samples before a band is trusted
BB_WARMUP_ENV* = "TR_BITBRAIN_WARMUP" ## (legacy; inert)
BB_ADAPT_ENV* = "TR_BITBRAIN_ADAPT" ## (legacy; inert)
BB_CALIB_ENV* = "TR_BITBRAIN_CALIB" ## (legacy; inert)
BB_DECAY_ENV* = "TR_BITBRAIN_DECAY" ## decay interval (samples)
BB_DECAY_FRAC_ENV* = "TR_BITBRAIN_DECAY_FRAC" ## fraction of words zeroed per decay
BB_SEED_ENV* = "TR_BITBRAIN_SEED" ## deterministic AD/decay seed
BB_DECAY_FRAC_ENV* = "TR_BITBRAIN_DECAY_FRAC" ## per-decay count shrink
BB_SEED_ENV* = "TR_BITBRAIN_SEED" ## (legacy; inert)
BB_RESET_ON_TARGET_ENV* = "TR_BITBRAIN_RESET_ON_TARGET"
## ── fixed geometry ────────────────────────────────────────────────────────
BB_WIDTHS* = [6, 8, 10, 12] ## the paper's multi-width ADs
BB_TARGET_RATE* = 0.01 ## the paper's ~1 % firing target
BB_PENDING_CAP* = 512 ## deferred-label queue (>= 4 buckets x 50 ticks)
## ── the gain learner ──────────────────────────────────────────────────────
BB_NBANDS* = 5 ## the ruler's range bands
BB_NHB* = 4 ## horizon buckets (for the per-tick label dedupe)
BB_BAND_LO* = [0.0, 100.0, 200.0, 300.0, 450.0]
BB_BAND_HI* = [100.0, 200.0, 300.0, 450.0, 1.0e18]
## The candidate lead gains the band selector picks from. 0.0 == HeadOn (aim
## at the current position) and 1.0 == Pattern (use the full lead).
BB_CAND* = [0.0, 0.25, 0.50, 0.75, 1.0]
BB_NCAND* = 5
BB_BB_RADIUS* = 18.0 ## hit-detection radius in px (ruler tolerance)
## Apply the correction only from this band up (range >= BB_BAND_LO[3] = 300).
## [MEASURED] below 300 Pattern's lead is informative and shrinking it loses
## hits; see the header note.
BB_GAIN_BAND_MIN* = 3
## ── shipped defaults ──────────────────────────────────────────────────────
BB_N_DEF = 32
BB_NADE_DEF = 256
@@ -90,19 +118,21 @@ type
bmPerRound, bmRetained, bmDecay
BbPending = object
## One deferred training sample. `lits` is the exact literal vector the ADs
## saw at fire time; the label is resolved `horizon` ticks later.
## One deferred training sample. `lead` is Pattern's lead over LOS at fire
## time (radians) and `tol` the target's angular half-width then; the label
## is resolved `horizon` ticks later.
fireTick: int
horizon: int
band: int
selfX*, selfY: float
baseBearing: float
lits: array[TMH_NLITS, uint8]
lead: float
tolDeg: float
BitBrainGun* = object
tmh: TmHorizonGun
bb: BitBrain
initialized: bool
# ── resolved config ──────────────────────────────────────────────────────
# ── resolved config (kept in the boot report) ─────────────────────────────
nClasses*: int
maxDeg*: float
nAde*: int
@@ -116,36 +146,24 @@ type
decayFrac*: float
seed*: int64
resetOnTarget*: bool
# ── AD calibration state ─────────────────────────────────────────────────
rng: Rand
hist: seq[seq[int32]] ## per-AD raw-score histogram (bins 2w+1)
histTotal: int
sampleCount*: int
sinceAdapt: int
sinceCalib: int
# ── gain learner: hit counts per (range band x candidate gain) ────────────
bandHits*: array[BB_NBANDS, array[BB_NCAND, float64]]
bandN*: array[BB_NBANDS, float64]
trained*: int
sinceDecay: int
decays*: int
# ── scratch (avoid per-sample allocation) ────────────────────────────────
scratch: seq[seq[int32]]
counts: seq[int]
# ── deferred labels ──────────────────────────────────────────────────────
decays*: int
# ── readout / accounting ──────────────────────────────────────────────────
lastGain*: array[BB_NBANDS, float]
corrections*: int
lastLogKey: string
# ── deferred labels ───────────────────────────────────────────────────────
pending: array[BB_PENDING_CAP, BbPending]
pendingCount*: int
pendingDropped*: int
# ── per-tick caches ──────────────────────────────────────────────────────
# ── per-tick caches ───────────────────────────────────────────────────────
lastTick: int
lastEnqTick: int
lastEnqBucket: int
cachedBits: array[TMH_N_BASE, uint8]
cachedBitsTick: int
bitsValid: bool
# ── accounting / logging ─────────────────────────────────────────────────
trained*: int
lastBest: int
lastShift*: float
corrections*: int
lastLogKey: string
lastLogTick: int
observedTargetId*: int
# ── small pure helpers ───────────────────────────────────────────────────────
@@ -162,8 +180,8 @@ proc memModeName*(m: BitMemMode): string =
of bmDecay: "decay"
proc parseMemMode*(value: string): BitMemMode =
## Empty / unknown values fall back to the shipped `perRound` (the measured
## winning regime), so a typo cannot silently select another regime.
## Empty / unknown values fall back to the shipped `perRound`, so a typo
## cannot silently select another regime.
case value.strip().toLowerAscii()
of "retained", "retain", "accum", "accumulate": bmRetained
of "decay", "forget", "age": bmDecay
@@ -186,20 +204,32 @@ proc envBoolBB(name: string, default: bool): bool =
else: default
proc bbCenterDeg*(k, nClasses: int, maxDeg: float): float =
## Centre (degrees) of correction class `k` over ±maxDeg.
## Centre (degrees) of correction class `k` over ±maxDeg. Retained for the
## registration guard test and the boot report; inert for the gain learner.
let w = 2.0 * maxDeg / float(nClasses)
-maxDeg + (float(k) + 0.5) * w
proc bbClassOf*(errRad: float, nClasses: int, maxDeg: float): int =
## Bin a signed angular error (radians) into one of `nClasses` bins over
## [−maxDeg, +maxDeg] (the gate test's `binOf`).
## [−maxDeg, +maxDeg]. Retained for the registration guard test; inert.
let x = radToDeg(errRad)
var k = int((x + maxDeg) / (2.0 * maxDeg) * float(nClasses))
if k < 0: k = 0
if k >= nClasses: k = nClasses - 1
k
# ── construction / lazy network build ────────────────────────────────────────
proc bbBandOf*(range: float): int {.inline.} =
## Range band (the ruler's bands), known causally at fire time.
for b in 0 ..< BB_NBANDS:
if range >= BB_BAND_LO[b] and range < BB_BAND_HI[b]: return b
BB_NBANDS - 1
proc bbTolDeg*(range: float): float {.inline.} =
## The target's angular half-width at `range` — atan(18/range) — i.e. the exact
## tolerance the offline ruler uses for its hit-probability proxy.
radToDeg(arctan2(BB_BB_RADIUS, max(range, 1e-9)))
# ── construction / lazy init ─────────────────────────────────────────────────
proc initBitBrainGun*(): BitBrainGun =
result.nClasses = clamp(envIntBB(BB_N_ENV, BB_N_DEF), 2, 512)
@@ -219,148 +249,58 @@ proc initBitBrainGun*(): BitBrainGun =
result.lastEnqTick = -1
result.lastEnqBucket = -1
result.observedTargetId = -1
result.rng = initRand(result.seed + 991)
proc resetThresholdsHeuristic(g: var BitBrainGun) =
## Cold-start thresholds: a small multiple of the raw-score standard deviation
## puts every ADE near the paper's ~1 % firing rate from the FIRST ticks, so
## the SBCs see a useful (sparse) coincidence set immediately and inference
## never degenerates into an O(nAde^2) dense scan. The running-histogram
## percentile calibration replaces these once warmup has passed.
for a in 0..<g.bb.ades.len:
let w = float(g.bb.ades[a].width)
let thr = int32(round(2.33 * sqrt(w * 0.28)))
let scaled = int32(g.bb.ades[a].scale) * thr
for e in 0..<g.bb.ades[a].nAde:
g.bb.ades[a].thresholds[e] = scaled
for b in 0 ..< BB_NBANDS: result.lastGain[b] = 1.0
proc ensureInit*(g: var BitBrainGun) =
## Build the AD/SBC network on first use. Never runs on the shipped default
## path (the rack does not admit BitBrain), so the default bot is untouched.
## Build the observation ring on first use. No network, no global-RNG use, so
## the shipped default path is untouched and construction stays cheap.
if g.initialized: return
g.initialized = true
g.tmh = initTmHorizonGun()
var rng = initRand(g.seed)
var ades: seq[AddressDecoder]
for w in BB_WIDTHS:
ades.add initRandomAddressDecoder(g.nAde, w, TMH_N_BITS, rng,
scale = DefaultScale, center = 0,
threshold = 0'i32)
g.bb = initBitBrain(ades, crossPairs(ades.len), g.nClasses)
g.hist = newSeq[seq[int32]](ades.len)
g.scratch = newSeq[seq[int32]](ades.len)
for a in 0..<ades.len:
g.hist[a] = newSeq[int32](2 * ades[a].width + 1)
g.counts = newSeq[int](g.nClasses)
g.resetThresholdsHeuristic()
# ── AD calibration (online percentile init + library homeostasis) ────────────
# ── the gain learner ─────────────────────────────────────────────────────────
proc calibrate(g: var BitBrainGun) =
## Set every ADE's threshold to the raw score whose `count >= t` is CLOSEST to
## `BB_TARGET_RATE * total` — the gate test's percentile init, run online over
## the running histogram. This is what pins the realised firing rate near 1 %.
if g.histTotal <= 0: return
let target = BB_TARGET_RATE * float(g.histTotal)
for a in 0..<g.bb.ades.len:
let w = g.bb.ades[a].width
let sc = g.bb.ades[a].scale
var cum = 0
var bestRaw = w
var bestDiff = Inf
for raw in countdown(w, -w):
cum += int(g.hist[a][raw + w])
let d = abs(float(cum) - target)
if d < bestDiff:
bestDiff = d
bestRaw = raw
let t = int32(bestRaw) * int32(sc)
for e in 0..<g.bb.ades[a].nAde:
g.bb.ades[a].thresholds[e] = t
proc bbAccumulate(g: var BitBrainGun, leadDeg, reqDeg, tolDeg: float, band: int) =
## Score every candidate gain on this resolved sample: a candidate "hits" when
## it would have put the aim within the target's angular half-width.
for ci in 0 ..< BB_NCAND:
if abs(BB_CAND[ci] * leadDeg - reqDeg) <= tolDeg:
g.bandHits[band][ci] += 1.0
g.bandN[band] += 1.0
inc g.trained
proc afterSample(g: var BitBrainGun) =
## Post-sample calibration/homeostasis schedule.
inc g.sampleCount
if g.sampleCount == g.warmupN:
g.calibrate()
g.sinceAdapt = 0
g.sinceCalib = 0
elif g.sampleCount > g.warmupN:
inc g.sinceAdapt
inc g.sinceCalib
if g.sinceAdapt >= g.adaptEvery:
for a in 0..<g.bb.ades.len:
g.bb.ades[a].adaptThresholds(g.adaptEvery, BB_TARGET_RATE, 1)
g.sinceAdapt = 0
if g.sinceCalib >= g.calibEvery:
g.calibrate()
g.sinceCalib = 0
# ── one AD pass: firing counts + histogram + inference ───────────────────────
proc bbObserve(g: var BitBrainGun, lits: array[TMH_NLITS, uint8]) =
## Drive every ADE: update its firing accumulator and the score histogram,
## collect the active list, then infer the class counts into `g.counts`.
for a in 0..<g.bb.ades.len:
let w = g.bb.ades[a].width
let sc = g.bb.ades[a].scale
g.scratch[a].setLen(0)
for e in 0..<g.bb.ades[a].nAde:
var raw = 0
let off = e * w
for j in 0..<w:
let c = g.bb.ades[a].codes[off + j]
let idx = if c > 0'i32: int(c) - 1 else: int(-c) - 1
let pol = if c > 0'i32: 1 else: -1
raw += pol * int(lits[idx])
inc g.hist[a][raw + w]
if raw * sc >= int(g.bb.ades[a].thresholds[e]):
g.scratch[a].add int32(e)
inc g.bb.ades[a].fireCounts[e]
inc g.histTotal
for k in 0..<g.counts.len: g.counts[k] = 0
for sl in 0..<g.bb.sbcs.len:
let spec = g.bb.specs[sl]
g.bb.sbcs[sl].infer(g.scratch[spec.row], g.scratch[spec.col], g.counts)
proc bbLearn(g: var BitBrainGun, lits: array[TMH_NLITS, uint8], cls: int) =
## Recompute the active lists for a resolved sample and set its class bits in
## every SBC (idempotent, so a repeat is a no-op).
for a in 0..<g.bb.ades.len:
let w = g.bb.ades[a].width
let sc = g.bb.ades[a].scale
g.scratch[a].setLen(0)
for e in 0..<g.bb.ades[a].nAde:
var raw = 0
let off = e * w
for j in 0..<w:
let c = g.bb.ades[a].codes[off + j]
let idx = if c > 0'i32: int(c) - 1 else: int(-c) - 1
let pol = if c > 0'i32: 1 else: -1
raw += pol * int(lits[idx])
if raw * sc >= int(g.bb.ades[a].thresholds[e]):
g.scratch[a].add int32(e)
for sl in 0..<g.bb.sbcs.len:
let spec = g.bb.specs[sl]
discard g.bb.sbcs[sl].learn(g.scratch[spec.row], g.scratch[spec.col], cls)
proc applyDecay(g: var BitBrainGun) =
## Age the SBC tensors: zero a fraction of their 32-bit words. This is the
## bounded-memory mechanism the gate test's diagnosis called for (the
## idempotent SBC otherwise only ADDS and saturates with stale class bits).
let cut = int(g.decayFrac * 1000.0)
if cut <= 0: return
for sl in 0..<g.bb.sbcs.len:
for wi in 0..<g.bb.sbcs[sl].bits.len:
if g.rng.rand(999) < cut:
g.bb.sbcs[sl].bits[wi] = 0'u32
proc bbApplyDecay(g: var BitBrainGun) =
## Forgetting for `TR_BITBRAIN_MEM=decay`: shrink the hit counts and, more
## strongly, pull them toward the gain-1.0 column so stale evidence ages out.
let f = 1.0 - g.decayFrac
if f >= 1.0: return
for b in 0 ..< BB_NBANDS:
for ci in 0 ..< BB_NCAND:
g.bandHits[b][ci] *= f
g.bandN[b] *= f
inc g.decays
proc bbGain(g: BitBrainGun, band: int): float =
## The band's gain is the candidate with the highest observed hit rate.
## Ties keep the SMALLER candidate (the scan is ascending), which is the
## conservative choice for the long-range regime this corrector targets.
## Returns 1.0 (Pattern) below the range gate or when the band is cold.
if band < BB_GAIN_BAND_MIN: return 1.0
if g.bandN[band] < float(g.minObs): return 1.0
var best = 4 # gain 1.0
var bestRate = -1.0
for ci in 0 ..< BB_NCAND:
let rate = g.bandHits[band][ci] / g.bandN[band]
if rate > bestRate:
bestRate = rate
best = ci
BB_CAND[best]
# ── deferred-label resolution (prequential learning) ─────────────────────────
proc resolvePending(g: var BitBrainGun, state: WorldState) =
var w = 0
for i in 0..<g.pendingCount:
for i in 0 ..< g.pendingCount:
let p = g.pending[i]
let due = p.fireTick + p.horizon
if due > state.tick:
@@ -370,12 +310,11 @@ proc resolvePending(g: var BitBrainGun, state: WorldState) =
let obs = tmhObservedAt(g.tmh, state.tick, p.selfX, p.selfY)
if obs.ok and (state.tick - obs.lastSeenTick) <= TMH_STALE_MAX:
let err = wrapRadBB(obs.bearing - p.baseBearing)
let cls = bbClassOf(err, g.nClasses, g.maxDeg)
g.bbLearn(p.lits, cls)
inc g.trained
let reqLead = wrapRadBB(err + p.lead)
g.bbAccumulate(radToDeg(p.lead), radToDeg(reqLead), p.tolDeg, p.band)
inc g.sinceDecay
if g.memMode == bmDecay and g.sinceDecay >= g.decayEvery:
g.applyDecay()
g.bbApplyDecay()
g.sinceDecay = 0
else:
inc g.pendingDropped
@@ -385,62 +324,49 @@ proc resolvePending(g: var BitBrainGun, state: WorldState) =
# ── logging ──────────────────────────────────────────────────────────────────
proc bbLog(g: var BitBrainGun, state: WorldState, h, bucket, total, best: int) =
## ONE change-gated `[bb]` line (behind TR_BITBRAIN_LOG=1) so the user tailing
## the GUI log sees what the corrector is thinking, not one line per tick.
proc bbLog(g: var BitBrainGun, state: WorldState, band: int, gain: float) =
## ONE change-gated `[bb]` line (behind TR_BITBRAIN_LOG=1) so a user tailing
## the GUI log sees the gain the corrector is applying.
if not g.logEnabled: return
let shift = bbCenterDeg(best, g.nClasses, g.maxDeg)
let key = fmt"{best}|{shift:.1f}"
let key = fmt"{gain:.2f}|{band}"
if key == g.lastLogKey: return
if state.tick == g.lastLogTick: return
g.lastLogKey = key
g.lastLogTick = state.tick
var nz = 0
for k in 0..<g.counts.len:
if g.counts[k] > 0: inc nz
echo fmt"[bb] t={state.tick} h={h} bucket={bucket} cls={best}/{g.nClasses} " &
fmt"shift={shift:+.1f}deg cnt={g.counts[best]}/{total} nz={nz} " &
fmt"trained={g.trained} samples={g.sampleCount} pend={g.pendingCount} " &
fmt"mode={memModeName(g.memMode)} warm={(if g.trained >= g.minObs: 1 else: 0)}"
var rate = 0.0
for ci in 0 ..< BB_NCAND:
if abs(BB_CAND[ci] - gain) < 1e-9: rate = g.bandHits[band][ci] / max(1.0, g.bandN[band])
echo fmt"[bb] t={state.tick} band={BB_BAND_LO[band]:.0f}+ gain={gain:.2f} " &
fmt"rate={rate:.3f} n={g.bandN[band]:.0f} trained={g.trained} " &
fmt"pend={g.pendingCount} dropped={g.pendingDropped} mode={memModeName(g.memMode)}"
# ── reset hooks (mirroring TmHorizonGun) ─────────────────────────────────────
proc resetRound(g: var BitBrainGun) =
## PER-ROUND wipe. Always clear the observation ring, deferred labels and
## per-tick caches (the bots teleport between rounds). In `perRound` mode the
## SBCs are wiped too; `retained`/`decay` keep them across the round.
## PER-ROUND reset: observation ring, deferred labels and per-tick caches (the
## bots teleport between rounds). The gain counts are deliberately KEPT — they
## are battle-scale and a new round is not a new enemy.
g.tmh.resetRoundState()
g.pendingCount = 0
g.lastTick = -1
g.lastEnqTick = -1
g.lastEnqBucket = -1
g.bitsValid = false
g.lastLogKey = ""
g.lastLogTick = -1
if g.memMode == bmPerRound:
g.bb.resetLearning()
g.trained = 0
proc resetRoundState*(g: var BitBrainGun) =
if not g.initialized: return
g.resetRound()
proc resetLearning*(g: var BitBrainGun, reason = "") =
## PER-BATTLE / PER-ENEMY wipe: SBCs, AD thresholds, histograms and counters.
## PER-BATTLE / PER-ENEMY wipe: gain counts, counters and the round state.
if not g.initialized: return
g.bb.resetLearning()
g.resetThresholdsHeuristic()
for a in 0..<g.hist.len:
for i in 0..<g.hist[a].len: g.hist[a][i] = 0
g.histTotal = 0
g.sampleCount = 0
g.sinceAdapt = 0
g.sinceCalib = 0
for b in 0 ..< BB_NBANDS:
for ci in 0 ..< BB_NCAND: g.bandHits[b][ci] = 0.0
g.bandN[b] = 0.0
g.lastGain[b] = 1.0
g.trained = 0
g.sinceDecay = 0
g.decays = 0
g.trained = 0
g.corrections = 0
g.observedTargetId = -1
g.rng = initRand(g.seed + 991)
g.resetRound()
if reason.len > 0 and g.logEnabled:
echo fmt"[bb-reset] reason={reason}"
@@ -462,8 +388,8 @@ proc targetChanged*(g: var BitBrainGun, enemyId: int): bool =
proc isWarmedUp*(g: BitBrainGun): bool {.inline.} = true
proc networkBytes*(g: BitBrainGun): int =
## Bytes held by the AD/SBC network (0 until the network is built).
if g.initialized: g.bb.memoryBytes else: 0
## No neural network is held any more; kept for the boot report / guard test.
0
proc predict*(g: var BitBrainGun, state: WorldState,
bulletSpeed: float): GunPrediction =
@@ -477,58 +403,43 @@ proc predict*(g: var BitBrainGun, state: WorldState,
tmhUpdateHistory(g.tmh, state)
g.resolvePending(state)
g.lastTick = state.tick
g.bitsValid = false
# The base prediction is Pattern; BitBrain only corrects its bearing.
# The base prediction is Pattern; BitBrain only scales its lead over LOS.
let base = g.tmh.pattern.predict(state, bulletSpeed)
if bulletSpeed <= 0.0: return base
let dist = hypot(state.enemyX - state.selfX, state.enemyY - state.selfY)
let h = tmhHorizonFor(dist, bulletSpeed)
let bucket = tmhHorizonBucket(h)
let hb = tmhHorizonBucket(h)
let band = bbBandOf(dist)
if not g.bitsValid or g.cachedBitsTick != state.tick:
g.cachedBits = tmhBaseBits(g.tmh, state)
g.cachedBitsTick = state.tick
g.bitsValid = true
let lits = tmhLits(g.cachedBits, bucket)
let los = arctan2(state.enemyY - state.selfY, state.enemyX - state.selfX)
let baseBearing = arctan2(base.y - state.selfY, base.x - state.selfX)
let lead = wrapRadBB(baseBearing - los)
# Observe this input (AD pass + inference) and advance the calibration clock.
g.bbObserve(lits)
g.afterSample()
# Enqueue one deferred sample per (tick, bucket): predict runs once per power
# bin, so all four horizons contribute evidence.
if g.lastEnqTick != state.tick or g.lastEnqBucket != bucket:
# Enqueue one deferred sample per (tick, horizon bucket): `predict` runs once
# per power bin, so all four horizons contribute evidence.
if g.lastEnqTick != state.tick or g.lastEnqBucket != hb:
if g.pendingCount < BB_PENDING_CAP:
g.pending[g.pendingCount] = BbPending(
fireTick: state.tick, horizon: h,
fireTick: state.tick, horizon: h, band: band,
selfX: state.selfX, selfY: state.selfY,
baseBearing: arctan2(base.y - state.selfY, base.x - state.selfX),
lits: lits)
baseBearing: baseBearing, lead: lead, tolDeg: bbTolDeg(dist))
inc g.pendingCount
else:
inc g.pendingDropped
g.lastEnqTick = state.tick
g.lastEnqBucket = bucket
g.lastEnqBucket = hb
# Readout: argmax class centre, zero correction with no evidence / cold.
var shiftDeg = 0.0
if g.trained >= g.minObs:
var total = 0
for k in 0..<g.counts.len: total += g.counts[k]
if total > 0:
var best = 0
for k in 1..<g.counts.len:
if g.counts[k] > g.counts[best]: best = k
shiftDeg = bbCenterDeg(best, g.nClasses, g.maxDeg)
g.lastBest = best
g.lastShift = shiftDeg
inc g.corrections
g.bbLog(state, h, bucket, total, best)
if shiftDeg == 0.0: return base
tmhApplyShift(state.selfX, state.selfY, base.x, base.y, shiftDeg)
# Readout: a fractional gain may be BELOW 1.0. When cold / gated out the
# learner returns 1.0 and the base prediction is returned unchanged.
let gain = g.bbGain(band)
g.lastGain[band] = gain
if abs(gain - 1.0) < 1e-9: return base
inc g.corrections
g.bbLog(state, band, gain)
tmhApplyShift(state.selfX, state.selfY, base.x, base.y,
radToDeg((gain - 1.0) * lead))
proc onResult*(g: var BitBrainGun, e: FeedbackEvent) =
## Labels come from our own observation ring, not from virtual-bullet
@@ -6,8 +6,8 @@ ruler : continuous (physically exact)
runs : 70
recorded ticks: 899607
tick x bin : 3598428
wall time : 406.88s (0.1131 ms per tick-bin)
per-arm speed : 0.1131 s per 1000 tick-bins per arm
wall time : 271.23s (0.0754 ms per tick-bin)
per-arm speed : 0.0754 s per 1000 tick-bins per arm
NOTE: offline OPEN-LOOP prediction quality only. No win/damage/survival claim.
@@ -18,7 +18,7 @@ VALIDATION -- the ruler must pass ALL of these before any number below is trus
ruler=continuous hits n=5480 mean|err|= 1.360 deg / 10.5 px | misses n=48304 mean|err|= 16.597 deg / 140.3 px | separation 12.20x deg / 13.34x px -> OK
ruler=integer hits n=5480 mean|err|= 1.478 deg / 11.4 px | misses n=48304 mean|err|= 16.724 deg / 141.3 px | separation 11.32x deg / 12.43x px -> OK
2. perfect-oracle gun max |err| over all tick-bins = 0.000000 deg -> OK
3. HeadOn (static LOS) mean|err| = 13.217 deg vs Pattern 16.609 / TMHorizon 16.627 / BitBrain 16.635
3. HeadOn (static LOS) mean|err| = 13.217 deg vs Pattern 16.609 / TMHorizon 16.627 / BitBrain 13.627
-> UNEXPECTED: a predictive gun is worse than static LOS
NaiveLinear mean|err| = 22.086 deg (over-leads; see the lead-gain sweep for why a larger
lead *response* does not mean a smaller angular error)
@@ -85,23 +85,47 @@ TMHorizon 200-300 74215 16.644 21.142 1.091 0.1786 84.4
TMHorizon 300-450 1119777 17.572 21.878 0.691 0.0996 84.01
TMHorizon 450+ 2311323 16.199 20.025 -0.687 0.0757 81.19
(TMHorizon: 63782 tick-bins had no valid interception)
BitBrain 0-100 4423 11.118 15.268 -1.507 0.6810 64.72
BitBrain 100-200 24908 15.107 19.782 1.289 0.3241 76.63
BitBrain 200-300 74215 16.838 21.426 1.341 0.1747 105.04
BitBrain 300-450 1119777 17.575 21.894 0.933 0.1025 100.98
BitBrain 450+ 2311323 16.200 20.032 -0.647 0.0767 96.38
BitBrain 0-100 4423 10.555 14.796 -0.404 0.6993 64.72
BitBrain 100-200 24908 14.745 19.510 1.276 0.3418 76.63
BitBrain 200-300 74215 16.610 21.174 1.350 0.1850 81.45
BitBrain 300-450 1119777 15.518 19.170 0.773 0.1073 85.65
BitBrain 450+ 2311323 12.608 15.431 -0.351 0.0957 79.92
(BitBrain: 63782 tick-bins had no valid interception)
PatternGain0.25 0-100 4423 16.984 19.724 -1.060 0.3914 46.00
PatternGain0.25 100-200 24908 17.791 20.374 -0.027 0.1661 51.44
PatternGain0.25 200-300 74215 15.950 18.761 0.698 0.1276 53.22
PatternGain0.25 300-450 1119777 14.131 16.972 0.755 0.1025 52.64
PatternGain0.25 450+ 2311323 12.261 14.888 -0.358 0.0931 51.24
(PatternGain0.25: 63782 tick-bins had no valid interception)
PatternGain0.50 0-100 4423 14.583 16.877 -0.841 0.4689 50.78
PatternGain0.50 100-200 24908 16.103 18.679 0.407 0.1673 58.27
PatternGain0.50 200-300 74215 15.316 18.261 0.915 0.1366 62.44
PatternGain0.50 300-450 1119777 14.491 17.566 0.819 0.1003 61.45
PatternGain0.50 450+ 2311323 12.938 15.799 -0.453 0.0883 59.63
(PatternGain0.50: 63782 tick-bins had no valid interception)
PatternGain0.75 0-100 4423 12.356 15.103 -0.622 0.6396 57.28
PatternGain0.75 100-200 24908 14.965 18.369 0.841 0.2337 66.41
PatternGain0.75 200-300 74215 15.528 19.120 1.133 0.1389 71.95
PatternGain0.75 300-450 1119777 15.639 19.275 0.884 0.0952 72.71
PatternGain0.75 450+ 2311323 14.280 17.588 -0.548 0.0813 68.72
(PatternGain0.75: 63782 tick-bins had no valid interception)
PatternBandGain 0-100 4423 10.555 14.796 -0.404 0.6993 64.72
PatternBandGain 100-200 24908 14.745 19.510 1.276 0.3418 76.63
PatternBandGain 200-300 74215 16.610 21.174 1.350 0.1850 81.45
PatternBandGain 300-450 1119777 14.607 17.606 0.690 0.1049 46.38
PatternBandGain 450+ 2311323 12.326 15.017 -0.263 0.0984 45.49
(PatternBandGain: 63782 tick-bins had no valid interception)
========================================================================================================================
HEADROOM -- the direct answer: how far each arm is from the oracle ceiling, per band
========================================================================================================================
band Pattern n Pattern|err| Pattern hpx Oracle hpx headroom pp naive hpx TMHoriz hpx BitBrain hpx
---------------------------------------------------------------------------------------------------------------
0-100 4423 10.555 0.6993 1.0000 0.3007 0.6505 0.6955 0.6810
100-200 24908 14.745 0.3418 1.0000 0.6582 0.3879 0.3418 0.3241
200-300 74215 16.610 0.1850 1.0000 0.8150 0.1924 0.1786 0.1747
300-450 1119777 17.531 0.1036 1.0000 0.8964 0.0995 0.0996 0.1025
450+ 2311323 16.193 0.0767 1.0000 0.9233 0.0537 0.0757 0.0767
0-100 4423 10.555 0.6993 1.0000 0.3007 0.6505 0.6955 0.6993
100-200 24908 14.745 0.3418 1.0000 0.6582 0.3879 0.3418 0.3418
200-300 74215 16.610 0.1850 1.0000 0.8150 0.1924 0.1786 0.1850
300-450 1119777 17.531 0.1036 1.0000 0.8964 0.0995 0.0996 0.1073
450+ 2311323 16.193 0.0767 1.0000 0.9233 0.0537 0.0757 0.0957
hitProxy = fraction of tick-bins aimed within atan(18/range) of the true interception point.
headroom pp = oracle hitProxy - Pattern hitProxy = the absolute hit-probability points available
@@ -129,6 +153,51 @@ band gain=1.0 gain=1.5 gain=2.0 gain=3.0 best-gain
300-450 17.531 23.279 29.955 44.274 1.0 (17.531)
450+ 16.193 21.245 26.997 39.271 1.0 (16.193)
========================================================================================================================
PHASE 1 — THE MISSING GAIN SWEEP: Pattern lead x gain in [0.00, 1.00] (0 = HeadOn, 1 = Pattern)
========================================================================================================================
Format per cell: mean|err| deg [hitProxy]. hitProxy is the objective. gain 0.0 is HeadOn,
gain 1.0 is Pattern. The per-band gain table IS causal to APPLY (range is known at fire time,
so a per-band lookup needs no learning); its ESTIMATION from these same runs is in-sample.
band |req| deg g=0.00 [hpx] g=0.25 [hpx] g=0.50 [hpx] g=0.75 [hpx] g=1.00 [hpx] bestHpx dHpx bestErr
--------------------------------------------------------------------------------------------------------------------
0-100 19.619 19.619 [0.342] 16.984 [0.391] 14.583 [0.469] 12.356 [0.640] 10.555 [0.699] 1.00 +0.0000 1.00
100-200 19.982 19.982 [0.172] 17.791 [0.166] 16.103 [0.167] 14.965 [0.234] 14.745 [0.342] 1.00 +0.0000 1.00
200-300 17.341 17.341 [0.133] 15.950 [0.128] 15.316 [0.137] 15.528 [0.139] 16.610 [0.185] 1.00 +0.0000 0.50
300-450 14.607 14.607 [0.105] 14.131 [0.102] 14.491 [0.100] 15.639 [0.095] 17.531 [0.104] 0.00 +0.0013 0.25
450+ 12.326 12.326 [0.098] 12.261 [0.093] 12.938 [0.088] 14.280 [0.081] 16.193 [0.077] 0.00 +0.0216 0.25
OPTIMAL GAIN CURVE (hitProxy-argmax per band) and its implied hit-probability gain vs Pattern:
0-100 best gain 1.00 hitProxy 0.6993 vs Pattern 0.6993 => +0.0000 pp
100-200 best gain 1.00 hitProxy 0.3418 vs Pattern 0.3418 => +0.0000 pp
200-300 best gain 1.00 hitProxy 0.1850 vs Pattern 0.1850 => +0.0000 pp
300-450 best gain 0.00 hitProxy 0.1049 vs Pattern 0.1036 => +0.0013 pp
450+ best gain 0.00 hitProxy 0.0984 vs Pattern 0.0767 => +0.0216 pp
gain = [ 0-100->1.00 100-200->1.00 200-300->1.00 300-450->0.00 450+->0.00 ]
DIRECT COMPARISON — Pattern vs the FIXED causal per-band rule [1,1,1,0.00,0.00] vs BitBrain (learned online):
band Pattern hpx fixed-band hpx BitBrain hpx fixed-Pat pp BB-Pat pp
--------------------------------------------------------------------------------
0-100 0.6993 0.6993 0.6993 +0.0000 +0.0000
100-200 0.3418 0.3418 0.3418 +0.0000 +0.0000
200-300 0.1850 0.1850 0.1850 +0.0000 +0.0000
300-450 0.1036 0.1049 0.1073 +0.0013 +0.0037
450+ 0.0767 0.0984 0.0957 +0.0216 +0.0190
fixed-band hpx = the [1,1,1,0,0] table applied causally; it was selected in-sample.
BitBrain is learned online from labels inside each run (cold start at gain 1.0).
LEAD CORRELATION PER GAIN (Pearson of applied lead with required lead). Pearson is invariant
under positive scaling, so every g>0 column must be IDENTICAL to Pattern; g=0 has no lead and
therefore no correlation. If they match, a shrinking gain does NOT add lead information — it
only shrinks the magnitude of an uninformative signal (the gain-sweep mechanism).
band corr(g) g=0.25 g=0.50 g=0.75 g=1.00
0-100 0.774 0.774 0.774 0.774
100-200 0.612 0.612 0.612 0.612
200-300 0.457 0.457 0.457 0.457
300-450 0.266 0.266 0.266 0.266
450+ 0.165 0.165 0.165 0.165
========================================================================================================================
LEAD INFORMATIVENESS -- capture slope (regression of applied lead on required lead) and lead correlation
========================================================================================================================
@@ -139,8 +208,8 @@ hits less' tension.
band HO |err| HO |req| Pat|req| Pat cap Pat corr Lin cap Lin corr TMH cap TMH corr BB cap BB corr
-----------------------------------------------------------------------------------------------------------------------
0-100 19.619 19.619 19.619 0.649 0.774 0.573 0.362 0.668 0.776 0.701 0.768
100-200 19.982 19.982 19.982 0.553 0.612 0.592 0.523 0.560 0.614 0.570 0.611
200-300 17.341 17.341 17.341 0.449 0.457 0.519 0.429 0.452 0.460 0.459 0.457
300-450 14.607 14.607 14.607 0.278 0.266 0.476 0.324 0.273 0.262 0.280 0.266
450+ 12.326 12.326 12.326 0.175 0.165 0.310 0.178 0.174 0.164 0.175 0.165
0-100 19.619 19.619 19.619 0.649 0.774 0.573 0.362 0.668 0.776 0.649 0.774
100-200 19.982 19.982 19.982 0.553 0.612 0.592 0.523 0.560 0.614 0.553 0.612
200-300 17.341 17.341 17.341 0.449 0.457 0.519 0.429 0.452 0.460 0.449 0.457
300-450 14.607 14.607 14.607 0.278 0.266 0.476 0.324 0.273 0.262 0.148 0.213
450+ 12.326 12.326 12.326 0.175 0.165 0.310 0.178 0.174 0.164 0.036 0.101
+98 -1
View File
@@ -30,8 +30,29 @@ const
A_NAIVE* = 7
A_TMH* = 8
A_BB* = 9
# ── Phase 1: the MISSING gain sweep. Gains >= 1 were measured worse at every
# band in Phase 0; the unexplored region is gain < 1. gain 0.0 is HeadOn
# (A_HEADON) and gain 1.0 is Pattern (A_PATTERN), so only 0.25/0.50/0.75 are
# new arms. Their `leadCorr` is IDENTICAL to Pattern's by construction (Pearson
# correlation is invariant under positive scaling) — printed only to prove it.
A_G025* = 10
A_G050* = 11
A_G075* = 12
# Phase 1 fixed causal per-band gain rule: the hitProxy-argmax curve measured by
# the sub-unity sweep ([1,1,1,0,0] == Pattern below 300 px, HeadOn above). This
# is the rule BitBrain must match; it needs no learning (range is known at fire
# time). The table was selected in-sample from this corpus.
A_BAND* = 13
BandGainTable* = [1.0, 1.0, 1.0, 0.0, 0.0]
ArmNames* = ["Oracle", "OracleQuant", "HeadOn", "Pattern", "PatternGain1.5",
"PatternGain2.0", "PatternGain3.0", "NaiveLinear", "TMHorizon", "BitBrain"]
"PatternGain2.0", "PatternGain3.0", "NaiveLinear", "TMHorizon", "BitBrain",
"PatternGain0.25", "PatternGain0.50", "PatternGain0.75", "PatternBandGain"]
const
## The five sub-unity gain arms, in increasing order, resolved to arm indices.
## gain 0.0 == HeadOn, gain 1.0 == Pattern.
GainArmIdx* = [A_HEADON, A_G025, A_G050, A_G075, A_PATTERN]
GainValues* = [0.0, 0.25, 0.50, 0.75, 1.0]
# ── the naive-linear control (job-95's LIN_M = 4 extrapolation) ──────────────
#
@@ -133,6 +154,13 @@ proc runRound(ctx: var Ctx, arms: var seq[ArmAcc], r: int) =
arms[A_G15].record(rng, wrap180(1.5 * plead - targetLead), 1.5 * plead, targetLead)
arms[A_G20].record(rng, wrap180(2.0 * plead - targetLead), 2.0 * plead, targetLead)
arms[A_G30].record(rng, wrap180(3.0 * plead - targetLead), 3.0 * plead, targetLead)
# sub-unity gains (Phase 1) — the region Phase 0 never covered
arms[A_G025].record(rng, wrap180(0.25 * plead - targetLead), 0.25 * plead, targetLead)
arms[A_G050].record(rng, wrap180(0.50 * plead - targetLead), 0.50 * plead, targetLead)
arms[A_G075].record(rng, wrap180(0.75 * plead - targetLead), 0.75 * plead, targetLead)
# the fixed causal per-band rule (Phase 1 hitProxy-argmax curve)
let bg = BandGainTable[bandOf(rng)]
arms[A_BAND].record(rng, wrap180(bg * plead - targetLead), bg * plead, targetLead)
# naive linear
let np = predict(ctx.naive, ctx.st, speed)
let nl = wrap180(bearingDeg(ox, oy, np.x, np.y) - los)
@@ -332,6 +360,75 @@ proc main() =
if g30 < bestV: bestV = g30; best = "3.0"
echo fmt"{BandLabels[b]:<9} {fmt3(g1):>10} {fmt3(g15):>10} {fmt3(g20):>10} {fmt3(g30):>10} {best} ({fmt3(bestV)})"
echo ""
echo "=".repeat(120)
echo "PHASE 1 — THE MISSING GAIN SWEEP: Pattern lead x gain in [0.00, 1.00] (0 = HeadOn, 1 = Pattern)"
echo "=".repeat(120)
echo "Format per cell: mean|err| deg [hitProxy]. hitProxy is the objective. gain 0.0 is HeadOn,"
echo "gain 1.0 is Pattern. The per-band gain table IS causal to APPLY (range is known at fire time,"
echo "so a per-band lookup needs no learning); its ESTIMATION from these same runs is in-sample."
echo ""
var hdrg = "band |req| deg"
for gi in 0 ..< GainValues.len: hdrg.add fmt" g={GainValues[gi]:.2f} [hpx]"
hdrg.add " bestHpx dHpx bestErr"
echo hdrg
echo "-".repeat(hdrg.len)
for b in 0 ..< NBands:
var line = fmt"{BandLabels[b]:<9} {fmt3(meanAbsReq(arms[A_PATTERN].bands[b])):>9}"
var bestHi = 0
var bestHp = -1.0
var bestEi = 0
var bestEr = Inf
for gi in 0 ..< GainValues.len:
let s = arms[GainArmIdx[gi]].bands[b]
let e = meanAbs(s)
let hp = s.hitProxy
line.add fmt"{fmt3(e):>7} [{fmt3(hp)}] "
if hp > bestHp: bestHp = hp; bestHi = gi
if e < bestEr: bestEr = e; bestEi = gi
let patHp = arms[A_PATTERN].bands[b].hitProxy
line.add fmt" {GainValues[bestHi]:.2f} {bestHp-patHp:+.4f} {GainValues[bestEi]:.2f}"
echo line
echo ""
echo "OPTIMAL GAIN CURVE (hitProxy-argmax per band) and its implied hit-probability gain vs Pattern:"
var curve = " gain = [ "
for b in 0 ..< NBands:
var bestHi = 0
var bestHp = -1.0
for gi in 0 ..< GainValues.len:
let hp = arms[GainArmIdx[gi]].bands[b].hitProxy
if hp > bestHp: bestHp = hp; bestHi = gi
curve.add fmt"{BandLabels[b]}->{GainValues[bestHi]:.2f} "
let patHp = arms[A_PATTERN].bands[b].hitProxy
echo fmt" {BandLabels[b]:<9} best gain {GainValues[bestHi]:.2f} hitProxy {bestHp:.4f} vs Pattern {patHp:.4f} => {bestHp-patHp:+.4f} pp"
echo curve & "]"
echo ""
echo "DIRECT COMPARISON — Pattern vs the FIXED causal per-band rule [1,1,1,0.00,0.00] vs BitBrain (learned online):"
let hdrd = "band Pattern hpx fixed-band hpx BitBrain hpx fixed-Pat pp BB-Pat pp"
echo hdrd
echo "-".repeat(hdrd.len)
for b in 0 ..< NBands:
let patHp = arms[A_PATTERN].bands[b].hitProxy
let fixHp = arms[A_BAND].bands[b].hitProxy
let bbHp = arms[A_BB].bands[b].hitProxy
echo fmt"{BandLabels[b]:<9} {patHp:>11.4f} {fixHp:>16.4f} {bbHp:>14.4f} {fixHp-patHp:>+14.4f} {bbHp-patHp:>+10.4f}"
echo "fixed-band hpx = the [1,1,1,0,0] table applied causally; it was selected in-sample."
echo "BitBrain is learned online from labels inside each run (cold start at gain 1.0)."
echo ""
echo "LEAD CORRELATION PER GAIN (Pearson of applied lead with required lead). Pearson is invariant"
echo "under positive scaling, so every g>0 column must be IDENTICAL to Pattern; g=0 has no lead and"
echo "therefore no correlation. If they match, a shrinking gain does NOT add lead information — it"
echo "only shrinks the magnitude of an uninformative signal (the gain-sweep mechanism)."
let hdrc = "band " & " corr(g) "
var hdrc2 = hdrc
for gi in 1 ..< GainValues.len: hdrc2.add fmt" g={GainValues[gi]:.2f}"
echo hdrc2
for b in 0 ..< NBands:
var line = fmt"{BandLabels[b]:<9}"
for gi in 1 ..< GainValues.len:
line.add fmt" {fmt3(leadCorr(arms[GainArmIdx[gi]].bands[b])):>8}"
echo line
echo ""
echo "=".repeat(120)
echo "LEAD INFORMATIVENESS -- capture slope (regression of applied lead on required lead) and lead correlation"
+164 -1
View File
@@ -253,6 +253,169 @@ Every D-item must end in a live A/B before any phase verdict.
---
## Phase 1 — *(unclaimed; append below)*
## Phase 1: the missing gain sweep *(owner: overnight job j100, committed)*
### Three negatives are on file (the morning reader must see these)
1. **BitBrain as previously shipped was statistically identical to Pattern
live.** 30 runs/arm, MDE 24.4 dmg/run: `docs/bitbrain_gun_verdict.md`,
commit `d93ce44`.
2. **Gun-mixing does not disrupt DrussGT.** TMHorizon+BitBrain mixed into the
rack produced no dodge disruption: `docs/gun_mix_disruption.md`, commit
`32a5e72`.
3. **BitBrain added no measurable aim information offline.** 450+: 16.200 deg
vs Pattern 16.193, hitProxy 0.0767 vs 0.0767 (`a82c864`, §0.3.5 above).
These bound the plausible upside: the previous BitBrain output — an ADDITIVE
angular shift — was information-free, so Phase 1 changes the output shape, not
the learning rate.
### 1.1 The missing sweep: Pattern x gain in [0.00, 1.00] **[MEASURED]**
`common_libs/tests/run_prediction_quality.nim` now carries sub-unity gain arms
(gain 0.0 is HeadOn, gain 1.0 is Pattern). Full 70-run sweep, 14 arms,
3 598 428 tick-bins, `common_libs/tests/prediction_quality_results.txt`.
Cells are `mean|err| deg [hitProxy]`; **hitProxy is the objective** (fraction of
tick-bins within `atan(18/range)`, the ruler's proxy for hit probability).
| band | |req| deg | g=0.00 | g=0.25 | g=0.50 | g=0.75 | g=1.00 | best hpx | dHpx | best |err| |
|---|---|---|---|---|---|---|---|---|
| 0–100 | 19.62 | 19.62 [.342] | 16.98 [.391] | 14.58 [.469] | 12.36 [.640] | **10.56 [.699]** | 1.00 | +.0000 | 1.00 |
| 100–200 | 19.98 | 19.98 [.172] | 17.79 [.166] | 16.10 [.167] | 14.97 [.234] | **14.75 [.342]** | 1.00 | +.0000 | 1.00 |
| 200–300 | 17.34 | 17.34 [.133] | 15.95 [.128] | 15.32 [.137] | 15.53 [.139] | **16.61 [.185]** | 1.00 | +.0000 | 0.50 |
| 300–450 | 14.61 | **14.61 [.105]** | 14.13 [.102] | 14.49 [.100] | 15.64 [.095] | 17.53 [.104] | 0.00 | +.0013 | 0.25 |
| 450+ | 12.33 | **12.33 [.098]** | 12.26 [.093] | 12.94 [.088] | 14.28 [.081] | 16.19 [.077] | 0.00 | +.0216 | 0.25 |
**The optimal gain curve (hitProxy-argmax per band) is**
**`[1.00, 1.00, 1.00, 0.00, 0.00]`** — use Pattern's full lead below 300 px and
**aim at the enemy's current position (zero lead, HeadOn) at 300+ px**. The
implied hit-probability gains vs Pattern are +0.00 pp (0–300), +0.13 pp
(300–450) and **+2.16 pp at 450+ (0.0767 -> 0.0984, a 28 % relative rise)**.
Weighted by tick-bin share (450+ is 64.2 %, 300–450 is 31.1 %) the whole-corpus
proxy rises ~+1.4 pp.
**[CAUSAL-SHIPPABILITY, important]** The rule only needs the *range*, which is
known at fire time, so applying `[1,1,1,0,0]` is causal and needs **no learning
at all** — a per-band table is shippable as a constant, exactly like the
Pattern radial-offset knob. What is *not* causal is the **estimation** of the
table from the same runs (it is in-sample here); a shipped table would be fitted
offline on past battles or learned online, which is what BitBrain does. The
table's value is robust to that caveat because the winning entries are the two
extremes (full Pattern lead, or zero lead), not a knife-edge intermediate value.
### 1.2 Shrinking the gain does NOT add lead information **[MEASURED]**
| band | corr @ g=0.25 | g=0.50 | g=0.75 | g=1.00 |
|---|---|---|---|---|
| 0–100 | 0.774 | 0.774 | 0.774 | 0.774 |
| 100–200 | 0.612 | 0.612 | 0.612 | 0.612 |
| 200–300 | 0.457 | 0.457 | 0.457 | 0.457 |
| 300–450 | 0.266 | 0.266 | 0.266 | 0.266 |
| 450+ | 0.165 | 0.165 | 0.165 | 0.165 |
Every `g > 0` column is **identical**: Pearson correlation is invariant under
positive scaling. A fractional gain therefore buys nothing on the
lead-information axis — it only shrinks the magnitude of an uninformative signal
toward the low-variance static aim. This confirms §0.3.3's reading and is the
mechanism behind the whole curve.
**The MSE-optimal gain and the hitProxy-optimal gain diverge [MEASURED, key].**
The `best |err|` column is the least-squares optimum `[1.00, 1.00, 0.50, 0.25,
0.25]`. At 200–300, MSE wants ~0.5 but hitProxy wants 1.0: Pattern's lead errors
are **bimodal** (it either nails the lead or is far off), so shrinking every
sample trades many small-within-tolerance hits for a smaller tail. Any *learned*
corrector that minimises squared error will therefore under-perform at mid range
— a fact this phase measured the hard way (an MSE-gain BitBrain dropped the
200–300 proxy from 0.190 to 0.131).
### 1.3 BitBrain rebuilt as a lead-gain corrector **[MEASURED]**
`common_libs/guns/bitbrain_gun.nim` is rewritten. The ADE+SBC network is gone
from the gun; the output is now a multiplicative gain on Pattern's lead,
`aim = LOS + gain * (patternAim - LOS)`, with `gain` chosen from a fixed
candidate set `{0, 0.25, 0.5, 0.75, 1.0}`.
* **Label path** (unchanged): at fire time we remember Pattern's lead and the
target's angular half-width `atan(18/range)`; `h = round(dist/speed)` ticks
later `tmhObservedAt` returns the enemy's observed bearing from the firing
position, giving `requiredLead = observedBearing - LOS`.
* **Learning rule** (changed): for every resolved sample we score *each*
candidate by whether `|gain*baseLead - requiredLead| <= atan(18/range)` and keep
the hit counts per range band; the band's gain is the **argmax hit rate** — the
hit-probability proxy itself, not squared error. This directly fixes the
bimodality failure above.
* **State / gate**: the state is the range band (causally known). The correction
is applied only for range >= 300 px (`BB_GAIN_BAND_MIN = 3`), below which
Pattern's lead is informative.
* **Stats are battle-scale**: a round boundary keeps the counts; a new battle or
target change (`resetLearning`/`targetChanged`) wipes them.
Offline 70-run result (same ruler, same run as §1.1):
| band | Pattern hpx | fixed rule [1,1,1,0,0] | BitBrain hpx | fixed−Pat | BB−Pat |
|---|---|---|---|---|---|
| 0–100 | 0.6993 | 0.6993 | 0.6993 | +0.0000 | +0.0000 |
| 100–200 | 0.3418 | 0.3418 | 0.3418 | +0.0000 | +0.0000 |
| 200–300 | 0.1850 | 0.1850 | 0.1850 | +0.0000 | +0.0000 |
| 300–450 | 0.1036 | 0.1049 | **0.1073** | +0.0013 | **+0.0037** |
| 450+ | 0.0767 | **0.0984** | 0.0957 | +0.0216 | **+0.0190** |
BitBrain's effective point estimates match the fixed table to within 0.27 pp at
450+ and actually exceed it at 300–450. The internal log shows why it is not
exactly equal: at long range the candidate hit rates are near-tied, so the
argmax flips between 0.00 and 0.25 (450+) / 0.00 and 0.75 (300–450); the
aggregate still lands on the right side. A fixed table is more stable; the
learned version needs no table and adapts per battle.
**Cost [MEASURED].** `/tmp/bench_bb` (200k ticks, 4 power bins) — Pattern
0.0224 ms/tick, BitBrain 0.0230 ms/tick, i.e. a **marginal ~0.0007 ms/tick**
(0.7 us/tick), versus the ~0.114 ms/tick of the old ADE+SBC gun and the
13.16 ms/tick budget. The whole 14-arm offline sweep now runs in 271 s
(0.0754 ms per tick-bin vs Phase 0's 0.1131 for 10 arms).
**Guard tests stay green [MEASURED].** `test_bitbrain` 32 checks,
`test_bitbrain_registration` 13, `test_rack_membership` 48,
`test_tm_pattern_registration` 20, `test_env_report` all pass. `env_report.nim`
is unchanged: the legacy BitBrain knobs are still resolved and reported. The
ruler still validates (recorded hits 1.360 deg vs misses 16.597 deg,
separation 12.20x deg / 13.34x px; oracle 0.000000 deg).
### 1.4 Verdict
**[MEASURED] Yes — a per-band lead-gain rule beats Pattern offline**, but only in
the two long bands: `[1,1,1,0,0]` gives +0.13 pp at 300–450 and **+2.16 pp at
450+** (a ~28 % relative rise in the hit-proxy, against Phase 0's realistic
causal ceiling of HeadOn's ~0.10, which it reaches). Below 300 px Pattern is
already optimal and the rule is a no-op.
**[MEASURED] Yes — BitBrain-as-gain-corrector matches that rule** (+1.90 pp at
450+, +0.37 pp at 300–450 over Pattern; within 0.27 pp of the fixed table at
450+, above it at 300–450), while being **causally learnable online** and
~163x cheaper per tick than the old gun.
**[INFERRED / honest caveat] BitBrain is not required to capture the gain** — a
constant `[1,1,1,0,0]` table would do it. BitBrain's value is that it discovers
the table per battle without one; its cost is cold-start (it begins at gain 1.0,
explaining the 0.27 pp gap at 450+) and the near-tie jitter in its point
estimates. So the honest ship decision is: **the per-band rule is the thing to
test live, and it can be shipped either as a constant table or as this online
learner** — the learner is redundant if a table is acceptable, and preferable
only if the optimum is expected to drift per enemy. The highest-value live
experiment is unchanged and stronger: a **range-gated Pattern<->HeadOn switch at
~300 px vs pure Pattern** (Phase 0's D1), left-running. **No live win is claimed
here** — the live gate is a separate phase.
### 1.5 Designs after Phase 1
| # | design | status |
|---|---|---|
| D1 | Live A/B: Pattern<->HeadOn switch at ~300 px vs Pattern | **TODO (highest value, live)** — offline now says +2.16 pp at 450+ |
| D2 | Raise Pattern lead *correlation* at long range (longer/multi-length keys, per-distance tables, k-NN) | TODO (offline-searchable) |
| D3 | Match-quality-conditional gain (keep full lead when the pattern match is good, shrink when bad) — attacks the bimodality that makes the MSE gain diverge | TODO (offline-searchable) |
| D4 | A causal predictability gate falling back to HeadOn when the future is unpredictable | TODO |
| D5 | Long-range power policy (already partly live) | TODO (offline proxy only) |
| D6 | Lead-gain sweep gains >= 1.0 | **DEAD — measured** (§0.3.2) |
| D7 | Constant sub-unity lead gain at all ranges | **DEAD — measured** (§1.1: gain < 1 hurts below 300) |
## Phase 2 — *(unclaimed; append below)*