56/60 opponent-level win deltas are non-negative and no opponent family
regresses reproducibly (the one WallAvoider loss reverses in Batch 2).
corr(dwins, d_hit_rate) = -0.09, corr(dwins, d_damage) = +0.38: the aggregate
win is a survival effect, but a per-opponent hit-rate gain does not predict a
per-opponent win gain - so hit rate stays an explanation, never a proxy.
- tfil is 4th of five on round wins, not last (ring is nominally 0.04 lower,
ns) - the Batch-1 commit message overstates one word; the correction is
recorded in the ledger rather than rewritten.
- the Batch-2 direct answer quoted three of four CIs excluding 0; it is four of
four ([+0.04,+0.63], [+0.16,+0.60], [+0.27,+0.89], [+0.22,+0.72]).
- added the cleanest aggression isolation of Batch 1 (ring - ring_notemp, same
engine and heat field, range weighting alone): +41.9 dmg/run, -0.11 wins/run,
+12.8 pp incoming hit rate at 236 vs 395 px.
Same frozen panel, same 3x3 design, new session on commit 8efa627 (no source file
changed since 1984a78, so the same code), 225 battles, 0 invalid runs. Arms:
tfil, strafe_notilt, strafe_325 + the tilt re-armed at 600px and 250px.
Paired vs tfil: strafe_325 +0.58 wins/run [CI +0.27,+0.89] 11/12 p=0.0063;
strafe_notilt +0.47 [+0.22,+0.72] 10/11 p=0.0117; tilt_600 +0.40 [+0.04,+0.76]
(sign test 8/11 p=0.23, sign-flip p=0.049); tilt_250 +0.38 [+0.07,+0.69] 10/12
p=0.039. Incoming hit rate -5.2..-6.9 pp with 0/15 opponents favouring tfil.
The range TARGET is not the lever: re-arming the tilt moved the achieved
distance from 459px (no steering) to 478px and 415px, and none of the three is
separable on wins. This overturns Batch 1's reading that the tilt costs wins -
the honest statement is that the tilt's win effect is below this design's
resolution. tfil reproduced to within 1.4 pp (40.7% -> 39.3% of rounds), so the
baseline itself is stable across sessions.
Ledger: Batch 2 section, the verbatim analyzer report, a data-driven
what-to-try-next, and the session log.
225 battles, one frozen binary, five env-only arms, the frozen panel, 0 invalid
runs. Paired per opponent vs the shipped tfil:
strafe_notilt wins/run +0.38 [CI +0.16,+0.60] 9/9 opponents p=0.0039
dmg/run -10.2 [CI -25.8,+5.5] p=0.61, MDE 20.4 (not detectable)
incoming hit rate 12.24% vs 18.17%, dmg taken 150 vs 200
strafe_325 wins/run +0.33 [CI +0.04,+0.63] 10/12 p=0.0386
ring dmg/run +31.2 [CI +11.5,+50.9] 13/15 p=0.0074, wins/run -0.04 (ns)
but hit rate 29.4% at 236 px: a damage/survival trade, not a win
ring_notemp indistinguishable from tfil on both primaries
Round wins in this harness are survival wins (in 216/219 attributable runs the
win count equals the rounds the opponent died in), and the winner takes ~1/3
fewer hits while fighting ~74 px farther out. The shipped tfil is last of five
on wins: the DrussGT-only picture did not generalize.
Also: tournament_analyze.py now prints BOTH readings of the pre-registered
'while the other does not go down' clause (strict: nothing is better;
substantive: the two strafe arms and ring are better on one metric each).
The repo's first multi-opponent gun measurement. Adds tools/ab/gauntlet_run.sh
(per-opponent A/B over the legacy roster, subject = frozen ModularBot),
tools/ab/gauntlet_analyze.py (paired per-opponent deltas, cross-opponent sign
test, style split, MDE) and the arm/opponent fixtures.
Result: BitBrain does NOT generalize beyond DrussGT. 32 opponents x 2 arms x
3 runs x 5 rounds = 192 battles / 960 rounds, 0 failed, 0 retries: damage/run
214.5 (pattern) vs 210.9 (bb), sign-flip p=0.53; round wins 237/480 vs 239/480,
p=0.91. Sign test: bb better on 13/32 opponents (damage). The DrussGT-only
penalty does not carry. The owner's 'killer vs regular movers' sub-claim is not
supported: regular bucket +1.3 dmg/run vs dodgers -0.2 (MW p=0.85), and the
measured movement predictability does not correlate with the delta.
Adds common_libs/tests/measure_melee_bitbrain_ab.nim (+ .sh driver, .py analyzer,
committed per-run fixtures) and docs/melee_bitbrain_ab.md.
Experiment: 4-bot Free-For-All (ModularBot + WaveSurfer + PatternMover +
RandomMover), 4 arms x 16 runs x 7 rounds, frozen ModularBot from git archive
HEAD (commit 0f5cfe3, binary 11bba27), shipped tfil movement in every run.
Arms differ only in the gun rack: pattern (shipped), bb_round, bb_ret, bb_learn.
Result: NOT DETECTABLE. Score (server round score = damage + survival bonus)
differs by -63..+33 pts (perm p=0.16-0.71) against an MDE of 151 (~5.1%).
Every arm finishes rank 1. Round wins hint BitBrain's way (112/112 and 111/112
vs 109/112) but p=0.225 (MW 0.080), half the 0.40-win MDE.
Liveness proven: rack boot lines flip (rack active melee = PATTERN / BITBRAIN),
every run faced 3 distinct targets and ~66-69 target changes, and the bb arms
logged one [bb-reset] reason=target_change per switch. The melee premise was
exercised; the fast adaptation bought no measurable score edge at this sample.
Live 3-arm x 15-run x 7-round A/B vs real DrussGT on commit 0f5cfe37.
Primary: surf ties strafe on round wins (37/105) and damage (255 vs 250/run),
both below tfil (45/105, 293/run; damage p=0.003). Incoming hit rate: surf
13.51% (worst) vs strafe 9.40% (best) and tfil 10.40%. So the plain surfer does
NOT dodge better and does NOT win more. Also records the j107 trap: strafe
dodges best yet wins fewer rounds than tfil. Next step: range/aggression A/B,
not a BitBrain upgrade.
4 arms x 15 runs x 7 rounds (60 battles, 0 failed) vs real DrussGT on the
shipped TFIL default, frozen at ed25ce2. bb_id (gain 1.0 identity) is
statistically indistinguishable from shipped Pattern -> plumbing validity
check passes. No BitBrain arm beats TMHorizon or Pattern: bb_learn (the config
the owner likely ran) is the worst arm (276 dmg/run, 39/105 wins), the only
comparison at alpha=0.05 is Pattern beating it on damage. Learned gains
(>=1.0, gated >=300px) over-lead and lose 1.61pp of hit rate at 300-450px.
MDE 29.5 dmg/run, 1.235 wins/run; a 6-4-sized effect needs ~39 runs/arm.
Runs the pre-registered A/B for the two movement changes in HEAD: the
time-indexed bullet heat (TR_TFIL_HEAT_TIME, fca8993) and the removal of the
invented virtual centre pillar (d0750ab). One frozen binary from HEAD vs real
DrussGT: 6 arms x 10 runs x 7 rounds = 60 battles, 420 rounds, 0 failed.
Judged on damage/run and ROUND WINS only (hit rate and hits-taken are context):
hit rate would have inverted the verdict again - tau3 has the best pooled hit
rate of all arms (11.56%) and the fewest round wins (20/70).
RESULT (vs the reconstructed pre-change mover "old"):
heat-time HURTS. tau3/tau5/tau9 lose 1.3-1.7 wins/run (p=0.0010-0.0125) and
deal 22-38 less damage/run (p=0.004-0.047); tau15 is a wash on wins (p=0.64)
and 22 damage/run lower (p=0.046). Nothing improves either metric.
pillar removal does nothing measurable. old vs pillaoff: +5.7 damage/run
(p=0.71), +0.5 wins/run (35 vs 30, p=0.43), 30.8 MORE damage taken/run
without the pillar (p=0.040). The mechanism check proves the knob works
(centre-box occupancy 0.09% -> 2.37%, p<0.0001; range 469 -> 443 px,
p=0.0002), so this is a real behaviour change that buys nothing. At n=10 the
pillar contrast is inside the MDE (33 damage/run, 1.2 wins/run), so this is
not a proven regression.
Flags that the shipped default (pillar removed) should be reverted to the
TR_TFIL_PILLAR_ON behaviour; heat-time stays off.
Adds tools/ab/arms_heat_pillar.txt and tools/ab/ab_mechanism.py (per-tick
mechanism check: central-box occupancy, range distribution, live enemy-bullet
proximity) plus the captured summary/report fixtures.
Adds a per-sample intrinsic-confidence field (GunPrediction.confidence,
threaded through FeedbackEvent/VirtualBullet, populated by Pattern, DecayGF,
KNN, GuessFactor, Tsetlin, TMHorizon) and an offline recorder + analyzer that
reproduce the paper's Figure 2 per gun and its Eq-8 composite.
Measured on 3 held-out tr-bridge DrussGT battles (33k ticks, ~133k samples/gun):
- FAITHFUL: DecayGF (rho +0.133), KNN (+0.090), Pattern (+0.064, weak).
- GuessFactor is ANTI-faithful (rho -0.067); Tsetlin c_max is useless (0.001).
- No pair of guns specialises complementarily: the same gun dominates both
high-confidence slices in every pair.
- Eq-8 alpha-normalised confidence-weighted composite: 18.41% vs Pattern
20.45% (McNemar p=3.1e-126). Faithful-only variant 18.68%, still loses.
Shuffle control passes weakly (composite > shuffle, p=4e-14) so ~0.7pp of
competence is real but ~2pp short. Offline veto: design is dead.
See docs/tmcomposites_gate.md.
Re-runs the SBC coincidence premise as a cheap veto test on the 70-battle
live-vs-DrussGT corpus (/tmp/tfil_ab2). Defines the wave-relative state
(lat/vlat/toa/room/turn, 5/7.9/10 bits at Q=2/3/4), quantises the miss offset
at the bullet's arrival into 7 bins, and sweeps window length K in
{1,4,8,16,32,48} with an interpolated suffix-backoff model under a BY-BATTLE
70/30 split (3 seeds).
Result: NO. On the pre-fire frame the window is worse than the single
fire-tick state at every K/Q/A (e.g. K=8 costs +0.35..+0.46 bits). On the
during-flight frame the entire apparent gain is the later decision tick, not
the window; the single state alone drops 2.70 -> 1.31 bits as K goes 1 -> 32.
The shuffle-order control confirms recency matters but the windows do not:
by Q=4/K=8 they average ~1 observation and never recur. The single state
survives as a strong predictor (log-loss 2.346 vs 2.698 majority; bin
accuracy 0.409 vs 0.235).
Gate only: no gun, no live-win claim.
Adds an smCounted storage mode alongside the default smBitset. Each
(i,j,class) cell becomes a saturating uint8 counter; learn increments it and
a global fractional decay (c -= c shr decayShift every decayEvery learns)
makes forgetting possible. infer sums raw counters; new inferProb sums the
per-cell posterior P(class|cell) (scale-free, recommended readout).
Bitset path is the default and byte-for-byte unchanged: test_bitbrain 56/56
(was 32), and test_bitbrain_mnist reproduces 97.210% corrected / 96.540%
bug-compatible exactly.
Counted mode configurable at runtime (TR_BITBRAIN_MODE / TR_BITBRAIN_DECAY_*)
and compile time (-d:bitbrainDecay*). Measured: forgetting (86.2% vs 48.9% on
a permuted-label stream), probabilities (rare-class balanced 0.998 vs 0.500),
and the stationary cost (counted hurts MNIST; see docs/bitbrain_counted_sbc.md).
Harness: common_libs/tests/measure_counted_sbc.nim
6 arms x 7 runs x 7 rounds vs real DrussGT on one frozen binary (2747ebd).
Validity check PASSES: g100 (fixed gain 1.0) is statistically indistinguishable
from control (dmg p=0.65, wins p=0.62, ALL hit rate +0.04pp p=0.92; zero [bb]
lines = provably no correction). Fixed gains above 1.0 LOSE at 300+: gfix150
-94 dmg/run (p=0.0006), 4/49 vs 16/49 wins (p=0.009), -3.85pp at 300-450
(p=0.0006) and -2.25pp at 450+ (p=0.0023). The hypothesis arm ghi (learner
allowed above 1) is directionally positive but inside the MDE (+14.3 dmg/run
p=0.41; +0.83pp at 450+ p=0.35). Kill the gain axis in both directions.
Includes the mandatory correction notice: Phase 1's [1,1,1,0,0] is LIVE-REFUTED
by 140fe25, and the standing rule that offline is veto-only / live decides.
Task A — the gain region Phase 0 never covered (gain < 1). Extend the
prediction-quality ruler with gain 0.25/0.50/0.75 arms and a fixed causal
per-band arm. Full 70-run result: the hitProxy-argmax curve is
[1.00, 1.00, 1.00, 0.00, 0.00] — Pattern below 300 px, HeadOn above —
worth +0.13 pp at 300-450 and +2.16 pp at 450+ (0.0767 -> 0.0984). Lead
correlation is identical for every g>0 (Pearson is scale-invariant), so a
shrinking gain adds no lead information; and the least-squares optimum
[1,1,1,1,0.25,0.25] diverges from the hitProxy optimum because Pattern's
lead errors are bimodal.
Task B — rebuild guns/bitbrain_gun.nim as a lead-gain corrector:
aim = LOS + gain*(patternAim - LOS), gain learned online per range band by
ranking candidate gains on the hit-probability proxy (the observed lead
label via tmhObservedAt), gated to range >= 300 px. ADE+SBC output removed.
Offline (70 runs): Pattern below 300 px, +0.37 pp at 300-450, +1.90 pp at
450+ (hitProxy 0.0957 vs 0.0767), matching the fixed rule to within 0.27 pp
at 450+ and exceeding it at 300-450. ~0.0007 ms/tick marginal (old gun
~0.114 ms/tick). Default off; rack membership, env report and guard tests
(bitbrain 32, registration 13, rack 48, tm_pattern 20, env_report) unchanged
and green.
Ledger: docs/bitbrain_campaign.md Phase 1, including the three on-file
negatives and the causal-shippability note. No live claim.
2 arms x 15 runs x 7 rounds, one frozen binary from HEAD a82c864, real DrussGT,
server-side events sidecar. Shipped rack is onlyPattern, so control=Pattern-only
and headon=HeadOn-only (TR_RACK_PATTERN=off TR_RACK_HEADON=both).
arm dmg/run dmgtk/run round wins shots/run
control 279 211 48/105 785
headon 14 228 0/105 580
Round wins and dmg/run both separate at p<0.0001 (MC permutation, se 0.0000),
~7x the damage MDE (35.8). Per range band (pooled, 15 runs):
300-450: Pattern 12.3% (4590 shots) vs HeadOn 0.6% (3701) p<0.0001, MDE 2.0pp
450+ : Pattern 9.2% (6671) vs HeadOn 0.4% (4177) p<0.0001, MDE 1.1pp
HeadOn loses EVERY long-range band by 20-23x, so the whole-battle loss is not a
close-range artefact.
The offline ruler (prediction_quality_results.txt) predicted the opposite: HeadOn
meanAbs 14.61 vs Pattern 17.53 at 300-450 and 12.33 vs 16.19 at 450+, hitProxy
.105/.104 and .098/.077 (+27%). That is an open-loop replay of a FIXED enemy
track, so it cannot see that a different bullet makes the surfer dodge
differently; live, the static gun does not lead at all.
TR_PATTERN_RAD_SCALE arms were skipped: applyRadial scales aim DISTANCE along an
unchanged bearing, so it cannot express 'less lead' (bearing is what firing uses).
HeadOn confirmed to ignore bulletSpeed (head_on.nim:9), liveness OK 15/15.
Adds the range-band analyzer tools/ab/ab_range_bands.py (reuses the lead-capture
Run alignment) and the captured fixtures. Does not touch bitbrain_gun.nim /
bitbrain_campaign.md (job-100).
New harness (common_libs/gun_harness/prediction_quality.nim +
common_libs/tests/run_prediction_quality.nim): per-gun single-tick aim error in
degrees against the true continuous interception point on the recorded
live-vs-real-DrussGT corpus (/tmp/tfil_ab2/out, 70 runs, 899607 ticks), per
range band, with the hit-probability proxy mean(|err|<=atan(18/range)).
Validated: recorded hits separate from misses 13.34x px (reference 11.59x),
perfect-oracle max |err| = 0, correct ordering on synthetic ground truth, two
full runs byte-identical. Fixed a wrap180 bug (Nim float mod keeps the dividend
sign) that inflated the negative error tail.
Bar (mean|err| deg [hitProxy] at 450+): Pattern 16.19 [0.077], naive-linear
22.86 [0.054], TMHorizon 16.20 [0.076], BitBrain 16.20 [0.077], static HeadOn
12.33 [0.098], oracle 0 [1.0]. Lead-gain sweep on Pattern is a dead end (1.0
wins every band). Naive-linear applies ~1.8x Pattern's lead but carries no more
lead information (corr 0.178 vs 0.165) and is strictly worse. Ledger:
docs/bitbrain_campaign.md. All verdicts remain live-only.
4 arms x 7 runs vs real DrussGT. mix alternates the two guns 476 times/7 runs
(liveness OK) but our bullets are no more varied (power sd / aim-offset sd flat)
and DrussGT's dodge quality is unchanged (miss/tick mix-pat +0.03, p=0.66; MDE
3.8%). mix wins 24/49 = the 49% baseline; the user's 6/10 has P=0.353 at 49%.
New tools/ab/ab_dodge_analyze.py splits the validated per-shot dodge instrument
by arm and adds gun-switch/power/bearing liveness; fixtures committed.