225 battles, one frozen binary, five env-only arms, the frozen panel, 0 invalid
runs. Paired per opponent vs the shipped tfil:
strafe_notilt wins/run +0.38 [CI +0.16,+0.60] 9/9 opponents p=0.0039
dmg/run -10.2 [CI -25.8,+5.5] p=0.61, MDE 20.4 (not detectable)
incoming hit rate 12.24% vs 18.17%, dmg taken 150 vs 200
strafe_325 wins/run +0.33 [CI +0.04,+0.63] 10/12 p=0.0386
ring dmg/run +31.2 [CI +11.5,+50.9] 13/15 p=0.0074, wins/run -0.04 (ns)
but hit rate 29.4% at 236 px: a damage/survival trade, not a win
ring_notemp indistinguishable from tfil on both primaries
Round wins in this harness are survival wins (in 216/219 attributable runs the
win count equals the rounds the opponent died in), and the winner takes ~1/3
fewer hits while fighting ~74 px farther out. The shipped tfil is last of five
on wins: the DrussGT-only picture did not generalize.
Also: tournament_analyze.py now prints BOTH readings of the pre-registered
'while the other does not go down' clause (strict: nothing is better;
substantive: the two strafe arms and ring are better on one metric each).
The repo's first multi-opponent gun measurement. Adds tools/ab/gauntlet_run.sh
(per-opponent A/B over the legacy roster, subject = frozen ModularBot),
tools/ab/gauntlet_analyze.py (paired per-opponent deltas, cross-opponent sign
test, style split, MDE) and the arm/opponent fixtures.
Result: BitBrain does NOT generalize beyond DrussGT. 32 opponents x 2 arms x
3 runs x 5 rounds = 192 battles / 960 rounds, 0 failed, 0 retries: damage/run
214.5 (pattern) vs 210.9 (bb), sign-flip p=0.53; round wins 237/480 vs 239/480,
p=0.91. Sign test: bb better on 13/32 opponents (damage). The DrussGT-only
penalty does not carry. The owner's 'killer vs regular movers' sub-claim is not
supported: regular bucket +1.3 dmg/run vs dodgers -0.2 (MW p=0.85), and the
measured movement predictability does not correlate with the delta.
4 arms x 15 runs x 7 rounds (60 battles, 0 failed) vs real DrussGT on the
shipped TFIL default, frozen at ed25ce2. bb_id (gain 1.0 identity) is
statistically indistinguishable from shipped Pattern -> plumbing validity
check passes. No BitBrain arm beats TMHorizon or Pattern: bb_learn (the config
the owner likely ran) is the worst arm (276 dmg/run, 39/105 wins), the only
comparison at alpha=0.05 is Pattern beating it on damage. Learned gains
(>=1.0, gated >=300px) over-lead and lose 1.61pp of hit rate at 300-450px.
MDE 29.5 dmg/run, 1.235 wins/run; a 6-4-sized effect needs ~39 runs/arm.
Runs the pre-registered A/B for the two movement changes in HEAD: the
time-indexed bullet heat (TR_TFIL_HEAT_TIME, fca8993) and the removal of the
invented virtual centre pillar (d0750ab). One frozen binary from HEAD vs real
DrussGT: 6 arms x 10 runs x 7 rounds = 60 battles, 420 rounds, 0 failed.
Judged on damage/run and ROUND WINS only (hit rate and hits-taken are context):
hit rate would have inverted the verdict again - tau3 has the best pooled hit
rate of all arms (11.56%) and the fewest round wins (20/70).
RESULT (vs the reconstructed pre-change mover "old"):
heat-time HURTS. tau3/tau5/tau9 lose 1.3-1.7 wins/run (p=0.0010-0.0125) and
deal 22-38 less damage/run (p=0.004-0.047); tau15 is a wash on wins (p=0.64)
and 22 damage/run lower (p=0.046). Nothing improves either metric.
pillar removal does nothing measurable. old vs pillaoff: +5.7 damage/run
(p=0.71), +0.5 wins/run (35 vs 30, p=0.43), 30.8 MORE damage taken/run
without the pillar (p=0.040). The mechanism check proves the knob works
(centre-box occupancy 0.09% -> 2.37%, p<0.0001; range 469 -> 443 px,
p=0.0002), so this is a real behaviour change that buys nothing. At n=10 the
pillar contrast is inside the MDE (33 damage/run, 1.2 wins/run), so this is
not a proven regression.
Flags that the shipped default (pillar removed) should be reverted to the
TR_TFIL_PILLAR_ON behaviour; heat-time stays off.
Adds tools/ab/arms_heat_pillar.txt and tools/ab/ab_mechanism.py (per-tick
mechanism check: central-box occupancy, range distribution, live enemy-bullet
proximity) plus the captured summary/report fixtures.
Task A of campaign phase 2: the lead-gain candidate set is now pure env, so the
live arms need no recompile.
- common_libs/guns/bitbrain_gun.nim: BB_GAINS_ENV (TR_BITBRAIN_GAINS); the
candidate list is parsed once at gun construction into a dynamic seq, so the
hit counts/hit rates are sized to it. Unset/unparsable -> the shipped
BB_CAND set [0,0.25,0.5,0.75,1.0] (byte-identical behaviour). Exactly ONE
candidate degenerates to a FIXED gain applied from the first shot (learning
bypassed), still gated to the long bands. parseGains clamps to [0,8],
de-dupes and sorts so the argmax tie rule is unchanged. The [bb] line now
prints the APPLIED gain AND the resulting angular shift, so a run's
correction is auditable from stdout.
- ModularBot_garage/src/env_report.nim: emit TR_BITBRAIN_GAINS (resolved
candidate set) and add BB_GAINS_ENV to the known-name list.
- tools/ab/arms_leadgain.txt: the 6-arm phase-2 sweep definition.
2 arms x 15 runs x 7 rounds, one frozen binary from HEAD a82c864, real DrussGT,
server-side events sidecar. Shipped rack is onlyPattern, so control=Pattern-only
and headon=HeadOn-only (TR_RACK_PATTERN=off TR_RACK_HEADON=both).
arm dmg/run dmgtk/run round wins shots/run
control 279 211 48/105 785
headon 14 228 0/105 580
Round wins and dmg/run both separate at p<0.0001 (MC permutation, se 0.0000),
~7x the damage MDE (35.8). Per range band (pooled, 15 runs):
300-450: Pattern 12.3% (4590 shots) vs HeadOn 0.6% (3701) p<0.0001, MDE 2.0pp
450+ : Pattern 9.2% (6671) vs HeadOn 0.4% (4177) p<0.0001, MDE 1.1pp
HeadOn loses EVERY long-range band by 20-23x, so the whole-battle loss is not a
close-range artefact.
The offline ruler (prediction_quality_results.txt) predicted the opposite: HeadOn
meanAbs 14.61 vs Pattern 17.53 at 300-450 and 12.33 vs 16.19 at 450+, hitProxy
.105/.104 and .098/.077 (+27%). That is an open-loop replay of a FIXED enemy
track, so it cannot see that a different bullet makes the surfer dodge
differently; live, the static gun does not lead at all.
TR_PATTERN_RAD_SCALE arms were skipped: applyRadial scales aim DISTANCE along an
unchanged bearing, so it cannot express 'less lead' (bearing is what firing uses).
HeadOn confirmed to ignore bulletSpeed (head_on.nim:9), liveness OK 15/15.
Adds the range-band analyzer tools/ab/ab_range_bands.py (reuses the lead-capture
Run alignment) and the captured fixtures. Does not touch bitbrain_gun.nim /
bitbrain_campaign.md (job-100).
4 arms x 7 runs vs real DrussGT. mix alternates the two guns 476 times/7 runs
(liveness OK) but our bullets are no more varied (power sd / aim-offset sd flat)
and DrussGT's dodge quality is unchanged (miss/tick mix-pat +0.03, p=0.66; MDE
3.8%). mix wins 24/49 = the 49% baseline; the user's 6/10 has P=0.353 at 49%.
New tools/ab/ab_dodge_analyze.py splits the validated per-shot dodge instrument
by arm and adds gun-switch/power/bearing liveness; fixtures committed.
- docs/bitbrain_gun_verdict.md: control vs bb_decay (decay SBC memory) vs a
provably-zero placebo, 30 runs/arm vs real DrussGT. Nothing separates
(bb_decay +3.3 dmg/run, p=0.71; round wins 97/210 vs 97/210, p=1.00); the
7-run shape does not replicate. TR_BITBRAIN_RANGE=0 is clamped to 1.0 deg
(bitbrain_gun.nim:207) so it is NOT a zero-shift placebo; TR_BITBRAIN_MIN_OBS
unreachable is used instead.
- tools/ab/ab_analyze.py: keep exact enumeration for C(n,na)<=20e6 (7v7), add
a seeded Monte-Carlo permutation test (1e6 draws, 0x5eed5eed) with its
standard error, a tie-corrected Mann-Whitney U cross-check, a minimum
detectable effect line, all-pairs comparisons, and a [bb] shift check.
- tools/ab/README.md: document the new analyzer output.
The user pushed back on "the TM can't be your best 1v1 gun", correctly, because two
decisive tests had never been run. Both are now run and they agree.
TASK 1 - THE GF HEAD vs ITS MAJORITY-CLASS BASELINE (offline, n=1,751,067):
label histogram [254286, 284578, 678879, 297055, 236269]
majority class = 2 (the CENTRE bucket) = 38.77%
RAW head accuracy = 36.69% -> margin **-2.08 pp, BELOW majority**
GATED head accuracy = 40.37% vs 38.75% majority -> +1.62 pp, BUT it predicts the
majority class on 62.4% of ticks and its minority recall is 13.6% / 12.9% - a
base-rate predictor wearing a classifier's clothes.
Shuffled control sits at its own majority (20.04% vs 20.12%), confirming chance.
**THE OLD "46% vs 20% CHANCE" FIGURE I QUOTED WAS WRONG ON TWO COUNTS:** the
baseline is 38.8%, not 20%, and the 46% predated the deferred-label fix. Against
the correct baseline the head is BELOW it.
TASK 2 - THE FIRST-EVER LIVE A/B OF THE TM GUN (7 runs x 7 rounds per arm, one
frozen binary from git archive HEAD = eb74f9b2, sha256 cb66d66b..., real DrussGT,
every arm forced alone with TR_RACK_<GUN>=both and all 14 others off, liveness
confirmed per run):
arm shots real % dmg/run round wins
onlyPattern 4610 10.74% 285 25/49
onlyTMPATTERN (radial) 3374 3.50% 71 0/49
onlyLinear 3218 3.23% 61 0/49
Pattern vs TM: +7.22 pp / +213.7 dmg, exact p=0.0006
TM vs Linear: +0.30 pp, p=0.659 (dmg p=0.438)
**The TM is statistically INDISTINGUISHABLE from its own Linear base live.** So it
is not "the TM works and we are aiming it wrong".
DIRECT ANSWER: **(c) It loses live AND sits at/below majority - the target carries
no learnable signal beyond the base rate, and that is the reason.** The reason is
not the machine, not the knobs, and not the application alone: the thing it was
asked to predict is dominated by the modal answer.
This closes the TM-as-gun thread. If a TM is wanted in the bot, a firing gate or a
movement decision is a better fit for a boolean-rule classifier than an aim point -
that is untested and is a different project.
A LIVE GF-MODE ARM WAS NOT RUN (stated as unmeasured): the task pinned one frozen
HEAD binary and HEAD registers the TM gun as radial only; Task 1 already makes GF
the unpromising candidate.
HARNESS FIX WORTH KEEPING: `tools/ab/which_gun_arm_env.sh` left the TARGET gun
unset, so with the now-Pattern-only default it silently fell back to the FULL rack
- an arm could appear to test a single gun while actually running the whole rack.
It now emits `TR_RACK_<GUN>=both` for the target and `=off` for all 14 others.
(Earlier which-gun results are unaffected: they ran before the Pattern-only default,
or - as in the melee/1v1 campaign - set the explicit `=both` themselves.)
tm_pattern.nim gains a per-class confusion matrix (warm samples only) to support the
majority baseline; no behaviour change. Adds Round 4 to
tm_pattern_sweep_results.md with both tasks and the interpretation rule.
Follow-up to e0666a5, which showed Pattern alone (10.78%) beats the full rack
(6.93%). That left two open questions: is a SMALL rack of good guns better than
Pattern alone, and does the selector add value on a good rack (rather than only
on the bloated one)? Both are now answered: NO and NO.
6 arms x 7 runs x 7 rounds, one frozen binary from CLEAN HEAD e0666a5 (built via
`git archive`, source verified byte-identical to the clean tree), rack knobs
only, 8 concurrent battles, real server-side hit rate vs the real DrussGT, exact
two-sided permutation test on per-run rates.
arm guns (selector active?) real % dmg/run p vs onlyPattern
onlyPattern Pattern, NO selection 10.36 264 --
lean8 HeadOn,Linear,Circular,Accel,Pattern,GF,KNN,WallBounce 6.31 146 0.0169
lean6 lean8 - HeadOn 8.83 212 0.0262
pairPC Pattern + Circular 8.23 185 0.0460
pairPK Pattern + KNN 9.80 264 0.3998
pairPL Pattern + Linear 8.23 200 0.0035
The control replicates the prior run (10.36% vs 10.78% before; same binary tree,
different build path).
THE MECHANISM, from the per-arm selected-gun mix - the virtual signal keeps
ranking the WRONG guns first, even on a two-gun rack:
lean8: HeadOn 46.2% of ticks at 2.0% REAL; Pattern only 14.4% (12.6% real)
lean6: Pattern 29.2% (10.4% real) vs KNN 25.3% (8.1%) and Linear 15.2% (8.0%)
pairPC: Circular 66.8% (6.9% real) vs Pattern 33.2% (11.2% real) - over-picks Circular
pairPL: Linear 57.7% (6.0% real) vs Pattern 42.3% (11.1% real) - over-picks Linear
pairPK: Pattern 86.4% - ties ONLY because the selector happens to pick Pattern
most of the time; it is numerically lower with identical dmg/run
So the failure is NOT rack size. Pruning does not fix it; the ranking is wrong.
VERDICT: ship `onlyPattern` - Pattern alone with selection bypassed - at 10.36%
real and 264 dmg/run, vs lean8 6.31%/146 and the prior full rack 6.93%/159.
This DIRECTLY CONTRADICTS the standing user directive to keep virtual-fitness
selection, so it is recorded here plainly rather than quietly acted on: disable
the selector (`TR_RACK_<every gun but PATTERN>=off`) pending a better fitness
signal. The mechanism itself is left intact and functional so it can be re-enabled
with one env var, and so it can be fixed rather than discarded.
REMAINING CAVEAT: ONE ADVERSARY. All of this is vs DrussGT. Pattern as the default
must be re-checked against other bots first - that is the next job.
Extends the reusable harness (tools/ab/which_gun_arm_env.sh now has lean8/lean6/
pairPC/pairPK/pairPL; which_gun_analyze.py is parameterised by WHICHGUN_OUT and
compares against both `full` and `onlyPattern`).
MEASURED against the real DrussGT, one frozen binary built from clean HEAD, rack
knobs only (no source edits), 5 arms x 7 runs x 7 rounds, 8 concurrent battles,
judged ONLY on server-side real hit rate from the events sidecar, exact
two-sided permutation test on per-run rates.
arm runs shots hits real % dmg/run p vs full
full (shipped) 7 3898 270 6.93 159 --
onlyPattern 7 4582 494 10.78 287 0.0012 <- BETTER
onlyKNN 7 4033 207 5.13 119 0.1340
onlyLinear 7 3215 105 3.27 65 0.0082
onlyGF 7 3193 72 2.25 45 0.0012
Firing Pattern ALONE gives +3.85pp pooled hit rate and +80% damage per run, and
it fires MORE shots (4582 vs 3898) - it dominates on rate and volume. This is not
"any single gun wins" (full beats Linear, GF and KNN); it is specifically
"Pattern alone beats the rack".
WHY - the virtual fitness signal mis-ranks guns against real outcomes:
- HeadOn is massively over-selected: 31.4% of ticks, the most real shots (1070),
but only 4.5% REAL. It alone drags the rack down.
- Pattern has the best virtual rank and near-best real rate (11.9%, rank 2), yet
is selected only 22.6% of the time.
- Linear's apparent strength was SELECTION BIAS: conditional on being selected it
looked like 15.2% (n=33), but its UNCONDITIONAL rate (onlyLinear) is 3.27%.
Every earlier per-gun "real rate" in this repo is conditional on selection and
is therefore confounded. This experiment is the clean measurement.
NOT YET SETTLED (do not overclaim):
- ONE ADVERSARY. All of this is vs DrussGT. Pattern must be re-checked against
other bots before it becomes the default on this evidence alone.
- Whether a SMALL rack of good guns beats Pattern alone. The selector is negative
value on the CURRENT bloated rack; that does not prove it is negative value on
a rack of only good guns. That is the next experiment and it decides whether
the selection apparatus is fixed or disabled.
- The user's standing directive is to KEEP virtual-fitness selection. This
measurement conflicts with it, so the next step tests the selector on a small
good rack rather than assuming either answer.
Context - three prior selection-side attempts all failed: hysteresis (7.02% ->
5.10%, p=0.002), commitment (7.17% -> 4.44%, p=0.0012), arrival-accuracy
tie-break (7.08%, p=0.88 null). The per-tick random draw is load-bearing on
three independent measurements. This experiment locates the real problem one
level up: which guns are in the rack, and that the virtual signal ranks them
wrongly.
Preserves the reusable harness (tools/ab/which_gun_run_one.sh,
which_gun_arm_env.sh, which_gun_analyze.py) and the full writeup
(docs/selector_negative_value.md).