The picker scored candidates on pathMaxHeat alone and then drew uniformly
among the survivors, so a mirror-side tile was as likely as a straight-ahead
one. Added a continuous turn cost as a DRAW WEIGHT applied only after the
hard heat filter:
w = max(1, round(1 + TR_TFIL_TURN_BIAS * (1 - max(0,|turn| - REF)/180)))
- TR_TFIL_TURN_BIAS (default 0) is the odds ratio straight-ahead vs 180 deg;
TR_TFIL_TURN_REF_DEG (default 45) is where the penalty starts. Both
default-off-effect: the default-path golden in test_tfil_commit_env.nim is
unchanged and still passes.
- Turn cost is NEVER folded into the heat score. The filter stays hard.
- The draw stays random (j51 measured an argmin worse); every weight is
floored at 1, so the pool can never be emptied and bias 0 is exactly the
shipped uniform draw.
- |turn| now travels on the ScoredTile, and the commit log gained turn /
minturn / promote so a caller can measure the regret of the draw.
Guards: 51 -> 66 checks (an absurd 99:1 bias never rescues an over-threshold
tile; mean |turn|, draw regret, >90 and mirror-side shares all fall; path
heat does not rise). env_report + .env.example updated.
Pooled verdict on the pre-registered rule is NOT DISTINGUISHABLE: the primary
cross-opponent sign test on round wins is 11/14, p=0.05737 for 'arrive', over the
0.05 line. Block 1 alone passed (p=0.01294); block 2 alone did not (p=0.0654).
Every delta is positive in both blocks for all three arms and damage/run is UP
on all three, so this reads as an under-powered null at the MDE boundary
(MDE 0.31 wins/run, observed 0.28) - but the verdict layer is not
reinterpreted, exactly as gate v1 was not.
The MECHANISM is established cleanly and replicates in both blocks: incoming hit
rate 18.07% -> 14.92%, sign-flip p=0.0013, CI [-5.29, -1.53] pp, damage taken
-24.57/run. That is precisely the pathology the owner watched.
Recommended for his own .env: TR_TFIL_COMMIT_ARRIVAL=1 and
TR_TFIL_NOREV_SPEED=4; leave TR_TFIL_COMMIT_MARGIN at 0 (weakest arm in both
blocks). Shipped default untouched: TR_MOVEMENT=strafe, all three new knobs
default off.
The owner's live-GUI report was correct on all four counts, and all four are
one bug: the commitment is cancelled by our own tile-boundary crossing
(96.1% of picks, 3793/3946, mean hold 5.06 ticks) while the bot is still
accelerating, and the picker is an unconstrained uniform draw over every
safe tile, so the new target can land in the mirror direction at |speed| < 4.
New knobs, all env-gated and default = today's behaviour (byte-for-byte
default parity guard re-run and green, 51 checks):
TR_TFIL_COMMIT_ARRIVAL hold the committed tile until we are ON it; the
tick knob becomes a MINIMUM dwell. 0 = shipped.
TR_TFIL_COMMIT_MARGIN leave only if the best alternative is at least
this much cooler on the same pathMaxHeat scale.
0 = shipped.
TR_TFIL_NOREV_SPEED while |speed| is below this, a mid-flight switch
may not take a tile >90 deg off the travel
direction. 0 = shipped. norevPool() never returns
an empty pool: with every candidate behind us it
takes the least-bad turn.
Offline gate (recorded DrussGT fixture, 20026 ticks): mean hold 4.1 -> 24.0
ticks, abandoned-before-arrival 92.8% -> 40.5%, committed tile actually
reached 3.3% -> 17.2%, opposite-direction slow mid-flight switches 394 -> 64
(-84%). 'TR_TFIL_TILE_REPLAN=off' alone - what cc11ede's arm B already tried -
only gets the hold to 13.6, which is why that A/B could not find this.
strafe is untouched: it imports only heatDecay/bulletMagScale/Pillar*, none
of which this touches. TR_MOVEMENT default stays strafe. Registered in
env_report.nim + knownEnvNames() + .env.example. Arms pre-registered in
docs/movement_campaign.md and tools/ab/arms_tfil_commit.txt.
Remove guns/bitbrain_net.nim (+README), test_bitbrain_net.nim,
measure_bitbrain_scaling.nim, rack id 17 and all of its plumbing in
selector.nim / ModularBot.nim / env_report.nim, the TR_BITBRAIN_NET switch
and the NEW-NETWORK TR_BITBRAIN_* knobs, and the BitBrainNet arm of
run_prediction_quality.nim.
With id 17 gone there is nothing to disambiguate, so the legacy namespace
becomes the ONLY one: TR_RACK_BITBRAIN always selects id 16 LEADGAIN and
every TR_BITBRAIN_<X> in the frozen 14-suffix alias set always means
TR_LEADGAIN_<X>. The alias layer and its [depr] line stay.
KEPT: the common_libs/bitbrain/ SBC library (learned_surfer imports
bitbrain/sbc), lead_gain at id 16 with env TR_LEADGAIN_* and log tag [lg],
and the c9b6753 crash fix (NumRackGuns widths + test_rack_stat_width).
Tombstone: docs/bitbrain_campaign.md ## RETIRED and one cross-reference line
in docs/gun_campaign.md. Shipped defaults unchanged: clean env -> rack
active 1v1 = PATTERN, movement default strafe.
A `#` preceded by whitespace and outside quotes now ends the value, so
`TR_DEBUG_DRAW=0 # hides the grid` resolves to `0` instead of the
whole tail. Values that are still not plain tokens (whitespace, `#`, an
unclosed quote) get one `[dotenv] WARNING` line naming file, key, raw value
and the fact that the reader falls back to its DEFAULT, instead of being
applied silently. Guard test 29 -> 47 checks.
56/60 opponent-level win deltas are non-negative and no opponent family
regresses reproducibly (the one WallAvoider loss reverses in Batch 2).
corr(dwins, d_hit_rate) = -0.09, corr(dwins, d_damage) = +0.38: the aggregate
win is a survival effect, but a per-opponent hit-rate gain does not predict a
per-opponent win gain - so hit rate stays an explanation, never a proxy.
- tfil is 4th of five on round wins, not last (ring is nominally 0.04 lower,
ns) - the Batch-1 commit message overstates one word; the correction is
recorded in the ledger rather than rewritten.
- the Batch-2 direct answer quoted three of four CIs excluding 0; it is four of
four ([+0.04,+0.63], [+0.16,+0.60], [+0.27,+0.89], [+0.22,+0.72]).
- added the cleanest aggression isolation of Batch 1 (ring - ring_notemp, same
engine and heat field, range weighting alone): +41.9 dmg/run, -0.11 wins/run,
+12.8 pp incoming hit rate at 236 vs 395 px.
Same frozen panel, same 3x3 design, new session on commit 8efa627 (no source file
changed since 1984a78, so the same code), 225 battles, 0 invalid runs. Arms:
tfil, strafe_notilt, strafe_325 + the tilt re-armed at 600px and 250px.
Paired vs tfil: strafe_325 +0.58 wins/run [CI +0.27,+0.89] 11/12 p=0.0063;
strafe_notilt +0.47 [+0.22,+0.72] 10/11 p=0.0117; tilt_600 +0.40 [+0.04,+0.76]
(sign test 8/11 p=0.23, sign-flip p=0.049); tilt_250 +0.38 [+0.07,+0.69] 10/12
p=0.039. Incoming hit rate -5.2..-6.9 pp with 0/15 opponents favouring tfil.
The range TARGET is not the lever: re-arming the tilt moved the achieved
distance from 459px (no steering) to 478px and 415px, and none of the three is
separable on wins. This overturns Batch 1's reading that the tilt costs wins -
the honest statement is that the tilt's win effect is below this design's
resolution. tfil reproduced to within 1.4 pp (40.7% -> 39.3% of rounds), so the
baseline itself is stable across sessions.
Ledger: Batch 2 section, the verbatim analyzer report, a data-driven
what-to-try-next, and the session log.
225 battles, one frozen binary, five env-only arms, the frozen panel, 0 invalid
runs. Paired per opponent vs the shipped tfil:
strafe_notilt wins/run +0.38 [CI +0.16,+0.60] 9/9 opponents p=0.0039
dmg/run -10.2 [CI -25.8,+5.5] p=0.61, MDE 20.4 (not detectable)
incoming hit rate 12.24% vs 18.17%, dmg taken 150 vs 200
strafe_325 wins/run +0.33 [CI +0.04,+0.63] 10/12 p=0.0386
ring dmg/run +31.2 [CI +11.5,+50.9] 13/15 p=0.0074, wins/run -0.04 (ns)
but hit rate 29.4% at 236 px: a damage/survival trade, not a win
ring_notemp indistinguishable from tfil on both primaries
Round wins in this harness are survival wins (in 216/219 attributable runs the
win count equals the rounds the opponent died in), and the winner takes ~1/3
fewer hits while fighting ~74 px farther out. The shipped tfil is last of five
on wins: the DrussGT-only picture did not generalize.
Also: tournament_analyze.py now prints BOTH readings of the pre-registered
'while the other does not go down' clause (strict: nothing is better;
substantive: the two strafe arms and ring are better on one metric each).
The repo's first multi-opponent gun measurement. Adds tools/ab/gauntlet_run.sh
(per-opponent A/B over the legacy roster, subject = frozen ModularBot),
tools/ab/gauntlet_analyze.py (paired per-opponent deltas, cross-opponent sign
test, style split, MDE) and the arm/opponent fixtures.
Result: BitBrain does NOT generalize beyond DrussGT. 32 opponents x 2 arms x
3 runs x 5 rounds = 192 battles / 960 rounds, 0 failed, 0 retries: damage/run
214.5 (pattern) vs 210.9 (bb), sign-flip p=0.53; round wins 237/480 vs 239/480,
p=0.91. Sign test: bb better on 13/32 opponents (damage). The DrussGT-only
penalty does not carry. The owner's 'killer vs regular movers' sub-claim is not
supported: regular bucket +1.3 dmg/run vs dodgers -0.2 (MW p=0.85), and the
measured movement predictability does not correlate with the delta.
Adds common_libs/tests/measure_melee_bitbrain_ab.nim (+ .sh driver, .py analyzer,
committed per-run fixtures) and docs/melee_bitbrain_ab.md.
Experiment: 4-bot Free-For-All (ModularBot + WaveSurfer + PatternMover +
RandomMover), 4 arms x 16 runs x 7 rounds, frozen ModularBot from git archive
HEAD (commit 0f5cfe3, binary 11bba27), shipped tfil movement in every run.
Arms differ only in the gun rack: pattern (shipped), bb_round, bb_ret, bb_learn.
Result: NOT DETECTABLE. Score (server round score = damage + survival bonus)
differs by -63..+33 pts (perm p=0.16-0.71) against an MDE of 151 (~5.1%).
Every arm finishes rank 1. Round wins hint BitBrain's way (112/112 and 111/112
vs 109/112) but p=0.225 (MW 0.080), half the 0.40-win MDE.
Liveness proven: rack boot lines flip (rack active melee = PATTERN / BITBRAIN),
every run faced 3 distinct targets and ~66-69 target changes, and the bb arms
logged one [bb-reset] reason=target_change per switch. The melee premise was
exercised; the fast adaptation bought no measurable score edge at this sample.
Live 3-arm x 15-run x 7-round A/B vs real DrussGT on commit 0f5cfe37.
Primary: surf ties strafe on round wins (37/105) and damage (255 vs 250/run),
both below tfil (45/105, 293/run; damage p=0.003). Incoming hit rate: surf
13.51% (worst) vs strafe 9.40% (best) and tfil 10.40%. So the plain surfer does
NOT dodge better and does NOT win more. Also records the j107 trap: strafe
dodges best yet wins fewer rounds than tfil. Next step: range/aggression A/B,
not a BitBrain upgrade.
4 arms x 15 runs x 7 rounds (60 battles, 0 failed) vs real DrussGT on the
shipped TFIL default, frozen at ed25ce2. bb_id (gain 1.0 identity) is
statistically indistinguishable from shipped Pattern -> plumbing validity
check passes. No BitBrain arm beats TMHorizon or Pattern: bb_learn (the config
the owner likely ran) is the worst arm (276 dmg/run, 39/105 wins), the only
comparison at alpha=0.05 is Pattern beating it on damage. Learned gains
(>=1.0, gated >=300px) over-lead and lose 1.61pp of hit rate at 300-450px.
MDE 29.5 dmg/run, 1.235 wins/run; a 6-4-sized effect needs ~39 runs/arm.