Commit Graph

91 Commits

Author SHA1 Message Date
SirStone 0e7e124c7f j144 ledger: the live A/B result, 600 battles, two independent blocks
Pooled verdict on the pre-registered rule is NOT DISTINGUISHABLE: the primary
cross-opponent sign test on round wins is 11/14, p=0.05737 for 'arrive', over the
0.05 line. Block 1 alone passed (p=0.01294); block 2 alone did not (p=0.0654).
Every delta is positive in both blocks for all three arms and damage/run is UP
on all three, so this reads as an under-powered null at the MDE boundary
(MDE 0.31 wins/run, observed 0.28) - but the verdict layer is not
reinterpreted, exactly as gate v1 was not.

The MECHANISM is established cleanly and replicates in both blocks: incoming hit
rate 18.07% -> 14.92%, sign-flip p=0.0013, CI [-5.29, -1.53] pp, damage taken
-24.57/run. That is precisely the pathology the owner watched.

Recommended for his own .env: TR_TFIL_COMMIT_ARRIVAL=1 and
TR_TFIL_NOREV_SPEED=4; leave TR_TFIL_COMMIT_MARGIN at 0 (weakest arm in both
blocks). Shipped default untouched: TR_MOVEMENT=strafe, all three new knobs
default off.
2026-09-26 21:43:03 +02:00
SirStone d2005abee9 j144 TFIL: arrival-based commitment + hysteresis + no mid-flight reversal
The owner's live-GUI report was correct on all four counts, and all four are
one bug: the commitment is cancelled by our own tile-boundary crossing
(96.1% of picks, 3793/3946, mean hold 5.06 ticks) while the bot is still
accelerating, and the picker is an unconstrained uniform draw over every
safe tile, so the new target can land in the mirror direction at |speed| < 4.

New knobs, all env-gated and default = today's behaviour (byte-for-byte
default parity guard re-run and green, 51 checks):
  TR_TFIL_COMMIT_ARRIVAL  hold the committed tile until we are ON it; the
                          tick knob becomes a MINIMUM dwell. 0 = shipped.
  TR_TFIL_COMMIT_MARGIN   leave only if the best alternative is at least
                          this much cooler on the same pathMaxHeat scale.
                          0 = shipped.
  TR_TFIL_NOREV_SPEED     while |speed| is below this, a mid-flight switch
                          may not take a tile >90 deg off the travel
                          direction. 0 = shipped. norevPool() never returns
                          an empty pool: with every candidate behind us it
                          takes the least-bad turn.

Offline gate (recorded DrussGT fixture, 20026 ticks): mean hold 4.1 -> 24.0
ticks, abandoned-before-arrival 92.8% -> 40.5%, committed tile actually
reached 3.3% -> 17.2%, opposite-direction slow mid-flight switches 394 -> 64
(-84%). 'TR_TFIL_TILE_REPLAN=off' alone - what cc11ede's arm B already tried -
only gets the hold to 13.6, which is why that A/B could not find this.

strafe is untouched: it imports only heatDecay/bulletMagScale/Pillar*, none
of which this touches. TR_MOVEMENT default stays strafe. Registered in
env_report.nim + knownEnvNames() + .env.example. Arms pre-registered in
docs/movement_campaign.md and tools/ab/arms_tfil_commit.txt.
2026-09-26 21:07:09 +02:00
SirStone 5e32ec16df j142 retire the ADE+SBC gun (rack id 17): the owner watched it, it does not learn, throw it away
Remove guns/bitbrain_net.nim (+README), test_bitbrain_net.nim,
measure_bitbrain_scaling.nim, rack id 17 and all of its plumbing in
selector.nim / ModularBot.nim / env_report.nim, the TR_BITBRAIN_NET switch
and the NEW-NETWORK TR_BITBRAIN_* knobs, and the BitBrainNet arm of
run_prediction_quality.nim.

With id 17 gone there is nothing to disambiguate, so the legacy namespace
becomes the ONLY one: TR_RACK_BITBRAIN always selects id 16 LEADGAIN and
every TR_BITBRAIN_<X> in the frozen 14-suffix alias set always means
TR_LEADGAIN_<X>. The alias layer and its [depr] line stay.

KEPT: the common_libs/bitbrain/ SBC library (learned_surfer imports
bitbrain/sbc), lead_gain at id 16 with env TR_LEADGAIN_* and log tag [lg],
and the c9b6753 crash fix (NumRackGuns widths + test_rack_stat_width).

Tombstone: docs/bitbrain_campaign.md ## RETIRED and one cross-reference line
in docs/gun_campaign.md. Shipped defaults unchanged: clean env -> rack
active 1v1 = PATTERN, movement default strafe.
2026-09-26 19:20:23 +02:00
SirStone 0dc5552c73 j139 dotenv: strip trailing inline comments and warn on non-token values
A `#` preceded by whitespace and outside quotes now ends the value, so
`TR_DEBUG_DRAW=0      # hides the grid` resolves to `0` instead of the
whole tail. Values that are still not plain tokens (whitespace, `#`, an
unclosed quote) get one `[dotenv] WARNING` line naming file, key, raw value
and the fact that the reader falls back to its DEFAULT, instead of being
applied silently. Guard test 29 -> 47 checks.
2026-09-26 15:18:04 +02:00
SirStone 189d990812 j136 [modules] inventory: also list the always-on body API 2026-09-26 14:48:43 +02:00
SirStone 8f44140783 j136 add TR_MODULE_* on/off switches + single [modules] boot inventory 2026-09-26 14:46:29 +02:00
SirStone 3903ec71b4 j134 ledger: note phantom_meteor is dead code (not selectable), out of scope 2026-09-26 13:01:36 +02:00
SirStone 3237b65f40 j134 ledger: propagated fire fix to all 5 movers via one shared fire_tracker (TR_FIRE_FIX), per-mover catch table 98.888->100, live tick alignment measured (correction belongs on event_tick+2) 2026-09-26 12:59:35 +02:00
SirStone 3bee4353b2 j133 ledger: appended 'Missed fires + the label question' (catch 98.888->100%, 746/67065 shots were blind: 456 by the server +3*power bonus, 290 by our own same-tick damage; exact label still -0.230, state-conditional model -0.347 -> the observable state is the constraint) + report fixtures 2026-09-26 12:22:31 +02:00
SirStone cc332138b3 j131 learned movement: real bullet-endpoint resolution (TR_LEARNED_REAL_EVENTS, default off) + exact-geometry Gate A/B (inversion NOT fixed; state still the constraint) 2026-09-26 11:54:11 +02:00
SirStone 8dd9b3b3b5 j130 learned movement outcome label: live panel results (no arm beats strafe; label swap is a dead heat; gap is information) + Gate A report 2026-09-26 11:35:47 +02:00
SirStone 61def1c3e9 j130 learned movement: outcome label (P(hit|state,g)) mode + Gate A; pre-registered outcome arms 2026-09-26 11:27:03 +02:00
SirStone 752d3a3829 j129: virtual-bullet debug overlay (TR_VBULLET_DEBUG, default off) - travelled path, predicted aim ring, hit/miss vector; shared turret colour table 2026-09-26 11:02:39 +02:00
SirStone a5c893ace6 learned movement: record the pre-registered prediction as partly wrong (right on wins, wrong on the hit-rate mechanism) 2026-09-26 10:56:02 +02:00
SirStone 4f332dcb95 learned movement (SBC): live panel results and the direct answer (module is a wash vs strafe on wins; state conditioning is a CI-separated hit-rate gain vs the same mover without it; the label is the wrong quantity) 2026-09-26 10:55:14 +02:00
SirStone ee98827a62 learned movement: danger-map alignment diagnostic (corr(P(arrival bin), P(hit|bin)) = -0.342) in the gate + ledger 2026-09-26 10:45:26 +02:00
SirStone b4e54fc3cd j127: allocation batch results - forced shares do not beat pattern; allocation not a lever (closed axis) 2026-09-26 10:45:10 +02:00
SirStone a436e9f16a j128 learned movement: state-conditional counted-SBC wave danger (TR_MOVEMENT=learned, default-off), offline gate + pre-registered panel arms 2026-09-26 10:42:10 +02:00
SirStone 6fd5fe6328 j127: forced-share allocator (TR_RACK_SHARE, default-off) + allocation batch pre-registration 2026-09-26 10:13:22 +02:00
SirStone 2d8b7d7875 j126: read bot settings from a .env file (default ./.env, --env-file flag, TR_ENV_FILE); file wins over shell leftovers, boot report labels (source: .env) 2026-09-26 09:08:39 +02:00
SirStone 8e109bae0b closing j125: verify movement ship end-to-end (clean build SuccessX, full 16-test guard suite all counts match, real default+revert battles) and restructure gun/movement ledgers outcome-first 2026-09-26 05:06:49 +02:00
SirStone cf77d0d647 j124: spinner + fair-melee results doc; analyzer convergence tests 2026-09-26 05:05:09 +02:00
SirStone 5fe28574ab gun j123 Batch 2 outcome + phase-2 final answer: len6 does not replicate (wash on 33 opponents); the shipped Pattern is already tuned on its match-length/history parameters; no default changed 2026-09-26 04:54:40 +02:00
SirStone 1b59b65810 gun j123 Batch 1 outcome + Batch 2 prereg: len6 is the only arm to beat pattern on wins (15 opp), but all 5 arms are wins-positive incl. the bearing-invariant rad_offset; pre-register the 33-opponent confirmation with a structural control 2026-09-26 04:22:49 +02:00
SirStone 2a98aba91b gun j123 Task A+B prereg: expose Pattern's TR_PATTERN_LEN/TR_PATTERN_DEPTH (default parity) and pre-register the 6-arm match-parameter sweep 2026-09-26 04:01:51 +02:00
SirStone 467e07a6a4 gun ledger: fix the arm count in the final answer (six configurations, not seven) 2026-09-26 03:55:32 +02:00
SirStone 2a5ea5c17b gun j121 Batch 2 outcome: the wider panel kills the damage hint (bitbrain wash, tmhorizon WORSE); onlyPattern confirmed; phase closed 2026-09-26 03:54:52 +02:00
SirStone 3fd6db97e7 movement ship: gate v2 fresh-data primary passed (sign-flip p=0.045, CI [+0.02,+0.58]); default TR_MOVEMENT flipped to strafe 2026-09-26 03:19:33 +02:00
SirStone 087e26955f gun j121 Batch 1 outcome: onlyPattern CONFIRMED across the 15-opponent panel; pre-register Batch 2 wider-panel calibration 2026-09-26 02:58:41 +02:00
SirStone 514674886d movement gate v2: pre-register the sign-flip primary test on fresh 300-battle data (before any battle) 2026-09-26 02:44:22 +02:00
SirStone feefc1912e movement final: confirmation result and ship decision (gate failed on the sign-test leg; default NOT flipped) 2026-09-26 02:41:08 +02:00
SirStone 1d8143a15e gun j121 Batch 1: pre-register the 6-arm Pattern-panel tournament and the decision rules 2026-09-26 02:22:38 +02:00
SirStone ff03e81591 movement final: pre-register the ship criterion before the confirmation battles 2026-09-26 02:21:09 +02:00
SirStone 47ce244740 movement j119 Batch 3+4 results: the reversal/dwell and heat-field axes are a clean negative; nothing beats strafe 2026-09-26 02:18:22 +02:00
SirStone 1256357f07 movement j119 Batch 3+4: pre-register the reversal/dwell and heat-field arms 2026-09-26 01:40:03 +02:00
SirStone c07e6d9599 movement ledger: the post-hoc structure of the strafe win (60 points, both batches)
56/60 opponent-level win deltas are non-negative and no opponent family
regresses reproducibly (the one WallAvoider loss reverses in Batch 2).
corr(dwins, d_hit_rate) = -0.09, corr(dwins, d_damage) = +0.38: the aggregate
win is a survival effect, but a per-opponent hit-rate gain does not predict a
per-opponent win gain - so hit rate stays an explanation, never a proxy.
2026-09-26 01:33:24 +02:00
SirStone 44e2d191e3 movement ledger: fix two claims in the Batch-1/2 prose
- tfil is 4th of five on round wins, not last (ring is nominally 0.04 lower,
  ns) - the Batch-1 commit message overstates one word; the correction is
  recorded in the ledger rather than rewritten.
- the Batch-2 direct answer quoted three of four CIs excluding 0; it is four of
  four ([+0.04,+0.63], [+0.16,+0.60], [+0.27,+0.89], [+0.22,+0.72]).
- added the cleanest aggression isolation of Batch 1 (ring - ring_notemp, same
  engine and heat field, range weighting alone): +41.9 dmg/run, -0.11 wins/run,
  +12.8 pp incoming hit rate at 236 vs 395 px.
2026-09-26 01:32:28 +02:00
SirStone 7d3645da5a movement Batch 2: the strafe win over shipped tfil replicates; the range target decides nothing
Same frozen panel, same 3x3 design, new session on commit 8efa627 (no source file
changed since 1984a78, so the same code), 225 battles, 0 invalid runs. Arms:
tfil, strafe_notilt, strafe_325 + the tilt re-armed at 600px and 250px.

Paired vs tfil: strafe_325 +0.58 wins/run [CI +0.27,+0.89] 11/12 p=0.0063;
strafe_notilt +0.47 [+0.22,+0.72] 10/11 p=0.0117; tilt_600 +0.40 [+0.04,+0.76]
(sign test 8/11 p=0.23, sign-flip p=0.049); tilt_250 +0.38 [+0.07,+0.69] 10/12
p=0.039. Incoming hit rate -5.2..-6.9 pp with 0/15 opponents favouring tfil.

The range TARGET is not the lever: re-arming the tilt moved the achieved
distance from 459px (no steering) to 478px and 415px, and none of the three is
separable on wins. This overturns Batch 1's reading that the tilt costs wins -
the honest statement is that the tilt's win effect is below this design's
resolution. tfil reproduced to within 1.4 pp (40.7% -> 39.3% of rounds), so the
baseline itself is stable across sessions.

Ledger: Batch 2 section, the verbatim analyzer report, a data-driven
what-to-try-next, and the session log.
2026-09-26 01:31:08 +02:00
SirStone 07766303f5 movement Batch 1: pure strafe (range tilt OFF) beats the shipped tfil on round wins across a 15-opponent panel
225 battles, one frozen binary, five env-only arms, the frozen panel, 0 invalid
runs. Paired per opponent vs the shipped tfil:

  strafe_notilt  wins/run +0.38  [CI +0.16,+0.60]  9/9 opponents p=0.0039
                 dmg/run  -10.2  [CI -25.8,+5.5]   p=0.61, MDE 20.4 (not detectable)
                 incoming hit rate 12.24% vs 18.17%, dmg taken 150 vs 200
  strafe_325     wins/run +0.33  [CI +0.04,+0.63]  10/12 p=0.0386
  ring           dmg/run  +31.2  [CI +11.5,+50.9]  13/15 p=0.0074, wins/run -0.04 (ns)
                 but hit rate 29.4% at 236 px: a damage/survival trade, not a win
  ring_notemp    indistinguishable from tfil on both primaries

Round wins in this harness are survival wins (in 216/219 attributable runs the
win count equals the rounds the opponent died in), and the winner takes ~1/3
fewer hits while fighting ~74 px farther out. The shipped tfil is last of five
on wins: the DrussGT-only picture did not generalize.

Also: tournament_analyze.py now prints BOTH readings of the pre-registered
'while the other does not go down' clause (strict: nothing is better;
substantive: the two strafe arms and ring are better on one metric each).
2026-09-26 01:13:29 +02:00
SirStone 8efa627c05 gauntlet: BitBrain vs Pattern across 32 legacy opponents (does not generalize)
The repo's first multi-opponent gun measurement. Adds tools/ab/gauntlet_run.sh
(per-opponent A/B over the legacy roster, subject = frozen ModularBot),
tools/ab/gauntlet_analyze.py (paired per-opponent deltas, cross-opponent sign
test, style split, MDE) and the arm/opponent fixtures.

Result: BitBrain does NOT generalize beyond DrussGT. 32 opponents x 2 arms x
3 runs x 5 rounds = 192 battles / 960 rounds, 0 failed, 0 retries: damage/run
214.5 (pattern) vs 210.9 (bb), sign-flip p=0.53; round wins 237/480 vs 239/480,
p=0.91. Sign test: bb better on 13/32 opponents (damage). The DrussGT-only
penalty does not carry. The owner's 'killer vs regular movers' sub-claim is not
supported: regular bucket +1.3 dmg/run vs dodgers -0.2 (MW p=0.85), and the
measured movement predictability does not correlate with the delta.
2026-09-26 00:55:57 +02:00
SirStone 1984a780f4 melee A/B doc: correct the per-arm [bb-reset] counts (perRound 68.75, retained 68.19, decay 66.75) 2026-09-26 00:43:04 +02:00
SirStone da4a971ca9 MELEE A/B: BitBrain vs shipped Pattern rack — repo's first melee measurement (null)
Adds common_libs/tests/measure_melee_bitbrain_ab.nim (+ .sh driver, .py analyzer,
committed per-run fixtures) and docs/melee_bitbrain_ab.md.

Experiment: 4-bot Free-For-All (ModularBot + WaveSurfer + PatternMover +
RandomMover), 4 arms x 16 runs x 7 rounds, frozen ModularBot from git archive
HEAD (commit 0f5cfe3, binary 11bba27), shipped tfil movement in every run.
Arms differ only in the gun rack: pattern (shipped), bb_round, bb_ret, bb_learn.

Result: NOT DETECTABLE. Score (server round score = damage + survival bonus)
differs by -63..+33 pts (perm p=0.16-0.71) against an MDE of 151 (~5.1%).
Every arm finishes rank 1. Round wins hint BitBrain's way (112/112 and 111/112
vs 109/112) but p=0.225 (MW 0.080), half the 0.40-win MDE.

Liveness proven: rack boot lines flip (rack active melee = PATTERN / BITBRAIN),
every run faced 3 distinct targets and ~66-69 target changes, and the bb arms
logged one [bb-reset] reason=target_change per switch. The melee premise was
exercised; the fast adaptation bought no measurable score edge at this sample.
2026-09-26 00:35:44 +02:00
SirStone 02b691dd5f docs: add per-run damage/wins series and liveness line to surfer_wiring_ab 2026-09-26 00:29:34 +02:00
SirStone 6d6648ccf8 docs: wave surfer wiring A/B vs TFIL and STRAFE (surf does not beat the shipped default)
Live 3-arm x 15-run x 7-round A/B vs real DrussGT on commit 0f5cfe37.
Primary: surf ties strafe on round wins (37/105) and damage (255 vs 250/run),
both below tfil (45/105, 293/run; damage p=0.003). Incoming hit rate: surf
13.51% (worst) vs strafe 9.40% (best) and tfil 10.40%. So the plain surfer does
NOT dodge better and does NOT win more. Also records the j107 trap: strafe
dodges best yet wins fewer rounds than tfil. Next step: range/aggression A/B,
not a BitBrain upgrade.
2026-09-26 00:27:34 +02:00
SirStone 4829f9ca13 BitBrain vs TMHorizon vs Pattern: live A/B on shipped TFIL (null result)
4 arms x 15 runs x 7 rounds (60 battles, 0 failed) vs real DrussGT on the
shipped TFIL default, frozen at ed25ce2. bb_id (gain 1.0 identity) is
statistically indistinguishable from shipped Pattern -> plumbing validity
check passes. No BitBrain arm beats TMHorizon or Pattern: bb_learn (the config
the owner likely ran) is the worst arm (276 dmg/run, 39/105 wins), the only
comparison at alpha=0.05 is Pattern beating it on damage. Learned gains
(>=1.0, gated >=300px) over-lead and lose 1.61pp of hit rate at 300-450px.
MDE 29.5 dmg/run, 1.235 wins/run; a 6-4-sized effect needs ~39 runs/arm.
2026-09-25 23:55:21 +02:00
SirStone 99cf9e5c82 tfil heat/pillar A/B: record the owner's decision to keep the virtual pillar removed 2026-09-25 22:24:47 +02:00
SirStone 48f38b80e7 TFIL heat-time + virtual pillar: live A/B (6 arms x 70 rounds) - neither change beats the pre-change mover
Runs the pre-registered A/B for the two movement changes in HEAD: the
time-indexed bullet heat (TR_TFIL_HEAT_TIME, fca8993) and the removal of the
invented virtual centre pillar (d0750ab). One frozen binary from HEAD vs real
DrussGT: 6 arms x 10 runs x 7 rounds = 60 battles, 420 rounds, 0 failed.

Judged on damage/run and ROUND WINS only (hit rate and hits-taken are context):
hit rate would have inverted the verdict again - tau3 has the best pooled hit
rate of all arms (11.56%) and the fewest round wins (20/70).

RESULT (vs the reconstructed pre-change mover "old"):
  heat-time HURTS. tau3/tau5/tau9 lose 1.3-1.7 wins/run (p=0.0010-0.0125) and
  deal 22-38 less damage/run (p=0.004-0.047); tau15 is a wash on wins (p=0.64)
  and 22 damage/run lower (p=0.046). Nothing improves either metric.
  pillar removal does nothing measurable. old vs pillaoff: +5.7 damage/run
  (p=0.71), +0.5 wins/run (35 vs 30, p=0.43), 30.8 MORE damage taken/run
  without the pillar (p=0.040). The mechanism check proves the knob works
  (centre-box occupancy 0.09% -> 2.37%, p<0.0001; range 469 -> 443 px,
  p=0.0002), so this is a real behaviour change that buys nothing. At n=10 the
  pillar contrast is inside the MDE (33 damage/run, 1.2 wins/run), so this is
  not a proven regression.

Flags that the shipped default (pillar removed) should be reverted to the
TR_TFIL_PILLAR_ON behaviour; heat-time stays off.

Adds tools/ab/arms_heat_pillar.txt and tools/ab/ab_mechanism.py (per-tick
mechanism check: central-box occupancy, range distribution, live enemy-bullet
proximity) plus the captured summary/report fixtures.
2026-09-25 22:23:11 +02:00
SirStone f58d65d2e8 TMComposites gate: per-gun confidence faithful for 3 guns; no pair composes
Adds a per-sample intrinsic-confidence field (GunPrediction.confidence,
threaded through FeedbackEvent/VirtualBullet, populated by Pattern, DecayGF,
KNN, GuessFactor, Tsetlin, TMHorizon) and an offline recorder + analyzer that
reproduce the paper's Figure 2 per gun and its Eq-8 composite.

Measured on 3 held-out tr-bridge DrussGT battles (33k ticks, ~133k samples/gun):
- FAITHFUL: DecayGF (rho +0.133), KNN (+0.090), Pattern (+0.064, weak).
- GuessFactor is ANTI-faithful (rho -0.067); Tsetlin c_max is useless (0.001).
- No pair of guns specialises complementarily: the same gun dominates both
  high-confidence slices in every pair.
- Eq-8 alpha-normalised confidence-weighted composite: 18.41% vs Pattern
  20.45% (McNemar p=3.1e-126). Faithful-only variant 18.68%, still loses.
  Shuffle control passes weakly (composite > shuffle, p=4e-14) so ~0.7pp of
  competence is real but ~2pp short. Offline veto: design is dead.

See docs/tmcomposites_gate.md.
2026-09-25 22:04:39 +02:00
SirStone 4270136948 docs: correct counted-SBC decay amortised cost figure 2026-09-25 08:43:48 +02:00
SirStone d85ff53d34 State-window gate: a temporal window of wave-relative states does NOT beat a single state
Re-runs the SBC coincidence premise as a cheap veto test on the 70-battle
live-vs-DrussGT corpus (/tmp/tfil_ab2).  Defines the wave-relative state
(lat/vlat/toa/room/turn, 5/7.9/10 bits at Q=2/3/4), quantises the miss offset
at the bullet's arrival into 7 bins, and sweeps window length K in
{1,4,8,16,32,48} with an interpolated suffix-backoff model under a BY-BATTLE
70/30 split (3 seeds).

Result: NO.  On the pre-fire frame the window is worse than the single
fire-tick state at every K/Q/A (e.g. K=8 costs +0.35..+0.46 bits).  On the
during-flight frame the entire apparent gain is the later decision tick, not
the window; the single state alone drops 2.70 -> 1.31 bits as K goes 1 -> 32.
The shuffle-order control confirms recency matters but the windows do not:
by Q=4/K=8 they average ~1 observation and never recur.  The single state
survives as a strong predictor (log-loss 2.346 vs 2.698 majority; bin
accuracy 0.409 vs 0.235).

Gate only: no gun, no live-win claim.
2026-09-25 08:43:36 +02:00