56/60 opponent-level win deltas are non-negative and no opponent family
regresses reproducibly (the one WallAvoider loss reverses in Batch 2).
corr(dwins, d_hit_rate) = -0.09, corr(dwins, d_damage) = +0.38: the aggregate
win is a survival effect, but a per-opponent hit-rate gain does not predict a
per-opponent win gain - so hit rate stays an explanation, never a proxy.
- tfil is 4th of five on round wins, not last (ring is nominally 0.04 lower,
ns) - the Batch-1 commit message overstates one word; the correction is
recorded in the ledger rather than rewritten.
- the Batch-2 direct answer quoted three of four CIs excluding 0; it is four of
four ([+0.04,+0.63], [+0.16,+0.60], [+0.27,+0.89], [+0.22,+0.72]).
- added the cleanest aggression isolation of Batch 1 (ring - ring_notemp, same
engine and heat field, range weighting alone): +41.9 dmg/run, -0.11 wins/run,
+12.8 pp incoming hit rate at 236 vs 395 px.
Same frozen panel, same 3x3 design, new session on commit 8efa627 (no source file
changed since 1984a78, so the same code), 225 battles, 0 invalid runs. Arms:
tfil, strafe_notilt, strafe_325 + the tilt re-armed at 600px and 250px.
Paired vs tfil: strafe_325 +0.58 wins/run [CI +0.27,+0.89] 11/12 p=0.0063;
strafe_notilt +0.47 [+0.22,+0.72] 10/11 p=0.0117; tilt_600 +0.40 [+0.04,+0.76]
(sign test 8/11 p=0.23, sign-flip p=0.049); tilt_250 +0.38 [+0.07,+0.69] 10/12
p=0.039. Incoming hit rate -5.2..-6.9 pp with 0/15 opponents favouring tfil.
The range TARGET is not the lever: re-arming the tilt moved the achieved
distance from 459px (no steering) to 478px and 415px, and none of the three is
separable on wins. This overturns Batch 1's reading that the tilt costs wins -
the honest statement is that the tilt's win effect is below this design's
resolution. tfil reproduced to within 1.4 pp (40.7% -> 39.3% of rounds), so the
baseline itself is stable across sessions.
Ledger: Batch 2 section, the verbatim analyzer report, a data-driven
what-to-try-next, and the session log.
225 battles, one frozen binary, five env-only arms, the frozen panel, 0 invalid
runs. Paired per opponent vs the shipped tfil:
strafe_notilt wins/run +0.38 [CI +0.16,+0.60] 9/9 opponents p=0.0039
dmg/run -10.2 [CI -25.8,+5.5] p=0.61, MDE 20.4 (not detectable)
incoming hit rate 12.24% vs 18.17%, dmg taken 150 vs 200
strafe_325 wins/run +0.33 [CI +0.04,+0.63] 10/12 p=0.0386
ring dmg/run +31.2 [CI +11.5,+50.9] 13/15 p=0.0074, wins/run -0.04 (ns)
but hit rate 29.4% at 236 px: a damage/survival trade, not a win
ring_notemp indistinguishable from tfil on both primaries
Round wins in this harness are survival wins (in 216/219 attributable runs the
win count equals the rounds the opponent died in), and the winner takes ~1/3
fewer hits while fighting ~74 px farther out. The shipped tfil is last of five
on wins: the DrussGT-only picture did not generalize.
Also: tournament_analyze.py now prints BOTH readings of the pre-registered
'while the other does not go down' clause (strict: nothing is better;
substantive: the two strafe arms and ring are better on one metric each).
The repo's first multi-opponent gun measurement. Adds tools/ab/gauntlet_run.sh
(per-opponent A/B over the legacy roster, subject = frozen ModularBot),
tools/ab/gauntlet_analyze.py (paired per-opponent deltas, cross-opponent sign
test, style split, MDE) and the arm/opponent fixtures.
Result: BitBrain does NOT generalize beyond DrussGT. 32 opponents x 2 arms x
3 runs x 5 rounds = 192 battles / 960 rounds, 0 failed, 0 retries: damage/run
214.5 (pattern) vs 210.9 (bb), sign-flip p=0.53; round wins 237/480 vs 239/480,
p=0.91. Sign test: bb better on 13/32 opponents (damage). The DrussGT-only
penalty does not carry. The owner's 'killer vs regular movers' sub-claim is not
supported: regular bucket +1.3 dmg/run vs dodgers -0.2 (MW p=0.85), and the
measured movement predictability does not correlate with the delta.
Adds common_libs/tests/measure_melee_bitbrain_ab.nim (+ .sh driver, .py analyzer,
committed per-run fixtures) and docs/melee_bitbrain_ab.md.
Experiment: 4-bot Free-For-All (ModularBot + WaveSurfer + PatternMover +
RandomMover), 4 arms x 16 runs x 7 rounds, frozen ModularBot from git archive
HEAD (commit 0f5cfe3, binary 11bba27), shipped tfil movement in every run.
Arms differ only in the gun rack: pattern (shipped), bb_round, bb_ret, bb_learn.
Result: NOT DETECTABLE. Score (server round score = damage + survival bonus)
differs by -63..+33 pts (perm p=0.16-0.71) against an MDE of 151 (~5.1%).
Every arm finishes rank 1. Round wins hint BitBrain's way (112/112 and 111/112
vs 109/112) but p=0.225 (MW 0.080), half the 0.40-win MDE.
Liveness proven: rack boot lines flip (rack active melee = PATTERN / BITBRAIN),
every run faced 3 distinct targets and ~66-69 target changes, and the bb arms
logged one [bb-reset] reason=target_change per switch. The melee premise was
exercised; the fast adaptation bought no measurable score edge at this sample.
Live 3-arm x 15-run x 7-round A/B vs real DrussGT on commit 0f5cfe37.
Primary: surf ties strafe on round wins (37/105) and damage (255 vs 250/run),
both below tfil (45/105, 293/run; damage p=0.003). Incoming hit rate: surf
13.51% (worst) vs strafe 9.40% (best) and tfil 10.40%. So the plain surfer does
NOT dodge better and does NOT win more. Also records the j107 trap: strafe
dodges best yet wins fewer rounds than tfil. Next step: range/aggression A/B,
not a BitBrain upgrade.
4 arms x 15 runs x 7 rounds (60 battles, 0 failed) vs real DrussGT on the
shipped TFIL default, frozen at ed25ce2. bb_id (gain 1.0 identity) is
statistically indistinguishable from shipped Pattern -> plumbing validity
check passes. No BitBrain arm beats TMHorizon or Pattern: bb_learn (the config
the owner likely ran) is the worst arm (276 dmg/run, 39/105 wins), the only
comparison at alpha=0.05 is Pattern beating it on damage. Learned gains
(>=1.0, gated >=300px) over-lead and lose 1.61pp of hit rate at 300-450px.
MDE 29.5 dmg/run, 1.235 wins/run; a 6-4-sized effect needs ~39 runs/arm.
Runs the pre-registered A/B for the two movement changes in HEAD: the
time-indexed bullet heat (TR_TFIL_HEAT_TIME, fca8993) and the removal of the
invented virtual centre pillar (d0750ab). One frozen binary from HEAD vs real
DrussGT: 6 arms x 10 runs x 7 rounds = 60 battles, 420 rounds, 0 failed.
Judged on damage/run and ROUND WINS only (hit rate and hits-taken are context):
hit rate would have inverted the verdict again - tau3 has the best pooled hit
rate of all arms (11.56%) and the fewest round wins (20/70).
RESULT (vs the reconstructed pre-change mover "old"):
heat-time HURTS. tau3/tau5/tau9 lose 1.3-1.7 wins/run (p=0.0010-0.0125) and
deal 22-38 less damage/run (p=0.004-0.047); tau15 is a wash on wins (p=0.64)
and 22 damage/run lower (p=0.046). Nothing improves either metric.
pillar removal does nothing measurable. old vs pillaoff: +5.7 damage/run
(p=0.71), +0.5 wins/run (35 vs 30, p=0.43), 30.8 MORE damage taken/run
without the pillar (p=0.040). The mechanism check proves the knob works
(centre-box occupancy 0.09% -> 2.37%, p<0.0001; range 469 -> 443 px,
p=0.0002), so this is a real behaviour change that buys nothing. At n=10 the
pillar contrast is inside the MDE (33 damage/run, 1.2 wins/run), so this is
not a proven regression.
Flags that the shipped default (pillar removed) should be reverted to the
TR_TFIL_PILLAR_ON behaviour; heat-time stays off.
Adds tools/ab/arms_heat_pillar.txt and tools/ab/ab_mechanism.py (per-tick
mechanism check: central-box occupancy, range distribution, live enemy-bullet
proximity) plus the captured summary/report fixtures.
Adds a per-sample intrinsic-confidence field (GunPrediction.confidence,
threaded through FeedbackEvent/VirtualBullet, populated by Pattern, DecayGF,
KNN, GuessFactor, Tsetlin, TMHorizon) and an offline recorder + analyzer that
reproduce the paper's Figure 2 per gun and its Eq-8 composite.
Measured on 3 held-out tr-bridge DrussGT battles (33k ticks, ~133k samples/gun):
- FAITHFUL: DecayGF (rho +0.133), KNN (+0.090), Pattern (+0.064, weak).
- GuessFactor is ANTI-faithful (rho -0.067); Tsetlin c_max is useless (0.001).
- No pair of guns specialises complementarily: the same gun dominates both
high-confidence slices in every pair.
- Eq-8 alpha-normalised confidence-weighted composite: 18.41% vs Pattern
20.45% (McNemar p=3.1e-126). Faithful-only variant 18.68%, still loses.
Shuffle control passes weakly (composite > shuffle, p=4e-14) so ~0.7pp of
competence is real but ~2pp short. Offline veto: design is dead.
See docs/tmcomposites_gate.md.
Re-runs the SBC coincidence premise as a cheap veto test on the 70-battle
live-vs-DrussGT corpus (/tmp/tfil_ab2). Defines the wave-relative state
(lat/vlat/toa/room/turn, 5/7.9/10 bits at Q=2/3/4), quantises the miss offset
at the bullet's arrival into 7 bins, and sweeps window length K in
{1,4,8,16,32,48} with an interpolated suffix-backoff model under a BY-BATTLE
70/30 split (3 seeds).
Result: NO. On the pre-fire frame the window is worse than the single
fire-tick state at every K/Q/A (e.g. K=8 costs +0.35..+0.46 bits). On the
during-flight frame the entire apparent gain is the later decision tick, not
the window; the single state alone drops 2.70 -> 1.31 bits as K goes 1 -> 32.
The shuffle-order control confirms recency matters but the windows do not:
by Q=4/K=8 they average ~1 observation and never recur. The single state
survives as a strong predictor (log-loss 2.346 vs 2.698 majority; bin
accuracy 0.409 vs 0.235).
Gate only: no gun, no live-win claim.
Adds an smCounted storage mode alongside the default smBitset. Each
(i,j,class) cell becomes a saturating uint8 counter; learn increments it and
a global fractional decay (c -= c shr decayShift every decayEvery learns)
makes forgetting possible. infer sums raw counters; new inferProb sums the
per-cell posterior P(class|cell) (scale-free, recommended readout).
Bitset path is the default and byte-for-byte unchanged: test_bitbrain 56/56
(was 32), and test_bitbrain_mnist reproduces 97.210% corrected / 96.540%
bug-compatible exactly.
Counted mode configurable at runtime (TR_BITBRAIN_MODE / TR_BITBRAIN_DECAY_*)
and compile time (-d:bitbrainDecay*). Measured: forgetting (86.2% vs 48.9% on
a permuted-label stream), probabilities (rare-class balanced 0.998 vs 0.500),
and the stationary cost (counted hurts MNIST; see docs/bitbrain_counted_sbc.md).
Harness: common_libs/tests/measure_counted_sbc.nim
6 arms x 7 runs x 7 rounds vs real DrussGT on one frozen binary (2747ebd).
Validity check PASSES: g100 (fixed gain 1.0) is statistically indistinguishable
from control (dmg p=0.65, wins p=0.62, ALL hit rate +0.04pp p=0.92; zero [bb]
lines = provably no correction). Fixed gains above 1.0 LOSE at 300+: gfix150
-94 dmg/run (p=0.0006), 4/49 vs 16/49 wins (p=0.009), -3.85pp at 300-450
(p=0.0006) and -2.25pp at 450+ (p=0.0023). The hypothesis arm ghi (learner
allowed above 1) is directionally positive but inside the MDE (+14.3 dmg/run
p=0.41; +0.83pp at 450+ p=0.35). Kill the gain axis in both directions.
Includes the mandatory correction notice: Phase 1's [1,1,1,0,0] is LIVE-REFUTED
by 140fe25, and the standing rule that offline is veto-only / live decides.
Task A — the gain region Phase 0 never covered (gain < 1). Extend the
prediction-quality ruler with gain 0.25/0.50/0.75 arms and a fixed causal
per-band arm. Full 70-run result: the hitProxy-argmax curve is
[1.00, 1.00, 1.00, 0.00, 0.00] — Pattern below 300 px, HeadOn above —
worth +0.13 pp at 300-450 and +2.16 pp at 450+ (0.0767 -> 0.0984). Lead
correlation is identical for every g>0 (Pearson is scale-invariant), so a
shrinking gain adds no lead information; and the least-squares optimum
[1,1,1,1,0.25,0.25] diverges from the hitProxy optimum because Pattern's
lead errors are bimodal.
Task B — rebuild guns/bitbrain_gun.nim as a lead-gain corrector:
aim = LOS + gain*(patternAim - LOS), gain learned online per range band by
ranking candidate gains on the hit-probability proxy (the observed lead
label via tmhObservedAt), gated to range >= 300 px. ADE+SBC output removed.
Offline (70 runs): Pattern below 300 px, +0.37 pp at 300-450, +1.90 pp at
450+ (hitProxy 0.0957 vs 0.0767), matching the fixed rule to within 0.27 pp
at 450+ and exceeding it at 300-450. ~0.0007 ms/tick marginal (old gun
~0.114 ms/tick). Default off; rack membership, env report and guard tests
(bitbrain 32, registration 13, rack 48, tm_pattern 20, env_report) unchanged
and green.
Ledger: docs/bitbrain_campaign.md Phase 1, including the three on-file
negatives and the causal-shippability note. No live claim.
2 arms x 15 runs x 7 rounds, one frozen binary from HEAD a82c864, real DrussGT,
server-side events sidecar. Shipped rack is onlyPattern, so control=Pattern-only
and headon=HeadOn-only (TR_RACK_PATTERN=off TR_RACK_HEADON=both).
arm dmg/run dmgtk/run round wins shots/run
control 279 211 48/105 785
headon 14 228 0/105 580
Round wins and dmg/run both separate at p<0.0001 (MC permutation, se 0.0000),
~7x the damage MDE (35.8). Per range band (pooled, 15 runs):
300-450: Pattern 12.3% (4590 shots) vs HeadOn 0.6% (3701) p<0.0001, MDE 2.0pp
450+ : Pattern 9.2% (6671) vs HeadOn 0.4% (4177) p<0.0001, MDE 1.1pp
HeadOn loses EVERY long-range band by 20-23x, so the whole-battle loss is not a
close-range artefact.
The offline ruler (prediction_quality_results.txt) predicted the opposite: HeadOn
meanAbs 14.61 vs Pattern 17.53 at 300-450 and 12.33 vs 16.19 at 450+, hitProxy
.105/.104 and .098/.077 (+27%). That is an open-loop replay of a FIXED enemy
track, so it cannot see that a different bullet makes the surfer dodge
differently; live, the static gun does not lead at all.
TR_PATTERN_RAD_SCALE arms were skipped: applyRadial scales aim DISTANCE along an
unchanged bearing, so it cannot express 'less lead' (bearing is what firing uses).
HeadOn confirmed to ignore bulletSpeed (head_on.nim:9), liveness OK 15/15.
Adds the range-band analyzer tools/ab/ab_range_bands.py (reuses the lead-capture
Run alignment) and the captured fixtures. Does not touch bitbrain_gun.nim /
bitbrain_campaign.md (job-100).
New harness (common_libs/gun_harness/prediction_quality.nim +
common_libs/tests/run_prediction_quality.nim): per-gun single-tick aim error in
degrees against the true continuous interception point on the recorded
live-vs-real-DrussGT corpus (/tmp/tfil_ab2/out, 70 runs, 899607 ticks), per
range band, with the hit-probability proxy mean(|err|<=atan(18/range)).
Validated: recorded hits separate from misses 13.34x px (reference 11.59x),
perfect-oracle max |err| = 0, correct ordering on synthetic ground truth, two
full runs byte-identical. Fixed a wrap180 bug (Nim float mod keeps the dividend
sign) that inflated the negative error tail.
Bar (mean|err| deg [hitProxy] at 450+): Pattern 16.19 [0.077], naive-linear
22.86 [0.054], TMHorizon 16.20 [0.076], BitBrain 16.20 [0.077], static HeadOn
12.33 [0.098], oracle 0 [1.0]. Lead-gain sweep on Pattern is a dead end (1.0
wins every band). Naive-linear applies ~1.8x Pattern's lead but carries no more
lead information (corr 0.178 vs 0.165) and is strictly worse. Ledger:
docs/bitbrain_campaign.md. All verdicts remain live-only.
4 arms x 7 runs vs real DrussGT. mix alternates the two guns 476 times/7 runs
(liveness OK) but our bullets are no more varied (power sd / aim-offset sd flat)
and DrussGT's dodge quality is unchanged (miss/tick mix-pat +0.03, p=0.66; MDE
3.8%). mix wins 24/49 = the 49% baseline; the user's 6/10 has P=0.353 at 49%.
New tools/ab/ab_dodge_analyze.py splits the validated per-shot dodge instrument
by arm and adds gun-switch/power/bearing liveness; fixtures committed.
- docs/bitbrain_gun_verdict.md: control vs bb_decay (decay SBC memory) vs a
provably-zero placebo, 30 runs/arm vs real DrussGT. Nothing separates
(bb_decay +3.3 dmg/run, p=0.71; round wins 97/210 vs 97/210, p=1.00); the
7-run shape does not replicate. TR_BITBRAIN_RANGE=0 is clamped to 1.0 deg
(bitbrain_gun.nim:207) so it is NOT a zero-shift placebo; TR_BITBRAIN_MIN_OBS
unreachable is used instead.
- tools/ab/ab_analyze.py: keep exact enumeration for C(n,na)<=20e6 (7v7), add
a seeded Monte-Carlo permutation test (1e6 draws, 0x5eed5eed) with its
standard error, a tie-corrected Mann-Whitney U cross-check, a minimum
detectable effect line, all-pairs comparisons, and a [bb] shift check.
- tools/ab/README.md: document the new analyzer output.
Measure, for every shot ModularBot fires at the real DrussGT, the lead we
actually applied vs the lead the enemy's motion required, from the recorded
live battles (/tmp/tfil_ab2, 70 battles / 490 rounds / 54926 shots, plus a
35-battle powtest replication of a different binary).
- requiredLead from an AIM-INDEPENDENT interception solve (bullet speed vs
enemy truth), appliedLead from the server-recorded bullet bearing.
- capture = applied/required, guarded at 2px lateral lead (1.6% excluded);
headline metric is the robust proportional slope.
- validation: hits 11.6px mean miss / 80.8% inside 18px, misses 134px,
11.6x separation; 496/496 death + 70/70 owner attributions correct.
Direct answer: capture falls with RANGE (capSlp 0.401 -> 0.135, and
|err|/tolerance 1.27 -> 7.54) but is FLAT across fired POWER within a band
(450+, enemy alive: 0.154 / 0.127 / 0.127). The sub-0.5 long-range shots
(1.18% hit) are finishKill endgame shots at a near-dead DrussGT, not a
lead-capture failure. A naive linear predictor captures 0.29-0.60; we reach
46-67% of that, so the under-lead is real but capture=1.0 is unattainable
against a dodger (oracle required lead).
The offline-harness audit (`e40c849`, `docs/offline_harness_trust.md`) found that the
claim "best offline hit rate of anything measured" for the ring mover was false.
The 20.28% figure is a LIVE number: `docs/feature_ab_results.md` and commit `bfdcdf8`
record 35 real-DrussGT bridge battles with a server-side event sidecar as the ground
truth, and the 6/49 round wins is likewise live. There is no offline measurement of
the ring mover anywhere - the offline harness scores GUNS, not movements, and has no
movement driver at all.
So this was NOT an offline-vs-live calibration failure, which is how it has been
described repeatedly (including by the orchestrator). It was a METRIC MISMATCH: a
movement arm judged on hit rate instead of damage/run and round wins - and hit rate is
precisely the metric that concealed its collapse.
The lesson previously attached to this result was therefore the wrong one. The
paragraph now says what actually happened and points at the audit.
Docs-only; no code touched.
AUDIT (docs/offline_harness_trust.md, new):
- Re-ran acceptance_offline_vs_online myself TWICE: 12/12 deterministic guns
exact both times (264 ticks/enemyId=1, 244 ticks/enemyId=2), death boundary
included. The offline range reproduces the live bot's own per-gun virtual
telemetry exactly.
- Re-verified the (fireTick, powerBin) wave-pairing fix: exact-key lookup,
collisions counted not silently mislabelled; test_wave_pairing 17/17 PASS.
- The offline score is the live TELEMETRY (last-100 virtual hit rate) but NOT
the live BATTLE score (damage/round wins). Two-level answer, documented.
- bmPoint scores up to one tick-step (~17px) PAST its documented aim distance,
while the tie-break probe scores exactly the aim point. Real, low-impact,
deliberately NOT fixed (point metric is non-default, measured negative, and
the committed point baselines would silently change).
- bmPoint/bmPath, perfect-info captures, conditional-on-selection live rates,
and hit-rate-as-objective-for-movement all catalogued as non-apples comparisons.
FIX (unambiguous, fail-before/pass-after):
- common_libs/tests/range_guns.nim: buildAllGunDrivers defaulted to
enableTmSelector=true, so run_range / analyze_selector / test_power_selection /
measure_power_policy spawned gun 13 (TMSelect) - a gun the shipped bot NEVER
spawns. The shared VirtualTracker ring is order-sensitive, so those 4
spawns/tick permuted the learning guns' resolution order (the exact confound
4cd5618 fixed for the acceptance test, left broken for every default caller).
Default is now false (mirror the shipped rack). Impact on
tr_drussgt_vs_modularbot: Tsetlin 18.8->18.5%, KNN 7.5->7.2%, TMSelect 15.2->0.
- New guard common_libs/tests/test_range_rack_parity.nim (3 checks); proven to
FAIL before and PASS after by stash-reverting the fix.
CALIBRATION (offline prediction vs live outcome, 9 usable arms):
- Direction agreement 3/9 = 33%. Split by domain: open-loop (single-tick
prediction / metric / threshold) 3/3; closed-loop (adaptation / range /
movement / selection) 0/6. Small, non-random, hand-assembled set - no
correlation coefficient is claimed.
- The four motivating "offline wins" re-attributed: ring mover was NEVER
offline (it is a live server-side hit rate, mislabelled "offline" in
env_reference.md:342 and commit 7f6ccfb); TMHorizon window/NSTATES and the TM
gun are the H3 classifier-accuracy harness (not hit rate); TFIL is the H2
open-loop movement replay, whose mechanism prediction was right and whose
outcome prediction was wrong.
- Open-loop hypothesis tested: TR_RACK_* knobs leave the offline range output
BYTE-IDENTICAL (the replay never calls the selector), and the range has no
driver for guns 14/15 (TMPATTERN/TMHORIZON). BUG vs LIMIT separated.
VERDICT: trust the harness for single-tick prediction quality only; never for
anything running through the closed loop. MEASURED vs INFERRED labelled.
Green counts unchanged: test_gun_harness 39, test_vbullet_metric 11,
test_power_selection 3, test_power_policy 58, test_adaptive_radar 41,
test_tfil_ring_weights 24, test_ram_decision 40, test_rack_membership 48,
test_selector_tiebreak 19, test_tm_pattern_registration 20,
test_vbullet_admit_gate 12, test_tm_horizon 104, test_tm_diag 48,
test_tm_automata_diag 55, test_tm_clause_shape 66, test_env_report 25,
test_tfil_commit_env 30 (as-is). New: test_range_rack_parity 3.
Answers the user's hypothesis that DrussGT dodges low-power shots better.
Measured on 70 live battles / 490 rounds / 54939 real shots vs real DrussGT
(/tmp/tfil_ab2) and replicated on 35 more battles / 24280 shots (/tmp/powtest).
Power is not randomly assigned - our policy caps it by RANGE
(TR_POWER_FAR_DIST=200 -> 1.0) and by OUR OWN ENERGY (the slope), so inside a
range band power is almost a deterministic function of our energy and a naive
low-vs-high comparison is secretly a losing-vs-healthy comparison. Everything
is stratified by range band and backed by a within-band shuffled-label null
(arrival re-derived, so the null keeps the kinematic channel), a round-cluster
bootstrap, and a within-shot CONTROL window 40 ticks later when the bullet is
long gone.
RESULT: no behavioural response. In band 450+ the raw miss distance at arrival
is +8.25 px [+5.39,+11.25] for HIGH power - but per flight tick it is 4.52 vs
4.51 px/tick (delta -0.01 [-0.12,+0.10]), i.e. entirely the 2.13-tick longer
flight window of the slower bullet. Fixed-12-tick lateral displacement is flat
(55.63 vs 55.47, -0.15 [-1.15,+0.84]) and turn rate / speed are flat. The whole
difference is already present 5 ticks after the trigger pull (+4.2 px) and is
just as large in the bullet-free control window (+5.6 px), so it is a property
of the low-energy situation, not of the shot. Hit rate is flat (0.10 vs 0.09).
Corpus/attribution notes: e*=DrussGT (subject), s*=ModularBot, per
TrBattleCapture.java; the Tank-Royale owner id is NOT stable across runs and is
recovered per battle from fire geometry + the energy decrement, cross-checked on
496/496 death events. Geometry validated on the server's own hits (mean miss
11.6 px, 80.6% inside the 18 px radius).
Reads docs/BitBrain_C_code.zip end to end (full_mnist_2048.c 568 lines +
bitarray.h), reconciles every weight/data file's byte size against its
loader, and answers two questions:
1. What the algorithm is: signed thresholded address decoders (random
projections) -> write-once sparse binary coincidence bit tensors ->
counting/argmax readout. Thresholds and ADs come from an unsupervised
program that is NOT in the zip; only the SBC tensors are learned here.
The supervised rule is genuinely online, single-pass, order-free, and
has no learning rate - so it satisfies ModularBot's reset-on-enemy-change
constraint. Found and isolated a real bug: read_from_sbc's uint8_t
bit_test truncates the 32-bit bit test, so only ADEs with i%%32 < 8 are
ever counted. Shipped reader: 96.540%% (reproduced exactly); with the bug
fixed: 97.210%%.
2. Whether it can become a gun: not on this evidence. Measured full
inference at ~0.56 ms/sample (cost is not the blocker), but the AD layer
is unobtainable, and this repo has already shown the binding constraint
is the target signal, not the learner - the TM head scored 2.08pp below
its own majority class and the AD/SBC primitives are already in the
16-family shootout as WiSARD/Bloom/SDM. Recommends a readout swap on the
existing WiSARD feature extractor and a Gate-2b-style >=80%% side-accuracy
gate before any port.
Also: recommends committing the 11 MB zip under its explicit name.
DOCS ONLY. The user set TMH_NSTATES, which is a compile-time {-d:intdefine.}
(-d:TMH_NSTATES=2, tm_horizon.nim:102), not an env var; the runtime var is
TR_TMHORIZON_NSTATES (tm_horizon.nim:139). This rewrites the reference so that
class of confusion cannot recur.
What was WRONG and is now fixed:
- "unparseable warns and falls back" was false for the numeric/bool knobs: only
the enumerated string knobs warn; envInt/envFloat/envBool fall back silently.
- TR_POWER_LOG/TR_RAM_LOG/TR_MOVEMENT_LOG/TR_RECORD_WORLDSTATE/
TR_RADAR_FORCE_SPIN/TR_RADAR_SCANLOG/TR_TRACKER_PROBE are read with existsEnv,
so TR_POWER_LOG=0 turns the log ON. Documented per knob.
- TR_POWER_ENERGY_MIN is a CAP at low energy, not a minimum-power floor.
- TR_TMHORIZON_* knobs are inert unless TR_RACK_TMHORIZON=both; the doc implied
they were live.
- GUN_SELECTOR_WINDOW is clamped 1..100; TR_MOVEMENT silently falls back to tfil
for any value other than tfil_ring.
- The ring file's own header comment (corridor 5 / wall 10) is stale; the code
defaults are 10.0/15.0 (commit 7f6ccfb).
- Compile-time section was incomplete and conflated the two TM modules:
tm_pattern uses TM_NCLAUSES/TM_NSTATES, tsetlin uses TM_N_CLAUSES/TM_N_STATES,
and -d:TM_S_DEF is defined in BOTH.
What was ADDED:
- "COMPILE-TIME vs RUNTIME: the two namespaces": the only define/env pair is
TMH_NSTATES <-> TR_TMHORIZON_NSTATES; everything else is compile-time only.
- "Did my env vars actually reach the bot?": the /proc exec-time check, the note
that grepping only TR_|GUN_ hides a wrongly-named var (grep -i tmh), and the
boot report described as an interface with the two sections + build identity.
- "Measured verdicts" table: window 26.5% vs 49.0% p=0.036 (harmful live),
TMHorizon N=2/8/64 42.9/42.9/53.1% (all p>=0.8), power floor 22/49 vs 21/49
p=1.0, sub-1.0 accuracy 10.80% vs 10.35% p=0.37, shipped bot 49% vs DrussGT.
- "Names that look real but do nothing": TMH_NSTATES (env), TR_VBULLET_METRIC,
TR_POWER_LOW_ENERGY, TR_TRACKER_RECONCILE.
- The adaptive-melee radar's compile-time constants (no env form).
The bot is spawned by the server/GUI, so it inherits the SERVER's
environment. The user could not tell whether their exports reached the
bot, so print a one-shot greppable report at boot:
grep '^\[env\]' /tmp/modularbot_stdout.log
Section A prints every TR_*/GUN_* this process actually received, the
count vs the total env size, a loud warning when nothing matched, and
the process identity (pid/ppid, cwd, self command line, and the PARENT
command line) so the spawn trap is obvious. Section B prints the
resolved effective value of every documented knob with its source
(env|default), including clamps and the rack's empty-set fallback.
Build identity (NimVersion, compile date/time, binary path/size/mtime)
pins the exact artifact. Suppress with TR_ENV_REPORT=0.
docs/env_reference.md: add the missing GUN_SHOTLOG_PATH,
GUN_SELECTOR_MINOBS/FLOOR/POOL/RANK/SHRINK/SEED, TR_ENV_REPORT and
-d:TM_NCLAUSES; record the measured TR_TMHORIZON_WINDOW verdict; and
add a prominent 'Did my env vars actually reach the bot?' section with
the boot report, the /proc/PID/environ no-code check, the correct GUI
launch recipe, and how to prove the trap deliberately.
The user's request: "when our bot is low OR enemy is low, it is useless to use high
power instead low fast bullets have more chances to finish the enemy. Let's do a
math slope: starting from some health down, the power goes down with it."
1. ENERGY SLOPE (`TR_POWER_ENERGY_*`), replacing the old hard step at 50 energy:
cap = ENERGY_MAX at/above ENERGY_HI, ENERGY_MIN at/below ENERGY_LO, LINEAR in
power between, clamped. Defaults HI=80 LO=20 MIN=0.5 MAX=3.0, so no cap >=80,
0.5 at <=20, and e.g. E=65 -> 2.375, E=50 -> 1.75, E=35 -> 1.125.
Rationale: bullet speed is 20-3p, so lower power = FASTER bullet (less lead
error, higher hit chance), fires more often (10+2p) and drains slower (p/shot).
E[dE] = p(3P-1) => break-even hit probability is 1/3 INDEPENDENT of power, and
our measured rates are 5-27%, far below it.
2. FINISHING CAP (`TR_POWER_FINISH_KILL`, default ON): cap power at the SMALLEST
bullet that still removes the enemy's remaining energy -
`E<=4 -> p=E/4` (min 0.1), `4<E<=16 -> p=(E+2)/6`, `E>16 -> no cap`.
Rationale, and it makes the user's instinct stronger than a heuristic: server
1.3.1 caps the damage SCORE at the energy ACTUALLY REMOVED, so overkill is
WASTED damage AND ~6x the energy for ZERO extra score. Damage is 4p (p<=1) /
6p-2 (p>1).
Both are min-composed with the existing far/below-average caps, may only LOWER
power (exhaustively tested), and are exempt while ramming.
`TR_POWER_POLICY=0` still returns the uncapped control exactly.
MEASURED ENERGY SAVING (offline replay of the DrussGT fixtures, 28,797 ticks):
arm shots energy meanP E/1k ticks vs cliff
control(uncapped) 1913 4646 2.43 161.4 -90.2%
cliff (today) 2363 2443 1.03 84.8 0.0%
slope 2404 2178 0.91 75.6 ** 10.9% LESS **
slope+finish 2404 2167 0.90 75.3 ** 11.3% LESS **
So the slope spends ~11% less energy than the cliff AND fires slightly MORE shots
(2404 vs 2363) - both directions at once.
HONEST NOTE on the finishing rule's reach here: ticks where the enemy is low
(0 < E <= 16) are only 2252/28797 = 7.8% of these fixtures, so finishing adds just
~11 energy of saving against DrussGT. It matters in CLOSER fights, not this one.
Verification: test_power_policy 58 (was 26) in BOTH the default and TR_POWER_POLICY=0
control arms - slope at E=100/80/65/50/35/20/5, powerToKill across E=0.1..100, the
inverse-cover property for E<=16, monotonicity, ram exemption, and an exhaustive
sweep proving power <= preference. Guards: test_gun_harness 39, test_vbullet_metric
11, test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24,
test_ram_decision 40, test_rack_membership 48, test_selector_tiebreak 19,
test_tm_pattern_registration 20, test_vbullet_admit_gate 12. acceptance
12/12 PASS. ModularBot compiles release.
Adds `common_libs/tests/measure_power_policy.nim` (the energy/histogram tool) and
updates docs/env_reference.md for the new `energySlope|finishKill` log reasons.
NOT MEASURED: the battle/hit-rate effect. The offline figures use the fixture
shooter's energy as a proxy, open-loop; the RELATIVE saving is the meaningful part.
Single place to answer 'what env vars exist, what do they default to, and what are
they for'. Grouped by area, with the measured reason each knob exists recorded
next to it, since several defaults are counterintuitive:
- The shipped rack is PATTERN ONLY, with the revert one-liner included.
- GUN_SELECTOR_TIEBREAK defaults off because it measured NEGATIVE on real hit
rate, and the per-tick random draw inside the tie band is load-bearing
(commitment cost 7.02% -> 5.10%, p=0.002).
- The power policy is ON because turning it off measured 10.6% -> 7.9% real hit
rate; TR_POWER_FINISH_KILL exists because server 1.3.1 caps the damage SCORE at
the energy actually removed, so overkill scores nothing.
- The ring mover is NOT the default and must not be shipped: best offline hit
rate of anything measured, but it halves survival (16/49 -> 6/49, p=0.012).
- The proactive ram is off because it converts 0/6 times.
- TR_PATTERN_RAD_* are kept but are structurally incapable of changing the shot
(the aim is bearing-only, measured byte-identical on the path metric).
- TR_TMHORIZON_NSTATES is the automata inertia ('mood'): lower = adapts faster.
- TR_VBULLET_ADMIT_ONLY=1 gave +68% tick rate (87 -> 146 ticks/s).
Documents the read-once-at-start convention and the operational trap that the bot
is spawned by the server/GUI, so a variable exported in an unrelated terminal does
NOT reach it. Also lists the compile-time -d: knobs and the test-harness jars.
Snapshot of the source at this commit; f9f8d84-era knobs added by the in-flight
jobs are included where already present.
Follows the diagnosis that proactive straight-line ramming CANNOT work: both bots
have MAX_SPEED=8, so a pursuit cannot catch an evading equal-speed opponent.
Measured over 49 rounds per arm, opportunity -> contact was **0/6** (base), 0/40
(ring), 0/12 (ringhot). The only proactive conversion in the whole corpus came
from a FINISHER, and only because a <20-energy DrussGT stops fleeing (that episode
closed at 6-8 px/tick). Opportunity episodes never got below ~80px; one ran the
full 60-tick duration cap and closed only 198->171px; a perfectly aligned
full-speed one closed 195->114px then plateaued.
CHANGES
- **Finisher-only default.** `finisher` (<20 energy, dist<300, we are healthier)
and the rare `desperation` (both <5, dist<150) are kept; `opportunity` and the
speculative `plan` are OFF. Both are env-reenableable with no rebuild:
`TR_RAM_OPPORTUNITY=1` (tune via TR_RAM_OPP_DIST/MARGIN) and `TR_RAM_PLAN=1`.
Justification: it removes 100+ non-converting episodes per fixture at zero
measured loss (oldram vs base was p=0.69, damage 279 vs 284, survival 17/49 vs
16/49) - and each of those episodes spent up to 60 ticks driving STRAIGHT at
the enemy, abandoning the mover's dodging and disrupting aim.
- **`desperation` KEPT** deliberately: it is cheap and rare, fires only when both
bots are nearly dead at short range (a coin-flip where 0.6 contact can decide
it), and it is not the refuted straight-line pursuit.
- **THE BULLET-RAIN ABORT WAS DEAD CODE AND IS NOW FIXED.** `onHitByBullet`
accumulated raw bullet FIREPOWER while `TR_RAM_ABORT_DMG = 0.5` was documented
as a DAMAGE rate - so the bar was implicitly "sum of power > 7.5 over 15 turns"
and the maximum rate ever observed was 0.27. It now accumulates REAL ENERGY via
a `bulletDamage(power)` helper matching the server's `4p` / `6p-2` formula, and
`TR_RAM_ABORT_DMG` defaults to **2.0 energy/turn** (~30 HP over 15 turns):
"abort an in-progress ram if we take > 2.0 energy per turn". Same effective bar
for normal firepower, and it can now actually fire - the live run reports
`dmgRate=1.07/turn` where the old units said 0.27.
- **`ramStuckTicks` REMOVED.** It required `dist < 5px`; contact occurs at ~36px
(two 18px radii) and position rewind prevents getting closer, so it could never
increment. Only the 60-tick duration cap can now self-end a ram.
LIVENESS (measured, default config, vs a charging Java RamFire, 3 rounds):
default -> `[ram] ON reason=finisher` x3, `reason=opportunity` x0
TR_RAM_OPPORTUNITY=1 -> `reason=opportunity` x4, `reason=finisher` x2
So the opportunity states DID occur and are suppressed by the new default - the
removal is real, not an arm that never fires. A line also read
`[ram] OFF reason=duration dmgRate=1.07/turn`, confirming the new energy units.
Adds docs/ramming_negative_result.md (70 lines) recording the question, the five
diagnostic answers, the geometric reason, the finisher exception, the two dead
code paths, and an explicit "do not re-attempt a proactive straight-line ram; if
point-blank forcing is ever wanted it is an INTERCEPTION/cornering movement
problem" note - the same pattern that stopped the corpse bug recurring.
Guards: test_ram_decision 40 (was 28), test_gun_harness 39, test_vbullet_metric 11,
test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24,
test_power_policy 26, test_rack_membership 48, test_selector_tiebreak 19,
test_tm_pattern_registration 20, test_vbullet_admit_gate 12,
acceptance_offline_vs_online 12/12. ModularBot compiles.
Honest note: the abort-threshold fix is a real (tiny) behaviour change, NOT
measured-neutral - it only bites while a finisher ram is under sustained fire,
which is exactly the user's stated wish. The finisher-only removal itself is
measured-neutral per the given A/B.
The user's plan was separate melee and 1v1 racks. The mechanism is built and
committed (TR_RACK_<GUN>=both|1v1|melee|off, mode from server truth). This job
produced the missing evidence: each candidate gun forced ALONE in BOTH modes,
one frozen binary, env-only arms, exact two-sided permutation tests.
1v1 vs the real DrussGT (6 arms x 5 runs x 7 rounds):
Pattern 10.80% 292 dmg/run 11/35 round wins
KNN 5.58% 128 1/35
Linear 2.99% 59 0/35
Circular 2.89% 60 0/35
WallBounce 2.72% 57 0/35
GuessFactor 2.14% 38 0/35
-> every arm differs from Pattern at p=0.0079 (the 5v5 floor). Per-gun damage
span 7.7x. In 1v1 the gun matters ENORMOUSLY.
Melee (+DrussGT, RandomMover, WaveSurfer, OscillatorBot; enemyCount 4):
WallBounce 25.79% 643 dmg/run 7/25
Linear 24.88% 624 7/25
GuessFactor 24.48% 613 4/25
Pattern 24.39% 599 5/25
Circular 25.88% 596 5/25
KNN 23.87% 539 3/25
-> NO arm beats Pattern with significance (p=0.42-0.96, fully overlapping).
WallBounce's nominal +7.3% is p=0.42. Per-gun damage span only 1.19x.
THE FINDING: in 1v1 the gun matters enormously (7.7x spread); in melee it barely
matters (1.19x). Melee is won by movement, survival and placement, not by which
gun you carry - every candidate lands in the same ~24-26% band. So a separate
melee rack has NO gun-level payoff and DefaultRackMembership stays Pattern-only.
The mechanism remains available if it is ever wanted.
Honest caveats: the melee comparison is UNDER-POWERED at 5 runs (detecting the
~44-damage WallBounce-Pattern gap would need ~2.5-3x the runs), so "no significant
difference" is NOT "no difference"; the only candidate worth re-testing is
WallBounce in melee, and it must not be shipped on this evidence.
SHIM LIMITATION FOUND, worth recording: run_bridge_battle.sh captures with
`--subject DrussGT`, and the shim DROPS every tick once the subject dies - losing
5-20% of ModularBot's shots in melee. The campaign was re-run with
`--subject ModularBot` so ModularBot's events are complete (fires match gun_stats
realShots exactly). Any future melee evidence through this shim must do the same
or it will silently under-count.
Adds docs/melee_vs_1v1_racks.md. No repository source changed.
Five arms x 7 runs x 7 rounds (35 real-DrusGT battles, 8 concurrent), one frozen
binary from git archive HEAD at 185a32e (includes the shipped Pattern-only rack),
env knobs only, server-side event sidecar, exact two-sided permutation tests.
arm real % dmg/run survival(rounds won) p vs base
base (shipped) 10.61 284.1 16/49 (32.7%) --
nopower (policy off) 7.88 247.1 9/49 (18.4%) 0.0012
oldram (old gates) 10.78 279.4 17/49 (34.7%) 0.6888
ring (tfil_ring) 20.28 267.4 6/49 (12.2%) 0.0006
ringhot (orig heat) 11.73 285.3 15/49 (30.6%) 0.0303
Every round ends with exactly one death (0 timeouts), so survival = round win.
1. POWER POLICY HELPS - KEEP. Turning it off drops real hit rate 10.61 -> 7.88
(p=0.0012), damage/run 284 -> 247, and wins FEWER rounds (16 -> 9). The cap
trades per-shot damage for many more shots and a higher per-shot rate; that
trade is a clear win. This was shipped on unit tests alone until now.
2. PROACTIVE RAMMING - INDISTINGUISHABLE. oldram vs base p=0.69, damage 279 vs
284, survival 17/49 vs 16/49 (p=1.0), ram contacts 1 vs 2. The lever IS live
(6 `opportunity` ON events vs 0 under the old 50px/+30 gates; base reached
<40px on 10 ticks vs 0) but converts to essentially no extra collisions and no
measurable outcome change. Safe to keep, but it is not earning its keep and
reverting it is equally defensible.
3. RING MOVER HURTS THE OBJECTIVE - DO NOT SHIP. And this is the important one.
*** METHODOLOGY CORRECTION - I HAD THIS WRONG ALL NIGHT ***
The ring arm has the BEST hit rate of anything measured tonight: 20.28% vs 10.61%
(+9.67pp, p=0.0006, non-overlapping). Read alone it says "ship it immediately".
It is a CONFOUND. The ring halves engagement range (median 460 -> 240px), which
halves round length (1542 -> 636 ticks) and shots (4466 -> 1834). So damage/run is
FLAT (284 -> 267) while survival/round-win MORE THAN HALVES (16/49 -> 6/49,
p=0.0122). It is a GLASS CANNON: same damage dealt, twice as many deaths. The
hit-rate gain is a geometric artefact of fighting closer, not an improvement.
** For MOVEMENT arms, hit rate alone INVERTS the verdict. ** A movement change
alters range, shots fired and round length simultaneously, so the objective
metrics are DAMAGE/RUN and ROUND-WIN RATE (survival) - report all three. For
gun/selection arms, where range and round length are held fixed, real hit rate
remains the right ground truth. I had been enforcing the hit-rate-only rule
without qualification; it is now qualified.
HEAT TAMING is the knob that moves the tradeoff: tamed (corridor 5 / wall 10)
lets the range weighting pull to ~240px (20.28% / 6 wins); original (20/30) keeps
it at ~400px (11.73% / 15 wins). So heat taming buys hit rate at the cost of
survival - a knob to keep conservative, and the reason the ring is not the default.
Nothing to revert: the shipped movement is already `tfil` and the ring is opt-in.
Caveats recorded: only adversary is the real DrussGT jar (the shim hosts only
jk.mega.DrussGT), so the ring-vs-winning result needs re-checking elsewhere; and
base (10.61%) is consistent with the committed onlyPattern result (9.99%),
validating the frozen binary and pipeline.
Adds docs/feature_ab_results.md. No source files changed by this job.
`DefaultRackMembership` now admits Pattern (id 5) and marks all 14 other guns
`rmOff`. **The selector mechanism is untouched** - `chooseFromFit`, the floor/band
logic, the hysteresis and the virtual-fitness plumbing are all intact and
functional. Only the rack membership changed, so this is reverted by env alone.
Evidence (measured, replicated three times, 10 adversaries): Pattern alone gives
10.36% real hit rate / 264 damage per run vs the full rack's 6.93% / 159. Pattern
significantly wins on DrussGT, Corners, Crazy and PatternMover, ties on three, and
the full rack never significantly beats it on ANY adversary. Mechanism: the
virtual signal keeps ranking the wrong guns first (HeadOn 46% of ticks at 2.0%
real; Linear 57.7% at 6.0% real while Pattern sits at 11.1%).
**THIS CONTRADICTS THE USER'S STANDING DIRECTIVE** to keep virtual-fitness
selection. Recorded plainly in docs/selector_negative_value.md with a SHIPPED
DECISION banner rather than done quietly: the mechanism is retained and one env
var away, because the measurement says it is negative value on every rack size
tested and on 10/10 adversaries.
Revert one-liner (no rebuild):
TR_RACK_PATTERN=both TR_RACK_HEADON=both TR_RACK_LINEAR=both TR_RACK_TSETLIN=both \
TR_RACK_CIRCULAR=both TR_RACK_GUESSFACTOR=both TR_RACK_WALLBOUNCE=both \
TR_RACK_ACCEL=both TR_RACK_STOPSHOT=both TR_RACK_DISPLACE=both TR_RACK_AVGLEAD=both \
TR_RACK_DECAYGF=both TR_RACK_KNN=both TR_RACK_TMSELECT=both ./out/ModularBot
The unit test `testRevertOverrideRestoresFullRack` exercises exactly this table.
FLOOR PATH, verified not assumed: `chooseFromFit` already returns `admitted[0]` on
the floor path, so it respects admission by construction. Cold field + shipped
default -> floor returns Pattern (id 5), NOT HeadOn. With an explicit all-`both`
membership the same cold field returns gun 0 (HeadOn) - the old behaviour. Four
assertions in `testFloorRespectsAdmission`.
LIVENESS: one 1-round battle with NO overrides -> Pattern selected 105/105 = 100%,
every other gun 0 including TMPattern.
Honesty caveat retained in the doc: 4 of the 10 opponents were Tank Royale
sample-bot PORTS rather than the original classic jars (only DrussGT is a real
classic jar through the shim).
Guards: test_rack_membership 48 (was 38; new floor/revert/default checks),
test_tm_pattern_registration 20 (5 checks hard-coded the old default and were
updated to assert the new one, with the TMPATTERN parity proof moved onto an
explicit old-rack table), test_gun_harness 39, test_vbullet_metric 11,
test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24,
test_power_policy 26, test_ram_decision 28, test_selector_tiebreak 19,
test_tm_pattern_rack_live 4, test_tm_pattern_learning 3,
acceptance_offline_vs_online 12/12 VERDICT PASS. ModularBot compiles.
FOLLOW-ON THIS EXPOSED: membership filters SELECTION but not virtual-bullet
SPAWNING, so under `onlyPattern` the 13 unselected guns still predict and spawn
every tick. Tsetlin alone is ~5.3 ms/tick (~41% of the 13.16 ms per-tick budget),
so we are still paying for it while never using it. Gating spawn on admission
would reclaim that; it was deliberately NOT done here because it would alter the
measurement protocol mid-A/B.