4 arms x 15 runs x 7 rounds (60 battles, 0 failed) vs real DrussGT on the
shipped TFIL default, frozen at ed25ce2. bb_id (gain 1.0 identity) is
statistically indistinguishable from shipped Pattern -> plumbing validity
check passes. No BitBrain arm beats TMHorizon or Pattern: bb_learn (the config
the owner likely ran) is the worst arm (276 dmg/run, 39/105 wins), the only
comparison at alpha=0.05 is Pattern beating it on damage. Learned gains
(>=1.0, gated >=300px) over-lead and lose 1.61pp of hit rate at 300-450px.
MDE 29.5 dmg/run, 1.235 wins/run; a 6-4-sized effect needs ~39 runs/arm.
Runs the pre-registered A/B for the two movement changes in HEAD: the
time-indexed bullet heat (TR_TFIL_HEAT_TIME, fca8993) and the removal of the
invented virtual centre pillar (d0750ab). One frozen binary from HEAD vs real
DrussGT: 6 arms x 10 runs x 7 rounds = 60 battles, 420 rounds, 0 failed.
Judged on damage/run and ROUND WINS only (hit rate and hits-taken are context):
hit rate would have inverted the verdict again - tau3 has the best pooled hit
rate of all arms (11.56%) and the fewest round wins (20/70).
RESULT (vs the reconstructed pre-change mover "old"):
heat-time HURTS. tau3/tau5/tau9 lose 1.3-1.7 wins/run (p=0.0010-0.0125) and
deal 22-38 less damage/run (p=0.004-0.047); tau15 is a wash on wins (p=0.64)
and 22 damage/run lower (p=0.046). Nothing improves either metric.
pillar removal does nothing measurable. old vs pillaoff: +5.7 damage/run
(p=0.71), +0.5 wins/run (35 vs 30, p=0.43), 30.8 MORE damage taken/run
without the pillar (p=0.040). The mechanism check proves the knob works
(centre-box occupancy 0.09% -> 2.37%, p<0.0001; range 469 -> 443 px,
p=0.0002), so this is a real behaviour change that buys nothing. At n=10 the
pillar contrast is inside the MDE (33 damage/run, 1.2 wins/run), so this is
not a proven regression.
Flags that the shipped default (pillar removed) should be reverted to the
TR_TFIL_PILLAR_ON behaviour; heat-time stays off.
Adds tools/ab/arms_heat_pillar.txt and tools/ab/ab_mechanism.py (per-tick
mechanism check: central-box occupancy, range distribution, live enemy-bullet
proximity) plus the captured summary/report fixtures.
Adds a per-sample intrinsic-confidence field (GunPrediction.confidence,
threaded through FeedbackEvent/VirtualBullet, populated by Pattern, DecayGF,
KNN, GuessFactor, Tsetlin, TMHorizon) and an offline recorder + analyzer that
reproduce the paper's Figure 2 per gun and its Eq-8 composite.
Measured on 3 held-out tr-bridge DrussGT battles (33k ticks, ~133k samples/gun):
- FAITHFUL: DecayGF (rho +0.133), KNN (+0.090), Pattern (+0.064, weak).
- GuessFactor is ANTI-faithful (rho -0.067); Tsetlin c_max is useless (0.001).
- No pair of guns specialises complementarily: the same gun dominates both
high-confidence slices in every pair.
- Eq-8 alpha-normalised confidence-weighted composite: 18.41% vs Pattern
20.45% (McNemar p=3.1e-126). Faithful-only variant 18.68%, still loses.
Shuffle control passes weakly (composite > shuffle, p=4e-14) so ~0.7pp of
competence is real but ~2pp short. Offline veto: design is dead.
See docs/tmcomposites_gate.md.
The default mover painted a 30/10 radiance blob on the arena centre even
though the arena has NO physical pillar there, creating a 4x4 tile
(144x144 px) exclusion zone over open centre floor. Set
PillarHotness/PillarRadiance to 0/0 in the shipped default (matching the
ring variant) and add TR_TFIL_PILLAR_ON=1 to restore the old 30/10 field
for A/B without a rebuild; registered in env_report.
Because the shipped default legitimately changed, the default-path parity
golden (fixtures/tfil_commit_default.golden) was regenerated from the NEW
default, with an explicit 'deliberate default change' note in the test so
a future failure is treated as a real regression.
Also register the three env reads job j102 added in common_libs/bitbrain
(TR_BITBRAIN_MODE / _DECAY_EVERY / _DECAY_SHIFT), which the env-report
guard was failing on.
Verification: test_env_report all green; test_tfil_commit_env 30/30.
Re-runs the SBC coincidence premise as a cheap veto test on the 70-battle
live-vs-DrussGT corpus (/tmp/tfil_ab2). Defines the wave-relative state
(lat/vlat/toa/room/turn, 5/7.9/10 bits at Q=2/3/4), quantises the miss offset
at the bullet's arrival into 7 bins, and sweeps window length K in
{1,4,8,16,32,48} with an interpolated suffix-backoff model under a BY-BATTLE
70/30 split (3 seeds).
Result: NO. On the pre-fire frame the window is worse than the single
fire-tick state at every K/Q/A (e.g. K=8 costs +0.35..+0.46 bits). On the
during-flight frame the entire apparent gain is the later decision tick, not
the window; the single state alone drops 2.70 -> 1.31 bits as K goes 1 -> 32.
The shuffle-order control confirms recency matters but the windows do not:
by Q=4/K=8 they average ~1 observation and never recur. The single state
survives as a strong predictor (log-loss 2.346 vs 2.698 majority; bin
accuracy 0.409 vs 0.235).
Gate only: no gun, no live-win claim.
2 arms x 15 runs x 7 rounds, one frozen binary from HEAD a82c864, real DrussGT,
server-side events sidecar. Shipped rack is onlyPattern, so control=Pattern-only
and headon=HeadOn-only (TR_RACK_PATTERN=off TR_RACK_HEADON=both).
arm dmg/run dmgtk/run round wins shots/run
control 279 211 48/105 785
headon 14 228 0/105 580
Round wins and dmg/run both separate at p<0.0001 (MC permutation, se 0.0000),
~7x the damage MDE (35.8). Per range band (pooled, 15 runs):
300-450: Pattern 12.3% (4590 shots) vs HeadOn 0.6% (3701) p<0.0001, MDE 2.0pp
450+ : Pattern 9.2% (6671) vs HeadOn 0.4% (4177) p<0.0001, MDE 1.1pp
HeadOn loses EVERY long-range band by 20-23x, so the whole-battle loss is not a
close-range artefact.
The offline ruler (prediction_quality_results.txt) predicted the opposite: HeadOn
meanAbs 14.61 vs Pattern 17.53 at 300-450 and 12.33 vs 16.19 at 450+, hitProxy
.105/.104 and .098/.077 (+27%). That is an open-loop replay of a FIXED enemy
track, so it cannot see that a different bullet makes the surfer dodge
differently; live, the static gun does not lead at all.
TR_PATTERN_RAD_SCALE arms were skipped: applyRadial scales aim DISTANCE along an
unchanged bearing, so it cannot express 'less lead' (bearing is what firing uses).
HeadOn confirmed to ignore bulletSpeed (head_on.nim:9), liveness OK 15/15.
Adds the range-band analyzer tools/ab/ab_range_bands.py (reuses the lead-capture
Run alignment) and the captured fixtures. Does not touch bitbrain_gun.nim /
bitbrain_campaign.md (job-100).
4 arms x 7 runs vs real DrussGT. mix alternates the two guns 476 times/7 runs
(liveness OK) but our bullets are no more varied (power sd / aim-offset sd flat)
and DrussGT's dodge quality is unchanged (miss/tick mix-pat +0.03, p=0.66; MDE
3.8%). mix wins 24/49 = the 49% baseline; the user's 6/10 has P=0.353 at 49%.
New tools/ab/ab_dodge_analyze.py splits the validated per-shot dodge instrument
by arm and adds gun-switch/power/bearing liveness; fixtures committed.
Measure, for every shot ModularBot fires at the real DrussGT, the lead we
actually applied vs the lead the enemy's motion required, from the recorded
live battles (/tmp/tfil_ab2, 70 battles / 490 rounds / 54926 shots, plus a
35-battle powtest replication of a different binary).
- requiredLead from an AIM-INDEPENDENT interception solve (bullet speed vs
enemy truth), appliedLead from the server-recorded bullet bearing.
- capture = applied/required, guarded at 2px lateral lead (1.6% excluded);
headline metric is the robust proportional slope.
- validation: hits 11.6px mean miss / 80.8% inside 18px, misses 134px,
11.6x separation; 496/496 death + 70/70 owner attributions correct.
Direct answer: capture falls with RANGE (capSlp 0.401 -> 0.135, and
|err|/tolerance 1.27 -> 7.54) but is FLAT across fired POWER within a band
(450+, enemy alive: 0.154 / 0.127 / 0.127). The sub-0.5 long-range shots
(1.18% hit) are finishKill endgame shots at a near-dead DrussGT, not a
lead-capture failure. A naive linear predictor captures 0.29-0.60; we reach
46-67% of that, so the under-lead is real but capture=1.0 is unattainable
against a dodger (oracle required lead).
Answers the user's hypothesis that DrussGT dodges low-power shots better.
Measured on 70 live battles / 490 rounds / 54939 real shots vs real DrussGT
(/tmp/tfil_ab2) and replicated on 35 more battles / 24280 shots (/tmp/powtest).
Power is not randomly assigned - our policy caps it by RANGE
(TR_POWER_FAR_DIST=200 -> 1.0) and by OUR OWN ENERGY (the slope), so inside a
range band power is almost a deterministic function of our energy and a naive
low-vs-high comparison is secretly a losing-vs-healthy comparison. Everything
is stratified by range band and backed by a within-band shuffled-label null
(arrival re-derived, so the null keeps the kinematic channel), a round-cluster
bootstrap, and a within-shot CONTROL window 40 ticks later when the bullet is
long gone.
RESULT: no behavioural response. In band 450+ the raw miss distance at arrival
is +8.25 px [+5.39,+11.25] for HIGH power - but per flight tick it is 4.52 vs
4.51 px/tick (delta -0.01 [-0.12,+0.10]), i.e. entirely the 2.13-tick longer
flight window of the slower bullet. Fixed-12-tick lateral displacement is flat
(55.63 vs 55.47, -0.15 [-1.15,+0.84]) and turn rate / speed are flat. The whole
difference is already present 5 ticks after the trigger pull (+4.2 px) and is
just as large in the bullet-free control window (+5.6 px), so it is a property
of the low-energy situation, not of the shot. Hit rate is flat (0.10 vs 0.09).
Corpus/attribution notes: e*=DrussGT (subject), s*=ModularBot, per
TrBattleCapture.java; the Tank-Royale owner id is NOT stable across runs and is
recovered per battle from fire geometry + the energy decrement, cross-checked on
496/496 death events. Geometry validated on the server's own hits (mean miss
11.6 px, 80.6% inside the 18 px radius).
Runs the five-arm commitment A/B that 19bf461 only implemented. One frozen
binary (git archive 19bf461, sha256 46e7ce19...) vs the real DrussGT through
tools/robocode_shim/run_bridge_battle.sh: 7 runs x 7 rounds per arm against
real DrussGT in the pre-registered block, plus an independent replication
block (runs 8-14) - 70 battles, 490 rounds, 5 arms in parallel.
A control (shipped) B TR_TFIL_TILE_REPLAN=off
C B + TR_TFIL_NO_REV=1 D TR_TFIL_TILE_REPLAN=enemy
E B + TR_TFIL_COMMIT_TICKS=30 (TR_TFIL_COMMIT_LOG=1 on every arm)
RESULT: no arm improves damage/run or round wins vs the shipped mover. Pooled
(14 runs/arm, exact two-sided permutation test over all C(28,14) relabellings):
arm damage/run (delta, p) round wins (delta, p)
A 274.44 2.79 (39/98)
B 274.07 (-0.37, p=0.97) 3.07 (+0.29, p=0.66)
C 257.34 (-17.10, p=0.16) 2.43 (-0.36, p=0.47)
D 286.61 (+12.17, p=0.36) 3.29 (+0.50, p=0.35)
E 265.29 (-9.15, p=0.56) 2.93 (+0.14, p=0.89)
The replication is what settles it: block 1 alone showed D at +20.99 damage
(p=0.26); block 2 put D at +3.36. Nothing replicates.
The PREMISE fails. Honouring the commitment does not reduce reversals - it
raises the reversal-pick rate from 33.9% (control) to 50.9% (B) / 57.0% (E);
the no-reversal arm C only claws part of it back (42.4%). The post-reversal
speed dip (~4.9 -> ~4.0 px/tick at +1..+2 calls) is identical in every arm, so
it is a property of turning around, not of the tile replan, and arm A - the
arm carrying the bug - has the LOWEST mean abs(speed) of all five.
Both premise premises are also caught by the metric the report is careful
about: arm C takes significantly FEWER hits (87.2 vs 96.5/run, p=0.0022) while
dealing the least damage and winning the fewest rounds - a movement change
alters hits taken, shots fired and round length at once, which is exactly why
damage and round wins are the primary metrics and hit rate is not.
Treatment verified per arm from the per-tick TR_TFIL_COMMIT_LOG: the pick
interval moves off ~5 calls only where the arm says it should (A 5.05 with
96.6% tile_self replans; B 14.77 with 0%; C 14.79; D 5.29 with 95.7%
tile_enemy; E 26.62), matching the offline fixture replay (5.07 / 14.65 /
14.65 / 4.94 / 27.57). Every arm ran.
Round wins are the final RESULTS firstPlaces, cross-checked against each
round's BotDeathEvent (35/35 and 35/35 runs agree). Damage is the server's
BulletHitBotEvent damage, attributed by numeric owner/victim id and verified
against the capture's own event counters.
measure_tfil_commit_ab.nim reads the raw run artifacts (JSONL states, events
sidecar, capture stdout, commit log) and prints the whole deliverable:
arm table, per-run values, the exact per-run permutation test, the exact
round-level hypergeometric test (flagged anti-conservative - rounds cluster
within a run), the treatment diagnostics, and the direct answer. It reproduces
either report offline from the committed summary:
nim c -r --path:common_libs common_libs/tests/measure_tfil_commit_ab.nim \
--from-summary common_libs/tests/fixtures/tfil_commit_ab_results_runs14.json
The raw ~200 MB of battle artifacts are not committed; the committed summary
holds every per-run value plus the arm-level diagnostics.
The tile-change replan cancels the 15-tick movement commitment whenever OUR
tile changes. With GridSize=36 and speed up to 8 px/tick that is every ~5
ticks, so the commitment is cancelled by the motion it commands (measured:
96.9% of picks were tile-change replans, 33.8% of picks reversed direction).
Adds four env knobs, every default reproducing the shipped mover
byte-for-byte:
TR_TFIL_TILE_REPLAN self (default) | off | enemy
TR_TFIL_COMMIT_TICKS 15 (default)
TR_TFIL_NO_REV 0 (default)
TR_TFIL_COMMIT_LOG off (default, JSONL per-tick diagnostics)
- `off` honours the commitment; the danger replan stays the safety valve.
- `enemy` keys the cancel to the TARGET's tile displacement (the intent the
original comment claimed).
- `TR_TFIL_NO_REV` down-weights (never filters) tiles >90 deg from the travel
direction; the pool can never be emptied.
Default-path parity is guarded by test_tfil_commit_env.nim, which replays
tools/fixtures/tr_drussgt_vs_modularbot.jsonl and diffs every move command
against a golden generated from the pre-change build (git archive f842ac0).
env_report known-name list updated for the four new names.