Files
SirRoboGarage/docs/gun_mix_disruption.md
SirStone 32a5e72fac Gun mixing (TMHorizon+BitBrain) vs DrussGT: clean negative, no dodge disruption
4 arms x 7 runs vs real DrussGT. mix alternates the two guns 476 times/7 runs
(liveness OK) but our bullets are no more varied (power sd / aim-offset sd flat)
and DrussGT's dodge quality is unchanged (miss/tick mix-pat +0.03, p=0.66; MDE
3.8%). mix wins 24/49 = the 49% baseline; the user's 6/10 has P=0.353 at 49%.
New tools/ab/ab_dodge_analyze.py splits the validated per-shot dodge instrument
by arm and adds gun-switch/power/bearing liveness; fixtures committed.
2026-09-24 23:07:05 +02:00

10 KiB

Does mixing guns disrupt DrussGT? No.

The user's hypothesis. "Maybe lucky, but sometimes the BitBrain entered and disrupted DrussGT's surfing data, allowing TMHorizon to hit more." The claim is that admitting both TR_RACK_TMHORIZON=both and TR_RACK_BITBRAIN=both poisons DrussGT's surfgun: two guns that disagree fire inconsistent bullets, its wave matching fails, and it dodges worse.

Verdict: CLEAN NEGATIVE. The mixing is real (the bot alternates the two guns 476 times per 7 runs — measured), but our bullets are not more varied (power and firing-bearing spread are flat), DrussGT's dodge quality is unchanged (miss/tick, fixed-12 and control-window lateral displacement all non-significant, run-cluster p = 0.28-0.80, and the one marginal result leans the wrong way), and the win difference is noise. The 6/10 was luck.

Design: shared A/B harness tools/ab/ab_run.sh (commit 17c5015), frozen ModularBot built from HEAD f91e121, 4 arms x 7 runs x 7 rounds = 196 rounds vs the real, unmodified DrussGT. Dodge instrument reuses common_libs/tests/analyze_drussgt_dodge_vs_power.py (commit 1adefab) per arm, driven by the new tools/ab/ab_dodge_analyze.py. Attribution cross-check: 199/199 death events have the mapped victim at ~0 energy.


1. The 6/10 reality check (do this first)

The user watched one 10-round battle and saw 6 round wins. Shipped baseline vs real DrussGT is 24/49 = 49% round wins (replicated). Under a true 49% rate:

quantity value
P(exactly 6 of 10) 0.197
P(at least 6 of 10) 0.353
P(at least 6 of 10) if the rate were a coin flip (0.50) 0.377

A 6/10 happens 35% of the time at the honest 49% rate — almost the same as at a pure coin flip. One 10-round battle is no evidence of anything; it is a single draw from a distribution whose mean is 4.9 wins. This section is not a caveat, it is the answer to the "6/10 proves it" reading: it does not.

The 4-arm experiment below ran 28 such battles; the mix arm won 24/49 = 49.0% — numerically the historical baseline, but its own in-session control pat won 42.9% and the difference is noise (Fisher p = 0.69, see §2).


2. Damage and round wins (the primary, weak metric)

TR_RACK_PATTERN=off in every non-pat arm.

arm configuration dmg/run dmg taken/run round wins win% shots/run hits taken/run
pat shipped default: Pattern only 291 213 21/49 42.9% 815 95.7
tmh TMHorizon only 265 211 18/49 36.7% 788 94.9
bb BitBrain only (MEM=decay) 278 183 22/49 44.9% 802 92.3
mix TMHorizon + BitBrain (user's config) 267 210 24/49 49.0% 804 93.6

Per-run round wins (wins cluster at 0/7, so the mean alone lies):

pat  r1..r7: 5 3 3 4 3 1 2      tmh: 1 4 2 2 3 2 4
bb : 2 3 4 3 3 4 3             mix: 2 5 2 3 3 4 5

Nothing separates. Against pat: mix damage/run -24.9 (perm p = 0.097), i.e. mix deals less damage; round wins +0.43/run (perm p = 0.68; pooled Fisher p = 0.69). The two point estimates even disagree in sign — the signature of noise.

Minimum detectable effect at 7 runs/arm (alpha 0.05, 80% power): dmg/run 39.7 (13.6% of the control mean) and round wins 1.93 of a 3.0 mean (64%). Seven runs cannot resolve anything smaller than a huge effect, so this metric is weak by construction and is not where the hypothesis is tested.


3. DrussGT's dodge quality, per arm (the strong, per-shot metric)

Per shot fired by us. miss/tick and lat12 are power-neutral; lat_ctrl is the same shot 40 ticks later, bullet gone (a real response is absent there; a geometry/phase difference is present). Disruption = DrussGT ends up closer to our aim line = these metrics FALL. 95% CI = run-cluster bootstrap (2 000 reps): 5 417-5 586 shots/arm, i.e. ~10x the power of the win counts.

arm shots miss px miss/tick [95% CI] lat12 px [95% CI] lat_ctrl px [95% CI] our hit%
pat 5586 117.0 4.64 [4.56, 4.72] 53.9 [52.5, 55.1] 51.4 [50.4, 52.3] 0.107
tmh 5417 121.6 4.74 [4.68, 4.80] 54.6 [53.4, 55.5] 52.7 [52.1, 53.3] 0.099
bb 5483 118.6 4.66 [4.59, 4.72] 54.0 [53.4, 54.7] 52.1 [51.3, 52.8] 0.099
mix 5518 117.9 4.67 [4.58, 4.76] 54.4 [53.8, 55.0] 52.6 [52.1, 53.1] 0.101

Between-arm contrast (exact run-cluster permutation, C(14,7)=3432; strat is range-stratified over 200-450 px+, n-weighted). mix vs the best single arm:

metric mix - pat perm p mix - bb perm p mix - tmh perm p
miss/tick +0.030 0.66 +0.014 0.80 -0.065 0.28
lat12 +0.50 0.58 +0.37 0.48 -0.19 0.79
lat_ctrl +1.24 0.056 +0.54 0.28 -0.10 0.80
our hit% -0.006 0.13 +0.001 0.73 +0.001 0.70

(The only result anywhere near significance, mix-pat lat_ctrl at p = 0.056, has mix higher — DrussGT drifting further off our line, i.e. dodging no worse.)

Against pat (the arm that makes DrussGT look worst-dodging) mix is, if anything, slightly higher on every metric — the opposite of disruption, and nowhere near significance. Against tmh/bb the sign flips. There is no dodge-quality degradation under mix.

Minimum detectable effect at 7 runs (run-cluster, 80% power): miss/tick 0.176 px/tick (3.8%), lat12 2.73 px (5.1%), lat_ctrl 2.06 px (4.0%), miss 3.65 px (3.1%). So the experiment bounds any dodge-quality change to <~4%; the observed mix-vs-best difference is +0.03 px/tick, ~6x below the floor.

The trap to avoid (this team already hit it once)

bb lands the fewest hits (0.099 vs pat 0.107, cluster p = 0.028 — real) yet wins more rounds (22/49 vs 21/49) and deals comparable damage (278 vs 291). Judging by hit rate alone would condemn bb; the objective metrics do not. Hit rate is not the verdict — dodge quality is the mechanism, damage/wins are the outcome, and all three are reported above.


4. Liveness: is the mix real, and are our bullets more varied?

Gun selection is real (MEASURED, from the bot's own [config] switch lines). The mix genuinely alternates the two guns; the single-gun arms never switch:

arm guns selected ([config] lines) gun switches (7 runs)
pat Pattern: 54 0
tmh TMHorizon: 54 0
bb BitBrain: 55 0
mix TMHorizon: 275, BitBrain: 258 476 (68/run)

Our bullets are NOT more varied (MEASURED). Power is set by the shared energy/range policy, independent of the selected gun; the fired power distribution is identical across arms. The firing-bearing spread (aim offset from the straight-at-target line, degrees) is also flat — the mix's variance (14.30^2) is not larger than either component (14.32^2, 14.33^2), so the two guns do not even have a measurably different marginal aim:

arm power mean power sd power IQR aim-offset sd aim-offset IQR fire range gun switches
pat 0.810 0.262 0.500 14.16 22.55 470 px 0
tmh 0.828 0.253 0.500 14.32 23.13 480 px 0
bb 0.829 0.248 0.500 14.33 22.82 473 px 0
mix 0.811 0.264 0.500 14.30 23.20 469 px 476

This is the hypothesis failing its own liveness test. The mechanism needs more varied / inconsistent bullets, and there are none: mixing guns changes nothing about the bullet stream (same powers, statistically identical aim spread, same range). Caveat: the aim-offset marginal is geometry-dominated (sd ~14 deg from range/target motion), so it is a weak discriminator for a few-degree gun disagreement; but the power channel — the one the hypothesis names ("mixed speeds/powers") — is structurally gun-independent and measured flat, and the gun-switch liveness proves the mix did fire both guns.


5. Rubric answer

The task set three possible outcomes. This is the third:

  • DrussGT dodge quality degrades under mix -> mechanism REAL. Not observed.
  • Dodge quality unchanged but our hit rate higher under mix -> just noisier aim. Not observed (hit rate flat: mix 0.101 vs pat 0.107, p = 0.13).
  • Nothing separates -> CLEAN NEGATIVE; the 6/10 was luck. <- THIS.

MEASURED

  • The two guns are both selected and alternate heavily under mix (476 switches / 7 runs; 0 for every single-gun arm).
  • DrussGT's power-neutral dodge metrics (miss/tick, fixed-12, control-window) are statistically indistinguishable across all four arms; mix - pat = +0.03 px/tick (p = 0.66), bounded by a 3.8% MDE.
  • Our own bullets are no more varied under mix: power sd 0.264 vs 0.248-0.262, aim-offset sd 14.30 vs 14.16-14.33, same range. Power is a gun-independent policy, so the "mixed powers" channel does not exist here at all.
  • Damage/run and round wins do not separate either (perm p = 0.097 and 0.68); mix won 24/49 = 49.0% vs in-session control pat 21/49 = 42.9% (Fisher p = 0.69), while dealing less damage. The 6/10 has probability 0.353.
  • The trap is live in this data: bb lands significantly fewer hits than pat (p = 0.028) yet wins at least as many rounds.

INFERRED

  • Two alternating guns sharing the same power policy and (marginal) aim distribution do not poison DrussGT's surfgun. Deliberate angle jitter or a randomized power band would be a different intervention, not this one; the liveness table says this mix does not even change the bullet statistics.
  • The 6/10 was a lucky draw, not a signal.

6. Reproducing

cat > /tmp/ab/mix_arms.txt <<'EOF'
pat | | shipped default: Pattern only
tmh | TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both | TMHorizon alone
bb | TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay | BitBrain alone
mix | TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay | user config
EOF
tools/ab/ab_run.sh --arms /tmp/ab/mix_arms.txt --runs 7 --outdir /tmp/ab/mix --conc 5
python3 tools/ab/ab_analyze.py /tmp/ab/mix
python3 tools/ab/ab_dodge_analyze.py /tmp/ab/mix --json /tmp/ab/mix_dodge.json --reps 2000

Captured outputs are committed (the ~246 MB of raw .jsonl captures are gitignored, as for drussgt_dodge_vs_power.md):

  • common_libs/tests/fixtures/gun_mix_ab_report.txt — ab_analyze.py output
  • common_libs/tests/fixtures/gun_mix_dodge_report.txt — per-arm dodge report
  • common_libs/tests/fixtures/gun_mix_dodge_results.json — the same numbers as JSON