4 arms x 7 runs vs real DrussGT. mix alternates the two guns 476 times/7 runs (liveness OK) but our bullets are no more varied (power sd / aim-offset sd flat) and DrussGT's dodge quality is unchanged (miss/tick mix-pat +0.03, p=0.66; MDE 3.8%). mix wins 24/49 = the 49% baseline; the user's 6/10 has P=0.353 at 49%. New tools/ab/ab_dodge_analyze.py splits the validated per-shot dodge instrument by arm and adds gun-switch/power/bearing liveness; fixtures committed.
10 KiB
Does mixing guns disrupt DrussGT? No.
The user's hypothesis. "Maybe lucky, but sometimes the BitBrain entered and
disrupted DrussGT's surfing data, allowing TMHorizon to hit more." The claim is
that admitting both TR_RACK_TMHORIZON=both and TR_RACK_BITBRAIN=both
poisons DrussGT's surfgun: two guns that disagree fire inconsistent bullets, its
wave matching fails, and it dodges worse.
Verdict: CLEAN NEGATIVE. The mixing is real (the bot alternates the two guns 476 times per 7 runs — measured), but our bullets are not more varied (power and firing-bearing spread are flat), DrussGT's dodge quality is unchanged (miss/tick, fixed-12 and control-window lateral displacement all non-significant, run-cluster p = 0.28-0.80, and the one marginal result leans the wrong way), and the win difference is noise. The 6/10 was luck.
Design: shared A/B harness tools/ab/ab_run.sh (commit 17c5015), frozen
ModularBot built from HEAD f91e121, 4 arms x 7 runs x 7 rounds = 196 rounds vs
the real, unmodified DrussGT. Dodge instrument reuses
common_libs/tests/analyze_drussgt_dodge_vs_power.py (commit 1adefab) per arm,
driven by the new tools/ab/ab_dodge_analyze.py. Attribution cross-check:
199/199 death events have the mapped victim at ~0 energy.
1. The 6/10 reality check (do this first)
The user watched one 10-round battle and saw 6 round wins. Shipped baseline vs real DrussGT is 24/49 = 49% round wins (replicated). Under a true 49% rate:
| quantity | value |
|---|---|
| P(exactly 6 of 10) | 0.197 |
| P(at least 6 of 10) | 0.353 |
| P(at least 6 of 10) if the rate were a coin flip (0.50) | 0.377 |
A 6/10 happens 35% of the time at the honest 49% rate — almost the same as at a pure coin flip. One 10-round battle is no evidence of anything; it is a single draw from a distribution whose mean is 4.9 wins. This section is not a caveat, it is the answer to the "6/10 proves it" reading: it does not.
The 4-arm experiment below ran 28 such battles; the mix arm won 24/49 = 49.0%
— numerically the historical baseline, but its own in-session control pat won
42.9% and the difference is noise (Fisher p = 0.69, see §2).
2. Damage and round wins (the primary, weak metric)
TR_RACK_PATTERN=off in every non-pat arm.
| arm | configuration | dmg/run | dmg taken/run | round wins | win% | shots/run | hits taken/run |
|---|---|---|---|---|---|---|---|
pat |
shipped default: Pattern only | 291 | 213 | 21/49 | 42.9% | 815 | 95.7 |
tmh |
TMHorizon only | 265 | 211 | 18/49 | 36.7% | 788 | 94.9 |
bb |
BitBrain only (MEM=decay) |
278 | 183 | 22/49 | 44.9% | 802 | 92.3 |
mix |
TMHorizon + BitBrain (user's config) | 267 | 210 | 24/49 | 49.0% | 804 | 93.6 |
Per-run round wins (wins cluster at 0/7, so the mean alone lies):
pat r1..r7: 5 3 3 4 3 1 2 tmh: 1 4 2 2 3 2 4
bb : 2 3 4 3 3 4 3 mix: 2 5 2 3 3 4 5
Nothing separates. Against pat: mix damage/run -24.9 (perm p = 0.097), i.e.
mix deals less damage; round wins +0.43/run (perm p = 0.68; pooled Fisher
p = 0.69). The two point estimates even disagree in sign — the signature of noise.
Minimum detectable effect at 7 runs/arm (alpha 0.05, 80% power): dmg/run 39.7 (13.6% of the control mean) and round wins 1.93 of a 3.0 mean (64%). Seven runs cannot resolve anything smaller than a huge effect, so this metric is weak by construction and is not where the hypothesis is tested.
3. DrussGT's dodge quality, per arm (the strong, per-shot metric)
Per shot fired by us. miss/tick and lat12 are power-neutral; lat_ctrl is the
same shot 40 ticks later, bullet gone (a real response is absent there; a
geometry/phase difference is present). Disruption = DrussGT ends up closer to our
aim line = these metrics FALL. 95% CI = run-cluster bootstrap (2 000 reps):
5 417-5 586 shots/arm, i.e. ~10x the power of the win counts.
| arm | shots | miss px | miss/tick [95% CI] | lat12 px [95% CI] | lat_ctrl px [95% CI] | our hit% |
|---|---|---|---|---|---|---|
pat |
5586 | 117.0 | 4.64 [4.56, 4.72] | 53.9 [52.5, 55.1] | 51.4 [50.4, 52.3] | 0.107 |
tmh |
5417 | 121.6 | 4.74 [4.68, 4.80] | 54.6 [53.4, 55.5] | 52.7 [52.1, 53.3] | 0.099 |
bb |
5483 | 118.6 | 4.66 [4.59, 4.72] | 54.0 [53.4, 54.7] | 52.1 [51.3, 52.8] | 0.099 |
mix |
5518 | 117.9 | 4.67 [4.58, 4.76] | 54.4 [53.8, 55.0] | 52.6 [52.1, 53.1] | 0.101 |
Between-arm contrast (exact run-cluster permutation, C(14,7)=3432; strat is
range-stratified over 200-450 px+, n-weighted). mix vs the best single arm:
| metric | mix - pat |
perm p | mix - bb |
perm p | mix - tmh |
perm p |
|---|---|---|---|---|---|---|
| miss/tick | +0.030 | 0.66 | +0.014 | 0.80 | -0.065 | 0.28 |
| lat12 | +0.50 | 0.58 | +0.37 | 0.48 | -0.19 | 0.79 |
| lat_ctrl | +1.24 | 0.056 | +0.54 | 0.28 | -0.10 | 0.80 |
| our hit% | -0.006 | 0.13 | +0.001 | 0.73 | +0.001 | 0.70 |
(The only result anywhere near significance, mix-pat lat_ctrl at p = 0.056, has
mix higher — DrussGT drifting further off our line, i.e. dodging no worse.)
Against pat (the arm that makes DrussGT look worst-dodging) mix is, if
anything, slightly higher on every metric — the opposite of disruption, and
nowhere near significance. Against tmh/bb the sign flips. There is no
dodge-quality degradation under mix.
Minimum detectable effect at 7 runs (run-cluster, 80% power): miss/tick 0.176 px/tick (3.8%), lat12 2.73 px (5.1%), lat_ctrl 2.06 px (4.0%), miss 3.65 px (3.1%). So the experiment bounds any dodge-quality change to <~4%; the observed mix-vs-best difference is +0.03 px/tick, ~6x below the floor.
The trap to avoid (this team already hit it once)
bb lands the fewest hits (0.099 vs pat 0.107, cluster p = 0.028 — real) yet
wins more rounds (22/49 vs 21/49) and deals comparable damage (278 vs 291).
Judging by hit rate alone would condemn bb; the objective metrics do not. Hit
rate is not the verdict — dodge quality is the mechanism, damage/wins are the
outcome, and all three are reported above.
4. Liveness: is the mix real, and are our bullets more varied?
Gun selection is real (MEASURED, from the bot's own [config] switch lines).
The mix genuinely alternates the two guns; the single-gun arms never switch:
| arm | guns selected ([config] lines) | gun switches (7 runs) |
|---|---|---|
pat |
Pattern: 54 | 0 |
tmh |
TMHorizon: 54 | 0 |
bb |
BitBrain: 55 | 0 |
mix |
TMHorizon: 275, BitBrain: 258 | 476 (68/run) |
Our bullets are NOT more varied (MEASURED). Power is set by the shared energy/range policy, independent of the selected gun; the fired power distribution is identical across arms. The firing-bearing spread (aim offset from the straight-at-target line, degrees) is also flat — the mix's variance (14.30^2) is not larger than either component (14.32^2, 14.33^2), so the two guns do not even have a measurably different marginal aim:
| arm | power mean | power sd | power IQR | aim-offset sd | aim-offset IQR | fire range | gun switches |
|---|---|---|---|---|---|---|---|
pat |
0.810 | 0.262 | 0.500 | 14.16 | 22.55 | 470 px | 0 |
tmh |
0.828 | 0.253 | 0.500 | 14.32 | 23.13 | 480 px | 0 |
bb |
0.829 | 0.248 | 0.500 | 14.33 | 22.82 | 473 px | 0 |
mix |
0.811 | 0.264 | 0.500 | 14.30 | 23.20 | 469 px | 476 |
This is the hypothesis failing its own liveness test. The mechanism needs more varied / inconsistent bullets, and there are none: mixing guns changes nothing about the bullet stream (same powers, statistically identical aim spread, same range). Caveat: the aim-offset marginal is geometry-dominated (sd ~14 deg from range/target motion), so it is a weak discriminator for a few-degree gun disagreement; but the power channel — the one the hypothesis names ("mixed speeds/powers") — is structurally gun-independent and measured flat, and the gun-switch liveness proves the mix did fire both guns.
5. Rubric answer
The task set three possible outcomes. This is the third:
- DrussGT dodge quality degrades under
mix-> mechanism REAL. Not observed. - Dodge quality unchanged but our hit rate higher under
mix-> just noisier aim. Not observed (hit rate flat: mix 0.101 vs pat 0.107, p = 0.13). - Nothing separates -> CLEAN NEGATIVE; the 6/10 was luck. <- THIS.
MEASURED
- The two guns are both selected and alternate heavily under
mix(476 switches / 7 runs; 0 for every single-gun arm). - DrussGT's power-neutral dodge metrics (miss/tick, fixed-12, control-window) are
statistically indistinguishable across all four arms;
mix-pat= +0.03 px/tick (p = 0.66), bounded by a 3.8% MDE. - Our own bullets are no more varied under
mix: power sd 0.264 vs 0.248-0.262, aim-offset sd 14.30 vs 14.16-14.33, same range. Power is a gun-independent policy, so the "mixed powers" channel does not exist here at all. - Damage/run and round wins do not separate either (perm p = 0.097 and 0.68);
mixwon 24/49 = 49.0% vs in-session controlpat21/49 = 42.9% (Fisher p = 0.69), while dealing less damage. The 6/10 has probability 0.353. - The trap is live in this data:
bblands significantly fewer hits thanpat(p = 0.028) yet wins at least as many rounds.
INFERRED
- Two alternating guns sharing the same power policy and (marginal) aim distribution do not poison DrussGT's surfgun. Deliberate angle jitter or a randomized power band would be a different intervention, not this one; the liveness table says this mix does not even change the bullet statistics.
- The 6/10 was a lucky draw, not a signal.
6. Reproducing
cat > /tmp/ab/mix_arms.txt <<'EOF'
pat | | shipped default: Pattern only
tmh | TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both | TMHorizon alone
bb | TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay | BitBrain alone
mix | TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay | user config
EOF
tools/ab/ab_run.sh --arms /tmp/ab/mix_arms.txt --runs 7 --outdir /tmp/ab/mix --conc 5
python3 tools/ab/ab_analyze.py /tmp/ab/mix
python3 tools/ab/ab_dodge_analyze.py /tmp/ab/mix --json /tmp/ab/mix_dodge.json --reps 2000
Captured outputs are committed (the ~246 MB of raw .jsonl captures are
gitignored, as for drussgt_dodge_vs_power.md):
common_libs/tests/fixtures/gun_mix_ab_report.txt—ab_analyze.pyoutputcommon_libs/tests/fixtures/gun_mix_dodge_report.txt— per-arm dodge reportcommon_libs/tests/fixtures/gun_mix_dodge_results.json— the same numbers as JSON