# Does mixing guns disrupt DrussGT? No. **The user's hypothesis.** *"Maybe lucky, but sometimes the BitBrain entered and disrupted DrussGT's surfing data, allowing TMHorizon to hit more."* The claim is that admitting **both** `TR_RACK_TMHORIZON=both` and `TR_RACK_BITBRAIN=both` poisons DrussGT's surfgun: two guns that disagree fire inconsistent bullets, its wave matching fails, and it dodges worse. **Verdict: CLEAN NEGATIVE.** The mixing is real (the bot alternates the two guns 476 times per 7 runs — measured), but our bullets are **not** more varied (power and firing-bearing spread are flat), DrussGT's dodge quality is **unchanged** (miss/tick, fixed-12 and control-window lateral displacement all non-significant, run-cluster p = 0.28-0.80, and the one marginal result leans the wrong way), and the win difference is noise. The 6/10 was luck. Design: shared A/B harness `tools/ab/ab_run.sh` (commit `17c5015`), frozen ModularBot built from HEAD `f91e121`, **4 arms x 7 runs x 7 rounds = 196 rounds vs the real, unmodified DrussGT**. Dodge instrument reuses `common_libs/tests/analyze_drussgt_dodge_vs_power.py` (commit `1adefab`) per arm, driven by the new `tools/ab/ab_dodge_analyze.py`. Attribution cross-check: **199/199** death events have the mapped victim at ~0 energy. --- ## 1. The 6/10 reality check (do this first) The user watched **one** 10-round battle and saw 6 round wins. Shipped baseline vs real DrussGT is 24/49 = **49%** round wins (replicated). Under a true 49% rate: | quantity | value | |---|---| | P(exactly 6 of 10) | **0.197** | | P(at least 6 of 10) | **0.353** | | P(at least 6 of 10) if the rate were a coin flip (0.50) | 0.377 | A 6/10 happens **35% of the time** at the honest 49% rate — almost the same as at a pure coin flip. One 10-round battle is **no evidence of anything**; it is a single draw from a distribution whose mean is 4.9 wins. This section is not a caveat, it is the answer to the "6/10 proves it" reading: **it does not**. The 4-arm experiment below ran 28 such battles; the mix arm won **24/49 = 49.0%** — numerically the historical baseline, but its own in-session control `pat` won 42.9% and the difference is noise (Fisher p = 0.69, see §2). --- ## 2. Damage and round wins (the primary, weak metric) `TR_RACK_PATTERN=off` in every non-`pat` arm. | arm | configuration | dmg/run | dmg taken/run | round wins | win% | shots/run | hits taken/run | |---|---|---|---|---|---|---|---| | `pat` | shipped default: Pattern only | 291 | 213 | 21/49 | 42.9% | 815 | 95.7 | | `tmh` | TMHorizon only | 265 | 211 | 18/49 | 36.7% | 788 | 94.9 | | `bb` | BitBrain only (`MEM=decay`) | 278 | 183 | 22/49 | 44.9% | 802 | 92.3 | | `mix` | TMHorizon **+** BitBrain (user's config) | 267 | 210 | **24/49** | **49.0%** | 804 | 93.6 | Per-run round wins (wins cluster at 0/7, so the mean alone lies): ``` pat r1..r7: 5 3 3 4 3 1 2 tmh: 1 4 2 2 3 2 4 bb : 2 3 4 3 3 4 3 mix: 2 5 2 3 3 4 5 ``` Nothing separates. Against `pat`: `mix` damage/run **-24.9** (perm p = 0.097), i.e. mix deals *less* damage; round wins **+0.43/run** (perm p = 0.68; pooled Fisher p = 0.69). The two point estimates even disagree in sign — the signature of noise. **Minimum detectable effect at 7 runs/arm (alpha 0.05, 80% power):** dmg/run **39.7** (13.6% of the control mean) and round wins **1.93 of a 3.0 mean (64%)**. Seven runs cannot resolve anything smaller than a huge effect, so this metric is weak by construction and is **not** where the hypothesis is tested. --- ## 3. DrussGT's dodge quality, per arm (the strong, per-shot metric) Per shot fired by us. `miss/tick` and `lat12` are power-neutral; `lat_ctrl` is the same shot 40 ticks later, bullet gone (a real response is absent there; a geometry/phase difference is present). **Disruption = DrussGT ends up closer to our aim line = these metrics FALL.** 95% CI = run-cluster bootstrap (2 000 reps): 5 417-5 586 shots/arm, i.e. ~10x the power of the win counts. | arm | shots | miss px | miss/tick [95% CI] | lat12 px [95% CI] | lat_ctrl px [95% CI] | our hit% | |---|---|---|---|---|---|---| | `pat` | 5586 | 117.0 | 4.64 [4.56, 4.72] | 53.9 [52.5, 55.1] | 51.4 [50.4, 52.3] | 0.107 | | `tmh` | 5417 | 121.6 | 4.74 [4.68, 4.80] | 54.6 [53.4, 55.5] | 52.7 [52.1, 53.3] | 0.099 | | `bb` | 5483 | 118.6 | 4.66 [4.59, 4.72] | 54.0 [53.4, 54.7] | 52.1 [51.3, 52.8] | 0.099 | | `mix` | 5518 | 117.9 | 4.67 [4.58, 4.76] | 54.4 [53.8, 55.0] | 52.6 [52.1, 53.1] | 0.101 | Between-arm contrast (exact **run-cluster** permutation, C(14,7)=3432; `strat` is range-stratified over 200-450 px+, n-weighted). `mix` vs the best single arm: | metric | `mix` - `pat` | perm p | `mix` - `bb` | perm p | `mix` - `tmh` | perm p | |---|---|---|---|---|---|---| | miss/tick | +0.030 | 0.66 | +0.014 | 0.80 | -0.065 | 0.28 | | lat12 | +0.50 | 0.58 | +0.37 | 0.48 | -0.19 | 0.79 | | lat_ctrl | +1.24 | 0.056 | +0.54 | 0.28 | -0.10 | 0.80 | | our hit% | -0.006 | 0.13 | +0.001 | 0.73 | +0.001 | 0.70 | (The only result anywhere near significance, `mix`-`pat` lat_ctrl at p = 0.056, has `mix` *higher* — DrussGT drifting **further** off our line, i.e. dodging no worse.) Against `pat` (the arm that makes DrussGT look *worst*-dodging) `mix` is, if anything, **slightly higher** on every metric — the opposite of disruption, and nowhere near significance. Against `tmh`/`bb` the sign flips. There is **no dodge-quality degradation under `mix`**. **Minimum detectable effect at 7 runs (run-cluster, 80% power):** miss/tick **0.176 px/tick (3.8%)**, lat12 2.73 px (5.1%), lat_ctrl 2.06 px (4.0%), miss 3.65 px (3.1%). So the experiment bounds any dodge-quality change to **<~4%**; the observed mix-vs-best difference is +0.03 px/tick, ~6x below the floor. ### The trap to avoid (this team already hit it once) `bb` lands the **fewest hits** (0.099 vs `pat` 0.107, cluster p = 0.028 — real) yet wins **more** rounds (22/49 vs 21/49) and deals comparable damage (278 vs 291). Judging by hit rate alone would condemn `bb`; the objective metrics do not. Hit rate is *not* the verdict — dodge quality is the mechanism, damage/wins are the outcome, and all three are reported above. --- ## 4. Liveness: is the mix real, and are our bullets more varied? **Gun selection is real (MEASURED, from the bot's own `[config]` switch lines).** The mix genuinely alternates the two guns; the single-gun arms never switch: | arm | guns selected ([config] lines) | gun switches (7 runs) | |---|---|---| | `pat` | Pattern: 54 | **0** | | `tmh` | TMHorizon: 54 | **0** | | `bb` | BitBrain: 55 | **0** | | `mix` | TMHorizon: 275, BitBrain: 258 | **476** (68/run) | **Our bullets are NOT more varied (MEASURED).** Power is set by the shared energy/range policy, independent of the selected gun; the fired power distribution is identical across arms. The firing-bearing spread (aim offset from the straight-at-target line, degrees) is also flat — the mix's variance (14.30^2) is *not* larger than either component (14.32^2, 14.33^2), so the two guns do not even have a measurably different marginal aim: | arm | power mean | power sd | power IQR | aim-offset sd | aim-offset IQR | fire range | gun switches | |---|---|---|---|---|---|---|---| | `pat` | 0.810 | 0.262 | 0.500 | 14.16 | 22.55 | 470 px | 0 | | `tmh` | 0.828 | 0.253 | 0.500 | 14.32 | 23.13 | 480 px | 0 | | `bb` | 0.829 | 0.248 | 0.500 | 14.33 | 22.82 | 473 px | 0 | | `mix` | 0.811 | 0.264 | 0.500 | 14.30 | 23.20 | 469 px | 476 | **This is the hypothesis failing its own liveness test.** The mechanism needs more varied / inconsistent bullets, and there are none: mixing guns changes *nothing* about the bullet stream (same powers, statistically identical aim spread, same range). Caveat: the aim-offset marginal is geometry-dominated (sd ~14 deg from range/target motion), so it is a *weak* discriminator for a few-degree gun disagreement; but the power channel — the one the hypothesis names ("mixed speeds/powers") — is structurally gun-independent and measured flat, and the gun-switch liveness proves the mix did fire both guns. --- ## 5. Rubric answer The task set three possible outcomes. This is the third: * DrussGT dodge quality **degrades** under `mix` -> mechanism REAL. **Not observed.** * Dodge quality unchanged but our hit rate higher under `mix` -> just noisier aim. **Not observed** (hit rate flat: mix 0.101 vs pat 0.107, p = 0.13). * **Nothing separates -> CLEAN NEGATIVE; the 6/10 was luck.** **<- THIS.** **MEASURED** * The two guns are both selected and alternate heavily under `mix` (476 switches / 7 runs; 0 for every single-gun arm). * DrussGT's power-neutral dodge metrics (miss/tick, fixed-12, control-window) are statistically indistinguishable across all four arms; `mix` - `pat` = +0.03 px/tick (p = 0.66), bounded by a 3.8% MDE. * Our own bullets are no more varied under `mix`: power sd 0.264 vs 0.248-0.262, aim-offset sd 14.30 vs 14.16-14.33, same range. Power is a gun-independent policy, so the "mixed powers" channel does not exist here at all. * Damage/run and round wins do not separate either (perm p = 0.097 and 0.68); `mix` won 24/49 = 49.0% vs in-session control `pat` 21/49 = 42.9% (Fisher p = 0.69), while dealing *less* damage. The 6/10 has probability 0.353. * The trap is live in this data: `bb` lands significantly fewer hits than `pat` (p = 0.028) yet wins at least as many rounds. **INFERRED** * Two alternating guns sharing the same power policy and (marginal) aim distribution do not poison DrussGT's surfgun. Deliberate angle jitter or a randomized power band would be a *different* intervention, not this one; the liveness table says this mix does not even change the bullet statistics. * The 6/10 was a lucky draw, not a signal. --- ## 6. Reproducing ```bash cat > /tmp/ab/mix_arms.txt <<'EOF' pat | | shipped default: Pattern only tmh | TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both | TMHorizon alone bb | TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay | BitBrain alone mix | TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay | user config EOF tools/ab/ab_run.sh --arms /tmp/ab/mix_arms.txt --runs 7 --outdir /tmp/ab/mix --conc 5 python3 tools/ab/ab_analyze.py /tmp/ab/mix python3 tools/ab/ab_dodge_analyze.py /tmp/ab/mix --json /tmp/ab/mix_dodge.json --reps 2000 ``` Captured outputs are committed (the ~246 MB of raw `.jsonl` captures are gitignored, as for `drussgt_dodge_vs_power.md`): * `common_libs/tests/fixtures/gun_mix_ab_report.txt` — `ab_analyze.py` output * `common_libs/tests/fixtures/gun_mix_dodge_report.txt` — per-arm dodge report * `common_libs/tests/fixtures/gun_mix_dodge_results.json` — the same numbers as JSON