32a5e72fac
4 arms x 7 runs vs real DrussGT. mix alternates the two guns 476 times/7 runs (liveness OK) but our bullets are no more varied (power sd / aim-offset sd flat) and DrussGT's dodge quality is unchanged (miss/tick mix-pat +0.03, p=0.66; MDE 3.8%). mix wins 24/49 = the 49% baseline; the user's 6/10 has P=0.353 at 49%. New tools/ab/ab_dodge_analyze.py splits the validated per-shot dodge instrument by arm and adds gun-switch/power/bearing liveness; fixtures committed.
216 lines
10 KiB
Markdown
216 lines
10 KiB
Markdown
# Does mixing guns disrupt DrussGT? No.
|
|
|
|
**The user's hypothesis.** *"Maybe lucky, but sometimes the BitBrain entered and
|
|
disrupted DrussGT's surfing data, allowing TMHorizon to hit more."* The claim is
|
|
that admitting **both** `TR_RACK_TMHORIZON=both` and `TR_RACK_BITBRAIN=both`
|
|
poisons DrussGT's surfgun: two guns that disagree fire inconsistent bullets, its
|
|
wave matching fails, and it dodges worse.
|
|
|
|
**Verdict: CLEAN NEGATIVE.** The mixing is real (the bot alternates the two guns
|
|
476 times per 7 runs — measured), but our bullets are **not** more varied (power
|
|
and firing-bearing spread are flat), DrussGT's dodge quality is **unchanged**
|
|
(miss/tick, fixed-12 and control-window lateral displacement all non-significant,
|
|
run-cluster p = 0.28-0.80, and the one marginal result leans the wrong way), and
|
|
the win difference is noise. The 6/10 was luck.
|
|
|
|
Design: shared A/B harness `tools/ab/ab_run.sh` (commit `17c5015`), frozen
|
|
ModularBot built from HEAD `f91e121`, **4 arms x 7 runs x 7 rounds = 196 rounds vs
|
|
the real, unmodified DrussGT**. Dodge instrument reuses
|
|
`common_libs/tests/analyze_drussgt_dodge_vs_power.py` (commit `1adefab`) per arm,
|
|
driven by the new `tools/ab/ab_dodge_analyze.py`. Attribution cross-check:
|
|
**199/199** death events have the mapped victim at ~0 energy.
|
|
|
|
---
|
|
|
|
## 1. The 6/10 reality check (do this first)
|
|
|
|
The user watched **one** 10-round battle and saw 6 round wins. Shipped baseline vs
|
|
real DrussGT is 24/49 = **49%** round wins (replicated). Under a true 49% rate:
|
|
|
|
| quantity | value |
|
|
|---|---|
|
|
| P(exactly 6 of 10) | **0.197** |
|
|
| P(at least 6 of 10) | **0.353** |
|
|
| P(at least 6 of 10) if the rate were a coin flip (0.50) | 0.377 |
|
|
|
|
A 6/10 happens **35% of the time** at the honest 49% rate — almost the same as at a
|
|
pure coin flip. One 10-round battle is **no evidence of anything**; it is a single
|
|
draw from a distribution whose mean is 4.9 wins. This section is not a caveat, it
|
|
is the answer to the "6/10 proves it" reading: **it does not**.
|
|
|
|
The 4-arm experiment below ran 28 such battles; the mix arm won **24/49 = 49.0%**
|
|
— numerically the historical baseline, but its own in-session control `pat` won
|
|
42.9% and the difference is noise (Fisher p = 0.69, see §2).
|
|
|
|
---
|
|
|
|
## 2. Damage and round wins (the primary, weak metric)
|
|
|
|
`TR_RACK_PATTERN=off` in every non-`pat` arm.
|
|
|
|
| arm | configuration | dmg/run | dmg taken/run | round wins | win% | shots/run | hits taken/run |
|
|
|---|---|---|---|---|---|---|---|
|
|
| `pat` | shipped default: Pattern only | 291 | 213 | 21/49 | 42.9% | 815 | 95.7 |
|
|
| `tmh` | TMHorizon only | 265 | 211 | 18/49 | 36.7% | 788 | 94.9 |
|
|
| `bb` | BitBrain only (`MEM=decay`) | 278 | 183 | 22/49 | 44.9% | 802 | 92.3 |
|
|
| `mix` | TMHorizon **+** BitBrain (user's config) | 267 | 210 | **24/49** | **49.0%** | 804 | 93.6 |
|
|
|
|
Per-run round wins (wins cluster at 0/7, so the mean alone lies):
|
|
|
|
```
|
|
pat r1..r7: 5 3 3 4 3 1 2 tmh: 1 4 2 2 3 2 4
|
|
bb : 2 3 4 3 3 4 3 mix: 2 5 2 3 3 4 5
|
|
```
|
|
|
|
Nothing separates. Against `pat`: `mix` damage/run **-24.9** (perm p = 0.097), i.e.
|
|
mix deals *less* damage; round wins **+0.43/run** (perm p = 0.68; pooled Fisher
|
|
p = 0.69). The two point estimates even disagree in sign — the signature of noise.
|
|
|
|
**Minimum detectable effect at 7 runs/arm (alpha 0.05, 80% power):** dmg/run
|
|
**39.7** (13.6% of the control mean) and round wins **1.93 of a 3.0 mean (64%)**.
|
|
Seven runs cannot resolve anything smaller than a huge effect, so this metric is
|
|
weak by construction and is **not** where the hypothesis is tested.
|
|
|
|
---
|
|
|
|
## 3. DrussGT's dodge quality, per arm (the strong, per-shot metric)
|
|
|
|
Per shot fired by us. `miss/tick` and `lat12` are power-neutral; `lat_ctrl` is the
|
|
same shot 40 ticks later, bullet gone (a real response is absent there; a
|
|
geometry/phase difference is present). **Disruption = DrussGT ends up closer to our
|
|
aim line = these metrics FALL.** 95% CI = run-cluster bootstrap (2 000 reps):
|
|
5 417-5 586 shots/arm, i.e. ~10x the power of the win counts.
|
|
|
|
| arm | shots | miss px | miss/tick [95% CI] | lat12 px [95% CI] | lat_ctrl px [95% CI] | our hit% |
|
|
|---|---|---|---|---|---|---|
|
|
| `pat` | 5586 | 117.0 | 4.64 [4.56, 4.72] | 53.9 [52.5, 55.1] | 51.4 [50.4, 52.3] | 0.107 |
|
|
| `tmh` | 5417 | 121.6 | 4.74 [4.68, 4.80] | 54.6 [53.4, 55.5] | 52.7 [52.1, 53.3] | 0.099 |
|
|
| `bb` | 5483 | 118.6 | 4.66 [4.59, 4.72] | 54.0 [53.4, 54.7] | 52.1 [51.3, 52.8] | 0.099 |
|
|
| `mix` | 5518 | 117.9 | 4.67 [4.58, 4.76] | 54.4 [53.8, 55.0] | 52.6 [52.1, 53.1] | 0.101 |
|
|
|
|
Between-arm contrast (exact **run-cluster** permutation, C(14,7)=3432; `strat` is
|
|
range-stratified over 200-450 px+, n-weighted). `mix` vs the best single arm:
|
|
|
|
| metric | `mix` - `pat` | perm p | `mix` - `bb` | perm p | `mix` - `tmh` | perm p |
|
|
|---|---|---|---|---|---|---|
|
|
| miss/tick | +0.030 | 0.66 | +0.014 | 0.80 | -0.065 | 0.28 |
|
|
| lat12 | +0.50 | 0.58 | +0.37 | 0.48 | -0.19 | 0.79 |
|
|
| lat_ctrl | +1.24 | 0.056 | +0.54 | 0.28 | -0.10 | 0.80 |
|
|
| our hit% | -0.006 | 0.13 | +0.001 | 0.73 | +0.001 | 0.70 |
|
|
|
|
(The only result anywhere near significance, `mix`-`pat` lat_ctrl at p = 0.056, has
|
|
`mix` *higher* — DrussGT drifting **further** off our line, i.e. dodging no worse.)
|
|
|
|
Against `pat` (the arm that makes DrussGT look *worst*-dodging) `mix` is, if
|
|
anything, **slightly higher** on every metric — the opposite of disruption, and
|
|
nowhere near significance. Against `tmh`/`bb` the sign flips. There is **no
|
|
dodge-quality degradation under `mix`**.
|
|
|
|
**Minimum detectable effect at 7 runs (run-cluster, 80% power):** miss/tick
|
|
**0.176 px/tick (3.8%)**, lat12 2.73 px (5.1%), lat_ctrl 2.06 px (4.0%), miss
|
|
3.65 px (3.1%). So the experiment bounds any dodge-quality change to **<~4%**;
|
|
the observed mix-vs-best difference is +0.03 px/tick, ~6x below the floor.
|
|
|
|
### The trap to avoid (this team already hit it once)
|
|
|
|
`bb` lands the **fewest hits** (0.099 vs `pat` 0.107, cluster p = 0.028 — real) yet
|
|
wins **more** rounds (22/49 vs 21/49) and deals comparable damage (278 vs 291).
|
|
Judging by hit rate alone would condemn `bb`; the objective metrics do not. Hit
|
|
rate is *not* the verdict — dodge quality is the mechanism, damage/wins are the
|
|
outcome, and all three are reported above.
|
|
|
|
---
|
|
|
|
## 4. Liveness: is the mix real, and are our bullets more varied?
|
|
|
|
**Gun selection is real (MEASURED, from the bot's own `[config]` switch lines).**
|
|
The mix genuinely alternates the two guns; the single-gun arms never switch:
|
|
|
|
| arm | guns selected ([config] lines) | gun switches (7 runs) |
|
|
|---|---|---|
|
|
| `pat` | Pattern: 54 | **0** |
|
|
| `tmh` | TMHorizon: 54 | **0** |
|
|
| `bb` | BitBrain: 55 | **0** |
|
|
| `mix` | TMHorizon: 275, BitBrain: 258 | **476** (68/run) |
|
|
|
|
**Our bullets are NOT more varied (MEASURED).** Power is set by the shared
|
|
energy/range policy, independent of the selected gun; the fired power distribution
|
|
is identical across arms. The firing-bearing spread (aim offset from the
|
|
straight-at-target line, degrees) is also flat — the mix's variance
|
|
(14.30^2) is *not* larger than either component (14.32^2, 14.33^2), so the two
|
|
guns do not even have a measurably different marginal aim:
|
|
|
|
| arm | power mean | power sd | power IQR | aim-offset sd | aim-offset IQR | fire range | gun switches |
|
|
|---|---|---|---|---|---|---|---|
|
|
| `pat` | 0.810 | 0.262 | 0.500 | 14.16 | 22.55 | 470 px | 0 |
|
|
| `tmh` | 0.828 | 0.253 | 0.500 | 14.32 | 23.13 | 480 px | 0 |
|
|
| `bb` | 0.829 | 0.248 | 0.500 | 14.33 | 22.82 | 473 px | 0 |
|
|
| `mix` | 0.811 | 0.264 | 0.500 | 14.30 | 23.20 | 469 px | 476 |
|
|
|
|
**This is the hypothesis failing its own liveness test.** The mechanism needs more
|
|
varied / inconsistent bullets, and there are none: mixing guns changes *nothing*
|
|
about the bullet stream (same powers, statistically identical aim spread, same
|
|
range). Caveat: the aim-offset marginal is geometry-dominated (sd ~14 deg from
|
|
range/target motion), so it is a *weak* discriminator for a few-degree gun
|
|
disagreement; but the power channel — the one the hypothesis names ("mixed
|
|
speeds/powers") — is structurally gun-independent and measured flat, and the
|
|
gun-switch liveness proves the mix did fire both guns.
|
|
|
|
---
|
|
|
|
## 5. Rubric answer
|
|
|
|
The task set three possible outcomes. This is the third:
|
|
|
|
* DrussGT dodge quality **degrades** under `mix` -> mechanism REAL. **Not observed.**
|
|
* Dodge quality unchanged but our hit rate higher under `mix` -> just noisier aim.
|
|
**Not observed** (hit rate flat: mix 0.101 vs pat 0.107, p = 0.13).
|
|
* **Nothing separates -> CLEAN NEGATIVE; the 6/10 was luck.** **<- THIS.**
|
|
|
|
**MEASURED**
|
|
|
|
* The two guns are both selected and alternate heavily under `mix` (476 switches /
|
|
7 runs; 0 for every single-gun arm).
|
|
* DrussGT's power-neutral dodge metrics (miss/tick, fixed-12, control-window) are
|
|
statistically indistinguishable across all four arms; `mix` - `pat` = +0.03
|
|
px/tick (p = 0.66), bounded by a 3.8% MDE.
|
|
* Our own bullets are no more varied under `mix`: power sd 0.264 vs 0.248-0.262,
|
|
aim-offset sd 14.30 vs 14.16-14.33, same range. Power is a gun-independent
|
|
policy, so the "mixed powers" channel does not exist here at all.
|
|
* Damage/run and round wins do not separate either (perm p = 0.097 and 0.68);
|
|
`mix` won 24/49 = 49.0% vs in-session control `pat` 21/49 = 42.9% (Fisher
|
|
p = 0.69), while dealing *less* damage. The 6/10 has probability 0.353.
|
|
* The trap is live in this data: `bb` lands significantly fewer hits than `pat`
|
|
(p = 0.028) yet wins at least as many rounds.
|
|
|
|
**INFERRED**
|
|
|
|
* Two alternating guns sharing the same power policy and (marginal) aim
|
|
distribution do not poison DrussGT's surfgun. Deliberate angle jitter or a
|
|
randomized power band would be a *different* intervention, not this one; the
|
|
liveness table says this mix does not even change the bullet statistics.
|
|
* The 6/10 was a lucky draw, not a signal.
|
|
|
|
---
|
|
|
|
## 6. Reproducing
|
|
|
|
```bash
|
|
cat > /tmp/ab/mix_arms.txt <<'EOF'
|
|
pat | | shipped default: Pattern only
|
|
tmh | TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both | TMHorizon alone
|
|
bb | TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay | BitBrain alone
|
|
mix | TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay | user config
|
|
EOF
|
|
tools/ab/ab_run.sh --arms /tmp/ab/mix_arms.txt --runs 7 --outdir /tmp/ab/mix --conc 5
|
|
python3 tools/ab/ab_analyze.py /tmp/ab/mix
|
|
python3 tools/ab/ab_dodge_analyze.py /tmp/ab/mix --json /tmp/ab/mix_dodge.json --reps 2000
|
|
```
|
|
|
|
Captured outputs are committed (the ~246 MB of raw `.jsonl` captures are
|
|
gitignored, as for `drussgt_dodge_vs_power.md`):
|
|
|
|
* `common_libs/tests/fixtures/gun_mix_ab_report.txt` — `ab_analyze.py` output
|
|
* `common_libs/tests/fixtures/gun_mix_dodge_report.txt` — per-arm dodge report
|
|
* `common_libs/tests/fixtures/gun_mix_dodge_results.json` — the same numbers as JSON
|