2 arms x 15 runs x 7 rounds, one frozen binary from HEAD a82c864, real DrussGT,
server-side events sidecar. Shipped rack is onlyPattern, so control=Pattern-only
and headon=HeadOn-only (TR_RACK_PATTERN=off TR_RACK_HEADON=both).
arm dmg/run dmgtk/run round wins shots/run
control 279 211 48/105 785
headon 14 228 0/105 580
Round wins and dmg/run both separate at p<0.0001 (MC permutation, se 0.0000),
~7x the damage MDE (35.8). Per range band (pooled, 15 runs):
300-450: Pattern 12.3% (4590 shots) vs HeadOn 0.6% (3701) p<0.0001, MDE 2.0pp
450+ : Pattern 9.2% (6671) vs HeadOn 0.4% (4177) p<0.0001, MDE 1.1pp
HeadOn loses EVERY long-range band by 20-23x, so the whole-battle loss is not a
close-range artefact.
The offline ruler (prediction_quality_results.txt) predicted the opposite: HeadOn
meanAbs 14.61 vs Pattern 17.53 at 300-450 and 12.33 vs 16.19 at 450+, hitProxy
.105/.104 and .098/.077 (+27%). That is an open-loop replay of a FIXED enemy
track, so it cannot see that a different bullet makes the surfer dodge
differently; live, the static gun does not lead at all.
TR_PATTERN_RAD_SCALE arms were skipped: applyRadial scales aim DISTANCE along an
unchanged bearing, so it cannot express 'less lead' (bearing is what firing uses).
HeadOn confirmed to ignore bulletSpeed (head_on.nim:9), liveness OK 15/15.
Adds the range-band analyzer tools/ab/ab_range_bands.py (reuses the lead-capture
Run alignment) and the captured fixtures. Does not touch bitbrain_gun.nim /
bitbrain_campaign.md (job-100).
6.4 KiB
HeadOn vs Pattern at long range — LIVE gate
Date: 2026-09-24
Branch: research/lead-targeting · commit: a82c864 · binary sha256: 9d20f530d6… (one frozen binary, git archive HEAD)
Session: /tmp/ab_headon · 2 arms × 15 runs × 7 rounds = 210 rounds, conc=8
Fixtures committed: headon_longrange_ab_report.txt (ab_analyze), headon_longrange_bands.txt
(range-band), headon_longrange_live_results.json (per-run + per-band counts). Arms: tools/ab/arms_headon.txt.
Direct answer
NO. The no-lead static gun does not beat Pattern live, overall or in any long-range band. It is not close: HeadOn won 0 of 105 rounds and dealt 14 dmg/run vs 279 (20× less), and at 450+ px it hits 0.4% vs Pattern's 9.2% (23× less). The offline ruler's headline — HeadOn has lower aim error and a better hit proxy than Pattern at 300+ px — is killed. The offline ruler was the only thing that predicted the opposite; this live A/B is the only thing that can confirm or kill it, and it killed it.
The TR_PATTERN_RAD_SCALE arms were skipped (the knob cannot express the lead hypothesis)
common_libs/guns/pattern_matcher.nim applyRadial moves the aim point along the unchanged
bearing (selfX + dx/d*nd, selfY + dy/d*nd) — the direction dx/d, dy/d is untouched, only the
distance changes. The fired bearing is aimAngle(self, pred) (ModularBot.nim:1208), so scaling the
aim distance cannot move the bearing. The knob is a radial range correction, not a lead
multiplier, and cannot express "less lead". Per the task rule, arms rad075/rad050 were skipped
and the experiment ran as 2 arms × 15 runs.
HeadOn really does ignore bulletSpeed
common_libs/guns/head_on.nim:9 returns GunPrediction(x: state.enemyX, y: state.enemyY), never
reading its bulletSpeed argument. The rack calls it for the selected power bin and the fire block
takes the bearing straight to that point. So the premise holds: HeadOn applies zero lead,
independent of power. The headon arm's liveness check passed (its env reached the boot report in
15/15 runs), so this was HeadOn and only HeadOn.
Primary result — damage/run and round wins [MEASURED]
| arm | runs | dmg/run | dmgtk/run | round wins | win% | shots/run | hits taken/run |
|---|---|---|---|---|---|---|---|
control (shipped Pattern-only) |
15 | 279 | 211 | 48/105 | 45.7 | 785 | 93 |
headon (HeadOn-only) |
15 | 14 | 228 | 0/105 | 0.0 | 580 | 72 |
Per-run damage (control): 263 307 294 221 290 235 327 262 330 312 274 292 288 281 217.
Per-run damage (headon): 4 8 4 17 14 10 28 18 0 36 17 8 14 14 15.
| metric | diff (control−headon) | perm p | method | MDE (abs) | MDE vs control mean |
|---|---|---|---|---|---|
| dmg/run | +265.7 | <0.0001 | MC/B=1,000,000 (se 0.0000) | 35.8 | 12.8% of 279.5 |
| round wins | +3.20 | <0.0001 | MC/B=1,000,000 (se 0.0000) | 1.29 | 40.4% of 3.2 |
Mann-Whitney U cross-check p = 0.0000 on both. Round-level Fisher p < 0.0001 (anti-conservative). The effect is ~7× the damage MDE; this is not an under-powered null.
Hit rate by range band — the load-bearing breakdown [MEASURED]
Pooled over each arm's 15 runs; band = shooter→target distance at the fire tick (same bands as the offline ruler). HeadOn loses the long-range bands by a landslide — the whole-battle defeat is not a close-range artefact.
| band (px) | control shots / hits / rate | headon shots / hits / rate |
|---|---|---|
| 0–100 | 3 / 2 / 66.7% | 1 / 0 / 0.0% |
| 100–200 | 28 / 4 / 14.3% | 97 / 8 / 8.2% |
| 200–300 | 211 / 26 / 12.3% | 469 / 12 / 2.6% |
| 300–450 | 4590 / 565 / 12.3% | 3701 / 21 / 0.6% |
| 450+ | 6671 / 616 / 9.2% | 4177 / 15 / 0.4% |
| ALL | 11503 / 1213 / 10.5% | 8445 / 56 / 0.7% |
Per-band permutation test on per-run band rates (arm − control):
| band | diff (pp) | p | method | MDE (pp) |
|---|---|---|---|---|
| 100–200 | +1.93 | 0.6609 | exact | 12.4 |
| 200–300 | −12.12 | <0.0001 | MC/B=200,000 | 12.1 |
| 300–450 | −11.84 | <0.0001 | MC/B=200,000 | 2.0 |
| 450+ | −8.89 | <0.0001 | MC/B=200,000 | 1.1 |
| ALL | −9.90 | <0.0001 | MC/B=200,000 | 1.0 |
The two long-range bands separate at 4–8× their MDE, overwhelmingly in HeadOn's favour being false. Note the 98%+ of shots that land at 300+ px, exactly as the offline ruler reported.
Offline ruler vs live [MEASURED]
| source | metric | Pattern 300–450 | HeadOn 300–450 | Pattern 450+ | HeadOn 450+ |
|---|---|---|---|---|---|
offline prediction_quality_results.txt |
meanAbs err (deg) | 17.53 | 14.61 | 16.19 | 12.33 |
| offline | hitProxy | 0.104 | 0.105 | 0.077 | 0.098 (+27%) |
| live (this run) | hit rate | 12.3% | 0.6% | 9.2% | 0.4% |
The ruler said HeadOn is equal-or-better at 300+; live it is 20×/23× worse. The ruler is a per-tick, open-loop replay of a fixed enemy track: a no-lead gun's aim error is scored against the enemy's true future position on that same pre-recorded dodge, which is not the distribution a different bullet induces. Live, the bullet changes the dodge, and the static gun simply does not lead the surfer. This is precisely the open-loop/live gap the ruler cannot see. The offline aim-error ruler is not a valid predictor of the live HeadOn-vs-Pattern question — do not use it to rank no-lead guns at range.
Caveats and what is inferred
- [MEASURED] All numbers above, from this session's server-side events sidecar; one frozen
binary; liveness OK for both arms;
wins == firstPlaces15/15 both arms. - [INFERRED] The rack-membership change also removes a gun's 4 virtual bullets/tick from the
shared order-sensitive ring (
docs/gun_rack_analysis.md§1.1). Both arms are single-gun racks and the effect size (0/105 rounds, ~7× MDE) is far too large for ring-order to explain. - [MEASURED] One adversary (the real DrussGT jar). The 70-run offline corpus this ruler was built on is the same live capture set, so the two disagree on the same adversary — the disagreement is the open-loop/live gap, not a matchup difference.
- Reproduce:
tools/ab/ab_run.sh --arms tools/ab/arms_headon.txt --runs 15 --outdir /tmp/ab_headon --conc 8 --rounds 7;python3 tools/ab/ab_analyze.py /tmp/ab_headon --reference control;python3 tools/ab/ab_range_bands.py /tmp/ab_headon --reference control.