Commit Graph

13 Commits

Author SHA1 Message Date
SirStone 07766303f5 movement Batch 1: pure strafe (range tilt OFF) beats the shipped tfil on round wins across a 15-opponent panel
225 battles, one frozen binary, five env-only arms, the frozen panel, 0 invalid
runs. Paired per opponent vs the shipped tfil:

  strafe_notilt  wins/run +0.38  [CI +0.16,+0.60]  9/9 opponents p=0.0039
                 dmg/run  -10.2  [CI -25.8,+5.5]   p=0.61, MDE 20.4 (not detectable)
                 incoming hit rate 12.24% vs 18.17%, dmg taken 150 vs 200
  strafe_325     wins/run +0.33  [CI +0.04,+0.63]  10/12 p=0.0386
  ring           dmg/run  +31.2  [CI +11.5,+50.9]  13/15 p=0.0074, wins/run -0.04 (ns)
                 but hit rate 29.4% at 236 px: a damage/survival trade, not a win
  ring_notemp    indistinguishable from tfil on both primaries

Round wins in this harness are survival wins (in 216/219 attributable runs the
win count equals the rounds the opponent died in), and the winner takes ~1/3
fewer hits while fighting ~74 px farther out. The shipped tfil is last of five
on wins: the DrussGT-only picture did not generalize.

Also: tournament_analyze.py now prints BOTH readings of the pre-registered
'while the other does not go down' clause (strict: nothing is better;
substantive: the two strafe arms and ring are better on one metric each).
2026-09-26 01:13:29 +02:00
SirStone 8efa627c05 gauntlet: BitBrain vs Pattern across 32 legacy opponents (does not generalize)
The repo's first multi-opponent gun measurement. Adds tools/ab/gauntlet_run.sh
(per-opponent A/B over the legacy roster, subject = frozen ModularBot),
tools/ab/gauntlet_analyze.py (paired per-opponent deltas, cross-opponent sign
test, style split, MDE) and the arm/opponent fixtures.

Result: BitBrain does NOT generalize beyond DrussGT. 32 opponents x 2 arms x
3 runs x 5 rounds = 192 battles / 960 rounds, 0 failed, 0 retries: damage/run
214.5 (pattern) vs 210.9 (bb), sign-flip p=0.53; round wins 237/480 vs 239/480,
p=0.91. Sign test: bb better on 13/32 opponents (damage). The DrussGT-only
penalty does not carry. The owner's 'killer vs regular movers' sub-claim is not
supported: regular bucket +1.3 dmg/run vs dodgers -0.2 (MW p=0.85), and the
measured movement predictability does not correlate with the delta.
2026-09-26 00:55:57 +02:00
SirStone 1984a780f4 melee A/B doc: correct the per-arm [bb-reset] counts (perRound 68.75, retained 68.19, decay 66.75) 2026-09-26 00:43:04 +02:00
SirStone 4829f9ca13 BitBrain vs TMHorizon vs Pattern: live A/B on shipped TFIL (null result)
4 arms x 15 runs x 7 rounds (60 battles, 0 failed) vs real DrussGT on the
shipped TFIL default, frozen at ed25ce2. bb_id (gain 1.0 identity) is
statistically indistinguishable from shipped Pattern -> plumbing validity
check passes. No BitBrain arm beats TMHorizon or Pattern: bb_learn (the config
the owner likely ran) is the worst arm (276 dmg/run, 39/105 wins), the only
comparison at alpha=0.05 is Pattern beating it on damage. Learned gains
(>=1.0, gated >=300px) over-lead and lose 1.61pp of hit rate at 300-450px.
MDE 29.5 dmg/run, 1.235 wins/run; a 6-4-sized effect needs ~39 runs/arm.
2026-09-25 23:55:21 +02:00
SirStone 48f38b80e7 TFIL heat-time + virtual pillar: live A/B (6 arms x 70 rounds) - neither change beats the pre-change mover
Runs the pre-registered A/B for the two movement changes in HEAD: the
time-indexed bullet heat (TR_TFIL_HEAT_TIME, fca8993) and the removal of the
invented virtual centre pillar (d0750ab). One frozen binary from HEAD vs real
DrussGT: 6 arms x 10 runs x 7 rounds = 60 battles, 420 rounds, 0 failed.

Judged on damage/run and ROUND WINS only (hit rate and hits-taken are context):
hit rate would have inverted the verdict again - tau3 has the best pooled hit
rate of all arms (11.56%) and the fewest round wins (20/70).

RESULT (vs the reconstructed pre-change mover "old"):
  heat-time HURTS. tau3/tau5/tau9 lose 1.3-1.7 wins/run (p=0.0010-0.0125) and
  deal 22-38 less damage/run (p=0.004-0.047); tau15 is a wash on wins (p=0.64)
  and 22 damage/run lower (p=0.046). Nothing improves either metric.
  pillar removal does nothing measurable. old vs pillaoff: +5.7 damage/run
  (p=0.71), +0.5 wins/run (35 vs 30, p=0.43), 30.8 MORE damage taken/run
  without the pillar (p=0.040). The mechanism check proves the knob works
  (centre-box occupancy 0.09% -> 2.37%, p<0.0001; range 469 -> 443 px,
  p=0.0002), so this is a real behaviour change that buys nothing. At n=10 the
  pillar contrast is inside the MDE (33 damage/run, 1.2 wins/run), so this is
  not a proven regression.

Flags that the shipped default (pillar removed) should be reverted to the
TR_TFIL_PILLAR_ON behaviour; heat-time stays off.

Adds tools/ab/arms_heat_pillar.txt and tools/ab/ab_mechanism.py (per-tick
mechanism check: central-box occupancy, range distribution, live enemy-bullet
proximity) plus the captured summary/report fixtures.
2026-09-25 22:23:11 +02:00
SirStone 2747ebd323 BitBrain: TR_BITBRAIN_GAINS env knob (candidate set + fixed-gain degenerate)
Task A of campaign phase 2: the lead-gain candidate set is now pure env, so the
live arms need no recompile.

- common_libs/guns/bitbrain_gun.nim: BB_GAINS_ENV (TR_BITBRAIN_GAINS); the
  candidate list is parsed once at gun construction into a dynamic seq, so the
  hit counts/hit rates are sized to it. Unset/unparsable -> the shipped
  BB_CAND set [0,0.25,0.5,0.75,1.0] (byte-identical behaviour). Exactly ONE
  candidate degenerates to a FIXED gain applied from the first shot (learning
  bypassed), still gated to the long bands. parseGains clamps to [0,8],
  de-dupes and sorts so the argmax tie rule is unchanged. The [bb] line now
  prints the APPLIED gain AND the resulting angular shift, so a run's
  correction is auditable from stdout.
- ModularBot_garage/src/env_report.nim: emit TR_BITBRAIN_GAINS (resolved
  candidate set) and add BB_GAINS_ENV to the known-name list.
- tools/ab/arms_leadgain.txt: the 6-arm phase-2 sweep definition.
2026-09-25 00:15:27 +02:00
SirStone 140fe2519a HeadOn (no-lead) vs Pattern LIVE at long range: clean negative, offline ruler killed
2 arms x 15 runs x 7 rounds, one frozen binary from HEAD a82c864, real DrussGT,
server-side events sidecar. Shipped rack is onlyPattern, so control=Pattern-only
and headon=HeadOn-only (TR_RACK_PATTERN=off TR_RACK_HEADON=both).

  arm      dmg/run  dmgtk/run  round wins  shots/run
  control      279        211     48/105       785
  headon        14        228      0/105       580

Round wins and dmg/run both separate at p<0.0001 (MC permutation, se 0.0000),
~7x the damage MDE (35.8). Per range band (pooled, 15 runs):
  300-450: Pattern 12.3% (4590 shots) vs HeadOn 0.6% (3701)  p<0.0001, MDE 2.0pp
  450+   : Pattern  9.2% (6671)       vs HeadOn 0.4% (4177)  p<0.0001, MDE 1.1pp
HeadOn loses EVERY long-range band by 20-23x, so the whole-battle loss is not a
close-range artefact.

The offline ruler (prediction_quality_results.txt) predicted the opposite: HeadOn
meanAbs 14.61 vs Pattern 17.53 at 300-450 and 12.33 vs 16.19 at 450+, hitProxy
.105/.104 and .098/.077 (+27%). That is an open-loop replay of a FIXED enemy
track, so it cannot see that a different bullet makes the surfer dodge
differently; live, the static gun does not lead at all.

TR_PATTERN_RAD_SCALE arms were skipped: applyRadial scales aim DISTANCE along an
unchanged bearing, so it cannot express 'less lead' (bearing is what firing uses).
HeadOn confirmed to ignore bulletSpeed (head_on.nim:9), liveness OK 15/15.

Adds the range-band analyzer tools/ab/ab_range_bands.py (reuses the lead-capture
Run alignment) and the captured fixtures. Does not touch bitbrain_gun.nim /
bitbrain_campaign.md (job-100).
2026-09-24 23:52:49 +02:00
SirStone 32a5e72fac Gun mixing (TMHorizon+BitBrain) vs DrussGT: clean negative, no dodge disruption
4 arms x 7 runs vs real DrussGT. mix alternates the two guns 476 times/7 runs
(liveness OK) but our bullets are no more varied (power sd / aim-offset sd flat)
and DrussGT's dodge quality is unchanged (miss/tick mix-pat +0.03, p=0.66; MDE
3.8%). mix wins 24/49 = the 49% baseline; the user's 6/10 has P=0.353 at 49%.
New tools/ab/ab_dodge_analyze.py splits the validated per-shot dodge instrument
by arm and adds gun-switch/power/bearing liveness; fixtures committed.
2026-09-24 23:07:05 +02:00
SirStone d93ce444c0 BitBrain verdict: clean negative at 30 runs/arm; analyzer gets MC + Mann-Whitney + MDE
- docs/bitbrain_gun_verdict.md: control vs bb_decay (decay SBC memory) vs a
  provably-zero placebo, 30 runs/arm vs real DrussGT. Nothing separates
  (bb_decay +3.3 dmg/run, p=0.71; round wins 97/210 vs 97/210, p=1.00); the
  7-run shape does not replicate. TR_BITBRAIN_RANGE=0 is clamped to 1.0 deg
  (bitbrain_gun.nim:207) so it is NOT a zero-shift placebo; TR_BITBRAIN_MIN_OBS
  unreachable is used instead.
- tools/ab/ab_analyze.py: keep exact enumeration for C(n,na)<=20e6 (7v7), add
  a seeded Monte-Carlo permutation test (1e6 draws, 0x5eed5eed) with its
  standard error, a tie-corrected Mann-Whitney U cross-check, a minimum
  detectable effect line, all-pairs comparisons, and a [bb] shift check.
- tools/ab/README.md: document the new analyzer output.
2026-09-24 23:02:59 +02:00
SirStone 17c50159ba tools/ab: reusable A/B runner + analyzer (frozen-HEAD build, exact permutation test, liveness check) 2026-09-24 22:08:59 +02:00
SirStone b0654d18eb TM verdict, settled: it loses LIVE and sits at/below its majority class - (c)
The user pushed back on "the TM can't be your best 1v1 gun", correctly, because two
decisive tests had never been run. Both are now run and they agree.

TASK 1 - THE GF HEAD vs ITS MAJORITY-CLASS BASELINE (offline, n=1,751,067):
  label histogram [254286, 284578, 678879, 297055, 236269]
  majority class = 2 (the CENTRE bucket) = 38.77%
  RAW head accuracy = 36.69%  ->  margin **-2.08 pp, BELOW majority**
  GATED head accuracy = 40.37% vs 38.75% majority -> +1.62 pp, BUT it predicts the
  majority class on 62.4% of ticks and its minority recall is 13.6% / 12.9% - a
  base-rate predictor wearing a classifier's clothes.
  Shuffled control sits at its own majority (20.04% vs 20.12%), confirming chance.
**THE OLD "46% vs 20% CHANCE" FIGURE I QUOTED WAS WRONG ON TWO COUNTS:** the
baseline is 38.8%, not 20%, and the 46% predated the deferred-label fix. Against
the correct baseline the head is BELOW it.

TASK 2 - THE FIRST-EVER LIVE A/B OF THE TM GUN (7 runs x 7 rounds per arm, one
frozen binary from git archive HEAD = eb74f9b2, sha256 cb66d66b..., real DrussGT,
every arm forced alone with TR_RACK_<GUN>=both and all 14 others off, liveness
confirmed per run):
  arm                     shots   real %   dmg/run   round wins
  onlyPattern              4610   10.74%     285      25/49
  onlyTMPATTERN (radial)   3374    3.50%      71       0/49
  onlyLinear               3218    3.23%      61       0/49
  Pattern vs TM:  +7.22 pp / +213.7 dmg, exact p=0.0006
  TM vs Linear:   +0.30 pp, p=0.659  (dmg p=0.438)
**The TM is statistically INDISTINGUISHABLE from its own Linear base live.** So it
is not "the TM works and we are aiming it wrong".

DIRECT ANSWER: **(c) It loses live AND sits at/below majority - the target carries
no learnable signal beyond the base rate, and that is the reason.** The reason is
not the machine, not the knobs, and not the application alone: the thing it was
asked to predict is dominated by the modal answer.

This closes the TM-as-gun thread. If a TM is wanted in the bot, a firing gate or a
movement decision is a better fit for a boolean-rule classifier than an aim point -
that is untested and is a different project.

A LIVE GF-MODE ARM WAS NOT RUN (stated as unmeasured): the task pinned one frozen
HEAD binary and HEAD registers the TM gun as radial only; Task 1 already makes GF
the unpromising candidate.

HARNESS FIX WORTH KEEPING: `tools/ab/which_gun_arm_env.sh` left the TARGET gun
unset, so with the now-Pattern-only default it silently fell back to the FULL rack
- an arm could appear to test a single gun while actually running the whole rack.
It now emits `TR_RACK_<GUN>=both` for the target and `=off` for all 14 others.
(Earlier which-gun results are unaffected: they ran before the Pattern-only default,
or - as in the melee/1v1 campaign - set the explicit `=both` themselves.)

tm_pattern.nim gains a per-class confusion matrix (warm samples only) to support the
majority baseline; no behaviour change. Adds Round 4 to
tm_pattern_sweep_results.md with both tasks and the interpretation rule.
2026-09-22 08:21:49 +02:00
SirStone a54ae6a162 SETTLED: no small rack beats Pattern alone; the selector is negative value on a GOOD rack
Follow-up to e0666a5, which showed Pattern alone (10.78%) beats the full rack
(6.93%). That left two open questions: is a SMALL rack of good guns better than
Pattern alone, and does the selector add value on a good rack (rather than only
on the bloated one)? Both are now answered: NO and NO.

6 arms x 7 runs x 7 rounds, one frozen binary from CLEAN HEAD e0666a5 (built via
`git archive`, source verified byte-identical to the clean tree), rack knobs
only, 8 concurrent battles, real server-side hit rate vs the real DrussGT, exact
two-sided permutation test on per-run rates.

  arm          guns (selector active?)                    real %  dmg/run  p vs onlyPattern
  onlyPattern  Pattern, NO selection                       10.36    264     --
  lean8        HeadOn,Linear,Circular,Accel,Pattern,GF,KNN,WallBounce  6.31  146  0.0169
  lean6        lean8 - HeadOn                               8.83    212     0.0262
  pairPC       Pattern + Circular                           8.23    185     0.0460
  pairPK       Pattern + KNN                                9.80    264     0.3998
  pairPL       Pattern + Linear                             8.23    200     0.0035

The control replicates the prior run (10.36% vs 10.78% before; same binary tree,
different build path).

THE MECHANISM, from the per-arm selected-gun mix - the virtual signal keeps
ranking the WRONG guns first, even on a two-gun rack:
  lean8: HeadOn 46.2% of ticks at 2.0% REAL; Pattern only 14.4% (12.6% real)
  lean6: Pattern 29.2% (10.4% real) vs KNN 25.3% (8.1%) and Linear 15.2% (8.0%)
  pairPC: Circular 66.8% (6.9% real) vs Pattern 33.2% (11.2% real) - over-picks Circular
  pairPL: Linear 57.7% (6.0% real) vs Pattern 42.3% (11.1% real) - over-picks Linear
  pairPK: Pattern 86.4% - ties ONLY because the selector happens to pick Pattern
          most of the time; it is numerically lower with identical dmg/run
So the failure is NOT rack size. Pruning does not fix it; the ranking is wrong.

VERDICT: ship `onlyPattern` - Pattern alone with selection bypassed - at 10.36%
real and 264 dmg/run, vs lean8 6.31%/146 and the prior full rack 6.93%/159.
This DIRECTLY CONTRADICTS the standing user directive to keep virtual-fitness
selection, so it is recorded here plainly rather than quietly acted on: disable
the selector (`TR_RACK_<every gun but PATTERN>=off`) pending a better fitness
signal. The mechanism itself is left intact and functional so it can be re-enabled
with one env var, and so it can be fixed rather than discarded.

REMAINING CAVEAT: ONE ADVERSARY. All of this is vs DrussGT. Pattern as the default
must be re-checked against other bots first - that is the next job.

Extends the reusable harness (tools/ab/which_gun_arm_env.sh now has lean8/lean6/
pairPC/pairPK/pairPL; which_gun_analyze.py is parameterised by WHICHGUN_OUT and
compares against both `full` and `onlyPattern`).
2026-09-22 01:35:51 +02:00
SirStone e0666a562d The gun selector is NEGATIVE value: Pattern alone beats the full rack (p=0.0012)
MEASURED against the real DrussGT, one frozen binary built from clean HEAD, rack
knobs only (no source edits), 5 arms x 7 runs x 7 rounds, 8 concurrent battles,
judged ONLY on server-side real hit rate from the events sidecar, exact
two-sided permutation test on per-run rates.

  arm             runs  shots  hits  real %  dmg/run  p vs full
  full (shipped)     7   3898   270    6.93     159      --
  onlyPattern        7   4582   494   10.78     287      0.0012  <- BETTER
  onlyKNN            7   4033   207    5.13     119      0.1340
  onlyLinear         7   3215   105    3.27      65      0.0082
  onlyGF             7   3193    72    2.25      45      0.0012

Firing Pattern ALONE gives +3.85pp pooled hit rate and +80% damage per run, and
it fires MORE shots (4582 vs 3898) - it dominates on rate and volume. This is not
"any single gun wins" (full beats Linear, GF and KNN); it is specifically
"Pattern alone beats the rack".

WHY - the virtual fitness signal mis-ranks guns against real outcomes:
- HeadOn is massively over-selected: 31.4% of ticks, the most real shots (1070),
  but only 4.5% REAL. It alone drags the rack down.
- Pattern has the best virtual rank and near-best real rate (11.9%, rank 2), yet
  is selected only 22.6% of the time.
- Linear's apparent strength was SELECTION BIAS: conditional on being selected it
  looked like 15.2% (n=33), but its UNCONDITIONAL rate (onlyLinear) is 3.27%.
  Every earlier per-gun "real rate" in this repo is conditional on selection and
  is therefore confounded. This experiment is the clean measurement.

NOT YET SETTLED (do not overclaim):
- ONE ADVERSARY. All of this is vs DrussGT. Pattern must be re-checked against
  other bots before it becomes the default on this evidence alone.
- Whether a SMALL rack of good guns beats Pattern alone. The selector is negative
  value on the CURRENT bloated rack; that does not prove it is negative value on
  a rack of only good guns. That is the next experiment and it decides whether
  the selection apparatus is fixed or disabled.
- The user's standing directive is to KEEP virtual-fitness selection. This
  measurement conflicts with it, so the next step tests the selector on a small
  good rack rather than assuming either answer.

Context - three prior selection-side attempts all failed: hysteresis (7.02% ->
5.10%, p=0.002), commitment (7.17% -> 4.44%, p=0.0012), arrival-accuracy
tie-break (7.08%, p=0.88 null). The per-tick random draw is load-bearing on
three independent measurements. This experiment locates the real problem one
level up: which guns are in the rack, and that the virtual signal ranks them
wrongly.

Preserves the reusable harness (tools/ab/which_gun_run_one.sh,
which_gun_arm_env.sh, which_gun_analyze.py) and the full writeup
(docs/selector_negative_value.md).
2026-09-22 01:21:14 +02:00