Gun mixing (TMHorizon+BitBrain) vs DrussGT: clean negative, no dodge disruption
4 arms x 7 runs vs real DrussGT. mix alternates the two guns 476 times/7 runs (liveness OK) but our bullets are no more varied (power sd / aim-offset sd flat) and DrussGT's dodge quality is unchanged (miss/tick mix-pat +0.03, p=0.66; MDE 3.8%). mix wins 24/49 = the 49% baseline; the user's 6/10 has P=0.353 at 49%. New tools/ab/ab_dodge_analyze.py splits the validated per-shot dodge instrument by arm and adds gun-switch/power/bearing liveness; fixtures committed.
This commit is contained in:
@@ -0,0 +1,69 @@
|
||||
# session /tmp/ab/mix
|
||||
# commit=f91e12196537c6c1e64353160e8885445e3e222b binary_sha256=54abc793b819832bf1becb366b3928f7db5c853668d18a525d63e677b7f89a73 rounds=7 runs=7 conc=5 ts=2026-09-24T22:54:59+02:00
|
||||
|
||||
ARM SUMMARY
|
||||
arm runs dmg/run dmgtk/run wins win% shots/run hitstk/run
|
||||
--------------------------------------------------------------------------
|
||||
pat 7 291 213 21/49 42.9 815 95.7
|
||||
tmh 7 265 211 18/49 36.7 788 94.9
|
||||
bb 7 278 183 22/49 44.9 802 92.3
|
||||
mix 7 267 210 24/49 49.0 804 93.6
|
||||
|
||||
PER-RUN (never just the mean)
|
||||
pat dmg: r1=319 r2=313 r3=270 r4=318 r5=281 r6=290 r7=250
|
||||
wins: r1=5/7 r2=3/7 r3=3/7 r4=4/7 r5=3/7 r6=1/7 r7=2/7
|
||||
tmh dmg: r1=237 r2=309 r3=249 r4=230 r5=288 r6=268 r7=274
|
||||
wins: r1=1/7 r2=4/7 r3=2/7 r4=2/7 r5=3/7 r6=2/7 r7=4/7
|
||||
bb dmg: r1=230 r2=271 r3=270 r4=282 r5=289 r6=307 r7=297
|
||||
wins: r1=2/7 r2=3/7 r3=4/7 r4=3/7 r5=3/7 r6=4/7 r7=3/7
|
||||
mix dmg: r1=260 r2=255 r3=220 r4=293 r5=284 r6=266 r7=288
|
||||
wins: r1=2/7 r2=5/7 r3=2/7 r4=3/7 r5=3/7 r6=4/7 r7=5/7
|
||||
|
||||
PAIRWISE PERMUTATION TEST (per-run values) + MANN-WHITNEY CROSS-CHECK
|
||||
permutation: exact when C(n,na) <= 20,000,000; otherwise Monte-Carlo 1,000,000 draws, seed=0x5eed5eed, p = (cnt+1)/(B+1), se = sqrt(p(1-p)/(B+1))
|
||||
metric A B diff(A-B) perm p method MC se MW p MW U
|
||||
-------------------------------------------------------------------------------------------------
|
||||
dmg/run pat tmh +26.284 0.0956 exact - 0.0736 10.0
|
||||
round wins pat tmh +0.429 0.6638 exact - 0.5542 19.5
|
||||
dmg/run pat bb +13.488 0.3450 exact - 0.4433 18.0
|
||||
round wins pat bb -0.143 1.0000 exact - 0.8368 22.5
|
||||
dmg/run pat mix +24.919 0.0973 exact - 0.1599 13.0
|
||||
round wins pat mix -0.429 0.6795 exact - 0.6439 20.5
|
||||
dmg/run tmh bb -12.796 0.3811 exact - 0.4433 18.0
|
||||
round wins tmh bb -0.571 0.4079 exact - 0.3157 16.5
|
||||
dmg/run tmh mix -1.365 0.9237 exact - 1.0000 24.0
|
||||
round wins tmh mix -0.857 0.2879 exact - 0.2346 15.0
|
||||
dmg/run bb mix +11.431 0.4103 exact - 0.3067 16.0
|
||||
round wins bb mix -0.286 0.7960 exact - 0.7880 22.0
|
||||
|
||||
MINIMUM DETECTABLE EFFECT (two-sample, alpha=0.05 two-sided, 80% power; MDE = 2.8016*sd*sqrt(2/n))
|
||||
metric n/arm sd(control) MDE(abs) MDE vs control mean
|
||||
----------------------------------------------------------------
|
||||
dmg/run 7 26.481 39.656 13.6% of 291.4
|
||||
round wins 7 1.291 1.933 64.4% of 3.0
|
||||
|
||||
ROUND-LEVEL TEST (pooled rounds, Fisher exact) vs `pat` — ANTI-CONSERVATIVE: rounds cluster within runs
|
||||
arm ref wins arm wins p
|
||||
----------------------------------------------
|
||||
tmh 21/49 18/49 0.6801
|
||||
bb 21/49 22/49 1.0000
|
||||
mix 21/49 24/49 0.6854
|
||||
|
||||
LIVENESS (arm env applied in the bot's own boot report)
|
||||
pat OK (7/7 runs: no arm env; report present)
|
||||
tmh OK (7/7 runs: TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both applied)
|
||||
bb OK (7/7 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay applied)
|
||||
mix OK (7/7 runs: TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay applied)
|
||||
|
||||
[bb] APPLIED-SHIFT CHECK (from bot stdout; needs TR_BITBRAIN_LOG=1). A provably-zero placebo emits ZERO [bb] lines.
|
||||
arm runs w/log lines min max zeros
|
||||
pat 0/7 0 - - -
|
||||
tmh 0/7 0 - - -
|
||||
bb 0/7 0 - - -
|
||||
mix 0/7 0 - - -
|
||||
|
||||
ROUND-WIN ATTRIBUTION (events primary; score tie-break for mutual-kill / timeout rounds)
|
||||
pat wins==firstPlaces 7/7 runs OK; single-death rounds agree with score 48/48 (1 tie-broken)
|
||||
tmh wins==firstPlaces 7/7 runs OK; single-death rounds agree with score 47/47 (2 tie-broken)
|
||||
bb wins==firstPlaces 7/7 runs OK; single-death rounds agree with score 49/49 (0 tie-broken)
|
||||
mix wins==firstPlaces 7/7 runs OK; single-death rounds agree with score 49/49 (0 tie-broken)
|
||||
@@ -0,0 +1,85 @@
|
||||
====================================================================================================
|
||||
PER-ARM DRUSSTGT DODGE QUALITY (session commit f91e12196, 7 runs x 7 rounds)
|
||||
attribution cross-check: 199/199 deaths have the mapped victim at ~0 energy
|
||||
shipped baseline vs DrussGT is ~49%% round wins; this instrument is per SHOT (~thousands/arm), not per round.
|
||||
|
||||
--- ARM `pat` --- 7 runs, 49 rounds, 5586 shots
|
||||
OUR shooting: power mean=0.810 sd=0.262 iqr=0.500 distinct=601 | aim-offset sd=14.16 deg iqr=22.55 | fire range mean=470 px
|
||||
power histogram (top): 1.00:3163 0.50:1470 0.10:116 0.65:28 0.68:27 0.55:27 0.60:25 0.71:24 0.97:23 0.51:22
|
||||
aim-offset histogram (deg bucket): -30:3 -25:319 -20:440 -15:558 -10:620 -5:657 +0:671 +5:605 +10:467 +15:462 +20:470 +25:312 +30:1 +45:1
|
||||
GUN SELECTION ([config] switch lines): Pattern:54 | switches=0 over 7/7 runs
|
||||
|
||||
--- ARM `tmh` --- 7 runs, 49 rounds, 5417 shots
|
||||
OUR shooting: power mean=0.828 sd=0.253 iqr=0.500 distinct=594 | aim-offset sd=14.32 deg iqr=23.13 | fire range mean=480 px
|
||||
power histogram (top): 1.00:3161 0.50:1434 0.82:25 0.78:22 0.54:22 0.57:22 0.63:21 0.81:20 0.52:20 0.74:20
|
||||
aim-offset histogram (deg bucket): -35:2 -30:49 -25:291 -20:485 -15:517 -10:559 -5:625 +0:640 +5:561 +10:519 +15:442 +20:483 +25:215 +30:29
|
||||
GUN SELECTION ([config] switch lines): TMHorizon:54 | switches=0 over 7/7 runs
|
||||
|
||||
--- ARM `bb` --- 7 runs, 49 rounds, 5483 shots
|
||||
OUR shooting: power mean=0.829 sd=0.248 iqr=0.500 distinct=647 | aim-offset sd=14.33 deg iqr=22.82 | fire range mean=473 px
|
||||
power histogram (top): 1.00:3212 0.50:1416 0.90:25 0.62:24 0.71:23 0.86:22 0.56:22 0.52:22 2.00:22 0.60:22
|
||||
aim-offset histogram (deg bucket): -60:1 -45:1 -35:3 -30:6 -25:345 -20:443 -15:529 -10:607 -5:610 +0:662 +5:577 +10:486 +15:442 +20:482 +25:282 +30:5 +40:1 +50:1
|
||||
GUN SELECTION ([config] switch lines): BitBrain:55 | switches=0 over 7/7 runs
|
||||
|
||||
--- ARM `mix` --- 7 runs, 49 rounds, 5518 shots
|
||||
OUR shooting: power mean=0.811 sd=0.264 iqr=0.500 distinct=651 | aim-offset sd=14.30 deg iqr=23.20 | fire range mean=469 px
|
||||
power histogram (top): 1.00:3074 0.50:1486 0.10:61 0.59:26 0.73:24 0.52:24 0.70:23 0.77:23 0.72:22 0.57:22
|
||||
aim-offset histogram (deg bucket): -30:38 -25:255 -20:478 -15:524 -10:594 -5:609 +0:642 +5:581 +10:521 +15:478 +20:521 +25:248 +30:29
|
||||
GUN SELECTION ([config] switch lines): TMHorizon:275 BitBrain:258 | switches=476 over 7/7 runs
|
||||
|
||||
====================================================================================================
|
||||
DRUSSGT DODGE QUALITY PER ARM (mean, 95%% CI = run-cluster bootstrap)
|
||||
arm shots | miss miss/tick [95%CI] | lat12 lat12 [95%CI] | lat_ctrl lat_ctrl [95%CI] | hit%
|
||||
------------------------------------------------------------------------------------------------------------------
|
||||
pat 5586 | 117.0 4.64[ 4.56, 4.72] | 53.9 [ 52.5, 55.1] | 51.4 [ 50.4, 52.3] | 0.107
|
||||
tmh 5417 | 121.6 4.74[ 4.68, 4.80] | 54.6 [ 53.4, 55.5] | 52.7 [ 52.1, 53.3] | 0.099
|
||||
bb 5483 | 118.6 4.66[ 4.59, 4.72] | 54.0 [ 53.4, 54.7] | 52.1 [ 51.3, 52.8] | 0.099
|
||||
mix 5518 | 117.9 4.67[ 4.58, 4.76] | 54.4 [ 53.8, 55.0] | 52.6 [ 52.1, 53.1] | 0.101
|
||||
|
||||
====================================================================================================
|
||||
MINIMUM DETECTABLE EFFECT (run-cluster level, n=7/arm, alpha=0.05 two-sided, 80% power; MDE = 2.8016*sd_perrun*sqrt(2/n))
|
||||
the shot counts are large but the RUNS are what set the between-arm uncertainty, so this is the honest floor
|
||||
metric sd(per-run) MDE(abs) MDE vs ref mean
|
||||
miss 2.437 3.650 3.1% of 117.029
|
||||
miss/tick 0.118 0.176 3.8% of 4.642
|
||||
lat12 1.824 2.732 5.1% of 53.898
|
||||
lat_ctrl 1.372 2.055 4.0% of 51.361
|
||||
hit% 0.006 0.008 0.1% of 10.705
|
||||
|
||||
====================================================================================================
|
||||
BETWEEN-ARM CONTRASTS (exact run-cluster permutation, C(14,7)=3432)
|
||||
pooled = pooled-shot difference A-B; strat = range-stratified (bands 200-300,300-450,450+, n-weighted)
|
||||
a mix that poisons DrussGT should show miss/tick, lat12 and lat_ctrl LOWER (worse dodging) than the best single gun
|
||||
metric A vs B pooled perm p | strat perm p
|
||||
-----------------------------------------------------------------------------------------
|
||||
miss mix - tmh -3.691 0.0789 | -2.108 0.1809 [exact]
|
||||
miss_per_tick mix - tmh -0.065 0.2805 | -0.080 0.1815 [exact]
|
||||
lat_fixed_abs mix - tmh -0.190 0.7862 | -0.054 0.9289 [exact]
|
||||
lat_ctrl_abs mix - tmh -0.101 0.7984 | -0.070 0.8584 [exact]
|
||||
hit mix - tmh +0.001 0.7040 | +0.000 0.9441 [exact]
|
||||
|
||||
miss mix - bb -0.683 0.7408 | +0.048 0.9790 [exact]
|
||||
miss_per_tick mix - bb +0.014 0.8048 | +0.003 0.9522 [exact]
|
||||
lat_fixed_abs mix - bb +0.368 0.4838 | +0.443 0.3865 [exact]
|
||||
lat_ctrl_abs mix - bb +0.537 0.2799 | +0.546 0.2869 [exact]
|
||||
hit mix - bb +0.001 0.7291 | +0.001 0.7932 [exact]
|
||||
|
||||
miss tmh - pat +4.584 0.0114 | +3.625 0.0189 [exact]
|
||||
miss_per_tick tmh - pat +0.095 0.1081 | +0.105 0.0673 [exact]
|
||||
lat_fixed_abs tmh - pat +0.686 0.4699 | +0.631 0.4926 [exact]
|
||||
lat_ctrl_abs tmh - pat +1.340 0.0498 | +1.356 0.0457 [exact]
|
||||
hit tmh - pat -0.008 0.0160 | -0.006 0.0685 [exact]
|
||||
|
||||
miss bb - pat +1.575 0.3632 | +1.375 0.3283 [exact]
|
||||
miss_per_tick bb - pat +0.016 0.7850 | +0.019 0.7373 [exact]
|
||||
lat_fixed_abs bb - pat +0.127 0.8858 | +0.121 0.8905 [exact]
|
||||
lat_ctrl_abs bb - pat +0.702 0.3067 | +0.738 0.2893 [exact]
|
||||
hit bb - pat -0.008 0.0277 | -0.008 0.0283 [exact]
|
||||
|
||||
miss mix - pat +0.893 0.6242 | +1.340 0.3318 [exact]
|
||||
miss_per_tick mix - pat +0.030 0.6644 | +0.020 0.7489 [exact]
|
||||
lat_fixed_abs mix - pat +0.495 0.5817 | +0.562 0.5322 [exact]
|
||||
lat_ctrl_abs mix - pat +1.239 0.0562 | +1.274 0.0481 [exact]
|
||||
hit mix - pat -0.006 0.1296 | -0.007 0.1151 [exact]
|
||||
|
||||
JSON written to /tmp/ab/mix_dodge.json
|
||||
@@ -0,0 +1,486 @@
|
||||
{
|
||||
"arms": {
|
||||
"bb": {
|
||||
"ci": {
|
||||
"lat_ctrl_abs": [
|
||||
51.342068468176116,
|
||||
52.805805223688346
|
||||
],
|
||||
"lat_fixed_abs": [
|
||||
53.391673798672606,
|
||||
54.727408738794885
|
||||
],
|
||||
"miss_per_tick": [
|
||||
4.593774088337814,
|
||||
4.716995057710113
|
||||
]
|
||||
},
|
||||
"dodge": {
|
||||
"hit": 0.09921575779682655,
|
||||
"lat_ctrl_abs": 52.06310968098528,
|
||||
"lat_disp_per_tick": 3.595491225896298,
|
||||
"lat_fixed_abs": 54.02567072842506,
|
||||
"miss": 118.60437036023144,
|
||||
"miss_per_tick": 4.657664761142575
|
||||
},
|
||||
"guns": {
|
||||
"counts": {
|
||||
"BitBrain": 55
|
||||
},
|
||||
"runs_with": 7,
|
||||
"switches": 0
|
||||
},
|
||||
"rounds": 49,
|
||||
"runs": 7,
|
||||
"shots": 5483,
|
||||
"variation": {
|
||||
"aimoff_hist": {
|
||||
"-60": 1,
|
||||
"-45": 1,
|
||||
"-35": 3,
|
||||
"-30": 6,
|
||||
"-25": 345,
|
||||
"-20": 443,
|
||||
"-15": 529,
|
||||
"-10": 607,
|
||||
"-5": 610,
|
||||
"0": 662,
|
||||
"5": 577,
|
||||
"10": 486,
|
||||
"15": 442,
|
||||
"20": 482,
|
||||
"25": 282,
|
||||
"30": 5,
|
||||
"40": 1,
|
||||
"50": 1
|
||||
},
|
||||
"aimoff_iqr": 22.82389184815574,
|
||||
"aimoff_mean": -0.6759743337347359,
|
||||
"aimoff_sd": 14.329539671665751,
|
||||
"power_distinct": 647,
|
||||
"power_hist": {
|
||||
"0.5": 1416,
|
||||
"0.52": 22,
|
||||
"0.56": 22,
|
||||
"0.6": 22,
|
||||
"0.62": 24,
|
||||
"0.71": 23,
|
||||
"0.86": 22,
|
||||
"0.9": 25,
|
||||
"1.0": 3212,
|
||||
"2.0": 22
|
||||
},
|
||||
"power_iqr": 0.5,
|
||||
"power_mean": 0.8289494984497539,
|
||||
"power_sd": 0.2477421650874806,
|
||||
"range_mean": 473.40395357200674
|
||||
}
|
||||
},
|
||||
"mix": {
|
||||
"ci": {
|
||||
"lat_ctrl_abs": [
|
||||
52.149269165916934,
|
||||
53.05168240489434
|
||||
],
|
||||
"lat_fixed_abs": [
|
||||
53.751035608179855,
|
||||
55.02334273976535
|
||||
],
|
||||
"miss_per_tick": [
|
||||
4.58467855277936,
|
||||
4.756884860165204
|
||||
]
|
||||
},
|
||||
"dodge": {
|
||||
"hit": 0.10057992026096411,
|
||||
"lat_ctrl_abs": 52.600241347777825,
|
||||
"lat_disp_per_tick": 3.7068191081033643,
|
||||
"lat_fixed_abs": 54.393824082815996,
|
||||
"miss": 117.92160657194027,
|
||||
"miss_per_tick": 4.671767229918404
|
||||
},
|
||||
"guns": {
|
||||
"counts": {
|
||||
"BitBrain": 258,
|
||||
"TMHorizon": 275
|
||||
},
|
||||
"runs_with": 7,
|
||||
"switches": 476
|
||||
},
|
||||
"rounds": 49,
|
||||
"runs": 7,
|
||||
"shots": 5518,
|
||||
"variation": {
|
||||
"aimoff_hist": {
|
||||
"-30": 38,
|
||||
"-25": 255,
|
||||
"-20": 478,
|
||||
"-15": 524,
|
||||
"-10": 594,
|
||||
"-5": 609,
|
||||
"0": 642,
|
||||
"5": 581,
|
||||
"10": 521,
|
||||
"15": 478,
|
||||
"20": 521,
|
||||
"25": 248,
|
||||
"30": 29
|
||||
},
|
||||
"aimoff_iqr": 23.19629241229177,
|
||||
"aimoff_mean": -0.19128061717982736,
|
||||
"aimoff_sd": 14.302914943275754,
|
||||
"power_distinct": 651,
|
||||
"power_hist": {
|
||||
"0.1": 61,
|
||||
"0.5": 1486,
|
||||
"0.52": 24,
|
||||
"0.57": 22,
|
||||
"0.59": 26,
|
||||
"0.7": 23,
|
||||
"0.72": 22,
|
||||
"0.73": 24,
|
||||
"0.77": 23,
|
||||
"1.0": 3074
|
||||
},
|
||||
"power_iqr": 0.5,
|
||||
"power_mean": 0.8109788329104748,
|
||||
"power_sd": 0.2636852717444021,
|
||||
"range_mean": 469.1328637380008
|
||||
}
|
||||
},
|
||||
"pat": {
|
||||
"ci": {
|
||||
"lat_ctrl_abs": [
|
||||
50.37731936384525,
|
||||
52.32848907082813
|
||||
],
|
||||
"lat_fixed_abs": [
|
||||
52.45312790721141,
|
||||
55.0568598598361
|
||||
],
|
||||
"miss_per_tick": [
|
||||
4.555371361130827,
|
||||
4.715270176768353
|
||||
]
|
||||
},
|
||||
"dodge": {
|
||||
"hit": 0.10705334765485142,
|
||||
"lat_ctrl_abs": 51.36125322431849,
|
||||
"lat_disp_per_tick": 3.5758709102416977,
|
||||
"lat_fixed_abs": 53.89836250020473,
|
||||
"miss": 117.0289207480357,
|
||||
"miss_per_tick": 4.641754269524353
|
||||
},
|
||||
"guns": {
|
||||
"counts": {
|
||||
"Pattern": 54
|
||||
},
|
||||
"runs_with": 7,
|
||||
"switches": 0
|
||||
},
|
||||
"rounds": 49,
|
||||
"runs": 7,
|
||||
"shots": 5586,
|
||||
"variation": {
|
||||
"aimoff_hist": {
|
||||
"-30": 3,
|
||||
"-25": 319,
|
||||
"-20": 440,
|
||||
"-15": 558,
|
||||
"-10": 620,
|
||||
"-5": 657,
|
||||
"0": 671,
|
||||
"5": 605,
|
||||
"10": 467,
|
||||
"15": 462,
|
||||
"20": 470,
|
||||
"25": 312,
|
||||
"30": 1,
|
||||
"45": 1
|
||||
},
|
||||
"aimoff_iqr": 22.547808248755246,
|
||||
"aimoff_mean": -0.5000302815020886,
|
||||
"aimoff_sd": 14.161575308720636,
|
||||
"power_distinct": 601,
|
||||
"power_hist": {
|
||||
"0.1": 116,
|
||||
"0.5": 1470,
|
||||
"0.51": 22,
|
||||
"0.55": 27,
|
||||
"0.6": 25,
|
||||
"0.65": 28,
|
||||
"0.68": 27,
|
||||
"0.71": 24,
|
||||
"0.97": 23,
|
||||
"1.0": 3163
|
||||
},
|
||||
"power_iqr": 0.5,
|
||||
"power_mean": 0.8101581095596133,
|
||||
"power_sd": 0.26204101790031575,
|
||||
"range_mean": 469.974151732332
|
||||
}
|
||||
},
|
||||
"tmh": {
|
||||
"ci": {
|
||||
"lat_ctrl_abs": [
|
||||
52.10724351808622,
|
||||
53.26490501694674
|
||||
],
|
||||
"lat_fixed_abs": [
|
||||
53.429352471589546,
|
||||
55.53570708426935
|
||||
],
|
||||
"miss_per_tick": [
|
||||
4.677616757831026,
|
||||
4.7978133827120795
|
||||
]
|
||||
},
|
||||
"dodge": {
|
||||
"hit": 0.09913236108547166,
|
||||
"lat_ctrl_abs": 52.70125777256337,
|
||||
"lat_disp_per_tick": 3.632106164750479,
|
||||
"lat_fixed_abs": 54.584278804670376,
|
||||
"miss": 121.61307223099864,
|
||||
"miss_per_tick": 4.736555181980398
|
||||
},
|
||||
"guns": {
|
||||
"counts": {
|
||||
"TMHorizon": 54
|
||||
},
|
||||
"runs_with": 7,
|
||||
"switches": 0
|
||||
},
|
||||
"rounds": 49,
|
||||
"runs": 7,
|
||||
"shots": 5417,
|
||||
"variation": {
|
||||
"aimoff_hist": {
|
||||
"-35": 2,
|
||||
"-30": 49,
|
||||
"-25": 291,
|
||||
"-20": 485,
|
||||
"-15": 517,
|
||||
"-10": 559,
|
||||
"-5": 625,
|
||||
"0": 640,
|
||||
"5": 561,
|
||||
"10": 519,
|
||||
"15": 442,
|
||||
"20": 483,
|
||||
"25": 215,
|
||||
"30": 29
|
||||
},
|
||||
"aimoff_iqr": 23.130255387415104,
|
||||
"aimoff_mean": -0.8124060010024183,
|
||||
"aimoff_sd": 14.316286775899426,
|
||||
"power_distinct": 594,
|
||||
"power_hist": {
|
||||
"0.5": 1434,
|
||||
"0.52": 20,
|
||||
"0.54": 22,
|
||||
"0.57": 22,
|
||||
"0.63": 21,
|
||||
"0.74": 20,
|
||||
"0.78": 22,
|
||||
"0.81": 20,
|
||||
"0.82": 25,
|
||||
"1.0": 3161
|
||||
},
|
||||
"power_iqr": 0.5,
|
||||
"power_mean": 0.8281302012183867,
|
||||
"power_sd": 0.25254388017875856,
|
||||
"range_mean": 479.94567493925
|
||||
}
|
||||
}
|
||||
},
|
||||
"commit": "f91e12196537c6c1e64353160e8885445e3e222b",
|
||||
"contrasts": {
|
||||
"bb-pat": {
|
||||
"hit": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.02767258957180309,
|
||||
"pooled": -0.007837589858024865,
|
||||
"strat": -0.007832565743913025,
|
||||
"strat_perm_p": 0.02825517040489368
|
||||
},
|
||||
"lat_ctrl_abs": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.3067288086221963,
|
||||
"pooled": 0.7018564566667891,
|
||||
"strat": 0.7377803893237963,
|
||||
"strat_perm_p": 0.2892513836294786
|
||||
},
|
||||
"lat_fixed_abs": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.8858141567142441,
|
||||
"pooled": 0.12730822822030063,
|
||||
"strat": 0.12122394700330835,
|
||||
"strat_perm_p": 0.8904748033789688
|
||||
},
|
||||
"miss": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.3632391494319837,
|
||||
"pooled": 1.5754496121956834,
|
||||
"strat": 1.3754074032538215,
|
||||
"strat_perm_p": 0.3282842994465482
|
||||
},
|
||||
"miss_per_tick": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.7850276725895718,
|
||||
"pooled": 0.015910491618221556,
|
||||
"strat": 0.019163315366865954,
|
||||
"strat_perm_p": 0.7372560442761433
|
||||
}
|
||||
},
|
||||
"mix-bb": {
|
||||
"hit": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.729099912612875,
|
||||
"pooled": 0.0013641624641375638,
|
||||
"strat": 0.0009968895263179599,
|
||||
"strat_perm_p": 0.7931838042528401
|
||||
},
|
||||
"lat_ctrl_abs": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.27993009030002913,
|
||||
"pooled": 0.5371316667925683,
|
||||
"strat": 0.5458010552403781,
|
||||
"strat_perm_p": 0.2869210602971162
|
||||
},
|
||||
"lat_fixed_abs": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.4838333818817361,
|
||||
"pooled": 0.36815335439101204,
|
||||
"strat": 0.4433817857043871,
|
||||
"strat_perm_p": 0.3865423827556073
|
||||
},
|
||||
"miss": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.7407515292746869,
|
||||
"pooled": -0.682763788291112,
|
||||
"strat": 0.048289040397066954,
|
||||
"strat_perm_p": 0.9790270900087387
|
||||
},
|
||||
"miss_per_tick": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.8048354209146519,
|
||||
"pooled": 0.014102468775829102,
|
||||
"strat": 0.003432018973777013,
|
||||
"strat_perm_p": 0.9522283716865715
|
||||
}
|
||||
},
|
||||
"mix-pat": {
|
||||
"hit": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.12962423536265658,
|
||||
"pooled": -0.006473427393887302,
|
||||
"strat": -0.006845548923353688,
|
||||
"strat_perm_p": 0.11505971453539178
|
||||
},
|
||||
"lat_ctrl_abs": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.056219050393242063,
|
||||
"pooled": 1.2389881234593574,
|
||||
"strat": 1.2737305776097125,
|
||||
"strat_perm_p": 0.048062918729973786
|
||||
},
|
||||
"lat_fixed_abs": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.5817069618409554,
|
||||
"pooled": 0.4954615826113127,
|
||||
"strat": 0.5615935091004379,
|
||||
"strat_perm_p": 0.5321875910282552
|
||||
},
|
||||
"miss": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.6242353626565686,
|
||||
"pooled": 0.8926858239045714,
|
||||
"strat": 1.3399115632023169,
|
||||
"strat_perm_p": 0.33177978444509176
|
||||
},
|
||||
"miss_per_tick": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.6644334401398194,
|
||||
"pooled": 0.03001296039405066,
|
||||
"strat": 0.01991009043907127,
|
||||
"strat_perm_p": 0.7489076609379551
|
||||
}
|
||||
},
|
||||
"mix-tmh": {
|
||||
"hit": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.7040489367899796,
|
||||
"pooled": 0.001447559175492455,
|
||||
"strat": 0.0003077780829409617,
|
||||
"strat_perm_p": 0.9440722400233033
|
||||
},
|
||||
"lat_ctrl_abs": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.7984270317506554,
|
||||
"pooled": -0.10101642478556272,
|
||||
"strat": -0.06964226580464877,
|
||||
"strat_perm_p": 0.8584328575589864
|
||||
},
|
||||
"lat_fixed_abs": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.786192834255753,
|
||||
"pooled": -0.1904547218543371,
|
||||
"strat": -0.053909913492993525,
|
||||
"strat_perm_p": 0.9289251383629479
|
||||
},
|
||||
"miss": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.07893970288377512,
|
||||
"pooled": -3.6914656590583093,
|
||||
"strat": -2.1077191144440586,
|
||||
"strat_perm_p": 0.1808913486746286
|
||||
},
|
||||
"miss_per_tick": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.28051267113311973,
|
||||
"pooled": -0.06478795206199361,
|
||||
"strat": -0.07950123064257822,
|
||||
"strat_perm_p": 0.1814739295077192
|
||||
}
|
||||
},
|
||||
"tmh-pat": {
|
||||
"hit": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.01602097290999126,
|
||||
"pooled": -0.007920986569379757,
|
||||
"strat": -0.006327715923026292,
|
||||
"strat_perm_p": 0.06845324788814448
|
||||
},
|
||||
"lat_ctrl_abs": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.04981066122924556,
|
||||
"pooled": 1.34000454824492,
|
||||
"strat": 1.3555453008528295,
|
||||
"strat_perm_p": 0.04573259539761142
|
||||
},
|
||||
"lat_fixed_abs": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.4698514418875619,
|
||||
"pooled": 0.6859163044656498,
|
||||
"strat": 0.631144398873892,
|
||||
"strat_perm_p": 0.49257209437809496
|
||||
},
|
||||
"miss": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.011360326245266531,
|
||||
"pooled": 4.584151482962881,
|
||||
"strat": 3.624637348567873,
|
||||
"strat_perm_p": 0.01893387707544422
|
||||
},
|
||||
"miss_per_tick": {
|
||||
"method": "exact",
|
||||
"perm_p": 0.10806874453830469,
|
||||
"pooled": 0.09480091245604427,
|
||||
"strat": 0.10491883548626751,
|
||||
"strat_perm_p": 0.06728808622196329
|
||||
}
|
||||
}
|
||||
},
|
||||
"rounds": 7,
|
||||
"runs": 7
|
||||
}
|
||||
@@ -0,0 +1,215 @@
|
||||
# Does mixing guns disrupt DrussGT? No.
|
||||
|
||||
**The user's hypothesis.** *"Maybe lucky, but sometimes the BitBrain entered and
|
||||
disrupted DrussGT's surfing data, allowing TMHorizon to hit more."* The claim is
|
||||
that admitting **both** `TR_RACK_TMHORIZON=both` and `TR_RACK_BITBRAIN=both`
|
||||
poisons DrussGT's surfgun: two guns that disagree fire inconsistent bullets, its
|
||||
wave matching fails, and it dodges worse.
|
||||
|
||||
**Verdict: CLEAN NEGATIVE.** The mixing is real (the bot alternates the two guns
|
||||
476 times per 7 runs — measured), but our bullets are **not** more varied (power
|
||||
and firing-bearing spread are flat), DrussGT's dodge quality is **unchanged**
|
||||
(miss/tick, fixed-12 and control-window lateral displacement all non-significant,
|
||||
run-cluster p = 0.28-0.80, and the one marginal result leans the wrong way), and
|
||||
the win difference is noise. The 6/10 was luck.
|
||||
|
||||
Design: shared A/B harness `tools/ab/ab_run.sh` (commit `17c5015`), frozen
|
||||
ModularBot built from HEAD `f91e121`, **4 arms x 7 runs x 7 rounds = 196 rounds vs
|
||||
the real, unmodified DrussGT**. Dodge instrument reuses
|
||||
`common_libs/tests/analyze_drussgt_dodge_vs_power.py` (commit `1adefab`) per arm,
|
||||
driven by the new `tools/ab/ab_dodge_analyze.py`. Attribution cross-check:
|
||||
**199/199** death events have the mapped victim at ~0 energy.
|
||||
|
||||
---
|
||||
|
||||
## 1. The 6/10 reality check (do this first)
|
||||
|
||||
The user watched **one** 10-round battle and saw 6 round wins. Shipped baseline vs
|
||||
real DrussGT is 24/49 = **49%** round wins (replicated). Under a true 49% rate:
|
||||
|
||||
| quantity | value |
|
||||
|---|---|
|
||||
| P(exactly 6 of 10) | **0.197** |
|
||||
| P(at least 6 of 10) | **0.353** |
|
||||
| P(at least 6 of 10) if the rate were a coin flip (0.50) | 0.377 |
|
||||
|
||||
A 6/10 happens **35% of the time** at the honest 49% rate — almost the same as at a
|
||||
pure coin flip. One 10-round battle is **no evidence of anything**; it is a single
|
||||
draw from a distribution whose mean is 4.9 wins. This section is not a caveat, it
|
||||
is the answer to the "6/10 proves it" reading: **it does not**.
|
||||
|
||||
The 4-arm experiment below ran 28 such battles; the mix arm won **24/49 = 49.0%**
|
||||
— numerically the historical baseline, but its own in-session control `pat` won
|
||||
42.9% and the difference is noise (Fisher p = 0.69, see §2).
|
||||
|
||||
---
|
||||
|
||||
## 2. Damage and round wins (the primary, weak metric)
|
||||
|
||||
`TR_RACK_PATTERN=off` in every non-`pat` arm.
|
||||
|
||||
| arm | configuration | dmg/run | dmg taken/run | round wins | win% | shots/run | hits taken/run |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| `pat` | shipped default: Pattern only | 291 | 213 | 21/49 | 42.9% | 815 | 95.7 |
|
||||
| `tmh` | TMHorizon only | 265 | 211 | 18/49 | 36.7% | 788 | 94.9 |
|
||||
| `bb` | BitBrain only (`MEM=decay`) | 278 | 183 | 22/49 | 44.9% | 802 | 92.3 |
|
||||
| `mix` | TMHorizon **+** BitBrain (user's config) | 267 | 210 | **24/49** | **49.0%** | 804 | 93.6 |
|
||||
|
||||
Per-run round wins (wins cluster at 0/7, so the mean alone lies):
|
||||
|
||||
```
|
||||
pat r1..r7: 5 3 3 4 3 1 2 tmh: 1 4 2 2 3 2 4
|
||||
bb : 2 3 4 3 3 4 3 mix: 2 5 2 3 3 4 5
|
||||
```
|
||||
|
||||
Nothing separates. Against `pat`: `mix` damage/run **-24.9** (perm p = 0.097), i.e.
|
||||
mix deals *less* damage; round wins **+0.43/run** (perm p = 0.68; pooled Fisher
|
||||
p = 0.69). The two point estimates even disagree in sign — the signature of noise.
|
||||
|
||||
**Minimum detectable effect at 7 runs/arm (alpha 0.05, 80% power):** dmg/run
|
||||
**39.7** (13.6% of the control mean) and round wins **1.93 of a 3.0 mean (64%)**.
|
||||
Seven runs cannot resolve anything smaller than a huge effect, so this metric is
|
||||
weak by construction and is **not** where the hypothesis is tested.
|
||||
|
||||
---
|
||||
|
||||
## 3. DrussGT's dodge quality, per arm (the strong, per-shot metric)
|
||||
|
||||
Per shot fired by us. `miss/tick` and `lat12` are power-neutral; `lat_ctrl` is the
|
||||
same shot 40 ticks later, bullet gone (a real response is absent there; a
|
||||
geometry/phase difference is present). **Disruption = DrussGT ends up closer to our
|
||||
aim line = these metrics FALL.** 95% CI = run-cluster bootstrap (2 000 reps):
|
||||
5 417-5 586 shots/arm, i.e. ~10x the power of the win counts.
|
||||
|
||||
| arm | shots | miss px | miss/tick [95% CI] | lat12 px [95% CI] | lat_ctrl px [95% CI] | our hit% |
|
||||
|---|---|---|---|---|---|---|
|
||||
| `pat` | 5586 | 117.0 | 4.64 [4.56, 4.72] | 53.9 [52.5, 55.1] | 51.4 [50.4, 52.3] | 0.107 |
|
||||
| `tmh` | 5417 | 121.6 | 4.74 [4.68, 4.80] | 54.6 [53.4, 55.5] | 52.7 [52.1, 53.3] | 0.099 |
|
||||
| `bb` | 5483 | 118.6 | 4.66 [4.59, 4.72] | 54.0 [53.4, 54.7] | 52.1 [51.3, 52.8] | 0.099 |
|
||||
| `mix` | 5518 | 117.9 | 4.67 [4.58, 4.76] | 54.4 [53.8, 55.0] | 52.6 [52.1, 53.1] | 0.101 |
|
||||
|
||||
Between-arm contrast (exact **run-cluster** permutation, C(14,7)=3432; `strat` is
|
||||
range-stratified over 200-450 px+, n-weighted). `mix` vs the best single arm:
|
||||
|
||||
| metric | `mix` - `pat` | perm p | `mix` - `bb` | perm p | `mix` - `tmh` | perm p |
|
||||
|---|---|---|---|---|---|---|
|
||||
| miss/tick | +0.030 | 0.66 | +0.014 | 0.80 | -0.065 | 0.28 |
|
||||
| lat12 | +0.50 | 0.58 | +0.37 | 0.48 | -0.19 | 0.79 |
|
||||
| lat_ctrl | +1.24 | 0.056 | +0.54 | 0.28 | -0.10 | 0.80 |
|
||||
| our hit% | -0.006 | 0.13 | +0.001 | 0.73 | +0.001 | 0.70 |
|
||||
|
||||
(The only result anywhere near significance, `mix`-`pat` lat_ctrl at p = 0.056, has
|
||||
`mix` *higher* — DrussGT drifting **further** off our line, i.e. dodging no worse.)
|
||||
|
||||
Against `pat` (the arm that makes DrussGT look *worst*-dodging) `mix` is, if
|
||||
anything, **slightly higher** on every metric — the opposite of disruption, and
|
||||
nowhere near significance. Against `tmh`/`bb` the sign flips. There is **no
|
||||
dodge-quality degradation under `mix`**.
|
||||
|
||||
**Minimum detectable effect at 7 runs (run-cluster, 80% power):** miss/tick
|
||||
**0.176 px/tick (3.8%)**, lat12 2.73 px (5.1%), lat_ctrl 2.06 px (4.0%), miss
|
||||
3.65 px (3.1%). So the experiment bounds any dodge-quality change to **<~4%**;
|
||||
the observed mix-vs-best difference is +0.03 px/tick, ~6x below the floor.
|
||||
|
||||
### The trap to avoid (this team already hit it once)
|
||||
|
||||
`bb` lands the **fewest hits** (0.099 vs `pat` 0.107, cluster p = 0.028 — real) yet
|
||||
wins **more** rounds (22/49 vs 21/49) and deals comparable damage (278 vs 291).
|
||||
Judging by hit rate alone would condemn `bb`; the objective metrics do not. Hit
|
||||
rate is *not* the verdict — dodge quality is the mechanism, damage/wins are the
|
||||
outcome, and all three are reported above.
|
||||
|
||||
---
|
||||
|
||||
## 4. Liveness: is the mix real, and are our bullets more varied?
|
||||
|
||||
**Gun selection is real (MEASURED, from the bot's own `[config]` switch lines).**
|
||||
The mix genuinely alternates the two guns; the single-gun arms never switch:
|
||||
|
||||
| arm | guns selected ([config] lines) | gun switches (7 runs) |
|
||||
|---|---|---|
|
||||
| `pat` | Pattern: 54 | **0** |
|
||||
| `tmh` | TMHorizon: 54 | **0** |
|
||||
| `bb` | BitBrain: 55 | **0** |
|
||||
| `mix` | TMHorizon: 275, BitBrain: 258 | **476** (68/run) |
|
||||
|
||||
**Our bullets are NOT more varied (MEASURED).** Power is set by the shared
|
||||
energy/range policy, independent of the selected gun; the fired power distribution
|
||||
is identical across arms. The firing-bearing spread (aim offset from the
|
||||
straight-at-target line, degrees) is also flat — the mix's variance
|
||||
(14.30^2) is *not* larger than either component (14.32^2, 14.33^2), so the two
|
||||
guns do not even have a measurably different marginal aim:
|
||||
|
||||
| arm | power mean | power sd | power IQR | aim-offset sd | aim-offset IQR | fire range | gun switches |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| `pat` | 0.810 | 0.262 | 0.500 | 14.16 | 22.55 | 470 px | 0 |
|
||||
| `tmh` | 0.828 | 0.253 | 0.500 | 14.32 | 23.13 | 480 px | 0 |
|
||||
| `bb` | 0.829 | 0.248 | 0.500 | 14.33 | 22.82 | 473 px | 0 |
|
||||
| `mix` | 0.811 | 0.264 | 0.500 | 14.30 | 23.20 | 469 px | 476 |
|
||||
|
||||
**This is the hypothesis failing its own liveness test.** The mechanism needs more
|
||||
varied / inconsistent bullets, and there are none: mixing guns changes *nothing*
|
||||
about the bullet stream (same powers, statistically identical aim spread, same
|
||||
range). Caveat: the aim-offset marginal is geometry-dominated (sd ~14 deg from
|
||||
range/target motion), so it is a *weak* discriminator for a few-degree gun
|
||||
disagreement; but the power channel — the one the hypothesis names ("mixed
|
||||
speeds/powers") — is structurally gun-independent and measured flat, and the
|
||||
gun-switch liveness proves the mix did fire both guns.
|
||||
|
||||
---
|
||||
|
||||
## 5. Rubric answer
|
||||
|
||||
The task set three possible outcomes. This is the third:
|
||||
|
||||
* DrussGT dodge quality **degrades** under `mix` -> mechanism REAL. **Not observed.**
|
||||
* Dodge quality unchanged but our hit rate higher under `mix` -> just noisier aim.
|
||||
**Not observed** (hit rate flat: mix 0.101 vs pat 0.107, p = 0.13).
|
||||
* **Nothing separates -> CLEAN NEGATIVE; the 6/10 was luck.** **<- THIS.**
|
||||
|
||||
**MEASURED**
|
||||
|
||||
* The two guns are both selected and alternate heavily under `mix` (476 switches /
|
||||
7 runs; 0 for every single-gun arm).
|
||||
* DrussGT's power-neutral dodge metrics (miss/tick, fixed-12, control-window) are
|
||||
statistically indistinguishable across all four arms; `mix` - `pat` = +0.03
|
||||
px/tick (p = 0.66), bounded by a 3.8% MDE.
|
||||
* Our own bullets are no more varied under `mix`: power sd 0.264 vs 0.248-0.262,
|
||||
aim-offset sd 14.30 vs 14.16-14.33, same range. Power is a gun-independent
|
||||
policy, so the "mixed powers" channel does not exist here at all.
|
||||
* Damage/run and round wins do not separate either (perm p = 0.097 and 0.68);
|
||||
`mix` won 24/49 = 49.0% vs in-session control `pat` 21/49 = 42.9% (Fisher
|
||||
p = 0.69), while dealing *less* damage. The 6/10 has probability 0.353.
|
||||
* The trap is live in this data: `bb` lands significantly fewer hits than `pat`
|
||||
(p = 0.028) yet wins at least as many rounds.
|
||||
|
||||
**INFERRED**
|
||||
|
||||
* Two alternating guns sharing the same power policy and (marginal) aim
|
||||
distribution do not poison DrussGT's surfgun. Deliberate angle jitter or a
|
||||
randomized power band would be a *different* intervention, not this one; the
|
||||
liveness table says this mix does not even change the bullet statistics.
|
||||
* The 6/10 was a lucky draw, not a signal.
|
||||
|
||||
---
|
||||
|
||||
## 6. Reproducing
|
||||
|
||||
```bash
|
||||
cat > /tmp/ab/mix_arms.txt <<'EOF'
|
||||
pat | | shipped default: Pattern only
|
||||
tmh | TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both | TMHorizon alone
|
||||
bb | TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay | BitBrain alone
|
||||
mix | TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay | user config
|
||||
EOF
|
||||
tools/ab/ab_run.sh --arms /tmp/ab/mix_arms.txt --runs 7 --outdir /tmp/ab/mix --conc 5
|
||||
python3 tools/ab/ab_analyze.py /tmp/ab/mix
|
||||
python3 tools/ab/ab_dodge_analyze.py /tmp/ab/mix --json /tmp/ab/mix_dodge.json --reps 2000
|
||||
```
|
||||
|
||||
Captured outputs are committed (the ~246 MB of raw `.jsonl` captures are
|
||||
gitignored, as for `drussgt_dodge_vs_power.md`):
|
||||
|
||||
* `common_libs/tests/fixtures/gun_mix_ab_report.txt` — `ab_analyze.py` output
|
||||
* `common_libs/tests/fixtures/gun_mix_dodge_report.txt` — per-arm dodge report
|
||||
* `common_libs/tests/fixtures/gun_mix_dodge_results.json` — the same numbers as JSON
|
||||
@@ -0,0 +1,417 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Per-arm DrussGT dodge quality from an `ab_run.sh` session.
|
||||
|
||||
The gun-mixing hypothesis is: admitting TMHorizon AND BitBrain makes DrussGT's
|
||||
surfgun dodge *worse*, because two guns that disagree on the same tick fire
|
||||
inconsistent bullets and poison its wave matching. The random 7-round win
|
||||
counts an A/B session produces cannot see that (a few rounds, huge variance);
|
||||
the shots can. This reads the SAME live captures `ab_analyze.py` uses, applies
|
||||
the SAME per-shot instrument as
|
||||
`common_libs/tests/analyze_drussgt_dodge_vs_power.py` (miss at arrival, miss per
|
||||
flight tick, fixed-window lateral displacement, and the bullet-free CONTROL
|
||||
window), and splits it PER ARM.
|
||||
|
||||
It is a thin driver: it imports the validated instrument, only adds (a) the
|
||||
per-arm split, (b) an aim-offset / power dispersion reading of OUR OWN shooting
|
||||
(the liveness check for the mechanism), and (c) between-arm tests whose
|
||||
uncertainty is clustered by RUN (7 runs per arm -> C(14,7)=3432 exact).
|
||||
The power-neutral metrics (miss/tick, fixed/control-window lateral displacement)
|
||||
are the ones to trust; raw miss distance is dominated by our own aim error and
|
||||
by the flight window (see docs/drussgt_dodge_vs_power.md).
|
||||
|
||||
Usage:
|
||||
python3 tools/ab/ab_dodge_analyze.py /tmp/ab/mix [--json out.json] [--reps 2000]
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import collections
|
||||
import glob
|
||||
import itertools
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import random
|
||||
import re
|
||||
import statistics
|
||||
import sys
|
||||
|
||||
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
REPO = os.path.dirname(os.path.dirname(HERE))
|
||||
sys.path.insert(0, os.path.join(REPO, "common_libs", "tests"))
|
||||
import analyze_drussgt_dodge_vs_power as dodge # noqa: E402
|
||||
|
||||
BAND_LABELS = dodge.BAND_LABELS # 0-100 .. 450+
|
||||
STRAT_BANDS = ["200-300", "300-450", "450+"] # bands with enough shots for a contrast
|
||||
DODGE_METRICS = ["miss", "miss_per_tick", "lat_fixed_abs", "lat_ctrl_abs",
|
||||
"lat_disp_per_tick", "hit"]
|
||||
TEST_METRICS = ["miss", "miss_per_tick", "lat_fixed_abs", "lat_ctrl_abs", "hit"]
|
||||
|
||||
|
||||
class ArmRun(dodge.Run):
|
||||
"""The validated instrument, plus the aim offset we fired at.
|
||||
|
||||
`aim_off` = (fired direction - bearing to DrussGT at the fire tick), in
|
||||
degrees, wrapped to (-180, 180]. It is the gun's answer relative to the
|
||||
straight-at-target line: its dispersion across shots is how *inconsistent*
|
||||
our bullets are, which is exactly the property that would poison a surfgun.
|
||||
"""
|
||||
|
||||
def _shot(self, ev, t0):
|
||||
s = super()._shot(ev, t0)
|
||||
if s is None:
|
||||
return None
|
||||
r0 = self.by_tick[t0]
|
||||
bearing = math.degrees(math.atan2(r0["ey"] - s["_y"], r0["ex"] - s["_x"]))
|
||||
s["aim_off"] = ((ev["dir"] - bearing + 180.0) % 360.0) - 180.0
|
||||
return s
|
||||
|
||||
|
||||
# ── discovery ────────────────────────────────────────────────────────────────
|
||||
def discover_arm(armdir):
|
||||
runs = []
|
||||
for cap in sorted(glob.glob(os.path.join(armdir, "run*.jsonl"))):
|
||||
if cap.endswith(".events.jsonl"):
|
||||
continue
|
||||
ev = cap[:-len(".jsonl")] + ".events.jsonl"
|
||||
rj = cap + ".rounds.json"
|
||||
if os.path.exists(ev) and os.path.exists(rj):
|
||||
runs.append(ArmRun(cap, ev, rj))
|
||||
return runs
|
||||
|
||||
|
||||
# ── shooting-variation (liveness for the mechanism) ──────────────────────────
|
||||
def iqr(xs):
|
||||
if len(xs) < 4:
|
||||
return float("nan")
|
||||
xs = sorted(xs)
|
||||
return dodge.pct(xs, 0.75) - dodge.pct(xs, 0.25)
|
||||
|
||||
|
||||
def shooting_variation(shots):
|
||||
pows = [s["power"] for s in shots]
|
||||
offs = [s["aim_off"] for s in shots]
|
||||
ranges = [s["range"] for s in shots]
|
||||
pwr_hist = collections.Counter(round(p, 2) for p in pows)
|
||||
off_hist = collections.Counter(round(o / 5.0) * 5 for o in offs)
|
||||
return {
|
||||
"power_mean": dodge.mean(pows),
|
||||
"power_sd": statistics.pstdev(pows) if len(pows) > 1 else 0.0,
|
||||
"power_iqr": iqr(pows),
|
||||
"power_distinct": len(set(pows)),
|
||||
"aimoff_mean": dodge.mean(offs),
|
||||
"aimoff_sd": statistics.pstdev(offs) if len(offs) > 1 else 0.0,
|
||||
"aimoff_iqr": iqr(offs),
|
||||
"range_mean": dodge.mean(ranges),
|
||||
"power_hist": dict(pwr_hist.most_common(10)),
|
||||
"aimoff_hist": dict(sorted(off_hist.items())),
|
||||
}
|
||||
|
||||
|
||||
# ── gun-selection liveness (does the mix actually alternate guns?) ────────────
|
||||
CONFIG_GUN_RE = re.compile(r"gun=([A-Za-z]+)")
|
||||
ANSI_RE = re.compile(r"\x1b\[[0-9;]*m")
|
||||
|
||||
|
||||
def gun_liveness(runs):
|
||||
"""From the `[config] gun=<name>` lines the bot emits when its selection
|
||||
changes: how often each gun was selected and how many times the selection
|
||||
switched. A one-gun arm never switches; a real two-gun mix switches often.
|
||||
"""
|
||||
total = collections.Counter()
|
||||
switches = runs_with = 0
|
||||
for r in runs:
|
||||
armdir = os.path.dirname(r.cap_path)
|
||||
num = os.path.basename(r.cap_path)[len("run"):-len(".jsonl")]
|
||||
path = os.path.join(armdir, "run%s.bot.stdout.log" % num)
|
||||
if not os.path.exists(path):
|
||||
continue
|
||||
seq = []
|
||||
for line in open(path, errors="replace"):
|
||||
if "[config]" not in line:
|
||||
continue
|
||||
m = CONFIG_GUN_RE.search(ANSI_RE.sub("", line))
|
||||
if m:
|
||||
seq.append(m.group(1))
|
||||
if seq:
|
||||
runs_with += 1
|
||||
total.update(seq)
|
||||
switches += sum(1 for i in range(1, len(seq)) if seq[i] != seq[i - 1])
|
||||
return total, switches, runs_with, len(runs)
|
||||
|
||||
|
||||
# ── per-run summaries + clustered tests ──────────────────────────────────────
|
||||
def run_summaries(runs_shots, key):
|
||||
"""Per run: (sum, count) for `key`, and per range band."""
|
||||
out = []
|
||||
for ss in runs_shots:
|
||||
vals = [s[key] for s in ss if s.get(key) is not None]
|
||||
bands = {b: [0.0, 0] for b in BAND_LABELS}
|
||||
for s in ss:
|
||||
v = s.get(key)
|
||||
if v is None:
|
||||
continue
|
||||
bands[s["band"]][0] += v
|
||||
bands[s["band"]][1] += 1
|
||||
out.append({"sum": sum(vals), "n": len(vals),
|
||||
"bands": {b: tuple(v) for b, v in bands.items()}})
|
||||
return out
|
||||
|
||||
|
||||
def cluster_ci(summary, reps=2000, seed=1):
|
||||
"""95% CI of the pooled mean, resampling RUNS with replacement."""
|
||||
rng = random.Random(seed)
|
||||
R = len(summary)
|
||||
if R == 0:
|
||||
return float("nan"), float("nan")
|
||||
means = []
|
||||
for _ in range(reps):
|
||||
s = c = 0
|
||||
for _ in range(R):
|
||||
it = summary[rng.randrange(R)]
|
||||
s += it["sum"]
|
||||
c += it["n"]
|
||||
if c:
|
||||
means.append(s / c)
|
||||
means.sort()
|
||||
return dodge.pct(means, 0.025), dodge.pct(means, 0.975)
|
||||
|
||||
|
||||
def _pooled(summary, idx):
|
||||
s = c = 0
|
||||
for i in idx:
|
||||
s += summary[i]["sum"]
|
||||
c += summary[i]["n"]
|
||||
return s, c
|
||||
|
||||
|
||||
def _strat(summary, idx, rest, bands):
|
||||
num = den = 0.0
|
||||
for b in bands:
|
||||
sa = ca = sb = cb = 0.0
|
||||
for i in idx:
|
||||
sa += summary[i]["bands"][b][0]
|
||||
ca += summary[i]["bands"][b][1]
|
||||
for i in rest:
|
||||
sb += summary[i]["bands"][b][0]
|
||||
cb += summary[i]["bands"][b][1]
|
||||
w = ca + cb
|
||||
if ca > 0 and cb > 0:
|
||||
num += w * (sa / ca - sb / cb)
|
||||
den += w
|
||||
return num / den if den else float("nan")
|
||||
|
||||
|
||||
EXACT_CAP = 20_000_000 # 7v7 -> C(14,7)=3432 exact; 30v30 -> Monte-Carlo
|
||||
MC_DRAWS = 1_000_000
|
||||
MC_SEED = 0x5EED5EED
|
||||
|
||||
|
||||
def cluster_perm(summary_a, summary_b, stat):
|
||||
"""Two-sided run-cluster permutation test.
|
||||
|
||||
`stat(summary, idxA, idxB)` is evaluated on the observed split and on
|
||||
relabellings of the pooled RUNS; the null is the difference itself, centred
|
||||
at 0, so p = P(|stat_perm| >= |stat_obs|). Clustering by run keeps the
|
||||
within-run shot correlation, which a shot-level test would ignore. Full
|
||||
enumeration when C(2R,R) <= EXACT_CAP (7v7 -> 3432, exact); otherwise a
|
||||
fixed-seed Monte-Carlo run (reported as such).
|
||||
"""
|
||||
summary = list(summary_a) + list(summary_b)
|
||||
n, na = len(summary), len(summary_a)
|
||||
obs = stat(summary, range(na), range(na, n))
|
||||
ncomb = math.comb(n, na)
|
||||
if ncomb <= EXACT_CAP:
|
||||
cnt = 0
|
||||
for combo in itertools.combinations(range(n), na):
|
||||
cset = set(combo)
|
||||
rest = [i for i in range(n) if i not in cset]
|
||||
v = stat(summary, combo, rest)
|
||||
if v == v and abs(v) >= abs(obs) - 1e-12:
|
||||
cnt += 1
|
||||
return obs, (cnt + 1) / (ncomb + 1), "exact"
|
||||
rng = random.Random(MC_SEED)
|
||||
cnt = 0
|
||||
for _ in range(MC_DRAWS):
|
||||
idx = rng.sample(range(n), na)
|
||||
cset = set(idx)
|
||||
rest = [i for i in range(n) if i not in cset]
|
||||
v = stat(summary, idx, rest)
|
||||
if v == v and abs(v) >= abs(obs) - 1e-12:
|
||||
cnt += 1
|
||||
return obs, (cnt + 1) / (MC_DRAWS + 1), "monte-carlo/%d" % MC_DRAWS
|
||||
|
||||
|
||||
def pooled_stat(summary, idx, rest):
|
||||
sa, ca = _pooled(summary, idx)
|
||||
sb, cb = _pooled(summary, rest)
|
||||
if ca == 0 or cb == 0:
|
||||
return float("nan")
|
||||
return sa / ca - sb / cb
|
||||
|
||||
|
||||
def strat_stat(summary, idx, rest):
|
||||
return _strat(summary, idx, rest, STRAT_BANDS)
|
||||
|
||||
|
||||
# ── report ───────────────────────────────────────────────────────────────────
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("outdir")
|
||||
ap.add_argument("--json", default=None)
|
||||
ap.add_argument("--reps", type=int, default=2000)
|
||||
ap.add_argument("--reference", default=None)
|
||||
a = ap.parse_args()
|
||||
|
||||
session = json.load(open(os.path.join(a.outdir, "session.json")))
|
||||
arm_names = [x["name"] for x in session["arms"]]
|
||||
ref = a.reference or arm_names[0]
|
||||
|
||||
arms = {}
|
||||
for name in arm_names:
|
||||
runs = discover_arm(os.path.join(a.outdir, name))
|
||||
runs_shots = [[s for s in r.shots()] for r in runs]
|
||||
shots = [s for ss in runs_shots for s in ss]
|
||||
rounds = sum(len(r.rounds) for r in runs)
|
||||
arms[name] = {"runs": runs, "runs_shots": runs_shots, "shots": shots,
|
||||
"rounds": rounds,
|
||||
"variation": shooting_variation(shots),
|
||||
"guns": gun_liveness(runs)}
|
||||
|
||||
out = {"commit": session["commit"], "runs": session["runs"],
|
||||
"rounds": session["rounds"], "arms": {}}
|
||||
|
||||
# ── corpus sanity: attribution is per battle, so restate the death check ──
|
||||
bad = tot = 0
|
||||
for name in arm_names:
|
||||
for r in arms[name]["runs"]:
|
||||
for ev in r.events:
|
||||
if ev.get("type") != "death":
|
||||
continue
|
||||
end = r.start[ev["round"]] + r.count[ev["round"]] - 1
|
||||
row = r.by_tick.get(end)
|
||||
if row is None:
|
||||
continue
|
||||
tot += 1
|
||||
vs = r.owner_side.get(ev["victim"])
|
||||
if (vs == "e" and row["ee"] > 1.0) or (vs == "s" and row["se"] > 1.0):
|
||||
bad += 1
|
||||
print("=" * 100)
|
||||
print("PER-ARM DRUSSTGT DODGE QUALITY (session commit %s, %d runs x %d rounds)"
|
||||
% (session["commit"][:9], session["runs"], session["rounds"]))
|
||||
print("attribution cross-check: %d/%d deaths have the mapped victim at ~0 energy"
|
||||
% (tot - bad, tot))
|
||||
print("shipped baseline vs DrussGT is ~49%% round wins; this instrument is per "
|
||||
"SHOT (~thousands/arm), not per round.")
|
||||
|
||||
for name in arm_names:
|
||||
A = arms[name]
|
||||
v = A["variation"]
|
||||
print("\n--- ARM `%s` --- %d runs, %d rounds, %d shots" %
|
||||
(name, len(A["runs"]), A["rounds"], len(A["shots"])))
|
||||
print(" OUR shooting: power mean=%.3f sd=%.3f iqr=%.3f distinct=%d | "
|
||||
"aim-offset sd=%.2f deg iqr=%.2f | fire range mean=%.0f px" %
|
||||
(v["power_mean"], v["power_sd"], v["power_iqr"], v["power_distinct"],
|
||||
v["aimoff_sd"], v["aimoff_iqr"], v["range_mean"]))
|
||||
print(" power histogram (top): " +
|
||||
" ".join("%.2f:%d" % (p, n) for p, n in v["power_hist"].items()))
|
||||
print(" aim-offset histogram (deg bucket): " +
|
||||
" ".join("%+d:%d" % (b, n) for b, n in v["aimoff_hist"].items()))
|
||||
gt, gsw, grw, gnr = A["guns"]
|
||||
print(" GUN SELECTION ([config] switch lines): %s | switches=%d over %d/%d runs"
|
||||
% (" ".join("%s:%d" % kv for kv in gt.most_common()), gsw, grw, gnr))
|
||||
|
||||
print("\n" + "=" * 100)
|
||||
print("DRUSSGT DODGE QUALITY PER ARM (mean, 95%% CI = run-cluster bootstrap)")
|
||||
hdr = (" %-6s %7s | %8s %16s | %10s %16s | %10s %16s | %8s" %
|
||||
("arm", "shots", "miss", "miss/tick [95%CI]", "lat12",
|
||||
"lat12 [95%CI]", "lat_ctrl", "lat_ctrl [95%CI]", "hit%"))
|
||||
print(hdr)
|
||||
print(" " + "-" * (len(hdr) - 2))
|
||||
for name in arm_names:
|
||||
A = arms[name]
|
||||
shots = A["shots"]
|
||||
d = A.setdefault("dodge", {})
|
||||
for key in DODGE_METRICS:
|
||||
vals = [s[key] for s in shots if s.get(key) is not None]
|
||||
d[key] = dodge.mean(vals)
|
||||
ci_mt = cluster_ci(run_summaries(A["runs_shots"], "miss_per_tick"), a.reps)
|
||||
ci_lf = cluster_ci(run_summaries(A["runs_shots"], "lat_fixed_abs"), a.reps)
|
||||
ci_lc = cluster_ci(run_summaries(A["runs_shots"], "lat_ctrl_abs"), a.reps)
|
||||
print(" %-6s %7d | %8.1f %7.2f[%5.2f,%6.2f] | %10.1f [%5.1f,%6.1f] | "
|
||||
"%10.1f [%5.1f,%6.1f] | %8.3f" % (
|
||||
name, len(shots), d["miss"], d["miss_per_tick"],
|
||||
ci_mt[0], ci_mt[1], d["lat_fixed_abs"], ci_lf[0], ci_lf[1],
|
||||
d["lat_ctrl_abs"], ci_lc[0], ci_lc[1], d["hit"]))
|
||||
gt, gsw, grw, gnr = A["guns"]
|
||||
out["arms"][name] = {
|
||||
"runs": len(A["runs"]), "rounds": A["rounds"], "shots": len(shots),
|
||||
"variation": A["variation"], "dodge": {k: d[k] for k in DODGE_METRICS},
|
||||
"guns": {"counts": dict(gt), "switches": gsw, "runs_with": grw},
|
||||
"ci": {"miss_per_tick": ci_mt, "lat_fixed_abs": ci_lf,
|
||||
"lat_ctrl_abs": ci_lc},
|
||||
}
|
||||
|
||||
# ── minimum detectable effect for the shot-level mechanism metrics ───────
|
||||
print("\n" + "=" * 100)
|
||||
print("MINIMUM DETECTABLE EFFECT (run-cluster level, n=%d/arm, alpha=0.05 two-sided,"
|
||||
" 80%% power; MDE = 2.8016*sd_perrun*sqrt(2/n))" % session["runs"])
|
||||
print(" the shot counts are large but the RUNS are what set the between-arm "
|
||||
"uncertainty, so this is the honest floor")
|
||||
print(" %-16s %14s %14s %-22s" % ("metric", "sd(per-run)", "MDE(abs)", "MDE vs ref mean"))
|
||||
Z_ALPHA_POWER = 1.959963984540054 + 0.8416212335729143 # 2.8016
|
||||
for label, key in (("miss", "miss"), ("miss/tick", "miss_per_tick"),
|
||||
("lat12", "lat_fixed_abs"), ("lat_ctrl", "lat_ctrl_abs"),
|
||||
("hit%", "hit")):
|
||||
summ = run_summaries(arms[ref]["runs_shots"], key)
|
||||
means = [it["sum"] / it["n"] for it in summ if it["n"]]
|
||||
if len(means) < 2:
|
||||
continue
|
||||
sd = statistics.stdev(means)
|
||||
mde = Z_ALPHA_POWER * sd * math.sqrt(2.0 / len(means))
|
||||
rm = arms[ref]["dodge"]
|
||||
refmean = rm[key] * (100.0 if key == "hit" else 1.0)
|
||||
print(" %-16s %14.3f %14.3f %-22s" % (
|
||||
label, sd, mde,
|
||||
"%.1f%% of %.3f" % (100 * mde / refmean, refmean) if refmean else "-"))
|
||||
|
||||
# ── between-arm tests, clustered by run ───────────────────────────────────
|
||||
print("\n" + "=" * 100)
|
||||
print("BETWEEN-ARM CONTRASTS (exact run-cluster permutation, C(14,7)=3432)")
|
||||
print(" pooled = pooled-shot difference A-B; strat = range-stratified "
|
||||
"(bands %s, n-weighted)" % ",".join(STRAT_BANDS))
|
||||
print(" a mix that poisons DrussGT should show miss/tick, lat12 and lat_ctrl "
|
||||
"LOWER (worse dodging) than the best single gun")
|
||||
hdr2 = (" %-24s %-24s %9s %8s | %9s %8s" %
|
||||
("metric", "A vs B", "pooled", "perm p", "strat", "perm p"))
|
||||
print(hdr2)
|
||||
print(" " + "-" * (len(hdr2) - 2))
|
||||
compare_pairs = [("mix", n) for n in arm_names if n != "mix" and n != ref]
|
||||
compare_pairs += [(n, ref) for n in arm_names if n not in (ref,)]
|
||||
seen = set()
|
||||
out["contrasts"] = {}
|
||||
for A, B in compare_pairs:
|
||||
if (A, B) in seen:
|
||||
continue
|
||||
seen.add((A, B))
|
||||
for key in TEST_METRICS:
|
||||
sa = run_summaries(arms[A]["runs_shots"], key)
|
||||
sb = run_summaries(arms[B]["runs_shots"], key)
|
||||
obs_p, p_p, meth = cluster_perm(sa, sb, pooled_stat)
|
||||
obs_s, p_s, _ = cluster_perm(sa, sb, strat_stat)
|
||||
print(" %-24s %-24s %+9.3f %8.4f | %+9.3f %8.4f [%s]" %
|
||||
(key, "%s - %s" % (A, B), obs_p, p_p, obs_s, p_s, meth))
|
||||
out["contrasts"].setdefault("%s-%s" % (A, B), {})[key] = {
|
||||
"pooled": obs_p, "perm_p": p_p, "method": meth,
|
||||
"strat": obs_s, "strat_perm_p": p_s}
|
||||
print()
|
||||
|
||||
if a.json:
|
||||
with open(a.json, "w") as f:
|
||||
json.dump(out, f, indent=1, sort_keys=True)
|
||||
print("JSON written to %s" % a.json)
|
||||
return out
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Reference in New Issue
Block a user