f91e12196537c6c1e64353160e8885445e3e222b
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
31c7c01d28 |
SHIPPED: the default rack is now Pattern-only (+49% hit rate, +66% damage on the boss)
`DefaultRackMembership` now admits Pattern (id 5) and marks all 14 other guns `rmOff`. **The selector mechanism is untouched** - `chooseFromFit`, the floor/band logic, the hysteresis and the virtual-fitness plumbing are all intact and functional. Only the rack membership changed, so this is reverted by env alone. Evidence (measured, replicated three times, 10 adversaries): Pattern alone gives 10.36% real hit rate / 264 damage per run vs the full rack's 6.93% / 159. Pattern significantly wins on DrussGT, Corners, Crazy and PatternMover, ties on three, and the full rack never significantly beats it on ANY adversary. Mechanism: the virtual signal keeps ranking the wrong guns first (HeadOn 46% of ticks at 2.0% real; Linear 57.7% at 6.0% real while Pattern sits at 11.1%). **THIS CONTRADICTS THE USER'S STANDING DIRECTIVE** to keep virtual-fitness selection. Recorded plainly in docs/selector_negative_value.md with a SHIPPED DECISION banner rather than done quietly: the mechanism is retained and one env var away, because the measurement says it is negative value on every rack size tested and on 10/10 adversaries. Revert one-liner (no rebuild): TR_RACK_PATTERN=both TR_RACK_HEADON=both TR_RACK_LINEAR=both TR_RACK_TSETLIN=both \ TR_RACK_CIRCULAR=both TR_RACK_GUESSFACTOR=both TR_RACK_WALLBOUNCE=both \ TR_RACK_ACCEL=both TR_RACK_STOPSHOT=both TR_RACK_DISPLACE=both TR_RACK_AVGLEAD=both \ TR_RACK_DECAYGF=both TR_RACK_KNN=both TR_RACK_TMSELECT=both ./out/ModularBot The unit test `testRevertOverrideRestoresFullRack` exercises exactly this table. FLOOR PATH, verified not assumed: `chooseFromFit` already returns `admitted[0]` on the floor path, so it respects admission by construction. Cold field + shipped default -> floor returns Pattern (id 5), NOT HeadOn. With an explicit all-`both` membership the same cold field returns gun 0 (HeadOn) - the old behaviour. Four assertions in `testFloorRespectsAdmission`. LIVENESS: one 1-round battle with NO overrides -> Pattern selected 105/105 = 100%, every other gun 0 including TMPattern. Honesty caveat retained in the doc: 4 of the 10 opponents were Tank Royale sample-bot PORTS rather than the original classic jars (only DrussGT is a real classic jar through the shim). Guards: test_rack_membership 48 (was 38; new floor/revert/default checks), test_tm_pattern_registration 20 (5 checks hard-coded the old default and were updated to assert the new one, with the TMPATTERN parity proof moved onto an explicit old-rack table), test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28, test_selector_tiebreak 19, test_tm_pattern_rack_live 4, test_tm_pattern_learning 3, acceptance_offline_vs_online 12/12 VERDICT PASS. ModularBot compiles. FOLLOW-ON THIS EXPOSED: membership filters SELECTION but not virtual-bullet SPAWNING, so under `onlyPattern` the 13 unselected guns still predict and spawn every tick. Tsetlin alone is ~5.3 ms/tick (~41% of the 13.16 ms per-tick budget), so we are still paying for it while never using it. Gating spawn on admission would reclaim that; it was deliberately NOT done here because it would alter the measurement protocol mid-A/B. |
||
|
|
394b3deeed |
SETTLED: Pattern alone is the best single gun IN GENERAL, not just vs DrussGT
Closes the one-adversary caveat that blocked shipping `onlyPattern`. Two arms
(`full` vs `onlyPattern`, rack knobs only), one frozen binary from clean HEAD
(
|
||
|
|
a54ae6a162 |
SETTLED: no small rack beats Pattern alone; the selector is negative value on a GOOD rack
Follow-up to |
||
|
|
e0666a562d |
The gun selector is NEGATIVE value: Pattern alone beats the full rack (p=0.0012)
MEASURED against the real DrussGT, one frozen binary built from clean HEAD, rack knobs only (no source edits), 5 arms x 7 runs x 7 rounds, 8 concurrent battles, judged ONLY on server-side real hit rate from the events sidecar, exact two-sided permutation test on per-run rates. arm runs shots hits real % dmg/run p vs full full (shipped) 7 3898 270 6.93 159 -- onlyPattern 7 4582 494 10.78 287 0.0012 <- BETTER onlyKNN 7 4033 207 5.13 119 0.1340 onlyLinear 7 3215 105 3.27 65 0.0082 onlyGF 7 3193 72 2.25 45 0.0012 Firing Pattern ALONE gives +3.85pp pooled hit rate and +80% damage per run, and it fires MORE shots (4582 vs 3898) - it dominates on rate and volume. This is not "any single gun wins" (full beats Linear, GF and KNN); it is specifically "Pattern alone beats the rack". WHY - the virtual fitness signal mis-ranks guns against real outcomes: - HeadOn is massively over-selected: 31.4% of ticks, the most real shots (1070), but only 4.5% REAL. It alone drags the rack down. - Pattern has the best virtual rank and near-best real rate (11.9%, rank 2), yet is selected only 22.6% of the time. - Linear's apparent strength was SELECTION BIAS: conditional on being selected it looked like 15.2% (n=33), but its UNCONDITIONAL rate (onlyLinear) is 3.27%. Every earlier per-gun "real rate" in this repo is conditional on selection and is therefore confounded. This experiment is the clean measurement. NOT YET SETTLED (do not overclaim): - ONE ADVERSARY. All of this is vs DrussGT. Pattern must be re-checked against other bots before it becomes the default on this evidence alone. - Whether a SMALL rack of good guns beats Pattern alone. The selector is negative value on the CURRENT bloated rack; that does not prove it is negative value on a rack of only good guns. That is the next experiment and it decides whether the selection apparatus is fixed or disabled. - The user's standing directive is to KEEP virtual-fitness selection. This measurement conflicts with it, so the next step tests the selector on a small good rack rather than assuming either answer. Context - three prior selection-side attempts all failed: hysteresis (7.02% -> 5.10%, p=0.002), commitment (7.17% -> 4.44%, p=0.0012), arrival-accuracy tie-break (7.08%, p=0.88 null). The per-tick random draw is load-bearing on three independent measurements. This experiment locates the real problem one level up: which guns are in the rack, and that the virtual signal ranks them wrongly. Preserves the reusable harness (tools/ab/which_gun_run_one.sh, which_gun_arm_env.sh, which_gun_analyze.py) and the full writeup (docs/selector_negative_value.md). |