MEASURED against the real DrussGT, one frozen binary built from clean HEAD, rack
knobs only (no source edits), 5 arms x 7 runs x 7 rounds, 8 concurrent battles,
judged ONLY on server-side real hit rate from the events sidecar, exact
two-sided permutation test on per-run rates.
arm runs shots hits real % dmg/run p vs full
full (shipped) 7 3898 270 6.93 159 --
onlyPattern 7 4582 494 10.78 287 0.0012 <- BETTER
onlyKNN 7 4033 207 5.13 119 0.1340
onlyLinear 7 3215 105 3.27 65 0.0082
onlyGF 7 3193 72 2.25 45 0.0012
Firing Pattern ALONE gives +3.85pp pooled hit rate and +80% damage per run, and
it fires MORE shots (4582 vs 3898) - it dominates on rate and volume. This is not
"any single gun wins" (full beats Linear, GF and KNN); it is specifically
"Pattern alone beats the rack".
WHY - the virtual fitness signal mis-ranks guns against real outcomes:
- HeadOn is massively over-selected: 31.4% of ticks, the most real shots (1070),
but only 4.5% REAL. It alone drags the rack down.
- Pattern has the best virtual rank and near-best real rate (11.9%, rank 2), yet
is selected only 22.6% of the time.
- Linear's apparent strength was SELECTION BIAS: conditional on being selected it
looked like 15.2% (n=33), but its UNCONDITIONAL rate (onlyLinear) is 3.27%.
Every earlier per-gun "real rate" in this repo is conditional on selection and
is therefore confounded. This experiment is the clean measurement.
NOT YET SETTLED (do not overclaim):
- ONE ADVERSARY. All of this is vs DrussGT. Pattern must be re-checked against
other bots before it becomes the default on this evidence alone.
- Whether a SMALL rack of good guns beats Pattern alone. The selector is negative
value on the CURRENT bloated rack; that does not prove it is negative value on
a rack of only good guns. That is the next experiment and it decides whether
the selection apparatus is fixed or disabled.
- The user's standing directive is to KEEP virtual-fitness selection. This
measurement conflicts with it, so the next step tests the selector on a small
good rack rather than assuming either answer.
Context - three prior selection-side attempts all failed: hysteresis (7.02% ->
5.10%, p=0.002), commitment (7.17% -> 4.44%, p=0.0012), arrival-accuracy
tie-break (7.08%, p=0.88 null). The per-tick random draw is load-bearing on
three independent measurements. This experiment locates the real problem one
level up: which guns are in the rack, and that the virtual signal ranks them
wrongly.
Preserves the reusable harness (tools/ab/which_gun_run_one.sh,
which_gun_arm_env.sh, which_gun_analyze.py) and the full writeup
(docs/selector_negative_value.md).
Hypothesis under test (from the gun audit, which named the tie-band as "the
lever that matters most"): `bmPath` is deliberately generous (2.3-3.6x
`bmPoint`), so a gun can sit in the tied band on a ray that sweeps the target's
path while its bullets ARRIVE badly. So: keep the `path`-ranked band (path beat
point on real hit rate 7.43% vs 4.70%, z=5.56), but narrow the random draw
inside it using a parallel `point` (arrival-accuracy) window.
RESULT: NO EFFECT. Real DrussGT, ONE frozen binary (/tmp/ModularBot_tieband,
md5 2c0c56e6...), env knobs only, 7 runs x 7 rounds per arm, server-side events
sidecar, exact two-sided permutation test on per-run rates.
arm runs shots real % dmg/run d p
tbbase (shipped) 7 4128 7.17 175 -- --
tbpt path-rank + point-narrow 7 3938 7.08 165 +0.14 0.88
tbpc =commit control 7 3759 4.44 98 +2.74 0.0012
tbpt25 point margin 0.25 7 3683 5.59 119 +1.65 0.20
tbtie05 / tbtie40 (band width) 7 3937/3917 5.84/6.28 133/144 1.49/1.00 0.11/0.25
tbwin50 (SelectorWindow=50) 7 3983 6.05 139 +1.20 0.11
tbfloor10 (FloorPeakFrac=0.10) 7 3829 5.33 118 +2.12 0.11
tbpt vs base: fully overlapping ranges, p=0.88. This is a REAL null, not a dead
arm - the mechanism was live, and it visibly changed the selected-gun mix
(Pattern 24%->16%, Accel 6%->16%, Tsetlin ~0%->13%).
CONTROL VALIDATED, AND THIS IS THE THIRD TIME: removing the random draw inside
the band is SIGNIFICANTLY WORSE (4.44%, p=0.0012). Combined with the earlier
hysteresis A/B (7.02% -> 5.10% for commitment) and the light-hysteresis result,
the selector's per-tick randomness is now load-bearing on three independent
measurements. Narrowing the band on ANY second virtual statistic has not helped.
Every knob swept (band width, floor, window) is nominally worse than shipped at
n=7; that is "no credible win" rather than "proven harm" (sd ~1.8pp, ~1pp
resolution, underpowered).
Shipped default stays `GUN_SELECTOR_TIEBREAK=off`; the feature is opt-in, fully
guarded, and costs zero extra work on the default path (point windows are scored
only when the mode is on).
Guards: test_selector_tiebreak 19 (new, pure), test_gun_harness 39,
test_vbullet_metric 11, test_adaptive_radar 41, test_tfil_ring_weights 24,
test_power_policy 26, test_ram_decision 28, test_rack_membership 38,
acceptance_offline_vs_online 12/12 PASS (offline path calls neither
chooseFromFit nor the tie-break).
STRATEGIC CONCLUSION: three selection-side attempts have now failed (hysteresis,
commitment, point tie-break). The selector is at a local optimum and the
remaining lever is the QUALITY OF THE GUNS, not the selection among them.
The belief "BotDeathEvent never reaches ModularBot, so enemyTracker keeps dead
enemies alive forever" was written into a code comment and then believed twice.
It is FALSE. Measured in a 7-bot melee with a per-tick probe comparing
enemyTracker's alive count against the server's getEnemyCount():
metric 1.3.1 (20 rd) 0.35.5 (15 rd)
observed enemy deaths 83 68
...non-round-ending 83 (100%) 66 (97%)
ekBotDeath events DROPPED 0 0
max dispatch lag (turns behind) 1 1
phantom ticks 1 / 16,820 1 / 12,596
MAX CORPSE LIFETIME 0 ticks 0 ticks
victims still alive at round end 0 0
onBotDeath fires for every death, including non-round-ending ones. The
API-level event-drop mechanism IS real (test_event_drop_mechanism.nim proves
it: ekBotDeath is not in isCritical and MAX_EVENTS_AGE=2) - the bot simply
never falls far enough behind for it to trigger (max lag 1 turn).
Removed:
- reconcileWithServer + ReconcilePersistTicks/mismatchTicks/sawServerAlive
(uncommitted, and ON BY DEFAULT despite the premise being false). Its own
comment admitted a shorter window once KILLED A LIVE ENEMY ("it fired three
more times after the tracker marked it dead") - a latent mis-prune path
defending against a bug that does not exist.
- The radar's CorpseTicks=40 filter and the same-class age>60 filter in
recordRadarStats, both carrying the false comment. Removal changes no real
behaviour: buildState feeds the radar enemyTracker.allAlive(), so a dead
enemy never reaches computeScan.
Kept:
- The TR_TRACKER_PROBE instrument (default OFF), which produced the table above.
- test_event_drop_mechanism.nim - the drop mechanism is a genuine library
behaviour worth guarding.
- isAlive/aliveCount on the tracker.
Added: docs/tracker_death_events.md (the durable negative, so this is not
re-invented a third time) and test_enemy_tracker_death.nim (13 checks) in place
of the test for the deleted feature.
Guards: test_gun_harness 39/39, test_vbullet_metric 11, test_power_selection 3,
test_adaptive_radar 41/41, test_event_drop_mechanism 6, test_enemy_tracker_death
13, acceptance 12/12, ModularBot compiles.
Three corrections, all prompted by later measurements:
1. The virtual-vs-real rank correlation is NOT robustly negative. Six
independent Spearman measurements now exist (-0.374, +0.335, +0.522,
-0.371, -0.073, -0.037) and the sign flips on large samples, so it is near
zero on average. The honest headline is that virtual hit rate is a POOR
RANKER, not an inverted one. The report said 'not weak - it is inverted' in
six places; it now says so in none. The practical conclusion (do not trust
it for ranking) is unchanged; the mechanism claimed was wrong.
2. The offline==online acceptance is FIXED, not flaky. Root cause was that the
replay spawned gun 13 (TMSelect) while the live rack has it disabled, and
the shared VirtualTracker ring is ORDER-SENSITIVE, so gun 13's extra 4
bullets/tick permuted the per-tick resolution order for every other gun and
shifted the learning guns' observations. After closing gun 13's ready gate
offline the live and offline KNN traces are byte-identical (904/904 lines,
empty diff). 5/5 consecutive runs now report 12/12 exact with the death
boundary included. Recorded with the lesson: a flaky proof was hiding a real
bug. Also records the general A/B confound - disabling a gun removes its 4
spawns/tick from the shared ring, perturbing resolution order for the rest.
3. Pruning was tested and does NOT help, so the verdict for Tsetlin and
Displace changes from an implied drop to BELOW OVERALL - KEEP. 15 paired
runs: baseline 6.18%, Tsetlin-off 5.76%, Tsetlin+Displace-off 5.46%;
paired permutation p=0.57 and p=0.21; distributions completely overlap; a
non-surfer control showed no separation. Being below average does not
justify removal.
Also records the tie-break randomness fix, and quotes run counts with every
rate (6.95% over 13 runs vs 6.18% over 15 runs, same binary) rather than
presenting a single figure as definitive.
Replaces the stale 2026-09-20 docs, which predated per-gun real attribution,
the offline gun range and the DrussGT boss, and whose verdicts were built on
virtual hit rates that turned out to be ANTI-correlated with reality.
docs/gun_rack_analysis.md (751 lines) covers: the test infrastructure
described honestly (offline range with its flaky-acceptance caveat, the 20
fixtures and what each set is good for, the live boss, and the A/B methodology
of per-run server-side real hit rate with an explicit overlap test); the
virtual-vs-real metric lesson with Spearman -0.374 and the point-vs-path A/B;
per-gun real performance and the 16-rule ranking A/B; the offline per-fixture
gun matrix; KEEP/MARGINAL/BELOW verdicts; and the five root-cause bugs with
before/after numbers.
docs/gun_rack_summary.md (58 lines) is the verdict table plus top actions.
The '~230 point' score-noise band that has been steering methodology all night
was re-derived from the artifacts rather than asserted: the 13 shipped-config
run scores span 175-526, s.d. ~105, i.e. a ~210-point 2-s.d. band.
Caveats recorded verbatim rather than softened: the offline==online acceptance
is flaky (typically 11/12 on unmodified HEAD), fixtures are perfect-information
and therefore optimistic vs live play, per-gun real N is small so single-gun
ordering is indicative, the headline numbers come from ONE wave-surfer
adversary, and HeadOn must stay despite being lowest because it is the floor
fallback (disabling it: 5.08% / 175 dmg vs 6.95% / 251 dmg).
Measured hazard counts for a Java->Nim port, with the correction that only
22 of 28 top-level classes ship source (6 do not: 5 in the gun package plus
dMove/Scan), so a full port would need a decompiler while a shim would not
care at all.
Real traps: 193 float / 35 casts / 147 literals / 60 float[] in the movement
closure (the danger histogram is float[171] -- porting 32-bit Java floats to
Nim's default float64 diverges silently); ~68 non-final statics; and 4 Java
single-& sites with side effects, which break under Nim's short-circuiting
'and'. Non-issues, correcting earlier assumptions: 0 sites of %-on-negative
(angle normalisation is floor-based) and FastTrig has no lookup tables, it is
7 coefficient-exact polynomials.
Movement scoping: 5007 LOC across 14 files, ~4.8-6.1k Nim LOC, 4-8 focused
agent-days to first-compiles. It can be ported WITHOUT the gun (data flows
movement->gun only), but the harness never routes HitByBullet into movement
modules, so the danger bins would never train -- that plumbing is the real
blocker, not the translation.
tm-learning-tracks.md covers three things, all marked [FACT]/[INFERENCE]/
[UNKNOWN]:
- Section A: the Tsetlin gun's label is measured against the wrong baseline.
predX = linearX + cx, so rx = actual - predX = delta - cx, and inside
tmLearnOne error = residual - predicted = (delta - cx) - cx = delta - 2cx.
The fixed point is cx = delta/2 -- HALF the correction needed, even with
perfect Granmo feedback. Fix: store linearX/linearY in TmTrace and train
on delta. Also: hits zero the label instead of carrying their true
residual, and the per-clause step is magnitude-blind.
- Section B: what a TM is actually good at (AND-clauses over binary
literals, readable output) and why this repo suits it -- the gun already
builds an 83-bit x 10-frame Gray-coded window (870 bits). Includes a
falsifiable known-rule benchmark proposal.
- Section C: delayed-reward learning belongs to the MOVEMENT layer, not the
gun. The gun's outcome is delayed but exactly pairable via
(fireTick, powerBin), so its effective lambda is 1 and discounting would
only destroy information.
Also records that docs/papers/tm-deb-paper.pdf was deleted by the user as
AI-generated and unverifiable, while the Granmo-based feedback diff in
tm-deb-assessment.md stands on its own.
Verdict: the document (docs/papers/tm-deb-paper.pdf, 'Generated by
Gemini Notebook') solves temporal credit assignment under delayed
reward, which is not our problem. Our gun's failure is clause
saturation (~131 of 1740 literals included per clause -> conjunction
fires with probability ~2^-131 -> correction identically 0), and
TM-DEB only scales update FREQUENCY by gamma^dt, so it would leave
the fixed point untouched and additionally delete the long-range
feedback we need: gamma^90 = 4.4e-7 at the paper's own gamma=0.85,
and those are the shots whose lead matters most.
Credibility signals recorded in the doc: reference [2] misattributes
authors/venue/year; Algorithm Spec 2 steps 9-12 are literal '%'
placeholders so the automata update is simply absent; Table 1 is
titled 'Expected' and reports never-measured accuracies; Eq. 12 is
not a faithful copy of Granmo's Lemma 2.
The audit also produced the actionable result: a line-by-line diff of
Granmo Table 2/3 feedback against tmLearnOne, identifying why the
automata saturate - Type I never conditions on the clause output so it
omits Granmo's c=0, lk=1 -> -1 w.p. 1/s counter-force (true literals
ratchet toward Include), Type II's guard cOut==1 AND lits[lit]==0 is
unsatisfiable and therefore dead code, and the (T - clip(v,-T,T))/(2T)
resource allocation is missing entirely.
Extracted concrete numbers from 13 papers in docs/papers/neuroevolution/.
Key findings: use CMA-ES or mutation-only truncation GA, mutate ALL weights
(not 5%), sigma=0.005-0.01, pop=64-200, no crossover, single elite.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Covers forward/reverse decision, proportional steering with speed-dependent
turn rate clamping, and deceleration using the existing getNewTargetSpeed util.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>