Replaces the stale 2026-09-20 docs, which predated per-gun real attribution,
the offline gun range and the DrussGT boss, and whose verdicts were built on
virtual hit rates that turned out to be ANTI-correlated with reality.
docs/gun_rack_analysis.md (751 lines) covers: the test infrastructure
described honestly (offline range with its flaky-acceptance caveat, the 20
fixtures and what each set is good for, the live boss, and the A/B methodology
of per-run server-side real hit rate with an explicit overlap test); the
virtual-vs-real metric lesson with Spearman -0.374 and the point-vs-path A/B;
per-gun real performance and the 16-rule ranking A/B; the offline per-fixture
gun matrix; KEEP/MARGINAL/BELOW verdicts; and the five root-cause bugs with
before/after numbers.
docs/gun_rack_summary.md (58 lines) is the verdict table plus top actions.
The '~230 point' score-noise band that has been steering methodology all night
was re-derived from the artifacts rather than asserted: the 13 shipped-config
run scores span 175-526, s.d. ~105, i.e. a ~210-point 2-s.d. band.
Caveats recorded verbatim rather than softened: the offline==online acceptance
is flaky (typically 11/12 on unmodified HEAD), fixtures are perfect-information
and therefore optimistic vs live play, per-gun real N is small so single-gun
ordering is indicative, the headline numbers come from ONE wave-surfer
adversary, and HeadOn must stay despite being lowest because it is the floor
fallback (disabling it: 5.08% / 175 dmg vs 6.95% / 251 dmg).
Measured hazard counts for a Java->Nim port, with the correction that only
22 of 28 top-level classes ship source (6 do not: 5 in the gun package plus
dMove/Scan), so a full port would need a decompiler while a shim would not
care at all.
Real traps: 193 float / 35 casts / 147 literals / 60 float[] in the movement
closure (the danger histogram is float[171] -- porting 32-bit Java floats to
Nim's default float64 diverges silently); ~68 non-final statics; and 4 Java
single-& sites with side effects, which break under Nim's short-circuiting
'and'. Non-issues, correcting earlier assumptions: 0 sites of %-on-negative
(angle normalisation is floor-based) and FastTrig has no lookup tables, it is
7 coefficient-exact polynomials.
Movement scoping: 5007 LOC across 14 files, ~4.8-6.1k Nim LOC, 4-8 focused
agent-days to first-compiles. It can be ported WITHOUT the gun (data flows
movement->gun only), but the harness never routes HitByBullet into movement
modules, so the danger bins would never train -- that plumbing is the real
blocker, not the translation.
tm-learning-tracks.md covers three things, all marked [FACT]/[INFERENCE]/
[UNKNOWN]:
- Section A: the Tsetlin gun's label is measured against the wrong baseline.
predX = linearX + cx, so rx = actual - predX = delta - cx, and inside
tmLearnOne error = residual - predicted = (delta - cx) - cx = delta - 2cx.
The fixed point is cx = delta/2 -- HALF the correction needed, even with
perfect Granmo feedback. Fix: store linearX/linearY in TmTrace and train
on delta. Also: hits zero the label instead of carrying their true
residual, and the per-clause step is magnitude-blind.
- Section B: what a TM is actually good at (AND-clauses over binary
literals, readable output) and why this repo suits it -- the gun already
builds an 83-bit x 10-frame Gray-coded window (870 bits). Includes a
falsifiable known-rule benchmark proposal.
- Section C: delayed-reward learning belongs to the MOVEMENT layer, not the
gun. The gun's outcome is delayed but exactly pairable via
(fireTick, powerBin), so its effective lambda is 1 and discounting would
only destroy information.
Also records that docs/papers/tm-deb-paper.pdf was deleted by the user as
AI-generated and unverifiable, while the Granmo-based feedback diff in
tm-deb-assessment.md stands on its own.
Verdict: the document (docs/papers/tm-deb-paper.pdf, 'Generated by
Gemini Notebook') solves temporal credit assignment under delayed
reward, which is not our problem. Our gun's failure is clause
saturation (~131 of 1740 literals included per clause -> conjunction
fires with probability ~2^-131 -> correction identically 0), and
TM-DEB only scales update FREQUENCY by gamma^dt, so it would leave
the fixed point untouched and additionally delete the long-range
feedback we need: gamma^90 = 4.4e-7 at the paper's own gamma=0.85,
and those are the shots whose lead matters most.
Credibility signals recorded in the doc: reference [2] misattributes
authors/venue/year; Algorithm Spec 2 steps 9-12 are literal '%'
placeholders so the automata update is simply absent; Table 1 is
titled 'Expected' and reports never-measured accuracies; Eq. 12 is
not a faithful copy of Granmo's Lemma 2.
The audit also produced the actionable result: a line-by-line diff of
Granmo Table 2/3 feedback against tmLearnOne, identifying why the
automata saturate - Type I never conditions on the clause output so it
omits Granmo's c=0, lk=1 -> -1 w.p. 1/s counter-force (true literals
ratchet toward Include), Type II's guard cOut==1 AND lits[lit]==0 is
unsatisfiable and therefore dead code, and the (T - clip(v,-T,T))/(2T)
resource allocation is missing entirely.
Extracted concrete numbers from 13 papers in docs/papers/neuroevolution/.
Key findings: use CMA-ES or mutation-only truncation GA, mutate ALL weights
(not 5%), sigma=0.005-0.01, pop=64-200, no crossover, single elite.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Covers forward/reverse decision, proportional steering with speed-dependent
turn rate clamping, and deceleration using the existing getNewTargetSpeed util.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>