Commit Graph

15 Commits

Author SHA1 Message Date
SirStone 013b9fe01e docs: correct the overstated 'inverted metric' claim; record tonight's fixes
Three corrections, all prompted by later measurements:

1. The virtual-vs-real rank correlation is NOT robustly negative. Six
   independent Spearman measurements now exist (-0.374, +0.335, +0.522,
   -0.371, -0.073, -0.037) and the sign flips on large samples, so it is near
   zero on average. The honest headline is that virtual hit rate is a POOR
   RANKER, not an inverted one. The report said 'not weak - it is inverted' in
   six places; it now says so in none. The practical conclusion (do not trust
   it for ranking) is unchanged; the mechanism claimed was wrong.
2. The offline==online acceptance is FIXED, not flaky. Root cause was that the
   replay spawned gun 13 (TMSelect) while the live rack has it disabled, and
   the shared VirtualTracker ring is ORDER-SENSITIVE, so gun 13's extra 4
   bullets/tick permuted the per-tick resolution order for every other gun and
   shifted the learning guns' observations. After closing gun 13's ready gate
   offline the live and offline KNN traces are byte-identical (904/904 lines,
   empty diff). 5/5 consecutive runs now report 12/12 exact with the death
   boundary included. Recorded with the lesson: a flaky proof was hiding a real
   bug. Also records the general A/B confound - disabling a gun removes its 4
   spawns/tick from the shared ring, perturbing resolution order for the rest.
3. Pruning was tested and does NOT help, so the verdict for Tsetlin and
   Displace changes from an implied drop to BELOW OVERALL - KEEP. 15 paired
   runs: baseline 6.18%, Tsetlin-off 5.76%, Tsetlin+Displace-off 5.46%;
   paired permutation p=0.57 and p=0.21; distributions completely overlap; a
   non-surfer control showed no separation. Being below average does not
   justify removal.

Also records the tie-break randomness fix, and quotes run counts with every
rate (6.95% over 13 runs vs 6.18% over 15 runs, same binary) rather than
presenting a single figure as definitive.
2026-09-21 06:59:00 +02:00
SirStone 19410164f1 docs: definitive gun-rack report on real measured numbers
Replaces the stale 2026-09-20 docs, which predated per-gun real attribution,
the offline gun range and the DrussGT boss, and whose verdicts were built on
virtual hit rates that turned out to be ANTI-correlated with reality.

docs/gun_rack_analysis.md (751 lines) covers: the test infrastructure
described honestly (offline range with its flaky-acceptance caveat, the 20
fixtures and what each set is good for, the live boss, and the A/B methodology
of per-run server-side real hit rate with an explicit overlap test); the
virtual-vs-real metric lesson with Spearman -0.374 and the point-vs-path A/B;
per-gun real performance and the 16-rule ranking A/B; the offline per-fixture
gun matrix; KEEP/MARGINAL/BELOW verdicts; and the five root-cause bugs with
before/after numbers.
docs/gun_rack_summary.md (58 lines) is the verdict table plus top actions.

The '~230 point' score-noise band that has been steering methodology all night
was re-derived from the artifacts rather than asserted: the 13 shipped-config
run scores span 175-526, s.d. ~105, i.e. a ~210-point 2-s.d. band.

Caveats recorded verbatim rather than softened: the offline==online acceptance
is flaky (typically 11/12 on unmodified HEAD), fixtures are perfect-information
and therefore optimistic vs live play, per-gun real N is small so single-gun
ordering is indicative, the headline numbers come from ONE wave-surfer
adversary, and HeadOn must stay despite being lowest because it is the floor
fallback (disabling it: 5.08% / 175 dmg vs 6.95% / 251 dmg).
2026-09-21 06:37:20 +02:00
SirStone 76ae6170f8 docs(research): portability audit of DrussGT 3.1.4159
Measured hazard counts for a Java->Nim port, with the correction that only
22 of 28 top-level classes ship source (6 do not: 5 in the gun package plus
dMove/Scan), so a full port would need a decompiler while a shim would not
care at all.

Real traps: 193 float / 35 casts / 147 literals / 60 float[] in the movement
closure (the danger histogram is float[171] -- porting 32-bit Java floats to
Nim's default float64 diverges silently); ~68 non-final statics; and 4 Java
single-& sites with side effects, which break under Nim's short-circuiting
'and'. Non-issues, correcting earlier assumptions: 0 sites of %-on-negative
(angle normalisation is floor-based) and FastTrig has no lookup tables, it is
7 coefficient-exact polynomials.

Movement scoping: 5007 LOC across 14 files, ~4.8-6.1k Nim LOC, 4-8 focused
agent-days to first-compiles. It can be ported WITHOUT the gun (data flows
movement->gun only), but the harness never routes HitByBullet into movement
modules, so the danger bins would never train -- that plumbing is the real
blocker, not the translation.
2026-09-20 23:44:42 +02:00
SirStone 54e9757567 docs(research): TM learning tracks + note that the TM-DEB source was deleted
tm-learning-tracks.md covers three things, all marked [FACT]/[INFERENCE]/
[UNKNOWN]:
- Section A: the Tsetlin gun's label is measured against the wrong baseline.
  predX = linearX + cx, so rx = actual - predX = delta - cx, and inside
  tmLearnOne error = residual - predicted = (delta - cx) - cx = delta - 2cx.
  The fixed point is cx = delta/2 -- HALF the correction needed, even with
  perfect Granmo feedback. Fix: store linearX/linearY in TmTrace and train
  on delta. Also: hits zero the label instead of carrying their true
  residual, and the per-clause step is magnitude-blind.
- Section B: what a TM is actually good at (AND-clauses over binary
  literals, readable output) and why this repo suits it -- the gun already
  builds an 83-bit x 10-frame Gray-coded window (870 bits). Includes a
  falsifiable known-rule benchmark proposal.
- Section C: delayed-reward learning belongs to the MOVEMENT layer, not the
  gun. The gun's outcome is delayed but exactly pairable via
  (fireTick, powerBin), so its effective lambda is 1 and discounting would
  only destroy information.

Also records that docs/papers/tm-deb-paper.pdf was deleted by the user as
AI-generated and unverifiable, while the Granmo-based feedback diff in
tm-deb-assessment.md stands on its own.
2026-09-20 23:28:31 +02:00
SirStone 669f9acd41 docs(research): assess TM-DEB paper against the Tsetlin gun's real failure
Verdict: the document (docs/papers/tm-deb-paper.pdf, 'Generated by
Gemini Notebook') solves temporal credit assignment under delayed
reward, which is not our problem. Our gun's failure is clause
saturation (~131 of 1740 literals included per clause -> conjunction
fires with probability ~2^-131 -> correction identically 0), and
TM-DEB only scales update FREQUENCY by gamma^dt, so it would leave
the fixed point untouched and additionally delete the long-range
feedback we need: gamma^90 = 4.4e-7 at the paper's own gamma=0.85,
and those are the shots whose lead matters most.

Credibility signals recorded in the doc: reference [2] misattributes
authors/venue/year; Algorithm Spec 2 steps 9-12 are literal '%'
placeholders so the automata update is simply absent; Table 1 is
titled 'Expected' and reports never-measured accuracies; Eq. 12 is
not a faithful copy of Granmo's Lemma 2.

The audit also produced the actionable result: a line-by-line diff of
Granmo Table 2/3 feedback against tmLearnOne, identifying why the
automata saturate - Type I never conditions on the clause output so it
omits Granmo's c=0, lk=1 -> -1 w.p. 1/s counter-force (true literals
ratchet toward Include), Type II's guard cOut==1 AND lits[lit]==0 is
unsatisfiable and therefore dead code, and the (T - clip(v,-T,T))/(2T)
resource allocation is missing entirely.
2026-09-20 22:57:28 +02:00
SirStone c034eb9d25 feat(testing): gun rack gauntlet + analysis reports
- fix(ModularBot): onBulletHitBot → onBulletHit (real hits were never tracked)
- feat(ModularBot): per-round gun stats dump to /tmp/gun_stats.jsonl
- feat(ModularBot): gun selection counter per round
- fix(tests): adversary paths _garage suffix removed from 7 test files
- feat(tests): test_gauntlet_5bots.nim — 10-round gauntlet vs all 5 adversaries
- feat(tests): analyze_gun_stats.nim — JSONL parser for gun performance tables
- docs: gun_rack_analysis.md — full per-gun performance report
- docs: gun_rack_summary.md — TL;DR verdict table (keep/drop/tune)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-09-20 21:14:12 +02:00
SirStone 1ed7797cb6 feat(ModularBot): 6 guns, pattern matcher, melee modules, adversarial bots
- New guns: guess-factor (GF histogram), pattern-matcher (movement tape replay)
- New modules: minimum-risk melee movement, spinning melee radar
- New test bots: PatternMover, RandomMover, WaveSurfer
- Fixed: FeedbackEvent now carries actualX/actualY for proper GF learning
- Fixed: TM gun warmup gating + directional residuals
- Fixed: circular gun integrated formula + multi-bin omega cache
- Fixed: oscillator wall-bounce lockout
- Fixed: phantom meteor perpendicular body orientation
- 6/6 battle wins across all enemy types
2026-09-20 00:59:53 +02:00
SirStone b509195ee9 chore: rename libs→common_libs, all bot dirs to _garage suffix, fix all path refs
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-27 18:18:41 +02:00
SirStone 8ba4bae21a chore: rename CLAUDE.md → AGENTS.md, move research doc to docs/ 2026-08-27 18:11:56 +02:00
SirStone d2d205e6f6 Merge branch 'research/ga-parameters' into research/goto-controller 2026-08-24 08:23:25 +02:00
SirStone 6241c41e52 Merge branch 'worktree-agent-a4c06ae3' into research/goto-controller 2026-08-24 08:23:19 +02:00
SirStone cb1bbc35dc docs(adr): Evo_Bot neuroevolution gun architecture (#63)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-08-24 00:02:47 +02:00
SirStone eae6fc15a2 research(Evo_Bot): GA/ES parameter recommendations for ~750-weight neuroevolution (#65)
Extracted concrete numbers from 13 papers in docs/papers/neuroevolution/.
Key findings: use CMA-ES or mutation-only truncation GA, mutate ALL weights
(not 5%), sigma=0.005-0.01, pop=64-200, no crossover, single elite.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-08-23 23:32:10 +02:00
SirStone 56e0b306c9 docs(research): goto controller algorithm for issue #20
Covers forward/reverse decision, proportional steering with speed-dependent
turn rate clamping, and deceleration using the existing getNewTargetSpeed util.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-17 17:04:57 +02:00
SirStone 64f73dd413 chore: add agent skills configuration
Add CLAUDE.md and docs/agents/ with issue tracker (Gitea), triage labels, and domain doc consumer rules.
2026-08-15 19:26:43 +02:00