6 Commits

Author SHA1 Message Date
SirStone f9f8d84671 TM diagnostics kit: VALIDATED (finds a known dead input), and it found a real bug
Built `common_libs/tm_diag/` as a first-class offline diagnostics kit for Tsetlin
work, BEFORE writing the new gun - because we hit two data problems tonight that no
amount of reading the TM's clauses would have revealed (a 38.8% majority answer,
and 36-58% mislabelled training samples).

WHAT IT PROVIDES
- `feature_spec.nim`: a NAMED feature container, so a learned clause prints as a
  sentence (`IF near-wall AND bullet-dead-on AND turn-left(t-2) THEN class=3`)
  instead of "feature 17". Includes the 49-bit draft spec from the design session
  and the shipped 40-bit encoding.
- `tm_core.nim`: a compact deterministic Granmo multiclass TM with an
  INTROSPECTABLE clause layout (mirrors the tm_pattern core).
- `diagnostics.nim`, six groups: (1) pre-flight DATA checks + shuffled-label
  control, (2) clause introspection (readable dump, per-clause vote counts, empty
  and never-fired clauses, length distribution, per-class balance), (3)
  per-feature contribution with an explicit DEAD-INPUT LIST and a ranked
  most-valuable list, (4) accuracy vs the majority baseline with per-class
  precision/recall and pred-majority share, (5) learning curve, (6) ablation hooks
  (drop a block / scramble a bit).

=== TASK 3: THE VALIDATION THAT GATES EVERYTHING - PASSED WITH NUMBERS ===
A diagnostic we never checked is worthless, so the kit was tested on a synthetic
set with a PLANTED RULE (class2 = A and B, class1 = A and not B, class0 = not A),
a deliberately IRRELEVANT block (US, 9 bits) and a PURE-NOISE bit (17).
- majority baseline 60.63% (class0); over-30% correctly flagged
- **the planted rule is recovered EXACTLY** via `necessaryLiterals`:
    class0 IF NOT dist-wall<50 | class1 IF dist-wall<50 AND NOT lat DEAD-ON |
    class2 IF dist-wall<50 AND lat DEAD-ON
- **DEAD-INPUT LIST = all 9 US bits AND the noise bit 17**, while the planted bits
  0 and 45 are correctly NOT listed
- top contributors: bit0 w=1241.7, bit45 w=583.3, then 49.8 - a 12-25x gap, so the
  relevant bits are unmistakable
- **ABLATION: drop WALLS -39.47pp, drop BULLETS -19.33pp, drop US 0.00pp**,
  scramble A -42.00pp, scramble the noise bit 0.00pp
- shuffled-label control 60.40% vs majority 60.63% = -0.23pp -> no leak
So the kit reliably finds a known dead input and a known relevant one.

=== TASK 4: THE REAL READING, AND A BUG IN THE SHIPPED GUN ===
`tm_pattern` GF head, 6 DrussGT fixtures, pooled 250,745 samples:
- label balance c2 = **34.4%** (majority-heavy, flagged); accuracy **35.72%** vs
  majority **34.24%** -> margin **+1.48pp**. On `tr_drussgt_vs_crazy` it is BELOW
  majority (33.81% vs 37.72%, -3.92pp).
- 200 clauses: **27 empty, 45 never fired**, mean length 19.17, max 57. The
  majority class is starved (class2: 22 non-empty, 18 empty, only 2 positive
  fired). Class4 fires 11-24-literal clauses -> memorisation signature.
- **REPRESENTATION BUG FOUND (reported, not silently fixed):** `tmBuildBits` writes
  only 38 raw bits into `var bits: array[TM_NBITS=40, uint8]` - bits 38 and 39 are
  NEVER ASSIGNED, so they are always 0 and their negated literals are always 1.
  The kit's `constantInputs` confirms 38/39 are constant, and **`UNUSED-38`/`39`
  rank #6 and #8 in the most-valuable-inputs list** - i.e. the model's
  highest-usage inputs are information-free. That is a representation bug, not a
  display artefact, and it is a concrete mechanism for part of the poor learning.

DRAFT ENCODING CHECKED: the 49-bit draft is arithmetically consistent
(4+4=8 walls, 6+3=9 us, 3+5+3+3+3+3=20 motion, 5+7=12 bullets = 49). No draft
inconsistency.

Guards: test_tm_diag 48 (new), diag_synthetic 17 (new), test_gun_harness 39,
test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41,
test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 40,
test_rack_membership 48, test_selector_tiebreak 19, test_tm_pattern_registration 20,
test_vbullet_admit_gate 12, acceptance_offline_vs_online 12/12; tm_pattern_learning
passes. The tm_pattern hook is additive and default-OFF (no behaviour change).

NOT YET INCLUDED (the automata metrics discussed for the next step): per-clause
automata settledness (distance from the flip point), clause diversity (pairwise
overlap), literal-set churn over time, and cross-clause vote disagreement. The kit
has clause-level diagnostics but not the automata-state ones.
2026-09-22 21:24:27 +02:00
SirStone b0654d18eb TM verdict, settled: it loses LIVE and sits at/below its majority class - (c)
The user pushed back on "the TM can't be your best 1v1 gun", correctly, because two
decisive tests had never been run. Both are now run and they agree.

TASK 1 - THE GF HEAD vs ITS MAJORITY-CLASS BASELINE (offline, n=1,751,067):
  label histogram [254286, 284578, 678879, 297055, 236269]
  majority class = 2 (the CENTRE bucket) = 38.77%
  RAW head accuracy = 36.69%  ->  margin **-2.08 pp, BELOW majority**
  GATED head accuracy = 40.37% vs 38.75% majority -> +1.62 pp, BUT it predicts the
  majority class on 62.4% of ticks and its minority recall is 13.6% / 12.9% - a
  base-rate predictor wearing a classifier's clothes.
  Shuffled control sits at its own majority (20.04% vs 20.12%), confirming chance.
**THE OLD "46% vs 20% CHANCE" FIGURE I QUOTED WAS WRONG ON TWO COUNTS:** the
baseline is 38.8%, not 20%, and the 46% predated the deferred-label fix. Against
the correct baseline the head is BELOW it.

TASK 2 - THE FIRST-EVER LIVE A/B OF THE TM GUN (7 runs x 7 rounds per arm, one
frozen binary from git archive HEAD = eb74f9b2, sha256 cb66d66b..., real DrussGT,
every arm forced alone with TR_RACK_<GUN>=both and all 14 others off, liveness
confirmed per run):
  arm                     shots   real %   dmg/run   round wins
  onlyPattern              4610   10.74%     285      25/49
  onlyTMPATTERN (radial)   3374    3.50%      71       0/49
  onlyLinear               3218    3.23%      61       0/49
  Pattern vs TM:  +7.22 pp / +213.7 dmg, exact p=0.0006
  TM vs Linear:   +0.30 pp, p=0.659  (dmg p=0.438)
**The TM is statistically INDISTINGUISHABLE from its own Linear base live.** So it
is not "the TM works and we are aiming it wrong".

DIRECT ANSWER: **(c) It loses live AND sits at/below majority - the target carries
no learnable signal beyond the base rate, and that is the reason.** The reason is
not the machine, not the knobs, and not the application alone: the thing it was
asked to predict is dominated by the modal answer.

This closes the TM-as-gun thread. If a TM is wanted in the bot, a firing gate or a
movement decision is a better fit for a boolean-rule classifier than an aim point -
that is untested and is a different project.

A LIVE GF-MODE ARM WAS NOT RUN (stated as unmeasured): the task pinned one frozen
HEAD binary and HEAD registers the TM gun as radial only; Task 1 already makes GF
the unpromising candidate.

HARNESS FIX WORTH KEEPING: `tools/ab/which_gun_arm_env.sh` left the TARGET gun
unset, so with the now-Pattern-only default it silently fell back to the FULL rack
- an arm could appear to test a single gun while actually running the whole rack.
It now emits `TR_RACK_<GUN>=both` for the target and `=off` for all 14 others.
(Earlier which-gun results are unaffected: they ran before the Pattern-only default,
or - as in the melee/1v1 campaign - set the explicit `=both` themselves.)

tm_pattern.nim gains a per-class confusion matrix (warm samples only) to support the
majority baseline; no behaviour change. Adds Round 4 to
tm_pattern_sweep_results.md with both tasks and the interpretation rule.
2026-09-22 08:21:49 +02:00
SirStone 9cd6e9b8ce Ablation: the radial TM is replaceable by a CONSTANT, and its avenue is dead on the
shipped metric

The radial TM beats Linear on bmPoint, but its head never beat the majority
baseline after the label bias was fixed - suggesting the win is a constant lean
rather than learning. So: sweep a stateless constant short-range offset (new
`common_libs/guns/radial_offset.nim`, no learning at all) against the learned TM.

VERDICT (measured, offline range, seeds=3, 18 paired runs, 231 TM rounds):
1. **bmPoint - REPLACE the TM with a constant.** `RO_s0.95` (aim distance x0.95)
   TIES it early (9/9, p=1.0) and BEATS it overall (15/3, p=0.0075; 7.47% vs
   6.89% per-run mean). A fixed -20px does the same. The head never beats its
   majority baseline (56.2% vs 57.2%).
2. **bmPath (the SHIPPED metric) - the radial avenue is a DEAD END.** TMRadial is
   a systematic LOSS there (2/16, p=0.0013); every constant is within +-0.2pp; the
   only real bmPath effect is the BotRadius clamp. So the radial shift cannot help
   the shipped configuration.
3. The per-adversary optimum DOES vary (fixed -10 for crazy, -30 for tr_crazy,
   scale 0.95 for three others) - but ONE GLOBAL CONSTANT still beats the
   adaptively-trained head, so the "fragility justifies learning" argument FAILS.

THE REAL FINDING UNDERNEATH, and it generalises beyond this gun: the base linear
prediction systematically OVERSHOOTS. Measured raw per-tick base radial error has
mean -71 to -100 px; the enemy is NEARER than the prediction in 63-81% of shots
and farther in only 4-14%, CONSISTENT ACROSS ALL SIX CAPTURES. Radial label
histogram [415166,126461,120325,44694,19351] = 57.2% majority class, mean label
-82.3 px, mean applied shift -37.6 px. So the net-short bias is a GENUINE property
of these range-holders against a constant-velocity extrapolation (they decelerate
and turn, so the true position is closer than the straight-line guess) - NOT a
fixture artefact. That is worth chasing for the guns that actually ship.

Caveat: bmPoint is not the shipped metric (bmPath won the real-hit-rate A/B for
SELECTION), so a bmPoint win is not yet evidence of a real win. That needs a live
test - and the natural target is Pattern, which is now the default and best gun.

Adds radial_offset.nim + sweep_radial_offset.nim; tm_pattern.nim gains additive
instrumentation only (radial label mean and applied-shift mean; no behaviour
change, and test_tm_pattern_registration still passes all 20 checks).
2026-09-22 02:08:38 +02:00
SirStone 589a230106 TM radial gun: registered (default OFF) + label-bias fix that removes the bias but
retracts its own earlier learning claim

=== TASK 1: REGISTERED AS GUN 14, DEFAULT `off` ===
The radial TM gun is now a first-class rack member (`TMPATTERN`, id 14), forceable
alone with `TR_RACK_TMPATTERN=both` plus every other `TR_RACK_*=off`.
DEFAULT IS `off`, and the justification matters: `both` would let it compete for
selection AND (because the shared VirtualTracker ring is order-sensitive) shift
every other gun's learning order, so it CANNOT leave the default path unchanged.
With `off` its predict and spawnBullets are additionally GATED on rack admission
(the only gun wired that way), so the shipped default never spawns it at all:
zero cost, zero ring perturbation.
Live proof: 1-round battle with only TMPATTERN racked ->
  `gun 14 (TMPattern): vShots=400 selected=104 other-gun selections=0`.
Default-path-unchanged proof: parity checks that the 15-gun default bestGun/
selectGun equals the old 14-gun rack RNG-draw-for-RNG-draw, that gun 14 is never
selected by default, and acceptance 12/12.
Cost: 0.36 ms/tick (predict 0.30 + onResult 0.05) ~= 3% of the 13.16 ms budget.
Tsetlin in the same harness is 1.62 ms/tick, so the new gun is ~4.5x cheaper.

=== TASK 2: THE LABEL-BIAS FIX - AND A RETRACTION ===
Root cause confirmed: under bmPoint a SHORT radial correction resolves the virtual
bullet BEFORE the base arrival tick, so the label was dropped (labelMisses).
Fix: defer the label in a pending queue and flush it once the arrival tick is
recorded; labels still come from the BASE arrival tick.
  labelMisses        4,281,695  ->  0
  training samples   1,071,824  ->  5,345,847  (x5)
  radial head acc         48.8% ->  57.0%   (shuffled control 20.0%)
  bmPoint hit rate     9.4/5.8% ->  9.1/5.7%  (unchanged, within noise)
So the fix IMPROVES LEARNING but NOT the metric.

**RETRACTION OF THE PREVIOUS JOB'S CLAIM.** It reported the radial head's 48.8%
against a 36.7% majority baseline and concluded "conditional learning, not a
constant bias". With the bias removed, the correctly-measured majority baseline is
**58.2%** - so the head at 57.0% is AT/BELOW majority. The earlier apparent
conditional learning was PARTLY AN ARTEFACT OF THE BIASED SAMPLE. The bmPoint
metric win is real (TMRadial > Linear early 16/2 p=0.0013, overall 18/0 p<0.0001;
> shuffled 18/0 p<0.0001) but it comes from a NET-POSITIVE AVERAGE RADIAL SHIFT,
not from beating a majority classifier. Recorded plainly rather than left standing.

Guards: test_tm_pattern_registration 20 (new), test_tm_pattern_rack_live 4 (new),
test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3 (the SIGSEGV is
gone - the knn_gun rewrite is now committed), test_adaptive_radar 41,
test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 28,
test_rack_membership 38, test_selector_tiebreak 19, test_tm_pattern_learning 3,
acceptance_offline_vs_online 12/12. ModularBot compiles (release).

Note: `common_libs/tests/range_guns.nim` still builds 14 offline drivers (the
offline sweep constructs TmPatternGun directly and acceptance only inspects ids
0..13), so nothing breaks - but a future job wanting it in the offline rack must
add a 15th driver and mirror the live admission gating. gun_stats.jsonl now emits
15 rows; downstream tooling should ignore id 14.
2026-09-22 01:58:33 +02:00
SirStone 1ea72c7f14 TM gun round 2: base was never behind; RADIAL target beats Linear on bmPoint
=== TASK 1: MY PREMISE WAS REFUTED ===
I instructed the job to "fix the baseline" because an earlier measurement said the
TM gun's base did not iterate flight time like `LinearGun`. MEASURED: the new gun's
base is BYTE-FOR-BYTE `LinearGun` - 18/18 runs tie exactly, p=1.000, every per-run
row byte-identical. The "non-iterating baseline" belonged to the OLD `tsetlin.nim`,
not this gun. So no fix was needed, and the earlier inference should not have been
generalised to the new gun. (It did still align the zero-correction clamp to
LinearGun's exact [0, arena] range, and reports the old BotRadius-inset base was a
wash/marginally better at 34.2%/24.7%.)

=== TASK 2: THE RADIAL TARGET - A CONTROL-VALIDATED WIN, BUT ONLY ON bmPoint ===
Instead of the lateral (GF-bucket) component - which the linear lead already
captures - the TM now predicts the RADIAL component: will the enemy be nearer or
farther than the base prediction when our bullet arrives? A 5-class radial head
sharing the same 40-bit context and TM core; the readout advances/retards the aim
distance along the base bearing.

  under bmPath (the SHIPPED metric): STRUCTURAL NO-OP
    synthetic 8/8 exact ties, p=1.0; real 33.9%/24.1% vs Linear 34.0%/24.3%
  under bmPoint: A WIN, control-validated
    TMRadial 9.4% (6013/63785) / 5.8% (42079/726652)
    Linear   7.2% / 4.7%          overall 17/1, p=0.0001
    Tsetlin  7.0% / 4.8%          overall 15/3, p=0.0075
    shuffled 7.0% / 3.6%          early 17/1 p=0.0001; overall 18/0, p<0.0001
  radial head online accuracy 48.8% vs 19.9% shuffled chance and 36.7% majority
  -> it is CONDITIONAL learning, not a constant short-range bias.
Best config: TM_RADIAL_RANGE=60, TM_RAD_MARGIN=0.25, 5 classes.

CAVEAT THAT MATTERS: a win on `bmPoint` is NOT yet evidence of a real win. `bmPath`
is the shipped SELECTION metric precisely because it beat `bmPoint` on real hit
rate (7.43% vs 4.70%). But that A/B was about which gun to PICK, not about gun
QUALITY - a gun can be better in reality while scoring worse on the selection
metric. So this needs a LIVE test, and it is the decisive one.

=== TASK 3: REVERSAL TARGET - CLEAN NEGATIVE ===
The label positive rate is only 9.7% (rev=[24772,2673]) and the head's 86.8%
accuracy is BELOW the 90.3% majority baseline: it does not learn the positive
class at all. Hit-rate effect neutral (bmPath 19.5%/18.4% vs shuffled 19.1%/17.8%,
p=0.24/0.82). Dropped.

=== OVERALL ===
Not competitive on the shipped bmPath metric (gated GF 28.3%/22.2% vs Linear
34.0%/24.3%, p=0.0075). Better than Linear on bmPoint via TMRadial (+2.2pp early,
+1.1pp overall). Per-enemy reset exists; a fresh gun per round; NO cross-battle
persistence (the user's non-negotiable).

MEASURED LIMITATION: radial mode has a high labelMiss because aiming short
resolves BEFORE the base arrival tick, biasing training toward resolvable samples.
The metric win is label-independent. A deferred-label fix is the next refinement.
INFERRED: the mechanism is surfers being NEARER than the base prediction
(range-holding); a constant-short-offset ablation would separate a learned
short-range bias from genuine per-tick conditional prediction.
2026-09-22 01:27:59 +02:00
SirStone ca82053a11 TM gun: the discrete-target diagnosis was RIGHT - it learns now. Still loses to Linear.
The user's goal: a TM gun that is the best 1v1 gun, starting from scratch every
battle but quickly overfitting the current enemy. The previous attempt (knob
tuning) failed: NO configuration beat its own shuffled-feedback control, and the
TM-off ablation scored the same as TM-on, i.e. the TM's correction was
near-zero-mean noise. Diagnosis then: a Tsetlin Machine is a CLASSIFIER, and we
were asking it for an absolute aim point - a regression target. So this attempt
gave it a DISCRETE target (multi-class over guess-factor buckets) with 40
binary/bucketed motion features, and measured it against Linear, the default
Tsetlin gun, and a MANDATORY shuffled control.

THE DIAGNOSIS IS CONFIRMED - THE TM LEARNS, DECISIVELY:
  online class accuracy     46.0%  vs shuffled control 20.0%   (2.3x chance)
  raw ungated argmax        21.2%/18.6% vs shuffled 15.2%/8.3% (18/18, p<0.0001)
  TMPattern > its shuffled control, overall   17/1 runs, p=0.0001
Compare the previous attempt, which could not beat shuffled feedback at all.
TMPattern also beats the default Tsetlin gun early (17/1, p=0.0001), so it is a
strictly better TM gun than the one in the rack.

BUT IT IS NOT COMPETITIVE WITH LINEAR ON REAL SURFERS:
  real DrussGT, bmPath (the shipped metric), 3 seeds, pooled early/overall
    Linear            34.0% (6358/18715)    24.3% (58297/239943)
    TMPattern (gated) 27.9% (15514/55535)   22.0% (158658/719681)
    TMPatternShuf     28.7%                 19.4%
  Linear > TMPattern: 15/18 early p=0.0075, 15/18 overall p=0.0075
  bmPoint: neutral (7.2%/4.6% vs Linear 7.2%/4.7%)
  synthetic controlled motion: matches/edges Linear (66.8%/60.6% vs 66.4%/59.6%,
    shuffled 55.7%/50.1%) - the mechanism works when motion is predictable.

So: the representation fix moved this from "learns nothing" to "learns strongly
but applies its knowledge badly". INFERRED reason for the residual loss: the
linear lead is already the modal GF bucket (the label histogram is centred), so
corrective excursions away from it are net-negative. The measured deficit lives
in the BASELINE and in RANGE, not in the TM knobs - which is why further knob
tuning was never going to work.

Best config: gated hard K=5, TM_CONF_MARGIN=0.25, TM_SHRINK=0.5.
NOT TRIED (time-boxed): the binary-reversal target, and a RADIAL (range-holding)
target - the latter is the top next step.

Adds `common_libs/guns/tm_pattern.nim` (NOT registered in the rack),
`common_libs/tests/sweep_tm_pattern.nim`, and a durable writeup at
`common_libs/tests/tm_pattern_sweep_results.md`.
2026-09-22 00:56:01 +02:00