Automata metrics: settledness alone does NOT separate learning from fidgeting

Added the four automata-level metrics to the TM diagnostics kit (settledness,
clause diversity, churn, vote disagreement) plus a state histogram, a per-input
confidence table and a one-line health summary, and validated them on a
learnable-vs-noise pair.

STATE CONVENTIONS, read off OUR code rather than from memory:
  range [-nStates, nStates] as int16; nStates = 64 for tm_pattern, 32 for tsetlin
  initial value 0 = the Exclude boundary
  INCLUDE iff state > 0; EXCLUDE iff state <= 0
  flip boundary sits between state 0 and 1; commitment = abs(st)/nStates in [0,1]

=== THE GATE, AND A RESULT THAT MATTERS ===
Case A (learnable planted rule) vs Case B (shuffled labels), 49 bits, N=64:
  metric                    A (learnable)     B (shuffled)
  settledness mean              0.970            0.719
  churn flip/sample        0.000055 FALLING  0.000788 FLAT
  clause-change/sample       0.00263 falling   0.0595 flat
  diversity (Jaccard)           0.176            0.014
  disagreement                  0.003            0.298
  verdict                    settling        mixed (NOT settling)

**SETTLEDNESS ALONE DOES NOT WORK.** On noise the automata still COMMIT (0.719) -
they just commit to the wrong thing. The decisive separators are **churn TREND
(falling vs flat)** and **vote DISAGREEMENT (0.003 vs 0.298)**. Had we built only
the settledness metric - the one that seems most obvious - we would have been
misled. That is now recorded in the README.

INERTIA SWEEP: A vs B separate at N=16/32/64/128. **Raising N raises A's
commitment but does NOT reduce B's noise-fitting** - so more inertia does not
rescue a noise-fitting TM.

=== REAL READING ON THE SHIPPED GUN, AND THE INFERENCE IT SUPPORTS ===
tm_pattern GF head over the DrussGT fixtures: settledness 0.484 (settling),
diversity 0.267 (moderate), churn 0.094/100 FALLING, disagreement 0.145
(coherent). **VERDICT: SETTLING** - not fidgeting, not collapsed. Constant inputs
flagged: 38/39 (the known never-written bits) plus 19/36/37.
Context: pooled warm accuracy 35.72% vs 34.24% majority = +1.48pp.
So: **the old gun was NOT failing because of inertia or instability - it settled
properly and its settled rules still barely beat a lazy guess.** Its settledness
(0.484) is LOWER than both synthetic cases (0.97/0.72), which is the signature of
WEAK OR CONFLICTING SIGNAL rather than too much inertia.
CONCLUSION: **N and s are not the observed bottleneck. The target/representation
is.** That is exactly why the new design changes the target and the label
pipeline rather than sweeping knobs - and it means we should NOT spend effort on
an N/s sweep expecting it to fix anything.

Also adds `diag_automata_validation.nim` (Case A/B/C + inertia sweep) and
`test_tm_automata_diag.nim` (55 pure checks); `test_tm_diag` 48 and
`diag_synthetic` 17 still pass, plus all other guards. acceptance_offline_vs_online
was NOT run (it needs a live battle and there is no tm_diag dependency).

Caveat: churn on the real gun is a PROXY (a tm_core retrain over captured samples
in live order) because the live gun exposes no per-sample state trace; the other
metrics are read directly off the exported teams.
This commit is contained in:
2026-09-22 21:40:27 +02:00
parent f9f8d84671
commit ab8d383121
5 changed files with 1109 additions and 5 deletions
@@ -177,5 +177,32 @@ proc main() =
echo &"# class{cls} votes={c.votes:<7} len={c.length} {c.text}"
echo &"# -> necessary literals: {spec.describeClause(necessaryLiterals(m, bestSamples, cls), cls)}"
# ── AUTOMATA-LEVEL metrics on the shipped GF head (Task 5) ──
# settledness / diversity / histogram / per-input confidence / disagreement
# are read DIRECTLY off the exported teams; churn is measured on a tm_core
# retrain (same algorithm) over the captured samples, because the live gun
# does not expose a per-sample state trace.
var ad = AutomataDiag()
ad.machine = m
ad.settledness = settledness(m, 0.5)
ad.diversity = clauseDiversity(m)
ad.histogram = stateHistogram(m, 9)
ad.inputConfidence = perInputConfidence(m, spec, bestSamples, 0.5)
ad.disagreement = voteDisagreement(m, bestSamples)
let churnN = min(bestSamples.len, 10000)
var churnSamples = newSeq[DiagSample](churnN)
for i in 0..<churnN: churnSamples[i] = bestSamples[i]
let tmpl = newMachine(TM_NBITS, TM_CLASSES, TM_NCLAUSES, TM_NSTATES, TM_S,
seed = 1)
# Mirror the live gun: one TEMPORAL pass (no shuffle) as bullets resolve.
ad.churn = churnTrace(tmpl, churnSamples, epochs = 1, seed = 777,
window = 200, shuffle = false)
ad.summary = healthLine(ad)
echo "\n## AUTOMATA-LEVEL metrics (shipped tm_pattern GF head)"
echo "# churn measured on a tm_core retrain over ", churnN,
" captured samples, ONE temporal pass (live-order proxy)"
echo formatAutomataReport(ad)
echo "# VERDICT: ", automataVerdict(ad)
when isMainModule:
main()