Automata metrics: settledness alone does NOT separate learning from fidgeting
Added the four automata-level metrics to the TM diagnostics kit (settledness, clause diversity, churn, vote disagreement) plus a state histogram, a per-input confidence table and a one-line health summary, and validated them on a learnable-vs-noise pair. STATE CONVENTIONS, read off OUR code rather than from memory: range [-nStates, nStates] as int16; nStates = 64 for tm_pattern, 32 for tsetlin initial value 0 = the Exclude boundary INCLUDE iff state > 0; EXCLUDE iff state <= 0 flip boundary sits between state 0 and 1; commitment = abs(st)/nStates in [0,1] === THE GATE, AND A RESULT THAT MATTERS === Case A (learnable planted rule) vs Case B (shuffled labels), 49 bits, N=64: metric A (learnable) B (shuffled) settledness mean 0.970 0.719 churn flip/sample 0.000055 FALLING 0.000788 FLAT clause-change/sample 0.00263 falling 0.0595 flat diversity (Jaccard) 0.176 0.014 disagreement 0.003 0.298 verdict settling mixed (NOT settling) **SETTLEDNESS ALONE DOES NOT WORK.** On noise the automata still COMMIT (0.719) - they just commit to the wrong thing. The decisive separators are **churn TREND (falling vs flat)** and **vote DISAGREEMENT (0.003 vs 0.298)**. Had we built only the settledness metric - the one that seems most obvious - we would have been misled. That is now recorded in the README. INERTIA SWEEP: A vs B separate at N=16/32/64/128. **Raising N raises A's commitment but does NOT reduce B's noise-fitting** - so more inertia does not rescue a noise-fitting TM. === REAL READING ON THE SHIPPED GUN, AND THE INFERENCE IT SUPPORTS === tm_pattern GF head over the DrussGT fixtures: settledness 0.484 (settling), diversity 0.267 (moderate), churn 0.094/100 FALLING, disagreement 0.145 (coherent). **VERDICT: SETTLING** - not fidgeting, not collapsed. Constant inputs flagged: 38/39 (the known never-written bits) plus 19/36/37. Context: pooled warm accuracy 35.72% vs 34.24% majority = +1.48pp. So: **the old gun was NOT failing because of inertia or instability - it settled properly and its settled rules still barely beat a lazy guess.** Its settledness (0.484) is LOWER than both synthetic cases (0.97/0.72), which is the signature of WEAK OR CONFLICTING SIGNAL rather than too much inertia. CONCLUSION: **N and s are not the observed bottleneck. The target/representation is.** That is exactly why the new design changes the target and the label pipeline rather than sweeping knobs - and it means we should NOT spend effort on an N/s sweep expecting it to fix anything. Also adds `diag_automata_validation.nim` (Case A/B/C + inertia sweep) and `test_tm_automata_diag.nim` (55 pure checks); `test_tm_diag` 48 and `diag_synthetic` 17 still pass, plus all other guards. acceptance_offline_vs_online was NOT run (it needs a live battle and there is no tm_diag dependency). Caveat: churn on the real gun is a PROXY (a tm_core retrain over captured samples in live order) because the live gun exposes no per-sample state trace; the other metrics are read directly off the exported teams.
This commit is contained in:
@@ -16,7 +16,7 @@ Everything is pure and offline — no battles, no Java, no harness.
|
||||
|---|---|
|
||||
| `feature_spec.nim` | `FeatureSpec`, `describe`, `describeClause`, `draftTMSpec()` (49-bit draft), `tmPatternSpec()` (40-bit shipped encoding) |
|
||||
| `tm_core.nim` | compact deterministic Granmo Table 2/3 multiclass TM (mirrors the tm_pattern core), introspectable clause layout |
|
||||
| `diagnostics.nim` | the six groups (re-exports the two above) |
|
||||
| `diagnostics.nim` | the seven groups (re-exports the two above) |
|
||||
|
||||
Import everything with:
|
||||
|
||||
@@ -93,6 +93,143 @@ let rs = ablateScrambleFeature(tmplMachine, train, eval, spec, bit)
|
||||
input earn its bits"). Pass `baselineAcc` from `ablateBaseline` to avoid
|
||||
recomputing it per block.
|
||||
|
||||
## Task 5 — the AUTOMATA level (settledness / diversity / churn / disagreement)
|
||||
|
||||
Task 2 reads the clauses. Task 5 reads the **automata inside them**: one
|
||||
automaton per input bit per clause, each holding a state that says how
|
||||
confident it is that its bit belongs in the clause. These are what tell us
|
||||
whether the TM is locking onto the enemy or just fidgeting, and whether the
|
||||
inertia `N` (the number of automata states) should go up or down.
|
||||
|
||||
### State conventions — from the ACTUAL code, not the textbook
|
||||
|
||||
Derived from `tm_core.nim` and `guns/tm_pattern.nim` (`tmEval`, `tmLearnDir`,
|
||||
`tmNewTeam`, `resetMachine`):
|
||||
|
||||
| property | value |
|
||||
|---|---|
|
||||
| state range | `[-nStates, nStates]` (int16); `nStates` = 64 for `tm_pattern`, 32 for `tsetlin.nim` |
|
||||
| initial value | `0` (both cores call it "the Exclude boundary") |
|
||||
| INCLUDE | `state > 0` (`tmEval` / `clauseLits`) |
|
||||
| EXCLUDE | `state <= 0` |
|
||||
| flip boundary | BETWEEN state `0` and state `1` — the middle of the range |
|
||||
| `commitment(st)` | `abs(st) / nStates` in `[0,1]`: 0 on the boundary, 1 at either extreme |
|
||||
|
||||
So a flip is exactly a change in the predicate `state > 0`.
|
||||
|
||||
### The four metrics
|
||||
|
||||
1. **SETTLEDNESS** — per clause and overall, the mean `commitment` and the
|
||||
fraction of automata settled at/above a threshold (default `0.5`). High and
|
||||
**rising** as training proceeds is healthy; low means the clause is wavering
|
||||
noise. `settledness(m, threshold)`, `settlednessTrend(early, late)`.
|
||||
2. **CLAUSE DIVERSITY** — mean pairwise **Jaccard** of the included-literal sets,
|
||||
within the same polarity. Low-to-moderate is healthy; ~1.0 = all clauses are
|
||||
one rule in 50 hats; ~0 = memorising ticks. Empty clauses carry no rule and
|
||||
are skipped by default. `clauseDiversity(m, skipEmpty = true)`, `jaccard(a,b)`.
|
||||
3. **CHURN** — both levels, per training sample: the fraction of **automata**
|
||||
crossing the flip boundary, and the fraction of **clauses** whose
|
||||
included-literal set changed. High early and **falling** is healthy; flat-high
|
||||
= fidgeting; zero from the start = never learned. `churnTrace(tmpl, samples,
|
||||
epochs, seed, window, shuffle)`. `shuffle = false` replays the given
|
||||
(temporal) order, mirroring a live gun.
|
||||
4. **VOTE DISAGREEMENT** — per class, over the samples where it casts a vote:
|
||||
the fraction of firing clauses whose polarity disagrees with the sign of the
|
||||
class's total vote. Low is healthy. `voteDisagreement(m, samples)`.
|
||||
|
||||
### The three readouts that make it readable
|
||||
|
||||
- **state histogram** — the distribution of automata states across the range:
|
||||
`stateHistogram(m, nBins)`, `histogramText(h)`.
|
||||
- **per-input confidence table** — for each bit, the mean commitment of its
|
||||
automata (and a `constant` flag for zero-variance inputs):
|
||||
`perInputConfidence(m, spec, samples, threshold)`,
|
||||
`rankedInputConfidence(conf)`.
|
||||
- **one-line health summary** — `healthLine(ad)` gives
|
||||
`settledness / diversity / churn trend / disagreement`, each with its own
|
||||
verdict word (`settling`, `fidgeting`, `frozen`, `coherent`, ...), and
|
||||
`automataVerdict(ad)` reduces the trajectory to one word. `formatAutomataReport(ad)`
|
||||
prints everything.
|
||||
|
||||
### API
|
||||
|
||||
```nim
|
||||
import tm_diag/diagnostics
|
||||
|
||||
let ad = automataDiagnostics(tmpl, samples, spec,
|
||||
epochs = 15, seed = 777,
|
||||
settleThreshold = 0.5, nHistBins = 9,
|
||||
window = 100, measureChurn = true)
|
||||
# ad.machine, ad.settledness, ad.diversity, ad.churn, ad.disagreement,
|
||||
# ad.histogram, ad.inputConfidence, ad.summary
|
||||
echo ad.summary # one line
|
||||
echo automataVerdict(ad) # "settling" | "fidgeting" | "collapsed" | ...
|
||||
echo formatAutomataReport(ad) # the full readout
|
||||
```
|
||||
|
||||
For a machine you already have (e.g. the shipped gun's `exportTeams()`), call
|
||||
`settledness` / `clauseDiversity` / `stateHistogram` / `perInputConfidence` /
|
||||
`voteDisagreement` directly and `churnTrace` on a fresh copy.
|
||||
|
||||
### Validation — Case A (learnable) vs Case B (noise)
|
||||
|
||||
`common_libs/tests/diag_automata_validation.nim` uses the planted rule from
|
||||
`diag_synthetic.nim` and the SAME inputs with shuffled labels. Measured
|
||||
(`nBits=49`, 3 classes, 40 clauses, `nStates=64`, 3000 samples, 15 epochs):
|
||||
|
||||
| metric | Case A (learnable) | Case B (noise) |
|
||||
|---|---|---|
|
||||
| settledness mean | **0.970** | 0.719 |
|
||||
| settled fraction | 0.999 | 0.799 |
|
||||
| churn trend | **falling** (`2.5e-4` -> `2e-6`) | **flat** (`1.0e-3` -> `7.2e-4`) |
|
||||
| clause-change trend | falling (`0.011` -> `2.3e-4`) | flat (`0.069` -> `0.057`) |
|
||||
| diversity (overall Jaccard) | 0.176 | 0.014 |
|
||||
| disagreement | **0.003** | **0.298** |
|
||||
| verdict | settling | mixed (not settling) |
|
||||
|
||||
Case A settledness RISES `0.475` (early prefix) -> `0.970` (full). The pair
|
||||
**separates learning from fidgeting**: the decisive signals are the churn trend
|
||||
(falling vs flat) and disagreement (0.003 vs 0.298). Settledness alone is NOT
|
||||
enough — on noise the automata still commit (0.719), just to the wrong thing.
|
||||
Case C (a forced-constant bit) is flagged: `perInputConfidence(...).constant`
|
||||
and `constantInputs` both surface it. Note a constant bit can show HIGH
|
||||
commitment (one literal is always 1), so the `constant` flag is what
|
||||
disambiguates.
|
||||
|
||||
### Inertia sweep (`N` = nStates, 10 epochs)
|
||||
|
||||
The validation also sweeps `N` to see whether inertia moves the metrics:
|
||||
|
||||
| N | Case A settled | A churn | A dis. | Case B settled | B churn | B dis. |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 16 | 0.904 | falling | 0.003 | 0.720 | flat | 0.363 |
|
||||
| 32 | 0.943 | falling | 0.000 | 0.700 | flat | 0.388 |
|
||||
| 64 | 0.970 | falling | 0.003 | 0.715 | flat | 0.368 |
|
||||
| 128 | 0.971 | falling | 0.000 | 0.660 | flat | 0.304 |
|
||||
|
||||
A and B are separated at EVERY N (churn falling vs flat, disagreement low vs
|
||||
high). Raising N only raises Case A's commitment (0.90 -> 0.97); it does NOT
|
||||
reduce noise-fitting in Case B. So **inertia is not the discriminator** the
|
||||
metrics identify — the churn trend and disagreement are.
|
||||
|
||||
### Real reading — the shipped `tm_pattern` GF head
|
||||
|
||||
`diag_tm_pattern_offline.nim` over the committed DrussGT fixtures (automata read
|
||||
directly off the exported teams from `tr_drussgt_vs_modularbot`; churn measured
|
||||
on a `tm_core` temporal one-pass proxy, live-order):
|
||||
|
||||
- **settledness 0.484** (`settling`), settled fraction 0.453
|
||||
- **diversity 0.267** (`moderate`) — positive 0.170, negative 0.322
|
||||
- **churn 0.094/100 per sample, FALLING** (0.00208 -> 0.00048); clause-change 5.31/100, falling
|
||||
- **disagreement 0.145** (`coherent`)
|
||||
- **verdict: settling** — not fidgeting, not collapsed
|
||||
- constant inputs flagged: bits 38/39 (the known never-written ones) plus 19/36/37 in this fixture
|
||||
- context: pooled warm accuracy 35.72% vs the 34.24% majority = **+1.48pp**
|
||||
|
||||
The gun settles onto the within-battle labels but its settled rules barely beat
|
||||
the majority class, which points at the TARGET / representation rather than the
|
||||
inertia `N`. See the Task 5 report for the inertia discussion.
|
||||
|
||||
## The default-off real-gun hook
|
||||
|
||||
`guns/tm_pattern.nim` gained only additive, default-off instrumentation:
|
||||
@@ -116,7 +253,9 @@ let m = machineFromTeams(TM_NBITS, TM_CLASSES, TM_NCLAUSES, TM_NSTATES, TM_S,
|
||||
|
||||
```sh
|
||||
nim c -r -d:release --path:common_libs common_libs/tests/test_tm_diag.nim # 48 pure unit checks
|
||||
nim c -r -d:release --path:common_libs common_libs/tests/diag_synthetic.nim # Task 3 proof
|
||||
nim c -r --path:common_libs common_libs/tests/test_tm_automata_diag.nim # 55 automata-metric unit checks
|
||||
nim c -r -d:release --path:common_libs common_libs/tests/diag_synthetic.nim # Task 3 proof (17 checks)
|
||||
nim c -r -d:release --path:common_libs common_libs/tests/diag_automata_validation.nim # Task 5 A/B/C proof
|
||||
nim c -r -d:release --path:common_libs common_libs/tests/diag_tm_pattern_offline.nim # Task 4 real reading
|
||||
```
|
||||
|
||||
|
||||
Reference in New Issue
Block a user