9064377740
=== TASK 1: THE DIRECTION QUESTION, ANSWERED WITH A DEMONSTRATION ===
I told the user `s` controls clause length but refused to claim the DIRECTION,
because I had seen it described both ways. It is now read out of our own code -
one site per core, in the Type I branch of `tmLearnDir` (`guns/tm_pattern.nim:258`,
`guns/tsetlin.nim:224`, `tm_diag/tm_core.nim:110`):
if pol * d > 0.0:
if lits[lit] == 1:
if cOut == 1: if rand < (s-1)/s: st += 1 # toward Include, w.p. (s-1)/s
else: if rand < 1/s: st -= 1 # toward Exclude, w.p. 1/s
else: if rand < 1/s: st -= 1 # toward Exclude, w.p. 1/s
=> **HIGHER `s` GIVES LONGER CLAUSES.** The include step runs w.p. (s-1)/s
(rising with s); both exclude steps run w.p. 1/s (falling with s).
DEMONSTRATED (49-bit draft, planted 2-literal rule, 3000 train / 1500 eval):
s=1.0 len 1.51 acc 100% s=2.0 len 1.82 acc 100% s=5.0 len 3.02 acc 100%
s=1.5 len 1.42 acc 100% s=3.0 len 2.19 acc 100% s=10 len 4.04 acc 99.7%
s=20 len 5.17 acc 94.5%
WHY s=1.0 DEGENERATES: (s-1)/s = 0 so the include step NEVER fires while 1/s = 1
so BOTH exclude steps always fire - Type I can only remove literals, so a clause
can grow only through the Type II penalty. (On random labels that leaves 43/120
non-empty clauses vs 120/120 at s>=3.)
USABLE RANGE ~[1.5, 5]. tm_pattern uses 3.0; tsetlin uses 1.5.
=== TASK 4: WOULD `s` HELP THE SHIPPED GUN? NO - MEASURED ===
Recompiling the offline driver with -d:TM_S_DEF=<v> (source untouched) retrains
the gun end to end:
s mean len verdict warm acc margin vs majority
1.5 15.82 too long 32.22% -2.03pp
2.0 14.99 too long 34.03% -0.22pp
3.0* 19.17 too long 35.72% +1.48pp (*shipped)
5.0 20.47 too long 34.72% +0.47pp
Lowering `s` shrinks the clauses and makes accuracy WORSE; raising it pads them
and also loses. The shipped 3.0 is the best of the four, and **no value comes
near the healthy 3-8 band.** Combined with the settledness finding, the shape is
consistent with "NO CONSISTENT SHORT RULE EXISTS in this representation/target".
So the bottleneck is the SIGNAL - now confirmed from a THIRD independent angle
(settledness, churn trend, and clause shape). This is the measurement behind the
decision not to spend effort sweeping N or s.
=== TASK 3: AN HONEST CORRECTION TO MY OWN HYPOTHESIS ===
I predicted that random labels would produce `too long` clauses (the TM padding).
MEASURED: on this encoding noise reads as **short / `collapsed`** (mean 1.88,
median 2.0, acc 33.3%) - the TM FAILS TO COMMIT rather than padding. So "too
long" is not the noise signature, which means the shipped gun's 19.17 mean is not
explained by label noise. Worth knowing.
Adds diagnostic group 8: the clause-shape checker - full length distribution
(min/median/p10/p90/std), per-polarity and per-class breakdowns, a
`clauseShapeVerdict` against a parameterised healthy band (default 3-8),
per-BLOCK length contributions, and clause coverage (mean firing clauses,
effectiveClauses = participation ratio, top3Share). `healthLine` now appends
`shape=<mean> (<verdict>)`.
Validation: `test_tm_clause_shape` 66 checks. A planted 2-literal rule reads
`healthy` with the literals recovered exactly; per-block correctly names the
planted blocks (WALLS 43.0%, BULLETS 28.1%) and buries an irrelevant block (3.4%,
below its uniform 8.3% share); random labels read `collapsed`.
REAL READING, shipped gun: mean 19.17 / median 16.00 / p90 44.80 / max 57,
173 non-empty of 200, 27 empty => **`too long`**; coverage firing/sample 48.53
(24.3%), effectiveClauses 97.85/200, top3Share 5.3% (voting NOT concentrated);
per-block is diffuse with no dominator, EXCEPT **UNUSED 6.4%** - the always-true
negations of the never-written bits 38/39 acting as FREE PADDING, the same bug the
kit found earlier now visible as clause bloat.
Guards: test_tm_clause_shape 66 (new), test_tm_diag 48, test_tm_automata_diag 55,
diag_synthetic 17, diag_automata_validation 11, test_gun_harness 39,
test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41,
test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 40,
test_rack_membership 48, test_selector_tiebreak 19, test_tm_pattern_registration 20,
test_vbullet_admit_gate 12. acceptance_offline_vs_online not run (needs a live
battle; no tm_diag dependency).
305 lines
14 KiB
Nim
305 lines
14 KiB
Nim
## Pure unit + synthetic-validation tests for the CLAUSE-SHAPE checker (group 8).
|
|
##
|
|
## Covers:
|
|
## * percentile / lengthStats / clauseSummary distribution fields;
|
|
## * clauseShapeVerdict against the healthy band (boundaries + custom band);
|
|
## * per-block length contribution and coverage/concentration;
|
|
## * synthetic ground truth: a SHORT planted rule reads healthy and is
|
|
## recovered near its true length; RANDOM labels are reported as measured;
|
|
## the block containing the planted rule dominates and an irrelevant block
|
|
## contributes ~0.
|
|
##
|
|
## Run: nim c -r -d:release --path:common_libs common_libs/tests/test_tm_clause_shape.nim
|
|
|
|
import std/[random, strformat, strutils]
|
|
import tm_diag/diagnostics
|
|
|
|
var checks = 0
|
|
var failures = 0
|
|
proc check(name: string, ok: bool) =
|
|
inc checks
|
|
if ok: echo "PASS: ", name
|
|
else: echo "FAIL: ", name; inc failures
|
|
|
|
# ── percentile / lengthStats ─────────────────────────────────────────────────
|
|
|
|
proc testPercentile() =
|
|
let s = @[1, 2, 3, 4]
|
|
check "percentile p0 = min", percentile(s, 0.0) == 1.0
|
|
check "percentile p100 = max", percentile(s, 1.0) == 4.0
|
|
check "percentile median interpolates", abs(percentile(s, 0.5) - 2.5) < 1e-9
|
|
check "percentile p25", abs(percentile(s, 0.25) - 1.75) < 1e-9
|
|
check "percentile of empty = 0", percentile(@[], 0.5) == 0.0
|
|
check "percentile clamps p", percentile(s, 2.0) == 4.0
|
|
|
|
proc testLengthStats() =
|
|
let st = lengthStats(@[1, 2, 2, 3, 3, 3, 4])
|
|
check "lengthStats n", st.n == 7
|
|
check "lengthStats min/max", st.minL == 1 and st.maxL == 4
|
|
check "lengthStats mean = 18/7", abs(st.mean - 18.0 / 7.0) < 1e-9
|
|
check "lengthStats median = 3", abs(st.median - 3.0) < 1e-9
|
|
check "lengthStats p90", abs(st.p90 - 3.4) < 1e-9
|
|
check "lengthStats histogram", st.hist == @[0, 1, 2, 3, 1]
|
|
let empty = lengthStats(@[])
|
|
check "lengthStats empty is well-defined",
|
|
empty.n == 0 and empty.mean == 0.0 and empty.maxL == 0
|
|
|
|
# ── clauseSummary distribution / polarity / per-class ────────────────────────
|
|
|
|
proc tinySpec(): FeatureSpec =
|
|
var s = FeatureSpec()
|
|
s.addBlock("b0", 4, @["b0.0", "b0.1", "b0.2", "b0.3"])
|
|
s.addBlock("b1", 4, @["b1.0", "b1.1", "b1.2", "b1.3"])
|
|
s
|
|
|
|
proc shapeModel(): TmMachine =
|
|
## nBits=8, 2 classes, 4 clauses/class (half=2 pos, 2 neg), 16 literals.
|
|
## class0 pos clause0 = {+b0.0, +b0.1, +b1.0} (len 3, block0 x2 + block1 x1)
|
|
## class0 pos clause1 = {+b0.2} (len 1, block0)
|
|
## class0 neg clause2 = {+b1.1, +b1.2} (len 2, block1)
|
|
## everything else empty.
|
|
result = newMachine(8, 2, 4, 64, 3.0, 1)
|
|
let nl = result.nLiterals
|
|
result.teams[0][0 * nl + 0] = 1
|
|
result.teams[0][0 * nl + 1] = 1
|
|
result.teams[0][0 * nl + 4] = 1
|
|
result.teams[0][1 * nl + 2] = 1
|
|
result.teams[0][2 * nl + 5] = 1 # negative clause (index 2 >= half=2)
|
|
result.teams[0][2 * nl + 6] = 1
|
|
|
|
proc testClauseSummaryShape() =
|
|
let m = shapeModel()
|
|
let spec = tinySpec()
|
|
let samples = @[makeSample(8, @[1, 1, 1, 0, 1, 1, 1, 0], 0)]
|
|
let summ = clauseSummary(clauseInfo(m, samples, spec))
|
|
check "summary nonEmpty = 3", summ.nonEmpty == 3
|
|
check "summary empty = 5", summ.emptyClauses == 5
|
|
check "summary mean length = 2.0", abs(summ.meanLength - 2.0) < 1e-9
|
|
check "summary median = 2.0", abs(summ.medianLength - 2.0) < 1e-9
|
|
check "summary min = 1", summ.minLength == 1
|
|
check "summary max = 3", summ.maxLength == 3
|
|
check "summary lengthHist", summ.lengthHist == @[0, 1, 1, 1]
|
|
check "summary posMean = (3+1)/2", abs(summ.posMeanLength - 2.0) < 1e-9
|
|
check "summary negMean = 2", abs(summ.negMeanLength - 2.0) < 1e-9
|
|
check "summary per-class mean = 2.0",
|
|
abs(summ.perClassMeanLength[0] - 2.0) < 1e-9
|
|
check "summary emptyFraction = 5/8",
|
|
abs(summ.emptyFraction - 5.0 / 8.0) < 1e-9
|
|
|
|
# ── healthy-band verdict ─────────────────────────────────────────────────────
|
|
|
|
proc verdictOf(mean: float, n = 4): string =
|
|
var s = ClauseSummary()
|
|
s.nonEmpty = n
|
|
s.totalClauses = 8
|
|
s.meanLength = mean
|
|
clauseShapeVerdict(s)
|
|
|
|
proc testShapeVerdict() =
|
|
check "mean 1.5 -> collapsed", verdictOf(1.5) == "collapsed"
|
|
check "mean 2.5 -> short", verdictOf(2.5) == "short"
|
|
check "mean 3.0 -> healthy (band low edge)", verdictOf(3.0) == "healthy"
|
|
check "mean 8.0 -> healthy (band high edge)", verdictOf(8.0) == "healthy"
|
|
check "mean 12.0 -> too long", verdictOf(12.0) == "too long"
|
|
check "mean 19.17 -> too long (the shipped gun)", verdictOf(19.17) == "too long"
|
|
check "no non-empty clauses -> n/a", verdictOf(0.0, 0) == "n/a"
|
|
# custom band is honoured
|
|
var s = ClauseSummary()
|
|
s.nonEmpty = 4
|
|
s.totalClauses = 8
|
|
s.meanLength = 2.5
|
|
check "custom band 2-6 makes mean 2.5 healthy",
|
|
clauseShapeVerdict(s, 2.0, 6.0, 1.0) == "healthy"
|
|
check "custom collapsedMax makes mean 2.5 collapsed",
|
|
clauseShapeVerdict(s, 2.0, 6.0, 3.0) == "collapsed"
|
|
|
|
proc testLengthHistText() =
|
|
let t = lengthHistText(@[0, 3, 1])
|
|
check "lengthHistText skips empty bins", "len 0" notin t
|
|
check "lengthHistText renders len 1", "len 1: 3" in t
|
|
check "lengthHistText renders len 2", "len 2: 1" in t
|
|
|
|
# ── per-block contribution ───────────────────────────────────────────────────
|
|
|
|
proc testBlockContribution() =
|
|
let m = shapeModel()
|
|
let spec = tinySpec()
|
|
let blocks = blockLengthContributions(m, spec)
|
|
check "one entry per block", blocks.len == 2
|
|
# block0 (bits 0-3): 2 in clause0 + 1 in clause1 = 3 over 3 clauses
|
|
check "block0 total = 3", blocks[0].totalLits == 3
|
|
check "block0 mean/clause = 1.0", abs(blocks[0].meanPerClause - 1.0) < 1e-9
|
|
check "block0 clausesUsing = 2", blocks[0].clausesUsing == 2
|
|
check "block0 max per clause = 2", blocks[0].perClauseMax == 2
|
|
# block1 (bits 4-7): 1 in clause0 + 2 in neg clause2 = 3 over 3 clauses
|
|
check "block1 total = 3", blocks[1].totalLits == 3
|
|
check "block1 mean/clause = 1.0", abs(blocks[1].meanPerClause - 1.0) < 1e-9
|
|
check "shares sum to 1", abs(blocks[0].share + blocks[1].share - 1.0) < 1e-9
|
|
# skipEmpty = false averages over all 8 clauses instead of 3
|
|
let blocksAll = blockLengthContributions(m, spec, skipEmpty = false)
|
|
check "skipEmpty=false dilutes the mean",
|
|
abs(blocksAll[0].meanPerClause - 3.0 / 8.0) < 1e-9
|
|
|
|
# ── coverage / concentration ─────────────────────────────────────────────────
|
|
|
|
proc testCoverage() =
|
|
## 2 classes x 2 clauses = 4 clauses. class0/clause0 = {+b0.0} fires on the
|
|
## all-ones sample; nothing else fires. So 1 of 4 clauses fires.
|
|
var m = newMachine(8, 2, 2, 64, 3.0, 1)
|
|
let nl = m.nLiterals
|
|
m.teams[0][0 * nl + 0] = 1
|
|
let samples = @[
|
|
makeSample(8, @[1, 0, 0, 0, 0, 0, 0, 0], 0),
|
|
makeSample(8, @[1, 1, 1, 1, 1, 1, 1, 1], 0),
|
|
]
|
|
let cov = clauseCoverage(m, samples)
|
|
check "coverage totalClauses = 4", cov.totalClauses == 4
|
|
check "coverage mean firing = 1", abs(cov.meanFiringClauses - 1.0) < 1e-9
|
|
check "coverage fraction = 1/4", abs(cov.meanFiringFraction - 0.25) < 1e-9
|
|
check "one firing clause", cov.firingClauseCount == 1
|
|
check "effective clauses = 1 (single clause does all voting)",
|
|
abs(cov.effectiveClauses - 1.0) < 1e-9
|
|
check "top3 share = 1 (all fires from one clause)",
|
|
abs(cov.top3Share - 1.0) < 1e-9
|
|
|
|
# ── report helpers ───────────────────────────────────────────────────────────
|
|
|
|
proc testReportHelpers() =
|
|
let m = shapeModel()
|
|
let spec = tinySpec()
|
|
let samples = @[makeSample(8, @[1, 1, 1, 0, 1, 1, 1, 0], 0)]
|
|
let d = clauseShapeDiagnostics(m, samples, spec)
|
|
check "shapeLine mentions shape/coverage",
|
|
"shape mean=" in shapeLine(d) and "coverage=" in shapeLine(d)
|
|
let rep = formatClauseShapeReport(d)
|
|
check "report has length histogram", "length histogram" in rep
|
|
check "report has per-block contribution", "per-block literal contribution" in rep
|
|
check "report has per-class mean", "per-class mean/median" in rep
|
|
|
|
# ── Task 3: synthetic validation with KNOWN geometry ─────────────────────────
|
|
|
|
const
|
|
NBits = 49
|
|
BitA = 0 ## dist-to-nearest-wall (WALLS block)
|
|
BitB = 45 ## bullet-lateral-offset (BULLETS block)
|
|
NoiseBit = 17
|
|
NClasses = 3
|
|
|
|
proc genRule(n, seed: int): seq[DiagSample] =
|
|
## class2 = A AND B ; class1 = A AND NOT B ; class0 = NOT A.
|
|
var rng = initRand(seed)
|
|
for i in 0..<n:
|
|
var raw = newSeq[int](NBits)
|
|
for b in 0..<NBits: raw[b] = (if rng.rand(1.0) < 0.5: 1 else: 0)
|
|
let a = if rng.rand(1.0) < 0.4: 1 else: 0
|
|
let bb = if rng.rand(1.0) < 0.5: 1 else: 0
|
|
raw[BitA] = a
|
|
raw[BitB] = bb
|
|
let label = if a == 1 and bb == 1: 2 elif a == 1: 1 else: 0
|
|
result.add makeSample(NBits, raw, label, i)
|
|
|
|
proc genRandom(n, seed: int): seq[DiagSample] =
|
|
var rng = initRand(seed)
|
|
for i in 0..<n:
|
|
var raw = newSeq[int](NBits)
|
|
for b in 0..<NBits: raw[b] = (if rng.rand(1.0) < 0.5: 1 else: 0)
|
|
result.add makeSample(NBits, raw, rng.rand(NClasses - 1), i)
|
|
|
|
proc blockByName(blocks: seq[BlockLengthContribution], name: string):
|
|
BlockLengthContribution =
|
|
for b in blocks:
|
|
if b.name == name: return b
|
|
|
|
when isMainModule:
|
|
testPercentile()
|
|
testLengthStats()
|
|
testClauseSummaryShape()
|
|
testShapeVerdict()
|
|
testLengthHistText()
|
|
testBlockContribution()
|
|
testCoverage()
|
|
testReportHelpers()
|
|
|
|
echo "\n## SYNTHETIC VALIDATION (Task 3) — MEASURED"
|
|
let spec = draftTMSpec()
|
|
let train = genRule(3000, 1)
|
|
let ev = genRule(1500, 2)
|
|
|
|
# s=8 is chosen so the TM's recovered clauses sit inside the healthy band; the
|
|
# rule itself is only 2 literals, and the checker must say so.
|
|
let tmpl = newMachine(NBits, NClasses, nClauses = 40, nStates = 64,
|
|
sValue = 8.0, seed = 1)
|
|
let m = trainModel(tmpl, train, epochs = 25, seed = 777)
|
|
let d = clauseShapeDiagnostics(m, ev, spec)
|
|
echo &"# SHORT rule (2 literals) s=8.0: mean={d.summary.meanLength:.2f} " &
|
|
&"median={d.summary.medianLength:.2f} p10={d.summary.p10Length:.2f} " &
|
|
&"p90={d.summary.p90Length:.2f} max={d.summary.maxLength} " &
|
|
&"verdict={d.verdict} acc={evalAcc(m, ev)*100:.1f}%"
|
|
for c in 0..<NClasses:
|
|
echo &"# recovered class{c}: {spec.describeClause(necessaryLiterals(m, ev, c), c)}"
|
|
check "SHORT rule reads healthy (default band 3-8)", d.verdict == "healthy"
|
|
check "SHORT rule mean is inside the healthy band",
|
|
d.summary.meanLength >= 3.0 and d.summary.meanLength <= 8.0
|
|
# the recovered necessary literals are the 2-literal rule
|
|
var c2ok = false
|
|
let c2 = necessaryLiterals(m, ev, 2)
|
|
var pos2: set[uint8]
|
|
for l in c2:
|
|
if l < NBits: pos2.incl uint8(l)
|
|
c2ok = pos2 == {uint8(BitA), uint8(BitB)}
|
|
check "SHORT rule recovered exactly (A AND B)", c2ok
|
|
check "SHORT rule is NOT called too long", d.verdict != "too long"
|
|
|
|
# ── per-block contribution: planted blocks dominate, irrelevant block dead ──
|
|
# Use the sparser s=3 model here: with less padding the irrelevant block is
|
|
# unambiguously negligible (below the uniform 1/12 = 8.3% share).
|
|
let btmpl = newMachine(NBits, NClasses, nClauses = 40, nStates = 64,
|
|
sValue = 3.0, seed = 1)
|
|
let bm = trainModel(btmpl, train, epochs = 25, seed = 777)
|
|
let bd = clauseShapeDiagnostics(bm, ev, spec)
|
|
let walls = blockByName(bd.blocks, "dist-to-nearest-wall")
|
|
let bullets = blockByName(bd.blocks, "bullet-lateral-offset")
|
|
let usBlock = blockByName(bd.blocks, "dist-from-us")
|
|
echo &"# per-block (s=3.0): walls(mean={walls.meanPerClause:.2f} share={walls.share*100:.1f}%) " &
|
|
&"bullets(mean={bullets.meanPerClause:.2f} share={bullets.share*100:.1f}%) " &
|
|
&"dist-from-us(mean={usBlock.meanPerClause:.2f} share={usBlock.share*100:.1f}%) " &
|
|
&"[uniform=8.3%]"
|
|
check "the WALLS block (holds the planted A) dominates", walls.share > 0.25
|
|
check "the BULLETS block (holds the planted B) is a top contributor",
|
|
bullets.share > 0.15
|
|
check "the irrelevant US block is negligible (below uniform share)",
|
|
usBlock.share < 0.083
|
|
check "the planted blocks outweigh the irrelevant block by >5x",
|
|
usBlock.totalLits > 0 and
|
|
walls.totalLits + bullets.totalLits > 5 * usBlock.totalLits
|
|
# the noise MOTION bit is not singled out as a top block
|
|
check "the pure-noise turn-direction block is far smaller than WALLS",
|
|
blockByName(bd.blocks, "turn-direction").totalLits < walls.totalLits
|
|
|
|
# ── RANDOM labels: report whatever the checker actually says ──
|
|
let rtrain = genRandom(3000, 5)
|
|
let rev = genRandom(1500, 6)
|
|
let rtmpl = newMachine(NBits, NClasses, nClauses = 40, nStates = 64,
|
|
sValue = 3.0, seed = 1)
|
|
let rm = trainModel(rtmpl, rtrain, epochs = 25, seed = 777)
|
|
let rd = clauseShapeDiagnostics(rm, rev, spec)
|
|
echo &"# RANDOM labels s=3.0: mean={rd.summary.meanLength:.2f} " &
|
|
&"median={rd.summary.medianLength:.2f} p90={rd.summary.p90Length:.2f} " &
|
|
&"max={rd.summary.maxLength} empty={rd.summary.emptyClauses}/" &
|
|
&"{rd.summary.totalClauses} verdict={rd.verdict} " &
|
|
&"acc={evalAcc(rm, rev)*100:.1f}% eff={rd.coverage.effectiveClauses:.1f} " &
|
|
&"top3={rd.coverage.top3Share*100:.1f}%"
|
|
# On THIS encoding noise does not pad clauses long; it fails to commit at all
|
|
# (short clauses, many firing, low concentration). Report and assert only that
|
|
# the verdict is a legal word and that it is not "healthy".
|
|
check "RANDOM labels are not reported healthy",
|
|
rd.verdict in ["collapsed", "short", "too long"]
|
|
check "RANDOM-label accuracy is near the majority baseline",
|
|
abs(evalAcc(rm, rev) - 1.0 / 3.0) < 0.10
|
|
|
|
echo ""
|
|
if failures > 0:
|
|
echo &"{failures} / {checks} check(s) FAILED"
|
|
quit(1)
|
|
echo "All ", checks, " clause-shape checks passed."
|