test(range): restore the 12/12 offline==online proof; measure TM clause readability

Task 1 - the acceptance proof was unrunnable because RecordWorldState was a
compile-time const set to false. It is now a RUNTIME switch
(let RecordWorldState* = existsEnv("TR_RECORD_WORLDSTATE")), default OFF, so
ordinary runs write no fixture, and acceptance_offline_vs_online.nim enables
it for the battle it spawns and clears it afterwards. Restored and run twice:
12/12 deterministic guns match exactly (128-tick and 546-tick battles), with
Tsetlin reported separately as stochastic. Both nimble build variants clean.

Task 2 - does a compact encoding turn the TM's 99.35% into a READABLE rule?
Measured across window sizes (fixed seed, no tuning):

  frames  TEST acc  eff.lits/clause  firing clauses  counterfactual low/high/mean
  10      99.35%    152.8            37              100/24/62.4%
  3       95.94%    54.9             35              96/20/58.6%
  2       99.48%    39.6             38              95/25/60.7%
  1       98.30%    19.2             45              100/24/62.3%

So 2 frames is strictly better than 10 on BOTH axes: +0.13 accuracy for 4x
smaller clauses. The 3-frame dip is non-monotonic and left unexplained rather
than smoothed over.

A readable rule WAS partially recovered. Five clauses carry the exact Gray
form !g10 ^ !g9 ^ !g8; g10 is inert in this data, so the effective rule is the
2-literal proposition !g9 ^ !g8, i.e. energy < 25.6. That is a genuine
threshold in readable propositional form - but at 25.6, NOT the labelled 30,
because 256 is a power-of-two Gray boundary expressible in two literals while
300 needs a longer conjunction. The TM found the nearest SIMPLE threshold.

The honest caveat: that threshold is not the ensemble's decision mechanism.
The counterfactual follow rate (high 24%, mean 62.3%) is statistically
identical at 1, 2 and 10 frames, so compactness did not make the model read
energy - its vote is carried by co-occurring bearing/velocity/heading/wall
literals. Also identified: clauses containing all 11 Gray energy bits are
satisfied at exactly one raw value (50, the dataset floor), so they are
'energy has hit the floor' detectors, not thresholds.

Methodological fix worth keeping: the earlier single-frame counterfactual
wrote energy into all 10 frame slots including the zeroed ones, reviving dead
clauses and producing a spurious 2% high-follow rate. setEnergyFrames now
rewrites only the exposed frames; the corrected figure is 24%.
This commit is contained in:
2026-09-21 00:14:46 +02:00
parent 89370008da
commit d5061ee215
3 changed files with 282 additions and 52 deletions
@@ -3,7 +3,9 @@
##
## Steps:
## 1. run ONE live ModularBot vs OscillatorBot round with the ModularBot
## recorder ON (compiled in via `const RecordWorldState = true`),
## recorder enabled for this battle only (the test exports
## TR_RECORD_WORLDSTATE=1, which the bot reads at RUNTIME; ordinary
## builds leave it unset and write no fixture),
## 2. read the online per-gun virtual fitness from /tmp/gun_stats.jsonl,
## 3. replay the recorded WorldState fixture offline through the same guns,
## 4. compare.
@@ -52,6 +54,12 @@ proc main() =
for p in [statsPath, recordPath]:
if fileExists(p): removeFile(p)
# Enable the ModularBot's runtime world-state recorder for THIS battle only.
# The env var is inherited by the battle-runner process and then by the bot
# processes it spawns, so a single invocation of this test is self-contained.
putEnv("TR_RECORD_WORLDSTATE", "1")
defer: delEnv("TR_RECORD_WORLDSTATE")
echo "=== live battle: ModularBot vs OscillatorBot, 1 round, max speed ==="
let battle = runBattle(@[modularBotDir, adversaryDir], rounds = 1,
timeout = 240000, maxSpeed = true)
@@ -59,7 +67,7 @@ proc main() =
echo fmt" {res.name:<14} rank={res.rank} score={res.totalScore}"
if not fileExists(recordPath):
echo "FAIL: recorder produced no fixture (is RecordWorldState true?)"
echo "FAIL: recorder produced no fixture (TR_RECORD_WORLDSTATE not inherited?)"
quit(1)
if not fileExists(statsPath):
echo "FAIL: no /tmp/gun_stats.jsonl"