Re-adaptation: the sliding WINDOW is a control-validated +9.3pp on late accuracy

The user's first-hand diagnosis: "once the pattern is learnt enough we hit DrussGT,
but as soon as it adapts we are not fast enough to re-adapt again." The gun kept
EVERY sample for the whole battle (which the user explicitly asked for), so stale
evidence weighed the same as new evidence - accumulation without forgetting.

TASK 1 - THE CURVE, MEASURED FIRST (prequential side accuracy, 15 rounds x 4
horizons, deciles of the eval stream):
  arm                 early%  late%  decay
  frozen (early only)   64.9   61.4   -3.5
  accum (shipped)       76.4   75.3   -1.1
  window N=150          84.7   84.6   -0.0
  resetdrop 5pp         83.2   84.3   +1.0
**The decay IS real but modest and LOCALIZED IN THE LAST ~30% of the round**
(last-third vs middle-third: frozen -7.7pp, accum -3.7, window -1.3, resetdrop -1.3).

TASK 3 - WHICH FIX HELPS (late-half accuracy, shuffled control in parens):
  accum (shipped)                       75.3
  **window N=150                       84.6  (51)**   +9.3pp
  **resetdrop 5pp                      84.3  (51)**   +9.0pp
  rehearse-all (retrain, NO forgetting) 79.5
- **Sliding window: +9.3pp late (within-round), +9.7 (cross-round), +10.3 (shield).**
  The shuffled control stays ~50-51%, so it learns the ENEMY, not noise.
- **Forgetting is the essential ingredient**: the same periodic retrain WITHOUT
  forgetting reaches only ~79.5%, so roughly half the gain is the retraining
  mechanics and half is the forgetting.
- Change detection ties the window on late accuracy and gives the best decay, but
  at a 5pp threshold it fired 300-500 times in the offline stream - noisy.
- INERTIA IS REDUNDANT WITH THE WINDOW: lower inertia helps the keep-everything
  model a lot (accum late 75.3 -> 83.3 at N=8) but leaves the window FLAT
  (84.6-84.7 at every N). Low inertia and forgetting are SUBSTITUTES, and
  forgetting is the robust one - `TR_TMHORIZON_NSTATES=8` is NOT the primary fix.

VERDICT: SHIPPING CANDIDATE = `TR_TMHORIZON_WINDOW=150`, kept at default 0 until a
live A/B confirms.

HONEST CAVEATS THAT SET EXPECTATIONS:
- The fixtures are OPEN-LOOP (DrussGT does not react to our bullets), so a true
  mid-round adaptation is NOT present; the dominant measured effect is the LEVEL
  gap, not the decay magnitude. The user's "it adapts" magnitude is still INFERRED.
- The harness uses a STRAIGHT-LINE base while the live gun uses Pattern's
  prediction, and it ignores the h-tick label delay, so its absolute side accuracy
  (75-85%) is INFLATED: the same gun measured ~52% = chance live against DrussGT.
  So +9.3pp is a real ARM DELTA, not a promise that the gun now clears the ~80%
  accuracy wall that hits need. A live A/B must decide.

Adds `measure_tm_readapt.nim` (prequential harness with the mandatory shuffled
control, within-round + cross-round protocols, inertia sweep) and its captured
results. test_tm_horizon 79 -> 104 (five new groups). test_tm_diag 48,
test_tm_automata_diag 55, test_tm_clause_shape 66, test_rack_membership 48 pass.
ModularBot compiles.

.gitignore: switched from a broad `measure_*` pattern to EXPLICIT binary names.
The broad rule was too blunt - it also excluded the `.txt` results file, which made
`git add` refuse the whole commit twice. Sources stay tracked; binaries do not.
This commit is contained in:
2026-09-23 00:29:37 +02:00
parent b68707c867
commit 69debbe347
5 changed files with 1267 additions and 7 deletions
+248 -6
View File
@@ -56,6 +56,31 @@
## per tick, one TM evaluation per horizon bucket), so it stays in the
## neighbourhood of the 0.36 ms/tick TM gun rather than Tsetlin's ~5.3 ms/tick.
##
## RE-ADAPTATION (all knobs default to the behaviour above — nothing regresses):
## The user's failure mode is "it learns the enemy, the enemy adapts, and we
## are too slow to re-adapt". The cause is that the machines accumulate EVERY
## resolved sample for the whole battle, so old evidence weighs as much as new.
## Three independent, off-by-default fixes:
## * TR_TMHORIZON_WINDOW=N (>0): every TR_TMHORIZON_RETRAIN_EVERY samples
## (default 50) rebuild BOTH heads from scratch and retrain on the last N
## resolved samples from a bounded ring. Stale evidence ages out.
## * TR_TMHORIZON_RESET_DROP=pp (>0): track the rolling-100 side accuracy;
## when it falls more than `pp` below its own recent peak, treat it as
## "the enemy changed", rebuild the heads and re-learn from the recent
## window. The rolling-100/300 accuracy and its per-round min/mean/end are
## always available in `roundSummary`; TR_TMHORIZON_ACCURVE=1 logs a
## `[tmh-acc]` curve line every 25 warm samples so the decay is visible.
## * TR_TMHORIZON_NSTATES=K: runtime automata state count (inertia), so the
## fast/slow trade-off can be swept without a rebuild. Default = the
## compile-time TMH_NSTATES.
## Offline measurement (`tests/measure_tm_readapt.nim`, prequential side
## accuracy on the DrussGT fixtures): on `tr_drussgt_vs_modularbot` the
## keep-everything arm sits at ~76% late accuracy, the sliding window at ~85%,
## the change-detection re-learn at ~84%, and a shuffled-label control stays at
## ~50%. Lowering the state count from 64 to 8 lifts the keep-everything arm to
## ~83% but leaves the window flat — i.e. forgetting and low inertia are
## substitutes, and forgetting is the stronger one.
##
## Coordinate system: 0° = East, CCW positive (Tank Royale standard).
import std/[math, os, strutils, strformat]
@@ -101,6 +126,34 @@ const
TMH_SHIFT_DEFAULT* = 2.0
TMH_BIG_MULT_DEFAULT* = 1.5
TMH_RESET_ON_TARGET_DEFAULT* = true
## ── re-adaptation knobs (all default to today's behaviour) ─────────────
## Sliding-window / forgetting: >0 rebuilds both heads on the most recent N
## resolved training samples every `TMH_RETRAIN_EVERY` new samples, so stale
## evidence ages out. 0 = keep everything (current shipped behaviour).
TMH_WINDOW_ENV* = "TR_TMHORIZON_WINDOW"
## Change detection: when the rolling side accuracy falls more than this many
## percentage points below its own recent peak, treat it as "the enemy
## changed" and rebuild + re-learn from the recent window. 0 = off.
TMH_RESET_DROP_ENV* = "TR_TMHORIZON_RESET_DROP"
## Runtime automata state count (inertia). Default = the compile-time value.
TMH_NSTATES_ENV* = "TR_TMHORIZON_NSTATES"
## 1 = log the rolling accuracy every `TMH_ACC_LOG_EVERY` warm samples.
TMH_ACCURVE_ENV* = "TR_TMHORIZON_ACCURVE"
## How often the sliding/window mode does its full retrain, and how many
## epochs that retrain runs. Defaults keep the retrain cheap.
TMH_RETRAIN_EVERY_ENV* = "TR_TMHORIZON_RETRAIN_EVERY"
TMH_EPOCHS_ENV* = "TR_TMHORIZON_EPOCHS"
## ── fixed instrumentation geometry ──────────────────────────────────────
TMH_ACC_CAP* = 512 ## rolling-accuracy ring (> the 300 window)
TMH_ACC_WIN1* = 100 ## fast rolling window (samples)
TMH_ACC_WIN2* = 300 ## slow rolling window (samples)
TMH_ACC_LOG_EVERY* = 25 ## ACCURVE: one line every N warm samples
TMH_RETRAIN_EVERY_DEF* = 50
TMH_EPOCHS_DEF* = 1
## When only RESET_DROP is on (WINDOW == 0) the re-learn uses this window.
TMH_RESET_WINDOW_DEF* = 200
TMH_RESET_MIN_SAMPLES* = 100 ## never judge a drop before this many samples
TMH_RESET_COOLDOWN* = 100 ## samples between two change-triggered resets
type
TmhTick* = object
@@ -117,6 +170,13 @@ type
sx, sy: float
active: bool
TmhSample = object
## One resolved training sample kept in the bounded ring that feeds the
## sliding-window rebuild and the change-detection re-learn.
lits: array[TMH_NLITS, uint8]
side: uint8
mag: uint8
TmhPending* = object
## One deferred training sample. `lits` is the exact literal vector the TM
## saw at fire time; the label is resolved h ticks later.
@@ -188,6 +248,29 @@ type
# magnitude-median histogram (TMH_ABS_RES deg bins)
absHist: array[TMH_ABS_BINS, int]
absCount: int
# ── re-adaptation: knobs + bounded sample ring ──────────────────────────
windowN*: int ## TR_TMHORIZON_WINDOW (0 = keep everything)
resetDrop*: float ## TR_TMHORIZON_RESET_DROP (0 = off), percentage points
retrainEvery*: int
retrainEpochs*: int
nStates*: int ## runtime automata state count (inertia)
buffer: seq[TmhSample]
bufCap: int
bufCount: int ## total resolved samples ever appended
sinceRetrain: int
# ── rolling side accuracy (change detection + curve) ────────────────────
accRing: array[TMH_ACC_CAP, uint8]
accPos: int
accCount*: int ## total warm samples recorded in the ring
accPeak*: float ## peak rolling-100 accuracy (percent)
sinceResetDrop: int
resetDrops*: int ## change-triggered re-learns
accurveEnabled*: bool
accSnapMin*: float
accSnapMax*: float
accSnapSum*: float
accSnapN*: int
lastAccLogCount: int
# ── instrumentation ─────────────────────────────────────────────────────
trained*: int
sideCorrect*, sideTotal*: int
@@ -226,6 +309,12 @@ proc envBoolT(name: string, default: bool): bool =
of "0", "false", "no", "off": false
else: default
proc envIntT(name: string, default: int): int =
let v = getEnv(name, "")
if v.len == 0: return default
try: parseInt(v.strip())
except ValueError: default
# ── horizon maths (pure, unit-tested) ────────────────────────────────────────
proc tmhHorizonFor*(dist, bulletSpeed: float): int =
@@ -275,11 +364,30 @@ proc tmhLits*(base: array[TMH_N_BASE, uint8],
# ── construction / reset ─────────────────────────────────────────────────────
proc ensureBuffer(g: var TmHorizonGun) =
## Grow the resolved-sample ring to hold whichever window the configured
## forgetting/re-learn mode needs. Never shrinks, so switching knobs at
## runtime cannot lose already-buffered samples.
let want = max(g.windowN,
(if g.resetDrop > 0.0: TMH_RESET_WINDOW_DEF else: 0))
if want > g.bufCap:
g.bufCap = want
g.buffer.setLen(want)
proc initTmHorizonGun*(): TmHorizonGun =
result.pattern = PatternMatcherGun()
result.sideMachine = newMachine(TMH_N_BITS, 2, TMH_NCLAUSES, TMH_NSTATES,
# Runtime inertia: read once here so the state count can be swept with an env
# knob instead of a rebuild. Default is the compile-time TMH_NSTATES.
result.nStates = clamp(envIntT(TMH_NSTATES_ENV, TMH_NSTATES), 2, 4096)
result.windowN = max(0, envIntT(TMH_WINDOW_ENV, 0))
result.resetDrop = max(0.0, envFloatT(TMH_RESET_DROP_ENV, 0.0))
result.retrainEvery = max(1, envIntT(TMH_RETRAIN_EVERY_ENV, TMH_RETRAIN_EVERY_DEF))
result.retrainEpochs = max(1, envIntT(TMH_EPOCHS_ENV, TMH_EPOCHS_DEF))
result.accurveEnabled = envBoolT(TMH_ACCURVE_ENV, false)
result.ensureBuffer()
result.sideMachine = newMachine(TMH_N_BITS, 2, TMH_NCLAUSES, result.nStates,
TMH_S, seed = 1)
result.magMachine = newMachine(TMH_N_BITS, 2, TMH_NCLAUSES, TMH_NSTATES,
result.magMachine = newMachine(TMH_N_BITS, 2, TMH_NCLAUSES, result.nStates,
TMH_S, seed = 2)
for c in 0..1:
result.sideScratch[c] = newSeq[uint8](TMH_NCLAUSES)
@@ -317,6 +425,23 @@ proc setResetOnTarget*(g: var TmHorizonGun, on: bool) =
g.resetOnTarget = on
g.shiftConfigured = true
proc setWindow*(g: var TmHorizonGun, n: int) =
## Explicit per-gun override (tests / offline sweeps): sliding-window width.
g.windowN = max(0, n)
g.ensureBuffer()
proc setResetDrop*(g: var TmHorizonGun, pp: float) =
## Explicit per-gun override (tests): change-detection threshold in pp.
g.resetDrop = max(0.0, pp)
g.ensureBuffer()
proc setRetrainConfig*(g: var TmHorizonGun, every, epochs: int) =
g.retrainEvery = max(1, every)
g.retrainEpochs = max(1, epochs)
proc setAccurve*(g: var TmHorizonGun, on: bool) =
g.accurveEnabled = on
proc resetRoundState*(g: var TmHorizonGun) =
## PER-ROUND wipe ONLY. Clears the observation ring, the deferred-label queue
## and every piece of motion / bullet / per-tick history that is meaningless
@@ -344,6 +469,12 @@ proc resetRoundState*(g: var TmHorizonGun) =
# A fresh round starts here: remember the cumulative count so the per-round
# summary can report `thisRound` while `trained` keeps climbing.
g.roundStartTrained = g.trained
# Per-round accuracy-curve snapshots (the rolling ring itself SURVIVES the
# boundary: change detection must see the whole battle).
g.accSnapMin = 0.0
g.accSnapMax = 0.0
g.accSnapSum = 0.0
g.accSnapN = 0
proc resetLearning*(g: var TmHorizonGun, reason = "") =
## PER-BATTLE / PER-ENEMY wipe. Wipes the Tsetlin machines and every learned
@@ -369,6 +500,16 @@ proc resetLearning*(g: var TmHorizonGun, reason = "") =
for i in 0..<4: g.quadHist[i] = 0
for i in 0..<TMH_ABS_BINS: g.absHist[i] = 0
g.absCount = 0
# Re-adaptation state: the bounded sample ring and the rolling-accuracy ring
# are part of "learning", so a new battle/enemy wipes them.
g.bufCount = 0
g.sinceRetrain = 0
g.accCount = 0
g.accPos = 0
g.accPeak = 0.0
g.sinceResetDrop = 0
g.resetDrops = 0
g.lastAccLogCount = 0
g.resetRoundState()
if reason.len > 0:
g.ensureConfig()
@@ -416,6 +557,10 @@ proc sideClauseStates*(g: TmHorizonGun): seq[int16] =
## states, so a test can prove the machines survive or are wiped by a reset.
g.sideMachine.teams[0] & g.sideMachine.teams[1]
proc headStates*(g: TmHorizonGun): int =
## Test seam: the runtime automata state count both heads were built with.
g.sideMachine.nStates
proc sideClausesAllExclude*(g: TmHorizonGun): bool =
## Test seam: true when every side clause is at the Exclude boundary.
for c in 0..<g.sideMachine.nClasses:
@@ -651,6 +796,80 @@ proc tmhTrainOne(m: var TmMachine, lits: array[TMH_NLITS, uint8], label: int,
let d = if c == label: 1.0 else: -1.0
tmLearnDir(m, m.teams[c], lits, caches[c], votes[c], d)
# ── sliding window / change detection ────────────────────────────────────────
proc bufferedCount*(g: TmHorizonGun): int =
## Test/observability seam: how many resolved samples the bounded ring has
## seen since the last battle/enemy reset.
g.bufCount
proc bufferCapacity*(g: TmHorizonGun): int =
## Test/observability seam: the allocated ring capacity (0 = buffering off).
g.bufCap
proc rollingAcc*(g: TmHorizonGun, n: int): float =
## Side accuracy over the last `min(n, accCount)` recorded warm samples.
## This is the measurement that shows whether accuracy decays when the
## enemy changes. Returns 0.0 when nothing has been recorded yet.
if g.accCount == 0: return 0.0
let k = min(n, g.accCount)
if k <= 0: return 0.0
var cor = 0
for i in 0..<k:
let idx = ((g.accPos - 1 - i) mod TMH_ACC_CAP + TMH_ACC_CAP) mod TMH_ACC_CAP
cor += int(g.accRing[idx])
cor.float / k.float
proc retrainFromBuffer*(g: var TmHorizonGun, n: int) =
## FULL rebuild of both heads from scratch, then train on the most recent
## `min(n, buffered)` resolved samples (oldest -> newest). `trained` is NOT
## reset: it counts resolved samples over the battle, so the cold gate and the
## per-round summary stay monotone while the MACHINES forget. The rebuild is
## DETERMINISTIC given the buffer (fixed seed), so an A/B rerun is repeatable.
g.sideMachine.resetMachine(seed = 1000'u64)
g.magMachine.resetMachine(seed = 2000'u64)
if g.bufCap <= 0: return
let k = min(n, min(g.bufCount, g.bufCap))
if k <= 0: return
for _ in 0..<g.retrainEpochs:
for j in 0..<k:
let idx = ((g.bufCount - k + j) mod g.bufCap + g.bufCap) mod g.bufCap
let s = g.buffer[idx]
tmhTrainOne(g.sideMachine, s.lits, int(s.side), g.sideScratch)
tmhTrainOne(g.magMachine, s.lits, int(s.mag), g.magScratch)
proc recordWarm(g: var TmHorizonGun, correct: bool) =
## Record one warm sample's side correctness, update the rolling windows and
## the change-detection peak, and fire the re-learn trigger when the rolling
## accuracy has fallen far enough below its own recent peak.
g.accRing[g.accPos] = (if correct: 1'u8 else: 0'u8)
g.accPos = (g.accPos + 1) mod TMH_ACC_CAP
inc g.accCount
let a100 = g.rollingAcc(TMH_ACC_WIN1) * 100.0
if g.accSnapN == 0:
g.accSnapMin = a100
g.accSnapMax = a100
else:
if a100 < g.accSnapMin: g.accSnapMin = a100
if a100 > g.accSnapMax: g.accSnapMax = a100
g.accSnapSum += a100
inc g.accSnapN
if a100 > g.accPeak: g.accPeak = a100
inc g.sinceResetDrop
if g.resetDrop > 0.0 and g.accCount >= TMH_RESET_MIN_SAMPLES and
(g.accPeak - a100) > g.resetDrop and g.sinceResetDrop >= TMH_RESET_COOLDOWN:
let n = if g.windowN > 0: g.windowN else: TMH_RESET_WINDOW_DEF
g.retrainFromBuffer(n)
g.accPeak = a100
g.sinceResetDrop = 0
inc g.resetDrops
if g.accurveEnabled and (g.accCount mod TMH_ACC_LOG_EVERY == 0) and
g.accCount != g.lastAccLogCount:
g.lastAccLogCount = g.accCount
let a300 = g.rollingAcc(TMH_ACC_WIN2) * 100.0
echo fmt"[tmh-acc] trained={g.trained} warm={g.accCount} acc100={a100:.1f} " &
fmt"acc300={a300:.1f} peak={g.accPeak:.1f} resets={g.resetDrops}"
# ── magnitude median (running histogram) ─────────────────────────────────────
proc recordAbs(g: var TmHorizonGun, a: float) =
@@ -691,12 +910,26 @@ proc tmhResolveOne(g: var TmHorizonGun, state: WorldState,
let magLabel = if absE > med: 1 else: 0
if p.warm:
inc g.sideTotal
if p.sidePred == sideLabel: inc g.sideCorrect
let correct = p.sidePred == sideLabel
if correct: inc g.sideCorrect
# Rolling accuracy + change detection (the re-adaptation measurement).
g.recordWarm(correct)
inc g.sideLabelHist[sideLabel]
inc g.magLabelHist[magLabel]
tmhTrainOne(g.sideMachine, p.lits, sideLabel, g.sideScratch)
tmhTrainOne(g.magMachine, p.lits, magLabel, g.magScratch)
inc g.trained
# Bounded ring feeding the sliding-window rebuild / change-detection
# re-learn. Appended AFTER the online update so a retrain triggered this
# tick sees this sample too.
if g.bufCap > 0:
g.buffer[g.bufCount mod g.bufCap] =
TmhSample(lits: p.lits, side: uint8(sideLabel), mag: uint8(magLabel))
inc g.bufCount
inc g.sinceRetrain
if g.windowN > 0 and g.sinceRetrain >= g.retrainEvery:
g.sinceRetrain = 0
g.retrainFromBuffer(g.windowN)
g.recordAbs(absE)
true
@@ -755,16 +988,25 @@ proc tmhLog(g: var TmHorizonGun, state: WorldState, bulletSpeed: float,
fmt"aim={aimDeg:.1f} gun=Pattern conf={sideConf:.2f} trained={g.trained}"
proc roundSummary*(g: var TmHorizonGun) =
## Per-round summary on round end (behind the same log switch), so the user can
## watch it learn across the round.
## Per-round summary on round end (behind `TR_TMHORIZON_LOG=1`, or
## `TR_TMHORIZON_ACCURVE=1` for the accuracy curve alone), so the user can
## watch it learn across the round. The rolling windows are the re-adaptation
## measurement: `acc100` = last 100 warm samples, `acc300` = last 300, and
## `accCurve[min/mean/end]` the per-round spread of the rolling-100 snapshot.
g.ensureConfig()
if not g.logEnabled: return
if not g.logEnabled and not g.accurveEnabled: return
let thisRound = g.trained - g.roundStartTrained
let acc = if g.sideTotal > 0: g.sideCorrect.float / g.sideTotal.float * 100.0
else: 0.0
let a100 = g.rollingAcc(TMH_ACC_WIN1) * 100.0
let a300 = g.rollingAcc(TMH_ACC_WIN2) * 100.0
let snapMean = if g.accSnapN > 0: g.accSnapSum / float(g.accSnapN) else: 0.0
echo fmt"[tmh-round] trained={g.trained} thisRound={thisRound} pending={g.pendingCount} " &
fmt"dropped={g.pendingDropped} sideAcc={g.sideCorrect}/{g.sideTotal} " &
fmt"({acc:.1f}%) " &
fmt"acc100={a100:.1f}% acc300={a300:.1f}% " &
fmt"accCurve[min={g.accSnapMin:.1f} mean={snapMean:.1f} end={a100:.1f} n={g.accSnapN}] " &
fmt"resets={g.resetDrops} window={g.windowN} resetDrop={g.resetDrop:.1f} " &
fmt"quad=[SR:{g.quadHist[0]} LR:{g.quadHist[1]} " &
fmt"SL:{g.quadHist[2]} LL:{g.quadHist[3]}] " &
fmt"sidePred=[R:{g.sidePredHist[0]} L:{g.sidePredHist[1]}] " &