Re-adaptation: the sliding WINDOW is a control-validated +9.3pp on late accuracy

The user's first-hand diagnosis: "once the pattern is learnt enough we hit DrussGT,
but as soon as it adapts we are not fast enough to re-adapt again." The gun kept
EVERY sample for the whole battle (which the user explicitly asked for), so stale
evidence weighed the same as new evidence - accumulation without forgetting.

TASK 1 - THE CURVE, MEASURED FIRST (prequential side accuracy, 15 rounds x 4
horizons, deciles of the eval stream):
  arm                 early%  late%  decay
  frozen (early only)   64.9   61.4   -3.5
  accum (shipped)       76.4   75.3   -1.1
  window N=150          84.7   84.6   -0.0
  resetdrop 5pp         83.2   84.3   +1.0
**The decay IS real but modest and LOCALIZED IN THE LAST ~30% of the round**
(last-third vs middle-third: frozen -7.7pp, accum -3.7, window -1.3, resetdrop -1.3).

TASK 3 - WHICH FIX HELPS (late-half accuracy, shuffled control in parens):
  accum (shipped)                       75.3
  **window N=150                       84.6  (51)**   +9.3pp
  **resetdrop 5pp                      84.3  (51)**   +9.0pp
  rehearse-all (retrain, NO forgetting) 79.5
- **Sliding window: +9.3pp late (within-round), +9.7 (cross-round), +10.3 (shield).**
  The shuffled control stays ~50-51%, so it learns the ENEMY, not noise.
- **Forgetting is the essential ingredient**: the same periodic retrain WITHOUT
  forgetting reaches only ~79.5%, so roughly half the gain is the retraining
  mechanics and half is the forgetting.
- Change detection ties the window on late accuracy and gives the best decay, but
  at a 5pp threshold it fired 300-500 times in the offline stream - noisy.
- INERTIA IS REDUNDANT WITH THE WINDOW: lower inertia helps the keep-everything
  model a lot (accum late 75.3 -> 83.3 at N=8) but leaves the window FLAT
  (84.6-84.7 at every N). Low inertia and forgetting are SUBSTITUTES, and
  forgetting is the robust one - `TR_TMHORIZON_NSTATES=8` is NOT the primary fix.

VERDICT: SHIPPING CANDIDATE = `TR_TMHORIZON_WINDOW=150`, kept at default 0 until a
live A/B confirms.

HONEST CAVEATS THAT SET EXPECTATIONS:
- The fixtures are OPEN-LOOP (DrussGT does not react to our bullets), so a true
  mid-round adaptation is NOT present; the dominant measured effect is the LEVEL
  gap, not the decay magnitude. The user's "it adapts" magnitude is still INFERRED.
- The harness uses a STRAIGHT-LINE base while the live gun uses Pattern's
  prediction, and it ignores the h-tick label delay, so its absolute side accuracy
  (75-85%) is INFLATED: the same gun measured ~52% = chance live against DrussGT.
  So +9.3pp is a real ARM DELTA, not a promise that the gun now clears the ~80%
  accuracy wall that hits need. A live A/B must decide.

Adds `measure_tm_readapt.nim` (prequential harness with the mandatory shuffled
control, within-round + cross-round protocols, inertia sweep) and its captured
results. test_tm_horizon 79 -> 104 (five new groups). test_tm_diag 48,
test_tm_automata_diag 55, test_tm_clause_shape 66, test_rack_membership 48 pass.
ModularBot compiles.

.gitignore: switched from a broad `measure_*` pattern to EXPLICIT binary names.
The broad rule was too blunt - it also excluded the `.txt` results file, which made
`git add` refuse the whole commit twice. Sources stay tracked; binaries do not.
This commit is contained in:
2026-09-23 00:29:37 +02:00
parent b68707c867
commit 69debbe347
5 changed files with 1267 additions and 7 deletions
+94 -1
View File
@@ -18,7 +18,7 @@
## Run with plain:
## nim c -r common_libs/tests/test_tm_horizon.nim
import std/[math]
import std/[math, os]
import gun_harness/gun_interface
import gun_harness/virtual_bullets
import gun_harness/offline_range
@@ -316,6 +316,94 @@ proc testWarmShiftMovesAim() =
check "warm: the applied rotation is the configured -3.0 deg",
approx(d, -3.0, 1e-6)
proc testReadaptDefaultKnobs() =
## The shipped defaults must be EXACTLY today's behaviour: no buffering, no
## change detection, compile-time state count, rolling curve off.
var g = initTmHorizonGun()
check "readapt: WINDOW defaults to 0 (keep everything)", g.windowN == 0
check "readapt: RESET_DROP defaults to 0 (off)", g.resetDrop == 0.0
check "readapt: the sample buffer is disabled by default", g.bufferCapacity() == 0
check "readapt: NSTATES defaults to the compile-time TMH_NSTATES",
g.nStates == TMH_NSTATES and g.headStates() == TMH_NSTATES
check "readapt: the accuracy curve is off by default", not g.accurveEnabled
drive(g, synthesizeCircular(ticks = 240))
check "readapt: no reset fires when RESET_DROP is off", g.resetDrops == 0
check "readapt: rolling accuracy is still recorded (measurement only)",
g.rollingAcc(100) >= 0.0 and g.rollingAcc(100) <= 1.0
proc testRuntimeStatesKnob() =
## Inertia is sweepable WITHOUT a rebuild via TR_TMHORIZON_NSTATES.
putEnv("TR_TMHORIZON_NSTATES", "16")
var g = initTmHorizonGun()
delEnv("TR_TMHORIZON_NSTATES")
check "readapt: NSTATES env sets the runtime state count", g.nStates == 16
check "readapt: both heads get the runtime state count",
g.headStates() == 16
var d = initTmHorizonGun()
check "readapt: an unset env cannot move the state count",
d.nStates == TMH_NSTATES
proc testWindowBuffersAndRebuilds() =
## Sliding-window mode must allocate a bounded ring, keep buffering resolved
## samples, and rebuild both heads deterministically from that ring.
var g = initTmHorizonGun()
g.setShift(0.0)
g.setWindow(50)
g.setRetrainConfig(50, 1)
check "window: the ring is allocated to the window width", g.bufferCapacity() == 50
drive(g, synthesizeCircular(ticks = 240))
check "window: more than the window width of samples were buffered",
g.bufferedCount() > 50 and g.trained > 0
let before = sideTeamCopy(g)
g.retrainFromBuffer(50)
let a = sideTeamCopy(g)
g.retrainFromBuffer(50)
let b = sideTeamCopy(g)
check "window: a rebuild is deterministic from the buffer", a == b
check "window: a rebuild actually resets + retrains the machine", a != before
check "window: the resolved-sample count keeps climbing across rebuilds",
g.trained > 0
proc testWindowRetrainKeepsDeferredState() =
## A window rebuild / re-learn must NOT touch the deferred-label queue or the
## observation ring; only `resetRoundState`/`resetLearning` may clear those.
var g = initTmHorizonGun()
g.setShift(0.0)
g.setWindow(50)
g.setRetrainConfig(50, 1)
drive(g, synthesizeCircular(ticks = 240))
let pend = g.pendingCount
let ringN = g.ringValidCount()
check "window: deferred labels exist before the rebuild", pend > 0
g.retrainFromBuffer(50)
check "window: retrain keeps the deferred label queue", g.pendingCount == pend
check "window: retrain keeps the observation ring", g.ringValidCount() == ringN
g.resetLearning()
check "window: a battle reset drops the buffered samples", g.bufferedCount() == 0
check "window: a battle reset drops the rolling accuracy ring", g.accCount == 0
proc testResetDropTriggersRelearn() =
## Change detection: with the knob on, a rolling-accuracy drop below its own
## peak fires a re-learn; with it off, nothing fires. The circular fixture
## mixes a cold start with warm tracking, so the peak/current gap is real.
var off = initTmHorizonGun()
off.setShift(0.0)
drive(off, synthesizeCircular(ticks = 240))
check "resetdrop: off by default fires nothing", off.resetDrops == 0
var g = initTmHorizonGun()
g.setShift(0.0)
g.setResetDrop(0.1)
g.setRetrainConfig(50, 1)
check "resetdrop: enabling it allocates the re-learn ring",
g.bufferCapacity() >= TMH_RESET_WINDOW_DEF
drive(g, synthesizeCircular(ticks = 240))
check "resetdrop: a rolling-accuracy drop triggers at least one re-learn",
g.resetDrops >= 1
check "resetdrop: the accuracy ring is populated", g.accCount == g.sideTotal
check "resetdrop: rolling accuracy stays a valid fraction",
g.rollingAcc(100) >= 0.0 and g.rollingAcc(100) <= 1.0
when isMainModule:
testHorizonMaths()
testHorizonBuckets()
@@ -334,6 +422,11 @@ when isMainModule:
testStaleObservationsDropped()
testColdModelEmitsNoShift()
testWarmShiftMovesAim()
testReadaptDefaultKnobs()
testRuntimeStatesKnob()
testWindowBuffersAndRebuilds()
testWindowRetrainKeepsDeferredState()
testResetDropTriggersRelearn()
if failures > 0:
echo "\n", failures, " check(s) FAILED"
quit(1)