Re-adaptation: the sliding WINDOW is a control-validated +9.3pp on late accuracy
The user's first-hand diagnosis: "once the pattern is learnt enough we hit DrussGT, but as soon as it adapts we are not fast enough to re-adapt again." The gun kept EVERY sample for the whole battle (which the user explicitly asked for), so stale evidence weighed the same as new evidence - accumulation without forgetting. TASK 1 - THE CURVE, MEASURED FIRST (prequential side accuracy, 15 rounds x 4 horizons, deciles of the eval stream): arm early% late% decay frozen (early only) 64.9 61.4 -3.5 accum (shipped) 76.4 75.3 -1.1 window N=150 84.7 84.6 -0.0 resetdrop 5pp 83.2 84.3 +1.0 **The decay IS real but modest and LOCALIZED IN THE LAST ~30% of the round** (last-third vs middle-third: frozen -7.7pp, accum -3.7, window -1.3, resetdrop -1.3). TASK 3 - WHICH FIX HELPS (late-half accuracy, shuffled control in parens): accum (shipped) 75.3 **window N=150 84.6 (51)** +9.3pp **resetdrop 5pp 84.3 (51)** +9.0pp rehearse-all (retrain, NO forgetting) 79.5 - **Sliding window: +9.3pp late (within-round), +9.7 (cross-round), +10.3 (shield).** The shuffled control stays ~50-51%, so it learns the ENEMY, not noise. - **Forgetting is the essential ingredient**: the same periodic retrain WITHOUT forgetting reaches only ~79.5%, so roughly half the gain is the retraining mechanics and half is the forgetting. - Change detection ties the window on late accuracy and gives the best decay, but at a 5pp threshold it fired 300-500 times in the offline stream - noisy. - INERTIA IS REDUNDANT WITH THE WINDOW: lower inertia helps the keep-everything model a lot (accum late 75.3 -> 83.3 at N=8) but leaves the window FLAT (84.6-84.7 at every N). Low inertia and forgetting are SUBSTITUTES, and forgetting is the robust one - `TR_TMHORIZON_NSTATES=8` is NOT the primary fix. VERDICT: SHIPPING CANDIDATE = `TR_TMHORIZON_WINDOW=150`, kept at default 0 until a live A/B confirms. HONEST CAVEATS THAT SET EXPECTATIONS: - The fixtures are OPEN-LOOP (DrussGT does not react to our bullets), so a true mid-round adaptation is NOT present; the dominant measured effect is the LEVEL gap, not the decay magnitude. The user's "it adapts" magnitude is still INFERRED. - The harness uses a STRAIGHT-LINE base while the live gun uses Pattern's prediction, and it ignores the h-tick label delay, so its absolute side accuracy (75-85%) is INFLATED: the same gun measured ~52% = chance live against DrussGT. So +9.3pp is a real ARM DELTA, not a promise that the gun now clears the ~80% accuracy wall that hits need. A live A/B must decide. Adds `measure_tm_readapt.nim` (prequential harness with the mandatory shuffled control, within-round + cross-round protocols, inertia sweep) and its captured results. test_tm_horizon 79 -> 104 (five new groups). test_tm_diag 48, test_tm_automata_diag 55, test_tm_clause_shape 66, test_rack_membership 48 pass. ModularBot compiles. .gitignore: switched from a broad `measure_*` pattern to EXPLICIT binary names. The broad rule was too blunt - it also excluded the `.txt` results file, which made `git add` refuse the whole commit twice. Sources stay tracked; binaries do not.
This commit is contained in:
@@ -18,7 +18,7 @@
|
||||
## Run with plain:
|
||||
## nim c -r common_libs/tests/test_tm_horizon.nim
|
||||
|
||||
import std/[math]
|
||||
import std/[math, os]
|
||||
import gun_harness/gun_interface
|
||||
import gun_harness/virtual_bullets
|
||||
import gun_harness/offline_range
|
||||
@@ -316,6 +316,94 @@ proc testWarmShiftMovesAim() =
|
||||
check "warm: the applied rotation is the configured -3.0 deg",
|
||||
approx(d, -3.0, 1e-6)
|
||||
|
||||
proc testReadaptDefaultKnobs() =
|
||||
## The shipped defaults must be EXACTLY today's behaviour: no buffering, no
|
||||
## change detection, compile-time state count, rolling curve off.
|
||||
var g = initTmHorizonGun()
|
||||
check "readapt: WINDOW defaults to 0 (keep everything)", g.windowN == 0
|
||||
check "readapt: RESET_DROP defaults to 0 (off)", g.resetDrop == 0.0
|
||||
check "readapt: the sample buffer is disabled by default", g.bufferCapacity() == 0
|
||||
check "readapt: NSTATES defaults to the compile-time TMH_NSTATES",
|
||||
g.nStates == TMH_NSTATES and g.headStates() == TMH_NSTATES
|
||||
check "readapt: the accuracy curve is off by default", not g.accurveEnabled
|
||||
drive(g, synthesizeCircular(ticks = 240))
|
||||
check "readapt: no reset fires when RESET_DROP is off", g.resetDrops == 0
|
||||
check "readapt: rolling accuracy is still recorded (measurement only)",
|
||||
g.rollingAcc(100) >= 0.0 and g.rollingAcc(100) <= 1.0
|
||||
|
||||
proc testRuntimeStatesKnob() =
|
||||
## Inertia is sweepable WITHOUT a rebuild via TR_TMHORIZON_NSTATES.
|
||||
putEnv("TR_TMHORIZON_NSTATES", "16")
|
||||
var g = initTmHorizonGun()
|
||||
delEnv("TR_TMHORIZON_NSTATES")
|
||||
check "readapt: NSTATES env sets the runtime state count", g.nStates == 16
|
||||
check "readapt: both heads get the runtime state count",
|
||||
g.headStates() == 16
|
||||
var d = initTmHorizonGun()
|
||||
check "readapt: an unset env cannot move the state count",
|
||||
d.nStates == TMH_NSTATES
|
||||
|
||||
proc testWindowBuffersAndRebuilds() =
|
||||
## Sliding-window mode must allocate a bounded ring, keep buffering resolved
|
||||
## samples, and rebuild both heads deterministically from that ring.
|
||||
var g = initTmHorizonGun()
|
||||
g.setShift(0.0)
|
||||
g.setWindow(50)
|
||||
g.setRetrainConfig(50, 1)
|
||||
check "window: the ring is allocated to the window width", g.bufferCapacity() == 50
|
||||
drive(g, synthesizeCircular(ticks = 240))
|
||||
check "window: more than the window width of samples were buffered",
|
||||
g.bufferedCount() > 50 and g.trained > 0
|
||||
let before = sideTeamCopy(g)
|
||||
g.retrainFromBuffer(50)
|
||||
let a = sideTeamCopy(g)
|
||||
g.retrainFromBuffer(50)
|
||||
let b = sideTeamCopy(g)
|
||||
check "window: a rebuild is deterministic from the buffer", a == b
|
||||
check "window: a rebuild actually resets + retrains the machine", a != before
|
||||
check "window: the resolved-sample count keeps climbing across rebuilds",
|
||||
g.trained > 0
|
||||
|
||||
proc testWindowRetrainKeepsDeferredState() =
|
||||
## A window rebuild / re-learn must NOT touch the deferred-label queue or the
|
||||
## observation ring; only `resetRoundState`/`resetLearning` may clear those.
|
||||
var g = initTmHorizonGun()
|
||||
g.setShift(0.0)
|
||||
g.setWindow(50)
|
||||
g.setRetrainConfig(50, 1)
|
||||
drive(g, synthesizeCircular(ticks = 240))
|
||||
let pend = g.pendingCount
|
||||
let ringN = g.ringValidCount()
|
||||
check "window: deferred labels exist before the rebuild", pend > 0
|
||||
g.retrainFromBuffer(50)
|
||||
check "window: retrain keeps the deferred label queue", g.pendingCount == pend
|
||||
check "window: retrain keeps the observation ring", g.ringValidCount() == ringN
|
||||
g.resetLearning()
|
||||
check "window: a battle reset drops the buffered samples", g.bufferedCount() == 0
|
||||
check "window: a battle reset drops the rolling accuracy ring", g.accCount == 0
|
||||
|
||||
proc testResetDropTriggersRelearn() =
|
||||
## Change detection: with the knob on, a rolling-accuracy drop below its own
|
||||
## peak fires a re-learn; with it off, nothing fires. The circular fixture
|
||||
## mixes a cold start with warm tracking, so the peak/current gap is real.
|
||||
var off = initTmHorizonGun()
|
||||
off.setShift(0.0)
|
||||
drive(off, synthesizeCircular(ticks = 240))
|
||||
check "resetdrop: off by default fires nothing", off.resetDrops == 0
|
||||
|
||||
var g = initTmHorizonGun()
|
||||
g.setShift(0.0)
|
||||
g.setResetDrop(0.1)
|
||||
g.setRetrainConfig(50, 1)
|
||||
check "resetdrop: enabling it allocates the re-learn ring",
|
||||
g.bufferCapacity() >= TMH_RESET_WINDOW_DEF
|
||||
drive(g, synthesizeCircular(ticks = 240))
|
||||
check "resetdrop: a rolling-accuracy drop triggers at least one re-learn",
|
||||
g.resetDrops >= 1
|
||||
check "resetdrop: the accuracy ring is populated", g.accCount == g.sideTotal
|
||||
check "resetdrop: rolling accuracy stays a valid fraction",
|
||||
g.rollingAcc(100) >= 0.0 and g.rollingAcc(100) <= 1.0
|
||||
|
||||
when isMainModule:
|
||||
testHorizonMaths()
|
||||
testHorizonBuckets()
|
||||
@@ -334,6 +422,11 @@ when isMainModule:
|
||||
testStaleObservationsDropped()
|
||||
testColdModelEmitsNoShift()
|
||||
testWarmShiftMovesAim()
|
||||
testReadaptDefaultKnobs()
|
||||
testRuntimeStatesKnob()
|
||||
testWindowBuffersAndRebuilds()
|
||||
testWindowRetrainKeepsDeferredState()
|
||||
testResetDropTriggersRelearn()
|
||||
if failures > 0:
|
||||
echo "\n", failures, " check(s) FAILED"
|
||||
quit(1)
|
||||
|
||||
Reference in New Issue
Block a user