69debbe347
The user's first-hand diagnosis: "once the pattern is learnt enough we hit DrussGT, but as soon as it adapts we are not fast enough to re-adapt again." The gun kept EVERY sample for the whole battle (which the user explicitly asked for), so stale evidence weighed the same as new evidence - accumulation without forgetting. TASK 1 - THE CURVE, MEASURED FIRST (prequential side accuracy, 15 rounds x 4 horizons, deciles of the eval stream): arm early% late% decay frozen (early only) 64.9 61.4 -3.5 accum (shipped) 76.4 75.3 -1.1 window N=150 84.7 84.6 -0.0 resetdrop 5pp 83.2 84.3 +1.0 **The decay IS real but modest and LOCALIZED IN THE LAST ~30% of the round** (last-third vs middle-third: frozen -7.7pp, accum -3.7, window -1.3, resetdrop -1.3). TASK 3 - WHICH FIX HELPS (late-half accuracy, shuffled control in parens): accum (shipped) 75.3 **window N=150 84.6 (51)** +9.3pp **resetdrop 5pp 84.3 (51)** +9.0pp rehearse-all (retrain, NO forgetting) 79.5 - **Sliding window: +9.3pp late (within-round), +9.7 (cross-round), +10.3 (shield).** The shuffled control stays ~50-51%, so it learns the ENEMY, not noise. - **Forgetting is the essential ingredient**: the same periodic retrain WITHOUT forgetting reaches only ~79.5%, so roughly half the gain is the retraining mechanics and half is the forgetting. - Change detection ties the window on late accuracy and gives the best decay, but at a 5pp threshold it fired 300-500 times in the offline stream - noisy. - INERTIA IS REDUNDANT WITH THE WINDOW: lower inertia helps the keep-everything model a lot (accum late 75.3 -> 83.3 at N=8) but leaves the window FLAT (84.6-84.7 at every N). Low inertia and forgetting are SUBSTITUTES, and forgetting is the robust one - `TR_TMHORIZON_NSTATES=8` is NOT the primary fix. VERDICT: SHIPPING CANDIDATE = `TR_TMHORIZON_WINDOW=150`, kept at default 0 until a live A/B confirms. HONEST CAVEATS THAT SET EXPECTATIONS: - The fixtures are OPEN-LOOP (DrussGT does not react to our bullets), so a true mid-round adaptation is NOT present; the dominant measured effect is the LEVEL gap, not the decay magnitude. The user's "it adapts" magnitude is still INFERRED. - The harness uses a STRAIGHT-LINE base while the live gun uses Pattern's prediction, and it ignores the h-tick label delay, so its absolute side accuracy (75-85%) is INFLATED: the same gun measured ~52% = chance live against DrussGT. So +9.3pp is a real ARM DELTA, not a promise that the gun now clears the ~80% accuracy wall that hits need. A live A/B must decide. Adds `measure_tm_readapt.nim` (prequential harness with the mandatory shuffled control, within-round + cross-round protocols, inertia sweep) and its captured results. test_tm_horizon 79 -> 104 (five new groups). test_tm_diag 48, test_tm_automata_diag 55, test_tm_clause_shape 66, test_rack_membership 48 pass. ModularBot compiles. .gitignore: switched from a broad `measure_*` pattern to EXPLICIT binary names. The broad rule was too blunt - it also excluded the `.txt` results file, which made `git add` refuse the whole commit twice. Sources stay tracked; binaries do not.
69 lines
1.6 KiB
Plaintext
69 lines
1.6 KiB
Plaintext
# Compiled output
|
|
*.out
|
|
|
|
# Python cache
|
|
__pycache__/
|
|
*.pyc
|
|
*.pyo
|
|
|
|
# Log files
|
|
*.log
|
|
|
|
# Training logs
|
|
*.jsonl
|
|
|
|
# Test binaries (but keep .nim source files)
|
|
**/tests/test_*
|
|
!**/tests/test_*.nim
|
|
!**/tests/test_*.nims
|
|
|
|
# User-specific dev environment
|
|
.envrc
|
|
devbox.json
|
|
devbox.lock
|
|
|
|
# Compiled bot binaries — all builds go to *_garage/out/
|
|
*_garage/out/
|
|
|
|
# Java compiled classes
|
|
*.class
|
|
|
|
# Log and snapshot directories
|
|
**/logs/
|
|
**/snapshots/
|
|
|
|
# Git worktrees
|
|
worktrees/
|
|
|
|
# Nimble puts the bot binary at the garage root (out/ is ignored, this was not)
|
|
*_garage/ModularBot
|
|
|
|
# Test fixtures are data, not logs - the *.jsonl rule above was written for
|
|
# training logs and silently excluded the entire gun-range fixture set.
|
|
!tools/fixtures/**/*.jsonl
|
|
|
|
# Compiled test harnesses have no extension; the test_* rule misses them
|
|
common_libs/tests/tm_measure
|
|
common_libs/tests/run_range
|
|
common_libs/tests/gen_synthetic_fixtures
|
|
common_libs/tests/acceptance_offline_vs_online
|
|
common_libs/tests/measure_power_policy
|
|
common_libs/tests/test_tsetlin_gun
|
|
common_libs/tests/test_tsetlin_live
|
|
common_libs/tests/test_tm_pattern_learning
|
|
|
|
# measurement tool binaries (no extension); sources are measure_*.nim
|
|
tr_bots/
|
|
|
|
# Measurement-tool binaries (extensionless; their .nim/.py/.txt sources stay tracked)
|
|
common_libs/tests/measure_tm_readapt
|
|
common_libs/tests/measure_tm_pattern_cost
|
|
common_libs/tests/measure_cornering_guns
|
|
common_libs/tests/measure_vbullet_admit_gate
|
|
common_libs/tests/measure_tm_miss_shrink
|
|
common_libs/tests/audit_wave_pairing
|
|
common_libs/tests/compare_pairing
|
|
common_libs/tests/diag_synthetic
|
|
common_libs/tests/diag_automata_validation
|
|
common_libs/tests/diag_tm_pattern_offline
|