Files
SirRoboGarage/common_libs/tests/measure_tm_readapt_results.txt
T
SirStone 69debbe347 Re-adaptation: the sliding WINDOW is a control-validated +9.3pp on late accuracy
The user's first-hand diagnosis: "once the pattern is learnt enough we hit DrussGT,
but as soon as it adapts we are not fast enough to re-adapt again." The gun kept
EVERY sample for the whole battle (which the user explicitly asked for), so stale
evidence weighed the same as new evidence - accumulation without forgetting.

TASK 1 - THE CURVE, MEASURED FIRST (prequential side accuracy, 15 rounds x 4
horizons, deciles of the eval stream):
  arm                 early%  late%  decay
  frozen (early only)   64.9   61.4   -3.5
  accum (shipped)       76.4   75.3   -1.1
  window N=150          84.7   84.6   -0.0
  resetdrop 5pp         83.2   84.3   +1.0
**The decay IS real but modest and LOCALIZED IN THE LAST ~30% of the round**
(last-third vs middle-third: frozen -7.7pp, accum -3.7, window -1.3, resetdrop -1.3).

TASK 3 - WHICH FIX HELPS (late-half accuracy, shuffled control in parens):
  accum (shipped)                       75.3
  **window N=150                       84.6  (51)**   +9.3pp
  **resetdrop 5pp                      84.3  (51)**   +9.0pp
  rehearse-all (retrain, NO forgetting) 79.5
- **Sliding window: +9.3pp late (within-round), +9.7 (cross-round), +10.3 (shield).**
  The shuffled control stays ~50-51%, so it learns the ENEMY, not noise.
- **Forgetting is the essential ingredient**: the same periodic retrain WITHOUT
  forgetting reaches only ~79.5%, so roughly half the gain is the retraining
  mechanics and half is the forgetting.
- Change detection ties the window on late accuracy and gives the best decay, but
  at a 5pp threshold it fired 300-500 times in the offline stream - noisy.
- INERTIA IS REDUNDANT WITH THE WINDOW: lower inertia helps the keep-everything
  model a lot (accum late 75.3 -> 83.3 at N=8) but leaves the window FLAT
  (84.6-84.7 at every N). Low inertia and forgetting are SUBSTITUTES, and
  forgetting is the robust one - `TR_TMHORIZON_NSTATES=8` is NOT the primary fix.

VERDICT: SHIPPING CANDIDATE = `TR_TMHORIZON_WINDOW=150`, kept at default 0 until a
live A/B confirms.

HONEST CAVEATS THAT SET EXPECTATIONS:
- The fixtures are OPEN-LOOP (DrussGT does not react to our bullets), so a true
  mid-round adaptation is NOT present; the dominant measured effect is the LEVEL
  gap, not the decay magnitude. The user's "it adapts" magnitude is still INFERRED.
- The harness uses a STRAIGHT-LINE base while the live gun uses Pattern's
  prediction, and it ignores the h-tick label delay, so its absolute side accuracy
  (75-85%) is INFLATED: the same gun measured ~52% = chance live against DrussGT.
  So +9.3pp is a real ARM DELTA, not a promise that the gun now clears the ~80%
  accuracy wall that hits need. A live A/B must decide.

Adds `measure_tm_readapt.nim` (prequential harness with the mandatory shuffled
control, within-round + cross-round protocols, inertia sweep) and its captured
results. test_tm_horizon 79 -> 104 (five new groups). test_tm_diag 48,
test_tm_automata_diag 55, test_tm_clause_shape 66, test_rack_membership 48 pass.
ModularBot compiles.

.gitignore: switched from a broad `measure_*` pattern to EXPLICIT binary names.
The broad rule was too blunt - it also excluded the `.txt` results file, which made
`git add` refuse the whole commit twice. Sources stay tracked; binaries do not.
2026-09-23 00:29:37 +02:00

193 lines
17 KiB
Plaintext

TM HORIZON RE-ADAPTATION — offline measurement results
Generated by: nim c -r -d:release --path:common_libs common_libs/tests/measure_tm_readapt.nim
Fixtures: tr_drussgt_vs_modularbot.jsonl (15 rounds), tr_drussgt_vs_modularbot_shield.jsonl (10 rounds)
NOTE: the primary-fixture by-horizon view omits the expensive rehearse-all arm (shown in the pooled table).
/home/davide/Projects/SirRoboGarage/common_libs/tests/measure_tm_readapt.nim(34, 50) Warning: imported and not used: 'algorithm' [UnusedImport]
====================================================================================================
TM HORIZON RE-ADAPTATION (offline, prequential side accuracy)
====================================================================================================
fixtures : tr_drussgt_vs_modularbot.jsonl
TM : 40 clauses, 64 states, s=3.0, warmEpochs=5
arms : frozen(none) | accum(keep all) | window(N=150, retrain 50) | resetdrop(drop<5.0pp -> retrain last 150)
protocol : within-round warmFrac=0.40; cross-round (warm R, stream R+1)
metric : prequential; early=1st half of eval, late=2nd half; decay=late-early
control : labels shuffled within each set (mandatory integrity check)
====================================================================================================
## FIXTURE tr_drussgt_vs_modularbot.jsonl: 15 rounds
## tr_drussgt_vs_modularbot.jsonl protocol=within (pooled over rounds/horizons)
h arm N early% late% decay total% curve(deciles of eval) resets
15 frozen 46842 64.9 61.4 -3.5 63.1 [ 61. 65. 63. 69. 66. 68. 58. 63. 59. 59.] resets=0
15 accum 46842 76.4 75.3 -1.1 75.8 [ 74. 78. 73. 77. 80. 76. 71. 75. 76. 78.] resets=0
15 accum-shuf 46842 50.9 50.7 -0.2 50.8 [ 50. 50. 51. 52. 51. 51. 51. 50. 51. 50.] resets=0
15 window 46842 84.7 84.6 -0.0 84.6 [ 78. 86. 86. 86. 86. 85. 85. 85. 84. 84.] resets=0
15 window-shuf 46842 50.5 51.0 +0.4 50.7 [ 50. 51. 50. 50. 52. 51. 52. 52. 50. 50.] resets=0
15 resetdrop 46842 83.2 84.3 +1.0 83.7 [ 74. 84. 86. 87. 86. 84. 84. 86. 83. 84.] resets=301
15 resetdrop-shuf 46842 50.1 50.6 +0.5 50.4 [ 50. 50. 50. 50. 50. 51. 51. 50. 51. 50.] resets=338
15 rehearse-all 46842 81.2 79.5 -1.6 80.4 [ 77. 83. 81. 83. 83. 79. 77. 80. 80. 81.] resets=0
15 rehearse-all-shuf 46842 50.6 50.7 +0.1 50.6 [ 50. 50. 51. 51. 51. 51. 51. 51. 50. 51.] resets=0
-- by horizon (accum / window / resetdrop, TRUE labels only) --
15 frozen 11778 63.0 59.1 -3.8 61.0 [ 61. 63. 60. 66. 65. 62. 60. 60. 56. 57.] resets=0
15 accum 11778 74.8 73.9 -0.9 74.4 [ 73. 78. 69. 77. 77. 75. 71. 73. 73. 76.] resets=0
15 window 11778 83.6 83.7 +0.1 83.6 [ 78. 86. 83. 85. 86. 84. 84. 84. 83. 82.] resets=0
15 resetdrop 11778 82.2 83.1 +0.9 82.7 [ 73. 83. 84. 86. 85. 83. 83. 85. 82. 82.] resets=77
20 frozen 11733 64.2 59.9 -4.3 62.0 [ 61. 65. 61. 68. 65. 68. 58. 60. 56. 58.] resets=0
20 accum 11733 76.4 74.4 -2.0 75.4 [ 74. 78. 73. 77. 80. 75. 72. 75. 74. 75.] resets=0
20 window 11733 84.7 84.3 -0.5 84.5 [ 78. 87. 87. 86. 86. 84. 86. 85. 83. 83.] resets=0
20 resetdrop 11733 83.0 83.9 +0.8 83.4 [ 74. 83. 87. 86. 86. 84. 85. 86. 81. 84.] resets=75
25 frozen 11688 65.0 63.4 -1.7 64.2 [ 61. 64. 66. 69. 66. 71. 57. 67. 63. 60.] resets=0
25 accum 11688 77.1 75.8 -1.3 76.4 [ 74. 77. 75. 78. 81. 77. 71. 74. 77. 80.] resets=0
25 window 11688 85.3 84.7 -0.6 85.0 [ 79. 87. 88. 87. 87. 86. 85. 85. 83. 85.] resets=0
25 resetdrop 11688 83.3 84.2 +1.0 83.8 [ 74. 83. 85. 87. 87. 84. 84. 85. 84. 85.] resets=72
30 frozen 11643 67.3 63.2 -4.0 65.2 [ 63. 69. 64. 72. 69. 70. 57. 67. 61. 61.] resets=0
30 accum 11643 77.4 77.0 -0.4 77.2 [ 75. 80. 74. 78. 81. 77. 69. 78. 80. 81.] resets=0
30 window 11643 85.0 85.9 +0.9 85.4 [ 78. 85. 87. 87. 87. 87. 85. 87. 85. 86.] resets=0
30 resetdrop 11643 84.4 85.8 +1.5 85.1 [ 75. 84. 88. 88. 87. 85. 85. 88. 85. 86.] resets=77
## tr_drussgt_vs_modularbot.jsonl protocol=cross (pooled over rounds/horizons)
h arm N early% late% decay total% curve(deciles of eval) resets
15 frozen 73164 65.4 68.8 +3.4 67.1 [ 58. 65. 67. 69. 67. 72. 70. 66. 71. 65.] resets=0
15 accum 73164 73.7 75.6 +2.0 74.6 [ 71. 73. 76. 74. 75. 76. 77. 75. 76. 75.] resets=0
15 accum-shuf 73164 49.6 50.6 +1.0 50.1 [ 49. 50. 50. 49. 49. 50. 51. 51. 51. 51.] resets=0
15 window 73164 85.2 85.3 +0.2 85.3 [ 82. 87. 87. 84. 86. 86. 87. 85. 86. 84.] resets=0
15 window-shuf 73164 50.3 50.2 -0.1 50.3 [ 50. 50. 51. 50. 50. 49. 50. 51. 50. 51.] resets=0
15 resetdrop 73164 82.9 84.7 +1.8 83.8 [ 74. 86. 86. 83. 86. 85. 86. 84. 85. 83.] resets=466
15 resetdrop-shuf 73164 50.5 50.7 +0.2 50.6 [ 50. 51. 51. 50. 51. 49. 51. 51. 51. 51.] resets=528
15 rehearse-all 73164 77.5 78.7 +1.3 78.1 [ 74. 76. 80. 76. 81. 79. 80. 77. 79. 79.] resets=0
15 rehearse-all-shuf 73164 50.3 50.7 +0.4 50.5 [ 49. 51. 51. 51. 50. 50. 51. 51. 51. 51.] resets=0
-- by horizon (accum / window / resetdrop, TRUE labels only) --
15 frozen 18396 64.2 66.5 +2.3 65.3 [ 56. 62. 67. 68. 68. 70. 68. 66. 68. 60.] resets=0
15 accum 18396 72.0 74.0 +2.0 73.0 [ 68. 70. 74. 73. 75. 73. 76. 75. 74. 71.] resets=0
15 window 18396 84.0 84.4 +0.4 84.2 [ 81. 85. 85. 84. 85. 85. 86. 84. 85. 82.] resets=0
15 resetdrop 18396 81.5 83.4 +2.0 82.4 [ 71. 84. 84. 83. 85. 84. 86. 84. 83. 81.] resets=121
20 frozen 18326 64.4 68.8 +4.4 66.6 [ 56. 62. 67. 69. 67. 72. 69. 67. 71. 64.] resets=0
20 accum 18326 73.5 75.8 +2.3 74.6 [ 69. 72. 75. 76. 76. 76. 76. 78. 76. 73.] resets=0
20 window 18326 85.1 85.5 +0.3 85.3 [ 82. 85. 87. 85. 87. 87. 86. 85. 86. 83.] resets=0
20 resetdrop 18326 82.7 84.7 +2.0 83.7 [ 73. 85. 86. 84. 86. 86. 86. 84. 85. 83.] resets=116
25 frozen 18256 66.7 69.8 +3.1 68.3 [ 60. 66. 68. 71. 68. 73. 71. 66. 74. 66.] resets=0
25 accum 18256 73.9 76.1 +2.2 75.0 [ 71. 74. 75. 74. 75. 77. 77. 73. 76. 77.] resets=0
25 window 18256 85.3 85.6 +0.3 85.5 [ 83. 87. 88. 84. 85. 85. 87. 86. 85. 84.] resets=0
25 resetdrop 18256 83.6 85.3 +1.7 84.4 [ 74. 86. 87. 84. 87. 85. 87. 84. 85. 84.] resets=115
30 frozen 18186 66.2 70.0 +3.8 68.1 [ 61. 68. 67. 69. 66. 72. 71. 66. 72. 69.] resets=0
30 accum 18186 75.2 76.7 +1.5 75.9 [ 74. 76. 78. 74. 74. 78. 77. 73. 78. 78.] resets=0
30 window 18186 86.3 86.0 -0.3 86.1 [ 84. 89. 88. 84. 87. 86. 87. 84. 87. 86.] resets=0
30 resetdrop 18186 84.0 85.4 +1.4 84.7 [ 77. 88. 86. 83. 86. 86. 86. 84. 87. 85.] resets=114
## INERTIA SWEEP (tr_drussgt_vs_modularbot.jsonl, within protocol, pooled over rounds x horizons)
states arm early%/late% (TRUE labels)
8 accum 84.3/ 83.3
8 window 85.3/ 84.7
16 accum 79.5/ 79.3
16 window 84.6/ 84.6
64 accum 76.4/ 75.3
64 window 84.7/ 84.6
256 accum 76.4/ 75.3
256 window 84.7/ 84.6
====================================================================================================
## INTEGRITY: shuffled-label controls should hover at chance (~50%)
If a fix only improves the shuffled rows, it is NOT learning the enemy.
rehearse-all = same periodic full retrain as window but WITHOUT forgetting
(isolates the effect of the sliding buffer from the effect of retraining).
## VERDICT INPUTS: late-half accuracy, TRUE vs shuffled (pp)
====================================================================================================
########################################################################
/home/davide/Projects/SirRoboGarage/common_libs/tests/measure_tm_readapt.nim(34, 50) Warning: imported and not used: 'algorithm' [UnusedImport]
====================================================================================================
TM HORIZON RE-ADAPTATION (offline, prequential side accuracy)
====================================================================================================
fixtures : tr_drussgt_vs_modularbot_shield.jsonl
TM : 40 clauses, 64 states, s=3.0, warmEpochs=5
arms : frozen(none) | accum(keep all) | window(N=150, retrain 50) | resetdrop(drop<5.0pp -> retrain last 150)
protocol : within-round warmFrac=0.40; cross-round (warm R, stream R+1)
metric : prequential; early=1st half of eval, late=2nd half; decay=late-early
control : labels shuffled within each set (mandatory integrity check)
====================================================================================================
## FIXTURE tr_drussgt_vs_modularbot_shield.jsonl: 10 rounds
## tr_drussgt_vs_modularbot_shield.jsonl protocol=within (pooled over rounds/horizons)
h arm N early% late% decay total% curve(deciles of eval) resets
15 frozen 28614 64.4 63.0 -1.5 63.7 [ 67. 58. 65. 72. 59. 72. 67. 57. 61. 58.] resets=0
15 accum 28614 75.3 75.9 +0.7 75.6 [ 74. 71. 75. 79. 78. 78. 75. 68. 79. 79.] resets=0
15 accum-shuf 28614 50.8 50.9 +0.1 50.8 [ 51. 51. 51. 50. 52. 50. 52. 51. 50. 51.] resets=0
15 window 28614 85.5 86.2 +0.7 85.9 [ 81. 87. 87. 85. 88. 86. 86. 87. 86. 87.] resets=0
15 window-shuf 28614 50.9 50.6 -0.3 50.7 [ 51. 51. 50. 51. 51. 49. 52. 50. 50. 52.] resets=0
15 resetdrop 28614 83.2 85.3 +2.1 84.3 [ 74. 82. 87. 84. 89. 85. 83. 87. 85. 86.] resets=177
15 resetdrop-shuf 28614 51.4 50.5 -0.9 51.0 [ 51. 52. 51. 51. 52. 50. 51. 51. 50. 50.] resets=202
15 rehearse-all 28614 82.0 80.2 -1.7 81.1 [ 78. 83. 82. 82. 85. 80. 79. 78. 80. 83.] resets=0
15 rehearse-all-shuf 28614 50.4 50.7 +0.3 50.6 [ 51. 51. 51. 50. 50. 49. 51. 50. 52. 51.] resets=0
-- by horizon (accum / window / resetdrop, TRUE labels only) --
15 frozen 7185 62.4 62.8 +0.4 62.6 [ 61. 58. 63. 74. 56. 68. 69. 57. 61. 59.] resets=0
15 accum 7185 75.2 75.8 +0.6 75.5 [ 71. 73. 74. 79. 79. 75. 74. 73. 78. 78.] resets=0
15 window 7185 84.2 85.3 +1.1 84.7 [ 80. 84. 87. 84. 86. 85. 84. 87. 86. 84.] resets=0
15 resetdrop 7185 82.0 83.6 +1.6 82.8 [ 71. 81. 86. 85. 87. 83. 83. 86. 83. 83.] resets=45
20 frozen 7164 62.7 62.5 -0.2 62.6 [ 63. 59. 62. 69. 60. 71. 67. 58. 58. 58.] resets=0
20 accum 7164 74.4 75.2 +0.9 74.8 [ 71. 73. 74. 78. 76. 77. 76. 65. 77. 80.] resets=0
20 window 7164 85.4 85.7 +0.3 85.6 [ 80. 87. 86. 85. 89. 85. 85. 87. 84. 88.] resets=0
20 resetdrop 7164 83.5 85.3 +1.8 84.4 [ 71. 84. 88. 85. 90. 85. 83. 88. 84. 86.] resets=42
25 frozen 7143 65.7 62.2 -3.5 64.0 [ 72. 56. 66. 73. 61. 73. 66. 56. 61. 56.] resets=0
25 accum 7143 76.1 75.8 -0.3 76.0 [ 78. 69. 75. 80. 79. 79. 76. 65. 81. 78.] resets=0
25 window 7143 86.1 86.9 +0.8 86.5 [ 81. 88. 88. 85. 89. 87. 86. 88. 86. 88.] resets=0
25 resetdrop 7143 84.1 85.7 +1.6 84.9 [ 78. 83. 86. 84. 89. 86. 83. 87. 86. 87.] resets=46
30 frozen 7122 66.9 64.4 -2.5 65.6 [ 74. 60. 70. 71. 60. 75. 65. 56. 65. 60.] resets=0
30 accum 7122 75.4 77.0 +1.6 76.2 [ 77. 68. 75. 78. 78. 81. 75. 69. 80. 80.] resets=0
30 window 7122 86.5 87.0 +0.6 86.7 [ 82. 89. 87. 86. 89. 87. 87. 87. 86. 88.] resets=0
30 resetdrop 7122 83.3 86.8 +3.5 85.0 [ 77. 81. 87. 83. 89. 86. 85. 87. 86. 89.] resets=44
## tr_drussgt_vs_modularbot_shield.jsonl protocol=cross (pooled over rounds/horizons)
h arm N early% late% decay total% curve(deciles of eval) resets
15 frozen 46902 59.6 62.9 +3.3 61.2 [ 57. 63. 63. 57. 59. 64. 65. 65. 58. 62.] resets=0
15 accum 46902 72.5 74.7 +2.2 73.6 [ 73. 71. 73. 75. 71. 73. 75. 77. 71. 77.] resets=0
15 accum-shuf 46902 50.3 50.4 +0.1 50.4 [ 50. 50. 51. 51. 48. 50. 51. 50. 50. 51.] resets=0
15 window 46902 85.4 86.2 +0.9 85.8 [ 82. 87. 85. 87. 86. 87. 87. 85. 85. 86.] resets=0
15 window-shuf 46902 50.4 50.2 -0.2 50.3 [ 50. 51. 51. 51. 49. 51. 50. 50. 50. 51.] resets=0
15 resetdrop 46902 83.7 85.8 +2.1 84.8 [ 76. 86. 85. 87. 86. 86. 87. 85. 85. 86.] resets=294
15 resetdrop-shuf 46902 50.6 50.3 -0.3 50.4 [ 50. 51. 50. 50. 51. 51. 49. 50. 50. 51.] resets=351
15 rehearse-all 46902 78.4 79.5 +1.1 78.9 [ 77. 77. 79. 81. 78. 79. 83. 78. 77. 81.] resets=0
15 rehearse-all-shuf 46902 50.2 50.9 +0.7 50.5 [ 50. 52. 50. 50. 49. 50. 51. 51. 51. 51.] resets=0
-- by horizon (accum / window / resetdrop, TRUE labels only) --
15 frozen 11793 58.2 61.7 +3.4 60.0 [ 55. 62. 62. 55. 58. 64. 63. 64. 58. 60.] resets=0
15 accum 11793 71.9 73.2 +1.3 72.6 [ 71. 70. 73. 76. 70. 71. 74. 77. 70. 75.] resets=0
15 window 11793 84.0 84.6 +0.6 84.3 [ 82. 84. 84. 86. 84. 86. 86. 84. 83. 84.] resets=0
15 resetdrop 11793 82.1 84.3 +2.2 83.2 [ 74. 85. 83. 85. 84. 85. 86. 83. 84. 83.] resets=76
20 frozen 11748 59.4 63.3 +3.9 61.3 [ 57. 62. 63. 57. 58. 64. 65. 67. 59. 62.] resets=0
20 accum 11748 72.6 74.7 +2.2 73.6 [ 74. 71. 73. 75. 70. 74. 74. 78. 69. 78.] resets=0
20 window 11748 85.7 86.1 +0.4 85.9 [ 82. 87. 85. 88. 87. 88. 87. 85. 84. 86.] resets=0
20 resetdrop 11748 83.3 85.5 +2.3 84.4 [ 77. 85. 84. 86. 85. 87. 85. 85. 84. 86.] resets=74
25 frozen 11703 60.8 63.2 +2.4 62.0 [ 57. 65. 64. 57. 61. 64. 67. 63. 58. 63.] resets=0
25 accum 11703 73.2 74.8 +1.7 74.0 [ 72. 73. 74. 76. 71. 75. 75. 78. 69. 77.] resets=0
25 window 11703 85.9 87.1 +1.2 86.5 [ 82. 87. 86. 88. 87. 88. 89. 85. 87. 87.] resets=0
25 resetdrop 11703 84.6 86.5 +1.9 85.5 [ 75. 88. 86. 88. 86. 87. 87. 86. 86. 86.] resets=72
30 frozen 11658 59.8 63.5 +3.7 61.7 [ 59. 62. 61. 57. 60. 63. 66. 65. 59. 64.] resets=0
30 accum 11658 72.3 75.9 +3.6 74.1 [ 74. 71. 74. 72. 71. 74. 77. 77. 74. 78.] resets=0
30 window 11658 86.0 87.2 +1.3 86.6 [ 82. 89. 85. 87. 88. 88. 89. 87. 87. 86.] resets=0
30 resetdrop 11658 84.9 87.0 +2.1 85.9 [ 77. 88. 85. 87. 88. 87. 88. 86. 87. 88.] resets=72
## INERTIA SWEEP (tr_drussgt_vs_modularbot_shield.jsonl, within protocol, pooled over rounds x horizons)
states arm early%/late% (TRUE labels)
8 accum 85.0/ 84.4
8 window 86.0/ 86.2
16 accum 79.5/ 79.8
16 window 85.5/ 86.2
64 accum 75.3/ 75.9
64 window 85.5/ 86.2
256 accum 75.3/ 75.9
256 window 85.5/ 86.2
====================================================================================================
## INTEGRITY: shuffled-label controls should hover at chance (~50%)
If a fix only improves the shuffled rows, it is NOT learning the enemy.
rehearse-all = same periodic full retrain as window but WITHOUT forgetting
(isolates the effect of the sliding buffer from the effect of retraining).
## VERDICT INPUTS: late-half accuracy, TRUE vs shuffled (pp)
====================================================================================================