docs(research): TM learning tracks + note that the TM-DEB source was deleted

tm-learning-tracks.md covers three things, all marked [FACT]/[INFERENCE]/
[UNKNOWN]:
- Section A: the Tsetlin gun's label is measured against the wrong baseline.
  predX = linearX + cx, so rx = actual - predX = delta - cx, and inside
  tmLearnOne error = residual - predicted = (delta - cx) - cx = delta - 2cx.
  The fixed point is cx = delta/2 -- HALF the correction needed, even with
  perfect Granmo feedback. Fix: store linearX/linearY in TmTrace and train
  on delta. Also: hits zero the label instead of carrying their true
  residual, and the per-clause step is magnitude-blind.
- Section B: what a TM is actually good at (AND-clauses over binary
  literals, readable output) and why this repo suits it -- the gun already
  builds an 83-bit x 10-frame Gray-coded window (870 bits). Includes a
  falsifiable known-rule benchmark proposal.
- Section C: delayed-reward learning belongs to the MOVEMENT layer, not the
  gun. The gun's outcome is delayed but exactly pairable via
  (fireTick, powerBin), so its effective lambda is 1 and discounting would
  only destroy information.

Also records that docs/papers/tm-deb-paper.pdf was deleted by the user as
AI-generated and unverifiable, while the Granmo-based feedback diff in
tm-deb-assessment.md stands on its own.
This commit is contained in:
2026-09-20 23:28:31 +02:00
parent 669f9acd41
commit 54e9757567
2 changed files with 264 additions and 0 deletions
+12
View File
@@ -1,5 +1,17 @@
# TM-DEB Paper Assessment — Adopt as a fix for the Tsetlin gun?
> **Editorial note — source document deleted.** The paper this file reviews,
> `docs/papers/tm-deb-paper.pdf`, was **deleted by the user** on the grounds that
> it is AI-generated and unverifiable (the concrete evidence is catalogued in §2:
> self-identifying "Generated by Gemini Notebook" provenance, a misattributed
> reference, an algorithm core rendered as literal `%`, and "Expected" rather
> than measured results). This assessment file is **kept** because its actionable
> content — the Granmo Table 2/3 feedback diff (§6) and the prioritised fix list
> (§6.5) — stands on its own from the *verified* sources
> (`docs/papers/1804.01508v15.pdf` and `common_libs/guns/tsetlin.nim`) and does
> not depend on the deleted document. Citations to the deleted PDF are retained
> only as the record of the review that was performed.
**Issue:** N/A (read-only review requested by orchestrator)
**Date:** 2026-09-20
**Paper under review:** `docs/papers/tm-deb-paper.pdf` — "Temporal Credit Assignment in Tsetlin Machines via Discounted Eligibility Experience Buffers (TM-DEB)"
+252
View File
@@ -0,0 +1,252 @@
# Tsetlin Machine learning tracks — target shape, high-level patterns, delayed reward
**Date:** 2026-09-20
**Scope:** Design notes only. No `.nim` file was edited, nothing was built, no battle was run, no `/tmp` artifact was touched. Every claim is tagged by its evidence class.
**Author's method:** Read `common_libs/guns/tsetlin.nim`, `common_libs/gun_harness/gun_interface.nim`, `common_libs/gun_harness/virtual_bullets.nim`, the `common_libs/movements/*` modules, `ModularBot_garage/src/ModularBot.nim`, `BNNBot_garage/src/{binary_encoding,tsetlin_predictor,wisard_predictor}.nim`, `SAC_LSTM_Bot_garage/src/SAC_LSTM_Bot/rewards.nim`, `docs/rl-algorithm-choice.md`, `docs/research/tm-deb-assessment.md`, and the installed `robocode_tankroyale_botapi` `schemas.nim` / `bot.nim`.
## Evidence legend
- **[FACT]** — verified by reading the cited code or paper at the cited location.
- **[INFERENCE]** — reasoned from facts; not measured.
- **[UNKNOWN]** — needs an experiment; no claim is made.
> **Provenance caveat.** `docs/papers/tm-deb-paper.pdf` (TM-DEB) was deleted by the user as AI-generated and unverifiable (see the editorial note in `docs/research/tm-deb-assessment.md`). **Nothing in this document uses it as evidence**, and the Section C delayed-reward design does **not** adopt any TM-DEB mechanism.
---
# SECTION A — The Tsetlin gun's target is the wrong shape (do this first)
## A.0 Correction to the brief, before anything else
The brief states that *"the gun currently learns a real-valued (cx,cy) pixel correction from a binary hit/miss outcome."* **That does not match HEAD.** The current training site already computes a signed residual. [FACT] `common_libs/guns/tsetlin.nim`, `proc onResult*`:
```nim
# Directional residual: actual enemy pos minus our prediction
# On hit residual is 0 (we were right); on miss we push toward actual position.
let rx = if e.hit: 0.0 else: clamp(e.actualX - t.predX, -TM_RESID_MAX, TM_RESID_MAX)
let ry = if e.hit: 0.0 else: clamp(e.actualY - t.predY, -TM_RESID_MAX, TM_RESID_MAX)
let lits = tmMakeLiterals(t.input)
g.net.tmLearnOne(0, lits, t.cache, rx)
g.net.tmLearnOne(1, lits, t.cache, ry)
```
The **previous** revision (`git show e536900~1:common_libs/guns/tsetlin.nim`) did the same. So the binary hit flag is **not** the label anywhere in the gun's history; it is only the *fitness* signal in `virtual_bullets.nim`. The brief's concern is therefore a non-issue here — but a *different and sharper* target-shape defect is present, derived below from the code.
## A.1 Why a hit/miss-only label *would* be ill-posed (stated for completeness)
If the label were only `hit ∈ {0,1}`:
- A hit says the total 2D prediction was good; a miss says it was bad. Neither says *which axis* was off, nor *by how much*, nor *which direction*. [FACT — by construction of a scalar label vs a 2D continuous target.]
- Fitting a signed 2D correction (two real outputs, each with sign and magnitude) to a scalar binary outcome would need the direction to be invented from something else (e.g. the predicted-vs-actual geometry). It would not be a well-posed supervised problem.
**This is not what the gun does**, per A.0.
## A.2 The harness hands the gun an exact signed residual
[FACT] `common_libs/gun_harness/gun_interface.nim` defines `FeedbackEvent` with `actualX`, `actualY`, `prediction`, plus `fireTick`, `powerBin`, `missDistance`, `hit`. At resolution, `common_libs/gun_harness/virtual_bullets.nim` (`tickBullets`) reads the target's actual position from the `enemies` table and builds the event:
```nim
let missDist = hypot(bx - ex, by - ey)
let fe = FeedbackEvent(
prediction: GunPrediction(x: b.aimX, y: b.aimY),
actualX: ex, actualY: ey, ... fireTick: b.fireTick, powerBin: b.powerBin, ...)
```
So `residual = actual − prediction` is known **exactly** at resolution time, per axis, with no ambiguity. [FACT]
[INFERENCE] Therefore the gun's learning task is **supervised regression with an exact continuous label**, not reinforcement learning. There is no hidden credit to assign, so no eligibility discounting is needed — the gun's effective `λ = 1`.
[FACT] The gun already exploits that: `onResult` retrieves the *exact* trace by `(fireTick, powerBin)` via `tmTraceSlot(e.fireTick, binIdx)` (ring of `TM_TRACE_SLOTS = 1024`), checking `t.fireTick == e.fireTick and t.powerBin == binIdx`. `trainedShots` counts successful pairings, `traceMisses` counts failures. This is the repo-native form of an eligibility buffer, and the pairing is exact. [FACT] `docs/research/tm-deb-assessment.md` §5 reaches the same conclusion: effective `λ = 1`.
## A.3 The *real* target-shape defect: the label is measured from the wrong baseline
This is the code-grounded finding that replaces the brief's premise.
Let, in `common_libs/guns/tsetlin.nim`:
- `L` = the linear extrapolation baseline, computed in `predict()` as `linearX = state.enemyX + cos(headingRad) * state.enemySpeed * ticksToArrive`. [FACT]
- `c` = the TM's output correction, `tmForwardWithCache` returns `(vx / TM_T * TM_RESID_MAX, …)`. [FACT]
- `P` = stored prediction, `predX = clamp(linearX + cx, …)`, and `TmTrace.predX = predX`. [FACT]
- `A` = `e.actualX` at resolution.
At training time `rx = A − P = A − L − c`. [FACT]
Inside `tmLearnOne`, `predicted` is the TM's own output recomputed from the cached clause outputs:
```nim
for c in 0..<TM_N_CLAUSES: vote += tmPolarity(c) * float(cache[outIdx * TM_N_CLAUSES + c])
vote = clamp(vote, -TM_T, TM_T)
let predicted = vote / TM_T * TM_RESID_MAX # == c
let error = residual - predicted # == rx - c
```
So `error = (A − L − c) − c = (A − L) − 2c`. [FACT — direct substitution.]
[INFERENCE] The update minimizes `|error|`, whose fixed point is `c* = (A − L)/2` — **half the correction that makes `P` hit `A`**. The code passes a label that already contains one copy of the TM's output, then subtracts a second copy. Even with a perfect feedback rule (Granmo Table 2/3), the TM is being asked for a systematically *shrunken* correction. This is a well-posedness bug in the target definition, not a learning-rate smell.
Corroborating detail: the same baseline pattern exists in the historical `BNNBot_garage/src/tsetlin_predictor.nim`, `proc learn` ("`residualX/Y: pixel correction needed (actual_target - aimed_point)`") and in `BNNBot_garage/src/TsetlinBot.nim` (`residualX = bot.lastEnemyX - bulletX`, where `bulletX` already includes the correction). [FACT] So this is not a one-off typo introduced by the current gun; it is a repeated design slip.
Three further target-shape defects at the same site:
1. **Hits are trained toward zero, not toward their true (small) residual.** [FACT] `rx = if e.hit: 0.0 else: …`. A hit means the 2D `missDistance < BotRadius = 18.0` [FACT], but each axis can still be off by up to 18 px. Zeroing both axes discards that information and biases the TM toward under-correction. A hit is not "residual exactly 0"; it is "`|(rx, ry)| < 18`".
2. **The update is magnitude-blind.** [FACT] `pFeedback = min(1.0, abs(error) / (2.0 * TM_RESID_MAX))` only scales the *probability* that a clause receives feedback; the per-clause state step is always ±1 (`st = min(st+1, …)` / `max(st-1, …)`). [INFERENCE] Direction is carried by `sign(error)` and clause polarity, but magnitude must be inferred statistically across many shots — slow, high-variance, and unable to express "this shot needed 40 px".
3. **Range/scale.** [FACT] `residual` is clamped to `±TM_RESID_MAX = 80` and the output vote to `±TM_T = 25` → `±80`. Because of the factor-2 defect above, the correction the TM converges to is **half** the baseline error: to correct a baseline error `δ` it must drive its output to `δ/2`, so it saturates at a baseline error of `±160` rather than the `±80` the constant name suggests. [INFERENCE]
## A.4 Consequence: the right target is the residual — measured from the baseline
[INFERENCE] The correct supervised target for the TM's correction is `δ = A − L` (the leftover error of the *linear* baseline), and `onResult` should pass `δ`, not `A − P`. Equivalently: store `linearX, linearY` in `TmTrace` alongside `predX, predY` and pass `e.actualX − t.linearX`. The current `TmTrace` stores `predX, predY` but **not** the baseline [FACT], so this requires a small struct change (but still no change to the harness — the baseline is computable at `predict` time and the actual position already arrives in the event).
## A.5 Learning a continuous residual with binary, gradient-free machinery
Tsetlin Machines are binary and gradient-free (a hard constraint for this project). [FACT — the automata in `tsetlin.nim` are discrete integer states updated by the Tables 2/3-style reward/penalty; there is no gradient.] Two honest ways to learn a continuous residual:
**Option A — quantise the residual and classify (recommended).**
- Split each axis's residual range `[−R, +R]` into `K` bins. Run two independent multi-class TMs (one per axis), each with `K` classes; each class gets its own clause team and vote, predict `argmax`, map the winning bin back to its centre pixel value. This is Granmo's standard multi-class formulation. [INFERENCE — standard TM construction; not present in this repo's code.]
- **Tradeoffs:**
- *Resolution vs data:* `K` classes need `K` teams and enough per-class samples. The residual distribution is heavily concentrated near 0, so most classes are rare. [INFERENCE] Finer bins = better magnitude resolution = more data per bin = slower convergence.
- *Magnitude vs direction:* a binned classifier captures both; a pure sign classifier (2 classes/axis) captures only direction and would still need a magnitude estimate (a fixed step, or a second head). [INFERENCE]
- *The no-correction class:* a hit should map to an explicit **zero bin** (e.g. `|residual| < BotRadius`). Modelling "no correction" as its own class lets the TM emit exactly 0 when it is right, instead of always nudging. Without it, hits trained toward zero (A.3 defect 1) will drag the whole output distribution toward 0.
**Option B — predict the SIGN per axis as an independent binary task.**
- 2 classes/axis → 2 clauses-team pairs; the smallest-data option and the most robust head. [INFERENCE] Magnitude must come from elsewhere (fixed step size, or a separate regression/binned head), so this is a partial solution but a good first milestone: prove the TM can learn *which way* the correction goes before asking it to learn *how far*.
[INFERENCE] Recommend **B as the first experiment, A as the destination**: the sign task is cheap and isolates the clause-learning question from the magnitude question.
## A.6 Recommended change, in order
1. **[HIGH] Fix the feedback rule first.** Implement Granmo Table 2/3 faithfully in `tmLearnOne` (the Type I `cOut` conditioning and the Type II direction). [FACT] The current Type I ignores `cOut`, and Type II is unsatisfiable dead code — see `docs/research/tm-deb-assessment.md` §6.3 and its §6.5 fixes 1–2. Changing the target shape while the update direction is broken is unmeasurable.
2. **[HIGH] Fix the label baseline.** Store `linearX, linearY` in `TmTrace`; pass `δ = actual − linear` to `tmLearnOne`. Remove the factor-2 shrink and stop zeroing residuals on hits (map hits into the zero bin / keep their true small `δ`).
3. **[MEDIUM] Choose the head.** Start with per-axis **sign** classification (Option B) to validate clause learning, then move to **binned multi-class** (Option A) with an explicit zero bin.
4. **[LOW] Keep the exact `(fireTick, powerBin)` pairing.** [FACT] It already works (`trainedShots` / `traceMisses`). Do **not** add discounting.
**Measurement that would prove it worked [INFERENCE]:**
- Build a deterministic oracle: for a constant-velocity enemy, `δ = 0` by construction, so a correct TM must output ≈ 0 (not oscillate). For a constant-acceleration / fixed-perpendicular-drift enemy, `δ = actual − linear` is analytically known; check the TM's output converges toward `δ`, **not `δ/2`**.
- Log per shot `(linear, actual, prediction, correction)`; metric = mean absolute `correction − δ` on held-out ticks, plus hit rate vs the `Linear` gun (the same adversary, same seed). Success = the factor-2 bias disappears and hit rate beats `Linear`. [UNKNOWN] the exact accuracy the TM can reach in the repo's data budget.
---
# SECTION B — Can a Tsetlin Machine learn high-level patterns?
## B.1 What a TM is actually good at
[FACT] A Tsetlin Machine learns a **DNF over binary literals**: each clause is a conjunction (`AND`) of a subset of the `2 × n_input` literals (a bit and its complement), and the output vote combines positive- and negative-polarity clauses. In this repo, `common_libs/guns/tsetlin.nim` `tmEvalClause` returns `1` iff *at least one* literal is included **and** *every* included literal's value is `1`; `tmForwardWithCache` sums `tmPolarity(c) * clause_output` into `vx`, `vy`. [FACT]
- Learning is by **discrete Tsetlin automata** (finite-state machines with `Reward`/`Penalty`), updated by Granmo's Tables 2/3; there is no gradient. [FACT]
- The learned model is **human-readable**: a clause is a list of included literals, which can be decoded back to `(frame index, field, bit)`. [FACT] The clause state array is indexed by `tmStateIdx(outIdx, clause, lit)`; `st > 0` = included. [FACT]
- [INFERENCE] This is the right tool when the target is genuinely a **conjunction of discrete conditions**, and the wrong tool when the target is a smooth signed magnitude that must be fitted precisely (Section A).
## B.2 The repo is already built for a temporal TM — [FACT]s
1. **Frame-stacked binary encoding already exists in the gun.** [FACT] `common_libs/guns/tsetlin.nim`:
- `TM_FRAME_BITS = 83`, `TM_SELF_BITS = 40`, `TM_WINDOW_SIZE = 10`, `TM_TOTAL_BITS = 83*10 + 40 = 870`.
- Gray coding via `tmToGray(value) = value xor (value shr 1)` and `tmToBits`.
- The window is LIFO (index 0 = newest) and is shifted **once per tick** (`if state.tick != g.frameTick`), with a stale comment documenting a past over-shift bug.
- [INFERENCE] Because every one of the 10 frames is laid out in the same fixed 83-bit block, a single clause can AND literals from *different frames*, i.e. it can express a **temporal conjunction** across the last 10 ticks.
2. **The per-frame field layout** (from `tmEncodeFrame`): bearing sin 8 + bearing cos 8 + distance 7 + velocity 5 + heading sin 8 + heading cos 8 + 4 wall distances × 7 + energy 11 = **83 bits**. [FACT]
3. **A clause over 10 frames is a high-level pattern.** [INFERENCE] "closing distance" = a distance-field bit pattern that changes across frames; "decelerating" = consecutive velocity-field codes ordered by Gray code; "low energy across the last N frames" = an energy-field bit repeated across frames. A conjunction of those literals is exactly the kind of rule a TM clause represents natively. [INFERENCE]
> **Mismatch to record.** [FACT] The gun's `tmEncodeFrame` call passes `state.selfEnergy` for the frame's energy field, with the comment `# use self energy as proxy (enemy energy not in WorldState)`. That comment is **stale**: `common_libs/gun_harness/gun_interface.nim` line 25 defines `WorldState.enemyEnergy*`. So the current gun encodes **self energy twice** and never sees enemy energy — which means a rule like "the enemy turns hard when *its* energy drops below X" is not even representable by the current gun, independent of its clause learner.
## B.3 What is reusable from `BNNBot_garage` today — proc-level
### `BNNBot_garage/src/binary_encoding.nim` — **encoder is fully reusable**
- Types: `BinaryVector = array[TOTAL_BITS, uint8]`, `EnemyScanFrame`, `SelfState`, `OutputVector`.
- Constants: `FRAME_BITS = 83`, `SELF_BITS = 40`, `WINDOW_SIZE = 10`, `TOTAL_BITS = 870`, `OUTPUT_BITS = 7`, `AIM_MIN = -60.0`, `AIM_MAX = 60.0`.
- **Reusable procs:** `encodeFrame(frame: EnemyScanFrame): array[FRAME_BITS, uint8]`, `encodeSelf(self: SelfState): array[SELF_BITS, uint8]`, `encodeFullVector(window, self): BinaryVector`, and the utility set `formatBinary`, `popcount`, `hammingDistance`, `bitwiseAnd`, `formatVectorBinary`, `formatVectorDecimal`, `formatFrameBinary`, `encodeOutput(angle): OutputVector`, `decodeOutput(vec): float`. [FACT]
- **NOT reusable (self-documented placeholders that return hardcoded values):** `formatFrameDecimal` (returns `"199,199,99,15,199,199,99,99,150"`) and `decodeFrameFields` (returns a fixed `[0,…,50,50,100]` array). [FACT] These are stubs, not decoders — do not build on them.
- [FACT] The encoder is **already inlined** into `common_libs/guns/tsetlin.nim` as `tmEncodeFrame` / `tmEncodeSelf` / `tmEncodeFullVector` (the gun's header says "adapted from BNNBot_garage/src/binary_encoding.nim"), so there is nothing to import; BNNBot is the reference copy. One real divergence: BNNBot's `EnemyScanFrame` carries `enemyEnergy`, while the gun's inlined copy takes `state.selfEnergy` (B.2 mismatch above).
### `BNNBot_garage/src/tsetlin_predictor.nim` — **reusable as scaffolding, but it is a regression TM with the same feedback defects**
- Types: `TsetlinNet`, `ClauseCache`, `EligibilityTrace`, `VirtualBullet`. Constants: `N_IN = 870`, `N_OUT = 2`, `N_CLAUSES = 64` (32 pos + 32 neg), `N_STATES = 15`, `T = 32.0`, `S = 4.0`, `RESID_MAX = 80.0`, `TRACE_MAX_AGE = 40`.
- **Procs:** `stateIdx`, `clausePolarity`, `evalClause`, `makeLiterals`, `computeVote`, `initTsetlinNet`, `forward`, `forwardWithCache`, `learnOne`, `learn`. [FACT]
- **Honest assessment of the learning core** [FACT]:
- It is a **regression** TM (continuous `(cx, cy)` scaled to `±RESID_MAX`), *not* Granmo's multi-class classifier; the output is `vote / T * RESID_MAX`.
- Its "Type Ia" (`cOut == 1`) and "Type Ib" (`else`) branches execute **identical bodies** (both grow toward the current input), so the split is cosmetic. The counteracting `(c=0, literal=1) → toward-Exclude` force is absent.
- Its "Bug 3 fix" Type II decrements **included** false literals (`if literals[lit] == 0 and st > 0`) rather than incrementing *excluded* false literals; when `cOut == 1` that condition is unsatisfiable, so the branch is dead code (same defect as the current gun — see `tm-deb-assessment.md` §6.3).
- **Verdict:** reusable as a **prototype scaffold** (types, vote, clause cache, eligibility trace), but **not** as a correct learning core. The fixed feedback rule from Section A.6 step 1 is a prerequisite.
### `BNNBot_garage/src/wisard_predictor.nim` — **not a TM**
- Types `WiSARDNet`, `WaveTrace`, `VirtualBullet`; procs `initWiSARD(seed)`, `computeAddresses`, `predictCorrection`, `learnCorrection`; constants `K = 14`, `N_NEURONS = 50`, `LUT_SIZE = 16384`. [FACT]
- It is a WiSARD/n-tuple **LUT regressor** averaging `(sumX/count, sumY/count)`. [FACT] Reusable only as a **non-TM baseline**; it regresses `(dx, dy)` directly, so it bypasses the binary-machinery question entirely. It has no interpretable clauses.
**Summary:** the *encoder* is reusable as-is (minus the two stub decoders). The TM code is reusable as reference/scaffolding but is a regression TM with broken feedback rules, not a correct classifier. Both already live inside the gun, so the practical value of `BNNBot_garage` today is *historical reference*, not an importable library.
## B.4 Concrete, falsifiable experiment
**Goal:** prove a TM can recover a *known* high-level temporal rule from the frame-stacked encoding — and measure whether it can do so sparsely.
**Setup [INFERENCE unless noted]:**
1. Define a synthetic enemy with a rule that is (a) deterministic, (b) temporal, (c) expressible as a conjunction over the encoded frame fields. Two candidates:
- **Rule E (energy-gated turn):** the enemy turns hard left/right on tick `t` iff its energy has **dropped below a fixed threshold within the last `k` frames**. [FACT] energy is an 11-bit Gray-coded field per frame; [FACT] the gun currently feeds *self* energy, so this rule needs the `enemyEnergy` fix (B.2).
- **Rule D (fixed perpendicular drift):** the enemy holds a constant heading offset relative to the bearing for `k` frames → the velocity/heading fields repeat across the window.
2. Feed the 10-frame stacked 870-bit vector to a **standalone classifier TM** (not the gun's 2-output regression head): multi-class over `{turn-left, turn-right, no-turn}` or `{drift, no-drift}`.
3. Train/test split, fixed seed, and a majority-class baseline reported alongside.
**Accuracy / sparsity bar that would prove the TM found the rule [INFERENCE — proposed thresholds, not measured]:**
- **Accuracy:** on held-out noiseless frames, a correct rule should be recovered near-ceiling (propose **≥ 95 %** on the deterministic rule, vs 50 % chance for a balanced binary task, and vs the majority-class rate). Anything at the majority-class rate means no rule was learned.
- **Sparsity:** the clauses that win should each include a **small single-digit** number of literals, and the count should be stable across runs. [FACT] The current gun's clauses saturate at ~131 literals/clause when the feedback rule is broken (assessment §6.4); that is the failure signature to watch for.
- **Interpretability check:** decode the highest-weight clause's literals via `tmStateIdx` back to `(frame, field, bit)` and verify they land on the energy/velocity/heading fields and the expected frames. This is the test that it found a *high-level* rule, not a coincidence.
**Can the current gun represent such a rule, once its feedback tables are fixed?**
- **Clause width:** [FACT] there is no width cap; a clause may include any of `TM_N_LITERALS = 1740` literals (`st > 0`). Representability is not blocked by clause width.
- **Feature count:** [FACT] 870 input bits / 1740 literals, with the energy field spanning 11 bits × 10 frames. [INFERENCE] Gray-coding means a monotone threshold ("energy < X") is **not one literal** — it is a sub-pattern of the 11-bit code; a single AND clause captures a specific bit pattern, and the class team as a whole is what learns the threshold. So the rule is representable by the team, not by a single clause.
- **Data volume:** [FACT] `docs/rl-algorithm-choice.md` states episodes are ~30–180 ticks; the gun only creates a trace on ticks with a valid power bin, and `trainedShots = 2141` (cited as observed in `tm-deb-assessment.md` §4) is the accumulated sample count. [UNKNOWN] whether that is enough for convergence — this is exactly what the experiment settles.
- **Head shape:** [FACT] the gun's head is `TM_N_OUT = 2` regression outputs (`cx, cy`), not discrete rule classes. So **the gun as-is cannot be pointed at the Rule E/D classification task** — the experiment must use a standalone classifier TM. The gun can only benefit from these findings indirectly.
---
# SECTION C — Delayed reward: the movement case
## C.1 Delayed supervision vs delayed reward — a hard distinction
- **Delayed supervision (the gun):** the outcome (a shot resolving) arrives 4–128 ticks after the fire tick. [FACT] `bulletSpeed(power) = 20 − 3·power` (`gun_interface.nim`) and `PowerBins = [1.0, 1.5, 2.0, 3.0]` → speeds `17, 15.5, 14, 11` px/tick (`virtual_bullets.nim`); at typical ≈400 px engagement that is ≈23–36 ticks, and the gun's own comment notes a power-3 long shot can take ~128 ticks. **But the outcome is exactly attributable per shot**: `FeedbackEvent` carries `fireTick` and `powerBin`, and `onResult` indexes `tmTraceSlot(e.fireTick, binIdx)` to the exact stored trace. [FACT] So the label is delayed, not confounded. **Effective `λ = 1`; no discounting needed.** [FACT, matching `tm-deb-assessment.md` §5]
- **Delayed reward (movement):** one sparse outcome per round (win/loss/rank/score) with **no** action-to-outcome mapping. You cannot say which of the hundreds of per-tick turn/thrust decisions caused a win. [INFERENCE] This is the case where credit assignment — discounting an eligibility buffer, or a TD/REINFORCE estimator — is the correct tool. [FACT] `tm-deb-assessment.md` §7 already identifies a learned movement module as the repo's genuine delayed-credit-assignment case.
[INFERENCE] **Discounting is only justified for the second case.** Applying it to the gun would delete exactly-attributable long-range samples (assessment §4: `γ^Δt = 4.4e-7` at `Δt = 90`).
## C.2 What the movement layer actually does today
[FACT] The movement modules in `common_libs/movements/` are: `the_floor_is_lava.nim`, `phantom_meteor.nim`, `wave_surfer.nim`, `minimum_risk.nim`, `oscillator.nim`, `random_oscillator.nim`, `rammer.nim`.
- **[FACT] The live mover is TFIL.** `ModularBot_garage/src/ModularBot.nim` declares `mover: TFILModule`, initialises `mover: TFILModule(debugGraphics: true)`, and imports `movements/the_floor_is_lava`.
- **[FACT] TFIL is hand-tuned with no learnable parameters.** Every coefficient is a `const`: `BulletCore = 10.0`, `BulletAura = 5.0`, `EnemyCore = 40.0`, `EnemyAura = 10.0`, `CorridorHeat = 20.0`, `WallHotness = 30.0`, `WallRadiance = 10.0`, `PillarHotness = 30.0`, `PillarRadiance = 10.0`, `CommitTicks = 15`, `MinCommitTicks = 5`, `DangerReplanThreshold = 25.0`, `CoolestLevels = 2`, `MaxTrackedBullets = 20`, `GridSize = 36.0`, `MaxSpeed = 8.0`. `computeMove` builds a lava grid, scores reachable tiles, commits for `CommitTicks`, and picks among safe tiles with a **random** tie-break (`rand(candidates.high)`). No reward ever reaches it; `resetRound` only clears transients.
- **[FACT] Other modules:** `minimum_risk.nim` — all scoring weights are `const` (`KEnemy = 1.0`, `KWall = 0.5`, `KCorner = 0.8`, `KTravel = 0.003`). `oscillator.nim` / `random_oscillator.nim` / `rammer.nim` — fixed period/threshold constants. `wave_surfer.nim` has an observed-danger histogram `bins: array[31, float64]` incremented as enemy waves pass, and **`resetRound` does not clear `bins`** (it persists across rounds). `phantom_meteor.nim` likewise keeps its `DangerHistogram` across rounds. [FACT]
- **[INFERENCE] These histograms are counters, not reward-optimized parameters:** no reward is ever propagated into them; they are incremented purely from observed bullet geometry. There is **no existing learnable movement parameter trained by any outcome signal**.
**Reward signal available from the bot API [FACT]:**
- `ResultsForBot` (`robocode_tankroyale_botapi/schemas.nim`) exposes per round: `rank`, `survival`, `lastSurvivorBonus`, `bulletDamage`, `bulletKillBonus`, `ramDamage`, `ramKillBonus`, `totalScore`, plus season counters `firstPlaces`/`secondPlaces`/`thirdPlaces`. Delivered in `onRoundEnded` via `RoundEndedEventForBot`. [FACT]
- `ModularBot.onRoundEnded` already reads `e.results.totalScore` and `e.results.bulletDamage` (it logs them to `/tmp/gun_stats.jsonl`), so the plumbing for a per-round reward already exists. [FACT]
- Dense per-tick proxies also exist: `getEnergy()` (`buildState` reads it), `onHitByBullet` (damage taken), `onBulletHit` (damage dealt), `onBotDeath`. [FACT] `SAC_LSTM_Bot_garage/src/SAC_LSTM_Bot/rewards.nim` shows a hand-designed `computeReward(...)` combining win/loss (+20/−10), bullet damage, ram penalty, charge penalty, and wall ticks. [FACT]
## C.3 Minimal delayed-reward learning design
[INFERENCE] A deliberately small first design, one live parameter at a time:
1. **One live knob.** Pick a single low-dimensional movement decision with a small discrete action set — e.g. a bias toward/away from a tile class in TFIL's scoring, or an "engagement distance / aggression" scalar (mirroring `wave_surfer`'s `WS_PrefDist`), or the reversal threshold in the oscillators. Keep it to a few discrete actions so a per-action clause team is trainable.
2. **Bounded experience buffer.** Each tick, store `(encodedStateBits, actionIdx)` in a fixed-size ring (a few hundred ticks). Encode the state with the existing 870-bit encoder from `binary_encoding.nim` (or a movement-relevant subset); this is the same eligibility-buffer shape the gun already uses, just keyed by tick instead of `(fireTick, powerBin)`. [INFERENCE]
3. **Dispatch at round end.** In `onRoundEnded`, compute a scalar reward `r` from `ResultsForBot` (`totalScore`/`rank`/`survival`; optionally plus a dense term as in `rewards.nim`), then replay the buffer: for each buffered `(state, action)` apply Type I/II feedback for the chosen action's clause team, weighted by `γ^(T − t)`. [INFERENCE]
4. **Credit-assignment honesty.** One reward per round spread over ~30–180 ticks is extremely sparse and high-variance. [FACT for the tick range from `docs/rl-algorithm-choice.md`; INFERENCE for the consequence.] The `γ^(T − t)` decay is exactly a heuristic for "earlier actions are less likely to be responsible", and it does not make the attribution correct — it only hedges. [INFERENCE]
**How many rounds before a signal is detectable [INFERENCE — standard two-proportion power calculation, computed by hand; no runtime measurement]:**
For a win-rate reward, two-proportion sample size per arm
`n = (z_{α/2} + z_β)² · [p₁(1−p₁) + p₂(1−p₂)] / (p₁ − p₂)²`, with `α = 0.05`, power `0.80` ⇒ `(1.96 + 0.8416)² ≈ 7.85`.
- Detect `0.50 → 0.55` (5 pts): `n ≈ 7.85 · (0.25 + 0.2475) / 0.0025 ≈ 1560` rounds **per arm**.
- Detect `0.50 → 0.60` (10 pts): `n ≈ 7.85 · (0.25 + 0.24) / 0.01 ≈ 385` rounds **per arm**.
These are *lower bounds* for a between-condition comparison; a single adaptive learner is worse (nonstationarity, exploration cost, interference), so treat them as optimistic. [INFERENCE]
**Feasibility.** Given the brief's figure of ~1 round per few seconds, 385 rounds ≈ 20–30 min, and 1560 rounds ≈ 1.5–2 h per arm. [FACT] `docs/rl-algorithm-choice.md` also states 10–35 rounds per match and a training window of "~few hundred milliseconds" between rounds. [INFERENCE] So an **offline, between-battle** experiment is feasible within hours; **online learning inside a single match is not** (a match has ~10–35 rounds). [UNKNOWN] the actual wall-clock duration of one round in this repo's harness — no file I read records it, so the minutes/hours above are conditional arithmetic, not measured.
## C.4 What is UNPROVEN, and the experiment that would settle it
- **[UNKNOWN] Does the per-round reward carry enough signal to be learnable at all?** It may be dominated by opponent identity, arena randomness, and starting positions rather than by the movement knob.
- **[UNKNOWN] The wall-clock cost per round**, hence total experiment time and whether the round budget is realistic.
- **[UNKNOWN] Whether γ-discounted per-tick feedback beats a fixed hand-tuned policy** (TFIL) at all, and by how much — no measurement exists.
- **[UNKNOWN] Whether one live knob is enough** to move `totalScore` measurably, or whether the effect is buried in variance.
**Experiment that would settle it [INFERENCE]:**
1. **Offline counterfactual first (cheap gate).** Replay a fixed corpus of recorded enemy movement tapes. For each candidate action bias, simulate the resulting movement and compute the reward proxy. If **no** single-knob policy beats TFIL's fixed policy offline, do not spend the online round budget — the signal is not there. If one does, use its effect size to size the online run.
2. **Online A/B, seeded.** Fixed opponent set and seeds; arm A = current TFIL, arm B = TM-biased TFIL. Run `N` rounds per arm from the power calculation above (start at ≈ 385/arm for a 10-pt effect, expand for smaller effects). Pre-register `totalScore`/`rank`/`survival` as the metrics and report confidence intervals, not point estimates.
3. **Instrument variance.** Log per-round reward and its spread *before* learning; if the round-to-round standard deviation is large relative to any plausible effect, report that directly as the blocker rather than running an under-powered test.
> **Explicit non-import.** None of the above relies on `docs/papers/tm-deb-paper.pdf`. The only external claim used is the standard power calculation (C.3). The TM-DEB idea of discounting a replayed buffer is *conceptually* the right shape for this case (as `tm-deb-assessment.md` §7 already noted), but no result, equation, or number from that deleted document is treated as evidence here.