Files
SirRoboGarage/docs/research/tm-learning-tracks.md
T
SirStone 54e9757567 docs(research): TM learning tracks + note that the TM-DEB source was deleted
tm-learning-tracks.md covers three things, all marked [FACT]/[INFERENCE]/
[UNKNOWN]:
- Section A: the Tsetlin gun's label is measured against the wrong baseline.
  predX = linearX + cx, so rx = actual - predX = delta - cx, and inside
  tmLearnOne error = residual - predicted = (delta - cx) - cx = delta - 2cx.
  The fixed point is cx = delta/2 -- HALF the correction needed, even with
  perfect Granmo feedback. Fix: store linearX/linearY in TmTrace and train
  on delta. Also: hits zero the label instead of carrying their true
  residual, and the per-clause step is magnitude-blind.
- Section B: what a TM is actually good at (AND-clauses over binary
  literals, readable output) and why this repo suits it -- the gun already
  builds an 83-bit x 10-frame Gray-coded window (870 bits). Includes a
  falsifiable known-rule benchmark proposal.
- Section C: delayed-reward learning belongs to the MOVEMENT layer, not the
  gun. The gun's outcome is delayed but exactly pairable via
  (fireTick, powerBin), so its effective lambda is 1 and discounting would
  only destroy information.

Also records that docs/papers/tm-deb-paper.pdf was deleted by the user as
AI-generated and unverifiable, while the Granmo-based feedback diff in
tm-deb-assessment.md stands on its own.
2026-09-20 23:28:31 +02:00

253 lines
32 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Tsetlin Machine learning tracks — target shape, high-level patterns, delayed reward
**Date:** 2026-09-20
**Scope:** Design notes only. No `.nim` file was edited, nothing was built, no battle was run, no `/tmp` artifact was touched. Every claim is tagged by its evidence class.
**Author's method:** Read `common_libs/guns/tsetlin.nim`, `common_libs/gun_harness/gun_interface.nim`, `common_libs/gun_harness/virtual_bullets.nim`, the `common_libs/movements/*` modules, `ModularBot_garage/src/ModularBot.nim`, `BNNBot_garage/src/{binary_encoding,tsetlin_predictor,wisard_predictor}.nim`, `SAC_LSTM_Bot_garage/src/SAC_LSTM_Bot/rewards.nim`, `docs/rl-algorithm-choice.md`, `docs/research/tm-deb-assessment.md`, and the installed `robocode_tankroyale_botapi` `schemas.nim` / `bot.nim`.
## Evidence legend
- **[FACT]** — verified by reading the cited code or paper at the cited location.
- **[INFERENCE]** — reasoned from facts; not measured.
- **[UNKNOWN]** — needs an experiment; no claim is made.
> **Provenance caveat.** `docs/papers/tm-deb-paper.pdf` (TM-DEB) was deleted by the user as AI-generated and unverifiable (see the editorial note in `docs/research/tm-deb-assessment.md`). **Nothing in this document uses it as evidence**, and the Section C delayed-reward design does **not** adopt any TM-DEB mechanism.
---
# SECTION A — The Tsetlin gun's target is the wrong shape (do this first)
## A.0 Correction to the brief, before anything else
The brief states that *"the gun currently learns a real-valued (cx,cy) pixel correction from a binary hit/miss outcome."* **That does not match HEAD.** The current training site already computes a signed residual. [FACT] `common_libs/guns/tsetlin.nim`, `proc onResult*`:
```nim
# Directional residual: actual enemy pos minus our prediction
# On hit residual is 0 (we were right); on miss we push toward actual position.
let rx = if e.hit: 0.0 else: clamp(e.actualX - t.predX, -TM_RESID_MAX, TM_RESID_MAX)
let ry = if e.hit: 0.0 else: clamp(e.actualY - t.predY, -TM_RESID_MAX, TM_RESID_MAX)
let lits = tmMakeLiterals(t.input)
g.net.tmLearnOne(0, lits, t.cache, rx)
g.net.tmLearnOne(1, lits, t.cache, ry)
```
The **previous** revision (`git show e536900~1:common_libs/guns/tsetlin.nim`) did the same. So the binary hit flag is **not** the label anywhere in the gun's history; it is only the *fitness* signal in `virtual_bullets.nim`. The brief's concern is therefore a non-issue here — but a *different and sharper* target-shape defect is present, derived below from the code.
## A.1 Why a hit/miss-only label *would* be ill-posed (stated for completeness)
If the label were only `hit ∈ {0,1}`:
- A hit says the total 2D prediction was good; a miss says it was bad. Neither says *which axis* was off, nor *by how much*, nor *which direction*. [FACT — by construction of a scalar label vs a 2D continuous target.]
- Fitting a signed 2D correction (two real outputs, each with sign and magnitude) to a scalar binary outcome would need the direction to be invented from something else (e.g. the predicted-vs-actual geometry). It would not be a well-posed supervised problem.
**This is not what the gun does**, per A.0.
## A.2 The harness hands the gun an exact signed residual
[FACT] `common_libs/gun_harness/gun_interface.nim` defines `FeedbackEvent` with `actualX`, `actualY`, `prediction`, plus `fireTick`, `powerBin`, `missDistance`, `hit`. At resolution, `common_libs/gun_harness/virtual_bullets.nim` (`tickBullets`) reads the target's actual position from the `enemies` table and builds the event:
```nim
let missDist = hypot(bx - ex, by - ey)
let fe = FeedbackEvent(
prediction: GunPrediction(x: b.aimX, y: b.aimY),
actualX: ex, actualY: ey, ... fireTick: b.fireTick, powerBin: b.powerBin, ...)
```
So `residual = actual − prediction` is known **exactly** at resolution time, per axis, with no ambiguity. [FACT]
[INFERENCE] Therefore the gun's learning task is **supervised regression with an exact continuous label**, not reinforcement learning. There is no hidden credit to assign, so no eligibility discounting is needed — the gun's effective `λ = 1`.
[FACT] The gun already exploits that: `onResult` retrieves the *exact* trace by `(fireTick, powerBin)` via `tmTraceSlot(e.fireTick, binIdx)` (ring of `TM_TRACE_SLOTS = 1024`), checking `t.fireTick == e.fireTick and t.powerBin == binIdx`. `trainedShots` counts successful pairings, `traceMisses` counts failures. This is the repo-native form of an eligibility buffer, and the pairing is exact. [FACT] `docs/research/tm-deb-assessment.md` §5 reaches the same conclusion: effective `λ = 1`.
## A.3 The *real* target-shape defect: the label is measured from the wrong baseline
This is the code-grounded finding that replaces the brief's premise.
Let, in `common_libs/guns/tsetlin.nim`:
- `L` = the linear extrapolation baseline, computed in `predict()` as `linearX = state.enemyX + cos(headingRad) * state.enemySpeed * ticksToArrive`. [FACT]
- `c` = the TM's output correction, `tmForwardWithCache` returns `(vx / TM_T * TM_RESID_MAX, …)`. [FACT]
- `P` = stored prediction, `predX = clamp(linearX + cx, …)`, and `TmTrace.predX = predX`. [FACT]
- `A` = `e.actualX` at resolution.
At training time `rx = A − P = A − L − c`. [FACT]
Inside `tmLearnOne`, `predicted` is the TM's own output recomputed from the cached clause outputs:
```nim
for c in 0..<TM_N_CLAUSES: vote += tmPolarity(c) * float(cache[outIdx * TM_N_CLAUSES + c])
vote = clamp(vote, -TM_T, TM_T)
let predicted = vote / TM_T * TM_RESID_MAX # == c
let error = residual - predicted # == rx - c
```
So `error = (A − L − c) − c = (A − L) − 2c`. [FACT — direct substitution.]
[INFERENCE] The update minimizes `|error|`, whose fixed point is `c* = (A − L)/2` — **half the correction that makes `P` hit `A`**. The code passes a label that already contains one copy of the TM's output, then subtracts a second copy. Even with a perfect feedback rule (Granmo Table 2/3), the TM is being asked for a systematically *shrunken* correction. This is a well-posedness bug in the target definition, not a learning-rate smell.
Corroborating detail: the same baseline pattern exists in the historical `BNNBot_garage/src/tsetlin_predictor.nim`, `proc learn` ("`residualX/Y: pixel correction needed (actual_target - aimed_point)`") and in `BNNBot_garage/src/TsetlinBot.nim` (`residualX = bot.lastEnemyX - bulletX`, where `bulletX` already includes the correction). [FACT] So this is not a one-off typo introduced by the current gun; it is a repeated design slip.
Three further target-shape defects at the same site:
1. **Hits are trained toward zero, not toward their true (small) residual.** [FACT] `rx = if e.hit: 0.0 else: …`. A hit means the 2D `missDistance < BotRadius = 18.0` [FACT], but each axis can still be off by up to 18 px. Zeroing both axes discards that information and biases the TM toward under-correction. A hit is not "residual exactly 0"; it is "`|(rx, ry)| < 18`".
2. **The update is magnitude-blind.** [FACT] `pFeedback = min(1.0, abs(error) / (2.0 * TM_RESID_MAX))` only scales the *probability* that a clause receives feedback; the per-clause state step is always ±1 (`st = min(st+1, …)` / `max(st-1, …)`). [INFERENCE] Direction is carried by `sign(error)` and clause polarity, but magnitude must be inferred statistically across many shots — slow, high-variance, and unable to express "this shot needed 40 px".
3. **Range/scale.** [FACT] `residual` is clamped to `±TM_RESID_MAX = 80` and the output vote to `±TM_T = 25` → `±80`. Because of the factor-2 defect above, the correction the TM converges to is **half** the baseline error: to correct a baseline error `δ` it must drive its output to `δ/2`, so it saturates at a baseline error of `±160` rather than the `±80` the constant name suggests. [INFERENCE]
## A.4 Consequence: the right target is the residual — measured from the baseline
[INFERENCE] The correct supervised target for the TM's correction is `δ = A − L` (the leftover error of the *linear* baseline), and `onResult` should pass `δ`, not `A − P`. Equivalently: store `linearX, linearY` in `TmTrace` alongside `predX, predY` and pass `e.actualX − t.linearX`. The current `TmTrace` stores `predX, predY` but **not** the baseline [FACT], so this requires a small struct change (but still no change to the harness — the baseline is computable at `predict` time and the actual position already arrives in the event).
## A.5 Learning a continuous residual with binary, gradient-free machinery
Tsetlin Machines are binary and gradient-free (a hard constraint for this project). [FACT — the automata in `tsetlin.nim` are discrete integer states updated by the Tables 2/3-style reward/penalty; there is no gradient.] Two honest ways to learn a continuous residual:
**Option A — quantise the residual and classify (recommended).**
- Split each axis's residual range `[−R, +R]` into `K` bins. Run two independent multi-class TMs (one per axis), each with `K` classes; each class gets its own clause team and vote, predict `argmax`, map the winning bin back to its centre pixel value. This is Granmo's standard multi-class formulation. [INFERENCE — standard TM construction; not present in this repo's code.]
- **Tradeoffs:**
- *Resolution vs data:* `K` classes need `K` teams and enough per-class samples. The residual distribution is heavily concentrated near 0, so most classes are rare. [INFERENCE] Finer bins = better magnitude resolution = more data per bin = slower convergence.
- *Magnitude vs direction:* a binned classifier captures both; a pure sign classifier (2 classes/axis) captures only direction and would still need a magnitude estimate (a fixed step, or a second head). [INFERENCE]
- *The no-correction class:* a hit should map to an explicit **zero bin** (e.g. `|residual| < BotRadius`). Modelling "no correction" as its own class lets the TM emit exactly 0 when it is right, instead of always nudging. Without it, hits trained toward zero (A.3 defect 1) will drag the whole output distribution toward 0.
**Option B — predict the SIGN per axis as an independent binary task.**
- 2 classes/axis → 2 clauses-team pairs; the smallest-data option and the most robust head. [INFERENCE] Magnitude must come from elsewhere (fixed step size, or a separate regression/binned head), so this is a partial solution but a good first milestone: prove the TM can learn *which way* the correction goes before asking it to learn *how far*.
[INFERENCE] Recommend **B as the first experiment, A as the destination**: the sign task is cheap and isolates the clause-learning question from the magnitude question.
## A.6 Recommended change, in order
1. **[HIGH] Fix the feedback rule first.** Implement Granmo Table 2/3 faithfully in `tmLearnOne` (the Type I `cOut` conditioning and the Type II direction). [FACT] The current Type I ignores `cOut`, and Type II is unsatisfiable dead code — see `docs/research/tm-deb-assessment.md` §6.3 and its §6.5 fixes 1–2. Changing the target shape while the update direction is broken is unmeasurable.
2. **[HIGH] Fix the label baseline.** Store `linearX, linearY` in `TmTrace`; pass `δ = actual − linear` to `tmLearnOne`. Remove the factor-2 shrink and stop zeroing residuals on hits (map hits into the zero bin / keep their true small `δ`).
3. **[MEDIUM] Choose the head.** Start with per-axis **sign** classification (Option B) to validate clause learning, then move to **binned multi-class** (Option A) with an explicit zero bin.
4. **[LOW] Keep the exact `(fireTick, powerBin)` pairing.** [FACT] It already works (`trainedShots` / `traceMisses`). Do **not** add discounting.
**Measurement that would prove it worked [INFERENCE]:**
- Build a deterministic oracle: for a constant-velocity enemy, `δ = 0` by construction, so a correct TM must output ≈ 0 (not oscillate). For a constant-acceleration / fixed-perpendicular-drift enemy, `δ = actual − linear` is analytically known; check the TM's output converges toward `δ`, **not `δ/2`**.
- Log per shot `(linear, actual, prediction, correction)`; metric = mean absolute `correction − δ` on held-out ticks, plus hit rate vs the `Linear` gun (the same adversary, same seed). Success = the factor-2 bias disappears and hit rate beats `Linear`. [UNKNOWN] the exact accuracy the TM can reach in the repo's data budget.
---
# SECTION B — Can a Tsetlin Machine learn high-level patterns?
## B.1 What a TM is actually good at
[FACT] A Tsetlin Machine learns a **DNF over binary literals**: each clause is a conjunction (`AND`) of a subset of the `2 × n_input` literals (a bit and its complement), and the output vote combines positive- and negative-polarity clauses. In this repo, `common_libs/guns/tsetlin.nim` `tmEvalClause` returns `1` iff *at least one* literal is included **and** *every* included literal's value is `1`; `tmForwardWithCache` sums `tmPolarity(c) * clause_output` into `vx`, `vy`. [FACT]
- Learning is by **discrete Tsetlin automata** (finite-state machines with `Reward`/`Penalty`), updated by Granmo's Tables 2/3; there is no gradient. [FACT]
- The learned model is **human-readable**: a clause is a list of included literals, which can be decoded back to `(frame index, field, bit)`. [FACT] The clause state array is indexed by `tmStateIdx(outIdx, clause, lit)`; `st > 0` = included. [FACT]
- [INFERENCE] This is the right tool when the target is genuinely a **conjunction of discrete conditions**, and the wrong tool when the target is a smooth signed magnitude that must be fitted precisely (Section A).
## B.2 The repo is already built for a temporal TM — [FACT]s
1. **Frame-stacked binary encoding already exists in the gun.** [FACT] `common_libs/guns/tsetlin.nim`:
- `TM_FRAME_BITS = 83`, `TM_SELF_BITS = 40`, `TM_WINDOW_SIZE = 10`, `TM_TOTAL_BITS = 83*10 + 40 = 870`.
- Gray coding via `tmToGray(value) = value xor (value shr 1)` and `tmToBits`.
- The window is LIFO (index 0 = newest) and is shifted **once per tick** (`if state.tick != g.frameTick`), with a stale comment documenting a past over-shift bug.
- [INFERENCE] Because every one of the 10 frames is laid out in the same fixed 83-bit block, a single clause can AND literals from *different frames*, i.e. it can express a **temporal conjunction** across the last 10 ticks.
2. **The per-frame field layout** (from `tmEncodeFrame`): bearing sin 8 + bearing cos 8 + distance 7 + velocity 5 + heading sin 8 + heading cos 8 + 4 wall distances × 7 + energy 11 = **83 bits**. [FACT]
3. **A clause over 10 frames is a high-level pattern.** [INFERENCE] "closing distance" = a distance-field bit pattern that changes across frames; "decelerating" = consecutive velocity-field codes ordered by Gray code; "low energy across the last N frames" = an energy-field bit repeated across frames. A conjunction of those literals is exactly the kind of rule a TM clause represents natively. [INFERENCE]
> **Mismatch to record.** [FACT] The gun's `tmEncodeFrame` call passes `state.selfEnergy` for the frame's energy field, with the comment `# use self energy as proxy (enemy energy not in WorldState)`. That comment is **stale**: `common_libs/gun_harness/gun_interface.nim` line 25 defines `WorldState.enemyEnergy*`. So the current gun encodes **self energy twice** and never sees enemy energy — which means a rule like "the enemy turns hard when *its* energy drops below X" is not even representable by the current gun, independent of its clause learner.
## B.3 What is reusable from `BNNBot_garage` today — proc-level
### `BNNBot_garage/src/binary_encoding.nim` — **encoder is fully reusable**
- Types: `BinaryVector = array[TOTAL_BITS, uint8]`, `EnemyScanFrame`, `SelfState`, `OutputVector`.
- Constants: `FRAME_BITS = 83`, `SELF_BITS = 40`, `WINDOW_SIZE = 10`, `TOTAL_BITS = 870`, `OUTPUT_BITS = 7`, `AIM_MIN = -60.0`, `AIM_MAX = 60.0`.
- **Reusable procs:** `encodeFrame(frame: EnemyScanFrame): array[FRAME_BITS, uint8]`, `encodeSelf(self: SelfState): array[SELF_BITS, uint8]`, `encodeFullVector(window, self): BinaryVector`, and the utility set `formatBinary`, `popcount`, `hammingDistance`, `bitwiseAnd`, `formatVectorBinary`, `formatVectorDecimal`, `formatFrameBinary`, `encodeOutput(angle): OutputVector`, `decodeOutput(vec): float`. [FACT]
- **NOT reusable (self-documented placeholders that return hardcoded values):** `formatFrameDecimal` (returns `"199,199,99,15,199,199,99,99,150"`) and `decodeFrameFields` (returns a fixed `[0,…,50,50,100]` array). [FACT] These are stubs, not decoders — do not build on them.
- [FACT] The encoder is **already inlined** into `common_libs/guns/tsetlin.nim` as `tmEncodeFrame` / `tmEncodeSelf` / `tmEncodeFullVector` (the gun's header says "adapted from BNNBot_garage/src/binary_encoding.nim"), so there is nothing to import; BNNBot is the reference copy. One real divergence: BNNBot's `EnemyScanFrame` carries `enemyEnergy`, while the gun's inlined copy takes `state.selfEnergy` (B.2 mismatch above).
### `BNNBot_garage/src/tsetlin_predictor.nim` — **reusable as scaffolding, but it is a regression TM with the same feedback defects**
- Types: `TsetlinNet`, `ClauseCache`, `EligibilityTrace`, `VirtualBullet`. Constants: `N_IN = 870`, `N_OUT = 2`, `N_CLAUSES = 64` (32 pos + 32 neg), `N_STATES = 15`, `T = 32.0`, `S = 4.0`, `RESID_MAX = 80.0`, `TRACE_MAX_AGE = 40`.
- **Procs:** `stateIdx`, `clausePolarity`, `evalClause`, `makeLiterals`, `computeVote`, `initTsetlinNet`, `forward`, `forwardWithCache`, `learnOne`, `learn`. [FACT]
- **Honest assessment of the learning core** [FACT]:
- It is a **regression** TM (continuous `(cx, cy)` scaled to `±RESID_MAX`), *not* Granmo's multi-class classifier; the output is `vote / T * RESID_MAX`.
- Its "Type Ia" (`cOut == 1`) and "Type Ib" (`else`) branches execute **identical bodies** (both grow toward the current input), so the split is cosmetic. The counteracting `(c=0, literal=1) → toward-Exclude` force is absent.
- Its "Bug 3 fix" Type II decrements **included** false literals (`if literals[lit] == 0 and st > 0`) rather than incrementing *excluded* false literals; when `cOut == 1` that condition is unsatisfiable, so the branch is dead code (same defect as the current gun — see `tm-deb-assessment.md` §6.3).
- **Verdict:** reusable as a **prototype scaffold** (types, vote, clause cache, eligibility trace), but **not** as a correct learning core. The fixed feedback rule from Section A.6 step 1 is a prerequisite.
### `BNNBot_garage/src/wisard_predictor.nim` — **not a TM**
- Types `WiSARDNet`, `WaveTrace`, `VirtualBullet`; procs `initWiSARD(seed)`, `computeAddresses`, `predictCorrection`, `learnCorrection`; constants `K = 14`, `N_NEURONS = 50`, `LUT_SIZE = 16384`. [FACT]
- It is a WiSARD/n-tuple **LUT regressor** averaging `(sumX/count, sumY/count)`. [FACT] Reusable only as a **non-TM baseline**; it regresses `(dx, dy)` directly, so it bypasses the binary-machinery question entirely. It has no interpretable clauses.
**Summary:** the *encoder* is reusable as-is (minus the two stub decoders). The TM code is reusable as reference/scaffolding but is a regression TM with broken feedback rules, not a correct classifier. Both already live inside the gun, so the practical value of `BNNBot_garage` today is *historical reference*, not an importable library.
## B.4 Concrete, falsifiable experiment
**Goal:** prove a TM can recover a *known* high-level temporal rule from the frame-stacked encoding — and measure whether it can do so sparsely.
**Setup [INFERENCE unless noted]:**
1. Define a synthetic enemy with a rule that is (a) deterministic, (b) temporal, (c) expressible as a conjunction over the encoded frame fields. Two candidates:
- **Rule E (energy-gated turn):** the enemy turns hard left/right on tick `t` iff its energy has **dropped below a fixed threshold within the last `k` frames**. [FACT] energy is an 11-bit Gray-coded field per frame; [FACT] the gun currently feeds *self* energy, so this rule needs the `enemyEnergy` fix (B.2).
- **Rule D (fixed perpendicular drift):** the enemy holds a constant heading offset relative to the bearing for `k` frames → the velocity/heading fields repeat across the window.
2. Feed the 10-frame stacked 870-bit vector to a **standalone classifier TM** (not the gun's 2-output regression head): multi-class over `{turn-left, turn-right, no-turn}` or `{drift, no-drift}`.
3. Train/test split, fixed seed, and a majority-class baseline reported alongside.
**Accuracy / sparsity bar that would prove the TM found the rule [INFERENCE — proposed thresholds, not measured]:**
- **Accuracy:** on held-out noiseless frames, a correct rule should be recovered near-ceiling (propose **≥ 95 %** on the deterministic rule, vs 50 % chance for a balanced binary task, and vs the majority-class rate). Anything at the majority-class rate means no rule was learned.
- **Sparsity:** the clauses that win should each include a **small single-digit** number of literals, and the count should be stable across runs. [FACT] The current gun's clauses saturate at ~131 literals/clause when the feedback rule is broken (assessment §6.4); that is the failure signature to watch for.
- **Interpretability check:** decode the highest-weight clause's literals via `tmStateIdx` back to `(frame, field, bit)` and verify they land on the energy/velocity/heading fields and the expected frames. This is the test that it found a *high-level* rule, not a coincidence.
**Can the current gun represent such a rule, once its feedback tables are fixed?**
- **Clause width:** [FACT] there is no width cap; a clause may include any of `TM_N_LITERALS = 1740` literals (`st > 0`). Representability is not blocked by clause width.
- **Feature count:** [FACT] 870 input bits / 1740 literals, with the energy field spanning 11 bits × 10 frames. [INFERENCE] Gray-coding means a monotone threshold ("energy < X") is **not one literal** — it is a sub-pattern of the 11-bit code; a single AND clause captures a specific bit pattern, and the class team as a whole is what learns the threshold. So the rule is representable by the team, not by a single clause.
- **Data volume:** [FACT] `docs/rl-algorithm-choice.md` states episodes are ~30–180 ticks; the gun only creates a trace on ticks with a valid power bin, and `trainedShots = 2141` (cited as observed in `tm-deb-assessment.md` §4) is the accumulated sample count. [UNKNOWN] whether that is enough for convergence — this is exactly what the experiment settles.
- **Head shape:** [FACT] the gun's head is `TM_N_OUT = 2` regression outputs (`cx, cy`), not discrete rule classes. So **the gun as-is cannot be pointed at the Rule E/D classification task** — the experiment must use a standalone classifier TM. The gun can only benefit from these findings indirectly.
---
# SECTION C — Delayed reward: the movement case
## C.1 Delayed supervision vs delayed reward — a hard distinction
- **Delayed supervision (the gun):** the outcome (a shot resolving) arrives 4–128 ticks after the fire tick. [FACT] `bulletSpeed(power) = 20 − 3·power` (`gun_interface.nim`) and `PowerBins = [1.0, 1.5, 2.0, 3.0]` → speeds `17, 15.5, 14, 11` px/tick (`virtual_bullets.nim`); at typical ≈400 px engagement that is ≈23–36 ticks, and the gun's own comment notes a power-3 long shot can take ~128 ticks. **But the outcome is exactly attributable per shot**: `FeedbackEvent` carries `fireTick` and `powerBin`, and `onResult` indexes `tmTraceSlot(e.fireTick, binIdx)` to the exact stored trace. [FACT] So the label is delayed, not confounded. **Effective `λ = 1`; no discounting needed.** [FACT, matching `tm-deb-assessment.md` §5]
- **Delayed reward (movement):** one sparse outcome per round (win/loss/rank/score) with **no** action-to-outcome mapping. You cannot say which of the hundreds of per-tick turn/thrust decisions caused a win. [INFERENCE] This is the case where credit assignment — discounting an eligibility buffer, or a TD/REINFORCE estimator — is the correct tool. [FACT] `tm-deb-assessment.md` §7 already identifies a learned movement module as the repo's genuine delayed-credit-assignment case.
[INFERENCE] **Discounting is only justified for the second case.** Applying it to the gun would delete exactly-attributable long-range samples (assessment §4: `γ^Δt = 4.4e-7` at `Δt = 90`).
## C.2 What the movement layer actually does today
[FACT] The movement modules in `common_libs/movements/` are: `the_floor_is_lava.nim`, `phantom_meteor.nim`, `wave_surfer.nim`, `minimum_risk.nim`, `oscillator.nim`, `random_oscillator.nim`, `rammer.nim`.
- **[FACT] The live mover is TFIL.** `ModularBot_garage/src/ModularBot.nim` declares `mover: TFILModule`, initialises `mover: TFILModule(debugGraphics: true)`, and imports `movements/the_floor_is_lava`.
- **[FACT] TFIL is hand-tuned with no learnable parameters.** Every coefficient is a `const`: `BulletCore = 10.0`, `BulletAura = 5.0`, `EnemyCore = 40.0`, `EnemyAura = 10.0`, `CorridorHeat = 20.0`, `WallHotness = 30.0`, `WallRadiance = 10.0`, `PillarHotness = 30.0`, `PillarRadiance = 10.0`, `CommitTicks = 15`, `MinCommitTicks = 5`, `DangerReplanThreshold = 25.0`, `CoolestLevels = 2`, `MaxTrackedBullets = 20`, `GridSize = 36.0`, `MaxSpeed = 8.0`. `computeMove` builds a lava grid, scores reachable tiles, commits for `CommitTicks`, and picks among safe tiles with a **random** tie-break (`rand(candidates.high)`). No reward ever reaches it; `resetRound` only clears transients.
- **[FACT] Other modules:** `minimum_risk.nim` — all scoring weights are `const` (`KEnemy = 1.0`, `KWall = 0.5`, `KCorner = 0.8`, `KTravel = 0.003`). `oscillator.nim` / `random_oscillator.nim` / `rammer.nim` — fixed period/threshold constants. `wave_surfer.nim` has an observed-danger histogram `bins: array[31, float64]` incremented as enemy waves pass, and **`resetRound` does not clear `bins`** (it persists across rounds). `phantom_meteor.nim` likewise keeps its `DangerHistogram` across rounds. [FACT]
- **[INFERENCE] These histograms are counters, not reward-optimized parameters:** no reward is ever propagated into them; they are incremented purely from observed bullet geometry. There is **no existing learnable movement parameter trained by any outcome signal**.
**Reward signal available from the bot API [FACT]:**
- `ResultsForBot` (`robocode_tankroyale_botapi/schemas.nim`) exposes per round: `rank`, `survival`, `lastSurvivorBonus`, `bulletDamage`, `bulletKillBonus`, `ramDamage`, `ramKillBonus`, `totalScore`, plus season counters `firstPlaces`/`secondPlaces`/`thirdPlaces`. Delivered in `onRoundEnded` via `RoundEndedEventForBot`. [FACT]
- `ModularBot.onRoundEnded` already reads `e.results.totalScore` and `e.results.bulletDamage` (it logs them to `/tmp/gun_stats.jsonl`), so the plumbing for a per-round reward already exists. [FACT]
- Dense per-tick proxies also exist: `getEnergy()` (`buildState` reads it), `onHitByBullet` (damage taken), `onBulletHit` (damage dealt), `onBotDeath`. [FACT] `SAC_LSTM_Bot_garage/src/SAC_LSTM_Bot/rewards.nim` shows a hand-designed `computeReward(...)` combining win/loss (+20/−10), bullet damage, ram penalty, charge penalty, and wall ticks. [FACT]
## C.3 Minimal delayed-reward learning design
[INFERENCE] A deliberately small first design, one live parameter at a time:
1. **One live knob.** Pick a single low-dimensional movement decision with a small discrete action set — e.g. a bias toward/away from a tile class in TFIL's scoring, or an "engagement distance / aggression" scalar (mirroring `wave_surfer`'s `WS_PrefDist`), or the reversal threshold in the oscillators. Keep it to a few discrete actions so a per-action clause team is trainable.
2. **Bounded experience buffer.** Each tick, store `(encodedStateBits, actionIdx)` in a fixed-size ring (a few hundred ticks). Encode the state with the existing 870-bit encoder from `binary_encoding.nim` (or a movement-relevant subset); this is the same eligibility-buffer shape the gun already uses, just keyed by tick instead of `(fireTick, powerBin)`. [INFERENCE]
3. **Dispatch at round end.** In `onRoundEnded`, compute a scalar reward `r` from `ResultsForBot` (`totalScore`/`rank`/`survival`; optionally plus a dense term as in `rewards.nim`), then replay the buffer: for each buffered `(state, action)` apply Type I/II feedback for the chosen action's clause team, weighted by `γ^(T − t)`. [INFERENCE]
4. **Credit-assignment honesty.** One reward per round spread over ~30–180 ticks is extremely sparse and high-variance. [FACT for the tick range from `docs/rl-algorithm-choice.md`; INFERENCE for the consequence.] The `γ^(T − t)` decay is exactly a heuristic for "earlier actions are less likely to be responsible", and it does not make the attribution correct — it only hedges. [INFERENCE]
**How many rounds before a signal is detectable [INFERENCE — standard two-proportion power calculation, computed by hand; no runtime measurement]:**
For a win-rate reward, two-proportion sample size per arm
`n = (z_{α/2} + z_β)² · [p₁(1−p₁) + p₂(1−p₂)] / (p₁ − p₂)²`, with `α = 0.05`, power `0.80` ⇒ `(1.96 + 0.8416)² ≈ 7.85`.
- Detect `0.50 → 0.55` (5 pts): `n ≈ 7.85 · (0.25 + 0.2475) / 0.0025 ≈ 1560` rounds **per arm**.
- Detect `0.50 → 0.60` (10 pts): `n ≈ 7.85 · (0.25 + 0.24) / 0.01 ≈ 385` rounds **per arm**.
These are *lower bounds* for a between-condition comparison; a single adaptive learner is worse (nonstationarity, exploration cost, interference), so treat them as optimistic. [INFERENCE]
**Feasibility.** Given the brief's figure of ~1 round per few seconds, 385 rounds ≈ 20–30 min, and 1560 rounds ≈ 1.5–2 h per arm. [FACT] `docs/rl-algorithm-choice.md` also states 10–35 rounds per match and a training window of "~few hundred milliseconds" between rounds. [INFERENCE] So an **offline, between-battle** experiment is feasible within hours; **online learning inside a single match is not** (a match has ~10–35 rounds). [UNKNOWN] the actual wall-clock duration of one round in this repo's harness — no file I read records it, so the minutes/hours above are conditional arithmetic, not measured.
## C.4 What is UNPROVEN, and the experiment that would settle it
- **[UNKNOWN] Does the per-round reward carry enough signal to be learnable at all?** It may be dominated by opponent identity, arena randomness, and starting positions rather than by the movement knob.
- **[UNKNOWN] The wall-clock cost per round**, hence total experiment time and whether the round budget is realistic.
- **[UNKNOWN] Whether γ-discounted per-tick feedback beats a fixed hand-tuned policy** (TFIL) at all, and by how much — no measurement exists.
- **[UNKNOWN] Whether one live knob is enough** to move `totalScore` measurably, or whether the effect is buried in variance.
**Experiment that would settle it [INFERENCE]:**
1. **Offline counterfactual first (cheap gate).** Replay a fixed corpus of recorded enemy movement tapes. For each candidate action bias, simulate the resulting movement and compute the reward proxy. If **no** single-knob policy beats TFIL's fixed policy offline, do not spend the online round budget — the signal is not there. If one does, use its effect size to size the online run.
2. **Online A/B, seeded.** Fixed opponent set and seeds; arm A = current TFIL, arm B = TM-biased TFIL. Run `N` rounds per arm from the power calculation above (start at ≈ 385/arm for a 10-pt effect, expand for smaller effects). Pre-register `totalScore`/`rank`/`survival` as the metrics and report confidence intervals, not point estimates.
3. **Instrument variance.** Log per-round reward and its spread *before* learning; if the round-to-round standard deviation is large relative to any plausible effect, report that directly as the blocker rather than running an under-powered test.
> **Explicit non-import.** None of the above relies on `docs/papers/tm-deb-paper.pdf`. The only external claim used is the standard power calculation (C.3). The TM-DEB idea of discounting a replayed buffer is *conceptually* the right shape for this case (as `tm-deb-assessment.md` §7 already noted), but no result, equation, or number from that deleted document is treated as evidence here.