Files
SirRoboGarage/docs/research/tm-learning-tracks.md
T
SirStone 54e9757567 docs(research): TM learning tracks + note that the TM-DEB source was deleted
tm-learning-tracks.md covers three things, all marked [FACT]/[INFERENCE]/
[UNKNOWN]:
- Section A: the Tsetlin gun's label is measured against the wrong baseline.
  predX = linearX + cx, so rx = actual - predX = delta - cx, and inside
  tmLearnOne error = residual - predicted = (delta - cx) - cx = delta - 2cx.
  The fixed point is cx = delta/2 -- HALF the correction needed, even with
  perfect Granmo feedback. Fix: store linearX/linearY in TmTrace and train
  on delta. Also: hits zero the label instead of carrying their true
  residual, and the per-clause step is magnitude-blind.
- Section B: what a TM is actually good at (AND-clauses over binary
  literals, readable output) and why this repo suits it -- the gun already
  builds an 83-bit x 10-frame Gray-coded window (870 bits). Includes a
  falsifiable known-rule benchmark proposal.
- Section C: delayed-reward learning belongs to the MOVEMENT layer, not the
  gun. The gun's outcome is delayed but exactly pairable via
  (fireTick, powerBin), so its effective lambda is 1 and discounting would
  only destroy information.

Also records that docs/papers/tm-deb-paper.pdf was deleted by the user as
AI-generated and unverifiable, while the Granmo-based feedback diff in
tm-deb-assessment.md stands on its own.
2026-09-20 23:28:31 +02:00

32 KiB
Raw Blame History

Tsetlin Machine learning tracks — target shape, high-level patterns, delayed reward

Date: 2026-09-20 Scope: Design notes only. No .nim file was edited, nothing was built, no battle was run, no /tmp artifact was touched. Every claim is tagged by its evidence class. Author's method: Read common_libs/guns/tsetlin.nim, common_libs/gun_harness/gun_interface.nim, common_libs/gun_harness/virtual_bullets.nim, the common_libs/movements/* modules, ModularBot_garage/src/ModularBot.nim, BNNBot_garage/src/{binary_encoding,tsetlin_predictor,wisard_predictor}.nim, SAC_LSTM_Bot_garage/src/SAC_LSTM_Bot/rewards.nim, docs/rl-algorithm-choice.md, docs/research/tm-deb-assessment.md, and the installed robocode_tankroyale_botapi schemas.nim / bot.nim.

Evidence legend

  • [FACT] — verified by reading the cited code or paper at the cited location.
  • [INFERENCE] — reasoned from facts; not measured.
  • [UNKNOWN] — needs an experiment; no claim is made.

Provenance caveat. docs/papers/tm-deb-paper.pdf (TM-DEB) was deleted by the user as AI-generated and unverifiable (see the editorial note in docs/research/tm-deb-assessment.md). Nothing in this document uses it as evidence, and the Section C delayed-reward design does not adopt any TM-DEB mechanism.


SECTION A — The Tsetlin gun's target is the wrong shape (do this first)

A.0 Correction to the brief, before anything else

The brief states that "the gun currently learns a real-valued (cx,cy) pixel correction from a binary hit/miss outcome." That does not match HEAD. The current training site already computes a signed residual. [FACT] common_libs/guns/tsetlin.nim, proc onResult*:

# Directional residual: actual enemy pos minus our prediction
# On hit residual is 0 (we were right); on miss we push toward actual position.
let rx = if e.hit: 0.0 else: clamp(e.actualX - t.predX, -TM_RESID_MAX, TM_RESID_MAX)
let ry = if e.hit: 0.0 else: clamp(e.actualY - t.predY, -TM_RESID_MAX, TM_RESID_MAX)
let lits = tmMakeLiterals(t.input)
g.net.tmLearnOne(0, lits, t.cache, rx)
g.net.tmLearnOne(1, lits, t.cache, ry)

The previous revision (git show e536900~1:common_libs/guns/tsetlin.nim) did the same. So the binary hit flag is not the label anywhere in the gun's history; it is only the fitness signal in virtual_bullets.nim. The brief's concern is therefore a non-issue here — but a different and sharper target-shape defect is present, derived below from the code.

A.1 Why a hit/miss-only label would be ill-posed (stated for completeness)

If the label were only hit ∈ {0,1}:

  • A hit says the total 2D prediction was good; a miss says it was bad. Neither says which axis was off, nor by how much, nor which direction. [FACT — by construction of a scalar label vs a 2D continuous target.]
  • Fitting a signed 2D correction (two real outputs, each with sign and magnitude) to a scalar binary outcome would need the direction to be invented from something else (e.g. the predicted-vs-actual geometry). It would not be a well-posed supervised problem.

This is not what the gun does, per A.0.

A.2 The harness hands the gun an exact signed residual

[FACT] common_libs/gun_harness/gun_interface.nim defines FeedbackEvent with actualX, actualY, prediction, plus fireTick, powerBin, missDistance, hit. At resolution, common_libs/gun_harness/virtual_bullets.nim (tickBullets) reads the target's actual position from the enemies table and builds the event:

let missDist = hypot(bx - ex, by - ey)
let fe = FeedbackEvent(
  prediction:  GunPrediction(x: b.aimX, y: b.aimY),
  actualX:     ex, actualY: ey, ... fireTick: b.fireTick, powerBin: b.powerBin, ...)

So residual = actual − prediction is known exactly at resolution time, per axis, with no ambiguity. [FACT]

[INFERENCE] Therefore the gun's learning task is supervised regression with an exact continuous label, not reinforcement learning. There is no hidden credit to assign, so no eligibility discounting is needed — the gun's effective λ = 1.

[FACT] The gun already exploits that: onResult retrieves the exact trace by (fireTick, powerBin) via tmTraceSlot(e.fireTick, binIdx) (ring of TM_TRACE_SLOTS = 1024), checking t.fireTick == e.fireTick and t.powerBin == binIdx. trainedShots counts successful pairings, traceMisses counts failures. This is the repo-native form of an eligibility buffer, and the pairing is exact. [FACT] docs/research/tm-deb-assessment.md §5 reaches the same conclusion: effective λ = 1.

A.3 The real target-shape defect: the label is measured from the wrong baseline

This is the code-grounded finding that replaces the brief's premise.

Let, in common_libs/guns/tsetlin.nim:

  • L = the linear extrapolation baseline, computed in predict() as linearX = state.enemyX + cos(headingRad) * state.enemySpeed * ticksToArrive. [FACT]
  • c = the TM's output correction, tmForwardWithCache returns (vx / TM_T * TM_RESID_MAX, …). [FACT]
  • P = stored prediction, predX = clamp(linearX + cx, …), and TmTrace.predX = predX. [FACT]
  • A = e.actualX at resolution.

At training time rx = A − P = A − L − c. [FACT]

Inside tmLearnOne, predicted is the TM's own output recomputed from the cached clause outputs:

for c in 0..<TM_N_CLAUSES: vote += tmPolarity(c) * float(cache[outIdx * TM_N_CLAUSES + c])
vote = clamp(vote, -TM_T, TM_T)
let predicted = vote / TM_T * TM_RESID_MAX      # == c
let error = residual - predicted                 # == rx - c

So error = (A − L − c) − c = (A − L) − 2c. [FACT — direct substitution.]

[INFERENCE] The update minimizes |error|, whose fixed point is c* = (A − L)/2 — half the correction that makes P hit A. The code passes a label that already contains one copy of the TM's output, then subtracts a second copy. Even with a perfect feedback rule (Granmo Table 2/3), the TM is being asked for a systematically shrunken correction. This is a well-posedness bug in the target definition, not a learning-rate smell.

Corroborating detail: the same baseline pattern exists in the historical BNNBot_garage/src/tsetlin_predictor.nim, proc learn ("residualX/Y: pixel correction needed (actual_target - aimed_point)") and in BNNBot_garage/src/TsetlinBot.nim (residualX = bot.lastEnemyX - bulletX, where bulletX already includes the correction). [FACT] So this is not a one-off typo introduced by the current gun; it is a repeated design slip.

Three further target-shape defects at the same site:

  1. Hits are trained toward zero, not toward their true (small) residual. [FACT] rx = if e.hit: 0.0 else: …. A hit means the 2D missDistance < BotRadius = 18.0 [FACT], but each axis can still be off by up to 18 px. Zeroing both axes discards that information and biases the TM toward under-correction. A hit is not "residual exactly 0"; it is "|(rx, ry)| < 18".
  2. The update is magnitude-blind. [FACT] pFeedback = min(1.0, abs(error) / (2.0 * TM_RESID_MAX)) only scales the probability that a clause receives feedback; the per-clause state step is always ±1 (st = min(st+1, …) / max(st-1, …)). [INFERENCE] Direction is carried by sign(error) and clause polarity, but magnitude must be inferred statistically across many shots — slow, high-variance, and unable to express "this shot needed 40 px".
  3. Range/scale. [FACT] residual is clamped to ±TM_RESID_MAX = 80 and the output vote to ±TM_T = 25 → ±80. Because of the factor-2 defect above, the correction the TM converges to is half the baseline error: to correct a baseline error δ it must drive its output to δ/2, so it saturates at a baseline error of ±160 rather than the ±80 the constant name suggests. [INFERENCE]

A.4 Consequence: the right target is the residual — measured from the baseline

[INFERENCE] The correct supervised target for the TM's correction is δ = A − L (the leftover error of the linear baseline), and onResult should pass δ, not A − P. Equivalently: store linearX, linearY in TmTrace alongside predX, predY and pass e.actualX − t.linearX. The current TmTrace stores predX, predY but not the baseline [FACT], so this requires a small struct change (but still no change to the harness — the baseline is computable at predict time and the actual position already arrives in the event).

A.5 Learning a continuous residual with binary, gradient-free machinery

Tsetlin Machines are binary and gradient-free (a hard constraint for this project). [FACT — the automata in tsetlin.nim are discrete integer states updated by the Tables 2/3-style reward/penalty; there is no gradient.] Two honest ways to learn a continuous residual:

Option A — quantise the residual and classify (recommended).

  • Split each axis's residual range [−R, +R] into K bins. Run two independent multi-class TMs (one per axis), each with K classes; each class gets its own clause team and vote, predict argmax, map the winning bin back to its centre pixel value. This is Granmo's standard multi-class formulation. [INFERENCE — standard TM construction; not present in this repo's code.]
  • Tradeoffs:
    • Resolution vs data: K classes need K teams and enough per-class samples. The residual distribution is heavily concentrated near 0, so most classes are rare. [INFERENCE] Finer bins = better magnitude resolution = more data per bin = slower convergence.
    • Magnitude vs direction: a binned classifier captures both; a pure sign classifier (2 classes/axis) captures only direction and would still need a magnitude estimate (a fixed step, or a second head). [INFERENCE]
    • The no-correction class: a hit should map to an explicit zero bin (e.g. |residual| < BotRadius). Modelling "no correction" as its own class lets the TM emit exactly 0 when it is right, instead of always nudging. Without it, hits trained toward zero (A.3 defect 1) will drag the whole output distribution toward 0.

Option B — predict the SIGN per axis as an independent binary task.

  • 2 classes/axis → 2 clauses-team pairs; the smallest-data option and the most robust head. [INFERENCE] Magnitude must come from elsewhere (fixed step size, or a separate regression/binned head), so this is a partial solution but a good first milestone: prove the TM can learn which way the correction goes before asking it to learn how far.

[INFERENCE] Recommend B as the first experiment, A as the destination: the sign task is cheap and isolates the clause-learning question from the magnitude question.

  1. [HIGH] Fix the feedback rule first. Implement Granmo Table 2/3 faithfully in tmLearnOne (the Type I cOut conditioning and the Type II direction). [FACT] The current Type I ignores cOut, and Type II is unsatisfiable dead code — see docs/research/tm-deb-assessment.md §6.3 and its §6.5 fixes 1–2. Changing the target shape while the update direction is broken is unmeasurable.
  2. [HIGH] Fix the label baseline. Store linearX, linearY in TmTrace; pass δ = actual − linear to tmLearnOne. Remove the factor-2 shrink and stop zeroing residuals on hits (map hits into the zero bin / keep their true small δ).
  3. [MEDIUM] Choose the head. Start with per-axis sign classification (Option B) to validate clause learning, then move to binned multi-class (Option A) with an explicit zero bin.
  4. [LOW] Keep the exact (fireTick, powerBin) pairing. [FACT] It already works (trainedShots / traceMisses). Do not add discounting.

Measurement that would prove it worked [INFERENCE]:

  • Build a deterministic oracle: for a constant-velocity enemy, δ = 0 by construction, so a correct TM must output ≈ 0 (not oscillate). For a constant-acceleration / fixed-perpendicular-drift enemy, δ = actual − linear is analytically known; check the TM's output converges toward δ, not δ/2.
  • Log per shot (linear, actual, prediction, correction); metric = mean absolute correction − δ on held-out ticks, plus hit rate vs the Linear gun (the same adversary, same seed). Success = the factor-2 bias disappears and hit rate beats Linear. [UNKNOWN] the exact accuracy the TM can reach in the repo's data budget.

SECTION B — Can a Tsetlin Machine learn high-level patterns?

B.1 What a TM is actually good at

[FACT] A Tsetlin Machine learns a DNF over binary literals: each clause is a conjunction (AND) of a subset of the 2 × n_input literals (a bit and its complement), and the output vote combines positive- and negative-polarity clauses. In this repo, common_libs/guns/tsetlin.nim tmEvalClause returns 1 iff at least one literal is included and every included literal's value is 1; tmForwardWithCache sums tmPolarity(c) * clause_output into vx, vy. [FACT]

  • Learning is by discrete Tsetlin automata (finite-state machines with Reward/Penalty), updated by Granmo's Tables 2/3; there is no gradient. [FACT]
  • The learned model is human-readable: a clause is a list of included literals, which can be decoded back to (frame index, field, bit). [FACT] The clause state array is indexed by tmStateIdx(outIdx, clause, lit); st > 0 = included. [FACT]
  • [INFERENCE] This is the right tool when the target is genuinely a conjunction of discrete conditions, and the wrong tool when the target is a smooth signed magnitude that must be fitted precisely (Section A).

B.2 The repo is already built for a temporal TM — [FACT]s

  1. Frame-stacked binary encoding already exists in the gun. [FACT] common_libs/guns/tsetlin.nim:
    • TM_FRAME_BITS = 83, TM_SELF_BITS = 40, TM_WINDOW_SIZE = 10, TM_TOTAL_BITS = 83*10 + 40 = 870.
    • Gray coding via tmToGray(value) = value xor (value shr 1) and tmToBits.
    • The window is LIFO (index 0 = newest) and is shifted once per tick (if state.tick != g.frameTick), with a stale comment documenting a past over-shift bug.
    • [INFERENCE] Because every one of the 10 frames is laid out in the same fixed 83-bit block, a single clause can AND literals from different frames, i.e. it can express a temporal conjunction across the last 10 ticks.
  2. The per-frame field layout (from tmEncodeFrame): bearing sin 8 + bearing cos 8 + distance 7 + velocity 5 + heading sin 8 + heading cos 8 + 4 wall distances × 7 + energy 11 = 83 bits. [FACT]
  3. A clause over 10 frames is a high-level pattern. [INFERENCE] "closing distance" = a distance-field bit pattern that changes across frames; "decelerating" = consecutive velocity-field codes ordered by Gray code; "low energy across the last N frames" = an energy-field bit repeated across frames. A conjunction of those literals is exactly the kind of rule a TM clause represents natively. [INFERENCE]

Mismatch to record. [FACT] The gun's tmEncodeFrame call passes state.selfEnergy for the frame's energy field, with the comment # use self energy as proxy (enemy energy not in WorldState). That comment is stale: common_libs/gun_harness/gun_interface.nim line 25 defines WorldState.enemyEnergy*. So the current gun encodes self energy twice and never sees enemy energy — which means a rule like "the enemy turns hard when its energy drops below X" is not even representable by the current gun, independent of its clause learner.

B.3 What is reusable from BNNBot_garage today — proc-level

BNNBot_garage/src/binary_encoding.nim — encoder is fully reusable

  • Types: BinaryVector = array[TOTAL_BITS, uint8], EnemyScanFrame, SelfState, OutputVector.
  • Constants: FRAME_BITS = 83, SELF_BITS = 40, WINDOW_SIZE = 10, TOTAL_BITS = 870, OUTPUT_BITS = 7, AIM_MIN = -60.0, AIM_MAX = 60.0.
  • Reusable procs: encodeFrame(frame: EnemyScanFrame): array[FRAME_BITS, uint8], encodeSelf(self: SelfState): array[SELF_BITS, uint8], encodeFullVector(window, self): BinaryVector, and the utility set formatBinary, popcount, hammingDistance, bitwiseAnd, formatVectorBinary, formatVectorDecimal, formatFrameBinary, encodeOutput(angle): OutputVector, decodeOutput(vec): float. [FACT]
  • NOT reusable (self-documented placeholders that return hardcoded values): formatFrameDecimal (returns "199,199,99,15,199,199,99,99,150") and decodeFrameFields (returns a fixed [0,…,50,50,100] array). [FACT] These are stubs, not decoders — do not build on them.
  • [FACT] The encoder is already inlined into common_libs/guns/tsetlin.nim as tmEncodeFrame / tmEncodeSelf / tmEncodeFullVector (the gun's header says "adapted from BNNBot_garage/src/binary_encoding.nim"), so there is nothing to import; BNNBot is the reference copy. One real divergence: BNNBot's EnemyScanFrame carries enemyEnergy, while the gun's inlined copy takes state.selfEnergy (B.2 mismatch above).

BNNBot_garage/src/tsetlin_predictor.nim — reusable as scaffolding, but it is a regression TM with the same feedback defects

  • Types: TsetlinNet, ClauseCache, EligibilityTrace, VirtualBullet. Constants: N_IN = 870, N_OUT = 2, N_CLAUSES = 64 (32 pos + 32 neg), N_STATES = 15, T = 32.0, S = 4.0, RESID_MAX = 80.0, TRACE_MAX_AGE = 40.
  • Procs: stateIdx, clausePolarity, evalClause, makeLiterals, computeVote, initTsetlinNet, forward, forwardWithCache, learnOne, learn. [FACT]
  • Honest assessment of the learning core [FACT]:
    • It is a regression TM (continuous (cx, cy) scaled to ±RESID_MAX), not Granmo's multi-class classifier; the output is vote / T * RESID_MAX.
    • Its "Type Ia" (cOut == 1) and "Type Ib" (else) branches execute identical bodies (both grow toward the current input), so the split is cosmetic. The counteracting (c=0, literal=1) → toward-Exclude force is absent.
    • Its "Bug 3 fix" Type II decrements included false literals (if literals[lit] == 0 and st > 0) rather than incrementing excluded false literals; when cOut == 1 that condition is unsatisfiable, so the branch is dead code (same defect as the current gun — see tm-deb-assessment.md §6.3).
    • Verdict: reusable as a prototype scaffold (types, vote, clause cache, eligibility trace), but not as a correct learning core. The fixed feedback rule from Section A.6 step 1 is a prerequisite.

BNNBot_garage/src/wisard_predictor.nim — not a TM

  • Types WiSARDNet, WaveTrace, VirtualBullet; procs initWiSARD(seed), computeAddresses, predictCorrection, learnCorrection; constants K = 14, N_NEURONS = 50, LUT_SIZE = 16384. [FACT]
  • It is a WiSARD/n-tuple LUT regressor averaging (sumX/count, sumY/count). [FACT] Reusable only as a non-TM baseline; it regresses (dx, dy) directly, so it bypasses the binary-machinery question entirely. It has no interpretable clauses.

Summary: the encoder is reusable as-is (minus the two stub decoders). The TM code is reusable as reference/scaffolding but is a regression TM with broken feedback rules, not a correct classifier. Both already live inside the gun, so the practical value of BNNBot_garage today is historical reference, not an importable library.

B.4 Concrete, falsifiable experiment

Goal: prove a TM can recover a known high-level temporal rule from the frame-stacked encoding — and measure whether it can do so sparsely.

Setup [INFERENCE unless noted]:

  1. Define a synthetic enemy with a rule that is (a) deterministic, (b) temporal, (c) expressible as a conjunction over the encoded frame fields. Two candidates:
    • Rule E (energy-gated turn): the enemy turns hard left/right on tick t iff its energy has dropped below a fixed threshold within the last k frames. [FACT] energy is an 11-bit Gray-coded field per frame; [FACT] the gun currently feeds self energy, so this rule needs the enemyEnergy fix (B.2).
    • Rule D (fixed perpendicular drift): the enemy holds a constant heading offset relative to the bearing for k frames → the velocity/heading fields repeat across the window.
  2. Feed the 10-frame stacked 870-bit vector to a standalone classifier TM (not the gun's 2-output regression head): multi-class over {turn-left, turn-right, no-turn} or {drift, no-drift}.
  3. Train/test split, fixed seed, and a majority-class baseline reported alongside.

Accuracy / sparsity bar that would prove the TM found the rule [INFERENCE — proposed thresholds, not measured]:

  • Accuracy: on held-out noiseless frames, a correct rule should be recovered near-ceiling (propose ≥ 95 % on the deterministic rule, vs 50 % chance for a balanced binary task, and vs the majority-class rate). Anything at the majority-class rate means no rule was learned.
  • Sparsity: the clauses that win should each include a small single-digit number of literals, and the count should be stable across runs. [FACT] The current gun's clauses saturate at ~131 literals/clause when the feedback rule is broken (assessment §6.4); that is the failure signature to watch for.
  • Interpretability check: decode the highest-weight clause's literals via tmStateIdx back to (frame, field, bit) and verify they land on the energy/velocity/heading fields and the expected frames. This is the test that it found a high-level rule, not a coincidence.

Can the current gun represent such a rule, once its feedback tables are fixed?

  • Clause width: [FACT] there is no width cap; a clause may include any of TM_N_LITERALS = 1740 literals (st > 0). Representability is not blocked by clause width.
  • Feature count: [FACT] 870 input bits / 1740 literals, with the energy field spanning 11 bits × 10 frames. [INFERENCE] Gray-coding means a monotone threshold ("energy < X") is not one literal — it is a sub-pattern of the 11-bit code; a single AND clause captures a specific bit pattern, and the class team as a whole is what learns the threshold. So the rule is representable by the team, not by a single clause.
  • Data volume: [FACT] docs/rl-algorithm-choice.md states episodes are ~30–180 ticks; the gun only creates a trace on ticks with a valid power bin, and trainedShots = 2141 (cited as observed in tm-deb-assessment.md §4) is the accumulated sample count. [UNKNOWN] whether that is enough for convergence — this is exactly what the experiment settles.
  • Head shape: [FACT] the gun's head is TM_N_OUT = 2 regression outputs (cx, cy), not discrete rule classes. So the gun as-is cannot be pointed at the Rule E/D classification task — the experiment must use a standalone classifier TM. The gun can only benefit from these findings indirectly.

SECTION C — Delayed reward: the movement case

C.1 Delayed supervision vs delayed reward — a hard distinction

  • Delayed supervision (the gun): the outcome (a shot resolving) arrives 4–128 ticks after the fire tick. [FACT] bulletSpeed(power) = 20 − 3·power (gun_interface.nim) and PowerBins = [1.0, 1.5, 2.0, 3.0] → speeds 17, 15.5, 14, 11 px/tick (virtual_bullets.nim); at typical ≈400 px engagement that is ≈23–36 ticks, and the gun's own comment notes a power-3 long shot can take ~128 ticks. But the outcome is exactly attributable per shot: FeedbackEvent carries fireTick and powerBin, and onResult indexes tmTraceSlot(e.fireTick, binIdx) to the exact stored trace. [FACT] So the label is delayed, not confounded. Effective λ = 1; no discounting needed. [FACT, matching tm-deb-assessment.md §5]
  • Delayed reward (movement): one sparse outcome per round (win/loss/rank/score) with no action-to-outcome mapping. You cannot say which of the hundreds of per-tick turn/thrust decisions caused a win. [INFERENCE] This is the case where credit assignment — discounting an eligibility buffer, or a TD/REINFORCE estimator — is the correct tool. [FACT] tm-deb-assessment.md §7 already identifies a learned movement module as the repo's genuine delayed-credit-assignment case.

[INFERENCE] Discounting is only justified for the second case. Applying it to the gun would delete exactly-attributable long-range samples (assessment §4: γ^Δt = 4.4e-7 at Δt = 90).

C.2 What the movement layer actually does today

[FACT] The movement modules in common_libs/movements/ are: the_floor_is_lava.nim, phantom_meteor.nim, wave_surfer.nim, minimum_risk.nim, oscillator.nim, random_oscillator.nim, rammer.nim.

  • [FACT] The live mover is TFIL. ModularBot_garage/src/ModularBot.nim declares mover: TFILModule, initialises mover: TFILModule(debugGraphics: true), and imports movements/the_floor_is_lava.
  • [FACT] TFIL is hand-tuned with no learnable parameters. Every coefficient is a const: BulletCore = 10.0, BulletAura = 5.0, EnemyCore = 40.0, EnemyAura = 10.0, CorridorHeat = 20.0, WallHotness = 30.0, WallRadiance = 10.0, PillarHotness = 30.0, PillarRadiance = 10.0, CommitTicks = 15, MinCommitTicks = 5, DangerReplanThreshold = 25.0, CoolestLevels = 2, MaxTrackedBullets = 20, GridSize = 36.0, MaxSpeed = 8.0. computeMove builds a lava grid, scores reachable tiles, commits for CommitTicks, and picks among safe tiles with a random tie-break (rand(candidates.high)). No reward ever reaches it; resetRound only clears transients.
  • [FACT] Other modules: minimum_risk.nim — all scoring weights are const (KEnemy = 1.0, KWall = 0.5, KCorner = 0.8, KTravel = 0.003). oscillator.nim / random_oscillator.nim / rammer.nim — fixed period/threshold constants. wave_surfer.nim has an observed-danger histogram bins: array[31, float64] incremented as enemy waves pass, and resetRound does not clear bins (it persists across rounds). phantom_meteor.nim likewise keeps its DangerHistogram across rounds. [FACT]
  • [INFERENCE] These histograms are counters, not reward-optimized parameters: no reward is ever propagated into them; they are incremented purely from observed bullet geometry. There is no existing learnable movement parameter trained by any outcome signal.

Reward signal available from the bot API [FACT]:

  • ResultsForBot (robocode_tankroyale_botapi/schemas.nim) exposes per round: rank, survival, lastSurvivorBonus, bulletDamage, bulletKillBonus, ramDamage, ramKillBonus, totalScore, plus season counters firstPlaces/secondPlaces/thirdPlaces. Delivered in onRoundEnded via RoundEndedEventForBot. [FACT]
  • ModularBot.onRoundEnded already reads e.results.totalScore and e.results.bulletDamage (it logs them to /tmp/gun_stats.jsonl), so the plumbing for a per-round reward already exists. [FACT]
  • Dense per-tick proxies also exist: getEnergy() (buildState reads it), onHitByBullet (damage taken), onBulletHit (damage dealt), onBotDeath. [FACT] SAC_LSTM_Bot_garage/src/SAC_LSTM_Bot/rewards.nim shows a hand-designed computeReward(...) combining win/loss (+20/−10), bullet damage, ram penalty, charge penalty, and wall ticks. [FACT]

C.3 Minimal delayed-reward learning design

[INFERENCE] A deliberately small first design, one live parameter at a time:

  1. One live knob. Pick a single low-dimensional movement decision with a small discrete action set — e.g. a bias toward/away from a tile class in TFIL's scoring, or an "engagement distance / aggression" scalar (mirroring wave_surfer's WS_PrefDist), or the reversal threshold in the oscillators. Keep it to a few discrete actions so a per-action clause team is trainable.
  2. Bounded experience buffer. Each tick, store (encodedStateBits, actionIdx) in a fixed-size ring (a few hundred ticks). Encode the state with the existing 870-bit encoder from binary_encoding.nim (or a movement-relevant subset); this is the same eligibility-buffer shape the gun already uses, just keyed by tick instead of (fireTick, powerBin). [INFERENCE]
  3. Dispatch at round end. In onRoundEnded, compute a scalar reward r from ResultsForBot (totalScore/rank/survival; optionally plus a dense term as in rewards.nim), then replay the buffer: for each buffered (state, action) apply Type I/II feedback for the chosen action's clause team, weighted by γ^(T − t). [INFERENCE]
  4. Credit-assignment honesty. One reward per round spread over ~30–180 ticks is extremely sparse and high-variance. [FACT for the tick range from docs/rl-algorithm-choice.md; INFERENCE for the consequence.] The γ^(T − t) decay is exactly a heuristic for "earlier actions are less likely to be responsible", and it does not make the attribution correct — it only hedges. [INFERENCE]

How many rounds before a signal is detectable [INFERENCE — standard two-proportion power calculation, computed by hand; no runtime measurement]: For a win-rate reward, two-proportion sample size per arm n = (z_{α/2} + z_β)² · [p₁(1−p₁) + p₂(1−p₂)] / (p₁ − p₂)², with α = 0.05, power 0.80 ⇒ (1.96 + 0.8416)² ≈ 7.85.

  • Detect 0.50 → 0.55 (5 pts): n ≈ 7.85 · (0.25 + 0.2475) / 0.0025 ≈ 1560 rounds per arm.
  • Detect 0.50 → 0.60 (10 pts): n ≈ 7.85 · (0.25 + 0.24) / 0.01 ≈ 385 rounds per arm.

These are lower bounds for a between-condition comparison; a single adaptive learner is worse (nonstationarity, exploration cost, interference), so treat them as optimistic. [INFERENCE]

Feasibility. Given the brief's figure of ~1 round per few seconds, 385 rounds ≈ 20–30 min, and 1560 rounds ≈ 1.5–2 h per arm. [FACT] docs/rl-algorithm-choice.md also states 10–35 rounds per match and a training window of "~few hundred milliseconds" between rounds. [INFERENCE] So an offline, between-battle experiment is feasible within hours; online learning inside a single match is not (a match has ~10–35 rounds). [UNKNOWN] the actual wall-clock duration of one round in this repo's harness — no file I read records it, so the minutes/hours above are conditional arithmetic, not measured.

C.4 What is UNPROVEN, and the experiment that would settle it

  • [UNKNOWN] Does the per-round reward carry enough signal to be learnable at all? It may be dominated by opponent identity, arena randomness, and starting positions rather than by the movement knob.
  • [UNKNOWN] The wall-clock cost per round, hence total experiment time and whether the round budget is realistic.
  • [UNKNOWN] Whether γ-discounted per-tick feedback beats a fixed hand-tuned policy (TFIL) at all, and by how much — no measurement exists.
  • [UNKNOWN] Whether one live knob is enough to move totalScore measurably, or whether the effect is buried in variance.

Experiment that would settle it [INFERENCE]:

  1. Offline counterfactual first (cheap gate). Replay a fixed corpus of recorded enemy movement tapes. For each candidate action bias, simulate the resulting movement and compute the reward proxy. If no single-knob policy beats TFIL's fixed policy offline, do not spend the online round budget — the signal is not there. If one does, use its effect size to size the online run.
  2. Online A/B, seeded. Fixed opponent set and seeds; arm A = current TFIL, arm B = TM-biased TFIL. Run N rounds per arm from the power calculation above (start at ≈ 385/arm for a 10-pt effect, expand for smaller effects). Pre-register totalScore/rank/survival as the metrics and report confidence intervals, not point estimates.
  3. Instrument variance. Log per-round reward and its spread before learning; if the round-to-round standard deviation is large relative to any plausible effect, report that directly as the blocker rather than running an under-powered test.

Explicit non-import. None of the above relies on docs/papers/tm-deb-paper.pdf. The only external claim used is the standard power calculation (C.3). The TM-DEB idea of discounting a replayed buffer is conceptually the right shape for this case (as tm-deb-assessment.md §7 already noted), but no result, equation, or number from that deleted document is treated as evidence here.