diff --git a/docs/research/goto-controller-algorithm.md b/docs/research/goto-controller-algorithm.md new file mode 100644 index 0000000..8ff8de0 --- /dev/null +++ b/docs/research/goto-controller-algorithm.md @@ -0,0 +1,224 @@ +# Goto Controller Algorithm — Research + +**Issue:** #20 +**Branch:** research/goto-controller +**Date:** 2026-08-17 + +--- + +## Problem Statement + +The PPO network will output a target position `(x, y)`. A goto controller must +translate that into per-tick `setTargetSpeed` and `setTurnRate` commands for +the Tank Royale Nim bot API. + +--- + +## Codebase Findings + +### Tank Royale Nim API — no built-in goto + +The library (`tankroyale_botapi` v1.0.1) provides: + +- `setTargetSpeed(speed: float)` — desired speed, clamped to ±8 units/tick. + Server auto-manages acceleration/deceleration via `getNewTargetSpeed`. +- `setTurnRate(rate: float)` — desired turn rate, clamped to `calcMaxTurnRate(speed) = 10 - 0.75 * abs(speed)`. +- `setForward(distance)` / `setBack(distance)` — blocking helpers that use + `gDistanceRemaining` + the server's deceleration model. These are blocking + (call `go()` internally) and therefore cannot be used in the non-blocking + per-tick run loop used by PPO_Bot. + +There is **no built-in `setDistanceRemaining`-style goto**. The controller must +be written from scratch. + +### Physics constants (from `constants.nim` / `utils.nim`) + +| Constant | Value | +|---|---| +| Max speed | 8 units/tick | +| Acceleration | +1 unit/tick² | +| Deceleration | −2 units/tick² (braking is twice as fast) | +| Max turn rate | `10 − 0.75 × |speed|` deg/tick | +| Min turn rate (at max speed) | `10 − 0.75 × 8 = 4` deg/tick | +| `getNewTargetSpeed(maxSpeed, speed, dist)` | already implemented in utils.nim | + +Key implication: **you can turn faster while slow**. Turn-then-drive lets the +bot use full 10°/tick turn rate, but wastes ticks stopped. Driving-while-turning +is smooth but limited to 4°/tick at top speed. + +### Coordinate system + +North = 0°, clockwise. `directionTo` in `utils.nim` returns a bearing in +`[0, 360)`. `bearingTo` returns a signed relative bearing in `(-180, 180]`. + +--- + +## Approaches Considered + +### A — Turn-then-drive (sequential) + +Stop → turn to face target → drive full speed → brake. + +- Simple to implement. +- Very slow: wastes ticks turning at zero speed then decelerating. +- Produces jerky, non-smooth movement — bad as a controller layer. + +### B — Proportional navigation (continuous per-tick) + +Each tick: compute bearing to target, set turn rate proportional to bearing +error, set speed based on distance remaining. + +- Standard Robocode idiom. Very common in published bots. +- Does not make the forward-vs-reverse decision optimally. +- Can overshoot if gains are too high; can be sluggish if too low. + +### C — Arc/pursuit steering (proportional + speed-dependent turn limit) + +Like B, but explicitly clamps turn rate to `calcMaxTurnRate(currentSpeed)` and +scales speed down when the heading error is large (so the bot slows to increase +turn authority). + +- Handles Tank Royale's speed-dependent turn rate correctly. +- Naturally smooth. +- Still needs explicit forward/reverse decision. + +### D — Forward-vs-reverse decision + proportional steering (recommended) + +Extend C with the classic Robocode "should I go backward?" heuristic: +if `|bearingError| > 90°`, it is faster to reverse and face the target with +the rear than to turn more than 90° forward. Flip target speed sign and add +180° to the bearing before computing turn rate. + +This is the approach used by high-quality Robocode 1 bots (e.g. RaikoMX, +Aristocles) and it trivially maps to Tank Royale's API. + +--- + +## Recommended Algorithm + +### Decision: forward or reverse? + +``` +bearing = normalizeRelativeAngle(directionTo(x, y) - direction) +if abs(bearing) > 90.0: + # Going backward is cheaper + direction_sign = -1 + effective_bearing = normalizeRelativeAngle(bearing + 180.0) +else: + direction_sign = +1 + effective_bearing = bearing +``` + +### Turn rate + +Apply full proportional turn rate toward the effective bearing: + +``` +max_turn = 10.0 - 0.75 * abs(currentSpeed) +turnRate = clamp(effective_bearing, -max_turn, max_turn) +``` + +`effective_bearing` acts as both direction and magnitude: if the error is +small, the turn rate is small (smooth approach); if large, it clamps to max +(fastest possible turn). + +### Target speed + +Use `getNewTargetSpeed` (already in `utils.nim`) to determine the speed +that will arrive at the target with zero velocity: + +``` +dist = distanceTo(x, y) +raw_speed = getNewTargetSpeed(MAX_SPEED, currentSpeed, dist) +targetSpeed = direction_sign * raw_speed +``` + +This reuses the exact deceleration model the server uses, so the bot always +brakes at the right time with no overshoot. + +### Stop condition + +``` +if dist < ARRIVAL_THRESHOLD: # e.g. 18.0 (= BOT_RADIUS) + targetSpeed = 0.0 + turnRate = 0.0 +``` + +### Full pseudocode (one tick) + +```nim +proc gotoTick*(tx, ty, x, y, direction, currentSpeed: float): + tuple[targetSpeed, turnRate: float] = + + let dist = distanceTo(x, y, tx, ty) + + if dist < ARRIVAL_THRESHOLD: + return (0.0, 0.0) + + let rawBearing = normalizeRelativeAngle(directionTo(x, y, tx, ty) - direction) + + let (dirSign, effBearing) = + if abs(rawBearing) > 90.0: + (-1.0, normalizeRelativeAngle(rawBearing + 180.0)) + else: + (1.0, rawBearing) + + let maxTurn = 10.0 - 0.75 * abs(currentSpeed) + let turnRate = effBearing.clamp(-maxTurn, maxTurn) + + let rawSpeed = getNewTargetSpeed(MAX_SPEED, abs(currentSpeed), dist) + let targetSpeed = dirSign * rawSpeed + + return (targetSpeed, turnRate) +``` + +Call once per tick from the `run` loop, pass results to `setTargetSpeed` / +`setTurnRate`. + +--- + +## Why not pure proportional navigation (option B)? + +Option B without the speed-dependent turn clamp will attempt to command more +turn rate than the server will honor at high speed — it does the right thing +emergently but wastes the gap. Explicitly scaling turn rate with +`calcMaxTurnRate(speed)` is more intentional and matches the physics exactly. +This is already coded in `actions.nim` (`r1 * (10.0 - 0.75 * abs(currentSpeed))`), +so the pattern is established in the codebase. + +--- + +## Why reuse `getNewTargetSpeed` from utils.nim? + +It already encodes the exact asymmetric acceleration/deceleration model +(accel +1, decel −2 per tick). Reimplementing distance-based speed management +from scratch would duplicate this and risk drift. Import it directly. + +--- + +## Forward/Reverse optimality + +The 90° threshold is the exact breakeven point: + +- Turning 91° forward takes ≥10 ticks at slow speed + travel time. +- Reversing 89° (i.e. 180−91=89° effective turn) takes fewer ticks total + for any distance large enough to matter. +- For very short distances (< ~36 units) the bot will decelerate before the + turn completes anyway; the threshold still works because the speed penalty + applies equally to both cases. + +For a controller layer that feeds a neural network's goto target, sub-optimal +behavior on very short hops is acceptable — the network will learn to avoid +issuing tiny hops. + +--- + +## Sources / References + +- Tank Royale Nim API source: `tankroyale_botapi/utils.nim`, `bot.nim`, + `constants.nim` (v1.0.1, installed at `~/.nimble/pkgs2/`). +- Robocode wiki — "Proportional navigation" and "Should I go backward?" + heuristic: widely documented in the Robocode community (e.g. RoboWiki + `BasicSurfer`, `RaikoMX` source). +- Tank Royale physics spec: confirmed against `ACCELERATION = 1.0`, + `ABS_DECELERATION = 2.0` in `constants.nim`.