docs(research): goto controller algorithm for issue #20
Covers forward/reverse decision, proportional steering with speed-dependent turn rate clamping, and deceleration using the existing getNewTargetSpeed util. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,224 @@
|
||||
# Goto Controller Algorithm — Research
|
||||
|
||||
**Issue:** #20
|
||||
**Branch:** research/goto-controller
|
||||
**Date:** 2026-08-17
|
||||
|
||||
---
|
||||
|
||||
## Problem Statement
|
||||
|
||||
The PPO network will output a target position `(x, y)`. A goto controller must
|
||||
translate that into per-tick `setTargetSpeed` and `setTurnRate` commands for
|
||||
the Tank Royale Nim bot API.
|
||||
|
||||
---
|
||||
|
||||
## Codebase Findings
|
||||
|
||||
### Tank Royale Nim API — no built-in goto
|
||||
|
||||
The library (`tankroyale_botapi` v1.0.1) provides:
|
||||
|
||||
- `setTargetSpeed(speed: float)` — desired speed, clamped to ±8 units/tick.
|
||||
Server auto-manages acceleration/deceleration via `getNewTargetSpeed`.
|
||||
- `setTurnRate(rate: float)` — desired turn rate, clamped to `calcMaxTurnRate(speed) = 10 - 0.75 * abs(speed)`.
|
||||
- `setForward(distance)` / `setBack(distance)` — blocking helpers that use
|
||||
`gDistanceRemaining` + the server's deceleration model. These are blocking
|
||||
(call `go()` internally) and therefore cannot be used in the non-blocking
|
||||
per-tick run loop used by PPO_Bot.
|
||||
|
||||
There is **no built-in `setDistanceRemaining`-style goto**. The controller must
|
||||
be written from scratch.
|
||||
|
||||
### Physics constants (from `constants.nim` / `utils.nim`)
|
||||
|
||||
| Constant | Value |
|
||||
|---|---|
|
||||
| Max speed | 8 units/tick |
|
||||
| Acceleration | +1 unit/tick² |
|
||||
| Deceleration | −2 units/tick² (braking is twice as fast) |
|
||||
| Max turn rate | `10 − 0.75 × |speed|` deg/tick |
|
||||
| Min turn rate (at max speed) | `10 − 0.75 × 8 = 4` deg/tick |
|
||||
| `getNewTargetSpeed(maxSpeed, speed, dist)` | already implemented in utils.nim |
|
||||
|
||||
Key implication: **you can turn faster while slow**. Turn-then-drive lets the
|
||||
bot use full 10°/tick turn rate, but wastes ticks stopped. Driving-while-turning
|
||||
is smooth but limited to 4°/tick at top speed.
|
||||
|
||||
### Coordinate system
|
||||
|
||||
North = 0°, clockwise. `directionTo` in `utils.nim` returns a bearing in
|
||||
`[0, 360)`. `bearingTo` returns a signed relative bearing in `(-180, 180]`.
|
||||
|
||||
---
|
||||
|
||||
## Approaches Considered
|
||||
|
||||
### A — Turn-then-drive (sequential)
|
||||
|
||||
Stop → turn to face target → drive full speed → brake.
|
||||
|
||||
- Simple to implement.
|
||||
- Very slow: wastes ticks turning at zero speed then decelerating.
|
||||
- Produces jerky, non-smooth movement — bad as a controller layer.
|
||||
|
||||
### B — Proportional navigation (continuous per-tick)
|
||||
|
||||
Each tick: compute bearing to target, set turn rate proportional to bearing
|
||||
error, set speed based on distance remaining.
|
||||
|
||||
- Standard Robocode idiom. Very common in published bots.
|
||||
- Does not make the forward-vs-reverse decision optimally.
|
||||
- Can overshoot if gains are too high; can be sluggish if too low.
|
||||
|
||||
### C — Arc/pursuit steering (proportional + speed-dependent turn limit)
|
||||
|
||||
Like B, but explicitly clamps turn rate to `calcMaxTurnRate(currentSpeed)` and
|
||||
scales speed down when the heading error is large (so the bot slows to increase
|
||||
turn authority).
|
||||
|
||||
- Handles Tank Royale's speed-dependent turn rate correctly.
|
||||
- Naturally smooth.
|
||||
- Still needs explicit forward/reverse decision.
|
||||
|
||||
### D — Forward-vs-reverse decision + proportional steering (recommended)
|
||||
|
||||
Extend C with the classic Robocode "should I go backward?" heuristic:
|
||||
if `|bearingError| > 90°`, it is faster to reverse and face the target with
|
||||
the rear than to turn more than 90° forward. Flip target speed sign and add
|
||||
180° to the bearing before computing turn rate.
|
||||
|
||||
This is the approach used by high-quality Robocode 1 bots (e.g. RaikoMX,
|
||||
Aristocles) and it trivially maps to Tank Royale's API.
|
||||
|
||||
---
|
||||
|
||||
## Recommended Algorithm
|
||||
|
||||
### Decision: forward or reverse?
|
||||
|
||||
```
|
||||
bearing = normalizeRelativeAngle(directionTo(x, y) - direction)
|
||||
if abs(bearing) > 90.0:
|
||||
# Going backward is cheaper
|
||||
direction_sign = -1
|
||||
effective_bearing = normalizeRelativeAngle(bearing + 180.0)
|
||||
else:
|
||||
direction_sign = +1
|
||||
effective_bearing = bearing
|
||||
```
|
||||
|
||||
### Turn rate
|
||||
|
||||
Apply full proportional turn rate toward the effective bearing:
|
||||
|
||||
```
|
||||
max_turn = 10.0 - 0.75 * abs(currentSpeed)
|
||||
turnRate = clamp(effective_bearing, -max_turn, max_turn)
|
||||
```
|
||||
|
||||
`effective_bearing` acts as both direction and magnitude: if the error is
|
||||
small, the turn rate is small (smooth approach); if large, it clamps to max
|
||||
(fastest possible turn).
|
||||
|
||||
### Target speed
|
||||
|
||||
Use `getNewTargetSpeed` (already in `utils.nim`) to determine the speed
|
||||
that will arrive at the target with zero velocity:
|
||||
|
||||
```
|
||||
dist = distanceTo(x, y)
|
||||
raw_speed = getNewTargetSpeed(MAX_SPEED, currentSpeed, dist)
|
||||
targetSpeed = direction_sign * raw_speed
|
||||
```
|
||||
|
||||
This reuses the exact deceleration model the server uses, so the bot always
|
||||
brakes at the right time with no overshoot.
|
||||
|
||||
### Stop condition
|
||||
|
||||
```
|
||||
if dist < ARRIVAL_THRESHOLD: # e.g. 18.0 (= BOT_RADIUS)
|
||||
targetSpeed = 0.0
|
||||
turnRate = 0.0
|
||||
```
|
||||
|
||||
### Full pseudocode (one tick)
|
||||
|
||||
```nim
|
||||
proc gotoTick*(tx, ty, x, y, direction, currentSpeed: float):
|
||||
tuple[targetSpeed, turnRate: float] =
|
||||
|
||||
let dist = distanceTo(x, y, tx, ty)
|
||||
|
||||
if dist < ARRIVAL_THRESHOLD:
|
||||
return (0.0, 0.0)
|
||||
|
||||
let rawBearing = normalizeRelativeAngle(directionTo(x, y, tx, ty) - direction)
|
||||
|
||||
let (dirSign, effBearing) =
|
||||
if abs(rawBearing) > 90.0:
|
||||
(-1.0, normalizeRelativeAngle(rawBearing + 180.0))
|
||||
else:
|
||||
(1.0, rawBearing)
|
||||
|
||||
let maxTurn = 10.0 - 0.75 * abs(currentSpeed)
|
||||
let turnRate = effBearing.clamp(-maxTurn, maxTurn)
|
||||
|
||||
let rawSpeed = getNewTargetSpeed(MAX_SPEED, abs(currentSpeed), dist)
|
||||
let targetSpeed = dirSign * rawSpeed
|
||||
|
||||
return (targetSpeed, turnRate)
|
||||
```
|
||||
|
||||
Call once per tick from the `run` loop, pass results to `setTargetSpeed` /
|
||||
`setTurnRate`.
|
||||
|
||||
---
|
||||
|
||||
## Why not pure proportional navigation (option B)?
|
||||
|
||||
Option B without the speed-dependent turn clamp will attempt to command more
|
||||
turn rate than the server will honor at high speed — it does the right thing
|
||||
emergently but wastes the gap. Explicitly scaling turn rate with
|
||||
`calcMaxTurnRate(speed)` is more intentional and matches the physics exactly.
|
||||
This is already coded in `actions.nim` (`r1 * (10.0 - 0.75 * abs(currentSpeed))`),
|
||||
so the pattern is established in the codebase.
|
||||
|
||||
---
|
||||
|
||||
## Why reuse `getNewTargetSpeed` from utils.nim?
|
||||
|
||||
It already encodes the exact asymmetric acceleration/deceleration model
|
||||
(accel +1, decel −2 per tick). Reimplementing distance-based speed management
|
||||
from scratch would duplicate this and risk drift. Import it directly.
|
||||
|
||||
---
|
||||
|
||||
## Forward/Reverse optimality
|
||||
|
||||
The 90° threshold is the exact breakeven point:
|
||||
|
||||
- Turning 91° forward takes ≥10 ticks at slow speed + travel time.
|
||||
- Reversing 89° (i.e. 180−91=89° effective turn) takes fewer ticks total
|
||||
for any distance large enough to matter.
|
||||
- For very short distances (< ~36 units) the bot will decelerate before the
|
||||
turn completes anyway; the threshold still works because the speed penalty
|
||||
applies equally to both cases.
|
||||
|
||||
For a controller layer that feeds a neural network's goto target, sub-optimal
|
||||
behavior on very short hops is acceptable — the network will learn to avoid
|
||||
issuing tiny hops.
|
||||
|
||||
---
|
||||
|
||||
## Sources / References
|
||||
|
||||
- Tank Royale Nim API source: `tankroyale_botapi/utils.nim`, `bot.nim`,
|
||||
`constants.nim` (v1.0.1, installed at `~/.nimble/pkgs2/`).
|
||||
- Robocode wiki — "Proportional navigation" and "Should I go backward?"
|
||||
heuristic: widely documented in the Robocode community (e.g. RoboWiki
|
||||
`BasicSurfer`, `RaikoMX` source).
|
||||
- Tank Royale physics spec: confirmed against `ACCELERATION = 1.0`,
|
||||
`ABS_DECELERATION = 2.0` in `constants.nim`.
|
||||
Reference in New Issue
Block a user