Files
SirRoboGarage/common_libs/tests/measure_parallel_vbullet_cost.nim
T
SirStone 78975a35c4 cost: parallel per-enemy virtual bullets are NOT affordable as proposed
Benchmark driving the real rack and the real VirtualTracker over 7 recorded
DrussGT fixtures synthesised into an N-enemy melee. Answers "the virtual bullets
are cheap, why not keep fitness for every enemy in parallel?" (the user's idea,
motivated by making kill-stealing target switches free).

BASELINE: exactly 4.0 predict calls per gun per tick (one per power bin) - 52/tick
for the shipped 13-gun rack (TMSelect is compiled out). The task's 56/tick was
the 14-gun figure.

VERDICT: NOT AFFORDABLE. Budget is 13.16 ms/tick (76 ticks/s measured live).
  N=1  46% of budget
  N=2  94%          <- already at the edge
  N=4  189%
  N=6  274%
Marginal cost ~= 5.9 ms per extra target, linear.

TWO FINDINGS THE PROPOSAL MISSED:

1. `onResult` TRAINING dominates, not predict. Tsetlin's onResult alone is
   3.40 ms/tick - ~99.5% of all 13-gun onResult cost - doing ~174k rand() calls
   per resolved bullet. Every spawned bullet that resolves triggers it, so it
   scales 1:1 with targets. The per-target cost is the Tsetlin training pass.

2. `MaxBullets=8192` is a HARD BLOCKER, not just CPU. Spawn rate is 56*N/tick and
   path-metric bullets live until they hit a wall (40-90 ticks). Measured dropped
   bullets/tick: N=1 -> 0, N=2 -> ~3, N=4 -> ~180, N=6 -> ~300. At N=6 the ring
   wraps every ~24 ticks, so most bullets are silently clobbered and never scored.
   A working N=6 pipeline needs MaxBullets ~30k-50k (~4-6 MB, cheap RAM).

ALSO MEASURED: Tsetlin and KNN do NOT cache per tick - they redo the full TM
forward pass / full KNN scan for EACH of the 4 power bins (Tsetlin 1.94 ms/tick
of predict, KNN 0.45). The earlier "tick-only cache" fix never touched the two
most expensive predicts. Pattern and TMSelect do cache fully.

MITIGATIONS (measured predict+spawn at N=6 vs 13.59 ms baseline):
  nearest-K=1 only        45% budget
  nearest-K=2             95%
  rotate every 3 ticks    95%
  drop Tsetlin for extras ~68% (INFERRED from Tsetlin's measured 90% share)
Tsetlin is ~90% of the per-target cost, so excluding it from non-primary targets
makes N=6 fit. "Resolve less often" is not a separate lever - resolution IS when
training happens.

ARCHITECTURAL CAVEAT (correctness, not cost - and not priced into the proposal):
the shared-rack topology is broken for this. The guns are global singletons, so
predicting for enemy B ADVANCES/OVERWRITES enemy A's velocity tracker, KNN
feature history and Tsetlin frame window in the SAME instance. Per-enemy fitness
with correct histories therefore requires PER-ENEMY GUN INSTANCES, which is what
this benchmark measured. That multiplies the (already dominant) Tsetlin cost.

CONSEQUENCE FOR THE PLAN: combined with the measured finding that the selector is
negative value and the rack should shrink to a few good guns, this work is much
less valuable than assumed - with a small rack (Pattern's predict is 40us and
fully cached) the cost falls proportionally. Priority lowered accordingly.

Caveat: the host was heavily loaded (load 15/16), so absolute ms carry ~30-50%
noise; min-of-2 and two independent runs agree on the trend, the Tsetlin
dominance, and the ring overflow. No melee fixture exists in the repo, so the
7 enemies are 7 distinct recorded trajectories (stated in the file header).
2026-09-22 01:26:12 +02:00

310 lines
13 KiB
Nim

## MEASUREMENT ONLY — offline cost benchmark for "keep per-enemy virtual bullets
## for ALL enemies in parallel" (kill-stealing feature).
##
## Drives the REAL 14-gun rack from `common_libs/guns/*` and the REAL
## `VirtualTracker`/`spawnBullets`/`tickBullets` from
## `common_libs/gun_harness/virtual_bullets.nim`, over recorded DrussGT movement
## (`tools/fixtures/*.jsonl`, READ-ONLY).
##
## NO shipped source is edited; NO bot is rebuilt; offline only.
##
## Synthesised melee: there is no melee fixture in the repo, so N "enemies" are
## N DISTINCT recorded trajectories (different fixtures), all placed in one
## 800x600 arena with the shooter pose taken from fixture 0. Stated explicitly —
## it is not a real melee capture.
##
## Two rack topologies are measured:
## * SHARED — one 14-gun rack, predict() called once per (target, gun, bin).
## The literal minimal change against the current bot, whose guns
## are global. Per-tick caches (keyed on state.tick) are SHARED
## across targets.
## * PER-ENEMY — one 14-gun rack PER target (N racks), so each enemy gets a
## clean gun history. Gun ids are flattened to target*14+gun.
##
## RING CAP: VirtualTracker.MaxBullets = 8192 (shipped const). At N targets the
## spawn rate is 56*N/tick and a path-metric bullet lives until it leaves the
## arena (tens of ticks), so N>=2 overflows the ring and silently drops bullets.
## The `dropped/tick` column measures that. The full-pipeline cost for a WORKING
## ring is therefore reported as a PROJECTION: measured predict+spawn(N) plus
## N x the measured N=1 onResult cost per tick (onResult is per resolved bullet,
## and a working pipeline resolves every spawned bullet).
##
## Usage:
## GUN_SELECTOR_MODE=absolute nim c -r common_libs/tests/measure_parallel_vbullet_cost.nim
import std/[os, strformat, times, tables, math, algorithm, monotimes, random]
import gun_harness/offline_range
import gun_harness/gun_interface
import gun_harness/virtual_bullets as vb
import range_guns
const
NGun = 14
NBins = len(vb.PowerBins)
TmGun = 13 ## TMSelect index (shipped DISABLED)
RepoRoot = currentSourcePath().parentDir.parentDir.parentDir
FixturesDir = RepoRoot / "tools" / "fixtures"
FixtureNames = [
"tr_drussgt_vs_corners.jsonl",
"tr_drussgt_vs_crazy.jsonl",
"tr_drussgt_vs_modularbot.jsonl",
"tr_drussgt_vs_modularbot_shield.jsonl",
"tr_drussgt_vs_spinbot.jsonl",
"drussgt_vs_ramfire.jsonl",
"drussgt_vs_drussgt.jsonl",
]
MaxTicks = 2000
# ── fixtures / state synthesis ────────────────────────────────────────────────
proc loadFixtures(): seq[Fixture] =
for n in FixtureNames:
let p = FixturesDir / n
if fileExists(p): result.add loadFixture(p)
proc buildStates(fx: seq[Fixture], maxTicks: int): seq[seq[WorldState]] =
## Per-target WorldState streams. Target k keeps its own recorded enemy pose
## but shares the shooter pose of fixture 0, and all targets share a single
## continuous tick counter `si` so the guns' per-tick caches align.
var ticks = maxTicks
for f in fx: ticks = min(ticks, f.states.len)
result = newSeq[seq[WorldState]](fx.len)
for k in 0..<fx.len:
result[k] = newSeq[WorldState](ticks)
for si in 0..<ticks:
var ws = fx[k].states[si]
let s0 = fx[0].states[si]
ws.selfX = s0.selfX
ws.selfY = s0.selfY
ws.selfHeading = s0.selfHeading
ws.selfSpeed = s0.selfSpeed
ws.selfRadarHeading = s0.selfRadarHeading
ws.selfEnergy = s0.selfEnergy
ws.tick = si
result[k][si] = ws
proc enemyTable(states: seq[seq[WorldState]], si: int):
Table[int, tuple[x, y: float, lastSeenTick: int, alive: bool]] =
for k in 0..<states.len:
let e = states[k][si]
result[k + 1] = (x: e.enemyX, y: e.enemyY, lastSeenTick: si, alive: true)
# ── full-pipeline scaling runner ──────────────────────────────────────────────
type
RunResult = object
ticks: int
predictSpawnNs: int64
resolveNs: int64
dropped: int
spawns: int
proc runPipeline(states: seq[seq[WorldState]], nTargets: int, perEnemy: bool,
warmup, measure: int, includeTm: bool, resolveEvery = 1,
kNearest = 0, rotateMod = 0): RunResult =
let n = nTargets
var racks: seq[seq[GunDriver]]
if perEnemy:
for k in 0..<n:
racks.add buildAllGunDrivers(seed = 1 + k, enableTmSelector = true)
else:
racks.add buildAllGunDrivers(seed = 1, enableTmSelector = true)
let numGuns = racks.len * NGun
var tracker = vb.initTracker(numGuns, bmPath)
let ticks = states[0].len
var dropStart = 0
var dropEnd = 0
for si in 0..<ticks:
var active: seq[int]
if kNearest > 0:
var ds: seq[tuple[d: float, k: int]]
for k in 0..<n:
let ws = states[k][si]
ds.add (hypot(ws.enemyX - ws.selfX, ws.enemyY - ws.selfY), k)
ds.sort(proc(a, b: tuple[d: float, k: int]): int = cmp(a.d, b.d))
for i in 0..<min(kNearest, n): active.add ds[i].k
else:
for k in 0..<n: active.add k
let t0 = getMonoTime()
for k in active:
if rotateMod > 0 and (si mod rotateMod) != (k mod rotateMod): continue
let ws = states[k][si]
let rack = racks[if perEnemy: k else: 0]
for gi in 0..<NGun:
if not includeTm and gi == TmGun: continue
var preds: array[NBins, GunPrediction]
for b in 0..<NBins:
preds[b] = rack[gi].predictCb(ws, bulletSpeed(vb.PowerBins[b]))
tracker.spawnBullets(if perEnemy: k * NGun + gi else: gi, preds, ws, k + 1)
if si >= warmup and si < warmup + measure:
inc result.spawns, NBins
let t1 = getMonoTime()
if (si mod resolveEvery) == 0:
let et = enemyTable(states, si)
tracker.tickBullets(states[0][si], et,
proc(gunId: GunId, binIdx: int, e: FeedbackEvent) =
if includeTm or gunId mod NGun != TmGun:
racks[gunId div NGun][gunId mod NGun].resultCb(e))
let t2 = getMonoTime()
if si == warmup:
dropStart = tracker.droppedBullets
if si >= warmup and si < warmup + measure:
inc result.ticks
result.predictSpawnNs += (t1 - t0).inNanoseconds
result.resolveNs += (t2 - t1).inNanoseconds
dropEnd = tracker.droppedBullets
# dropped during the MEASUREMENT window only
result.dropped = dropEnd - dropStart
# ── per-gun breakdown (N = 1, timed per call) ─────────────────────────────────
type
PerGun = object
calls: array[NGun, array[NBins, int]]
predNs: array[NGun, array[NBins, int64]]
onCalls: array[NGun, int]
onNs: array[NGun, int64]
measure: int
proc runPerGun(states: seq[seq[WorldState]], warmup, measure: int): ref PerGun =
let rack = buildAllGunDrivers(seed = 1, enableTmSelector = true)
var tracker = vb.initTracker(NGun, bmPath)
let r = new(PerGun)
r.measure = measure
let ticks = states[0].len
for si in 0..<ticks:
let ws = states[0][si]
for gi in 0..<NGun:
var preds: array[NBins, GunPrediction]
for b in 0..<NBins:
let t0 = getMonoTime()
preds[b] = rack[gi].predictCb(ws, bulletSpeed(vb.PowerBins[b]))
let dt = (getMonoTime() - t0).inNanoseconds
if si >= warmup and si < warmup + measure:
r.predNs[gi][b] += dt
inc r.calls[gi][b]
tracker.spawnBullets(gi, preds, ws, 1)
let et = enemyTable(states, si)
tracker.tickBullets(ws, et,
proc(gunId: GunId, binIdx: int, e: FeedbackEvent) =
let t0 = getMonoTime()
rack[gunId].resultCb(e)
let dt = (getMonoTime() - t0).inNanoseconds
if si >= warmup and si < warmup + measure:
r.onNs[gunId] += dt
inc r.onCalls[gunId])
result = r
const GunNames = ["HeadOn", "Linear", "Tsetlin", "Circular", "GuessFactor",
"Pattern", "WallBounce", "Accel", "StopShot", "Displace", "AvgLead",
"DecayGF", "KNN", "TMSelect"]
proc printPerGun(pg: ref PerGun) =
echo ""
echo "=================== BASELINE PER-GUN (N=1, 14-gun rack) ==================="
echo "(times include one getMonoTime pair ~50ns/call; cheap guns are overhead-bound)"
echo "gun calls/tick bin0 ns/call bin1-3 ns/call predict ns/tick onResult ns/call onRes/tick"
for gi in 0..<NGun:
let c0 = pg.calls[gi][0]
let c13 = pg.calls[gi][1] + pg.calls[gi][2] + pg.calls[gi][3]
let b0 = if c0 > 0: pg.predNs[gi][0].float / c0.float else: 0.0
let b13 = if c13 > 0:
(pg.predNs[gi][1] + pg.predNs[gi][2] + pg.predNs[gi][3]).float /
c13.float else: 0.0
let pt = (pg.predNs[gi][0] + pg.predNs[gi][1] + pg.predNs[gi][2] +
pg.predNs[gi][3]).float / max(1, pg.measure).float
let onc = if pg.onCalls[gi] > 0: pg.onNs[gi].float / pg.onCalls[gi].float else: 0.0
let onpt = pg.onNs[gi].float / max(1, pg.measure).float
echo fmt"{GunNames[gi]:<12} {float(c0 + c13) / max(1, pg.measure).float:10.1f} " &
fmt"{b0:12.0f} {b13:14.0f} {pt:14.0f} {onc:16.0f} {onpt:9.0f}"
proc onNsPerTick(pg: ref PerGun, includeTm: bool): float =
## Total onResult cost per tick at N=1 for the selected rack.
for gi in 0..<NGun:
if not includeTm and gi == TmGun: continue
result += pg.onNs[gi].float / max(1, pg.measure).float
# ── main ──────────────────────────────────────────────────────────────────────
proc main() =
randomize(1)
let fx = loadFixtures()
echo "fixtures loaded: ", fx.len
for i, f in fx: echo fmt" [{i}] {f.meta.adversary:<12} ticks={f.states.len}"
let states = buildStates(fx, MaxTicks)
echo "common ticks: ", states[0].len, " (synthesised ", fx.len, "-enemy melee)"
const warmup = 150
const measure = 300
let pg = runPerGun(states, warmup, measure)
printPerGun(pg)
let on14 = onNsPerTick(pg, true)
let on13 = onNsPerTick(pg, false)
echo ""
echo fmt"onResult cost/tick at N=1: 14-gun={on14 / 1.0e6:.2f} ms " &
fmt"shipped 13-gun={on13 / 1.0e6:.2f} ms"
const BudgetMs = 1000.0 / 76.0 # ~76 ticks/s measured bridge throughput
echo fmt"per-tick budget (76 ticks/s) = {BudgetMs:.2f} ms/tick"
echo ""
echo "=================== N-SCALING ==================="
echo "(warmup=" & $warmup & ", measured=" & $measure & " ticks; budget " &
fmt"{BudgetMs:.2f} ms/tick)"
echo "rack topo N predict+spawn resolve(meas) dropped/tick PROJECTED full ticks/s %budget"
const Reps = 2
for includeTm in [false, true]:
let rk = if includeTm: "14-gun" else: "13-gun"
let ont = if includeTm: on14 else: on13
for perEnemy in [false, true]:
let topo = if perEnemy: "per-enemy" else: "shared "
for n in [1, 2, 4, 6]:
var psMs = 1.0e18
var rMs = 1.0e18
var dropPt = 0.0
for rep in 0..<Reps:
let r = runPipeline(states, n, perEnemy, warmup, measure, includeTm)
psMs = min(psMs, r.predictSpawnNs.float / 1.0e6 / max(1, r.ticks).float)
rMs = min(rMs, r.resolveNs.float / 1.0e6 / max(1, r.ticks).float)
dropPt = max(dropPt, r.dropped.float / max(1, r.ticks).float)
let projMs = psMs + float(n) * ont / 1.0e6
let tps = if projMs > 0: 1000.0 / projMs else: 0.0
echo fmt"{rk:<10} {topo} {n:>2} {psMs:11.2f} {rMs:11.2f} {dropPt:11.1f} " &
fmt"{projMs:13.2f} {tps:7.0f} {projMs / BudgetMs * 100.0:6.0f}%"
# ── mitigations at N = 6 (shipped 13-gun shared rack) ───────────────────────
echo ""
echo "=================== MITIGATIONS at N=6 (shipped 13-gun, shared rack) ==================="
var basePs = 1.0e18
for rep in 0..<Reps:
let base = runPipeline(states, 6, false, warmup, measure, includeTm = false)
basePs = min(basePs, base.predictSpawnNs.float / 1.0e6 / max(1, base.ticks).float)
let baseProj = basePs + 6.0 * on13 / 1.0e6
echo fmt"baseline N=6: predict+spawn={basePs:.2f} ms/tick projected full={baseProj:.2f} ms/tick"
echo "variant predict+spawn saving% spawns/tick"
for k in [1, 2, 3]:
var ps = 1.0e18
var sp = 0
for rep in 0..<Reps:
let r = runPipeline(states, 6, false, warmup, measure, includeTm = false, kNearest = k)
ps = min(ps, r.predictSpawnNs.float / 1.0e6 / max(1, r.ticks).float)
sp = r.spawns div max(1, r.ticks)
echo fmt"nearest-K spawn (K={k}) {ps:11.2f} {(1.0 - ps / basePs) * 100.0:7.1f} " &
fmt"{sp:9d}"
for m in [2, 3]:
var ps = 1.0e18
var sp = 0
for rep in 0..<Reps:
let r = runPipeline(states, 6, false, warmup, measure, includeTm = false, rotateMod = m)
ps = min(ps, r.predictSpawnNs.float / 1.0e6 / max(1, r.ticks).float)
sp = r.spawns div max(1, r.ticks)
echo fmt"rotate targets (every {m} ticks) {ps:11.2f} {(1.0 - ps / basePs) * 100.0:7.1f} " &
fmt"{sp:9d}"
when isMainModule:
main()