cost: parallel per-enemy virtual bullets are NOT affordable as proposed
Benchmark driving the real rack and the real VirtualTracker over 7 recorded DrussGT fixtures synthesised into an N-enemy melee. Answers "the virtual bullets are cheap, why not keep fitness for every enemy in parallel?" (the user's idea, motivated by making kill-stealing target switches free). BASELINE: exactly 4.0 predict calls per gun per tick (one per power bin) - 52/tick for the shipped 13-gun rack (TMSelect is compiled out). The task's 56/tick was the 14-gun figure. VERDICT: NOT AFFORDABLE. Budget is 13.16 ms/tick (76 ticks/s measured live). N=1 46% of budget N=2 94% <- already at the edge N=4 189% N=6 274% Marginal cost ~= 5.9 ms per extra target, linear. TWO FINDINGS THE PROPOSAL MISSED: 1. `onResult` TRAINING dominates, not predict. Tsetlin's onResult alone is 3.40 ms/tick - ~99.5% of all 13-gun onResult cost - doing ~174k rand() calls per resolved bullet. Every spawned bullet that resolves triggers it, so it scales 1:1 with targets. The per-target cost is the Tsetlin training pass. 2. `MaxBullets=8192` is a HARD BLOCKER, not just CPU. Spawn rate is 56*N/tick and path-metric bullets live until they hit a wall (40-90 ticks). Measured dropped bullets/tick: N=1 -> 0, N=2 -> ~3, N=4 -> ~180, N=6 -> ~300. At N=6 the ring wraps every ~24 ticks, so most bullets are silently clobbered and never scored. A working N=6 pipeline needs MaxBullets ~30k-50k (~4-6 MB, cheap RAM). ALSO MEASURED: Tsetlin and KNN do NOT cache per tick - they redo the full TM forward pass / full KNN scan for EACH of the 4 power bins (Tsetlin 1.94 ms/tick of predict, KNN 0.45). The earlier "tick-only cache" fix never touched the two most expensive predicts. Pattern and TMSelect do cache fully. MITIGATIONS (measured predict+spawn at N=6 vs 13.59 ms baseline): nearest-K=1 only 45% budget nearest-K=2 95% rotate every 3 ticks 95% drop Tsetlin for extras ~68% (INFERRED from Tsetlin's measured 90% share) Tsetlin is ~90% of the per-target cost, so excluding it from non-primary targets makes N=6 fit. "Resolve less often" is not a separate lever - resolution IS when training happens. ARCHITECTURAL CAVEAT (correctness, not cost - and not priced into the proposal): the shared-rack topology is broken for this. The guns are global singletons, so predicting for enemy B ADVANCES/OVERWRITES enemy A's velocity tracker, KNN feature history and Tsetlin frame window in the SAME instance. Per-enemy fitness with correct histories therefore requires PER-ENEMY GUN INSTANCES, which is what this benchmark measured. That multiplies the (already dominant) Tsetlin cost. CONSEQUENCE FOR THE PLAN: combined with the measured finding that the selector is negative value and the rack should shrink to a few good guns, this work is much less valuable than assumed - with a small rack (Pattern's predict is 40us and fully cached) the cost falls proportionally. Priority lowered accordingly. Caveat: the host was heavily loaded (load 15/16), so absolute ms carry ~30-50% noise; min-of-2 and two independent runs agree on the trend, the Tsetlin dominance, and the ring overflow. No melee fixture exists in the repo, so the 7 enemies are 7 distinct recorded trajectories (stated in the file header).
This commit is contained in:
@@ -0,0 +1,309 @@
|
||||
## MEASUREMENT ONLY — offline cost benchmark for "keep per-enemy virtual bullets
|
||||
## for ALL enemies in parallel" (kill-stealing feature).
|
||||
##
|
||||
## Drives the REAL 14-gun rack from `common_libs/guns/*` and the REAL
|
||||
## `VirtualTracker`/`spawnBullets`/`tickBullets` from
|
||||
## `common_libs/gun_harness/virtual_bullets.nim`, over recorded DrussGT movement
|
||||
## (`tools/fixtures/*.jsonl`, READ-ONLY).
|
||||
##
|
||||
## NO shipped source is edited; NO bot is rebuilt; offline only.
|
||||
##
|
||||
## Synthesised melee: there is no melee fixture in the repo, so N "enemies" are
|
||||
## N DISTINCT recorded trajectories (different fixtures), all placed in one
|
||||
## 800x600 arena with the shooter pose taken from fixture 0. Stated explicitly —
|
||||
## it is not a real melee capture.
|
||||
##
|
||||
## Two rack topologies are measured:
|
||||
## * SHARED — one 14-gun rack, predict() called once per (target, gun, bin).
|
||||
## The literal minimal change against the current bot, whose guns
|
||||
## are global. Per-tick caches (keyed on state.tick) are SHARED
|
||||
## across targets.
|
||||
## * PER-ENEMY — one 14-gun rack PER target (N racks), so each enemy gets a
|
||||
## clean gun history. Gun ids are flattened to target*14+gun.
|
||||
##
|
||||
## RING CAP: VirtualTracker.MaxBullets = 8192 (shipped const). At N targets the
|
||||
## spawn rate is 56*N/tick and a path-metric bullet lives until it leaves the
|
||||
## arena (tens of ticks), so N>=2 overflows the ring and silently drops bullets.
|
||||
## The `dropped/tick` column measures that. The full-pipeline cost for a WORKING
|
||||
## ring is therefore reported as a PROJECTION: measured predict+spawn(N) plus
|
||||
## N x the measured N=1 onResult cost per tick (onResult is per resolved bullet,
|
||||
## and a working pipeline resolves every spawned bullet).
|
||||
##
|
||||
## Usage:
|
||||
## GUN_SELECTOR_MODE=absolute nim c -r common_libs/tests/measure_parallel_vbullet_cost.nim
|
||||
|
||||
import std/[os, strformat, times, tables, math, algorithm, monotimes, random]
|
||||
import gun_harness/offline_range
|
||||
import gun_harness/gun_interface
|
||||
import gun_harness/virtual_bullets as vb
|
||||
import range_guns
|
||||
|
||||
const
|
||||
NGun = 14
|
||||
NBins = len(vb.PowerBins)
|
||||
TmGun = 13 ## TMSelect index (shipped DISABLED)
|
||||
RepoRoot = currentSourcePath().parentDir.parentDir.parentDir
|
||||
FixturesDir = RepoRoot / "tools" / "fixtures"
|
||||
FixtureNames = [
|
||||
"tr_drussgt_vs_corners.jsonl",
|
||||
"tr_drussgt_vs_crazy.jsonl",
|
||||
"tr_drussgt_vs_modularbot.jsonl",
|
||||
"tr_drussgt_vs_modularbot_shield.jsonl",
|
||||
"tr_drussgt_vs_spinbot.jsonl",
|
||||
"drussgt_vs_ramfire.jsonl",
|
||||
"drussgt_vs_drussgt.jsonl",
|
||||
]
|
||||
MaxTicks = 2000
|
||||
|
||||
# ── fixtures / state synthesis ────────────────────────────────────────────────
|
||||
|
||||
proc loadFixtures(): seq[Fixture] =
|
||||
for n in FixtureNames:
|
||||
let p = FixturesDir / n
|
||||
if fileExists(p): result.add loadFixture(p)
|
||||
|
||||
proc buildStates(fx: seq[Fixture], maxTicks: int): seq[seq[WorldState]] =
|
||||
## Per-target WorldState streams. Target k keeps its own recorded enemy pose
|
||||
## but shares the shooter pose of fixture 0, and all targets share a single
|
||||
## continuous tick counter `si` so the guns' per-tick caches align.
|
||||
var ticks = maxTicks
|
||||
for f in fx: ticks = min(ticks, f.states.len)
|
||||
result = newSeq[seq[WorldState]](fx.len)
|
||||
for k in 0..<fx.len:
|
||||
result[k] = newSeq[WorldState](ticks)
|
||||
for si in 0..<ticks:
|
||||
var ws = fx[k].states[si]
|
||||
let s0 = fx[0].states[si]
|
||||
ws.selfX = s0.selfX
|
||||
ws.selfY = s0.selfY
|
||||
ws.selfHeading = s0.selfHeading
|
||||
ws.selfSpeed = s0.selfSpeed
|
||||
ws.selfRadarHeading = s0.selfRadarHeading
|
||||
ws.selfEnergy = s0.selfEnergy
|
||||
ws.tick = si
|
||||
result[k][si] = ws
|
||||
|
||||
proc enemyTable(states: seq[seq[WorldState]], si: int):
|
||||
Table[int, tuple[x, y: float, lastSeenTick: int, alive: bool]] =
|
||||
for k in 0..<states.len:
|
||||
let e = states[k][si]
|
||||
result[k + 1] = (x: e.enemyX, y: e.enemyY, lastSeenTick: si, alive: true)
|
||||
|
||||
# ── full-pipeline scaling runner ──────────────────────────────────────────────
|
||||
|
||||
type
|
||||
RunResult = object
|
||||
ticks: int
|
||||
predictSpawnNs: int64
|
||||
resolveNs: int64
|
||||
dropped: int
|
||||
spawns: int
|
||||
|
||||
proc runPipeline(states: seq[seq[WorldState]], nTargets: int, perEnemy: bool,
|
||||
warmup, measure: int, includeTm: bool, resolveEvery = 1,
|
||||
kNearest = 0, rotateMod = 0): RunResult =
|
||||
let n = nTargets
|
||||
var racks: seq[seq[GunDriver]]
|
||||
if perEnemy:
|
||||
for k in 0..<n:
|
||||
racks.add buildAllGunDrivers(seed = 1 + k, enableTmSelector = true)
|
||||
else:
|
||||
racks.add buildAllGunDrivers(seed = 1, enableTmSelector = true)
|
||||
let numGuns = racks.len * NGun
|
||||
var tracker = vb.initTracker(numGuns, bmPath)
|
||||
let ticks = states[0].len
|
||||
var dropStart = 0
|
||||
var dropEnd = 0
|
||||
|
||||
for si in 0..<ticks:
|
||||
var active: seq[int]
|
||||
if kNearest > 0:
|
||||
var ds: seq[tuple[d: float, k: int]]
|
||||
for k in 0..<n:
|
||||
let ws = states[k][si]
|
||||
ds.add (hypot(ws.enemyX - ws.selfX, ws.enemyY - ws.selfY), k)
|
||||
ds.sort(proc(a, b: tuple[d: float, k: int]): int = cmp(a.d, b.d))
|
||||
for i in 0..<min(kNearest, n): active.add ds[i].k
|
||||
else:
|
||||
for k in 0..<n: active.add k
|
||||
|
||||
let t0 = getMonoTime()
|
||||
for k in active:
|
||||
if rotateMod > 0 and (si mod rotateMod) != (k mod rotateMod): continue
|
||||
let ws = states[k][si]
|
||||
let rack = racks[if perEnemy: k else: 0]
|
||||
for gi in 0..<NGun:
|
||||
if not includeTm and gi == TmGun: continue
|
||||
var preds: array[NBins, GunPrediction]
|
||||
for b in 0..<NBins:
|
||||
preds[b] = rack[gi].predictCb(ws, bulletSpeed(vb.PowerBins[b]))
|
||||
tracker.spawnBullets(if perEnemy: k * NGun + gi else: gi, preds, ws, k + 1)
|
||||
if si >= warmup and si < warmup + measure:
|
||||
inc result.spawns, NBins
|
||||
let t1 = getMonoTime()
|
||||
if (si mod resolveEvery) == 0:
|
||||
let et = enemyTable(states, si)
|
||||
tracker.tickBullets(states[0][si], et,
|
||||
proc(gunId: GunId, binIdx: int, e: FeedbackEvent) =
|
||||
if includeTm or gunId mod NGun != TmGun:
|
||||
racks[gunId div NGun][gunId mod NGun].resultCb(e))
|
||||
let t2 = getMonoTime()
|
||||
if si == warmup:
|
||||
dropStart = tracker.droppedBullets
|
||||
if si >= warmup and si < warmup + measure:
|
||||
inc result.ticks
|
||||
result.predictSpawnNs += (t1 - t0).inNanoseconds
|
||||
result.resolveNs += (t2 - t1).inNanoseconds
|
||||
dropEnd = tracker.droppedBullets
|
||||
# dropped during the MEASUREMENT window only
|
||||
result.dropped = dropEnd - dropStart
|
||||
|
||||
# ── per-gun breakdown (N = 1, timed per call) ─────────────────────────────────
|
||||
|
||||
type
|
||||
PerGun = object
|
||||
calls: array[NGun, array[NBins, int]]
|
||||
predNs: array[NGun, array[NBins, int64]]
|
||||
onCalls: array[NGun, int]
|
||||
onNs: array[NGun, int64]
|
||||
measure: int
|
||||
|
||||
proc runPerGun(states: seq[seq[WorldState]], warmup, measure: int): ref PerGun =
|
||||
let rack = buildAllGunDrivers(seed = 1, enableTmSelector = true)
|
||||
var tracker = vb.initTracker(NGun, bmPath)
|
||||
let r = new(PerGun)
|
||||
r.measure = measure
|
||||
let ticks = states[0].len
|
||||
for si in 0..<ticks:
|
||||
let ws = states[0][si]
|
||||
for gi in 0..<NGun:
|
||||
var preds: array[NBins, GunPrediction]
|
||||
for b in 0..<NBins:
|
||||
let t0 = getMonoTime()
|
||||
preds[b] = rack[gi].predictCb(ws, bulletSpeed(vb.PowerBins[b]))
|
||||
let dt = (getMonoTime() - t0).inNanoseconds
|
||||
if si >= warmup and si < warmup + measure:
|
||||
r.predNs[gi][b] += dt
|
||||
inc r.calls[gi][b]
|
||||
tracker.spawnBullets(gi, preds, ws, 1)
|
||||
let et = enemyTable(states, si)
|
||||
tracker.tickBullets(ws, et,
|
||||
proc(gunId: GunId, binIdx: int, e: FeedbackEvent) =
|
||||
let t0 = getMonoTime()
|
||||
rack[gunId].resultCb(e)
|
||||
let dt = (getMonoTime() - t0).inNanoseconds
|
||||
if si >= warmup and si < warmup + measure:
|
||||
r.onNs[gunId] += dt
|
||||
inc r.onCalls[gunId])
|
||||
result = r
|
||||
|
||||
const GunNames = ["HeadOn", "Linear", "Tsetlin", "Circular", "GuessFactor",
|
||||
"Pattern", "WallBounce", "Accel", "StopShot", "Displace", "AvgLead",
|
||||
"DecayGF", "KNN", "TMSelect"]
|
||||
|
||||
proc printPerGun(pg: ref PerGun) =
|
||||
echo ""
|
||||
echo "=================== BASELINE PER-GUN (N=1, 14-gun rack) ==================="
|
||||
echo "(times include one getMonoTime pair ~50ns/call; cheap guns are overhead-bound)"
|
||||
echo "gun calls/tick bin0 ns/call bin1-3 ns/call predict ns/tick onResult ns/call onRes/tick"
|
||||
for gi in 0..<NGun:
|
||||
let c0 = pg.calls[gi][0]
|
||||
let c13 = pg.calls[gi][1] + pg.calls[gi][2] + pg.calls[gi][3]
|
||||
let b0 = if c0 > 0: pg.predNs[gi][0].float / c0.float else: 0.0
|
||||
let b13 = if c13 > 0:
|
||||
(pg.predNs[gi][1] + pg.predNs[gi][2] + pg.predNs[gi][3]).float /
|
||||
c13.float else: 0.0
|
||||
let pt = (pg.predNs[gi][0] + pg.predNs[gi][1] + pg.predNs[gi][2] +
|
||||
pg.predNs[gi][3]).float / max(1, pg.measure).float
|
||||
let onc = if pg.onCalls[gi] > 0: pg.onNs[gi].float / pg.onCalls[gi].float else: 0.0
|
||||
let onpt = pg.onNs[gi].float / max(1, pg.measure).float
|
||||
echo fmt"{GunNames[gi]:<12} {float(c0 + c13) / max(1, pg.measure).float:10.1f} " &
|
||||
fmt"{b0:12.0f} {b13:14.0f} {pt:14.0f} {onc:16.0f} {onpt:9.0f}"
|
||||
|
||||
proc onNsPerTick(pg: ref PerGun, includeTm: bool): float =
|
||||
## Total onResult cost per tick at N=1 for the selected rack.
|
||||
for gi in 0..<NGun:
|
||||
if not includeTm and gi == TmGun: continue
|
||||
result += pg.onNs[gi].float / max(1, pg.measure).float
|
||||
|
||||
# ── main ──────────────────────────────────────────────────────────────────────
|
||||
|
||||
proc main() =
|
||||
randomize(1)
|
||||
let fx = loadFixtures()
|
||||
echo "fixtures loaded: ", fx.len
|
||||
for i, f in fx: echo fmt" [{i}] {f.meta.adversary:<12} ticks={f.states.len}"
|
||||
let states = buildStates(fx, MaxTicks)
|
||||
echo "common ticks: ", states[0].len, " (synthesised ", fx.len, "-enemy melee)"
|
||||
|
||||
const warmup = 150
|
||||
const measure = 300
|
||||
|
||||
let pg = runPerGun(states, warmup, measure)
|
||||
printPerGun(pg)
|
||||
|
||||
let on14 = onNsPerTick(pg, true)
|
||||
let on13 = onNsPerTick(pg, false)
|
||||
echo ""
|
||||
echo fmt"onResult cost/tick at N=1: 14-gun={on14 / 1.0e6:.2f} ms " &
|
||||
fmt"shipped 13-gun={on13 / 1.0e6:.2f} ms"
|
||||
|
||||
const BudgetMs = 1000.0 / 76.0 # ~76 ticks/s measured bridge throughput
|
||||
echo fmt"per-tick budget (76 ticks/s) = {BudgetMs:.2f} ms/tick"
|
||||
|
||||
echo ""
|
||||
echo "=================== N-SCALING ==================="
|
||||
echo "(warmup=" & $warmup & ", measured=" & $measure & " ticks; budget " &
|
||||
fmt"{BudgetMs:.2f} ms/tick)"
|
||||
echo "rack topo N predict+spawn resolve(meas) dropped/tick PROJECTED full ticks/s %budget"
|
||||
const Reps = 2
|
||||
for includeTm in [false, true]:
|
||||
let rk = if includeTm: "14-gun" else: "13-gun"
|
||||
let ont = if includeTm: on14 else: on13
|
||||
for perEnemy in [false, true]:
|
||||
let topo = if perEnemy: "per-enemy" else: "shared "
|
||||
for n in [1, 2, 4, 6]:
|
||||
var psMs = 1.0e18
|
||||
var rMs = 1.0e18
|
||||
var dropPt = 0.0
|
||||
for rep in 0..<Reps:
|
||||
let r = runPipeline(states, n, perEnemy, warmup, measure, includeTm)
|
||||
psMs = min(psMs, r.predictSpawnNs.float / 1.0e6 / max(1, r.ticks).float)
|
||||
rMs = min(rMs, r.resolveNs.float / 1.0e6 / max(1, r.ticks).float)
|
||||
dropPt = max(dropPt, r.dropped.float / max(1, r.ticks).float)
|
||||
let projMs = psMs + float(n) * ont / 1.0e6
|
||||
let tps = if projMs > 0: 1000.0 / projMs else: 0.0
|
||||
echo fmt"{rk:<10} {topo} {n:>2} {psMs:11.2f} {rMs:11.2f} {dropPt:11.1f} " &
|
||||
fmt"{projMs:13.2f} {tps:7.0f} {projMs / BudgetMs * 100.0:6.0f}%"
|
||||
|
||||
# ── mitigations at N = 6 (shipped 13-gun shared rack) ───────────────────────
|
||||
echo ""
|
||||
echo "=================== MITIGATIONS at N=6 (shipped 13-gun, shared rack) ==================="
|
||||
var basePs = 1.0e18
|
||||
for rep in 0..<Reps:
|
||||
let base = runPipeline(states, 6, false, warmup, measure, includeTm = false)
|
||||
basePs = min(basePs, base.predictSpawnNs.float / 1.0e6 / max(1, base.ticks).float)
|
||||
let baseProj = basePs + 6.0 * on13 / 1.0e6
|
||||
echo fmt"baseline N=6: predict+spawn={basePs:.2f} ms/tick projected full={baseProj:.2f} ms/tick"
|
||||
echo "variant predict+spawn saving% spawns/tick"
|
||||
for k in [1, 2, 3]:
|
||||
var ps = 1.0e18
|
||||
var sp = 0
|
||||
for rep in 0..<Reps:
|
||||
let r = runPipeline(states, 6, false, warmup, measure, includeTm = false, kNearest = k)
|
||||
ps = min(ps, r.predictSpawnNs.float / 1.0e6 / max(1, r.ticks).float)
|
||||
sp = r.spawns div max(1, r.ticks)
|
||||
echo fmt"nearest-K spawn (K={k}) {ps:11.2f} {(1.0 - ps / basePs) * 100.0:7.1f} " &
|
||||
fmt"{sp:9d}"
|
||||
for m in [2, 3]:
|
||||
var ps = 1.0e18
|
||||
var sp = 0
|
||||
for rep in 0..<Reps:
|
||||
let r = runPipeline(states, 6, false, warmup, measure, includeTm = false, rotateMod = m)
|
||||
ps = min(ps, r.predictSpawnNs.float / 1.0e6 / max(1, r.ticks).float)
|
||||
sp = r.spawns div max(1, r.ticks)
|
||||
echo fmt"rotate targets (every {m} ticks) {ps:11.2f} {(1.0 - ps / basePs) * 100.0:7.1f} " &
|
||||
fmt"{sp:9d}"
|
||||
|
||||
when isMainModule:
|
||||
main()
|
||||
Reference in New Issue
Block a user