Energy economy: the cliff becomes a SLOPE, plus a finishing cap. 11% less energy.
The user's request: "when our bot is low OR enemy is low, it is useless to use high power instead low fast bullets have more chances to finish the enemy. Let's do a math slope: starting from some health down, the power goes down with it." 1. ENERGY SLOPE (`TR_POWER_ENERGY_*`), replacing the old hard step at 50 energy: cap = ENERGY_MAX at/above ENERGY_HI, ENERGY_MIN at/below ENERGY_LO, LINEAR in power between, clamped. Defaults HI=80 LO=20 MIN=0.5 MAX=3.0, so no cap >=80, 0.5 at <=20, and e.g. E=65 -> 2.375, E=50 -> 1.75, E=35 -> 1.125. Rationale: bullet speed is 20-3p, so lower power = FASTER bullet (less lead error, higher hit chance), fires more often (10+2p) and drains slower (p/shot). E[dE] = p(3P-1) => break-even hit probability is 1/3 INDEPENDENT of power, and our measured rates are 5-27%, far below it. 2. FINISHING CAP (`TR_POWER_FINISH_KILL`, default ON): cap power at the SMALLEST bullet that still removes the enemy's remaining energy - `E<=4 -> p=E/4` (min 0.1), `4<E<=16 -> p=(E+2)/6`, `E>16 -> no cap`. Rationale, and it makes the user's instinct stronger than a heuristic: server 1.3.1 caps the damage SCORE at the energy ACTUALLY REMOVED, so overkill is WASTED damage AND ~6x the energy for ZERO extra score. Damage is 4p (p<=1) / 6p-2 (p>1). Both are min-composed with the existing far/below-average caps, may only LOWER power (exhaustively tested), and are exempt while ramming. `TR_POWER_POLICY=0` still returns the uncapped control exactly. MEASURED ENERGY SAVING (offline replay of the DrussGT fixtures, 28,797 ticks): arm shots energy meanP E/1k ticks vs cliff control(uncapped) 1913 4646 2.43 161.4 -90.2% cliff (today) 2363 2443 1.03 84.8 0.0% slope 2404 2178 0.91 75.6 ** 10.9% LESS ** slope+finish 2404 2167 0.90 75.3 ** 11.3% LESS ** So the slope spends ~11% less energy than the cliff AND fires slightly MORE shots (2404 vs 2363) - both directions at once. HONEST NOTE on the finishing rule's reach here: ticks where the enemy is low (0 < E <= 16) are only 2252/28797 = 7.8% of these fixtures, so finishing adds just ~11 energy of saving against DrussGT. It matters in CLOSER fights, not this one. Verification: test_power_policy 58 (was 26) in BOTH the default and TR_POWER_POLICY=0 control arms - slope at E=100/80/65/50/35/20/5, powerToKill across E=0.1..100, the inverse-cover property for E<=16, monotonicity, ram exemption, and an exhaustive sweep proving power <= preference. Guards: test_gun_harness 39, test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24, test_ram_decision 40, test_rack_membership 48, test_selector_tiebreak 19, test_tm_pattern_registration 20, test_vbullet_admit_gate 12. acceptance 12/12 PASS. ModularBot compiles release. Adds `common_libs/tests/measure_power_policy.nim` (the energy/histogram tool) and updates docs/env_reference.md for the new `energySlope|finishKill` log reasons. NOT MEASURED: the battle/hit-rate effect. The offline figures use the fixture shooter's energy as a proxy, open-loop; the RELATIVE saving is the meaningful part.
This commit is contained in:
@@ -0,0 +1,210 @@
|
||||
## Offline energy-economy measurement for the power policy (TR_POWER_*).
|
||||
##
|
||||
## Replays the committed DrussGT movement fixtures through the real VirtualTracker
|
||||
## (same path the range/acceptance tests use) and, per tick, computes the power
|
||||
## the SHIPPED rack's only gun (Pattern, id 5) would prefer (`bestPower`). It then
|
||||
## simulates four policy arms against that SAME preference sequence:
|
||||
##
|
||||
## control : uncapped (TR_POWER_POLICY=0) — the baseline
|
||||
## cliff : TODAY's shipped rule (far 1.0, self energy < 50 -> 1.0, belowAvg 2.0)
|
||||
## slope : the new linear self-energy cap (finishing OFF)
|
||||
## finish : slope + the new smallest-killing-bullet cap (finishing ON)
|
||||
##
|
||||
## The fire schedule uses the server's gun heat: a shot adds `1 + p/5` heat and
|
||||
## the gun cools `0.1`/tick, so the interval is `10 + 2p` ticks — LOWER power fires
|
||||
## more often. `energySpent` sums the fired power over shots; the histogram counts
|
||||
## shots at each 0.1-wide power bucket.
|
||||
##
|
||||
## IMPORTANT (labelled in the output): the trajectory is recorded (open-loop,
|
||||
## perfect-information) so this is NOT a closed-loop hit-rate A/B. What is
|
||||
## MEASURED is the policy's energy draw over real recorded movement; what is
|
||||
## INFERRED is the resulting battle outcome. Hit rates are NOT modelled here.
|
||||
##
|
||||
## Run: nim c -r common_libs/tests/measure_power_policy.nim
|
||||
|
||||
import std/[os, strformat, strutils, math, tables, algorithm]
|
||||
import gun_harness/offline_range
|
||||
import gun_harness/virtual_bullets
|
||||
import range_guns
|
||||
|
||||
const
|
||||
repoRoot = currentSourcePath().parentDir.parentDir.parentDir
|
||||
fixturesDir = repoRoot / "tools" / "fixtures"
|
||||
ShippedGunId = 5 ## Pattern — the only gun in the shipped rack
|
||||
CoolRate = 0.1 ## gun heat lost per tick (server config)
|
||||
OldCliffEnergy = 50.0 ## TR_POWER_LOW_ENERGY's old hard threshold
|
||||
|
||||
const FixtureNames = [
|
||||
"drussgt_vs_spinbot.jsonl",
|
||||
"drussgt_vs_ramfire.jsonl",
|
||||
"drussgt_vs_crazy.jsonl",
|
||||
"drussgt_vs_corners.jsonl",
|
||||
"drussgt_vs_drussgt.jsonl",
|
||||
]
|
||||
|
||||
type
|
||||
Arm* = enum
|
||||
aControl ## uncapped — the control arm
|
||||
aCliff ## today's shipped cliff
|
||||
aSlope ## new energy slope only
|
||||
aFinish ## energy slope + finishing
|
||||
|
||||
TickRec = object
|
||||
dist, selfE, enemyE: float
|
||||
prefPower: float
|
||||
pEst, pRef: float
|
||||
|
||||
SimResult = object
|
||||
shots: int
|
||||
energy: float
|
||||
hist: Table[int, int] ## key = round(power*10)
|
||||
|
||||
proc armName(a: Arm): string =
|
||||
case a
|
||||
of aControl: "control"
|
||||
of aCliff: "cliff "
|
||||
of aSlope: "slope "
|
||||
of aFinish: "finish "
|
||||
|
||||
proc firedPower(arm: Arm, r: TickRec): float =
|
||||
## The power this arm would fire given the gun's own preference and the
|
||||
## per-tick state. `applyPowerPolicy` already returns `min(preference, cap)`.
|
||||
case arm
|
||||
of aControl:
|
||||
r.prefPower
|
||||
of aCliff:
|
||||
var cap = 3.0
|
||||
if r.dist > PowerFarDist: cap = PowerFarCap
|
||||
elif r.selfE < OldCliffEnergy: cap = PowerFarCap
|
||||
elif r.pEst <= r.pRef: cap = PowerMidCap
|
||||
min(r.prefPower, cap)
|
||||
of aSlope:
|
||||
applyPowerPolicy(r.prefPower, r.dist, r.selfE, r.pEst, r.pRef, false,
|
||||
enemyEnergy = r.enemyE, enabled = true,
|
||||
finishKill = false).power
|
||||
of aFinish:
|
||||
applyPowerPolicy(r.prefPower, r.dist, r.selfE, r.pEst, r.pRef, false,
|
||||
enemyEnergy = r.enemyE, enabled = true,
|
||||
finishKill = true).power
|
||||
|
||||
proc collect(fx: Fixture): seq[TickRec] =
|
||||
## One pass through the real tracker, recording the per-tick preference and
|
||||
## policy inputs. The preference sequence is arm-INDEPENDENT (virtual bullets
|
||||
## are spawned for every bin regardless of what we fire), so all four arms are
|
||||
## compared against the exact same sequence.
|
||||
var recs: seq[TickRec]
|
||||
var tickIdx = 0
|
||||
let drivers = buildAllGunDrivers(seed = 1)
|
||||
let cb = proc(t: ptr VirtualTracker) =
|
||||
let si = tickIdx
|
||||
inc tickIdx
|
||||
if si >= fx.states.len: return
|
||||
let st = fx.states[si]
|
||||
let tid = fx.enemyId
|
||||
let (prefBin, prefPower) = t[].bestPower(ShippedGunId, tid)
|
||||
let fit = t[].fitnessFor(tid)
|
||||
let pRef = gunRate(fit[ShippedGunId], pooled = true)
|
||||
let pEst =
|
||||
if fit[ShippedGunId].bins[prefBin].count == 0: pRef
|
||||
else: fit[ShippedGunId].bins[prefBin].hitRate()
|
||||
recs.add TickRec(
|
||||
dist: hypot(st.enemyX - st.selfX, st.enemyY - st.selfY),
|
||||
selfE: st.selfEnergy,
|
||||
enemyE: st.enemyEnergy,
|
||||
prefPower: prefPower,
|
||||
pEst: pEst, pRef: pRef)
|
||||
discard replayFixture(fx, drivers, metric = bmPath, tickCb = cb)
|
||||
recs
|
||||
|
||||
proc simulate(recs: seq[TickRec], arm: Arm): SimResult =
|
||||
## Fire whenever the gun is cool (heat <= 0), drawing `power` energy per shot
|
||||
## and adding `1 + p/5` heat. Mirrors the live `setFire` + `getEnergy() > power`
|
||||
## guard, so a shot is skipped if our energy cannot cover it.
|
||||
var heat = 0.0
|
||||
for r in recs:
|
||||
heat = max(0.0, heat - CoolRate)
|
||||
if heat > 1e-9: continue
|
||||
let p = firedPower(arm, r)
|
||||
if r.selfE <= p: continue
|
||||
result.energy += p
|
||||
inc result.shots
|
||||
let key = int(round(p * 10.0))
|
||||
result.hist[key] = result.hist.getOrDefault(key) + 1
|
||||
heat = 1.0 + p / 5.0
|
||||
|
||||
proc addHist(dst: var Table[int, int], src: Table[int, int]) =
|
||||
for k, v in src: dst[k] = dst.getOrDefault(k) + v
|
||||
|
||||
proc histLine(h: Table[int, int]): string =
|
||||
var keys: seq[int]
|
||||
for k in h.keys: keys.add k
|
||||
keys.sort()
|
||||
for k in keys:
|
||||
if result.len > 0: result.add " "
|
||||
result.add fmt"p={k.float/10.0:.1f}:{h[k]}"
|
||||
|
||||
proc main() =
|
||||
echo "=== offline energy-economy measurement (power policy) ==="
|
||||
echo "fixtures: ", FixtureNames.len, " gun: Pattern(id=", ShippedGunId, ")"
|
||||
echo "heat model: +1+p/5 per shot, -0.1/tick => interval 10+2p ticks"
|
||||
echo ""
|
||||
|
||||
var totalTicks = 0
|
||||
var lowEnemyTicks = 0
|
||||
var agg: array[Arm, SimResult]
|
||||
|
||||
echo "fixture ticks arm shots energy meanP"
|
||||
echo "-".repeat(72)
|
||||
for name in FixtureNames:
|
||||
let path = fixturesDir / name
|
||||
if not fileExists(path):
|
||||
echo " (missing: ", path, ")"
|
||||
continue
|
||||
let fx = loadFixture(path)
|
||||
let recs = collect(fx)
|
||||
totalTicks += recs.len
|
||||
for r in recs:
|
||||
if r.enemyE > 0.0 and r.enemyE <= 16.0: inc lowEnemyTicks
|
||||
for arm in Arm:
|
||||
let s = simulate(recs, arm)
|
||||
agg[arm].shots += s.shots
|
||||
agg[arm].energy += s.energy
|
||||
agg[arm].hist.addHist(s.hist)
|
||||
let meanP = if s.shots > 0: s.energy / s.shots.float else: 0.0
|
||||
echo fmt"{name:<26} {recs.len:>6} {armName(arm):<8} {s.shots:>6} " &
|
||||
fmt"{s.energy:>8.0f} {meanP:>6.2f}"
|
||||
echo ""
|
||||
|
||||
echo "=== AGGREGATE over all fixtures (", totalTicks, " ticks) ==="
|
||||
echo "arm shots energy meanP E/1k ticks vs control vs cliff"
|
||||
echo "-".repeat(72)
|
||||
let base = agg[aControl].energy
|
||||
let cliff = agg[aCliff].energy
|
||||
for arm in Arm:
|
||||
let s = agg[arm]
|
||||
let meanP = if s.shots > 0: s.energy / s.shots.float else: 0.0
|
||||
let per1k = if totalTicks > 0: s.energy / totalTicks.float * 1000.0 else: 0.0
|
||||
let vsControl = if base > 0: (base - s.energy) / base * 100.0 else: 0.0
|
||||
let vsCliff = if cliff > 0: (cliff - s.energy) / cliff * 100.0 else: 0.0
|
||||
echo fmt"{armName(arm):<8} {s.shots:>7} {s.energy:>9.0f} {meanP:>7.2f} " &
|
||||
fmt"{per1k:>11.1f} {vsControl:>10.1f}% {vsCliff:>9.1f}%"
|
||||
|
||||
echo ""
|
||||
echo fmt"low-enemy ticks (0 < E <= 16, where the finishing rule can bind): " &
|
||||
fmt"{lowEnemyTicks}/{totalTicks} ({lowEnemyTicks.float/max(1,totalTicks).float*100.0:.1f}%)"
|
||||
echo ""
|
||||
echo "=== POWER HISTOGRAM (shots per 0.1-wide power bucket, all fixtures) ==="
|
||||
for arm in Arm:
|
||||
echo armName(arm), ": ", histLine(agg[arm].hist)
|
||||
|
||||
echo ""
|
||||
echo "=== HIT-CHANCE / BREAK-EVEN REASONING (INFERRED, not measured here) ==="
|
||||
echo "E[dE] = p(3P-1): the break-even hit probability is 1/3 INDEPENDENT of power."
|
||||
echo "Our measured real hit rates are 5-27% (far below 1/3), so every point of"
|
||||
echo "power costs more energy than it returns. A smaller bullet needs MORE hits"
|
||||
echo "(ceil(E/damage)) but each hit is MORE LIKELY (speed 20-3p => less lead"
|
||||
echo "error) and shots come FASTER (interval 10+2p). This tool measures only the"
|
||||
echo "ENERGY side; which effect wins for hit rate needs the battle A/B."
|
||||
|
||||
when isMainModule:
|
||||
main()
|
||||
Reference in New Issue
Block a user