Energy economy: the cliff becomes a SLOPE, plus a finishing cap. 11% less energy.

The user's request: "when our bot is low OR enemy is low, it is useless to use high
power instead low fast bullets have more chances to finish the enemy. Let's do a
math slope: starting from some health down, the power goes down with it."

1. ENERGY SLOPE (`TR_POWER_ENERGY_*`), replacing the old hard step at 50 energy:
   cap = ENERGY_MAX at/above ENERGY_HI, ENERGY_MIN at/below ENERGY_LO, LINEAR in
   power between, clamped. Defaults HI=80 LO=20 MIN=0.5 MAX=3.0, so no cap >=80,
   0.5 at <=20, and e.g. E=65 -> 2.375, E=50 -> 1.75, E=35 -> 1.125.
   Rationale: bullet speed is 20-3p, so lower power = FASTER bullet (less lead
   error, higher hit chance), fires more often (10+2p) and drains slower (p/shot).
   E[dE] = p(3P-1) => break-even hit probability is 1/3 INDEPENDENT of power, and
   our measured rates are 5-27%, far below it.

2. FINISHING CAP (`TR_POWER_FINISH_KILL`, default ON): cap power at the SMALLEST
   bullet that still removes the enemy's remaining energy -
   `E<=4 -> p=E/4` (min 0.1), `4<E<=16 -> p=(E+2)/6`, `E>16 -> no cap`.
   Rationale, and it makes the user's instinct stronger than a heuristic: server
   1.3.1 caps the damage SCORE at the energy ACTUALLY REMOVED, so overkill is
   WASTED damage AND ~6x the energy for ZERO extra score. Damage is 4p (p<=1) /
   6p-2 (p>1).

Both are min-composed with the existing far/below-average caps, may only LOWER
power (exhaustively tested), and are exempt while ramming.
`TR_POWER_POLICY=0` still returns the uncapped control exactly.

MEASURED ENERGY SAVING (offline replay of the DrussGT fixtures, 28,797 ticks):
  arm              shots  energy  meanP  E/1k ticks   vs cliff
  control(uncapped) 1913    4646   2.43    161.4      -90.2%
  cliff (today)     2363    2443   1.03     84.8       0.0%
  slope             2404    2178   0.91     75.6    ** 10.9% LESS **
  slope+finish      2404    2167   0.90     75.3    ** 11.3% LESS **
So the slope spends ~11% less energy than the cliff AND fires slightly MORE shots
(2404 vs 2363) - both directions at once.

HONEST NOTE on the finishing rule's reach here: ticks where the enemy is low
(0 < E <= 16) are only 2252/28797 = 7.8% of these fixtures, so finishing adds just
~11 energy of saving against DrussGT. It matters in CLOSER fights, not this one.

Verification: test_power_policy 58 (was 26) in BOTH the default and TR_POWER_POLICY=0
control arms - slope at E=100/80/65/50/35/20/5, powerToKill across E=0.1..100, the
inverse-cover property for E<=16, monotonicity, ram exemption, and an exhaustive
sweep proving power <= preference. Guards: test_gun_harness 39, test_vbullet_metric
11, test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24,
test_ram_decision 40, test_rack_membership 48, test_selector_tiebreak 19,
test_tm_pattern_registration 20, test_vbullet_admit_gate 12. acceptance
12/12 PASS. ModularBot compiles release.

Adds `common_libs/tests/measure_power_policy.nim` (the energy/histogram tool) and
updates docs/env_reference.md for the new `energySlope|finishKill` log reasons.

NOT MEASURED: the battle/hit-rate effect. The offline figures use the fixture
shooter's energy as a proxy, open-loop; the RELATIVE saving is the meaningful part.
This commit is contained in:
2026-09-23 00:12:04 +02:00
parent 81af5854df
commit b68707c867
7 changed files with 582 additions and 95 deletions
+210
View File
@@ -0,0 +1,210 @@
## Offline energy-economy measurement for the power policy (TR_POWER_*).
##
## Replays the committed DrussGT movement fixtures through the real VirtualTracker
## (same path the range/acceptance tests use) and, per tick, computes the power
## the SHIPPED rack's only gun (Pattern, id 5) would prefer (`bestPower`). It then
## simulates four policy arms against that SAME preference sequence:
##
## control : uncapped (TR_POWER_POLICY=0) — the baseline
## cliff : TODAY's shipped rule (far 1.0, self energy < 50 -> 1.0, belowAvg 2.0)
## slope : the new linear self-energy cap (finishing OFF)
## finish : slope + the new smallest-killing-bullet cap (finishing ON)
##
## The fire schedule uses the server's gun heat: a shot adds `1 + p/5` heat and
## the gun cools `0.1`/tick, so the interval is `10 + 2p` ticks — LOWER power fires
## more often. `energySpent` sums the fired power over shots; the histogram counts
## shots at each 0.1-wide power bucket.
##
## IMPORTANT (labelled in the output): the trajectory is recorded (open-loop,
## perfect-information) so this is NOT a closed-loop hit-rate A/B. What is
## MEASURED is the policy's energy draw over real recorded movement; what is
## INFERRED is the resulting battle outcome. Hit rates are NOT modelled here.
##
## Run: nim c -r common_libs/tests/measure_power_policy.nim
import std/[os, strformat, strutils, math, tables, algorithm]
import gun_harness/offline_range
import gun_harness/virtual_bullets
import range_guns
const
repoRoot = currentSourcePath().parentDir.parentDir.parentDir
fixturesDir = repoRoot / "tools" / "fixtures"
ShippedGunId = 5 ## Pattern — the only gun in the shipped rack
CoolRate = 0.1 ## gun heat lost per tick (server config)
OldCliffEnergy = 50.0 ## TR_POWER_LOW_ENERGY's old hard threshold
const FixtureNames = [
"drussgt_vs_spinbot.jsonl",
"drussgt_vs_ramfire.jsonl",
"drussgt_vs_crazy.jsonl",
"drussgt_vs_corners.jsonl",
"drussgt_vs_drussgt.jsonl",
]
type
Arm* = enum
aControl ## uncapped — the control arm
aCliff ## today's shipped cliff
aSlope ## new energy slope only
aFinish ## energy slope + finishing
TickRec = object
dist, selfE, enemyE: float
prefPower: float
pEst, pRef: float
SimResult = object
shots: int
energy: float
hist: Table[int, int] ## key = round(power*10)
proc armName(a: Arm): string =
case a
of aControl: "control"
of aCliff: "cliff "
of aSlope: "slope "
of aFinish: "finish "
proc firedPower(arm: Arm, r: TickRec): float =
## The power this arm would fire given the gun's own preference and the
## per-tick state. `applyPowerPolicy` already returns `min(preference, cap)`.
case arm
of aControl:
r.prefPower
of aCliff:
var cap = 3.0
if r.dist > PowerFarDist: cap = PowerFarCap
elif r.selfE < OldCliffEnergy: cap = PowerFarCap
elif r.pEst <= r.pRef: cap = PowerMidCap
min(r.prefPower, cap)
of aSlope:
applyPowerPolicy(r.prefPower, r.dist, r.selfE, r.pEst, r.pRef, false,
enemyEnergy = r.enemyE, enabled = true,
finishKill = false).power
of aFinish:
applyPowerPolicy(r.prefPower, r.dist, r.selfE, r.pEst, r.pRef, false,
enemyEnergy = r.enemyE, enabled = true,
finishKill = true).power
proc collect(fx: Fixture): seq[TickRec] =
## One pass through the real tracker, recording the per-tick preference and
## policy inputs. The preference sequence is arm-INDEPENDENT (virtual bullets
## are spawned for every bin regardless of what we fire), so all four arms are
## compared against the exact same sequence.
var recs: seq[TickRec]
var tickIdx = 0
let drivers = buildAllGunDrivers(seed = 1)
let cb = proc(t: ptr VirtualTracker) =
let si = tickIdx
inc tickIdx
if si >= fx.states.len: return
let st = fx.states[si]
let tid = fx.enemyId
let (prefBin, prefPower) = t[].bestPower(ShippedGunId, tid)
let fit = t[].fitnessFor(tid)
let pRef = gunRate(fit[ShippedGunId], pooled = true)
let pEst =
if fit[ShippedGunId].bins[prefBin].count == 0: pRef
else: fit[ShippedGunId].bins[prefBin].hitRate()
recs.add TickRec(
dist: hypot(st.enemyX - st.selfX, st.enemyY - st.selfY),
selfE: st.selfEnergy,
enemyE: st.enemyEnergy,
prefPower: prefPower,
pEst: pEst, pRef: pRef)
discard replayFixture(fx, drivers, metric = bmPath, tickCb = cb)
recs
proc simulate(recs: seq[TickRec], arm: Arm): SimResult =
## Fire whenever the gun is cool (heat <= 0), drawing `power` energy per shot
## and adding `1 + p/5` heat. Mirrors the live `setFire` + `getEnergy() > power`
## guard, so a shot is skipped if our energy cannot cover it.
var heat = 0.0
for r in recs:
heat = max(0.0, heat - CoolRate)
if heat > 1e-9: continue
let p = firedPower(arm, r)
if r.selfE <= p: continue
result.energy += p
inc result.shots
let key = int(round(p * 10.0))
result.hist[key] = result.hist.getOrDefault(key) + 1
heat = 1.0 + p / 5.0
proc addHist(dst: var Table[int, int], src: Table[int, int]) =
for k, v in src: dst[k] = dst.getOrDefault(k) + v
proc histLine(h: Table[int, int]): string =
var keys: seq[int]
for k in h.keys: keys.add k
keys.sort()
for k in keys:
if result.len > 0: result.add " "
result.add fmt"p={k.float/10.0:.1f}:{h[k]}"
proc main() =
echo "=== offline energy-economy measurement (power policy) ==="
echo "fixtures: ", FixtureNames.len, " gun: Pattern(id=", ShippedGunId, ")"
echo "heat model: +1+p/5 per shot, -0.1/tick => interval 10+2p ticks"
echo ""
var totalTicks = 0
var lowEnemyTicks = 0
var agg: array[Arm, SimResult]
echo "fixture ticks arm shots energy meanP"
echo "-".repeat(72)
for name in FixtureNames:
let path = fixturesDir / name
if not fileExists(path):
echo " (missing: ", path, ")"
continue
let fx = loadFixture(path)
let recs = collect(fx)
totalTicks += recs.len
for r in recs:
if r.enemyE > 0.0 and r.enemyE <= 16.0: inc lowEnemyTicks
for arm in Arm:
let s = simulate(recs, arm)
agg[arm].shots += s.shots
agg[arm].energy += s.energy
agg[arm].hist.addHist(s.hist)
let meanP = if s.shots > 0: s.energy / s.shots.float else: 0.0
echo fmt"{name:<26} {recs.len:>6} {armName(arm):<8} {s.shots:>6} " &
fmt"{s.energy:>8.0f} {meanP:>6.2f}"
echo ""
echo "=== AGGREGATE over all fixtures (", totalTicks, " ticks) ==="
echo "arm shots energy meanP E/1k ticks vs control vs cliff"
echo "-".repeat(72)
let base = agg[aControl].energy
let cliff = agg[aCliff].energy
for arm in Arm:
let s = agg[arm]
let meanP = if s.shots > 0: s.energy / s.shots.float else: 0.0
let per1k = if totalTicks > 0: s.energy / totalTicks.float * 1000.0 else: 0.0
let vsControl = if base > 0: (base - s.energy) / base * 100.0 else: 0.0
let vsCliff = if cliff > 0: (cliff - s.energy) / cliff * 100.0 else: 0.0
echo fmt"{armName(arm):<8} {s.shots:>7} {s.energy:>9.0f} {meanP:>7.2f} " &
fmt"{per1k:>11.1f} {vsControl:>10.1f}% {vsCliff:>9.1f}%"
echo ""
echo fmt"low-enemy ticks (0 < E <= 16, where the finishing rule can bind): " &
fmt"{lowEnemyTicks}/{totalTicks} ({lowEnemyTicks.float/max(1,totalTicks).float*100.0:.1f}%)"
echo ""
echo "=== POWER HISTOGRAM (shots per 0.1-wide power bucket, all fixtures) ==="
for arm in Arm:
echo armName(arm), ": ", histLine(agg[arm].hist)
echo ""
echo "=== HIT-CHANCE / BREAK-EVEN REASONING (INFERRED, not measured here) ==="
echo "E[dE] = p(3P-1): the break-even hit probability is 1/3 INDEPENDENT of power."
echo "Our measured real hit rates are 5-27% (far below 1/3), so every point of"
echo "power costs more energy than it returns. A smaller bullet needs MORE hits"
echo "(ceil(E/damage)) but each hit is MORE LIKELY (speed 20-3p => less lead"
echo "error) and shots come FASTER (interval 10+2p). This tool measures only the"
echo "ENERGY side; which effect wins for hit rate needs the battle A/B."
when isMainModule:
main()