Files
SirRoboGarage/docs/env_reference.md
T
SirStone b68707c867 Energy economy: the cliff becomes a SLOPE, plus a finishing cap. 11% less energy.
The user's request: "when our bot is low OR enemy is low, it is useless to use high
power instead low fast bullets have more chances to finish the enemy. Let's do a
math slope: starting from some health down, the power goes down with it."

1. ENERGY SLOPE (`TR_POWER_ENERGY_*`), replacing the old hard step at 50 energy:
   cap = ENERGY_MAX at/above ENERGY_HI, ENERGY_MIN at/below ENERGY_LO, LINEAR in
   power between, clamped. Defaults HI=80 LO=20 MIN=0.5 MAX=3.0, so no cap >=80,
   0.5 at <=20, and e.g. E=65 -> 2.375, E=50 -> 1.75, E=35 -> 1.125.
   Rationale: bullet speed is 20-3p, so lower power = FASTER bullet (less lead
   error, higher hit chance), fires more often (10+2p) and drains slower (p/shot).
   E[dE] = p(3P-1) => break-even hit probability is 1/3 INDEPENDENT of power, and
   our measured rates are 5-27%, far below it.

2. FINISHING CAP (`TR_POWER_FINISH_KILL`, default ON): cap power at the SMALLEST
   bullet that still removes the enemy's remaining energy -
   `E<=4 -> p=E/4` (min 0.1), `4<E<=16 -> p=(E+2)/6`, `E>16 -> no cap`.
   Rationale, and it makes the user's instinct stronger than a heuristic: server
   1.3.1 caps the damage SCORE at the energy ACTUALLY REMOVED, so overkill is
   WASTED damage AND ~6x the energy for ZERO extra score. Damage is 4p (p<=1) /
   6p-2 (p>1).

Both are min-composed with the existing far/below-average caps, may only LOWER
power (exhaustively tested), and are exempt while ramming.
`TR_POWER_POLICY=0` still returns the uncapped control exactly.

MEASURED ENERGY SAVING (offline replay of the DrussGT fixtures, 28,797 ticks):
  arm              shots  energy  meanP  E/1k ticks   vs cliff
  control(uncapped) 1913    4646   2.43    161.4      -90.2%
  cliff (today)     2363    2443   1.03     84.8       0.0%
  slope             2404    2178   0.91     75.6    ** 10.9% LESS **
  slope+finish      2404    2167   0.90     75.3    ** 11.3% LESS **
So the slope spends ~11% less energy than the cliff AND fires slightly MORE shots
(2404 vs 2363) - both directions at once.

HONEST NOTE on the finishing rule's reach here: ticks where the enemy is low
(0 < E <= 16) are only 2252/28797 = 7.8% of these fixtures, so finishing adds just
~11 energy of saving against DrussGT. It matters in CLOSER fights, not this one.

Verification: test_power_policy 58 (was 26) in BOTH the default and TR_POWER_POLICY=0
control arms - slope at E=100/80/65/50/35/20/5, powerToKill across E=0.1..100, the
inverse-cover property for E<=16, monotonicity, ram exemption, and an exhaustive
sweep proving power <= preference. Guards: test_gun_harness 39, test_vbullet_metric
11, test_power_selection 3, test_adaptive_radar 41, test_tfil_ring_weights 24,
test_ram_decision 40, test_rack_membership 48, test_selector_tiebreak 19,
test_tm_pattern_registration 20, test_vbullet_admit_gate 12. acceptance
12/12 PASS. ModularBot compiles release.

Adds `common_libs/tests/measure_power_policy.nim` (the energy/histogram tool) and
updates docs/env_reference.md for the new `energySlope|finishKill` log reasons.

NOT MEASURED: the battle/hit-rate effect. The offline figures use the fixture
shooter's energy as a proxy, open-loop; the RELATIVE saving is the meaningful part.
2026-09-23 00:12:04 +02:00

197 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Environment variable reference
Every knob the bot reads. **All are read ONCE at process start.**
> **The bot is spawned by the server (or by the GUI if it starts its own server).**
> So the variable must be set on **that** process, not in an unrelated terminal.
> Export it in the shell that launches the server/GUI.
An unset (or empty, or unparseable) value means **the shipped default** — a typo
can never silently change behaviour, it warns on stderr and falls back.
---
## The ones you'll actually use live
| variable | default | what it does |
|---|---|---|
| `TR_MOVEMENT` | `tfil` | movement engine: `tfil` (shipped) or `tfil_ring` (range-weighted variant) |
| `TR_RACK_<GUN>` | see below | move a gun in/out of the rack per mode |
| `TR_POWER_POLICY` | `1` | `0` = uncapped control arm (today's behaviour without the energy policy) |
| `TR_POWER_LOG` | off | `1` = log each power decision with its reason |
| `TR_RAM_LOG` | off | `1` = log ram on/off with the reason |
| `TR_MOVEMENT_LOG` | off | `1` = log movement band/class changes |
| `TR_TMHORIZON_LOG` | off | `1` = let the horizon TM gun log its thinking per shot |
---
## 1. Gun rack — which gun(s) may be chosen
| variable | default | what it does |
|---|---|---|
| `TR_RACK_<GUN>` | `PATTERN=both`, all others `off` | membership: `both` \| `1v1` \| `melee` \| `off` |
| `GUN_RACK_DISABLE` | empty | comma-separated gun **ids** to remove entirely (older mechanism; the gun never even spawns virtual bullets) |
| `GUN_STATS_PATH` | `/tmp/gun_stats.jsonl` | where the per-round per-gun stats are written |
`<GUN>` names: `HEADON LINEAR TSETLIN CIRCULAR GUESSFACTOR PATTERN WALLBOUNCE ACCEL
STOPSHOT DISPLACE AVGLEAD DECAYGF KNN TMSELECT TMPATTERN TMHORIZON`.
**The shipped default is Pattern only.** To restore the full old rack:
```sh
TR_RACK_PATTERN=both TR_RACK_HEADON=both TR_RACK_LINEAR=both TR_RACK_TSETLIN=both \
TR_RACK_CIRCULAR=both TR_RACK_GUESSFACTOR=both TR_RACK_WALLBOUNCE=both \
TR_RACK_ACCEL=both TR_RACK_STOPSHOT=both TR_RACK_DISPLACE=both TR_RACK_AVGLEAD=both \
TR_RACK_DECAYGF=both TR_RACK_KNN=both TR_RACK_TMSELECT=both ./out/ModularBot
```
To force one gun alone: `TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both`.
If **every** gun is `off`, the rack falls back to admitting all of them.
## 2. The selector (virtual-bullet fitness)
| variable | default | what it does |
|---|---|---|
| `GUN_VBULLET_METRIC` | `path` | how a virtual bullet is scored: `path` (swept ray — the better coarse signal, 7.4% vs 4.7% real) or `point` (arrival accuracy) |
| `GUN_SELECTOR_MODE` | `relative` | tie-band model: `relative` (scale-aware) or `absolute` (legacy fixed bars) |
| `GUN_SELECTOR_WINDOW` | `100` | rolling window (ticks) for the fitness rates |
| `GUN_SELECTOR_TIE` | `0.20` | tie band width: tied if `rate >= bestRate*(1-this)` |
| `GUN_SELECTOR_DWELL` | `10` | ticks the incumbent gun is held before it may be displaced |
| `GUN_SELECTOR_MARGIN` | `0.05` | a challenger must beat the incumbent by this fraction to displace it |
| `GUN_SELECTOR_TIEBREAK` | `off` | `off` (shipped) / `point` / `pointCommit`. Measured **negative** on real hit rate; opt-in only |
| `GUN_SELECTOR_POINT_TIE` | `0.5` | margin used by the `point` tie-break |
**Note:** the per-tick random draw inside the tie band is **load-bearing** —
replacing it with commitment to the virtual-best measurably *lowered* real hit
rate (7.02% → 5.10%, p=0.002). Three attempts to "smarten" the band all failed.
## 3. Power policy (energy economy)
| variable | default | what it does |
|---|---|---|
| `TR_POWER_POLICY` | `1` | `0` = uncapped control arm. **Measured: turning it off drops real hit rate 10.6%→7.9%** |
| `TR_POWER_FAR_DIST` | `200` px | beyond this = "bad chances zone" |
| `TR_POWER_FAR_CAP` | `1.0` | power cap beyond that distance |
| `TR_POWER_MID_CAP` | `2.0` | cap when close + healthy but chances are not above average |
| `TR_POWER_REF` | `0.0` | reference probability: `0` = the gun's own mean, `>0` = a fixed threshold |
| `TR_POWER_ENERGY_HI` | `80` | at/above this energy, no energy-slope cap |
| `TR_POWER_ENERGY_LO` | `20` | at/below this energy, cap = `ENERGY_MIN` |
| `TR_POWER_ENERGY_MIN` | `0.5` | the cap at/below `ENERGY_LO` |
| `TR_POWER_ENERGY_MAX` | `3.0` | the cap at/above `ENERGY_HI` (3.0 = effectively uncapped) |
| `TR_POWER_FINISH_KILL` | `1` | cap power at the **smallest bullet that can still kill** the enemy |
| `TR_POWER_LOG` | off | `1` = log each decision: `[power] p=… cap=… reason=far\|energySlope\|finishKill\|belowAvg\|full\|ram` |
WHY the slope: bullet speed is `20-3p`, so **lower power = faster bullet** (less
lead error, higher hit chance), fires more often (`10+2p` ticks) and drains energy
more slowly (`p`/shot). `E[dE] = p(3P-1)`, so break-even is **1/3 regardless of
power** — and our measured hit rates are 5–27%, far below it.
WHY the finishing cap: server 1.3.1 caps the damage **score** at the energy
actually removed, so **overkill scores nothing**. Damage is `4p` (p≤1) / `6p-2`
(p>1), so the smallest killing power is `E/4` for E≤4 and `(E+2)/6` for 4<E≤16.
## 4. Movement
| variable | default | what it does |
|---|---|---|
| `TR_MOVEMENT` | `tfil` | `tfil` (shipped) or `tfil_ring` (the range-weighted variant) |
| `TR_MOVEMENT_LOG` | off | `1` = log band/range-class changes for the ring mover |
| `TR_TFIL_RANGE_LO` / `_HI` | `100` / `200` px | the target-range band the ring mover prefers |
| `TR_TFIL_RANGE_TEMP` | `0.4` | peakiness of the range weighting. `0` = plain uniform draw = exact control |
| `TR_TFIL_RANGE_K` | `60` | softness of the falloff outside the band |
| `TR_TFIL_CORRIDOR_HEAT` | `10.0` | heat added along a bullet's corridor to the wall |
| `TR_TFIL_WALL_HOTNESS` | `15.0` | wall radiance |
**Ring mover is NOT the default and should not be shipped as-is:** it had the best
offline hit rate of anything measured, but it halves engagement range and **halves
survival** (round wins 16/49 → 6/49, p=0.012) — a glass cannon.
## 5. Ramming
| variable | default | what it does |
|---|---|---|
| `TR_RAM_OPPORTUNITY` | off | enable the proactive ram. **Measured: converts 0/6 times** — a straight-line pursuit cannot catch an equal-speed enemy |
| `TR_RAM_OPP_DIST` | `200` px | opportunity trigger distance |
| `TR_RAM_OPP_MARGIN` | `15` | how much MORE energy we must have than the enemy |
| `TR_RAM_ABORT_DMG` | `2.0` | abort an in-progress ram above this incoming **energy per turn** (the "bullet rain" abort) |
| `TR_RAM_PLAN` | off | the speculative "change of plan" trigger |
| `TR_RAM_PLAN_DIST` / `_MARGIN` / `_HITRATE` | `250` / `20` / `0.05` | its thresholds |
| `TR_RAM_LOG` | off | `1` = log `[ram] ON/OFF` with reason |
The **finisher** ram (enemy <20 energy, we are healthier) is always on and is the
only path that converts.
## 6. The horizon TM gun (`TMHORIZON`)
| variable | default | what it does |
|---|---|---|
| `TR_TMHORIZON_LOG` | off | `1` = per-shot thinking log + per-round summary |
| `TR_TMHORIZON_SHIFT` | `2.0` deg | how far it may move the aim off Pattern's answer. **`0` = predict but never move the aim** (isolates prediction from application) |
| `TR_TMHORIZON_BIG_MULT` | `1.5` | multiplier when the predicted correction is BIG |
| `TR_TMHORIZON_NSTATES` | `64` | **automata inertia** — the "mood". Lower = adapts faster, higher = more stubborn (2…4096) |
| `TR_TMHORIZON_RESET_ON_TARGET` | on | wipe the model when the target changes (no-op in 1v1) |
| `TR_TMHORIZON_WINDOW` | `0` | sliding window: retrain on only the most recent N samples. `0` = keep everything |
| `TR_TMHORIZON_RESET_DROP` | `0` | change detection: if rolling accuracy falls this many points below its own peak, treat it as "the enemy changed" and re-learn |
| `TR_TMHORIZON_ACCURVE` | off | `1` = log the rolling accuracy periodically |
| `TR_TMHORIZON_RETRAIN_EVERY` | `50` | samples between windowed retrains |
| `TR_TMHORIZON_EPOCHS` | `1` | epochs per windowed retrain |
Learning **persists across rounds** of the same battle (so it can overfit the
current enemy) and wipes only on a new battle or a target change. **No learning is
written to disk — nothing carries between battles.**
## 7. Other gun knobs
| variable | default | what it does |
|---|---|---|
| `TR_PATTERN_RAD_OFFSET` | `0.0` px | shift Pattern's aim along its own bearing (negative = aim short) |
| `TR_PATTERN_RAD_SCALE` | `1.0` | multiplier on Pattern's aim distance |
Both are **structurally incapable of changing the shot** — the live aim is
bearing-only, so a purely radial offset is invisible. Measured byte-identical on
`bmPath`. Kept for experiments; leave at defaults.
## 8. Measurement / instrumentation (all off by default, all zero-cost when off)
| variable | default | what it does |
|---|---|---|
| `TR_VBULLET_ADMIT_ONLY` | `1` | only admitted guns spawn virtual bullets. **Turning it on gave +68% tick rate** (87→146 ticks/s); `0` restores the old behaviour |
| `TR_RECORD_WORLDSTATE` | off | append every tick's WorldState to `/tmp/worldstate_record.jsonl` for offline replay |
| `TR_RADAR_FORCE_SPIN` | off | force the old stateless full-spin melee radar (for A/B on one binary) |
| `TR_RADAR_SCANLOG` | off | append per-round radar scan/coverage JSON |
| `TR_RADAR_SCAN_LOG_PATH` | `/tmp/radar_scan_log.jsonl` | where that goes |
| `TR_TRACKER_PROBE` | off | per-tick enemy tracker vs server enemy count |
| `TR_TRACKER_PROBE_PATH` | `/tmp/tracker_probe.jsonl` | where that goes |
## 9. Test harness only (not read by the bot itself)
| variable | default | what it does |
|---|---|---|
| `TR_SERVER_JAR` | the 1.3.1 jar in `~/Downloads/robocode-tankroyale/` | which server the tests launch. Set to the 0.35.5 jar to reproduce old numbers |
| `TR_BATTLE_RUNNER` | the 1.0.2 runner jar | which runner the tests launch |
| `TR_BATTLE_RUNNER_DIR` | derived | directory form of the runner |
## 10. Compile-time knobs (`-d:` flags, need a rebuild)
| flag | default | file |
|---|---|---|
| `-d:TMH_NSTATES=N` | 64 | `guns/tm_horizon.nim` — automata inertia (the runtime env overrides it) |
| `-d:TMH_NCLAUSES=N` | 40 | `guns/tm_horizon.nim` |
| `-d:TMH_MIN_OBS=N` | 24 | cold gate: below this many samples the TM emits no correction |
| `-d:TMH_STALE_MAX=N` | 8 | a sample is dropped if the enemy was not seen this recently |
| `-d:TM_S_DEF=x` | `"3.0"` | `guns/tm_pattern.nim` specificity (`s`). **Higher = LONGER clauses**, measured. `s=1.0` is degenerate |
| `-d:TM_NSTATES=N` | 64 | `guns/tm_pattern.nim` |
## Adding / finding knobs
```sh
# everything the code reads
grep -rn 'TR_[A-Z_]*"' --include=*.nim ModularBot_garage/src common_libs | sort -u
# what a specific knob defaults to
grep -rn 'TR_POWER_ENERGY_MIN' --include=*.nim .
```
Convention followed throughout: read once at module init, into a `let` with an
`EnvVar` name constant beside it, so one compiled binary can A/B every arm by
environment alone.