9bf3005850
The bot is spawned by the server/GUI, so it inherits the SERVER's environment. The user could not tell whether their exports reached the bot, so print a one-shot greppable report at boot: grep '^\[env\]' /tmp/modularbot_stdout.log Section A prints every TR_*/GUN_* this process actually received, the count vs the total env size, a loud warning when nothing matched, and the process identity (pid/ppid, cwd, self command line, and the PARENT command line) so the spawn trap is obvious. Section B prints the resolved effective value of every documented knob with its source (env|default), including clamps and the rack's empty-set fallback. Build identity (NimVersion, compile date/time, binary path/size/mtime) pins the exact artifact. Suppress with TR_ENV_REPORT=0. docs/env_reference.md: add the missing GUN_SHOTLOG_PATH, GUN_SELECTOR_MINOBS/FLOOR/POOL/RANK/SHRINK/SEED, TR_ENV_REPORT and -d:TM_NCLAUSES; record the measured TR_TMHORIZON_WINDOW verdict; and add a prominent 'Did my env vars actually reach the bot?' section with the boot report, the /proc/PID/environ no-code check, the correct GUI launch recipe, and how to prove the trap deliberately.
295 lines
15 KiB
Markdown
295 lines
15 KiB
Markdown
# Environment variable reference
|
||
|
||
Every knob the bot reads. **All are read ONCE at process start.**
|
||
|
||
> **The bot is spawned by the server (or by the GUI if it starts its own server).**
|
||
> So the variable must be set on **that** process, not in an unrelated terminal.
|
||
> Export it in the shell that launches the server/GUI.
|
||
|
||
An unset (or empty, or unparseable) value means **the shipped default** — a typo
|
||
can never silently change behaviour, it warns on stderr and falls back.
|
||
|
||
Every process prints a one-shot, greppable boot report to stdout; see
|
||
[Did my env vars actually reach the bot?](#did-my-env-vars-actually-reach-the-bot)
|
||
below. `TR_ENV_REPORT=0` suppresses it.
|
||
|
||
---
|
||
|
||
## Did my env vars actually reach the bot?
|
||
|
||
**The spawn trap:** the bot is spawned by the **server/GUI**, so it inherits the
|
||
**server's** environment — exporting a variable in an unrelated terminal does
|
||
nothing, because that terminal is not the bot's parent.
|
||
|
||
### 1. Ask the bot itself (boot-time report)
|
||
|
||
Every process prints a one-shot, greppable report to **stdout**
|
||
(`/tmp/modularbot_stdout.log` when launched by the GUI, or the GUI/server
|
||
console). Recover just the report with:
|
||
|
||
```sh
|
||
grep '^\[env\]' /tmp/modularbot_stdout.log
|
||
```
|
||
|
||
It has three parts, all behind `[env] === ENVIRONMENT ... ===`:
|
||
|
||
- **A. raw process environment** — every `TR_*`/`GUN_*` this process *actually*
|
||
received, sorted, followed by
|
||
`raw: N TR_*/GUN_* of M total environment variables`. It then prints the
|
||
process identity: `pid`/`ppid`, the cwd, the bot's own command line, and the
|
||
**parent's command line**. The parent command line is the conclusive proof of
|
||
which process spawned the bot (the server/GUI — not your shell). If nothing
|
||
matched it prints a loud
|
||
`[env] WARNING: this process has NO TR_*/GUN_* variables - the env you exported did NOT reach the bot.`
|
||
- **B. effective values** — for every knob, the value the bot will *actually*
|
||
use and whether it came from `env` or the shipped `default`, including the
|
||
effects of parsing, clamping and fallback (e.g. `TR_TMHORIZON_NSTATES` is
|
||
clamped to `2..4096`; `TR_RACK_*` shows the empty-set fallback). Where a value
|
||
is only resolvable later, it is labelled `(raw - not resolved here)` rather
|
||
than guessed.
|
||
- **build identity** — `NimVersion`, compile date/time, and the binary path +
|
||
size + mtime, so a stale binary is obvious.
|
||
|
||
Suppress the report with `TR_ENV_REPORT=0`.
|
||
|
||
### 2. The no-code check (authoritative)
|
||
|
||
`/proc/PID/environ` is the environment **at exec time** — exactly what the bot
|
||
started with, before it read anything:
|
||
|
||
```sh
|
||
for p in $(pgrep -f ModularBot); do
|
||
echo "== $p"
|
||
tr '\0' '\n' < /proc/$p/environ | grep -E '^(TR_|GUN_)'
|
||
done
|
||
```
|
||
|
||
### 3. Launching the GUI with knobs
|
||
|
||
Export in the **same shell** that launches the GUI/server, then start it from
|
||
that shell. Because the launcher **appends** to the log, clear it first so you
|
||
read only the new run:
|
||
|
||
```sh
|
||
rm -f /tmp/modularbot_stdout.log
|
||
export TR_POWER_ENERGY_MIN=1.0 TR_RACK_PATTERN=off
|
||
./start-gui.sh # whatever launches the server/GUI
|
||
grep '^\[env\]' /tmp/modularbot_stdout.log
|
||
```
|
||
|
||
### 4. Prove the trap deliberately
|
||
|
||
Terminal A — where you *exported*, but never launched anything:
|
||
|
||
```sh
|
||
export TR_POWER_ENERGY_MIN=1.0
|
||
```
|
||
|
||
Terminal B — where you launch the GUI:
|
||
|
||
```sh
|
||
./start-gui.sh
|
||
grep '^\[env\]' /tmp/modularbot_stdout.log | grep WARNING
|
||
# [env] raw: 0 TR_*/GUN_* of ... total environment variables
|
||
# [env] WARNING: this process has NO TR_*/GUN_* variables - the env you exported did NOT reach the bot.
|
||
```
|
||
|
||
Terminal A's variable never reached the bot because the bot's parent is the
|
||
server, not terminal A. Fix it by exporting in terminal B before launching.
|
||
|
||
---
|
||
|
||
## The ones you'll actually use live
|
||
|
||
| variable | default | what it does |
|
||
|---|---|---|
|
||
| `TR_MOVEMENT` | `tfil` | movement engine: `tfil` (shipped) or `tfil_ring` (range-weighted variant) |
|
||
| `TR_RACK_<GUN>` | see below | move a gun in/out of the rack per mode |
|
||
| `TR_POWER_POLICY` | `1` | `0` = uncapped control arm (today's behaviour without the energy policy) |
|
||
| `TR_POWER_LOG` | off | `1` = log each power decision with its reason |
|
||
| `TR_RAM_LOG` | off | `1` = log ram on/off with the reason |
|
||
| `TR_MOVEMENT_LOG` | off | `1` = log movement band/class changes |
|
||
| `TR_TMHORIZON_LOG` | off | `1` = let the horizon TM gun log its thinking per shot |
|
||
| `TR_ENV_REPORT` | `1` | print the boot-time `[env]` process-environment report to stdout; `0` suppresses it |
|
||
|
||
---
|
||
|
||
## 1. Gun rack — which gun(s) may be chosen
|
||
|
||
| variable | default | what it does |
|
||
|---|---|---|
|
||
| `TR_RACK_<GUN>` | `PATTERN=both`, all others `off` | membership: `both` \| `1v1` \| `melee` \| `off` |
|
||
| `GUN_RACK_DISABLE` | empty | comma-separated gun **ids** to remove entirely (older mechanism; the gun never even spawns virtual bullets) |
|
||
| `GUN_STATS_PATH` | `/tmp/gun_stats.jsonl` | where the per-round per-gun stats are written |
|
||
| `GUN_SHOTLOG_PATH` | `/tmp/shot_log.jsonl` | where the per-shot JSONL (Task A) is written |
|
||
|
||
`<GUN>` names: `HEADON LINEAR TSETLIN CIRCULAR GUESSFACTOR PATTERN WALLBOUNCE ACCEL
|
||
STOPSHOT DISPLACE AVGLEAD DECAYGF KNN TMSELECT TMPATTERN TMHORIZON`.
|
||
|
||
**The shipped default is Pattern only.** To restore the full old rack:
|
||
|
||
```sh
|
||
TR_RACK_PATTERN=both TR_RACK_HEADON=both TR_RACK_LINEAR=both TR_RACK_TSETLIN=both \
|
||
TR_RACK_CIRCULAR=both TR_RACK_GUESSFACTOR=both TR_RACK_WALLBOUNCE=both \
|
||
TR_RACK_ACCEL=both TR_RACK_STOPSHOT=both TR_RACK_DISPLACE=both TR_RACK_AVGLEAD=both \
|
||
TR_RACK_DECAYGF=both TR_RACK_KNN=both TR_RACK_TMSELECT=both ./out/ModularBot
|
||
```
|
||
|
||
To force one gun alone: `TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both`.
|
||
If **every** gun is `off`, the rack falls back to admitting all of them.
|
||
|
||
## 2. The selector (virtual-bullet fitness)
|
||
|
||
| variable | default | what it does |
|
||
|---|---|---|
|
||
| `GUN_VBULLET_METRIC` | `path` | how a virtual bullet is scored: `path` (swept ray — the better coarse signal, 7.4% vs 4.7% real) or `point` (arrival accuracy) |
|
||
| `GUN_SELECTOR_MODE` | `relative` | tie-band model: `relative` (scale-aware) or `absolute` (legacy fixed bars) |
|
||
| `GUN_SELECTOR_WINDOW` | `100` | rolling window (ticks) for the fitness rates |
|
||
| `GUN_SELECTOR_TIE` | `0.20` | tie band width: tied if `rate >= bestRate*(1-this)` |
|
||
| `GUN_SELECTOR_FLOOR` | `0.25` | relative floor: fire HeadOn only if best rate < this × its peak |
|
||
| `GUN_SELECTOR_MINOBS` | `50` | min observations before a gun×bin may compete (clamped to ≥1) |
|
||
| `GUN_SELECTOR_POOL` | `on` | pool the gun's power bins when computing its rate |
|
||
| `GUN_SELECTOR_RANK` | `mean` | window statistic: `mean` \| `wilson` \| `ucb` \| `thompson` \| `shrunk` |
|
||
| `GUN_SELECTOR_SHRINK` | `20.0` | empirical-Bayes shrink strength (`shrunk` rank only) |
|
||
| `GUN_SELECTOR_DWELL` | `10` | ticks the incumbent gun is held before it may be displaced |
|
||
| `GUN_SELECTOR_MARGIN` | `0.05` | a challenger must beat the incumbent by this fraction to displace it |
|
||
| `GUN_SELECTOR_SEED` | unset | integer seed for the tie-break RNG; unset = time+pid per process |
|
||
| `GUN_SELECTOR_TIEBREAK` | `off` | `off` (shipped) / `point` / `pointCommit`. Measured **negative** on real hit rate; opt-in only |
|
||
| `GUN_SELECTOR_POINT_TIE` | `0.5` | margin used by the `point` tie-break |
|
||
|
||
**Note:** the per-tick random draw inside the tie band is **load-bearing** —
|
||
replacing it with commitment to the virtual-best measurably *lowered* real hit
|
||
rate (7.02% → 5.10%, p=0.002). Three attempts to "smarten" the band all failed.
|
||
|
||
## 3. Power policy (energy economy)
|
||
|
||
| variable | default | what it does |
|
||
|---|---|---|
|
||
| `TR_POWER_POLICY` | `1` | `0` = uncapped control arm. **Measured: turning it off drops real hit rate 10.6%→7.9%** |
|
||
| `TR_POWER_FAR_DIST` | `200` px | beyond this = "bad chances zone" |
|
||
| `TR_POWER_FAR_CAP` | `1.0` | power cap beyond that distance |
|
||
| `TR_POWER_MID_CAP` | `2.0` | cap when close + healthy but chances are not above average |
|
||
| `TR_POWER_REF` | `0.0` | reference probability: `0` = the gun's own mean, `>0` = a fixed threshold |
|
||
| `TR_POWER_ENERGY_HI` | `80` | at/above this energy, no energy-slope cap |
|
||
| `TR_POWER_ENERGY_LO` | `20` | at/below this energy, cap = `ENERGY_MIN` |
|
||
| `TR_POWER_ENERGY_MIN` | `0.5` | the cap at/below `ENERGY_LO` |
|
||
| `TR_POWER_ENERGY_MAX` | `3.0` | the cap at/above `ENERGY_HI` (3.0 = effectively uncapped) |
|
||
| `TR_POWER_FINISH_KILL` | `1` | cap power at the **smallest bullet that can still kill** the enemy |
|
||
| `TR_POWER_LOG` | off | `1` = log each decision: `[power] p=… cap=… reason=far\|energySlope\|finishKill\|belowAvg\|full\|ram` |
|
||
|
||
WHY the slope: bullet speed is `20-3p`, so **lower power = faster bullet** (less
|
||
lead error, higher hit chance), fires more often (`10+2p` ticks) and drains energy
|
||
more slowly (`p`/shot). `E[dE] = p(3P-1)`, so break-even is **1/3 regardless of
|
||
power** — and our measured hit rates are 5–27%, far below it.
|
||
|
||
WHY the finishing cap: server 1.3.1 caps the damage **score** at the energy
|
||
actually removed, so **overkill scores nothing**. Damage is `4p` (p≤1) / `6p-2`
|
||
(p>1), so the smallest killing power is `E/4` for E≤4 and `(E+2)/6` for 4<E≤16.
|
||
|
||
## 4. Movement
|
||
|
||
| variable | default | what it does |
|
||
|---|---|---|
|
||
| `TR_MOVEMENT` | `tfil` | `tfil` (shipped) or `tfil_ring` (the range-weighted variant) |
|
||
| `TR_MOVEMENT_LOG` | off | `1` = log band/range-class changes for the ring mover |
|
||
| `TR_TFIL_RANGE_LO` / `_HI` | `100` / `200` px | the target-range band the ring mover prefers |
|
||
| `TR_TFIL_RANGE_TEMP` | `0.4` | peakiness of the range weighting. `0` = plain uniform draw = exact control |
|
||
| `TR_TFIL_RANGE_K` | `60` | softness of the falloff outside the band |
|
||
| `TR_TFIL_CORRIDOR_HEAT` | `10.0` | heat added along a bullet's corridor to the wall |
|
||
| `TR_TFIL_WALL_HOTNESS` | `15.0` | wall radiance |
|
||
|
||
**Ring mover is NOT the default and should not be shipped as-is:** it had the best
|
||
offline hit rate of anything measured, but it halves engagement range and **halves
|
||
survival** (round wins 16/49 → 6/49, p=0.012) — a glass cannon.
|
||
|
||
## 5. Ramming
|
||
|
||
| variable | default | what it does |
|
||
|---|---|---|
|
||
| `TR_RAM_OPPORTUNITY` | off | enable the proactive ram. **Measured: converts 0/6 times** — a straight-line pursuit cannot catch an equal-speed enemy |
|
||
| `TR_RAM_OPP_DIST` | `200` px | opportunity trigger distance |
|
||
| `TR_RAM_OPP_MARGIN` | `15` | how much MORE energy we must have than the enemy |
|
||
| `TR_RAM_ABORT_DMG` | `2.0` | abort an in-progress ram above this incoming **energy per turn** (the "bullet rain" abort) |
|
||
| `TR_RAM_PLAN` | off | the speculative "change of plan" trigger |
|
||
| `TR_RAM_PLAN_DIST` / `_MARGIN` / `_HITRATE` | `250` / `20` / `0.05` | its thresholds |
|
||
| `TR_RAM_LOG` | off | `1` = log `[ram] ON/OFF` with reason |
|
||
|
||
The **finisher** ram (enemy <20 energy, we are healthier) is always on and is the
|
||
only path that converts.
|
||
|
||
## 6. The horizon TM gun (`TMHORIZON`)
|
||
|
||
| variable | default | what it does |
|
||
|---|---|---|
|
||
| `TR_TMHORIZON_LOG` | off | `1` = per-shot thinking log + per-round summary |
|
||
| `TR_TMHORIZON_SHIFT` | `2.0` deg | how far it may move the aim off Pattern's answer. **`0` = predict but never move the aim** (isolates prediction from application) |
|
||
| `TR_TMHORIZON_BIG_MULT` | `1.5` | multiplier when the predicted correction is BIG |
|
||
| `TR_TMHORIZON_NSTATES` | `64` | **automata inertia** — the "mood". Lower = adapts faster, higher = more stubborn (2…4096) |
|
||
| `TR_TMHORIZON_RESET_ON_TARGET` | on | wipe the model when the target changes (no-op in 1v1) |
|
||
| `TR_TMHORIZON_WINDOW` | `0` | sliding window: retrain on only the most recent N samples. `0` = keep everything. **ON is MEASURED HARMFUL live: 26.5% vs 49.0% round wins, p=0.036 — keep it at `0`.** |
|
||
| `TR_TMHORIZON_RESET_DROP` | `0` | change detection: if rolling accuracy falls this many points below its own peak, treat it as "the enemy changed" and re-learn |
|
||
| `TR_TMHORIZON_ACCURVE` | off | `1` = log the rolling accuracy periodically |
|
||
| `TR_TMHORIZON_RETRAIN_EVERY` | `50` | samples between windowed retrains |
|
||
| `TR_TMHORIZON_EPOCHS` | `1` | epochs per windowed retrain |
|
||
|
||
Learning **persists across rounds** of the same battle (so it can overfit the
|
||
current enemy) and wipes only on a new battle or a target change. **No learning is
|
||
written to disk — nothing carries between battles.**
|
||
|
||
## 7. Other gun knobs
|
||
|
||
| variable | default | what it does |
|
||
|---|---|---|
|
||
| `TR_PATTERN_RAD_OFFSET` | `0.0` px | shift Pattern's aim along its own bearing (negative = aim short) |
|
||
| `TR_PATTERN_RAD_SCALE` | `1.0` | multiplier on Pattern's aim distance |
|
||
|
||
Both are **structurally incapable of changing the shot** — the live aim is
|
||
bearing-only, so a purely radial offset is invisible. Measured byte-identical on
|
||
`bmPath`. Kept for experiments; leave at defaults.
|
||
|
||
## 8. Measurement / instrumentation (all off by default, all zero-cost when off)
|
||
|
||
| variable | default | what it does |
|
||
|---|---|---|
|
||
| `TR_VBULLET_ADMIT_ONLY` | `1` | only admitted guns spawn virtual bullets. **Turning it on gave +68% tick rate** (87→146 ticks/s); `0` restores the old behaviour |
|
||
| `TR_ENV_REPORT` | `1` | print the boot-time `[env]` report to stdout; `0` suppresses it |
|
||
| `TR_RECORD_WORLDSTATE` | off | append every tick's WorldState to `/tmp/worldstate_record.jsonl` for offline replay |
|
||
| `TR_RADAR_FORCE_SPIN` | off | force the old stateless full-spin melee radar (for A/B on one binary) |
|
||
| `TR_RADAR_SCANLOG` | off | append per-round radar scan/coverage JSON |
|
||
| `TR_RADAR_SCAN_LOG_PATH` | `/tmp/radar_scan_log.jsonl` | where that goes |
|
||
| `TR_TRACKER_PROBE` | off | per-tick enemy tracker vs server enemy count |
|
||
| `TR_TRACKER_PROBE_PATH` | `/tmp/tracker_probe.jsonl` | where that goes |
|
||
|
||
## 9. Test harness only (not read by the bot itself)
|
||
|
||
| variable | default | what it does |
|
||
|---|---|---|
|
||
| `TR_SERVER_JAR` | the 1.3.1 jar in `~/Downloads/robocode-tankroyale/` | which server the tests launch. Set to the 0.35.5 jar to reproduce old numbers |
|
||
| `TR_BATTLE_RUNNER` | the 1.0.2 runner jar | which runner the tests launch |
|
||
| `TR_BATTLE_RUNNER_DIR` | derived | directory form of the runner |
|
||
|
||
## 10. Compile-time knobs (`-d:` flags, need a rebuild)
|
||
|
||
| flag | default | file |
|
||
|---|---|---|
|
||
| `-d:TMH_NSTATES=N` | 64 | `guns/tm_horizon.nim` — automata inertia (the runtime env overrides it) |
|
||
| `-d:TMH_NCLAUSES=N` | 40 | `guns/tm_horizon.nim` |
|
||
| `-d:TMH_MIN_OBS=N` | 24 | cold gate: below this many samples the TM emits no correction |
|
||
| `-d:TMH_STALE_MAX=N` | 8 | a sample is dropped if the enemy was not seen this recently |
|
||
| `-d:TM_S_DEF=x` | `"3.0"` | `guns/tm_pattern.nim` specificity (`s`). **Higher = LONGER clauses**, measured. `s=1.0` is degenerate |
|
||
| `-d:TM_NSTATES=N` | 64 | `guns/tm_pattern.nim` |
|
||
| `-d:TM_NCLAUSES=N` | 40 | `guns/tm_pattern.nim` — clauses per class |
|
||
|
||
## Adding / finding knobs
|
||
|
||
```sh
|
||
# everything the code reads
|
||
grep -rn 'TR_[A-Z_]*"' --include=*.nim ModularBot_garage/src common_libs | sort -u
|
||
# what a specific knob defaults to
|
||
grep -rn 'TR_POWER_ENERGY_MIN' --include=*.nim .
|
||
```
|
||
|
||
Convention followed throughout: read once at module init, into a `let` with an
|
||
`EnvVar` name constant beside it, so one compiled binary can A/B every arm by
|
||
environment alone.
|