j145 TFIL: a turn-cost tiebreak among the SAFE tiles (default-off)

The picker scored candidates on pathMaxHeat alone and then drew uniformly
among the survivors, so a mirror-side tile was as likely as a straight-ahead
one. Added a continuous turn cost as a DRAW WEIGHT applied only after the
hard heat filter:

  w = max(1, round(1 + TR_TFIL_TURN_BIAS * (1 - max(0,|turn| - REF)/180)))

- TR_TFIL_TURN_BIAS (default 0) is the odds ratio straight-ahead vs 180 deg;
  TR_TFIL_TURN_REF_DEG (default 45) is where the penalty starts. Both
  default-off-effect: the default-path golden in test_tfil_commit_env.nim is
  unchanged and still passes.
- Turn cost is NEVER folded into the heat score. The filter stays hard.
- The draw stays random (j51 measured an argmin worse); every weight is
  floored at 1, so the pool can never be emptied and bias 0 is exactly the
  shipped uniform draw.
- |turn| now travels on the ScoredTile, and the commit log gained turn /
  minturn / promote so a caller can measure the regret of the draw.

Guards: 51 -> 66 checks (an absurd 99:1 bias never rescues an over-threshold
tile; mean |turn|, draw regret, >90 and mirror-side shares all fall; path
heat does not rise). env_report + .env.example updated.
This commit is contained in:
2026-09-26 21:56:54 +02:00
parent 0e7e124c7f
commit 39c90fd930
7 changed files with 595 additions and 20 deletions
+74
View File
@@ -2978,3 +2978,77 @@ He should expect the dodge to look *smoother and more deliberate* (fewer, longer
commitments) rather than twitchy, and he should see fewer bullets connect. He
should NOT expect a step change in his score from this alone: the measured
outcome effect is +0.28 wins/run with a p of 0.057 on the primary test.
---
# Batch 6 — TFIL turn-cost tiebreak (j145)
*Pre-registered BEFORE any battle of this batch was launched. No battle of this
batch existed when this section was written; the frozen binary for it is the
commit that adds the tiebreak.*
## The cause this batch fixes
The `tfil` picker's `ScoredTile` carried **one** term, `pathMaxHeat`. After the
hard filter (`pathMaxHeat <= PathDangerThreshold` = 10) the pick was a plain
`rand()` over the survivors, so a far-cooler tile on the OPPOSITE side was drawn
exactly as readily as a marginally-cooler one straight ahead. The only
heading-aware influence in the mover is `NoRevForwardWeight = 3` under
`TR_TFIL_NO_REV` — **binary** (it cannot tell 20 deg from 90, nor 91 from 179)
and **off by default**. This batch adds the continuous version.
## The treatment
Two knobs, both **off by default** (the shipped default path is byte-for-byte
identical — the golden in `test_tfil_commit_env.nim` still passes):
| knob | default | meaning |
|---|---|---|
| `TR_TFIL_TURN_BIAS` | `0.0` | the tiebreak's **odds ratio**: a straight-ahead safe tile is drawn `1 + bias` times as often as a 180 deg one |
| `TR_TFIL_TURN_REF_DEG` | `45.0` | the turn below which no penalty applies |
The draw weight of a safe candidate is
w = max(1, round(1 + bias * (1 - max(0, |turn| - refDeg) / 180)))
**The safety filter is untouched and stays hard.** Turn cost is never added to
the heat score (`heat + k*turnDeg` would trade dodging for smoothness, which is
backwards in a bullet-dodging game); the bias is applied *only* to the
weight of a draw *among tiles that already passed the filter*. Guard:
`test_tfil_commit_env.nim` runs an absurd bias (99:1) and asserts that **no**
over-threshold tile is ever chosen unless the mover's own "fewer than two tiles
are safe" fallback promoted it.
**Randomness is preserved.** Job j51 (`3142b70`) measured that randomness in
this tie is load-bearing for this bot — a deterministic argmin scored worse.
The pick is therefore a **weighted draw**, not an argmin; every weight is
floored at 1 so the pool can never be emptied, and at bias 0 every weight is 1,
i.e. exactly the shipped uniform draw.
## Arms (frozen, all `TR_MOVEMENT=tfil`)
1. `tfil_shipped` — stock defaults. **The reference.**
2. `arrive_norev` — the two knobs j144 recommends (`COMMIT_ARRIVAL=1`,
`NOREV_SPEED=4`, `MARGIN=0`).
3. `arrive_norev_turn` — arm 2 + `TURN_BIAS=9 TURN_REF_DEG=0`.
4. `turn_only` — `TURN_BIAS=9 TURN_REF_DEG=0` alone; isolates the turn fix.
Panel: the FROZEN 15-opponent `tools/ab/panel_movement.txt`. Harness:
`tools/ab/tournament_run.sh` + `tournament_analyze.py`.
## Pre-registered prediction and decision rule
* **Prediction.** Arm 3 > arm 2 > arm 1 on damage/run and round wins, because a
smaller commanded turn is a faster arrival and a shorter exposure. Arm 4 sits
between arm 1 and arm 3. The **incoming hit rate is the mechanism, not the
verdict** — the verdict is damage/run and round wins under the campaign's
pre-registered rule 2 (cross-opponent sign test p < 0.05 on one primary metric
with the other not down), with the SD/SE/95% CI/MDE reported alongside.
* **If nothing separates**, the verdict is *not distinguishable* and it is
**not shipped**. The pre-registered bar is not re-interpreted afterwards.
* **A null here does NOT undo j144.** j144's result is a *mechanism* result
(the incoming hit rate fell 18.07% → 14.92%, sign-flip p = 0.0013, in two
independent blocks) plus an under-powered outcome null. This batch can only
add to or fail to add to that; it cannot retract it.
*(results appended below after the battles)*