j146 ledger: the field shape restores a real safe set offline (filter broken 63.5%->30.4%) and is a live null on damage/run and round wins; 375 battles, 5 arms

This commit is contained in:
2026-09-26 22:35:17 +02:00
parent 298ea6d586
commit de5d02ba3f
3 changed files with 385 additions and 0 deletions
+143
View File
@@ -3311,3 +3311,146 @@ Panel: the FROZEN 15-opponent `tools/ab/panel_movement.txt`. Harness:
shipped `TR_MOVEMENT=strafe` default.
*(results appended below after the battles)*
### MEASURED — gate A: the offline mechanism ruler (cheap, first)
`common_libs/tests/measure_tfil_arrival.nim`, replaying the recorded DrussGT
fixture (20 026 ticks) with the j144 base on in every arm, so only the SHAPE
moves. `filter broken` = the share of picks that had to promote a hot tile
because FEWER THAN TWO tiles were safe; `mean safe candidates` = the size of the
set the picker actually drew from.
| arm (corr / wall / rad / core / aura) | picks | filter broken | mean safe candidates | mean path heat | path heat >10 | mean \|turn\| | >90 deg |
|---|---:|---:|---:|---:|---:|---:|---:|
| `shipped` 20/30/10/10/5 | 529 | 63.5% | 15.12 | 26.10 | 62.9% | 73.6 | 35.7% |
| **`middle` 10/15/5/20/10** | 457 | **30.4%** | **34.45** | **16.41** | 30.2% | 77.4 | 36.1% |
| `corr10` 10/30/10/10/5 | 448 | 32.1% | 30.94 | 14.31 | 31.7% | 73.2 | 33.0% |
| `bullets` 20/30/10/20/10 | 531 | 65.0% | 16.60 | 26.93 | 65.0% | 71.9 | 34.5% |
| `nofield` 0/0/0/20/10 | 414 | 3.9% | 109.87 | 3.19 | 3.9% | 67.7 | 27.3% |
**The j145 diagnosis is confirmed and the middle shape fixes it — the way the task
predicted.** Halving the corridor to 10 takes the filter-break rate **63.5% ->
30.4%** and the safe set from **15.1 to 34.5 candidates**, and it does it by
**removing** lava, not by trading safety: the mean heat of the path the bot was
told to walk **falls 26.10 -> 16.41 (-37%)**. The "no safe set to tie-break in"
problem is genuinely an artefact of corridor 20 being twice the threshold.
Interestingly `corr10` alone does nearly as well as the whole middle shape
offline (32.1% / 30.9 candidates), while `bullets` alone does **not** (65.0% /
16.6) — raising the bullet's own heat while the corridor is still 20 just swaps
one saturated source for another.
### MEASURED — gate B: the guard (`test_tfil_commit_env.nim`, 66 -> 77 checks)
All 77 pass. The j146 ones:
* **default-off-effect, byte-for-byte**: the shipped shape written out in full as
env (`20/30/10/10/5`) and the knobs left UNSET give **the same move commands
over all 20 026 ticks**; the pre-existing golden is untouched and still green.
* **the mechanism claim, measured on a synthetic single bullet**: with the
shipped core 10 the bullet's core tile is **10.0, not over** `PathDangerThreshold`
(=10, and `safe` is `<=`) — i.e. a bullet is never dangerous on its own, exactly
as j145 said. With `TR_TFIL_BULLET_CORE=20` the same bullet paints **20.0, over
the threshold alone**, while a corridor at 10 paints **10.0 and never reaches
it**.
* **the shape really restores a safe set** in the full replay: filter breaks
63.5% -> 30.4%, and the mean path heat FALLS 26.10 -> 16.41.
### MEASURED — gate C: the live A/B, 375 battles
> **Provenance.** Session `/tmp/ab/j146_shape`, frozen binary `298ea6d` (sha256
> `0c6fd6c3…`), panel `tools/ab/panel_movement.txt` (15 opponents, FROZEN),
> `TR_MOVEMENT=tfil` pinned on every arm with the j144 knobs ON in every arm
> (`TR_TFIL_COMMIT_ARRIVAL=1 TR_TFIL_NOREV_SPEED=4`), arms file
> `tools/ab/arms_tfil_shape.txt` registered above BEFORE any of these battles
> ran. **5 arms x 15 opponents x 5 runs x 3 rounds = 375 battles, 0 failed, 0
> never started, 1 excluded (BlitzBat/shape_shipped run5, owner attribution
> failed), 988 s.** Reference `shape_shipped`. No budget cut: the panel and RUNS
> are both full.
> ```sh
> TOURNAMENT_NIMCACHE=/tmp/nc_j146 tools/ab/tournament_run.sh \
> --arms tools/ab/arms_tfil_shape.txt --panel tools/ab/panel_movement.txt \
> --runs 5 --rounds 3 --conc 6 --wait-arena 45 --reference shape_shipped \
> --outdir /tmp/ab/j146_shape
> python3 tools/ab/tournament_analyze.py /tmp/ab/j146_shape --reference shape_shipped
> ```
**Pooled dashboard (descriptive, NOT the verdict):**
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `shape_shipped` | 74 | 122.2 | 176.6 | 1.50 | 111/222 | 50.0% | 15.02% | 397 |
| `shape_middle` | 75 | 116.2 | 166.3 | 1.53 | 115/225 | 51.1% | 14.28% | 412 |
| `shape_corr10` | 75 | 123.0 | 175.7 | 1.53 | 115/225 | 51.1% | 15.86% | 393 |
| `shape_bullets` | 75 | 119.0 | 168.9 | 1.55 | 116/225 | 51.6% | 14.02% | 414 |
| `shape_nofield` | 75 | 108.0 | 165.4 | 1.49 | 112/225 | 49.8% | 15.05% | 427 |
**Verdict layer (paired per opponent against `shape_shipped`):**
| arm | metric | mean Δ | SD | SE | 95% CI | sign test | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|---|---|---:|---:|---:|---|---:|---:|---:|---:|---:|
| `shape_middle` | damage | -5.18 | 17.71 | 4.57 | [-14.99, +4.63] | 6/15 | 0.6072 | 0.27 | 0.2681 | 12.81 |
| `shape_middle` | wins | +0.02 | 0.38 | 0.10 | [-0.19, +0.23] | 8/13 | 0.5811 | 0.8945 | 0.726 | 0.28 |
| `shape_middle` | hit_rate | -0.43 | 2.75 | 0.71 | [-1.96, +1.09] | 5/15 | 0.3018 | 0.5649 | 0.182 | 1.99 |
| `shape_corr10` | damage | +1.57 | 18.61 | 4.80 | [-8.74, +11.87] | 8/15 | 1 | 0.7516 | 0.7548 | 13.46 |
| `shape_corr10` | wins | +0.02 | 0.35 | 0.09 | [-0.18, +0.21] | 5/9 | 1 | 0.8906 | 0.7211 | 0.25 |
| `shape_corr10` | hit_rate | +0.44 | 2.15 | 0.56 | [-0.76, +1.63] | 8/15 | 1 | 0.4528 | 0.712 | 1.56 |
| `shape_bullets` | damage | -2.38 | 16.89 | 4.36 | [-11.73, +6.97] | 5/15 | 0.3018 | 0.6035 | 0.3487 | 12.22 |
| `shape_bullets` | wins | +0.03 | 0.34 | 0.09 | [-0.16, +0.22] | 7/13 | 1 | 0.7686 | 0.506 | 0.25 |
| `shape_bullets` | hit_rate | -0.62 | 1.82 | 0.47 | [-1.63, +0.39] | 3/15 | 0.03516 | 0.2069 | 0.1183 | 1.32 |
| `shape_nofield` | damage | **-13.45** | 17.62 | 4.55 | **[-23.21, -3.69]** | 2/15 | **0.007385** | **0.00946** | **0.01149** | 12.75 |
| `shape_nofield` | wins | -0.02 | 0.46 | 0.12 | [-0.28, +0.23] | 6/12 | 1 | 0.9126 | 0.9686 | 0.33 |
| `shape_nofield` | hit_rate | -0.83 | 3.18 | 0.82 | [-2.59, +0.94] | 6/15 | 0.6072 | 0.3331 | 0.3487 | 2.30 |
**Offline `filter broken` per arm (the mechanism, from gate A):** shipped 63.5% ·
middle 30.4% · corr10 32.1% · bullets 65.0% · nofield 3.9%.
### VERDICT — plain
1. **Does tfil want the same field shape strafe won on? NOT DEMONSTRATED.** The
middle shape does not beat today's shape on either primary metric:
**+0.02 wins/run** (sign 8/13, p = 0.5811, sign-flip p = 0.8945) and
**-5.18 dmg/run** (sign 6/15, p = 0.6072). Round wins 115/225 (51.1%) vs
111/222 (50.0%). Nothing reaches the pre-registered bar, so **nothing is
changed**: the shipped shape stays 20/30/10/10/5 and the new knobs stay
default-off-effect.
2. **The prediction that was WRONG, recorded as wrong.** The batch predicted
`shape_nofield` would be the weakest arm and that it would separate from the
middle. The first half held — `nofield` is the only arm that separates at all
(**-13.45 dmg/run, 95% CI [-23.2, -3.7], sign-flip p = 0.0095**), exactly
replicating strafe's `field_off` (j119 batch 4, -0.47 wins/run p=0.0063) and
confirming the false-winner control works. But the middle was NOT
distinguishable from the shipped shape, and `corr10` and `bullets` were not
either, so the shape decomposition cannot be resolved live: the middle is not
better than shipped, and nothing in the family is.
3. **The mechanism moved, the outcome did not — and that is the real finding.**
Offline the middle shape is a large, unambiguous win of the mechanism j145
said was missing: filter breaks **63.5% -> 30.4%**, safe candidates
**15.1 -> 34.5**, mean path heat **26.10 -> 16.41**. Live the *incoming* hit
rate moves only 15.02% -> 14.28% (p = 0.56) — and compare j144/j145, where the
*same* metric moved 17.13% -> 15.05% with sign-flip p = 0.006. So on tfil a
restored safe set is **not** worth measurable incoming hits, unlike a restored
commitment. Three jobs in a row now show the same thing: tfil's field
(corridor/wall/bullet heat) is a second-order knob behind the commitment, and
the tie-break/field layers are exactly where the live outcome stops responding.
4. **No convergence recommendation.** Because the middle shape did NOT win, the
honest answer is the opposite of "converge": **tfil and strafe must keep their
own heat constants for now.** Merging them on the strength of a strafe-only
result is exactly the cross-mover extrapolation this campaign has refused
twice. The shared observation that IS worth writing down is the one this
batch *measured in both movers*: the field is a **huge** mechanism
(`filter broken` 63.5% -> 30.4% offline for one constant) and a **nil**
outcome, and the "no safe set to tie-break in" pathology is real and is caused
by `corridor 20 > PathDangerThreshold 10`.
5. **What it would take to resolve it.** The MDE at 5 runs/opponent is
**0.28 wins/run** and **12.8 dmg/run**; the middle's observed effect
(+0.02 wins, -5.2 dmg) is an order of magnitude inside that, so this is an
**under-powered null, not evidence of no effect**. MDE scales as 1/sqrt(runs),
so detecting a 0.10 wins/run effect would take **~40 runs per opponent
(~2900 battles, ~2.2 h at the observed 988 s / 375)** and a 5 dmg/run effect
~33 runs (~2400 battles). That is a decision for the owner, not a default:
the cheap offline evidence is already unambiguous about the mechanism and
flat about everything the bot actually scores on.
**The shipped default is untouched.** `TR_MOVEMENT=strafe` remains the default;
`TR_TFIL_BULLET_CORE` / `TR_TFIL_BULLET_AURA` default to today's `10.0` / `5.0`,
so `TR_MOVEMENT=tfil` still means today's tfil, byte-for-byte (guard check 1).