j146 ledger: the field shape restores a real safe set offline (filter broken 63.5%->30.4%) and is a live null on damage/run and round wins; 375 battles, 5 arms
This commit is contained in:
@@ -3311,3 +3311,146 @@ Panel: the FROZEN 15-opponent `tools/ab/panel_movement.txt`. Harness:
|
||||
shipped `TR_MOVEMENT=strafe` default.
|
||||
|
||||
*(results appended below after the battles)*
|
||||
|
||||
### MEASURED — gate A: the offline mechanism ruler (cheap, first)
|
||||
|
||||
`common_libs/tests/measure_tfil_arrival.nim`, replaying the recorded DrussGT
|
||||
fixture (20 026 ticks) with the j144 base on in every arm, so only the SHAPE
|
||||
moves. `filter broken` = the share of picks that had to promote a hot tile
|
||||
because FEWER THAN TWO tiles were safe; `mean safe candidates` = the size of the
|
||||
set the picker actually drew from.
|
||||
|
||||
| arm (corr / wall / rad / core / aura) | picks | filter broken | mean safe candidates | mean path heat | path heat >10 | mean \|turn\| | >90 deg |
|
||||
|---|---:|---:|---:|---:|---:|---:|---:|
|
||||
| `shipped` 20/30/10/10/5 | 529 | 63.5% | 15.12 | 26.10 | 62.9% | 73.6 | 35.7% |
|
||||
| **`middle` 10/15/5/20/10** | 457 | **30.4%** | **34.45** | **16.41** | 30.2% | 77.4 | 36.1% |
|
||||
| `corr10` 10/30/10/10/5 | 448 | 32.1% | 30.94 | 14.31 | 31.7% | 73.2 | 33.0% |
|
||||
| `bullets` 20/30/10/20/10 | 531 | 65.0% | 16.60 | 26.93 | 65.0% | 71.9 | 34.5% |
|
||||
| `nofield` 0/0/0/20/10 | 414 | 3.9% | 109.87 | 3.19 | 3.9% | 67.7 | 27.3% |
|
||||
|
||||
**The j145 diagnosis is confirmed and the middle shape fixes it — the way the task
|
||||
predicted.** Halving the corridor to 10 takes the filter-break rate **63.5% ->
|
||||
30.4%** and the safe set from **15.1 to 34.5 candidates**, and it does it by
|
||||
**removing** lava, not by trading safety: the mean heat of the path the bot was
|
||||
told to walk **falls 26.10 -> 16.41 (-37%)**. The "no safe set to tie-break in"
|
||||
problem is genuinely an artefact of corridor 20 being twice the threshold.
|
||||
Interestingly `corr10` alone does nearly as well as the whole middle shape
|
||||
offline (32.1% / 30.9 candidates), while `bullets` alone does **not** (65.0% /
|
||||
16.6) — raising the bullet's own heat while the corridor is still 20 just swaps
|
||||
one saturated source for another.
|
||||
|
||||
### MEASURED — gate B: the guard (`test_tfil_commit_env.nim`, 66 -> 77 checks)
|
||||
|
||||
All 77 pass. The j146 ones:
|
||||
|
||||
* **default-off-effect, byte-for-byte**: the shipped shape written out in full as
|
||||
env (`20/30/10/10/5`) and the knobs left UNSET give **the same move commands
|
||||
over all 20 026 ticks**; the pre-existing golden is untouched and still green.
|
||||
* **the mechanism claim, measured on a synthetic single bullet**: with the
|
||||
shipped core 10 the bullet's core tile is **10.0, not over** `PathDangerThreshold`
|
||||
(=10, and `safe` is `<=`) — i.e. a bullet is never dangerous on its own, exactly
|
||||
as j145 said. With `TR_TFIL_BULLET_CORE=20` the same bullet paints **20.0, over
|
||||
the threshold alone**, while a corridor at 10 paints **10.0 and never reaches
|
||||
it**.
|
||||
* **the shape really restores a safe set** in the full replay: filter breaks
|
||||
63.5% -> 30.4%, and the mean path heat FALLS 26.10 -> 16.41.
|
||||
|
||||
### MEASURED — gate C: the live A/B, 375 battles
|
||||
|
||||
> **Provenance.** Session `/tmp/ab/j146_shape`, frozen binary `298ea6d` (sha256
|
||||
> `0c6fd6c3…`), panel `tools/ab/panel_movement.txt` (15 opponents, FROZEN),
|
||||
> `TR_MOVEMENT=tfil` pinned on every arm with the j144 knobs ON in every arm
|
||||
> (`TR_TFIL_COMMIT_ARRIVAL=1 TR_TFIL_NOREV_SPEED=4`), arms file
|
||||
> `tools/ab/arms_tfil_shape.txt` registered above BEFORE any of these battles
|
||||
> ran. **5 arms x 15 opponents x 5 runs x 3 rounds = 375 battles, 0 failed, 0
|
||||
> never started, 1 excluded (BlitzBat/shape_shipped run5, owner attribution
|
||||
> failed), 988 s.** Reference `shape_shipped`. No budget cut: the panel and RUNS
|
||||
> are both full.
|
||||
> ```sh
|
||||
> TOURNAMENT_NIMCACHE=/tmp/nc_j146 tools/ab/tournament_run.sh \
|
||||
> --arms tools/ab/arms_tfil_shape.txt --panel tools/ab/panel_movement.txt \
|
||||
> --runs 5 --rounds 3 --conc 6 --wait-arena 45 --reference shape_shipped \
|
||||
> --outdir /tmp/ab/j146_shape
|
||||
> python3 tools/ab/tournament_analyze.py /tmp/ab/j146_shape --reference shape_shipped
|
||||
> ```
|
||||
|
||||
**Pooled dashboard (descriptive, NOT the verdict):**
|
||||
|
||||
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|
||||
|---|---:|---:|---:|---:|---:|---:|---:|---:|
|
||||
| `shape_shipped` | 74 | 122.2 | 176.6 | 1.50 | 111/222 | 50.0% | 15.02% | 397 |
|
||||
| `shape_middle` | 75 | 116.2 | 166.3 | 1.53 | 115/225 | 51.1% | 14.28% | 412 |
|
||||
| `shape_corr10` | 75 | 123.0 | 175.7 | 1.53 | 115/225 | 51.1% | 15.86% | 393 |
|
||||
| `shape_bullets` | 75 | 119.0 | 168.9 | 1.55 | 116/225 | 51.6% | 14.02% | 414 |
|
||||
| `shape_nofield` | 75 | 108.0 | 165.4 | 1.49 | 112/225 | 49.8% | 15.05% | 427 |
|
||||
|
||||
**Verdict layer (paired per opponent against `shape_shipped`):**
|
||||
|
||||
| arm | metric | mean Δ | SD | SE | 95% CI | sign test | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|
||||
|---|---|---:|---:|---:|---|---:|---:|---:|---:|---:|
|
||||
| `shape_middle` | damage | -5.18 | 17.71 | 4.57 | [-14.99, +4.63] | 6/15 | 0.6072 | 0.27 | 0.2681 | 12.81 |
|
||||
| `shape_middle` | wins | +0.02 | 0.38 | 0.10 | [-0.19, +0.23] | 8/13 | 0.5811 | 0.8945 | 0.726 | 0.28 |
|
||||
| `shape_middle` | hit_rate | -0.43 | 2.75 | 0.71 | [-1.96, +1.09] | 5/15 | 0.3018 | 0.5649 | 0.182 | 1.99 |
|
||||
| `shape_corr10` | damage | +1.57 | 18.61 | 4.80 | [-8.74, +11.87] | 8/15 | 1 | 0.7516 | 0.7548 | 13.46 |
|
||||
| `shape_corr10` | wins | +0.02 | 0.35 | 0.09 | [-0.18, +0.21] | 5/9 | 1 | 0.8906 | 0.7211 | 0.25 |
|
||||
| `shape_corr10` | hit_rate | +0.44 | 2.15 | 0.56 | [-0.76, +1.63] | 8/15 | 1 | 0.4528 | 0.712 | 1.56 |
|
||||
| `shape_bullets` | damage | -2.38 | 16.89 | 4.36 | [-11.73, +6.97] | 5/15 | 0.3018 | 0.6035 | 0.3487 | 12.22 |
|
||||
| `shape_bullets` | wins | +0.03 | 0.34 | 0.09 | [-0.16, +0.22] | 7/13 | 1 | 0.7686 | 0.506 | 0.25 |
|
||||
| `shape_bullets` | hit_rate | -0.62 | 1.82 | 0.47 | [-1.63, +0.39] | 3/15 | 0.03516 | 0.2069 | 0.1183 | 1.32 |
|
||||
| `shape_nofield` | damage | **-13.45** | 17.62 | 4.55 | **[-23.21, -3.69]** | 2/15 | **0.007385** | **0.00946** | **0.01149** | 12.75 |
|
||||
| `shape_nofield` | wins | -0.02 | 0.46 | 0.12 | [-0.28, +0.23] | 6/12 | 1 | 0.9126 | 0.9686 | 0.33 |
|
||||
| `shape_nofield` | hit_rate | -0.83 | 3.18 | 0.82 | [-2.59, +0.94] | 6/15 | 0.6072 | 0.3331 | 0.3487 | 2.30 |
|
||||
|
||||
**Offline `filter broken` per arm (the mechanism, from gate A):** shipped 63.5% ·
|
||||
middle 30.4% · corr10 32.1% · bullets 65.0% · nofield 3.9%.
|
||||
|
||||
### VERDICT — plain
|
||||
|
||||
1. **Does tfil want the same field shape strafe won on? NOT DEMONSTRATED.** The
|
||||
middle shape does not beat today's shape on either primary metric:
|
||||
**+0.02 wins/run** (sign 8/13, p = 0.5811, sign-flip p = 0.8945) and
|
||||
**-5.18 dmg/run** (sign 6/15, p = 0.6072). Round wins 115/225 (51.1%) vs
|
||||
111/222 (50.0%). Nothing reaches the pre-registered bar, so **nothing is
|
||||
changed**: the shipped shape stays 20/30/10/10/5 and the new knobs stay
|
||||
default-off-effect.
|
||||
2. **The prediction that was WRONG, recorded as wrong.** The batch predicted
|
||||
`shape_nofield` would be the weakest arm and that it would separate from the
|
||||
middle. The first half held — `nofield` is the only arm that separates at all
|
||||
(**-13.45 dmg/run, 95% CI [-23.2, -3.7], sign-flip p = 0.0095**), exactly
|
||||
replicating strafe's `field_off` (j119 batch 4, -0.47 wins/run p=0.0063) and
|
||||
confirming the false-winner control works. But the middle was NOT
|
||||
distinguishable from the shipped shape, and `corr10` and `bullets` were not
|
||||
either, so the shape decomposition cannot be resolved live: the middle is not
|
||||
better than shipped, and nothing in the family is.
|
||||
3. **The mechanism moved, the outcome did not — and that is the real finding.**
|
||||
Offline the middle shape is a large, unambiguous win of the mechanism j145
|
||||
said was missing: filter breaks **63.5% -> 30.4%**, safe candidates
|
||||
**15.1 -> 34.5**, mean path heat **26.10 -> 16.41**. Live the *incoming* hit
|
||||
rate moves only 15.02% -> 14.28% (p = 0.56) — and compare j144/j145, where the
|
||||
*same* metric moved 17.13% -> 15.05% with sign-flip p = 0.006. So on tfil a
|
||||
restored safe set is **not** worth measurable incoming hits, unlike a restored
|
||||
commitment. Three jobs in a row now show the same thing: tfil's field
|
||||
(corridor/wall/bullet heat) is a second-order knob behind the commitment, and
|
||||
the tie-break/field layers are exactly where the live outcome stops responding.
|
||||
4. **No convergence recommendation.** Because the middle shape did NOT win, the
|
||||
honest answer is the opposite of "converge": **tfil and strafe must keep their
|
||||
own heat constants for now.** Merging them on the strength of a strafe-only
|
||||
result is exactly the cross-mover extrapolation this campaign has refused
|
||||
twice. The shared observation that IS worth writing down is the one this
|
||||
batch *measured in both movers*: the field is a **huge** mechanism
|
||||
(`filter broken` 63.5% -> 30.4% offline for one constant) and a **nil**
|
||||
outcome, and the "no safe set to tie-break in" pathology is real and is caused
|
||||
by `corridor 20 > PathDangerThreshold 10`.
|
||||
5. **What it would take to resolve it.** The MDE at 5 runs/opponent is
|
||||
**0.28 wins/run** and **12.8 dmg/run**; the middle's observed effect
|
||||
(+0.02 wins, -5.2 dmg) is an order of magnitude inside that, so this is an
|
||||
**under-powered null, not evidence of no effect**. MDE scales as 1/sqrt(runs),
|
||||
so detecting a 0.10 wins/run effect would take **~40 runs per opponent
|
||||
(~2900 battles, ~2.2 h at the observed 988 s / 375)** and a 5 dmg/run effect
|
||||
~33 runs (~2400 battles). That is a decision for the owner, not a default:
|
||||
the cheap offline evidence is already unambiguous about the mechanism and
|
||||
flat about everything the bot actually scores on.
|
||||
|
||||
**The shipped default is untouched.** `TR_MOVEMENT=strafe` remains the default;
|
||||
`TR_TFIL_BULLET_CORE` / `TR_TFIL_BULLET_AURA` default to today's `10.0` / `5.0`,
|
||||
so `TR_MOVEMENT=tfil` still means today's tfil, byte-for-byte (guard check 1).
|
||||
|
||||
Reference in New Issue
Block a user