e0666a562d
MEASURED against the real DrussGT, one frozen binary built from clean HEAD, rack knobs only (no source edits), 5 arms x 7 runs x 7 rounds, 8 concurrent battles, judged ONLY on server-side real hit rate from the events sidecar, exact two-sided permutation test on per-run rates. arm runs shots hits real % dmg/run p vs full full (shipped) 7 3898 270 6.93 159 -- onlyPattern 7 4582 494 10.78 287 0.0012 <- BETTER onlyKNN 7 4033 207 5.13 119 0.1340 onlyLinear 7 3215 105 3.27 65 0.0082 onlyGF 7 3193 72 2.25 45 0.0012 Firing Pattern ALONE gives +3.85pp pooled hit rate and +80% damage per run, and it fires MORE shots (4582 vs 3898) - it dominates on rate and volume. This is not "any single gun wins" (full beats Linear, GF and KNN); it is specifically "Pattern alone beats the rack". WHY - the virtual fitness signal mis-ranks guns against real outcomes: - HeadOn is massively over-selected: 31.4% of ticks, the most real shots (1070), but only 4.5% REAL. It alone drags the rack down. - Pattern has the best virtual rank and near-best real rate (11.9%, rank 2), yet is selected only 22.6% of the time. - Linear's apparent strength was SELECTION BIAS: conditional on being selected it looked like 15.2% (n=33), but its UNCONDITIONAL rate (onlyLinear) is 3.27%. Every earlier per-gun "real rate" in this repo is conditional on selection and is therefore confounded. This experiment is the clean measurement. NOT YET SETTLED (do not overclaim): - ONE ADVERSARY. All of this is vs DrussGT. Pattern must be re-checked against other bots before it becomes the default on this evidence alone. - Whether a SMALL rack of good guns beats Pattern alone. The selector is negative value on the CURRENT bloated rack; that does not prove it is negative value on a rack of only good guns. That is the next experiment and it decides whether the selection apparatus is fixed or disabled. - The user's standing directive is to KEEP virtual-fitness selection. This measurement conflicts with it, so the next step tests the selector on a small good rack rather than assuming either answer. Context - three prior selection-side attempts all failed: hysteresis (7.02% -> 5.10%, p=0.002), commitment (7.17% -> 4.44%, p=0.0012), arrival-accuracy tie-break (7.08%, p=0.88 null). The per-tick random draw is load-bearing on three independent measurements. This experiment locates the real problem one level up: which guns are in the rack, and that the virtual signal ranks them wrongly. Preserves the reusable harness (tools/ab/which_gun_run_one.sh, which_gun_arm_env.sh, which_gun_analyze.py) and the full writeup (docs/selector_negative_value.md).
28 lines
999 B
Bash
Executable File
28 lines
999 B
Bash
Executable File
#!/usr/bin/env bash
|
|
# Run ONE bridge battle for one arm/run against the frozen binary.
|
|
# run_one.sh <arm> <run> <rounds>
|
|
set -u
|
|
ARM="$1"; RUN="$2"; ROUNDS="${3:-7}"
|
|
ROOT=/home/davide/Projects/SirRoboGarage
|
|
OUT=/tmp/whichgun
|
|
BOTDIR="$OUT/drussgt_bots/${ARM}_${RUN}/DrussGT"
|
|
mkdir -p "$BOTDIR" "$OUT/botlog"
|
|
|
|
# arm env knobs (empty for `full`)
|
|
# shellcheck disable=SC2046
|
|
ARMV=$($OUT/arm_env.sh "$ARM")
|
|
|
|
rm -f "$OUT/gun_stats_${ARM}_r${RUN}.jsonl" "$OUT/events_${ARM}_r${RUN}.json" "$OUT/cap_${ARM}_r${RUN}.jsonl"
|
|
|
|
# shellcheck disable=SC2086
|
|
env $ARMV \
|
|
DRUSSGT_BOTDIR="$BOTDIR" \
|
|
DRUSSGT_DATA="$OUT/drussgt_data/${ARM}_${RUN}" \
|
|
GUN_STATS_PATH="$OUT/gun_stats_${ARM}_r${RUN}.jsonl" \
|
|
TR_EVENTS_OUT="$OUT/events_${ARM}_r${RUN}.json" \
|
|
WHICHGUN_LOG="$OUT/botlog/${ARM}_r${RUN}.log" \
|
|
timeout 600 "$ROOT/tools/robocode_shim/run_bridge_battle.sh" \
|
|
"$OUT/bots/ModularBot" "$ROUNDS" "$OUT/cap_${ARM}_r${RUN}.jsonl" \
|
|
> "$OUT/battle_${ARM}_r${RUN}.log" 2>&1
|
|
echo "$? ${ARM} r${RUN}"
|