Files
SirRoboGarage/tools/ab/which_gun_run_one.sh
T
SirStone e0666a562d The gun selector is NEGATIVE value: Pattern alone beats the full rack (p=0.0012)
MEASURED against the real DrussGT, one frozen binary built from clean HEAD, rack
knobs only (no source edits), 5 arms x 7 runs x 7 rounds, 8 concurrent battles,
judged ONLY on server-side real hit rate from the events sidecar, exact
two-sided permutation test on per-run rates.

  arm             runs  shots  hits  real %  dmg/run  p vs full
  full (shipped)     7   3898   270    6.93     159      --
  onlyPattern        7   4582   494   10.78     287      0.0012  <- BETTER
  onlyKNN            7   4033   207    5.13     119      0.1340
  onlyLinear         7   3215   105    3.27      65      0.0082
  onlyGF             7   3193    72    2.25      45      0.0012

Firing Pattern ALONE gives +3.85pp pooled hit rate and +80% damage per run, and
it fires MORE shots (4582 vs 3898) - it dominates on rate and volume. This is not
"any single gun wins" (full beats Linear, GF and KNN); it is specifically
"Pattern alone beats the rack".

WHY - the virtual fitness signal mis-ranks guns against real outcomes:
- HeadOn is massively over-selected: 31.4% of ticks, the most real shots (1070),
  but only 4.5% REAL. It alone drags the rack down.
- Pattern has the best virtual rank and near-best real rate (11.9%, rank 2), yet
  is selected only 22.6% of the time.
- Linear's apparent strength was SELECTION BIAS: conditional on being selected it
  looked like 15.2% (n=33), but its UNCONDITIONAL rate (onlyLinear) is 3.27%.
  Every earlier per-gun "real rate" in this repo is conditional on selection and
  is therefore confounded. This experiment is the clean measurement.

NOT YET SETTLED (do not overclaim):
- ONE ADVERSARY. All of this is vs DrussGT. Pattern must be re-checked against
  other bots before it becomes the default on this evidence alone.
- Whether a SMALL rack of good guns beats Pattern alone. The selector is negative
  value on the CURRENT bloated rack; that does not prove it is negative value on
  a rack of only good guns. That is the next experiment and it decides whether
  the selection apparatus is fixed or disabled.
- The user's standing directive is to KEEP virtual-fitness selection. This
  measurement conflicts with it, so the next step tests the selector on a small
  good rack rather than assuming either answer.

Context - three prior selection-side attempts all failed: hysteresis (7.02% ->
5.10%, p=0.002), commitment (7.17% -> 4.44%, p=0.0012), arrival-accuracy
tie-break (7.08%, p=0.88 null). The per-tick random draw is load-bearing on
three independent measurements. This experiment locates the real problem one
level up: which guns are in the rack, and that the virtual signal ranks them
wrongly.

Preserves the reusable harness (tools/ab/which_gun_run_one.sh,
which_gun_arm_env.sh, which_gun_analyze.py) and the full writeup
(docs/selector_negative_value.md).
2026-09-22 01:21:14 +02:00

28 lines
999 B
Bash
Executable File

#!/usr/bin/env bash
# Run ONE bridge battle for one arm/run against the frozen binary.
# run_one.sh <arm> <run> <rounds>
set -u
ARM="$1"; RUN="$2"; ROUNDS="${3:-7}"
ROOT=/home/davide/Projects/SirRoboGarage
OUT=/tmp/whichgun
BOTDIR="$OUT/drussgt_bots/${ARM}_${RUN}/DrussGT"
mkdir -p "$BOTDIR" "$OUT/botlog"
# arm env knobs (empty for `full`)
# shellcheck disable=SC2046
ARMV=$($OUT/arm_env.sh "$ARM")
rm -f "$OUT/gun_stats_${ARM}_r${RUN}.jsonl" "$OUT/events_${ARM}_r${RUN}.json" "$OUT/cap_${ARM}_r${RUN}.jsonl"
# shellcheck disable=SC2086
env $ARMV \
DRUSSGT_BOTDIR="$BOTDIR" \
DRUSSGT_DATA="$OUT/drussgt_data/${ARM}_${RUN}" \
GUN_STATS_PATH="$OUT/gun_stats_${ARM}_r${RUN}.jsonl" \
TR_EVENTS_OUT="$OUT/events_${ARM}_r${RUN}.json" \
WHICHGUN_LOG="$OUT/botlog/${ARM}_r${RUN}.log" \
timeout 600 "$ROOT/tools/robocode_shim/run_bridge_battle.sh" \
"$OUT/bots/ModularBot" "$ROUNDS" "$OUT/cap_${ARM}_r${RUN}.jsonl" \
> "$OUT/battle_${ARM}_r${RUN}.log" 2>&1
echo "$? ${ARM} r${RUN}"