tune(PPO_Bot): logStd=-2.0 (std≈0.135), entropy=0, ceiling=-1.0

Stochastic eval at std≈0.37 was 0/10 vs Corners (deterministic: 10/10).
Warm-start policy is correct but brittle — any noise breaks it.
- log_std initialized to -2.0 (std≈0.135) for moderate exploration
- entropy_coeff=0.0 (no push toward exploration during fine-tuning)
- logStd ceiling=-1.0 (cap at std≈0.37)
This commit is contained in:
2026-08-20 15:28:15 +02:00
parent 0d35646dc9
commit ca3e3d2272
2 changed files with 5 additions and 5 deletions
+3 -3
View File
@@ -37,10 +37,10 @@ for name in unchanged:
np.save(DST / f"{name}.npy", data)
print(f" {name}: {data.shape} copied")
# Initialize log_std to -1.0 (std ≈ 0.37) — snapshot values (2.27–4.68) are too high for fine-tuning
log_std = np.full(6, -1.0, dtype=np.float32)
# Initialize log_std to -2.0 (std ≈ 0.135) — tighter than -1.0, proven workable
log_std = np.full(6, -2.0, dtype=np.float32)
np.save(DST / "log_std.npy", log_std)
print(f" log_std: initialized to -1.0 (std≈0.37), shape={log_std.shape}")
print(f" log_std: initialized to -2.0 (std≈0.135), shape={log_std.shape}")
# Copy unchanged Adam moments (all except w1, which were handled above)
unchanged_adam = [