0d35646dc99fc40cc240f9be331b07f17b3ca8c3
- actorForward: deterministic param, uses mean-only when PPOB_EVAL_ONLY=1 (eval was adding unit Gaussian noise to every action — unreliable scores) - warm_start.py: log_std initialized to -1.0 (std≈0.37) instead of copying snapshot values (were 2.27-4.68 → std 9-108, completely drowning signal) - training.env: LOG_STD_CEILING 0.0→-0.5 (cap exploration at std≈0.6)
Description
No description provided
Languages
Nim
73.7%
Python
18%
Shell
3.7%
Java
3.5%
HTML
1%
Other
0.1%