Exploration noise #10
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Question
Does PPO need explicit exploration noise on top of the stochastic policy, or is entropy in the action distribution sufficient?
Options:
If entropy bonus is enough, what coefficient? If explicit noise, what kind and decay?
Resolution
No explicit exploration noise. PPO's stochastic Gaussian policy (learned log-std per action dimension, decided in #9) provides exploration naturally. Entropy bonus in the PPO loss prevents premature convergence.
Log-std floor:
log_std = max(log_std, -3.0)(std ≈ 0.05 minimum). Prevents total policy collapse when switching opponents — the bot always retains a minimum level of randomness. Cheap safety net; multi-opponent curriculum training is the real fix (remains in fog).Entropy bonus coefficient: 0.01 (standard PPO default). Tunable hyperparameter.
What was ruled out: