Exploration noise #10

Closed
opened 2026-08-16 11:51:25 +02:00 by SirStone · 1 comment
Owner

Question

Does PPO need explicit exploration noise on top of the stochastic policy, or is entropy in the action distribution sufficient?

Options:

  • No noise — PPO's stochastic policy (Gaussian with learned std) provides exploration naturally. Entropy bonus in the loss prevents premature convergence.
  • Ornstein-Uhlenbeck noise — temporally correlated, common in DDPG but unusual in PPO.
  • Action-space noise with decay schedule — start noisy, anneal to pure policy.

If entropy bonus is enough, what coefficient? If explicit noise, what kind and decay?

## Question Does PPO need explicit exploration noise on top of the stochastic policy, or is entropy in the action distribution sufficient? Options: - No noise — PPO's stochastic policy (Gaussian with learned std) provides exploration naturally. Entropy bonus in the loss prevents premature convergence. - Ornstein-Uhlenbeck noise — temporally correlated, common in DDPG but unusual in PPO. - Action-space noise with decay schedule — start noisy, anneal to pure policy. If entropy bonus is enough, what coefficient? If explicit noise, what kind and decay?
SirStone added the wayfinder:grilling label 2026-08-16 11:51:25 +02:00
Author
Owner

Resolution

No explicit exploration noise. PPO's stochastic Gaussian policy (learned log-std per action dimension, decided in #9) provides exploration naturally. Entropy bonus in the PPO loss prevents premature convergence.

Log-std floor: log_std = max(log_std, -3.0) (std ≈ 0.05 minimum). Prevents total policy collapse when switching opponents — the bot always retains a minimum level of randomness. Cheap safety net; multi-opponent curriculum training is the real fix (remains in fog).

Entropy bonus coefficient: 0.01 (standard PPO default). Tunable hyperparameter.

What was ruled out:

  • Ornstein-Uhlenbeck noise — DDPG technique, fights the learned std in PPO
  • Action-space noise with decay schedule — redundant with learned log-std
  • Opponent recognition / fingerprinting — not in scope; the bot adapts within a battle via the 5-tick enemy history window in the state vector, not via explicit opponent identification
## Resolution **No explicit exploration noise.** PPO's stochastic Gaussian policy (learned log-std per action dimension, decided in #9) provides exploration naturally. Entropy bonus in the PPO loss prevents premature convergence. **Log-std floor:** `log_std = max(log_std, -3.0)` (std ≈ 0.05 minimum). Prevents total policy collapse when switching opponents — the bot always retains a minimum level of randomness. Cheap safety net; multi-opponent curriculum training is the real fix (remains in fog). **Entropy bonus coefficient:** 0.01 (standard PPO default). Tunable hyperparameter. **What was ruled out:** - Ornstein-Uhlenbeck noise — DDPG technique, fights the learned std in PPO - Action-space noise with decay schedule — redundant with learned log-std - Opponent recognition / fingerprinting — not in scope; the bot adapts within a battle via the 5-tick enemy history window in the state vector, not via explicit opponent identification
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#10