Define the tunable parameter set #31

Closed
opened 2026-08-17 19:48:37 +02:00 by SirStone · 1 comment
Owner

Parent: #29

Question

Which hyperparameters from PPO_Bot's full inventory should be exposed for LLM-driven tuning?

Candidates (easy to expose, no structural impact):

  • lr (learning rate) — 3e-4
  • clipEpsilon (PPO clip ratio) — 0.2
  • entropyCoeff (exploration bonus) — 0.01
  • valueLossCoeff — 0.5
  • epochs (PPO updates per batch) — 4
  • miniBatchSize — 64
  • gamma (discount factor) — 0.99
  • lam (GAE lambda) — 0.95
  • maxGradNorm (gradient clipping) — 0.5

Candidates (medium risk):

  • logStd floor clamp (-3.0) — affects minimum exploration; duplicated in 4 places, needs consolidation first
  • initialLogStd (0.0) — only matters at fresh init, not resume

Exclude (structural/dangerous):

  • Hidden layer size (64) — invalidates saved weights on change
  • STATE_DIM / ACTION_DIM — compile-time, architectural
  • Reward shaping formula — too structural for automated tuning (see fog)

The LLM needs a bounded, safe set it can experiment with without breaking training.

Parent: #29 ## Question Which hyperparameters from PPO_Bot's full inventory should be exposed for LLM-driven tuning? **Candidates (easy to expose, no structural impact):** - `lr` (learning rate) — 3e-4 - `clipEpsilon` (PPO clip ratio) — 0.2 - `entropyCoeff` (exploration bonus) — 0.01 - `valueLossCoeff` — 0.5 - `epochs` (PPO updates per batch) — 4 - `miniBatchSize` — 64 - `gamma` (discount factor) — 0.99 - `lam` (GAE lambda) — 0.95 - `maxGradNorm` (gradient clipping) — 0.5 **Candidates (medium risk):** - `logStd` floor clamp (-3.0) — affects minimum exploration; duplicated in 4 places, needs consolidation first - `initialLogStd` (0.0) — only matters at fresh init, not resume **Exclude (structural/dangerous):** - Hidden layer size (64) — invalidates saved weights on change - STATE_DIM / ACTION_DIM — compile-time, architectural - Reward shaping formula — too structural for automated tuning (see fog) The LLM needs a bounded, safe set it can experiment with without breaking training.
SirStone added the wayfinder:grilling label 2026-08-17 19:48:37 +02:00
Author
Owner

Resolution

Decision: expose all non-structural hyperparameters.

Full tunable set (11 params, all via PPOB_* env vars with sensible defaults):

Easy (already wired):

  • PPOB_LR (3e-4)
  • PPOB_CLIP_EPSILON (0.2)
  • PPOB_ENTROPY_COEFF (0.01)
  • PPOB_VALUE_LOSS_COEFF (0.5)
  • PPOB_EPOCHS (4)
  • PPOB_MINI_BATCH_SIZE (64)
  • PPOB_GAMMA (0.99)
  • PPOB_LAM (0.95)
  • PPOB_MAX_GRAD_NORM (0.5)

Added in this ticket (with logStd floor consolidation from 4 sites → 1):

  • PPOB_LOG_STD_FLOOR (-3.0) — minimum exploration clamp
  • PPOB_INITIAL_LOG_STD (0.0) — initial logStd on fresh start

Excluded (structural/dangerous — would break saved weights or architecture):

  • Hidden layer size (64)
  • STATE_DIM / ACTION_DIM
  • Reward shaping formula
## Resolution **Decision: expose all non-structural hyperparameters.** Full tunable set (11 params, all via `PPOB_*` env vars with sensible defaults): **Easy (already wired):** - `PPOB_LR` (3e-4) - `PPOB_CLIP_EPSILON` (0.2) - `PPOB_ENTROPY_COEFF` (0.01) - `PPOB_VALUE_LOSS_COEFF` (0.5) - `PPOB_EPOCHS` (4) - `PPOB_MINI_BATCH_SIZE` (64) - `PPOB_GAMMA` (0.99) - `PPOB_LAM` (0.95) - `PPOB_MAX_GRAD_NORM` (0.5) **Added in this ticket (with logStd floor consolidation from 4 sites → 1):** - `PPOB_LOG_STD_FLOOR` (-3.0) — minimum exploration clamp - `PPOB_INITIAL_LOG_STD` (0.0) — initial logStd on fresh start **Excluded (structural/dangerous — would break saved weights or architecture):** - Hidden layer size (64) - STATE_DIM / ACTION_DIM - Reward shaping formula
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#31