State vector additions #21

Closed
opened 2026-08-17 16:58:12 +02:00 by SirStone · 1 comment
Owner

Blocked by: Output space design for goto/aimTo

Parent map: #18

Question

What new inputs should be added to the state vector to close the feedback loop on in-flight commands?

Candidates:

  • Remaining distance to goto target (scalar)
  • Remaining gun angle to aimTo target (scalar)
  • Bearing to goto target relative to current heading?
  • Time (ticks) since current command was issued?
  • Current goto target (x, y) — or is remaining distance sufficient?

How should these be normalized? Current state vector is 42-dim — what's the new size?

**Blocked by:** [Output space design for goto/aimTo](#19) Parent map: #18 ## Question What new inputs should be added to the state vector to close the feedback loop on in-flight commands? Candidates: - Remaining distance to goto target (scalar) - Remaining gun angle to aimTo target (scalar) - Bearing to goto target relative to current heading? - Time (ticks) since current command was issued? - Current goto target (x, y) — or is remaining distance sufficient? How should these be normalized? Current state vector is 42-dim — what's the new size?
SirStone added the wayfinder:grilling label 2026-08-17 16:58:12 +02:00
Author
Owner

Resolution

2 new state inputs (42 → 44 dims):

New input Normalization
Remaining distance to goto target ÷ battlefield diagonal → [0, 1]
Remaining gun angle to aimTo target ÷ 180° → [0, 1]

Decided against adding:

  • Previous goto target (x, y) — the network sees its own position + remaining distance each tick, which is sufficient signal to learn consistency. Feeding back the previous target adds complexity the network must learn to interpret.
  • Bearing to goto target — derivable from position + heading already in state.
  • Ticks since command — redundant with remaining distance.
  • Current aimTo target — remaining angle is sufficient, gun turns are fast.
## Resolution **2 new state inputs (42 → 44 dims):** | New input | Normalization | |-----------|--------------| | Remaining distance to goto target | ÷ battlefield diagonal → [0, 1] | | Remaining gun angle to aimTo target | ÷ 180° → [0, 1] | **Decided against adding:** - Previous goto target (x, y) — the network sees its own position + remaining distance each tick, which is sufficient signal to learn consistency. Feeding back the previous target adds complexity the network must learn to interpret. - Bearing to goto target — derivable from position + heading already in state. - Ticks since command — redundant with remaining distance. - Current aimTo target — remaining angle is sufficient, gun turns are fast.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#21