jaxdem.rl.environments.single_roller#
Environment where a single agent rolls towards a target on the floor.
Functions
|
Normal, frictional, and restitution forces for a sphere on a \(z = 0\) plane. |
Classes
|
Single-agent 3D navigation via torque-controlled rolling. |
- class jaxdem.rl.environments.single_roller.SingleRoller(state: State, system: System, env_params: dict[str, Any])#
Bases:
EnvironmentSingle-agent 3D navigation via torque-controlled rolling.
The agent is a sphere resting on a \(z = 0\) floor under gravity. Actions are 3-D torque vectors; translational motion arises from frictional contact with the floor (see
frictional_wall_force()). A viscous drag-friction * veland a fixed angular damping of-friction * ang_velare applied each step.The reward uses potential-based shaping with a proximity-gated kinetic-energy term:
\[\varphi(d, K) = \exp\!\left(-2 d - \frac{K}{\text{ke\_tau}}\,e^{-\text{ke\_gate} \cdot d}\right)\]where \(d\) is the distance to the objective, \(K\) is the total (translational + rotational) kinetic energy,
ke_tauis the KE scale that sets the overall strength of the penalty, andke_gatecontrols how sharply KE sensitivity falls off with distance — largerke_gatemeans KE only matters very close to the objective.The shaping credit is \(F_t = \varphi(d_t, K_t) - \varphi(d_{t-1}, K_{t-1})\), so kinetic energy is penalised only near the objective — far away the gate \(e^{-\text{ke\_gate} \cdot d} \to 0\) and fast motion is free.
Per-step reward:
\[\mathrm{rew}_t = \frac{F_t + b \cdot \mathbb{1}[d_t \le r]}{b}\]where \(b\) is the near-goal bonus and \(r\) is the agent radius.
Notes
The observation vector per agent is:
Feature
Size
Unit direction to objective
2
Clamped displacement (x, y)
2
Velocity (x, y)
2
Angular velocity
3
If one wants some realistic parameters for training,
skip_frames = 50will give a response rate of 200 Hz, meaning thatnum_steps_epoch = 100gives a horizon of 0.5 seconds.- classmethod Create(min_box_size: float = 40.0, max_box_size: float = 40.0, max_steps: int = 20000, friction: float = 0.2, near_goal_bonus: float = 0.1, ke_tau: float = 5.0, ke_gate: float = 4.0) SingleRoller[source]#
Create a single-agent roller environment.
- Parameters:
min_box_size (float) – Range for the random square domain side length.
max_box_size (float) – Range for the random square domain side length.
max_steps (int) – Episode length in physics steps.
friction (float) – Viscous drag coefficient applied as
-friction * vel.ke_tau (float) – Overall strength of the KE term in the potential (larger = less important). See class docstring.
ke_gate (float) – Distance decay rate of KE sensitivity (larger = KE only matters very close to the goal). See class docstring.
- Returns:
A freshly constructed environment (call
reset()before use).- Return type:
- static reset(env: SingleRoller, key: Array | ndarray | bool | number | bool | int | float | complex) Environment[source]#
Randomly place the agent and objective on the floor.
- Parameters:
env (Environment) – Current environment instance.
key (ArrayLike) – JAX PRNG key.
- Returns:
Freshly initialised environment.
- Return type:
- static step(env: SingleRoller, action: Array) Environment[source]#
Apply a torque action, advance physics by one step.
- Parameters:
env (Environment) – Current environment.
action (jax.Array) – 3-D torque vector per agent.
- Returns:
Updated environment after one physics step.
- Return type:
- static observation(env: SingleRoller) Array[source]#
Per-agent observation vector.
Contents per agent:
Unit displacement to objective projected to x-y (shape
(2,)).Clamped displacement to objective projected to x-y (shape
(2,)).Velocity projected to x-y (shape
(2,)).Angular velocity (shape
(3,)).
- Returns:
Shape
(N, 9).- Return type:
jax.Array
- static reward(env: SingleRoller) Array[source]#
Returns a vector of per-agent rewards.
Potential-based shaping with a proximity-gated KE term:
\[\varphi(d, K) = \exp\!\left(-2 d - \frac{K}{\text{ke\_tau}}\,e^{-\text{ke\_gate} \cdot d}\right)\]The gate \(e^{-\text{ke\_gate} \cdot d}\) suppresses the KE term away from the objective, so fast motion is free until the agent is close;
ke_tausets the overall strength of the penalty.Per-step reward:
\[\mathrm{rew}_t = \frac{\varphi(d_t, K_t) - \varphi(d_{t-1}, K_{t-1}) + b \cdot \mathbb{1}[d_t \le r]}{b}\]where \(b\) is the near-goal bonus and \(r\) is the agent radius.
- Returns:
Shape
(N,).- Return type:
jax.Array
- static done(env: SingleRoller) Array[source]#
Truewhenstep_countexceedsmax_steps.