Policies¶
The Policy contract, the provider matrix generated from the registry, create_policy, and swapping a policy without touching the robot.
After this page you can pick a provider for what you have, build one with create_policy, write your own in twenty lines, and swap providers by changing one string; one run_policy call takes the provider name and a policy_config.
Which provider¶
| you have | start with | because |
|---|---|---|
| nothing yet, a laptop | mock |
no download; it proves the loop and its report says it ignored the instruction |
| a Hub checkpoint for your arm (ACT, SmolVLA, Pi0, MolmoAct2) | lerobot_local with embodiment= |
runs in this process; the embodiment map speaks sim and real |
| a GPU on another machine | remote, create_policy("ws://gpu:8765") |
the robot host keeps the loop and the gate, the GPU host runs the model |
| a G1 or another humanoid | wbc, holosoma |
velocity commands in, whole-body joint targets out |
| a target pose, no model | curobo |
a planner reads target_pose, not words |
| a Cosmos endpoint | cosmos3 |
a world model behind one URL |
| a task the arm has never seen | Teach it | no checkpoint learns your task from a page; record and train first |
One gap is open: over the mesh policy_config travels but embodiment does not (#4180), so run a Hub checkpoint on a real arm from the process that owns it. Deprecated providers (moveit2, kimodo, protomotions) stay in the table until 0.7; their pages name the replacement.
The contract¶
class Policy(ABC):
control_frequency: float | None = None # set by the runtime
rtc_observed_delay_steps: int | None = None
reads_instruction: bool = True # False: the words never shape actions
instruction_free_actions: str | None = None # what a non-reader does
requires_action_controller: ClassVar[str | None] = None # the engine installs it or refuses
@abstractmethod
async def get_actions(self, observation_dict: dict[str, Any], instruction: str, **kwargs: Any) -> list[dict[str, Any]]: ...
@abstractmethod
def set_robot_state_keys(self, robot_state_keys: list[str]) -> None: ...
def reset(self, seed: int | None = None) -> None: ...
@classmethod
def preflight(cls, observation_keys: set[str], **policy_config: Any) -> None: ...
@property
def requires_images(self) -> bool: ... # planners return False
@property
def required_bodies(self) -> tuple[str, ...]: ... # a tracker names its anchor
@property
def children(self) -> tuple[Policy, ...]: ... # a wrapper lists what it drives
@property
@abstractmethod
def provider_name(self) -> str: ...
API reference. get_actions returns the chunk: one dict per control tick, joint name to a float. Planners read target_pose or target_joints, not words.
Providers¶
From the registry at build time; "Also spelled" lists the shorthands create_policy accepts.
| provider | class | install extra | also spelled | what it drives | trainer |
|---|---|---|---|---|---|
mock |
MockPolicy |
none | mock, random, test |
Sinusoidal test actions (no deps) | yes |
lerobot_local |
LerobotLocalPolicy |
[lerobot] |
lerobot |
Direct HuggingFace inference (no server), incl. GR00T N1.7 (policy_type='groot') | yes |
cosmos3 |
Cosmos3Policy |
[cosmos3-service] |
cosmos3, c3 |
NVIDIA Cosmos 3 omnimodal VLA policy (DROID/UMI/AV/bridge/OpenArm) via Cosmos Framework | yes |
moveit2 |
MoveIt2Policy |
[moveit2] |
moveit2, moveit |
MoveIt2 motion planning via ZMQ sidecar | no |
curobo |
CuroboPolicy |
[curobo] (empty: install cuRobo yourself) |
curobo, cumotion |
NVIDIA cuRobo collision-aware motion planning (in-process, CUDA) | no |
wbc |
WBCPolicy |
[wbc] |
wbc, sonic |
NVIDIA GR00T Whole-Body-Control (SONIC) humanoid locomotion (in-process, ONNX) | no |
holosoma |
HolosomaPolicy |
[holosoma] |
holosoma |
Amazon FAR Holosoma whole-body locomotion for the Unitree G1 (in-process, ONNX, Apache-2.0 weights) | no |
wbc_gait |
WBCGaitPolicy |
[wbc] |
wbc_gait, sonic_gait |
NVIDIA GR00T Whole-Body-Control gait-clock variant (single ONNX policy, 95-dim obs + phase clock) | no |
wbc_latent |
WBCLatentPolicy |
[wbc] |
wbc_latent, sonic_latent |
A VLA's SONIC motion tokens decoded into Unitree G1 joint targets (nvidia/GEAR-SONIC decoder, in-process ONNX) and tracked with SONIC's per-joint PD | no |
kimodo |
KimodoPolicy |
[kimodo] |
kimodo, kimodo_g1, text2motion |
NVIDIA Kimodo text-to-motion diffusion for the Unitree G1 (prompt-driven kinematic motion, HuggingFace) | no |
protomotions |
ProtoMotionsPolicy |
[protomotions] |
protomotions, gtp, gtp_g1, protomotions_g1 |
ProtoMotions Generalist Tracking Policy (ONNX, BeyondMimic-trained) for the Unitree G1 , consumes a MotionPlayer reference and emits PD joint targets | no |
remote |
RemotePolicy |
[inference] |
remote |
Remote inference over WebSocket (WS-JSON) - forwards observations to a PolicyServer and returns action chunks | no |
rl |
RLCheckpointPolicy |
[rl] |
rl |
Deterministic rollout of an RL training checkpoint's actor: policy.pt + policy_meta.json written by create_trainer('ppo'/'fast_sac'/'fast_td3'), or an rsl_rl run from Isaac Lab (a model_ |
no |
microduck |
MicroduckPolicy |
[microduck] |
microduck |
Pollen Microduck locomotion policy (ONNX, normaliser fused into the graph) for the 14-DOF open biped , self-configures from ONNX metadata, feeds obs raw, decodes DEFAULT_POSE + action*scale | no |
flux3_action |
Flux3ActionPolicy |
[flux3] |
flux3_action, flux3, f3a |
FLUX 3 Action (Black Forest Labs) flow-matching VLA, in-process via flux_action: 2 cameras + 8-tick history + text -> 42-step absolute joint chunk at 30 Hz; SO-101 checkpoint speaks lerobot degrees + gripper percent, UnitAdapter converts to the robot's radians | no |
rsl_rl_onnx |
RslRlOnnxPolicy |
[sim-mjlab] |
rsl_rl_onnx, mjlab_onnx |
Deterministic rollout of an rsl_rl actor exported to ONNX by mjlab (metadata-driven observation rebuild; JointPositionAction decode) | yes |
composite and persistent resolve by module name; construct PersistentPolicy(provider="mock") directly.
Build one¶
from strands_robots.policies import create_policy, list_providers
print(list_providers())
policy = create_policy("mock")
policy.set_robot_state_keys(["shoulder_pan", "elbow_flex"])
actions = policy.get_actions_sync({"shoulder_pan": 0.0, "elbow_flex": 0.0}, "wave")
print(len(actions), actions[0])
Smart strings: a Hub id resolves to lerobot_local, ws:// to remote. A misspelled keyword raises TypeError before download; lerobot_local and kimodo need STRANDS_TRUST_REMOTE_CODE=1.
Run one, then swap it¶
run_policy builds the policy, runs preflight on set(sim.get_observation(robot)) before any download, then drives the loop; a registered class is one more string:
from typing import Any
from strands_robots.policies import Policy, register_policy
from strands_robots.simulation import create_simulation
class HoldPolicy(Policy):
reads_instruction = False
instruction_free_actions = "a fixed pose on every joint"
def __init__(self, angle: float = 0.3, **kwargs: Any) -> None:
self.angle = angle
self.keys: list[str] = []
def set_robot_state_keys(self, robot_state_keys: list[str]) -> None:
self.keys = list(robot_state_keys)
async def get_actions(self, observation_dict: dict[str, Any], instruction: str, **kwargs: Any) -> list[dict[str, Any]]:
return [{k: self.angle for k in self.keys}]
@property
def requires_images(self) -> bool:
return False
@property
def provider_name(self) -> str:
return "hold"
register_policy("hold", lambda: HoldPolicy, aliases=["freeze"])
sim = create_simulation("mujoco")
sim.create_world()
sim.add_robot("so101")
for provider, config in (("mock", None), ("freeze", {"angle": 0.5})):
result = sim.run_policy(robot_name="so101", policy_provider=provider, policy_config=config, instruction="wave",
n_steps=50, control_frequency=50.0)
print(provider, result["status"], round(sim.get_observation("so101", skip_images=True)["1"], 3))
print(result["content"][0]["text"].splitlines()[-1])
sim.cleanup()
You should see:
mock success 0.493
Note: MockPolicy does not read the instruction. Its actions - a test motion on every joint - were commanded to the robot whatever the task says; nothing above means the task was performed.
freeze success 0.501
Note: HoldPolicy does not read the instruction. Its actions - a fixed pose on every joint - were commanded to the robot whatever the task says; nothing above means the task was performed.
The notes come from reads_instruction = False: a policy that never reads the words says so, so an agent cannot relay a test motion as done. Registering a built-in name needs overwrite=True.
On hardware, start_task(instruction, policy_provider=..., **policy_config) takes the provider string and run_policy(create_policy(...)) a built object; the operator gate sits in front of both.