Skip to content

Policies

The Policy contract, the provider matrix generated from the registry, create_policy, and swapping a policy without touching the robot.

After this page you can pick a provider for what you have, build one with create_policy, write your own in twenty lines, and swap providers by changing one string; one run_policy call takes the provider name and a policy_config.

Which provider

you have start with because
nothing yet, a laptop mock no download; it proves the loop and its report says it ignored the instruction
a Hub checkpoint for your arm (ACT, SmolVLA, Pi0, MolmoAct2) lerobot_local with embodiment= runs in this process; the embodiment map speaks sim and real
a GPU on another machine remote, create_policy("ws://gpu:8765") the robot host keeps the loop and the gate, the GPU host runs the model
a G1 or another humanoid wbc, holosoma velocity commands in, whole-body joint targets out
a target pose, no model curobo a planner reads target_pose, not words
a Cosmos endpoint cosmos3 a world model behind one URL
a task the arm has never seen Teach it no checkpoint learns your task from a page; record and train first

One gap is open: over the mesh policy_config travels but embodiment does not (#4180), so run a Hub checkpoint on a real arm from the process that owns it. Deprecated providers (moveit2, kimodo, protomotions) stay in the table until 0.7; their pages name the replacement.

The contract

Left card, the Policy: declared attributes reads_instruction, instruction_free_actions, requires_images, required_bodies, requires_action_controller and provider_name; implemented methods get_actions, set_robot_state_keys, reset and preflight. Right card, the runtime inside run_policy: it sets control_frequency and rtc_observed_delay_steps, calls set_robot_state_keys with the robot's joint names, runs preflight against the observation keys before the first tick, and installs the action controller the policy names or refuses the rollout. Between the cards, every control tick: an observation dict and the instruction go in; a chunk, one dict per tick, joint name to a float, comes back, the one green wire. Bottom: create_policy(provider, **policy_config) builds any of them; swapping providers is one string. Footnote: planners read target_pose or target_joints and ignore the words; a non-reader says so in its report.Left card, the Policy: declared attributes reads_instruction, instruction_free_actions, requires_images, required_bodies, requires_action_controller and provider_name; implemented methods get_actions, set_robot_state_keys, reset and preflight. Right card, the runtime inside run_policy: it sets control_frequency and rtc_observed_delay_steps, calls set_robot_state_keys with the robot's joint names, runs preflight against the observation keys before the first tick, and installs the action controller the policy names or refuses the rollout. Between the cards, every control tick: an observation dict and the instruction go in; a chunk, one dict per tick, joint name to a float, comes back, the one green wire. Bottom: create_policy(provider, **policy_config) builds any of them; swapping providers is one string. Footnote: planners read target_pose or target_joints and ignore the words; a non-reader says so in its report.
strands_robots/policies/base.py (abridged)
class Policy(ABC):
    control_frequency: float | None = None          # set by the runtime
    rtc_observed_delay_steps: int | None = None
    reads_instruction: bool = True                   # False: the words never shape actions
    instruction_free_actions: str | None = None             # what a non-reader does
    requires_action_controller: ClassVar[str | None] = None # the engine installs it or refuses

    @abstractmethod
    async def get_actions(self, observation_dict: dict[str, Any], instruction: str, **kwargs: Any) -> list[dict[str, Any]]: ...

    @abstractmethod
    def set_robot_state_keys(self, robot_state_keys: list[str]) -> None: ...

    def reset(self, seed: int | None = None) -> None: ...

    @classmethod
    def preflight(cls, observation_keys: set[str], **policy_config: Any) -> None: ...

    @property
    def requires_images(self) -> bool: ...            # planners return False
    @property
    def required_bodies(self) -> tuple[str, ...]: ... # a tracker names its anchor
    @property
    def children(self) -> tuple[Policy, ...]: ...     # a wrapper lists what it drives

    @property
    @abstractmethod
    def provider_name(self) -> str: ...

API reference. get_actions returns the chunk: one dict per control tick, joint name to a float. Planners read target_pose or target_joints, not words.

Providers

From the registry at build time; "Also spelled" lists the shorthands create_policy accepts.

provider class install extra also spelled what it drives trainer
mock MockPolicy none mock, random, test Sinusoidal test actions (no deps) yes
lerobot_local LerobotLocalPolicy [lerobot] lerobot Direct HuggingFace inference (no server), incl. GR00T N1.7 (policy_type='groot') yes
cosmos3 Cosmos3Policy [cosmos3-service] cosmos3, c3 NVIDIA Cosmos 3 omnimodal VLA policy (DROID/UMI/AV/bridge/OpenArm) via Cosmos Framework yes
moveit2 MoveIt2Policy [moveit2] moveit2, moveit MoveIt2 motion planning via ZMQ sidecar no
curobo CuroboPolicy [curobo] (empty: install cuRobo yourself) curobo, cumotion NVIDIA cuRobo collision-aware motion planning (in-process, CUDA) no
wbc WBCPolicy [wbc] wbc, sonic NVIDIA GR00T Whole-Body-Control (SONIC) humanoid locomotion (in-process, ONNX) no
holosoma HolosomaPolicy [holosoma] holosoma Amazon FAR Holosoma whole-body locomotion for the Unitree G1 (in-process, ONNX, Apache-2.0 weights) no
wbc_gait WBCGaitPolicy [wbc] wbc_gait, sonic_gait NVIDIA GR00T Whole-Body-Control gait-clock variant (single ONNX policy, 95-dim obs + phase clock) no
wbc_latent WBCLatentPolicy [wbc] wbc_latent, sonic_latent A VLA's SONIC motion tokens decoded into Unitree G1 joint targets (nvidia/GEAR-SONIC decoder, in-process ONNX) and tracked with SONIC's per-joint PD no
kimodo KimodoPolicy [kimodo] kimodo, kimodo_g1, text2motion NVIDIA Kimodo text-to-motion diffusion for the Unitree G1 (prompt-driven kinematic motion, HuggingFace) no
protomotions ProtoMotionsPolicy [protomotions] protomotions, gtp, gtp_g1, protomotions_g1 ProtoMotions Generalist Tracking Policy (ONNX, BeyondMimic-trained) for the Unitree G1 , consumes a MotionPlayer reference and emits PD joint targets no
remote RemotePolicy [inference] remote Remote inference over WebSocket (WS-JSON) - forwards observations to a PolicyServer and returns action chunks no
rl RLCheckpointPolicy [rl] rl Deterministic rollout of an RL training checkpoint's actor: policy.pt + policy_meta.json written by create_trainer('ppo'/'fast_sac'/'fast_td3'), or an rsl_rl run from Isaac Lab (a model_.pt directory, one such file, or a HuggingFace repo id holding either) no
microduck MicroduckPolicy [microduck] microduck Pollen Microduck locomotion policy (ONNX, normaliser fused into the graph) for the 14-DOF open biped , self-configures from ONNX metadata, feeds obs raw, decodes DEFAULT_POSE + action*scale no
flux3_action Flux3ActionPolicy [flux3] flux3_action, flux3, f3a FLUX 3 Action (Black Forest Labs) flow-matching VLA, in-process via flux_action: 2 cameras + 8-tick history + text -> 42-step absolute joint chunk at 30 Hz; SO-101 checkpoint speaks lerobot degrees + gripper percent, UnitAdapter converts to the robot's radians no
rsl_rl_onnx RslRlOnnxPolicy [sim-mjlab] rsl_rl_onnx, mjlab_onnx Deterministic rollout of an rsl_rl actor exported to ONNX by mjlab (metadata-driven observation rebuild; JointPositionAction decode) yes

composite and persistent resolve by module name; construct PersistentPolicy(provider="mock") directly.

Build one

from strands_robots.policies import create_policy, list_providers

print(list_providers())
policy = create_policy("mock")
policy.set_robot_state_keys(["shoulder_pan", "elbow_flex"])
actions = policy.get_actions_sync({"shoulder_pan": 0.0, "elbow_flex": 0.0}, "wave")
print(len(actions), actions[0])

Smart strings: a Hub id resolves to lerobot_local, ws:// to remote. A misspelled keyword raises TypeError before download; lerobot_local and kimodo need STRANDS_TRUST_REMOTE_CODE=1.

Run one, then swap it

run_policy builds the policy, runs preflight on set(sim.get_observation(robot)) before any download, then drives the loop; a registered class is one more string:

from typing import Any

from strands_robots.policies import Policy, register_policy
from strands_robots.simulation import create_simulation


class HoldPolicy(Policy):
    reads_instruction = False
    instruction_free_actions = "a fixed pose on every joint"

    def __init__(self, angle: float = 0.3, **kwargs: Any) -> None:
        self.angle = angle
        self.keys: list[str] = []

    def set_robot_state_keys(self, robot_state_keys: list[str]) -> None:
        self.keys = list(robot_state_keys)

    async def get_actions(self, observation_dict: dict[str, Any], instruction: str, **kwargs: Any) -> list[dict[str, Any]]:
        return [{k: self.angle for k in self.keys}]

    @property
    def requires_images(self) -> bool:
        return False

    @property
    def provider_name(self) -> str:
        return "hold"


register_policy("hold", lambda: HoldPolicy, aliases=["freeze"])

sim = create_simulation("mujoco")
sim.create_world()
sim.add_robot("so101")
for provider, config in (("mock", None), ("freeze", {"angle": 0.5})):
    result = sim.run_policy(robot_name="so101", policy_provider=provider, policy_config=config, instruction="wave",
                            n_steps=50, control_frequency=50.0)
    print(provider, result["status"], round(sim.get_observation("so101", skip_images=True)["1"], 3))
    print(result["content"][0]["text"].splitlines()[-1])
sim.cleanup()

You should see:

mock success 0.493
Note: MockPolicy does not read the instruction. Its actions - a test motion on every joint - were commanded to the robot whatever the task says; nothing above means the task was performed.
freeze success 0.501
Note: HoldPolicy does not read the instruction. Its actions - a fixed pose on every joint - were commanded to the robot whatever the task says; nothing above means the task was performed.

The notes come from reads_instruction = False: a policy that never reads the words says so, so an agent cannot relay a test motion as done. Registering a built-in name needs overwrite=True.

On hardware, start_task(instruction, policy_provider=..., **policy_config) takes the provider string and run_policy(create_policy(...)) a built object; the operator gate sits in front of both.

Edit page