Skip to content

Policies

The Policy contract, how a provider string resolves through create_policy, and the cache that keeps a loaded model.

A Policy turns an observation into actions. Providers live in registry/policies.json and resolve through create_policy: the Policy contract, how a provider string resolves, the persistent cache that keeps a loaded model between calls.

Contract

Abstract base class for robot policies (VLA, motion planners, MPC, scripted).

The :class:Policy ABC is intentionally agnostic about how actions are produced. Built-in providers (mock, lerobot_local, remote) are VLA-style, but the same interface is the right shape for:

  • Classical motion planners - cuRobo, MoveIt2, OMPL, RRT*: take a goal pose and joint state, return a collision-free trajectory.
  • Model-predictive controllers (MPC) - solve a finite-horizon optimal control problem each tick.
  • Scripted / pure-IK trajectories - analytic IK followed by interpolation; zero learning involved.

Non-VLA implementations typically set :attr:Policy.requires_images to False to skip camera rendering (~10x throughput win at 500Hz) and read their goal from the well-known **kwargs keys documented on :meth:Policy.get_actions rather than parsing the natural-language instruction string.

See :class:~strands_robots.policies.mock.MockPolicy for the canonical non-VLA reference implementation.

Policy

Bases: ABC

Abstract base class for robot policies (VLA, motion planners, MPC, scripted).

All policies implement async :meth:get_actions. For convenience, a synchronous wrapper :meth:get_actions_sync is provided.

The interface is general enough to cover both VLA-style providers (consume images + instruction, output joint targets) and non-VLA providers such as classical motion planners (cuRobo, MoveIt2), model-predictive controllers, and pure-IK / scripted trajectories. Non-VLA providers typically set :attr:requires_images to False and read their goal from the well-known **kwargs keys documented on :meth:get_actions.

All providers MUST honour the per-tick action value convention documented on :meth:get_actions: each action value is a python float (single-DOF) or list[float] (multi-DOF group), never a raw np.ndarray, so downstream consumers handle every provider's output uniformly regardless of its internal compute backend. See MockPolicy for the canonical reference.

children property

children: tuple[Policy, ...]

The policies this one delegates to, in the order it consults them.

Default () - a leaf policy that runs its own inference. A wrapper returns the policies it drives: :class:~strands_robots.policies.composite.CompositePolicy its lower and upper children, :class:~strands_robots.policies.persistent.PersistentPolicy the single policy it holds warm.

This is the same "policy declares, runtime supplies" contract as :attr:requires_images and :attr:required_bodies, applied to a capability probe rather than an observation. A probe answers about the object it is handed, and a wrapper is a different object than the policy inside it, so an isinstance test against a wrapper reports the wrapped policy's capability as absent. The MuJoCo backend's WBC torque shim is the motivating case: it is required for a :class:~strands_robots.policies.wbc.WBCPolicy to hold a stable gait on a position-servo scene, and the physics does not change when that policy is wrapped - only the type of the object the probe sees does. Declaring the children lets one probe walk to the policy that answers, instead of every probe having to learn the name of every wrapper.

Returns:

Type Description
tuple[Policy, ...]

The child policies. Empty (the default) means this policy is a leaf.

execution_horizon property

execution_horizon: int

Number of actions the SIM consumes from one get_actions chunk before re-querying.

This is the SINGLE source of truth for the re-query interval; a chunk consumer (the single-policy runner, the multi-episode eval loop, the synchronized multi-robot loop) reads it via :func:resolve_chunk_length and never inspects actions_per_step directly. Distinguishing the re-query interval from the trained chunk length is what makes Real-Time Chunking (RTC) actually engage:

  • RTC policy -> the RTC execution_horizon (typically << the trained chunk). The policy is re-queried mid-chunk so it can blend the unexecuted tail of the previous chunk (prev_chunk_left_over) into the next one. Re-querying only after the full trained chunk drains leaves that tail permanently empty and silently degrades RTC to plain open-loop replay.
  • chunked open-loop (ACT, diffusion, pi0/SmolVLA without RTC) -> actions_per_step (the trained chunk; truncating drops its tail and forces an out-of-distribution re-query).
  • single-step (MockPolicy, classical planners) -> 1.

The default derives from actions_per_step (1 when undeclared), so a single-step or chunked open-loop policy needs no override; only a policy with an inference-time budget distinct from its trained chunk (RTC) overrides this.

provider_name abstractmethod property

provider_name: str

Get provider name for identification.

required_bodies property

required_bodies: tuple[str, ...]

Named rigid bodies whose world pose this policy needs in its observation.

Default () - most policies are driven by joint state alone and pay nothing for this. A whole-body motion-mimic tracker (ProtoMotions GTP, PHC, OmniH2O and the text-to-motion pipelines built on them) is the motivating case: its network consumes the world orientation of a single anchor link - torso_link on a Unitree G1 - which is NOT derivable from the observation's floating-base signals. base_quat is the pelvis, and the torso differs from it by the three waist joints, so a tracker written against base_quat silently feeds the network the wrong frame whenever the waist is not neutral.

Declaring the bodies here is the same "policy declares, runtime supplies" contract as :attr:requires_images: the runtime (:class:~strands_robots.simulation.policy_runner.PolicyRunner) resolves the names ONCE before the rollout and merges the pose of each into every observation it hands to :meth:get_actions, under the keys documented on :meth:~strands_robots.simulation.base.SimEngine.get_observation::

body.<name>.pos      # world x, y, z (m)
body.<name>.quat     # world orientation w, x, y, z
body.<name>.lin_vel  # world linear velocity x, y, z (m/s)
body.<name>.ang_vel  # world angular velocity x, y, z (rad/s)

A policy that declares a body the scene does not contain fails at the start of the rollout with the available body names, rather than reading a missing key as a zero pose on every tick.

Every surface that reads this collects it over the whole policy tree through :func:collect_required_bodies, so a policy that declares a body is honoured when it is wrapped: a wrapper which does not override this property does not hide its child's declaration. That one owner is shared with the remote-inference handshake, so a policy served over the wire declares the same bodies it does in-process.

Returns:

Type Description
str

Ordered, de-duplicated body names. Empty (the default) means the

...

observation is left exactly as the backend produced it.

requires_images property

requires_images: bool

Whether this policy needs camera frames in its observation.

Default True (most VLA policies do). Subclasses that only consume joint state (e.g. MockPolicy, classical motion planners such as cuRobo / MoveIt2, MPC, pure-IK controllers, scripted trajectories) can return False to let the simulation skip expensive camera rendering - a ~10x throughput win at 500Hz when no cameras are needed.

get_actions abstractmethod async

get_actions(observation_dict: dict[str, Any], instruction: str, **kwargs: Any) -> list[dict[str, Any]]

Get actions from policy given observation and instruction.

Parameters:

Name Type Description Default
observation_dict dict[str, Any]

Robot observation (cameras + state). VLA providers consume both observation.images.* and observation.state. Non-VLA providers typically consume observation.state only and set :attr:requires_images to False to skip camera rendering.

required
instruction str

Natural language instruction. Required by the signature for VLA providers; non-VLA providers (motion planners, MPC, scripted) may ignore it and read the goal from **kwargs instead.

required
**kwargs Any

Provider-specific parameters. The following keys are well-known and SHOULD be honoured by non-VLA providers when present so callers don't have to JSON-encode goals into the instruction string:

  • target_pose: list[float] - Cartesian goal as [x, y, z, qw, qx, qy, qz] (position in metres, orientation as a unit quaternion in the robot base frame).
  • target_joints: dict[str, float] - joint-space goal keyed by joint name; values are in radians (revolute) or metres (prismatic).
  • target_velocity: list[float] - base-velocity goal as [vx, vy, omega] (linear m/s, angular rad/s, in the robot base frame), read by locomotion providers. Unlike the keys above, the component count is deliberately NOT fixed by this contract, because shipped receivers disagree: whole-body controllers require the three components named above, while planar drivers read [vx, vy]. Each receiver states its own arity and refuses a shape it cannot use, so this key crosses a transport on the per-component domain alone.
  • world_update: dict | None - per-call world refresh for collision-aware planners (e.g. point cloud / depth image / mesh updates). None means "reuse the world configured at init time".

Providers MUST ignore unknown **kwargs rather than raising, so callers can pass shared keys across providers.

{}

Returns:

Type Description
list[dict[str, Any]]

List of action dicts for robot execution. Each dict maps a

list[dict[str, Any]]

robot state key (joint/actuator name) to its target value

list[dict[str, Any]]

for that tick.

list[dict[str, Any]]

Values MUST be JSON / python-native: a python float for

list[dict[str, Any]]

a single-DOF actuator, or a list[float] for a multi-DOF

list[dict[str, Any]]

actuator group. Implementations MUST NOT return raw

list[dict[str, Any]]

np.ndarray objects -- coerce with .tolist() /

list[dict[str, Any]]

float(...) before returning -- so downstream consumers can

list[dict[str, Any]]

treat every provider's output uniformly (e.g. float(v) on a

list[dict[str, Any]]

scalar, len(v) on a group) regardless of the policy's

list[dict[str, Any]]

internal compute backend.

list[dict[str, Any]]

The list length is the action-chunk horizon; consumers execute

list[dict[str, Any]]

it at a fixed control rate (e.g. 50Hz).

get_actions_sync

get_actions_sync(observation_dict: dict[str, Any], instruction: str, **kwargs: Any) -> list[dict[str, Any]]

Synchronous convenience wrapper around :meth:get_actions.

Safe to call from sync code, event loops, or notebooks. Resolution is delegated to :mod:strands_robots._async_utils, the one owner of "resolve a policy coroutine in a sync context": with no running loop it is asyncio.run, and inside one the coroutine is offloaded to that module's single reused worker thread.

Delegating rather than re-deriving is what keeps this callable at control rate from a running loop -- a notebook cell and any async host take the offload branch, and constructing a private executor there starts and joins one OS thread per call. Every in-process rollout path already resolves through the same owner, so the wrapper and the runner cannot drift on which branch a caller lands in.

is_chunk_emitting

is_chunk_emitting() -> bool

Whether this policy returns multi-action chunks per get_actions.

A chunk-emitting policy (ACT, diffusion, pi0, pi0.5, pi0-FAST, SmolVLA, MolmoAct2) returns more than one action per inference, so its inference latency can be hidden behind the EXECUTION of the current chunk while the next chunk is computed in the background. :meth:PolicyRunner.run auto-enables that overlap (run_policy(async_rtc=None)) only when this is true AND the policy blends the seam (supports_rtc); single-step policies (MockPolicy, classical planners) gain nothing from overlap and stay on the synchronous loop.

The default derives the answer from the re-query interval the consumer actually drives - :attr:execution_horizon - so ANY policy that emits a chunk longer than one action is detected without enumerating provider names: a model under RTC reports its RTC horizon (> 1), a chunked open-loop model reports its trained chunk length (> 1), and a single-step policy reports 1. Providers whose chunk shape is not visible through execution_horizon (e.g. a model that must be driven via predict_action_chunk) override this.

:class:~strands_robots.policies.mock.MockPolicy returns eight actions per call and still declares 1 on purpose: its sinusoid is a function of a step counter, so a re-query continues the same curve wherever it happens, there is no inference latency for the async pipeline to hide, and 1 keeps the reference policy on the synchronous loop every tutorial and test reads. A provider that pays for its chunk (Cosmos 3, every LeRobot checkpoint) declares actions_per_step instead.

Returns:

Type Description
bool

True when the policy emits multi-action chunks; False for

bool

single-step policies.

preflight classmethod

preflight(observation_keys: set[str], **policy_config: Any) -> None

Cheap pre-construction validation hook (no download, no instantiation).

Called by the simulation's run_policy / eval_policy BEFORE :func:~strands_robots.policies.create_policy builds the policy - and therefore before any model weight download - with the set of observation keys the runtime will feed the policy. Override this to fail fast on a misconfiguration (e.g. sim camera names that cannot be routed to the model's declared image inputs) instead of surfacing it as a confusing failure deep inside inference after a multi-minute weight download.

The default implementation is a no-op. Implementations MUST be cheap: no network access, no model instantiation - only local metadata (policy_config plus packaged JSON such as the embodiment registry) and the provided observation_keys.

Parameters:

Name Type Description Default
observation_keys set[str]

Keys present in the runtime observation dict the policy will receive (joint state names + attached camera names), as returned by SimEngine.get_observation.

required
**policy_config Any

The same provider kwargs that will be forwarded to the policy constructor by create_policy.

{}

Raises:

Type Description
ValueError

When the configuration cannot consume the runtime observation (e.g. a required camera source key is absent and no override maps an available key onto the model's image feature).

reset

reset(seed: int | None = None) -> None

Reset per-episode policy state.

Default implementation is a no-op. Policies that hold per-episode state (e.g. diffusion sampler RNG, action chunk caches, KV-caches) should override to apply the reset.

For SERVICE-mode policies (e.g. Cosmos3Policy(host=...) over WebSocket), the override forwards the call to the server so its per-episode RNG state can be re-initialised - without this, set_eval_seed only seeds the client-side process, leaving the server's diffusion sampler RNG drifting across calls and breaking reproducibility (#187).

Parameters:

Name Type Description Default
seed int | None

Optional master seed forwarded to the policy's random-number generators. When None, implementations may apply a default seed or leave RNG state untouched.

None

set_control_frequency

set_control_frequency(hz: float) -> None

Tell the policy the control rate (Hz) of the executing loop.

The runtime that drives the policy (PolicyRunner.run / evaluate) calls this once before the rollout loop so providers that estimate an inference delay in action steps (Real-Time Chunking) can convert their measured wall-clock latency into the correct number of steps. Without it, such providers fall back to a hardcoded rate and silently mis-blend chunks at any other control frequency.

Parameters:

Name Type Description Default
hz float

Finite positive control frequency in Hz.

required

Raises:

Type Description
ValueError

If hz is not a finite positive number. The rate is the multiplier that converts a measured latency into a step count, so it has to be checked where it arrives rather than where it is read: nan and inf both survive a bare hz <= 0 test (neither compares <= to anything) and are only discovered later, inside the provider, as a bare ValueError/OverflowError out of the int() that converts the delay - and not on the first inference, because the estimator returns 0 until it has a latency sample. bool is refused for the same reason it is everywhere else in this domain: True would install a silent 1 Hz clock.

set_robot_state_keys abstractmethod

set_robot_state_keys(robot_state_keys: list[str]) -> None

Configure the policy with robot state keys.

These are the ordered joint/motor names the policy emits as its action-dict keys, so they decide which actuator each action value is sent to. An implementation must refuse a malformed list rather than bind it. Most do so through the shared domain :func:~strands_robots.utils.name_list_error, gated on a truthy value because an empty list already means "auto-detect" on the providers that support it. :class:~strands_robots.policies.wbc.policy.WBCPolicy is already total without it: it resolves every joint it drives BY NAME inside the caller's list, so any malformed shape fails that membership check instead - and it deliberately tolerates a repeated name, which resolves to its first occurrence.

Unlike :meth:set_control_frequency and :meth:set_rtc_observed_delay, this setter has no shared implementation to carry the domain: each provider binds the names into its own layout, so each refuses at its own entry. That parity is pinned structurally by the policy state-key name-list contract tests.

Parameters:

Name Type Description Default
robot_state_keys list[str]

Ordered list of distinct non-blank joint/motor names.

required

Raises:

Type Description
ValueError

If robot_state_keys is not such a list.

set_rtc_observed_delay

set_rtc_observed_delay(steps: int | None) -> None

Tell the policy how many control steps elapse during inference.

The runtime that drives the policy calls this before each get_actions so Real-Time Chunking providers can compute the chunk-seam offset deterministically instead of deriving it from wall-clock latency. In a synchronous eval loop the world is paused during inference, so exactly 0 steps elapse; in the async overlap pipeline the count is the number of still-pending steps of the chunk being executed. Either way it is a known integer, not a measurement.

Parameters:

Name Type Description Default
steps int | None

Non-negative control-step count, or None to clear the override and let the provider fall back to its wall-clock estimate.

required

Raises:

Type Description
ValueError

If steps is neither None nor a non-negative int. The count is an offset into the action chunk, so a fractional value is not a smaller offset and bool is not a count of one - both were previously coerced by the int() below into a neighbouring value the caller never asked for, which moves the chunk seam silently.

ChunkedPolicy

Bases: Protocol

Introspection contract for policies that emit ACTION CHUNKS.

A chunked policy returns more than one action per :meth:Policy.get_actions call: a model trained for N-step open-loop replay (ACT, diffusion, pi0, SmolVLA, MolmoAct2) emits a length-N chunk that a consumer executes before re-querying. The chunk PRODUCER is the existing async :meth:Policy.get_actions - this protocol deliberately does NOT add a second chunk-producing method (that would split one contract across two code paths); it only surfaces the metadata a consumer needs to drive an already-produced chunk correctly.

Every consumer of a chunk (the single-policy runner, the multi-episode eval loop, and the synchronized multi-robot loop) must size the chunk the same way - see :func:resolve_chunk_length. Routing all of them through one helper that reads this contract keeps a chunk-emitting policy from being truncated differently depending on which loop happens to drive it.

The protocol is runtime_checkable so a consumer can branch on isinstance(policy, ChunkedPolicy) and a type checker rejects a non-chunked policy where a chunked one is required.

Attributes:

Name Type Description
actions_per_step int

Number of actions the policy intends a consumer to execute open-loop from one get_actions chunk before re-querying (the policy's trained chunk length). Truncating below this drops the chunk tail and forces an out-of-distribution re-query.

supports_rtc bool

Whether the policy blends chunk seams internally via Real-Time Chunking - it carries prev-chunk state across re-queries so consecutive chunks join smoothly. Introspection only; a consumer never has to drive RTC, the policy does it inside get_actions.

resolve_chunk_length

resolve_chunk_length(policy: Policy, action_horizon: int) -> int

Effective number of actions to consume from one get_actions chunk.

Centralizes the single re-query rule every consumer must apply identically. The number of actions consumed before re-querying is the policy's :attr:Policy.execution_horizon - the single source of truth - never actions_per_step read directly. How action_horizon interacts with it depends on whether the policy carries cross-chunk state (RTC):

  • RTC policy (supports_rtc is true): the policy hard-decides the interval and is re-queried at exactly its execution_horizon so it can blend the unexecuted tail of the previous chunk into the next one. A caller-supplied action_horizon must NOT stretch (or shrink) this interval - doing so leaves prev_chunk_left_over empty and silently degrades RTC to plain open-loop replay. action_horizon is ignored.
  • non-RTC (open-loop chunked or single-step): consume max(action_horizon, execution_horizon) so a model trained for N-step replay (execution_horizon == actions_per_step == N) keeps its FULL chunk - clamping to a smaller action_horizon drops the chunk tail and forces an out-of-distribution re-query. Single-action providers (MockPolicy) have execution_horizon == 1 so the result is just max(action_horizon, 1).

Before this helper existed each consumer inlined the same max(action_horizon, getattr(policy, "actions_per_step", 1)) expression and they drifted; worse, all of them keyed off actions_per_step, so an RTC policy was re-queried only after its full trained chunk drained and its cross-chunk blending never engaged.

Parameters:

Name Type Description Default
policy Policy

Any policy. The re-query interval is read from :attr:Policy.execution_horizon (falling back to a raw actions_per_step attribute for duck-typed objects that are not Policy subclasses); a policy that declares neither is treated as single-action.

required
action_horizon int

Consumer-requested actions per chunk (clamped to >= 1). Ignored for RTC policies, which decide their own interval.

required

Returns:

Type Description
int

The number of leading chunk actions to execute before re-querying.

align_action_values

align_action_values(values: Sequence[float] | ndarray, action_keys: Sequence[str], *, pad_short: bool = False) -> tuple[list[float], list[str]]

Pair a model's ordered action vector with the actuator keys it drives.

Every provider maps a policy's flat action vector onto actuator names BY INDEX, and the two lengths are not guaranteed to agree: a checkpoint trained for a 6-DOF arm can be pointed at a 7-actuator robot, or an embodiment can declare a gripper the checkpoint never learned. This centralizes the single rule for that mismatch so providers cannot drift.

  • More values than keys - the trailing values are dropped. There is no actuator to receive them.
  • Fewer values than keys (the default) - only the leading keys the model actually produced a value for are returned. The unmatched actuators are left out of the action dict entirely, so they receive no command and hold their current position.
  • Fewer values than keys with pad_short=True - the unmatched keys are returned carrying 0.0. That is a COMMAND, not an omission: where the action space is absolute position - a LeRobot <motor>.pos follower, a MuJoCo position actuator - 0.0 means "travel to zero", so those actuators MOVE, at whatever rate the servo will do it. Opt in only when the consumer needs a fixed-width action dict and zero is a meaningful target for every key it pads.

Parameters:

Name Type Description Default
values Sequence[float] | ndarray

The model's per-step action vector. Any sized, indexable numeric sequence - a list or a 1-D array, as the two providers hand over a NumPy row; entries are coerced with float.

required
action_keys Sequence[str]

Ordered actuator keys the vector maps onto, index 0 first.

required
pad_short bool

Emit 0.0 for keys past the end of values instead of omitting them. See the note above before enabling this.

False

Returns:

Type Description
list[float]

(values, keys), equal length and aligned 1:1 by index - ready to zip

list[str]

into an action dict after any unit conversion has been applied to the

tuple[list[float], list[str]]

values.

Factory

Policy factory - create_policy() and runtime registration.

UntrustedRemoteCodeError

UntrustedRemoteCodeError(message: str = '', *, code: str | None = None, subject: str | None = None)

Bases: RuntimeError

Raised when a HF model requires trust_remote_code but the user has not opted in.

Carries a stable machine-readable :attr:code (:data:~strands_robots.refusal_codes.TRUST_REMOTE_CODE_REQUIRED) and the :attr:subject provider, so a consumer offering the operator the opt-in classifies on identity instead of matching the message text. The message is unchanged by this. See :mod:strands_robots.refusal_codes.

Parameters:

Name Type Description Default
message str

The operator-facing reason, unchanged by the code.

''
code str | None

A member of :data:~strands_robots.refusal_codes.REFUSAL_CODES.

None
subject str | None

The policy provider the gate refused.

None

Attributes:

Name Type Description
code

The stable identifier for this refusal, or None.

subject

The policy provider the gate refused, or None.

create_policy

create_policy(provider: str, /, **kwargs) -> Policy

Create a policy instance.

Accepts either a provider name or a smart string:

  • Provider name: create_policy("lerobot_local", pretrained_name_or_path="lerobot/act_aloha_sim")
  • Server URL: create_policy("ws://gpu-box:8765")
  • Checkpoint: create_policy("lerobot/act_aloha_sim") or a path such as create_policy("outputs/train/act/checkpoints/last/pretrained_model")
  • Shorthand: create_policy("mock")

Any other spelling is a provider name; one carrying stray punctuation ("wbc/", "protomotions:") is refused as an unknown provider with the nearest names, not forwarded to lerobot_local as a checkpoint id.

All provider definitions live in registry/policies.json.

Parameters:

Name Type Description Default
provider str

Provider name, HF model ID, or server URL. Positional-only, so a provider keyword named provider reaches the policy's constructor (and its refusal) instead of colliding with this one.

required
**kwargs

Provider-specific parameters.

{}

Returns:

Type Description
Policy

Policy instance ready for get_actions().

Warns:

Type Description
DeprecationWarning

If provider is in _REMOVED_IN_0_7, naming its replacement.

Raises:

Type Description
TypeError

If provider is not a string (a pre-built policy goes to policy_object= instead), or if a keyword misspells one the provider's constructor binds, names one it cannot bind at all (no **kwargs), omits one it requires, or the provider needs its own provider argument (persistent, which is constructed directly) - see :func:policy_kwargs_error. Raised before construction and before the trust-remote-code gate, so no model is downloaded, no server dialled and no opt-in asked for on a typo.

UntrustedRemoteCodeError

If the provider loads HF models with trust_remote_code=True and STRANDS_TRUST_REMOTE_CODE is not set.

register_policy

register_policy(name: str, loader: Callable[[], type[Policy]], aliases: list[str] | None = None, *, overwrite: bool = False)

Register a custom policy provider at runtime.

Use this to add providers without editing policies.json.

Example::

from strands_robots.policies import register_policy

register_policy("my_provider", lambda: MyPolicy, aliases=["my"])
policy = create_policy("my_provider", ...)

Parameters:

Name Type Description Default
name str

Provider name :func:create_policy will accept. Surrounding whitespace is stripped before it is stored, as for aliases.

required
loader Callable[[], type[Policy]]

Zero-argument callable returning the :class:Policy subclass.

required
aliases list[str] | None

Extra spellings that resolve to name.

None
overwrite bool

Allow name or an alias to reuse a spelling a built-in provider in policies.json already answers to. A runtime entry wins over the built-in for the rest of the process.

False

Raises:

Type Description
TypeError

name or an alias is not a non-empty string, aliases is not a list of them, or loader is not callable. Nothing is registered.

ValueError

name or an alias is a built-in provider name or alias and overwrite is false. Nothing is registered.

list_providers

list_providers() -> list[str]

List the canonical policy provider names (JSON + runtime); aliases are in :func:list_aliases.

list_aliases

list_aliases() -> dict[str, str]

Return every provider alias and the canonical name it resolves to.

:func:create_policy accepts a provider's declared aliases and shorthands as readily as its canonical name, but :func:list_providers reports the canonical names from the JSON registry. Together the two surfaces enumerate every spelling the registries hold::

registered = set(list_providers()) | set(list_aliases())

That is every registered spelling, not every spelling :func:create_policy resolves. :func:import_policy_class falls back to auto-discovery, so a module under strands_robots.policies that exports a :class:~strands_robots.policies.base.Policy subclass resolves under its own module name with no registry entry. Two ship, and neither is a registry provider because each wraps a policy the caller already holds rather than building one from config:

  • composite (:class:~strands_robots.policies.composite.CompositePolicy) builds through this factory -- create_policy("composite", lower=..., upper=...) -- and is the one spelling registered above omits.
  • persistent (:class:~strands_robots.policies.persistent.PersistentPolicy) resolves but cannot be built here: its first parameter is named provider, which :func:create_policy has already bound, so it is constructed directly. :func:create_policy refuses it with a TypeError that says so.

Covers both registries, matching the union :func:list_providers reports: aliases declared in policies.json and aliases passed to :func:register_policy at runtime. A runtime alias shadows a JSON alias of the same name, which is the precedence :func:create_policy applies.

Returns:

Type Description
dict[str, str]

Mapping of alias to the canonical provider name it resolves to.

import_policy_class

import_policy_class(provider: str) -> type

Dynamically import and return the Policy class for a provider.

Uses the module + class paths from policies.json. Falls back to auto-discovery (strands_robots.policies.) if not in JSON.

Parameters:

Name Type Description Default
provider str

Canonical provider name.

required

Returns:

Type Description
type

The Policy subclass.

Raises:

Type Description
ValueError

If the provider does not exist, or was removed - a removed spelling (groot) is refused with the sentence :data:~strands_robots.registry.policies.REMOVED_PROVIDERS holds for it, never rerouted to another provider.

ImportError

If the provider exists but its module cannot be imported, naming the provider, the missing module and the remedy (see :func:_provider_import_error). A provider whose module is present but whose optional dependency is missing reports that rather than being misreported as an unknown provider.

preflight_policy

preflight_policy(provider: str, observation_keys: set[str], **kwargs) -> None

Run a provider's class-level :meth:Policy.preflight check, if any.

Resolves provider to its policy class WITHOUT instantiating it (so no model weights are downloaded) and invokes the class's preflight hook with the runtime observation_keys and the provider kwargs. Providers that do not override :meth:Policy.preflight are a no-op.

This is the fail-fast seam used by SimEngine.run_policy / eval_policy to catch a misconfiguration (e.g. sim camera names that cannot be routed to the model's declared image inputs) BEFORE the expensive create_policy download, instead of crashing deep inside the first inference. Resolution failures are swallowed (the matching error is surfaced authoritatively by the subsequent create_policy); only the provider's own preflight ValueError propagates.

Parameters:

Name Type Description Default
provider str

Provider name, HF model ID, or server URL (as passed to create_policy).

required
observation_keys set[str]

Keys the runtime observation will contain (joint names + camera names).

required
**kwargs

Provider-specific parameters (the policy_config).

{}

Raises:

Type Description
TypeError

When provider is not a string - a caller bug, not a resolution failure, so it is not swallowed.

ValueError

When the resolved provider's preflight rejects the configuration.

preflight_reason

preflight_reason(provider: str, read_observation_keys: Callable[[], Iterable[str]], /, **kwargs: Any) -> str | None

Why provider refuses this configuration, or None.

The whole pre-build check in one call, so the three entry points that owe it - the simulation engine, the physical arm and a native driver's task verb - read one rule instead of keeping three copies of it in step:

  • the observation is read only when the resolved class actually overrides :meth:Policy.preflight (:func:policy_overrides_preflight). That read is not cheap - the sim renders every camera in the scene, an arm warms and grabs a frame from each configured camera, a driver crosses the wire - and for every shipped provider but lerobot_local the result is gathered only to be discarded.
  • a read that fails, or answers nothing, is not a verdict on the policy configuration and does not become one here: the check is skipped and that read stays the caller's own to report.
  • the provider's ValueError comes back as text, because two of the three callers answer a refusal envelope rather than raise.

Parameters:

Name Type Description Default
provider str

Provider name, HF model ID, or server URL (as passed to :func:create_policy). Positional-only, as is the reader below, so a policy kwarg spelled either way reaches the hook instead of binding here.

required
read_observation_keys Callable[[], Iterable[str]]

Answers the keys the runtime observation will carry (joint names plus camera names). Called at most once, and only when there is a hook to feed.

required
**kwargs Any

Provider-specific parameters (the policy_config), judged as the mapping :func:create_policy will be given.

{}

Returns:

Type Description
str | None

The provider's refusal text, or None when the configuration passes,

str | None

when there is no hook to run, or when the observation could not be read.

Raises:

Type Description
TypeError

When provider is not a string, from :func:policy_overrides_preflight.

Built-in policies

strands_robots.policies.mock.MockPolicy

MockPolicy(amplitude: float = 0.5, seed: int | None = None, **kwargs: Any)

Bases: Policy

Mock policy for testing - generates smooth sinusoidal trajectories.

Parameters:

Name Type Description Default
amplitude float

Peak of the per-joint sinusoid in radians (default 0.5), before the actuator range learnt in :meth:set_sim_context clips it. Must be a finite number of at least 0.

0.5
seed int | None

Accepted for parity with the other providers' policy_config bags and recorded on the instance; the sinusoid is deterministic, so it changes nothing.

None

The constructor declares its keywords, so the factory's near-miss screen applies to the mock as to every other provider: policy_config={"amplitud": 0.5} is refused with a did-you-mean instead of running on the default (#4165), and a bag passed whole as create_policy("mock", policy_config={...}) is refused naming the unpack; docs/learn/policies/index.md promises a misspelled keyword is refused before anything runs. The **kwargs sink stays because the hardware drivers hand every provider the server address (host, and port when one is given) whether or not it dials one; a name that is not close to a declared keyword is forwarded there and logged at DEBUG, which is the pass-through contract every provider with a sink has.

Raises:

Type Description
ValueError

If amplitude is not a finite number, or is negative.

provider_name property

provider_name: str

Provider name for identification (always "mock").

requires_images property

requires_images: bool

Mock policy only consumes joint state - skip camera rendering.

get_actions async

get_actions(observation_dict: dict[str, Any], instruction: str, **kwargs: Any) -> list[dict[str, Any]]

Return smooth sinusoidal actions.

Canonical reference for the per-tick action value convention documented on :meth:Policy.get_actions: every value is a python float (single-DOF joint target), never a raw np.ndarray.

reset

reset(seed: int | None = None) -> None

Rewind the sinusoid to its first step.

The mock is deterministic but not stateless: _step advances by one chunk per :meth:get_actions, and it is the only per-episode state the mock holds. Left where the previous episode ended, two episodes seeded alike began at different phases of the sinusoid, so the reference implementation of the ABC broke the reproducibility its own :meth:Policy.reset docstring asks providers to keep. The seed is not read: the trajectory has no random draw for it to reach.

set_robot_state_keys

set_robot_state_keys(robot_state_keys: list[str]) -> None

Record the ordered joint keys used to name the sinusoidal action dict.

Raises:

Type Description
ValueError

If robot_state_keys is not an ordered list of distinct non-blank names, per :func:~strands_robots.utils.name_list_error. A single name passed as a bare string is the mistake this catches: str is iterable per character, so it would bind one joint per letter.

set_sim_context

set_sim_context(model: Any, namespace: str) -> None

Learn the range each driven actuator is held to, so the sinusoid stays inside it.

Called by the MuJoCo engine's bind_policy_sim_context right after :meth:set_robot_state_keys, with the compiled MjModel and the robot's namespace prefix ("so100/"). The mock's ±0.5 rad sinusoid was written for a generic joint; on a real model some actuators do not span it - the SO-100 Pitch ctrlrange is [-3.32, 0.174] and its Jaw is [-0.174, 1.75] - so the value was held to the range and the engine warned that the commanded trajectory was NOT reproduced. That warning was the first thing examples/01_sim_hello_world.py printed. Knowing the ranges, the mock clips its own output so what it commands is what the actuator does.

Which range holds the command is :func:~strands_robots.simulation.mujoco.scene_ops.effective_ctrl_range's rule, read here rather than re-derived: an actuator whose MJCF authors neither ctrlrange nor inheritrange compiles to ctrlrange == (0, 0) with actuator_ctrllimited == 0, and for a position servo its ctrl IS the joint target, so the driven joint's limits bound the pose. Reading only actuator_ctrllimited therefore learned nothing at all on the so101 - all six of its actuators are in that second case - and the mock commanded -0.433 to a jaw whose joint range is [-0.1745, 1.745], which is the warning this method exists to prevent.

Actuators that are held to no range and names that resolve to no actuator are left alone; an error while reading the model leaves the policy exactly as configured, never fails the rollout.

strands_robots.policies.composite.CompositePolicy

CompositePolicy(lower: Policy, upper: Policy, *, lower_joints: Sequence[str] | None = None, upper_joints: Sequence[str] | None = None, lower_obs_keys: Sequence[str] | None = None, upper_obs_keys: Sequence[str] | None = None)

Bases: Policy

Compose a lower and an upper policy on one robot's joint set.

Each :meth:get_actions queries both children with the (optionally per-child filtered) observation and the same instruction + kwargs, then merges their per-tick action dicts by joint name:

  • lower contributes the names in lower_joints (or all names it emits when lower_joints is None).
  • upper contributes the names in upper_joints (or, when None, every name it emits that the lower policy did not already claim - lower precedence).

An explicit group is EXCLUSIVE, either way round: the policy it names is the only one allowed to command those joints. A defaulted lower_joints may not command into an explicit upper_joints, and a defaulted upper_joints may not command into an explicit lower_joints - that tick is refused, on every tick alike, rather than resolved by whichever child emitted the name. Precedence only decides between two DEFAULTED groups, where the caller declared no owner.

Routing that discards a child's ENTIRE action dict raises: the composite would otherwise silently be the surviving child alone. Two children that drive the same joint set cannot be composed, only cascaded (module docstring). An observation subset that selects NOTHING raises for the mirror reason: a child queried with an empty dict acts on no reading at all.

The merged chunk length is the shorter of the two children's chunks, so the more frequently re-querying child sets the re-query cadence (:attr:execution_horizon). This keeps a per-tick controller (WBC, execution_horizon == 1) closed-loop even when paired with a chunk-emitting manipulation policy.

Parameters:

Name Type Description Default
lower Policy

Policy driving the lower joint group (e.g. legs+waist locomotion).

required
upper Policy

Policy driving the upper joint group (e.g. arms manipulation).

required
lower_joints Sequence[str] | None

Joint/actuator names the lower policy is authoritative for. Exclusive: no other child may command these, whichever names the lower policy emits on a given tick. None (default) accepts every name the lower policy emits, minus any that upper_joints assigns to the upper policy - commanding into an explicit group is refused, not silently resolved.

None
upper_joints Sequence[str] | None

Joint/actuator names the upper policy is authoritative for. Exclusive: no other child may command these, whichever names the upper policy emits on a given tick. None (default) accepts every upper name not already owned by the lower policy.

None
lower_obs_keys Sequence[str] | None

Observation keys to forward to the lower policy. None (default) forwards the full observation (children read by name). A subset that shares no key with the observation is refused rather than forwarded as an empty dict - the names belong to whatever produces the observation, so a subset written in the wrong namespace would otherwise starve the child on every tick.

None
upper_obs_keys Sequence[str] | None

Observation keys to forward to the upper policy. None (default) forwards the full observation. Held to the same rule as lower_obs_keys.

None

Raises:

Type Description
ValueError

If lower or upper is None, or if lower_joints and upper_joints are both given and share a name (ambiguous ownership).

children property

children: tuple[Policy, ...]

Both child policies, lower first - the order they are merged in.

Lets a runtime capability probe reach the concrete policies inside the composite; see :attr:Policy.children.

execution_horizon property

execution_horizon: int

Re-query interval: the shorter of the two children's horizons.

The merged chunk is truncated to the shorter child's length, so the consumer must re-query at the faster child's cadence to keep that child closed-loop (a per-tick locomotion controller must not be starved by a slower chunk-emitting manipulation policy).

lower property

lower: Policy

The lower-body (e.g. locomotion) child policy.

provider_name property

provider_name: str

Provider name for identification (always "composite").

requires_images property

requires_images: bool

True if EITHER child consumes camera frames.

The composite cannot skip rendering unless both children opt out; a manipulation upper body typically needs images even when the locomotion lower body does not.

upper property

upper: Policy

The upper-body (e.g. manipulation) child policy.

get_actions async

get_actions(observation_dict: dict[str, Any], instruction: str, **kwargs: Any) -> list[dict[str, Any]]

Query both children and merge their per-tick action dicts by joint name.

Both children receive the same instruction and kwargs (each ignores keys it does not use, per the :class:Policy contract), and the observation filtered to its configured key subset. The two action chunks are merged element-wise up to the shorter length.

Returns:

Type Description
list[dict[str, Any]]

The merged action chunk (length == the shorter child's chunk). Each

list[dict[str, Any]]

dict maps a joint/actuator name to its target value, routed from the

list[dict[str, Any]]

child that owns that name.

Raises:

Type Description
ValueError

If a configured observation subset shares no key with the observation (the child would be queried blind), if either child returns an empty chunk, if routing discards a child's entire action dict (the composite would be the other child alone), or if either child commands a joint the other child's explicit joint group assigns to it.

reset

reset(seed: int | None = None) -> None

Reset per-episode state on both children.

set_control_frequency

set_control_frequency(hz: float) -> None

Set the control rate on the composite and forward it to both children.

set_robot_state_keys

set_robot_state_keys(robot_state_keys: list[str]) -> None

Forward the robot's state-key list to both children.

set_rtc_observed_delay

set_rtc_observed_delay(steps: int | None) -> None

Forward the RTC observed-delay step count to the composite and both children.

Persistent cache

Persistent, reusable policy handles - load the model once, reuse everywhere.

Loading a VLA / LeRobot checkpoint (a MolmoAct2 SO-100/101 build reads ~1300 weight files into GPU memory) costs on the order of a minute or two. A naive multi-episode loop that calls :func:create_policy per rollout pays that cost every episode; the dominant fix already exists - a process-level model cache in :mod:strands_robots.policies.lerobot_local.policy shares the resident weights across instances - but two ergonomic gaps remained:

  1. The win was implicit. There was no first-class "load this once and hand me a handle I reuse" object, so an LLM harness driving the API blind had no obvious way to express the intent and would re-call create_policy.
  2. There was no provider-agnostic way to warm the cache ahead of a run, see what is resident, or free a checkpoint between runs of different policies.

This module closes both. :class:PersistentPolicy is a thin, thread-safe wrapper that builds the underlying policy ONCE at construction and is meant to be passed to every run_policy/eval_policy call via policy_object=::

from strands_robots.policies import PersistentPolicy

policy = PersistentPolicy("lerobot_local", pretrained_name_or_path="...")
for _ in range(20):
    sim.run_policy(robot_name="arm", policy_object=policy)  # zero reload
    sim.save_episode()
    sim.reset()

The :func:preload, :func:list_cached, and :func:evict helpers are the agent-facing cache controls: warm before a run, introspect what is hot, free memory before switching checkpoints.

This is a SYNCHRONOUS persistent worker: the model lives in-process and is shared via the module-level cache. It deliberately does not spawn a background daemon or expose cross-process IPC - inference is GIL- and GPU-serialised, so a per-call lock gives correct concurrent reuse without the complexity (and races) of a separate worker process. Cross-process sharing is a separable concern.

PersistentPolicy

PersistentPolicy(provider: str, *, policy_object: Policy | None = None, **config: Any)

Bases: Policy

A persistent, reusable handle around an underlying policy.

Builds the wrapped policy ONCE (warming the process-level model cache) and delegates every :class:Policy operation to it, so the same object can be passed to many run_policy/eval_policy calls without ever reloading weights. Inference calls are serialised by a per-call lock, so two threads sharing one handle never corrupt the wrapped model's per-episode state.

The wrapper is transparent: chunk-shape introspection (execution_horizon, is_chunk_emitting, actions_per_step, supports_rtc), RTC delay / control-frequency hooks, reset, and load telemetry (load_time_s, load_cache_hit) all forward to the wrapped policy, so the runtime drives it exactly as it would the bare policy.

Parameters:

Name Type Description Default
provider str

Provider name or smart string forwarded to :func:create_policy (e.g. "lerobot_local", "mock").

required
policy_object Policy | None

An already-constructed policy to wrap instead of building a new one. When given, both provider and **config are ignored; :attr:provider_name reports the wrapped policy's own provider, which is what identifies the resident model.

None
**config Any

Provider-specific keyword arguments forwarded to :func:create_policy.

{}

children property

children: tuple[Policy, ...]

The single wrapped policy.

Keeps the wrapper transparent to a runtime capability probe as well as to the delegated :class:Policy operations; see :attr:Policy.children.

control_frequency property writable

control_frequency: float | None

The control rate the wrapped policy was told (None until set).

execution_horizon property

execution_horizon: int

Actions consumed from one get_actions chunk, from the wrapped policy.

inner property

inner: Policy

The wrapped policy instance.

provider_name property

provider_name: str

Provider name of the wrapped policy (identifies the resident model).

requires_images property

requires_images: bool

Whether the wrapped policy needs camera frames in its observation.

rtc_observed_delay_steps property writable

rtc_observed_delay_steps: int | None

The RTC delay the wrapped policy was told (None until set).

get_actions async

get_actions(observation_dict: dict[str, Any], instruction: str, **kwargs: Any) -> list[dict[str, Any]]

Delegate to the wrapped policy under the shared thread lock (see :meth:Policy.get_actions).

Acquires the same threading.Lock the sync path uses, so the sync and async entry points mutually exclude on the wrapped model's per-episode state. The blocking acquire runs in the loop's default executor so a shared running loop is never frozen while another caller holds the lock (each runtime call still runs on its own per-call asyncio.run loop).

Cancellation-safe: if this coroutine is cancelled (e.g. by asyncio.wait_for) while the acquire is still pending, the lock is released as soon as the executor thread obtains it, so a cancelled call never poisons the shared handle. See :class:_LockHandoff.

get_actions_sync

get_actions_sync(observation_dict: dict[str, Any], instruction: str, **kwargs: Any) -> list[dict[str, Any]]

Delegate to the wrapped policy under the per-call lock (see :meth:Policy.get_actions_sync).

is_chunk_emitting

is_chunk_emitting() -> bool

Whether the wrapped policy returns multi-action chunks per get_actions.

reset

reset(seed: int | None = None) -> None

Reset the wrapped policy's per-episode state (weights stay resident).

set_control_frequency

set_control_frequency(hz: float) -> None

Forward the executing loop's control rate to the wrapped policy.

set_robot_state_keys

set_robot_state_keys(robot_state_keys: list[str]) -> None

Forward the robot state keys to the wrapped policy.

set_rtc_observed_delay

set_rtc_observed_delay(steps: int | None) -> None

Forward the observed inference delay (RTC steps) to the wrapped policy.

preload

preload(provider: str, **config: Any) -> dict[str, Any]

Warm the model cache for a provider and report the load cost.

Builds a :class:PersistentPolicy (which loads the model once into the process-level cache) and measures the wall time and resident-memory delta. Call this before a multi-episode run so every subsequent run_policy with the returned policy is a zero-reload cache hit.

Parameters:

Name Type Description Default
provider str

Provider name or smart string (see :func:create_policy).

required
**config Any

Provider-specific keyword arguments.

{}

Returns:

Type Description
dict[str, Any]

A dict with: policy: the ready :class:PersistentPolicy to pass as policy_object= to run_policy/eval_policy. load_time_s: wall-clock seconds spent constructing the policy. load_cache_hit: whether the heavy weight load was served from the process-level cache (a warm preload). resident_rss_mb: process RSS in MB after the load (None if unmeasurable on this platform). rss_delta_mb: change in process RSS across the load (None if unmeasurable); ~0 on a cache hit, large on a cold load.

list_cached

list_cached() -> list[dict[str, Any]]

List the heavy models currently resident in the process-level cache.

Provider-agnostic introspection for deciding whether to evict before loading a different checkpoint. Currently the lerobot_local provider is the only one with a process-level weight cache; the list is empty when its optional dependencies are not installed.

Returns:

Type Description
list[dict[str, Any]]

One dict per cached entry (see

list[dict[str, Any]]

func:strands_robots.policies.lerobot_local.list_cached_models), or an

list[dict[str, Any]]

empty list when no cache is available.

evict

evict(pretrained_name_or_path: str | None = None) -> int

Free cached models, returning held GPU/CPU memory.

Parameters:

Name Type Description Default
pretrained_name_or_path str | None

When None (default), evict every cached model. When set, evict only the entries loaded from that checkpoint - free one policy before switching to another without dropping the rest.

None

Returns:

Type Description
int

Number of cache entries evicted (0 when no cache is available).

LeRobot policy types

strands_robots.policies.lerobot_local.resolution.list_policy_types

list_policy_types() -> list[str]

List the LeRobot policy type strings resolvable in this environment.

These are exactly the values accepted as the policy_type argument to :func:resolve_policy_class_by_name -- and therefore as create_policy("lerobot_local", policy_type=...). The list reflects the installed lerobot: it is sourced from lerobot's own draccus choice registry (PreTrainedConfig.get_known_choices()) after every policy config module has been imported via :func:_ensure_policy_configs_registered, so a brand-new policy a newer lerobot ships appears automatically and a slimmer install reports fewer.

This is the discovery surface for the lerobot_local provider: rather than reading lerobot internals to learn which policy_type strings are valid, a caller (or an agent) can enumerate them here.

Returns:

Type Description
list[str]

Sorted list of policy type strings (e.g. ``["act", "diffusion",

list[str]

"smolvla", ...]``). Empty when lerobot is not installed -- a discovery

list[str]

surface yields an empty list on a missing dependency rather than

list[str]

raising.

Edit page