Policies¶
The Policy contract, how a provider string resolves through create_policy, and the cache that keeps a loaded model.
A Policy turns an observation into actions. Providers live in registry/policies.json and resolve through create_policy: the Policy contract, how a provider string resolves, the persistent cache that keeps a loaded model between calls.
Contract¶
Abstract base class for robot policies (VLA, motion planners, MPC, scripted).
The :class:Policy ABC is intentionally agnostic about how actions are
produced. Built-in providers (mock, lerobot_local, remote) are VLA-style,
but the same interface is the right shape for:
- Classical motion planners - cuRobo, MoveIt2, OMPL, RRT*: take a goal pose and joint state, return a collision-free trajectory.
- Model-predictive controllers (MPC) - solve a finite-horizon optimal control problem each tick.
- Scripted / pure-IK trajectories - analytic IK followed by interpolation; zero learning involved.
Non-VLA implementations typically set :attr:Policy.requires_images to
False to skip camera rendering (~10x throughput win at 500Hz) and read
their goal from the well-known **kwargs keys documented on
:meth:Policy.get_actions rather than parsing the natural-language
instruction string.
See :class:~strands_robots.policies.mock.MockPolicy for the canonical
non-VLA reference implementation.
Policy ¶
Bases: ABC
Abstract base class for robot policies (VLA, motion planners, MPC, scripted).
All policies implement async :meth:get_actions. For convenience, a
synchronous wrapper :meth:get_actions_sync is provided.
The interface is general enough to cover both VLA-style providers
(consume images + instruction, output joint targets) and non-VLA
providers such as classical motion planners (cuRobo, MoveIt2),
model-predictive controllers, and pure-IK / scripted trajectories.
Non-VLA providers typically set :attr:requires_images to False
and read their goal from the well-known **kwargs keys documented
on :meth:get_actions.
All providers MUST honour the per-tick action value convention
documented on :meth:get_actions: each action value is a python
float (single-DOF) or list[float] (multi-DOF group), never a
raw np.ndarray, so downstream consumers handle every provider's
output uniformly regardless of its internal compute backend. See
MockPolicy for the canonical reference.
children
property
¶
The policies this one delegates to, in the order it consults them.
Default () - a leaf policy that runs its own inference. A wrapper
returns the policies it drives:
:class:~strands_robots.policies.composite.CompositePolicy its lower
and upper children,
:class:~strands_robots.policies.persistent.PersistentPolicy the single
policy it holds warm.
This is the same "policy declares, runtime supplies" contract as
:attr:requires_images and :attr:required_bodies, applied to a
capability probe rather than an observation. A probe answers about the
object it is handed, and a wrapper is a different object than the policy
inside it, so an isinstance test against a wrapper reports the
wrapped policy's capability as absent. The MuJoCo backend's WBC torque
shim is the motivating case: it is required for a
:class:~strands_robots.policies.wbc.WBCPolicy to hold a stable gait on
a position-servo scene, and the physics does not change when that policy
is wrapped - only the type of the object the probe sees does. Declaring
the children lets one probe walk to the policy that answers, instead of
every probe having to learn the name of every wrapper.
Returns:
| Type | Description |
|---|---|
tuple[Policy, ...]
|
The child policies. Empty (the default) means this policy is a leaf. |
execution_horizon
property
¶
Number of actions the SIM consumes from one get_actions chunk before re-querying.
This is the SINGLE source of truth for the re-query interval; a chunk
consumer (the single-policy runner, the multi-episode eval loop, the
synchronized multi-robot loop) reads it via
:func:resolve_chunk_length and never inspects actions_per_step
directly. Distinguishing the re-query interval from the trained chunk
length is what makes Real-Time Chunking (RTC) actually engage:
- RTC policy -> the RTC
execution_horizon(typically << the trained chunk). The policy is re-queried mid-chunk so it can blend the unexecuted tail of the previous chunk (prev_chunk_left_over) into the next one. Re-querying only after the full trained chunk drains leaves that tail permanently empty and silently degrades RTC to plain open-loop replay. - chunked open-loop (ACT, diffusion, pi0/SmolVLA without RTC) ->
actions_per_step(the trained chunk; truncating drops its tail and forces an out-of-distribution re-query). - single-step (
MockPolicy, classical planners) ->1.
The default derives from actions_per_step (1 when undeclared),
so a single-step or chunked open-loop policy needs no override; only a
policy with an inference-time budget distinct from its trained chunk
(RTC) overrides this.
required_bodies
property
¶
Named rigid bodies whose world pose this policy needs in its observation.
Default () - most policies are driven by joint state alone and pay
nothing for this. A whole-body motion-mimic tracker (ProtoMotions
GTP, PHC, OmniH2O and the text-to-motion pipelines built on them) is the
motivating case: its network consumes the world orientation of a single
anchor link - torso_link on a Unitree G1 - which is NOT derivable
from the observation's floating-base signals. base_quat is the
pelvis, and the torso differs from it by the three waist joints, so a
tracker written against base_quat silently feeds the network the
wrong frame whenever the waist is not neutral.
Declaring the bodies here is the same "policy declares, runtime
supplies" contract as :attr:requires_images: the runtime
(:class:~strands_robots.simulation.policy_runner.PolicyRunner)
resolves the names ONCE before the rollout and merges the pose of each
into every observation it hands to :meth:get_actions, under the keys
documented on
:meth:~strands_robots.simulation.base.SimEngine.get_observation::
body.<name>.pos # world x, y, z (m)
body.<name>.quat # world orientation w, x, y, z
body.<name>.lin_vel # world linear velocity x, y, z (m/s)
body.<name>.ang_vel # world angular velocity x, y, z (rad/s)
A policy that declares a body the scene does not contain fails at the start of the rollout with the available body names, rather than reading a missing key as a zero pose on every tick.
Every surface that reads this collects it over the whole policy tree
through :func:collect_required_bodies, so a policy that declares a
body is honoured when it is wrapped: a wrapper which does not override
this property does not hide its child's declaration. That one owner is
shared with the remote-inference handshake, so a policy served over the
wire declares the same bodies it does in-process.
Returns:
| Type | Description |
|---|---|
str
|
Ordered, de-duplicated body names. Empty (the default) means the |
...
|
observation is left exactly as the backend produced it. |
requires_images
property
¶
Whether this policy needs camera frames in its observation.
Default True (most VLA policies do). Subclasses that only
consume joint state (e.g. MockPolicy, classical motion planners
such as cuRobo / MoveIt2, MPC, pure-IK controllers, scripted
trajectories) can return False to let the simulation skip
expensive camera rendering - a ~10x throughput win at 500Hz when
no cameras are needed.
get_actions
abstractmethod
async
¶
get_actions(observation_dict: dict[str, Any], instruction: str, **kwargs: Any) -> list[dict[str, Any]]
Get actions from policy given observation and instruction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
observation_dict
|
dict[str, Any]
|
Robot observation (cameras + state). VLA
providers consume both |
required |
instruction
|
str
|
Natural language instruction. Required by the
signature for VLA providers; non-VLA providers (motion
planners, MPC, scripted) may ignore it and read the goal
from |
required |
**kwargs
|
Any
|
Provider-specific parameters. The following keys
are well-known and SHOULD be honoured by non-VLA
providers when present so callers don't have to JSON-encode
goals into the
Providers MUST ignore unknown |
{}
|
Returns:
| Type | Description |
|---|---|
list[dict[str, Any]]
|
List of action dicts for robot execution. Each dict maps a |
list[dict[str, Any]]
|
robot state key (joint/actuator name) to its target value |
list[dict[str, Any]]
|
for that tick. |
list[dict[str, Any]]
|
Values MUST be JSON / python-native: a python |
list[dict[str, Any]]
|
a single-DOF actuator, or a |
list[dict[str, Any]]
|
actuator group. Implementations MUST NOT return raw |
list[dict[str, Any]]
|
|
list[dict[str, Any]]
|
|
list[dict[str, Any]]
|
treat every provider's output uniformly (e.g. |
list[dict[str, Any]]
|
scalar, |
list[dict[str, Any]]
|
internal compute backend. |
list[dict[str, Any]]
|
The list length is the action-chunk horizon; consumers execute |
list[dict[str, Any]]
|
it at a fixed control rate (e.g. 50Hz). |
get_actions_sync ¶
get_actions_sync(observation_dict: dict[str, Any], instruction: str, **kwargs: Any) -> list[dict[str, Any]]
Synchronous convenience wrapper around :meth:get_actions.
Safe to call from sync code, event loops, or notebooks. Resolution is
delegated to :mod:strands_robots._async_utils, the one owner of
"resolve a policy coroutine in a sync context": with no running loop it
is asyncio.run, and inside one the coroutine is offloaded to that
module's single reused worker thread.
Delegating rather than re-deriving is what keeps this callable at control rate from a running loop -- a notebook cell and any async host take the offload branch, and constructing a private executor there starts and joins one OS thread per call. Every in-process rollout path already resolves through the same owner, so the wrapper and the runner cannot drift on which branch a caller lands in.
is_chunk_emitting ¶
Whether this policy returns multi-action chunks per get_actions.
A chunk-emitting policy (ACT, diffusion, pi0, pi0.5, pi0-FAST, SmolVLA,
MolmoAct2) returns more than one action per inference, so its inference
latency can be hidden behind the EXECUTION of the current chunk while the
next chunk is computed in the background. :meth:PolicyRunner.run
auto-enables that overlap (run_policy(async_rtc=None)) only when this
is true AND the policy blends the seam (supports_rtc); single-step
policies (MockPolicy, classical planners) gain nothing from overlap
and stay on the synchronous loop.
The default derives the answer from the re-query interval the consumer
actually drives - :attr:execution_horizon - so ANY policy that emits a
chunk longer than one action is detected without enumerating provider
names: a model under RTC reports its RTC horizon (> 1), a chunked
open-loop model reports its trained chunk length (> 1), and a single-step
policy reports 1. Providers whose chunk shape is not visible through
execution_horizon (e.g. a model that must be driven via
predict_action_chunk) override this.
:class:~strands_robots.policies.mock.MockPolicy returns eight actions
per call and still declares 1 on purpose: its sinusoid is a function
of a step counter, so a re-query continues the same curve wherever it
happens, there is no inference latency for the async pipeline to hide,
and 1 keeps the reference policy on the synchronous loop every
tutorial and test reads. A provider that pays for its chunk (Cosmos 3,
every LeRobot checkpoint) declares actions_per_step instead.
Returns:
| Type | Description |
|---|---|
bool
|
|
bool
|
single-step policies. |
preflight
classmethod
¶
Cheap pre-construction validation hook (no download, no instantiation).
Called by the simulation's run_policy / eval_policy BEFORE
:func:~strands_robots.policies.create_policy builds the policy - and
therefore before any model weight download - with the set of
observation keys the runtime will feed the policy. Override this to
fail fast on a misconfiguration (e.g. sim camera names that cannot be
routed to the model's declared image inputs) instead of surfacing it as
a confusing failure deep inside inference after a multi-minute weight
download.
The default implementation is a no-op. Implementations MUST be cheap:
no network access, no model instantiation - only local metadata
(policy_config plus packaged JSON such as the embodiment registry)
and the provided observation_keys.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
observation_keys
|
set[str]
|
Keys present in the runtime observation dict the
policy will receive (joint state names + attached camera
names), as returned by |
required |
**policy_config
|
Any
|
The same provider kwargs that will be forwarded to
the policy constructor by |
{}
|
Raises:
| Type | Description |
|---|---|
ValueError
|
When the configuration cannot consume the runtime observation (e.g. a required camera source key is absent and no override maps an available key onto the model's image feature). |
reset ¶
Reset per-episode policy state.
Default implementation is a no-op. Policies that hold per-episode state (e.g. diffusion sampler RNG, action chunk caches, KV-caches) should override to apply the reset.
For SERVICE-mode policies (e.g. Cosmos3Policy(host=...) over
WebSocket), the override forwards the call to the server so its
per-episode RNG state can be re-initialised - without this,
set_eval_seed only seeds the client-side process, leaving
the server's diffusion sampler RNG drifting across calls and
breaking reproducibility (#187).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
seed
|
int | None
|
Optional master seed forwarded to the policy's
random-number generators. When |
None
|
set_control_frequency ¶
Tell the policy the control rate (Hz) of the executing loop.
The runtime that drives the policy (PolicyRunner.run / evaluate)
calls this once before the rollout loop so providers that estimate an
inference delay in action steps (Real-Time Chunking) can convert their
measured wall-clock latency into the correct number of steps. Without
it, such providers fall back to a hardcoded rate and silently mis-blend
chunks at any other control frequency.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
hz
|
float
|
Finite positive control frequency in Hz. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
set_robot_state_keys
abstractmethod
¶
Configure the policy with robot state keys.
These are the ordered joint/motor names the policy emits as its
action-dict keys, so they decide which actuator each action value is
sent to. An implementation must refuse a malformed list rather than
bind it. Most do so through the shared domain
:func:~strands_robots.utils.name_list_error, gated on a truthy value
because an empty list already means "auto-detect" on the providers that
support it. :class:~strands_robots.policies.wbc.policy.WBCPolicy is
already total without it: it resolves every joint it drives BY NAME
inside the caller's list, so any malformed shape fails that membership
check instead - and it deliberately tolerates a repeated name, which
resolves to its first occurrence.
Unlike :meth:set_control_frequency and
:meth:set_rtc_observed_delay, this setter has no shared
implementation to carry the domain: each provider binds the names into
its own layout, so each refuses at its own entry. That parity is pinned
structurally by the policy state-key name-list contract tests.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
robot_state_keys
|
list[str]
|
Ordered list of distinct non-blank joint/motor names. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
set_rtc_observed_delay ¶
Tell the policy how many control steps elapse during inference.
The runtime that drives the policy calls this before each
get_actions so Real-Time Chunking providers can compute the
chunk-seam offset deterministically instead of deriving it from
wall-clock latency. In a synchronous eval loop the world is paused
during inference, so exactly 0 steps elapse; in the async overlap
pipeline the count is the number of still-pending steps of the chunk
being executed. Either way it is a known integer, not a measurement.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
steps
|
int | None
|
Non-negative control-step count, or |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
ChunkedPolicy ¶
Bases: Protocol
Introspection contract for policies that emit ACTION CHUNKS.
A chunked policy returns more than one action per
:meth:Policy.get_actions call: a model trained for N-step open-loop replay
(ACT, diffusion, pi0, SmolVLA, MolmoAct2) emits a length-N chunk that a
consumer executes before re-querying. The chunk PRODUCER is the existing
async :meth:Policy.get_actions - this protocol deliberately does NOT add a
second chunk-producing method (that would split one contract across two
code paths); it only surfaces the metadata a consumer needs to drive an
already-produced chunk correctly.
Every consumer of a chunk (the single-policy runner, the multi-episode eval
loop, and the synchronized multi-robot loop) must size the chunk the same
way - see :func:resolve_chunk_length. Routing all of them through one
helper that reads this contract keeps a chunk-emitting policy from being
truncated differently depending on which loop happens to drive it.
The protocol is runtime_checkable so a consumer can branch on
isinstance(policy, ChunkedPolicy) and a type checker rejects a
non-chunked policy where a chunked one is required.
Attributes:
| Name | Type | Description |
|---|---|---|
actions_per_step |
int
|
Number of actions the policy intends a consumer to
execute open-loop from one |
supports_rtc |
bool
|
Whether the policy blends chunk seams internally via
Real-Time Chunking - it carries prev-chunk state across re-queries
so consecutive chunks join smoothly. Introspection only; a consumer
never has to drive RTC, the policy does it inside |
resolve_chunk_length ¶
Effective number of actions to consume from one get_actions chunk.
Centralizes the single re-query rule every consumer must apply identically.
The number of actions consumed before re-querying is the policy's
:attr:Policy.execution_horizon - the single source of truth - never
actions_per_step read directly. How action_horizon interacts with it
depends on whether the policy carries cross-chunk state (RTC):
- RTC policy (
supports_rtcis true): the policy hard-decides the interval and is re-queried at exactly itsexecution_horizonso it can blend the unexecuted tail of the previous chunk into the next one. A caller-suppliedaction_horizonmust NOT stretch (or shrink) this interval - doing so leavesprev_chunk_left_overempty and silently degrades RTC to plain open-loop replay.action_horizonis ignored. - non-RTC (open-loop chunked or single-step): consume
max(action_horizon, execution_horizon)so a model trained for N-step replay (execution_horizon == actions_per_step == N) keeps its FULL chunk - clamping to a smalleraction_horizondrops the chunk tail and forces an out-of-distribution re-query. Single-action providers (MockPolicy) haveexecution_horizon == 1so the result is justmax(action_horizon, 1).
Before this helper existed each consumer inlined the same
max(action_horizon, getattr(policy, "actions_per_step", 1)) expression
and they drifted; worse, all of them keyed off actions_per_step, so an
RTC policy was re-queried only after its full trained chunk drained and its
cross-chunk blending never engaged.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
policy
|
Policy
|
Any policy. The re-query interval is read from
:attr: |
required |
action_horizon
|
int
|
Consumer-requested actions per chunk (clamped to >= 1). Ignored for RTC policies, which decide their own interval. |
required |
Returns:
| Type | Description |
|---|---|
int
|
The number of leading chunk actions to execute before re-querying. |
align_action_values ¶
align_action_values(values: Sequence[float] | ndarray, action_keys: Sequence[str], *, pad_short: bool = False) -> tuple[list[float], list[str]]
Pair a model's ordered action vector with the actuator keys it drives.
Every provider maps a policy's flat action vector onto actuator names BY INDEX, and the two lengths are not guaranteed to agree: a checkpoint trained for a 6-DOF arm can be pointed at a 7-actuator robot, or an embodiment can declare a gripper the checkpoint never learned. This centralizes the single rule for that mismatch so providers cannot drift.
- More values than keys - the trailing values are dropped. There is no actuator to receive them.
- Fewer values than keys (the default) - only the leading keys the model actually produced a value for are returned. The unmatched actuators are left out of the action dict entirely, so they receive no command and hold their current position.
- Fewer values than keys with
pad_short=True- the unmatched keys are returned carrying0.0. That is a COMMAND, not an omission: where the action space is absolute position - a LeRobot<motor>.posfollower, a MuJoCo position actuator -0.0means "travel to zero", so those actuators MOVE, at whatever rate the servo will do it. Opt in only when the consumer needs a fixed-width action dict and zero is a meaningful target for every key it pads.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
values
|
Sequence[float] | ndarray
|
The model's per-step action vector. Any sized, indexable
numeric sequence - a list or a 1-D array, as the two providers hand
over a NumPy row; entries are coerced with |
required |
action_keys
|
Sequence[str]
|
Ordered actuator keys the vector maps onto, index 0 first. |
required |
pad_short
|
bool
|
Emit |
False
|
Returns:
| Type | Description |
|---|---|
list[float]
|
|
list[str]
|
into an action dict after any unit conversion has been applied to the |
tuple[list[float], list[str]]
|
values. |
Factory¶
Policy factory - create_policy() and runtime registration.
UntrustedRemoteCodeError ¶
Bases: RuntimeError
Raised when a HF model requires trust_remote_code but the user has not opted in.
Carries a stable machine-readable :attr:code
(:data:~strands_robots.refusal_codes.TRUST_REMOTE_CODE_REQUIRED) and the
:attr:subject provider, so a consumer offering the operator the opt-in
classifies on identity instead of matching the message text. The message
is unchanged by this. See :mod:strands_robots.refusal_codes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
message
|
str
|
The operator-facing reason, unchanged by the code. |
''
|
code
|
str | None
|
A member of :data: |
None
|
subject
|
str | None
|
The policy provider the gate refused. |
None
|
Attributes:
| Name | Type | Description |
|---|---|---|
code |
The stable identifier for this refusal, or |
|
subject |
The policy provider the gate refused, or |
create_policy ¶
Create a policy instance.
Accepts either a provider name or a smart string:
- Provider name:
create_policy("lerobot_local", pretrained_name_or_path="lerobot/act_aloha_sim") - Server URL:
create_policy("ws://gpu-box:8765") - Checkpoint:
create_policy("lerobot/act_aloha_sim")or a path such ascreate_policy("outputs/train/act/checkpoints/last/pretrained_model") - Shorthand:
create_policy("mock")
Any other spelling is a provider name; one carrying stray punctuation
("wbc/", "protomotions:") is refused as an unknown provider with the
nearest names, not forwarded to lerobot_local as a checkpoint id.
All provider definitions live in registry/policies.json.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
provider
|
str
|
Provider name, HF model ID, or server URL. Positional-only,
so a provider keyword named |
required |
**kwargs
|
Provider-specific parameters. |
{}
|
Returns:
| Type | Description |
|---|---|
Policy
|
Policy instance ready for get_actions(). |
Warns:
| Type | Description |
|---|---|
DeprecationWarning
|
If |
Raises:
| Type | Description |
|---|---|
TypeError
|
If |
UntrustedRemoteCodeError
|
If the provider loads HF models with
|
register_policy ¶
register_policy(name: str, loader: Callable[[], type[Policy]], aliases: list[str] | None = None, *, overwrite: bool = False)
Register a custom policy provider at runtime.
Use this to add providers without editing policies.json.
Example::
from strands_robots.policies import register_policy
register_policy("my_provider", lambda: MyPolicy, aliases=["my"])
policy = create_policy("my_provider", ...)
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
Provider name :func: |
required |
loader
|
Callable[[], type[Policy]]
|
Zero-argument callable returning the :class: |
required |
aliases
|
list[str] | None
|
Extra spellings that resolve to |
None
|
overwrite
|
bool
|
Allow |
False
|
Raises:
| Type | Description |
|---|---|
TypeError
|
|
ValueError
|
|
list_providers ¶
List the canonical policy provider names (JSON + runtime); aliases are in :func:list_aliases.
list_aliases ¶
Return every provider alias and the canonical name it resolves to.
:func:create_policy accepts a provider's declared aliases and
shorthands as readily as its canonical name, but
:func:list_providers reports the canonical names from the JSON
registry. Together the two surfaces enumerate every spelling the
registries hold::
registered = set(list_providers()) | set(list_aliases())
That is every registered spelling, not every spelling
:func:create_policy resolves.
:func:import_policy_class falls back
to auto-discovery, so a module under strands_robots.policies that
exports a :class:~strands_robots.policies.base.Policy subclass resolves
under its own module name with no registry entry. Two ship, and neither is
a registry provider because each wraps a policy the caller already holds
rather than building one from config:
composite(:class:~strands_robots.policies.composite.CompositePolicy) builds through this factory --create_policy("composite", lower=..., upper=...)-- and is the one spellingregisteredabove omits.persistent(:class:~strands_robots.policies.persistent.PersistentPolicy) resolves but cannot be built here: its first parameter is namedprovider, which :func:create_policyhas already bound, so it is constructed directly. :func:create_policyrefuses it with aTypeErrorthat says so.
Covers both registries, matching the union :func:list_providers
reports: aliases declared in policies.json and aliases passed to
:func:register_policy at runtime. A runtime alias shadows a JSON
alias of the same name, which is the precedence
:func:create_policy applies.
Returns:
| Type | Description |
|---|---|
dict[str, str]
|
Mapping of alias to the canonical provider name it resolves to. |
import_policy_class ¶
Dynamically import and return the Policy class for a provider.
Uses the module + class paths from policies.json. Falls back to
auto-discovery (strands_robots.policies.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
provider
|
str
|
Canonical provider name. |
required |
Returns:
| Type | Description |
|---|---|
type
|
The Policy subclass. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the provider does not exist, or was removed - a removed
spelling ( |
ImportError
|
If the provider exists but its module cannot be imported,
naming the provider, the missing module and the remedy (see
:func: |
preflight_policy ¶
Run a provider's class-level :meth:Policy.preflight check, if any.
Resolves provider to its policy class WITHOUT instantiating it (so no
model weights are downloaded) and invokes the class's preflight hook
with the runtime observation_keys and the provider kwargs. Providers
that do not override :meth:Policy.preflight are a no-op.
This is the fail-fast seam used by SimEngine.run_policy /
eval_policy to catch a misconfiguration (e.g. sim camera names that
cannot be routed to the model's declared image inputs) BEFORE the
expensive create_policy download, instead of crashing deep inside the
first inference. Resolution failures are swallowed (the matching error is
surfaced authoritatively by the subsequent create_policy); only the
provider's own preflight ValueError propagates.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
provider
|
str
|
Provider name, HF model ID, or server URL (as passed to
|
required |
observation_keys
|
set[str]
|
Keys the runtime observation will contain (joint names + camera names). |
required |
**kwargs
|
Provider-specific parameters (the policy_config). |
{}
|
Raises:
| Type | Description |
|---|---|
TypeError
|
When |
ValueError
|
When the resolved provider's |
preflight_reason ¶
preflight_reason(provider: str, read_observation_keys: Callable[[], Iterable[str]], /, **kwargs: Any) -> str | None
Why provider refuses this configuration, or None.
The whole pre-build check in one call, so the three entry points that owe it - the simulation engine, the physical arm and a native driver's task verb - read one rule instead of keeping three copies of it in step:
- the observation is read only when the resolved class actually overrides
:meth:
Policy.preflight(:func:policy_overrides_preflight). That read is not cheap - the sim renders every camera in the scene, an arm warms and grabs a frame from each configured camera, a driver crosses the wire - and for every shipped provider butlerobot_localthe result is gathered only to be discarded. - a read that fails, or answers nothing, is not a verdict on the policy configuration and does not become one here: the check is skipped and that read stays the caller's own to report.
- the provider's
ValueErrorcomes back as text, because two of the three callers answer a refusal envelope rather than raise.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
provider
|
str
|
Provider name, HF model ID, or server URL (as passed to
:func: |
required |
read_observation_keys
|
Callable[[], Iterable[str]]
|
Answers the keys the runtime observation will carry (joint names plus camera names). Called at most once, and only when there is a hook to feed. |
required |
**kwargs
|
Any
|
Provider-specific parameters (the policy_config), judged as
the mapping :func: |
{}
|
Returns:
| Type | Description |
|---|---|
str | None
|
The provider's refusal text, or |
str | None
|
when there is no hook to run, or when the observation could not be read. |
Raises:
| Type | Description |
|---|---|
TypeError
|
When |
Built-in policies¶
strands_robots.policies.mock.MockPolicy ¶
Bases: Policy
Mock policy for testing - generates smooth sinusoidal trajectories.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
amplitude
|
float
|
Peak of the per-joint sinusoid in radians (default |
0.5
|
seed
|
int | None
|
Accepted for parity with the other providers' |
None
|
The constructor declares its keywords, so the factory's near-miss screen
applies to the mock as to every other provider: policy_config={"amplitud":
0.5} is refused with a did-you-mean instead of running on the default
(#4165), and a bag passed whole as
create_policy("mock", policy_config={...}) is refused naming the unpack;
docs/learn/policies/index.md promises a misspelled keyword is
refused before anything runs. The **kwargs sink stays because the
hardware drivers hand every provider the server address (host, and
port when one is given) whether or not it dials one; a name that is
not close to a declared keyword is forwarded there and logged at DEBUG,
which is the pass-through contract every provider with a sink has.
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
requires_images
property
¶
Mock policy only consumes joint state - skip camera rendering.
get_actions
async
¶
get_actions(observation_dict: dict[str, Any], instruction: str, **kwargs: Any) -> list[dict[str, Any]]
Return smooth sinusoidal actions.
Canonical reference for the per-tick action value convention
documented on :meth:Policy.get_actions: every value is a python
float (single-DOF joint target), never a raw np.ndarray.
reset ¶
Rewind the sinusoid to its first step.
The mock is deterministic but not stateless: _step advances by one
chunk per :meth:get_actions, and it is the only per-episode state the
mock holds. Left where the previous episode ended, two episodes seeded
alike began at different phases of the sinusoid, so the reference
implementation of the ABC broke the reproducibility its own
:meth:Policy.reset docstring asks providers to keep. The seed is not
read: the trajectory has no random draw for it to reach.
set_robot_state_keys ¶
Record the ordered joint keys used to name the sinusoidal action dict.
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
set_sim_context ¶
Learn the range each driven actuator is held to, so the sinusoid stays inside it.
Called by the MuJoCo engine's bind_policy_sim_context right after
:meth:set_robot_state_keys, with the compiled MjModel and the
robot's namespace prefix ("so100/"). The mock's ±0.5 rad sinusoid
was written for a generic joint; on a real model some actuators do not
span it - the SO-100 Pitch ctrlrange is [-3.32, 0.174] and its
Jaw is [-0.174, 1.75] - so the value was held to the range and
the engine warned that the commanded trajectory was NOT reproduced.
That warning was the first thing examples/01_sim_hello_world.py
printed. Knowing the ranges, the mock clips its own output so what it
commands is what the actuator does.
Which range holds the command is
:func:~strands_robots.simulation.mujoco.scene_ops.effective_ctrl_range's
rule, read here rather than re-derived: an actuator whose MJCF authors
neither ctrlrange nor inheritrange compiles to
ctrlrange == (0, 0) with actuator_ctrllimited == 0, and for a
position servo its ctrl IS the joint target, so the driven joint's
limits bound the pose. Reading only actuator_ctrllimited therefore
learned nothing at all on the so101 - all six of its actuators are in
that second case - and the mock commanded -0.433 to a jaw whose joint
range is [-0.1745, 1.745], which is the warning this method exists
to prevent.
Actuators that are held to no range and names that resolve to no actuator are left alone; an error while reading the model leaves the policy exactly as configured, never fails the rollout.
strands_robots.policies.composite.CompositePolicy ¶
CompositePolicy(lower: Policy, upper: Policy, *, lower_joints: Sequence[str] | None = None, upper_joints: Sequence[str] | None = None, lower_obs_keys: Sequence[str] | None = None, upper_obs_keys: Sequence[str] | None = None)
Bases: Policy
Compose a lower and an upper policy on one robot's joint set.
Each :meth:get_actions queries both children with the (optionally
per-child filtered) observation and the same instruction + kwargs, then
merges their per-tick action dicts by joint name:
lowercontributes the names inlower_joints(or all names it emits whenlower_jointsisNone).uppercontributes the names inupper_joints(or, whenNone, every name it emits that the lower policy did not already claim - lower precedence).
An explicit group is EXCLUSIVE, either way round: the policy it names is the
only one allowed to command those joints. A defaulted lower_joints may not
command into an explicit upper_joints, and a defaulted upper_joints may
not command into an explicit lower_joints - that tick is refused, on every
tick alike, rather than resolved by whichever child emitted the name.
Precedence only decides between two DEFAULTED groups, where the caller
declared no owner.
Routing that discards a child's ENTIRE action dict raises: the composite would otherwise silently be the surviving child alone. Two children that drive the same joint set cannot be composed, only cascaded (module docstring). An observation subset that selects NOTHING raises for the mirror reason: a child queried with an empty dict acts on no reading at all.
The merged chunk length is the shorter of the two children's chunks, so the
more frequently re-querying child sets the re-query cadence
(:attr:execution_horizon). This keeps a per-tick controller (WBC,
execution_horizon == 1) closed-loop even when paired with a chunk-emitting
manipulation policy.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
lower
|
Policy
|
Policy driving the lower joint group (e.g. legs+waist locomotion). |
required |
upper
|
Policy
|
Policy driving the upper joint group (e.g. arms manipulation). |
required |
lower_joints
|
Sequence[str] | None
|
Joint/actuator names the lower policy is authoritative for.
Exclusive: no other child may command these, whichever names the lower
policy emits on a given tick. |
None
|
upper_joints
|
Sequence[str] | None
|
Joint/actuator names the upper policy is authoritative for.
Exclusive: no other child may command these, whichever names the upper
policy emits on a given tick. |
None
|
lower_obs_keys
|
Sequence[str] | None
|
Observation keys to forward to the lower policy. |
None
|
upper_obs_keys
|
Sequence[str] | None
|
Observation keys to forward to the upper policy. |
None
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
children
property
¶
Both child policies, lower first - the order they are merged in.
Lets a runtime capability probe reach the concrete policies inside the
composite; see :attr:Policy.children.
execution_horizon
property
¶
Re-query interval: the shorter of the two children's horizons.
The merged chunk is truncated to the shorter child's length, so the consumer must re-query at the faster child's cadence to keep that child closed-loop (a per-tick locomotion controller must not be starved by a slower chunk-emitting manipulation policy).
requires_images
property
¶
True if EITHER child consumes camera frames.
The composite cannot skip rendering unless both children opt out; a manipulation upper body typically needs images even when the locomotion lower body does not.
get_actions
async
¶
get_actions(observation_dict: dict[str, Any], instruction: str, **kwargs: Any) -> list[dict[str, Any]]
Query both children and merge their per-tick action dicts by joint name.
Both children receive the same instruction and kwargs (each
ignores keys it does not use, per the :class:Policy contract), and the
observation filtered to its configured key subset. The two action chunks
are merged element-wise up to the shorter length.
Returns:
| Type | Description |
|---|---|
list[dict[str, Any]]
|
The merged action chunk (length == the shorter child's chunk). Each |
list[dict[str, Any]]
|
dict maps a joint/actuator name to its target value, routed from the |
list[dict[str, Any]]
|
child that owns that name. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If a configured observation subset shares no key with the observation (the child would be queried blind), if either child returns an empty chunk, if routing discards a child's entire action dict (the composite would be the other child alone), or if either child commands a joint the other child's explicit joint group assigns to it. |
set_control_frequency ¶
Set the control rate on the composite and forward it to both children.
set_robot_state_keys ¶
Forward the robot's state-key list to both children.
set_rtc_observed_delay ¶
Forward the RTC observed-delay step count to the composite and both children.
Persistent cache¶
Persistent, reusable policy handles - load the model once, reuse everywhere.
Loading a VLA / LeRobot checkpoint (a MolmoAct2 SO-100/101 build reads ~1300
weight files into GPU memory) costs on the order of a minute or two. A naive
multi-episode loop that calls :func:create_policy per rollout pays that cost
every episode; the dominant fix already exists - a process-level model cache in
:mod:strands_robots.policies.lerobot_local.policy shares the resident weights
across instances - but two ergonomic gaps remained:
- The win was implicit. There was no first-class "load this once and hand me a
handle I reuse" object, so an LLM harness driving the API blind had no
obvious way to express the intent and would re-call
create_policy. - There was no provider-agnostic way to warm the cache ahead of a run, see what is resident, or free a checkpoint between runs of different policies.
This module closes both. :class:PersistentPolicy is a thin, thread-safe
wrapper that builds the underlying policy ONCE at construction and is meant to
be passed to every run_policy/eval_policy call via policy_object=::
from strands_robots.policies import PersistentPolicy
policy = PersistentPolicy("lerobot_local", pretrained_name_or_path="...")
for _ in range(20):
sim.run_policy(robot_name="arm", policy_object=policy) # zero reload
sim.save_episode()
sim.reset()
The :func:preload, :func:list_cached, and :func:evict helpers are the
agent-facing cache controls: warm before a run, introspect what is hot, free
memory before switching checkpoints.
This is a SYNCHRONOUS persistent worker: the model lives in-process and is shared via the module-level cache. It deliberately does not spawn a background daemon or expose cross-process IPC - inference is GIL- and GPU-serialised, so a per-call lock gives correct concurrent reuse without the complexity (and races) of a separate worker process. Cross-process sharing is a separable concern.
PersistentPolicy ¶
Bases: Policy
A persistent, reusable handle around an underlying policy.
Builds the wrapped policy ONCE (warming the process-level model cache) and
delegates every :class:Policy operation to it, so the same object can be
passed to many run_policy/eval_policy calls without ever reloading
weights. Inference calls are serialised by a per-call lock, so two threads
sharing one handle never corrupt the wrapped model's per-episode state.
The wrapper is transparent: chunk-shape introspection (execution_horizon,
is_chunk_emitting, actions_per_step, supports_rtc), RTC delay /
control-frequency hooks, reset, and load telemetry (load_time_s,
load_cache_hit) all forward to the wrapped policy, so the runtime drives
it exactly as it would the bare policy.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
provider
|
str
|
Provider name or smart string forwarded to
:func: |
required |
policy_object
|
Policy | None
|
An already-constructed policy to wrap instead of building
a new one. When given, both |
None
|
**config
|
Any
|
Provider-specific keyword arguments forwarded to
:func: |
{}
|
children
property
¶
The single wrapped policy.
Keeps the wrapper transparent to a runtime capability probe as well as to
the delegated :class:Policy operations; see :attr:Policy.children.
control_frequency
property
writable
¶
The control rate the wrapped policy was told (None until set).
execution_horizon
property
¶
Actions consumed from one get_actions chunk, from the wrapped policy.
provider_name
property
¶
Provider name of the wrapped policy (identifies the resident model).
requires_images
property
¶
Whether the wrapped policy needs camera frames in its observation.
rtc_observed_delay_steps
property
writable
¶
The RTC delay the wrapped policy was told (None until set).
get_actions
async
¶
get_actions(observation_dict: dict[str, Any], instruction: str, **kwargs: Any) -> list[dict[str, Any]]
Delegate to the wrapped policy under the shared thread lock (see :meth:Policy.get_actions).
Acquires the same threading.Lock the sync path uses, so the sync and
async entry points mutually exclude on the wrapped model's per-episode
state. The blocking acquire runs in the loop's default executor so a
shared running loop is never frozen while another caller holds the lock
(each runtime call still runs on its own per-call asyncio.run loop).
Cancellation-safe: if this coroutine is cancelled (e.g. by
asyncio.wait_for) while the acquire is still pending, the lock is
released as soon as the executor thread obtains it, so a cancelled call
never poisons the shared handle. See :class:_LockHandoff.
get_actions_sync ¶
get_actions_sync(observation_dict: dict[str, Any], instruction: str, **kwargs: Any) -> list[dict[str, Any]]
Delegate to the wrapped policy under the per-call lock (see :meth:Policy.get_actions_sync).
is_chunk_emitting ¶
Whether the wrapped policy returns multi-action chunks per get_actions.
reset ¶
Reset the wrapped policy's per-episode state (weights stay resident).
set_control_frequency ¶
Forward the executing loop's control rate to the wrapped policy.
set_robot_state_keys ¶
Forward the robot state keys to the wrapped policy.
set_rtc_observed_delay ¶
Forward the observed inference delay (RTC steps) to the wrapped policy.
preload ¶
Warm the model cache for a provider and report the load cost.
Builds a :class:PersistentPolicy (which loads the model once into the
process-level cache) and measures the wall time and resident-memory delta.
Call this before a multi-episode run so every subsequent run_policy with
the returned policy is a zero-reload cache hit.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
provider
|
str
|
Provider name or smart string (see :func: |
required |
**config
|
Any
|
Provider-specific keyword arguments. |
{}
|
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
A dict with:
|
list_cached ¶
List the heavy models currently resident in the process-level cache.
Provider-agnostic introspection for deciding whether to evict before
loading a different checkpoint. Currently the lerobot_local provider is
the only one with a process-level weight cache; the list is empty when its
optional dependencies are not installed.
Returns:
| Type | Description |
|---|---|
list[dict[str, Any]]
|
One dict per cached entry (see |
list[dict[str, Any]]
|
func: |
list[dict[str, Any]]
|
empty list when no cache is available. |
evict ¶
Free cached models, returning held GPU/CPU memory.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pretrained_name_or_path
|
str | None
|
When |
None
|
Returns:
| Type | Description |
|---|---|
int
|
Number of cache entries evicted ( |
LeRobot policy types¶
strands_robots.policies.lerobot_local.resolution.list_policy_types ¶
List the LeRobot policy type strings resolvable in this environment.
These are exactly the values accepted as the policy_type argument to
:func:resolve_policy_class_by_name -- and therefore as
create_policy("lerobot_local", policy_type=...). The list reflects the
installed lerobot: it is sourced from lerobot's own draccus choice
registry (PreTrainedConfig.get_known_choices()) after every policy
config module has been imported via
:func:_ensure_policy_configs_registered, so a brand-new policy a newer
lerobot ships appears automatically and a slimmer install reports fewer.
This is the discovery surface for the lerobot_local provider: rather
than reading lerobot internals to learn which policy_type strings are
valid, a caller (or an agent) can enumerate them here.
Returns:
| Type | Description |
|---|---|
list[str]
|
Sorted list of policy type strings (e.g. ``["act", "diffusion", |
list[str]
|
"smolvla", ...]``). Empty when lerobot is not installed -- a discovery |
list[str]
|
surface yields an empty list on a missing dependency rather than |
list[str]
|
raising. |