Simulation¶
The engine contract, the backend factory, what SimWorld carries, and how to register your own backend.
strands_robots.simulation holds the engine contract, the backend factory and the data types a simulation returns: how backends are created, what SimWorld carries, registering your own.
Factory¶
Simulation factory - create_simulation() and runtime backend registration.
Mirrors the policy factory pattern: JSON-driven defaults with runtime override capability. Backends are lazy-loaded on first use.
Usage::
from strands_robots.simulation import create_simulation
# Default backend (MuJoCo)
sim = create_simulation()
# Explicit backend
sim = create_simulation("mujoco", timestep=0.001)
# GPU-native built-in backends
sim = create_simulation("isaac", num_envs=1, headless=True)
sim = create_simulation("newton")
# Custom backend (runtime-registered)
from strands_robots.simulation.factory import register_backend
register_backend("my_sim", lambda: MySimBackend, aliases=["custom"])
sim = create_simulation("custom")
Third-party packages may also register backends out-of-tree via the
strands_robots.backends entry-point group (see create_simulation).
create_simulation ¶
Create a simulation backend instance.
This is the primary entry point for creating simulations. Backend classes are lazy-loaded on first call.
Resolution order for backend:
- Runtime-registered backends (see
register_backend). - Built-in backends (currently
mujoco,newton,isaac). Built-ins always win over entry-point plugins of the same name, so a third-party plugin can never accidentally shadow a built-in backend. -
Entry-point plugins. Third-party packages (e.g.
strands-robots-sim <https://github.com/strands-labs/robots-sim>_) register heavy out-of-tree backends - Isaac Sim, Newton - by declaring them under thestrands_robots.backendsentry-point group in theirpyproject.toml::[project.entry-points."strands_robots.backends"] newton = "strands_robots_sim.newton.simulation:NewtonSimulation" warp = "strands_robots_sim.newton.simulation:NewtonSimulation"
so they can be discovered on pip install without patching this
package. A plugin may map several entry-point names to the same class
(newton and warp above) - whichever name is requested resolves
cleanly. Plugins are discovered lazily on the first
create_simulation / list_backends call (not at import time),
and a plugin that fails to import is logged and skipped rather than
crashing the factory. See the Python packaging spec for details:
https://packaging.python.org/en/latest/specifications/entry-points/
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
backend
|
str
|
Backend name or alias. Defaults to |
DEFAULT_BACKEND
|
**kwargs
|
Any
|
Backend-specific keyword arguments passed to the
constructor (e.g., |
{}
|
Returns:
| Type | Description |
|---|---|
SimEngine
|
A |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the backend name is not recognized. The message lists
all available backends (built-in + plugin) and, for known
out-of-tree backends, a |
ImportError
|
If the backend's dependencies are missing
(e.g., |
Examples::
# Default (MuJoCo)
sim = create_simulation()
sim.create_world()
sim.add_robot("so100")
# With alias
sim = create_simulation("mj")
# Pass kwargs to backend constructor
sim = create_simulation("mujoco", tool_name="my_sim")
# GPU-native built-in backend (requires strands-robots[sim-isaac])
sim = create_simulation("isaac", num_envs=1, headless=True)
list_backends ¶
List all available backend names (built-in + plugin + runtime).
Merges the built-in registry (and its aliases), entry-point plugin
backends discovered via importlib.metadata (see
create_simulation), and any runtime-registered backends/aliases.
Discovering plugins triggers a one-time lazy scan of the
strands_robots.backends entry-point group.
Returns:
| Type | Description |
|---|---|
list[str]
|
Sorted list of unique backend identifiers and aliases. |
Example::
>>> list_backends()
['mj', 'mjc', 'mjx', 'mujoco']
register_backend ¶
register_backend(name: str, loader: Callable[[], type[SimEngine]], aliases: list[str] | None = None, force: bool = False) -> None
Register a custom simulation backend at runtime.
Use this to add backends without editing source code.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
Backend identifier (e.g., |
required |
loader
|
Callable[[], type[SimEngine]]
|
Zero-arg callable that returns the backend class
(not instance). Called lazily on first |
required |
aliases
|
list[str] | None
|
Optional short names that resolve to |
None
|
force
|
bool
|
If False (default), raises ValueError when |
False
|
Raises:
| Type | Description |
|---|---|
TypeError
|
If loader is the backend class itself rather than a
callable returning it. Calling a class produces an instance,
and the backend is instantiated later by |
ValueError
|
If |
Example::
from strands_robots.simulation.factory import register_backend
register_backend(
"bullet",
lambda: BulletSimulation,
aliases=["pybullet", "pb"],
)
sim = create_simulation("bullet")
Engine contract¶
stop_policy returns a json block with robot, was_running and exited; exited is null with no worker to join, the first stop after a rollout ended on its own adds last_result, and an empty name means the one rollout in flight. MuJoCo waits 1 s (MuJoCoSimulation._POLICY_STOP_JOIN_TIMEOUT) for the worker first.
strands_robots.simulation.base.SimEngine ¶
Bases: ABC
Abstract base class for simulation engines.
Defines the contract that all backends (MuJoCo, Isaac, Newton) must implement. This is the programmatic API - the AgentTool layer wraps it with tool_spec/stream for LLM access.
Method categories:
Required (@abstractmethod): Core simulation loop - world
lifecycle, entity management, observation/action, rendering, robot
discovery. Every physics engine must implement these to be usable.
Provided (concrete base-class methods): Policy orchestration
(run_policy / start_policy / replay_episode / eval_policy)
is implemented once in this ABC as a facade over the abstract primitives.
Backends inherit them for free by implementing the primitives. They
may override for backend-specific optimisations (e.g. GPU-batched
policy inference on Isaac).
Optional (default raises NotImplementedError): Higher-level
features - scene loading, domain randomization, contact queries.
Backends opt in by overriding only what they support.
Lifecycle::
sim = SomeEngine()
sim.create_world()
sim.add_robot("so100", data_config="so100")
sim.add_object("cube", shape="box", position=[0.3, 0, 0.05])
# Control loop
obs = sim.get_observation("so100")
sim.send_action({"joint_0": 0.5}, robot_name="so100")
sim.step(n_steps=10)
# Render
result = sim.render(camera_name="default")
# Cleanup
sim.destroy()
Concrete engines must set self._init_complete = True as the final
statement of their __init__. :meth:__del__ consults it and skips an
instance that never finished construction, so a half-built engine is never
reported as a cleanup failure.
foxglove_url
property
¶
The live Foxglove WebSocket URL, or None when no Foxglove bridge runs.
foxglove_link
property
¶
A foxglove:// deep link to this engine's server, or None.
robot_name
property
¶
The name of the one robot in this world, as its methods accept it.
Robot("so100").robot_name is "so100" - the string passed to
Robot(), not the agent tool name ("so100_sim") - and matches
robot_name on the real-hardware return. None when the world
holds no robot or several, since then no single name answers it.
predicate_robot
property
¶
The robot an unnamed base_* clause reads ON THIS THREAD, or None.
Read-only; set through :meth:bind_predicate_robot. The binding is
thread-scoped, not scene-scoped: a rollout binds on the thread that
drives it and every per-step read of the binding happens on that same
thread, so two rollouts on two robots each read their own robot. See
:meth:bind_predicate_robot for why a scene-wide attribute could not
carry this.
capabilities ¶
Return CAPABILITIES, or derive it: the default set plus each overridden optional method.
Returns:
| Type | Description |
|---|---|
frozenset[str]
|
Frozen set of names from :mod: |
create_world
abstractmethod
¶
create_world(timestep: float | None = None, gravity: list[float] | None = None, ground_plane: bool = True, terrain: str | None = None, difficulty: float = 1.0) -> dict[str, Any]
Create a new simulation world.
terrain ("rough" = value-noise bumps, "stairs" = discrete
step plateaus rising along +x, "pyramid" = concentric step plateaus
rising toward the centre, "slope" = a constant-grade inclined ramp;
see :mod:strands_robots.simulation.terrain) lays down a
deterministic heightfield instead of the flat ground plane so a
locomotion policy can be spawned/evaluated on non-flat ground; it is
only meaningful when ground_plane=True and defaults to None (a
flat plane). Backends without heightfield support reject a non-None
terrain with an actionable error rather than silently ignoring it.
difficulty scales the terrain's peak elevation (1.0 = full
height, <1 gentler, >1 harsher) so a curriculum can ramp
terrain magnitude across resets without changing the terrain kind.
It is only meaningful with a terrain; setting difficulty != 1.0
with no terrain is rejected with an actionable error rather than
silently having no effect. Must be a finite value > 0.
A floating-base robot added to a terrain world is spawned seated on
the local terrain surface (raised by the heightfield height beneath
it) at add_robot and on reset(), so its feet are not buried
below the raised terrain.
timestep (seconds) and gravity must be values the engine can
honor, on the same terms the set_timestep / set_gravity setters
enforce: timestep a finite number > 0 (0 is rejected, never
coalesced to the engine default), gravity a 3-element vector of
finite numbers or a real scalar taken as the z-component. A value the
backend cannot apply is rejected with a structured error rather than
compiled into the world - a world built around a negative or nan
dt integrates backwards or to nan while every subsequent call
still reports status="success". None means "use the engine
default".
destroy
abstractmethod
¶
Destroy the simulation world and release resources.
reset
abstractmethod
¶
Reset simulation to its initial state.
Contract: on return the world must be left in a fully consistent,
observation-ready state - derived kinematics (Cartesian body/site/geom
poses and camera transforms) must reflect the reset pose WITHOUT
requiring a subsequent step(). eval_policy calls
get_observation() immediately after reset() and before the
first action, so a backend that leaves derived state stale would feed
the policy's first inference of every episode a degenerate observation.
The MuJoCo backend enforces this by running mj_forward after
mj_resetData (which alone zeroes all derived quantities). It also
re-applies any per-robot home pose captured from an
add_robot(keyframe=...) spawn, so a keyframe pose survives a reset
instead of collapsing to the zero configuration.
step
abstractmethod
¶
Advance simulation by n physics steps.
When the backend exposes an engine lock (self._lock, all in-tree
backends), implementations must not hold it for the whole count: they
release it at least every :attr:_STEPS_PER_BATCH steps, and re-check
that the world still exists on each batch boundary before advancing it,
aborting with a structured error naming the steps completed if it does
not. Releasing the lock is what makes a concurrent teardown reachable
mid-call, so the two halves are one contract rather than two - the same
pairing _primitive_abort_reason already makes for the motion-primitive
loops, which release the lock on the same schedule.
add_robot
abstractmethod
¶
add_robot(name: str, urdf_path: str | None = None, data_config: str | None = None, position: list[float] | None = None, orientation: list[float] | None = None, keyframe: str | int | None = None) -> dict[str, Any]
Add a robot to the simulation.
keyframe optionally spawns the robot in a canonical pose declared
by a <keyframe> in its source model (e.g. panda "home", aloha
"neutral_pose") instead of the default all-zero configuration.
Pass the keyframe name (str) or index (int). The pose is
applied to the robot's joints by name and stored so :meth:reset
restores it (a keyframe spawn is sticky across resets, matching how a
benchmark restores its canonical start each episode). None (the
default) keeps the historical zero-pose spawn. An unknown keyframe
name/index is a hard error that names the available keyframes; it
never silently falls back to zeros.
Refused while a dataset recording is live, on every backend
(:meth:~strands_robots.simulation.recording.DatasetRecordingMixin._recording_schema_frozen_error):
the recorder's columns were declared at start_recording and a robot
added into them has nowhere to be written.
remove_robot
abstractmethod
¶
Remove a robot from the simulation.
list_robots
abstractmethod
¶
Return ordered list of robot names currently in the world.
Used by the backend-agnostic PolicyRunner to resolve a
default robot when the caller omits robot_name.
robot_joint_names
abstractmethod
¶
Return ordered joint names for robot_name.
This is every joint, in the backend's order, including a floating
base's free joint (floating_base_joint on g1), which has seven
position coordinates and no scalar column. A LeRobotDataset recording
writes observation.state in this order with free joints left out
(the base goes to the base_* columns), so on a floating-base robot
this list is one wider than that vector. Bind a policy with
:meth:robot_action_keys (Policy.set_robot_state_keys,
send_action with a numeric vector, PolicyRunner.replay): it
skips the free joint and names actuators, which are not always joints.
Raises:
| Type | Description |
|---|---|
ValueError
|
|
robot_action_keys ¶
Return the action keys send_action resolves for robot_name.
These are the names a policy should emit as its action-dict keys: the
robot's actuators, which are NOT always its joints. A robot can have
passive/mimic joints with no driving actuator (gripper finger
followers) and tendon-driven actuators that are not joints at all (a
grasp tendon). Keying a policy by robot_joint_names in those cases
emits keys that send_action cannot resolve, so the affected
actuators never move and the robot silently no-ops.
The default mirrors :meth:robot_joint_names for backends whose
actuator set matches their joint set. Two kinds of backend override it.
One has a distinct actuator namespace (MuJoCo tendon grippers) and
returns actuator short-names instead. The other shares the namespace but
commands a subset of it: the Newton engine drops a floating base's
6-DoF free joint, which is a joint and not a commandable scalar, so its
action keys are the joint names minus that one. An override may
therefore rename or narrow this list, and a caller must not assume it
has the same width as :meth:robot_joint_names.
An override orders the keys by the joint each actuator drives, because
this list also orders the observation.state vector a policy reads
(Policy.set_robot_state_keys) and a recording writes those columns in
joint order. An actuator that drives no single joint has no joint to be
ordered by and keeps its backend-declared position.
actuator_ranges ¶
Return the (low, high) control range of each range-limited actuator of robot_name.
Keyed like :meth:robot_action_keys. A backend clamps a send_action
target past these bounds silently, so a driver reports the clamp from
this map rather than from the backend's model. An unlimited actuator is
absent; the default is {} for backends that cannot report ranges.
Raises:
| Type | Description |
|---|---|
ValueError
|
|
saturated_actuators ¶
Return the actuators of robot_name whose applied force sits at its force limit now.
Keyed like :meth:robot_action_keys. A position servo pinned at its
limit is pushing against contact or a joint stop instead of reaching its
command, which key resolution cannot see. [] means none is pinned;
the default None means the backend cannot tell.
Raises:
| Type | Description |
|---|---|
ValueError
|
|
bind_predicate_robot ¶
Bind the robot an unnamed base_* clause reads, for the calling thread.
Benchmark and stop_when clauses default robot to "the sole
robot". In a multi-robot scene that used to resolve to the FIRST
registered robot, so evaluate_benchmark(benchmark_name='go2_walk_forward',
robot_name='go2') with an arm registered first probed the arm ("has no
floating base") and, with two floating-base robots, would have scored the
wrong one silently. run_policy / eval_policy / evaluate_benchmark
call this with the robot they resolved; the predicate readers consult it
through :func:~strands_robots.simulation.predicates._bound_robot.
Concurrency contract: the binding is per thread. Rollouts are
per-robot and explicitly concurrent - start_policy submits each to
the engine's executor, and "policies on different robots can execute
concurrently" is a documented surface - so one scene-wide attribute
would make the last bind win: from that instant the OTHER rollout's
unnamed clauses (evaluated every step) would read the wrong robot,
silently, under status=success. All three surfaces bind on the
thread that then drives the rollout, and every reader of the binding
(the base_* predicates, a benchmark's on_episode_start
compatibility check) runs on that same thread, so a thread-local slot
is exactly the scope the binding needs. A refused or concurrent call
therefore cannot disturb a rollout in flight on another thread. The
binding stays until the same thread rebinds; a stale one (its robot
since removed) is dropped by the reader.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
robot_name
|
str | None
|
The robot to bind, or |
required |
bind_policy_sim_context ¶
Give a policy the backend sim context it needs to close the loop.
Default no-op. The MuJoCo engine overrides this to hand a policy that
opts in - by exposing a callable set_sim_context - the compiled
MjModel + the robot's namespace, so an eef/cartesian-delta policy can
auto-configure its IK end-effector frame with zero manual wiring.
MockPolicy opts in to keep its sinusoid inside each actuator's
ctrlrange; this is also the extension point an out-of-tree policy uses.
Policies that do not expose set_sim_context are unaffected.
add_object
abstractmethod
¶
add_object(name: str, shape: str = 'box', position: list[float] | None = None, orientation: list[float] | None = None, size: list[float] | None = None, color: list[float] | None = None, mass: float = 0.1, is_static: bool | None = None, mesh_path: str | None = None, material: dict[str, Any] | None = None) -> dict[str, Any]
Add a primitive or mesh object to the scene.
size is the full extent in meters on every backend: a box's
[x, y, z] edge lengths, a sphere's diameter, a cylinder's or
capsule's [diameter, unused, length]. size=[0.05, 0.05, 0.05]
is a 5 cm cube on MuJoCo, Newton, mjlab and Isaac alike; each backend
halves it into its engine's own half-extents. See the concrete
backend's add_object docstring for the exact per-shape semantics.
Returns an agent-tool status dict.
A backend MUST NOT discard size components the caller did supply.
When the vector is shorter than the shape consumes it either rejects it
(MuJoCo: the per-shape component count is part of the contract) or pads
only the missing trailing components from a documented default
(Isaac). Replacing the whole vector with a backend default compiles a
differently-sized object while reporting success -- and the reported
size echoes what was asked for, not what was built.
The same rule applies to color: a backend either honors the
component count it was given or rejects it, and may complete only
components it documents a default for (MuJoCo completes an RGB triple
with an opaque alpha, and rejects every other count). Falling back to
the backend's default colour paints a surface the caller never asked
for under a success result.
mass must be a finite number greater than zero for a dynamic
object. A backend MUST NOT establish a body on a mass its own
set_body_properties would refuse: a non-finite mass makes the first
integration step produce nan and, because the solver shares one
state vector, poisons every other body in the world too.
is_static is tri-state, and None is the default because it
is the only value that means "the caller did not specify". That is what
lets a backend derive the answer from shape: MuJoCo forces a
shape="plane" static -- a plane is infinite and cannot carry a
dynamic mass -- and refuses an explicit is_static=False there
rather than quietly overriding it. A backend with no shape-derived rule
resolves None to False (dynamic). Declaring the default as
False would state a value the default backend does not deliver, and
would make restating that declared default a hard error for the one
shape whose whole point is being static.
The two chosen values select a posture -- welded to the world, or a
free body the solver integrates -- so a supplied is_static is
checked rather than read by truthiness: anything that is neither a
boolean nor None is refused. Reading it by truthiness inverts both
halves. 0 is the same value as the False a backend may refuse
for a shape it forces static, so it reaches the quiet override that
refusal exists to prevent, and every non-empty string is truthy, so
"false" welds a body the caller asked to be dynamic and lands on
:class:SimObject.is_static, which is annotated bool and read by
list_objects, the scene rebuild and domain randomization.
material (optional): backend-specific visual material/texture
spec. None keeps the flat color rgba (unchanged); a backend
that supports it (MuJoCo) attaches a real material so surfaces can be
matte or textured. Backends that do not support it should reject a
non-None material loudly rather than silently ignore it. A
supporting backend must likewise reject material keys it cannot honor
(a typo, or a field from another renderer) instead of dropping them --
a dropped key renders the backend default while reporting success.
remove_object
abstractmethod
¶
Remove an object from the scene.
get_observation
abstractmethod
¶
Get full observation for a robot: joint state + all attached cameras.
Unified observation consumed by :class:Policy and
:class:~strands_robots.simulation.policy_runner.PolicyRunner.
Backends MUST return a dict with the following schema; extra keys
are allowed.
Schema
"<joint_name>"(float): One entry per joint on the robot, keyed by the model's joint name with any multi-robot namespace stripped ("joint1"on the Panda,"1"on the SO-101). The schema is stable regardless of multi-robot namespacing at the physics-engine level. A registryjoint_labelsname ("shoulder_pan") is a write-side alias:send_actiontakes it, the observation keeps the model's name, so recorded datasets and trained checkpoints keep one column per joint."<joint_name>.vel"(float): The same joint's velocity (rad/s or m/s), one entry per scalar joint, additive beside the position key so position-only consumers are unaffected. Velocity-feedback controllers (WBC's balance loop, the microduck and ProtoMotions observation packers, an RL env with.velin itsactor_obs_keys) read these to close the loop; a backend that omits them feeds those consumers zeros or aKeyErrorwhile the identical policy works elsewhere, which is exactly the portability break this schema exists to prevent. This entry was previously undocumented here and lived only in the MuJoCo implementation, which is how two backends shipped without it. A free-joint (floating) base is NOT a scalar joint and reports its twist viabase_lin_vel/base_ang_velbelow, never as"<name>.vel"."<camera_name>"(np.ndarray): One RGB uint8 frame per camera associated with the robot, keyed by camera name. Shape(H, W, 3). A key MUST carry the view of the camera it names; a backend that cannot render that camera MUST omit the key rather than substitute another view (the free/overview camera in particular), because every consumer of this schema - a policy readingobservation.images.<name>, a recorded dataset column - reads the key as a promise about which camera it is looking through and has no way to detect a substitution. Cameras whose render fails MAY be omitted; joint state MUST still be returned.- Floating base: a robot whose root is a 6-DoF free joint (a
humanoid's named
floating_base_jointor a mobile base's unnamed<freejoint>) does NOT report that free joint as a scalar"<joint_name>"entry - its qpos is [xyz + quat], so a scalar would report the base x-coordinate as a joint angle and drop the rest. Instead it surfaces the full base pose + twist as"base_pos"(world x,y,z incl. height),"base_quat"(w,x,y,z),"base_lin_vel"and"base_ang_vel", matching :meth:get_robot_state's"base"entry. Absent for fixed-base arms. "body.<name>.pos"/".quat"/".lin_vel"/".ang_vel"(list[float]): World pose + twist of a NAMED body, present only when the running policy declared that body in :attr:~strands_robots.policies.base.Policy.required_bodies. Backends do not emit these fromget_observationitself - the runtime (:class:~strands_robots.simulation.policy_runner.PolicyRunner) merges them in from :meth:get_body_statefor the declared bodies only, so the default observation is unchanged and nothing pays for a link nobody asked for. Motion-mimic trackers need them because their anchor link (torso_linkon a G1) is separated frombase_quat(the pelvis) by the waist joints.
Single-camera rendering is :meth:render's job, not this method's.
For batched multi-robot observation (future Isaac / Newton), add a
separate get_observations(robot_names) method - do NOT extend
this one.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
robot_name
|
str | None
|
Which robot to observe. If |
None
|
skip_images
|
bool
|
Skip camera rendering and return joint state only.
Rendering dominates the per-step cost, so every consumer that
reads joint values alone passes |
False
|
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
Observation dict per schema above. Returns |
dict[str, Any]
|
WARNING naming the cause - when there is no world (never created, |
dict[str, Any]
|
or after :meth: |
dict[str, Any]
|
|
dict[str, Any]
|
write methods ( |
dict[str, Any]
|
condition with |
get_ground_height ¶
Query the terrain surface height (world z) beneath world (x, y).
Public counterpart of the internal :meth:_ground_height_at hook: a
create_world(terrain=...) heightfield raises the local ground up to
TERRAIN_ELEVATION * difficulty above z=0, and there was no public
way to ask where that surface is. Callers building a terrain scene need
it to place an object / camera / goal on the surface -- an object added
at a flat-ground z (computed as if the support were at z=0) on a
raised plateau spawns buried in the heightfield and sinks through
instead of resting on it. The same local-height sampler already backs the
terrain-relative locomotion predicates (base_below_z) and the
spawn/reset base-seating; this exposes it as a facade query.
Returns 0.0 for a flat ground plane, for any backend without a
heightfield, and before create_world (a world-less engine has no
terrain), so a non-terrain -- or not-yet-built -- world reports a flat
surface rather than raising, unlike the world-scoped physics queries.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
SupportsFloat
|
World x coordinate. Any object convertible to |
required |
y
|
SupportsFloat
|
World y coordinate. Same accepted types as |
required |
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
Agent-tool status dict. On success |
dict[str, Any]
|
|
dict[str, Any]
|
surface height in meters. Errors when |
dict[str, Any]
|
real number. Accepts any real scalar, including NumPy scalar |
dict[str, Any]
|
types ( |
dict[str, Any]
|
coordinates naturally come from |
dict[str, Any]
|
(a NumPy array), not hand-typed Python floats. |
send_action
abstractmethod
¶
send_action(action: dict[str, Any] | Sequence[float], robot_name: str | None = None, n_substeps: int = 1) -> dict[str, Any]
Apply action and advance physics by n_substeps.
Contract: each call writes actuator/ctrl values and then runs
n_substeps physics steps (e.g. mj_step). PolicyRunner.run()
relies on this - it calls send_action once per control step and
does NOT call sim.step() separately.
n_substeps is a positive whole number, on the shared
:func:~strands_robots.utils.positive_whole_number_error domain every
backend applies. A NumPy or float count with an integral value is
honored and coerced; a fractional, zero, negative, non-finite, boolean
or non-numeric count is refused as a structured error, and nothing is
written when it is - a refusal arriving after the write would leave the
robot commanded and the world un-advanced, which is the one state this
surface must never report an error from. The floor is 1 rather than
:meth:step's 0 precisely because of the write: "advance nothing"
is step(0), an accepted no-op that commands nothing, while a
send_action advancing nothing leaves a target the world never
integrates. It is also the floor both producers of this count already
enforce - PolicyRunner._control_substeps returns >= 1 and
raises otherwise, and training.rl.env.SimEnv refuses an
n_substeps below 1 - so this surface was the only member of that
chain without the guarantee.
Backends are responsible for internal thread-safety (e.g. MuJoCo acquires self._lock here). PolicyRunner does not manage locks.
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
Dict with |
dict[str, Any]
|
not at all: when any action key cannot be resolved, nothing is |
dict[str, Any]
|
written, the world does not advance, and the |
dict[str, Any]
|
includes a |
dict[str, Any]
|
|
dict[str, Any]
|
|
physics_timestep ¶
Return the physics integration timestep in seconds, or None.
Used by :class:PolicyRunner to convert a policy's control_frequency
into the number of physics substeps per control step
(round(1 / control_frequency / physics_timestep)) so a
position-servo robot actually tracks each action's target before the
next action overwrites ctrl. Backends that cannot report a fixed
timestep return None and the runner falls back to n_substeps=1.
render
abstractmethod
¶
render(camera_name: str = 'default', width: int | None = None, height: int | None = None) -> dict[str, Any]
Render a camera view.
Returns an agent-tool dict with status and a content list. On
success the content holds an image block carrying PNG bytes
({"image": {"format": "png", "source": {"bytes": ...}}}); the raw
RGB numpy arrays are available per-camera via :meth:get_observation.
Resolution comes from the named camera's configuration (set via
add_camera) unless width/height are given; the free camera
and model-only cameras fall back to the engine default.
run_policy ¶
run_policy(robot_name: str | None = None, policy_provider: str = 'mock', policy_config: dict[str, Any] | None = None, instruction: str = '', duration: float = 10.0, control_frequency: float | None = None, action_horizon: int = 8, fast_mode: bool = False, video: dict[str, Any] | None = None, policy_object: Policy | None = None, n_steps: int | None = None, max_steps: int | None = None, max_onframe_failures: int | None = None, control_substeps: int | None = None, policy_kwargs: dict[str, Any] | None = None, seed: int | None = None, n_episodes: int = 1, reset_between: bool = True, async_rtc: bool | None = None, rtc_inference_timeout_s: float | None = None, wbc_install_torque_control: bool = True, stop_when: dict[str, Any] | Callable[[SimEngine], bool] | None = None, observer: RunPolicyObserver | None = None) -> dict[str, Any]
Run a policy loop in the simulation (blocking).
Default implementation delegates to the backend-agnostic
:class:~strands_robots.simulation.policy_runner.PolicyRunner.
Backends MAY override for backend-specific optimisations
(e.g. GPU-batched policy inference on Isaac).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
robot_name
|
str | None
|
Robot to control. |
None
|
policy_provider
|
str
|
Name passed to
:func: |
'mock'
|
policy_config
|
dict[str, Any] | None
|
Opaque dict of provider-specific kwargs
( |
None
|
instruction
|
str
|
Natural-language instruction for the policy. |
''
|
duration
|
float
|
Wall-clock seconds to run, honored as such: the loop
paces on a deadline, so a step's own cost comes out of the
period rather than being added to it ( |
10.0
|
control_frequency
|
float | None
|
Target Hz for policy queries. Must be a positive number; a non-positive, non-numeric, or bool value is reported as a structured caller error. |
None
|
control_substeps
|
int | None
|
Explicit physics steps to integrate per applied
action, overriding the |
None
|
action_horizon
|
int
|
Lower bound on actions consumed from each
policy chunk before re-querying. The effective interval is
|
8
|
fast_mode
|
bool
|
Skip the real-time pacing and run as fast as inference
and physics allow. When False (default) the loop is paced on a
deadline at |
False
|
video
|
dict[str, Any] | None
|
Optional video-recording config dict. Accepted keys:
|
None
|
policy_object
|
Policy | None
|
Already-constructed
:class: |
None
|
n_steps
|
int | None
|
Exact control-step horizon. When given it REPLACES
|
None
|
max_steps
|
int | None
|
Legacy alias for |
None
|
max_onframe_failures
|
int | None
|
Maximum consecutive |
None
|
seed
|
int | None
|
Optional master RNG seed for a reproducible single rollout.
When set, reseeds Python / NumPy / torch / cuDNN and forwards
|
None
|
policy_kwargs
|
dict[str, Any] | None
|
Optional per-call goal payload forwarded verbatim to
every |
None
|
n_episodes
|
int
|
Number of sequential episode rollouts to run in this
single call (default |
1
|
reset_between
|
bool
|
When running multiple episodes, reset the sim to its
initial state between episodes (default |
True
|
async_rtc
|
bool | None
|
When |
None
|
rtc_inference_timeout_s
|
float | None
|
Optional hard per-chunk timeout (seconds)
for the async-RTC prefetch. When set, a stuck inference surfaces
as a structured |
None
|
wbc_install_torque_control
|
bool
|
When |
True
|
stop_when
|
dict[str, Any] | Callable[[SimEngine], bool] | None
|
Optional semantic early-return condition: end the
rollout as soon as the WORLD reaches a state, not only when
the step budget runs out - which turns a monolithic rollout
into a retryable primitive an agent can invoke -> inspect ->
re-invoke. A predicate-DSL clause in the same schema as a
benchmark spec's |
None
|
observer
|
RunPolicyObserver | None
|
Optional read-only rollout observer, forwarded verbatim to
:meth: This is a SECOND lane, not the backend's hook: the hook slot is
filled from |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict[str, Any]
|
The standard agent-tool envelope |
|
dict[str, Any]
|
``{"status": "success"|"error", "content": [{"text": ...}, |
|
dict[str, Any]
|
{"json": {...}}]} |
|
dict[str, Any]
|
rollout report; |
|
dict[str, Any]
|
Read the json block by SCANNING |
|
dict[str, Any]
|
with a report = next(b["json"] for b in result["content"] if "json" in b) |
|
dict[str, Any]
|
An early caller-error return (a rejected |
|
dict[str, Any]
|
robot) carries a |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
caller most needs to read. |
|
dict[str, Any]
|
IMPORTANT - |
|
dict[str, Any]
|
whether the CALL was accepted and the loop ran; it does not say the |
|
dict[str, Any]
|
robot did anything useful. A rollout that drove only a SUBSET of |
|
dict[str, Any]
|
the robot's actuators is deliberately |
|
dict[str, Any]
|
operational), so |
|
dict[str, Any]
|
of a Panda's 8 actuators returns |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
on |
|
dict[str, Any]
|
binding-degradation flags below to decide whether a rollout is worth |
|
dict[str, Any]
|
anything. Coarse backend errors remain in |
|
dict[str, Any]
|
excluded from action-rate denominators rather than fabricated as |
|
dict[str, Any]
|
confirmed misses. A TOTAL |
|
dict[str, Any]
|
failure - no emitted key resolving to any actuator - is reported as |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
Fields in the json block: |
|
Identity |
dict[str, Any]
|
|
dict[str, Any]
|
name), |
|
dict[str, Any]
|
policy never read it - |
|
dict[str, Any]
|
the task says, and the |
|
see |
dict[str, Any]
|
attr: |
Horizon |
dict[str, Any]
|
|
dict[str, Any]
|
(alias of |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
horizon was exhausted; |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
deciding whether to retry knows WHY the rollout ended). |
|
dict[str, Any]
|
|
|
attribution |
dict[str, Any]
|
the clause is evaluated only AFTER an applied action, |
dict[str, Any]
|
so one the scene's initial state already satisfies fires on the |
|
dict[str, Any]
|
first step whatever the policy commands, making |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
from a rollout that drove the world to the condition - the mirror |
|
dict[str, Any]
|
of the never-fires case the pre-rollout entity probe refuses. |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
when the flag is |
|
dict[str, Any]
|
every other figure is left as measured: domain randomisation |
|
dict[str, Any]
|
legitimately draws an initial state that satisfies a clause. |
|
dict[str, Any]
|
Action health: |
|
dict[str, Any]
|
an error), |
|
dict[str, Any]
|
of the robot's keys - NOT the number of |
|
dict[str, Any]
|
an action naming no key reaches the backend like any other; a |
|
dict[str, Any]
|
rollout whose count is |
|
dict[str, Any]
|
step and is returned as |
|
dict[str, Any]
|
all-keys-unresolved refusal), |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
so a joint stuck at |
|
dict[str, Any]
|
known step) and |
|
dict[str, Any]
|
of the robot's DOF not confirmed driven across those known steps; |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
only 1 of 6). A coarse backend error is excluded from both rate |
|
dict[str, Any]
|
denominators instead of being counted as a physical miss; it remains |
|
dict[str, Any]
|
visible in |
|
dict[str, Any]
|
step whose applied keys name driven JOINTS rather than actuators is |
|
dict[str, Any]
|
excluded on the same terms: |
|
dict[str, Any]
|
(it looks the joint's driving actuator up), but it reports no |
|
dict[str, Any]
|
actuator per key, so the step is unknown for per-actuator purposes |
|
dict[str, Any]
|
rather than a miss - a rollout keyed entirely that way reports an |
|
dict[str, Any]
|
empty map and |
|
Achievement |
dict[str, Any]
|
resolution says a command reached an actuator, not that |
dict[str, Any]
|
the actuator achieved it, so an arm driven into the table reads |
|
dict[str, Any]
|
healthy on every field above. |
|
dict[str, Any]
|
fraction of steps on which any of the robot's actuators sat at its |
|
dict[str, Any]
|
force limit, and |
|
dict[str, Any]
|
meth: |
|
dict[str, Any]
|
tell). A short burst is a fast move; most of a rollout is a stall. |
|
Video |
dict[str, Any]
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
the requested |
|
dict[str, Any]
|
rollout renders at most one frame per control step). |
|
Episodes |
dict[str, Any]
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
episode indices this call flushed, empty without a recording). |
|
dict[str, Any]
|
Policy binding: |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
means the driving policy could not bind the observation to the |
|
dict[str, Any]
|
model's inputs by name and silently fell back (a camera routed to a |
|
dict[str, Any]
|
model image slot positionally, or |
|
dict[str, Any]
|
from the observation's own scalar keys because none of |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
inputs. |
|
dict[str, Any]
|
Policy load: |
|
dict[str, Any]
|
( |
|
dict[str, Any]
|
rebuilt the policy instead of reusing |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
Chunk-prefetch telemetry, so latency masking is provable from the |
|
dict[str, Any]
|
payload instead of from logs: |
|
dict[str, Any]
|
background chunk pipeline was on - this is NOT the policy's RTC |
|
dict[str, Any]
|
algorithm, which |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
one release with the same values. |
|
dict[str, Any]
|
Across episodes ( |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
already knows - the identity fields, the policy-binding flags and |
|
dict[str, Any]
|
the policy-load telemetry, all read off the ONE policy object the |
|
dict[str, Any]
|
episodes shared. Per-episode action health ( |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
the per-episode horizon/video fields are reported by each record in |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
aggregate an N-episode call can report without choosing a summary:: worst = max(e["partial_action_failure_rate"] for e in report["episodes"]) |
|
dict[str, Any]
|
So the binding-degradation gate reads the same way at any episode |
|
dict[str, Any]
|
count, and the action-health gate is per episode. |
|
dict[str, Any]
|
Fail-fast: if EVERY action step in the opening probe window drives |
|
dict[str, Any]
|
zero actuators - none of the policy's emitted keys resolve to any of |
|
dict[str, Any]
|
the robot's actuators - the rollout can never move the robot, so it |
|
dict[str, Any]
|
returns |
|
dict[str, Any]
|
the full episode (and every remaining model inference call + |
|
dict[str, Any]
|
recording write). The error enumerates the unresolved keys and the |
|
dict[str, Any]
|
robot's valid actuator names. A PARTIAL failure runs to completion, |
|
dict[str, Any]
|
surfaced via |
run_multi_policy ¶
run_multi_policy(policies: dict[str, Policy], instructions: dict[str, str] | str = '', duration: float = 10.0, control_frequency: float | None = None, action_horizon: int | dict[str, int] = 8, n_steps: int | None = None, max_steps: int | None = None, *, fast_mode: bool = False) -> dict[str, Any]
Drive MULTIPLE robots, each with its own policy, in ONE synchronized loop.
The backend-agnostic contract for concurrent multi-robot rollout
(e.g. two arms doing a handover, or a bimanual setup). A backend that
implements it must honour every clause below - they are what
distinguishes this driver from launching one :meth:start_policy
thread per robot, which steps physics per robot and interleaves
single-robot recording frames:
- Per-robot policies:
policiesmaps each driven robot to its own :class:~strands_robots.policies.Policy. Every key must name a robot in the scene;policiesorder defines the merged state/action column order. - Per-robot instructions:
instructionsis either one string applied to all robots or a{robot_name: instruction}mapping. A mapping key naming no driven robot is rejected rather than silently dropped; a robot omitted from the mapping gets an empty instruction (see :meth:_normalize_multi_policy_instructions). - Per-robot action_horizon:
action_horizonis either one int applied to all robots or a{robot_name: horizon}mapping. Every horizon must be a positive integer, and the effective per-robot chunk length is resolved through :func:~strands_robots.policies.base.resolve_chunk_lengthexactly as :meth:run_policyresolves its own (see :meth:_normalize_multi_policy_horizons). - Shared control_frequency: one target Hz for every robot's policy queries, so the robots stay phase-aligned.
- Lockstep physics: each loop iteration applies EVERY robot's control, then steps physics ONCE - regardless of each robot's individual re-query cadence.
- One merged recording frame per timestep: when a dataset
recording is active, each timestep records a single frame carrying
ALL robots' prefixed state/action (
alice__shoulder_pan...) plus all camera images - never one interleaved frame per robot.
The step horizon follows :meth:run_policy's resolution: n_steps
(then its legacy alias max_steps) overrides duration, via
:meth:_resolve_horizon on the shared positive-count domain.
This base implementation is a documented refusal, not a fallback: a backend that has no synchronized multi-robot loop must say so rather than silently driving robots one at a time (which would interleave frames and break the merged-frame contract above). The MuJoCo and Isaac backends override it with full implementations; backends that do not yet (Newton) inherit this structured error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
policies
|
dict[str, Policy]
|
Mapping |
required |
instructions
|
dict[str, str] | str
|
Single instruction string for all robots, or a
|
''
|
duration
|
float
|
Episode length in seconds (steps = duration x freq).
Used only when no |
10.0
|
control_frequency
|
float | None
|
Target Hz for policy queries / physics. Must be a positive number. |
None
|
action_horizon
|
int | dict[str, int]
|
Actions consumed from each policy's chunk before re-querying it, as one int or a per-robot mapping (see contract above). |
8
|
n_steps
|
int | None
|
Exact step horizon (overrides |
None
|
max_steps
|
int | None
|
Legacy alias for |
None
|
fast_mode
|
bool
|
Skip the real-time pacing and run as fast as inference
and physics allow, as :meth: |
False
|
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
A structured |
dict[str, Any]
|
backend class and stating that it does not implement synchronized |
dict[str, Any]
|
multi-robot rollout. Implementing backends return the standard |
dict[str, Any]
|
status dict with per-robot step counts. |
verify_dataset_episodes ¶
Verify the recorded dataset holds exactly expected episodes.
Reads the LeRobot dataset parquet (the ground truth) for the active or
most-recently-recorded session AND cross-checks it against the
meta/info.json total_episodes header. Both must agree with
expected; a parquet that matches expected but disagrees with
info.json (an internally inconsistent dataset) still fails. Reports the
actual episode count.
Call this AFTER :meth:stop_recording for a definitive check that a
collection run produced N distinct episodes rather than one merged
episode_index=0 mega-episode.
Episodes are flushed to meta/episodes/**/*.parquet only at
save_episode / stop_recording (finalize) time, so this reads
the canonical on-disk truth - it does not trust the recorder's in-memory
bookkeeping (which is what :meth:run_policy reports while a session is
still open).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
expected
|
int
|
The episode count the caller intended to record. A non-negative int; anything else is reported as an error dict. |
required |
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
Standard status dict. |
dict[str, Any]
|
holds exactly |
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
metadata sources must agree, never just one - and when any episode |
dict[str, Any]
|
parquet file could not be read ( |
dict[str, Any]
|
since the episodes found are then only a lower bound. An unreadable |
dict[str, Any]
|
or corrupt parquet is reported as this same error dict, never raised. |
save_episode ¶
Flush the current recording episode and begin a fresh one.
Backends that support dataset recording override this (see the MuJoCo
RecordingMixin). The base has no recorder, so it returns a
structured error rather than pretending to flush.
start_policy ¶
start_policy(robot_name: str | None = None, policy_provider: str = 'mock', policy_config: dict[str, Any] | None = None, instruction: str = '', duration: float = 10.0, control_frequency: float | None = None, action_horizon: int = 8, fast_mode: bool = False, video: dict[str, Any] | None = None, policy_object: Policy | None = None, n_steps: int | None = None, max_steps: int | None = None, max_onframe_failures: int | None = None, control_substeps: int | None = None, policy_kwargs: dict[str, Any] | None = None, seed: int | None = None, n_episodes: int = 1, reset_between: bool = True, async_rtc: bool | None = None, rtc_inference_timeout_s: float | None = None, wbc_install_torque_control: bool = True, stop_when: dict[str, Any] | Callable[[SimEngine], bool] | None = None, observer: RunPolicyObserver | None = None) -> dict[str, Any]
Run a policy rollout, in the background where the backend has one.
DEFAULT IMPLEMENTATION IS SYNCHRONOUS: it passes through to
:meth:run_policy and returns only after the rollout has finished, so
its result reports a COMPLETED rollout ("Policy complete on ...") and
the call blocks for the whole duration. The summary line used to
promise a background thread outright, which is what MuJoCo's override
does, not what a caller of this default gets - and the two backends
shipped on this default are the ones whose callers most need to know.
Backends with true background execution override this (MuJoCo, via the
ThreadPoolExecutor it owns) and return as soon as the rollout is
submitted.
Either way :meth:stop_policy is the counterpart, and a caller can tell
which of the two it holds without reading the source: this method's
entry in :meth:describe states which one this engine implements, and a
backend that also tracks rollouts in flight advertises
list_policies_running there beside it.
Takes exactly the keywords :meth:run_policy takes, with the same
meaning and the same refusals, so a blocking call becomes a background
one by changing the method name and nothing else.
stop_policy ¶
Stop robot_name's rollout (cooperative) and report what was in flight.
The counterpart to :meth:start_policy, and the verb that OWNS the
question "was a rollout halted" for every backend. It lived only on the
MuJoCo engine, so on the other backends the attribute did not exist at
all: :meth:~strands_robots.mesh.Mesh._dispatch probes for it with
hasattr and answered "peer exposes no stop_task" for a sim it could
in fact have stopped, and the Device Connect stop RPC re-derived the
answer inline from the per-robot flag - a second construction of the
verdict that
:meth:~strands_robots.simulation.models.SimRobot.request_policy_stop
exists to prevent ("EVERY stop path goes through here ... so they cannot
drift to different answers about whether a rollout was halted"). This is
the same promotion :meth:run_multi_policy had (#2157): a capability
every backend is asked for answers in the tool envelope on all of them,
never with AttributeError because there was no contract.
The flag write itself is backend-owned, because the per-robot rollout
claim is: :meth:_request_policy_stop is the seam, the mirror of the
:meth:_make_run_policy_hook / :meth:_release_run_policy_hook pair
that raises and lowers the same flag around a rollout driven here.
This is a cooperative stop, not a join: it moves the robot's claim out
of date so the rollout's next frame ends it. It cannot interrupt a
rollout that is blocked inside a single send_action or a single
policy inference, and on a backend whose :meth:start_policy is the
synchronous default the caller's own thread is the one inside the
rollout - so the callers that reach this verb usefully are the ones on
another thread (the Device Connect stop RPC and the mesh fanout).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
robot_name
|
str
|
The robot whose rollout to stop. An empty name means
the only rollout in flight when there is exactly one - the
remedy every rollout gate names - and is otherwise refused
naming what is running; it is never silently matched against
the sole robot, because a stop aimed at the wrong robot reads
as a stop that worked (:meth: |
''
|
Returns:
| Name | Type | Description |
|---|---|---|
dict[str, Any]
|
The agent-tool envelope. On success the |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
stop arrived - so a caller aggregating several answers reads the |
|
dict[str, Any]
|
verdict rather than matching on the sentence. Idempotent: a robot |
|
dict[str, Any]
|
with nothing running is |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
backend that keeps no durable per-robot claim to move - that refusal |
|
dict[str, Any]
|
names the class, because "nothing was running" would be an |
|
dict[str, Any]
|
affirmative answer given on no evidence. Isaac is on that default |
|
today |
dict[str, Any]
|
its per-robot record carries a bare |
dict[str, Any]
|
and not the durable counter, and a bare flag write is the exact |
|
dict[str, Any]
|
thing a worker that has not reached its first frame overwrites |
|
dict[str, Any]
|
(#2833), so it refuses rather than reporting a stop it cannot keep. |
policy_result ¶
The envelope the last asynchronous rollout on robot_name ended with.
:meth:start_policy returns "Policy started" and nothing else, so the
report :meth:run_policy would have returned (steps, action health,
video_path, the per-episode list) reached no caller once the
worker finished; :meth:stop_policy after natural completion answered
"Was not running" with no report either (#4162). A backend that keeps a
worker table records the finished envelope, success or error, where
:meth:_rollouts_ended_in_error records the failure reason, and
answers here; the default is None for a backend whose
:meth:start_policy is the synchronous default (its caller already
holds the result).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
robot_name
|
str
|
The robot the rollout was started on. |
required |
Returns:
| Type | Description |
|---|---|
dict[str, Any] | None
|
The |
dict[str, Any] | None
|
that robot, or |
dict[str, Any] | None
|
is still in flight, or when the backend keeps no record. A later |
dict[str, Any] | None
|
rollout on the same robot replaces the entry once it completes. |
list_policies_running ¶
Name the robots a rollout is driving right now.
The public reader of the in-flight population, promoted here from the
MuJoCo engine so it answers on every backend: docs/reference/simulation/rollouts.md
lists it in the Policy action table with no backend qualifier, and
documents :meth:stop_policy -- on this ABC since a robot's stop became
a base contract -- as deriving its verdict from "the same in-flight
population list_policies_running reads". Only one of that documented
pair existed on Newton and Isaac; asking for the other raised
AttributeError. docs/reference/device-connect.md makes the same promise for
the Device Connect stop, whose driver runs on any backend.
The population comes from :meth:_rollouts_in_flight, the one seam the
mesh's reporting surfaces also ask, so a peer polled over the wire and a
caller holding the engine cannot be told different things about the same
instant. MuJoCo's override of that seam delegates to its own registry
reader, so the prune this verb used to perform still happens.
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
|
dict[str, Any]
|
none are, when this backend reports a population. |
dict[str, Any]
|
|
dict[str, Any]
|
running" is an affirmative claim, and a backend that cannot |
dict[str, Any]
|
enumerate its rollouts has no evidence for it. That mirrors |
dict[str, Any]
|
meth: |
dict[str, Any]
|
cannot stand behind, and names the seam to override. |
replay_episode ¶
replay_episode(repo_id: str, robot_name: str | None = None, episode: int = 0, root: str | None = None, speed: float = 1.0, action_key_map: list[str] | None = None) -> dict[str, Any]
Replay a LeRobotDataset episode via PolicyRunner.replay.
robot_name is resolved by the rule :meth:run_policy,
:meth:eval_policy and :meth:evaluate_benchmark share: None
picks the sole loaded robot, a scene holding several of them returns
an error listing the candidates, and a name that IS supplied - the
empty string included - is never re-resolved, so one the scene does
not hold is reported by name. Replay is the one policy surface that
drives the actuators from a recording rather than a policy, so a
substituted robot here is a robot the caller never chose being moved.
episode must be a non-negative whole number - the shared domain the
replay_episode teleop knob uses - and is rejected with a structured
error before the dataset is downloaded. A bool is refused rather than
read as an index: episode=True previously resolved episode 1 and
replayed it under a "success" status.
speed is a playback-rate multiplier (1.0 = real time) and must be a
positive number; a non-positive or non-numeric value is rejected with a
structured error rather than raising or silently playing back at full
speed. speed scales only the wall-clock playback rate: each recorded
frame always advances physics for a full control period (derived from
the dataset fps), so a position-servo robot reproduces the recorded
trajectory instead of under-integrating it.
action_key_map binds recorded action-vector indices to action keys
(default: :meth:robot_action_keys). It must be a non-empty list/tuple
of unique strings whose length matches the recorded action width; a bare
string, a non-string entry, a duplicate key or a width mismatch is
rejected rather than truncated to fit. A "success" status therefore
means at least one recorded action actually reached the actuators and
every frame that carried one was applied - a frame that send_action
could not apply aborts the replay with the frame index, the frames
applied so far and the unresolved keys, and an episode whose frames
carry no action value at all aborts naming the columns they do
carry rather than reporting a replay that commanded nothing. The
json block reports frames_with_action beside frames_applied.
Override per backend for optimised replay (e.g. direct ctrl writes) only when measured necessary.
eval_policy ¶
eval_policy(robot_name: str | None = None, policy_provider: str = 'mock', policy_config: dict[str, Any] | None = None, instruction: str = '', n_episodes: int = 1, max_steps: int = 300, success_fn: str | None = None, success_when: dict[str, Any] | None = None, policy_object: Policy | None = None, control_frequency: float | None = None, control_substeps: int | None = None, action_horizon: int = 8, seed: int | None = None, async_rtc: bool = False, rtc_inference_timeout_s: float | None = None, wbc_install_torque_control: bool = True, on_frame: Callable[[int, dict[str, Any], dict[str, Any]], None] | None = None, max_onframe_failures: int | None = None, policy_kwargs: dict[str, Any] | None = None, video: dict[str, Any] | None = None) -> dict[str, Any]
Multi-episode policy evaluation via PolicyRunner.evaluate.
robot_name resolves like :meth:run_policy: None (the
default) auto-selects the sole robot in a single-robot scene and
errors with the candidate list only when the choice is ambiguous
(multiple robots) or impossible (empty scene). This keeps the two
sibling entry points consistent - a policy you just ran with
run_policy() evals the same way with eval_policy().
n_episodes default lowered from 10 to 1 (callers opt in to
longer evals explicitly).
seed pins the eval the way it pins a single :meth:run_policy
rollout: the client RNGs are reseeded once from it and then per episode
from a master RNG derived from it, and each per-episode seed is
forwarded to policy.reset so a service-mode policy can reseed its
own process. Two evals at the same seed replay identically for a
state-only policy (bit-exact), and to a render tolerance for a camera
policy on GPU rendering, where MUJOCO_GL=egl renders a static scene
with 1 LSB differences between frames and a VLA's trajectory drifts
from them (see :meth:run_policy); compare such evals by
success_rate, not frame by frame. None
leaves RNG state untouched. Only a non-negative integer can seed those
RNGs, so anything else is refused here rather than at the first draw.
Each episode's record in the returned episodes list reports the
seed that attempt ran on, so a caller reading a single failed
episode out of a batch can replay that one rather than the whole eval;
it is None when no seed was given, because an unseeded eval
derives no per-episode seed to report.
policy_object mirrors :meth:run_policy: pass an already-built
Policy to skip the create_policy round-trip (e.g. a loaded
SmolVLA checkpoint you want to evaluate without re-instantiating).
When omitted, the policy is built from policy_provider /
policy_config.
control_frequency / control_substeps flow through to
:meth:PolicyRunner.evaluate so the eval loop steps physics for the
full control period per action (same servo-tracking semantics as
:meth:run_policy). Without these the arm under-steps and the policy
looks like a no-op (the arm under-steps each control period). An explicit
control_substeps must be a positive integer - 0/negative/float
is rejected with a structured error instead of collapsing to a single
physics step, which would reinstate that same no-op.
async_rtc (default False) opts into overlapping policy
inference with action-chunk execution, evaluating a chunk-emitting
policy under the realistic control latency it faces in deployment.
It is forwarded to :meth:PolicyRunner.evaluate; the default keeps
the success-rate synchronous and bit-stable. It must be a boolean -
a value of any other type is reported as a structured caller error
rather than read by truthiness, since a truthy "false" would
otherwise evaluate under the latency it reads as declining, and a
success rate measured that way is not the one the caller asked for.
rtc_inference_timeout_s
bounds each async inference (structured error instead of a hung
rollout). For benchmark-style latency masking use
:meth:run_policy (async_rtc=...).
wbc_install_torque_control is the posture :meth:run_policy
declares under that name, applied here through the same reader
(:meth:_install_action_controller) and checked as the same boolean
domain: True (default) installs the torque shim a
:class:~strands_robots.policies.wbc.WBCPolicy needs on a
position-servo scene for the duration of the call, then uninstalls it.
One install covers every episode - both halves of it survive the
per-episode reset. A scored rollout that drove the scene differently
from an unscored one published the difference as the policy's own
success rate, with no field saying which pipeline produced it.
on_frame is an optional (step, observation, action) -> None
hook fired per applied control step on the eval thread, immediately
after sim.send_action - the success-rate analogue of the
:meth:run_policy / :meth:evaluate_benchmark hook. step is a
monotonic index that continues across episode boundaries. Use it to
record frames or stream telemetry synchronously on the eval thread
(e.g. paired with start_cameras_recording_synchronous) so a
daemon-thread recorder does not race mjData mutations. A hook
exception other than CooperativeStop or
:class:~strands_robots.recording_errors.RecordingFrameError is logged
at WARN and tolerated up to max_onframe_failures consecutive
failures (default 5, the :meth:run_policy ceiling and domain),
after which the eval stops with status="error" and onframe_error
naming the last failure; a RecordingFrameError is data loss
rather than telemetry and propagates on the first occurrence, so the
caller learns the episode is incomplete instead of reading a successful
eval. Raising :class:~strands_robots.simulation.policy_runner.CooperativeStop
stops the evaluation gracefully after the episodes completed so far
(the result carries stopped_early=True and episodes_completed),
matching :meth:run_policy. That best-effort posture covers a hook that
FAILS, not one that cannot be called at all: a non-callable on_frame
is a caller error, refused up front with a structured error like
:meth:run_policy's observer, because absorbing it per frame would
return a success rate the caller's telemetry had watched none of.
n_episodes and max_steps must be positive integers and
control_frequency must be > 0; a non-positive value is
rejected with a structured error at the entry point (before
create_policy) rather than running a degenerate eval that
reports a fabricated success rate over zero/negative episodes.
policy_kwargs is the per-call goal payload forwarded verbatim to
every policy.get_actions(obs, instruction, **policy_kwargs) call,
exactly as on :meth:run_policy. Goal-conditioned providers read their
target from these well-known keys (target_velocity for WBC and other
locomotion policies; target_pose / target_joints / world_update
for cuRobo / MoveIt2 - the issue #300 contract). Without it the eval ran
such a policy with an empty goal and reported a meaningless success rate.
success_fn defaults to None. With no success_fn (and no
benchmark spec) there is no criterion by which an episode can be marked
successful, so success_rate reports a hard 0.0 for every episode
regardless of what the policy does - indistinguishable from a policy that
genuinely failed every episode. This case logs a warning and sets
success_measured=false in the returned json; pass
success_fn="contact" (or a callable) to measure real task success.
success_when is the other way to say what success IS: the same
predicate DSL as :meth:run_policy's stop_when and a benchmark
spec's success clause - {'predicate': 'body_above_z', 'body':
'cube', 'z': 0.2} or an all / any group - compiled through the
closed predicate registry and probed against the live scene before the
first episode, so a body the scene does not have is refused up front
instead of scoring every episode a miss. success_fn (the named
'contact' criterion) and success_when are alternatives; passing
both is refused. Before this the only criterion an agent-tool call could
express was 'contact', and a predicate spelled as a string
('base_beyond_x:0.5') was refused without saying what IS accepted.
video optionally records one rollout MP4 PER EPISODE so an eval can
be watched to see WHY episodes fail, not just read as an aggregate
success rate. Same dict schema as :meth:run_policy (path enables
it; fps / camera / width / height); the path is
validated and the camera probed up-front. _ep{i} is inserted into
the filename per episode (eval.mp4 -> eval_ep0.mp4,
eval_ep1.mp4, ...) so episodes never overwrite each other, and the
written files are returned in the result json video_paths. Recording
is unsupported on the benchmark (evaluate_benchmark) path.
Returns:
| Name | Type | Description |
|---|---|---|
dict[str, Any]
|
The standard agent-tool envelope |
|
dict[str, Any]
|
``{"status": "success"|"error", "content": [{"text": ...}, |
|
dict[str, Any]
|
{"json": {...}}]} |
|
dict[str, Any]
|
by scanning |
|
dict[str, Any]
|
never by a fixed index (an early caller-error return carries a |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
policy succeeded: an evaluation in which every episode failed is |
|
dict[str, Any]
|
still |
|
dict[str, Any]
|
two things that make it |
|
dict[str, Any]
|
keep (see |
|
dict[str, Any]
|
which the policy never commanded the robot at all (see |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
did and measures nothing. |
|
dict[str, Any]
|
Fields in the json block: |
|
Outcome |
dict[str, Any]
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
criterion already held at reset, before any action was applied. The |
|
dict[str, Any]
|
criterion is sampled only after an applied action, so such an episode |
|
dict[str, Any]
|
succeeds on its first step whatever the policy commands and its |
|
dict[str, Any]
|
contribution to |
|
dict[str, Any]
|
initial state rather than the policy - the mirror of |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
of reason. Usually a threshold on the wrong side of the initial state |
|
dict[str, Any]
|
(a lift height below where the object already rests). Every reported |
|
dict[str, Any]
|
figure is left as measured; |
|
dict[str, Any]
|
qualifying text ( |
|
dict[str, Any]
|
record carries its own |
|
error |
dict[str, Any]
|
domain randomisation legitimately draws initial states per |
dict[str, Any]
|
episode. |
|
dict[str, Any]
|
Commanded actions: |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
evaluation advanced), and |
|
dict[str, Any]
|
evaluation that commanded the robot at least once, and the reason |
|
dict[str, Any]
|
string when it never did. A policy call that returns an empty |
|
dict[str, Any]
|
action chunk is tolerated per step (physics advances so a |
|
dict[str, Any]
|
degenerate policy cannot hang the episode), so the two counts |
|
dict[str, Any]
|
differ whenever any call came back empty and |
|
dict[str, Any]
|
cannot tell a scored zero from an unexercised policy. When |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
scene's initial state rather than the policy and |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
refused, since some empty calls are real policy behaviour. Each |
|
dict[str, Any]
|
per-episode record in |
|
dict[str, Any]
|
|
|
Horizon |
dict[str, Any]
|
|
dict[str, Any]
|
ran with) and |
|
Recording |
dict[str, Any]
|
|
dict[str, Any]
|
evaluation, and the reason string when a per-episode dataset flush |
|
dict[str, Any]
|
failed. A failed flush closes the recorder, after which |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
stops at that episode and |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
than |
|
dict[str, Any]
|
instead of averaging over ones whose data does not exist. |
|
Physics |
dict[str, Any]
|
|
dict[str, Any]
|
and the backend's divergence report (episode, step, joint) when the |
|
dict[str, Any]
|
physics diverged and the backend reset the world mid-episode. The |
|
dict[str, Any]
|
evaluation stops there, the diverged episode is not counted and its |
|
dict[str, Any]
|
unsaved recording frames are discarded, and |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
Video |
dict[str, Any]
|
|
dict[str, Any]
|
recording was requested). |
|
dict[str, Any]
|
Policy binding: |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
means the policy silently fell back to positional camera routing or |
|
dict[str, Any]
|
to observation-derived state keys, so the robot moved on |
|
dict[str, Any]
|
meaningless inputs and the success rate measures nothing about the |
|
dict[str, Any]
|
policy. See :meth: |
|
dict[str, Any]
|
Policy load: |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
Chunk-prefetch telemetry: |
|
dict[str, Any]
|
background chunk pipeline, not the policy's RTC algorithm, which |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
one release with the same values. |
evaluate_benchmark ¶
evaluate_benchmark(benchmark_name: str, robot_name: str | None = None, policy_provider: str = 'mock', policy_config: dict[str, Any] | None = None, instruction: str = '', n_episodes: int = 1, seed: int | None = None, action_horizon: int = 8, on_frame: Callable[[int, dict[str, Any], dict[str, Any]], None] | None = None, max_onframe_failures: int | None = None, policy_kwargs: dict[str, Any] | None = None, control_frequency: float | None = None, control_substeps: int | None = None, policy_object: Policy | None = None, video: dict[str, Any] | None = None, wbc_install_torque_control: bool = True, async_rtc: bool = False, rtc_inference_timeout_s: float | None = None) -> dict[str, Any]
Run a registered :class:BenchmarkProtocol against the current sim.
Benchmark-agnostic evaluation entry point. Looks up benchmark_name
in the global benchmark registry, validates robot compatibility, and
forwards to :meth:PolicyRunner.evaluate with the spec.
max_steps comes from the benchmark (not a parameter here), so it
is validated where it is read rather than at this signature: a
benchmark declaring a horizon that is not a positive integer is
rejected with a structured error, for the same reason n_episodes
is. Both are bounds of the same nested loop, and a non-positive one
runs episodes of zero length and then reports a 0% success rate over
them.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
benchmark_name
|
str
|
Key from :func: |
required |
robot_name
|
str | None
|
Robot to evaluate, resolved by the same rule
:meth: |
None
|
policy_provider
|
str
|
Policy provider name (forwarded to
:func: |
'mock'
|
policy_config
|
dict[str, Any] | None
|
Provider-specific kwargs. |
None
|
instruction
|
str
|
Natural-language instruction for the policy. |
''
|
n_episodes
|
int
|
Number of episodes. Must be a positive integer; a zero/negative/non-int value is rejected with a structured error rather than fabricating a 0%-success report over an empty rollout loop. |
1
|
seed
|
int | None
|
Master RNG seed for per-episode reproducibility. |
None
|
action_horizon
|
int
|
How many actions to consume from each
|
8
|
on_frame
|
Callable[[int, dict[str, Any], dict[str, Any]], None] | None
|
Optional |
None
|
max_onframe_failures
|
int | None
|
The :meth: |
None
|
policy_kwargs
|
dict[str, Any] | None
|
Per-call goal payload forwarded verbatim to every
|
None
|
control_frequency
|
float | None
|
Target Hz for |
None
|
control_substeps
|
int | None
|
Explicit physics substeps per action, overriding
the |
None
|
policy_object
|
Policy | None
|
An already-built :class: |
None
|
video
|
dict[str, Any] | None
|
Optional per-episode rollout MP4 config (same dict schema as
:meth: |
None
|
wbc_install_torque_control
|
bool
|
The same posture :meth: |
True
|
async_rtc
|
bool
|
Accepted for the same shape as :meth: |
False
|
rtc_inference_timeout_s
|
float | None
|
The async prefetch deadline of
:meth: |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
dict[str, Any]
|
Standard status dict. On success, carries per-episode cumulative |
|
dict[str, Any]
|
reward + aggregate success_rate / avg_reward / avg_steps in the |
|
dict[str, Any]
|
JSON payload, plus |
|
dict[str, Any]
|
when |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
carries the reason when a per-episode dataset flush failed, in |
|
dict[str, Any]
|
which case the benchmark stops at that episode and |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
way. |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
divergence report when the physics diverged mid-episode; the |
|
dict[str, Any]
|
benchmark stops the same way :meth: |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
reports under those names, and each per-episode record carries its |
|
dict[str, Any]
|
own |
|
dict[str, Any]
|
chunk is tolerated per step, so the counts differ whenever any call |
|
dict[str, Any]
|
came back empty; when |
|
dict[str, Any]
|
never commanded the robot, |
|
dict[str, Any]
|
describe the scene's initial state rather than the policy, and |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
dict[str, Any]
|
criterion already held at reset, before any action was applied. The |
|
dict[str, Any]
|
criterion is sampled only after an applied action, so such an episode |
|
dict[str, Any]
|
succeeds on its first step whatever the policy commands and its |
|
dict[str, Any]
|
contribution to |
|
dict[str, Any]
|
initial state rather than the policy - the mirror of |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
of reason. Usually a threshold on the wrong side of the initial state |
|
dict[str, Any]
|
(a lift height below where the object already rests). Every reported |
|
dict[str, Any]
|
figure is left as measured; |
|
dict[str, Any]
|
qualifying text ( |
|
dict[str, Any]
|
record carries its own |
|
error |
dict[str, Any]
|
domain randomisation legitimately draws initial states per |
dict[str, Any]
|
episode. |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
reports under those names. |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
|
way |
dict[str, Any]
|
a clause already satisfied at reset ends the episode on its |
dict[str, Any]
|
first step whatever the policy commands, so |
|
dict[str, Any]
|
a hard 0.0 for it. It costs more than the success mirror, because the |
|
dict[str, Any]
|
eval loop reads |
|
dict[str, Any]
|
is scored a failure with the success criterion never consulted, and in |
|
dict[str, Any]
|
the report it is indistinguishable from a policy that immediately did |
|
dict[str, Any]
|
something catastrophic. Usually a threshold on the wrong side of the |
|
dict[str, Any]
|
initial state (a fall height above where the object already rests, or |
|
dict[str, Any]
|
a base-collapse height above the robot's spawned stance). Every |
|
dict[str, Any]
|
reported figure is left as measured; |
|
dict[str, Any]
|
the qualifying text ( |
|
dict[str, Any]
|
per-episode record carries its own |
|
dict[str, Any]
|
count is not an error, for the reason the success mirror's is not. |
|
dict[str, Any]
|
Only this route reports it: :meth: |
|
dict[str, Any]
|
and has no failure criterion to sample. |
list_benchmarks ¶
Enumerate registered benchmarks.
Returns a standard status dict whose JSON payload contains the
:func:~strands_robots.simulation.benchmark.list_benchmarks
metadata snapshot. Safe to call from any backend; the registry is
engine-agnostic.
register_benchmark_from_file ¶
Load a declarative benchmark spec from disk and register it.
Wraps :func:strands_robots.simulation.benchmark_spec.register_benchmark_from_file
so agents can author benchmarks as YAML / JSON at runtime. Parsing
errors surface as structured error dicts rather than exceptions.
register_builtin_benchmarks ¶
Register the built-in benchmark specs shipped with strands_robots.
Wraps :func:strands_robots.simulation.builtin_benchmarks.register_builtin_benchmarks
so the shipped specs become discoverable via :meth:list_benchmarks and
runnable via :meth:evaluate_benchmark without hand-authoring a spec
file. Ships a canonical velocity-tracking locomotion benchmark
(go2_walk_forward) composed from the floating-base predicate/reward
DSL. Opt-in and idempotent (mirrors the on-demand LIBERO suite
registration); importing strands_robots performs no registry mutation.
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
A status dict whose JSON payload carries the |
dict[str, Any]
|
benchmark names now available to :meth: |
load_scene ¶
Load a complete scene from file. Override per backend.
randomize ¶
Apply domain randomization.
Concrete backends define their own parameter signatures. Because this
base signature is **kwargs-typed, an override inherits a sink that
would swallow any keyword it does not declare; backends must reject the
residual keys (see :func:unknown_kwargs_error) so a misspelled axis
cannot report success while leaving that axis untouched.
Override per backend.
set_obs_noise ¶
Configure additive sensor noise on observations.
Models real-sensor measurement noise (joint encoders, camera frames)
so policies are not trained on noise-free observations. Concrete
backends define their own parameter signatures and, as for
:meth:randomize, must reject keywords they do not declare rather than
let this **kwargs-typed signature swallow them. Override per backend.
get_contacts ¶
Get contact information. Override per backend.
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
The agent-tool envelope -- |
dict[str, Any]
|
whose |
dict[str, Any]
|
per-contact records. The payload lives in that block, not on the |
dict[str, Any]
|
envelope itself, so a caller reading |
dict[str, Any]
|
directly always misses. The predicate DSL's |
dict[str, Any]
|
factories (see |
dict[str, Any]
|
mod: |
dict[str, Any]
|
readers; |
dict[str, Any]
|
meth: |
dict[str, Any]
|
shares them. |
Raises:
| Type | Description |
|---|---|
NotImplementedError
|
Backends that expose no contact list. |
get_frame ¶
get_frame(camera_name: str = 'default', width: int | None = None, height: int | None = None) -> tuple[np.ndarray, np.ndarray | None]
Render a camera to raw (rgb, depth) ndarrays.
The numeric-array counterpart of :meth:render (which wraps pixels in
the agent-tool PNG envelope). In-process consumers -- the hybrid
compositor, dataset recorders, video writers -- use this to get pixels
without a PNG round-trip.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
camera_name
|
str
|
name of a camera previously added via |
'default'
|
width
|
int | None
|
image width in pixels; |
None
|
height
|
int | None
|
image height in pixels; |
None
|
Returns:
| Type | Description |
|---|---|
ndarray
|
|
ndarray | None
|
|
tuple[ndarray, ndarray | None]
|
backends with no depth path (Newton). Backends must never |
tuple[ndarray, ndarray | None]
|
substitute silently wrong pixels -- failures raise. |
Raises:
| Type | Description |
|---|---|
KeyError
|
unknown camera name. |
ValueError
|
invalid render dimensions. |
RuntimeError
|
no world / renderer unavailable / backend render failure. |
NotImplementedError
|
backend has no raw-frame path. |
get_camera_params ¶
get_camera_params(camera_name: str = 'default', width: int | None = None, height: int | None = None) -> CameraParams
Return pinhole intrinsics/extrinsics for a named camera.
The returned :class:strands_robots.rendering.CameraParams carries
the intrinsic matrix K (pixels), the world-from-camera SE(3) pose
T_world_cam in the OpenGL optical convention (+X right, +Y up,
-Z forward), the image size, and the clip planes. Backends whose
native camera basis differs (e.g. Isaac's USD camera prim) apply the
fixed basis correction here so consumers never see a backend-specific
frame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
camera_name
|
str
|
name of a camera previously added via |
'default'
|
width
|
int | None
|
image width to compute |
None
|
height
|
int | None
|
image height to compute |
None
|
Raises:
| Type | Description |
|---|---|
KeyError
|
unknown camera name. |
ValueError
|
a camera whose projection no pinhole |
RuntimeError
|
no world created. |
NotImplementedError
|
backend has no camera-params path. |
get_world_point ¶
get_world_point(camera_name: str = 'default', pixels: Sequence[Sequence[SupportsFloat]] | None = None, width: int | None = None, height: int | None = None) -> dict[str, Any]
Ground image pixels to metric world coordinates via the depth buffer.
The perception half of deployment-shaped grounding (Harness VLA,
arXiv:2607.08448, Appendix E.2): instead of reading privileged object
poses (:meth:get_body_state -- sim-only oracle truth), the agent
picks pixels on the visible surface of the target in the RGB frame and
this call unprojects each one through the pixel-aligned metric depth
buffer -- p_cam = depth * K^-1 @ [u, v, 1] in the OpenGL optical
frame, then p_world = T_world_cam @ p_cam. The same call shape
works on hardware with an RGB-D camera, so grounding built on it
transfers.
Guidance for agents (the paper's localization rule):
- Render the camera first (
render/get_frame) and pick pixels ON the visible surface of the target object. - Avoid rims, edges, reflections, transparent surfaces, and background pixels -- depth there is unstable or belongs to something else.
- Sample SEVERAL pixels on the same surface (typically 3-9): the
returned
pointis the median over the valid samples, which rejects stray outliers. The median is PER-COMPONENT, so on a strongly tilted surface the combined[x, y, z]may lie on no single sampled point - treat it as a robust surface estimate, not as one ofpoints. - Pixels with no depth (background / far plane) are dropped, not
zero-filled; check
n_validagainst the count you sent. - Re-localize after any robot, camera, or object motion -- world points are snapshots, not tracks.
Depth samples are treated as z-depth (distance along the optical
axis), the convention every in-tree backend emits. Pixels are indexed
[u, v] with u the column from the left and v the row from
the top; the unprojection uses the pixel center (u + 0.5, v + 0.5).
Atomicity: when the backend exposes an engine lock (self._lock,
all in-tree backends), the frame render and the camera-params read
happen under it, so a concurrent scene mutation cannot slip between
the two. All failures return a structured error dict -- this is a
tool-envelope method and never raises.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
camera_name
|
str
|
a camera previously added via |
'default'
|
pixels
|
Sequence[Sequence[SupportsFloat]] | None
|
non-empty list of |
None
|
width
|
int | None
|
image width; |
None
|
height
|
int | None
|
image height; |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
On success ``{"status": "success", "content": [{"text": ...}, |
dict[str, Any]
|
{"json": {"point": [x, y, z], "points": [...], "n_valid": int, |
dict[str, Any]
|
"n_requested": int, "camera": str, "width": int, "height": int}}]}`` |
dict[str, Any]
|
where |
dict[str, Any]
|
samples and |
dict[str, Any]
|
( |
dict[str, Any]
|
metric-depth path (Newton), all-invalid pixel sets, out-of-bounds |
dict[str, Any]
|
pixels, malformed input, and a failed frame render or |
dict[str, Any]
|
camera-params read all return |
dict[str, Any]
|
|
dict[str, Any]
|
two backend reads reporting distinguishable text so a caller |
dict[str, Any]
|
knows which one failed. The camera-params read can fail on input |
dict[str, Any]
|
this call already accepted and a frame it already rendered -- |
dict[str, Any]
|
most notably a camera whose projection no pinhole |
dict[str, Any]
|
represent, such as MuJoCo's orthographic free camera, which |
dict[str, Any]
|
renders normally but has no intrinsics. So check |
dict[str, Any]
|
rather than inferring success from a valid pixel set. |
describe ¶
Return a machine-readable summary of this engine's live contract.
Agents should call this first to learn what robots exist, what cameras are attached, and the signatures of the methods most commonly needed -- in a single call, instead of guessing method names.
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
Plain dict with keys: robots, capabilities, cameras, methods, note. |
cleanup ¶
Release all resources.
Called on context exit, and best-effort from :meth:__del__ for an
engine whose __init__ ran to completion. Implementations are
written against a fully-constructed instance: a caller whose
__init__ raised part-way must release whatever it acquired itself
rather than relying on the finalizer.
Capabilities¶
Capability vocabulary a :class:~strands_robots.simulation.base.SimEngine backend declares.
The names form a closed set; a backend may add "<vendor>:<name>" extras. A
missing capability is reported with :data:UNSUPPORTED_BY_BACKEND, which no
operator grant can lift, so it is not a continuable refusal code. A declaration
must name the four :data:CORE_CAPABILITIES; an undeclared backend is credited
with all eight names of :data:DEFAULT_CAPABILITIES. Pure stdlib, and it
imports nothing from the package, so the engine is described by
:class:CapabilityReporter rather than by the base class.
CapabilityNotSupported ¶
Bases: NotImplementedError
A member was called on a backend that lacks the capability it needs.
Raised only by list-returning members; others return :func:unsupported_result.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
capability
|
str
|
The missing capability name. |
required |
member
|
str
|
The |
required |
CapabilityReporter ¶
Bases: Protocol
Anything that reports its capabilities, such as a SimEngine.
ManipulationOptional ¶
Mix in before SimEngine for a backend with no joints, objects, rendering or rollout.
Declares only the four core capabilities and supplies the manipulation-only
abstract members as standard refusals, so such a backend implements only
what it can honour. It is for backends without these features, not for
layering over a full backend. To support one of them, override the member
and add its name to CAPABILITIES; declaring the name while the refusal
is still inherited raises TypeError at class creation.
add_object ¶
add_object(name: str, shape: str = 'box', position: list[float] | None = None, orientation: list[float] | None = None, size: list[float] | None = None, color: list[float] | None = None, mass: float = 0.1, is_static: bool | None = None, mesh_path: str | None = None, material: dict[str, Any] | None = None) -> dict[str, Any]
Refuse with :func:unsupported_result: no :data:OBJECTS capability. Arguments are ignored.
remove_object ¶
Refuse with :func:unsupported_result: no :data:OBJECTS capability. Arguments are ignored.
render ¶
render(camera_name: str = 'default', width: int | None = None, height: int | None = None) -> dict[str, Any]
Refuse with :func:unsupported_result: no :data:RENDER capability. Arguments are ignored.
robot_joint_names ¶
Raise, since an empty list would read as a robot with no joints.
Raises:
| Type | Description |
|---|---|
CapabilityNotSupported
|
Always; this backend has no :data: |
check_capabilities ¶
check_capabilities(sim: CapabilityReporter, required: Iterable[str], *, caller: str) -> dict[str, Any] | None
Check that sim has every capability in required.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sim
|
CapabilityReporter
|
The engine to check; any object with a |
required |
required
|
Iterable[str]
|
Capability names the caller needs. |
required |
caller
|
str
|
The member or tool doing the check, named in the result. |
required |
Returns:
| Type | Description |
|---|---|
dict[str, Any] | None
|
|
dict[str, Any] | None
|
whose |
Raises:
| Type | Description |
|---|---|
TypeError
|
|
unsupported_result ¶
Build the standard error result for a member the backend does not support.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
capability
|
str
|
The missing capability name. |
required |
member
|
str
|
The |
required |
backend
|
str
|
The backend's name, usually its class name. |
required |
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
A tool-result dict with |
dict[str, Any]
|
|
World model¶
Dataclasses for simulation state.
These dataclasses provide a backend-independent typed state representation consumed by simulation engine implementations (e.g. MuJoCo, Isaac Sim, PyBullet).
They enable
- Type-safe state tracking across simulation steps.
- Serialisation for checkpoints and trajectory recording.
- A backend-independent interface for agent tools.
They are defined alongside the SimEngine ABC because its method
signatures reference them (e.g. create_world() → SimWorld).
SimWorld
dataclass
¶
SimWorld(robots: dict[str, SimRobot] = dict(), objects: dict[str, SimObject] = dict(), cameras: dict[str, SimCamera] = dict(), timestep: float = 0.002, gravity: list[float] = (lambda: [0.0, 0.0, -9.81])(), ground_plane: bool = True, terrain: str | None = None, terrain_difficulty: float = 1.0, status: SimStatus = SimStatus.IDLE, sim_time: float = 0.0, step_count: int = 0, _model: Any = None, _data: Any = None, _backend_state: dict[str, Any] = dict(), _checkpoints: dict[str, Any] = dict(), _recompile_generation: int = 0)
Complete simulation world state.
Backend-independent state with engine-specific internals kept in three escape hatches, each with a distinct role so backend implementers know which to use:
_model: the physics engine's core model handle - the single compiled/loaded representation of the scene (e.g.mujoco.MjModel, Isaac'sScene, PyBullet's body registry). Every backend has one._data: the physics engine's core simulation state handle - the mutable per-step state companion to_model(e.g.mujoco.MjData, Isaac'sWorld). Every backend has one._backend_state: a catch-all dict for everything else the backend needs to persist - generated XML, temp dirs, recording buffers, caches, etc. Prefer this over adding new fields here.
All three are typed Any/dict so nothing leaks engine-specific
types into this base module.
SimRobot
dataclass
¶
SimRobot(name: str, urdf_path: str, position: list[float] = (lambda: [0.0, 0.0, 0.0])(), orientation: list[float] = (lambda: [1.0, 0.0, 0.0, 0.0])(), data_config: str | None = None, body_id: int = -1, joint_ids: list[int] = list(), joint_names: list[str] = list(), actuator_ids: list[int] = list(), namespace: str = '', home_qpos: dict[str, list[float]] = dict(), home_actuators: dict[str, tuple[float, list[float]]] = dict(), policy_running: bool = False, policy_stops: int = 0, policy_claim_stops: int | None = None, policy_steps: int = 0, policy_instruction: str = '', mesh: Any = None, peer_id: str = '', _world: Any = None, _sim_parent: Any = None)
A robot instance within the simulation.
mesh / peer_id: when the parent Simulation is
itself attached to a Zenoh mesh, every robot added via add_robot
auto-joins as its own peer so the agent can address it directly
(e.g. robot_mesh tell target=<peer_id>) instead of having to talk to
the sim container and then route by robot name. Both fields stay
None / "" for stand-alone sims that are not on a mesh.
request_policy_stop ¶
Cooperatively stop this robot's rollout, durably, and report what was in flight.
A stop is a fact about the rollout and not only a flag: lowering
policy_running alone was overwritten by a worker that had not yet
reached its first frame, so the stop was reported and then discarded and
the rollout ran to full duration (#2833). Moving policy_stops puts
the launcher's claim out of date, which is what stops the worker raising
the flag back over this stop.
EVERY stop path goes through here rather than assigning
policy_running = False - the stop_policy action, remove_robot,
teardown, and the Device Connect stop / emergency-stop handlers - so they
cannot drift to different answers about whether a rollout was halted.
Returns:
| Type | Description |
|---|---|
bool
|
Whether this robot's |
bool
|
arrived. A caller that also tracks rollout Futures should treat its |
bool
|
own registry as part of the answer (see |
bool
|
meth: |
SimObject
dataclass
¶
SimObject(name: str, shape: str, position: list[float] = (lambda: [0.0, 0.0, 0.0])(), orientation: list[float] = (lambda: [1.0, 0.0, 0.0, 0.0])(), size: list[float] = (lambda: [0.05, 0.05, 0.05])(), color: list[float] = (lambda: [0.5, 0.5, 0.5, 1.0])(), mass: float = 0.1, mesh_path: str | None = None, material: dict[str, Any] | None = None, body_id: int = -1, is_static: bool = False, _original_position: list[float] = list(), _original_color: list[float] = list())
An object in the simulation scene.
SimCamera
dataclass
¶
SimCamera(name: str, position: list[float] = (lambda: [1.0, 1.0, 1.0])(), target: list[float] = (lambda: [0.0, 0.0, 0.0])(), fov: float = 60.0, width: int = 640, height: int = 480, camera_id: int = -1, origin_robot: str = '', parent_body: str = '')
A camera in the simulation.
origin_robot: when the camera was discovered inside a
robot's URDF during add_robot, this is set to the robot's name so the
scene builder knows NOT to re-add the camera at the top level (it'll be
re-introduced via spec.attach(robot_spec)). For user-added cameras
(via the add_camera tool action) this stays empty.
A discovered camera belongs to exactly one robot: the one whose namespace
prefixes name. Removing a robot removes its cameras and only its
cameras, so origin_robot must never name a robot outside name's
namespace - otherwise the wrong robot's departure strands or drops the
entry.
SimStatus ¶
Bases: Enum
Simulation execution status.
TrajectoryStep
dataclass
¶
TrajectoryStep(timestamp: float, sim_time: float, robot_name: str, observation: dict[str, Any], action: dict[str, Any], instruction: str = '')
A single step in a recorded trajectory.
Run-policy observers¶
Read-only rollout events for :meth:PolicyRunner.run.
Why this is a second lane rather than a use of on_frame¶
on_frame looks like the observation seam and is not one. It is owned by the
backend for the duration of a rollout: MuJoCo's hook raises
:class:~strands_robots.simulation.policy_runner.CooperativeStop so
stop_policy can interrupt, appends the trajectory mirror, publishes mesh step
telemetry and drives the LeRobot dataset recorder; Isaac's and Newton's do the
recording half of the same job. There is exactly one of them -
:meth:~strands_robots.simulation.base.SimEngine.run_policy obtains it from
_make_run_policy_hook and does not accept one from the caller - so a consumer
that supplied its own would not add observation, it would silently remove
cancellation and recording.
These events are therefore emitted beside that hook, and deliberately report the three things the hook's own signature cannot carry:
- what the backend answered -
send_action's per-key verdict, normalised to :data:ActionResolutionso a consumer never parses a backend envelope; - how old the observation was -
observation_age_stepsis the authoritative number of completed rollout action attempts since the snapshot was sampled;observation_is_chunk_reusedonly says that the same chunk-start snapshot is being used by a later action in that chunk; - what the legacy hook did - including the step it aborted on, which the
legacy step accounting excludes (see :class:
RunPolicyStep).
The payload-ownership rule¶
RunPolicyStep.observation and RunPolicyStep.action are borrowed: the
same objects the legacy hook received, not copies. Copying them per step would
put an image-sized deep copy on the control path of every rollout that enables
the lane, which is the opposite of what an observability lane should cost. So the
contract is placed on the consumer instead, and it is narrow:
- Treat both as read-only. A backend may reuse the same buffers next step, and the dataset recorder reads them after you do.
- Do not retain them past the call. Snapshot the few fields you need (synchronously, inside the callback) if you intend to hand them to a queue, a thread or a socket.
An event is dispatched synchronously on the rollout thread. That makes the
lane cheap and ordered, and it means a consumer that blocks, blocks the robot.
Ordinary exceptions and CooperativeStop are contained - those raises never
change the rollout's outcome, and never reach the on_frame consecutive-failure
watchdog, which exists for a recorder losing dataset frames (GH #117) rather than
for a visualiser that cannot draw. Process-control and cancellation
BaseException classes propagate after terminal dispatch is attempted. Containment
is not isolation: this is telemetry, not a sandbox. Keep the callback short and
non-blocking.
Ordering guarantees¶
event_seq is dense and 0-based within one run_id, so a gap is
observable. monotonic_ns is the ordering clock (a date -s or an NTP
correction cannot move it); utc_ns is derived from a single rollout anchor so
a wall-clock label never reorders the stream.
Scope¶
Single-policy simulation rollouts through :meth:PolicyRunner.run and every
episode of :meth:~strands_robots.simulation.base.SimEngine.run_policy (one
lifecycle and run_id per episode). eval_policy, evaluate_benchmark,
run_multi_policy and hardware carry no observer yet - they are separate
loops with different step semantics, and claiming them here would promise
coverage this module does not have.
Example::
from strands_robots.simulation.observers import RunPolicyStep
def watch(event):
if isinstance(event, RunPolicyStep) and event.action_resolution != "full":
print(event.applied_action_index, event.unresolved_action_keys)
sim.run_policy(robot_name="arm", observer=watch)
RunPolicyEvent
module-attribute
¶
StoppedReason
module-attribute
¶
ActionResolution
module-attribute
¶
RunPolicyStarted
dataclass
¶
RunPolicyStarted(schema_version: int, run_id: str, event_seq: int, monotonic_ns: int, utc_ns: int, robot_name: str, policy: str, instruction: str, control_frequency: float, action_horizon: int, total_steps: int, async_rtc: bool)
Opens a rollout. Emitted once, after setup, before the first observation.
Emitted only for a rollout that actually begins: a request refused in pre-flight (a bad horizon, an unusable seed, a video path that cannot be opened) raises or returns before this event, so no lifecycle is opened and none has to be closed.
Attributes:
| Name | Type | Description |
|---|---|---|
schema_version |
int
|
:data: |
run_id |
str
|
Identifies this rollout. Every event of one
:meth: |
event_seq |
int
|
|
monotonic_ns |
int
|
Ordering clock, from :func: |
utc_ns |
int
|
Wall-clock label for the same instant, derived from the rollout's
single |
robot_name |
str
|
Robot being driven. |
policy |
str
|
Class name of the driving policy (e.g. |
instruction |
str
|
Natural-language instruction forwarded to the policy. |
control_frequency |
float
|
Target Hz of the control loop. |
action_horizon |
int
|
Max actions consumed per policy call, as requested. The
effective chunk may be longer when the policy declares a larger
|
total_steps |
int
|
Step budget resolved for this rollout. |
async_rtc |
bool
|
Whether inference is overlapped with action execution. This
is the resolved value, so a rollout that auto-detected a
chunk-emitting policy reports |
RunPolicyStep
dataclass
¶
RunPolicyStep(schema_version: int, run_id: str, event_seq: int, monotonic_ns: int, utc_ns: int, applied_action_index: int, legacy_step_index: int, observation: dict[str, Any], action: Any, observation_is_chunk_reused: bool, observation_age_steps: int, action_resolution: ActionResolution, applied_action_keys: tuple[str, ...], unresolved_action_keys: tuple[str, ...], elapsed_s: float, sim_time_s: float | None, legacy_hook_outcome: LegacyHookOutcome)
One completed send_action call.
Emitted after send_action has returned and after the legacy on_frame
hook has run, whatever that hook did. The backend's physical state is known
to have advanced only when its result says so; on a coarse atomic refusal the
event reports action_resolution="unknown" rather than inventing applied
or unresolved keys.
The hook runs before the legacy step_count increments. Consequently
:attr:applied_action_index and :attr:legacy_step_index identify the same
zero-based action, including the action whose hook cancels or loses a dataset
frame. The abort is identified by :attr:legacy_hook_outcome. Terminal
accounting is intentionally different: :attr:RunPolicyEnded.applied_actions
counts calls made to send_action, while
:attr:RunPolicyEnded.legacy_steps_used excludes a hook-aborted final step.
Attributes:
| Name | Type | Description |
|---|---|---|
schema_version |
int
|
:data: |
run_id |
str
|
The rollout this step belongs to. |
event_seq |
int
|
Dense, monotonic position in the rollout's event stream. |
monotonic_ns |
int
|
Ordering clock, sampled after |
utc_ns |
int
|
Wall-clock label for the same instant. |
applied_action_index |
int
|
0-based count of |
legacy_step_index |
int
|
The index this step's |
observation |
dict[str, Any]
|
Borrowed pre-action observation - the same object the legacy hook received. Read-only; do not retain past the call. See the module docstring. |
action |
Any
|
Borrowed action sent to the backend. Usually a
|
observation_is_chunk_reused |
bool
|
|
observation_age_steps |
int
|
Authoritative nonnegative count of completed
rollout action attempts since |
action_resolution |
ActionResolution
|
A :data: |
applied_action_keys |
tuple[str, ...]
|
Keys that drove an actuator this step. |
unresolved_action_keys |
tuple[str, ...]
|
Keys the backend could not absorb. Empty on the
success path and on |
elapsed_s |
float
|
Seconds since the rollout's monotonic start. |
sim_time_s |
float | None
|
Backend simulation clock after the action, when the backend
exposes one cheaply; |
legacy_hook_outcome |
LegacyHookOutcome
|
A :data: |
RunPolicyEnded
dataclass
¶
RunPolicyEnded(schema_version: int, run_id: str, event_seq: int, monotonic_ns: int, utc_ns: int, outcome: RunPolicyOutcome, stopped_reason: StoppedReason, applied_actions: int, legacy_steps_used: int, action_errors: int, elapsed_s: float, error_type: str | None, error_message: str | None, observer_failures: int = 0)
Closes a rollout. Emitted once, if and only if :class:RunPolicyStarted was.
Attempted on every exit path a started rollout can take - budget exhausted,
predicate fired, cooperative stop, or any error - so a consumer can always
pair an open with a close. The one thing it cannot survive is the process
dying under it (SIGKILL, OOM, a consumer that blocks forever), which is
why this lane is telemetry rather than a durable record.
Attributes:
| Name | Type | Description |
|---|---|---|
schema_version |
int
|
:data: |
run_id |
str
|
The rollout being closed. |
event_seq |
int
|
Final position in the rollout's event stream. |
monotonic_ns |
int
|
Ordering clock at close. |
utc_ns |
int
|
Wall-clock label for the same instant. |
outcome |
RunPolicyOutcome
|
|
stopped_reason |
StoppedReason
|
A :data: |
applied_actions |
int
|
Total |
legacy_steps_used |
int
|
The rollout's own |
action_errors |
int
|
Steps whose |
elapsed_s |
float
|
Rollout duration, measured on the monotonic clock. |
error_type |
str | None
|
Exception class name when |
error_message |
str | None
|
Exception message when |
observer_failures |
int
|
Contained failures before this terminal dispatch. The returned result payload is authoritative and also includes a contained failure raised while consuming this Ended event itself. |
Models and assets¶
Robot model resolution - URDF registry + asset manager.
Bridges the robot registry with actual URDF/MJCF files on disk.
Resolution order for :func:resolve_model:
1. User-registered URDFs (:func:register_urdf)
2. URDF search paths (STRANDS_ASSETS_DIR, CWD, etc.)
3. Asset manager (robot_descriptions - fallback for standard robots)
resolve_model ¶
Resolve a robot name or data_config to an MJCF/URDF model path.
Resolution order (local assets take priority): 1. User-registered URDFs (custom user registrations) 2. URDF search paths (STRANDS_ASSETS_DIR, CWD, etc.) 3. Asset manager (robot_descriptions - fallback for standard robots)
Step 3 fetches an asset that is not on disk - the right default for a caller
about to load the model. allow_download=False hands the same decline to
:func:~strands_robots.assets.manager.resolve_model_path, so a caller that
reports on assets reads the disk and reaches neither the network nor the
robot_descriptions import that clones on a cold cache. Steps 1 and 2 are
filesystem reads either way.
resolve_urdf ¶
Resolve a data_config name to a URDF file path.
Also checks the registry's legacy_urdf field - a backward-compatible
path for robots that were registered before the MJCF asset system
was introduced (e.g. robots originally configured with raw URDF paths).
register_urdf ¶
Register a URDF/MJCF file for a data_config name.
list_registered_urdfs ¶
List all registered URDF mappings and their resolved paths.
list_available_models ¶
List all available robot models (Menagerie + custom).
Both halves are always reported. The asset-manager table alone used to be
returned whenever the asset manager was importable - which is every normal
install - so a caller who had just registered an asset with
:func:register_urdf was told by the discovery surface that it did not
exist, while :func:resolve_urdf resolved it and add_robot spawned it.
The registered section is omitted entirely when nothing is registered, so a
default install's listing is unchanged.
Returns:
| Name | Type | Description |
|---|---|---|
str
|
The built-in robot table, followed by a |
|
when |
str
|
func: |
str
|
only the registered section is available, so that is returned alone. |
Predicates and benchmarks¶
Named-predicate library for declarative :class:BenchmarkProtocol specs.
Each entry in :data:PREDICATE_REGISTRY is a factory (**kwargs) -> callable
where the returned callable takes a :class:SimEngine and returns either
bool (for success/failure predicates) or float (for reward terms).
The registry is a closed set - the YAML/JSON loader in
:mod:strands_robots.simulation.benchmark_spec refuses predicates whose
name is not in this registry, so spec files are safe to parse from
untrusted / LLM-authored input. No eval is ever called. User-defined
predicates must be registered programmatically via :func:register_predicate
before loading the spec.
Predicates are backend-aware but not backend-specific: they exclusively call
SimEngine methods (abstract) or probe for MuJoCo-only methods via
getattr and return a safe fallback (False / 0.0) when the
backend does not support them. A predicate that silently evaluates to
False because of an unimplemented backend call is a bug in the
predicate, not the benchmark - file an issue.
Contact predicates count a geom pair only when the physics engine reports it
as a real touch. get_contacts also lists pairs inside the detection range
that carry no force -- see :func:contact_is_active -- and counting those
would make contact_any / contact_between / grasped /
body_on(require_contact=True) fire for bodies that are visibly apart.
When the backend does support a lookup but the referenced body /
joint name cannot be resolved (almost always a spec typo), the term still
degrades to a constant (False / 0.0) but the offending name is logged
once at WARNING (see :func:_warn_unresolved), so a broken spec surfaces
instead of silently preventing episode success or emitting a dead reward.
Available predicates (bool):
body_above_z(body, z)
body_below_z(body, z)
joint_above(joint, value)
joint_below(joint, value)
distance_less_than(body_a, body_b, threshold)
inside_region(body, min, max)
contact_between(geom_a, geom_b)
contact_any()
body_on(body_a, body_b, z_offset=0.02, xy_tol=0.15, require_contact=False)
body_inside(body, container, xy_tol=0.15, z_tol=0.15)
particles_inside(particles, container, min_fraction=1.0, xy_tol=0.15, z_tol=0.15)
particles_spilled(particles, containers, max_spilled=0, xy_tol=0.15, z_tol=0.15)
body_upright(body, tol=0.15)
grasped(body, gripper_prefix)
base_tipped(tol=0.15, robot=None)
base_below_z(z, robot=None)
base_beyond_x(x, robot=None)
base_beyond_y(y, robot=None)
base_yaw_beyond(yaw, robot=None)
Available reward terms (float):
distance_neg(body_a, body_b, weight=1.0)
joint_progress(joint, target, weight=1.0)
particles_inside_fraction(particles, container, xy_tol=0.15, z_tol=0.15, weight=1.0)
base_velocity(vx=0.0, vy=0.0, wz=0.0, weight=1.0, robot=None)
base_velocity_tracking(vx=0.0, vy=0.0, wz=0.0, lin_weight=1.0, ang_weight=0.5, tracking_sigma=0.25, robot=None)
base_height(target, weight=1.0, robot=None)
base_orientation(weight=1.0, robot=None)
base_lin_vel_z(weight=1.0, robot=None)
base_ang_vel_xy(weight=1.0, robot=None)
staged_reward(stages)
constant(value)
Register custom predicates with :func:register_predicate.
make_predicate ¶
Instantiate a predicate from its name + kwargs.
This is the single entry point the DSL loader uses - it never touches
eval or exec. Unknown names produce a ValueError listing
the valid set; a keyword the factory does not take, or a required one
that is missing, produces a ValueError naming the keywords the
predicate accepts (read from the factory's signature), so no surface
ever shows the factory's own TypeError text.
Every numeric kwarg is held to a finite domain here - a tolerance kwarg
additionally to a non-negative one and a heading kwarg to the measurable
(-pi, pi) range - and every kwarg naming a scene entity to a non-empty
string, rather than in the
spec compiler, because this is the only choke point every predicate call
passes through: staged_reward builds its per-stage reward /
advance_when calls by calling back into this function, so a guard in
:func:~strands_robots.simulation.benchmark_spec._compile_call would
leave nested stage calls unchecked. See :func:_kwarg_domain_error for
what a non-finite threshold or weight does when it compiles.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
Predicate name. Must be registered in :data: |
required |
**kwargs
|
Any
|
Forwarded verbatim to the factory. |
{}
|
Returns:
| Type | Description |
|---|---|
Callable[[SimEngine], Any]
|
A callable |
Callable[[SimEngine], Any]
|
predicate. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
register_predicate ¶
Register a user-defined predicate factory.
Must be called before loading a spec that references name. Factories
registered at runtime are NOT sandboxed - by registering, you opt into
running the factory with kwargs parsed from the spec. Only register
predicates from trusted code paths; anything LLM-authored should use the
built-in DSL exclusively.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
Predicate name used in spec files. Must not shadow a built-in. |
required |
factory
|
PredicateFactory
|
Callable that takes DSL kwargs and returns a predicate
|
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
TypeError
|
If |
Benchmark-agnostic evaluation protocol for any SimEngine.
Every standard benchmark (LIBERO, Meta-World, RoboSuite, ManiSkill, user-authored tasks) has a different notion of "what a task is" - sparse-success, dense-reward, procedural scenes, BDDL predicates, hardcoded robots, etc. The correct abstraction is the protocol the eval loop calls into, not a benchmark-specific schema.
:class:BenchmarkProtocol is that protocol. Each adapter implements a handful of
lifecycle hooks (on_episode_start, on_step, is_success, is_failure)
and declares the robots it is compatible with. The evaluation loop
(:meth:~strands_robots.simulation.policy_runner.PolicyRunner.evaluate) drives
the protocol without knowing anything about the underlying benchmark.
An adapter that needs a heavyweight simulator declares it in an optional
extra, so the core package stays dependency-free. A reference :class:DeclarativeBenchmark
shipped in :mod:strands_robots.simulation.benchmark_spec turns a YAML/JSON
spec into a fully functional BenchmarkProtocol instance - LLMs can author
and register benchmarks at runtime without writing Python code.
Registry: a module-level dict[str, BenchmarkProtocol] keyed by name,
mirroring the shape of :func:~strands_robots.simulation.model_registry.register_urdf.
Registration is idempotent-by-overwrite: re-registering the same name replaces
the previous entry and logs a warning. This matches how users iterate on a
spec file during development.
Thread safety: the registry is guarded by an internal lock so concurrent
registrations from agent threads do not race. The benchmark instances
themselves are expected to be immutable after registration - adapters that
keep per-episode state MUST put it on the rng-scoped call, not on self.
register_benchmark ¶
Register a :class:BenchmarkProtocol under name.
Idempotent-by-overwrite: re-registering the same name replaces the previous entry and logs a warning. This matches how users iterate on a spec file during development.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
String key. Must be non-empty; any other validation is up to the caller (lowercase / underscores / hyphens are all fine). |
required |
benchmark
|
BenchmarkProtocol
|
An instantiated :class: |
required |
Raises:
| Type | Description |
|---|---|
TypeError
|
If |
ValueError
|
If |
unregister_benchmark ¶
Remove a benchmark from the registry.
Returns the removed benchmark or None if it was not registered.
Primarily used by tests for cleanup; user code is rarely expected to
unregister benchmarks at runtime.
get_benchmark ¶
Return the registered benchmark or None if not found.
list_benchmarks ¶
Enumerate registered benchmarks with their metadata.
Returns a shallow-copy snapshot keyed by name. Each value is a dict
with class, supported_robots, default_robot, max_steps
- enough for an LLM to pick an appropriate benchmark without
instantiating one. Reads a snapshot under the registry lock so a
concurrent registration does not corrupt the returned dict.