Skip to content

Simulation

The engine contract, the backend factory, what SimWorld carries, and how to register your own backend.

strands_robots.simulation holds the engine contract, the backend factory and the data types a simulation returns: how backends are created, what SimWorld carries, registering your own.

Factory

Simulation factory - create_simulation() and runtime backend registration.

Mirrors the policy factory pattern: JSON-driven defaults with runtime override capability. Backends are lazy-loaded on first use.

Usage::

from strands_robots.simulation import create_simulation

# Default backend (MuJoCo)
sim = create_simulation()

# Explicit backend
sim = create_simulation("mujoco", timestep=0.001)

# GPU-native built-in backends
sim = create_simulation("isaac", num_envs=1, headless=True)
sim = create_simulation("newton")

# Custom backend (runtime-registered)
from strands_robots.simulation.factory import register_backend
register_backend("my_sim", lambda: MySimBackend, aliases=["custom"])
sim = create_simulation("custom")

Third-party packages may also register backends out-of-tree via the strands_robots.backends entry-point group (see create_simulation).

create_simulation

create_simulation(backend: str = DEFAULT_BACKEND, **kwargs: Any) -> SimEngine

Create a simulation backend instance.

This is the primary entry point for creating simulations. Backend classes are lazy-loaded on first call.

Resolution order for backend:

  1. Runtime-registered backends (see register_backend).
  2. Built-in backends (currently mujoco, newton, isaac). Built-ins always win over entry-point plugins of the same name, so a third-party plugin can never accidentally shadow a built-in backend.
  3. Entry-point plugins. Third-party packages (e.g. strands-robots-sim <https://github.com/strands-labs/robots-sim>_) register heavy out-of-tree backends - Isaac Sim, Newton - by declaring them under the strands_robots.backends entry-point group in their pyproject.toml::

    [project.entry-points."strands_robots.backends"] newton = "strands_robots_sim.newton.simulation:NewtonSimulation" warp = "strands_robots_sim.newton.simulation:NewtonSimulation"

so they can be discovered on pip install without patching this package. A plugin may map several entry-point names to the same class (newton and warp above) - whichever name is requested resolves cleanly. Plugins are discovered lazily on the first create_simulation / list_backends call (not at import time), and a plugin that fails to import is logged and skipped rather than crashing the factory. See the Python packaging spec for details: https://packaging.python.org/en/latest/specifications/entry-points/

Parameters:

Name Type Description Default
backend str

Backend name or alias. Defaults to "mujoco". Built-in: "mujoco" (aliases: "mj", "mjc", "mjx"). May also be any entry-point plugin name (see above).

DEFAULT_BACKEND
**kwargs Any

Backend-specific keyword arguments passed to the constructor (e.g., tool_name, timestep).

{}

Returns:

Type Description
SimEngine

A SimEngine instance ready for create_world().

Raises:

Type Description
ValueError

If the backend name is not recognized. The message lists all available backends (built-in + plugin) and, for known out-of-tree backends, a pip install hint.

ImportError

If the backend's dependencies are missing (e.g., pip install mujoco).

Examples::

# Default (MuJoCo)
sim = create_simulation()
sim.create_world()
sim.add_robot("so100")

# With alias
sim = create_simulation("mj")

# Pass kwargs to backend constructor
sim = create_simulation("mujoco", tool_name="my_sim")

# GPU-native built-in backend (requires strands-robots[sim-isaac])
sim = create_simulation("isaac", num_envs=1, headless=True)

list_backends

list_backends() -> list[str]

List all available backend names (built-in + plugin + runtime).

Merges the built-in registry (and its aliases), entry-point plugin backends discovered via importlib.metadata (see create_simulation), and any runtime-registered backends/aliases. Discovering plugins triggers a one-time lazy scan of the strands_robots.backends entry-point group.

Returns:

Type Description
list[str]

Sorted list of unique backend identifiers and aliases.

Example::

>>> list_backends()
['mj', 'mjc', 'mjx', 'mujoco']

register_backend

register_backend(name: str, loader: Callable[[], type[SimEngine]], aliases: list[str] | None = None, force: bool = False) -> None

Register a custom simulation backend at runtime.

Use this to add backends without editing source code.

Parameters:

Name Type Description Default
name str

Backend identifier (e.g., "my_physics").

required
loader Callable[[], type[SimEngine]]

Zero-arg callable that returns the backend class (not instance). Called lazily on first create_simulation().

required
aliases list[str] | None

Optional short names that resolve to name.

None
force bool

If False (default), raises ValueError when name or an alias is already registered. Set True to overwrite.

False

Raises:

Type Description
TypeError

If loader is the backend class itself rather than a callable returning it. Calling a class produces an instance, and the backend is instantiated later by create_simulation, so the mistake is refused here rather than at that later call.

ValueError

If name or an alias conflicts with an existing registration and force is False.

Example::

from strands_robots.simulation.factory import register_backend

register_backend(
    "bullet",
    lambda: BulletSimulation,
    aliases=["pybullet", "pb"],
)
sim = create_simulation("bullet")

Engine contract

stop_policy returns a json block with robot, was_running and exited; exited is null with no worker to join, the first stop after a rollout ended on its own adds last_result, and an empty name means the one rollout in flight. MuJoCo waits 1 s (MuJoCoSimulation._POLICY_STOP_JOIN_TIMEOUT) for the worker first.

strands_robots.simulation.base.SimEngine

Bases: ABC

Abstract base class for simulation engines.

Defines the contract that all backends (MuJoCo, Isaac, Newton) must implement. This is the programmatic API - the AgentTool layer wraps it with tool_spec/stream for LLM access.

Method categories:

Required (@abstractmethod): Core simulation loop - world lifecycle, entity management, observation/action, rendering, robot discovery. Every physics engine must implement these to be usable.

Provided (concrete base-class methods): Policy orchestration (run_policy / start_policy / replay_episode / eval_policy) is implemented once in this ABC as a facade over the abstract primitives. Backends inherit them for free by implementing the primitives. They may override for backend-specific optimisations (e.g. GPU-batched policy inference on Isaac).

Optional (default raises NotImplementedError): Higher-level features - scene loading, domain randomization, contact queries. Backends opt in by overriding only what they support.

Lifecycle::

sim = SomeEngine()
sim.create_world()
sim.add_robot("so100", data_config="so100")
sim.add_object("cube", shape="box", position=[0.3, 0, 0.05])

# Control loop
obs = sim.get_observation("so100")
sim.send_action({"joint_0": 0.5}, robot_name="so100")
sim.step(n_steps=10)

# Render
result = sim.render(camera_name="default")

# Cleanup
sim.destroy()

Concrete engines must set self._init_complete = True as the final statement of their __init__. :meth:__del__ consults it and skips an instance that never finished construction, so a half-built engine is never reported as a cleanup failure.

foxglove_url property

foxglove_url: str | None

The live Foxglove WebSocket URL, or None when no Foxglove bridge runs.

foxglove_link: str | None

A foxglove:// deep link to this engine's server, or None.

robot_name property

robot_name: str | None

The name of the one robot in this world, as its methods accept it.

Robot("so100").robot_name is "so100" - the string passed to Robot(), not the agent tool name ("so100_sim") - and matches robot_name on the real-hardware return. None when the world holds no robot or several, since then no single name answers it.

predicate_robot property

predicate_robot: str | None

The robot an unnamed base_* clause reads ON THIS THREAD, or None.

Read-only; set through :meth:bind_predicate_robot. The binding is thread-scoped, not scene-scoped: a rollout binds on the thread that drives it and every per-step read of the binding happens on that same thread, so two rollouts on two robots each read their own robot. See :meth:bind_predicate_robot for why a scene-wide attribute could not carry this.

capabilities

capabilities() -> frozenset[str]

Return CAPABILITIES, or derive it: the default set plus each overridden optional method.

Returns:

Type Description
frozenset[str]

Frozen set of names from :mod:strands_robots.simulation.capabilities.

create_world abstractmethod

create_world(timestep: float | None = None, gravity: list[float] | None = None, ground_plane: bool = True, terrain: str | None = None, difficulty: float = 1.0) -> dict[str, Any]

Create a new simulation world.

terrain ("rough" = value-noise bumps, "stairs" = discrete step plateaus rising along +x, "pyramid" = concentric step plateaus rising toward the centre, "slope" = a constant-grade inclined ramp; see :mod:strands_robots.simulation.terrain) lays down a deterministic heightfield instead of the flat ground plane so a locomotion policy can be spawned/evaluated on non-flat ground; it is only meaningful when ground_plane=True and defaults to None (a flat plane). Backends without heightfield support reject a non-None terrain with an actionable error rather than silently ignoring it.

difficulty scales the terrain's peak elevation (1.0 = full height, <1 gentler, >1 harsher) so a curriculum can ramp terrain magnitude across resets without changing the terrain kind. It is only meaningful with a terrain; setting difficulty != 1.0 with no terrain is rejected with an actionable error rather than silently having no effect. Must be a finite value > 0.

A floating-base robot added to a terrain world is spawned seated on the local terrain surface (raised by the heightfield height beneath it) at add_robot and on reset(), so its feet are not buried below the raised terrain.

timestep (seconds) and gravity must be values the engine can honor, on the same terms the set_timestep / set_gravity setters enforce: timestep a finite number > 0 (0 is rejected, never coalesced to the engine default), gravity a 3-element vector of finite numbers or a real scalar taken as the z-component. A value the backend cannot apply is rejected with a structured error rather than compiled into the world - a world built around a negative or nan dt integrates backwards or to nan while every subsequent call still reports status="success". None means "use the engine default".

destroy abstractmethod

destroy() -> dict[str, Any]

Destroy the simulation world and release resources.

reset abstractmethod

reset() -> dict[str, Any]

Reset simulation to its initial state.

Contract: on return the world must be left in a fully consistent, observation-ready state - derived kinematics (Cartesian body/site/geom poses and camera transforms) must reflect the reset pose WITHOUT requiring a subsequent step(). eval_policy calls get_observation() immediately after reset() and before the first action, so a backend that leaves derived state stale would feed the policy's first inference of every episode a degenerate observation. The MuJoCo backend enforces this by running mj_forward after mj_resetData (which alone zeroes all derived quantities). It also re-applies any per-robot home pose captured from an add_robot(keyframe=...) spawn, so a keyframe pose survives a reset instead of collapsing to the zero configuration.

step abstractmethod

step(n_steps: int = 1) -> dict[str, Any]

Advance simulation by n physics steps.

When the backend exposes an engine lock (self._lock, all in-tree backends), implementations must not hold it for the whole count: they release it at least every :attr:_STEPS_PER_BATCH steps, and re-check that the world still exists on each batch boundary before advancing it, aborting with a structured error naming the steps completed if it does not. Releasing the lock is what makes a concurrent teardown reachable mid-call, so the two halves are one contract rather than two - the same pairing _primitive_abort_reason already makes for the motion-primitive loops, which release the lock on the same schedule.

get_state abstractmethod

get_state() -> dict[str, Any]

Get full simulation state summary.

add_robot abstractmethod

add_robot(name: str, urdf_path: str | None = None, data_config: str | None = None, position: list[float] | None = None, orientation: list[float] | None = None, keyframe: str | int | None = None) -> dict[str, Any]

Add a robot to the simulation.

keyframe optionally spawns the robot in a canonical pose declared by a <keyframe> in its source model (e.g. panda "home", aloha "neutral_pose") instead of the default all-zero configuration. Pass the keyframe name (str) or index (int). The pose is applied to the robot's joints by name and stored so :meth:reset restores it (a keyframe spawn is sticky across resets, matching how a benchmark restores its canonical start each episode). None (the default) keeps the historical zero-pose spawn. An unknown keyframe name/index is a hard error that names the available keyframes; it never silently falls back to zeros.

Refused while a dataset recording is live, on every backend (:meth:~strands_robots.simulation.recording.DatasetRecordingMixin._recording_schema_frozen_error): the recorder's columns were declared at start_recording and a robot added into them has nowhere to be written.

remove_robot abstractmethod

remove_robot(name: str) -> dict[str, Any]

Remove a robot from the simulation.

list_robots abstractmethod

list_robots() -> list[str]

Return ordered list of robot names currently in the world.

Used by the backend-agnostic PolicyRunner to resolve a default robot when the caller omits robot_name.

robot_joint_names abstractmethod

robot_joint_names(robot_name: str) -> list[str]

Return ordered joint names for robot_name.

This is every joint, in the backend's order, including a floating base's free joint (floating_base_joint on g1), which has seven position coordinates and no scalar column. A LeRobotDataset recording writes observation.state in this order with free joints left out (the base goes to the base_* columns), so on a floating-base robot this list is one wider than that vector. Bind a policy with :meth:robot_action_keys (Policy.set_robot_state_keys, send_action with a numeric vector, PolicyRunner.replay): it skips the free joint and names actuators, which are not always joints.

Raises:

Type Description
ValueError

robot_name is not in the world. The message names the robots that are, so a robot added as "g1" and asked for as "unitree_g1" is told its real name instead of handed an empty roster that binds a policy to nothing.

robot_action_keys

robot_action_keys(robot_name: str) -> list[str]

Return the action keys send_action resolves for robot_name.

These are the names a policy should emit as its action-dict keys: the robot's actuators, which are NOT always its joints. A robot can have passive/mimic joints with no driving actuator (gripper finger followers) and tendon-driven actuators that are not joints at all (a grasp tendon). Keying a policy by robot_joint_names in those cases emits keys that send_action cannot resolve, so the affected actuators never move and the robot silently no-ops.

The default mirrors :meth:robot_joint_names for backends whose actuator set matches their joint set. Two kinds of backend override it. One has a distinct actuator namespace (MuJoCo tendon grippers) and returns actuator short-names instead. The other shares the namespace but commands a subset of it: the Newton engine drops a floating base's 6-DoF free joint, which is a joint and not a commandable scalar, so its action keys are the joint names minus that one. An override may therefore rename or narrow this list, and a caller must not assume it has the same width as :meth:robot_joint_names.

An override orders the keys by the joint each actuator drives, because this list also orders the observation.state vector a policy reads (Policy.set_robot_state_keys) and a recording writes those columns in joint order. An actuator that drives no single joint has no joint to be ordered by and keeps its backend-declared position.

actuator_ranges

actuator_ranges(robot_name: str) -> dict[str, tuple[float, float]]

Return the (low, high) control range of each range-limited actuator of robot_name.

Keyed like :meth:robot_action_keys. A backend clamps a send_action target past these bounds silently, so a driver reports the clamp from this map rather than from the backend's model. An unlimited actuator is absent; the default is {} for backends that cannot report ranges.

Raises:

Type Description
ValueError

robot_name is not in the world, with the message :meth:robot_joint_names raises, so an alias is not read as a robot with no limits.

saturated_actuators

saturated_actuators(robot_name: str) -> list[str] | None

Return the actuators of robot_name whose applied force sits at its force limit now.

Keyed like :meth:robot_action_keys. A position servo pinned at its limit is pushing against contact or a joint stop instead of reaching its command, which key resolution cannot see. [] means none is pinned; the default None means the backend cannot tell.

Raises:

Type Description
ValueError

robot_name is not in the world, with the message :meth:robot_joint_names raises, so an alias is not read as "cannot tell".

bind_predicate_robot

bind_predicate_robot(robot_name: str | None) -> None

Bind the robot an unnamed base_* clause reads, for the calling thread.

Benchmark and stop_when clauses default robot to "the sole robot". In a multi-robot scene that used to resolve to the FIRST registered robot, so evaluate_benchmark(benchmark_name='go2_walk_forward', robot_name='go2') with an arm registered first probed the arm ("has no floating base") and, with two floating-base robots, would have scored the wrong one silently. run_policy / eval_policy / evaluate_benchmark call this with the robot they resolved; the predicate readers consult it through :func:~strands_robots.simulation.predicates._bound_robot.

Concurrency contract: the binding is per thread. Rollouts are per-robot and explicitly concurrent - start_policy submits each to the engine's executor, and "policies on different robots can execute concurrently" is a documented surface - so one scene-wide attribute would make the last bind win: from that instant the OTHER rollout's unnamed clauses (evaluated every step) would read the wrong robot, silently, under status=success. All three surfaces bind on the thread that then drives the rollout, and every reader of the binding (the base_* predicates, a benchmark's on_episode_start compatibility check) runs on that same thread, so a thread-local slot is exactly the scope the binding needs. A refused or concurrent call therefore cannot disturb a rollout in flight on another thread. The binding stays until the same thread rebinds; a stale one (its robot since removed) is dropped by the reader.

Parameters:

Name Type Description Default
robot_name str | None

The robot to bind, or None to restore the sole-robot default on this thread.

required

bind_policy_sim_context

bind_policy_sim_context(policy: Any, robot_name: str) -> None

Give a policy the backend sim context it needs to close the loop.

Default no-op. The MuJoCo engine overrides this to hand a policy that opts in - by exposing a callable set_sim_context - the compiled MjModel + the robot's namespace, so an eef/cartesian-delta policy can auto-configure its IK end-effector frame with zero manual wiring. MockPolicy opts in to keep its sinusoid inside each actuator's ctrlrange; this is also the extension point an out-of-tree policy uses. Policies that do not expose set_sim_context are unaffected.

add_object abstractmethod

add_object(name: str, shape: str = 'box', position: list[float] | None = None, orientation: list[float] | None = None, size: list[float] | None = None, color: list[float] | None = None, mass: float = 0.1, is_static: bool | None = None, mesh_path: str | None = None, material: dict[str, Any] | None = None) -> dict[str, Any]

Add a primitive or mesh object to the scene.

size is the full extent in meters on every backend: a box's [x, y, z] edge lengths, a sphere's diameter, a cylinder's or capsule's [diameter, unused, length]. size=[0.05, 0.05, 0.05] is a 5 cm cube on MuJoCo, Newton, mjlab and Isaac alike; each backend halves it into its engine's own half-extents. See the concrete backend's add_object docstring for the exact per-shape semantics. Returns an agent-tool status dict.

A backend MUST NOT discard size components the caller did supply. When the vector is shorter than the shape consumes it either rejects it (MuJoCo: the per-shape component count is part of the contract) or pads only the missing trailing components from a documented default (Isaac). Replacing the whole vector with a backend default compiles a differently-sized object while reporting success -- and the reported size echoes what was asked for, not what was built.

The same rule applies to color: a backend either honors the component count it was given or rejects it, and may complete only components it documents a default for (MuJoCo completes an RGB triple with an opaque alpha, and rejects every other count). Falling back to the backend's default colour paints a surface the caller never asked for under a success result.

mass must be a finite number greater than zero for a dynamic object. A backend MUST NOT establish a body on a mass its own set_body_properties would refuse: a non-finite mass makes the first integration step produce nan and, because the solver shares one state vector, poisons every other body in the world too.

is_static is tri-state, and None is the default because it is the only value that means "the caller did not specify". That is what lets a backend derive the answer from shape: MuJoCo forces a shape="plane" static -- a plane is infinite and cannot carry a dynamic mass -- and refuses an explicit is_static=False there rather than quietly overriding it. A backend with no shape-derived rule resolves None to False (dynamic). Declaring the default as False would state a value the default backend does not deliver, and would make restating that declared default a hard error for the one shape whose whole point is being static.

The two chosen values select a posture -- welded to the world, or a free body the solver integrates -- so a supplied is_static is checked rather than read by truthiness: anything that is neither a boolean nor None is refused. Reading it by truthiness inverts both halves. 0 is the same value as the False a backend may refuse for a shape it forces static, so it reaches the quiet override that refusal exists to prevent, and every non-empty string is truthy, so "false" welds a body the caller asked to be dynamic and lands on :class:SimObject.is_static, which is annotated bool and read by list_objects, the scene rebuild and domain randomization.

material (optional): backend-specific visual material/texture spec. None keeps the flat color rgba (unchanged); a backend that supports it (MuJoCo) attaches a real material so surfaces can be matte or textured. Backends that do not support it should reject a non-None material loudly rather than silently ignore it. A supporting backend must likewise reject material keys it cannot honor (a typo, or a field from another renderer) instead of dropping them -- a dropped key renders the backend default while reporting success.

remove_object abstractmethod

remove_object(name: str) -> dict[str, Any]

Remove an object from the scene.

get_observation abstractmethod

get_observation(robot_name: str | None = None, *, skip_images: bool = False) -> dict[str, Any]

Get full observation for a robot: joint state + all attached cameras.

Unified observation consumed by :class:Policy and :class:~strands_robots.simulation.policy_runner.PolicyRunner. Backends MUST return a dict with the following schema; extra keys are allowed.

Schema
  • "<joint_name>" (float): One entry per joint on the robot, keyed by the model's joint name with any multi-robot namespace stripped ("joint1" on the Panda, "1" on the SO-101). The schema is stable regardless of multi-robot namespacing at the physics-engine level. A registry joint_labels name ("shoulder_pan") is a write-side alias: send_action takes it, the observation keeps the model's name, so recorded datasets and trained checkpoints keep one column per joint.
  • "<joint_name>.vel" (float): The same joint's velocity (rad/s or m/s), one entry per scalar joint, additive beside the position key so position-only consumers are unaffected. Velocity-feedback controllers (WBC's balance loop, the microduck and ProtoMotions observation packers, an RL env with .vel in its actor_obs_keys) read these to close the loop; a backend that omits them feeds those consumers zeros or a KeyError while the identical policy works elsewhere, which is exactly the portability break this schema exists to prevent. This entry was previously undocumented here and lived only in the MuJoCo implementation, which is how two backends shipped without it. A free-joint (floating) base is NOT a scalar joint and reports its twist via base_lin_vel / base_ang_vel below, never as "<name>.vel".
  • "<camera_name>" (np.ndarray): One RGB uint8 frame per camera associated with the robot, keyed by camera name. Shape (H, W, 3). A key MUST carry the view of the camera it names; a backend that cannot render that camera MUST omit the key rather than substitute another view (the free/overview camera in particular), because every consumer of this schema - a policy reading observation.images.<name>, a recorded dataset column - reads the key as a promise about which camera it is looking through and has no way to detect a substitution. Cameras whose render fails MAY be omitted; joint state MUST still be returned.
  • Floating base: a robot whose root is a 6-DoF free joint (a humanoid's named floating_base_joint or a mobile base's unnamed <freejoint>) does NOT report that free joint as a scalar "<joint_name>" entry - its qpos is [xyz + quat], so a scalar would report the base x-coordinate as a joint angle and drop the rest. Instead it surfaces the full base pose + twist as "base_pos" (world x,y,z incl. height), "base_quat" (w,x,y,z), "base_lin_vel" and "base_ang_vel", matching :meth:get_robot_state's "base" entry. Absent for fixed-base arms.
  • "body.<name>.pos" / ".quat" / ".lin_vel" / ".ang_vel" (list[float]): World pose + twist of a NAMED body, present only when the running policy declared that body in :attr:~strands_robots.policies.base.Policy.required_bodies. Backends do not emit these from get_observation itself - the runtime (:class:~strands_robots.simulation.policy_runner.PolicyRunner) merges them in from :meth:get_body_state for the declared bodies only, so the default observation is unchanged and nothing pays for a link nobody asked for. Motion-mimic trackers need them because their anchor link (torso_link on a G1) is separated from base_quat (the pelvis) by the waist joints.

Single-camera rendering is :meth:render's job, not this method's. For batched multi-robot observation (future Isaac / Newton), add a separate get_observations(robot_names) method - do NOT extend this one.

Parameters:

Name Type Description Default
robot_name str | None

Which robot to observe. If None and exactly one robot exists, that robot is used; otherwise returns {}.

None
skip_images bool

Skip camera rendering and return joint state only. Rendering dominates the per-step cost, so every consumer that reads joint values alone passes True - the predicate / reward DSL (:mod:~strands_robots.simulation.predicates), the LIBERO adapter's state reads, and the ROS 2 bridge when it publishes joint_states without image_raw. Camera keys are then absent from the result rather than present and empty, so a caller must not read a missing frame as a render failure. True renders nothing, recording or not: a rollout loop that records the observation it reads asks for the cameras itself while a dataset recording keeps them. Defaults to False (render every attached camera).

False

Returns:

Type Description
dict[str, Any]

Observation dict per schema above. Returns {} - and logs a

dict[str, Any]

WARNING naming the cause - when there is no world (never created,

dict[str, Any]

or after :meth:destroy), the robot cannot be resolved, or

dict[str, Any]

robot_name is unknown. An empty dict is not a heartbeat: the

dict[str, Any]

write methods (send_action, step) answer the same

dict[str, Any]

condition with status="error".

get_ground_height

get_ground_height(x: SupportsFloat, y: SupportsFloat) -> dict[str, Any]

Query the terrain surface height (world z) beneath world (x, y).

Public counterpart of the internal :meth:_ground_height_at hook: a create_world(terrain=...) heightfield raises the local ground up to TERRAIN_ELEVATION * difficulty above z=0, and there was no public way to ask where that surface is. Callers building a terrain scene need it to place an object / camera / goal on the surface -- an object added at a flat-ground z (computed as if the support were at z=0) on a raised plateau spawns buried in the heightfield and sinks through instead of resting on it. The same local-height sampler already backs the terrain-relative locomotion predicates (base_below_z) and the spawn/reset base-seating; this exposes it as a facade query.

Returns 0.0 for a flat ground plane, for any backend without a heightfield, and before create_world (a world-less engine has no terrain), so a non-terrain -- or not-yet-built -- world reports a flat surface rather than raising, unlike the world-scoped physics queries.

Parameters:

Name Type Description Default
x SupportsFloat

World x coordinate. Any object convertible to float (SupportsFloat), including NumPy scalars, that is a finite real number.

required
y SupportsFloat

World y coordinate. Same accepted types as x.

required

Returns:

Type Description
dict[str, Any]

Agent-tool status dict. On success content carries a

dict[str, Any]

{"json": {"x": ..., "y": ..., "height": ...}} block with the

dict[str, Any]

surface height in meters. Errors when x / y is not a finite

dict[str, Any]

real number. Accepts any real scalar, including NumPy scalar

dict[str, Any]

types (np.float32 / np.int64 / ...), since terrain

dict[str, Any]

coordinates naturally come from mj_data / an observation

dict[str, Any]

(a NumPy array), not hand-typed Python floats.

send_action abstractmethod

send_action(action: dict[str, Any] | Sequence[float], robot_name: str | None = None, n_substeps: int = 1) -> dict[str, Any]

Apply action and advance physics by n_substeps.

Contract: each call writes actuator/ctrl values and then runs n_substeps physics steps (e.g. mj_step). PolicyRunner.run() relies on this - it calls send_action once per control step and does NOT call sim.step() separately.

n_substeps is a positive whole number, on the shared :func:~strands_robots.utils.positive_whole_number_error domain every backend applies. A NumPy or float count with an integral value is honored and coerced; a fractional, zero, negative, non-finite, boolean or non-numeric count is refused as a structured error, and nothing is written when it is - a refusal arriving after the write would leave the robot commanded and the world un-advanced, which is the one state this surface must never report an error from. The floor is 1 rather than :meth:step's 0 precisely because of the write: "advance nothing" is step(0), an accepted no-op that commands nothing, while a send_action advancing nothing leaves a target the world never integrates. It is also the floor both producers of this count already enforce - PolicyRunner._control_substeps returns >= 1 and raises otherwise, and training.rl.env.SimEnv refuses an n_substeps below 1 - so this surface was the only member of that chain without the guarantee.

Backends are responsible for internal thread-safety (e.g. MuJoCo acquires self._lock here). PolicyRunner does not manage locks.

Returns:

Type Description
dict[str, Any]

Dict with status and content. A batch is applied whole or

dict[str, Any]

not at all: when any action key cannot be resolved, nothing is

dict[str, Any]

written, the world does not advance, and the content list

dict[str, Any]

includes a json block with unresolved_keys and an empty

dict[str, Any]

applied so callers can self-correct and resend. status is

dict[str, Any]

"error" when n_substeps is outside its domain.

physics_timestep

physics_timestep() -> float | None

Return the physics integration timestep in seconds, or None.

Used by :class:PolicyRunner to convert a policy's control_frequency into the number of physics substeps per control step (round(1 / control_frequency / physics_timestep)) so a position-servo robot actually tracks each action's target before the next action overwrites ctrl. Backends that cannot report a fixed timestep return None and the runner falls back to n_substeps=1.

render abstractmethod

render(camera_name: str = 'default', width: int | None = None, height: int | None = None) -> dict[str, Any]

Render a camera view.

Returns an agent-tool dict with status and a content list. On success the content holds an image block carrying PNG bytes ({"image": {"format": "png", "source": {"bytes": ...}}}); the raw RGB numpy arrays are available per-camera via :meth:get_observation. Resolution comes from the named camera's configuration (set via add_camera) unless width/height are given; the free camera and model-only cameras fall back to the engine default.

run_policy

run_policy(robot_name: str | None = None, policy_provider: str = 'mock', policy_config: dict[str, Any] | None = None, instruction: str = '', duration: float = 10.0, control_frequency: float | None = None, action_horizon: int = 8, fast_mode: bool = False, video: dict[str, Any] | None = None, policy_object: Policy | None = None, n_steps: int | None = None, max_steps: int | None = None, max_onframe_failures: int | None = None, control_substeps: int | None = None, policy_kwargs: dict[str, Any] | None = None, seed: int | None = None, n_episodes: int = 1, reset_between: bool = True, async_rtc: bool | None = None, rtc_inference_timeout_s: float | None = None, wbc_install_torque_control: bool = True, stop_when: dict[str, Any] | Callable[[SimEngine], bool] | None = None, observer: RunPolicyObserver | None = None) -> dict[str, Any]

Run a policy loop in the simulation (blocking).

Default implementation delegates to the backend-agnostic :class:~strands_robots.simulation.policy_runner.PolicyRunner. Backends MAY override for backend-specific optimisations (e.g. GPU-batched policy inference on Isaac).

Parameters:

Name Type Description Default
robot_name str | None

Robot to control.

None
policy_provider str

Name passed to :func:strands_robots.policies.create_policy.

'mock'
policy_config dict[str, Any] | None

Opaque dict of provider-specific kwargs (observation_mapping, action_mapping, host, port, api_token, pretrained_name_or_path, trust_remote_code, actions_per_step, use_processor, processor_overrides, device, ...). Forwarded verbatim to create_policy.

None
instruction str

Natural-language instruction for the policy.

''
duration float

Wall-clock seconds to run, honored as such: the loop paces on a deadline, so a step's own cost comes out of the period rather than being added to it (fast_mode=True removes the pacing entirely). Used only when no n_steps / max_steps is given (the step count wins and duration is recomputed from it). Must be a finite positive number; a non-positive, non-finite, non-numeric, or bool value is reported as a structured caller error rather than running a zero-step rollout that reports success - as is a span shorter than one control period (1 / control_frequency), which resolves to no steps at that rate.

10.0
control_frequency float | None

Target Hz for policy queries. Must be a positive number; a non-positive, non-numeric, or bool value is reported as a structured caller error.

None
control_substeps int | None

Explicit physics steps to integrate per applied action, overriding the control_frequency-derived value. Must be a positive integer; 0, a negative value, a float, or a bool is reported as a structured caller error rather than collapsing to a single physics step (which under-integrates each control period so the arm barely moves while the rollout still reports success). None (default) derives it from control_frequency and the backend's physics timestep.

None
action_horizon int

Lower bound on actions consumed from each policy chunk before re-querying. The effective interval is max(action_horizon, policy.execution_horizon) (see strands_robots.policies.resolve_chunk_length): a chunk-emitting policy always keeps its full trained chunk, so a value below that chunk length (e.g. a VLA whose execution_horizon is 50) has no effect. RTC policies own their own interval and ignore this entirely. Must be a positive integer (>= 1); a non-positive or non-int value is reported as a caller error.

8
fast_mode bool

Skip the real-time pacing and run as fast as inference and physics allow. When False (default) the loop is paced on a deadline at control_frequency, so the wall clock a step spends working is subtracted from the period rather than added to it and duration is honored whatever a step costs. Must be a boolean: it selects a posture rather than scaling a quantity, so a value of any other type is reported as a structured caller error rather than read by truthiness - a truthy "false" would otherwise run the loop unpaced.

False
video dict[str, Any] | None

Optional video-recording config dict. Accepted keys: path (str, output MP4 - required to enable recording), fps (int, default 30), camera (str, default backend default), width (int, default 640), height (int, default 480). See :class:~strands_robots.simulation.policy_runner.VideoConfig. Any other key - and any fps/width/height that is not a positive whole number - is reported as a caller error naming the offending key, rather than being dropped (which used to turn a mistyped path into a "successful" rollout with no MP4, and a mistyped size into a default-resolution recording). For extension points beyond video (custom telemetry, dataset recording), backends plug into PolicyRunner.run's on_frame hook via :meth:_make_run_policy_hook.

None
policy_object Policy | None

Already-constructed :class:~strands_robots.policies.Policy to drive the rollout, bypassing create_policy entirely (passing it beside a policy_provider / policy_config is refused). Reuse one instance across calls so the checkpoint is not reloaded per rollout - the policy_load_cache_hit field below reports when a caller rebuilt it instead. None (default) builds the policy from policy_provider / policy_config.

None
n_steps int | None

Exact control-step horizon. When given it REPLACES duration, which is recomputed as n_steps / control_frequency, and it bypasses the lossy int(duration * control_frequency) conversion so the rollout executes exactly this many control steps. Must be a positive integer; 0, a negative value, a bool, or a fractional or non-numeric value is reported as a structured caller error rather than truncated to a horizon the caller never asked for. None (default) falls back to max_steps, then to duration.

None
max_steps int | None

Legacy alias for n_steps, kept for callers written against the older name. Consulted only when n_steps is None - n_steps wins when both are passed - and refused under its own name on the same domain.

None
max_onframe_failures int | None

Maximum consecutive on_frame-hook exceptions tolerated before the rollout aborts the episode. That hook is where a backend attaches dataset recording and video capture, so a broken recorder otherwise fills an empty dataset behind a successful-looking rollout. Two classes are exempt from the count rather than tolerated by it: CooperativeStop is the documented graceful stop, and :class:~strands_robots.recording_errors.RecordingFrameError is data loss rather than telemetry, so a lost dataset frame aborts on the FIRST occurrence whatever this is set to. None (default) uses the runner's own limit (currently 5); non-consecutive failures reset the counter. Must otherwise be a positive integer - the same domain as n_steps and n_episodes above, since the runner compares it against a plain integer counter. nan and inf make that comparison false forever and so disable the abort entirely; 0 aborts on the first failure exactly as 1 does while reporting a count of zero. Both are reported as a structured caller error rather than silencing the watchdog - see :meth:~strands_robots.simulation.base.SimEngine._validate_onframe_failure_limit. Forwarded verbatim to :meth:~strands_robots.simulation.policy_runner.PolicyRunner.run.

None
seed int | None

Optional master RNG seed for a reproducible single rollout. When set, reseeds Python / NumPy / torch / cuDNN and forwards policy.reset(seed=...) so a stochastic policy (VLA action- chunk sampling, diffusion noise) produces the same trajectory on re-run of the same scene. Reproducible means bit-exact for a state-only policy (mock re-runs byte for byte), and to a render tolerance for a policy that reads camera frames on GPU rendering: under MUJOCO_GL=egl a static scene renders with 1 LSB differences between frames (measured 2 to 9 pixels per 640x480 frame), a VLA reads them, and two seeded ACT rollouts ended 0.006 to 0.3 rad apart after 90 to 150 steps. Compare seeded camera-policy runs by their outcome, not by equality; a policy that reads no frame is exact on any renderer. None (default) leaves RNG state untouched. Mirrors the per-episode reseed in :meth:eval_policy.

None
policy_kwargs dict[str, Any] | None

Optional per-call goal payload forwarded verbatim to every policy.get_actions(obs, instruction, **policy_kwargs) call. Carries the well-known #300 goal keys (target_pose / target_joints / target_velocity / world_update) to non-VLA providers (cuRobo, MoveIt2, WBC) that read their goal from kwargs rather than the instruction. This is the local-sim analogue of the mesh tell() path, which already forwards these keys. VLA providers ignore unknown kwargs per the #300 contract, so forwarding is always safe.

None
n_episodes int

Number of sequential episode rollouts to run in this single call (default 1 - the historical single-rollout behaviour, unchanged). IMPORTANT: calling with the default n_episodes=1 produces exactly ONE dataset episode, no matter how many "episodes" you intend in natural language. To record N DISTINCT dataset episodes pass n_episodes=N in a single call - do NOT loop this call N times narrating "N episodes" (that buffers all frames into one merged episode_index=0 mega-episode). After stop_recording, confirm the count with :meth:verify_dataset_episodes. When > 1, each episode runs one rollout for the configured horizon, then a dataset episode boundary is flushed via :meth:save_episode (only when a recording is active) so the dataset ends up with N correctly delimited episodes instead of one merged episode. This is the first-class multi-episode collection API; it removes the need for a manual for _ in range(n): run_policy(); save_episode(); reset() loop. seed (when set) is offset per episode (seed + i) for reproducible-yet-distinct rollouts, and video (when set) is written per episode to a path with _ep{i} inserted before the extension so episodes do not overwrite one another.

1
reset_between bool

When running multiple episodes, reset the sim to its initial state between episodes (default True). The reset never fires after the final episode. Set False to chain episodes from the end state of the previous one. Must be a boolean, reported as a structured caller error otherwise - a falsy 0 would otherwise chain the episodes without being a declared spelling of that.

True
async_rtc bool | None

When True, overlap policy inference with action execution so the next action chunk is computed in the background while the current chunk is still draining (latency masking). False keeps the synchronous chunk-then-drain loop. None (default) enables it only for a chunk-emitting policy that blends the seam (supports_rtc), since a prefetched chunk without RTC starts from a stale observation; every other policy stays synchronous; an explicit True/False always wins, and a supplied value must be one of those two: any other type is reported as a structured caller error rather than read by truthiness, since a truthy "false" would otherwise start the background inference thread it reads as declining. Forwarded verbatim to :meth:PolicyRunner.run; see its docstring for the full contract (provider-agnostic, RTC-policy seam blending, thread safety).

None
rtc_inference_timeout_s float | None

Optional hard per-chunk timeout (seconds) for the async-RTC prefetch. When set, a stuck inference surfaces as a structured status=error result (carrying the RTC telemetry) instead of hanging the sim. None (default) waits without a deadline. Forwarded verbatim to :meth:PolicyRunner.run; ignored on the synchronous path.

None
wbc_install_torque_control bool

When True (default), a :class:~strands_robots.policies.wbc.WBCPolicy run on a position-servo scene (the stock Robot("unitree_g1")) gets the torque shim auto-installed for the duration of this call, then uninstalled. WBC emits joint-position targets; the stock G1's uniform kp=500 servo would override SONIC's tuned per-joint PD and the gait diverges, so the documented quickstart silently falls over without it. Set False to manage the controller yourself or to drive a torque-actuated scene directly. No-op for non-WBC policies and on backends without the hook. Must be a boolean, reported as a structured caller error otherwise - a falsy 0 would otherwise skip the shim without being a declared spelling of that.

True
stop_when dict[str, Any] | Callable[[SimEngine], bool] | None

Optional semantic early-return condition: end the rollout as soon as the WORLD reaches a state, not only when the step budget runs out - which turns a monolithic rollout into a retryable primitive an agent can invoke -> inspect -> re-invoke. A predicate-DSL clause in the same schema as a benchmark spec's success clause: a single call {"predicate": "grasped", "body": "cube", "gripper_prefix": "so101/gripper"} or an {"all": [...]} / {"any": [...]} group of bool predicate calls. Compiled via :func:~strands_robots.simulation.benchmark_spec.compile_stop_when against the closed predicate registry (never eval / exec; an unknown predicate name is rejected up front with the valid list), and the clause's referenced body/joint names are probed against the LIVE scene before the rollout starts - a typo'd name (or a backend without body lookups) is an up-front structured error instead of a clause that silently never fires and burns the whole budget. The compiled clause is evaluated against the SIM after every applied action - matching the benchmark semantics, not the observation dict - on both the synchronous and async-RTC paths, so the stop lands within one control step of the condition holding. Composes with an active recording session: frames are captured up to the stop, so a recorded episode's frame count equals the result's steps_used. Programmatic callers may pass a callable (sim) -> bool instead of a dict (the tool surface accepts dicts only). None (default) keeps the pure step-budget horizon. The result json reports why the rollout ended via stopped_reason. A clause the scene's initial state ALREADY satisfies is the mirror of the typo above: evaluated only after an applied action, it fires on the first step whatever the policy commands, and the resulting stopped_reason="predicate" after one step is indistinguishable from a rollout that drove the world to the condition. That is reported rather than refused - stop_when_true_at_reset (bool, always present) and stop_when_reset_warning (the qualifying text, None otherwise) in the result json, plus a logged warning - because a randomised initial state legitimately satisfies a clause on some draws, and every other reported figure is left as measured.

None
observer RunPolicyObserver | None

Optional read-only rollout observer, forwarded verbatim to :meth:PolicyRunner.run. Receives one :class:~strands_robots.simulation.observers.RunPolicyStarted, one :class:~strands_robots.simulation.observers.RunPolicyStep per completed send_action call and one :class:~strands_robots.simulation.observers.RunPolicyEnded, carrying the per-key send_action verdict, authoritative observation_age_steps (with the narrower observation_is_chunk_reused chunk-position flag), and what the backend's own on_frame hook did. Must be None or callable; another value is refused before policy construction, backend hook creation, or rollout side effects.

This is a SECOND lane, not the backend's hook: the hook slot is filled from _make_run_policy_hook (cancellation, trajectory, mesh telemetry, dataset recording) and is not available to callers, which is exactly why a read-only consumer needs this. Installing one changes neither the actions applied nor any existing field of the result json; it adds observer_failures. With n_episodes > 1 each episode is its own lifecycle with its own run_id. Called synchronously, so a blocking observer blocks the rollout, and payloads are borrowed rather than copied - see :mod:strands_robots.simulation.observers for the ownership rules. Not yet wired into eval_policy, evaluate_benchmark or run_multi_policy.

None

Returns:

Name Type Description
dict[str, Any]

The standard agent-tool envelope

dict[str, Any]

``{"status": "success"|"error", "content": [{"text": ...},

dict[str, Any]

{"json": {...}}]}. Thejson`` block is the machine-readable

dict[str, Any]

rollout report; text carries the same facts for humans.

dict[str, Any]

Read the json block by SCANNING content for the first block

dict[str, Any]

with a "json" key, never by a fixed index::

report = next(b["json"] for b in result["content"] if "json" in b)

dict[str, Any]

An early caller-error return (a rejected duration, an unknown

dict[str, Any]

robot) carries a text block ONLY, so a hardcoded

dict[str, Any]

content[1] raises IndexError on exactly the results a

dict[str, Any]

caller most needs to read.

dict[str, Any]

IMPORTANT - status is not the rollout verdict. It reports

dict[str, Any]

whether the CALL was accepted and the loop ran; it does not say the

dict[str, Any]

robot did anything useful. A rollout that drove only a SUBSET of

dict[str, Any]

the robot's actuators is deliberately success (it is

dict[str, Any]

operational), so status alone cannot see it: a policy driving 1

dict[str, Any]

of a Panda's 8 actuators returns status="success" with

dict[str, Any]

action_errors=0 and partial_action_failure_rate=0.875. Gate

dict[str, Any]

on action_errors, partial_action_failure_rate and the

dict[str, Any]

binding-degradation flags below to decide whether a rollout is worth

dict[str, Any]

anything. Coarse backend errors remain in action_errors but are

dict[str, Any]

excluded from action-rate denominators rather than fabricated as

dict[str, Any]

confirmed misses. A TOTAL

dict[str, Any]

failure - no emitted key resolving to any actuator - is reported as

dict[str, Any]

status="error".

dict[str, Any]

Fields in the json block:

Identity dict[str, Any]

robot_name, policy (the driving policy's class

dict[str, Any]

name), instruction, instruction_read (False when the

dict[str, Any]

policy never read it - MockPolicy drives a test motion whatever

dict[str, Any]

the task says, and the text block then carries a note saying so;

see dict[str, Any]

attr:~strands_robots.policies.base.Policy.reads_instruction).

Horizon dict[str, Any]

n_steps (control steps executed), steps_used

dict[str, Any]

(alias of n_steps under the retry-loop name), elapsed_s,

dict[str, Any]

sim_time_s (when the backend reports sim time),

dict[str, Any]

stopped_early, stopped_reason ("predicate" - the

dict[str, Any]

stop_when condition fired; "budget" - the step/duration

dict[str, Any]

horizon was exhausted; "cancelled" - a cooperative stop, e.g.

dict[str, Any]

stop_policy; "error" on error results - so an agent

dict[str, Any]

deciding whether to retry knows WHY the rollout ended).

dict[str, Any]

stop_when_true_at_reset (bool, always present) qualifies that

attribution dict[str, Any]

the clause is evaluated only AFTER an applied action,

dict[str, Any]

so one the scene's initial state already satisfies fires on the

dict[str, Any]

first step whatever the policy commands, making

dict[str, Any]

stopped_reason="predicate" after one step indistinguishable

dict[str, Any]

from a rollout that drove the world to the condition - the mirror

dict[str, Any]

of the never-fires case the pre-rollout entity probe refuses.

dict[str, Any]

stop_when_reset_warning carries the qualifying text (None

dict[str, Any]

when the flag is False). Reported rather than refused, and

dict[str, Any]

every other figure is left as measured: domain randomisation

dict[str, Any]

legitimately draws an initial state that satisfies a clause.

dict[str, Any]

Action health: action_errors (steps where the backend reported

dict[str, Any]

an error), actions_applied (actions that commanded at least one

dict[str, Any]

of the robot's keys - NOT the number of send_action calls, since

dict[str, Any]

an action naming no key reaches the backend like any other; a

dict[str, Any]

rollout whose count is 0 never commanded the robot for a single

dict[str, Any]

step and is returned as status="error", the mirror of the

dict[str, Any]

all-keys-unresolved refusal), action_resolution_rate (an

dict[str, Any]

{actuator_name: fraction_of_resolution-known_steps_driven} map,

dict[str, Any]

so a joint stuck at 0.0 names an actuator not confirmed on any

dict[str, Any]

known step) and partial_action_failure_rate (the mean fraction

dict[str, Any]

of the robot's DOF not confirmed driven across those known steps;

dict[str, Any]

0.0 == every actuator confirmed every known step, ~0.83 ==

dict[str, Any]

only 1 of 6). A coarse backend error is excluded from both rate

dict[str, Any]

denominators instead of being counted as a physical miss; it remains

dict[str, Any]

visible in action_errors and the human-readable diagnostic. A

dict[str, Any]

step whose applied keys name driven JOINTS rather than actuators is

dict[str, Any]

excluded on the same terms: send_action resolves that spelling

dict[str, Any]

(it looks the joint's driving actuator up), but it reports no

dict[str, Any]

actuator per key, so the step is unknown for per-actuator purposes

dict[str, Any]

rather than a miss - a rollout keyed entirely that way reports an

dict[str, Any]

empty map and 0.0, not the 1.0 of a robot that never moved.

Achievement dict[str, Any]

resolution says a command reached an actuator, not that

dict[str, Any]

the actuator achieved it, so an arm driven into the table reads

dict[str, Any]

healthy on every field above. saturated_step_rate is the

dict[str, Any]

fraction of steps on which any of the robot's actuators sat at its

dict[str, Any]

force limit, and saturation_rate the same per actuator (from

dict[str, Any]

meth:saturated_actuators; both None on a backend that cannot

dict[str, Any]

tell). A short burst is a fast move; most of a rollout is a stall.

Video dict[str, Any]

video_path (None when no MP4 was written),

dict[str, Any]

video_frames and video_fps (the rate the MP4 plays at -

dict[str, Any]

the requested fps capped to control_frequency, since a

dict[str, Any]

rollout renders at most one frame per control step).

Episodes dict[str, Any]

n_episodes_requested, n_episodes_completed,

dict[str, Any]

episodes_saved and dataset_episode_indices (the dataset

dict[str, Any]

episode indices this call flushed, empty without a recording).

dict[str, Any]

Policy binding: positional_fallback_used,

dict[str, Any]

generic_state_keys_used and missing_state_keys_used. True

dict[str, Any]

means the driving policy could not bind the observation to the

dict[str, Any]

model's inputs by name and silently fell back (a camera routed to a

dict[str, Any]

model image slot positionally, or observation.state composed

dict[str, Any]

from the observation's own scalar keys because none of

dict[str, Any]

robot_state_keys matched). A True flag on an otherwise

dict[str, Any]

success run is the signature of a robot moving on meaningless

dict[str, Any]

inputs.

dict[str, Any]

Policy load: policy_load_time_s, policy_load_cache_hit

dict[str, Any]

(False on episode 2+ of a loop is a smell that the caller

dict[str, Any]

rebuilt the policy instead of reusing policy_object=) and

dict[str, Any]

policy_resident_rss_mb.

dict[str, Any]

Chunk-prefetch telemetry, so latency masking is provable from the

dict[str, Any]

payload instead of from logs: chunk_prefetch_enabled (the

dict[str, Any]

background chunk pipeline was on - this is NOT the policy's RTC

dict[str, Any]

algorithm, which policy_rtc_enabled reports),

dict[str, Any]

chunk_prefetch_chunks_acquired, chunk_prefetch_hits,

dict[str, Any]

chunk_prefetch_blocks, avg_inference_ms and

dict[str, Any]

max_inference_ms. The pre-rename spellings

dict[str, Any]

rtc_async_enabled, rtc_chunks_acquired,

dict[str, Any]

rtc_prefetch_hits, rtc_prefetch_blocks,

dict[str, Any]

rtc_avg_inference_ms and rtc_max_inference_ms are kept for

dict[str, Any]

one release with the same values.

dict[str, Any]

Across episodes (n_episodes > 1): the aggregate payload adds

dict[str, Any]

total_steps, the per-episode episodes records,

dict[str, Any]

stopped_reasons (aligned with episodes) and

dict[str, Any]

video_paths, and keeps every field above whose value the call

dict[str, Any]

already knows - the identity fields, the policy-binding flags and

dict[str, Any]

the policy-load telemetry, all read off the ONE policy object the

dict[str, Any]

episodes shared. Per-episode action health (action_errors,

dict[str, Any]

action_resolution_rate, partial_action_failure_rate) and

dict[str, Any]

the per-episode horizon/video fields are reported by each record in

dict[str, Any]

episodes rather than summarised, because a rate has no single

dict[str, Any]

aggregate an N-episode call can report without choosing a summary::

worst = max(e["partial_action_failure_rate"] for e in report["episodes"])

dict[str, Any]

So the binding-degradation gate reads the same way at any episode

dict[str, Any]

count, and the action-health gate is per episode.

dict[str, Any]

Fail-fast: if EVERY action step in the opening probe window drives

dict[str, Any]

zero actuators - none of the policy's emitted keys resolve to any of

dict[str, Any]

the robot's actuators - the rollout can never move the robot, so it

dict[str, Any]

returns status="error" at the probe boundary instead of running

dict[str, Any]

the full episode (and every remaining model inference call +

dict[str, Any]

recording write). The error enumerates the unresolved keys and the

dict[str, Any]

robot's valid actuator names. A PARTIAL failure runs to completion,

dict[str, Any]

surfaced via partial_action_failure_rate.

run_multi_policy

run_multi_policy(policies: dict[str, Policy], instructions: dict[str, str] | str = '', duration: float = 10.0, control_frequency: float | None = None, action_horizon: int | dict[str, int] = 8, n_steps: int | None = None, max_steps: int | None = None, *, fast_mode: bool = False) -> dict[str, Any]

Drive MULTIPLE robots, each with its own policy, in ONE synchronized loop.

The backend-agnostic contract for concurrent multi-robot rollout (e.g. two arms doing a handover, or a bimanual setup). A backend that implements it must honour every clause below - they are what distinguishes this driver from launching one :meth:start_policy thread per robot, which steps physics per robot and interleaves single-robot recording frames:

  • Per-robot policies: policies maps each driven robot to its own :class:~strands_robots.policies.Policy. Every key must name a robot in the scene; policies order defines the merged state/action column order.
  • Per-robot instructions: instructions is either one string applied to all robots or a {robot_name: instruction} mapping. A mapping key naming no driven robot is rejected rather than silently dropped; a robot omitted from the mapping gets an empty instruction (see :meth:_normalize_multi_policy_instructions).
  • Per-robot action_horizon: action_horizon is either one int applied to all robots or a {robot_name: horizon} mapping. Every horizon must be a positive integer, and the effective per-robot chunk length is resolved through :func:~strands_robots.policies.base.resolve_chunk_length exactly as :meth:run_policy resolves its own (see :meth:_normalize_multi_policy_horizons).
  • Shared control_frequency: one target Hz for every robot's policy queries, so the robots stay phase-aligned.
  • Lockstep physics: each loop iteration applies EVERY robot's control, then steps physics ONCE - regardless of each robot's individual re-query cadence.
  • One merged recording frame per timestep: when a dataset recording is active, each timestep records a single frame carrying ALL robots' prefixed state/action (alice__shoulder_pan ...) plus all camera images - never one interleaved frame per robot.

The step horizon follows :meth:run_policy's resolution: n_steps (then its legacy alias max_steps) overrides duration, via :meth:_resolve_horizon on the shared positive-count domain.

This base implementation is a documented refusal, not a fallback: a backend that has no synchronized multi-robot loop must say so rather than silently driving robots one at a time (which would interleave frames and break the merged-frame contract above). The MuJoCo and Isaac backends override it with full implementations; backends that do not yet (Newton) inherit this structured error.

Parameters:

Name Type Description Default
policies dict[str, Policy]

Mapping {robot_name: Policy} of the robots to drive.

required
instructions dict[str, str] | str

Single instruction string for all robots, or a {robot_name: instruction} mapping (see contract above).

''
duration float

Episode length in seconds (steps = duration x freq). Used only when no n_steps / max_steps is given. Must be a finite positive number of at least one control period, so that product is at least one step.

10.0
control_frequency float | None

Target Hz for policy queries / physics. Must be a positive number.

None
action_horizon int | dict[str, int]

Actions consumed from each policy's chunk before re-querying it, as one int or a per-robot mapping (see contract above).

8
n_steps int | None

Exact step horizon (overrides duration when set).

None
max_steps int | None

Legacy alias for n_steps.

None
fast_mode bool

Skip the real-time pacing and run as fast as inference and physics allow, as :meth:run_policy does. When False (default) the loop is paced on a deadline at control_frequency. Must be a boolean: a value of any other type is refused rather than read by truthiness.

False

Returns:

Type Description
dict[str, Any]

A structured {"status": "error", ...} dict naming this

dict[str, Any]

backend class and stating that it does not implement synchronized

dict[str, Any]

multi-robot rollout. Implementing backends return the standard

dict[str, Any]

status dict with per-robot step counts.

verify_dataset_episodes

verify_dataset_episodes(expected: int) -> dict[str, Any]

Verify the recorded dataset holds exactly expected episodes.

Reads the LeRobot dataset parquet (the ground truth) for the active or most-recently-recorded session AND cross-checks it against the meta/info.json total_episodes header. Both must agree with expected; a parquet that matches expected but disagrees with info.json (an internally inconsistent dataset) still fails. Reports the actual episode count. Call this AFTER :meth:stop_recording for a definitive check that a collection run produced N distinct episodes rather than one merged episode_index=0 mega-episode.

Episodes are flushed to meta/episodes/**/*.parquet only at save_episode / stop_recording (finalize) time, so this reads the canonical on-disk truth - it does not trust the recorder's in-memory bookkeeping (which is what :meth:run_policy reports while a session is still open).

Parameters:

Name Type Description Default
expected int

The episode count the caller intended to record. A non-negative int; anything else is reported as an error dict.

required

Returns:

Type Description
dict[str, Any]

Standard status dict. status is "success" when the parquet

dict[str, Any]

holds exactly expected episodes, else "error". The

dict[str, Any]

{"json": {...}} block carries expected, actual,

dict[str, Any]

info_total_episodes, info_problems, sources_agree,

dict[str, Any]

episode_indices,

dict[str, Any]

total_frames, total_frames_per_ep, unreadable_files and

dict[str, Any]

root so a caller (or CI) can fail loudly programmatically.

dict[str, Any]

status is "error" when the parquet count differs from

dict[str, Any]

expected, when the parquet disagrees with meta/info.json's

dict[str, Any]

total_episodes (sources_agree is then False) - the two

dict[str, Any]

metadata sources must agree, never just one - and when any episode

dict[str, Any]

parquet file could not be read (unreadable_files non-empty),

dict[str, Any]

since the episodes found are then only a lower bound. An unreadable

dict[str, Any]

or corrupt parquet is reported as this same error dict, never raised.

save_episode

save_episode() -> dict[str, Any]

Flush the current recording episode and begin a fresh one.

Backends that support dataset recording override this (see the MuJoCo RecordingMixin). The base has no recorder, so it returns a structured error rather than pretending to flush.

start_policy

start_policy(robot_name: str | None = None, policy_provider: str = 'mock', policy_config: dict[str, Any] | None = None, instruction: str = '', duration: float = 10.0, control_frequency: float | None = None, action_horizon: int = 8, fast_mode: bool = False, video: dict[str, Any] | None = None, policy_object: Policy | None = None, n_steps: int | None = None, max_steps: int | None = None, max_onframe_failures: int | None = None, control_substeps: int | None = None, policy_kwargs: dict[str, Any] | None = None, seed: int | None = None, n_episodes: int = 1, reset_between: bool = True, async_rtc: bool | None = None, rtc_inference_timeout_s: float | None = None, wbc_install_torque_control: bool = True, stop_when: dict[str, Any] | Callable[[SimEngine], bool] | None = None, observer: RunPolicyObserver | None = None) -> dict[str, Any]

Run a policy rollout, in the background where the backend has one.

DEFAULT IMPLEMENTATION IS SYNCHRONOUS: it passes through to :meth:run_policy and returns only after the rollout has finished, so its result reports a COMPLETED rollout ("Policy complete on ...") and the call blocks for the whole duration. The summary line used to promise a background thread outright, which is what MuJoCo's override does, not what a caller of this default gets - and the two backends shipped on this default are the ones whose callers most need to know. Backends with true background execution override this (MuJoCo, via the ThreadPoolExecutor it owns) and return as soon as the rollout is submitted.

Either way :meth:stop_policy is the counterpart, and a caller can tell which of the two it holds without reading the source: this method's entry in :meth:describe states which one this engine implements, and a backend that also tracks rollouts in flight advertises list_policies_running there beside it.

Takes exactly the keywords :meth:run_policy takes, with the same meaning and the same refusals, so a blocking call becomes a background one by changing the method name and nothing else.

stop_policy

stop_policy(robot_name: str = '') -> dict[str, Any]

Stop robot_name's rollout (cooperative) and report what was in flight.

The counterpart to :meth:start_policy, and the verb that OWNS the question "was a rollout halted" for every backend. It lived only on the MuJoCo engine, so on the other backends the attribute did not exist at all: :meth:~strands_robots.mesh.Mesh._dispatch probes for it with hasattr and answered "peer exposes no stop_task" for a sim it could in fact have stopped, and the Device Connect stop RPC re-derived the answer inline from the per-robot flag - a second construction of the verdict that :meth:~strands_robots.simulation.models.SimRobot.request_policy_stop exists to prevent ("EVERY stop path goes through here ... so they cannot drift to different answers about whether a rollout was halted"). This is the same promotion :meth:run_multi_policy had (#2157): a capability every backend is asked for answers in the tool envelope on all of them, never with AttributeError because there was no contract.

The flag write itself is backend-owned, because the per-robot rollout claim is: :meth:_request_policy_stop is the seam, the mirror of the :meth:_make_run_policy_hook / :meth:_release_run_policy_hook pair that raises and lowers the same flag around a rollout driven here.

This is a cooperative stop, not a join: it moves the robot's claim out of date so the rollout's next frame ends it. It cannot interrupt a rollout that is blocked inside a single send_action or a single policy inference, and on a backend whose :meth:start_policy is the synchronous default the caller's own thread is the one inside the rollout - so the callers that reach this verb usefully are the ones on another thread (the Device Connect stop RPC and the mesh fanout).

Parameters:

Name Type Description Default
robot_name str

The robot whose rollout to stop. An empty name means the only rollout in flight when there is exactly one - the remedy every rollout gate names - and is otherwise refused naming what is running; it is never silently matched against the sole robot, because a stop aimed at the wrong robot reads as a stop that worked (:meth:_stop_policy_target).

''

Returns:

Name Type Description
dict[str, Any]

The agent-tool envelope. On success the json block reports

dict[str, Any]

was_running - whether a rollout really was in flight when the

dict[str, Any]

stop arrived - so a caller aggregating several answers reads the

dict[str, Any]

verdict rather than matching on the sentence. Idempotent: a robot

dict[str, Any]

with nothing running is status="success" with

dict[str, Any]

was_running=False, so a caller may stop unconditionally.

dict[str, Any]

status="error" for an empty or unknown robot_name, and for a

dict[str, Any]

backend that keeps no durable per-robot claim to move - that refusal

dict[str, Any]

names the class, because "nothing was running" would be an

dict[str, Any]

affirmative answer given on no evidence. Isaac is on that default

today dict[str, Any]

its per-robot record carries a bare policy_running flag

dict[str, Any]

and not the durable counter, and a bare flag write is the exact

dict[str, Any]

thing a worker that has not reached its first frame overwrites

dict[str, Any]

(#2833), so it refuses rather than reporting a stop it cannot keep.

policy_result

policy_result(robot_name: str) -> dict[str, Any] | None

The envelope the last asynchronous rollout on robot_name ended with.

:meth:start_policy returns "Policy started" and nothing else, so the report :meth:run_policy would have returned (steps, action health, video_path, the per-episode list) reached no caller once the worker finished; :meth:stop_policy after natural completion answered "Was not running" with no report either (#4162). A backend that keeps a worker table records the finished envelope, success or error, where :meth:_rollouts_ended_in_error records the failure reason, and answers here; the default is None for a backend whose :meth:start_policy is the synchronous default (its caller already holds the result).

Parameters:

Name Type Description Default
robot_name str

The robot the rollout was started on.

required

Returns:

Type Description
dict[str, Any] | None

The run_policy envelope of the most recent completed rollout on

dict[str, Any] | None

that robot, or None when none has completed, when the rollout

dict[str, Any] | None

is still in flight, or when the backend keeps no record. A later

dict[str, Any] | None

rollout on the same robot replaces the entry once it completes.

list_policies_running

list_policies_running() -> dict[str, Any]

Name the robots a rollout is driving right now.

The public reader of the in-flight population, promoted here from the MuJoCo engine so it answers on every backend: docs/reference/simulation/rollouts.md lists it in the Policy action table with no backend qualifier, and documents :meth:stop_policy -- on this ABC since a robot's stop became a base contract -- as deriving its verdict from "the same in-flight population list_policies_running reads". Only one of that documented pair existed on Newton and Isaac; asking for the other raised AttributeError. docs/reference/device-connect.md makes the same promise for the Device Connect stop, whose driver runs on any backend.

The population comes from :meth:_rollouts_in_flight, the one seam the mesh's reporting surfaces also ask, so a peer polled over the wire and a caller holding the engine cannot be told different things about the same instant. MuJoCo's override of that seam delegates to its own registry reader, so the prune this verb used to perform still happens.

Returns:

Type Description
dict[str, Any]

status="success" naming the robots in flight, or reporting that

dict[str, Any]

none are, when this backend reports a population.

dict[str, Any]

status="error" when it reports none at all: "no policies

dict[str, Any]

running" is an affirmative claim, and a backend that cannot

dict[str, Any]

enumerate its rollouts has no evidence for it. That mirrors

dict[str, Any]

meth:stop_policy, which refuses rather than reporting a halt it

dict[str, Any]

cannot stand behind, and names the seam to override.

replay_episode

replay_episode(repo_id: str, robot_name: str | None = None, episode: int = 0, root: str | None = None, speed: float = 1.0, action_key_map: list[str] | None = None) -> dict[str, Any]

Replay a LeRobotDataset episode via PolicyRunner.replay.

robot_name is resolved by the rule :meth:run_policy, :meth:eval_policy and :meth:evaluate_benchmark share: None picks the sole loaded robot, a scene holding several of them returns an error listing the candidates, and a name that IS supplied - the empty string included - is never re-resolved, so one the scene does not hold is reported by name. Replay is the one policy surface that drives the actuators from a recording rather than a policy, so a substituted robot here is a robot the caller never chose being moved.

episode must be a non-negative whole number - the shared domain the replay_episode teleop knob uses - and is rejected with a structured error before the dataset is downloaded. A bool is refused rather than read as an index: episode=True previously resolved episode 1 and replayed it under a "success" status.

speed is a playback-rate multiplier (1.0 = real time) and must be a positive number; a non-positive or non-numeric value is rejected with a structured error rather than raising or silently playing back at full speed. speed scales only the wall-clock playback rate: each recorded frame always advances physics for a full control period (derived from the dataset fps), so a position-servo robot reproduces the recorded trajectory instead of under-integrating it.

action_key_map binds recorded action-vector indices to action keys (default: :meth:robot_action_keys). It must be a non-empty list/tuple of unique strings whose length matches the recorded action width; a bare string, a non-string entry, a duplicate key or a width mismatch is rejected rather than truncated to fit. A "success" status therefore means at least one recorded action actually reached the actuators and every frame that carried one was applied - a frame that send_action could not apply aborts the replay with the frame index, the frames applied so far and the unresolved keys, and an episode whose frames carry no action value at all aborts naming the columns they do carry rather than reporting a replay that commanded nothing. The json block reports frames_with_action beside frames_applied.

Override per backend for optimised replay (e.g. direct ctrl writes) only when measured necessary.

eval_policy

eval_policy(robot_name: str | None = None, policy_provider: str = 'mock', policy_config: dict[str, Any] | None = None, instruction: str = '', n_episodes: int = 1, max_steps: int = 300, success_fn: str | None = None, success_when: dict[str, Any] | None = None, policy_object: Policy | None = None, control_frequency: float | None = None, control_substeps: int | None = None, action_horizon: int = 8, seed: int | None = None, async_rtc: bool = False, rtc_inference_timeout_s: float | None = None, wbc_install_torque_control: bool = True, on_frame: Callable[[int, dict[str, Any], dict[str, Any]], None] | None = None, max_onframe_failures: int | None = None, policy_kwargs: dict[str, Any] | None = None, video: dict[str, Any] | None = None) -> dict[str, Any]

Multi-episode policy evaluation via PolicyRunner.evaluate.

robot_name resolves like :meth:run_policy: None (the default) auto-selects the sole robot in a single-robot scene and errors with the candidate list only when the choice is ambiguous (multiple robots) or impossible (empty scene). This keeps the two sibling entry points consistent - a policy you just ran with run_policy() evals the same way with eval_policy(). n_episodes default lowered from 10 to 1 (callers opt in to longer evals explicitly).

seed pins the eval the way it pins a single :meth:run_policy rollout: the client RNGs are reseeded once from it and then per episode from a master RNG derived from it, and each per-episode seed is forwarded to policy.reset so a service-mode policy can reseed its own process. Two evals at the same seed replay identically for a state-only policy (bit-exact), and to a render tolerance for a camera policy on GPU rendering, where MUJOCO_GL=egl renders a static scene with 1 LSB differences between frames and a VLA's trajectory drifts from them (see :meth:run_policy); compare such evals by success_rate, not frame by frame. None leaves RNG state untouched. Only a non-negative integer can seed those RNGs, so anything else is refused here rather than at the first draw. Each episode's record in the returned episodes list reports the seed that attempt ran on, so a caller reading a single failed episode out of a batch can replay that one rather than the whole eval; it is None when no seed was given, because an unseeded eval derives no per-episode seed to report.

policy_object mirrors :meth:run_policy: pass an already-built Policy to skip the create_policy round-trip (e.g. a loaded SmolVLA checkpoint you want to evaluate without re-instantiating). When omitted, the policy is built from policy_provider / policy_config.

control_frequency / control_substeps flow through to :meth:PolicyRunner.evaluate so the eval loop steps physics for the full control period per action (same servo-tracking semantics as :meth:run_policy). Without these the arm under-steps and the policy looks like a no-op (the arm under-steps each control period). An explicit control_substeps must be a positive integer - 0/negative/float is rejected with a structured error instead of collapsing to a single physics step, which would reinstate that same no-op.

async_rtc (default False) opts into overlapping policy inference with action-chunk execution, evaluating a chunk-emitting policy under the realistic control latency it faces in deployment. It is forwarded to :meth:PolicyRunner.evaluate; the default keeps the success-rate synchronous and bit-stable. It must be a boolean - a value of any other type is reported as a structured caller error rather than read by truthiness, since a truthy "false" would otherwise evaluate under the latency it reads as declining, and a success rate measured that way is not the one the caller asked for. rtc_inference_timeout_s bounds each async inference (structured error instead of a hung rollout). For benchmark-style latency masking use :meth:run_policy (async_rtc=...).

wbc_install_torque_control is the posture :meth:run_policy declares under that name, applied here through the same reader (:meth:_install_action_controller) and checked as the same boolean domain: True (default) installs the torque shim a :class:~strands_robots.policies.wbc.WBCPolicy needs on a position-servo scene for the duration of the call, then uninstalls it. One install covers every episode - both halves of it survive the per-episode reset. A scored rollout that drove the scene differently from an unscored one published the difference as the policy's own success rate, with no field saying which pipeline produced it.

on_frame is an optional (step, observation, action) -> None hook fired per applied control step on the eval thread, immediately after sim.send_action - the success-rate analogue of the :meth:run_policy / :meth:evaluate_benchmark hook. step is a monotonic index that continues across episode boundaries. Use it to record frames or stream telemetry synchronously on the eval thread (e.g. paired with start_cameras_recording_synchronous) so a daemon-thread recorder does not race mjData mutations. A hook exception other than CooperativeStop or :class:~strands_robots.recording_errors.RecordingFrameError is logged at WARN and tolerated up to max_onframe_failures consecutive failures (default 5, the :meth:run_policy ceiling and domain), after which the eval stops with status="error" and onframe_error naming the last failure; a RecordingFrameError is data loss rather than telemetry and propagates on the first occurrence, so the caller learns the episode is incomplete instead of reading a successful eval. Raising :class:~strands_robots.simulation.policy_runner.CooperativeStop stops the evaluation gracefully after the episodes completed so far (the result carries stopped_early=True and episodes_completed), matching :meth:run_policy. That best-effort posture covers a hook that FAILS, not one that cannot be called at all: a non-callable on_frame is a caller error, refused up front with a structured error like :meth:run_policy's observer, because absorbing it per frame would return a success rate the caller's telemetry had watched none of.

n_episodes and max_steps must be positive integers and control_frequency must be > 0; a non-positive value is rejected with a structured error at the entry point (before create_policy) rather than running a degenerate eval that reports a fabricated success rate over zero/negative episodes.

policy_kwargs is the per-call goal payload forwarded verbatim to every policy.get_actions(obs, instruction, **policy_kwargs) call, exactly as on :meth:run_policy. Goal-conditioned providers read their target from these well-known keys (target_velocity for WBC and other locomotion policies; target_pose / target_joints / world_update for cuRobo / MoveIt2 - the issue #300 contract). Without it the eval ran such a policy with an empty goal and reported a meaningless success rate.

success_fn defaults to None. With no success_fn (and no benchmark spec) there is no criterion by which an episode can be marked successful, so success_rate reports a hard 0.0 for every episode regardless of what the policy does - indistinguishable from a policy that genuinely failed every episode. This case logs a warning and sets success_measured=false in the returned json; pass success_fn="contact" (or a callable) to measure real task success.

success_when is the other way to say what success IS: the same predicate DSL as :meth:run_policy's stop_when and a benchmark spec's success clause - {'predicate': 'body_above_z', 'body': 'cube', 'z': 0.2} or an all / any group - compiled through the closed predicate registry and probed against the live scene before the first episode, so a body the scene does not have is refused up front instead of scoring every episode a miss. success_fn (the named 'contact' criterion) and success_when are alternatives; passing both is refused. Before this the only criterion an agent-tool call could express was 'contact', and a predicate spelled as a string ('base_beyond_x:0.5') was refused without saying what IS accepted.

video optionally records one rollout MP4 PER EPISODE so an eval can be watched to see WHY episodes fail, not just read as an aggregate success rate. Same dict schema as :meth:run_policy (path enables it; fps / camera / width / height); the path is validated and the camera probed up-front. _ep{i} is inserted into the filename per episode (eval.mp4 -> eval_ep0.mp4, eval_ep1.mp4, ...) so episodes never overwrite each other, and the written files are returned in the result json video_paths. Recording is unsupported on the benchmark (evaluate_benchmark) path.

Returns:

Name Type Description
dict[str, Any]

The standard agent-tool envelope

dict[str, Any]

``{"status": "success"|"error", "content": [{"text": ...},

dict[str, Any]

{"json": {...}}]}`, read the same way as :meth:run_policy`'s -

dict[str, Any]

by scanning content for the first block with a "json" key,

dict[str, Any]

never by a fixed index (an early caller-error return carries a

dict[str, Any]

text block only).

dict[str, Any]

status reports whether the evaluation RAN, not whether the

dict[str, Any]

policy succeeded: an evaluation in which every episode failed is

dict[str, Any]

still status="success" with success_rate=0.0. The one

dict[str, Any]

two things that make it "error" are a recording it could not

dict[str, Any]

keep (see recording_save_error below) and an evaluation in

dict[str, Any]

which the policy never commanded the robot at all (see

dict[str, Any]

uncommanded_error). Read

dict[str, Any]

success_measured and episodes_successful_at_reset first - it is False when no

dict[str, Any]

success_fn / benchmark spec was supplied, in which case

dict[str, Any]

success_rate is 0.0 for every policy regardless of what it

dict[str, Any]

did and measures nothing.

dict[str, Any]

Fields in the json block:

Outcome dict[str, Any]

success_rate, n_success, success_measured,

dict[str, Any]

episodes_completed, episodes (the per-episode records) and

dict[str, Any]

avg_steps.

dict[str, Any]

episodes_successful_at_reset (int) counts episodes whose success

dict[str, Any]

criterion already held at reset, before any action was applied. The

dict[str, Any]

criterion is sampled only after an applied action, so such an episode

dict[str, Any]

succeeds on its first step whatever the policy commands and its

dict[str, Any]

contribution to success_rate / pass_hat_k describes the scene's

dict[str, Any]

initial state rather than the policy - the mirror of

dict[str, Any]

success_measured=False, which reports a hard 0.0 for the same kind

dict[str, Any]

of reason. Usually a threshold on the wrong side of the initial state

dict[str, Any]

(a lift height below where the object already rests). Every reported

dict[str, Any]

figure is left as measured; reset_success_warning carries the

dict[str, Any]

qualifying text (None when the count is zero) and each per-episode

dict[str, Any]

record carries its own success_at_reset. A partial count is not an

error dict[str, Any]

domain randomisation legitimately draws initial states per

dict[str, Any]

episode.

dict[str, Any]

Commanded actions: actions_applied (actions actually handed to

dict[str, Any]

send_action) beside steps_advanced (control steps the

dict[str, Any]

evaluation advanced), and uncommanded_error - None on every

dict[str, Any]

evaluation that commanded the robot at least once, and the reason

dict[str, Any]

string when it never did. A policy call that returns an empty

dict[str, Any]

action chunk is tolerated per step (physics advances so a

dict[str, Any]

degenerate policy cannot hang the episode), so the two counts

dict[str, Any]

differ whenever any call came back empty and avg_steps alone

dict[str, Any]

cannot tell a scored zero from an unexercised policy. When

dict[str, Any]

actions_applied is 0 the outcome figures describe the

dict[str, Any]

scene's initial state rather than the policy and status is

dict[str, Any]

"error"; a PARTIAL shortfall is reported as a count rather than

dict[str, Any]

refused, since some empty calls are real policy behaviour. Each

dict[str, Any]

per-episode record in episodes carries its own

dict[str, Any]

actions_applied.

Horizon dict[str, Any]

n_episodes, max_steps (the values the evaluation

dict[str, Any]

ran with) and stopped_early.

Recording dict[str, Any]

recording_save_error - None on every healthy

dict[str, Any]

evaluation, and the reason string when a per-episode dataset flush

dict[str, Any]

failed. A failed flush closes the recorder, after which

dict[str, Any]

add_frame writes nothing and counts no drop, so the evaluation

dict[str, Any]

stops at that episode and status is "error":

dict[str, Any]

episodes_completed is then the last episode attempted rather

dict[str, Any]

than n_episodes, and the aggregate covers only those episodes

dict[str, Any]

instead of averaging over ones whose data does not exist.

Physics dict[str, Any]

physics_error - None on every healthy evaluation,

dict[str, Any]

and the backend's divergence report (episode, step, joint) when the

dict[str, Any]

physics diverged and the backend reset the world mid-episode. The

dict[str, Any]

evaluation stops there, the diverged episode is not counted and its

dict[str, Any]

unsaved recording frames are discarded, and status is

dict[str, Any]

"error". onframe_error reports an on_frame hook that hit

dict[str, Any]

max_onframe_failures the same way.

Video dict[str, Any]

video_paths (one MP4 per episode, empty when no

dict[str, Any]

recording was requested).

dict[str, Any]

Policy binding: positional_fallback_used,

dict[str, Any]

generic_state_keys_used and missing_state_keys_used - True

dict[str, Any]

means the policy silently fell back to positional camera routing or

dict[str, Any]

to observation-derived state keys, so the robot moved on

dict[str, Any]

meaningless inputs and the success rate measures nothing about the

dict[str, Any]

policy. See :meth:run_policy for the full contract.

dict[str, Any]

Policy load: policy_load_time_s, policy_load_cache_hit and

dict[str, Any]

policy_resident_rss_mb.

dict[str, Any]

Chunk-prefetch telemetry: chunk_prefetch_enabled (the

dict[str, Any]

background chunk pipeline, not the policy's RTC algorithm, which

dict[str, Any]

policy_rtc_enabled reports), chunk_prefetch_chunks_acquired,

dict[str, Any]

chunk_prefetch_hits, chunk_prefetch_blocks,

dict[str, Any]

avg_inference_ms and max_inference_ms; the pre-rename

dict[str, Any]

rtc_async_enabled, rtc_chunks_acquired,

dict[str, Any]

rtc_prefetch_hits, rtc_prefetch_blocks,

dict[str, Any]

rtc_avg_inference_ms and rtc_max_inference_ms are kept for

dict[str, Any]

one release with the same values.

evaluate_benchmark

evaluate_benchmark(benchmark_name: str, robot_name: str | None = None, policy_provider: str = 'mock', policy_config: dict[str, Any] | None = None, instruction: str = '', n_episodes: int = 1, seed: int | None = None, action_horizon: int = 8, on_frame: Callable[[int, dict[str, Any], dict[str, Any]], None] | None = None, max_onframe_failures: int | None = None, policy_kwargs: dict[str, Any] | None = None, control_frequency: float | None = None, control_substeps: int | None = None, policy_object: Policy | None = None, video: dict[str, Any] | None = None, wbc_install_torque_control: bool = True, async_rtc: bool = False, rtc_inference_timeout_s: float | None = None) -> dict[str, Any]

Run a registered :class:BenchmarkProtocol against the current sim.

Benchmark-agnostic evaluation entry point. Looks up benchmark_name in the global benchmark registry, validates robot compatibility, and forwards to :meth:PolicyRunner.evaluate with the spec. max_steps comes from the benchmark (not a parameter here), so it is validated where it is read rather than at this signature: a benchmark declaring a horizon that is not a positive integer is rejected with a structured error, for the same reason n_episodes is. Both are bounds of the same nested loop, and a non-positive one runs episodes of zero length and then reports a 0% success rate over them.

Parameters:

Name Type Description Default
benchmark_name str

Key from :func:register_benchmark / :func:register_benchmark_from_file.

required
robot_name str | None

Robot to evaluate, resolved by the same rule :meth:run_policy and :meth:eval_policy apply: None picks the sole loaded robot, and a scene holding several of them returns an error listing the candidates. A name that IS supplied is never re-resolved, so one the scene does not hold - including an empty string - is reported by name rather than replaced with a robot the caller did not ask for. Compatibility with the benchmark's own supported_robots is then enforced per episode by :meth:~strands_robots.simulation.benchmark.BenchmarkProtocol.on_episode_start.

None
policy_provider str

Policy provider name (forwarded to :func:create_policy).

'mock'
policy_config dict[str, Any] | None

Provider-specific kwargs.

None
instruction str

Natural-language instruction for the policy.

''
n_episodes int

Number of episodes. Must be a positive integer; a zero/negative/non-int value is rejected with a structured error rather than fabricating a 0%-success report over an empty rollout loop.

1
seed int | None

Master RNG seed for per-episode reproducibility.

None
action_horizon int

How many actions to consume from each policy.get_actions(...) chunk before re-querying the policy. Default 8 matches NVIDIA's upstream GR00T LIBERO eval (MultiStepWrapper with n_action_steps=8) - the policy commits to 8 actions before re-observing, which is what GR00T-N1.7-LIBERO checkpoints were trained against. Set to 1 for closed-loop receding-horizon control (re-observe every step; matches OpenVLA-style eval) ONLY for single-action policies: the interval is clamped up to the policy's execution_horizon (resolve_chunk_length), so a chunk-emitting policy (e.g. a VLA) still consumes its full chunk open-loop regardless of this value. Values < 1 are rejected with a structured error. on_step and success/failure checks run after EACH applied action, so per-step rewards and early termination work correctly regardless of horizon.

8
on_frame Callable[[int, dict[str, Any], dict[str, Any]], None] | None

Optional (step, observation, action) -> None hook fired per applied control step on the eval thread, immediately after sim.send_action. Use this for synchronous recording or telemetry when the eval is dispatched from a thread distinct from the script main (e.g. Strands Agent tool dispatch under asyncio) - the daemon-thread recorder (:meth:~strands_robots.simulation.mujoco.simulation.Simulation.start_cameras_recording) races mjData mutations on the eval thread under that pattern and produces 2-3% frame-capture rates with greenish GL clear-colour artifacts. Pair with :meth:~strands_robots.simulation.mujoco.simulation.Simulation.start_cameras_recording_synchronous for the recorder side. See #191. Raising :class:~strands_robots.simulation.policy_runner.CooperativeStop from the hook ends the benchmark gracefully after the episodes completed so far - the result json carries stopped_early=True and episodes_completed (matching :meth:run_policy / :meth:eval_policy); any in-progress episode's partial video is closed cleanly and is NOT listed in video_paths. A hook exception other than CooperativeStop or :class:~strands_robots.recording_errors.RecordingFrameError is logged at WARN and tolerated up to max_onframe_failures consecutive failures, after which the benchmark stops with status="error" and onframe_error set; a RecordingFrameError is data loss rather than telemetry and propagates on the first occurrence. A value that is not callable at all is refused up front instead, as in :meth:eval_policy.

None
max_onframe_failures int | None

The :meth:run_policy on_frame failure ceiling: a positive integer, or None (default) for 5.

None
policy_kwargs dict[str, Any] | None

Per-call goal payload forwarded verbatim to every policy.get_actions(obs, instruction, **policy_kwargs) call (same contract as :meth:run_policy / :meth:eval_policy). Goal-conditioned providers read their target from these keys (target_velocity / target_pose / target_joints / world_update); a benchmark that drives such a policy must pass them or the policy runs with an empty goal.

None
control_frequency float | None

Target Hz for policy.get_actions calls, used to derive the physics substeps executed per action (round(1 / control_frequency / physics_timestep)) so the benchmark loop steps a full control period per action. Must be > 0; a non-positive value is rejected with a structured error. Defaults to 50.0 (same default as :meth:eval_policy). Set it to the rate the policy was trained/evaluated at - a benchmark's max_steps maps to a wall-clock episode length that depends on this rate, so a mismatched frequency changes the effective episode horizon.

None
control_substeps int | None

Explicit physics substeps per action, overriding the control_frequency-derived value (mirrors :meth:eval_policy). Must be a positive integer; 0, negative, float and bool values are rejected with a structured error rather than collapsing to a single under-integrated physics step. None (default) derives it from control_frequency.

None
policy_object Policy | None

An already-built :class:Policy to evaluate, skipping the create_policy round-trip (mirrors :meth:run_policy / :meth:eval_policy). Use it to benchmark a checkpoint you have already loaded - e.g. a multi-GB VLA - once per process instead of reloading it on every benchmark call. When None the policy is built from policy_provider / policy_config.

None
video dict[str, Any] | None

Optional per-episode rollout MP4 config (same dict schema as :meth:run_policy / :meth:eval_policy: path enables it, plus fps / camera / width / height). One file per episode with _ep{i} inserted into the filename so a benchmark eval can be WATCHED to see why episodes fail, not just read as an aggregate success_rate. Frames are captured synchronously on the eval thread (render is read-only over mjData), so recording does not perturb the bit-stable benchmark rollout. Written paths are returned in the result json video_paths. None (default) records nothing.

None
wbc_install_torque_control bool

The same posture :meth:run_policy takes, checked and applied the same way: when True (default) a :class:~strands_robots.policies.wbc.WBCPolicy evaluated on a position-servo scene gets the torque shim installed for the duration of this call, then uninstalled. One install covers every episode - the controller registration and the actuator mode both survive the per-episode reset. Set False to manage the controller yourself or to drive a torque-actuated scene directly. No-op for non-WBC policies and on backends without the hook.

True
async_rtc bool

Accepted for the same shape as :meth:eval_policy, but only False runs: a benchmark rollout stays synchronous so its success rate is bit-stable. True is refused before any policy is built, with the reason and the sibling that does overlap inference (:meth:run_policy). A non-boolean is refused like every other posture flag.

False
rtc_inference_timeout_s float | None

The async prefetch deadline of :meth:eval_policy, checked against the same domain (None or a positive finite number). The synchronous benchmark path has no prefetch, so a valid value changes nothing.

None

Returns:

Name Type Description
dict[str, Any]

Standard status dict. On success, carries per-episode cumulative

dict[str, Any]

reward + aggregate success_rate / avg_reward / avg_steps in the

dict[str, Any]

JSON payload, plus video_paths (the per-episode MP4s written

dict[str, Any]

when video is set).

dict[str, Any]

recording_save_error is None on every healthy run and

dict[str, Any]

carries the reason when a per-episode dataset flush failed, in

dict[str, Any]

which case the benchmark stops at that episode and status is

dict[str, Any]

"error" - see :meth:eval_policy, which reports it the same

dict[str, Any]

way.

dict[str, Any]

physics_error is None on every healthy run and carries the

dict[str, Any]

divergence report when the physics diverged mid-episode; the

dict[str, Any]

benchmark stops the same way :meth:eval_policy does, and on an

dict[str, Any]

on_frame hook that hit max_onframe_failures (onframe_error).

dict[str, Any]

actions_applied (actions actually handed to send_action),

dict[str, Any]

steps_advanced (control steps the benchmark advanced) and

dict[str, Any]

uncommanded_error report the same fact :meth:eval_policy

dict[str, Any]

reports under those names, and each per-episode record carries its

dict[str, Any]

own actions_applied. A policy call returning an empty action

dict[str, Any]

chunk is tolerated per step, so the counts differ whenever any call

dict[str, Any]

came back empty; when actions_applied is 0 the benchmark

dict[str, Any]

never commanded the robot, success_rate / avg_reward

dict[str, Any]

describe the scene's initial state rather than the policy, and

dict[str, Any]

status is "error".

dict[str, Any]

episodes_successful_at_reset (int) counts episodes whose success

dict[str, Any]

criterion already held at reset, before any action was applied. The

dict[str, Any]

criterion is sampled only after an applied action, so such an episode

dict[str, Any]

succeeds on its first step whatever the policy commands and its

dict[str, Any]

contribution to success_rate / pass_hat_k describes the scene's

dict[str, Any]

initial state rather than the policy - the mirror of

dict[str, Any]

success_measured=False, which reports a hard 0.0 for the same kind

dict[str, Any]

of reason. Usually a threshold on the wrong side of the initial state

dict[str, Any]

(a lift height below where the object already rests). Every reported

dict[str, Any]

figure is left as measured; reset_success_warning carries the

dict[str, Any]

qualifying text (None when the count is zero) and each per-episode

dict[str, Any]

record carries its own success_at_reset. A partial count is not an

error dict[str, Any]

domain randomisation legitimately draws initial states per

dict[str, Any]

episode. episodes_successful_at_reset and

dict[str, Any]

reset_success_warning report the same fact :meth:eval_policy

dict[str, Any]

reports under those names.

dict[str, Any]

episodes_failed_at_reset (int) is its mirror for the spec's

dict[str, Any]

failure clause, sampled at the same probe and reported the same

way dict[str, Any]

a clause already satisfied at reset ends the episode on its

dict[str, Any]

first step whatever the policy commands, so success_rate reports

dict[str, Any]

a hard 0.0 for it. It costs more than the success mirror, because the

dict[str, Any]

eval loop reads is_failure BEFORE is_success - such an episode

dict[str, Any]

is scored a failure with the success criterion never consulted, and in

dict[str, Any]

the report it is indistinguishable from a policy that immediately did

dict[str, Any]

something catastrophic. Usually a threshold on the wrong side of the

dict[str, Any]

initial state (a fall height above where the object already rests, or

dict[str, Any]

a base-collapse height above the robot's spawned stance). Every

dict[str, Any]

reported figure is left as measured; reset_failure_warning carries

dict[str, Any]

the qualifying text (None when the count is zero) and each

dict[str, Any]

per-episode record carries its own failure_at_reset. A partial

dict[str, Any]

count is not an error, for the reason the success mirror's is not.

dict[str, Any]

Only this route reports it: :meth:eval_policy takes a success_fn

dict[str, Any]

and has no failure criterion to sample.

list_benchmarks

list_benchmarks() -> dict[str, Any]

Enumerate registered benchmarks.

Returns a standard status dict whose JSON payload contains the :func:~strands_robots.simulation.benchmark.list_benchmarks metadata snapshot. Safe to call from any backend; the registry is engine-agnostic.

register_benchmark_from_file

register_benchmark_from_file(benchmark_name: str, spec_path: str) -> dict[str, Any]

Load a declarative benchmark spec from disk and register it.

Wraps :func:strands_robots.simulation.benchmark_spec.register_benchmark_from_file so agents can author benchmarks as YAML / JSON at runtime. Parsing errors surface as structured error dicts rather than exceptions.

register_builtin_benchmarks

register_builtin_benchmarks() -> dict[str, Any]

Register the built-in benchmark specs shipped with strands_robots.

Wraps :func:strands_robots.simulation.builtin_benchmarks.register_builtin_benchmarks so the shipped specs become discoverable via :meth:list_benchmarks and runnable via :meth:evaluate_benchmark without hand-authoring a spec file. Ships a canonical velocity-tracking locomotion benchmark (go2_walk_forward) composed from the floating-base predicate/reward DSL. Opt-in and idempotent (mirrors the on-demand LIBERO suite registration); importing strands_robots performs no registry mutation.

Returns:

Type Description
dict[str, Any]

A status dict whose JSON payload carries the registered list of

dict[str, Any]

benchmark names now available to :meth:evaluate_benchmark.

load_scene

load_scene(scene_path: str) -> dict[str, Any]

Load a complete scene from file. Override per backend.

randomize

randomize(**kwargs: Any) -> dict[str, Any]

Apply domain randomization.

Concrete backends define their own parameter signatures. Because this base signature is **kwargs-typed, an override inherits a sink that would swallow any keyword it does not declare; backends must reject the residual keys (see :func:unknown_kwargs_error) so a misspelled axis cannot report success while leaving that axis untouched. Override per backend.

set_obs_noise

set_obs_noise(**kwargs: Any) -> dict[str, Any]

Configure additive sensor noise on observations.

Models real-sensor measurement noise (joint encoders, camera frames) so policies are not trained on noise-free observations. Concrete backends define their own parameter signatures and, as for :meth:randomize, must reject keywords they do not declare rather than let this **kwargs-typed signature swallow them. Override per backend.

get_contacts

get_contacts() -> dict[str, Any]

Get contact information. Override per backend.

Returns:

Type Description
dict[str, Any]

The agent-tool envelope -- {"status": ..., "content": [...]} --

dict[str, Any]

whose json content block carries contacts, a list of

dict[str, Any]

per-contact records. The payload lives in that block, not on the

dict[str, Any]

envelope itself, so a caller reading result["contacts"]

dict[str, Any]

directly always misses. The predicate DSL's contact_*

dict[str, Any]

factories (see

dict[str, Any]

mod:strands_robots.simulation.predicates) are the supported

dict[str, Any]

readers; success_fn="contact" on

dict[str, Any]

meth:~strands_robots.simulation.policy_runner.PolicyRunner.evaluate

dict[str, Any]

shares them.

Raises:

Type Description
NotImplementedError

Backends that expose no contact list.

get_frame

get_frame(camera_name: str = 'default', width: int | None = None, height: int | None = None) -> tuple[np.ndarray, np.ndarray | None]

Render a camera to raw (rgb, depth) ndarrays.

The numeric-array counterpart of :meth:render (which wraps pixels in the agent-tool PNG envelope). In-process consumers -- the hybrid compositor, dataset recorders, video writers -- use this to get pixels without a PNG round-trip.

Parameters:

Name Type Description Default
camera_name str

name of a camera previously added via add_camera (backends supporting a free camera also accept their free-cam tokens for the RGB path).

'default'
width int | None

image width in pixels; None uses the camera's configured resolution.

None
height int | None

image height in pixels; None uses the camera's configured resolution.

None

Returns:

Type Description
ndarray

(rgb, depth) where rgb is (H, W, 3) uint8 and

ndarray | None

depth is (H, W) float32 metric meters, or None on

tuple[ndarray, ndarray | None]

backends with no depth path (Newton). Backends must never

tuple[ndarray, ndarray | None]

substitute silently wrong pixels -- failures raise.

Raises:

Type Description
KeyError

unknown camera name.

ValueError

invalid render dimensions.

RuntimeError

no world / renderer unavailable / backend render failure.

NotImplementedError

backend has no raw-frame path.

get_camera_params

get_camera_params(camera_name: str = 'default', width: int | None = None, height: int | None = None) -> CameraParams

Return pinhole intrinsics/extrinsics for a named camera.

The returned :class:strands_robots.rendering.CameraParams carries the intrinsic matrix K (pixels), the world-from-camera SE(3) pose T_world_cam in the OpenGL optical convention (+X right, +Y up, -Z forward), the image size, and the clip planes. Backends whose native camera basis differs (e.g. Isaac's USD camera prim) apply the fixed basis correction here so consumers never see a backend-specific frame.

Parameters:

Name Type Description Default
camera_name str

name of a camera previously added via add_camera. Backends with a free camera also accept their free-cam tokens here, reporting the same view :meth:get_frame renders, so the two APIs stay symmetric (MuJoCo: None / "" / "default" / "free").

'default'
width int | None

image width to compute K for; None uses the camera's configured resolution.

None
height int | None

image height to compute K for; None uses the camera's configured resolution.

None

Raises:

Type Description
KeyError

unknown camera name.

ValueError

a camera whose projection no pinhole K can represent (e.g. an orthographic camera), or a resolution the backend cannot honor.

RuntimeError

no world created.

NotImplementedError

backend has no camera-params path.

get_world_point

get_world_point(camera_name: str = 'default', pixels: Sequence[Sequence[SupportsFloat]] | None = None, width: int | None = None, height: int | None = None) -> dict[str, Any]

Ground image pixels to metric world coordinates via the depth buffer.

The perception half of deployment-shaped grounding (Harness VLA, arXiv:2607.08448, Appendix E.2): instead of reading privileged object poses (:meth:get_body_state -- sim-only oracle truth), the agent picks pixels on the visible surface of the target in the RGB frame and this call unprojects each one through the pixel-aligned metric depth buffer -- p_cam = depth * K^-1 @ [u, v, 1] in the OpenGL optical frame, then p_world = T_world_cam @ p_cam. The same call shape works on hardware with an RGB-D camera, so grounding built on it transfers.

Guidance for agents (the paper's localization rule):

  1. Render the camera first (render / get_frame) and pick pixels ON the visible surface of the target object.
  2. Avoid rims, edges, reflections, transparent surfaces, and background pixels -- depth there is unstable or belongs to something else.
  3. Sample SEVERAL pixels on the same surface (typically 3-9): the returned point is the median over the valid samples, which rejects stray outliers. The median is PER-COMPONENT, so on a strongly tilted surface the combined [x, y, z] may lie on no single sampled point - treat it as a robust surface estimate, not as one of points.
  4. Pixels with no depth (background / far plane) are dropped, not zero-filled; check n_valid against the count you sent.
  5. Re-localize after any robot, camera, or object motion -- world points are snapshots, not tracks.

Depth samples are treated as z-depth (distance along the optical axis), the convention every in-tree backend emits. Pixels are indexed [u, v] with u the column from the left and v the row from the top; the unprojection uses the pixel center (u + 0.5, v + 0.5).

Atomicity: when the backend exposes an engine lock (self._lock, all in-tree backends), the frame render and the camera-params read happen under it, so a concurrent scene mutation cannot slip between the two. All failures return a structured error dict -- this is a tool-envelope method and never raises.

Parameters:

Name Type Description Default
camera_name str

a camera previously added via add_camera (backends with a free camera also accept their free-cam tokens, as for :meth:get_frame).

'default'
pixels Sequence[Sequence[SupportsFloat]] | None

non-empty list of [u, v] pixel coordinates (integer-valued; at most _WORLD_POINT_MAX_PIXELS).

None
width int | None

image width; None uses the camera's configured resolution.

None
height int | None

image height; None uses the camera's configured resolution.

None

Returns:

Type Description
dict[str, Any]

On success ``{"status": "success", "content": [{"text": ...},

dict[str, Any]

{"json": {"point": [x, y, z], "points": [...], "n_valid": int,

dict[str, Any]

"n_requested": int, "camera": str, "width": int, "height": int}}]}``

dict[str, Any]

where point is the per-component median over the valid

dict[str, Any]

samples and points is aligned with the input pixels

dict[str, Any]

(None where the pixel had no valid depth). Backends without a

dict[str, Any]

metric-depth path (Newton), all-invalid pixel sets, out-of-bounds

dict[str, Any]

pixels, malformed input, and a failed frame render or

dict[str, Any]

camera-params read all return

dict[str, Any]

{"status": "error", "content": [{"text": ...}]}, with the

dict[str, Any]

two backend reads reporting distinguishable text so a caller

dict[str, Any]

knows which one failed. The camera-params read can fail on input

dict[str, Any]

this call already accepted and a frame it already rendered --

dict[str, Any]

most notably a camera whose projection no pinhole K can

dict[str, Any]

represent, such as MuJoCo's orthographic free camera, which

dict[str, Any]

renders normally but has no intrinsics. So check status

dict[str, Any]

rather than inferring success from a valid pixel set.

describe

describe() -> dict[str, Any]

Return a machine-readable summary of this engine's live contract.

Agents should call this first to learn what robots exist, what cameras are attached, and the signatures of the methods most commonly needed -- in a single call, instead of guessing method names.

Returns:

Type Description
dict[str, Any]

Plain dict with keys: robots, capabilities, cameras, methods, note.

cleanup

cleanup() -> None

Release all resources.

Called on context exit, and best-effort from :meth:__del__ for an engine whose __init__ ran to completion. Implementations are written against a fully-constructed instance: a caller whose __init__ raised part-way must release whatever it acquired itself rather than relying on the finalizer.

Capabilities

Capability vocabulary a :class:~strands_robots.simulation.base.SimEngine backend declares.

The names form a closed set; a backend may add "<vendor>:<name>" extras. A missing capability is reported with :data:UNSUPPORTED_BY_BACKEND, which no operator grant can lift, so it is not a continuable refusal code. A declaration must name the four :data:CORE_CAPABILITIES; an undeclared backend is credited with all eight names of :data:DEFAULT_CAPABILITIES. Pure stdlib, and it imports nothing from the package, so the engine is described by :class:CapabilityReporter rather than by the base class.

CapabilityNotSupported

CapabilityNotSupported(capability: str, member: str)

Bases: NotImplementedError

A member was called on a backend that lacks the capability it needs.

Raised only by list-returning members; others return :func:unsupported_result.

Parameters:

Name Type Description Default
capability str

The missing capability name.

required
member str

The SimEngine member that was called.

required

CapabilityReporter

Bases: Protocol

Anything that reports its capabilities, such as a SimEngine.

capabilities

capabilities() -> frozenset[str]

Return the capability names this object supports.

ManipulationOptional

Mix in before SimEngine for a backend with no joints, objects, rendering or rollout.

Declares only the four core capabilities and supplies the manipulation-only abstract members as standard refusals, so such a backend implements only what it can honour. It is for backends without these features, not for layering over a full backend. To support one of them, override the member and add its name to CAPABILITIES; declaring the name while the refusal is still inherited raises TypeError at class creation.

add_object

add_object(name: str, shape: str = 'box', position: list[float] | None = None, orientation: list[float] | None = None, size: list[float] | None = None, color: list[float] | None = None, mass: float = 0.1, is_static: bool | None = None, mesh_path: str | None = None, material: dict[str, Any] | None = None) -> dict[str, Any]

Refuse with :func:unsupported_result: no :data:OBJECTS capability. Arguments are ignored.

remove_object

remove_object(name: str) -> dict[str, Any]

Refuse with :func:unsupported_result: no :data:OBJECTS capability. Arguments are ignored.

render

render(camera_name: str = 'default', width: int | None = None, height: int | None = None) -> dict[str, Any]

Refuse with :func:unsupported_result: no :data:RENDER capability. Arguments are ignored.

robot_joint_names

robot_joint_names(robot_name: str) -> list[str]

Raise, since an empty list would read as a robot with no joints.

Raises:

Type Description
CapabilityNotSupported

Always; this backend has no :data:JOINTS capability.

check_capabilities

check_capabilities(sim: CapabilityReporter, required: Iterable[str], *, caller: str) -> dict[str, Any] | None

Check that sim has every capability in required.

Parameters:

Name Type Description Default
sim CapabilityReporter

The engine to check; any object with a capabilities() method.

required
required Iterable[str]

Capability names the caller needs.

required
caller str

The member or tool doing the check, named in the result.

required

Returns:

Type Description
dict[str, Any] | None

None when every capability is present, otherwise an error result

dict[str, Any] | None

whose json block lists the absent names under missing.

Raises:

Type Description
TypeError

required is a single str rather than a collection.

unsupported_result

unsupported_result(capability: str, member: str, backend: str) -> dict[str, Any]

Build the standard error result for a member the backend does not support.

Parameters:

Name Type Description Default
capability str

The missing capability name.

required
member str

The SimEngine member that was called.

required
backend str

The backend's name, usually its class name.

required

Returns:

Type Description
dict[str, Any]

A tool-result dict with status="error" and a json block carrying

dict[str, Any]

code, capability, member and backend.

World model

Dataclasses for simulation state.

These dataclasses provide a backend-independent typed state representation consumed by simulation engine implementations (e.g. MuJoCo, Isaac Sim, PyBullet).

They enable
  • Type-safe state tracking across simulation steps.
  • Serialisation for checkpoints and trajectory recording.
  • A backend-independent interface for agent tools.

They are defined alongside the SimEngine ABC because its method signatures reference them (e.g. create_world() → SimWorld).

SimWorld dataclass

SimWorld(robots: dict[str, SimRobot] = dict(), objects: dict[str, SimObject] = dict(), cameras: dict[str, SimCamera] = dict(), timestep: float = 0.002, gravity: list[float] = (lambda: [0.0, 0.0, -9.81])(), ground_plane: bool = True, terrain: str | None = None, terrain_difficulty: float = 1.0, status: SimStatus = SimStatus.IDLE, sim_time: float = 0.0, step_count: int = 0, _model: Any = None, _data: Any = None, _backend_state: dict[str, Any] = dict(), _checkpoints: dict[str, Any] = dict(), _recompile_generation: int = 0)

Complete simulation world state.

Backend-independent state with engine-specific internals kept in three escape hatches, each with a distinct role so backend implementers know which to use:

  • _model: the physics engine's core model handle - the single compiled/loaded representation of the scene (e.g. mujoco.MjModel, Isaac's Scene, PyBullet's body registry). Every backend has one.
  • _data: the physics engine's core simulation state handle - the mutable per-step state companion to _model (e.g. mujoco.MjData, Isaac's World). Every backend has one.
  • _backend_state: a catch-all dict for everything else the backend needs to persist - generated XML, temp dirs, recording buffers, caches, etc. Prefer this over adding new fields here.

All three are typed Any/dict so nothing leaks engine-specific types into this base module.

SimRobot dataclass

SimRobot(name: str, urdf_path: str, position: list[float] = (lambda: [0.0, 0.0, 0.0])(), orientation: list[float] = (lambda: [1.0, 0.0, 0.0, 0.0])(), data_config: str | None = None, body_id: int = -1, joint_ids: list[int] = list(), joint_names: list[str] = list(), actuator_ids: list[int] = list(), namespace: str = '', home_qpos: dict[str, list[float]] = dict(), home_actuators: dict[str, tuple[float, list[float]]] = dict(), policy_running: bool = False, policy_stops: int = 0, policy_claim_stops: int | None = None, policy_steps: int = 0, policy_instruction: str = '', mesh: Any = None, peer_id: str = '', _world: Any = None, _sim_parent: Any = None)

A robot instance within the simulation.

mesh / peer_id: when the parent Simulation is itself attached to a Zenoh mesh, every robot added via add_robot auto-joins as its own peer so the agent can address it directly (e.g. robot_mesh tell target=<peer_id>) instead of having to talk to the sim container and then route by robot name. Both fields stay None / "" for stand-alone sims that are not on a mesh.

request_policy_stop

request_policy_stop() -> bool

Cooperatively stop this robot's rollout, durably, and report what was in flight.

A stop is a fact about the rollout and not only a flag: lowering policy_running alone was overwritten by a worker that had not yet reached its first frame, so the stop was reported and then discarded and the rollout ran to full duration (#2833). Moving policy_stops puts the launcher's claim out of date, which is what stops the worker raising the flag back over this stop.

EVERY stop path goes through here rather than assigning policy_running = False - the stop_policy action, remove_robot, teardown, and the Device Connect stop / emergency-stop handlers - so they cannot drift to different answers about whether a rollout was halted.

Returns:

Type Description
bool

Whether this robot's policy_running flag was raised when the stop

bool

arrived. A caller that also tracks rollout Futures should treat its

bool

own registry as part of the answer (see

bool

meth:~strands_robots.simulation.mujoco.simulation.MuJoCoSimEngine.stop_policy).

SimObject dataclass

SimObject(name: str, shape: str, position: list[float] = (lambda: [0.0, 0.0, 0.0])(), orientation: list[float] = (lambda: [1.0, 0.0, 0.0, 0.0])(), size: list[float] = (lambda: [0.05, 0.05, 0.05])(), color: list[float] = (lambda: [0.5, 0.5, 0.5, 1.0])(), mass: float = 0.1, mesh_path: str | None = None, material: dict[str, Any] | None = None, body_id: int = -1, is_static: bool = False, _original_position: list[float] = list(), _original_color: list[float] = list())

An object in the simulation scene.

SimCamera dataclass

SimCamera(name: str, position: list[float] = (lambda: [1.0, 1.0, 1.0])(), target: list[float] = (lambda: [0.0, 0.0, 0.0])(), fov: float = 60.0, width: int = 640, height: int = 480, camera_id: int = -1, origin_robot: str = '', parent_body: str = '')

A camera in the simulation.

origin_robot: when the camera was discovered inside a robot's URDF during add_robot, this is set to the robot's name so the scene builder knows NOT to re-add the camera at the top level (it'll be re-introduced via spec.attach(robot_spec)). For user-added cameras (via the add_camera tool action) this stays empty.

A discovered camera belongs to exactly one robot: the one whose namespace prefixes name. Removing a robot removes its cameras and only its cameras, so origin_robot must never name a robot outside name's namespace - otherwise the wrong robot's departure strands or drops the entry.

SimStatus

Bases: Enum

Simulation execution status.

TrajectoryStep dataclass

TrajectoryStep(timestamp: float, sim_time: float, robot_name: str, observation: dict[str, Any], action: dict[str, Any], instruction: str = '')

A single step in a recorded trajectory.

Run-policy observers

Read-only rollout events for :meth:PolicyRunner.run.

Why this is a second lane rather than a use of on_frame

on_frame looks like the observation seam and is not one. It is owned by the backend for the duration of a rollout: MuJoCo's hook raises :class:~strands_robots.simulation.policy_runner.CooperativeStop so stop_policy can interrupt, appends the trajectory mirror, publishes mesh step telemetry and drives the LeRobot dataset recorder; Isaac's and Newton's do the recording half of the same job. There is exactly one of them - :meth:~strands_robots.simulation.base.SimEngine.run_policy obtains it from _make_run_policy_hook and does not accept one from the caller - so a consumer that supplied its own would not add observation, it would silently remove cancellation and recording.

These events are therefore emitted beside that hook, and deliberately report the three things the hook's own signature cannot carry:

  • what the backend answered - send_action's per-key verdict, normalised to :data:ActionResolution so a consumer never parses a backend envelope;
  • how old the observation was - observation_age_steps is the authoritative number of completed rollout action attempts since the snapshot was sampled; observation_is_chunk_reused only says that the same chunk-start snapshot is being used by a later action in that chunk;
  • what the legacy hook did - including the step it aborted on, which the legacy step accounting excludes (see :class:RunPolicyStep).

The payload-ownership rule

RunPolicyStep.observation and RunPolicyStep.action are borrowed: the same objects the legacy hook received, not copies. Copying them per step would put an image-sized deep copy on the control path of every rollout that enables the lane, which is the opposite of what an observability lane should cost. So the contract is placed on the consumer instead, and it is narrow:

  • Treat both as read-only. A backend may reuse the same buffers next step, and the dataset recorder reads them after you do.
  • Do not retain them past the call. Snapshot the few fields you need (synchronously, inside the callback) if you intend to hand them to a queue, a thread or a socket.

An event is dispatched synchronously on the rollout thread. That makes the lane cheap and ordered, and it means a consumer that blocks, blocks the robot. Ordinary exceptions and CooperativeStop are contained - those raises never change the rollout's outcome, and never reach the on_frame consecutive-failure watchdog, which exists for a recorder losing dataset frames (GH #117) rather than for a visualiser that cannot draw. Process-control and cancellation BaseException classes propagate after terminal dispatch is attempted. Containment is not isolation: this is telemetry, not a sandbox. Keep the callback short and non-blocking.

Ordering guarantees

event_seq is dense and 0-based within one run_id, so a gap is observable. monotonic_ns is the ordering clock (a date -s or an NTP correction cannot move it); utc_ns is derived from a single rollout anchor so a wall-clock label never reorders the stream.

Scope

Single-policy simulation rollouts through :meth:PolicyRunner.run and every episode of :meth:~strands_robots.simulation.base.SimEngine.run_policy (one lifecycle and run_id per episode). eval_policy, evaluate_benchmark, run_multi_policy and hardware carry no observer yet - they are separate loops with different step semantics, and claiming them here would promise coverage this module does not have.

Example::

from strands_robots.simulation.observers import RunPolicyStep

def watch(event):
    if isinstance(event, RunPolicyStep) and event.action_resolution != "full":
        print(event.applied_action_index, event.unresolved_action_keys)

sim.run_policy(robot_name="arm", observer=watch)

RunPolicyObserver module-attribute

RunPolicyObserver = Callable[[RunPolicyEvent], None]

RunPolicyEvent module-attribute

RunPolicyEvent = RunPolicyStarted | RunPolicyStep | RunPolicyEnded

RunPolicyOutcome module-attribute

RunPolicyOutcome = Literal['success', 'error']

StoppedReason module-attribute

StoppedReason = Literal['budget', 'predicate', 'cancelled', 'error']

ActionResolution module-attribute

ActionResolution = Literal['full', 'partial', 'none', 'unknown']

RunPolicyStarted dataclass

RunPolicyStarted(schema_version: int, run_id: str, event_seq: int, monotonic_ns: int, utc_ns: int, robot_name: str, policy: str, instruction: str, control_frequency: float, action_horizon: int, total_steps: int, async_rtc: bool)

Opens a rollout. Emitted once, after setup, before the first observation.

Emitted only for a rollout that actually begins: a request refused in pre-flight (a bad horizon, an unusable seed, a video path that cannot be opened) raises or returns before this event, so no lifecycle is opened and none has to be closed.

Attributes:

Name Type Description
schema_version int

:data:SCHEMA_VERSION at emission time.

run_id str

Identifies this rollout. Every event of one :meth:PolicyRunner.run call shares it; a multi-episode run_policy produces one run_id per episode, because each episode is its own runner call.

event_seq int

0 - the first event of the rollout.

monotonic_ns int

Ordering clock, from :func:time.monotonic_ns.

utc_ns int

Wall-clock label for the same instant, derived from the rollout's single (utc, monotonic) anchor.

robot_name str

Robot being driven.

policy str

Class name of the driving policy (e.g. "MockPolicy").

instruction str

Natural-language instruction forwarded to the policy.

control_frequency float

Target Hz of the control loop.

action_horizon int

Max actions consumed per policy call, as requested. The effective chunk may be longer when the policy declares a larger actions_per_step.

total_steps int

Step budget resolved for this rollout.

async_rtc bool

Whether inference is overlapped with action execution. This is the resolved value, so a rollout that auto-detected a chunk-emitting policy reports True even though the caller passed None.

RunPolicyStep dataclass

RunPolicyStep(schema_version: int, run_id: str, event_seq: int, monotonic_ns: int, utc_ns: int, applied_action_index: int, legacy_step_index: int, observation: dict[str, Any], action: Any, observation_is_chunk_reused: bool, observation_age_steps: int, action_resolution: ActionResolution, applied_action_keys: tuple[str, ...], unresolved_action_keys: tuple[str, ...], elapsed_s: float, sim_time_s: float | None, legacy_hook_outcome: LegacyHookOutcome)

One completed send_action call.

Emitted after send_action has returned and after the legacy on_frame hook has run, whatever that hook did. The backend's physical state is known to have advanced only when its result says so; on a coarse atomic refusal the event reports action_resolution="unknown" rather than inventing applied or unresolved keys.

The hook runs before the legacy step_count increments. Consequently :attr:applied_action_index and :attr:legacy_step_index identify the same zero-based action, including the action whose hook cancels or loses a dataset frame. The abort is identified by :attr:legacy_hook_outcome. Terminal accounting is intentionally different: :attr:RunPolicyEnded.applied_actions counts calls made to send_action, while :attr:RunPolicyEnded.legacy_steps_used excludes a hook-aborted final step.

Attributes:

Name Type Description
schema_version int

:data:SCHEMA_VERSION at emission time.

run_id str

The rollout this step belongs to.

event_seq int

Dense, monotonic position in the rollout's event stream.

monotonic_ns int

Ordering clock, sampled after send_action completed.

utc_ns int

Wall-clock label for the same instant.

applied_action_index int

0-based count of send_action calls, including this one. Dense across the whole rollout.

legacy_step_index int

The index this step's on_frame call received, or the index it would have received had a hook been installed. Equal to applied_action_index; use legacy_hook_outcome to identify an aborting step.

observation dict[str, Any]

Borrowed pre-action observation - the same object the legacy hook received. Read-only; do not retain past the call. See the module docstring.

action Any

Borrowed action sent to the backend. Usually a dict[str, float]; a numeric vector for policies that bind positionally. Same ownership rule as observation.

observation_is_chunk_reused bool

True only when this is a later action using the same chunk-start snapshot. It is not authoritative freshness: the first action after an async prefetch swap has not reused that snapshot within its new chunk, but the snapshot is already old.

observation_age_steps int

Authoritative nonnegative count of completed rollout action attempts since observation was sampled. 0 means no intervening action attempt; a positive value means the snapshot is older in control-step terms. It does not claim physical advancement when action_resolution is "unknown". Dataset recording refreshes the observation per step and therefore reports 0 throughout.

action_resolution ActionResolution

A :data:ActionResolution. "partial" and "none" require a complete explicit backend per-key breakdown; a coarse error is "unknown".

applied_action_keys tuple[str, ...]

Keys that drove an actuator this step.

unresolved_action_keys tuple[str, ...]

Keys the backend could not absorb. Empty on the success path and on "unknown" resolutions without a breakdown.

elapsed_s float

Seconds since the rollout's monotonic start.

sim_time_s float | None

Backend simulation clock after the action, when the backend exposes one cheaply; None otherwise. Never fetched at the cost of an extra backend call.

legacy_hook_outcome LegacyHookOutcome

A :data:LegacyHookOutcome.

RunPolicyEnded dataclass

RunPolicyEnded(schema_version: int, run_id: str, event_seq: int, monotonic_ns: int, utc_ns: int, outcome: RunPolicyOutcome, stopped_reason: StoppedReason, applied_actions: int, legacy_steps_used: int, action_errors: int, elapsed_s: float, error_type: str | None, error_message: str | None, observer_failures: int = 0)

Closes a rollout. Emitted once, if and only if :class:RunPolicyStarted was.

Attempted on every exit path a started rollout can take - budget exhausted, predicate fired, cooperative stop, or any error - so a consumer can always pair an open with a close. The one thing it cannot survive is the process dying under it (SIGKILL, OOM, a consumer that blocks forever), which is why this lane is telemetry rather than a durable record.

Attributes:

Name Type Description
schema_version int

:data:SCHEMA_VERSION at emission time.

run_id str

The rollout being closed.

event_seq int

Final position in the rollout's event stream.

monotonic_ns int

Ordering clock at close.

utc_ns int

Wall-clock label for the same instant.

outcome RunPolicyOutcome

"success" or "error", matching the rollout result's status.

stopped_reason StoppedReason

A :data:StoppedReason, matching the result payload's own field.

applied_actions int

Total send_action calls made. Equals the number of :class:RunPolicyStep events whose dispatch was attempted, and may exceed :attr:legacy_steps_used by one when the legacy hook aborted the final step. On action_resolution="unknown" this count does not claim that physical state changed.

legacy_steps_used int

The rollout's own steps_used / n_steps.

action_errors int

Steps whose send_action reported an error, partial resolutions included.

elapsed_s float

Rollout duration, measured on the monotonic clock.

error_type str | None

Exception class name when outcome == "error", else None.

error_message str | None

Exception message when outcome == "error", else None. Truncated; the full traceback stays in the log.

observer_failures int

Contained failures before this terminal dispatch. The returned result payload is authoritative and also includes a contained failure raised while consuming this Ended event itself.

Models and assets

Robot model resolution - URDF registry + asset manager.

Bridges the robot registry with actual URDF/MJCF files on disk.

Resolution order for :func:resolve_model: 1. User-registered URDFs (:func:register_urdf) 2. URDF search paths (STRANDS_ASSETS_DIR, CWD, etc.) 3. Asset manager (robot_descriptions - fallback for standard robots)

resolve_model

resolve_model(name: str, prefer_scene: bool = True, *, allow_download: bool = True) -> str | None

Resolve a robot name or data_config to an MJCF/URDF model path.

Resolution order (local assets take priority): 1. User-registered URDFs (custom user registrations) 2. URDF search paths (STRANDS_ASSETS_DIR, CWD, etc.) 3. Asset manager (robot_descriptions - fallback for standard robots)

Step 3 fetches an asset that is not on disk - the right default for a caller about to load the model. allow_download=False hands the same decline to :func:~strands_robots.assets.manager.resolve_model_path, so a caller that reports on assets reads the disk and reaches neither the network nor the robot_descriptions import that clones on a cold cache. Steps 1 and 2 are filesystem reads either way.

resolve_urdf

resolve_urdf(data_config: str) -> str | None

Resolve a data_config name to a URDF file path.

Also checks the registry's legacy_urdf field - a backward-compatible path for robots that were registered before the MJCF asset system was introduced (e.g. robots originally configured with raw URDF paths).

register_urdf

register_urdf(data_config: str, urdf_path: str) -> None

Register a URDF/MJCF file for a data_config name.

list_registered_urdfs

list_registered_urdfs() -> dict[str, str | None]

List all registered URDF mappings and their resolved paths.

list_available_models

list_available_models() -> str

List all available robot models (Menagerie + custom).

Both halves are always reported. The asset-manager table alone used to be returned whenever the asset manager was importable - which is every normal install - so a caller who had just registered an asset with :func:register_urdf was told by the discovery surface that it did not exist, while :func:resolve_urdf resolved it and add_robot spawned it. The registered section is omitted entirely when nothing is registered, so a default install's listing is unchanged.

Returns:

Name Type Description
str

The built-in robot table, followed by a Registered URDFs: section

when str

func:register_urdf has been called. Without the asset manager

str

only the registered section is available, so that is returned alone.

Predicates and benchmarks

Named-predicate library for declarative :class:BenchmarkProtocol specs.

Each entry in :data:PREDICATE_REGISTRY is a factory (**kwargs) -> callable where the returned callable takes a :class:SimEngine and returns either bool (for success/failure predicates) or float (for reward terms).

The registry is a closed set - the YAML/JSON loader in :mod:strands_robots.simulation.benchmark_spec refuses predicates whose name is not in this registry, so spec files are safe to parse from untrusted / LLM-authored input. No eval is ever called. User-defined predicates must be registered programmatically via :func:register_predicate before loading the spec.

Predicates are backend-aware but not backend-specific: they exclusively call SimEngine methods (abstract) or probe for MuJoCo-only methods via getattr and return a safe fallback (False / 0.0) when the backend does not support them. A predicate that silently evaluates to False because of an unimplemented backend call is a bug in the predicate, not the benchmark - file an issue.

Contact predicates count a geom pair only when the physics engine reports it as a real touch. get_contacts also lists pairs inside the detection range that carry no force -- see :func:contact_is_active -- and counting those would make contact_any / contact_between / grasped / body_on(require_contact=True) fire for bodies that are visibly apart.

When the backend does support a lookup but the referenced body / joint name cannot be resolved (almost always a spec typo), the term still degrades to a constant (False / 0.0) but the offending name is logged once at WARNING (see :func:_warn_unresolved), so a broken spec surfaces instead of silently preventing episode success or emitting a dead reward.

Available predicates (bool):

body_above_z(body, z)
body_below_z(body, z)
joint_above(joint, value)
joint_below(joint, value)
distance_less_than(body_a, body_b, threshold)
inside_region(body, min, max)
contact_between(geom_a, geom_b)
contact_any()
body_on(body_a, body_b, z_offset=0.02, xy_tol=0.15, require_contact=False)
body_inside(body, container, xy_tol=0.15, z_tol=0.15)
particles_inside(particles, container, min_fraction=1.0, xy_tol=0.15, z_tol=0.15)
particles_spilled(particles, containers, max_spilled=0, xy_tol=0.15, z_tol=0.15)
body_upright(body, tol=0.15)
grasped(body, gripper_prefix)
base_tipped(tol=0.15, robot=None)
base_below_z(z, robot=None)
base_beyond_x(x, robot=None)
base_beyond_y(y, robot=None)
base_yaw_beyond(yaw, robot=None)

Available reward terms (float):

distance_neg(body_a, body_b, weight=1.0)
joint_progress(joint, target, weight=1.0)
particles_inside_fraction(particles, container, xy_tol=0.15, z_tol=0.15, weight=1.0)
base_velocity(vx=0.0, vy=0.0, wz=0.0, weight=1.0, robot=None)
base_velocity_tracking(vx=0.0, vy=0.0, wz=0.0, lin_weight=1.0, ang_weight=0.5, tracking_sigma=0.25, robot=None)
base_height(target, weight=1.0, robot=None)
base_orientation(weight=1.0, robot=None)
base_lin_vel_z(weight=1.0, robot=None)
base_ang_vel_xy(weight=1.0, robot=None)
staged_reward(stages)
constant(value)

Register custom predicates with :func:register_predicate.

make_predicate

make_predicate(name: str, **kwargs: Any) -> Callable[[SimEngine], Any]

Instantiate a predicate from its name + kwargs.

This is the single entry point the DSL loader uses - it never touches eval or exec. Unknown names produce a ValueError listing the valid set; a keyword the factory does not take, or a required one that is missing, produces a ValueError naming the keywords the predicate accepts (read from the factory's signature), so no surface ever shows the factory's own TypeError text.

Every numeric kwarg is held to a finite domain here - a tolerance kwarg additionally to a non-negative one and a heading kwarg to the measurable (-pi, pi) range - and every kwarg naming a scene entity to a non-empty string, rather than in the spec compiler, because this is the only choke point every predicate call passes through: staged_reward builds its per-stage reward / advance_when calls by calling back into this function, so a guard in :func:~strands_robots.simulation.benchmark_spec._compile_call would leave nested stage calls unchecked. See :func:_kwarg_domain_error for what a non-finite threshold or weight does when it compiles.

Parameters:

Name Type Description Default
name str

Predicate name. Must be registered in :data:PREDICATE_REGISTRY.

required
**kwargs Any

Forwarded verbatim to the factory.

{}

Returns:

Type Description
Callable[[SimEngine], Any]

A callable (sim) -> bool or (sim) -> float depending on the

Callable[[SimEngine], Any]

predicate.

Raises:

Type Description
ValueError

If name is unknown, a keyword is one the factory does not take or a required keyword is missing (the message names the accepted keywords), a kwarg the factory annotates as numeric is not a finite number, a kwarg that names a tolerance is negative, a kwarg that names a heading lies outside the (-pi, pi) range a heading can be measured in, or a kwarg that names a scene entity is not a non-empty string.

register_predicate

register_predicate(name: str, factory: PredicateFactory) -> None

Register a user-defined predicate factory.

Must be called before loading a spec that references name. Factories registered at runtime are NOT sandboxed - by registering, you opt into running the factory with kwargs parsed from the spec. Only register predicates from trusted code paths; anything LLM-authored should use the built-in DSL exclusively.

Parameters:

Name Type Description Default
name str

Predicate name used in spec files. Must not shadow a built-in.

required
factory PredicateFactory

Callable that takes DSL kwargs and returns a predicate (sim) -> bool or reward term (sim) -> float.

required

Raises:

Type Description
ValueError

If name shadows a built-in predicate.

TypeError

If factory is not callable.

Benchmark-agnostic evaluation protocol for any SimEngine.

Every standard benchmark (LIBERO, Meta-World, RoboSuite, ManiSkill, user-authored tasks) has a different notion of "what a task is" - sparse-success, dense-reward, procedural scenes, BDDL predicates, hardcoded robots, etc. The correct abstraction is the protocol the eval loop calls into, not a benchmark-specific schema.

:class:BenchmarkProtocol is that protocol. Each adapter implements a handful of lifecycle hooks (on_episode_start, on_step, is_success, is_failure) and declares the robots it is compatible with. The evaluation loop (:meth:~strands_robots.simulation.policy_runner.PolicyRunner.evaluate) drives the protocol without knowing anything about the underlying benchmark.

An adapter that needs a heavyweight simulator declares it in an optional extra, so the core package stays dependency-free. A reference :class:DeclarativeBenchmark shipped in :mod:strands_robots.simulation.benchmark_spec turns a YAML/JSON spec into a fully functional BenchmarkProtocol instance - LLMs can author and register benchmarks at runtime without writing Python code.

Registry: a module-level dict[str, BenchmarkProtocol] keyed by name, mirroring the shape of :func:~strands_robots.simulation.model_registry.register_urdf. Registration is idempotent-by-overwrite: re-registering the same name replaces the previous entry and logs a warning. This matches how users iterate on a spec file during development.

Thread safety: the registry is guarded by an internal lock so concurrent registrations from agent threads do not race. The benchmark instances themselves are expected to be immutable after registration - adapters that keep per-episode state MUST put it on the rng-scoped call, not on self.

register_benchmark

register_benchmark(name: str, benchmark: BenchmarkProtocol) -> None

Register a :class:BenchmarkProtocol under name.

Idempotent-by-overwrite: re-registering the same name replaces the previous entry and logs a warning. This matches how users iterate on a spec file during development.

Parameters:

Name Type Description Default
name str

String key. Must be non-empty; any other validation is up to the caller (lowercase / underscores / hyphens are all fine).

required
benchmark BenchmarkProtocol

An instantiated :class:BenchmarkProtocol subclass.

required

Raises:

Type Description
TypeError

If benchmark is not a :class:BenchmarkProtocol.

ValueError

If name is empty.

unregister_benchmark

unregister_benchmark(name: str) -> BenchmarkProtocol | None

Remove a benchmark from the registry.

Returns the removed benchmark or None if it was not registered. Primarily used by tests for cleanup; user code is rarely expected to unregister benchmarks at runtime.

get_benchmark

get_benchmark(name: str) -> BenchmarkProtocol | None

Return the registered benchmark or None if not found.

list_benchmarks

list_benchmarks() -> dict[str, dict[str, Any]]

Enumerate registered benchmarks with their metadata.

Returns a shallow-copy snapshot keyed by name. Each value is a dict with class, supported_robots, default_robot, max_steps - enough for an LLM to pick an appropriate benchmark without instantiating one. Reads a snapshot under the registry lock so a concurrent registration does not corrupt the returned dict.

Edit page