Skip to content

How it works

How Strands Robots works in three minutes: one Robot object that is an agent tool, one Policy object that never sees a robot, and the backends under them.

Strands Robots is one interface to any robot, simulated or physical, that never moves a real arm until a person says yes. Three objects carry the whole design.

Left column, top to bottom: a Strands Agent with the robot in its tools makes a tool call; it passes the operator gate, the one green element, where run_policy and send_action through the tool wait for a yes (an interrupt, fail closed without an operator, an audit row); after a yes it reaches Robot("so101"), the execution target and the tool, with run_policy, send_action, get_observation, get_state, render and status; a dashed result wire returns to the agent. Under the robot a dashed policy runtime layer holds create_policy, the embodiment map and the action chunk, fed by the observation and returning actions. Right column, the backends the same calls reach: simulation with MuJoCo, Newton or Isaac Sim; a hardware driver, lerobot or native; and a PolicyServer on a GPU host that a RemotePolicy talks to over a WebSocket. Footnote: one interface, get_observation, send_action, run_policy; the backend changes, the call does not.Left column, top to bottom: a Strands Agent with the robot in its tools makes a tool call; it passes the operator gate, the one green element, where run_policy and send_action through the tool wait for a yes (an interrupt, fail closed without an operator, an audit row); after a yes it reaches Robot("so101"), the execution target and the tool, with run_policy, send_action, get_observation, get_state, render and status; a dashed result wire returns to the agent. Under the robot a dashed policy runtime layer holds create_policy, the embodiment map and the action chunk, fed by the observation and returning actions. Right column, the backends the same calls reach: simulation with MuJoCo, Newton or Isaac Sim; a hardware driver, lerobot or native; and a PolicyServer on a GPU host that a RemotePolicy talks to over a WebSocket. Footnote: one interface, get_observation, send_action, run_policy; the backend changes, the call does not.

Robot: the execution target

Robot("so101") is a factory, not a class. In mode="sim", the default, it returns a simulation engine with the world built and the robot in it; in mode="real" it returns the lerobot driver or a native driver for the same name. Both objects are Strands agent tools: they carry a tool_name, a tool_spec the model reads and one action enum, so Agent(tools=[robot]) works with either. Every method returns the same envelope, status plus a content list of text, JSON and image blocks, which is what the model sees when it calls the tool and what you print when you call the method. Robots goes deeper.

Policy: behaviour as a runtime component

A Policy turns an observation and an instruction into a chunk of actions. create_policy(provider, **config) builds one from a provider name (16 providers, from mock through lerobot_local to remote) and a configuration, usually a Hub checkpoint. The policy never sees a robot: it sees observation.state in the units it was trained on and returns action in the same units. The robot object owns the control loop, injects the state keys and the control frequency, and applies each action with send_action. Policies chooses a provider; Embodiments explains the map between the policy's numbers and the robot's joints.

Backend: where the robot is

Under the same get_observation, send_action and run_policy sits MuJoCo on a CPU, Newton or Isaac on a GPU, the lerobot driver on a USB port, or a native driver speaking a serial bus, DDS or a vendor API. The call does not change; the backend does. Simulation and hardware lists them and what each refuses.

The gate

Anything that moves a physical robot passes gate_motion first: an allowlist variable, then BYPASS_TOOL_CONSENT, then the operator through a Strands interrupt, and with nobody to ask the call fails closed. Every decision is written to the audit log. Agents and robots draws the chain and names the one native command that skips it today.

Around the spine

A fleet puts many such robots on one mesh with one e-stop; the data flywheel records what a robot did, trains a checkpoint from it and runs the checkpoint back on the robot; remote inference moves the policy to a GPU host while the robot host keeps the control loop and the gate. The glossary defines the words these pages use.

Edit page