lerobot_local¶
lerobot_local runs any LeRobot checkpoint in process. Install extra, constructor keywords, embodiments, camera naming rules, and limits.
This page runs a HuggingFace LeRobot checkpoint (ACT, diffusion, pi0, SmolVLA, GR00T N1.7, MolmoAct2, each behind an extra below) on a simulated or real arm and teaches the two naming rules deciding whether the model sees your cameras and joints.
pip install 'strands-robots[lerobot]' # lerobot[feetech,dataset] + psutil
pip install 'strands-robots[smolvla]' # adds lerobot[smolvla]
pip install 'strands-robots[molmoact2]' # adds lerobot[molmoact2]
pip install 'strands-robots[groot]' # adds lerobot[groot] (GR00T N1.7)
pip install 'lerobot[diffusion]' # diffusion: lerobot's own extra, not ours; pi0: lerobot[pi]
export STRANDS_TRUST_REMOTE_CODE=1 # required: models load with trust_remote_code=True
What it is¶
LerobotLocalPolicy hands the checkpoint to LeRobot's own factory: the policy type is read from the model's config.json, so any class LeRobot registers works unchanged. The processor pipeline (preprocessor.json / postprocessor.json) normalises observations, unnormalises actions. Flow-matching models get Real-Time Chunking when the config declares it: the runtime reports its control rate and the steps consumed during inference; the policy blends the next chunk onto the seam.
Build it by name or smart string:
from strands_robots.policies import create_policy
policy = create_policy("lerobot_local", pretrained_name_or_path="robotfuel/act_so101_t16b", embodiment="so101")
policy = create_policy("robotfuel/act_so101_t16b", embodiment="so101") # same thing
Constructor keywords¶
| keyword | type | default |
|---|---|---|
pretrained_name_or_path |
str |
'' |
policy_type |
str \| None |
None |
device |
str \| None |
None |
actions_per_step |
int |
1 |
use_processor |
bool |
True |
processor_overrides |
dict \| None |
None |
tokenizer_max_length |
int |
48 |
tokenizer_padding_side |
str |
'right' |
rtc_enabled |
bool \| None |
None |
rtc_execution_horizon |
int \| None |
None |
rtc_max_guidance_weight |
float \| None |
None |
inference_kwargs |
dict \| None |
None |
embodiment |
str \| dict \| Any \| None |
None |
norm_tag |
str \| None |
None |
image_keys |
list[str] \| None |
None |
inference_action_mode |
str |
'continuous' |
camera_key_map |
dict[str, str] \| None |
None |
obs_rename_override |
dict[str, str \| None] \| None |
None |
strict_keys |
bool |
False |
pad_short_actions |
bool |
False |
out_of_range_actions |
str |
'warn' |
cache_model |
bool |
True |
revision |
str \| None |
None |
compile_model |
bool \| None |
None |
**ignored_kwargs |
unknown keywords are ignored |
Inert normalization has a two-part remedy: processor_overrides={"normalizer_processor": {"stats": ...}} replaces the stats, and the embodiment's state_units / action_units (degrees or native) say which frame they were recorded in; native is what the robot emits, radians in MuJoCo. so100 and so101 declare degrees. Neither is a constructor keyword; create_policy refuses both.
pretrained_name_or_path is required. actions_per_step=1 becomes the trained n_action_steps; above 1 pins it. cache_model=True shares weights in-process (clear_model_cache(), list_cached_models()). Without device= it uses CUDA if present; torch.compile stays off unless compile_model=True.
Embodiments¶
An embodiment is a declared key map from what the robot emits to what the model was trained on: state_keys, action_keys, obs_rename, and a dim_policy (strict, pad, or truncate) for a state width unlike the robot's. They live in strands_robots/policies/lerobot_local/embodiments.json; sim entries use MuJoCo joint names, *_real entries LeRobot motor names with .pos. Known embodiments and aliases:
panda_libero, so101, so100, so_real, koch, koch_real, lekiwi_real, lekiwi_sim, panda, panda_droid, fr3, ur5e, kinova_gen3, xarm7, vx300s, wx250s, yam, piper, aloha, unitree_g1, unitree_g1_arms, unitree_g1_sonic, unitree_g1_real, unitree_g1_real_arms, unitree_h1, unitree_h1_2, omx_real, bi_so_real, openarm_real, bi_openarm_real, rebot_b601_real, bi_rebot_b601_real, reachy2_real, hope_jr_arm_real, hope_jr_hand_real, earthrover_real, franka_libero, so100_real, so101_real, franka, franka_fr3, gen3, kinova, viperx_300s, widowx_250s, lekiwi, g1, g1_arms, h1, h1_2, omx_follower, bi_so_follower, bi_so, openarm, openarm_follower, bi_openarm, bi_openarm_follower, rebot_b601, rebot_b601_follower, bi_rebot_b601, bi_rebot_b601_follower, reachy2, hope_jr_arm, hope_jr_hand, earthrover, earthrover_mini_plus, so100_follower, so101_follower, koch_follower, lekiwi_client, franka_droid, g1_sonic
Rule 1: state keys¶
Without set_robot_state_keys, the policy infers the state vector from the observation's insertion order of numeric scalars. The sim backends write obs[joint] then obs[f"{joint}.vel"], so strands_robots.policies._state_keys.drop_velocity_siblings removes each .vel whose position companion is present, keeping one that has none (LeKiwi declares x.vel, y.vel, theta.vel as state). Every provider that infers an ordering shares this rule; an explicit robot_state_keys list is kept.
Rule 2: camera names¶
A checkpoint declares image features such as observation.images.image. The embodiment's obs_rename maps the camera key you attach onto that feature. Name a sim camera after the model card (realsense_top) rather than the embodiment's source key (front) and the rename never fires; preflight refuses before any download, naming the expected source keys.
From embodiments.json:
| embodiment | camera key you attach | model image feature | aliases |
|---|---|---|---|
panda_libero |
imagewrist_image |
observation.images.imageobservation.images.wrist_image |
franka_libero |
so101 |
frontwrist |
observation.images.imageobservation.images.wrist_image |
|
so100 |
frontwrist |
observation.images.imageobservation.images.wrist_image |
|
so_real |
frontwrist |
observation.images.imageobservation.images.wrist_image |
so100_real, so101_real, so100_follower, so101_follower |
koch |
frontwrist |
observation.images.imageobservation.images.wrist_image |
|
koch_real |
frontwrist |
observation.images.imageobservation.images.wrist_image |
koch_follower |
lekiwi_real |
frontwrist |
observation.images.imageobservation.images.wrist_image |
lekiwi, lekiwi_client |
lekiwi_sim |
frontwrist |
observation.images.imageobservation.images.wrist_image |
|
panda |
imagewrist_image |
observation.images.imageobservation.images.wrist_image |
franka |
panda_droid |
exteriorwrist |
observation.images.base_0_rgbobservation.images.left_wrist_0_rgb |
franka_droid |
fr3 |
imagewrist_image |
observation.images.imageobservation.images.wrist_image |
franka_fr3 |
ur5e |
imagewrist_image |
observation.images.imageobservation.images.wrist_image |
|
kinova_gen3 |
imagewrist_image |
observation.images.imageobservation.images.wrist_image |
gen3, kinova |
xarm7 |
imagewrist_image |
observation.images.imageobservation.images.wrist_image |
|
vx300s |
imagewrist_image |
observation.images.imageobservation.images.wrist_image |
viperx_300s |
wx250s |
imagewrist_image |
observation.images.imageobservation.images.wrist_image |
widowx_250s |
yam |
imagewrist_image |
observation.images.imageobservation.images.wrist_image |
|
piper |
imagewrist_image |
observation.images.imageobservation.images.wrist_image |
|
aloha |
top |
observation.images.top |
|
unitree_g1 |
ego_view |
observation.images.ego_view |
g1 |
unitree_g1_arms |
ego_view |
observation.images.ego_view |
g1_arms |
unitree_g1_sonic |
ego_viewleft_wristright_wrist |
observation.images.ego_viewobservation.images.left_wristobservation.images.right_wrist |
g1_sonic |
unitree_g1_real |
ego_view |
observation.images.ego_view |
|
unitree_g1_real_arms |
ego_view |
observation.images.ego_view |
|
unitree_h1 |
ego_view |
observation.images.ego_view |
h1 |
unitree_h1_2 |
ego_view |
observation.images.ego_view |
h1_2 |
omx_real |
frontwrist |
observation.images.imageobservation.images.wrist_image |
omx_follower |
bi_so_real |
frontwrist |
observation.images.imageobservation.images.wrist_image |
bi_so_follower, bi_so |
openarm_real |
frontwrist |
observation.images.imageobservation.images.wrist_image |
openarm, openarm_follower |
bi_openarm_real |
frontwrist |
observation.images.imageobservation.images.wrist_image |
bi_openarm, bi_openarm_follower |
rebot_b601_real |
frontwrist |
observation.images.imageobservation.images.wrist_image |
rebot_b601, rebot_b601_follower |
bi_rebot_b601_real |
frontwrist |
observation.images.imageobservation.images.wrist_image |
bi_rebot_b601, bi_rebot_b601_follower |
reachy2_real |
head |
observation.images.image |
reachy2 |
hope_jr_arm_real |
frontwrist |
observation.images.imageobservation.images.wrist_image |
hope_jr_arm |
hope_jr_hand_real |
hand |
observation.images.image |
hope_jr_hand |
earthrover_real |
frontrear |
observation.images.imageobservation.images.image2 |
earthrover, earthrover_mini_plus |
Two remedies:
# 1. Name the cameras as the embodiment expects.
sim.add_camera(name="front", position=[0.22, 0.025, 0.6], target=[0.22, 0.025, 0])
sim.add_camera(name="wrist", parent_body="so101/gripper", position=[0.058, 0.0, -0.029], target=[-0.024, 0.0, -0.297])
# 2. Keep your names and route them (camera_key_map, then obs_rename_override, merge over obs_rename).
sim.run_policy(
robot_name="so101",
policy_provider="lerobot_local",
policy_config={
"pretrained_name_or_path": "allenai/MolmoAct2-SO100_101",
"embodiment": "so101",
"obs_rename_override": {"realsense_top": "observation.images.image", "realsense_side": "observation.images.wrist_image"},
},
instruction="pick up the cube",
)
parent_body mounts a camera on a link (a wrist view rides with the arm); position and target are then in that frame, both required. It works on every backend.
Run it¶
Needs the extra and an 865 MB download. smolvla_base declares camera1..3 and ships no SO-101 stats, so embodiment="so101" (degrees) is refused; an inline native-unit embodiment runs:
import os
os.environ["STRANDS_TRUST_REMOTE_CODE"] = "1"
from strands_robots.simulation import create_simulation
sim = create_simulation("mujoco", mesh=False)
sim.create_world()
sim.add_robot("so101")
sim.add_camera(name="front", position=[0.22, 0.025, 0.6], target=[0.22, 0.025, 0])
sim.add_camera(name="wrist", parent_body="so101/gripper", position=[0.058, 0.0, -0.029], target=[-0.024, 0.0, -0.297])
joints = sim.robot_joint_names("so101")
embodiment = {"name": "so101_native", "state_keys": joints, "action_keys": joints, "dim_policy": "pad",
"obs_rename": {"front": "observation.images.camera1", "wrist": "observation.images.camera2",
"default": "observation.images.camera3"}}
result = sim.run_policy(robot_name="so101", policy_provider="lerobot_local",
policy_config={"pretrained_name_or_path": "lerobot/smolvla_base", "embodiment": embodiment},
instruction="pick up the cube", n_steps=30, control_frequency=30.0)
print(result["status"])
sim.cleanup()
An SO-101 fine-tune carries degree stats. On the sim joints 1..6 the so101 embodiment applies even unnamed, converting both ways; other radian state is refused before the first action. On hardware it binds the .pos keys. robotfuel/act_so101_t16b applied 90 of 90 steps at 30 Hz through run_policy(policy_object=...) on a laptop CPU; arm moved, no cube lifted. A real arm's tool takes the same policy_config dict:
{"action": "execute", "policy_provider": "lerobot_local",
"policy_config": {"pretrained_name_or_path": "robotfuel/act_so101_t16b", "embodiment": "so101",
"obs_rename_override": {"front": null, "wrist": "observation.images.wrist"}}}
GR00T N1.7 through lerobot¶
nvidia/GR00T-N1.7-3B and its fine-tunes are lerobot's native groot type and load like any checkpoint, without an Isaac-GR00T checkout or ZMQ service. embodiment_tag comes from the checkpoint config; install the groot extra.
from strands_robots.policies import create_policy
policy = create_policy("nvidia/GR00T-N1.7-3B", policy_type="groot", embodiment="so101")
print(policy.provider_name)
The 3B model wants a GPU: run PolicyServer there, dialled with remote.
cfg = {"pretrained_name_or_path": "nvidia/GR00T-N1.7-3B", "policy_type": "groot", "embodiment": "so101"}
PolicyServer(policy_provider="lerobot_local", policy_config=cfg, port=8765).start() # GPU host
policy = create_policy("ws://gpu-box:8765") # robot host
Limits¶
trust_remote_code=Trueis unconditional here, hence the environment gate; load only checkpoints from organisations you trust.dim_policy="pad"/"truncate"adapt the state width and take the first N values of a wider action (32-D pi0/pi0.5);strictrefuses. One the pipeline cannot take is refused at load.- An embodiment not in
embodiments.jsonneeds its own entry (state keys, action keys, camera renames); training shows how a checkpoint carries them.