LeRobot training¶
Post-tune a LeRobot policy from a recorded dataset: the lerobot_train tool, the train_policy tool, LerobotTrainer through TrainSpec, and the provider knobs for lerobot (GR00T N1.7 included) and Cosmos 3.
By the end of this page you can turn a recorded LeRobotDataset into a fine-tuned checkpoint that create_policy loads back, from Python or from an agent, and you know which knobs are strands' and which belong to lerobot.
pip install 'strands-robots[lerobot]' # lerobot[feetech,dataset]
pip install 'lerobot[training]' # accelerate for train(); add peft for method="lora"
Python: TrainSpec and LerobotTrainer¶
LerobotTrainer builds a typed lerobot.configs.train.TrainPipelineConfig and calls lerobot's train(cfg) in this interpreter. The training logic is lerobot's; the adapter translates the spec, manages resume, and parses the run for a verdict. Needs a GPU for anything past a smoke test.
from strands_robots.training import TrainSpec, create_trainer
spec = TrainSpec(
dataset_root="~/.cache/huggingface/lerobot/me/so101_pick", # has meta/info.json; what stop_recording writes
base_model="lerobot/smolvla_base",
output_dir="/tmp/so101_smolvla",
steps=20_000,
global_batch_size=32,
save_freq=2_000,
val_episodes=2, # hold out the last 2 episodes, log eval loss
method="lora", lora_r=16,
extra={"policy_type": "smolvla", "policy.freeze_vision_encoder": False},
)
trainer = create_trainer("lerobot_local")
assert trainer.validate(spec) == []
result = trainer.train(spec)
print(result.status, result.checkpoint_dir, result.metrics)
Then, with STRANDS_TRUST_REMOTE_CODE=1 exported (the checkpoint's config is code the loader will not run unasked), create_policy("lerobot_local", pretrained_name_or_path=result.checkpoint_dir, embodiment="so101") runs it; see lerobot-local for the naming rules the checkpoint now carries.
What strands adds on top of lerobot¶
lerobot owns the training knobs (what each policy freezes, sample weighting, relative actions, quantile normalisation, streaming) and documents them at huggingface.co/docs/lerobot. Four things are strands':
extrareaches any field of lerobot's config tree by dotted key:policy.*,dataset.*,wandb.*. Values may be typed or text (decoded by lerobot's draccus, so"false"is a boolean);Noneclears an optional field;extra["policy_type"]picks the architecture (act,diffusion,smolvla,pi0,pi05,pi0_fast,groot,xvla, ...).methodis strands' tuning strategy, not lerobot's per-policy defaults:full,lora(peft, needslora_r),expert_only(freezes the VLM; only policies whose lerobot config exposestrain_expert_only, read live).loraandexpert_onlyare mutually exclusive.validate()refuses before launch what lerobot would fail on inside the run: exactly one data source (dataset_rootordataset_repo_id);val_episodeson a streamed, multi-task or count-less dataset (lerobot splits by fraction per task); acodebase_versionthe installed lerobot cannot read (it names the converter); quantile-normalised policy types (pi05,molmoact2) on a dataset withoutq01..q99stats;extra["relative_actions"]on a policy other thangroot,pi0,pi05,pi0_fast;num_nodes > 1; missingaccelerateorpeft.- A fresh start clears an empty leftover
output_dirand nothing else;resume=Truecontinues from<output_dir>/checkpoints/last.
extra["reward_model"] switches the run to a reward model (sarm) instead of a policy; extra["sample_weighting"] configures RA-BC weighting. Both take friendly field dicts checked against lerobot's live dataclasses.
From an agent: two tools¶
train_policy is provider-agnostic and in process. Its arguments are TrainSpec field for field plus action and job_id:
train_policy(action="validate", provider="lerobot_local", dataset_root="...", base_model="lerobot/smolvla_base", output_dir="/tmp/ft")
train_policy(action="train", provider="lerobot_local", dataset_root="...", base_model="lerobot/smolvla_base", output_dir="/tmp/ft", steps=20000, method="lora", lora_r=16)
train_policy(action="status", provider="lerobot_local", job_id="...") # RUNNING != learning
train_policy(action="export", provider="lerobot_local", output_dir="/tmp/ft")
train_policy(action="list")
lerobot_train is the detached alternative: it launches python -m lerobot.scripts.lerobot_train (or accelerate launch for num_gpus > 1) as a background process and tracks it in the same on-disk session store as lerobot_teleoperate, so the run outlives the agent turn.
| argument | meaning |
|---|---|
action |
start (default), status (pid, uptime, log tail), stop (SIGTERM then SIGKILL), list |
dataset_root, policy_type="act", pretrained_path, output_dir, job_name |
what to train on and where |
steps=20000, batch_size=8, save_freq=5000, device="cuda", dtype, num_gpus=1 |
run size and placement |
lora, lora_r, lora_alpha, lora_target_modules, train_expert_only, gradient_checkpointing |
memory-fit levers; lora and train_expert_only are mutually exclusive |
val_episodes |
holds out the last N episodes via --dataset.eval_split plus --eval_steps; refused on multi-task datasets |
resume, push_to_hub, session_name, extra_flags |
resume only when checkpoints/last exists; raw --key=value passthrough |
The boolean levers are checked, not read by truthiness: "false" is refused rather than turning a lever on.
GR00T N1.7 and Cosmos 3 knobs¶
GR00T N1.7 trains through lerobot_local with extra={"policy_type": "groot"} and base_model="nvidia/GR00T-N1.7-3B". embodiment is required there (it becomes GrootConfig.embodiment_tag) and refused for a policy type with no such field; tune toggles the components GrootConfig exposes (llm, visual, projector, diffusion), and method="expert_only" is refused because the config has no single switch for it.
create_trainer("cosmos3") needs prepare(spec) first: it converts the HF checkpoint to DCP with convert_model_to_dcp. train loads a TOML recipe and applies extra as Hydra key.path=value overrides; extra["cosmos_root"] names the checkout and extra["sft_toml"] the recipe TOML; validate() refuses a spec without either by name. The floor is eight 80 GB GPUs.
create_trainer("sagemaker", image_uri=..., role_arn=..., instance_type="ml.g5.xlarge") ships the same TrainSpec to a managed job and waits; the image packages one of the trainers above.
Limits¶
LerobotTrainerandtrain_policyare single-node;num_nodes > 1is refused. Usesagemakerorlerobot_trainwithacceleratefor scale-out.- The trainer imports lerobot in this process, which pins
transformers>=5; the[groot]extra adds the peft, diffusers and timm pins GR00T N1.7 needs on top. - No training runs on this documentation machine; every fence above is a sketch by design.