kimodo¶
kimodo samples whole-body Unitree G1 motion from an English prompt with NVIDIA's text-to-motion diffusion model.
Deprecated, removed in 0.7
No in-tree replacement: generate the motion offline and replay it as joint targets.
This page turns a sentence such as "a person walking forward with confident strides" into a 29-joint kinematic reference for the Unitree G1 and plays it in simulation.
pip install 'strands-robots[kimodo]' # torch, diffusers, transformers, huggingface_hub, accelerate, scipy
export STRANDS_TRUST_REMOTE_CODE=1 # the diffusers pipeline loads with trust_remote_code=True
What it is¶
KimodoPolicy wraps NVIDIA's Kimodo (nvidia/Kimodo-G1-RP-v1), a text-conditioned motion diffusion model. One diffusion pass samples a per-frame full-body qpos sequence for the G1; the policy resamples it from the model's native frame rate to the control rate and emits one action dict per tick, keyed by the canonical WBC joint ordering for all 29 leg, waist and arm joints. Any English motion description is a goal; there is no fixed clip vocabulary. requires_images is False.
Standalone in MuJoCo the targets are applied directly: the faithful kinematic reference. Closing the loop through physics needs a controller tracking that reference over the same 29 joints, a cascade CompositePolicy (which merges disjoint joint groups) cannot express. protomotions is the tracker; wbc cannot track a pose, its only input is a base velocity.
from strands_robots.policies import create_policy
policy = create_policy("kimodo", diffusion_steps=100, guidance_scale=7.5, num_frames=120)
policy = create_policy("text2motion") # same provider, all defaults
Constructor keywords¶
| keyword | type | default |
|---|---|---|
config |
KimodoConfig \| dict[str, Any] \| None |
None |
motion_agent |
KimodoMotionAgent \| None |
None |
model_id |
str \| None |
None |
diffusion_steps |
int \| None |
None |
guidance_scale |
float \| None |
None |
num_frames |
int \| None |
None |
transition_frames |
int \| None |
None |
native_fps |
int \| None |
None |
tracker_fps |
int \| None |
None |
device |
str \| None |
None |
dtype |
str \| None |
None |
cache_dir |
str \| None |
None |
trust_remote_code |
bool \| None |
None |
seed |
int \| None |
None |
no **kwargs: an unknown keyword is a TypeError |
Everything after config is keyword-only and mirrors a KimodoConfig field; an explicit keyword wins over the config object. No **kwargs: a misspelled keyword is a TypeError, while a misspelled key inside config={...} is dropped silently by KimodoConfig.from_dict. Defaults: model_id="nvidia/Kimodo-G1-RP-v1", diffusion_steps=100, guidance_scale=7.5, num_frames=120, dtype="fp16" (bf16 and fp32 accepted). motion_agent injects a sampler; it is how NVIDIA's bare-weights checkpoint, built for its own runtime rather than a diffusers pipeline, is driven.
Per-call keywords¶
| keyword | meaning |
|---|---|
text_prompt |
overrides instruction as the prompt |
diffusion_steps, guidance_scale |
override the config for this sample |
seed |
seeds the sampler for this episode; reset(seed) does the same between episodes |
Run it¶
Needs the extra, the environment gate, a CUDA GPU and a first-use download.
from strands_robots.simulation import create_simulation
sim = create_simulation("mujoco")
sim.create_world()
sim.add_robot("g1")
result = sim.run_policy(
robot_name="g1",
policy_provider="kimodo",
policy_config={"diffusion_steps": 100, "guidance_scale": 7.5, "num_frames": 120},
instruction="a person walking forward with confident strides",
n_steps=200,
control_frequency=50.0,
)
print(result["status"])
On hardware, strands_robots.policies.kimodo.hardware.kimodo_action_to_lerobot_g1 renames the 29 canonical joints onto the LeRobot G1 driver's motor names.
Driving the NVIDIA checkpoint¶
The bare weights load through NVIDIA's own kimodo runtime, shipped with the model rather than on PyPI. Wrap it as the motion_agent and hand the policy to run_policy as a built object; seed must reach torch.manual_seed, because the runtime draws from the global generator and an adapter that ignores it defeats the per-episode seed:
import numpy as np
import torch
from strands_robots.policies.kimodo import KimodoPolicy
class NativeKimodoAgent:
"""Samples through NVIDIA's kimodo runtime instead of diffusers."""
def __init__(self, device: str = "cuda") -> None:
from kimodo.exports.mujoco import MujocoQposConverter
from kimodo.model.load_model import load_model
self._model = load_model("kimodo-g1-rp", device=device)
self._converter = MujocoQposConverter(self._model.skeleton)
self._device = device
def sample(self, prompt, num_frames, diffusion_steps, guidance_scale, seed):
if seed is not None:
torch.manual_seed(seed)
output = self._model([prompt.strip().rstrip(".") + "."], [num_frames], num_denoising_steps=diffusion_steps, num_samples=1, return_numpy=True)
qpos = np.asarray(self._converter.dict_to_qpos(output, self._device))
return qpos[0].astype(np.float32) if qpos.ndim == 3 else qpos.astype(np.float32)
sim.run_policy(robot_name="g1", policy_object=KimodoPolicy(motion_agent=NativeKimodoAgent()), instruction="a person walking forward", n_steps=200, control_frequency=50)
MujocoQposConverter turns the runtime's rotation matrices into the (num_frames, 7 + 29) array; guidance_scale has no counterpart there and is ignored.
Limits¶
- Unitree G1 only; the output is the G1's 29-joint layout.
- Kinematic, not dynamic. Played directly, the reference ignores contact and balance. Track it with protomotions for physics.
- A sample is one diffusion pass over
num_framesat the native rate, so the first call is long;transition_framesblends consecutive samples. trust_remote_code=Trueis how thediffuserspipeline loads, hence the environment gate.