Skip to content

Label and judge

Predicate verdicts on every episode, a judge agent's grade and failure tag on top, and a filter for the episodes worth training on.

At the end of this page every episode in a dataset carries a verdict the simulator's own predicates decided, a quality grade and failure-mode tag a judge agent added on top, and a filter that picks the episodes worth training on. The judge can annotate a verdict; it can never overturn one.

This runs without lerobot, against the three-episode toy dataset from verify:

import json, pathlib
import pyarrow as pa, pyarrow.parquet as pq
from strands_robots.episode_labels import (
    annotate_episode, filter_episodes, labels_path, record_deterministic_verdicts,
)

root = pathlib.Path("/tmp/toy_dataset")                 # the dataset the verify page builds
(root / "meta" / "episodes" / "chunk-000").mkdir(parents=True, exist_ok=True)
pq.write_table(pa.table({"episode_index": [0, 1, 2], "length": [90, 90, 90]}),
               root / "meta/episodes/chunk-000/file-000.parquet")
(root / "meta/info.json").write_text(json.dumps({"total_episodes": 3, "total_frames": 270, "fps": 30, "features": {}}))
(root / "episode_labels.json").unlink(missing_ok=True)
verdicts = [   # the per-episode list evaluate_benchmark returns
    {"episode": 0, "success": True,  "failure": False, "steps": 90, "cumulative_reward": 1.0, "seed": 0},
    {"episode": 1, "success": True,  "failure": False, "steps": 90, "cumulative_reward": 1.0, "seed": 1},
    {"episode": 2, "success": False, "failure": True,  "steps": 90, "cumulative_reward": 0.0, "seed": 2},
]
record_deterministic_verdicts(root, verdicts, benchmark="so101_reach")         # stage one
annotate_episode(root, 0, quality="high", note="clean approach", model="judge-vlm")
annotate_episode(root, 1, quality="medium", failure_mode="near_miss", success_opinion=False, model="judge-vlm")

print(filter_episodes(root, require_success=True, min_quality="medium"))                          # [0, 1]
print(filter_episodes(root, require_success=True, min_quality="medium", exclude_disputed=True))   # [0]
print(labels_path(root))                                                                          # /tmp/toy_dataset/episode_labels.json

Two stages, one precedence

stage who writes can it change the other block
deterministic the benchmark's predicates, from simulator state (evaluate_benchmark) the deterministic block: success, failure, steps, cumulative_reward, seed no
judge a VLM agent reading the recorded episode the judge block: quality, failure_mode, note, success_opinion, disputes_verdict, model, labeled_at no

A recorded frame is the state before its action, so a stop_when episode ends one step short of the state that fired it. annotate_episode refuses an episode with no deterministic verdict yet. A success_opinion that contradicts the predicate is recorded as disputes_verdict: true for a human to review; the deterministic block stays byte-identical. That precedence is structural, not advice.

The sidecar

Labels live in episode_labels.json at the dataset root, next to LeRobot's meta/, data/ and videos/, so training can filter episodes without rewriting the dataset. schema_version is 1; read_labels(root) returns the document, deterministic_verdict(root, episode) one verdict.

Vocabulary is fixed so filters match on identity:

field values
quality low, medium, high
failure_mode jerky_motion, near_miss, camera_occlusion, wrong_but_lucky, drift, collision, incomplete, other

near_miss and wrong_but_lucky are legal on a deterministically successful episode on purpose: they mark a success worth excluding from training data.

The judge agent

Four @tools in strands_robots.tools.episode_judge drive a judge; create_judge_agent(model=None) assembles them with a system prompt carrying the two-stage doctrine. Pass any strands model object (a Bedrock VLM, an OpenAI-compatible endpoint) or none for the default:

from strands_robots.tools.episode_judge import create_judge_agent

judge = create_judge_agent()
judge(f"Label every episode of the dataset at {root}. Sample four frames each, with images.")
tool returns
load_episode(root, episode) frame count, features, whether a verdict and a label exist yet
sample_frames(root, episode, n_frames=4, include_images=False) evenly spaced frames: observation.state and timestamps always, one decoded image per camera (all cameras, sorted key order) when asked, plus rms_state_jerk over the episode so a text-only judge can ground jerky_motion
read_predicate_verdict(root, episode) the authoritative deterministic verdict
write_label(root, episode, quality, failure_mode=None, note="", success_opinion=None, judge_model="") the judge block, through annotate_episode

Every tool returns the {"status", "content"} envelope and never raises. sample_frames with images needs the [lerobot] extra to decode video.

Filtering for training

filter_episodes(root, require_success=True, min_quality="medium", exclude_disputed=False) returns the episode indices that clear the bar. Unlabeled episodes are excluded: an episode with no judge block has no quality to compare, and admitting it would make the answer depend on how far the judge got. Pass the list to lerobot's --dataset.episodes or to StreamingDatasetReader.open(episodes=...) (stream and sync).

Checking the judge

measure_agreement(root, human_labels) compares the judge's labels with a human-labeled holdout and reports agreement per field against a constant-answer baseline, so a judge that always says medium cannot look accurate.

Edit page