Skip to the content.

Chapter 15 — From Demonstrations to a Dataset

Part VI closes the loop: your demonstrations become a fine-tuned policy that drives the robot back. This chapter is the first half — turning a folder of recorded episodes into a dataset a model can train on. It assumes you know the episode format, what fine-tuning is, and the idea of an embodiment.


A folder of episodes is not yet a dataset

You have recorded episodes — maybe eighty of them, each a folder of frames and a steps.jsonl (Chapter 13). That is data, but it is not yet a dataset in the shape a model’s training code expects. Between the two sits a conversion, and the conversion is more interesting than it sounds, because it is where you tell the model three things it must know before it can learn from your episodes:

  1. what each number means (this vector is joint positions, in this order),
  2. how big the numbers typically are (so it can normalize them), and
  3. which body it is looking at (a robot it already knows, or a brand-new one).

Get those three right and training just works. Get any of them wrong and the model trains on nonsense that looks like data. So the conversion is not mere reformatting — it is another contract with the model, exactly like the wire protocol was a contract for live control.


The conversion

The target format is LeRobot (specifically the v2.1 flavor GR00T expects) — a standard, efficient layout for robot demonstration datasets. A converter walks your episode folders and produces:

data/chunk-000/episode_000000.parquet          — the per-step numbers (state, action, timestamp)
videos/chunk-000/observation.images.ego_view/
  episode_000000.mp4                            — the camera frames, encoded as H.264 video
meta/info.json                                  — dataset shape and version
meta/stats.json                                 — normalization statistics (below)
meta/episodes.jsonl, meta/tasks.jsonl           — episode and task bookkeeping
meta/modality.json                              — what each part of the data *is* (below)

The frames become a compressed MP4 per episode rather than thousands of loose JPEGs (much smaller, and what the training pipeline reads), the per-step numbers become an efficient Parquet table, and the meta/ folder carries the metadata that makes the numbers meaningful. One small, real detail: the Unity recorder writes its text files with a byte-order mark, so the converter reads them as utf-8-sig — the kind of unglamorous compatibility fix that separates a pipeline that works from one that almost works.


modality.json: telling the model what its numbers mean

The recorder stored state and action as bare lists of 29 numbers. Bare lists are meaningless to a model that could be trained on many robots — is number 12 a joint angle, a gripper width, a velocity? modality.json is the answer key:

{
  "state":      { "joint_position": { "start": 0, "end": 29 } },
  "action":     { "joint_position": { "start": 0, "end": 29 } },
  "video":      { "ego_view": { "original_key": "observation.images.ego_view" } },
  "annotation": { "human.task_description": { "original_key": "task_index" } }
}

It says: the 29-number state is joint positions; the 29-number action is joint positions (absolute targets, matching how Unity applies them); the video is the ego-view camera; the language annotation is the task description. This is the same lesson as the embodiment profile in Chapter 10 — a model needs its inputs named and shaped, not just handed over — now applied to the training data instead of the live observation. The wire protocol and the dataset format are two faces of one requirement: the model must know what each number is.


Normalization: the statistics the model needs

Neural networks learn best when their inputs are standardized — centered near zero and scaled to a consistent range — rather than raw radians that might run from −2.5 to +2.5 on one joint and −0.1 to +0.1 on another. So the converter computes, per dimension, the normalization statistics and stores them in meta/stats.json:

These are computed over your actual episodes — the concatenated actions, states, and timestamps — so they describe your data specifically. At training and inference time, the model uses them to normalize inputs and de-normalize its outputs back into real joint angles. This is why fine-tuning on your own demonstrations needs your own statistics: a model normalized to someone else’s data would systematically mis-scale yours.

Insight: the dataset format is a contract, exactly like the wire protocol. Chapter 5’s lesson — the model must be handed inputs it recognizes — returns here in a different costume. Live, the contract is the observation shape; offline, it is modality.json plus the normalization stats. Both encode the same fact: a model’s numbers are only meaningful with their names and scales attached. Nearly every “the fine-tune trained but the robot does nonsense” failure is a broken version of this contract — a mislabeled modality, wrong-scale stats, an off-by-one action column.


How much does the model already know your robot?

The third thing the conversion declares is which body the demonstrations are on — and this determines how much work fine-tuning has to do. GR00T sorts robots into three tiers, easiest to hardest:

Tier What it means Effort
Pretrained (e.g. REAL_G1) The base checkpoint already drives this robot zero-shot. Fine-tuning gets a warm start — it is nudging a model that already half-knows the body. Fewest demonstrations, best few-shot results
Fine-tune-only (e.g. Fourier GR-1) The model ships the modality config for this robot but no checkpoint actually drives it. You must fine-tune, but you do not have to invent the body description. No zero-shot, but the config is given
New embodiment (e.g. Unitree H1, and the 29-DoF push G1) The model has no idea this robot exists. You author the modality config from scratch, and the projection layer that reads this body’s numbers trains from zero. No warm start. Most data-hungry

Insight: the tier decides your data budget before you record a single demo. A warm-start (pretrained) robot might reach a good policy on a few dozen demonstrations, because the model is adapting, not learning from scratch. A new-embodiment robot must learn the very meaning of its 29 numbers, and needs more. Knowing your robot’s tier up front tells you how many demonstrations to plan for and whether zero-shot is even an option. For hand-equipped manipulation, this is a concrete reason to prefer the Dex3-equipped G1 — it rides the pretrained REAL_G1 tier and its hands are full-fidelity (Chapter 7) — over robots that demand a cold-start fine-tune.

The push campaign in the next chapter is a new-embodiment case: a 29-DoF G1 configured as a body the base model had never seen, with a projection layer trained from scratch.


Two ecosystems

One more thing the conversion decides: which training ecosystem you are feeding. The two humanoid VLAs and the arm VLA do not share a dataset format:

You do not need the details of RLDS to take the point: the record step produced one universal format (Chapter 13), and the convert step fans that one format out to whichever ecosystem the target model uses. The universal recording format is what makes it a fan-out rather than a fork all the way back to Unity — every model’s converter starts from the same episode folders.


What you now understand

The dataset is ready. The next chapter runs the actual fine-tune — eighty push demonstrations, one overnight training run on the owned GPU box, and the honest number for how well it worked.

Continue to Chapter 16 — The First Fine-Tune.


The GR00T converter is tools/convert_g1_push_to_lerobot.py (episodes → LeRobot v2.1, modality.json, stats); the new-embodiment modality config is scripts/spark/g1_push_config.py. The OpenVLA/RLDS path uses a separate converter (server/tools/convert_to_lerobot.py for the 7-DoF LeRobot variant; OpenVLA’s own fine-tune consumes RLDS). Embodiment tiers are described in the repository’s PROGRESS.md.