Chapter 15 — From Demonstrations to a Dataset
Part VI closes the loop: your demonstrations become a fine-tuned policy that drives the robot back. This chapter is the first half — turning a folder of recorded episodes into a dataset a model can train on. It assumes you know the episode format, what fine-tuning is, and the idea of an embodiment.
A folder of episodes is not yet a dataset
You have recorded episodes — maybe eighty of them, each a folder of frames and a steps.jsonl (Chapter 13). That is data, but it is not yet a dataset in the shape a model’s training code expects. Between the two sits a conversion, and the conversion is more interesting than it sounds, because it is where you tell the model three things it must know before it can learn from your episodes:
- what each number means (this vector is joint positions, in this order),
- how big the numbers typically are (so it can normalize them), and
- which body it is looking at (a robot it already knows, or a brand-new one).
Get those three right and training just works. Get any of them wrong and the model trains on nonsense that looks like data. So the conversion is not mere reformatting — it is another contract with the model, exactly like the wire protocol was a contract for live control.
The conversion
The target format is LeRobot (specifically the v2.1 flavor GR00T expects) — a standard, efficient layout for robot demonstration datasets. A converter walks your episode folders and produces:
data/chunk-000/episode_000000.parquet — the per-step numbers (state, action, timestamp)
videos/chunk-000/observation.images.ego_view/
episode_000000.mp4 — the camera frames, encoded as H.264 video
meta/info.json — dataset shape and version
meta/stats.json — normalization statistics (below)
meta/episodes.jsonl, meta/tasks.jsonl — episode and task bookkeeping
meta/modality.json — what each part of the data *is* (below)
The frames become a compressed MP4 per episode rather than thousands of loose JPEGs (much smaller, and what the training pipeline reads), the per-step numbers become an efficient Parquet table, and the meta/ folder carries the metadata that makes the numbers meaningful. One small, real detail: the Unity recorder writes its text files with a byte-order mark, so the converter reads them as utf-8-sig — the kind of unglamorous compatibility fix that separates a pipeline that works from one that almost works.
modality.json: telling the model what its numbers mean
The recorder stored state and action as bare lists of 29 numbers. Bare lists are meaningless to a model that could be trained on many robots — is number 12 a joint angle, a gripper width, a velocity? modality.json is the answer key:
{
"state": { "joint_position": { "start": 0, "end": 29 } },
"action": { "joint_position": { "start": 0, "end": 29 } },
"video": { "ego_view": { "original_key": "observation.images.ego_view" } },
"annotation": { "human.task_description": { "original_key": "task_index" } }
}
It says: the 29-number state is joint positions; the 29-number action is joint positions (absolute targets, matching how Unity applies them); the video is the ego-view camera; the language annotation is the task description. This is the same lesson as the embodiment profile in Chapter 10 — a model needs its inputs named and shaped, not just handed over — now applied to the training data instead of the live observation. The wire protocol and the dataset format are two faces of one requirement: the model must know what each number is.
Normalization: the statistics the model needs
Neural networks learn best when their inputs are standardized — centered near zero and scaled to a consistent range — rather than raw radians that might run from −2.5 to +2.5 on one joint and −0.1 to +0.1 on another. So the converter computes, per dimension, the normalization statistics and stores them in meta/stats.json:
- mean and standard deviation — to center and scale each dimension;
- min and max — the observed range;
- 1st and 99th percentiles (
q01,q99) — robust range bounds that ignore rare outliers.
These are computed over your actual episodes — the concatenated actions, states, and timestamps — so they describe your data specifically. At training and inference time, the model uses them to normalize inputs and de-normalize its outputs back into real joint angles. This is why fine-tuning on your own demonstrations needs your own statistics: a model normalized to someone else’s data would systematically mis-scale yours.
Insight: the dataset format is a contract, exactly like the wire protocol. Chapter 5’s lesson — the model must be handed inputs it recognizes — returns here in a different costume. Live, the contract is the observation shape; offline, it is
modality.jsonplus the normalization stats. Both encode the same fact: a model’s numbers are only meaningful with their names and scales attached. Nearly every “the fine-tune trained but the robot does nonsense” failure is a broken version of this contract — a mislabeled modality, wrong-scale stats, an off-by-one action column.
How much does the model already know your robot?
The third thing the conversion declares is which body the demonstrations are on — and this determines how much work fine-tuning has to do. GR00T sorts robots into three tiers, easiest to hardest:
| Tier | What it means | Effort |
|---|---|---|
Pretrained (e.g. REAL_G1) |
The base checkpoint already drives this robot zero-shot. Fine-tuning gets a warm start — it is nudging a model that already half-knows the body. | Fewest demonstrations, best few-shot results |
| Fine-tune-only (e.g. Fourier GR-1) | The model ships the modality config for this robot but no checkpoint actually drives it. You must fine-tune, but you do not have to invent the body description. | No zero-shot, but the config is given |
| New embodiment (e.g. Unitree H1, and the 29-DoF push G1) | The model has no idea this robot exists. You author the modality config from scratch, and the projection layer that reads this body’s numbers trains from zero. No warm start. | Most data-hungry |
Insight: the tier decides your data budget before you record a single demo. A warm-start (pretrained) robot might reach a good policy on a few dozen demonstrations, because the model is adapting, not learning from scratch. A new-embodiment robot must learn the very meaning of its 29 numbers, and needs more. Knowing your robot’s tier up front tells you how many demonstrations to plan for and whether zero-shot is even an option. For hand-equipped manipulation, this is a concrete reason to prefer the Dex3-equipped G1 — it rides the pretrained
REAL_G1tier and its hands are full-fidelity (Chapter 7) — over robots that demand a cold-start fine-tune.
The push campaign in the next chapter is a new-embodiment case: a 29-DoF G1 configured as a body the base model had never seen, with a projection layer trained from scratch.
Two ecosystems
One more thing the conversion decides: which training ecosystem you are feeding. The two humanoid VLAs and the arm VLA do not share a dataset format:
- GR00T trains on LeRobot datasets (what this chapter described), plus the
modality.jsonand stats. - OpenVLA (the arm model) trains on RLDS, a different standard from the TensorFlow world, and there is a separate, older converter for its 7-DoF end-effector-delta episodes.
You do not need the details of RLDS to take the point: the record step produced one universal format (Chapter 13), and the convert step fans that one format out to whichever ecosystem the target model uses. The universal recording format is what makes it a fan-out rather than a fork all the way back to Unity — every model’s converter starts from the same episode folders.
What you now understand
- A folder of episodes is data, not yet a dataset. The conversion declares three things to the model: what each number means, how big the numbers are, and which body it is.
- The target is a LeRobot v2.1 dataset: an MP4 per episode, a Parquet table of per-step numbers, and a
meta/folder of metadata (including a smallutf-8-sigfix for Unity’s byte-order mark). modality.jsonnames the numbers (state/action=joint_position[0:29], the ego-view video, the task annotation) — the offline twin of the live observation contract.stats.jsoncarries per-dimension mean/std/min/max/q01/q99 so inputs and outputs are normalized to your data.- Robots fall into three tiers — pretrained (warm start, fewest demos), fine-tune-only (config ships, no checkpoint), and new embodiment (author the config, projector trains from scratch, most data). The tier sets your data budget before you record.
- The record step produces one universal format; the convert step fans it out to the target model’s ecosystem — LeRobot for GR00T, RLDS for OpenVLA.
The dataset is ready. The next chapter runs the actual fine-tune — eighty push demonstrations, one overnight training run on the owned GPU box, and the honest number for how well it worked.
Continue to Chapter 16 — The First Fine-Tune.
The GR00T converter is tools/convert_g1_push_to_lerobot.py (episodes → LeRobot v2.1, modality.json, stats); the new-embodiment modality config is scripts/spark/g1_push_config.py. The OpenVLA/RLDS path uses a separate converter (server/tools/convert_to_lerobot.py for the 7-DoF LeRobot variant; OpenVLA’s own fine-tune consumes RLDS). Embodiment tiers are described in the repository’s PROGRESS.md.