Skip to the content.

Chapter 13 — The Recorder That Doesn’t Care Who’s Driving

Part V is the heart of the book’s thesis. Everything so far has been about getting a robot to move; this is where you start teaching it. And the teaching begins with one component the entire project is designed around. It assumes you know what a demonstration, an episode, and a rollout are.


The pivot of the whole loop

Chapter 1 drew the loop: drive → record → fine-tune → serve. The record step is the pivot. It is where a robot doing a task turns into data a model can learn from. And the single design decision that makes the loop close is this: the recorder is decoupled from whatever is driving the robot.

Read that carefully, because it is the thesis made concrete. The recorder does not know, and does not ask, whether the actions it is capturing came from a policy, a keyboard, or a pair of tracked hands. It exposes one method — record this step — and anything driving the robot calls it. A model rollout and a human demonstration flow through the exact same recorder and come out as the exact same files.

Insight: decoupling the recorder from the driver is what makes “one format, many drivers” real. If the recorder had to know who was driving — one format for policy rollouts, another for keyboard, another for hand-tracking — then every new way of driving the robot would need its own data pipeline, and demonstrations would not be interchangeable with rollouts. Instead there is one recorder with one output, and the driver is just whoever happens to be calling it. This is why a demonstration you record by hand is already, with no conversion, in the same shape as the model’s own output. The loop closes here.


What a recorded step is

An episode on disk is a folder, and its contents are deliberately plain — readable, inspectable, no exotic format:

vla_recordings/episode_000000/
  meta.json          — {"instruction", "fps", "robot_type", "success", "num_steps"}
  steps.jsonl        — one line per step (see below)
  frames/000000.jpg  — one camera frame per step, index-aligned with steps.jsonl

meta.json describes the whole episode: the instruction that was in effect, the frame rate, which robot, whether the attempt succeeded, and how many steps it ran. The frames are ordinary JPEGs, one per step. And the substance is steps.jsonl — a text file with one JSON object per line, each object a single recorded step:

{"timestep": 0, "state": [ ... ], "action": [ ... ], "joints": [ ... ], "timestamp": 0.0000}

That is the whole format. It is intentionally N-dimensional and generic: state and action are just lists of numbers, so the identical recorder handles a 7-DoF arm and a 29-DoF humanoid with no change. One recorder, every robot.


The decision-boundary contract

There is one rule the recorder is strict about, and it is the difference between clean training data and subtly corrupt data:

The recorded state and frame must be the observation the action was decided from — not the state after the action was applied.

A training example is a claim: given this observation, this was the right action. If you record the observation after the action has already moved the robot, the pairing is a lie — the model would learn to predict an action from the situation that action already produced. So the recorder captures the frame and proprioception at the decision boundary: the instant before the action is sent, using exactly the observation that produced it. The arm bridge records (image, proprio, action) as one aligned triple at that boundary; the humanoid bridge records one row per request — one observation-to-action pair at the rate the model actually thinks — which is precisely the alignment fine-tuning expects.

This is easy to get wrong and impossible to see later — a dataset with off-by-one observation/action pairing looks perfectly well-formed and trains a subtly worse model. Getting the alignment right in the recorder, once, means every demonstration is honest by construction.


The correction flag: teaching by fixing

One optional field on a step hints at a more advanced way of teaching, and it is a single boolean:

{"timestep": 57, "state": [...], "action": [...], "correction": true, "timestamp": 11.4}

When a human is watching a policy drive and takes over to fix a mistake — nudging the robot back on track — the steps recorded during that takeover are marked "correction": true. This is the seed of a technique called DAgger (Dataset Aggregation): instead of only demonstrating tasks from scratch, you let the policy try, correct it exactly where it goes wrong, and train on those corrections. The corrections target the policy’s actual failure modes, which is often far more data-efficient than fresh demonstrations. The flag is a one-field extension precisely because the recorder was decoupled from the driver — “a human corrected here” is just another thing a driver can tell the recorder, and the downstream converter can later filter to episodes that contain corrections.

Insight: a demonstration and a rollout are the same file, and that is the whole point. Open two episode folders — one recorded from a policy, one recorded from your own hands — and you cannot tell which is which. Same meta.json, same steps.jsonl shape, same frames. The only difference is a fact you’d have to remember, not read. That indistinguishability is not a coincidence to admire; it is the mechanism by which your effort at the keyboard or in front of a camera becomes training data with zero glue. Every driver you add — and the next chapter adds your hands — inherits the entire downstream pipeline for free.


What you now understand

The recorder is ready to capture a demonstration. The only thing missing is a way for you to drive — to become the policy for a while. That is the next chapter: teleoperation, from a keyboard to your own hands seen through a camera.

Continue to Chapter 14 — Becoming the Robot.


The recorder is DevVLA/Assets/Scripts/VLA/EpisodeRecorder.cs, driven through a single RecordStep(jpeg, state, action, joints) call from VLABridge.cs (arm), GrootBridge.cs (humanoid), and the teleop components. Episodes are written under Unity’s persistent data path; the correctionActive flag emits the DAgger marker. The on-disk format is consumed by the dataset converters in Chapter 15.