Chapter 13 — The Recorder That Doesn’t Care Who’s Driving
Part V is the heart of the book’s thesis. Everything so far has been about getting a robot to move; this is where you start teaching it. And the teaching begins with one component the entire project is designed around. It assumes you know what a demonstration, an episode, and a rollout are.
The pivot of the whole loop
Chapter 1 drew the loop: drive → record → fine-tune → serve. The record step is the pivot. It is where a robot doing a task turns into data a model can learn from. And the single design decision that makes the loop close is this: the recorder is decoupled from whatever is driving the robot.
Read that carefully, because it is the thesis made concrete. The recorder does not know, and does not ask, whether the actions it is capturing came from a policy, a keyboard, or a pair of tracked hands. It exposes one method — record this step — and anything driving the robot calls it. A model rollout and a human demonstration flow through the exact same recorder and come out as the exact same files.
Insight: decoupling the recorder from the driver is what makes “one format, many drivers” real. If the recorder had to know who was driving — one format for policy rollouts, another for keyboard, another for hand-tracking — then every new way of driving the robot would need its own data pipeline, and demonstrations would not be interchangeable with rollouts. Instead there is one recorder with one output, and the driver is just whoever happens to be calling it. This is why a demonstration you record by hand is already, with no conversion, in the same shape as the model’s own output. The loop closes here.
What a recorded step is
An episode on disk is a folder, and its contents are deliberately plain — readable, inspectable, no exotic format:
vla_recordings/episode_000000/
meta.json — {"instruction", "fps", "robot_type", "success", "num_steps"}
steps.jsonl — one line per step (see below)
frames/000000.jpg — one camera frame per step, index-aligned with steps.jsonl
meta.json describes the whole episode: the instruction that was in effect, the frame rate, which robot, whether the attempt succeeded, and how many steps it ran. The frames are ordinary JPEGs, one per step. And the substance is steps.jsonl — a text file with one JSON object per line, each object a single recorded step:
{"timestep": 0, "state": [ ... ], "action": [ ... ], "joints": [ ... ], "timestamp": 0.0000}
timestep— the step index.state— the proprioception the model was given (the arm’s 7-number pose, or the humanoid’s 29 joint angles).action— the action taken this step.joints— optionally, the full measured joint vector, so the episode can be replayed as a posed robot in the visualizer later.timestamp— seconds since the episode began.
That is the whole format. It is intentionally N-dimensional and generic: state and action are just lists of numbers, so the identical recorder handles a 7-DoF arm and a 29-DoF humanoid with no change. One recorder, every robot.
The decision-boundary contract
There is one rule the recorder is strict about, and it is the difference between clean training data and subtly corrupt data:
The recorded
stateandframemust be the observation the action was decided from — not the state after the action was applied.
A training example is a claim: given this observation, this was the right action. If you record the observation after the action has already moved the robot, the pairing is a lie — the model would learn to predict an action from the situation that action already produced. So the recorder captures the frame and proprioception at the decision boundary: the instant before the action is sent, using exactly the observation that produced it. The arm bridge records (image, proprio, action) as one aligned triple at that boundary; the humanoid bridge records one row per request — one observation-to-action pair at the rate the model actually thinks — which is precisely the alignment fine-tuning expects.
This is easy to get wrong and impossible to see later — a dataset with off-by-one observation/action pairing looks perfectly well-formed and trains a subtly worse model. Getting the alignment right in the recorder, once, means every demonstration is honest by construction.
The correction flag: teaching by fixing
One optional field on a step hints at a more advanced way of teaching, and it is a single boolean:
{"timestep": 57, "state": [...], "action": [...], "correction": true, "timestamp": 11.4}
When a human is watching a policy drive and takes over to fix a mistake — nudging the robot back on track — the steps recorded during that takeover are marked "correction": true. This is the seed of a technique called DAgger (Dataset Aggregation): instead of only demonstrating tasks from scratch, you let the policy try, correct it exactly where it goes wrong, and train on those corrections. The corrections target the policy’s actual failure modes, which is often far more data-efficient than fresh demonstrations. The flag is a one-field extension precisely because the recorder was decoupled from the driver — “a human corrected here” is just another thing a driver can tell the recorder, and the downstream converter can later filter to episodes that contain corrections.
Insight: a demonstration and a rollout are the same file, and that is the whole point. Open two episode folders — one recorded from a policy, one recorded from your own hands — and you cannot tell which is which. Same
meta.json, samesteps.jsonlshape, same frames. The only difference is a fact you’d have to remember, not read. That indistinguishability is not a coincidence to admire; it is the mechanism by which your effort at the keyboard or in front of a camera becomes training data with zero glue. Every driver you add — and the next chapter adds your hands — inherits the entire downstream pipeline for free.
What you now understand
- The record step is the pivot of the loop, and the recorder is decoupled from the driver: it exposes one record-this-step call, and a policy, a keyboard, or tracked hands all invoke it identically.
- An episode is a plain folder:
meta.json(instruction, fps, robot type, success, step count),steps.jsonl(one JSON object per step:timestep,state,action, optionaljoints,timestamp), and index-aligned JPEG frames. The format is N-dimensional and generic — the same recorder serves a 7-DoF arm and a 29-DoF humanoid. - The recorder enforces the decision-boundary contract: the recorded observation is the one the action was decided from, so every training pair is an honest claim. The humanoid records one row per model request.
- An optional
"correction": trueflag marks human takeovers, seeding DAgger-style training on a policy’s actual mistakes. - Because a demonstration and a rollout are the same file, any new way of driving the robot inherits the whole downstream pipeline for free — which is exactly “one format, many drivers.”
The recorder is ready to capture a demonstration. The only thing missing is a way for you to drive — to become the policy for a while. That is the next chapter: teleoperation, from a keyboard to your own hands seen through a camera.
Continue to Chapter 14 — Becoming the Robot.
The recorder is DevVLA/Assets/Scripts/VLA/EpisodeRecorder.cs, driven through a single RecordStep(jpeg, state, action, joints) call from VLABridge.cs (arm), GrootBridge.cs (humanoid), and the teleop components. Episodes are written under Unity’s persistent data path; the correctionActive flag emits the DAgger marker. The on-disk format is consumed by the dataset converters in Chapter 15.