Chapter 01 — The Big Picture
Welcome. This is the first chapter of a workbook that starts from nothing and ends with a robot learning a new skill from demonstrations you record by hand. You do not need a robotics or programming background. Every term will be defined the first time it appears.
What is this project about?
Point a camera at a table with some blocks on it. Type a sentence: push the objects off the table. A moment later, a robot arm — or a humanoid robot the size of a teenager — starts moving, reaching out, and sweeping the blocks off the edge. Nobody wrote instructions for how to move the joints. Nobody choreographed the reach. The only inputs were a picture and a sentence in plain English; the output was motion.
The thing that turned the picture and the sentence into motion is called a vision-language-action model — a VLA for short. It is a single piece of software, trained on enormous amounts of robot data, that takes in vision (a camera image), language (an instruction), and produces actions (numbers that drive a robot’s joints). This project is about putting one of those models to work.
But a model on its own does nothing useful. It needs a body to drive, a way to talk to that body, a way for you to take over and show it what you meant, and a way to turn what you showed it back into a better model. This workbook is about building that whole loop — and the loop, not the model, is the interesting part.
This chapter unpacks three big ideas that the rest of the book rests on:
- What it means for a robot to be instructed rather than programmed or even trained from scratch.
- Why the system is split into a body and a brain with a wire between them.
- What the full loop is — drive, record, fine-tune, serve — and the one design decision that makes it work.
Aim for intuition here. The precise vocabulary comes in Chapter 2.
Programmed, learned, instructed: three ways to make a robot move
Imagine you want a robot to pick up a cup.
The oldest approach is to program it. An engineer writes down, in painstaking detail, exactly what each joint should do at each instant: the shoulder rotates to 32°, the elbow to 87°, the gripper closes when the fingers are 4 cm apart. This works, and it is precise, but it is brittle. Move the cup two inches and the whole script is wrong. Every new object, every new table, every new task is a new program.
A newer approach is to learn the motion from scratch, by trial and error, inside a simulator — a robot practicing millions of times, getting a score, and gradually discovering a skill. This is powerful (it is exactly how the companion project to this one teaches a humanoid to walk), but it is expensive: each new skill is its own long training campaign, and the robot starts every one knowing nothing.
This project is mostly about a third way: the robot is instructed. A vision-language-action model has already been trained — once, by someone with a very large budget — on thousands of hours of robots doing thousands of different tasks. It arrives already knowing, in a rough and general way, what “pick up,” “push,” and “put down” tend to look like. You do not program it and you do not train it from zero. You instruct it: you hand it a camera image and a sentence, and it produces a plausible action. Change the sentence and you get different behavior, from the same model, with no new code.
Insight: the instruction is the interface. With a programmed robot, you change behavior by editing joint angles. With a from-scratch learned robot, you change behavior by designing a score and retraining. With an instructed robot, you change behavior by changing the sentence — and the same model handles a task it has a rough idea about without any retraining at all. The engineering has moved again: from specifying motion, to specifying a goal, to specifying an instruction and a body for the model to drive.
There is a catch, and it is the thread that runs through the whole book. A general model that has seen thousands of robots is a generalist. It will do something reasonable on your robot and your task — but “reasonable” is often not “good.” Getting from reasonable to good is where your demonstrations come in, and that is the second half of this book.
Body, brain, and the wire between them
The system has two halves that could not be more different, and understanding why they are split is half of understanding the project.
The body is a simulated robot living inside Unity — a game engine, the same kind of software used to build video games, here taken seriously as a physics simulator. Inside Unity there is a digital robot with joints, motors, a camera, and a world to act in. It runs comfortably on a laptop. When the robot falls over, nothing breaks and nothing costs money. (Chapter 3 is all about the body.)
The brain is the vision-language-action model. These models are large — one of the three in this project has 7 billion internal numbers it consults to make each decision — and they need a specialized graphics processor (a GPU) to run at any usable speed. That GPU lives in a separate box: sometimes a rented cloud machine, sometimes an owned desktop supercomputer sitting on a desk. The brain does not run on your laptop. (Chapter 4 is all about the brain.)
Between them runs a wire: a network connection carrying small text messages back and forth. Unity takes a picture, packages it with the current instruction, and sends it over the wire. The brain thinks for a fraction of a second and sends back an action. Unity applies the action to the robot’s joints, takes the next picture, and the cycle repeats. (Chapter 5 is all about the wire.)
Insight: the split is not an accident, it is the design. Keeping the body and the brain in separate programs, talking over a network, means the body never needs to know what the brain is. It could be a giant model on a supercomputer, a scripted stand-in running on the same laptop, or a human moving their hands in front of a webcam. As long as the thing on the other end of the wire sends back actions in the agreed format, Unity cannot tell the difference — and does not care. This one fact is what the entire pipeline is built on.
The loop: drive, record, fine-tune, serve
Here is the whole project in one paragraph. You drive the simulated robot from a model and watch what it does. When the model is not good enough — and at first it never is — you take over and record yourself doing the task correctly, by keyboard or by moving your own hands. Those recordings are demonstrations. You fine-tune the model on your demonstrations — a short, cheap training run that adjusts the general model toward your specific robot and task. Then you serve the improved model back to the same Unity robot and watch it do better. Then you do it again.
flowchart LR
A["VLA model<br/>(the brain)"] -->|action over the wire| B["Unity robot<br/>(the body)"]
B -->|"image + instruction"| A
B -.->|you take over| C["Record a<br/>demonstration"]
A -.->|"policy rollout"| C
C --> D["Fine-tune<br/>on your demos"]
D -->|"serve the<br/>improved model"| A
The magic is in the dashed lines. The demonstration you record when you drive the robot comes out in the exact same format as the actions the model produces when it drives the robot. A recording is a recording, whether a policy or a human generated it. That is not a coincidence; it is the single most important design choice in the codebase, and it is why the recorder in Chapter 13 is deliberately built to not care who is driving.
Insight: one action format, many drivers. A mock stand-in, a human at a keyboard, a human waving their hands at a webcam, and a 7-billion-parameter model all speak the same action language. So every way of driving the robot produces the same kind of data, and every demonstration you record is already in the shape a model trains on. The loop closes because the format never changes. This is the sentence the whole book circles: the model is the easy part; the pipeline — held together by one shared format — is the product.
What you now understand
- A vision-language-action model (VLA) takes a camera image and a language instruction and produces robot actions. It is trained once, at great expense, and then instructed — you change behavior by changing the sentence, not the code.
- An instructed robot is a third thing next to a programmed robot (hand-coded joint angles) and a from-scratch learned robot (trial and error in a simulator). It starts out generally capable but rarely task-perfect — which is why your demonstrations matter.
- The system is a body (a simulated robot in Unity, running on a laptop), a brain (the model, running on a GPU box), and a wire (a network connection carrying small messages). They are split on purpose, so the body never needs to know what is on the other end of the wire.
- The whole project is a loop: drive → record → fine-tune → serve. It closes because a demonstration and a model’s output share one action format — one format, many drivers.
These four ideas are the scaffold everything else rests on. The next chapter makes them precise: it defines the dozen exact words — observation, action, policy, embodiment, and the rest — that let us describe the loop without hand-waving.
Continue to Chapter 02 — The Vocabulary.
This project wires three vision-language-action models to simulated Unity robots — a robot arm and several humanoids — over one WebSocket protocol, on a Mac for development and an NVIDIA GPU box for real inference: GR00T drives Unity end to end; LingBot-VLA is served on the Spark; OpenVLA is so far only tested against a stub. Everything in this chapter is expanded, with real code and real messages, in the chapters that follow.