Chapter 04 — The Model as the Brain
This chapter is the brain: the vision-language-action model that turns an observation into an action. It assumes you know what a policy, a VLA, and an action are. We stay conceptual — you will not need any mathematics — because the point is to understand what kind of thing the model is and why it lives on its own machine.
What a model is, without the mystique
Strip away the word “model” and here is what sits on the other end of the wire: a very large function. It takes numbers in — the pixels of a camera image, the robot’s joint angles, the letters of an instruction — and returns numbers out — an action. Between the input and the output are billions of internal numbers, called parameters or weights, arranged in layers. The function’s behavior is entirely determined by those parameters.
Parameters (weights) — the internal numbers of a model that determine what it does. One of the models in this project has about 7 billion of them. They are not written by a programmer; they are learned — set automatically during training so that the function produces good actions.
That is the only bit of machine-learning intuition you need. A model is a function whose behavior is stored in billions of tunable numbers, and training is the process that tuned them. Once trained, the numbers are frozen into a file — the checkpoint — and running the model means loading that file and evaluating the function. Evaluating it once, to turn one observation into one action, is called inference.
Insight: the checkpoint is the skill. Everything the model “knows” lives in one file of frozen parameters. To give the robot a different skill, you do not rewrite code — you load a different checkpoint, or you adjust the current one with fine-tuning. This is why the model is swappable: the body and the wire never change, only which file of weights is answering.
What the training bought
The three models in this project were each trained, once, by a well-funded team, on an enormous corpus of robot experience — thousands of hours of many different robots doing many different tasks, often pooled from public datasets that dozens of labs contributed to. That training is what makes them worth using: a model that has watched thousands of robots pick things up and push things around has absorbed a rough, general sense of what those verbs look like as motion.
The result is a generalist. Show it a camera image and the words push the objects off the table, and it will produce a plausible pushing motion on a robot it has some familiarity with — without anyone writing pushing code, and without training it on your specific setup. That is the “instructed robot” from Chapter 1, and it is genuinely remarkable that it works at all.
But generalist is not specialist, and this is the honest catch that motivates the second half of the book:
Insight: zero-shot is a starting point, not a finish line. Running a pretrained model on your robot with no adaptation is called zero-shot — zero examples of your task. Zero-shot usually produces something reasonable and rarely something good: the motion is in the right spirit but imprecise, hesitant, or subtly wrong for your exact robot and camera. Closing that gap is what your demonstrations and fine-tuning are for. Expect mediocre zero-shot; the authors of these models expect it too. The pipeline exists precisely because the model alone is not enough.
Why the brain needs its own machine
Seven billion parameters is a lot of arithmetic. Running that function fast enough to drive a robot requires a GPU (graphics processing unit) — a chip built to do millions of simple calculations in parallel, originally for rendering video-game graphics, now the workhorse of machine learning. The largest model here needs roughly 16 gigabytes of GPU memory just to hold its weights, and a fraction of a second per inference even on capable hardware.
A laptop cannot do this at any useful speed, so the brain lives on a separate machine with a real GPU. In this project that is one of two places:
- an owned desktop supercomputer — an NVIDIA DGX Spark — sitting on a desk, reached over a private network so Unity can talk to it directly;
- or a rented cloud GPU — a Lambda instance — reached through a secure tunnel, which bills by the hour and so is torn down when idle.
The body stays on your laptop; the brain runs on the GPU box; the wire (next chapter) carries messages between them. Because the split is clean, moving the brain from the cloud to the desktop changes one thing — the address Unity dials — and nothing else.
The three brains
This project runs not one model but three, because there is no single VLA that drives every robot well, and running more than one is itself instructive. Here they are at a glance, with an honest status — GR00T drives Unity end to end; LingBot-VLA is served on the Spark; OpenVLA is so far only tested against a stub; later chapters give each its due.
| Model | Drives | Action style | Roughly |
|---|---|---|---|
| OpenVLA-7B | robot arms (Franka / WidowX) | a single 7-number end-effector delta, ~5 times a second | 7 billion parameters; predicts the action as a sequence of tokens, like a language model finishing a sentence |
| GR00T N1.7 | humanoids (Unitree G1, +hands; Fourier GR-1; Unitree H1) | a chunk of ~40 absolute joint targets, played back at 50 Hz | NVIDIA’s humanoid foundation model; a vision-language core with a diffusion-style action head; the primary model in this book |
| LingBot-VLA-V2 | other humanoids (G1, AgiBot A2, Fourier GR-2) | a chunk in a 55-number canonical action space | a Qwen3-VL vision-language core with a mixture-of-experts flow-matching action head |
You do not need to know what “diffusion,” “flow-matching,” or “tokens” mean to use these — those are the internal machinery each team chose for turning an observation into an action. What matters for the pipeline is the shape of what comes out: the arm model returns one small delta; the humanoid models return a chunk of absolute targets. That difference drives the design of everything downstream, and Chapter 8 is where it gets unpacked.
Insight: three models, one interface. The three brains could not be more different inside — different sizes, different action machinery, different robots. Yet all three sit behind the same wire protocol, so Unity talks to each of them with identical code. That is not luck; it is the design decision the next chapter is about. The variety of the brains is exactly why the uniformity of the interface is so valuable.
What you now understand
- A model is a large function whose behavior lives in billions of learned parameters, frozen into a checkpoint file. Running it on one observation is inference.
- Training on thousands of hours of many robots made each model a generalist — able to produce a plausible action for a familiar verb on a familiar robot, with no task-specific code.
- Zero-shot performance (no adaptation) is usually reasonable, not good. That gap is expected, and closing it is what demonstrations and fine-tuning are for.
- The model needs a GPU and so runs on its own machine — an owned desktop box or a rented cloud instance — while the body stays on the laptop.
- The project runs three VLAs: OpenVLA (arms, single delta), GR00T (humanoids, chunk — the primary one here), and LingBot (other humanoids, 55-dim chunk). They differ completely inside but share one interface. Status: GR00T drives Unity end to end; LingBot-VLA is served on the Spark; OpenVLA is so far only tested against a stub.
That shared interface is the wire. The next chapter is the single most important design decision in the project: one message format that makes the mock, the model, and even a human indistinguishable to Unity.
Continue to Chapter 05 — The Wire Between Them.
OpenVLA is loaded as openvla/openvla-7b; GR00T as nvidia/GR00T-N1.7-3B; LingBot as Robbyant/lingbot-vla-v2. In bfloat16, OpenVLA occupies ~16 GB of GPU memory (~8 GB when 4-bit quantized). Each model installs into its own environment on the GPU box; the per-model install notes — including the Blackwell-generation GPU quirks — live in the repository’s CLAUDE.md.