Chapter 02 — The Vocabulary
This chapter builds the words. Chapter 1 gave you the shape of the project in plain language; this one makes it precise. Every term here comes back constantly, so it is worth reading slowly. Nothing new is built — we are just naming the parts.
Why a vocabulary chapter
In Chapter 1 we said the model “takes a picture and a sentence and produces motion.” That is true, but it is too vague to build on. What exactly is “a picture and a sentence”? What exactly is “motion” — a position, a speed, a nudge? When a humanoid hands back forty numbers, which number is which?
The field of robot learning has precise words for all of this, and once you have them, the rest of the book reads easily. We will define about a dozen, in the order they show up in one pass through the loop. The running example is the one this project actually ran: a Unitree G1 — a humanoid robot the size of a teenager — being told to push the objects off the table.
The body: embodiment, degrees of freedom, end-effectors
Embodiment — the specific robot body a model is driving: how many joints it has, in what arrangement, what sensors it carries, and what its actions are allowed to be. A robot arm and a humanoid are different embodiments; so are two humanoids with different hands.
Embodiment is the first word because it is the one everything else is relative to. A model trained to drive one embodiment cannot simply drive another — the number of joints is different, the sensors are different, the very meaning of “action number 12” is different. Much of the hard work in this project is telling a model, precisely, what embodiment it is looking at. (Chapter 10 is where that stops being abstract.)
Degree of freedom (DoF) — one independently controllable axis of motion. A hinge that bends one way is one degree of freedom. A shoulder that can rotate three ways is three.
The G1 in this project has 29 degrees of freedom: six in each leg, three in the waist, and seven in each arm. When we say “a 29-DoF humanoid,” we mean the robot’s pose is fully described by 29 numbers, one per joint. A robot arm like the Franka Panda has 7. Count the degrees of freedom and you know how many numbers an observation and an action need.
End-effector — the business end of a limb: the part that touches the world. On an arm it is the gripper; on a hand it is the fingertips; on a leg it is the foot.
Often you care less about every joint angle and more about where the end-effector is — where the gripper sits in space. That distinction (joints versus end-effector) turns out to matter a great deal, and Chapter 8 is built on it.
What the model sees: observation and proprioception
Observation — everything the model is given at one moment to decide its next action. In this project an observation is a camera image, the robot’s proprioception, and the instruction.
Every trip around the loop starts with an observation. Unity assembles one — snaps a picture from the robot’s camera, reads the robot’s current pose, attaches the instruction — and sends it over the wire. The model’s entire view of the world is that one message. If something the model needs is not in the observation, the model cannot use it.
Proprioception — a body’s sense of its own configuration. For a robot, it is the current reading of its joints (their angles) and sometimes the position of its end-effector. It is the robot’s answer to “where are my limbs right now?”
The word comes from human physiology — it is how you can touch your nose with your eyes closed. For the arm in this project, proprioception is a seven-number vector: the end-effector’s position and orientation plus how open the gripper is. For the humanoid, it is the 29 joint angles. Proprioception is half of the observation; the camera image is the other half.
What the model is and does: policy, action, VLA
Policy — the function that turns an observation into an action. Give it what the robot sees; it gives back what the robot should do. The policy is the “brain” from Chapter 1.
Policy is the central noun of the whole field. A mock policy is a few lines of arithmetic. A learned policy is a neural network with billions of numbers inside. Both are policies because both answer the same question: observation in, action out. This shared shape is exactly why a mock and a real model are interchangeable behind the wire.
Action — what the policy outputs: the numbers that command the robot for the next step. An action can be a delta (a small change: move 2 cm forward, rotate a little) or an absolute target (put joint 12 at 0.4 radians).
The delta-versus-absolute distinction is not a detail; it is a genuine fork in how the two robot tracks work. The arm gets deltas — nudges relative to where it is now. The humanoid gets absolute joint targets — exact angles to move to. Chapter 8 is devoted to why.
Vision-language-action model (VLA) — a policy whose observation includes an image (vision) and a natural-language instruction (language), and whose output is an action. The three models in this project — OpenVLA, GR00T, and LingBot — are all VLAs.
A VLA is just a policy with a particular kind of observation: one that includes a sentence. That sentence is what makes the robot instructable — the same model, same weights, produces different behavior when you change the words.
How you teach it: episode, demonstration, teleoperation, fine-tuning
Episode — one complete attempt at a task, from start to reset: a sequence of (observation, action) pairs recorded step by step. Also called a rollout when a policy generated it.
An episode is the unit of recorded experience. Push the blocks off the table once, from first frame to last, and you have one episode: a folder of camera frames and a list of the action taken at each frame.
Demonstration — an episode produced by a human doing the task correctly, recorded to teach a model what “correct” looks like. A demonstration and a policy rollout have the exact same shape on disk; the only difference is who was driving.
That last sentence is the whole thesis in miniature, and Chapter 13 makes it real.
Teleoperation — a human directly driving the robot in real time to produce demonstrations: by keyboard, by moving their own tracked hands, or eventually through a headset. “Teleop” for short.
Teleoperation is how you become the policy for a while — you supply the actions, the robot obeys, and the recorder captures it as a demonstration. Chapter 14 is all about it.
Fine-tuning — taking a model that was already trained, at great expense, on huge amounts of data, and running a short, cheap additional training pass on your demonstrations to adapt it to your specific robot and task. You are not training from scratch; you are nudging a capable generalist toward your specialty.
Fine-tuning is the payoff of the whole loop. It is what turns “the general model does something reasonable” into “the model does my task well.” Chapters 15 and 16 do it for real, on 80 recorded demonstrations.
The chunk, and the goal: sim2real
Action chunk — a whole batch of future actions predicted in one shot. Instead of returning a single next action, a chunked policy returns, say, the next forty actions at once, and the robot plays them back in order while the policy thinks about the batch after that.
The humanoid models in this project are chunked; the arm model is not. A chunk is how a slow model (one that takes half a second to think) can still drive a robot that needs commands fifty times a second. Chapter 8 works through the mechanics, and Chapter 11 is the cautionary tale of getting the chunk’s timing wrong.
Sim2real (simulation-to-reality) — the discipline of producing behavior in simulation that will still work when transferred to a physical robot. The whole point of practicing in Unity is eventually to drive real hardware, and the gap between the two is where a great deal of robotics effort goes.
This project lives on the simulation side of that gap — everything here runs against a Unity robot — but the entire design points across it. The data formats, the action conventions, the coordinate frames: all of them are chosen to match how real robots and real datasets work, so that nothing has to be rebuilt when hardware arrives.
What you now understand
You now have the working vocabulary for the rest of the book:
- Embodiment (which body), degree of freedom (one controllable axis), and end-effector (the part that touches the world) describe the robot.
- Observation (image + proprioception + instruction) is what the model sees; proprioception is the robot’s sense of its own pose.
- Policy (observation → action) is the brain; an action is either a delta or an absolute target; a VLA is a policy whose observation includes an image and a sentence.
- Episode / demonstration / rollout are recorded attempts; teleoperation is a human producing demonstrations; fine-tuning adapts a pretrained model to your data.
- An action chunk is a batch of future actions predicted at once; sim2real is the goal every convention is chosen to serve.
With the words in hand, we can look at the parts. The next three chapters take the body, the brain, and the wire one at a time — starting with the body, the simulated robot living inside Unity.
Continue to Chapter 03 — Unity as the Body.
The running example — a 29-DoF Unitree G1 told to “push the objects off the table” — is a real task in this repository: 80 demonstrations of it were recorded, converted, and used to fine-tune a model, the campaign documented in Chapters 15 and 16.