A sim2real workbook — Learning Series
A twenty-chapter workbook that takes a reader with no robotics or machine-learning background from “what is a simulator?” to fine-tuning a vision-language-action model on your own demonstrations and serving it back to a robot you drive in Unity. Every term is defined the first time it appears. Read the chapters in order — each one builds on the concepts the previous one introduced, and the series is designed as one cohesive arc, not a collection of standalone notes.
Part I — What this is
The two questions a newcomer brings before anything else can land: what kind of project is this, and what are the words I need before I read further?
| Chapter | Hook |
|---|---|
| Chapter 1 — The Big Picture | A robot that reads a picture and a sentence and answers with motion — here is the whole loop at thirty thousand feet, and why “instructed” is a third thing next to “programmed” and “learned.” |
| Chapter 2 — The Vocabulary | The dozen precise terms — observation, action, policy, embodiment, degree of freedom, end-effector, proprioception, teleoperation, demonstration, fine-tuning, chunk, sim2real — that let us talk about the loop without hand-waving. |
Part II — The two halves and the wire
The system is a body, a brain, and a wire. One chapter each on what they are and why they are split the way they are.
| Chapter | Hook |
|---|---|
| Chapter 3 — Unity as the Body | The simulator is a robot you can crash for free: what a game engine taken seriously as physics gives you, how a robot description becomes an articulated puppet, and why we simulate before we touch hardware. |
| Chapter 4 — The Model as the Brain | What a vision-language-action model actually is — image plus instruction in, action out — and a first look at the three we use, why they need a GPU, and what “7 billion parameters” is doing for you. |
| Chapter 5 — The Wire Between Them | One JSON message format on a WebSocket connects body to brain — and the single most important design decision in the project: the mock and the real model are indistinguishable to Unity, so switching between them is a flag. |
Part III — Speaking the robot’s language
Three chapters on the conventions that decide whether the smartest model flails or the dumbest mock looks real.
| Chapter | Hook |
|---|---|
| Chapter 6 — Left, Right, and Why They Disagree | The robot and Unity disagree about which way is left — one is right-handed, the other left-handed — and a single wrong sign flips a robot inside out. The coordinate-frame conversion, worked through both directions. |
| Chapter 7 — Joints in a Row | A robot’s brain hands back a list of numbers, and everything depends on which number drives which joint. Document order, the 29-joint humanoid, and the mimic-joint gotcha that cost real debugging. |
| Chapter 8 — Deltas, Targets, and the Chunk | Two robots, two dialects: the arm is told to move a little five times a second; the humanoid is handed forty exact poses at once and plays them back at fifty. Why the difference exists and what a chunk buys you. |
Part IV — Driving the robot
The system running for real — including the most instructive failure in the book.
| Chapter | Hook |
|---|---|
| Chapter 9 — First Light | Before any GPU, a scripted mock policy proves the entire pipeline on a laptop — observation to transport to action buffer to control to recording — so that when a real model arrives, only one thing is new. |
| Chapter 10 — A Real Brain on a Real Box | A real GR00T model on an owned GPU box drives the Unity humanoid end to end — and the reality that its observation is not a flat array but a nested structure keyed by embodiment, the first real friction point. |
| Chapter 11 — The Cadence Bug | The bridge advertised fifty hertz; the model was trained at five. The motion played ten times too fast and the action buffer starved — a protocol number that lied, and what it teaches about the seams between systems. |
| Chapter 12 — Watching the Robot Think | You cannot debug what you cannot see. The visualization layer that draws the posed robot, every action and proprioception channel, and the end-effector’s path through space — the instrument panel for the whole loop. |
Part V — Teaching it your way
The heart of the thesis: one recorder, many drivers.
| Chapter | Hook |
|---|---|
| Chapter 13 — The Recorder That Doesn’t Care Who’s Driving | The single component the whole project is designed around — a recorder decoupled from the action source, so a policy rollout, a keyboard, and a pair of tracked hands all write the exact same training data. |
| Chapter 14 — Becoming the Robot | Teleoperation: your hands, seen through a webcam or a pair of camera glasses, retargeted onto a robot’s fingers and arms. What transfers cleanly, what monocular vision gets wrong, and the honest gap between the two. |
Part VI — Closing the loop
From your demonstrations to a fine-tuned policy and back to the robot.
| Chapter | Hook |
|---|---|
| Chapter 15 — From Demonstrations to a Dataset | Recorded episodes become a training dataset — the conversion, the “new embodiment” description a model needs to accept a body it has never seen, and the three tiers of how much a model already knows about your robot. |
| Chapter 16 — The First Fine-Tune | Eighty push demonstrations, one overnight training run on an owned GPU, and the honest result: the model learned the sweep to about one and a half degrees of error — and why open-loop demonstrations cap how good it can get. |
Part VII — Beyond imitation
Three tracks around the humanoid spine.
| Chapter | Hook |
|---|---|
| Chapter 17 — A Second Opinion | Why the project runs more than one model: a second VLA for robot arms that speaks deltas, a third for other humanoids that speaks a fifty-five-number canonical action — and the embodiment gate that decides what runs today versus after a fine-tune. |
| Chapter 18 — When Copying Isn’t Enough | Imitation cannot teach balance. A position-controlled biped falls over in physics without a controller it had to learn — the reinforcement-learning track, and the sim-to-sim transfer that carries a policy from one physics engine to another. |
| Chapter 19 — The Classical Counterpoint | A reminder that not everything should be learned. A classical motion planner solves for a smooth, jerk-limited trajectory in closed form — and cuts peak command jerk by a factor of three hundred thousand — where a learned policy would only approximate it. |
Part VIII — Synthesis
| Chapter | Hook |
|---|---|
| Chapter 20 — The Pipeline Is the Product | Strip nineteen chapters to one sentence: the model is the easy part, one action format is what makes the loop possible, and the loop — not the model — is the thing you build. |
Appendix
Reference material for readers who want the exact numbers.
| Methods and protocol reference | One lookup surface for the whole toolkit: the WebSocket protocol message-by-message, the port map, the embodiment profiles, every server flag, the frame and joint conventions, the recording format, and the known gotchas. |
Figures
Diagrams are drawn inline. Screenshots, Rerun captures, and demonstration clips
are marked in the chapter source with a hidden 📷 Figure — placeholder comment that names the exact scene or
command that produces them; drop the captured media in under
assets/ and replace the placeholder. Everything the book describes is
runnable from the main project’s code (private repository; available on request) — its main README
and TESTING.md have the per-robot recipes.