Skip to the content.

A sim2real workbook — Learning Series

A twenty-chapter workbook that takes a reader with no robotics or machine-learning background from “what is a simulator?” to fine-tuning a vision-language-action model on your own demonstrations and serving it back to a robot you drive in Unity. Every term is defined the first time it appears. Read the chapters in order — each one builds on the concepts the previous one introduced, and the series is designed as one cohesive arc, not a collection of standalone notes.


Part I — What this is

The two questions a newcomer brings before anything else can land: what kind of project is this, and what are the words I need before I read further?

Chapter Hook
Chapter 1 — The Big Picture A robot that reads a picture and a sentence and answers with motion — here is the whole loop at thirty thousand feet, and why “instructed” is a third thing next to “programmed” and “learned.”
Chapter 2 — The Vocabulary The dozen precise terms — observation, action, policy, embodiment, degree of freedom, end-effector, proprioception, teleoperation, demonstration, fine-tuning, chunk, sim2real — that let us talk about the loop without hand-waving.

Part II — The two halves and the wire

The system is a body, a brain, and a wire. One chapter each on what they are and why they are split the way they are.

Chapter Hook
Chapter 3 — Unity as the Body The simulator is a robot you can crash for free: what a game engine taken seriously as physics gives you, how a robot description becomes an articulated puppet, and why we simulate before we touch hardware.
Chapter 4 — The Model as the Brain What a vision-language-action model actually is — image plus instruction in, action out — and a first look at the three we use, why they need a GPU, and what “7 billion parameters” is doing for you.
Chapter 5 — The Wire Between Them One JSON message format on a WebSocket connects body to brain — and the single most important design decision in the project: the mock and the real model are indistinguishable to Unity, so switching between them is a flag.

Part III — Speaking the robot’s language

Three chapters on the conventions that decide whether the smartest model flails or the dumbest mock looks real.

Chapter Hook
Chapter 6 — Left, Right, and Why They Disagree The robot and Unity disagree about which way is left — one is right-handed, the other left-handed — and a single wrong sign flips a robot inside out. The coordinate-frame conversion, worked through both directions.
Chapter 7 — Joints in a Row A robot’s brain hands back a list of numbers, and everything depends on which number drives which joint. Document order, the 29-joint humanoid, and the mimic-joint gotcha that cost real debugging.
Chapter 8 — Deltas, Targets, and the Chunk Two robots, two dialects: the arm is told to move a little five times a second; the humanoid is handed forty exact poses at once and plays them back at fifty. Why the difference exists and what a chunk buys you.

Part IV — Driving the robot

The system running for real — including the most instructive failure in the book.

Chapter Hook
Chapter 9 — First Light Before any GPU, a scripted mock policy proves the entire pipeline on a laptop — observation to transport to action buffer to control to recording — so that when a real model arrives, only one thing is new.
Chapter 10 — A Real Brain on a Real Box A real GR00T model on an owned GPU box drives the Unity humanoid end to end — and the reality that its observation is not a flat array but a nested structure keyed by embodiment, the first real friction point.
Chapter 11 — The Cadence Bug The bridge advertised fifty hertz; the model was trained at five. The motion played ten times too fast and the action buffer starved — a protocol number that lied, and what it teaches about the seams between systems.
Chapter 12 — Watching the Robot Think You cannot debug what you cannot see. The visualization layer that draws the posed robot, every action and proprioception channel, and the end-effector’s path through space — the instrument panel for the whole loop.

Part V — Teaching it your way

The heart of the thesis: one recorder, many drivers.

Chapter Hook
Chapter 13 — The Recorder That Doesn’t Care Who’s Driving The single component the whole project is designed around — a recorder decoupled from the action source, so a policy rollout, a keyboard, and a pair of tracked hands all write the exact same training data.
Chapter 14 — Becoming the Robot Teleoperation: your hands, seen through a webcam or a pair of camera glasses, retargeted onto a robot’s fingers and arms. What transfers cleanly, what monocular vision gets wrong, and the honest gap between the two.

Part VI — Closing the loop

From your demonstrations to a fine-tuned policy and back to the robot.

Chapter Hook
Chapter 15 — From Demonstrations to a Dataset Recorded episodes become a training dataset — the conversion, the “new embodiment” description a model needs to accept a body it has never seen, and the three tiers of how much a model already knows about your robot.
Chapter 16 — The First Fine-Tune Eighty push demonstrations, one overnight training run on an owned GPU, and the honest result: the model learned the sweep to about one and a half degrees of error — and why open-loop demonstrations cap how good it can get.

Part VII — Beyond imitation

Three tracks around the humanoid spine.

Chapter Hook
Chapter 17 — A Second Opinion Why the project runs more than one model: a second VLA for robot arms that speaks deltas, a third for other humanoids that speaks a fifty-five-number canonical action — and the embodiment gate that decides what runs today versus after a fine-tune.
Chapter 18 — When Copying Isn’t Enough Imitation cannot teach balance. A position-controlled biped falls over in physics without a controller it had to learn — the reinforcement-learning track, and the sim-to-sim transfer that carries a policy from one physics engine to another.
Chapter 19 — The Classical Counterpoint A reminder that not everything should be learned. A classical motion planner solves for a smooth, jerk-limited trajectory in closed form — and cuts peak command jerk by a factor of three hundred thousand — where a learned policy would only approximate it.

Part VIII — Synthesis

Chapter Hook
Chapter 20 — The Pipeline Is the Product Strip nineteen chapters to one sentence: the model is the easy part, one action format is what makes the loop possible, and the loop — not the model — is the thing you build.

Appendix

Reference material for readers who want the exact numbers.

   
Methods and protocol reference One lookup surface for the whole toolkit: the WebSocket protocol message-by-message, the port map, the embodiment profiles, every server flag, the frame and joint conventions, the recording format, and the known gotchas.

Figures

Diagrams are drawn inline. Screenshots, Rerun captures, and demonstration clips are marked in the chapter source with a hidden 📷 Figure — placeholder comment that names the exact scene or command that produces them; drop the captured media in under assets/ and replace the placeholder. Everything the book describes is runnable from the main project’s code (private repository; available on request) — its main README and TESTING.md have the per-robot recipes.