Skip to the content.

Chapter 03 — Unity as the Body

Part II opens here. Chapters 1 and 2 gave you the shape of the project and its vocabulary. Now we look at the three parts one at a time. This chapter is the body: the simulated robot living inside Unity. It assumes you know what an embodiment, a degree of freedom, and proprioception are.


A robot you can crash for free

The body in this project is not a physical robot. It is a simulated one, living inside Unity — a game engine, the same category of software used to build video games, here used as a physics simulator.

Why simulate at all? Because a real robot is a bad place to make mistakes. A Unitree G1 humanoid costs tens of thousands of dollars. Every fall risks a motor or a sensor. It runs in real time — one second of practice takes one second — and repairs take days. A learning loop that involves the robot doing the wrong thing thousands of times is simply not viable on hardware.

A simulator removes every one of those constraints. Inside Unity there is a digital G1 with the same joints, the same proportions, and a physics engine applying the same gravity and contact forces a real one would feel. It falls: nothing breaks, nothing costs money, and you reset it with a keystroke. It runs on a laptop. And crucially, because the body is software, you can swap what is driving it — a mock, a model, your own hands — without touching the robot at all.

Insight: the simulator is the training ground, not the destination. The goal is always a policy that works on a real robot. We simulate because trial-and-error on hardware is too slow and too expensive. This creates the sim2real gap: a behavior that works in Unity may not transfer perfectly to a physical robot, because the simulation’s physics is only an approximation. Everything in this book lives on the simulation side of that gap — but every convention is chosen so that nothing has to be rebuilt when hardware arrives.


From a description to a puppet: URDF and ArticulationBody

A robot in Unity does not start as a Unity object. It starts as a URDF — a Unified Robot Description Format file, the standard way the robotics world writes down a robot’s body. A URDF is a text file listing every link (a rigid piece — an upper arm, a forearm, a hand) and every joint connecting two links (a hinge, a rotation), with the exact dimensions, axes, and limits of each. It is the robot’s blueprint, and the same blueprint is used everywhere: in Unity, in the Python servers, and in the visualization tools.

Unity imports that blueprint into a tree of ArticulationBody components — Unity’s built-in representation of a physically-simulated jointed mechanism. Each joint in the URDF becomes an ArticulationBody that PhysX (Unity’s physics engine) knows how to simulate: it has mass, it responds to forces, and it can be driven toward a target angle. The result is an articulated puppet: a hierarchy of rigid bodies connected by motors, obeying physics.

One detail matters enough to state now, because the whole codebase depends on it: joints are identified by their URDF name, not by their position in the Unity scene. When the scene-setup code wires up the G1’s 29 joints, it looks each one up by its blueprint name — left_shoulder_pitch, waist_yaw, and so on — never by where it happens to sit in the object tree. This is what lets the Python server and the Unity body agree, exactly, on which number drives which joint. (Chapter 7 is entirely about that agreement.)

(One sharp edge, recorded here so it does not bite you: on Apple Silicon Macs the scripted importer must use Unity’s built-in convex decomposer, not the package’s V-HACD option — the V-HACD library ships only an Intel binary and throws on an arm64 Mac, silently producing a robot with no collision shapes.)


Driving a joint: targets, stiffness, and the degrees-versus-radians trap

An ArticulationBody joint is driven the way a real robot’s joint is: not by teleporting it to an angle, but by giving its motor a target and letting a controller push it there. Three numbers govern that push, set once when the scene is built:

Together these make a joint that tracks its target firmly but not instantly — like a real geared motor. Set the target to a new angle and the joint swings there over a few physics steps.

Now the trap. Unity’s physics runs on a fixed timestep of 0.005 seconds — two hundred physics updates per second — set explicitly by every scene-setup routine, because a robot simulated at a jittery or slow rate behaves badly. But the more painful subtlety is units. Unity’s joint drive target is expressed in degrees. The robot’s actual measured joint angle is read back in radians. The two ends of the same joint speak different angular units:

// Writing a target: convert radians → degrees
float deg = targetRadians * Mathf.Rad2Deg;
drive.target = deg;              // xDrive.target is DEGREES

// Reading the current angle: it is already radians
float actual = joint.jointPosition[0];   // RADIANS

Everything the models produce and consume is in radians (and meters), matching the robotics convention. So the Unity side is dotted with Rad2Deg on the way out to a drive and nothing on the way in from a sensor. Forget the conversion in one place and a joint that should move a tenth of a radian instead tries to move a tenth of a degree — or, worse, treats 0.4 radians as 0.4 degrees and barely twitches. It is the single most common unit bug in the project, and now you will recognize it.


The pinned base: why the humanoid is bolted down

Watch the G1 in this project and you will notice its feet never move. The legs are held rigid, the pelvis is fixed in the air, and only the arms and waist act. This is deliberate, and the reason is honest: a position-controlled biped will not balance in PhysX without a controller it had to learn.

Standing on two legs is a continuous, split-second act of balance — a real humanoid runs a whole-body controller adjusting every joint hundreds of times a second just to stay upright. Our manipulation robots do not have one. If you let the legs go free and the pelvis fall under gravity, the G1 simply topples. So for every manipulation scene, the base is pinned — the pelvis is marked immovable, bolting the robot in place — and a control mask holds the twelve leg joints at a fixed standing pose while leaving the seventeen upper-body joints (waist and both arms) free for the policy to drive.

Insight: pin what you are not solving. Arms-only-with-a-pinned-base is not a limitation to apologize for — it is the correct way to study manipulation without also having to solve balance. The two problems are genuinely separate. Manipulation is what the vision-language-action models are for; balance is a reinforcement-learning problem in its own right, and it gets its own track (Chapter 18) where the base is finally un-pinned and the robot has to hold itself up. Until then, pinning the base is how we keep one hard problem from contaminating another.

The control mask is also how teleoperation works: to let a human drive only the fingers, or only the arms, the scene simply flips which joints the mask leaves free. Same body, different subset under control. (Chapter 14.)


The camera is part of the body

One last piece of the body is easy to forget: the camera. The model’s entire visual input is a picture, and that picture comes from a camera mounted in the Unity scene — on the robot’s head for a humanoid, over the shoulder for the arm. Each step, Unity renders that camera’s view to a small texture (224×224 pixels, the size the models were trained on), compresses it to a JPEG, and packs it into the observation.

That rendering has a cost — reading pixels back from the GPU briefly stalls it — but at five observations per second it is comfortably affordable. The camera is as much part of the embodiment as the joints: change where it points, and you change what the model can know.


What you now understand

That is the body. The next chapter is the brain: what a vision-language-action model actually is, and why it lives on a different machine.

Continue to Chapter 04 — The Model as the Brain.


The Unity project is DevVLA/ in the repository — Unity 6, URP, the new Input System. Scenes are built by the VLA/… editor menu (Assets/Editor/VLASceneSetup.cs); the joint-driving and base-pinning logic lives in HumanoidJointController.cs and RobotController.cs. Drive gains 10000/100/1000 and the 0.005 s timestep are set by the scene-setup routines.