Skip to the content.

Chapter 18 — When Copying Isn’t Enough

The whole book so far has taught robots by showing them. This chapter is about the one thing you cannot show: balance. It assumes you know why the humanoid’s base is pinned, the difference between open-loop chunks and closed-loop control, and what a policy is.


The one thing you cannot demonstrate

Every humanoid in this book has stood with its base bolted in place and its legs frozen (Chapter 3). That was not laziness; it was honesty. Un-pin the base, let the legs go free, and the G1 falls over instantly, because a position-controlled biped cannot balance in physics without a controller it had to learn.

Balance is not a pose you can record. Standing upright is a continuous, split-second act of correction — the ankle rolls a degree to catch a lean, the hip counter-rotates, the whole body negotiates with gravity a hundred times a second. There is no fixed sequence of joint targets that is balance; the right target at each instant depends on exactly how the body is tipping right now. So you cannot demonstrate balance the way you demonstrate a push — there is nothing static to copy. Imitation, the engine of the whole VLA half of the book, has met the one skill it structurally cannot teach.

Insight: imitation teaches what to do; it cannot teach reactions you never recorded. A demonstration is a fixed trajectory: this observation, that action. It works beautifully for tasks whose correct motion is roughly repeatable — reaching, pushing, pouring. It fails for tasks that are pure reaction to disturbance, because the disturbances in the demonstration are not the disturbances at run time. Balance is the purest example: the whole skill is reacting to a lean you have not seen yet. To get a reactive skill, you cannot copy a trajectory — you must learn a policy that responds to states, through trial and error.

That is a different paradigm entirely, and it is the subject of this project’s reinforcement-learning track.


Learning balance the hard way

Reinforcement learning (RL) is the from-scratch-learned approach from Chapter 1: no demonstrations, no reference motion, just a reward and millions of practice attempts. The robot tries something, gets a score for how upright and on-command it stayed, and gradually — over billions of simulated steps across thousands of parallel robots — discovers a policy that balances and walks. Nobody demonstrates the gait; it emerges from chasing the reward. (This is exactly the subject of the companion book, Teaching a Humanoid to Move — an entire volume on getting this right, and a good next read if this chapter whets the appetite.)

The training itself happens in a specialized simulator — Isaac Lab, built for massively parallel RL — where thousands of G1s practice at once on a GPU. Out of that training comes a policy file (exported to a portable format, ONNX) that maps the robot’s current state to joint targets. And now the interesting problem for this project begins: that policy was trained in Isaac Lab’s physics, but our robot lives in Unity’s PhysX. Getting one to drive the other is a transfer.


Sim-to-sim: from one physics engine to another

We have talked about the sim2real gap — sim to hardware. This is its cousin:

Sim-to-sim — transferring a policy trained in one simulator to a different simulator with different physics. Isaac Lab (where the policy learned) and Unity’s PhysX (where our robot lives) do not compute contacts, friction, and joint dynamics identically. A policy that balances perfectly in one may wobble in the other. Sim-to-sim is a rehearsal for sim-to-real: if a policy survives the jump between two simulators, it is a better bet for surviving the jump to hardware — and if it doesn’t, you learn why somewhere cheap.

Making the transfer work means matching the two worlds as closely as possible. The Unity RL scene un-pins the base (the whole point is to make the robot hold itself up) and sets the joint drive gains to match the actuators the policy trained against — much softer than the stiff manipulation gains (legs around 150–200 stiffness rather than 10,000), because a balance policy expects compliant, real-motor-like joints, not rigid servos. Get the gains wrong and the transfer fails even if the policy is perfect.


A different wire: per-tick control

Here is where the RL track breaks, deliberately, from everything else in the book. The humanoid VLAs are chunked: predict forty actions, play them open-loop, refill. That is fine for a push. It is fatal for balance.

Insight: match the control paradigm to the problem — open-loop chunks cannot stabilize a biped. A chunk is open-loop: the robot plays a second of pre-computed motion without looking. A falling robot cannot afford to stop looking for even a fraction of a second — by the time a stale chunk finishes, the lean it should have caught has become a fall. Balance demands closed-loop control: observe, act, observe, act, every single tick, so each action responds to the freshest state. So the RL track does not use the chunked protocol at all. It uses a per-tick protocol on its own port (8769): one observation in, one action out, fifty times a second, no buffer, no horizon.

The observation and action are shaped for control, not for language. Each tick, Unity sends the robot’s base linear and angular velocity, the direction of gravity in the body frame, the velocity command (where you want it to walk, from the keyboard), and every joint’s position and velocity. The policy returns one action, which is added to the robot’s default standing pose to get the joint targets (target = default_pose + 0.5 · action). There is no image and no instruction — this is not a VLA, it is a balance controller. The 50 Hz rate is not a coincidence; it is Unity’s physics cadence (a 0.005-second step, acted on every fourth tick), matched exactly to the rate the policy trained at.


Two honest frictions

Two real details make the transfer concrete, and both are the “conventions” lesson of Part III returning in a new setting:

The current status is honest: the server is built and tested in the sandbox, the Unity side is written but not yet compiled, and the success bar — a G1 that stands, then walks under keyboard command, in Unity physics with no pin — is the next thing to verify. It is a track in progress, and the sim-to-sim gap study is itself the deliverable.


What you now understand

Balance is what imitation cannot buy. The last track in Part VII is the opposite lesson: a skill you should not pay a model to learn at all, because it can be solved.

Continue to Chapter 19 — The Classical Counterpoint.


The RL track is scripts/spark/rl/ (Isaac Lab setup, PPO training of Isaac-Velocity-Flat-G1-v0, ONNX export) and server/rl_policy_server.py (the per-tick server on port 8769, with --policy stand and --policy onnx). The Unity side is RlBridge.cs and the VLA/Setup RL Locomotion Scene menu. The 123-number observation, the target = default + 0.5·action rule, and the 37→29 alias map are documented in the repository’s docs/humanoid-rl-roadmap.md.