Skip to the content.

Chapter 09 — First Light

Part IV opens here: the system running for real. Before we put a real model on a real GPU, this chapter proves the entire pipeline with no GPU at all — a scripted stand-in that exercises every seam. It assumes you understand the wire, the chunk, and that anything speaking the protocol is a valid brain.


The cheapest possible robot brain

“First light” is what astronomers call the moment a new telescope first collects an image — not a scientific result, just proof the whole instrument works end to end. Before pointing an expensive telescope at a distant galaxy, you check that light gets from the mirror to the sensor at all.

This project’s first light is a mock policy: a few lines of arithmetic that speak the act protocol and return scripted actions. It knows nothing, sees nothing, learns nothing. Its entire job is to prove that an observation can leave Unity, cross the wire, come back as an action (or a chunk), drive the joints, and be recorded — with no GPU, no model download, and no cloud. Everything except the intelligence.

That “everything except the intelligence” is most of the system. Chapter 5 promised that anything speaking the protocol is a valid brain; the mock cashes that promise the moment you need it, on a laptop, for free.

Insight: prove the plumbing before you add the water. A real model introduces two things at once — a new brain and a live GPU box, network, and multi-gigabyte download. If the loop breaks, you cannot tell which half failed. The mock removes the model from the equation entirely: run it, and if the robot moves and the recorder writes files, then the observation-building, the transport, the chunk buffer, the joint control, and the recorder are all correct. When the real model arrives (next chapter), it is the only new variable. Debugging one new thing at a time is the whole reason the mock exists.


What the mock actually does

The mock is deliberately not random — random motion would be useless for spotting bugs. It produces a gentle, deterministic, recognizable pattern, so that if the robot does something else, you know a seam is broken.

For the humanoid, the mock reads the robot’s joint names from the URDF (so it automatically produces the right number of joints for whatever robot is loaded), then builds each chunk from a resting standing pose plus two scripted motions:

Each joint gets its own phase (derived from its name), so the motion looks organic rather than robotic-in-lockstep, and the phase carries across chunk boundaries so the playback is continuous — chunk B picks up exactly where chunk A left off. That continuity is itself a test: if the robot hitches at chunk boundaries, the buffer’s hand-off (Chapter 8) is wrong.

The arm mock is the same idea in seven numbers: small sinusoidal end-effector deltas — dx = 0.01·sin(0.15·t), and so on — with the gripper toggling open/closed every few seconds, so you can watch a slow reaching-and-grasping loop.

Because the mock returns raw joint targets (not the model’s structured action space), the Unity humanoid must be told to send flat measured joints and expect flat targets back — a single flag, useStructuredObs = OFF. This is the same flag the teleop server and motion planner rely on, and it is why all three are drop-in interchangeable with the mock.


Running it end to end

The whole thing runs from the base environment, which is Mac-friendly by design — no GPU packages, just the pieces needed for mock and visualization:

cd server
uv sync                                        # base env: mock + Rerun, no GPU

uv run groot_server.py --backend mock --rerun  # humanoid: 40×29 scripted chunks + posed mesh
uv run test_client_groot.py                    # a fake Unity client: verify chunking, no editor

The test_client_groot.py step is worth pausing on: it is a fake Unity written in Python that connects to the server, sends synthetic observations, and checks the chunks come back well-formed. You can validate the entire server side — protocol, chunk shape, timing — before opening Unity at all. Then, in Unity, you point the humanoid scene at ws://localhost:8766, press Play, and watch the G1’s arms sway and its hands curl. The arm track is the same pattern on port 8765 with inference_server.py --mock.


What First Light proves — and what it doesn’t

When the mock loop runs cleanly, you have verified, concretely:

What it does not prove is the only thing left: that a real model produces useful actions. The mock’s motion is meaningless by construction. But that is precisely the point of doing it first — every part of the machine except the mind is now known-good, so when the mind arrives, it is the sole suspect if anything looks wrong.

Insight: the mock is not throwaway — it is permanent infrastructure. It is tempting to see the mock as scaffolding you delete once the real model works. It is the opposite. Every time you change the protocol, the buffer, the recorder, or a scene, the mock re-proves the plumbing in seconds on a laptop, with no GPU bill and no network. A mock that speaks the real protocol is the fastest test you own, and it stays useful for the life of the project.


What you now understand

The plumbing is proven. Now we add the water — a real 3-billion-parameter humanoid model on a real GPU box — and meet the first thing the mock could not warn us about: the model’s observation is not a flat list of numbers.

Continue to Chapter 10 — A Real Brain on a Real Box.


The humanoid mock is MockG1Policy in server/groot_server.py (--backend mock); the arm mock is MockPolicy in server/inference_server.py (--mock). The fake-Unity clients are server/test_client_groot.py and server/test_client.py. The base uv environment (websockets, numpy, pillow, Rerun, trimesh) runs the whole mock-and-visualize loop on a Mac.