Chapter 09 — First Light
Part IV opens here: the system running for real. Before we put a real model on a real GPU, this chapter proves the entire pipeline with no GPU at all — a scripted stand-in that exercises every seam. It assumes you understand the wire, the chunk, and that anything speaking the protocol is a valid brain.
The cheapest possible robot brain
“First light” is what astronomers call the moment a new telescope first collects an image — not a scientific result, just proof the whole instrument works end to end. Before pointing an expensive telescope at a distant galaxy, you check that light gets from the mirror to the sensor at all.
This project’s first light is a mock policy: a few lines of arithmetic that speak the act protocol and return scripted actions. It knows nothing, sees nothing, learns nothing. Its entire job is to prove that an observation can leave Unity, cross the wire, come back as an action (or a chunk), drive the joints, and be recorded — with no GPU, no model download, and no cloud. Everything except the intelligence.
That “everything except the intelligence” is most of the system. Chapter 5 promised that anything speaking the protocol is a valid brain; the mock cashes that promise the moment you need it, on a laptop, for free.
Insight: prove the plumbing before you add the water. A real model introduces two things at once — a new brain and a live GPU box, network, and multi-gigabyte download. If the loop breaks, you cannot tell which half failed. The mock removes the model from the equation entirely: run it, and if the robot moves and the recorder writes files, then the observation-building, the transport, the chunk buffer, the joint control, and the recorder are all correct. When the real model arrives (next chapter), it is the only new variable. Debugging one new thing at a time is the whole reason the mock exists.
What the mock actually does
The mock is deliberately not random — random motion would be useless for spotting bugs. It produces a gentle, deterministic, recognizable pattern, so that if the robot does something else, you know a seam is broken.
For the humanoid, the mock reads the robot’s joint names from the URDF (so it automatically produces the right number of joints for whatever robot is loaded), then builds each chunk from a resting standing pose plus two scripted motions:
- arms sway — each arm joint follows a slow sine wave,
0.25 · sin(0.8·t + phase), so the arms rock smoothly; - hands curl — each finger joint follows
0.3 · (1 − cos(1.2·t + phase)), a one-sided curl from open toward closed.
Each joint gets its own phase (derived from its name), so the motion looks organic rather than robotic-in-lockstep, and the phase carries across chunk boundaries so the playback is continuous — chunk B picks up exactly where chunk A left off. That continuity is itself a test: if the robot hitches at chunk boundaries, the buffer’s hand-off (Chapter 8) is wrong.
The arm mock is the same idea in seven numbers: small sinusoidal end-effector deltas — dx = 0.01·sin(0.15·t), and so on — with the gripper toggling open/closed every few seconds, so you can watch a slow reaching-and-grasping loop.
Because the mock returns raw joint targets (not the model’s structured action space), the Unity humanoid must be told to send flat measured joints and expect flat targets back — a single flag, useStructuredObs = OFF. This is the same flag the teleop server and motion planner rely on, and it is why all three are drop-in interchangeable with the mock.
Running it end to end
The whole thing runs from the base environment, which is Mac-friendly by design — no GPU packages, just the pieces needed for mock and visualization:
cd server
uv sync # base env: mock + Rerun, no GPU
uv run groot_server.py --backend mock --rerun # humanoid: 40×29 scripted chunks + posed mesh
uv run test_client_groot.py # a fake Unity client: verify chunking, no editor
The test_client_groot.py step is worth pausing on: it is a fake Unity written in Python that connects to the server, sends synthetic observations, and checks the chunks come back well-formed. You can validate the entire server side — protocol, chunk shape, timing — before opening Unity at all. Then, in Unity, you point the humanoid scene at ws://localhost:8766, press Play, and watch the G1’s arms sway and its hands curl. The arm track is the same pattern on port 8765 with inference_server.py --mock.
What First Light proves — and what it doesn’t
When the mock loop runs cleanly, you have verified, concretely:
- Unity builds a valid observation (camera JPEG, proprioception, instruction, timestep) and sends it;
- the server parses it, produces a well-shaped chunk, and returns it inside the 16 MB frame limit;
- the timestamped buffer plays the chunk back smoothly at the advertised rate, seamlessly across chunk boundaries;
- the joint controller applies each target, in the right units (radians in, degrees to the drive), on the fixed timestep;
- the recorder writes an episode to disk in the training format.
What it does not prove is the only thing left: that a real model produces useful actions. The mock’s motion is meaningless by construction. But that is precisely the point of doing it first — every part of the machine except the mind is now known-good, so when the mind arrives, it is the sole suspect if anything looks wrong.
Insight: the mock is not throwaway — it is permanent infrastructure. It is tempting to see the mock as scaffolding you delete once the real model works. It is the opposite. Every time you change the protocol, the buffer, the recorder, or a scene, the mock re-proves the plumbing in seconds on a laptop, with no GPU bill and no network. A mock that speaks the real protocol is the fastest test you own, and it stays useful for the life of the project.
What you now understand
- First light is proving the whole instrument works end to end before chasing a real result. Here it is a mock policy — scripted arithmetic that speaks the
actprotocol and returns deterministic actions, with no GPU. - The mock produces recognizable motion (arms sway on a sine, hands curl, continuous across chunks) so that any deviation flags a broken seam. It returns raw joint targets, so Unity runs with
useStructuredObs = OFF— the same mode the teleop server and motion planner use. - It runs entirely in the Mac-friendly base environment; a Python fake-Unity client validates the server side before the editor is even opened.
- Running the mock loop proves everything except the model’s usefulness — observation-building, transport, chunk buffering, joint control, and recording are all confirmed. When the real model arrives it is the only new variable.
- The mock is permanent infrastructure, not scaffolding: the fastest, cheapest regression test in the project.
The plumbing is proven. Now we add the water — a real 3-billion-parameter humanoid model on a real GPU box — and meet the first thing the mock could not warn us about: the model’s observation is not a flat list of numbers.
Continue to Chapter 10 — A Real Brain on a Real Box.
The humanoid mock is MockG1Policy in server/groot_server.py (--backend mock); the arm mock is MockPolicy in server/inference_server.py (--mock). The fake-Unity clients are server/test_client_groot.py and server/test_client.py. The base uv environment (websockets, numpy, pillow, Rerun, trimesh) runs the whole mock-and-visualize loop on a Mac.