Chapter 05 — The Wire Between Them
This chapter closes Part II. You have the body (Chapter 3) and the brain (Chapter 4); this is the wire that connects them — a single message format that carries the whole project’s most important idea. It assumes you know what an observation, an action, and an action chunk are.
One connection, small messages, both directions
The body is in Unity on your laptop. The brain is a model on a GPU box. Between them runs a WebSocket — a kind of network connection that stays open and lets both sides send messages whenever they like, in either direction. It is the same technology a chat app uses to keep a live connection to its server. We chose it for three plain reasons: it is bidirectional (Unity asks, the model answers, on one persistent connection), it carries text (so the messages are human-readable and easy to debug), and it is spoken by every language (C# on the Unity side, Python on the server side, with no special glue).
The messages themselves are JSON — the ubiquitous text format of {"key": value} pairs. Each message has a "type" field that says what it is, and the rest of the fields depend on the type. There are only three types you need, and two of them are trivial:
{"type": "ping"}→ the server replies{"type": "pong"}. A heartbeat, to check the connection is alive.{"type": "reset"}→ the server replies{"type": "reset_ok"}. Start a fresh episode: clear the model’s internal state and the visualization’s history.{"type": "act", ...}→ the real one. Here is an observation; give me an action.
Everything interesting is in the act message and its reply.
The act round-trip
Here is a real exchange for the humanoid track. Unity assembles an observation and sends:
{
"type": "act",
"image": "<base64-encoded JPEG, ~224×224>",
"instruction": "push the objects off the table",
"state": [ -0.10, 0.03, ..., 0.41 ],
"timestep": 42
}
The image is the robot’s camera view, JPEG-compressed and encoded as text so it fits in a JSON string. The instruction is the plain-English task. The state is the robot’s proprioception — here, its 29 joint angles in radians. The timestep is a counter so the reply can be matched to the request.
The server runs the model — a fraction of a second of GPU work — and replies:
{
"type": "action_chunk",
"actions_flat": [ /* 40 × 29 = 1160 numbers, row-major */ ],
"horizon": 40,
"dim": 29,
"chunk_hz": 50.0,
"timestep": 42,
"inference_ms": 619.2
}
This is a chunk: horizon × dim numbers, laid out one action after another. horizon: 40 says there are forty future actions; dim: 29 says each is a full set of 29 joint targets; chunk_hz: 50.0 says play them back fifty per second; inference_ms reports how long the model took. Unity reads actions_flat in blocks of 29 — action k is the slice from k × 29 to (k+1) × 29 — and feeds them to the joints in order. Chapter 8 is about why the humanoid gets a whole chunk at once, and Chapter 11 is about the day chunk_hz lied.
The arm track uses the exact same act request (with proprio instead of state) and a simpler reply — a single action rather than a chunk:
{ "type": "action", "action": [dx, dy, dz, droll, dpitch, dyaw, gripper],
"timestep": 7, "inference_ms": 210.5 }
sequenceDiagram
participant U as Unity (body)
participant S as Server (brain)
loop every observation
U->>S: {"type":"act", image, instruction, state, timestep}
Note right of S: run the model (inference)
S->>U: {"type":"action_chunk", actions_flat, horizon, dim, chunk_hz, ...}
Note left of U: play the chunk on the joints
end
If something goes wrong on the server, it replies {"type": "error", "message": "..."} and keeps the connection open — one bad observation does not kill the session. That small robustness matters when you are iterating live.
The one idea this chapter exists for
Look again at what the server actually has to be. It must accept an act message, and it must return an action or action_chunk. That is the entire contract. Nothing in it says how the action was produced. And that is the hinge the whole project turns on:
Insight: the wire hides what is on the other end. Unity sends an observation and receives an action. It cannot tell — and does not care — whether the action came from a 7-billion-parameter model on a supercomputer, a five-line scripted mock on the same laptop, a classical motion planner solving equations, or a human moving their hands in front of a webcam. Anything that speaks the
actprotocol is a valid brain. One wire, many speakers.
This is not a hypothetical. In this project, five different programs answer the humanoid’s port, 8766, with byte-for-byte the same protocol:
- the real GR00T model on the GPU box (Chapter 10),
- a scripted mock that emits a gentle scripted sway, so the whole loop can be tested with no GPU (Chapter 9),
- a teleoperation server that turns your tracked hands into joint targets (Chapter 14),
- a classical motion planner that streams a computed trajectory (Chapter 19),
- and, on a neighboring port, the fine-tuned version of GR00T serving your own trained policy (Chapter 16).
Unity’s humanoid code is identical for all five. Switching between them is a matter of which address it dials — and in the editor, that is literally a menu item. The mock and the real model differ by a flag.
Insight: uniformity upstream is what makes the pipeline possible. Because every driver speaks one format, every driver produces the same kind of data, and the recorder (Chapter 13) can capture a demonstration without knowing or caring who drove. This is the first concrete payoff of “one format, many drivers,” and it is why the wire — not the model — is the load-bearing piece of the architecture.
The port map, and where the box lives
Each track gets its own port, so several can run at once (with one exception, noted):
| Port | Speaks | Style |
|---|---|---|
| 8765 | OpenVLA (arm) | single action |
| 8766 | GR00T (humanoid) — also the mock, teleop, and motion planner | chunk |
| 8767 | LingBot (humanoid) | chunk |
| 8769 | reinforcement-learning locomotion policy | single action, per tick |
Port 8766 is shared by everything that speaks the humanoid chunk protocol, which is exactly the point — they are interchangeable, so they use the same door (and only one runs at a time).
Where is the GPU box? Two arrangements, and Unity’s code is blind to the difference:
- The owned DGX Spark sits on a private network (Tailscale), so Unity dials it by name —
ws://chaotic-spark:8766— with no tunnel. - A rented cloud GPU is reached through an SSH tunnel that maps the remote port onto
localhost, so Unity dialsws://localhost:8766and the tunnel forwards it.
The servers have no authentication — they will hand robot actions to anyone who connects — so they are bound to localhost or a private network and never exposed to the open internet. That is a safety rule, not a convenience: a port that returns robot commands is not something to leave open.
What you now understand
- The body and brain talk over a WebSocket carrying JSON text messages. There are three message types:
ping/pong,reset/reset_ok, and the real one,act. - An
actrequest carries the image, instruction, proprioception (proprioorstate), and a timestep. The reply is a singleaction(arm) or anaction_chunk(humanoid) withhorizon,dim, andchunk_hz. Errors reply without dropping the connection. - The protocol says nothing about how the action was produced — so anything that speaks it is a valid brain. In this project the same humanoid port is answered by the real model, a mock, a teleop server, a motion planner, and a fine-tuned policy, with identical Unity code. One wire, many speakers.
- Ports separate the tracks (8765 arm, 8766 humanoid, 8767 LingBot, 8769 locomotion). The GPU box is reached by private-network name or through an SSH tunnel; the servers are never exposed to the internet.
You now understand the three parts — body, brain, wire — and the idea that unifies them. Part III goes one level deeper, into the conventions the messages assume: coordinate frames, joint order, and the delta-versus-chunk distinction. We start where the most bugs hide: the fact that Unity and the robot disagree about which way is left.
Continue to Chapter 06 — Left, Right, and Why They Disagree.
The protocol is implemented in the Python servers (server/inference_server.py, server/groot_server.py, server/lingbot_server.py) and consumed by the Unity bridges (VLABridge.cs, GrootBridge.cs). JSON frames are capped at 16 MB — comfortably above a 224×224 JPEG. The full message-by-message specification is in the methods reference.