Skip to the content.

Chapter 17 — A Second Opinion

Part VII steps out from the humanoid VLA spine to the three tracks that surround it. This chapter is the other two VLAs — a model for robot arms and a model for other humanoids — and the recurring question of what runs today versus after a fine-tune. It assumes you know the delta-vs-chunk distinction, the embodiment tiers, and that anything speaking the protocol is a valid brain.


Why more than one model

The book so far has driven one model, GR00T, on one family of robots. But the project deliberately runs three VLAs, and the reason is not indecision — it is that no single model is best at everything, and running more than one is itself the strongest possible demonstration of the thesis. If three completely different models, built by three different teams, with three different action styles, all plug into the same Unity body with the same recorder and the same loop, then the interface really is the product.

So this chapter is a tour of the other two brains — what they are, what robots they drive, and the one thing they all share.


OpenVLA: the arm’s second opinion

The first is OpenVLA-7B, and it drives robot arms — a Franka Panda or a WidowX — rather than humanoids. It is the model behind the arm track that has quietly been the counter-example throughout the book: the one that speaks deltas, not chunks.

Recall from Chapter 8: OpenVLA returns a single 7-number end-effector delta — move the gripper a little, rotate a little, open or close — about five times a second, and Unity turns that into joint motion with inverse kinematics. Internally it works unlike the humanoid models: it predicts its action the way a language model finishes a sentence, emitting the action as a short sequence of discrete tokens (OpenVLA is built on a vision-language language model). You do not need that detail to use it — what matters at the interface is the shape: one small delta, at 5 Hz, over the same act protocol, on port 8765.

OpenVLA is the project’s arm-manipulation option, and it fine-tunes on your data too — through the RLDS ecosystem (Chapter 15) rather than LeRobot. Same loop, different action dialect, different dataset format, same recorder feeding both.


LingBot: a third dialect behind the same door

The second is LingBot-VLA-V2, another humanoid model, built on a Qwen3-VL vision-language core with a mixture-of-experts flow-matching action head. It drives other humanoids — the G1, the AgiBot A2, the Fourier GR-2 — and it speaks yet a third action dialect: a 55-number canonical action that spans arm, end-effector, gripper, waist, head, base, and hand channels (notably, no legs — it is an upper-body manipulation model).

LingBot also introduces a wrinkle worth seeing, because it shows the bridge pattern doing real work. LingBot is served not over our JSON protocol but over openpi, a different serving stack that speaks WebSocket with msgpack (a binary message format) on its own port. So the LingBot bridge (lingbot_server.py, port 8767) sits in the middle and translates: Unity sends the ordinary JSON act message, the bridge repackages it into openpi’s msgpack format, forwards it to the real model process, and repackages the reply back into our action_chunk. From Unity’s side, it is identical to GR00T — same request, same chunk reply — so the Unity humanoid code is unchanged.

Insight: the bridge is where a foreign model becomes a native one. LingBot’s real serving stack is nothing like ours. Yet Unity talks to it with the exact code that talks to GR00T, because a thin bridge absorbs the difference — JSON in, msgpack out, and back. This is the same move the GR00T bridge made for nested observations (Chapter 10) and the teleop server made for hand tracking (Chapter 14). The pattern is the point: put the model-specific weirdness in a bridge that speaks our protocol on one side, so the body never has to learn a third language. A new model is a new bridge, not a new Unity.


The embodiment gate

The theme that has recurred since Chapter 10 — what does the model already know about this robot? — returns for LingBot with a twist. The G1, GR-2, and A2 are all among the twenty robots LingBot was pretrained on, so there is no fundamental barrier to driving them. But “pretrained on” is not the same as “runs out of the box,” because what actually gates a robot is a shipped robot config plus normalization statistics — and only a couple of robots ship those complete (the ones from LingBot’s own benchmark, like an Aloha-style dual-arm setup).

So for the Unity G1, LingBot needs a config and norm-stats before it drives well — and, echoing Chapter 15, those can be computed over a small recording through the same LeRobot conversion path, without a full fine-tune. The pattern generalizes across all three models:

Model Robots What runs today What a new robot needs
GR00T G1 (+Dex3), GR-1, H1 G1 zero-shot (REAL_G1) fine-tune for others; tier sets the cost
OpenVLA Panda, WidowX arms not yet run here — only a mock stub has driven the Unity Panda (zero-shot on its pretraining arms is expected, untested) RLDS fine-tune for your task
LingBot G1, AgiBot A2, GR-2 robots with shipped config + norm-stats (e.g. Aloha dual-arm) a config + norm-stats (small recording), or a fine-tune

Insight: “supported” is a spectrum, and the config is the gate. Across all three models, whether a robot runs is rarely a binary. It is a spectrum: fully zero-shot, needs-just-norm-stats, needs-a-fine-tune, needs-a-from-scratch-embodiment. And the thing that moves a robot along that spectrum is almost always a piece of description — a modality config, a set of normalization statistics — not new model weights. Knowing that reframes “does model X support robot Y?” into the more useful “what description does model X need before it drives robot Y, and can I produce it from a small recording?”


One interface, a family of brains

Look at what these three models have in common, because it is almost nothing except the one thing that matters. OpenVLA is a 7-billion-parameter token-predicting arm model. GR00T is a 3-billion-parameter diffusion humanoid model. LingBot is a mixture-of-experts flow-matching model served over a foreign stack. They drive different robots, output different action shapes, and train on different dataset formats. And every one of them plugs into the same Unity body, is captured by the same recorder, and is swapped by changing an address.

Insight: a uniform interface turns models into interchangeable parts. Because the wire hides what is behind it (Chapter 5), the three models are not three integrations — they are three speakers on one interface, each behind a small bridge. That is what makes a “model marketplace” practical: when a better VLA ships next month, adopting it is writing a bridge, not rebuilding the stack. The variety of the brains is exactly what makes the uniformity of the interface valuable — which is the whole reason the project runs more than one.


What you now understand

Imitation, in all three flavors, has a hard limit: it can only teach what can be demonstrated. The next chapter is about the one thing that cannot — balance — and the reinforcement-learning track that has to learn it instead.

Continue to Chapter 18 — When Copying Isn’t Enough.


The three VLA servers are server/inference_server.py (OpenVLA, port 8765), server/groot_server.py (GR00T, 8766), and server/lingbot_server.py (LingBot, 8767 — a bridge to an openpi deploy server on 8010). LingBot is Robbyant/lingbot-vla-v2 (Qwen3-VL-4B + MoE flow-matching); its per-robot config and norm-stats path is documented in the repository’s CLAUDE.md and scripts/spark/lingbot/.