Skip to the content.

Chapter 16 — The First Fine-Tune

This chapter closes the loop. The dataset is ready (Chapter 15); now we fine-tune the model on it and serve the result back to the same Unity robot. It assumes you know what fine-tuning is, what a checkpoint is, the new-embodiment tier, and the cadence lesson — which was learned serving this very model.


Eighty demonstrations

The task was the running example of this whole book: a 29-DoF Unitree G1, told to push the objects off the table. The dataset was 80 episodes of it — converted, as Chapter 15 described, into a LeRobot v2.1 dataset with a new-embodiment modality config and normalization stats computed over those eighty episodes.

One honest detail about how those demonstrations were produced, because it determines the ceiling of everything that follows: they were an open-loop scripted sweep — a programmed demonstrator that pushed through the motion without reacting to what it saw. Not teleoperation, not a reactive policy; a repeatable canned motion, recorded eighty times with the objects in varying positions. That was enough to prove the pipeline end to end, which was the goal of the first fine-tune. It is also, as we will see, exactly what limits how good the resulting policy can be.


What fine-tuning actually adjusts

You do not retrain a 3-billion-parameter model from scratch on eighty episodes — that would be hopeless, and pointless. Fine-tuning adjusts a small, targeted part of the model and leaves the rest frozen. For GR00T’s new-embodiment recipe, specifically:

Because the 29-DoF G1 is a new embodiment (Chapter 15), that projection layer starts from random values and learns, over the run, what this body’s numbers mean and how to turn the frozen backbone’s understanding into good pushes.

Insight: this is not LoRA — it is selective tuning of the parts that must change. A common way to fine-tune large models is LoRA, which threads small trainable adapters throughout a frozen network. GR00T’s new-embodiment recipe is different and simpler to reason about: freeze the giant general-purpose backbone entirely, and train only the two pieces that are specific to this robot and task — the input projector (which literally did not exist for this body before) and the action head. You are not adapting the model’s understanding of the world; you are teaching it to speak a new robot’s language and to produce a new task’s motion. Naming what is frozen and what is trained tells you exactly what fine-tuning can and cannot fix.


The overnight run

The training ran on the owned GPU box — the DGX Spark — entirely locally, no cloud, no bill. The numbers, for concreteness:

Training loss going down means the model is getting better at reproducing the demonstrations — predicting, for each recorded observation, the action the demonstrator took. That is necessary but, as this book keeps insisting, not sufficient: a low loss is a proxy for a good policy, not a proof. So before serving it to the robot, we checked it a different way.


Checking before you serve: offline evaluation

Rather than immediately wiring the fresh checkpoint into Unity, the project first evaluates it offline — no simulator, no live loop. It replays held-out demonstrations through the model and measures how far the model’s predicted joint targets fall from the demonstrated ones. On the push checkpoint:

An average of one-and-a-half degrees of error across the joints means the model learned to reproduce the demonstrated sweep closely. This is the right thing to check first because it is fast, deterministic, and isolates the model from every live-loop variable (frames, cadence, buffering). If the offline error had been large, there would be no point serving it; because it was small, serving was worth doing.


Serving it back

Then the loop closed. The fine-tuned checkpoint was served on the GPU box under the new-embodiment tag, with its own bridge profile, and Unity — in exactly the flat-observation mode the mock used (useStructuredObs = OFF, since this is a plain 29-number joint vector) — pointed at it and pressed Play. The G1 executed the learned push.

This is also, not coincidentally, the model whose cadence bug was Chapter 11: its profile declares a native rate of 5 Hz, because the demonstrations were recorded at 5 frames per second, and serving it at the wrong 50 Hz is what produced the racing, starving motion. Served at its true rate — 16-action chunks at 5 Hz, about 3.2 seconds of motion each, roughly 620 ms per inference — it plays the push at the tempo it was taught. The two chapters are one story: the first fine-tune and the cadence lesson came from the same served checkpoint.

The milestone is real and worth stating plainly: the first fine-tune in the project, run entirely on an owned box, closed the full loop — demonstrations recorded in Unity, converted, trained, evaluated, and served back to drive the same robot. Drive → record → fine-tune → serve, all the way around, for the first time.


The honest ceiling

And now the honest part, because this book does not sand its results smooth. The served policy reproduces the push — but it reproduces the push it was shown, and it was shown an open-loop scripted sweep. A model trained on demonstrations that never reacted to the scene cannot learn to react to the scene. It learned the shape of a good push, not the judgment of when to push harder, adjust for an object in a new spot, or recover from a miss. The demonstrations’ quality is the policy’s ceiling.

Insight: you get out what you put in. A fine-tuned policy is bounded by the demonstrations it learned from. Open-loop, non-reactive demonstrations produce an open-loop, non-reactive policy, no matter how low the training loss goes. To get a better policy, the lever is not more training steps or a bigger model — it is better data: reactive, closed-loop demonstrations, ideally teleoperated by a human (Chapter 14), and refined by correcting the policy’s own mistakes (the DAgger correction flag, Chapter 13). The pipeline is proven; raising the ceiling is a data problem, and the pipeline is exactly what makes collecting better data cheap.

That is not a disappointment — it is the roadmap. The whole apparatus exists so that recording better demonstrations is a matter of driving the robot differently, and everything downstream — convert, train, evaluate, serve — is already built and proven. The first fine-tune bought the machine. Better data is what you feed it next.


What you now understand

The humanoid VLA spine is now complete, end to end. Part VII steps out from that spine to the tracks around it: a second and third VLA for other robots, a reinforcement-learning track for the one thing imitation cannot teach, and a classical planner that is a reminder not everything should be learned.

Continue to Chapter 17 — A Second Opinion.


The fine-tune pipeline is scripts/spark/: convert_g1_push.sh, train_g1_push.sh (GR00T’s launch_finetune.py, backbone frozen / projector + diffusion tuned), eval_g1_push.sh (offline open-loop eval), and serve_finetuned_g1_push.sh. The evaluated/served result — train_loss 0.874, MAE ~0.024 rad — is at checkpoint-1000. The full campaign log is in the repository’s PROGRESS.md and trials/.