Chapter 11 — The Cadence Bug
This is the most instructive failure in the book, and like the best bugs it was a single wrong number. It assumes you understand the chunk and its buffer — how a chunk’s actions are stamped with times by chunk_hz, played back at that rate, and refilled when the buffer runs low. This chapter is that machinery meeting a lie.
The setup: a served model that moved wrong
By this point in the project we had done something new: fine-tuned GR00T on 80 of our own demonstrations of the push task and served the result back to Unity (that campaign is Chapters 15 and 16). The loop connected. The model produced chunks. And the robot moved — badly, in two distinct and confusing ways at once.
First, the motion was frantic. The arms snapped through the push far too fast, a blur where the demonstrations had been a deliberate sweep. Second, the motion stuttered: a quick burst, then a freeze, then another burst — the robot lurching rather than flowing. Fast and stuttering is a strange combination. A slow model might stutter; a jittery controller might race; but both symptoms together, from a model that had trained cleanly, pointed at something more specific than “the model is bad.”
It was not the model. It was one number in the reply: chunk_hz.
One number, two symptoms
Recall from Chapter 8 what chunk_hz does. When a chunk of actions arrives, the buffer stamps action k with the time start + k / chunk_hz and plays each one back when its moment comes. chunk_hz is the playback rate: how many of the chunk’s actions to execute per second.
The served model returned chunks of 16 actions, and the bridge advertised chunk_hz = 50 — the same 50 Hz that the base REAL_G1 model used. But this fine-tuned model had been trained on demonstrations recorded at 5 frames per second. Its 16 actions were meant to represent 16 ÷ 5 = 3.2 seconds of motion. Played back at the advertised 50 Hz, they represented 16 ÷ 50 = 0.32 seconds instead.
That one mismatch produced both symptoms, for two different reasons.
Why the motion raced. The model learned the push as a sequence of poses spaced a fifth of a second apart. Playing those same poses a fiftieth of a second apart compresses three-and-a-fifth seconds of intended motion into a third of a second — a tenfold speed-up. The push was correct in shape and ten times too fast in time. The robot was faithfully reproducing what it learned, at the wrong tempo.
Why the motion stuttered. Chapter 8’s arithmetic was the whole point of the chunk: a chunk must hold more seconds of motion than the model takes to compute the next one, or the buffer empties before the refill arrives. This model took about 0.62 seconds per inference. At the true rate, 16 actions at 5 Hz is 3.2 seconds of buffered motion — plenty of cushion. At the advertised rate, 16 actions at 50 Hz is 0.32 seconds — shorter than the 0.62 seconds the next chunk needs to arrive. The buffer drained to empty every single cycle and sat starved, waiting, before the next burst landed.
Advertised (wrong): 16 actions @ 50 Hz = 0.32 s of motion
inference = 0.62 s
→ buffer empties (0.32) long before refill (0.62): STARVE + 10× fast
Native (correct): 16 actions @ 5 Hz = 3.20 s of motion
inference = 0.62 s
→ 3.2 s cushion over a 0.62 s refill: smooth, correct tempo
One wrong scalar, and the same 16 numbers were simultaneously played too fast and run out of before the next batch arrived. The code was correct. The buffer did exactly what it was told. It was told a lie.
The number was a fact about the data, wearing the disguise of a setting
Here is the deeper reason this bug was possible. chunk_hz looks like a configuration knob — a rate you pick. It is not. It is a fact about the data the model was trained on. A model trained on 5-frames-per-second demonstrations produces actions spaced a fifth of a second apart, and there is exactly one correct playback rate: 5 Hz. Any other value desynchronizes the model’s intended motion from the robot’s clock.
The base REAL_G1 model happened to be trained at 50 Hz, so 50 had been quietly baked in as “the” chunk rate. When a differently-trained model arrived, that baked-in default was silently wrong — not because anyone chose 50 for the new model, but because nobody had told the system that the new model’s native rate was different.
Insight: a number that encodes a fact about the data must travel with the data. The native playback rate is not the bridge’s to choose — it belongs to the checkpoint, because it was fixed the moment the training data’s frame rate was fixed. A value like this that lives as a global default instead of a property of the model is a desync waiting for the second model to arrive. The fix is to make it a fact the model carries, not a setting the server guesses.
The fix: let the model carry its own rate
The repair was small and exactly matched the lesson. Each embodiment profile now carries its native chunk rate as a property: the fine-tuned push profile declares chunk_hz = 5.0; the base REAL_G1 profile declares 50. When the server assembles a reply, it resolves the rate in order: an explicit command-line override if you gave one, else the profile’s own native rate, else the old default of 50. The number that describes the data now lives with the description of the data.
Re-served, the same model produced the same 16-action chunks — now advertised at 5 Hz, i.e. 3.2 seconds of motion each. The push played at the tempo it was demonstrated, and the buffer stayed comfortably full. Two symptoms, one number, gone.
Why this is an honest failure worth a chapter
Nothing here was a coding mistake in the usual sense. The buffer’s timestamp math was correct. The refill logic was correct. The model was correct — it had trained cleanly to a low error. The bug lived entirely in a mismatch of provenance: a rate that was true for one model, silently inherited by another for which it was false. That is a subtler and more common class of failure than a crash, and it is invisible in code review — every line is right — and glaringly obvious the instant you watch the robot.
It echoes a refrain this book keeps returning to: the numbers can lie. A chunk of 16 correct actions, timed by a wrong rate, is a correct action sequence that produces incorrect motion. No error was thrown. The only signal was a robot moving wrong, and the only way to catch it early is to look — which is exactly why the next chapter is about the tool that lets you look.
Insight: the seams between systems are where the honest bugs live. The model was right and the bridge was right; the bug was in the handshake between them — the shared assumption about how fast a chunk plays. As systems compose, more and more of the real failures move out of any single component and into the contracts between components. A number that both sides must agree on, where one side quietly assumes a default, is precisely such a seam. Distrust shared scalars; make them travel with the thing they describe.
What you now understand
- A served, cleanly-trained model moved frantically and in stutters at once — from a single wrong number,
chunk_hz. chunk_hzis the playback rate. The model was trained at 5 fps but the reply advertised 50 Hz. Playing 16 actions at 50 Hz instead of 5 made the motion 10× too fast (three seconds compressed into a third of a second) and starved the buffer (0.32 s of motion against a 0.62 s inference).chunk_hzonly looks like a setting; it is a fact about the training data’s frame rate, with exactly one correct value per model. Baked in as a global default, it was silently wrong for the second model.- The fix: each profile carries its own native rate (push = 5 Hz, REAL_G1 = 50 Hz), resolved before any default. The number that describes the data now travels with it.
- The failure was in the seam between a correct model and a correct bridge — the class of honest bug that no line-by-line code review catches, and that only watching the robot reveals.
To watch the robot — really watch it, every channel, every joint, the camera it saw and the path its hand took — you need an instrument panel. That is the next chapter.
Continue to Chapter 12 — Watching the Robot Think.
The native-rate fix is the chunk_hz field on Gr00tObsProfile in server/groot_server.py, resolved as args.chunk_hz or profile.chunk_hz or 50.0. The push profile (g1_push) sets chunk_hz = 5.0; the fine-tune campaign that produced the model is Chapters 15–16. The re-verified result on the wire: 16×29 chunks at 5 Hz — 3.2 seconds of motion each.