Skip to the content.

Chapter 08 — Deltas, Targets, and the Chunk

This chapter closes Part III. The two robot tracks speak different action dialects, and the difference explains a large fraction of the codebase. It assumes you know what an action, an end-effector, and an action chunk are, and that actions cross the wire as flat lists.


Two dialects

Both robot tracks receive actions over the same wire, but the meaning of the numbers is fundamentally different. The arm is told to move a little — a change relative to where it is now. The humanoid is told to be exactly here — a full set of target joint angles. One speaks in deltas, the other in absolute targets, and the choice is not arbitrary: it follows from what each model produces and what each robot needs.

Delta action — an action expressed as a change from the current state: move 2 cm forward, rotate 5° clockwise. The robot adds the delta to where it is. Deltas are relative and self-correcting in a rough way (a small error just means the next delta starts from a slightly different place), but they accumulate: applied blindly forever, small errors drift.

Absolute target action — an action expressed as a destination: put joint 12 at 0.41 radians. The robot moves there regardless of where it was. Absolute targets are exact and do not drift, but they demand that the model know the whole pose it wants, every time.


The arm speaks deltas

The arm model, OpenVLA, returns a single 7-number action:

[ dx, dy, dz, droll, dpitch, dyaw, gripper ]

The first six are a delta on the end-effector — move the gripper dx meters forward, dy left, dz up, and rotate it by the three small angles — expressed in the robot’s base frame (which is why Chapter 6’s frame conversion runs on every one of them). The seventh is the gripper: roughly 1.0 for open, 0.0 for closed. There are no joint angles here at all; the model reasons about where the hand should go, not how the elbow should bend.

Unity turns that end-effector delta into joint motion with inverse kinematics — the geometry of “given where I want the hand, what angles put it there.” It adds the delta to a running target pose and solves for the joint angles each physics step:

targetPos = targetPos + baseRotation * (dx, dy, dz)   // accumulate the nudge
targetRot = (small rotation) * targetRot
// then solve the arm's joints to reach targetPos/targetRot

This is why deltas suit the arm. Manipulation is naturally described in end-effector space — reach toward the cup — and small relative nudges are forgiving: if the hand ends up a centimeter off, the next observation shows that, and the next delta corrects from there. The model runs about five times a second, and between updates the last action simply holds. Five corrections a second is plenty for a reaching arm.


The humanoid speaks absolute targets

The humanoid model, GR00T, returns something quite different: a full vector of 29 absolute joint angles, the exact pose it wants the whole upper body in. Not a nudge — a destination. And it does not return one such pose; it returns forty of them at once, a chunk.

Why absolute rather than delta? Because a humanoid’s motion is a whole-body coordination problem, not a single end-effector reach. The model has been trained to output complete target poses, and absolute targets do not accumulate drift — each one stands on its own, so a chunk of forty can be played back open-loop (without a fresh observation between them) and still land where intended. That last property is essential, because of the rate problem.


The rate problem, and why chunks exist

Here is the tension the chunk resolves. The humanoid’s joints want fresh commands fifty times a second — that is the rate at which a smooth 50 Hz control loop runs. But the model takes roughly half a second to produce an answer (inference_ms ≈ 600). If Unity asked the model for one action and waited, the robot would get a new command about twice a second and lurch between long-held poses. Fifty-per-second control from a twice-per-second brain is impossible if each request yields one action.

The chunk breaks the deadlock:

Insight: a chunk decouples the thinking rate from the acting rate. Instead of one action, the model returns forty — nearly a full second of motion — in a single inference. Unity plays those forty back smoothly at 50 Hz while the model, in parallel, thinks about the next chunk. The robot acts at 50 Hz; the brain thinks at ~2 Hz; the chunk is the buffer between the two rates. This only works because the actions are absolute targets that can be played open-loop — a chunk of drifting deltas would wander before the next thought arrived.

model thinks:   [====chunk A (0.8s)====]      [====chunk B====]
robot plays:    a0 a1 a2 ... a39  ┃ b0 b1 b2 ...
                └ 50 Hz playback ──┘└ next chunk seamlessly continues

The numbers have to line up for this to be seamless, and they do: forty actions at 50 Hz is 0.8 seconds of motion, comfortably longer than the ~0.6 seconds the model takes to produce the next chunk. There is always motion queued to play while the next batch is computed. Chapter 11 is the story of the day those numbers did not line up.


The buffer that makes it smooth

The piece that plays a chunk back at the right times is a small timestamped buffer. When a chunk arrives, each of its forty actions is stamped with the moment it should execute — action k at start + k / chunk_hz — and dropped into a queue. Fifty times a second, Unity asks the buffer “what should I be doing now?” and it hands back the action whose time has come.

Two behaviors of that buffer are worth knowing, because they encode real judgment:

Together these give a robot that moves continuously and smoothly, driven by a brain that answers only twice a second — as long as the timing is honest. The buffer trusts the chunk_hz number in the reply to stamp the actions. If that number is wrong, everything downstream is wrong in a way that is invisible in the code and obvious on the robot. That is Chapter 11.


What you now understand

That completes the conventions: frames, joints, and dialects. You now know enough to watch the whole thing run. Part IV does exactly that — starting with the cheapest possible version, a scripted mock that proves the entire pipeline on a laptop with no GPU at all.

Continue to Chapter 09 — First Light.


The arm’s delta handling is in RobotController.cs (ApplyAction, with a CCD inverse-kinematics solver); the humanoid’s chunk handling is split between GrootBridge.cs (request pacing) and ChunkedActionBuffer.cs (timestamped playback, stale-chunk clearing). The ~0.35 s refill threshold and the 4 Hz request cap are GrootBridge fields; the 40-action horizon at 50 Hz gives 0.8 s of motion per chunk.