Skip to the content.

Chapter 12 — Watching the Robot Think

This chapter closes Part IV. The last three chapters kept ending on the same note — the bug is invisible in the numbers and obvious on the robot — so this one is about the instrument that lets you look. It assumes you have met the frame, joint-order, and cadence bugs, all of which share one property: nothing crashes.


The problem with invisible bugs

Every failure in Part IV was silent. A wrong coordinate frame produces a valid vector that means the wrong direction. A scrambled joint order produces a valid action that drives the wrong joint. A wrong chunk_hz produces a valid chunk that plays at the wrong tempo. None of them throws an error. None of them shows up in a log. Each one is a robot doing something almost right, and the only reliable way to catch it is to see the robot and its data at the same time.

Reading raw numbers off a terminal will not do it. A stream of 29 joint angles scrolling past tells you nothing about whether the elbow is where it should be. You need to see: the robot posed in 3D, the camera image it actually received, every action and proprioception channel plotted over time, and the path its hand traced through space — all synchronized to the same clock. That instrument is Rerun, a visualization tool built for exactly this kind of robot data, and wiring it into the loop is one of the highest-leverage things in the project.

Insight: observability is not a luxury for a system whose bugs are all silent. When failures crash, a stack trace points at the bug. When failures are plausible wrong numbers, there is nothing to point at — so you must build the ability to see. For this project, “watch the robot” is not a nice-to-have; it is the primary debugging tool, because it is the only tool that catches the entire class of frame, order, and cadence bugs. Time spent on the instrument panel pays back on the first silent bug it reveals.


An instrument panel for the loop

When a server runs with visualization on, every step logs to a shared layout with a place for each part of the observation and action. The panels are:

Everything shares one timeline, so you can scrub to any moment and see the camera frame, the pose, the action, and the latency for that exact step together. A frame bug shows up as the 3D robot leaning the wrong way while the action channel says “forward.” A joint scramble shows up as the wrong limb moving when a channel spikes. A cadence problem shows up in the inference_ms and buffer timing. The panel turns invisible bugs visible.


The posed robot

The centerpiece is world/robot: the robot drawn in 3D, in the pose it is actually in, updated every step. This is not a cartoon — it is the same URDF blueprint the body and server use (Chapter 3), loaded mesh by mesh, and posed from the joint angles by forward kinematics (the geometry of “given these joint angles, where does every link end up”).

Getting there took a small custom tool, because at the time no off-the-shelf loader could animate a URDF in Rerun — the community ones posed it once and froze. So the project parses the URDF itself, loads each link’s mesh, and each step applies the current joint angles down the kinematic tree. Two details from earlier chapters reappear here, confirming they are truly universal:

Because the posed robot is driven by the joint numbers that crossed the wire, it is a faithful mirror of what the robot is actually doing — which is exactly what you want when the question is “is the model’s output sane?”

One profile detail matters: for the humanoid, the proprioception is a vector of joint angles, not an end-effector pose, so the end-effector dot-and-trajectory view is turned off (it would be meaningless). The arm keeps it on, because its proprioception is an end-effector pose. The instrument adapts to what each robot’s numbers actually mean.


Three ways to look

The same logging feeds three viewing modes, chosen by how you are working:

All three are one flag on the server. The recordings are real: a saved session with the Panda arm carries its eleven meshes posed correctly; one with the G1 carries thirty-five. You can hand someone a .rrd and they can scrub through exactly what your robot did, frame by frame.

Insight: the same refrain, one more time — the number tells you where to look, the picture tells you the truth. A reward curve, an error metric, a latency plot: each is a proxy for whether the robot is doing the right thing, and any proxy can be satisfied by the wrong behavior. The instrument panel exists so that no number is ever the final word. When the plot says “good” and you are tempted to move on, the posed robot and the camera frame are there to confirm or contradict it. Trust the picture over the number, every time.


What you now understand

You can now drive the robot from a real model and see what it does. But zero-shot is only reasonable, not good. Part V is the turn where you stop being a spectator and start teaching — beginning with the one component the whole project is designed around: a recorder that does not care who is driving.

Continue to Chapter 13 — The Recorder That Doesn’t Care Who’s Driving.


The visualization is server/rerun_logging.py (the entity layout and blueprint) and server/urdf_rerun.py (the stdlib-XML URDF parser + posed-mesh forward kinematics). It targets Rerun 0.34. Turn it on with --rerun (spawn), --rerun-save FILE.rrd (headless), or --rerun-connect URL (stream to a remote viewer) on any server.