Chapter 14 — Becoming the Robot
This chapter closes Part V. The recorder (Chapter 13) will capture a demonstration from anyone driving; this is how *you drive — from a keyboard, or with your own hands seen through a camera. It assumes you know that teleoperation means a human supplying the actions, that the recorder doesn’t care who drives, and that Unity’s control mask can free a subset of joints.*
Supplying the actions yourself
Zero-shot behavior is only reasonable, not good (Chapters 4 and 10). To make it better you need demonstrations — episodes of the task done right — and to get those, you have to drive the robot yourself. Teleoperation is exactly that: for a while, you are the policy. The observation still comes from Unity, but the action comes from you, and the recorder captures the result as a demonstration.
The elegant part, which you already have the pieces for: from the robot’s and the recorder’s point of view, you are just another driver on the wire. Whatever turns your intent into joint targets simply speaks the same act protocol the mock and the model speak, on the same port, and Unity cannot tell the difference. So teleoperation adds no new plumbing — it adds a new source of actions and inherits everything downstream.
The simplest teleop: a keyboard
The most basic driver is the keyboard. Press a key, nudge the robot: WASD-style keys map to end-effector nudges for the arm, or to velocity commands for a locomotion policy. It is crude — you would not demonstrate a delicate grasp with arrow keys — but it works today, it needs no camera, and it is enough to record simple demonstrations and to prove the teleop path end to end. Keyboard teleop is the “hello world” of becoming the robot: the observation flows from Unity, your keystrokes become the action, the recorder writes an episode. Everything past that keystroke is the pipeline you already have.
Your hands as the controller
The richer, and far more natural, way to demonstrate a manipulation task is to use your own hands. A webcam watches you; software estimates the position of every finger joint; and those human joint angles are mapped onto the robot’s fingers and arm. Move your hand, the robot’s hand moves. This is the teleoperation prototype in the project, and it has three stages worth understanding, because each is a seam that later work swaps out.
See the hand. A camera frame goes into MediaPipe, a hand-and-body tracking library, which returns 21 landmark points per hand (knuckles, joints, fingertips) and, optionally, a full-body pose. This is the “vision” of teleop — turning pixels into a skeleton.
Retarget the fingers. A human hand and a robot hand are not the same shape, so you cannot copy joint angles directly. Instead the software computes each finger’s curl — the interior angle at the knuckle, straight (0) to fully curled (1) — which has a lovely property:
Insight: retarget by angle, not by position. An interior joint angle — how bent a finger is — is invariant to how far your hand is from the camera and how it is rotated. A fist is a fist whether it is near or far, tilted or straight. So mapping human curl onto robot curl sidesteps the hardest part of monocular vision (absolute distance and scale, which a single camera estimates poorly) and keeps only the part it does well (relative angles). The retargeter is a set of ranges — this human curl maps to that robot joint’s radian range — that you tune once per hand design.
Two retargeters cover the project’s hands: one for the 6-degree-of-freedom hands (the H1’s Inspire and the GR-1’s Fourier hands), and one for the G1’s 3-fingered, 7-degree-of-freedom Dex3. Each outputs joint targets in the exact URDF actuated order (Chapter 7), so the robot’s fingers land on the right joints.
Retarget the arm. The arm is harder, because it needs the position of the wrist, not just angles — and position is exactly what one camera struggles with. The webcam path estimates arm joint angles from the body pose in a torso-relative frame, through a tuning table (gain and offset per joint). It works, with honest caveats spelled out below.
The output of all three stages is a chunk of joint targets that simply holds your current pose — the robot mirrors where your hand is now, and updates as you move. That chunk goes out on port 8766 in the ordinary protocol, and the recorder captures it as a demonstration.
What transfers cleanly, and what doesn’t
The prototype is honest about its limits, and the limits are instructive because they are inherent to monocular (single-camera) vision, not bugs to be fixed:
- Finger curl and elbow bend transfer well. These are joint angles, and angles are what one camera recovers reliably. A closed fist reads as a closed fist; a bent elbow reads as a bent elbow.
- Absolute position transfers poorly. Where exactly your wrist is in space — its distance from the camera — is the thing a single lens estimates worst. So the arm follows your configuration faithfully but its reach in metric space is approximate.
- Shoulder yaw and wrist orientation are rough. These depend on subtle 3D cues a monocular view barely captures.
And one limitation that surprises people until it is named: with camera glasses (an egocentric, first-person view from the headset), the fingers track beautifully but the arms cannot be tracked at all.
Insight: the glasses’ arm gap is occlusion, not depth. It is tempting to blame the missing arms on the single camera’s poor depth sense. That is not it. The forward-facing glasses camera simply never sees your own shoulders and torso — they are outside its field of view. There is no body pose to estimate because the body is not in frame. The fix is therefore not a better depth network but a different source of arm information: infer the arm from the wrist’s position via inverse kinematics, or add a sensor that can see the body. Diagnosing “occlusion, not depth” correctly is what points you at the right fix.
That diagnosis is exactly what the project’s third teleop pipeline acts on.
Three pipelines, one seam
The teleop prototype is built as three pipelines that share almost everything and differ only at the input:
| Pipeline | Input | Fingers | Arms | Status |
|---|---|---|---|---|
| Webcam (third-person) | a webcam facing you | ✅ tracked | ✅ from body pose | works |
| Glasses (egocentric) | the headset’s forward camera | ✅ tracked | ❌ your torso is out of frame | fingers work; arms N/A |
| XR (head + hands + IK) | the headset app streams head pose + hand landmarks | ✅ tracked | ✅ via inverse kinematics from wrist position | built and tested in sim; live calibration pending |
The third pipeline answers the glasses’ arm gap directly: since the forward camera cannot see your arms, it reconstructs them. The headset reports where your head and hands are; the software anchors a virtual torso beneath your head, works out where each wrist sits relative to a virtual shoulder, and solves a two-link inverse-kinematics problem (upper arm, forearm) with an “elbow-down” preference — the relaxed way a human elbow hangs. Crucially, that IK path reuses the same tuning table as the webcam path, so the two share calibration: a fix in one improves both.
Insight: build the teleop as swappable seams, and every future sensor is a drop-in. Notice what stayed constant across all three pipelines: the retargeters, the joint layouts, the protocol, the recorder. Only the input — how the hand and arm are sensed — changed. That is deliberate. A better 3D sensor (a depth camera, a headset with body tracking) does not require a new teleop system; it swaps in at one seam and feeds the same downstream chain. The honest, limited webcam prototype is not a dead end — it is the first driver in a family that shares everything but its eyes.
Still one format
Step back and see what teleoperation actually added: nothing downstream. The teleop server is a drop-in backend on the humanoid’s port, speaking the identical chunk protocol as the mock and the model, with Unity in the same useStructuredObs = OFF mode the mock uses. Your hands became a driver; the recorder captured a demonstration; and that demonstration is byte-for-byte the same kind of file a policy rollout produces. You are now a first-class speaker on the wire — which is the sentence this whole book keeps earning.
What you now understand
- Teleoperation makes you the policy: the observation comes from Unity, the action comes from you, and the recorder captures a demonstration. The teleop driver speaks the same
actprotocol as the mock and the model, so it adds a source of actions, not new plumbing. - The simplest driver is a keyboard; the natural one is your hands, tracked by MediaPipe into 21 landmarks per hand, then retargeted by angle (finger curl → robot joint range) — invariant to distance and rotation, which is what makes monocular vision usable.
- Monocular vision transfers angles well and absolute position poorly; the glasses’ missing arms are occlusion, not depth (your torso is out of the forward camera’s frame), which points to reconstructing arms via inverse kinematics rather than better depth.
- Three pipelines (webcam, glasses, XR-with-IK) share retargeters, layouts, protocol, and recorder, and differ only at the input seam — so a better sensor is a drop-in.
- Teleop adds nothing downstream: your hands are a first-class driver on the wire, and the demonstration you record is identical to a policy rollout — one format, many drivers, made literal.
You can now record demonstrations — by keyboard or by hand. Part VI closes the loop: turn a pile of those demonstrations into a training dataset, fine-tune the model on them, and serve the improved policy back to the robot.
Continue to Chapter 15 — From Demonstrations to a Dataset.
The teleop prototype is the handtracking/ directory — teleop_server.py (the drop-in :8766 backend), retarget.py (finger retargeters), arm_retarget.py and arm_ik.py (arm angle estimation and IK), xr_pose.py (the headset stream), and robot_layouts.py (per-robot joint order). It uses MediaPipe’s Tasks API; the camera loop runs on the main thread (a macOS constraint). Keyboard teleop is KeyboardTeleop.cs in the Unity project. The input-pipeline roadmap (P1–P6) is handtracking/docs/TELEOP_PIPELINES.md.