Methods & Protocol Reference
A lookup handbook for the sim2real workbook. Companion to the book index.
Who this is for: anyone using this repository and wanting one place to look up the WebSocket protocol, ports, embodiment profiles, server flags, frame and joint conventions, the recording format, and known gotchas. No robotics or ML background assumed — every term is defined on first use (or in the chapter cited).
How to use it: jump to the section you need. Nothing here is a narrative; it is designed to be scanned, not read top to bottom. Values are quoted from the code as of this writing — when in doubt, the source file named at the end of each section is ground truth.
Contents
- The WebSocket protocol
- Port map
- Frame conventions
- Joint orders
- GR00T embodiment profiles
- Server flags
- The recording format
- Fine-tuning quick reference
- Rerun entity layout
- Environments & ops
- Known gotchas
1. The WebSocket protocol
JSON text frames, capped at 16 MB (comfortably above a 224×224 JPEG). Every server accepts the same three message types. Introduced in Chapter 5.
Common (all servers):
| Request | Reply |
|---|---|
{"type":"ping"} |
{"type":"pong"} |
{"type":"reset"} |
{"type":"reset_ok"} (clears policy + visualization state) |
| malformed / unknown | {"type":"error","message":"..."} — connection stays open |
Arm act (OpenVLA, port 8765):
// request
{"type":"act", "image":"<b64 jpeg>", "instruction":"...",
"proprio":[x,y,z,roll,pitch,yaw,gripper], "joints":[...optional...], "timestep":7}
// reply
{"type":"action", "action":[dx,dy,dz,droll,dpitch,dyaw,gripper],
"timestep":7, "inference_ms":210.5}
Action = end-effector deltas (meters / radians) in the robot base frame; gripper ~1.0 open, ~0.0 closed. proprio = current EE pose + gripper. joints is optional and used only for the posed visualization.
Humanoid act (GR00T 8766 / LingBot 8767 / motion 8766 / teleop 8766):
// request
{"type":"act", "image":"<b64 jpeg>", "instruction":"...",
"state":[q0..qN-1] OR {"left_arm":[...], "waist":[...], ...}, "timestep":42}
// reply
{"type":"action_chunk", "actions_flat":[K*M row-major], "horizon":K, "dim":M,
"chunk_hz":50.0, "timestep":42, "inference_ms":619.2,
"action_layout":[["left_arm",7],...] // optional; GR00T zmq path only
}
Action = absolute joint targets (radians). Read action k as actions_flat[k*M : (k+1)*M]. chunk_hz is the playback rate — a fact about the model’s training data, not a free knob (Chapter 11). state may be a flat vector or a per-part dict.
RL locomotion obs (port 8769, per-tick, NOT chunked):
// request
{"type":"obs", "timestep":t, "base_lin_vel":[3], "base_ang_vel":[3],
"projected_gravity":[3], "commands":[vx,vy,wz],
"joint_names":[...], "joint_pos":[rad], "joint_vel":[rad/s]}
// reply
{"type":"action", "joint_names":[mapped subset], "targets":[abs radians],
"timestep":t, "inference_ms":...}
One observation → one action, 50 Hz, closed-loop (Chapter 18). Also supports {"type":"spec"} → the policy’s dims/joint order.
2. Port map
| Port | Server | Robot | Style | Bind |
|---|---|---|---|---|
| 8765 | inference_server.py (OpenVLA) |
arm (Panda / WidowX) | single delta, ~5 Hz | 0.0.0.0 |
| 8766 | groot_server.py (GR00T) |
humanoid | chunk | 0.0.0.0 |
| 8766 | motion_server.py (classical) |
Panda arm | chunk, dim 9 | shares 8766 |
| 8766 | teleop_server.py (hand teleop) |
humanoid | chunk (holds pose) | 0.0.0.0 |
| 8767 | lingbot_server.py (LingBot) |
humanoid | chunk | 127.0.0.1 |
| 8769 | rl_policy_server.py (RL) |
G1 locomotion | per-tick | — |
| 8010 | openpi deploy (behind LingBot bridge) | — | msgpack | 127.0.0.1 |
| 5555 | Isaac-GR00T PolicyServer (behind GR00T bridge) | — | ZMQ | 127.0.0.1 |
Port 8766 is shared by everything speaking the humanoid chunk protocol (GR00T, mock, teleop, motion) — they are interchangeable, so only one runs at a time. Servers have no authentication: bind to loopback or a private network, reach remote boxes via SSH tunnel, never expose to the internet.
3. Frame conventions
Robot/ROS frame: right-handed, x forward, y left, z up. Unity: left-handed, x right, y up, z forward. Opposite handedness ⇒ conversion is a relabel plus a sign flip. Same mapping as Unity’s ROS-TCP-Connector. Detail in Chapter 6.
Position: unity = (-y, z, x) robot = ( uz, -ux, uy)
Quaternion: unity = ( qy, -qz, -qx, qw) from ROS (qx,qy,qz,qw)
Angular velocity (RL only): robot = (-uz, ux, -uy) ← every sign flipped vs position
Source: DevVLA/Assets/Scripts/VLA/FrameConversions.cs.
4. Joint orders
Everything is URDF document order. Server, Unity, and visualization all derive the order from the same URDF, resolving joints by URDF jointName. Action arrays contain only actuated joints (movable, non-mimic). Detail in Chapter 7.
Unitree G1, 29-DoF (g1_29dof_rev_1_0.urdf):
| Index | Block |
|---|---|
| 0–5 | left leg: hip_pitch, hip_roll, hip_yaw, knee, ankle_pitch, ankle_roll |
| 6–11 | right leg (same) |
| 12–14 | waist: yaw, roll, pitch |
| 15–21 | left arm: shoulder_pitch, shoulder_roll, shoulder_yaw, elbow, wrist_roll, wrist_pitch, wrist_yaw |
| 22–28 | right arm (same) |
Manipulation scenes: legs (0–11) held, upper body (12–28) driven, base pinned.
Other bodies (joints → actuated after mimic removal):
| Robot | URDF joints | Actuated | Hand DoF/side | Arm DoF/side |
|---|---|---|---|---|
| Franka Panda | — | 7 arm + 2 finger (1 mimic) | — | 7 |
| G1 + Dex3 | 43 | 43 (no mimic — full fidelity) | 7 (3-finger) | 7 |
| Fourier GR-1 | 54 | 44 | 6 (mimic-coupled) | 7 |
| Unitree H1 + Inspire | 45 | 33 | 6 (mimic-coupled) | 4 (no wrist) |
Mimic-coupled hands mean the actuated finger DoF become contiguous; miscounting scrambles the whole hand.
5. GR00T embodiment profiles
Selected with --embodiment <name>. The observation the model receives is nested: {video:{cam:(1,T,H,W,C)}, state:{part:(1,1,D)}, language:{key:[[instruction]]}}. Only real_g1 runs zero-shot on the base GR00T-N1.7-3B checkpoint. Detail in Chapter 10.
--embodiment |
State dim | Action dim | Camera | Lang key | Native chunk_hz |
Status |
|---|---|---|---|---|---|---|
real_g1 |
49 | 53 | ego_view (T=2) |
annotation.human.task_description |
50 | verified, zero-shot |
unitree_g1_fullbody |
43 | 35 | ego_view (T=2) |
same | 50 | unverified |
gr1_tabletop |
29 | 29 | ..._res256_freq20 (T=2) |
task |
50 | finetune-only (--image-size 256) |
h1 |
21 | 21 | ego_view (T=2) |
same | 50 | template only |
g1_push |
29 | 29 | ego_view (T=1) |
same | 5 | new-embodiment (fine-tuned) |
g1_dex3_task |
43 | 43 | ego_view (T=1) |
same | 5 | new-embodiment |
real_g1 state layout (49): left_wrist_eef_9d(9) · right_wrist_eef_9d(9) · left_hand(7) · right_hand(7) · left_arm(7) · right_arm(7) · waist(3). The g1_push/g1_dex3_task profiles use a single flat joint_position key (Unity sends flat measured joints, useStructuredObs = OFF).
Gotcha: an all-zero *_wrist_eef_9d slot makes the model raise “SVD did not converge” — the bridge reseeds it with identity rotation [0,0,0, 1,0,0, 0,1,0]. The video slot needs ≥2 frames for real N1.7 embodiments (left-padded at the start).
6. Server flags
inference_server.py (OpenVLA, 8765): --mock · --model openvla/openvla-7b · --unnorm-key bridge_orig (use dataset name after fine-tune) · --quantize {none,8bit,4bit} · --attn {sdpa,eager,flash_attention_2} · --rerun / --rerun-save FILE.rrd / --rerun-connect URL / --rerun-urdf {auto,none,PATH}.
groot_server.py (GR00T, 8766): --backend {mock,zmq} · --embodiment <name> (see §5) · --profile-json PATH · --urdf PATH · --horizon 40 (mock only) · --chunk-hz N (override; else profile native, else 50) · --zmq-host/--zmq-port 5555 · --image-size 224 · --video-key/--annotation-key · --action-keys ... · --sonic-onnx DECODER.onnx / --sonic-mock · --rerun*.
lingbot_server.py (LingBot, 8767): --deploy-host/--deploy-port 8010 · --robo-name robotwin · --camera-keys ... (Unity’s single frame is duplicated to each) · --state-key observation.state · --image-size 256 · --chunk-hz 15 · --lingbot-dir PATH. No --mock — requires the upstream deploy server.
motion_server.py (classical, 8766): --profile {raw,quintic,trapezoid,scurve,uniform,toppra} · --program {home,pick_place,figure8} · --horizon 32 · --hz 50 · --plan-hz 200 · --transition-s 1.5 · --record (planned-vs-measured JSONL) · --rerun*.
rl_policy_server.py (RL, 8769): --policy {stand,onnx} · --onnx POLICY.onnx · --spec g1_flat_env_spec.json.
7. The recording format
EpisodeRecorder writes one directory per episode. Decoupled from the action source — policy, keyboard, and teleop all call the same RecordStep. Detail in Chapter 13.
vla_recordings/episode_000000/
meta.json {"instruction","fps","robot_type","success","num_steps"}
steps.jsonl one JSON object per step (below)
frames/000000.jpg one camera frame per step, index-aligned
steps.jsonl row keys, in order: timestep (int), state ([N floats]), action ([N floats]), joints ([N floats], only if provided), correction (true, only during human takeover — the DAgger marker), timestamp (seconds since episode start). Contract: the recorded state/frame is the observation the action was decided from. N-dimensional and generic — same recorder for a 7-DoF arm and a 29-DoF humanoid. Files are UTF-8 (with a BOM; converters read utf-8-sig).
8. Fine-tuning quick reference
Episodes → dataset → fine-tune → serve. Detail in Chapters 15–16.
Embodiment tiers (GR00T):
| Tier | Example | Zero-shot? | Cost |
|---|---|---|---|
| Pretrained | REAL_G1 |
yes | warm start, fewest demos |
| Fine-tune-only | Fourier GR-1 (gr1_tabletop) |
no | config ships, no checkpoint |
| New embodiment | Unitree H1; 29-DoF push G1 | no | author config, projector trains from scratch |
The push campaign (worked example): 80 episodes of “push the objects off the table” → LeRobot v2.1 + 29-DoF NEW_EMBODIMENT → fine-tune GR00T-N1.7-3B (backbone frozen, projector + diffusion head trained — not LoRA), ~1.5 s/step, 26/121 GB → checkpoint-1000, train_loss 0.874 → offline eval MAE ≈ 0.0238 rad (~1.4°) → served at native 5 Hz (16×29 chunks, ~620 ms/inference). Converter output includes modality.json (names each part — state/action = joint_position[0:29]) and stats.json (per-dim mean/std/min/max/q01/q99). GR00T trains on LeRobot; OpenVLA trains on RLDS.
9. Rerun entity layout
Targets Rerun 0.34. One flag on any server: --rerun (spawn window), --rerun-save FILE.rrd (headless), --rerun-connect URL (remote). Detail in Chapter 12.
| Entity | Content |
|---|---|
camera |
the exact JPEG the model received (EncodedImage passthrough) |
world/robot |
the posed URDF mesh, forward-kinematics from the joint array (document order; mimic joints followed) |
world/ee, world/ee_trajectory |
end-effector point (green open / red closed) and accumulated path (arm only) |
action, proprio |
per-channel time series (named lines) |
task |
the instruction (logged on change) |
diagnostics/inference_ms |
model latency per step |
Humanoid disables the EE view (ee_from_proprio = False) — its proprioception is joint angles, not an EE pose. Validate a saved recording with rerun rrd stats FILE.rrd.
10. Environments & ops
Mac (dev): cd server && uv sync → base env (websockets, numpy, pillow, rerun-sdk ≥0.34,<0.35, trimesh, scipy). Runs mock + Rerun for both tracks, no GPU. Extras: gpu (torch, transformers), quant (bitsandbytes, Linux), lerobot (Python 3.12), groot, motion (Pinocchio).
DGX Spark (owned GPU box, GB10, aarch64 CUDA 13): the durable real-inference box. scripts/spark/{setup,launch,stop}. Unity reaches it at ws://chaotic-spark:8766 over Tailscale — no tunnel. Blackwell/sm_121 quirks handled (NVRTC + PTXAS redirects). Stop = kill two processes (no billing).
Lambda (rented fallback): A10 (24 GB, fits OpenVLA bf16) or A100 (fine-tuning). scripts/lambda/*.sh (launch/setup/terminate — bills until terminated). Reach via SSH tunnel (ssh -N -L 8766:localhost:8766 ubuntu@IP); Unity keeps ws://localhost:8766.
Unity endpoint switch: VLA/Server Endpoint/{DGX Spark | Lambda | Local hand-teleop} rewrites every bridge’s URL in the scene, no retyping.
11. Known gotchas
- Degrees vs radians (Unity): drive targets (
xDrive.target) are degrees; measured angles (jointPosition[0]) are radians. Convert on the way out (Rad2Deg), not in. chunk_hzis a data fact, not a setting. A model trained at 5 fps must be served at 5 Hz. Wrong rate ⇒ motion 10× too fast and buffer starvation. Profiles now carry their native rate (Chapter 11).- Frame signs are silent when wrong — a valid vector meaning the wrong direction. Angular velocity uses a different sign pattern than position (Chapter 6).
- Mimic joints are excluded from the action space: GR-1 54→44 actuated, H1 45→33. The actuated count is the contract, not the joint count (Chapter 7).
- All-zero wrist-pose slot → GR00T “SVD did not converge”; reseed with identity rotation. Video window needs ≥2 frames for real N1.7 embodiments.
useStructuredObs: ON for real GR00T/LingBot structured obs; OFF for the mock, teleop, motion planner, and the flat-vector fine-tuned profiles (g1_push).- Apple Silicon URDF import: the scripted importer must use
ImportSettings.convexDecomposer.unity, not.vHACD(the V-HACD library ships an Intel-only binary; arm64 Unity throws and generates no colliders). - Toppra (motion time-parameterization) has no macOS wheels — Linux/Spark only; the planner falls back to the uniform profile on a Mac.
- Port 8766 is shared by GR00T, the mock, teleop, and the motion planner — run one at a time.
- Servers have no auth — bind loopback / private network, tunnel remote, never expose.
- Base pinning: manipulation scenes pin the pelvis (immovable) and hold the legs — a position-controlled biped won’t balance in PhysX without a learned controller (Chapter 18). The RL scene is the only one that un-pins the base.
This reference is a companion to the twenty-chapter book — it is the toolkit; the book is the story behind it. When a value here disagrees with the code, the code (and the repository’s CLAUDE.md) is ground truth.