Skip to the content.

Methods & Protocol Reference

A lookup handbook for the sim2real workbook. Companion to the book index.


Who this is for: anyone using this repository and wanting one place to look up the WebSocket protocol, ports, embodiment profiles, server flags, frame and joint conventions, the recording format, and known gotchas. No robotics or ML background assumed — every term is defined on first use (or in the chapter cited).

How to use it: jump to the section you need. Nothing here is a narrative; it is designed to be scanned, not read top to bottom. Values are quoted from the code as of this writing — when in doubt, the source file named at the end of each section is ground truth.


Contents

  1. The WebSocket protocol
  2. Port map
  3. Frame conventions
  4. Joint orders
  5. GR00T embodiment profiles
  6. Server flags
  7. The recording format
  8. Fine-tuning quick reference
  9. Rerun entity layout
  10. Environments & ops
  11. Known gotchas

1. The WebSocket protocol

JSON text frames, capped at 16 MB (comfortably above a 224×224 JPEG). Every server accepts the same three message types. Introduced in Chapter 5.

Common (all servers):

Request Reply
{"type":"ping"} {"type":"pong"}
{"type":"reset"} {"type":"reset_ok"} (clears policy + visualization state)
malformed / unknown {"type":"error","message":"..."} — connection stays open

Arm act (OpenVLA, port 8765):

// request
{"type":"act", "image":"<b64 jpeg>", "instruction":"...",
 "proprio":[x,y,z,roll,pitch,yaw,gripper], "joints":[...optional...], "timestep":7}
// reply
{"type":"action", "action":[dx,dy,dz,droll,dpitch,dyaw,gripper],
 "timestep":7, "inference_ms":210.5}

Action = end-effector deltas (meters / radians) in the robot base frame; gripper ~1.0 open, ~0.0 closed. proprio = current EE pose + gripper. joints is optional and used only for the posed visualization.

Humanoid act (GR00T 8766 / LingBot 8767 / motion 8766 / teleop 8766):

// request
{"type":"act", "image":"<b64 jpeg>", "instruction":"...",
 "state":[q0..qN-1]  OR  {"left_arm":[...], "waist":[...], ...}, "timestep":42}
// reply
{"type":"action_chunk", "actions_flat":[K*M row-major], "horizon":K, "dim":M,
 "chunk_hz":50.0, "timestep":42, "inference_ms":619.2,
 "action_layout":[["left_arm",7],...]  // optional; GR00T zmq path only
}

Action = absolute joint targets (radians). Read action k as actions_flat[k*M : (k+1)*M]. chunk_hz is the playback rate — a fact about the model’s training data, not a free knob (Chapter 11). state may be a flat vector or a per-part dict.

RL locomotion obs (port 8769, per-tick, NOT chunked):

// request
{"type":"obs", "timestep":t, "base_lin_vel":[3], "base_ang_vel":[3],
 "projected_gravity":[3], "commands":[vx,vy,wz],
 "joint_names":[...], "joint_pos":[rad], "joint_vel":[rad/s]}
// reply
{"type":"action", "joint_names":[mapped subset], "targets":[abs radians],
 "timestep":t, "inference_ms":...}

One observation → one action, 50 Hz, closed-loop (Chapter 18). Also supports {"type":"spec"} → the policy’s dims/joint order.


2. Port map

Port Server Robot Style Bind
8765 inference_server.py (OpenVLA) arm (Panda / WidowX) single delta, ~5 Hz 0.0.0.0
8766 groot_server.py (GR00T) humanoid chunk 0.0.0.0
8766 motion_server.py (classical) Panda arm chunk, dim 9 shares 8766
8766 teleop_server.py (hand teleop) humanoid chunk (holds pose) 0.0.0.0
8767 lingbot_server.py (LingBot) humanoid chunk 127.0.0.1
8769 rl_policy_server.py (RL) G1 locomotion per-tick —
8010 openpi deploy (behind LingBot bridge) — msgpack 127.0.0.1
5555 Isaac-GR00T PolicyServer (behind GR00T bridge) — ZMQ 127.0.0.1

Port 8766 is shared by everything speaking the humanoid chunk protocol (GR00T, mock, teleop, motion) — they are interchangeable, so only one runs at a time. Servers have no authentication: bind to loopback or a private network, reach remote boxes via SSH tunnel, never expose to the internet.


3. Frame conventions

Robot/ROS frame: right-handed, x forward, y left, z up. Unity: left-handed, x right, y up, z forward. Opposite handedness ⇒ conversion is a relabel plus a sign flip. Same mapping as Unity’s ROS-TCP-Connector. Detail in Chapter 6.

Position:   unity = (-y,  z,  x)        robot = ( uz, -ux,  uy)
Quaternion: unity = ( qy, -qz, -qx, qw)   from ROS (qx,qy,qz,qw)
Angular velocity (RL only): robot = (-uz, ux, -uy)   ← every sign flipped vs position

Source: DevVLA/Assets/Scripts/VLA/FrameConversions.cs.


4. Joint orders

Everything is URDF document order. Server, Unity, and visualization all derive the order from the same URDF, resolving joints by URDF jointName. Action arrays contain only actuated joints (movable, non-mimic). Detail in Chapter 7.

Unitree G1, 29-DoF (g1_29dof_rev_1_0.urdf):

Index Block
0–5 left leg: hip_pitch, hip_roll, hip_yaw, knee, ankle_pitch, ankle_roll
6–11 right leg (same)
12–14 waist: yaw, roll, pitch
15–21 left arm: shoulder_pitch, shoulder_roll, shoulder_yaw, elbow, wrist_roll, wrist_pitch, wrist_yaw
22–28 right arm (same)

Manipulation scenes: legs (0–11) held, upper body (12–28) driven, base pinned.

Other bodies (joints → actuated after mimic removal):

Robot URDF joints Actuated Hand DoF/side Arm DoF/side
Franka Panda — 7 arm + 2 finger (1 mimic) — 7
G1 + Dex3 43 43 (no mimic — full fidelity) 7 (3-finger) 7
Fourier GR-1 54 44 6 (mimic-coupled) 7
Unitree H1 + Inspire 45 33 6 (mimic-coupled) 4 (no wrist)

Mimic-coupled hands mean the actuated finger DoF become contiguous; miscounting scrambles the whole hand.


5. GR00T embodiment profiles

Selected with --embodiment <name>. The observation the model receives is nested: {video:{cam:(1,T,H,W,C)}, state:{part:(1,1,D)}, language:{key:[[instruction]]}}. Only real_g1 runs zero-shot on the base GR00T-N1.7-3B checkpoint. Detail in Chapter 10.

--embodiment State dim Action dim Camera Lang key Native chunk_hz Status
real_g1 49 53 ego_view (T=2) annotation.human.task_description 50 verified, zero-shot
unitree_g1_fullbody 43 35 ego_view (T=2) same 50 unverified
gr1_tabletop 29 29 ..._res256_freq20 (T=2) task 50 finetune-only (--image-size 256)
h1 21 21 ego_view (T=2) same 50 template only
g1_push 29 29 ego_view (T=1) same 5 new-embodiment (fine-tuned)
g1_dex3_task 43 43 ego_view (T=1) same 5 new-embodiment

real_g1 state layout (49): left_wrist_eef_9d(9) · right_wrist_eef_9d(9) · left_hand(7) · right_hand(7) · left_arm(7) · right_arm(7) · waist(3). The g1_push/g1_dex3_task profiles use a single flat joint_position key (Unity sends flat measured joints, useStructuredObs = OFF).

Gotcha: an all-zero *_wrist_eef_9d slot makes the model raise “SVD did not converge” — the bridge reseeds it with identity rotation [0,0,0, 1,0,0, 0,1,0]. The video slot needs ≥2 frames for real N1.7 embodiments (left-padded at the start).


6. Server flags

inference_server.py (OpenVLA, 8765): --mock · --model openvla/openvla-7b · --unnorm-key bridge_orig (use dataset name after fine-tune) · --quantize {none,8bit,4bit} · --attn {sdpa,eager,flash_attention_2} · --rerun / --rerun-save FILE.rrd / --rerun-connect URL / --rerun-urdf {auto,none,PATH}.

groot_server.py (GR00T, 8766): --backend {mock,zmq} · --embodiment <name> (see §5) · --profile-json PATH · --urdf PATH · --horizon 40 (mock only) · --chunk-hz N (override; else profile native, else 50) · --zmq-host/--zmq-port 5555 · --image-size 224 · --video-key/--annotation-key · --action-keys ... · --sonic-onnx DECODER.onnx / --sonic-mock · --rerun*.

lingbot_server.py (LingBot, 8767): --deploy-host/--deploy-port 8010 · --robo-name robotwin · --camera-keys ... (Unity’s single frame is duplicated to each) · --state-key observation.state · --image-size 256 · --chunk-hz 15 · --lingbot-dir PATH. No --mock — requires the upstream deploy server.

motion_server.py (classical, 8766): --profile {raw,quintic,trapezoid,scurve,uniform,toppra} · --program {home,pick_place,figure8} · --horizon 32 · --hz 50 · --plan-hz 200 · --transition-s 1.5 · --record (planned-vs-measured JSONL) · --rerun*.

rl_policy_server.py (RL, 8769): --policy {stand,onnx} · --onnx POLICY.onnx · --spec g1_flat_env_spec.json.


7. The recording format

EpisodeRecorder writes one directory per episode. Decoupled from the action source — policy, keyboard, and teleop all call the same RecordStep. Detail in Chapter 13.

vla_recordings/episode_000000/
  meta.json          {"instruction","fps","robot_type","success","num_steps"}
  steps.jsonl        one JSON object per step (below)
  frames/000000.jpg  one camera frame per step, index-aligned

steps.jsonl row keys, in order: timestep (int), state ([N floats]), action ([N floats]), joints ([N floats], only if provided), correction (true, only during human takeover — the DAgger marker), timestamp (seconds since episode start). Contract: the recorded state/frame is the observation the action was decided from. N-dimensional and generic — same recorder for a 7-DoF arm and a 29-DoF humanoid. Files are UTF-8 (with a BOM; converters read utf-8-sig).


8. Fine-tuning quick reference

Episodes → dataset → fine-tune → serve. Detail in Chapters 15–16.

Embodiment tiers (GR00T):

Tier Example Zero-shot? Cost
Pretrained REAL_G1 yes warm start, fewest demos
Fine-tune-only Fourier GR-1 (gr1_tabletop) no config ships, no checkpoint
New embodiment Unitree H1; 29-DoF push G1 no author config, projector trains from scratch

The push campaign (worked example): 80 episodes of “push the objects off the table” → LeRobot v2.1 + 29-DoF NEW_EMBODIMENT → fine-tune GR00T-N1.7-3B (backbone frozen, projector + diffusion head trained — not LoRA), ~1.5 s/step, 26/121 GB → checkpoint-1000, train_loss 0.874 → offline eval MAE ≈ 0.0238 rad (~1.4°) → served at native 5 Hz (16×29 chunks, ~620 ms/inference). Converter output includes modality.json (names each part — state/action = joint_position[0:29]) and stats.json (per-dim mean/std/min/max/q01/q99). GR00T trains on LeRobot; OpenVLA trains on RLDS.


9. Rerun entity layout

Targets Rerun 0.34. One flag on any server: --rerun (spawn window), --rerun-save FILE.rrd (headless), --rerun-connect URL (remote). Detail in Chapter 12.

Entity Content
camera the exact JPEG the model received (EncodedImage passthrough)
world/robot the posed URDF mesh, forward-kinematics from the joint array (document order; mimic joints followed)
world/ee, world/ee_trajectory end-effector point (green open / red closed) and accumulated path (arm only)
action, proprio per-channel time series (named lines)
task the instruction (logged on change)
diagnostics/inference_ms model latency per step

Humanoid disables the EE view (ee_from_proprio = False) — its proprioception is joint angles, not an EE pose. Validate a saved recording with rerun rrd stats FILE.rrd.


10. Environments & ops

Mac (dev): cd server && uv sync → base env (websockets, numpy, pillow, rerun-sdk ≥0.34,<0.35, trimesh, scipy). Runs mock + Rerun for both tracks, no GPU. Extras: gpu (torch, transformers), quant (bitsandbytes, Linux), lerobot (Python 3.12), groot, motion (Pinocchio).

DGX Spark (owned GPU box, GB10, aarch64 CUDA 13): the durable real-inference box. scripts/spark/{setup,launch,stop}. Unity reaches it at ws://chaotic-spark:8766 over Tailscale — no tunnel. Blackwell/sm_121 quirks handled (NVRTC + PTXAS redirects). Stop = kill two processes (no billing).

Lambda (rented fallback): A10 (24 GB, fits OpenVLA bf16) or A100 (fine-tuning). scripts/lambda/*.sh (launch/setup/terminate — bills until terminated). Reach via SSH tunnel (ssh -N -L 8766:localhost:8766 ubuntu@IP); Unity keeps ws://localhost:8766.

Unity endpoint switch: VLA/Server Endpoint/{DGX Spark | Lambda | Local hand-teleop} rewrites every bridge’s URL in the scene, no retyping.


11. Known gotchas


This reference is a companion to the twenty-chapter book — it is the toolkit; the book is the story behind it. When a value here disagrees with the code, the code (and the repository’s CLAUDE.md) is ground truth.