A robot foundation model maps camera frames, proprioception, and a language instruction to a chunk of motor commands. The architectural choices behind that mapping decide your latency budget, your fine-tuning cost, and whether you can even download the weights. This guide compares the four families you will actually evaluate in late 2026: NVIDIA Isaac GR00T, Physical Intelligence π, Google DeepMind Gemini Robotics, and the open-weight alternatives (OpenVLA, Octo, RDT-1B, SmolVLA). It also gives you the control-loop math and the glue code to run them.
Quick Takeaways
- GR00T is the most deployable open option for humanoids. It is commercially licensed, runs on Jetson Thor, and plugs into LeRobot.
- π (openpi) gives you the strongest open-weight manipulation checkpoints (π0, π0-FAST, π0.5). The newest models (π*0.6, π0.7) are not released.
- Gemini Robotics 2 is a three-model stack. Only the reasoning model (ER 2) is publicly callable. The VLA and On-Device variants are gated.
- SmolVLA, Octo, and OpenVLA win on cost and hackability. You can fine-tune them on a single GPU and a SO-101 arm.
| Model family | Latest (as of Oct 2026) | Weights | Best fit |
|---|---|---|---|
| GR00T N-series | N1.7 Early Access | Open, commercial license | Humanoids, Jetson edge |
| π (Physical Intelligence) | π0.7 (research); π0.5 open | π0 / π0-FAST / π0.5 open (Apache-2.0) | Dexterous manipulation fine-tuning |
| Gemini Robotics | 2 / ER 2 / On-Device 2 | ER 2 via API; others gated | High-level planning, multi-robot orchestration |
| Open VLAs | SmolVLA, OpenVLA, Octo, RDT-1B | Fully open | Research, low-cost arms |
What a Robot Foundation Model Actually Computes
Every model in this comparison is a vision-language-action (VLA) policy or a component of one. The policy is a conditional distribution:
a_{t:t+H} ~ π_θ( · | I_t^{1..N}, q_t, ℓ )
I_t^{1..N} : N camera images at time t
q_t : proprioceptive state (joint angles, gripper width)
ℓ : language instruction (tokenized)
H : action chunk horizon (typically 16–50 steps)
Predicting a chunk of H actions instead of one amortizes slow inference across many control ticks. It also smooths trajectories. The families differ in how they decode that chunk:
- Autoregressive tokenization (OpenVLA, π0-FAST) discretizes actions into tokens. It is simple, but decoding is slow.
- Flow matching / diffusion heads (π0, π0.5, GR00T’s DiT) denoise a continuous chunk in a few steps. This is the dominant choice for dexterous work.
- Dual-system designs (GR00T, Gemini Robotics) split slow reasoning (System 2) from fast motor control (System 1).
Flow-Matching Action Head in Plain Text
The π0-style action expert trains on interpolated noisy actions:
A^τ = τ·A + (1 − τ)·ε, ε ~ N(0, I), τ ∈ [0, 1]
Loss L(θ) = E[ || v_θ(A^τ, o, τ) − (A − ε) ||² ]
Inference (Euler integration, N_steps ≈ 10):
A^0 = ε
for k = 0 .. N_steps−1:
A^{τ+δ} = A^τ + δ · v_θ(A^τ, o, τ), δ = 1 / N_steps
Ten integration steps on a small action expert cost far less than ten full VLM forward passes. The observation prefix is cached once, and only the action expert reruns.
NVIDIA Isaac GR00T
NVIDIA’s GR00T line targets humanoid whole-body control. The latest release, GR00T N1.7 Early Access (April 17, 2026), is a 3B-parameter open, commercially licensed VLA built on a Cosmos-Reason2-2B backbone with a 32-layer DiT for low-level motor control. That is a System 2 / System 1 split: a reasoning VLM plans, and a diffusion transformer emits motor actions.
The ecosystem matters as much as the weights. N1.6 is built on Cosmos Reason and enables coordinated whole-body movement, so humanoids can manipulate objects while moving. On June 1, NVIDIA introduced an open GR00T Reference Humanoid that combines a Unitree H2 Plus body, Sharpa dexterous hands, and Jetson Thor onboard compute. The LeRobot integration also means you can load GR00T checkpoints through the same dataset and training tooling as the small open models.
Strengths: commercial license, edge hardware path (Jetson Thor), simulation tooling (Isaac Sim/Lab, Cosmos world models for synthetic data).
Trade-offs: N1.7 is Early Access, and the best results assume NVIDIA’s full stack.
Physical Intelligence (π)
Physical Intelligence has published the most influential recipe for dexterous generalist policies. Three tiers exist:
- π0 / π0-FAST / π0.5: open weights and code in
openpi. The weights, code, and fine-tuned checkpoints ship under Apache-2.0. - π*0.6: trained with the RECAP reinforcement-learning method. It is not released, and community reimplementations exist.
- π0.7: published April 16, 2026 as a steerable model with emergent capabilities. It pairs Google’s open Gemma3 4B language model with an 860M-parameter action expert that generates the robot motions. A Japanese analysis from July 2026 says π*0.6 and π0.7 remain closed, so plan commercial validation around a partnership rather than a download (see the source at ai-souken.com).
Strengths: best open checkpoints for contact-rich tasks (laundry folding, box building), proven fine-tuning on ALOHA, DROID/Franka, and UR5e platforms.
Trade-offs: the frontier model is closed, and openpi expects JAX (a PyTorch port exists).
Gemini Robotics 2
Google DeepMind ships its stack as separate models. Gemini Robotics 2 (announced July 30, 2026) ships as three separate models that work as one system.
| Component | Role | Access |
|---|---|---|
| Gemini Robotics 2 | Full-body VLA (vision + language → motor commands) | Private preview |
| Gemini Robotics ER 2 | Embodied reasoning, planning, tool calls | Gemini API, AI Studio |
| Gemini Robotics On-Device 2 | Lightweight local VLA | Trusted testers |
Unlike its predecessor, the VLA now controls the entire body, so robots can walk, squat, bend, reach, and lift objects in confined spaces. The On-Device model was also demonstrated on smaller bi-arm research robots including Dexmate, Trossen, and SO101.
The only piece you can call today is the reasoning layer. Per its model card, ER 2 is based on Gemini 3.5 Flash and integrates with the Gemini Live API through a bidirectional streaming endpoint. Use it as a high-level planner above whatever low-level policy you run.
# pip install google-genai
from google import genai
from google.genai import types
client = genai.Client() # reads GEMINI_API_KEY from the environment
with open("workspace.jpg", "rb") as f:
img = types.Part.from_bytes(data=f.read(), mime_type="image/jpeg")
prompt = (
"Point to every cup on the table. "
'Return JSON: [{"point": [y, x], "label": "<name>"}] '
"with coordinates normalized to 0-1000."
)
resp = client.models.generate_content(
model="gemini-robotics-er-2-preview", # preview model string
contents=[img, prompt],
config=types.GenerateContentConfig(temperature=0.5),
)
print(resp.text) # parse JSON, then back-project (y, x) through camera intrinsics
Back-project the returned 2D points with a depth image and the camera’s pinhole model:
X = (u − c_x) · Z / f_x
Y = (v − c_y) · Z / f_y
P_base = T_base←cam · [X, Y, Z, 1]^T
Open Alternatives
| Model | Params | Action decoding | Fine-tune hardware | Notes |
|---|---|---|---|---|
| SmolVLA | ~450M | Flow-matching expert | Single consumer GPU | Native in LeRobot; targets SO-101 |
| OpenVLA | 7B | Autoregressive tokens | 1–8 A100-class GPUs (LoRA helps) | Trained on Open X-Embodiment |
| Octo | 27M / 93M | Diffusion head | Single GPU | Small, fast, weaker language grounding |
| RDT-1B | 1.2B | Diffusion transformer | Multi-GPU | Bimanual focus |
Choose these when you need to inspect and modify every layer, run on a budget, or publish reproducible research.
Head-to-Head Comparison
| Criterion | GR00T N1.7 | π0.5 (open) | Gemini Robotics 2 | SmolVLA |
|---|---|---|---|---|
| Access | Open, Early Access | Open (Apache-2.0) | ER 2 API; VLA gated | Open |
| Embodiment focus | Humanoid | Arms, mobile manipulators | Arms to humanoids | Low-cost arms |
| Architecture | Reasoning VLM + DiT | VLM + flow-matching expert | VLM planner + VLA | Small VLM + flow expert |
| Edge deployment | Jetson Thor | Desktop GPU (RTX-class) | On-Device 2 (gated) | Laptop / Jetson-class |
| Fine-tuning cost | Medium | Medium | Not available publicly | Low |
| Commercial use | Yes | Yes (check checkpoint terms) | Partner agreement | Yes |
The Control-Loop Math You Cannot Skip
Inference latency, not model quality, breaks most first deployments. Let f be the control rate, H the chunk length, and T_inf the end-to-end inference time.
Stale actions while inferring: s = ceil(T_inf · f)
Executable actions per chunk: H_exec = H − s
Constraint for continuous motion: H_exec ≥ 1 → H / f > T_inf
Example: f = 50 Hz, H = 50, T_inf = 120 ms.
s = ceil(0.12 × 50) = 6 stale steps, so you can execute up to 44 actions per chunk. Run inference asynchronously, and blend the old and new chunks over the overlap window to avoid velocity discontinuities:
a_blend[i] = (1 − w_i) · a_old[i] + w_i · a_new[i], w_i = i / s, i = 0..s
Re-planning every 10–25 steps (a receding horizon) trades smoothness against reactivity. Shorter intervals recover from slips faster.
Practical Workflow 1: Serve π0.5 to a ROS 2 Robot
Run the policy server on a GPU workstation and a thin rclpy node on the robot. This keeps the real-time loop out of the Python inference process.
# Workstation: serve a pretrained openpi checkpoint over WebSocket
git clone --recurse-submodules https://github.com/Physical-Intelligence/openpi
cd openpi && uv sync
uv run scripts/serve_policy.py policy:checkpoint \
--policy.config=pi05_droid \
--policy.dir=gs://openpi-assets/checkpoints/pi05_droid
# ROS 2 node: queries the policy and publishes joint targets at 50 Hz
import rclpy, numpy as np
from rclpy.node import Node
from sensor_msgs.msg import JointState
from openpi_client import websocket_client_policy
class PiBridge(Node):
def __init__(self):
super().__init__("pi_bridge")
self.policy = websocket_client_policy.WebsocketClientPolicy("192.168.1.50", 8000)
self.pub = self.create_publisher(JointState, "/joint_targets", 10)
self.chunk, self.idx, self.REPLAN = None, 0, 15 # re-plan every 15 steps
self.create_timer(1 / 50.0, self.tick) # 50 Hz control timer
def get_obs(self):
# Fill from your camera and joint topics. Key names must match the checkpoint config.
return {
"observation/exterior_image_1_left": self.ext_img, # HxWx3 uint8
"observation/wrist_image_left": self.wrist_img, # HxWx3 uint8
"observation/joint_position": self.q, # (7,) float32
"observation/gripper_position": self.g, # (1,) float32
"prompt": "put the mug in the bin",
}
def tick(self):
if self.chunk is None or self.idx >= self.REPLAN:
self.chunk = self.policy.infer(self.get_obs())["actions"] # (H, action_dim)
self.idx = 0
a = self.chunk[self.idx]; self.idx += 1
msg = JointState(); msg.position = a[:7].tolist() # joint targets (rad)
self.pub.publish(msg)
def main():
rclpy.init(); rclpy.spin(PiBridge())
The synchronous infer call blocks the timer while it runs. For production, move it to a worker thread and apply the blending formula above.
Practical Workflow 2: Fine-Tune SmolVLA on an SO-101 Dataset
pip install lerobot
# Train from the pretrained base on your recorded teleoperation dataset
lerobot-train \
--policy.path=lerobot/smolvla_base \
--dataset.repo_id=YOUR_HF_USER/so101_pick_place \
--batch_size=64 \
--steps=20000 \
--output_dir=outputs/smolvla_so101 \
--policy.device=cuda
Record 50 or more clean demonstrations per task before training. Keep camera placement identical between data collection and deployment. Camera shift is the most common cause of poor transfer for every model in this guide.
Which Model Should You Pick?
| Scenario | Pick | Why |
|---|---|---|
| Humanoid with Jetson Thor | GR00T N1.7 | Edge path, whole-body focus, commercial license |
| Dexterous tabletop arm, own data | π0.5 via openpi | Strong open checkpoints, proven fine-tunes |
| Multi-step task planning above any policy | Gemini Robotics ER 2 | Public API, tool orchestration, multi-robot coordination |
| Budget research on SO-101 | SmolVLA | Low fine-tuning cost, LeRobot-native |
| Full architectural control | OpenVLA / Octo | Fully open, simple to modify |
A hybrid works well: ER 2 decomposes the task into subgoals, and a local policy (GR00T, π0.5, or SmolVLA) executes each subgoal. Each layer is replaceable.
Evaluation Pitfalls
- Success rate without trial counts hides variance. Run at least 20 trials per task.
- Retrieval versus generalization. Reviewers of π0.7 point out that separating true skill recombination from nearest-neighbor recall is now the central evaluation problem.
- Latency measured on the server only. Include image encoding, network transfer, and ROS message serialization.
- Simulation-to-real gaps. Validate in Isaac Sim or MuJoCo first, then budget hardware time for friction and calibration errors.
FAQ
What is a robot foundation model?
A robot foundation model is a large pretrained network that maps images, robot state, and language instructions to actions across many tasks and robot bodies. Most are vision-language-action (VLA) models built on a pretrained vision-language backbone with an action head.
Is NVIDIA GR00T open source?
The GR00T N-series models are open and commercially licensed, distributed through Hugging Face. The latest, N1.7, is in Early Access. Check the license on each checkpoint before commercial use.
Can I run Gemini Robotics on my own robot?
Partly. Gemini Robotics ER 2 is available through the Gemini API and Google AI Studio. The full-body VLA is in private preview, and On-Device 2 is limited to trusted testers.
What is the best open-source VLA for a low-cost robot arm?
SmolVLA is the practical starting point. It is small, integrated into LeRobot, and fine-tunes on a single GPU. Move to π0.5 via openpi when you need stronger dexterity and have a larger GPU.




