Robot Foundation Models: GR00T, pi, Gemini Robotics, and Open Alternatives Compared

A robot foundation model maps camera frames, proprioception, and a language instruction to a chunk of motor commands. The architectural choices behind that mapping decide your latency budget, your fine-tuning cost, and whether you can even download the weights. This guide compares the four families you will actually evaluate in late 2026: NVIDIA Isaac GR00T, Physical Intelligence π, Google DeepMind Gemini Robotics, and the open-weight alternatives (OpenVLA, Octo, RDT-1B, SmolVLA). It also gives you the control-loop math and the glue code to run them.

Quick Takeaways

  • GR00T is the most deployable open option for humanoids. It is commercially licensed, runs on Jetson Thor, and plugs into LeRobot.
  • π (openpi) gives you the strongest open-weight manipulation checkpoints (π0, π0-FAST, π0.5). The newest models (π*0.6, π0.7) are not released.
  • Gemini Robotics 2 is a three-model stack. Only the reasoning model (ER 2) is publicly callable. The VLA and On-Device variants are gated.
  • SmolVLA, Octo, and OpenVLA win on cost and hackability. You can fine-tune them on a single GPU and a SO-101 arm.
Model family Latest (as of Oct 2026) Weights Best fit
GR00T N-series N1.7 Early Access Open, commercial license Humanoids, Jetson edge
π (Physical Intelligence) π0.7 (research); π0.5 open π0 / π0-FAST / π0.5 open (Apache-2.0) Dexterous manipulation fine-tuning
Gemini Robotics 2 / ER 2 / On-Device 2 ER 2 via API; others gated High-level planning, multi-robot orchestration
Open VLAs SmolVLA, OpenVLA, Octo, RDT-1B Fully open Research, low-cost arms

What a Robot Foundation Model Actually Computes

Every model in this comparison is a vision-language-action (VLA) policy or a component of one. The policy is a conditional distribution:

a_{t:t+H} ~ π_θ( · | I_t^{1..N}, q_t, ℓ )

I_t^{1..N} : N camera images at time t
q_t        : proprioceptive state (joint angles, gripper width)
ℓ          : language instruction (tokenized)
H          : action chunk horizon (typically 16–50 steps)

Predicting a chunk of H actions instead of one amortizes slow inference across many control ticks. It also smooths trajectories. The families differ in how they decode that chunk:

  • Autoregressive tokenization (OpenVLA, π0-FAST) discretizes actions into tokens. It is simple, but decoding is slow.
  • Flow matching / diffusion heads (π0, π0.5, GR00T’s DiT) denoise a continuous chunk in a few steps. This is the dominant choice for dexterous work.
  • Dual-system designs (GR00T, Gemini Robotics) split slow reasoning (System 2) from fast motor control (System 1).

Flow-Matching Action Head in Plain Text

The π0-style action expert trains on interpolated noisy actions:

A^τ = τ·A + (1 − τ)·ε,        ε ~ N(0, I),  τ ∈ [0, 1]

Loss  L(θ) = E[ || v_θ(A^τ, o, τ) − (A − ε) ||² ]

Inference (Euler integration, N_steps ≈ 10):
    A^0 = ε
    for k = 0 .. N_steps−1:
        A^{τ+δ} = A^τ + δ · v_θ(A^τ, o, τ),   δ = 1 / N_steps

Ten integration steps on a small action expert cost far less than ten full VLM forward passes. The observation prefix is cached once, and only the action expert reruns.

NVIDIA Isaac GR00T

NVIDIA’s GR00T line targets humanoid whole-body control. The latest release, GR00T N1.7 Early Access (April 17, 2026), is a 3B-parameter open, commercially licensed VLA built on a Cosmos-Reason2-2B backbone with a 32-layer DiT for low-level motor control. That is a System 2 / System 1 split: a reasoning VLM plans, and a diffusion transformer emits motor actions.

The ecosystem matters as much as the weights. N1.6 is built on Cosmos Reason and enables coordinated whole-body movement, so humanoids can manipulate objects while moving. On June 1, NVIDIA introduced an open GR00T Reference Humanoid that combines a Unitree H2 Plus body, Sharpa dexterous hands, and Jetson Thor onboard compute. The LeRobot integration also means you can load GR00T checkpoints through the same dataset and training tooling as the small open models.

Strengths: commercial license, edge hardware path (Jetson Thor), simulation tooling (Isaac Sim/Lab, Cosmos world models for synthetic data).
Trade-offs: N1.7 is Early Access, and the best results assume NVIDIA’s full stack.

Physical Intelligence (π)

Physical Intelligence has published the most influential recipe for dexterous generalist policies. Three tiers exist:

  1. π0 / π0-FAST / π0.5: open weights and code in openpi. The weights, code, and fine-tuned checkpoints ship under Apache-2.0.
  2. π*0.6: trained with the RECAP reinforcement-learning method. It is not released, and community reimplementations exist.
  3. π0.7: published April 16, 2026 as a steerable model with emergent capabilities. It pairs Google’s open Gemma3 4B language model with an 860M-parameter action expert that generates the robot motions. A Japanese analysis from July 2026 says π*0.6 and π0.7 remain closed, so plan commercial validation around a partnership rather than a download (see the source at ai-souken.com).

Strengths: best open checkpoints for contact-rich tasks (laundry folding, box building), proven fine-tuning on ALOHA, DROID/Franka, and UR5e platforms.
Trade-offs: the frontier model is closed, and openpi expects JAX (a PyTorch port exists).

Gemini Robotics 2

Google DeepMind ships its stack as separate models. Gemini Robotics 2 (announced July 30, 2026) ships as three separate models that work as one system.

Component Role Access
Gemini Robotics 2 Full-body VLA (vision + language → motor commands) Private preview
Gemini Robotics ER 2 Embodied reasoning, planning, tool calls Gemini API, AI Studio
Gemini Robotics On-Device 2 Lightweight local VLA Trusted testers

Unlike its predecessor, the VLA now controls the entire body, so robots can walk, squat, bend, reach, and lift objects in confined spaces. The On-Device model was also demonstrated on smaller bi-arm research robots including Dexmate, Trossen, and SO101.

The only piece you can call today is the reasoning layer. Per its model card, ER 2 is based on Gemini 3.5 Flash and integrates with the Gemini Live API through a bidirectional streaming endpoint. Use it as a high-level planner above whatever low-level policy you run.

# pip install google-genai
from google import genai
from google.genai import types

client = genai.Client()  # reads GEMINI_API_KEY from the environment

with open("workspace.jpg", "rb") as f:
    img = types.Part.from_bytes(data=f.read(), mime_type="image/jpeg")

prompt = (
    "Point to every cup on the table. "
    'Return JSON: [{"point": [y, x], "label": "<name>"}] '
    "with coordinates normalized to 0-1000."
)

resp = client.models.generate_content(
    model="gemini-robotics-er-2-preview",  # preview model string
    contents=[img, prompt],
    config=types.GenerateContentConfig(temperature=0.5),
)
print(resp.text)  # parse JSON, then back-project (y, x) through camera intrinsics

Back-project the returned 2D points with a depth image and the camera’s pinhole model:

X = (u − c_x) · Z / f_x
Y = (v − c_y) · Z / f_y
P_base = T_base←cam · [X, Y, Z, 1]^T

Open Alternatives

Model Params Action decoding Fine-tune hardware Notes
SmolVLA ~450M Flow-matching expert Single consumer GPU Native in LeRobot; targets SO-101
OpenVLA 7B Autoregressive tokens 1–8 A100-class GPUs (LoRA helps) Trained on Open X-Embodiment
Octo 27M / 93M Diffusion head Single GPU Small, fast, weaker language grounding
RDT-1B 1.2B Diffusion transformer Multi-GPU Bimanual focus

Choose these when you need to inspect and modify every layer, run on a budget, or publish reproducible research.

Head-to-Head Comparison

Criterion GR00T N1.7 π0.5 (open) Gemini Robotics 2 SmolVLA
Access Open, Early Access Open (Apache-2.0) ER 2 API; VLA gated Open
Embodiment focus Humanoid Arms, mobile manipulators Arms to humanoids Low-cost arms
Architecture Reasoning VLM + DiT VLM + flow-matching expert VLM planner + VLA Small VLM + flow expert
Edge deployment Jetson Thor Desktop GPU (RTX-class) On-Device 2 (gated) Laptop / Jetson-class
Fine-tuning cost Medium Medium Not available publicly Low
Commercial use Yes Yes (check checkpoint terms) Partner agreement Yes

The Control-Loop Math You Cannot Skip

Inference latency, not model quality, breaks most first deployments. Let f be the control rate, H the chunk length, and T_inf the end-to-end inference time.

Stale actions while inferring:   s = ceil(T_inf · f)
Executable actions per chunk:    H_exec = H − s
Constraint for continuous motion: H_exec ≥ 1   →   H / f > T_inf

Example: f = 50 Hz, H = 50, T_inf = 120 ms.
s = ceil(0.12 × 50) = 6 stale steps, so you can execute up to 44 actions per chunk. Run inference asynchronously, and blend the old and new chunks over the overlap window to avoid velocity discontinuities:

a_blend[i] = (1 − w_i) · a_old[i] + w_i · a_new[i],   w_i = i / s,  i = 0..s

Re-planning every 10–25 steps (a receding horizon) trades smoothness against reactivity. Shorter intervals recover from slips faster.

Practical Workflow 1: Serve π0.5 to a ROS 2 Robot

Run the policy server on a GPU workstation and a thin rclpy node on the robot. This keeps the real-time loop out of the Python inference process.

# Workstation: serve a pretrained openpi checkpoint over WebSocket
git clone --recurse-submodules https://github.com/Physical-Intelligence/openpi
cd openpi && uv sync
uv run scripts/serve_policy.py policy:checkpoint \
  --policy.config=pi05_droid \
  --policy.dir=gs://openpi-assets/checkpoints/pi05_droid
# ROS 2 node: queries the policy and publishes joint targets at 50 Hz
import rclpy, numpy as np
from rclpy.node import Node
from sensor_msgs.msg import JointState
from openpi_client import websocket_client_policy

class PiBridge(Node):
    def __init__(self):
        super().__init__("pi_bridge")
        self.policy = websocket_client_policy.WebsocketClientPolicy("192.168.1.50", 8000)
        self.pub = self.create_publisher(JointState, "/joint_targets", 10)
        self.chunk, self.idx, self.REPLAN = None, 0, 15   # re-plan every 15 steps
        self.create_timer(1 / 50.0, self.tick)             # 50 Hz control timer

    def get_obs(self):
        # Fill from your camera and joint topics. Key names must match the checkpoint config.
        return {
            "observation/exterior_image_1_left": self.ext_img,   # HxWx3 uint8
            "observation/wrist_image_left": self.wrist_img,      # HxWx3 uint8
            "observation/joint_position": self.q,                # (7,) float32
            "observation/gripper_position": self.g,              # (1,) float32
            "prompt": "put the mug in the bin",
        }

    def tick(self):
        if self.chunk is None or self.idx >= self.REPLAN:
            self.chunk = self.policy.infer(self.get_obs())["actions"]  # (H, action_dim)
            self.idx = 0
        a = self.chunk[self.idx]; self.idx += 1
        msg = JointState(); msg.position = a[:7].tolist()   # joint targets (rad)
        self.pub.publish(msg)

def main():
    rclpy.init(); rclpy.spin(PiBridge())

The synchronous infer call blocks the timer while it runs. For production, move it to a worker thread and apply the blending formula above.

Practical Workflow 2: Fine-Tune SmolVLA on an SO-101 Dataset

pip install lerobot
# Train from the pretrained base on your recorded teleoperation dataset
lerobot-train \
  --policy.path=lerobot/smolvla_base \
  --dataset.repo_id=YOUR_HF_USER/so101_pick_place \
  --batch_size=64 \
  --steps=20000 \
  --output_dir=outputs/smolvla_so101 \
  --policy.device=cuda

Record 50 or more clean demonstrations per task before training. Keep camera placement identical between data collection and deployment. Camera shift is the most common cause of poor transfer for every model in this guide.

Which Model Should You Pick?

Scenario Pick Why
Humanoid with Jetson Thor GR00T N1.7 Edge path, whole-body focus, commercial license
Dexterous tabletop arm, own data π0.5 via openpi Strong open checkpoints, proven fine-tunes
Multi-step task planning above any policy Gemini Robotics ER 2 Public API, tool orchestration, multi-robot coordination
Budget research on SO-101 SmolVLA Low fine-tuning cost, LeRobot-native
Full architectural control OpenVLA / Octo Fully open, simple to modify

A hybrid works well: ER 2 decomposes the task into subgoals, and a local policy (GR00T, π0.5, or SmolVLA) executes each subgoal. Each layer is replaceable.

Evaluation Pitfalls

  • Success rate without trial counts hides variance. Run at least 20 trials per task.
  • Retrieval versus generalization. Reviewers of π0.7 point out that separating true skill recombination from nearest-neighbor recall is now the central evaluation problem.
  • Latency measured on the server only. Include image encoding, network transfer, and ROS message serialization.
  • Simulation-to-real gaps. Validate in Isaac Sim or MuJoCo first, then budget hardware time for friction and calibration errors.

FAQ

What is a robot foundation model?

A robot foundation model is a large pretrained network that maps images, robot state, and language instructions to actions across many tasks and robot bodies. Most are vision-language-action (VLA) models built on a pretrained vision-language backbone with an action head.

Is NVIDIA GR00T open source?

The GR00T N-series models are open and commercially licensed, distributed through Hugging Face. The latest, N1.7, is in Early Access. Check the license on each checkpoint before commercial use.

Can I run Gemini Robotics on my own robot?

Partly. Gemini Robotics ER 2 is available through the Gemini API and Google AI Studio. The full-body VLA is in private preview, and On-Device 2 is limited to trusted testers.

What is the best open-source VLA for a low-cost robot arm?

SmolVLA is the practical starting point. It is small, integrated into LeRobot, and fine-tunes on a single GPU. Move to π0.5 via openpi when you need stronger dexterity and have a larger GPU.

Hot this week

The State of Robotics in 2026: 10 Biggest Developments

The 10 biggest robotics developments of 2026: whole-body VLA models, humanoid safety, ROS 2 Lyrical Luth, and Jetson Thor. Get the data and code.

EU Machinery Regulation 2027: What Robot Builders Need to Know

Building robots for the EU? Regulation (EU) 2023/1230 applies from 20 Jan 2027. Get the cybersecurity, AI, and CE marking checklist now.

ISO 10218:2025 Explained: The New Industrial Robot Safety Standard

ISO 10218:2025 rewrites industrial robot safety: Class I/II robots, built-in cobot limits, cybersecurity. Get the checklist and ROS2 code. Read now.

NVIDIA Jetson Orin Nano, AGX Orin, and Thor: Which One for Your Robot?

Jetson Orin Nano vs AGX Orin vs Thor: compare TOPS, memory bandwidth, power, and price to pick the right robot compute. Read the guide.

ROS 2 Distributions Explained: Humble, Jazzy, Kilted, and Lyrical (Which to Use)

Compare ROS 2 Humble, Jazzy, Kilted, and Lyrical by EOL date, platform support, and features. Pick the right distro for your robot. Read the guide.

Topics

The State of Robotics in 2026: 10 Biggest Developments

The 10 biggest robotics developments of 2026: whole-body VLA models, humanoid safety, ROS 2 Lyrical Luth, and Jetson Thor. Get the data and code.

EU Machinery Regulation 2027: What Robot Builders Need to Know

Building robots for the EU? Regulation (EU) 2023/1230 applies from 20 Jan 2027. Get the cybersecurity, AI, and CE marking checklist now.

ISO 10218:2025 Explained: The New Industrial Robot Safety Standard

ISO 10218:2025 rewrites industrial robot safety: Class I/II robots, built-in cobot limits, cybersecurity. Get the checklist and ROS2 code. Read now.

NVIDIA Jetson Orin Nano, AGX Orin, and Thor: Which One for Your Robot?

Jetson Orin Nano vs AGX Orin vs Thor: compare TOPS, memory bandwidth, power, and price to pick the right robot compute. Read the guide.

ROS 2 Distributions Explained: Humble, Jazzy, Kilted, and Lyrical (Which to Use)

Compare ROS 2 Humble, Jazzy, Kilted, and Lyrical by EOL date, platform support, and features. Pick the right distro for your robot. Read the guide.

Build a Low-Cost AI Robot Arm With SO-101 and LeRobot

Build an SO-101 robot arm under $250, calibrate it, record demos, and train an ACT policy with LeRobot. Follow the full guide and start building.

How Much Does a Humanoid Robot Cost? Prices, Subscriptions, and Hidden Costs

Humanoid robot cost in 2026: prices from $4,900, $499/mo subscriptions, and hidden fees. See the full TCO breakdown and compare models now.

Humanoid Robot Companies Compared: Tesla, Figure, Boston Dynamics, Unitree, 1X, and More

Compare Tesla Optimus, Figure 03, Atlas, Unitree & 1X NEO on specs, price, control stacks, and availability. Pick the right humanoid today.

Related Articles

Popular Categories