Your robot’s compute module sets the ceiling for perception rate, model size, sensor count, and battery life. Pick too small and your VSLAM pipeline drops frames. Pick too large and you burn 100 W of battery on a platform that spends most of its time idle. This guide compares the three Jetson tiers by the numbers that decide a build: memory bandwidth, usable precision, power envelope, and I/O.
Quick Takeaways
- Jetson Orin Nano (Super) fits hobby and light-duty robots: 8GB, up to 67 sparse INT8 TOPS, 7W to 25W. Best for one or two cameras and small detection models.
- Jetson AGX Orin fits multi-sensor mobile robots and AMRs: up to 64GB, 275 sparse INT8 TOPS, 15W to 60W. It is the proven middle tier with the deepest software maturity.
- Jetson AGX Thor fits humanoids, manipulators, and robots running large vision-language-action models: 128GB, 2070 sparse FP4 TFLOPS, roughly 40W to 130W.
- Memory bandwidth, not TOPS, limits transformer inference on the edge. Compare GB/s first.
| Module | AI Compute (vendor peak) | Memory | Bandwidth | Power Range | Launch Dev Kit Price* |
|---|---|---|---|---|---|
| Orin Nano 8GB (Super) | 67 TOPS (INT8, sparse) | 8GB LPDDR5 | 102 GB/s | 7W-25W | ~$249 |
| AGX Orin 64GB | 275 TOPS (INT8, sparse) | 64GB LPDDR5 | 204.8 GB/s | 15W-60W | ~$1,999 |
| AGX Thor (T5000) | 2070 TFLOPS (FP4, sparse) | 128GB LPDDR5X | 273 GB/s | ~40W-130W | ~$3,499 |
*Launch-era dev kit pricing. Verify current pricing and availability with NVIDIA or a distributor before budgeting.
Reading the Spec Sheet Without Getting Fooled
NVIDIA quotes peak throughput at different precisions per generation. The headline numbers are not directly comparable.
- Orin generation: INT8 with 2:4 structured sparsity.
- Thor generation: FP4 with sparsity, enabled by the Blackwell tensor cores.
Thor’s 2070 figure does not mean it is 7.5x faster than AGX Orin on every workload. That ratio holds only for models quantized to FP4 and pruned for sparsity. A dense FP16 YOLO model sees a much smaller gain. Benchmark your own model before trusting any ratio.
Architecture at a Glance
| Feature | Orin Nano 8GB | AGX Orin 64GB | AGX Thor |
|---|---|---|---|
| GPU Architecture | Ampere | Ampere | Blackwell |
| CUDA Cores | 1024 | 2048 | 2560 |
| Tensor Cores | 32 | 64 | 96 (5th gen) |
| CPU | 6x Cortex-A78AE | 12x Cortex-A78AE | 14x Neoverse-V3AE |
| Safety Island Features | No | Yes (functional safety capable) | Yes |
| Network I/O | 1x GbE | 1x 10GbE | 4x25GbE via QSFP28 |
| Software Stack | JetPack 6.x | JetPack 6.x | JetPack 7.x |
Thor’s QSFP28 networking matters for sensor-heavy platforms. A single 25GbE lane carries several uncompressed camera streams, which removes a common PCIe/USB bottleneck on humanoids and large AMRs.
The Real Bottleneck: Memory Bandwidth
Autoregressive decoding for LLMs and VLAs is memory-bound. Each generated token reads nearly all model weights once. The theoretical ceiling is:
tokens_per_second_max ≈ B / S
B = memory bandwidth (GB/s)
S = model size in memory (GB)
For an 8B-parameter model quantized to 4-bit (about 4.5 GB with overhead):
| Module | Bandwidth (B) | Theoretical Max | Realistic (~50-65%) |
|---|---|---|---|
| Orin Nano 8GB | 102 GB/s | ~22 tok/s | ~11-14 tok/s |
| AGX Orin 64GB | 204.8 GB/s | ~45 tok/s | ~23-29 tok/s |
| AGX Thor | 273 GB/s | ~60 tok/s | ~30-39 tok/s |
Two points follow:
- Thor’s bandwidth gain over AGX Orin is only about 1.33x. Its real advantage on large models comes from FP4/FP8 support (smaller weights, so a smaller S) and 128GB capacity (bigger models fit at all).
- On Orin Nano, an 8B model barely fits next to the OS, ROS 2 nodes, and camera buffers. Plan on models of 3B parameters or fewer there.
Perception Latency Budget
Treat the pipeline as a latency sum, not a TOPS contest:
L_total = L_sensor + L_preproc + L_infer + L_postproc + L_plan + L_actuate
For a 30 Hz perception loop: L_total <= 33.3 ms
For a 60 Hz loop: L_total <= 16.7 ms
On a camera-based AMR, a typical allocation looks like:
| Stage | Budget (30 Hz) |
|---|---|
| Sensor exposure + transport | 8 ms |
| Preprocess (resize, normalize on GPU) | 2 ms |
| Inference (TensorRT, FP16/INT8) | 10 ms |
| Postprocess (NMS, tracking) | 3 ms |
| Planning/costmap update | 7 ms |
| Margin | ~3 ms |
If inference alone exceeds its slot, move up a tier or quantize harder.
Orin Nano: The Entry Platform
The Jetson Orin Nano Super dev kit runs the same module as the earlier Orin Nano 8GB, but with a higher-clock software mode that lifts it to 67 sparse INT8 TOPS and 102 GB/s of bandwidth. Power modes span 7W to 25W.
Strengths
- Low cost and low power. Runs from a 12V to 19V supply.
- Full CUDA, TensorRT, and Isaac ROS support.
- Handles 1-2 cameras with detection, segmentation, or AprilTag pipelines.
Limits
- 8GB shared between CPU and GPU. Memory fills quickly with a ROS 2 stack plus Nav2 plus a vision model.
- 6 CPU cores constrain multi-process ROS 2 graphs.
- Limited camera lanes and I/O for multi-sensor builds.
Typical robots: line followers with vision, educational manipulators, small differential-drive rovers, drone companion computers.
AGX Orin: The Proven Workhorse
AGX Orin remains the balanced choice for production mobile robots. It offers up to 64GB of memory, 12 CPU cores, and mature drivers for GMSL cameras, CAN, and PCIe.
Strengths
- Runs a full Nav2 stack, 4-6 camera pipelines, and a mid-size detection model concurrently.
- Large ecosystem: Isaac ROS packages, community carrier boards, and years of field-tested reliability data.
- 15W to 60W lets you tune thermals for sealed enclosures.
Limits
- Large VLAs and 13B+ LLMs run slowly or do not fit.
- No Blackwell FP4 path.
Typical robots: warehouse AMRs, agricultural platforms, inspection robots, outdoor delivery vehicles.
AGX Thor: The Physical AI Platform
Jetson AGX Thor targets robots that run foundation models onboard: humanoids, dexterous manipulators, and generalist mobile manipulators. The Blackwell GPU supports FP4 and FP8 transformer inference, and the 128GB memory pool holds a VLA, a perception stack, and a planner at the same time.
Strengths
- Capacity for large multimodal models without cloud round-trips.
- 25GbE-class networking for sensor-rich platforms.
- Multi-Instance GPU (MIG) support to partition the GPU between workloads, such as isolating a safety-relevant perception task from a language model.
Limits
- Power draw reaches the 100W+ class. Battery sizing, heat sinking, and a regulated 12V to 24V supply path need real design effort.
- JetPack 7.x moves to a newer base OS and CUDA release. Check that your drivers, containers, and third-party libraries support it before committing.
- Highest cost of the three.
Typical robots: humanoids, mobile manipulators, multi-arm cells with onboard VLA policies.
Decision Matrix
| Your Requirement | Pick |
|---|---|
| Under $500 compute budget, 1-2 cameras | Orin Nano |
| Battery under 100 Wh, runtime over 4 hours | Orin Nano |
| 4+ cameras, Nav2, 3D LiDAR, mid-size detector | AGX Orin |
| Sealed enclosure, 60W thermal cap | AGX Orin |
| Onboard 7B-30B parameter model (VLM/VLA) | Thor |
| Humanoid or whole-body control with foundation model | Thor |
| Need proven, stable software today | AGX Orin |
Power Budget Math
Size the battery against the whole compute plus sensor load:
t_runtime (h) = (E_batt (Wh) x eta) / (P_compute + P_sensors + P_drive_avg)
eta ~ 0.85 (DC-DC conversion and usable battery depth)
Example: a 24V, 10Ah pack stores 240 Wh. With eta = 0.85, usable energy is 204 Wh. Take a mobile base drawing 60W average for drive and 12W for sensors:
| Compute Module (typical load) | Total Load | Runtime |
|---|---|---|
| Orin Nano (15W) | 87W | ~2.3 h |
| AGX Orin (40W) | 112W | ~1.8 h |
| Thor (90W) | 162W | ~1.3 h |
Drive power dominates here, so compute choice shifts runtime by under an hour. On a lightweight, low-drive-power platform such as a drone or a stationary arm, compute becomes the main load, and the gap widens sharply.
Configuring Power Modes
Lock a known power mode before benchmarking. Otherwise results vary run to run.
# Show the active power mode
sudo nvpmodel -q
# List all modes available on this module
sudo nvpmodel -q --verbose
# Set a mode by its index (read the index from the verbose listing)
sudo nvpmodel -m 0
# Pin CPU/GPU/EMC clocks to their maximums for repeatable benchmarks
sudo jetson_clocks
# Live telemetry: CPU/GPU load, memory, clocks, temperature, power rails
tegrastats --interval 500
Mode indices differ between modules and JetPack releases, so read them from the --verbose output instead of copying numbers from another platform.
ROS 2 Thermal Monitor Node
Thermal throttling silently ruins latency. This rclpy node publishes every SoC thermal zone as a standard diagnostic_msgs message, so you can alert before throttling begins.
#!/usr/bin/env python3
import glob
import rclpy
from rclpy.node import Node
from diagnostic_msgs.msg import DiagnosticArray, DiagnosticStatus, KeyValue
class JetsonThermalMonitor(Node):
def __init__(self):
super().__init__('jetson_thermal_monitor')
# Warn and error thresholds in degrees Celsius
self.declare_parameter('warn_c', 80.0)
self.declare_parameter('error_c', 95.0)
self.declare_parameter('period_s', 1.0)
self.pub = self.create_publisher(DiagnosticArray, '/diagnostics', 10)
period = self.get_parameter('period_s').value
self.timer = self.create_timer(period, self.publish_status)
def read_zones(self):
zones = {}
for zone in glob.glob('/sys/class/thermal/thermal_zone*'):
try:
with open(f'{zone}/type') as f:
name = f.read().strip()
with open(f'{zone}/temp') as f:
# The kernel reports millidegrees Celsius
zones[name] = int(f.read().strip()) / 1000.0
except (OSError, ValueError):
continue # Skip zones that are unreadable
return zones
def publish_status(self):
warn = self.get_parameter('warn_c').value
err = self.get_parameter('error_c').value
zones = self.read_zones()
status = DiagnosticStatus()
status.name = 'jetson/thermal'
status.hardware_id = 'jetson_soc'
peak = max(zones.values()) if zones else 0.0
# DiagnosticStatus levels: OK=0, WARN=1, ERROR=2
if peak >= err:
status.level = DiagnosticStatus.ERROR
elif peak >= warn:
status.level = DiagnosticStatus.WARN
else:
status.level = DiagnosticStatus.OK
status.message = f'peak {peak:.1f} C'
status.values = [KeyValue(key=k, value=f'{v:.1f}') for k, v in zones.items()]
msg = DiagnosticArray()
msg.header.stamp = self.get_clock().now().to_msg()
msg.status.append(status)
self.pub.publish(msg)
def main():
rclpy.init()
rclpy.spin(JetsonThermalMonitor())
if __name__ == '__main__':
main()
Run it with ros2 run <your_pkg> jetson_thermal_monitor and watch it with ros2 topic echo /diagnostics. Zone names differ between modules, which is why the node reads them dynamically.
Real-World Deployment Scenarios
Scenario 1: Indoor AMR on AGX Orin
A 150 kg warehouse cart carries four GMSL cameras, a 3D LiDAR, and a wheel-odometry stack.
- Run Nav2 with the costmap on CPU cores 0-5.
- Run a TensorRT INT8 detector on the GPU for pallet and person detection at 30 Hz.
- Feed LiDAR and camera depth into a
nvblox-style 3D reconstruction node for dynamic obstacle layers. - Cap the module at 40W to keep the sealed enclosure under 70°C ambient-inclusive.
- Keep motor control on a separate MCU (STM32 with micro-ROS) running a 1 kHz inner loop.
Scenario 2: Hobby Rover on Orin Nano
A 4-wheel rover with one RGB-D camera.
- Run RTAB-Map in a lightweight configuration for visual SLAM.
- Run a small YOLO variant in INT8 at 15-20 Hz.
- Limit swap use. Mount an NVMe drive and keep camera buffers small.
- Use the 15W mode for 5+ hours on a small pack.
Scenario 3: Mobile Manipulator on Thor
A single-arm platform runs a VLA policy for language-conditioned grasping.
- Quantize the VLA to FP8 or FP4 using TensorRT tooling and validate task success rate against the FP16 baseline.
- Partition GPU resources so perception and policy inference do not starve each other.
- Stream wrist and head cameras over the high-speed network path.
- Keep the joint-level control loop (1 kHz or faster) on a real-time MCU or EtherCAT master. Jetson is not a hard real-time controller.
Control Loop Placement Rule
Jetson modules run Linux, so scheduling jitter can reach milliseconds under load. Split responsibilities:
| Layer | Rate | Runs On |
|---|---|---|
| Joint/motor current and velocity loops | 1-20 kHz | MCU / servo drive |
| Whole-body or arm controller | 200-1000 Hz | Real-time MCU, or Jetson with PREEMPT_RT and CPU isolation |
| Perception and VLA inference | 10-60 Hz | Jetson GPU |
| Planning and navigation | 5-20 Hz | Jetson CPU |
FAQ
Is Jetson Thor worth it over AGX Orin?
Yes, if you run transformer-class models onboard. Thor’s 128GB memory, FP4/FP8 support, and faster networking suit VLAs and multimodal models. If your robot runs CNN detectors, SLAM, and Nav2, AGX Orin delivers similar results at lower power and cost.
Can Orin Nano run an LLM?
Yes, small ones. A 4-bit quantized model of 3B parameters or fewer fits in the 8GB pool and generates roughly 10 to 20 tokens per second, depending on the model and runtime. Larger models run out of memory once ROS 2 and camera buffers load.
Which Jetson is best for ROS 2?
All three run ROS 2 through JetPack and NVIDIA’s Isaac ROS packages. AGX Orin offers the most mature support today. Thor uses JetPack 7.x on a newer base OS, so confirm your ROS 2 distribution and package versions match before you start.
How much power does a Jetson Thor robot need?
Plan for up to about 130W for the module at full load, plus converter losses and cooling. Use a regulated supply with headroom, and size the battery with the runtime formula above.




