Build a Low-Cost AI Robot Arm With SO-101 and LeRobot

The SO-101 is an open-source 6-DoF desktop arm designed for one purpose: collecting manipulation data cheaply and training neural policies on it. Pair it with LeRobot, Hugging Face’s PyTorch robotics library, and you get a full pipeline from teleoperation to autonomous pick-and-place for the price of a mid-range GPU fan. This guide covers the hardware, wiring, kinematics, calibration, data recording, ACT training, and deployment. Command-line flags change between LeRobot releases, so verify them against lerobot-record --help on your installed version.

Quick Takeaways

  • Cost: A leader-follower pair runs roughly $200-$300 in parts, plus a USB webcam or two.
  • Hardware: Follower arm uses 6x Feetech STS3215 serial bus servos (7.4V or 12V variants). The leader arm uses the same servos with different gear ratios.
  • Software: lerobot handles calibration, teleoperation, dataset recording, and policy inference through one Python API.
  • Result: 50 clean demonstrations typically produce a working ACT policy for single-object pick-and-place in a fixed scene.
Spec SO-100 (predecessor) SO-101
Degrees of freedom 5 + gripper 5 + gripper
Servo STS3215 STS3215 (improved wiring)
Wiring Daisy-chain, tight routing Cleaner cable path, no gear-slip prone parts
Leader gear ratios Single ratio Mixed ratios per joint for lighter feel
Assembly Moderate Easier, fewer fasteners
Estimated cost ~$110/arm ~$110-$130/arm

Bill of Materials

Part Qty Notes Approx. Cost
STS3215 servo (follower) 6 12V, 1:345 gearing $15 each
STS3215 servo (leader) 6 Mixed ratios (1:191, 1:345, 1:147) $15 each
Waveshare serial bus servo adapter 2 USB-to-TTL for the Feetech bus $6 each
5V 3A PSU (leader) / 12V 5A PSU (follower) 1 each Match your servo voltage $15 total
3D printed parts 2 sets PLA or PETG, 0.2mm layers, 15-20% infill $10-$20 filament
USB webcam (720p+) 1-2 Fixed mount plus optional wrist cam $10-$30
Table clamps 2 Arm base stability matters for repeatability $8

Match your power supply to the servo variant. Running a 7.4V servo from a 12V supply destroys the board.

Kinematic Foundations

The follower is a serial chain: shoulder pan, shoulder lift, elbow flex, wrist flex, wrist roll, and gripper. You do not need inverse kinematics for imitation learning because the policy operates in joint space. Still, understanding the forward model helps with debugging and safety limits.

For a planar slice (shoulder lift, elbow, wrist flex) with link lengths L1, L2, L3 and joint angles θ1, θ2, θ3:

x = L1*cos(θ1) + L2*cos(θ1 + θ2) + L3*cos(θ1 + θ2 + θ3)
z = L1*sin(θ1) + L2*sin(θ1 + θ2) + L3*sin(θ1 + θ2 + θ3)

The shoulder pan θ0 rotates this plane about the vertical axis:

X = x * cos(θ0)
Y = x * sin(θ0)
Z = z

A servo’s raw position runs 0-4095 ticks per revolution, so one tick is:

resolution = 360° / 4096 ≈ 0.088°

At a 0.15 m reach, one tick at the shoulder translates to about 0.23 mm of end-effector motion. That bounds your theoretical precision. Backlash and 3D-printed compliance in practice push real repeatability to roughly 2-5 mm.

Hardware Assembly and Wiring

Step 1: Set Servo IDs Before Assembly

Every servo on a bus needs a unique ID. Set them one at a time, with only that servo connected to the adapter. LeRobot provides a helper for this:

# Find the serial port of each adapter (run before and after plugging in)
lerobot-find-port

# Assign IDs 1-6 interactively; the script prompts you to connect each motor
lerobot-setup-motors --robot.type=so101_follower --robot.port=/dev/ttyACM0
lerobot-setup-motors --teleop.type=so101_leader --teleop.port=/dev/ttyACM1

Label each servo with tape immediately. Mixed-up IDs are the most common build error.

Step 2: Joint-to-Servo Map

ID Joint Leader Gear Ratio
1 Shoulder pan 1:191
2 Shoulder lift 1:345
3 Elbow flex 1:191
4 Wrist flex 1:147
5 Wrist roll 1:147
6 Gripper 1:147

Verify this table against the current SO-101 assembly documentation, since printed-part revisions occasionally change the recommended ratios.

Step 3: Power and Bus Topology

The Feetech bus is a half-duplex TTL line. Daisy-chain all six servos with the 3-pin cables, then connect the first servo to the Waveshare adapter. Power the bus through the adapter’s DC jack, not from USB. Keep the 12V supply on the follower and the 5V supply on the leader if you use 7.4V-class leader servos.

Software Environment

# Create an isolated environment (Python 3.10+ per current LeRobot requirements)
conda create -y -n lerobot python=3.10
conda activate lerobot

# Install ffmpeg for video encoding of recorded datasets
conda install ffmpeg -c conda-forge

# Install LeRobot with Feetech motor support
git clone https://github.com/huggingface/lerobot.git
cd lerobot
pip install -e ".[feetech]"

Log in to Hugging Face so you can push datasets and policies:

huggingface-cli login --token <YOUR_WRITE_TOKEN> --add-to-git-credential

Calibration

Calibration maps raw servo ticks to a normalized range so the leader and follower agree on pose. Skipping or rushing this step ruins every downstream recording.

lerobot-calibrate \
  --robot.type=so101_follower \
  --robot.port=/dev/ttyACM0 \
  --robot.id=my_follower_arm

lerobot-calibrate \
  --teleop.type=so101_leader \
  --teleop.port=/dev/ttyACM1 \
  --teleop.id=my_leader_arm

The procedure has two phases. First, move the arm to the middle of its range and confirm. Second, move each joint through its full range of motion while the script records min/max ticks. Rotate the wrist roll through its full travel too.

Calibration files save under ~/.cache/huggingface/lerobot/calibration/. Keep the same id strings every time or LeRobot will prompt you to recalibrate.

Teleoperation and Camera Setup

Test teleoperation before recording anything. The follower should mirror the leader with no visible lag.

lerobot-teleoperate \
  --robot.type=so101_follower \
  --robot.port=/dev/ttyACM0 \
  --robot.id=my_follower_arm \
  --teleop.type=so101_leader \
  --teleop.port=/dev/ttyACM1 \
  --teleop.id=my_leader_arm

Add cameras by index, then confirm which index maps to which device:

lerobot-find-cameras opencv

A minimal camera config for a front view plus wrist view looks like this:

--robot.cameras="{ front: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}, wrist: {type: opencv, index_or_path: 1, width: 640, height: 480, fps: 30}}"
Camera Choice Latency Cost Best For
Fixed overhead webcam Low $10-$25 Scene layout, object location
Wrist-mounted camera Low $15-$30 Grasp alignment, fine approach
Depth camera (RealSense) Medium $150+ Not required for ACT baselines

Use two cameras if you can. The wrist view noticeably improves grasp success.

Recording a Dataset

A good dataset beats a clever model. Record demonstrations that are consistent in speed, approach direction, and grasp style.

lerobot-record \
  --robot.type=so101_follower \
  --robot.port=/dev/ttyACM0 \
  --robot.id=my_follower_arm \
  --robot.cameras="{ front: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}}" \
  --teleop.type=so101_leader \
  --teleop.port=/dev/ttyACM1 \
  --teleop.id=my_leader_arm \
  --dataset.repo_id=${HF_USER}/so101_pick_cube \
  --dataset.num_episodes=50 \
  --dataset.single_task="Pick up the cube and place it in the bin" \
  --dataset.episode_time_s=20 \
  --dataset.reset_time_s=10 \
  --display_data=true

Keyboard shortcuts during recording: Right arrow ends the current episode early, Left arrow re-records it, Esc stops the session.

Data Quality Rules

Rule Reason
Keep the camera fixed between sessions Policies overfit to viewpoint
Vary object start position within a defined zone Prevents memorizing a single trajectory
Match lighting between recording and deployment Vision encoders are sensitive to exposure shifts
Discard failed or hesitant episodes ACT imitates mistakes faithfully
Aim for 50+ episodes per task Fewer episodes often yield jittery policies

Policy Architecture: How ACT Works

ACT (Action Chunking with Transformers) predicts a chunk of k future actions from the current observation, instead of one action at a time. This reduces compounding error because the policy commits to short trajectories.

The training objective combines a reconstruction loss with a KL regularizer from a conditional VAE:

L = L1(a_pred, a_true) + β * KL( q(z | a, o) || N(0, I) )

where a is the action chunk, o is the observation (images plus joint state), z is the latent style variable, and β is the KL weight (commonly 10 in the original paper).

At inference, temporal ensembling smooths overlapping chunks. For a chunk size k = 100 at 30 Hz, the policy plans about 3.3 seconds ahead:

planning horizon = k / fps = 100 / 30 ≈ 3.33 s
Policy Strength Weakness Typical Data Need
ACT Stable, fast training, good for fixed scenes Weak language grounding 50 episodes
Diffusion Policy Handles multimodal behavior Slower inference 100+ episodes
SmolVLA Language-conditioned, multi-task Heavier, more data Larger multi-task sets

Training the Policy

lerobot-train \
  --dataset.repo_id=${HF_USER}/so101_pick_cube \
  --policy.type=act \
  --output_dir=outputs/train/act_so101_pick_cube \
  --job_name=act_so101_pick_cube \
  --policy.device=cuda \
  --wandb.enable=true \
  --policy.repo_id=${HF_USER}/act_so101_pick_cube

On a single consumer GPU with 8-12 GB VRAM, expect several hours for a 50-episode dataset at default step counts. Apple Silicon works with --policy.device=mps, and CPU training is possible but impractical. Watch the L1 loss curve in Weights & Biases. A plateau well above zero with erratic rollouts usually signals inconsistent demonstrations rather than a model problem.

Evaluating on the Real Arm

Run the policy while recording an evaluation dataset, which doubles as a log of every rollout:

lerobot-record \
  --robot.type=so101_follower \
  --robot.port=/dev/ttyACM0 \
  --robot.id=my_follower_arm \
  --robot.cameras="{ front: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}}" \
  --dataset.repo_id=${HF_USER}/eval_act_so101_pick_cube \
  --dataset.single_task="Pick up the cube and place it in the bin" \
  --dataset.num_episodes=10 \
  --policy.path=${HF_USER}/act_so101_pick_cube

Keep a hand near the power supply for the first runs. A policy that has never seen your emergency situations will not stop on its own.

Real-World Workflow: Pick-and-Place Cube Sorting

A practical sequence from unboxed parts to a working policy:

  1. Day 1: Print parts, set servo IDs, assemble both arms, and run teleoperation.
  2. Day 2: Mount cameras on a rigid frame, calibrate, and rehearse the task 10 times by hand.
  3. Day 3: Record 50 episodes with the object start position randomized inside a 10 cm square.
  4. Day 4: Train ACT, evaluate 10 rollouts, and log success rate.
  5. Day 5: Record 20 more episodes targeting observed failures (for example, grasps near the zone edge), then retrain.

Track a simple success metric:

success rate = successful_rollouts / total_rollouts

A baseline of 60-80% on a constrained scene is realistic for a first iteration. Scores climb with targeted data, not with longer training.

Troubleshooting Table

Symptom Likely Cause Fix
Servo not detected Duplicate ID or loose bus cable Re-run lerobot-setup-motors, reseat 3-pin cables
Follower jitters in teleoperation Voltage sag or poor calibration Check PSU rating, recalibrate
Gripper crushes objects No torque limit on gripper Lower gripper max torque in the robot config
Policy moves then freezes Camera index swapped between record and eval Confirm the same camera names and indices
Arm drifts after reboot Calibration id mismatch Use identical --robot.id strings
Overheating servos Holding load at stall Add rest periods, reduce episode duration

Safety and Maintenance

  • Clamp the base to the table. Dynamic motion tips an unclamped arm.
  • Print structural parts in PETG or add extra perimeters to joints that carry load.
  • Re-tighten set screws after the first hour of operation. Printed horns loosen.
  • Unplug power before touching wiring.

FAQ

How much does an SO-101 robot arm cost to build?

A single follower arm costs about $110-$130 in parts, and a leader-follower pair needed for teleoperation costs roughly $200-$300 including power supplies and filament. Cameras add $10-$60 depending on how many you use.

Can the SO-101 run without a leader arm?

Yes, but you need the leader to record imitation-learning demonstrations efficiently. Once a policy is trained, the follower runs autonomously with no leader attached.

How many demonstrations does an ACT policy need?

Roughly 50 consistent episodes handle a single-object, fixed-scene task. More varied scenes or objects need proportionally more data, and quality matters more than raw count.

Do I need a GPU to train a LeRobot policy?

A CUDA GPU with 8 GB or more of VRAM is strongly recommended. Apple Silicon (MPS) works for smaller runs, and CPU-only training is possible but very slow.

Hot this week

The State of Robotics in 2026: 10 Biggest Developments

The 10 biggest robotics developments of 2026: whole-body VLA models, humanoid safety, ROS 2 Lyrical Luth, and Jetson Thor. Get the data and code.

EU Machinery Regulation 2027: What Robot Builders Need to Know

Building robots for the EU? Regulation (EU) 2023/1230 applies from 20 Jan 2027. Get the cybersecurity, AI, and CE marking checklist now.

ISO 10218:2025 Explained: The New Industrial Robot Safety Standard

ISO 10218:2025 rewrites industrial robot safety: Class I/II robots, built-in cobot limits, cybersecurity. Get the checklist and ROS2 code. Read now.

NVIDIA Jetson Orin Nano, AGX Orin, and Thor: Which One for Your Robot?

Jetson Orin Nano vs AGX Orin vs Thor: compare TOPS, memory bandwidth, power, and price to pick the right robot compute. Read the guide.

ROS 2 Distributions Explained: Humble, Jazzy, Kilted, and Lyrical (Which to Use)

Compare ROS 2 Humble, Jazzy, Kilted, and Lyrical by EOL date, platform support, and features. Pick the right distro for your robot. Read the guide.

Topics

The State of Robotics in 2026: 10 Biggest Developments

The 10 biggest robotics developments of 2026: whole-body VLA models, humanoid safety, ROS 2 Lyrical Luth, and Jetson Thor. Get the data and code.

EU Machinery Regulation 2027: What Robot Builders Need to Know

Building robots for the EU? Regulation (EU) 2023/1230 applies from 20 Jan 2027. Get the cybersecurity, AI, and CE marking checklist now.

ISO 10218:2025 Explained: The New Industrial Robot Safety Standard

ISO 10218:2025 rewrites industrial robot safety: Class I/II robots, built-in cobot limits, cybersecurity. Get the checklist and ROS2 code. Read now.

NVIDIA Jetson Orin Nano, AGX Orin, and Thor: Which One for Your Robot?

Jetson Orin Nano vs AGX Orin vs Thor: compare TOPS, memory bandwidth, power, and price to pick the right robot compute. Read the guide.

ROS 2 Distributions Explained: Humble, Jazzy, Kilted, and Lyrical (Which to Use)

Compare ROS 2 Humble, Jazzy, Kilted, and Lyrical by EOL date, platform support, and features. Pick the right distro for your robot. Read the guide.

Robot Foundation Models: GR00T, pi, Gemini Robotics, and Open Alternatives Compared

Compare robot foundation models: NVIDIA GR00T, Physical Intelligence π, Gemini Robotics 2, and open VLAs. Get latency math, code, and a pick guide.

How Much Does a Humanoid Robot Cost? Prices, Subscriptions, and Hidden Costs

Humanoid robot cost in 2026: prices from $4,900, $499/mo subscriptions, and hidden fees. See the full TCO breakdown and compare models now.

Humanoid Robot Companies Compared: Tesla, Figure, Boston Dynamics, Unitree, 1X, and More

Compare Tesla Optimus, Figure 03, Atlas, Unitree & 1X NEO on specs, price, control stacks, and availability. Pick the right humanoid today.

Related Articles

Popular Categories