The SO-101 is an open-source 6-DoF desktop arm designed for one purpose: collecting manipulation data cheaply and training neural policies on it. Pair it with LeRobot, Hugging Face’s PyTorch robotics library, and you get a full pipeline from teleoperation to autonomous pick-and-place for the price of a mid-range GPU fan. This guide covers the hardware, wiring, kinematics, calibration, data recording, ACT training, and deployment. Command-line flags change between LeRobot releases, so verify them against lerobot-record --help on your installed version.
Quick Takeaways
- Cost: A leader-follower pair runs roughly $200-$300 in parts, plus a USB webcam or two.
- Hardware: Follower arm uses 6x Feetech STS3215 serial bus servos (7.4V or 12V variants). The leader arm uses the same servos with different gear ratios.
- Software:
lerobothandles calibration, teleoperation, dataset recording, and policy inference through one Python API. - Result: 50 clean demonstrations typically produce a working ACT policy for single-object pick-and-place in a fixed scene.
| Spec | SO-100 (predecessor) | SO-101 |
|---|---|---|
| Degrees of freedom | 5 + gripper | 5 + gripper |
| Servo | STS3215 | STS3215 (improved wiring) |
| Wiring | Daisy-chain, tight routing | Cleaner cable path, no gear-slip prone parts |
| Leader gear ratios | Single ratio | Mixed ratios per joint for lighter feel |
| Assembly | Moderate | Easier, fewer fasteners |
| Estimated cost | ~$110/arm | ~$110-$130/arm |
Bill of Materials
| Part | Qty | Notes | Approx. Cost |
|---|---|---|---|
| STS3215 servo (follower) | 6 | 12V, 1:345 gearing | $15 each |
| STS3215 servo (leader) | 6 | Mixed ratios (1:191, 1:345, 1:147) | $15 each |
| Waveshare serial bus servo adapter | 2 | USB-to-TTL for the Feetech bus | $6 each |
| 5V 3A PSU (leader) / 12V 5A PSU (follower) | 1 each | Match your servo voltage | $15 total |
| 3D printed parts | 2 sets | PLA or PETG, 0.2mm layers, 15-20% infill | $10-$20 filament |
| USB webcam (720p+) | 1-2 | Fixed mount plus optional wrist cam | $10-$30 |
| Table clamps | 2 | Arm base stability matters for repeatability | $8 |
Match your power supply to the servo variant. Running a 7.4V servo from a 12V supply destroys the board.
Kinematic Foundations
The follower is a serial chain: shoulder pan, shoulder lift, elbow flex, wrist flex, wrist roll, and gripper. You do not need inverse kinematics for imitation learning because the policy operates in joint space. Still, understanding the forward model helps with debugging and safety limits.
For a planar slice (shoulder lift, elbow, wrist flex) with link lengths L1, L2, L3 and joint angles θ1, θ2, θ3:
x = L1*cos(θ1) + L2*cos(θ1 + θ2) + L3*cos(θ1 + θ2 + θ3)
z = L1*sin(θ1) + L2*sin(θ1 + θ2) + L3*sin(θ1 + θ2 + θ3)
The shoulder pan θ0 rotates this plane about the vertical axis:
X = x * cos(θ0)
Y = x * sin(θ0)
Z = z
A servo’s raw position runs 0-4095 ticks per revolution, so one tick is:
resolution = 360° / 4096 ≈ 0.088°
At a 0.15 m reach, one tick at the shoulder translates to about 0.23 mm of end-effector motion. That bounds your theoretical precision. Backlash and 3D-printed compliance in practice push real repeatability to roughly 2-5 mm.
Hardware Assembly and Wiring
Step 1: Set Servo IDs Before Assembly
Every servo on a bus needs a unique ID. Set them one at a time, with only that servo connected to the adapter. LeRobot provides a helper for this:
# Find the serial port of each adapter (run before and after plugging in)
lerobot-find-port
# Assign IDs 1-6 interactively; the script prompts you to connect each motor
lerobot-setup-motors --robot.type=so101_follower --robot.port=/dev/ttyACM0
lerobot-setup-motors --teleop.type=so101_leader --teleop.port=/dev/ttyACM1
Label each servo with tape immediately. Mixed-up IDs are the most common build error.
Step 2: Joint-to-Servo Map
| ID | Joint | Leader Gear Ratio |
|---|---|---|
| 1 | Shoulder pan | 1:191 |
| 2 | Shoulder lift | 1:345 |
| 3 | Elbow flex | 1:191 |
| 4 | Wrist flex | 1:147 |
| 5 | Wrist roll | 1:147 |
| 6 | Gripper | 1:147 |
Verify this table against the current SO-101 assembly documentation, since printed-part revisions occasionally change the recommended ratios.
Step 3: Power and Bus Topology
The Feetech bus is a half-duplex TTL line. Daisy-chain all six servos with the 3-pin cables, then connect the first servo to the Waveshare adapter. Power the bus through the adapter’s DC jack, not from USB. Keep the 12V supply on the follower and the 5V supply on the leader if you use 7.4V-class leader servos.
Software Environment
# Create an isolated environment (Python 3.10+ per current LeRobot requirements)
conda create -y -n lerobot python=3.10
conda activate lerobot
# Install ffmpeg for video encoding of recorded datasets
conda install ffmpeg -c conda-forge
# Install LeRobot with Feetech motor support
git clone https://github.com/huggingface/lerobot.git
cd lerobot
pip install -e ".[feetech]"
Log in to Hugging Face so you can push datasets and policies:
huggingface-cli login --token <YOUR_WRITE_TOKEN> --add-to-git-credential
Calibration
Calibration maps raw servo ticks to a normalized range so the leader and follower agree on pose. Skipping or rushing this step ruins every downstream recording.
lerobot-calibrate \
--robot.type=so101_follower \
--robot.port=/dev/ttyACM0 \
--robot.id=my_follower_arm
lerobot-calibrate \
--teleop.type=so101_leader \
--teleop.port=/dev/ttyACM1 \
--teleop.id=my_leader_arm
The procedure has two phases. First, move the arm to the middle of its range and confirm. Second, move each joint through its full range of motion while the script records min/max ticks. Rotate the wrist roll through its full travel too.
Calibration files save under ~/.cache/huggingface/lerobot/calibration/. Keep the same id strings every time or LeRobot will prompt you to recalibrate.
Teleoperation and Camera Setup
Test teleoperation before recording anything. The follower should mirror the leader with no visible lag.
lerobot-teleoperate \
--robot.type=so101_follower \
--robot.port=/dev/ttyACM0 \
--robot.id=my_follower_arm \
--teleop.type=so101_leader \
--teleop.port=/dev/ttyACM1 \
--teleop.id=my_leader_arm
Add cameras by index, then confirm which index maps to which device:
lerobot-find-cameras opencv
A minimal camera config for a front view plus wrist view looks like this:
--robot.cameras="{ front: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}, wrist: {type: opencv, index_or_path: 1, width: 640, height: 480, fps: 30}}"
| Camera Choice | Latency | Cost | Best For |
|---|---|---|---|
| Fixed overhead webcam | Low | $10-$25 | Scene layout, object location |
| Wrist-mounted camera | Low | $15-$30 | Grasp alignment, fine approach |
| Depth camera (RealSense) | Medium | $150+ | Not required for ACT baselines |
Use two cameras if you can. The wrist view noticeably improves grasp success.
Recording a Dataset
A good dataset beats a clever model. Record demonstrations that are consistent in speed, approach direction, and grasp style.
lerobot-record \
--robot.type=so101_follower \
--robot.port=/dev/ttyACM0 \
--robot.id=my_follower_arm \
--robot.cameras="{ front: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}}" \
--teleop.type=so101_leader \
--teleop.port=/dev/ttyACM1 \
--teleop.id=my_leader_arm \
--dataset.repo_id=${HF_USER}/so101_pick_cube \
--dataset.num_episodes=50 \
--dataset.single_task="Pick up the cube and place it in the bin" \
--dataset.episode_time_s=20 \
--dataset.reset_time_s=10 \
--display_data=true
Keyboard shortcuts during recording: Right arrow ends the current episode early, Left arrow re-records it, Esc stops the session.
Data Quality Rules
| Rule | Reason |
|---|---|
| Keep the camera fixed between sessions | Policies overfit to viewpoint |
| Vary object start position within a defined zone | Prevents memorizing a single trajectory |
| Match lighting between recording and deployment | Vision encoders are sensitive to exposure shifts |
| Discard failed or hesitant episodes | ACT imitates mistakes faithfully |
| Aim for 50+ episodes per task | Fewer episodes often yield jittery policies |
Policy Architecture: How ACT Works
ACT (Action Chunking with Transformers) predicts a chunk of k future actions from the current observation, instead of one action at a time. This reduces compounding error because the policy commits to short trajectories.
The training objective combines a reconstruction loss with a KL regularizer from a conditional VAE:
L = L1(a_pred, a_true) + β * KL( q(z | a, o) || N(0, I) )
where a is the action chunk, o is the observation (images plus joint state), z is the latent style variable, and β is the KL weight (commonly 10 in the original paper).
At inference, temporal ensembling smooths overlapping chunks. For a chunk size k = 100 at 30 Hz, the policy plans about 3.3 seconds ahead:
planning horizon = k / fps = 100 / 30 ≈ 3.33 s
| Policy | Strength | Weakness | Typical Data Need |
|---|---|---|---|
| ACT | Stable, fast training, good for fixed scenes | Weak language grounding | 50 episodes |
| Diffusion Policy | Handles multimodal behavior | Slower inference | 100+ episodes |
| SmolVLA | Language-conditioned, multi-task | Heavier, more data | Larger multi-task sets |
Training the Policy
lerobot-train \
--dataset.repo_id=${HF_USER}/so101_pick_cube \
--policy.type=act \
--output_dir=outputs/train/act_so101_pick_cube \
--job_name=act_so101_pick_cube \
--policy.device=cuda \
--wandb.enable=true \
--policy.repo_id=${HF_USER}/act_so101_pick_cube
On a single consumer GPU with 8-12 GB VRAM, expect several hours for a 50-episode dataset at default step counts. Apple Silicon works with --policy.device=mps, and CPU training is possible but impractical. Watch the L1 loss curve in Weights & Biases. A plateau well above zero with erratic rollouts usually signals inconsistent demonstrations rather than a model problem.
Evaluating on the Real Arm
Run the policy while recording an evaluation dataset, which doubles as a log of every rollout:
lerobot-record \
--robot.type=so101_follower \
--robot.port=/dev/ttyACM0 \
--robot.id=my_follower_arm \
--robot.cameras="{ front: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}}" \
--dataset.repo_id=${HF_USER}/eval_act_so101_pick_cube \
--dataset.single_task="Pick up the cube and place it in the bin" \
--dataset.num_episodes=10 \
--policy.path=${HF_USER}/act_so101_pick_cube
Keep a hand near the power supply for the first runs. A policy that has never seen your emergency situations will not stop on its own.
Real-World Workflow: Pick-and-Place Cube Sorting
A practical sequence from unboxed parts to a working policy:
- Day 1: Print parts, set servo IDs, assemble both arms, and run teleoperation.
- Day 2: Mount cameras on a rigid frame, calibrate, and rehearse the task 10 times by hand.
- Day 3: Record 50 episodes with the object start position randomized inside a 10 cm square.
- Day 4: Train ACT, evaluate 10 rollouts, and log success rate.
- Day 5: Record 20 more episodes targeting observed failures (for example, grasps near the zone edge), then retrain.
Track a simple success metric:
success rate = successful_rollouts / total_rollouts
A baseline of 60-80% on a constrained scene is realistic for a first iteration. Scores climb with targeted data, not with longer training.
Troubleshooting Table
| Symptom | Likely Cause | Fix |
|---|---|---|
| Servo not detected | Duplicate ID or loose bus cable | Re-run lerobot-setup-motors, reseat 3-pin cables |
| Follower jitters in teleoperation | Voltage sag or poor calibration | Check PSU rating, recalibrate |
| Gripper crushes objects | No torque limit on gripper | Lower gripper max torque in the robot config |
| Policy moves then freezes | Camera index swapped between record and eval | Confirm the same camera names and indices |
| Arm drifts after reboot | Calibration id mismatch |
Use identical --robot.id strings |
| Overheating servos | Holding load at stall | Add rest periods, reduce episode duration |
Safety and Maintenance
- Clamp the base to the table. Dynamic motion tips an unclamped arm.
- Print structural parts in PETG or add extra perimeters to joints that carry load.
- Re-tighten set screws after the first hour of operation. Printed horns loosen.
- Unplug power before touching wiring.
FAQ
How much does an SO-101 robot arm cost to build?
A single follower arm costs about $110-$130 in parts, and a leader-follower pair needed for teleoperation costs roughly $200-$300 including power supplies and filament. Cameras add $10-$60 depending on how many you use.
Can the SO-101 run without a leader arm?
Yes, but you need the leader to record imitation-learning demonstrations efficiently. Once a policy is trained, the follower runs autonomously with no leader attached.
How many demonstrations does an ACT policy need?
Roughly 50 consistent episodes handle a single-object, fixed-scene task. More varied scenes or objects need proportionally more data, and quality matters more than raw count.
Do I need a GPU to train a LeRobot policy?
A CUDA GPU with 8 GB or more of VRAM is strongly recommended. Apple Silicon (MPS) works for smaller runs, and CPU-only training is possible but very slow.




