Environment setup eats hours of every data science project. A fresh pip install of pandas, scikit-learn, and PyTorch can stall for minutes, and a mismatched dependency can break a notebook that ran fine yesterday. uv, a Python package and project manager written in Rust by Astral, collapses pip, pip-tools, virtualenv, pyenv, and pipx into a single binary that resolves and installs packages roughly 10 to 100 times faster than pip.
Quick Takeaways
- uv replaces
pip,venv,pip-tools,pyenv, andpipxwith one tool and one CLI. - Its resolver and global cache make installs of heavy stacks (NumPy, pandas, PyTorch) dramatically faster, especially on repeat installs.
- The uv.lock file gives you cross-platform, reproducible environments, which is critical for ML experiments and production pipelines.
- You can adopt it gradually through the uv pip interface without rewriting existing workflows.
| Task | Traditional Tool | uv Command |
|---|---|---|
| Install Python 3.12 | pyenv install 3.12 |
uv python install 3.12 |
| Create environment | python -m venv .venv |
uv venv |
| Install packages | pip install pandas |
uv add pandas |
| Lock dependencies | pip-compile |
uv lock |
| Run a script | source .venv/bin/activate && python x.py |
uv run x.py |
| Run a CLI tool | pipx run ruff |
uvx ruff |
What Is uv and Why Does It Matter for Data Science?
uv is an extremely fast Python package and project manager. It uses a Rust-based dependency resolver, parallel downloads, and a global content-addressed cache. When a package already exists in the cache, uv links it into your environment instead of copying files, so a second environment with the same dependencies is created almost instantly.
Data science stacks amplify the problem uv solves. They include large binary wheels (NumPy, SciPy, PyTorch, TensorFlow), tight version constraints between libraries, and frequent environment switching between projects. A slow, non-deterministic installer costs real time at every one of those points.
How uv Differs from pip, Conda, and Poetry
| Feature | pip + venv | Conda | Poetry | uv |
|---|---|---|---|---|
| Install Speed | Slow | Slow to moderate | Moderate | Very fast |
| Python Version Management | No | Yes | No | Yes |
| Lock File | No (needs pip-tools) | Partial (explicit exports) | Yes | Yes (uv.lock) |
| Cross-Platform Lock | No | No | Yes | Yes |
| Non-Python Dependencies (CUDA, MKL) | No | Yes | No | No (wheels only) |
| Single Binary | No | No | No | Yes |
| Learning Curve | Low | Moderate | Moderate | Low |
Trade-off to know: uv installs Python wheels from PyPI. If your workflow depends on Conda-specific binaries (for example, system-level GDAL or custom CUDA toolkits from conda-forge), Conda remains the better choice. For most pandas, scikit-learn, XGBoost, and PyTorch workloads, PyPI wheels are sufficient.
Installing uv
Use the standalone installer on macOS or Linux:
# Install uv (macOS / Linux)
curl -LsSf https://astral.sh/uv/install.sh | sh
On Windows, use PowerShell:
# Install uv (Windows)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
You can also install it with pipx install uv or brew install uv. Verify the install:
uv --version
Keep it current with:
uv self update
Setting Up a Data Science Project with uv
Step 1: Initialize the Project
# Create a new project with a pyproject.toml and starter files
uv init churn-model
cd churn-model
This generates a pyproject.toml, a .python-version file, and a starter script. The pyproject.toml is the single source of truth for your dependencies.
Step 2: Pin the Python Version
# Download and pin Python 3.12 for this project
uv python install 3.12
uv python pin 3.12
uv downloads a managed Python build, so you do not need pyenv or a system-wide install. Teammates who run uv sync get the same interpreter version.
Step 3: Add Dependencies
# Core analytics stack
uv add pandas numpy scikit-learn matplotlib seaborn
# Development-only dependencies go in a separate group
uv add --dev jupyterlab ipykernel pytest ruff
Each uv add command updates pyproject.toml, resolves the full dependency tree, writes uv.lock, and installs everything into a project-local .venv.
Step 4: Run Code Without Activating Anything
# uv run auto-syncs the environment, then executes the command
uv run python train.py
uv run jupyter lab
uv run pytest
uv run checks that the environment matches the lock file before executing. This removes the classic “I forgot to activate the venv” error.
Reproducible Environments: pyproject.toml and uv.lock
Reproducibility separates a notebook experiment from a deployable model. uv gives you two files that work together.
| File | Purpose | Commit to Git? |
|---|---|---|
| pyproject.toml | Declares direct dependencies and version ranges | Yes |
| uv.lock | Records exact resolved versions and hashes for every package, across platforms | Yes |
| .python-version | Pins the interpreter version | Yes |
| .venv/ | Local environment, rebuilt from the lock file | No |
A typical pyproject.toml for a data project looks like this:
[project]
name = "churn-model"
version = "0.1.0"
requires-python = ">=3.12"
dependencies = [
"pandas>=2.2",
"numpy>=1.26",
"scikit-learn>=1.5",
"matplotlib>=3.9",
]
[dependency-groups]
dev = ["jupyterlab", "ipykernel", "pytest", "ruff"]
To rebuild the exact environment on another machine or in CI:
# Install exactly what uv.lock specifies
uv sync --frozen
Use –frozen when you want uv to fail rather than silently update the lock file. Use –locked to fail if the lock file is out of date relative to pyproject.toml.
To upgrade a single package without touching the rest:
uv lock --upgrade-package scikit-learn
uv sync
Migrating an Existing Project
You do not need to rewrite your workflow. The uv pip interface mirrors pip’s commands.
# Create a venv and install from an existing requirements file
uv venv
uv pip install -r requirements.txt
# Compile a pinned requirements file, like pip-compile
uv pip compile requirements.in -o requirements.txt
For a full migration, import your dependencies into a native uv project:
uv init
uv add -r requirements.txt
Replace pip install with uv pip install in your existing scripts first. Move to uv add and uv sync once your team is comfortable.
Using uv with Jupyter Notebooks
Data scientists live in notebooks, so the Jupyter workflow matters. There are two clean approaches.
Option 1: Launch Jupyter from the project environment.
uv add --dev jupyterlab ipykernel
uv run jupyter lab
Option 2: Run Jupyter in isolation and attach a project kernel. This keeps JupyterLab out of your project dependencies.
# Register the project's kernel
uv add --dev ipykernel
uv run ipython kernel install --user --name=churn-model
# Launch JupyterLab as a standalone tool with the project kernel available
uvx jupyter lab
For quick, throwaway analysis, run a temporary notebook with extra packages without modifying the project:
uv run --with pandas,matplotlib --with jupyter jupyter lab
Inline Script Dependencies (PEP 723)
For one-off data scripts, uv supports inline script metadata. Dependencies live inside the file, so the script is self-contained and shareable.
# /// script
# requires-python = ">=3.12"
# dependencies = [
# "pandas",
# "pyarrow",
# ]
# ///
import pandas as pd
# Read a Parquet file and print summary statistics
df = pd.read_parquet("sales.parquet")
print(df.describe())
Run it directly:
uv run summarize.py
uv builds an ephemeral environment, installs the declared packages, and executes the script. This pattern suits data engineers who share ad hoc ETL utilities.
Benchmarking uv Against pip
Measure the difference on your own machine. This script times a cold install of a typical analytics stack:
#!/usr/bin/env bash
# benchmark_install.sh: compare pip and uv install times
PACKAGES="pandas numpy scikit-learn scipy matplotlib seaborn xgboost"
# pip baseline
python -m venv .venv-pip
time .venv-pip/bin/pip install --no-cache-dir $PACKAGES
# uv with a cold cache
uv venv .venv-uv
time uv pip install --no-cache --python .venv-uv/bin/python $PACKAGES
# uv with a warm cache (the realistic daily case)
rm -rf .venv-uv && uv venv .venv-uv
time uv pip install --python .venv-uv/bin/python $PACKAGES
Expect the largest gains on the warm-cache run, where uv links files from its cache instead of downloading and unpacking them. Results vary with network speed, disk, and package mix, so treat published multipliers as directional and run the test yourself.
Real-World Use Cases
Predicting Customer Churn with a Reproducible Environment
A churn model typically uses pandas for feature engineering, scikit-learn for the classifier, and AUC-ROC for evaluation. Lock the stack so every retraining run uses identical library versions.
uv init churn-model && cd churn-model
uv add pandas scikit-learn
# train.py
import pandas as pd
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import roc_auc_score
from sklearn.model_selection import train_test_split
# Load and split the data (assumes a binary 'churned' column)
df = pd.read_csv("customers.csv")
X = df.drop(columns="churned")
y = df["churned"]
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, stratify=y, random_state=42 # stratify preserves class balance
)
# Train and evaluate
model = RandomForestClassifier(n_estimators=300, random_state=42, n_jobs=-1)
model.fit(X_train, y_train)
auc = roc_auc_score(y_test, model.predict_proba(X_test)[:, 1])
print(f"AUC-ROC: {auc:.3f}")
uv run train.py
Because uv.lock pins every transitive dependency, a retraining job six months later produces comparable results, free of silent library drift.
Containerizing a Model Service
Fast installs shorten Docker build times and CI runs. A minimal multi-layer Dockerfile:
FROM python:3.12-slim
# Copy the uv binary from the official image
COPY --from=ghcr.io/astral-sh/uv:latest /uv /usr/local/bin/uv
WORKDIR /app
# Install dependencies first to maximize layer caching
COPY pyproject.toml uv.lock ./
RUN uv sync --frozen --no-dev --no-install-project
# Copy the application code and finish the install
COPY . .
RUN uv sync --frozen --no-dev
CMD ["uv", "run", "python", "serve.py"]
Separating the dependency layer from the code layer means Docker only reinstalls packages when uv.lock changes.
Installing PyTorch with the Right CUDA Build
PyTorch ships separate wheels per accelerator. Configure an index in pyproject.toml so uv selects the right build:
[[tool.uv.index]]
name = "pytorch-cu124"
url = "https://download.pytorch.org/whl/cu124"
explicit = true
[tool.uv.sources]
torch = { index = "pytorch-cu124" }
uv add torch
Check the PyTorch and uv documentation for the index URL that matches your CUDA version, since available builds change over time.
Common Pitfalls and Fixes
| Problem | Cause | Fix |
|---|---|---|
| Package fails to build from source | No prebuilt wheel for your Python version | Pin an older Python with uv python pin 3.11 |
| Resolver conflict | Incompatible version constraints | Loosen bounds in pyproject.toml; run uv lock -v for details |
| Wrong CUDA build installed | Default PyPI index used | Configure an explicit index as shown above |
| Slow first install | Cold cache | Subsequent installs reuse the cache |
| Cache grows large | Many environments over time | Run uv cache clean or uv cache prune |
FAQ
Is uv faster than pip?
Yes. Astral’s benchmarks report speedups of 10 to 100 times over pip, driven by a Rust implementation, parallel downloads, and a global cache. Actual gains depend on your network, hardware, and whether the cache is warm.
Can uv replace Conda for data science?
For most Python-only workflows, yes. uv manages Python versions, environments, and locked dependencies using PyPI wheels. Conda remains the better option when you need non-Python binaries from conda-forge, such as system libraries or specific CUDA toolkits.
Does uv work with Jupyter notebooks?
Yes. Add ipykernel and jupyterlab as dev dependencies and launch with uv run jupyter lab. You can also register a project kernel and run Jupyter separately with uvx.
How do I migrate from requirements.txt to uv?
Run uv init followed by uv add -r requirements.txt to import dependencies into pyproject.toml and generate uv.lock. For a gradual move, keep your file and use uv pip install -r requirements.txt.




