The Global Interpreter Lock (GIL) has limited CPython to one thread executing bytecode at a time for over 30 years. Python 3.13 introduced an experimental build without it. Python 3.14 promoted that build to officially supported status. Python 3.15 adds the packaging and ABI infrastructure to make it practical for C extensions. This guide covers what each release changes, what still breaks, and how to test your code.
Quick Takeaways
- Python 3.14 moved the free-threaded build (
python3.14t) from experimental to supported under PEP 779. The adaptive specializing interpreter now runs in that build, which cuts single-threaded overhead to roughly 5-10%. - Python 3.15 keeps the GIL on by default. The free-threaded build is a separate binary, python3.15t, and it stays opt-in.
- Python 3.15 adds abi3t, a Stable ABI for free-threaded builds (PEP 803, 820, 793). C extensions targeting the Stable ABI can be compiled to work with free-threaded builds.
- Free-threading removes the lock that hid race conditions. Compound operations such as check-then-act on shared state still need explicit locks.
| Feature | Python 3.13 | Python 3.14 | Python 3.15 |
|---|---|---|---|
| Free-threaded build status | Experimental | Supported (PEP 779) | Supported, still opt-in |
| Default build | GIL | GIL | GIL |
Specializing interpreter in t build |
Disabled | Enabled | Enabled |
Stable ABI for t build |
None | None | abi3t (PEP 803) |
| Multiple interpreters in stdlib | No | concurrent.interpreters (PEP 734) |
Yes |
| Extension import re-enables GIL | Yes | Yes | Yes |
What the GIL Actually Does
The GIL is a mutex around the interpreter. It protects reference counts and internal structures from concurrent mutation. The cost is that CPU-bound threads cannot run in parallel. I/O-bound threads still benefit, because the lock is released during blocking calls.
The free-threaded build removes that mutex and replaces it with finer-grained mechanisms:
- Biased reference counting keeps a fast, non-atomic count for the owning thread and a slower shared count for other threads.
- Deferred reference counting skips counting for hot, long-lived objects such as functions and modules.
- Per-object locks protect built-in containers like
listanddictduring individual operations. - mimalloc serves as the thread-friendly memory allocator.
Check Which Build You Are Running
Never assume. A python3.14 binary and a python3.14t binary behave differently.
import sys
import sysconfig
# True only if this binary was compiled without the GIL
built_without_gil = bool(sysconfig.get_config_var("Py_GIL_DISABLED"))
# Reflects the runtime state: an extension import can re-enable the GIL
gil_active = sys._is_gil_enabled()
print(f"Free-threaded build: {built_without_gil}")
print(f"GIL currently active: {gil_active}")
On a standard build, this prints False and True. On python3.14t with only pure-Python imports, it prints True and False. The two values can diverge on a free-threaded build, which is the key detail. sys._is_gil_enabled() is the runtime truth.
What Changed in Python 3.14
From Experimental to Supported
PEP 779 set the criteria for the “supported” label: stable performance, acceptable memory overhead, and a working C API story. The build is still optional. Supported means the core team commits to maintaining it and fixing bugs, not that it is the default.
Single-Thread Penalty Shrinks
In 3.13, the free-threaded build disabled the specializing adaptive interpreter, so it ran noticeably slower on single-threaded code. In 3.14, specialization is thread-safe and enabled. The documented single-threaded overhead is roughly 5-10% compared with the GIL build, depending on workload. Memory use is also higher, because of per-object metadata and allocator behavior.
Context-Aware Warnings and Thread Inheritance
Two smaller changes matter for real applications:
- The
warningsmodule can store its filter state in a context variable in the free-threaded build, sowarnings.catch_warnings()no longer bleeds across threads. threading.Threadcan inherit the parent’scontextvarscontext viasys.flags.thread_inherit_context, which is on by default in the free-threaded build.
Multiple Interpreters in the Standard Library
PEP 734 exposes sub-interpreters through concurrent.interpreters. Each interpreter has its own GIL (or none), so you get isolation plus parallelism on a standard build. This is the alternative route to multi-core CPU work.
from concurrent.futures import InterpreterPoolExecutor
def cpu_task(n: int) -> int:
"""Pure-Python CPU work. Arguments and results are pickled between interpreters."""
return sum(i * i for i in range(n))
if __name__ == "__main__":
with InterpreterPoolExecutor(max_workers=4) as pool:
results = list(pool.map(cpu_task, [5_000_000] * 4))
print(results)
This runs in parallel on a standard 3.14 build. Each worker is a separate interpreter, so data crosses the boundary by copying rather than by sharing objects.
What Changes in Python 3.15
Python 3.15.0 final is scheduled for October 9, 2026, after the third and final release candidate. If you are reading this before that date, treat the RC as the testing target, because the ABI is frozen.
abi3t: One Wheel Instead of Many
Before 3.15, a package with C extensions had to build a separate cp3XXt wheel for each free-threaded Python version. abi3t lets one wheel cover multiple free-threaded releases, the way abi3 already does for GIL builds. Extensions switch from a PyInit_ function to a new export hook, PyModExport_*, defined in PEP 793, using the PySlot structure from PEP 820.
For library maintainers, this lowers the wheel matrix. For application developers, it means faster arrival of free-threaded wheels for popular packages.
Allocator and Installer Changes
mimalloc is now the default allocator for raw memory allocations such as PyMem_RawMalloc(), which improves performance on free-threaded builds. The official macOS binaries now install free-threading support by default.
Lazy Imports and the JIT
Two non-GIL features change how you benchmark. PEP 810 adds explicit lazy imports, which defer module loading until first use and cut startup time. The experimental JIT also improved: reported gains are 8-9% on x86-64 Linux over the standard interpreter. When you compare GIL and free-threaded builds, hold these settings constant.
The GIL Is Still the Default
Your python3.15 still has the GIL, and the ecosystem now gates adoption more than CPython does. No release has been committed to PEP 703’s “Phase III”, where free-threading becomes the default.
Practical Implementation: CPU-Bound Threads
This benchmark shows the behavioral difference. Run it on a standard build and on a free-threaded build.
import sys
import time
from concurrent.futures import ThreadPoolExecutor
def burn(n: int) -> int:
"""CPU-bound loop. Holds the GIL on standard builds."""
total = 0
for i in range(n):
total += i * i % 7
return total
def run(workers: int, n: int = 10_000_000) -> float:
start = time.perf_counter()
with ThreadPoolExecutor(max_workers=workers) as pool:
list(pool.map(burn, [n] * workers))
return time.perf_counter() - start
if __name__ == "__main__":
print(f"GIL enabled: {sys._is_gil_enabled()}")
t1 = run(1)
t4 = run(4)
print(f"1 thread: {t1:.2f}s")
print(f"4 threads: {t4:.2f}s (4x the total work)")
On a standard build, t4 lands near 4 * t1, because the threads take turns. On a free-threaded build with at least four physical cores, t4 should land close to t1. Actual scaling depends on core count, memory bandwidth, and allocation patterns.
Thread Safety Without the GIL
Built-in operations like list.append() and dict.__setitem__() stay individually atomic through per-object locks. Sequences of operations are not atomic. The GIL often masked this because thread switches happened infrequently. Free-threading exposes it.
import threading
counter = 0
lock = threading.Lock()
def increment(times: int) -> None:
global counter
for _ in range(times):
with lock: # Makes the read-modify-write sequence atomic
counter += 1
threads = [threading.Thread(target=increment, args=(100_000,)) for _ in range(8)]
for t in threads:
t.start()
for t in threads:
t.join()
print(counter) # 800000
Without the with lock: block, counter += 1 compiles to separate load, add, and store steps. Two threads can read the same value and both write back value + 1, losing an update.
Extensions That Re-Enable the GIL
If you import a C extension that has not declared free-threading support, CPython re-enables the GIL at runtime and emits a warning. You can override this for testing:
# Force the GIL off even if an extension asks for it (unsafe for untested extensions)
PYTHON_GIL=0 python3.14t app.py
# Equivalent command-line flag
python3.14t -X gil=0 app.py
Use this only in test environments. An extension that is not thread-safe can corrupt memory under real parallelism.
Free-Threading vs. Multiprocessing vs. Sub-Interpreters
| Approach | CPU Parallelism | Memory Overhead | Data Sharing | Best For |
|---|---|---|---|---|
| Threads (GIL build) | None | Low | Direct | I/O-bound work |
| Threads (free-threaded) | Yes | Moderate | Direct, needs locks | Shared-state CPU work |
multiprocessing |
Yes | High | Pickling or shared memory | Isolated CPU jobs |
| Sub-interpreters | Yes | Moderate | Copy between interpreters | Isolated work, low startup cost |
asyncio |
None | Low | Direct | High-concurrency I/O |
Best Practices: Anti-Patterns vs. Refactored Code
Bad Code: Check-Then-Act on a Shared Dict
# Anti-pattern: the GIL made this look safe. Free-threading does not.
cache = {}
def get_or_compute(key):
if key not in cache: # Thread A and B both see "missing"
cache[key] = expensive(key) # Both compute and overwrite
return cache[key]
Two threads can both pass the in check and compute the value twice. For idempotent work, that wastes time. For side-effecting work, it causes bugs.
Good Code: Atomic Compute Under a Lock
import threading
cache: dict[str, int] = {}
cache_lock = threading.Lock()
def get_or_compute(key: str) -> int:
# Fast path: a plain read of a present key is safe
value = cache.get(key)
if value is not None:
return value
with cache_lock:
# Re-check inside the lock to avoid duplicate work
value = cache.get(key)
if value is None:
value = expensive(key)
cache[key] = value
return value
The double-checked pattern keeps the common path lock-free and makes the slow path atomic. Use dict.setdefault() only when the default value is cheap to build, because the default is evaluated before the call.
Bad Code: Module-Level Mutable State
# Anti-pattern: hidden shared state across threads
results = []
def worker(x):
results.append(process(x)) # Order is nondeterministic
Good Code: Return Values, Collect in One Place
from concurrent.futures import ThreadPoolExecutor
def worker(x):
return process(x)
with ThreadPoolExecutor() as pool:
results = list(pool.map(worker, inputs)) # Ordered, no shared mutation
Migration Checklist
- Install
python3.14tor the 3.15 RC free-threaded build alongside your current interpreter. - Run your test suite with
PYTHON_GIL=0and watch for crashes and hangs. - Audit dependencies against the free-threading compatibility tracker. Packages without a
cp3XXtwheel will trigger the GIL fallback. - Search for shared mutable globals, check-then-act patterns, and lazily initialized singletons.
- Benchmark single-threaded and multi-threaded paths separately. The 5-10% single-threaded cost only pays off if you actually parallelize.
- Run under a thread sanitizer (TSan builds of CPython exist) for C-extension-heavy code.
FAQ
Is the GIL removed in Python 3.14?
No. Python 3.14 ships two builds. The standard python3.14 keeps the GIL. The optional python3.14t runs without it and is officially supported under PEP 779.
Will Python 3.15 make free-threading the default?
No. Python 3.15 keeps the GIL on in the default build and ships free-threading as a separate python3.15t binary. PEP 703’s third phase, where the GIL-free build becomes the default, has no committed release.
Does free-threaded Python make single-threaded code slower?
Slightly. In 3.14 the overhead is roughly 5-10% versus the GIL build, down from larger penalties in 3.13, because the specializing interpreter is now enabled. Memory use is also somewhat higher.
Do I need to rewrite my code for free-threaded Python?
Usually not. Pure-Python code runs unchanged. You must add locks or use thread-safe structures wherever multiple threads perform compound operations on shared state. Your C extensions need to declare free-threading support, or CPython re-enables the GIL when they load.




