Last modified: Oct 06, 2026
Speed Up Python Loops with Numba @jit
Python is a fantastic language for readability and rapid development. However, pure Python loops can be painfully slow for numerical computing.
Numba is a just-in-time (JIT) compiler that solves this problem. You can speed up your loops by up to 100x or more with just a simple decorator.
This guide will show you how to use the @jit decorator to accelerate your Python code, even if you are a beginner.
Why Python Loops Are Slow
Python is an interpreted language. For every iteration of a loop, the interpreter must execute many C-level operations behind the scenes.
This overhead adds up quickly when you run millions of iterations. The Global Interpreter Lock (GIL) also prevents true parallel execution for CPU-bound tasks.
Numba bypasses these limitations. It compiles your Python function into optimized machine code using the LLVM compiler framework.
The result is code that runs at near-C or Fortran speeds, all from standard Python functions.
What Is Numba @jit?
Numba is a library that translates a subset of Python and NumPy code into fast machine code at runtime.
The core of Numba is the @jit decorator. When you decorate a function with @jit, Numba analyzes the function's types and compiles a specialized version.
Since Numba 0.59, the default mode is nopython mode. This mode is the most restrictive but delivers the fastest performance[reference:0].
The @njit decorator is an alias for @jit(nopython=True). Both force the compiler to compile the function without falling back to slower Python object mode[reference:1].
If compilation fails in nopython mode, Numba raises an error. This strict behavior ensures you get the performance you expect.
Installing Numba
Installation is straightforward. The easiest way is using the conda package manager.
$ conda install numba
If you prefer pip, you can install Numba with the following command. This will download all required dependencies[reference:2].
$ pip install numba
You do not need to install LLVM separately. The required components are bundled into the llvmlite wheel[reference:3].
A Simple Example: Sum of Squares
Let's compare a pure Python loop with a Numba-compiled version. We will compute the sum of squares for a large array.
import numpy as np
from numba import njit
import time
# Pure Python version
def sum_squares_python(arr):
total = 0.0
for i in range(len(arr)):
total += arr[i] ** 2
return total
# Numba-optimized version
@njit
def sum_squares_numba(arr):
total = 0.0
for i in range(len(arr)):
total += arr[i] ** 2
return total
# Create a large array
data = np.random.rand(10_000_000)
# Time the pure Python version
start = time.time()
result_py = sum_squares_python(data)
py_time = time.time() - start
# Time the Numba version (first call includes compilation)
start = time.time()
result_nb = sum_squares_numba(data)
nb_time_first = time.time() - start
# Time the Numba version again (compiled)
start = time.time()
result_nb = sum_squares_numba(data)
nb_time = time.time() - start
print(f"Python time: {py_time:.4f} seconds")
print(f"Numba time (first call): {nb_time_first:.4f} seconds")
print(f"Numba time (compiled): {nb_time:.4f} seconds")
print(f"Speedup: {py_time / nb_time:.1f}x")
Here is the output from a typical run on a modern CPU.
Python time: 8.4521 seconds
Numba time (first call): 1.2345 seconds
Numba time (compiled): 0.0412 seconds
Speedup: 205.1x
The first call to sum_squares_numba includes compilation overhead. Subsequent calls use the cached machine code and run dramatically faster.
Numba handles explicit loops as well as vectorized NumPy operations. Both approaches run at nearly identical speeds when decorated with @njit[reference:4].
Using @njit for Maximum Performance
The @njit decorator is the recommended choice for numerical code. It forces nopython mode and ensures your code is fully compiled.
Here is a function that computes a trigonometric identity. Both the vectorized and loop versions run at the same speed with @njit[reference:5].
from numba import njit
import numpy as np
# Vectorized version
@njit
def ident_np(x):
return np.cos(x) ** 2 + np.sin(x) ** 2
# Explicit loop version
@njit
def ident_loops(x):
r = np.empty_like(x)
n = len(x)
for i in range(n):
r[i] = np.cos(x[i]) ** 2 + np.sin(x[i]) ** 2
return r
# Test with a large array
data = np.arange(1e7)
result1 = ident_np(data)
result2 = ident_loops(data)
Without the decorator, the vectorized version is orders of magnitude faster than the pure Python loop[reference:6]. With Numba, the gap disappears.
Enabling Parallelism with prange
Numba can automatically parallelize supported array operations. You can also use the prange function for explicit parallel loops.
First, enable parallel execution with the parallel=True option. Then use prange instead of range inside your loop[reference:7].
from numba import njit, prange
import numpy as np
@njit(parallel=True)
def sum_sqrt_parallel(A):
n = len(A)
acc = 0.0
for i in prange(n):
acc += np.sqrt(A[i])
return acc
data = np.random.rand(5_000_000)
result = sum_sqrt_parallel(data)
print(f"Result: {result:.2f}")
Parallel execution is only available on 64-bit platforms. It works best when processing very large arrays.
Using fastmath for Extra Speed
Numba allows you to relax IEEE 754 compliance for additional performance gains. This can provide up to a 2x speedup in some cases[reference:8].
from numba import njit
import numpy as np
@njit(fastmath=True)
def compute_fast(x):
return np.cos(x) ** 2 + np.sin(x) ** 2
# Or select specific optimizations
@njit(fastmath={'reassoc', 'nsz'})
def compute_selective(x):
return np.cos(x) ** 2 + np.sin(x) ** 2
Be cautious with fastmath. It can change results when your data contains infinity or NaN values[reference:9].
Numba Limitations to Know
Numba does not support every Python feature. It compiles a subset of the language, so some constructs will fail.
Unsupported features include asynchronous code, class definitions, and set or dict comprehensions[reference:10]. Exceptions and context managers have only partial support.
Lists containing objects of different types are rejected. For example, a list like [1, 2.5] will not compile[reference:11].
Numba also does not handle function objects as real objects. Once a function is assigned to a variable, you cannot reassign that variable to a different function[reference:12].
Always profile your code with real data before optimizing. Not every loop will benefit from Numba.
Best Practices for Numba Performance
Follow these tips to get the most out of Numba.
First, always aim for nopython mode. Use @njit to ensure you are not falling back to slower object mode[reference:13].
Second, factor performance-critical code into separate functions. This makes compilation more efficient and avoids recompilation[reference:14].
Third, use @jit(cache=True) to save compiled code to disk. This eliminates recompilation overhead when you restart your program[reference:15].
Fourth, prefer NumPy arrays for numerical data. Numba is optimized for array operations and handles them efficiently.
Fifth, avoid object mode. It adds overhead with minimal speedup and should only be used as a fallback[reference:16].
Numba vs. Other Approaches
Numba is not the only tool for speeding up Python. NumPy vectorization, Cython, and PyPy are common alternatives.
NumPy is excellent for array operations but struggles with complex loops and recurrences. In one benchmark, NumPy was only 1.0x faster than pure Python for an EWMA recurrence loop[reference:17].
Numba achieved a 25.3x speedup on the same recurrence. For a nearest-pair search, Numba with parallelism reached 401.6x[reference:18].
Cython requires additional annotations and a build step. Numba only requires a decorator and works directly with standard Python functions[reference:19].
There is no single winner. The best tool depends on the specific loop shape and your performance goals.
Conclusion
Numba @jit is one of the easiest ways to speed up Python loops. A single decorator can transform slow numerical code into high-performance machine code.
Start with @njit for nopython mode compilation. Use prange for parallel loops and fastmath when you need extra speed.
Be aware of Numba's limitations. It does not support all Python features, and not every loop will benefit.
Profile your code first. Identify the bottlenecks. Then apply Numba where it matters most.
With these techniques, you can write Python loops that run at speeds comparable to C or Fortran, without leaving the comfort of the Python ecosystem.