Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
To speed up Python with Numba, first profile representative workloads, then compile the numerical hot path with @njit, test parallel loops where iterations can run independently, and use cache=True when repeated program starts make compilation a startup cost. Measure cold-start and warmed execution separately: none of these techniques guarantees a speedup for every workload.
When Numba can help—and what to measure first
Numba is most useful for numerical functions and loop-heavy sections whose operations and types it supports. It does not make arbitrary Python code native automatically; code that depends on unsupported Python features may fail to compile in nopython mode.
Start by profiling the application with real, representative data and find a function that materially contributes to runtime. Numba’s performance guide recommends profiling real data to guide tuning and cautions that its examples demonstrate features rather than canonical performance guidance: Numba Performance Tips.
- Record the Numba version, machine, input size, and threading configuration.
- Time first-call compilation separately from warmed execution.
- Check outputs against the original implementation, including relevant edge cases.
1. Compile the hot numeric function with @njit
Nopython mode compiles supported operations to native code without relying on Python objects for the function’s execution. Use @njit to make that requirement explicit:
#1 Best Overall
import numba
def sum_squares(values):
total = 0.0
for value in values:
total += value * value
return total
sum_squares_jit = numba.njit(sum_squares)
Alternatively, decorate the function directly with @numba.njit. If the function uses an unsupported operation or type, compilation can fail; simplify the compiled kernel or keep Python-level orchestration outside it. Numba’s @jit decorator has defaulted to nopython mode since version 0.59.0, but @njit clearly signals the intended mode. See the Numba JIT reference.
Do not assume loops need to be rewritten as array expressions. Numba compiles ordinary loops, and its guide’s teaching example reports similar performance for compiled loop and compiled vector-expression versions. That result applies to that example, not every function or array operation.
Rank #2
2. Try parallel loops only when the work fits
When loop iterations are independent, Numba can parallelize supported work using parallel=True and prange. For example, a kernel that computes each output element from the matching input element can express that independent iteration structure:
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchimport numba
import numpy as np
@numba.njit(parallel=True)
def square_values(values):
result = np.empty_like(values)
for i in numba.prange(values.size):
result[i] = values[i] * values[i]
return result
Parallel execution adds overhead, so it may not pay off for small inputs; dependencies between iterations can also make a loop unsuitable. Benchmark the serial and parallel versions at the sizes your application actually processes, and verify results before adopting the parallel version. Numba documents parallel compilation and prange in its performance guide.
3. Cache compilation across program runs
Set cache=True when the same supported function is compiled in separate program invocations and startup time matters:
import numba
@numba.njit(cache=True)
def sum_squares(values):
total = 0.0
for value in values:
total += value * value
return total
Numba persists cached compilation results so later invocations can avoid some compilation work. Cache files are normally placed in the source file’s __pycache__ directory; if that location is not writable, Numba can use a user-wide fallback. Some functions are not cacheable, and cache behavior depends on filesystem support. Consult the Numba caching documentation.
Caching targets repeated compilation overhead, not the kernel’s steady-state runtime. Compare the first invocation in a fresh process with later invocations, and time warmed function execution separately.
Optional tuning: treat fastmath=True as a numerical trade-off
fastmath=True permits floating-point transformations that Numba otherwise treats as unsafe. Those transformations can change numerical results, so enable the option only when the application can tolerate the behavior and tests confirm the results meet its accuracy requirements. The option is documented in the performance guide.
Best Value
Also note that Numba’s JIT reference says bounds checking is off by default. An out-of-range index may produce garbage or a segmentation fault; enabling bounds checking makes such an access raise IndexError. Use the JIT reference for the bounds-checking option and consider it when diagnosing indexing bugs.
How to interpret Numba’s published timing example
Numba’s stable performance guide reports timings for a contrived trigonometric-identity example based on np.arange(1.e7), run on an Intel i7-4790 with four hardware threads:
| Version in the example | Published time |
|---|---|
| Uncompiled NumPy expression | 0.581 s |
| Compiled NumPy expression | 0.659 s |
| Uncompiled loop | 25.2 s |
| Compiled loop | 0.670 s |
These are the Numba project’s illustrative timings for that particular example and setup, not a forecast for another program. The guide labels its figures indicative and its examples pedagogical; use your own representative workload to decide whether a technique helps.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Quick Recap
A practical tuning sequence
- Profile real application data and identify a numerical hot function.
- Compile that function with
@njit; resolve unsupported constructs or leave them outside the compiled kernel. - Benchmark warmed execution against the original, while tracking first-call compilation separately.
- If iterations are independent, try
parallel=Truewithprangeand compare representative input sizes. - If repeated process startup is costly, test
cache=Trueand compare cold and later invocations. - Consider
fastmath=Trueonly after validating numerical behavior against application requirements.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

