Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

To speed up Python with Numba, first profile representative workloads, then compile the numerical hot path with @njit, test parallel loops where iterations can run independently, and use cache=True when repeated program starts make compilation a startup cost. Measure cold-start and warmed execution separately: none of these techniques guarantees a speedup for every workload.

When Numba can help—and what to measure first

Numba is most useful for numerical functions and loop-heavy sections whose operations and types it supports. It does not make arbitrary Python code native automatically; code that depends on unsupported Python features may fail to compile in nopython mode.

Start by profiling the application with real, representative data and find a function that materially contributes to runtime. Numba’s performance guide recommends profiling real data to guide tuning and cautions that its examples demonstrate features rather than canonical performance guidance: Numba Performance Tips.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Record the Numba version, machine, input size, and threading configuration.
  • Time first-call compilation separately from warmed execution.
  • Check outputs against the original implementation, including relevant edge cases.

1. Compile the hot numeric function with @njit

Nopython mode compiles supported operations to native code without relying on Python objects for the function’s execution. Use @njit to make that requirement explicit:

import numba

def sum_squares(values):
    total = 0.0
    for value in values:
        total += value * value
    return total

sum_squares_jit = numba.njit(sum_squares)

Alternatively, decorate the function directly with @numba.njit. If the function uses an unsupported operation or type, compilation can fail; simplify the compiled kernel or keep Python-level orchestration outside it. Numba’s @jit decorator has defaulted to nopython mode since version 0.59.0, but @njit clearly signals the intended mode. See the Numba JIT reference.

Do not assume loops need to be rewritten as array expressions. Numba compiles ordinary loops, and its guide’s teaching example reports similar performance for compiled loop and compiled vector-expression versions. That result applies to that example, not every function or array operation.

2. Try parallel loops only when the work fits

When loop iterations are independent, Numba can parallelize supported work using parallel=True and prange. For example, a kernel that computes each output element from the matching input element can express that independent iteration structure:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import numba
import numpy as np

@numba.njit(parallel=True)
def square_values(values):
    result = np.empty_like(values)
    for i in numba.prange(values.size):
        result[i] = values[i] * values[i]
    return result

Parallel execution adds overhead, so it may not pay off for small inputs; dependencies between iterations can also make a loop unsuitable. Benchmark the serial and parallel versions at the sizes your application actually processes, and verify results before adopting the parallel version. Numba documents parallel compilation and prange in its performance guide.

3. Cache compilation across program runs

Set cache=True when the same supported function is compiled in separate program invocations and startup time matters:

import numba

@numba.njit(cache=True)
def sum_squares(values):
    total = 0.0
    for value in values:
        total += value * value
    return total

Numba persists cached compilation results so later invocations can avoid some compilation work. Cache files are normally placed in the source file’s __pycache__ directory; if that location is not writable, Numba can use a user-wide fallback. Some functions are not cacheable, and cache behavior depends on filesystem support. Consult the Numba caching documentation.

Caching targets repeated compilation overhead, not the kernel’s steady-state runtime. Compare the first invocation in a fresh process with later invocations, and time warmed function execution separately.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Optional tuning: treat fastmath=True as a numerical trade-off

fastmath=True permits floating-point transformations that Numba otherwise treats as unsafe. Those transformations can change numerical results, so enable the option only when the application can tolerate the behavior and tests confirm the results meet its accuracy requirements. The option is documented in the performance guide.

Also note that Numba’s JIT reference says bounds checking is off by default. An out-of-range index may produce garbage or a segmentation fault; enabling bounds checking makes such an access raise IndexError. Use the JIT reference for the bounds-checking option and consider it when diagnosing indexing bugs.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to interpret Numba’s published timing example

Numba’s stable performance guide reports timings for a contrived trigonometric-identity example based on np.arange(1.e7), run on an Intel i7-4790 with four hardware threads:

Version in the example Published time
Uncompiled NumPy expression 0.581 s
Compiled NumPy expression 0.659 s
Uncompiled loop 25.2 s
Compiled loop 0.670 s

These are the Numba project’s illustrative timings for that particular example and setup, not a forecast for another program. The guide labels its figures indicative and its examples pedagogical; use your own representative workload to decide whether a technique helps.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A practical tuning sequence

  1. Profile real application data and identify a numerical hot function.
  2. Compile that function with @njit; resolve unsupported constructs or leave them outside the compiled kernel.
  3. Benchmark warmed execution against the original, while tracking first-call compilation separately.
  4. If iterations are independent, try parallel=True with prange and compare representative input sizes.
  5. If repeated process startup is costly, test cache=True and compare cold and later invocations.
  6. Consider fastmath=True only after validating numerical behavior against application requirements.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.