iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
numpy.sum() adds array elements. With no arguments it adds every element and returns one scalar. The axis argument chooses which dimensions get collapsed, keepdims=True keeps those collapsed dimensions at length one so the result broadcasts back against the input, and dtype sets the type used for accumulation as well as the returned value. For a sum of squares, the square must be computed in a wide enough type before sum() runs, because integer overflow in x ** 2 happens before the sum sees the values. NumPy does not raise an error when integer totals overflow, so the wrong answer can look normal.
What numpy.sum() does by default
The function is documented in the numpy.sum reference for the NumPy stable release labelled v2.5, checked in October 2026. Its signature is:
numpy.sum(a, axis=None, dtype=None, out=None, keepdims=<no value>, initial=<no value>, where=<no value>)
With the default axis=None, every element of the array is summed and the result is a scalar. Most real questions come down to three arguments: axis, keepdims and dtype. The sections below take them in that order, followed by sum of squares and floating-point behaviour.
Choosing an axis
An axis is a dimension of the array. In a two-dimensional array, axis=0 runs down the rows, so it reduces each column to one value. axis=1 runs across the columns, so it reduces each row. Reduced dimensions are removed from the output shape by default.
#1 Best Overall
Using the reference’s example array [[0, 1], [0, 5]]:
| Call | Result | Output shape | What was added |
|---|---|---|---|
np.sum(a) or axis=None |
6 | scalar, () | all four elements |
np.sum(a, axis=0) |
[0, 6] | (2,) | each column: 0+0 and 1+5 |
np.sum(a, axis=1) |
[1, 5] | (2,) | each row: 0+1 and 0+5 |
np.sum(a, axis=(0, 1)) |
6 | scalar, () | both axes, same as axis=None |
np.sum(a, axis=-1) |
[1, 5] | (2,) | last axis, same as axis=1 here |
A tuple of axes reduces all of the listed dimensions at once. A negative axis counts from the last dimension, so -1 is always the final axis regardless of how many dimensions the array has. When you are unsure which axis you want, print a.shape first and check which entry you intend to collapse.
Keeping reduced dimensions with keepdims
Setting keepdims=True leaves each reduced dimension in place with length one. The result then has the same number of dimensions as the input, which is what makes broadcasting work when you divide or subtract the total back into the original array.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesimport numpy as np
x = np.array([[1.0, 3.0],
[2.0, 6.0]])
row_totals = np.sum(x, axis=1)
print(row_totals.shape) # (2,)
row_totals_k = np.sum(x, axis=1, keepdims=True)
print(row_totals_k.shape) # (2, 1)
print(x / row_totals_k)
# [[0.25 0.75]
# [0.25 0.75]]
Without keepdims, the division would pair a shape (2, 2) array with a shape (2,) array, and broadcasting would align the totals with columns rather than rows. The row sums [4.0, 8.0] would then be divided across the wrong entries. Keeping the dimension makes the intent explicit: each row is divided by its own total.
Controlling dtype: the result type and the accumulator
The dtype argument sets the type that values are accumulated in, and the returned array uses that type too. It is not just a cast applied after the sum. If you leave it at None, NumPy chooses a type from the input:
- Signed integers narrower than the platform integer are promoted to the platform integer. On most 64-bit systems that is int64, but check
np.int_on your own machine. - Unsigned integers narrower than the platform integer are promoted to the platform unsigned integer.
- Floating-point and complex inputs keep their own dtype, so a
float32array sums infloat32unless you say otherwise.
a8 = np.ones(10, dtype=np.int8)
print(np.sum(a8).dtype) # platform integer, e.g. int64
print(np.sum(a8, dtype=np.int8).dtype) # int8
print(np.sum(a8, dtype=np.int8)) # 10, fits in int8
Passing a narrower dtype is an explicit choice to accept the narrower range. Passing a wider one raises the ceiling before values are added, which is the main tool for avoiding overflow.
Integer overflow wraps silently
NumPy integer arithmetic is modular. When a running total passes the largest value the type can hold, it wraps around to the negative end and no exception is raised. The reference’s own example shows this:
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →np.ones(128, dtype=np.int8).sum(dtype=np.int8) # -128
The true total is 128, which is one more than the largest int8 value of 127, so the result wraps to -128. NumPy’s data-types documentation describes numeric types as having fixed sizes and finite limits, unlike Python’s int, which grows as needed.
Before choosing a dtype, estimate the largest possible total, not the typical one. A sum of a million values, each up to 1,000, can reach one billion, which exceeds int32’s maximum of 2,147,483,647 only if values are larger; but a sum of 3 million values of 1,000 (3 billion) exceeds it. Pick an accumulator whose maximum comfortably covers that bound.
Rank #4
Sum of squares
The expression np.sum(x ** 2) computes the sum of squares, but its safety depends on the order of operations. NumPy squares every element first, storing the results in the dtype of x. Only then does sum() accumulate them. The dtype argument on sum() therefore cannot repair squares that already overflowed.
Why widening only the sum fails
x = np.array([100, 120], dtype=np.int8)
print(np.sum(x ** 2)) # 80 (wrong)
print(np.sum(x ** 2, dtype=np.int64)) # 80 (still wrong)
Here 100² = 10,000 wraps to 16 in int8, and 120² = 14,400 wraps to 64, so both lines report 80 instead of 24,400.
Widen before squaring
x = np.array([100, 120], dtype=np.int8)
total = np.sum(x.astype(np.int64) ** 2, dtype=np.int64)
print(total) # 24400
The conversion has to happen before the power operator. Check that int64 can hold the largest single square and the full total. The largest int64 value is 9,223,372,036,854,775,807, so each square must be at most that value, which means each input must be below roughly 3.04 billion in absolute value. The total also needs to stay within that limit, so very long arrays of large values may need a different approach.
Best Value
Floating-point input
For floating-point arrays, overflow to infinity is a larger concern than wraparound, and rounding error accumulates instead. Choose a wider accumulator where it matters, as described in the next section.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Floating-point accuracy
Floating-point addition is not exact, so the order and grouping of additions can change the last digits of a result. When summing many float32 values, passing dtype=np.float64 reduces accumulation error. NumPy’s documentation notes that this improvement depends on summing along the fast axis in memory, and that exact precision can vary with other parameters. math.fsum from Python’s standard library is slower but gives a more precisely rounded result.
Do not expect bitwise-identical floating-point results when you change the axis, the memory layout, or the dtype. Compare results with a tolerance such as np.isclose rather than exact equality.
Recommended Free Tools
Quick Recap
Other parameters
outwrites the result into an existing array, casting values to that array’s type where necessary.whereincludes only the elements where its boolean mask isTrue.initialsets the starting value of the sum, which matters for empty reductions and for adding a constant offset.
Troubleshooting common results
- The result is negative or much smaller than expected for an integer array. The total wrapped around. Pass a wider
dtype, and convert before squaring if you are computing a sum of squares. - The result has the wrong shape for broadcasting. The reduced dimension was removed. Add
keepdims=True. - A scalar came back when you wanted one value per row.
axiswas omitted or set toNone. Useaxis=1for rows in a two-dimensional array. - Two floating-point totals differ in the last digits. This is normal rounding. Use a wider accumulator and compare with a tolerance.
“
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

