Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Interviewers use NumPy questions to find out whether a candidate can predict what an array operation returns: its shape, its dtype, and whether the result shares memory with the original. Reciting function names rarely separates strong candidates from weak ones. The 75 questions below are grouped in the order a working data scientist meets these ideas, from array basics through indexing, views, broadcasting, dtypes, reductions, sorting, random generation, and linear algebra. Each answer states the behavior, gives a short example with shapes written next to the arrays, and names the trap the question is designed to expose.
The behavior described follows NumPy’s stable documentation, including the v2.5 manual pages for fundamentals, quickstart, and indexing. Output comments are shown for NumPy 2.x. Check np.__version__ against your environment before quoting an exact printed format.
Array foundations (questions 1 to 10)
1. What is an ndarray, and why use it instead of a Python list?
An ndarray is NumPy’s N-dimensional array type. Every element has the same dtype, and the values are stored in one contiguous block of memory rather than as separate Python objects. Operations on the whole array run in compiled code, so you can write a * 2 instead of a Python loop. The homogeneity is the reason the speed is possible, and it is also why dtype questions come up so often later in an interview.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minute2. What do ndim, shape, and size report?
ndim is the number of dimensions, shape is a tuple giving the length of each axis, and size is the total number of elements.
#1 Best Overall
a = np.zeros((2, 3, 4))
a.ndim # 3
a.shape # (2, 3, 4)
a.size # 24
3. What is a dtype, and why does it matter?
The dtype is the type of every element, such as int64, float64, bool, or a fixed-width string like <U5. It fixes how many bytes each element uses, which operations are meaningful, and which values can be stored. Read it with a.dtype. Interviewers often ask this to check whether a candidate notices that integer and float results behave differently.
4. What are itemsize and nbytes?
itemsize is the number of bytes per element, and nbytes is the total memory the data occupies (size * itemsize). A 1000 by 1000 array of float64 has an itemsize of 8 and an nbytes of 8,000,000. The figure covers the element data only, not the small object overhead of the array itself.
5. How do you create an array from Python sequences?
Pass a nested list or tuple to np.array(). The shape comes from the nesting, and the dtype is inferred from the contents unless you pass dtype= explicitly.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →np.array([[1, 2], [3, 4]]) # shape (2, 2), int64 on most platforms
np.array([[1, 2], [3]]) # ragged: raises ValueError in current NumPy
Ragged nested lists raise an error unless you explicitly request dtype=object, which gives up most of NumPy’s performance. Mention that trade-off if the question is about irregular data.
6. When do you use zeros, ones, empty, and full?
np.zeros(shape) and np.ones(shape) fill with 0 and 1. np.full(shape, value) fills with any constant. np.empty(shape) allocates memory without initializing it, so its contents are arbitrary. Use empty only when every element will be overwritten before it is read; it is not a faster substitute for zeros in ordinary code.
7. What is the difference between arange and linspace?
np.arange(start, stop, step) excludes stop and produces values spaced by step. With floating-point steps, rounding can make the length differ from what you expect. np.linspace(start, stop, num) includes both endpoints and returns exactly num points.
np.linspace(0, 1, 5) # [0., 0.25, 0.5, 0.75, 1.]
np.arange(0, 1, 0.25) # [0., 0.25, 0.5, 0.75] (stop excluded)
8. How does reshape work, and what can go wrong?
reshape changes the shape without changing the element count. The product of the new dimensions must equal size. You may pass -1 for one dimension, and NumPy infers it. When possible the result is a view of the original data; when memory layout forbids it, NumPy returns a copy.
Recommended Free Tools
a = np.arange(12).reshape(3, 4) # shape (3, 4)
a.reshape(2, -1).shape # (2, 6)
a.reshape(5, -1) # ValueError: 12 is not divisible by 5
9. What is the difference between flatten() and ravel()?
flatten() always returns a new copy. ravel() returns a flattened view when the memory layout allows it and a copy otherwise. Use flatten() when you need an independent result and ravel() when you only need to read the data. Writing to the result of ravel() can change the original, which is a common follow-up question.
10. What happens when a Python list mixes types?
NumPy chooses one dtype that can hold every element. Integers and floats give float64, and a boolean becomes an integer when mixed with integers. Mixing strings with numbers produces a string dtype, so the numbers are converted to text.
np.array([1, 2.5]).dtype # float64
np.array([1, True]).dtype # int64 (True becomes 1)
np.array([1, 'a']).dtype # a unicode string dtype, e.g. <U21
When the meaning of a column depends on its type, pass dtype= explicitly so the conversion does not happen silently.
Indexing and selection (questions 11 to 22)
11. How do you read a single element from a 2D array?
Use a[i, j]. The chained form a[i][j] returns the same value but first builds an intermediate row view, so the comma form is the idiomatic answer. A single element comes back as a NumPy scalar, not a Python int.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
12. How do negative indices work?
A negative index counts from the end. a[-1] is the last element along that axis, and a[-2:] is the last two. The rule applies to each dimension separately.
13. How does slicing work, and is the stop value included?
A slice has the form start:stop:step. The stop value is excluded, and omitted values take defaults.
a = np.arange(10)
a[2:8:3] # [2 5] (indices 2 and 5; 8 is excluded)
a[::-1] # reversed copy of the order, still a view
14. How do multidimensional slices work?
Separate one index or slice per axis with commas. In a 3 by 4 array, a[0, -1] is the top-right value, and a[:, 1:3] keeps every row and columns 1 and 2, giving shape (3, 2).
15. How do the shapes of a[0], a[0:1], and a[:, 0] differ?
With a of shape (3, 4), a[0] has shape (4,), a[0:1] has shape (1, 4), and a[:, 0] has shape (3,). An integer index removes that axis, while a slice keeps it. Candidates who predict these shapes correctly usually avoid the shape errors that follow later in broadcasting work.
16. How do boolean masks select data?
A boolean array with the same shape as the target selects the elements where it is True. The result is always one-dimensional, in row-major order.
a = np.array([[1, 5], [3, 7]]) # shape (2, 2)
a[a > 2] # [5 3 7], shape (3,)
17. What happens when a boolean mask does not match the array?
If the mask’s length along an axis differs from the array’s length along that axis, NumPy raises an IndexError. For example, indexing a length-3 array with a length-2 mask fails. Check mask.shape against a.shape before indexing, especially when the mask came from a filtered dataset.
Rank #2
- Careercup, Easy To Read
- Condition : Good
- Compact for travelling
18. How does integer-array indexing differ from slicing?
Passing a list or array of integers selects elements by position rather than by range. A single integer list reorders or repeats rows, and two integer arrays pick individual elements pairwise.
a = np.array([[10, 11, 12], [20, 21, 22], [30, 31, 32]]) # shape (3, 3)
a[[2, 0]] # rows 2 then 0, shape (2, 3)
a[[0, 1], [2, 0]] # elements (0,2) and (1,0): [12 20], shape (2,)
This form returns a copy, not a view, which matters for the questions in the next section.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →19. How do you assign values through a selection?
Assign to a slice or mask on the left side of =. A scalar on the right is broadcast over every selected position.
a = np.array([[1, -2], [3, -4]])
a[a < 0] = 0 # [[1 0] [3 0]]
a[:, 0] = 9 # sets the first column to 9
Assignment through a basic slice writes into the original array, so this is an in-place change.
20. What does ... (Ellipsis) do in an index?
The Ellipsis stands in for as many full slices as needed to reach the axes you name. In a 4D array, a[..., 0] means a[:, :, :, 0]. It makes code independent of the number of dimensions, which is useful in batch-processing functions.
21. How do you add a new axis?
Insert None (also written np.newaxis) in the index. A 1D array of length 5 becomes a column vector with x[:, None], shape (5, 1), or a row vector with x[None, :], shape (1, 5). This is the standard way to prepare shapes for broadcasting.
Free tools Windows power users keep installed
One-click scans. No signup required.
22. How do you select a sub-grid of specific rows and columns?
Use np.ix_ to build an open mesh from two index lists, so the result is the cross-product of those rows and columns.
a = np.arange(12).reshape(3, 4)
rows = [0, 2]
cols = [1, 3]
a[np.ix_(rows, cols)] # shape (2, 2): rows 0 and 2, columns 1 and 3
Without np.ix_, passing two plain lists pairs them element by element, which is a common mistake.
Views, copies, and memory (questions 23 to 29)
23. What is the difference between a view and a copy?
A view is a new array object that refers to the same memory as the original, so changes through either one are visible in both. A copy owns separate memory. Views are cheap to create, while copies cost time and memory proportional to the data size.
24. Does slicing share memory with the original?
Basic slicing with integers, slices, None, and ... returns a view. You can confirm this with np.shares_memory(a, b), or by checking that b.base is the original array.
a = np.arange(6)
b = a[1:4]
np.shares_memory(a, b) # True
25. What side effect does mutating a view have?
Writing to a view changes the original data, and the original’s later reads see the change. This is the most common accidental-mutation bug in interview code.
a = np.array([1, 2, 3, 4])
b = a[1:3]
b[0] = 99
a # [ 1 99 3 4]
26. Do integer-array and boolean-mask indexing return views?
No. Advanced indexing, meaning integer arrays and boolean masks, returns a copy. Changing the result leaves the original untouched.
a = np.array([1, 2, 3])
b = a[[0, 2]]
b[0] = 100
a # [1 2 3]
Candidates should state this distinction explicitly, because treating every index as a view is an easy mistake to make.
27. How do you force an independent copy?
Call a.copy(). The result has its own memory and can be modified safely. Note that b = a creates no copy at all: it is a second name for the same object.
28. What is contiguity, and why does it matter?
An array is C-contiguous when its elements are laid out row by row in memory, which is NumPy’s default. Some views, such as a transposed array or a strided slice, are not contiguous. You can check this with a.flags['C_CONTIGUOUS']. Contiguity affects speed and whether reshape or ravel can return a view, so it is a useful follow-up when a candidate has explained views.
29. What are safe mutation patterns?
Work on an explicit copy whenever you are about to change data you will reuse. Check for sharing with np.shares_memory when a function receives an array from elsewhere. Avoid in-place operations on slices inside loops unless you intend to change the source. In function design, either document that the input is modified or copy it on entry.
Broadcasting and vectorization (questions 30 to 39)
30. What are the broadcasting rules?
To operate on two arrays, NumPy aligns their shapes from the right. For each pair of dimensions, the sizes must be equal, or one of them must be 1. Missing leading dimensions are treated as 1. The result takes the larger size in each position.
Rank #3
31. Which operations fail under broadcasting?
Adding a shape (3, 4) to a shape (3,) fails, because the trailing dimension is 4 against 3. Adding a shape (4,) works, and so does a shape (3, 1).
a = np.zeros((3, 4))
a + np.ones(4) # shape (3, 4): works
a + np.ones(3) # ValueError: operands could not be broadcast together
32. How does scalar broadcasting work?
A scalar has shape () and is compatible with any shape, so a * 2 multiplies every element. The scalar is not copied in memory; NumPy reuses it along each axis.
33. What happens when two arrays both have a singleton dimension?
Both singleton dimensions expand, which produces a grid. A column of shape (3, 1) plus a row of shape (1, 4) gives shape (3, 4), with every row-column pair combined. This is how you compute an outer sum or outer product without a loop.
col = np.array([[1], [2], [3]]) # (3, 1)
row = np.array([10, 20, 30, 40]) # (4,)
(col + row).shape # (3, 4)
34. How do you use broadcasting to subtract a per-feature mean?
For a dataset with samples on rows and features on columns, the shape is (n_samples, n_features). Subtracting the column means (shape (n_features,)) works directly, because the trailing dimension matches. When the operand is per-sample, add an axis first with [:, None], as in question 21.
35. Why do interviewers ask about vectorization?
Vectorized code expresses an operation on whole arrays instead of element-by-element Python loops. The gain is usually large for numeric work because the loop runs in compiled code. A good answer explains the rewrite, for example replacing a loop that scales each row with x * scale[:, None], and checks that both versions agree on a small test.
36. How do you diagnose a broadcasting error?
Read the shapes in the message, then align them from the right. The message operands could not be broadcast together with shapes (3,4) (3,) means the last dimensions are 4 and 3, so the failure is in the last axis. The usual fix is to reshape the smaller operand with None or reshape so its length lines up with the axis you intended.
37. What does np.broadcast_to return?
np.broadcast_to(x, shape) returns a read-only view of x with the requested shape, and it allocates no new data. Writing to it raises an error. Use it when you need the expanded shape only for reading, such as passing it to a function that expects matching shapes.
38. How do you predict the output shape of a broadcast operation?
Pad the shorter shape with leading 1s, then take the larger value per position when the two values differ and one of them is 1. For shapes (5, 1, 3) and (4, 1), pad the second to (1, 4, 1), which gives (5, 4, 3).
39. How can broadcasting produce a silently wrong result?
Suppose x has shape (n,) and you write x - x[:, None]. The result has shape (n, n) and contains all pairwise differences, not the per-element differences you may have meant. No error is raised. Always check the output shape of a broadcast expression before using it.
Dtypes and missing or non-finite values (questions 40 to 46)
40. How do you choose a dtype?
Choose the smallest type that represents the values correctly. float64 is the default for floating-point data and keeps the most precision. int64 is the default for integers on most platforms. Narrower types such as float32 or int32 halve or quarter the memory, but they reduce precision or range. Pick the dtype based on the data’s range and the accuracy the analysis requires, not on memory alone.
41. What does astype do with floats converted to integers?
a.astype(np.int64) truncates toward zero, so 2.9 becomes 2 and -2.9 becomes -2. It does not round. Use np.round first if rounding is intended. Values outside the target range or NaN values give undefined integer results, so check them before converting.
42. How do / and // differ on integers?
/ is true division and always returns floating-point values. // is floor division and keeps the integer dtype. For negative numbers, floor division rounds toward negative infinity, so -7 // 2 is -4, not -3.
a = np.array([7, -7])
a / 2 # [ 3.5 -3.5]
a // 2 # [ 3 -4]
43. How do you find and handle NaN values?
Use np.isnan(a) to get a boolean mask. Do not test with a == np.nan, because NaN is not equal to anything, including itself, so that comparison is always False. To replace missing values, use np.nan_to_num or assign through the mask, as in question 19. Reductions such as np.sum return NaN when any input is NaN, which is covered in question 53.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
44. How do you handle infinity?
np.isinf(a) detects positive and negative infinity. np.isfinite(a) is True only for values that are neither NaN nor infinite. A filter such as a[np.isfinite(a)] keeps only usable numbers. Infinity often appears from division by zero in a feature-engineering step, so checking after such operations is a reasonable habit.
45. How does type promotion work?
When two arrays of different dtypes meet, NumPy chooses a dtype that can hold both. An integer array plus a float array gives float64. In NumPy 2, Python scalars no longer silently widen a typed array. The scalar takes the array’s dtype, so an integer value that does not fit raises an error instead of upcasting.
x = np.array([1, 2], dtype=np.int8)
x + 1000 # OverflowError: Python integer 1000 out of bounds for int8
Check promotion behavior when reviewing code that mixes narrow dtypes with constants.
46. What precision and overflow traps should a candidate know?
Two are common. First, fixed-width integers wrap around silently inside arrays: np.array([255], dtype=np.uint8) + 1 gives 0, with no error. Second, float64 cannot represent every integer above 253, so large int64 identifiers converted to floats can change. Floating-point results should also be compared with np.isclose rather than ==.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallRank #4
Aggregations and axes (questions 47 to 54)
47. How do sum, mean, min, and max work?
Called without an axis, each reduces the whole array to one value. Methods and functions are interchangeable: a.sum() and np.sum(a) give the same result. The mean of an integer array is returned as a float.
48. What does axis mean?
axis names the dimension that is collapsed. Along that axis, the function combines the values, and the other axes remain in the output. Think of the axis as the one you disappear.
49. What is the difference between axis=0 and axis=1?
For a 2D array with shape (rows, cols), sum(axis=0) collapses rows and returns one total per column. sum(axis=1) collapses columns and returns one total per row.
a = np.array([[1, 2, 3],
[4, 5, 6]]) # shape (2, 3)
a.sum(axis=0) # [5 7 9], shape (3,): column totals
a.sum(axis=1) # [ 6 15], shape (2,): row totals
Confusion between these two is among the most frequent errors in interview exercises, so it is worth asking candidates to predict the output shape before running the code.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problems50. What does keepdims=True change?
With keepdims=True, the collapsed axis stays in the result with length 1. The output then has the same number of dimensions as the input, which lets it broadcast back against the original array. For example, a.mean(axis=1, keepdims=True) on the array in question 49 has shape (2, 1), and a - a.mean(axis=1, keepdims=True) centers each row.
51. How do you predict the output shape of a reduction on a 3D array?
Remove the reduced axis from the shape. For b of shape (2, 3, 4), b.sum(axis=1) has shape (2, 4), b.sum(axis=(0, 2)) has shape (3,), and b.sum() is a scalar. Adding keepdims=True keeps the removed axes as length 1.
52. How do you reduce across batches and features?
Assume a batch of shape (batch, features). x.mean(axis=0) gives one mean per feature across the batch, and x.mean(axis=1) gives one mean per sample across features. Stating which of these you need, and why, is usually the point of the question.
53. How do you aggregate arrays that contain NaN?
Ordinary reductions return NaN if any input is NaN. The NaN-aware functions np.nansum, np.nanmean, np.nanmin, and np.nanmax ignore NaN values. Use them only when ignoring the missing values is the intended treatment, and record how many values were dropped if the result matters.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →54. How do you find the index of the maximum in each row?
Use np.argmax(a, axis=1). It returns one index per row, with shape (rows,). Without axis, the function flattens the array and returns a single index into the flattened data, which is a frequent source of wrong answers.
Sorting, uniqueness, and conditional operations (questions 55 to 60)
55. What is the difference between sort and argsort?
np.sort(a) returns the sorted values. np.argsort(a) returns the positions that would sort the array. Use argsort when you need to reorder a second array to match, or to recover the rank of each element.
a = np.array([30, 10, 20])
np.sort(a) # [10 20 30]
np.argsort(a) # [1 2 0]
56. Does sorting change the original array?
np.sort(a) returns a sorted copy and leaves a unchanged. a.sort() sorts in place and returns None, so assigning its result is a bug. Note the trap: in-place sorting modifies every view that shares the memory.
57. How do you count unique values?
np.unique(a, return_counts=True) returns the sorted unique values and how often each occurs. It flattens multidimensional input unless you pass axis, so use the axis argument when you want unique rows.
vals, counts = np.unique(np.array([3, 1, 3, 2, 3]), return_counts=True)
# vals = [1 2 3], counts = [1 1 3]
58. How does np.where work?
With three arguments, np.where(cond, x, y) picks from x where the condition is true and from y elsewhere. The arrays are broadcast together. With a single argument, it returns the indices where the condition is true, as a tuple of arrays, one per axis.
a = np.array([-1, 4, -3, 2])
np.where(a > 0, a, 0) # [0 4 0 2]
np.where(a > 0) # (array([1, 3]),)
59. How does clip work?
np.clip(a, low, high) bounds every value to the range, so values below low become low and values above high become high. Either bound can be None to leave that side open. Clipping is a common way to limit outliers before modeling, but the choice of bounds is a domain decision the candidate should be able to justify.
60. How do you get the top k values?
For a small k, sort the indices and slice: idx = np.argsort(a)[::-1][:k]. For large arrays where only the top k matter, np.argpartition(a, -k)[-k:] finds the top k in linear time, but those k items are not themselves sorted. Sort them afterward if the order matters.
Random generation and reproducibility (questions 61 to 65)
61. What is the modern random workflow?
Create a generator with rng = np.random.default_rng(), then call methods on it, such as rng.random(size) or rng.normal(size=...). The Generator object owns its state, so you can pass it into functions explicitly.
62. How do you make results repeatable?
Pass an integer seed: rng = np.random.default_rng(42). The same seed and the same sequence of calls produce the same numbers on the same NumPy version. The results of a seeded generator are not guaranteed to be identical across every NumPy release, so pin the version in shared code.
63. How do integer ranges work?
rng.integers(low, high, size) excludes high by default. A six-sided die is rng.integers(1, 7, size=10). Set endpoint=True if you want the upper bound included.
64. How do you sample and shuffle?
rng.choice(a, size, replace=False) draws without replacement, and replace=True draws with it. rng.shuffle(a) shuffles in place and returns None. rng.permutation(a) returns a shuffled copy and leaves the input unchanged.
65. Why prefer Generator over the legacy global functions?
The legacy functions such as np.random.seed and np.random.rand share one hidden global state. Any library call can advance that state, so results depend on what ran earlier. A Generator is an explicit object you own, so two parts of a program can have independent, reproducible streams. NumPy’s documentation directs new code to the Generator workflow, though the legacy functions remain available.
Recommended Free Tools
Linear algebra and practical data tasks (questions 66 to 75)
66. How do *, np.dot, and @ differ?
* is elementwise multiplication with broadcasting. @ is matrix multiplication, equivalent to np.matmul for the usual 2D and batched cases. np.dot performs the same matrix product for 2D inputs but treats higher-dimensional and 1D inputs differently, so @ is the clearer choice in new code.
A = np.array([[1, 2], [3, 4]])
B = np.array([[0, 1], [1, 0]])
A * B # [[0 2] [3 0]]: elementwise
A @ B # [[2 1] [4 3]]: matrix product
67. What shape rule does matrix multiplication follow?
For A @ B, the last axis of A must equal the second-to-last axis of B. An (m, n) matrix times an (n, p) matrix gives (m, p). A mismatch raises a ValueError, and checking the inner dimensions first is the quickest way to debug it.
68. What does a dot product of 1D arrays return?
For two vectors of the same length, np.dot(u, v) and u @ v return a scalar, the inner product. A 2D matrix times a 1D vector returns a 1D vector whose length is the number of rows.
69. How does transpose work?
a.T reverses the order of the axes. For a shape (2, 3) matrix, the transpose has shape (3, 2). For a 1D array, .T changes nothing, which surprises candidates who expect a column vector. Use x[:, None] from question 21 to create one.
Free tools Windows power users keep installed
One-click scans. No signup required.
70. How do you solve a linear system?
Use np.linalg.solve(A, b) for a square, invertible matrix A and right-hand side b. It finds x satisfying A @ x = b without computing an explicit inverse, which is more accurate and faster. If A is singular, NumPy raises numpy.linalg.LinAlgError.
71. What do you use for a non-square system?
For an overdetermined system, where there are more equations than unknowns, use np.linalg.lstsq(A, b, rcond=None). It returns the least-squares solution, along with residuals, the rank, and the singular values. solve does not accept non-square matrices.
72. How do norms work?
np.linalg.norm(x) returns the Euclidean (2-) norm for a vector and the Frobenius norm for a matrix when no ord is given. Pass axis to compute one norm per row, for example np.linalg.norm(X, axis=1) for sample lengths. Norms are the usual basis for distance and normalization tasks.
73. How do you compute pairwise distances without loops?
For a matrix X of shape (n, d), subtract with broadcasting to get all pairwise differences, then reduce.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →diff = X[:, None, :] - X[None, :, :] # shape (n, n, d)
dist = np.sqrt((diff ** 2).sum(axis=-1)) # shape (n, n)
The intermediate array has n * n * d elements, so the approach can use a lot of memory for large n. Mention that trade-off when the interviewer asks about scale.
74. How do you normalize features by column?
Standardize each feature to zero mean and unit variance with keepdims-shaped statistics, so broadcasting lines them up with the columns.
mu = X.mean(axis=0, keepdims=True) # shape (1, d)
sd = X.std(axis=0, keepdims=True) # shape (1, d), population std (ddof=0)
sd = np.where(sd == 0, 1.0, sd) # avoid dividing by zero on constant columns
Z = (X - mu) / sd
The guard on zero standard deviation is the detail interviewers look for. A constant column would otherwise produce NaN or infinity.
75. How do you clean a numeric matrix with missing values?
A concise answer covers two strategies. To drop incomplete rows, keep rows with no NaN: X = X[~np.isnan(X).any(axis=1)]. To impute with column means, compute the means ignoring NaN and fill the gaps through the index arrays.
col_means = np.nanmean(X, axis=0)
rows, cols = np.where(np.isnan(X))
X[rows, cols] = col_means[cols]
The better answer names the cost of each choice: dropping rows can discard a large share of the data, while mean imputation shrinks variance and ignores relationships between features.
The Bottom Line
Practice by predicting before you run. For each snippet, write down the output shape, the dtype, and whether the result is a view, then check your answer in the interpreter. That habit is what the questions above are built to test, and it is the fastest way to catch the broadcasting, axis, and view traps before an interview does.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

