Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use np.where(condition, value_if_true, value_if_false) to create conditional values from pandas data. For example, assign its result to a DataFrame column to label each row according to a test on that row. If you need to keep original values, filter rows, or handle more than two outcomes, pandas and NumPy offer different operations better suited to those jobs.

Use np.where to create a conditional column

Import NumPy as np, build a Boolean condition from a DataFrame column, and pass that condition and the two possible values to np.where. The result can be assigned to a new or existing column:

import numpy as np
import pandas as pd

df = pd.DataFrame({'col2': ['Z', 'A', 'Z']})
df['color'] = np.where(df['col2'] == 'Z', 'green', 'red')

Rows where col2 equals 'Z' get 'green'; the other rows get 'red'. The condition is evaluated element by element, so the output provides one choice for each row. This is the pattern shown in the pandas “Indexing and selecting data” guide.

In general, the arguments are:

  • condition: a Boolean Series or array that identifies which positions meet the test.
  • value_if_true: the value selected where the condition is true.
  • value_if_false: the value selected where the condition is false.

Combine multiple conditions safely

For a row-level test involving more than one column, use elementwise operators and put parentheses around each comparison:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
condition = (df['score'] > 0) & (df['group'] == 'x')
df['result'] = np.where(condition, 'include', 'exclude')

Use & for elementwise AND and | for elementwise OR. Python’s scalar and and or do not combine the individual True/False values in pandas Series.

Choose the operation that matches the result you want

Goal Use Result
Create a value from a two-way condition np.where(condition, true_value, false_value) A conditional result suitable for assigning to a column.
Replace values that fail a condition while retaining the object’s shape Series.where or DataFrame.where Original values remain where the condition is true; false positions receive other, or null if no replacement is supplied.
Return only rows that meet a condition Boolean selection, such as df[df['Age'] > 35] A subset of rows, rather than a same-shape result.
Choose among several conditions numpy.select Values selected from corresponding ordered conditions and choices, with a default for unmatched positions.

The pandas guide describes df1.where(mask, df2) as roughly equivalent to np.where(mask, df1, df2). The difference is the framing: call pandas where on the values to keep, while np.where takes both alternatives as arguments. See the pandas guide for the comparison and its examples.

Use numpy.select for more than two outcomes

When a condition has several possible outcomes, numpy.select accepts a list of conditions and a matching list of choices. Provide an explicit default so rows that match none of the conditions receive an intentional value:

conditions = [df['score'] >= 90, df['score'] >= 70]
choices = ['high', 'medium']
df['level'] = np.select(conditions, choices, default='low')

Conditions and choices correspond by position. Order the conditions deliberately, since a row satisfying multiple conditions is assigned according to the first matching condition. The pandas guide documents this as the multi-choice alternative to np.where.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep pandas alignment and dtype behavior in mind

A pandas Series carries an index, while a raw NumPy array is positional. Ensure a condition and any replacement values correspond to the intended rows; mixing label-aware pandas objects with positional arrays can otherwise produce unexpected row assignments. The DataFrame.where API reference documents alignment for its condition and replacement value.

Also check the resulting column’s dtype when the two choices have different types. NumPy may produce a dtype different from what you expect for mixed choices. For pandas DataFrame.where, the caller’s dtype takes precedence, and a replacement is cast when that can be done losslessly. Consult the API reference matching your installed pandas version if dtype or alignment behavior is important.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Filter rows instead of labeling them

If you want to discard rows that do not meet a condition, use a Boolean mask in the DataFrame brackets rather than using np.where to create replacement labels:

older = df[df['Age'] > 35]

This returns only rows whose age is greater than 35. The pandas tutorial “How do I select a subset of a DataFrame?” demonstrates Boolean-mask selection.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pandas documentation pages are version-sensitive. The guide and tutorial results are labeled pandas 3.0.5 and 3.0.6, respectively, while the DataFrame.where API reference is development documentation. For exact behavior, use documentation for the pandas and NumPy versions installed in your environment.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.