Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

For a rapid first-pass exploratory data analysis (EDA), use df.info() to check structure and missingness, df.select_dtypes() to group columns by the types pandas sees, and value_counts() to find common or unusual values. They complement describe(): each answers a different question about an unfamiliar DataFrame.

1. What is this table made of? Use df.info()

DataFrame.info() prints a concise structural summary. It shows the index and columns, non-null counts, data types, and memory information, making it a useful opening check before choosing analyses. See the pandas DataFrame.info() reference.

Compare each column’s non-null count with the number of rows to spot columns that may contain missing values. Treat this as a clue, not a complete data-quality audit: the display can vary with the method’s arguments and pandas display options, and it does not establish whether the data is valid for your use case.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Which columns need type-specific attention? Use select_dtypes()

select_dtypes(include=..., exclude=...) returns a DataFrame containing columns selected by dtype. This lets you focus a check on numeric fields or review text and category fields separately. The pandas basics guide also documents using df.dtypes.value_counts() to count how many columns have each dtype.

df.dtypes.value_counts()

numeric = df.select_dtypes(include="number")
textual = df.select_dtypes(include=["object", "string", "category"])

If a field that should be numeric appears among text columns, investigate how it was read or encoded before using it in calculations. Dtype selection shows the types pandas currently sees; it cannot determine whether a type matches the field’s meaning.

3. Which values or combinations dominate? Use value_counts()

For a single column, Series.value_counts() counts the frequency of each value. It is useful for checking category balance, finding rare labels, or noticing unexpected values. The pandas user guide describes it as a histogram of a one-dimensional array. By default, missing values are excluded; set dropna=False when you want them included.

# Replace "status" with a categorical column in your DataFrame.
df["status"].value_counts(dropna=False)

# Count combinations across two columns.
df.value_counts(subset=["status", "region"], dropna=False)

DataFrame.value_counts() counts distinct combinations across rows; subset limits the combination to selected columns. Refer to the pandas basics guide for the Series and DataFrame methods. High-cardinality columns can produce long results, and a frequency count alone cannot explain why a value occurs or whether it is erroneous.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How these checks complement describe()

On a mixed-type DataFrame, the default describe() behavior summarizes numeric columns, or categorical columns when there are no numeric columns. Its include and exclude arguments let you choose which types to summarize. It remains useful for descriptive statistics; the three checks above make structure, dtype composition, and value frequencies visible through separate, targeted views. See the pandas basics guide.

For an initial scan, run the structural check first, group columns by type next, then count values in columns that merit closer inspection. These outputs are diagnostic clues, not substitutes for checking domain meaning, collection methods, and whether the data is suitable for the analysis you plan to perform.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.