Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use pandas to load, inspect, clean, analyze, and reshape tabular data in Python. This cheatsheet covers the core objects and everyday operations, with links to the official guide for details. The current pandas documentation identifies version 3.0.6, dated September 17, 2026. See the pandas documentation.

What kind of data does pandas handle?

pandas is an open-source Python library for working with data, especially tables like those found in spreadsheets and databases. Its common tasks include exploring, cleaning, and processing data, calculating summaries, grouping records, and reshaping tables. The pandas project overview describes these uses.

The two main objects are:

  • Series: a one-dimensional labeled array, such as one table column.
  • DataFrame: a two-dimensional labeled table whose columns can hold different types of data.

Labels are central to pandas. Each row and column can have an index or name, and pandas uses labels to align data during operations. That makes it useful to check indexes and column names when results do not line up as expected. Read about pandas data structures.

How do I install and import pandas?

The official installation guide recommends installing and running pandas in a virtual environment. Choose the command that matches your package manager:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • conda: conda install -c conda-forge pandas
  • pip: pip install pandas
  • From source: follow the source-installation instructions in the official installation guide.

In a Python script or notebook, import it with the customary alias:

import pandas as pd

How do I create and inspect a table?

You can build a small DataFrame from Python data. A dictionary maps each column name to its values:

import pandas as pd

df = pd.DataFrame({
    "item": ["notebook", "pen", "folder"],
    "price": [4.50, 1.25, 3.00],
    "in_stock": [True, True, False],
})

For a quick first look, inspect the rows, column types, and summary statistics:

df.head()       # first rows
 df.info()       # column names, non-missing counts, and types
 df.describe()   # summary statistics for numeric columns

In a notebook, entering df by itself displays the table. Use df.columns to check column names and df.shape to see the number of rows and columns.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I read and write tabular data?

CSV files are a common starting point. pd.read_csv returns a DataFrame:

df = pd.read_csv("sales.csv")
df.to_csv("sales_clean.csv", index=False)

The index=False argument prevents the DataFrame index from being written as an extra CSV column. pandas also supports common formats and sources including Excel, SQL, JSON, and Parquet. Reader functions generally follow the read_* naming pattern; check the relevant function documentation for required options and dependencies. See the official read-and-write tutorial.

How do I select rows and columns?

Use square brackets for straightforward column selection or filtering. Use loc when selecting by labels and iloc when selecting by integer positions. The accessors at and iat are optimized for retrieving a single value; the 10 Minutes to pandas guide recommends these access methods for production code. Review its indexing examples.

df["price"]                         # one column (a Series)
df[["item", "price"]]              # selected columns (a DataFrame)
df[df["price"] > 2]                 # rows matching a condition
df.loc[0, "item"]                   # row label 0, column label "item"
df.iloc[0, 1]                        # first row, second column

Whether an integer passed to loc means a particular row depends on the index labels; iloc always refers to positions. Choose based on whether your selection is label-based or position-based.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I handle missing data and transform columns?

Check for missing values with isna. You can remove rows with missing values or fill them with a value appropriate to your data; the right choice depends on what a missing entry means.

df.isna().sum()                       # missing values per column
df.dropna()                           # return rows with no missing values
df["price"] = df["price"].fillna(0)  # fill missing prices with 0

Apply operations to a column directly for elementwise transformations. For example, create a tax-inclusive price column:

df["price_with_tax"] = df["price"] * 1.08

Filling missing prices with zero is only suitable when zero accurately represents the missing value in your context; otherwise, select a more appropriate treatment.

How do I calculate summaries and group data?

Use a Series or DataFrame method for common summary statistics, then use groupby to calculate them for categories:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
df["price"].mean()                   # average price
df.groupby("in_stock")["price"].mean()

The grouped expression returns the average price for each value in in_stock. You can replace mean with another suitable aggregation, such as sum or count. See the official statistics tutorial.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When should I merge or reshape tables?

Combine related tables with merge

Use merge when two tables share a key and you want to combine their columns. For example, if products and orders both have a product_id column:

combined = orders.merge(products, on="product_id", how="left")

A left merge keeps every row from orders and attaches matching product data where available. Choose the join type based on which unmatched rows should remain.

Change table layout with reshape operations

Reshaping changes how values are organized, for example, converting columns into rows with melt or turning row categories into columns with pivot. Use these when the data’s current layout does not suit the analysis or output you need. See the official reshaping tutorial.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Where should I continue learning?

The pandas User Guide recommends that people new to the library start with 10 Minutes to pandas. It introduces the core structures, creating and viewing data, selection, missing data, operations, merging, grouping, reshaping, time series, categorical data, plotting, and import/export. Treat it as an overview, then use the User Guide for deeper explanations of specific topics.

If you prefer a book-length resource, the pandas project also recommends Python for Data Analysis by Wes McKinney. It is optional; the official online guides provide a free starting point. Find the project’s learning resources.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.