iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Pandas is a Python library for working with structured data: it lets you load tables, inspect and select records, handle missing values, reshape information, and save results. Its core objects are the one-dimensional Series and the two-dimensional, labeled DataFrame. This guide walks through a small workflow from installation to analysis.
What is pandas in Python?
Pandas is an open-source library for analyzing and manipulating structured data. It is especially useful for tabular or heterogeneous data: a DataFrame can have labeled rows and columns, with different data types in different columns. A Series is a one-dimensional labeled sequence; a DataFrame is a two-dimensional table made up of labeled columns.
This focus differs from NumPy’s typical use for homogeneous numerical arrays. Pandas builds on Python’s data ecosystem and is commonly used to prepare, explore, and summarize data before further analysis or visualization. O’Reilly’s sample chapter describes pandas as designed for tabular or heterogeneous data.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →How do I install pandas?
Install pandas into the same Python environment that will run your script or notebook. The two common package-management routes are pip and Conda. For pip, open a terminal associated with that environment and run:
#1 Best Overall
python -m pip install pandas
Using python -m pip ties pip to the selected Python interpreter more explicitly than a bare pip command. If your setup uses Conda, install through the environment’s Conda package manager instead. Exact compatibility and release details can change, so consult the official pandas installation guide for your operating system and environment.
Then import the library, conventionally under the short name pd:
import pandas as pd
How do I create a DataFrame and read a CSV?
Create a small DataFrame
You can build a DataFrame from a dictionary whose values are column data. Each list below becomes a column, and pandas assigns a default integer index to the rows.
import pandas as pd
data = {
"name": ["Ari", "Bea", "Chen"],
"team": ["North", "South", "North"],
"sales": [120, 95, 140],
}
df = pd.DataFrame(data)
print(df)
Read a CSV file
For a comma-separated file, pass its path to read_csv. The path is interpreted relative to the script’s current working directory unless you provide an absolute path.
Rank #2
df = pd.read_csv("sales.csv")
CSV files can contain quirks such as a different delimiter, text encoding, or missing-value markers. If the parsed columns or values look wrong, check the file’s format and the relevant read_csv options in the pandas documentation. For Excel files, pandas also provides Excel-reading and writing functions; a suitable optional engine may be needed for a particular workbook format. Data can also be read from SQL sources or URLs, subject to the right database driver, access, and format.
How do I inspect data before changing it?
Inspect a newly loaded table before making assumptions about its contents. These methods answer different questions:
df.head()shows the first rows;df.tail()shows the last rows.df.shapereturns a pair: row count followed by column count.df.info()summarizes columns, non-null counts, and data types.df.describe()provides descriptive statistics for eligible numeric columns by default.
For example, a column that should contain numbers may have been read as text because of symbols or inconsistent entries. Checking info() alongside a few actual rows can reveal that before calculations produce confusing results.
How do I select rows and columns with loc and iloc?
loc selects by index and column labels; iloc selects by integer position. These are different addressing systems, so choose based on whether you mean a label or an ordinal position.
Select by label with loc
# The row index label is 1; the column label is "name".
second_name = df.loc[1, "name"]
# Select all rows whose team label is "North" and return two columns.
north_names_and_sales = df.loc[df["team"] == "North", ["name", "sales"]]
Select by position with iloc
# Row position 1 (the second row), column position 0 (the first column).
second_row_first_column = df.iloc[1, 0]
For straightforward column selection, use the column label directly: df["sales"] returns that column as a Series. A comparison such as df["sales"] > 100 creates a Boolean condition that can filter the table:
high_sales = df[df["sales"] > 100]
How do I handle missing values in pandas?
Use isna() to identify missing entries and sum() to count them by column:
missing_by_column = df.isna().sum()
There are two common responses, and they change the data in different ways:
Free tools Windows power users keep installed
One-click scans. No signup required.
- Drop incomplete rows:
df.dropna()removes rows with missing values under its default behavior. This is simple, but can discard useful observations and reduce the data available for analysis. - Fill missing values:
df.fillna({"sales": 0})replaces missing values in the named column with zero. Do this only when zero is a defensible meaning for the missing entry; otherwise, filling can distort totals and averages.
Decide what a missing value means in the context of the data before choosing a treatment. A blank may mean “not recorded,” “not applicable,” or something else; those cases should not automatically be treated as equivalent.
How do I summarize, combine, and reshape data?
Group rows with groupby
Use groupby to calculate summaries for categories. This example computes total sales by team:
sales_by_team = df.groupby("team")["sales"].sum()
Combine tables with merge or concat
Use merge when two tables share a key and you want to match related records, much like a database join. Use concat when you want to place compatible tables together along a row or column axis.
# Match each order to a customer using the shared customer_id column.
orders_with_customers = orders.merge(customers, on="customer_id")
# Stack tables with compatible columns vertically.
all_months = pd.concat([january, february], ignore_index=True)
Before combining, check whether the key is unique where you expect it to be: duplicate keys can multiply rows in a merge. Also confirm that the tables’ column names and data types align with the result you intend.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsReshape with a pivot table
A pivot table turns categories into a compact summary layout. For example, if a table has team, month, and sales columns, you can summarize sales by team and month:
Best Value
monthly_sales = df.pivot_table(
index="team",
columns="month",
values="sales",
aggfunc="sum",
)
Grouping and pivoting are useful when the question is about totals or other summaries rather than individual records. Choose the aggregation deliberately; a pivot needs an aggregation when multiple input rows map to the same output cell.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How do I save results, plot, or work with dates?
Save a DataFrame to CSV with to_csv. Setting index=False avoids writing the row index as an extra column when that index is not part of the data you want to export.
high_sales.to_csv("high_sales.csv", index=False)
Pandas can also read and write Excel data, and it includes tools for date and time series work. For a date column, parse it as dates when loading where appropriate, or convert it before date-based operations. Pandas plotting methods can produce quick charts and use Matplotlib for plotting; Matplotlib provides additional control when you need to customize a visualization.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallWhat should I check when pandas behaves unexpectedly?
- Unexpected row counts: inspect the input file, filtering condition, and index; a merge on duplicate keys can expand the result.
- Wrong-looking values or types: compare a few rows with
df.info()and check how the source file represents delimiters, missing entries, and numbers. - Missing-value surprises: count nulls before and after cleaning, and verify that a fill value is meaningful for that column.
- Slow work on larger tables: inspect only the columns and rows you need, prefer vectorized column operations over row-by-row Python loops, and avoid loading unnecessary data when the file-reading options can limit it.
- Installation or import errors: confirm that the package was installed in the environment used by the interpreter or notebook, then use the official installation guide for environment-specific instructions.
Where can I continue learning pandas?
Python Guides offers a free pandas course covering installation with pip and Conda, core objects, file input, selection, missing data, grouping, dates, and visualization: Python Pandas Training Course. It is a guided sequence rather than a substitute for release-specific documentation.
For a longer-form reference, Wes McKinney’s Python for Data Analysis, 3rd Edition covers pandas alongside broader data-analysis topics, including loading, cleaning, merging, grouping, visualization, and time series. O’Reilly says this edition is updated for Python 3.10 and pandas 1.4, so use current pandas documentation when you need API guidance for later releases: O’Reilly book listing.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

