Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a pandas DataFrame workflow, the five common routes are pandas.read_csv() for delimited text, pandas.read_json() for JSON, pandas.read_excel() for spreadsheets, pandas.read_sql() or its query/table variants for databases, and pandas.read_parquet() for Parquet files. Choose based on the source format and the structure you need; CSV’s built-in csv module is a good alternative when you want to handle records directly instead of building a DataFrame.

How pandas loading works

pandas groups these operations under its I/O API. Its reader functions, such as pandas.read_csv(), generally return pandas objects that you can inspect and analyze. The source format determines which reader to use, while the input’s structure, parser options, and installed dependencies determine how to use it. See the pandas I/O guide for the current API and format-specific details.

1. Load CSV and other delimited text

Read a file into a DataFrame

Use read_csv() for comma-separated files and other delimited text. It accepts paths, URLs, and file-like objects.

import pandas as pd

df = pd.read_csv("data.csv")

If the file uses a delimiter other than a comma, set sep. For example, a tab-delimited file can be read with pd.read_csv("data.tsv", sep="t"). Check the file’s header, quoting, encoding, and missing-value conventions when the resulting columns or values look unexpected; producers can differ in these details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Handle rows directly with the standard library

When you want to process records one at a time rather than create a DataFrame, Python includes the csv module. csv.reader returns rows, while csv.DictReader lets you access row values by field name.

import csv

with open("data.csv", newline="", encoding="utf-8") as file:
    for row in csv.DictReader(file):
        print(row["name"])

The Python documentation recommends opening CSV files with newline="". CSV is widely used for exchanging spreadsheet and database data, but it does not have one tightly defined standard, so applications may handle quoting and other dialect details differently. See the Python CSV module documentation.

2. Load JSON into pandas

Read a JSON source

Use read_json() when you want JSON data represented as a pandas object.

import pandas as pd

df = pd.read_json("data.json")

JSON can be nested or organized in different ways. After loading it, inspect the result’s columns, index, and data types to confirm that its structure suits your analysis. If the shape is not what you expect, consult the pandas I/O guide for the options that apply to your input.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

3. Read an Excel workbook

Select a sheet

Use read_excel() and specify the worksheet when you want a particular sheet:

import pandas as pd

df = pd.read_excel("workbook.xlsx", sheet_name="Sheet1")

The file format affects which reader engine pandas uses, and the required engine must be installed in your environment. The pandas 3.0.6 I/O documentation describes openpyxl for .xlsx, xlrd for .xls, and pyxlsb for .xlsb; it also describes calamine as supporting the listed Excel and OpenDocument formats. Check the current pandas Excel documentation for the format and setup you use.

4. Load data from a SQL database

Choose a query or table

Use read_sql_query() when you want to provide a SQL query, or read_sql_table() when you want to read a table. read_sql() is a convenience wrapper.

import pandas as pd
import sqlite3

with sqlite3.connect("app.db") as connection:
    df = pd.read_sql_query("SELECT name, email FROM users", connection)

SQLite connections can use Python’s standard-library sqlite3 module. Other databases need a suitable connection layer, such as SQLAlchemy together with the driver for the database. Keep credentials out of source code, and parameterize values in application queries rather than joining untrusted input into SQL strings. See the pandas SQL documentation for connection and reader details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

5. Read a Parquet file

Load columnar data

For a Parquet file, use read_parquet():

import pandas as pd

df = pd.read_parquet("data.parquet")

Parquet is a columnar file format. Reading it with pandas may require a compatible Parquet engine to be installed; check the current pandas Parquet documentation for the setup instructions that match your environment.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Compare the five methods

Method Source Typical result Setup to check Useful controls or considerations
read_csv() CSV or other delimited text DataFrame pandas Set sep for a non-comma delimiter; check header, quoting, encoding, and missing values. Paths, URLs, and file-like objects are accepted.
read_json() JSON Pandas object pandas Inspect the input’s nesting and the resulting columns, index, and types.
read_excel() Excel workbook DataFrame for a selected sheet pandas plus an installed engine appropriate to the workbook format Choose a sheet with sheet_name; engine requirements vary by format.
read_sql_query(), read_sql_table(), or read_sql() SQL database query or table DataFrame SQLite uses Python’s standard library; other databases need suitable connection support and often a driver Use a query for selected or filtered data, or a table reader when reading a table.
read_parquet() Parquet file DataFrame pandas plus a compatible Parquet engine as required by the environment Check current engine installation guidance for your setup.

How to choose a loading method

  • Choose read_csv() for delimited text, especially when you need to tune the separator or related parsing assumptions.
  • Choose read_json() when the source is JSON, then verify that the loaded structure fits your analysis.
  • Choose read_excel() for workbook data and confirm the sheet name and installed engine.
  • Choose a SQL reader when the data lives in a database and you need a query or table; make sure a suitable connection is available.
  • Choose read_parquet() for Parquet files and check the environment’s engine requirements.
  • Choose Python’s csv module instead of pandas when direct row-by-row handling is the better fit.

These methods address different formats and workflows. The documentation does not establish a controlled, apples-to-apples speed ranking across all five, so choose by compatibility and the structure you need rather than an assumed performance order.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.