Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To make a pandas DataFrame use less memory, first measure its columns with df.memory_usage(deep=True), then selectively convert repeated text to categorical data, downcast numeric columns only when their ranges and precision allow it, or use sparse types for genuinely sparse data. If you mean a smaller saved file, optimize Parquet separately: compression reduces file bytes, not necessarily the memory required when the data is loaded.

How to find which columns use the most memory

Start by measuring the current frame. The result is an estimate of bytes per column; the index is included by default. deep=True inspects values stored in object-dtype columns, which ordinary accounting may undercount. The deeper inspection can take more time, and its total is not a measurement of the entire Python process’s resident memory.

usage = df.memory_usage(deep=True).sort_values(ascending=False)
print(usage)
print(f"Total: {usage.sum():,} bytes")

To exclude the index from the report, pass index=False. Pandas’ FAQ explains that the “+” shown in some memory reports means actual use could be higher because values in object columns are not counted by ordinary accounting. See the memory_usage API documentation and the pandas memory FAQ.

Which dtype changes can reduce DataFrame memory?

Convert repeated, low-cardinality text to category

A categorical stores the distinct categories separately and represents rows with codes. It is often a good candidate when a long column repeats a relatively small set of labels, such as status, region, or group. The benefit depends on both the number of rows and the number of distinct categories: a near-unique text column may not shrink and can use more memory.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sandisk 2TB Extreme Portable SSD, Up to 1050MB/s, USB-C, USB 3.2 Gen 2, IP65 Water and Dust Resistance, Updated Firmware, External Solid State Drive, SDSSDE61-2T00-G25
  • Get NVMe solid state performance with up to 1050MB/s read and 1000MB/s write speeds in a portable, high-capacity drive(1) (Based on internal testing; performance may be lower depending on host device & other factors. 1MB=1,000,000 bytes.)
  • Up to 3-meter drop protection and IP65 water and dust resistance mean this tough drive can take a beating(3) (Previously rated for 2-meter drop protection and IP55 rating. Now qualified for the higher, stated specs.)
  • Use the handy carabiner loop to secure it to your belt loop or backpack for extra peace of mind.
  • Help keep private content private with the included password protection featuring 256‐bit AES hardware encryption.(3)
  • Easily manage files and automatically free up space with the SanDisk Memory Zone app.(5). Non-Operating Temperature -20°C to 85°C
before = df.memory_usage(deep=True).sum()
df["group"] = df["group"].astype("category")
after = df.memory_usage(deep=True).sum()
print(f"Before: {before:,} bytes")
print(f"After:  {after:,} bytes")

Compare the measured result on your own data, and retain the conversion only if categorical semantics suit the column and the full workload. See pandas’ categorical data guide.

Downcast numeric columns only after checking their requirements

Smaller integer and floating-point dtypes can reduce memory, but a narrower dtype may not represent every value or the precision your calculations require. Check each column’s minimum and maximum, missing-value behavior, and required numerical precision before converting. Pandas provides pd.to_numeric(..., downcast=...) to select a smaller compatible numeric dtype; measure again after the conversion.

# Example only: choose downcast options after validating each column.
df["id"] = pd.to_numeric(df["id"], downcast="unsigned")
df["amount"] = pd.to_numeric(df["amount"], downcast="float")
print(df.memory_usage(deep=True).sort_values(ascending=False))

The exact safe type depends on the data and calculations, so do not treat a downcast choice as universally suitable. Pandas’ scaling guide demonstrates downcasting on a particular generated dataset; its result is an illustration, not a general memory or speed guarantee.

Rank #2
Sandisk 1TB Portable SSD, Up to 800MB/s Read Speeds, Black (Old Model)
  • Solid state performance with up to 800MB/s read speeds in a portable drive. (Based on internal testing; performance may be lower depending on host device, interface, usage conditions and other factors. 1MB=1,000,000 bytes.)
  • Back up your content and memories on a storage solution that fits seamlessly into your mobile lifestyle.
  • Take it with you on your adventures—up to two-meter drop protection means this durable drive can take a beating. (Based on internal testing.)
  • Secure it to your belt loop or backpack for extra peace of mind thanks to the tough rubber hook.
  • From Sandisk, a brand professional photographers trust to take on assignments.

Use sparse types when most values are the fill value

Sparse storage is intended for columns or matrices in which most entries are the fill value, commonly zero. For dense data, it may not save memory, and storage savings do not guarantee that every operation will be faster or supported in the same way.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
print(df.sparse.density)

Check that the data are sparse, measure memory after using a suitable SparseDtype, and test the operations your application actually needs. Pandas documents the sparse accessor and density in its sparse API.

What does pandas’ scaling example show?

The pandas 3.0.6 scaling guide uses a generated frame with 1,051,201 rows. After converting a repeated name field to category and downcasting numeric columns, it reports a new-to-original deep-memory ratio of 0.42—about 42% of the original memory for that example. The same passage describes the result as “1/5” of the original, which does not agree with the printed ratio. The ratio is the consistent figure to use; neither number predicts the result for another dataset.

Rank #3
Sale
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
  • Easily store and access 2TB to content on the go with the Seagate Portable Drive, a USB external hard drive
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition no software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.

How to make a saved Parquet file smaller

File size and in-memory DataFrame size are different measurements. Parquet is a columnar binary format, and compression can reduce bytes on disk without producing an equivalent reduction in memory after loading. Pandas’ to_parquet requires either pyarrow or fastparquet; compare the resulting file size, load time, and loaded dtypes for your actual use case.

  1. Choose Parquet and an available engine, then try an appropriate compression option with DataFrame.to_parquet.

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  2. Decide explicitly whether to write the index; index serialization changes the saved representation.

    Rank #4
    Sale
    Sandisk 1TB Extreme Portable SSD, Up to 2000MB/s Transfer Speeds-New Model
    • NEARLY 2X FASTER THAN OUR PREVIOUS GENERATION(8) – move 1,000 high-res photos in under 60 seconds(6) with up to 2000MB/s transfer speeds(2).
    • IP65 RATING AND UP TO 3M DROP PROTECTION(3) – protects against spills and drops.
    • POCKET-SIZED – fits easily in pockets and small bags.
    • SPACE TO OWN YOUR AI CONTENT – speed and capacity to download your high-res clips and photo edits.
    • 256-BIT AES ENCRYPTION(4) – helps keep private files secure with password protection.
  3. If categorical columns contain unused categories, consider removing them before writing. The extra category metadata can enlarge the output.

  4. Read the file back and check its size, dtypes, and behavior against the application’s needs.

See the to_parquet API documentation and pandas’ Parquet I/O guide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
  • Easily store and access 5TB of content on the go with the Seagate portable drive, a USB external hard Drive
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

A practical order of operations

  1. Record per-column and total deep memory usage, including or excluding the index deliberately.

  2. Identify repeated text, numeric columns with safely reducible ranges or precision, and genuinely sparse data.

  3. Make one dtype change at a time and measure again, checking missing values and value fidelity.

  4. Run representative operations, such as grouping or serialization, before keeping a conversion.

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  5. If disk space is the problem, optimize and validate the Parquet representation separately.

Chunked processing can help with some large-data workflows, but pandas notes that operations such as DataFrame.groupby() are harder to perform chunkwise; chunking is not a universal fix for memory limits.

Quick Recap

Bestseller No. 2
Sandisk 1TB Portable SSD, Up to 800MB/s Read Speeds, Black (Old Model)
Sandisk 1TB Portable SSD, Up to 800MB/s Read Speeds, Black (Old Model)
From Sandisk, a brand professional photographers trust to take on assignments.
$165.70
SaleBestseller No. 3
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$119.99
SaleBestseller No. 4
Sandisk 1TB Extreme Portable SSD, Up to 2000MB/s Transfer Speeds-New Model
Sandisk 1TB Extreme Portable SSD, Up to 2000MB/s Transfer Speeds-New Model
IP65 RATING AND UP TO 3M DROP PROTECTION(3) – protects against spills and drops.; POCKET-SIZED – fits easily in pockets and small bags.
$209.99
Bestseller No. 5
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$227.38

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.