The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
DuckDB can make some pandas-based analytics faster by running SQL directly against a DataFrame, scanning only the columns a query needs, and using multiple CPU threads. But “10x faster” is not a general guarantee: the gain depends on the operation, data, file layout, hardware, and whether loading and result conversion count in the timing. The practical question is whether DuckDB improves your complete workflow, not just one query.
What changes when you use DuckDB with pandas?
You can query a pandas DataFrame from Python with DuckDB SQL without first copying it into a separate database table. DuckDB’s replacement-scan behavior resolves a Python variable by name, reads its columns and types, and lets you return the result as another DataFrame.
pip install duckdb
import duckdb
result = duckdb.query("SELECT sum(a) FROM mydf").to_df()
Here, mydf is an existing pandas DataFrame. This is useful when an operation is naturally expressed as an aggregation, filter, join, or other SQL query. It does not translate arbitrary pandas code into SQL, nor does it make every pandas API operation interchangeable with DuckDB.
When is a switch most likely to help?
DuckDB is worth testing when your script spends substantial time on analytical operations such as aggregations, joins, sorts, or window calculations. It can also query supported files such as Parquet directly, avoiding the need to load every column into a pandas DataFrame when only a subset is required.
#1 Best Overall
- Large analytical queries: SQL execution and parallel processing may help with operations over many rows.
- Column-selective file scans: A query against Parquet can read only the columns it needs.
- Repeated analysis: If you repeatedly query the same data and have storage available, loading it into a DuckDB database may suit the workload better than repeatedly scanning files.
- Small or simple transformations: A switch may not help. In a 2025 academic evaluation, pandas performed best for small datasets in that study; the evaluation did not establish a universal DuckDB-versus-pandas ranking.
File layout and row-group size, filtering, data size relative to RAM, thread count, and whether results are converted back to pandas can all affect the outcome. DuckDB’s file-format guide reports that, in its own TPC-H microbenchmark, queries on Parquet files ran approximately 1.1–5.0× slower than queries on a DuckDB database. That is a DuckDB-versus-Parquet comparison, not a DuckDB-versus-pandas result; it illustrates why data storage and query patterns matter.
What DuckDB’s published benchmarks actually show
DuckDB’s 2021 pandas comparison used TPC-H’s lineitem and orders tables, approximately 1 GB of uncompressed CSV data, and Google Colab. It compared selected aggregations and a join over DataFrames, with DuckDB set to one or two threads because the Colab environment supported two. The article also compared direct Parquet queries with reading Parquet into pandas.
Rank #2
That setup is evidence for those tested queries and conditions, not proof that a typical analytics script will become 10× faster. Dataset size, query shape, format, engine versions, memory, thread count, and whether input loading and output conversion are included can change the result.
DuckDB’s 2024 benchmark-history article makes another useful distinction: it measured raw query speed separately from import and export performance across pandas, Arrow, and Parquet. Its replacement-scan benchmark read one column from a 100-million-row, 5 GB dataset and calculated one aggregate, focusing on scan speed rather than aggregation or output conversion. A benchmark result is only meaningful when you know what the timer includes.
How to benchmark the switch fairly
Compare the same expected output on representative data. If pandas reads a file and creates a DataFrame while DuckDB scans the file directly, report that workflow difference rather than presenting the query-only times as equivalent. For an end-to-end comparison, include all work required to produce the result in each version.
- Choose a representative workload. Use the real input format, data size, and operation—such as a join or aggregation—that matters to your script.
- Record the environment. Note CPU, memory, Python, pandas, and DuckDB versions; DuckDB’s thread count; and whether caches are warm or cold.
- Separate the timings. Measure file reading, conversion or loading, query execution, and output conversion individually. Also report end-to-end time and peak memory.
- Repeat runs and explain the statistic. Run each version more than once and state whether you report a median, mean, or another measure.
- Check the plan when results surprise you. Use
EXPLAINto inspect the plan andEXPLAIN ANALYZEto profile execution. DuckDB’s tuning guide notes that multithreaded step times can add up to more than total wall-clock time.
A credible “10× faster” claim should identify the specific query, data, environment, and timing boundary. Without those details, treat 10× as a result to verify on your own workload, not an expected outcome.
Rank #4
What happens when the data does not fit in memory?
DuckDB can spill some larger-than-memory work to disk, including grouping, joins, sorting, and window operations. That can make it useful when a workload exceeds available RAM, but it is not a guarantee that every query will complete: multiple blocking operators in one query can still cause out-of-memory errors, and some aggregates, including list() and string_agg(), do not support disk offload.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Disk spilling also means temporary storage capacity can matter. DuckDB documents a configurable temporary directory; an external drive is not a requirement, and the documentation does not promise that buying one will make a query faster. The tuning guide also warns that increasing the thread count can slow some workloads, so measure thread settings rather than assuming more is better.
Best Value
Which tool should you keep?
Keep pandas when its APIs fit your workflow and your operations are already fast enough, especially for smaller data. Try DuckDB when profiling identifies substantial analytical query work, when direct file scans can avoid unnecessary loading, or when SQL makes a complex transformation clearer. You can also combine them: use DuckDB for a query and return a DataFrame for downstream pandas code.
Make the decision using the same output and complete work on both sides. Compare the operation, input format and layout, data size relative to RAM, thread configuration, loading and conversion costs, peak memory, and the effort of changing code. A faster central query alone does not establish that the whole script is faster.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

