Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Exploratory data analysis (EDA) is how you learn what a dataset contains before deciding how to model it: inspect its structure, summarize variables, visualize distributions and relationships, and investigate surprises. EDA is an open-ended approach, not a fixed checklist. It helps reveal patterns and shape questions; it does not by itself confirm a hypothesis or show that one factor caused another.

What is exploratory data analysis?

EDA is an approach to understanding data through direct inspection, numerical summaries, and—especially—visualization. The National Institute of Standards and Technology (NIST) describes it as a philosophy rather than a prescribed set of techniques. Its goals include identifying important variables, finding outliers or anomalies, examining assumptions, and informing a parsimonious model. See NIST’s overview of EDA.

EDA combines graphics with simple statistics. A mean or median can orient you, but it cannot show every feature of a distribution. A plot may reveal skew, gaps, multiple clusters, extreme observations, or subgroups that a single summary conceals. Numerical and visual views complement one another; neither makes the other unnecessary.

Why explore before choosing a model?

EDA helps you understand what the data can reasonably support before you commit to a model or interpret its results. NIST contrasts EDA’s sequence—problem, data, analysis, model, conclusions—with classical analysis, in which a model is specified before analysis. In NIST’s words, “For EDA, the data collection is not followed by a model imposition; rather it is followed immediately by analysis with a goal of inferring what model would be appropriate.” Read NIST’s comparison of EDA and classical analysis.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Storytelling with Data: A Data Visualization Guide for Business Professionals
  • Wiley
  • Language: english
  • Book - storytelling with data: a data visualization guide for business professionals

Exploration is not a substitute for formal testing. If you search a dataset, discover a pattern, and then test that same pattern as though it had been specified in advance, the search has influenced the claim. Treat exploratory findings as leads for hypotheses, models, or follow-up studies—not as automatic confirmation or evidence of causation.

A practical first-pass EDA workflow

This sequence is a useful way to organize a first pass, not a universal standard. Adapt it to the question, data source, and consequences of being wrong.

1. Establish what the rows and columns mean

Before interpreting values, establish what one row represents, what each column measures, and the units, time period, collection method, and intended population. Inspect the dataset’s dimensions, column names, data types, and plausible ranges. A column labeled “age,” for example, is difficult to interpret responsibly without knowing whether it is measured in years, at what date, and for whom.

2. Check data quality and coverage

Look for missing values, duplicate records, inconsistent category labels, and values that appear implausible. Also ask who or what is absent: incomplete coverage or a biased sampling process can distort the picture even when individual entries are valid. Treat anomalies as prompts to investigate rather than proof that a record is wrong.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

3. Summarize and plot variables individually

For categorical columns, inspect counts and proportions. For numerical columns, choose measures of location and spread that suit the data, then look at the distribution rather than relying on a single statistic. Histograms can expose shape, gaps, and multiple modes; box plots can make spread and unusual values easier to compare. These are options, not a requirement to create every plot for every column.

4. Explore relationships relevant to the question

Compare variables with displays suited to their types and structure. Look for relationships that matter to the problem, and check whether an apparent pattern changes across groups or over time. Notice when a relationship depends on only a few observations or is obscured by overplotting. NIST’s gallery of EDA techniques organizes graphical and quantitative methods around different problem types.

5. Record findings and follow-up questions

Keep a record of the choices you made, surprising observations, possible explanations, and analyses to try next. This makes it easier to distinguish ideas that emerged during exploration from questions specified in advance. If an exploratory pattern is important, plan an appropriate confirmatory analysis or collect independent evidence rather than treating repeated searches of the same data as confirmation.

Which plots should you use for EDA?

Choose a display by the question it should answer, not by a rule that one chart is best. Consider the variable types, how many variables you are comparing, whether observations have a meaningful order or time dimension, and whether the goal is to inspect a distribution or a relationship. Sample size matters too: dense plots can hide observations, while a display that separates groups may make subgroup structure clearer.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Histogram: inspect the shape of a numerical distribution, including skew, gaps, or multiple modes.
  • Box plot: compare distributions compactly and identify observations that stand apart from the central spread.
  • Probability plot: assess how observed values compare with a reference distribution; it can help examine distributional assumptions.
  • Plots of simple statistics: compare summaries where that comparison answers the question, while remembering that summaries omit detail from the raw observations.

NIST’s handbook covers raw-data plots, including histograms and probability plots, as well as plots of simple statistics such as box plots. Its technique gallery offers further examples; a dataset rarely needs every technique.

How should you investigate outliers and assumptions?

An unusual value may be a data-entry mistake, a unit mismatch, a sensor problem, a join error, a member of a meaningful subgroup, or a rare but valid event. Check its source and context before deciding whether to correct, exclude, or retain it. If you change or remove observations, document what you did and why.

EDA can also expose assumptions worth examining before modeling—for example, whether a distribution or relationship behaves as expected. A visual signal is a reason to investigate, not a verdict on its own. Use methods appropriate to the question to evaluate assumptions, and keep exploratory observations distinct from later formal conclusions.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Using pandas to support an EDA workflow

Python is one option, not a prerequisite. The pandas project describes its core structures as Series and DataFrame and documents common tasks for cleaning and analyzing data and preparing results for plots or tables. Its official documentation includes guidance on missing data, descriptive statistics, and chart visualization. Documentation is version-specific; the surfaced documentation is for pandas 3.0.6, so check which version you use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A small first inspection in pandas might look like this:

import pandas as pd

# Replace "data.csv" with the path to your dataset.
df = pd.read_csv("data.csv")

print(df.shape)          # rows and columns
print(df.dtypes)         # inferred column types
print(df.head())         # sample records
print(df.isna().sum())   # missing-value count by column
print(df.describe(include="all"))

These commands provide an initial view, not an assessment of whether the data is representative or the analysis is statistically sound. Verify the row meaning, type assumptions, units, and missing-data patterns against the data’s context. Consult pandas’ user guide for details on missing data, descriptive statistics, and visualization.

Further reading

NIST identifies John W. Tukey’s 1977 book Exploratory Data Analysis as the seminal work on the subject. Its handbook chapter on EDA was published June 1, 2003, and was written by N. Alan Heckert and James J. Filliben; see the NIST publication record.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.