What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Sweetviz 2.0 made automated exploratory data analysis (EDA) more useful in notebooks by adding show_notebook(), report scaling, vertical layouts, and optional HTML output. The package is no longer at version 2.0—the PyPI release history lists Sweetviz 2.3.3, released April 11, 2026—so the examples below use current-compatible syntax while explaining what changed in 2.0. Install the current release unless you are reproducing an older environment.
Sweetviz is a pandas-focused, open-source profiler. It creates a visual, self-contained HTML report that helps you inspect schema, missing data, distributions, duplicates, feature relationships, train/test differences, and selected target behavior. It accelerates the first pass; it does not replace domain knowledge, statistical testing, data cleaning, or production monitoring.
What exploratory data analysis means
EDA is the structured inspection you perform before modeling or drawing conclusions. You examine column types, unique values, missingness, distributions, outliers, duplicates, relationships, and differences between datasets or groups. The purpose is to find assumptions and anomalies worth investigating—not to prove causation or automatically decide what to do with the data.
What Sweetviz does
Sweetviz accepts pandas DataFrames and produces a high-density visual report. Its project description highlights rapid dataset characterization, target analysis, and training-versus-test comparison. The report combines numerical and categorical summaries with mixed-type association measures. See the current package description at PyPI.
#1 Best Overall
- Wiley
- Language: english
- Book - storytelling with data: a data visualization guide for business professionals
- Dataset-level row, feature, missing-value, duplicate, and type summaries
- Unique-value counts and frequent values
- Numerical statistics such as range, quartiles, mean, median, standard deviation, skewness, kurtosis, and coefficient of variation
- Visual distributions and feature-level details
- Comparisons between datasets or subgroups
- Associations using Pearson correlation, uncertainty coefficient, and correlation ratio, depending on feature types
An association score is descriptive. It is not causation, statistical significance, or model feature importance.
What changed in Sweetviz 2.0—and what changed later
The 2.0 release addressed a practical limitation of earlier versions: notebook users no longer had to rely only on an externally opened HTML file. It added:
show_notebook()for embedded Jupyter and Google Colab display- Iframe sizing and fractional scaling
- A vertical report layout
- Optional saving of the displayed report to an HTML path
Those are historical 2.0 additions, not a claim that 2.0 is current. Later project notes identify Comet.ml support in 2.1, compatibility updates in 2.2, and a verbosity parameter and fixes in 2.3.0. The checked PyPI history lists 2.3.3 as released April 11, 2026. Review the current release page when pinning an environment.
Install Sweetviz in the right Python environment
For a project, create and activate a virtual environment, then install with the interpreter that will run your code:
python -m pip install sweetviz
In a notebook, install into the active kernel environment:
!pip install sweetviz
Verify the interpreter and package after installation:
Rank #2
import sys
import sweetviz as sv
print(sys.executable)
print(sv.__version__)
Pin the version in a project requirements file when reproducibility matters. Minimum Python and pandas requirements have changed across releases, so use the metadata for the exact version you select rather than relying on an old tutorial.
Create your first standalone EDA report
Load a DataFrame, create a report object, then render it:
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsimport pandas as pd
import sweetviz as sv
df = pd.read_csv("data.csv")
report = sv.analyze(df)
report.show_html("eda_report.html")
show_html() writes a self-contained HTML report. If you omit the path, the documented default is SWEETVIZ_REPORT.html. Open the resulting file in a browser to inspect the dataset overview and feature sections.
Display the report in Jupyter or Colab
Sweetviz 2.0’s notebook-oriented method embeds the report in the output area:
report = sv.analyze(df)
report.show_notebook()
For a large or narrow notebook cell, control the iframe and layout:
report.show_notebook(
w="100%",
h=700,
scale=0.8,
layout="vertical",
filepath="eda_report.html"
)
wcontrols width, such as"100%"or a pixel value.hcontrols height, such as700or"Full".scalechanges the displayed report size.layoutselects the documented widescreen or vertical presentation; defaults can vary between release lines.filepathoptionally saves an HTML copy.
If embedding fails, use show_html() with an explicit path. Notebook frontends, browser security policies, and package versions do not all render iframes identically.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Analyze a target column
Pass a target feature when it is Boolean or numerical:
report = sv.analyze(df, target_feat="target")
report.show_html("target_report.html")
Current package documentation describes Boolean and numerical target support. Do not assume that every multiclass categorical target is accepted. For a multiclass outcome, compare groups explicitly or use a profiling workflow that documents multiclass support.
Target analysis is diagnostic: it can reveal distribution differences and associations around the target, but it does not establish leakage, causality, or predictive value.
Compare training and test data
Name each DataFrame so the report labels remain unambiguous:
Free tools Windows power users keep installed
One-click scans. No signup required.
comparison = sv.compare(
[train_df, "Training"],
[test_df, "Test"]
)
comparison.show_html("train_test_report.html")
Look for changed distributions, missing-value rates, category appearance, feature presence, and suspiciously different cardinality. These are clues, not proof that a split is invalid or unrepresentative. Also inspect post-outcome columns, target-derived variables, and duplicate records crossing the split.
You can include a supported target in the comparison:
Rank #4
comparison = sv.compare(
[train_df, "Training"],
[test_df, "Test"],
"target"
)
comparison.show_html("comparison_with_target.html")
Compare two subgroups in one DataFrame
compare_intra() takes a DataFrame and a Boolean mask, then compares the matching and non-matching rows:
group_report = sv.compare_intra(
df,
df["gender"] == "female",
["Female", "Male"]
)
group_report.show_html("group_comparison.html")
Use this for a descriptive subgroup check. Differences may reflect sampling, measurement, missingness, or real population variation; investigate with domain context.
Correct automatic type inference
Inference is convenient but can misclassify identifiers, codes, dates, ordinal categories, or text. Configure feature meaning explicitly:
feature_config = sv.FeatureConfig(
skip="PassengerId",
force_text=["Age"]
)
report = sv.analyze(df, feat_cfg=feature_config)
report.show_html("configured_report.html")
Available controls include skip, force_cat, force_num, and force_text. Exclude columns such as customer_id, row_number, or transaction_id unless their values have analytical meaning. Convert raw date strings into features such as year, month, weekday, or elapsed time before profiling them.
Reduce pairwise work for wide data
Pairwise associations can make generation slower and the report harder to read:
report = sv.analyze(
df,
pairwise_analysis="off"
)
report.show_html("basic_report.html")
For large or wide data, also consider excluding irrelevant identifiers, profiling a representative sample, or analyzing related feature groups separately. Sampling can hide rare categories and tail behavior, so validate important findings against the full data.
How to read the report responsibly
- Start with integrity: check row counts, types, duplicates, missingness, and impossible values.
- Inspect severe missingness: determine whether missing values are structural, informative, or caused by collection problems.
- Review distributions and outliers: confirm that unusual values are not unit, parsing, or data-entry errors.
- Check train/test or subgroup shifts: investigate differences before treating them as leakage or drift.
- Review associations: use them to choose follow-up plots or tests, not to claim causation.
- Challenge target-related features: verify that every feature would be available at prediction time.
Sweetviz surfaces signals; pandas, SciPy, Matplotlib, Seaborn, and domain-specific methods remain useful for detailed follow-up.
Troubleshooting common failures
ModuleNotFoundError
pip may have installed into a different interpreter. Run python -m pip install sweetviz, compare sys.executable with the notebook kernel, and restart the kernel.
sweetviz has no attribute analyze
Rename a local script or folder named sweetviz.py, remove related __pycache__ or .pyc files, and retry. Local names can shadow the installed package.
The notebook iframe is blank or clipped
Use show_html() as a fallback, provide an explicit filepath, increase h or w, reduce scale, and confirm that the file exists. Reinstall Sweetviz in the same environment as the kernel if necessary.
NumPy compatibility errors
Environment-specific failures have been reported, including an issue involving numpy.warnings. Use a clean virtual environment, pin compatible versions, and consult the project issue when the traceback is version-specific.
The report is slow or enormous
Disable pairwise analysis, remove non-analytic identifiers, reduce the feature set, or profile a carefully chosen sample. Large self-contained HTML files can also delay browser rendering.
Privacy and sharing
A self-contained report is portable, but it may expose raw values, category labels, rare groups, and sensitive distributions. Review the HTML before sending it outside your team, storing it in a shared artifact system, or attaching it to an issue.
Sweetviz compared with alternatives
| Tool | Best fit | Trade-off |
|---|---|---|
| Sweetviz | Fast visual first-pass reports, train/test and subgroup comparisons | Preselected summaries; limited production validation and customization |
| ydata-profiling | Broader automated profiling and data-quality exploration | Reports can be heavier and more computationally demanding |
| DataPrep | Convenient automated exploratory reports | Check current maintenance, Python support, and report behavior |
| D-Tale | Interactive browser inspection and manipulation of DataFrames | Less suited to a portable, self-contained report |
| pandas with Seaborn/Matplotlib | Full control over transformations, tests, visuals, and business logic | More code and manual work |
| Great Expectations or similar validators | Repeatable checks in pipelines and CI/CD | Answers whether rules pass, not what exploratory patterns exist |
When Sweetviz is the right choice
- Your data already fits comfortably in a pandas DataFrame.
- You need a quick, shareable HTML overview.
- Train/test or subgroup comparison is central to the investigation.
- You want useful visual output with minimal code.
Choose another approach when data is distributed or out-of-core, reports must become customized production components, automated CI checks are the primary goal, formal statistical tests are required, or sensitive data cannot be safely written into an HTML artifact.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →The Bottom Line
Sweetviz is best used as a fast first-pass EDA accelerator. Sweetviz 2.0’s notebook display made that workflow more convenient, while current releases add later compatibility and feature changes. Use it to find questions, then verify important findings with targeted analysis, domain knowledge, and appropriate validation.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

