What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
An outlier is an observation that differs substantially from the prevailing pattern in a dataset. It may be a measurement error, a data-quality problem, a rare but valid event, fraud, equipment failure, or a previously unseen behavior—not automatically something to delete.
PyOD (Python Outlier Detection) is an open-source Python toolkit that gives many outlier and anomaly-detection algorithms a broadly consistent interface. You can install it with pip, fit a detector, rank observations by anomaly score, and review thresholded labels.
What is an outlier?
An outlier is a data point that departs markedly from the expected pattern. “Far from the mean” is only one simple case: useful detectors also examine combinations of variables, local neighborhoods, density, projections, learned representations, and temporal context.
Common types of outliers
- Univariate: unusual in one feature, such as a transaction amount far above the usual range.
- Multivariate: each value looks ordinary alone, but the combination is unusual for the population.
- Global: unusual compared with the entire dataset.
- Local: unusual within a nearby neighborhood, even when it is not globally extreme.
- Contextual: abnormal under a particular time, location, customer segment, or operating condition. A temperature can be normal in summer but anomalous in winter.
- Collective: a sequence or group is abnormal together, although each individual point may look normal.
Detection identifies observations that satisfy a model’s criterion. It does not establish the cause. A flagged row might be a bad sensor reading, a new customer population, a process change, or the event your investigation is intended to find.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
- Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
- Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
- Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
- Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
- 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
Why detect outliers?
- Fraud, abuse, and unusual transactions
- Manufacturing faults and predictive maintenance
- Network intrusion and security monitoring
- Medical, laboratory, and scientific data review
- Data-quality and pipeline monitoring
- Customer behavior analysis and rare-event discovery
- Distribution-shift and process-change detection
The safe default is to flag and investigate rather than automatically remove. Preserve the original row, record why an action was taken, and use trusted operational evidence before correcting or excluding a case. An overview of PyOD-based outlier detection is available in Analytics Vidhya’s historical tutorial; its package and compatibility claims are now dated.
What is PyOD?
PyOD is a Python library for outlier and anomaly detection. Its common estimator-style workflow uses methods such as fit, predict, and decision_function, making it familiar to users of scikit-learn.
The current documentation describes PyOD 3.6.5 and more than 60 detectors as of August 18, 2026. The project spans tabular data plus time-series, graph, text and image embeddings, audio, ensembles, thresholding utilities, lifecycle orchestration through ADEngine, and agent-oriented workflows. Detector availability and dependencies vary by method; see the PyOD documentation and the source repository.
Most everyday PyOD usage is unsupervised: the training data has no trusted anomaly labels. The project also includes supervised or label-assisted methods such as XGBOD and DevNet. PyOD is distributed under the BSD-2-Clause license; current package metadata is listed on PyPI.
Recommended Free Tools
PyOD versus scikit-learn
scikit-learn already provides IsolationForest, LocalOutlierFactor, OneClassSVM, SGDOneClassSVM, and EllipticEnvelope. It is therefore incorrect to say that scikit-learn cannot detect outliers.
PyOD’s advantage is breadth and a dedicated ecosystem: many additional statistical, proximity, density, ensemble, neural, graph, and specialized detectors share similar usage conventions. Choose scikit-learn when its built-in estimators and pipeline tooling are sufficient; choose PyOD when you need to compare a wider set of detectors or use methods outside scikit-learn’s catalog. See scikit-learn’s outlier and novelty-detection guide.
Install PyOD
Current PyPI metadata requires Python 3.9 or newer. The basic installation is:
python -m pip install pyod
To upgrade an existing installation:
python -m pip install --upgrade pyod
A virtual environment is general Python practice, not a PyOD-specific requirement:
Rank #2
- 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
- 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
- Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
- 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
- What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.
python -m venv .venv
On macOS or Linux:
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install pyod pandas scikit-learn
On Windows PowerShell:
.venvScriptsActivate.ps1
python -m pip install --upgrade pip
python -m pip install pyod pandas scikit-learn
Optional extras include capabilities such as torch, suod, xgboost, combo, pythresh, embedding, openai, huggingface, graph, mcp, audio, and all. Install and verify the extra required by a particular detector instead of assuming the base package includes every neural or embedding dependency.
The basic PyOD workflow
- Prepare features: address missing values, encode categories, remove identifiers, and prevent target leakage.
- Choose a detector: match the method to geometry, scale, dimensionality, and data context.
- Set a threshold policy: configure
contaminationonly when its assumption is defensible, or establish a threshold through validation and review. - Fit on training data: keep held-out or future data separate.
- Inspect scores and labels: scores rank abnormality; labels are thresholded decisions.
- Investigate and validate: decide whether to correct, segment, retain, escalate, or monitor each case.
Minimal Isolation Forest example
import numpy as np
from pyod.models.iforest import IForest
X_train = np.array([
[10.0, 1.0],
[11.0, 1.2],
[10.5, 0.9],
[12.0, 1.1],
[11.2, 1.0],
[50.0, 8.0],
])
detector = IForest(contamination=0.10, random_state=42)
detector.fit(X_train)
labels = detector.labels_
scores = detector.decision_scores_
X_new = np.array([[10.8, 1.1], [48.0, 7.5]])
new_scores = detector.decision_function(X_new)
new_labels = detector.predict(X_new)
print(labels)
print(scores)
print(new_labels)
print(new_scores)
decision_scores_ contains scores for observations used in fitting. decision_function(X_new) scores new observations, and predict(X_new) returns thresholded labels. Exact score direction and semantics are detector-specific, so consult the selected class documentation rather than assuming every algorithm exposes identical raw values.
Preserve an investigation table
import pandas as pd
results = pd.DataFrame({
"row_id": row_ids,
"anomaly_score": scores,
"is_outlier": labels == 1,
})
results = results.sort_values("anomaly_score", ascending=False)
Confirm the ordering for your detector and installed PyOD version. Scores from unrelated algorithms are not automatically comparable.
Prepare data before fitting
- Missing values: impute or otherwise handle them before the detector sees the data.
- Categorical variables: encode them appropriately; PyOD detectors generally expect numerical feature matrices rather than raw strings.
- Scaling: standardize features for distance-, covariance-, PCA-, and SVM-based methods when units differ. Tree-based Isolation Forest is generally less dependent on scale.
- Skew: consider log or other domain-appropriate transforms for heavily skewed positive variables.
- Identifiers: remove row IDs, account numbers, and other columns that merely memorize identity.
- Leakage: do not include future information or a target-derived feature.
- Traceability: retain the original row identifier outside the feature matrix.
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from pyod.models.knn import KNN
model = make_pipeline(
StandardScaler(),
KNN(contamination=0.05)
)
model.fit(X_train)
predictions = model.predict(X_test)
Fit preprocessing on training data and apply the learned transformation to held-out or future data. A production-style evaluation should not calculate scaling statistics from the complete dataset.
Choosing a PyOD detector
| Requirement | Starting point | Main caveat |
|---|---|---|
| General tabular baseline | Isolation Forest | Validate features and threshold; representation still matters. |
| Local-density anomalies | LOF or kNN | Sensitive to scaling, neighborhood size, distance metric, and differing cluster densities. |
| Fast, relatively interpretable baseline | ECOD, COPOD, or HBOS | Distribution and feature-dependence assumptions can fail. |
| Low-dimensional linear structure | PCA | Weak for strongly nonlinear relationships or poorly scaled features. |
| Gaussian-like data | Elliptic Envelope or MCD | Sensitive to non-Gaussian data and high dimensionality. |
| Many candidate models | SUOD or an ensemble | More complexity and harder explanations. |
| Known representative labels | Supervised model, XGBOD, or DevNet | Requires label quality and strict leakage controls. |
| Time series | Time-series detectors or windowed features | Pointwise tabular methods can ignore temporal context. |
| Graphs | Graph-specific PyOD detectors | Requires graph structures and may be transductive. |
| Text or images | Embeddings followed by detection | Embedding quality may dominate detector quality. |
Isolation Forest
IForest is a strong first baseline for many tabular datasets. It handles nonlinear structure and usually scales better than neighborhood methods without requiring a covariance model. It can still be affected by feature representation and an unsuitable contamination assumption.
from pyod.models.iforest import IForest
clf = IForest(contamination=0.05, random_state=42)
Local Outlier Factor
LOF is useful when an observation is sparse relative to nearby points. Scaling, n_neighbors, and unequal cluster densities matter. Ordinary LOF is intended for outlier detection on fitted data; scoring unseen data requires a novelty-detection configuration and careful interpretation. The scikit-learn documentation explains the distinction between training-set fit_predict behavior and novelty scoring.
from pyod.models.lof import LOF
lof = LOF(contamination=0.05, n_neighbors=20)
ECOD, COPOD, and HBOS
These provide fast, non-neural baselines with relatively transparent distributional logic. HBOS is most appropriate when treating features independently is reasonable; important interaction-based anomalies can be missed.
from pyod.models.ecod import ECOD
from pyod.models.copod import COPOD
from pyod.models.hbos import HBOS
kNN and PCA
kNN works when distance to neighbors has a meaningful interpretation, but it needs scaling, a distance metric, and a sensible neighborhood size. PCA is useful when normal observations lie near a lower-dimensional linear structure and anomalies produce large projection or reconstruction errors. Neither is a universal solution.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #3
- Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
- Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
- Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
- Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
- What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
from pyod.models.knn import KNN
from pyod.models.pca import PCA
Deep and ensemble methods
Autoencoders, variational models, DeepSVDD, SUOD, and other ensembles can help with large or complex datasets after simpler baselines have been tested. They add dependencies, tuning, training instability, computational cost, and explanation challenges.
Understanding contamination, scores, and labels
contamination=0.02 configures the workflow around approximately 2% anomalous observations or uses that proportion to establish a threshold. It is not proof that exactly 2% of the rows are truly anomalous. If the rate is unknown, compare settings and validate them through labeled examples, expert review, stability, downstream cost, or operational capacity.
- Score: a continuous abnormality measure produced by the detector.
- Label: a thresholded inlier/outlier decision.
- Explanation: evidence about which features, neighbors, residuals, or conditions made a row suspicious.
- Action: review, correction, segmentation, retention, escalation, or monitoring.
A high score means “more suspicious according to this detector,” not “fraud,” “bad data,” or a causal explanation. Do not compare raw score magnitudes across different algorithms as if they were calibrated probabilities.
Evaluate detections
When trusted labels exist
- Precision and recall
- Precision at the review budget your team can handle
- PR-AUC for rare-event problems
- ROC-AUC when its assumptions fit the decision
- Cost-weighted false-positive and false-negative analysis
- Performance by customer, device, geography, or other important segment
- Threshold and calibration analysis
When labels do not exist
- Have domain experts review the highest-ranked cases.
- Check stability across random seeds, resamples, and reasonable preprocessing choices.
- Compare agreement among different detector families.
- Use temporal holdouts and monitor score distributions for drift.
- Track investigation outcomes and the false-positive burden on operators.
Do not report accuracy on an unlabeled dataset: there is no verified target against which to calculate it.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallFailure modes to avoid
Deleting every flagged row
Rare observations may be the fraud, failure, or new population you need to find. Flag first, preserve evidence, and document any correction or removal.
Assuming contamination is ground truth
The parameter controls thresholding; it does not reveal the actual prevalence of anomalies.
Ignoring scale and mixed units
Distance and covariance methods can be dominated by a feature measured in larger numerical units. Scale where appropriate and compare sensitivity.
Fitting on all data
Using future or held-out rows to fit preprocessing or the detector can make offline results optimistic and hide train/test mismatch.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsRank #4
- Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
- Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
- Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
- Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
- Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
Using one global model for multimodal data
Several legitimate clusters can make a global detector label an entire subgroup. Segment by product, geography, device, customer type, or operating regime, or use local methods.
Ignoring high dimensionality
As dimensions grow, distance becomes less informative. Remove irrelevant features, use domain-driven selection or dimensionality reduction, compare detector families, and inspect stability.
Confusing historical normality with future normality
A genuine process change can make ordinary new observations look anomalous. Use time-based validation, drift checks, a retraining policy, and monitoring of score distributions.
Using LOF training and novelty APIs interchangeably
Outlier detection on fitted rows and novelty detection on unseen rows are different operations. Configure and evaluate the method for the way it will be used.
PyOD, scikit-learn, or a managed platform?
| Option | Best fit | Trade-off |
|---|---|---|
| PyOD | Local Python development, research, batch scoring, and custom pipelines needing many detector families. | Teams must own deployment, monitoring, review workflows, and detector-specific dependencies. |
| scikit-learn | A smaller set of established estimators with strong preprocessing, pipeline, and model-selection integration. | Fewer dedicated outlier algorithms than PyOD. |
| Managed observability such as Datadog | Continuous infrastructure and application monitoring, dashboards, alerting, logs, traces, metrics, and on-call ownership. | It is not a drop-in replacement for a local tabular detector and can add platform cost and operational complexity. |
PyOD is open source and has no normal subscription. Datadog’s pricing page, checked August 18, 2026, lists products such as APM starting at $31 per host/month with annual billing or $36 on demand, and Universal Service Monitoring at $9 per infrastructure host/month annually or $13 on demand. Those are observability products, not prices for a PyOD-equivalent library; see Datadog’s pricing page.
What happens after detection?
Choose treatment based on evidence and business context:
- Correct a confirmed measurement or processing error.
- Cap or transform a value only when the modeling rationale is documented.
- Remove a row only when trusted evidence establishes that it is invalid.
- Keep the observation and use a robust downstream model.
- Segment a new population or operating regime.
- Escalate a possible fraud, safety, or reliability event.
The durable workflow is a ranking-and-review system with monitoring, not a one-click data-cleaning command.
Frequently Asked Questions
Is PyOD supervised or unsupervised?
Most common PyOD workflows are unsupervised, but the project also includes supervised or label-assisted detectors such as XGBOD and DevNet.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Best Value
- 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
- Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
- Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
- HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
- What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
Is PyOD free to use?
The PyOD package is open source under the BSD-2-Clause license. Optional integrations may have their own dependencies or service terms.
What is the best PyOD algorithm?
There is no universal winner. Isolation Forest is a practical tabular baseline; LOF or kNN suit local-density problems, while ECOD, COPOD, and HBOS provide fast distribution-based baselines.
Does PyOD replace pandas or scikit-learn?
No. pandas and NumPy prepare data, scikit-learn supplies preprocessing and pipelines, and PyOD supplies a broader outlier-detection catalog.
Can PyOD detect time-series anomalies?
Yes, current PyOD documentation includes time-series capabilities, but pointwise tabular methods can miss temporal context. Windowed features or time-series-specific detectors may be needed.
How should I choose contamination?
Treat it as a thresholding assumption, not truth. Compare plausible values and validate with labels, expert review, stability, and the cost of false alerts.
Should detected outliers be removed?
Not by default. A flagged observation may be a valid rare event, fraud, a regime change, or a new population; investigate before changing the data.
How do I score new observations?
Fit the detector on training data, then call its documented decision_function and predict methods on new rows. Detector-specific novelty behavior must be checked, especially for LOF.
Does PyOD accept categorical features directly?
Typical detectors expect numerical feature matrices, so encode categorical variables using a representation appropriate to the algorithm and domain.
Which Python versions does the current package support?
PyPI metadata checked August 18, 2026 requires Python 3.9 or newer.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

