Data scientists use techniques across the full workflow: collecting and checking data, exploring patterns, preparing variables, building models, evaluating results, and delivering findings. There is no single official list of exactly 40 techniques; the ones below are an editorially selected guide to common methods and what each is for. If you are asking how data scientists analyze data or which machine-learning methods they use, start with the question being answered: describe, estimate, predict, group, detect, or communicate. The workflow is iterative—findings in one stage can send an analyst back to an earlier one.
How to choose among data science techniques
Techniques are not interchangeable, and categories overlap. Before choosing one, pin down the task and the conditions around it:
- Question: Are you describing a population, estimating an effect, predicting an outcome, grouping observations, finding unusual cases, or reducing dimensions?
- Data: Are outcomes labeled? Is the data missing, imbalanced, time-ordered, or sampled in a way that affects analysis?
- Assumptions and representation: Do variables need scaling or encoding? Does the method suit the data’s shape and measurement?
- Evaluation: Which errors matter most, and what validation strategy reflects how the model will be used?
- Practical constraints: Consider interpretability, compute, latency, monitoring, and reproducibility. Compare methods against a simple baseline where useful.
The techniques below are arranged by workflow stage, not as a ranking or a recipe that assigns one method to each problem.
Data acquisition, quality, and exploration
1. Data ingestion and joining
Bring data from source systems into an analysis environment, then combine related records using appropriate keys. Joining can make a richer analysis possible, but incorrect keys or many-to-many matches can duplicate rows and distort counts. Microsoft’s Fabric data-science tutorial illustrates source ingestion and an iterative workflow.
#1 Best Overall
2. Schema and type validation
Check that fields have expected names, meanings, and representations—for example, that a date is parsed as a date and a quantity is numeric. A value can be technically valid but semantically wrong, so validate units and definitions as well as data types. Google’s data-analysis guidance emphasizes understanding what the measurements represent.
3. Missing-value handling
Decide whether to retain nulls, remove affected rows or columns, or impute values. The right choice depends on why data is missing and how the method handles it; an imputed value is not an observed fact. Document the decision because different treatments can change both the sample and the result. See scikit-learn’s preprocessing guidance.
4. Duplicate detection and removal
Find repeated records and determine whether they are accidental copies or legitimate repeated events before removing them. Deduplicating on too few fields can erase real observations; failing to deduplicate true copies can overcount them. Microsoft’s tutorial demonstrates duplicate removal in its example workflow.
5. Unit and spelling normalization
Standardize inconsistent spellings, labels, and units so equivalent values can be compared—for instance, converting measurements to a common unit. Keep a record of corrections: normalization can fix representation, but it should not silently redefine what a measurement means. Google’s data-quality guidance discusses documenting data corrections.
Recommended Free Tools
6. Summary statistics
Use measures such as the mean, median, and standard deviation to summarize a variable. The mean can be sensitive to extreme values, while a median alone says little about spread; neither reveals the full distribution. Treat summaries as an initial view, not a substitute for inspecting the data.
7. Histograms and empirical distributions
Plot values or their empirical distribution to see shape, spread, multiple peaks, and possible outliers. These views can reveal structure hidden by a single average. Bin width and axis choices affect what a histogram appears to show, so inspect more than one view when the shape matters.
8. Quantile-quantile plots
A quantile-quantile (Q–Q) plot compares quantiles from a sample with those from a reference distribution, helping assess differences in distributional shape. It is a diagnostic rather than a proof that data follow a particular distribution; patterns can also reflect sampling variation or outliers.
9. Time slicing and trend checks
Inspect measurements across time to identify changing collection practices, system breaks, seasonality, or unusual periods. An unusual day might be an error or a real event; deleting it without investigating can remove meaningful information. Google’s analysis guidance recommends checking time-based behavior and investigating unexpected periods.
Rank #2
10. Filtering and cohort definition
Define which records belong in the analysis and apply filters transparently. State the criteria and count how many records remain after each filtering stage; otherwise, it is hard to see how the analyzed population differs from the original data. A cohort is meaningful only in relation to the question and the data collection process.
11. Ratio definition
Specify both the numerator and denominator for a rate or ratio. “Conversion rate,” for example, can differ depending on whether the denominator is visitors, sessions, or eligible users. Changing the denominator changes the population the figure describes, even when the label stays the same.
12. Repeated measurement
Measure a phenomenon in multiple ways or compare independent sources when possible, then investigate discrepancies. Agreement can increase confidence that a pattern is not an artifact of one measurement method, but it does not guarantee that the measures are unbiased or independent.
Statistical analysis and feature preparation
13. Correlation and covariance analysis
Correlation summarizes the direction and strength of a relationship on a standardized scale; covariance describes how two variables vary together in their original units. Both are measures of association, not evidence that one variable causes another. Nonlinear relationships and outliers can also make a single summary misleading.
Free tools Windows power users keep installed
One-click scans. No signup required.
14. Regression analysis
Regression models a numeric outcome in relation to one or more predictors. Linear regression describes a conditional mean under its modeling assumptions; quantile regression can model a selected conditional quantile instead. A fitted relationship does not, on its own, establish a causal effect.
15. Logistic regression
Logistic regression models class probabilities, commonly for binary outcomes, from predictor values. It offers a statistical model for classification and can provide probability estimates, but those estimates depend on model fit and data representativeness. A decision threshold is a separate choice from fitting the model.
16. Hypothesis testing and uncertainty estimation
Use confidence intervals or significance procedures to quantify uncertainty around a clearly defined estimate, given a sampling and measurement process. A visible difference in a plot is not by itself evidence of a reliable difference. Statistical significance also does not automatically imply practical importance or causation.
17. Outlier handling
Investigate unusually high, low, or distant observations. Correct a confirmed recording error, but retain legitimate extremes when they are part of the phenomenon being studied. Automatically deleting outliers can bias analysis, while leaving errors unexamined can distort it.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchRank #3
- Wiley
- Language: english
- Book - storytelling with data: a data visualization guide for business professionals
18. Categorical encoding
Convert categories into a representation a model can use. One-hot encoding creates indicator variables for categories and avoids imposing a numeric order where none exists. It can create many columns for high-cardinality fields, so consider the data size and model’s needs. AWS describes one-hot encoding and other feature-engineering operations in its Machine Learning Lens.
19. Binning and discretization
Turn a continuous variable into intervals or categories, such as age bands. This can make patterns easier to describe or suit a particular model, but cut points discard within-bin differences and may obscure useful information. Choose boundaries for a stated reason rather than treating bins as naturally occurring facts.
20. Feature construction
Calculate new predictors from existing data using domain knowledge—for example, derive elapsed time from two timestamps. A constructed feature is useful only if its inputs are valid and available at the time a prediction would be made; otherwise it can introduce leakage or encode an unintended proxy.
21. Feature imputation and transformation
Prepare predictor values by filling missing entries or applying transformations such as scaling or a logarithm. The operation should suit the variable, missingness, and model. Fit learned transformations using training data only, then apply them to validation, test, or live data to avoid information leakage. See the scikit-learn user guide.
22. Feature selection
Choose a subset of predictors using univariate tests, sequential procedures, or model-based criteria. Selection can simplify a model and reduce noise, but selecting features using the full dataset before evaluation leaks information into the test result. Perform selection within the training or cross-validation process.
23. Dimensionality reduction
Represent many variables with fewer derived dimensions. Principal component analysis (PCA) is a common linear approach; other methods can be suited to different data structures. A reduced representation may help analysis or modeling, but its components are not automatically easy to interpret, and some information is lost.
Modeling and pattern discovery
24. Linear and regularized regression
Ordinary least squares fits a linear relationship by minimizing squared residuals. Ridge, lasso, and elastic net add penalties that constrain coefficients; they can help manage complexity, with lasso also capable of shrinking some coefficients to zero. Regularization does not repair poor measurement, omitted structure, or an unsuitable target.
25. Decision trees
A decision tree repeatedly splits data into groups using feature-based rules, for classification or regression. Its rule structure can be inspected, but deep trees can fit noise and perform poorly on new data. Depth and other settings should be evaluated rather than assumed.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
26. Random forests
A random forest combines predictions from multiple randomized decision trees. It often captures nonlinear patterns without requiring the same scaling as some other models, but it is less directly interpretable than a small tree and still needs validation on representative data.
27. Gradient boosting
Gradient boosting builds an ensemble in sequence, with later learners addressing errors made by earlier ones according to the training objective. It can be effective for structured prediction tasks, but tuning and overfitting are concerns; compare it with simpler baselines using an appropriate validation procedure.
28. Support vector machines
Support vector machines (SVMs) are supervised methods for classification and regression. Their behavior depends on the chosen kernel and settings; feature scaling is often important, and large datasets can make some implementations computationally demanding. Choose the variant and evaluation strategy for the task.
29. Neural networks
Neural networks learn layered transformations of inputs and can model complex relationships. They may need substantial data, computation, and tuning, depending on the task and architecture; they are not automatically better than simpler methods. Establish a baseline and evaluate on held-out data.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitches30. Naive Bayes
Naive Bayes classifiers estimate class probabilities using Bayes’ rule with a simplifying conditional-independence assumption among features. They can be useful for classification, including some text tasks, but the assumption may be unrealistic and probability estimates may need separate calibration.
31. Nearest-neighbor methods
Nearest-neighbor methods classify, regress, or retrieve examples by proximity in a chosen feature representation. The result depends on the distance measure, feature scaling, and what counts as a meaningful neighbor. High-dimensional data can make distances less informative.
32. Clustering
Clustering groups observations without known target labels. K-means, hierarchical clustering, DBSCAN, and HDBSCAN represent different approaches and assumptions about cluster shape, density, and number. A cluster is a pattern under a selected representation and method, not necessarily a naturally distinct real-world group.
33. Association rules
Association-rule methods identify items or events that co-occur, often in transaction-like data. A rule can describe a recurring pattern, but co-occurrence does not show that one item causes another or that the pattern will generalize beyond the data used to find it.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →34. Anomaly or novelty detection
Anomaly detection identifies observations unusual relative to a dataset or modeled baseline; novelty detection asks whether new observations depart from a baseline learned earlier. A flagged case is a candidate for investigation, not necessarily an error or threat. The relevant baseline can also change over time.
35. Matrix factorization
Matrix factorization decomposes a data matrix into lower-dimensional factors. PCA, non-negative matrix factorization (NMF), and latent semantic analysis are examples with different assumptions and uses. The factors can reveal structure, but their meaning depends on the inputs and method.
36. Text feature extraction
Convert text into numerical features for analysis or modeling, using representations such as token counts or other vectorizations. Choices about tokenization, vocabulary, and preprocessing affect what information remains; a representation can lose context, order, or nuance. Scikit-learn documents text feature extraction in its user guide.
37. Time-related feature engineering
Derive predictors such as calendar fields or lagged measurements for a time-dependent task. Make sure each predictor would have been available at prediction time and split data in time order when that reflects deployment. Otherwise future information can leak into training and make evaluation overoptimistic.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →38. Ensemble learning
Combine predictions through methods such as bagging, voting, or stacking. These approaches can blend model strengths, but add complexity and do not guarantee better results. In stacking, train the combination step without using predictions that leak information from the same observations.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Evaluation, interpretation, and delivery
39. Train, validation, and test separation
Keep held-out data separate from model fitting and tuning so evaluation estimates performance on unseen cases. Use a split compatible with the data: time-ordered splits for future prediction, group-aware splits when records from one person or entity could cross partitions, and sampling strategies suited to class balance where appropriate. Leakage can make a test score look better than actual performance. Scikit-learn’s model-selection guidance covers splitting and validation.
40. Cross-validation, metrics, thresholds, tuning, calibration, and delivery
These are distinct techniques commonly used after and around model fitting; they are grouped here because they answer connected questions about whether a model is useful and how its outputs reach users.
- Cross-validation: evaluate across multiple folds to compare models or estimate performance more robustly than a single split. Keep all preprocessing and feature selection inside each fold.
- Metrics: use classification measures that reflect class balance and error costs rather than relying on accuracy alone; use regression error measures suited to the scale and consequences of numeric errors.
- Threshold tuning: choose the probability cutoff for a classification decision based on the intended trade-off between types of errors. The cutoff is not inherent to the probability model.
- Hyperparameter tuning: compare configurations using validation or cross-validation procedures, not by repeatedly optimizing against the final test set.
- Calibration: check whether predicted probabilities match observed outcome frequencies; a model can rank cases well while its probabilities are poorly calibrated.
- Feature inspection: permutation importance and partial-dependence tools can help examine model behavior, but correlated features complicate interpretations and these tools do not establish causation.
- Visualization: plots help inspect distributions and model behavior and communicate results. Microsoft’s Fabric tutorial uses matplotlib, seaborn, and plotly in its workflow.
- Experiment tracking and model registration: record runs, settings, and artifacts, then manage selected models for later use. Microsoft’s Fabric tutorial demonstrates MLflow integration.
- Batch scoring and reporting: generate predictions for a batch of records and make them available to downstream reporting or visualization. Monitoring remains important because data and relationships can change after deployment.
Microsoft’s end-to-end Fabric tutorial presents these activities as part of an iterative lifecycle, not a one-way sequence.
What makes a result trustworthy?
Data quality, definitions, sampling, and analytic choices shape what an output means. A polished chart or high-performing model cannot compensate for erroneous measurements or a population that does not match the intended use. Google’s data-quality guide states: “No matter how beautiful or striking or persuasive the end products are, if the underlying data was erroneous, badly collected, or low-quality, the resulting model, prediction, visualization, or conclusion will likewise be of low quality.”
In practice, document data corrections, filter criteria, measurement definitions, split strategy, and evaluation choices. Predictive performance answers how well a model predicts under an evaluation setup; it does not establish that changing an input will cause an outcome to change.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

