Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Data scientists use techniques across the full workflow: collecting and checking data, exploring patterns, preparing variables, building models, evaluating results, and delivering findings. There is no single official list of exactly 40 techniques; the ones below are an editorially selected guide to common methods and what each is for. If you are asking how data scientists analyze data or which machine-learning methods they use, start with the question being answered: describe, estimate, predict, group, detect, or communicate. The workflow is iterative—findings in one stage can send an analyst back to an earlier one.

How to choose among data science techniques

Techniques are not interchangeable, and categories overlap. Before choosing one, pin down the task and the conditions around it:

  • Question: Are you describing a population, estimating an effect, predicting an outcome, grouping observations, finding unusual cases, or reducing dimensions?
  • Data: Are outcomes labeled? Is the data missing, imbalanced, time-ordered, or sampled in a way that affects analysis?
  • Assumptions and representation: Do variables need scaling or encoding? Does the method suit the data’s shape and measurement?
  • Evaluation: Which errors matter most, and what validation strategy reflects how the model will be used?
  • Practical constraints: Consider interpretability, compute, latency, monitoring, and reproducibility. Compare methods against a simple baseline where useful.

The techniques below are arranged by workflow stage, not as a ranking or a recipe that assigns one method to each problem.

Data acquisition, quality, and exploration

1. Data ingestion and joining

Bring data from source systems into an analysis environment, then combine related records using appropriate keys. Joining can make a richer analysis possible, but incorrect keys or many-to-many matches can duplicate rows and distort counts. Microsoft’s Fabric data-science tutorial illustrates source ingestion and an iterative workflow.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Schema and type validation

Check that fields have expected names, meanings, and representations—for example, that a date is parsed as a date and a quantity is numeric. A value can be technically valid but semantically wrong, so validate units and definitions as well as data types. Google’s data-analysis guidance emphasizes understanding what the measurements represent.

3. Missing-value handling

Decide whether to retain nulls, remove affected rows or columns, or impute values. The right choice depends on why data is missing and how the method handles it; an imputed value is not an observed fact. Document the decision because different treatments can change both the sample and the result. See scikit-learn’s preprocessing guidance.

4. Duplicate detection and removal

Find repeated records and determine whether they are accidental copies or legitimate repeated events before removing them. Deduplicating on too few fields can erase real observations; failing to deduplicate true copies can overcount them. Microsoft’s tutorial demonstrates duplicate removal in its example workflow.

5. Unit and spelling normalization

Standardize inconsistent spellings, labels, and units so equivalent values can be compared—for instance, converting measurements to a common unit. Keep a record of corrections: normalization can fix representation, but it should not silently redefine what a measurement means. Google’s data-quality guidance discusses documenting data corrections.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

6. Summary statistics

Use measures such as the mean, median, and standard deviation to summarize a variable. The mean can be sensitive to extreme values, while a median alone says little about spread; neither reveals the full distribution. Treat summaries as an initial view, not a substitute for inspecting the data.

7. Histograms and empirical distributions

Plot values or their empirical distribution to see shape, spread, multiple peaks, and possible outliers. These views can reveal structure hidden by a single average. Bin width and axis choices affect what a histogram appears to show, so inspect more than one view when the shape matters.

8. Quantile-quantile plots

A quantile-quantile (Q–Q) plot compares quantiles from a sample with those from a reference distribution, helping assess differences in distributional shape. It is a diagnostic rather than a proof that data follow a particular distribution; patterns can also reflect sampling variation or outliers.

9. Time slicing and trend checks

Inspect measurements across time to identify changing collection practices, system breaks, seasonality, or unusual periods. An unusual day might be an error or a real event; deleting it without investigating can remove meaningful information. Google’s analysis guidance recommends checking time-based behavior and investigating unexpected periods.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

10. Filtering and cohort definition

Define which records belong in the analysis and apply filters transparently. State the criteria and count how many records remain after each filtering stage; otherwise, it is hard to see how the analyzed population differs from the original data. A cohort is meaningful only in relation to the question and the data collection process.

11. Ratio definition

Specify both the numerator and denominator for a rate or ratio. “Conversion rate,” for example, can differ depending on whether the denominator is visitors, sessions, or eligible users. Changing the denominator changes the population the figure describes, even when the label stays the same.

12. Repeated measurement

Measure a phenomenon in multiple ways or compare independent sources when possible, then investigate discrepancies. Agreement can increase confidence that a pattern is not an artifact of one measurement method, but it does not guarantee that the measures are unbiased or independent.

Statistical analysis and feature preparation

13. Correlation and covariance analysis

Correlation summarizes the direction and strength of a relationship on a standardized scale; covariance describes how two variables vary together in their original units. Both are measures of association, not evidence that one variable causes another. Nonlinear relationships and outliers can also make a single summary misleading.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

14. Regression analysis

Regression models a numeric outcome in relation to one or more predictors. Linear regression describes a conditional mean under its modeling assumptions; quantile regression can model a selected conditional quantile instead. A fitted relationship does not, on its own, establish a causal effect.

15. Logistic regression

Logistic regression models class probabilities, commonly for binary outcomes, from predictor values. It offers a statistical model for classification and can provide probability estimates, but those estimates depend on model fit and data representativeness. A decision threshold is a separate choice from fitting the model.

16. Hypothesis testing and uncertainty estimation

Use confidence intervals or significance procedures to quantify uncertainty around a clearly defined estimate, given a sampling and measurement process. A visible difference in a plot is not by itself evidence of a reliable difference. Statistical significance also does not automatically imply practical importance or causation.

17. Outlier handling

Investigate unusually high, low, or distant observations. Correct a confirmed recording error, but retain legitimate extremes when they are part of the phenomenon being studied. Automatically deleting outliers can bias analysis, while leaving errors unexamined can distort it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Storytelling with Data: A Data Visualization Guide for Business Professionals
  • Wiley
  • Language: english
  • Book - storytelling with data: a data visualization guide for business professionals

18. Categorical encoding

Convert categories into a representation a model can use. One-hot encoding creates indicator variables for categories and avoids imposing a numeric order where none exists. It can create many columns for high-cardinality fields, so consider the data size and model’s needs. AWS describes one-hot encoding and other feature-engineering operations in its Machine Learning Lens.

19. Binning and discretization

Turn a continuous variable into intervals or categories, such as age bands. This can make patterns easier to describe or suit a particular model, but cut points discard within-bin differences and may obscure useful information. Choose boundaries for a stated reason rather than treating bins as naturally occurring facts.

20. Feature construction

Calculate new predictors from existing data using domain knowledge—for example, derive elapsed time from two timestamps. A constructed feature is useful only if its inputs are valid and available at the time a prediction would be made; otherwise it can introduce leakage or encode an unintended proxy.

21. Feature imputation and transformation

Prepare predictor values by filling missing entries or applying transformations such as scaling or a logarithm. The operation should suit the variable, missingness, and model. Fit learned transformations using training data only, then apply them to validation, test, or live data to avoid information leakage. See the scikit-learn user guide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

22. Feature selection

Choose a subset of predictors using univariate tests, sequential procedures, or model-based criteria. Selection can simplify a model and reduce noise, but selecting features using the full dataset before evaluation leaks information into the test result. Perform selection within the training or cross-validation process.

23. Dimensionality reduction

Represent many variables with fewer derived dimensions. Principal component analysis (PCA) is a common linear approach; other methods can be suited to different data structures. A reduced representation may help analysis or modeling, but its components are not automatically easy to interpret, and some information is lost.

Modeling and pattern discovery

24. Linear and regularized regression

Ordinary least squares fits a linear relationship by minimizing squared residuals. Ridge, lasso, and elastic net add penalties that constrain coefficients; they can help manage complexity, with lasso also capable of shrinking some coefficients to zero. Regularization does not repair poor measurement, omitted structure, or an unsuitable target.

25. Decision trees

A decision tree repeatedly splits data into groups using feature-based rules, for classification or regression. Its rule structure can be inspected, but deep trees can fit noise and perform poorly on new data. Depth and other settings should be evaluated rather than assumed.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

26. Random forests

A random forest combines predictions from multiple randomized decision trees. It often captures nonlinear patterns without requiring the same scaling as some other models, but it is less directly interpretable than a small tree and still needs validation on representative data.

27. Gradient boosting

Gradient boosting builds an ensemble in sequence, with later learners addressing errors made by earlier ones according to the training objective. It can be effective for structured prediction tasks, but tuning and overfitting are concerns; compare it with simpler baselines using an appropriate validation procedure.

28. Support vector machines

Support vector machines (SVMs) are supervised methods for classification and regression. Their behavior depends on the chosen kernel and settings; feature scaling is often important, and large datasets can make some implementations computationally demanding. Choose the variant and evaluation strategy for the task.

29. Neural networks

Neural networks learn layered transformations of inputs and can model complex relationships. They may need substantial data, computation, and tuning, depending on the task and architecture; they are not automatically better than simpler methods. Establish a baseline and evaluate on held-out data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

30. Naive Bayes

Naive Bayes classifiers estimate class probabilities using Bayes’ rule with a simplifying conditional-independence assumption among features. They can be useful for classification, including some text tasks, but the assumption may be unrealistic and probability estimates may need separate calibration.

31. Nearest-neighbor methods

Nearest-neighbor methods classify, regress, or retrieve examples by proximity in a chosen feature representation. The result depends on the distance measure, feature scaling, and what counts as a meaningful neighbor. High-dimensional data can make distances less informative.

32. Clustering

Clustering groups observations without known target labels. K-means, hierarchical clustering, DBSCAN, and HDBSCAN represent different approaches and assumptions about cluster shape, density, and number. A cluster is a pattern under a selected representation and method, not necessarily a naturally distinct real-world group.

33. Association rules

Association-rule methods identify items or events that co-occur, often in transaction-like data. A rule can describe a recurring pattern, but co-occurrence does not show that one item causes another or that the pattern will generalize beyond the data used to find it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

34. Anomaly or novelty detection

Anomaly detection identifies observations unusual relative to a dataset or modeled baseline; novelty detection asks whether new observations depart from a baseline learned earlier. A flagged case is a candidate for investigation, not necessarily an error or threat. The relevant baseline can also change over time.

35. Matrix factorization

Matrix factorization decomposes a data matrix into lower-dimensional factors. PCA, non-negative matrix factorization (NMF), and latent semantic analysis are examples with different assumptions and uses. The factors can reveal structure, but their meaning depends on the inputs and method.

36. Text feature extraction

Convert text into numerical features for analysis or modeling, using representations such as token counts or other vectorizations. Choices about tokenization, vocabulary, and preprocessing affect what information remains; a representation can lose context, order, or nuance. Scikit-learn documents text feature extraction in its user guide.

37. Time-related feature engineering

Derive predictors such as calendar fields or lagged measurements for a time-dependent task. Make sure each predictor would have been available at prediction time and split data in time order when that reflects deployment. Otherwise future information can leak into training and make evaluation overoptimistic.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

38. Ensemble learning

Combine predictions through methods such as bagging, voting, or stacking. These approaches can blend model strengths, but add complexity and do not guarantee better results. In stacking, train the combination step without using predictions that leak information from the same observations.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Evaluation, interpretation, and delivery

39. Train, validation, and test separation

Keep held-out data separate from model fitting and tuning so evaluation estimates performance on unseen cases. Use a split compatible with the data: time-ordered splits for future prediction, group-aware splits when records from one person or entity could cross partitions, and sampling strategies suited to class balance where appropriate. Leakage can make a test score look better than actual performance. Scikit-learn’s model-selection guidance covers splitting and validation.

40. Cross-validation, metrics, thresholds, tuning, calibration, and delivery

These are distinct techniques commonly used after and around model fitting; they are grouped here because they answer connected questions about whether a model is useful and how its outputs reach users.

  • Cross-validation: evaluate across multiple folds to compare models or estimate performance more robustly than a single split. Keep all preprocessing and feature selection inside each fold.
  • Metrics: use classification measures that reflect class balance and error costs rather than relying on accuracy alone; use regression error measures suited to the scale and consequences of numeric errors.
  • Threshold tuning: choose the probability cutoff for a classification decision based on the intended trade-off between types of errors. The cutoff is not inherent to the probability model.
  • Hyperparameter tuning: compare configurations using validation or cross-validation procedures, not by repeatedly optimizing against the final test set.
  • Calibration: check whether predicted probabilities match observed outcome frequencies; a model can rank cases well while its probabilities are poorly calibrated.
  • Feature inspection: permutation importance and partial-dependence tools can help examine model behavior, but correlated features complicate interpretations and these tools do not establish causation.
  • Visualization: plots help inspect distributions and model behavior and communicate results. Microsoft’s Fabric tutorial uses matplotlib, seaborn, and plotly in its workflow.
  • Experiment tracking and model registration: record runs, settings, and artifacts, then manage selected models for later use. Microsoft’s Fabric tutorial demonstrates MLflow integration.
  • Batch scoring and reporting: generate predictions for a batch of records and make them available to downstream reporting or visualization. Monitoring remains important because data and relationships can change after deployment.

Microsoft’s end-to-end Fabric tutorial presents these activities as part of an iterative lifecycle, not a one-way sequence.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What makes a result trustworthy?

Data quality, definitions, sampling, and analytic choices shape what an output means. A polished chart or high-performing model cannot compensate for erroneous measurements or a population that does not match the intended use. Google’s data-quality guide states: “No matter how beautiful or striking or persuasive the end products are, if the underlying data was erroneous, badly collected, or low-quality, the resulting model, prediction, visualization, or conclusion will likewise be of low quality.”

In practice, document data corrections, filter criteria, measurement definitions, split strategy, and evaluation choices. Predictive performance answers how well a model predicts under an evaluation setup; it does not establish that changing an input will cause an outcome to change.

Quick Recap

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.