Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no universally correct way to replace missing values. First find out what each absence means and how missingness is patterned; then choose a method whose assumptions fit your analysis and report how sensitive your conclusions are to those assumptions. A blank may mean “not asked,” “not applicable,” “not recorded,” or “refused”—and those states should not automatically be treated as the same thing.

What does a missing value mean?

Before calculating how to handle missing data, check the variable definition, missing-value codes, and data-collection process. A blank is not necessarily an unknown measurement that should be estimated. It may represent a skipped question, a value that was withheld, a recording or import error, or a question that did not apply to that person.

Structural missingness is especially important: if a follow-up question was not asked because an earlier answer made it irrelevant, filling the blank with an estimate can change the meaning of the variable. Preserve distinct states where they carry different information, and correct data-entry or coding errors when the intended value can be established rather than treating them as ordinary missing observations.

How should you describe the missingness?

For the variables needed in your analysis, count missing values and calculate their proportions. Examine which variables tend to be missing together, how missingness varies across records, and whether observed characteristics differ between complete and incomplete records. Use knowledge of how the data were collected to investigate plausible causes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A model or comparison showing that observed variables predict missingness can help describe the pattern, but it cannot establish that missingness is ignorable or rule out dependence on unseen values. The UCLA Office of Advanced Research Computing emphasizes that MCAR is a strong assumption and that the right approach depends on the missingness mechanism and analytic goal (UCLA, “Multiple Imputation in Stata”).

MCAR, MAR, and MNAR are assumptions

  • MCAR (missing completely at random): whether a value is missing is unrelated to both observed and unobserved data.
  • MAR (missing at random): after conditioning on the observed information included in the analysis, missingness does not further depend on the unseen value.
  • MNAR (missing not at random): even after conditioning on observed information, missingness still depends on the unseen value or another unobserved factor.

These describe assumptions about the process that produced the missingness; they are not labels that can generally be proven by a convenient test of observed data. In particular, observed data alone generally cannot distinguish MAR from MNAR. Heymans and Twisk discuss this limitation in their clinical-research guidance, which is useful methodological context but not a prescription for every field (“Handling missing data in clinical research,” 2022).

Which method fits your analysis?

Choose based on the quantity you want to estimate (the estimand), plausible missingness assumptions, information loss, and whether the method represents uncertainty from missing values. The methods below are not interchangeable defaults.

Method What it does Key trade-off or assumption
Complete-case analysis (listwise deletion) Uses only records complete for all variables required by the analysis. Simple, but discards incomplete records and can reduce precision. It may avoid bias under MCAR for some parameter estimates; outside suitable conditions it can be biased.
Available-case (pairwise) analysis Uses all available observations separately for each calculation. Can retain more data for some descriptive calculations, but different results may use different subsets, complicating comparisons and some multivariate analyses.
Single-value imputation Fills each missing entry with one value, such as a mean, median, mode, or model prediction. Convenient, but acts as if the imputed value were known; it can distort relationships and standard errors by concealing imputation uncertainty.
Multiple imputation Creates multiple plausible completed datasets, analyzes each, and combines estimates so imputation uncertainty is carried forward. Can be appropriate under MAR assumptions, but depends on a suitable, well-specified model compatible with the analysis.
Likelihood-based analysis Fits a model using the observed portions of the data directly. May suit some data structures and analytic models better than imputation, but relies on its own modeling assumptions.
MNAR-sensitive analysis Explicitly examines scenarios in which missingness depends on unseen values, using approaches such as selection, pattern-mixture, or tipping-point analyses. Requires specifying plausible departures from MAR; consequential decisions may warrant specialist statistical input.
Missing-value handling in machine learning Some algorithms can route or otherwise handle missing inputs within the model. Behavior varies by implementation. Verify it for the exact algorithm and prevent leakage in evaluation; prediction handling does not by itself answer inferential questions.

When is deleting incomplete rows defensible?

Complete-case analysis can be reasonable when its validity conditions are plausible for the target analysis and the amount of information lost is acceptable. Under MCAR, it may avoid bias in parameter estimates, but the smaller sample can still increase standard errors. It is not automatically safe just because the missing fraction appears small. The VA Health Economics Resource Center outlines the sample-size and power costs of listwise deletion (VA HERC, “Dealing with Missing Data”).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can you fill missing values with the mean?

Mean imputation is easy to implement but usually inadequate when the goal is valid inference: it inserts a fixed value without reflecting the uncertainty about what was missing, and can distort relationships and uncertainty estimates. The same central issue applies to a single median, mode, or prediction. Multiple imputation can represent that uncertainty, but is not automatically correct; its model and assumptions still need to fit the data and analysis.

A peer-reviewed review specifically cautions that multiple imputation is not always the answer (“Accounting for missing data in statistical analyses: multiple imputation is not always the answer,” International Journal of Epidemiology, 2019). As the UCLA guidance puts it, “The goal is not to advocate for one universal method.”

Rank #3
Sale
Storytelling with Data: A Data Visualization Guide for Business Professionals
  • Wiley
  • Language: english
  • Book - storytelling with data: a data visualization guide for business professionals

When should you consider multiple imputation?

Consider it when a suitable model can generate plausible values under defensible assumptions and you can include relevant information that predicts missingness or incomplete values. Ensure the imputation setup is compatible with the planned analysis. The number and type of variables, transformations, and data structure matter; a poorly specified model can still mislead.

When might likelihood or an MNAR analysis fit better?

Likelihood-based approaches can use observed data directly and may be a better fit for some models; imputation is not mandatory. If values may be missing because of the unseen values themselves—for example, if higher values are less likely to be recorded—an MAR-based method does not resolve that concern on its own. State plausible MNAR scenarios and examine whether the result changes under them. Clinical guidance recommends addressing the mechanism, method, and sensitivity to MNAR scenarios explicitly (Heymans and Twisk, 2022).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

A practical sequence for treating missing values

  1. Define the absence. Review codes, documentation, and collection logic. Separate not-applicable or not-asked states from genuinely unrecorded measurements, refusals, and errors.
  2. Map the pattern. Report missing counts and percentages for relevant variables; inspect co-occurrence and compare observed characteristics of complete and incomplete records.
  3. Specify the analysis target. Decide which estimate, prediction, or descriptive result matters and which variables it requires. The method should serve that target rather than a generic preference for keeping or deleting rows.
  4. Choose a method and state its assumptions. Consider deletion, available-case calculations, imputation, likelihood, or an MNAR-sensitive analysis in light of the estimand, data structure, and plausible missingness process.
  5. Check robustness. Where assumptions are uncertain and conclusions matter, compare plausible alternatives, including departures from MAR. Note whether the substantive conclusion changes.

For multiple imputation, include useful auxiliary information that predicts missingness or incomplete values, and make the imputation model compatible with the analysis. For machine-learning prediction, check the exact implementation’s missing-value behavior and ensure that preprocessing or imputation does not use information from an evaluation set that should remain held out.

What should you report?

Readers need enough detail to understand the scale of missingness, the assumptions behind the method, and how the result might change under alternatives. Report:

  • Missing counts and proportions for important variables, along with patterns or plausible collection causes.
  • The complete-case sample size when deletion is used, and which variables defined completeness.
  • The method, its assumptions, and the rationale for using it with the analysis target.
  • For imputation, the software and version, variables and transformations in the imputation model, and the number of imputed datasets and iterations when applicable.
  • Robustness or sensitivity checks, including any MNAR scenarios examined, and whether conclusions changed.

Do not present a test of observed-data associations as proof that missingness is MCAR or MAR. For an advanced treatment of missing-data methodology, Wiley lists Roderick J. A. Little and Donald B. Rubin’s Statistical Analysis with Missing Data, Third Edition, first published in 2019 (Wiley book listing).

Quick Recap

SaleBestseller No. 3
Storytelling with Data: A Data Visualization Guide for Business Professionals
Storytelling with Data: A Data Visualization Guide for Business Professionals
Wiley; Language: english; Book - storytelling with data: a data visualization guide for business professionals
$14.87

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.