Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

There is no single best method for incomplete data. A defensible choice links three things: the analysis you need to run, the reason values are missing, and whether a method’s assumptions fit that reason. The sections below follow that order, from defining the analysis to reporting the decision.

Start with the analysis, not the missing values

Missing data matter only in relation to a specific analysis. A gap that barely affects one question can distort another. Before examining any gaps, write down:

  • The outcome, and whether it is continuous, binary, time-to-event, or measured repeatedly.
  • The exposure or main predictors.
  • The covariates the model will adjust for.
  • The estimand: the specific quantity you want to estimate, such as a difference between groups in the target population, or an association conditional on the model’s covariates.
  • The data structure: a single cross-section, a panel, or repeated measures within participants over time.

Where the gap sits changes the problem. A missing outcome, a missing predictor, and a missing repeated measurement each affect the analysis differently. A single study often has more than one of these, and each needs its own assessment.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Describe where and why values are missing

Describe the missingness before choosing any method. Check the following:

  • The count of missing values for each variable in the planned model, and how many records would remain under complete-case analysis.
  • Whether variables tend to be missing together, and whether missingness is concentrated in particular subgroups, sites or time points.
  • For longitudinal data, dropout by wave and by exposure or treatment group.
  • The reason recorded for each gap at collection or follow-up, such as a refused question, a site that did not collect a measure, or a participant who moved away.

Missing data can reduce power, introduce bias, increase uncertainty, and make the analysed sample less representative of the target population, according to the ENCEPP methodological guide, Chapter 6, section 6.3. Each of these consequences is a reason to examine the pattern rather than only counting gaps.

Make the mechanism assumptions explicit

Missingness mechanisms are usually classified as MCAR, MAR, or MNAR. These labels describe assumptions about the process that produced the gaps. They are not properties that can simply be read off the dataset.

MCAR: missing completely at random

Missingness is unrelated to the variables in the analysis, including the value that is missing. An example is a set of lab samples lost to a freezer failure that affected samples at random. MCAR is a strong assumption, so it has to be justified from how the data were collected rather than assumed by default.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

MAR: missing at random

Systematic differences between missing and observed values can be explained by observed data included in the analysis process. Suppose older participants are less likely to return a follow-up questionnaire, and age is recorded for everyone. Once age is accounted for, the chance of a missing response no longer depends on the response itself.

MNAR: missing not at random

Systematic differences remain after the observed data are taken into account. Missingness depends on the unobserved value itself or on other unobserved causes. For example, people with high incomes may be more likely to skip an income question, and the survey holds no measure that explains that tendency once age, education and region are known.

What the observed data can and cannot show

Observed data cannot decide between MAR and MNAR. The ENCEPP guide states: “It is however not feasible to assess MAR versus MNAR based on the observed data.” Observed predictors of missingness can still challenge MCAR, because they show that missingness is not unrelated to measured variables. They do not show that the remaining gaps are MAR. The judgement rests on study knowledge, including:

  • The collection protocol, follow-up procedures, and field or site records.
  • Reasons participants gave for non-response or dropout.
  • Subject-matter knowledge about why the outcome would be absent.
  • Comparisons with external information about the population, where available.

State the assumption in plain terms, for example: “We assume dropout depends on baseline severity but not on follow-up severity once baseline severity is adjusted for.” A written statement like this can be checked against the study team’s knowledge and reviewed by others.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Compare the methods against your assumptions

No method is defensible in every setting. The table lists the key condition behind each family and when it can make sense. The subsections that follow give the cautions for each.

Method Key assumption or condition When it can make sense
Complete-case analysis (CCA) Selection of complete records does not bias the target analysis Selection is unrelated to the outcome once the model’s variables are considered; can be appropriate in specific settings, including some cases involving MNAR covariates
Multiple imputation (MI) MAR, with an imputation model that includes the relevant observed data Missingness is explained by observed variables, which can be included in the imputation model along with auxiliary variables
Likelihood / maximum likelihood The model and missingness assumptions are stated and suit the estimand Longitudinal outcomes, and models that can use incomplete records under their assumptions
Weighting / inverse probability weighting Observation probabilities can be modelled from observed covariates, with adequate support The chance of a record being observed can be credibly modelled from observed variables
MNAR-oriented models and sensitivity analysis Additional assumptions or subject-matter knowledge about the missing values Missingness may depend on unobserved values, or the mechanism remains uncertain

Use these questions to narrow the field:

  • Is the gap in the outcome, a predictor, or repeated measures? Each location points to different options.
  • Can the missingness be explained by variables you have recorded? If yes, MAR-based methods such as MI become candidates.
  • Can the probability of observation be modelled from observed covariates, with enough records across their range? If yes, weighting becomes a candidate.
  • Is the outcome longitudinal? If yes, consider likelihood methods or the MI approaches described below.
  • Could missingness depend on the missing value itself? If yes, an MNAR model or sensitivity analysis belongs in the plan from the start.

When you compare the shortlisted methods, make these axes explicit:

  • Assumptions about the missingness mechanism.
  • Compatibility with the estimand and model.
  • Ability to use incomplete cases and auxiliary information.
  • Risk of bias.
  • Precision and uncertainty.
  • Sensitivity to alternative plausible mechanisms.

Complete-case analysis

CCA keeps only records with every variable the analysis needs. It is simple and does not model the missingness, but it discards incomplete records, which can reduce precision and power. A small amount of missingness does not make CCA valid, and missingness that is not MCAR does not make it invalid. The question is how complete-case selection relates to the outcome and the covariates. In a regression, one route to validity is that the chance of a record being complete depends only on variables already in the model, not on the outcome. The ENCEPP guide notes that CCA may be appropriate in specific settings, including some cases involving MNAR covariates.

Multiple imputation

MI replaces each missing value with several plausible values, producing multiple completed datasets. Each dataset is analysed and the results are combined, so the uncertainty introduced by imputation appears in the final estimates. Under MAR, the imputation model should include every variable in the analysis (outcome, exposure and covariates) plus auxiliary variables that help explain missingness or predict the missing values, even when those auxiliary variables are not in the analysis model. The ENCEPP guide and the 2019 International Journal of Epidemiology article both describe this framework.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

MI is not a universal fix. Its results depend on the MAR assumption and on how the imputation model is specified, and when MAR is wrong, MI based on it can be biased. The 2019 article, titled “Accounting for missing data in statistical analyses: multiple imputation is not always the answer,” develops this point in detail.

Likelihood and maximum likelihood

Likelihood methods estimate model parameters from all available observations under a stated model, without first filling in the missing values. They are particularly relevant to longitudinal outcomes. NIH guidance on missing outcomes recommends considering maximum likelihood or MI methods that can condition on prior outcomes and baseline variables. Before using either, state the model and missingness assumptions, and confirm that the approach suits your estimand and data structure.

Weighting

Inverse probability weighting gives each observed record a weight based on its estimated probability of being observed, so that the observed records better resemble the full target population. It is among the principled approaches described in the literature. It requires a credible model of observation probability built from observed covariates, and adequate support: enough observed records across the range of covariate values. If some records have very small estimated probabilities of being observed, their weights become extreme and the estimates unstable. Document which variables enter the weight model and why. Roderick Little’s 2024 review of missing data analysis in Annual Review of Clinical Psychology covers the wider field.

MNAR-oriented models and sensitivity analysis

When missingness may depend on the unobserved values, MAR-based methods are not enough. The ENCEPP guide points to pattern-mixture models and other specialised MNAR models. These add assumptions that observed data cannot check, so their parameters have to be set from subject-matter knowledge or varied across plausible values in a sensitivity analysis.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Watch out for three common shortcuts

  • Choosing by the proportion missing. The ENCEPP guide points to a published discussion holding that the proportion missing should not guide the choice of MI method. A small gap in the exposure can matter more than a large gap in a variable outside the model.
  • Mean substitution and last observation carried forward. These are not generally valid fixes. The ENCEPP guide says simple methods can produce misleading inferences when their assumptions fail.
  • Missing-indicator categories. Adding a “missing” category to a variable is not an automatic solution. The ENCEPP guide warns that this approach can be invalid, including under MCAR.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Check robustness with sensitivity analysis

When the mechanism is uncertain, the goal is to show whether the conclusion holds under plausible alternatives. NIH guidance states: “If there is considerable uncertainty about the missing-data mechanism, investigators should consider a sensitivity analysis (Baker, 2019), which may include a worst-case scenario.” A practical sequence:

  1. Run the primary analysis under the assumption you judge most plausible, and record the estimate and its interval.
  2. Add alternative analyses: for instance, MI with and without a key auxiliary variable, a complete-case comparison, and an MNAR scenario in which missing outcomes are assumed to differ from observed ones by a stated amount.
  3. Decide before looking at the results which change would matter, such as a reversal of direction or a result crossing a threshold of interest.
  4. Present the results side by side and label the assumption behind each one.

A worst-case scenario sets each missing value to the least favourable plausible outcome for its group. The NIH guidance mentions it as one option, particularly in a clinical-trial planning context.

Report the decision

A reader should be able to see how the method was chosen and what would change the result. Report:

  1. The missingness: counts by variable, patterns across variables and time, and the reasons recorded at collection or follow-up.
  2. The mechanism assumptions, and the study knowledge behind each.
  3. The model and estimand, with the variables in the analysis.
  4. The auxiliary information used, and the imputation or weighting model in enough detail to reproduce it.
  5. Uncertainty, including how imputation or weighting variability was handled.
  6. Sensitivity results alongside the primary result, and the limitations that remain.

Further reading

The ENCEPP guide names Little and Rubin’s Statistical Analysis with Missing Data as a useful reference for the underlying methods.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.