Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

MICE (Multiple Imputation by Chained Equations) replaces missing cells with several plausible values, creates a completed dataset for each set of replacements, and combines the results of the same analysis across those datasets. It does not discover the one “true” value or make missingness irrelevant. Your conclusions still depend on the imputation models, predictors, missingness assumptions, and diagnostics.

What MICE does

The mice package implements Fully Conditional Specification (FCS), also called chained equations. Instead of one model for the entire dataset, it fits a conditional imputation model for each variable with missing values. The models can handle continuous, binary, unordered categorical, ordered categorical, and some two-level continuous data. Passive imputation can maintain deterministic relationships among variables.

MICE generates multiple completed datasets because one imputation would hide uncertainty about the missing values. Each dataset contains different plausible replacements. The variation among results is part of the uncertainty reported after pooling.

How to use mice in R for missing data

1. Describe the data and missingness

Start with the analysis question, the variables required by the scientific model, measurement levels, and the locations and patterns of missing cells. A pattern summary is descriptive; it cannot, by itself, identify why values are missing or prove that an imputation model is valid.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
library(mice)

# Inspect missingness patterns
md.pattern(dat)

# A compact overview of missing counts
sapply(dat, function(x) sum(is.na(x)))

Look for structural issues before imputing: impossible codes, inconsistent units, duplicate records, and variables that should be derived rather than independently imputed.

2. Choose an imputation model for each incomplete variable

The method vector selects a model for each target. With documented defaults, MICE uses predictive mean matching (pmm) for continuous variables, logistic regression (logreg) for binary variables, polytomous regression (polyreg) for unordered categorical variables, and proportional-odds logistic regression (polr) for ordered categorical variables. These are defaults, not guarantees that a method is suitable for your data.

Target type Documented default Questions to check
Continuous pmm Are the observed values and bounds plausible? Is predictive mean matching appropriate for skewness or outliers?
Binary logreg Are both outcome levels represented, and is separation a concern?
Unordered categorical polyreg Are categories sufficiently represented for a multinomial model?
Ordered categorical polr Is the ordering meaningful and is the proportional-odds assumption defensible?

Include variables that predict the incomplete value, the probability of being missing, or the outcome in the substantive analysis when justified. The imputation model should be compatible with the final analysis; omitting important predictors can distort associations and uncertainty.

3. Set predictors, methods, and other controls

The predictor matrix specifies which columns predict each target: a nonzero entry means that the column is used for that row’s imputation model. Do not automatically use every variable. Exclude identifiers, post-outcome variables that would not be available at the relevant time, or predictors that create leakage, while retaining variables needed to preserve important relationships.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
# Start from documented defaults, then edit deliberately
ini <- mice(dat, maxit = 0, printFlag = FALSE)
method <- ini$method
pred   <- ini$predictorMatrix

# Example: do not impute an ID column and do not use it as a predictor
method["record_id"] <- ""
pred[, "record_id"] <- 0
pred["record_id", ] <- 0

imp <- mice(dat, m = 20, maxit = 10,
            method = method,
            predictorMatrix = pred,
            seed = 20261002)

m is the number of completed datasets and maxit is the number of iterations through the chained models. The documented function defaults are m = 5 and maxit = 5; they are starting defaults, not evidence that five datasets or five iterations are sufficient for every analysis. Increase or otherwise justify them according to the amount of missing information, model complexity, and convergence behavior.

For more complex designs, MICE also exposes:

  • Blocks and formulas: group targets or specify model formulas when a simple predictor matrix is not enough.
  • Visit sequence: control the order in which target models are visited during an iteration.
  • where: choose the cells to impute, including selected observed cells for overimputation checks.

These controls are method-dependent. Some multivariate methods do not honor ignore; external imputation methods may require a complete predictor space and may not permit custom where matrices. Check the documentation for the method you select rather than assuming every control applies.

Inspect imputations and diagnostics

Check the iteration traces

plot(imp)

Trace plots help reveal chains that drift, mix poorly, or behave very differently from one another. A stable-looking trace does not prove that the model or missingness assumptions are correct.

Compare observed and imputed values

# Density and strip plots for a variable
 densityplot(imp, ~ age)
 stripplot(imp, age ~ .imp, pch = 20, cex = 1.2)

Compare ranges, centers, tails, category frequencies, and relationships with important predictors. Investigate imputed values outside physical or logical limits, implausible category combinations, and distributions that are inconsistent with observed data. Diagnostics are checks for problems, not proof that the replacements are true.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use overimputation when useful

An observed value can be temporarily treated as missing through a suitable where matrix. Comparing its imputed value with the known observed value provides a predictive check of the imputation model. This is an evaluation of model behavior on observed data, not a guarantee about unobserved cells.

Analyze every imputed dataset, then pool

Fit the intended scientific model separately to each completed dataset. The usual workflow is with() followed by pool():

# Example scientific model
fit <- with(imp, lm(outcome ~ exposure + age + sex))

# Combine coefficient estimates and uncertainty
pooled <- pool(fit)
summary(pooled, conf.int = TRUE)

pool() combines estimates from the repeated complete-data analyses using Rubin’s rules. Its output can include pooled estimates, standard errors, confidence intervals, p-values, relative increase in variance, degrees of freedom, proportion of total variance attributable to missingness, and fraction of missing information.

Do not average or stack the completed datasets and fit one model. Pooling datasets before fitting reverses the required sequence and can produce incorrect estimates, intervals, and p-values because it discards the between-imputation variation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When automatic pooling does not work

Pooling requires extractable estimates, standard errors, and residual degrees of freedom from each fitted model. Many common models are supported through broom methods. For mixed models, you may need the broom.mixed package. If your model has no supported tidier, extract the needed quantities explicitly or use an appropriate scalar-pooling approach; do not treat an unsupported object as pooled merely because with() ran.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Reporting checklist

  • Describe the dataset, analysis question, incomplete variables, and missingness patterns.
  • State the method used for each incomplete variable and why it matches the variable’s scale and structure.
  • Document predictors, formulas, blocks, visit sequence, exclusions, and any where or passive-imputation rules.
  • Report the number of imputations (m), iterations (maxit), and random seed when reproducibility matters.
  • Describe trace, distribution, plausibility, and any overimputation diagnostics, including problems found and changes made.
  • Name the substantive model and explain that it was fitted separately to every completed dataset and pooled with Rubin’s rules.
  • State limitations: imputed values are model-based, and conclusions depend on the missingness assumptions and model specification.

Common mistakes and how to avoid them

Mistake Why it is a problem Better practice
Using one deterministic replacement, such as the mean It understates uncertainty and can weaken relationships. Use multiple imputations with models appropriate to each variable.
Accepting defaults without review Defaults cannot encode your design, measurement levels, or substantive relationships. Inspect and justify methods and predictors.
Pooling completed rows before analysis Between-imputation uncertainty is lost. Fit the same model to each dataset, then call pool().
Treating a pattern table as a missingness diagnosis Patterns describe what is missing, not why it is missing. Use subject-matter knowledge and sensitivity checks.
Ignoring structural constraints Imputations can violate bounds, ordering, or multilevel relationships. Choose compatible methods, formulas, and passive rules; inspect results.

How do I impute missing values with mice?

In practice: load mice, inspect missingness, initialize with mice(data, maxit = 0), review methods and predictors, run mice() to create multiple imputations, inspect plots and distributions, then analyze with with(). The imputation is complete only when the resulting values are plausible for the analysis and the chosen model is defensible—not merely when the function returns an object.

How do I pool results after multiple imputation in R?

Store the repeated analyses returned by with(imp, ...) and pass that object to pool(). Summarize the pooled object with confidence intervals as needed. Pool model results, not the imputed datasets themselves. If extraction fails, install the relevant tidier or explicitly provide the estimates, standard errors, and degrees of freedom required for pooling.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.