Choose the method based first on how the cases and controls were sampled and what you want to estimate. Geocoded case and control locations treated as point patterns call for a different approach from binary outcomes measured within neighborhoods or other spatial clusters. And testing whether cases cluster is not the same task as estimating an exposure effect. A spatial model cannot compensate for controls who fail to represent the population that produced the cases.
Start by identifying the data structure and question
“Spatial dependence” can describe more than one feature of a case–control dataset. Before choosing a model, identify the observational unit, how the controls were selected, and the inferential target.
- Point-pattern data: Case and control locations are represented as spatial point patterns across a defined study region. The question may be whether the patterns differ, how relative risk varies across space, or whether an exposure is associated with case status.
- Clustered binary data: Individuals have binary outcomes and are grouped within villages, neighborhoods, or other spatial clusters. Outcomes within a cluster may be dependent, and the analysis must specify whether it targets population-average or subject-specific effects.
- Area-level data: Outcomes or counts are summarized by geographic area. This is a distinct data structure; do not assume that a point-pattern method or an individual-level clustered-binary model automatically fits it.
Also distinguish three goals: detecting clustering, estimating a spatial relative-risk surface, and estimating an adjusted exposure association. They are not interchangeable analyses.
For case–control point patterns, model or compare the spatial patterns
When cases and controls are observed locations across a study region, one approach is to compare their spatial intensity patterns. A risk surface can be represented by the ratio of the case and control intensity functions. This describes how the spatial patterns differ; it should not be presented as an absolute disease probability without an appropriate design and interpretation.
#1 Best Overall
Spatial point-process models
A multivariate log-Gaussian Cox process (LGCP) is one Bayesian option described for case–control point-pattern data. In the documented approach, covariates and residual spatial variation are represented with fixed and spatial random effects. The method was implemented using INLA through the R package inlabru.
The published implementation example uses the Chorley–Ribble dataset in Lancashire, England. It demonstrates one practical modeling route, not evidence that an LGCP is best for every case–control sampling design. The model, spatial domain, covariates, and treatment of the control process must match the data-generating and sampling setup.
Rank #2
- This guide is a perfect overview for the topics covered in introductory statistics courses.
For clustered binary outcomes, choose the estimand before the model
When binary outcomes are sampled within spatial clusters, a marginal model and a random-effects model can answer different questions. Do not choose between them solely because one includes a spatial term.
| Approach | Dependence representation | Interpretation supported in the cited work |
|---|---|---|
| Generalized estimating equations (GEE) | Marginal dependence can be represented using distance-related pairwise odds ratios; a 2018 paper uses hybrid pairwise likelihood. | Population-average effects. |
| Spatial random-effects model | Spatial random effects represent dependence or residual spatial variation. | Subject-specific inference. |
The 2018 pairwise-odds-ratio approach concerns spatially clustered binary prevalence data. It is not a universal recipe for matched case–control point patterns. Select a model whose assumptions and dependence structure fit the sampling design, and report the target interpretation.
Recommended Free Tools
Rank #3
If the goal is clustering detection, use a clustering method
A test for spatial clustering answers whether cases show a clustering pattern under the method’s comparison scheme. It does not by itself estimate an adjusted exposure effect or provide a general-purpose regression model.
Rogerson’s 2006 case–control methods include global and local tests. Examples include statistics based on cases closer to a given control than to other controls, cases within a specified distance, and a local statistic around a prespecified focus. The choice of statistic should follow the specific clustering question, including whether the analysis is global or focused on a location.
Rank #4
Protect validity when selecting controls and handling matching
Spatial adjustment is not a substitute for a sound comparison group. CDC case–control guidance recommends selecting controls that reflect the source population and its expected exposure, independently of the exposure under evaluation. Neighborhood matching may be useful in some designs, but excessive matching can be counterproductive; spatial proximity alone does not establish that controls are appropriate.
If cases and controls were matched, the analysis must account for that design. CDC field epidemiology guidance states that case–control data must account for matching in the analysis when matching was used. Conditional logistic regression is particularly appropriate for pair-matched data. A spatial random effect does not remove selection bias, confounding, or the need to respect matching.
Best Value
A practical decision path
- Describe the sampling unit. State whether the observations are geocoded individuals represented as point patterns, binary outcomes grouped within clusters, or area-level summaries.
- Define the study region and control process. Specify the geographic domain and how controls were obtained, including the source population they represent.
- Name the target. Decide whether the aim is cluster detection, a relative-risk surface, or an adjusted exposure association. For clustered binary outcomes, state whether the target is population-average or subject-specific.
- Choose a method that fits both design and target. For point patterns, consider case/control intensity comparisons or a point-process model. For clustered binary outcomes, consider marginal GEE or spatial random effects in light of the intended interpretation. For a clustering question, choose a global or local test suited to that question.
- Account for matching and covariates. Preserve the matched design in the analysis where applicable, and distinguish measured covariate adjustment from residual spatial variation.
- Check assumptions and uncertainty. State the dependence representation and estimation approach; assess whether the chosen model is appropriate for the sampling scheme and spatial domain. Report uncertainty rather than presenting a spatial pattern as certain.
What to report so readers can interpret the analysis
A reproducible account should give enough information for readers to understand what the spatial model represents and what its estimate means. Include:
- Case and control definitions, observational unit, and study region.
- Geographic coordinates or scale used, and how control locations were sampled.
- Whether matching was used, the matching variables, and how matching entered the analysis.
- The inferential target: clustering, relative-risk surface, or adjusted exposure association; for clustered outcomes, population-average or subject-specific.
- The dependence structure, covariates, model or test, and estimation method, including software where relevant.
- Uncertainty summaries and the assumptions needed to interpret the results.
These reporting details follow from the information needed to understand the design and models described above; they are not a quoted formal reporting standard.
Quick Recap
How to avoid common analytical errors
- Do not conflate data types. Individual point patterns and binary observations grouped in clusters require different modeling choices.
- Do not treat a cluster test as an exposure model. A clustering statistic does not automatically adjust for confounders or estimate an exposure association.
- Do not call a spatial intensity ratio absolute risk without justification. Its meaning depends on the case–control sampling design and model.
- Do not assume spatial adjustment fixes design problems. Control selection, matching, and confounding still require attention.
- Do not declare a universally best method. The cited methods address different data structures and targets; none establishes one approach as best for all case–control studies.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

