Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For two pandas columns, call Series.corr(): r = df["height"].corr(df["weight"]). It calculates Pearson correlation by default, aligns the columns by index, and excludes rows where either value is missing. Use Spearman or Kendall when a rank-based measure better fits the relationship, and use SciPy when you also need a p-value.

Calculate correlation between two pandas columns

Use Series.corr() when each row represents a paired observation and the Series indexes correctly identify those rows:

r = df["height"].corr(df["weight"])
print(r)

The default method is Pearson. To calculate rank correlations instead, pass a method:

r_spearman = df["height"].corr(df["weight"], method="spearman")
r_kendall = df["height"].corr(df["weight"], method="kendall")

Pandas aligns Series by index before calculating the result. If the indexes do not represent the same observations, reset or otherwise correct them before calling corr(); otherwise values may be paired differently than intended. Missing values are excluded pairwise. See the pandas Series.corr documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Calculate a correlation matrix in pandas

To compare every numeric column with every other numeric column, use DataFrame.corr(). The default is Pearson; the method can be changed, and min_periods sets a minimum number of paired observations for a result.

corr_matrix = df.corr()  # Pearson
spearman_matrix = df.corr(method="spearman")
kendall_matrix = df.corr(method="kendall")
minimum_n_matrix = df.corr(min_periods=10)

Pandas calculates each pair using pairwise complete observations, so different cells in the matrix can be based on different sample sizes when columns have missing data. min_periods=10 requires at least 10 non-missing pairs for a correlation to be returned; it does not make the sample size equal across pairs. Consult pandas DataFrame.corr for supported methods and parameters.

Choose Pearson, Spearman, or Kendall

The methods measure different kinds of association. Choose based on the question and data, rather than selecting the method that produces the largest coefficient.

Method Relationship measured Python call Returns a p-value? Main cautions
Pearson Linear association between quantitative variables scipy.stats.pearsonr(x, y) or df.corr() Yes with SciPy; no with pandas Sensitive to influential outliers; a near-zero value can miss nonlinear patterns; constant inputs make the result undefined.
Spearman Monotonic association measured using ranks scipy.stats.spearmanr(x, y) or df.corr(method="spearman") Yes with SciPy; no with pandas Use for ordinal data or monotonic relationships that are not necessarily linear. Ties and missing pairs need attention.
Kendall Rank or ordinal association, expressed as Kendall’s tau scipy.stats.kendalltau(x, y) or df.corr(method="kendall") Yes with SciPy; no with pandas Consider ties and sample size when interpreting the estimate and test.

When Pearson fits

Pearson is the usual choice when the question is whether two quantitative variables have a linear relationship. Its coefficient summarizes how closely the values vary together in a straight-line pattern. A curved relationship can exist even when Pearson correlation is near zero, so inspect a scatter plot rather than relying on the coefficient alone. SciPy describes Pearson correlation as measuring the linear relationship between two datasets: scipy.stats.pearsonr.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When Spearman fits

Spearman converts observations to ranks and measures whether higher values of one variable tend to accompany higher (or lower) values of the other. It is suitable for ordinal data and for monotonic relationships that are not linear. A monotonic relationship consistently moves in one direction, though its rate of change may vary. See scipy.stats.spearmanr and the Python statistics documentation.

When Kendall fits

Kendall’s tau is another rank-based measure for ordinal association. Use it when tau is the measure you want to report; SciPy’s function also provides a test result. Review its treatment of tied values and the suitability of the test for your data: scipy.stats.kendalltau.

Calculate Pearson correlation with NumPy

For NumPy arrays, np.corrcoef() returns Pearson product-moment correlation coefficients. With two one-dimensional arrays, select the off-diagonal entry for the correlation between them:

import numpy as np

r = np.corrcoef(x, y)[0, 1]

If a two-dimensional array has observations in rows and variables in columns, set rowvar=False so NumPy treats columns as variables:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
matrix = np.corrcoef(array, rowvar=False)

See the NumPy corrcoef documentation for the function’s input and output conventions.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Get a correlation p-value with SciPy

Pandas correlation methods return coefficients, not p-values. Use SciPy’s association-test functions when you need both a test statistic and p-value:

from scipy.stats import pearsonr, spearmanr, kendalltau

pearson_result = pearsonr(x, y)
spearman_result = spearmanr(x, y)
kendall_result = kendalltau(x, y)

print(pearson_result.statistic, pearson_result.pvalue)

Each test has assumptions and tests a null hypothesis of no association according to that method. A small p-value is not a measure of how strong or practically important the relationship is, and it does not show that one variable causes the other. Report the coefficient, p-value, and number of paired observations, and explain the sampling context.

Interpret the coefficient and check the data

Correlation coefficients range from −1 to +1. Positive values indicate that larger values in one variable tend to accompany larger values in the other; negative values indicate that larger values tend to accompany smaller ones. A value near zero suggests little linear association for Pearson or little rank-based monotonic association for Spearman and Kendall. It does not rule out a curved pattern or other structure in the data. The meaning of “near zero” also depends on uncertainty and context.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Verify paired rows. Confirm that each x value belongs with the y value in the same observation. For pandas Series, check that index alignment is intentional.
  2. Check missingness and sample size. Count rows where both variables are present. State this effective paired sample size with the reported coefficient; pairwise deletion means it can differ by variable pair in a matrix.
  3. Check for constant or nearly constant inputs. Correlation is undefined when a variable has no variation. SciPy can issue a ConstantInputWarning and return NaN for constant inputs; nearly constant values can also cause numerical inaccuracy. Review the warnings and data rather than treating such a result as an ordinary coefficient.
  4. Plot the paired values. A scatter plot can reveal curvature, clusters, or influential observations that a single number hides.
  5. Separate association from cause. Correlation describes association; the coefficient and p-value alone do not establish a causal effect.

For the documented constant-input behavior and method details, see the relevant SciPy pages for Pearson and Spearman.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.