Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchFor two pandas columns, call Series.corr(): r = df["height"].corr(df["weight"]). It calculates Pearson correlation by default, aligns the columns by index, and excludes rows where either value is missing. Use Spearman or Kendall when a rank-based measure better fits the relationship, and use SciPy when you also need a p-value.
Calculate correlation between two pandas columns
Use Series.corr() when each row represents a paired observation and the Series indexes correctly identify those rows:
r = df["height"].corr(df["weight"])
print(r)
The default method is Pearson. To calculate rank correlations instead, pass a method:
r_spearman = df["height"].corr(df["weight"], method="spearman")
r_kendall = df["height"].corr(df["weight"], method="kendall")
Pandas aligns Series by index before calculating the result. If the indexes do not represent the same observations, reset or otherwise correct them before calling corr(); otherwise values may be paired differently than intended. Missing values are excluded pairwise. See the pandas Series.corr documentation.
#1 Best Overall
Calculate a correlation matrix in pandas
To compare every numeric column with every other numeric column, use DataFrame.corr(). The default is Pearson; the method can be changed, and min_periods sets a minimum number of paired observations for a result.
corr_matrix = df.corr() # Pearson
spearman_matrix = df.corr(method="spearman")
kendall_matrix = df.corr(method="kendall")
minimum_n_matrix = df.corr(min_periods=10)
Pandas calculates each pair using pairwise complete observations, so different cells in the matrix can be based on different sample sizes when columns have missing data. min_periods=10 requires at least 10 non-missing pairs for a correlation to be returned; it does not make the sample size equal across pairs. Consult pandas DataFrame.corr for supported methods and parameters.
Rank #2
- Python Data Science Handbook
Choose Pearson, Spearman, or Kendall
The methods measure different kinds of association. Choose based on the question and data, rather than selecting the method that produces the largest coefficient.
| Method | Relationship measured | Python call | Returns a p-value? | Main cautions |
|---|---|---|---|---|
| Pearson | Linear association between quantitative variables | scipy.stats.pearsonr(x, y) or df.corr() |
Yes with SciPy; no with pandas | Sensitive to influential outliers; a near-zero value can miss nonlinear patterns; constant inputs make the result undefined. |
| Spearman | Monotonic association measured using ranks | scipy.stats.spearmanr(x, y) or df.corr(method="spearman") |
Yes with SciPy; no with pandas | Use for ordinal data or monotonic relationships that are not necessarily linear. Ties and missing pairs need attention. |
| Kendall | Rank or ordinal association, expressed as Kendall’s tau | scipy.stats.kendalltau(x, y) or df.corr(method="kendall") |
Yes with SciPy; no with pandas | Consider ties and sample size when interpreting the estimate and test. |
When Pearson fits
Pearson is the usual choice when the question is whether two quantitative variables have a linear relationship. Its coefficient summarizes how closely the values vary together in a straight-line pattern. A curved relationship can exist even when Pearson correlation is near zero, so inspect a scatter plot rather than relying on the coefficient alone. SciPy describes Pearson correlation as measuring the linear relationship between two datasets: scipy.stats.pearsonr.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Rank #3
When Spearman fits
Spearman converts observations to ranks and measures whether higher values of one variable tend to accompany higher (or lower) values of the other. It is suitable for ordinal data and for monotonic relationships that are not linear. A monotonic relationship consistently moves in one direction, though its rate of change may vary. See scipy.stats.spearmanr and the Python statistics documentation.
When Kendall fits
Kendall’s tau is another rank-based measure for ordinal association. Use it when tau is the measure you want to report; SciPy’s function also provides a test result. Review its treatment of tied values and the suitability of the test for your data: scipy.stats.kendalltau.
Rank #4
Calculate Pearson correlation with NumPy
For NumPy arrays, np.corrcoef() returns Pearson product-moment correlation coefficients. With two one-dimensional arrays, select the off-diagonal entry for the correlation between them:
import numpy as np
r = np.corrcoef(x, y)[0, 1]
If a two-dimensional array has observations in rows and variables in columns, set rowvar=False so NumPy treats columns as variables:
Recommended Free Tools
Best Value
matrix = np.corrcoef(array, rowvar=False)
See the NumPy corrcoef documentation for the function’s input and output conventions.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Get a correlation p-value with SciPy
Pandas correlation methods return coefficients, not p-values. Use SciPy’s association-test functions when you need both a test statistic and p-value:
from scipy.stats import pearsonr, spearmanr, kendalltau
pearson_result = pearsonr(x, y)
spearman_result = spearmanr(x, y)
kendall_result = kendalltau(x, y)
print(pearson_result.statistic, pearson_result.pvalue)
Each test has assumptions and tests a null hypothesis of no association according to that method. A small p-value is not a measure of how strong or practically important the relationship is, and it does not show that one variable causes the other. Report the coefficient, p-value, and number of paired observations, and explain the sampling context.
Interpret the coefficient and check the data
Correlation coefficients range from −1 to +1. Positive values indicate that larger values in one variable tend to accompany larger values in the other; negative values indicate that larger values tend to accompany smaller ones. A value near zero suggests little linear association for Pearson or little rank-based monotonic association for Spearman and Kendall. It does not rule out a curved pattern or other structure in the data. The meaning of “near zero” also depends on uncertainty and context.
- Verify paired rows. Confirm that each x value belongs with the y value in the same observation. For pandas Series, check that index alignment is intentional.
- Check missingness and sample size. Count rows where both variables are present. State this effective paired sample size with the reported coefficient; pairwise deletion means it can differ by variable pair in a matrix.
- Check for constant or nearly constant inputs. Correlation is undefined when a variable has no variation. SciPy can issue a
ConstantInputWarningand return NaN for constant inputs; nearly constant values can also cause numerical inaccuracy. Review the warnings and data rather than treating such a result as an ordinary coefficient. - Plot the paired values. A scatter plot can reveal curvature, clusters, or influential observations that a single number hides.
- Separate association from cause. Correlation describes association; the coefficient and p-value alone do not establish a causal effect.
For the documented constant-input behavior and method details, see the relevant SciPy pages for Pearson and Spearman.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

