The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →A week-to-week change in Cohen’s kappa does not, by itself, mean that raters have become better or worse. Kappa depends both on how often the raters agree and on the category frequencies each rater assigns. A shift in those frequencies—or a change in the mix of easy and difficult cases—can move kappa even when the underlying rating process has not simply improved or declined.
Start by comparing the weekly case counts, contingency tables, raw agreement, and each rater’s category proportions. Then interpret the kappa change in that context.
Why Cohen’s kappa changes from one week to the next
For two raters assigning nominal categories, Cohen’s kappa is calculated as κ = (Po − Pe) / (1 − Pe). Here, Po is the observed proportion of matching ratings, while Pe is the agreement expected from the raters’ category proportions. The formula is described in the foundational discussion of Byrt, Bishop, and Carlin’s 1993 paper, “Bias, prevalence and kappa”.
That adjustment means kappa can change for more than one reason. If the raters match on a different proportion of cases, Po changes. If the distribution of categories assigned by either rater shifts, Pe can change. Both can happen at once. As a result, the same raw agreement in two weeks can produce different kappa values when the category marginals differ.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
The cases themselves matter too. One week may contain a greater share of straightforward examples, and another may contain more ambiguous ones. Agreement can differ with that case mix; it is a possibility to investigate, not proof that a change in kappa is harmless. Vach discusses the role of sample composition and prevalence in a 2005 article on Cohen’s kappa.
Diagnose the change before calling it improvement or decline
- Check that the weeks are comparable. Confirm that category definitions, inclusion rules, rater pairing, treatment of missing or duplicate ratings, and the kappa variant are consistent.
- Compare the denominators and case mix. Record how many cases have ratings from both raters. Check whether case types, sources, or difficulty mix changed.
- Put the contingency tables side by side. Review each cell, not just the total agreement. Note which categories account for matches and where disagreements occur.
- Compare the ingredients of kappa. For each week, calculate or report Po, Pe, and each rater’s category proportions. This shows whether the change tracks observed matches, the marginal-based adjustment, or both. Byrt, Bishop, and Carlin recommend reporting prevalence and bias information alongside kappa in their discussion.
- Show uncertainty. Report a suitable confidence interval for each weekly estimate and consider uncertainty in the difference between weeks. A small difference between point estimates alone is not enough to establish a meaningful change. The appropriate interval method depends on the study design; there is no single method established for every repeated weekly comparison.
- Investigate operational changes if the ratings changed. Check for rater turnover, retraining, revised instructions, changed tools, or a new kind of borderline case when raw agreement or particular disagreement cells move.
- Confirm that the measure fits the data. Cohen’s kappa is for two raters. For ordered categories where near and far disagreements should count differently, consider weighted kappa. Data from more than two raters require a method designed for that setting.
How to report weekly agreement
Give readers enough information to see what moved, rather than presenting kappa as a standalone score. A useful weekly report includes:
Rank #2
- This guide is a perfect overview for the topics covered in introductory statistics courses.
- the number of jointly rated cases, n;
- the two-rater contingency table;
- observed agreement, Po;
- Cohen’s kappa and its uncertainty interval;
- each rater’s category proportions; and
- when relevant, the case-type distribution and a brief note of protocol or rater changes.
When comparing weeks, describe whether raw agreement, category proportions, case mix, or more than one of these changed. Avoid labelling the result “reliability improved” or “reliability declined” until you have examined those inputs.
Why high raw agreement can coexist with low kappa
Raw agreement answers how often the raters gave matching labels. Kappa adjusts that observed agreement using the agreement expected from the raters’ marginal category frequencies. When the prevalence pattern makes the expected-agreement adjustment influential, a high match rate and a comparatively low kappa can coexist.
Rank #3
This does not make kappa a direct diagnosis of why the raters disagree, nor does prevalence dependence automatically make the statistic defective. Interpretation depends on the population being rated and the question the measure is meant to answer. Vach’s discussion distinguishes observed marginal prevalence from latent prevalence and emphasizes population composition; see the 2005 article. Report raw agreement and context with kappa rather than relying on a universal “good” threshold.
If prevalence-related interpretation is central to your use case, Gwet’s AC1 can be included as a sensitivity comparison, with an explanation of what it estimates and why it is relevant. An open-access paper on the high-agreement, high-prevalence paradox argues that AC1 is more robust in the scenarios it discusses; that is not a universal reason to replace kappa or select whichever statistic gives a more favourable result.
Rank #4
Choose an agreement measure that matches the ratings
The right statistic depends on the design and the meaning of disagreement, not on which one produces the most reassuring number.
| Rating design or question | Measure to consider | Why it may fit |
|---|---|---|
| Two raters; nominal categories | Cohen’s kappa | Chance-corrected agreement for two raters assigning categories without an ordering. |
| Two raters; ordered categories, with disagreement severity relevant | Weighted kappa | Can treat close disagreements differently from more substantial ones. |
| More than two raters | A method suited to multi-rater data | Cohen’s kappa is a two-rater measure. |
| Prevalence-related interpretation is a key concern | AC1 as a sensitivity comparison, where justified | Provides another perspective; explain the measure’s assumptions and do not use it simply to obtain a higher score. |
For a deeper methods treatment, Wiley’s description of Measuring Agreement: Models, Methods, and Applications notes that it covers Cohen’s kappa and other categorical-data measures, sample-size determination, case studies, and R resources.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsQuick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

