Inter-rater agreement measures how consistently raters give the same label to the same items, beyond what chance alone would produce. For a pair of raters, Cohen’s kappa is κ = (p_o − p_e) / (1 − p_e), where p_o is the observed agreement and p_e the agreement expected from each rater’s label frequencies (Cohen, A Coefficient of Agreement for Nominal Scales); Krippendorff’s alpha extends the idea to more raters and missing labels.
In FDE interviews
It comes up whenever human labels become ground truth, in a review queue or when a is checked against experts. Raw agreement flatters a skewed label set: if defects are rare, raters who mark nearly everything “fine” still agree on most items, and kappa exposes that. Two reviewers who each mark 5% of items defective and agree on 95% of items reach only κ ≈ 0.47, because chance alone gives p_e = 0.905. Published bands such as Landis and Koch’s (0.61–0.80 “substantial”) are conventions, not a standard; the working bar is the kappa your own reviewers reach with each other on this label set.
A strong candidate has reviewers label an overlapping sample, rewrites the guideline wherever they disagree, and holds a model judge to the agreement the humans reach with each other, not to a perfect score.
Related: LLM-as-judge, golden set.