Also known as: Inter rater Agreement
Inter-rater agreement is a measure of the extent to which multiple human annotators or evaluators make consistent judgments. It is used to assess labeling quality, evaluation reliability, and whether task instructions are sufficiently clear.
A measurement of how often human raters agree when doing a task. If raters disagree, the task instructions may need to be improved. Also sometimes called inter-annotator agreement or inter-rater reliability . See also Cohen's kappa, which is one of the most popular inter-rater agreement measurements. See Categorical data: Common issues in Machine Learning Crash Course for more information.