Human Evaluation ist der Prozess, bei dem Menschen Qualität, Sicherheit, Nützlichkeit oder Korrektheit von Modellausgaben beurteilen. Sie ist besonders wichtig für Aufgaben, bei denen es keine einzelne richtige Antwort gibt, etwa Übersetzungsqualität, Zusammenfassung, Dialog oder kreative Generierung.
A process in which people judge the quality of an ML model's output; for example, having bilingual people judge the quality of an ML translation model. Human evaluation is particularly useful for judging models that have no one right answer. Contrast with automatic evaluation and autorater evaluation.