Caesar AI Atlas

Human Evaluation

Caesar AI Atlas Definition

Human evaluation is the process of having people judge the quality, safety, usefulness, or correctness of model outputs. It is especially important for tasks where there is no single correct answer, such as translation quality, summarization, dialogue, or creative generation.

Other Definitions

Human Evaluation Source

A process in which people judge the quality of an ML model's output; for example, having bilingual people judge the quality of an ML translation model. Human evaluation is particularly useful for judging models that have no one right answer. Contrast with automatic evaluation and autorater evaluation.

Related Terms