Autorater evaluation is an automated evaluation process in which a model or scoring system judges the quality of model outputs. Autoraters may be prebuilt or fine-tuned for a specific evaluation task, and their reliability depends on the quality of their criteria, training data, and validation against human judgments.