Caesar AI Atlas

Evals

Also known as: Evals (Evaluation)

Caesar AI Atlas Definition

Evals are structured evaluations used to measure the quality, safety, robustness, or performance of an AI system. The term is often used for LLM evaluations, but it can also refer more broadly to repeatable tests and metrics for any model or system.

Other Definitions

Evals Source

Primarily used as an abbreviation for LLM evaluations. More broadly, evals is an abbreviation for any form of evaluation.

Evals Source

Systematic tests for quality, safety and robustness (e.g., accuracy, bias, jailbreak resistance). Should be repeatable and linked to risk level.

Related Terms