Also known as: Evals (Evaluation)
Evals are structured evaluations used to measure the quality, safety, robustness, or performance of an AI system. The term is often used for LLM evaluations, but it can also refer more broadly to repeatable tests and metrics for any model or system.