Caesar AI Atlas

Side-by-side Evaluation

Also known as: Side by side Evaluation

Caesar AI Atlas Definition

Side-by-side evaluation compares two models or outputs by presenting responses to the same prompt or task for judgment. Human or automated raters assess which output is better according to criteria such as accuracy, usefulness, safety, style, or preference.

Other Definitions

Side-by-side Evaluation Source

Comparing the quality of two models by judging their responses to the same prompt . For example, suppose the following prompt is given to two different models: Create an image of a cute dog juggling three balls. In a side-by-side evaluation, a rater would pick which image was "better" (More accurate? More beautiful? Cuter?).

Related Terms