RealToxicityPrompts — набор начальных фрагментов предложений, используемый для оценки того, генерируют ли языковые модели токсичные продолжения. Он часто применяется для измерения и сравнения рисков токсичности в сгенерированном тексте, обычно с использованием автоматических классификаторов токсичности.
A dataset that contains a set of sentence beginnings that might contain toxic content. Use this dataset to evaluate an LLM's ability to generate non-toxic text to complete the sentence. Typically, you use the Perspective API to determine how well the LLM performed at this task. See RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models for details.