Toxizität ist der Grad, in dem Inhalte beleidigend, bedrohlich, hasserfüllt, belästigend oder anderweitig anstößig sind. KI-Systeme können Toxizität über verschiedene Dimensionen hinweg erkennen oder bewerten, um Moderation, Sicherheitsevaluation und Inhaltsgovernance zu unterstützen.
The degree to which content is abusive, threatening, or offensive. Many machine learning models can identify, measure, and classify toxicity. Most of these models identify toxicity along multiple parameters, such as the level of abusive language and the level of threatening language.