Le red teaming est un processus structuré de test adversarial utilisé pour identifier les faiblesses, comportements dangereux ou sorties nuisibles d’un système d’IA. Il consiste à sonder délibérément le système avec des entrées difficiles, malveillantes ou limites afin que les développeurs puissent évaluer et améliorer les garde-fous.
Red-teaming means a structured testing effort to find flaws and vulnerabilities in an AI system, often in a controlled environment and in collaboration with developers of AI. AI red-teaming is most often performed by dedicated 'red-teams' that adopt adversarial methods to identify flaws and vulnerabilities, such as harmful or discriminatory outputs from an AI system, unforeseen or undesirable system behaviors, limitations or potential risks associated with the misuse of the system.