Red teaming es un proceso estructurado de pruebas adversariales usado para identificar debilidades, comportamiento inseguro o salidas dañinas en un sistema de IA. Consiste en sondear deliberadamente el sistema con entradas difíciles, maliciosas o de casos límite para que los desarrolladores puedan evaluar y mejorar las salvaguardias.
Red-teaming means a structured testing effort to find flaws and vulnerabilities in an AI system, often in a controlled environment and in collaboration with developers of AI. AI red-teaming is most often performed by dedicated 'red-teams' that adopt adversarial methods to identify flaws and vulnerabilities, such as harmful or discriminatory outputs from an AI system, unforeseen or undesirable system behaviors, limitations or potential risks associated with the misuse of the system.