Also known as: Prompt Injection and Jailbreak Β· Prompt Injection Jailbreak
Prompt injection or jailbreak refers to adversarial input intended to bypass a model's instructions, safety policies, or operational controls. Mitigations include input filtering, retrieval isolation, least-privilege tool access, policy enforcement, and layered guardrails.
Adversarial inputs that subvert a model's instructions or controls. Mitigate with input filtering, retrieval isolation, guardrails and least-privilege design.