Caesar AI Atlas

Prompt Injection

Caesar AI Atlas Definition

Prompt injection is an adversarial technique that manipulates a language model by placing malicious or unintended instructions in user input, documents, web pages, or other content the model processes. It can cause the model to ignore prior instructions, disclose restricted information, alter outputs, or trigger unauthorized actions.

Other Definitions

Prompt Injection Source

Prompt injection manipulates an LLM's behaviour (for example to alter response style, retrieve hidden or restricted data, or disrupt interactions) by embedding specific instructions within a prompt. This approach exploits the model's tendency to follow instructions within the prompt sequence, even if instructions are unintended or malicious.

Also Referenced In

Concept Comparisons

Related Terms