Responsible AI / Safety · Level 2 of 5
Prompt Injection
Untrusted content attempting to redirect an instruction-following system.
Separating trusted instructions from external data is a core defense principle.
Example
A retrieved page tells an assistant to ignore its user's request.
Listen to the definition and example
Audio transcript
Prompt Injection. Untrusted content attempting to redirect an instruction-following system. Separating trusted instructions from external data is a core defense principle. For example: A retrieved page tells an assistant to ignore its user's request.
Explore this concept
Why it matters
This helps you identify a concrete failure mode or evaluate the limits of a proposed control.
Start with
Quick recall question
Try answering before looking back at the definition.