Responsible AI / Safety · Level 2 of 5
Adversarial Example
An input crafted to cause a model to behave incorrectly or undesirably.
Attack assumptions and allowed changes define the threat model.
Example
A small image perturbation changes a classifier's prediction.
Listen to the definition and example
Audio transcript
Adversarial Example. An input crafted to cause a model to behave incorrectly or undesirably. Attack assumptions and allowed changes define the threat model. For example: A small image perturbation changes a classifier's prediction.
Explore this concept
Why it matters
This helps you identify a concrete failure mode or evaluate the limits of a proposed control.
Start with
Related concepts
Quick recall question
Try answering before looking back at the definition.