Responsible AI / Safety · Level 2 of 5
Alignment
Making system behavior better match intended goals, constraints, or human preferences.
The intended values and evaluation criteria must be specified.
Example
A model is trained and evaluated to follow task instructions within defined boundaries.
Listen to the definition and example
Audio transcript
Alignment. Making system behavior better match intended goals, constraints, or human preferences. The intended values and evaluation criteria must be specified. For example: A model is trained and evaluated to follow task instructions within defined boundaries.
Explore this concept
Why it matters
This helps you identify a concrete failure mode or evaluate the limits of a proposed control.
Start with
Related concepts
Quick recall question
Try answering before looking back at the definition.