Alignment
Making sure AI does what people actually intend, and does it safely.
In everyday terms
An aligned assistant is helpful, honest and avoids harm, even when instructions are vague or someone tries to misuse it.
For professionals
Techniques such as RLHF, constitutional training, evaluations and interpretability research aimed at steering models toward intended goals and values.
Think of it like…
Teaching a powerful new employee your organisation's values, not just the tasks.
You've already seen it
Chatbots declining harmful requests; AI safety research news.
Myth vs reality
Myth: Alignment is just censorship.
Reality: It's broader: being genuinely helpful, honest and safe, including refusing less when refusal isn't needed.
Quick check
Alignment is about AI…
- Running faster
- Pursuing what humans actually intend, safely
- Using less memory
- Looking nicer
Show answer
Pursuing what humans actually intend, safely: Intent and safety.