Risks & Society

Alignment

Making sure AI does what people actually intend, and does it safely.

In everyday terms

An aligned assistant is helpful, honest and avoids harm, even when instructions are vague or someone tries to misuse it.

For professionals

Techniques such as RLHF, constitutional training, evaluations and interpretability research aimed at steering models toward intended goals and values.

Think of it like…

Teaching a powerful new employee your organisation's values, not just the tasks.

You've already seen it

Chatbots declining harmful requests; AI safety research news.

Myth vs reality

Myth: Alignment is just censorship.

Reality: It's broader: being genuinely helpful, honest and safe, including refusing less when refusal isn't needed.

Quick check

Alignment is about AI…

Show answer

Pursuing what humans actually intend, safely: Intent and safety.

Builds on

Large Language Model (LLM)

Related

AI Bias · Responsible AI · AI Agent

🔎esc