Language Models

Transformer

The design behind modern language AIs. The "T" in GPT.

In everyday terms

Introduced by Google researchers in 2017, transformers read a whole passage at once and work out which words relate to which. That made today's chatbots possible.

For professionals

Architecture built on self-attention and feed-forward layers, parallelisable across sequence positions. Scales efficiently on GPUs.

Think of it like…

Reading a whole page at a glance and drawing lines between related words, rather than reading one word at a time.

You've already seen it

GPT = Generative Pre-trained Transformer.

Myth vs reality

Myth: Transformers were invented for chatbots.

Reality: They were designed for translation. Chatbots came later.

Quick check

What does the "T" in GPT stand for?

Show answer

Transformer: Generative Pre-trained Transformer.

Builds on

Deep Learning

Related

Attention · Large Language Model (LLM) · Neural Network

🔎esc