Language Models

Multimodal

AI that handles more than one kind of input or output, such as text, images, audio and video.

In everyday terms

You can show a multimodal AI a photo of your fridge and ask what to cook, or speak to it and hear it reply.

For professionals

Models with encoders (or shared tokenisation) for several modalities, mapped into a common representation space.

Think of it like…

A colleague who can read, look at pictures and listen, not just read emails.

You've already seen it

Uploading screenshots to ChatGPT, live voice mode, describing an image for accessibility.

Myth vs reality

Myth: Multimodal AI "sees" like we do.

Reality: Images are converted into numbers it processes. It can miss things humans find obvious.

Quick check

Which is a multimodal task?

Show answer

Describe what's in a photo: Image in, text out = two modalities.

Builds on

Generative AI

Related

Generative AI · Diffusion Model

🔎esc