Retrieval-Augmented Generation (RAG)
Letting the AI look things up in your documents before it answers.
In everyday terms
Instead of relying on memory, the system first searches a knowledge base, pastes the relevant bits into the prompt, then the AI answers using them, ideally with sources.
For professionals
Retrieve top-k chunks (often via embeddings), inject into context, generate a grounded answer. Quality hinges on chunking, retrieval and citation.
Think of it like…
An open-book exam instead of a closed-book one.
You've already seen it
Company chatbots that answer from internal policies, AI search with citations.
Myth vs reality
Myth: RAG means the AI is retrained on your documents.
Reality: Nothing is retrained. Relevant text is fetched and shown to the model at question time.
Quick check
What does RAG add before the AI answers?
- A retraining step
- A search for relevant information
- A higher temperature
- A human reviewer
Show answer
A search for relevant information: Retrieve, then generate.
Builds on
Related
Embedding · Vector Database · Hallucination · Context Window