Retrieval-augmented generation, usually shortened to RAG, is a way of using a language model in which relevant documents are retrieved first and the model writes its answer from them. The retrieval step runs when the question arrives, so the model can answer about things it never saw during training: yesterday's news or a product page published last week.
The term comes from a 2020 paper by Patrick Lewis and colleagues at Facebook AI Research, who combined a retrieval system with a generative model and showed that the pair answered factual questions better than the model alone. The idea has since become the standard design for chat over documents, customer support bots and, at web scale, AI search.
Why does RAG exist?
A large language model on its own has two limits. Its knowledge ends at the date of its training data, and it has seen little text about rare or new topics. Asked about either, it tends to produce a fluent answer with a wrong fact in it.
Retraining the model every day is not practical. Putting the relevant text in front of the model at the time of the question is. If the prompt contains the paragraph that answers the question, the model only has to read and rephrase it, which is a task it does reliably. RAG also makes citations possible: because the system knows which passages it gave the model, it can show the reader where each claim came from.
How does RAG work?
A RAG system has an offline part that prepares the documents and an online part that runs for every question.
Indexing
Documents are collected and split into passages, a step called chunking. A passage is typically a paragraph or a section, a few hundred tokens long. Each passage is stored with a keyword index and, in most systems, an embedding: a list of numbers that places the passage in a space where texts with similar meaning sit close together.
Retrieval
When a question comes in, the system turns it into one or more search queries and looks up candidate passages. Keyword search finds passages that share terms with the query. Vector search finds passages whose embeddings are close to the query's embedding, which catches matches that use different words for the same idea. Most production systems use both, a combination called hybrid search.
The first pass usually returns more passages than the model can read. A reranker then scores each candidate against the question more carefully and keeps the top few.
Generation
The selected passages are placed in the prompt together with the question and instructions such as "answer only from the sources below and cite them". The model writes the answer. If the system tracks which passage each sentence was drawn from, it can attach a citation to it.
What makes a RAG answer good or bad?
Most failures happen before the model writes a word. If the retrieval step returns passages that are off topic, out of date or missing the fact that answers it, the model has nothing good to work with. It will either say it cannot answer or, worse, fill the gap from its training data.
Three things decide retrieval quality. The index has to contain the right documents. The passages have to be cut so that each one makes sense on its own. And the scoring has to prefer passages that answer the question over passages that merely mention its words. How AI search chooses which sources to cite looks at what that scoring rewards.
What does RAG mean for a website?
AI search products are RAG systems whose document collection is the web. Every design decision above has a consequence for the pages they read.
Because retrieval works on passages, a page is judged section by section. A section that opens with a direct answer and keeps to one topic is easy to retrieve and easy to quote. A section that depends on context from elsewhere on the page is not.
Embeddings match meaning rather than exact wording, so a page does not need to repeat the question's phrasing. It does need to say the thing clearly. The model only reads what the retriever returns, so a page that cannot be crawled or indexed is outside the system entirely, however well it is written.