Retrieval-augmented generation (RAG) is a method in which an AI system first searches a set of documents for material relevant to a question, then passes that material to a large language model to write the answer. Instead of relying only on what the model absorbed in training, the answer is built from sources fetched at the moment of asking.
How retrieval-augmented generation works
A RAG system has two halves. The retrieval half finds the passages most likely to answer the question. The generation half writes a response using those passages as its evidence.
- Preparation. Documents (web pages, help articles, PDFs, product data) are split into chunks, often a few paragraphs each.
- Indexing. Each chunk is converted into an embedding, a list of numbers that represents its meaning, and stored in a searchable index. Many systems also keep an ordinary keyword index alongside it.
- Retrieval. When a question arrives, it is converted the same way and the closest-matching chunks are pulled back.
- Generation. The model receives the question plus those chunks alongside its instructions and writes an answer, often with links to the sources it used.
This is the principle behind AI answers that cite web pages. When ChatGPT search, Perplexity, Microsoft Copilot or Google’s AI features answer with links, they are retrieving pages from a search index and generating a summary from them. The exact pipelines are not published, but the shape is the same. Supplying sources this way is also called grounding.
Why it matters
RAG matters to UK businesses in two quite different ways.
Visibility in AI search. If an assistant answers “Who does emergency boiler repairs in Leeds?” by retrieving pages, your business can only appear if your pages can be crawled, are in the index the assistant uses, and contain a passage that plainly answers the question. A page that buries its location and services under vague slogans is hard to retrieve, however well designed it looks. Clear, self-contained passages with the facts stated in words, not only in images, give retrieval something to match.
Your own AI tools. If you add an AI assistant to your website or use one internally, RAG is how it answers from your actual policies, prices and product details rather than general knowledge. A chatbot that retrieves from your up-to-date returns policy is far less likely to promise a refund you do not offer. It does not remove the risk of hallucination entirely: the model can still misread or overstate what it retrieved.
Common mistakes
- Blocking the crawlers AI search tools rely on Often by accident through a security setting or a blanket robots.txt rule, then wondering why the business never appears in answers.
- Feeding a RAG system outdated documents. An assistant retrieving last year’s price list will quote last year’s prices with confidence.
- Chunks that make no sense alone. A passage that says “as mentioned above, we cover it” gives retrieval nothing to work with.
- Assuming retrieval means accuracy. Retrieved sources reduce errors; they do not eliminate them.
- Putting personal data into the index. Anything indexed may be returned to whoever asks the right question.
How to act on it
For your website, check that your important pages are crawlable and indexed, then read each key page as if it were cut into short passages. Does the passage about your service say what it is, who it is for and where you work, in plain words? Put facts such as service areas, opening hours and prices in text rather than images, and keep them consistent across your site and profiles.
If you are building an AI assistant on your own content, start with a small, current, well-organised set of documents, assign someone to keep them updated, and test the assistant with real customer questions before launch. For the search side, my work on generative engine optimisation (GEO) covers how assistants retrieve and cite businesses, from crawler access to the way your pages are written.
