An embedding is a representation of a piece of content, usually text, as a long list of numbers (a vector) produced by a machine learning model, arranged so that items with similar meanings end up with similar numbers. It lets software compare meaning rather than matching exact words.
How embeddings work
A model trained on huge amounts of text learns which words and phrases appear in similar contexts. When you pass it a sentence, a paragraph or a whole page, it outputs a vector, often several hundred or a few thousand numbers long. Each number on its own means nothing to a person; what matters is the position of the vector relative to others.
To compare two pieces of text, you compare their vectors, typically with cosine similarity. “How much does a loft conversion cost in London” and “price of converting an attic in the capital” share almost no words, but their embeddings will sit close together because they mean nearly the same thing. A paragraph about loft insulation will be further away, and one about car insurance further still.
Embeddings are the engine behind semantic search. Search systems embed documents, or chunks of them, in advance and store the vectors in an index. When a query arrives, it is embedded too, and the system retrieves the nearest passages. In retrieval-augmented generation, those passages are then handed to a language model to write an answer.
Why it matters
Google has used neural matching and related techniques for years to connect queries with pages that do not repeat the exact words, and AI answer features rely on retrieving relevant passages before they generate a response. Precise details of Google’s systems are not public, but the direction is clear: pages are matched on meaning, often passage by passage, not just on keywords.
For a UK business that has two practical consequences. First, writing the same phrase over and over does not help; covering the topic clearly and completely does. Second, each section of a page may be judged on its own. A clear, self-contained answer to “do I need planning permission for a garden room” is easier to retrieve and quote than the same information scattered across a long sales page.
Common mistakes
- Treating it as a new keyword trick. Some tools sell “vector optimisation” as a score to push up. A high similarity between your page and a query does not mean the page will rank.
- Vague headings and passages. Sections that wander across several questions produce muddled vectors and are less likely to be retrieved for any of them.
- Ignoring context. Pronouns and references like “this service” mean little when a passage is pulled out on its own.
- Over-trusting one model. Different embedding models place text differently; none of them is Google’s.
How to act on it
Write pages so each section answers one question in plain words and makes sense if read alone. Name the subject rather than relying on “it” and “this”. Cover the concepts a knowledgeable person would expect, the kind of thing topic clusters help you plan, without padding.
On your own site, embeddings are a useful diagnostic. Embedding every page and comparing them shows which pages overlap so heavily they may compete, and which questions your audience asks that no page answers. I use this kind of analysis in AI search optimisation work alongside conventional keyword research.
