A foundation model is a large AI model trained on a very broad body of data, such as text, images and code, so that it can be adapted to many different tasks rather than built for one. The chat assistants, writing tools and AI search answers most people now use are applications built on top of a handful of these models.
How foundation models work
Building one happens in stages. First comes pre-training: the model learns from an enormous collection of public web pages, books, code and licensed material, typically by predicting the next piece of text again and again. That produces a base model with broad general knowledge and language ability, but no particular sense of how to behave in a conversation.
Next, the developer trains it further to follow instructions, answer helpfully and refuse harmful requests, often using human feedback. The result is the assistant model people interact with. Businesses and software companies can then adapt it again, through prompting, fine-tuning, or by connecting it to their own documents and tools.
Most foundation models today are large language models, but the term is wider. Many are now multimodal models that read and produce images, audio or video as well as text. “Frontier model” is used for the most capable models from the largest developers at any given time.
A foundation model’s knowledge stops at its training cut-off date. To answer questions about recent events or a specific business, it needs fresh information supplied at the time of the question, a process called grounding. AI search products do this by running web searches and passing the results to the model.
Why it matters
At the time of writing (October 2026), a small number of foundation models from OpenAI, Google, Anthropic, Meta and a few others sit behind a very large share of AI tools. Thousands of products that look different are often the same underlying model with a different interface and instructions. Knowing that helps you judge tools: the useful questions are which model a product uses, what data it can access, and what happens to the information you put in.
For search visibility, it matters in two ways. What a model learned in training shapes how it describes your industry and, sometimes, your brand. What it retrieves at the time of a question decides whether your pages are cited in its answer. A business that is clearly described across its own site, its Google Business Profile and reputable third-party sources gives both stages consistent material to work with.
Common mistakes
- Assuming the model knows your business. Unless you are widely written about, it probably knows little, and may guess.
- Treating output as current. Without grounding, answers reflect the training data, which may be a year or more old.
- Choosing tools on branding alone. Two products with different names may run the same model; compare data handling and features instead.
- Ignoring data terms. Check whether a tool’s provider can use your inputs to train future models, and switch that off where you can for business data.
How to act on it
Ask the main AI assistants what they know about your business and your services in your area, and note what is wrong or missing. Make sure your website states clearly who you are, what you do, where you work and how to contact you, in plain text rather than only in images. Keep your Google Business Profile and key directory listings consistent. When choosing AI tools for your team, ask which foundation model they use and read the data terms before you upload anything sensitive.
Checking how AI models describe and cite a business, and fixing the gaps, is what my generative engine optimisation service covers.
