A context window is the maximum amount of text an AI language model can take into account at one time, counting your instructions, any documents you paste in, the conversation so far and the reply it writes. Anything outside that window is invisible to the model for that request.
How a context window works
Models do not read words; they read tokens, which are chunks of text. In English, a token is roughly three quarters of a word on average, so a 1,000-word article uses something like 1,300 tokens. Every large language model has a fixed limit on how many tokens it can handle in one go, and that limit is its context window.
Everything competes for that space. A long system instruction, a pasted brand guide, three competitor pages and a long chat history all use tokens before the model writes a single word of its answer. When a conversation grows beyond the limit, older messages are dropped or summarised, which is why a chatbot sometimes forgets what you agreed at the start.
At the time of writing (October 2026), the leading commercial models offer windows measured in hundreds of thousands of tokens, and some reach around a million. Those figures change often, and the limit you get can depend on the product and plan you use, so check the current documentation rather than relying on a number you read last year.
A bigger window is not the same as better reading. Models tend to use information at the start and end of a long input more reliably than material buried in the middle. Large inputs also cost more and take longer to process.
Why it matters
For a UK business using AI day to day, the context window decides what you can sensibly ask for in one request. Summarising a 40-page tender document, comparing five supplier contracts or checking a whole website section for consistent tone all depend on whether the material fits, and on whether the model can still pay attention to the details once it does.
It matters for search visibility too. AI search tools typically retrieve passages from web pages and place them in the model’s context before writing an answer, a method called retrieval-augmented generation. They cannot load every page in full, so they work with selected chunks. A page where each section makes sense on its own, with the answer near the top, gives those systems a better passage to work with.
Common mistakes
- Pasting everything in. Dumping a whole website export into a prompt often produces vaguer answers than sending the three pages that matter.
- Assuming the model read it all. A fact buried on page 30 of a long upload can be missed or misquoted. Ask the model to quote the passage it relied on.
- Running very long chats. After hours of back and forth, early instructions about tone or audience may have dropped out. Start a fresh conversation with a short summary instead.
- Confusing context with memory. Unless a product has a separate memory feature, nothing carries over between sessions.
- Uploading data you should not. Customer records pasted into a public AI tool may be processed in ways your privacy notice does not cover.
How to act on it
Treat context as a budget. Put the instruction first, give only the material the task needs, and state the output you want. For long documents, ask for a summary section by section rather than in one pass, then ask follow-up questions about specific parts.
On your website, write sections that stand alone: a clear heading, a direct answer in the first sentence or two, then the detail. That helps human readers who skim, and it gives AI systems a clean passage to retrieve. Making pages easy for AI tools to find, quote and cite is the focus of my generative engine optimisation service.
