Automation and AI

Context Window

Also called context length

The maximum amount of text, measured in tokens, that an AI language model can take into account in a single request, including its own reply.

Quick facts: Context Window

Category
Automation and AI
Also called
context length
Level
Intermediate
Affects
AI prompt quality, long-document tasks, AI search retrieval, AI tool costs
Where to see it
AI model documentation, token counters in AI developer consoles, chat tool settings
In this article4
  1. How a context window works
  2. Why it matters
  3. Common mistakes
  4. How to act on it

A context window is the maximum amount of text an AI language model can take into account at one time, counting your instructions, any documents you paste in, the conversation so far and the reply it writes. Anything outside that window is invisible to the model for that request.

How a context window works

Models do not read words; they read tokens, which are chunks of text. In English, a token is roughly three quarters of a word on average, so a 1,000-word article uses something like 1,300 tokens. Every large language model has a fixed limit on how many tokens it can handle in one go, and that limit is its context window.

Everything competes for that space. A long system instruction, a pasted brand guide, three competitor pages and a long chat history all use tokens before the model writes a single word of its answer. When a conversation grows beyond the limit, older messages are dropped or summarised, which is why a chatbot sometimes forgets what you agreed at the start.

At the time of writing (October 2026), the leading commercial models offer windows measured in hundreds of thousands of tokens, and some reach around a million. Those figures change often, and the limit you get can depend on the product and plan you use, so check the current documentation rather than relying on a number you read last year.

A bigger window is not the same as better reading. Models tend to use information at the start and end of a long input more reliably than material buried in the middle. Large inputs also cost more and take longer to process.

Why it matters

For a UK business using AI day to day, the context window decides what you can sensibly ask for in one request. Summarising a 40-page tender document, comparing five supplier contracts or checking a whole website section for consistent tone all depend on whether the material fits, and on whether the model can still pay attention to the details once it does.

It matters for search visibility too. AI search tools typically retrieve passages from web pages and place them in the model’s context before writing an answer, a method called retrieval-augmented generation. They cannot load every page in full, so they work with selected chunks. A page where each section makes sense on its own, with the answer near the top, gives those systems a better passage to work with.

Common mistakes

  • Pasting everything in. Dumping a whole website export into a prompt often produces vaguer answers than sending the three pages that matter.
  • Assuming the model read it all. A fact buried on page 30 of a long upload can be missed or misquoted. Ask the model to quote the passage it relied on.
  • Running very long chats. After hours of back and forth, early instructions about tone or audience may have dropped out. Start a fresh conversation with a short summary instead.
  • Confusing context with memory. Unless a product has a separate memory feature, nothing carries over between sessions.
  • Uploading data you should not. Customer records pasted into a public AI tool may be processed in ways your privacy notice does not cover.

How to act on it

Treat context as a budget. Put the instruction first, give only the material the task needs, and state the output you want. For long documents, ask for a summary section by section rather than in one pass, then ask follow-up questions about specific parts.

On your website, write sections that stand alone: a clear heading, a direct answer in the first sentence or two, then the detail. That helps human readers who skim, and it gives AI systems a clean passage to retrieve. Making pages easy for AI tools to find, quote and cite is the focus of my generative engine optimisation service.

Do and do not

Do

  • Give the model only the material the task needs
  • Ask the model to quote the passage it relied on
  • Write page sections that make sense on their own

Do not

  • Assume a long upload was read evenly
  • Run one chat for hours and expect it to remember the start
  • Paste customer personal data into public AI tools

Questions people ask about this

What happens when you exceed the context window?

It depends on the tool. Some refuse the request and ask you to shorten it, while chat products usually drop or summarise the oldest messages so the conversation can continue. Either way, the model can no longer see the material that fell outside the window, which can lead to answers that ignore earlier instructions.

Is a bigger context window always better?

Not always. A larger window lets you include whole documents rather than extracts, which helps with tasks such as comparing contracts. But you pay for every token you send, and long requests run more slowly, so filling the window by habit gets expensive. For most marketing tasks, choosing the right material matters more than the size of the window.

How many words fit in a context window?

As a rough rule, one token is about three quarters of an English word, so 100,000 tokens is in the region of 75,000 words. Tables, code, unusual words and other languages use more tokens per word. Remember that the model's reply also uses part of the window.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.