A token is the unit of text an AI language model works with: usually a short word, a piece of a longer word, a number or a punctuation mark. Models read your prompt as tokens, write their answers as tokens, and providers measure limits and charge for usage in tokens.
How tokens work
Before a model can process text, a tokeniser splits it into pieces from a fixed vocabulary. Common words like “the” or “shop” are usually one token each. Longer or rarer words are broken up, so a word like “refurbishment” might become several tokens. Spaces, capital letters and punctuation all affect the split. Numbers, product codes and postcodes are often broken into many small tokens, which is one reason models can make slips with figures.
Each provider uses its own tokeniser, so the same paragraph can produce different token counts in different models. Most providers publish a token counter or an estimate in their documentation, and the token figure in your usage dashboard is the one that counts for billing.
Tokens matter in three places:
- The context window The maximum number of tokens a model can consider at once, covering your instructions, any documents you paste in, the conversation so far and the reply.
- Output limits A separate cap on how long a single reply can be.
- Price API usage is billed per token, usually quoted per million tokens, with output tokens typically costing more than input tokens.
The word also has unrelated meanings in marketing. In payments, tokenisation replaces a card number with a stand-in value. In email tools, a personalisation token inserts a field such as a first name. This entry is about AI tokens.
Why it matters
For most people using a chat app on a flat monthly subscription, tokens show up as limits rather than costs: a long document that is cut off, a conversation where the model seems to forget early instructions, or a usage cap reached mid-afternoon.
Once you automate, tokens become a budget line. Suppose you run every new product description, every customer review summary or every inbound enquiry through a model. The monthly cost is roughly the number of jobs multiplied by the tokens in each prompt and reply, priced at the provider’s rate in US dollars and converted to sterling on your card statement. A prompt that pastes your entire brand guide into every request can cost many times more than one that sends only the rules that apply.
Tokens also explain some quality problems. If a long brief, a pasted document and a long chat history push past the context window, older material may be dropped or given less attention, and the model starts ignoring instructions you gave earlier.
Common mistakes
- Equating tokens with words. Token counts are usually higher than word counts, and much higher for tables, code and numbers.
- Pasting everything “just in case”. Irrelevant material costs tokens and can distract the model from what matters.
- Ignoring output costs Which are often the larger share when replies are long.
- Endless chat threads. Starting a fresh conversation with a clean brief is often better than continuing one that has grown huge.
- Not setting a spending limit on API accounts before running an automation at scale.
How to act on it
If you use AI through a subscription, keep prompts focused and start new conversations for new tasks. If you use an API or an automation tool, estimate tokens before you scale: run ten typical jobs, note the input and output tokens from the usage dashboard, and multiply by your monthly volume. Set a hard spending cap and a usage alert in the provider’s billing settings.
Trim what you send. Pass only the relevant rules and source material, not every document you own; a system prompt with your standing rules is more efficient than repeating them in full. If you need answers from a large library of documents, retrieval-augmented generation sends only the relevant passages. Working out where AI is worth its running cost in your marketing is part of the digital marketing strategy and consulting I offer.
