Temperature is a setting on an AI language model that controls how predictable or varied its output is. A low temperature makes the model stick to its most likely wording, so the same request gives very similar answers; a higher temperature lets it pick less likely words, so answers vary more.
How temperature works
A large language model writes one small piece of text at a time. At each step it calculates a probability for every possible next piece. Temperature adjusts how those probabilities are used. At a very low setting the model almost always takes the top choice. As the setting rises, the probabilities are flattened, so second and third choices get picked more often, and the text becomes more varied and sometimes stranger.
A simple example. Asked to finish “Our bakery in Bristol is known for its…”, a low-temperature model will keep saying “sourdough” or “fresh bread”. A higher setting might offer “cardamom buns”, “Sunday queues” or something odd that does not fit at all.
The scale differs by provider. Many APIs accept values from 0 to 1, some go up to 2, and defaults usually sit somewhere in the middle. There is often a related setting called top-p, which limits the model to the most likely options that together make up a chosen share of the probability. Providers generally advise adjusting one of the two, not both.
Not every model honours the setting. Some reasoning models fix it or ignore it, so check the provider’s documentation before building a process around a particular value.
Why it matters
If you use AI inside a repeatable process, temperature decides whether that process behaves consistently. Tasks with one right answer, such as sorting enquiries into categories, pulling fields out of a form, or rewriting meta descriptions to a strict length, benefit from a low setting. Tasks where you want a range of options, such as brainstorming ad headlines or naming a campaign, benefit from a higher one.
It also explains why two people get different answers to the same question. That is expected behaviour, not a fault. It is one reason to treat a single AI answer as a draft and to test a prompt several times before trusting it.
Common mistakes
- Believing low temperature means accurate. A model at 0 can state a false fact with complete consistency. Temperature controls variety, not truth, and does not prevent hallucination.
- Assuming 0 is perfectly repeatable. Outputs at the lowest setting are very similar but not always identical, and they change when the provider updates the model.
- Turning it up for “better” creative copy. High settings produce more surprising text, which often means more rambling and more errors, not better ideas.
- Tweaking temperature instead of fixing the prompt. Dull or off-brand output is usually a briefing problem.
- Changing temperature and top-p together, which makes results hard to interpret.
How to act on it
If your tool exposes the setting, match it to the job. For classification, extraction, summaries and anything feeding a spreadsheet or CRM, start low. For ideas, start in the middle and generate several options rather than pushing the setting high. Write the chosen value down with the prompt so results can be reproduced.
Whatever the setting, keep a person reviewing anything customers will see. If you are building AI steps into your marketing processes, deciding which tasks need predictable output and which need variety is part of the marketing strategy and consulting work I do, alongside prompt engineering and testing.
