Gemini's Temperature Range: Why It Goes Above 1.0
7 min read · updated August 11, 2026
The Gemini API documents temperature as accepting values from 0.0 to 2.0. That upper half exists because of what the parameter divides into, and knowing the mechanism tells you why 2.0 is a sensible ceiling and why the range is not comparable across providers.
The documented range
temperature is a field of generationConfig, documented in Google’s generateContent API reference with a range of 0.0 to 2.0 for the current Gemini models, and a default that is set per model rather than globally. The default for the recent Gemini models is 1.0.
models.get response and the model page both carry the current values; read them for the exact model id you call rather than assuming.Values outside the range are rejected with an INVALID_ARGUMENT error rather than clamped, so a configuration bug that sets 3.0 fails loudly.
What temperature divides
The model’s final layer produces a logit for every token in the vocabulary — an unnormalised score. Those are turned into probabilities by a softmax, and temperature is the divisor applied to the logits first:
p_i = exp(z_i / T) / sum_j exp(z_j / T) where z_i is the logit for token i and T is the temperature.
Read the arithmetic and the whole behaviour falls out:
- T = 1.0 leaves the logits untouched. The distribution is exactly what the model produced.
- T < 1.0 magnifies the differences between logits. The gap between the best token and the second-best widens, so the top token takes more of the probability mass and output becomes more predictable.
- T → 0 pushes the top token toward probability 1. This is greedy decoding, and it is why
temperature: 0is near-deterministic rather than exactly deterministic — floating-point non-determinism in batched inference can still flip a near-tie. - T > 1.0 shrinks the differences. Low-scoring tokens gain mass relative to high-scoring ones and the distribution flattens toward uniform.
What happens above 1.0
The interesting question is what the extra range buys, and the honest answer is: access to the tail of the distribution, with all the consequences of that.
At 1.0 you sample from the model’s beliefs as they are. Above 1.0 you are deliberately over-weighting tokens the model thought were unlikely. Modest values — 1.1 to 1.3 — produce genuinely more varied phrasing and are useful for brainstorming, naming, or generating a diverse set of candidates. Values approaching 2.0 flatten the distribution far enough that grammar and coherence degrade, and in long outputs the failure is not subtle: repetition loops, invented words, and language switching mid-sentence.
The upper bound is therefore a practical one rather than a mathematical one. Nothing in the softmax breaks at 2.0; it is roughly where output stops being useful for any purpose, so there is no reason to accept more.
Two interactions worth knowing. Temperature is applied before topP and topK truncate the distribution, so a tight topP can mask a high temperature almost entirely — the flattened tail gets cut off before sampling. And on a reasoning model the temperature also applies to the thinking tokens, so raising it makes the model’s internal chain less reliable, not just its prose more varied.
Under constrained decoding the parameter loses most of its meaning. When a response schema masks the sampler down to a handful of legal tokens — an enum with four members, say — flattening or sharpening the distribution over an already-tiny candidate set changes very little. Raising temperature to get “more creative” classification does nothing useful; it only makes the occasional near-tie flip. If a constrained output is unstable, the fix is the schema or the prompt, not the temperature.
Why temperature 0 is not reproducibility
Setting temperature: 0 is the standard move for tests and fixtures, and it is worth being precise about what it buys, because the gap between “greedy” and “reproducible” is where flaky test suites come from.
Greedy decoding removes the sampler’s randomness: at every step the highest-scoring token is taken. What it cannot remove is variation in the scores themselves. Inference runs in batches on floating-point hardware, and the arithmetic is not associative — the same request batched with different neighbours can produce logits that differ in the last bits. Where two tokens are nearly tied, that is enough to pick a different one, and once one token differs everything after it can diverge. So the output is stable most of the time and occasionally is not, which is worse for a test suite than something that is reliably variable.
There is a second source of drift that has nothing to do with arithmetic: the model behind an unpinned alias can change underneath you. A fixture recorded against gemini-2.5-flash is a fixture recorded against whatever that alias pointed at on the day.
What to do instead: assert on properties rather than on exact strings — the JSON parses, the enum value is one of four, the number is within tolerance — and pin the dated model version for anything you compare byte for byte. Where the API offers a seed in generationConfig, setting it alongside a temperature makes repeated identical requests much more likely to agree, but it is a best-effort guarantee across serving changes rather than a promise.
Why providers document different ranges
The ranges are conventions about what a provider is willing to serve, not statements about different mathematics. Every one of these models divides logits by T in the same softmax.
- Google documents 0.0 to 2.0 for the Gemini API, in the generateContent reference.
- OpenAI also documents 0 to 2 for
temperaturein the Chat Completions API reference, with a default of 1. - Anthropic documents 0.0 to 1.0 in the Messages API reference, a narrower window over the same mechanism.
The consequence for anyone calling more than one provider is the one that actually costs money: temperature is not portable. A value of 1.0 is the neutral midpoint on Gemini and the documented maximum on Anthropic. Copying a number across providers changes what it means, and a router that passes temperature through unchanged is passing through a value with a different meaning at the other end. Tune per model, and where you must map, map by intent — “deterministic”, “default”, “creative” — rather than by number. Anthropic’s own range is covered in Claude’s temperature parameter range.
Choosing a value
- 0.0 for extraction, classification, structured output and anything you will compare against a fixture. Not bit-exact reproducibility, but as close as the API offers.
- 0.2 to 0.5 for factual answering and code, where you want the model’s best guess with a little room.
- 1.0 for conversation and general writing. This is the model as trained and it is the default for a reason.
- 1.1 to 1.4 for generating variety on purpose — several options for a human to choose from. Pair it with candidateCount rather than issuing separate requests.
- Change one knob. Google’s guidance, and the general practice, is to adjust temperature or
topPbut not both — two overlapping truncations of the same distribution make the effect of either impossible to reason about.