Temperature and top_p Together in the OpenAI API: What Changing Both Does
9 min read · updated August 11, 2026
OpenAI’s API reference says the same thing under both parameters: we generally recommend altering this or top_p but not both. It does not say why. The reason is that the two are not independent knobs at all — one of them changes the input the other one operates on, so setting both means the second value no longer means what its documentation says it means.
The order the two are applied in
A model’s forward pass ends with a vector of logits, one real number per token in the vocabulary. Nothing in that vector is a probability yet, and nothing about it has been influenced by your sampling parameters — the model does not know what you set. The sampler runs afterwards, and it runs in a fixed order:
logits → raw scores from the model, one per vocab token ÷ temperature → every logit divided by T softmax → scores become a probability distribution top_p filter → keep the smallest prefix summing to ≥ p, drop the rest renormalise → the survivors are rescaled to sum to 1 sample → draw one token
The important line is the third one. top_p is defined over the distribution that comes out of the softmax, and the softmax has already been reshaped by temperature. So temperature is upstream of top_p, and any change to temperature changes how many tokens top_p keeps — without top_p having been touched.
Both default to 1 in the OpenAI API. temperature accepts 0 to 2 and top_p accepts 0 to 1. At the defaults neither does anything: dividing by one leaves the logits alone, and a nucleus of 1.0 includes every token. That is worth knowing because it means the pair has a genuine no-op setting, which is the right thing to send when you want the model’s own distribution.
What top_p actually selects
Nucleus sampling sorts the tokens by probability, walks down the list accumulating, and stops as soon as the running total reaches p. Everything from that point down is discarded and the survivors are renormalised. The size of the surviving set is not fixed. On a confident distribution — the model is nearly certain of the next token — top_p: 0.9 might keep one or two tokens. On a flat distribution, in the middle of an open-ended sentence, the same 0.9 might keep forty.
That variable width is the point of the parameter and the reason it is usually preferred to a fixed top_k — which OpenAI does not expose at all, though most open-weights serving stacks do. A fixed k keeps twenty tokens whether the model is certain or lost; a fixed p keeps as many as it takes to cover the same amount of probability mass, which is closer to what you meant. It is also exactly why top_p cannot be reasoned about independently of temperature: temperature is the knob that decides how confident or how flat the distribution is before top_p looks at it.
One consequence of the definition that surprises people: top_p can never exclude the most likely token, no matter how small you set it. The nucleus is built by taking the top token first and stopping as soon as the running total reaches p, so at top_p: 0.001 the set is the single best token and the sampler is deterministic. The parameter has a floor of behaviour, not a floor of quality, and it cannot be used to make output more varied — only less.
A worked example on five tokens
Take a model whose next-token logits, for the five candidates that matter, are 2.0, 1.0, 0.0, −1.0 and −2.0. Call them A through E. Everything below is the softmax arithmetic, done on those numbers, with top_p: 0.9 held constant throughout.
At temperature: 1 the logits are unchanged and the softmax gives:
token logit exp(logit) probability cumulative
A 2.0 7.389 0.6364 0.6364
B 1.0 2.718 0.2341 0.8705
C 0.0 1.000 0.0861 0.9566 ← crosses 0.9 here
D -1.0 0.368 0.0317 0.9883
E -2.0 0.135 0.0117 1.0000
top_p = 0.9 → nucleus is {A, B, C}, three tokens
renormalised → A 0.665, B 0.245, C 0.090Now lower the temperature to 0.5 and change nothing else. Dividing each logit by 0.5 doubles it, which is what “sharpening the distribution” means concretely:
token logit/T exp() probability cumulative
A 4.0 54.598 0.8647 0.8647
B 2.0 7.389 0.1170 0.9817 ← crosses 0.9 here
C 0.0 1.000 0.0158 0.9975
D -2.0 0.135 0.0021 0.9997
E -4.0 0.018 0.0003 1.0000
top_p = 0.9 → nucleus is {A, B}, two tokens
renormalised → A 0.881, B 0.119The nucleus shrank from three tokens to two. Nobody changed top_p. If you were tuning top_p to control how much of the tail the model may reach for, your control just moved under you.
Push the other way, to temperature: 2, and the effect inverts. Halving the logits gives probabilities of roughly 0.370, 0.225, 0.136, 0.083 and 0.050. The cumulative total after four tokens is 0.814 — still under 0.9 — so the nucleus is all five tokens and top_p: 0.9 excludes nothing whatsoever. The parameter is still in your request, still documented as a truncation, and on this distribution it has become a no-op.
Why the pair confounds itself
Three separate problems come out of that arithmetic, and they are worth separating because only the first is obvious.
- The effects are not additive, so tuning is not searchable. If you lower temperature and lower top_p together and the output gets worse, you cannot attribute it. The second change may have done nothing at all, because the first change already put the whole nucleus inside the new threshold.
- There are dead bands at both ends. At
temperature: 0the sampler takes the argmax and top_p is irrelevant at any value, because the top token alone always reaches any p. Attop_p: 0.01the nucleus is almost always the single most likely token, so temperature is irrelevant at any value. Two different parameter settings produce the same behaviour, and a reader of your config cannot tell which one is doing the work. - The interaction is distribution-dependent, so it is prompt-dependent. The example above shrank from three tokens to two. On a different prompt, at a different position, the same temperature change might take a nucleus from forty tokens to four. There is no fixed conversion between the two settings, which means a pair tuned on one prompt does not transfer to another.
None of this makes a request with both parameters set invalid. The API accepts it and the model answers. The guidance in the reference is about tractability, not legality: two knobs onto one effect is a configuration you cannot reason about, cannot bisect when it goes wrong, and cannot hand to somebody else with an explanation of what it does.
Which one to set
Pick on the basis of what you want to hold constant across prompts.
- Temperature is a shape control. It rescales the whole distribution and never removes a token — every token in the vocabulary keeps a non-zero probability at any T above 0. Use it when you want a continuous creativity dial and you are willing for rare tokens to remain reachable.
- top_p is a truncation control. It hard-excludes the tail so the worst tokens become impossible rather than merely unlikely. Use it when the failure you care about is the occasional wildly wrong token rather than the general register of the output.
Then send the other one at its default explicitly. Writing "temperature": 0.7, "top_p": 1 costs nothing and documents the decision in the request itself, which is a great deal better than the next person finding a bare temperature and wondering whether top_p was omitted deliberately or forgotten. If you need reproducibility rather than conservatism, the parameter for that is seed, and it is orthogonal to both of these.
One thing that does not follow from any of the above: penalties. frequency_penalty and presence_penalty modify the logits before the softmax rather than filtering after it, so they compose with temperature in a different and more predictable way than top_p does.