Models & LLMsGuides & Tutorials

LLM Temperature Settings Explained: Top-p, Top-k, and Penalties

Stop guessing at API settings. This guide shows what each sampling knob really does - temperature, top-p, top-k, penalties, max tokens - and gives you the exact ranges that work for coding, chat, and creative writing.

Toolbit AI - Team
11 min read
LLM Temperature Settings Explained: Top-p, Top-k, and Penalties

Every API dashboard and half the AI tools you open show the same little sliders: LLM temperature, top-p, sometimes top-k, penalties. Most people copy the numbers from some blog post and hope for the best. Here is the real answer up front: temperature, top-p, and top-k control HOW the model picks its next word - not WHAT it knows. Low settings make it pick the most likely word. High settings let surprises through. Penalties punish repetition, max tokens caps the length of the reply, and on the newest Claude models these knobs are gone for good - send anything but the backwards-compatible values and the API answers with a 400 error. Once you see the mechanism behind each slider, picking settings stops being guesswork.

In short:

  • Temperature reshapes the probability list of possible next words. Low means sharp and focused, high means flat and varied.
  • Top-p (nucleus sampling) keeps the smallest set of words whose combined probability reaches p. It is an adaptive cutoff.
  • Top-k is a hard cap: only the K most likely words stay in the race.
  • Frequency and presence penalties discourage repetition, and they work word by word.
  • Max tokens caps reply length. It is not a randomness setting at all.
  • Temperature 0 is not fully deterministic - Anthropic says so in its own docs - and Claude models released after Opus 4.6 reject these settings with 400 errors, with only a couple of backwards-compatible exceptions.

What does temperature actually do to the model's output?

How low and high temperature reshape the probability of the next token

Temperature in AI models controls how random the model's next word choice is: low values make it pick the most likely word almost every time, high values give unlikely words real chances.

Picture the model mid-sentence. Before it writes the next word, it scores every word it knows. Those raw scores are called logits. A function called softmax turns the logits into probabilities that all add up to 100 percent. Temperature steps in and reshapes that probability list before the model rolls the dice.

Low temperature makes the list sharp: the favorite word gets almost all the probability, and everything else fades. High temperature flattens the list: unlikely words get real chances. In the limit, as temperature approaches 0, the model just picks the top-scoring word every time - that is called greedy decoding. The standard formulation in ML literature is that the logits get divided by the temperature value before softmax, and the HuggingFace how-to-generate walkthrough shows the same effect from the outside: lower the temperature of the softmax and the distribution gets sharper.

The official docs agree on the effect, but not on the numbers. OpenAI's reference says higher values like 0.8 make output "more random," while lower values like 0.2 make it "more focused and deterministic" - its range is 0 to 2. Anthropic says to use temperature "closer to 0.0 for analytical / multiple choice, and closer to 1.0 for creative and generative tasks," with a range of 0 to 1. Gemini spans 0 to 2 like OpenAI. Same idea everywhere, but the same number does not mean the same behavior across providers, because the ranges differ.

What do top-p and top-k change that temperature doesn't?

Five LLM sampling knobs and what each one controls

Temperature reshapes the whole probability list. Top-p and top-k instead decide which words get to compete at all.

Top-p, also called nucleus sampling, works like this: sort all words by probability, biggest first, and keep the smallest group whose combined probability reaches p. In OpenAI's example, top_p 0.1 means only the words making up the top 10 percent of probability mass are considered. Anthropic describes the same thing as a cumulative distribution that gets cut off once it reaches the value you set.

The clever part is that the group size adapts. When the model is unsure and the list is flat, a top-p of 0.9 might keep dozens of words in play. When the next word is obvious and the list is sharp, the same 0.9 might keep one or two. That is exactly why researchers proposed nucleus sampling (Holtzman et al., 2019): a fixed word count can either keep gibberish in play or throw away good options, depending on the shape of the list.

Top-k is that fixed word count. It keeps the K most likely words, redistributes their probability among themselves, and zeroes the rest (Fan et al., 2018). Anthropic's docs describe it as a way to remove "long tail" low-probability responses. Simple and predictable, but blind to whether the list is flat or sharp.

Availability differs by provider. Anthropic and Gemini expose top_k; OpenAI's hosted API does not, with no top_k parameter anywhere in its official reference. Gemini adds a twist: models running on pure nucleus sampling reject topK settings outright.

One more thing: OpenAI explicitly recommends altering temperature OR top_p, not both. Two knobs reshaping the same list at once is hard to reason about.

Top-k vs Top-p: How Cutoffs Reshape Candidate Tokens

What do the penalty settings and max tokens actually control?

Frequency and presence penalties control repetition by pushing down words the model has already written, while max tokens caps the reply length without touching randomness. Temperature and top-p reshape the whole probability list no matter what came before. Penalties look at the history - what the model already wrote - and nudge individual words.

Frequency penalty scales with how many times a word already appeared. Said it five times? It gets pushed down five times as hard as a word said once. Presence penalty is binary: did the word appear at all? If yes, it gets the same nudge whether it showed up once or twenty times, which pushes the model toward new topics. Gemini's reference states this distinction precisely: presencePenalty is "binary on/off and not dependant on the number of times the token is used," while frequencyPenalty is "multiplied by the number of times each token has been seen." Technically, both adjust the logits for each next word, which is why OpenAI describes its logit_bias as added "to the logits generated by the model prior to sampling."

On OpenAI, penalties range from -2.0 to 2.0. Gemini cautions that negative values can force the model into repetition loops - the exact opposite of what you want. And Anthropic's Messages API does not expose penalty parameters at all.

Max tokens is the simplest one, and the most misunderstood. It caps the LENGTH of the reply, not its randomness - it is not a sampling setting. The names differ per API: OpenAI's Chat Completions moved from max_tokens to max_completion_tokens, the Responses API uses max_output_tokens (which counts reasoning tokens too), Gemini calls it maxOutputTokens, and Anthropic keeps max_tokens with a minimum of 0. That 0 is not useless on Anthropic: setting max_tokens to 0 pre-warms the prompt cache without generating anything, a trick worth knowing if you care about caching your prompts without breaking the prefix.

Why is temperature 0 not fully deterministic?

Temperature 0 collapses the model to always picking the top-scoring word, yet the output still is not guaranteed to repeat. Anthropic's reference says it verbatim: "even with temperature of 0.0, the results will not be fully deterministic."

Why? The picked word is only as stable as the scores underneath it. The scores are computed on shifting hardware: community analysis of hosted APIs suggests that batch sizes vary between your requests, and since floating-point math is not associative, the logits can differ in the last decimal places. When two words are near-tied, those tiny shifts flip which one wins. Treat that root-cause explanation as informed community analysis rather than provider-documented fact.

The providers themselves hedge in writing. OpenAI's seed parameter only triggers "a best effort to sample deterministically" - and they tell you to watch the system_fingerprint field to detect backend changes. OpenAI's text-generation guide adds that model output "is non-deterministic" and recommends pinning production apps to specific model snapshots. Gemini uses a randomly generated seed whenever you do not set one, so the dice are quietly rolling even when you thought they were fixed.

The practical takeaway: for reproducibility, pin the model snapshot, set a seed where available, and monitor system_fingerprint. Do not treat temperature 0 as a determinism guarantee, because nobody selling you the API does.

Which settings should you use for coding, extraction, chat, and creative writing?

The short answer: 0 to 0.3 for code, data extraction, and classification; 0.5 to 0.7 for chat and assistants; 0.7 to 1.0 for creative writing and brainstorming. Treat these as community convention, not official guidance - no provider doc says "use 0.2 for JSON extraction." The ranges that work well in practice:

Use caseTemperatureWhy
Code, data extraction, classification0 to 0.3You want the same answer to the same input
Chat and assistants0.5 to 0.7Some variety without losing the thread
Creative writing and brainstorming0.7 to 1.0Long-tail words get real chances

The confirmed anchor points line up with this. Anthropic points analytical and multiple-choice work toward 0.0 and creative work toward 1.0. HuggingFace's walkthrough notes the failure mode at the bottom end: with greedy decoding "the model quickly starts repeating itself," while sampling produces more fluent text on open-ended generation. So do not reflexively slam temperature to 0 for every structured task - slight sampling or a penalty handles repetition better than pure greed. And for creative writing, Gemini's positive penalties "increase vocabulary" by pushing the model away from words it already used.

Two rules of thumb keep this simple. One: change ONE knob. Pick temperature or top-p, not both - that is OpenAI's own recommendation. Two: remember that settings cannot rescue a wrong model choice. If the model you picked does not fit the job, no slider fixes that, so choose the right model first and tune after.

Why did Anthropic remove these settings from its newest Claude models?

Models released after Claude Opus 4.6 do not support setting temperature - the Anthropic Messages API reference spells the deprecation out. A value of 1.0 is still accepted for backwards compatibility, and any other value is rejected with a 400 error. Top-p only accepts values >= 0.99 for backwards compatibility, and any top_k value at all returns a 400. This is the angle most "temperature explained" posts miss entirely.

What does that mean in practice? The provider-side defaults become the only behavior you get. You trade a knob for whatever consistency the provider tuned the model for. If your code sends sampling parameters to Claude, strip them out or guard them by model, because post-Opus-4.6 models will refuse anything but the backwards-compatible values.

There is a broader signal here too. The newest frontier models are shipping with the knobs removed, and it is a fair bet that newer behavior controls - like the effort modes in our guide to using AI effort modes like a pro - take their place. Memorizing slider values has a short shelf life. Understanding the mechanism from earlier in this post does not.

What are the default settings on OpenAI, Claude, and Gemini?

There is no single shared default: temperature spans 0 to 2 on OpenAI and Gemini but only 0 to 1 on Anthropic, OpenAI's hosted API does not expose top-k at all, and some defaults are not printed in the current references. This table pulls together what each official reference actually prints today:

SettingOpenAIAnthropic (pre-deprecation)Gemini
Temperature range0 to 20 to 10 to 2
Temperature defaultNo number printed in the current reference (historically 1)1.0Per model, via Model.temperature from getModel
Top-pYes, nucleus samplingYes, range 0 to 1Yes, per model via Model.top_p
Top-kNot exposed in the hosted APIYes (rejected on post-Opus-4.6 models)Yes, but nucleus-only models reject topK
Penalties-2.0 to 2.0 (frequency and presence)Not exposed in the Messages APIYes, on next-token logprobs
SeedYes, best-effort determinism plus system_fingerprintNone documentedOptional; random if unset
Max tokens namingmax_completion_tokens / max_output_tokensmax_tokens (0 pre-warms the cache)maxOutputTokens

Two cells deserve honesty. OpenAI's current reference marks temperature and top_p as optional without printing numeric defaults - the famous "default 1" comes from older docs, so treat that number as history rather than current print. And Gemini's blanket "default 1.0" is only what the docs examples show; the reference itself says the default varies by model, which is why getModel is the real source of truth. The OpenAI Completions reference and the Gemini GenerationConfig reference are the places to check the current state, since these details do drift.

FAQ

Should I set temperature AND top-p together?

No. OpenAI explicitly recommends altering temperature OR top_p, not both. Both reshape the same probability list, so tuning them jointly is mostly guesswork. Pick one knob and leave the other at its default.

Does OpenAI's API expose a top-k setting?

No. There is no top_k parameter anywhere in the official OpenAI reference - the closest tool is logit_bias, which nudges individual word scores before sampling. Top-k exists on Anthropic, Gemini, and local open models instead.

My code sends temperature or top_k to Claude - will newer models break?

Yes, if you are calling models released after Claude Opus 4.6. Any top_k value returns a 400, temperature only accepts 1.0, and top_p below 0.99 is rejected too. Remove the sampling parameters from those calls or pin your requests to pre-4.6 models.

Share this article

Related articles

Continue exploring similar guides and insights

Featured image for Space Bunny Alpha and MiniMax M3.1-Flash: How to Read a Stealth Coding Model Before It Names Itself
9 min read
5 views

Space Bunny Alpha and MiniMax M3.1-Flash: How to Read a Stealth Coding Model Before It Names Itself

OpenRouter listed free Space Bunny Alpha on Sep 23, 2026; MiniMax announced M3.1-Flash-Preview on MiniMax Code and Token Plan on Sep 27. Keep MiniMax = Space Bunny unconfirmed. Map confirmed listing fields, quote retention carefully, run a short non-sensitive eval while Free lasts (pricing may change), and wait for a named card before a production default.

  • Models & LLMs
  • AI Infrastructure