Model Behavior Questions (2026)

Covers model parameters and behavior. See also all interview topics. These assume you already know the concept — if a section here is unfamiliar, read its linked concept page first; the questions test judgment on top of the concept, not the concept itself.

In this guide
  1. Temperature
  2. Hallucination
  3. Reasoning Models

Temperature

Full concept page →

You want to generate five candidate solutions and have a separate check pick the best one. What temperature fits, and why?

Higher than you'd use for a single answer — the whole point is meaningful variety between the five candidates. If temperature were near 0, most of the five would come back nearly identical, defeating the purpose of generating multiple options to choose from.

A teammate says temperature 0 guarantees identical output for identical input. Why might they still see different results?

Floating-point math on real hardware isn't perfectly reproducible across runs, especially under parallel or batched execution. Some model architectures also route a request through different internal paths depending on system load, which can shift output even when nothing about the request changed. Temperature 0 makes output close to deterministic, not guaranteed to be identical.

Why might you pick temperature 0.3 instead of 0 for a classification task?

To leave room for genuinely close calls. At strict 0, the model always picks its single top-scoring label even when a second label was nearly as likely — a small amount of temperature lets borderline cases surface instead of being silently forced to whichever label happened to score highest.

If you're already using a low top-p to restrict the vocabulary, does raising temperature still do much?

Less than it would without top-p in play. Top-p has already cut the candidate pool down to a small set of likely tokens; raising temperature can only reshuffle probability within whatever's left in that smaller set, so its effect is more limited than when it's working against the model's full vocabulary.

You're porting a prompt that used temperature 1.7 from one provider's API to another. What should you check first?

Whether the new provider's range even goes that high — ranges and defaults aren't standardized (Anthropic's Messages API tops out at 1; OpenAI's goes to 2). The same number can mean a very different level of randomness on a different provider, so check the current docs rather than porting the value directly.

You're A/B testing two prompts for a creative-writing feature. Should temperature be fixed or left to vary between runs?

Fixed, and the same value in both variants. If temperature is left to vary, differences between the two prompts' results could just be random variation rather than a real effect of the prompt change — a fair comparison needs everything held constant except the one thing being tested.

A chatbot's answers feel repetitive and generic. Would raising temperature fix that?

Only if the repetition is coming from the sampling itself. Often it isn't — a weak or overly rigid system prompt produces the same kind of flat answer regardless of temperature, and raising it just makes that same weak answer more erratically worded instead of actually better. Check the prompt before reaching for the temperature setting.

A model gives a confidently wrong answer. Would lowering temperature to make the output more focused fix it?

No, and it can make things look worse in a specific way: if the top-scoring token happens to be the wrong one, a lower temperature just returns that same wrong answer more consistently. Temperature changes how sharply the model commits to its top candidate — it has no way to tell whether that candidate is true, so it can't fix a wrong answer, only make a wrong answer more repeatable.

Hallucination

Full concept page →

Your team upgrades to a newer, more capable reasoning model, expecting fewer hallucinations. Is that a safe assumption?

No — OpenAI's own system card for o3 and o4-mini measured both hallucinating more often than the older o1 they replaced, and states that the cause isn't well understood. Newer models are often better at this, but not dependably, and a stronger model's fluency can make a wrong answer harder to catch, not easier.

A prompt says "don't make anything up." The model still fabricates a source. Why didn't the instruction work?

A general instruction to be truthful has nothing to check itself against — the model isn't comparing its answer to a source, because it doesn't have one. A specific instruction that names an acceptable alternative, like "say so if the documents don't cover this," works far better, because it gives the model an actual output to produce instead of a guess.

A model says a library's latest version is 3.2, but 4.0 shipped last month. Is that a hallucination?

No — that's a knowledge cutoff problem, not a fabrication. The model is accurately reporting what was true when its training data was collected; nothing was invented. The fix is different too: give it current information or a way to look it up, not the fixes that target genuine invention.

Your RAG system retrieves the exact right passage, but the model still gives a wrong answer that contradicts it. Doesn't grounding the answer in real text prevent this?

It reduces hallucination, it doesn't eliminate it. A model can still answer from what it already believes and ignore the material it was just handed — retrieval controls what the model is given, not what it actually uses. That's a faithfulness failure, and the fix is different from a retrieval fix: a stricter instruction to answer only from the given text, not a change to the retrieval step.

A classification prompt lists five allowed labels. Does that guarantee the model can't invent a sixth?

Not by itself. Listing allowed labels in the prompt narrows what's likely, but the model can still generate something outside that list. A real guarantee needs constrained or schema-enforced output, where the response is mechanically forced to match one of the declared options — that's a structural limit, not a request the model is merely likely to honor.

Reasoning Models

Full concept page →

A reasoning model takes noticeably longer and costs more per request than your current model. Is that extra reasoning always buying you something?

Not always. Reasoning models can keep extending their intermediate steps past the point where the extra steps stop adding real accuracy, since a longer visible reasoning trace tends to look more thorough whether or not it actually is. On a task that didn't need multi-step reasoning in the first place, that cost buys nothing.

You need a model to classify incoming support tickets into one of five categories, fast and cheap. Should you reach for a reasoning model?

Probably not. Classification is closer to pattern recognition than multi-step logic — there's no chain of dependent steps for extended reasoning to help with. A reasoning model's advantage shows up on problems where getting the right answer depends on getting several steps right in sequence, which this task doesn't have.

A reasoning model shows its intermediate reasoning steps, and they look thorough and logical. Does that mean the final answer is reliable?

Not necessarily. A plausible-looking chain of steps and a correct final answer are different things — the reasoning trace can look methodical while still arriving somewhere wrong, the same way a confident, fluent hallucination can look correct. Treat a visible reasoning trace as insight into how the model got there, not as proof that where it got to is right.