Reasoning Models
A reasoning model works through a problem in explicit intermediate steps before giving a final answer, instead of predicting an answer directly. Ask a standard model a multi-step math problem and it jumps straight to a result; ask a reasoning model the same problem and it works through the steps first — closer to a person using scratch paper before writing down the final answer.
This isn't just a longer prompt. The model itself was trained — often with reinforcement learning — to generate this chain of intermediate reasoning and get better at it over training, learning which reasoning steps actually lead to a correct answer rather than one that merely looks plausible.
Where this actually helps
Multi-step math, planning, and non-trivial debugging are the cases this genuinely helps with: problems where getting the right answer depends on getting several steps right in sequence, not just recognizing a familiar pattern. Asked "why is this function returning the wrong value," a reasoning model can trace through the logic step by step instead of guessing at a fix that merely looks reasonable.
The cost is real
Every intermediate reasoning step is itself made of tokens, and a reasoning model reads and writes far more of them than a direct-answer model would for the same question — more cost and more time, on every single request. Some reasoning models also keep extending their reasoning past the point where extra steps stop adding real accuracy, since more visible reasoning tends to look at least a little more thorough, whether or not it actually is.
A reasoning model isn't simply "more accurate"
It trades speed and cost for a different way of arriving at an answer, and that trade genuinely pays off on the multi-step problems above. It does not make the model less likely to be confidently wrong elsewhere. The hallucination page's own findings are the sharpest version of this: OpenAI's system card for two of its own reasoning models measured them hallucinating more often than the plainer model they replaced, not less.
When to reach for one
Use a reasoning model when the task genuinely has multiple dependent steps to get right — real planning, math, tracing a subtle logic error. Skip it for a simple lookup, a classification task, or anything a direct answer already handles reliably; the extra reasoning adds cost and latency — the time it takes to get a response back — without buying anything back.
In this guide
FAQ
Does a reasoning model always produce a better answer than a standard model?
No. It's more reliable on problems that genuinely need multiple dependent steps done correctly in sequence, but it isn't a general accuracy upgrade — see the hallucination page's own finding that two reasoning models measured worse, not better, than the plainer model they replaced.
Can you see a reasoning model's intermediate steps?
It depends on the provider. Some show the reasoning trace alongside the final answer; others generate it internally and return only the final answer, treating the reasoning itself as something the model uses but the application doesn't need to display.