Tool Calling

A model on its own can only write text back. Tool calling gives it a second option: a structured request naming a specific action and the arguments to run it with. Your code is the one that actually carries that request out. The model never touches a database, calls an API, or runs code itself. It only decides what should happen next, and describes it in a shape your code can act on directly.

The request, execute, respond loop

Every tool call follows the same round trip:

  1. You describe a tool — a name, a plain-language description of what it does, and a schema for what arguments it needs.
  2. The model reads that description and decides, on its own, whether the current task calls for it.
  3. If it does, the model doesn't run anything itself — it returns a structured call naming the tool and filling in arguments that match your schema.
  4. Your code reads that request and actually executes it: hits the real API, queries the real database, runs the real function.
  5. Your code sends the result back to the model as a new message.
  6. The model reads the result and continues — answering the person, or calling another tool if the task isn't finished.

Anthropic's Claude API shows this shape directly:


            // 1. You define the tool
            {
              "name": "get_weather",
              "description": "Get the current weather for a location.",
              "input_schema": {
                "type": "object",
                "properties": { "location": { "type": "string" } },
                "required": ["location"]
              }
            }

            // 2. The model's response, when it decides to call it
            {
              "type": "tool_use",
              "id": "toolu_01a2b3",
              "name": "get_weather",
              "input": { "location": "San Francisco, CA" }
            }

            // 3. You run the real lookup, then send the result back,
            //    matched to the call by its id
            {
              "type": "tool_result",
              "tool_use_id": "toolu_01a2b3",
              "content": "15°C, partly cloudy"
            }
            

Those are Claude's key field names. A real request also needs a couple of bookkeeping fields, like the ones shown above, so the result gets matched back to the call that produced it. OpenAI's API does the same round trip under different labels — and isn't even consistent about them across its own two current APIs. The newer Responses API returns a function_call item with a JSON-string arguments field. The older Chat Completions API nests the call inside a tool_calls array instead. Either way, the shape is what matters, not the exact field names.

Why not just ask for this in plain text?

You could prompt a model to write something like {"tool": "get_weather", "location": "San Francisco"} as plain text and parse it yourself. It's technically possible, but fragile — free text can drift from the format you expect, wrap the answer in commentary, or get a field name slightly wrong. Tool calling constrains the model's output to a schema you control, so your code can trust the shape of what comes back instead of guessing at it.

Two kinds of tools

Tools split into two kinds, and the kind decides who runs the code. A tool you define yourself — a function that queries your own database, say — runs inside your own application. A tool the vendor built and runs on its own infrastructure, like a web search, runs there instead. You just receive the result, without writing any execution code at all. Either way, the request-and-response shape the model works with stays the same.

Tool calling or function calling?

The two terms are used interchangeably in practice, and Anthropic's own documentation describes tool use as "also called function calling." This page uses "tool calling" throughout. Nothing changes if the next doc or article you read calls it "function calling" instead — see function calling for when the more specific term actually matters.

When a model needs a tool

Reach for tool calling whenever a model needs something outside its own text-generation ability. That covers current information its training data doesn't have, a real action to take, or a calculation it shouldn't be trusted to do in its head. Skip it when the model can already answer correctly from the conversation alone. Every tool you define adds to what the model has to read before answering, and a tool it doesn't need is one more thing that can get called by mistake. An MCP server relies on this exact mechanism once a model decides to call one of its tools.

In this guide
  1. The request, execute, respond loop
  2. Why not just ask for this in plain text?
  3. Two kinds of tools
  4. Tool calling or function calling?
  5. When a model needs a tool
  6. FAQ

FAQ

What happens if the model gets the arguments wrong or leaves one out?

It depends on how capable the model is and how ambiguous the request was. A model that's missing a required argument often asks the person for it rather than guessing. A vaguer request, or a less capable model, is more likely to fill in a plausible-looking value on its own. That's why validating arguments in your own code before executing them matters, rather than trusting them blindly.