Sam Austin on October 10, 2026

Function Calling and Tool Use: How LLMs Take Actions

Function Calling and Tool Use: How LLMs Take Actions
Contents

Figure 1: Your code sits between the model's request and the real action

Ask a language model for the weather and it can only guess from training data. Give it a weather function and it can fetch the real answer. That ability, usually called function calling or tool use, is what turns a text generator into something that can query databases, call APIs, and take actions. It's also widely misunderstood, and the misunderstanding causes both bugs and security holes.

The key fact is this: the model never runs anything. It outputs a structured request, and your code decides whether and how to act on it. For a full agent built on this loop from scratch, see the from-scratch agent guide earlier in this series; for the standard packaging layer on top, see the MCP explainer.

The Core Mechanism

Every provider implements the same four-step round trip, with different field names:

  1. You describe your tools to the model: a name, a plain-language description, and a JSON Schema for the arguments.
  2. The model decides whether a tool would help. If so, instead of final text it emits a structured request, such as "call get_weather with {"city": "Paris"}."
  3. Your code executes the function (or refuses to) and sends the result back.
  4. The model reads the result and either answers or requests another tool.

The model never executes functions; it only outputs which function to call and with what arguments, and your code parses the arguments and actually runs the function. Everything else in this article follows from that. Because your code sits between the model's request and the real action, that's where validation, authorization, and approval belong.

Anatomy of a Tool Definition

A tool definition has three parts, and the model sees only these:

  • Name: short, specific, verb-first, such as search_orders rather than orders.
  • Description: what the tool does, when to use it, when not to, and what it returns. This is the most influential text you'll write.
  • Parameters: a JSON Schema listing argument names, types, descriptions, which are required, and constraints like enums and ranges.
weather_tool = {
    "name": "get_weather",
    "description": (
        "Get current weather for a city. Use when the user asks about "
        "present conditions. Returns temperature in Celsius and a short summary."
    ),
    "input_schema": {
        "type": "object",
        "properties": {
            "city": {"type": "string", "description": "City name, e.g. 'Paris'"},
            "units": {"type": "string", "enum": ["celsius", "fahrenheit"]},
        },
        "required": ["city"],
    },
}

Tight schemas prevent whole classes of errors. Use enums for fixed choices, describe formats and units, and mark only truly necessary fields as required.

One Round Trip in Code

Here's the full cycle with Anthropic's API, where tool requests arrive as tool_use content blocks with already-parsed arguments:

import anthropic

client = anthropic.Anthropic()
MODEL = "claude-sonnet-5-5"        # check current model names

def get_weather(city, units="celsius"):
    return {"city": city, "temp": 18, "summary": "light rain"}   # stub

messages = [{"role": "user", "content": "Is it raining in Paris?"}]

resp = client.messages.create(model=MODEL, max_tokens=512,
                              tools=[weather_tool], messages=messages)

if resp.stop_reason == "tool_use":
    messages.append({"role": "assistant", "content": resp.content})
    results = []
    for block in resp.content:
        if block.type == "tool_use":
            output = get_weather(**block.input)
            results.append({"type": "tool_result",
                            "tool_use_id": block.id,
                            "content": str(output)})
    messages.append({"role": "user", "content": results})

    final = client.messages.create(model=MODEL, max_tokens=512,
                                   tools=[weather_tool], messages=messages)
    print(final.content[0].text)

Every tool_use must be answered with a tool_result carrying the matching ID. In production you'd wrap this in the loop from the from-scratch agent guide, with step limits and error handling.

How the Providers Differ

The concepts match, but the plumbing varies, which matters if you support more than one provider:

  • OpenAI originally returned tool_calls with arguments as a JSON string you must parse. For new work, comparisons recommend its Responses API, which represents function calls as output items, supports built-in tools and remote MCP tools alongside custom functions, and exposes tool_choice, allowed-tool constraints, and strict argument validation, in place of the deprecated Assistants API.
  • Anthropic returns tool_use content blocks, and results go back as tool_result blocks inside a user message.
  • Google Gemini returns functionCall parts. Gemini has no tool call ID concept, and requires tool results in the exact same order as the tool use parts.

One more Anthropic-specific rule: when the model makes parallel calls, all of the results must be answered together in a single response. A robust codebase puts a thin adapter between your agent logic and each provider, as the from-scratch guide suggests.

Steering Behavior: Tool Choice

By default the model decides whether to call a tool. You can override that:

  • Auto: the model decides (the default).
  • Required or any: it must call some tool, useful when you always want structured output through a tool.
  • A specific tool: force one named tool.
  • None: no tool calls, text only.

The idioms differ by provider, with OpenAI using tool_choice values like "required", Anthropic using {"type": "any"}, and Gemini using a tool_config mode of ANY. Notably, the semantics of "none" have not been identical everywhere, so test that the behavior you expect actually holds on your provider.

Parallel Tool Calls

Models can request several independent tools in one response, such as the weather in three cities. Running them concurrently cuts latency, with 2026 guides claiming improvements of several-fold in practice — though that depends on your workload.

You may want to disable parallelism when calls depend on each other, or when ordering matters for side effects. Anthropic exposes a disable_parallel_tool_use option, OpenAI exposes parallel_tool_calls, and Gemini has no way to disable parallel tool calls at present. One source also reports that OpenAI's strict schema mode does not combine with parallel calls, so you must set parallel_tool_calls to false when you need strict guarantees. Check current docs, since these interactions change.

Schemas, Strict Mode, and Structured Outputs

Without schema enforcement, models produce arguments on a best-effort basis — occasional missing fields, wrong types, or invented values. Two defenses work together:

  • Strict or constrained decoding, where the provider guarantees output matches your schema. Support and syntax vary by provider and model snapshot, so confirm in the current docs.
  • Your own validation, always. Validate every argument against the schema and against business rules before executing, whatever the provider promises.

It also helps to separate two tasks that often get conflated. Structured outputs constrain the model's reply format, such as returning an extracted record or a classification, and need no execution loop. Function calling is for invoking actions or fetching live data. If you only need JSON in a particular shape, structured output is simpler and cheaper.

Designing Tools the Model Can Use Well

  • Keep the toolset small and distinct. Overlapping tools like search_docs, find_documents, and lookup_documents confuse selection. Every tool definition also consumes context tokens, so large toolsets cost money and degrade accuracy. Some providers and frameworks now offer dynamic tool discovery to load only relevant tools, which is worth checking if you have dozens.
  • Write descriptions like documentation. State when to use the tool, when not to, and what comes back.
  • Return compact, informative results. Include identifiers the model will need for follow-up calls, trim noise, and cap length. A 50,000-character JSON blob wastes context and distracts the model.
  • Make errors actionable. "Invalid date format; expected YYYY-MM-DD" lets the model fix its call, while "error 500" doesn't.
  • Prefer coarse, task-shaped tools over raw access. get_customer_orders(customer_id) is safer and easier to use than run_sql(query).
  • Design for idempotency. Retries happen. Use idempotency keys for anything that writes or charges.

Security: The Model Is an Untrusted Client

Treat tool calls like requests from an untrusted caller, because functionally that's what they are. The model's arguments can be shaped by user input, retrieved documents, web pages, and earlier tool outputs, any of which may contain adversarial text.

  • Validate arguments against schema and business rules, including path sandboxing, allowlists, and numeric limits.
  • Enforce authorization outside the model. Check that the user is permitted to do what the tool does. Never rely on the prompt to enforce permissions.
  • Apply least privilege. Recommended mitigations include strict JSON Schema validation, structured outputs, read-only tools by default, logging every call, and a regression eval on each model snapshot.
  • Require human approval for irreversible actions, such as payments, deletions, and outbound messages.
  • Treat tool results as data, not instructions. A web page or email returned by a tool can contain text telling the model to do something harmful — indirect prompt injection, covered in the red-teaming angle of the frameworks discussion.
  • Keep secrets out of the model's context and out of tool results.
  • Log every call, with arguments and results, for audit and debugging.

Common Failure Modes

Models can call a tool that doesn't exist, pass wrong or missing arguments, choose the tool, ignore results, repeat the same failing call, or loop until budgets run out. Streaming adds another wrinkle: arguments arrive in fragments, so accumulate them and parse only when complete. Mitigations include precise descriptions, tool choice constraints, validation with informative errors, step limits, and evaluations that run on every model snapshot, since behavior shifts between versions.

Testing Tool Use

Test the pieces separately:

  • Unit-test tool functions as ordinary code.
  • Contract-test schemas to ensure the definitions match implementations.
  • Test your loop with a scripted fake model, so approval, error handling, and limits are verified deterministically.
  • Evaluate with the real model on a set of tasks, measuring tool selection accuracy, argument accuracy, task success, steps, and cost, with repeated runs because outputs vary.

Beyond Basic Function Calling

MCP (Model Context Protocol) is often confused with function calling, but they sit at different layers. MCP uses function calling under the hood: the model calls tools through its own API, and the MCP client routes the call to the right server. Function calling is the invocation mechanism, and MCP is a standard way to package and share tools across clients — the subject of the full MCP guide in this series.

Server-side or built-in tools, such as web search and code execution offered by some providers, run on the provider's infrastructure, so you skip implementing them but also cede some control.

Computer-use style tools let a model operate a browser or desktop through screenshots and clicks, which makes sandboxing and approval even more important.

Common Pitfalls

  • Assuming the model executes tools, and so skipping validation and authorization.
  • Vague tool descriptions, then blaming the model for choosing badly.
  • Exposing broad capabilities like arbitrary SQL or shell when a narrow tool would do.
  • Trusting provider schema guarantees without validating arguments yourself.
  • Forgetting to answer every tool call, which breaks the conversation on the next request.
  • Returning huge or unstructured tool results.
  • Skipping per-snapshot regression evals, then finding that a model update changed behavior.
CoverBookDescriptionGet it
Cover of “AI Agents in Action” AI Agents in Action - hands-on builds where every agent action rides this function calling loop. View on Amazon
Cover of “Build a Large Language Model (From Scratch)” Build a Large Language Model (From Scratch) - the model side of the wire: how tool requests get generated in the first place. View on Amazon
Cover of “Designing Agentic AI Systems” Designing Agentic AI Systems - architecture decisions for tool boundaries, permissions, and approval flows. View on Amazon

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What is function calling in simple terms?

The model never runs anything. You describe your tools (name, description, JSON Schema arguments); if a tool would help, the model emits a structured request like "call get_weather with {"city": "Paris"}"; your code validates, executes or refuses, and sends the result back; the model then answers or requests another tool. Because your code sits between the request and the real action, that's where validation, authorization, and approval belong.

How do OpenAI, Anthropic, and Google Gemini differ?

Concepts match, plumbing varies. OpenAI originally returned tool_calls with JSON-string arguments (its Responses API is recommended for new work); Anthropic returns tool_use content blocks with already-parsed arguments, answered with tool_result blocks in a user message, and all parallel-call results must go back together in one response; Gemini returns functionCall parts with no tool-call ID concept, so results must come back in the exact same order as the tool-use parts. A thin adapter between your agent logic and each provider absorbs the differences.

When should I use structured outputs instead of function calling?

Structured outputs constrain the model's reply format — an extracted record, a classification — and need no execution loop. Function calling is for invoking actions or fetching live data. If you only need JSON in a particular shape, structured output is simpler and cheaper. Both benefit from strict or constrained decoding plus your own validation.

How should I secure tool-using LLM applications?

Treat tool calls like requests from an untrusted caller — the model's arguments can be shaped by user input, retrieved documents, and earlier tool outputs, any of which may contain adversarial text. Validate every argument against schema and business rules, enforce authorization outside the model, expose only read-only tools by default, require human approval for irreversible actions, treat tool results as data rather than instructions (indirect prompt injection), keep secrets out of the model's context, and log every call.

Wrapping This Up

Function calling is a simple protocol with big consequences: you describe tools, the model requests them in structured form, your code runs them, and the results flow back into the conversation. The model supplies judgment about what to call, while your code retains control over whether and how, which is exactly where validation, permissions, and human approval belong.

How do you get reliable tool use? Keep toolsets small and well described, validate everything the model sends, return concise and actionable results, and test with both fake and real models. The same pattern powers everything from a weather lookup to a multi-step agent, and getting the basics right makes the advanced cases much easier to trust.

What are You Looking For?

esc