…
Skip to content
Topics
On this page

Tool Calling (Function Calling) in LLM

Tool calling, also called function calling, is a capability of large language models in which the model returns a structured request to execute a named function instead of an ordinary text response. The application executes that function, returns the result to the model, and the model incorporates the result to continue the task.

  • Tool: A function the application exposes to the model, such as retrieving logs, querying metrics or creating an issue.
  • Tool schema: A name, a description and a JSON Schema that specifies the arguments, their data types and the required fields.
  • Tool call: A structured message from the model that identifies one tool and supplies its arguments as a JSON object.
  • Tool result: The output of the function, returned to the model in a message associated with the original call by an identifier.
  • Execution boundary: The model never executes code itself, because the application determines whether and how each request is executed.
The tool calling round trip between an application and an LLMThe application sends tool schemas and a question to the LLM. The LLM replies with a tool call for get_error_rate. The application validates the arguments and executes the call against a metrics API, then sends the tool result, an error rate of 0.042, back to the LLM. The LLM replies with the final answer. The error rate is illustrative.Applicationvalidate andexecuteLLM1. tool schemas + question2. tool_use: get_error_rate3. tool_result: 0.0424. final answermetrics API
The tool calling round trip between an application and an LLM

For example, a DevOps agent receives the question "Is checkout-api failing more than usual?" and responds with a request to get_error_rate containing the service identifier and a monitoring interval.

Key Characteristics of Tool Calling

  • Schema-driven: Each tool is described with JSON Schema, the same contract that governs structured output.
  • Model-selected: The model determines which tool to invoke, or whether to invoke any tool, from the descriptions and the conversation history.
  • Application-executed: Application code executes the function, so permissions, timeouts and audit logging remain under the developer's control.
  • Multi-turn: A complex task frequently requires several iterations of requests and results before the model produces a final response.
  • Parallel invocation: Some model APIs allow several independent tool requests within a single response.
  • Foundation for agents: Tool calling is the mechanism behind LLM tool use, and it allows an AI agent to operate on external systems.

How Tool Calling Works

  1. Define the tools: Write a name, a precise description and a JSON Schema for the arguments of each available function.
  2. Send the request: Provide the tool definitions to the model together with the system prompt and the user message.
  3. Receive a tool call: The model responds with a tool name, a unique identifier and a JSON object of arguments.
  4. Validate the arguments: Verify the required fields, data types and permitted values before executing anything.
  5. Execute the function: Invoke the tool with a timeout and capture either its output or its error message.
  6. Return the result: Send a tool result message carrying the same identifier, then request the next model response.
  7. Finish: Repeat the cycle until the model returns a text response instead of another tool request.

Example: Tool Calling with a Tool Registry in Python

The program below validates two hard-coded sample tool calls against a JSON Schema, executes the valid request and constructs the tool result messages.

Python
import json

# Tools offered to the model, each with a JSON Schema for its arguments
TOOLS = {
    "get_error_rate": {
        "description": "Return the 5xx error rate of a service over recent minutes.",
        "input_schema": {
            "type": "object",
            "properties": {
                "service": {"type": "string", "enum": ["checkout-api", "search-api"]},
                "minutes": {"type": "integer", "minimum": 1, "maximum": 60},
            },
            "required": ["service", "minutes"],
        },
    },
}
METRICS = {"checkout-api": 0.042, "search-api": 0.001}  # sample data
HANDLERS = {"get_error_rate": lambda service, minutes: {"service": service, "error_rate": METRICS[service]}}
TYPES = {"string": str, "integer": int}

def validate(args, schema):
    errors = [f"missing argument: {n}" for n in schema["required"] if n not in args]
    for name, value in args.items():
        spec = schema["properties"].get(name)
        if spec is None:
            errors.append(f"unknown argument: {name}")
        elif not isinstance(value, TYPES[spec["type"]]):
            errors.append(f"{name} must be {spec['type']}")
        elif "enum" in spec and value not in spec["enum"]:
            errors.append(f"{name} '{value}' not in {spec['enum']}")
        elif "maximum" in spec and not spec["minimum"] <= value <= spec["maximum"]:
            errors.append(f"{name} {value} outside {spec['minimum']}..{spec['maximum']}")
    return errors

def execute(call):
    tool = TOOLS.get(call["name"])
    errors = validate(call["input"], tool["input_schema"]) if tool else ["unknown tool"]
    if errors:
        return {"tool_use_id": call["id"], "is_error": True, "content": "; ".join(errors)}
    result = HANDLERS[call["name"]](**call["input"])
    return {"tool_use_id": call["id"], "is_error": False, "content": json.dumps(result)}

# Sample model tool calls, hard-coded for illustration (no model is called)
good = {"id": "call_1", "name": "get_error_rate", "input": {"service": "checkout-api", "minutes": 15}}
bad = {"id": "call_2", "name": "get_error_rate", "input": {"service": "billing-api", "minutes": 90}}

print("Tools sent to the model:", list(TOOLS))
for call in (good, bad):
    print(f"Model requested {call['name']}({call['input']})")
    print("  tool_result:", execute(call))
Output
Tools sent to the model: ['get_error_rate']
Model requested get_error_rate({'service': 'checkout-api', 'minutes': 15})
  tool_result: {'tool_use_id': 'call_1', 'is_error': False, 'content': '{"service": "checkout-api", "error_rate": 0.042}'}
Model requested get_error_rate({'service': 'billing-api', 'minutes': 90})
  tool_result: {'tool_use_id': 'call_2', 'is_error': True, 'content': "service 'billing-api' not in ['checkout-api', 'search-api']; minutes 90 outside 1..60"}
  • Validation before execution: The second request specifies an unknown service and an interval above the maximum, so the function is never executed.
  • Errors return to the model: An error result with is_error enabled allows the model to correct its arguments on the following turn.
  • Identifiers: Each result carries the identifier of its request, so the model can associate every result with the correct call.

With the Anthropic Python SDK, the same schema is supplied through the tools parameter. This sketch requires an API key and is not executed here.

Python
import anthropic

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY from the environment
response = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    tools=[{"name": "get_error_rate", **TOOLS["get_error_rate"]}],
    messages=[{"role": "user", "content": "Is checkout-api failing more than usual?"}],
)
for block in response.content:
    if block.type == "tool_use":
        print(block.id, block.name, block.input)

Applications of Tool Calling

  • Observability queries: Retrieving error rates, latency percentiles and recent log entries for a production service.
  • Code operations: Executing tests, reading source files and opening a pull request with a proposed correction.
  • Issue management: Creating or updating tickets in an issue tracker with validated, structured fields.
  • Data access: Executing read-only SQL queries against an analytics database.
  • Retrieval: Searching documentation or operational runbooks when the model requires current information.
  • Standard connectors: Exposing tools through a protocol such as MCP, compared in MCP vs API vs function calling.

Advantages

  • Current information: The model can consult the live system state instead of depending entirely on its training data.
  • Controlled actions: Every operation passes through application code that can validate, record and restrict it.
  • Typed integration: JSON arguments correspond directly to existing functions and internal service interfaces.
  • Composability: Individual tool calls combine into iterative loops such as the ReAct agent pattern.

Limitations

  • Incorrect arguments: The model can select an inappropriate tool or supply invalid values, so validation is essential.
  • Security exposure: Tool results can contain malicious instructions, a category of prompt injection.
  • Latency and cost: Every iteration of requests and results adds another model invocation and additional input tokens.
  • Context growth: Lengthy tool outputs consume the context window, which memory in AI agents must manage.
  • Description quality: Ambiguous tool names or descriptions produce incorrect or missing tool calls.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. What does the model return when it decides to use a tool?

Frequently Asked Questions

What is the difference between tool calling and function calling?

They are two names for the same capability. Some providers say function calling and others say tool use or tool calling, but in each case the model returns a named call with JSON arguments for the application to run.

Does the LLM execute the function itself?

No. The model only produces the name of the tool and its arguments. The application validates the request, runs the function and returns the result to the model.

What happens if the model sends invalid arguments?

The application should reject the call before execution and return an error result that explains the problem. The model can then send a corrected call on the next turn.

How many tools can an LLM use at once?

Many APIs accept dozens of tool definitions in one request, but accuracy drops when descriptions overlap. Grouping tools by task and exposing only the relevant set keeps selection reliable.

Is tool calling the same as MCP?

No. Tool calling is the model capability, while MCP is a protocol that standardises how tools are described and served to many applications. An MCP client still presents tools to the model through tool calling.