Lesson 11.2 · 25 min
Function Calling: Giving LLMs Tools to Act on the World
If a language model can only output text, how does it “check the weather”, “query a database” or “book a meeting”?
In short: Function calling (also called tool calling) lets an LLM ask our application to run a function. We describe the available functions with names, descriptions and JSON Schemas; the model replies with a structured request naming a function and its arguments; our code runs it and sends the result back; the model then answers using that result. The model never executes anything itself, which is exactly what makes the pattern safe and controllable.
What is function calling?
Function calling is a feature of modern LLM APIs where, instead of replying with ordinary text, the model can reply with a structured request to call a function that we told it about. The request contains the function's name and the arguments to pass, written as JSON (JavaScript Object Notation, a simple text format for data like {"city": "Paris"}).
Vendors use slightly different names: OpenAI says function calling or tool calling, Anthropic says tool use, Google says function calling. The idea is the same everywhere. In this lesson, tool and function mean the same thing: a piece of our code the model may ask us to run.
Think of it like a head chef and a kitchen The head chef (the LLM) never touches the stove. They shout precise orders: “Two eggs, fried, three minutes.” The kitchen staff (our code) cook and bring the plate back. The chef tastes it and decides what to say to the customer. The orders must follow the menu (the function definitions), and the kitchen can refuse an order that is not on it.
Why we need function calling
An LLM on its own has three hard limits:
- Frozen knowledge: it only knows what was in its training data, up to a cutoff date. It cannot know today's weather or your latest order.
- No access to private data: it has never seen your company's database, calendar or tickets.
- No ability to act: it cannot send an email, create a ticket or run code. It can only write text about doing so.
Before function calling was built into APIs (OpenAI added it in June 2023, and other vendors followed), developers asked the model in the prompt to “reply in this exact format if you need the weather”, then parsed the text with regular expressions. It worked, but the model often drifted from the format and parsing broke. Function calling turned this into a first-class, trained behaviour with a predictable, machine-readable output.
Pause and think: A user asks a plain LLM (no tools) “What is the price of order #4812?” What is the best case and the worst case?
Best case: it says it cannot access order data. Worst case: it invents a plausible price (a hallucination). With a get_order(order_id) tool, the model can ask our code for the real record instead of guessing.
The key insight: the model does not run the function
This is the single most important point of the lesson, and the most common misunderstanding. The LLM never executes your function. It cannot reach your servers. All it does is produce text (structured as a tool call) saying which function it would like called and with what arguments.
Our application receives that request and decides what to do with it. Usually we run the function, but we are in full control. We can:
- Validate the arguments (is
amounta positive number below the limit?). - Check permissions (is this user allowed to see that account?).
- Ask a human to approve before anything irreversible (sending money, deleting data).
- Refuse and send back an error message, which the model can read and react to.
- Log every call for debugging and auditing.
Common mistake Treating the model's arguments as trusted. They are generated text: they can be wrong, malformed, or even shaped by a malicious document the model just read (prompt injection). Always validate tool arguments the same way you would validate input from an untrusted web form.
How does the model know how to produce these requests? Models are fine-tuned on many examples of tool definitions and tool calls. When we send tool definitions through the API, the provider renders them into the model's prompt in a special format, and the model has learned to emit a tool call in a matching format when it judges a tool is needed. The API then parses that output and hands it to us as a clean JSON object.
How function calling works step by step
One function call, end to end
- Define the tools: We describe each function: a
name, a plain-languagedescriptionof what it does and when to use it, and a JSON Schema for its parameters (types, required fields, allowed values). - Send the request: We call the LLM API with the conversation messages plus the tool definitions.
- The model decides: The model reads the question and the tool descriptions. It either answers directly with text, or returns one or more tool calls, each with an
id, aname, andarguments. - Our code executes: We parse the arguments, validate them, and run the real function (an API call, a database query, anything).
- Send the result back: We append the model's tool-call message and a tool-result message (linked by the call
id) to the conversation, and call the API again. - The model answers: Now the model sees the real data and writes the final reply for the user, or asks for another tool if it still needs more.
Most APIs also let us steer whether tools are used with a setting usually called tool_choice: let the model decide (auto), forbid tools (none), require some tool, or force one specific tool. Forcing a tool is a handy trick for reliable data extraction.
A concrete example: get_weather(city)
Here is what the messages look like on the wire. The exact field names differ by vendor, so we show the two most common shapes side by side. Read them as data, not as something to memorise.
Now let us run the whole round trip. To keep it runnable offline, fake_model plays the role of the API, returning exactly the kind of message a real model would. The rest of the code (tool schema, argument parsing, validation, the tool message, the second call) is what you would write in production.
function_calling.py
import json
# 1. The tool definition we send to the model (name, description, JSON Schema).
TOOLS = [{"name": "get_weather",
"description": "Current weather for a city.",
"parameters": {"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]}}]
def get_weather(city): # our real code (fake data here)
data = {"Paris": (18, "cloudy"), "Tokyo": (24, "sunny")}
temp, sky = data.get(city, (None, "unknown"))
return {"city": city, "temp_c": temp, "sky": sky}
def fake_model(messages, tools):
"""Stands in for the LLM API. Turn 1: ask for tools. Turn 2: answer."""
if messages[-1]["role"] == "user":
return {"role": "assistant", "tool_calls": [
{"id": "call_1", "name": "get_weather", "arguments": '{"city": "Paris"}'},
{"id": "call_2", "name": "get_weather", "arguments": '{"city": "Tokyo"}'}]}
w = [json.loads(m["content"]) for m in messages if m["role"] == "tool"]
return {"role": "assistant", "content": " | ".join(
f"{x['city']}: {x['temp_c']} C, {x['sky']}" for x in w)}
messages = [{"role": "user", "content": "Weather in Paris and Tokyo?"}]
reply = fake_model(messages, TOOLS)
messages.append(reply)
for call in reply.get("tool_calls", []): # 2. the APP executes each call
args = json.loads(call["arguments"]) # arguments arrive as a JSON string
missing = [k for k in TOOLS[0]["parameters"]["required"] if k not in args]
result = {"error": f"missing {missing}"} if missing else get_weather(**args)
print("model asked:", call["name"], args, "->", result)
messages.append({"role": "tool", "tool_call_id": call["id"],
"content": json.dumps(result)})
final = fake_model(messages, TOOLS) # 3. model sees results, answers
print("final answer:", final["content"])
print("messages in history:", len(messages))Output:
model asked: get_weather {'city': 'Paris'} -> {'city': 'Paris', 'temp_c': 18, 'sky': 'cloudy'}
model asked: get_weather {'city': 'Tokyo'} -> {'city': 'Tokyo', 'temp_c': 24, 'sky': 'sunny'}
final answer: Paris: 18 C, cloudy | Tokyo: 24 C, sunny
messages in history: 4The conversation loop
A single question may need several rounds: the model asks for a tool, reads the result, then asks for another. So in practice we wrap the API call in a loop:
- Call the model with the messages and tools.
- If the reply contains tool calls, run them, append the results, and go back to step 1.
- If the reply is plain text with no tool calls, that is the final answer: stop.
- As a safety net, stop after a maximum number of rounds.
That loop is the heart of every AI agent. The next lesson studies it in detail. One rule to keep in mind now: the history must stay consistent. Every tool call the model makes needs exactly one matching result message with the same id, in order. APIs usually reject a history where a tool call has no result.
Errors are results too If a function fails, do not crash the loop. Send the error back as the tool result, for example {"error": "city not found: Pariss"}. Models are good at reading an error and retrying with corrected arguments.
Multi-step and parallel function calling
There are two different ways a model can need more than one call:
- Multi-step (sequential): each call depends on the previous result. Example:
find_user(email)returns an id, thenget_orders(user_id)needs that id. These must happen in separate rounds. - Parallel: the calls are independent. Example: weather in Paris and Tokyo. Many models can return several tool calls in one reply, and our code can run them at the same time.
Our Python example above already used parallel calls: one reply asked for both Paris and Tokyo. Parallel calling is a big latency win, but only for independent calls. If the second call needs the first call's output, the model must wait for it.
Pause and think: A model is asked: “Find the customer with email ana@example.com and list her last three orders.” Can both tool calls be made in parallel?
No. get_orders needs the customer id that find_customer returns, so the calls are sequential: two separate rounds. Parallel calling only applies when the calls do not depend on each other.
Relation to structured outputs and JSON mode
Function calling is one member of a family of features that make LLM output machine-readable. It helps to separate three ideas:
Constrained decoding means that while generating, the model is only allowed to pick tokens that keep the output valid against a grammar or schema. That guarantees shape, not truth: a perfectly valid {"city": "Paris"} can still be the wrong city.
Real-world use: the backbone of AI agents
Almost every LLM product that does something runs on function calling:
| Product | Example tools | Why function calling fits |
|---|---|---|
| Support assistant | get_order, check_policy, create_ticket | Needs private, live data and safe, limited actions. |
| Coding agent | read_file, edit_file, run_tests, search_code | Every step depends on what the last one revealed. |
| Data assistant | run_sql, plot_chart | The model writes the query; the database computes the exact answer. |
| Personal assistant | list_events, create_event, send_email | Actions need user approval, which the app enforces between call and execution. |
Writing good tool definitions Treat descriptions as documentation for a new colleague. Say what the tool does, when to use it, and when not to. Use clear parameter names (order_id, not x), enums for fixed choices ("unit": "celsius" | "fahrenheit"), and keep the tool list short. Too many similar tools confuse models and bloat every prompt.
Quick summary. We describe functions with a name, a description and a JSON Schema. The model replies with a structured call, never running anything itself. Our code validates and executes the call, returns the result tagged with the call id, and calls the model again. Independent calls can run in parallel; dependent calls need multiple rounds. Wrap that in a loop with a step limit, and you have the core of an AI agent.
Common mistakes and how to spot them
Most function-calling bugs show up in one of two places: the tool call the model wrote, or the history we send back. The good news is that each bug leaves a clear trace. If we log every tool call and every tool result, we can usually name the problem in a minute.
| Symptom | Likely cause | Fix |
|---|---|---|
| The model answers from memory and never calls the tool | The description does not say when to use the tool, or the question does not match it | Rewrite the description with a clear “use this when…” line; force the tool if it must always run |
| A call names a tool that does not exist | Similar tool names, or a tool the model has seen elsewhere | Return an error result listing the valid names; keep names distinct |
Arguments have the wrong type, such as "ten" for a number | Loose schema, or vague parameter names | Tighter schema with types and enums; validate in code; use strict mode where the API offers it |
| The arguments string does not parse as JSON | Output was cut off by the token limit, or the model drifted | Parse inside error handling; raise the output limit; send the parse error back |
| The API rejects our second request | A tool call in the history has no result with the same id | Append exactly one result per call, including for failed calls |
| The model calls the right tool with a made-up id or name | It had to guess a value it never saw | Give it a lookup tool first, or tell it to ask the user when a value is unknown |
When a call is broken, the order of our checks matters. Each check assumes the one before it passed.
A safe order for checking one tool call
- Is the tool known?: Look the name up in our own table of functions. Never build a function name from model text and run it.
- Do the arguments parse?: Turn the JSON string into data. A parse failure is a normal event, not a crash.
- Are the fields present and the right type?: Compare against the schema: required fields, types, allowed values.
- Is the call allowed?: Business rules and permissions: may this user do this, and is the amount within limits?
- Run it and report back: Whether a check failed or the function ran, send one result with the call id. A clear error message is what lets the model fix its own call.
Practice: try it yourself
We will build the part of the app that sits between the model and our functions: a dispatcher. It takes a tool call, checks it in the safe order above, runs it, and always returns a result message. We feed it one good call and four broken ones to see every path.
practice_tool_dispatcher.py
import json
# One tool: its schema says which arguments it needs and their Python types.
SCHEMAS = {"convert": {"amount": float, "to": str}}
RATES = {"EUR": 0.5, "INR": 80.0} # made-up rates, easy to check
def convert(amount, to): # the real function
return {"amount": amount * RATES[to], "currency": to}
FUNCS = {"convert": convert}
def dispatch(call):
"""Validate one tool call, run it, and always return a result message."""
name, raw = call["name"], call["arguments"]
try:
if name not in FUNCS:
raise ValueError(f"unknown tool '{name}'")
args = json.loads(raw) # arguments arrive as a JSON string
for key, typ in SCHEMAS[name].items():
if key not in args:
raise ValueError(f"missing argument '{key}'")
if isinstance(args[key], int) and typ is float:
args[key] = float(args[key]) # 10 is a fine float
if not isinstance(args[key], typ):
raise ValueError(f"'{key}' must be {typ.__name__}")
if args["to"] not in RATES:
raise ValueError(f"'to' must be one of {sorted(RATES)}")
result = FUNCS[name](**args)
except (ValueError, json.JSONDecodeError) as e:
result = {"error": str(e)} # errors go back as results
return {"role": "tool", "tool_call_id": call["id"], "content": json.dumps(result)}
# Five calls a model might produce: one good, four broken in different ways.
calls = [{"id": "c1", "name": "convert", "arguments": '{"amount": 10, "to": "EUR"}'},
{"id": "c2", "name": "convert", "arguments": '{"amount": "ten", "to": "EUR"}'},
{"id": "c3", "name": "convert", "arguments": '{"amount": 10}'},
{"id": "c4", "name": "convert", "arguments": '{"amount": 10, "to": "EUR"'},
{"id": "c5", "name": "exchange", "arguments": '{"amount": 10, "to": "INR"}'}]
for c in calls:
print(c["id"], "->", dispatch(c)["content"])Output:
c1 -> {"amount": 5.0, "currency": "EUR"}
c2 -> {"error": "'amount' must be float"}
c3 -> {"error": "missing argument 'to'"}
c4 -> {"error": "Expecting ',' delimiter: line 1 column 27 (char 26)"}
c5 -> {"error": "unknown tool 'exchange'"}Now change it:
- Change call
c1to"to": "USD". Predict which check catches it and the exact error text before you run it. - Change call
c3to'{"amount": 10, "to": 5}'. Predict whether the message will talk about a missing argument or a wrong type. - Remove
json.JSONDecodeErrorfrom theexceptline, leaving onlyValueError, and run again. Predict what happens atc4. (Hint: check whetherJSONDecodeErroris a kind ofValueError.) Then remove the wholetry/exceptand predict again.
Pause and think: Call c2 and call c4 both fail, but they fail at different checks. Why is it useful that the two error messages are different?
The error text is what the model reads on the next turn. “'amount' must be float” tells it to send a number. A parse error tells it the JSON itself was malformed, which often means the output was cut off. A single vague message like “bad call” would leave the model guessing and it might repeat the same mistake.
Pause and think: Suppose the dispatcher raised an exception for c5 instead of returning an error result. What two things would go wrong?
First, the program would stop, so one bad call would end the whole conversation. Second, the history would hold a tool call c5 with no matching result, and many APIs reject such a history. Returning an error result keeps the history consistent and gives the model a chance to pick the correct tool name.
Key takeaways
- Function calling lets a model request that our code run a function; it returns a name plus JSON arguments.
- The model never executes anything; our app validates, authorises, runs and logs every call.
- Results go back as tool messages linked by call id, then the model is called again.
- Independent calls can be parallel in one turn; dependent calls need sequential rounds.
- Structured outputs fix the shape of a response; function calling lets the model decide to act.
Key terms
- Function calling (tool use): An LLM API feature where the model returns a structured request to call a described function instead of plain text.
- Tool definition: The name, description and JSON Schema of parameters we send so the model knows what it can call.
- JSON Schema: A standard way to describe the shape of JSON data: field names, types, required fields and allowed values.
- Tool result: The message we send back with the output (or error) of a function, linked to the call by its id.
- Parallel function calling: The model returning several independent tool calls in one reply so they can run together.
- Constrained decoding: Restricting which tokens the model may generate so the output always matches a grammar or schema.
← 11.1 AI Agents: Autonomous Decision-Making Systems · 11.3 The Agent Loop: Observe, Think, Act, Repeat →