Function Calling and Structured Outputs
How LLMs call tools and return machine-readable data, how the request and response cycle works, how to design good tools, and how to make outputs reliably parseable.
From text to actions
On its own, an LLM only produces text. Function calling (also called tool use) lets it ask your system to do something:
- You send the conversation plus tool definitions: name, description and a JSON schema for the parameters.
- The model either answers or returns one or more tool calls with arguments.
- Your code validates the arguments, executes the call, and appends the result to the conversation.
- The model continues, answering or calling more tools.
The model never touches your systems directly. Your code is the enforcement point for validation, permissions and rate limits. This is the most important security fact about tool use.
Structured outputs
Much LLM integration isn't agents at all. It is "extract these fields from this email" or "classify this ticket". You need output your code can parse every time.
| Approach | Reliability |
|---|---|
| "Respond in JSON" in the prompt | Usually works, sometimes doesn't (extra prose, trailing commas, missing fields) |
| JSON mode | Valid JSON, but not necessarily your schema |
| Structured outputs / constrained decoding | Output is guaranteed to match the JSON schema: the decoder masks out tokens that would break it |
Use schema-constrained output wherever the provider or engine supports it. Self-hosted engines (vLLM, SGLang) support it through grammar-constrained decoding. Then validate semantics in code, such as dates in range, IDs that exist, and enums that make sense. The schema guarantees shape, not truth.
Designing good tools
The model reads your tool definitions as instructions, so they are prompts as much as APIs.
- Names and descriptions:
search_orders_by_customer_emailbeatsquery. Say when to use the tool and when not to. - Few, typed parameters: use enums instead of free text, and required fields where possible.
- Right granularity: avoid 50 tiny tools (too many choices and round trips) and one giant
run_sql(too much power). Model tools on user tasks, such asrefund_order(order_id, reason). - Concise results: return what the model needs, not a 5,000-token raw API response. Every result token is paid for on every later step.
- Helpful errors: "Order 8841 not found. Did you mean 8814?" lets the model recover. A stack trace doesn't.
- Idempotency: the model may retry. Make side-effecting tools safe to call twice, for example with an idempotency key.
Cost and latency of tools
With 50+ tools, consider tool retrieval: embed tool descriptions and include only the most relevant 10–15 per request, or group tools behind sub-agents.
Forcing and restricting tool use
APIs typically let you:
- Require a tool call (tool choice "required", or a specific tool). This is useful for extraction pipelines, where the "tool" is just your output schema.
- Disallow tools for a turn.
- Limit which tools are available per user or role. This is part of your permission model.
Key takeaways
- With function calling, the model returns a structured request to call a tool. Your code executes it and sends back the result. The model never runs code itself.
- Structured outputs (JSON-schema constrained decoding) make responses reliably parseable, which is the foundation of LLM-to-system integration.
- Tool design is API design. Use clear names and descriptions, few, well-typed parameters, and concise, informative results and errors.
- Tool definitions cost prompt tokens on every call, and every tool call adds a round trip of latency.