GenAI System Design
5. Agents and Tool Use

Function Calling and Structured Outputs

How LLMs call tools and return machine-readable data, how the request and response cycle works, how to design good tools, and how to make outputs reliably parseable.

Lesson 1 of 7 9 min

From text to actions

On its own, an LLM only produces text. Function calling (also called tool use) lets it ask your system to do something:

Your app
sends messages + tool schemas
LLM
returns tool_call: get_order(id=8841)
Your app executes
validates + runs the call
Tool / API
returns result
LLM continues
uses result → answer or next call
  1. You send the conversation plus tool definitions: name, description and a JSON schema for the parameters.
  2. The model either answers or returns one or more tool calls with arguments.
  3. Your code validates the arguments, executes the call, and appends the result to the conversation.
  4. The model continues, answering or calling more tools.

The model never touches your systems directly. Your code is the enforcement point for validation, permissions and rate limits. This is the most important security fact about tool use.

Structured outputs

Much LLM integration isn't agents at all. It is "extract these fields from this email" or "classify this ticket". You need output your code can parse every time.

ApproachReliability
"Respond in JSON" in the promptUsually works, sometimes doesn't (extra prose, trailing commas, missing fields)
JSON modeValid JSON, but not necessarily your schema
Structured outputs / constrained decodingOutput is guaranteed to match the JSON schema: the decoder masks out tokens that would break it

Use schema-constrained output wherever the provider or engine supports it. Self-hosted engines (vLLM, SGLang) support it through grammar-constrained decoding. Then validate semantics in code, such as dates in range, IDs that exist, and enums that make sense. The schema guarantees shape, not truth.

Designing good tools

The model reads your tool definitions as instructions, so they are prompts as much as APIs.

  • Names and descriptions: search_orders_by_customer_email beats query. Say when to use the tool and when not to.
  • Few, typed parameters: use enums instead of free text, and required fields where possible.
  • Right granularity: avoid 50 tiny tools (too many choices and round trips) and one giant run_sql (too much power). Model tools on user tasks, such as refund_order(order_id, reason).
  • Concise results: return what the model needs, not a 5,000-token raw API response. Every result token is paid for on every later step.
  • Helpful errors: "Order 8841 not found. Did you mean 8814?" lets the model recover. A stack trace doesn't.
  • Idempotency: the model may retry. Make side-effecting tools safe to call twice, for example with an idempotency key.

Cost and latency of tools

With 50+ tools, consider tool retrieval: embed tool descriptions and include only the most relevant 10–15 per request, or group tools behind sub-agents.

Forcing and restricting tool use

APIs typically let you:

  • Require a tool call (tool choice "required", or a specific tool). This is useful for extraction pipelines, where the "tool" is just your output schema.
  • Disallow tools for a turn.
  • Limit which tools are available per user or role. This is part of your permission model.

Key takeaways

  • With function calling, the model returns a structured request to call a tool. Your code executes it and sends back the result. The model never runs code itself.
  • Structured outputs (JSON-schema constrained decoding) make responses reliably parseable, which is the foundation of LLM-to-system integration.
  • Tool design is API design. Use clear names and descriptions, few, well-typed parameters, and concise, informative results and errors.
  • Tool definitions cost prompt tokens on every call, and every tool call adds a round trip of latency.

Go deeper

Finished reading? Mark it done to track your progress.