GenAI System Design
5. Agents and Tool Use

The Agent Loop

What makes a system an agent, the observe-think-act loop, workflows vs agents, stopping conditions, and the failure modes to design against.

Lesson 2 of 7 10 min

The loop

Goal + context
LLM decides
answer or tool call
Execute tool
search, code, API
Observe result
append to context
Done?
no → loop
Also known as ReAct (reason + act). The model's own judgement decides the path and when to stop.

In code, the core is small:

messages = [system_prompt, user_goal] for step in range(MAX_STEPS): reply = llm(messages, tools=TOOLS) if not reply.tool_calls: return reply.text # model decided it's done for call in reply.tool_calls: result = execute(call) # validated, permission-checked messages += [reply, tool_result(call.id, result)] raise StepLimitExceeded

Everything hard is outside this loop: tool design, context management, state, error handling, cost control and evaluation.

Workflows vs agents

Not every multi-step LLM system should be an agent.

PatternControl flowUse when
Prompt chainFixed sequence: extract → draft → checkSteps are known in advance
RouterA classifier picks one of several pathsDistinct request types
Parallel / map-reduceSame step over many inputs, then combineSummarising many docs, multi-aspect review
Evaluator-optimiserGenerate → critique → revise, N timesQuality-sensitive output (code, writing)
AgentThe model picks each stepOpen-ended tasks with unpredictable paths

Workflows are cheaper, faster, more predictable and easier to test. A good default is to start with a workflow and give the model autonomy only where the task needs it. Many "agents" in production are workflows with one agentic step.

Stopping conditions

An unbounded loop is a runaway cost and a reliability risk. Enforce:

  • Max steps, for example 10–30, with a graceful "here's what I found so far" when the limit hits.
  • Token and cost budget per task, checked before each call.
  • Wall-clock timeout per task and per tool call.
  • A clear definition of done in the prompt, and ideally a tool for it (submit_answer), so completion is explicit and structured.
  • Loop detection: the same tool called with the same arguments repeatedly means the agent is stuck, so break out or escalate.

Failure modes and mitigations

FailureLooks likeMitigation
LoopingRepeats the same search or editDetect repeated calls, step limits, better tool errors
DriftWanders into unrelated subtasksRestate the goal in context, planning steps, narrower tools
Compounding errorsA wrong early assumption poisons later stepsVerification steps, checkpoints, self-critique before acting
Tool misuseWrong arguments, wrong toolClear descriptions, enums, examples, validation errors
Premature "done"Claims success without verifyingRequire evidence (test output, citations) before submitting
Context overflowQuality drops as history growsTrim tool outputs, compaction, sub-agents (context lesson)

Humans in the loop

For side effects that are costly or irreversible, such as payments, emails, deletions or production changes, insert an approval step. The agent proposes the action and a human confirms it. Design the UX and the state machine so the task can pause waiting for approval and resume later. That requires durable state (lesson 5).

Evaluating agents

Agents need task-level evals:

  • Did it complete the task? Check with a verifiable outcome where possible (tests pass, the right record changed).
  • How many steps, tokens and dollars did it take?
  • Did it take any unsafe or unnecessary actions?

Run a suite of representative tasks on every prompt, tool or model change, and track success rate against cost per task.

Key takeaways

  • An agent is an LLM in a loop. It chooses the next action (tool call), observes the result, and repeats until it decides it's done.
  • Prefer fixed workflows (prompt chains, routing, parallel steps) when the path is known. Use agents when the steps can't be predicted.
  • Every loop needs stopping conditions, such as max steps, tokens, wall-clock time and cost, plus a clear definition of done.
  • Agents fail by looping, drifting off task, misusing tools and compounding errors. Design checks and human checkpoints.

Go deeper