The Agent Loop
What makes a system an agent, the observe-think-act loop, workflows vs agents, stopping conditions, and the failure modes to design against.
The loop
In code, the core is small:
messages = [system_prompt, user_goal]
for step in range(MAX_STEPS):
reply = llm(messages, tools=TOOLS)
if not reply.tool_calls:
return reply.text # model decided it's done
for call in reply.tool_calls:
result = execute(call) # validated, permission-checked
messages += [reply, tool_result(call.id, result)]
raise StepLimitExceededEverything hard is outside this loop: tool design, context management, state, error handling, cost control and evaluation.
Workflows vs agents
Not every multi-step LLM system should be an agent.
| Pattern | Control flow | Use when |
|---|---|---|
| Prompt chain | Fixed sequence: extract → draft → check | Steps are known in advance |
| Router | A classifier picks one of several paths | Distinct request types |
| Parallel / map-reduce | Same step over many inputs, then combine | Summarising many docs, multi-aspect review |
| Evaluator-optimiser | Generate → critique → revise, N times | Quality-sensitive output (code, writing) |
| Agent | The model picks each step | Open-ended tasks with unpredictable paths |
Workflows are cheaper, faster, more predictable and easier to test. A good default is to start with a workflow and give the model autonomy only where the task needs it. Many "agents" in production are workflows with one agentic step.
Stopping conditions
An unbounded loop is a runaway cost and a reliability risk. Enforce:
- Max steps, for example 10–30, with a graceful "here's what I found so far" when the limit hits.
- Token and cost budget per task, checked before each call.
- Wall-clock timeout per task and per tool call.
- A clear definition of done in the prompt, and ideally a tool for it (
submit_answer), so completion is explicit and structured. - Loop detection: the same tool called with the same arguments repeatedly means the agent is stuck, so break out or escalate.
Failure modes and mitigations
| Failure | Looks like | Mitigation |
|---|---|---|
| Looping | Repeats the same search or edit | Detect repeated calls, step limits, better tool errors |
| Drift | Wanders into unrelated subtasks | Restate the goal in context, planning steps, narrower tools |
| Compounding errors | A wrong early assumption poisons later steps | Verification steps, checkpoints, self-critique before acting |
| Tool misuse | Wrong arguments, wrong tool | Clear descriptions, enums, examples, validation errors |
| Premature "done" | Claims success without verifying | Require evidence (test output, citations) before submitting |
| Context overflow | Quality drops as history grows | Trim tool outputs, compaction, sub-agents (context lesson) |
Humans in the loop
For side effects that are costly or irreversible, such as payments, emails, deletions or production changes, insert an approval step. The agent proposes the action and a human confirms it. Design the UX and the state machine so the task can pause waiting for approval and resume later. That requires durable state (lesson 5).
Evaluating agents
Agents need task-level evals:
- Did it complete the task? Check with a verifiable outcome where possible (tests pass, the right record changed).
- How many steps, tokens and dollars did it take?
- Did it take any unsafe or unnecessary actions?
Run a suite of representative tasks on every prompt, tool or model change, and track success rate against cost per task.
Key takeaways
- An agent is an LLM in a loop. It chooses the next action (tool call), observes the result, and repeats until it decides it's done.
- Prefer fixed workflows (prompt chains, routing, parallel steps) when the path is known. Use agents when the steps can't be predicted.
- Every loop needs stopping conditions, such as max steps, tokens, wall-clock time and cost, plus a clear definition of done.
- Agents fail by looping, drifting off task, misusing tools and compounding errors. Design checks and human checkpoints.