State, Memory and Durable Execution
Where agent state lives, how short-term and long-term memory work, and how durable execution lets long-running agents survive crashes, retries and human approval pauses.
State lives outside the model
Every agent step is a fresh, stateless LLM call. Everything the agent "remembers" is state your system stores and re-sends:
| State | What | Where it typically lives |
|---|---|---|
| Conversation / transcript | Messages, tool calls, results | Database (Postgres, DynamoDB), keyed by session |
| Task state | Plan, current step, intermediate outputs, status | Workflow engine or DB |
| Artifacts | Files, drafts, code changes | Object storage, sandbox filesystem |
| Long-term memory | Facts about users, preferences, past outcomes | DB + vector index |
Short-term vs long-term memory
Short-term (working) memory is what is in the context window right now. Manage it with the techniques from the context windows lesson: trim, summarise, compact, or offload to files.
Long-term memory persists across sessions:
- What to store: stable facts and preferences ("prefers Python", "account is on the EU region"), not every sentence.
- Updating: new facts may contradict old ones, so store timestamps and resolve conflicts, or let the LLM merge.
- Privacy: users must be able to see and delete what's remembered. Memory is personal data under GDPR and similar laws.
- Scope: per user, per team or per tenant. Never let memories cross those boundaries.
Why agents need durable execution
A research agent that runs for 20 minutes, or a workflow waiting two days for a manager's approval, will run into server restarts, deploys, timeouts, rate limits and crashes. If state lives only in process memory, the whole task restarts: you pay for every LLM call again, and side effects may repeat.
Durable execution means:
- Checkpoint after every step: store the transcript and task state before moving on.
- Resume from the last checkpoint on failure, on any worker.
- Pause and resume for human approvals, scheduled waits or external events, without holding a process open.
- Retry with backoff on transient errors (429s, timeouts), per step rather than per task.
Workflow engines such as Temporal, AWS Step Functions and Restate, plus agent frameworks with checkpointing, provide this. A simpler option is a queue plus a state table that a worker updates after each step.
Idempotency: retries without double side effects
If the agent crashes after calling charge_card but before checkpointing, a naive resume charges twice. Protect against this:
- Give each tool call an idempotency key, for example task ID + step number, and have side-effecting services deduplicate on it.
- Record intent before acting and outcome after, so recovery can check whether the action happened.
- Prefer tools that are naturally idempotent, such as "set status to X" rather than "toggle status".
Concurrency and ownership
- Two agents, or two runs of the same agent, editing the same record need optimistic locking or clear ownership.
- User messages may arrive while an agent step is running. Decide whether to queue them, interrupt the run, or reject them.
Key takeaways
- The LLM is stateless. The conversation, the task progress and the memory all live in your storage.
- Short-term memory is the working context. Long-term memory is stored facts retrieved into the context when relevant.
- Long-running agents need durable execution. Checkpoint after each step so a crash or deploy resumes rather than restarts.
- Make tool calls idempotent so retries after failures don't repeat side effects.