TL;DR
An Agent Loop is a bounded runtime protocol for iterative AI work. On each turn, trusted code supplies the goal, state, and observations; the model proposes a typed answer, tool batch, or escalation; the runtime validates and executes it, commits the resulting event, and either continues or returns an explicit terminal reason. Production quality comes from this control protocol, not from an unbounded while true around an LLM.
Table of Contents
- Key Takeaways
- Agent Loop Definition
- Agent Loop Execution Protocol
- ReAct and Other Agent Patterns
- Production State Machine
- Runnable Go Controller
- Parallel Tool Calls
- Termination and Progress
- Recovery and Side Effects
- Evaluating an Agent Loop
- Failure Modes
- Production Checklist
- FAQ
- Summary
- Primary Sources
Key Takeaways
- The model proposes; the runtime disposes. A model output is an untrusted proposal until code validates its type, schema, permission, budget, and effect policy.
- A loop turn is a protocol transaction. Decide, validate, execute, correlate, commit, and resume are separate stages with different failure semantics.
- Termination is typed.
completed,cancelled,approval_required, andno_progressare not interchangeable forms of failure. - Parallel calls require correlation. Preserve a unique call ID from request through result, then commit one complete observation batch.
- A checkpoint is not an effect ledger. Durable conversation state cannot by itself prevent duplicate emails, payments, deletes, or other external writes.
- Evaluate trajectories, not only final prose. Correct tool selection, effect safety, progress, cost, latency, and stop reason are part of task quality.
Agent Loop Definition
An Agent Loop is the repeated control cycle an AI Agent uses to complete one invocation:
goal and committed state
→ model decision
→ runtime validation
→ tool execution
→ correlated observations
→ durable commit
→ continue or terminate
This definition is narrower than "everything an agent system does." Queue scheduling, worker leases, long-term memory, deployment, and fleet management belong to the broader Agent Runtime. Offline prompt improvement and regression analysis belong to Loop Engineering. The Agent Loop covered here is the inner control protocol for one task execution.
Agent Loop versus a normal while loop
A normal while loop evaluates deterministic conditions written in code. An Agent Loop lets a probabilistic model propose the next semantic action, but the loop boundary must remain deterministic.
| Responsibility | Normal while loop | Production Agent Loop |
|---|---|---|
| Next action | Program logic | Model proposal |
| Allowed transitions | Program logic | Runtime state machine |
| Input validity | Usually typed locally | Schema validation is mandatory |
| Permissions | Called code | Runtime and downstream service |
| Stop condition | Boolean expression | Typed terminal policy |
| Recovery | Program-specific | Checkpoint plus effect reconciliation |
| Correctness evidence | Return value and tests | Final result plus trajectory and effects |
The useful mental model is not "the LLM owns the loop." It is "the LLM is one decision component inside a loop owned by trusted code."
Agent Loop Execution Protocol
A robust turn has eight stages, and the ordering is part of correctness.
| Stage | Runtime responsibility | Persisted evidence |
|---|---|---|
| 1. Load | Read goal, policy version, checkpoint, and remaining budget | Invocation ID and state version |
| 2. Build context | Select relevant history, tool schemas, and current observations | Context references or hashes |
| 3. Decide | Ask the model for one typed decision | Model response ID and decision type |
| 4. Validate | Check schema, call IDs, tools, authorization, and budget | Validation result and policy reason |
| 5. Execute | Run the approved call batch with cancellation and timeouts | Attempt and effect identifiers |
| 6. Correlate | Match every result to its originating call ID | Complete result batch |
| 7. Commit | Atomically append the turn event and advance state version | New durable checkpoint |
| 8. Continue or stop | Re-enter with committed state or emit a terminal reason | Terminal reason and summary |
The commit boundary matters. If the runtime asks the model for the next decision before tool results are durably attached to state, a crash can produce a different history on resume. Google ADK's runtime documentation similarly models state changes as part of committed events before execution continues.
Use a closed decision type
Do not ask the model to return an arbitrary paragraph and infer control flow from its wording. Use a closed union such as:
final(answer)
tool_calls([{call_id, tool_name, arguments}])
request_approval(proposal_digest)
clarification(question)
cannot_continue(reason)
Provider APIs represent these choices differently. The adapter should normalize them into one internal contract. That keeps provider-specific message blocks outside the state machine and makes terminal behavior testable.
Separate protocol state from business state
Protocol state answers "where is execution?" Business state answers "what has the task learned or changed?"
| Protocol state | Business state |
|---|---|
| invocation ID | user request |
| turn and state version | extracted facts |
| pending call IDs | domain objects |
| remaining budgets | validation findings |
| terminal reason | approved result |
| worker lease | external effect status |
Mixing them makes recovery dangerous. A summary such as "refund completed" is not proof that the payment service committed a refund.
ReAct and Other Agent Patterns
ReAct, Plan-and-Execute, and Reflexion can all run inside an Agent Loop; none of them replaces the runtime protocol.
| Pattern | What it changes | What the runtime still owns |
|---|---|---|
| ReAct | Interleaves reasoning and actions | Validation, tools, state, budgets, termination |
| Plan-and-Execute | Creates a plan before executing steps | Plan versioning, permissions, replanning limits |
| Reflexion | Uses feedback to revise behavior | Feedback provenance, retry bounds, acceptance tests |
| State machine | Restricts decisions to explicit states | Tool execution, durable commit, effect safety |
| Graph workflow | Routes work across predefined nodes | Per-node loops, shared state, cancellation |
The original ReAct paper showed the value of interleaving reasoning traces with actions and observations. That is a cognitive pattern, not a production authorization model. You do not need to store hidden chain-of-thought to operate a reliable loop; record typed decisions, evidence references, tool outcomes, policy decisions, and stop reasons instead.
Anthropic's guidance on building effective agents distinguishes predefined workflows from agents that dynamically direct their own process and recommends starting with the simplest pattern that works. The runtime controls in this article apply to either form whenever a model can choose the next action.
Production State Machine
A production Agent Loop should make every continuation and exit path explicit.
The state machine should not expose a generic success: false. Operators and callers need a stable terminal vocabulary.
| Terminal reason | Meaning | Typical caller action |
|---|---|---|
completed |
A validated final answer was committed | Return result |
approval_required |
A valid action crossed a policy gate | Present exact proposal |
cancelled |
Caller or lease cancelled execution | Do not start new effects |
deadline_exceeded |
Wall-clock deadline expired | Retry only under caller policy |
max_steps |
Turn budget was exhausted | Inspect trajectory or narrow goal |
tool_budget_exhausted |
Approved call count reached its cap | Increase only with evidence |
no_progress |
Equivalent action batches repeated | Escalate, replan, or stop |
invalid_decision |
Model output violated the protocol | Repair once or fail explicitly |
model_error |
Model request failed | Apply provider retry policy |
checkpoint_failed |
Durable state could not be committed | Reconcile attempts before resume |
Runnable Go Controller
The following standard-library Go program implements the control plane, not a vendor SDK adapter. It validates typed decisions, preserves call IDs, executes independent tools concurrently, commits observations before the next model turn, enforces budgets, detects repeated plans, handles cancellation, and returns typed terminal reasons.
To keep the example focused, its model adapter accepts only final and tool_calls; approval is enforced by runtime policy. A full adapter can add clarification and model-declared refusal variants without changing the commit or execution invariants.
package main
import (
"bytes"
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"errors"
"fmt"
"sort"
"strings"
"sync"
"time"
)
type DecisionKind string
const (
DecisionFinal DecisionKind = "final"
DecisionTools DecisionKind = "tool_calls"
)
type TerminalReason string
const (
Completed TerminalReason = "completed"
ApprovalRequired TerminalReason = "approval_required"
Cancelled TerminalReason = "cancelled"
DeadlineExceeded TerminalReason = "deadline_exceeded"
MaxSteps TerminalReason = "max_steps"
ToolBudgetExhausted TerminalReason = "tool_budget_exhausted"
NoProgress TerminalReason = "no_progress"
InvalidDecision TerminalReason = "invalid_decision"
ModelError TerminalReason = "model_error"
CheckpointFailed TerminalReason = "checkpoint_failed"
)
var ErrApprovalRequired = errors.New("approval required")
type ToolCall struct {
ID string `json:"id"`
Name string `json:"name"`
Arguments json.RawMessage `json:"arguments"`
}
type Decision struct {
Kind DecisionKind `json:"kind"`
Answer string `json:"answer,omitempty"`
Calls []ToolCall `json:"calls,omitempty"`
}
type ToolResult struct {
CallID string `json:"call_id"`
Name string `json:"name"`
Output json.RawMessage `json:"output,omitempty"`
Error string `json:"error,omitempty"`
}
type Event struct {
Step int `json:"step"`
StateVersion int `json:"state_version"`
Decision Decision `json:"decision"`
Results []ToolResult `json:"results,omitempty"`
}
type State struct {
Goal string
Step int
StateVersion int
ToolCalls int
Events []Event
LastPlan string
RepeatStreak int
}
type Budget struct {
MaxSteps int
MaxToolCalls int
MaxRepeatedPlan int
Deadline time.Time
}
type Outcome struct {
Reason TerminalReason
Answer string
State State
Err error
}
type Tool interface {
Execute(context.Context, json.RawMessage) (json.RawMessage, error)
}
type Controller struct {
Budget Budget
Decide func(context.Context, State) (Decision, error)
Tools map[string]Tool
Authorize func(State, []ToolCall) error
Append func(context.Context, Event) error
}
func contextOutcome(state State, err error) Outcome {
reason := Cancelled
if errors.Is(err, context.DeadlineExceeded) {
reason = DeadlineExceeded
}
return Outcome{Reason: reason, State: state, Err: err}
}
func (c Controller) Run(ctx context.Context, state State) Outcome {
runCtx := ctx
var cancel context.CancelFunc = func() {}
if !c.Budget.Deadline.IsZero() {
runCtx, cancel = context.WithDeadline(ctx, c.Budget.Deadline)
}
defer cancel()
for {
if err := runCtx.Err(); err != nil {
return contextOutcome(state, err)
}
if state.Step >= c.Budget.MaxSteps {
return Outcome{Reason: MaxSteps, State: state}
}
decision, err := c.Decide(runCtx, state)
if err != nil {
if ctxErr := runCtx.Err(); ctxErr != nil {
return contextOutcome(state, ctxErr)
}
return Outcome{Reason: ModelError, State: state, Err: err}
}
switch decision.Kind {
case DecisionFinal:
if strings.TrimSpace(decision.Answer) == "" || len(decision.Calls) != 0 {
return Outcome{Reason: InvalidDecision, State: state}
}
if err := c.commit(runCtx, &state, decision, nil); err != nil {
if ctxErr := runCtx.Err(); ctxErr != nil {
return contextOutcome(state, ctxErr)
}
return Outcome{Reason: CheckpointFailed, State: state, Err: err}
}
return Outcome{Reason: Completed, Answer: decision.Answer, State: state}
case DecisionTools:
if err := c.validateCalls(decision.Calls); err != nil {
return Outcome{Reason: InvalidDecision, State: state, Err: err}
}
if state.ToolCalls+len(decision.Calls) > c.Budget.MaxToolCalls {
return Outcome{Reason: ToolBudgetExhausted, State: state}
}
if c.Authorize != nil {
if err := c.Authorize(state, decision.Calls); err != nil {
if errors.Is(err, ErrApprovalRequired) {
return Outcome{Reason: ApprovalRequired, State: state, Err: err}
}
return Outcome{Reason: InvalidDecision, State: state, Err: err}
}
}
fingerprint, err := batchFingerprint(decision.Calls)
if err != nil {
return Outcome{Reason: InvalidDecision, State: state, Err: err}
}
if fingerprint == state.LastPlan {
state.RepeatStreak++
} else {
state.LastPlan = fingerprint
state.RepeatStreak = 1
}
if state.RepeatStreak > c.Budget.MaxRepeatedPlan {
return Outcome{Reason: NoProgress, State: state}
}
results := c.executeParallel(runCtx, decision.Calls)
if err := c.commit(runCtx, &state, decision, results); err != nil {
if ctxErr := runCtx.Err(); ctxErr != nil {
return contextOutcome(state, ctxErr)
}
return Outcome{Reason: CheckpointFailed, State: state, Err: err}
}
state.ToolCalls += len(decision.Calls)
default:
return Outcome{Reason: InvalidDecision, State: state}
}
}
}
func (c Controller) validateCalls(calls []ToolCall) error {
if len(calls) == 0 {
return errors.New("empty tool batch")
}
seen := make(map[string]struct{}, len(calls))
for _, call := range calls {
if call.ID == "" || call.Name == "" || !json.Valid(call.Arguments) {
return fmt.Errorf("invalid tool call %q", call.ID)
}
if _, ok := seen[call.ID]; ok {
return fmt.Errorf("duplicate call ID %q", call.ID)
}
if _, ok := c.Tools[call.Name]; !ok {
return fmt.Errorf("unknown tool %q", call.Name)
}
seen[call.ID] = struct{}{}
}
return nil
}
func (c Controller) executeParallel(ctx context.Context, calls []ToolCall) []ToolResult {
type indexedResult struct {
index int
result ToolResult
}
ch := make(chan indexedResult, len(calls))
var wg sync.WaitGroup
for index, call := range calls {
wg.Add(1)
go func(index int, call ToolCall) {
defer wg.Done()
output, err := c.Tools[call.Name].Execute(ctx, call.Arguments)
result := ToolResult{CallID: call.ID, Name: call.Name, Output: output}
if err != nil {
result.Output = nil
result.Error = err.Error()
}
ch <- indexedResult{index: index, result: result}
}(index, call)
}
wg.Wait()
close(ch)
indexed := make([]indexedResult, 0, len(calls))
for result := range ch {
indexed = append(indexed, result)
}
sort.Slice(indexed, func(i, j int) bool {
return indexed[i].index < indexed[j].index
})
results := make([]ToolResult, 0, len(indexed))
for _, result := range indexed {
results = append(results, result.result)
}
return results
}
func (c Controller) commit(
ctx context.Context,
state *State,
decision Decision,
results []ToolResult,
) error {
event := Event{
Step: state.Step + 1,
StateVersion: state.StateVersion + 1,
Decision: decision,
Results: results,
}
if err := c.Append(ctx, event); err != nil {
return err
}
state.Step = event.Step
state.StateVersion = event.StateVersion
state.Events = append(state.Events, event)
return nil
}
func batchFingerprint(calls []ToolCall) (string, error) {
type normalizedCall struct {
Name string `json:"name"`
Arguments json.RawMessage `json:"arguments"`
}
normalized := make([]normalizedCall, 0, len(calls))
for _, call := range calls {
var compacted bytes.Buffer
if err := json.Compact(&compacted, call.Arguments); err != nil {
return "", err
}
normalized = append(normalized, normalizedCall{
Name: call.Name,
Arguments: append(json.RawMessage(nil), compacted.Bytes()...),
})
}
sort.Slice(normalized, func(i, j int) bool {
if normalized[i].Name == normalized[j].Name {
return string(normalized[i].Arguments) < string(normalized[j].Arguments)
}
return normalized[i].Name < normalized[j].Name
})
payload, err := json.Marshal(normalized)
if err != nil {
return "", err
}
sum := sha256.Sum256(payload)
return hex.EncodeToString(sum[:]), nil
}
type lookupTool struct {
values map[string]string
}
func (t lookupTool) Execute(
ctx context.Context,
arguments json.RawMessage,
) (json.RawMessage, error) {
select {
case <-ctx.Done():
return nil, ctx.Err()
default:
}
var input struct {
Key string `json:"key"`
}
if err := json.Unmarshal(arguments, &input); err != nil {
return nil, err
}
value, ok := t.values[input.Key]
if !ok {
return nil, fmt.Errorf("key %q not found", input.Key)
}
return json.Marshal(struct {
Key string `json:"key"`
Value string `json:"value"`
}{Key: input.Key, Value: value})
}
func scriptedDecision(_ context.Context, state State) (Decision, error) {
if len(state.Events) == 0 {
return Decision{
Kind: DecisionTools,
Calls: []ToolCall{
{ID: "call-profile", Name: "lookup", Arguments: json.RawMessage(`{"key":"profile"}`)},
{ID: "call-policy", Name: "lookup", Arguments: json.RawMessage(`{"key":"policy"}`)},
},
}, nil
}
values := make(map[string]string)
for _, result := range state.Events[len(state.Events)-1].Results {
var output struct {
Key string `json:"key"`
Value string `json:"value"`
}
if result.Error != "" {
return Decision{}, errors.New(result.Error)
}
if err := json.Unmarshal(result.Output, &output); err != nil {
return Decision{}, err
}
values[output.Key] = output.Value
}
return Decision{
Kind: DecisionFinal,
Answer: fmt.Sprintf("profile=%s; policy=%s", values["profile"], values["policy"]),
}, nil
}
func main() {
events := make([]Event, 0, 2)
controller := Controller{
Budget: Budget{
MaxSteps: 4,
MaxToolCalls: 4,
MaxRepeatedPlan: 1,
},
Decide: scriptedDecision,
Tools: map[string]Tool{
"lookup": lookupTool{values: map[string]string{
"profile": "active",
"policy": "refund-under-100",
}},
},
Append: func(_ context.Context, event Event) error {
events = append(events, event)
return nil
},
}
outcome := controller.Run(context.Background(), State{
Goal: "Load the customer profile and refund policy",
})
fmt.Println(outcome.Reason)
fmt.Println(outcome.Answer)
fmt.Printf(
"steps=%d tool_calls=%d events=%d\n",
outcome.State.Step,
outcome.State.ToolCalls,
len(events),
)
}
Expected output:
completed
profile=active; policy=refund-under-100
steps=2 tool_calls=2 events=2
The sample intentionally treats both tools as read-only. A production adapter should also validate each tool's argument schema, establish a per-call timeout, redact persisted payloads, and classify effectful calls before concurrency is allowed.
Deterministic control tests
The model itself may vary, but the controller's invariants should not. Inject scripted decisions and fake tools to test correlation, budgets, cancellation, approvals, repeated plans, and checkpoint failures.
package main
import (
"context"
"encoding/json"
"sync/atomic"
"testing"
"time"
)
type delayedLookup struct {
values map[string]string
}
func (t delayedLookup) Execute(
ctx context.Context,
arguments json.RawMessage,
) (json.RawMessage, error) {
var input struct {
Key string `json:"key"`
}
if err := json.Unmarshal(arguments, &input); err != nil {
return nil, err
}
if input.Key == "profile" {
select {
case <-ctx.Done():
return nil, ctx.Err()
case <-time.After(10 * time.Millisecond):
}
}
return json.Marshal(struct {
Key string `json:"key"`
Value string `json:"value"`
}{Key: input.Key, Value: t.values[input.Key]})
}
func newTestController() Controller {
return Controller{
Budget: Budget{
MaxSteps: 4,
MaxToolCalls: 4,
MaxRepeatedPlan: 1,
},
Decide: scriptedDecision,
Tools: map[string]Tool{
"lookup": delayedLookup{values: map[string]string{
"profile": "active",
"policy": "refund-under-100",
}},
},
Append: func(context.Context, Event) error {
return nil
},
}
}
type countingTool struct {
executions *atomic.Int32
}
func (t countingTool) Execute(
_ context.Context,
_ json.RawMessage,
) (json.RawMessage, error) {
t.executions.Add(1)
return json.RawMessage(`{"ok":true}`), nil
}
func newRepeatingTestController(executions *atomic.Int32) Controller {
return Controller{
Budget: Budget{
MaxSteps: 4,
MaxToolCalls: 4,
MaxRepeatedPlan: 1,
},
Decide: func(context.Context, State) (Decision, error) {
return Decision{
Kind: DecisionTools,
Calls: []ToolCall{{
ID: "same-call",
Name: "count",
Arguments: json.RawMessage(`{"key":"same"}`),
}},
}, nil
},
Tools: map[string]Tool{
"count": countingTool{executions: executions},
},
Append: func(context.Context, Event) error {
return nil
},
}
}
func TestParallelResultsPreserveCallOrder(t *testing.T) {
controller := newTestController()
outcome := controller.Run(context.Background(), State{Goal: "test"})
if outcome.Reason != Completed {
t.Fatalf("reason = %s, want %s", outcome.Reason, Completed)
}
results := outcome.State.Events[0].Results
if results[0].CallID != "call-profile" || results[1].CallID != "call-policy" {
t.Fatalf("unexpected result order: %#v", results)
}
}
func TestRepeatedPlanStopsBeforeSecondExecution(t *testing.T) {
var executions atomic.Int32
controller := newRepeatingTestController(&executions)
outcome := controller.Run(context.Background(), State{Goal: "test"})
if outcome.Reason != NoProgress {
t.Fatalf("reason = %s, want %s", outcome.Reason, NoProgress)
}
if executions.Load() != 1 {
t.Fatalf("executions = %d, want 1", executions.Load())
}
}
This test style is more reliable than asserting an exact natural-language trace. It verifies the boundary that must remain deterministic even when a real model supplies decisions.
Parallel Tool Calls
Parallel tool execution is safe only when independence is proven by policy, not merely because the model emitted multiple calls.
OpenAI's Agents SDK loop documentation describes a model turn that may emit multiple tool calls. Anthropic's tool-use documentation uses a unique tool_use ID and requires each tool_result to reference it. These wire formats differ, but they imply the same runtime invariants:
- Every call receives a unique, stable ID.
- The result repeats that call ID.
- A result never attaches by array position alone.
- The next model turn sees the complete result batch.
- Concurrency is a runtime decision, not a model permission.
When calls may run concurrently
Calls may run together when they are read-only or commute: neither changes data read or written by the other, ordering is irrelevant, and both can be retried independently.
| Batch | Run concurrently | Reason |
|---|---|---|
| Read customer profile + read policy | Usually yes | Independent reads |
| Search two unrelated indexes | Usually yes | No shared mutation |
| Create order + charge that order | No | Charge depends on created ID |
| Update balance + send final receipt | No | Receipt depends on committed balance |
| Two writes to the same document | No | Order and conflict policy matter |
For mixed batches, build a dependency graph or reject the batch and ask the planner for explicit sequencing. Do not guess dependencies from tool names.
Do not collapse partial failures
A batch can contain two successes and one failure. Preserve every result:
[
{"call_id":"call-a","name":"read_profile","output":{"tier":"pro"}},
{"call_id":"call-b","name":"read_policy","error":"deadline exceeded"},
{"call_id":"call-c","name":"read_balance","output":{"amount":42}}
]
The next decision can retry only call-b, use the successful evidence, or stop. A single batch-level failed flag destroys that option and encourages duplicate work.
Termination and Progress
An Agent Loop should terminate by protocol, policy, or infrastructure reason; only the first category is a normal model answer.
Protocol termination
- The model emits
finaland the answer passes output validation. - The model emits
cannot_continuewith a structured reason. - The user must clarify an ambiguity.
Policy termination
- A proposed tool requires human approval.
- The call exceeds tenant, object, amount, or capability permissions.
- Step, token, tool-call, cost, or wall-clock budget is exhausted.
- The same or equivalent plan repeats without new evidence.
Infrastructure termination
- The model provider fails beyond its retry budget.
- A required tool remains unavailable.
- The checkpoint cannot be committed.
- The worker loses its lease or receives cancellation.
Treat max_steps as a safety net, not the primary definition of progress. A loop that makes eight distinct but useless calls is still stuck.
Detect semantic no-progress
A practical progress detector combines several signals:
| Signal | Example |
|---|---|
| Repeated action fingerprint | Same tool and normalized arguments recur |
| State delta | No new fact, artifact, or validation result was committed |
| Goal predicate | Required acceptance condition remains unchanged |
| Error cycle | The same error class follows the same repair action |
| Evidence novelty | New output duplicates already committed evidence |
Start with exact normalized fingerprints because they are deterministic. Add semantic similarity only as a secondary signal, with replay data and a false-stop review process. A semantic detector can incorrectly stop valid polling or pagination.
Recovery and Side Effects
Crash recovery is usually at-least-once, so exactly-once business effects must be created at the service boundary.
Google ADK's resume documentation explicitly describes at-least-once resume behavior and duplicate-tool prevention concerns. The general lesson is framework-independent: after a crash, the runtime may know that it attempted a call without knowing whether the remote system committed it.
Use three durable identities
| Identity | Scope | Purpose |
|---|---|---|
| Invocation ID | Entire task | Groups all turns and retries |
| Call ID | One logical tool request | Correlates proposal, attempts, and results |
| Effect key | One business mutation | Deduplicates the external side effect |
The effect key must be stable across transport retries and process restarts. The downstream service that owns the mutation must enforce it; an in-memory map in the Agent process is insufficient.
Model the unknown outcome
Suppose a payment call times out after the request reaches the server. The correct state is not automatically failed:
prepared → dispatched → committed
↘ outcome_unknown → reconcile → committed or rejected
Retrying from outcome_unknown without querying by effect key can double-charge. First reconcile with the owning service. If reconciliation is unavailable, stop for review instead of claiming success or failure.
Checkpoint and effect ledger solve different problems
- Checkpoint: what decisions and observations the Agent Loop has durably committed.
- Effect ledger: what external mutations the business service accepted, rejected, or cannot yet resolve.
Persist the call proposal before dispatch, persist the correlated result after completion, and resume only from committed state. For high-risk writes, use a transactional outbox or service-owned idempotency record so that state and dispatch cannot silently diverge.
Cancellation follows the same ownership rule. Cancel the invocation, revoke or expire its worker lease, and require each tool to check context.Context before starting a new effect. Cancellation cannot retroactively undo an effect that already committed.
Evaluating an Agent Loop
An Agent Loop evaluation needs task, trajectory, effect, and operational metrics. Final-answer grading alone can reward unsafe or wasteful paths.
| Dimension | Metric | Question answered |
|---|---|---|
| Task | Success rate by scenario | Did the requested outcome occur? |
| Decision | Correct tool and argument rate | Did the model propose the right action? |
| Efficiency | Turns, calls, tokens, and cost per success | How much work did success require? |
| Progress | No-progress and repeated-plan rate | Did the loop advance toward the goal? |
| Reliability | Tool error recovery and resume success | Did controlled failures recover correctly? |
| Effects | Unauthorized or duplicate effect rate | Did the loop mutate only what it should? |
| Termination | Stop-reason precision | Did it stop for the right reason? |
| Operations | p50, p95, and p99 latency by terminal reason | Where does execution time accumulate? |
Report distributions by task class and terminal reason. A global average hides the difference between a fast approval_required, a slow successful research task, and a repeated-plan incident.
Build an evaluation set from state transitions
Include cases that force each important branch:
- direct answer without tools
- one valid tool call
- independent parallel calls returned out of order
- malformed arguments and unknown tools
- partial batch failure
- approval-required write
- repeated equivalent plan
- cancellation while tools are running
- crash after dispatch but before result commit
- resume with an existing effect key
- final answer that fails a deterministic validator
For deeper trace design, use Agent Observability Engineering. It covers event contracts, privacy, OpenTelemetry boundaries, and release metrics without requiring storage of private reasoning.
Failure Modes
The model controls the loop boundary
Symptom: The prompt says "stop when done," but there is no enforced step, time, or tool budget.
Fix: Put budgets and terminal reasons in trusted code. Prompt instructions are hints, not resource controls.
Tool results are concatenated without IDs
Symptom: Parallel results are attached by order, and a delayed response is interpreted as another tool's output.
Fix: Require unique call IDs and reject missing, duplicate, or unknown IDs before committing the batch.
Errors become plain text
Symptom: The model cannot distinguish validation failure, timeout, denial, unknown outcome, and permanent business rejection.
Fix: Preserve structured error class, retryability, attempt ID, and effect status.
Every failure is retried
Symptom: Invalid arguments, authorization denials, and uncertain writes are treated like transient read failures.
Fix: Define retry policy by operation and error class. Never let the model turn a denial into permission.
The checkpoint claims an external effect
Symptom: State says an email or refund completed, but there is no service-owned receipt.
Fix: Store the authoritative effect ID and reconcile against the owning service.
Final text is the only evaluation
Symptom: A plausible answer passes even though the loop used the wrong customer record, duplicated a write, or exceeded budget.
Fix: Grade state transitions, evidence, effects, budgets, and terminal reason alongside the final answer.
Production Checklist
Before shipping one Agent Loop, verify:
- [ ] The model returns a closed, versioned decision schema.
- [ ] Unknown decision types, tools, fields, and duplicate call IDs are rejected.
- [ ] Tool permissions are checked for actor, tenant, object, and operation.
- [ ] Independent calls are identified by policy before parallel execution.
- [ ] Results retain call IDs and partial failures.
- [ ] A complete turn event is committed before the next model call.
- [ ] Step, tool, token, cost, and wall-clock budgets are enforced in code.
- [ ] Cancellation prevents the worker from starting another effect.
- [ ] Exact repeated plans and unchanged goal predicates trigger no-progress handling.
- [ ] External writes use stable idempotency keys and service-owned receipts.
- [ ] Unknown outcomes enter reconciliation instead of blind retry.
- [ ] Terminal reasons are typed and visible to callers and operators.
- [ ] Replay tests cover every critical transition and failure branch.
- [ ] Traces store bounded evidence, not hidden chain-of-thought by default.
FAQ
What is an Agent Loop
An Agent Loop is the bounded runtime cycle that turns a model from a one-shot responder into an iterative task executor. Each turn loads committed state, asks the model for a typed decision, validates that decision, executes approved tools, commits correlated observations, and either repeats or returns an explicit terminal reason.
How is an Agent Loop different from a normal while loop
A normal while loop owns both its condition and next action in deterministic code. In an Agent Loop, a model proposes the next semantic action. Trusted code still owns the state machine, schema validation, permissions, budgets, effect policy, commit order, cancellation, and termination. Wrapping an LLM call in while true without those controls is not a production Agent Loop.
How should an Agent Loop terminate
Use a closed set of terminal reasons rather than one success flag. At minimum distinguish completed, approval_required, cancelled, deadline_exceeded, max_steps, tool_budget_exhausted, no_progress, invalid_decision, model_error, and checkpoint_failed. The caller can then decide whether to return, request approval, retry, reconcile, or escalate.
Can an Agent Loop execute tools in parallel
Yes, if the calls are independent and runtime policy permits it. Assign each call a unique ID, execute the batch under cancellation and timeout controls, retain partial failures, restore a deterministic order, and commit the full correlated batch before the next model turn. Dependent or conflicting writes must be sequenced.
How can an Agent Loop resume without duplicating effects
Persist invocation, call, attempt, and effect identities. On resume, load the last committed event and query the owning service for any dispatched-but-unresolved effect key. Do not assume a timeout means failure, and do not rely on a conversation checkpoint as proof of a business mutation. At-least-once orchestration requires idempotent downstream effects or explicit reconciliation.
Summary
An Agent Loop is best understood as a protocol-controlled state machine, not an autonomous while loop. The model proposes typed decisions; trusted code validates permissions and budgets, executes and correlates tools, commits state, detects progress, and emits a precise terminal reason. Parallelism improves latency only when dependencies are known, and durable resume is safe only when checkpoints are paired with service-owned effect identities.
For the outer process that improves prompts, tools, and policies across many runs, see Agent Loop vs Loop Engineering. For execution-wide scheduling, leases, and durable lifecycle management, continue with the Agent Runtime glossary.