TL;DR

An Agent Loop runs inside one invocation. It repeatedly calls a model, executes tools or handoffs, commits observations, and stops with a typed result or limit.

Loop Engineering designs the external system that discovers work, launches one or more agent runs, verifies artifacts, records durable state, and decides whether to retry, release, escalate, or stop. It is an emerging engineering label rather than a finished standard.

The distinction matters because an inner loop can report “done” while the outer goal is still false. A test may fail after the agent's final answer, a deployment may be unhealthy, or a write may have an unknown outcome. The outer controller owns those facts.

Three Loops That Teams Commonly Confuse

There are at least three different cycles in an AI delivery system:

Cycle Unit of work Controller Terminal evidence
Agent runtime loop One invocation Runner or Agent Runtime Final output, handoff, pause, error, cancellation, or turn limit
Engineered automation loop Repeated attempts toward one durable goal Scheduler or Loop Controller Independent verifier, budget, deadline, block, or human decision
Product improvement lifecycle Many tasks and released versions Engineering and product process Eval results, incidents, user outcomes, and release decision

The previous version of this article treated the third row as Loop Engineering itself. That is too narrow. Evaluation and release governance are essential, but a loop specification also governs how recurring work is found, isolated, attempted, checked, remembered, and stopped.

Agent Loop: One Invocation's Runtime Protocol

The Agent Loop answers:

Given the committed state and current observation, what transition happens next?

OpenAI's Agents SDK documents an inner loop that calls the model, then either returns final output, performs a handoff, executes tool calls and loops again, or raises after a turn limit. Google ADK documents an event loop in which execution yields an event, the Runner processes and commits it, and execution resumes afterward. These are framework-specific implementations of a shared runtime shape.

flowchart LR S["Start"] --> L["Load committed state"] L --> M["Call model"] M --> V["Validate decision"] V -->|Tool calls| T["Execute tools"] V -->|Handoff| H["Switch agent"] V -->|Final output| C["Complete"] T --> O["Commit observations"] H --> O O --> M M -->|Approval required| P["Paused"] M -->|Invalid output or limit| F["Failed"] P -->|Resume| L

The runtime owns correlation IDs, tool-result ordering, state commitment, cancellation, limits, and recovery. The model proposes a transition; it does not get to rewrite identity, authorization, or committed business state.

Loop Engineering: A Bounded External Loop

Loop Engineering answers:

What durable system can repeatedly invoke agents and checks until a verifiable goal reaches a terminal state?

A 2026 preprint proposes the term loop specification for a bounded reusable artifact containing a trigger, goal, verification step, stopping rule, and memory. Practitioner descriptions add automation, isolated worktrees, skills, connectors, subagents, and external state. These sources describe an emerging practice, not a universal standards-body definition.

A production loop specification should include:

Contract Required question
Trigger What event or schedule creates work, and how is duplicate discovery suppressed?
Goal Which observable state must become true?
Work identity Which stable ID binds attempts, artifacts, approvals, and side effects?
Isolation Which branch, worktree, sandbox, tenant, and credentials belong to the attempt?
Executor Which Agent Harness may run, with which tools and limits?
Verifier Which independent check decides whether the artifact satisfies the goal?
Budget Maximum attempts, wall time, tokens, money, files, and external actions
State What survives between runs, and which revision is authoritative?
Terminal states succeeded, failed, blocked, cancelled, expired, outcome_unknown
Escalation Which condition requires a human or a different system?

An automation that simply re-prompts until the model says “finished” is not a reliable loop. It has repetition, but no independent truth.

How the Inner and Outer Loops Compose

The outer loop may invoke many inner loops. Each attempt has its own runtime state; the durable work item stores only reviewed artifacts and evidence needed by the next attempt.

flowchart TD T["Trigger creates work item"] --> L["Lease isolated workspace"] L --> R["Run one Agent invocation"] R --> A["Persist artifact and trace"] A --> V["Independent verifier"] V -->|Goal true| S["Succeeded"] V -->|Retryable failure| B{"Budget remains?"} B -->|Yes| R B -->|No| F["Failed"] V -->|Needs decision| H["Human escalation"] V -->|Effect uncertain| U["Outcome unknown / reconcile"]

Do not pass an unlimited transcript from one attempt into the next. Persist the goal, artifact revision, verifier output, unresolved evidence, and next action. Reconstruct context deliberately so failed guesses do not become authoritative memory.

A Runnable Go Loop Controller

The example below separates an inner executor from an outer verifier. The executor can claim success, but only the verifier can move the work item to succeeded.

go
package main

import (
	"errors"
	"fmt"
)

type Result struct {
	Revision string
	Claim    string
}

type Executor func(attempt int) (Result, error)
type Verifier func(Result) (bool, string)

func RunLoop(maxAttempts int, execute Executor, verify Verifier) (string, error) {
	if maxAttempts < 1 {
		return "", errors.New("invalid attempt budget")
	}
	for attempt := 1; attempt <= maxAttempts; attempt++ {
		result, err := execute(attempt)
		if err != nil {
			continue
		}
		ok, evidence := verify(result)
		fmt.Printf("attempt=%d revision=%s verified=%t evidence=%s\n",
			attempt, result.Revision, ok, evidence)
		if ok {
			return result.Revision, nil
		}
	}
	return "", errors.New("goal not verified within budget")
}

func main() {
	revision, err := RunLoop(
		3,
		func(attempt int) (Result, error) {
			return Result{
				Revision: fmt.Sprintf("rev-%d", attempt),
				Claim:    "tests pass",
			}, nil
		},
		func(result Result) (bool, string) {
			if result.Revision != "rev-2" {
				return false, "independent tests failed"
			}
			return true, "independent tests passed"
		},
	)
	fmt.Printf("terminal_revision=%s err=%v\n", revision, err)
}

Expected final line:

text
terminal_revision=rev-2 err=<nil>

The sample is intentionally small. A real controller must use durable leases, cancellation, artifact hashes, verifier versioning, idempotency keys, and reconciliation for writes that may have completed before a timeout.

Verification Must Be Independent of the Attempt

The same model, prompt, and context that produced an artifact share its blind spots. Independence is a spectrum:

  1. deterministic compiler, parser, schema, or policy check;
  2. executable tests whose assertions were not written by the current attempt;
  3. structural or security oracle;
  4. a separately configured reviewer model;
  5. accountable human review;
  6. staged production observation.

Use the strongest affordable evidence for the risk. A verifier should bind to the exact artifact revision and return a typed result. “Looks good” and a model-reported confidence score are not completion conditions.

State, Retry, and Side-Effect Rules

An engineered loop must survive process restarts without repeating unsafe work.

  • Give every work item, attempt, run, tool call, and effect a stable identity.
  • Lease work with an owner and expiry; do not let two attempts mutate the same workspace.
  • Classify failures as retryable, permanent, blocked, cancelled, or unknown.
  • Retry only from committed state and preserve the previous evidence.
  • Make external writes idempotent where possible.
  • If a timeout happens after dispatch, enter outcome_unknown and reconcile before retrying.
  • Invalidate approval when the artifact, arguments, actor, policy, or target changes.
  • Keep secrets and raw sensitive payloads out of durable memory.

These controls belong to the controller and Agent Runtime, not to a prompt asking the model to be careful.

Where Product Improvement Fits

Evaluation, failure analysis, and release governance form a third cycle around the deployed loop:

text
sample tasks
-> classify failures
-> change prompt/context/tools/controller
-> regression evaluation
-> staged release
-> observe outcomes
-> update the task set

This lifecycle can improve both the inner runtime loop and the outer automation loop. Keep it separate in diagrams and metrics: per-run success, per-work-item convergence, and per-release quality answer different questions.

Choose the Smallest Sufficient Control Structure

Not every task needs an autonomous outer loop.

Task shape Prefer
One deterministic transformation Ordinary code
One bounded model response Prompt plus schema validation
One tool-using invocation Agent Loop
Repeated attempts against an objective verifier Engineered loop
Multiple owners, approvals, long waits, and compensation Durable workflow or graph
Ambiguous high-impact decision Human-led process

Autonomy is not a maturity score. Add repetition only when verification is stronger than the additional failure surface.

Failure Modes

  • confusing a model's final answer with verified goal completion;
  • allowing an outer loop to inherit an unbounded failed transcript;
  • retrying non-idempotent effects after an ambiguous timeout;
  • scheduling duplicate work without a stable work-item key;
  • letting two agents edit the same checkout or deploy the same target;
  • using the maker model as the only judge;
  • defining “keep improving” without a measurable terminal condition;
  • hiding attempt cost, discarded artifacts, or manual intervention;
  • treating a successful demo as evidence of safe unattended operation;
  • calling every product evaluation process “Loop Engineering.”

Frequently Asked Questions

What is the difference between an Agent Loop and Loop Engineering?

An Agent Loop is the runtime transition cycle inside one invocation. Loop Engineering specifies and operates an external loop that can trigger and verify multiple attempts toward a durable goal. The outer controller owns work identity, budgets, terminal states, and escalation.

Is Loop Engineering just another name for ReAct?

No. ReAct interleaves reasoning and action and can be used inside an Agent Loop. Loop Engineering governs the system outside that invocation: triggers, isolated execution, persistent evidence, independent verification, retry, and stop policy.

Does Loop Engineering replace prompt or harness engineering?

No. Prompts express tasks, context engineering selects evidence, and harness engineering supplies tools, state, and execution controls. Loop Engineering composes those pieces over repeated runs. A weak prompt or unsafe harness remains weak inside a loop.

What is the minimum safe loop specification?

Define a stable goal, trigger, work key, isolated workspace, independent verifier, retry classes, attempt and cost budgets, durable state, terminal outcomes, cancellation, and escalation. Add idempotency and reconciliation before the loop can perform external writes.

Can an agent decide that its own loop is complete?

It can propose completion. For meaningful effects, a separate deterministic or governed verifier should decide. Bind the decision to the exact artifact hash, test output, policy revision, and approval rather than to the agent's final message.

Summary

Agent Loop and Loop Engineering operate at different boundaries. The inner loop advances one invocation through model decisions, tools, observations, and terminal outcomes. The engineered outer loop repeatedly creates and verifies attempts against a durable goal.

Keep a third product-improvement lifecycle for evals and release learning. When these layers are named separately, teams can debug the right controller, measure the right unit, and stop unsafe repetition before it becomes automation debt.

Primary Sources