TL;DR
An Agent Loop runs inside one invocation. It repeatedly calls a model, executes tools or handoffs, commits observations, and stops with a typed result or limit.
Loop Engineering designs the external system that discovers work, launches one or more agent runs, verifies artifacts, records durable state, and decides whether to retry, release, escalate, or stop. It is an emerging engineering label rather than a finished standard.
The distinction matters because an inner loop can report “done” while the outer goal is still false. A test may fail after the agent's final answer, a deployment may be unhealthy, or a write may have an unknown outcome. The outer controller owns those facts.
Three Loops That Teams Commonly Confuse
There are at least three different cycles in an AI delivery system:
| Cycle | Unit of work | Controller | Terminal evidence |
|---|---|---|---|
| Agent runtime loop | One invocation | Runner or Agent Runtime | Final output, handoff, pause, error, cancellation, or turn limit |
| Engineered automation loop | Repeated attempts toward one durable goal | Scheduler or Loop Controller | Independent verifier, budget, deadline, block, or human decision |
| Product improvement lifecycle | Many tasks and released versions | Engineering and product process | Eval results, incidents, user outcomes, and release decision |
The previous version of this article treated the third row as Loop Engineering itself. That is too narrow. Evaluation and release governance are essential, but a loop specification also governs how recurring work is found, isolated, attempted, checked, remembered, and stopped.
Agent Loop: One Invocation's Runtime Protocol
The Agent Loop answers:
Given the committed state and current observation, what transition happens next?
OpenAI's Agents SDK documents an inner loop that calls the model, then either returns final output, performs a handoff, executes tool calls and loops again, or raises after a turn limit. Google ADK documents an event loop in which execution yields an event, the Runner processes and commits it, and execution resumes afterward. These are framework-specific implementations of a shared runtime shape.
The runtime owns correlation IDs, tool-result ordering, state commitment, cancellation, limits, and recovery. The model proposes a transition; it does not get to rewrite identity, authorization, or committed business state.
Loop Engineering: A Bounded External Loop
Loop Engineering answers:
What durable system can repeatedly invoke agents and checks until a verifiable goal reaches a terminal state?
A 2026 preprint proposes the term loop specification for a bounded reusable artifact containing a trigger, goal, verification step, stopping rule, and memory. Practitioner descriptions add automation, isolated worktrees, skills, connectors, subagents, and external state. These sources describe an emerging practice, not a universal standards-body definition.
A production loop specification should include:
| Contract | Required question |
|---|---|
| Trigger | What event or schedule creates work, and how is duplicate discovery suppressed? |
| Goal | Which observable state must become true? |
| Work identity | Which stable ID binds attempts, artifacts, approvals, and side effects? |
| Isolation | Which branch, worktree, sandbox, tenant, and credentials belong to the attempt? |
| Executor | Which Agent Harness may run, with which tools and limits? |
| Verifier | Which independent check decides whether the artifact satisfies the goal? |
| Budget | Maximum attempts, wall time, tokens, money, files, and external actions |
| State | What survives between runs, and which revision is authoritative? |
| Terminal states | succeeded, failed, blocked, cancelled, expired, outcome_unknown |
| Escalation | Which condition requires a human or a different system? |
An automation that simply re-prompts until the model says “finished” is not a reliable loop. It has repetition, but no independent truth.
How the Inner and Outer Loops Compose
The outer loop may invoke many inner loops. Each attempt has its own runtime state; the durable work item stores only reviewed artifacts and evidence needed by the next attempt.
Do not pass an unlimited transcript from one attempt into the next. Persist the goal, artifact revision, verifier output, unresolved evidence, and next action. Reconstruct context deliberately so failed guesses do not become authoritative memory.
A Runnable Go Loop Controller
The example below separates an inner executor from an outer verifier. The executor can claim success, but only the verifier can move the work item to succeeded.
package main
import (
"errors"
"fmt"
)
type Result struct {
Revision string
Claim string
}
type Executor func(attempt int) (Result, error)
type Verifier func(Result) (bool, string)
func RunLoop(maxAttempts int, execute Executor, verify Verifier) (string, error) {
if maxAttempts < 1 {
return "", errors.New("invalid attempt budget")
}
for attempt := 1; attempt <= maxAttempts; attempt++ {
result, err := execute(attempt)
if err != nil {
continue
}
ok, evidence := verify(result)
fmt.Printf("attempt=%d revision=%s verified=%t evidence=%s\n",
attempt, result.Revision, ok, evidence)
if ok {
return result.Revision, nil
}
}
return "", errors.New("goal not verified within budget")
}
func main() {
revision, err := RunLoop(
3,
func(attempt int) (Result, error) {
return Result{
Revision: fmt.Sprintf("rev-%d", attempt),
Claim: "tests pass",
}, nil
},
func(result Result) (bool, string) {
if result.Revision != "rev-2" {
return false, "independent tests failed"
}
return true, "independent tests passed"
},
)
fmt.Printf("terminal_revision=%s err=%v\n", revision, err)
}
Expected final line:
terminal_revision=rev-2 err=<nil>
The sample is intentionally small. A real controller must use durable leases, cancellation, artifact hashes, verifier versioning, idempotency keys, and reconciliation for writes that may have completed before a timeout.
Verification Must Be Independent of the Attempt
The same model, prompt, and context that produced an artifact share its blind spots. Independence is a spectrum:
- deterministic compiler, parser, schema, or policy check;
- executable tests whose assertions were not written by the current attempt;
- structural or security oracle;
- a separately configured reviewer model;
- accountable human review;
- staged production observation.
Use the strongest affordable evidence for the risk. A verifier should bind to the exact artifact revision and return a typed result. “Looks good” and a model-reported confidence score are not completion conditions.
State, Retry, and Side-Effect Rules
An engineered loop must survive process restarts without repeating unsafe work.
- Give every work item, attempt, run, tool call, and effect a stable identity.
- Lease work with an owner and expiry; do not let two attempts mutate the same workspace.
- Classify failures as retryable, permanent, blocked, cancelled, or unknown.
- Retry only from committed state and preserve the previous evidence.
- Make external writes idempotent where possible.
- If a timeout happens after dispatch, enter
outcome_unknownand reconcile before retrying. - Invalidate approval when the artifact, arguments, actor, policy, or target changes.
- Keep secrets and raw sensitive payloads out of durable memory.
These controls belong to the controller and Agent Runtime, not to a prompt asking the model to be careful.
Where Product Improvement Fits
Evaluation, failure analysis, and release governance form a third cycle around the deployed loop:
sample tasks
-> classify failures
-> change prompt/context/tools/controller
-> regression evaluation
-> staged release
-> observe outcomes
-> update the task set
This lifecycle can improve both the inner runtime loop and the outer automation loop. Keep it separate in diagrams and metrics: per-run success, per-work-item convergence, and per-release quality answer different questions.
Choose the Smallest Sufficient Control Structure
Not every task needs an autonomous outer loop.
| Task shape | Prefer |
|---|---|
| One deterministic transformation | Ordinary code |
| One bounded model response | Prompt plus schema validation |
| One tool-using invocation | Agent Loop |
| Repeated attempts against an objective verifier | Engineered loop |
| Multiple owners, approvals, long waits, and compensation | Durable workflow or graph |
| Ambiguous high-impact decision | Human-led process |
Autonomy is not a maturity score. Add repetition only when verification is stronger than the additional failure surface.
Failure Modes
- confusing a model's final answer with verified goal completion;
- allowing an outer loop to inherit an unbounded failed transcript;
- retrying non-idempotent effects after an ambiguous timeout;
- scheduling duplicate work without a stable work-item key;
- letting two agents edit the same checkout or deploy the same target;
- using the maker model as the only judge;
- defining “keep improving” without a measurable terminal condition;
- hiding attempt cost, discarded artifacts, or manual intervention;
- treating a successful demo as evidence of safe unattended operation;
- calling every product evaluation process “Loop Engineering.”
Frequently Asked Questions
What is the difference between an Agent Loop and Loop Engineering?
An Agent Loop is the runtime transition cycle inside one invocation. Loop Engineering specifies and operates an external loop that can trigger and verify multiple attempts toward a durable goal. The outer controller owns work identity, budgets, terminal states, and escalation.
Is Loop Engineering just another name for ReAct?
No. ReAct interleaves reasoning and action and can be used inside an Agent Loop. Loop Engineering governs the system outside that invocation: triggers, isolated execution, persistent evidence, independent verification, retry, and stop policy.
Does Loop Engineering replace prompt or harness engineering?
No. Prompts express tasks, context engineering selects evidence, and harness engineering supplies tools, state, and execution controls. Loop Engineering composes those pieces over repeated runs. A weak prompt or unsafe harness remains weak inside a loop.
What is the minimum safe loop specification?
Define a stable goal, trigger, work key, isolated workspace, independent verifier, retry classes, attempt and cost budgets, durable state, terminal outcomes, cancellation, and escalation. Add idempotency and reconciliation before the loop can perform external writes.
Can an agent decide that its own loop is complete?
It can propose completion. For meaningful effects, a separate deterministic or governed verifier should decide. Bind the decision to the exact artifact hash, test output, policy revision, and approval rather than to the agent's final message.
Summary
Agent Loop and Loop Engineering operate at different boundaries. The inner loop advances one invocation through model decisions, tools, observations, and terminal outcomes. The engineered outer loop repeatedly creates and verifies attempts against a durable goal.
Keep a third product-improvement lifecycle for evals and release learning. When these layers are named separately, teams can debug the right controller, measure the right unit, and stop unsafe repetition before it becomes automation debt.
Primary Sources
- ReAct: Synergizing Reasoning and Acting in Language Models
- OpenAI Agents SDK: The Agent Loop
- Google ADK Runtime Event Loop
- Loop Engineering
- Stop Hand-Holding Your Coding Agent: Engineering the Loops