核心摘要
Agent Loop 运行在一次 Invocation 内部:反复调用模型,执行 Tool 或 Handoff,提交 Observation,并以类型化结果或限制结束。
Loop Engineering 设计外部系统:发现工作、启动一次或多次 Agent Run、验证 Artifact、记录持久状态,并决定 Retry、Release、Escalate 或 Stop。它是正在形成的工程术语,并不是已经完成标准化的学科。
区分二者很重要,因为内循环报告“完成”时,外部 Goal 仍可能为假。Agent Final Answer 之后测试可能失败,部署可能不健康,外部写入也可能处于结果未知状态。外部 Controller 才拥有这些事实。
团队最容易混淆的三种循环
AI 交付系统里至少存在三种不同 Cycle:
| 循环 | 工作单元 | Controller | 终态证据 |
|---|---|---|---|
| Agent Runtime Loop | 一次 Invocation | Runner 或 Agent Runtime | Final Output、Handoff、Pause、Error、Cancel 或 Turn Limit |
| 工程化自动循环 | 围绕持久 Goal 的多次 Attempt | Scheduler 或 Loop Controller | 独立 Verifier、Budget、Deadline、Block 或人工决策 |
| 产品改进 Lifecycle | 多个 Task 与已发布 Version | Engineering 与 Product 流程 | Eval Result、Incident、用户结果与发布决策 |
旧版文章把第三行直接等同 Loop Engineering,范围过窄。评测和发布治理非常重要,但 Loop Specification 还要规定怎样发现重复工作、隔离执行、发起尝试、验证、记忆和停止。
Agent Loop:一次 Invocation 的运行协议
Agent Loop回答:
基于已提交 State 和当前 Observation,下一次 Transition 应是什么?
OpenAI Agents SDK 记录的内循环会调用模型,然后返回 Final Output、执行 Handoff、运行 Tool Call 后继续,或在超过 Turn Limit 时抛错。Google ADK 的 Event Loop 则由 Execution Logic 产生 Event,Runner 处理并提交后,再恢复执行。这些是同一 Runtime 形态的不同框架实现。
Runtime 负责 Correlation ID、Tool Result Ordering、State Commit、Cancellation、Limit 和 Recovery。模型可以提议 Transition,但不能改写 Identity、Authorization 或已经提交的 Business State。
Loop Engineering:有边界的外部循环
Loop Engineering 回答:
怎样让一个持久系统重复调用 Agent 与 Checker,直到可验证 Goal 进入终态?
一篇 2026 年预印本用 Loop Specification 描述由 Trigger、Goal、Verification Step、Stopping Rule 与 Memory 组成的有界可复用制品。工程实践还加入 Automation、隔离 Worktree、Skill、Connector、Sub-agent 和 External State。这些来源描述的是新兴实践,不是标准组织给出的统一定义。
生产级 Loop Specification 至少包含:
| 契约 | 必须回答的问题 |
|---|---|
| Trigger | 哪个 Event 或 Schedule 创建 Work,怎样抑制重复发现? |
| Goal | 哪个可观察 State 必须变成真? |
| Work Identity | 哪个稳定 ID 绑定 Attempt、Artifact、Approval 和 Effect? |
| Isolation | Attempt 使用哪个 Branch、Worktree、Sandbox、Tenant 和 Credential? |
| Executor | 哪个 Agent Harness 可以运行,拥有哪些 Tool 和 Limit? |
| Verifier | 哪个独立检查判断 Artifact 是否满足 Goal? |
| Budget | 最大 Attempt、Wall Time、Token、Money、File 与 External Action |
| State | 哪些信息跨 Run 保留,哪个 Revision 才是权威? |
| Terminal State | succeeded、failed、blocked、cancelled、expired、outcome_unknown |
| Escalation | 哪种条件必须交给人或其他系统? |
只会反复 Prompt,直到模型声称“完成”的 Automation 并不可靠。它有重复,却没有独立真值。
内循环与外循环如何组合
外循环可以调用多个内循环。每个 Attempt 拥有独立 Runtime State;持久 Work Item 只保存下一次 Attempt 真正需要的已审 Artifact 与 Evidence。
不要把一个 Attempt 的无限 Transcript 原样塞给下一个 Attempt。持久化 Goal、Artifact Revision、Verifier Output、未决证据和 Next Action,再有意识地重建 Context,避免失败猜测变成权威 Memory。
用 Go 实现外部 Loop Controller
下面示例把内部 Executor 与外部 Verifier 分开。Executor 可以声称成功,但只有 Verifier 能把 Work Item 推进到 succeeded。
package main
import (
"errors"
"fmt"
)
type Result struct {
Revision string
Claim string
}
type Executor func(attempt int) (Result, error)
type Verifier func(Result) (bool, string)
func RunLoop(maxAttempts int, execute Executor, verify Verifier) (string, error) {
if maxAttempts < 1 {
return "", errors.New("invalid attempt budget")
}
for attempt := 1; attempt <= maxAttempts; attempt++ {
result, err := execute(attempt)
if err != nil {
continue
}
ok, evidence := verify(result)
fmt.Printf("attempt=%d revision=%s verified=%t evidence=%s\n",
attempt, result.Revision, ok, evidence)
if ok {
return result.Revision, nil
}
}
return "", errors.New("goal not verified within budget")
}
func main() {
revision, err := RunLoop(
3,
func(attempt int) (Result, error) {
return Result{
Revision: fmt.Sprintf("rev-%d", attempt),
Claim: "tests pass",
}, nil
},
func(result Result) (bool, string) {
if result.Revision != "rev-2" {
return false, "independent tests failed"
}
return true, "independent tests passed"
},
)
fmt.Printf("terminal_revision=%s err=%v\n", revision, err)
}
最终一行应为:
terminal_revision=rev-2 err=<nil>
示例刻意保持精简。真实 Controller 还需要 Durable Lease、Cancellation、Artifact Hash、Verifier Version、Idempotency Key,以及针对 Timeout 前可能已完成的外部写入进行 Reconciliation。
Verification 必须独立于 Attempt
生成 Artifact 的同一 Model、Prompt 和 Context 会共享同样的盲区。独立性可以逐级增强:
- 确定性的 Compiler、Parser、Schema 或 Policy Check;
- Assertion 不是由本次 Attempt 编写的可执行 Test;
- Structural 或 Security Oracle;
- 单独配置的 Reviewer Model;
- 可追责的 Human Review;
- 分阶段 Production Observation。
按风险采用能负担的最强证据。Verifier 必须绑定精确 Artifact Revision 并返回类型化结果。“看起来不错”和模型自报 Confidence 都不是完成条件。
State、Retry 与副作用规则
工程化循环必须能在进程重启后恢复,同时避免重复执行危险操作。
- 为 Work Item、Attempt、Run、Tool Call 和 Effect 分配稳定身份;
- 用 Owner 与 Expiry 租用工作,不让两个 Attempt 修改同一 Workspace;
- 将失败分类为 Retryable、Permanent、Blocked、Cancelled 或 Unknown;
- 只从已提交 State 重试,并保留上一轮 Evidence;
- 尽可能让外部写入幂等;
- Dispatch 后 Timeout 时进入
outcome_unknown,对账后才能重试; - Artifact、Argument、Actor、Policy 或 Target 变化时使 Approval 失效;
- 不把 Secret 与完整敏感 Payload 写入持久 Memory。
这些控制属于 Controller 和 Agent Runtime,不属于一句“请谨慎执行”的 Prompt。
产品改进 Lifecycle 放在哪里
评测、失败分析和发布治理构成部署 Loop 外面的第三种 Cycle:
抽样任务
-> 分类失败
-> 修改 Prompt/Context/Tool/Controller
-> 回归评测
-> 分阶段发布
-> 观察结果
-> 更新任务集
这套 Lifecycle 可以同时改进内部 Runtime Loop 与外部 Automation Loop。架构图与指标应把三者分开:每次 Run 的成功、每个 Work Item 的收敛,以及每个 Release 的质量,回答的是不同问题。
选择足够用的最小控制结构
并非所有任务都需要自主外循环。
| 任务形态 | 优先选择 |
|---|---|
| 一次确定性转换 | 普通代码 |
| 一次有边界的模型响应 | Prompt 加 Schema Validation |
| 一次使用 Tool 的 Invocation | Agent Loop |
| 针对客观 Verifier 的多次 Attempt | 工程化 Loop |
| 多 Ownership、审批、长等待和补偿 | Durable Workflow 或 Graph |
| 含糊且高影响的决策 | 人主导流程 |
Autonomy 不是成熟度分数。只有当 Verification 强于新增 Failure Surface 时,才增加重复执行。
常见失败模式
- 把模型 Final Answer 当作 Goal 已验证;
- 让外循环继承无限增长的失败 Transcript;
- 在 Timeout 结果不明时重试非幂等 Effect;
- 没有稳定 Work Item Key 就重复调度;
- 让两个 Agent 修改同一 Checkout 或部署同一 Target;
- 让 Maker Model 成为唯一 Judge;
- 用“持续改进”代替可测量 Terminal Condition;
- 隐藏 Attempt Cost、被丢弃 Artifact 或人工介入;
- 把成功 Demo 当作无人值守安全证据;
- 把所有产品评测流程都称为 Loop Engineering。
常见问题
Agent Loop 与 Loop Engineering 有什么区别?
Agent Loop 是一次 Invocation 内部的 Runtime Transition Cycle。Loop Engineering 则定义并运行一个外部循环,针对持久 Goal 触发和验证多次 Attempt。外部 Controller 拥有 Work Identity、Budget、Terminal State 和 Escalation。
Loop Engineering 只是 ReAct 的另一个名称吗?
不是。ReAct 让推理与行动交替,可以作为 Agent Loop 的内部模式。Loop Engineering 管理 Invocation 之外的 Trigger、隔离执行、持久 Evidence、独立 Verification、Retry 与 Stop Policy。
Loop Engineering 会取代 Prompt Engineering 或 Harness Engineering 吗?
不会。Prompt 表达任务,Context Engineering 选择证据,Harness Engineering 提供 Tool、State 和执行控制;Loop Engineering 在多次 Run 上组合这些能力。弱 Prompt 或不安全 Harness 放进循环后仍然弱。
最小安全 Loop Specification 应包含什么?
定义稳定 Goal、Trigger、Work Key、隔离 Workspace、独立 Verifier、Retry 分类、Attempt 与 Cost Budget、Durable State、Terminal Outcome、Cancellation 和 Escalation。外部 Effect 还要有 Idempotency 与 Reconciliation。
Agent 能否自己决定循环已经完成?
它可以提议完成。对于有实质影响的结果,应由独立确定性检查或受治理的 Reviewer 决策,并将结论绑定 Artifact Hash、测试输出、Policy Revision 与 Approval,而不是 Agent Final Message。
总结
Agent Loop 与 Loop Engineering 工作在不同边界。内循环通过模型决策、Tool、Observation 与 Terminal Outcome 推进一次 Invocation;工程化外循环则围绕持久 Goal 重复创建并验证 Attempt。
把 Eval 和 Release Learning 留给第三层产品改进 Lifecycle。只有分开命名这些层,团队才能调试正确的 Controller、测量正确的 Unit,并在不安全重复演变成自动化债务前停止。
一手来源
- ReAct:协同语言模型的推理与行动
- OpenAI Agents SDK:Agent Loop
- Google ADK Runtime Event Loop
- Loop Engineering 工程实践
- Stop Hand-Holding Your Coding Agent:Loop Specification