TL;DR

Multi-agent orchestration is not a choice between three fashionable diagrams. It is a control system that owns transitions, state, budgets, effects, termination, and recovery. Start with a single agent or deterministic workflow. Add a supervisor when one component should retain control, a handoff when a specialist should own the next turn, hierarchy when independently governed subteams need nested coordination, and fan-out only when work is truly independent. Keep model decisions inside code-enforced boundaries.

Table of Contents

  1. Key Takeaways
  2. Start With the Smallest Control Plane
  3. Separate Topology From Execution Flow
  4. The Main Orchestration Patterns
  5. Define the Control Contract Before Prompts
  6. Runnable Go Orchestrator
  7. Failure Recovery and Side Effects
  8. Security Boundaries
  9. Evaluate the System Not the Demo
  10. Map the Contract to Current Frameworks
  11. A Production Migration Path
  12. FAQ
  13. Summary
  14. Primary Sources

Key Takeaways

  • Multi-agent is an optimization, not a starting requirement. Compare it with a single-agent and deterministic baseline.
  • Topology and execution flow are different choices. A supervisor can run a pipeline, router, loop, or parallel fan-out.
  • Handoff is a transfer of logical control. It does not remove the runtime, state store, or model provider as failure domains.
  • The contract matters more than the prompt. State ownership, allowed transitions, termination, budgets, and effect semantics must be machine-enforced.
  • Retries are unsafe without idempotency. A timeout after an external action creates an unknown outcome, not proof that nothing happened.
  • Evaluate coordination-specific failures. Measure routing, missing or duplicate work, loops, state loss, unsafe effects, tail latency, cost, and escalation.

This article focuses on orchestration topology and transition contracts. For the earlier decision about whether specialization is justified at all, read Multi-Agent Systems: When and How to Build Them.

Start With the Smallest Control Plane

The default architecture should be the least dynamic system that can satisfy the task. Every extra agent adds another prompt, context boundary, failure surface, latency source, and authorization decision.

Use this progression:

Baseline Use it when Add more control only when
Deterministic code Steps and dependencies are known A step requires semantic judgment that rules cannot express reliably
One agent with tools One context can hold the task and tool set Tool selection, context, or permissions become too broad
One agent plus deterministic workflow Most steps are fixed, with a few model decisions The next owner cannot be known in advance
Multiple agents Roles need independent context, tools, policy, or parallel ownership Measurements show better task outcomes than the simpler baseline

This ordering is not anti-agent. It isolates where probabilistic reasoning creates value. A fixed parser, authorization check, payment write, merge function, or termination counter should remain code even when an AI agent proposes the surrounding plan.

Agent count is not a sizing metric. Two agents with cyclic handoffs and write access to production can be harder to operate than twenty read-only workers behind a deterministic fan-out. Select a pattern from dependency shape, control ownership, risk, and evaluation evidence rather than fixed thresholds such as "five agents means swarm."

Separate Topology From Execution Flow

Topology answers who may decide; execution flow answers how work progresses. Treating them as the same decision produces misleading labels and brittle migrations.

flowchart LR A["User objective"] --> B["Control owner"] B --> C["Execution flow"] C --> D["State and effect runtime"] D --> E["Verified result or escalation"] B --> B1["Code"] B --> B2["Supervisor"] B --> B3["Active specialist"] B --> B4["Nested managers"] C --> C1["Pipeline"] C --> C2["Router"] C --> C3["Fan-out and merge"] C --> C4["Bounded loop"]

The choices are independent:

  • Control owner: deterministic code, a central supervisor, the currently active specialist, or nested managers.
  • Execution flow: sequential pipeline, conditional router, parallel fan-out, or bounded review loop.
  • Runtime semantics: state persistence, cancellation, retries, idempotency, authorization, and observability.

A "supervisor architecture" can still dispatch independent workers in parallel. A handoff network can still run inside one centralized process. A hierarchy can contain deterministic pipelines inside each team. The popular term "swarm" is especially ambiguous: it may mean peer handoffs, group conversation, decentralized planning, or merely many workers. Name the actual control and transition semantics in design documents.

The Main Orchestration Patterns

The useful pattern catalog contains four primitives plus hybrids, not a universal ranking from simple to enterprise.

Pattern Who chooses the next step State sent to workers Primary strength Primary risk Good fit
Code-orchestrated pipeline or router Application code Typed task slice Predictability and testability Rules become rigid Stable workflows, regulated effects
Supervisor with agents as tools Supervisor retains control Per-call scoped context Central synthesis and policy Bottleneck and supervisor bias Research, analysis, delegated subtasks
Handoff network Active agent transfers control Handoff envelope plus selected history Specialist owns the conversation Loops, context drift, unclear ownership Support triage, domain escalation
Hierarchical or nested teams Managers coordinate subteams Team-local state and artifacts Policy and context isolation Coordination overhead and hidden loss Independently governed domains
Hybrid Code constrains model-selected edges Explicit artifacts Flexibility within hard bounds More semantics to test Most production systems

Code-orchestrated pipelines routers and fan-out

Code orchestration is the safest default when dependencies are known. The model may classify an intent or produce an artifact, but code selects the allowed branch, starts parallel work, applies timeouts, and merges results.

Use a pipeline when each stage consumes a defined predecessor artifact. Use a router when exactly one or a bounded subset of workers should run. Use fan-out and merge when subtasks are independent and the merge rule is explicit. Do not parallelize work that writes the same resource or depends on a shared mutable conversation.

Chain orchestration explains fixed stage contracts, while workflow orchestration covers durable scheduling and recovery. Neither requires every node to be an agent.

Supervisor with agents as tools

A supervisor agent retains ownership of the task and invokes specialists like tools. The specialists return artifacts; they do not become the active conversational owner.

This pattern works when:

  • one component must synthesize the final answer;
  • workers should not see the entire conversation;
  • a central policy decides which specialist and tool scopes are available;
  • the task benefits from iterative delegation or review.

The supervisor must not be an unconstrained superuser. Code should validate its proposed route, cap calls, filter worker context, and require typed outputs. If every request follows the same path, replace the model router with code.

Handoff networks

A handoff transfers active ownership to a specialist. The target receives a structured reason, the current objective, selected state, and a bounded history view. It then answers the user or performs another allowed handoff.

This pattern fits conversations where the specialist needs direct control, such as billing support transferring to technical support after discovering a product defect. It does not imply there is no central runtime or single point of failure. A runner still persists state, enforces the transition graph, applies guardrails, and records the trace.

A safe handoff is closer to an API call than forwarding an entire chat transcript. The target should receive the minimum context needed for its role, and both the source and destination must be authorized for the transition.

Hierarchical and nested teams

Hierarchy is useful when subteams have independent policy, context, tools, or lifecycle, not simply because the system has many agents. A top-level manager should exchange compact artifacts with a subteam manager instead of copying every worker transcript upward.

Good boundaries include separate security domains, data regions, business units, or long-lived specialist workflows. Poor boundaries imitate an organization chart without reducing context or policy complexity. Each management layer adds routing decisions and can compress away evidence, so measure whether the layer improves decomposition and verification.

Hybrid orchestration

Most production systems are hybrids: deterministic code owns the graph, a model proposes one of a small set of routes, a supervisor delegates analysis, selected workers run in parallel, and a human approves sensitive effects.

flowchart TD A["Validated request"] --> B{"Code router"} B -->|"Known path"| C["Deterministic workflow"] B -->|"Open-ended analysis"| D["Bounded supervisor"] D --> E["Read-only worker A"] D --> F["Read-only worker B"] E --> G["Deterministic merge"] F --> G G --> H{"Sensitive effect"} H -->|"No"| I["Return result"] H -->|"Yes"| J["Human approval"] J --> K["Idempotent executor"]

The key boundary is simple: the model may propose; trusted code authorizes, persists, executes, and terminates.

Define the Control Contract Before Prompts

A production orchestration contract specifies what may change at every transition. Prompts can explain a role, but they cannot reliably enforce invariants.

Task and artifact contract

Define a stable envelope with these fields:

Field Purpose Required invariant
run_id and task_id Correlation and replay Immutable across the run
from and to Proposed control transfer Edge exists in an allowlist
goal Current bounded objective Cannot silently expand scope
artifact_refs or typed artifacts Evidence passed between roles Schema and provenance validated
hop and depth Loop and recursion control Hard maximum enforced by code
remaining_budget Token, call, cost, or work budget Monotonically decreases
deadline End-to-end time bound Propagated to every child
idempotency_key Stable identity for an effect Unique at the effect boundary
principal and policy context Authorization Never inferred from agent text
status and reason Completion or escalation Must match an allowed terminal state

Prefer artifact references or task-scoped summaries over a shared transcript. A researcher owns evidence artifacts; a reviewer owns review findings; the orchestrator owns routing and budgets. Avoid two agents concurrently writing the same field. Version state updates or use compare-and-swap when concurrent workers must contribute to one aggregate.

For JSON-based envelopes, JSON Formatter helps inspect captured payloads, and JSON Schema Generator can bootstrap a schema from a representative sample. Production validation still belongs in the service boundary and must reject unknown or malformed fields.

Transition and termination contract

Do not ask an LLM to "stop when done" as the only termination rule. Enforce:

  1. an allowlisted directed graph of valid source-to-destination transitions;
  2. maximum hops, nested depth, model calls, tokens, cost, and wall-clock deadline;
  3. a terminal predicate that requires the expected artifact type;
  4. no-progress detection, such as repeated state hashes or unchanged required fields;
  5. an explicit failed, partial, or needs_human terminal path.

Budgets should be reserved before dispatch. Otherwise parallel workers can each observe the same remaining budget and collectively overspend it.

Context contract

Context engineering is access control as well as token management. For each agent, define:

  • instructions it may see;
  • conversation turns and artifacts it may read;
  • tools and resource scopes it may invoke;
  • fields it may write;
  • secrets that must be redacted;
  • retention and deletion policy.

Treat every peer artifact as untrusted input. Do not paste a worker's output into a higher-priority instruction channel. Preserve provenance so the final synthesizer can distinguish user input, retrieved evidence, model claims, tool results, and policy decisions.

Runnable Go Orchestrator

The following standard-library Go program implements a bounded supervisor flow. The worker functions are deterministic stand-ins for model or tool adapters, which keeps the control behavior reproducible. The engine, not an agent prompt, enforces allowed transitions, typed artifacts, hop and work budgets, context deadlines, state copying, and in-process deduplication of successful steps.

go
package main

import (
	"context"
	"errors"
	"fmt"
	"sync"
	"time"
)

type Artifact struct {
	Kind   string
	Body   string
	Source string
}

type Handoff struct {
	RunID           string
	TaskID          string
	From            string
	To              string
	Goal            string
	Hop             int
	RemainingBudget int
	Artifacts       []Artifact
}

type Outcome struct {
	Artifact Artifact
	Next     string
	Cost     int
	Done     bool
}

type Agent func(context.Context, Handoff) (Outcome, error)

type journalEntry struct {
	done    chan struct{}
	outcome Outcome
	err     error
}

type Journal struct {
	mu      sync.Mutex
	entries map[string]*journalEntry
}

func NewJournal() *Journal {
	return &Journal{entries: make(map[string]*journalEntry)}
}

func (j *Journal) Do(
	ctx context.Context,
	key string,
	fn func() (Outcome, error),
) (Outcome, error) {
	j.mu.Lock()
	if entry, ok := j.entries[key]; ok {
		j.mu.Unlock()
		select {
		case <-entry.done:
			return entry.outcome, entry.err
		case <-ctx.Done():
			return Outcome{}, ctx.Err()
		}
	}

	entry := &journalEntry{done: make(chan struct{})}
	j.entries[key] = entry
	j.mu.Unlock()

	entry.outcome, entry.err = fn()

	j.mu.Lock()
	if entry.err != nil {
		delete(j.entries, key)
	}
	close(entry.done)
	j.mu.Unlock()
	return entry.outcome, entry.err
}

type Engine struct {
	Agents  map[string]Agent
	Allowed map[string]map[string]bool
	MaxHops int
	Journal *Journal
}

func (e Engine) Run(ctx context.Context, h Handoff) (Artifact, error) {
	if h.RunID == "" || h.TaskID == "" || h.Goal == "" {
		return Artifact{}, errors.New("run_id, task_id, and goal are required")
	}

	for {
		if err := ctx.Err(); err != nil {
			return Artifact{}, fmt.Errorf("run deadline: %w", err)
		}
		if h.Hop >= e.MaxHops {
			return Artifact{}, fmt.Errorf("maximum hops reached: %d", e.MaxHops)
		}
		if h.RemainingBudget <= 0 {
			return Artifact{}, errors.New("work budget exhausted")
		}
		if !e.Allowed[h.From][h.To] {
			return Artifact{}, fmt.Errorf("transition denied: %s -> %s", h.From, h.To)
		}

		agent, ok := e.Agents[h.To]
		if !ok {
			return Artifact{}, fmt.Errorf("unknown agent: %s", h.To)
		}

		input := h
		input.Artifacts = append([]Artifact(nil), h.Artifacts...)
		stepKey := fmt.Sprintf("%s/%s/%d/%s", h.RunID, h.TaskID, h.Hop, h.To)

		outcome, err := e.Journal.Do(ctx, stepKey, func() (Outcome, error) {
			return agent(ctx, input)
		})
		if err != nil {
			return Artifact{}, fmt.Errorf("agent %s: %w", h.To, err)
		}
		if outcome.Cost <= 0 || outcome.Cost > h.RemainingBudget {
			return Artifact{}, fmt.Errorf("invalid cost from %s: %d", h.To, outcome.Cost)
		}
		if outcome.Artifact.Kind == "" || outcome.Artifact.Body == "" {
			return Artifact{}, fmt.Errorf("invalid artifact from %s", h.To)
		}

		remaining := h.RemainingBudget - outcome.Cost
		fmt.Printf(
			"hop=%d agent=%s artifact=%s remaining=%d\n",
			h.Hop,
			h.To,
			outcome.Artifact.Kind,
			remaining,
		)

		if outcome.Done {
			if outcome.Next != "" || outcome.Artifact.Kind != "final" {
				return Artifact{}, errors.New("invalid terminal outcome")
			}
			return outcome.Artifact, nil
		}
		if outcome.Next == "" {
			return Artifact{}, errors.New("non-terminal outcome has no next agent")
		}

		h.Artifacts = append(h.Artifacts, outcome.Artifact)
		h.From, h.To = h.To, outcome.Next
		h.Hop++
		h.RemainingBudget = remaining
	}
}

func hasArtifact(artifacts []Artifact, kind string) bool {
	for _, artifact := range artifacts {
		if artifact.Kind == kind {
			return true
		}
	}
	return false
}

func main() {
	researcher := func(ctx context.Context, h Handoff) (Outcome, error) {
		if err := ctx.Err(); err != nil {
			return Outcome{}, err
		}
		return Outcome{
			Artifact: Artifact{
				Kind:   "evidence",
				Body:   "primary sources collected",
				Source: "researcher",
			},
			Next: "reviewer",
			Cost: 2,
		}, nil
	}

	reviewer := func(ctx context.Context, h Handoff) (Outcome, error) {
		if !hasArtifact(h.Artifacts, "evidence") {
			return Outcome{}, errors.New("evidence artifact is required")
		}
		return Outcome{
			Artifact: Artifact{
				Kind:   "review",
				Body:   "evidence and scope approved",
				Source: "reviewer",
			},
			Next: "supervisor",
			Cost: 1,
		}, nil
	}

	supervisor := func(ctx context.Context, h Handoff) (Outcome, error) {
		if !hasArtifact(h.Artifacts, "evidence") ||
			!hasArtifact(h.Artifacts, "review") {
			return Outcome{}, errors.New("verified evidence is required")
		}
		return Outcome{
			Artifact: Artifact{
				Kind:   "final",
				Body:   "architecture brief ready",
				Source: "supervisor",
			},
			Cost: 1,
			Done: true,
		}, nil
	}

	engine := Engine{
		Agents: map[string]Agent{
			"researcher": researcher,
			"reviewer":   reviewer,
			"supervisor": supervisor,
		},
		Allowed: map[string]map[string]bool{
			"supervisor": {"researcher": true},
			"researcher": {"reviewer": true},
			"reviewer":   {"supervisor": true},
		},
		MaxHops: 4,
		Journal: NewJournal(),
	}

	ctx, cancel := context.WithTimeout(context.Background(), 2*time.Second)
	defer cancel()

	final, err := engine.Run(ctx, Handoff{
		RunID:           "run-42",
		TaskID:          "architecture-brief",
		From:            "supervisor",
		To:              "researcher",
		Goal:            "produce a source-backed architecture brief",
		RemainingBudget: 4,
	})
	if err != nil {
		fmt.Println("run failed:", err)
		return
	}
	fmt.Println("final:", final.Body)
}

Expected output:

text
hop=0 agent=researcher artifact=evidence remaining=2
hop=1 agent=reviewer artifact=review remaining=1
hop=2 agent=supervisor artifact=final remaining=0
final: architecture brief ready

The example intentionally keeps model calls outside the core. In a real adapter, parse model output into Outcome, reject unknown fields and destinations, and let the engine make the transition. Its journal only deduplicates a successful step while this process is alive; it cannot make an external effect exactly once. Use a durable effect record, a unique idempotency key, and provider-side reconciliation before handling payments, messages, tickets, or other external effects.

Failure Recovery and Side Effects

Recovery policy must depend on what is known about the failed step. A generic "retry three times" rule can repeat purchases, emails, or writes.

Failure class What is known Correct response
Rejected before execution No effect started Fix input, reroute, or fail
Transient read failure Operation is read-only Retry with deadline and backoff
Invalid semantic output Model returned unusable content Repair once, replan, or escalate
Confirmed effect failure Effect did not commit Retry with the same idempotency key
Unknown effect outcome Timeout occurred after dispatch Query by idempotency key; do not issue a new effect
Policy denial Requested action is forbidden Stop or request authorized human approval
Budget or deadline exhausted Run may be incomplete Return a typed partial result or escalate

Make side effects explicit

Separate reasoning from effects:

  1. an agent proposes an action with typed arguments;
  2. policy code authorizes the principal, resource, and limits;
  3. an effect executor records an idempotency key;
  4. the executor performs or reconciles the action;
  5. the result becomes an immutable artifact.

For multi-step effects, prefer a saga with explicit compensation over pretending the whole workflow is one transaction. Compensation is a business action, not a database rollback: refunding a payment, closing a duplicate ticket, or sending a correction can fail and needs its own audit record.

Define fan-out completion before dispatch

Parallel branches need a completion policy:

  • all: every required branch must succeed;
  • quorum: a defined threshold is sufficient;
  • best effort: return successful artifacts plus typed failures;
  • first valid: accept the first result that passes validation and cancel the rest.

The merge should be deterministic where possible. Deduplicate by stable task and artifact IDs, preserve provenance, and propagate cancellation to children. A model may summarize merged evidence, but it should not decide silently whether a missing mandatory branch matters.

Persist checkpoints at semantic boundaries, not after every token. A useful checkpoint contains the run version, current node, accepted artifacts, reserved budget, pending effects, and next allowed actions. See AI Agent Observability for trace and cost event design.

Security Boundaries

Every transition is a trust-boundary crossing. The receiving agent must not inherit authority merely because another agent requested a transfer.

Apply these controls:

  • carry the authenticated principal and policy context separately from natural-language messages;
  • grant each agent the minimum tool, resource, tenant, region, and spend scope;
  • authorize handoffs and tool calls in trusted code before side effects;
  • treat peer outputs, retrieved content, and tool results as untrusted data;
  • keep secrets out of shared transcripts and summaries;
  • require human approval for high-impact or irreversible actions;
  • record policy decision, agent proposal, approved arguments, effect result, and actor.

A supervisor should not have every worker's credentials. Give it delegation authority over task types, while the executor independently verifies the user and resource authorization. AI Agent Tool Security covers permission design and tool-result poisoning in depth.

OpenAI's handoff documentation also highlights an important lifecycle detail: authorization or data loading required by a handoff should happen before the target performs effects, and input guardrails attached only to the initial agent do not automatically cover every later agent. Treat guardrail scope as an implementation property to verify, not a blanket framework guarantee.

Evaluate the System Not the Demo

Multi-agent evaluation must compare end-to-end outcomes against simpler baselines and isolate coordination failures. A successful transcript is anecdotal evidence, not a release gate.

Build an evaluation set containing routine tasks, ambiguous routes, missing evidence, conflicting artifacts, repeated handoffs, tool failures, slow workers, malicious peer output, and unknown effect outcomes. Record expected routes, required artifacts, forbidden actions, acceptable terminal states, and human-escalation conditions.

Measure at least:

Dimension Example metric
Task outcome Success rate and rubric score versus single-agent and deterministic baselines
Routing Correct destination, unnecessary delegation, and quality loss versus the best valid route
Coverage Missing required subtasks or artifacts
Duplication Duplicate work and duplicate external effects
Termination Loop rate, max-hop exhaustion, and no-progress exits
State integrity Lost, stale, conflicting, or overexposed context
Safety Unauthorized proposals, blocked effects, and approval bypasses
Reliability Retry recovery, partial-result quality, and escalation precision
Efficiency Model calls, tokens, tool calls, cost, and p50/p95/p99 latency

The MAST study analyzed more than 1,600 multi-agent traces across seven frameworks and reported fourteen failure modes spanning system design, inter-agent misalignment, and task verification. That evidence is useful as a test taxonomy, not proof that any fixed architecture or framework fails at a universal rate.

Run fault injection deliberately: time out a worker after an effect, return a malformed artifact, repeat an old handoff, deny a tool permission, cancel a parent while children run, and make two branches conflict. Release only when the runtime reaches the expected terminal state without duplicate effects.

Map the Contract to Current Frameworks

Frameworks package useful primitives, but they do not replace application policy. The mapping below reflects the linked official documentation; APIs can change, so verify the current version before implementation.

Framework Relevant primitives Design implication
OpenAI Agents SDK Agents as tools, handoffs, code-driven orchestration Choose whether the manager retains control or transfers it; filter handoff input and verify guardrail scope
LangChain multi-agent Subagents, handoffs, skills, router, custom workflow Select a pattern from context and control needs; a single agent may still be sufficient
AutoGen AgentChat Round-robin, selector, swarm, and Magentic-One teams; termination conditions Team presets still require explicit stop conditions and task-specific evaluation
CrewAI Sequential and hierarchical processes Hierarchical mode requires a manager model or manager agent; process choice does not define your effect semantics

OpenAI's original Swarm repository describes itself as experimental and educational and points production users to the Agents SDK. It remains useful for understanding handoffs, but new production examples should target the supported SDK rather than treating Swarm as the current production API.

Magentic-One demonstrates an orchestrator that plans, tracks progress, and replans across specialist agents. Its published benchmark results are evidence for that system under its tested tasks and environment, not a general guarantee that hierarchy outperforms simpler workflows.

A Production Migration Path

A safe rollout adds autonomy only after the surrounding contract is observable.

  1. Establish baselines. Implement the task with deterministic code and one agent. Capture quality, cost, latency, and error classes.
  2. Extract typed boundaries. Define task, artifact, transition, terminal-state, and effect schemas before adding roles.
  3. Add one specialization. Delegate the clearest bounded task with read-only tools and compare outcomes.
  4. Constrain dynamic routing. Let a model propose only allowlisted destinations; keep budgets and authorization in code.
  5. Persist durable state. Add checkpoints, leases, idempotency records, cancellation, and replay-safe recovery.
  6. Test hostile failures. Inject loops, stale context, malformed outputs, partial fan-out, timeouts, and policy denials.
  7. Introduce sensitive effects last. Require approval, narrow credentials, reconciliation, and a kill switch.
  8. Expand only on evidence. Add handoffs, parallelism, or hierarchy when evaluation shows a specific measurable gain.

Do not make topology hot-switchable in the middle of an in-flight run unless the run records a versioned graph and migration semantics. Resume against the graph version that created the checkpoint, or perform an explicit state migration.

FAQ

What is multi-agent orchestration

Multi-agent orchestration is the control layer that coordinates specialized agents. It decides the next actor, limits the state each actor receives, validates outputs, authorizes transitions and tools, merges parallel results, persists progress, and determines completion or escalation. A topology diagram alone is not an orchestration design.

When should I use multiple agents instead of one agent

Use multiple agents when specialization creates a measurable advantage: isolated context, disjoint tools, different permissions, parallel independent work, or separate ownership boundaries. First compare with a single agent and deterministic workflow. If the same model, context, tools, and policy are merely split into role prompts, coordination cost may increase without adding capability.

What is the difference between a supervisor and a handoff

A supervisor retains control and calls specialists for bounded artifacts. A handoff changes the active owner, so the target specialist handles the next turn and may later transfer control again. Both patterns can use the same centralized runner and durable state store. Logical peer control is not the same as physically decentralized infrastructure.

How do you prevent infinite handoffs

Use an allowlisted transition graph, maximum hops and depth, end-to-end deadline, decreasing work budget, explicit completion predicates, repeated-state detection, and a human escalation state. Log every proposed and accepted transition. Do not rely on a prompt instruction such as "avoid loops."

Which multi-agent framework should I choose

Define your control contract first, then choose the framework that maps cleanly to it. OpenAI Agents SDK directly distinguishes manager-style agents-as-tools from handoffs. LangChain offers several patterns plus custom workflows. AutoGen provides team presets and termination conditions. CrewAI provides sequential and hierarchical processes. Prototype the same evaluation set with the simplest viable option rather than comparing feature lists alone.

Summary

Production multi-agent orchestration is controlled state transition under uncertainty. Start with deterministic code or one agent, separate topology from execution flow, and add autonomy only where it improves a measured task. Keep transition allowlists, budgets, termination, authorization, persistence, and effect idempotency outside model prompts. The best architecture is not the one with the most agents; it is the smallest system that reaches the required outcome and fails in a bounded, explainable way.

Primary Sources