AI Agent Engineering

A practical engineering path for building production AI agents, covering ReAct, tool use, memory, multi-agent orchestration, observability, safety boundaries, self-driving codebases, Agent Shepherd workflows, and review gates for reliable autonomous systems.

24 Articles in This Series · 创建于 2026-02-06
1

How to Build an AI Agent: Production Architecture Guide

Learn how to build a production AI agent with typed tools, durable state, guardrails, human approval, tracing, and outcome evaluation. This architecture guide includes a runnable Go runtime, framework selection criteria, security boundaries, and a deployment checklist.

2

Multi-Agent Systems: When and How to Build Them

A production-oriented guide to multi-agent systems: when multiple agents are justified, manager-worker and peer architectures, explicit message contracts, least-privilege tools, bounded parallel execution, framework selection, evaluation, and a runnable Go coordinator.

3

CrewAI Tutorial: JSON-First Crews and Production Flows

Build a current CrewAI workflow with JSON-first agents and tasks, understand sequential versus hierarchical processes, wrap crews in a Flow for production control, constrain tools and delegation, verify outputs, and evaluate whether multiple agents outperform a simpler baseline.

5

Open Source AI Agent Ecosystem: A Framework Selection Guide

Map the open-source AI agent ecosystem by responsibility rather than popularity. Compare direct model loops, agent SDKs, graph runtimes, role-based orchestration, MCP integrations, evaluation, and deployment through control ownership, task fit, failure recovery, portability, governance, and exit cost.

6

ReAct Agent Pattern: Production Loop and Safety Guide

Understand the ReAct agent pattern as a reasoning-action control loop, not a software framework. Build structured tool calls, policy gates, bounded observations, side-effect recovery, and trajectory evaluations without treating visible reasoning as audit evidence.

7

AI Agent Memory: Production Architecture and Evaluation

Design AI agent memory as a governed lifecycle rather than a vector database. Separate thread state, semantic facts, episodic evidence, and procedural knowledge; implement consent-aware writes, temporal updates, conflict resolution, secure retrieval, deletion, and LongMemEval-style evaluation.

8

Claude Code Agent Workflows: CLI, SDK, and CI/CD

Build reliable Claude Code workflows across the terminal, Agent SDK, and GitHub Actions. Learn task contracts, permission boundaries, session control, isolation, credential handling, deterministic validation, and release gates for production coding agents.

9

Cloud Coding Agents: Runtime Architecture, Security, and Gates

A vendor-neutral architecture guide to cloud coding agents: design per-run identity, isolated workspaces, egress and secret boundaries, branch-scoped delivery, tamper-evident evidence, acceptance gates, cancellation and recovery, and a Go policy validator for safe remote execution.

12

GitHub Agentic Workflows: Secure CI/CD Automation Guide

Build GitHub Agentic Workflows without handing repository authority to a model. Learn the public-preview Markdown and compiled lock-file model, read-only permissions, safe outputs, credential isolation, prompt-injection controls, deterministic quality gates, OIDC deployment boundaries, cost caps, observability, and evaluation for issue triage, CI investigation, documentation, testing, and pull-request automation.

14

Self-Driving Codebase: A Governance Guide for Agent PRs

A production governance guide for self-driving codebases and agent-authored pull requests. Learn how to classify eligible tasks, choose an autonomy level, define executable acceptance contracts, collect review evidence, restrict credentials and network access, protect reviewer capacity, measure accepted goodput, and roll back unsafe changes without treating vendor adoption claims as universal benchmarks.

15

A2UI: Building Safe Agent-Driven User Interface Contracts

A practical engineering guide to A2UI-style Agent-to-UI contracts. Learn how to pin a protocol version, treat generated UI payloads as untrusted data, constrain a component catalog and renderer, authorize every server-side action, protect users from injection and phishing, and test accessibility, retries, privacy, and rollback before production.

16

A2UI, AG-UI, and AI SDK: Choose the Right UI Boundary

A practical comparison of A2UI-style UI contracts, AG-UI-style event protocols, and framework-owned AI UI runtimes. Evaluate payload contracts, transport, rendering, trust boundaries, action authorization, accessibility, recovery, and upgrade risk without treating evolving packages or model output as production guarantees.

20

Agent Observability: Traces, Evals, and Debugging

Design an observability system for production AI agents that explains what happened without collecting hidden chain-of-thought or unnecessary personal data. This guide defines an event contract, OpenTelemetry boundaries, cost and quality signals, offline and online evaluation, replay-safe debugging, sampling, retention, and failure tests.

21

Loop Engineering: Designing Reliable Agent Automation Loops

A production-oriented guide to Loop Engineering: design triggers, durable state, bounded agent execution, independent verification, budgets, stop conditions, recovery, and human escalation without turning repeated prompts into uncontrolled automation.

22

What Is an Agent Loop? AI Agent Runtime Guide

Understand an Agent Loop as a production runtime protocol: typed decisions, parallel tool calls, stopping rules, durable state, recovery, and evaluation.

23

Agent Loop vs Loop Engineering: Runtime and Automation

Distinguish an Agent Loop inside one invocation from Loop Engineering across repeated agent runs. Learn the contracts for triggers, goals, state, verification, stopping, escalation, idempotency, and release evidence, with a runnable Go controller that keeps model execution separate from acceptance.

24

AI Agent Observability: Privacy-Safe Traces, Evaluation, and Cost

Design AI Agent observability around bounded event contracts rather than raw reasoning capture. This guide separates traces, evaluation, and cost accounting, shows what to redact and hash, explains LLM-as-a-Judge limits, and provides a workload-specific rollout plan for debugging, regression detection, privacy, and budget control.