核心摘要: MCP Gateway 是可选代理或聚合层,不是官方 MCP Participant。MCP 2026-07-28 的现代 Request 是 Stateless,可到达任意兼容 Replica;Session Affinity 只属于 Legacy 兼容问题。生产设计必须校验 Routing Header 与 JSON-RPC Body、保持 OAuth Audience 边界、限制并发、只重试安全操作,并通过真实工作负载证明容量。
核心要点
- 官方 MCP 架构 只定义 Host、Client 与 Server;Gateway 是可选部署模式
- Transparent Proxy 保留一段 Server 关系,Aggregator 则对下游扮演 Server、对各上游扮演独立 Client
- MCP 2026-07-28 移除现代路径的 Protocol Session,并增加
Mcp-Method与Mcp-Name供基础设施路由 - Gateway Deny 只是其中一层策略;每个上游 Server 仍要授权具体 Principal、Tenant、Object、Argument、Purpose 与 Side Effect
- 有界 Queue、Per-upstream Bulkhead、Cancellation、Idempotency-aware Retry 与故障压测比宣传连接数更重要
为什么需要 MCP Gateway
只有共享控制减少的风险和运维工作大于新增跳数时,MCP Gateway 才值得引入。一个 Host、一个 Server 或本地 stdio 场景继续让 MCP Client 直连 MCP Server,通常更简单。
Fleet Policy 重复:多个 Host 会重复维护 Server Inventory、Trust Review、Protocol-version Policy、Egress Rule、Quota 与 Telemetry 配置。
Capability 冲突:聚合 tools/list、resources/list 或 prompts/list 时,需要稳定 Namespace 与 Provenance;裸 Tool Name 只在单个 Server 内唯一。
故障压力不均:若缺少 Upstream Isolation,一个慢速或高成本 Server 会耗尽共享 Queue、Socket、Memory 与 Retry Budget。
跨 Hop 安全边界:Gateway 可以集中执行入口 Admission 与 Deny Control,但下游 Server 仍需实施对象级和参数级授权,也不能把入口 Bearer Token 透传给另一个 Upstream Resource。
不要因为 Server 数量超过某个固定阈值就部署 Gateway。决策应基于可测的 Policy Duplication、Blast Radius、Latency Budget、Availability Target 与 Ownership。MCP Gateway 术语页 给出简明角色边界,本文聚焦生产数据路径。
MCP Gateway 核心架构
更稳妥的架构把经过审查的 Control Plane 与可水平扩展的 Data Plane 分开。Control Plane 管理 Server Identity、Policy、Descriptor Revision 与 Rollout;Data Plane 使用不可变 Snapshot 校验并路由 Request。
对 Aggregating Gateway 而言,下游连接终止于 Gateway。它校验 Request、解析已配置 Upstream、执行 Admission Policy,再创建独立的上游 MCP Request。这是协议边界,不是逐 Byte 透传:
Mcp-Method + Mcp-Name GW->>Auth: Validate JWT Token Auth-->>GW: Principal + Tenant GW->>Valid: Compare headers with JSON-RPC body Valid-->>GW: Version + capability + name GW->>Router: Resolve namespace and policy Router->>CB: Check circuit state alt Circuit Open CB-->>GW: Reject with bounded error GW-->>Client: HTTP / JSON-RPC failure else Circuit Closed/Half-Open CB->>Server: New upstream MCP request Server-->>CB: JSON or request-scoped SSE CB-->>GW: Response GW-->>Client: Validated result end
优先路由现代无状态请求
MCP 2026-07-28 Streamable HTTP 规范把普通远程路径定义为独立 HTTP POST,而不是复用 Protocol Session。每个 Request 都在 _meta 中携带 Protocol Version 与相关 Client Capabilities。Client 可以在其他操作前调用 server/discover 获取 Server Version 与 Capability,但这不是连接握手。
Streamable HTTP 把部分路由字段映射到 Header:
POST /mcp HTTP/1.1
MCP-Protocol-Version: 2026-07-28
Mcp-Method: tools/call
Mcp-Name: finance.payments.create
Content-Type: application/json
Accept: application/json,text/event-stream
{"jsonrpc":"2.0","id":41,"method":"tools/call","params":{"name":"finance.payments.create","arguments":{"invoiceId":"inv_8f2","idempotencyKey":"op_7b16f3a0"},"_meta":{"io.modelcontextprotocol/protocolVersion":"2026-07-28","io.modelcontextprotocol/clientCapabilities":{}}}}
Gateway 可以依据 Mcp-Method 与 Mcp-Name 路由或计量,但 JSON-RPC Body 仍是权威内容。终止协议的 Gateway 必须拒绝 Version、Method 或 Name 不一致,不能选择其中一个方便的副本来信任。只有 Tool Property 通过 x-mcp-header 显式标注时才能生成 Mcp-Param-*,不得镜像任意 Argument 或 Secret。
普通 Request 可以复用 HTTP Keep-alive Connection,但 Connection 不是 Protocol Identity、Conversation 或 Authorization Context。兼容 Replica 没有隐藏应用状态时,Round-robin Routing 就成立。Workflow 若需要连续性,应在协议数据中暴露受限业务 Handle,并在每个 Request 上重新授权。
隔离旧版连接处理
旧 MCP Revision 可能仍需要 Initialize、Mcp-Session-Id、GET/SSE、Affinity 或 Stream Resume。应将这些行为放在明确的版本化 Route,避免现代流量继承 Legacy 状态假设。
下面的 Go 风格连接池只用于说明旧版兼容问题。它省略 Handshake、Authentication、Parser、Heartbeat、Cancellation 与 Shutdown,不是当前 MCP 实现,也不能直接用于生产:
package gateway
import (
"context"
"fmt"
"net/http"
"sync"
"time"
)
type ConnState int
const (
ConnIdle ConnState = iota
ConnActive
ConnDraining
)
type SSEConn struct {
ID string
ServerURL string
State ConnState
CreatedAt time.Time
LastUsed time.Time
mu sync.Mutex
client *http.Client
eventCh chan []byte
closeCh chan struct{}
}
type ConnPool struct {
mu sync.RWMutex
conns map[string][]*SSEConn // serverURL -> connections
maxPerHost int
idleTimeout time.Duration
maxLifetime time.Duration
}
func NewConnPool(maxPerHost int, idleTimeout, maxLifetime time.Duration) *ConnPool {
pool := &ConnPool{
conns: make(map[string][]*SSEConn),
maxPerHost: maxPerHost,
idleTimeout: idleTimeout,
maxLifetime: maxLifetime,
}
go pool.evictLoop()
return pool
}
func (p *ConnPool) Acquire(ctx context.Context, serverURL string) (*SSEConn, error) {
p.mu.Lock()
defer p.mu.Unlock()
conns := p.conns[serverURL]
for _, conn := range conns {
conn.mu.Lock()
if conn.State == ConnIdle && time.Since(conn.CreatedAt) < p.maxLifetime {
conn.State = ConnActive
conn.LastUsed = time.Now()
conn.mu.Unlock()
return conn, nil
}
conn.mu.Unlock()
}
if len(conns) >= p.maxPerHost {
return nil, fmt.Errorf("connection pool exhausted for %s", serverURL)
}
conn, err := p.dial(ctx, serverURL)
if err != nil {
return nil, err
}
conn.State = ConnActive
p.conns[serverURL] = append(p.conns[serverURL], conn)
return conn, nil
}
func (p *ConnPool) Release(conn *SSEConn) {
conn.mu.Lock()
defer conn.mu.Unlock()
conn.State = ConnIdle
conn.LastUsed = time.Now()
}
func (p *ConnPool) evictLoop() {
ticker := time.NewTicker(30 * time.Second)
defer ticker.Stop()
for range ticker.C {
p.mu.Lock()
for url, conns := range p.conns {
alive := conns[:0]
for _, conn := range conns {
conn.mu.Lock()
expired := conn.State == ConnIdle &&
(time.Since(conn.LastUsed) > p.idleTimeout ||
time.Since(conn.CreatedAt) > p.maxLifetime)
if expired {
close(conn.closeCh)
conn.mu.Unlock()
continue
}
conn.mu.Unlock()
alive = append(alive, conn)
}
p.conns[url] = alive
}
p.mu.Unlock()
}
}
func (p *ConnPool) dial(ctx context.Context, serverURL string) (*SSEConn, error) {
conn := &SSEConn{
ID: fmt.Sprintf("conn-%d", time.Now().UnixNano()),
ServerURL: serverURL,
CreatedAt: time.Now(),
LastUsed: time.Now(),
client: &http.Client{Timeout: 0},
eventCh: make(chan []byte, 256),
closeCh: make(chan struct{}),
}
req, err := http.NewRequestWithContext(ctx, "GET", serverURL+"/sse", nil)
if err != nil {
return nil, err
}
req.Header.Set("Accept", "text/event-stream")
go conn.readLoop(req)
return conn, nil
}
func (c *SSEConn) readLoop(req *http.Request) {
resp, err := c.client.Do(req)
if err != nil {
return
}
defer resp.Body.Close()
buf := make([]byte, 4096)
for {
select {
case <-c.closeCh:
return
default:
n, err := resp.Body.Read(buf)
if err != nil {
return
}
if n > 0 {
data := make([]byte, n)
copy(data, buf[:n])
select {
case c.eventCh <- data:
case <-c.closeCh:
return
}
}
}
}
}
该示例会扫描并清除空闲 Legacy Connection,其中的数值只是占位符。真实兼容 Adapter 必须按 Authenticated Principal、Configured Server、Negotiated Legacy Revision 与 Authorization Context 隔离状态,并限制 Parser Memory、Event Queue、Descriptor、Redirect 与 Shutdown Time。
http.Client{Timeout: 0} 正说明这段代码不能复制使用。Legacy Stream 仍需明确 Context Cancellation、Heartbeat / Idle Limit、Connection Cap、经过测试的 Drain Deadline 与目标 Revision 的 Reconnect Semantics。现代 Request-scoped SSE 在最终 Response 后结束;长期 Change Notification 使用独立 subscriptions/listen Request,而且不支持 Last-Event-ID Resume。
请求路由与负载均衡
Gateway 只能依据经过审查、并与 Configured Server Identity 绑定的 Registry 路由,不能信任最近一个声明某 Tool Name 的 Server。Transparent Proxy 的配置 Endpoint 已确定上游;Aggregator 则必须为重名 Tool 分配 Namespace,构造确定性 Composite List,并把每个对外 Descriptor Revision 映射回唯一 Upstream。
路由策略分为三层:
- 协议校验:核对
MCP-Protocol-Version、Mcp-Method、Mcp-Name与 Body - Namespace 与 Policy Lookup:把暴露的 Capability Revision 解析到允许访问的 Upstream Server
- Replica 选择:在该 Server Cluster 内选择健康且版本兼容的实例
package gateway
import (
"fmt"
"hash/crc32"
"sort"
"sync"
)
type ServerInfo struct {
URL string
Weight int
Tools []string
Healthy bool
}
type ToolRouter struct {
mu sync.RWMutex
toolMap map[string][]*ServerInfo // toolName -> servers
hashRing *ConsistentHash
}
type ConsistentHash struct {
ring map[uint32]*ServerInfo
keys []uint32
replicas int
}
func NewConsistentHash(replicas int) *ConsistentHash {
return &ConsistentHash{
ring: make(map[uint32]*ServerInfo),
replicas: replicas,
}
}
func (ch *ConsistentHash) Add(server *ServerInfo) {
for i := 0; i < ch.replicas; i++ {
key := crc32.ChecksumIEEE([]byte(fmt.Sprintf("%s-%d", server.URL, i)))
ch.ring[key] = server
ch.keys = append(ch.keys, key)
}
sort.Slice(ch.keys, func(i, j int) bool { return ch.keys[i] < ch.keys[j] })
}
func (ch *ConsistentHash) Get(key string) *ServerInfo {
if len(ch.keys) == 0 {
return nil
}
hash := crc32.ChecksumIEEE([]byte(key))
idx := sort.Search(len(ch.keys), func(i int) bool { return ch.keys[i] >= hash })
if idx >= len(ch.keys) {
idx = 0
}
return ch.ring[ch.keys[idx]]
}
func NewToolRouter() *ToolRouter {
return &ToolRouter{
toolMap: make(map[string][]*ServerInfo),
hashRing: NewConsistentHash(150),
}
}
func (r *ToolRouter) Register(server *ServerInfo) {
r.mu.Lock()
defer r.mu.Unlock()
for _, tool := range server.Tools {
r.toolMap[tool] = append(r.toolMap[tool], server)
}
r.hashRing.Add(server)
}
func (r *ToolRouter) Route(toolName, routingKey string) (*ServerInfo, error) {
r.mu.RLock()
defer r.mu.RUnlock()
servers, ok := r.toolMap[toolName]
if !ok || len(servers) == 0 {
return nil, fmt.Errorf("no server registered for tool: %s", toolName)
}
healthy := make([]*ServerInfo, 0, len(servers))
for _, s := range servers {
if s.Healthy {
healthy = append(healthy, s)
}
}
if len(healthy) == 0 {
return nil, fmt.Errorf("all servers for tool %s are unhealthy", toolName)
}
if len(healthy) == 1 {
return healthy[0], nil
}
target := r.hashRing.Get(routingKey + ":" + toolName)
if target != nil && target.Healthy {
return target, nil
}
return healthy[0], nil
}
一致性哈希是可选实现。它可减少 Tenant、Shard 或显式业务状态 Handle 的重映射,但 Connection 或自报 Client Name 都不是安全 Routing Identity。现代 Stateless Request 默认应使用普通健康 Replica 选择,只有应用契约确实要求 Affinity 时才引入哈希。
聚合能力时保留来源
Aggregation 会改变协议表面,因此必须拥有独立契约。Gateway 不能简单拼接多个上游 List,再假设 Name 天然唯一。
| 问题 | Gateway 必须执行的行为 |
|---|---|
| Server Identity | 将 Route 绑定 Configured Endpoint 或不可变 Deployment Identity;serverInfo 只作显示 Metadata |
| Name Collision | 使用稳定 Server Namespace 前缀或映射,并拒绝 Ambiguous Call |
| Version Support | 只声明 Gateway 能校验且忠实 Relay 或 Translate 的 Revision |
| Capability | 只暴露完整 Downstream-to-upstream Path 真正支持的交集 |
| Discovery Cache | Cache Key 绑定 Principal、Tenant、Server、Protocol Revision、Policy Revision、cacheScope 与 TTL |
| Descriptor Change | 比较 Tool、Resource、Prompt Hash,高影响能力重新审查后再暴露 |
| Fan-out | 限制并发、Deadline、Byte,并定义 Partial Failure 行为 |
Notification 是 Cache Invalidation Signal,不是 Permission Grant。现代 Gateway 针对所需 Filter 显式打开 subscriptions/listen Stream,按 Subscription ID 关联 Event,使受影响 Cache 失效,并在断线后重新订阅。
在两个 Hop 之间保留授权边界
Aggregating Gateway 会终止两段独立安全关系。入口侧可能把它视为受保护的 MCP Resource Server;在每个上游侧,它又是持有该 Server 专属 Credential 的 MCP Client。
MCP Authorization 规范与安全指南都禁止把同一个 Bearer Token 当作跨 Resource Boundary 通用凭证。禁止把入口 Token 透传给上游 MCP Server。
Gateway 必须校验 Issuer 与 Audience,把 Credential 绑定 RFC 8707 Resource Indicator,只申请被 Challenge 的 Least-privilege Scope,并按 Principal、Tenant、Issuer 与 Server 隔离 Token。Gateway 可以执行 Allowlist 或 Contextual Deny,但上游 Server 仍要授权具体 Object、Argument、Purpose 与 Side Effect。
对于 tools/call,用户 Approval 要绑定有效 Server、Descriptor Revision、关键 Argument、Destination 与 Policy Revision。Tool Annotation 与自然语言 Description 都是不可信 Hint。这种分工让 Gateway 承担 OAuth Transport Security,却不会成为唯一 Authorization Authority。
并发控制与背压机制
高并发场景下,如果不对请求速率进行控制,下游 MCP Server 极易被流量洪峰击垮。Gateway 需要实现两层防护:信号量控制并发数 + 令牌桶控制请求速率。
package gateway
import (
"context"
"fmt"
"sync"
"time"
)
type RateLimiter struct {
tokens chan struct{}
maxTokens int
refillRate time.Duration
stopCh chan struct{}
}
func NewRateLimiter(maxTokens int, refillRate time.Duration) *RateLimiter {
rl := &RateLimiter{
tokens: make(chan struct{}, maxTokens),
maxTokens: maxTokens,
refillRate: refillRate,
stopCh: make(chan struct{}),
}
for i := 0; i < maxTokens; i++ {
rl.tokens <- struct{}{}
}
go rl.refill()
return rl
}
func (rl *RateLimiter) refill() {
ticker := time.NewTicker(rl.refillRate)
defer ticker.Stop()
for {
select {
case <-rl.stopCh:
return
case <-ticker.C:
select {
case rl.tokens <- struct{}{}:
default:
}
}
}
}
func (rl *RateLimiter) Allow(ctx context.Context) bool {
select {
case <-rl.tokens:
return true
case <-ctx.Done():
return false
}
}
type BackpressureController struct {
semaphore chan struct{}
rateLimiter *RateLimiter
queueSize int64
mu sync.Mutex
metrics *BackpressureMetrics
}
type BackpressureMetrics struct {
Accepted int64
Rejected int64
Queued int64
}
func NewBackpressureController(maxConcurrent, maxRPS int) *BackpressureController {
return &BackpressureController{
semaphore: make(chan struct{}, maxConcurrent),
rateLimiter: NewRateLimiter(maxRPS, time.Second/time.Duration(maxRPS)),
metrics: &BackpressureMetrics{},
}
}
func (bp *BackpressureController) Execute(
ctx context.Context,
fn func(context.Context) (any, error),
) (any, error) {
if !bp.rateLimiter.Allow(ctx) {
bp.mu.Lock()
bp.metrics.Rejected++
bp.mu.Unlock()
return nil, fmt.Errorf("rate limit exceeded")
}
select {
case bp.semaphore <- struct{}{}:
defer func() { <-bp.semaphore }()
case <-ctx.Done():
bp.mu.Lock()
bp.metrics.Rejected++
bp.mu.Unlock()
return nil, ctx.Err()
}
bp.mu.Lock()
bp.metrics.Accepted++
bp.mu.Unlock()
return fn(ctx)
}
令牌桶限制 Admission Rate,信号量限制已经进入某个受保护资源池的并发工作。这段代码刻意保持精简:生产实现还要拒绝非法 Limit、停止 Refill Goroutine、暴露 Queue Deadline、按 Tenant 与 Upstream 划分配额,并避免让一个慢 Server 通过全局 Semaphore 阻塞无关流量。有限 Queue 可以吸收经过测量的短时 Burst;无界 Queue 只会把 Overload 转化成延迟与内存压力。
分离无状态路由与应用状态
现代 MCP 水平扩展不需要共享 Protocol Session Store。每个 2026-07-28 Request 都携带处理所需的 Version 与相关 Client Capabilities,因此可以由任意兼容 Gateway Replica 校验和路由。HTTP Keep-alive、Request-scoped SSE 与 subscriptions/listen Connection 都是 Transport Resource,不是持久 Conversation Identity。
部分应用仍需要连续状态,但必须显式建模:
| 状态 | 正确 Owner 与 Key | 扩展规则 |
|---|---|---|
| Workflow 或 Job | Application 或 Upstream Server,以不透明受限 Handle 为 Key | 每次 Request 都授权该 Handle,并定义 TTL、Concurrency 与 Replay Semantics |
| MRTR Continuation | Upstream Server,以不透明 requestState 表示 |
保持 Principal 与 Server Binding;不得解析或把它当作 Permission |
| Discovery Cache | Gateway,Key 绑定 Identity、Authorization Context、Server、Version、Policy Revision 与 cacheScope |
遵循 TTL 与 Invalidation;Private Entry 不得跨 Principal 或 Tenant |
| Subscription Stream | 发起 subscriptions/listen 的 Gateway Replica |
断线后重新订阅,并按 Subscription ID 与 Event ID 去重 |
| Legacy Session | 隔离的 Compatibility Adapter | 遵循目标 Revision 的 Session ID、Affinity、Expiry、Ordering 与 Reconnect Contract |
Redis 可以承载显式应用状态或 Legacy Registry,但不是现代 MCP 要求。Redis Pub/Sub 本身也不提供 Durable Delivery;若业务依赖不丢失、顺序或 Replay,应选择与应用契约匹配的 Store 或 Broker。不得仅用 Connection ID、Request ID、自报 clientInfo 或调用方提供的 Business Handle 作为授权或 Cache Access Key。
可观测性与监控
调试 MCP 请求时,应使用协议测试和所选 SDK 校验 JSON-RPC framing。生产 Gateway 应提供有界 Metrics、结构化 Logs 和分布式 Traces,不记录原始凭证或无界 Tool Payload。
核心指标应采用有界 Cardinality Label。原始 Principal、Request ID、URL、Argument 或无界 Tool Name 应进入采样并脱敏的 Trace / Log,而不是 Prometheus Label:
| 指标名称 | 类型 | 说明 |
|---|---|---|
mcp_gateway_requests_total |
Counter | 按 Version、Method Class、Route 与 Outcome 统计 Accepted、Queued、Rejected、Cancelled、Retried 和 Completed Request |
mcp_gateway_queue_duration_seconds |
Histogram | 按 Route 与 Upstream Class 统计 Admission 等待时间 |
mcp_gateway_service_duration_seconds |
Histogram | 排除 Queue 和 Upstream 后的 Gateway 处理时间 |
mcp_gateway_upstream_duration_seconds |
Histogram | 按稳定 Server Identity 与 Outcome 统计 Upstream Latency |
mcp_gateway_streams_active |
Gauge | 按 Stream Class 统计活跃 Request-scoped Response Stream 与 Subscription Stream |
mcp_gateway_cache_operations_total |
Counter | 按 Cache Class 统计 Hit、Miss、Bypass、Expiry 与 Invalidation |
mcp_gateway_circuit_state |
Gauge | 按稳定 Upstream Identity 记录 Circuit State |
mcp_gateway_effect_outcomes_total |
Counter | 统计 Known Success、Known Failure 与 Unknown Write Effect |
链路追踪使用 OpenTelemetry:继续有效 Inbound Context 或创建新 Trace,再分别建立 Admission、Routing、Authorization、Upstream 与 Streaming Span。
记录 Gateway / Policy Revision、Protocol Version、稳定 Route / Server Identity、Descriptor Hash、Request Correlation、Retry / Idempotency Decision 与 Final Effect Status。不能假设每个 Upstream 都会传播 Trace Context,还需单独关联 MCP Request ID。
审计日志应记录经过认证的 Principal 与 Tenant、目标 Server 与带 Namespace 的 Capability、Descriptor / Policy Revision、Authorization Decision、脱敏 Argument Digest、Response Type、Byte、Cancellation 与 Outcome。身份必须来自经过校验的 Credential,不能只解码 JWT。Access Token、原始 Secret、完整 Prompt 与敏感 Tool / Resource Payload 不得进入日志。
生产级容错设计
MCP Server 可能因部署更新、资源耗尽或网络分区而暂时不可用。Gateway 必须实现熔断器模式来隔离故障,避免级联失败。
package gateway
import (
"fmt"
"sync"
"time"
)
type CircuitState int
const (
StateClosed CircuitState = iota // 正常放行
StateOpen // 熔断,拒绝请求
StateHalfOpen // 半开,试探性放行
)
type CircuitBreaker struct {
mu sync.Mutex
state CircuitState
failureCount int
successCount int
failureThreshold int
successThreshold int
timeout time.Duration
lastFailureTime time.Time
onStateChange func(from, to CircuitState)
}
type CircuitBreakerConfig struct {
FailureThreshold int
SuccessThreshold int
Timeout time.Duration
OnStateChange func(from, to CircuitState)
}
func NewCircuitBreaker(cfg CircuitBreakerConfig) *CircuitBreaker {
return &CircuitBreaker{
state: StateClosed,
failureThreshold: cfg.FailureThreshold,
successThreshold: cfg.SuccessThreshold,
timeout: cfg.Timeout,
onStateChange: cfg.OnStateChange,
}
}
func (cb *CircuitBreaker) Allow() (bool, error) {
cb.mu.Lock()
defer cb.mu.Unlock()
switch cb.state {
case StateClosed:
return true, nil
case StateOpen:
if time.Since(cb.lastFailureTime) > cb.timeout {
cb.transitionTo(StateHalfOpen)
return true, nil
}
return false, fmt.Errorf("circuit breaker is open")
case StateHalfOpen:
return true, nil
}
return false, fmt.Errorf("unknown circuit state")
}
func (cb *CircuitBreaker) RecordSuccess() {
cb.mu.Lock()
defer cb.mu.Unlock()
switch cb.state {
case StateClosed:
cb.failureCount = 0
case StateHalfOpen:
cb.successCount++
if cb.successCount >= cb.successThreshold {
cb.transitionTo(StateClosed)
}
}
}
func (cb *CircuitBreaker) RecordFailure() {
cb.mu.Lock()
defer cb.mu.Unlock()
cb.lastFailureTime = time.Now()
switch cb.state {
case StateClosed:
cb.failureCount++
if cb.failureCount >= cb.failureThreshold {
cb.transitionTo(StateOpen)
}
case StateHalfOpen:
cb.transitionTo(StateOpen)
}
}
func (cb *CircuitBreaker) transitionTo(newState CircuitState) {
oldState := cb.state
cb.state = newState
cb.failureCount = 0
cb.successCount = 0
if cb.onStateChange != nil {
cb.onStateChange(oldState, newState)
}
}
func (cb *CircuitBreaker) State() CircuitState {
cb.mu.Lock()
defer cb.mu.Unlock()
return cb.state
}
熔断器状态机包含三个状态:Closed(正常转发并统计失败)→ Open(针对该 Upstream 快速失败)→ Half-Open(仅放行有界 Probe Budget)。上面的紧凑示例只展示状态转换,并未限制 Half-open 并发 Probe、分类 Failure、持久化状态或协调 Replica;生产实现必须补齐这些边界。一个 Upstream Circuit 也不能演变为全 Fleet 的 Global Circuit。
Retry Policy 依据 MCP Operation Semantics,而不是 HTTP POST 本身:
- Discovery 与 Read 只能在 Deadline、Freshness、Authorization 和 Consistency Contract 允许时重试
- 带 Side Effect 的
tools/call只有在 Upstream 实施 Idempotency Key 或 Transactional Deduplication 时才能重试 - 多次 Attempt 保持同一个 Logical Operation Key,同时使用正确的 Per-attempt Transport Correlation
- Timeout、Cancellation、Connection Loss 与 Circuit Open 都应视为 Unknown Effect,除非 Server 能证明 Outcome
- 使用 Exponential Backoff、Jitter 与总 Attempt Budget,而且 Retry 必须重新进入 Admission Control
只有 cacheScope、TTL、Authorization Context 与 Freshness Policy 均允许时,才能返回缓存的 Resource 或 Read Result,并明确标记它来自 Cache。不得把历史 Tool Output 伪装成新 Tool Execution 成功。若 Effect Unknown,应返回保留不确定性的有界错误;切换备用 Tool 是新的、需要独立授权的 Operation,也可能重复原来的 Side Effect。
常见问题 (FAQ)
Q: MCP Gateway 和 API Gateway 有什么区别? A: 两者都能终止 TLS、认证调用方、路由 HTTP、限流并输出 Telemetry。MCP-aware Gateway 还会校验 MCP Version 与 Routing Header,理解 JSON-RPC Method、Capability Discovery、Result Envelope、MRTR 与 Subscription Stream。若不需要这些控制,现有 API Gateway 可能已经足够。
Q: 单个 MCP Gateway 节点能支撑多少并发连接? A: 没有可迁移的固定数字。2026-07-28 普通 Request 是独立 POST,Request-scoped SSE 与 Subscription 才占用较长连接。应在目标硬件上测量 Request Rate、Concurrent Stream、Byte、Queue Time、p95/p99 Latency、下游饱和、Cancellation 与 Recovery。
Q: MCP 2026-07-28 需要 Sticky Session 或共享 Session Store 吗? A: 不需要。现代 Core 没有 Protocol Session、Initialize Handshake、GET Stream 或 Stream Resumption;Self-describing Request 可以到达任意兼容实例。只有隔离的 Legacy Route 或由受限 Handle 标识的显式应用状态才需要 Affinity 或共享状态。
Q: Gateway 层会增加多少延迟? A: 没有可迁移到所有部署的固定数字。应分离 Queue、Gateway Service、Upstream 与 Stream Duration,并覆盖 Token Validation、Policy Lookup、Serialization、Payload Size、Observability、Retry 与 Cache Decision;在目标 Transport 和硬件上报告 p50/p95/p99,不能假设额外 Hop 可以忽略。
Q: 如何平滑迁移现有的 MCP Server 到 Gateway 架构? A: 先选择 Transparent 或 Aggregating Behavior,再注册不可变 Server Identity,核对 Revision 与 Transport,给 Descriptor 分配 Namespace 并计算 Hash,配置独立 OAuth Audience 与 Credential,并测试 Authorization 与 Failure Behavior。通过显式 Route 切换小比例 Client,保持 Legacy 流量隔离;只有 Rollback、Duplicate Effect、Cancellation 与 Recovery 测试通过后,才下线直连。
总结
只有共享路由、Capability Governance、OAuth Mediation、Backpressure 与 Observability 的价值高于额外 Hop 和更大 Blast Radius 时,才应部署 MCP Gateway。面向 MCP 2026-07-28,应保持常规 Data Plane Stateless、隔离 Legacy Session 行为、保留 Server Authorization,并把 Aggregation 视为拥有明确 Identity 与 Namespace 规则的新协议表面。容量、重试安全与故障恢复都必须通过真实工作负载验证。
如需先理解底层 Participant 与 Transport 模型,可阅读 MCP 协议完全指南;MCP 协议高阶实战 则说明 Gateway 后方仍由 Server 负责的实现问题。