LLM Evaluation & Security

A production-focused series on LLM evaluation, governance, and security, covering Harness Engineering, LLM-as-a-Judge, red teaming, prompt injection defense, guardrails, jailbreak analysis, OWASP Agentic Top 10, audit trails, and quality gates.

13 Articles in This Series · 创建于 2026-04-01
1

AI Agent Harness Engineering: Runtime Control, Scope, and Boundaries

Define Harness Engineering for AI agents without mechanical-engineering ambiguity. Learn how runtime policy, tool governance, state, budgets, approvals, observability, evaluation, and recovery bound model-driven actions, and where prompts, MCP, sandboxes, DevOps, architecture, implementation, and testing fit.

3

LLM Jailbreak Defense: Threat Model, Guardrails, and Evaluation

Learn how to defend LLM applications against jailbreak attempts with layered guardrails, least-privilege tools, output controls, red-team evaluation, and incident response. Clarifies how jailbreaks differ from prompt injection.

4

Agent Harness Evaluation: Test AI Agents for Production

Design a reproducible Agent Harness for AI systems. Learn how to isolate tools, replay scenarios, inject failures, enforce step and cost budgets, evaluate task outcomes and safety, and compare judge-assisted scoring without exposing private chain-of-thought.

7

When AI Benchmarks Mislead: A Practical Model Evaluation Method

A rigorous guide to evaluating language models when public scores are incomplete evidence. Learn how contamination, saturation, annotator disagreement, Goodhart effects, judge bias, and workload mismatch distort benchmarks, then build versioned task sets, deterministic oracles, calibrated human or model review, cost and latency measurements, and release gates.

9

AI Agent Memory and the Right to Erasure

Design AI Agent memory around purpose limitation, data minimization, provenance, retention, access control, deletion propagation, and evidence. Explains why vector deletion, summaries, caches, backups, fine-tuning, and model outputs need separate treatment under GDPR-style privacy programs.

10

EU AI Act Compliance: Developer Safety Checklist

A practical EU AI Act engineering checklist covering applicability, risk classification, audit logging, evaluation, human oversight, technical documentation, and conformity-assessment evidence. Uses the Digital Omnibus timeline: Article 50 from August 2026, Annex III high-risk rules from December 2027, and Annex I product-embedded rules from August 2028.