A Spec-Driven Workflow for AI Coding Agents, and the Hook That Enforces It

A spec-driven workflow for Claude Code agents: a numbered spec that gets attacked before any code, checks set up before features, three roles with walls between them, a PreToolUse hook that enforces the rules, and injected bugs to grade the tests. Written after running it once and logging every confident mistake.

September 2026

Agent Evals: How to Know a Change Made Your Agent Worse

Normal code has tests; agents need evals. What an eval harness is, the vocabulary that goes with it, how I would build one in Go for an SRE agent working against Kubernetes, why pass^k matters more than pass@k for ops, the tools that exist (and the one OpenAI is shutting down), the ways a harness's numbers lie to you, and the habits that keep a suite useful.

October 2026

Agent Harness: Everything That Isn't the Model

What an agent harness actually is, where the word comes from, and the dozen or so primitives hiding inside the term. A companion to the loop engineering post.

August 2026

Loop Engineering: How to Create Loops in Claude Code

What loop engineering actually means, where the term came from, and how /loop and /goal turn Claude Code into something that babysits deployments and works until the tests pass.

August 2026