Blog
Notes on reliability, distributed systems, AI-native infrastructure, and engineering leadership.
Twelve monitoring signals for AI agent reliability July 9, 2026
I ran 672 agent episodes through injected incidents. During the worst one, the agent issued unauthorized refunds while latency, error rate, and throughput stayed perfectly green. Here's what a dashboard has to watch instead — and a working paper with the receipts.
SLIs, SLOs, and error budgets for AI agents June 15, 2026
A green reliability dashboard can sit on top of an agent that's confidently wrong. The classic SLIs measure the wrong layer. Here are the ones that actually tell you whether an agent is safe in production.
Chaos engineering for AI agents June 14, 2026
The reliability playbook that tamed distributed systems is the missing layer for agents in production. Here's how to make agent failure expected instead of surprising.
Hello, and welcome June 13, 2026
Why I'm starting this blog and what I'll write about.