Series
Production failures
Find the failures that surface after an agent has already changed something.
8 articles
Reading list
In this series
- 01
· Operations · 8 min
Five things that keep breaking when AI agents move into production
The operator-pain pattern that keeps recurring in CISO discovery calls, with concrete scenarios and the standard fixes that don't close them.
- 02
· Operations · 5 min
Discovery lag: how long before you find out what your agents did?
Agents act at machine speed; teams find out at meeting speed. Naming the gap, measuring it, and the one mechanism that makes it negative.
- 03
· Operations · 8 min
Stopping a running agent
Three containment outcomes (stop, scope, quarantine), the credentials-in-flight problem, the sub-agent cascade, and what a kill switch actually means at runtime.
- 04
· Engineering · 7 min
A 1-second submit hid a 30-minute outage
How a bounded, read-only agent investigation separated a compound runtime outage into five causal defects, and what we built to make that method repeatable.
- 05
· Engineering · 6 min
After the outage: making agent completion provable
An agent turn can be accepted, scheduled, streamed, and still fail. We changed the runtime so intermediate progress can no longer masquerade as completion.
- 06
· Operations · 5 min
The last mile is where AI content gets expensive
An accepted AI draft still needs metadata, links, accessibility, approvals, and publishing checks. Measure the work between acceptance and a live asset.
- 07
· Operations · 4 min
Your AI program has a review queue, even if you call it approval
Measure waiting time, reviewer effort, and rework in the queue between an AI-assisted draft and a published marketing asset.
- 08
· Operations · 5 min
Brand review is a production cost, not a final opinion
Turn brand review into clear checks, useful return reasons, and measurable work instead of a late subjective pass on every AI-assisted asset.