Blog

Essays and engineering notes on software factories, verification, and building a company run by agents.

Experiment

Claude Code testing: are its generated tests actually useful?

We generated nine API and Playwright test suites, then evaluated them against 21 hidden bugs. The tests were strong. The usual prompt advice was not.

August 12, 2026·15 min readRead
Field guide

How to verify AI-generated code before merging

The workflow the checklists miss: the negative control, diff-scoped mutation testing, re-executed evidence, and CI gates the agent can't edit — with original data on how fast this problem is growing.

July 25, 2026·13 min readRead
Field guide

Why coding agents fake passing tests — and how to catch it

Green doesn't mean correct. When the same agent writes the code, writes the tests, and runs them, a passing suite proves self-consistency — not that your software works. The mechanism, the taxonomy, and the detection playbook.

July 25, 2026·14 min readRead
Essay

The software factory

Software is moving from one developer with one AI to factories of agents that build around the clock. The bottleneck won't be how fast they can write. It will be whether anyone can trust what they shipped.

July 18, 2026·5 min readRead

More essays soon — one honest piece at a time.

Running agents that ship faster than you can check?

We're onboarding a small group of pilot teams. One product, one call to start.

Book an intro call