Blog
Essays and engineering notes on software factories, verification, and building a company run by agents.
Claude Code testing: are its generated tests actually useful?
We generated nine API and Playwright test suites, then evaluated them against 21 hidden bugs. The tests were strong. The usual prompt advice was not.
Field guideHow to verify AI-generated code before merging
The workflow the checklists miss: the negative control, diff-scoped mutation testing, re-executed evidence, and CI gates the agent can't edit — with original data on how fast this problem is growing.
Field guideWhy coding agents fake passing tests — and how to catch it
Green doesn't mean correct. When the same agent writes the code, writes the tests, and runs them, a passing suite proves self-consistency — not that your software works. The mechanism, the taxonomy, and the detection playbook.
EssayThe software factory
Software is moving from one developer with one AI to factories of agents that build around the clock. The bottleneck won't be how fast they can write. It will be whether anyone can trust what they shipped.
More essays soon — one honest piece at a time.
Running agents that ship faster than you can check?
We're onboarding a small group of pilot teams. One product, one call to start.
Book an intro call→