Writing
You can’t install an AI-SDLC
One human, one AI agent, over 500 production pull requests in four months. This series is the delivery process that made that safe: every rule in it dated, and traceable to something that went wrong.
More about the series
Over four months I built and shipped a production SaaS with an unusual team structure: one human and an AI agent as the primary implementer. More than 500 pull requests, hundreds of backlog items, and a delivery process that did not exist on day one, because it could not have. Nearly every rule in it is scar tissue: dated, and traceable to a specific incident where the agent, or I, did something that had to become a rule.
That is the series thesis: you can't install an AI-SDLC, mine or anyone else's. The rules you need are functions of the failures your org, your stack, and your agents actually produce. What transfers is the loop: notice the failure, write the rule, harden it into code, audit compliance.
The series runs in two tracks. One covers process and governance: interrogation gates, review gates, agent identity, audits, and where the human must stay. The other covers the architecture underneath: codebases agents can navigate, decisions as data, enforcement as code, least-privilege operations. Alongside both, I draw on rolling out an AI-enabled SDLC at a global enterprise in a regulated industry, where the same questions arrive with compliance attached.
Get the next article
The series runs to 15–20 articles, published roughly weekly. One email when each one goes up, and nothing in between.
You'll get one email asking you to confirm. Unsubscribe in one click, any time. Prefer a feed? Subscribe to the series over RSS.
Interlude: Jev and the economics of no-training
I re-ran an old AWS Comprehend project against TypeSafe's Jev. With no training data it beat the trained classifier at roughly a two-hundredth of the cost. It also showed where the risk goes when you stop training models and start writing questions.
Architecting for agent legibility
An AI agent is a permanently new team member: it re-pays your codebase's discovery cost every single session. What that means for repo structure, docs as code, restraint about abstraction, and conventions machines can check.
Build the quality floor before the autonomy
Nineteen phantom products, the pull request that pivoted a proof of concept into a product, and what a ratcheting coverage floor actually buys when agents write most of the code: regression detection at machine speed, so the human can follow the bottleneck.
Constrain, observe, review: agent safety as infrastructure
A self-hosted agent sandbox with exactly three egress paths, the breakout attempt it contained, and the reference architecture enterprise agent platforms should be heading for: every path an agent can take runs through a control point you own.
Show me what you think we built
Generated customer updates, questionnaires with assumptions attached, and a solution design that audits its own citations: reading what the AI thinks you built is how you test architecture and finish the last 10 per cent.
Engineered disagreement
Four AI personas that changed nothing, one that changed everything, and what it actually takes to stop two agents on the same model from agreeing with each other.
You can't install an AI-SDLC
Why an AI-enabled software development lifecycle (SDLC) has to be grown, not installed: eighteen weeks of evidence from one human directing AI agents on a production system.
AI code review A/B testing series
A weekly series benchmarking frontier LLMs on code review of a pinned 40,000-line codebase against a known-bug register.
Series finale: Opus 5 flips the result, at a quarter of the cost
The series started with Fable beating Opus at code review. Opus 5 just flipped it: full like-for-like numbers, plus honest caveats about flakiness.
My cheap AI code-review trick broke in a week
Has pxpipe with Fable been nerfed? On silent model-capability drift between two Tuesdays.
The $3.44 frontier-grade code review
Running code review through the pxpipe image-compression proxy, and the failed attempt to reproduce the result.
GLM-5.2: dead last as one agent, competitive as a swarm
Testing an open-weight model on the benchmark: last place as a single agent, competitive as a ten-agent swarm at half the price.
An exhaustive deep code review lost to a quick one. Three times.
Testing Claude Sonnet 5 on launch day against the known-bug register.
A tool promised to cut my AI token bill by 60–95%. It cut 4%.
Headroom context compression versus the Anthropic prompt cache, and why a 4% saving is not a bug.
Fable 5 taken offline three days after I said I'd use it
Model availability as supply-chain risk, after a US export-control order took my chosen model offline.
Fable 5 vs Opus 4.8: the first 24 hours
The series opener. A/B testing Claude Fable 5 against Opus 4.8 on a full review of a 40,000-line codebase: zero false positives, consistent runs, and it caught a race condition I didn't know about.
LinkedIn posts
An AI detector flagged the writing I did 15 years ago
Substack is rolling out AI detection with Pangram Labs, and it flagged personal writing of mine from well before LLMs existed. Detection is the wrong tool: the real question is trust, and the answer is being open about how you actually use AI.
Research
From Rules as Code to Mindset Strategies and Aligned Interpretive Approaches
Peer-reviewed journal article with QUT Law, co-authored with Mark Burdon, Anna Huggins, Nic Godfrey, Rhyle Simcock, Josh Buckley and Síobháine Slevin. Develops the rules-as-code work into mindset strategies and aligned interpretive approaches for legal and regulatory coding.
From Rules as Code To Legal and Regulatory Coding Strategies
Conference paper with QUT Law, co-authored with Mark Burdon, Anna Huggins, Rhyle Simcock, Nicholas Godfrey, Joshua Buckley and Síobháine Slevin. Coding an Australian statute and ASIC Regulatory Guide 274 into machine-executable form, showing the limits of plain-reading legal coding and proposing distinct legal and regulatory coding strategies.
Other writing
Motivating Your Dev Team With Real-Time Feedback...And A Fish Tank
Dev teams lack fast, visible feedback loops. Why an office marine aquarium, automated with Home Assistant and CI/CD, makes both a stunning centerpiece and a live testbed for DevOps practices.
I've Quit My Job and Gone to Greenland...
Trip report from the 2013 "North of Disko" expedition: sailing from Ireland to Greenland and putting up new climbing routes among the icebergs, including an E5 6a named after the expedition.