Post

You can't install an AI-SDLC

Why an AI-enabled software development lifecycle (SDLC) has to be grown, not installed: eighteen weeks of evidence from one human directing AI agents on a production system.

Eighteen weeks ago, PR #1 was merged into my project, TrolleyRelay. This week, we surpassed PR #596. The team consists of myself and a group of AI agents as primary implementers. During this period, we merged 579 pull requests, created 372 backlog items (169 closed), recorded 21 architecture decisions, completed 24 sprint close-outs, and built a codebase with more lines of test code than production code (78,000 versus 74,000).

To clarify, this is not vibecoding. Vibecoding is suitable for prototypes, spikes, and temporary explorations where code is disposable or long-term maintenance is unnecessary. The figures above represent production code for a real business, involving other people’s money and data, where ownership and accountability are critical. While the same agent and model are used, the key difference lies in the surrounding process, which is the focus of this series.

Workflow overview: one delivery pass across human, AI agent, and automation swimlanes, with the evolution loop beneath

When people see these numbers, they often ask about the process: “Which agent setup, which rules file, can they have a copy?” I can provide the rulebook, a single markdown file, but it is important to clarify what is actually being shared.

The CLAUDE.md file began as 15 lines with two headings: Backlog and Workflow. Today, it spans 165 lines across 17 sections, including a test-coverage ratchet, an interrogation gate the agent must pass before modifying code, review findings with mandatory categories, human-only merge paths for changes involving secrets, pricing, or decision records, and a dedicated GitHub identity for the agent to distinguish its actions from mine during audits.

Very little of this originated from a template, and not due to lack of preparation. With 18 years of software experience, I began this project with the safeguards I knew were necessary. The rulebook documents what experience could not predict: the failure modes of a new type of colleague, a fleet of AI workers. Nearly every section is dated, each reflecting a specific agent action. For example, after the agent merged on partially green CI, every check must now pass. When it shipped scaffolding based on an invented API payload shape, guessing shapes was prohibited. After it cut corners on seemingly routine tasks, the interrogation gate was applied to all work, regardless of size. I will detail each of these, with dates, as the series continues.

The origins predate this project. In March, before this codebase existed, I spent a month building an agent containment product: a sandboxed stack designed with only three controlled exit points, ensuring the agent never accessed real secrets. Its first commit message was: “First commit, overly complex. Time to simplify.” and nicely sums up why it failed. Although that product was ultimately unsuccessful, its core principle “do not trust the agent, constrain it, and make all actions observable and reviewable” became the foundation for everything that followed.

Network membership of the agent-containment stack: three isolated networks whose only overlaps are the API gateway, the LLM proxy, and the web proxy - the three controlled ways out

TrolleyRelay has had both busy and quiet months, with monthly commits as follows: April 74, May 189, June 73, July 187, and August 79 to date. May was the busiest month and also when most review processes were implemented, but the process was not an overhead; it enabled me to increase the agent’s autonomy while maintaining control.

But commit counts and PR counts are volume, not velocity, and the difference matters, because there is sixty years of prior work on it. Little’s Law (John Little, 1961) is the queueing result every delivery pipeline obeys: time in system equals work in progress divided by throughput. Hold throughput constant and every extra item in flight makes everything slower. The Toyota Production System (Taiichi Ohno) built lean manufacturing on that arithmetic: shrink the batches, cap the work in progress, and flow follows. The DORA research programme (Forsgren, Humble and Kim, Accelerate, 2018) established the software version across thousands of organisations: short lead times, small batches and high deployment frequency predict throughput and stability together. Speed and quality are not a trade-off; they have a common cause.

Measured that way, this project genuinely got faster. The median time from a PR opening to that PR merging fell from hours in April (a 22.7-hour weekly median at the April peak) to between 4 and 25 minutes for every full-throughput week since mid-May, a sustained trend of roughly -17% per week. Deploys went from zero in April, before a pipeline existed, to 28 in a single week by August. Those are DORA’s lead time and deployment frequency, and they moved the way the research says they should, because work in progress stays near one: each PR arrives already reviewed, merges alone, and deploys as a release tag.

Median PR open-to-merge time per week on a log scale: hours in April, then 4 to 25 minutes every full-throughput week from mid-May onward, trending -17% per week

So here is my thesis on AI and velocity: AI changes none of the underlying law. Minimising work in progress is still what buys you speed, exactly as it did on Ohno’s production line. What AI changes is the toolbox. The expensive parts of working in small batches - the review that queued on a human calendar, the tests nobody had time to write alongside the change, the deploy that needed a person watching - become cheap enough to run on every small change. The agent did not repeal Little’s Law; it lowered the cost of complying with it.

However, this acceleration presents a risk. AI speeds up whichever step it is applied to, and most teams focus first on code generation. As Goldratt’s Theory of Constraints (The Goal, 1984) explains, increasing speed at one stage does not improve throughput if the constraint remains elsewhere; it only creates inventory before the bottleneck. In an AI-assisted SDLC, the constraint is often verification and integration: review capacity, testing, deployment, and human attention. If the agent accelerates code generation but the rest of the process is unchanged, unreviewed branches and PRs accumulate, and unfinished work goes stale. This accumulation reduces the velocity you expect to gain. My experience in April reflected this: the agent produced work faster than I could safely absorb. The improvement in May was not due to a better agent, but to focusing acceleration on the constraint (reviews, tests, and deployments) and capping work in progress with gates. Each gate in this series serves as a WIP limit. Work does not start until it is reviewed, is not merged until all checks pass, and is not complete until deployed and audited. The process is what transformed production speed into consistent flow.

To be clear, I did not build all of this from scratch, and you should not either. Parts of my workflow use Matt Pocock’s published agent skills; his grilling skills power my interrogation gate, along with other adopted conventions. These are excellent, but they also demonstrate the thesis: even with the best available skills, I still needed to develop a governance layer tailored to my project’s unique failure modes.

There is a second thread running through this series. Alongside this project, I am rolling out AI tooling and an AI-enabled SDLC at a global enterprise in a regulated industry, where governance is not a style preference but a regulatory fact.

This is why my project includes more governance than a typical one-person team requires: it serves as a testing ground. Every gate described here was trialed in an environment where I could observe failures before recommending them elsewhere. Patterns from the enterprise rollout, what translates and what must change when the operator is not the owner, will be discussed throughout the series.

The thesis of this series is that you cannot simply install an AI-SDLC, whether mine or anyone else’s. The necessary rules depend on the specific failures your organisation, your stack, and agents encounter. What is transferable is the process: identify failures, write rules, implement them in code, and audit compliance.

This weekly (planned) series runs in two tracks. Track A covers process and governance: quality standards, interrogation gates, review gates for externally written code, the autonomy dial, and areas requiring human oversight. Track B addresses the underlying architecture: agent safety as infrastructure, designing navigable codebases, decisions as data, enforcement as code, and least-privilege agent operations. I also aim to provide a candid overview of the agent’s mistakes, as the failure log has become the roadmap.

If you are running agents in your SDLC: which of your rules exists because of a specific incident, and would you recognise the next incident as a rule waiting to be written?

This post is licensed under CC BY 4.0 by the author.