Post

Architecting for agent legibility

An AI agent is a permanently new team member: it re-pays your codebase's discovery cost every single session. What that means for repo structure, docs as code, restraint about abstraction, and conventions machines can check.

Architecting for agent legibility

Seventeen days into building TrolleyRelay, pull request #36 moved a perfectly functional codebase into a pnpm monorepo: apps under apps/, shared libraries under packages/ with each workspace carrying its own tests, coverage thresholds and configuration, and commands that fan out from the root. At that moment the whole project was two apps and one freshly extracted logging package.

The push did not come from me, it was a natural progression for the project but it was actually the product owner agent and the architect agent recommending that the workspace structure land before the first new app did, not after. So, just like many times in my career, a team decided it was the right time to bring structure to an early project…but this time the team was AI agents. Their argument was about what would happen later: shared code duplicated through brittle relative paths, a single test configuration going meaningless as new apps at zero percent coverage sat beside an established one in the nineties, cross-app changes that could not be reviewed as one atomic pull request. So this was still based on the basic principles I had implemented in the project already but applied pragmatically.

Twenty weeks later, seven apps and eleven shared packages sit inside that skeleton, and it hasn’t needed to be restructured, yet. Every new app since has landed into a place that already existed for it. Structure is cheapest at the moment it looks least necessary but do it too early and you can create a rod for your own back…yes the agents suggested these changes but I still think it is important that a human is in the loop for key decisions like this.

Pull request #36's skeleton (one app, one package) and the same skeleton twenty weeks later (seven apps, eleven packages), unchanged in shape

The permanently new team member

As a codebase grows, the reason structure matters even more for agents than for humans is not that agents are worse at reading code, it is that an agent is a permanently new team member.

A human joining a team pays a large discovery cost once: weeks of asking where things live, which document is current, which convention is real etc. Then that cost pays off over that person’s time working with that company/project. With an agent, every session starts from zero, or from whatever the codebase itself can teach in the first few seconds. Whatever your repository’s discovery cost is, an agent pays it again, in full, across hundreds of sessions.

So your codebase has to carry its own onboarding. In practice, what the agent gets in this project is a rulebook whose opening section is a map of the repository, one tree in which search finds everything, workspaces that describe themselves, and root commands that fan out to every workspace the same way. Predictability is what is important, the agent shouldn’t have to discover the layout; we should spoonfeed it.

I’m not saying that monorepos are the only solution. The monorepo is the mechanism in this project, not the requirement. The actual requirement is one navigable universe with explicit boundaries (and the second last section translates that for organisations where one repository is never going to happen). So, tests living beside the code they test (src/**/*.test.ts, in every workspace) is a legibility feature before it is a quality feature; the nearest documentation of intended behaviour is alongside the code, rather than requiring tribal knowledge.

The planning context ships with the code

The second structural decision I made is that everything the agent plans from lives in the same tree as the code it plans for: 386 backlog items with structured frontmatter, 24 architecture decision records, 26 sprint close-out documents, the environment-naming conventions, the operator runbooks. All versioned, all at the same commit as the code they describe. I follow the LLM-wiki pattern, I didn’t always but what I had built organically was so close to that pattern it made sense to adopt it properly.

This sounds like documentation hygiene but it is actually a statement about interfaces. When the backlog is a SaaS tool and the decisions are a wiki and the runbooks are on someone’s drive, an agent needs three integrations and three freshness judgments before it can plan. When they are files, planning context arrives through the same interface as everything else…the git repo.

The abstractions you refuse

Legibility is mostly described as things you add: structure, documents, conventions…but just as much of it is things you refuse to add.

My CLAUDE.md’s version is one sentence: “Do not extract a shared package into packages/* until a second consumer exists.” Speculative abstraction is exactly what an eager agent generates endlessly if you let it. Every proposed helper library, every “this might be reused” module is a plausible pull request, and an agent can produce an unlimited supply of plausible pull requests.

Twenty weeks in, an audit says that this rule worked: of the nine shared packages consumed as workspace dependencies today, eight have two or more consumers, and the ninth is a deliberate layering, not a premature extraction: everything reaches it through one operator toolkit. Under sustained agent-generated extraction pressure, nothing sits in packages/ waiting for a consumer that never came.

The same restraint shows up one layer down, in the migration’s own design notes, task orchestration is plain pnpm -r, chosen over dedicated build orchestrators with the recorded rationale that zero extra dependencies was sufficient at this scale. Every decision you have made not to do something is as important to document as the initiatives you did implement. Every piece of structure that was not built is discovery the agent never has to do.

Conventions the agent can check

A convention the agent cannot check is a convention the agent will eventually break. Not out of defiance; out of statistics. A rule that lives in one person’s head fails on the first session that never sees that head. So enforcement of these conventions progresses on a ladder, from conventions that are merely written down to conventions that cannot be violated.

The enforcement ladder: documented in one canonical place, centralised in code, schema'd and lintable, pipeline-enforced in CI

Four rungs, one example each, all live in this repository:

  • Documented, once, canonically. An environments document owns every naming convention for deployment targets/apps, storage prefixes and key aliases, and opens by telling the reader to reference it rather than reinvent it. The point is not that it is written down; it is that there is exactly one of it.
  • Centralised in code. Environment names are computed by a small shared package, not just recalled from memory or a rulebook, so the convention became an import. This isn’t perfect though, the module still ended up with mirrors, such as a human-readable table in the docs…but the documentation records the precedence rule for these details: “If the table and the module disagree, the module wins (the module is exercised by tests, the table is not).” Centralising a convention does not stop copies growing but legibility includes writing down which copy wins.
  • Schema’d. Backlog items carry structured frontmatter governed by a template document that the planning tooling uses. Format drift stopped being a style opinion and became something a script can flag. This schema can be used to lint documentation…so now automated testing isn’t just about code, it is also checking my documentation.
  • Pipeline-enforced. The framework this project’s admin app uses treats every file in its routes directory as a route, which means a stray export from a route file silently drags server code into the client bundle. That rule was implicit, invisible to agent and human alike, until the day it broke the build. The fix wasn’t adding to the rulebook, it is a test that fails on any route file exporting anything beyond the framework’s reserved names. The convention became agent-testable and enforced. A later article walks through the full enforcement ladder.

Each rung up converts a convention from something the agent must remember into something the agent cannot get wrong. When you find yourself writing the same review comment twice, you are really being told that it’s time to automate your enforcement better.

If you have two hundred repositories

Almost no established organisation gets to do this in one repository, and the monorepo was never the point. Every rung above has polyrepo equivalents:

  • Predictable layout becomes repository templates: every service repo shaped the same way, so an agent that has seen one has seen all of them. The test is whether your agents’ instructions transfer between repos unchanged. Make use of GitHub’s template repositories for the per-repo skeleton, and the organisation-level .github repository for the defaults every repo should inherit.
  • One navigable universe becomes a service catalogue that is itself machine-readable: the answer to “what exists, who owns it, where does it run” has to be a query, not a person.
  • Canonical conventions become a single versioned conventions repository that per-team documents reference instead of restating. The failure mode being designed against is the same one the monorepo avoids: five copies, each slightly wrong.
  • Schema’d metadata becomes enforced frontmatter or manifests on every repo, so fleet-wide tooling (and fleet-wide agents) can lint what they update.

The agent’s questions are identical in both shapes: can I find it, can I trust it, can I check it. An organisation that can answer those three across two hundred repositories is legible; a monorepo that cannot answer them is not.

Nothing is new but the noise is getting louder

Yet again, we are hitting the same recurring theme…none of this is new. Predictable layout, co-located tests, versioned docs, restraint about abstraction, conventions that tighten into checks. These are what good engineering organisations already say they want, and mostly deprioritise, because humans compensate. People absorb illegibility with context, memory and hallway questions, so the cost stays diffuse. But we still see the fallout when a P1 drags every senior engineer into the room.

Agents do not compensate, they amplify. Point them at a legible codebase and the structure pays out on every one of the hundreds of sessions that follow. Point them at an illegible one and every session re-pays the discovery cost, re-guesses the same ambiguities, and occasionally guesses wrong at machine speed. Legibility for agents is legibility for humans but on a much grander scale, and the codebase is now the primary interface to the team member doing most of the typing. Architect it like one.

This post is licensed under CC BY 4.0 by the author.