Multi-Agent Orchestration for Large Codebases
Specialization and git isolation let parallel agents work together without architectural chaos.

Multi-agent orchestration has become the default way serious engineering teams approach large codebases, a response to a structural ceiling that single agents hit on any codebase of real size.
Why multi-agent coding dominates large codebases
A single agent working alone runs into a hard limit long before it runs out of ideas. The failure is a lack of room, not a lack of intelligence.
The second limit appears when a single agent tries to do too many jobs at once. An agent asked to handle the data layer, the API, the UI, and the test suite in one pass produces noticeably worse work on all four than three focused agents would produce each handling one. Specialization helps a coding agent the same way it helps a team of people: narrower scope means sharper attention to the conventions and constraints that matter for that particular layer of the system.
The natural response, splitting work across several agents, carries its own risk. Teams have found two practical mechanisms that make parallel agent work usable instead of chaotic: isolating each agent's work in its own git worktree so parallel changes do not step on each other mid-task, and splitting roles into one agent that produces code and another that reviews it before it merges. Together, these two patterns are what let a team run agents in parallel without every merge turning into a pileup.
The scale this now reaches is documented. Anthropic's 2026 report on agentic coding describes an agent completing autonomous work across vLLM, a 12.5-million-line open-source library, in a single seven-hour run with minimal human intervention. That is a concrete demonstration that the ceiling on what agents can do to a real, large, actively maintained codebase has moved a long way in a short time, and it's the fact that sets up everything the rest of this piece has to say about what that scale costs.
What the coordination literature misses
The body of work on orchestrating multi-agent systems has identified a real and difficult set of engineering problems, and it deserves credit for naming them clearly. Failure recovery matters too, so one agent crashing does not stall an entire pipeline, and so does observability: knowing which agent acted, on what context, and why, a layer that production orchestration platforms in 2026 now treat as a baseline requirement rather than a luxury.
All of this is accurate, and all of it is worth solving. These mechanisms govern whether agents run smoothly together but say nothing about whether what those agents produce, collectively, holds together as a coherent system. A multi-agent setup can hit every one of these coordination goals, agents never collide, state passes cleanly, the merge queue clears without a hitch, and still produce code that is architecturally incoherent, because each agent completed its individual task correctly while the group of them quietly dismantled the design they were supposed to be building toward. Coordination governs how agents relate to each other. It has no mechanism for checking how agents relate to the codebase itself, to the architecture the team actually intends. The literature's gap sits exactly there: coordination is an agent-to-agent problem, architectural integrity is an agent-to-codebase problem, and most of the writing on orchestration treats the two as if solving the first automatically handles the second.
How parallel agents cause architecture drift faster than any reviewer can track
The mechanism behind this drift is ordinary, well-intentioned behavior happening at a speed no human review process was built to match. Agents are routinely instructed to look at existing code, follow its conventions, and reuse what is already there, which is sound practice for a single agent working alone. In a multi-agent system, that same instinct turns into a propagation mechanism: if the existing code contains one bad pattern, several agents working in parallel can all copy it independently in the same afternoon, each one reasonably following the example the codebase handed it. In a multi-agent setup, multiple agents can copy that same mistake at the same time, across different parts of the system, with no single one of them doing anything a reviewer would flag as suspicious.
That is what makes this hard to catch at the diff level. The problem lives one level up, in what the change does to the shape of the system, not in anything visible in the lines that changed.
Review processes built around a handful of diffs a week cannot scale to meet thousands of concurrent, machine-generated commits. That is a mismatch in rate: human review, by its nature, serializes one diff at a time through one person's attention, and agent output does not serialize itself to match. Only something structural, not more headcount, can close a gap like that.
The ContextCov research (arXiv:2603.00822) gives this pattern a name and a mechanism. Without immediate feedback at the point where a violation happens, these violations do not get caught and corrected. They accumulate quietly, across every parallel workstream running at once, until the gap between the intended architecture and the actual one is wide enough that no one can trace how it got there. Yegge's "50 First Dates" problem is the same pathology in a different outfit: agents with no memory between sessions do not just forget prior decisions, they produce conflicting records of what the codebase is supposed to look like, and each new session inherits that confusion instead of a clean picture.
The diff as the wrong interface for catching architectural problems
Engineers reach for the diff because it is the tool code review has always used, but the diff was built to show what changed line by line, not what a change does to the shape of the system around it. Architecture is not a property of any single line. It lives in the relationships between modules, between layers, in the shape of the call graph, none of which a diff view was ever designed to surface.
A single new import statement takes up one line in a diff and looks trivial. The number of files a long agent session can touch in a single run means the diff a reviewer is handed is too large to hold as one coherent object in working memory, however careful that reviewer is.
The common fix, pointing a second set of AI agents at the first set's output to review it, moves the bottleneck without closing it. A reviewer agent checking AI-generated code can catch local problems well: tests that fail, style that does not conform, bugs that are obvious on inspection. That model was never given to the reviewer agent to check against, so none of that tells you whether the change conforms to the architectural model the team actually holds. A reviewer, human or agent, can only judge a change against something it has been shown. Judging a change for internal consistency is a different task from judging it against an intended design nobody wrote down anywhere the reviewer could see.
What closes that gap is a representation of the codebase, built from its modules, dependencies, and call relationships, that can be checked mechanically against what the architecture is supposed to be. That representation is a graph, and it is the subject the next two sections turn to.
What structural enforcement requires
A structural enforcement layer is achievable without slowing agents down, but only if the check happens at the point where agents commit work, not later in a review queue someone has to clear by hand. The checks that hold up are the ones that exit non-zero and block the merge automatically. A gate that runs without anyone interpreting it is worth more than any amount of review conducted after the fact.
The fair objection here is that encoding architectural rules explicitly is itself expensive: rules go stale, teams argue about what the right boundary even is, and enforcing a wrong rule can do more damage than enforcing none. The fix is to make the architecture a graph that lives in the repository itself, versions alongside the code, and gets checked in continuous integration the same way a test suite does. Encoded that way, the cost of maintaining the constraint is a normal part of maintaining the codebase, not a separate standing cost, and it is considerably lower than the cost of the drift it exists to prevent.
The ContextCov validator pattern (arXiv:2603.00822) shows what this looks like in working form: build a dependency graph, enforce the layered structure that graph defines, and give agents immediate feedback that names the specific violation and where in the code it happened, at the moment the code gets generated rather than after it has already merged. The Thoughtworks Technology Radar describes the same operating model from a different angle: deterministic analysis finds structural problems, a verification loop checks that finding, and language models get used to help fix the violations that analysis turns up, with the fixes kept small and focused rather than sprawling. That ratchet tightens gradually, rather than acting as a one-time audit that catches a snapshot and then goes stale.
Meta's RADAR system shows what happens when this pressure builds at real scale. The RADAR paper treats automated diff triage as a necessary response to AI-generated volume, not an optional enhancement to a review process that was already working.
None of this requires slowing agents down, adding a human checkpoint at every step, or limiting how much work can run in parallel. The goal of a structural enforcement layer is to let agents run at full speed inside a clearly defined space, not to throttle the speed itself.
The graph as the load-bearing layer: from living map to CI gate
The step that actually separates a team managing drift from a team quietly accumulating it is turning the architectural graph from something people look at into something that blocks a merge. A graph that only gets drawn once and pinned to a wall is a snapshot. A graph wired into continuous integration is a gate, and that difference is the one that matters.
Used this way, the graph changes how developers direct agent work. Instead of writing a prompt and hoping the resulting structure happens to survive contact with the rest of the codebase, developers can make the architectural decision at the level of the map itself, then hand agents a bounded task defined by that map. The steering happens before the code gets written, not after.
The enforcement half of this runs as a check, commonly described as gr check, inside CI: it runs automatically, exits non-zero the moment it finds an architectural constraint violated, and blocks the merge without anyone needing to read a report and make a judgment call. It is versioned with the code the same way a test suite is, and it requires no human interpretation to act on.
That combination resolves the core failure mode described earlier in this piece. Agents can no longer introduce a quiet architectural violation and have it sit undetected for weeks until someone runs a drift audit and finds the damage after the fact. The violation gets caught at the exact boundary where an agent's work tries to enter the shared codebase, which is the earliest point it could possibly be caught. This is the Thoughtworks ratchet made concrete rather than described in the abstract: every CI run that passes tightens the constraint a little further, every refactoring pull request raises the baseline the next change has to meet, and the architecture improves in small increments instead of drifting in small increments. For a team running multiple agents in parallel, the practical payoff is that agents keep running at full speed and full parallelism. The enforcement layer is the thing that defines the space the agents are allowed to move in.
The developer's role in an agent-orchestrated codebase
Once structural enforcement is carrying the weight that human review no longer can, a developer's central job becomes designing and maintaining the constraints that decide what agents are allowed to build, a real change in what the work consists of rather than a smaller version of the old job.
The "Supervisor Class" framing, the idea that developers become high-level orchestrators directing autonomous agents rather than people producing code by hand, gets the direction of the shift right but leaves out what makes that supervision actually work. Supervision without a structural model behind it is just watching agents produce output and hoping it holds together, the exact condition this piece has been describing as the problem. What makes supervision a real form of control, rather than an after-the-fact audit, is having an enforceable graph that makes a violation detectable the moment it happens rather than weeks later.
Seen this way, the developer's role in an agent-orchestrated codebase is to build and maintain the map the agents are checked against, to decide what the architecture should be and encode it where a machine can verify it, and to treat the CI gate as the place where that decision gets enforced rather than merely suggested. The agents can move fast. Whether that speed produces a coherent system or an accumulating mess depends on whether someone defined, in a form a machine can check, what the system was supposed to look like.
