Spec-Driven Development With AI Coding Agents
Executable specifications keep AI-generated code aligned with intent across iterations.

A pull request lands in the queue, five hundred lines long, touching nineteen files, and the reviewer opens it knowing the first honest question is not "is this correct" but "who do I ask." Nobody wrote it, in the sense that matters: an agent generated it from a ticket and a prompt, and the person whose name sits on the PR is accountable for code they did not reason through line by line. Software has always run on an unspoken deal: the person who wrote the code can explain why it looks the way it does, what alternatives got rejected, and what constraint forced the odd workaround buried deep in the file. AI coding agents did not break that deal through bad output. They broke it by removing the decision history that used to travel with every line, so a reviewer now has to reconstruct intent from a ticket, a codebase, and the agent's final answer, a far more expensive task than reading a colleague's reasoning. Better prompting will not fix this, because the problem has already moved past whether the agent can write code to whether anyone can review how much of it there is, and the volume issue and the intent issue feed each other. Left alone, this produces a specific structural failure: the code keeps changing, nothing written down about intent changes with it, and the gap between what the system does and what anyone can say it is supposed to do only becomes visible once it is expensive to fix.
What "vibe coding" costs on a maintained codebase
Loose, conversational prompting, often called vibe coding, works fine for a weekend prototype and turns dangerous the moment a codebase has to be maintained by more than one person over more than a few weeks. The danger comes from two failures that compound as a repository grows. The first is context explosion: once an agent has to reason across an entire codebase at once, its output quality degrades as its context window fills, so large, sweeping changes come out less coherent, not more, the opposite of what a team hopes for when it hands an agent a big task. The second is silent drift. Without some machine-checkable specification pinning down intent, whatever design the team had in mind at the start dissolves a little with every iteration, and no single commit looks wrong in isolation, only the sum of them does. Agent scope creep makes both failures worse in practice: agents tend to add features nobody requested, polish a solution well past what was asked, and swap dependencies in or out without flagging the change, turning what started as a small feature request into an uninvited refactor. On a codebase of real size, a change touching a large share of the repository's files is not an outlier under these conditions, it is close to the default outcome of letting an agent work without a spec to hold it to.
Spec-driven development: inverting the source of truth between spec and code
Spec-driven development answers this by flipping which artifact gets trusted. In the traditional model, code is the ground truth and any design document is a courtesy, something written once and left to rot. SDD makes the specification the authoritative record: when required behavior changes, the specification changes first, and the code is regenerated, adapted, or audited against it. Code shifts from being the thing a team trusts to being the thing it verifies. The distinction that gets skipped in most casual descriptions of SDD is what kind of document the spec actually is. A traditional spec is prose that a human reads and may or may not follow. An SDD spec is executable, taking the form of BDD scenarios, API contract tests, or model simulations that run in a pipeline and fail when the code diverges from them, which makes SDD an enforcement model. A spec that cannot cause a test to fail is not an SDD spec, it is a wiki page, and the distinction matters because wiki pages are exactly the design docs everyone already knows go stale, sitting in a tool nobody opens while the code moves on without them. SDD specs live in the repository and in CI/CD, getting checked on every build. This changes what counts as valuable human work. The highest-leverage thing an engineer produces in an agent-first workflow is no longer the implementation, it is a specification precise enough that a machine cannot misread it. And a well-built spec does more than describe behavior: structured specs act as what some practitioners call super-prompts, breaking a complex problem into modular pieces sized to fit an agent's context window, which is what lets an agent handle a problem too large for one prompt without the coherence collapse described above.
The three levels of SDD rigor
SDD is not one practice with one setting, and picking the wrong rigor level is its own way to fail. Several serious accounts of the discipline from 2026 converge on the same three-tier structure, and they agree on where most teams should sit: the middle tier, not the top one. The lightest tier, spec-first, uses a specification to seed the initial code generation and then lets the code drift afterward; it suits prototypes and AI-assisted one-off features, costs little to set up, and offers no guarantee that spec and code still agree a month later. The tier in the middle, spec-anchored, treats the specification and the code as living together: both evolve, and automated tests enforce that they stay in step. One of the primary papers on the practice calls spec-anchored "the sweet spot for most production systems," and the claim holds up because it is the first tier that delivers real drift protection without requiring a team to trust an agent to generate all its production code unsupervised. The top tier, spec-as-source, lets humans edit only the specification while code is generated in full and never hand-touched; it removes drift by construction, but it demands generation tooling mature enough to be trusted at that level, and for most teams that tooling is not there. The industry's attention concentrates on spec-as-source because it is the cleanest idea, but the value available right now sits at spec-anchored, and teams that skip ahead inherit a generation-tooling burden stacked on top of the specification challenge they were trying to solve. The objection that this is too heavy for a fast-moving team does not hold against the middle tier specifically: spec-anchored does not require generating all the code, only that the spec travels in the same repository as the code and that the CI pipeline fails when the two disagree.
The canonical SDD workflow
The practical SDD workflow runs as a loop, not a handoff. A team starts with a constitution, a set of governing principles for the project, written once. From there, each feature moves through specify, then plan, then tasks, then implement, with clarify, checklist, and analyze available as quality gates a team adds wherever a feature carries real ambiguity. The rule stated by the people who build this process is to review the plan before it gets broken into tasks, and review the tasks before an agent starts implementing them. Each of those checkpoints is where a human catches drift before it compounds across files, which is the entire point of the loop. The clarify step deserves particular respect, because an agent that pauses to ask a question before acting is far cheaper to work with than one that charges ahead on a confident wrong assumption and spreads that assumption across dozens of files before anyone notices. Some teams run this as a multi-agent pipeline that mirrors the phases directly: one agent investigates whether a plan is technically feasible, a second designs the architecture, a third writes the implementation, and a fourth checks the result strictly against the spec, each agent doing one job. The clearest evidence that this loop pays off at scale comes from CRED, a fintech platform serving more than 15 million users in India, which put Claude Code to work across its entire development lifecycle and saw execution speed for delivering features and fixes double. That gain did not come from removing people from the loop. It came from moving developers toward the parts of the process, specification and review, where their judgment is worth the most.
Sharing Architectural Context Across a Multi-File Codebase
None of this works if the specification only lives in one person's head or one feature's folder. Once a codebase grows past a handful of files, the specification needs a place agents can find and read reliably every time they start a task, and two conventions have emerged to cover two different needs: one for how a project runs day to day, one for the architectural rules that cut across the whole thing. The first is AGENTS.md, a file that sits in a repository and tells an agent the operating rules of the project: how to build it, how to test it, what conventions to follow. In December 2025, the specification behind AGENTS.md was donated to the Linux Foundation's Agentic AI Foundation, a sign of how far the convention has spread as a shared standard. AGENTS.md has a hard limit: as natural language, it can only offer guidance. It is the right place to explain how a project works, and the wrong place to put a rule that must hold every single time, because nothing stops an agent from misreading or ignoring prose. ARCHITECTURE.md fills a different gap: it holds the architectural invariants that apply everywhere at once, principles, architecture decision records, security policies, naming conventions, and instead of belonging to one team or one folder, it sits transversally across the whole ownership tree and gets prepended automatically to every context bundle an agent works from, so no part of the codebase can claim it does not apply. Some teams push this further with a CI-driven generation pattern: a script runs in the pipeline, queries the data catalog, formats the result to the AGENTS.md spec, and replaces the relevant sections automatically, so the file describing the system stays synced to what the system actually does instead of becoming one more document that quietly goes stale. As AGENTS.md adoption spreads across tens of thousands of repositories, a new question is becoming measurable in its own right: how closely agents actually follow the instructions these files give them.
What enforces architectural constraints when natural-language specs cannot
Every piece of this system points to the same limit: prose, no matter how carefully written, cannot force an outcome. A human or an agent can read an instruction in AGENTS.md and still misunderstand it, skip it under time pressure, or apply it inconsistently across files, because natural language describes intent without any mechanism that stops a contradictory action from happening. SDD's real architecture has two layers. The first layer is guidance: AGENTS.md and ARCHITECTURE.md, written in plain language, telling agents and humans alike what the project intends and why. The second layer is enforcement: executable specifications, BDD scenarios, contract tests, CI checks, that run automatically and fail the build the moment code and intent disagree. Guidance shapes what an agent attempts. Enforcement catches what guidance fails to prevent. A rule that only lives as a sentence in a markdown file is a request; the same rule compiled into a contract test that fails on violation is a constraint, and that gap is the entire reason SDD exists as a discipline. Teams that build only the first layer get agents that behave well most of the time and drift silently the rest of the time, which is the exact failure this article started with. Teams that build both layers get something closer to what spec-anchored development promises: a codebase where the specification is still written by a human in language a human can check, but where the system itself, not a reviewer's memory or a wiki page, is what catches the moment code stops matching intent.
Sources
- Spec-Driven Development for Agentic Software Engineering: Harnessing Human-Agent Teamwork
- The Spec Growth Engine: Spec-Anchored, Code-Coupled, Drift-Enforced Architecture for AI-Assisted Software Development
- Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal
- A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities
- Spec-Driven Development:From Code to Contract in the Age of AI Coding Assistants
- Spec-driven development with AI: Get started with a new open source toolkit - The GitHub Blog
- Codified Context: Infrastructure for AI Agents in a Complex Codebase


