Est.

Context Files for AI Coding Agents

Context files speed up AI agents but don't yet improve the quality of their code.

Staff Writer · · 11 min read
Cover illustration for “Context Files for AI Coding Agents”
AI Agent Workflows · October 8, 2026 · 11 min read · 2,371 words

AI coding agents do not remember anything between sessions. Each invocation starts cold: it has no awareness of a project's conventions, its prior mistakes, or the decisions a team made the week before. The industry's answer to that problem is the context file, and understanding what it can and cannot do is now a basic requirement for running agents on production code.

The emergence of context files for persistent project knowledge

An agent that opens a repository for the first time knows nothing about that repository's history. It does not know that a function was renamed last sprint for a reason, or that a particular module is off-limits for a planned refactor, or that the team prefers a specific commit message format. Early workflows tried to compensate with prompt engineering: you craft careful instructions at the start of a session. But that approach cannot hold up across long-running, multi-file tasks, because prompts vanish when the session ends, and nothing carries forward to the next one.

The fix that emerged across the industry is the version-controlled, repository-level instruction file: a document, checked into source control alongside the code it describes, that the agent reads at the start of every session. CLAUDE.md, AGENTS.md, and.github/copilot-instructions.md are the names that have stuck, and each functions the same way: a persistent briefing, rewritten by humans, read automatically by the agent before it touches a single line of code. Research on agent harness design (Galster et al.) frames these files as configuration mechanisms, artifacts that customize how the agent behaves on a given project, shifting guidance away from one-off prompting and toward something inspectable and collaboratively maintained.

The scale of the problem these files solve is easiest to see at the extreme. A production codebase running 108,000 lines of C#, as documented in Vasilopoulos's account of building a distributed system, cannot be summarized in a single prompt. Someone has to tell the agent, repeatedly and in a form it can act on, how the project is organized, which patterns to follow, and which mistakes it has already made and should not repeat. A context file is the mechanism for delivering that briefing without re-explaining it by hand every time.

The fragmented format landscape: AGENTS.md, CLAUDE.md, and the push toward a shared standard

No single format governs this space yet. Claude Code reads CLAUDE.md. Codex reads AGENTS.md, and will defer to an AGENTS.override.md where one exists in the same directory, since the override file takes precedence. GitHub Copilot reads.github/copilot-instructions.md. Gemini reads GEMINI.md. If a team runs more than one of these tools against the same repository, they end up maintaining parallel files that often say the same thing in slightly different words, simply because each tool was built to look for its own filename.

That fragmentation has not stopped a pattern of convergence from appearing in the data. Galster et al. find that context files dominate the configuration landscape overall and are frequently the only configuration mechanism present in a repository. Among the competing formats, AGENTS.md has emerged as the closest thing to an interoperable standard: tools other than one vendor's now recognize it. The same study finds that Claude Code users tend to reach for the broadest range of configuration mechanisms of any tool's user base, yet the typical repository, whatever tool it uses, relies on a context file as its sole configuration artifact. Few repositories adopt more advanced mechanisms such as skills or subagents.

For a team running multiple agent tools against one codebase, the practical path follows directly from that pattern: a tool-agnostic AGENTS.md at the repository root carries whatever conventions apply across all tools, and tool-specific files extend it only where a particular agent needs something the others don't. That keeps the shared material in one place and limits duplication to the genuinely tool-specific parts.

What context files contain and conspicuously leave out

A typical context file contains a predictable set of contents: a project description, the commands to build and test the code, a summary of coding conventions, some notes on architecture, and a specified commit message format, all arranged in a shallow Markdown hierarchy. That consistency across projects is itself informative. It tells you what developers believe an agent needs to know, and it tells you, by omission, what they have not yet learned to specify.

Galster et al. categorize agent configuration into eight distinct mechanism types, and context files, static and written in prose, dominate that set overwhelmingly. Executable scripts and external integrations could actually enforce a rule instead of just stating it, but they are rare by comparison. Even where a context file documents something resembling a skill or a specific capability, the instruction is almost always static prose rather than an executable check, so nothing confirms whether the agent actually followed it.

Architecture sections appear in a large majority of these files, but the architecture they describe is a prose summary, naming layers, naming patterns, describing intent, not a queryable map of which modules actually import which, or which boundaries currently hold and which have already eroded. Security instructions and CI/CD guidance appear in only a small minority of files, a gap that matters more than it might seem at first, because these are exactly the categories of constraint an agent can violate without anyone noticing until much later.

Vasilopoulos's project shows what happens when a team pushes this approach to its limit. A single CLAUDE.md proved insufficient for a 108,000-line codebase, so the project evolved into a tiered infrastructure: a hot-memory constitution encoding core conventions, 19 specialized domain-expert agents, and a cold-memory knowledge base of 34 on-demand specification documents, roughly an order of magnitude beyond what a typical manifest contains. That scale of investment is itself evidence that a single prose file, however carefully written, runs out of road well before a codebase of that size is fully described.

The empirical evidence on whether context files improve agent performance

Diagram: What Context Files Do — and Don't — Improve. Visualizes: Show the contrast between two outcome dimensions measured across three studies: efficiency (median runtime drops, output tokens fall) versus accuracy (task completion rates stay…

When you ask whether context files actually work, the research points to a real but narrow benefit, bounded by conditions that matter as much as the headline result. Lulla et al., presented at ICSE JAWs 2026, studied repositories and pull requests both with and without an AGENTS.md file, and found that when it is present, median runtime drops meaningfully and output token consumption falls, while task completion rates stay about the same. Agents finish faster and burn fewer tokens getting there. That is an efficiency result, not an accuracy result, and the distinction carries real weight: the file doesn't make agents produce more correct code, it makes them waste less effort finding their way to the code they were going to produce anyway.

That efficiency gain also turns out to be conditional on the quality of the file itself. The ETH Zurich study (Gloaguen et al., arXiv:2602.11988), examining 138 real-world instances, found that context files in general consistently increase the cost and number of steps required to complete a task. Inside that overall pattern, the source of the file matters: files generated by an LLM produced a marginal negative effect on task success, actually performing worse than having no context file at all, while files written by developers produced a marginal positive gain. The improvement from a human-authored file is modest, well short of the dramatic jump in capability that the format's advocates sometimes imply.

The sharpest challenge to the whole premise comes from a further ablation study (Khatri, arXiv:2607.27250, 2026), which argues that agents mostly fail on implementation skill, meaning feature design, pattern selection, the exact wiring of a change, rather than on any gap in repository knowledge that a context file could fill. A manipulation probe in that study found that the real AGENTS.md never converted a near-miss into a pass for either agent tested. Across all three studies, the consistent thread is that context files shape behavior and cut down wasted inference, but they do not reliably raise the ceiling on what the agent actually gets right.

Why prose constraints decay as architectural requirements accumulate

Diagram: How Constraint Decay Compounds as Rules Accumulate. Visualizes: Visualize the stepped decline in agent assertion pass rate as architectural constraints stack up: unconstrained generation performs best, adding one structured architectural…

That ceiling has a mechanism behind it, and it is not a problem that better writing can fix. The decline in an agent's compliance with architectural rules as more of them accumulate reflects a basic property of how large language models handle layered instructions under the pressure of generating output, not a failure of whoever wrote the file.

A study gave this effect a name: constraint decay. As architectural requirements pile up in prose form, correctness degrades in a measurable way. Agents perform best when generation is unconstrained. Introducing a single structured architectural pattern already costs a meaningful amount of ground in assertion pass rate, and stacking constraints up to a fully specified configuration drops even capable setups by a substantial margin. The decay is not linear convenience lost at the margins; it compounds as the list of rules grows.

Two distinct failure modes produce that overall decline. In some cases, instructions are underspecified: the constraint exists in the file, but it is too ambiguous for the agent to turn into a concrete action. In other cases, the instruction is perfectly clear and the agent violates it anyway, a case of non-compliance. Both failure modes occur in fully autonomous runs and in human-in-the-loop settings alike, so adding a human reviewer to the pipeline does not close the gap on its own, particularly once an agent's edit velocity outpaces the bandwidth of the humans meant to be reviewing it.

What a context file cannot do, no matter how carefully it is written, is represent the live structural relationships inside the codebase it describes: which modules currently depend on which, which boundaries are intact today, and which have already drifted since the file was last updated. Vasilopoulos's own infrastructure grew to an extraordinary volume of documentation because prose alone kept failing to stop the agent from repeating the same architectural mistakes across sessions, and that is a practitioner-side confirmation of the same limit the research identifies from the outside.

What machine-enforced structural checks catch that context files cannot

A constraint that fails a build and blocks a merge is categorically more reliable than the same constraint stated as an instruction, because the former does not depend on the agent choosing to follow it. That distinction, between an enforced rule and an instructed one, is the whole answer to the limit the prior section describes.

Software delivery pipelines already enforce a set of deterministic rules at the merge gate: linting, automated tests, security scanning, build verification, CODEOWNERS assignments, required reviews, and branch protection policies. None of these depend on anyone remembering to follow a written guideline; they run automatically and stop a change that fails them. Treating merge-gate enforcement as the right home for coding standards is itself established engineering practice, one that predates the current generation of AI agents by years (Thoughtworks has documented this approach as a standard part of delivery pipelines), and the agent era has simply made the question of what belongs on that list more urgent, since the volume of code an agent can generate in an afternoon has no equivalent in human-paced development.

Context files and CI checks do different jobs; they do not compete for the same one. A context file steers the agent during generation, producing the efficiency gains and convention adherence the Lulla et al. study measured. A structural check verifies the output after generation and stops a drift before it merges. Both are necessary, and neither substitutes for the other.

Even running both together, a gap remains. A CI check reports that a constraint was broken only after the agent has already finished the work, the violation is already written, and someone now has to unwind it. A live architectural map works differently: it shows which parts of the codebase's structure an agent is about to touch before it touches them, giving a developer the chance to steer the work before the violation ever gets written.

The mental model teams need: context files for conventions, structural maps for architecture

Context files and architectural maps are not two ways of solving the same problem. They operate on entirely different objects. A context file encodes what the agent should do: which commands to run, which conventions to follow, which commit format to use. A structural map encodes what the codebase actually is, in a form a machine can query and a CI pipeline can enforce.

Context files do a well-defined job, and they do it well. They persist the build and test commands so the agent never has to guess at them. They record coding conventions, commit message formats, and security considerations in one place where everyone can read them. They give the agent a project overview and a description of its layers, cutting down on wasted inference before a single line gets written. They give an agent enough orientation to avoid the most obviously wrong moves early in a task.

What a context file cannot do, regardless of how much care goes into writing it, is represent the live dependency graph of the codebase: which modules import which, which boundaries exist right now, and which have already drifted since the last time someone updated the file. It cannot detect that an agent's edit just introduced a cross-layer dependency that violates the very architecture the file describes. And it cannot update itself as the codebase changes. The prose goes stale over time, and stale context tells the agent something confidently that is no longer true.

Closing that gap requires a different category of tool entirely: one that maps every module, file, class, function, and call relationship into a live, queryable graph, so that architectural constraints can be checked against the codebase's actual structure rather than a prose description of what that structure was meant to be, and enforced through a check that fails the build when it finds a violation. For a team running agents against production code, the practical stack follows from everything above: write a well-crafted AGENTS.md, plus tool-specific files where a particular agent needs them, to handle conventions and commands; maintain a live architectural graph to handle structural oversight; and enforce both in CI, treating each as load-bearing for the class of constraint it was actually built to hold.

Sources

  1. Codified Context: Infrastructure for AI Agents in a Complex Codebase
  2. On the Impact of AGENTS.md Files on the Efficiency of AI Coding Agents
  3. Harness Engineering for Agentic AI Coding Tools: An Exploratory Study
  4. Instruction Adherence in Coding Agent Configuration Files: A Factorial Study of Four File-Structure Variables
  5. [2607.27250] Do Context Files Help Coding Agents? A Two-Agent Ablation Study on Real Repositories

More in AI Agent Workflows