Est.
Code ReviewLong read

Reviewing Large AI-Generated Diffs Effectively

Line-by-line reading misses the architectural failures that make AI-generated code dangerous.

Senior Writer · · 11 min read
Cover illustration for “Reviewing Large AI-Generated Diffs Effectively”
Code Review · October 2, 2026 · 11 min read · 2,431 words

Diff size is a secondary, not primary, signal: an agent can refactor hundreds of repetitive lines without changing much risk, so volume is useful as an attention-allocator but should not determine review depth. Now it's routine, the output of a single prompt to a coding agent that read the repository, made its changes, ran its own tests, and opened the PR already confident in itself. This code is not worse than what a human would have written. The unit of output has outgrown the unit of review: line-by-line reading was built for patches a person wrote by hand, one decision at a time, and it does not scale to a diff a system generated in one pass.

The mismatch is qualitative before it is quantitative. Agentic pull requests differ from human ones not just in size but in shape: analysis of tens of thousands of merged agentic PRs against a smaller set of merged human PRs from the MSR 2026 Mining Challenge dataset found substantial differences in commit count and moderate differences in files touched and lines deleted. A diff from an agent tends to touch more of the system at once and leave a different footprint than a comparable human change. The review habits built around human-shaped patches are already out of date before diff size even enters the picture.

Defect detection degrades once a change crosses roughly 400 lines, the point past which reviewer working memory starts to strain, and AI-assisted changes land past that threshold at the 75th percentile far more often than unassisted changes, whose median sits much lower. Code churn has climbed sharply alongside agentic adoption, so more code is arriving, faster, compounding a review bottleneck that was already strained. The paradox appears in a 2026 LinkedIn report: developers using AI were on average slower to reach a verified correct solution than those working manually, even though they perceived themselves as more productive. The discrepancy between how fast the code arrives and how fast a correct, shippable solution actually ships is review burden, and it is the problem this piece is about solving.

Three failure modes that line-by-line reading cannot catch

Diagram: Where AI-Generated Code Fails: Three Failure Modes Above the Line Level. Visualizes: Show three named failure modes that line-by-line review cannot catch, arranged as a vertical or horizontal sequence: (1) Missing Intent — no rationale…

Large AI diffs fail in three specific ways that a careful line-by-line read does not surface, because each failure lives above the level of the individual line.

The first is missing intent. A human developer who writes a patch carries a rationale for it: the approach chosen, the alternatives discarded, the constraints that shaped the decision, often recoverable from commit messages, comments, or a conversation with the author. AI-generated code arrives with none of that history attached, so a reviewer is left reconstructing intent from the ticket and the observed behavior alone. The second is behavioral failure hidden behind syntactic correctness. AI-generated code tends to compile cleanly, follow the style guide, and pass the baseline tests, all while remaining fragile under load, at the edges of input, or when given something invalid. A 470-PR analysis found AI-co-authored pull requests carried substantially more logic and correctness issues than their human-authored counterparts, with security vulnerabilities markedly elevated, error-handling gaps nearly doubled, and concurrency and dependency correctness issues running at roughly twice the rate. The third is context drift: agent-generated PRs drift away from the ticket they were meant to resolve, touch files outside the stated scope, or introduce architectural patterns that do not belong in the rest of the codebase, because the agent implements the obvious mechanism convincingly while missing a constraint that only exists as domain knowledge or a product rule nobody wrote down.

A practitioner's analysis of a getUser function makes the second and third failures concrete in a single example. An agent asked to fetch a user record produces a function using raw fetch and console.error, syntactically clean and unremarkable in isolation. The rest of the codebase, though, already routes HTTP through a shared apiClient with authentication, retries, and tracing built in, and the generated function introduces a second HTTP pattern, a second error-handling strategy, and in doing so may bypass authentication entirely. Nothing in the function itself is wrong. The wrongness exists only in relation to a system the function does not know about.

The same shape appears in testing. An order-cancellation feature is specified with the rule that a user cannot cancel an order after it ships; an agent generates an implementation that checks for a "delivered" status instead of a "shipped" one, then generates a test suite that confirms this behavior is green. The tests pass because they were written by the same reasoning that misunderstood the rule in the first place, and all they prove is that the implementation matches that misunderstanding. All three failures, missing intent, hidden behavioral faults, and context drift, share a single shape: the code is correct wherever you look at it directly, and wrong only once you widen the frame to the system it has to live inside.

Why AI review of AI code does not close the gap

If an AI system wrote the code, an obvious fix is to have another AI system review it. Automated review does catch a real class of problems, but it cannot substitute for the architectural and domain judgment that ultimately decides whether a change belongs in the system it was written for.

The category of AI code review tools has matured rapidly: PR-level reviewers, security scanners, and general-purpose assistants are all running in production today, solving different slices of the problem, and are often used in combination rather than as a single gate. Each does a real job. Security scanners catch the deterministic classes of problems, secrets, known CVEs, style deviations, reliably and fast. But AI reviewers also produce significantly more suggestions per PR than human reviewers do, and a larger volume of flagged lines is not the same thing as coverage of architectural risk. A tool can flag forty style issues and miss the one decision that matters: whether this change duplicates a pattern the system already solved, or violates a constraint nobody encoded anywhere a model could read it.

Addy Osmani, citing Greg Foster of Graphite, puts the stakes plainly: "If we're shipping code that's never actually read or understood by a fellow human, we're running a huge risk." Human sign-off is not disappearing from this picture; its purpose is shifting toward exactly the failures automated tools are least equipped to catch. That shift will only sharpen as more tools move toward fully autonomous remediation, scanning code, generating fixes, and opening PRs without a human in the loop at the drafting stage. When that happens, the reviewer stops catching individual bugs and instead certifies that the system's design intent survived an interaction the reviewer did not watch happen. The question that decides whether a change belongs, does this fit the architecture of this system as it is meant to evolve, depends on a model of design intent that does not live inside the diff, no matter how many tools read that diff first.

Reading the requirement before the diff: the first structural habit

The most consequential change a reviewer can make to their own process is a change of sequence: read what the change was supposed to accomplish before opening the diff that claims to accomplish it. Coding agents are very good at following instructions, and the instructions they are given are often incomplete. A requirement as simple as "only employees should be able to register" can produce a sophisticated email-validation utility covering international domains, quoted addresses, and a long tail of RFC edge cases, when the actual rule the business needed was a single check for whether the address ends in @company.com. The code is impressive. It answers a question nobody asked.

Before touching the diff, the reviewer's job is to ask whether the implementation missed an important business rule, solved a problem that was never posed, or built on assumptions that were never part of the original request. None of these questions are new; they were worth asking before any agent existed. What has changed is the cost of skipping them, because an agent can now generate a large volume of polished, confident code around a wrong assumption far faster than a person ever could.

One concrete check belongs here: does the PR description actually describe what the diff does? Agent-generated PRs can drift from the ticket that spawned them and touch files beyond the original scope, so the description and the diff need to be treated as independent claims that have to agree with each other, not one claim restating the other. Where a reviewer has access to it, a project's AGENTS.md or CLAUDE.md file is useful evidence at this stage to read before the diff. Anthropic added AGENTS.md support to Claude Code in version 2.1.277, released September 18, 2026: where a folder has no CLAUDE.md, Claude Code now checks for and uses AGENTS.md instead, following a cross-tool convention already adopted by other coding tools. These files persist in version control as a standing record of what constraints the agent was actually given, and a reviewer who reads them before the diff has a documented baseline to check the implementation against, rather than reconstructing intent from the ticket alone.

Reviewing against the architecture, not against the diff

None of this points toward reading every line more slowly. Effective review of a large AI-generated diff is a targeted audit of whether the change preserves the system's module boundaries, respects patterns the codebase has already settled on, and fits the direction the system is meant to evolve in.

The clearest signal to look for is pattern duplication: a second way of solving a problem the codebase has already solved once. A new HTTP client alongside an existing one, a new error-handling convention next to an established one, a utility function that duplicates one that already exists elsewhere in the repo. As the practitioner analysis behind this observation puts it, AI did not invent architectural inconsistency, it simply made it far easier to produce more of it, faster. Agents are also prone to altering files outside the scope the task called for, so scanning the full list of touched files against the stated change, before reading any individual diff, catches scope creep that a line-by-line pass would miss.

The getUser example from earlier shows what this looks like in practice: a function that passes every local check, every type signature satisfied, every test green, while quietly violating the codebase's conventions for HTTP calls, error handling, authentication, and observability. The right question a reviewer asks of that function is whether it belongs in this system at all.

Some of what a change needs to respect never appears in the repository in any form a model could read. A deliberate decision to avoid global state, an unusual retry pattern built around one external service known to be unreliable, an abstraction that looks clumsy until you learn that three separate teams depend on it, a field that looks obsolete but is still read by an older mobile client still in the field: these live in institutional memory, not in code, and an agent generating a plausible-looking change has no way to know they exist. Reviewers who have been through that history carry it; codebases do not store it. That gap is precisely why some categories of change cannot be delegated to review tooling. Where code touches authentication, payments, secrets, or any untrusted input, Osmani's framing is the right one: treat the agent as a fast, capable intern, and require both a human threat-model review and a pass from dedicated security tooling before the change merges. An agent can implement the obvious mechanism convincingly and still miss the one constraint that only exists as domain knowledge nobody wrote down.

Diff size still matters, but only as a way of allocating attention, not as a measure of how deep the review needs to go. An agent can rewrite hundreds of repetitive lines in a mechanical refactor without changing the risk profile of the system at all, and the architectural questions above apply with equal force to a ten-file change and a forty-file one. Size tells a reviewer where to look first. It says nothing about how carefully to look once they get there.

Visibility into the system's actual structure is what makes this kind of review possible at scale, beyond the diff's own claims about it. A living map of a codebase, tracking every module, file, class, function, and the call relationships between them, changes what a reviewer can see before they ever open a pull request. Instead of inferring the system's architecture from the diff in front of them (a diff written by the same agent whose judgment is under review), the reviewer can check directly whether a change shifts a module boundary, introduces an unexpected dependency across layers, or brings in a pattern inconsistent with everything already mapped in the graph. That is the structural counterpart to reading the requirement before the diff: reading the system before the change, so the change can be judged against it rather than against itself.

Handling tests when both code and tests are generated

Test coverage carries a specific meaning when a human writes the implementation and a different human writes the tests: it's evidence that two independent readings of the requirement agree. That meaning collapses when the same agent generates both from the same prompt. All a passing suite then proves is that the agent was consistent with its own understanding of the task, not that the understanding was correct.

The order-cancellation example from earlier shows how this plays out. The requirement says a user cannot cancel an order once it has shipped; the agent's implementation checks for "delivered" instead of "shipped," and the agent's own tests confirm that this (wrong) behavior works as implemented. The tests are green. The feature is broken. Nothing in the test run would tell a reviewer that, because the tests were never checked against the requirement, only against the code written to satisfy a misreading of it.

A test suite built this way can report high line coverage while verifying nothing that matters, which makes it a source of false confidence rather than a safety net. The fix is procedural: review the tests against the ticket or the specification directly, not against the implementation they were generated alongside. A test that can only be understood by reading the implementation first is testing the implementation's own assumptions, including the ones that are wrong; a green checkmark next to that kind of test is evidence that the agent was thorough about being wrong in a consistent way.

Sources

  1. The State of AI Code Review in 2026 - Trends, Tools, and What's Next - DEV Community
  2. Code Review in the Age of AI - by Addy Osmani - Elevate
  3. AI Code Review in 2026: The Hard Part Isn't the Code - DEV Community
  4. Code review best practices for AI-generated diffs
  5. How to Review AI-Generated Code: The Complete Developers Guide - Software Testing and Development Company
  6. How AI Agents Are Changing Software Development in 2026 - DEV Community
  7. How AI Coding Agents Modify Code: A Large-Scale Study of GitHub Pull Requests
Filed underCode Review