Dekita

Code review is your last line of defense against agent pull requests

ai agents testing ci

If more than a quarter of your pull requests are written by an agent, review is no longer a quality gate. It is the only gate.

A dev.to analysis by Max Quimby, published early October 2026, puts agent-generated pull requests at 27.6% of GitHub PRs, up from under 1% fourteen months earlier. Anthropic's 2026 Agentic Coding Trends Report, cited in the same piece, puts AI's share of newly written code at 41%.

The uncomfortable part: acceptance rates look fine. A forensic study of 33,000 agent-authored PRs referenced in the analysis found agents reach an 83.77% acceptance rate against 91.01% for humans. That gap is small enough to wave off.

The failures are not where you are looking.

What agent PRs actually break

The study's finding is that agents rarely introduce typos or syntax errors. They produce code that compiles cleanly and violates something else:

The analysis groups the recurring problems into four failure modes requiring different defenses. Reviewing an agent PR as if a human's reasoning stands behind it misses three of the four.

Failure mode 1: the agent cannot see the whole system

The first mode is contextual. Citing a FeatBit analysis of the 2026 productivity paradox, the piece splits missing context into four kinds: what the business actually needs, how the codebase fits together, what earlier reviewers flagged, and the unwritten conventions inside a team.

The damage is measurable. An MSR 2026 empirical study cited in the post found 23% of rejected agent PRs were duplicates, submitted while another contributor was already working the same issue. Stack Overflow's engineering blog reports AI-generated code carries 1.7 times as many bugs as human code, with logic and correctness errors running 1.75 times higher and security findings 1.57 times more frequent.

The same study found about 28% of agent PRs merge almost instantly. Small, well-scoped changes are the safe zone. Anything crossing services or architectural boundaries is where agents stumble.

The gate: contract tests at CI, run with tools like Pact or Specmatic. Plus architecture rules enforced ArchUnit-style. Greptile data cited in the piece puts Codex at 5-6% rework against a 10% human baseline.

Failure mode 2: polished diffs, fragile deployments

The second mode is a paradox of presentation. New Relic's 2026 State of AI Coding report, cited in the post, found 94% of engineering leaders rate AI-generated code as higher quality than human code at review time. Yet 78% of the same respondents report more incidents once it ships, 82% hit at least one production failure tied to AI-generated code within six months, and 74% say at least a quarter of it needs significant rework within a year.

Meanwhile 62% of teams now ship AI-generated code without line-by-line manual verification.

The structural explanation: agents have absorbed what well-written code looks like on the page, not how correct behavior holds up under concurrent load, stale caches, partial network failures, or a third-party API returning errors instead of success responses. Addy Osmani, quoted in the piece, summarizes it: "Code generation became cheap while understanding stayed expensive."

The gate: property-based testing with tools like Hypothesis or fast-check, which check invariants rather than fixed input-output pairs agents can memorize. Canary deployments that let production traffic judge behavior. Semantic diffing that surfaces behavioral changes rather than textual ones.

Failure mode 3: review queues outpacing reviewers

The third mode is arithmetic. Faros AI telemetry cited in the piece shows AI adoption correlates with 98% more PRs that are 154% larger, review times up 91%, and zero-review merges up 31%.

Reviewer instincts trained on human-authored diffs misfire on agent code, because the bugs sit in assumptions rather than implementation, and reviewers have not had time to build new instincts. The piece references the running debate between ThePrimeagen and Theo over agent-generated "slop" PRs, junior engineers who never build intuition by deferring everything to an agent, and auto-merge tooling that skips a human gate.

The gate: put a human in the loop with a narrower scope. Auto-merge only small, well-scoped changes. Any PR crossing services or architectural boundaries goes to a reviewer who knows the system.

Failure mode 4: the ones the analysis could not see

The fourth mode is not detailed in the text the way the first three are. What is documented: agents rarely introduce typos or syntax errors. They violate contracts, break cross-service dependencies, and roll back architectural decisions. That is a class of failure that CI, linters, and formatting gates do not catch, because the code is clean and the assumption is wrong.

What to change this week

The four modes point to changes in CI gates, deployment practice, and reviewer training, not exhortations to write better prompts.

  1. Contract tests at the gate. The failure is a contract you did not encode. If you cannot express the boundary in a test, the agent will not see it either.
  2. Property-based tests, not example tests. Example tests give agents the exact pairs to memorize. Invariants give you something to check when the model does something you did not predict.
  3. Canary before merge for anything risky. New Relic's numbers say polished diffs ship failures. Let production traffic decide, with the blast radius contained.
  4. A human gate with a scope. Reviewer time is the constraint. Auto-merge small well-scoped changes, and route everything crossing boundaries to a human who knows the system.
  5. Measure rework, not acceptance rate. The acceptance gap looks benign because the failure is downstream. Track incidents, rework, and production failures tied to AI-generated code instead.

Where I am unsure

The 27.6% figure and the 33,000-PR study are secondhand through Quimby's analysis and my reading of it. The vendor surveys (New Relic, Stack Overflow, Faros) are self-reported and should be read as direction rather than precision. I have not run these gates myself on a large agent-PR population; the numbers I would want before betting on them are my own repo's rework rate and production failure count tied to AI-generated code.

The structural argument does not need precise numbers. If code generation became cheap while understanding stayed expensive, then the gate that verifies understanding is the one that matters. That gate is code review, and it needs tooling built for the failures above.

Sources

https://dev.to

This post was written with AI assistance. The author is responsible for its content.