Stackness
Alternatives to line-by-line code review: six verification layers for agent-written pull requests, and what each one catches

Alternatives to line-by-line code review: six verification layers for agent-written pull requests, and what each one catches

As of 11 October 2026, teams use six verification layers in place of reading every line of an agent-written pull request: AI review agents, risk routing, mutation testing, independent oracles such as property tests, plan and test review, and provenance gates. None replaces reading on its own. The best AI reviewer on GitHub's ReviewBench, published on 5 October, finds 26% of the issues in the benchmark's answer key. For most teams the pick is an AI first pass on every pull request plus risk routing, so a human reads only changes to auth, money, schema or a public API.

Pull requests grew 64% in a year, and Meta's diffs doubled

The median pull request grew 64% in a year, from 44 lines in July 2025 to 72 in June 2026, in DX telemetry from more than 400 organizations (Justin Reock, 17 June 2026). At Meta, significant lines of code per human-landed diff rose 105.9% year over year and diffs per developer per month rose 51%, according to Meta's own RADAR paper. Bigger diffs, and more of them, land on the same reviewers.

Brian Houck of DX collected those numbers in a 5 August newsletter, and Rachel Laycock, CTO of Thoughtworks, disagreed with his conclusion on martinfowler.com on 2 September. On GitHub, monthly commits went from 1.4 billion to 2.9 billion since April. On r/ChatGPTCoding, a reviewer wrote on 6 October that they review about 20 pull requests a day that nobody on their team wrote.

What did line-by-line review provide beyond catching bugs?

Line-by-line review provided understanding more than bug finding. In Bacchelli and Bird's 2013 Microsoft study, 44% of 873 developers named finding defects as their first reason to review, yet defects were 14% of 570 review comments, fourth of nine categories. Code improvements led with 29%, ahead of understanding and social communication.

Google's 2018 case study found the same. Developers expected education, shared norms, gatekeeping and accident prevention from review, and only 2 survey respondents said review comments had found a bug. Knowledge transfer shows up mostly in interviews: 12 of Microsoft's 570 comments were about it. Every benchmark below measures bugs, the smaller part of what review produced.

How I picked the six layers

I counted a layer when at least two practitioner sources from the last six weeks recommend it, or when someone published a measurement of it. The sources are Addy Osmani's The Code Nobody Reads (28 September), Laycock's piece, Gergely Orosz's Pragmatic Engineer issue (8 September) and two r/ChatGPTCoding threads. Osmani discloses that he is "not neutral": he helps build a tool that writes code.

I left out the CodeRabbit-alternatives listicles, which rank products inside one layer, and feature flags with fast rollback, which nobody has measured.

Six layers teams use instead of reading every line

The six layers are AI review agents, risk routing, mutation testing, independent oracles, plan and test review, and provenance gates. Each takes over one job reading did: finding bugs, choosing what a human looks at, checking the tests, checking intent, or recording who approved what. Each block below lists tools, measurement, cost and when to pick it.

AI review agents

ReviewBench scores reviewers on 219 public pull requests from 187 repositories in 19 languages, with an MIT-licensed repo. Its answer key comes from human review comments, follow-up commits, linters and LLM review agents, judged by Claude Sonnet 5, which agreed with human labels 96.6% of the time. On the leaderboard I read on 11 October, precision sits between 83.8% and 90.2% for every reviewer. Recall is the gap:

Reviewer Grounded recall Precision F1 on high-severity issues
Copilot code review, Balanced 26.0% 87.8% 63
Devin 23.8% 84.0% 50
Qodo 22.1% 85.3% 45
Codex, GPT-5.6 Sol, max reasoning 18.5% 86.0% 81
Greptile (one round, June) 16.2% 86.1% 36
Cursor Bugbot 9.6% 87.7% 38

GitHub and Microsoft built the benchmark, and their product ranks first. The methodology does not say which LLM agents seeded the answer key, and CodeRabbit, Claude Code Review and Graphite are not on the board. Two other benchmarks also put recall around a third. LangChain's separate benchmark, also called ReviewBench (31 July, 59 tasks from its own monorepo), scored the best models at 0.30. On c-CRAB, Claude Code alone passed 32.1% of 234 tests built from human review comments, and four tools together passed 41.5%.

When to pick it: first, on every pull request, as a filter rather than an approver. Copilot can approve pull requests in a preview since 1 September, off by default.

Risk routing and review by exception

  • Tools: GitHub CODEOWNERS and rulesets, Graphite stacked pull requests, merge queues; Meta's RADAR is internal
  • Measured: Meta RADAR, June 2026, 331,000+ diffs landed
  • Cost: building and tuning the risk rules
  • Misses: whatever the risk score gets wrong

Risk routing sends a diff to a human only when it crosses a risk line, and lands the rest on automated checks. Meta's RADAR scores each diff, runs LLM review and deterministic validation, then lands it or routes it to people. RADAR diffs were reverted a third as often as other diffs and caused a fiftieth of the production incidents, and median review wall time fell 35%. The authors warn that diffs "are not randomly assigned": low-risk diffs go to RADAR by design.

Orosz reports the same triage at AI labs: public API and MCP changes, auth, the design system, non-additive schema changes and agent skills stay with humans. Reddit practitioners cap pull requests at around 250 lines.

When to pick it: once an AI first pass and CI exist and the review backlog is what slows you down.

Mutation testing

  • Tools: Stryker for JavaScript, TypeScript, C# and Scala, PIT for Java, cargo-mutants for Rust, mutmut for Python, all free and open source
  • Measured: Google, 17 million mutants on 760,000 changes (2021)
  • Cost: roughly mutants times suite runtime; no 2026 source measured CI hours or dollars
  • Misses: bugs outside its fault types, wrong specs

Mutation testing injects small faults and reruns the tests. A surviving fault marks behaviour the tests do not pin down, which matters when the agent that wrote the code also wrote the tests. One Reddit reviewer has "seen models write tests that satisfy the coverage but test the wrong thing". Google runs mutants only on lines changed in a review, and developers rated 82% to 89% of them worth fixing.

CircleCI's example is illustrative, not measured: 500 mutants on a 30-second suite take about 4 hours. Diff scoping cuts that. Stryker's incremental mode reused 3,731 of 3,965 results in its docs example, and cargo-mutants has --in-diff.

Stryker's downloads grew four times faster than React's over the year. These are npm API weekday medians, which count CI installs, not teams:

Package Oct 2025, per day Sep 2026, per day Growth
@stryker-mutator/core 28,367 460,968 16.3x
fast-check 911,179 6,892,115 7.6x
React, as a baseline 8,319,264 31,664,395 3.8x

When to pick it: when agents write the tests. Start on the diff of the riskiest modules with no threshold, and run the full set nightly.

Independent oracles: property tests, reference implementations, fitness functions

An independent oracle gets its authority from something other than the code's author: a spec, a property, a reference implementation or an architecture rule run as a test. Osmani cites aviation's DO-178C rule that critical verification is done by someone other than the author, and warns that a second copy of the same model "shares the first one's blind spots". Kiro's docs say property tests are "not formal verification". The related move is writing TLA+ specifications alongside agentic implementation.

When to pick it: for parsers, serializers, money and date maths, and anything with a reference implementation.

Plan, spec and test review

  • Tools: GitHub Spec Kit, Matt Pocock's /grill-me skill, Kiro specs
  • Measured: nothing
  • Cost: human time before the code exists
  • Misses: whether the implementation follows the plan

Plan review moves the human read onto the smallest artifact that still carries intent: the plan, the schema, the tests. antiburn's order is plan review, agent implementation, a developer reading tests and acceptance criteria, then a second developer at "the principle level". One Reddit rule: "if a test got weaker, the whole change goes back." This layer recovers the "alternative solutions" outcome of the Microsoft study, and the move is documenting system architecture and intent.

When to pick it: when the backlog comes from large agent pull requests.

Provenance and commit gates

  • Tools: gitsign from Sigstore, GitHub artifact attestations, SLSA, the Linux kernel's Assisted-by trailer, pre-commit hooks
  • Measured: no defect outcomes
  • Cost: low; Sigstore is free
  • Misses: correctness, entirely

Provenance gates prove who or what made and approved a change, and that nothing changed after approval. They do not prove the change is correct, and no source claims they replace reading. GitHub's docs say attestations "are not a guarantee that an artifact is secure". SLSA's v1.2 Source Level 4 requires "two or more trusted persons", so a Copilot approval cannot count as one of them. Erik Hill's receipts hook left a file unchecked for four days because of a head -1 bug, and he writes that it "never judges whether a claim is true".

When to pick it: for compliance, and to bind an approval to one commit so routing by author can be trusted.

The six layers side by side

The six layers differ most in evidence. Only AI review and mutation testing have large measurements, and only risk routing has outcome data, from one non-random study at Meta.

Layer Takes over from reading Misses Cost Evidence
AI review agent Correctness, security, test gaps ~74% of known issues, conventions, intent $0.05-$25 a review Three benchmarks, all with builder interests
Risk routing Choosing what a human reads Mis-scored diffs Building the rules Meta RADAR, observational
Mutation testing Checking that tests catch wrong code Spec errors Mutants x suite time Google at scale
Independent oracles Edge cases, spec violations, drift Missing properties Writing properties Vendor docs
Plan and test review Wrong approach, schema mistakes Unfaithful implementation Up-front human time Anecdotes
Provenance gates Who approved which commit Correctness Low Docs only

Which layers can you stack, and in what order?

The order the sources agree on runs from intent to merge, with a human at both ends. Osmani's rule is to build the checks before you stop reading, because stopping first is "the worst possible order":

  1. Review the plan, the schema and the acceptance criteria before the agent writes code.
  2. Run deterministic CI: types, linters, tests, security scans.
  3. Run an AI first pass on every pull request.
  4. Run mutation testing on the diff of tested, risky modules.
  5. Route by risk: humans read auth, money, data deletion, schema and public API changes.
  6. A human approves the merge of the exact commit that was checked.

Sources disagree on whether step 6 can go. Martin Monperrus's position paper says yes, without new data. Duckbill Group dropped most human review, Orosz reports, and its merged pull requests went from 353 to 684. Osmani, Laycock, the kernel's sign-off rule and SLSA Level 4 keep a person.

On Stackness, 7 real profiles list Claude Code and one lists GitHub Copilot, as of 11 October 2026 (data sources). None lists CodeRabbit, Graphite or Greptile. The numbers are small. The review agents sit in the AI coding tools developers list, the gates with the CI and DevOps tools, and the workflow is the pull request code review move.

Which layers for which team

  • Small team on a web app: AI first pass on every pull request, CI, and a human read for auth and payments.
  • Team whose agents write the tests: add mutation testing on the diff, and reject any change that weakens a test.
  • Library, parser or anything with money and dates: property tests and a reference oracle before more reviewers.
  • Large codebase with a review backlog: risk routing with explicit rules, plus plan review for big pull requests.
  • Team under SOC 2 or aiming at SLSA Level 4: two human approvers stay, and provenance gates bind them to the commit.

Tools in this post

and 19 more from this post

Use any of these tools?

Put them on a Stackness profile, say how you use each one and see who pairs them the same way. It takes a couple of minutes.

Show my stack