
How to keep coding agent commits reviewable: turn, commit or pull request as the review unit
To keep coding agent commits reviewable, pick the review unit before the session starts. Read each agent turn as it lands, cut a commit only at a green test run that includes every caller the change broke, and ship the commits as a stack of small pull requests. Reading the whole session at the end stopped working when diffs grew: DX measured median pull request size up 64%, from 44 to 72 lines, between July 2025 and June 2026, and Meta reported 105.9% more significant lines per human-landed diff in a year. The 2006 Cisco review study put the ceiling at 400 lines a sitting.
This post treats the review unit as a setting, like the model or the sandbox in a coding agent harness. The whole protocol fits in five lines:
before the session pin current behaviour in tests, list the signatures that may change
every green run commit, with every caller the change touched
every commit read git show HEAD before the next prompt
every few commits open one stack layer as a pull request
always a person reads auth, schema, public API and money paths line by line
What is the review unit, and why does it change when an agent writes the code?
The review unit is the slice of change a person reads in one sitting. With human authors it was the pull request, written at human speed by someone who could explain it. A coding agent produces three boundaries, the turn, the commit and the pull request, and writes faster than anyone reads. So the unit has to be chosen on purpose instead of inherited from the PR template.
| Unit | Boundary | Tools that cut there |
|---|---|---|
| Turn | One prompt and the agent's work on it | Claude Code checkpoints: "Every prompt you send that starts a turn creates a new checkpoint". Cursor checkpoints |
| Commit | A commit, by you or the agent | Aider commits every edit by default. Jujutsu snapshots the working copy on most commands |
| Pull request or stack layer | A branch opened for review | GitHub stacked pull requests, in public preview since 30 July 2026. Graphite, ghstack |
Checkpoints are not commits. Claude Code's checkpointing docs say they do not track files changed by Bash commands and are "Not a replacement for version control". Cursor's say checkpoints "are stored locally and separate from Git".
Two September essays disagree on where review goes. Addy Osmani, who says he helps build a tool that writes code, argues in The code nobody reads (28 September) that "Line-by-line review is going away." Thoughtworks CTO Rachel Laycock wants the judgment moved before implementation: "why are we waiting until code review to do all of those things?"
How much bigger did agent-era diffs get, and who measured it?
Between 18% and 106%, depending on who measured what. Meta's figure comes from its own telemetry. DX, Faros AI, Jellyfish and Greptile sell measurement or review tools. An independent study points the other way: pull requests opened by autonomous agents on public GitHub are smaller than human ones. The growth numbers describe human-landed, AI-assisted pull requests inside companies.
| Source | Change | Metric | Sample | Who measured |
|---|---|---|---|---|
| Meta, RADAR paper, June 2026 | +105.9% | Significant lines per human-landed diff, year over year. The term is not defined | Meta-wide | Meta, on itself |
| DX, 17 June 2026 | +64%, 44 to 72 lines | Median PR size, July 2025 to June 2026 | Over 400 companies | DX, vendor telemetry |
| Faros AI, 12 April 2026 | +51.3% | Average PR size, lowest against highest AI-adoption quarters | 22,000 developers | Faros AI, vendor telemetry |
| Greptile, Q2 2026 | +79% | Median PR size, baseline not stated | PRs it reviewed | Greptile, vendor |
| Jellyfish, 2 September 2025 | +18.2% | Additions per PR, 0% against 100% AI adoption | Over 500 companies | Jellyfish, vendor |
| UNLV, arXiv 2601.17581, April 2026 | Smaller | Agent PRs against human PRs, merged | 24,014 agent and 5,081 human PRs | Independent |
DX's Brian Houck paired the Meta and DX figures on 5 August, and Laycock quoted them. DX's own headline says PR size "nearly doubled", which overstates 64%.
A 72-line median is still a readable unit. Google's 2018 study of 9 million changes found a median of 24 lines. The Cisco study, run by the review-tool vendor SmartBear over 2,500 reviews, said "LOC under review should be under 200, not to exceed 400", read at under 300 lines an hour. The load comes from volume: Meta's diffs per developer per month rose 51%, and in Faros's data median time in review rose 441.5% and PRs merged with no review at all rose 31.3%.
What does reviewing one agent turn catch that a pull request review misses?
Turn review catches local damage while the context is fresh: a removed guard, a file the agent had no reason to touch, an unrequested cleanup. It misses anything that spans turns, such as a signature changed in turn four that breaks a call written in turn two, because neither turn's diff shows the break. No study has measured per-turn review. The evidence is practitioner reports.
The failure, in the shape a developer described on r/ChatGPTCoding on 29 September (illustrative code, not theirs):
# turn 2: read, approved
total = charge(order)
# turn 4: read, approved
def charge(order, currency):
...
# at commit time
TypeError: charge() missing 1 required positional argument: 'currency'
They read every turn and objected to nothing. "An early-turn call no longer matched a signature a later turn had changed", and failing tests found it. In the same thread, another developer's rename was correct everywhere "except a call built from a string, which no test covered and no turn diff showed as a change":
handler = getattr(handlers, f"on_{event}")
A third failure sits above the diff. One solo developer had "a reviewer subagent say 'ship' on a 47-green-tests diff that had changed a 404 response body". Comparing real HTTP responses before and after caught it. Osmani's explanation: "a second copy of the same model shares the first one's blind spots".
| Unit | Catches | Misses |
|---|---|---|
| Turn | Removed guards, scope creep, wrong files | Cross-turn signature drift, string-built calls |
| Green commit | Cross-turn breaks inside the commit, if tests cover them | Behaviour no test pins |
| Pull request | Design, scope, blast radius | Line detail past a few hundred lines |
Where do people cut commits inside a long agent session?
At a green test run, and at an interface change together with every caller it broke, not at the end of each turn. Developers in the late September threads commit when the suite passes, read git show HEAD instead of the session diff, and abort the turn when a signature change cannot be committed with its callers. One commit can span several turns.
1. Pin behaviour, and lock the pins
Write tests that freeze current behaviour first, full response bodies and error bodies included, and stop the agent from editing them. In Claude Code, a PreToolUse hook that exits with code 2 blocks the tool call (hooks docs):
#!/usr/bin/env bash
# .claude/hooks/protect-contract-tests.sh
path=$(jq -r '.tool_input.file_path // empty')
case "$path" in
*/tests/contract/*) echo "contract tests are frozen for this session" >&2; exit 2 ;;
esac
The matcher covers Edit and Write. An agent that edits through Bash with sed gets past it, so CI should also fail when git diff --name-only origin/main... -- tests/contract is not empty.
2. List the signatures that may change
Before the session, write down the function names, route shapes and table columns the task may touch. A change is then either on the list or visibly off it.
3. Commit on green, automatically
A Stop hook runs when Claude finishes responding. This one commits only when the suite passes:
#!/usr/bin/env bash
# .claude/hooks/commit-on-green.sh
cd "$CLAUDE_PROJECT_DIR" || exit 0
[ -z "$(git status --porcelain)" ] && exit 0
make test >/dev/null 2>&1 || exit 0
git add -A
git commit -q -m "agent: green at $(date -u +%H:%M:%S)"
{
"hooks": {
"PreToolUse": [
{ "matcher": "Edit|Write",
"hooks": [{ "type": "command", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/protect-contract-tests.sh" }] }
],
"Stop": [
{ "hooks": [{ "type": "command", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/commit-on-green.sh" }] }
]
}
}
A red turn stays uncommitted. The next turn fixes the callers, or /rewind drops it. Aider gets you a commit per edit with no test gate, and /undo reverts the last one.
4. Read the commit, not the session
git show HEAD after each commit, and git diff HEAD plus git status for a red turn still in the working tree. One developer will not let "the agent start the next step until I've reviewed and approved the last diff". Keep your own edits in separate commits from the agent's.
5. Stack the commits as pull requests
| Tool | How | Status, 5 October 2026 |
|---|---|---|
| GitHub stacked PRs | gh-stack extension and agent skill. Review "only the diff for that specific layer" | Public preview since 30 July 2026 |
| Graphite | gt create, gt submit --stack |
Shipped, CLI free on every plan |
| ghstack | Each commit becomes its own PR | Open source. No forks, needs write access |
| Jujutsu | jj split, jj squash to recut history |
Open source |
GitHub's own guidance of 7 May says to split an agent PR that "touches more than five unrelated files" or whose purpose cannot be said in one sentence.
What does an AI reviewer cover when there is no second human?
A second pass, not a sign-off. On Martian's Code Review Bench for the month to 28 September, the top five reviewers scored between 60.7% and 65.1% F1 across 15,753 scored pull requests. Their recall sits at roughly 50% to 60%, so 40% to 50% of the fixes developers made after review had not been flagged. The merge decision should rest on a mechanical signal.
| Reviewer | Martian F1, month to 28 September | Price on 5 October 2026 |
|---|---|---|
| Cubic | 65.1%, first | Not checked here |
| Greptile | 62.3% | Pro $30 a seat a month with 50 credits |
| CodeRabbit | 61.7% | From $24 a developer a month, free on public repos |
| GitHub Copilot code review | 60.8% | In Copilot Pro at $10 a month and up, $0.05 to $5 a review in credits |
| Claude Code Review | 60.7% | $15 to $25 a review on average, research preview, Team and Enterprise |
| Cursor Bugbot | 57.5%, tenth | Usage billing, $1 to $1.50 a run on average |
Every vendor quotes its own denominator. GitHub says "71% of reviews surface actionable feedback". Cursor ranks Bugbot first on a 78.13% resolution rate, a metric it defines. Anthropic says "less than 1% of findings are marked incorrect". Copilot leaves comments without approving by default, and Claude Code Review's check run is always neutral, so neither blocks a merge.
The 28 September thread from a solo developer whose product has "a bit over 300k users", with CodeRabbit as the only reviewer, collected the controls that work without a second person:
- No production credentials in the agent's environment.
- A read-only reviewer session that can only read and grep.
- Merge on tests plus a migration dry run, "not 'the review looked fine'".
- A smoke script against staging. One commenter's four curls: "bots can both say lgtm, that script still caught a broken auth cookie last week."
Which changes still get read line by line?
Changes that cross a security boundary or are hard to undo: authentication and permissions, non-additive schema changes, public APIs, money paths, and the agent's own instructions and skills. The sources that let AI-only review through, Meta's RADAR, Duckbill Group and Osmani among them, all route these to a person. Make the rule mechanical so nobody has to remember it.
Duckbill Group requires human review when a change touches "the public API/MCP, auth, design system, non-additive database schema changes, or agent skills", enforced by a script that adds a GitHub label (Pragmatic Engineer, 8 September). Its self-reported results for five engineers: merged PRs up from 353 to 684, median merge time 26 hours with human review and 1 hour without, and no defect data. Meta's RADAR auto-reviews only the lowest-risk 5% of human diffs by default and excludes any diff showing secrets exposure, SQL injection or an auth bypass.
Code owners do not work for a solo maintainer, because GitHub does not let you approve your own pull request. A CI check does, with PR set to the pull request number:
changed=$(git diff --name-only origin/main...HEAD)
if grep -qE '^(auth/|db/migrations/|api/public/|billing/|\.claude/skills/)' <<<"$changed"; then
gh pr view "$PR" --json labels -q '.labels[].name' | grep -qx human-read \
|| { echo "sensitive path changed: read it line by line, then add the human-read label"; exit 1; }
fi
On Stackness, as of 5 October 2026, 8 real profiles list Claude Code and 4 list Git. One lists GitHub Copilot, and none lists CodeRabbit or another review bot (data sources). The numbers are small, but among the AI coding tools developers list on Stackness, the reviewer slot is empty. The pull request code review move is where people describe how they do it.
Key numbers
- +64%: median PR size, 44 to 72 lines, July 2025 to June 2026, across over 400 companies (DX).
- +105.9%: significant lines per human-landed diff at Meta, year over year, in a paper revised on 12 June 2026 (arXiv 2605.30208).
- 400 lines: the most a reviewer should read at once, from the 2006 Cisco study (SmartBear).
- 24 lines: Google's median change across 9 million reviewed changes, 2018 (ICSE paper).
- 60.7% to 65.1% F1 for the top five AI reviewers, month to 28 September 2026 (Martian).
- 8 real Stackness profiles list Claude Code and 0 list an AI reviewer, as of 5 October 2026 (data sources).
Quick answers
How do you keep coding agent commits reviewable? Commit only on a green test run, include every caller a change touched, and read git show HEAD before the next prompt. Ship the commits as a stack of small pull requests instead of one session-sized diff.
Should I review every agent turn? Yes, while the context is fresh, but not as the only review. Turn review misses breaks that span turns, so pair it with tests that pin behaviour and a commit gate on green.
Are agent pull requests bigger than human ones? It depends on the population. AI-assisted PRs that humans land inside companies grew 18% to 106% in vendor and company telemetry. PRs opened by autonomous agents on public GitHub were smaller than human PRs in an independent study of 24,014 of them.
Can an AI reviewer replace a second human? Not on its own. The top five score 60.7% to 65.1% F1 on Martian's benchmark and miss 40% to 50% of the fixes made after review. Use one alongside tests, a staging smoke check and a human read of sensitive paths.
Which changes still need a human to read every line? Authentication and permissions, non-additive schema changes, public APIs, payment code and agent skills or instructions.
Tools in this post
Use any of these tools?
Put them on a Stackness profile, say how you use each one and see who pairs them the same way. It takes a couple of minutes.


