provefab

How to verify AI-generated pull requests

An AI-generated pull request is verified when it can show, not just claim, that it works: a reproduction for a bug that now passes, the repository’s own format, lint and test commands passing, no test removed or disabled to get there, and a review by a model from a different family than the one that wrote it. The evidence travels with the pull request, so the person who merges reads proof instead of redoing the checks.

Why AI pull requests pile up

Agents now open pull requests faster than teams can check them. Faros AI, across more than 10,000 developers, found that teams with high AI adoption “merge 98% more pull requests”, while review time grows by 91% and “AI adoption is consistently associated with a 154% increase in average PR size” (Faros AI). LinearB, over 8.1 million pull requests, reports that AI-assisted pull requests “merge at just 32.7%”, against about 84.5% for human ones (LinearB).

The failure mode is well documented. When the Copilot coding agent opened pull requests on dotnet/runtime in May 2025, some arrived with failing tests, and one Hacker News commenter summed it up: “It’s like you have a junior developer except they don’t even read what you’re telling them” (Hacker News). The bottleneck moved from writing code to trusting it.

The five things a pull request must prove

Check Why it matters How to automate it
A reproduction for a bug A fix without a failing case first may fix nothing The plan names a command that fails before the change and passes after
Your checks pass The repository’s own rules decide, not the agent’s opinion Run the same format, lint and test commands as CI, in the agent’s worktree
No weakened test The easiest way to turn a check green is to delete or skip the test Diff the tests: flag deleted tests, new skips and disabled assertions
A review by another model family A model tends to accept its own blind spots Route the review to a different vendor than the implementer
Evidence attached The reviewer should read results, not rerun them Write the plan, routing, check results and review findings into the pull request

A second model family is not magic. One July 2026 study of 116 tasks found that Claude reviewing Codex drafts raised the pass rate “from 71.6% to 89.7%”, while “Codex reviewing Claude drafts drops the pass rate from 91.4% to 82.8%” (arXiv 2607.21656). It is one study with two models, so measure the effect on your own repositories.

Doing it by hand

You can apply the checklist yourself on every agent pull request: check out the branch, run the commands, read the test diff, then ask a second tool for a review. It works, and it is exactly the work that makes review time grow. The value of the checklist comes from applying it every time, before a person looks.

Doing it with Provefab

Provefab applies the checklist to every labelled GitHub issue, on your Mac, with your own Claude Code, Codex or Pi sign-in:

  1. You label the issue. The label is the only authorization.
  2. An agent plans, read-only. For a bug, the plan must give a reproduction command that fails before the fix; if it already passes, Provefab stops.
  3. An agent implements in an isolated worktree, and a guard filters every tool call.
  4. Your checks run: the gates you configure, plus the reproduction command for a bug.
  5. A model from a different provider reviews. Each finding goes back to the implementer, who must write a test for it.
  6. The pull request lists the evidence: checks, routing, and any deleted or disabled test. A person merges; Provefab Pro can merge small, tested changes under a policy you set.

You can see the result on Provefab’s own repository: every pull request Provefab opened on itself, with its checks and reviewer.

Limits

Verification is only as good as the checks that exist. A repository without tests gets pull requests that pass nothing meaningful, and no reviewer model replaces a missing test suite. Provefab runs on macOS only today, and it uses your own model plans or API keys; for work, use plans whose terms allow commercial use.

Try it: the open core on GitHub, orProvefab Pro.