James Coombs, a design engineer specializing in AI agent code-review skills, pulled four weeks of data across 49 repositories to test whether code review infrastructure was keeping pace with AI-driven generation. The results expose a critical disconnect: while dashboard metrics suggest healthy review activity, the human element has effectively vanished from half the workflow. Between August 24 and September 20, 2026, 1,239 pull requests merged in a single engineering org, with 90.7% flagged as AI-assisted. Crucially, 50.6% of these PRs merged without any review from a human account, meaning the verification step was entirely automated.

The Bot Review Illusion

The metrics look healthier than they are because most tools count bot activity as review coverage. In the final week of the analysis period, bots generated 635 of the 898 total reviews. When Coombs isolated human-only reviews, the data inverted: the share of PRs with no human review rose from 47.7% in July to 61.6% in the most recent week. This creates a dangerous feedback loop where agents review agents, and the system launders its own output. Coombs notes that when at least half the merges have no human reviewer, review has quietly become a signature, and increasingly a machine's.

Why Mandating Rigor Fails

The intuitive fix—mandating stricter human review or adding checklists—fails because it ignores adoption costs. Coombs cites an ablation study where behavioral rules in AI agent configurations had a 0% compliance rate when unenforced. The same logic applies to developers: if a review checklist adds five minutes to a two-line change, it gets skipped. The gap isn’t a rigor problem; it’s an adoption problem. A check that is routinely skipped provides less assurance than a lighter check that actually runs, especially when the skipped check manufactures false confidence by appearing in the audit log.

Re-Engineering Review for Volume

To survive the volume of AI-generated code, review processes must be scaled to risk rather than applied uniformly. Coombs proposes four structural changes: tiering review requirements based on blast radius, ensuring every required evidence item has a practical capture path, prioritizing artifacts like logs and screenshots over prose claims, and refusing to let systems verify their own output without explicit flags. The goal is to prevent unsatisfiable items that discredit the entire review system, ensuring that human attention is reserved for high-risk changes where it actually matters.

Key Takeaways

  • 50.6% of AI-assisted PRs in the studied org merged without any human review.
  • Bot reviews inflate coverage metrics; human-only review coverage dropped, with 61.6% of PRs having no human review in the final week.
  • Mandating detailed checklists for low-risk changes leads to 0% compliance and erodes trust in the process.
  • Effective verification requires risk-tiered requirements and artifact-based evidence, not prose claims.

The Bottom Line

We automated code generation but forgot to scale human verification, and relying on bot reviews to fill the gap is just laundering our own output. Until teams split their metrics between humans and bots and demand artifact-based evidence for high-risk changes, the 'verification gap' will remain a silent, systemic risk.