Short answer: Code review is the slowest step in the delivery cycle because AI made code cheap to write and did nothing for the cost of reading it.
Teams keep review for two reasons: to verify the change works, and to share knowledge. Reading AI-generated code line by line is a weak tool for both. Verification belongs in the running product, on the pull request. Knowledge belongs in the ticket, the spec and the evidence attached to the PR. The human reviewer stays, with a shorter question to answer.
What 440 engineers said in Copenhagen
On 1 October 2026, during Nathen Harvey's keynote at GOTO Copenhagen, 400+ people in the room were asked which stage of their SDLC causes the most friction today. Half of them picked reviewing changes. Building, the stage everyone has spent two years accelerating, got 3%. A conference poll, not a study, but 440 engineers is a decent room.
Live Slido poll, 440 responses, GOTO Copenhagen, 1 October 2026.
Harvey leads DORA at Google Cloud. In the 2025 DORA report on AI-assisted software development, 90% of nearly 5,000 respondents use AI at work, over 80% believe it has made them more productive, and 30% say they have little or no trust in the code it produces. The report finds AI adoption now correlates positively with throughput and still negatively with delivery stability.
So half the room says its time goes on reading code a model wrote, which, going by DORA's numbers, a fair share of them don't much trust.
Why review got slower
Nobody's reviewers got worse. What changed is what lands in front of them: more pull requests, bigger ones, and a growing share written by something the reviewer has no shared history with. Faros AI followed 22,000 developers through their teams' AI adoption and found the median review now takes five times as long as before, on pull requests that are half again as large. The code didn't get harder to read. There is more of it per change, and the person reading it has less context about why it was written that way.
That lack of context shows up as effort. In Sonar's January 2026 survey, 38% of developers said reviewing AI code is more work than reviewing a colleague's. DORA's 2026 ROI report calls the same thing a "verification tax" and lists it as one of the three reasons productivity dips after teams adopt AI tools. The gains from generating code are real, and a good part of them is being handed straight back at the review step.
We hear it on calls every week, in less tidy language. One HR-tech company we spoke to estimates more than 80% of its code is now AI-generated, and its CTO described the review side like this: "there was this one engineer who would review 10 pull requests a day, which is like, come on, somebody help me out here." Another team we work with went from about 15 pull requests per release to somewhere between 150 and 200, and nobody added reviewers.
What code review was ever for
Microsoft studied this in 2013. Of 873 developers surveyed, finding defects was the top motivation for code review, named first by 44%. Then the researchers read 570 review comments to see what review actually produced. 14% of comments were about defects and 29% were code improvements. About a dozen, out of 570, carried any knowledge transfer.
A follow-up from the same company two years later was titled "Code Reviews Do Not Find Bugs". About 15% of reviewer comments indicate a possible defect, far fewer a blocking one, and at least half of all comments are about long-term maintainability. Reviewers who had never seen the codebase before wrote comments the author found useful a third of the time, and once a change went past 20 files, usefulness fell off for everyone.
Both studies predate coding agents. Review was already a weak defect finder on human code at human volume. Its knowledge value was never in the comments either; it was in the reading, in one engineer seeing how another had solved something. Since then pull requests have grown by half, a growing share come from an agent the reviewer has no shared history with, and the median review takes five times as long. Models haven't changed the underlying problem either: they're good at writing and fixing code and much worse at finding bugs in it, for the same reason a developer is. Looking at code alone, it is hard to know whether the assembled thing will work.
Reading the diff is the wrong tool for both jobs
When teams say they can't cut review, they mean two things: we need to verify the change works, and we need people to know what's in the codebase. (The maintainability feedback that made up half the comments in the Microsoft data is a third job, and it's the one a code review agent with your rules handles well. More on that below.) The first two are real, and reading model-written code is a poor way to get either.
Verification. Whoever reads a pull request, person or model, is inspecting an artefact that hasn't run yet. Done well, that catches a missing null check or a secret committed to a config file before anything executes. It cannot catch what happens when a user clicks through the change against real data in a real browser. We've written about this distinction at length, so one sentence here: a reviewer reads the diff, and the only way to know what the product does is to run it.
W. Edwards Deming made the same argument about factories. Point three of his fourteen points: "Cease dependence on inspection to achieve quality. Eliminate the need for inspection on a mass basis by building quality into the product in the first place." (Out of the Crisis, 1982) Reading every AI pull request line by line is inspection on a mass basis. The only thing that changed since 1982 is how much there is to inspect.
Knowledge sharing. Reviewing a colleague's code used to teach you how that colleague thinks, and it caught the moment their mental model drifted from yours. Reviewing an agent's code teaches you how a model phrased something it will never phrase the same way again. Mob programming had a role for this, the typist: one person types what the group decides, and nobody learned anything by reading the typist's keystrokes back afterwards. A team with a coding agent is a mob with a very fast typist.
The knowledge that matters now, what the change was for, what must not break, which architectural rules apply, lives upstream of the diff, in the ticket, the spec and the PR description, and in whatever rules you've managed to write down. A lot of engineering knowledge still lives in people's heads, and the useful work is getting it into something the pipeline can check on every change, so it stops depending on who happens to be reviewing that day.
Where the review budget should go
Our view, and the way we run our own pipeline, is to spread what review was doing across four automated layers, and give the human a different job at the end.
Start with the rules you keep repeating in review comments. Any rule you can state precisely belongs in an invariant: a linter, a type, an architecture check that fails the build. These are cheap, deterministic, and you don't want an agent renegotiating them while it builds. Rules that need judgement, how data flows between services, what counts as a safe migration, go to a judge: a code review agent with your project's instructions loaded. A code review agent is useful even with no instructions at all, because it catches generally bad practice. With your rules in it, it becomes a check on the things that actually matter in your codebase, which a generic reviewer never was.
Then verify behaviour, on the pull request, against the deployed build. Review by itself is a floor. Add behavioural verification and you have a guardrail, since one catches bad code and the other catches broken products. The verifier should also not be the agent that wrote the change. There was always a bias in letting developers test their own work, and an agent checking its own pull request inherits it. Not the whole regression suite on every PR.
What a behaviour-verification verdict on a pull request looks like.
The flows the change can actually reach, worked out from the diff, the ticket and what's known about the product, run in a preview environment, with a verdict posted back to the PR. The verdict needs three states: pass, fail with evidence, and unverified when the test cases ran but never touched the change. A green check that means nothing was looked at is worse than no check, because somebody will merge on it.
Alongside that, keep a small regression suite of the flows that must never break, sign-up, payment, whatever your business dies without. We mean tens of test cases, run on a schedule, and when one of those goes red you treat it as an incident. Everything else gets tested when it changes. If your suite has grown into the thousands, the job is to shrink it, because at that size a few test cases fail for unrelated reasons on every run and the team stops reading the results.
Then let humans review, with a different question in front of them. Nobody needs to read every line, but the reviewer needs enough artefacts, checks and collected evidence that when they look at the PR the reaction is "yes, this makes sense" or "no, and here's why". Given that the invariants passed, the judge raised nothing serious and the behaviour was verified against the preview, does this change make sense for the product? A senior engineer can answer that in minutes, and it's the one question in the chain that should stay with a person. Skip this and the choice becomes a fork: either spend more time reviewing and offset the coding gains until you're shipping the same number of features as before, or trust that it works without validating it and let the bugs reach customers.
Where QA.tech fits
No test run catches everything, and nothing described here replaces your code reviewer. QA.tech, an AI testing platform, is the behaviour-verification step in that split. What it removes is the part of review where a human opens the preview, clicks through the change and tries to break it, because an agent has already done that and attached the evidence to the PR.
On each pull request the agents read the diff, the PR description and the linked ticket, decide which user flows the change can affect, generate test cases for that change, drive them in the preview deployment the way a user would, and post pass, fail or unverified back to GitHub as a check that can block the merge. Typically 5 to 15 test cases per PR. Teams without preview environments can run the same review after merge against staging. How dynamic PR testing works.
FAQ
Why has code review become the biggest bottleneck in the SDLC? Pull requests got bigger and more frequent once teams adopted coding agents, and developers say AI-written code takes more effort to review than a colleague's. Writing speed went up several times over. Review speed is bounded by how fast a person can read and understand a change, which has not moved.
Does AI code review fix the review bottleneck? Partly. An AI code reviewer reads the diff faster and more consistently than a person and catches generally bad practice, and with project-specific rules it catches your particular mistakes. It still only reads code. Whether the change works for a user is a separate question that stays open after the reviewer is done, and that is where most of the remaining review time goes.
Is reading AI-generated code a good way to share knowledge? No. Reviewing a colleague's code taught you how they think. Reviewing an agent's output doesn't, because the next change will be written differently. The knowledge worth sharing, what the change is for and what must not break, belongs in the ticket, the spec, the PR description and in rules the pipeline enforces.
Should humans still review pull requests? Yes, but the job changes. Instead of reading every line, a human looks at the collected evidence: which invariants passed, what the code review agent raised, and whether the behaviour was verified against the preview with proof attached. Then they answer the one question only a person can: does this change make sense for the product?
What is the difference between code review and dynamic testing on a pull request? Code review, human or AI, reads the changed code and judges how it is written. Dynamic testing runs the changed application and observes what a user gets. A reviewer can wave through a change that breaks the payment flow, because the payment flow was never in front of it. You want both on a pull request, and the second one is the one most teams don't have yet.
What should run on every pull request instead of a full regression suite? Invariants that fail the build on rule violations, a code review agent carrying your project's rules, and behaviour verification scoped to the flows the change can reach, run against a preview deployment. A small regression suite of critical flows runs on a schedule; everything else is tested when it changes.

