# How to Verify the Work of Coding Agents: Four Layers, and the One Your Pipeline Doesn't Have

> AI code verification in four layers: static checks, the agent's own tests, review agents, and the running product. Only the last one opens the app.

Source: https://qa.tech/blog/how-to-verify-coding-agent-work · Published: 2026-10-09

---
A VP of engineering at a European marketplace told us in September that more than 80% of his code is now written by agents, and that one of his engineers had started reviewing ten pull requests a day. "Come on, somebody help me out here," was how he put the engineer's complaint. A head of engineering at an ad-tech company said his releases used to carry 15 PRs and now carry 150 to 200, with the same QA headcount. He described each production push as rolling a sketchier dice than the last.

That is the situation this guide is for. The agent writes the change. Somebody, or something, has to establish that the change works before it ships. There are four places that check can happen, and they catch different things.

<div class="qa-highlight">

Verification of agent-written code happens in four layers: static checks that run in seconds, tests the agent wrote for itself, a review agent that reads the diff, and a check on the behaviour of the running product. The first three read code. The fourth opens the app. Both teams above had the first three wired in and a person doing the fourth, which is the layer that doesn't scale with PR volume.

</div>

## Layer 1: static checks

Linters, type checkers, formatters, structural tests that enforce architecture rules. They're deterministic, they run in seconds, and an agent can run them itself before it opens a PR. Stripe's Minions run a lint selection on every push that [takes under five seconds](https://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agents), and OpenAI's team shipping an internal product with no hand-written code enforces its conventions [mechanically through custom linters and structural tests](https://openai.com/index/harness-engineering/).

Vilhelm von Ehrenheim, our co-founder and Chief AI Officer, calls the linter the simplest invariant. Something that has to stay true about the codebase, encoded once, checked on every change, never decaying. If you keep thinking "oh no, this again" when you read agent output, that's a rule waiting to be written down as a check instead of a comment. A passing linter tells you the code is shaped correctly and nothing about whether the feature works.

## Layer 2: the tests the agent wrote

Ask a coding agent to build a feature and it writes tests for it, runs them, and reports green. This is the layer teams lean on hardest, and the one with the weakest evidence behind it.

An engineering lead at a logistics SaaS builds his frontend with Claude, tests included. His team started seeing regressions anyway, and the bugs were, in his words, logical things the agent had missed: functionality that simply hadn't been built, which is hard for an agent to notice from inside the task. Our account executive on that call summed it up as the agent marking its own homework.

Daniel Mauno Pettersson, our CEO, put the structural problem this way in a webinar in June: whenever the agent changes the code, it changes the test alongside it. If you don't trust the change, you now have to trust the change in the test too.

The people writing about agent harnesses say the same. Birgitta Böckeler's [harness-engineering article](https://martinfowler.com/articles/harness-engineering.html) on martinfowler.com lists an AI-generated test suite as the feedback sensor for application behaviour and then concedes, a paragraph later, that this asks too much of those tests for now. Simon Willison's note on StrongDM's agent-only setup makes the narrower point: agent-written tests only help if the agent isn't [asserting true](https://simonwillison.net/2026/Feb/7/software-factory/) to get them green.

This layer is worth keeping for the agent's own regressions during the task, as long as nobody reads its green as a verdict on the PR.

## Layer 3: review agents

CodeRabbit and its peers are worth having. They read the diff, spot the missing null check, flag the function that's grown to 400 lines, and they do it before a human opens the PR. We compared the field in [our roundup of AI PR code review tools](https://qa.tech/blog/top-5-ai-pr-code-reviewers-2025).

Their limit is the input. A reviewer, human or agent, reads the change. It doesn't run the product. It can't know that the new health badge renders on the pipeline board and the deals list and is missing from the deal's own page, because that fact only exists when the three components meet in a browser. Vilhelm's line from his CTO Craft talk in September: the reviewer reads the diff, a verifier can actually run the product. Review is the floor. The big labs are drawing the same line in their own pipelines: [verifier agents in Dots and Muse](https://qa.tech/blog/verifier-agents-dots-muse-codex-security) check the agent's actions, and nothing checks the product.

Review agents also have a volume problem of their own. The reviewer's comments land on a human who has to read them, and when Nathen Harvey polled the room during his GOTO Copenhagen keynote on 1 October, [half of 440 engineers](https://qa.tech/blog/code-review-bottleneck-ai-generated-code) named reviewing changes as the stage causing the most friction in their cycle. More review output doesn't shrink that queue.

## Layer 4: the behaviour of the running product

This is the layer a human does today. An engineer at a fintech startup described their process to us: there's a staging environment, the product manager sits with one of the developers, and they go through it. Daniel described the same thing from the other side: you check out your colleague's branch and run through a bunch of user journeys by hand, and if something looks unintended the back-and-forth that follows takes longer than the fix. The check is the right one; doing it by hand doesn't survive 150 PRs a release.

Here is the same check done by an agent. A pull request opens and a preview deploys. An agent that didn't write the change [reads the PR description, the linked Jira or Linear ticket, and the changed files](https://docs.qa.tech/best-practices/pr-testing), works out which user-facing flows the diff can affect, [selects the existing test cases that cover them](https://docs.qa.tech/pr-testing/overview), sets how deep to test from the blast radius of the change, and writes new test cases only where there's a gap. It runs them against the preview and posts a review on the PR with a verdict, a summary and a results table, with the screenshots and step trace for each row one click away in the run. The verdict has three values: pass, fail, or [unable to verify](https://docs.qa.tech/pr-testing/overview), meaning the run never exercised the change. The third one matters, because a green run that never touched the diff tells you nothing about the change.

It won't catch everything: the agent checks behaviour through the UI and API responses, so a side effect that never surfaces there, a stored procedure three services deep, is out of scope. What it removes is the hand-run pass through staging. That's PR testing at [QA.tech](https://qa.tech/product/pr-testing), an AI testing platform with GitHub PR integration, and it's also, in our experience, the one layer where the question "did we build the right thing" gets asked by something other than a person. The three layers above it answer narrower questions: is the code well-formed, does it pass the tests it wrote, does the diff look right.

The verdict is advice. You can require the [status check in branch protection](https://docs.qa.tech/pr-testing/github), but the merge is still your decision, which is where Vilhelm lands too: low-risk PRs can go through on their own, everything else should still be looked at by a person who now has evidence in front of them instead of a diff.

## Which layers for which pull request

Not every PR needs all four, and the cost of each layer is different. Vilhelm's test for whether a check belongs in the agent's loop at all is whether the agent can run it without a human and get back output it can act on; he sets that out in [the harness piece](https://qa.tech/blog/harness-engineering-behaviour-sensor). For layer 4 that means the failing assertion, the screenshot of the broken state and the network call that returned a 500. A red badge on its own, as Vilhelm put it, is very hard to dig the reason out of; make the fail actionable and the agent can climb towards what good looks like.

A rough split:

- Refactors and dependency bumps: layers 1 and 3, plus the existing regression tests at layer 4. Zero new test cases. The product shouldn't behave differently, so the check is that it doesn't.
- Bug fixes: all four, but layer 4 is the handful of existing tests that cover the affected flow.
- New features, especially ones a product manager or designer shipped through an agent: all four, and layer 4 is where the new test cases get written, because nothing else in the pipeline knows what the feature is supposed to do for a user.
- Docs and infra-only changes: layer 1 and a human glance. Nothing to run.

One customer told our sales team they'd check 80% of the agent's findings by hand at the start and drop to 20% once the hit rate held, which is the ramp we'd suggest for any new verifier.


## FAQ

### How do you test code written by an AI agent?

Static checks (lint, types, structural rules) run in seconds and the agent runs them itself. The agent's own unit tests catch its regressions during the task. A review agent reads the diff before a human does. A verification agent opens the deployed preview and checks the user-facing behaviour the change affects. The first three are standard CI; the fourth is the one your pipeline hands to a person.

### Can a coding agent test its own work?

It can and does, and that's the problem with relying on it. Daniel's framing above applies: if you don't trust the change, you can't trust the test the same agent wrote for it. Keep agent-written tests as a development aid and get the behaviour checked by something else.

### Is an AI code review agent enough?

No. A review agent sees the diff and nothing of what the diff does once it runs against real state in a browser. Keep it, because it catches real defects cheaply, and pair it with a check on the running product.

### What is dynamic PR testing?

Testing the pull request's preview deployment, with the test cases chosen from what the diff can affect. [What is dynamic testing](https://qa.tech/blog/what-is-dynamic-testing) covers the term; the short form is that static analysis reads code and dynamic testing runs it.

### Should agent-written PRs be auto-merged?

Some. Low-risk changes with all four layers green are the candidates. Anything touching a user-facing flow should still get a person. Even Stripe's Minions, which run with no human-written code, have [a human read every PR before merge](https://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agents-part-2).

### How does this relate to the agent harness?

The harness is everything around the model: tools, memory, permissions, and the checks that run on its output. The four layers are the checks. We've written about [the sensor most harnesses are missing](https://qa.tech/blog/harness-engineering-behaviour-sensor) and about where verification sits in a software factory.
