Short answer: The most interesting thing OpenAI shipped at DevDay this week is the agent standing behind each Dot, judging what it's allowed to do. Meta did the same thing three weeks earlier with Muse: one agent that acts, a second one, sealed off from the first, that approves or blocks every action before it reaches the internet. The pattern has a name now, the verifier agent, and in 2026 it has become a standard part of how anyone serious builds an agent harness.
Which raises an obvious question for anyone shipping software with AI. Your coding agent has a verifier. Your security scanner has a verifier. Who's verifying the product?
What actually launched
Dots, announced 29 September, are OpenAI's always-on agents. Each one runs on its own cloud computer, browses, uses connected apps, works on long projects without being prompted. The safety layer, reported by Axios as a system OpenAI calls Guardian internally and Auto-review publicly, evaluates each proposed action and decides whether a human has to approve it. OpenAI first shipped that in Codex earlier this year; their own write-up describes it as sending the planned action and recent context to an auto-approval subagent that can wave through low-risk steps on its own.
Muse, Meta's personal agent, launched in early September. Same architecture, stated more bluntly: Muse runs on its own dedicated cloud computer, and a separate Sentinel agent runs on the same machine, kept apart at the system level. Nothing Muse does reaches the internet unless the Sentinel approves it. Permissions are scoped per service: Muse can read your mail, or read and send on your behalf, and you choose which. Forbes reported staff flagging security flaws at launch anyway, which is its own lesson about verifiers.
The code side had already been moving for a year. OpenAI's Aardvark, an agent that reads a repository and hunts for vulnerabilities, identified 92% of known and seeded vulnerabilities in its benchmark repos and became Codex Security in March, where Axios reported it finding nearly 800 critical and over 10,500 high-severity issues during testing. Google DeepMind's CodeMender has upstreamed 72 security fixes to open-source projects. Anthropic's Claude Code gained a security review command and a GitHub Action that runs on every pull request. At DARPA's AI Cyber Challenge final, the top systems found 54 of 63 planted vulnerabilities and patched 43, at roughly $152 a task.
The pattern underneath all of it
Strip the branding and every one of these is the same shape. One agent does the work. A second agent, with a narrower job and no stake in the outcome, checks it before it counts.
Addy Osmani's framing is agent equals model plus harness, and he cites a team that moved a coding agent from the top thirty to the top five on a benchmark by changing nothing but the harness. The verifier is the harness component that's grown fastest this year, and I think the reason is simple. Everyone who ran agents in production in 2025 learned the same thing: an agent's own confidence in its output is worth very little. The findings had to be checked by something else before anyone would act on them.
You can see that lesson written into the security tools. Aardvark doesn't report a vulnerability until it has tried to trigger it in a sandbox. CodeMender runs an LLM critique, differential testing and fuzzing on every patch before proposing it. Even XBOW, the autonomous pentester that hit the top of HackerOne's US leaderboard, routes every submission through a human security team. Unverified findings turned out to be noise, and the category learned it the hard way.
The thing none of them check
Line them up and almost every verifier in that list reads code. Aardvark, Codex Security, CodeMender, Claude's security review, the AIxCC entrants: they start from the repository or the diff. Guardian and Sentinel judge actions before they execute. XBOW is the one that works against a live target, and it's looking for a way in. Whether the checkout works is nobody's question.
Nobody in the lineup opens the product you just shipped and tries to use it.
That's a strange gap, because it's the oldest verification problem in software. The code review passed. The security scan passed. The unit tests are green. And the invoice export still produces an empty PDF, because the bug lives in how four components behave together when a real user does a real thing in a real browser, and nothing upstream ever ran that. In my experience models are better at fixing code than at finding bugs in it, and that isn't strange: looking at the code alone, even a human developer can't be sure the whole thing works, because so many parts only meet at runtime.
Code review reads the diff. Dynamic testing opens the app. Those are different kinds of evidence, and the second one is what's missing from the harness most teams have built.
What a verifier for the product looks like
Same trigger as a security review, different evidence. The pull request opens, a preview build deploys, and an agent that didn't write the change goes and uses it. It works out from the diff which parts of the running product the change can affect, generates test cases for those parts, typically 5 to 15, runs them against the preview, and reports back with screenshots and a trace. If the test cases ran without ever exercising the change, the verdict is "unable to verify", posted as a comment on the PR, and there's no green to misread.
Three properties matter, and they're the same three the Guardian and Sentinel designs settled on.
It has to be separate from the thing it's checking. A coding agent verifying its own output has the same bias a developer testing their own code does. The verifier is a different agent, with different context, looking from the outside in.
It has to reproduce. A verifier that says "this might break the checkout" is a code reviewer with extra steps. One that tried the checkout on the preview build and has the screenshot is evidence.
It has to leave the consequence to a human. Every system above does this. Dots don't send messages unless asked, and hand consequential money moves back to the user. Sentinel asks for permission. Big Sleep's findings get a human before disclosure. The honest level of autonomy in 2026 is autonomous investigation with a human-approved consequence, and for a product verifier the consequence is the merge.
That's what PR testing is at QA.tech, an agentic testing platform: the verifier for the running product, slotted into the same place in the pipeline as the security review, producing behaviour as evidence where the others produce a reading of the source. It won't catch everything. A side effect buried in a stored procedure three services away is beyond what a diff-scoped agent can reach from the UI. What it removes is the gap between "the code looks right" and "a user can do the thing", which is where most of the embarrassing bugs have always lived.
Where this goes
I said in a talk earlier this year that to move fast and be safe when thousands of agents are building things for you, you can't sit around and check everything. You need a system where what good looks like is already codified. Linting was the first verifier. Types were the second. Unit tests, security scans, and now a guardian agent in the harness. The verifier for the product is the next one in the row, and the companies shipping Dots and Muse have just made the case for it better than any testing vendor could.
Frequently asked questions
What is a verifier agent? A second AI agent whose only job is to check another agent's proposed actions or outputs before they take effect. OpenAI's Guardian (publicly Auto-review) in Codex and Dots, and Meta's Sentinel in Muse, are current examples. The verifier is isolated from the agent it checks and typically escalates to a human for high-risk actions.
Is a security review agent the same as testing? No. Security review agents such as Codex Security, CodeMender and Claude's security review read source code and look for vulnerability patterns, often confirming them in a sandbox. Testing, in the sense used here, runs the deployed application and checks that it behaves correctly for a user. They catch different classes of bug and sit at the same point in the pipeline.
What is an agent harness? Everything around the model that makes it an agent: the tools it can call, the memory, the loop that plans and retries, the permissions, and the verifier that checks its work. The same model with a better harness performs materially better, which is why harness design is now where much of the engineering effort goes.
Does QA.tech's PR testing replace code review or security scanning? No. It runs alongside them on the same pull request and produces a different kind of evidence: what the application did when the change was exercised in a browser. The scanner covers what the source looks like. Keep the code reviewer and the scanner.
Why do these agents still need a human to approve things? Because the companies shipping them have concluded that autonomous investigation is reliable enough to run unattended, and autonomous consequence is not. Dots are not meant to send messages unless asked, and hand some consequential financial transactions back to the user; Muse's Sentinel asks for permission. For software, the equivalent is that the agent tests and reports, and a person decides whether to merge.
