Dynamic testing is the practice of evaluating software by executing it – running the actual code with real inputs and watching how it behaves – as opposed to static testing, which examines code, requirements, and design documents without running anything.
That's the textbook definition, and it's still correct. In 2026 the phrase has picked up a second meaning though, and it's the one worth paying attention to: testing that is itself dynamic, performed by AI agents that adapt to the application at runtime instead of replaying fixed scripts. This guide covers both. The first explains why the second matters.
The classical definition: dynamic vs. static testing
Testing has always come in two halves.
Static testing looks at software without running it. Code reviews, linting, static analysis, requirements walkthroughs, architecture reviews. It catches defects early and cheaply, but only the ones visible in the artifact itself.
Dynamic testing runs the software and checks what it actually does. Unit tests, integration tests, system tests, end-to-end tests, performance and security testing. It's the only way to catch the defects that only show up in real execution: race conditions, integration failures, broken user flows, the environment-specific bug that never appears on paper.
| Static testing | Dynamic testing | |
|---|---|---|
| Code executed? | No | Yes |
| When | Earliest – before code even runs | Throughout development and CI/CD |
| Finds | Syntax issues, standards violations, design flaws | Functional defects, integration bugs, broken flows, performance issues |
| Examples | Code review, linting, static analysis | Unit, integration, system, E2E, regression testing |
| Cost of the bugs it catches | Low (caught on paper) | Higher – but these are the bugs users would actually hit |
Types of dynamic testing
Dynamic testing splits two ways: by visibility and by level.
By visibility, you've got black-box testing, which checks behavior against requirements with no knowledge of the internals (most functional and E2E testing); white-box testing, which exercises the internal structures directly (unit tests, coverage work); and grey-box, which mixes the two.
By level, it runs from unit testing (individual functions) up through integration testing (components working together), system testing (the full assembled product), and acceptance or end-to-end testing (whole user journeys in production-like conditions). Non-functional testing – performance, load, security, usability – is dynamic too. It executes the software, but it asks "how well" rather than "does it work."
The problem that broke classical dynamic testing
For two decades, dynamic testing at the E2E level meant scripts. Selenium first, then Cypress and Playwright. A script encodes one path with exact selectors – button#submit, div.checkout-total – and checks for exact expected states.
Here's the catch, and it's structural: these tests are dynamic in name only. They execute the application, sure. But the tests themselves are static artifacts. The UI changes and the selectors break. The flow evolves and the script is now testing yesterday's product. What you get is a maintenance queue that grows with every release, and the usual industry answer – "self-healing" that patches selectors after they've already failed – just treats the symptom while leaving the script-based cause in place.
Breakage is the visible cost. The quieter one is what a scripted suite does not cover. Every test is a mine you placed by hand, and the bugs live in the gaps between them. Adding more tests narrows the gaps and widens the maintenance queue at the same rate, which is why suites grow for years without anyone feeling safer.
Coding agents did not cause this. Scripted suites were already failing to keep up, because you cannot script a product that changes faster than you can write the scripts. What agents changed is the rate. Take an approach that was already broken, run it twice as fast, and you ship broken twice as fast.
The cost moved with it. Writing code got cheap; validating that the change works did not. That gap between how much gets written and how much gets verified is the actual problem, and it does not close by generating more test scripts.
Dynamic testing in 2026: tests that are themselves dynamic
The answer is a different architecture – call it agentic or autonomous testing – that makes the testing dynamic, not just the thing being tested.
It works from goals instead of scripts. A QA agent gets an intent, something like "verify a returning user can check out with a saved card," and works out the steps at runtime by reading the application the way a person would. No selectors are stored, so a DOM change does not fail the test by itself. That is not the same as no maintenance. You still write the goals, and you still correct the agent when it reads your product wrong. It is closer to onboarding a new colleague than to running a script: present at the start, more autonomous once it knows the product.
It keeps a living model of the product rather than a folder of scripts – a knowledge graph of how the app actually behaves, updated as the app changes. Coverage compounds instead of decaying. And because the agent explores instead of replaying, it turns up paths and edge cases nobody thought to script, which is exactly the class of regression that script-based dynamic testing structurally can't reach. All of it runs where the change happens, on the pull request, where the result is a check you can require before anyone merges, rather than a nightly suite someone triages over coffee the next morning. That does not mean the scheduled suite disappears. It means it gets small, and how small a regression suite should be is a question worth answering deliberately rather than by accumulation.
This is the sense in which QA.tech uses the term. Dynamic testing as testing done by QA agents that adapt to the product at runtime – the old goal, validate real behavior by running the software, finally done without the static-script bottleneck. For the wider picture, see what AI in quality assurance means today.
Scoping a run to the change
Running everything on every change is how suites get slow enough that people stop running them. Scoping means deciding, per change, what is worth executing.
In QA.tech that runs in five steps. The change is classified first, so a docs-only or infrastructure-only pull request skips testing and gets an informational comment instead of a run. User-facing changes go through coverage assessment, where existing tests are matched semantically to what the diff touched, typically 5 to 15 of them. Where coverage is genuinely missing, 1 to 3 new tests get generated, and most pull requests generate none. Those tests run against the preview deployment, and the verdict comes back as a review and a commit check.
The tests generated this way persist. They become part of the regression suite for future changes, which means the suite assembles itself out of what real changes actually needed, rather than out of what someone guessed at during a planning session.
The environment and the data around it
Scripts also constrain the environment they run in. A script says "click the first row"; the agent says "click the thing I just created", because it knows which one that is. Seeded, torn-down, mocked-out environments exist mostly because scripts need that determinism to stay green, and the sterility hides an entire class of bug. Stop populating a field upstream and the dashboard breaks in production while the seeded fixture keeps supplying it.
What it does not do
Two things it does not do. It needs an interface, so a backend-only service with no front end is out of scope for this kind of testing. And blocking a merge needs somewhere to run the change before it lands, which in practice means a preview deployment; without one you can still test after the merge, you just lose the block and the clean attribution.
Dynamic testing and AI code review are not the same thing
An AI reviewer that reads your pull request is doing static testing. A good one finds real problems, and it finds them before anything runs. But it reads the diff. It never opens the application.
The code is a blueprint for a building. It tells you where the rooms are meant to be. It does not tell you what the building looks like, whether the door opens, or whether the lift reaches the fourth floor. Dynamic testing at the pull request is the part that walks in and tries the door.
So the two answer different questions. Code review answers whether the change is written correctly. Dynamic testing answers whether the product still works for a user once the change is in. If you are running coding agents, you almost certainly need both, and most teams have only bought the first one.
Classical vs. agent-based dynamic testing
| Script-based (Selenium/Cypress/Playwright) | Agent-based dynamic testing | |
|---|---|---|
| Test definition | Code with selectors | Goals in natural language |
| When the UI changes | Tests break; engineers repair | Agent adapts at runtime |
| Coverage source | What someone scripted | Defined goals + autonomous discovery |
| Maintenance | Grows with the suite | Moves from repairing selectors to directing an agent |
| Best at | Custom logic, precise assertions | Broad, high-churn UI surface |
Most mature teams run both. A compact code-based layer for the deeply custom logic, and QA agents for the broad E2E surface where all the maintenance used to pile up.
For a tool-by-tool view of which platforms test dynamically and which replay scripts, see the best agentic QA tools in 2026.
See how QA.tech runs dynamic tests from your tickets, PRs, and coding agents.
