# A large regression suite is a liability, not an asset

> Big regression suites are slow, produce a spurious red on every run and rarely find the bug. Keep 50 to 200 critical-path test cases on a schedule and cover each pull request with dynamic testing.

Source: https://qa.tech/blog/large-regression-suite-is-a-liability · Published: 2026-10-01

---
<div class="qa-highlight">

**Short answer:** A regression suite stops being an asset somewhere around the point where nobody on the team can say what a red run means. Our position: keep the scheduled suite small, somewhere between 50 and 200 test cases depending on the product, and cover everything else with dynamic testing scoped to the change in front of you. A thousand tests re-run every night is the shape we argue against.

</div>

That's the whole argument. The rest of this piece is why we hold it, with the numbers.

## What a big suite actually costs

A big suite costs you in three places, and only the first one shows up on an invoice.

The first is time and compute. A suite that takes two hours to run gets run less often, and when it does run, the feedback arrives long after the developer has moved on. <a href="https://circleci.com/landing-pages/assets/2025-state-of-software-delivery-report.pdf" rel="noopener" target="_blank">CircleCI's 2025 State of Software Delivery</a>, built on 14 million workflows from September 2024, puts the median workflow at 2m 43s and the slowest five percent at more than 25 minutes. The report doesn't break the tail down by stage, but in our experience the test jobs are where most of it sits.

The second is flakiness, and this one is pure arithmetic. Google's test infrastructure team reported in 2016 that <a href="https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html" rel="noopener" target="_blank">about 1.5% of all test runs</a> across their corpus report a flaky result, and almost 16% of their tests have some level of flakiness. A year later they added the part that matters for suite size: <a href="https://testing.googleblog.com/2017/04/where-do-our-flaky-tests-come-from.html" rel="noopener" target="_blank">14% of their large tests were flaky</a>, against 0.5% of small ones. End-to-end regression tests are large tests by definition.

Now do the sum for your own suite. This is an illustration. Give each end-to-end test a 1% chance of failing for reasons that have nothing to do with the code, a slow response, a race, a third-party widget that timed out. Run 100 of them and you'll see roughly one spurious red per run. Run 1,000 and you'll see ten. <a href="https://www.linkedin.com/in/danielmaunopettersson/" rel="noopener" target="_blank">Daniel Mauno Pettersson</a>, QA.tech's co-founder and CEO, put it this way on a customer call in August:

> "If you have 100 tests to run every time, and you've only changed something that touches five of them but you still run all 100, and there's a 1% flake risk in it, then statistically you'll always have a test that fails. The best way to get rid of that is to not always run. Just run what matters."

The third cost is the one that hurts most and gets measured least. When every run has a red in it, the team stops reading reds. Slack's engineering team wrote up what that looked like for them before they fixed it: a main branch passing <a href="https://slack.engineering/handling-flaky-tests-at-scale-auto-detection-suppression/" rel="noopener" target="_blank">roughly 20% of the time, with 57% of build failures</a> coming from test jobs, and about 28 minutes of manual triage per failure. At that point the suite has stopped being a signal and become a tax.

## Why suites get big anyway

Nobody sets out to build a thousand-test suite. It accumulates. A bug ships, so someone writes a test for it. A feature launches, so someone adds the happy path. Someone runs a coverage report, it says 61%, and a sprint gets spent pushing it to 80%, which [doesn't tell you anything about what's being asserted](https://qa.tech/blog/why-100-test-coverage-isnt-the-goal). Then a test starts failing on a flow nobody owns any more, and because deleting a test feels like deleting coverage, it gets marked skip instead. Three years in, you've got a suite that's half skips, runs overnight, and that the newest engineer on the team is frightened of.

<a href="https://www.linkedin.com/in/vilhelm-von-ehrenheim/" rel="noopener" target="_blank">Vilhelm von Ehrenheim</a>, co-founder and Chief AI Officer, has a blunter version of this: you can have 10,000 tests and still have an extremely bad product, because they can all be green and test nothing your users care about.

The uncomfortable part is that deleting from the suite is the right move and almost nobody does it. We wrote a whole piece on [which tests to prune first](https://qa.tech/blog/when-kill-test-how-prune-test-suite-that-s-slowing). The short version: tests nobody can explain, tests that test the framework, tests duplicated three layers down, and tests that only ever fail for reasons unrelated to the code.

## What belongs in the suite that's left

Flows that must never break and that a user would notice within the hour. Sign-up. Login. Checkout or payment. The two or three things your product exists to do. For most products that's somewhere between 50 and 200 test cases. If you're in a regulated industry with documented evidence requirements it'll be at the top of that or above. If you're a four-person startup it might be ten.

The test for inclusion is simple. If this test went red at 2am, would someone get out of bed? If the answer is no, it doesn't belong in the scheduled suite.

This is also the suite where you want the same steps run the same way every time. If you already have a lot of Playwright covering these flows, keep it: you've paid for it, it's cheap to run per test, and it's good at exactly this job. If you don't, a scheduled suite of plain-language test cases on an agentic platform does the same job without anyone writing or maintaining scripts.

What you don't want in there is a test that checks that the dots are green in a particular view. Too fine-grained, too likely to block a release because something looks slightly off, and too expensive relative to what it tells you.

## What covers everything else

The change, and specifically the pull request. When a PR opens, something has to look at what changed, work out which parts of the running product it can affect, and go and try those parts. Deeply, once, against the preview build. The generated test cases are kept as drafts outside the main list; a few get promoted into the scheduled suite after merge, and the rest stay out of the nightly run. If the same area changes again next week, generate again from next week's diff. [Dynamic testing](https://qa.tech/blog/what-is-dynamic-testing) in our sense means the test cases are derived from the change in front of you instead of replayed from a script written for an earlier version of the product.

The reason this works where the big suite doesn't is scope. A PR that touches the invoice export doesn't need the sign-up flow re-verified. It needs invoice export tried properly, with the edge cases, on the build that contains the change. QA.tech, an agentic testing platform, generates five to fifteen targeted test cases per PR on [PR testing](https://qa.tech/product/pr-testing), and because they're scoped to the diff the flake arithmetic above works for you.

It won't catch everything. Side effects buried in a stored procedure three services away are not something a diff-scoped agent will find, and nothing in this article changes the need for unit and integration tests underneath. What it removes is the reason to keep growing the end-to-end suite forever.

## The objection we hear most

"Once a test passes, I want it static and dumb. Same steps, same way, every time. Anything adaptive feels like instability."

That's a fair position, and it's exactly right for the small scheduled suite. For the rest, the question to ask is what the suite is for. If it exists to catch regressions in code that changed, running everything that didn't change is noise, and at a thousand tests the noise is guaranteed. Most "flaky" failures we see in customer suites turn out to be one of three things: a real bug, a misconfigured suite (parallel tests logging each other out is the classic), or a race condition in the product. Very few are the agent.

## Where to start if your suite is already too big

Pick the 50 flows that would get someone out of bed. Move them into a scheduled suite that runs on main and on release. Everything else goes into a quarantine suite that runs weekly, and each week you delete whatever failed for a reason you can't explain. Put dynamic testing on the PR so the long tail is covered by the change that touches it, and watch the quarantine suite shrink. After a couple of months the quarantine suite is usually a fraction of what it was and nobody misses it.

The companion to this piece is our [regression testing pillar](https://qa.tech/blog/what-is-regression-testing), which goes into suite sizing, what to do with the tests you remove, and how to pair a small suite with PR testing.

## Frequently asked questions

**How big should a regression suite be?**
Big enough to cover the flows that must never break, and no bigger. For most SaaS products that's 50 to 200 end-to-end test cases. Hand-written scripted suites tend to sit at the low end of that, because every script is a maintenance commitment.

**Isn't a bigger suite safer?**
Only if every test in it is trusted. Past a certain size the suite produces a spurious failure on every run, the team stops reading failures, and real bugs ride through under the noise. A suite nobody trusts is less safe than a small one everybody does.

**What should I do with the tests I remove?**
Don't delete them on day one. Move them to a quarantine suite that runs weekly, and drop whatever keeps failing for reasons unrelated to the code. Cover the areas they used to cover with testing scoped to each pull request.

**Does dynamic testing replace the regression suite?**
No. The model is a small scheduled suite for the critical path plus dynamic testing on every PR. The two cover different risks: the suite catches "something that always worked broke", the PR tests catch "the thing you just changed doesn't work."

**Why do large tests flake more than small ones?**
More moving parts. <a href="https://testing.googleblog.com/2017/04/where-do-our-flaky-tests-come-from.html" rel="noopener" target="_blank">Google's 2017 analysis</a> found flakiness correlated with test binary size and memory use, and that 14% of their large tests were flaky against 0.5% of small ones. End-to-end tests touch the network, the browser, third-party services and real timing, and each of those is a way to fail that has nothing to do with the code.
