Quality Assurance·

The Pull Request Checklist: What Should You Test on Every PR?

What to run before a pull request merges (build, unit, scoped E2E, a small regression suite, secret scanning), what to schedule instead, and how to keep the gate under ten minutes.

QA

QA.tech

Contents

Short answer: a pull request checklist has two halves: test the change in depth, and run a small regression suite. For the change: unit and integration tests for the new logic, and end-to-end tests for the flows it touches, run against a preview environment. For the suite: the handful of critical paths that must never break, and nothing else. Add build, lint and secret scanning because they are cheap. Everything slow or environment-dependent runs on a schedule. The whole gate should finish in under ten minutes.

Why this question is back

Pull request volume moved. One engineering lead told us a release used to carry 15 PRs and now carries 150 to 200, on the same QA headcount. Another estimated that over 80% of his team's code is AI-generated, with a single engineer reviewing about 10 PRs a day. Faros AI's 2026 engineering report, drawn from 22,000 developers, put the median time a PR spends in review up 441.5% year on year.

When PRs sit for weeks, the usual cause is a regression suite that grew for years and now runs in full before every merge. Every team we talk to that has this problem believes the alternative is bugs in production. It isn't the only alternative. The gate can stay fast if it tests the right things.

Three questions the gate has to answer

  1. Does the new code work?
  2. Did it break something that must never break?
  3. Did it introduce a vulnerability?

The first question needs the most work and is where the tests should be new each time. The second needs a small, stable suite. The third needs a fast scan. Anything that doesn't answer one of these three doesn't belong on the merge gate.

The pull request checklist: what runs on every PR

Build

If the branch doesn't compile or a dependency fails to resolve, stop. A build check takes seconds and saves twenty minutes of downstream testing on a broken branch.

Smoke

Does the app boot, and do the main pages load? Run this before anything expensive. A PR that fails smoke shouldn't consume an E2E run.

Unit and integration tests for the change

This is where you answer "does the new code work." Unit tests for the new logic; integration tests for how it talks to the database or an external API. Cover the happy path first, then empty states, bad input and the edge cases the ticket mentions.

End-to-end tests for the changed flows

Once the logic works, check the user-facing behaviour: states, redirects, what the customer sees. E2E is slow, so scope it. If the PR touched payment, test purchases and retries and leave the rest of the product to the regression suite.

The problem is that hand-writing a fresh E2E suite for every PR doesn't scale. Teams that try it end up with thin coverage and week-long reviews, which is the original problem. This is where dynamic PR testing fits. The agent reads the diff, works out the blast radius of the change, generates test cases for the flows that change can reach, and runs them against the preview environment. The test cases are written for this change, so there is no script to maintain when the UI moves next sprint.

A small regression suite

Your regression suite should contain the flows that, if broken, would stop the business: sign-up, checkout, login, whatever your critical path is. It should be small enough to run on every PR and stable enough that a red is always worth reading.

That second property is why size matters. A suite of a few hundred tests with a 1% flake rate fails almost every run for reasons unrelated to the change. Developers learn to ignore it within a week, and then it catches nothing. A suite of twenty critical-path tests that never flakes gets read every time.

A suite bigger than that is a liability, and the job is to shrink it. Keep the critical paths, move what the change-scoped tests already cover out, and delete what neither protects. Our guide on when to kill a test covers the mechanics; our regression testing explainer covers what belongs in the suite in the first place.

Lint and formatting

Run both on every PR if your team has a style standard. A linter catches unused variables, shadowed names and small correctness traps. A formatter settles arguments before they reach review. Neither takes long.

Secret scanning and static analysis

Read the diff for exposed keys, tokens and passwords before it merges; a leaked credential in git history is expensive to revoke. Static analysis (SAST) for common patterns like injection and unsafe deserialisation is also fast enough for the gate. Deeper security work runs later.

Coverage threshold

Coverage tells you which lines executed under test. It's useful for spotting gaps. As a target it produces trivial tests, and a 100% figure says nothing about whether the suite would catch a real regression. Set a threshold that flags a PR adding untested logic and leave it there.

Green is not verified

A green gate means the checks you defined passed. It says nothing about behaviour no check covers. This is the gap that lets a PR merge clean and break checkout: the regression suite tested what the team already knew about, and the change introduced something new.

That is the argument for testing the change itself. The suite answers "did we break the old thing." A test written for this diff answers "does the new thing work." You need both answers before merging.

An example PR pipeline

In GitHub Actions or GitLab CI, on the pull_request event:

  1. Build. Compile, resolve dependencies.
  2. Smoke. App boots, main routes respond.
  3. Unit and integration. New logic and its interfaces.
  4. Lint, format, secret scan. In parallel with 3.
  5. Deploy preview. The E2E steps need a running instance of the change.
  6. E2E for the changed flows. Scoped to the diff.
  7. Critical-path regression. The small suite.
  8. Write the verdict back. As a check on the PR, with evidence attached.

Steps 1 to 4 should finish in a couple of minutes; 5 to 8 in under ten. We wrote up the workflow files for step 5 onward in our guide to running E2E tests on every pull request in GitHub Actions, and the GitLab equivalent side by side.

QA.tech integrations: GitHub App, GitHub Actions, GitLab CI, GitLab MR reviews, Jira, Linear

What not to run on the PR

Some tests are slow, need a stable environment, or produce results a human has to triage. Run them on a schedule or before release, where a re-run costs nothing.

Deep security scans. Dynamic scanning against the running app, full dependency and container audits, penetration testing. Noisy, slow, and someone needs to read the output.

Load and performance tests. They need production-like infrastructure and time to apply real load, and results vary run to run.

The full suite you haven't shrunk yet. If you still have a large regression suite, run it nightly while you cut it down. Treat every nightly failure that the PR gate would have caught as a test you can delete.

Choosing your own checklist

Match the product. A payments product needs more on the security side than a publishing tool. An internal admin panel needs less coverage than the customer-facing app. Decide by what breaks the business.

Give the gate a time budget. Ten minutes is the practical limit; past that, developers merge around the check. A budget also tells you when the regression suite has bloated: it's the day the gate misses the budget.

No flaky tests on the required list. A test that fails at random teaches the team to ignore red. Fix it or move it off the gate. We covered the two kinds and how to tell them apart in flaky tests: the two kinds.

Gate on the small set only. Make the critical-path suite and the change-scoped E2E run required checks. Report everything else without blocking until the team trusts its real-red rate. GitHub's branch protection and GitLab's merge checks both support this split.

Where QA.tech fits

As with any agentic tool, it won't catch everything. What it takes off the gate is the part that doesn't scale by hand: writing a fresh set of E2E test cases for every change.

QA.tech is an AI testing tool with GitHub PR integration, and the review it runs is exploratory: no script exists for the change until the agent writes one. When a PR opens, the agent classifies the change (a docs-only diff skips testing and posts a note), checks which existing test cases cover the affected flows, creates test cases for the gaps, and runs them in the preview environment. Typically that's 5 to 15 test cases per PR. The verdict lands on the pull request with the evidence: steps taken, what was expected, what happened, screenshots. Test cases it created are saved with an ephemeral label, so they don't pile into your regression suite unless you promote them.

QA.tech test case run showing goal, steps, and a passed assessment report with observations and final screenshot

The agent drives the app the way a user would, so there are no selectors to break when the markup changes. Where a flow must never drift, you pin explicit steps on the test case; where it just needs to work, you give it a goal. It runs through the GitHub App or GitHub Actions, on GitLab merge requests, and can open Jira or Linear issues from findings.

It also runs after merge. Teams without per-branch previews point the same review at staging once the PR lands; the change is still tested against a live application, the feedback just arrives after the merge instead of before it. Most teams start there and move the gate forward once their preview environments are stable.

QA.tech test case list with status and last-run history per test case

If your PRs are queueing behind a suite, book a demo and bring one.

Common questions

How long should PR checks take? Under ten minutes. Past that, developers stop waiting and merge around the check. Scoping tests to what the change touched is what keeps the gate inside the budget.

Should I run the full regression suite on every pull request? No. Run a small critical-path suite on every PR and test the change itself in depth. A full suite on every PR is slow, flaky at scale, and still misses behaviour no existing test covers.

Should a flaky test block a merge? No. A test that fails at random teaches developers to ignore failures. Fix it or take it off the required list.

What is the difference between a smoke test and a regression test on a PR? A smoke test checks the app starts and its main areas load. A regression test checks that existing behaviour, like sign-up or payment, still works after the change. On a PR, a small regression suite proves nothing critical broke.

Do I need a preview environment for PR testing? For end-to-end tests, yes. They need a running instance of the change. A per-PR preview is cleanest; shared staging works but serialises the pipeline. Without previews, run the same tests after merge against staging. The feedback arrives later, and the change still gets tested against a live application.

Your team moves fast. Can your testing keep up?

QA.tech agents test your product autonomously, so moving fast never means shipping broken. See how it works in a 30-minute demo.

Get a demo