Voice testing

Test user flows that run on voice.

Your app has a microphone button, and until now that's where test automation stopped. QA.tech agents speak into your app's mic input and verify what happens next, on the same AI QA testing platform that powers our web, mobile and API agents.

Trusted by high-performing engineering teams at:

From "user speaks" to "it worked".

Put voice flows in the pipeline

The agent plays your audio clip into the browser microphone during a test run. Your app hears a user speaking. The flow runs on every trigger you already use: a PR, a deploy, a schedule.

Verify the outcome, not the vibe

Assert on what the speech should cause: the transcribed text in the field, the search results, the created note, the request your app sent to its transcription endpoint.

No brittle audio API scripts

Describe the step in natural language: “play the audio into the microphone, then check the transcript.” The agent grants mic permission, streams the file, and moves on. No WebRTC mocking, no fake-device plumbing.

320h/mo
QA time saved

"We have replaced over 320h of manual testing every month with QA.tech."

Fredrik Seidl
Fredrik Seidl – CTO at Upsales
Upsales is a SaaS platform for CRM, marketing automation & Sales in B2B. ~150 employees.
Read their story

Voice flows tested unattended.

Speak any phrase into your app's mic

Attach an Audio Input config to your project with the speech clip you want the agent to say. In the test, one natural-language step plays it into the browser microphone at exactly the right moment, right after the mic feature is activated.

  • Upload your own clips: phrases, commands, long-form speech
  • The agent handles mic permission grants on its own
  • Reuse the same clip across dictation, search and command tests

Test every flow the microphone triggers

If the mic starts it, the agent can test it. Speak a phrase and confirm the dictated text lands in the field. Feed a longer clip and check the produced transcript. Say a search query and verify the results.

  • Dictation and speech-to-text accuracy in real flows
  • Voice search, voice notes and voice messages
  • Assistant-style commands and the UI actions they trigger

Assert on UI state and network traffic

The agent supplies the input; you verify the outcome the same way as any other test. Check the visible result, or assert that the browser sent the right request to your speech endpoint: method, status, payload and response.

  • UI assertions on transcripts, results and triggered actions
  • Network assertions on your transcription/ASR calls
  • Failures come with the full evidence trail, like every QA.tech run

Runs where your other tests run

Voice tests are ordinary test cases. Put them in the plan that runs on every pull request, trigger them from CI after a deploy, or run them on a schedule next to your critical-path regression plan. Testing a voice feature in a native app? See mobile testing.

  • Trigger from GitHub, GitLab or any CI via the REST API
  • Include voice flows in AI PR reviews on preview deployments
  • Results land in Slack, Teams or your tracker like any other run
More on PR testing

Everything a voice flow needs to run unattended.

Audio Input configs

Attach speech clips to a project once, reference them from any test with a natural-language step.

Built for CI/CD and PR workflows

Run voice tests on every PR or merge or trigger them from any pipeline via the REST API.

Network request validation

Assert your app called its speech endpoint with the right method, payload and response.

Chat-built tests

Describe the voice flow in chat and the agent drafts the test, audio step included.

Protected environments

Test voice features on gated previews and staging via Vercel bypass, WAF rules or IP allowlisting.

Alerts where you work

Slack and Teams notifications on completion or failure, per project or per plan.

What is voice testing?

01

The category

Voice testing is verifying the flows a microphone starts: dictation, speech-to-text, voice search, voice notes, and assistant-style commands. More products listen than ever, and the testing hasn't caught up – clicks and forms are automated, while the mic flow gets a developer talking at a laptop before release, or nothing at all.

02

Where it gets hard

Voice input sits outside what a scripted framework can reach. Faking a microphone means WebRTC mocking and fake-device plumbing, and the interesting part – what the app does with the speech – depends on a model, a prompt and an SDK that all change under you. Every model swap can quietly break the path between “user speaks” and “the right thing happens”, and no scheduled suite will catch it.

03

How agents change it

QA.tech agents grant mic permission, stream your audio clip into the browser as live microphone input, and then verify the outcome the way a tester would: the transcript in the UI, the triggered action, the request sent to your speech endpoint. It runs on the same AI QA testing platform as our web, mobile and API agents, so voice flows live in the same plans and the same PR checks.

Where voice testing goes next.

Today, QA.tech voice testing is about audio going in: simulating a user speaking into a web app's microphone in Chrome, with the outcome verified through UI state and network traffic. Teams building AI products keep asking us for the other direction, and it's where this is headed.

  • Audio output validation – confirming that your app actually played the right sound, TTS response or call audio.
  • Video playback validation – asserting on what a media player rendered, not just that the page around it works.
  • Voice conversations – multi-turn spoken exchanges with assistant-style products, including interactive avatars.

Building something that speaks back? Talk to us – these capabilities are being shaped with design partners now.

Frequently asked questions

Can't find what you're looking for? Reach out.

How does QA.tech test voice input?
You attach an audio file to your project as an Audio Input config. During a test run, the agent grants microphone permission and streams that file into the browser's mic input, so any feature that listens behaves exactly as if a real person spoke. You then assert on the result: the transcript in the UI, the triggered action, or the network request your app sent.
What voice features can I test?
Any browser flow the microphone triggers: dictation and speech-to-text, transcription of longer clips, voice notes and voice messages, voice search, and voice commands or assistant-style input.
Can QA.tech verify that my app played audio correctly?
Not yet. Voice testing currently covers audio input only, so the agent can't assert on playback quality or what came out of the speakers. Audio output validation is on the roadmap; if that's your use case, talk to us about the design-partner program.
Does this work for phone systems or IVR?
No. Voice testing runs in Chrome against a web page's microphone input. Telephony and IVR flows outside the browser aren't in scope.
Do I need to record my own audio clips?
Yes. You supply the speech clip in the format your app expects, which also means the agent says exactly what your test needs: your product's vocabulary, your users' phrasing, your edge cases.
Can voice tests run on every pull request?
Yes. Voice tests are regular test cases, so they run wherever your other tests run: in AI PR reviews against preview deployments, triggered from CI, or on a schedule.

Your app listens. Your tests should too.

See a QA.tech agent speak into your app and verify the result, live, in a 30-minute demo.

Get a demo