VoiceOver Smoke Tests With guidepup on macOS

Permalink to "VoiceOver Smoke Tests With guidepup on macOS"

guidepup is an open-source library that drives real screen readers — VoiceOver on macOS and NVDA on Windows — from test code, and exposes what they spoke. Combined with Playwright, it lets you write a smoke test that opens a data grid, sorts a column with VoiceOver running, and asserts that VoiceOver actually said “sorted ascending”. It is slower and more fragile than a DOM test, and it catches a class of regression nothing else can: markup that is valid, passes axe, and is not announced.

This page sets up VoiceOver automation and writes a grid smoke test. It belongs to screen reader smoke testing and complements the NVDA and virtual-screen-reader approach in writing a screen reader smoke test with Playwright.

Spec reference

Permalink to "Spec reference"

guidepup (@guidepup/guidepup, @guidepup/playwright) provides:

  • voiceOver.start() / stop(); voiceOver.perform(voiceOver.keyboardCommands.<command>) for VoiceOver commands (move to next, interact with item, find next table…);
  • voiceOver.press(key) / type(text) for keys;
  • voiceOver.spokenPhraseLog() / lastSpokenPhrase() / clearSpokenPhraseLog() to read output;
  • @guidepup/setup — a one-time setup that enables AppleScript control of VoiceOver and required permissions on macOS (including CI runners).

macOS requirements: VoiceOver’s “Allow VoiceOver to be controlled with AppleScript” setting, Accessibility and Automation permissions for the terminal or runner process, and a logged-in GUI session. GitHub’s macOS runners support this after running the setup.

Criteria under test are typically SC 4.1.2 (names, roles, states announced), SC 4.1.3 (status messages heard), SC 1.3.1 (table headers).

A VoiceOver smoke test run Flow of an automated VoiceOver smoke test: setup permissions, start VoiceOver, open the page, drive the journey with VoiceOver and keys, assert on spoken phrases, stop VoiceOver. A VoiceOver smoke test runSetupguidepup setup,permissionsStartVoiceOver onOpen pageSafari or ChromiumDriveVO commands +keysAssertspoken phrase log
The assertions are on speech, not on the DOM — that is the whole point of the test.

When to automate VoiceOver — and when not to

Permalink to "When to automate VoiceOver — and when not to"

Automate a small set of critical journeys per release: sort announces, filter count announces, dialog opens and returns focus, treegrid expands. Run them on a macOS runner nightly or before release, not on every pull request — each test takes seconds to minutes and macOS runners are slower and costlier.

Do not try to automate exploratory testing or iOS VoiceOver; guidepup drives macOS VoiceOver only, and touch gestures are out of scope. Keep manual device testing for iOS, as in VoiceOver on iOS: the rotor and data tables.

The misapplication to name is exact-string assertions on full phrase logs. VoiceOver’s wording, punctuation and timing vary between macOS versions; tests that compare entire logs verbatim break on every OS update. Assert on the key phrase with tolerant matching.

Annotated code example

Permalink to "Annotated code example"
// voiceover.grid.spec.js
import { voiceOverTest as test } from '@guidepup/playwright';
import { expect } from '@playwright/test';

test.use({ browserName: 'webkit' });                 // Safari engine: VoiceOver's home

test('sorting the invoice grid is announced by VoiceOver', async ({ page, voiceOver }) => {
  await page.goto('http://localhost:8080/invoices');
  await page.getByRole('heading', { name: 'Open invoices' }).waitFor();

  // Move VoiceOver to the Amount sort button
  await page.getByRole('button', { name: 'Amount' }).focus();
  await voiceOver.clearSpokenPhraseLog();

  // Activate with the keyboard, as a user would
  await voiceOver.press('Enter');

  // Wait for speech rather than sleeping: poll the log for the status message
  await expect.poll(async () => (await voiceOver.spokenPhraseLog()).join(' | '), {
    timeout: 10000,
  }).toMatch(/sorted by amount,? ascending/i);        // tolerant: case, punctuation

  // Header state is also exposed
  const log = (await voiceOver.spokenPhraseLog()).join(' | ');
  expect(log).toMatch(/ascending/i);
});

test('filter count is announced once', async ({ page, voiceOver }) => {
  await page.goto('http://localhost:8080/invoices');
  await page.getByRole('searchbox', { name: 'Filter invoices' }).focus();
  await voiceOver.clearSpokenPhraseLog();
  await voiceOver.type('overdue');
  await expect.poll(async () => (await voiceOver.spokenPhraseLog()).join(' | '))
    .toMatch(/\d+ invoices? shown/i);
  const counts = (await voiceOver.spokenPhraseLog()).filter((p) => /invoices? shown/i.test(p));
  expect(counts.length).toBe(1);                      // debounced: exactly one
});
# .github/workflows/voiceover.yml (excerpt)
jobs:
  voiceover:
    runs-on: macos-14
    steps:
      - uses: actions/checkout@v4
      - run: npm ci && npx playwright install webkit
      - run: npx @guidepup/setup --ci              # enables VoiceOver automation and permissions
      - run: npm run serve & npx wait-on http://localhost:8080
      - run: npx playwright test voiceover.*.spec.js --workers=1   # one screen reader at a time

--workers=1 matters: there is one VoiceOver per machine, and parallel tests would fight over it.

Keyboard & AT behaviour

Permalink to "Keyboard & AT behaviour"
Journey step Assert on Tolerance
Sort by Amount /sorted by amount,? ascending/i Case, comma, extra words around
Filter typed Exactly one /\d+ invoices? shown/i Number normalised
Dialog opened /edit invoice .*, dialog/i Word order varies by macOS version
Dialog closed Trigger name re-read Allow “button” before or after
Treegrid expand /expanded/i Allow “row” and level text around it
Flaky patterns and stable replacements Matrix of flaky VoiceOver test patterns and the stable replacement for each. Flaky patterns and stable replacementsFlaky patternWhy it flakesStable replacementFixed sleep then readSpeech timing variesexpect.poll on the logExact full-log matchWording changes by OSRegex on key phraseParallel workersOne VoiceOver per machineworkers = 1No log clear between stepsOld phrases matchclearSpokenPhraseLog()Real network dataNumbers varySeeded data + normalisation
Most flakiness comes from timing and exact wording — poll for phrases, match tolerantly.

Integration context

Permalink to "Integration context"

The VoiceOver suite sits beside the NVDA and virtual-screen-reader suite from writing a screen reader smoke test with Playwright, sharing journeys and normalisation. VoiceOver’s handling of polite messages differs from NVDA’s — it can drop a polite message if another arrives quickly — which is why assertions poll rather than read once; see VoiceOver versus NVDA aria-live politeness handling.

When a VoiceOver test fails, capture the full phrase log as a test artefact; it is the equivalent of the NVDA speech logs in capturing NVDA speech logs for manual testing.

Virtual screen reader versus real VoiceOver Comparison of smoke testing with a virtual screen reader that reads the accessibility tree against driving real VoiceOver with guidepup. Virtual screen reader versus real VoiceOverVirtual screen readerAny OS, fast, per pull requestReads the accessibility tree directlyDeterministic outputMisses VoiceOver-specific quirksReal VoiceOver (guidepup)macOS runner, slower, per releaseReal speech, real timingCatches dropped messages and wordingNeeds setup, one test at a time
Run the virtual reader per pull request and real VoiceOver per release — they catch different things.

Gotchas

Permalink to "Gotchas"

VoiceOver left running. A failed test that skips teardown leaves VoiceOver on and speaking on a shared machine. Always stop it in afterEach/fixture teardown.

Safari versus Chromium. VoiceOver users mostly use Safari. Test WebKit first; Chromium with VoiceOver has different quirks.

First-run dialogs. VoiceOver’s welcome dialog blocks automation on fresh machines; the setup action disables it.

Design system notes

Permalink to "Design system notes"

Run a VoiceOver suite against the design system’s reference components on a macOS runner before each component release, and publish the phrase logs with the release notes. Product teams then know which announcements are guaranteed by the components and which they must test themselves.

Testing checklist

Permalink to "Testing checklist"

FAQ

Permalink to "FAQ"
Can VoiceOver be automated for accessibility testing?

Yes, on macOS. guidepup can start and stop VoiceOver, send VoiceOver commands and keys, and read the spoken phrase log, and it integrates with Playwright for browser control.

Can guidepup test VoiceOver on iPhone?

No. It drives macOS VoiceOver only. iOS VoiceOver, with its gestures and rotor, needs manual testing on a device.

Why do my VoiceOver tests flake?

Usually because of fixed waits, exact-string matching or parallel workers. Poll the phrase log for the expected phrase, match with tolerant regular expressions, and run one test at a time.

What should a failing VoiceOver test save?

The full spoken phrase log and a screenshot of the page at the failure. Together they show what VoiceOver said and what was on screen, which is usually enough to tell a timing problem from a markup regression.

Should VoiceOver tests run on every pull request?

Usually not — they are slower and need macOS runners. Run a virtual screen reader per pull request and real VoiceOver nightly or per release.

Permalink to "Related"

← Back to Screen Reader Smoke Testing