pa11y-ci for Multi-State Data Pages

Permalink to "pa11y-ci for Multi-State Data Pages"

pa11y-ci runs automated accessibility scans over a list of URLs and fails CI when issues are found. It is simple to adopt: a JSON config, a list of pages, one command. Its default — load the page, scan it — is also its weakness for data interfaces, where most of the accessibility surface only exists after interaction: the sorted state with aria-sort, the filter popover, the edit dialog, the expanded treegrid row, the “no results” empty state. A scan of the initial load tests maybe a fifth of what users will meet.

pa11y’s actions script those states before scanning. This page configures pa11y-ci to scan data pages in all their important states. It belongs to automated accessibility testing pipelines.

Spec reference

Permalink to "Spec reference"

pa11y-ci (v3) reads .pa11yci or a JS config with defaults and urls. Each URL can be a string or an object with:

  • url — the page;
  • actions — an array of strings such as click element #sort-amount, set field #filter to overdue, wait for element #status to be visible, wait for element [aria-busy] to be removed, screen capture path.png;
  • runners — ['axe'], ['htmlcs'] or both;
  • threshold — the number of issues allowed before failing (keep at 0);
  • hideElements — elements to exclude (use sparingly, with a reason);
  • standard — for htmlcs, e.g. WCAG2AA.

pa11y drives a headless Chromium through Puppeteer. The same limits as any automated scanner apply: it finds structural issues, not behavioural ones.

One URL, five scanned states Flow of pa11y-ci scanning a single data page in five states: initial load, sorted, filtered to empty, dialog open, row expanded. One URL, five scanned statesInitialpage load, settledSortedclick header; waitEmptyfilter to no resultsDialogopen edit; waitExpandedexpand a row; wait
Each state is a separate entry with its own actions — five scans, five chances to catch state-specific failures.

When pa11y-ci fits — and when Playwright tests fit better

Permalink to "When pa11y-ci fits — and when Playwright tests fit better"

pa11y-ci fits when you want broad, low-effort coverage of many pages and a handful of states each, maintained as configuration rather than code. It suits teams without a large end-to-end test suite.

Playwright (or Cypress) tests with axe fit better when states need real logic — logging in, creating data, waiting on network responses, asserting announcements as well as scanning. If you already have end-to-end tests that reach the states, add scans there instead; see axe scans of grid states with Playwright.

The misapplication to name is hideElements: '.data-grid' to get a noisy grid out of the report. That removes the most important component from the scan entirely.

Annotated code example

Permalink to "Annotated code example"
// .pa11yci.js
module.exports = {
  defaults: {
    runners: ['axe', 'htmlcs'],           // broader coverage; expect some overlap
    standard: 'WCAG2AA',
    threshold: 0,                          // zero new issues
    timeout: 60000,
    chromeLaunchConfig: { args: ['--no-sandbox'] },
  },
  urls: [
    // 1. Initial state, after data has loaded
    { url: 'http://localhost:8080/invoices',
      actions: ['wait for element #invoice-grid [role="row"]:nth-child(2) to be visible'] },

    // 2. Sorted: aria-sort and the status message exist only now
    { url: 'http://localhost:8080/invoices',
      actions: [
        'wait for element #invoice-grid to be visible',
        'click element #invoice-grid th:nth-child(4) button',
        'wait for element #grid-status to be visible',
      ] },

    // 3. Empty state after filtering
    { url: 'http://localhost:8080/invoices',
      actions: [
        'set field #filter-q to zzz-no-match',
        'wait for element .empty-state to be visible',
      ] },

    // 4. Edit dialog open
    { url: 'http://localhost:8080/invoices',
      actions: [
        'click element [data-edit="INV-1042"]',
        'wait for element dialog[open] to be visible',
      ] },

    // 5. Treegrid row expanded (children loaded)
    { url: 'http://localhost:8080/accounts',
      actions: [
        'click element [data-row="finance"] .toggle',
        'wait for element [aria-level="2"] to be visible',
      ] },
  ],
};
# .github/workflows/a11y.yml (excerpt)
- run: npm run build && npx http-server dist -p 8080 &
- run: npx wait-on http://localhost:8080
- run: npx pa11y-ci --config .pa11yci.js --json > pa11y.json || (cat pa11y.json && exit 1)
- uses: actions/upload-artifact@v4
  with: { name: pa11y-report, path: pa11y.json }

Every state waits for a condition that proves the state exists — a status element, an open dialog, a level-2 row. Without the waits, pa11y scans the transition and reports issues that users never see, or misses the ones they do.

Keyboard & AT behaviour

Permalink to "Keyboard & AT behaviour"

pa11y does not test keyboard or screen reader behaviour. It can, however, reach states through keyboard actions, which doubles as a smoke check that those states are keyboard-reachable.

Action What it proves What it does not prove
click element … button The control exists and opens the state That it works with Enter or Space
set field … to … The input accepts text That results are announced
wait for element dialog[open] The dialog opened That focus moved into it
Scan after expansion Child rows have valid ARIA That expansion is announced
axe runner versus htmlcs runner Matrix comparing pa11y's axe and htmlcs runners on rule source, strengths for data UIs, noise level and maintenance. axe runner versus htmlcs runnerAspectaxe runnerhtmlcs runnerRule sourceaxe-coreHTML_CodeSnifferARIA validityStrongWeakerTable header checksGoodGood, stricterNoiseLowMore notices and warningsNeeds—Filter notices out
Running both finds more, at the cost of overlap to deduplicate.

Integration context

Permalink to "Integration context"

Rule decisions and suppressions should follow the same ledger as axe-core elsewhere — configuring axe-core rules and handling false positives — so pa11y and Playwright scans do not diverge. For a score-based signal across pages, Lighthouse CI accessibility budgets complements pa11y’s issue list.

States that pa11y actions cannot reach reliably — anything needing authentication flows, seeded data or network mocks — belong in Playwright tests.

Choosing the states to scan Steps for choosing which states of a data page to scan: list interactions, identify states that add markup, pick one representative per kind, and add a wait condition for each. Choosing the states to scanList interactionssort, filter, page, select, edit, expand, deleteper pageKeep states that change markuparia-sort, dialogs, errors, empty states, childrenskip pure re-rendersOne per kindone dialog, one empty state, one expansionenough for structureAdd a waita selector that exists only in the settled stateno fixed sleeps
Scan the states that add or change markup — that is where new violations come from.

Gotchas

Permalink to "Gotchas"

Fixed waits. wait for 2000 is flaky in CI. Always wait for an element or attribute that proves the state.

Stateful back ends. Actions that create or delete data change the next run’s starting point. Use a seeded, resettable environment.

htmlcs notices. HTML_CodeSniffer reports notices and warnings that are not failures; configure includeNotices: false and includeWarnings: false, or review them separately.

Design system notes

Permalink to "Design system notes"

The design system’s documentation site is an ideal pa11y target: each component example page, in each documented state, scanned on every change. Products then scan only their own compositions.

Testing checklist

Permalink to "Testing checklist"

FAQ

Permalink to "FAQ"
Can pa11y-ci test pages after interaction?

Yes. Each URL entry can have actions — click, set field, check, wait for element — that run before the scan, so you can scan sorted, filtered, dialog-open and expanded states.

Should I use pa11y's axe runner, htmlcs runner or both?

axe alone gives low noise and strong ARIA checks. Adding htmlcs finds some additional issues, especially around tables, at the cost of more warnings to filter and duplicates to ignore.

What threshold should pa11y-ci use?

Zero. Track any accepted known issues in a separate suppression list with reasons and expiry dates, so the threshold catches everything new.

How do I scan pages behind a login with pa11y-ci?

Use actions to fill in the login form and submit it before navigating to the page, or point pa11y at a seeded preview environment that does not need authentication. For complex flows, a Playwright test is usually easier.

Is pa11y-ci enough to test data UI accessibility?

No. It finds structural issues in the states you script. Keyboard behaviour, focus management and announcements need smoke tests and manual checks.

Permalink to "Related"

← Back to Automated Accessibility Testing Pipelines