UCDL Accessibility Use Case Definition Language
Contents

Accessibility Use Case Definition Language (UCDL) and Runner

9A. Contrast and Language Verbs

This section is normative.

The contrast and lang_check verbs are delegating verbs (like audit): they read DOM/CSSOM state and run a published assertion library against it. Both delegate to @afixt/test-utils, which is an OPTIONAL peer dependency.

- contrast: text "Login" # default WCAG AA
- contrast: button "Submit" level "AAA"
- contrast: region "Sidebar" # iterates text descendants internally
- contrast: self # inside scope.for_each

Semantics. The verb resolves the target to an element (§6), then walks the target's subtree (inclusive) and runs testContrast(el, { level }) on every descendant that has direct text content. The result is an aggregate record:

interface ContrastResult {
  level: 'AA' | 'AAA';
  tested: number;
  passed: number;
  failed: number;
  skipped: number; // background image / mix-blend-mode / off-screen
}

The step status is fail if failed > 0, otherwise pass. Skipped elements do NOT count toward failure but ARE reported in accessibility_notes, which are score-affecting (§11). When every tested element was skipped the step MUST report outcome: "inapplicable": nothing was measured.

Modifier. level "AA" (default) or level "AAA". Any other value MUST be a parse error.

- lang_check: page # walks <body>
- lang_check: region "Article"
- lang_check: self # inside scope.for_each

Semantics. The verb resolves the target (or the page's <body> for the page form), reads its textContent plus the nearest declared lang (or xml:lang) attribute on the element or any ancestor, and runs franc on the text. The detected ISO 639-3 code is mapped to ISO 639-1 and compared to the declared 2-letter code. The result is:

interface LangCheckResult {
  tested: number; // 0 when text < 10 chars or no declared lang
  passed: number;
  failed: number;
}

The step status is pass when tested === 0 (insufficient signal — not a failure) or when failed === 0. Otherwise fail. When tested === 0 the step MUST also report outcome: "inapplicable" (§11): there was no declared language to compare against, which is not the same as the declared language being right.

- read_image: image "Banner" # any non-empty text
- read_image: image "Banner" has_text "Sale 50%" # substring expectation
- read_image: image "Logo" matches "^AFix" # regex expectation
- read_image: image "Affiche" lang "fra" # custom Tesseract lang
- read_image: self has_text "Submit" # inside scope.for_each

Semantics. The verb resolves the target to a Playwright locator, takes a screenshot of that element (locator.screenshot()), and hands the buffer to getImageText from @afixt/test-utils, which wraps Tesseract.js. The extracted text is then compared against an optional expectation:

interface ReadImageResult {
  lang: string; // Tesseract lang pack used
  extracted_text: string | null; // null when Tesseract found nothing
  expectation: 'substring' | 'regex' | 'any';
  expected: string | null; // pattern or substring; null for 'any'
  matched: boolean;
}

The step status is pass if matched === true, otherwise fail. The default expectation (any) passes whenever any non-empty text was extracted.

Modifiers.

Modifier Form Meaning
has_text has_text "..." Substring expectation. Pass iff extracted_text.includes(...).
matches matches "regex" Regex expectation. Pass iff a RegExp constructed from ... matches.
lang lang "eng+fra" Tesseract language pack(s). Default eng.

has_text and matches are two spellings of one expectation, and a step carrying both MUST be rejected at parse time (§7.5). A processor that accepted both would have to pick one, which makes a step reading as two expectations assert one, with nothing distinguishing that from both holding. The error names the step and says which clause to keep.

Either expectation alone is valid, with or without lang:

read_image: image "Banner" has_text "Sale 50%"
read_image: image "Banner" matches "Sale [0-9]+%" lang "eng"

Both together is a parse error, not a step that checks two things:

read_image: image "Banner" has_text "Sale 50%" matches "Sale [0-9]+%"

Performance constraint. Each read_image invocation runs a full Tesseract recognition pass (~1–3 seconds on small images). When used inside scope.for_each the runner MUST run iterations sequentially regardless of any future concurrency setting — Tesseract is CPU-heavy and parallel invocations starve each other.

- sr_says: '"Welcome to the dashboard"' # substring against full log
- sr_says: matches "(Welcome|Hello).*dashboard" # regex form
- sr_says: role "alert" "Email is required" # role-filtered substring
- sr_says: role "alert" matches "(Email|Name) is required"
- sr_says: '"3 results found" after activate button "Search"' # diff around step
- sr_says: role "alert" "Saved" after activate button "Save"
- sr_says: '"Loading complete" within 5s' # poll up to 5 seconds
- sr_says: '"Saved" after activate button "Save" within 2s' # combine after + within

Semantics. sr_says reads the spoken-phrase log of the active screen-reader session and asserts that the expected substring (or regex) appears. When role "X" is supplied the assertion is filtered to the phrases that belong to that role.

How a role is expressed in the log is driver-specific, and a processor MUST recognise the forms its driver produces rather than assume one. For the virtual driver there are two, both observed against @guidepup/virtual-screen-reader:

Situation Spoken as Example
A live region updates <politeness>: <text> assertive: Email is required
An element is navigated to <role>[, <name>[, <detail>]] button, Sign In

A live announcement is labelled by politeness, not by role: an update to a role="alert" element is spoken assertive: … and one to a role="status" element polite: …. A processor filtering role "alert" MUST therefore treat assertive: phrases as belonging to it, and polite: phrases as belonging to role "status" and role "log". Matching on the role token alone excludes exactly the phrases the assertion exists to check.

The reference runner never issues navigation commands, so in practice its log holds live announcements. Where a container is navigated to, its role, its text, and its end of <role> boundary are separate entries; an assertion combining a role filter with text spoken inside such a container is not supported.

The after <step> form. When supplied, the assertion is restricted to phrases produced during or after the most-recent prior step matching <step>. This lets a use-case author write "after I clicked Search the SR announced '3 results found'" without conflating it with earlier announcements. The descriptor is a partial step — keyword [<target>]:

  • after activate button "Search" — match keyword + role + name (most specific)
  • after activate button — match keyword + role
  • after navigate — match keyword only

If no prior step matches, the verb fails with a descriptive error. When multiple prior steps match, the most recent one wins.

The within <duration> form. Optional and combinable with after. Accepted units: s (seconds), ms (milliseconds). The processor MUST poll spokenPhraseLog() at a small interval (the reference runner uses 50 ms) and pass the assertion as soon as a match appears, or fail at the deadline. Without within, the assertion is one-shot. The duration MUST be a positive integer; within 0s and within Xunit for unrecognized units MUST be rejected at parse time.

type SrDriver = 'virtual' | 'voiceover' | 'nvda';

interface SrSaysResult {
  driver: SrDriver; // which driver produced the log
  expectation: 'substring' | 'regex';
  expected: string; // pattern or substring
  role: string | null;
  matched: boolean;
  phrases: string[]; // trailing 50 phrases considered
  after_step?: {
    keyword: string;
    role: string | null;
    name: string | null;
  } | null;
}

Driver selection. Configured via RunOptions.srDriver:

Value Behavior
virtual Default. @guidepup/virtual-screen-reader, in-process, no OS deps. Works on every CI platform.
auto Autodetect a running real SR (VoiceOver on macOS, NVDA on Windows); fall back to virtual.
voiceover Force macOS VoiceOver via @guidepup/guidepup. Fail if unavailable.
nvda Force Windows NVDA via @guidepup/guidepup. Fail if unavailable.

The chosen driver MUST be recorded on every SrSaysResult so virtual-SR results are never silently conflated with real-SR results.

Where the virtual session runs. The virtual driver reads a live DOM, so a processor MUST construct its session inside the page, against the document body. Constructed outside the page against a stand-in container the library does not fail — it reports an empty log for every step, which presents as every sr_says assertion failing rather than as a driver that is not working. A processor SHOULD therefore have a test that asserts a non-empty log for a page known to announce something; a test that only asserts the session constructs will pass against a driver that reads nothing.

Because the session lives in the page, a navigation destroys it. A processor MUST still present one continuous log for the run, as the lifecycle constraint below requires: re-establishing the in-page session and carrying the phrases already spoken across the navigation satisfies this.

Scope of the log. The session covers the whole run, including the before block (§4.6) — its entries use the same step language and may themselves be sr_says steps. Phrases spoken during setup are therefore in the log an sr_says step reads, and an assertion with no after clause considers them along with everything else. Authors who mean "announced during the behaviour under test, not during setup" should say so with an after <step> clause rather than relying on the log starting empty.

Lifecycle constraint. A conforming runner MUST start the screen-reader session once per use-case run and reuse it across all iterations. Starting / stopping a real SR per element is seconds-scale; a 20-element iterated use case with per-element setup would take minutes.

Modifiers.

Modifier Form Meaning
phrase "..." Substring expectation. Synthesized when a bare quoted phrase is given.
matches matches "regex" Regex expectation.
after_keyword after <keyword> Required head of the after clause; resolves the prior step's keyword.
after_role after <keyword> <role> Optional role component of the after clause's target.
after_name after <keyword> <role> "<name>" Optional accessible-name component.

phrase and matches are mutually exclusive. A processor MUST raise a parse error when an sr_says step is missing both. The after_* modifiers are populated only when the after <step> clause is present and are populated together as a unit by the parser — authors write after <step> and the parser splits it into the three keys.

If @afixt/test-utils cannot be resolved, contrast and read_image MUST fail with a descriptive error and SHOULD include installation guidance. Likewise for lang_check if franc cannot be resolved (it is a transitive dependency of @afixt/test-utils). For sr_says, @guidepup/virtual-screen-reader is required for the default virtual driver and @guidepup/guidepup is required for the real drivers; both are optional peer dependencies. Every such failure MUST set failure_reason: "dependency_unavailable" on the step result (§10.1): a runner without the library is an environment problem, and a report that presents it as the page failing is wrong. Other steps in the same use case MUST proceed according to continue_on_failure.

A driver that cannot start MUST fail the sr_says steps that need it and MUST NOT abort the run. A single sr_says step taking down every use case queued behind it — including cases that never mention a screen reader — is non-conforming.

The diagnostic MUST distinguish a library that is absent from one that is present but not the shape the processor supports. Reporting the latter as the former sends a reader to reinstall a package that is already installed.