9A. Contrast and Language Verbs
Permalink to 9A. Contrast and Language VerbsThis section is normative.
The contrast and lang_check verbs are delegating verbs (like audit): they
read DOM/CSSOM state and run a published assertion library against it. Both
delegate to @afixt/test-utils, which is an OPTIONAL peer dependency.
9A.1 contrast
Permalink to 9A.1 contrast- contrast: text "Login" # default WCAG AA
- contrast: button "Submit" level "AAA"
- contrast: region "Sidebar" # iterates text descendants internally
- contrast: self # inside scope.for_each
Semantics. The verb resolves the target to an element (§6), then walks the
target's subtree (inclusive) and runs testContrast(el, { level }) on every
descendant that has direct text content. The result is an aggregate record:
interface ContrastResult {
level: 'AA' | 'AAA';
tested: number;
passed: number;
failed: number;
skipped: number; // background image / mix-blend-mode / off-screen
}
The step status is fail if failed > 0, otherwise pass. Skipped elements do
NOT count toward failure but ARE reported in accessibility_notes, which are
score-affecting (§11). When every tested element was skipped the
step MUST report outcome: "inapplicable": nothing was measured.
Modifier. level "AA" (default) or level "AAA". Any other value MUST be a
parse error.
9A.2 lang_check
Permalink to 9A.2 lang_check- lang_check: page # walks <body>
- lang_check: region "Article"
- lang_check: self # inside scope.for_each
Semantics. The verb resolves the target (or the page's <body> for the
page form), reads its textContent plus the nearest declared lang (or
xml:lang) attribute on the element or any ancestor, and runs franc on the
text. The detected ISO 639-3 code is mapped to ISO 639-1 and compared to the
declared 2-letter code. The result is:
interface LangCheckResult {
tested: number; // 0 when text < 10 chars or no declared lang
passed: number;
failed: number;
}
The step status is pass when tested === 0 (insufficient signal — not a
failure) or when failed === 0. Otherwise fail. When tested === 0 the step
MUST also report outcome: "inapplicable" (§11): there was no
declared language to compare against, which is not the same as the declared
language being right.
9A.3 read_image
Permalink to 9A.3 read_image- read_image: image "Banner" # any non-empty text
- read_image: image "Banner" has_text "Sale 50%" # substring expectation
- read_image: image "Logo" matches "^AFix" # regex expectation
- read_image: image "Affiche" lang "fra" # custom Tesseract lang
- read_image: self has_text "Submit" # inside scope.for_each
Semantics. The verb resolves the target to a Playwright locator, takes a
screenshot of that element (locator.screenshot()), and hands the buffer to
getImageText from @afixt/test-utils, which wraps Tesseract.js. The extracted
text is then compared against an optional expectation:
interface ReadImageResult {
lang: string; // Tesseract lang pack used
extracted_text: string | null; // null when Tesseract found nothing
expectation: 'substring' | 'regex' | 'any';
expected: string | null; // pattern or substring; null for 'any'
matched: boolean;
}
The step status is pass if matched === true, otherwise fail. The default
expectation (any) passes whenever any non-empty text was extracted.
Modifiers.
| Modifier | Form | Meaning |
|---|---|---|
has_text |
has_text "..." |
Substring expectation. Pass iff extracted_text.includes(...). |
matches |
matches "regex" |
Regex expectation. Pass iff a RegExp constructed from ... matches. |
lang |
lang "eng+fra" |
Tesseract language pack(s). Default eng. |
has_text and matches are two spellings of one expectation, and a step
carrying both MUST be rejected at parse time (§7.5). A
processor that accepted both would have to pick one, which makes a step reading
as two expectations assert one, with nothing distinguishing that from both
holding. The error names the step and says which clause to keep.
Either expectation alone is valid, with or without lang:
read_image: image "Banner" has_text "Sale 50%"
read_image: image "Banner" matches "Sale [0-9]+%" lang "eng"
Both together is a parse error, not a step that checks two things:
read_image: image "Banner" has_text "Sale 50%" matches "Sale [0-9]+%"
Performance constraint. Each read_image invocation runs a full Tesseract
recognition pass (~1–3 seconds on small images). When used inside
scope.for_each the runner MUST run iterations sequentially regardless of any
future concurrency setting — Tesseract is CPU-heavy and parallel invocations
starve each other.
9A.4 sr_says
Permalink to 9A.4 sr_says- sr_says: '"Welcome to the dashboard"' # substring against full log
- sr_says: matches "(Welcome|Hello).*dashboard" # regex form
- sr_says: role "alert" "Email is required" # role-filtered substring
- sr_says: role "alert" matches "(Email|Name) is required"
- sr_says: '"3 results found" after activate button "Search"' # diff around step
- sr_says: role "alert" "Saved" after activate button "Save"
- sr_says: '"Loading complete" within 5s' # poll up to 5 seconds
- sr_says: '"Saved" after activate button "Save" within 2s' # combine after + within
Semantics. sr_says reads the spoken-phrase log of the active
screen-reader session and asserts that the expected substring (or regex)
appears. When role "X" is supplied the assertion is filtered to the phrases
that belong to that role.
How a role is expressed in the log is driver-specific, and a processor MUST
recognise the forms its driver produces rather than assume one. For the virtual
driver there are two, both observed against @guidepup/virtual-screen-reader:
| Situation | Spoken as | Example |
|---|---|---|
| A live region updates | <politeness>: <text> |
assertive: Email is required |
| An element is navigated to | <role>[, <name>[, <detail>]] |
button, Sign In |
A live announcement is labelled by politeness, not by role: an update to a
role="alert" element is spoken assertive: … and one to a role="status"
element polite: …. A processor filtering role "alert" MUST therefore treat
assertive: phrases as belonging to it, and polite: phrases as belonging to
role "status" and role "log". Matching on the role token alone excludes
exactly the phrases the assertion exists to check.
The reference runner never issues navigation commands, so in practice its log
holds live announcements. Where a container is navigated to, its role, its
text, and its end of <role> boundary are separate entries; an assertion
combining a role filter with text spoken inside such a container is not
supported.
The after <step> form. When supplied, the assertion is restricted to
phrases produced during or after the most-recent prior step matching <step>.
This lets a use-case author write "after I clicked Search the SR announced '3
results found'" without conflating it with earlier announcements. The
descriptor is a partial step — keyword [<target>]:
after activate button "Search"— match keyword + role + name (most specific)after activate button— match keyword + roleafter navigate— match keyword only
If no prior step matches, the verb fails with a descriptive error. When multiple prior steps match, the most recent one wins.
The within <duration> form. Optional and combinable with after. Accepted
units: s (seconds), ms (milliseconds). The processor MUST poll
spokenPhraseLog() at a small interval (the reference runner uses 50 ms) and
pass the assertion as soon as a match appears, or fail at the deadline. Without
within, the assertion is one-shot. The duration MUST be a positive integer;
within 0s and within Xunit for unrecognized units MUST be rejected at parse
time.
type SrDriver = 'virtual' | 'voiceover' | 'nvda';
interface SrSaysResult {
driver: SrDriver; // which driver produced the log
expectation: 'substring' | 'regex';
expected: string; // pattern or substring
role: string | null;
matched: boolean;
phrases: string[]; // trailing 50 phrases considered
after_step?: {
keyword: string;
role: string | null;
name: string | null;
} | null;
}
Driver selection. Configured via RunOptions.srDriver:
| Value | Behavior |
|---|---|
virtual |
Default. @guidepup/virtual-screen-reader, in-process, no OS deps. Works on every CI platform. |
auto |
Autodetect a running real SR (VoiceOver on macOS, NVDA on Windows); fall back to virtual. |
voiceover |
Force macOS VoiceOver via @guidepup/guidepup. Fail if unavailable. |
nvda |
Force Windows NVDA via @guidepup/guidepup. Fail if unavailable. |
The chosen driver MUST be recorded on every SrSaysResult so virtual-SR results
are never silently conflated with real-SR results.
Where the virtual session runs. The virtual driver reads a live DOM, so a
processor MUST construct its session inside the page, against the document
body. Constructed outside the page against a stand-in container the library does
not fail — it reports an empty log for every step, which presents as every
sr_says assertion failing rather than as a driver that is not working. A
processor SHOULD therefore have a test that asserts a non-empty log for a
page known to announce something; a test that only asserts the session
constructs will pass against a driver that reads nothing.
Because the session lives in the page, a navigation destroys it. A processor MUST still present one continuous log for the run, as the lifecycle constraint below requires: re-establishing the in-page session and carrying the phrases already spoken across the navigation satisfies this.
Scope of the log. The session covers the whole run, including the before
block (§4.6) — its entries use the same step language
and may themselves be sr_says steps. Phrases spoken during setup are therefore
in the log an sr_says step reads, and an assertion with no after clause
considers them along with everything else. Authors who mean "announced during
the behaviour under test, not during setup" should say so with an after <step>
clause rather than relying on the log starting empty.
Lifecycle constraint. A conforming runner MUST start the screen-reader session once per use-case run and reuse it across all iterations. Starting / stopping a real SR per element is seconds-scale; a 20-element iterated use case with per-element setup would take minutes.
Modifiers.
| Modifier | Form | Meaning |
|---|---|---|
phrase |
"..." |
Substring expectation. Synthesized when a bare quoted phrase is given. |
matches |
matches "regex" |
Regex expectation. |
after_keyword |
after <keyword> |
Required head of the after clause; resolves the prior step's keyword. |
after_role |
after <keyword> <role> |
Optional role component of the after clause's target. |
after_name |
after <keyword> <role> "<name>" |
Optional accessible-name component. |
phrase and matches are mutually exclusive. A processor MUST raise a parse
error when an sr_says step is missing both. The after_* modifiers are
populated only when the after <step> clause is present and are populated
together as a unit by the parser — authors write after <step> and the parser
splits it into the three keys.
9A.5 Library absence
Permalink to 9A.5 Library absenceIf @afixt/test-utils cannot be resolved, contrast and read_image MUST fail
with a descriptive error and SHOULD include installation guidance. Likewise for
lang_check if franc cannot be resolved (it is a transitive dependency of
@afixt/test-utils). For sr_says, @guidepup/virtual-screen-reader is
required for the default virtual driver and @guidepup/guidepup is required
for the real drivers; both are optional peer dependencies. Every such failure
MUST set failure_reason: "dependency_unavailable" on the step result
(§10.1): a runner without the library is an environment
problem, and a report that presents it as the page failing is wrong. Other steps
in the same use case MUST proceed according to continue_on_failure.
A driver that cannot start MUST fail the sr_says steps that need it and MUST
NOT abort the run. A single sr_says step taking down every use case queued
behind it — including cases that never mention a screen reader — is
non-conforming.
The diagnostic MUST distinguish a library that is absent from one that is present but not the shape the processor supports. Reporting the latter as the former sends a reader to reinstall a package that is already installed.