5. Step Language
Permalink to 5. Step LanguageThis section is normative.
5.1 Keywords
Permalink to 5.1 KeywordsThe Step Language defines the following keywords. Core keywords (1–8) are inherited from the manual use case methodology. Supplementary keywords (9–20) support real-world testing scenarios. Delegating keywords (21–25) read DOM/CSSOM state and delegate to companion engines.
| # | Keyword | Class | Has target? | Has modifiers? |
|---|---|---|---|---|
| 1 | locate |
core | yes | yes |
| 2 | focus |
core | yes | no |
| 3 | enter |
core | yes | yes (value) |
| 4 | select |
core | yes | yes |
| 5 | deselect |
core (extended) | yes | no |
| 6 | activate |
core | yes | yes |
| 7 | toggle |
core (extended) | yes | yes (attribute) |
| 8 | verify |
core | sub-type-dependent | yes |
| 9 | wait |
supplementary | no | duration only |
| 10 | wait_for |
supplementary | yes | no |
| 11 | navigate |
supplementary | no | url only |
| 12 | screenshot |
supplementary | no | name only |
| 13 | keyboard |
supplementary | no | keys only |
| 14 | type |
supplementary | no | text only |
| 15 | scroll |
supplementary | yes | no |
| 16 | hover |
supplementary | yes | no |
| 17 | hover_out |
supplementary | yes | no |
| 18 | note |
supplementary | no | text only |
| 19 | viewport |
supplementary | no | preset or <width>x<height> |
| 20 | anchor |
supplementary | no | label only |
| 21 | audit |
delegating | optional | level only |
| 22 | contrast |
delegating | yes | level (AA/AAA) |
| 23 | lang_check |
delegating | optional | no |
| 24 | read_image |
delegating | yes | has_text/matches/lang |
| 25 | sr_says |
delegating | optional (role) | phrase/matches, optional after / within |
A processor MUST reject any document containing a step with a keyword not in this table, and the rejection MUST name the running version and the keywords it supports, so a file that is merely newer than the processor is distinguishable from a typo.
A processor MUST also reject a step whose modifiers it does not recognise, rather than ignoring the unrecognised part. Silently discarding a clause turns a step that reads as an assertion into one that asserts nothing, and no diagnostic distinguishes that from the check passing.
Note.
deselectis provided in addition to the six core keywords from the manual methodology to allow expressing the inverse of aselecton a checkbox. It is a normative extension to the manual rubric.
5.2 Step grammar (overview)
Permalink to 5.2 Step grammar (overview)The general form is:
keyword: role_or_type "accessible_name" [modifiers]
The full ABNF grammar appears in Appendix A. This section describes each keyword in human-readable form.
5.3 Quoting
Permalink to 5.3 QuotingAccessible names and modifier values that may contain spaces MUST be enclosed in
straight double quotes ("). Inside a quoted string, the sequence \" denotes
a literal double quote and \\ denotes a literal backslash. A backslash before
any other character denotes both characters literally: "a\nb" is the four
characters a, \, n, b, not a newline. There are no other escapes.
A quoted string MAY contain any Unicode scalar value other than the double
quote, the backslash, and the control characters — C0 and DEL — (Appendix A,
qchar). Accessible names are routinely non-ASCII — button "Suche",
link "Café", heading "検索結果" — and a processor MUST accept them. The code
points are preserved as authored: a processor MUST NOT apply Unicode
normalisation, case folding, or whitespace collapsing to the string before
handing it to the locator. What the locator does with it is the matching rule in
§6.1.
Tokens that do not contain spaces (role tokens, modifier keys, integers, booleans) appear unquoted.
5.4 Role tokens
Permalink to 5.4 Role tokensThe following role tokens are defined and map to user-agent operations as
follows. The curated list covers the concrete ARIA 1.2 widget, composite widget,
document-structure, landmark, live-region, and window roles. Implementations
MUST use the documented Playwright API for every role token; using a CSS
selector or XPath as the locator strategy for one is non-conforming. The id
and data-* escape hatches (§5.4.2) are
the only tokens that resolve through a CSS locator, and they are defined as
escape hatches precisely because they do.
The accessible-name modifier is optional for every role that resolves to
getByRole(). When absent, the locator omits { name } and matches any element
with the given role. The field and text tokens still require a name —
getByLabel() and getByText() have no "any element" mode, and a step with one
of those tokens and an empty name MUST be rejected.
Only the getByRole() path consults the accessibility tree. Playwright's role
query excludes elements hidden from assistive technology — an aria-hidden
subtree, or an element the user agent does not expose — so a role token that
resolves is evidence the element is exposed. getByLabel() and getByText()
apply no such exclusion, and the id/data-* locators consult nothing but the
DOM: a field, text, id or data-* target inside an aria-hidden subtree
still resolves, and locate on it asserts CSS visibility only. Where exposure
is the point of the step, authors SHOULD target by role. See
§15.
| Token | Resolution |
|---|---|
link |
getByRole('link', { name? }) |
button |
getByRole('button', { name? }) |
field |
getByLabel(name) — name required |
text |
getByText(name) — name required |
heading |
getByRole('heading', { name? }), with optional level: N |
image |
getByRole('img', { name? }) |
dialog |
getByRole('dialog', { name? }) |
alertdialog |
getByRole('alertdialog', { name? }) |
checkbox |
getByRole('checkbox', { name? }) |
radio |
getByRole('radio', { name? }) |
switch |
getByRole('switch', { name? }) |
option |
getByRole('option', { name? }) |
select |
getByRole('combobox', { name? }), with fallback to getByLabel(name) (§6.3) |
combobox |
getByRole('combobox', { name? }) — ARIA combobox without the <select> fallback |
searchbox |
getByRole('searchbox', { name? }) |
textbox |
getByRole('textbox', { name? }) |
slider |
getByRole('slider', { name? }) |
spinbutton |
getByRole('spinbutton', { name? }) |
progressbar |
getByRole('progressbar', { name? }) |
meter |
getByRole('meter', { name? }) |
scrollbar |
getByRole('scrollbar', { name? }) |
separator |
getByRole('separator', { name? }) |
tab |
getByRole('tab', { name? }) |
tabpanel |
getByRole('tabpanel', { name? }) |
treeitem |
getByRole('treeitem', { name? }) |
menu |
getByRole('menu', { name? }) |
menubar |
getByRole('menubar', { name? }) |
menuitem |
getByRole('menuitem', { name? }) |
radiogroup |
getByRole('radiogroup', { name? }) |
tablist |
getByRole('tablist', { name? }) |
listbox |
getByRole('listbox', { name? }) |
tree |
getByRole('tree', { name? }) |
treegrid |
getByRole('treegrid', { name? }) |
grid |
getByRole('grid', { name? }) |
article |
getByRole('article', { name? }) |
feed |
getByRole('feed', { name? }) |
figure |
getByRole('figure', { name? }) |
group |
getByRole('group', { name? }) |
list |
getByRole('list', { name? }) |
listitem |
getByRole('listitem', { name? }) |
table |
getByRole('table', { name? }) |
row |
getByRole('row', { name? }) |
rowgroup |
getByRole('rowgroup', { name? }) |
cell |
getByRole('cell', { name? }) |
gridcell |
getByRole('gridcell', { name? }) |
columnheader |
getByRole('columnheader', { name? }) |
rowheader |
getByRole('rowheader', { name? }) |
toolbar |
getByRole('toolbar', { name? }) |
tooltip |
getByRole('tooltip', { name? }) |
banner |
getByRole('banner', { name? }) — landmark |
navigation |
getByRole('navigation', { name? }) — landmark |
main |
getByRole('main', { name? }) — landmark |
complementary |
getByRole('complementary', { name? }) — landmark |
contentinfo |
getByRole('contentinfo', { name? }) — landmark |
region |
getByRole('region', { name? }) — landmark |
search |
getByRole('search', { name? }) — landmark |
form |
getByRole('form', { name? }) — landmark |
alert |
getByRole('alert', { name? }) — also a verify sub-type (see §5.7) |
log |
getByRole('log', { name? }) |
status |
getByRole('status', { name? }) |
role |
getByRole(rolename, { name? }) — see §5.4.1 |
id |
locator('#' + name) — see §5.4.2 |
data-* |
locator('[data-*="<name>"]') — see §5.4.2 |
self |
the current iteration element — see §5.8 |
5.4.1 Custom ARIA roles
Permalink to 5.4.1 Custom ARIA rolesFor ARIA roles outside the curated list (e.g., an app-defined
custom-treegrid), the authoring surface is the two-token form:
- locate: role "custom-treegrid" name "File Browser"
- locate: role "custom-treegrid" # name optional — existence by role alone
5.4.2 ID and data attribute targets
Permalink to 5.4.2 ID and data attribute targetsTwo escape hatches exist for cases where a target genuinely cannot be identified by role and accessible name. They are normative but their use is discouraged because they bypass the accessibility-first principle.
- locate: id "main"
- locate: data-testid "checkout-cta"
A data-* token is any token beginning with data- and consisting of the
characters allowed in HTML attribute names. The processor MUST resolve id "X"
to page.locator('#X') and data-foo "Y" to page.locator('[data-foo="Y"]').
Use of these tokens SHOULD be flagged in reports. Processors MAY add a warning
to the step's accessibility_notes.
Resolving through one of these tokens establishes DOM presence, not exposure to
assistive technology (§5.4): the element may sit inside an
aria-hidden subtree and still resolve. A step targeted this way asserts what
its verb asserts about the DOM, and nothing about the accessibility tree. That
is what makes them escape hatches rather than alternative spellings.
5.5 Scoping with within
Permalink to 5.5 Scoping with withinAny role-bearing step MAY scope its target to a parent region using the within
modifier. The keyword inside is accepted as a synonym so iteration expressions
read naturally:
- locate: field "Email" within region "Sign Up"
- focus: button "Submit" within dialog "Confirm"
- locate: link inside navigation "Breadcrumb" # synonym for within
within (or inside) consumes a complete target descriptor (role + optional
name) and, recursively, its own optional within. The resolved locator is the
inner target located relative to the scope, semantically equivalent to
page.<scope>.<innerLocator>.
5.6 Keywords in detail
Permalink to 5.6 Keywords in detail5.6.1 locate
Permalink to 5.6.1 locatePurpose. Assert that the named target is present and visible in the accessibility tree.
Modifiers. level N (only for heading), within ….
Behavior. The processor MUST evaluate the element's visibility through
Playwright's toBeVisible() assertion (codegen) or
waitFor({ state: 'visible' }) (direct execution). A timeout MUST be reported
as a failure of locate.
5.6.2 focus
Permalink to 5.6.2 focusPurpose. Move keyboard focus to the target and verify the target actually received focus.
Behavior. The processor MUST call focus() on the resolved locator and
assert that the document's activeElement is now the target. If focus could not
be moved — the element is hidden, disabled, inert, or not focusable at all —
the step MUST fail with a message identifying the target.
What focus establishes. That the element can receive programmatic focus.
It does not establish that a keyboard user can reach it: an element with
tabindex="-1" accepts focus() and passes this step while being excluded from
the sequential tab order, so a workflow written as a sequence of focus steps
can pass against a page whose tab order is broken. Reachability is a separate
assertion, made by driving the keyboard and observing where focus lands:
- focus: field "Email"
- keyboard: Tab
- verify: focus field "Password"
The manual rubric's gain focus on verb is the programmatic step. Authors who
mean tab to SHOULD write the keyboard / verify: focus pair.
5.6.3 enter
Permalink to 5.6.3 enterPurpose. Type a value into a text input or textarea.
Required modifier. value "<text>".
Optional modifiers.
type_slowly true— usespressSequentially()instead offill(). Useful for input masks and autosuggest.clear_first false— appends instead of overwriting. Default behavior is to clear (becausefill()clears).
Both take a literal true or false; any other token is a parse error
(§7.5).
The cross-product is normative:
type_slowly |
clear_first |
Resolved operation |
|---|---|---|
| absent/false | absent/true | fill(value) |
| absent/false | false |
pressSequentially(value) — appends |
true |
absent/true | clear() then pressSequentially(value) |
true |
false |
pressSequentially(value) |
5.6.4 select
Permalink to 5.6.4 selectPurpose. Make an affirmative choice from a checkbox, radio, or combobox/select control.
Behavior. The operation depends on the target's role token:
| Target | Resolved operation |
|---|---|
checkbox, radio |
check(). Idempotent (§5.6.5). |
option |
click(), then assert aria-selected="true" on the option. An ARIA listbox option has no checked state for check() to drive. |
self |
selectOption(option) when an option modifier is present, otherwise check(): the iteration element is taken to be a native select or a checkable control respectively. |
any other role, with option "…" |
Resolve per §6.3 — by accessible name, without using the role token — then selectOption(option). Written for select and combobox; another role here names an element the fallback will not find. |
any other role, without option |
MUST fail with an error naming the role and the accepted forms. There is no meaningful way to "select a select", and a processor MUST NOT attempt check() on one, which can only time out. |
selectOption() operates on a native <select> element. A custom ARIA combobox
— an input or button with role="combobox" driving a listbox — is not driven
by it: Playwright rejects a non-<select> element and the step fails with that
error, not with an accessibility finding. UCDL defines no interaction sequence
for custom comboboxes. A case that exercises one composes it from the verbs the
pattern actually requires — activate or keyboard to open the popup, then
select: option "…" on the listbox option — which also asserts the widget's
states along the way.
Optional modifier (option targets only). force true — drive the click at
an option that is expected to refuse selection, e.g. one with
aria-disabled="true". The processor MUST bypass the actionability check
(click({ force: true })) and MUST suppress the post-click aria-selected
success assertion: an option that correctly refuses selection would fail that
assertion by definition, so the negative case asserts aria-selected itself
with verify. force true on any other select target MUST be rejected at
parse time — the other branches (check(), selectOption()) assert the
resulting state internally, so forcing them can never produce a passing negative
case; those cases use activate ... force true instead.
- select: option "Durian" force true
- verify: option "Durian" attribute "aria-selected" is "false"
5.6.5 deselect
Permalink to 5.6.5 deselectPurpose. Make a negative choice — uncheck a checkbox.
Behavior. The resolved operation is uncheck().
select and deselect are set-to-on and set-to-off, and both are
idempotent: they return immediately when the control is already in the requested
state. Neither asserts that anything changed, so neither can express "this
control flips state" — a case that calls select twice expecting a flip is
asserting nothing on the second call. Use toggle (§5.6.6) for that.
5.6.6 toggle
Permalink to 5.6.6 togglePurpose. Activate a control and assert its state actually inverted.
toggle: <target> [attribute "<name>"]
- toggle: checkbox "Subscribe" # was false -> MUST now be true
- toggle: checkbox "Subscribe" # was true -> MUST now be false
- toggle: button "Mute" attribute "aria-pressed"
Behavior. The processor MUST read the governing state, click() the target,
re-read the state, and fail if the value did not change. The failure MUST name
both the attribute and the value it was stuck at.
State resolution. When no attribute modifier is given, the governing
attribute is inferred from the role:
| Role | Attribute |
|---|---|
checkbox, switch, radio |
aria-checked, falling back to checked |
option, tab |
aria-selected |
treeitem |
aria-expanded |
button, menuitem, link |
whichever candidate is present (see below) |
A native <input type="checkbox"> exposes no aria-checked at all, so for the
native-backed roles the processor MUST fall back to the checked property;
reading the attribute alone would see null both times and report a working
control as stuck.
button is ambiguous — it carries aria-pressed as a toggle button and
aria-expanded as a disclosure — so the processor MUST use whichever candidate
is actually present on the element, and MUST fail with a message naming the
candidates when none is. An attribute "<name>" modifier overrides inference
entirely, which is the escape hatch for roles whose state attribute is unusual.
Change, not strict inversion. aria-checked="mixed" is valid for a
checkbox, and the APG's mixed -> false -> true cycle inverts nothing on a
single step, so the assertion is that the value changed. Two boolean values
that differ are an inversion anyway, so this still catches the defect the verb
exists for: a handler written setChecked(true) rather than
setChecked(!checked).
5.6.7 activate
Permalink to 5.6.7 activatePurpose. Activate a button or link.
Optional modifiers.
via keyboard— usespress('Enter')instead ofclick(). RECOMMENDED whenever the manual rubric specifies keyboard activation.keyboardis the only argumentviaaccepts; any other token, and an absent one, MUST be rejected at parse time (§7.5). Discarding it silently turns the one modifier that distinguishes pointer operability from keyboard operability into a mouse click that still reads as keyboard activation.new_tab true— expects activation to open a new tab; the processor MUST await the new page event and ensure it has loaded. As with every boolean-valued modifier, a value other than a literaltrue/falseis a parse error (§7.5).force true— bypasses the actionability check (click({ force: true })). Playwright counts anaria-disabledelement as not enabled, so a plainactivateat one times out on a page that is behaving exactly as a negative case requires;forcelets the case express "activating this disabled control does nothing" and assert the outcome withverify. The processor MUST NOT infer forcing from the element'saria-disabledstate — a case that clicks a disabled control by accident must keep failing rather than pass quietly.force truecombined withvia keyboardMUST be rejected at parse time:press()applies no enabled check, so there is nothing to force.
- activate: tab "Tab 3" force true
- verify: tab "Tab 3" attribute "aria-selected" is "false"
5.6.8 verify
Permalink to 5.6.8 verifyverify has 19 sub-types. Each sub-type's syntax and runtime semantics is
defined in §5.7.
5.6.9 wait
Permalink to 5.6.9 waitwait: <integer> — a fixed sleep, in milliseconds. Use SHOULD be rare; prefer
wait_for.
The duration is a whole number of milliseconds, optionally carrying an ms or
s suffix (Appendix A, duration): wait: 2000, wait: 2000ms and wait: 2s
are the same step. A processor MUST reject anything else — a decimal, a
negative, an unrecognised or uppercase suffix, or digits and letters mixed — as
a parse error naming the offending token. Reading the leading digits and
discarding the rest is specifically forbidden: wait: 2s then means two
milliseconds for a step that reads as two seconds, and a case built on it
passes or fails on timing luck with no diagnostic to distinguish the two.
The same spellings are accepted wherever a duration appears, including
sr_says ... within (§9A.4). Opposite conventions — one
requiring a unit, the other forbidding one — would be an accident of when each
was written rather than a distinction worth making, and would mean an author who
had read within 5s hit a parse error writing wait: 5s for a form the
language uses a line away.
5.6.10 wait_for
Permalink to 5.6.10 wait_forwait_for: <target> — block until the target is visible, with the configured
step timeout.
5.6.11 navigate
Permalink to 5.6.11 navigatenavigate: "<url>" — direct page navigation. The processor MUST emit
page.goto(url). The argument MAY contain template variables.
5.6.12 screenshot
Permalink to 5.6.12 screenshotscreenshot: "<name>" — capture a full-page screenshot to <name>.png, written
beneath the configured screenshot_dir
(§12.1) in the directory
§8.2 gives that step's failure screenshot:
<screenshot_dir>/<use-case-id>/, with before/ beneath it for an entry of the
before block and iter-<i>/ for a step in iteration i of an iterated case,
the use case id reduced as §8.2 reduces it. Both execution modes MUST write it
there (§8.3). A generated test takes the configured
directory as it was when the test was generated and, as direct execution does,
reads a relative one against the directory the run is started from. Namespacing
by use case and by iteration keeps two captures with one name — the same name in
two documents of a batch, or one step run once per element — from overwriting
each other.
<name> names one capture; it is not a path. A processor MUST reject a name
containing anything outside A–Z a–z 0–9 . _ -, and MUST reject . and ..,
as a parse error naming the accepted characters. The value may come from a
template variable, so accepting a path lets a data file choose where the run
writes; and a name resolved against the working directory rather than
screenshot_dir escapes the directory an implementation was asked to keep its
artifacts in.
5.6.13 keyboard
Permalink to 5.6.13 keyboardkeyboard: "<key-or-sequence>" — raw keyboard input. Whitespace separates
discrete key presses; modifier+key combinations use + (e.g., Shift+Tab).
- keyboard: Tab
- keyboard: Shift+Tab
- keyboard: ArrowDown ArrowDown Enter
Every token MUST name a key. A token that is not a key name is rejected at parse
time rather than compiled to a press() that fails at run time; use type: for
literal text.
5.6.14 type
Permalink to 5.6.14 typetype: "<text>" — type literal text into whatever currently holds keyboard
focus, one character at a time.
- focus: field "Title"
- type: Hello accessible world
type differs from enter in what it targets, not in what it does: enter
names a field and fills it, while type types into the focused element. Use
type when there is no nameable field to address — a contenteditable region,
a canvas-backed editor — or when per-keystroke handlers (live word counts,
markdown shortcuts, autocomplete) have to fire.
5.6.15 scroll
Permalink to 5.6.15 scrollscroll: to <target> — scroll the target into view. The leading to is
optional but RECOMMENDED.
5.6.16 hover
Permalink to 5.6.16 hoverhover: <target> — dispatch a pointer hover over the target. The resolved
operation is locator.hover().
hover is mouse-only and has no keyboard equivalent. It exists for patterns
that gate behavior on pointer hover (e.g., tooltip exposure, carousel
auto-rotation pause). Tests using hover are expected to fail under
keyboard-only profiles; that is intentional. Authors SHOULD pair hover with a
focus-driven companion step when WCAG 1.4.13 (Content on Hover or Focus) also
requires keyboard parity.
5.6.17 hover_out
Permalink to 5.6.17 hover_outhover_out: <target> — exit a pointer hover state on the target. Playwright has
no unhover primitive; the processor MUST dispatch mouseleave and mouseout
directly on the element rather than moving the mouse to an arbitrary coordinate.
5.6.18 note
Permalink to 5.6.18 notenote: "<text>" — informational only; no runtime action. In codegen, it MUST
emit a comment. In direct execution, it MUST be a no-op. Notes MAY be carried
into reports.
5.6.19 viewport
Permalink to 5.6.19 viewportviewport: <preset> or viewport: <width>x<height> — set the browser viewport
in CSS pixels for the remainder of the use case. Applies from the step onward,
the same way navigate does; there is no automatic restore.
| Preset | Size | Why |
|---|---|---|
mobile |
375x812 | Below the common sm/md breakpoint, where mobile navigation renders. |
tablet |
768x1024 | The usual tablet boundary. |
desktop |
1280x720 | Matches the viewport default in the generated config, so it restores the size a run started at. |
reflow |
320x256 | WCAG 2.2 SC 1.4.10: 320 CSS pixels wide (1280px at 400% zoom) for content that scrolls vertically, and 256 CSS pixels tall for content that scrolls horizontally. The preset applies both bounds at once. |
Setting the viewport evaluates nothing by itself. viewport: reflow puts the
page at the size where SC 1.4.10 binds; whether content reflows without
two-dimensional scrolling is what the audit step that follows checks.
A preset SHOULD be preferred to an explicit size. A size records a number; a
preset records the intent. Neither tracks a project's breakpoints — the presets
are fixed values — but when a project moves a breakpoint, viewport: mobile
still says what the step was for, whereas 375x812 silently stops crossing the
boundary it was chosen for and nothing says so. Explicit sizes exist for the
cases where a specific boundary is itself the subject.
The value MUST be a listed preset (matched case-insensitively) or two decimal
integers separated by x, with no spaces, each between 1 and 10000. Anything
else MUST be a parse error naming the accepted presets. A processor MUST NOT
fall back to a default size: a viewport step that quietly did nothing would
leave the journey at the previous size while reading as though it had changed
it, and the resulting failure would name a missing element rather than the
viewport.
Audits run at the current viewport. audit, contrast and the other
delegating keywords resolve against the page as it stands, so a viewport step
ahead of one is how the viewport-sensitive criteria are exercised — reflow
(1.4.10), target size (2.5.8) and orientation (1.3.4) are only meaningful at a
size where they bind:
- viewport: reflow
- audit: page
This is a runtime concern only; it does not change how a document is parsed.
5.6.20 anchor
Permalink to 5.6.20 anchoranchor: <label> — name a position in this case's flow, for an extension to
splice at with steps_override.from_anchor
(§4.3).
A label may use letters, digits, _ and -. Names MUST be unique within a use
case; a duplicate is a parse error rather than a silent last-one-wins, since a
child's from_anchor would otherwise splice at whichever happened to be last.
An anchor is not a step. A processor MUST strip anchors from the resolved step list and record them separately, so that they:
- never execute, never appear in a report, and never emit generated code; and
- do not occupy a step index.
The second property is what makes adoption safe: a parent can be annotated with
anchors while its children still splice by from_step, and none of them moves.
5.7 The verify keyword
Permalink to 5.7 The verify keywordverify dispatches to one of the sub-types listed below. The sub-type token
appears immediately after verify:.
5.7.1 Standalone sub-types (no target element)
Permalink to 5.7.1 Standalone sub-types (no target element)| Sub-type | Syntax | Runtime semantics |
|---|---|---|
url |
verify: url "<value>" |
If <value> starts with http:// or https://, exact match; otherwise, contains-match. |
title |
verify: title "<value>" |
Exact match against page.title(). |
download |
verify: download "<filename>" |
Consume the oldest download captured since the page was created that no earlier download verify has consumed, waiting up to the step timeout for one to arrive if none is buffered; assert suggestedFilename() === <filename>. |
live_region |
verify: live_region [role "<role>" | politeness "<level>"] "<text>" |
Poll every ARIA live region on the page — an element with role="status", role="alert", role="log", role="marquee" or role="timer", or an explicit aria-live other than off — until one contains <text> or the live-region timeout elapses (5 s in the reference implementation). With role "<role>", only regions of that role are examined; a region declared by aria-live alone has no role to name, so it is skipped. With politeness "polite" or politeness "assertive", the regions examined are those whose politeness matches — the roles that imply it (alert is assertive; status, log, marquee and timer are polite) together with elements declaring that value in aria-live. The two qualifiers are alternatives, not combinable. Every region is examined; none is preferred by role. |
iteration_summary |
verify: iteration_summary.<field> is <N> |
Terminal aggregate assertion after a scope.for_each block — see §5.8.4. |
Event capture. A processor MUST start capturing download events when the
page is created — before navigation to start_location, and before any step
runs — and MUST keep every download until a verify: download step consumes it.
Arming the listener inside the verify step is non-conforming: a download that
starts with the activating click and completes before the verify step is reached
would never be observed, and the step would time out on exactly the fast server
a CI run has. Each verify step consumes one download, in arrival order, so two
verify steps check two downloads. Both execution modes MUST capture identically
(§8.3); in code generation the capture is emitted ahead
of step 1 whenever the document contains a download verify.
A region declared by aria-live alone, on an element with no live-region role,
is a live region to every screen reader and is examined like any other; it has
no role for a role "<role>" qualifier to name, so a qualified assertion skips
it. Selecting by the aria-live attribute is not the CSS hook §15 forbids: the
attribute is the declaration that makes an element a live region.
5.7.2 Element-bearing sub-types
Permalink to 5.7.2 Element-bearing sub-types| Sub-type | Syntax | Runtime semantics |
|---|---|---|
text |
verify: text "<value>" |
Assert getByText(value) is visible. |
heading |
verify: heading "<value>" [within <scope>] [level N] |
Assert getByRole('heading', { name, [level] }) is visible, resolved inside <scope> when one is given. |
alert |
verify: alert "<value>" |
Assert at least one getByRole('alert') is visible whose text contains <value>. |
field_error |
verify: field_error "<fieldname>" or verify: field_error <role> ["<name>"] |
Resolve the field — through getByLabel(fieldname) in the quoted form, or as §6 resolves any target in the role form — and assert it has aria-invalid="true" AND has a non-empty aria-describedby or aria-errormessage. The element referenced by the first id in that list MUST also be visible. |
heading is the only sub-type in this table that always resolves a real
target: it is a role token, so verify: heading "Cart" and
verify: visible heading "Cart" describe the same element and both accept
within/inside scoping. The other three build a synthetic target from
their quoted string — text resolves through getByText, alert filters every
role="alert" by text, and field_error in its quoted form resolves through
getByLabel — and take no scope. A scope clause on those three MUST be rejected
as a parse error rather than ignored; a processor MAY add scoping support for
them in a later minor version, since accepting a clause that is currently an
error is backward-compatible.
field_error also takes a target in place of its quoted string: a role token
with an optional name, or any other form of role-and-name (Appendix A),
resolved as §6 resolves any target. The runtime
semantics are the same in both forms. The role form exists because a label's
text is not always the field's accessible name: a label that marks a required
field with an aria-hidden asterisk has text ending in * and a name that
does not, so getByLabel with the name finds nothing. Without it, the advice
in §5.4 — where exposure is the point of the step, target
by role — could not be followed for an error check. The role form takes no
scope either: a scope clause after it MUST be rejected as a parse error, and a
later minor version MAY accept one, as above.
verify: field_error textbox "Confirm Password"
Where a sub-type does accept scope, the clause binds to the target, not to the modifier list — see Appendix A. It therefore follows the target directly and precedes any modifier:
verify: heading "Cart" within main "" level 2
Putting the modifier first is a parse error, not a differently-ordered spelling of the same step:
verify: heading "Cart" level 2 within main ""
5.7.3 State sub-types (target follows the sub-type)
Permalink to 5.7.3 State sub-types (target follows the sub-type)The target's accessible name is optional for all state sub-types when the
role resolves to getByRole() (per §5.4). For count, this means
verify: count main is 1 is valid and counts every element with role="main"
regardless of name.
| Sub-type | Syntax | Runtime semantics |
|---|---|---|
visible |
verify: visible <target> |
toBeVisible() |
hidden |
verify: hidden <target> |
toBeHidden() |
enabled |
verify: enabled <target> |
toBeEnabled() / !isDisabled() |
disabled |
verify: disabled <target> |
toBeDisabled() / isDisabled() |
checked |
verify: checked <target> |
toBeChecked() |
unchecked |
verify: unchecked <target> |
not.toBeChecked() |
focus |
verify: focus <target> |
toBeFocused() |
count |
verify: count <target> is <integer> |
toHaveCount(integer) |
5.7.4 Composite sub-types (target plus modifier)
Permalink to 5.7.4 Composite sub-types (target plus modifier)| Sub-type | Syntax | Runtime semantics |
|---|---|---|
field_value |
verify: <target> has_value "<value>" |
toHaveValue(value) |
attribute |
verify: <target> attribute "<attr>" <predicate> |
See §5.7.5 |
A processor MUST raise a parse error for any verify: whose first token is not
a sub-type defined in this section, a role token (which produces field_value,
attribute, or visible), or id/data-*.
5.7.5 Attribute predicates
Permalink to 5.7.5 Attribute predicatesThe attribute sub-type supports eight predicate forms. A predicate keyword is
required: is "Y" is written out like the rest, and a bare value after the
attribute name is a parse error (§7.5, Appendix A
attr-predicate).
| Predicate | Runtime semantics |
|---|---|
is "<value>" |
Strict equality against getAttribute(attr). Equivalent to toHaveAttribute(attr, value). |
present |
Attribute exists with a non-empty value. Equivalent to toHaveAttribute(attr, /.+/). |
absent |
getAttribute(attr) returns null — the attribute is not on the element. Equivalent to not.toHaveAttribute(attr). |
is_or_absent "<value>" |
getAttribute(attr) is null or strictly equals <value>. An empty-string attribute is present, so it fails. |
starts_with "<prefix>" |
Attribute value begins with the literal prefix. Equivalent to toHaveAttribute(attr, /^<escaped-prefix>/). |
matches "<regex>" |
Attribute value matches the JS RegExp constructed from <regex>. Equivalent to toHaveAttribute(attr, new RegExp(regex)). |
references_existing_id |
The attribute is split on whitespace; every token MUST resolve to an existing element via [id="<token>"]. |
within_range_of "<min-attr>" "<max-attr>" |
All three attributes are read from the same target and parsed as numbers; assert min <= now <= max. |
present and absent are not complements. An attribute present with an empty
value — alt="" — satisfies neither: it is not absent, because it is on the
element, and not present, because the predicate requires a value. That third
state is asserted with is "", which is how the companion to C.5 — every
decorative image has an empty alt — is written. No predicate currently means
"on the element, whatever its value".
is_or_absent covers ARIA state attributes whose spec default is a real value
(aria-invalid, aria-disabled, aria-required, aria-readonly,
aria-busy), where omitting the attribute is a conformant way to express the
default. is MUST NOT resolve ARIA defaults: for attributes defaulting to
undefined (aria-expanded, aria-pressed, aria-checked, aria-selected,
aria-hidden) absence is distinct from "false" and MUST keep failing.
references_existing_id exists to verify the semantic relationship implied by
aria-labelledby, aria-describedby, aria-errormessage, aria-controls, and
similar attributes — the meaningful check is "does the reference resolve," not
"does the attribute hold this specific id literal."
within_range_of takes attribute names (not values) for the bounds; typical
use is aria-valuenow between aria-valuemin and aria-valuemax.
5.8 Iteration with scope.for_each and self
Permalink to 5.8 Iteration with scope.for_each and selfThis section is normative.
The base step language is flow-shaped: a use case is a singular workflow,
each step targets a singular element, and execution is sequential. Many
accessibility checks are audit-shaped — "for every element matching some
criteria, the following must hold". The scope.for_each field promotes a flow
case into an iteration over a collection of elements; self refers to the
current iteration element.
5.8.1 The scope object
Permalink to 5.8.1 The scope objectA use case MAY include a top-level scope field. When present:
| Key | Type | Required | Description |
|---|---|---|---|
for_each |
non-empty string | yes | Source expression for the iteration set. See §5.8.2. |
concurrency |
positive integer | no | Maximum number of pure-DOM iterations to run concurrently. Reserved for future use; conformant runners MAY ignore it. |
When scope is omitted, execution semantics are unchanged from prior versions
of this specification: steps run once, sequentially.
When scope.for_each is set, the runner MUST:
- Resolve
for_eachto an ordered list of zero or more iteration elements (Playwright locators). - Replay the case's
stepsonce per iteration element, with theselftoken bound to that element for the duration of the iteration. - Emit a per-iteration result and an aggregate summary (§10.1).
5.8.2 The for_each grammar
Permalink to 5.8.2 The for_each grammarfor_each accepts two forms:
Target form. A relaxed step-language target descriptor in which the accessible-name modifier is optional:
scope:
for_each: 'role "img"' # every element with ARIA role "img"
# for_each: image # equivalent shorthand for a curated role
# for_each: 'image "Logo"' # narrowed to images with that accessible name
# for_each: 'field "Email"' # all fields labeled "Email" (typically one)
The runner MUST enumerate the matching elements via Playwright's .all() API
(or the equivalent in the chosen automation engine).
The target expression MAY be scoped to a parent region using within (or its
synonym inside), so iteration enumerates only matching descendants of the
scope locator. Without scoping, "every link in the breadcrumb" iterates every
link on the page — the scoping form prevents that mis-attribution:
scope:
for_each: 'link within navigation "Breadcrumb"'
# for_each: 'link inside navigation "Breadcrumb"'
# for_each: 'role "option" within role "listbox" name "Tags"'
Criteria form. A reference to a consumer-supplied detector:
scope:
for_each: 'criteria headings_without_h1'
The runner MUST consult an injected resolver (in the reference implementation,
RunOptions.criteriaResolver) which returns either a CSS selector string or an
array of locators. If no resolver is configured the runner MUST fail the use
case with a clear error, not silently skip.
5.8.3 The self token
Permalink to 5.8.3 The self tokenself is a role token (§5.4) usable wherever a target
descriptor appears. It has no accessible-name argument; the iteration element
supplied by scope.for_each is the resolution.
scope:
for_each: 'role "img"'
steps:
- locate: self
- verify: self attribute "alt" is ""
- verify: visible self
- focus: self
A processor MUST reject a document in which self is referenced outside a
scope.for_each iteration — that is, one whose case has no scope. Whether a
document declares one is knowable without running it, so this is a validation
error rather than a runtime error, and the diagnostic MUST name the use case and
the offending step. self reached through a scope root
(locate: link "X" within self) is the same reference and MUST be rejected
alike.
5.8.4 Post-iteration aggregate verify
Permalink to 5.8.4 Post-iteration aggregate verifyverify: iteration_summary.<field> is <N> is a terminal assertion that runs
once after all iterations complete, against the aggregate counts the runner
emits as iteration_summary. Supported fields: failed, passed, total,
inapplicable.
A processor MUST:
- Partition
verify: iteration_summary.*steps out of the iteration body at execution time. They MUST NOT run per-element. - After
runIterated(or equivalent) completes, evaluate each such step against the computediteration_summaryand append itsStepResultto the case's flatstep_resultsso it participates in scoring. - Fail the step with a clear error when used in a use case that has no
scope.for_eachblock.
Code generation enforces these steps. The generated loop collects each
iteration's outcome into a summary of the same shape and asserts the aggregate
after the loop, so iteration_summary.total is 3 fails in a generated test
exactly where it fails in direct execution.
scope:
for_each: 'option within listbox "Tags"'
steps:
- select: self
- verify: self attribute "aria-selected" is "true"
- verify: iteration_summary.failed is 0 # terminal aggregate
5.8.5 Empty iteration sets
Permalink to 5.8.5 Empty iteration setsIf for_each resolves to zero elements, the runner MUST report the use case as
inapplicable (iteration_summary.inapplicable = 1, total = 0) rather than
as a failure, and the case scores inapplicable (§11) rather
than pass. Deleting every matching element from a page must not make a case
that sweeps them go green.
5.8.6 Failure handling within iterations
Permalink to 5.8.6 Failure handling within iterationscontinue_on_failure (§12.1) applies
within an iteration only. A failed step in iteration N aborts the
remainder of iteration N when continue_on_failure: false, but iteration
N+1 still runs. This guarantees that audit-shaped use cases gather diagnostic
information for every element.
The guarantee holds in both execution modes. A generated test runs each
iteration in its own try/catch, records the outcome, and continues, so it
reports every defective element rather than the first. The collected failures
are asserted after the loop, so the test still fails — with all of them named —
and the per-iteration results are attached to the Playwright report in the shape
§10.1 defines.
5.8.7 Per-iteration scoring
Permalink to 5.8.7 Per-iteration scoringAn iteration's status is fail if any of its step results is fail, otherwise
pass. The iteration_summary.passed and iteration_summary.failed counts are
populated from per-iteration statuses. The use case's overall score
(§11) is computed from the flat list of step results across
all iterations: any failed step makes the overall score fail.
5.8.8 Concurrency
Permalink to 5.8.8 ConcurrencyConformant runners MUST execute iterations sequentially by default. When
scope.concurrency > 1 is set, the runner MAY run iterations in parallel
batches of size concurrency, provided that the iteration body contains no
verb that mutates shared state. The mutating set comprises activate,
enter, focus, select, deselect, toggle, keyboard, type,
navigate, scroll, hover, hover_out, viewport and sr_says; an
iteration body containing any of these MUST be downgraded to sequential
execution regardless of the requested concurrency. hover, hover_out and
focus are in the set because a page has one pointer and one focus; viewport
because it resizes the page every iteration shares; sr_says because the
screen-reader log is one per run.
Read-only verbs (locate, wait_for, screenshot, note, wait, anchor,
all verify sub-types, audit, contrast, lang_check) observe DOM/ARIA
state without changing it and are safe to run in parallel. read_image is also
read-only but MUST run sequentially, for the reason given in
§9A.3: its recognition pass is CPU-bound and parallel passes
starve one another. A processor therefore treats it as a member of the
sequential set.
Each batch awaits all its members before the next begins, so the concurrency
value is a strict ceiling on simultaneous browser activity. Step numbers MUST
remain deterministic across iterations (numbered by
iteration_index * step_count + step_index + 1) so reports are reproducible
regardless of completion order.
The runner MUST set iteration_summary.concurrency_used to the effective
parallelism level — the requested concurrency value, or 1 when downgraded
because the body is state-mutating. This lets reports distinguish a parallel run
from a sequential one even when the YAML requested parallelism.