UCDL Accessibility Use Case Definition Language
Contents

Generate for CI, Run for Right Now

Part 9 of "Testing the Way Users Actually Experience the Web" — a tutorial series on @afixt/usecase-runner.

So far every example has used usecase-runner run, which executes a use case directly against a live browser. That's one of two execution modes. This post covers both, when to use which, and how a run is scored.

npx usecase-runner generate ./usecases --outdir ./tests/generated

For each .uc.yaml you get a standalone Playwright .spec.ts. Here's the start of one, generated from the login case in Part 2:

// tests/generated/login-success.spec.ts
// Auto-generated by @afixt/usecase-runner from login-success.uc.yaml
// Do not edit manually — changes will be overwritten on next generation.

import { test, expect } from '@playwright/test';

test.describe('Log In To The System', () => {
  test.describe.configure({ mode: 'serial' });

  test('login-success: Log In To The System (positive)', async ({ page }) => {
    // Preconditions:
    //   - User has a valid username and password
    //   - User is not already logged in

    // Step 1: Access start location
    await page.goto('https://www.example.com');

    // Step 2: Locate link "Client Sign In"
    await expect(
      page.getByRole('link', { name: 'Client Sign In' }),
    ).toBeVisible();

    // Step 3: Gain focus on link "Client Sign In"
    const step3_el = page.getByRole('link', { name: 'Client Sign In' });
    await step3_el.focus();
    await expect(step3_el).toBeFocused();

    // Step 4: Activate link "Client Sign In"
    await page.getByRole('link', { name: 'Client Sign In' }).click();
    ...

Three things to notice.

It's readable. Every step is a comment in the original DSL followed by the Playwright it expands to. A developer who has never seen a .uc.yaml can read the failure in their CI output, see // Step 3: Gain focus on link "Client Sign In", and understand both the intent and the mechanism. The preconditions are carried in as comments for the same reason.

It's not meant to be edited. The header says so. The YAML is the source of truth; the spec file is a build artifact. Commit the generated files if your CI setup needs them present, but treat a hand edit the way you'd treat a hand edit to dist/. If the generated code is wrong, the fix is in the YAML or in the generator.

It's ordinary Playwright. Run it with npx playwright test, in any Playwright config, with any reporter, on any browser Playwright supports. The generated file doesn't import usecase-runner at runtime (delegating verbs like audit import the engine they delegate to, and nothing else). This is the mode for CI: it fits into whatever you already have.

generate --run does both steps in one go.

npx usecase-runner run ./usecases/login-success.uc.yaml --headed --browser chromium

No intermediate files. The runner parses the YAML and drives a Playwright Page in-process, step by step. This is the mode for:

  • Writing a new case and iterating on it until it passes.

  • Ad-hoc audits: point a case at a staging URL and see what breaks.

  • Overriding data on the command line without touching the file:

    npx usecase-runner run login-success.uc.yaml --set [email protected]
    
  • Producing the HTML/DOCX deliverable reports directly.

A rule in the spec (§8.3) says the two modes must produce the same verdict for the same case. It sounds obvious; it's also the reason several features exist only when both sides implement them — interaction profiles (Part 7) are not runner-only, and the parser strictness (Part 5) was rejected in the form "just make validate fail" precisely because the direct runner would have stayed silent. If run says pass and generate says fail, one of them is lying, and you can't tell which.

Scoring: four outcomes, not two

A use case doesn't just pass or fail. It gets one of four scores — the three the manual rubric uses, plus one for a run that never got started:

Score Meaning
Pass Every step passed, with no skips and no notes attached.
Pass w/ Conditions Every step completed, but at least one was skipped or carried advisory accessibility_notes — for example an audit that found only low-priority issues.
Fail At least one step failed.
Error A step in the before block failed, so the steps under test never ran.

The middle category matters. A manual tester's report is full of "passed, but note that…" comments, and those notes are often where the most useful findings are. The automated score preserves that channel instead of collapsing everything into green or red.

Error matters for a different reason, and it is the one CI meets. A case can declare a before block — setup the case depends on but is not itself asserting, typically a log-in:

before:
  - 'enter: field "Email" value "{{ username }}"'
  - 'enter: field "Password" value "{{ password }}"'
  - 'activate: button "Sign in"'

If any of it fails, the steps under test do not run and the score is Error rather than Fail. The distinction is worth knowing: Fail means the suite looked and found a defect; Error means it never got far enough to look. Both exit non-zero, so neither slips through a pipeline, but only one of them is evidence about the page.

By default, continue_on_failure is true. When step 4 fails, steps 5 through 14 still run. Some of them will fail as a consequence (you can't verify the dashboard if login didn't happen) and the report says so; some will pass and tell you something independent.

This mirrors manual practice — a tester doesn't close the laptop at the first defect — and it's the difference between one finding per run and all of them. A run with six failures is usually one root cause plus five consequences, and the report's step order makes that obvious. Set continue_on_failure: false only when later steps would be destructive or meaningless after a failure.

npx usecase-runner run ./usecases --report json,html

JSON is the machine-readable form: per-step status, duration, error text, accessibility_notes, and a screenshot path captured automatically on failure. HTML mirrors the two-column step/comments table from the manual deliverable. (DOCX is reserved in the specification and not yet emitted.) Both carry the same tester string — usecase-runner / Playwright / chromium — and an environment record with the processor, Playwright, browser, OS and Node versions, so a result always says how it was produced.

  • Writing or debugging a case → run --headed.
  • Every commit → generate into your Playwright suite.
  • Client deliverable → run --report html.
  • Checking whether a flow is pointer-dependent → either, with --profile no-pointer; they'll agree.

Next: Beyond the Step: Delegating to Audit Engines