Generate for CI, Run for Right Now
Part 9 of "Testing the Way Users Actually Experience the Web" — a tutorial
series on @afixt/usecase-runner.
So far every example has used usecase-runner run, which executes a use case
directly against a live browser. That's one of two execution modes. This post
covers both, when to use which, and how a run is scored.
Mode 1: Code generation
Permalink to Mode 1: Code generationnpx usecase-runner generate ./usecases --outdir ./tests/generated
For each .uc.yaml you get a standalone Playwright .spec.ts. Here's the start
of one, generated from the login case in Part 2:
// tests/generated/login-success.spec.ts
// Auto-generated by @afixt/usecase-runner from login-success.uc.yaml
// Do not edit manually — changes will be overwritten on next generation.
import { test, expect } from '@playwright/test';
test.describe('Log In To The System', () => {
test.describe.configure({ mode: 'serial' });
test('login-success: Log In To The System (positive)', async ({ page }) => {
// Preconditions:
// - User has a valid username and password
// - User is not already logged in
// Step 1: Access start location
await page.goto('https://www.example.com');
// Step 2: Locate link "Client Sign In"
await expect(
page.getByRole('link', { name: 'Client Sign In' }),
).toBeVisible();
// Step 3: Gain focus on link "Client Sign In"
const step3_el = page.getByRole('link', { name: 'Client Sign In' });
await step3_el.focus();
await expect(step3_el).toBeFocused();
// Step 4: Activate link "Client Sign In"
await page.getByRole('link', { name: 'Client Sign In' }).click();
...
Three things to notice.
It's readable. Every step is a comment in the original DSL followed by the
Playwright it expands to. A developer who has never seen a .uc.yaml can read
the failure in their CI output, see
// Step 3: Gain focus on link "Client Sign In", and understand both the intent
and the mechanism. The preconditions are carried in as comments for the same
reason.
It's not meant to be edited. The header says so. The YAML is the source of
truth; the spec file is a build artifact. Commit the generated files if your CI
setup needs them present, but treat a hand edit the way you'd treat a hand edit
to dist/. If the generated code is wrong, the fix is in the YAML or in the
generator.
It's ordinary Playwright. Run it with npx playwright test, in any
Playwright config, with any reporter, on any browser Playwright supports. The
generated file doesn't import usecase-runner at runtime (delegating verbs like
audit import the engine they delegate to, and nothing else). This is the mode
for CI: it fits into whatever you already have.
generate --run does both steps in one go.
Mode 2: Direct execution
Permalink to Mode 2: Direct executionnpx usecase-runner run ./usecases/login-success.uc.yaml --headed --browser chromium
No intermediate files. The runner parses the YAML and drives a Playwright Page
in-process, step by step. This is the mode for:
-
Writing a new case and iterating on it until it passes.
-
Ad-hoc audits: point a case at a staging URL and see what breaks.
-
Overriding data on the command line without touching the file:
npx usecase-runner run login-success.uc.yaml --set [email protected] -
Producing the HTML/DOCX deliverable reports directly.
The modes must agree
Permalink to The modes must agreeA rule in the spec (§8.3) says the two modes must produce the same verdict for
the same case. It sounds obvious; it's also the reason several features exist
only when both sides implement them — interaction profiles (Part 7) are not
runner-only, and the parser strictness (Part 5) was rejected in the form "just
make validate fail" precisely because the direct runner would have stayed
silent. If run says pass and generate says fail, one of them is lying, and
you can't tell which.
Scoring: four outcomes, not two
Permalink to Scoring: four outcomes, not twoA use case doesn't just pass or fail. It gets one of four scores — the three the manual rubric uses, plus one for a run that never got started:
| Score | Meaning |
|---|---|
| Pass | Every step passed, with no skips and no notes attached. |
| Pass w/ Conditions | Every step completed, but at least one was skipped or carried advisory accessibility_notes — for example an audit that found only low-priority issues. |
| Fail | At least one step failed. |
| Error | A step in the before block failed, so the steps under test never ran. |
The middle category matters. A manual tester's report is full of "passed, but note that…" comments, and those notes are often where the most useful findings are. The automated score preserves that channel instead of collapsing everything into green or red.
Error matters for a different reason, and it is the one CI meets. A case can
declare a before block — setup the case depends on but is not itself
asserting, typically a log-in:
before:
- 'enter: field "Email" value "{{ username }}"'
- 'enter: field "Password" value "{{ password }}"'
- 'activate: button "Sign in"'
If any of it fails, the steps under test do not run and the score is Error
rather than Fail. The distinction is worth knowing: Fail means the suite
looked and found a defect; Error means it never got far enough to look. Both
exit non-zero, so neither slips through a pipeline, but only one of them is
evidence about the page.
Keep going after a failure
Permalink to Keep going after a failureBy default, continue_on_failure is true. When step 4 fails, steps 5 through
14 still run. Some of them will fail as a consequence (you can't verify the
dashboard if login didn't happen) and the report says so; some will pass and
tell you something independent.
This mirrors manual practice — a tester doesn't close the laptop at the first
defect — and it's the difference between one finding per run and all of them. A
run with six failures is usually one root cause plus five consequences, and the
report's step order makes that obvious. Set continue_on_failure: false only
when later steps would be destructive or meaningless after a failure.
The report
Permalink to The reportnpx usecase-runner run ./usecases --report json,html
JSON is the machine-readable form: per-step status, duration, error text,
accessibility_notes, and a screenshot path captured automatically on failure.
HTML mirrors the two-column step/comments table from the manual deliverable.
(DOCX is reserved in the specification and not yet emitted.) Both carry the same
tester string — usecase-runner / Playwright / chromium — and an
environment record with the processor, Playwright, browser, OS and Node
versions, so a result always says how it was produced.
Which mode, when
Permalink to Which mode, when- Writing or debugging a case →
run --headed. - Every commit →
generateinto your Playwright suite. - Client deliverable →
run --report html. - Checking whether a flow is pointer-dependent → either, with
--profile no-pointer; they'll agree.