Skip to content

QA

/flow-next:qa runs a live-app, real-user QA pass. It drives the running app like an unforgiving customer, derives its test scenarios directly from the spec, files structured findings with evidence, and ends with a YES/NO ship verdict.

It is the one review surface that isn’t static. impl-review, spec-completion-review, quality-auditor, and code-review all read code or specs; /flow-next:qa exercises the deployed app. It is forbidden from marking PASS by reading source - a scenario passes only on captured evidence (screenshot / console / URL), never on inspection.

The doctrine, the spec-as-intent advantage, and the driver ladder are in Live-app QA & driving UIs.

QA never re-implements driving - it consumes Flow-Next Drive’s surface-aware driver ladder (agent-browser → chrome-devtools-mcp → Playwright → cursor-ide-browser → manual; native surfaces go Cua Driver → Computer Use, or Cua Sandbox headless). Whatever rung Flow-Next Drive resolves for the surface is what QA inherits.

Failures are filed immediately as structured P0 / P1 / P2 reports - persona, steps to reproduce, expected vs actual, and evidence (console, screenshots, full URL). Findings feed the bug memory track (track: bug) with overlap dedup, and can be promoted to flow specs or tasks for the fix.

The pass ends with a verdict carried as a proof-of-work receipt (type: qa_verdict), with four outcomes:

qa_outcomeMeaningShip?
SHIPAll scenarios pass, zero open P0/P1, R-ID coverage completeYes
NEEDS_WORKAny open P0/P1, or incomplete coverageNo
BLOCKEDNo live deploy or no driver - could not verifyNo
NASpec has no driveable user-visible ACn/a

The receipt’s verdict field projects onto the existing review-receipt enum (BLOCKED → NEEDS_WORK, NA → SHIP), so the verdict can feed Spec Completion Review - “does the live app satisfy the AC, not just the code?”

QA needs a live deploy + a driver (Flow-Next Drive). With neither, it surfaces a BLOCKED verdict rather than failing; a spec with no driveable UI yields a clean NA. Opt-in - it adds nothing to the base flow when unused.

pipeline.qa is a string enum off | on | auto; any other value is off.

ValueWhat runs
off (default)QA runs only when you invoke /flow-next:qa <spec>.
onOne live pass at all-tasks-done, before make-pr, on every spec, under attended flow and flow --auto alike.
auto/flow-next:flow runs the pass when the spec’s acceptance describes UI behaviour on a drivable surface and a target can be started (a documented start command, a deploy URL, or .flow/features/). Otherwise the stage records skipped(config: pipeline.qa=auto: <reason>) and the route advances. flow --auto reads the same gate-selection reference, so the skip line and the advance to make-pr are identical unattended.

Enable it with flowctl config set pipeline.qa auto (or on). Setup asks the question once on an unset key, and separately recommends /flow-next:features whatever the answer, since the live pass reads .flow/features/ to navigate the app. QA reads .flow/features/ itself to find the feature before driving. QA never hard-blocks a loop. NEEDS_WORK and BLOCKED advance to make-pr, and their findings become open items on a draft PR.

The QA discipline - the P0/P1/P2 taxonomy, evidence rules, and session-hygiene practices - is a lean borrow from Ray Fernando’s running-bug-review-board (Apache-2.0). Thank you, Ray.

/flow-next:qa fn-12-export-json-flag
Deriving scenarios from spec: R1-R3 -> 4 scenarios (2 runtime, 1 CLI contract, 1 error path)
Driving live app via flow-next-drive...
R1 export --json emits valid JSON PASS (captured stdout, parsed)
R2 null-owner rows serialize FAIL - P1: row dropped (captured output attached)
qa_outcome: NEEDS_WORK (1 P1) - findings filed; PR still advances with a Live QA section.
Receipt: qa_verdict (head_sha e4f5a6b, rid_coverage 3/3, open_p0p1: 1)

The verdict rests on captured evidence from the running app - the skill is forbidden from passing by reading source.

Recipes that compose with qa in the cookbook:

  • Evidence-first - the qa_verdict receipt is idempotent per branch head; your scripts can key on it.
  • Autonomy dial - set pipeline.qa to on and flow --auto runs the live pass before every PR; auto lets attended flow and flow --auto decide per spec from the gate-selection reference.
Terminal window
/flow-next:make-pr <spec-id>