# QA

Source: https://flow-next.dev/skills/qa/

Live-app real-user QA pass that drives the running app and derives its test scenarios from the spec.

`/flow-next:qa` runs a live-app, real-user QA pass. It drives the **running** app like an unforgiving customer, derives its test scenarios **directly from the spec**, files structured findings with evidence, and ends with a YES/NO ship verdict.

It is the one review surface that isn’t static. `impl-review`, `spec-completion-review`, `quality-auditor`, and `code-review` all read code or specs; `/flow-next:qa` exercises the deployed app. It is **forbidden from marking PASS by reading source** - a scenario passes only on captured evidence (screenshot / console / URL), never on inspection.

`/flow-next:qa` is the **cheap first live pass** - the app already runs on the dev’s machine during `/flow-next:work`, so this drives it like a real user before a human ever opens the PR, catching obvious runtime breakage early. Like everything in Flow-Next, it **reduces human work agentically and surfaces problems to humans** - it does **not** stand in for CI/staging QA or manual QA, which still happen downstream. Its findings are advisory. They ride the PR as open items, which opens it as a draft, and the bug-memory track, and the human reviewer + the [land](https://flow-next.dev/autonomy/land/) gate decide. Use it to shrink the QA humans have to do, not to remove them from the loop.

The doctrine, the spec-as-intent advantage, and the driver ladder are in [Live-app QA & driving UIs](https://flow-next.dev/guides/live-qa/).

## How it drives the app

QA never re-implements driving - it consumes [Flow-Next Drive](https://flow-next.dev/skills/flow-next-drive/)’s surface-aware driver ladder (agent-browser → chrome-devtools-mcp → Playwright → cursor-ide-browser → manual; native surfaces go Cua Driver → Computer Use, or Cua Sandbox headless). Whatever rung Flow-Next Drive resolves for the surface is what QA inherits.

## Findings

Failures are filed immediately as structured **P0 / P1 / P2** reports - persona, steps to reproduce, expected vs actual, and evidence (console, screenshots, full URL). Findings feed the **bug memory track** (`track: bug`) with overlap dedup, and can be promoted to flow specs or tasks for the fix.

## The verdict

The pass ends with a verdict carried as a proof-of-work receipt (`type: qa_verdict`), with four outcomes:

| `qa_outcome` | Meaning                                                     | Ship? |
| ------------ | ----------------------------------------------------------- | ----- |
| `SHIP`       | All scenarios pass, zero open P0/P1, R-ID coverage complete | Yes   |
| `NEEDS_WORK` | Any open P0/P1, or incomplete coverage                      | No    |
| `BLOCKED`    | No live deploy or no driver - could not verify              | No    |
| `NA`         | Spec has no driveable user-visible AC                       | n/a   |

The receipt’s `verdict` field projects onto the existing review-receipt enum (`BLOCKED → NEEDS_WORK`, `NA → SHIP`), so the verdict can feed [Spec Completion Review](https://flow-next.dev/skills/spec-completion-review/) - “does the *live app* satisfy the AC, not just the code?”

## Requirements

QA needs a **live deploy + a driver** ([Flow-Next Drive](https://flow-next.dev/skills/flow-next-drive/)). With neither, it surfaces a `BLOCKED` verdict rather than failing; a spec with no driveable UI yields a clean `NA`. Opt-in - it adds nothing to the base flow when unused.

## As a pipeline stage

`pipeline.qa` is a string enum `off | on | auto`; any other value is `off`.

| Value           | What runs                                                                                                                                                                                                                                                                                                                                                                                                                                                                  |
| --------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `off` (default) | QA runs only when you invoke `/flow-next:qa <spec>`.                                                                                                                                                                                                                                                                                                                                                                                                                       |
| `on`            | One live pass at all-tasks-done, before make-pr, on every spec, under attended flow and `flow --auto` alike.                                                                                                                                                                                                                                                                                                                                                               |
| `auto`          | [`/flow-next:flow`](https://flow-next.dev/skills/flow/) runs the pass when the spec’s acceptance describes UI behaviour on a drivable surface and a target can be started (a documented start command, a deploy URL, or `.flow/features/`). Otherwise the stage records `skipped(config: pipeline.qa=auto: <reason>)` and the route advances. `flow --auto` reads the same gate-selection reference, so the skip line and the advance to make-pr are identical unattended. |

Enable it with `flowctl config set pipeline.qa auto` (or `on`). Setup asks the question once on an unset key, and separately recommends [`/flow-next:features`](https://flow-next.dev/skills/features/) whatever the answer, since the live pass reads `.flow/features/` to navigate the app. QA reads `.flow/features/` itself to find the feature before driving. QA never hard-blocks a loop. `NEEDS_WORK` and `BLOCKED` advance to make-pr, and their findings become open items on a draft PR.

## Credit

The QA discipline - the P0/P1/P2 taxonomy, evidence rules, and session-hygiene practices - is a lean borrow from Ray Fernando’s [running-bug-review-board](https://github.com/RayFernando1337/rayfernando-skills) (Apache-2.0). Thank you, Ray.

## Worked example

```plaintext
/flow-next:qa fn-12-export-json-flag
```

```text
Deriving scenarios from spec: R1-R3 -> 4 scenarios (2 runtime, 1 CLI contract, 1 error path)
Driving live app via flow-next-drive...
  R1 export --json emits valid JSON        PASS (captured stdout, parsed)
  R2 null-owner rows serialize             FAIL - P1: row dropped (captured output attached)
qa_outcome: NEEDS_WORK (1 P1) - findings filed; PR still advances with a Live QA section.
Receipt: qa_verdict (head_sha e4f5a6b, rid_coverage 3/3, open_p0p1: 1)
```

The verdict rests on captured evidence from the running app - the skill is forbidden from passing by reading source.

* Keep the app running from your work session; QA drives what you already have up rather than owning a boot procedure.
* QA augments CI and staging, never replaces them - it catches the spec-level “does it actually behave” gaps automated suites were not written for.
* Deterministic checks the worker already re-ran are subtracted; runtime, UI, and integration criteria are always exercised live.

## Dynamic usage

Recipes that compose with qa in the [cookbook](https://flow-next.dev/guides/cookbook/):

* [Evidence-first](https://flow-next.dev/guides/cookbook/#evidence-first) - the `qa_verdict` receipt is idempotent per branch head; your scripts can key on it.
* [Autonomy dial](https://flow-next.dev/guides/cookbook/#autonomy-dial) - set `pipeline.qa` to `on` and `flow --auto` runs the live pass before every PR; `auto` lets attended flow and `flow --auto` decide per spec from the gate-selection reference.

## Next step

```bash
/flow-next:make-pr <spec-id>
```
