# Work

Source: https://flow-next.dev/skills/work/

Execute a flow-next spec - one task inline, several in parallel lanes, reviewed by risk, evidence-gated completion.

Direct execution is the default for a ready spec. `/flow-next:work <spec-id> --no-plan` gives one execution owner the full spec through the configured gates. A task graph is the exception, chosen on a positive signal (you asked for a plan, separate human owners, or delivery staged across several PRs). [Start without a plan](https://flow-next.dev/skills/work/#start-without-a-plan) explains consent and review.

A spec with one task (the direct route’s owner, a one-task plan, or a task-id run) is implemented inline by the conversation: no worker and no worktree unless you ask for one (`--worktree`). A worker takes a single task only when you named an implementer model or tier, the task’s review mode resolves to `host`, or the conversation’s context is too full, and the run says which reason applied.

Before building a large spec inline, start from a fresh session (or `/clear`) so the build has room, or name an implementer model and a fresh-context worker takes the task.

With several tasks, `/flow-next:work` schedules them on the **rolling frontier** by default. The moment any in-flight worker returns, the next eligible task is admitted, in its own isolated workspace, and the conductor integrates, reviews, and completes each task as it lands. It falls back to the wave loop only when the spec’s shape gives rolling nothing to schedule. Every dispatched task gets a fresh-context worker and the same review, evidence, and completion gates on either route.

## Worker subagent model

On a multi-task spec, each task runs in its own subagent. The benefits compound:

* Fresh context per task prevents bleed between unrelated changes.
* Re-anchor information stays bundled with its implementation.
* Review cycles stay isolated to the task they belong to.
* The main conversation stays lean even on long specs.

```mermaid
flowchart TB
  Main["Main session"] --> Pick["Admit ready tasks (cap 3)"]
  Pick --> Spawn["Dispatch worker per task"]
  Spawn --> Anchor["Worker re-reads spec"]
  Anchor --> Impl["Implement + commit in isolated workspace"]
  Impl --> Ret["Worker returns"]
  Ret --> Join["Integrate that task"]
  Join --> Review["Per-task review (conductor-owned)"]
  Review --> Done["Mark task done"]
  Done --> Pick
```

The re-anchor is a single `flowctl anchor <task-id>` call: one deterministic bundle carrying the task, the parent spec, git state, memory and glossary indices, and dependency summaries, so each worker starts in one round-trip instead of eight. Each section is the verbatim output of the command it names. Since 6.1.0 those commands are the narrow forms: the memory index as text, only the glossary entries the task’s title or description names (`glossary list --match`), `git status --short --branch`, and the spec record without its review and tracker ledgers. A typical bundle dropped from about 160 KB to about 60 KB, and both bundles scored 7 of 7 on the comprehension eval across three task sets. The bundle is a floor, not a ceiling; the worker still reads anything else it needs.

**A worker never returns while a command it started is still running.** It launches its gates in the foreground, and when the host moves a long command to the background it waits for that command to exit and reads the exit code before it reports. A result therefore always describes a finished run. The conductor checks the same thing from its side: before it accepts a return, it reads the task status and looks for a live command in that worker’s workspace. If it finds one, it waits within the dispatch timebox, then sends a re-anchoring continuation worker into the same workspace once the command exits. The early return is not counted as a failed attempt. To look at another tree state a worker uses a temporary worktree, never `git stash`.

## Scheduling: rolling frontier by default

Two schedulers exist inside work, and the route is decided once, from spec state, at the start of the task loop:

| Route                 | When                                                                                                                                                                                                                      | What happens                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Rolling** (default) | Everything else                                                                                                                                                                                                           | A ready task is admitted at every worker-return event (same five fail-closed conditions as a wave: same spec, cap 3, no dependency path, disjoint declared paths, nothing in the always-serial set - judged against the in-flight set). Each worker gets an isolated worktree; the conductor integrates, reviews, and completes each task as it returns; a shared outside-tree notes surface lets workers leave findings for siblings, read by pointer. |
| **Wave** (fallback)   | The run was given a task id (only that task runs); `planSync.enabled` is not `false`; the spec has fewer than two open tasks (a no-plan implicit task, or one task left); or no two open tasks are dependency-independent | Concurrent safe subsets joined at wave boundaries, or a single worker in the checkout. Plan-sync runs after each joined wave.                                                                                                                                                                                                                                                                                                                           |

The route prints once, before the first task is claimed, as `Scheduling: rolling` or `Scheduling: wave (<reason>)`. There is no config knob and no host table. A host whose subagent dispatch is measured to block reports `Scheduling: degraded to wave (host lacks non-blocking dispatch)` and keeps the rest of the rolling lifecycle. `flow --auto` and land dispatch plain `/flow-next:work` and inherit the route.

**Why rolling, and why isolated worktrees.** A pre-registered three-arm eval compared architectures on the same specs. Rolling admission with a worktree per task saved **52.1% of work-phase wall clock at quality parity** with zero uncontained correctness incidents. A shared-checkout arm was faster still (69.2%) and **failed quality parity**: part of its speed came from roughly 30% thinner test artifacts, because making the declared-paths list the commit boundary disincentivized new test files. The per-task worktree pool removes that pressure, so speed is never bought with under-testing. The scheduler shipped first as the experimental `/flow-next:work-rolling` beta, ran end-to-end on Claude Code, Cursor, and Grok Build, and then graduated into work as the default; the beta command is gone.

**Playing nice with other runs.** Task claims are spec-scoped: two runs on the same spec contend on the same claims and fail closed against each other - a task claimed by another actor is dropped from the run’s admissible set, never stolen. Since 5.5.0 a task you hold yourself is no different at the claim: a second `flowctl start` as the same actor on an `in_progress` task refuses with a typed error instead of reading as a crash resume, and a genuine resume passes `--reclaim` after confirming the prior run ended. A merge conflict at per-task integration retries that task serially: one re-run, never a correctness loss.

## Quick commands: what gets verified, and how often

The spec’s `## Quick commands` block names the focused checks for the code a task touches. A worker (several tasks) runs them once before its first edit, to prove the tree was green *before* the change, and again before it may mark the task done; a single task built inline runs its failing test first and focused tests after. The full gates run only when the repository’s instructions or you ask for a full suite: once, at the end, and not again after later fixes. CI owns regressions.

The scaffold’s convention is that per-task commands name focused suites for the files that task touches. Repeats are cheap by construction. A green receipt keyed to the exact commit and the exact command string skips a re-run at unchanged HEAD, and a docs-only diff drops to lint and format only.

**Code in sibling repos.** In a home-base workspace, where `.flow/` lives in a planning repo and the product code in sibling clones, work records the base of each sibling the spec changes (your project instructions or the spec name them) and classifies each repo’s diff too. The docs-only tier applies only when every repo’s diff is docs-only, and a sibling it cannot read runs its full gates. The feature-map update reads the sibling diffs as well, so a route moved in a sibling updates the map. A single-repo workspace works exactly as before.

Which means the tier is yours to set. Author focused commands and keep the full-suite entrypoint as the final-gate command if you want narrow per-task runs; list the full suite if your project would rather pay it every task; put the policy in `CLAUDE.md` / `AGENTS.md` if it should hold across specs. Selection is authored, never computed - it keeps needless full runs out of the loop, but it will not infer which distant suite your change actually touches. See [the verification spine](https://flow-next.dev/understand/how-work-gets-proven/).

## Start without a plan

Direct execution through `/flow-next:work <spec-id> --no-plan` is the default for a ready spec. Plan is chosen only on a positive signal: you asked for a plan, separate human owners will implement, or delivery is staged across several PRs. Risk, size, and file count never trigger plan on their own. Design risk routes to plan-review, which reviews a spec with zero tasks; unresolved product or authority choices route to refine. Explicit `--no-plan` or stated intent answers the zero-task route choice; without either, work’s ask on a zero-task spec reads the same plan-versus-no-plan rule the [flow](https://flow-next.dev/skills/flow/) skill owns and recommends its result. On an already-planned spec, the flag is ignored with a one-line notice.

```bash
/flow-next:work fn-N --no-plan
```

### Idea text and single tasks

`/flow-next:work "rename the config key"` accepts idea text and mints the minimal spec and task itself, so a spec always exists underneath and the same evidence gates apply. `/flow-next:work fn-N.M` instead selects one existing planned task.

An explicit request to review the design, or a recorded design review that still needs work or human judgment, stops direct work before owner creation or resume. Run `/flow-next:plan-review` for the spec, resolve its findings, then invoke work again. Work’s `--review` option selects implementation review and does not clear this gate.

The strong spec remains the implementation contract. One implicit execution owner covers every acceptance criterion and follows the same configured implementation review, coverage, completion policy, and evidence gates. Optional live QA stays opt-in and augments staging and manual QA; a `NEEDS_WORK` QA finding can accompany the PR and does not authorize merge. Work records the explicit direct-route choice before creating its single execution owner, so a resumed run preserves the route. `flow --auto` and work can resume that same actor’s sole owner only after confirming the prior run ended, including when work receives a task ID; active or unknown ownership remains `NEEDS_HUMAN`. The owner implements the full spec, with execution-time decomposition or justified delegation as needed. `flowctl spec set-no-plan fn-N` or capture with `--no-plan` records the same choice for `flow --auto`, and a `flow --auto` run that meets a ready zero-task spec with no recorded route applies the plan-versus-no-plan rule and records the result on the spec before any mint. Work takes consent only from the flag, stated intent, or the field. Without one of them, work stops on a zero-task spec. See [How Flow chooses](https://flow-next.dev/choosing-your-route/#no-plan-route).

## Fixing a reported defect

When a task fixes a reported defect, whether flow routed a bug report here or the spec’s requirement is that reported behaviour stops happening, the worker runs four steps before and around the fix: it checks for prior fixes, reproduces the symptom reliably and confirms its cause with runtime evidence, bisects when a known-good revision exists, and proves the fix by running the same reproduction on the base (failing) and the head (passing). A cheap reproduction test is committed failing before the fix. An existing fix, or a fix someone visibly owns, stops the task rather than getting a competing one. Other open pull requests that touch the area are noted in the done summary, and the task continues. The done summary records each step, including the ones skipped and why, and make-pr turns that record into proof cells. [How Flow chooses](https://flow-next.dev/choosing-your-route/#bug-or-defect) describes the steps in full.

## Climbing a metric toward a target

When a spec asks for one number moved toward a target through repeated attempts, the worker runs a measured loop instead of a single change. It sets up the experiment itself, using your numbers where you gave them, and writes the setup at the top of the ledger before the first attempt; only a spec with no target stops it. It builds a harness, proves it can tell better from worse and rejects a wrong output, freezes it with a hash, then tries one hypothesis per attempt against the current best. An attempt is kept, as exactly one commit, only when it beats the best by more than the noise and the regression gate is green; anything else is reverted to a clean tree. Every attempt writes a ledger row, and the run stops on the target plus the attempt floor, on the budget, when the supported ideas run out, or when the harness breaks. The target is yours and is never relaxed. Review runs once over the kept commits, and the done summary carries the full ledger. [How Flow chooses](https://flow-next.dev/choosing-your-route/#hill-climb) walks through the loop and a worked example.

## Branch modes

| Mode       | Use                                                                                                                    |
| ---------- | ---------------------------------------------------------------------------------------------------------------------- |
| `current`  | Stay on the active branch. The default when you are not on the default branch.                                         |
| `new`      | Create a fresh feature branch off the current base. The default on the default branch, and always under `flow --auto`. |
| `worktree` | Fully isolated parallel work via `flow-next-worktree-kit`.                                                             |

Without a branch option, work asks nothing: it takes the default and says which in one line. A worktree is used only when you ask for one.

Worktree mode is the right choice when more than one spec is in flight, or when a review cycle on a separate branch must not disturb the work in progress.

### Dependent specs branch from the parent

A spec whose parent (`depends_on_epics`) has every task done and its branch on origin starts now instead of after the parent’s PR merges. Work asks `flowctl spec chain <id>` once, fetches the parent’s branch, and creates the spec branch from `origin/<parent-branch>`; the recorded spec base is the merge-base with that tip, so gate classification and the quality auditor see only this spec’s own diff. `--branch=current` on such a spec requires the current branch to already contain the parent tip, otherwise the run stops with `BLOCKED` naming the missing ancestry; worktree mode applies the same base. A spec the predicate refuses (parent branch not pushed, two open parents, a sibling already chained on the same parent, or a failed remote query) stops before any task starts with `BLOCKED: <reason>`. Make-pr later opens the PR against the parent’s branch and, on GitHub, links it into the parent’s stack: [chains and stacks](https://flow-next.dev/skills/make-pr/#dependent-specs-chains-and-stacks). A spec with no open parent branches exactly as before.

## Plan-sync after each joined wave

When `planSync.enabled` is true, work takes the wave route (plan-sync’s per-wave barrier is a fail-closed rule) and downstream task specs are checked for stale references after a serial task or after the conductor has joined and resolved an entire parallel wave. This avoids syncing against a partial integration. If a task changed an API or path that later tasks assumed, the drift surfaces with a reason. The user decides whether to update the downstream specs, regenerate them, or accept the drift explicitly.

Plan-sync is conservative. It never silently rewrites a spec; it only proposes.

## Quality audit (large or risky changes)

In the quality phase - after tasks complete, alongside the full gates - the conductor may dispatch the [`quality-auditor`](https://flow-next.dev/subagents/execution/#quality-auditor) subagent over the spec’s whole diff. The trigger is a judgment call, not a flag. It runs when the conductor judges the change large or risky, and is skipped for small routine changes. It dispatches as two parallel axis-scoped runs (correctness and standards), both reports come back verbatim, and only correctness-axis Criticals block.

## Feature-map update

When the repository has a feature map at `.flow/features/`, the quality phase also checks whether this spec’s change altered how a user reaches a mapped feature. If it did, work edits only those feature files, proves each new route with one live drive, refreshes their `**Last proven:**` lines, and retires any drift notes naming a route it proved; the map diff rides the same commit as the code. If the app cannot start, or a changed surface cannot be tied to one feature file, work leaves the map unchanged and files a drift note instead, without blocking the run. It never adds new features and never runs the full [`/flow-next:features`](https://flow-next.dev/skills/features/) pass. Without a map the step costs one existence check. See [Keep the feature map current](https://flow-next.dev/guides/keep-feature-map-current/).

## Completion review gate

When all tasks are done, an optional `/flow-next:spec-completion-review` runs to verify the combined implementation matches the spec end-to-end. This is the place to catch criteria that no individual task fully owned.

## Optional review

```bash
/flow-next:work <spec-id> --review=codex
```

`--review=codex|rp|copilot|cursor|claude|host|none` runs a per-task adversarial review, sized by risk: three reviewers from another model family for a change that touches persisted or shared state, concurrency, security, data layout or migrations, or a multi-file feature; one reviewer otherwise; none for a small local output, wording or display fix, with the skip recorded. Attended, a `NEEDS_WORK` gets one fix pass and one re-review scoped to the fixes. Under `flow --auto` the fix and scoped re-review repeat until SHIP before the task is marked done. See [Impl Review](https://flow-next.dev/skills/impl-review/#review-by-risk-then-one-scoped-re-review).

Without `--review`, work reads the configured backend for each task from the repository root, so a task’s own `review:` pin wins over the project default. Work never asks about review. With no backend configured, review is off and the handoff says so once (“no review backend set; run setup or set review\.backend”); set one with [`/flow-next:setup`](https://flow-next.dev/skills/setup/), `flowctl config set review.backend <backend>`, or in the prompt.

## Implementation offload

Handing the token-heavy part - writing code - to a second CLI agent is a **routing decision you write, not a subsystem you configure**. Name an `implementer` tier in your `CLAUDE.md` / `AGENTS.md` [routing block](https://flow-next.dev/guides/model-routing/#the-routing-block), or just say it in the moment; the host drives the other CLI through a headless bridge for the draft. With no implementer tier, the session model implements - that is the default and needs no configuration.

Two rules are not optional:

* **The bridged child owns the task; the worker keeps judgment.** The child writes code, commits checkpoints on the branch it was given, and decides its own delegation under the same license an in-host worker holds. It never pushes, never rebases or rewrites history, never decides scope, never issues a review verdict, and never spawns a bridge of its own. The worker hands the task over (Phase 1b), reviews the child’s commit range from the recorded base, runs the gates, dispatches review, and owns `flowctl done`; the conductor never bridges. Your done summary records which model implemented and how many subagents the child dispatched.
* **Send clear, well-scoped tasks to the value tier.** On well-specified work a value tier matches a strong tier on correctness for meaningfully less wall clock; escalate to the strong tier only for genuinely gnarly ones. Spec quality is what makes the trade safe - a vague brief burns the saving on rework.

Bridge recipes ship into your repo and are read on demand with `flowctl usage` (`## Orchestration & model steering`). Full doctrine: [Orchestration → Implementation offload](https://flow-next.dev/guides/model-routing/#implementation-offload-the-bridge-route).

The packaged delegation mode and its `work.delegate*` config keys were removed in 4.0.0. Leftover keys in `.flow/config.json` are inert - flowctl names them once in a non-blocking advisory and otherwise ignores them. Run [`/flow-next:setup`](https://flow-next.dev/skills/setup/), accept the routing-block scaffold, and use the bridge recipes above.

## Worked example

```plaintext
/flow-next:work fn-12-export-json-flag
```

```text
Scheduling: rolling
ready frontier: [.1, .2, .3]   in-flight: []   admitted: [.1, .2]   held: [.3: depends on .1]
  .2 -> returned: commit e4f5a6b, tests green; integrated; review SHIP; done
  admitted: [.3]
  .1 -> returned: commit a1b2c3d, tests green; integrated; review SHIP; done
  .3 -> returned: commit c7d8e9f, tests green; integrated; review SHIP; done
quiesce: full suite green; completion review: SHIP - receipt on disk
Spec fn-12-export-json-flag: all tasks done.
```

Fresh context per task plus mandatory evidence is what makes the loop repeatable instead of lucky.

* `flowctl done` refuses completion without evidence JSON - if a worker “finished” without commits and test commands, it did not finish.
* The planner reports execution waves as a hint. The work conductor admits from the live ready frontier per return event, with isolation, integration order, and worker count its own call; it holds a task when the fail-closed conditions are not sound.
* Atomic claims prevent duplicate ownership. They do not make simultaneous edits to one Git index or filesystem safe.
* Use `--branch=worktree` to isolate a spec’s changes in `.worktrees/<name>/` without touching your current branch.

## Dynamic usage

Recipes that compose with work in the [cookbook](https://flow-next.dev/guides/cookbook/):

* [Parallelize](https://flow-next.dev/guides/cookbook/#parallelize) - both supported forms of task-level parallelism, with the guardrails that keep them safe.
* [Model routing](https://flow-next.dev/guides/cookbook/#model-routing) - an `implementer` tier hands implementation to a second CLI while the host keeps git, review, and judgment.
* [Skip & lighten](https://flow-next.dev/guides/cookbook/#skip--lighten) - plan + work alone is a complete, evidence-backed workflow for small changes.

## Next step

```bash
/flow-next:make-pr <spec-id>
```
