Work
Direct execution is the default for a ready spec. /flow-next:work <spec-id> --no-plan gives one execution owner the full spec through the configured gates. A task graph is the exception, chosen on a positive signal (you asked for a plan, separate human owners, or delivery staged across several PRs). Start without a plan explains consent and review.
A spec with one task (the direct route’s owner, a one-task plan, or a task-id run) is implemented inline by the conversation: no worker and no worktree unless you ask for one (--worktree). A worker takes a single task only when you named an implementer model or tier, the task’s review mode resolves to host, or the conversation’s context is too full, and the run says which reason applied.
Before building a large spec inline, start from a fresh session (or /clear) so the build has room, or name an implementer model and a fresh-context worker takes the task.
With several tasks, /flow-next:work schedules them on the rolling frontier by default. The moment any in-flight worker returns, the next eligible task is admitted, in its own isolated workspace, and the conductor integrates, reviews, and completes each task as it lands. It falls back to the wave loop only when the spec’s shape gives rolling nothing to schedule. Every dispatched task gets a fresh-context worker and the same review, evidence, and completion gates on either route.
Worker subagent model
Section titled “Worker subagent model”On a multi-task spec, each task runs in its own subagent. The benefits compound:
- Fresh context per task prevents bleed between unrelated changes.
- Re-anchor information stays bundled with its implementation.
- Review cycles stay isolated to the task they belong to.
- The main conversation stays lean even on long specs.
flowchart TB Main["Main session"] --> Pick["Admit ready tasks (cap 3)"] Pick --> Spawn["Dispatch worker per task"] Spawn --> Anchor["Worker re-reads spec"] Anchor --> Impl["Implement + commit in isolated workspace"] Impl --> Ret["Worker returns"] Ret --> Join["Integrate that task"] Join --> Review["Per-task review (conductor-owned)"] Review --> Done["Mark task done"] Done --> Pick
The re-anchor is a single flowctl anchor <task-id> call: one deterministic bundle carrying the task, the parent spec, git state, memory and glossary indices, and dependency summaries, so each worker starts in one round-trip instead of eight. Each section is the verbatim output of the command it names. Since 6.1.0 those commands are the narrow forms: the memory index as text, only the glossary entries the task’s title or description names (glossary list --match), git status --short --branch, and the spec record without its review and tracker ledgers. A typical bundle dropped from about 160 KB to about 60 KB, and both bundles scored 7 of 7 on the comprehension eval across three task sets. The bundle is a floor, not a ceiling; the worker still reads anything else it needs.
A worker never returns while a command it started is still running. It launches its gates in the foreground, and when the host moves a long command to the background it waits for that command to exit and reads the exit code before it reports. A result therefore always describes a finished run. The conductor checks the same thing from its side: before it accepts a return, it reads the task status and looks for a live command in that worker’s workspace. If it finds one, it waits within the dispatch timebox, then sends a re-anchoring continuation worker into the same workspace once the command exits. The early return is not counted as a failed attempt. To look at another tree state a worker uses a temporary worktree, never git stash.
Scheduling: rolling frontier by default
Section titled “Scheduling: rolling frontier by default”Two schedulers exist inside work, and the route is decided once, from spec state, at the start of the task loop:
| Route | When | What happens |
|---|---|---|
| Rolling (default) | Everything else | A ready task is admitted at every worker-return event (same five fail-closed conditions as a wave: same spec, cap 3, no dependency path, disjoint declared paths, nothing in the always-serial set - judged against the in-flight set). Each worker gets an isolated worktree; the conductor integrates, reviews, and completes each task as it returns; a shared outside-tree notes surface lets workers leave findings for siblings, read by pointer. |
| Wave (fallback) | The run was given a task id (only that task runs); planSync.enabled is not false; the spec has fewer than two open tasks (a no-plan implicit task, or one task left); or no two open tasks are dependency-independent | Concurrent safe subsets joined at wave boundaries, or a single worker in the checkout. Plan-sync runs after each joined wave. |
The route prints once, before the first task is claimed, as Scheduling: rolling or Scheduling: wave (<reason>). There is no config knob and no host table. A host whose subagent dispatch is measured to block reports Scheduling: degraded to wave (host lacks non-blocking dispatch) and keeps the rest of the rolling lifecycle. flow --auto and land dispatch plain /flow-next:work and inherit the route.
Why rolling, and why isolated worktrees. A pre-registered three-arm eval compared architectures on the same specs. Rolling admission with a worktree per task saved 52.1% of work-phase wall clock at quality parity with zero uncontained correctness incidents. A shared-checkout arm was faster still (69.2%) and failed quality parity: part of its speed came from roughly 30% thinner test artifacts, because making the declared-paths list the commit boundary disincentivized new test files. The per-task worktree pool removes that pressure, so speed is never bought with under-testing. The scheduler shipped first as the experimental /flow-next:work-rolling beta, ran end-to-end on Claude Code, Cursor, and Grok Build, and then graduated into work as the default; the beta command is gone.
Playing nice with other runs. Task claims are spec-scoped: two runs on the same spec contend on the same claims and fail closed against each other - a task claimed by another actor is dropped from the run’s admissible set, never stolen. Since 5.5.0 a task you hold yourself is no different at the claim: a second flowctl start as the same actor on an in_progress task refuses with a typed error instead of reading as a crash resume, and a genuine resume passes --reclaim after confirming the prior run ended. A merge conflict at per-task integration retries that task serially: one re-run, never a correctness loss.
Quick commands: what gets verified, and how often
Section titled “Quick commands: what gets verified, and how often”The spec’s ## Quick commands block names the focused checks for the code a task touches. A worker (several tasks) runs them once before its first edit, to prove the tree was green before the change, and again before it may mark the task done; a single task built inline runs its failing test first and focused tests after. The full gates run only when the repository’s instructions or you ask for a full suite: once, at the end, and not again after later fixes. CI owns regressions.
The scaffold’s convention is that per-task commands name focused suites for the files that task touches. Repeats are cheap by construction. A green receipt keyed to the exact commit and the exact command string skips a re-run at unchanged HEAD, and a docs-only diff drops to lint and format only.
Code in sibling repos. In a home-base workspace, where .flow/ lives in a planning repo and the product code in sibling clones, work records the base of each sibling the spec changes (your project instructions or the spec name them) and classifies each repo’s diff too. The docs-only tier applies only when every repo’s diff is docs-only, and a sibling it cannot read runs its full gates. The feature-map update reads the sibling diffs as well, so a route moved in a sibling updates the map. A single-repo workspace works exactly as before.
Which means the tier is yours to set. Author focused commands and keep the full-suite entrypoint as the final-gate command if you want narrow per-task runs; list the full suite if your project would rather pay it every task; put the policy in CLAUDE.md / AGENTS.md if it should hold across specs. Selection is authored, never computed - it keeps needless full runs out of the loop, but it will not infer which distant suite your change actually touches. See the verification spine.
Start without a plan
Section titled “Start without a plan”Direct execution through /flow-next:work <spec-id> --no-plan is the default for a ready spec. Plan is chosen only on a positive signal: you asked for a plan, separate human owners will implement, or delivery is staged across several PRs. Risk, size, and file count never trigger plan on their own. Design risk routes to plan-review, which reviews a spec with zero tasks; unresolved product or authority choices route to refine. Explicit --no-plan or stated intent answers the zero-task route choice; without either, work’s ask on a zero-task spec reads the same plan-versus-no-plan rule the flow skill owns and recommends its result. On an already-planned spec, the flag is ignored with a one-line notice.
/flow-next:work fn-N --no-planIdea text and single tasks
Section titled “Idea text and single tasks”/flow-next:work "rename the config key" accepts idea text and mints the minimal spec and task itself, so a spec always exists underneath and the same evidence gates apply. /flow-next:work fn-N.M instead selects one existing planned task.
An explicit request to review the design, or a recorded design review that still needs work or human judgment, stops direct work before owner creation or resume. Run /flow-next:plan-review for the spec, resolve its findings, then invoke work again. Work’s --review option selects implementation review and does not clear this gate.
The strong spec remains the implementation contract. One implicit execution owner covers every acceptance criterion and follows the same configured implementation review, coverage, completion policy, and evidence gates. Optional live QA stays opt-in and augments staging and manual QA; a NEEDS_WORK QA finding can accompany the PR and does not authorize merge. Work records the explicit direct-route choice before creating its single execution owner, so a resumed run preserves the route. flow --auto and work can resume that same actor’s sole owner only after confirming the prior run ended, including when work receives a task ID; active or unknown ownership remains NEEDS_HUMAN. The owner implements the full spec, with execution-time decomposition or justified delegation as needed. flowctl spec set-no-plan fn-N or capture with --no-plan records the same choice for flow --auto, and a flow --auto run that meets a ready zero-task spec with no recorded route applies the plan-versus-no-plan rule and records the result on the spec before any mint. Work takes consent only from the flag, stated intent, or the field. Without one of them, work stops on a zero-task spec. See How Flow chooses.
Fixing a reported defect
Section titled “Fixing a reported defect”When a task fixes a reported defect, whether flow routed a bug report here or the spec’s requirement is that reported behaviour stops happening, the worker runs four steps before and around the fix: it checks for prior fixes, reproduces the symptom reliably and confirms its cause with runtime evidence, bisects when a known-good revision exists, and proves the fix by running the same reproduction on the base (failing) and the head (passing). A cheap reproduction test is committed failing before the fix. An existing fix, or a fix someone visibly owns, stops the task rather than getting a competing one. Other open pull requests that touch the area are noted in the done summary, and the task continues. The done summary records each step, including the ones skipped and why, and make-pr turns that record into proof cells. How Flow chooses describes the steps in full.
Climbing a metric toward a target
Section titled “Climbing a metric toward a target”When a spec asks for one number moved toward a target through repeated attempts, the worker runs a measured loop instead of a single change. It sets up the experiment itself, using your numbers where you gave them, and writes the setup at the top of the ledger before the first attempt; only a spec with no target stops it. It builds a harness, proves it can tell better from worse and rejects a wrong output, freezes it with a hash, then tries one hypothesis per attempt against the current best. An attempt is kept, as exactly one commit, only when it beats the best by more than the noise and the regression gate is green; anything else is reverted to a clean tree. Every attempt writes a ledger row, and the run stops on the target plus the attempt floor, on the budget, when the supported ideas run out, or when the harness breaks. The target is yours and is never relaxed. Review runs once over the kept commits, and the done summary carries the full ledger. How Flow chooses walks through the loop and a worked example.
Branch modes
Section titled “Branch modes”| Mode | Use |
|---|---|
current | Stay on the active branch. The default when you are not on the default branch. |
new | Create a fresh feature branch off the current base. The default on the default branch, and always under flow --auto. |
worktree | Fully isolated parallel work via flow-next-worktree-kit. |
Without a branch option, work asks nothing: it takes the default and says which in one line. A worktree is used only when you ask for one.
Worktree mode is the right choice when more than one spec is in flight, or when a review cycle on a separate branch must not disturb the work in progress.
Dependent specs branch from the parent
Section titled “Dependent specs branch from the parent”A spec whose parent (depends_on_epics) has every task done and its branch on origin starts now instead of after the parent’s PR merges. Work asks flowctl spec chain <id> once, fetches the parent’s branch, and creates the spec branch from origin/<parent-branch>; the recorded spec base is the merge-base with that tip, so gate classification and the quality auditor see only this spec’s own diff. --branch=current on such a spec requires the current branch to already contain the parent tip, otherwise the run stops with BLOCKED naming the missing ancestry; worktree mode applies the same base. A spec the predicate refuses (parent branch not pushed, two open parents, a sibling already chained on the same parent, or a failed remote query) stops before any task starts with BLOCKED: <reason>. Make-pr later opens the PR against the parent’s branch and, on GitHub, links it into the parent’s stack: chains and stacks. A spec with no open parent branches exactly as before.
Plan-sync after each joined wave
Section titled “Plan-sync after each joined wave”When planSync.enabled is true, work takes the wave route (plan-sync’s per-wave barrier is a fail-closed rule) and downstream task specs are checked for stale references after a serial task or after the conductor has joined and resolved an entire parallel wave. This avoids syncing against a partial integration. If a task changed an API or path that later tasks assumed, the drift surfaces with a reason. The user decides whether to update the downstream specs, regenerate them, or accept the drift explicitly.
Plan-sync is conservative. It never silently rewrites a spec; it only proposes.
Quality audit (large or risky changes)
Section titled “Quality audit (large or risky changes)”In the quality phase - after tasks complete, alongside the full gates - the conductor may dispatch the quality-auditor subagent over the spec’s whole diff. The trigger is a judgment call, not a flag. It runs when the conductor judges the change large or risky, and is skipped for small routine changes. It dispatches as two parallel axis-scoped runs (correctness and standards), both reports come back verbatim, and only correctness-axis Criticals block.
Feature-map update
Section titled “Feature-map update”When the repository has a feature map at .flow/features/, the quality phase also checks whether this spec’s change altered how a user reaches a mapped feature. If it did, work edits only those feature files, proves each new route with one live drive, refreshes their **Last proven:** lines, and retires any drift notes naming a route it proved; the map diff rides the same commit as the code. If the app cannot start, or a changed surface cannot be tied to one feature file, work leaves the map unchanged and files a drift note instead, without blocking the run. It never adds new features and never runs the full /flow-next:features pass. Without a map the step costs one existence check. See Keep the feature map current.
Completion review gate
Section titled “Completion review gate”When all tasks are done, an optional /flow-next:spec-completion-review runs to verify the combined implementation matches the spec end-to-end. This is the place to catch criteria that no individual task fully owned.
Optional review
Section titled “Optional review”/flow-next:work <spec-id> --review=codex--review=codex|rp|copilot|cursor|claude|host|none runs a per-task adversarial review, sized by risk: three reviewers from another model family for a change that touches persisted or shared state, concurrency, security, data layout or migrations, or a multi-file feature; one reviewer otherwise; none for a small local output, wording or display fix, with the skip recorded. Attended, a NEEDS_WORK gets one fix pass and one re-review scoped to the fixes. Under flow --auto the fix and scoped re-review repeat until SHIP before the task is marked done. See Impl Review.
Without --review, work reads the configured backend for each task from the repository root, so a task’s own review: pin wins over the project default. Work never asks about review. With no backend configured, review is off and the handoff says so once (“no review backend set; run setup or set review.backend”); set one with /flow-next:setup, flowctl config set review.backend <backend>, or in the prompt.
Implementation offload
Section titled “Implementation offload”Handing the token-heavy part - writing code - to a second CLI agent is a routing decision you write, not a subsystem you configure. Name an implementer tier in your CLAUDE.md / AGENTS.md routing block, or just say it in the moment; the host drives the other CLI through a headless bridge for the draft. With no implementer tier, the session model implements - that is the default and needs no configuration.
Two rules are not optional:
- The bridged child owns the task; the worker keeps judgment. The child writes code, commits checkpoints on the branch it was given, and decides its own delegation under the same license an in-host worker holds. It never pushes, never rebases or rewrites history, never decides scope, never issues a review verdict, and never spawns a bridge of its own. The worker hands the task over (Phase 1b), reviews the child’s commit range from the recorded base, runs the gates, dispatches review, and owns
flowctl done; the conductor never bridges. Your done summary records which model implemented and how many subagents the child dispatched. - Send clear, well-scoped tasks to the value tier. On well-specified work a value tier matches a strong tier on correctness for meaningfully less wall clock; escalate to the strong tier only for genuinely gnarly ones. Spec quality is what makes the trade safe - a vague brief burns the saving on rework.
Bridge recipes ship into your repo and are read on demand with flowctl usage (## Orchestration & model steering). Full doctrine: Orchestration → Implementation offload.
Worked example
Section titled “Worked example”/flow-next:work fn-12-export-json-flagScheduling: rollingready frontier: [.1, .2, .3] in-flight: [] admitted: [.1, .2] held: [.3: depends on .1] .2 -> returned: commit e4f5a6b, tests green; integrated; review SHIP; done admitted: [.3] .1 -> returned: commit a1b2c3d, tests green; integrated; review SHIP; done .3 -> returned: commit c7d8e9f, tests green; integrated; review SHIP; donequiesce: full suite green; completion review: SHIP - receipt on diskSpec fn-12-export-json-flag: all tasks done.Fresh context per task plus mandatory evidence is what makes the loop repeatable instead of lucky.
Dynamic usage
Section titled “Dynamic usage”Recipes that compose with work in the cookbook:
- Parallelize - both supported forms of task-level parallelism, with the guardrails that keep them safe.
- Model routing - an
implementertier hands implementation to a second CLI while the host keeps git, review, and judgment. - Skip & lighten - plan + work alone is a complete, evidence-backed workflow for small changes.
Next step
Section titled “Next step”/flow-next:make-pr <spec-id>