Skip to content

Flow and the road ahead

Flow-Next 5.0.0 adds one command that reads whatever you have and decides what happens next, and 5.1.0 gives that command an unattended shape, flow --auto. This page explains why the command exists, what it does with your work, what shipped in 5.1.0, and how 7.0 made the route lean.

Before 5.0.0, four places each held a copy of the same decision: which stage runs next. Capture printed its own Recommended next: line, plan printed its own menu, work asked its own question on a zero-task spec, and a separate guide skill kept a third copy of the matrix. Copies drift. A guide recommendation and a capture closer could disagree, and either way you read the advice and typed the command yourself.

The 5.0.0 answer is one routing reference and one conductor that reads it. The reference is a set of files owned by the flow skill, one per rule, among them the route matrix, the spec-count rule, the plan-versus-no-plan rule, gate selection, prototype-before-ask, and the tail rule that says where an attended run ends. Capture’s closer, plan’s menu, work’s zero-task ask, and flow --explain read the same files at the step that needs them. The recommendation you read and the route that runs are the same rule, and a test fails when any pointer names a file that does not exist.

/flow-next:flow accepts anything: nothing at all, an idea, a spec or task id, a tracker issue, a branch, a path, a pasted bug report, a how or why question, something slow, a cleanup that must keep behaviour, a design fork. It routes on content and context, never on input kind.

$ /flow-next:flow "the export endpoint takes four seconds on the demo tenant"
Flow stopped at: PR exists
Route taken: work -> make-pr
stage: work - ran [2026-09-12T09:02:11Z..2026-09-12T09:41:58Z]
stage: qa - skipped(config: pipeline.qa=auto: no UI-observable criteria)
stage: make-pr - ran [2026-09-12T09:41:59Z..2026-09-12T09:42:31Z]
Next: review https://github.com/acme/api/pull/214, then merge or send it back

The ran [<start>..<end>] ranges are the stage’s start and end times where the run knows them; the whole report grammar is on the Flow page. Since 7.1.0 an attended run like this one first stops after work with the change committed on a local branch, and runs make-pr when you say “open the PR”; the report above is the one you get after that.

The route matched the performance row (a metric and a surface you can name). Work measured first (p95 4.1 s on the demo tenant over 3 runs), made that baseline and its target the requirement, and recorded the post-change measurement (p95 0.6 s) as the task’s evidence; the review inside work read both numbers. The same release added two other routes. A why question dispatches the read-only why-scout, which answers from git blame, the PRs behind the commits, and the bug and decision memory with each finding tiered direct, supported, inferred, or unknown; and a ready spec that names a library the repo does not already use runs /flow-next:refine --scope=research first, satisfied by an existing ## Resolved via Research section for what it already covers; a library named since then reruns the delta only, and --force replaces the section.

Three properties hold on every run:

  • Every stage keeps its own contract. Flow re-implements no stage. Capture, refine, plan, plan-review, work, qa, make-pr, and resolve-pr keep their own contracts, receipts, and gates. Every stage flow reaches records ran, skipped(reason), or failed(reason), so a skipped stage is an event with a reason, never an absence.
  • It stops at your decisions. A pick among options a stage produced is asked inline and the run continues. A decision that ends the run stops it: missing merge consent, a review verdict that needs a person, a product question no stage framed as options. Before any “which approach” question, flow classifies the fork; an observable answer is settled by running something, and the report names what ran.
  • Direct execution is the default. A ready spec with no tasks routes to work --no-plan. Plan is chosen only on a positive signal: you asked for one, separate human owners will implement, or delivery is staged across several PRs. Risk, size, and file count never trigger a plan on their own. Internal benchmarking found the direct route can produce higher-scoring implementations with capable frontier models, because one owner sees the whole task; decomposition is now the exception with a stated reason.

flow --explain prints the route, its positive signal, the safe skip and its kind, and why not the alternatives, and writes nothing. How Flow chooses walks the situations it answers.

Attended flow refuses to run under any autonomy marker (FLOW_AUTONOMOUS, AUTONOMOUS=1, a mode:autonomous token). Under one it stops before any read or write with this line:

NEEDS_HUMAN: /flow-next:flow is attended - run /flow-next:flow --auto for unattended runs

flow --auto has no marker refusal, because it sets FLOW_AUTONOMOUS and mode:autonomous itself for the stages it dispatches. Attended flow and flow --auto are two drivers and are never nested; since 7.0 they are the only two.

By default, an attended run from intent ends with the change committed on a local branch and opens the pull request when you ask (since 7.1.0); flow --auto ends when the PR exists. Neither asks about landing a pull request the run just opened. Plain attended flow on an existing PR offers landing and asks once unless that item is already authorized. --until=merge authorizes continuation through land for the selected item, independently of attended or autonomous mode. Land owns every merge gate; make-pr closes the spec before the pull request opens, so nothing is written after the merge. Flow never fabricates a review, QA, or completion verdict; recursive drivers remain forbidden.

Flow-Next 5.1.0 made flow --auto the unattended driver. One /flow-next:flow --auto [<spec-id>] invocation carries a ready spec to a pull request, hop after hop; --tick runs exactly one hop and stops, which is what a pilot tick was. Both shapes classify each hop from the same routing references the attended conductor reads (route-matrix.md for the spec-state row, plan-vs-no-plan.md for a ready zero-task spec with no recorded route, gate-selection.md for the review, QA, and completion-review gates), so the route --explain shows is the route the unattended run takes. Both end with one PILOT_VERDICT line; a long-horizon run names every dispatched stage joined by + (stage=work+qa+make-pr) and carries the last hop’s verdict.

/flow-next:pilot became a one-release alias for flow --auto --tick in 5.1.0, and 6.1.0 removed the command. Pilot’s rails carry over unchanged: the ready-flag consent boundary, the dirty-tree refusal at run start and after every hop, the two-strike ledger (flowctl pilot strikes list|clear), the all-done PR probe, the never-merge boundary, the decision log (flowctl pilot-log, .flow/pilot-runs/), and PRs that stop short of merge. Since 7.0 those PRs open ready for review unless they carry open items, and the body carries the run’s Decisions list. Every hop ends with the same receipts, evidence echo, and ledger write a pilot tick ended with, so a run that dies mid-way resumes from disk on the next invocation.

pipeline.qa=auto now takes effect unattended through the same gate-selection rule. A skipped hop records stage: qa - skipped(config: pipeline.qa=auto: <reason>) and the route advances to make-pr; on and off are unchanged.

7.0 made the route lean. Flow picks it in seconds and asks no branch or readiness question before a build: it takes a sensible default and says so in one line. A small local change goes direct, with no spec, and a PR only if you ask. A single-task spec is built inline in the conversation; workers and the scheduler run only for several tasks or a task you send to a chosen model, and a worktree only when you ask. Review goes by risk, and a stage is added only where the risk calls for it. The work comes back in about the time plain Claude Code, or your harness, takes, often less, across more than 170 full end-to-end runs; the result is better than the plain agent’s even before any review, and the optional stages (cross-model review, live QA) widen the gap to up to 25% better outcomes, especially on large and long-horizon work.

Use /flow-next:flow fn-12 --until=merge, or add --auto for an unattended run, to carry one item through landing. Without that destination, unattended flow keeps the pre-merge stop. The Flow reference covers consent across retries and fresh sessions, waits, revocation, and recovery after a successful merge. Standalone land stays available.

Model and harness recommendations are deferred. The existing orchestration block already lets you choose who runs each stage.

Deferred: model and harness recommendations

Section titled “Deferred: model and harness recommendations”

Use the orchestration block to choose models, harnesses, and effort today. Automatic implementer selection and a model-capability scorer are not planned.

Since 5.2.1, choosing another implementation model or harness does not itself require task decomposition. Plan when you request it, separate human owners will implement, or delivery spans multiple PRs. Research and review retain their own routing rules.

Today: one dial from interactive to autonomous. /flow-next:flow while you are at the keyboard, /flow-next:flow --auto fn-12 --until=merge for one item or standalone land for scheduled babysitting when you are not, the same gates at every rung. One conductor now sits at both ends of that dial. Your orchestration block chooses who does the work; additional recommendation tooling remains deferred.

  • Flow - the skill reference: inputs, the report shape, --explain, the routing files, the refusals.
  • How Flow chooses - the flow --explain walkthrough, situation by situation.
  • Compose the pipeline - the doctrine flow applies, and what holds on every route.
  • Flow —auto - the build loop: the consent boundary, what one hop does, the strikes ledger, backlog mode.
  • Going autonomous - the unattended tier as it ships today.