Skip to content

Refine

/flow-next:refine (interview) sharpens an existing spec before anyone builds it. By default it is one question pass: it asks only the open decisions that would change what gets built, product, technical, or anything else, and writes each answer into the section it belongs in on the same artifact. Under --scope=research it asks nothing: four read-only scouts resolve the libraries and APIs the spec names into one cited section.

Refine when you can name at least one open decision that would change what gets built and that only you can make. Skip it when:

  • the acceptance criteria state the intended behaviour and what remains is how;
  • the touched area already has established patterns;
  • the only gaps are technical detail, performance, or edge cases that implementation, review, and QA will surface;
  • the only uncertainty is criteria capture inferred itself;
  • the work is a defect, a structural cleanup, or a request to move a measured number.

A technical decision is worth a session only when it is costly to reverse and the code does not answer it: a data model or migration, a public contract, a security boundary. Flow, capture, and plan apply the same rule before they recommend refine.

Capture the ticket first; that is capture’s job (it proposes the spec count and tags every line with its provenance). Refine earns its keep after the spec exists, pointed at what is actually soft:

  • Burn down the guesses. Capture leaves your own words untagged and tags what it authored [paraphrase] or [inferred] - so aim the interview at exactly the inferred ones: “interview me on this spec regarding anything that has been inferred”. Questions then land on lines the agent guessed, not on things you already stated, and only on the guesses that would change the build.
  • Pressure-test one requirement. “In-depth technical interview regarding R9” - the whole question machinery (recommendation-led questions, codebase investigation before asking, experiments instead of questions) concentrates on the single criterion that worries you instead of re-walking the spec.
  • Settle a fork the code cannot. Refine is worth running when a costly-to-reverse choice is still open - a data model, a public contract, a security boundary. Edge cases and performance detail are left to implementation, review, and QA.

Capture marks the guesses; refine burns them down.

A pattern we see in the field: teams treat each question’s proposed options as a multiple-choice test and blindly pick one - usually the recommended one. That inverts the tool. The AI proposes answers to help you think about the question; it should not be making these judgment calls for you. The recommendation is a starting position with a stated confidence, not a verdict.

When none of the options fit, that is not friction - it is the interview working. Use the Other option, and remember it is fully agentic, not just a text field for a fourth answer:

  • “None of these - here’s what actually matters: …” - answer in your own words; the spec records your words untagged, as yours.
  • “Go do more research on the codebase before asking me this” - the interview investigates and comes back with a better-grounded question.
  • “This question isn’t the right one - focus on the migration risk instead” - redirect the whole line of questioning.
  • “Not decidable yet, park it” - it lands in ## Open Questions instead of forcing a fake decision.

An interview where you picked the recommended option every time didn’t sharpen anything - it laundered the AI’s guesses into requirements with your name on them. The questions where you push back are the ones paying for the session.

Do not reopen discovery as chart just because a valid spec has judgment gaps - refine owns those. Route backward to chart only when the questions reveal the effort itself is not yet specifiable. Chart also has an attended interview decision type that uses the same question machinery on a chart rather than a spec; that path never writes .flow/specs/.

--scope=research selects the read-only pass described below. Every other use of refine is the same question pass, with no scope question up front.

A question is asked only when all three hold:

  1. A wrong guess would build the wrong thing or ship behaviour you would reject.
  2. The code, the docs, a quick experiment, or implementation itself cannot settle it.
  3. It is your call.

Everything else the agent resolves, records, or leaves to work. Refine stops as soon as no question that passes the test remains. Asking nothing is a good outcome: the run reports “Nothing worth asking; the spec is clear enough to build.”

Answers keep the precision you gave them. A preference such as “performance matters here” goes to ## Decision Context as guidance, not an acceptance criterion. A number or measurable commitment becomes a criterion only when you stated it; the agent’s recommended options never become thresholds. Refine does not ask for success metrics or latency budgets unless the spec is about them.

The interview carries a short check-list of what to look for in the spec: who it is for, what done looks like, what is explicitly out, a constraint the domain implies (a regulation, a contract, a partner commitment), an irreversible data or contract change (a data model, a migration, a public contract), an external interface, a security boundary. A topic on the list is asked about only when the spec leaves it unclear and the question passes the one test; the list is not a set of questions to walk.

--scope=<anything> is an optional free-text lens: business, technical, qa, security, ops, or any other audience. There is no predefined list. The agent interprets the lens and focuses the questions on that audience’s open decisions, which still have to pass the one test. --biz and --tech are aliases for --scope=business and --scope=technical, and plain language such as “run a business interview” works too. No scope means no filter.

Terminal window
/flow-next:refine fn-1
/flow-next:refine fn-1 --biz
/flow-next:refine fn-1 --tech --strategy --docs
/flow-next:refine fn-1 --scope=qa

Every session writes to the same spec file, and R-IDs are append-only across sessions, so a product owner’s session followed by a developer’s produces one continuous chain. Sections no answer belongs in come back byte-for-byte, and the read-back before writing names every section the session changed. A product owner and a developer refining the same spec in separate sessions each see any change outside their layer. Run a session only when it has an open decision to settle.

Existing specs keep their layout. A spec that already carries ### Motivation / ### Implementation Tradeoffs sub-headings or <!-- scope: ... --> markers loads and refines without being rewritten.

The interview maps the spec as a design tree - every decision branches into the decisions that hang off it - and asks in rounds over the tree’s frontier: the questions whose prerequisites are already settled, askable now without guessing at answers not yet heard.

  • Each round asks the whole frontier, split across question calls of up to 4 questions each, grouped by topic and announced as one round (“Round 2 - part 1/2”).
  • A question is never asked alongside its own prerequisite. Anything that depends on an answer still open in the current round waits for a later round.
  • A frontier slot is earned by passing the one test. Failure modes, concurrency, scale, portability, and testing qualify only when they pass it; implementation, review, and QA surface the rest. Pure-cosmetic polish (message wording, label spelling) never gets its own question: it folds into a related question’s options or a stated default the user can veto at write-back.
  • The frontier is recomputed between rounds. Answers reshape the tree: settled decisions unblock their dependents, pruned branches are announced at the next round’s opener (“Skipping persistence questions - you said no DB”), and the interview is done when no question that passes the test remains. A follow-up is asked only when an answer opens a new question that passes it.

Standalone checkpoints - the code-mismatch question, write-back consent, the mark-ready offer - sit outside rounds and are never counted against them.

Write-back is a summary and one ask. The refined draft is written once to a temporary file, and you see a compact summary (title, criteria, the source tally, a split proposal when one exists, the recommended route) with one ask: approve and write, open in editor, or abort, plus free text for edits. Edit cycles print only the diff; the full draft prints only on request. Ratification still precedes every .flow/ write.

While the user answers a round, the interviewer may dispatch a read-only fact-scout to resolve in the background the codebase lookups gating the next round’s questions - investigation latency hides inside answer time instead of stalling the interview between rounds.

Refine does not ask about deadlines, sprint cadence, or “ship before X”. Two reasons:

  • Agents can’t reliably estimate their own work, so any answer is a guess that anchors the rest of the interview.
  • Time-pressure framing collapsed interviews into brutal-prioritization debates instead of surfacing requirements.

If you volunteer a deadline, refine acknowledges it without chasing it through follow-up questions. What is explicitly out is settled by value, not by the clock.

Refine integrates with the repo’s existing documents. Before its first question, every interview reads STRATEGY.md and searches the rest of the project docs (README, changelog, glossary, recent decision records, open specs, docs/) and the codebase for what the spec touches, instead of reading them end to end. An answer found in the docs lands under ## Resolved via Project Docs, one found in the code under ## Resolved via Codebase, and neither becomes a question.

  • Resolves vocabulary against GLOSSARY.md.
  • Surfaces foreign-file references when answers cite paths.
  • Flags contradictions with active STRATEGY.md tracks.
  • Writes a decision-record entry when a strategy track is intentionally overridden.

Doc-aware meta-questions carry a per-round budget: at most one glossary question, one strategy-conflict question, and three doc-aware questions combined per round. A meta-question deferred by the budget is held for a later round, not dropped - the one sanctioned hold-back in the rounds protocol.

Refine tags the acceptance criteria it newly writes with the same vocabulary /flow-next:capture uses. The words of whoever ran the session (the PO in a product session, the tech lead in a technical one) stay untagged and must be findable in their answers; [paraphrase] marks their meaning in the agent’s wording, [inferred] the agent’s own fill-in, and [strategy:<track>] a line that traces to a STRATEGY.md track.

That makes “which of these did the agent fill in?” a grep rather than a re-read - see the tally recipe. Three rules to rely on:

  • A session tags only what it authors and never adds, changes or removes a tag on an existing bullet, so provenance is frozen exactly like the R-ID number. On a spec refined in several sessions, each tag reflects the session that wrote it.
  • A criterion someone answered is untagged or [paraphrase]. Only genuine gap-fill is [inferred].
  • The read-back will not recommend approve while unverified [inferred] items remain - narrowed here to inferred criteria that no question covered, since an answered question has already done the verifying.

Tags apply to a spec’s ## Acceptance Criteria bullets. Task acceptance is a plain checklist and carries none, and a pass over a loose markdown file leaves that file’s shape alone.

When refinement pushes a spec past 8 counted requirements (business and technical only - standing criteria and process items never count) or an answer reveals a second independently shippable outcome, refine proposes a split before writing back, applying the same spec-count rule flow and capture read: proposed titles, requirement allocation, and dependency edges, with keep-single as the default. Criteria a review cycle has already judged are never moved or renumbered - for those the proposal is recorded in the spec’s Decision Context instead. Autonomous runs never split.

After the write-back, when readiness is adopted in the repo (≥1 spec already marked ready) and tracker.readyState is not configured, refine offers once to mark the refined spec ready for execution. Default is keep-draft - re-read the refined spec on disk before blessing it. The offer applies to flow-spec inputs only (task ids and file paths carry no spec readiness), and refinement never auto-resets a previously-blessed spec - only /flow-next:capture --rewrite does. Non-adopters see no question anywhere; tracker-connected repos set readiness on the board.

Some questions are not yours to answer. They are facts the agent can observe by running something: how an input behaves, how long something takes, whether a layout fits at 320 px, what a command outputs, whether an eval separates two options. Refine classifies every question before asking it. When running something can answer the question, refine runs a throwaway experiment instead of asking and writes the result under ## Resolved via Experiment: the question, what ran, what it observed, and the decision that follows. Later sessions keep that section as written, like the other Resolved via sections.

Refine runs an experiment on its own only when it is read-only or fully disposable. Its files live in .flow/tmp/experiments/, which git ignores, and are thrown away or quoted in the spec as evidence; none of it ships. An experiment that needs live or shared state, credentials, the network, or anything destructive becomes a question to you. So does a result too noisy to decide, and the question comes with the data attached. Product and preference calls always come to you.

Example: refining a spec for a new search box raises three questions.

  • Should results update on every keystroke or after a pause? Running something can answer it. Refine builds a throwaway page with both behaviours behind a toggle, types a 12-character query against the local index, and measures 180 ms per keystroke update with visible stutter, while a 150 ms pause feels immediate. It records the numbers and writes “update after a 150 ms pause” into the spec.
  • Does the existing tokenizer handle accented names? It runs the tokenizer on five sample names, sees two split wrongly, and records that as a constraint.
  • Should search include archived projects? That is a product call, so refine asks you.

Each question leads with the recommended option and a confidence tier:

  • [high] - the agent is confident in the recommendation.
  • [judgment-call] - reasonable people disagree; the user should weigh in.
  • [your-call] - the agent has no view; the user owns the decision.

Questions arrive in rounds - up to 4 topically grouped questions per call, the whole frontier per round.

Skips are not answers (2.9.0). Only an explicit answer or an explicit “you decide” delegation resolves a question - the agent’s recommendation never silently becomes spec content. Every skipped, declined, or “I don’t know” question parks under ## Open Questions with an owner hint and the agent’s unconfirmed leaning. When at least one question was skipped, a consent checkpoint runs before the spec is written back:

  • park-open (default) - skipped items land under Open Questions only; nothing skipped becomes a decision.
  • fill-assumptions - the agent’s recommendations are written into the spec, each marked inline *(assumed - unconfirmed)*, with an Open Questions pointer for later ratification.
  • re-ask - walk the skipped questions once more.

/flow-next:refine <spec-id> --scope=research is the read-first pass. It asks no questions and runs no rounds. It dispatches four read-only scouts in parallel over the spec’s (or task’s) surface:

ScoutReturns
docs-scoutOfficial docs, version anchored on the repo’s manifest scan
practice-scoutCurrent best practices and the pitfalls that bite at implementation time
docs-gap-scoutThe repo’s own docs that must change
memory-scoutThe memory entries that apply

github-scout joins when scouts.github is on. The pass writes one ## Resolved via Research section on the spec (or the task body for a task target): one sub-block per scout that ran, one bullet per finding, a source on every line. It lands through the same write-back and read-back contract as the interview, so you ratify the section before it is written.

The skip rule, shared with plan. The pass is observably skipped, with the reason printed and nothing written, when the target already carries ## Resolved via Research or when the spec’s tasks carry plan’s scout findings (it names the task that holds them). When either case would skip but the spec now names a library or API that neither mentions, the pass reruns for that delta only and appends under the right sub-block. --force reruns the whole pass and replaces the section. Plan’s research step applies the same rule in reverse: when the section exists it skips its research scouts (docs, practice and docs-gap; repo-scout, spec-scout, and the gap analyst still run), and when it runs them it writes the section too. The pass never runs twice.

When flow routes to it. Flow runs the research scope before work when a ready spec names a library or API the repo does not already use. The signal is satisfied by the section or by a plan that ran the scouts, so it never runs by default.

Terminal window
/flow-next:refine fn-12 --scope=research
/flow-next:refine fn-12.3 --scope=research
/flow-next:refine fn-12 --scope=research --force

A file-path target prints research: skipped(policy: research writes a spec or task section; give a spec or task id) and stops.

/flow-next:refine fn-14-rate-limits
Reading spec fn-14-rate-limits... 3 gaps found: limit scope (per-user? per-key?),
burst behavior, and the 429 response contract.
Round 1: 4 questions (multiple choice + free text)...
Round 2: 2 follow-ups on the sliding-window choice...
Spec updated: Boundaries sharpened, R3 rewritten, R6 added, 0 open questions remain.

Refine extracts decisions you already half-made and pins them into the spec before any code exists.

Recipes that compose with refine in the cookbook:

  • Prompt into a stage - “interview me only about the error-handling section” scopes the rounds.
  • Team patterns - refine + Spec-as-PR is the async replacement for a refinement meeting.
Terminal window
/flow-next:work <spec-id> --no-plan

Direct execution is the default for a ready spec. Plan only on a positive signal: you asked for one, separate human owners will implement, or delivery is staged across several PRs. Design risk routes to /flow-next:plan-review, which works before any tasks exist. Unsure: /flow-next:flow --explain fn-N prints the route, and plain /flow-next:flow fn-N runs it.

If the pass reveals fundamental ambiguity, return to /flow-next:capture --rewrite or /flow-next:strategy before building.