Skip to content

Flow

/flow-next:flow chooses the next step so you do not have to. It reads what it was given, matches the starting state against the shared routing reference, runs the routed stage skill, re-evaluates after each hop, and stops at the next decision that belongs to a human. It re-implements no stage; capture, refine, plan, plan-review, work, qa, make-pr, resolve-pr, and land keep their own contracts, receipts, and gates. With --auto the same conductor runs unattended, and the route --explain shows is the route the unattended run takes.

Flow replaced /flow-next:guide in 5.0.0. Guide printed a recommendation and refused to run anything; flow --explain prints the same recommendation, and plain flow runs it. There is no alias for guide.

/flow-next:flow <anything>. The argument is any starting point, and the agent decides what it is:

  • nothing at all (“what should I do next”)
  • an idea, or a request for a change
  • a spec or task id
  • a tracker issue id or URL, read through the access the session already has (the sync bridge, an MCP, gh, glab)
  • a branch or a path
  • a pasted bug report or console output
  • a how or why question about the code
  • something slow to speed up, a cleanup that keeps behaviour, or a design fork to settle
  • the live conversation

Flow adds no input adapter and no classifier. It routes on content and context, never on input kind. With no argument it resolves the item from the most recent thing it can see, first match wins, then routes it as if the id had been typed: the item this conversation last touched (the spec capture wrote, the task work closed, the PR make-pr opened), then the spec whose branch_name matches the current branch, then intent in the conversation that no spec captures yet (it asks whether to capture), then the next open spec in .flow by its judgement with an inline pick when several are equally plausible, then a single question about what to work on. So a bare flow right after a capture proceeds to the step capture recommended. On that no-argument reading flow also runs flowctl features status once; when the feature map is due a maintain pass, it prints Also recommended: /flow-next:features - feature map due a maintain pass (<reasons>) after the Next: line. It recommends and never dispatches: see Keep the feature map current.

A request that names no skill reaches flow. On hosts that match skills by description, “this endpoint is slow” or “why does this guard exist” routes through flow without a slash command. Where the host needs the command, /flow-next:flow <the same words> is the spelling; on OpenCode it is /flow-next-flow, on Codex $flow-next-flow.

read the starting point -> match the route matrix -> run the routed skill -> re-evaluate -> repeat

Each hop matches the matrix afresh; there is no fixed conveyor. Advancement is judged on observed state (flowctl show, the PR probe, the receipts), never on a stage’s narration. By default, an attended run from intent ends with the change committed on a local branch and one line: “Say ‘open the PR’ when you want it.” Nothing is pushed until you say so; then flow runs make-pr. Under --auto the run opens the PR itself and ends when it exists. A tiny, local, low-risk change takes the direct route instead: changed, checked the way a user would, and handed back, with no spec and no pull request unless you ask for one. On an existing PR, attended flow offers continuation through merge and asks once unless that item is already authorized. A PR the run just opened gets no landing question. With --until=merge, flow invokes land for the selected item and continues until landing completes or an existing stop condition applies. Flow never fabricates a review, QA, or completion verdict to pass a gate.

Routing takes seconds. Flow asks no branch or readiness question before a build: it takes a sensible default and says so in one line. A single-task spec is built inline in the conversation; workers run only for several tasks or a task sent to a chosen model. Across more than 170 full end-to-end runs, the change comes back in about the time plain Claude Code, or your harness, takes, often less. The optional stages (cross-model review, live QA) then add quality, and flow adds one only where the risk calls for it.

Every stop prints one report:

Flow stopped at: <the human decision, or "PR exists">
Route taken: capture -> work -> make-pr
stage: capture - ran [2026-09-12T09:02:11Z..2026-09-12T09:05:40Z]
stage: refine - skipped(policy: no unresolved product questions)
stage: plan - skipped(policy: no positive plan signal; route recorded no_plan)
stage: work - ran [2026-09-12T09:05:41Z..2026-09-12T09:41:58Z]
stage: qa - skipped(config: pipeline.qa=auto: no UI-observable criteria)
stage: make-pr - ran [2026-09-12T09:41:59Z..2026-09-12T09:42:31Z]
Next: review https://github.com/acme/api/pull/214, then merge or send it back

That report is from a run that reached make-pr: an attended run after you said “open the PR”, or a flow --auto run. Before you ask, an attended run stops after work (and QA, when it runs) with the change committed and Next: telling you to say “open the PR” when you want it.

Route taken carries the hops that executed, plus any inline pick. Every stage flow reaches gets one stage: line: ran [<start>..<end>] with the stage’s start and end times where the run knows them, skipped(<policy|config|empty|error>: <detail>), or failed(<reason>: <detail>). A skipped stage is an event with a reason, never a silent absence; flowctl usage --stages <spec-id> summarizes the task-scoped lines.

Add --until=merge to carry the selected item through land:

/flow-next:flow fn-12 --until=merge
/flow-next:flow --auto fn-12 --until=merge
/flow-next:flow --auto --tick fn-12 --until=merge

Destination and interaction mode are independent. Without the flag, unattended flow keeps its pre-merge stop. Plain attended flow on an existing PR offers landing and asks once unless you have already authorized that item; a PR the same run just opened gets no landing question. Declining or leaving the question unanswered causes no landing mutation. A missing or invalid destination value fails closed before selection or mutation.

Authorization covers the selected spec and PR through retries and waits. It never includes unrelated eligible PRs. A fresh session needs the flag again or current explicit authorization; a historical transcript or receipt alone is insufficient. Revocation stops subsequent mutations.

Flow binds the one pull request of the selected spec and invokes /flow-next:land <PR>, passing the pull request and the current authorization as ordinary arguments. Land owns review resolution, CI repair, the merge gates, and the merge; flow copies none of those steps. One tick performs at most one land invocation; a long-running invocation waits at the driver’s cadence (the caller’s interval, otherwise 30 minutes) and invokes land on the same pull request again. Waiting consumes no pilot strikes. Several plausible targets, a failed probe, conflicting identity, and closed-unmerged targets stop safely, and a missing target is never replaced.

The spec is already closed when landing starts: make-pr commits the close on the pull request branch, and the merge carries it to the base. Land writes nothing to the repository after the merge and needs no checkout of yours. An open pull request routes to landing even when its spec is closed or no longer ready. A confirmed merge ends the run, including when the tracker touchpoint failed; the failure is reported with the merge commit. Releases stay outside land.

--explain: the recommendation without the run

Section titled “--explain: the recommendation without the run”

/flow-next:flow --explain <situation> prints the route and stops. It writes nothing under .flow/ and dispatches nothing, so an explain run leaves the repository byte-identical.

Next: work fn-12 through Flow-Next without task decomposition
Route: work --no-plan
Signal: ready, cohesive spec; a capable coding agent can own the whole acceptance contract
Skip/narrow: plan only on a positive signal (asked, separate owners, staged PRs)
Skip kind: signal absent
Why not the alternatives: no unresolved product questions, so refine adds nothing; no design risk named, so plan-review is optional

Routing never asks the optional Jev judge, so a run with a TypeSafe key and one without take the same route. Code decides the lifecycle rows from flowctl show, the PR probe, and the task counts, and prints Route: <route> (code); the host decides the rest from the route matrix and prints Route: <route> (host). A run without a key prints judge: off once and skips the calls that cannot answer. Optional Jev judgments lists the narrow questions a key still speeds up.

Drop --explain and the same words run the route. How Flow chooses walks the situations --explain answers. Under --auto, --explain prints the unattended classification instead; see the unattended shape.

Picks are inline; only run-ending decisions stop it

Section titled “Picks are inline; only run-ending decisions stop it”

A pick among options a stage produced is asked inline and the run continues: prospect’s ranked candidates, a chart briefing’s capture-or-split question, refine’s choices when it hands back. The answer lands on the Route taken line as prospect [picked: <candidate>]. At most one question per hop.

The default attended run stops when the change is committed and ready for a pull request, when the PR you asked for exists, or when a landing offer on an existing PR needs an answer. The default --auto run stops when the PR exists. An authorized merge-destination run stops when landing completes. Either run stops when a stage surfaces a decision that ends the run (a NEEDS_HUMAN, a review verdict that needs a person, a product question no stage framed as options), or a blocking question this hop must ask is not a pick.

Before any “which approach” question, flow classifies the fork. An answer that is observable (behaviour, output, timing, layout, a failing case, a measurement) is settled by running something, and the report names what ran. Only a product or preference call becomes a question.

When the fork has more than one viable answer, the prototype builds the variants behind one switcher: a toggle, flag or keypress that swaps between them, each variant labelled, so you compare them in one place under the same data. A search box that could refresh on every keystroke or after a pause gets both behind a toggle; typing the same query shows 180 ms per keystroke update with visible stutter against a pause that feels immediate. A fork with only one viable answer gets a single prototype. When the options are still open, flow first gathers prior art, meaning how this repository and comparable products or libraries solve the problem, and lets you pick a direction before it builds. Picking a direction is a preference call, so under --auto it stops with NEEDS_HUMAN and the references. A prototype is evidence: it lives in scratch space and never ships.

Files under the flow skill’s references/ carry the rules, one per rule:

FileRule
route-matrix.mdWhich starting state routes where, its positive signal, its safe skip, and the skip kind
route-matrix-more.mdThe rarer starting states (strategy, prospect, and similar), read from route-matrix.md
spec-count.mdWhether one intent is one spec or several (the 8-criteria tripwire, independence over size)
plan-vs-no-plan.mdDirect execution is the default; plan needs a positive signal; refine needs a named open decision
gate-selection.mdWhich review, QA, and completion gate applies, and from which config key or flag
prototype-before-ask.mdAn observable fork is settled by an experiment, never by a question; competing variants sit behind one switcher
tail.mdExisting-PR continuation and the scoped landing handoff
no-argument.mdWhat a bare flow routes, first match wins
defect-intake.mdFinding a reported defect’s feature on the map before reproducing it
explain.mdWhat --explain prints
qa-stage.md, backlog-mode.mdThe unattended QA freshness probe and backlog mode, read only when active

The worked variants live on How Flow chooses.

Flow’s route step, flow --explain, flow --auto, capture’s closer, plan’s next-steps menu, and work’s zero-task ask read the same files, and each reads only the file its current step needs. The recommendation you read and the route that runs are the same rule, attended or unattended. Router staleness is a defect. Adding or removing a skill updates the matrix in the same change, and a test asserts every pointer names a file that exists.

The matrix rows carry the boundary that keeps neighbouring rows apart:

  • Bug or defect. The route checks for prior fixes, reproduces and diagnoses, bisects when a known-good revision exists, then fixes and proves on base and head (the four steps). Refine is the wrong instrument for a defect. Capture only when the diagnosis conversation itself carries decisions worth locking down. With a feature map, a report that does not say where the problem is starts from the mapped feature (bug intake).
  • Structured brief. A brief with its business and technical choices resolved goes to capture and skips refine unless a named product or authority decision is open.
  • Refine. A spec routes to refine only when the open decision can be named. Reopen discovery as chart only when the answers show the effort itself is not yet specifiable.
  • Refactoring. New behaviour named anywhere makes it a feature with cleanup inside, and it routes to capture or work.
  • Performance. No nameable metric or surface routes to the investigation row first. A fix motivated by reading source instead of a measurement is not evidence.
  • Hill climb. One expected fix is the performance row. The target is never relaxed to meet it.
  • Investigation. When the answer is a prerequisite for a change already asked for, flow routes the change and lets its stage read.
  • Prototype. No decision means no prototype. Only a product or preference call no experiment can settle becomes a question.
  • Docs or chore. The flowctl triage-skip receipt, never the label, is what justifies the review skip.

For a ready spec with no tasks and no recorded route, the route is /flow-next:work <spec-id> --no-plan. Plan is chosen only on one of three positive signals:

  1. You asked for a plan.
  2. Separate human owners will implement.
  3. Delivery is staged across several PRs.

Risk, size, and file count never trigger plan on their own. Design risk routes to /flow-next:plan-review, which reviews a spec with zero tasks. Unresolved product or authority choices route to /flow-next:refine.

When to refine. Flow routes to refine only when it can name at least one open decision that would change what gets built and that only you can make. It skips refine when the acceptance criteria state the intended behaviour and what remains is how; when the touched area has established patterns; when the only gaps are technical detail, performance, or edge cases that implementation, review, and QA will surface; when the only uncertainty is criteria capture inferred itself; and for a defect, a structural cleanup, or a request to move a measured number. The technical pass needs a named technical fork that is costly to reverse and that the code does not answer: a data model or migration, a public contract, a security boundary. The recommendation names the decision: Recommended next: /flow-next:refine <spec-id> - <the named open decision>.

Flow records the route (flowctl spec set-no-plan for direct, spec clear-no-plan for plan) before any stage runs, and only after the explain stop. Under flow, capture applies the same rule itself and sets no_plan when it resolves to direct. User-invoked capture keeps the explicit --no-plan opt-in and never sets the field on its own judgment. The direct route keeps implementation review per review.backend, acceptance coverage from the single implicit owner task, the single-task completion-review skip, and QA per pipeline.qa.

Read first when the spec names something new

Section titled “Read first when the spec names something new”

The ready-spec row carries one more signal: the spec names a library or API the repo does not already use. On that signal flow runs /flow-next:refine <spec-id> --scope=research before work on either route. An existing ## Resolved via Research section, or plan findings, satisfy the signal for the libraries and APIs they cover. A newly named library or API triggers a delta pass for that one. --force reruns the whole pass and replaces the section. The pass never runs by default.

A pasted bug report, console output, or screenshot routes to the defect row: check for prior fixes, reproduce and diagnose, bisect when a known-good revision exists, then fix and prove the fix on base and head (the four steps). When the repo has a feature map at .flow/features/ and the report does not say where the problem is (a screenshot with no page title, “this thing in my list”, a symptom with no location), flow uses the map to find the broken place before driving the reproduction. A report that names the page, control or command goes there directly, as before, and falls back to the map only if that lookup fails:

  1. Resolve the report to one feature. Flow reads the map’s index and matches the report’s wording against each feature’s description and sub-features. A screenshot is matched on what it shows (headings, table columns, labels, controls), so a screenshot with no page title or URL can still land on the right feature. The pick is a **Surface:** identifier plus a sub-feature ID.
  2. Drive along the file. Flow hands drive that feature file, so the reproduction follows its route, preconditions and gotchas instead of discovering them.

When nothing matches, flow reproduces by live discovery, as before; a missing feature is a hint for the next maintain pass. When several features plausibly match, flow names them and tries the most specific first. When the mapped route no longer matches the app, flow files a drift note and continues live. Flow never edits the map. Without a map, the only added cost is one existence check.

This gate applies to bug intake only. The stages that drive the running app later on (the fix’s live proof on base and head, a performance baseline and its post-change measurement, QA) always read the map when there is one: Which routes read the map.

The map says how to reach a feature, never what the bug is: the report and the reproduction stay the evidence for the defect.

Why the gate, measured: on a fixture app with reported defects, including screenshots with no page title, the map ended long searches on untitled screenshots (38 turns to about 12), with cost and wall time flat. Expect faster, more predictable where-is-it reports.

Flow reads pipeline.qa at all-tasks-done through gate-selection.md, attended and under --auto alike. on runs one live pass on every spec. auto runs it only when the spec’s acceptance describes UI behaviour on a drivable surface and a target can be started (a documented start command, a deploy URL, or .flow/features/); otherwise the stage records skipped(config: pipeline.qa=auto: <no UI-observable criteria | no drivable surface | no startable target>) and the route advances to make-pr. Since 5.1.0 auto takes effect unattended too; before that an unattended tick read only the literal on. Enable it with flowctl config set pipeline.qa auto. QA never hard-blocks. NEEDS_WORK and BLOCKED advance to make-pr, and their findings become open items on a draft PR.

/flow-next:flow --auto [<spec-id>] selects one ready spec and drives it hop after hop until a terminal, so one invocation carries a spec from ready to a pull request without a scheduled loop. The ready flag authorizes build selection; --until=merge separately authorizes landing the selected item. --auto accepts no intent, no path, no branch, and no free text; a spec id scope-locks the run, and without one the run walks the open, ready specs in stable id order and skips any whose depends_on_epics are unsatisfied. A dependency is satisfied when it is done, or when it is the spec’s chain parent: open with every task done and its branch on origin, as flowctl spec chain reports. Such a spec is dispatched chained, its work branch forked from the parent’s tip and its PR opened against the parent’s branch, and the verdict reason begins chained on <parent-id>;. A spec with two open parents, or a sibling already chained on the same parent, parks with that reason and no strike. Details: chains and stacks. An unknown flag warns to stderr and is ignored, except a missing or invalid --until value, which fails closed.

Each hop classifies from the same routing references the attended conductor reads. route-matrix.md supplies the spec-state row, plan-vs-no-plan.md decides a ready zero-task spec with no recorded route (the run records the route with spec set-no-plan or spec clear-no-plan before creating any task and prints why), and gate-selection.md selects the review, QA, and completion-review gates. Pilot’s private stage table is gone. What the reference cannot carry lives in auto.md, read only when --auto is parsed: selection order, the collision and re-bless checks, the all-done PR probe, the branch matrix, the evidence echo, and the strikes ledger.

The hop itself is classify, dispatch exactly one stage with mode:autonomous, verify from observed state (flowctl status fields, task counts, the receipt, the gh-confirmed PR URL), record the hop (receipts, the evidence echo, the ledger write), and re-classify. The build stages are plan, plan-review, work, qa, and make-pr, with qa only when gate-selection selected it. By default, ADVANCED continues until make-pr creates the PR. With --until=merge, the run continues to land and handles its progress and waits through the scoped handoff.

Flags under --auto:

  • --tick runs exactly one hop and stops, which is what a pilot tick was. It is the shape for hosts without stable long sessions and the replacement for the /flow-next:pilot command, removed in 6.1.0 (argument mapping).
  • --explain (or --dry-run, its one-release alias) prints the selected spec, the classified stage with its routing row and gate section, the review backend, task counts, the consulted status fields, the resolved zero-task route as would-record, the PR probe result, skipped candidates, would-clear ledger entries, and chain=<off|on>. It writes nothing, checks out nothing, and dispatches nothing. In ready mode its terminal line is PILOT_VERDICT=NO_WORK spec=<id> stage=<stage> reason="dry-run: classification only, nothing dispatched". In backlog mode an explain run skips the tracker-sync dispatches (reconcile, list-open, list-comments, list-relations), selects from the flow-side ready --all facts alone, triages the picked item, and ends with PILOT_VERDICT=TRIAGED spec=<id> stage=triage reason="dry-run: classified <class>, nothing dispatched or parked".
  • --backlog (or pilot.autonomy=backlog) widens selection to the whole open backlog. It never authors a spec or sets ready, and still selects one item per run. Backlog mode alone grants no merge authority; landing requires --until=merge for the selected item. In long-horizon mode a backlog run drives its one selected item to a terminal and stops; the next invocation selects the next item. The workflow is on Backlog mode.
  • --review=<backend> passes through unchanged to the plan, plan-review, and work dispatches. --research=<grep|rp> and --depth=<level> reach only the plan dispatch; qa and make-pr take no passthrough. Defaults are research=grep, depth=short, and the backend from flowctl review-backend.

There is no branch flag (the run resolves the branch from the spec’s branch_name) and no --no-plan flag (the route is recorded on the spec, as above).

The terminal line. Every run ends with one PILOT_VERDICT line, the last line of output, in the grammar drivers already parse. A long-horizon run names every dispatched stage in order joined by + and carries the last hop’s verdict; a --tick run names one stage.

PILOT_VERDICT=ADVANCED spec=fn-12 stage=work+qa+make-pr reason="make-pr: open PR https://github.com/acme/api/pull/214"

Without a merge destination, the terminals are the PR exists (ADVANCED with make-pr), DEFERRED_TO_LAND (an all-done spec with an open PR that land owns), ASKED (backlog mode), BLOCKED (the two-strike ledger, flowctl pilot strikes list|clear), NEEDS_HUMAN, and NO_WORK. The full verdict table is on the flow —auto page.

Landing evidence. With --until=merge, read the original LAND_VERDICT and the selected PR’s identity, state, and merge commit alongside PILOT_VERDICT:

Observed land resultFlow outcome
MERGED, with GitHub confirming the merged state and a merge commitADVANCED; destination complete
FIXING_CI / RESOLVING / AWAITING_REVIEW with observed workADVANCED; progress, not completion
Those same states without progressDEFERRED_TO_LAND when the invocation stops; long runs may wait and continue
BLOCKED / NEEDS_HUMANPreserve the blocker and stop
NO_WORK, a failed identity probe, or a MERGED verdict GitHub does not confirmNEEDS_HUMAN; never treat it as completion

A landing reason carries the original land verdict, PR identity, and waiting, progress, or merged=<sha>. Consumers must not infer a completed merge from ADVANCED alone. Landing outcomes do not spend pilot strikes.

Resume from disk. Every hop ends with the same receipts, evidence echo, and ledger write a pilot tick ended with. A run that dies mid-way (crash, kill, session limit) leaves committed receipts, a ledger entry, and a branch; the next invocation classifies from disk, and nothing is resumed from transcript. The dirty-tree refusal (dirty outside .flow/) runs at run start and again after every hop.

Nobody is waiting. An unattended run never asks. It decides from evidence, stops only for a call only a human can make or an irreversible action, and prints a Decisions: list before its terminal line: each default it chose, finding it declined and review it skipped, with the evidence behind it. The same list goes into the pull request body. A human-only call that blocks nothing else (refreshing a frozen fixture, a requirement only CI can prove) becomes an open item on a draft pull request instead of a stop. Under --until=merge the run also holds the merge, so it makes such a call itself when it is reversible, inside the spec and backed by evidence (updating a snapshot or golden file the change legitimately altered), and records it in the Decisions list. A call that is irreversible, a product choice the spec does not settle, or one that makes merging unsafe stops the run NEEDS_HUMAN before the merge, and land never marks an unattended draft ready. Review loops fix and a scoped re-review until SHIP before the handoff, as Impl Review describes.

Default boundary. Without --until=merge, --auto stops before landing, at a pull request that opens ready unless make-pr found open items. With the flag, it consumes land as one scoped stage; it never copies land’s gates or dispatches a recursive driver. pipeline.chainStages was removed in 7.0.0: under --tick, make-pr runs on the next tick, and a long-horizon run already runs QA and make-pr as consecutive hops. flowctl ignores a leftover key with a one-line note. What a long-horizon run removes at each stage boundary is the repeated run initialisation (the skill re-read and hard guards), the config capture, the selection pass, and the external loop interval. Classification and branch resolution still run before every hop.

The run’s rails, the consent boundary, and readiness as the control plane are on Flow —auto - the build loop. Driver snippets per host are on Driving a loop.

Attended flow and flow --auto are two drivers and are never nested. Without --auto, flow is attended and stops before any read under any autonomy marker (FLOW_AUTONOMOUS, AUTONOMOUS=1, a mode:autonomous token):

NEEDS_HUMAN: /flow-next:flow is attended - run /flow-next:flow --auto for unattended runs

--auto has no marker refusal, because it sets FLOW_AUTONOMOUS and mode:autonomous for the stages it dispatches. The sole composition exception is flow invoking land for its authorized item; recursive drivers remain forbidden. The default DEFERRED_TO_LAND behavior is preserved. A /loop 30m /flow-next:pilot recipe needs /flow-next:flow --auto --tick since 6.1.0, which removed the pilot command. Where flow is heading is on Flow and the road ahead.

Terminal window
/flow-next:flow # the next step for the most recent item, or ask
/flow-next:flow fn-12 # route a spec from its current state
/flow-next:flow "why does the parser guard against empty headers"
/flow-next:flow --explain I have a spec with an open PR, what next?
/flow-next:flow fn-12 --review=codex # forwarded unchanged to every stage it dispatches
/flow-next:flow --auto fn-12 # unattended: hop from ready to a PR, then stop
/flow-next:flow --auto fn-12 --until=merge # continue through scoped, gated landing
/flow-next:flow --auto --tick # unattended: one hop on the next ready spec, then stop
FlagUse
--explainPrint the recommendation shape and stop. No .flow/ write, no dispatch. Under --auto, print the unattended classification and end with the NO_WORK dry-run verdict (ready mode) or the TRIAGED diagnostic line (backlog mode)
--review=<backend>Passed through unchanged to every stage flow dispatches
--auto [<spec-id>]Unattended: select one ready spec (or the given one) and hop until a terminal; ends with one PILOT_VERDICT line
--until=mergeAuthorize landing the selected item through land, independently of --auto
--tickWith --auto: one hop, including at most one landing tick, then stop
--backlogWith --auto: widen selection to the whole open backlog; ASKED parks an item behind an async question
--research=<grep|rp>With --auto: passed to the plan dispatch only; default grep
--depth=<level>With --auto: passed to the plan dispatch only; default short

Every copy-pasteable command flow prints uses the spelling your host invokes: the flat /flow-next-<name> form on an OpenCode install, otherwise as written here.

A pre-registered non-inferiority study (September 2026, 90 draws on one frontier model at medium effort) compared the retired guide matrix against the shared routing reference read through flow’s pointers on 27 fixture situations. On the 9 verdict-bearing discriminating items across 3 draws each, the retired matrix scored 24/27 and flow’s reference 27/27, so the reference is non-inferior under the registered rule; the candidate read the required reference file on every one of its 45 draws. Superiority was never the claim and is not confirmed. See Evidence and evals.