Skip to content

Changelog

Flow-Next changed shape with 5.0.0. One command reads whatever you have and picks the route, and the same route runs unattended with --auto. Arriving from 4.x? Read the 5.0.0 entry first, or start with Flow, then come back to the items below.

7.0.0 - Roadrunner: faster on the actual work

Section titled “7.0.0 - Roadrunner: faster on the actual work”

Flow-Next now hands you the working change in about the time plain Claude Code, or your harness, takes, often less, and large features in about half the time. The result is better than the plain agent’s even before any review, and its optional cross-model review and live QA widen the gap to up to 25% better outcomes, especially on large and long-horizon work. This is a breaking release: Ralph and the HTML render lenses are gone, and the upgrade steps are below.

Detail

The work is fast now, and the quality stages are where the extra time goes. /flow-next:flow picks the route in seconds from what you give it. A small local fix goes straight to the change, with no spec. A spec with one task is built right in your conversation. You get the result back first; a reviewer from another model family then checks it in the background, and after any fixes the same reviewer looks only at what changed. Risky changes (persisted or shared state, concurrency, security, data layout or migrations, multi-file features) get three reviewers; small ones get one; a change to output, wording or display gets none, with the reason recorded.

Unattended runs finish on their own. flow --auto never stops to ask. It keeps fixing until the reviewer signs off (the author may decline hardening, scope creep and problems that predate the change, and when only those are left, all below Major, the loop ends with each disagreement in the pull request; a broken stated requirement is always fixed), writes every decision it made on your behalf (defaults it picked, findings it declined, reviews it skipped) into the pull request, and puts a call only you can make into the pull request as an open item instead of stopping. The pull request opens ready for review, or as a draft when an open item is left for you. In our runs with --until=merge, a large feature and a hard bug both went from spec to a merged pull request with nobody watching and no stops, the large feature in about half the time plain Claude Code took.

How it got here. Flow-Next started on December 26, 2025, as a plugin called flow: a plan command, a few scouts and a quality auditor. Claude Opus 4.5 was the model of the day, and models have come a long way since. It was the first to run autonomous cross-model review loops, where a model from another family argues with your agent’s work until it holds up, and one of the first to interview you before building. The goal was always to let R&D teams work together on big, messy codebases and get better work out of their agents. Over this year that meant hundreds of features for attended and unattended runs, the building blocks for code factories, and support for six hosts: Claude Code, Codex, Factory Droid, Cursor, Grok Build and OpenCode. The output was consistently better than a plain agent’s, but all those features made Flow-Next slower and hungrier for tokens than I liked. So for 7.0 I went back through the whole plugin and rebuilt it for speed, measuring every change against plain Claude Code before keeping it. I’m going to keep improving it, in what it does and in how fast it does it.

Benchmarks. More than 170 full end-to-end runs against plain Claude Code on the same model, each case run several times, across a simple bug, a hard bug, a small feature and a large feature, attended and unattended, plus a held-out large feature from a repository and stack the tuning never touched. Hidden tests the agent never sees check every result, and a blind judge scores the handoff.

Large feature

up to 1.9×faster than the default harness

+25%better result, blind-judged

Three cross-model reviewers

Plain Claude Code shipped a data-integrity bug in every run. Flow-Next's reviewers caught it every time, before the pull request.

Held-out large feature (Rust)

1.2×faster than the default harness

+48%better result, blind-judged

Three reviewers, then the repo's full test suite

Four real bugs fixed, including a race condition. Flow-Next passed every hidden test; plain Claude Code failed one run in three.

Hard bug

about 2×faster than the default harness

+9%better result, blind-judged

Three cross-model reviewers

Found the real cause and fixed it there, instead of loosening the flaky test.

Simple bug

1.2×faster than the default harness

+10%better result, blind-judged

No review: a small, local fix

Better even without review: a failing test first, a fix at the cause, then a check that it works for the user.

Small feature

1.1–1.3×faster than the default harness

+8%better result, blind-judged

One cross-model reviewer

The reviewer caught a setup check the new option broke. Fixed before handoff.

Against the default harness (plain Claude Code, or yours) on the same model. Speed is time to the working change, before any review or QA. More than 170 full end-to-end runs, each case run several times. Hidden tests the agent never sees check every result; a blind judge scores each handoff.

What changes when you upgrade. Ralph is removed, and /flow-next:flow --auto is the one unattended mode. If you still run Ralph, pin flow-next 6.7.x. Otherwise delete scripts/ralph/ and any ralph-guard hook entries from your project settings, and use flow --auto for unattended runs. The HTML render lenses are removed too: a config that still sets artifacts.html.enabled keeps working and prints a one-line note, and for a visual view of a spec, a plan or a diff you run /flow-next:visual or ask the agent for an HTML page. Your old spec.html and pr.html files stay where they are, and you can delete them. The hidden pilot and interview skill stubs are gone; use /flow-next:flow --auto and /flow-next:refine. pipeline.chainStages is removed: flowctl ignores the key with a one-line note, and under --tick make-pr runs on the next tick, so delete it from .flow/config.json. An unattended --until=merge now waits 10 minutes after the last push before merging (land.patienceMinutes, was 30); set the key if your review bots are slower. The Copilot reviewer now runs read-only, without write and shell tools. The Codex installer no longer adds hooks = true to ~/.codex/config.toml (only Ralph needed it); a hooks line already there stays. You don’t need to re-run setup.

What you keep. Specs and their state in your repository, cross-model review from the backend you configured, live QA under pipeline.qa, the pull request briefing with its evidence, land, and every receipt. The optional Jev judge still offers its fork hint, memory reordering and task tier; it no longer picks the route or decides whether QA runs, so a run with a key and one without take the same route. Keyed calls now also work behind an HTTPS proxy, where before they quietly fell back to no judge.

Under the hood. One shared set of working rules covers every stage: the smallest change the evidence justifies (including the sibling case the same cause breaks), focused tests for what changed and the full suite only when your repository or you ask, a check you name always run and waited for, and a handoff that marks each claim measured, inferred or a guess. Refine (interview) asks only what would change the build, in one question pass. flowctl judge --preset route is code-only and sends no request; the qa-gate preset is retired.

7.1.1 - Clean upgrades and every memory lesson kept

Section titled “7.1.1 - Clean upgrades and every memory lesson kept”

Codex users: re-run ./scripts/install-codex.sh once. Upgrading from an older checkout then leaves exactly the skills and commands this release ships, memory-migrate keeps every lesson from very old memory files, and make-pr accepts a Bitbucket pull request link.

Detail

Clean upgrades. When a release removes a skill, git pull can leave its folder behind if untracked files such as __pycache__ are still in it. The Codex, OpenCode and Cursor installers treated that folder as a skill, and the Codex installer copied it over the real installed skill. A folder now counts as a skill only when it has a SKILL.md. The Codex installer also retires the interview and pilot prompts 7.0 removed, which 7.1’s cleanup missed. Everything it retires goes to ~/.codex/.flow-next-retired/, nothing is deleted, and prompts you wrote yourself are left alone.

Every memory lesson kept. Memory files from flow-next before 0.33 put each lesson under a ## <date> manual [<type>] header with nothing between them. Memory migrate read each of those files as one entry, so it merged the lessons and made an empty entry for a file holding only its header. Each lesson is now its own entry. Thanks @TechupBusiness (#509).

Bitbucket pull request links. A FLOW_PR_CREATE_CMD that prints a …/pull-requests/<n> link no longer reads as a failed create. Thanks @CWayman (#504).

At the keyboard, Flow hands the change back committed on a local branch and opens the pull request when you say so, and a run you hand the merge stops before merging on any call that is yours. A bare flowctl typed outside a skill now needs its path, and RepoPrompt and the review.backend setting leave in 8.0.0.

Detail

The pull request opens when you ask. An attended build used to push and open a pull request at the end of the run. Now it commits on a local branch and ends with one line: say “open the PR” when you want it. Nothing is pushed until you do. Flow no longer asks “land it now?” about a pull request it just opened; a later /flow-next:flow on that open pull request still offers landing once. flow --auto is unchanged and still stops at an open pull request. See Flow.

A run you hand the merge decides what it can. With --until=merge the run holds the merge, so it makes a call itself when that call is reversible, inside the spec and backed by evidence (updating a snapshot your change legitimately altered, say), and records it in the pull request’s Decisions list. An irreversible step, a product choice the spec leaves open, or anything that makes merging unsafe stops the run with NEEDS_HUMAN before the merge. Land no longer marks a draft ready on an unattended merge; it stops and names the open items. You saying “land it” still merges a draft. Without a merge to hold, a reviewer’s question that blocks nothing else no longer ends an unattended run empty-handed: the run fixes everything else, opens a draft pull request and lists the question as an open item.

Land takes more pull requests through. A pull request opened without a spec now lands through the same gates as any other, where land used to stop with NO_WORK. When resolve-pr hands a review thread to a person, land stops once with NEEDS_HUMAN instead of retrying the same thread every 30 minutes. Repairs push again: land fixes review threads and red CI in a detached worktree at the pull request’s head and pushes to its branch by name. See Land.

Fewer questions before the build. Work and plan no longer stop to ask setup questions when no reviewer is configured. Plan uses its default depth, review is off, and the handoff says so once; set a backend in setup, with review.backend, or in the prompt. A check you ask for by name, such as the full suite or a command, finishes before the reply instead of running on in the background. Running something to settle a question stays read-only or throwaway on every route, and anything that needs live state, credentials, the network or a destructive command becomes a question when you are there and a human call when you are not. Audit’s autofix commits to a local docs/audit-memory-<date> branch and names it, where it used to open a pull request on its own.

Less for the agent to read. Instructions that only matter in some runs (tracker steps, unattended rules, stacked pull requests, the no-plan route, feature-map writes and more) now load only when the run needs them, and each skill reads the working rules once per run. A typical attended run reads about a quarter less instruction text than in 7.0, roughly 19,000 words down to 14,600.

What changes when you upgrade. Attended runs stop opening pull requests unless you ask. The plugin’s top-level bin/flowctl is gone, so a bare flowctl typed outside a skill needs the path to the plugin install’s scripts/flowctl (where it lives); skills resolve it themselves, and nothing needs re-running. The same change lets claude.ai, Cowork and organization sync install Flow-Next, since they refuse a plugin with a top-level bin/ directory. Thanks @sn-furali (#506). Codex, OpenCode and Cursor-script installs pick all of this up when you re-run their installer, as usual, and the Codex installer now moves skills and prompts a release removed into ~/.codex/.flow-next-retired/, leaving your own files alone. You don’t need to re-run setup.

Deprecated, leaving in 8.0.0. RepoPrompt support goes: the rp review backend and the Export Context skill. Projects already on rp keep working until then and get one notice per run. To move now, pick another backend with flowctl config set review.backend codex (or host, claude, copilot, cursor). The review.backend setting goes too. From 8.0.0 the reviewer is chosen in the model-routing block of your CLAUDE.md or AGENTS.md, the one setup proposes, the same way implementers and scouts already are, and that block can route review to another model family. Until then review.backend works exactly as it does today, and a reviewer named in the prompt still wins. See Review backends.

Fixed. Work could check the review backend from the skill’s own folder, read “nothing configured” and skip review on a repository that had a reviewer; it now resolves the backend per task from the repository root. Reviewer instructions no longer contradict themselves: a pre-existing problem that blocks shipping counts against the change, an unaddressed requirement blocks at any severity, and completion review grades the finished build on its own scale. A one-task build whose review was skipped skips completion review too. An attended flow run commits QA’s verdict on the spec branch. A reviewer’s NEEDS_HUMAN stops the run on every backend. flowctl show keeps a spec’s tracker link (thanks @sn-furali, #484, #483). A second GitHub blocker is queued for you instead of stopping tracker sync (thanks @TechupBusiness, #502). Tracker pulls stop duplicating comments, and a spec-only pull request carrying old render files no longer counts as shipped work (thanks @sn-furali, #501). Worktree Kit works from inside a worktree (thanks @sn-furali, #469). A formatter’s blank line no longer marks your setup block as customized. Codex setup keeps agent files you edited and asks once about any that differ. Memory migration keeps the lessons it did not migrate. Prospect promotes the idea you picked. Prime asks before creating CI or devcontainer files and never runs a full suite to assess. Drive closes only apps it opened. The repository changelog lists the smaller fixes.

What you keep. flow --auto still opens the pull request itself, ready for review or as a draft when an open item is left for you. A merge you authorize goes through land’s gates as before. Review by risk, live QA under pipeline.qa, receipts and the pull request briefing are unchanged.

Under the hood. Land’s NO_WORK now means only a closed, unmerged pull request. It has two new NEEDS_HUMAN stops: a review thread resolve-pr handed to a person, and a draft pull request under a merge that flow authorized without a person in the session.

Codex reviews and the Codex agents run on GPT-6.1 Sol, and Claude reviews can fall back to Opus 5.5 and Sonnet 5.5. Re-run ./scripts/install-codex.sh to update your Codex agents.

A Codex review now starts on gpt-6.1-sol at high effort and steps down to gpt-6-astra, then gpt-6-sol, when your account cannot serve it. The generated Codex scouts, planning helpers and quality auditor move to gpt-6.1-sol at their usual effort. A Claude review steps down from Fable 5.1 through Opus 5.5, Opus 5, Sonnet 5.5 and Sonnet 5; a Claude Code older than 2.1.284 skips the 5.5 models. A model you name explicitly still wins, and Copilot and Cursor reviews are unchanged.

6.6.0 - Scouts on Sonnet 5.5, sibling repos gated

Section titled “6.6.0 - Scouts on Sonnet 5.5, sibling repos gated”

Plan and prime run their scouts on Sonnet 5.5, Anthropic’s lower-cost complement to Opus 5.5 that scores close to it on coding benchmarks. Teams that keep product code in sibling repos get full gates when a sibling’s code changes, and a feature-map maintain pass ships on repos with ticket-key naming rules or a non-GitHub host. Update Claude Code to 2.1.284 or later so the scouts get Sonnet 5.5.

Detail

Scouts are the read-only helpers that plan, prime and the other planning steps send out to read your repo, docs and memory. They now run on Sonnet 5.5 instead of Opus. Anthropic reports Sonnet 5.5 more than 30% faster than Sonnet 5 and within a few points of Opus 5.5 on its published coding and knowledge-work benchmarks (for example 55.5% against 57.8% on CursorBench 4.0), on a cheaper model, so scout-heavy steps cost less and return sooner. The quality auditor stays on Opus, because a bug it misses goes unnoticed. See the model tiers.

If you keep .flow/ in a planning repo and your code in sibling clones next to it, work now checks the siblings too. It records the base of each sibling the spec changes, as named in your project instructions or the spec, and classifies every repo’s diff. A spec that only touched specs in the planning repo but changed TypeScript in two siblings used to run lint-only gates; now the docs-only shortcut applies only when every repo’s diff is docs-only, and a sibling work cannot read runs its full gates. The feature-map update reads the sibling diffs, and flowctl features status --repo <path> counts sibling commits toward a feature’s age, so the map comes due when the product changes. See work and code in sibling repos. Thanks to @CWayman for the report.

A feature-map maintain pass can now finish on repos with naming rules or a host other than GitHub. It reads your branch and commit naming rules before any proof work and asks then for a value only you have, such as a Jira key, so a pass never proves every route and then fails at the commit hook. It opens the PR through the same FLOW_PR_CREATE_CMD make-pr uses, and when no create command reaches your host it stops after the push and names the branch. You can also ask it to only commit, or to leave the edits uncommitted. When the commit, push or PR create fails after the proofs, the proven edits stay on the branch or in your working tree instead of being discarded. See shipping a maintain pass. Thanks to @CWayman for the report.

make-pr now opens the pull request when the agent’s shell is zsh, the macOS default, where the default create command used to fail with command not found: gh pr create.

What changes when you upgrade. Update Claude Code to 2.1.284 or later: older versions resolve the scouts’ sonnet alias to Sonnet 5. To keep scouts on another model, name it on the fast scout and thinking scout lines of your routing block. You don’t need to re-run setup.

What you keep. A single-repo workspace gates and ages the map exactly as before, maintain’s default branch and commit names are unchanged when your repo states no rules, the worker and PR comment resolver still follow your session model, and the generated Codex agents are unchanged.

6.5.0 - Fewer questions before the build, the same checks after

Section titled “6.5.0 - Fewer questions before the build, the same checks after”

Refine asks only what would change what gets built, often nothing in a codebase whose patterns already answer it, and flow sends you there only when such a decision is open. Hill climbs start from the target you state, and unattended bug fixes stop less often. Scripts that call flowctl scope or flowctl done --resolved-feature need updating.

Detail

Refine is a shorter conversation. It asks a question only when a wrong guess would build the wrong thing, nothing but you can settle it, and the call is yours. Anything the code, the docs, a quick experiment or the implementation itself can settle, the agent settles, records, or leaves to the build. When nothing like that is open, refine says the spec is clear enough to build, and that is a good result. Your answers keep the precision you gave them. “Performance matters here” stays in Decision Context as guidance and never comes back as “must load in 20 ms”; a number becomes an acceptance criterion only when you said the number.

It is also one interview now. You no longer choose between business, technical or both before anything is asked. Name an audience when you want one, with --scope=qa, --scope=security, --biz, --tech, or plain words such as “run a business interview”, and the questions focus on that audience’s decisions. The read-back names every section the session changed, so a product owner and a tech lead refining the same spec in separate sessions both see what moved. See refine.

Flow, capture and plan now recommend refine only when they can name an open decision that would change the build and that only you can make (when to refine). An established pattern, a performance detail, an edge case the build and QA will surface, or a criterion capture inferred is no longer a reason to stop and interview. A structured brief goes straight to capture, and prospect points its ideas at /flow-next:flow so the router picks the next step.

A hill climb starts from the target you stated. The agent picks how to measure it, proves and freezes the benchmark, and writes its setup at the top of the attempt ledger before the first attempt, where you review it in the pull request. It stops to ask only when the spec has no target (hill climb).

An unattended bug fix no longer stops because other pull requests touch the same files. It notes them and carries on, and still stops when one of them already fixes the bug or someone owns a fix in progress. One reproduction that fires reliably is enough. Live-app stages read the feature map directly instead of passing a recorded match along.

What changes when you upgrade. Scripts that call flowctl scope (resolve, bank, write-policy) or pass flowctl done --resolved-feature need updating, because both are removed. A repo-root SPEC.md copied from an older template keeps working, with its scope comments ignored; re-copy the bundled template to pick up the cleaner version. You don’t need to re-run setup. The details are on compatibility.

What you keep. Implementation review, QA, the completion gate, receipts and the pull request evidence are unchanged. --scope=research still asks nothing and has the research scouts write what the spec depends on. Old specs load as they are.

6.4.0 - Move a number with measured attempts

Section titled “6.4.0 - Move a number with measured attempts”

Ask flow to move one number toward a target and it runs a measured loop, keeping only the changes that beat the noise; on a fixture CLI, five attempts cut --version cold start from 124 ms to 8 ms. Scouts now run on Opus, which costs more per plan or prime run.

Detail

You write the metric, the target, a minimum number of attempts and a budget in a ## Hill-climb pre-registration section of the spec, along with the command that measures it and the tests that must stay green. Work first proves the benchmark can tell a known-worse and a known-better variant from the baseline, then freezes it. Each attempt tries one idea. It is kept, as its own commit, only when it beats the best so far by more than the measured noise, a second measurement agrees and your tests pass; anything else is reverted before the next attempt. The run stops when the target is met and the minimum attempts have run, or when the budget is spent. A missed target is reported as missed, never relaxed. The pull request shows the baseline and final values, every attempt including the reverted and inconclusive ones, and the best idea not yet tried. The loop is on choosing your route. The 124 ms to 8 ms run is one fixture CLI on one machine: three attempts kept, one reverted, one inconclusive.

Refine asks you fewer questions. When a question is a fact it can observe, such as how long a query takes, whether a layout fits at 320 px, whether a parser accepts an input or whether an eval separates two variants, refine runs a throwaway experiment and records the question, what ran, what it saw and the decision under ## Resolved via Experiment in the spec. An experiment that would touch live state, credentials or the network becomes a question for you, and so does a result too noisy to decide, with the data attached. Product and preference calls still come to you. When flow settles a design fork with a prototype, it now builds the competing variants behind one labelled switcher, and when the options are still open it gathers prior art first and lets you pick a direction. See refine and flow.

Codex reviews keep the model they started with. Since codex-cli 0.154 a resumed session ran on the model in your Codex config, so a review pinned to one model re-reviewed on another. Re-reviews, validator passes and deep passes now re-pin the model the review started on. Thanks to @TechupBusiness for the report.

What changes when you upgrade. The bundled scouts run on Opus instead of Haiku and Sonnet, so plan and prime, which fan out scouts, cost more per run. They pin Opus rather than your session model, so planning on a pricier model does not make every scout run on it. To trade the cost back, name a cheaper model on the fast scout and thinking scout lines of your routing block. The generated Codex agents move to the gpt-6 models. You don’t need to re-run setup.

What you keep. Specs without a pre-registration section run exactly as before, the worker and the PR comment resolver still follow your session model, and explicit routing still wins over every default.

Under the hood. Set CODEX_MODEL_INTELLIGENT or CODEX_MODEL_FAST when regenerating the Codex mirror to choose other Codex models. The codex triage judge defaults to gpt-6-luna; the copilot one stays on claude-haiku-4.5.

6.3.0 - Bug fixes arrive with their cause and proof

Section titled “6.3.0 - Bug fixes arrive with their cause and proof”

Hand flow a bug and the pull request shows the cause, the commit that introduced it, and one reproduction failing on the base and passing on the head. Stages that drive your app now start from the feature map, which took about 40% fewer turns when the task didn’t say where its target was (one fixture app, one model, 66 runs).

Detail

You hand flow a bug report, console output or a screenshot as before. Before it writes a fix, flow looks for work that already exists: open pull requests and branches that touch the failing path, recent commits and reverts, the bug track in your project memory, and tracker issues. When it finds a fix, it runs your reproduction against that fix, records the result and stops, and it never writes a competing one. When someone visibly owns a fix in progress, or several open pull requests touch the area, flow stops and hands the bug back to you with the list. An unattended run ends with NEEDS_HUMAN at any of these stops. A reverted attempt counts as a cause already ruled out.

Flow then reproduces the symptom twice, lists the candidate causes, and eliminates them with runtime evidence such as logging, instrumentation or a narrowed input. It designs the fix only after one cause survives. When the report or the history names a revision where the bug did not exist, flow bisects with the reproduction to find the commit that introduced it. A symptom that will not reproduce is reported as not reproduced, with what was tried, and no fix ships for it.

The proof is one reproduction run twice. It must fail on the base revision and pass on the head, and it runs on the live app when the bug lives in one. When the reproduction already passes on the base, it does not capture the bug, and flow says so and does not claim the fix. The pull request briefing lists the prior-fix findings, the confirmed cause, the introducing commit, and the base and head results. A step that did not run appears as not done, with its reason. The steps are on choosing your route.

Stages that drive your running app now read the feature map first. That covers performance baselines and post-change measurements, the live proof of a bug fix, QA, and the live checks reported in pull requests. Each stage reads the map’s index and the one matching feature file, including its notes on controls that misbehave, and later stages on the same spec reuse the matched feature. The briefing names that feature, or says unmapped. On one fixture app with one model (66 runs), tasks that did not say where their target was took about 40% fewer turns and a third to half less wall time, with no drop in success. When the task named the page, reading the map cost 0 to 3 extra turns.

What you keep. Bug intake reads the map only when the report doesn’t say where, as in 6.2.0. Planning, questions, refactors and other routes that never drive an app never read the map, and a repository without a map pays one existence check. You don’t need to re-run setup.

Under the hood. A script that writes QA results can pass the matched feature to flowctl qa receipt as resolved_feature, in the same shape flowctl done --resolved-feature takes. A stage matches the feature again when the recorded file, its sub-feature or its last-proven line has changed. Details are on the CLI reference.

6.2.0 - The feature map keeps itself current

Section titled “6.2.0 - The feature map keeps itself current”

Your feature map now stays current as a side effect of normal work. A bug report that doesn’t say where the problem is starts from the map, which took the hardest untitled screenshot on a fixture app from 38 turns to about 12 (48 runs, same model).

Detail

A spec that renames a button or moves a page now updates the matching feature file in the same pull request. Work proves each new route with one live drive, and you review the map diff beside the code diff. When work can’t start the app, or can’t tie the change to one feature file, it leaves the map alone and files a drift note. QA, drive and bug intake file the same note when a mapped route no longer matches the live app, and a later proof of that route closes it, so the open notes are the open drift.

Flow tells you when a full pass is due. Run /flow-next:flow with no argument and it adds Also recommended: /flow-next:features when a drift note is open or a feature was last proven 50 or more code-changing commits ago. Each feature file now records the date and commit it was last proven at. Setup and prime recommend seeding a map on every run, whether or not live QA is on. The keep the map current guide covers the three mechanisms and how to run maintain on a schedule.

Bug intake uses the map when the report doesn’t say where. Hand flow an untitled screenshot or a report like “get this thing out of my list”, and it matches the report to one mapped feature, using what the screenshot shows as well as its text, then drives the reproduction along that feature’s route. On a fixture app with six reported defects (48 runs, same model), the hardest untitled screenshot went from 38 turns to about 12 and the other one did not change, with every defect reproduced in both arms and flat cost and wall time. A report that names the page or control goes there directly, as before, because reading the map first only added turns there. No match falls back to live discovery, and flow never edits the map during intake. The matched feature travels with the fix, so QA and the pull request briefing start from it. See bug intake.

What you keep. /flow-next:features stays the only command that seeds or fully maintains the map, and it runs only when you invoke it; no unattended run starts it. Existing maps load unchanged. A file without a last-proven line reads as never proven, so the first due report asks for one maintain pass. A repository without a map pays one existence check. You don’t need to re-run setup.

Under the hood. flowctl features status [--json] reports each file’s last-proven line, its age, the open drift notes and a seed, maintain or none recommendation. flowctl done --resolved-feature records the matched feature in the task’s evidence. Change the due threshold with flowctl config set features.staleAfterCommits <n> (default 50), documented on the configuration reference.

6.1.1 - Re-run setup for the prose reminder

Section titled “6.1.1 - Re-run setup for the prose reminder”

Re-run /flow-next:setup once in each repository, so the block it writes into CLAUDE.md and AGENTS.md points agents at the prose skill by an id they can call. Plan also dispatches each scout one step sooner.

Detail

6.1.0 made /flow-next:* commands typed-only, so the old block’s instruction to invoke /flow-next:prose pointed agents at a command they can no longer call. The refreshed block names the skill id, flow-next:flow-next-prose. The block’s version moves to 3, and setup offers the refresh once per repository.

Plan no longer asks the tier judge before each scout. Scout tiers are fixed: memory-scout runs as the fast scout and every other scout as a thinking scout, so the call cost a step and never changed the result. Work still asks the judge for task dispatches.

pipeline.chainStages is deprecated with no removal release set. The key still works under flow --auto --tick.

An unattended run now keeps the spec you named and stops on blocked work with a line saying what it needs, and a typo in .flow/config.json no longer resets your settings. Each worker loads about 60 KB of context per task instead of 160 KB, at the same comprehension score. Two retired commands are gone, so read the upgrade steps.

Detail

A run you leave alone works on the spec you gave it, under the review routing you set. It resolves a tracker key through the same lookup every other command uses, keeps per-task review overrides, creates the branch from the resolved default base, and stops on a git failure. Blocked work stops the run where it used to be tried again. A review result whose verdict or finding counts disagree with the review it records is refused, and every guard stops when its evidence is missing. You come back to a stop line that names what the run needs from you.

Your settings and your teammates’ edits survive. A typo in .flow/config.json used to reset every setting to its default. Now each reader names the file, with line and column for a syntax error, and config set refuses to write over it. When someone edits a spec’s tracker issue after the last sync, a push leaves their edit alone. An attended run asks whether to reconcile both sides (the recommended answer), overwrite, or leave it. An unattended run never overwrites; it queues the decision with the other deferred work.

Agents read less. A worker’s context for one task drops from about 160 KB to about 60 KB. It carries the memory index as text, only the glossary entries its task names, and a short git status, and both versions scored 7 of 7 on the comprehension eval across three task sets. The skill listing every session loads shows each skill once where it showed it twice. Review prompts, review and QA results, prospect output, tracker bodies and memory audit changes now come from flowctl, so a skill writes only its judgment. A bad input reports every error at once and writes nothing. Setup remembers the optional questions you declined and doesn’t ask them again.

Reviews need fewer arguments. Plan review runs without a file list, impl review uses the repository’s default branch when you leave out the base, and plan and completion review results go under each checkout’s own .flow/tmp/, so two repositories never share one.

What changes when you upgrade. The /flow-next:pilot and /flow-next:interview commands are removed. Replace them with /flow-next:flow --auto --tick and /flow-next:refine in loop recipes, scripts and instruction files; the loop form is on Driving a loop. You still type slash commands as before, but an agent can no longer call one, so a prompt that tells an agent to run a skill should name it by id (flow-next:flow-next-<name>). A script that reads review_attempts or tracker from flowctl show <spec> --json should call flowctl review-rounds attempts and flowctl sync get-state instead. A script that pushes tracker bodies should handle the new tracker_diverged conflict (Tracker operations). flowctl done now refuses evidence with none of commits, tests or prs. No config key changed. The setup block refresh this release needed arrived in 6.1.1.

What you keep. Typed slash commands, every config key including the pilot.* keys, the flowctl pilot verbs, and the PILOT_VERDICT grammar keep their names. Existing .flow/ state, review results and config files load unchanged. Explicit inputs that skills passed before (a --files list, a --base, hand-built evidence) still work.

Under the hood. flowctl gains verbs for one-call state reads, rendered artifacts, memory audits, rolling admission and commit-range evidence, among them glossary list --match, done --range and tracker sync --op push --overwrite-diverged. The full list is on the CLI reference.

6.0.2 - Specs made from an issue stay linked

Section titled “6.0.2 - Specs made from an issue stay linked”

A spec you create from a tracker issue is linked to that issue the moment it exists, so the next sync updates the issue and opens no duplicate.

The plan, capture, work, refine and QA skills pass the issue’s id and URL for you. A script that calls flowctl spec create --tracker-first should add --tracker-id and --tracker-url; with only the key, the spec stores the key alone, as before. An issue already linked to another spec is refused before anything is written. Thanks to @sn-furali for the report (#464).

Specs and comments pushed to Jira now show real headings, bold, lists, code blocks and tables. Turn on Jira’s Wiki Style Renderer for the description and comment fields.

flowctl converts each body to Jira wiki markup on the way in and back to Markdown on the way out, so your specs, merges and comment history stay in Markdown. An issue pushed before this release is converted by its next push or reconcile, and until then a pull refuses it with a jira_body_unconverted conflict. Thanks to @flecamos for the report (#465).

6.0.0 - Pull requests brief the reviewer, 60% faster

Section titled “6.0.0 - Pull requests brief the reviewer, 60% faster”

Your reviewer gets a briefing instead of a file list: why the change exists, the review steps in order, the few files to read, and only the checks that passed. make-pr writes it 60% faster (251 seconds to 100 on one 23-file pull request, three cold runs per point). This is a major release, so read the upgrade steps if you call land from a script or a schedule.

Detail

Your reviewer opens a body that starts with why the change exists and what changes for a user or operator. Scope gives the review steps in reading order. Each step says what to check and links up to ten files that must be read, with each file’s purpose and the requirement it serves, and everything else is counted in one line. Verification ticks only what passed; a failed or unverified check stays unticked with its note. Blast radius, tradeoffs and open items follow, and an empty section is left out. A 23-file pull request comes to about 800 to 1,000 words, where the old form ran 2,700 to 3,700 words for 25 to 39 files (#447; thanks to @flecamos for the report). A house style in your AGENTS.md or CLAUDE.md shortens it further (how). A branch that closes several specs gets one body, with requirement ids qualified per spec such as fn-250:R4, and lands as one pull request.

make-pr gets there faster. On a fixed 23-file pull request the median run went from 251 seconds to 100, from 22 tool calls to 13, and from 21,627 output tokens to 8,174 (three cold runs per point). The agent writes only the judgment and flowctl fills in the rest. The spec closes as the last commit on the branch, so the close reaches a protected base through the merge itself.

Land handles one pull request per run. Name it with /flow-next:land <PR>, and land resolves the review threads, makes one focused CI fix or one flake rerun per failure, catches the branch up on the server, and squash-merges pinned to the full head SHA. It merges only when you authorized the merge in the session; without that, it repairs and stops at AWAITING_REVIEW with reason merge-ready; authorization required. It never rebases, force-pushes or retargets, keeps no state between runs, and writes nothing to the repository after the merge; the tracker update is the one step left. This release’s own pull request (#459, 175 files, four specs) landed that way.

A worker now waits for every command it started before it reports back. When one is still running, work waits and sends a continuation worker into the same workspace, and the early return does not count as a failed attempt.

What changes when you upgrade. This is a major release.

  • Replace every bare or scheduled /flow-next:land with a loop over named pull requests: list them, call /flow-next:land <PR> on each, and put the instruction to merge them in the prompt, because a list grants no merge authority. The copy-paste blocks are on the land loop.

  • The default merge gate is green checks, a nonblocking GitHub review decision and zero unresolved threads, and a bot comment neither satisfies nor blocks it. If you relied on land.reviewSignal, move the requirement to your instruction file, to branch protection, or to land.mergeVerdictCommand.

  • Eight land.* keys are retired. They still load, land names them once and ignores them, and you can remove them when convenient.

  • Land no longer releases, requests reviewers or honors the FLOW_PR_MERGE_CMD override. Run releases yourself after LAND_VERDICT=MERGED and request reviewers outside land.

  • Drop --no-mermaid from any make-pr script.

  • If your repository already tracks pull request aid files, untrack them once from the repository root and commit the index change. Your local files stay.

    Terminal window
    flowctl init
    git rm --cached --ignore-unmatch -- '.flow/artifacts/*/pr-cognitive-aid/*.json' '.flow/artifacts/*/pr-cognitive-aid/.write.lock'

The rest of the upgrade notes are on Compatibility.

What you keep. The LAND_VERDICT grammar and its vocabulary; RELEASED stays for parsers and is never emitted. Existing config files load. Old land files under .git/ are inert. The stored aid artifact keeps schema version 1, and the HTML lens is unchanged. flow --auto <spec> --until=merge still carries one item through landing.

Under the hood. New read-only verb flowctl spec closed-in-range --base <ref> prints the specs a branch closes. A dependency squash-merged into a non-default base now reads as landed in flowctl spec chain. The clean-review judge preset is removed with the land gate it served, which leaves five judged sites.

5.6.1 - Fresh ideas route in about a second

Section titled “5.6.1 - Fresh ideas route in about a second”

With a TypeSafe key set, a new idea or brief now gets its route in about 0.6 seconds on the first try, where 5.6.0 could spend minutes before falling back to the host’s own judgment.

Specs that already exist were never affected. The state for an idea or brief now holds only its text, and flowctl assembles the repository, lifecycle and pull request facts itself, so a host can’t invent them. A state file missing several fields names all of them in one error.

5.6.0 - Optional judgments for six pipeline decisions

Section titled “5.6.0 - Optional judgments for six pipeline decisions”

With a TypeSafe API key in your environment, six narrow pipeline decisions each become one judgment request made by the same flowctl command on every host, about 0.6 seconds for routing. Without the key nothing changes.

Detail

To turn it on, run export TYPESAFE_API_KEY=<key> in the shell that runs your coding agent. There is no SDK, model selector or config to write. Setup prints Judge: off (TYPESAFE_API_KEY is not set). when the key is absent, and flowctl config set judge.enabled false turns the judge off while keeping the key.

What you see. Flow’s route step and flow --explain print the route with its confidence, such as Route: defect (jev 0.91), or Route: host (jev below floor: ...) when the judge hands the decision back. At the 0.7 floor the judge routed 85% of held-out intents on its own at 95% agreement with the labels; below the floor, the host picks from the top three candidates. Under pipeline.qa=auto, flowctl supplies whether the app can start and the judge answers only whether the acceptance criteria describe a UI, so a skipped QA stage names which half failed. The fork check confirms a fork exists before it classifies one, which cut invented forks from 12 to 1 in the evaluation. Memory search reranks its keyword hits in one request, and plan and work skip the memory scout when it is available. A confidently mechanical task goes to your fast-scout tier, and a confidently long-running one gets a bridge recommendation; everything else stays on the session model. Land asked the judge whether a bot review was clean until 6.0.0 retired that gate.

What you keep. Every review, QA and merge gate keeps its own contract, and no judgment predicts a verdict. An explicit IMPLEMENTER: line in the invocation wins over the tier answer. The floors are fixed. Every stage line names which path ran, and when the judge can’t answer (no key, a timeout, a bad response, a state over budget) the decision takes its previous path and says so. The key is read at call time and never written to config, receipts, stage lines or logs.

Under the hood. One flowctl judge --preset <name> command carries the request, retries, validation and decision rule for every site. The evaluation numbers, their bounds and the question text are on Optional Jev judgments.

5.5.1 - Your spec scaffold decides every section

Section titled “5.5.1 - Your spec scaffold decides every section”

A customized SPEC.md now controls every section of a captured spec, so you can drop the 15 to 20 lines of quoted prompt text capture used to add. If you wrote your SPEC.md by hand, copy the auxiliary_sections list from the bundled template first.

Detail

Capture now writes only the sections your scaffold names, where your scaffold puts them. It used to add a ## Conversation Evidence block at the top and a ## Requirement coverage table at the end whatever the scaffold said, and a hand deletion came back on the next rewrite. To drop the evidence block, delete Conversation Evidence from the auxiliary_sections list in your SPEC.md. The scaffold guide has the details. A repository without a custom SPEC.md sees no change.

What changes when you upgrade. If your SPEC.md was written by hand and has no auxiliary_sections list, captured specs stop getting the strategy, parked-unknowns, evidence and coverage sections. Copy the list from the bundled template to keep them. A scaffold copied from the bundled template already has it.

What you keep. Capture still collects your verbatim turns during the run and checks every [user] tag against them before it saves, so a tag still means you said it. With the evidence block dropped, a reviewer who doubts a tag can no longer look the quote up in the spec.

Two smaller fixes. When you answer one of capture’s questions by picking an option, the evidence line quotes the option label exactly and marks it as a selection. The prose contract now says that a length budget or reading level in your AGENTS.md or CLAUDE.md reaches every artifact the agent drafts, and that the contract sets no length rule of its own. Thanks to @flecamos for the report and the source reading.

5.5.0 - Maintainability questions at plan time

Section titled “5.5.0 - Maintainability questions at plan time”

Plan review and the technical refine pass now ask two advisory questions before code exists: does the plan repeat an edit or a decision in several places, and does it bend the intended dependency direction. A second run as the same person on one clone now refuses a task the first run holds.

Detail

Plan review adds a Maintainability criterion, answered from the plan as written. Does the plan make the same edit or decision in more than one place? Does it add a back-edge against the intended dependency direction, or new branching to a function that is already the hottest in its module? Each answer is a concrete finding or none identified, shown as an advisory block in the verdict, and a named finding also lands as one line in the spec’s Decision Context. The technical refine pass asks the same two questions once, so a spec on the direct route, which skips plan review, still answers them before build. A finding pushes the verdict to NEEDS_WORK only when it names concrete duplication, a specific back-edge, or a specific function absorbing the branching. “Could be cleaner” and suggested abstractions are out of scope.

The public claim now says what the pipeline does not prove. The README, the docs home and the verification spine carry the same sentence: the pipeline proves the change does what was asked and records what it did; it does not prove the codebase stays maintainable. The findings are advisory, never a gate, score or threshold.

Two runs as the same person no longer share a task. A second terminal, a scheduled flow --auto tick or a second machine on a shared checkout used to read its own flowctl start as a crash resume and send a second worker onto a task in progress. Now flowctl start refuses with an error naming the task, the claimant and the two ways forward: confirm the earlier run ended and re-run with --reclaim, or leave the task to it. Work and flow --auto pass --reclaim only after they have evidence that the earlier run ended. Claims by other people and --force behave as before, and one conductor per clone stays the rule. Two rolling runs on one checkout also stop overwriting each other’s notes.

5.4.0 - An implementer on another CLI owns the task

Section titled “5.4.0 - An implementer on another CLI owns the task”

When your routing block sends implementation to a model behind another CLI, such as codex exec, that model now owns the task. It commits checkpoints and runs its own subagents, and the worker reviews its commits before the usual review and gates.

Detail

Until now the worker ignored an implementer tier reached through another CLI and implemented on the session model. A bridge prompted by hand told the child it could not commit or spawn agents, so a long task became a chain of returns and re-briefs; one reported task took 19 dispatches. The next work run under such a routing block takes the new path. With no implementer tier, or one your harness reaches in-host, nothing changes.

The worker hands the child a short prompt with the task’s identities and rules plus the long-task brief from flowctl usage, and runs it as one foreground call. The child writes code, commits checkpoints on the branch it was given, and decides its own parallelization the way an in-host worker does. It never pushes, rewrites history, changes scope, issues a review verdict or starts a nested bridge. On return the worker commits anything left uncommitted, reviews the child’s commit range against every acceptance criterion, runs the focused gates, and continues into review and flowctl done as before. Your done summary says which model implemented and how many subagents it dispatched. To override the model for one run, add IMPLEMENTER: <model> at <effort> to the work invocation.

On Codex, workspace-write keeps .git/ read-only, so a child that commits checkpoints runs with danger-full-access inside the repository root, or the host commits between one-run-per-scope dispatches. flowctl usage describes both.

What you keep. The worker owns the range review, the gates, review dispatch and the done record; the conductor never bridges; land still needs its own consent.

Under the hood. The docs no longer tell you to keep a Codex child flat. The upstream decode bug behind that advice is still open for codex-cli 0.144 to 0.145, but its own minimal repro ran clean three of three on 0.153.4, and a month of spawning runs on the maintainer’s machine showed zero decode errors. Thanks to @DanielKillenberger for #431 and the diagnosis in #437.

5.3.0 - Dependent specs build on their parent’s branch

Section titled “5.3.0 - Dependent specs build on their parent’s branch”

A spec that depends on another now starts on the parent’s branch as soon as the parent is built, instead of waiting for its merge. Its pull request shows only its own layer, joins a GitHub stacked pull request where available, and land merges the chain from the bottom.

Detail

Take specs B and C that depend on A. The build loop used to park B until A’s pull request merged, then C until B’s did, so each layer cost a review, a merge and an idle wait. Now B becomes selectable once A’s tasks are done and its branch is on origin. Work forks B’s branch from A’s tip, and make-pr opens B’s pull request against A’s branch. The dependency edges your plans already record are the only input, and a spec with no dependencies behaves as before.

Your reviewer gets a chain of small pull requests. On GitHub, make-pr links them into a stack (in public preview since 2026-07-30, announcement): the merge box shows the stack map, each layer’s diff holds only its own change, and you review and merge from the bottom. Off GitHub, or where the preview is unavailable, the same chain works as plain dependent pull requests. flow --auto picks chained specs up in ready and backlog mode. A spec whose parent is still in progress, or a second child of the same parent, parks with a stated reason and no strike. Chains are linear.

Land merges only the bottom open layer, so a stack never collapses from the top. On a GitHub stack it uses the asynchronous stack merge with a head pin the server enforces, and a stale pin is refused before anything merges. A merged branch stays until no open pull request targets it. Merge judgment stays yours, and nothing in this release merges on its own. This release’s own two pull requests, #432 and #433, formed the first live chain and merged from the GitHub stack UI.

Since 6.0.0, land uses GitHub stacks when they are available and leaves a conflicted layer of a plain chain for you to rebase by hand. See Land, Make PR, and Work.

5.2.2 - Five reported defects, no new knobs

Section titled “5.2.2 - Five reported defects, no new knobs”

If you keep a glossary, land PRs from a zsh or worktree setup, or project specs to a GitHub or GitLab tracker, five things that silently went wrong now behave; nothing new to configure.

Detail

Nothing to do first. Update the plugin and the fixes apply.

A glossary entry that carried both an Avoid line and a Relates to line used to lose part of its relationships on every glossary add, even when the command touched a different term, and the list command kept reporting the right count so nothing warned you. The parser now removes those two lines in the right order and the file round-trips unchanged.

Land’s merge step failed with “command not found” when the agent’s shell was zsh, which is the macOS default. The command is now built as an array and works under bash and zsh alike; the FLOW_PR_MERGE_CMD override keeps its contract. In the same skill, when the PR branch is already checked out in another worktree (the Worktree Kit shape), the ci-fix step now tells the agent to run the fix in that worktree instead of failing on the checkout. No worktree is created or removed for it.

On GitHub and GitLab trackers, every planned spec recorded a “status conflict” on its first push because a spec with all tasks still to do was compared against the status:backlog label from capture. Those are the same “not started” bucket and now agree. And a merged PR whose whole diff was the spec’s own text (the one-PR-per-gate convention) no longer counts as shipped work, so a still-open spec stops jumping to In Review on its first claim; each merged PR is checked for files outside .flow/specs/ and .flow/tasks/ before it counts as evidence.

Under the hood: each fix ships with its own regression test, including one that runs the real merge fence under both shells; the status-sync reference now says that perTracker.statusMap is read only by the Jira and Linear providers. Thanks to @flecamos, @TechupBusiness and @sn-furali for the reports.

5.2.1 - Change implementers, keep the route

Section titled “5.2.1 - Change implementers, keep the route”

Choosing another implementation model or harness no longer forces a task breakdown; a ready cohesive spec keeps the direct route.

Plan when you ask for it, separate human owners will implement, or delivery spans multiple PRs. Existing plans, research, refinement, review, and QA retain their own rules. The orchestration block still controls model and harness choices. This builds on the 5.0.0 conductor and the 5.2.0 merge destination.

Carry one approved spec through review and a gated merge in the same flow run, while keeping the choice to stop at the PR. It builds on the 5.0.0 conductor below.

Detail

Add --until=merge to /flow-next:flow <spec> or /flow-next:flow --auto <spec> when you want the selected item landed. Returning to an existing PR with plain attended flow offers that next step and asks once unless you have already authorized it. Without the destination, unattended flow keeps its pre-merge stop.

Flow calls land for the selected PR. Land keeps its CI, review, dependency and branch-protection gates; waits do not spend pilot strikes. --tick performs at most one landing tick. Consent stays with that item during retries and waits, and a fresh session needs current authorization again. A completed merge and an unfinished post-merge tail are reported separately, so recovery never repeats the merge. Releases and tracker writes retain their own authorization and configuration.

Capture also saves the spec before offering the saved file in your editor. The redundant approve-and-write question is gone. Product questions, split choices and readiness consent remain; plan/refine approval and scripted autofix’s --yes gate remain.

Existing verdict names and standalone land remain supported. The deprecated pilot and interview aliases are retained for compatibility in this release; use flow --auto --tick and refine in new recipes. See Flow, Land and Capture for the full contracts.

Plan review no longer asks for a task split. Read the 5.1.0 entry for flow --auto, and the 5.0.0 entry for the release that introduced Flow.

A spec with zero tasks or one owner task is the default route, so plan review now reviews the spec’s content and treats task count, decomposition, and dispatch shape as the owner’s decision. The consistency and Touches checks run only when task specs exist. A missing approach order or test sequence is still a finding against the spec, and the owner decides whether it changes the split.

5.1.0 - One unattended invocation, from ready spec to draft PR

Section titled “5.1.0 - One unattended invocation, from ready spec to draft PR”

Teams that run Flow-Next unattended get one driver instead of two. One /flow-next:flow --auto invocation carries a ready spec to a draft PR, hop after hop, classifying each hop from the same routing references the attended conductor reads, so the route you see in --explain is the route the unattended run takes. Pilot users keep working through the alias for one release. /flow-next:pilot is retired in 5.1.0, forwards to flow --auto --tick with one deprecation line, and the release after 5.1.0 removes it. It builds on the 5.0.0 conductor below.

Detail

Do these first. Update /loop, /goal, and Ralph recipes now. The default recipe is one /flow-next:flow --auto invocation per item; the loop recipe is /loop 30m /flow-next:flow --auto --tick. An existing /flow-next:pilot recipe keeps working for this release. The shim rewrites --spec <id> to the positional id, rewrites --dry-run to --explain, maps both --backlog and pilot’s old --auto backlog switch to --backlog, passes --review, --research, and --depth through unchanged, prints pilot is now flow --auto --tick; this alias is removed in the next release to stderr, and behaves byte-for-byte as a pilot tick. pipeline.chainStages is deprecated with the alias. It is honoured under --tick and ignored with one stderr notice in long-horizon mode.

What was hard before. An unattended run meant a driver looping /flow-next:pilot and paying a full re-anchor (skill re-read, config snapshot, selection, classification, branch resolution) plus the loop interval at every stage boundary, with only qa into make-pr allowed to chain. Pilot classified from its own private stage table, so the route it took and the route flow --explain showed could differ.

What you get now. The driver invokes /flow-next:flow --auto [<spec-id>] once. The conductor classifies the hop, dispatches the stage with mode:autonomous, verifies from observed state, records the hop, and re-classifies until it reaches a terminal: the PR exists, deferred to land, asked, blocked, needs human, or no work. Every hop ends with the same receipts, evidence echo, and ledger write a pilot tick ended with, so a run that dies at hour six resumes from disk on the next invocation; nothing is resumed from transcript. --tick runs exactly one hop and stops, which is what a pilot tick was, for hosts without stable long sessions. Both shapes end with the one PILOT_VERDICT line drivers already parse; a long-horizon run names every dispatched stage joined by + (stage=work+qa+make-pr) and carries the last hop’s verdict. Every rail pilot had moves across unchanged: the strikes ledger, the dirty-tree refusal, the all-done PR probe, the never-merge boundary, the decision log. The operator still reads one verdict line and still owns the merge. Flow —auto, the build loop is the page; Driving a loop has the recipes per host.

Unattended classification reads the routing reference. --auto reads route-matrix.md for the spec-state rows, plan-vs-no-plan.md for a ready spec with no tasks and no recorded route (it records the route with spec set-no-plan or spec clear-no-plan before any mint and echoes the deciding signal), and gate-selection.md for the review, QA, and completion-review gates. --explain under --auto prints the selected spec, the classified stage with its routing row and gate, the consulted fields, the PR probe, and would-clear ledger entries, with no write, no checkout, and no dispatch.

pipeline.qa=auto now takes effect unattended. A skipped QA hop records stage: qa - skipped(config: pipeline.qa=auto: <reason>) and advances to make-pr. on and off are byte-for-byte unchanged.

The refusal is inverted for --auto. Attended flow still refuses under every autonomy marker; its line now reads NEEDS_HUMAN: /flow-next:flow is attended - run /flow-next:flow --auto for unattended runs. --auto refuses only under Ralph (FLOW_RALPH, REVIEW_RECEIPT_PATH), because it sets FLOW_AUTONOMOUS and mode:autonomous for the stages it dispatches. Flow, flow --auto, and Ralph are three drivers and are never nested; --auto never dispatches land or a second driver. Land is unchanged and DEFERRED_TO_LAND keeps its meaning.

Backlog mode is flow --auto --backlog. pilot.autonomy=backlog still enables it and every safety invariant stays. In long-horizon mode a backlog run drives one selected item to its terminal and stops; the next invocation selects the next item.

What did not change. Config keys (pilot.autonomy, pilot.gateClasses, pipeline.qa, pipeline.chainStages), flowctl verbs (flowctl pilot strikes list|clear, flowctl pilot-log), the ledger path, .flow/pilot-runs/, and the PILOT_VERDICT name are not renamed; a rename is a separate, deliberate break for a later major. Published counts drop by one skill and one command.

Pending. Terminal-parity and wall-clock results for long-horizon versus tick execution are not yet measured. Full detail in the repository changelog.

A bare /flow-next:flow now proceeds to the next best step on its own, and gate receipts no longer fail because of a stray .git above your working tree. Read the 5.0.0 entry for the release itself.

With no argument, flow resolves the item this conversation last touched, then the spec matching the current branch, then asks to capture intent no spec holds yet, then picks the next open spec by judgement with an inline pick on ties, and only then asks what to work on. The gate receipt probe stops at any GIT_CEILING_DIRECTORIES entry, the same boundary git uses, so an empty /tmp/.git no longer turns an outside-a-repo check into a tooling error. The README and plugin docs now describe the shipped 5.0.0 routing contracts.

5.0.0 - Hand it anything, it picks the route

Section titled “5.0.0 - Hand it anything, it picks the route”

You no longer choose the next command. Hand Flow-Next whatever you have, from nothing to a pasted bug report to a spec with an open PR, and it chooses the smallest sufficient route, runs it, and stops at the next decision that is yours. The recommendation you read and the route that runs are now the same rule.

Detail

Do these first. Two command names changed, so update scripts, /loop and /goal prompts, CLAUDE.md or AGENTS.md policy paragraphs, and team docs:

  • /flow-next:guide is removed with no alias. Use /flow-next:flow --explain <the same words> for the recommendation, or plain /flow-next:flow to run the route.
  • /flow-next:interview is now /flow-next:refine, same scopes and flags. The old name forwards with one deprecation line for this release only; the release after removes it.
  • Live QA can now decide for itself: flowctl config set pipeline.qa auto runs the live pass only where the spec describes UI behaviour on a surface the QA skill can reach. Existing off and on values keep their meaning. Pilot and land invocations are unchanged, so an unattended /loop 30m /flow-next:pilot recipe needs no edit.

What was hard before. Every stage skill printed its own next-step advice, the guide skill kept a third copy of the same matrix, and the three drifted. You read a recommendation, then typed the command yourself, and a ready spec still met the plan-or-not question on every route.

What you get now. /flow-next:flow takes anything: an idea, a spec or task id, a tracker issue, a branch, a pasted console dump, a how or why question, something slow, a cleanup that must keep behaviour, or a design fork. It matches the starting state against one shared routing reference, runs the routed skill by name, re-evaluates from observed state after each hop, and stops with a report that names every stage as ran, skipped(reason), or failed(reason). A run from intent ends when the PR exists; a run on an open PR converges it and stops when merge is the only step left. It never merges, never closes a spec, and never invents a review or QA verdict. Picks a stage produced (a ranked candidate, a split proposal) are asked inline and the run continues; only a decision that ends the run stops it. --explain prints the route and leaves .flow/ byte-identical. On hosts that match skills by description, “this endpoint is slow” reaches flow with no slash command at all. Choosing your route is now the flow --explain walkthrough.

Direct execution is the default. A ready spec with no tasks routes to work --no-plan. Plan is chosen only on a positive signal: you asked for one, separate human owners will implement, delivery is staged across several PRs, or the implementer is routed out of the session model. Risk, size, and file count never trigger a plan on their own. Capture’s and plan’s closers print the same rule’s result, and under flow, capture records the route itself.

The routing reference grew. Beside the six worked routes the docs already carried, the matrix names refactoring (pin the contract with a characterization test first), performance (baseline on a real surface before any change), hill climb (a frozen harness and a target that is never relaxed), investigation (a cited read-only answer, with a new why-scout that tiers each finding as direct, supported, inferred, or unknown), and prototype (an observable fork is settled by running something, never by asking you to guess). Every row names its positive signal, its safe skip, and the skip kind.

Refine gains a read-first pass. /flow-next:refine <spec> --scope=research asks nothing: four read-only scouts resolve the libraries and APIs the spec names into one ## Resolved via Research section with a source on every line. It skips itself, visibly, when the section exists or a plan already ran the scouts, and flow routes to it only when a spec names a library the repo does not already use.

Less to read at the read-back. Capture, refine, and plan write the draft once and show you a summary, one ask (approve and write, open in editor, abort), and only the diff on each edit cycle. Before, three full copies of the spec printed per edit. Ratification before every write is unchanged.

What did not change. Pilot is byte-for-byte the same and activates its QA stage only on the literal on; auto is flow’s call, because judging whether a surface is drivable is attended work. Flow never runs under pilot or Ralph. The unattended conductor and harness or model autorouting are the road ahead, on Flow and the road ahead, with no dates promised.

Evidence. A pre-registered non-inferiority study (90 draws, one frontier model, 27 fixture situations) scored the retired guide matrix 24/27 and the shared reference 27/27 on the nine discriminating items; the reference read its required file on every draw. Superiority was never the claim. Both implementation tasks reached SHIP under cross-family review, on rounds 2 and 3.

Under the hood. Six reference files under the flow skill (route-matrix, spec-count, plan-vs-no-plan, gate-selection, prototype-before-ask, tail), each opening with a decision record; capture’s closer, plan’s menu, work’s zero-task ask, and flow --explain read them by pointer, and a test fails on any pointer that names a missing file. pipeline.qa is a string enum off | on | auto; any other value is off. Published counts stay at 28 commands and 32 stable skills. Skill prose no longer carries spec-provenance markers. Full detail in the repository changelog.

4.18.0 - Better recommendations for your next step

Section titled “4.18.0 - Better recommendations for your next step”

Flow-Next now recommends working directly from a ready, cohesive spec when a separate task plan would add little value.

Detail

After capture or an interview, Flow-Next now helps you choose whether to refine the spec, review its design, split it into tasks, or start implementation. Its guide follows the same approach. When the spec is ready and a capable agent can own the work, the recommendation is work --no-plan, with a reason for that choice.

Previously, the guidance leaned too heavily toward creating a task plan first. Planning is now recommended when dependencies, separate owners, staged delivery, or execution constraints make a breakdown useful. A large change or a risky design alone does not make task planning mandatory: unresolved decisions call for clarification, and design risk can call for a separate review.

You still choose the route. The recommendation does not silently skip your requested reviews or change your configuration. Direct work carries the whole spec through the same configured implementation review, acceptance coverage, completion policy, and evidence requirements. Optional live QA remains available alongside staging and manual QA.

Internal benchmarking found that direct execution can produce higher-scoring implementations with capable frontier models. That informs the recommendation; it is not a guarantee for every spec or agent.

4.17.0 - Reviews through your managed host

Section titled “4.17.0 - Reviews through your managed host”

Developers working in a compatible managed host can keep reviews on the host’s selected provider account while retaining Flow-Next’s findings, fix loop and review receipts.

Detail

Update Flow-Next before enabling required managed reviews in your host. Standalone users need no new configuration and keep their existing CLI review path.

A managed session previously needed a separately launched reviewer CLI, which could pick up a different account from the one selected in the host. A compatible host can now supply the review execution path for that session. You choose the review backend as before, inspect the returned findings and receipt, and retain the same review and merge gates. Flow-Next still owns the review prompt, round accounting and verdict.

If the configured managed provider is unavailable, rejects the session scope or returns an invalid response, the review stops. It never silently switches to a local CLI. Provider and account availability remain the host’s responsibility; this release does not add cloud credential management or make every backend available in every host.

Under the hood. The installed launcher advertises its managed-execution capability. Hosts pass a local endpoint and session token; users do not configure those variables themselves. Completion reviews can require the managed path before reserving a review round. See managed review execution.

Reinstalling flow-next into Codex could leave a duplicate agent registration, lose a setting that sat after a commented table header, or fail outright on a Windows machine whose locale was not UTF-8. All three are fixed, and a forced task takeover with a custom note now hands the task over instead of leaving it with the previous owner.

Nothing to do on upgrade beyond re-running the Codex installer if you use it. The same release removes 141 lines of unused Python found by a five-reviewer pass over the CLI; prompts and public contracts are unchanged, and the pass also surfaced the tracker fixes listed in the repository changelog.

Teams conducting Flow-Next from Codex, Cursor, Grok Build, Droid, or OpenCode could not get a Claude-family review through the packaged review path - the closest thing was a hand-typed claude -p prompt with no receipt, no model ladder, and no fix loop. claude is now a review backend like codex, copilot, and cursor: set it once and every plan, implementation, and completion review carries the same receipt, round counter, and fix loop.

Detail

Nothing to do on upgrade. To use it, run /flow-next:setup and pick Claude Code CLI from the review menu (it appears when the claude CLI is on your PATH), or set it directly with flowctl config set review.backend claude - the spec form claude:<model>:<effort> pins a model and one of the CLI’s effort levels (low, medium, high, xhigh, max), and a per-task pin or --review=claude on any review command overrides for one run.

The reviewer is read-only by construction rather than by policy. The prompt arrives on standard input, the child process gets exactly three tools (Read, Grep, Glob) with every MCP server excluded, and there is no shell at all - because a pre-approved git diff would still accept flags that write files. The diff under review is handed over as a file path under .flow/tmp/claude-review/, one file per reviewed range, so a re-review after a fix commit reads the new range while the earlier evidence stays byte-identical. Deep passes and validator passes resume the primary review’s session by id instead of starting cold, which is what the shared deep-pass prompt has always assumed. When the pinned model is not available to your account, the resolution ladder steps down the ranking on the CLI’s real signature (it exits 0 with an error payload, not a non-zero code) at most twice and records what actually ran in the receipt.

Same-family reviews are allowed and receipted, not refused. On a Claude Code host a claude review is a fresh-process second opinion from the same family; the receipt names the backend and model, and the skill prose says so. Independence is judged on the model family that wrote the diff, not on the host name: Cursor, Droid, and OpenCode can run Claude writers too, so the docs condition every “cross-family” claim on the writer’s model. The host backend keeps its fail-closed cross-family rule; the first-round three-draw fan-out stays codex-only; the bridge recipes remain the way to have Claude write code from another host.

Ralph users get the same rows: the harness menu, config.env, the prompt templates, and the guard that blocks --force and other human-only recovery arguments all recognise flowctl claude review commands exactly as they do the cursor spelling.

Under the hood. The backend is a registry entry with the fn-76 ranking invariant, a stdin runner, an explicit resolver that rejects a foreign --spec and coerces foreign configured defaults (Claude ids do not cross over), and the five subcommands routed through the shared review driver. Plan review ran to SHIP over six rounds on two model generations; the implementation had four per-task cross-model reviews and a completion review on GPT-6 Astra. Reference: review backends, CLI reference, model routing.

A spec created from the command line used to arrive with different headings than a captured one, so plan review, R-ID coverage, completion review, and your PR body all read it half-empty. Every spec now starts from the same template - the one your repo controls.

Detail

Nothing to do on upgrade, and your existing spec files are never rewritten.

Flow-Next has two ways into a spec. /flow-next:capture and /flow-next:interview wrote the canonical template - Goal & Context, Architecture & Data Models, API Contracts, Edge Cases & Constraints, Acceptance Criteria, Boundaries, Decision Context - while flowctl spec create wrote a six-heading skeleton of its own. Every reader downstream keys on the template’s headings, so a CLI-born spec exported an empty goal and empty boundaries, and /flow-next:make-pr fell back to the spec id as the pull-request title. Three specs in a private repo shipped that way before the pattern was noticed.

flowctl spec create and flowctl spec skeleton now render the canonical template, resolved through the same cascade the skills already used: a repo-root SPEC.md, then spec.md, then the bundled template, first match wins. Point a repo at its own SPEC.md and command-line specs pick up your house sections too - no config key, no second scaffold to keep in step. The old skeleton is gone.

You keep every spec you already wrote. Specs with the older headings still export through read-only synonyms (Overview or Context, Boundaries / non-goals, Decision context in any case), and flowctl validate prints one legacy spec headings warning per such spec as the nudge to migrate - a warning, never a block. Two export bugs went with it: template guidance written in HTML comments no longer exports as acceptance criteria, and fenced code inside a section - a diagram, a snippet - no longer vanishes from the exported body. A sweep of 217 existing specs against the old parser found zero regressions and one accidental correction, a spec that had been exporting example text from inside a fence as its boundaries.

One fix rides along for anyone who renames specs. flowctl spec set-title renamed a spec’s id and its files but left branch_name at the old slug, so a retitled spec dropped out of land’s pull-request discovery and autonomous runs named the wrong branch. A branch_name still equal to the old spec id now follows the rename; a value you set yourself is kept, and the JSON result reports both the name and whether it was re-derived.

Under the hood. The byte-for-byte skeleton baseline moved onto the template itself and is hash-pinned, so a scaffold edit is always a deliberate bump. Deliberately not done: flowctl prospect promote and the plan skill’s own scaffold step still write the legacy headings, and are a follow-up. The implementation was written over a headless bridge to a different model family, then reviewed five rounds in-host - rounds two, three, and four each caught a regression the previous fix pass had introduced (phantom decision bullets from a trailing template comment, fenced code erased from exported bodies, fence-first sections exporting empty), and round five shipped only after the 217-spec sweep came back clean. Reference: CLI reference, customizing the spec scaffold.

Multi-task specs used to wait at every wave barrier until the slowest worker finished. /flow-next:work now schedules on the rolling frontier by default - the next task starts the moment any worker returns - and the wave loop survives only as a structural fallback. Nothing to enable, nothing to configure.

Detail

If you ran the experimental /flow-next:work-rolling command, switch to plain /flow-next:work: the beta command is gone, with no alias, because the scheduler it carried is now the default. Every other invocation is unchanged.

A wave-scheduled spec dispatched a safe group of tasks, then held the next group until the whole wave had joined, so one slow task idled every finished sibling. The rolling frontier admits a new ready task at every worker return, each worker in its own isolated worktree, and the conductor integrates, reviews, and completes each task as it lands. The pre-registered eval behind it measured a 52.1% work-phase wall-clock saving at quality parity, and the beta ran end-to-end on Claude Code, Cursor, and Grok Build before it graduated. The faster shared-checkout variant stays rejected: its speed came partly from thinner tests, and that is not a trade this project makes.

You keep the wave loop where rolling has nothing to schedule, chosen from the shape of the run rather than a knob: a task-id run (only that task runs), plan-sync switched on (its per-wave barrier is a fail-closed rule), a spec with fewer than two open tasks, or a fully sequential dependency chain. The route prints once at the start of the task loop - Scheduling: rolling or Scheduling: wave (<reason>) - so a run report always says which scheduler it used. A host whose subagent dispatch is measured to block still reports degraded to wave and keeps the rest of the rolling lifecycle. Pilot, land, and Ralph dispatch plain work and inherit the change; every gate, receipt, and review surface behaves exactly as before.

Under the hood. The scheduler reference moved to skills/flow-next-work/references/rolling-scheduler.md and is read only on the rolling route, so the always-loaded work surface grew by the route decision and one pointer. Reference: work, choosing your route.

If you run pilot and land unattended, two of the intervals the loop spends waiting on nothing can now be switched off: pilot can open the draft PR in the same tick as a live QA verdict, and land can measure its merge wait from the bot’s clean review instead of from the last push. Both are opt-in keys, off by default, and neither changes a gate.

Detail

Nothing changes until you set a key. flowctl config set pipeline.chainStages on and flowctl config set land.patienceMinutesAfterReview 15 are the two switches; leave either unset and the loops behave exactly as before.

Pilot advances one stage per tick, so a spec whose QA pass just finished used to wait a full driver interval, plus a re-anchor, before make-pr ran, even though make-pr was always the next stage. With pipeline.chainStages on, a tick whose live QA stage produced a fresh terminal verdict runs make-pr before it exits, with its own evidence block and stage line, and the verdict reads stage=qa+make-pr. The chain table is closed to that one row. The research finding also named plan → plan-review, and the plan review dissolved it: pilot’s plan dispatch already runs its own review loop to SHIP, so there was no idle interval there to remove. Nothing chains into work, and the switch is inert unless the QA stage is on.

Land’s default silence signal waits a patience window measured from the last push, restarted by every fix push, even once the bot has reviewed the current head clean and no threads are open. With land.patienceMinutesAfterReview set, that wait is measured from the review event instead, and only under silence, only while the review is head-current with zero unresolved threads. A fix push moves the head, the review stops being head-current, and the push window applies again until the bot re-reviews. The window is your human-objection grace period, which is why this stays opt-in and why the report line names which anchor bound: anchor=push or anchor=review. The approve and login signals, the merge command, and the ledger are untouched.

One fix rides along for every pilot user. The make-pr verification probe piped gh through jq | head, so a GitHub outage read as “no PR” and recorded a strike; two of those unready the spec. The probe now captures the gh exit status, and a probe failure is the crash-class NEEDS_HUMAN pilot already uses elsewhere, never a strike.

Under the hood. Both keys are seeded defaults in the published config schema (pipeline.chainStages as a strict off | on enum, land.patienceMinutesAfterReview as integer-or-null with minimum: 0; unset, null, and 0 are off). The skill bash is pinned by contract tests that execute the fences under a POSIX bash. Reference: configuration, pilot, land.

When your agent keeps re-learning a lesson from a memory entry that is already correct, the problem is retrieval. /flow-next:audit now repairs that entry’s title, tags, module, and filing, so the next search surfaces it.

Detail

Nothing to switch on. The next /flow-next:audit run classifies this way on its own.

A recurring lesson already had a graduation path: Harden turns it into a lint rule, a CI step, or an instruction-file rule, so it fires on its own instead of riding the context window. That path needs a rule a machine can check. A lesson stated as judgment - a convention, a naming call, an ordering constraint - has no such rule, so a correct entry that kept coming back landed on Keep, and the next run re-learned the same thing while the entry sat unread in the store.

The audit now reads that pattern as a findability problem and gives it the third answer. When an entry recurs, resists mechanization, and carries a nameable defect in how it is found, the audit classifies it as an Update with a retrieval fix: it repairs the entry’s title, tags, module, and applies_when, and moves a misfiled entry into the category the lesson belongs to, since a category-scoped search never reaches an entry filed somewhere else. Your report counts these inside Updated, with the retrieval fixes named, so a run tells you how much of the sweep went to findability rather than content.

The repair stays in its lane. The retrieval rationale never rewrites what an entry says. An entry that also carries plain reference drift - a renamed path, a dead link, a stale snippet - still gets that repaired on its own evidence, in the same Update, and you keep the same interactive confirmation you had before.

Under the hood. No new status and no new field. The signal is the same write-side recurrence scan Harden already reads - ## Update headings and entry-file commits - never a usage count, because nothing records that an entry was read during a run. The named defect is required: recurrence alone does not trigger the fix, since the recurrence counters only grow and a repair is itself a commit the next scan would count, so a recurrence-only branch would churn the same entry’s metadata once per audit forever. A placement move is a git mv into the new category directory plus the frontmatter category set to match, with every related_to naming the old entry id re-pointed in the same edit.

Waiting through review used to mean waiting through rounds: fix three findings, re-review, get three more, repeat. The first review round now draws three reviewers at once, each reading for a different kind of defect, and hands you one merged list to fix in a single pass.

Detail

Nothing to enable and nothing to upgrade. On the codex and host review backends the change is already the default the next time you run /flow-next:work or /flow-next:impl-review.

Before, a review round was one reviewer’s single look at your diff, and what it happened to notice set your next hour: you fixed those findings, sent the diff back, and the next round surfaced a different set the first pass had walked straight past. The findings were real, so the loop felt productive while it was mostly re-reading the same code.

The first round of a scope now runs three concurrent draws of the same reviewer, each carrying one axis lens: correctness and logic, contracts and consistency (do the docs, tests, and stated promises agree with what the code actually does), and integration with the code you did not touch. The coordinator merges the three finding sets into one, drops duplicates, ranks what is left, and runs a single fix pass over the whole thing. Most of what used to trickle out across serial rounds is in front of you in the first merged one.

Two pre-registered evals set the shape rather than an intuition. A single review pass surfaced roughly 45% of the validated finding pool, which makes one reviewer a sample rather than a sweep. The union of three axis-differentiated draws reached 1.56x single-draw recall against a pre-registered 1.5x bar, at flat validity, so the extra findings are genuine defects rather than noise.

The bounds are worth knowing before you rely on it. A clean diff that would have shipped in one round pays roughly 3x review tokens for findings one draw would have surfaced anyway. And roughly a third of validated findings eluded every draw, so round 2 shrinks rather than disappears.

Which is why the topology is yours to steer, in the invocation, in a sentence. There is no flag and no config key to learn. Say use 1 reviewer instead of 3 on a small clean diff and the round collapses to a single draw with the token cost that goes with it. Say use three different model families for the review fan-out before a high-stakes merge and each draw routes to a different family, so blind spots decorrelate by model as well as by axis. Reviews are optional to begin with, so all of this only ever applies to a layer you already chose to switch on. The recipes are on Steering the fan-out.

Under the hood. The merged round consumes one review round against the deterministic cap, not three, and the receipt records each draw honestly. Re-review after the fix pass stays a single dispatch carrying the full merged finding container, because verifying fixes needs continuity rather than breadth. Scope is the codex and host backends: rp keeps its single stateful chat, copilot and cursor keep single dispatch every round, and completion review, land, and external PR bots are untouched. On the codex backend a cross-family fan-out keeps the primary draw on codex while secondary draws may name codex, copilot, or cursor; on the host backend the per-draw pins are unconstrained. Partial failures fail open from whichever draws returned a verdict.

A mixed backlog no longer forces a choice between a blanket skip-planning flag and hand-running the small stuff: you mark the individual spec, and the loop builds it straight through while everything else stays planned.

Detail

If you drive pilot over a backlog, the no-plan route used to be an invocation flag - /flow-next:pilot --no-plan applied to whatever spec the tick happened to select, which made it useless the moment your backlog mixed trivial fixes with real features. That flag is gone. The decision now lives on the spec itself, next to the ready flag it resembles: flowctl spec set-no-plan fn-N, or --no-plan at capture time, records the human judgment “decomposing this would convert no unknown,” and a ready spec carrying it goes through pilot straight to work’s direct route - one implicit task, no plan or plan-review stage, every review gate and receipt unchanged.

The control stays yours in both directions. No autonomous path ever sets the field; setting it is refused once a spec has tasks; clear-no-plan undoes it any time; and a stale marker on a spec that later got planned is inert - the planned tasks run, with a one-line notice. Scripts still passing pilot’s old flag degrade safely: an unknown-flag notice, and the affected specs simply route through planning, the safe default. Direct /flow-next:work fn-N --no-plan works exactly as before.

Under the hood: the field mirrors the ready flag’s lazy contract (absent reads false, idempotent toggles), surfaces on every JSON read surface including noPlan on ready --all rows, stays flow-local (never tracker-projected), and work-rolling’s refusal of the route now covers the field as well as the flag. See the migration note.

A captured spec can no longer put words in your mouth: the strongest provenance tag now means “findable in the quoted evidence”, and capture’s own process rules can’t masquerade as things you asked for.

Detail

When capture turns a conversation into a spec, every line it writes carries a provenance tag - your words, a restatement, or the agent’s own fill-in - so you can see at read-back what you actually asked for. A field report showed the strongest tag leaking: a close restatement could be stamped as your words, and after the duplicate check agents sometimes invented a “user turn” out of capture’s own process rules (“this is a new spec, don’t mark ready”) and wrote it into the spec’s boundaries as if you’d said it.

Both holes are closed. Your-words now means the exact content is findable in the quoted conversation evidence the spec carries - a restatement is labeled as one. Evidence lines must quote things you actually typed; process rules stay process. A correction you give during the read-back edit loop becomes evidence before the redraft, so your own fresh words are never downgraded, and when a large capture splits into several specs, each one is checked against its own evidence slice. The read-back now verifies all of this before asking you to approve.

4.10.1 - Unattended landing sees today’s Codex

Section titled “4.10.1 - Unattended landing sees today’s Codex”

A converged PR no longer stalls waiting for a human just because the review bot changed how it says “all clear” - and skills stop hand-rolling their own memory dedup.

Detail

Codex’s PR reviewer used to post a “Didn’t find any issues. Reviewed commit: …” comment when a review came back clean. It now delivers the same verdict as an edited-in-place summary table naming the reviewed commit. The land loop’s clean-review gate only knew the old phrasing, so a PR with green CI, zero open threads, and a demonstrably clean review of the current head still ended in “needs human” - the one outcome an unattended ship loop exists to avoid. Land now recognizes both forms. The safety posture is unchanged: the comment must come from an allowlisted automated reviewer and name the current head commit, a stale verdict never counts, and setting the pattern to an empty string still switches the comment path off entirely. Repos that installed before this release are covered without any action - the old default stored in .flow/config.json upgrades itself at read time, while customized patterns are never touched.

Two smaller fixes ride along. Skills that file recurring memory notes (like the feature map’s drift memos) now use one deterministic flowctl memory upsert call - it updates the existing note when the title matches exactly, creates one when nothing matches, and refuses to guess when two candidates exist - instead of each skill hand-rolling its own find-or-create logic. And in repos with a markdown formatter, /flow-next:features runs it over the map files before finishing, so seeding the map no longer produces follow-up “formatter artifact” commits.

4.10.0: Navigation knowledge that survives the run

Section titled “4.10.0: Navigation knowledge that survives the run”

Live verification stops paying the same navigation tax on every pass. A new committed map at .flow/features/ records, from the user’s point of view, how a user reaches each feature, how an agent drives it, and which traps waste a run - and /flow-next:qa and the driver read it when it exists. /flow-next:features seeds the map with every route proven by one live drive before it lands, then keeps it honest with an audit-shaped maintain pass on your own cadence. The spec still says what to prove, and captured live evidence is still the only ship basis.

Detail
  • The map is committed, and it is navigation only. .flow/features/ sits beside .flow/memory/ because its whole value is surviving the session and the machine - deliberately the reverse of /flow-next:map’s local .clawpatch/ code index, which stays local-per-developer. An index carries the operating rules (baseline preconditions, driving conventions, proof standards) so a cold agent can drive from the map alone; each feature file carries a **Surface:** identifier and exactly four sections: Sub-features, How to get to it (user POV), Driving it, Gotchas.
  • Nothing undriven lands. Seed interviews the repo rather than you - surface, run command, drive mechanism, observable evidence, isolation - then proves each route with one live drive before writing it. A partial seed lands the proven features and names the failures instead of discarding the run; a pure library, a host with no usable driver, or a checkout that will not build is refused with the reason rather than mapped from guesswork.
  • Maintain is audit-shaped, and its edits stay in their lane. Index hygiene, one read-only source reader per feature, reconcile, one live pass over every feature even when source looks clean, then triage: wrong user-POV description is doc drift (fix the map), working behavior the harness cannot drive is a harness gap (fix it and re-drive before shipping), broken app behavior is a product bug (report it, keep it out of the PR). Outcomes are clean (no branch, no PR), changed (one chore PR of proven map and harness corrections, never a merge), or blocked with the blocker named. Product code is never edited.
  • Doctor gates every drive. One read-only check - right build, port owned by this run, valid auth - before the first drive, on each fresh session, and after any failed drive. Never drive an instance this run did not start, never kill by process name; an orphaned port from a crashed prior run blocks with a reclaim instruction rather than being reaped.
  • Consumers cost nothing when the map is absent. QA and drive gate on directory existence only - no config key, no registration - so a repo without .flow/features/ behaves exactly as it did. A stale route QA finds is filed as a feature-map-drift memory entry for the next maintain pass; QA never edits the map mid-run.
  • Cadence belongs to you. /flow-next:features is never a pilot stage, a land tail step, a Ralph iteration, or a post-merge hook: any autonomy marker in the environment refuses the run. Every invocation ends with a typed FEATURES_VERDICT= line, so a host loop such as /loop 1d /flow-next:features can read the outcome.

4.9.1: Reviews judge the work, not the paperwork

Section titled “4.9.1: Reviews judge the work, not the paperwork”

Two review-rubric fixes. A workaround wearing a well-written justifying comment no longer reads as “well documented”: the reviewer now treats that comment as a flag on the underlying code, and rewriting the comment without fixing the code does not resolve the finding. And a review bot can no longer hold your merge hostage on process ceremony - decisions recorded in the spec are settled, and checklist or handoff paperwork is feedback, never a merge gate.

Detail
  • Comment-as-alibi is a named finding class in the implementation and standalone review prompts. Severity is judged from the workaround, not the prose; the fix is the code, or the constraint encoded as an assert, a test, or a lint rule. Licensed comments are never flagged: license headers, external-constraint notes, lint suppressions with reasons, public API contracts, issue links - the same keep-list the worker’s authoring rule has carried since 4.8.0, so author and reviewer never disagree.
  • Recorded decisions are settled. The plan-review and completion-review prompts gain the rule the implementation review already had: a finding that re-litigates a decision recorded in the spec’s Decision Context is FYI, never blocking. All three prompts now also state that process-compliance observations - checklist ceremony, dogfood records, handoff paperwork - never gate a merge. The recommended land.reviewTrigger text tells external bots the same up front.

4.9.0: Start work without the planning stage

Section titled “4.9.0: Start work without the planning stage”

A fully-known spec can now go straight to implementation as a recorded choice. /flow-next:work on a spec with no tasks used to fall through and could end green having built nothing; it now forks - plan first, or work directly - with a recommendation judged from that spec’s size and blast radius. Say “skip planning” or pass --no-plan and the ask never appears; the direct route mints one minimal implicit task and runs the same pipeline, so evidence, review, and receipts hold unchanged. The same release removes an unexplained permission ban that had silently kept dispatched workers from spawning their own subagents.

Detail
  • You choose, and the choice is recorded. The fork’s recommendation is judged per spec (size, independent surfaces, blast radius) with its reason stated - no static default. Contradictory signals ask instead of guessing, and the flag on an already-planned spec is ignored with a one-line notice.
  • The fast route keeps every contract. The minted task is deliberately minimal (“implement this spec”, satisfying every R-ID, its review contract pointed at the parent spec) - never a copied plan. Receipts, impl review, done evidence, and the single-task completion-review skip compose unchanged, and the mint is atomic: two concurrent runs can never double-implement the spec.
  • Autonomous loops keep planning. Without an explicit no-plan instruction, a zero-task spec under autonomy stops with a typed report instead of asking; pilot forwards an explicit --no-plan to its work dispatch and never decides it; Ralph stops typed on the new needs_tasks signal instead of spinning; /flow-next:work-rolling refuses the route - a single task degenerates the rolling frontier.
  • Workers can dispatch subagents again. Three writing agents carried a disallowedTools: Task denial from their first commit with no recorded rationale - Task is the subagent-dispatch tool, not a planning feature. The audit removes it from the writers, keeps it on every read-only agent with the reason written inline, and the direct route’s worker uses the restored capability under a judicious license: parallel implementation, research, scouting, shape chosen at execution time, every subagent joined before commit.
  • Under the hood: flowctl next surfaces zero-task specs as status: plan, reason: needs_tasks; flowctl task create gains --require-empty-spec (the atomic mint guard); the guide, capture, and interview routers can recommend the route for near-zero-risk fully-known specs; the Codex mirror’s transform roster and guard grew to cover every new copy-pasteable command.

4.8.0: The autonomy prose stops failing quietly

Section titled “4.8.0: The autonomy prose stops failing quietly”

Autonomous runs get thirty-four hardened rules across the worker, the land conductor, and the review rubrics. Closing failure classes banked from real overnight runs: budgets burned against stale state, gates trusted on narration instead of evidence, locks that outlive their tick, and cleanup that could sweep away work a human left uncommitted. The pass was pressure-tested on its own PR: fourteen cross-model review rounds forced twenty-nine further repairs before merge.

Detail
  • Workers can no longer grade their own homework. A gate, test, or baseline is never edited to make it pass; an assertion is never weakened to match a wrong implementation; an errored or wrong-surface gate observation records as inconclusive, not green. And a suspiciously fast or zero-case pass gets its log read before any green receipt is minted, so one false pass can’t poison every later run that honors the receipt. Deliberate baseline updates stay legal: a pin update the task’s acceptance names ships in the same commit with its rationale.
  • Your uncommitted work is safe from the machinery. Both the pause path and blocked-tree cleanup commit only what the run itself produced; pre-existing uncommitted changes are left in place and named in the handoff, never swept into a pause commit or reverted away.
  • The land conductor spends its budgets on real problems. Red-CI triage reads merge state and open threads before burning a fix attempt; a repeat identical failure is reclassified instead of re-run blind, and an infra-shaped repeat escalates instead of consuming the fix budget; flake signatures bind to the head they diagnosed; siblings re-gate after any base-moving action; spec dependencies are honored at merge, not just at select.
  • Ticks stop losing state to each other. Each land tick claims its clone atomically before the first ledger read. Owner-aware, so a live tick is never reaped by the age threshold, an idle tick never leaves a lock behind, and a session reclaims its own abandoned claim. Dry runs take no claim and mutate nothing.
  • Reviews name their evidence. A five-label evidence scale (claimed / cited / walked / executed / reproduced) plus new judgment probes: structure-over-instruction, wire-type leakage, legacy dual-path, the shallow-module smell with its falsifiable sign, and a mechanical 1000-line-crossing check capped at Should-Fix.
  • Questions the machine can answer never reach you. Interview and plan resolve empirically answerable forks with a throwaway probe. But only when the probe is non-mutating or fully disposable; a stateful question still asks. Wildly divergent independent opinions trigger a reframe, never an average.
  • Also: the setup-installed docs snippet now carries the standing prose-contract line for chat replies (sentinel v2; existing repos pick it up on their next /flow-next:setup run).

The agent now drafts substantial replies under the contract without being asked, on every host. /flow-next:prose’s description was reshaped from user phrases to the agent’s own drafting moment, and Codex joins the ambient behavior - the skill enters the implicit catalog with a dieted entry after a measurement showed the shared catalog sitting under half its budget, so the earlier explicit-only carve-out was protecting headroom that was never at risk. Manual invocation with a draft to tighten is unchanged.

4.7.0: Agent-written prose answers to a contract

Section titled “4.7.0: Agent-written prose answers to a contract”

PR bodies, tracker comments, spec prose, done summaries, and changelog entries now draft against one shipped ten-rule contract instead of whatever register the model reached for, so filler that could describe any project, feelings standing in for numbers, and invented outcome lines get caught while the text is being written rather than in review.

Detail
  • Nothing to switch on. All 23 durable emission points carry a one-line pointer to the contract file and read it at the moment they write: make-pr bodies, capture and plan spec prose, interview write-backs, tracker-sync and resolve-pr comments, chart briefings, strategy sections, qa findings, land verdict comments, prospect candidates, prime glossary definitions, audit memory entries, and worker done summaries. The pointer is non-blocking - a host that cannot find the doc proceeds unchanged.
  • Your structural contracts still win. Where a rule collides with the shape of the surface being written, the surface wins: tracker dedup markers stay first-line and byte-unchanged, projection-only source truth is never overridden, and outcome-first ordering never invents an outcome that was not in the payload.
  • The scope is deliberately narrow. The contract governs how artifact prose reads. It makes no claim about code quality or maintainability, and cites SlopCodeBench (arXiv 2603.24755) for exactly what that evidence shows.
  • It caught its own author first. The first enforcement pass found the contract’s own bridged author breaking rule 9, the em-dash ban, in the contract’s opening sentence. The review gate held and the sentence was rewritten before it shipped. Two further bot-review rounds grew the pointer set from 19 emission points to 23.
  • New skill: /flow-next:prose extends the same contract to substantial chat replies on request - opportunistic rather than guaranteed. Plain-language triggers (“tighten this reply”) work on hosts that match skill descriptions (Claude Code, Cursor, Droid, Grok); Codex is explicit-only, the same carve-out visual gets. The visual digest itself stays excluded from the contract, because its output is ephemeral chat rendering rather than a written artifact.
  • Scout reports stop asserting absence for free. A scout that reports something is missing now states the search it ran to conclude that, so you can judge the negative claim instead of taking it.

Under the hood. The Codex mirror’s docs-link rewrite and its hard-fail link guard now cover agents/*.md, so agent-file pointers mirror to paths that resolve instead of shipping dangling.

4.6.1: The assumption that fell to a five-minute probe, twice

Section titled “4.6.1: The assumption that fell to a five-minute probe, twice”

Cursor and Grok Build now run the rolling-frontier scheduler for real instead of quietly degrading to waves. The scheduler’s prose named both hosts as blocking-dispatch by assumption; a five-minute probe on each (dispatch two sleeping agents, watch when control returns) measured non-blocking dispatch with per-completion notifications, and both hosts then drove a full rolling run end-to-end - simultaneous three-task admission, each task integrated the moment it returned while siblings kept running. The clause now binds on measured dispatch behavior with dated provenance, never on host name; a genuinely blocking host still degrades honestly, unchanged.

4.6.0: Excused reviews close the loop, reviewers stop re-running the world

Section titled “4.6.0: Excused reviews close the loop, reviewers stop re-running the world”

Two field-report sweeps in one release: a completion review that policy deliberately excused is now a recorded state every gate honors - excused specs ship instead of looping - and both reviewer prompts gained a verification budget, so a review round re-checks what a finding disputes instead of re-running the suite the run’s final gate already owns. Plus three more field fixes: land’s post-merge bookkeeping survives its own ignore rules, dead-end review modes are refused before any work happens, and OpenCode users can run every command a closer prints.

Detail
  • completion_review_status: not_required (#371, thanks @sn-furali). Work’s single-task policy skip records its decision instead of leaving unknown; the tracker projection, flowctl next --require-completion-review, pilot routing, and Ralph’s completion gate all decide through one satisfying set {ship, not_required}. Unrecognized values still fail closed; the write is a compare-and-set from unknown (a real verdict is never overwritten) that also refuses a surface no longer single-task; adding a task, rewriting the plan, editing a task’s contract, or resetting a task re-arms the review.
  • Reviewer verification budget. Both reviewer rubrics now assign focused, evidence-named suites (plus the exact test a finding disputes) to review rounds and the full suite to the run’s final gate - closing a measured tail-latency mode where a re-reviewer ran a ten-minute full discover on a rolling run’s critical path. Token deltas in the rebaseline evidence are recorded measurements, never an enforced ceiling.
  • Land’s sync-state commit survives (#367, thanks @sn-furali). The post-merge tail no longer names the auto-ignored receipts directory in its git add; the sidecar commit is guarded on a staged diff, so an unchanged sidecar (including resume-tail re-entry) is success.
  • --review=export refused at parse time (#366) at both mouths - work’s surfaces and impl-review’s own parser - with a pointer to plan-review, where export actually lives.
  • Host-correct closer commands (#364). Every closer that prints a copy-pasteable next step emits OpenCode’s flat /flow-next-<name> form there and the colon form elsewhere; the Codex mirror’s rewrite guard gained a positive expected-output check so a reworded literal fails the sync instead of silently staling.

New installs stop paying a reconciliation pass after every completed task: on most specs the automatic plan-sync pass finds nothing to change, and /flow-next:sync keeps the full capability for the moment a task genuinely invalidates a downstream assumption. Existing configs keep their setting, and setup still asks the question with the trade-off explained. A welcome side effect: the rolling-frontier beta’s prerequisite (planSync.enabled=false) is now the default state, so a fresh repo can invoke /flow-next:work-rolling directly. Opt back in with flowctl config set planSync.enabled true.

4.5.0: Multi-task runs finish ~50% faster at the same quality bar (beta)

Section titled “4.5.0: Multi-task runs finish ~50% faster at the same quality bar (beta)”

/flow-next:work-rolling is an experimental beta variant of /flow-next:work: instead of pausing at wave boundaries until the slowest in-flight task finishes, it admits the next ready task the moment any task returns - each in its own isolated worktree, with the conductor reviewing every return under the unchanged canonical review rules. The architecture was picked by a pre-registered three-arm eval: this arm cut work-phase wall-clock 52.1% at quality parity with zero uncontained correctness incidents - and the eval rejected a faster arm, which is the part worth reading.

Detail
  • Same inputs, same workers, same review. Invoke it exactly like /flow-next:work. Reviewer identity, rubric, diff scope, the SHIP-before-done gate, and the fix-loop cap are canonical work’s, byte-unchanged; the concurrency cap stays at 3. Workers share an outside-tree run-notes surface they read by pointer - notes content is never embedded into a dispatch prompt.
  • User-invoked only. Pilot, land, and Ralph always dispatch canonical /flow-next:work; the beta never enters an autonomous loop on its own.
  • Prerequisite: plan-sync off. planSync.enabled=true (the shipped default) fail-closes the run to serial, canonical behavior - rolling admission needs flowctl config set planSync.enabled false. An interactive run offers that command once; an autonomous run only reports it and never mutates config.
  • The eval that said no. A shared-checkout arm was faster still - 69.2% - but failed quality parity: roughly 30% thinner test artifacts. The mechanism was structural, not random: making the declared-paths list the commit boundary disincentivized new test files (tests are the artifact that spawns new files), so part of the speed was bought with under-testing. The per-task worktree pool removes that pressure, which is why the slower-but-parity arm shipped.

Capture and plan tell you the smallest sufficient next step at the moment you choose it. Field-requested by a flow-next team: the pipeline-variations selector (risk and unknowns, never size) was documentation - and documentation doesn’t fire at the decision point. Now each closer prints one advisory Recommended next: line above its unchanged menu, judged from the spec or plan just produced. The menu stays a menu; autonomous runs are untouched. And the PR’s own external review pass hardened the Codex distribution chain along the way.

Detail
  • Capture (base, rewrite, and per-spec split footers) judges the just-written spec: open [inferred] criteria and parked unknowns lean interview; real design risk leans plan; near-zero risk leans a minimal plan with plan-review typically ceremony. Legal targets are interview/plan/guide only - chart is upstream of capture and never recommended; genuinely conflicting signals recommend /flow-next:guide.
  • Plan adds the same line on the interactive next-steps menu only (re-judged after every go-deeper round): plan-review vs straight to work, with a review skip legal only for the two documented ceremony shapes. steps.md is untouched, so autonomous output carries no recommendation.
  • Single rubric home: the closers judge against pipeline-variations and link it - no copied rubric, no config keys, no persisted classification.
  • Installed Codex hardening (from this PR’s six-round external review): docs now ship to Codex installs under an owned docs/flow-next/ namespace - the installer can never delete or overwrite non-owned files (sentinel regression test); every mirror docs link resolves on disk or is an absolute URL (hard-fail closure guard); actionable commands render in Codex’s $flow-next-* syntax; and 7 long-broken mirror links from earlier releases are repaired.

4.3.1: Stop paying twice for the same tests

Section titled “4.3.1: Stop paying twice for the same tests”

When a review sends work back and the fix passes the full gate, the pipeline now remembers that instead of running the identical suite again a few minutes later - and on a multi-task spec, a task can inherit the previous task’s proven-green baseline.

Detail

Nothing to enable and nothing to re-run - this is default behavior from 4.3.1 on.

A NEEDS_WORK verdict used to cost you the test suite twice: once when the fix loop proved its fix good, and again when the Verify stage re-ran the same commands against the same commit a few minutes later. The fix loop now writes the same green receipt the pipeline already trusts, at the committed fix HEAD, so Verify honors it rather than repeating work it can prove already passed. That is 2-10 minutes back per NEEDS_WORK to SHIP cycle; mining real receipts found 15 to 18 affected tasks in a typical repo.

The second half is the baseline. On a multi-task run, worker N+1 can be handed the previous task’s result instead of re-proving a tree nobody touched - but only when the prior Verify ran the same Quick commands green, HEAD has not moved, and the two tasks’ file sets are disjoint. The worker records where the baseline came from, so the provenance stays in the receipt.

Both halves fail closed, deliberately. A focused suite never mints a full-gate receipt, any doubt means the gate simply runs. The first task of every spec still runs its own baseline - that is the check that catches a red your local environment has and CI does not, which is exactly how this shipped (a test file red at base here while CI was green). Lint always runs.

4.3.0: OpenCode joins the first-class roster

Section titled “4.3.0: OpenCode joins the first-class roster”

If OpenCode is your daily driver, flow-next now installs once from the canonical repo and the whole pipeline follows - planning fan-outs, cross-model reviews, receipts - riding every release. The stale community port is superseded and archived.

Detail
  • One installer, no port to chase. OpenCode has no plugin format, so ./scripts/install-opencode.sh scatters the canonical files into ~/.config/opencode/: skills as-is, generated agents whose tool denials translate to OpenCode’s permission map, flat /flow-next-<name> command stubs, and the support dirs that make flowctl and the spec-template cascade resolve unchanged. A deterministic ownership manifest scopes every re-run and --uninstall to installed paths only - a colliding user directory aborts the install instead of being deleted.
  • Setup runs like every other host. The ownership manifest doubles as setup’s platform-detection signal, so /flow-next-setup writes AGENTS.md instructions in the flat slash form - never Codex-shaped snippets.
  • Host review works, and it earned a hard rule everywhere. OpenCode subagents take their model from their own agent definition or inherit the session model - there is no dispatch-time override. Pinning a tier’s model is one 5-line user agent file; verified live, the conductor matched the routing block’s reviewer model to the pinned agent unhinted and the receipt recorded the real reviewer. Without a pin the degraded reviewer self-reports and the review fail-closes rather than letting the session model grade its own work. And on every harness, the host backend now carries the rule the name implies: it never shells out to another CLI.
  • Verified on opencode 1.18.19: 29/29 skills and 20/20 subagents discovered with correct permission maps, a full plan scout fan-out, codex-backend review, and host-backend review - all driven end-to-end from an OpenCode session.

4.2.2: Runs finish sooner, reviews check the same things

Section titled “4.2.2: Runs finish sooner, reviews check the same things”

First batch of fixes from a measured wall-clock pass over the whole pipeline: small specs stop paying a duplicate review, multi-task waves stop paying repeated plan-sync passes, and plans that could run tasks in parallel stop silently running them one by one.

Detail
  • Single-task specs skip the completion review when it would re-review the same diff - the spec’s only task already holds a per-task SHIP and covers every requirement. The skip is recorded as an explicit stage line, never a silent absence; multi-task specs keep the full completion review, whose cross-task integration value is the point.
  • Plan-sync runs once per resolved wave with the full completed-task set instead of once per task - same downstream scope, same per-task drift verdicts, k× fewer dispatches.
  • A missing Touches: line is now a plan-review finding on multi-task specs (omission silently forces serial dispatch - waves are fail-closed on the line), and pilot’s evidence echo repeats the work stage’s Sequential fallback: reason so a driver loop can see when a spec ran serial and why.
  • Parallel agents in sibling worktrees: flowctl’s runtime claim state lives in the git common dir and is shared by every worktree - now documented, with the FLOW_STATE_DIR override (set it outside the repo tree) for concurrent same-repo pipelines.

Setup prices each opt-in before you answer it. The review-backend question now says where the time goes (each review round is a serial pass the pipeline waits on - usually the largest wall-clock item in a run), the None option spells out what still gates a run and what stops being checked, and every host’s menu offers Host - the reviewer that keeps every gate with no second CLI, configured by one reviewer: line in the routing block. Memory, plan-sync, GitHub-scout, and HTML-artifact questions carry the same one-line cost shape. The full dial from a cross-model backend down to host or none is priced in Running Lean.

4.2.0: Land asks the human reviewer when it is their turn

Section titled “4.2.0: Land asks the human reviewer when it is their turn”

Teams whose merge gate is a human - a code-owner review required by a ruleset, sharpest when the PR author is a GitHub App that cannot be a code owner - no longer watch a converged PR sit idle until someone happens to notice it. Opt-in land.requestReviewers makes land request the right people at the one moment their review is the only thing left. Thanks @sn-furali (#359).

Detail
  • The ask fires on a sharp predicate, not “at convergence.” CI green, zero unresolved threads, and a human review as the sole missing merge input - so it never deadlocks under reviewSignal: approve (where convergence is the approval) and never notifies a human about a PR that merges the same tick.
  • The list is yours: csv of GitHub logins and/or org/team slugs and/or the literal codeowners - the codeowners token rides the draft→ready flip and GitHub resolves the owners itself (no local CODEOWNERS parsing). The PR author is filtered out; “ready” keeps meaning “a human may review this now.”
  • Exactly once per PR per head SHA. The request is recorded in the land ledger and claimed atomically, so overlapping ticks cannot double-ask; a CI-fix push moves the head and re-asks only if the human’s review is again missing - a genuine re-request, not spam. A failed request records the head anyway and surfaces reviewers=failed:<reason>, never a retry loop.
  • Never a merge gate. land.reviewSignal still decides what counts as reviewed; the key only asks. Default off - with it unset, every gate, action, and ledger write is byte-identical, and the evidence line gains one additive reviewers=off field. --dry-run reports would-request and mutates nothing.
  • Part 2 of the report (land.draftOnChangesRequested) is deferred until this has proven itself in the field.

Pilot loops finally start inside a git worktree, stop refusing branches that carry merged gate PRs, and your tracker’s workflow states are checkable without touching config. Three field reports, each verified against 4.0.0 before fixing - thanks @sn-furali (#354, #355, #356).

Detail
  • Plan stages run where you stand (#354): pilot’s plan and plan-review stages demanded a checkout of the default branch, which git refuses in any secondary worktree - so the documented one-worktree-per-spec setup could never start the loop. The rule now checks the actual hazard instead: an open PR on the current branch. No open PR means plan in place; only a branch with an open PR still steps aside, and only a failed step-aside asks for a human.
  • Merged gate PRs are history, not an inconsistency (#355): a team opening one PR per gate (capture, plan, review) reaches make-pr with several merged PRs and unshipped work on the same branch. Pilot called that state inconsistent while make-pr’s own rules called it normal. Pilot now compares the branch head against the newest merged PR’s head - squash-merge-safe, unlike counting commits against the default branch - and refuses only when the branch truly holds nothing new.
  • flowctl tracker wire list-states (#356): a read-only verb listing Linear workflow states or Jira statuses (id, name, type) with an explicit complete flag, so a truncated page can never masquerade as the whole board. Answers “does every configured state id still name a live state?” without a second tracker client - and without tracker resolve, which repairs the mapping and writes config. Detection never writes; repair stays where it was.
  • A failed review round can no longer wedge a spec’s reviews forever. When a backend returned no parseable verdict, flowctl refunded the round but left bookkeeping that could never finish - every later review of that spec then refused with REPLAY_REQUIRED, with no way out short of hand-editing .flow state. Refunded rounds now clean up after themselves, and repos already wedged by an earlier version heal silently on their next review. Found dogfooding this release’s own pipeline.

4.0.0: Your repo stops carrying flow-next around

Section titled “4.0.0: Your repo stops carrying flow-next around”

Installs are copy-less. Setup no longer copies the CLI, the agent guide, or the spec template into your repository, so updating the plugin is the entire update - no per-repo re-run, no stale copy quietly running last month’s behavior. If an older install left .flow/bin/, .flow/templates/spec.md, or .flow/usage.md behind, delete them; nothing reads them. And the second thing you used to maintain per project is gone too: choosing which model does what is now a short block of your own words in your own instruction file, not a configuration ceremony that stored model names it could not keep true.

Detail

Do these first (all optional - nothing breaks if you skip them):

  1. Delete your copies. If your repo has .flow/bin/, .flow/templates/spec.md, or .flow/usage.md, remove them (git rm if they are tracked). Nothing reads them any more, and a stale copied flowctl can shadow the current one. /flow-next:setup offers to delete them for you, and /flow-next:plan prints a one-line nudge when it sees them.
  2. If you used delegate:codex, run /flow-next:setup and accept the model-routing scaffold, then drive the other CLI through the bridge recipes in flowctl usage. The packaged delegation subsystem is gone.
  3. Leftover config keys are inert, not dangerous. work.delegate* and models.* keys still sitting in .flow/config.json are named once in a non-blocking advisory and otherwise ignored; delete them when convenient.

The docs-snippet schema did not change in this release, so nobody has to re-run setup to stay current.

Nothing lives in your repo except your work. Previously most projects ran in “copy mode”: flowctl, the agent guide, and the spec template landed as snapshots under .flow/, and every plugin update meant re-running setup in each project - or silently running an old CLI, which is the failure mode behind most “this flag should exist” reports. Now every host resolves flowctl from the plugin install itself: Claude Code and Factory Droid through their plugin-root environment variables, Codex through its own home, and Cursor and Grok by deriving the plugin root from the absolute skill path those hosts already hand the agent. Update the plugin and you are done. /flow-next:setup is now for the first run, a configuration change, or the rare release that says the docs-snippet schema bumped.

Routing is something you say, not something you configure. Setup used to ask which model should play which role, probe a CLI for the model ids it served, and store what it found - configuration that claimed what was installed, and then aged into configuration that lied. That ceremony, the role map, and its staleness stamps are gone. In their place: four named tiers - reviewer, implementer, fast scout, thinking scout - and a routing block in your own CLAUDE.md / AGENTS.md, written in your words with model names you can verify against your own account. Every session reads it, including unattended pilot, land, and Ralph ticks, and applies it with judgment. If you never routed anything, nothing changes: each tier falls back to the session model, exactly as flow-next has always run out of the box.

A tier says which model executes a stage, never which stages run. Planning, capture, interview, every verdict, and the worker stay on the session model unless you say otherwise. Where the harness exposes it, a stage records the model that actually ran, so a routing preference written in prose leaves evidence instead of a hope - and unavailable provenance is recorded as unknown, never as the configured value.

Can I even do that here? has an answer per harness now. The new reach pages state, for each supported host, which mechanisms exist (in-session model, in-host subagent, another CLI over a bridge), which do not, and what the degradation is when one is missing. Skills ask for a tier and name no spawn primitive, CLI flag, or vendor path; a harness flow-next cannot identify resolves to the generic page and says so.

What is gone. The packaged codex-delegation subsystem: six work.delegate* config keys, the delegate role pin, the delegate:codex / delegate:local work arguments, the delegation gating and result-classification path in the work skill, two flowctl codex subcommands, and setup’s delegation option. The standard work loop is unchanged - delegation was off by default, and the host always owned gating, git, review, and commit. Two rules carry over to the bridge route and are not optional: the bridged child writes code while the host keeps git, judgment, and the verdict.

Under the hood: the plugin-root derivation is one added rung in every skill’s resolution chain - environment variable first, the skill file’s own absolute path second, the legacy .flow/bin/flowctl third as a silent backstop for repos that have not deleted their copies yet. The backstop is not going away, but nothing is documented, tested, or designed around keeping copies. flowctl setup-mode is gone from the CLI, the help text, and the docs; old setup_mode / setup_version stamps in .flow/meta.json are tolerated as inert metadata and never read.

3.34.0: Seven field reports, six issues closed, zero new config keys

Section titled “3.34.0: Seven field reports, six issues closed, zero new config keys”

A two-day batch of forensic-grade reports, every claim verified before fixing. Locked-down repos get a merge-identity seam and a tail that finishes; land loses its force-push; fresh clones can finally pass validate; a phantom lock holder stops eating 10 seconds. Thanks @sn-furali (five reports) and @TechupBusiness.

Detail
  • FLOW_PR_MERGE_CMD (#337): env-only merge-identity seam mirroring the create side’s, with a stderr-verbatim contract so land’s race-vs-refusal judgment survives interposition.
  • Server-side catch-up (#342): gh pr update-branch replaces checkout+rebase+force-push - commit SHAs survive, removing the cause of orphaned evidence; repeated refusals escalate bounded.
  • The post-merge tail finishes (#345): release and board update run before the one push a PR-only base can refuse - a refused push is a bookkeeping note, not a stuck board.
  • Fresh-clone validate is meaningful (#347): committed-snapshot status warns instead of erroring (582 structurally-false errors on flow-next’s own clone became warnings); real inconsistencies still fail.
  • done names the receipt it wrote (#346): modified_paths + a dirty-file advisory; the commit-done-stage ordering is canonical on every documented surface.
  • Honest lock diagnostics (#340): a lock that can never be created fails fast with errno and path; “holder appears alive” is reserved for an owner that was actually read.
  • The attempt ledger answers “which model reviewed this?” (#338, recording half): rows carry the model/effort that actually ran; the publication half was declined - flow-next never writes its own verdict into the channel land treats as independent-review evidence.

3.33.0: Independent tasks finally run at the same time

Section titled “3.33.0: Independent tasks finally run at the same time”

When a spec’s tasks do not depend on each other, flow-next can implement them side by side - and does, but far less often than your task graph allows. The rule needs each task to declare which paths it will touch, and the planning guidance said to skip that line whenever it was hard to predict, so a default behaved like an exception. It is written on every task now, and two independent tasks that took 187 seconds one after the other take 96 seconds together, on the same tokens.

Detail

Nothing to enable and nothing to configure - plan a spec and the tasks carry the declaration. What changed is the default: planning writes a Touches line on every task and declares wider rather than omitting when it is unsure. Previously how often waves fired came down to how boldly a given planning pass read the omit-when-unsure advice, which is why the same feature could run in one repo and stay dormant in another. Both choices err the same way (a declaration that is too wide overlaps a sibling and keeps the tasks serial, exactly like omitting it), but only the declared one can ever become a wave once the overlap turns out to be false.

A wave still refuses unless everything lines up: same spec, at most three tasks, no dependency path between any pair in either direction, declarations that do not overlap, and no task touching the always-serial set (flow state, lockfiles, migrations, generated output, spec and task files). Any doubt sends the whole wave serial, which is the old behavior. Getting a declaration wrong is cheap on purpose: workers write in isolated workspaces, so an overlap shows up as a merge conflict when their commits are joined, costing one serial re-run and never correctness. Troubleshooting now has an entry for exactly that moment.

Two things that broke the first live wave are fixed where the conductor reads them. A wave workspace is branched from a commit, so a spec that was just planned and not yet committed does not exist inside it, and every parallel worker fails to re-anchor while looking like a broken worker. And each worker’s handover evidence names commits that live only on its own workspace branch - recorded unchanged, every finished task ends up pointing at commits that disappear when the workspace does, which validation then reports as orphaned. Both are now stated preconditions.

Under the hood: measured before shipping by running the path end to end - two workers dispatched concurrently into linked worktrees, both honoring their declarations exactly, both deferring completion and review to the conductor, join clean, both tasks verified done. 96s concurrent against 187s serial, 16s of join overhead, no token penalty. Both preconditions are pinned by tests on the canonical prose and the generated Codex mirror.

3.32.2: Cursor’s in-IDE browser drives for real

Section titled “3.32.2: Cursor’s in-IDE browser drives for real”

On Cursor, agents skipped the built-in browser because the instructions were wrong. They now probe it by id, and if that probe fails they ask you once to type @Browser (no space) so the pane is connected, then try again. A vanished MCP mid-run is a partial pass, not “the rung does not exist.” Console and network from the driven page stay unverified, so a QA pass routed here must record BLOCKED (could not verify) rather than PASS.

3.32.1: Model guidance catches Grok 4.6 on day one

Section titled “3.32.1: Model guidance catches Grok 4.6 on day one”

Grok 4.6 shipped 2026-08-12; within a day the setup scaffold, bridge recipes, and orchestration docs carry an evidence-based reprofile - built from independent benchmarks and real user reports, not the launch post.

Detail
  • The model-routing scaffold’s grok tier moves to grok-4.6: intelligence up (Artificial Analysis Index 61, tied with GPT-5.6 Sol Max; real-user consensus ~Opus 4.8-tier), raw speed down but ~2x turn efficiency, taste unchanged.
  • Routing sharpened along the split the independent evals exposed: supervised editor-shaped implementation is its strong surface (CursorBench 69.9%, day-one Cursor with a permanent 2x usage pool - bridge slug cursor-grok-4.6-high verified live); long unsupervised terminal loops are its weak one (Terminal-Bench v3 26%).
  • The never-the-gate posture stands, now with numbers: AA-Omniscience measures it inventing ~1/3 of the time when it doesn’t know; API cache reads cost 67% more than 4.5 on long sessions.

3.32.0: The plan you can read at a glance - and the skill that reviewed itself

Section titled “3.32.0: The plan you can read at a glance - and the skill that reviewed itself”

After /flow-next:plan, reviewing meant reading a spec plus seven task files and rebuilding the structure in your head. /flow-next:visual restates it as one screen of compact markdown - and its own dogfood pass caught a real grounding bug before it shipped, while the PR that shipped it carried the first diff-fenced structural sketch in place of an edge-less mermaid diagram.

Detail
  • /flow-next:visual - point it at a spec, a task, a git range, or the conversation; it restates the thing with a fixed 8-shape vocabulary (call trees, file trees, diff-fenced sketches, type sketches, tiny tables; mermaid last) in plain fenced blocks that colorize natively on every host. Read-only, chat-only.
  • Post-plan digest (the primary mode): thesis, task tree in dependency order, planned file-layout diff with owning tasks annotated, R-ID coverage line where uncovered requirements jump out, IS/IS-NOT boundaries.
  • Hard grounding: every path from task files/spec/git diff --name-status, every edge from real code or a real dependency, coverage from declared satisfies frontmatter - never invented; missing state degrades to the nearest viable mode with a one-line notice.
  • Offered where the text walls are: capture, plan, and interview closers suggest the digest at their read-back moment - an option you pick, never auto-run.
  • make-pr learns the sketch: ## Structural changes may emit a diff-fenced file-tree or call-tree sketch where mermaid is weakest (cap-forced collapse, or a diagram under four nodes) - same hallucination guardrails, no silent-rendering-failure risk.
  • Review dividend: the PR’s bot-review round exposed a long-standing sync bug flattening fenced-block indentation across the whole Codex mirror - fixed fence-aware, 159 mirror files restored.
  • On Codex the digest is explicit-only ($flow-next-visual) by design - its trigger-rich description stays out of the shared skill-catalog budget.

3.31.0: Your repo’s own command can now gate the merge

Section titled “3.31.0: Your repo’s own command can now gate the merge”

On a free-plan private repo, branch protection and rulesets 403 - no required status check can exist - so land’s gate tree read a server with nothing to say, and a repo-local gate of record had no way to bind the merge. Thanks @TechupBusiness for the precise gap analysis.

Detail
  • land.mergeVerdictCommand (opt-in, fail-closed): once every other gate passes and the planned action is merge, land runs the repo’s command once - exit 0 merges; missing, unexecutable, timed out, or signal death blocks with NEEDS_HUMAN and never skips.
  • Context arrives as environment only (FLOW_HEAD_SHA, FLOW_BASE_REF, FLOW_PR_NUMBER, FLOW_SPEC_ID); the command string is never built from PR-derived text.
  • The verdict binds the (head, base) pair it judged: a push after the verdict refuses server-side at --match-head-commit; a base that moved re-ticks; a trust guard refuses to execute from a non-base checkout.
  • --dry-run reports would-run and executes nothing; unset/null/empty all mean off.
  • Also documents gate classify’s known fail-open (CI-guarded generated docs, #334) with the conductor-prose remedy.

3.30.0: Codex reviews get their verdicts back

Section titled “3.30.0: Codex reviews get their verdicts back”

review.backend codex could return no verdict 13 times in a row while every probe said the backend was healthy - the reviewer subprocess was inheriting instructions never meant for it (the host repo’s auto-loaded AGENTS.md, and the plugin’s own coordinator skills saying “never self-declare a verdict”). Thanks @sn-furali for the isolating-controls table that separated the two routes.

Detail
  • Persona override on for codex: the counter-instruction preamble (cursor-only since fn-90) now rides every codex review; the codex False was inherited from a refactor, never decided.
  • Host project docs suppressed: -c project_doc_max_bytes=0 on both the fresh and resume codex exec dispatch - the reporter’s measured fix.
  • The plan-review prompt states its role: it was the only review prompt missing the “You ARE the reviewer” anchor, which is why plan reviews failed first.
  • Honest failure classes: a healthy exit-0 run with no verdict journals as missing_verdict (the old ladder matched the word “timeout” in the reviewer’s own prose), and a streak of them terminates with instruction-contamination guidance instead of “repair the backend”.
  • Repo review invariants belong in .flow/criteria.md, which rides the review prompts to every backend - now documented.
  • Release-gated on the reporter’s own reproducer shape: a live plan review on flow-next’s own large AGENTS.md returned a verdict on the first attempt.

3.29.0: The worktree kit works without a keyboard

Section titled “3.29.0: The worktree kit works without a keyboard”

cleanup - the kit’s only sanctioned removal path - read two answers from stdin unconditionally and died silently with no terminal; create left the new branch tracking the base, so under push.default=upstream a bare git push aimed at the base branch. Thanks @sn-furali.

Detail
  • cleanup [<name>...] [--yes]: names skip the prompt, --yes skips the confirmation (required off a terminal), EOF-guarded reads fail loudly naming the remedy. Interactive behavior unchanged.
  • create --no-track: an isolated per-worktree branch no longer tracks the base’s remote branch; first push under default config needs -u.
  • Invocation story told honestly: the five phrase-triggered skills answer to plain language and - on hosts that surface skills as commands - their full skill name, e.g. /flow-next:flow-next-worktree-kit.

3.28.1: Your memory entries, back exactly as you wrote them

Section titled “3.28.1: Your memory entries, back exactly as you wrote them”

A memory entry edited with memory add --update could silently come back different: a mid-string " #" was written unquoted (YAML comment syntax - any conforming parser truncates the value there, while flowctl’s own fallback reader hides the damage), and without PyYAML installed, quoted list items containing commas were split apart and the mis-parse written back. Both fixed, with the one damaged entry in flow-next’s own repo repaired. Thanks @sn-furali for the forensic report.

Detail
  • Writer: any whitespace-then-hash scalar is now quoted; non-comment hashes (C#, issue#140) stay plain.
  • Reader (no-PyYAML fallback): flow lists and mappings split quote- and depth-aware via one shared helper - quoted scalars recognized at item starts, after mapping key separators, and at nested collection boundaries; double-quoted keys/values unescape correctly.
  • Deliberately not done: teaching the fallback to strip comments (it would turn latent damage into active truncation for repos with unquoted " #" already on disk).

3.28.0: The strikeout you can actually recover from

Section titled “3.28.0: The strikeout you can actually recover from”

On a board-armed repo, a spec that struck out of the pilot loop could read ready everywhere a human looks while staying permanently invisible to pilot - and the docs described a recovery that has been impossible since fn-87. A measured three-phase report proved both halves. Thanks @sn-furali.

Detail

Nothing to enable. Default behavior.

What was hard before. Pilot’s two-strike guard takes a failing spec out of selection. With tracker.readyState armed, the next pull re-readies it from the board - but since fn-87, deliberately, that projection-set ready never clears the strike: the board echo re-grants readiness with nobody acting, and clearing on it would re-dispatch the same failing spec every tick forever. The docs still described the old behavior, and the skill’s own escape clause (“an explicit re-ready, not a projection echo”) had nothing to key on - the reporter measured that a deliberate out-and-back board move is byte-identical to an echo in every durable artifact. The only real escape was hand-editing an undocumented file under .git/.

What you get now. flowctl pilot strikes list shows what is struck and why; flowctl pilot strikes clear <spec-id> (or clear --all) is the recognized human recovery - atomic, shared across worktrees, with a distinct not-found for unknown ids. The strike 2/2 verdict names the command, so the transcript carries its own way out. Clearing a strike never changes spec readiness: strikes are pilot state, the board keeps owning readiness.

And the docs tell the truth. Every surface that claimed a board move clears strikes now states the fn-87 rule, troubleshooting documents the ledger, and the board-native alternative (clearing when a tick observes the issue leave and re-enter the ready lane) is recorded as a deferred decision - it narrows but cannot remove the ambiguity the reporter proved, and silently misses a fast out-and-back between ticks. (#325)

3.27.0: The tracker bridge stops racing, lying, and losing Projects

Section titled “3.27.0: The tracker bridge stops racing, lying, and losing Projects”

Two agents promoting the same intake issue could each end up with their own spec; a dedup query against a populated Linear board could look healthy while blind; and a Linear issue could not be placed in a Project at all. The tracker-bridge batch closes all three and writes down the abandon path. Thanks @sn-furali for the measured reports.

Detail

Nothing to enable - Project placement is per-spec opt-in; everything else is default behavior.

One winner per candidate. The create-first mint claim is now compare-and-set: sync create-first-put --if-absent records the minted spec only while the claim slot is free, and the loser of a concurrent promotion exits with a distinct conflict naming the winner to adopt - instead of silently overwriting it. The tracker-sync ceremony wires the CAS into the canonical path, teaches adopt-the-winner (retire the duplicate with spec close, never re-put), refuses a claim when the record is already promoted and cleared, and defers the whole collision to a human under autonomous operation. (#310)

A refusal you can handle, not a lie you cannot detect. Linear wire list-open with tracker.readyState unset returned an empty success - indistinguishable from a genuinely empty board. It now returns an explicit error naming the unresolved key and how to set it; leaving it unset remains a valid, deliberate configuration, and backlog automation treats the refusal as “no ready lane configured”. (#311)

Issues land in their Project. Optional per-spec sidecar fields tracker.projectId / tracker.projectMilestoneId are sent on issue creation and reconciled on every sync push. Absent means unmanaged: payloads stay byte-identical and a Project set on the Linear side is never cleared by flow-next, which carries exactly the id it is given and never creates or manages Projects. Verified live against the Linear sandbox, including the never-clears contract. (#315)

And the abandon path is written down. For a candidate that will never be promoted: close the issue in the tracker first (that side is yours), then clear the local record - ordering stated so a live intake issue is never left without a trace. (#309)

3.26.0: Coverage that tells the truth twice

Section titled “3.26.0: Coverage that tells the truth twice”

A fully-planned spec that had not shipped code yet looked like 0% coverage - and make-pr refused to open the draft with advice you could not follow. And a rebase could orphan every evidence commit a spec recorded while validate stayed green over the dead links. Coverage now answers the plan-gate and merge-gate questions separately, and validate tells you when history rewrites have voided your evidence. Thanks @sn-furali for both measured reports.

Detail

Nothing to enable. All default behavior.

Two coverage questions, two answers. The export payload gains undeclared_r_ids - criteria no task claims at any status - beside the unchanged uncovered_r_ids (criteria no done task evidences). make-pr’s coverage abort now fires only on undeclared coverage, the one state where “go declare coverage” is advice you can act on. A plan-gate spec renders honestly: the coverage table gains a third state (claimed, not yet evidenced beside evidenced and undeclared), the warning marker belongs only to genuinely unclaimed criteria, and the ratio stays evidenced-only with the claimed/undeclared counts appended when non-zero. (#301)

Evidence links stop dying silently. A rebase, amend, or squash-merge leaves recorded evidence SHAs present in the object store but unreachable from HEAD. flowctl validate now warns per orphaned commit (“recorded value left as-is”) while reachable commits stay silent and tokens that are not commits in this repo - tracker UUIDs, foreign SHAs - are ignored by design: flagging them would corrupt exactly the evidence the record exists to hold. Nothing is rewritten, the run never fails, and the whole pass costs two batched git reads regardless of commit count, so the land loop can keep calling validate freely. make-pr marks orphaned SHAs as annotated text instead of rendering commit links that 404. (#302)

3.25.0: Six measured reports, six fixes, nothing silent

Section titled “3.25.0: Six measured reports, six fixes, nothing silent”

A criterion your spec wrote should never vanish without a trace, a wrong platform guess should not survive on the primary host, and a half-resolved tracker map should never read as resolved. Six field-reported defects fixed in one pass - each with the reporter’s verified repro as the acceptance fixture. Thanks @sn-furali.

Detail

Nothing to enable. All default behavior.

Criteria stop disappearing. The export parser now reads title-form (**R14 - title**) and parenthetical-form (**R15 (note):**) acceptance criteria and keeps text wrapped across lines - the reported repro parses 5 of 5, and 25 criteria were recovered across flow-next’s own specs. Anything still unparseable is counted and surfaced as acceptance_criteria_residue in the export payload, so a short coverage denominator is visible instead of silent. Suffixed R-IDs (R4a) are accepted by the PR cognitive aid’s validator too - the last straggler of the grammar widening. (#300, #303)

Setup guesses right. A single SPEC.md on a case-insensitive filesystem no longer counts twice and prints a bogus both-files warning (files are counted by inode now). And Claude Code - the primary host - detects as Claude Code: the old signal never reached a plugin skill’s environment, so setup fell through to the Codex fallback and wrote the wrong snippet syntax. The cascade now keys on CLAUDECODE paired with the Claude plugin manifest, positioned so hosts that prove themselves with their own signals still win over the inherited marker. (#305, #306)

Tracker resolution finishes the job. tracker resolve --select used to persist just the slot you picked, leaving the rest unfilled while the scope stamped fresh. It now runs the normal assignment over the remaining slots and persists the union; a map still missing a required slot is kept but reported as CONFLICT and left unstamped, so a later plain resolve repairs it. Configs already half-stamped by this bug self-repair on the next --select. (#308)

Repair is not takeover. flowctl start --reclaim rewrites a task’s claimant when it is held by a stale or wrong identity, recording Reclaimed from <identity> (identity repair) - distinct from --force, which keeps its takeover meaning and note. Only the claim-ownership gates relax. (#316)

And one promised answer. The gate-classify path taxonomy is deliberately closed to config: per-repo gate policy belongs in your conductor instructions (CLAUDE.md / AGENTS.md), with pilot.gateClasses as the open vocabulary, and classifier reason strings are not a stable contract. (#313, docs)

3.24.1: Judgment stays on a judgment model

Section titled “3.24.1: Judgment stays on a judgment model”

Two field-reported routing papercuts on non-Claude hosts. The plan skill’s gap analyst and judgment scouts could silently run on the host’s fast default when the Claude model alias in their agent files wasn’t resolved - the plan prose now states they run on the session model there (scanner scouts may still ride the fast tier). And setup could write a Cursor model pin with an id your account doesn’t serve (ids vary per account - composer-2.5 vs composer-2.5-fast): pins are now restricted to ids seen verbatim in this run’s probe output, and a failed probe writes no pin.

3.24.0: Review verdicts you can audit, not just believe

Section titled “3.24.0: Review verdicts you can audit, not just believe”

A resumed reviewer session can answer from its previous round’s context in about a kilobyte, with zero tool calls, while asserting “measured” facts that happen to be true - and the verdict text is indistinguishable from a real review. Every review attempt now records how its verdict was produced: how much output it cost, how many tool calls it actually made where that could be measured, and exactly which commits it judged. You ask “was this measured?” of the ledger instead of taking the narration on faith.

Detail

Nothing to enable. This is default ledger behavior. Every new attempt row carries the fields; nothing about how reviews run changes.

What was hard before. The attempts ledger recorded what was reviewed - backend, verdict, output hash - but not how the verdict was produced. A reporter measured resumed review sessions returning SHIP in 1.1-1.6 KB with zero tool calls, stating facts that were true but answered from the previous round’s context while claiming fresh measurement. Verdict-text inspection cannot catch this by construction: the fabricated verdict states true facts, and the resumed session even reuses the same thread id. Only work volume separates a review that measured the repo from one that remembered it.

What you see now. Every new review_attempts[] row records output_bytes (always - the size, never the output itself), tool_calls where the codex event stream let the dispatcher genuinely count them (a recorded 0 is the signal itself; plain-text paths carry no key at all), head_sha_observed marking whether the reviewed commit came from a pre-dispatch snapshot or the finalize-time fallback, and base_sha beside head_sha wherever the review snapshot ran - so the judged diff can be located and re-rendered. flowctl review-rounds attempts --json surfaces all of it.

Absence means unknown, never zero. Rows written by older versions carry none of the new fields and read back untouched. The tool-call count is measured only from a genuine codex exec --json stream - a review from another backend whose prose happens to quote codex-shaped event lines never gets a fabricated count - and a crash between the write-ahead journal and the ledger write replays the row with its measured provenance intact.

What this does not change. No verdict-validity rules, no re-review policy, no reviewer behavior change, no new commands. The consumer asks the question; flowctl only makes it askable. Fixes #312 - thanks @sn-furali for the measured report.

3.23.0: Status answers say where they came from

Section titled “3.23.0: Status answers say where they came from”

A status read that answered from a stale snapshot looked exactly like a right answer - a review sandbox once burned three review rounds arguing with a spec that was already done. Status output now names its source, the pre-work commands warn when your checkout is behind, and reviewers are told task lifecycle is not theirs to judge from committed files.

Detail

Two ways a wrong answer used to dress as a right one. Task status lives in a runtime store shared across your worktrees, but when that store is not reachable - a fresh clone, a review sandbox scoped to the diff - flowctl fell back to the committed snapshot without saying so; a reviewer reading that snapshot marked a finished spec NEEDS_WORK at full confidence, three rounds running. And the checkout itself can be behind the shared truth, so the commands you run just before starting work answered from yesterday’s state without comment.

What changes for you: flowctl show and list now carry status_source on every task - flow-state when the authoritative store answered, committed when you are reading a snapshot that may be stale - and plain output adds one advisory line when runtime state is absent entirely. ready and anchor tell you once when HEAD is behind its upstream, because a wrong answer right before you start work is the most expensive one. Both shared review prompts now state that committed task files are snapshots and that a task looking unfinished there is never grounds for a finding.

What does not change: nothing fetches, nothing blocks, no freshness is enforced - a stale checkout still gets its computed answer, now qualified. The high-frequency polls (list, status, next) deliberately gained no upstream check, so the fast-poll performance stays intact; the advisory costs one read-only git probe on the two pre-work commands only.

Under the hood: provenance is stamped at the single merge point and stripped on every persisted write; the upstream probe is one git --no-optional-locks status spawn; 24 behavioral tests include spawn-count assertions locking the hot-path exclusion. Fixes #304 and #307 - thanks @sn-furali for the measured reports.

3.22.0: Handovers point at the work, not a retelling

Section titled “3.22.0: Handovers point at the work, not a retelling”

When one agent finished and handed off to the next, it wrote the story of what it had just done a second time - and that retelling started aging the moment anything moved, cost you a full re-read on every consumer, and could disagree with the files it was describing. A finishing agent now hands over pointers: the task, its status, where the summary and evidence live, what changed, and the verdict. Whoever picks it up reads the files themselves, which are the current truth, and every finished run ends with a next step you can actually run.

Detail

Nothing to enable. This is default behavior. Run the same commands and the handoffs get shorter and stop drifting.

What was hard before. A worker would finish a task, write its summary and evidence to disk, and then narrate the same thing back to whoever dispatched it: what it implemented, which files it touched, which tests it ran. Two copies of one story. The copy in the handoff was written from memory of the work rather than from the artifact, so the two could disagree - and when they did, the one you read first was usually the wrong one. It also meant the same content was paid for twice, once to write and once to read back, on every single task in a run.

What you see now. A finishing worker reports where the outcome lives: the task id, the terminal status, the paths to its summary and evidence, the range of commits it produced, and - where its path produced one - the review verdict. It no longer restates what it built. The conductor opens the files, which are the thing that actually exists, so the account you read is the account on disk. Workers running in a parallel wave also name the workspace they were given and their gate results, which is what the join needs to reconcile them. A return that restates content the files already carry now counts as a broken contract, not a stylistic preference.

The end of a run is a runnable line. The work skill’s final summary closes with a Next: line you can execute - open the pull request, or run QA first when your pipeline asks for it - instead of a description of what you might do next. The chart-to-capture handoff is the same idea: it hands you a paste-ready command carrying the briefing path, so you run the handoff rather than reconstruct it from a paragraph.

The doctrine is written down. The teams page now carries pointer-shaped handover as its fifth handover property, with the two carve-outs stated honestly rather than left as folklore. A consumer without access to the repo still gets content, because there content is the only transport available. And a bounded control signal - a verdict enum, an id, a strike class - stays inline and counts as a pointer, not a retelling, so a driver reading only the transcript never has to open a file to learn whether something passed.

What this does not change. The files were always the record; what changed is that the handover now defers to them instead of competing with them. Nothing about how summaries and evidence are written moves, and merge judgment stays where it was.

One reviewer reading your whole change has one pool of attention, and naming conventions are far easier to spot than a requirement quietly implemented wrong - so a run could come back with a tidy list of hygiene notes while a behavioral defect sat unremarked. The in-host quality audit now runs as two reviewers at once with separate jobs: one asks only whether the code does what the spec said, the other only whether it is code you would want to keep. You get both reports side by side, in full, and only the correctness reviewer can call something Critical or decide the change is shippable.

Detail

Nothing to enable. This is how the audit behaves by default. Run the same commands and the review phase reports back the way it always did, with two headings instead of one.

What was hard before. A single generalist reviewer had to hold the whole change and decide, in one pass, what mattered most about it. Hygiene findings are cheap to see and easy to justify; a silent regression of something an earlier task got right, or an assertion that was weakened rather than fixed, takes real reading to notice. When both compete for the same attention budget, the cheap findings win often enough to matter. The failure was rarely a bad review - it was a review that spent itself on the wrong axis and never got to the thing that would have bitten you.

What you see now. Two reviewers work the same change at the same time, each told exactly one question to answer. The correctness reviewer looks at spec conformance, places the spec was interpreted differently than you meant, silent regressions of earlier work, weakened assertions, security, and test coverage. The standards reviewer looks at simplicity, duplication, dead code, over-engineering, naming, vocabulary, and performance shape, against a rubric it carries with it. Both reports come back verbatim under their own headings. They are never merged, never reranked, and never summarized into a single list - which means you can read the correctness report first, and read the standards report knowing it is not competing for the same slot.

Severity has an owner. Only the correctness reviewer can raise a Critical finding or say the change is ready to ship. The standards reviewer is capped at Should-Fix by construction: it cannot mark something Critical, and it cannot issue a ship verdict. That is deliberate, and it is the part that makes the split worth having - a duplication complaint and a broken requirement should not be able to arrive wearing the same badge. If the standards reviewer does spot something it believes is outage-grade, it is not silenced: it hands the suspicion across as a short untiered note, outside its own axis, for the correctness reviewer’s judgment rather than as a verdict of its own.

Neither reviewer can flood the fix loop. Each axis has a hard limit on how many findings it may report, and when it has more than that it says so rather than padding the list. The point is that a review with thirty style observations does not turn into thirty rounds of fixing - you get the ones that reviewer judged most worth your time, plus an honest note that there were more. Merge judgment stays yours in all of it: the audit reports, it does not gate.

Honest bounds. The two reviewers are the same auditor given different instructions, not different models with different strengths - the gain here is undivided attention, not a second opinion from a second mind. And the standards axis is genuinely narrower than a generalist reviewer was: it will not tell you a requirement is missing, because that is not its job anymore.

Under the hood. The audit dispatches two axis-scoped runs in parallel; a dispatch that arrives without its axis line defaults visibly to correctness rather than guessing. Caps are eight tiered findings for correctness, five for standards, three Considers each, with any overflow declared in the report. Cross-axis handoffs from standards are limited to two untiered lines. Changes to reviewer behavior are now gated on replaying a banked corpus of reviewer regressions - two historical catches have to survive the change, and both did.

Three things kept costing you the same argument twice. A feature you had already refused on principle came back weeks later as a fresh proposal, because nothing remembered the refusal. Genuinely-open questions got written into specs as though they had been settled, so an unknown read as a decision. And the prose steering your agents leaned on capital letters where it should have stated a rule the agent could check its own output against. Now a refusal is recorded with its reasoning and every date it was re-asked, planning consults that record before it proposes scope, and only you can reopen a declined idea; unknowns get parked as unknowns until someone actually resolves them; and the shouting is replaced by rules that name what a failure looks like, plus about 60 new completion bounds that tell a procedure step when it is genuinely done.

Detail

Nothing to do. No config, no defaults, no commands changed. Update the plugin and the same commands behave the same way, with less scope drift and less half-specified content in what they hand you.

Your “no” now has a memory. The first time a feature is refused on policy grounds - not “we already have that”, but a real judgment call about what this product is - the refusal gets written down: what was declined, why, and a dated list of every time it has been asked for since. Planning reads that record before it proposes scope, so the idea you turned down in March does not arrive in June wearing a new name and consume the same conversation. Two boundaries keep it honest. Only you reopen a declined concept - an agent may surface that a decline is being re-requested, and it may not decide the answer has changed. And “declined because it already exists” never earns a file, because the record exists to hold judgment, not history. The recurrence list is the useful part: when the same request shows up for the fourth time, you can see that, and decide with the evidence in front of you.

Unknowns stop impersonating decisions. Specs used to have nowhere to put a genuine unknown, so unknowns got written up as content - half-specified, confidently phrased, and indistinguishable from a decision someone actually made. A spec can now park them in a section of their own, gated by a test that is easy to apply: if it can be decided now, decide it; if it is known work that just has not been scheduled, make it a task; only what is truly fogged gets parked. Parked items graduate into real sections the moment interview or planning resolves them, so the section empties as the spec matures. What you gain is the ability to read a spec and tell the difference between what was settled and what nobody knows yet.

Specs describe contracts, not file paths. A spec that named files and line numbers was accurate for exactly as long as it took someone to refactor, and then it generated churn: plan-sync chasing renames, reviewers reconciling paths that had moved, and a spec that read as wrong when the behavior it described was still exactly right. Specs now state types, signatures, and behaviors - the things that survive a move. One deliberate exception stays for decision-rich snippets where the location genuinely is the decision. Tasks are untouched and stay path-bearing: naming the files you are about to modify is a task’s job, and always was.

Rules an agent can check itself against. Hundreds of CRITICAL / MUST / FORBIDDEN blocks across the skills were restated as plain declaratives that describe the failure rather than raise the volume - “a SHIP verdict with no backend response behind it has broken this” tells an agent what to look for in its own output in a way that a capitalized MUST does not. The conversion was meaning-preserving line by line, and what stayed capitalized stayed for a named reason: literals that tests pin, fences that get executed, anchors that host mirrors transform, and blocks that evaluations guard. Alongside it, roughly 60 procedure steps across qa, map, setup, drive, work, pilot, land, prospect, and plan gained a Done when: bound that makes a demand - every item accounted for, every finding filed - rather than merely describing what finishing looks like. The practical effect is fewer steps an agent can consider complete while something is still outstanding.

The repo’s own glossary became a dictionary. What had grown into an encyclopedia is now a short list of the terms whose synonyms cause real ambiguity - twelve of them, each with one definition and the aliases to stop using (a spec, never an epic or a ticket or a story; plan-sync, never tracker-sync). The long-form text is archived rather than deleted. This is flow-next’s contributor vocabulary only; the glossary feature that /flow-next:setup seeds into your project is unchanged.

Routing stops going stale silently. The guide that recommends which workflow to run is now held to a rule: recommending a skill that no longer exists, or missing one that does, is a defect, and any change that adds or removes a skill has to update the router in the same breath. The rule found its first gap on the day it landed - arriving with no written direction at all now routes you to strategy instead of into a plan. Install instructions are also single-sourced from one place now, so the copy you read cannot disagree with itself depending on where you found it, and the judgment-boundary prose in plan, chart, and guide names the specific pull toward just doing the work yourself as the signal that you are standing on one.

Fixes worth knowing. The Codex host mirror’s plan rewrite had been quietly doing nothing since an earlier rework moved the prose its anchors targeted - Codex hosts were getting Claude-specific wording where multi-agent phrasing belonged. The anchors are repaired. A report skeleton in memory-migrate had drifted from its canonical copy and was missing two sections; the duplicate is now a pointer to the one source. Two reference files reachable from nothing at all are deleted.

Under the hood. Every conversion was checked against the test corpus and the host-mirror transform before the edit, and test pins were retargeted in the same commits rather than loosened. Prose-contract tests, per-skill conduct checklists, and the full suite gated each wave of the rewrite.

3.19.0: Skills read only the path you take

Section titled “3.19.0: Skills read only the path you take”

Every skill you ran used to hand your agent its whole instruction set up front, including the rules for branches that session was never going to take - and a model reading four sets of conditions it does not need is exactly where instruction-following quietly erodes in long prose. Ten skills now load a lean spine and pull a branch’s instructions at the moment they reach it: up to 63% less always-loaded prose where the modes are genuinely exclusive, 15-40% across the heavy skills, with every safety net and every-run contract still inline. You get cheaper sessions and better adherence on the long, condition-heavy skills - not faster runs.

Detail

Nothing to do. No config, no defaults, no commands changed. Update the plugin and the skills you already run behave the same, from a smaller starting load.

What changed for your session. Skills grew many-branched over dozens of releases of features people asked for, and the disclosure discipline flow-next started with did not grow with them. A chart run that is going to take one mode was reading all of them; an implementation review pinned to one backend was reading the prose for the others. Now what every invocation needs stays inline, and what only some paths reach lives in a reference the skill reads when it gets there. Where a config decides the branch, the probe that reads it is fail-open: if the probe cannot answer, you get the branch rather than silence.

What deliberately did not move. Every safety net, every every-run contract, and every calibration block stays inline where the agent always sees it - a skill that only sometimes remembers its guardrails would be a worse trade than any prose it saved. The refactor moved text verbatim rather than rewording it, and each of the ten skills was checked against a written list of observable behaviors before it shipped.

The discipline now has a keeper. Every skill carries a conduct checklist - four to six falsifiable observables that prose changes are reviewed and dogfooded against. That is the part meant to outlast this release: the reason the disclosure rotted the first time was that nothing failed when it did.

Two internal surfaces are gone. The rp-explorer exploration skill was dispatched by nothing, and planning always used repo-scout in practice - so context-scout is removed too, and with it an entire research-mode question branch from planning. Planning now has one codebase-research path instead of a fork you never chose consciously. RepoPrompt is unaffected as a review backend and remains fully supported everywhere it was before.

Fixes worth knowing. Interview’s doc-aware autodetect was fail-closed, so a probe failure silently switched doc-aware behavior off - it now fails open like every other gate. Make-pr’s inline mermaid recap had drifted to eight rules while the canonical checklist carries nine, including the subgraph/node-id collision rule that a real PR caught; the duplicate recap is gone. Codex hosts now receive the host-native review invocation on work’s reference paths too, not just the main phase file, so the wave-join and host-deferred paths stopped emitting a Claude-only slash command.

Running lean is now a documented choice. flow-next has always run fully as spec then plan then work, with everything else optional, but nothing said what turning a layer on actually costs you. The new Running Lean page frames two operating profiles - human-driven, where you are present and can be the reviewer, the tracker, and the QA; and autonomous, where those same layers are what replace you - and prices every optional layer with the same four fields: what it automates away, what it costs structurally, when it earns its keep, and the manual invocation if you want the capability without the standing cost. A layer you skip still leaves a record: stage receipts carry skipped(reason), readable with flowctl usage --stages <spec-id>.

Deprecated, not removed: Ralph and packaged codex delegation. Nothing is removed, no defaults change, and existing installs keep working - these are signals so new setups stop adopting a path intended for retirement. A host loop or cron calling /flow-next:pilot and /flow-next:land does Ralph’s job without the scaffold and the guard-hook registration; the setup model-routing scaffold plus the bridge recipes in .flow/usage.md cover what packaged codex delegation was built for. Both reference pages stay maintained for current users.

Under the hood. Prose-contract tests now pin content and reachability rather than file location, so a verbatim move to a reachable reference no longer breaks the suite while content disappearing or becoming unreachable still does. A new encoding guard keeps every reference file and its Codex mirror twin clean UTF-8.

3.18.0: The pipeline stops overbuilding your requests

Section titled “3.18.0: The pipeline stops overbuilding your requests”

Ask an autonomous planner for a small feature and it builds you a subsystem: in a replay campaign against real shipped work, two unguided runs of the same request each invented a 500-900-line risk-management layer nobody asked for. This release lands the disciplines that campaign measured, as one batch. Plans now bind to scope minimality - every task traces to a requirement, every requirement to your request, and overengineering is a review finding rather than a taste note; the guided replay arm delivered 43% fewer output tokens at 57% lower cost with reviewed quality above the unguided one. Tasks become lean delegation payloads that reference the spec instead of restating it. A pipeline stage that silently does nothing - the class behind the plan-sync bug that no-oped for weeks - now announces itself in the receipts it already writes. And same-spec worker concurrency runs on an explicit fail-closed rule instead of a judgment call that almost never fired.

Detail

Plans stop growing past the request. Every task must trace to a requirement and every requirement to what you actually asked for; capabilities nobody requested become one-line Boundaries exclusions instead of tasks, and planners are steered to eliminate risks structurally (a closed schema, an inert format, an unexposed capability) before building machinery to manage them. The discipline trims scope, never rigor: error-case enumeration and filesystem/permission/concurrency guards are explicitly exempt - an eliminated guard is not an eliminated feature. Both review rubric copies treat overengineering as a finding with three concrete patterns to flag.

Tasks carry the how, not a retelling of the why. Replay agents wrote tasks at roughly three times the fleet norm, and the bloat was paraphrased spec context that drifts out of date. Tasks now reference the spec’s requirement IDs and carry the concrete implementation plan - named files, approach, ordering, task-scoped acceptance - which is exactly what lets a cheaper implementer build without re-deriving design decisions. Nothing is lost: executors always receive the task together with the full parent spec. Tasks can also declare the paths they expect to modify, which feeds the new concurrency rule.

Planning documents are files, not heredocs. A plan that goes through review fix loops is edited in place with span edits instead of being regenerated wholesale into the command string - measured 13% cheaper in both replay A/B pairs. Spec examples are now the contract: the fields an example shows are exhaustive, closing a deviation class caught twice where an implementer “helpfully” extended a shown shape. Workers run focused tests while iterating and the full suite only where a gate already requires it - the campaign measured 54% of full-suite runs as redundant mid-loop re-runs.

No stage can silently do nothing. Every optional or delegated stage records ran, skipped with a reason, or failed with a reason in the receipts it already writes - a skipped stage is an event, never an absence, and a stage with no line is treated by review as failed. flowctl usage --stages <spec> summarizes them per spec, plain or JSON; malformed lines are counted, never a crash. No new stores, no dashboards, and token telemetry is explicitly out of scope.

Concurrency by rule, not vibes. 85% of 684 measured worker dispatches ran with zero overlapping sibling because the wave trigger was a judgment call. Now tasks run concurrently only when their declared write-surfaces are disjoint, no dependency path connects them, the wave is at most three, and nothing touches the always-serial set - anything missing or doubtful stays serial, exactly as today. A join conflict is never auto-resolved: the losing task re-runs serially and the collision lands in the receipt. Verified by a sequential-equivalence replay: wave and serial runs produced identical test outcomes and identical trees.

3.17.0: Say it was wrong, and know when it drifted

Section titled “3.17.0: Say it was wrong, and know when it drifted”

Two kinds of state you could not correct are now correctable. When discovery disproves the fact a chart started from, the decision that disproves it can carry the correction - and that correction travels into the briefing your next spec is written from, instead of the refuted claim sitting above the ledger that contradicts it. Separately, the flow-next-managed block in your instruction files can now be one of several in a single file, and a new read-only command tells you (or CI) whether anyone has hand-edited one, without writing anything to find out. Both shipped from field reports by @sn-furali (#292, #294).

Detail

Do this first if you wrote your own setup-block template. A template must now be exactly its marker-pair block: the BEGIN marker on the first line, the END marker on the last, a trailing newline, and no marker token repeated inside the body. Every template flow-next ships already conforms, so nothing to do for the standard install - but a hand-rolled template with a heading or a note outside the markers now fails with a clear message instead of drifting. It has to: prose outside the pair was written and hashed on apply but never seen by the comparison, so a block you had just applied could report itself as edited.

Charts can now admit they were wrong. A chart seeds ## Notes with the grounding facts it starts from, and discovery routinely disproves one of them - that is what discovery is for. Until now there was nowhere to record that. The note was write-once, so a refuted claim survived into the immutable briefing that /flow-next:capture reads, sitting above the ledger entry that contradicts it. Now the resolve that closes the decision can carry the correction with it: a dated bullet is appended to the notes, the original text is left byte-for-byte alone, and the next briefing renders both. Nothing becomes mutable - corrections are append-only and stamped by the tool, so the chart still reads as a record of what was believed and when.

The same report caught a quieter failure. A sharpen file with a key the tool did not recognize - including the notes key you would naturally reach for - used to be accepted and silently ignored, so a correction you thought you had recorded simply was not. Any unrecognized key now fails the whole resolve, names what it did not recognize and what it accepts, and does so before anything is written. A typo can no longer look like success.

Managed blocks are addressable, and drift is checkable. flow-next owns a marker-delimited block inside files you also edit, and until now it could track exactly one such block per file: pointing it at a second block in the same file overwrote the first one’s recorded state, so one of them lost its ability to tell “pristine” from “you edited this”. Blocks now carry an id, each with its own markers and its own recorded state, and operating on one never touches another’s - including a stray or corrupt one sitting in the same file. Omit the id and everything behaves exactly as before.

The other half is a read-only verdict. If you keep the block as an ordinary tracked file (setup_mode: copy), an edit to it is a normal reviewable diff - but there was no way to ask “is this still what flow-next generated?” without running the write path. The new check answers exactly that and writes nothing on any branch: clean exits zero, drift exits two, a structurally broken block exits three, so a CI job can gate on it with no jq and no risk. A block someone edited and then reverted reads clean, and a CRLF-only difference is never drift.

Under the hood. Recorded state moved from one hash per file to one per (file, block id); a hash written by an older version is read transparently as the default block’s state and upgraded on the next write, with no migration step to run. Review rounds on the way in tightened the fail-closed contract in three more places: keeping a customized block used to record that decision without ever validating the block, a duplicate or orphaned marker after the first valid pair escaped the corruption scan, and the check released its lock between reading state and reading the file, so a concurrent apply could skew its verdict.

3.16.3: The split path, hardened by running it

Section titled “3.16.3: The split path, hardened by running it”

A live end-to-end run of 3.16.2’s spec-split found what section testing could not: leftover copy that made “split” readable as “abort”, split specs authored after approval instead of shown before it, and a write step that silently dropped spec titles. All fixed - you now see every composed spec document before anything is written, and every copy of the workflow agrees on what the split answer does.

3.16.2: Capture tells you how many specs it should be

Section titled “3.16.2: Capture tells you how many specs it should be”

If you capture epics, briefing packages, or large features, the hardest question was never the content - it was “is this one spec or four?” Capture now answers it. Past 8 real requirements (standing rules and “tests must pass” items don’t count), or when the requirements clearly serve more than one shippable outcome, the read-back shows the actual split: proposed titles, which requirements go where, and how the specs depend on each other. One answer writes the whole linked set. The judgment is independence, not size - a big-but-cohesive spec is recommended to stay one spec, small captures see nothing new, and nothing ever splits without your say-so. Interview makes the same call when refinement outgrows a spec.

If you switched on plan-sync so completed work updates the tasks that come after it, that update was silently never happening. The work loop misread the task list’s JSON shape, the error went where nobody looks, and the empty result was indistinguishable from “nothing downstream to update” - so every run looked clean. Fixed, and a failed extraction now announces itself instead of impersonating an empty list. Thanks to the field report that caught it, real drift surfaced on the very first plan-sync run after the fix.

3.16.0: Your reviewer reads the whole change

Section titled “3.16.0: Your reviewer reads the whole change”

Cross-model review used to be handed a copy of your diff inside the prompt, capped at 50 KB. On a large change that meant the reviewer judged your code having been shown about a tenth of it, then went and read the rest off disk anyway. The copy is gone. A review now gets the commit range, the exact list of changed files, and the paths to the spec and tasks, and reads whatever it needs from your checkout. Nothing is trimmed to fit, and a review that cannot read its evidence now stops instead of returning a verdict based on nothing.

Detail

Set expectations honestly first: this makes reviews better informed, not cheaper. The prompt itself shrank by 83% on the release’s own largest review, but a reviewer that fetches spends turns on tool calls instead, and measured input tokens came out above the previous numbers. Most of that is cached, so billed cost does not track the raw figure, but no saving is claimed. A fetching reviewer is also slower in wall-clock on a big diff, which is why the dispatch bound moved from 600 to 1800 seconds - override with FLOW_REVIEW_EXEC_TIMEOUT if your changes are larger still.

What you get for that is a reviewer working from complete evidence rather than a truncated sample, and a changed-file list you can trust. That list is load-bearing now, and git abbreviates in three separate ways that all had to be switched off: --stat elides long paths behind an ellipsis, plain --numstat collapses a rename into {old => new} so neither real path appears, and without -z any non-ASCII filename comes back escaped. A scope map you cannot resolve to real paths cannot bound a review.

Re-reviews changed too. The reviewer now continues its own session instead of being re-briefed from scratch, so when it checks whether your fixes landed it is comparing against findings it actually remembers making. If a session cannot be resumed, the findings travel in the prompt exactly as before - the fallback is deliberate and loud, because a reviewer handed a lean prompt with no memory would silently produce a fresh blind review.

Convergence got the same treatment. Loops that were converging no longer get cut off and handed to you to verify by hand: the re-review prompt states the exact format for reporting prior findings, the parser accepts every token that format advertises, and the two rules that guessed at convergence from finding counts and severity trends are removed - they escalated three healthy loops in a row and caught no stuck ones. What remains is the reviewer explicitly calling the same finding unfixed twice, plus a round cap you own via review.maxIterations.

Under the hood. Removing the payload also removed everything built to make it fit: three prompt fitters, the 50 KB diff cap, the reviewer-facing “truncated to fit” markers, and the interim guard that distrusted an all-clear from a backend whose prompt could be shortened. That last one is retired by construction rather than gated, because no backend truncates now. One size guard survives, renamed to say what it is - a transport boundary for the one backend that delivers its prompt as a command-line argument, which refuses loudly instead of trimming.

This decision was made once before, in 3.2.0, and undone twice by later work that each had a good local reason. So it is enforced by a test that drives the real dispatch path and was verified to fail when a re-embed is simulated, plus pinned builder signatures so a new payload parameter cannot slip in under a different name.

If you work on Windows - or merge work from someone who does - a green build now means the same thing on all three OSes. The Windows CI leg used to skip six “incompatible” test files; a change could pass Linux and macOS while quietly breaking behavior hidden behind that filter. The filter is gone, and the bugs it was hiding are fixed.

Detail

Every one of the six excluded files turned out to be hiding a real defect, and none matched its recorded excuse. The failures were locale-dependent text reads (Windows writes cp1252 where production expects UTF-8), a POSIX-only permission check that crashed at import time on NT, and backend subprocesses that could block forever waiting on an inherited stdin no automated caller answers.

The infamous “900-second hang” was never the backend it was blamed on: the test runner killed only its direct child on timeout, then waited forever on output pipes held by grandchildren - on every platform. The runner now kills whole process trees (process groups on POSIX, Job Objects on Windows), reports leaked descendants instead of hiding them, and bounds its timeout diagnostics.

Concurrent tracker writes also stopped intermittently refusing legitimate paths on Windows - under load, path resolution can return two spellings of the same directory, which looked like an escape attempt. Containment is now derived from a single resolve, and the review cycle hardened it further: NTFS junctions (which are not symlinks and evaded the symlink check) are now rejected fail-closed during the containment walk, alongside pointer-width handle declarations for the Windows kill path.

Nothing was weakened to get there: no raised timeouts, no blanket platform skips, no broadened assertions. The full corpus runs green on windows-latest in parallel, serial, and shuffled order on the same commit as the Linux and macOS gates.

3.15.0: Flow on every task, without the tax

Section titled “3.15.0: Flow on every task, without the tax”

Running flow-next on small tasks used to cost more ceremony than the task: ~20 CLI calls to author a spec with tasks, a pile of file reads every time a fresh session re-oriented, and the one quality miss that kept repeating - the untested error path. This release cuts authoring to 2 calls, re-anchoring to 1, and moves error-case thinking to plan time where it is cheap.

Detail

Benchmark evidence made the overhead concrete: on identical work items, the flow-next pipeline spent 2.4x the wall-clock of a no-flow control, and the biggest blocks were ceremony calls and re-read context - fixed costs that dominate exactly the small tasks you most want tracked. The losses that were not overhead traced to a single pattern: an error path nobody enumerated, which a green test suite then certified forever.

Authoring is now a fast path. spec create --plan-file plan.md creates the spec with its plan in one call, and task create --from-json tasks.json materializes the whole task set in another - descriptions, acceptance, satisfies, dependencies (tasks in the same batch can reference each other by position). Validation is all-or-nothing: one invalid item rejects the whole batch with zero writes, so a half-created plan cannot exist. The granular verbs are unchanged and remain how you edit. The canonical spec-plus-3-tasks flow now measures 8 calls, down from ~20, and a test counts the real subprocess invocations to keep it honest.

A fresh session runs flowctl brief and is oriented: open specs with one-line goals, which tasks are actually ready (dependency-aware, with claim state), the last five completions with an evidence flag, the memory index, and pointers for going deeper. The output is deterministic and capped at roughly 2k tokens regardless of repo size - what gets dropped is marked, --full lifts the cap, --json is the machine form. Your context window stops paying for spec bodies you did not need.

New specs now carry their error cases in the acceptance criteria themselves - each criterion states its invalid-input and boundary handling, or says “no error surface beyond X” outright, so a reviewer can tell considered-and-none from forgot. The plan skill derives these during AC writing, the interview probes when they are missing, and workers treat every enumerated case as a required test before done. Existing specs are untouched.

Under the hood: brief performs no git calls and no writes (identical state renders identical bytes), the bulk-create path holds one lock per batch, and receipts, evidence, and start/done validation are byte-for-byte unchanged.

3.14.0: Review loops that end the way a human lead would end them

Section titled “3.14.0: Review loops that end the way a human lead would end them”

Converging review work gets its room, a loop that has stopped improving reaches you early instead of burning its whole budget, and a reviewer facing a genuine judgment call can hand it to you directly - evidence trail intact. This is the mechanism the 3.13.3 cap raise promised.

Detail

The review cap counts dispatches, and a dispatch counter cannot tell “genuinely stuck” from “nearly there.” Now the loop also reads the structured findings each verdict already persists and ends a measurably stuck loop early: an open finding chain two rounds fail to resolve, severity and count both failing to improve, or each fix introducing a fresh P0/P1 - exit 4 with ESCALATE: review loop stalled (<rule>), the same exit code and marker family drivers already handle.

Reviewers get a new terminal verdict, NEEDS_HUMAN - “a human must adjudicate,” distinct from NEEDS_WORK (fixable) and MAJOR_RETHINK (redesign). It writes a real needs_human status and its receipt before escalating, so nothing stops silently.

Repeat reviews stop wasting rounds on unchanged content: dispatching the same artifact that just received a verdict is refused before it costs anything (NOT_RETRYABLE, exit 1). The identity is per-surface - a completion re-review after an implementation-only fix dispatches cleanly - and SHIP or an explicit human re-plan starts a fresh epoch.

Each review surface now blocks on what it can actually break: plan review blocks only on findings naming a concrete bad downstream outcome, impl review treats recorded Decision Context decisions as settled, and the land-loop PR bot is scoped and triaged as a safety net, never a second gate. Autonomous loops cannot grant themselves more rounds - every new terminal is shorten-only, and Ralph blocks the reset commands and --force as human-only recovery.

If you run Codex from more than one home - a work account, a client sandbox, a second instance - flow-next could only ever live in one of them. Now it installs into whichever home you point it at, and each install stays self-contained.

Point CODEX_HOME at the home you want and run the installer once for it.

Three ways a chart could quietly hold the wrong state: a reopened chart with no route back to capture, a supersession that wired a replacement to the premise it had just superseded, and an ambiguous initial map that pointed edges at the wrong decision. None raised, none logged - each one persisted a chart that looked correct and answered every later question from the wrong state.

“Make our CLI more deterministic” reads like exactly the big unclear idea chart was built for, and chart would have taken it. It has no finish line, so the map could never close. Chart now names the test it was always applying and says no before spending a discovery pass on it.

One oversized idea wrapped in unknowns no longer has to become a half-guessed spec or a meeting that evaporates. Optional chart finds the route one decision at a time, then hands capture a briefing - and the short first-run path stays short.

Detail

Teams kept hitting the same gap: prospect ranked ideas, capture wanted intent you could write down, and between them large unclear efforts either got captured too early (a wall of inferred criteria) or lived only in conversations. Chart is the optional pre-capture route for that situation - not a new mandatory stage.

You describe the outcome in plain language. The agent grounds a bounded snapshot against the repo (safe citations only; nothing invents a resolved decision), reads back the smallest visible frontier and attended/unattended cost, and only then persists a chart. Work resolves one decision per session via evidence-first routes - research, probe, eval, prototype, interview, or an enabling task. Attended routes never self-answer under autonomous drivers (NEEDS_HUMAN). Prototypes attach a throwaway artefact before the human reacts; wrong turns stay struck-through via supersession so the briefing keeps the path that failed.

When nothing material remains to decide, a confirmed briefing package (one or more clusters, shared context named once) hands off to capture. Chart never writes a spec and never sits inside pilot. Guide recommends the smallest sufficient next step when you are unsure; tracker projection and pasted URL re-entry are optional conveniences over the local ledger, never the source of truth.

3.12.0: Your editor understands .flow/config.json now

Section titled “3.12.0: Your editor understands .flow/config.json now”

flow-next’s config file carries a published JSON Schema: editors validate and autocomplete every setting, scaffolded configs reference it automatically, and drift between the schema and what flowctl actually reads is a failing test, not a documentation promise.

The schema lives in the repo and at the stable URL this site now serves: https://flow-next.dev/schema/flow-config.schema.json

3.11.0: Trust and identity fixes for the autonomous PR path

Section titled “3.11.0: Trust and identity fixes for the autonomous PR path”

Two community-reported gaps closed: land can no longer mistake a PR that merely talks about flow-next for one it authored, and repos that require bot-authored PRs get a documented seam for supplying their own PR-create identity.

3.10.0: Standing team rules, checked on every spec

Section titled “3.10.0: Standing team rules, checked on every spec”

Write your project-wide acceptance criteria down once - “every route change regenerates the contract”, “no new dependency without a health check” - and the completion review you already run judges every spec against them, with the verdicts recorded in the receipt.

Detail

Every team has standing rules that outlive any single feature. Until now they lived in instruction files and reviewer memory: applied when someone remembered, invisible when they were not.

Flow-Next 3.10.0 gives them a home. Put one bullet per rule in .flow/criteria.md using the familiar requirement grammar (- **G1:** ...), and every spec completion review - on every review backend - judges each criterion against the whole implementation. The verdicts (met, violated, not applicable) land in the ordinary review receipt, so compliance is a recorded fact rather than a feeling. Violations also appear as normal review findings, so nothing new needs watching.

There is no separate audit pass, no rule engine, and no scoring - the reviewer that already reads your diff simply gets your standing rules alongside the spec. Repos without a criteria file pay nothing: not a token of prompt content changes until the file exists. Setup offers to scaffold a documented template on request, and declining leaves no trace.

The input boundary fails closed. A criteria file that exists but is broken - typo’d bullets in any Markdown style, duplicate or malformed ids, an unreadable file, a dangling symlink - surfaces a validation error before a review round is spent, instead of silently running the review without your rules. On the receipt side, ambiguous or contradictory reviewer output degrades the compliance array to absent rather than ever recording a wrong verdict, and the recorded ids must exactly match your configured criteria before anything attaches.

New plumbing, for the curious: flowctl criteria list --json validates the file; flowctl criteria prompt-block composes the injection for the RepoPrompt and host review paths; receipts gain an additive criteria: [{id, status, note?}] array documented beside the structured findings schema.

Full model: Standing Criteria.

2026-07-31: Front doors that lead with the problem (docs release, no version bump)

Section titled “2026-07-31: Front doors that lead with the problem (docs release, no version bump)”

The landing page and the README now open on the problem they solve and show the measured evidence for it, so you can judge Flow-Next in one screen instead of reading two thirds of a manual first.

A new Evidence page carries the full argument.

3.9.0: Pull requests that guide the review

Section titled “3.9.0: Pull requests that guide the review”

Reviewers get a guided journey through the change: what changed, why each step exists, what deliberately stayed untouched, where the risk sits, and which evidence supports each claim.

Detail

A large pull request usually arrives in file order. The reviewer has to rebuild the story: find the important decisions, work out which files belong together, separate meaningful changes from generated churn, and decide what still needs human judgment.

Flow-Next 3.9.0 does that preparation before handover. The pull request now walks through the change in logical steps. Each step explains its purpose, groups the files that implement it, links back to the relevant requirement or task, and names deliberate non-changes that protect the boundary of the work. Generated and mechanical files stay with the step they support instead of forming a distracting pile at the end.

The existing risk-ranked review plan remains independent from that journey. The walkthrough explains how the change fits together; the review plan points the human reviewer at the decisions and code that deserve attention. Evidence is attached to the claims it supports, so a reviewer can distinguish what the pipeline already proved from what still calls for experience and judgment.

Review findings now keep their identity across fix rounds. A finding can be traced to the review that raised it, its current status, and an optional snapshot-bound code location. Resolved, superseded, and current findings no longer blur together when the branch moves.

The same review journey can appear in the GitHub pull request and the optional local HTML view. Structured JSON carries the same meaning for downstream tools, including Flow Swarm, without forcing them to scrape prose or guess the order of the change.

Under the hood, /flow-next:make-pr stores this journey as a versioned changeWalkthrough, while review receipts can carry versioned structured findings. Existing receipts still work. Invalid or stale structured data falls back to the original reviewer prose, and parsing adds no model or network call.

RepoPrompt CE now supplies its review context and response through one direct Context Builder result. Fix rounds stay attached to that returned context and chat, with no hidden setup conversation or dependency on a Classic-style tab. Parser and render benchmarks enforce a strict <100 ms p95 ceiling over 30 warm runs.

3.8.0: Interview-written criteria now say where they came from

Section titled “3.8.0: Interview-written criteria now say where they came from”

You can finally tell which acceptance criteria came from a human and which the agent guessed, on specs that came out of an interview rather than a capture.

Interview now emits the same four tags everywhere it writes criteria.

3.7.0: Your own spec sections get filled in

Section titled “3.7.0: Your own spec sections get filled in”

Add a section to your project’s spec template and the interview now writes it for you, instead of leaving your heading empty or quietly ignoring it.

Detail

Do this first if you already keep a repo-root SPEC.md: put a scope marker under any section you added, e.g. <!-- scope: business -->. That one line is the difference between a section the interview fills and a section it leaves alone.

A project has always been able to override the spec scaffold with a repo-root SPEC.md - add a risk register, user stories, a rollout runbook, whatever your post-mortems justify. What was missing is that the interview passes only knew about the seven sections we ship, so your own headings sat outside the contract: nothing promised to preserve them, and nothing would ever fill them.

Ownership now comes from the section itself:

  • marker naming the pass you are running - written and refined, like any section we ship
  • marker naming the other pass - preserved exactly as-is
  • <!-- scope: both --> - written by either pass
  • no marker - preserved exactly as-is, and the read-back tells you it was skipped

One consequence worth knowing: a marked section is rewritable, so hand-written content under a marker you own will be refined by the next pass of that scope. Drop the marker to freeze it.

Also new in this release: the customization route itself is documented properly for the first time, including which four headings you must not rename (Acceptance Criteria, Boundaries, Goal & Context, Decision Context - renaming them does not error, it silently drops the feature that reads them). See writing specs.

Honest bound: we tried widening the default template with user-story and test-seam sections and did not ship it. The first measurement looked good, a pre-registered replication did not hold, and the wider scaffold ran about a third longer - which every implementer and reviewer downstream pays to read. Section preferences are project-specific, so the override is the right place for them rather than the default.

3.6.1: Tracker conflicts follow your chosen policy

Section titled “3.6.1: Tracker conflicts follow your chosen policy”

When Flow and your tracker disagree about status, the policy you chose now decides the outcome. Sync no longer stalls because the setting was documented but ignored.

Tracker updates now behave consistently across GitHub, GitLab, Jira, and Linear, including retries and partial failures, so sync no longer depends on which provider or agent happens to run it.

Detail
  • No migration step. Existing tracker configuration and lifecycle settings continue to work. The change is inside the execution boundary: skills decide what a body or comment means, then make one flowctl tracker sync call.
  • Provider requests, pagination, create-first recovery, status policy, dependency links, comment deduplication, and receipts now run through one tested implementation. A retry can prove what already landed instead of asking the agent to reconstruct provider state from prose.
  • Backlog autonomy now uses the same executable layer. Dependency ordering reads normalized directed edges, and parked questions carry stable identities so retries do not post duplicates.
  • The failure boundary is explicit. Authentication, rate limits, stale ids, capability gaps, conflicts, and transport failures return structured classes. The host still decides whether to ask, defer, continue through MCP, or correct local input; provider mechanics no longer consume its judgment budget.
  • The old transport recipes and tracker-runner agent are gone. Lifecycle callers retain their silent inactive gate, so repositories without tracker sync do not pay a new process or output cost.

The tracker-sync batch was initially published as 3.5.2, then republished unchanged as 3.6.0 the same day because three substantial specs belong in a minor release, not a patch.

The 3.5.2 artifact remains available as an accurate historical record. Upgrade to 3.6.0; there is no runtime difference between the two versions beyond corrected version metadata.

3.5.1: The review error stops suggesting the wrong fix

Section titled “3.5.1: The review error stops suggesting the wrong fix”

If your agent has been reporting that a review failed for “sandbox” reasons and needs retrying with wider permissions, it was reading our error message, not your repo. Reviewers are read-only on purpose, so a reviewer that hits the sandbox is a scoping bug - the message now says so instead of telling you to hand the reviewer write access.

3.5.0: Two agents, one spec number, and the collision stops being your problem

Section titled “3.5.0: Two agents, one spec number, and the collision stops being your problem”

fn-7 twice is not bad luck, it is arithmetic: spec ids were allocated by counting the files in your working tree, so two branches cut from the same base both saw the same number and both took it. Allocation now sees every worktree and every ref, and teams with a tracker can hand the job to the tracker instead - one setting, and new specs are keyed WOR-17 or gh-123 from the start.

Detail
  • The collision was structural. Allocation counted only the current working tree, so parallel spec creation was guaranteed to collide, not merely likely - and in an agent-heavy workflow parallel is the normal case. It now takes the maximum across the working tree, every registered worktree, and every ref, and it is monotonic: a retired number is never handed out again, because reusing it would resurrect an ambiguous reference in your commit history and release notes.
  • The number stays. Dropping fn-N would have been a vocabulary migration across changelogs, commits, tags, tracker comments and notes, to fix a symptom. The full fn-N-slug was always the real identity; what actually broke was a validate error and ambiguity in prose.
  • Honest bound, stated in the spec itself: sequential allocation without coordination is unsolvable in general. This shrinks the window; it does not close it. Two clones that have never fetched each other can still collide - which is what the tracker route is for.
  • Tracker-keyed ids are now a setting, not a flag you had to remember. flowctl config set tracker.specIds tracker, or just answer the question setup asks once when a tracker is configured. A tracker is a real distributed allocator, and collaborative repos are exactly the ones that have one.
  • GitHub and GitLab are no longer second-class. Linear and Jira ship a KEY-N that mints directly; #123 and group/project#456 do not, so they mint through a synthetic key derived from your configured tracker type - gh-123-slug, gl-456-slug. Unambiguous because a repo has exactly one configured tracker, and guarded so a minted id can never collide with a historical one.
  • New: create-first. Every previous tracker operation needed a local spec first, so “make the issue, then key the spec from it” could not be expressed at all. It can now, with the failure case designed in rather than discovered later: if the issue is created and something downstream fails, a retry links to that issue instead of creating a second one.
  • Network cost is stated accurately, not flattered. An earlier draft claimed tracker-first adds no network cost. That is false when your lifecycle events are off, which is the default - so the docs and the setup question now say plainly that choosing tracker-keyed ids makes spec creation contact your tracker immediately.
  • Security fix. flowctl no longer writes through a symlink anywhere between .flow and the file being written. An untrusted checkout could previously redirect a write outside your workspace, or onto another managed file. A legitimately symlinked .flow directory still works.
  • Incidental: flowctl task set-title updates the JSON title and the markdown heading together, so the two cannot drift apart.

3.4.5: Recurring lessons graduate into enforced gates

Section titled “3.4.5: Recurring lessons graduate into enforced gates”

When your agent keeps re-learning the same lesson every run, the lesson is in the wrong place. The memory audit now has a sixth outcome, Harden: a correct, recurring, mechanizable lesson gets proposed as a lint rule, a CI step, or a rule in your CLAUDE.md/AGENTS.md - and only after the gate is verified to actually fire does the memory entry retire into a pointer at it.

3.4.4: Claude Opus 5 joins the routing menu

Section titled “3.4.4: Claude Opus 5 joins the routing menu”

Claude Opus 5 launched yesterday at near-frontier intelligence for half Fable 5’s price - the recommended routing now leads with it, and absorbing a new model generation took a table edit and two registry rungs, not a pipeline change.

3.4.3: Faster skills, honest review retries

Section titled “3.4.3: Faster skills, honest review retries”

The skills you run most now load only the instructions their current route needs, while a broken review transport no longer wastes the three-round correctness budget.

A memory title starting with a quote or - could write an entry the CLI itself could not read back - the file sat on disk while memory list and search silently pretended it never existed. Both halves are fixed: those titles now write valid frontmatter, and any unreadable entry is reported instead of hidden.

Thanks to @TechupBusiness for the exceptional report (#235).

Independent tasks can now move together when your host can isolate them safely, without turning every plan into a hand-built scheduler.

Detail
  • Plans show dependency-ordered execution waves, so you can see which tasks are candidates to run together.
  • Work evaluates the complete ready frontier and may dispatch a safe concurrent subset. The host chooses worker count, isolation, and integration from its live capabilities; when those conditions are uncertain, it explains why and continues sequentially.
  • Concurrent workers return task-specific handovers. The conductor joins and integrates the whole wave before review, completion, tracker updates, and downstream plan sync.
  • Atomic task claims prevent duplicate ownership only. They do not make a shared Git index or filesystem race-safe.
  • Grok documentation is corrected from live 0.2.111 evidence: type /flow-next: to open the plugin command namespace, including plan and work. /flow-next- searches the separate hyphen-named skill surface, and argument hints appear after autocomplete selection.

If you run flow-next in xAI’s Grok Build, setup now recognizes it instead of mistaking it for Codex - so you get proper slash commands and Claude-format instructions, not Codex $flow-next- syntax written into the wrong file.

Rewriting a draft no longer asks you to mark it ready just because some other spec is ready.

3.3.2: Capture understands old compactions

Section titled “3.3.2: Capture understands old compactions”

A conversation that was compacted earlier no longer blocks capture when the feature you are capturing is still fully visible.

Your flow-next slash commands showed up in the Claude Code menu with the plugin name stuttered three times (/flow-next:flow-next:flow-next:qa). They now read the way they always should have: /flow-next:qa.

Detail
  • The fix. The command shims lived in a subfolder named after the plugin and carried a legacy namespaced name: field, so a recent Claude Code namespacing change stacked the prefix three deep. The shims are now a flat set of files with bare command names, and Claude Code prepends the plugin prefix exactly once. Nothing about how you invoke a command changes; the menu just reads correctly.
  • Cursor and Codex stay in lockstep. Every command keeps the name + description Cursor’s marketplace review requires, the Cursor manifest and installers point at the new flat layout, and the Codex prompt install is unchanged. A retired epic-review alias (superseded by spec-completion-review back in 2.0) is cleaned off older Codex installs on upgrade, non-destructively - it is moved aside, never deleted, and a same-named file of your own is left untouched.
  • Messaging tidy-up. The plugin’s own store/marketplace descriptions now match the wording on this site, and the bundled component counts are current.

If your team runs flow-next in Cursor, the rough edges are gone: install it org-wide from a repo instead of a per-person script-and-restart, get cross-family review without reaching for an external CLI, and stop wondering whether the stale “autocomplete doesn’t list commands” warnings were still true.

Detail
  • Install once, for the whole team. A Cursor Teams/Enterprise admin imports the GitHub repo as a team marketplace (Default Off / On / Required, auto-refresh on push) - every engineer gets flow-next with no local install and no restart dance. The install-cursor.sh / .ps1 scripts stay as the individual path. Admin runbook on the platforms page.
  • review.backend host - cross-family review from inside Cursor, no external CLI. Review runs as a fresh-context subagent pinned to a model from a different family than the one that wrote the code (Cursor honors in-prompt slug pins). It fails closed rather than quietly reviewing your code with the same model that wrote it - if no cross-family pin is set it asks (interactive) or reports NEEDS_HUMAN (autonomous). Works on Claude Code and Codex too; the other backends are unchanged.
  • Setup understands it’s running in Cursor. It detects a marketplace or local install correctly, leads the review menu with Host, and writes a model-routing block into AGENTS.md with live Cursor model slugs - a cheap one for read-only scouts, a cross-family one for review. Model tiering by alias (haiku/sonnet/opus) falls back to your session model on Cursor; the explicit slug pins are how you steer.
  • Read-only agents are actually read-only on Cursor. Cursor ignores the tool blacklist other hosts use, so review and scout agents now carry Cursor’s native read-only flag - closing a gap where a “read-only” reviewer could still edit files.
  • Approval prompts you can read. When capture or interview asks you to approve a draft, the full draft now prints as normal text first and the question stays short - no more multi-paragraph spec collapsed into an unreadable one-line prompt.
  • Docs match reality. Slash autocomplete lists the commands (hyphenated form), plain-English requests trigger the right skill, and native structured questions work including multi-question batches. Ralph autonomous mode is still Claude-Code/Codex only - Cursor has the hooks, flow-next just doesn’t wire them there.

Automated validation now tells you exactly which native spec IDs collide, without forcing a second text-mode run to find the missing errors.

Detail

flowctl validate --all --json now includes every counted native spec-ID collision in root_errors, in deterministic order, while keeping total_errors exactly aligned with the returned diagnostics. Existing text output and exit behavior are unchanged.

flowctl now gets out of the agent’s way without trading away trust: common entry points start faster, large repositories scale linearly, and concurrent agents cannot silently lose tasks or reuse stale model choices.

Detail
  • Upgrade first: Flow-Next now requires Python 3.11 or newer. Every launcher rejects an older-but-working interpreter before loading the CLI and tells you how to select or install a supported one.
  • Root help and flowctl usage are about 68-73% faster on the measured macOS baseline. flowctl specs is 27.5% faster and Prime classification 18.8% faster. The acceleration stays source-authoritative: no opaque or stale executable cache becomes a new source of truth.
  • Portable cross-process locks now cover task creation, setup/runtime state, and model-cache updates on POSIX and Windows. Paired task JSON/Markdown publication is transactional, and explicit model pins never silently downgrade.
  • Large-repository work now shares one task inventory and one reverse-dependency graph. Status/list read each eligible task once and spawn no subprocesses; Prime, cognitive-aid export, memory, pilot logging, and frontmatter parsing shed repeated scans and reads.
  • RepoPrompt Community Edition is now the primary integration. Flow-Next prefers rpce-cli, retains discontinued Classic only as the final compatibility fallback, and understands CE’s current window, repository-root, and chat response shapes. Repeated review setup reuses the existing repository window instead of cloning the workspace. Thanks @aidancurry for the precise #228 report.
  • Active docs, skills, smoke labels, and the Codex mirror now match the live post-3.1 command and payload surface; confirmed dead helpers and test-only production surfaces are gone.

3.1.2: Say it in prose, Codex finds the skill

Section titled “3.1.2: Say it in prose, Codex finds the skill”

On Codex, “plan this feature”, “work on fn-12”, or “pilot this to completion” now resolves the matching flow-next skill by itself. Before, the model-facing skill catalog carried six internal helper skills and hid every user-facing verb, so prose-invoked runs depended on the model rediscovering skills from disk - or silently improvising without the skill contract. Setup also stopped assuming you know flow-next: every ceremony question now explains what it decides and links these docs.

Detail
  • Catalog policy un-inverted: all 22 user-facing skills are now implicit-invocable on Codex (its naming rule then requires their use when you name or clearly describe one); the 6 skill-dispatched internals (drive, sync, export-context, rp-explorer, worktree-kit, deps) are explicitly hidden - still invocable by name, out of the catalog budget.
  • Catalog descriptions dieted to fit Codex’s shared skills context budget (min of 8,000 chars and 2% of the context window, shared with every other skill on your machine): 2.9k chars total for all 22, vs ~7.6k undieted. New sync guards hard-fail on a missing catalog policy or an oversized surfaced description.
  • /flow-next:setup rewritten for newcomers: each question states its stakes in plain language (what a setup mode decides, what a review backend is, what plan-sync/memory/Ralph actually do) and links the relevant page here for the longer answer. Same options, same defaults, same config keys.
  • Codex installer fix: the memory-track templates (templates/memory/*.tpl) now actually land in ~/.codex/templates/ instead of flowctl silently falling back to embedded defaults.
  • Update: git pull && ./scripts/install-codex.sh.

If you route bulk implementation to Grok 4.5, the recipe you copy now unlocks the whole CLI: it edits files headlessly like the codex and cursor bridges do - the old guidance sold it as print-only, and hid a flag trap (-p swallows the next flag as its prompt) that made the write mode look broken. Docs-only patch; copy the corrected invocation and you get a third editing delegate on its own quota.

Detail
  • Corrected form, flags first: grok --permission-mode acceptEdits -m grok-4.5-high -p "<task>" (write mode; --always-approve is the blanket variant). The old grok -p --always-approve "..." shape misparses - live-verified.
  • Extras now documented: --check (self-verify loop), --best-of-n N (parallel attempts, best picked), --json-schema (structured output).
  • Same discipline as every bridge: run inside a trusted git dir; the host reviews and commits - Grok stays routed to bulk implementation, never final taste-critical work.
  • Updated in the usage guide (flowctl usage serves it live in plugin-mode repos) and the model-routing scaffold’s grok route.

3.1.0: Set up once, never again (Claude Code)

Section titled “3.1.0: Set up once, never again (Claude Code)”

“Re-run setup in every project after every update” stops being a rule you have to remember. On Claude Code, setup now asks one question per repo - and if the repo is Claude-Code-only, it copies nothing at all: flowctl is simply on your agent’s PATH, the guide is one flowctl usage away, and plugin updates land silently. The nag that fired in 15 skills after every release goes quiet, permanently.

Detail
  • Plugin mode (new, Claude Code): the only thing written to your repo is a slim versioned block in CLAUDE.md. No .flow/bin/, no .flow/usage.md, no snapshots to drift. Bare flowctl list works in any agent shell; flowctl usage prints the always-current CLI cheatsheet + orchestration recipes straight from the installed plugin.
  • Copy mode (unchanged): repos with Codex/Cursor/Droid teammates, CI, or plain-terminal flowctl use keep the committed snapshots - that is what makes a teammate’s clone work with no plugin installed. The update-then-re-run rule still applies there, exactly as before.
  • Switching is consented, never silent: moving a copy-mode repo to plugin mode lists the leftover snapshots and asks before removing them; the mode stamp itself is written by a new flowctl setup-mode set command that refuses to declare plugin mode unless the CLAUDE.md rail is in place and no snapshots remain - so a half-finished switch cannot leave you in a broken in-between.
  • Two layers of steering, written down: the orchestration page now spells out session steering (prompts and per-task pins - “implement via grok-4.5 and review with sol” just works and persists nothing) vs machinery steering (config that pilot/Ralph resolve at 3am when nobody is prompting), with the full precedence chain.

3.0.0: Smaller, quieter, and only what you use

Section titled “3.0.0: Smaller, quieter, and only what you use”

You stop paying for machinery you never asked for. Every install used to run Ralph’s guard on every shell command and file edit in every session, whether or not you ever touched autonomous mode - that is gone. The CLI dropped thousands of lines of commands nobody called, review verdicts get harder to fool, and model choices move into config you can actually see and refresh. Three breaking changes, each with a short documented path forward.

Detail

If you upgrade, do these first:

  • Using Ralph? Run /flow-next:ralph-init once per project. Hooks are no longer installed by the plugin - nothing Ralph-related runs anywhere until you opt a project in, and without re-init your guard will not fire. Everyone else: do nothing, and enjoy sessions with zero flow-next hook overhead.
  • Still on a pre-1.0 .flow/epics/ layout? The automated migrator is gone; porting by hand is three short steps, listed in .flow/usage.md under “Pre-1.0 layout porting”.
  • Scripts or tools reading epic fields from flowctl JSON, or passing --epic flags? Switch them to the spec forms before upgrading - the legacy aliases and duplicate JSON keys no longer exist. (The depends_on_epics field in spec files is real schema, not an alias, and is unchanged.)

Why this release exists: an audit asked one question of every deterministic line in the CLI - would this still be needed if the model were smarter? What failed the question was removed or handed back to the agent; what passed got engineered properly. Concretely:

  • The CLI is honest about what exists. Dead commands are gone rather than half-documented, and the reference docs now match the CLI exactly - if a command is documented, it works.
  • Reviews are harder to fool and easier to extend. All nine review commands share one engine, review prompts are visible markdown files you can read and diff instead of strings buried in Python, and reviewers report their tallies in a structured block that resists prompt injection from the code under review. Receipts your automation reads are byte-for-byte unchanged.
  • The agent decides; the CLI stores. memory add no longer silently rewrites an existing entry when a new one looks similar - it always creates unless you explicitly say which entry to update, and shows you the near-matches so you (or your agent) make the call. Review-quality judgment calls in interactive sessions go to the agent in front of you; autonomous runs keep the deterministic behavior their receipts depend on.
  • Model choices live in config, not code. Which model judges triage, reviews, or implements delegated work is now a small table in .flow/config.json that /flow-next:setup offers to refresh by probing what is actually installed - so pins stop rotting when providers ship new tiers. flowctl models resolve <role> shows you what will actually run.

Waiting 15+ minutes for a test run is where verification discipline goes to die - so the suite now runs in about 90 seconds, and the faster gate immediately caught real problems the old setup had been hiding for months.

Detail
  • A new parallel runner (scripts/run_tests_parallel.py, pure stdlib) runs test files concurrently and proves it returns exactly the same results as the serial run. A hung test file fails loudly with its name instead of stalling the whole suite.
  • The honest surprise: CI turned out to be running only 29 of 87 test files - the old hand-enumerated steps had silently drifted. The parallel runner discovers everything, and the first full runs on Windows surfaced six real latent issues, each now tracked with its cause instead of swept aside.
  • The two slowest test files got 2-3x faster without losing a single test, and per-task verification now runs just the focused suites for the files you touched - the full suite runs once at the end, where it belongs.
  • What it means for you: if you dogfood flow-next’s pipeline pattern in your own repos, this is the template - a fast full-suite entrypoint plus focused per-task suites is what makes run-the-tests-every-task actually sustainable.

2.21.0: Fewer round-trips in every pipeline run

Section titled “2.21.0: Fewer round-trips in every pipeline run”

The skills you run most (plan, land, pilot, make-pr) now read configuration once and create tasks in one call instead of three - less waiting, fewer places for a half-written task to exist.

Detail
  • flowctl config get can return a single value, a whole section, or the entire config in one call - so a skill that used to issue seven config reads issues one.
  • task create accepts the description, acceptance criteria, and requirement links up front. A freshly planned task is complete the moment it exists; there is no window where a task file is created but empty.
  • A hardening rider: the rule that review commands must run in the foreground (backgrounded review calls die silently and stall workers - observed live, twice) is now embedded at every point where a review is invoked, and pinned by tests so it cannot quietly erode.

2.20.0: Skip-what-you-proved, now actually working

Section titled “2.20.0: Skip-what-you-proved, now actually working”

2.18.0 promised the work loop would stop re-running test suites it had already proven green. An audit of a real run showed the promise was structurally broken - zero receipts were ever honored. This release makes it real, and the measured ~20-25% wall-clock saving on multi-task specs arrives.

Detail
  • The bug: the work loop’s own bookkeeping commit changed HEAD right after every green run, so the receipt recorded for the previous commit never matched. Receipts now survive bookkeeping-only commits through a strictly bounded ancestor check that fails closed on anything suspicious - a receipt is honored only when nothing that could affect the tests has changed.
  • Two worker rules close the other measured leak (suites re-run just to look at a result the exit code already carried): greenness is read from the captured exit code, and gate suites run as one blocking foreground call.
  • Honest bounds, stated plainly: about 35% of runs still deliberately force a full suite as the safety floor. The next lever below that is parallelization - which shipped in 2.22.0.

On a large repo, flowctl list took 30 seconds - and your autonomous loops paid that on every single tick. Now it is under half a second, with every guarantee intact.

Detail
  • The cause: every task load spawned two git subprocesses - 809 process spawns per list on a 100-spec repo. Repo-root lookups are now cached safely (a directory change or a transient git failure never poisons the cache), and several smaller repeat offenders were swept in the same pass.
  • Measured on flow-next’s own repo: list 30.8s to 0.48s, status 32s to 0.41s. Every pilot, land, and Ralph tick gets that time back, every time.
  • This was the first installment of a standing audit with one question - would this still be needed if the model were smarter? Mechanisms that survive the question (locks, receipts, atomic writes) get engineered like the hot paths they are; the rest gets deleted. The releases that follow are the same program continuing.

Delegating a task to another model used to involve composing a multi-kilobyte briefing document per task - 5 to 17 minutes of prep. A controlled experiment showed the briefing added nothing the task file didn’t already carry. So now the task file IS the brief.

Detail
  • The experiment replayed three real tasks with the historical hand-composed briefs as the control, blind judges, and pre-registered pass bars. The tiny fixed prompt (“read the task and spec files, implement, follow the rails”) tied the composed briefs on every measure; the one gap that appeared closed with a single added sentence in the task template.
  • What it means for you: plan quality moved to where it belongs - the task file. Plans now require named files, named test cases, and named acceptance criteria, because whatever executes the task receives that file as its entire instruction. A task too thin to delegate safely is implemented in-session instead - automatically.
  • Every safety rail (pre-flight checks, rollback, failure classification, the circuit breaker) is unchanged and machine-verified.

2.18.0: Stop re-proving what you already know

Section titled “2.18.0: Stop re-proving what you already know”

A trace of one real pipeline run found the test suite executing ten times in a single work stage - half the wall-clock - while actual implementation was 4%. A docs-only task spent 84% of its time on gates its diff physically could not break. Both stop here.

Detail
  • Green test runs now leave a receipt keyed to the exact commit and command; an identical later check honors the receipt instead of re-running. Anything at all suspicious - a dirty file, a changed command, a stale timestamp - and the suite runs in full. Fail-closed, always.
  • A mechanical classifier recognizes when a change touches only docs and prose, and skips the gates that change cannot affect. No semantic guessing - purely file-type rules, with skill prose (which IS shipped behavior) always running full gates.
  • Every skip is loud: it lands in the task’s evidence and the run summary, so you can always see what was skipped and why. CI never consults receipts - remote gates always run everything.

2.17.0: Tracker updates stop crowding your session

Section titled “2.17.0: Tracker updates stop crowding your session”

At 2.17.0, linked tracker updates moved to a background runner to isolate about 124k measured tokens from the working session. That runner was later removed when deterministic provider operations moved into flowctl tracker; current lifecycle callers invoke one compact facade command instead.

Detail
  • Historical behavior: the four routine tracker touchpoints dispatched to a small background subagent. Current releases keep semantic judgment in the caller and execute deterministic provider operations through the facade.
  • Anything that needs judgment, including discovery choices, 3-way body conflicts, comment synthesis, and structured-error recovery, stays with the host agent.
  • Zero added wall-clock in the live proof, and duplicate-protection hardening came straight from that dogfood run (a comment that lands but whose confirmation is lost is re-checked, not re-posted).

/flow-next:interview used to trickle questions a few at a time across many turns. It now asks everything that is ready to be asked in each round - and while you type, an optional background scout looks up the codebase facts the next round depends on.

Detail
  • Rounds are computed from what actually blocks what: a question never appears alongside its own prerequisite, and dropped lines of questioning are announced, not silently vanished.
  • Question quality has a bar - “a slot is earned”: failure modes, concurrency, scale, and testing always qualify; cosmetic polish gets folded into option text instead of costing you a question.
  • The background fact-scout was validated the hard way: the fastest model tier missed a load-bearing architecture fact the mid tier found on the identical brief, so the mid tier is the floor.

2.15.0: Leaner always-loaded docs that teach the one thing that matters

Section titled “2.15.0: Leaner always-loaded docs that teach the one thing that matters”

Every token in your CLAUDE.md is context you pay for in every session. An experiment isolated the single thing the flow-next setup block measurably buys - agents filing completion evidence in the right shape - so the block now teaches exactly that, at half the size.

Detail
  • The old block named the evidence-JSON flag but never showed the shape; capable agents reliably completed tasks without valid evidence. The new block shows the schema inline. Measured across three model families: the failure mode disappears, nothing else regresses.
  • .flow/usage.md dropped from ~5.4k to ~1.9k tokens by cutting what --help already teaches. It is read on demand, not always loaded.
  • Setup re-runs now silently refresh blocks you never customized (a stored hash tells the difference) - so improvements like this reach existing projects without a prompt storm, while your edits are never overwritten unasked.

2.14.0: Multi-model defaults you can trust

Section titled “2.14.0: Multi-model defaults you can trust”

Which model should implement delegated work? We ran the experiment instead of guessing: gpt-5.6-terra at medium effort matched the stronger tier’s correctness at two-thirds the wall-clock on well-specified tasks - so that is now the default, and setup scaffolds the whole recommended multi-model pipeline when it detects the CLIs to run it.

Detail
  • The recommended shape, now stated concretely everywhere: your session model authors specs (that is where quality is made), a value-tier model implements against them, and a reviewer from a different model family reviews - uncorrelated blind spots by construction.
  • Setup only offers what can actually run: routing questions appear when the bridge CLIs are installed, cross-family review is recommended only when it IS cross-family for your setup, and switching an existing review backend is always explicit, never silent.
  • For unattended loops, the docs now carry the self-healing wrapper pattern and the known sharp edges of each bridge CLI (silent-refusal modes, workspace checks) so overnight runs fail loudly instead of mysteriously.

2.13.1: Portfolio triage stays quiet on ordinary repos

Section titled “2.13.1: Portfolio triage stays quiet on ordinary repos”

Dogfooding 2.13.0’s classifier on real machines found it flagging noise - your own repo referencing itself, ~/Downloads, cache directories - as “cross-repo signals”. Patched same-day: --classify-only sweeps now stay quiet unless there is a real constellation to report.

2.13.0: Prime tells you the truth about your repo

Section titled “2.13.0: Prime tells you the truth about your repo”

The old readiness assessment could award “Level 5” to an empty template with a broken build - it checked whether files existed, not whether anything worked. The rebuilt /flow-next:prime runs your build, runs your tests, boots your app, and leads with a verdict and a ranked list of what to fix first.

Detail
  • Prime now starts by classifying what it is looking at - lifecycle, monorepo topology, size, stacks (including legacy ones like Delphi and COBOL), delivery shape - and judges against the right yardstick for that kind of project. A greenfield toy and a 20-year brownfield monolith stop getting the same checklist.
  • Hard gates cap the score: if the build does not run, tests are not discoverable, or the commands in your agent docs do not actually execute, no amount of nice file structure rescues the rating. Executed evidence beats existence.
  • --classify-only sweeps a portfolio of 100+ repos in seconds when you want triage rather than a deep assessment.
  • The claims survived 29 adversarial review waves before shipping - capped scans admit they are capped, CI steps only count if they actually gate, and secrets never leave the machine.

See the rewritten Prime skill page.

2.12.4: A stale setup finally gets your attention

Section titled “2.12.4: A stale setup finally gets your attention”

If your project’s local flow-next files lag the plugin you are running, skills now ask you once - Refresh now / Remind me next version / Skip - instead of hiding a one-line note you were never going to see. Your answer is remembered per version, autonomous runs are never interrupted by the question, and pilot/land print a grep-able SETUP_STALE: line for loop drivers. Also fixed: tracker-sync no longer reports false conflicts when Linear rewrites your markdown on save - the sync now remembers what the tracker actually stored, not what was sent.

2.12.3: Know what each model is actually good for

Section titled “2.12.3: Know what each model is actually good for”

The optional model-routing table gains a speed column and a Grok 4.5 row - fast and cheap with strong coding, weaker on UI and more prone to invention, so the guidance routes it to bulk implementation and never to final taste-critical work. The fourth headless bridge (grok -p) joins the recipes, and the cost column now says plainly what it measures: how lightly a model rides your subscription quota, not list price. Nothing activates unless you ask for routing - defaults are unchanged.

2.12.2: Delegated work goes to a strong model, deliberately

Section titled “2.12.2: Delegated work goes to a strong model, deliberately”

Delegated implementation is real work that ships, so its default model moved up a tier rather than down - the counterpart to routing cheap models at bulk reads. One sharp edge documented honestly: the delegation path has no fallback ladder, so this default needs a current codex CLI; the one-line downgrade for older CLIs is in the notes.

2.12.1: Fresh model numbers the day after a launch

Section titled “2.12.1: Fresh model numbers the day after a launch”

GPT-5.6 went GA and the scaffolded routing table was still handing new users last-generation numbers - so the table got current rows and an explicit “as of” staleness stamp. These are starting opinions you edit after scaffolding, not runtime defaults; nothing about live model resolution changed.

2.12.0: PRs that tell reviewers where to look

Section titled “2.12.0: PRs that tell reviewers where to look”

As AI-assisted PRs grow, reviewers lose the thread - which came from a field report in exactly those words. PR bodies now open with what the pipeline already verified mechanically versus what genuinely needs human judgment, and a risk-ranked review plan: Must review / Spot-check / Safe to skim, capped at roughly 30% of the diff, each item saying why and what to check.

Detail

Validated blind before implementation: reviewers scored the old PR bodies 7/10 on effort-targeting and 5/10 on trust calibration; the shipped format scores 9/9. Underneath, the PR export gained deterministic traceability - which symbols changed per file, which files are generated copies a human should never re-review, and which deleted exports still have references (conservative candidates worth a look). No claim in a PR body is invented: every line traces to a computed field.

2.11.0: A model launch never breaks your reviews again

Section titled “2.11.0: A model launch never breaks your reviews again”

GPT-5.6’s launch reproduced a familiar failure live: review backends erroring because their pinned model did not exist yet on your CLI. Model resolution now dispatches the best model directly, and when a model is genuinely unavailable it steps down a quality-ranked ladder to the CLI’s own default - your review runs, always, and the receipt records what actually ran.

Detail

The outcome is cached per CLI version so the fallback probe costs one retry per upgrade at worst. Your explicit pins bypass everything, unchanged. Unknown model names warn and proceed instead of failing - no more waiting for a plugin release to try a model that launched this morning. Three review rounds before merge caught three real bugs, including the fallback being silently dead on the paths that mattered; each is now pinned by a regression test shaped like the actual failure.

2.10.3: Cursor reviews on GPT-5.6 Sol, verified first

Section titled “2.10.3: Cursor reviews on GPT-5.6 Sol, verified first”

The cursor review backend moved to GPT-5.6 Sol with 1M context - after live-verifying the model actually resolves on the current CLI. Codex and copilot deliberately stayed put: live probes showed both would reject the new model, and swapping an unverified default is precisely the failure 2.11.0 then eliminated for good. This was the verified interim.

2.10.2: Capture read-back: plain-language ratification, no self-blessing

Section titled “2.10.2: Capture read-back: plain-language ratification, no self-blessing”

The capture approval question - where a human ratifies the synthesized spec - now opens with what approving means, lists every requirement in one plain line inside the question, translates its own machinery (“[inferred] = something I added that you didn’t say outright”), and never recommends approve while unverified inferred items exist - the agent doesn’t pre-bless its own guesses. The companion to 2.10.1, aimed at the highest-stakes dialogue moment in the system. Blind-eval before shipping found the real bug was beyond language: “Recommended: approve - the inferred items are reasonable” was steering users to rubber-stamp exactly what they exist to verify (ratification-safety 4/10). The shipped contract scores 9/9/10 on legibility / ratification-safety / precision.

2.10.1: Interview questions in plain language

Section titled “2.10.1: Interview questions in plain language”

Every interview question now opens with one sentence of stakes (what this decides, in your words), glosses any term of art in plain words at first use, and states each option’s consequence (“Choose this if…”) - field feedback from a team where jargon-dense questions were disempowering the product people the interview exists to hear. Sizing is expressed as priorities (always-keep / trim-first lists with a target shape), not a length cap - deliberately aligned with OpenAI’s GPT-5.6 guidance that generic brevity instructions make capable models drop required content. Eval-validated blind before shipping: baseline questions scored 4/10 legibility for a second-language PM; the shipped contract scores 7.5+ with precision held and ~30% fewer tokens per question. The same eval tested extending this to spec prose and rejected it (business sections already read at 9/10 - the contract only made them longer), so specs are unchanged by design.

2.10.0: Review-loop runaway root fix: honest verdicts, convergence ratchet, deterministic cap

Section titled “2.10.0: Review-loop runaway root fix: honest verdicts, convergence ratchet, deterministic cap”

A field-reported 17-round plan-review runaway on the Cursor backend is fixed at its roots: codex/copilot verdicts are now parsed only from the reviewer’s final message (tool-output <verdict> literals can no longer produce a false SHIP/NEEDS_WORK), re-reviews follow a shrink-only convergence ratchet (prior findings injected; only NEW ≥Major blocks; all-fixed ⇒ MUST SHIP), and MAX_REVIEW_ITERATIONS (default now 4) is enforced deterministically by flowctl - the counter survives fresh invocations and refuses at the cap with an ESCALATE marker instead of looping. Cursor reviews also carry a persona override superseding cursor-agent’s built-in review rubric and auto-attached AGENTS.md guidance. Review receipts are spec/task-scoped (concurrent reviews no longer collide); flowctl spec reset-review-rounds resets on re-plan. Round counting includes failed dispatch attempts by design (anti-runaway bias). Validated live on a pinned five-review baseline dataset: a cursor fix→re-review cycle converges to SHIP in 2 rounds.

2.9.1: Fix: completion-review tracker audit was dead (event-key mismatch)

Section titled “2.9.1: Fix: completion-review tracker audit was dead (event-key mismatch)”

The work skill’s completion-review tracker touchpoint dispatched and audited the event as work.completionReview, but the config leaf is top-level tracker.perEvent.completionReview - so flowctl sync check resolved no leaf, treated the event as configured-off, and could never report the touchpoint missing (a dead dispatch could never retro-fire). The event key is now completionReview everywhere; existing configs keep working unchanged.

Detail

Found during the fn-90 review-loop investigation: five independent cross-backend reviews (cursor + codex) of a planned spec each flagged the key mismatch, and a code check confirmed it live. Only the event tag moved - the config leaf written by the discovery ceremony stays top-level, so no user action is needed beyond updating the plugin. Regression coverage landed in test_sync_check.py: the top-level key round-trips the audit (enabled + no receipt → MISSING; a tagged receipt clears), the old work.-prefixed shape demonstrably resolves no leaf, and a prose guard fails the suite if any canonical skill or doc reintroduces the mismatched tag.

2.9.0: Interview: scope question + skips are not answers

Section titled “2.9.0: Interview: scope question + skips are not answers”

A bare /flow-next:interview no longer silently runs the technical question bank - it asks which pass to run (business / technical / both) - and skipped questions no longer become silent decisions: they park under ## Open Questions, with a consent checkpoint before write-back.

Detail

Both fixes come from a downstream field report: a product manager ran a bare interview, skipped the technical questions, and the agent filled architecture/stack/API sections with project-rails-derived defaults written as settled decisions.

  • Scope question (interview): when no --scope / --biz / --tech flag is passed, the interview asks one upfront question with a recommendation derived from the spec’s current state (both layers empty → both; business populated → technical; 1.0.2-shape tech-only spec → technical). An explicit flag skips the question; technical remains the fallback when the question can’t be asked. Plumbing: flowctl scope resolve --json now emits a defaulted boolean. The Codex mirror gets the same question via the plain-text prompt transform.
  • Skip contract: only an explicit answer or an explicit “you decide” delegation resolves a question. Skips/declines/“I don’t know” park under ## Open Questions with the agent’s unconfirmed leaning; a skipped judgment question never demotes to codebase-/docs-derived backfill. When anything was skipped, a consent checkpoint fires before write-back: park-open (default) / fill-assumptions (inline *(assumed, unconfirmed)* markers for later ratification) / re-ask. The completion summary reports the disposition.

After updating, re-run /flow-next:setup in your projects so the local .flow/bin/flowctl copy picks up the new defaulted field.

2.8.1: Model-routing scaffold: named tiers, menu wiring, freedom grant

Section titled “2.8.1: Model-routing scaffold: named tiers, menu wiring, freedom grant”

The setup scaffold’s routing table now names concrete models (fable-5, opus-4.8, gpt-5.5, composer-2.5, sonnet-5, haiku-4.5 - “session model” is a role, whichever row conducts), and the wiring is a per-role menu instead of fixed pairings: implementation via native opus/sonnet subagents, delegate:codex, or the cursor-agent bridge; reviews cross-family or prompted same-family; reads native or via cursor scouts. Plus an explicit freedom grant: unless prompted otherwise the harness routes as it judges best, and an explicit user instruction always overrides the table. Probe-gating unchanged - routes to CLIs you don’t have are never active. The orchestration page block mirrors the shipped scaffold verbatim.

2.8.0: Orchestration in your repo: usage.md steering recipes + setup routing scaffold

Section titled “2.8.0: Orchestration in your repo: usage.md steering recipes + setup routing scaffold”

The orchestration story now ships into the repo you work in: every installed .flow/usage.md carries dogfooded headless-bridge recipes (codex exec, cursor-agent, and the reverse claude -p - every direction works), and /flow-next:setup optionally scaffolds an opinionated model-routing table into CLAUDE.md/AGENTS.md.

Detail

Closes the use-time discoverability gap from 2.7.2’s orchestration doc: agents read .flow/usage.md and instruction files, not the plugin’s doc tree.

  • usage.md ## Orchestration & model steering (unconditional, ships in every project): codex exec recipes with the real gotchas baked in (read-only default sandbox, -o output capture, the </dev/null stdin-hang guard), cursor-agent (-p, --force to apply, volatile model IDs via --list-models), the claude -p reverse bridge (prompt before the variadic --allowedTools), flow-next shortcuts (delegate:codex, review.backend, per-task review:), and prompted-orchestration examples. Every recipe was verified against the live CLIs before shipping.
  • Optional setup ceremony (setup): scaffolds the cost/intelligence/taste scores table + routing rules + flow-next wiring into your instruction file - probe-annotated for the CLIs actually installed (<!-- probe:codex/cursor --> sentinel lines, deterministic composition), shown in full before writing, marker-fenced for idempotent re-runs (probe drift counts as drift), platform-correct invocation syntax per target file. Delegation opt-in sets work.delegate but never pre-sets the consent gate.
  • /flow-next:uninstall removes the scaffold via a deterministic damaged-marker algorithm; 23 new tests + smoke prose contracts pin the template shape, four-state probe composition, and removal.
  • Codex installs: the mirror’s usage.md now renders commands as $flow-next-<cmd> (generator rewrite + regression guard).

Defaults stay pre-tuned and unchanged - steering remains a capability, not a prerequisite; headless/Ralph setups skip the new question silently.

Given the trend toward frontier-model orchestration - Fable 5 conducting while implementation, reviews, and bulk reads route to cheaper/faster models - a new Orchestration & Model Routing page maps every routing dial flow-next already ships. Two composable methodologies: deterministic parameters (review-backend grammar + precedence, delegate:codex offload, subagent tiers) and prompted orchestration - the host’s own intelligence routing per item by complexity, escalating conditionally, even prompting capabilities into existence that no parameter encodes. Plus a copy-paste CLAUDE.md model-routing table and the pilot+land loop-chaining recipe. The frame throughout: the defaults are pre-tuned to work well out of the box - steering is a capability, not a prerequisite.

The installed Codex ~/.codex/hooks.json carried a top-level description key that Codex’s hooks parser (stable since 0.142.x) rejects - a warning on every invocation and the Ralph guard hooks silently disabled; the generated mirror no longer emits the key. Reproduced and verified clean against Codex CLI 0.142.5. Codex users: re-run scripts/install-codex.sh to replace the broken file. Thanks to @TechupBusiness for the report and root-cause analysis (#198).

2.7.0: Fleet-wide capability & efficiency review

Section titled “2.7.0: Fleet-wide capability & efficiency review”

Two adversarial-review passes over the whole skill/agent fleet - one new feature (make-pr --update), broad correctness and autonomy-safety fixes, progressive-disclosure efficiency, and seven A/B-verified opus→sonnet model downgrades - with every judgment call fable-reviewed before it shipped.

Detail

The engine was adversarial review, not prose-squeezing: six-reviewer fable audits found the gaps, a fable judge verified each fix, and every model-tier change was proven head-to-head. Across ~46 improvements in 25 skills and agents:

  • New capability - make-pr --update refreshes a stale PR body after review/land fix rounds; /prime gained a real evidence + scoring contract (a failed scout no longer silently drops or fabricates a pillar, verification is mandatory before a “runnable” pass, inapplicable criteria no longer deflate the score) plus create-or-augment CLAUDE.md/AGENTS.md handling; completion-review scope-creep detection, plan real-anchor derivation, prospect’s genuinely-isolated critique, qa evidence enforcement, and /deps surfacing deadlocks it used to hide.
  • Autonomy safety - land stops auto-merging a PR whose QA verdict is NEEDS_WORK; pilot’s strike limit survives a tracker re-projection that had it re-dispatching a failing spec forever; quality-auditor fails loudly instead of reporting a false-clean audit over an empty diff.
  • Model tiers - plan-sync, flow-gap-analyst, and the five retrieval scouts moved opus→sonnet, each A/B-verified; opus is now used by a single agent. Cheaper per call, quality held.
  • Efficiency - progressive-disclosure splits (interview, make-pr, impl-review, capture) and common-path short-circuits (tracker-sync, audit) take ~5k-17k tokens off the paths that run most.
  • Correctness spine - a review-diff base...HEAD fix (13 sites) that had a fast-moving base branch showing its own commits as false reversions, and a signal refresh so /prime recognizes 2025-era stacks (uv, bun, compose.yaml, mise, monorepo layouts, GitLab, goreleaser, Biome).

No breaking changes. Verified by a 1425-test suite green across Linux/macOS/Windows.

2.6.3: Single-call worker anchor + plan-sync gate shelved

Section titled “2.6.3: Single-call worker anchor + plan-sync gate shelved”

The /flow-next:work worker now re-anchors in one flowctl anchor call instead of ~8 separate reads (proven zero information loss), plus a CROSS_SPEC caller bug-fix - while the other half of the work, a deterministic plan-sync skip-gate, was proven non-viable by cross-repo eval and deliberately shelved rather than shipped.

Detail

flowctl anchor <task-id> assembles the worker’s Phase-1 re-anchor from the verbatim stdout of the same production commands it already runs - byte-for-byte superset test plus a comprehension-equivalence eval (bundle 7/7 = status-quo on frozen real tasks) prove no information is lost. The plan-sync skip-gate (a deterministic probe to skip the post-task drift check) was built, eval’d against the real plan-sync agent across three external repos, and killed by its own evidence: a genuine false skip from semantic drift no path/token probe can see, plus a 6.7% skip-rate against a ≥50% bar. It is shelved with a decision record, not shipped - plan-sync still runs after every task, and the gate machinery was removed from the CLI. Also fixed: a Windows encoding bug in the anchor render, surfaced by wiring the anchor guardrail test into CI.

flowctl ready --spec now honors spec-level dependencies - a spec blocked by unfinished depends_on_epics no longer reports its tasks as ready. It returns empty lists plus blocked_by_specs (legacy alias epic_blocked_by), matching the gate next and ready --all already applied. Latent in the default workflow; hit by external consumers calling ready per-spec. Thanks to Mike Bannister (#95).

Setup and install-codex.sh could leave a Codex config.toml with a duplicate hooks key (invalid TOML - Codex silently stops loading hooks) or the deprecated codex_hooks spelling (a warning on every run); both paths now converge through one idempotent, dedup-safe normalizer that guarantees exactly one hooks = true under [features]. Regression-tested (10 cases including the both-keys scenario); everything outside [features] is byte-preserved and re-running is a no-op.

2.6.0: Skill efficiency: single-emission writes + prompt diet

Section titled “2.6.0: Skill efficiency: single-emission writes + prompt diet”

Two paired specs cut the token cost of every skill run with zero quality loss - fn-81 eliminates runtime re-emission (spec bodies, review prompts, and responses materialized once instead of two-or-three times, plus 13 redundant CLI round-trips removed across 12 skills), and fn-82 trims the always-loaded prompt weight the hot-path skills carry on every invocation (−10.7k tokens across 11 skills). Skill-markdown only - no flowctl behavior change, no new commands, read-backs stay mandatory and user-authoritative.

Detail

Follows a fleet survey of all 28 skills (2026-07-02) and lands behind a full behavioral regression pass (gate matrix, two eval-suite re-runs at full score, smoke 138/138, pytest 1393 passed).

  • Runtime plumbing. The drafted spec body is materialized exactly once via the Write tool (the Write render is the user-visible read-back), revised via Edit deltas, and consumed by spec set-plan --file <path> - no Phase-5 heredoc re-authoring (capture, interview). RP review prompts are built by deterministic file composition (rp prompt-get > file, quoted-heredoc criteria, flowctl show >> file) - every [PASTE …] content-retype placeholder is gone and untrusted reviewer/spec content never transits a shell var (the injection surface is closed). RP review responses enter context exactly once (redirect → single Read). Round-trips removed: single LEAF= config read per tracker gate (7 sites), plan drops a post-write show+cat, deps runs one per-spec loop instead of two, make-pr’s §4.6b live gh pr view fires only on the local-assertion miss, tracker-sync reconcile passes the on-disk spec to set-merge-base --flow-file. Guards hardened: the fix-loop cap (MAX_REVIEW_ITERATIONS, default 3) now bounds all review backends, and both RP fix loops replace git add -A with snapshot-scoped staging (pre-existing dirty paths are never swept in).
  • Prompt diet. Default-OFF machinery moved behind a forcing-sentinel gate into references/*.md - zero tokens until Read (Anthropic Agent Skills 3-level loading): work’s tracker touchpoints and pilot’s QA-stage freshness probe. Each gate emits an imperative the agent must act on (GATE ACTIVE, STOP. Read <ref> …), fails open on probe/parse error, and no-ops silently on the default path; the safety nets (work’s Phase-5 sync check + four-state summary, pilot’s QA routing) stay inline. Duplicated explanatory blocks collapse to one authoritative site (the review pair now resolves the backend once - killing a double review-backend round-trip); build-time fn-N provenance and flowctl.py line-refs are stripped from always-loaded prose; make-pr folds its per-phase Done-when checklists inline (body eval held 5/5, −4.5k tok/run) and capture single-sources its biz-routing table at the consumer (suite held 15/15).

2.5.4: Section-write hardening + rp-gate completion

Section titled “2.5.4: Section-write hardening + rp-gate completion”

flowctl task-section writes are now normalization-hardened (the H2-layering bug caught in fn-78’s own autonomous dogfood is fixed, with self-heal for already-damaged files), and the fn-78 RepoPrompt eligibility gate now covers all four review skills - impl-review and spec-completion-review stop steering toward rp on hosts where it can’t run.

Detail
  • Task-section normalization. Agents routinely pass section content that starts with its own ## Acceptance Criteria … H2; task create --acceptance-file embedded it as a rogue sibling section and every later set-acceptance layered a new block above the old one. All task-section write sites now normalize through one helper: a leading H2 is stripped only when it matches the section’s known-title-variant grammar (## Acceptance Tests is content - demoted, never stripped), remaining H2s demote to H3 outside code fences, writes are byte-idempotent, and an on-write self-heal folds contiguous rogue sections from already-damaged files (a byte-exact duplicate ## Acceptance still raises). Fence-awareness extended end-to-end via one shared tracker (patch_task_section, get_task_section, heading validation, set-spec scaffold check).
  • RP_ELIGIBLE gate completed. The 2.5.3 gate covered plan/plan-review; now impl-review and spec-completion-review compute the same guard locally in every gated file and, when ineligible (non-macOS, no rp-cli), steer only to codex/copilot/cursor (+ none). Explicit --review=rp / env / config / per-task overrides still resolve; eligible hosts render byte-for-byte as before.

2.5.3: RepoPrompt proposal gate + review-call hardening

Section titled “2.5.3: RepoPrompt proposal gate + review-call hardening”

/flow-next:plan and /flow-next:plan-review no longer offer the RepoPrompt path on hosts where it can’t run (non-macOS with no rp-cli on PATH) - explicit --review=rp / config still resolves as before - and the review skills now pin an explicit Foreground rule so agents never background a review CLI call and idle on a finished verdict.

Detail
  • RepoPrompt eligibility gate. RepoPrompt is a macOS-only GUI app, yet both skills proactively dangled the rp option in their interactive setup on every host - on Linux/Windows without rp-cli, picking it was a guaranteed runtime failure. Both now compute one POSIX guard - RP_ELIGIBLE ⟺ uname == "Darwin" OR rp-cli on PATH - and, when ineligible, drop every RepoPrompt proposal (plan’s research question defaults silently to repo-scout; plan-review steers only to the runnable codex / copilot / cursor + none). Suppression is not a ban: explicit --research=rp / --review=rp / FLOW_REVIEW_BACKEND=rp / review.backend=rp still resolve; eligible hosts render byte-for-byte as before.
  • Foreground rule for review CLI calls. Found in fn-78’s own autonomous dogfood: a worker subagent backgrounded its cursor impl-review and idled on the already-finished verdict (background completion doesn’t reliably resume a subagent). The backend CLI was flawless - 8/8 verdicts - so impl-review / plan-review / spec-completion-review and the worker agent now pin the calling discipline: one blocking foreground call, generous timeout, never background + monitor.

The scout subagents move off a frozen claude-sonnet-4-6 pin to family aliases matched to each task: the 8 pure config-scanners drop to fast, cheap haiku (Haiku 4.5 - which out-scores the gpt-5.4-mini the Codex mirror already runs them on), while only the 3 judgment scouts (spec-scout, claude-md-scout, docs-gap-scout) stay on sonnet; heavy agents keep opus, worker/pr-comment-resolver inherit. No version pins, cheaper + faster scouts, and Claude finally matches the FAST/INTELLIGENT tiering the Codex mirror already encoded.

flowctl now just works on Windows when python3 resolves to the Microsoft Store App Execution Alias stub - a 0-byte reparse point that’s on PATH but exits 9009 - by probing interpreter functionality (<cand> -c "import sys") instead of presence, across every invocation context (Git Bash / WSL, cmd.exe / PowerShell, Claude Desktop, native Codex / Cursor), with a companion flowctl.cmd launcher and no mac/linux regression.

Detail
  • Probe over presence. A shared resolver (scripts/lib/pick-python.sh) and the self-contained launchers probe interpreter functionality in order $PYTHON_BIN → py -3 → python3 → python; the 9009 stub is skipped even though it’s on PATH, while a machine with a working python3 (and no py launcher) still picks python3 first. The old launchers hardcoded exec python3 and the prior pick_python helper tested command -v (presence, not function) - both selected the broken stub.
  • Dual launcher. A flowctl.cmd batch shim ships alongside the extensionless bash flowctl, running the same probe under cmd.exe / PowerShell where the bash shebang is never honored (py -3 preferred). CRLF/LF pinned so Git Bash doesn’t regress.
  • init self-heal. flowctl init re-stamps both .flow/bin/flowctl and .flow/bin/flowctl.cmd, so an existing (pre-fix) install refreshes on the next init - no full re-setup. A broken bash launcher is reached via the new .cmd, a plugin auto-update, or py -3 .flow/bin/flowctl.py init.
  • Swept everywhere + covered. Ralph hooks, watch-filter.py, and the qa/prospect agent heredocs all resolve a working interpreter (Ralph mode requires Git Bash on Windows). A fake-9009-stub regression harness plus a real windows-latest CI job (proper .exe/.cmd stub) exercise both launchers against the stub; docs/troubleshooting.md + docs/platforms.md document the fix, the probe order, and both recovery paths (re-stamp via init, or disable the App Execution Aliases).

A fourth cross-model review backend - cursor (Cursor-billed cursor-agent CLI) - joins rp / codex / copilot; all agentic backends now read changed files from disk instead of embedding them (smaller, cheaper prompts); the review rubric itself gets eval-validated tuning - an always-on code-smell baseline lifts impl detection 7 → 10/10 at ~27% fewer prompt tokens, and plan reviews gain a spec-quality checklist (8.0 → 9.7); and per-task / per-spec review: overrides now route correctly instead of silently falling back to the project default.

Detail

Cursor review backend. A parity port of the copilot backend - no new review features, same Carmack-level criteria, same receipt schema, same verdict grammar, same --deep / --validate passes - wired through /flow-next:impl-review, /flow-next:plan-review, /flow-next:spec-completion-review, and /flow-next:setup. Select it the usual ways: flowctl config set review.backend cursor, FLOW_REVIEW_BACKEND=cursor, --review=cursor, or a per-task/spec cursor:<model>.

  • Cursor-billed, no extra key. Runs cursor-agent -p --output-format json --trust --mode ask against the workspace (read-only Q&A - it never mutates the tree). Reaches reviewer models the others can’t in one place: gpt-5.5-high (1M ctx, the default), the gpt-5.3-codex family, composer-2.5, Opus 4.8 thinking. Auth is your stored cursor-agent login or CURSOR_API_KEY.
  • Resume-only sessions. The first review omits --resume and persists Cursor’s generated session_id; a re-review resumes it (only when the prior receipt’s mode == "cursor" - a cross-backend receipt starts fresh).
  • Effort folds into the model name (Cursor convention), so a spec is cursor:<model> with no :effort rung - cursor:gpt-5.5-high, not cursor:gpt-5.5:high.
  • Triage judge unchanged. The opt-in LLM triage judge (FLOW_TRIAGE_LLM=1, default off) stays codex|copilot; with it off cursor reviews use the deterministic trivial-diff whitelist, zero extra dependency.

Review backends read changed files from disk. The agentic backends - codex, copilot, cursor - no longer embed changed-file contents (previously up to ~500 KB) into the reviewer prompt. They read from disk the way rp’s Builder already did (codex sandbox, copilot --add-dir, cursor --mode ask), so prompts are smaller and cheaper and cursor no longer trips its argv limit on non-trivial diffs. Verified equivalent on a ground-truth planted-bug test (codex’s own audit: QUALITY=PRESERVED). The per-backend FLOW_*_EMBED_MAX_BYTES budget knobs are removed.

Sharper, leaner review prompts. The Carmack rubric gains an always-on code-smell baseline (Fowler Refactoring ch.3 - Feature Envy, Data Clumps, Primitive Obsession, Long Method, Duplicated Code, …) on impl + standalone reviews, with its rubric blocks tightened and every machine-parsed marker preserved. Applied to every backend - codex/copilot/cursor and RepoPrompt. Eval-validated on a ground-truth corpus (correctness bugs + planted smells): detection rose 7 → 10/10 (the old rubric reliably missed Feature Envy / Data Clumps / Primitive Obsession) while the prompt shrank ~27% (−950 tokens), correctness detection held at 5/5, and clean code was not over-flagged - confirmed on both codex (GPT-5.5-high) and RepoPrompt. Plan reviews additionally gain a targeted spec-quality checklist (a stated test strategy, observability for async/batch work, each task sized-for-one-iteration and correctly dependency-ordered, non-functional requirements) - eval-validated 8.0 → 9.7/10 for +74 tokens, no over-flagging of good specs.

Per-task / per-spec review-backend overrides route correctly. A task’s review: <backend>:... (or a spec’s default_review) is now honored end-to-end: flowctl review-backend resolves the per-task/epic override above env/config (canonicalizing short/tracker handles first), and every review skill + /flow-next:work’s per-task worker passes it - so a task set to review: cursor:... under a codex project default actually reviews with cursor. Every backend command also defensively coerces a foreign stored spec to its own default, so an explicit --review=<backend> / flowctl <backend> always wins over a stored cross-backend spec instead of shelling a foreign model.

Copilot CLI 1.0.65 compatibility. The default copilot model moves gpt-5.2 → gpt-5.5, and gpt-5.2 / gpt-5.2-codex are dropped from the accepted set (1.0.65 rejects them), so copilot:gpt-5.2 is now rejected. Session creation is fixed for the CLI’s resume-only --resume change - the first call now uses --session-id (marker-tracked) and re-reviews resume it.

Tracker-sync gained GitLab and Jira as its 3rd and 4th providers, so teams on the dominant self-managed (GitLab) and enterprise (Jira) trackers could mirror Flow-Next specs to their board with zero special setup. The prose-driven provider implementation described in this historical entry was superseded by the deterministic flowctl tracker boundary; current behavior is documented on Tracker Sync.

Detail

The supported-tracker set became Linear, GitHub, GitLab, Jira. Current releases normalize these providers behind flowctl tracker; the skill retains semantic merge and recovery judgment.

GitLab (the 3rd tracker - a large share of self-managed and EU/regulated shops). Modelled on the GitHub adapter:

  • Historical GitLab implementation. At 2.4.0 the skill described direct glab and REST choices. Current releases resolve GitLab once, persist destination and capability facts under tracker.resolved, and execute provider operations through flowctl tracker.
  • Reduced-fidelity status, like GitHub. Open/closed plus a configurable board label, not a rich workflow.
  • License-gated dependency projection. depends_on_epics edges project as native is_blocked_by links on a Premium/Ultimate namespace; a Free or personal namespace (where the API returns 403 Blocked issues not available for current license) degrades to a directionless relates_to link plus a provenance-fenced <!-- flow:deps --> body block for direction.

Jira (the 4th tracker - the enterprise default). REST-only by design, the most adapter-specific weight of the four:

  • Historical Jira implementation. At 2.4.0 Cloud used API version 3 and Data Center / Server used version 2. Current releases pin both deployment families to API version 2 for plain-string body fidelity and migrate a legacy configured version 3.
  • No MCP. The official Atlassian MCP is read-mostly - it can’t transition status, update fields, or set links - so the bridge uses the REST + token path directly, headless-native with the fewest moving parts.
  • Workflow-aware status. A change goes through the transitions API against a configurable statusMap; an unmapped or unreachable transition defers with a receipt rather than forcing a lane. The fn-66 terminal invariant holds - a locally-done spec stays In Review until the PR is MERGED.
  • Current Jira fidelity. Dependencies project as native directional Blocks issue links. API version 2 keeps bodies as plain strings, and backlog enumeration runs through flowctl tracker wire list-open.

The new adapter behavior is documented on Tracker Sync.

/flow-next:pilot gains an opt-in backlog mode (pilot.autonomy=backlog, default off): instead of advancing one already-ready spec, pilot widens to a standing scheduler for the entire open backlog - enumerating flow specs + tracker issues, triaging the top dep-ordered item, and either advancing it one stage or surfacing a precise async question and parking it (ASKED). The consent boundary moves from before the loop to inside the loop, on block, while every safety boundary holds: it never authors a spec, never promotes, and never merges.

Detail

By default pilot’s consent boundary sits before the loop - it only picks from the already-ready queue. Backlog mode (flowctl config set pilot.autonomy backlog, or per-run --backlog / --auto) makes each tick enumerate everything open (flow specs via flowctl ready --all plus tracker issues at the promoted lane, unioned in from the tracker-sync adapter), select the top dep-ordered actionable item, triage it agentically, and advance it along the same plan → plan-review → work → [qa] → make-pr pipeline. When it can’t safely proceed it surfaces an async question into the spec’s ## Open Questions + a tracker comment and parks the item - “stuck” becomes a question a human answers async, not a stall.

  • Same single-tick conductor, widened left. One /loop//goal target, one verdict grammar (adds ASKED <id> (<n>), keeps NO_WORK/DEFERRED_TO_LAND verbatim), one mental model - not a new skill or command, and not a prospect-style idea generator (it manages the existing backlog).
  • Boundaries hold. Never authors a spec (a thin/missing spec is a surfaced “run /flow-next:capture or /flow-next:interview” gap); never sets the ready flag (promotion is the human’s board act; un-promoted items are skipped silently); never merges (land stays human-gated). Readiness stays the human’s explicit signal, never an agent-inferred score.
  • Substrate. A backlog-wide eligibility scan (flowctl ready --all → deterministic facts only), a per-tick decision log (flowctl pilot-log → the factory-efficiency readout), and a tracker-sync autonomy-parity fix + the async question-valve so a per-tick sync never hangs the loop. The agentic/deterministic line holds: flowctl enumerates + checks hard fields; the host agent judges and formulates the question.

Off by default - existing pilot/land/Ralph users are unaffected until they opt in.

2.2.0: QA pipeline stage + Cua native driver

Section titled “2.2.0: QA pipeline stage + Cua native driver”

Two opt-in additions to the autonomous pipeline: /flow-next:qa becomes a config-gated (pipeline.qa, default off) pilot stage that live-tests the complete build before make-pr, and Flow-Next Drive’s native rung gains the Cua driver + sandbox for provider-agnostic, headless/CI computer-use. Both augment, never replace, existing tooling.

Optional QA pipeline stage

/flow-next:qa already did the hard part - derive scenarios from the spec, drive the live app, file P0/P1/P2 findings, emit a qa_verdict - but it lived outside the build loop. fn-72 wires it in as an opt-in pilot stage: flowctl config set pipeline.qa on inserts a qa stage at the all-tasks-done juncture, so the autonomous span becomes plan → plan-review → work → qa → make-pr. Default off - with the gate off, pilot’s stage set is byte-for-byte unchanged.

  • Augments, never replaces. The app is already up on the dev’s machine during work, so this is the cheap first live pass that catches obvious runtime breakage before a human opens the PR. Like everything in Flow-Next it reduces human work agentically and surfaces problems to humans - it does not stand in for CI/staging QA or manual QA, which still happen downstream.
  • Lean + agentic, evidence-aware. Net-new flowctl is a single pipeline.qa config-key default - no new subcommand, engine, or persisted artifact. The host derives scenarios in-context and drives the local running app, reusing the existing executor. It reads work’s recorded evidence first and subtracts only AC proven by a deterministic re-runnable check (a real test/lint/build command), always live-running every runtime/UI/integration AC even when work narrated it done.
  • Surfaced, not loop-blocking. The stage is idempotent (a head_sha freshness gate runs it at most once per branch head) and the pilot gate routes on qa_outcome, not the Ralph-guard verdict projection: SHIP/NA/BLOCKED advance cleanly, and NEEDS_WORK still advances to the draft PR - make-pr surfaces the findings in a ## Live QA section, plus the bug-memory track and a tracker comment when the bridge is active. QA never hard-blocks the loop; merge stays the human’s + land’s decision.
  • Principled reversal. Pilot’s “QA is never a stage” is reversed only under the gate; capture/interview/resolve-pr/merge/release stay forbidden for their distinct loop-ownership / consent reasons.
Cua native driver rung

The native rung of the surface-aware driver ladder was served only by Computer Use (Codex CU / Anthropic Claude CU) - provider-locked, macOS/Windows-only, focus-stealing, and never reachable on a headless / CI / Linux path. trycua/cua (MIT) is added as a detected, opt-in driver with two surfaces, never a hard dependency:

  • Cua Driver (cua-driver mcp) - background computer-use on the local machine over an MCP server: no focus steal, accessibility-tree-based (drives structured element_index elements, not pixels), and provider-agnostic. On macOS the load-bearing TCC permission split is documented - Accessibility unlocks driving, Screen Recording unlocks screenshots - so when Screen Recording is absent the rung surfaces “AX-only evidence, no screenshot” rather than emitting an empty one.
  • Cua Sandbox - drives an app inside a disposable VM/container (any OS), the only native option on a headless/CI host with no display. Opt-in per run, torn down each run; local lume/QEMU/Docker is the default backend, the cua.ai cloud is explicit opt-in (bills + data-egress, never auto-selected).

Detect-and-instruct, never auto-install - the same consent rule /flow-next:map applies to clawpatch. The base install stays zero-dependency, agent-browser remains the only assumed-present driver, and flowctl never imports Cua. The default driving path (background cua-driver MCP) uses only MIT components; the optional cua-agent[omni] (ultralytics AGPL-3.0) / OmniParser (CC-BY-4.0) extras are documented and never auto-installed. A pass still completes with no Cua installed (fall to Computer Use → documented-limitation). No new skill or command - a rung, not a re-architecture; /flow-next:qa accepts cua-driver / cua-sandbox as evidence driver_rung values with no schema change.

2.1.3: Resolve-pr keeps null-state threads in scope

Section titled “2.1.3: Resolve-pr keeps null-state threads in scope”

/flow-next:resolve-pr now treats only literal true as resolved - GitHub/GraphQL can surface a newly-created unresolved inline thread as isResolved: null (not just false), and those Codex/Bugbot findings were being silently dropped; fetch observability (counts + previews across all three feedback surfaces) is now mandatory in full mode and watch loops.

Tracker-sync now reserves Linear Done for merge-confirmed PRs - an open PR maps to In Review, completion-review never completes the issue, and pilot never declares NO_WORK for an all-done spec that hasn’t shipped.

Detail

Done is a claim that the work shipped, so projecting it from local completion (all tasks done + completion-review SHIP) was a correctness bug - a spec with no PR could land on the board as Done and a human had to drag it back. The flow→tracker status map is now a function of (spec status, completion_review_status, **PR-merge-evidence**): terminal Done requires a GitHub MERGED probe result on every write path (automatic touchpoints and a manual reconcile, which can still recover Done once a merge exists). An open PR projects In Review on make-pr’s unconditional bridge-active link path; completion-review is now a verdict comment only; and land.merged - active by default when the bridge is active - is the sole Done driver. Pilot mirrors it: an all-done spec with no merged PR routes to make-pr, or reports the new DEFERRED_TO_LAND verdict when an open PR exists, instead of silently collapsing to NO_WORK. See Status lifecycle.

Land’s silence merge signal now recognizes a review bot’s clean-pass comment (naming the reviewed commit), not just formal reviews - so a Codex-reviewed PR with no findings actually merges instead of stalling at NEEDS_HUMAN.

Detail

Codex (and bots like it) only file a formal review when they have findings; a clean pass is an issue comment - "Didn't find any major issues. Reviewed commit abc1234" - that never reaches the reviews API land reads. So a converged-clean PR could sit unmerged forever. Under silence, land now also scans PR comments: a comment from an automated reviewer matching land.cleanReviewCommentPattern and naming the current head SHA counts as a head-current review. It only ever adds evidence - never overrides a formal review, an open thread, or a red check - and the SHA must be the current head, so a land-authored fix push still forces a fresh clean comment before merge. Set land.cleanReviewCommentPattern to an empty string to disable the comment path. Found dogfooding the fn-64 land. See The merge gate, precisely.

2.1.0: Dependency projection to the tracker

Section titled “2.1.0: Dependency projection to the tracker”

Tracker-sync now projects a spec’s depends_on_epics edges onto the board as blocked-by relations - on both Linear and GitHub - idempotently, provenance-tracked, and without ever clobbering a relation a human added by hand.

Detail

Dependency projection first shipped as skill-side provider prose. Current releases run flowctl tracker relate: Linear and Jira use native directional relations; GitHub records the blocked issue as a sub_issue hierarchy proxy with structured degradation; GitLab uses native blocked-by when the resolved capability allows it and otherwise preserves direction in a fenced body block.

Provenance over diff-reconcile. Neither platform records who created a relation, so Flow tracks the edges it created in a per-spec depRelations ledger (native) or the fenced marker (GitHub fallback). A relation Flow can’t prove it created is never removed; a ledgered edge a tracker user deleted is deferred (queued receipt), never silently recreated. The projected flag keys off the directed tracker edge, so a relinked issue reads un-projected.

Safe by construction. A dependency with no linked issue surfaces a named warning and the sync proceeds; a done dependency keeps its relation visible but never re-gates ready=true; self-edges are skipped and cycles project as independent direct edges (no traversal); unreachable transport writes a noop receipt and never blocks. New flowctl sync list-dep-relations / set-dep-relation / clear-dep-relation own the deterministic ledger plumbing. See Dependency projection.

Opt-in HTML artifact mode: capture, plan, and make-pr now also emit self-contained HTML render lenses - a spec visualizer for business and plan review, and a read-only PR review instrument - while markdown (and tracker-sync) stays 100% the source of truth.

Detail

Render lens, never record. One config key - flowctl config set artifacts.html.enabled true (OFF by default, offered once by /flow-next:setup) - switches the lifecycle skills into artifact mode. When active they load a shared disclosure reference carrying all generation rules plus an explicit anti-slop design contract (own instrument-panel house style, local-only fonts, zero external requests), and write self-contained single-file HTML to fixed paths under .flow/artifacts/<spec-id>/ - never timestamped, regenerated in place, never parsed back as state, each with a staleness stamp in the footer. Mode off (the default) means zero new steps, zero token cost, zero behavior change. See Visual Aids - specs.

The spec lens. One generation pathway, state-dependent rendering: /flow-next:capture renders the spec-only business-review view (thesis, acceptance criteria with source-tag provenance chips, boundaries, decision context); /flow-next:plan regenerates the same file with the plan layer - task dependency DAG with critical path and the R-ID → task coverage matrix. The spec markdown carries an idempotent artifact link line, replaced in place on every regeneration.

The PR lens. /flow-next:make-pr emits a read-only review instrument: diff-derived (never from commit messages), verified against the spec’s R-ID export - mismatches render as visibly flagged rows, warn-in-artifact, never blocking. It lands in one narrow chore(flow): pr artifact <spec-id> commit so the PR body’s SHA-pinned blob link resolves; --dry-run writes nothing, generation failure is non-fatal, and Ralph’s PR_URL= stdout contract is untouched.

Lavish annotation (optional). lavish-axi is detected on PATH and never required: spec artifacts open as browser annotation sessions, and feedback maps to edits of the markdown source followed by lens regeneration. Pull-only and session-spanning (annotations queue in ~/.lavish-axi/state.json and survive agent death). The PR lens never enters the annotate loop, and autonomous runs generate but never poll.

Breaking. The deprecated planSync.crossEpic config alias (1.x deprecation, readable through the 1.x line) is removed - use planSync.crossSpec.

New /flow-next:land skill - a cadence-driven, fully autonomous babysitter that takes the build loop’s draft PRs the rest of the way: CI kept green, automated reviews converged via resolve-pr, a gated explicit merge, spec close, and your project’s own release process - closing the lifecycle end to end.

Detail

The tick. Each /flow-next:land invocation discovers the open PRs the build loop authored (spec branch_name match and the make-pr breadcrumb - both signals required before any mutation; hand-opened PRs are never touched), walks each through a read-only gate tree - CI tri-state over all checks, a reviewer patience window anchored to the last push (land.patienceMinutes, default 30), unresolved threads, the review signal, mergeStateStatus - and takes at most one action class per PR: a bounded CI fix (land.ciFixBudget, default 3, with a durable flow-next:needs-human label on exhaustion), a /flow-next:resolve-pr dispatch, a mechanical rebase (any conflict hunk → honest BLOCKED), or the merge. Every tick ends with LAND_VERDICT=<MERGED|RELEASED|FIXING_CI|AWAITING_REVIEW|RESOLVING|BLOCKED|NEEDS_HUMAN|NO_WORK> prs=<n> pr=<url|-> reason="…" - worst severity across PRs, last line of output. Drive it on a cadence: /loop 30m /flow-next:land.

The merge gate. Land is the one confined exception to flow-next’s “no gh pr merge from skills” rule. It flips the draft to ready and merges explicitly - gh pr merge --squash --delete-branch --match-head-commit, never --auto - only after CI is green, threads are addressed, and land.reviewSignal is satisfied: silence (default - an automated review present + zero unresolved threads + the window elapsed; built for bot reviewers that never file formal APPROVEs), approve, or a named reviewer login. No automated review ever and no signal configured → it never merges unreviewed (NEEDS_HUMAN).

The tail. After merge: flowctl spec close (the build loop never re-selects merged work), the opt-in tracker.perEvent.land.merged touchpoint (issue → terminal state + verdict comment), then release-follow of your project’s own release docs (RELEASING.md et al.) with an idempotency probe - or stop at merge. A merged-but-unclosed spec re-enters idempotently. --dry-run reports the full gate classification with zero mutations.

Autonomous resolve-pr. /flow-next:resolve-pr now honors the mode:autonomous token (plus FLOW_AUTONOMOUS=1 env): needs-human cases become NEEDS_HUMAN: report lines instead of a blocking question, and the run ends with the machine-readable RESOLVE_PR_VERDICT=<RESOLVED|PENDING|NEEDS_HUMAN> threads=<n> fixed=<n> needs_human=<n> line land gates on. The 2 fix-verify cycle bound is unchanged, and the land.* config keys ship with seeded flowctl defaults.

Land was, fittingly, the first spec pilot drove end-to-end. See Going Autonomous for the three-loop picture.

1.13.0: Pilot: host-driven autonomous loop

Section titled “1.13.0: Pilot: host-driven autonomous loop”

New /flow-next:pilot skill - a single-tick build-loop conductor that advances one ready spec by one pipeline stage (plan → plan-review → work → make-pr) per invocation and ends with a machine-greppable PILOT_VERDICT line, so your host’s /loop or /goal owns the iteration instead of an external shell script.

Detail

The tick. Each /flow-next:pilot invocation selects the first open + ready spec with satisfied dependencies and no other-actor claims, classifies its stage from flowctl state, dispatches exactly one existing stage skill autonomously, verifies advancement (flowctl review-status fields + status transitions; a gh-confirmed OPEN PR URL for make-pr), and prints the terminal verdict: PILOT_VERDICT=<ADVANCED|NO_WORK|BLOCKED|NEEDS_HUMAN> spec=<id> stage=<stage> reason="<one line>". /goal validators are transcript-blind, so the evidence is echoed into the conversation and stop conditions are phrased against the grammar - e.g. /goal keep running /flow-next:pilot until it prints PILOT_VERDICT=NO_WORK, or stop after 20 turns.

Drivers. Claude Code /goal (v2.1.139+), Claude Code /loop (v2.1.72+; loops expire after 7 days), and Codex /goal (opt-in [features] goals = true, CLI ≥ 0.128.0, plain-text objective - no $skill-in-goal syntax). Caps and budgets belong to the driver; a tick has no timeout machinery. For unattended runs the rp backend works headlessly while the Repo Prompt app is running on the same Mac; on machines without it use --review=codex, --review=copilot, or --review=none.

Autonomous sub-skills. plan, work, and make-pr now honor a mode:autonomous token (plus FLOW_AUTONOMOUS=1 env) that suppresses user questions and picks safe defaults - work branches deterministically, make-pr forces a draft PR and hard-errors instead of prompting. Deliberately distinct from FLOW_RALPH: no ralph-guard hooks, no receipt choreography.

Don’t-thrash. A spec that fails to advance on two healthy ticks is taken out of selection (flowctl spec unready) with the reason in the BLOCKED verdict; re-blessing via flowctl spec ready clears its strikes. Pilot and Ralph are alternative drivers, never nested - pilot refuses to run under FLOW_RALPH.

A spec now carries a human-owned ready flag - the “complete enough to hand to an agent” gate that autonomous loops will consume - set via flowctl spec ready, projected one-way from your tracker (tracker.readyState), and surfaced through adoption-gated prompts in capture/interview/plan; invisible until you opt in.

Detail

The flag. flowctl spec ready <id> / spec unready <id> toggle a ready boolean on the spec record (default false) - orthogonal to status (a ready spec stays open through planning and work), human-owned or tracker-projected, never agent-inferred. Both verbs are idempotent (no write, no updated_at bump when the flag already matches), and the on-disk flag is lazy - written only after a toggle actually changes state, so non-adopters’ sidecars stay byte-identical. Every JSON read surface (show, specs, list) emits an explicit "ready": <bool>, and ready specs carry a [ready] badge in listings (badge only when set - no draft-noise). See Before planning - the ready flag.

Tracker projection. For tracker-connected repos, the /flow-next:tracker-sync discovery ceremony asks one optional, skippable question: which tracker workflow state means “ready for work”? (Linear: a workflow-state name, matched case-insensitive - names, not state.type; GitHub: a label, pre-created idempotently - present ⇒ ready, absent ⇒ not ready.) Every pull-side sync then projects that state onto the local ready flag - one-way, tracker → local, tracker authoritative - with change-only event-tagged receipts and graceful stale-config degradation (warn + noop receipt + flag untouched + sync continues). See Readiness projection.

Adoption-gated prompting. One in-use gate (≥1 ready spec OR tracker.readyState configured) governs every new prompt - non-adopters see zero new questions anywhere. /flow-next:capture and /flow-next:interview offer an optional end-of-authoring “Mark ready?” consent (default keep-draft; gated off when the tracker is authoritative). /flow-next:plan soft-checks readiness before the scout fan-out - warn, never block, default proceed. capture --rewrite resets a previously-ready spec to draft (a full re-authoring re-opens the blessing); interview refinement never auto-resets.

New regression suite (test_spec_ready.py) wired into CI; Codex mirror regenerated with all three net-new prompts verified.

1.11.0: Tracker-sync forcing + self-improving glossary

Section titled “1.11.0: Tracker-sync forcing + self-improving glossary”

Tracker lifecycle touchpoints can no longer silently not fire - receipts are event-tagged and work/capture/make-pr end every run with a read-only sync check + one bounded retro-fire - and the glossary now compounds through normal work (prime seeds it, capture adds to it, plan/work/review actually read it).

Detail

Tracker-sync became observable and forcing. This release added event-tagged receipts, the read-only flowctl sync check audit, one bounded retro-fire, and the mandatory Tracker sync: summary slot. Its original rule that flowctl contained no tracker mutation code was superseded by the deterministic flowctl tracker command surface; the receipt and audit semantics remain current.

Self-improving glossary. The same principle - the system gets better as you use it, never via a manual ceremony - applied to project vocabulary: /flow-next:prime seeds GLOSSARY.md from the repo when it’s absent or a husk (~10-20 evidence-backed terms, read-back gated, never rewrites a populated glossary), /flow-next:capture joins interview as a writer (new conversation-surfaced terms offered at read-back), and the read path widens to where wrong-concept errors get built: repo-scout / context-scout surface request-matched terms (max 5, budget-capped), the work worker’s re-anchor reads task-relevant terms, and impl-/plan-review prompts gain a Vocabulary criterion. Every gate is total_terms == 0 → silent skip. The compounding surfaces (memory, glossary, decisions, strategy) now have a dedicated Self-improving page, a STRATEGY.md track, and a “Self-improving” pillar in the redesigned six-pillar hero grid.

New regression suites (test_sync_check.py, --event coverage in test_tracker_receipts.py) wired into CI; Codex mirror regenerated.

The plugin homepage (Claude marketplace + the Claude/Codex plugin manifests, and the Codex websiteURL) now points at https://flow-next.dev instead of the stale mickel.tech/apps/flow-next. .cursor-plugin was already correct; the rest were drift. The flow-next-tui package homepage and the README “Visual overview” doc row were aligned too. author.url / owner.url (personal site / GitHub) are unchanged.

flowctl impl-review and console output no longer crash on a non-UTF-8 source subtree or a legacy console codepage (e.g. Windows cp1252).

Detail

Read side (#167). flowctl copilot impl-review could abort with UnicodeDecodeError on a repo containing a non-UTF-8 file anywhere in the tree. find_references() (behind gather_context_hints) ran git grep repo-wide and decoded hits with a hard encoding="utf-8" and no errors= - so a single legacy cp1252 file (e.g. a German C/C++ subtree carrying ü / ä / ö / ß) aborted context gathering even when every file you actively edited was UTF-8. The collector now captures git grep output as bytes and decodes with errors="replace". Reported with measured data by VGottselig (304 of ~5400 C/C++ files non-UTF-8 in a large Windows CAD codebase).

Write side (#167). flowctl now forces its own stdout/stderr to UTF-8 at startup, so non-ASCII output (→, umlauts) no longer aborts on a legacy console codepage (UnicodeEncodeError: 'charmap' codec can't encode character '→'). This removes the need for the PYTHONIOENCODING=utf-8 workaround.

/flow-next:work Verify-Completion recovery. When the host drops a long-running worker’s completion report, phase 3d no longer blocks waiting for a result that will never arrive: it diagnoses from ground truth (flowctl show + git log + git status) and classifies - already done → plan-sync; code present but unfinished → re-anchoring continuation worker; nothing landed → retry.

8 read-only scout/analyst agents got a feature-preserving output budget - ~40-71% leaner output into the planner/work-loop context, with accuracy held (proven by per-target evals + an end-to-end smoke).

Detail

Rolled the external “autoresearch” eval loop (baseline → one mutation → keep-if-better ratchet) across the read-only agents whose free-form output flows into the planner and work-loop context. Each gained a hard output budget - the reductions are at runtime (the rendered output), not in prompt size:

  • repo-scout (83→100% on its eval set, ~40-50% leaner) · context-scout (60→93%, ~60-70%; dropped the prescribed Code-Signatures block) · flow-gap-analyst (~50-70%, 26/27 gaps held) · quality-auditor (~63%) · spec-scout (No-Relationship → count, scale-robust) · docs-scout (~48-69%) · github-scout (~71%, the biggest) · practice-scout (~52%).

Feature-preservation is the guarantee, not a hope. A mutation was kept only if a per-target coverage/accuracy eval held (the ratchet): grounding (context-scout’s cited paths test -f-verified against a real 442k-LOC app repo), findings (quality-auditor against a 7-planted-issue testbed - Major bug + all slop still caught, clean code stays ✅), gaps (per-input answer keys), and docs/APIs/gotchas (the “pointer-not-paste” rule: name the API inline, drop code blocks, the link carries depth). The leaner research scouts even surfaced extra real issues a verbose baseline missed (a current CVE; an extra security gotcha).

End-to-end verified. The optimized scouts → a planner produced a correct, ship-quality build plan for a deliberately hard, cross-cutting feature (org-scoped rate limiting) reading only the budgeted scout output - features preserved at the consumer level, not just at scout-output level. The method lives in agent_docs/optimizing-skills.md.

Also: /flow-next:make-pr shed stale build-scaffolding archaeology (render output byte-identical). /flow-next:capture is unchanged - a trim was tried and reverted (the ratchet caught a routing regression), and its no-silent-overwrite guard was verified intact. No flowctl logic touched; Codex mirror regenerated.

1.9.1: Cursor setup detection + tracker merge-base

Section titled “1.9.1: Cursor setup detection + tracker merge-base”

/flow-next:setup now detects Cursor instead of mis-treating it as Codex, and a comment-first tracker auto-link snapshots its merge-base so a later sync can’t fast-forward over tracker edits.

Detail

Cursor setup detection. Setup keyed platform off plugin-root env vars (DROID_PLUGIN_ROOT → Droid, CLAUDE_PLUGIN_ROOT → Claude Code, else → Codex). Cursor exposes neither, so it fell into the Codex branch - writing the $flow-next-* Codex command syntax + running .codex/ setup, while the installer advertises /flow-next:*. Setup now branches on CURSOR_AGENT + the .cursor-plugin/plugin.json manifest + no codex/ mirror dir → PLATFORM=cursor, applied at every platform-branch point (detection, docs-status template, the Docs question, and the write mapping): it writes the /flow-next:plan snippet into AGENTS.md (which Cursor reads), resolves flowctl via .flow/bin/flowctl, and skips the Codex-only .codex/ copy. The triple guard is hardening against CURSOR_AGENT being inherited by child processes (Codex launched from a Cursor shell) and against the shared repo source tree carrying all manifests - and the installers now --delete-excluded / Remove-Item excluded paths so the codex/-absence proof holds on re-install. Hardened across five rounds of automated cross-model review.

Tracker merge-base snapshot. When the first lifecycle touchpoint for an unlinked spec was a comment op, create-if-unlinked attached the tracker id but didn’t snapshot the merge-base - leaving the issue base-less, so a later body sync could fast-forward and silently overwrite tracker-side edits. The auto-create path now set-merge-base (both halves) + set-last-synced at create time (the written issue body is the render, so the base is exact).

Flow-Next now installs into Cursor via a one-shot local plugin (./scripts/install-cursor.sh), and a lifecycle event on an unlinked spec now creates + links the tracker issue first instead of silently no-opping.

Detail

Cursor support (macOS / Linux + Windows). Cursor has its own .cursor-plugin/ plugin namespace and does not auto-read Claude Code plugins the way Grok Build does, so Flow-Next ships a Cursor-native manifest plus a one-shot installer on both platform families - ./scripts/install-cursor.sh (rsync) and install-cursor.ps1 (robocopy). Both copy the plugin into ~/.cursor/plugins/local/flow-next as a real directory (Cursor’s loader rejects symlinks escaping ~/.cursor/), exclude the Codex mirror + tests, and are a re-runnable snapshot. Verified end-to-end, multi-agent flows included - a full /flow-next:plan fanned out the scouts in parallel and drove flowctl to create the spec + tasks; flowctl resolves via the project-local .flow/bin/flowctl. Caveats: no grouped plugin card and the slash autocomplete under-lists the commands (both cosmetic - they run when typed), and Ralph autonomous mode is unsupported (Cursor’s afterFileEdit / beforeShellExecution hooks don’t map to the Claude PreToolUse + Bash|Execute matchers the Ralph guard relies on). See Install → Cursor.

Tracker create-if-unlinked. Previously only capture flow-first-pushed an unlinked spec; every other lifecycle touchpoint (plan / interview / work / make-pr / resolve-pr / completion-review) no-op’d when the spec had no tracker id, so starting a spec with /flow-next:plan left it orphaned from Linear / GitHub. Create-if-unlinked is now part of the flowctl tracker sync facade. unlink remains the only operation that no-ops on an unlinked spec; provider failures return a structured class.

New /flow-next:qa drives the running app like an unforgiving real user - deriving its test scenarios straight from the spec, filing P0/P1/P2 findings with evidence, and ending with a YES/NO ship verdict.

Detail

Every other Flow-Next review is static - impl-review, spec-completion-review, quality-auditor, code-review read code or specs. /flow-next:qa is the live-app gate: it drives the deployed app via Flow-Next Drive’s surface-aware driver ladder (it never re-implements driving) and is forbidden from marking PASS by reading source - a scenario passes only on captured evidence (screenshot / console / URL).

The differentiator vs spec-less QA tools: scenarios derive directly from the spec - acceptance criteria → scenarios, R-IDs → a coverage table, boundaries → what not to test, decision context → expected behavior - so the host already encodes intent instead of reconstructing it. Findings feed the bug memory track (track: bug, with overlap dedup) and can be promoted to specs/tasks. The pass emits a qa_verdict receipt with four outcomes (SHIP / NEEDS_WORK / BLOCKED / NA) projected onto the review-receipt enum, so it can feed spec-completion-review - “does the live app satisfy the AC, not just the code?”. Runs interactively and autonomously; not a hard Ralph-block; opt-in tracker verdict-post via tracker.perEvent.qa. Requires a live deploy + a driver - with neither it surfaces a BLOCKED verdict rather than failing; a spec with no driveable UI yields a clean NA. The QA discipline (P0/P1/P2 taxonomy, evidence rules, session hygiene) is a lean, credited borrow from Ray Fernando’s running-bug-review-board (Apache-2.0).

1.7.1: Codex delegation: cheaper on non-Claude hosts

Section titled “1.7.1: Codex delegation: cheaper on non-Claude hosts”

The opt-in Codex delegation reference no longer loads into a Codex / Droid / OpenCode session - the host-platform check moved into the cheap value-check, so delegation short-circuits before the ~45k reference is ever read. Byte-identical for Claude Code users.

1.7.0: Optional Codex delegation for /flow-next:work

Section titled “1.7.0: Optional Codex delegation for /flow-next:work”

/flow-next:work gains an opt-in delegate:codex mode that offloads code implementation to codex exec (gpt-5.5/medium) while Claude keeps orchestration, review, and all git - a second efficiency lever that offloads work, not prompt size.

Detail

When activated (delegate:codex arg or flowctl config set work.delegate codex), the Claude host stays the orchestrator - plan-reading, review, all git, and decisions - and delegates the token-heavy implementation to codex exec. Default model gpt-5.5, default effort medium, with proven per-batch risk escalation. It’s a different lever than prompt-trimming: it moves implementation tokens to a separate Codex budget.

Strictly opt-in and progressive-disclosure: with delegation off (the default) the work flow is byte-identical to before - one cheap config get, zero new steps. All mechanics live in a reference loaded only when active. The safety surface is the headline work: a one-time sandbox consent; a value-aware recursion guard (the flow-next CODEX_SANDBOX=auto review knob never trips it); mandatory --output-schema with MCP isolation (--ignore-user-config); a deterministic result classifier + sanitized scoped rollback that never touches .flow/; “Codex never touches git” enforced by a post-run HEAD-unchanged assertion (not just the prompt); and a host-owned 3-strike circuit breaker that always falls back to standard mode. The ralph-guard PreToolUse hook is rebuilt to a tokenized argv allowlist that admits only the full canonical delegation shape. Runs in interactive and Ralph mode (consent pre-granted in config for headless). Configure via the work.delegate* keys; see the flowctl reference.

Hooking up the tracker bridge via /flow-next:tracker-sync now activates the whole lifecycle pipeline by default - you opt out of events, not in.

Detail

Previously every tracker.perEvent.* touchpoint defaulted off, so after the discovery ceremony you had to opt each lifecycle event in by hand. That inverted the intent - connecting a tracker means you want it kept in sync. The discovery ceremony now activates every event on confirmation: capture / interview / plan → reconcile, work.firstClaim → push, work.done / makePr / resolvePr → comment, completionReview → reconcile. Exclude events at ceremony time, or turn any off later with flowctl config set tracker.perEvent.<event> off.

The accidental-enable guard is preserved: the config schema default for each leaf stays off, so a bare tracker.enabled=true set by hand or a script - without running the ceremony - fires no lifecycle-event sync (make-pr’s unconditional PR↔issue link is the one exception). Activation is ceremony-gated, not flag-gated. No config-schema change; docs updated across every surface.

.flow/sync-runs/ (per-run tracker-sync receipts) is now auto-gitignored, and flowctl’s managed .flow/.gitignore self-upgrades so existing repos pick up new patterns on the next init.

1.5.2: Tracker-sync projects the full spec

Section titled “1.5.2: Tracker-sync projects the full spec”

/flow-next:tracker-sync now guarantees the issue body mirrors the entire spec - a render guardrail stops the host agent from pushing a summarized body instead of the full projection.

Fresh /flow-next:setup now ships the tracker-sync CLI reference it dropped in 1.5.0, a CI guard keeps the bundled template and the dogfood copy in lockstep, and a flaky Windows migration race is fixed.

Detail
  • usage.md shipped without tracker-sync docs. 1.5.0 added the flowctl sync / --tracker-first command block to the repo’s lived-in .flow/usage.md but not to the bundled template /flow-next:setup actually copies - so every fresh setup documented the whole CLI except the tracker-sync bridge that shipped in the same release. The canonical template is now byte-synced (Codex mirror regenerated).
  • Drift guard so it can’t recur. A new parity test hard-asserts .flow/usage.md ≡ its canonical setup template (and .flow/templates/spec.md ≡ templates/spec.md) across the ubuntu/macos/windows CI matrix. Edit the dogfood copy and forget the template → CI fails instead of consumers getting stale docs.
  • Flaky Windows CI fixed. Parallel migrate-rename could leave a concurrent writability probe’s temp file (.rw-probe-*.tmp) visible to the backup copy’s directory scan, then have it vanish before the copy opened it (FileNotFoundError) - a TOCTOU that Windows’s slower unlink widened. The backup copy now skips those transients and tolerates any file that disappears mid-copy.

New /flow-next:tracker-sync projects a spec to a Linear or GitHub issue and reconciles body, status, and comments two-way - projection, not coordination - and a tracker key (wor-17) now resolves as a spec id everywhere fn-NN does.

Detail

/flow-next:tracker-sync mirrors a .flow/specs/<id>.md spec onto a tracker issue (Linear first, GitHub next). Projection, not coordination: the spec stays the single source of truth and the quality layer; the tracker is a co-editable mirror that never drives flow state or spawns agents (contrast OpenAI Symphony, where the board is the control plane). It is distinct from /flow-next:sync, which is plan-sync. New docs: Tracker Sync.

  • Discovery ceremony (detect → surface → ask → never-assume) probes a Linear MCP / LINEAR_API_KEY / GitHub auth / a Jira host and writes tracker.* config only on confirmation (env > config > ask, the same ladder as flowctl review-backend). The bridge is off until explicitly enabled and active when tracker.enabled == true or tracker.type ∈ {linear, github}.
  • Historical provider implementation. The first release encoded provider request choices in skill prose. Current releases persist normalized resolution facts and capabilities under tracker.resolved, execute deterministic operations through flowctl tracker, and return structured error classes instead of a prose-selected no-op path.
  • Hybrid id model: tracker-first specs are canonically wor-17-slug (tasks wor-17-slug.M; bare wor-17 resolves); flow-first specs keep fn-NN plus a resolvable WOR-17 display alias. show / work / plan wor-17 resolve case-insensitively, the native fn- scheme is reserved, one tracker team per repo, and ids never rename on link. See Spec & task ids.
  • Seven lifecycle skills gain opt-in tracker-sync touchpoints - capture, interview, plan, work (first-claim + done), make-pr, resolve-pr, spec-completion-review. Each tracker.perEvent.* leaf defaults off; the no-tracker workflow is unchanged and remains the documented default.
  • PRs are Diffs-ready. When the bridge is active, make-pr unconditionally links the new PR to its tracker issue (no makePr opt-in). For Linear it writes a non-closing Ref WOR-N (plus a rich attachmentLinkURL on the GraphQL transport) so the PR renders as a Linear Diff inside the issue; for GitHub it is a native Refs #N cross-link. Non-closing is deliberate - merge never auto-completes the issue, spec-completion-review owns Done.
  • make-pr now creates the PR autonomously - no confirm gate. Invoking the skill is the intent; the body is deterministic; the default is a reversible smart-draft. --dry-run prints the body without creating, --ready / --draft set the draft state.
  • Setup proposes the bridge. /flow-next:setup never touches tracker config (keeping the zero-dep base clean) and now proposes running /flow-next:tracker-sync as an optional next step when it finishes.
  • Ralph-safe: every run emits a receipt; genuine conflicts queue to the review deferred-findings sink rather than block. An always-ask tiebreak resolves to queue in autonomous mode.

Sync-engine shape (discovery ceremony, per-item lastSyncedAt, surface-diffs-never-overwrite) adapted from Ray Fernando’s rayfernando-skills running-bug-review-board issue-trackers.md (Apache-2.0). Thank you, Ray.

1.4.0: flow-next-drive surface-aware automation

Section titled “1.4.0: flow-next-drive surface-aware automation”

The browser skill is renamed flow-next-drive and rebuilt as a surface-aware driver ladder - it detects the UI surface (web, Chromium-backed desktop, or true-native app) and picks the best available driver, degrading gracefully.

Detail

The skill is no longer hardwired to a single browser driver. It detects the target surface and branches: (a) web app → web ladder; (b) Chromium-backed desktop app (Electron / Windows WebView2) → the same web ladder, attaching over CDP to the app’s remote-debugging port (agent-browser --cdp <port> / --auto-connect; chrome-devtools-mcp --browser-url); (c) true-native / non-CDP surface (macOS AppKit/SwiftUI, or a webview exposing no CDP - e.g. macOS WKWebView / Tauri-on-macOS) → Computer Use. All surfaces share one universal flow (observe → snapshot → act on fresh refs → verify → capture → release); only the actuation differs.

The web ladder, in priority order: agent-browser (default rung, the only assumed-present driver, CDP-based + headless-safe) → chrome-devtools-mcp (auto-wait + attach-to-real-signed-in-Chrome) → Playwright → cursor-ide-browser MCP → manual screenshot relay. The native rung is Computer Use - driver-agnostic across Codex Computer Use and/or Anthropic “Claude” Computer Use (the API computer tool via its own harness); detected and optional, never a hard dependency, never on a headless path. When no Computer Use is present, a Chromium-backed app still drives via the web-ladder CDP attach; a genuinely native app documents the limitation rather than fails. The existing agent-browser references fold into the default-rung reference - no capability regression.

The driver ladder + universal-flow structure is adapted from Ray Fernando’s rayfernando-skills running-bug-review-board skill (Apache-2.0). Migration: /flow-next:browser is gone - the skill is now /flow-next:flow-next-drive (and flow-next-drive on the Codex mirror, fixing the prior agent-browser rename that collided with the user’s global agent-browser skill and Codex-native browser skills). An orphaned browser / agent-browser skill in a cached install auto-clears within ~7 days or immediately by deleting the stale cached marketplace directory under ~/.claude/plugins/cache/<marketplace>. See the Drive skill page.

The review-output R-ID parser now preserves single-letter suffixes (R4a / R4b) - they were being silently dropped from the coverage gate and fix-loop targeting.

Detail

parse_unaddressed_rids read R-IDs from a reviewer’s Unaddressed R-IDs: summary line (_extract_rids) and from the ## Requirements coverage table fallback using bare \bR(\d+)\b. fn-49.1 (1.2.1) taught the spec acceptance-criteria parser the R\d+[a-z]? suffix form but left this review-output path behind - so a reviewer reporting Unaddressed R-IDs: [R4a, R4b] parsed to ['R5'], dropping the suffixed IDs from the R-ID coverage gate and fix-loop targeting. Both review-output regexes are now \bR(\d+[a-z]?)\b, in lockstep with the spec parser; multi-letter suffixes (R4ab) and separators (R-4) stay rejected. New test_unaddressed_rids_parser.py (10 cases) wired into the ubuntu/macos/windows CI matrix. Surfaced by a live impl-review A/B - the current review prompt caught it (the experimental slop rubric being tested was shelved as unproven).

Scouts fall back to the bundled .flow/bin/flowctl so .clawpatch/ feature enrichment fires even when dispatched subagents don’t inherit the plugin-root env var.

Detail

When repo-scout / context-scout run as dispatched subagents they may not inherit CLAUDE_PLUGIN_ROOT / DROID_PLUGIN_ROOT, which left their Step 0 flowctl repo-map list --json call resolving to a broken /scripts/flowctl - the scout then silently grep-degraded and features_anchored never fired even with a populated .clawpatch/. Both scouts now fall back to the bundled .flow/bin/flowctl ([ -x "$FLOWCTL" ] || FLOWCTL=".flow/bin/flowctl"), so a /flow-next:setup-installed repo resolves regardless of subprocess env - across Claude Code, Factory Droid, and Codex. Also makes sync-codex.sh’s agent-body fallback injection idempotent (no duplicate line in the Codex mirror) and the scout-fallback contract test hermetic (runs in a throwaway git repo so a local dogfood .clawpatch/ can’t break it). Surfaced by full live end-to-end testing - mapping Flow-Next’s own repo via --source=agent (codex) produced 9 features, then the scout enrichment was exercised against them.

/flow-next:map now explains why the provider-free mapper found 0 features on an unconventional repo, and points at --source=auto|agent.

Detail

Live-testing on Flow-Next’s own repo showed clawpatch’s heuristic detectors target conventional app/framework layouts (npm bins, Next.js routes, Python packages, Rails / Laravel / Django, Go / Rust, JVM, .NET, SwiftPM, Phoenix); a plugin + markdown-skill + flowctl.py-CLI + bun-TUI repo matches none, so heuristic returns 0 features while clawpatch flags coverage as “weak.” The Phase 5 summary previously printed a silent “Mapped: 0 feature(s)” - it now explains the conventional-layout targeting and points at --source=auto (heuristic-first, provider only if weak) or --source=agent (always provider-backed; needs CLAWPATCH_PROVIDER + tokens). For reference, --source=agent via codex produced 9 well-scoped features for Flow-Next’s repo (Flowctl CLI Core, Ralph Guardrails, Flow Memory System, TUI shell/theme/integration, Plugin Packaging, Strategy/Docs). Also root-ignores .clawpatch/ in the dev repo so the local feature map never gets staged. See the skill page for the conventional-vs-unconventional note.

The /flow-next:map install hint is now conditional and pnpm-version-agnostic - it no longer presumes an install already happened.

Detail

Live-testing 1.3.0 (clawpatch never installed, pnpm 10) showed the hint presumed an install had already happened (“install succeeds but PATH unchanged”) and attributed the PATH wiring to pnpm v11 specifically - misleading for a first-time or pnpm-10 user whose global bin sits at ~/.local/share/pnpm. Now conditional and version-agnostic: “if you already ran pnpm add -g clawpatch and still see this, pnpm installs globals under $PNPM_HOME and needs a one-time pnpm setup.” Same correction in docs/troubleshooting.md + the skill page above. Logic unchanged.

New opt-in /flow-next:map skill wraps clawpatch for a semantic feature index that scouts and /flow-next:prime can read - flowctl core stays zero-dep.

Detail

Wraps clawpatch map to produce a semantic feature index (~20 languages, persisted at .clawpatch/features/*.json, Zod-validated schemaVersion: 1). Opt-in convenience throughout - flowctl core never imports or requires clawpatch; the skill is the only flow-next surface that touches it. Default invocation is provider-free (--source heuristic, zero LLM calls, deterministic mapper); --source auto|agent flows through as passthrough. Missing binary → skill prints pnpm add -g clawpatch install instructions verbatim and exits cleanly (no auto-install); pnpm-installed-but-not-on-PATH → skill prints the PNPM_HOME bin/ hint. Single-source SUPPORTED_CLAWPATCH=">=0.4.0 <0.5.0" version pin lives in skill prose; outside-range → one-line stderr warning + degrade, never block. Ralph-block (decline-to-run, no receipt write) under FLOW_RALPH=1 / REVIEW_RECEIPT_PATH.

Companion flowctl repo-map list / show / since-ref reader subcommands parse the index directly; readers bypass ensure_flow_exists() and gate on .clawpatch/ presence instead so the prime detection branch works without special-casing. since-ref uses three-dot <ref>...HEAD semantics (fixed during PR review - two-dot <ref>..HEAD polluted overlap results with upstream-only advancement). repo-scout and context-scout call flowctl repo-map list --json as Step 0 when .clawpatch/ is present and emit an optional features_anchored: [...] field with a last_mapped staleness timestamp; scouts remain useful with the existing grep/glob path when .clawpatch/ is absent (fallback contract is load-bearing). /flow-next:prime adds a DE7 informational sub-criterion under Pillar 5 (Dev Environment) surfacing /flow-next:map in Top Recommendations when the index is missing - pillar count stays at 8, scored criteria stay at 48, total criteria become 48 → 49.

The feature index is local-only by design: .clawpatch/.gitignore skeleton is * + !.gitignore, so the index is regenerable-per-developer rather than committed (avoids PR review noise + merge conflicts on a pre-1.0 schema; full sharing-contract trade-off table on the skill page). Docs: platforms.md gains an “Optional skill requirements” section; troubleshooting.md gains a clawpatch-failure-modes section. GLOSSARY.md entries added for “feature map” and “features_anchored”. CI matrix wires test_repo_map.py (22 tests) + test_scout_fallback_contract.py (14 tests) + test_pnpm_home_hint_prose.py (5 tests) + map_smoke_test.sh (75 cases) on ubuntu / macos / windows using checked-in fixtures (no Node 22+ or clawpatch needed in CI).

Two spec export-cognitive-aid parser bugs fixed so /flow-next:make-pr bodies stop silently dropping content.

Detail

(1) The R-ID parser regex R\d+ was extended to R\d+[a-z]? so sub-scoped sibling criteria like R4a / R4b (introduced by /flow-next:capture when revising specs in-flight) are no longer dropped from acceptance_criteria / uncovered_r_ids; the suffix form is now blessed in templates/spec.md as canonical. (2) The memory_during_spec time-window filter now has a deterministic null-safe fallback chain (spec.created_at → earliest task created_at → branch first-commit via git log {base}..{branch}) so memory entries surface correctly when spec.created_at is null. Both surfaced during fn-48’s make-pr where PR #146 carried workaround prose. Two further bugs in the branch-first-commit fallback (returning the branch tip not the root commit; walking inherited mainline history) were caught by chatgpt-codex-connector[bot] review on PR #147 and fixed with regression tests. Unit suite 624 → 646.

Review skills split by backend so codex/copilot load only their own slice (impl-review 1126 → 70 LOC on codex); FLOWCTL prelude consolidated.

Detail

spec-completion-review drops from 645 to 41 LOC on codex (14×), impl-review from 1126 to 70 LOC (16×). RP keeps its cohesive prompt template since it only loads under the RP backend. resolve-pr was evaluated and kept inline (its parallel-vs-serial divergence sits below the 50-line split threshold codified in agent_docs/adding-skills.md). Also drops the dead DROID_PLUGIN_ROOT:-CLAUDE_PLUGIN_ROOT fallback from the Codex mirror’s FLOWCTL prelude (neither env var is set in Codex) and consolidates the canonical prelude to once-per-skill-file. Mechanical refactor only - bash, gating, and verdict semantics unchanged; smoke 127/2 baseline-equivalent. Factory Droid contract re-verified 2026-05-25 - DROID_PLUGIN_ROOT fallback + Bash|Execute matcher stay; .factory-plugin/plugin.json fallback dropped as dead code per Factory’s Claude-Code interop guarantee.

flowctl init no longer silently flips pre-1.1.3 users’ planSync.crossEpic from on to off - it mirrors the legacy value to the canonical planSync.crossSpec key before the default-merge when canonical is absent. Legacy key preserved through 1.x.

.flow/usage.md template promoted to a comprehensive CLI reference (100 → 212 lines) - adds status, config get/set, per-spec/task set-backend, checkpoint, ralph control, lifecycle commands, and a corrected file-structure diagram.

1.1.9: Copilot review on Windows (real fix)

Section titled “1.1.9: Copilot review on Windows (real fix)”

flowctl copilot *-review now works on native Windows by delivering the prompt via stdin (subprocess.run(input=…) with --session-id / --resume), sidestepping the CreateProcessW 32,767-char argv cap entirely. Verified by a real-subprocess Windows CI smoke round-tripping a 60 KB prompt. Supersedes the 1.1.8 WSL workaround. Upstream: github/copilot-cli#3398.

Fail-fast guard + WSL pointer for the Windows Copilot argv-cap failure (cryptic OSError winerror 206 before; clean error after). Reported by Simon Flauger (SEMA-CAD). Real fix landed in 1.1.9.

Stripped request_user_input from 6 Codex-mirror SKILL.md frontmatters - fn-45 rewrote the prose reference but left it in frontmatter, so Codex agents called the unavailable tool and reintroduced the Default-mode failure. sync-codex.sh guard tightened so it can’t regress.

1.1.6: prime ruleset-based branch protection

Section titled “1.1.6: prime ruleset-based branch protection”

/flow-next:prime SE1 now detects ruleset-based enforcement in addition to classic branch protection, so GHE Enterprise repos protected via repo / org / enterprise rulesets correctly show SE1 ✅. Reported by Georg Keller (SEMA-CAD).

Removed deadline / time-budget / sprint-cadence questions from /flow-next:interview --scope=business (agents can’t estimate their own work, and time pressure collapsed the interview into brutal-prioritization). MVP-scope cuts reframed by feature value; budget envelope scoped to infra / vendor / licensing.

1.1.4: Canonical ## Acceptance Criteria heading

Section titled “1.1.4: Canonical ## Acceptance Criteria heading”

## Acceptance Criteria is the canonical spec heading. Parsing stays tolerant of legacy ## Acceptance and lowercase ## Acceptance criteria - existing specs need no migration.

1.1.3: crossSpec alias + SPEC.md discovery

Section titled “1.1.3: crossSpec alias + SPEC.md discovery”

Cross-spec plan-sync aligned on planSync.crossSpec. Repo-root SPEC.md / spec.md template discovery for project-customized spec scaffolds; /flow-next:setup can opt into a root SPEC.md without clobbering custom templates.

The 1.0 line that stabilized the vocabulary and core workflow:

  • Spec vocabulary stabilized (Spec / Task / R-ID / Handover / Receipt / Ralph)
  • Symmetric --scope=business|technical|both interview
  • Source-tagged capture with mandatory read-back
  • PR-as-cognitive-aid generation
  • Agent-native memory audit and migration
  • PR feedback resolver
  • Strategy and glossary grounding
  • Ralph autonomous mode with receipts

Maintaining this page (for contributors)

Register first: entries are customer-facing. Use one narrative order: human outcome first, changed journey second, portable data and implementation details third. Lead with what the release gives the reader and its number, then what changes when they upgrade. The development story (what was tried and dropped, what measured worse first) stays out; a bound sits beside the number it bounds. For review features, make the human reviewer the protagonist: orientation, logical review journey, risk focus, evidence, and retained judgment. Machinery stays in the repository changelog or an “Under the hood” tail. Upgrade actions open the detail block, imperative. Numbers are outcomes (”30s to half a second”), not inventory (LOC, test counts). Candid about bounds and what did not change; zero hype; plain hyphens.

Hard rejection test: hide the technical tail. A user must still be able to explain why the release matters, how their workflow changes, and what control they retain. Rewrite any entry whose title or opening sentence leads with a command, schema/type name, artifact, fixture, parser, hash, or benchmark. Concrete capability creates excitement; adjectives do not.

What belongs here - human-readable release highlights:

  • new slash commands
  • changes to spec or task semantics
  • review receipt changes
  • migration requirements
  • breaking or deprecated behavior
  • docs, team workflow, or Ralph changes that alter how people should work

The repository CHANGELOG.md remains canonical - this page summarizes the current public story, not every commit.

Per-release entry format - add to the top of ## Latest. The full writing rules are in Writing an entry readers can use:

### X.Y.Z - short title (3-6 words)
**One or two short sentences: what the reader can do now or what got better, by how much, and an upgrade action only when one exists.**
<details>
<summary>Detail</summary>
Start with the changed user journey and retained control. Put commands, schemas,
artifacts, fixtures, parsers, and benchmarks in an "Under the hood" tail. Blank
lines around this block so MDX renders the markdown inside.
</details>

Rules: version ### X.Y.Z - title heading (never a bare bullet - it’s what makes the right-sidebar TOC a version index); bold one-liner mandatory; <details> only for verbose multi-paragraph releases (trivial patches skip it); newest at the top of ## Latest; migrate the oldest ## Latest entries down to ## Earlier releases once it grows past ~10 (migrated entries keep their heading and bold one-liner, and keep a <details> block only for major behavior changes); bump src/lib/site.ts FLOW_NEXT_VERSION + package.json in the same commit. Full runbook: agent_docs/releasing.md -> “Docs-site changelog entry”.

Release flow:

flowchart LR
  Change["Behavior change"] --> Docs["Update docs"]
  Change --> Tests["Run tests"]
  Docs --> Changelog["Update changelog"]
  Tests --> Release["Cut release"]

If Flow-Next behavior changes and the docs site does not, assume the release is incomplete until proven otherwise.