# Changelog

Source: https://flow-next.dev/releases/changelog/

Human-readable release highlights for Flow-Next.

Flow-Next changed shape with 5.0.0. One command reads whatever you have and picks the route, and the same route runs unattended with `--auto`. Arriving from 4.x? Read [the 5.0.0 entry](https://flow-next.dev/releases/changelog/#flow-next-5-0-0) first, or start with [Flow](https://flow-next.dev/skills/flow/), then come back to the items below.

## Featured release

### 7.0.0 - Roadrunner: faster on the actual work

**Flow-Next now hands you the working change in about the time plain Claude Code, or your harness, takes, often less, and large features in about half the time. The result is better than the plain agent’s even before any review, and its optional cross-model review and live QA widen the gap to up to 25% better outcomes, especially on large and long-horizon work. This is a breaking release: Ralph and the HTML render lenses are gone, and the upgrade steps are below.**

Detail

**The work is fast now, and the quality stages are where the extra time goes.** `/flow-next:flow` picks the route in seconds from what you give it. A small local fix goes straight to the change, with no spec. A spec with one task is built right in your conversation. You get the result back first; a reviewer from another model family then checks it in the background, and after any fixes the same reviewer looks only at what changed. Risky changes (persisted or shared state, concurrency, security, data layout or migrations, multi-file features) get three reviewers; small ones get one; a change to output, wording or display gets none, with the reason recorded.

**Unattended runs finish on their own.** `flow --auto` never stops to ask. It keeps fixing until the reviewer signs off (the author may decline hardening, scope creep and problems that predate the change, and when only those are left, all below Major, the loop ends with each disagreement in the pull request; a broken stated requirement is always fixed), writes every decision it made on your behalf (defaults it picked, findings it declined, reviews it skipped) into the pull request, and puts a call only you can make into the pull request as an open item instead of stopping. The pull request opens ready for review, or as a draft when an open item is left for you. In our runs with `--until=merge`, a large feature and a hard bug both went from spec to a merged pull request with nobody watching and no stops, **the large feature in about half the time plain Claude Code took**.

**How it got here.** Flow-Next started on December 26, 2025, as a plugin called flow: a plan command, a few scouts and a quality auditor. Claude Opus 4.5 was the model of the day, and models have come a long way since. It was the first to run autonomous cross-model review loops, where a model from another family argues with your agent’s work until it holds up, and one of the first to interview you before building. The goal was always to let R\&D teams work together on big, messy codebases and get better work out of their agents. Over this year that meant hundreds of features for attended and unattended runs, the building blocks for code factories, and support for six hosts: Claude Code, Codex, Factory Droid, Cursor, Grok Build and OpenCode. The output was consistently better than a plain agent’s, but all those features made Flow-Next slower and hungrier for tokens than I liked. So for 7.0 I went back through the whole plugin and rebuilt it for speed, measuring every change against plain Claude Code before keeping it. I’m going to keep improving it, in what it does and in how fast it does it.

**Benchmarks.** More than 170 full end-to-end runs against plain Claude Code on the same model, each case run several times, across a simple bug, a hard bug, a small feature and a large feature, attended and unattended, plus a held-out large feature from a repository and stack the tuning never touched. Hidden tests the agent never sees check every result, and a blind judge scores the handoff.

Large feature

**up to 1.9×**&#x66;aster than the default harness

**+25%**&#x62;etter result, blind-judged

Three cross-model reviewers

Plain Claude Code shipped a data-integrity bug in every run. Flow-Next's reviewers caught it every time, before the pull request.

Held-out large feature (Rust)

**1.2×**&#x66;aster than the default harness

**+48%**&#x62;etter result, blind-judged

Three reviewers, then the repo's full test suite

Four real bugs fixed, including a race condition. Flow-Next passed every hidden test; plain Claude Code failed one run in three.

Hard bug

**about 2×**&#x66;aster than the default harness

**+9%**&#x62;etter result, blind-judged

Three cross-model reviewers

Found the real cause and fixed it there, instead of loosening the flaky test.

Simple bug

**1.2×**&#x66;aster than the default harness

**+10%**&#x62;etter result, blind-judged

No review: a small, local fix

Better even without review: a failing test first, a fix at the cause, then a check that it works for the user.

Small feature

**1.1–1.3×**&#x66;aster than the default harness

**+8%**&#x62;etter result, blind-judged

One cross-model reviewer

The reviewer caught a setup check the new option broke. Fixed before handoff.

Against the default harness (plain Claude Code, or yours) on the same model. Speed is time to the working change, before any review or QA. More than 170 full end-to-end runs, each case run several times. Hidden tests the agent never sees check every result; a blind judge scores each handoff.

**What changes when you upgrade.** Ralph is removed, and `/flow-next:flow --auto` is the one unattended mode. If you still run Ralph, pin flow-next 6.7.x. Otherwise delete `scripts/ralph/` and any `ralph-guard` hook entries from your project settings, and use [`flow --auto`](https://flow-next.dev/autonomy/pilot/) for unattended runs. The HTML render lenses are removed too: a config that still sets `artifacts.html.enabled` keeps working and prints a one-line note, and for a visual view of a spec, a plan or a diff you run [`/flow-next:visual`](https://flow-next.dev/skills/visual/) or ask the agent for an HTML page. Your old `spec.html` and `pr.html` files stay where they are, and you can delete them. The hidden pilot and interview skill stubs are gone; use `/flow-next:flow --auto` and `/flow-next:refine`. `pipeline.chainStages` is removed: flowctl ignores the key with a one-line note, and under `--tick` make-pr runs on the next tick, so delete it from `.flow/config.json`. An unattended `--until=merge` now waits 10 minutes after the last push before merging (`land.patienceMinutes`, was 30); set the key if your review bots are slower. The Copilot reviewer now runs read-only, without write and shell tools. The Codex installer no longer adds `hooks = true` to `~/.codex/config.toml` (only Ralph needed it); a `hooks` line already there stays. You don’t need to re-run setup.

**What you keep.** Specs and their state in your repository, cross-model review from the backend you configured, live QA under `pipeline.qa`, the pull request briefing with its evidence, land, and every receipt. The optional Jev judge still offers its fork hint, memory reordering and task tier; it no longer picks the route or decides whether QA runs, so a run with a key and one without take the same route. Keyed calls now also work behind an HTTPS proxy, where before they quietly fell back to no judge.

**Under the hood.** One shared set of working rules covers every stage: the smallest change the evidence justifies (including the sibling case the same cause breaks), focused tests for what changed and the full suite only when your repository or you ask, a check you name always run and waited for, and a handoff that marks each claim measured, inferred or a guess. Refine (interview) asks only what would change the build, in one question pass. `flowctl judge --preset route` is code-only and sends no request; the `qa-gate` preset is retired.

## Latest

### 7.1.1 - Clean upgrades and every memory lesson kept

**Codex users: re-run `./scripts/install-codex.sh` once. Upgrading from an older checkout then leaves exactly the skills and commands this release ships, memory-migrate keeps every lesson from very old memory files, and make-pr accepts a Bitbucket pull request link.**

Detail

**Clean upgrades.** When a release removes a skill, `git pull` can leave its folder behind if untracked files such as `__pycache__` are still in it. The Codex, OpenCode and Cursor installers treated that folder as a skill, and the Codex installer copied it over the real installed skill. A folder now counts as a skill only when it has a `SKILL.md`. The Codex installer also retires the `interview` and `pilot` prompts 7.0 removed, which 7.1’s cleanup missed. Everything it retires goes to `~/.codex/.flow-next-retired/`, nothing is deleted, and prompts you wrote yourself are left alone.

**Every memory lesson kept.** Memory files from flow-next before 0.33 put each lesson under a `## <date> manual [<type>]` header with nothing between them. [Memory migrate](https://flow-next.dev/skills/memory-migrate/) read each of those files as one entry, so it merged the lessons and made an empty entry for a file holding only its header. Each lesson is now its own entry. Thanks @TechupBusiness (#509).

**Bitbucket pull request links.** A [`FLOW_PR_CREATE_CMD`](https://flow-next.dev/skills/make-pr/#appbot-authored-prs-flow_pr_create_cmd) that prints a `…/pull-requests/<n>` link no longer reads as a failed create. Thanks @CWayman (#504).

### 7.1.0 - The PR opens when you ask

**At the keyboard, Flow hands the change back committed on a local branch and opens the pull request when you say so, and a run you hand the merge stops before merging on any call that is yours. A bare `flowctl` typed outside a skill now needs its path, and RepoPrompt and the `review.backend` setting leave in 8.0.0.**

Detail

**The pull request opens when you ask.** An attended build used to push and open a pull request at the end of the run. Now it commits on a local branch and ends with one line: say “open the PR” when you want it. Nothing is pushed until you do. Flow no longer asks “land it now?” about a pull request it just opened; a later `/flow-next:flow` on that open pull request still offers landing once. `flow --auto` is unchanged and still stops at an open pull request. See [Flow](https://flow-next.dev/skills/flow/#what-it-does).

**A run you hand the merge decides what it can.** With `--until=merge` the run holds the merge, so it makes a call itself when that call is reversible, inside the spec and backed by evidence (updating a snapshot your change legitimately altered, say), and records it in the pull request’s Decisions list. An irreversible step, a product choice the spec leaves open, or anything that makes merging unsafe stops the run with `NEEDS_HUMAN` before the merge. Land no longer marks a draft ready on an unattended merge; it stops and names the open items. You saying “land it” still merges a draft. Without a merge to hold, a reviewer’s question that blocks nothing else no longer ends an unattended run empty-handed: the run fixes everything else, opens a draft pull request and lists the question as an open item.

**Land takes more pull requests through.** A pull request opened without a spec now lands through the same gates as any other, where land used to stop with `NO_WORK`. When resolve-pr hands a review thread to a person, land stops once with `NEEDS_HUMAN` instead of retrying the same thread every 30 minutes. Repairs push again: land fixes review threads and red CI in a detached worktree at the pull request’s head and pushes to its branch by name. See [Land](https://flow-next.dev/autonomy/land/#what-one-invocation-does).

**Fewer questions before the build.** Work and plan no longer stop to ask setup questions when no reviewer is configured. Plan uses its default depth, review is off, and the handoff says so once; set a backend in setup, with `review.backend`, or in the prompt. A check you ask for by name, such as the full suite or a command, finishes before the reply instead of running on in the background. Running something to settle a question stays read-only or throwaway on every route, and anything that needs live state, credentials, the network or a destructive command becomes a question when you are there and a human call when you are not. Audit’s autofix commits to a local `docs/audit-memory-<date>` branch and names it, where it used to open a pull request on its own.

**Less for the agent to read.** Instructions that only matter in some runs (tracker steps, unattended rules, stacked pull requests, the no-plan route, feature-map writes and more) now load only when the run needs them, and each skill reads the working rules once per run. A typical attended run reads about a quarter less instruction text than in 7.0, roughly 19,000 words down to 14,600.

**What changes when you upgrade.** Attended runs stop opening pull requests unless you ask. The plugin’s top-level `bin/flowctl` is gone, so a bare `flowctl` typed outside a skill needs the path to the plugin install’s `scripts/flowctl` ([where it lives](https://flow-next.dev/install/#optional-cli-access)); skills resolve it themselves, and nothing needs re-running. The same change lets claude.ai, Cowork and organization sync install Flow-Next, since they refuse a plugin with a top-level `bin/` directory. Thanks @sn-furali (#506). Codex, OpenCode and Cursor-script installs pick all of this up when you re-run their installer, as usual, and the Codex installer now moves skills and prompts a release removed into `~/.codex/.flow-next-retired/`, leaving your own files alone. You don’t need to re-run setup.

**Deprecated, leaving in 8.0.0.** RepoPrompt support goes: the `rp` review backend and the [Export Context](https://flow-next.dev/skills/export-context/) skill. Projects already on `rp` keep working until then and get one notice per run. To move now, pick another backend with `flowctl config set review.backend codex` (or `host`, `claude`, `copilot`, `cursor`). The `review.backend` setting goes too. From 8.0.0 the reviewer is chosen in the model-routing block of your CLAUDE.md or AGENTS.md, the one setup proposes, the same way implementers and scouts already are, and that block can route review to another model family. Until then `review.backend` works exactly as it does today, and a reviewer named in the prompt still wins. See [Review backends](https://flow-next.dev/reference/review-backends/#deprecated-in-710).

**Fixed.** Work could check the review backend from the skill’s own folder, read “nothing configured” and skip review on a repository that had a reviewer; it now resolves the backend per task from the repository root. Reviewer instructions no longer contradict themselves: a pre-existing problem that blocks shipping counts against the change, an unaddressed requirement blocks at any severity, and completion review grades the finished build on its own scale. A one-task build whose review was skipped skips completion review too. An attended flow run commits QA’s verdict on the spec branch. A reviewer’s `NEEDS_HUMAN` stops the run on every backend. `flowctl show` keeps a spec’s tracker link (thanks @sn-furali, #484, #483). A second GitHub blocker is queued for you instead of stopping tracker sync (thanks @TechupBusiness, #502). Tracker pulls stop duplicating comments, and a spec-only pull request carrying old render files no longer counts as shipped work (thanks @sn-furali, #501). Worktree Kit works from inside a worktree (thanks @sn-furali, #469). A formatter’s blank line no longer marks your setup block as customized. Codex setup keeps agent files you edited and asks once about any that differ. Memory migration keeps the lessons it did not migrate. Prospect promotes the idea you picked. Prime asks before creating CI or devcontainer files and never runs a full suite to assess. Drive closes only apps it opened. The [repository changelog](https://github.com/gmickel/flow-next/blob/main/CHANGELOG.md) lists the smaller fixes.

**What you keep.** `flow --auto` still opens the pull request itself, ready for review or as a draft when an open item is left for you. A merge you authorize goes through land’s gates as before. Review by risk, live QA under `pipeline.qa`, receipts and the pull request briefing are unchanged.

**Under the hood.** Land’s `NO_WORK` now means only a closed, unmerged pull request. It has two new `NEEDS_HUMAN` stops: a review thread resolve-pr handed to a person, and a draft pull request under a merge that flow authorized without a person in the session.

### 6.7.0 - Codex reviews on GPT-6.1 Sol

**Codex reviews and the Codex agents run on GPT-6.1 Sol, and Claude reviews can fall back to Opus 5.5 and Sonnet 5.5. Re-run `./scripts/install-codex.sh` to update your Codex agents.**

A Codex review now starts on `gpt-6.1-sol` at high effort and steps down to `gpt-6-astra`, then `gpt-6-sol`, when your account cannot serve it. The generated Codex scouts, planning helpers and quality auditor move to `gpt-6.1-sol` at their usual effort. A Claude review steps down from Fable 5.1 through Opus 5.5, Opus 5, Sonnet 5.5 and Sonnet 5; a Claude Code older than 2.1.284 skips the 5.5 models. A model you name explicitly still wins, and Copilot and Cursor reviews are unchanged.

### 6.6.0 - Scouts on Sonnet 5.5, sibling repos gated

**Plan and prime run their scouts on Sonnet 5.5, Anthropic’s lower-cost complement to Opus 5.5 that scores close to it on coding benchmarks. Teams that keep product code in sibling repos get full gates when a sibling’s code changes, and a feature-map maintain pass ships on repos with ticket-key naming rules or a non-GitHub host. Update Claude Code to 2.1.284 or later so the scouts get Sonnet 5.5.**

Detail

Scouts are the read-only helpers that plan, prime and the other planning steps send out to read your repo, docs and memory. They now run on Sonnet 5.5 instead of Opus. Anthropic reports Sonnet 5.5 more than 30% faster than Sonnet 5 and within a few points of Opus 5.5 on its published coding and knowledge-work benchmarks (for example 55.5% against 57.8% on CursorBench 4.0), on a cheaper model, so scout-heavy steps cost less and return sooner. The quality auditor stays on Opus, because a bug it misses goes unnoticed. See [the model tiers](https://flow-next.dev/subagents/overview/).

If you keep `.flow/` in a planning repo and your code in sibling clones next to it, work now checks the siblings too. It records the base of each sibling the spec changes, as named in your project instructions or the spec, and classifies every repo’s diff. A spec that only touched specs in the planning repo but changed TypeScript in two siblings used to run lint-only gates; now the docs-only shortcut applies only when every repo’s diff is docs-only, and a sibling work cannot read runs its full gates. The feature-map update reads the sibling diffs, and `flowctl features status --repo <path>` counts sibling commits toward a feature’s age, so the map comes due when the product changes. See [work](https://flow-next.dev/skills/work/) and [code in sibling repos](https://flow-next.dev/guides/keep-feature-map-current/#code-in-sibling-repos). Thanks to @CWayman for the report.

A feature-map maintain pass can now finish on repos with naming rules or a host other than GitHub. It reads your branch and commit naming rules before any proof work and asks then for a value only you have, such as a Jira key, so a pass never proves every route and then fails at the commit hook. It opens the PR through the same `FLOW_PR_CREATE_CMD` make-pr uses, and when no create command reaches your host it stops after the push and names the branch. You can also ask it to only commit, or to leave the edits uncommitted. When the commit, push or PR create fails after the proofs, the proven edits stay on the branch or in your working tree instead of being discarded. See [shipping a maintain pass](https://flow-next.dev/skills/features/#shipping-a-maintain-pass). Thanks to @CWayman for the report.

make-pr now opens the pull request when the agent’s shell is zsh, the macOS default, where the default create command used to fail with `command not found: gh pr create`.

**What changes when you upgrade.** Update Claude Code to 2.1.284 or later: older versions resolve the scouts’ `sonnet` alias to Sonnet 5. To keep scouts on another model, name it on the `fast scout` and `thinking scout` lines of your [routing block](https://flow-next.dev/guides/model-routing/). You don’t need to re-run setup.

What you keep. A single-repo workspace gates and ages the map exactly as before, maintain’s default branch and commit names are unchanged when your repo states no rules, the worker and PR comment resolver still follow your session model, and the generated Codex agents are unchanged.

### 6.5.0 - Fewer questions before the build, the same checks after

**Refine asks only what would change what gets built, often nothing in a codebase whose patterns already answer it, and flow sends you there only when such a decision is open. Hill climbs start from the target you state, and unattended bug fixes stop less often. Scripts that call `flowctl scope` or `flowctl done --resolved-feature` need updating.**

Detail

Refine is a shorter conversation. It asks a question only when a wrong guess would build the wrong thing, nothing but you can settle it, and the call is yours. Anything the code, the docs, a quick experiment or the implementation itself can settle, the agent settles, records, or leaves to the build. When nothing like that is open, refine says the spec is clear enough to build, and that is a good result. Your answers keep the precision you gave them. “Performance matters here” stays in Decision Context as guidance and never comes back as “must load in 20 ms”; a number becomes an acceptance criterion only when you said the number.

It is also one interview now. You no longer choose between business, technical or both before anything is asked. Name an audience when you want one, with `--scope=qa`, `--scope=security`, `--biz`, `--tech`, or plain words such as “run a business interview”, and the questions focus on that audience’s decisions. The read-back names every section the session changed, so a product owner and a tech lead refining the same spec in separate sessions both see what moved. See [refine](https://flow-next.dev/skills/refine/).

Flow, capture and plan now recommend refine only when they can name an open decision that would change the build and that only you can make ([when to refine](https://flow-next.dev/skills/refine/#when-to-refine)). An established pattern, a performance detail, an edge case the build and QA will surface, or a criterion capture inferred is no longer a reason to stop and interview. A structured brief goes straight to capture, and prospect points its ideas at `/flow-next:flow` so the router picks the next step.

A hill climb starts from the target you stated. The agent picks how to measure it, proves and freezes the benchmark, and writes its setup at the top of the attempt ledger before the first attempt, where you review it in the pull request. It stops to ask only when the spec has no target ([hill climb](https://flow-next.dev/choosing-your-route/#hill-climb)).

An unattended bug fix no longer stops because other pull requests touch the same files. It notes them and carries on, and still stops when one of them already fixes the bug or someone owns a fix in progress. One reproduction that fires reliably is enough. Live-app stages read the feature map directly instead of passing a recorded match along.

**What changes when you upgrade.** Scripts that call `flowctl scope` (`resolve`, `bank`, `write-policy`) or pass `flowctl done --resolved-feature` need updating, because both are removed. A repo-root `SPEC.md` copied from an older template keeps working, with its scope comments ignored; re-copy the bundled template to pick up the cleaner version. You don’t need to re-run setup. The details are on [compatibility](https://flow-next.dev/releases/compatibility/).

What you keep. Implementation review, QA, the completion gate, receipts and the pull request evidence are unchanged. `--scope=research` still asks nothing and has the research scouts write what the spec depends on. Old specs load as they are.

### 6.4.0 - Move a number with measured attempts

**Ask flow to move one number toward a target and it runs a measured loop, keeping only the changes that beat the noise; on a fixture CLI, five attempts cut `--version` cold start from 124 ms to 8 ms. Scouts now run on Opus, which costs more per plan or prime run.**

Detail

You write the metric, the target, a minimum number of attempts and a budget in a `## Hill-climb pre-registration` section of the spec, along with the command that measures it and the tests that must stay green. Work first proves the benchmark can tell a known-worse and a known-better variant from the baseline, then freezes it. Each attempt tries one idea. It is kept, as its own commit, only when it beats the best so far by more than the measured noise, a second measurement agrees and your tests pass; anything else is reverted before the next attempt. The run stops when the target is met and the minimum attempts have run, or when the budget is spent. A missed target is reported as missed, never relaxed. The pull request shows the baseline and final values, every attempt including the reverted and inconclusive ones, and the best idea not yet tried. The loop is on [choosing your route](https://flow-next.dev/choosing-your-route/#hill-climb). The 124 ms to 8 ms run is one fixture CLI on one machine: three attempts kept, one reverted, one inconclusive.

Refine asks you fewer questions. When a question is a fact it can observe, such as how long a query takes, whether a layout fits at 320 px, whether a parser accepts an input or whether an eval separates two variants, refine runs a throwaway experiment and records the question, what ran, what it saw and the decision under `## Resolved via Experiment` in the spec. An experiment that would touch live state, credentials or the network becomes a question for you, and so does a result too noisy to decide, with the data attached. Product and preference calls still come to you. When flow settles a design fork with a prototype, it now builds the competing variants behind one labelled switcher, and when the options are still open it gathers prior art first and lets you pick a direction. See [refine](https://flow-next.dev/skills/refine/) and [flow](https://flow-next.dev/skills/flow/).

Codex reviews keep the model they started with. Since codex-cli 0.154 a resumed session ran on the model in your Codex config, so a review pinned to one model re-reviewed on another. Re-reviews, validator passes and deep passes now re-pin the model the review started on. Thanks to @TechupBusiness for the report.

**What changes when you upgrade.** The bundled scouts run on Opus instead of Haiku and Sonnet, so plan and prime, which fan out scouts, cost more per run. They pin Opus rather than your session model, so planning on a pricier model does not make every scout run on it. To trade the cost back, name a cheaper model on the `fast scout` and `thinking scout` lines of your [routing block](https://flow-next.dev/guides/model-routing/). The generated Codex agents move to the gpt-6 models. You don’t need to re-run setup.

What you keep. Specs without a pre-registration section run exactly as before, the worker and the PR comment resolver still follow your session model, and explicit routing still wins over every default.

Under the hood. Set `CODEX_MODEL_INTELLIGENT` or `CODEX_MODEL_FAST` when regenerating the Codex mirror to choose other Codex models. The codex triage judge defaults to `gpt-6-luna`; the copilot one stays on `claude-haiku-4.5`.

### 6.3.0 - Bug fixes arrive with their cause and proof

**Hand flow a bug and the pull request shows the cause, the commit that introduced it, and one reproduction failing on the base and passing on the head. Stages that drive your app now start from the feature map, which took about 40% fewer turns when the task didn’t say where its target was (one fixture app, one model, 66 runs).**

Detail

You hand flow a bug report, console output or a screenshot as before. Before it writes a fix, flow looks for work that already exists: open pull requests and branches that touch the failing path, recent commits and reverts, the bug track in your project memory, and tracker issues. When it finds a fix, it runs your reproduction against that fix, records the result and stops, and it never writes a competing one. When someone visibly owns a fix in progress, or several open pull requests touch the area, flow stops and hands the bug back to you with the list. An unattended run ends with `NEEDS_HUMAN` at any of these stops. A reverted attempt counts as a cause already ruled out.

Flow then reproduces the symptom twice, lists the candidate causes, and eliminates them with runtime evidence such as logging, instrumentation or a narrowed input. It designs the fix only after one cause survives. When the report or the history names a revision where the bug did not exist, flow bisects with the reproduction to find the commit that introduced it. A symptom that will not reproduce is reported as not reproduced, with what was tried, and no fix ships for it.

The proof is one reproduction run twice. It must fail on the base revision and pass on the head, and it runs on the live app when the bug lives in one. When the reproduction already passes on the base, it does not capture the bug, and flow says so and does not claim the fix. The pull request briefing lists the prior-fix findings, the confirmed cause, the introducing commit, and the base and head results. A step that did not run appears as not done, with its reason. The steps are on [choosing your route](https://flow-next.dev/choosing-your-route/).

Stages that drive your running app now read the [feature map](https://flow-next.dev/skills/features/) first. That covers performance baselines and post-change measurements, the live proof of a bug fix, QA, and the live checks reported in pull requests. Each stage reads the map’s index and the one matching feature file, including its notes on controls that misbehave, and later stages on the same spec reuse the matched feature. The briefing names that feature, or says `unmapped`. On one fixture app with one model (66 runs), tasks that did not say where their target was took about 40% fewer turns and a third to half less wall time, with no drop in success. When the task named the page, reading the map cost 0 to 3 extra turns.

What you keep. Bug intake reads the map only when the report doesn’t say where, as in 6.2.0. Planning, questions, refactors and other routes that never drive an app never read the map, and a repository without a map pays one existence check. You don’t need to re-run setup.

Under the hood. A script that writes QA results can pass the matched feature to `flowctl qa receipt` as `resolved_feature`, in the same shape `flowctl done --resolved-feature` takes. A stage matches the feature again when the recorded file, its sub-feature or its last-proven line has changed. Details are on the [CLI reference](https://flow-next.dev/flowctl/cli-reference/).

### 6.2.0 - The feature map keeps itself current

**Your feature map now stays current as a side effect of normal work. A bug report that doesn’t say where the problem is starts from the map, which took the hardest untitled screenshot on a fixture app from 38 turns to about 12 (48 runs, same model).**

Detail

A spec that renames a button or moves a page now updates the matching [feature file](https://flow-next.dev/skills/features/) in the same pull request. Work proves each new route with one live drive, and you review the map diff beside the code diff. When work can’t start the app, or can’t tie the change to one feature file, it leaves the map alone and files a drift note. QA, drive and bug intake file the same note when a mapped route no longer matches the live app, and a later proof of that route closes it, so the open notes are the open drift.

Flow tells you when a full pass is due. Run `/flow-next:flow` with no argument and it adds `Also recommended: /flow-next:features` when a drift note is open or a feature was last proven 50 or more code-changing commits ago. Each feature file now records the date and commit it was last proven at. Setup and prime recommend seeding a map on every run, whether or not live QA is on. The [keep the map current](https://flow-next.dev/guides/keep-feature-map-current/) guide covers the three mechanisms and how to run maintain on a schedule.

Bug intake uses the map when the report doesn’t say where. Hand flow an untitled screenshot or a report like “get this thing out of my list”, and it matches the report to one mapped feature, using what the screenshot shows as well as its text, then drives the reproduction along that feature’s route. On a fixture app with six reported defects (48 runs, same model), the hardest untitled screenshot went from 38 turns to about 12 and the other one did not change, with every defect reproduced in both arms and flat cost and wall time. A report that names the page or control goes there directly, as before, because reading the map first only added turns there. No match falls back to live discovery, and flow never edits the map during intake. The matched feature travels with the fix, so QA and the pull request briefing start from it. See [bug intake](https://flow-next.dev/skills/flow/#bug-intake).

What you keep. `/flow-next:features` stays the only command that seeds or fully maintains the map, and it runs only when you invoke it; no unattended run starts it. Existing maps load unchanged. A file without a last-proven line reads as never proven, so the first due report asks for one maintain pass. A repository without a map pays one existence check. You don’t need to re-run setup.

Under the hood. `flowctl features status [--json]` reports each file’s last-proven line, its age, the open drift notes and a seed, maintain or none recommendation. `flowctl done --resolved-feature` records the matched feature in the task’s evidence. Change the due threshold with `flowctl config set features.staleAfterCommits <n>` (default 50), documented on the [configuration reference](https://flow-next.dev/flowctl/configuration/).

### 6.1.1 - Re-run setup for the prose reminder

**Re-run `/flow-next:setup` once in each repository, so the block it writes into CLAUDE.md and AGENTS.md points agents at the prose skill by an id they can call. Plan also dispatches each scout one step sooner.**

Detail

6.1.0 made `/flow-next:*` commands typed-only, so the old block’s instruction to invoke `/flow-next:prose` pointed agents at a command they can no longer call. The refreshed block names the skill id, `flow-next:flow-next-prose`. The block’s version moves to 3, and setup offers the refresh once per repository.

Plan no longer asks the tier judge before each scout. Scout tiers are fixed: memory-scout runs as the fast scout and every other scout as a thinking scout, so the call cost a step and never changed the result. Work still asks the judge for task dispatches.

`pipeline.chainStages` is deprecated with no removal release set. The key still works under `flow --auto --tick`.

### 6.1.0 - Unattended runs stop and say why

**An unattended run now keeps the spec you named and stops on blocked work with a line saying what it needs, and a typo in `.flow/config.json` no longer resets your settings. Each worker loads about 60 KB of context per task instead of 160 KB, at the same comprehension score. Two retired commands are gone, so read the upgrade steps.**

Detail

A run you leave alone works on the spec you gave it, under the review routing you set. It resolves a tracker key through the same lookup every other command uses, keeps per-task review overrides, creates the branch from the resolved default base, and stops on a git failure. Blocked work stops the run where it used to be tried again. A review result whose verdict or finding counts disagree with the review it records is refused, and every guard stops when its evidence is missing. You come back to a stop line that names what the run needs from you.

Your settings and your teammates’ edits survive. A typo in `.flow/config.json` used to reset every setting to its default. Now each reader names the file, with line and column for a syntax error, and `config set` refuses to write over it. When someone edits a spec’s tracker issue after the last sync, a push leaves their edit alone. An attended run asks whether to reconcile both sides (the recommended answer), overwrite, or leave it. An unattended run never overwrites; it queues the decision with the other deferred work.

Agents read less. A worker’s context for one task drops from about 160 KB to about 60 KB. It carries the memory index as text, only the glossary entries its task names, and a short git status, and both versions scored 7 of 7 on the comprehension eval across three task sets. The skill listing every session loads shows each skill once where it showed it twice. Review prompts, review and QA results, prospect output, tracker bodies and memory audit changes now come from flowctl, so a skill writes only its judgment. A bad input reports every error at once and writes nothing. Setup remembers the optional questions you declined and doesn’t ask them again.

Reviews need fewer arguments. Plan review runs without a file list, impl review uses the repository’s default branch when you leave out the base, and plan and completion review results go under each checkout’s own `.flow/tmp/`, so two repositories never share one.

What changes when you upgrade. The `/flow-next:pilot` and `/flow-next:interview` commands are removed. Replace them with `/flow-next:flow --auto --tick` and `/flow-next:refine` in loop recipes, scripts and instruction files; the loop form is on [Driving a loop](https://flow-next.dev/autonomy/driving-a-loop/). You still type slash commands as before, but an agent can no longer call one, so a prompt that tells an agent to run a skill should name it by id (`flow-next:flow-next-<name>`). A script that reads `review_attempts` or `tracker` from `flowctl show <spec> --json` should call `flowctl review-rounds attempts` and `flowctl sync get-state` instead. A script that pushes tracker bodies should handle the new `tracker_diverged` conflict ([Tracker operations](https://flow-next.dev/integrations/tracker-operations/#edits-made-on-the-tracker)). `flowctl done` now refuses evidence with none of `commits`, `tests` or `prs`. No config key changed. The setup block refresh this release needed arrived in [6.1.1](https://flow-next.dev/releases/changelog/#flow-next-6-1-1).

What you keep. Typed slash commands, every config key including the `pilot.*` keys, the `flowctl pilot` verbs, and the `PILOT_VERDICT` grammar keep their names. Existing `.flow/` state, review results and config files load unchanged. Explicit inputs that skills passed before (a `--files` list, a `--base`, hand-built evidence) still work.

Under the hood. flowctl gains verbs for one-call state reads, rendered artifacts, memory audits, rolling admission and commit-range evidence, among them `glossary list --match`, `done --range` and `tracker sync --op push --overwrite-diverged`. The full list is on the [CLI reference](https://flow-next.dev/flowctl/cli-reference/#rendered-artifacts).

### 6.0.2 - Specs made from an issue stay linked

**A spec you create from a tracker issue is linked to that issue the moment it exists, so the next sync updates the issue and opens no duplicate.**

The plan, capture, work, refine and QA skills pass the issue’s id and URL for you. A script that calls `flowctl spec create --tracker-first` should add `--tracker-id` and `--tracker-url`; with only the key, the spec stores the key alone, as before. An issue already linked to another spec is refused before anything is written. Thanks to @sn-furali for the report ([#464](https://github.com/gmickel/flow-next/issues/464)).

### 6.0.1 - Specs render formatted in Jira

**Specs and comments pushed to Jira now show real headings, bold, lists, code blocks and tables. Turn on Jira’s Wiki Style Renderer for the description and comment fields.**

flowctl converts each body to Jira wiki markup on the way in and back to Markdown on the way out, so your specs, merges and comment history stay in Markdown. An issue pushed before this release is converted by its next push or reconcile, and until then a pull refuses it with a `jira_body_unconverted` conflict. Thanks to @flecamos for the report ([#465](https://github.com/gmickel/flow-next/issues/465)).

### 6.0.0 - Pull requests brief the reviewer, 60% faster

**Your reviewer gets a briefing instead of a file list: why the change exists, the review steps in order, the few files to read, and only the checks that passed. make-pr writes it 60% faster (251 seconds to 100 on one 23-file pull request, three cold runs per point). This is a major release, so read the upgrade steps if you call land from a script or a schedule.**

Detail

Your reviewer opens a body that starts with why the change exists and what changes for a user or operator. Scope gives the review steps in reading order. Each step says what to check and links up to ten files that must be read, with each file’s purpose and the requirement it serves, and everything else is counted in one line. Verification ticks only what passed; a failed or unverified check stays unticked with its note. Blast radius, tradeoffs and open items follow, and an empty section is left out. A 23-file pull request comes to about 800 to 1,000 words, where the old form ran 2,700 to 3,700 words for 25 to 39 files ([#447](https://github.com/gmickel/flow-next/issues/447); thanks to @flecamos for the report). A house style in your `AGENTS.md` or `CLAUDE.md` shortens it further ([how](https://flow-next.dev/skills/make-pr/#house-style-shortens-the-body)). A branch that closes several specs gets one body, with requirement ids qualified per spec such as `fn-250:R4`, and lands as one pull request.

make-pr gets there faster. On a fixed 23-file pull request the median run went from 251 seconds to 100, from 22 tool calls to 13, and from 21,627 output tokens to 8,174 (three cold runs per point). The agent writes only the judgment and flowctl fills in the rest. The spec closes as the last commit on the branch, so the close reaches a protected base through the merge itself.

Land handles one pull request per run. Name it with `/flow-next:land <PR>`, and land resolves the review threads, makes one focused CI fix or one flake rerun per failure, catches the branch up on the server, and squash-merges pinned to the full head SHA. It merges only when you authorized the merge in the session; without that, it repairs and stops at `AWAITING_REVIEW` with reason `merge-ready; authorization required`. It never rebases, force-pushes or retargets, keeps no state between runs, and writes nothing to the repository after the merge; the tracker update is the one step left. This release’s own pull request ([#459](https://github.com/gmickel/flow-next/pull/459), 175 files, four specs) landed that way.

A worker now waits for every command it started before it reports back. When one is still running, work waits and sends a continuation worker into the same workspace, and the early return does not count as a failed attempt.

What changes when you upgrade. This is a major release.

* Replace every bare or scheduled `/flow-next:land` with a loop over named pull requests: list them, call `/flow-next:land <PR>` on each, and put the instruction to merge them in the prompt, because a list grants no merge authority. The copy-paste blocks are on [the land loop](https://flow-next.dev/autonomy/land/#the-land-loop).

* The default merge gate is green checks, a nonblocking GitHub review decision and zero unresolved threads, and a bot comment neither satisfies nor blocks it. If you relied on `land.reviewSignal`, move the requirement to your instruction file, to branch protection, or to `land.mergeVerdictCommand`.

* Eight `land.*` keys are retired. They still load, land names them once and ignores them, and you can remove them when convenient.

* Land no longer releases, requests reviewers or honors the `FLOW_PR_MERGE_CMD` override. Run releases yourself after `LAND_VERDICT=MERGED` and request reviewers outside land.

* Drop `--no-mermaid` from any make-pr script.

* If your repository already tracks pull request aid files, untrack them once from the repository root and commit the index change. Your local files stay.

  ```sh
  flowctl init
  git rm --cached --ignore-unmatch -- '.flow/artifacts/*/pr-cognitive-aid/*.json' '.flow/artifacts/*/pr-cognitive-aid/.write.lock'
  ```

The rest of the upgrade notes are on [Compatibility](https://flow-next.dev/releases/compatibility/).

What you keep. The `LAND_VERDICT` grammar and its vocabulary; `RELEASED` stays for parsers and is never emitted. Existing config files load. Old land files under `.git/` are inert. The stored aid artifact keeps schema version 1, and the HTML lens is unchanged. `flow --auto <spec> --until=merge` still carries one item through landing.

Under the hood. New read-only verb `flowctl spec closed-in-range --base <ref>` prints the specs a branch closes. A dependency squash-merged into a non-default base now reads as landed in `flowctl spec chain`. The `clean-review` judge preset is removed with the land gate it served, which leaves [five judged sites](https://flow-next.dev/guides/jev-judge/).

### 5.6.1 - Fresh ideas route in about a second

**With a TypeSafe key set, a new idea or brief now gets its route in about 0.6 seconds on the first try, where 5.6.0 could spend minutes before falling back to the host’s own judgment.**

Specs that already exist were never affected. The state for an idea or brief now holds only its text, and flowctl assembles the repository, lifecycle and pull request facts itself, so a host can’t invent them. A state file missing several fields names all of them in one error.

### 5.6.0 - Optional judgments for six pipeline decisions

**With a TypeSafe API key in your environment, six narrow pipeline decisions each become one judgment request made by the same flowctl command on every host, about 0.6 seconds for routing. Without the key nothing changes.**

Detail

To turn it on, run `export TYPESAFE_API_KEY=<key>` in the shell that runs your coding agent. There is no SDK, model selector or config to write. Setup prints `Judge: off (TYPESAFE_API_KEY is not set).` when the key is absent, and `flowctl config set judge.enabled false` turns the judge off while keeping the key.

What you see. Flow’s route step and `flow --explain` print the route with its confidence, such as `Route: defect (jev 0.91)`, or `Route: host (jev below floor: ...)` when the judge hands the decision back. At the 0.7 floor the judge routed 85% of held-out intents on its own at 95% agreement with the labels; below the floor, the host picks from the top three candidates. Under `pipeline.qa=auto`, flowctl supplies whether the app can start and the judge answers only whether the acceptance criteria describe a UI, so a skipped QA stage names which half failed. The fork check confirms a fork exists before it classifies one, which cut invented forks from 12 to 1 in the evaluation. Memory search reranks its keyword hits in one request, and plan and work skip the memory scout when it is available. A confidently mechanical task goes to your fast-scout tier, and a confidently long-running one gets a bridge recommendation; everything else stays on the session model. Land asked the judge whether a bot review was clean until 6.0.0 retired that gate.

What you keep. Every review, QA and merge gate keeps its own contract, and no judgment predicts a verdict. An explicit `IMPLEMENTER:` line in the invocation wins over the tier answer. The floors are fixed. Every stage line names which path ran, and when the judge can’t answer (no key, a timeout, a bad response, a state over budget) the decision takes its previous path and says so. The key is read at call time and never written to config, receipts, stage lines or logs.

Under the hood. One `flowctl judge --preset <name>` command carries the request, retries, validation and decision rule for every site. The evaluation numbers, their bounds and the question text are on [Optional Jev judgments](https://flow-next.dev/guides/jev-judge/).

### 5.5.1 - Your spec scaffold decides every section

**A customized `SPEC.md` now controls every section of a captured spec, so you can drop the 15 to 20 lines of quoted prompt text capture used to add. If you wrote your `SPEC.md` by hand, copy the `auxiliary_sections` list from the bundled template first.**

Detail

Capture now writes only the sections your scaffold names, where your scaffold puts them. It used to add a `## Conversation Evidence` block at the top and a `## Requirement coverage` table at the end whatever the scaffold said, and a hand deletion came back on the next rewrite. To drop the evidence block, delete `Conversation Evidence` from the `auxiliary_sections` list in your `SPEC.md`. The [scaffold guide](https://flow-next.dev/guides/spec-scaffold/#leaving-a-section-out) has the details. A repository without a custom `SPEC.md` sees no change.

What changes when you upgrade. If your `SPEC.md` was written by hand and has no `auxiliary_sections` list, captured specs stop getting the strategy, parked-unknowns, evidence and coverage sections. Copy the list from the bundled template to keep them. A scaffold copied from the bundled template already has it.

What you keep. Capture still collects your verbatim turns during the run and checks every `[user]` tag against them before it saves, so a tag still means you said it. With the evidence block dropped, a reviewer who doubts a tag can no longer look the quote up in the spec.

Two smaller fixes. When you answer one of capture’s questions by picking an option, the evidence line quotes the option label exactly and marks it as a selection. The prose contract now says that a length budget or reading level in your `AGENTS.md` or `CLAUDE.md` reaches every artifact the agent drafts, and that the contract sets no length rule of its own. Thanks to @flecamos for the report and the source reading.

### 5.5.0 - Maintainability questions at plan time

**Plan review and the technical refine pass now ask two advisory questions before code exists: does the plan repeat an edit or a decision in several places, and does it bend the intended dependency direction. A second run as the same person on one clone now refuses a task the first run holds.**

Detail

Plan review adds a Maintainability criterion, answered from the plan as written. Does the plan make the same edit or decision in more than one place? Does it add a back-edge against the intended dependency direction, or new branching to a function that is already the hottest in its module? Each answer is a concrete finding or `none identified`, shown as an advisory block in the verdict, and a named finding also lands as one line in the spec’s Decision Context. The technical refine pass asks the same two questions once, so a spec on the direct route, which skips plan review, still answers them before build. A finding pushes the verdict to NEEDS\_WORK only when it names concrete duplication, a specific back-edge, or a specific function absorbing the branching. “Could be cleaner” and suggested abstractions are out of scope.

The public claim now says what the pipeline does not prove. The README, the docs home and the [verification spine](https://flow-next.dev/understand/how-work-gets-proven/) carry the same sentence: the pipeline proves the change does what was asked and records what it did; it does not prove the codebase stays maintainable. The findings are advisory, never a gate, score or threshold.

Two runs as the same person no longer share a task. A second terminal, a scheduled `flow --auto` tick or a second machine on a shared checkout used to read its own `flowctl start` as a crash resume and send a second worker onto a task in progress. Now `flowctl start` refuses with an error naming the task, the claimant and the two ways forward: confirm the earlier run ended and re-run with `--reclaim`, or leave the task to it. Work and `flow --auto` pass `--reclaim` only after they have evidence that the earlier run ended. Claims by other people and `--force` behave as before, and one conductor per clone stays the rule. Two rolling runs on one checkout also stop overwriting each other’s notes.

### 5.4.0 - An implementer on another CLI owns the task

**When your routing block sends implementation to a model behind another CLI, such as `codex exec`, that model now owns the task. It commits checkpoints and runs its own subagents, and the worker reviews its commits before the usual review and gates.**

Detail

Until now the worker ignored an implementer tier reached through another CLI and implemented on the session model. A bridge prompted by hand told the child it could not commit or spawn agents, so a long task became a chain of returns and re-briefs; one reported task took 19 dispatches. The next work run under such a routing block takes the new path. With no implementer tier, or one your harness reaches in-host, nothing changes.

The worker hands the child a short prompt with the task’s identities and rules plus the long-task brief from `flowctl usage`, and runs it as one foreground call. The child writes code, commits checkpoints on the branch it was given, and decides its own parallelization the way an in-host worker does. It never pushes, rewrites history, changes scope, issues a review verdict or starts a nested bridge. On return the worker commits anything left uncommitted, reviews the child’s commit range against every acceptance criterion, runs the focused gates, and continues into review and `flowctl done` as before. Your done summary says which model implemented and how many subagents it dispatched. To override the model for one run, add `IMPLEMENTER: <model> at <effort>` to the work invocation.

On Codex, `workspace-write` keeps `.git/` read-only, so a child that commits checkpoints runs with `danger-full-access` inside the repository root, or the host commits between one-run-per-scope dispatches. `flowctl usage` describes both.

What you keep. The worker owns the range review, the gates, review dispatch and the done record; the conductor never bridges; land still needs its own consent.

Under the hood. The docs no longer tell you to keep a Codex child flat. The upstream decode bug behind that advice is still open for codex-cli 0.144 to 0.145, but its own minimal repro ran clean three of three on 0.153.4, and a month of spawning runs on the maintainer’s machine showed zero decode errors. Thanks to @DanielKillenberger for [#431](https://github.com/gmickel/flow-next/issues/431) and the diagnosis in [#437](https://github.com/gmickel/flow-next/issues/437).

### 5.3.0 - Dependent specs build on their parent’s branch

**A spec that depends on another now starts on the parent’s branch as soon as the parent is built, instead of waiting for its merge. Its pull request shows only its own layer, joins a GitHub [stacked pull request](https://docs.github.com/en/pull-requests/get-started/about-stacked-prs) where available, and land merges the chain from the bottom.**

Detail

Take specs B and C that depend on A. The build loop used to park B until A’s pull request merged, then C until B’s did, so each layer cost a review, a merge and an idle wait. Now B becomes selectable once A’s tasks are done and its branch is on origin. Work forks B’s branch from A’s tip, and make-pr opens B’s pull request against A’s branch. The dependency edges your plans already record are the only input, and a spec with no dependencies behaves as before.

Your reviewer gets a chain of small pull requests. On GitHub, make-pr links them into a stack (in public preview since 2026-07-30, [announcement](https://github.blog/changelog/2026-07-30-stacked-pull-requests-are-now-in-public-preview/)): the merge box shows the stack map, each layer’s diff holds only its own change, and you review and merge from the bottom. Off GitHub, or where the preview is unavailable, the same chain works as plain dependent pull requests. `flow --auto` picks chained specs up in ready and backlog mode. A spec whose parent is still in progress, or a second child of the same parent, parks with a stated reason and no strike. Chains are linear.

Land merges only the bottom open layer, so a stack never collapses from the top. On a GitHub stack it uses the [asynchronous stack merge](https://docs.github.com/en/rest/pulls/pulls#merge-a-pull-request-asynchronously) with a head pin the server enforces, and a stale pin is refused before anything merges. A merged branch stays until no open pull request targets it. Merge judgment stays yours, and nothing in this release merges on its own. This release’s own two pull requests, #432 and #433, formed the first live chain and merged from the GitHub stack UI.

Since 6.0.0, land uses GitHub stacks when they are available and leaves a conflicted layer of a plain chain for you to rebase by hand. See [Land](https://flow-next.dev/autonomy/land/), [Make PR](https://flow-next.dev/skills/make-pr/), and [Work](https://flow-next.dev/skills/work/).

## Earlier releases

### 5.2.2 - Five reported defects, no new knobs

**If you keep a glossary, land PRs from a zsh or worktree setup, or project specs to a GitHub or GitLab tracker, five things that silently went wrong now behave; nothing new to configure.**

Detail

Nothing to do first. Update the plugin and the fixes apply.

A glossary entry that carried both an *Avoid* line and a *Relates to* line used to lose part of its relationships on every `glossary add`, even when the command touched a different term, and the list command kept reporting the right count so nothing warned you. The parser now removes those two lines in the right order and the file round-trips unchanged.

Land’s merge step failed with “command not found” when the agent’s shell was zsh, which is the macOS default. The command is now built as an array and works under bash and zsh alike; the `FLOW_PR_MERGE_CMD` override keeps its contract. In the same skill, when the PR branch is already checked out in another worktree (the Worktree Kit shape), the ci-fix step now tells the agent to run the fix in that worktree instead of failing on the checkout. No worktree is created or removed for it.

On GitHub and GitLab trackers, every planned spec recorded a “status conflict” on its first push because a spec with all tasks still to do was compared against the `status:backlog` label from capture. Those are the same “not started” bucket and now agree. And a merged PR whose whole diff was the spec’s own text (the one-PR-per-gate convention) no longer counts as shipped work, so a still-open spec stops jumping to In Review on its first claim; each merged PR is checked for files outside `.flow/specs/` and `.flow/tasks/` before it counts as evidence.

Under the hood: each fix ships with its own regression test, including one that runs the real merge fence under both shells; the status-sync reference now says that `perTracker.statusMap` is read only by the Jira and Linear providers. Thanks to @flecamos, @TechupBusiness and @sn-furali for the reports.

### 5.2.1 - Change implementers, keep the route

**Choosing another implementation model or harness no longer forces a task breakdown; a ready cohesive spec keeps the direct route.**

Plan when you ask for it, separate human owners will implement, or delivery spans multiple PRs. Existing plans, research, refinement, review, and QA retain their own rules. The orchestration block still controls model and harness choices. This builds on the [5.0.0 conductor](https://flow-next.dev/releases/changelog/#flow-next-5-0-0) and the [5.2.0 merge destination](https://flow-next.dev/releases/changelog/#520---carry-a-spec-through-merge).

### 5.2.0 - Carry a spec through merge

**Carry one approved spec through review and a gated merge in the same flow run, while keeping the choice to stop at the PR. It builds on the [5.0.0 conductor](https://flow-next.dev/releases/changelog/#flow-next-5-0-0) below.**

Detail

Add `--until=merge` to `/flow-next:flow <spec>` or `/flow-next:flow --auto <spec>` when you want the selected item landed. Returning to an existing PR with plain attended flow offers that next step and asks once unless you have already authorized it. Without the destination, unattended flow keeps its pre-merge stop.

Flow calls land for the selected PR. Land keeps its CI, review, dependency and branch-protection gates; waits do not spend pilot strikes. `--tick` performs at most one landing tick. Consent stays with that item during retries and waits, and a fresh session needs current authorization again. A completed merge and an unfinished post-merge tail are reported separately, so recovery never repeats the merge. Releases and tracker writes retain their own authorization and configuration.

Capture also saves the spec before offering the saved file in your editor. The redundant approve-and-write question is gone. Product questions, split choices and readiness consent remain; plan/refine approval and scripted autofix’s `--yes` gate remain.

Existing verdict names and standalone land remain supported. The deprecated pilot and interview aliases are retained for compatibility in this release; use `flow --auto --tick` and `refine` in new recipes. See [Flow](https://flow-next.dev/skills/flow/#continue-through-merge), [Land](https://flow-next.dev/autonomy/land/) and [Capture](https://flow-next.dev/skills/capture/#mandatory-read-back) for the full contracts.

### 5.1.1 - A patch on 5.1.0

**Plan review no longer asks for a task split. [Read the 5.1.0 entry](https://flow-next.dev/releases/changelog/#flow-next-5-1-0) for `flow --auto`, and [the 5.0.0 entry](https://flow-next.dev/releases/changelog/#flow-next-5-0-0) for the release that introduced Flow.**

A spec with zero tasks or one owner task is the default route, so plan review now reviews the spec’s content and treats task count, decomposition, and dispatch shape as the owner’s decision. The consistency and `Touches` checks run only when task specs exist. A missing approach order or test sequence is still a finding against the spec, and the owner decides whether it changes the split.

### 5.1.0 - One unattended invocation, from ready spec to draft PR

**Teams that run Flow-Next unattended get one driver instead of two. One `/flow-next:flow --auto` invocation carries a ready spec to a draft PR, hop after hop, classifying each hop from the same routing references the attended conductor reads, so the route you see in `--explain` is the route the unattended run takes. Pilot users keep working through the alias for one release. `/flow-next:pilot` is retired in 5.1.0, forwards to `flow --auto --tick` with one deprecation line, and the release after 5.1.0 removes it. It builds on the [5.0.0 conductor](https://flow-next.dev/releases/changelog/#flow-next-5-0-0) below.**

Detail

**Do these first.** Update `/loop`, `/goal`, and Ralph recipes now. The default recipe is one `/flow-next:flow --auto` invocation per item; the loop recipe is `/loop 30m /flow-next:flow --auto --tick`. An existing `/flow-next:pilot` recipe keeps working for this release. The shim rewrites `--spec <id>` to the positional id, rewrites `--dry-run` to `--explain`, maps both `--backlog` and pilot’s old `--auto` backlog switch to `--backlog`, passes `--review`, `--research`, and `--depth` through unchanged, prints `pilot is now flow --auto --tick; this alias is removed in the next release` to stderr, and behaves byte-for-byte as a pilot tick. `pipeline.chainStages` is deprecated with the alias. It is honoured under `--tick` and ignored with one stderr notice in long-horizon mode.

**What was hard before.** An unattended run meant a driver looping `/flow-next:pilot` and paying a full re-anchor (skill re-read, config snapshot, selection, classification, branch resolution) plus the loop interval at every stage boundary, with only `qa` into `make-pr` allowed to chain. Pilot classified from its own private stage table, so the route it took and the route `flow --explain` showed could differ.

**What you get now.** The driver invokes `/flow-next:flow --auto [<spec-id>]` once. The conductor classifies the hop, dispatches the stage with `mode:autonomous`, verifies from observed state, records the hop, and re-classifies until it reaches a terminal: the PR exists, deferred to land, asked, blocked, needs human, or no work. Every hop ends with the same receipts, evidence echo, and ledger write a pilot tick ended with, so a run that dies at hour six resumes from disk on the next invocation; nothing is resumed from transcript. `--tick` runs exactly one hop and stops, which is what a pilot tick was, for hosts without stable long sessions. Both shapes end with the one `PILOT_VERDICT` line drivers already parse; a long-horizon run names every dispatched stage joined by `+` (`stage=work+qa+make-pr`) and carries the last hop’s verdict. Every rail pilot had moves across unchanged: the strikes ledger, the dirty-tree refusal, the all-done PR probe, the never-merge boundary, the decision log. The operator still reads one verdict line and still owns the merge. [Flow —auto, the build loop](https://flow-next.dev/autonomy/pilot/) is the page; [Driving a loop](https://flow-next.dev/autonomy/driving-a-loop/) has the recipes per host.

**Unattended classification reads the routing reference.** `--auto` reads `route-matrix.md` for the spec-state rows, `plan-vs-no-plan.md` for a ready spec with no tasks and no recorded route (it records the route with `spec set-no-plan` or `spec clear-no-plan` before any mint and echoes the deciding signal), and `gate-selection.md` for the review, QA, and completion-review gates. `--explain` under `--auto` prints the selected spec, the classified stage with its routing row and gate, the consulted fields, the PR probe, and would-clear ledger entries, with no write, no checkout, and no dispatch.

**`pipeline.qa=auto` now takes effect unattended.** A skipped QA hop records `stage: qa - skipped(config: pipeline.qa=auto: <reason>)` and advances to make-pr. `on` and `off` are byte-for-byte unchanged.

**The refusal is inverted for `--auto`.** Attended `flow` still refuses under every autonomy marker; its line now reads `NEEDS_HUMAN: /flow-next:flow is attended - run /flow-next:flow --auto for unattended runs`. `--auto` refuses only under Ralph (`FLOW_RALPH`, `REVIEW_RECEIPT_PATH`), because it sets `FLOW_AUTONOMOUS` and `mode:autonomous` for the stages it dispatches. Flow, `flow --auto`, and Ralph are three drivers and are never nested; `--auto` never dispatches land or a second driver. Land is unchanged and `DEFERRED_TO_LAND` keeps its meaning.

**Backlog mode is `flow --auto --backlog`.** `pilot.autonomy=backlog` still enables it and every safety invariant stays. In long-horizon mode a backlog run drives one selected item to its terminal and stops; the next invocation selects the next item.

**What did not change.** Config keys (`pilot.autonomy`, `pilot.gateClasses`, `pipeline.qa`, `pipeline.chainStages`), flowctl verbs (`flowctl pilot strikes list|clear`, `flowctl pilot-log`), the ledger path, `.flow/pilot-runs/`, and the `PILOT_VERDICT` name are not renamed; a rename is a separate, deliberate break for a later major. Published counts drop by one skill and one command.

**Pending.** Terminal-parity and wall-clock results for long-horizon versus tick execution are not yet measured. Full detail in the [repository changelog](https://github.com/gmickel/flow-next/blob/main/CHANGELOG.md).

### 5.0.1 - A patch on 5.0.0

**A bare `/flow-next:flow` now proceeds to the next best step on its own, and gate receipts no longer fail because of a stray `.git` above your working tree. [Read the 5.0.0 entry](https://flow-next.dev/releases/changelog/#flow-next-5-0-0) for the release itself.**

With no argument, flow resolves the item this conversation last touched, then the spec matching the current branch, then asks to capture intent no spec holds yet, then picks the next open spec by judgement with an inline pick on ties, and only then asks what to work on. The gate receipt probe stops at any `GIT_CEILING_DIRECTORIES` entry, the same boundary git uses, so an empty `/tmp/.git` no longer turns an outside-a-repo check into a tooling error. The README and plugin docs now describe the shipped 5.0.0 routing contracts.

### 5.0.0 - Hand it anything, it picks the route

**You no longer choose the next command. Hand Flow-Next whatever you have, from nothing to a pasted bug report to a spec with an open PR, and it chooses the smallest sufficient route, runs it, and stops at the next decision that is yours. The recommendation you read and the route that runs are now the same rule.**

Detail

**Do these first.** Two command names changed, so update scripts, `/loop` and `/goal` prompts, `CLAUDE.md` or `AGENTS.md` policy paragraphs, and team docs:

* `/flow-next:guide` is removed with no alias. Use `/flow-next:flow --explain <the same words>` for the recommendation, or plain `/flow-next:flow` to run the route.
* `/flow-next:interview` is now `/flow-next:refine`, same scopes and flags. The old name forwards with one deprecation line for this release only; the release after removes it.
* Live QA can now decide for itself: `flowctl config set pipeline.qa auto` runs the live pass only where the spec describes UI behaviour on a surface the QA skill can reach. Existing `off` and `on` values keep their meaning. Pilot and land invocations are unchanged, so an unattended `/loop 30m /flow-next:pilot` recipe needs no edit.

**What was hard before.** Every stage skill printed its own next-step advice, the guide skill kept a third copy of the same matrix, and the three drifted. You read a recommendation, then typed the command yourself, and a ready spec still met the plan-or-not question on every route.

**What you get now.** [`/flow-next:flow`](https://flow-next.dev/skills/flow/) takes anything: an idea, a spec or task id, a tracker issue, a branch, a pasted console dump, a how or why question, something slow, a cleanup that must keep behaviour, or a design fork. It matches the starting state against one shared routing reference, runs the routed skill by name, re-evaluates from observed state after each hop, and stops with a report that names every stage as `ran`, `skipped(reason)`, or `failed(reason)`. A run from intent ends when the PR exists; a run on an open PR converges it and stops when merge is the only step left. It never merges, never closes a spec, and never invents a review or QA verdict. Picks a stage produced (a ranked candidate, a split proposal) are asked inline and the run continues; only a decision that ends the run stops it. `--explain` prints the route and leaves `.flow/` byte-identical. On hosts that match skills by description, “this endpoint is slow” reaches flow with no slash command at all. [Choosing your route](https://flow-next.dev/choosing-your-route/) is now the `flow --explain` walkthrough.

**Direct execution is the default.** A ready spec with no tasks routes to `work --no-plan`. Plan is chosen only on a positive signal: you asked for one, separate human owners will implement, delivery is staged across several PRs, or the implementer is routed out of the session model. Risk, size, and file count never trigger a plan on their own. Capture’s and plan’s closers print the same rule’s result, and under flow, capture records the route itself.

**The routing reference grew.** Beside the six worked routes the docs already carried, the matrix names refactoring (pin the contract with a characterization test first), performance (baseline on a real surface before any change), hill climb (a frozen harness and a target that is never relaxed), investigation (a cited read-only answer, with a new `why-scout` that tiers each finding as direct, supported, inferred, or unknown), and prototype (an observable fork is settled by running something, never by asking you to guess). Every row names its positive signal, its safe skip, and the skip kind.

**Refine gains a read-first pass.** [`/flow-next:refine <spec> --scope=research`](https://flow-next.dev/skills/refine/#the-research-scope) asks nothing: four read-only scouts resolve the libraries and APIs the spec names into one `## Resolved via Research` section with a source on every line. It skips itself, visibly, when the section exists or a plan already ran the scouts, and flow routes to it only when a spec names a library the repo does not already use.

**Less to read at the read-back.** Capture, refine, and plan write the draft once and show you a summary, one ask (approve and write, open in editor, abort), and only the diff on each edit cycle. Before, three full copies of the spec printed per edit. Ratification before every write is unchanged.

**What did not change.** Pilot is byte-for-byte the same and activates its QA stage only on the literal `on`; `auto` is flow’s call, because judging whether a surface is drivable is attended work. Flow never runs under pilot or Ralph. The unattended conductor and harness or model autorouting are the road ahead, on [Flow and the road ahead](https://flow-next.dev/understand/flow-and-the-road-ahead/), with no dates promised.

**Evidence.** A pre-registered non-inferiority study (90 draws, one frontier model, 27 fixture situations) scored the retired guide matrix 24/27 and the shared reference 27/27 on the nine discriminating items; the reference read its required file on every draw. Superiority was never the claim. Both implementation tasks reached SHIP under cross-family review, on rounds 2 and 3.

**Under the hood.** Six reference files under the flow skill (`route-matrix`, `spec-count`, `plan-vs-no-plan`, `gate-selection`, `prototype-before-ask`, `tail`), each opening with a decision record; capture’s closer, plan’s menu, work’s zero-task ask, and `flow --explain` read them by pointer, and a test fails on any pointer that names a missing file. `pipeline.qa` is a string enum `off | on | auto`; any other value is `off`. Published counts stay at 28 commands and 32 stable skills. Skill prose no longer carries spec-provenance markers. Full detail in the [repository changelog](https://github.com/gmickel/flow-next/blob/main/CHANGELOG.md).

### 4.18.0 - Better recommendations for your next step

**Flow-Next now recommends working directly from a ready, cohesive spec when a separate task plan would add little value.**

Detail

After capture or an interview, Flow-Next now helps you choose whether to refine the spec, review its design, split it into tasks, or start implementation. Its guide follows the same approach. When the spec is ready and a capable agent can own the work, the recommendation is `work --no-plan`, with a reason for that choice.

Previously, the guidance leaned too heavily toward creating a task plan first. Planning is now recommended when dependencies, separate owners, staged delivery, or execution constraints make a breakdown useful. A large change or a risky design alone does not make task planning mandatory: unresolved decisions call for clarification, and design risk can call for a separate review.

You still choose the route. The recommendation does not silently skip your requested reviews or change your configuration. Direct work carries the whole spec through the same configured implementation review, acceptance coverage, completion policy, and evidence requirements. Optional live QA remains available alongside staging and manual QA.

Internal benchmarking found that direct execution can produce higher-scoring implementations with capable frontier models. That informs the recommendation; it is not a guarantee for every spec or agent.

### 4.17.0 - Reviews through your managed host

**Developers working in a compatible managed host can keep reviews on the host’s selected provider account while retaining Flow-Next’s findings, fix loop and review receipts.**

Detail

Update Flow-Next before enabling required managed reviews in your host. Standalone users need no new configuration and keep their existing CLI review path.

A managed session previously needed a separately launched reviewer CLI, which could pick up a different account from the one selected in the host. A compatible host can now supply the review execution path for that session. You choose the review backend as before, inspect the returned findings and receipt, and retain the same review and merge gates. Flow-Next still owns the review prompt, round accounting and verdict.

If the configured managed provider is unavailable, rejects the session scope or returns an invalid response, the review stops. It never silently switches to a local CLI. Provider and account availability remain the host’s responsibility; this release does not add cloud credential management or make every backend available in every host.

**Under the hood.** The installed launcher advertises its managed-execution capability. Hosts pass a local endpoint and session token; users do not configure those variables themselves. Completion reviews can require the managed path before reserving a review round. See [managed review execution](https://flow-next.dev/reference/review-backends/#managed-review-execution).

### 4.16.1 - Safer Codex reinstalls

**Reinstalling flow-next into Codex could leave a duplicate agent registration, lose a setting that sat after a commented table header, or fail outright on a Windows machine whose locale was not UTF-8. All three are fixed, and a forced task takeover with a custom note now hands the task over instead of leaving it with the previous owner.**

Nothing to do on upgrade beyond re-running the Codex installer if you use it. The same release removes 141 lines of unused Python found by a five-reviewer pass over the CLI; prompts and public contracts are unchanged, and the pass also surfaced the tracker fixes listed in the [repository changelog](https://github.com/gmickel/flow-next/blob/main/CHANGELOG.md).

### 4.16.0 - A Claude review from any harness

**Teams conducting Flow-Next from Codex, Cursor, Grok Build, Droid, or OpenCode could not get a Claude-family review through the packaged review path - the closest thing was a hand-typed `claude -p` prompt with no receipt, no model ladder, and no fix loop. `claude` is now a review backend like `codex`, `copilot`, and `cursor`: set it once and every plan, implementation, and completion review carries the same receipt, round counter, and fix loop.**

Detail

Nothing to do on upgrade. To use it, run `/flow-next:setup` and pick Claude Code CLI from the review menu (it appears when the `claude` CLI is on your PATH), or set it directly with `flowctl config set review.backend claude` - the spec form `claude:<model>:<effort>` pins a model and one of the CLI’s effort levels (`low`, `medium`, `high`, `xhigh`, `max`), and a per-task pin or `--review=claude` on any review command overrides for one run.

The reviewer is read-only by construction rather than by policy. The prompt arrives on standard input, the child process gets exactly three tools (Read, Grep, Glob) with every MCP server excluded, and there is no shell at all - because a pre-approved `git diff` would still accept flags that write files. The diff under review is handed over as a file path under `.flow/tmp/claude-review/`, one file per reviewed range, so a re-review after a fix commit reads the new range while the earlier evidence stays byte-identical. Deep passes and validator passes resume the primary review’s session by id instead of starting cold, which is what the shared deep-pass prompt has always assumed. When the pinned model is not available to your account, the resolution ladder steps down the ranking on the CLI’s real signature (it exits 0 with an error payload, not a non-zero code) at most twice and records what actually ran in the receipt.

Same-family reviews are allowed and receipted, not refused. On a Claude Code host a `claude` review is a fresh-process second opinion from the same family; the receipt names the backend and model, and the skill prose says so. Independence is judged on the model family that wrote the diff, not on the host name: Cursor, Droid, and OpenCode can run Claude writers too, so the docs condition every “cross-family” claim on the writer’s model. The `host` backend keeps its fail-closed cross-family rule; the first-round three-draw fan-out stays `codex`-only; the bridge recipes remain the way to have Claude write code from another host.

Ralph users get the same rows: the harness menu, `config.env`, the prompt templates, and the guard that blocks `--force` and other human-only recovery arguments all recognise `flowctl claude` review commands exactly as they do the cursor spelling.

**Under the hood.** The backend is a registry entry with the fn-76 ranking invariant, a stdin runner, an explicit resolver that rejects a foreign `--spec` and coerces foreign configured defaults (Claude ids do not cross over), and the five subcommands routed through the shared review driver. Plan review ran to SHIP over six rounds on two model generations; the implementation had four per-task cross-model reviews and a completion review on GPT-6 Astra. Reference: [review backends](https://flow-next.dev/reference/review-backends/), [CLI reference](https://flow-next.dev/flowctl/cli-reference/), [model routing](https://flow-next.dev/guides/model-routing/).

### 4.15.0 - One spec shape everywhere

**A spec created from the command line used to arrive with different headings than a captured one, so plan review, R-ID coverage, completion review, and your PR body all read it half-empty. Every spec now starts from the same template - the one your repo controls.**

Detail

Nothing to do on upgrade, and your existing spec files are never rewritten.

Flow-Next has two ways into a spec. [`/flow-next:capture`](https://flow-next.dev/skills/capture/) and [`/flow-next:interview`](https://flow-next.dev/skills/interview/) wrote the canonical template - Goal & Context, Architecture & Data Models, API Contracts, Edge Cases & Constraints, Acceptance Criteria, Boundaries, Decision Context - while `flowctl spec create` wrote a six-heading skeleton of its own. Every reader downstream keys on the template’s headings, so a CLI-born spec exported an empty goal and empty boundaries, and [`/flow-next:make-pr`](https://flow-next.dev/skills/make-pr/) fell back to the spec id as the pull-request title. Three specs in a private repo shipped that way before the pattern was noticed.

`flowctl spec create` and `flowctl spec skeleton` now render the canonical template, resolved through the same cascade the skills already used: a repo-root `SPEC.md`, then `spec.md`, then the bundled template, first match wins. Point a repo at its own `SPEC.md` and command-line specs pick up your house sections too - no config key, no second scaffold to keep in step. The old skeleton is gone.

You keep every spec you already wrote. Specs with the older headings still export through read-only synonyms (`Overview` or `Context`, `Boundaries / non-goals`, `Decision context` in any case), and `flowctl validate` prints one `legacy spec headings` warning per such spec as the nudge to migrate - a warning, never a block. Two export bugs went with it: template guidance written in HTML comments no longer exports as acceptance criteria, and fenced code inside a section - a diagram, a snippet - no longer vanishes from the exported body. A sweep of 217 existing specs against the old parser found zero regressions and one accidental correction, a spec that had been exporting example text from inside a fence as its boundaries.

One fix rides along for anyone who renames specs. `flowctl spec set-title` renamed a spec’s id and its files but left `branch_name` at the old slug, so a retitled spec dropped out of [land](https://flow-next.dev/autonomy/land/)’s pull-request discovery and autonomous runs named the wrong branch. A `branch_name` still equal to the old spec id now follows the rename; a value you set yourself is kept, and the JSON result reports both the name and whether it was re-derived.

**Under the hood.** The byte-for-byte skeleton baseline moved onto the template itself and is hash-pinned, so a scaffold edit is always a deliberate bump. Deliberately not done: `flowctl prospect promote` and the plan skill’s own scaffold step still write the legacy headings, and are a follow-up. The implementation was written over a headless bridge to a different model family, then reviewed five rounds in-host - rounds two, three, and four each caught a regression the previous fix pass had introduced (phantom decision bullets from a trailing template comment, fenced code erased from exported bodies, fence-first sections exporting empty), and round five shipped only after the 217-spec sweep came back clean. Reference: [CLI reference](https://flow-next.dev/flowctl/cli-reference/), [customizing the spec scaffold](https://flow-next.dev/guides/spec-scaffold/).

### 4.14.0 - Faster builds by default

**Multi-task specs used to wait at every wave barrier until the slowest worker finished. `/flow-next:work` now schedules on the rolling frontier by default - the next task starts the moment any worker returns - and the wave loop survives only as a structural fallback. Nothing to enable, nothing to configure.**

Detail

If you ran the experimental `/flow-next:work-rolling` command, switch to plain `/flow-next:work`: the beta command is gone, with no alias, because the scheduler it carried is now the default. Every other invocation is unchanged.

A wave-scheduled spec dispatched a safe group of tasks, then held the next group until the whole wave had joined, so one slow task idled every finished sibling. The rolling frontier admits a new ready task at every worker return, each worker in its own isolated worktree, and the conductor integrates, reviews, and completes each task as it lands. The pre-registered eval behind it measured a 52.1% work-phase wall-clock saving at quality parity, and the beta ran end-to-end on Claude Code, Cursor, and Grok Build before it graduated. The faster shared-checkout variant stays rejected: its speed came partly from thinner tests, and that is not a trade this project makes.

You keep the wave loop where rolling has nothing to schedule, chosen from the shape of the run rather than a knob: a task-id run (only that task runs), plan-sync switched on (its per-wave barrier is a fail-closed rule), a spec with fewer than two open tasks, or a fully sequential dependency chain. The route prints once at the start of the task loop - `Scheduling: rolling` or `Scheduling: wave (<reason>)` - so a run report always says which scheduler it used. A host whose subagent dispatch is measured to block still reports `degraded to wave` and keeps the rest of the rolling lifecycle. Pilot, land, and Ralph dispatch plain work and inherit the change; every gate, receipt, and review surface behaves exactly as before.

**Under the hood.** The scheduler reference moved to `skills/flow-next-work/references/rolling-scheduler.md` and is read only on the rolling route, so the always-loaded work surface grew by the route decision and one pointer. Reference: [work](https://flow-next.dev/skills/work/#scheduling-rolling-frontier-by-default), [choosing your route](https://flow-next.dev/choosing-your-route/).

### 4.13.1 - Two idle intervals, opt-in

**If you run pilot and land unattended, two of the intervals the loop spends waiting on nothing can now be switched off: pilot can open the draft PR in the same tick as a live QA verdict, and land can measure its merge wait from the bot’s clean review instead of from the last push. Both are opt-in keys, off by default, and neither changes a gate.**

Detail

Nothing changes until you set a key. `flowctl config set pipeline.chainStages on` and `flowctl config set land.patienceMinutesAfterReview 15` are the two switches; leave either unset and the loops behave exactly as before.

Pilot advances one stage per tick, so a spec whose QA pass just finished used to wait a full driver interval, plus a re-anchor, before make-pr ran, even though make-pr was always the next stage. With `pipeline.chainStages` on, a tick whose live QA stage produced a fresh terminal verdict runs make-pr before it exits, with its own evidence block and stage line, and the verdict reads `stage=qa+make-pr`. The chain table is closed to that one row. The research finding also named `plan → plan-review`, and the plan review dissolved it: pilot’s plan dispatch already runs its own review loop to SHIP, so there was no idle interval there to remove. Nothing chains into `work`, and the switch is inert unless the QA stage is on.

Land’s default `silence` signal waits a patience window measured from the last push, restarted by every fix push, even once the bot has reviewed the current head clean and no threads are open. With `land.patienceMinutesAfterReview` set, that wait is measured from the review event instead, and only under `silence`, only while the review is head-current with zero unresolved threads. A fix push moves the head, the review stops being head-current, and the push window applies again until the bot re-reviews. The window is your human-objection grace period, which is why this stays opt-in and why the report line names which anchor bound: `anchor=push` or `anchor=review`. The `approve` and login signals, the merge command, and the ledger are untouched.

One fix rides along for every pilot user. The make-pr verification probe piped `gh` through `jq | head`, so a GitHub outage read as “no PR” and recorded a strike; two of those unready the spec. The probe now captures the `gh` exit status, and a probe failure is the crash-class `NEEDS_HUMAN` pilot already uses elsewhere, never a strike.

**Under the hood.** Both keys are seeded defaults in the published config schema (`pipeline.chainStages` as a strict `off | on` enum, `land.patienceMinutesAfterReview` as integer-or-null with `minimum: 0`; unset, `null`, and `0` are off). The skill bash is pinned by contract tests that execute the fences under a POSIX bash. Reference: [configuration](https://flow-next.dev/flowctl/configuration/), [pilot](https://flow-next.dev/autonomy/pilot/), [land](https://flow-next.dev/autonomy/land/).

### 4.13.0 - Fix how a lesson is found

**When your agent keeps re-learning a lesson from a memory entry that is already correct, the problem is retrieval. `/flow-next:audit` now repairs that entry’s title, tags, module, and filing, so the next search surfaces it.**

Detail

Nothing to switch on. The next `/flow-next:audit` run classifies this way on its own.

A recurring lesson already had a graduation path: Harden turns it into a lint rule, a CI step, or an instruction-file rule, so it fires on its own instead of riding the context window. That path needs a rule a machine can check. A lesson stated as judgment - a convention, a naming call, an ordering constraint - has no such rule, so a correct entry that kept coming back landed on Keep, and the next run re-learned the same thing while the entry sat unread in the store.

The audit now reads that pattern as a findability problem and gives it the third answer. When an entry recurs, resists mechanization, and carries a nameable defect in how it is found, the audit classifies it as an Update with a retrieval fix: it repairs the entry’s `title`, `tags`, `module`, and `applies_when`, and moves a misfiled entry into the category the lesson belongs to, since a category-scoped search never reaches an entry filed somewhere else. Your report counts these inside Updated, with the retrieval fixes named, so a run tells you how much of the sweep went to findability rather than content.

The repair stays in its lane. The retrieval rationale never rewrites what an entry says. An entry that also carries plain reference drift - a renamed path, a dead link, a stale snippet - still gets that repaired on its own evidence, in the same Update, and you keep the same interactive confirmation you had before.

**Under the hood.** No new status and no new field. The signal is the same write-side recurrence scan Harden already reads - `## Update` headings and entry-file commits - never a usage count, because nothing records that an entry was read during a run. The named defect is required: recurrence alone does not trigger the fix, since the recurrence counters only grow and a repair is itself a commit the next scan would count, so a recurrence-only branch would churn the same entry’s metadata once per audit forever. A placement move is a `git mv` into the new category directory plus the frontmatter `category` set to match, with every `related_to` naming the old entry id re-pointed in the same edit.

### 4.12.0 - Review findings arrive together

**Waiting through review used to mean waiting through rounds: fix three findings, re-review, get three more, repeat. The first review round now draws three reviewers at once, each reading for a different kind of defect, and hands you one merged list to fix in a single pass.**

Detail

Nothing to enable and nothing to upgrade. On the `codex` and `host` review backends the change is already the default the next time you run `/flow-next:work` or `/flow-next:impl-review`.

Before, a review round was one reviewer’s single look at your diff, and what it happened to notice set your next hour: you fixed those findings, sent the diff back, and the next round surfaced a different set the first pass had walked straight past. The findings were real, so the loop felt productive while it was mostly re-reading the same code.

The first round of a scope now runs three concurrent draws of the same reviewer, each carrying one axis lens: correctness and logic, contracts and consistency (do the docs, tests, and stated promises agree with what the code actually does), and integration with the code you did not touch. The coordinator merges the three finding sets into one, drops duplicates, ranks what is left, and runs a single fix pass over the whole thing. Most of what used to trickle out across serial rounds is in front of you in the first merged one.

Two pre-registered evals set the shape rather than an intuition. A single review pass surfaced roughly 45% of the validated finding pool, which makes one reviewer a sample rather than a sweep. The union of three axis-differentiated draws reached 1.56x single-draw recall against a pre-registered 1.5x bar, at flat validity, so the extra findings are genuine defects rather than noise.

The bounds are worth knowing before you rely on it. A clean diff that would have shipped in one round pays roughly 3x review tokens for findings one draw would have surfaced anyway. And roughly a third of validated findings eluded every draw, so round 2 shrinks rather than disappears.

Which is why the topology is yours to steer, in the invocation, in a sentence. There is no flag and no config key to learn. Say `use 1 reviewer instead of 3` on a small clean diff and the round collapses to a single draw with the token cost that goes with it. Say `use three different model families for the review fan-out` before a high-stakes merge and each draw routes to a different family, so blind spots decorrelate by model as well as by axis. Reviews are optional to begin with, so all of this only ever applies to a layer you already chose to switch on. The recipes are on [Steering the fan-out](https://flow-next.dev/guides/model-routing/#steering-the-fan-out).

**Under the hood.** The merged round consumes one review round against the deterministic cap, not three, and the receipt records each draw honestly. Re-review after the fix pass stays a single dispatch carrying the full merged finding container, because verifying fixes needs continuity rather than breadth. Scope is the `codex` and `host` backends: `rp` keeps its single stateful chat, `copilot` and `cursor` keep single dispatch every round, and completion review, land, and external PR bots are untouched. On the codex backend a cross-family fan-out keeps the primary draw on codex while secondary draws may name codex, copilot, or cursor; on the host backend the per-draw pins are unconstrained. Partial failures fail open from whichever draws returned a verdict.

### 4.11.0 - Mark a spec too small to plan

**A mixed backlog no longer forces a choice between a blanket skip-planning flag and hand-running the small stuff: you mark the individual spec, and the loop builds it straight through while everything else stays planned.**

Detail

If you drive pilot over a backlog, the no-plan route used to be an invocation flag - `/flow-next:pilot --no-plan` applied to whatever spec the tick happened to select, which made it useless the moment your backlog mixed trivial fixes with real features. That flag is gone. The decision now lives on the spec itself, next to the ready flag it resembles: `flowctl spec set-no-plan fn-N`, or `--no-plan` at capture time, records the human judgment “decomposing this would convert no unknown,” and a ready spec carrying it goes through pilot straight to work’s direct route - one implicit task, no plan or plan-review stage, every review gate and receipt unchanged.

The control stays yours in both directions. No autonomous path ever sets the field; setting it is refused once a spec has tasks; `clear-no-plan` undoes it any time; and a stale marker on a spec that later got planned is inert - the planned tasks run, with a one-line notice. Scripts still passing pilot’s old flag degrade safely: an unknown-flag notice, and the affected specs simply route through planning, the safe default. Direct `/flow-next:work fn-N --no-plan` works exactly as before.

Under the hood: the field mirrors the ready flag’s lazy contract (absent reads false, idempotent toggles), surfaces on every JSON read surface including `noPlan` on `ready --all` rows, stays flow-local (never tracker-projected), and work-rolling’s refusal of the route now covers the field as well as the flag. See [the migration note](https://flow-next.dev/releases/compatibility/#411-pilots---no-plan-flag-removed).

### 4.10.2 - Specs only quote what you said

**A captured spec can no longer put words in your mouth: the strongest provenance tag now means “findable in the quoted evidence”, and capture’s own process rules can’t masquerade as things you asked for.**

Detail

When capture turns a conversation into a spec, every line it writes carries a provenance tag - your words, a restatement, or the agent’s own fill-in - so you can see at read-back what you actually asked for. A field report showed the strongest tag leaking: a close restatement could be stamped as your words, and after the duplicate check agents sometimes invented a “user turn” out of capture’s own process rules (“this is a new spec, don’t mark ready”) and wrote it into the spec’s boundaries as if you’d said it.

Both holes are closed. Your-words now means the exact content is findable in the quoted conversation evidence the spec carries - a restatement is labeled as one. Evidence lines must quote things you actually typed; process rules stay process. A correction you give during the read-back edit loop becomes evidence before the redraft, so your own fresh words are never downgraded, and when a large capture splits into several specs, each one is checked against its own evidence slice. The read-back now verifies all of this before asking you to approve.

### 4.10.1 - Unattended landing sees today’s Codex

**A converged PR no longer stalls waiting for a human just because the review bot changed how it says “all clear” - and skills stop hand-rolling their own memory dedup.**

Detail

Codex’s PR reviewer used to post a “Didn’t find any issues. Reviewed commit: …” comment when a review came back clean. It now delivers the same verdict as an edited-in-place summary table naming the reviewed commit. The land loop’s clean-review gate only knew the old phrasing, so a PR with green CI, zero open threads, and a demonstrably clean review of the current head still ended in “needs human” - the one outcome an unattended ship loop exists to avoid. Land now recognizes both forms. The safety posture is unchanged: the comment must come from an allowlisted automated reviewer and name the current head commit, a stale verdict never counts, and setting the pattern to an empty string still switches the comment path off entirely. Repos that installed before this release are covered without any action - the old default stored in `.flow/config.json` upgrades itself at read time, while customized patterns are never touched.

Two smaller fixes ride along. Skills that file recurring memory notes (like the feature map’s drift memos) now use one deterministic `flowctl memory upsert` call - it updates the existing note when the title matches exactly, creates one when nothing matches, and refuses to guess when two candidates exist - instead of each skill hand-rolling its own find-or-create logic. And in repos with a markdown formatter, `/flow-next:features` runs it over the map files before finishing, so seeding the map no longer produces follow-up “formatter artifact” commits.

### 4.10.0: Navigation knowledge that survives the run

**Live verification stops paying the same navigation tax on every pass. A new committed map at `.flow/features/` records, from the user’s point of view, how a user reaches each feature, how an agent drives it, and which traps waste a run - and `/flow-next:qa` and the driver read it when it exists. `/flow-next:features` seeds the map with every route proven by one live drive before it lands, then keeps it honest with an audit-shaped maintain pass on your own cadence. The spec still says what to prove, and captured live evidence is still the only ship basis.**

Detail

* **The map is committed, and it is navigation only.** `.flow/features/` sits beside `.flow/memory/` because its whole value is surviving the session and the machine - deliberately the reverse of `/flow-next:map`’s local `.clawpatch/` code index, which stays local-per-developer. An index carries the operating rules (baseline preconditions, driving conventions, proof standards) so a cold agent can drive from the map alone; each feature file carries a `**Surface:**` identifier and exactly four sections: Sub-features, How to get to it (user POV), Driving it, Gotchas.
* **Nothing undriven lands.** Seed interviews the repo rather than you - surface, run command, drive mechanism, observable evidence, isolation - then proves each route with one live drive before writing it. A partial seed lands the proven features and names the failures instead of discarding the run; a pure library, a host with no usable driver, or a checkout that will not build is refused with the reason rather than mapped from guesswork.
* **Maintain is audit-shaped, and its edits stay in their lane.** Index hygiene, one read-only source reader per feature, reconcile, one live pass over every feature even when source looks clean, then triage: wrong user-POV description is doc drift (fix the map), working behavior the harness cannot drive is a harness gap (fix it and re-drive before shipping), broken app behavior is a product bug (report it, keep it out of the PR). Outcomes are `clean` (no branch, no PR), `changed` (one chore PR of proven map and harness corrections, never a merge), or `blocked` with the blocker named. Product code is never edited.
* **Doctor gates every drive.** One read-only check - right build, port owned by this run, valid auth - before the first drive, on each fresh session, and after any failed drive. Never drive an instance this run did not start, never kill by process name; an orphaned port from a crashed prior run blocks with a reclaim instruction rather than being reaped.
* **Consumers cost nothing when the map is absent.** QA and drive gate on directory existence only - no config key, no registration - so a repo without `.flow/features/` behaves exactly as it did. A stale route QA finds is filed as a `feature-map-drift` memory entry for the next maintain pass; QA never edits the map mid-run.
* **Cadence belongs to you.** `/flow-next:features` is never a pilot stage, a land tail step, a Ralph iteration, or a post-merge hook: any autonomy marker in the environment refuses the run. Every invocation ends with a typed `FEATURES_VERDICT=` line, so a host loop such as `/loop 1d /flow-next:features` can read the outcome.

### 4.9.1: Reviews judge the work, not the paperwork

**Two review-rubric fixes. A workaround wearing a well-written justifying comment no longer reads as “well documented”: the reviewer now treats that comment as a flag on the underlying code, and rewriting the comment without fixing the code does not resolve the finding. And a review bot can no longer hold your merge hostage on process ceremony - decisions recorded in the spec are settled, and checklist or handoff paperwork is feedback, never a merge gate.**

Detail

* **Comment-as-alibi is a named finding class** in the implementation and standalone review prompts. Severity is judged from the workaround, not the prose; the fix is the code, or the constraint encoded as an assert, a test, or a lint rule. Licensed comments are never flagged: license headers, external-constraint notes, lint suppressions with reasons, public API contracts, issue links - the same keep-list the worker’s authoring rule has carried since 4.8.0, so author and reviewer never disagree.
* **Recorded decisions are settled.** The plan-review and completion-review prompts gain the rule the implementation review already had: a finding that re-litigates a decision recorded in the spec’s Decision Context is FYI, never blocking. All three prompts now also state that process-compliance observations - checklist ceremony, dogfood records, handoff paperwork - never gate a merge. The recommended `land.reviewTrigger` text tells external bots the same up front.

### 4.9.0: Start work without the planning stage

**A fully-known spec can now go straight to implementation as a recorded choice. `/flow-next:work` on a spec with no tasks used to fall through and could end green having built nothing; it now forks - plan first, or work directly - with a recommendation judged from that spec’s size and blast radius. Say “skip planning” or pass `--no-plan` and the ask never appears; the direct route mints one minimal implicit task and runs the same pipeline, so evidence, review, and receipts hold unchanged. The same release removes an unexplained permission ban that had silently kept dispatched workers from spawning their own subagents.**

Detail

* **You choose, and the choice is recorded.** The fork’s recommendation is judged per spec (size, independent surfaces, blast radius) with its reason stated - no static default. Contradictory signals ask instead of guessing, and the flag on an already-planned spec is ignored with a one-line notice.
* **The fast route keeps every contract.** The minted task is deliberately minimal (“implement this spec”, satisfying every R-ID, its review contract pointed at the parent spec) - never a copied plan. Receipts, impl review, done evidence, and the single-task completion-review skip compose unchanged, and the mint is atomic: two concurrent runs can never double-implement the spec.
* **Autonomous loops keep planning.** Without an explicit no-plan instruction, a zero-task spec under autonomy stops with a typed report instead of asking; pilot forwards an explicit `--no-plan` to its work dispatch and never decides it; Ralph stops typed on the new `needs_tasks` signal instead of spinning; `/flow-next:work-rolling` refuses the route - a single task degenerates the rolling frontier.
* **Workers can dispatch subagents again.** Three writing agents carried a `disallowedTools: Task` denial from their first commit with no recorded rationale - Task is the subagent-dispatch tool, not a planning feature. The audit removes it from the writers, keeps it on every read-only agent with the reason written inline, and the direct route’s worker uses the restored capability under a judicious license: parallel implementation, research, scouting, shape chosen at execution time, every subagent joined before commit.
* Under the hood: `flowctl next` surfaces zero-task specs as `status: plan, reason: needs_tasks`; `flowctl task create` gains `--require-empty-spec` (the atomic mint guard); the guide, capture, and interview routers can recommend the route for near-zero-risk fully-known specs; the Codex mirror’s transform roster and guard grew to cover every new copy-pasteable command.

### 4.8.0: The autonomy prose stops failing quietly

**Autonomous runs get thirty-four hardened rules across the worker, the land conductor, and the review rubrics. Closing failure classes banked from real overnight runs: budgets burned against stale state, gates trusted on narration instead of evidence, locks that outlive their tick, and cleanup that could sweep away work a human left uncommitted. The pass was pressure-tested on its own PR: fourteen cross-model review rounds forced twenty-nine further repairs before merge.**

Detail

* **Workers can no longer grade their own homework.** A gate, test, or baseline is never edited to make it pass; an assertion is never weakened to match a wrong implementation; an errored or wrong-surface gate observation records as inconclusive, not green. And a suspiciously fast or zero-case pass gets its log read before any green receipt is minted, so one false pass can’t poison every later run that honors the receipt. Deliberate baseline updates stay legal: a pin update the task’s acceptance names ships in the same commit with its rationale.
* **Your uncommitted work is safe from the machinery.** Both the pause path and blocked-tree cleanup commit only what the run itself produced; pre-existing uncommitted changes are left in place and named in the handoff, never swept into a pause commit or reverted away.
* **The land conductor spends its budgets on real problems.** Red-CI triage reads merge state and open threads before burning a fix attempt; a repeat identical failure is reclassified instead of re-run blind, and an infra-shaped repeat escalates instead of consuming the fix budget; flake signatures bind to the head they diagnosed; siblings re-gate after any base-moving action; spec dependencies are honored at merge, not just at select.
* **Ticks stop losing state to each other.** Each land tick claims its clone atomically before the first ledger read. Owner-aware, so a live tick is never reaped by the age threshold, an idle tick never leaves a lock behind, and a session reclaims its own abandoned claim. Dry runs take no claim and mutate nothing.
* **Reviews name their evidence.** A five-label evidence scale (claimed / cited / walked / executed / reproduced) plus new judgment probes: structure-over-instruction, wire-type leakage, legacy dual-path, the shallow-module smell with its falsifiable sign, and a mechanical 1000-line-crossing check capped at Should-Fix.
* **Questions the machine can answer never reach you.** Interview and plan resolve empirically answerable forks with a throwaway probe. But only when the probe is non-mutating or fully disposable; a stateful question still asks. Wildly divergent independent opinions trigger a reframe, never an average.
* Also: the setup-installed docs snippet now carries the standing prose-contract line for chat replies (sentinel `v2`; existing repos pick it up on their next `/flow-next:setup` run).

### 4.7.1: The prose skill applies itself

**The agent now drafts substantial replies under the contract without being asked, on every host.** `/flow-next:prose`’s description was reshaped from user phrases to the agent’s own drafting moment, and Codex joins the ambient behavior - the skill enters the implicit catalog with a dieted entry after a measurement showed the shared catalog sitting under half its budget, so the earlier explicit-only carve-out was protecting headroom that was never at risk. Manual invocation with a draft to tighten is unchanged.

### 4.7.0: Agent-written prose answers to a contract

**PR bodies, tracker comments, spec prose, done summaries, and changelog entries now draft against one shipped ten-rule contract instead of whatever register the model reached for, so filler that could describe any project, feelings standing in for numbers, and invented outcome lines get caught while the text is being written rather than in review.**

Detail

* **Nothing to switch on.** All 23 durable emission points carry a one-line pointer to the contract file and read it at the moment they write: make-pr bodies, capture and plan spec prose, interview write-backs, tracker-sync and resolve-pr comments, chart briefings, strategy sections, qa findings, land verdict comments, prospect candidates, prime glossary definitions, audit memory entries, and worker done summaries. The pointer is non-blocking - a host that cannot find the doc proceeds unchanged.
* **Your structural contracts still win.** Where a rule collides with the shape of the surface being written, the surface wins: tracker dedup markers stay first-line and byte-unchanged, projection-only source truth is never overridden, and outcome-first ordering never invents an outcome that was not in the payload.
* **The scope is deliberately narrow.** The contract governs how artifact prose reads. It makes no claim about code quality or maintainability, and cites SlopCodeBench (arXiv 2603.24755) for exactly what that evidence shows.
* **It caught its own author first.** The first enforcement pass found the contract’s own bridged author breaking rule 9, the em-dash ban, in the contract’s opening sentence. The review gate held and the sentence was rewritten before it shipped. Two further bot-review rounds grew the pointer set from 19 emission points to 23.
* **New skill: [`/flow-next:prose`](https://flow-next.dev/skills/prose/)** extends the same contract to substantial chat replies on request - opportunistic rather than guaranteed. Plain-language triggers (“tighten this reply”) work on hosts that match skill descriptions (Claude Code, Cursor, Droid, Grok); Codex is explicit-only, the same carve-out [visual](https://flow-next.dev/skills/visual/) gets. The visual digest itself stays excluded from the contract, because its output is ephemeral chat rendering rather than a written artifact.
* **Scout reports stop asserting absence for free.** A scout that reports something is missing now states the search it ran to conclude that, so you can judge the negative claim instead of taking it.

**Under the hood.** The Codex mirror’s docs-link rewrite and its hard-fail link guard now cover `agents/*.md`, so agent-file pointers mirror to paths that resolve instead of shipping dangling.

### 4.6.1: The assumption that fell to a five-minute probe, twice

**Cursor and Grok Build now run the rolling-frontier scheduler for real instead of quietly degrading to waves.** The scheduler’s prose named both hosts as blocking-dispatch by assumption; a five-minute probe on each (dispatch two sleeping agents, watch when control returns) measured non-blocking dispatch with per-completion notifications, and both hosts then drove a full rolling run end-to-end - simultaneous three-task admission, each task integrated the moment it returned while siblings kept running. The clause now binds on measured dispatch behavior with dated provenance, never on host name; a genuinely blocking host still degrades honestly, unchanged.

### 4.6.0: Excused reviews close the loop, reviewers stop re-running the world

**Two field-report sweeps in one release: a completion review that policy deliberately excused is now a recorded state every gate honors - excused specs ship instead of looping - and both reviewer prompts gained a verification budget, so a review round re-checks what a finding disputes instead of re-running the suite the run’s final gate already owns.** Plus three more field fixes: land’s post-merge bookkeeping survives its own ignore rules, dead-end review modes are refused before any work happens, and OpenCode users can run every command a closer prints.

Detail

* **`completion_review_status: not_required`** (#371, thanks @sn-furali). Work’s single-task policy skip records its decision instead of leaving `unknown`; the tracker projection, `flowctl next --require-completion-review`, pilot routing, and Ralph’s completion gate all decide through one satisfying set `{ship, not_required}`. Unrecognized values still fail closed; the write is a compare-and-set from `unknown` (a real verdict is never overwritten) that also refuses a surface no longer single-task; adding a task, rewriting the plan, editing a task’s contract, or resetting a task re-arms the review.
* **Reviewer verification budget.** Both reviewer rubrics now assign focused, evidence-named suites (plus the exact test a finding disputes) to review rounds and the full suite to the run’s final gate - closing a measured tail-latency mode where a re-reviewer ran a ten-minute full discover on a rolling run’s critical path. Token deltas in the rebaseline evidence are recorded measurements, never an enforced ceiling.
* **Land’s sync-state commit survives** (#367, thanks @sn-furali). The post-merge tail no longer names the auto-ignored receipts directory in its `git add`; the sidecar commit is guarded on a staged diff, so an unchanged sidecar (including resume-tail re-entry) is success.
* **`--review=export` refused at parse time** (#366) at both mouths - work’s surfaces and impl-review’s own parser - with a pointer to plan-review, where export actually lives.
* **Host-correct closer commands** (#364). Every closer that prints a copy-pasteable next step emits OpenCode’s flat `/flow-next-<name>` form there and the colon form elsewhere; the Codex mirror’s rewrite guard gained a positive expected-output check so a reworded literal fails the sync instead of silently staling.

### 4.5.1: Plan-sync becomes opt-in

**New installs stop paying a reconciliation pass after every completed task: on most specs the automatic plan-sync pass finds nothing to change, and [`/flow-next:sync`](https://flow-next.dev/skills/sync/) keeps the full capability for the moment a task genuinely invalidates a downstream assumption. Existing configs keep their setting, and setup still asks the question with the trade-off explained. A welcome side effect: the rolling-frontier beta’s prerequisite (`planSync.enabled=false`) is now the default state, so a fresh repo can invoke `/flow-next:work-rolling` directly. Opt back in with `flowctl config set planSync.enabled true`.**

### 4.5.0: Multi-task runs finish \~50% faster at the same quality bar (beta)

**`/flow-next:work-rolling` is an experimental beta variant of `/flow-next:work`: instead of pausing at wave boundaries until the slowest in-flight task finishes, it admits the next ready task the moment any task returns - each in its own isolated worktree, with the conductor reviewing every return under the unchanged canonical review rules. The architecture was picked by a pre-registered three-arm eval: this arm cut work-phase wall-clock 52.1% at quality parity with zero uncontained correctness incidents - and the eval rejected a faster arm, which is the part worth reading.**

Detail

* **Same inputs, same workers, same review.** Invoke it exactly like `/flow-next:work`. Reviewer identity, rubric, diff scope, the SHIP-before-done gate, and the fix-loop cap are canonical work’s, byte-unchanged; the concurrency cap stays at 3. Workers share an outside-tree run-notes surface they read by pointer - notes content is never embedded into a dispatch prompt.
* **User-invoked only.** Pilot, land, and Ralph always dispatch canonical `/flow-next:work`; the beta never enters an autonomous loop on its own.
* **Prerequisite: plan-sync off.** `planSync.enabled=true` (the shipped default) fail-closes the run to serial, canonical behavior - rolling admission needs `flowctl config set planSync.enabled false`. An interactive run offers that command once; an autonomous run only reports it and never mutates config.
* **The eval that said no.** A shared-checkout arm was faster still - 69.2% - but failed quality parity: roughly 30% thinner test artifacts. The mechanism was structural, not random: making the declared-paths list the commit boundary disincentivized new test files (tests are the artifact that spawns new files), so part of the speed was bought with under-testing. The per-task worktree pool removes that pressure, which is why the slower-but-parity arm shipped.

### 4.4.0: The menu now coaches

**Capture and plan tell you the smallest sufficient next step at the moment you choose it.** Field-requested by a flow-next team: the [pipeline-variations selector](https://flow-next.dev/choosing-your-route/) (risk and unknowns, never size) was documentation - and documentation doesn’t fire at the decision point. Now each closer prints one advisory `Recommended next:` line above its unchanged menu, judged from the spec or plan just produced. The menu stays a menu; autonomous runs are untouched. And the PR’s own external review pass hardened the Codex distribution chain along the way.

Detail

* **Capture** (base, rewrite, and per-spec split footers) judges the just-written spec: open `[inferred]` criteria and parked unknowns lean interview; real design risk leans plan; near-zero risk leans a minimal plan with plan-review typically ceremony. Legal targets are interview/plan/guide only - `chart` is upstream of capture and never recommended; genuinely conflicting signals recommend `/flow-next:guide`.
* **Plan** adds the same line on the interactive next-steps menu only (re-judged after every go-deeper round): plan-review vs straight to work, with a review skip legal only for the two documented ceremony shapes. `steps.md` is untouched, so autonomous output carries no recommendation.
* **Single rubric home:** the closers judge against pipeline-variations and link it - no copied rubric, no config keys, no persisted classification.
* **Installed Codex hardening** (from this PR’s six-round external review): docs now ship to Codex installs under an owned `docs/flow-next/` namespace - the installer can never delete or overwrite non-owned files (sentinel regression test); every mirror docs link resolves on disk or is an absolute URL (hard-fail closure guard); actionable commands render in Codex’s `$flow-next-*` syntax; and 7 long-broken mirror links from earlier releases are repaired.

### 4.3.1: Stop paying twice for the same tests

**When a review sends work back and the fix passes the full gate, the pipeline now remembers that instead of running the identical suite again a few minutes later - and on a multi-task spec, a task can inherit the previous task’s proven-green baseline.**

Detail

Nothing to enable and nothing to re-run - this is default behavior from 4.3.1 on.

A NEEDS\_WORK verdict used to cost you the test suite twice: once when the fix loop proved its fix good, and again when the Verify stage re-ran the same commands against the same commit a few minutes later. The fix loop now writes the same green receipt the pipeline already trusts, at the committed fix HEAD, so Verify honors it rather than repeating work it can prove already passed. That is 2-10 minutes back per NEEDS\_WORK to SHIP cycle; mining real receipts found 15 to 18 affected tasks in a typical repo.

The second half is the baseline. On a multi-task run, worker N+1 can be handed the previous task’s result instead of re-proving a tree nobody touched - but only when the prior Verify ran the same Quick commands green, HEAD has not moved, and the two tasks’ file sets are disjoint. The worker records where the baseline came from, so the provenance stays in the receipt.

Both halves fail closed, deliberately. A focused suite never mints a full-gate receipt, any doubt means the gate simply runs. The first task of every spec still runs its own baseline - that is the check that catches a red your local environment has and CI does not, which is exactly how this shipped (a test file red at base here while CI was green). Lint always runs.

### 4.3.0: OpenCode joins the first-class roster

**If OpenCode is your daily driver, flow-next now installs once from the canonical repo and the whole pipeline follows - planning fan-outs, cross-model reviews, receipts - riding every release. The stale community port is superseded and archived.**

Detail

* **One installer, no port to chase.** OpenCode has no plugin format, so `./scripts/install-opencode.sh` scatters the canonical files into `~/.config/opencode/`: skills as-is, generated agents whose tool denials translate to OpenCode’s permission map, flat `/flow-next-<name>` command stubs, and the support dirs that make `flowctl` and the spec-template cascade resolve unchanged. A deterministic ownership manifest scopes every re-run and `--uninstall` to installed paths only - a colliding user directory aborts the install instead of being deleted.
* **Setup runs like every other host.** The ownership manifest doubles as setup’s platform-detection signal, so `/flow-next-setup` writes AGENTS.md instructions in the flat slash form - never Codex-shaped snippets.
* **Host review works, and it earned a hard rule everywhere.** OpenCode subagents take their model from their own agent definition or inherit the session model - there is no dispatch-time override. Pinning a tier’s model is one 5-line user agent file; verified live, the conductor matched the routing block’s reviewer model to the pinned agent unhinted and the receipt recorded the real reviewer. Without a pin the degraded reviewer self-reports and the review fail-closes rather than letting the session model grade its own work. And on every harness, the host backend now carries the rule the name implies: it never shells out to another CLI.
* Verified on opencode 1.18.19: 29/29 skills and 20/20 subagents discovered with correct permission maps, a full plan scout fan-out, codex-backend review, and host-backend review - all driven end-to-end from an OpenCode session.

### 4.2.2: Runs finish sooner, reviews check the same things

**First batch of fixes from a measured wall-clock pass over the whole pipeline: small specs stop paying a duplicate review, multi-task waves stop paying repeated plan-sync passes, and plans that could run tasks in parallel stop silently running them one by one.**

Detail

* **Single-task specs skip the completion review when it would re-review the same diff** - the spec’s only task already holds a per-task SHIP and covers every requirement. The skip is recorded as an explicit stage line, never a silent absence; multi-task specs keep the full completion review, whose cross-task integration value is the point.
* **Plan-sync runs once per resolved wave** with the full completed-task set instead of once per task - same downstream scope, same per-task drift verdicts, k× fewer dispatches.
* **A missing `Touches:` line is now a plan-review finding** on multi-task specs (omission silently forces serial dispatch - waves are fail-closed on the line), and pilot’s evidence echo repeats the work stage’s `Sequential fallback:` reason so a driver loop can see when a spec ran serial and why.
* **Parallel agents in sibling worktrees:** flowctl’s runtime claim state lives in the git common dir and is shared by every worktree - now documented, with the `FLOW_STATE_DIR` override (set it outside the repo tree) for concurrent same-repo pipelines.

### 4.2.1: Setup prices each opt-in

**Setup prices each opt-in before you answer it.** The review-backend question now says where the time goes (each review round is a serial pass the pipeline waits on - usually the largest wall-clock item in a run), the `None` option spells out what still gates a run and what stops being checked, and every host’s menu offers `Host` - the reviewer that keeps every gate with no second CLI, configured by one `reviewer:` line in the routing block. Memory, plan-sync, GitHub-scout, and HTML-artifact questions carry the same one-line cost shape. The full dial from a cross-model backend down to `host` or `none` is priced in Running Lean.

### 4.2.0: Land asks the human reviewer when it is their turn

**Teams whose merge gate is a human - a code-owner review required by a ruleset, sharpest when the PR author is a GitHub App that cannot be a code owner - no longer watch a converged PR sit idle until someone happens to notice it. Opt-in `land.requestReviewers` makes land request the right people at the one moment their review is the only thing left. Thanks @sn-furali (#359).**

Detail

* **The ask fires on a sharp predicate, not “at convergence.”** CI green, zero unresolved threads, and a human review as the sole missing merge input - so it never deadlocks under `reviewSignal: approve` (where convergence *is* the approval) and never notifies a human about a PR that merges the same tick.
* **The list is yours: csv of GitHub logins and/or `org/team` slugs and/or the literal `codeowners`** - the `codeowners` token rides the draft→ready flip and GitHub resolves the owners itself (no local CODEOWNERS parsing). The PR author is filtered out; “ready” keeps meaning “a human may review this now.”
* **Exactly once per PR per head SHA.** The request is recorded in the land ledger and claimed atomically, so overlapping ticks cannot double-ask; a CI-fix push moves the head and re-asks only if the human’s review is again missing - a genuine re-request, not spam. A failed request records the head anyway and surfaces `reviewers=failed:<reason>`, never a retry loop.
* **Never a merge gate.** `land.reviewSignal` still decides what counts as reviewed; the key only asks. Default off - with it unset, every gate, action, and ledger write is byte-identical, and the evidence line gains one additive `reviewers=off` field. `--dry-run` reports `would-request` and mutates nothing.
* Part 2 of the report (`land.draftOnChangesRequested`) is deferred until this has proven itself in the field.

### 4.1.0: Pilot composes with worktrees

**Pilot loops finally start inside a git worktree, stop refusing branches that carry merged gate PRs, and your tracker’s workflow states are checkable without touching config. Three field reports, each verified against 4.0.0 before fixing - thanks @sn-furali (#354, #355, #356).**

Detail

* **Plan stages run where you stand** (#354): pilot’s plan and plan-review stages demanded a checkout of the default branch, which git refuses in any secondary worktree - so the documented one-worktree-per-spec setup could never start the loop. The rule now checks the actual hazard instead: an open PR on the current branch. No open PR means plan in place; only a branch with an open PR still steps aside, and only a failed step-aside asks for a human.
* **Merged gate PRs are history, not an inconsistency** (#355): a team opening one PR per gate (capture, plan, review) reaches make-pr with several merged PRs and unshipped work on the same branch. Pilot called that state inconsistent while make-pr’s own rules called it normal. Pilot now compares the branch head against the newest merged PR’s head - squash-merge-safe, unlike counting commits against the default branch - and refuses only when the branch truly holds nothing new.
* **`flowctl tracker wire list-states`** (#356): a read-only verb listing Linear workflow states or Jira statuses (id, name, type) with an explicit `complete` flag, so a truncated page can never masquerade as the whole board. Answers “does every configured state id still name a live state?” without a second tracker client - and without `tracker resolve`, which repairs the mapping and writes config. Detection never writes; repair stays where it was.
* **A failed review round can no longer wedge a spec’s reviews forever.** When a backend returned no parseable verdict, flowctl refunded the round but left bookkeeping that could never finish - every later review of that spec then refused with `REPLAY_REQUIRED`, with no way out short of hand-editing `.flow` state. Refunded rounds now clean up after themselves, and repos already wedged by an earlier version heal silently on their next review. Found dogfooding this release’s own pipeline.

### 4.0.0: Your repo stops carrying flow-next around

**Installs are copy-less. Setup no longer copies the CLI, the agent guide, or the spec template into your repository, so updating the plugin is the entire update - no per-repo re-run, no stale copy quietly running last month’s behavior. If an older install left `.flow/bin/`, `.flow/templates/spec.md`, or `.flow/usage.md` behind, delete them; nothing reads them. And the second thing you used to maintain per project is gone too: choosing which model does what is now a short block of your own words in your own instruction file, not a configuration ceremony that stored model names it could not keep true.**

Detail

**Do these first (all optional - nothing breaks if you skip them):**

1. **Delete your copies.** If your repo has `.flow/bin/`, `.flow/templates/spec.md`, or `.flow/usage.md`, remove them (`git rm` if they are tracked). Nothing reads them any more, and a stale copied `flowctl` can shadow the current one. `/flow-next:setup` offers to delete them for you, and `/flow-next:plan` prints a one-line nudge when it sees them.
2. **If you used `delegate:codex`,** run `/flow-next:setup` and accept the model-routing scaffold, then drive the other CLI through the bridge recipes in `flowctl usage`. The packaged delegation subsystem is gone.
3. **Leftover config keys are inert, not dangerous.** `work.delegate*` and `models.*` keys still sitting in `.flow/config.json` are named once in a non-blocking advisory and otherwise ignored; delete them when convenient.

The docs-snippet schema did not change in this release, so nobody has to re-run setup to stay current.

**Nothing lives in your repo except your work.** Previously most projects ran in “copy mode”: `flowctl`, the agent guide, and the spec template landed as snapshots under `.flow/`, and every plugin update meant re-running setup in each project - or silently running an old CLI, which is the failure mode behind most “this flag should exist” reports. Now every host resolves `flowctl` from the plugin install itself: Claude Code and Factory Droid through their plugin-root environment variables, Codex through its own home, and Cursor and Grok by deriving the plugin root from the absolute skill path those hosts already hand the agent. Update the plugin and you are done. `/flow-next:setup` is now for the first run, a configuration change, or the rare release that says the docs-snippet schema bumped.

**Routing is something you say, not something you configure.** Setup used to ask which model should play which role, probe a CLI for the model ids it served, and store what it found - configuration that claimed what was installed, and then aged into configuration that lied. That ceremony, the role map, and its staleness stamps are gone. In their place: four named tiers - reviewer, implementer, fast scout, thinking scout - and a routing block in your own `CLAUDE.md` / `AGENTS.md`, written in your words with model names you can verify against your own account. Every session reads it, including unattended pilot, land, and Ralph ticks, and applies it with judgment. If you never routed anything, nothing changes: each tier falls back to the session model, exactly as flow-next has always run out of the box.

**A tier says which model executes a stage, never which stages run.** Planning, capture, interview, every verdict, and the worker stay on the session model unless you say otherwise. Where the harness exposes it, a stage records the model that actually ran, so a routing preference written in prose leaves evidence instead of a hope - and unavailable provenance is recorded as unknown, never as the configured value.

**Can I even do that here?** has an answer per harness now. The new reach pages state, for each supported host, which mechanisms exist (in-session model, in-host subagent, another CLI over a bridge), which do not, and what the degradation is when one is missing. Skills ask for a tier and name no spawn primitive, CLI flag, or vendor path; a harness flow-next cannot identify resolves to the generic page and says so.

**What is gone.** The packaged codex-delegation subsystem: six `work.delegate*` config keys, the `delegate` role pin, the `delegate:codex` / `delegate:local` work arguments, the delegation gating and result-classification path in the work skill, two `flowctl codex` subcommands, and setup’s delegation option. The standard work loop is unchanged - delegation was off by default, and the host always owned gating, git, review, and commit. Two rules carry over to the bridge route and are not optional: the bridged child writes code while the host keeps git, judgment, and the verdict.

**Under the hood:** the plugin-root derivation is one added rung in every skill’s resolution chain - environment variable first, the skill file’s own absolute path second, the legacy `.flow/bin/flowctl` third as a silent backstop for repos that have not deleted their copies yet. The backstop is not going away, but nothing is documented, tested, or designed around keeping copies. `flowctl setup-mode` is gone from the CLI, the help text, and the docs; old `setup_mode` / `setup_version` stamps in `.flow/meta.json` are tolerated as inert metadata and never read.

### 3.34.0: Seven field reports, six issues closed, zero new config keys

**A two-day batch of forensic-grade reports, every claim verified before fixing. Locked-down repos get a merge-identity seam and a tail that finishes; land loses its force-push; fresh clones can finally pass `validate`; a phantom lock holder stops eating 10 seconds. Thanks @sn-furali (five reports) and @TechupBusiness.**

Detail

* **`FLOW_PR_MERGE_CMD`** (#337): env-only merge-identity seam mirroring the create side’s, with a stderr-verbatim contract so land’s race-vs-refusal judgment survives interposition.
* **Server-side catch-up** (#342): `gh pr update-branch` replaces checkout+rebase+force-push - commit SHAs survive, removing the *cause* of orphaned evidence; repeated refusals escalate bounded.
* **The post-merge tail finishes** (#345): release and board update run before the one push a PR-only base can refuse - a refused push is a bookkeeping note, not a stuck board.
* **Fresh-clone `validate` is meaningful** (#347): committed-snapshot status warns instead of erroring (582 structurally-false errors on flow-next’s own clone became warnings); real inconsistencies still fail.
* **`done` names the receipt it wrote** (#346): `modified_paths` + a dirty-file advisory; the commit-done-stage ordering is canonical on every documented surface.
* **Honest lock diagnostics** (#340): a lock that can never be created fails fast with errno and path; “holder appears alive” is reserved for an owner that was actually read.
* **The attempt ledger answers “which model reviewed this?”** (#338, recording half): rows carry the model/effort that actually ran; the publication half was declined - flow-next never writes its own verdict into the channel land treats as independent-review evidence.

### 3.33.0: Independent tasks finally run at the same time

**When a spec’s tasks do not depend on each other, flow-next can implement them side by side - and does, but far less often than your task graph allows. The rule needs each task to declare which paths it will touch, and the planning guidance said to skip that line whenever it was hard to predict, so a default behaved like an exception. It is written on every task now, and two independent tasks that took 187 seconds one after the other take 96 seconds together, on the same tokens.**

Detail

Nothing to enable and nothing to configure - plan a spec and the tasks carry the declaration. What changed is the default: planning writes a `Touches` line on every task and declares wider rather than omitting when it is unsure. Previously how often waves fired came down to how boldly a given planning pass read the omit-when-unsure advice, which is why the same feature could run in one repo and stay dormant in another. Both choices err the same way (a declaration that is too wide overlaps a sibling and keeps the tasks serial, exactly like omitting it), but only the declared one can ever become a wave once the overlap turns out to be false.

A wave still refuses unless everything lines up: same spec, at most three tasks, no dependency path between any pair in either direction, declarations that do not overlap, and no task touching the always-serial set (flow state, lockfiles, migrations, generated output, spec and task files). Any doubt sends the whole wave serial, which is the old behavior. Getting a declaration wrong is cheap on purpose: workers write in isolated workspaces, so an overlap shows up as a merge conflict when their commits are joined, costing one serial re-run and never correctness. Troubleshooting now has an entry for exactly that moment.

Two things that broke the first live wave are fixed where the conductor reads them. A wave workspace is branched from a commit, so a spec that was just planned and not yet committed does not exist inside it, and every parallel worker fails to re-anchor while looking like a broken worker. And each worker’s handover evidence names commits that live only on its own workspace branch - recorded unchanged, every finished task ends up pointing at commits that disappear when the workspace does, which validation then reports as orphaned. Both are now stated preconditions.

**Under the hood:** measured before shipping by running the path end to end - two workers dispatched concurrently into linked worktrees, both honoring their declarations exactly, both deferring completion and review to the conductor, join clean, both tasks verified done. 96s concurrent against 187s serial, 16s of join overhead, no token penalty. Both preconditions are pinned by tests on the canonical prose and the generated Codex mirror.

### 3.32.2: Cursor’s in-IDE browser drives for real

**On Cursor, agents skipped the built-in browser because the instructions were wrong. They now probe it by id, and if that probe fails they ask you once to type `@Browser` (no space) so the pane is connected, then try again. A vanished MCP mid-run is a partial pass, not “the rung does not exist.” Console and network from the driven page stay unverified, so a QA pass routed here must record BLOCKED (could not verify) rather than PASS.**

### 3.32.1: Model guidance catches Grok 4.6 on day one

**Grok 4.6 shipped 2026-08-12; within a day the setup scaffold, bridge recipes, and orchestration docs carry an evidence-based reprofile - built from independent benchmarks and real user reports, not the launch post.**

Detail

* The model-routing scaffold’s grok tier moves to `grok-4.6`: intelligence up (Artificial Analysis Index 61, tied with GPT-5.6 Sol Max; real-user consensus \~Opus 4.8-tier), raw speed down but \~2x turn efficiency, taste unchanged.
* Routing sharpened along the split the independent evals exposed: supervised editor-shaped implementation is its strong surface (CursorBench 69.9%, day-one Cursor with a permanent 2x usage pool - bridge slug `cursor-grok-4.6-high` verified live); long unsupervised terminal loops are its weak one (Terminal-Bench v3 26%).
* The never-the-gate posture stands, now with numbers: AA-Omniscience measures it inventing \~1/3 of the time when it doesn’t know; API cache reads cost 67% more than 4.5 on long sessions.

### 3.32.0: The plan you can read at a glance - and the skill that reviewed itself

**After `/flow-next:plan`, reviewing meant reading a spec plus seven task files and rebuilding the structure in your head. `/flow-next:visual` restates it as one screen of compact markdown - and its own dogfood pass caught a real grounding bug before it shipped, while the PR that shipped it carried the first diff-fenced structural sketch in place of an edge-less mermaid diagram.**

Detail

* **`/flow-next:visual`** - point it at a spec, a task, a git range, or the conversation; it restates the thing with a fixed 8-shape vocabulary (call trees, file trees, diff-fenced sketches, type sketches, tiny tables; mermaid last) in plain fenced blocks that colorize natively on every host. Read-only, chat-only.
* **Post-plan digest** (the primary mode): thesis, task tree in dependency order, planned file-layout diff with owning tasks annotated, R-ID coverage line where uncovered requirements jump out, IS/IS-NOT boundaries.
* **Hard grounding**: every path from task files/spec/`git diff --name-status`, every edge from real code or a real dependency, coverage from declared `satisfies` frontmatter - never invented; missing state degrades to the nearest viable mode with a one-line notice.
* **Offered where the text walls are**: capture, plan, and interview closers suggest the digest at their read-back moment - an option you pick, never auto-run.
* **make-pr learns the sketch**: `## Structural changes` may emit a diff-fenced file-tree or call-tree sketch where mermaid is weakest (cap-forced collapse, or a diagram under four nodes) - same hallucination guardrails, no silent-rendering-failure risk.
* **Review dividend**: the PR’s bot-review round exposed a long-standing sync bug flattening fenced-block indentation across the whole Codex mirror - fixed fence-aware, 159 mirror files restored.
* On Codex the digest is explicit-only (`$flow-next-visual`) by design - its trigger-rich description stays out of the shared skill-catalog budget.

### 3.31.0: Your repo’s own command can now gate the merge

**On a free-plan private repo, branch protection and rulesets 403 - no required status check can exist - so land’s gate tree read a server with nothing to say, and a repo-local gate of record had no way to bind the merge. Thanks @TechupBusiness for the precise gap analysis.**

Detail

* **`land.mergeVerdictCommand`** (opt-in, fail-closed): once every other gate passes and the planned action is merge, land runs the repo’s command once - exit 0 merges; missing, unexecutable, timed out, or signal death blocks with `NEEDS_HUMAN` and never skips.
* Context arrives as environment only (`FLOW_HEAD_SHA`, `FLOW_BASE_REF`, `FLOW_PR_NUMBER`, `FLOW_SPEC_ID`); the command string is never built from PR-derived text.
* **The verdict binds the (head, base) pair it judged**: a push after the verdict refuses server-side at `--match-head-commit`; a base that moved re-ticks; a trust guard refuses to execute from a non-base checkout.
* `--dry-run` reports `would-run` and executes nothing; unset/null/empty all mean off.
* Also documents `gate classify`’s known fail-open (CI-guarded generated docs, #334) with the conductor-prose remedy.

### 3.30.0: Codex reviews get their verdicts back

**`review.backend codex` could return no verdict 13 times in a row while every probe said the backend was healthy - the reviewer subprocess was inheriting instructions never meant for it (the host repo’s auto-loaded `AGENTS.md`, and the plugin’s own coordinator skills saying “never self-declare a verdict”). Thanks @sn-furali for the isolating-controls table that separated the two routes.**

Detail

* **Persona override on for codex**: the counter-instruction preamble (cursor-only since fn-90) now rides every codex review; the codex `False` was inherited from a refactor, never decided.
* **Host project docs suppressed**: `-c project_doc_max_bytes=0` on both the fresh and resume `codex exec` dispatch - the reporter’s measured fix.
* **The plan-review prompt states its role**: it was the only review prompt missing the “You ARE the reviewer” anchor, which is why plan reviews failed first.
* **Honest failure classes**: a healthy exit-0 run with no verdict journals as `missing_verdict` (the old ladder matched the word “timeout” in the reviewer’s own prose), and a streak of them terminates with instruction-contamination guidance instead of “repair the backend”.
* Repo review invariants belong in `.flow/criteria.md`, which rides the review prompts to every backend - now documented.
* Release-gated on the reporter’s own reproducer shape: a live plan review on flow-next’s own large `AGENTS.md` returned a verdict on the first attempt.

### 3.29.0: The worktree kit works without a keyboard

**`cleanup` - the kit’s only sanctioned removal path - read two answers from stdin unconditionally and died silently with no terminal; `create` left the new branch tracking the base, so under `push.default=upstream` a bare `git push` aimed at the base branch. Thanks @sn-furali.**

Detail

* **`cleanup [<name>...] [--yes]`**: names skip the prompt, `--yes` skips the confirmation (required off a terminal), EOF-guarded reads fail loudly naming the remedy. Interactive behavior unchanged.
* **`create --no-track`**: an isolated per-worktree branch no longer tracks the base’s remote branch; first push under default config needs `-u`.
* **Invocation story told honestly**: the five phrase-triggered skills answer to plain language and - on hosts that surface skills as commands - their full skill name, e.g. `/flow-next:flow-next-worktree-kit`.

### 3.28.1: Your memory entries, back exactly as you wrote them

**A memory entry edited with `memory add --update` could silently come back different: a mid-string `" #"` was written unquoted (YAML comment syntax - any conforming parser truncates the value there, while flowctl’s own fallback reader hides the damage), and without PyYAML installed, quoted list items containing commas were split apart and the mis-parse written back. Both fixed, with the one damaged entry in flow-next’s own repo repaired. Thanks @sn-furali for the forensic report.**

Detail

* **Writer**: any whitespace-then-hash scalar is now quoted; non-comment hashes (`C#`, `issue#140`) stay plain.
* **Reader (no-PyYAML fallback)**: flow lists and mappings split quote- and depth-aware via one shared helper - quoted scalars recognized at item starts, after mapping key separators, and at nested collection boundaries; double-quoted keys/values unescape correctly.
* Deliberately not done: teaching the fallback to strip comments (it would turn latent damage into active truncation for repos with unquoted `" #"` already on disk).

### 3.28.0: The strikeout you can actually recover from

**On a board-armed repo, a spec that struck out of the pilot loop could read ready everywhere a human looks while staying permanently invisible to pilot - and the docs described a recovery that has been impossible since fn-87. A measured three-phase report proved both halves. Thanks @sn-furali.**

Detail

**Nothing to enable.** Default behavior.

**What was hard before.** Pilot’s two-strike guard takes a failing spec out of selection. With `tracker.readyState` armed, the next pull re-readies it from the board - but since fn-87, deliberately, that projection-set ready never clears the strike: the board echo re-grants readiness with nobody acting, and clearing on it would re-dispatch the same failing spec every tick forever. The docs still described the old behavior, and the skill’s own escape clause (“an explicit re-ready, not a projection echo”) had nothing to key on - the reporter measured that a deliberate out-and-back board move is byte-identical to an echo in every durable artifact. The only real escape was hand-editing an undocumented file under `.git/`.

**What you get now.** `flowctl pilot strikes list` shows what is struck and why; `flowctl pilot strikes clear <spec-id>` (or `clear --all`) is the recognized human recovery - atomic, shared across worktrees, with a distinct not-found for unknown ids. The `strike 2/2` verdict names the command, so the transcript carries its own way out. Clearing a strike never changes spec readiness: strikes are pilot state, the board keeps owning readiness.

**And the docs tell the truth.** Every surface that claimed a board move clears strikes now states the fn-87 rule, troubleshooting documents the ledger, and the board-native alternative (clearing when a tick observes the issue leave and re-enter the ready lane) is recorded as a deferred decision - it narrows but cannot remove the ambiguity the reporter proved, and silently misses a fast out-and-back between ticks. (#325)

### 3.27.0: The tracker bridge stops racing, lying, and losing Projects

**Two agents promoting the same intake issue could each end up with their own spec; a dedup query against a populated Linear board could look healthy while blind; and a Linear issue could not be placed in a Project at all. The tracker-bridge batch closes all three and writes down the abandon path. Thanks @sn-furali for the measured reports.**

Detail

**Nothing to enable** - Project placement is per-spec opt-in; everything else is default behavior.

**One winner per candidate.** The create-first mint claim is now compare-and-set: `sync create-first-put --if-absent` records the minted spec only while the claim slot is free, and the loser of a concurrent promotion exits with a distinct conflict naming the winner to adopt - instead of silently overwriting it. The tracker-sync ceremony wires the CAS into the canonical path, teaches adopt-the-winner (retire the duplicate with `spec close`, never re-put), refuses a claim when the record is already promoted and cleared, and defers the whole collision to a human under autonomous operation. (#310)

**A refusal you can handle, not a lie you cannot detect.** Linear `wire list-open` with `tracker.readyState` unset returned an empty success - indistinguishable from a genuinely empty board. It now returns an explicit error naming the unresolved key and how to set it; leaving it unset remains a valid, deliberate configuration, and backlog automation treats the refusal as “no ready lane configured”. (#311)

**Issues land in their Project.** Optional per-spec sidecar fields `tracker.projectId` / `tracker.projectMilestoneId` are sent on issue creation and reconciled on every sync push. Absent means unmanaged: payloads stay byte-identical and a Project set on the Linear side is never cleared by flow-next, which carries exactly the id it is given and never creates or manages Projects. Verified live against the Linear sandbox, including the never-clears contract. (#315)

**And the abandon path is written down.** For a candidate that will never be promoted: close the issue in the tracker first (that side is yours), then clear the local record - ordering stated so a live intake issue is never left without a trace. (#309)

### 3.26.0: Coverage that tells the truth twice

**A fully-planned spec that had not shipped code yet looked like 0% coverage - and make-pr refused to open the draft with advice you could not follow. And a rebase could orphan every evidence commit a spec recorded while validate stayed green over the dead links. Coverage now answers the plan-gate and merge-gate questions separately, and validate tells you when history rewrites have voided your evidence. Thanks @sn-furali for both measured reports.**

Detail

**Nothing to enable.** All default behavior.

**Two coverage questions, two answers.** The export payload gains `undeclared_r_ids` - criteria no task claims at any status - beside the unchanged `uncovered_r_ids` (criteria no done task evidences). make-pr’s coverage abort now fires only on undeclared coverage, the one state where “go declare coverage” is advice you can act on. A plan-gate spec renders honestly: the coverage table gains a third state (`claimed, not yet evidenced` beside evidenced and undeclared), the warning marker belongs only to genuinely unclaimed criteria, and the ratio stays evidenced-only with the claimed/undeclared counts appended when non-zero. (#301)

**Evidence links stop dying silently.** A rebase, amend, or squash-merge leaves recorded evidence SHAs present in the object store but unreachable from HEAD. `flowctl validate` now warns per orphaned commit (“recorded value left as-is”) while reachable commits stay silent and tokens that are not commits in this repo - tracker UUIDs, foreign SHAs - are ignored by design: flagging them would corrupt exactly the evidence the record exists to hold. Nothing is rewritten, the run never fails, and the whole pass costs two batched git reads regardless of commit count, so the land loop can keep calling validate freely. make-pr marks orphaned SHAs as annotated text instead of rendering commit links that 404. (#302)

### 3.25.0: Six measured reports, six fixes, nothing silent

**A criterion your spec wrote should never vanish without a trace, a wrong platform guess should not survive on the primary host, and a half-resolved tracker map should never read as resolved. Six field-reported defects fixed in one pass - each with the reporter’s verified repro as the acceptance fixture. Thanks @sn-furali.**

Detail

**Nothing to enable.** All default behavior.

**Criteria stop disappearing.** The export parser now reads title-form (`**R14 - title**`) and parenthetical-form (`**R15 (note):**`) acceptance criteria and keeps text wrapped across lines - the reported repro parses 5 of 5, and 25 criteria were recovered across flow-next’s own specs. Anything still unparseable is counted and surfaced as `acceptance_criteria_residue` in the export payload, so a short coverage denominator is visible instead of silent. Suffixed R-IDs (`R4a`) are accepted by the PR cognitive aid’s validator too - the last straggler of the grammar widening. (#300, #303)

**Setup guesses right.** A single `SPEC.md` on a case-insensitive filesystem no longer counts twice and prints a bogus both-files warning (files are counted by inode now). And Claude Code - the primary host - detects as Claude Code: the old signal never reached a plugin skill’s environment, so setup fell through to the Codex fallback and wrote the wrong snippet syntax. The cascade now keys on `CLAUDECODE` paired with the Claude plugin manifest, positioned so hosts that prove themselves with their own signals still win over the inherited marker. (#305, #306)

**Tracker resolution finishes the job.** `tracker resolve --select` used to persist just the slot you picked, leaving the rest unfilled while the scope stamped fresh. It now runs the normal assignment over the remaining slots and persists the union; a map still missing a required slot is kept but reported as CONFLICT and left unstamped, so a later plain `resolve` repairs it. Configs already half-stamped by this bug self-repair on the next `--select`. (#308)

**Repair is not takeover.** `flowctl start --reclaim` rewrites a task’s claimant when it is held by a stale or wrong identity, recording `Reclaimed from <identity> (identity repair)` - distinct from `--force`, which keeps its takeover meaning and note. Only the claim-ownership gates relax. (#316)

**And one promised answer.** The gate-classify path taxonomy is deliberately closed to config: per-repo gate policy belongs in your conductor instructions (CLAUDE.md / AGENTS.md), with `pilot.gateClasses` as the open vocabulary, and classifier reason strings are not a stable contract. (#313, docs)

### 3.24.1: Judgment stays on a judgment model

**Two field-reported routing papercuts on non-Claude hosts. The plan skill’s gap analyst and judgment scouts could silently run on the host’s fast default when the Claude model alias in their agent files wasn’t resolved - the plan prose now states they run on the session model there (scanner scouts may still ride the fast tier). And setup could write a Cursor model pin with an id your account doesn’t serve (ids vary per account - `composer-2.5` vs `composer-2.5-fast`): pins are now restricted to ids seen verbatim in this run’s probe output, and a failed probe writes no pin.**

### 3.24.0: Review verdicts you can audit, not just believe

**A resumed reviewer session can answer from its previous round’s context in about a kilobyte, with zero tool calls, while asserting “measured” facts that happen to be true - and the verdict text is indistinguishable from a real review. Every review attempt now records how its verdict was produced: how much output it cost, how many tool calls it actually made where that could be measured, and exactly which commits it judged. You ask “was this measured?” of the ledger instead of taking the narration on faith.**

Detail

**Nothing to enable.** This is default ledger behavior. Every new attempt row carries the fields; nothing about how reviews run changes.

**What was hard before.** The attempts ledger recorded what was reviewed - backend, verdict, output hash - but not how the verdict was produced. A reporter measured resumed review sessions returning SHIP in 1.1-1.6 KB with zero tool calls, stating facts that were true but answered from the previous round’s context while claiming fresh measurement. Verdict-text inspection cannot catch this by construction: the fabricated verdict states true facts, and the resumed session even reuses the same thread id. Only work volume separates a review that measured the repo from one that remembered it.

**What you see now.** Every new `review_attempts[]` row records `output_bytes` (always - the size, never the output itself), `tool_calls` where the codex event stream let the dispatcher genuinely count them (a recorded `0` is the signal itself; plain-text paths carry no key at all), `head_sha_observed` marking whether the reviewed commit came from a pre-dispatch snapshot or the finalize-time fallback, and `base_sha` beside `head_sha` wherever the review snapshot ran - so the judged diff can be located and re-rendered. `flowctl review-rounds attempts --json` surfaces all of it.

**Absence means unknown, never zero.** Rows written by older versions carry none of the new fields and read back untouched. The tool-call count is measured only from a genuine codex `exec --json` stream - a review from another backend whose prose happens to quote codex-shaped event lines never gets a fabricated count - and a crash between the write-ahead journal and the ledger write replays the row with its measured provenance intact.

**What this does not change.** No verdict-validity rules, no re-review policy, no reviewer behavior change, no new commands. The consumer asks the question; flowctl only makes it askable. Fixes #312 - thanks @sn-furali for the measured report.

### 3.23.0: Status answers say where they came from

**A status read that answered from a stale snapshot looked exactly like a right answer - a review sandbox once burned three review rounds arguing with a spec that was already done. Status output now names its source, the pre-work commands warn when your checkout is behind, and reviewers are told task lifecycle is not theirs to judge from committed files.**

Detail

Two ways a wrong answer used to dress as a right one. Task status lives in a runtime store shared across your worktrees, but when that store is not reachable - a fresh clone, a review sandbox scoped to the diff - flowctl fell back to the committed snapshot without saying so; a reviewer reading that snapshot marked a finished spec NEEDS\_WORK at full confidence, three rounds running. And the checkout itself can be behind the shared truth, so the commands you run just before starting work answered from yesterday’s state without comment.

What changes for you: `flowctl show` and `list` now carry `status_source` on every task - `flow-state` when the authoritative store answered, `committed` when you are reading a snapshot that may be stale - and plain output adds one advisory line when runtime state is absent entirely. `ready` and `anchor` tell you once when HEAD is behind its upstream, because a wrong answer right before you start work is the most expensive one. Both shared review prompts now state that committed task files are snapshots and that a task looking unfinished there is never grounds for a finding.

What does not change: nothing fetches, nothing blocks, no freshness is enforced - a stale checkout still gets its computed answer, now qualified. The high-frequency polls (`list`, `status`, `next`) deliberately gained no upstream check, so the fast-poll performance stays intact; the advisory costs one read-only git probe on the two pre-work commands only.

Under the hood: provenance is stamped at the single merge point and stripped on every persisted write; the upstream probe is one `git --no-optional-locks status` spawn; 24 behavioral tests include spawn-count assertions locking the hot-path exclusion. Fixes #304 and #307 - thanks @sn-furali for the measured reports.

### 3.22.0: Handovers point at the work, not a retelling

**When one agent finished and handed off to the next, it wrote the story of what it had just done a second time - and that retelling started aging the moment anything moved, cost you a full re-read on every consumer, and could disagree with the files it was describing. A finishing agent now hands over pointers: the task, its status, where the summary and evidence live, what changed, and the verdict. Whoever picks it up reads the files themselves, which are the current truth, and every finished run ends with a next step you can actually run.**

Detail

**Nothing to enable.** This is default behavior. Run the same commands and the handoffs get shorter and stop drifting.

**What was hard before.** A worker would finish a task, write its summary and evidence to disk, and then narrate the same thing back to whoever dispatched it: what it implemented, which files it touched, which tests it ran. Two copies of one story. The copy in the handoff was written from memory of the work rather than from the artifact, so the two could disagree - and when they did, the one you read first was usually the wrong one. It also meant the same content was paid for twice, once to write and once to read back, on every single task in a run.

**What you see now.** A finishing worker reports where the outcome lives: the task id, the terminal status, the paths to its summary and evidence, the range of commits it produced, and - where its path produced one - the review verdict. It no longer restates what it built. The conductor opens the files, which are the thing that actually exists, so the account you read is the account on disk. Workers running in a parallel wave also name the workspace they were given and their gate results, which is what the join needs to reconcile them. A return that restates content the files already carry now counts as a broken contract, not a stylistic preference.

**The end of a run is a runnable line.** The work skill’s final summary closes with a `Next:` line you can execute - open the pull request, or run QA first when your pipeline asks for it - instead of a description of what you might do next. The chart-to-capture handoff is the same idea: it hands you a paste-ready command carrying the briefing path, so you run the handoff rather than reconstruct it from a paragraph.

**The doctrine is written down.** The teams page now carries pointer-shaped handover as its fifth handover property, with the two carve-outs stated honestly rather than left as folklore. A consumer without access to the repo still gets content, because there content is the only transport available. And a bounded control signal - a verdict enum, an id, a strike class - stays inline and counts as a pointer, not a retelling, so a driver reading only the transcript never has to open a file to learn whether something passed.

**What this does not change.** The files were always the record; what changed is that the handover now defers to them instead of competing with them. Nothing about how summaries and evidence are written moves, and merge judgment stays where it was.

### 3.21.0: Nits stop crowding out defects

**One reviewer reading your whole change has one pool of attention, and naming conventions are far easier to spot than a requirement quietly implemented wrong - so a run could come back with a tidy list of hygiene notes while a behavioral defect sat unremarked. The in-host quality audit now runs as two reviewers at once with separate jobs: one asks only whether the code does what the spec said, the other only whether it is code you would want to keep. You get both reports side by side, in full, and only the correctness reviewer can call something Critical or decide the change is shippable.**

Detail

**Nothing to enable.** This is how the audit behaves by default. Run the same commands and the review phase reports back the way it always did, with two headings instead of one.

**What was hard before.** A single generalist reviewer had to hold the whole change and decide, in one pass, what mattered most about it. Hygiene findings are cheap to see and easy to justify; a silent regression of something an earlier task got right, or an assertion that was weakened rather than fixed, takes real reading to notice. When both compete for the same attention budget, the cheap findings win often enough to matter. The failure was rarely a bad review - it was a review that spent itself on the wrong axis and never got to the thing that would have bitten you.

**What you see now.** Two reviewers work the same change at the same time, each told exactly one question to answer. The correctness reviewer looks at spec conformance, places the spec was interpreted differently than you meant, silent regressions of earlier work, weakened assertions, security, and test coverage. The standards reviewer looks at simplicity, duplication, dead code, over-engineering, naming, vocabulary, and performance shape, against a rubric it carries with it. Both reports come back verbatim under their own headings. They are never merged, never reranked, and never summarized into a single list - which means you can read the correctness report first, and read the standards report knowing it is not competing for the same slot.

**Severity has an owner.** Only the correctness reviewer can raise a Critical finding or say the change is ready to ship. The standards reviewer is capped at Should-Fix by construction: it cannot mark something Critical, and it cannot issue a ship verdict. That is deliberate, and it is the part that makes the split worth having - a duplication complaint and a broken requirement should not be able to arrive wearing the same badge. If the standards reviewer does spot something it believes is outage-grade, it is not silenced: it hands the suspicion across as a short untiered note, outside its own axis, for the correctness reviewer’s judgment rather than as a verdict of its own.

**Neither reviewer can flood the fix loop.** Each axis has a hard limit on how many findings it may report, and when it has more than that it says so rather than padding the list. The point is that a review with thirty style observations does not turn into thirty rounds of fixing - you get the ones that reviewer judged most worth your time, plus an honest note that there were more. Merge judgment stays yours in all of it: the audit reports, it does not gate.

**Honest bounds.** The two reviewers are the same auditor given different instructions, not different models with different strengths - the gain here is undivided attention, not a second opinion from a second mind. And the standards axis is genuinely narrower than a generalist reviewer was: it will not tell you a requirement is missing, because that is not its job anymore.

**Under the hood.** The audit dispatches two axis-scoped runs in parallel; a dispatch that arrives without its axis line defaults visibly to correctness rather than guessing. Caps are eight tiered findings for correctness, five for standards, three Considers each, with any overflow declared in the report. Cross-axis handoffs from standards are limited to two untiered lines. Changes to reviewer behavior are now gated on replaying a banked corpus of reviewer regressions - two historical catches have to survive the change, and both did.

### 3.20.0: Declined scope stays declined

**Three things kept costing you the same argument twice. A feature you had already refused on principle came back weeks later as a fresh proposal, because nothing remembered the refusal. Genuinely-open questions got written into specs as though they had been settled, so an unknown read as a decision. And the prose steering your agents leaned on capital letters where it should have stated a rule the agent could check its own output against. Now a refusal is recorded with its reasoning and every date it was re-asked, planning consults that record before it proposes scope, and only you can reopen a declined idea; unknowns get parked as unknowns until someone actually resolves them; and the shouting is replaced by rules that name what a failure looks like, plus about 60 new completion bounds that tell a procedure step when it is genuinely done.**

Detail

**Nothing to do.** No config, no defaults, no commands changed. Update the plugin and the same commands behave the same way, with less scope drift and less half-specified content in what they hand you.

**Your “no” now has a memory.** The first time a feature is refused on policy grounds - not “we already have that”, but a real judgment call about what this product is - the refusal gets written down: what was declined, why, and a dated list of every time it has been asked for since. Planning reads that record before it proposes scope, so the idea you turned down in March does not arrive in June wearing a new name and consume the same conversation. Two boundaries keep it honest. Only you reopen a declined concept - an agent may surface that a decline is being re-requested, and it may not decide the answer has changed. And “declined because it already exists” never earns a file, because the record exists to hold judgment, not history. The recurrence list is the useful part: when the same request shows up for the fourth time, you can see that, and decide with the evidence in front of you.

**Unknowns stop impersonating decisions.** Specs used to have nowhere to put a genuine unknown, so unknowns got written up as content - half-specified, confidently phrased, and indistinguishable from a decision someone actually made. A spec can now park them in a section of their own, gated by a test that is easy to apply: if it can be decided now, decide it; if it is known work that just has not been scheduled, make it a task; only what is truly fogged gets parked. Parked items graduate into real sections the moment interview or planning resolves them, so the section empties as the spec matures. What you gain is the ability to read a spec and tell the difference between what was settled and what nobody knows yet.

**Specs describe contracts, not file paths.** A spec that named files and line numbers was accurate for exactly as long as it took someone to refactor, and then it generated churn: plan-sync chasing renames, reviewers reconciling paths that had moved, and a spec that read as wrong when the behavior it described was still exactly right. Specs now state types, signatures, and behaviors - the things that survive a move. One deliberate exception stays for decision-rich snippets where the location genuinely is the decision. Tasks are untouched and stay path-bearing: naming the files you are about to modify is a task’s job, and always was.

**Rules an agent can check itself against.** Hundreds of CRITICAL / MUST / FORBIDDEN blocks across the skills were restated as plain declaratives that describe the failure rather than raise the volume - “a SHIP verdict with no backend response behind it has broken this” tells an agent what to look for in its own output in a way that a capitalized MUST does not. The conversion was meaning-preserving line by line, and what stayed capitalized stayed for a named reason: literals that tests pin, fences that get executed, anchors that host mirrors transform, and blocks that evaluations guard. Alongside it, roughly 60 procedure steps across qa, map, setup, drive, work, pilot, land, prospect, and plan gained a `Done when:` bound that makes a demand - every item accounted for, every finding filed - rather than merely describing what finishing looks like. The practical effect is fewer steps an agent can consider complete while something is still outstanding.

**The repo’s own glossary became a dictionary.** What had grown into an encyclopedia is now a short list of the terms whose synonyms cause real ambiguity - twelve of them, each with one definition and the aliases to stop using (a spec, never an epic or a ticket or a story; plan-sync, never tracker-sync). The long-form text is archived rather than deleted. This is flow-next’s contributor vocabulary only; the glossary feature that `/flow-next:setup` seeds into your project is unchanged.

**Routing stops going stale silently.** The guide that recommends which workflow to run is now held to a rule: recommending a skill that no longer exists, or missing one that does, is a defect, and any change that adds or removes a skill has to update the router in the same breath. The rule found its first gap on the day it landed - arriving with no written direction at all now routes you to strategy instead of into a plan. Install instructions are also single-sourced from one place now, so the copy you read cannot disagree with itself depending on where you found it, and the judgment-boundary prose in plan, chart, and guide names the specific pull toward just doing the work yourself as the signal that you are standing on one.

**Fixes worth knowing.** The Codex host mirror’s plan rewrite had been quietly doing nothing since an earlier rework moved the prose its anchors targeted - Codex hosts were getting Claude-specific wording where multi-agent phrasing belonged. The anchors are repaired. A report skeleton in memory-migrate had drifted from its canonical copy and was missing two sections; the duplicate is now a pointer to the one source. Two reference files reachable from nothing at all are deleted.

**Under the hood.** Every conversion was checked against the test corpus and the host-mirror transform before the edit, and test pins were retargeted in the same commits rather than loosened. Prose-contract tests, per-skill conduct checklists, and the full suite gated each wave of the rewrite.

### 3.19.0: Skills read only the path you take

**Every skill you ran used to hand your agent its whole instruction set up front, including the rules for branches that session was never going to take - and a model reading four sets of conditions it does not need is exactly where instruction-following quietly erodes in long prose. Ten skills now load a lean spine and pull a branch’s instructions at the moment they reach it: up to 63% less always-loaded prose where the modes are genuinely exclusive, 15-40% across the heavy skills, with every safety net and every-run contract still inline. You get cheaper sessions and better adherence on the long, condition-heavy skills - not faster runs.**

Detail

**Nothing to do.** No config, no defaults, no commands changed. Update the plugin and the skills you already run behave the same, from a smaller starting load.

**What changed for your session.** Skills grew many-branched over dozens of releases of features people asked for, and the disclosure discipline flow-next started with did not grow with them. A chart run that is going to take one mode was reading all of them; an implementation review pinned to one backend was reading the prose for the others. Now what every invocation needs stays inline, and what only some paths reach lives in a reference the skill reads when it gets there. Where a config decides the branch, the probe that reads it is fail-open: if the probe cannot answer, you get the branch rather than silence.

**What deliberately did not move.** Every safety net, every every-run contract, and every calibration block stays inline where the agent always sees it - a skill that only sometimes remembers its guardrails would be a worse trade than any prose it saved. The refactor moved text verbatim rather than rewording it, and each of the ten skills was checked against a written list of observable behaviors before it shipped.

**The discipline now has a keeper.** Every skill carries a **conduct checklist** - four to six falsifiable observables that prose changes are reviewed and dogfooded against. That is the part meant to outlast this release: the reason the disclosure rotted the first time was that nothing failed when it did.

**Two internal surfaces are gone.** The rp-explorer exploration skill was dispatched by nothing, and planning always used `repo-scout` in practice - so `context-scout` is removed too, and with it an entire research-mode question branch from planning. Planning now has one codebase-research path instead of a fork you never chose consciously. RepoPrompt is unaffected as a review backend and remains fully supported everywhere it was before.

**Fixes worth knowing.** Interview’s doc-aware autodetect was fail-closed, so a probe failure silently switched doc-aware behavior off - it now fails open like every other gate. Make-pr’s inline mermaid recap had drifted to eight rules while the canonical checklist carries nine, including the subgraph/node-id collision rule that a real PR caught; the duplicate recap is gone. Codex hosts now receive the host-native review invocation on work’s reference paths too, not just the main phase file, so the wave-join and host-deferred paths stopped emitting a Claude-only slash command.

**Running lean is now a documented choice.** flow-next has always run fully as spec then plan then work, with everything else optional, but nothing said what turning a layer on actually costs you. The new [Running Lean](https://flow-next.dev/understand/what-each-layer-costs/) page frames two operating profiles - human-driven, where you are present and can be the reviewer, the tracker, and the QA; and autonomous, where those same layers are what replace you - and prices every optional layer with the same four fields: what it automates away, what it costs structurally, when it earns its keep, and the manual invocation if you want the capability without the standing cost. A layer you skip still leaves a record: stage receipts carry `skipped(reason)`, readable with `flowctl usage --stages <spec-id>`.

**Deprecated, not removed: Ralph and packaged codex delegation.** Nothing is removed, no defaults change, and existing installs keep working - these are signals so new setups stop adopting a path intended for retirement. A host loop or `cron` calling `/flow-next:pilot` and `/flow-next:land` does Ralph’s job without the scaffold and the guard-hook registration; the setup model-routing scaffold plus the bridge recipes in `.flow/usage.md` cover what packaged codex delegation was built for. Both reference pages stay maintained for current users.

**Under the hood.** Prose-contract tests now pin content and reachability rather than file location, so a verbatim move to a reachable reference no longer breaks the suite while content disappearing or becoming unreachable still does. A new encoding guard keeps every reference file and its Codex mirror twin clean UTF-8.

### 3.18.0: The pipeline stops overbuilding your requests

**Ask an autonomous planner for a small feature and it builds you a subsystem: in a replay campaign against real shipped work, two unguided runs of the same request each invented a 500-900-line risk-management layer nobody asked for. This release lands the disciplines that campaign measured, as one batch. Plans now bind to scope minimality - every task traces to a requirement, every requirement to your request, and overengineering is a review finding rather than a taste note; the guided replay arm delivered 43% fewer output tokens at 57% lower cost with reviewed quality above the unguided one. Tasks become lean delegation payloads that reference the spec instead of restating it. A pipeline stage that silently does nothing - the class behind the plan-sync bug that no-oped for weeks - now announces itself in the receipts it already writes. And same-spec worker concurrency runs on an explicit fail-closed rule instead of a judgment call that almost never fired.**

Detail

**Plans stop growing past the request.** Every task must trace to a requirement and every requirement to what you actually asked for; capabilities nobody requested become one-line Boundaries exclusions instead of tasks, and planners are steered to eliminate risks structurally (a closed schema, an inert format, an unexposed capability) before building machinery to manage them. The discipline trims scope, never rigor: error-case enumeration and filesystem/permission/concurrency guards are explicitly exempt - an eliminated guard is not an eliminated feature. Both review rubric copies treat overengineering as a finding with three concrete patterns to flag.

**Tasks carry the how, not a retelling of the why.** Replay agents wrote tasks at roughly three times the fleet norm, and the bloat was paraphrased spec context that drifts out of date. Tasks now reference the spec’s requirement IDs and carry the concrete implementation plan - named files, approach, ordering, task-scoped acceptance - which is exactly what lets a cheaper implementer build without re-deriving design decisions. Nothing is lost: executors always receive the task together with the full parent spec. Tasks can also declare the paths they expect to modify, which feeds the new concurrency rule.

**Planning documents are files, not heredocs.** A plan that goes through review fix loops is edited in place with span edits instead of being regenerated wholesale into the command string - measured 13% cheaper in both replay A/B pairs. Spec examples are now the contract: the fields an example shows are exhaustive, closing a deviation class caught twice where an implementer “helpfully” extended a shown shape. Workers run focused tests while iterating and the full suite only where a gate already requires it - the campaign measured 54% of full-suite runs as redundant mid-loop re-runs.

**No stage can silently do nothing.** Every optional or delegated stage records ran, skipped with a reason, or failed with a reason in the receipts it already writes - a skipped stage is an event, never an absence, and a stage with no line is treated by review as failed. `flowctl usage --stages <spec>` summarizes them per spec, plain or JSON; malformed lines are counted, never a crash. No new stores, no dashboards, and token telemetry is explicitly out of scope.

**Concurrency by rule, not vibes.** 85% of 684 measured worker dispatches ran with zero overlapping sibling because the wave trigger was a judgment call. Now tasks run concurrently only when their declared write-surfaces are disjoint, no dependency path connects them, the wave is at most three, and nothing touches the always-serial set - anything missing or doubtful stays serial, exactly as today. A join conflict is never auto-resolved: the losing task re-runs serially and the collision lands in the receipt. Verified by a sequential-equivalence replay: wave and serial runs produced identical test outcomes and identical trees.

### 3.17.0: Say it was wrong, and know when it drifted

**Two kinds of state you could not correct are now correctable. When discovery disproves the fact a chart started from, the decision that disproves it can carry the correction - and that correction travels into the briefing your next spec is written from, instead of the refuted claim sitting above the ledger that contradicts it. Separately, the flow-next-managed block in your instruction files can now be one of several in a single file, and a new read-only command tells you (or CI) whether anyone has hand-edited one, without writing anything to find out. Both shipped from field reports by @sn-furali (#292, #294).**

Detail

**Do this first if you wrote your own setup-block template.** A template must now be exactly its marker-pair block: the BEGIN marker on the first line, the END marker on the last, a trailing newline, and no marker token repeated inside the body. Every template flow-next ships already conforms, so nothing to do for the standard install - but a hand-rolled template with a heading or a note outside the markers now fails with a clear message instead of drifting. It has to: prose outside the pair was written and hashed on apply but never seen by the comparison, so a block you had just applied could report itself as edited.

**Charts can now admit they were wrong.** A chart seeds `## Notes` with the grounding facts it starts from, and discovery routinely disproves one of them - that is what discovery is for. Until now there was nowhere to record that. The note was write-once, so a refuted claim survived into the immutable briefing that `/flow-next:capture` reads, sitting above the ledger entry that contradicts it. Now the resolve that closes the decision can carry the correction with it: a dated bullet is appended to the notes, the original text is left byte-for-byte alone, and the next briefing renders both. Nothing becomes mutable - corrections are append-only and stamped by the tool, so the chart still reads as a record of what was believed and when.

The same report caught a quieter failure. A sharpen file with a key the tool did not recognize - including the `notes` key you would naturally reach for - used to be accepted and silently ignored, so a correction you thought you had recorded simply was not. Any unrecognized key now fails the whole resolve, names what it did not recognize and what it accepts, and does so before anything is written. A typo can no longer look like success.

**Managed blocks are addressable, and drift is checkable.** flow-next owns a marker-delimited block inside files you also edit, and until now it could track exactly one such block per file: pointing it at a second block in the same file overwrote the first one’s recorded state, so one of them lost its ability to tell “pristine” from “you edited this”. Blocks now carry an id, each with its own markers and its own recorded state, and operating on one never touches another’s - including a stray or corrupt one sitting in the same file. Omit the id and everything behaves exactly as before.

The other half is a read-only verdict. If you keep the block as an ordinary tracked file (`setup_mode: copy`), an edit to it is a normal reviewable diff - but there was no way to ask “is this still what flow-next generated?” without running the write path. The new check answers exactly that and writes nothing on any branch: clean exits zero, drift exits two, a structurally broken block exits three, so a CI job can gate on it with no jq and no risk. A block someone edited and then reverted reads clean, and a CRLF-only difference is never drift.

**Under the hood.** Recorded state moved from one hash per file to one per (file, block id); a hash written by an older version is read transparently as the default block’s state and upgraded on the next write, with no migration step to run. Review rounds on the way in tightened the fail-closed contract in three more places: keeping a customized block used to record that decision without ever validating the block, a duplicate or orphaned marker after the first valid pair escaped the corruption scan, and the check released its lock between reading state and reading the file, so a concurrent apply could skew its verdict.

### 3.16.3: The split path, hardened by running it

**A live end-to-end run of 3.16.2’s spec-split found what section testing could not: leftover copy that made “split” readable as “abort”, split specs authored after approval instead of shown before it, and a write step that silently dropped spec titles. All fixed - you now see every composed spec document before anything is written, and every copy of the workflow agrees on what the split answer does.**

### 3.16.2: Capture tells you how many specs it should be

**If you capture epics, briefing packages, or large features, the hardest question was never the content - it was “is this one spec or four?” Capture now answers it. Past 8 real requirements (standing rules and “tests must pass” items don’t count), or when the requirements clearly serve more than one shippable outcome, the read-back shows the actual split: proposed titles, which requirements go where, and how the specs depend on each other. One answer writes the whole linked set. The judgment is independence, not size - a big-but-cohesive spec is recommended to stay one spec, small captures see nothing new, and nothing ever splits without your say-so. Interview makes the same call when refinement outgrows a spec.**

### 3.16.1: Plan-sync actually runs

**If you switched on plan-sync so completed work updates the tasks that come after it, that update was silently never happening. The work loop misread the task list’s JSON shape, the error went where nobody looks, and the empty result was indistinguishable from “nothing downstream to update” - so every run looked clean. Fixed, and a failed extraction now announces itself instead of impersonating an empty list. Thanks to the field report that caught it, real drift surfaced on the very first plan-sync run after the fix.**

### 3.16.0: Your reviewer reads the whole change

**Cross-model review used to be handed a copy of your diff inside the prompt, capped at 50 KB. On a large change that meant the reviewer judged your code having been shown about a tenth of it, then went and read the rest off disk anyway. The copy is gone. A review now gets the commit range, the exact list of changed files, and the paths to the spec and tasks, and reads whatever it needs from your checkout. Nothing is trimmed to fit, and a review that cannot read its evidence now stops instead of returning a verdict based on nothing.**

Detail

Set expectations honestly first: this makes reviews **better informed, not cheaper**. The prompt itself shrank by 83% on the release’s own largest review, but a reviewer that fetches spends turns on tool calls instead, and measured input tokens came out *above* the previous numbers. Most of that is cached, so billed cost does not track the raw figure, but no saving is claimed. A fetching reviewer is also slower in wall-clock on a big diff, which is why the dispatch bound moved from 600 to 1800 seconds - override with `FLOW_REVIEW_EXEC_TIMEOUT` if your changes are larger still.

What you get for that is a reviewer working from complete evidence rather than a truncated sample, and a changed-file list you can trust. That list is load-bearing now, and git abbreviates in three separate ways that all had to be switched off: `--stat` elides long paths behind an ellipsis, plain `--numstat` collapses a rename into `{old => new}` so neither real path appears, and without `-z` any non-ASCII filename comes back escaped. A scope map you cannot resolve to real paths cannot bound a review.

Re-reviews changed too. The reviewer now continues its own session instead of being re-briefed from scratch, so when it checks whether your fixes landed it is comparing against findings it actually remembers making. If a session cannot be resumed, the findings travel in the prompt exactly as before - the fallback is deliberate and loud, because a reviewer handed a lean prompt with no memory would silently produce a fresh blind review.

Convergence got the same treatment. Loops that were converging no longer get cut off and handed to you to verify by hand: the re-review prompt states the exact format for reporting prior findings, the parser accepts every token that format advertises, and the two rules that guessed at convergence from finding counts and severity trends are removed - they escalated three healthy loops in a row and caught no stuck ones. What remains is the reviewer explicitly calling the same finding unfixed twice, plus a round cap you own via `review.maxIterations`.

**Under the hood.** Removing the payload also removed everything built to make it fit: three prompt fitters, the 50 KB diff cap, the reviewer-facing “truncated to fit” markers, and the interim guard that distrusted an all-clear from a backend whose prompt could be shortened. That last one is retired by construction rather than gated, because no backend truncates now. One size guard survives, renamed to say what it is - a transport boundary for the one backend that delivers its prompt as a command-line argument, which refuses loudly instead of trimming.

This decision was made once before, in 3.2.0, and undone twice by later work that each had a good local reason. So it is enforced by a test that drives the real dispatch path and was verified to fail when a re-embed is simulated, plus pinned builder signatures so a new payload parameter cannot slip in under a different name.

### 3.15.1: Windows runs everything now

**If you work on Windows - or merge work from someone who does - a green build now means the same thing on all three OSes. The Windows CI leg used to skip six “incompatible” test files; a change could pass Linux and macOS while quietly breaking behavior hidden behind that filter. The filter is gone, and the bugs it was hiding are fixed.**

Detail

Every one of the six excluded files turned out to be hiding a real defect, and none matched its recorded excuse. The failures were locale-dependent text reads (Windows writes cp1252 where production expects UTF-8), a POSIX-only permission check that crashed at import time on NT, and backend subprocesses that could block forever waiting on an inherited stdin no automated caller answers.

The infamous “900-second hang” was never the backend it was blamed on: the test runner killed only its direct child on timeout, then waited forever on output pipes held by grandchildren - on every platform. The runner now kills whole process trees (process groups on POSIX, Job Objects on Windows), reports leaked descendants instead of hiding them, and bounds its timeout diagnostics.

Concurrent tracker writes also stopped intermittently refusing legitimate paths on Windows - under load, path resolution can return two spellings of the same directory, which looked like an escape attempt. Containment is now derived from a single resolve, and the review cycle hardened it further: NTFS junctions (which are not symlinks and evaded the symlink check) are now rejected fail-closed during the containment walk, alongside pointer-width handle declarations for the Windows kill path.

Nothing was weakened to get there: no raised timeouts, no blanket platform skips, no broadened assertions. The full corpus runs green on `windows-latest` in parallel, serial, and shuffled order on the same commit as the Linux and macOS gates.

### 3.15.0: Flow on every task, without the tax

**Running flow-next on small tasks used to cost more ceremony than the task: \~20 CLI calls to author a spec with tasks, a pile of file reads every time a fresh session re-oriented, and the one quality miss that kept repeating - the untested error path. This release cuts authoring to 2 calls, re-anchoring to 1, and moves error-case thinking to plan time where it is cheap.**

Detail

Benchmark evidence made the overhead concrete: on identical work items, the flow-next pipeline spent 2.4x the wall-clock of a no-flow control, and the biggest blocks were ceremony calls and re-read context - fixed costs that dominate exactly the small tasks you most want tracked. The losses that were not overhead traced to a single pattern: an error path nobody enumerated, which a green test suite then certified forever.

Authoring is now a fast path. `spec create --plan-file plan.md` creates the spec with its plan in one call, and `task create --from-json tasks.json` materializes the whole task set in another - descriptions, acceptance, satisfies, dependencies (tasks in the same batch can reference each other by position). Validation is all-or-nothing: one invalid item rejects the whole batch with zero writes, so a half-created plan cannot exist. The granular verbs are unchanged and remain how you edit. The canonical spec-plus-3-tasks flow now measures 8 calls, down from \~20, and a test counts the real subprocess invocations to keep it honest.

A fresh session runs `flowctl brief` and is oriented: open specs with one-line goals, which tasks are actually ready (dependency-aware, with claim state), the last five completions with an evidence flag, the memory index, and pointers for going deeper. The output is deterministic and capped at roughly 2k tokens regardless of repo size - what gets dropped is marked, `--full` lifts the cap, `--json` is the machine form. Your context window stops paying for spec bodies you did not need.

New specs now carry their error cases in the acceptance criteria themselves - each criterion states its invalid-input and boundary handling, or says “no error surface beyond X” outright, so a reviewer can tell considered-and-none from forgot. The plan skill derives these during AC writing, the interview probes when they are missing, and workers treat every enumerated case as a required test before done. Existing specs are untouched.

Under the hood: brief performs no git calls and no writes (identical state renders identical bytes), the bulk-create path holds one lock per batch, and receipts, evidence, and start/done validation are byte-for-byte unchanged.

### 3.14.0: Review loops that end the way a human lead would end them

**Converging review work gets its room, a loop that has stopped improving reaches you early instead of burning its whole budget, and a reviewer facing a genuine judgment call can hand it to you directly - evidence trail intact. This is the mechanism the 3.13.3 cap raise promised.**

Detail

The review cap counts dispatches, and a dispatch counter cannot tell “genuinely stuck” from “nearly there.” Now the loop also reads the structured findings each verdict already persists and ends a measurably stuck loop early: an open finding chain two rounds fail to resolve, severity and count both failing to improve, or each fix introducing a fresh P0/P1 - exit `4` with `ESCALATE: review loop stalled (<rule>)`, the same exit code and marker family drivers already handle.

Reviewers get a new terminal verdict, `NEEDS_HUMAN` - “a human must adjudicate,” distinct from `NEEDS_WORK` (fixable) and `MAJOR_RETHINK` (redesign). It writes a real `needs_human` status and its receipt before escalating, so nothing stops silently.

Repeat reviews stop wasting rounds on unchanged content: dispatching the same artifact that just received a verdict is refused before it costs anything (`NOT_RETRYABLE`, exit `1`). The identity is per-surface - a completion re-review after an implementation-only fix dispatches cleanly - and `SHIP` or an explicit human re-plan starts a fresh epoch.

Each review surface now blocks on what it can actually break: plan review blocks only on findings naming a concrete bad downstream outcome, impl review treats recorded Decision Context decisions as settled, and the land-loop PR bot is scoped and triaged as a safety net, never a second gate. Autonomous loops cannot grant themselves more rounds - every new terminal is shorten-only, and Ralph blocks the reset commands and `--force` as human-only recovery.

### 3.13.3: Codex, in every home you use

**If you run Codex from more than one home - a work account, a client sandbox, a second instance - flow-next could only ever live in one of them. Now it installs into whichever home you point it at, and each install stays self-contained.**

Point `CODEX_HOME` at the home you want and run the installer once for it.

### 3.13.2: Three silent chart defects

**Three ways a chart could quietly hold the wrong state: a reopened chart with no route back to capture, a supersession that wired a replacement to the premise it had just superseded, and an ambiguous initial map that pointed edges at the wrong decision. None raised, none logged - each one persisted a chart that looked correct and answered every later question from the wrong state.**

### 3.13.1: Chart refuses a direction

**“Make our CLI more deterministic” reads like exactly the big unclear idea chart was built for, and chart would have taken it. It has no finish line, so the map could never close. Chart now names the test it was always applying and says no before spending a discovery pass on it.**

### 3.13.0: Chart: discovery before the spec

**One oversized idea wrapped in unknowns no longer has to become a half-guessed spec or a meeting that evaporates. Optional chart finds the route one decision at a time, then hands capture a briefing - and the short first-run path stays short.**

Detail

Teams kept hitting the same gap: prospect ranked ideas, capture wanted intent you could write down, and between them large unclear efforts either got captured too early (a wall of inferred criteria) or lived only in conversations. Chart is the optional pre-capture route for **that** situation - not a new mandatory stage.

You describe the outcome in plain language. The agent grounds a bounded snapshot against the repo (safe citations only; nothing invents a resolved decision), reads back the smallest visible frontier and attended/unattended cost, and only then persists a chart. Work resolves **one decision per session** via evidence-first routes - research, probe, eval, prototype, interview, or an enabling task. Attended routes never self-answer under autonomous drivers (`NEEDS_HUMAN`). Prototypes attach a throwaway artefact before the human reacts; wrong turns stay struck-through via supersession so the briefing keeps the path that failed.

When nothing material remains to decide, a confirmed briefing package (one or more clusters, shared context named once) hands off to capture. Chart never writes a spec and never sits inside pilot. Guide recommends the smallest sufficient next step when you are unsure; tracker projection and pasted URL re-entry are optional conveniences over the local ledger, never the source of truth.

### 3.12.0: Your editor understands .flow/config.json now

**flow-next’s config file carries a published JSON Schema: editors validate and autocomplete every setting, scaffolded configs reference it automatically, and drift between the schema and what flowctl actually reads is a failing test, not a documentation promise.**

The schema lives in the repo and at the stable URL this site now serves: <https://flow-next.dev/schema/flow-config.schema.json>

### 3.11.0: Trust and identity fixes for the autonomous PR path

**Two community-reported gaps closed: land can no longer mistake a PR that merely talks about flow-next for one it authored, and repos that require bot-authored PRs get a documented seam for supplying their own PR-create identity.**

### 3.10.0: Standing team rules, checked on every spec

**Write your project-wide acceptance criteria down once - “every route change regenerates the contract”, “no new dependency without a health check” - and the completion review you already run judges every spec against them, with the verdicts recorded in the receipt.**

Detail

Every team has standing rules that outlive any single feature. Until now they lived in instruction files and reviewer memory: applied when someone remembered, invisible when they were not.

Flow-Next 3.10.0 gives them a home. Put one bullet per rule in `.flow/criteria.md` using the familiar requirement grammar (`- **G1:** ...`), and every spec completion review - on every review backend - judges each criterion against the whole implementation. The verdicts (`met`, `violated`, `not applicable`) land in the ordinary review receipt, so compliance is a recorded fact rather than a feeling. Violations also appear as normal review findings, so nothing new needs watching.

There is no separate audit pass, no rule engine, and no scoring - the reviewer that already reads your diff simply gets your standing rules alongside the spec. Repos without a criteria file pay nothing: not a token of prompt content changes until the file exists. Setup offers to scaffold a documented template on request, and declining leaves no trace.

The input boundary fails closed. A criteria file that exists but is broken - typo’d bullets in any Markdown style, duplicate or malformed ids, an unreadable file, a dangling symlink - surfaces a validation error before a review round is spent, instead of silently running the review without your rules. On the receipt side, ambiguous or contradictory reviewer output degrades the compliance array to absent rather than ever recording a wrong verdict, and the recorded ids must exactly match your configured criteria before anything attaches.

New plumbing, for the curious: `flowctl criteria list --json` validates the file; `flowctl criteria prompt-block` composes the injection for the RepoPrompt and host review paths; receipts gain an additive `criteria: [{id, status, note?}]` array documented beside the structured findings schema.

Full model: [Standing Criteria](https://flow-next.dev/reference/standing-criteria/).

### 2026-07-31: Front doors that lead with the problem (docs release, no version bump)

**The landing page and the README now open on the problem they solve and show the measured evidence for it, so you can judge Flow-Next in one screen instead of reading two thirds of a manual first.**

A new [Evidence](https://flow-next.dev/project/evidence/) page carries the full argument.

### 3.9.0: Pull requests that guide the review

**Reviewers get a guided journey through the change: what changed, why each step exists, what deliberately stayed untouched, where the risk sits, and which evidence supports each claim.**

Detail

A large pull request usually arrives in file order. The reviewer has to rebuild the story: find the important decisions, work out which files belong together, separate meaningful changes from generated churn, and decide what still needs human judgment.

Flow-Next 3.9.0 does that preparation before handover. The pull request now walks through the change in logical steps. Each step explains its purpose, groups the files that implement it, links back to the relevant requirement or task, and names deliberate non-changes that protect the boundary of the work. Generated and mechanical files stay with the step they support instead of forming a distracting pile at the end.

The existing risk-ranked review plan remains independent from that journey. The walkthrough explains how the change fits together; the review plan points the human reviewer at the decisions and code that deserve attention. Evidence is attached to the claims it supports, so a reviewer can distinguish what the pipeline already proved from what still calls for experience and judgment.

Review findings now keep their identity across fix rounds. A finding can be traced to the review that raised it, its current status, and an optional snapshot-bound code location. Resolved, superseded, and current findings no longer blur together when the branch moves.

The same review journey can appear in the GitHub pull request and the optional local HTML view. Structured JSON carries the same meaning for downstream tools, including Flow Swarm, without forcing them to scrape prose or guess the order of the change.

Under the hood, `/flow-next:make-pr` stores this journey as a versioned `changeWalkthrough`, while review receipts can carry versioned structured `findings`. Existing receipts still work. Invalid or stale structured data falls back to the original reviewer prose, and parsing adds no model or network call.

RepoPrompt CE now supplies its review context and response through one direct Context Builder result. Fix rounds stay attached to that returned context and chat, with no hidden setup conversation or dependency on a Classic-style tab. Parser and render benchmarks enforce a strict `<100 ms p95` ceiling over 30 warm runs.

### 3.8.0: Interview-written criteria now say where they came from

**You can finally tell which acceptance criteria came from a human and which the agent guessed, on specs that came out of an interview rather than a capture.**

Interview now emits the same four tags everywhere it writes criteria.

### 3.7.0: Your own spec sections get filled in

**Add a section to your project’s spec template and the interview now writes it for you, instead of leaving your heading empty or quietly ignoring it.**

Detail

Do this first if you already keep a repo-root `SPEC.md`: put a scope marker under any section you added, e.g. `<!-- scope: business -->`. That one line is the difference between a section the interview fills and a section it leaves alone.

A project has always been able to override the spec scaffold with a repo-root `SPEC.md` - add a risk register, user stories, a rollout runbook, whatever your post-mortems justify. What was missing is that the interview passes only knew about the seven sections we ship, so your own headings sat outside the contract: nothing promised to preserve them, and nothing would ever fill them.

Ownership now comes from the section itself:

* marker naming the pass you are running - **written and refined**, like any section we ship
* marker naming the other pass - preserved exactly as-is
* `<!-- scope: both -->` - written by either pass
* no marker - preserved exactly as-is, and the read-back tells you it was skipped

One consequence worth knowing: a marked section is rewritable, so hand-written content under a marker you own will be refined by the next pass of that scope. Drop the marker to freeze it.

Also new in this release: the customization route itself is documented properly for the first time, including which four headings you must not rename (`Acceptance Criteria`, `Boundaries`, `Goal & Context`, `Decision Context` - renaming them does not error, it silently drops the feature that reads them). See [writing specs](https://flow-next.dev/guides/spec-scaffold/#adding-your-own-sections).

Honest bound: we tried widening the *default* template with user-story and test-seam sections and did not ship it. The first measurement looked good, a pre-registered replication did not hold, and the wider scaffold ran about a third longer - which every implementer and reviewer downstream pays to read. Section preferences are project-specific, so the override is the right place for them rather than the default.

### 3.6.1: Tracker conflicts follow your chosen policy

**When Flow and your tracker disagree about status, the policy you chose now decides the outcome. Sync no longer stalls because the setting was documented but ignored.**

### 3.6.0: Tracker sync stops improvising

**Tracker updates now behave consistently across GitHub, GitLab, Jira, and Linear, including retries and partial failures, so sync no longer depends on which provider or agent happens to run it.**

Detail

* **No migration step.** Existing tracker configuration and lifecycle settings continue to work. The change is inside the execution boundary: skills decide what a body or comment means, then make one `flowctl tracker sync` call.
* Provider requests, pagination, create-first recovery, status policy, dependency links, comment deduplication, and receipts now run through one tested implementation. A retry can prove what already landed instead of asking the agent to reconstruct provider state from prose.
* Backlog autonomy now uses the same executable layer. Dependency ordering reads normalized directed edges, and parked questions carry stable identities so retries do not post duplicates.
* The failure boundary is explicit. Authentication, rate limits, stale ids, capability gaps, conflicts, and transport failures return structured classes. The host still decides whether to ask, defer, continue through MCP, or correct local input; provider mechanics no longer consume its judgment budget.
* The old transport recipes and tracker-runner agent are gone. Lifecycle callers retain their silent inactive gate, so repositories without tracker sync do not pay a new process or output cost.

### 3.5.2: Superseded by 3.6.0

**The tracker-sync batch was initially published as 3.5.2, then republished unchanged as 3.6.0 the same day because three substantial specs belong in a minor release, not a patch.**

The 3.5.2 artifact remains available as an accurate historical record. Upgrade to 3.6.0; there is no runtime difference between the two versions beyond corrected version metadata.

### 3.5.1: The review error stops suggesting the wrong fix

**If your agent has been reporting that a review failed for “sandbox” reasons and needs retrying with wider permissions, it was reading our error message, not your repo. Reviewers are read-only on purpose, so a reviewer that hits the sandbox is a scoping bug - the message now says so instead of telling you to hand the reviewer write access.**

### 3.5.0: Two agents, one spec number, and the collision stops being your problem

**`fn-7` twice is not bad luck, it is arithmetic: spec ids were allocated by counting the files in your working tree, so two branches cut from the same base both saw the same number and both took it. Allocation now sees every worktree and every ref, and teams with a tracker can hand the job to the tracker instead - one setting, and new specs are keyed `WOR-17` or `gh-123` from the start.**

Detail

* **The collision was structural.** Allocation counted only the current working tree, so parallel spec creation was guaranteed to collide, not merely likely - and in an agent-heavy workflow parallel is the normal case. It now takes the maximum across the working tree, every registered worktree, and every ref, and it is monotonic: a retired number is never handed out again, because reusing it would resurrect an ambiguous reference in your commit history and release notes.
* **The number stays.** Dropping `fn-N` would have been a vocabulary migration across changelogs, commits, tags, tracker comments and notes, to fix a symptom. The full `fn-N-slug` was always the real identity; what actually broke was a `validate` error and ambiguity in prose.
* **Honest bound, stated in the spec itself:** sequential allocation without coordination is unsolvable in general. This shrinks the window; it does not close it. Two clones that have never fetched each other can still collide - which is what the tracker route is for.
* **Tracker-keyed ids are now a setting, not a flag you had to remember.** `flowctl config set tracker.specIds tracker`, or just answer the question setup asks once when a tracker is configured. A tracker is a real distributed allocator, and collaborative repos are exactly the ones that have one.
* **GitHub and GitLab are no longer second-class.** Linear and Jira ship a `KEY-N` that mints directly; `#123` and `group/project#456` do not, so they mint through a synthetic key derived from your configured tracker type - `gh-123-slug`, `gl-456-slug`. Unambiguous because a repo has exactly one configured tracker, and guarded so a minted id can never collide with a historical one.
* **New: create-first.** Every previous tracker operation needed a local spec first, so “make the issue, then key the spec from it” could not be expressed at all. It can now, with the failure case designed in rather than discovered later: if the issue is created and something downstream fails, a retry **links** to that issue instead of creating a second one.
* **Network cost is stated accurately, not flattered.** An earlier draft claimed tracker-first adds no network cost. That is false when your lifecycle events are off, which is the default - so the docs and the setup question now say plainly that choosing tracker-keyed ids makes spec creation contact your tracker immediately.
* **Security fix.** `flowctl` no longer writes through a symlink anywhere between `.flow` and the file being written. An untrusted checkout could previously redirect a write outside your workspace, or onto another managed file. A legitimately symlinked `.flow` directory still works.
* Incidental: `flowctl task set-title` updates the JSON title and the markdown heading together, so the two cannot drift apart.

### 3.4.5: Recurring lessons graduate into enforced gates

**When your agent keeps re-learning the same lesson every run, the lesson is in the wrong place. The memory audit now has a sixth outcome, Harden: a correct, recurring, mechanizable lesson gets proposed as a lint rule, a CI step, or a rule in your CLAUDE.md/AGENTS.md - and only after the gate is verified to actually fire does the memory entry retire into a pointer at it.**

### 3.4.4: Claude Opus 5 joins the routing menu

**Claude Opus 5 launched yesterday at near-frontier intelligence for half Fable 5’s price - the recommended routing now leads with it, and absorbing a new model generation took a table edit and two registry rungs, not a pipeline change.**

### 3.4.3: Faster skills, honest review retries

**The skills you run most now load only the instructions their current route needs, while a broken review transport no longer wastes the three-round correctness budget.**

### 3.4.2: Memory entries stop vanishing

**A memory title starting with a quote or `- `could write an entry the CLI itself could not read back - the file sat on disk while `memory list` and search silently pretended it never existed. Both halves are fixed: those titles now write valid frontmatter, and any unreadable entry is reported instead of hidden.**

Thanks to [@TechupBusiness](https://github.com/TechupBusiness) for the exceptional report ([#235](https://github.com/gmickel/flow-next/issues/235)).

### 3.4.1: Parallel work waves

**Independent tasks can now move together when your host can isolate them safely, without turning every plan into a hand-built scheduler.**

Detail

* Plans show dependency-ordered execution waves, so you can see which tasks are candidates to run together.
* Work evaluates the complete ready frontier and may dispatch a safe concurrent subset. The host chooses worker count, isolation, and integration from its live capabilities; when those conditions are uncertain, it explains why and continues sequentially.
* Concurrent workers return task-specific handovers. The conductor joins and integrates the whole wave before review, completion, tracker updates, and downstream plan sync.
* Atomic task claims prevent duplicate ownership only. They do not make a shared Git index or filesystem race-safe.
* Grok documentation is corrected from live 0.2.111 evidence: type `/flow-next:` to open the plugin command namespace, including plan and work. `/flow-next-` searches the separate hyphen-named skill surface, and argument hints appear after autocomplete selection.

### 3.4.0: Grok Build, detected

**If you run flow-next in xAI’s Grok Build, setup now recognizes it instead of mistaking it for Codex - so you get proper slash commands and Claude-format instructions, not Codex `$flow-next-` syntax written into the wrong file.**

### 3.3.3: Readiness asks the right spec

**Rewriting a draft no longer asks you to mark it ready just because some other spec is ready.**

### 3.3.2: Capture understands old compactions

**A conversation that was compacted earlier no longer blocks capture when the feature you are capturing is still fully visible.**

### 3.3.1: Clean command names

**Your flow-next slash commands showed up in the Claude Code menu with the plugin name stuttered three times (`/flow-next:flow-next:flow-next:qa`). They now read the way they always should have: `/flow-next:qa`.**

Detail

* **The fix.** The command shims lived in a subfolder named after the plugin and carried a legacy namespaced `name:` field, so a recent Claude Code namespacing change stacked the prefix three deep. The shims are now a flat set of files with bare command names, and Claude Code prepends the plugin prefix exactly once. Nothing about how you invoke a command changes; the menu just reads correctly.
* **Cursor and Codex stay in lockstep.** Every command keeps the `name` + `description` Cursor’s marketplace review requires, the Cursor manifest and installers point at the new flat layout, and the Codex prompt install is unchanged. A retired `epic-review` alias (superseded by `spec-completion-review` back in 2.0) is cleaned off older Codex installs on upgrade, non-destructively - it is moved aside, never deleted, and a same-named file of your own is left untouched.
* **Messaging tidy-up.** The plugin’s own store/marketplace descriptions now match the wording on this site, and the bundled component counts are current.

### 3.3.0: Cursor, first-class

**If your team runs flow-next in Cursor, the rough edges are gone: install it org-wide from a repo instead of a per-person script-and-restart, get cross-family review without reaching for an external CLI, and stop wondering whether the stale “autocomplete doesn’t list commands” warnings were still true.**

Detail

* **Install once, for the whole team.** A Cursor Teams/Enterprise admin imports the GitHub repo as a team marketplace (Default Off / On / Required, auto-refresh on push) - every engineer gets flow-next with no local install and no restart dance. The `install-cursor.sh` / `.ps1` scripts stay as the individual path. Admin runbook on the [platforms page](https://flow-next.dev/install/#cursor).
* **`review.backend host` - cross-family review from inside Cursor, no external CLI.** Review runs as a fresh-context subagent pinned to a model from a different family than the one that wrote the code (Cursor honors in-prompt slug pins). It fails closed rather than quietly reviewing your code with the same model that wrote it - if no cross-family pin is set it asks (interactive) or reports `NEEDS_HUMAN` (autonomous). Works on Claude Code and Codex too; the other backends are unchanged.
* **Setup understands it’s running in Cursor.** It detects a marketplace or local install correctly, leads the review menu with Host, and writes a model-routing block into `AGENTS.md` with live Cursor model slugs - a cheap one for read-only scouts, a cross-family one for review. Model tiering by alias (`haiku`/`sonnet`/`opus`) falls back to your session model on Cursor; the explicit slug pins are how you steer.
* **Read-only agents are actually read-only on Cursor.** Cursor ignores the tool blacklist other hosts use, so review and scout agents now carry Cursor’s native read-only flag - closing a gap where a “read-only” reviewer could still edit files.
* **Approval prompts you can read.** When capture or interview asks you to approve a draft, the full draft now prints as normal text first and the question stays short - no more multi-paragraph spec collapsed into an unreadable one-line prompt.
* **Docs match reality.** Slash autocomplete lists the commands (hyphenated form), plain-English requests trigger the right skill, and native structured questions work including multi-question batches. Ralph autonomous mode is still Claude-Code/Codex only - Cursor has the hooks, flow-next just doesn’t wire them there.

### 3.2.1: Complete validation diagnostics

**Automated validation now tells you exactly which native spec IDs collide, without forcing a second text-mode run to find the missing errors.**

Detail

`flowctl validate --all --json` now includes every counted native spec-ID collision in `root_errors`, in deterministic order, while keeping `total_errors` exactly aligned with the returned diagnostics. Existing text output and exit behavior are unchanged.

### 3.2.0: Fast plumbing, stronger guarantees

**flowctl now gets out of the agent’s way without trading away trust: common entry points start faster, large repositories scale linearly, and concurrent agents cannot silently lose tasks or reuse stale model choices.**

Detail

* **Upgrade first:** Flow-Next now requires Python 3.11 or newer. Every launcher rejects an older-but-working interpreter before loading the CLI and tells you how to select or install a supported one.
* Root help and `flowctl usage` are about 68-73% faster on the measured macOS baseline. `flowctl specs` is 27.5% faster and Prime classification 18.8% faster. The acceleration stays source-authoritative: no opaque or stale executable cache becomes a new source of truth.
* Portable cross-process locks now cover task creation, setup/runtime state, and model-cache updates on POSIX and Windows. Paired task JSON/Markdown publication is transactional, and explicit model pins never silently downgrade.
* Large-repository work now shares one task inventory and one reverse-dependency graph. Status/list read each eligible task once and spawn no subprocesses; Prime, cognitive-aid export, memory, pilot logging, and frontmatter parsing shed repeated scans and reads.
* RepoPrompt Community Edition is now the primary integration. Flow-Next prefers `rpce-cli`, retains discontinued Classic only as the final compatibility fallback, and understands CE’s current window, repository-root, and chat response shapes. Repeated review setup reuses the existing repository window instead of cloning the workspace. Thanks [@aidancurry](https://github.com/aidancurry) for the precise [#228](https://github.com/gmickel/flow-next/issues/228) report.
* Active docs, skills, smoke labels, and the Codex mirror now match the live post-3.1 command and payload surface; confirmed dead helpers and test-only production surfaces are gone.

### 3.1.2: Say it in prose, Codex finds the skill

**On Codex, “plan this feature”, “work on fn-12”, or “pilot this to completion” now resolves the matching flow-next skill by itself. Before, the model-facing skill catalog carried six internal helper skills and hid every user-facing verb, so prose-invoked runs depended on the model rediscovering skills from disk - or silently improvising without the skill contract. Setup also stopped assuming you know flow-next: every ceremony question now explains what it decides and links these docs.**

Detail

* Catalog policy un-inverted: all 22 user-facing skills are now implicit-invocable on Codex (its naming rule then requires their use when you name or clearly describe one); the 6 skill-dispatched internals (`drive`, `sync`, `export-context`, `rp-explorer`, `worktree-kit`, `deps`) are explicitly hidden - still invocable by name, out of the catalog budget.
* Catalog descriptions dieted to fit Codex’s shared skills context budget (min of 8,000 chars and 2% of the context window, shared with every other skill on your machine): 2.9k chars total for all 22, vs \~7.6k undieted. New sync guards hard-fail on a missing catalog policy or an oversized surfaced description.
* `/flow-next:setup` rewritten for newcomers: each question states its stakes in plain language (what a setup mode decides, what a review backend is, what plan-sync/memory/Ralph actually do) and links the relevant page here for the longer answer. Same options, same defaults, same config keys.
* Codex installer fix: the memory-track templates (`templates/memory/*.tpl`) now actually land in `~/.codex/templates/` instead of flowctl silently falling back to embedded defaults.
* Update: `git pull && ./scripts/install-codex.sh`.

### 3.1.1: The grok bridge, at full strength

**If you route bulk implementation to Grok 4.5, the recipe you copy now unlocks the whole CLI: it edits files headlessly like the codex and cursor bridges do - the old guidance sold it as print-only, and hid a flag trap (`-p` swallows the next flag as its prompt) that made the write mode look broken. Docs-only patch; copy the corrected invocation and you get a third editing delegate on its own quota.**

Detail

* Corrected form, flags first: `grok --permission-mode acceptEdits -m grok-4.5-high -p "<task>"` (write mode; `--always-approve` is the blanket variant). The old `grok -p --always-approve "..."` shape misparses - live-verified.
* Extras now documented: `--check` (self-verify loop), `--best-of-n N` (parallel attempts, best picked), `--json-schema` (structured output).
* Same discipline as every bridge: run inside a trusted git dir; the host reviews and commits - Grok stays routed to bulk implementation, never final taste-critical work.
* Updated in the usage guide (`flowctl usage` serves it live in plugin-mode repos) and the model-routing scaffold’s grok route.

### 3.1.0: Set up once, never again (Claude Code)

**“Re-run setup in every project after every update” stops being a rule you have to remember. On Claude Code, setup now asks one question per repo - and if the repo is Claude-Code-only, it copies nothing at all: `flowctl` is simply on your agent’s PATH, the guide is one `flowctl usage` away, and plugin updates land silently. The nag that fired in 15 skills after every release goes quiet, permanently.**

Detail

* **Plugin mode (new, Claude Code):** the only thing written to your repo is a slim versioned block in `CLAUDE.md`. No `.flow/bin/`, no `.flow/usage.md`, no snapshots to drift. Bare `flowctl list` works in any agent shell; `flowctl usage` prints the always-current CLI cheatsheet + orchestration recipes straight from the installed plugin.
* **Copy mode (unchanged):** repos with Codex/Cursor/Droid teammates, CI, or plain-terminal flowctl use keep the committed snapshots - that is what makes a teammate’s clone work with no plugin installed. The update-then-re-run rule still applies there, exactly as before.
* **Switching is consented, never silent:** moving a copy-mode repo to plugin mode lists the leftover snapshots and asks before removing them; the mode stamp itself is written by a new `flowctl setup-mode set` command that refuses to declare plugin mode unless the CLAUDE.md rail is in place and no snapshots remain - so a half-finished switch cannot leave you in a broken in-between.
* **Two layers of steering, written down:** the [orchestration page](https://flow-next.dev/guides/model-routing/) now spells out session steering (prompts and per-task pins - “implement via grok-4.5 and review with sol” just works and persists nothing) vs machinery steering (config that pilot/Ralph resolve at 3am when nobody is prompting), with the full precedence chain.

### 3.0.0: Smaller, quieter, and only what you use

**You stop paying for machinery you never asked for. Every install used to run Ralph’s guard on every shell command and file edit in every session, whether or not you ever touched autonomous mode - that is gone. The CLI dropped thousands of lines of commands nobody called, review verdicts get harder to fool, and model choices move into config you can actually see and refresh. Three breaking changes, each with a short documented path forward.**

Detail

**If you upgrade, do these first:**

* **Using Ralph?** Run `/flow-next:ralph-init` once per project. Hooks are no longer installed by the plugin - nothing Ralph-related runs anywhere until you opt a project in, and without re-init your guard will not fire. Everyone else: do nothing, and enjoy sessions with zero flow-next hook overhead.
* **Still on a pre-1.0 `.flow/epics/` layout?** The automated migrator is gone; porting by hand is three short steps, listed in `.flow/usage.md` under “Pre-1.0 layout porting”.
* **Scripts or tools reading `epic` fields from flowctl JSON, or passing `--epic` flags?** Switch them to the `spec` forms before upgrading - the legacy aliases and duplicate JSON keys no longer exist. (The `depends_on_epics` field in spec files is real schema, not an alias, and is unchanged.)

**Why this release exists:** an audit asked one question of every deterministic line in the CLI - would this still be needed if the model were smarter? What failed the question was removed or handed back to the agent; what passed got engineered properly. Concretely:

* **The CLI is honest about what exists.** Dead commands are gone rather than half-documented, and the reference docs now match the CLI exactly - if a command is documented, it works.
* **Reviews are harder to fool and easier to extend.** All nine review commands share one engine, review prompts are visible markdown files you can read and diff instead of strings buried in Python, and reviewers report their tallies in a structured block that resists prompt injection from the code under review. Receipts your automation reads are byte-for-byte unchanged.
* **The agent decides; the CLI stores.** `memory add` no longer silently rewrites an existing entry when a new one looks similar - it always creates unless you explicitly say which entry to update, and shows you the near-matches so you (or your agent) make the call. Review-quality judgment calls in interactive sessions go to the agent in front of you; autonomous runs keep the deterministic behavior their receipts depend on.
* **Model choices live in config, not code.** Which model judges triage, reviews, or implements delegated work is now a small table in `.flow/config.json` that `/flow-next:setup` offers to refresh by probing what is actually installed - so pins stop rotting when providers ship new tiers. `flowctl models resolve <role>` shows you what will actually run.

### 2.22.0: The full test suite in 90 seconds

**Waiting 15+ minutes for a test run is where verification discipline goes to die - so the suite now runs in about 90 seconds, and the faster gate immediately caught real problems the old setup had been hiding for months.**

Detail

* A new parallel runner (`scripts/run_tests_parallel.py`, pure stdlib) runs test files concurrently and proves it returns exactly the same results as the serial run. A hung test file fails loudly with its name instead of stalling the whole suite.
* The honest surprise: CI turned out to be running only 29 of 87 test files - the old hand-enumerated steps had silently drifted. The parallel runner discovers everything, and the first full runs on Windows surfaced six real latent issues, each now tracked with its cause instead of swept aside.
* The two slowest test files got 2-3x faster without losing a single test, and per-task verification now runs just the focused suites for the files you touched - the full suite runs once at the end, where it belongs.
* **What it means for you:** if you dogfood flow-next’s pipeline pattern in your own repos, this is the template - a fast full-suite entrypoint plus focused per-task suites is what makes run-the-tests-every-task actually sustainable.

### 2.21.0: Fewer round-trips in every pipeline run

**The skills you run most (plan, land, pilot, make-pr) now read configuration once and create tasks in one call instead of three - less waiting, fewer places for a half-written task to exist.**

Detail

* `flowctl config get` can return a single value, a whole section, or the entire config in one call - so a skill that used to issue seven config reads issues one.
* `task create` accepts the description, acceptance criteria, and requirement links up front. A freshly planned task is complete the moment it exists; there is no window where a task file is created but empty.
* A hardening rider: the rule that review commands must run in the foreground (backgrounded review calls die silently and stall workers - observed live, twice) is now embedded at every point where a review is invoked, and pinned by tests so it cannot quietly erode.

### 2.20.0: Skip-what-you-proved, now actually working

**2.18.0 promised the work loop would stop re-running test suites it had already proven green. An audit of a real run showed the promise was structurally broken - zero receipts were ever honored. This release makes it real, and the measured \~20-25% wall-clock saving on multi-task specs arrives.**

Detail

* The bug: the work loop’s own bookkeeping commit changed HEAD right after every green run, so the receipt recorded for the previous commit never matched. Receipts now survive bookkeeping-only commits through a strictly bounded ancestor check that fails closed on anything suspicious - a receipt is honored only when nothing that could affect the tests has changed.
* Two worker rules close the other measured leak (suites re-run just to *look at* a result the exit code already carried): greenness is read from the captured exit code, and gate suites run as one blocking foreground call.
* Honest bounds, stated plainly: about 35% of runs still deliberately force a full suite as the safety floor. The next lever below that is parallelization - which shipped in 2.22.0.

### 2.19.1: list and status in half a second

**On a large repo, `flowctl list` took 30 seconds - and your autonomous loops paid that on every single tick. Now it is under half a second, with every guarantee intact.**

Detail

* The cause: every task load spawned two git subprocesses - 809 process spawns per `list` on a 100-spec repo. Repo-root lookups are now cached safely (a directory change or a transient git failure never poisons the cache), and several smaller repeat offenders were swept in the same pass.
* Measured on flow-next’s own repo: `list` 30.8s to 0.48s, `status` 32s to 0.41s. Every pilot, land, and Ralph tick gets that time back, every time.
* This was the first installment of a standing audit with one question - would this still be needed if the model were smarter? Mechanisms that survive the question (locks, receipts, atomic writes) get engineered like the hot paths they are; the rest gets deleted. The releases that follow are the same program continuing.

### 2.19.0: Delegation without the paperwork

**Delegating a task to another model used to involve composing a multi-kilobyte briefing document per task - 5 to 17 minutes of prep. A controlled experiment showed the briefing added nothing the task file didn’t already carry. So now the task file IS the brief.**

Detail

* The experiment replayed three real tasks with the historical hand-composed briefs as the control, blind judges, and pre-registered pass bars. The tiny fixed prompt (“read the task and spec files, implement, follow the rails”) tied the composed briefs on every measure; the one gap that appeared closed with a single added sentence in the task template.
* **What it means for you:** plan quality moved to where it belongs - the task file. Plans now require named files, named test cases, and named acceptance criteria, because whatever executes the task receives that file as its entire instruction. A task too thin to delegate safely is implemented in-session instead - automatically.
* Every safety rail (pre-flight checks, rollback, failure classification, the circuit breaker) is unchanged and machine-verified.

### 2.18.0: Stop re-proving what you already know

**A trace of one real pipeline run found the test suite executing ten times in a single work stage - half the wall-clock - while actual implementation was 4%. A docs-only task spent 84% of its time on gates its diff physically could not break. Both stop here.**

Detail

* Green test runs now leave a receipt keyed to the exact commit and command; an identical later check honors the receipt instead of re-running. Anything at all suspicious - a dirty file, a changed command, a stale timestamp - and the suite runs in full. Fail-closed, always.
* A mechanical classifier recognizes when a change touches only docs and prose, and skips the gates that change cannot affect. No semantic guessing - purely file-type rules, with skill prose (which IS shipped behavior) always running full gates.
* Every skip is loud: it lands in the task’s evidence and the run summary, so you can always see what was skipped and why. CI never consults receipts - remote gates always run everything.

### 2.17.0: Tracker updates stop crowding your session

**At 2.17.0, linked tracker updates moved to a background runner to isolate about 124k measured tokens from the working session.** That runner was later removed when deterministic provider operations moved into `flowctl tracker`; current lifecycle callers invoke one compact facade command instead.

Detail

* Historical behavior: the four routine tracker touchpoints dispatched to a small background subagent. Current releases keep semantic judgment in the caller and execute deterministic provider operations through the facade.
* Anything that needs judgment, including discovery choices, 3-way body conflicts, comment synthesis, and structured-error recovery, stays with the host agent.
* Zero added wall-clock in the live proof, and duplicate-protection hardening came straight from that dogfood run (a comment that lands but whose confirmation is lost is re-checked, not re-posted).

### 2.16.0: Interviews that respect your time

**`/flow-next:interview` used to trickle questions a few at a time across many turns. It now asks everything that is ready to be asked in each round - and while you type, an optional background scout looks up the codebase facts the next round depends on.**

Detail

* Rounds are computed from what actually blocks what: a question never appears alongside its own prerequisite, and dropped lines of questioning are announced, not silently vanished.
* Question quality has a bar - “a slot is earned”: failure modes, concurrency, scale, and testing always qualify; cosmetic polish gets folded into option text instead of costing you a question.
* The background fact-scout was validated the hard way: the fastest model tier missed a load-bearing architecture fact the mid tier found on the identical brief, so the mid tier is the floor.

### 2.15.0: Leaner always-loaded docs that teach the one thing that matters

**Every token in your `CLAUDE.md` is context you pay for in every session. An experiment isolated the single thing the flow-next setup block measurably buys - agents filing completion evidence in the right shape - so the block now teaches exactly that, at half the size.**

Detail

* The old block *named* the evidence-JSON flag but never showed the shape; capable agents reliably completed tasks without valid evidence. The new block shows the schema inline. Measured across three model families: the failure mode disappears, nothing else regresses.
* `.flow/usage.md` dropped from \~5.4k to \~1.9k tokens by cutting what `--help` already teaches. It is read on demand, not always loaded.
* Setup re-runs now silently refresh blocks you never customized (a stored hash tells the difference) - so improvements like this reach existing projects without a prompt storm, while your edits are never overwritten unasked.

### 2.14.0: Multi-model defaults you can trust

**Which model should implement delegated work? We ran the experiment instead of guessing: `gpt-5.6-terra` at medium effort matched the stronger tier’s correctness at two-thirds the wall-clock on well-specified tasks - so that is now the default, and setup scaffolds the whole recommended multi-model pipeline when it detects the CLIs to run it.**

Detail

* The recommended shape, now stated concretely everywhere: your session model authors specs (that is where quality is made), a value-tier model implements against them, and a reviewer from a different model family reviews - uncorrelated blind spots by construction.
* Setup only offers what can actually run: routing questions appear when the bridge CLIs are installed, cross-family review is recommended only when it IS cross-family for your setup, and switching an existing review backend is always explicit, never silent.
* For unattended loops, the docs now carry the self-healing wrapper pattern and the known sharp edges of each bridge CLI (silent-refusal modes, workspace checks) so overnight runs fail loudly instead of mysteriously.

### 2.13.1: Portfolio triage stays quiet on ordinary repos

**Dogfooding 2.13.0’s classifier on real machines found it flagging noise - your own repo referencing itself, `~/Downloads`, cache directories - as “cross-repo signals”. Patched same-day: `--classify-only` sweeps now stay quiet unless there is a real constellation to report.**

### 2.13.0: Prime tells you the truth about your repo

**The old readiness assessment could award “Level 5” to an empty template with a broken build - it checked whether files existed, not whether anything worked. The rebuilt `/flow-next:prime` runs your build, runs your tests, boots your app, and leads with a verdict and a ranked list of what to fix first.**

Detail

* Prime now starts by classifying what it is looking at - lifecycle, monorepo topology, size, stacks (including legacy ones like Delphi and COBOL), delivery shape - and judges against the right yardstick for that kind of project. A greenfield toy and a 20-year brownfield monolith stop getting the same checklist.
* Hard gates cap the score: if the build does not run, tests are not discoverable, or the commands in your agent docs do not actually execute, no amount of nice file structure rescues the rating. Executed evidence beats existence.
* `--classify-only` sweeps a portfolio of 100+ repos in seconds when you want triage rather than a deep assessment.
* The claims survived 29 adversarial review waves before shipping - capped scans admit they are capped, CI steps only count if they actually gate, and secrets never leave the machine.

See the rewritten [Prime skill page](https://flow-next.dev/skills/prime/).

### 2.12.4: A stale setup finally gets your attention

**If your project’s local flow-next files lag the plugin you are running, skills now ask you once - Refresh now / Remind me next version / Skip - instead of hiding a one-line note you were never going to see.** Your answer is remembered per version, autonomous runs are never interrupted by the question, and pilot/land print a grep-able `SETUP_STALE:` line for loop drivers. Also fixed: tracker-sync no longer reports false conflicts when Linear rewrites your markdown on save - the sync now remembers what the tracker actually stored, not what was sent.

### 2.12.3: Know what each model is actually good for

**The optional model-routing table gains a speed column and a Grok 4.5 row - fast and cheap with strong coding, weaker on UI and more prone to invention, so the guidance routes it to bulk implementation and never to final taste-critical work.** The fourth headless bridge (`grok -p`) joins the recipes, and the cost column now says plainly what it measures: how lightly a model rides your subscription quota, not list price. Nothing activates unless you ask for routing - defaults are unchanged.

### 2.12.2: Delegated work goes to a strong model, deliberately

**Delegated implementation is real work that ships, so its default model moved up a tier rather than down - the counterpart to routing cheap models at bulk reads.** One sharp edge documented honestly: the delegation path has no fallback ladder, so this default needs a current codex CLI; the one-line downgrade for older CLIs is in the notes.

### 2.12.1: Fresh model numbers the day after a launch

**GPT-5.6 went GA and the scaffolded routing table was still handing new users last-generation numbers - so the table got current rows and an explicit “as of” staleness stamp.** These are starting opinions you edit after scaffolding, not runtime defaults; nothing about live model resolution changed.

### 2.12.0: PRs that tell reviewers where to look

**As AI-assisted PRs grow, reviewers lose the thread - which came from a field report in exactly those words. PR bodies now open with what the pipeline already verified mechanically versus what genuinely needs human judgment, and a risk-ranked review plan: Must review / Spot-check / Safe to skim, capped at roughly 30% of the diff, each item saying why and what to check.**

Detail

Validated blind before implementation: reviewers scored the old PR bodies 7/10 on effort-targeting and 5/10 on trust calibration; the shipped format scores 9/9. Underneath, the PR export gained deterministic traceability - which symbols changed per file, which files are generated copies a human should never re-review, and which deleted exports still have references (conservative candidates worth a look). No claim in a PR body is invented: every line traces to a computed field.

### 2.11.0: A model launch never breaks your reviews again

**GPT-5.6’s launch reproduced a familiar failure live: review backends erroring because their pinned model did not exist yet on your CLI. Model resolution now dispatches the best model directly, and when a model is genuinely unavailable it steps down a quality-ranked ladder to the CLI’s own default - your review runs, always, and the receipt records what actually ran.**

Detail

The outcome is cached per CLI version so the fallback probe costs one retry per upgrade at worst. Your explicit pins bypass everything, unchanged. Unknown model names warn and proceed instead of failing - no more waiting for a plugin release to try a model that launched this morning. Three review rounds before merge caught three real bugs, including the fallback being silently dead on the paths that mattered; each is now pinned by a regression test shaped like the actual failure.

### 2.10.3: Cursor reviews on GPT-5.6 Sol, verified first

**The cursor review backend moved to GPT-5.6 Sol with 1M context - after live-verifying the model actually resolves on the current CLI.** Codex and copilot deliberately stayed put: live probes showed both would reject the new model, and swapping an unverified default is precisely the failure 2.11.0 then eliminated for good. This was the verified interim.

### 2.10.2: Capture read-back: plain-language ratification, no self-blessing

**The capture approval question - where a human ratifies the synthesized spec - now opens with what approving means, lists every requirement in one plain line inside the question, translates its own machinery (“\[inferred] = something I added that you didn’t say outright”), and never recommends approve while unverified inferred items exist - the agent doesn’t pre-bless its own guesses.** The companion to 2.10.1, aimed at the highest-stakes dialogue moment in the system. Blind-eval before shipping found the real bug was beyond language: “Recommended: approve - the inferred items are reasonable” was steering users to rubber-stamp exactly what they exist to verify (ratification-safety 4/10). The shipped contract scores 9/9/10 on legibility / ratification-safety / precision.

### 2.10.1: Interview questions in plain language

**Every interview question now opens with one sentence of stakes (what this decides, in your words), glosses any term of art in plain words at first use, and states each option’s consequence (“Choose this if…”) - field feedback from a team where jargon-dense questions were disempowering the product people the interview exists to hear.** Sizing is expressed as priorities (always-keep / trim-first lists with a target shape), not a length cap - deliberately aligned with OpenAI’s GPT-5.6 guidance that generic brevity instructions make capable models drop required content. Eval-validated blind before shipping: baseline questions scored 4/10 legibility for a second-language PM; the shipped contract scores 7.5+ with precision held and \~30% fewer tokens per question. The same eval tested extending this to spec prose and rejected it (business sections already read at 9/10 - the contract only made them longer), so specs are unchanged by design.

### 2.10.0: Review-loop runaway root fix: honest verdicts, convergence ratchet, deterministic cap

**A field-reported 17-round plan-review runaway on the Cursor backend is fixed at its roots: codex/copilot verdicts are now parsed only from the reviewer’s final message (tool-output `<verdict>` literals can no longer produce a false SHIP/NEEDS\_WORK), re-reviews follow a shrink-only convergence ratchet (prior findings injected; only NEW ≥Major blocks; all-fixed ⇒ MUST SHIP), and `MAX_REVIEW_ITERATIONS` (default now 4) is enforced deterministically by flowctl - the counter survives fresh invocations and refuses at the cap with an `ESCALATE` marker instead of looping.** Cursor reviews also carry a persona override superseding cursor-agent’s built-in review rubric and auto-attached AGENTS.md guidance. Review receipts are spec/task-scoped (concurrent reviews no longer collide); `flowctl spec reset-review-rounds` resets on re-plan. Round counting includes failed dispatch attempts by design (anti-runaway bias). Validated live on a pinned five-review baseline dataset: a cursor fix→re-review cycle converges to SHIP in 2 rounds.

### 2.9.1: Fix: completion-review tracker audit was dead (event-key mismatch)

**The work skill’s completion-review tracker touchpoint dispatched and audited the event as `work.completionReview`, but the config leaf is top-level `tracker.perEvent.completionReview` - so `flowctl sync check` resolved no leaf, treated the event as configured-off, and could never report the touchpoint missing (a dead dispatch could never retro-fire). The event key is now `completionReview` everywhere; existing configs keep working unchanged.**

Detail

Found during the fn-90 review-loop investigation: five independent cross-backend reviews (cursor + codex) of a planned spec each flagged the key mismatch, and a code check confirmed it live. Only the event tag moved - the config leaf written by the discovery ceremony stays top-level, so no user action is needed beyond updating the plugin. Regression coverage landed in `test_sync_check.py`: the top-level key round-trips the audit (enabled + no receipt → MISSING; a tagged receipt clears), the old `work.`-prefixed shape demonstrably resolves no leaf, and a prose guard fails the suite if any canonical skill or doc reintroduces the mismatched tag.

### 2.9.0: Interview: scope question + skips are not answers

**A bare `/flow-next:interview` no longer silently runs the technical question bank - it asks which pass to run (business / technical / both) - and skipped questions no longer become silent decisions: they park under `## Open Questions`, with a consent checkpoint before write-back.**

Detail

Both fixes come from a downstream field report: a product manager ran a bare interview, skipped the technical questions, and the agent filled architecture/stack/API sections with project-rails-derived defaults written as settled decisions.

* **Scope question** ([interview](https://flow-next.dev/skills/interview/)): when no `--scope` / `--biz` / `--tech` flag is passed, the interview asks one upfront question with a recommendation derived from the spec’s current state (both layers empty → `both`; business populated → `technical`; 1.0.2-shape tech-only spec → `technical`). An explicit flag skips the question; `technical` remains the fallback when the question can’t be asked. Plumbing: `flowctl scope resolve --json` now emits a `defaulted` boolean. The Codex mirror gets the same question via the plain-text prompt transform.
* **Skip contract**: only an explicit answer or an explicit “you decide” delegation resolves a question. Skips/declines/“I don’t know” park under `## Open Questions` with the agent’s *unconfirmed* leaning; a skipped judgment question never demotes to codebase-/docs-derived backfill. When anything was skipped, a consent checkpoint fires before write-back: `park-open` (default) / `fill-assumptions` (inline `*(assumed, unconfirmed)*` markers for later ratification) / `re-ask`. The completion summary reports the disposition.

After updating, re-run `/flow-next:setup` in your projects so the local `.flow/bin/flowctl` copy picks up the new `defaulted` field.

### 2.8.1: Model-routing scaffold: named tiers, menu wiring, freedom grant

**The setup scaffold’s routing table now names concrete models (`fable-5`, `opus-4.8`, `gpt-5.5`, `composer-2.5`, `sonnet-5`, `haiku-4.5` - “session model” is a role, whichever row conducts), and the wiring is a per-role menu instead of fixed pairings: implementation via native opus/sonnet subagents, `delegate:codex`, or the `cursor-agent` bridge; reviews cross-family or prompted same-family; reads native or via cursor scouts.** Plus an explicit freedom grant: unless prompted otherwise the harness routes as it judges best, and an explicit user instruction always overrides the table. Probe-gating unchanged - routes to CLIs you don’t have are never active. The [orchestration page](https://flow-next.dev/guides/model-routing/) block mirrors the shipped scaffold verbatim.

### 2.8.0: Orchestration in your repo: usage.md steering recipes + setup routing scaffold

**The orchestration story now ships into the repo you work in: every installed `.flow/usage.md` carries dogfooded headless-bridge recipes (`codex exec`, `cursor-agent`, and the reverse `claude -p` - every direction works), and `/flow-next:setup` optionally scaffolds an opinionated model-routing table into `CLAUDE.md`/`AGENTS.md`.**

Detail

Closes the use-time discoverability gap from 2.7.2’s [orchestration doc](https://flow-next.dev/guides/model-routing/): agents read `.flow/usage.md` and instruction files, not the plugin’s doc tree.

* **usage.md `## Orchestration & model steering`** (unconditional, ships in every project): `codex exec` recipes with the real gotchas baked in (read-only default sandbox, `-o` output capture, the `</dev/null` stdin-hang guard), `cursor-agent` (`-p`, `--force` to apply, volatile model IDs via `--list-models`), the `claude -p` reverse bridge (prompt before the variadic `--allowedTools`), flow-next shortcuts (`delegate:codex`, `review.backend`, per-task `review:`), and prompted-orchestration examples. Every recipe was verified against the live CLIs before shipping.
* **Optional setup ceremony** ([setup](https://flow-next.dev/skills/setup/)): scaffolds the cost/intelligence/taste scores table + routing rules + flow-next wiring into your instruction file - probe-annotated for the CLIs actually installed (`<!-- probe:codex/cursor -->` sentinel lines, deterministic composition), shown in full before writing, marker-fenced for idempotent re-runs (probe drift counts as drift), platform-correct invocation syntax per target file. Delegation opt-in sets `work.delegate` but never pre-sets the consent gate.
* **`/flow-next:uninstall`** removes the scaffold via a deterministic damaged-marker algorithm; 23 new tests + smoke prose contracts pin the template shape, four-state probe composition, and removal.
* **Codex installs**: the mirror’s usage.md now renders commands as `$flow-next-<cmd>` (generator rewrite + regression guard).

Defaults stay pre-tuned and unchanged - steering remains a capability, not a prerequisite; headless/Ralph setups skip the new question silently.

### 2.7.2: Orchestration & model routing doc

**Given the trend toward frontier-model orchestration - Fable 5 conducting while implementation, reviews, and bulk reads route to cheaper/faster models - a new [Orchestration & Model Routing](https://flow-next.dev/guides/model-routing/) page maps every routing dial flow-next already ships.** Two composable methodologies: deterministic parameters (review-backend grammar + precedence, `delegate:codex` offload, subagent tiers) and *prompted orchestration* - the host’s own intelligence routing per item by complexity, escalating conditionally, even prompting capabilities into existence that no parameter encodes. Plus a copy-paste CLAUDE.md model-routing table and the pilot+land loop-chaining recipe. The frame throughout: the defaults are pre-tuned to work well out of the box - steering is a capability, not a prerequisite.

### 2.7.1: Codex hooks.json parse fix

**The installed Codex `~/.codex/hooks.json` carried a top-level `description` key that Codex’s hooks parser (stable since 0.142.x) rejects - a warning on every invocation and the Ralph guard hooks silently disabled; the generated mirror no longer emits the key.** Reproduced and verified clean against Codex CLI 0.142.5. Codex users: re-run `scripts/install-codex.sh` to replace the broken file. Thanks to [@TechupBusiness](https://github.com/TechupBusiness) for the report and root-cause analysis ([#198](https://github.com/gmickel/flow-next/issues/198)).

### 2.7.0: Fleet-wide capability & efficiency review

**Two adversarial-review passes over the whole skill/agent fleet - one new feature (`make-pr --update`), broad correctness and autonomy-safety fixes, progressive-disclosure efficiency, and seven A/B-verified `opus→sonnet` model downgrades - with every judgment call fable-reviewed before it shipped.**

Detail

The engine was adversarial review, not prose-squeezing: six-reviewer fable audits found the gaps, a fable judge verified each fix, and every model-tier change was proven head-to-head. Across \~46 improvements in 25 skills and agents:

* **New capability** - `make-pr --update` refreshes a stale PR body after review/land fix rounds; `/prime` gained a real evidence + scoring contract (a failed scout no longer silently drops or fabricates a pillar, verification is mandatory before a “runnable” pass, inapplicable criteria no longer deflate the score) plus create-or-**augment** CLAUDE.md/AGENTS.md handling; completion-review scope-creep detection, plan real-anchor derivation, prospect’s genuinely-isolated critique, qa evidence enforcement, and `/deps` surfacing deadlocks it used to hide.
* **Autonomy safety** - `land` stops auto-merging a PR whose QA verdict is NEEDS\_WORK; `pilot`’s strike limit survives a tracker re-projection that had it re-dispatching a failing spec forever; `quality-auditor` fails loudly instead of reporting a false-clean audit over an empty diff.
* **Model tiers** - `plan-sync`, `flow-gap-analyst`, and the five retrieval scouts moved `opus→sonnet`, each A/B-verified; `opus` is now used by a single agent. Cheaper per call, quality held.
* **Efficiency** - progressive-disclosure splits (`interview`, `make-pr`, `impl-review`, `capture`) and common-path short-circuits (`tracker-sync`, `audit`) take \~5k-17k tokens off the paths that run most.
* **Correctness spine** - a review-diff `base...HEAD` fix (13 sites) that had a fast-moving base branch showing its own commits as false reversions, and a signal refresh so `/prime` recognizes 2025-era stacks (uv, bun, `compose.yaml`, mise, monorepo layouts, GitLab, goreleaser, Biome).

No breaking changes. Verified by a 1425-test suite green across Linux/macOS/Windows.

### 2.6.3: Single-call worker anchor + plan-sync gate shelved

**The `/flow-next:work` worker now re-anchors in one `flowctl anchor` call instead of \~8 separate reads (proven zero information loss), plus a CROSS\_SPEC caller bug-fix - while the other half of the work, a deterministic plan-sync skip-gate, was proven non-viable by cross-repo eval and deliberately shelved rather than shipped.**

Detail

`flowctl anchor <task-id>` assembles the worker’s Phase-1 re-anchor from the verbatim stdout of the same production commands it already runs - byte-for-byte superset test plus a comprehension-equivalence eval (bundle 7/7 = status-quo on frozen real tasks) prove no information is lost. The plan-sync skip-gate (a deterministic probe to skip the post-task drift check) was built, eval’d against the real plan-sync agent across three external repos, and killed by its own evidence: a genuine false skip from semantic drift no path/token probe can see, plus a 6.7% skip-rate against a ≥50% bar. It is shelved with a decision record, not shipped - plan-sync still runs after every task, and the gate machinery was removed from the CLI. Also fixed: a Windows encoding bug in the anchor render, surfaced by wiring the anchor guardrail test into CI.

### 2.6.2: ready honors spec-level deps

**`flowctl ready --spec` now honors spec-level dependencies - a spec blocked by unfinished `depends_on_epics` no longer reports its tasks as ready.** It returns empty lists plus `blocked_by_specs` (legacy alias `epic_blocked_by`), matching the gate `next` and `ready --all` already applied. Latent in the default workflow; hit by external consumers calling `ready` per-spec. Thanks to Mike Bannister ([#95](https://github.com/gmickel/flow-next/pull/95)).

### 2.6.1: Codex hooks config fix

**Setup and `install-codex.sh` could leave a Codex `config.toml` with a duplicate `hooks` key (invalid TOML - Codex silently stops loading hooks) or the deprecated `codex_hooks` spelling (a warning on every run); both paths now converge through one idempotent, dedup-safe normalizer that guarantees exactly one `hooks = true` under `[features]`.** Regression-tested (10 cases including the both-keys scenario); everything outside `[features]` is byte-preserved and re-running is a no-op.

### 2.6.0: Skill efficiency: single-emission writes + prompt diet

**Two paired specs cut the token cost of every skill run with zero quality loss - fn-81 eliminates *runtime* re-emission (spec bodies, review prompts, and responses materialized once instead of two-or-three times, plus 13 redundant CLI round-trips removed across 12 skills), and fn-82 trims the *always-loaded* prompt weight the hot-path skills carry on every invocation (−10.7k tokens across 11 skills). Skill-markdown only - no `flowctl` behavior change, no new commands, read-backs stay mandatory and user-authoritative.**

Detail

Follows a fleet survey of all 28 skills (2026-07-02) and lands behind a full behavioral regression pass (gate matrix, two eval-suite re-runs at full score, smoke 138/138, pytest 1393 passed).

* **Runtime plumbing.** The drafted spec body is materialized exactly once via the Write tool (the Write render *is* the user-visible read-back), revised via Edit deltas, and consumed by `spec set-plan --file <path>` - no Phase-5 heredoc re-authoring (capture, interview). RP review prompts are built by deterministic file composition (`rp prompt-get > file`, quoted-heredoc criteria, `flowctl show >> file`) - every `[PASTE …]` content-retype placeholder is gone and untrusted reviewer/spec content never transits a shell var (the injection surface is closed). RP review responses enter context exactly once (redirect → single Read). Round-trips removed: single `LEAF=` config read per tracker gate (7 sites), plan drops a post-write `show`+`cat`, deps runs one per-spec loop instead of two, make-pr’s §4.6b live `gh pr view` fires only on the local-assertion miss, tracker-sync reconcile passes the on-disk spec to `set-merge-base --flow-file`. Guards hardened: the fix-loop cap (`MAX_REVIEW_ITERATIONS`, default 3) now bounds **all** review backends, and both RP fix loops replace `git add -A` with snapshot-scoped staging (pre-existing dirty paths are never swept in).
* **Prompt diet.** Default-OFF machinery moved behind a **forcing-sentinel gate** into `references/*.md` - zero tokens until Read (Anthropic Agent Skills 3-level loading): work’s tracker touchpoints and pilot’s QA-stage freshness probe. Each gate emits an imperative the agent must act on (`GATE ACTIVE, STOP. Read <ref> …`), **fails open** on probe/parse error, and no-ops silently on the default path; the safety nets (work’s Phase-5 `sync check` + four-state summary, pilot’s QA routing) stay inline. Duplicated explanatory blocks collapse to one authoritative site (the review pair now resolves the backend once - killing a double `review-backend` round-trip); build-time `fn-N` provenance and `flowctl.py` line-refs are stripped from always-loaded prose; make-pr **folds** its per-phase Done-when checklists inline (body eval held 5/5, −4.5k tok/run) and capture single-sources its biz-routing table at the *consumer* (suite held 15/15).

### 2.5.4: Section-write hardening + rp-gate completion

**`flowctl` task-section writes are now normalization-hardened (the H2-layering bug caught in fn-78’s own autonomous dogfood is fixed, with self-heal for already-damaged files), and the fn-78 RepoPrompt eligibility gate now covers all four review skills - `impl-review` and `spec-completion-review` stop steering toward rp on hosts where it can’t run.**

Detail

* **Task-section normalization.** Agents routinely pass section content that starts with its own `## Acceptance Criteria …` H2; `task create --acceptance-file` embedded it as a rogue sibling section and every later `set-acceptance` *layered* a new block above the old one. All task-section write sites now normalize through one helper: a leading H2 is stripped only when it matches the section’s known-title-variant grammar (`## Acceptance Tests` is content - demoted, never stripped), remaining H2s demote to H3 outside code fences, writes are byte-idempotent, and an on-write self-heal folds contiguous rogue sections from already-damaged files (a byte-exact duplicate `## Acceptance` still raises). Fence-awareness extended end-to-end via one shared tracker (`patch_task_section`, `get_task_section`, heading validation, `set-spec` scaffold check).
* **RP\_ELIGIBLE gate completed.** The 2.5.3 gate covered `plan`/`plan-review`; now `impl-review` and `spec-completion-review` compute the same guard locally in every gated file and, when ineligible (non-macOS, no `rp-cli`), steer only to `codex`/`copilot`/`cursor` (+ `none`). Explicit `--review=rp` / env / config / per-task overrides still resolve; eligible hosts render byte-for-byte as before.

### 2.5.3: RepoPrompt proposal gate + review-call hardening

**`/flow-next:plan` and `/flow-next:plan-review` no longer *offer* the RepoPrompt path on hosts where it can’t run (non-macOS with no `rp-cli` on `PATH`) - explicit `--review=rp` / config still resolves as before - and the review skills now pin an explicit Foreground rule so agents never background a review CLI call and idle on a finished verdict.**

Detail

* **RepoPrompt eligibility gate.** RepoPrompt is a macOS-only GUI app, yet both skills proactively dangled the rp option in their interactive setup on every host - on Linux/Windows without `rp-cli`, picking it was a guaranteed runtime failure. Both now compute one POSIX guard - `RP_ELIGIBLE ⟺ uname == "Darwin" OR rp-cli on PATH` - and, when ineligible, drop every RepoPrompt *proposal* (plan’s research question defaults silently to `repo-scout`; plan-review steers only to the runnable `codex` / `copilot` / `cursor` + `none`). **Suppression is not a ban:** explicit `--research=rp` / `--review=rp` / `FLOW_REVIEW_BACKEND=rp` / `review.backend=rp` still resolve; eligible hosts render byte-for-byte as before.
* **Foreground rule for review CLI calls.** Found in fn-78’s own autonomous dogfood: a worker subagent backgrounded its cursor impl-review and idled on the already-finished verdict (background completion doesn’t reliably resume a subagent). The backend CLI was flawless - 8/8 verdicts - so `impl-review` / `plan-review` / `spec-completion-review` and the `worker` agent now pin the calling discipline: one blocking foreground call, generous timeout, never background + monitor.

### 2.5.2: Scout models tiered by task

**The scout subagents move off a frozen `claude-sonnet-4-6` pin to family aliases matched to each task: the 8 pure config-scanners drop to fast, cheap `haiku` (Haiku 4.5 - which out-scores the `gpt-5.4-mini` the Codex mirror already runs them on), while only the 3 judgment scouts (`spec-scout`, `claude-md-scout`, `docs-gap-scout`) stay on `sonnet`; heavy agents keep `opus`, `worker`/`pr-comment-resolver` `inherit`. No version pins, cheaper + faster scouts, and Claude finally matches the FAST/INTELLIGENT tiering the Codex mirror already encoded.**

### 2.5.1: Windows python3 Store-stub fix

**`flowctl` now *just works* on Windows when `python3` resolves to the Microsoft Store **App Execution Alias** stub - a 0-byte reparse point that’s on `PATH` but exits 9009 - by probing interpreter *functionality* (`<cand> -c "import sys"`) instead of presence, across every invocation context (Git Bash / WSL, cmd.exe / PowerShell, Claude Desktop, native Codex / Cursor), with a companion `flowctl.cmd` launcher and no mac/linux regression.**

Detail

* **Probe over presence.** A shared resolver (`scripts/lib/pick-python.sh`) and the self-contained launchers probe interpreter *functionality* in order `$PYTHON_BIN` → `py -3` → `python3` → `python`; the 9009 stub is skipped even though it’s on `PATH`, while a machine with a working `python3` (and no `py` launcher) still picks `python3` first. The old launchers hardcoded `exec python3` and the prior `pick_python` helper tested `command -v` (presence, not function) - both selected the broken stub.
* **Dual launcher.** A `flowctl.cmd` batch shim ships alongside the extensionless bash `flowctl`, running the same probe under cmd.exe / PowerShell where the bash shebang is never honored (`py -3` preferred). CRLF/LF pinned so Git Bash doesn’t regress.
* **`init` self-heal.** `flowctl init` re-stamps both `.flow/bin/flowctl` and `.flow/bin/flowctl.cmd`, so an existing (pre-fix) install refreshes on the next `init` - no full re-setup. A broken bash launcher is reached via the new `.cmd`, a plugin auto-update, or `py -3 .flow/bin/flowctl.py init`.
* **Swept everywhere + covered.** Ralph hooks, `watch-filter.py`, and the qa/prospect agent heredocs all resolve a working interpreter (Ralph mode requires Git Bash on Windows). A fake-9009-stub regression harness plus a real `windows-latest` CI job (proper `.exe`/`.cmd` stub) exercise both launchers against the stub; `docs/troubleshooting.md` + `docs/platforms.md` document the fix, the probe order, and both recovery paths (re-stamp via `init`, or disable the App Execution Aliases).

### 2.5.0: Cursor backend + sharper reviews

**A fourth cross-model [review](https://flow-next.dev/guides/review-workflow/) backend - `cursor` (Cursor-billed `cursor-agent` CLI) - joins `rp` / `codex` / `copilot`; all agentic backends now read changed files from disk instead of embedding them (smaller, cheaper prompts); the review rubric itself gets eval-validated tuning - an always-on code-smell baseline lifts impl detection 7 → 10/10 at \~27% fewer prompt tokens, and plan reviews gain a spec-quality checklist (8.0 → 9.7); and per-task / per-spec `review:` overrides now route correctly instead of silently falling back to the project default.**

Detail

**Cursor review backend.** A parity port of the `copilot` backend - no new review *features*, same Carmack-level criteria, same receipt schema, same verdict grammar, same `--deep` / `--validate` passes - wired through [`/flow-next:impl-review`](https://flow-next.dev/guides/review-workflow/), `/flow-next:plan-review`, `/flow-next:spec-completion-review`, and `/flow-next:setup`. Select it the usual ways: `flowctl config set review.backend cursor`, `FLOW_REVIEW_BACKEND=cursor`, `--review=cursor`, or a per-task/spec `cursor:<model>`.

* **Cursor-billed, no extra key.** Runs `cursor-agent -p --output-format json --trust --mode ask` against the workspace (read-only Q\&A - it never mutates the tree). Reaches reviewer models the others can’t in one place: `gpt-5.5-high` (1M ctx, the default), the `gpt-5.3-codex` family, `composer-2.5`, Opus 4.8 thinking. Auth is your stored `cursor-agent` login or `CURSOR_API_KEY`.
* **Resume-only sessions.** The first review omits `--resume` and persists Cursor’s generated `session_id`; a re-review resumes it (only when the prior receipt’s `mode == "cursor"` - a cross-backend receipt starts fresh).
* **Effort folds into the model name** (Cursor convention), so a spec is `cursor:<model>` with no `:effort` rung - `cursor:gpt-5.5-high`, not `cursor:gpt-5.5:high`.
* **Triage judge unchanged.** The opt-in LLM triage judge (`FLOW_TRIAGE_LLM=1`, default off) stays `codex|copilot`; with it off cursor reviews use the deterministic trivial-diff whitelist, zero extra dependency.

**Review backends read changed files from disk.** The agentic backends - `codex`, `copilot`, `cursor` - no longer embed changed-file contents (previously up to \~500 KB) into the reviewer prompt. They read from disk the way `rp`’s Builder already did (codex sandbox, copilot `--add-dir`, cursor `--mode ask`), so prompts are smaller and cheaper and `cursor` no longer trips its argv limit on non-trivial diffs. Verified equivalent on a ground-truth planted-bug test (codex’s own audit: QUALITY=PRESERVED). The per-backend `FLOW_*_EMBED_MAX_BYTES` budget knobs are removed.

**Sharper, leaner review prompts.** The Carmack rubric gains an always-on **code-smell baseline** (Fowler *Refactoring* ch.3 - Feature Envy, Data Clumps, Primitive Obsession, Long Method, Duplicated Code, …) on impl + standalone reviews, with its rubric blocks tightened and every machine-parsed marker preserved. Applied to **every backend** - codex/copilot/cursor and RepoPrompt. Eval-validated on a ground-truth corpus (correctness bugs + planted smells): detection rose **7 → 10/10** (the old rubric reliably missed Feature Envy / Data Clumps / Primitive Obsession) while the prompt shrank **\~27% (−950 tokens)**, correctness detection held at **5/5**, and clean code was not over-flagged - confirmed on both codex (GPT-5.5-high) and RepoPrompt. Plan reviews additionally gain a targeted **spec-quality checklist** (a stated test strategy, observability for async/batch work, each task sized-for-one-iteration and correctly dependency-ordered, non-functional requirements) - eval-validated **8.0 → 9.7/10** for **+74 tokens**, no over-flagging of good specs.

**Per-task / per-spec review-backend overrides route correctly.** A task’s `review: <backend>:...` (or a spec’s `default_review`) is now honored end-to-end: `flowctl review-backend` resolves the per-task/epic override **above env/config** (canonicalizing short/tracker handles first), and every review skill + `/flow-next:work`’s per-task worker passes it - so a task set to `review: cursor:...` under a `codex` project default actually reviews with **cursor**. Every backend command also defensively **coerces a foreign stored spec to its own default**, so an explicit `--review=<backend>` / `flowctl <backend>` always wins over a stored cross-backend spec instead of shelling a foreign model.

**Copilot CLI 1.0.65 compatibility.** The default copilot model moves `gpt-5.2` → `gpt-5.5`, and `gpt-5.2` / `gpt-5.2-codex` are dropped from the accepted set (1.0.65 rejects them), so `copilot:gpt-5.2` is now rejected. Session creation is fixed for the CLI’s resume-only `--resume` change - the first call now uses `--session-id` (marker-tracked) and re-reviews resume it.

### 2.4.0: GitLab + Jira tracker adapters

**Tracker-sync gained GitLab and Jira as its 3rd and 4th providers, so teams on the dominant self-managed (GitLab) and enterprise (Jira) trackers could mirror Flow-Next specs to their board with zero special setup.** The prose-driven provider implementation described in this historical entry was superseded by the deterministic `flowctl tracker` boundary; current behavior is documented on [Tracker Sync](https://flow-next.dev/integrations/tracker-sync/).

Detail

The supported-tracker set became **Linear, GitHub, GitLab, Jira**. Current releases normalize these providers behind `flowctl tracker`; the skill retains semantic merge and recovery judgment.

**GitLab** (the 3rd tracker - a large share of self-managed and EU/regulated shops). Modelled on the GitHub adapter:

* **Historical GitLab implementation.** At 2.4.0 the skill described direct `glab` and REST choices. Current releases resolve GitLab once, persist destination and capability facts under `tracker.resolved`, and execute provider operations through `flowctl tracker`.
* **Reduced-fidelity status, like GitHub.** Open/closed plus a configurable board label, not a rich workflow.
* **License-gated dependency projection.** `depends_on_epics` edges project as native `is_blocked_by` links on a Premium/Ultimate namespace; a Free or personal namespace (where the API returns `403 Blocked issues not available for current license`) degrades to a directionless `relates_to` link plus a provenance-fenced `<!-- flow:deps -->` body block for direction.

**Jira** (the 4th tracker - the enterprise default). REST-only by design, the most adapter-specific weight of the four:

* **Historical Jira implementation.** At 2.4.0 Cloud used API version 3 and Data Center / Server used version 2. Current releases pin both deployment families to API version 2 for plain-string body fidelity and migrate a legacy configured version 3.
* **No MCP.** The official Atlassian MCP is read-mostly - it can’t transition status, update fields, or set links - so the bridge uses the REST + token path directly, headless-native with the fewest moving parts.
* **Workflow-aware status.** A change goes through the transitions API against a configurable `statusMap`; an unmapped or unreachable transition defers with a receipt rather than forcing a lane. The fn-66 terminal invariant holds - a locally-done spec stays In Review until the PR is **MERGED**.
* **Current Jira fidelity.** Dependencies project as native directional `Blocks` issue links. API version 2 keeps bodies as plain strings, and backlog enumeration runs through `flowctl tracker wire list-open`.

The new adapter behavior is documented on [Tracker Sync](https://flow-next.dev/integrations/tracker-sync/).

### 2.3.0: Pilot backlog mode

**[`/flow-next:pilot`](https://flow-next.dev/skills/pilot/) gains an opt-in [backlog mode](https://flow-next.dev/autonomy/pilot/#backlog-mode) (`pilot.autonomy=backlog`, default off): instead of advancing one already-ready spec, pilot widens to a standing scheduler for the entire open backlog - enumerating flow specs + tracker issues, triaging the top dep-ordered item, and either advancing it one stage or surfacing a precise async question and parking it (`ASKED`). The consent boundary moves from *before* the loop to *inside the loop, on block*, while every safety boundary holds: it never authors a spec, never promotes, and never merges.**

Detail

By default pilot’s consent boundary sits *before* the loop - it only picks from the already-ready queue. Backlog mode (`flowctl config set pilot.autonomy backlog`, or per-run `--backlog` / `--auto`) makes each tick enumerate everything open (flow specs via `flowctl ready --all` plus tracker issues at the promoted lane, unioned in from the tracker-sync adapter), select the top **dep-ordered** actionable item, **triage** it agentically, and advance it along the same `plan → plan-review → work → [qa] → make-pr` pipeline. When it can’t safely proceed it surfaces an **async question** into the spec’s `## Open Questions` + a tracker comment and parks the item - “stuck” becomes a question a human answers async, not a stall.

* **Same single-tick conductor, widened left.** One `/loop`/`/goal` target, one verdict grammar (adds `ASKED <id> (<n>)`, keeps `NO_WORK`/`DEFERRED_TO_LAND` verbatim), one mental model - not a new skill or command, and not a [prospect](https://flow-next.dev/skills/prospect/)-style idea generator (it manages the *existing* backlog).
* **Boundaries hold.** Never authors a spec (a thin/missing spec is a surfaced *“run [`/flow-next:capture`](https://flow-next.dev/skills/capture/) or [`/flow-next:interview`](https://flow-next.dev/skills/interview/)”* gap); never sets the `ready` flag (promotion is the human’s board act; un-promoted items are skipped silently); never merges ([land](https://flow-next.dev/autonomy/land/) stays human-gated). Readiness stays the human’s explicit signal, never an agent-inferred score.
* **Substrate.** A backlog-wide eligibility scan (`flowctl ready --all` → deterministic facts only), a per-tick decision log (`flowctl pilot-log` → the factory-efficiency readout), and a tracker-sync autonomy-parity fix + the async question-valve so a per-tick sync never hangs the loop. The agentic/deterministic line holds: flowctl enumerates + checks hard fields; the host agent judges and formulates the question.

Off by default - existing pilot/land/Ralph users are unaffected until they opt in.

### 2.2.0: QA pipeline stage + Cua native driver

**Two opt-in additions to the autonomous pipeline: [`/flow-next:qa`](https://flow-next.dev/skills/qa/) becomes a config-gated ([`pipeline.qa`](https://flow-next.dev/guides/live-qa/#lifecycle-position), default off) [pilot](https://flow-next.dev/skills/pilot/) stage that live-tests the complete build before make-pr, and [Flow-Next Drive](https://flow-next.dev/skills/flow-next-drive/)’s native rung gains the [Cua](https://github.com/trycua/cua) driver + sandbox for provider-agnostic, headless/CI computer-use. Both augment, never replace, existing tooling.**

Optional QA pipeline stage

`/flow-next:qa` already did the hard part - derive scenarios from the spec, drive the live app, file P0/P1/P2 findings, emit a `qa_verdict` - but it lived **outside** the build loop. fn-72 wires it in as an **opt-in pilot stage**: `flowctl config set pipeline.qa on` inserts a `qa` stage at the all-tasks-done juncture, so the autonomous span becomes `plan → plan-review → work → qa → make-pr`. Default **off** - with the gate off, pilot’s stage set is byte-for-byte unchanged.

* **Augments, never replaces.** The app is already up on the dev’s machine during `work`, so this is the cheap first live pass that catches obvious runtime breakage before a human opens the PR. Like everything in Flow-Next it **reduces human work agentically and surfaces problems to humans** - it does not stand in for CI/staging QA or manual QA, which still happen downstream.
* **Lean + agentic, evidence-aware.** Net-new flowctl is a single `pipeline.qa` config-key default - no new subcommand, engine, or persisted artifact. The host derives scenarios in-context and drives the local running app, reusing the existing executor. It reads `work`’s recorded evidence first and **subtracts only AC proven by a deterministic re-runnable check** (a real test/lint/build command), always live-running every runtime/UI/integration AC even when work narrated it done.
* **Surfaced, not loop-blocking.** The stage is idempotent (a `head_sha` freshness gate runs it at most once per branch head) and the pilot gate routes on `qa_outcome`, not the Ralph-guard `verdict` projection: `SHIP`/`NA`/`BLOCKED` advance cleanly, and **`NEEDS_WORK` still advances** to the draft PR - make-pr surfaces the findings in a `## Live QA` section, plus the bug-memory track and a tracker comment when the bridge is active. QA never hard-blocks the loop; merge stays the human’s + [land](https://flow-next.dev/autonomy/land/)’s decision.
* **Principled reversal.** Pilot’s “QA is never a stage” is reversed **only under the gate**; capture/interview/resolve-pr/merge/release stay forbidden for their distinct loop-ownership / consent reasons.

Cua native driver rung

The native rung of the [surface-aware driver ladder](https://flow-next.dev/guides/live-qa/#step-4-native-rung-cua-driver--computer-use--cua-sandbox) was served **only** by Computer Use (Codex CU / Anthropic Claude CU) - provider-locked, macOS/Windows-only, focus-stealing, and **never reachable on a headless / CI / Linux path**. [trycua/cua](https://github.com/trycua/cua) (MIT) is added as a detected, opt-in driver with two surfaces, never a hard dependency:

* **Cua Driver** (`cua-driver mcp`) - background computer-use on the *local* machine over an MCP server: **no focus steal**, **accessibility-tree-based** (drives structured `element_index` elements, not pixels), and provider-agnostic. On macOS the load-bearing **TCC permission split** is documented - Accessibility unlocks *driving*, Screen Recording unlocks *screenshots* - so when Screen Recording is absent the rung surfaces “AX-only evidence, no screenshot” rather than emitting an empty one.
* **Cua Sandbox** - drives an app inside a disposable VM/container (any OS), the **only** native option on a headless/CI host with no display. Opt-in per run, torn down each run; **local `lume`/QEMU/Docker is the default backend, the `cua.ai` cloud is explicit opt-in** (bills + data-egress, never auto-selected).

**Detect-and-instruct, never auto-install** - the same consent rule `/flow-next:map` applies to `clawpatch`. The base install stays zero-dependency, `agent-browser` remains the only assumed-present driver, and `flowctl` never imports Cua. The default driving path (background `cua-driver` MCP) uses only MIT components; the optional `cua-agent[omni]` (ultralytics AGPL-3.0) / OmniParser (CC-BY-4.0) extras are documented and never auto-installed. A pass still completes with no Cua installed (fall to Computer Use → documented-limitation). No new skill or command - a rung, not a re-architecture; `/flow-next:qa` accepts `cua-driver` / `cua-sandbox` as evidence `driver_rung` values with no schema change.

### 2.1.3: Resolve-pr keeps null-state threads in scope

**`/flow-next:resolve-pr` now treats only literal `true` as resolved - GitHub/GraphQL can surface a newly-created unresolved inline thread as `isResolved: null` (not just `false`), and those Codex/Bugbot findings were being silently dropped; fetch observability (counts + previews across all three feedback surfaces) is now mandatory in full mode and watch loops.**

### 2.1.2: `Done` means merged

**Tracker-sync now reserves Linear `Done` for merge-confirmed PRs - an open PR maps to `In Review`, completion-review never completes the issue, and pilot never declares `NO_WORK` for an all-done spec that hasn’t shipped.**

Detail

`Done` is a claim that the work *shipped*, so projecting it from local completion (all tasks done + completion-review `SHIP`) was a correctness bug - a spec with no PR could land on the board as `Done` and a human had to drag it back. The flow→tracker status map is now a function of `(spec status, completion_review_status, **PR-merge-evidence**)`: terminal `Done` requires a GitHub `MERGED` probe result on **every** write path (automatic touchpoints *and* a manual reconcile, which can still recover `Done` once a merge exists). An open PR projects `In Review` on make-pr’s unconditional bridge-active link path; completion-review is now a verdict comment only; and `land.merged` - active by default when the bridge is active - is the sole `Done` driver. Pilot mirrors it: an all-done spec with no merged PR routes to make-pr, or reports the new `DEFERRED_TO_LAND` verdict when an open PR exists, instead of silently collapsing to `NO_WORK`. See [Status lifecycle](https://flow-next.dev/integrations/tracker-operations/#status-lifecycle-done-means-merged-212).

### 2.1.1: Land sees clean-review comments

**Land’s `silence` merge signal now recognizes a review bot’s clean-pass *comment* (naming the reviewed commit), not just formal reviews - so a Codex-reviewed PR with no findings actually merges instead of stalling at `NEEDS_HUMAN`.**

Detail

Codex (and bots like it) only file a formal review when they have findings; a clean pass is an *issue comment* - `"Didn't find any major issues. Reviewed commit `abc1234`"` - that never reaches the reviews API land reads. So a converged-clean PR could sit unmerged forever. Under `silence`, land now also scans PR comments: a comment from an automated reviewer matching `land.cleanReviewCommentPattern` **and naming the current head SHA** counts as a head-current review. It only ever *adds* evidence - never overrides a formal review, an open thread, or a red check - and the SHA must be the current head, so a land-authored fix push still forces a fresh clean comment before merge. Set `land.cleanReviewCommentPattern` to an empty string to disable the comment path. Found dogfooding the fn-64 land. See [The merge gate, precisely](https://flow-next.dev/autonomy/land/#the-merge-gate-precisely).

### 2.1.0: Dependency projection to the tracker

**Tracker-sync now projects a spec’s `depends_on_epics` edges onto the board as blocked-by relations - on both Linear and GitHub - idempotently, provenance-tracked, and without ever clobbering a relation a human added by hand.**

Detail

**Dependency projection first shipped as skill-side provider prose.** Current releases run `flowctl tracker relate`: Linear and Jira use native directional relations; GitHub records the blocked issue as a `sub_issue` hierarchy proxy with structured degradation; GitLab uses native blocked-by when the resolved capability allows it and otherwise preserves direction in a fenced body block.

**Provenance over diff-reconcile.** Neither platform records who created a relation, so Flow tracks the edges it created in a per-spec `depRelations` ledger (native) or the fenced marker (GitHub fallback). A relation Flow can’t prove it created is never removed; a ledgered edge a tracker user deleted is deferred (`queued` receipt), never silently recreated. The `projected` flag keys off the directed tracker edge, so a relinked issue reads un-projected.

**Safe by construction.** A dependency with no linked issue surfaces a named warning and the sync proceeds; a `done` dependency keeps its relation visible but never re-gates `ready=true`; self-edges are skipped and cycles project as independent direct edges (no traversal); unreachable transport writes a `noop` receipt and never blocks. New `flowctl sync list-dep-relations` / `set-dep-relation` / `clear-dep-relation` own the deterministic ledger plumbing. See [Dependency projection](https://flow-next.dev/integrations/tracker-operations/#dependency-projection).

### 2.0.0: HTML artifact mode & render lenses

**Opt-in HTML artifact mode: capture, plan, and make-pr now also emit self-contained HTML render lenses - a spec visualizer for business and plan review, and a read-only PR review instrument - while markdown (and tracker-sync) stays 100% the source of truth.**

Detail

**Render lens, never record.** One config key - `flowctl config set artifacts.html.enabled true` (OFF by default, offered once by [`/flow-next:setup`](https://flow-next.dev/skills/setup/)) - switches the lifecycle skills into artifact mode. When active they load a shared disclosure reference carrying all generation rules plus an explicit anti-slop design contract (own instrument-panel house style, local-only fonts, zero external requests), and write self-contained single-file HTML to fixed paths under `.flow/artifacts/<spec-id>/` - never timestamped, regenerated in place, never parsed back as state, each with a staleness stamp in the footer. Mode off (the default) means zero new steps, zero token cost, zero behavior change. See [Visual Aids - specs](https://flow-next.dev/guides/render-lenses/).

**The spec lens.** One generation pathway, state-dependent rendering: [`/flow-next:capture`](https://flow-next.dev/skills/capture/) renders the spec-only business-review view (thesis, acceptance criteria with source-tag provenance chips, boundaries, decision context); [`/flow-next:plan`](https://flow-next.dev/skills/plan/) regenerates the same file with the plan layer - task dependency DAG with critical path and the R-ID → task coverage matrix. The spec markdown carries an idempotent artifact link line, replaced in place on every regeneration.

**The PR lens.** [`/flow-next:make-pr`](https://flow-next.dev/skills/make-pr/) emits a [read-only review instrument](https://flow-next.dev/guides/render-lenses/): diff-derived (never from commit messages), verified against the spec’s R-ID export - mismatches render as visibly flagged rows, warn-in-artifact, never blocking. It lands in one narrow `chore(flow): pr artifact <spec-id>` commit so the PR body’s SHA-pinned blob link resolves; `--dry-run` writes nothing, generation failure is non-fatal, and Ralph’s `PR_URL=` stdout contract is untouched.

**Lavish annotation (optional).** [`lavish-axi`](https://www.npmjs.com/package/lavish-axi) is detected on PATH and never required: spec artifacts open as browser annotation sessions, and feedback maps to edits of the markdown source followed by lens regeneration. Pull-only and session-spanning (annotations queue in `~/.lavish-axi/state.json` and survive agent death). The PR lens never enters the annotate loop, and autonomous runs generate but never poll.

**Breaking.** The deprecated `planSync.crossEpic` config alias (1.x deprecation, readable through the 1.x line) is removed - use `planSync.crossSpec`.

### 1.14.0: Land: the autonomous ship loop

**New `/flow-next:land` skill - a cadence-driven, fully autonomous babysitter that takes the build loop’s draft PRs the rest of the way: CI kept green, automated reviews converged via resolve-pr, a gated explicit merge, spec close, and your project’s own release process - closing the lifecycle end to end.**

Detail

**The tick.** Each [`/flow-next:land`](https://flow-next.dev/autonomy/land/) invocation discovers the open PRs the build loop authored (spec `branch_name` match **and** the make-pr breadcrumb - both signals required before any mutation; hand-opened PRs are never touched), walks each through a read-only gate tree - CI tri-state over **all** checks, a reviewer patience window anchored to the last push (`land.patienceMinutes`, default 30), unresolved threads, the review signal, `mergeStateStatus` - and takes at most one action class per PR: a bounded CI fix (`land.ciFixBudget`, default 3, with a durable `flow-next:needs-human` label on exhaustion), a [`/flow-next:resolve-pr`](https://flow-next.dev/skills/resolve-pr/) dispatch, a mechanical rebase (any conflict hunk → honest `BLOCKED`), or the merge. Every tick ends with `LAND_VERDICT=<MERGED|RELEASED|FIXING_CI|AWAITING_REVIEW|RESOLVING|BLOCKED|NEEDS_HUMAN|NO_WORK> prs=<n> pr=<url|-> reason="…"` - worst severity across PRs, last line of output. Drive it on a cadence: `/loop 30m /flow-next:land`.

**The merge gate.** Land is the one confined exception to flow-next’s “no `gh pr merge` from skills” rule. It flips the draft to ready and merges **explicitly** - `gh pr merge --squash --delete-branch --match-head-commit`, never `--auto` - only after CI is green, threads are addressed, and `land.reviewSignal` is satisfied: `silence` (default - an automated review present + zero unresolved threads + the window elapsed; built for bot reviewers that never file formal APPROVEs), `approve`, or a named reviewer login. No automated review ever and no signal configured → it never merges unreviewed (`NEEDS_HUMAN`).

**The tail.** After merge: `flowctl spec close` (the build loop never re-selects merged work), the opt-in `tracker.perEvent.land.merged` touchpoint (issue → terminal state + verdict comment), then release-follow of **your project’s own** release docs (`RELEASING.md` et al.) with an idempotency probe - or stop at merge. A merged-but-unclosed spec re-enters idempotently. `--dry-run` reports the full gate classification with zero mutations.

**Autonomous resolve-pr.** [`/flow-next:resolve-pr`](https://flow-next.dev/skills/resolve-pr/) now honors the `mode:autonomous` token (plus `FLOW_AUTONOMOUS=1` env): needs-human cases become `NEEDS_HUMAN:` report lines instead of a blocking question, and the run ends with the machine-readable `RESOLVE_PR_VERDICT=<RESOLVED|PENDING|NEEDS_HUMAN> threads=<n> fixed=<n> needs_human=<n>` line land gates on. The 2 fix-verify cycle bound is unchanged, and the `land.*` config keys ship with seeded flowctl defaults.

Land was, fittingly, the first spec [pilot](https://flow-next.dev/skills/pilot/) drove end-to-end. See [Going Autonomous](https://flow-next.dev/autonomy/going-autonomous/) for the three-loop picture.

### 1.13.0: Pilot: host-driven autonomous loop

**New `/flow-next:pilot` skill - a single-tick build-loop conductor that advances one ready spec by one pipeline stage (plan → plan-review → work → make-pr) per invocation and ends with a machine-greppable `PILOT_VERDICT` line, so your host’s `/loop` or `/goal` owns the iteration instead of an external shell script.**

Detail

**The tick.** Each [`/flow-next:pilot`](https://flow-next.dev/skills/pilot/) invocation selects the first `open` + [`ready`](https://flow-next.dev/guides/writing-specs/#the-ready-flag) spec with satisfied dependencies and no other-actor claims, classifies its stage from flowctl state, dispatches exactly one existing stage skill autonomously, verifies advancement (flowctl review-status fields + status transitions; a gh-confirmed OPEN PR URL for make-pr), and prints the terminal verdict: `PILOT_VERDICT=<ADVANCED|NO_WORK|BLOCKED|NEEDS_HUMAN> spec=<id> stage=<stage> reason="<one line>"`. `/goal` validators are transcript-blind, so the evidence is echoed into the conversation and stop conditions are phrased against the grammar - e.g. `/goal keep running /flow-next:pilot until it prints PILOT_VERDICT=NO_WORK, or stop after 20 turns`.

**Drivers.** Claude Code `/goal` (v2.1.139+), Claude Code `/loop` (v2.1.72+; loops expire after 7 days), and Codex `/goal` (opt-in `[features] goals = true`, CLI ≥ 0.128.0, plain-text objective - no $skill-in-goal syntax). Caps and budgets belong to the driver; a tick has no timeout machinery. For unattended runs the `rp` backend works headlessly while the Repo Prompt app is running on the same Mac; on machines without it use `--review=codex`, `--review=copilot`, or `--review=none`.

**Autonomous sub-skills.** [`plan`](https://flow-next.dev/skills/plan/), [`work`](https://flow-next.dev/skills/work/), and [`make-pr`](https://flow-next.dev/skills/make-pr/) now honor a `mode:autonomous` token (plus `FLOW_AUTONOMOUS=1` env) that suppresses user questions and picks safe defaults - work branches deterministically, make-pr forces a draft PR and hard-errors instead of prompting. Deliberately distinct from `FLOW_RALPH`: no ralph-guard hooks, no receipt choreography.

**Don’t-thrash.** A spec that fails to advance on two healthy ticks is taken out of selection (`flowctl spec unready`) with the reason in the `BLOCKED` verdict; re-blessing via `flowctl spec ready` clears its strikes. Pilot and [Ralph](https://flow-next.dev/autonomy/ralph/) are alternative drivers, never nested - pilot refuses to run under `FLOW_RALPH`.

### 1.12.0: Spec readiness signal

**A spec now carries a human-owned `ready` flag - the “complete enough to hand to an agent” gate that autonomous loops will consume - set via `flowctl spec ready`, projected one-way from your tracker (`tracker.readyState`), and surfaced through adoption-gated prompts in capture/interview/plan; invisible until you opt in.**

Detail

**The flag.** `flowctl spec ready <id>` / `spec unready <id>` toggle a `ready` boolean on the spec record (default `false`) - orthogonal to `status` (a ready spec stays `open` through planning and work), human-owned or tracker-projected, **never agent-inferred**. Both verbs are idempotent (no write, no `updated_at` bump when the flag already matches), and the on-disk flag is **lazy** - written only after a toggle actually changes state, so non-adopters’ sidecars stay byte-identical. Every JSON read surface (`show`, `specs`, `list`) emits an explicit `"ready": <bool>`, and ready specs carry a `[ready]` badge in listings (badge only when set - no draft-noise). See [Before planning - the ready flag](https://flow-next.dev/guides/writing-specs/#the-ready-flag).

**Tracker projection.** For tracker-connected repos, the `/flow-next:tracker-sync` discovery ceremony asks one optional, skippable question: which tracker workflow state means “ready for work”? (Linear: a workflow-state **name**, matched case-insensitive - names, not `state.type`; GitHub: a **label**, pre-created idempotently - present ⇒ ready, absent ⇒ not ready.) Every pull-side sync then projects that state onto the local `ready` flag - **one-way, tracker → local, tracker authoritative** - with change-only event-tagged receipts and graceful stale-config degradation (warn + `noop` receipt + flag untouched + sync continues). See [Readiness projection](https://flow-next.dev/integrations/tracker-operations/#readiness-projection).

**Adoption-gated prompting.** One in-use gate (≥1 ready spec OR `tracker.readyState` configured) governs every new prompt - non-adopters see zero new questions anywhere. [`/flow-next:capture`](https://flow-next.dev/skills/capture/) and [`/flow-next:interview`](https://flow-next.dev/skills/interview/) offer an optional end-of-authoring “Mark ready?” consent (default keep-draft; gated off when the tracker is authoritative). [`/flow-next:plan`](https://flow-next.dev/skills/plan/) soft-checks readiness before the scout fan-out - warn, never block, default proceed. `capture --rewrite` resets a previously-ready spec to draft (a full re-authoring re-opens the blessing); interview refinement never auto-resets.

New regression suite (`test_spec_ready.py`) wired into CI; Codex mirror regenerated with all three net-new prompts verified.

### 1.11.0: Tracker-sync forcing + self-improving glossary

**Tracker lifecycle touchpoints can no longer silently not fire - receipts are event-tagged and work/capture/make-pr end every run with a read-only `sync check` + one bounded retro-fire - and the glossary now compounds through normal work (prime seeds it, capture adds to it, plan/work/review actually read it).**

Detail

**Tracker-sync became observable and forcing.** This release added event-tagged receipts, the read-only `flowctl sync check` audit, one bounded retro-fire, and the mandatory `Tracker sync:` summary slot. Its original rule that flowctl contained no tracker mutation code was superseded by the deterministic `flowctl tracker` command surface; the receipt and audit semantics remain current.

**Self-improving glossary.** The same principle - the system gets better as you use it, never via a manual ceremony - applied to project vocabulary: `/flow-next:prime` **seeds** GLOSSARY.md from the repo when it’s absent or a husk (\~10-20 evidence-backed terms, read-back gated, never rewrites a populated glossary), `/flow-next:capture` **joins interview as a writer** (new conversation-surfaced terms offered at read-back), and the **read path widens to where wrong-concept errors get built**: repo-scout / context-scout surface request-matched terms (max 5, budget-capped), the work worker’s re-anchor reads task-relevant terms, and impl-/plan-review prompts gain a Vocabulary criterion. Every gate is `total_terms == 0 → silent skip`. The compounding surfaces (memory, glossary, decisions, strategy) now have a dedicated [Self-improving](https://flow-next.dev/understand/how-it-compounds/) page, a STRATEGY.md track, and a “Self-improving” pillar in the redesigned six-pillar hero grid.

New regression suites (`test_sync_check.py`, `--event` coverage in `test_tracker_receipts.py`) wired into CI; Codex mirror regenerated.

### 1.10.2: Homepage points at flow-next.dev

**The plugin `homepage` (Claude marketplace + the Claude/Codex plugin manifests, and the Codex `websiteURL`) now points at `https://flow-next.dev` instead of the stale `mickel.tech/apps/flow-next`.** `.cursor-plugin` was already correct; the rest were drift. The `flow-next-tui` package homepage and the README “Visual overview” doc row were aligned too. `author.url` / `owner.url` (personal site / GitHub) are unchanged.

### 1.10.1: cp1252 / non-UTF-8 robustness

**`flowctl` impl-review and console output no longer crash on a non-UTF-8 source subtree or a legacy console codepage (e.g. Windows cp1252).**

Detail

**Read side (#167).** `flowctl copilot impl-review` could abort with `UnicodeDecodeError` on a repo containing a non-UTF-8 file *anywhere* in the tree. `find_references()` (behind `gather_context_hints`) ran `git grep` repo-wide and decoded hits with a hard `encoding="utf-8"` and no `errors=` - so a single legacy cp1252 file (e.g. a German C/C++ subtree carrying `ü` / `ä` / `ö` / `ß`) aborted context gathering even when every file you actively edited was UTF-8. The collector now captures `git grep` output as **bytes** and decodes with `errors="replace"`. Reported with measured data by VGottselig (304 of \~5400 C/C++ files non-UTF-8 in a large Windows CAD codebase).

**Write side (#167).** `flowctl` now forces its own stdout/stderr to UTF-8 at startup, so non-ASCII output (`→`, umlauts) no longer aborts on a legacy console codepage (`UnicodeEncodeError: 'charmap' codec can't encode character '→'`). This removes the need for the `PYTHONIOENCODING=utf-8` workaround.

**`/flow-next:work` Verify-Completion recovery.** When the host drops a long-running worker’s completion report, phase 3d no longer blocks waiting for a result that will never arrive: it diagnoses from ground truth (`flowctl show` + `git log` + `git status`) and classifies - already done → plan-sync; code present but unfinished → re-anchoring continuation worker; nothing landed → retry.

### 1.10.0: Eval-optimized scout agents

**8 read-only scout/analyst agents got a feature-preserving output budget - \~40-71% leaner output into the planner/work-loop context, with accuracy held (proven by per-target evals + an end-to-end smoke).**

Detail

Rolled the external “autoresearch” eval loop (baseline → one mutation → keep-if-better ratchet) across the read-only agents whose free-form output flows into the planner and work-loop context. Each gained a hard **output budget** - the reductions are at *runtime* (the rendered output), not in prompt size:

* `repo-scout` (83→100% on its eval set, \~40-50% leaner) · `context-scout` (60→93%, \~60-70%; dropped the prescribed Code-Signatures block) · `flow-gap-analyst` (\~50-70%, 26/27 gaps held) · `quality-auditor` (\~63%) · `spec-scout` (No-Relationship → count, scale-robust) · `docs-scout` (\~48-69%) · `github-scout` (\~71%, the biggest) · `practice-scout` (\~52%).

**Feature-preservation is the guarantee, not a hope.** A mutation was kept only if a per-target coverage/accuracy eval held (the ratchet): grounding (`context-scout`’s cited paths `test -f`-verified against a real 442k-LOC app repo), findings (`quality-auditor` against a 7-planted-issue testbed - Major bug + all slop still caught, clean code stays ✅), gaps (per-input answer keys), and docs/APIs/gotchas (the “pointer-not-paste” rule: name the API inline, drop code blocks, the link carries depth). The leaner research scouts even surfaced *extra* real issues a verbose baseline missed (a current CVE; an extra security gotcha).

**End-to-end verified.** The optimized scouts → a planner produced a correct, ship-quality build plan for a deliberately hard, cross-cutting feature (org-scoped rate limiting) reading *only* the budgeted scout output - features preserved at the *consumer* level, not just at scout-output level. The method lives in `agent_docs/optimizing-skills.md`.

Also: `/flow-next:make-pr` shed stale build-scaffolding archaeology (render output byte-identical). `/flow-next:capture` is unchanged - a trim was tried and *reverted* (the ratchet caught a routing regression), and its no-silent-overwrite guard was verified intact. No `flowctl` logic touched; Codex mirror regenerated.

### 1.9.1: Cursor setup detection + tracker merge-base

**`/flow-next:setup` now detects Cursor instead of mis-treating it as Codex, and a `comment`-first tracker auto-link snapshots its merge-base so a later sync can’t fast-forward over tracker edits.**

Detail

**Cursor setup detection.** Setup keyed platform off plugin-root env vars (`DROID_PLUGIN_ROOT` → Droid, `CLAUDE_PLUGIN_ROOT` → Claude Code, else → Codex). Cursor exposes neither, so it fell into the Codex branch - writing the `$flow-next-*` Codex command syntax + running `.codex/` setup, while the installer advertises `/flow-next:*`. Setup now branches on **`CURSOR_AGENT` + the `.cursor-plugin/plugin.json` manifest + no `codex/` mirror dir** → `PLATFORM=cursor`, applied at every platform-branch point (detection, docs-status template, the Docs question, and the write mapping): it writes the `/flow-next:plan` snippet into AGENTS.md (which Cursor reads), resolves `flowctl` via `.flow/bin/flowctl`, and skips the Codex-only `.codex/` copy. The triple guard is hardening against `CURSOR_AGENT` being **inherited by child processes** (Codex launched from a Cursor shell) and against the shared repo source tree carrying all manifests - and the installers now `--delete-excluded` / `Remove-Item` excluded paths so the `codex/`-absence proof holds on re-install. Hardened across five rounds of automated cross-model review.

**Tracker merge-base snapshot.** When the first lifecycle touchpoint for an unlinked spec was a `comment` op, create-if-unlinked attached the tracker id but didn’t snapshot the merge-base - leaving the issue base-less, so a later body sync could fast-forward and silently overwrite tracker-side edits. The auto-create path now `set-merge-base` (both halves) + `set-last-synced` at create time (the written issue body is the render, so the base is exact).

### 1.9.0: Cursor install + tracker auto-link

**Flow-Next now installs into Cursor via a one-shot local plugin (`./scripts/install-cursor.sh`), and a lifecycle event on an *unlinked* spec now creates + links the tracker issue first instead of silently no-opping.**

Detail

**Cursor support (macOS / Linux + Windows).** Cursor has its own `.cursor-plugin/` plugin namespace and does **not** auto-read Claude Code plugins the way Grok Build does, so Flow-Next ships a Cursor-native manifest plus a one-shot installer on **both** platform families - `./scripts/install-cursor.sh` (rsync) and `install-cursor.ps1` (robocopy). Both copy the plugin into `~/.cursor/plugins/local/flow-next` as a **real directory** (Cursor’s loader rejects symlinks escaping `~/.cursor/`), exclude the Codex mirror + tests, and are a re-runnable snapshot. **Verified end-to-end, multi-agent flows included** - a full `/flow-next:plan` fanned out the scouts in parallel and drove `flowctl` to create the spec + tasks; `flowctl` resolves via the project-local `.flow/bin/flowctl`. Caveats: no grouped plugin card and the slash autocomplete under-lists the commands (both cosmetic - they run when typed), and **Ralph autonomous mode is unsupported** (Cursor’s `afterFileEdit` / `beforeShellExecution` hooks don’t map to the Claude `PreToolUse` + `Bash|Execute` matchers the Ralph guard relies on). See [Install → Cursor](https://flow-next.dev/install/#cursor).

**Tracker create-if-unlinked.** Previously only `capture` flow-first-pushed an unlinked spec; every other lifecycle touchpoint (`plan` / `interview` / `work` / `make-pr` / `resolve-pr` / `completion-review`) no-op’d when the spec had no tracker id, so starting a spec with `/flow-next:plan` left it orphaned from Linear / GitHub. Create-if-unlinked is now part of the `flowctl tracker sync` facade. `unlink` remains the only operation that no-ops on an unlinked spec; provider failures return a structured class.

### 1.8.0: Live-app QA pass (`/flow-next:qa`)

**New `/flow-next:qa` drives the *running* app like an unforgiving real user - deriving its test scenarios straight from the spec, filing P0/P1/P2 findings with evidence, and ending with a YES/NO ship verdict.**

Detail

Every other Flow-Next review is static - `impl-review`, `spec-completion-review`, `quality-auditor`, `code-review` read code or specs. `/flow-next:qa` is the live-app gate: it drives the deployed app via [Flow-Next Drive](https://flow-next.dev/skills/flow-next-drive/)’s surface-aware driver ladder (it never re-implements driving) and is **forbidden from marking PASS by reading source** - a scenario passes only on captured evidence (screenshot / console / URL).

The differentiator vs spec-less QA tools: scenarios derive **directly from the spec** - acceptance criteria → scenarios, R-IDs → a coverage table, boundaries → what *not* to test, decision context → expected behavior - so the host already encodes intent instead of reconstructing it. Findings feed the **bug memory track** (`track: bug`, with overlap dedup) and can be promoted to specs/tasks. The pass emits a `qa_verdict` receipt with four outcomes (`SHIP` / `NEEDS_WORK` / `BLOCKED` / `NA`) projected onto the review-receipt enum, so it can feed `spec-completion-review` - “does the *live app* satisfy the AC, not just the code?”. Runs interactively and autonomously; **not a hard Ralph-block**; opt-in tracker verdict-post via `tracker.perEvent.qa`. Requires a live deploy + a driver - with neither it surfaces a `BLOCKED` verdict rather than failing; a spec with no driveable UI yields a clean `NA`. The QA discipline (P0/P1/P2 taxonomy, evidence rules, session hygiene) is a lean, credited borrow from Ray Fernando’s [running-bug-review-board](https://github.com/RayFernando1337/rayfernando-skills) (Apache-2.0).

### 1.7.1: Codex delegation: cheaper on non-Claude hosts

**The opt-in Codex delegation reference no longer loads into a Codex / Droid / OpenCode session - the host-platform check moved into the cheap value-check, so delegation short-circuits *before* the \~45k reference is ever read. Byte-identical for Claude Code users.**

### 1.7.0: Optional Codex delegation for `/flow-next:work`

**`/flow-next:work` gains an opt-in `delegate:codex` mode that offloads code implementation to `codex exec` (gpt-5.5/medium) while Claude keeps orchestration, review, and all git - a second efficiency lever that offloads *work*, not prompt size.**

Detail

When activated (`delegate:codex` arg or `flowctl config set work.delegate codex`), the Claude host stays the orchestrator - plan-reading, review, all git, and decisions - and delegates the token-heavy implementation to `codex exec`. Default model **gpt-5.5**, default effort **medium**, with proven per-batch risk escalation. It’s a different lever than prompt-trimming: it moves implementation tokens to a separate Codex budget.

Strictly opt-in and progressive-disclosure: with delegation **off** (the default) the work flow is byte-identical to before - one cheap `config get`, zero new steps. All mechanics live in a reference loaded only when active. The safety surface is the headline work: a one-time sandbox consent; a value-aware recursion guard (the flow-next `CODEX_SANDBOX=auto` review knob never trips it); mandatory `--output-schema` with MCP isolation (`--ignore-user-config`); a deterministic result classifier + sanitized scoped rollback that never touches `.flow/`; “Codex never touches git” enforced by a post-run HEAD-unchanged assertion (not just the prompt); and a host-owned 3-strike circuit breaker that always falls back to standard mode. The `ralph-guard` PreToolUse hook is rebuilt to a tokenized argv allowlist that admits only the full canonical delegation shape. Runs in interactive **and** Ralph mode (consent pre-granted in config for headless). Configure via the `work.delegate*` keys; see the flowctl reference.

### 1.6.0: Tracker-sync is opt-out by default

**Hooking up the tracker bridge via `/flow-next:tracker-sync` now activates the whole lifecycle pipeline by default - you opt *out* of events, not in.**

Detail

Previously every `tracker.perEvent.*` touchpoint defaulted `off`, so after the discovery ceremony you had to opt each lifecycle event in by hand. That inverted the intent - connecting a tracker means you want it kept in sync. The discovery ceremony now activates every event on confirmation: capture / interview / plan → `reconcile`, work.firstClaim → `push`, work.done / makePr / resolvePr → `comment`, completionReview → `reconcile`. Exclude events at ceremony time, or turn any off later with `flowctl config set tracker.perEvent.<event> off`.

The accidental-enable guard is preserved: the config *schema* default for each leaf stays `off`, so a bare `tracker.enabled=true` set by hand or a script - without running the ceremony - fires no lifecycle-event sync (make-pr’s unconditional PR↔issue link is the one exception). Activation is ceremony-gated, not flag-gated. No config-schema change; docs updated across every surface.

### 1.5.3: Tracker-sync receipts auto-ignored

**`.flow/sync-runs/` (per-run tracker-sync receipts) is now auto-gitignored, and flowctl’s managed `.flow/.gitignore` self-upgrades so existing repos pick up new patterns on the next `init`.**

### 1.5.2: Tracker-sync projects the full spec

**`/flow-next:tracker-sync` now guarantees the issue body mirrors the **entire** spec - a render guardrail stops the host agent from pushing a summarized body instead of the full projection.**

### 1.5.1: Setup docs fix + Windows de-flake

**Fresh `/flow-next:setup` now ships the tracker-sync CLI reference it dropped in 1.5.0, a CI guard keeps the bundled template and the dogfood copy in lockstep, and a flaky Windows migration race is fixed.**

Detail

* **`usage.md` shipped without tracker-sync docs.** 1.5.0 added the `flowctl sync` / `--tracker-first` command block to the repo’s lived-in `.flow/usage.md` but not to the **bundled** template `/flow-next:setup` actually copies - so every fresh setup documented the whole CLI except the tracker-sync bridge that shipped in the same release. The canonical template is now byte-synced (Codex mirror regenerated).
* **Drift guard so it can’t recur.** A new parity test hard-asserts `.flow/usage.md` ≡ its canonical setup template (and `.flow/templates/spec.md` ≡ `templates/spec.md`) across the ubuntu/macos/windows CI matrix. Edit the dogfood copy and forget the template → CI fails instead of consumers getting stale docs.
* **Flaky Windows CI fixed.** Parallel `migrate-rename` could leave a concurrent writability probe’s temp file (`.rw-probe-*.tmp`) visible to the backup copy’s directory scan, then have it vanish before the copy opened it (`FileNotFoundError`) - a TOCTOU that Windows’s slower unlink widened. The backup copy now skips those transients and tolerates any file that disappears mid-copy.

### 1.5.0: Tracker sync bridge

**New `/flow-next:tracker-sync` projects a spec to a Linear or GitHub issue and reconciles body, status, and comments two-way - projection, not coordination - and a tracker key (`wor-17`) now resolves as a spec id everywhere `fn-NN` does.**

Detail

`/flow-next:tracker-sync` mirrors a `.flow/specs/<id>.md` spec onto a tracker issue (Linear first, GitHub next). **Projection, not coordination:** the spec stays the single source of truth and the quality layer; the tracker is a co-editable mirror that **never drives flow state or spawns agents** (contrast OpenAI Symphony, where the board is the control plane). It is distinct from `/flow-next:sync`, which is plan-sync. New docs: [Tracker Sync](https://flow-next.dev/integrations/tracker-sync/).

* **Discovery ceremony** (detect → surface → ask → never-assume) probes a Linear MCP / `LINEAR_API_KEY` / GitHub auth / a Jira host and writes `tracker.*` config only on confirmation (env > config > ask, the same ladder as `flowctl review-backend`). The bridge is **off until explicitly enabled** and active when `tracker.enabled == true` **or** `tracker.type ∈ {linear, github}`.
* **Historical provider implementation.** The first release encoded provider request choices in skill prose. Current releases persist normalized resolution facts and capabilities under `tracker.resolved`, execute deterministic operations through `flowctl tracker`, and return structured error classes instead of a prose-selected no-op path.
* **Hybrid id model:** tracker-first specs are canonically `wor-17-slug` (tasks `wor-17-slug.M`; bare `wor-17` resolves); flow-first specs keep `fn-NN` plus a resolvable `WOR-17` display alias. `show` / `work` / `plan wor-17` resolve case-insensitively, the native `fn-` scheme is reserved, one tracker team per repo, and **ids never rename** on link. See [Spec & task ids](https://flow-next.dev/reference/spec-schema/#spec-and-task-ids).
* **Seven lifecycle skills gain opt-in tracker-sync touchpoints** - capture, interview, plan, work (first-claim + done), make-pr, resolve-pr, spec-completion-review. Each `tracker.perEvent.*` leaf defaults `off`; the no-tracker workflow is unchanged and remains the documented default.
* **PRs are Diffs-ready.** When the bridge is active, `make-pr` **unconditionally** links the new PR to its tracker issue (no `makePr` opt-in). For Linear it writes a **non-closing** `Ref WOR-N` (plus a rich `attachmentLinkURL` on the GraphQL transport) so the PR renders as a [Linear Diff](https://flow-next.dev/integrations/tracker-sync/#linear-diffs-review-the-pr-inside-the-issue) inside the issue; for GitHub it is a native `Refs #N` cross-link. Non-closing is deliberate - merge never auto-completes the issue, spec-completion-review owns Done.
* **`make-pr` now creates the PR autonomously** - no confirm gate. Invoking the skill is the intent; the body is deterministic; the default is a reversible smart-draft. `--dry-run` prints the body without creating, `--ready` / `--draft` set the draft state.
* **Setup proposes the bridge.** `/flow-next:setup` never touches tracker config (keeping the zero-dep base clean) and now proposes running `/flow-next:tracker-sync` as an optional next step when it finishes.
* **Ralph-safe:** every run emits a receipt; genuine conflicts queue to the review deferred-findings sink rather than block. An `always-ask` tiebreak resolves to *queue* in autonomous mode.

Sync-engine shape (discovery ceremony, per-item `lastSyncedAt`, surface-diffs-never-overwrite) adapted from Ray Fernando’s [`rayfernando-skills`](https://github.com/RayFernando1337/rayfernando-skills) `running-bug-review-board` `issue-trackers.md` (Apache-2.0). Thank you, Ray.

### 1.4.0: `flow-next-drive` surface-aware automation

**The `browser` skill is renamed `flow-next-drive` and rebuilt as a surface-aware driver ladder - it detects the UI surface (web, Chromium-backed desktop, or true-native app) and picks the best available driver, degrading gracefully.**

Detail

The skill is no longer hardwired to a single browser driver. It detects the target surface and branches: (a) **web app** → web ladder; (b) **Chromium-backed desktop app** (Electron / Windows WebView2) → the *same* web ladder, attaching over CDP to the app’s remote-debugging port (`agent-browser --cdp <port>` / `--auto-connect`; chrome-devtools-mcp `--browser-url`); (c) **true-native / non-CDP surface** (macOS AppKit/SwiftUI, or a webview exposing no CDP - e.g. macOS WKWebView / Tauri-on-macOS) → Computer Use. All surfaces share one universal flow (`observe → snapshot → act on fresh refs → verify → capture → release`); only the actuation differs.

The **web ladder**, in priority order: **agent-browser** (default rung, the only assumed-present driver, CDP-based + headless-safe) → **chrome-devtools-mcp** (auto-wait + attach-to-real-signed-in-Chrome) → **Playwright** → **cursor-ide-browser** MCP → **manual** screenshot relay. The **native rung** is Computer Use - driver-agnostic across **Codex Computer Use** and/or **Anthropic “Claude” Computer Use** (the API `computer` tool via its own harness); detected and optional, never a hard dependency, never on a headless path. When no Computer Use is present, a Chromium-backed app still drives via the web-ladder CDP attach; a genuinely native app documents the limitation rather than fails. The existing agent-browser references fold into the default-rung reference - no capability regression.

The driver ladder + universal-flow structure is adapted from Ray Fernando’s [`rayfernando-skills`](https://github.com/RayFernando1337/rayfernando-skills) `running-bug-review-board` skill (Apache-2.0). **Migration:** `/flow-next:browser` is gone - the skill is now `/flow-next:flow-next-drive` (and `flow-next-drive` on the Codex mirror, fixing the prior `agent-browser` rename that collided with the user’s global `agent-browser` skill and Codex-native browser skills). An orphaned `browser` / `agent-browser` skill in a cached install auto-clears within \~7 days or immediately by deleting the stale cached marketplace directory under `~/.claude/plugins/cache/<marketplace>`. See the [Drive skill page](https://flow-next.dev/skills/flow-next-drive/).

### 1.3.4: Review-output R-ID suffix fix

**The review-output R-ID parser now preserves single-letter suffixes (`R4a` / `R4b`) - they were being silently dropped from the coverage gate and fix-loop targeting.**

Detail

`parse_unaddressed_rids` read R-IDs from a reviewer’s `Unaddressed R-IDs:` summary line (`_extract_rids`) and from the `## Requirements coverage` table fallback using bare `\bR(\d+)\b`. fn-49.1 (1.2.1) taught the *spec* acceptance-criteria parser the `R\d+[a-z]?` suffix form but left this *review-output* path behind - so a reviewer reporting `Unaddressed R-IDs: [R4a, R4b]` parsed to `['R5']`, dropping the suffixed IDs from the R-ID coverage gate and fix-loop targeting. Both review-output regexes are now `\bR(\d+[a-z]?)\b`, in lockstep with the spec parser; multi-letter suffixes (`R4ab`) and separators (`R-4`) stay rejected. New `test_unaddressed_rids_parser.py` (10 cases) wired into the ubuntu/macos/windows CI matrix. Surfaced by a live impl-review A/B - the current review prompt caught it (the experimental slop rubric being tested was shelved as unproven).

### 1.3.3: Scout `flowctl` fallback

**Scouts fall back to the bundled `.flow/bin/flowctl` so `.clawpatch/` feature enrichment fires even when dispatched subagents don’t inherit the plugin-root env var.**

Detail

When `repo-scout` / `context-scout` run as dispatched subagents they may not inherit `CLAUDE_PLUGIN_ROOT` / `DROID_PLUGIN_ROOT`, which left their Step 0 `flowctl repo-map list --json` call resolving to a broken `/scripts/flowctl` - the scout then silently grep-degraded and `features_anchored` never fired even with a populated `.clawpatch/`. Both scouts now fall back to the bundled `.flow/bin/flowctl` (`[ -x "$FLOWCTL" ] || FLOWCTL=".flow/bin/flowctl"`), so a `/flow-next:setup`-installed repo resolves regardless of subprocess env - across Claude Code, Factory Droid, and Codex. Also makes `sync-codex.sh`’s agent-body fallback injection idempotent (no duplicate line in the Codex mirror) and the scout-fallback contract test hermetic (runs in a throwaway git repo so a local dogfood `.clawpatch/` can’t break it). Surfaced by full live end-to-end testing - mapping Flow-Next’s own repo via `--source=agent` (codex) produced 9 features, then the scout enrichment was exercised against them.

### 1.3.2: Heuristic-0-features hint

**`/flow-next:map` now explains why the provider-free mapper found 0 features on an unconventional repo, and points at `--source=auto|agent`.**

Detail

Live-testing on Flow-Next’s own repo showed clawpatch’s heuristic detectors target conventional app/framework layouts (npm bins, Next.js routes, Python packages, Rails / Laravel / Django, Go / Rust, JVM, .NET, SwiftPM, Phoenix); a plugin + markdown-skill + `flowctl.py`-CLI + bun-TUI repo matches none, so heuristic returns 0 features while clawpatch flags coverage as “weak.” The Phase 5 summary previously printed a silent “Mapped: 0 feature(s)” - it now explains the conventional-layout targeting and points at `--source=auto` (heuristic-first, provider only if weak) or `--source=agent` (always provider-backed; needs `CLAWPATCH_PROVIDER` + tokens). For reference, `--source=agent` via codex produced 9 well-scoped features for Flow-Next’s repo (Flowctl CLI Core, Ralph Guardrails, Flow Memory System, TUI shell/theme/integration, Plugin Packaging, Strategy/Docs). Also root-ignores `.clawpatch/` in the dev repo so the local feature map never gets staged. See the [skill page](https://flow-next.dev/skills/map/) for the conventional-vs-unconventional note.

### 1.3.1: PNPM\_HOME hint reword

**The `/flow-next:map` install hint is now conditional and pnpm-version-agnostic - it no longer presumes an install already happened.**

Detail

Live-testing 1.3.0 (clawpatch never installed, pnpm 10) showed the hint presumed an install had already happened (“install succeeds but PATH unchanged”) and attributed the PATH wiring to pnpm v11 specifically - misleading for a first-time or pnpm-10 user whose global bin sits at `~/.local/share/pnpm`. Now conditional and version-agnostic: “if you already ran `pnpm add -g clawpatch` and still see this, pnpm installs globals under `$PNPM_HOME` and needs a one-time `pnpm setup`.” Same correction in `docs/troubleshooting.md` + the skill page above. Logic unchanged.

### 1.3.0: `/flow-next:map` skill

**New opt-in `/flow-next:map` skill wraps [clawpatch](https://github.com/openclaw/clawpatch) for a semantic feature index that scouts and `/flow-next:prime` can read - `flowctl` core stays zero-dep.**

Detail

Wraps `clawpatch map` to produce a semantic feature index (`~20` languages, persisted at `.clawpatch/features/*.json`, Zod-validated `schemaVersion: 1`). Opt-in convenience throughout - `flowctl` core never imports or requires clawpatch; the skill is the only flow-next surface that touches it. Default invocation is provider-free (`--source heuristic`, zero LLM calls, deterministic mapper); `--source auto|agent` flows through as passthrough. Missing binary → skill prints `pnpm add -g clawpatch` install instructions verbatim and exits cleanly (no auto-install); pnpm-installed-but-not-on-PATH → skill prints the `PNPM_HOME` `bin/` hint. Single-source `SUPPORTED_CLAWPATCH=">=0.4.0 <0.5.0"` version pin lives in skill prose; outside-range → one-line stderr warning + degrade, never block. Ralph-block (decline-to-run, no receipt write) under `FLOW_RALPH=1` / `REVIEW_RECEIPT_PATH`.

Companion `flowctl repo-map list / show / since-ref` reader subcommands parse the index directly; readers bypass `ensure_flow_exists()` and gate on `.clawpatch/` presence instead so the prime detection branch works without special-casing. `since-ref` uses three-dot `<ref>...HEAD` semantics (fixed during PR review - two-dot `<ref>..HEAD` polluted overlap results with upstream-only advancement). `repo-scout` and `context-scout` call `flowctl repo-map list --json` as Step 0 when `.clawpatch/` is present and emit an optional `features_anchored: [...]` field with a `last_mapped` staleness timestamp; scouts remain useful with the existing grep/glob path when `.clawpatch/` is absent (fallback contract is load-bearing). `/flow-next:prime` adds a `DE7` informational sub-criterion under Pillar 5 (Dev Environment) surfacing `/flow-next:map` in Top Recommendations when the index is missing - pillar count stays at 8, scored criteria stay at 48, total criteria become 48 → 49.

The feature index is **local-only by design**: `.clawpatch/.gitignore` skeleton is `*` + `!.gitignore`, so the index is regenerable-per-developer rather than committed (avoids PR review noise + merge conflicts on a pre-1.0 schema; full sharing-contract trade-off table on the [skill page](https://flow-next.dev/skills/map/)). Docs: `platforms.md` gains an “Optional skill requirements” section; `troubleshooting.md` gains a clawpatch-failure-modes section. `GLOSSARY.md` entries added for “feature map” and “features\_anchored”. CI matrix wires `test_repo_map.py` (22 tests) + `test_scout_fallback_contract.py` (14 tests) + `test_pnpm_home_hint_prose.py` (5 tests) + `map_smoke_test.sh` (75 cases) on ubuntu / macos / windows using checked-in fixtures (no Node 22+ or clawpatch needed in CI).

### 1.2.1: make-pr parser fixes

**Two `spec export-cognitive-aid` parser bugs fixed so `/flow-next:make-pr` bodies stop silently dropping content.**

Detail

(1) The R-ID parser regex `R\d+` was extended to `R\d+[a-z]?` so sub-scoped sibling criteria like `R4a` / `R4b` (introduced by `/flow-next:capture` when revising specs in-flight) are no longer dropped from `acceptance_criteria` / `uncovered_r_ids`; the suffix form is now blessed in `templates/spec.md` as canonical. (2) The `memory_during_spec` time-window filter now has a deterministic null-safe fallback chain (spec.created\_at → earliest task created\_at → branch first-commit via `git log {base}..{branch}`) so memory entries surface correctly when `spec.created_at` is null. Both surfaced during fn-48’s make-pr where PR #146 carried workaround prose. Two further bugs in the branch-first-commit fallback (returning the branch tip not the root commit; walking inherited mainline history) were caught by `chatgpt-codex-connector[bot]` review on PR #147 and fixed with regression tests. Unit suite 624 → 646.

### 1.2.0: Backend-split review workflows

**Review skills split by backend so codex/copilot load only their own slice (impl-review 1126 → 70 LOC on codex); FLOWCTL prelude consolidated.**

Detail

`spec-completion-review` drops from 645 to 41 LOC on codex (14×), `impl-review` from 1126 to 70 LOC (16×). RP keeps its cohesive prompt template since it only loads under the RP backend. `resolve-pr` was evaluated and kept inline (its parallel-vs-serial divergence sits below the 50-line split threshold codified in `agent_docs/adding-skills.md`). Also drops the dead `DROID_PLUGIN_ROOT:-CLAUDE_PLUGIN_ROOT` fallback from the Codex mirror’s FLOWCTL prelude (neither env var is set in Codex) and consolidates the canonical prelude to once-per-skill-file. Mechanical refactor only - bash, gating, and verdict semantics unchanged; smoke 127/2 baseline-equivalent. Factory Droid contract re-verified 2026-05-25 - `DROID_PLUGIN_ROOT` fallback + `Bash|Execute` matcher stay; `.factory-plugin/plugin.json` fallback dropped as dead code per Factory’s Claude-Code interop guarantee.

### 1.1.11: Cross-spec plan-sync setup fix

`flowctl init` no longer silently flips pre-1.1.3 users’ `planSync.crossEpic` from on to off - it mirrors the legacy value to the canonical `planSync.crossSpec` key before the default-merge when canonical is absent. Legacy key preserved through 1.x.

### 1.1.10: `usage.md` reference

`.flow/usage.md` template promoted to a comprehensive CLI reference (100 → 212 lines) - adds `status`, `config get/set`, per-spec/task `set-backend`, `checkpoint`, `ralph` control, lifecycle commands, and a corrected file-structure diagram.

### 1.1.9: Copilot review on Windows (real fix)

`flowctl copilot *-review` now works on native Windows by delivering the prompt via stdin (`subprocess.run(input=…)` with `--session-id` / `--resume`), sidestepping the `CreateProcessW` 32,767-char argv cap entirely. Verified by a real-subprocess Windows CI smoke round-tripping a 60 KB prompt. Supersedes the 1.1.8 WSL workaround. Upstream: [github/copilot-cli#3398](https://github.com/github/copilot-cli/issues/3398).

### 1.1.8: Copilot Windows fail-fast guard

Fail-fast guard + WSL pointer for the Windows Copilot argv-cap failure (cryptic `OSError winerror 206` before; clean error after). Reported by Simon Flauger (SEMA-CAD). Real fix landed in 1.1.9.

### 1.1.7: Codex mirror frontmatter cleanup

Stripped `request_user_input` from 6 Codex-mirror `SKILL.md` frontmatters - fn-45 rewrote the prose reference but left it in frontmatter, so Codex agents called the unavailable tool and reintroduced the Default-mode failure. `sync-codex.sh` guard tightened so it can’t regress.

### 1.1.6: prime ruleset-based branch protection

`/flow-next:prime` SE1 now detects ruleset-based enforcement in addition to classic branch protection, so GHE Enterprise repos protected via repo / org / enterprise rulesets correctly show SE1 ✅. Reported by Georg Keller (SEMA-CAD).

### 1.1.5: interview business-scope reframe

Removed deadline / time-budget / sprint-cadence questions from `/flow-next:interview --scope=business` (agents can’t estimate their own work, and time pressure collapsed the interview into brutal-prioritization). MVP-scope cuts reframed by feature value; budget envelope scoped to infra / vendor / licensing.

### 1.1.4: Canonical `## Acceptance Criteria` heading

`## Acceptance Criteria` is the canonical spec heading. Parsing stays tolerant of legacy `## Acceptance` and lowercase `## Acceptance criteria` - existing specs need no migration.

### 1.1.3: crossSpec alias + SPEC.md discovery

Cross-spec plan-sync aligned on `planSync.crossSpec`. Repo-root `SPEC.md` / `spec.md` template discovery for project-customized spec scaffolds; `/flow-next:setup` can opt into a root `SPEC.md` without clobbering custom templates.

### 1.0: Spec-driven foundation

The 1.0 line that stabilized the vocabulary and core workflow:

* Spec vocabulary stabilized (Spec / Task / R-ID / Handover / Receipt / Ralph)
* Symmetric `--scope=business|technical|both` interview
* Source-tagged capture with mandatory read-back
* PR-as-cognitive-aid generation
* Agent-native memory audit and migration
* PR feedback resolver
* Strategy and glossary grounding
* Ralph autonomous mode with receipts

***

Maintaining this page (for contributors)

**Register first**: entries are customer-facing. Use one narrative order: human outcome first, changed journey second, portable data and implementation details third. Lead with what the release gives the reader and its number, then what changes when they upgrade. The development story (what was tried and dropped, what measured worse first) stays out; a bound sits beside the number it bounds. For review features, make the human reviewer the protagonist: orientation, logical review journey, risk focus, evidence, and retained judgment. Machinery stays in the [repository changelog](https://github.com/gmickel/flow-next/blob/main/CHANGELOG.md) or an “Under the hood” tail. Upgrade actions open the detail block, imperative. Numbers are outcomes (”30s to half a second”), not inventory (LOC, test counts). Candid about bounds and what did not change; zero hype; plain hyphens.

**Hard rejection test**: hide the technical tail. A user must still be able to explain why the release matters, how their workflow changes, and what control they retain. Rewrite any entry whose title or opening sentence leads with a command, schema/type name, artifact, fixture, parser, hash, or benchmark. Concrete capability creates excitement; adjectives do not.

**What belongs here** - human-readable release highlights:

* new slash commands
* changes to spec or task semantics
* review receipt changes
* migration requirements
* breaking or deprecated behavior
* docs, team workflow, or Ralph changes that alter how people should work

The [repository `CHANGELOG.md`](https://github.com/gmickel/flow-next/blob/main/CHANGELOG.md) remains canonical - this page summarizes the current public story, not every commit.

**Per-release entry format** - add to the top of `## Latest`. The full writing rules are in [Writing an entry readers can use](https://github.com/gmickel/flow-next/blob/main/agent_docs/releasing.md#writing-an-entry-readers-can-use):

```mdx
### X.Y.Z - short title (3-6 words)


**One or two short sentences: what the reader can do now or what got better, by how much, and an upgrade action only when one exists.**


<details>
<summary>Detail</summary>


Start with the changed user journey and retained control. Put commands, schemas,
artifacts, fixtures, parsers, and benchmarks in an "Under the hood" tail. Blank
lines around this block so MDX renders the markdown inside.


</details>
```

Rules: version `### X.Y.Z - title` heading (never a bare bullet - it’s what makes the right-sidebar TOC a version index); bold one-liner mandatory; `<details>` only for verbose multi-paragraph releases (trivial patches skip it); newest at the top of `## Latest`; migrate the oldest `## Latest` entries down to `## Earlier releases` once it grows past \~10 (migrated entries keep their heading and bold one-liner, and keep a `<details>` block only for major behavior changes); bump `src/lib/site.ts` `FLOW_NEXT_VERSION` + `package.json` in the same commit. Full runbook: `agent_docs/releasing.md` -> “Docs-site changelog entry”.

**Release flow:**

```mermaid
flowchart LR
  Change["Behavior change"] --> Docs["Update docs"]
  Change --> Tests["Run tests"]
  Docs --> Changelog["Update changelog"]
  Tests --> Release["Cut release"]
```

If Flow-Next behavior changes and the docs site does not, assume the release is incomplete until proven otherwise.
