Optional Jev judgments
Flow-Next works the same with or without a TypeSafe API key. Every decision has a working default path: code decides lifecycle facts, and the host decides the rest from the route matrix and the repository. With a key, Jev (TypeSafe’s System One model) answers a few narrow questions with one HTTP request per decision point, so the host can reach the same decision faster or cheaper. Jev may change how long a decision takes and what it costs; it must not change which decision is made.
Routing and the QA gate never ask Jev. A run with a key and a run without one take the same route and reach the same QA and memory decisions. What Jev still does: hint at the kind of a fork the host already found, reorder memory search hits without dropping any, and pick a task tier before dispatch.
The judge never predicts whether a review, QA pass, or landing will succeed, and every review, QA, and merge gate keeps its own contract.
Turn it on
Section titled “Turn it on”Export TYPESAFE_API_KEY in the environment of the host that runs Flow-Next. The command reads it at call time only; it never lands in .flow/config.json, a state file, a receipt, a stage line, or a prompt.
export TYPESAFE_API_KEY=<your key> # in the shell that runs your coding agentflowctl config get judge.enabled # true by defaultflowctl config set judge.enabled false # force off while keeping the keyflowctl config set judge.enabled truejudge.enabled defaults to true; a non-boolean value warns and behaves as true. A missing key or false sends no request and leaves the default path. Setup prints Judge: off (TYPESAFE_API_KEY is not set). when the key is absent. There is no SDK to install and no model selector: requests go to jev-latest, and every floor is a preset constant, not a config key.
Where Jev decides and where it advises
Section titled “Where Jev decides and where it advises”| Decision | With a key | The host’s part |
|---|---|---|
| A live spec’s lifecycle (PR tail, all done, recorded work route, direct or plan) | Code decides; Jev is not asked | None; the route prints as (code) |
| Intake route for an idea, a bug report, a question | Jev is not asked | The host decides from the route matrix, so a keyless run takes the same route |
QA under pipeline.qa=auto: are the criteria UI-observable | Jev is not asked | The host decides from the acceptance and the repo |
QA under pipeline.qa=auto: a startable target | Code resolves a documented target; Jev is not asked | None; no target is invented |
| Research before work on a ready spec | Jev is not asked | The route matrix’s read-first rule |
| Fork: observable or preference | An optional hint on the host’s own fork sentence | Decides; a hint never removes a fork the host found |
| Memory relevance | Reorders the top 15 BM25 hits; drops none | Picks the entries that apply from titles and snippets |
| Task tier | The tier preset below | Unchanged |
Runs without a key
Section titled “Runs without a key”/flow-next:flow checks once per run whether the judge can run: the key is present (checked without printing it) and judge.enabled is not false. When it cannot, flow prints judge: off once and makes no fork-gate call for the rest of the run. The live-spec route call runs either way, because its lifecycle decision and PR observation come from code. Memory search runs the same command either way; without a key --rerank returns BM25 order and sends nothing.
The host prints Route: <route> (host) at intake and records the QA stage lines, key or no key. A judge that is on but fails keeps the same default path and names the reason, jev-unavailable(<reason>).
judge: offRoute: work_planned (code)Route: build (host)stage: qa - ran (target: <cmd>)stage: qa - skipped(config: pipeline.qa=auto: no UI-observable criteria)The three sites that ask Jev
Section titled “The three sites that ask Jev”Fork hint
Section titled “Fork hint”When the host finds an open design or behaviour fork, it may ask Jev one question about its own fork sentence: does the answer depend on something observable (behaviour, output, timing, a failing case, a measurement) or on a product_or_preference call (scope, priority, authority, taste, a business rule)? An answer at 0.5 or above is a hint; anything else returns host. The hint never cancels a fork: the host classifies its own fork, and under an autonomy marker a preference fork still stops with NEEDS_HUMAN. The one-question-per-hop budget is unchanged.
fork-gate: observable (host)fork-gate: preference (host, jev hint product_or_preference 0.71)Memory rerank
Section titled “Memory rerank”Plan and workers run one shape, with or without a key:
flowctl memory search "<task sentence>" --limit 15 --rerank --jsonThe search returns up to 15 BM25 hits. With a key, one request scores them and the command reorders them without dropping any; --limit applies after the reorder. The judge receives each entry’s id, title, track, category, module, tags and snippet, never its path or BM25 score. The host reads titles and snippets and keeps the entries that apply, on both paths. Plan renders the result itself and does not spawn the memory scout to refine a keyless result. See Planning scouts.
memory: reranked (jev, 15 entries)memory: bm25 (jev-unavailable(no_key))Tier at dispatch
Section titled “Tier at dispatch”Before spawning a worker or a scout, the conductor asks which tier a senior engineer would assign the task to (mechanical, moderate, intelligent, long_running) plus two yes/no checks. Code supplies the task title, body, acceptance, the touch count, and whether quick commands exist. Only the two confident extremes act: mechanical at 0.8 or above selects the routing block’s fast-scout model, and long_running at 0.8 or above adds a bridge recommendation. Every other answer leaves routing unchanged, an explicit IMPLEMENTER: in the invocation always wins, and a block that names no fast model, or a host with neither model steering nor a bridge, keeps the current model and says so. The done summary records the model that actually ran. See Orchestration.
Tier: mechanical (jev 0.88) -> <model>Tier: long_running (jev 0.86) - bridge recommendedTier: session (jev moderate 0.61)Decision order for a live spec
Section titled “Decision order for a live spec”Code reads lifecycle facts in this order, with no judge call:
- An observed PR routes to the existing PR tail. The probe keeps its four observations (open, merged, closed, failed); a failed probe never reads as “no PR”.
- Tasks exist and all are done: make-pr, after the QA decision.
- An intentional plan with tasks, or one implicit owner under
no_plan: true: continue work on the recorded route, including resume admission. - A ready spec with no tasks: a recorded
no_plan: truegoes direct outright. Otherwise the three positive plan signals (an explicit ask for a plan, separate owners, staged PRs) select plan and their absence selects direct work. Whether to read before work is the host’s, by the route matrix’s read-first rule. - A spec that is not ready stays with the host for refinement, plan review, or proceeding.
Intake without a spec has no route call: the host picks the row from the route matrix. flow --explain prints its Next:, Route:, Signal:, Skip/narrow:, and Why not the alternatives: lines from that row, without a judge call.
What happens when the judge cannot answer
Section titled “What happens when the judge cannot answer”Availability is a fact, never an error. flowctl judge returns available: false with exit 0 and one of these reasons: no_key, disabled, http_<status>, transport, timeout, bad_answer (a response missing a question or a sent option), or over_budget. The caller records the reason and takes its default path: the host, the static routing block, or BM25 order. An unknown preset or an invalid state file is a command error and exits non-zero.
Requests use stdlib HTTP with a 10-second timeout per attempt and reach TypeSafe through HTTPS_PROXY when one is set, respecting NO_PROXY. Only HTTP 429 and 529 retry, twice, after 1 and 2 seconds. State plus questions are estimated at four characters per token against a budget of about 32k tokens; an oversized request returns over_budget without sending. The judge writes neither state nor answers. Credentials never appear in command output, receipts, stage lines, or logs.
A request sends its supplied artifact text to TypeSafe. Route state is never sent. Tier state carries task_title, task_body, acceptance, touches_count, has_quick_commands, and repo.
flowctl judge --preset tier --task fn-1.1 --jsonflowctl memory search "windows subprocess" --limit 15 --rerank --jsonThe CLI reference has the JSON envelope.
Numbers of record
Section titled “Numbers of record”From the September 2026 evaluation, one directional pass per site. These are historical measurements, not a price quote or a latency guarantee.
| Site | Result and bound |
|---|---|
| Fork | 0.88 against 0.76 for the fork-present and kind pair on spec text; the hint on the host’s own fork sentence has not been measured |
| Memory rerank | Precision@5 0.66 against BM25’s 0.48, measured with the earlier score levels and floor; the reorder-only shape has not been measured |
| Tier | 0.91 exact and 121/121 within one tier; at 0.8, mechanical 20/20 and long-running 13/13 matched labels |
| Route kind (retired) | 0.95 raw agreement on 196 stable samples. Retired because keyed intake routes sent small features down heavier routes than keyless runs, and routing was no faster for it |
| QA gate (retired) | 0.88 against a 0.71 baseline. Retired because keyless runs reached the same QA decision in the same step, so the call added a few seconds and changed nothing |
What these numbers do not show: that a cheaper worker completes every task labeled mechanical, or that production runs match the evaluation’s latency. The implementation review and verification gates remain authoritative.
TypeSafe references
Section titled “TypeSafe references”- System One and primitives
- Choice, Noul, and Score
- Confidence and confidence routing
- Fan-out and building with System One
- Structured instructions, HTTP API, and the documentation index
Related
Section titled “Related”- Flow - the route step and
--explain - Orchestration & routing - tiers, the routing block, and the bridge
- Configuration -
judge.enabled