{"_id":"@benkamber/purser","name":"@benkamber/purser","dist-tags":{"latest":"2.9.0"},"versions":{"2.9.0":{"name":"@benkamber/purser","version":"2.9.0","publishConfig":{"access":"public"},"description":"Installer for purser — a Claude Code Dynamic Workflow that splits an implementation task across parallel worktree agents (preflight → explore → plan → parallel implementers → merge → adversarial review → PR). Places the workflow file into ~/.claude/workfl","type":"module","bin":{"purser":"bin/cli.mjs","purser-checks":"phase0-checks/run-checks.mjs","purser-ledger":"scripts/ledger-aggregate.mjs"},"scripts":{"test":"node --test \"tests/*.test.mjs\" && node phase0-checks/run-checks.mjs","checks":"node phase0-checks/run-checks.mjs"},"keywords":["claude-code","claude","workflow","dynamic-workflows","agents","orchestration","purser"],"homepage":"https://github.com/benkamber/purser#readme","bugs":{"url":"https://github.com/benkamber/purser/issues"},"repository":{"type":"git","url":"git+https://github.com/benkamber/purser.git"},"author":{"name":"Ben Kamber"},"license":"MIT","engines":{"node":">=18"},"gitHead":"427f2e32464892a74540698c3aed792d68b3038e","_id":"@benkamber/purser@2.9.0","_nodeVersion":"25.9.0","_npmVersion":"11.12.1","dist":{"integrity":"sha512-7/RBKiqtN7NPlNbdfTDt+1G2bhQMQUUNksbRoJ7DIInajk07sL8+OvcgNIw9r0W1oMlUD3VjdR6vjDCVlqJqLg==","shasum":"4a7d801d518a17b8419ef7f9a4e037f4229c023d","tarball":"https://registry.npmjs.org/@benkamber/purser/-/purser-2.9.0.tgz","fileCount":9,"unpackedSize":272977,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIB4GMGpmL6RQxljyKpuIusgVYTmzPSGuPHmPNNrF5HfyAiEAoP6F4eXAnvHvmJnVbl4ynj/B8n5W5mdskJQuxTfaT9E="}]},"_npmUser":{"name":"benkamber","email":"benkamber@gmail.com"},"directories":{},"maintainers":[{"name":"benkamber","email":"benkamber@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/purser_2.9.0_1784863117024_0.8101252447180387"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-24T03:18:36.883Z","2.9.0":"2026-07-24T03:18:37.164Z","modified":"2026-07-24T03:18:37.376Z"},"maintainers":[{"name":"benkamber","email":"benkamber@gmail.com"}],"description":"Installer for purser — a Claude Code Dynamic Workflow that splits an implementation task across parallel worktree agents (preflight → explore → plan → parallel implementers → merge → adversarial review → PR). Places the workflow file into ~/.claude/workfl","homepage":"https://github.com/benkamber/purser#readme","keywords":["claude-code","claude","workflow","dynamic-workflows","agents","orchestration","purser"],"repository":{"type":"git","url":"git+https://github.com/benkamber/purser.git"},"author":{"name":"Ben Kamber"},"bugs":{"url":"https://github.com/benkamber/purser/issues"},"license":"MIT","readme":"# purser\n\n[![tests](https://github.com/benkamber/purser/actions/workflows/test.yml/badge.svg)](https://github.com/benkamber/purser/actions/workflows/test.yml)\n\n**Per-work-item model-tier governance for Claude Code pipelines, built so\ntop-tier model spend is never a surprise. Invoke it as `/purser`. Hand it a coding\ntask: it reviews the repo, plans, implements in parallel git worktrees, merges,\nruns your full suite, adversarially reviews the diff, and opens a PR. Opt into\ninnovate mode and it also decides, item by item and under a dollar budget, when\nthe top model tier is worth paying for, and lets you overrule it before any\nmoney moves.**\n\nPurser is a single-file [Claude Code Dynamic Workflow](https://code.claude.com/docs/en/workflows).\nTwo design decisions differentiate it:\n\n1. **A disinterested selector gates escalation.** The selector chooses between\n   the `opus` and `fable` tiers per work item, so purser defaults the selector to\n   `sonnet`, a model outside that candidate set. An interested model has a\n   conflict in recommending itself; a disinterested one does not. `opus` and\n   `fable` remain available as A/B arms (`--selector=`) precisely so runs generate\n   first-party evidence for or against the self-interest hypothesis. To be plain:\n   the conflict-of-interest rationale is a design principle with an eval protocol\n   attached ([docs/eval-protocol.md](./docs/eval-protocol.md)), not yet a measured\n   effect. The results will be published whichever way they land.\n2. **A per-item human tier preview that actually binds.** `/purser --quote\n   <task>` plans the work, prints each item with the selector's recommendation,\n   reasoning, and a deterministic dollar estimate, and saves the quote.\n   `/purser --execute` (optionally with `--fable-items=`) builds exactly the\n   quoted plan: Explore, Plan, and Select are all skipped (zero re-spend), the\n   plan hash guarantees your picks bind to the plan you read, and an unknown\n   item id fails loudly with the valid id list and the hash, never silently.\n   The structured form (`planOnly: true` plus passing `plan`/`verdicts`/`recon`\n   back) remains fully supported for scripts.\n\n## The pipeline\n\nEight phases. Each runs on the cheapest model that does its job well, and the top\ntier is spent only where reasoning compounds.\n\n| # | Phase | Model / effort | What it does |\n|---|---|---|---|\n| 0 | Preflight | haiku / low | Repo sanity, default-branch and verify-command detection, acquires the per-repo lock. Read-only except the lock file. |\n| 1 | Explore | haiku / low (parallel scouts) | Read-only recon of the task surface and blast radius. Skipped when a previewed plan is passed back in. |\n| 2 | Plan | fable / high | The one stage always worth the top tier. Emits file-disjoint work items with per-item fast-verify commands. Skipped when a previewed plan is passed back in. |\n| 3 | Select | sonnet / medium (innovate only) | The disinterested selector rules opus-vs-fable per item. Skipped when previewed verdicts are passed back in. |\n| 3b | Quote | haiku / low (quote only) | Persists the quote (plan, verdicts, hash) under `.git/purser/previews/` so `--execute` can load it. |\n| 4 | Implement | sonnet / medium (opus / high for `complex` items; retries escalate) | One fresh git worktree per item. Implement, verify, commit, publish a branch. |\n| 5 | Integrate | sonnet / medium | Merge item branches, run the full suite once. A red build gets one opus repair attempt. |\n| 6 | Review | opus / high, read-only (sonnet fallback when the token budget is low) | Adversarial diff review. Opus fixes serious findings, sonnet re-verifies the fix. A green build is never left unreviewed. |\n| 7 | Finalize | haiku / low | `push --force-with-lease`, open the PR, delete only fully-merged item branches, release the lock. |\n\nThe model names (`haiku`, `sonnet`, `opus`, `fable`) are aliases the runtime maps\nto real models. See [Install](#install).\n\n## You are always in control\n\npurser never takes a model decision away from you. It has four modes, and they\ndiffer only in who decides which work items get the premium (Fable) tier:\n\n| Mode | Who decides | Command |\n|---|---|---|\n| Off (default) | nobody escalates; standard casting runs | `/purser <task>` |\n| Advisor | the selector decides, within your budget | `/purser --innovate <task>` |\n| You pick | purser quotes, you choose per item | `/purser --quote <task>` then `/purser --execute --fable-items=a,c` |\n| Everything premium | you, with one flag | `/purser --innovate --fable-items=all <task>` |\n\nWhatever the mode, every run prints what was decided, by whom, and what it is\nestimated to cost, and the PR body records it permanently.\n\n## The quote flow (pick items yourself)\n\nThink of it as getting a contractor's quote before approving the work.\n\nStep 1, get the quote. Nothing is built; this costs roughly a dollar:\n\n    /purser --quote migrate the auth system\n\npurser plans the work and prints a table: each work item, its model plan,\nwhat the premium tier would buy you on that item, and a per-item dollar\nestimate. It also prints the whole-task cost line:\n\n    cost:  standard ~$4 | recommended ~$11 | everything-Fable ~$26\n\nStep 2, approve. You never compose the execute command; you reply to a menu.\nEvery quote ends with:\n\n    what next? reply with a word:\n      advisor - Fable on the advisor's picks: auth, tokens (~$11 total)\n      all     - everything on Fable (~$26)\n      none    - no Fable; recommended items fall back to Opus (~$5)\n      or name the items you want on Fable - e.g. \"auth, tokens\"\n    (internals, only needed for scripting: quote id <hash>, branch slug\n    <slug>; execute later with /purser --execute, an older quote with\n    /purser --execute=<id>, item picks with --fable-items=<ids>|all, no\n    Fable with --fable-budget=0)\n\nReplying \"advisor\", \"none\", or \"just auth\" works because the quote lands in\nyour Claude Code conversation and the session translates your reply into the\ncommand. Exclusions (\"all except docs\") are not a reply form; name the items\nyou want instead. A bare --execute always means \"my latest quote in this\nrepo\"; the quote id lives in the footnote and is only needed to reach an\nolder quote. The syntax exists for precision and scripting; reply words and\nplain English are the everyday path.\n\npurser builds exactly the plan you were quoted (the hash guarantees it; a\nmistyped item id fails immediately with the valid list) and your picks always\nrun premium-first. Picks are honored even past the budget, with the overage\nlogged, because an explicit human decision outranks the cap.\n\nQuotes are saved per repo under `.git/purser/previews/<hash>.json` (plus a\n`latest` pointer), never committed, and written and read by agents, since\nworkflow scripts have no filesystem access (see [HARNESS.md](./HARNESS.md)).\nA missing, corrupted, or mismatched quote fails loudly with what was found and\nhow to re-quote, and `--execute` with a different task than the one quoted\nrefuses to build.\n\nWhen would you use everything-premium (`--fable-items=all`) instead? Three\ncases. When stakes dwarf cost: a production incident, a one-time migration, a\nresearch spike where a missed insight costs more than the run; you want\nmaximum capability everywhere and zero risk of a wrong demotion (hard\nexclusions still apply, since those are safety routing, not economy). When\nmeasuring: an all-premium run on a representative task shows you what the\nadvisor's selectivity would have saved. And when starting out: `all`\nreproduces \"just use the best model for everything\" inside purser, so you\nadopt the pipeline with your old cost profile, and every run's cost line\nquietly shows what the advisor would have picked. Start at `all`, watch\nthe deltas, graduate to advisor mode when the evidence convinces you.\n\n### A complete example, end to end\n\nReal output, captured 2026-07-06 in this repository: purser quoting one of\nits own v2.7 roadmap items. Rendered verbatim from the run's journal; long\nlines are wrapped and the log prefix stripped.\n\n    you>  /purser --quote make retries and the opus-first cascade escalate on\n          failed verification, not just failure to commit - an implementer\n          that commits with a red verify must be retried or escalated\n\n    MODE: quote - plan + per-item recommendations only, nothing built; the\n    quote is saved for --execute (Fable $ budget: $8); selector=sonnet\n    (disinterested: outside the Opus/Fable candidate set)\n    ⚠ main checkout is dirty (on main); scouts will verify facts against\n    main, implementers build from main\n    explore: 2/2 scouts, 24 facts\n    plan: 1 work item (single item: fast path, no integrator)\n    selector (sonnet): 1 escalation candidate(s); 1 on Fable - $1.05 reserved\n    of the $8 Fable $ budget (worst-case $3.00 if every cascade fires):\n    verify-trigger\n\n      item            model plan                              what Fable buys you                                           est\n      verify-trigger  Opus first, Fable as backup on failure  Miss: retry silently drops red-verify work, ships wrong; ...  $3.00\n\n    cost:  standard ~$1.50 | recommended ~$2.55 | everything-Fable ~$3.00\n    (implementer-spend estimates from the pinned 2026-07-06 rate table; live\n    runs so far have come in under estimates)\n    risk if you skip Fable entirely: MEDIUM - verify-trigger: Fable is\n    recommended but verification is strong; a cheap-model miss would surface\n    as a failed verify (a retry, not a silent defect).\n    quote ready - 1 work item (saved; reply below, or execute any time later)\n    what next? reply with a word:\n      advisor - Fable on the advisor's picks: verify-trigger (~$2.55 total)\n      all     - everything on Fable (~$3.00)\n      none    - no Fable; recommended items fall back to Opus (~$1.50)\n      or name the items you want on Fable - e.g. \"verify-trigger\"\n    (internals, only needed for scripting: quote id 65a7070d, branch slug\n    verify-aware-cascade; execute later with /purser --execute, an older\n    quote with /purser --execute=<id>, item picks with\n    --fable-items=<ids>|all, no Fable with --fable-budget=0)\n\n(the selector's summary line has since been reworded to plain language)\n\nQuote dollars price implementer spend; whole-run cost adds planning, review,\nand orchestration overhead (real figures from the second capture below:\n$0.54 quoted implementer spend, $1.97 leg total).\n\nThe table truncates the consequence column at 60 characters; the full\nverdict reads: \"Miss: retry silently drops red-verify work, ships wrong;\ntests catch it\". That is the column's one job - the consequence of a\ncheap-model miss on this item, not a description of the item - and because\nverification here is strong, the risk line grades skipping Fable as MEDIUM:\na miss surfaces as a red verify, not a silent defect.\n\nAnd here is the advisor declining. Same day, same repo, purser quoted its\nother pending fix (preamble and internals footnote as above, elided):\n\n    you>  /purser --quote fix the finalize flake - retry the finalizer once\n          on spawn failure, and log loudly when a run ends with the lock\n          still held\n\n      item            model plan         what Fable buys you                                       est\n      finalize-retry  standard (Sonnet)  Detailed spec + strong prior art; Opus handles this fine  -\n\n    cost:  standard ~$0.54 | recommended ~$0.54 | everything-Fable ~$1.80\n    risk if you skip Fable entirely: LOW - no items were recommended for\n    Fable.\n    quote ready - 1 work item (saved; reply below, or execute any time later)\n    what next? reply with a word:\n      advisor - Fable on the advisor's picks: none recommended (same as \"none\")\n      all     - everything on Fable (~$1.80)\n      none    - no Fable; recommended items fall back to Opus (~$0.54)\n      or name the items you want on Fable - e.g. \"finalize-retry\"\n\nReal output: the advisor declining the premium tier on a routine fix -\ndetailed spec, strong in-repo prior art, three files, tests alongside;\nbenefit \"low\", nothing recommended, risk LOW. A selector with a stake in the\npremium tier does not write that quote. In the four live selector runs\nrecorded to date it has never once recommended the top tier outright;\n[docs/eval-protocol.md](./docs/eval-protocol.md) keeps the running count\n(direction noted, no conclusions).\n\nThe reply to that quote was \"none\" - build it standard. The session\ntranslated the reply to `/purser --execute --fable-budget=0`; purser loaded\nthe saved quote, verified its hash, skipped every planning stage (zero\nre-quote spend), built the fix with a sonnet implementer, and the opus\nadversarial review returned zero findings - 169 tests and 16 checks green.\nThat fix is merged as [PR #1](https://github.com/benkamber/purser/pull/1),\nand the run that built it also reproduced the very flake it fixes: its own\nfinalizer, still running the pre-fix code, never spawned, and the run ended\nsilently with the repo lock held. The push and PR were finished by hand;\nwith the fix merged, that same failure now retries once and then says\nloudly what state it left behind.\n\n> **Why the risk line exists.** In an earlier real run on a private repo\n> (identifiers genericized), a routine-rated item in a multi-tenant\n> workspace app - built on the standard tier - passed the entire verify\n> suite, and the adversarial review then caught a cross-tenant data leak in\n> it: project ids derived from user-visible names onto a global primary\n> key, with the collision guard turning same-name creation into an\n> information leak across tenants. That is the silent miss the risk line\n> prices - the one verification would never have caught.\n\n### purser is not /model\n\n`/model fable` and `/purser --innovate --fable-items=all` sound similar and do\ndifferent things. `/model` changes which model you are chatting with: one\nmodel, working serially in your conversation, no worktree isolation, no\nparallel implementers, no adversarial review, no ledger, and every mechanical\nstep (scouting, merging, pushing) billed at the premium rate too.\n`--fable-items=all` changes who implements the work items inside the\npipeline: the premium model writes the code, while the cheap tiers still do\nthe mechanical stages, the standard tiers still integrate and review, hard\nexclusions still route safety-sensitive items away, and the ledger still\nrecords everything.\n\nOne gotcha worth stating plainly: purser's stage casting is pinned by design\nand never inherits your session's model, in either direction. Running\n`/model fable` and then `/purser` does not put Fable on the implementers,\nand `/model haiku` does not downgrade them either. The only levers over\npurser's model usage are its own flags, which is exactly what keeps its\ncosts predictable.\n\n#### If your session default is already the premium model\n\nThe pipeline is not a downgrade; it is a reallocation. One concrete task,\nthree ways (costs illustrative; the shape is the point):\n\n| Setup | What runs where | Rough cost |\n|---|---|---|\n| Fable default, no purser | Fable does everything serially in your chat: scouting, implementing, merging, pushing, all premium-priced, no worktree isolation, no adversarial review | ~$25 |\n| Fable default, `/purser` | haiku scouts, fable plans, sonnet implements routine items, opus handles complex items and the adversarial review, haiku finalizes | ~$5 |\n| Fable default, `/purser --quote` then execute | the same pipeline, plus Fable returns for exactly the items where the advisor or you decided it earns its rate | ~$7 |\n\nYour session default sets what you talk to; purser sets what you pay for,\nper work item, and the quality-critical stages get more scrutiny than a\nsingle-model chat gives them, not less.\n\n## Learn Mode (`--explain`) — learn system design as you build\n\n`--explain` (opt-in, off by default) turns any run into a system-design lesson.\nAn in-workflow Sonnet tutor reads the task and the plan and explains the **system\ndesign of the feature you're building** — its components and data flow, the key\ndecisions with their tradeoffs and the alternatives not taken, and the general\nCS/SWE/ML concepts it exercises. It explains the *thing being built, not purser\nitself*, so doing real work across your projects doubles as interview-grade\nsystem-design practice — the breadth comes from your work (a RAG feature teaches\nML system design; a schema change teaches data modeling and consistency).\n\n- `--explain` — standard explanation.\n- `--explain=deep` — also maps each design choice to the canonical system-design\n  **interview topic** it exemplifies and what an interviewer would probe.\n- `--explain=<focus>` — emphasize a free-form angle, e.g. `--explain=scaling`.\n\nIt is returned as `designExplanation` — in `--quote`/`planOnly` previews it comes\nback **before anything is built**, so it also informs your go/no-go — and as a\n\"System design of this change\" section in the PR body. It reads the plan as input\nand never alters it (the plan hash is identical with and without it), so it never\nchanges what gets built; it costs one Sonnet agent. Tradeoffs are enforced by the\noutput schema (at least one decision, each with an explicit tradeoff) plus a\nretry, so a normal run reliably teaches tradeoffs; depth is the model's judgment.\nComposes with every mode, e.g. `--quote --explain` or `--innovate --explain`.\n\n## How innovate mode works\n\nNormal `/purser <task>` runs `fable` for planning only; everything else is\nhaiku/sonnet/opus, and that cost path is covered by a regression test that fails\non any drift. `--innovate` (or `innovate: true`) adds the selector and the\nbudget. It is entirely opt-in.\n\n### The selector rubric\n\nThe selector escalates an item to `fable` when any of these fire (each recorded\nin the verdict's `signals`):\n\n- estimated effort over about an hour human-equivalent, or explicitly multi-stage\n- touches more than 3 files, or must hold an invariant spanning modules\n- no test coverage exists for the affected code\n- production-critical, high blast radius, or expensive to reverse\n- genuine novelty with no in-repo prior art, judged against scout recon\n\n### Hard exclusions (never fable, forced to opus)\n\n- **security, cryptography, privacy, bio, chem.** This rests on a pinned,\n  dated fact about the model version (the `FABLE_BEHAVIOR` constant, currently\n  as of 2026-07-04 for `claude-fable-5`): the tier's safety classifier reroutes\n  this work to an opus-class path while still billing fable rates, so escalating\n  pays about 2x for opus-grade output. If a model bump changes that behavior,\n  the constant is a one-line audit.\n- **zero-data-retention-sensitive items**, best-effort: flagged only when an item\n  explicitly handles sensitive data, since ZDR need is often invisible in a\n  description.\n- interactive, trivial, single-function, or boilerplate items.\n- underspecified items (the top tier over-explores loose prompts).\n\nHard exclusions beat an explicit human pick: force-select an excluded item and it\nstays on opus, with the block logged.\n\n### How the estimates work (deterministic dollars, not model-invented ones)\n\nThe selector never emits dollar figures. It estimates the tokens each item\nwould consume on the premium tier (`estTokensK`) and ranks benefit ordinally\n(`high`, `medium`, `low`); the harness prices tokens against a pinned\n`PRICING` table with an as-of date. The dollars are estimates for ranking and\ncapping, not a bill. The benefit judgment is the selector's honest opinion,\nprinted so you can disagree with it. Every dollar in logs, previews, and PR\nbodies traces to that arithmetic. If you execute a quote whose verdicts were\npriced under an older table, purser keeps the previewed dollars (your picks\nwere made against them) and logs the drift loudly.\n\n### Why purser will not give you a probability\n\nNo one, purser included, can compute the probability that the premium model\nfinds an insight the standard one would miss on your specific novel task.\nThat is unknowable before the work runs, for four reasons that stack:\n\n1. It is a counterfactual. Measuring it for one item means running the item\n   on both models and comparing, which spends what you were deciding whether\n   to spend, and even then compares single samples of stochastic processes,\n   not the models themselves.\n2. The tasks where the question matters most are the tasks with no oracle.\n   Routine work has tests that score the output cheaply, and routine work is\n   exactly where the cheap model suffices. Genuinely novel work lacks an\n   automatic judge, so you often cannot score the outcome even after the\n   fact, let alone before.\n3. Base rates do not transfer. Published benchmark gaps between model tiers\n   are averages over a task distribution. A per-item probability needs a\n   reference class of similar items, and a genuinely novel item, by\n   definition, has none.\n4. A made-up number is worse than no number. A language model asked for a\n   probability here will produce one, and it will be confabulated precision\n   that invites expected-value arithmetic on invented inputs. An ordinal\n   judgment with stated reasoning is the selector's true epistemic state:\n   rankable, auditable, and easy to overrule.\n\nThis is exactly why the quote flow keeps the call human, and why the ledger\nexists: over time it accumulates empirical frequencies by item class\n(first-attempt success by tier, escalation-on-failure rates), which is the\nhonest route to numbers. Those are measured after the fact and published as\ncalibration data, never asserted per item in advance.\n\nWhat purser gives you instead of a probability is the risk line in every\nquote: a qualitative judgment of where a decision NOT to escalate is most\nlikely to be regretted, and where a quality miss would go undetected because\nverification is weak. It is derived from the same signals as the\nrecommendations, stated with its reasoning, and easy to overrule. When the\nodds cannot be computed, purser names the blind spots and keeps the hedge\n(the \"all\" reply, everything premium) one keystroke away.\n\n### Budget semantics: a ceiling on the luxury option, never a wall\n\nThe Fable budget (default $8) decides which items are admitted to the premium\ntier before any work starts. It is not a meter running while work happens.\nSince the 2026-07-06 price correction, the default $8 admits roughly 3x more\nescalation than pre-correction runs; that is the corrected, as-documented\nbehavior (kept deliberately, not retuned).\n\n- Nothing gets interrupted, ever. An item admitted to Fable runs on Fable to\n  completion, even if it turns out bigger than estimated. The failure mode\n  \"the model was mid-breakthrough and got cut off at $19.50\" cannot happen,\n  because the budget gates admission, not execution.\n- Nothing fails for budget reasons. If the budget cannot cover every\n  recommended item, the highest-benefit items are funded first and the rest\n  run on Opus, the same strong model they would have used anyway. The log\n  says exactly which items were demoted and why.\n- You can always override. Raise the ceiling for one run\n  (`--fable-budget=$40`), remove it (`--fable-budget=unlimited`, logged as\n  \"unlimited\", never as a giant number), or hand-pick the items that matter\n  (`--fable-items=...`), which bypasses the cap entirely for those items.\n\nOne honest note on when to override: the budget is a cost-control tool for\nroutine engineering, where value is predictable. For exploratory work where\nmissing an insight is unacceptable, do not delegate the decision. Use the\nquote flow and pick the items yourself, or set the budget high. The estimates\nexist to inform you, not to decide for you. (Pre-correction live data ran about\n3x conservative because the placeholder price table sat roughly 3x above list;\nrepriced 2026-07-06, so quote dollars should now be roughly calibrated.)\n\nMechanically: candidates are ranked by benefit, then by cost. Fable-first\nitems reserve their full estimate (they will spend it); opus-first-cascade\nitems reserve at a discount (`--cascade-reserve=0.35` by default), because\nthey spend fable dollars only if the opus attempt fails verification. The\nworst-case exposure (every cascade firing) is logged alongside the\nreservation. Items past the boundary demote to opus, each demotion logged\nwith its estimate and benefit rank.\n\nHuman picks are honored **outside** the budget: an explicit human decision beats\nthe cap, and every override dollar is tallied, logged, and surfaced in the PR\nbody. A pick always means fable-first for any non-excluded item, even when the\nselector said otherwise; the override is logged and the ledger keeps the\nselector's original verdict next to the forced assignment, so the eval data\nstays honest.\n\n### Escalation strategy and graceful degradation\n\n- Item has tests and normal blast radius: opus first; only a verification\n  failure escalates to fable, with the failure notes as context (the autopsy,\n  not just the task).\n- Item is weak-verification, high-blast, long-horizon, or novel: fable first;\n  on failure it falls back to opus with the degradation noted. A fable outage\n  never breaks a run.\n\n### The ledger\n\nEvery run returns `telemetry.itemLedger`: a per-item join of the selector\nverdict (including `estTokensK`), the assigned strategy, the attempt sequence\nwith models, the final model, whether attempt one was committed and verified,\nwhether the item escalated on failure, and unresolved review findings.\nAccumulated across runs, this table is the selector-calibration eval.\n`scripts/ledger-aggregate.mjs` turns a directory of run results into the\ncalibration table.\n\n### The nudge (normal runs)\n\nAfter planning on a normal run, purser runs a pure-code, zero-model check: if\nany item is complex, touches more than 3 files, or no test suite was detected\n(user-supplied verify commands are respected), it may print one informational\nline:\n\n    innovate hint: 2 item(s) matched innovate criteria (touches >3 files; no\n    test suite detected); rough Fable projection ~$3.40. Re-run with --quote\n    to see per-item recommendations and pick yourself (nothing is built), or\n    --innovate to let the selector decide, if quality matters more than cost\n    on this task (suppress with --no-nudge).\n\nThat line is the nudge. It is on by default, costs nothing (no model is\nconsulted; it is computed from file counts and the pinned price table, and a\npremium model must never recommend its own use), and changes nothing about\nthe run. `--no-nudge` (or `nudge: false`) turns it off, and that is the\nflag's only job. The `--no-` prefix follows the CLI convention for disabling\nsomething that defaults to on.\n\n### The Fable trust audit (`--audit`)\n\n`--audit` (innovate only, opt-in) answers \"was Fable actually worth it over\nplain Opus?\" After the build, a **disinterested, skeptical Sonnet auditor — never\nFable, so it cannot grade its own homework** — rules each Fable item `warranted`,\n`unclear`, or `likely-unnecessary`. Evidence is routed by what is knowable:\nescalate-on-failure items are **measured** from the ledger (Opus verifiably\nfailed, then Fable ran); fable-first items are **assessed from the diff Fable\nactually landed** (there is no Opus counterfactual). Items that fell back to Opus\nare reported honestly rather than judged as Fable's work, and reserved cascade\nitems where Opus succeeded are shown as **$0 premium**.\n\nCrucially it can return a negative verdict — an audit that can only vindicate the\ntool is advocacy, not trust — and those negatives feed selector calibration. An\nagent fires only for items where Fable actually ran, so cost is proportional. It\nsurfaces as `telemetry.audit` plus a PR-body section. Dollars are estimates\n(the runtime exposes no per-Fable token counts).\n\n## Threat-aware (`.threatlas.json`)\n\nAlways on, but inert unless a signed `.threatlas.json` policy sits at the root\nof the target repo. Threatlas is a separate advisory tool (its own repo) that\nturns a threat-model intake — IP sensitivity, jurisdiction constraints,\ncompliance regime, budget — into a policy file naming which paths are\nsensitive enough to require the top execution tier. Purser is the consumer,\nnever the author, of that file.\n\n- **Verify-or-nothing ingest.** `.threatlas.json` is a signed JWS envelope, so\n  purser never reads it directly — it asks the authoritative verifier\n  (`guard/verify-policy.mjs`, expected at `$HOME/threatlas/guard/`) and honors\n  its exit code. **Verified**: the policy's `tier0Paths` glob list is trusted.\n  **Refused** (tampered, unpinned key, or a policy the verifier can't run\n  against): purser fails CLOSED to most-restrictive — every work item is\n  locked to Opus and Fable escalation is disabled, loudly, never as a silent\n  \"policy loaded\". **Absent**: the repo is ungoverned and behavior is\n  unchanged, whether or not the verifier is even installed. A one-time human\n  `--pin` (via the verifier, never purser) establishes trust in a policy per\n  repo.\n- **The risk override.** After Select, any work item whose files glob-match a\n  verified policy's `tier0Paths` is forced onto Opus — deterministically,\n  code-side, beating the selector's cost/benefit verdict *and* an explicit\n  human `--fable-items` pick. Recorded on the ledger as `riskForced` plus the\n  matched paths, so the override is auditable after the fact.\n- **Advisory only, always executed on Anthropic tiers.** The harness has no\n  per-agent `base_url`, so purser cannot actually route work to external\n  providers (GLM, Kimi, Codex) even when the profile recommends one for cheap,\n  low-sensitivity items. It reports what the profile *would* pick\n  (`recommendedTier`, mapped by capability/role — Threatlas's own \"Fable\"\n  label names Opus 4.8, never purser's `fable` tier, so the two names are\n  never conflated) next to what actually ran, and `--audit` rolls this up:\n  this run's estimated spend against the profile's `monthlyBudgetUsd` (context\n  only — purser sees one run, not a month) plus recommended-vs-executed tier\n  per item.\n- **Harness-block enforcement.** A verified policy's harness requirements are\n  enforced by purser itself, not just the Claude Code guard hook (which sees\n  tool calls, not whether the runner honors them). `networkMode: \"deny\"`\n  suppresses the git push and the PR — labeled honestly as partial, since\n  purser controls its own push, not an implementer's runtime egress, and\n  worktrees are `workspace_only`, never OS-sandboxed. `autonomy: t0`,\n  `approvalGate: every_action`, or `sandbox: isolated` refuse to build at all;\n  `planOnly`/`--quote` stay legal and are the compliant alternative. An\n  `unknown` or unrecognized field fails closed to the strictest reading.\n- **Supervision gate.** Runs on every build, policy or not: purser must never\n  write a policy file itself (`.threatlas.json`, `.claude/settings*.json`,\n  `guard/**`, `HANDOFF.md`, `~/.config/threatlas/**`) — an agent that can\n  rewrite its own permissions is ungoverned. A haiku agent reports the diff\n  verbatim; the match and the verdict happen code-side, before the opus\n  reviewer, so a build that tried it dies cheap rather than shipping.\n\nNo flag turns this on or off — presence of a verified `.threatlas.json` is the\nonly switch. See [Known limits](#known-limits) for what's advisory-only and\nwhat's still blocked upstream.\n\n## Evidence\n\nThis section summarizes what has actually been tested and how. The full plan is\n[docs/TESTPLAN.md](./docs/TESTPLAN.md) and the full accepted report is\n[docs/TESTREPORT.md](./docs/TESTREPORT.md).\n\n**Three layers.** Layer 1 exercises every deterministic decision with zero model\nspend (mocked runtime globals). Layer 2 runs all eleven end-to-end scenarios in\nsimulation with scripted agent responses but real git effects (worktrees,\nbranches, merges, the actual lock script) in throwaway fixture repos. Layer 3 is\nthe ongoing selector eval protocol. The suites in `tests/` plus the\n`phase0-checks/` regression pack total 264 tests plus 16 checks, all green in CI\non every push, at $0.\n\n**Five bugs found by adversarial review before install.** The v2.6 candidate was\nreviewed against its own change list with the null hypothesis that at least one\nclaim was subtly wrong. It was right five times: a path-boundary bug in review\nfinding attribution that would have corrupted the eval data, interior newlines\ncollapsing in multi-line tasks, an empty pick list silently enabling innovate\nmode with its default budget, stale preview pricing winning silently, and one\nambiguous budget log line. All five were fixed and regression-tested before the\nworkflow ever ran. Two further fixes came from human review of the findings\n(picks always mean fable-first; invalid selector values fail loudly). Finding\nbugs before install is the point of the process, so they are listed rather than\nhidden.\n\n**Live validation.** The preview-pick-execute loop was validated with real\nmodels against a seeded sandbox repo: four legs passed (a small golden-path run,\nthe preview, the execute leg bound to the previewed plan hash with zero\nExplore/Plan/Select agents and the human pick running fable-first over an\nescalate=false verdict, and the wrong-id leg failing loudly after preflight\nonly). Every pass/fail was determined from the workflow journal and repo state,\nnot from any model's account of itself. Total live spend was $6.05. The first\ncalibration row is seeded from those ledgers: one run, in which the\ndisinterested sonnet arm declined to escalate both items including a seeded\nfable-first candidate. That is a direction worth noting, and one run supports no\nconclusion; the table grows from here.\n\n## Who saves what\n\n- **API-billed users** get deterministic per-item dollar estimates before any\n  spend, a hard dollar budget on top-tier escalation, and demotion (never silent\n  overrun) when candidates exceed it.\n- **Subscription users** conserve premium-tier allowance: opus and fable fire\n  only where the rubric or an explicit pick says the quality is worth it, and\n  everything mechanical runs on cheap tiers.\n\n## Install\n\nPurser is a Claude Code workflow — a single `purser.js` that Claude Code runs\ninside its Dynamic Workflows harness (it is **not** a standalone Node program; it\nhas no CLI of its own and does nothing when run with `node`). \"Installing\" it\ntherefore means one thing: getting `purser.js` into a directory Claude Code scans\nfor workflows. There are two supported paths, and you only need one.\n\n**Requirements (both paths).** Claude Code with\n[Dynamic Workflows](https://code.claude.com/docs/en/workflows) enabled.\n\n### Path A — npm (recommended)\n\nThe npm package is an **installer**, not a runner. It ships `purser.js` and a\nsmall `purser` command that copies it into your workflows directory. There is\n**no npm `postinstall` hook** — writing into your Claude Code config is a visible\nstep you run yourself, never a side effect of `npm install`.\n\n```bash\nnpm install -g @benkamber/purser         # or, no global install: npx @benkamber/purser install\npurser install                           # copy purser.js -> ~/.claude/workflows/\n```\n\n> The package is **scoped** (`@benkamber/purser`) because the unscoped name\n> `purser` on npm belongs to an unrelated project. After a global install the\n> command is just `purser`; the one-shot form is `npx @benkamber/purser <cmd>`.\n\nThat copies `purser.js` into `~/.claude/workflows/purser.js`, backing up any\nexisting file first (to `purser.js.bak`, never `*.bak.js`), and prints exactly\nwhere it landed. Then in any project: `/purser --quote <your task>`.\n\nInstaller commands and flags:\n\n| command | what it does |\n|---|---|\n| `purser install` | Copy `purser.js` into `~/.claude/workflows/` (global, for you). |\n| `purser install --project` | Copy into `<cwd>/.claude/workflows/` instead (shared via the repo). |\n| `purser install --symlink` | Symlink instead of copy, so `npm update -g purser` refreshes it automatically. |\n| `purser install --dir <path>` | Install into an explicit directory (also honored via `$PURSER_WORKFLOWS_DIR`). |\n| `purser install --force` | Overwrite without leaving a backup. |\n| `purser status` | Show what is installed, where, and whether it matches the bundled version. |\n| `purser uninstall` | Remove the installed file (a hand-edited copy is protected unless `--force`). |\n\nCopy is the default because the README's \"edit it in place\" advice (below)\nassumes a real file you own; a symlink points back into `node_modules`, so edits\nthere are lost on the next `npm update`. Use `--symlink` only if you want\nauto-refresh and won't hand-edit.\n\nThe package also exposes two genuinely standalone tools as bins (they are real\nNode scripts, no harness needed): `purser-checks` runs the zero-cost adversarial\nregression pack against the bundled `purser.js`, and `purser-ledger <dir>`\naggregates purser run telemetry.\n\n### Path B — manual one-file copy (GitHub Release)\n\nNo npm, nothing runs against your config. Download `purser.js` from a\n[Release](https://github.com/benkamber/purser/releases) (or copy it from a clone)\ninto a workflows folder:\n\n- `~/.claude/workflows/purser.js` is available in every project, just for you.\n- `<repo>/.claude/workflows/purser.js` is shared with everyone who clones the repo.\n\n> If both exist with the same name, the project copy wins and the two silently\n> drift. Pick one home per machine and edit it in place.\n\n> **Every `.js` file in a workflows directory is registered as a live\n> command** under its `meta.name`, whatever the file is called. A backup\n> like `purser-old.js` whose `meta.name` is still `purser` (or an older\n> `forge.js`-era copy) will register alongside or instead of the real one.\n> Keep backups out of the directory, or give them a non-`.js` extension\n> (`purser.js.bak`), never `.bak.js`. (The npm installer follows this rule when\n> it backs up.)\n\n### After either path\n\n1. **Map the model aliases.** Purser routes stages to `haiku`, `sonnet`, `opus`,\n   and `fable`. Those aliases must resolve on your plan or provider. If your\n   account lacks a tier, edit the `model:` values in `purser.js` (and the\n   `phases` casting table in `meta`).\n2. **Monitor with `/workflows`**: arrow to the run, press Enter for the phase\n   view, drill into any agent to read its prompt and result.\n\n## Usage\n\nThe primary flow is the quote flow, two commands:\n\n    /purser --quote implement the recurrence engine and fix the pagination bug\n    /purser --execute --fable-items=pagination\n\nThe quote prints the per-item table, the cost line, the risk line, and the\nreply-word menu, and saves itself; `--execute` builds exactly the quoted plan\n(bare form = the repo's latest quote; `--execute=<hash>` reaches an older\none). A wrong item id throws with the valid id list and the plan hash before\nany implementer runs.\n\nThe same round trip as structured arguments, for scripts and automation:\n\n```javascript\n// 1. Preview: plan + per-item verdicts + deterministic dollar estimates,\n//    zero implementer spend. (--quote is this plus persistence.)\nconst preview = Workflow({ name: \"purser\", args: {\n  task: \"implement the recurrence engine and fix the pagination bug\",\n  innovate: true, planOnly: true,\n} })\n// preview returns: { plan, planHash, selector, recon, estUsdByItem,\n//                    assignments, fableBudget, pricingAsOf, rerun }\n\n// 2. Pick: read the verdicts, choose the items you want on the top tier.\n\n// 3. Execute: pass the preview back with your picks. Explore, Plan, and\n//    Select are skipped; your picks bind to the exact previewed plan.\nWorkflow({ name: \"purser\", args: {\n  task: \"implement the recurrence engine and fix the pagination bug\",\n  innovate: true,\n  plan: preview.plan, verdicts: preview.selector, recon: preview.recon,\n  fableItems: [\"pagination\"],\n} })\n```\n\nEveryday forms:\n\n```javascript\n// Full pipeline, auto-detected base branch and verify commands\nWorkflow({ name: \"purser\", args: { task: \"add rate limiting to the public API\" } })\n\n// small: 1 scout, 1 implementer, no integrator (about 1/4 the work)\nWorkflow({ name: \"purser\", args: { task: \"fix the off-by-one in paginate()\", small: true } })\n\n// innovate without a preview: selector + budget, inline flags in the task string\nWorkflow({ name: \"purser\", args: { task: \"--innovate --fable-budget=$12 <task>\" } })\n\n// plain-language launch also works: \"run purser: <task>\"\n```\n\nFlag reference (inline flags are recognized only at the leading and trailing\nedges of the task string, so a task like \"fix the parser so --innovate is\ndocumented\" keeps its middle intact — and a recognized flag that lands in the\n*middle* is now reported as unapplied via a note, instead of being silently\nignored):\n\n| Flag / arg | Meaning |\n|---|---|\n| `--explain` / `--explain=deep` / `--explain=<focus>` / `explain: true` | Learn Mode (off by default): an in-workflow Sonnet tutor explains the **system design of the feature this run builds** (from the plan) — components, tradeoffs, and the CS/SWE/ML concepts it exercises — so real work doubles as system-design-interview practice. It explains the *thing being built, not purser*. `=deep` also maps each choice to the canonical interview topic; `=<focus>` emphasizes an angle (e.g. `scaling`). Returned as `designExplanation` + a PR section; reads the plan without changing it; costs one agent |\n| `--audit` / `audit: true` | Innovate only, opt-in: a disinterested **skeptical** Sonnet auditor (never Fable) rules each Fable item **warranted / unclear / likely-unnecessary** — escalate-on-failure = *measured* (Opus failed → Fable fixed, judged from the ledger), fable-first = *assessed from the diff* (no counterfactual), reserved cascade = $0. Adds `telemetry.audit` + a PR-body section; fires an agent only where Fable actually ran; dollars are estimates |\n| `--quote` / `quote: true` | plan + recommendations + persistence, nothing built (implies innovate and planOnly) |\n| `--execute` / `execute: true` or `execute: \"<hash>\"` | build a saved quote (bare = latest in this repo; `--execute=<hash>` or a bare hash reaches an older one); loud failures name what was found and how to re-quote |\n| `--innovate` / `innovate: true` | enable the selector and the fable budget |\n| `--fable-budget=$X` or `=unlimited` / `fableBudget` | dollar budget for escalation (default $8); `unlimited` removes the cap and is logged as \"unlimited\" |\n| `--fable-items=a,b` or `=all` / `fableItems: []` or `\"all\"` | human tier picks; `all` = every non-hard-exclusion item; binds to a supplied plan; unknown ids fail loudly; auto-enables innovate |\n| `--selector=sonnet\\|opus\\|fable` / `selector` | selector arm (sonnet default, disinterested; opus and fable are self-interested A/B arms); invalid object-arg values throw |\n| `--cascade-reserve=0.35` / `cascadeReserve` | budget-reservation discount for opus-first-cascade items, clamped to [0,1] |\n| `--no-nudge` / `nudge: false` | silence the normal-run innovate hint |\n| `plan`, `verdicts`, `recon` | pass a preview back to skip Explore, Plan, and Select |\n| `planOnly: true` | stop after planning (and selecting, under innovate) with zero implementer spend |\n| `small: true` | 1 scout, 1 implementer, no integrator |\n| `verify` | override detected full-suite verify commands (string or array) |\n| `base`, `pr: false`, `maxItems`, `explorers`, `force`, `lockStaleMinutes` | base branch, skip the PR, parallelism caps, steal a stale lock, staleness window |\n\n### Syntax reference: this run vs every run\n\nFlags are typed inline with the task, before or after the text, and apply to\nthat run only:\n\n    /purser --innovate --fable-budget=$20 add rate limiting to the API\n\n| Setting | This run only (flag) | Every run (edit purser.js) | Default |\n|---|---|---|---|\n| Innovate mode | `--innovate` | opt-in by design | off |\n| Fable budget | `--fable-budget=$20` or `=unlimited` | `FABLE_BUDGET` default | $8 |\n| Pick items | `--fable-items=a,c` or `=all` | n/a (per-run by nature) | none |\n| Selector model | `--selector=opus` | `SELECTOR_MODEL` default | sonnet |\n| Cascade reserve | `--cascade-reserve=0.5` | `CASCADE_RESERVE` default | 0.35 |\n| Nudge | `--no-nudge` | `NUDGE` default | on |\n| Get a quote | `--quote` (implies innovate + plan-only) | n/a | off |\n| Execute a quote | `--execute` (bare = latest quote here) | n/a | n/a |\n\nPermanent changes are one-line edits: open `~/.claude/workflows/purser.js`,\nfind the default near the top (each is commented), change the value, save.\nFlags always win over file defaults for the run they are typed on.\n\nStructured arguments (`Workflow({name:'purser', args:{...}})`) exist for\nscripts and automation; every capability above is reachable without them.\n\n## Cost model\n\nCosts are relative because the absolute number depends on provider pricing, repo\nsize, and the task. Use `small: true` as the 1x baseline. The casting was tuned\nagainst measured per-stage cost across a batch of real runs: implementers were\nroughly 60% of spend, and about half of that implement spend came from items\nthat inherited an expensive session default because their `agent()` call left\nthe model unset. Pinning a model on every stage cut the measured average by\nroughly 40% with quality held by escalate-on-retry. The rule purser encodes:\nmatch the model to the work, and never let a stage inherit an expensive default.\n\nOn the first live data, the deterministic estimates ran about 3x conservative:\nthe root cause was the placeholder price table sitting roughly 3x above list\nprices, corrected 2026-07-06. Post-correction quote dollars should be roughly\ncalibrated, to be confirmed against upcoming ledger data. Estimate accuracy is\nbeing measured, not assumed; see Known limits.\n\n## Known limits\n\n- **Cascade and retry trigger on failure to commit, not failed verification.**\n  An implementer that commits despite a red verify (against its instructions) is\n  not retried or cascaded; the integrator's full-suite run is the backstop.\n  Moving to a verified-based trigger is a named v2.7 roadmap item.\n- **The finalize stage can flake at spawn.** Intermittently the runtime fails to\n  start the finalizer agent; purser now retries the finalizer once on that spawn\n  flake. If the retry also fails to spawn, purser reports an honest\n  `pushed: false` (never an invented PR url) and all branches survive, but the\n  per-repo lock stays held; the run loudly logs the lock path and remediation\n  (`rm` it or pass `force: true`) so it isn't a silent hunt. A legitimate\n  finalizer run that returns `pushed: false` (e.g. a rejected push) is never\n  retried - it is a successful execution reporting a real outcome, and it too\n  gets a loud warning that the lock is likely still held. A stale `purser.lock`\n  is safe to `rm` and self-expires after `lockStaleMinutes` (default 20).\n- **PRICING and FABLE_BEHAVIOR are pinned by date.** The rate table and the\n  classifier-rerouting fact carry as-of dates and must be re-verified against\n  current list prices and model behavior before each release. They are constants\n  precisely so a change is a one-line audit.\n- **Estimate accuracy is being measured, not assumed.** First live data showed\n  estimates running about 3x conservative; the root cause was the placeholder\n  price table sitting roughly 3x above list prices, corrected 2026-07-06. With\n  the corrected table, quote dollars should be roughly calibrated, to be\n  confirmed against upcoming ledger data. The ledger records `estTokensK` per\n  verdict so accuracy can be computed once the runtime exposes observed per-item\n  tokens; until then the aggregator reports the column honestly as n/a.\n- **Claude Code only.** Purser targets the Dynamic Workflows runtime and its\n  primitives (see [HARNESS.md](./HARNESS.md)). Runtime behavior can shift across\n  Claude Code updates; re-check after upgrading.\n- **Single-machine lock.** The per-repo lock serializes runs on one machine; it\n  does not coordinate two machines against the same remote.\n- **Quote persistence is agent-mediated.** Workflow scripts have no filesystem\n  access (verified against the runtime, see HARNESS.md), so `--quote` writes\n  its file through a haiku agent and `--execute` reads it back through\n  preflight. The quote file's integrity is hash-checked on load, and every\n  failure mode (missing, corrupted, wrong hash, different task) is loud. Saved\n  quotes live under `.git/purser/previews/` and are never committed.\n- **No visual verification.** Green means tests and typecheck, not pixels.\n- **A PR you still have to review.** Ten runs a day is ten PRs a day.\n- **Threat-aware execution is advisory only.** Purser can recommend a cheaper\n  or riskier tier from a `.threatlas.json` profile, but the harness has no\n  per-agent `base_url`, so it always executes on Anthropic tiers regardless of\n  the recommendation. Cross-provider execution waits for a harness provider\n  seam.\n- **The Codex lane's per-dispatch policy check isn't built.** The supervision\n  gate covers purser's own Anthropic implementers today; the equivalent\n  per-dispatch check for a future Codex lane is blocked on an upstream\n  Threatlas PR that hasn't landed.\n- **The threat-policy verifier is a local dev path.** `THREATLAS_VERIFIER`\n  points at `$HOME/threatlas/guard/verify-policy.mjs` (invoked, never\n  vendored, so a forked implementation can't drift from the authoritative\n  one). This moves to an npm binary once Threatlas ships one; until then,\n  threat-awareness needs a local Threatlas checkout at that path to do\n  anything beyond \"no policy present.\"\n- **The risk override forces by path, not by profile posture.** A verified\n  policy's `tier0Paths` is the only thing that forces an item onto Opus.\n  Threatlas's intake also carries a project-level posture (`ipSensitivity`,\n  `sensitiveCore`), and it is carried on `THREATLAS.profile` and reported in\n  `--audit`, but purser never uses it to force anything: a profile that says\n  \"everything here is sensitive\" without naming specific paths gets no\n  enforcement beyond what `tier0Paths` names. This is a real gap, not just an\n  undocumented one — it was scoped down during implementation from a design\n  that called for a profile-posture fallback, and the concrete rule for that\n  fallback (force every item on `ipSensitivity: high`? some threshold on\n  `sensitiveCore`?) was never actually specified. See Roadmap.\n\n## Troubleshooting\n\n- **\"another purser run appears active\"**: a live run holds the lock, or a\n  crashed or finalize-flaked run left it behind. Wait, `rm` the named lock file,\n  or pass `force: true`. Locks self-expire after `lockStaleMinutes`.\n- **A run ended with `pushed: false` and no error**: either the finalize spawn\n  flake (above) fired twice - purser already retried once and gives up after\n  the second failure, loudly logging the lock path and that push/PR were not\n  completed - or the finalizer ran but the push was legitimately rejected\n  (`notes` explains why), which is never retried. Either way the integration\n  branch exists and is verified (honest `pushed: false`, never an invented PR\n  url); push it yourself or re-run finalize by hand, and remove the lock at the\n  logged path.\n- **\"fableItems: unknown item id(s)\"**: item ids are minted by the planner, so a\n  fresh run re-mints them. To execute the exact plan you were quoted, use\n  `--execute` against the saved quote (the normal path), or pass the preview's\n  `plan`, `verdicts`, and `recon` back as structured args; the error message\n  includes the plan hash and the valid ids.\n- **\"--execute: no saved quote\"**: there is no quote saved in this repo (or the\n  `latest` pointer dangles). The error lists the hashes that are saved. Run\n  `/purser --quote <task>` first; quotes are per repo, under\n  `.git/purser/previews/`.\n- **Leftover worktrees under `.claude/worktrees`**: the runtime retains agent\n  worktrees that contain changes; one can even hold a published branch checked\n  out, which blocks a later `git branch -f` on the same name. `git worktree\n  remove` them between runs; the finalizer's notes flag leftovers.\n\n## Roadmap\n\nv2.7 is evidence-based tier **downshift**, calibrated from published ledger\ndata: opus-to-sonnet demotion for items the planner over-marks `complex`,\nreview-tier scaling to diff size, and four-tier selector verdicts. Plus the\nverified-based cascade trigger named above (the finalizer retry landed).\nThe selector A/B (sonnet vs opus vs fable arms) keeps accumulating\n`itemLedger` data, and the results will be published whichever way they land,\nper [docs/eval-protocol.md](./docs/eval-protocol.md).\n\nThreat-aware execution (v2.9) has two items left. The Codex lane's\nper-dispatch policy check is blocked on an upstream Threatlas PR. **Open and\nunblocked, needs a design pass before it's built:** a profile-posture fallback\nfor the risk override — today only `tier0Paths` forces anything, so a profile\nthat declares high sensitivity without naming paths gets no enforcement (see\nKnown limits). The open question is the actual rule, not the plumbing: does\n`ipSensitivity: high` force every item to Opus regardless of triviality (blunt,\nexpensive, but matches \"everything here is sensitive\" literally), or does it\nneed a threshold on `sensitiveCore` or a category-based rule to avoid pricing\ntrivial items at the top tier just because the project as a whole is\nsensitive? Whichever rule is picked needs its own `riskForced` reason distinct\nfrom path-matching, so the ledger can tell the two apart. Once a harness\nprovider seam exists, the natural follow-on is turning `recommendedTier` from\nadvisory into actual cross-provider execution for the tiers a profile clears\nfor cheap work.\n\n## Files in this repo\n\n- `purser.js` is the workflow script.\n- `HARNESS.md` is the runtime contract purser depends on, including operational\n  notes for running it from nested or headless sessions.\n- `docs/TESTPLAN.md`, `docs/TESTREPORT.md`, `docs/eval-protocol.md`, and\n  `docs/naming-map.md` are the evidence and provenance trail.\n- `tests/` and `phase0-checks/` are the zero-spend suites CI runs on every push.\n- `scripts/ledger-aggregate.mjs` builds the selector calibration table from run\n  telemetry; `scripts/sync-from-local.mjs` is the mechanical naming-map sync.\n- `example-PURSER.md` is a longer operator guide; `example-settings.json` holds\n  placeholder harness settings.\n\n## License\n\n[MIT](./LICENSE).\n","readmeFilename":"README.md","_rev":"1-1f1ad5b91ca057cf07e2db3764f12060"}