{"_id":"@bnecko/orqlaude","name":"@bnecko/orqlaude","dist-tags":{"latest":"0.12.1"},"versions":{"0.12.1":{"name":"@bnecko/orqlaude","version":"0.12.1","description":"Multi-agent orchestrator for Claude Code. One primary session decomposes a complex task, gets a single budget approval, then dispatches N parallel Agnets via the Desktop app's native spawn_task (with explicit fallbacks). Tracks cost/status via JSONL tails","type":"module","bin":{"orqlaude":"dist/cli.js","orql":"dist/cli.js","orqlaude-mcp":"dist/server.js"},"main":"./dist/server.js","scripts":{"build":"tsc && chmod +x dist/cli.js dist/server.js","dev":"tsc --watch","start":"node dist/server.js","typecheck":"tsc --noEmit","test":"npm run build && node --test dist/__tests__/*.test.js","prepublishOnly":"npm run build"},"keywords":["claude","claude-code","mcp","orchestrator","multi-agent","agents","anthropic"],"author":{"name":"matthew-demidoff"},"license":"MIT","repository":{"type":"git","url":"git+https://github.com/matthew-demidoff/orqlaude.git"},"bugs":{"url":"https://github.com/matthew-demidoff/orqlaude/issues"},"homepage":"https://github.com/matthew-demidoff/orqlaude#readme","dependencies":{"@modelcontextprotocol/sdk":"^1.0.4","zod":"^3.23.8"},"devDependencies":{"@types/node":"^22.7.5","typescript":"^5.6.3"},"engines":{"node":">=22"},"publishConfig":{"access":"public"},"gitHead":"570a7530c7fb692c6c9b67f93d9199ef05e5fed1","types":"./dist/server.d.ts","_id":"@bnecko/orqlaude@0.12.1","_nodeVersion":"26.0.0","_npmVersion":"11.12.1","dist":{"integrity":"sha512-qNFhJy4pwSWCJve+X4Rz4fL61ZFbb4KDK6vx1yCa8zgBhatnXyDix5AR4lJ91tLzpV4PRPbkdZ/GeNc6RcXkgw==","shasum":"b2e95030f2850f40a1a7f96c74c6d6f6ee2d1d76","tarball":"https://registry.npmjs.org/@bnecko/orqlaude/-/orqlaude-0.12.1.tgz","fileCount":271,"unpackedSize":1468321,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCICbqFnklKous/gHHV8Dr6eZwgbQwoOwAaoRpoR6lsDD0AiAY3VCT1ibU2d+NoTqoRFWjCjaWmspzCVnFP4zREvdJvg=="}]},"_npmUser":{"name":"synaplink","email":"demidovmatwey@gmail.com"},"directories":{},"maintainers":[{"name":"synaplink","email":"demidovmatwey@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/orqlaude_0.12.1_1780659512610_0.7917714369495503"},"_hasShrinkwrap":false}},"time":{"created":"2026-06-05T11:38:32.448Z","0.12.1":"2026-06-05T11:38:32.776Z","modified":"2026-06-05T11:38:32.991Z"},"maintainers":[{"name":"synaplink","email":"demidovmatwey@gmail.com"}],"description":"Multi-agent orchestrator for Claude Code. One primary session decomposes a complex task, gets a single budget approval, then dispatches N parallel Agnets via the Desktop app's native spawn_task (with explicit fallbacks). Tracks cost/status via JSONL tails","homepage":"https://github.com/matthew-demidoff/orqlaude#readme","keywords":["claude","claude-code","mcp","orchestrator","multi-agent","agents","anthropic"],"repository":{"type":"git","url":"git+https://github.com/matthew-demidoff/orqlaude.git"},"author":{"name":"matthew-demidoff"},"bugs":{"url":"https://github.com/matthew-demidoff/orqlaude/issues"},"license":"MIT","readme":"# @synaplink/orqlaude\n\nMulti-agent orchestrator for Claude Code. One primary Claude session decomposes a complex task into N parallel **Agnets** (orqlaude's name for spawned workers), gets a single user approval, then dispatches each Agnet (in its own session and worktree) via the Claude Desktop app's native `mcp__ccd_session__spawn_task`. Tracks cost/tokens via JSONL tails, brokers messages between Agnets, detects hallucination, manages PRs, streams updates to your Telegram, and can spawn a reviewer Agnet per PR at the end.\n\nThe name is **orq**hestrator + **Claude**.\n\n> Status: **v0.12.1** — 35+ tools, 236 tests passing, CI green. Live HTML dashboard (`orql web` — keyboard shortcuts, click-to-copy, SSE w/ heartbeat + CSP), cost analytics with sparklines (`orql cost`), goal quickstart wizard (`orql goal new`), token-first budgets (Max-friendly), self-registering child agents, hallucination detection, file-claim broker, durable memory + backlog, autopilot daemon, audit log, resumability, auto-review pipeline, and a **Telegram bot** for fleet notifications + remote control.\n\n## Why orqlaude exists\n\nA single Claude agent is great at focused work but slow at multi-region refactors. You can manually spawn parallel sessions via `spawn_task`, but you lose budget oversight, cross-agent coordination, and a single place to see \"what's the fleet doing right now?\"\n\norqlaude is the thin layer that adds those things. It never spawns processes itself — the Desktop app's `spawn_task` does that — but it owns the *plan*, the *budget*, the *broker*, and the *aggregation*.\n\n## How it fits together\n\n```\n                          ┌──────────────────────┐\n                          │   PRIMARY CLAUDE     │\n                          └─────────┬────────────┘\n   ┌─── orqlaude.create_plan ──────►│\n   │   orqlaude.request_approval ──►│  (relays via AskUserQuestion)\n   │   orqlaude.confirm        ────►│\n   │   orqlaude.next_task      ────►│\n   │   ccd_session.spawn_task  ────►├─── chip ─► ┌──────────┐\n   │                                │            │ child #1 │ ─► auto-registers via checkin\n   │                                │            │ session  │\n   │   orqlaude.next_task      ────►│            └────┬─────┘\n   │   ccd_session.spawn_task  ────►├─── chip ─►      │\n   │                                │            ┌────▼─────┐\n   │                                │            │ child #2 │\n   │   orqlaude.status         ────►│ ◄────── claim_files, post_note\n   │   orqlaude.poll_notes     ────►│ ◄────── PR url via post_note\n   │   orqlaude.send_message   ────►│\n   │   orqlaude.collect        ────►│\n   │   orqlaude.review_prs     ────►├─── chip ─► reviewer #1\n   │                                ├─── chip ─► reviewer #2\n   └────────────────────────────────┘\n```\n\n## Install\n\n```sh\nnpm install -g @synaplink/orqlaude   # CLI + MCP server\n```\n\nThen wire it into Claude Desktop's MCP config in one command:\n\n```sh\ncd /path/to/your/project\norql setup\n```\n\n`orql setup` **patches** Claude Desktop's `claude_desktop_config.json` in place — adds an `orqlaude` MCP server entry pointed at this project's state dir, and **preserves every other server you have configured** (lm-studio, etc.) plus the entire `preferences` block. Writes a timestamped `.bak` before changing anything. Re-runs are idempotent — if the entry is already correct, it reports `already correct; nothing to do` and exits without touching the file.\n\nFlags:\n- `--state-dir <path>` override the default (which walks up from cwd for a `.git`)\n- `--config-path <path>` override Claude Desktop's config path\n- `--yes` skip prompts\n\nFully quit and relaunch Claude Desktop after running. The `mcp__orqlaude__*` tools will then appear in your sessions.\n\nIf you'd rather edit the config yourself, the entry should look like:\n\n```json\n{\n  \"mcpServers\": {\n    \"orqlaude\": {\n      \"command\": \"npx\",\n      \"args\": [\"-y\", \"-p\", \"@synaplink/orqlaude\", \"orqlaude-mcp\"],\n      \"env\": {\n        \"ORQLAUDE_STATE_DIR\": \"/absolute/path/to/your/project/.orqlaude\"\n      }\n    }\n  }\n}\n```\n\nThe `ORQLAUDE_STATE_DIR` env var is important — it pins state to a specific path so the MCP server (running with `cwd=/` on some hosts) and the Telegram bot (running in your project) share the same state file.\n\n## Spawning Agnets: which tool to use\n\norqlaude itself doesn't spawn processes — it returns prompts and lets the orchestrator pick a spawn tool. Use them in this priority order:\n\n| Priority | Tool | Isolation | Visibility | When to use |\n|---|---|---|---|---|\n| **1** | `mcp__ccd_session__spawn_task` | git worktree per Agnet | Claude Desktop Code sidebar | **Default.** Worktree-isolated, sandbox-clean, the Agnet shows up as its own session you can switch into. |\n| 2 | Host's `Agent` tool (Claude Code built-in) | none — shares your cwd | tool-use only, not a separate session | Faster, no chip-click. **Loses worktree isolation** — Agnets may race on shared files. `claim_files` from the broker is your only collision signal. |\n| 3 | Shell out `claude -p --worktree …` | explicit `--worktree` flag | JSONL on disk — not in sidebar until Desktop restart | Headless / cron / non-Desktop hosts. |\n\n`next_task` returns a `spawn_strategies[]` array with ready-to-call args for each option, so the orchestrator can pick deliberately. **Picking by habit is the most common way to bypass orqlaude's isolation guarantees** — check the returned strategies and make a conscious choice.\n\n### Orphan detection\n\nIf an Agnet is dispatched but doesn't call `mcp__orqlaude__checkin` within 60 s, `status()` flags it in `orphan_alerts[]`. Common cause: the orchestrator used a non-`ccd_session__spawn_task` path and the Agnet skipped (or never reached) the protocol footer that tells it to register.\n\n## Tool reference\n\n### Planning (primary Claude)\n\n| Tool | Purpose |\n|---|---|\n| `create_plan(root_task, tasks[], budget_cap_tokens?, model_for_estimate?, effort_multiplier?)` | Register a fleet. Returns `plan_id`. Budget is in TOKENS (Max-plan friendly); USD is informational. |\n| `estimate(plan_id, model?, effort_multiplier?)` | Recompute cost/duration estimates. |\n| `request_approval(plan_id)` | Returns `approval_token` and a prebuilt `ask_user_question` payload. Surfaces your daily token usage from the Desktop app's `buddy-tokens.json`. |\n| `confirm(plan_id, approval_token)` | Lock the plan after user approves. |\n\n### Dispatch (primary Claude)\n\n| Tool | Purpose |\n|---|---|\n| `next_task(plan_id)` | Pull the next pending task. Returned `prompt` embeds `plan_id` + `task_id` and instructs the agent to self-register via `checkin` on its first turn. |\n| `status(plan_id)` | Per-agent live snapshot: cost, tokens, last activity, current tool, terminated yes/no, **hallucination report**. Auto-cancels and STOPs all agents if total tokens exceed the cap. |\n| `collect(plan_id)` | Aggregated PR URLs, summaries, costs, exit reasons. |\n| `review_prs(plan_id, auto_approve?, budget_cap_tokens?)` | Spawn a reviewer agent against each PR produced by `plan_id`. Creates a new \"review plan\" auto-approved by default. |\n| `register_spawn(plan_id, task_id, session_id)` | Manual fallback if a child fails to self-register. Rarely needed. |\n\n### Broker\n\n| Tool | Caller | Purpose |\n|---|---|---|\n| `checkin(session_id, task_id?)` | child agent | **First call**: pass `task_id` to self-register. Subsequent calls: pull queued messages, see STOP signals, ack state of blocking notes. |\n| `post_note(session_id, text, blocking?, pr_url?)` | child agent | Share findings or report a PR URL. `blocking: true` pauses until acked. |\n| `claim_files(session_id, paths[], reason?)` | child agent | Register intent to edit specific files. Conflicting claims by other agents surface to the caller. |\n| `release_files(session_id, paths[])` | child agent | Release claims after finishing. |\n| `poll_notes(plan_id, since_ts?, mark_acked?)` | primary Claude | Read agent notes; ack blocking ones to unblock posters. |\n| `send_message(plan_id, to_session_id, text, from_task_id?, kind?)` | primary Claude | Queue a directed message. `kind: \"stop\"` triggers child commit-and-exit. |\n\n### Lifecycle\n\n| Tool | Purpose |\n|---|---|\n| `kill_task(plan_id, task_id, reason)` | Queue STOP broker message; returns session_id ready for `archive_session`. Use for hallucinating/looping agents. |\n| `resume_plan(plan_id)` | Pick up an in-flight plan after a Desktop-app restart or new session. Refreshes per-task status from JSONL, returns a \"do this next\" hint. |\n| `list_plans(include_collected?)` | All plans known to orqlaude in this project, active first. |\n\n### Broker-to-user (v0.4+, expanded in v0.5)\n\nThese let primary Claude push messages to and pull answers from the **user** (via Telegram if configured, with a `/respond` text fallback).\n\n| Tool | Purpose |\n|---|---|\n| `notify_user(plan_id, text, urgency?, task_id?)` | One-way push to user's Telegram. urgency = `low`/`normal`/`high` (affects emoji). Returns immediately. |\n| `request_user_response(plan_id, prompt, options?[], timeout_sec?, task_id?)` | Ask the user a question. With `options`, Telegram shows inline-keyboard buttons; without, user replies via `/respond <short_id> <text>`. Returns `request_id` + `short_id`. Defaults to a 10-minute timeout. |\n| `poll_user_response(request_id)` | Returns `status: pending\\|answered\\|timed_out\\|cancelled` + `response` once available. Safe to poll repeatedly. |\n| `stream_to_user_start(plan_id, title, initial_content?, task_id?)` | **v0.5+** Open a streaming Telegram message. Returns `stream_id`. |\n| `stream_to_user_append(stream_id, chunk)` | **v0.5+** Append a chunk; notifier edits the Telegram message in place (throttled ~1 edit/1.5s). |\n| `stream_to_user_end(stream_id, final_chunk?)` | **v0.5+** Finalize the stream — adds a `✓` marker to the message. |\n\nWithout a running `orqlaude tg start`, `notify_user` queues silently, `request_user_response` will always `timed_out`, and streaming tools accumulate content in state but no message lands on Telegram. Fall back to `AskUserQuestion` if Telegram is unavailable.\n\n#### Streaming transport\n\norqlaude streams by opening a single Telegram message with `sendMessage`, then calling `editMessageText` on each append. The final `stream_to_user_end` does one more edit appending a `✓` marker.\n\n(v0.5.1 briefly used `sendMessageDraft` for animated, ephemeral previews — reverted in v0.5.4 after that endpoint proved unreliable in the standard Bot API.)\n\nLimits to know:\n- A Telegram message tops out at 4096 chars. orqlaude caps stream content at 3800 to leave room for the title + completion marker; further appends are silently truncated.\n- Edits are rate-limited (~1/sec per message); orqlaude throttles to 1.5 s between edits per stream.\n- If you need to stream more than 4 kb of output, start a new stream when you're approaching the cap.\n\n### Health\n\n| Tool | Purpose |\n|---|---|\n| `ping(echo?)` | Returns version, cwd, state_dir, state_dir_source, warnings[], node, pid. First call after install to verify wiring + state-dir resolution. |\n\n## End-to-end walkthrough\n\nUser says: *\"Refactor the auth system — magic-link login, update the docs, add tests.\"* You judge it as parallelizable.\n\n1. `orqlaude.create_plan` with 3 subtasks (auth-core, docs, tests), `budget_cap_tokens: 600000`.\n2. `orqlaude.request_approval` → returns `approval_token` and a prebuilt question payload showing your remaining daily quota.\n3. You call `AskUserQuestion` with that payload. User picks \"Approve and spawn\".\n4. `orqlaude.confirm`.\n5. Loop three times:\n   - `orqlaude.next_task` → returns a task with the wrapped prompt\n   - `mcp__ccd_session__spawn_task` with the title/prompt/tldr → user clicks the chip\n   - The spawned agent calls `orqlaude.checkin(session_id, task_id)` on first turn → self-registers\n6. Periodically: `orqlaude.status` (shows hallucination scores) + `orqlaude.poll_notes`. Forward cross-cutting info via `send_message`. If an agent goes off the rails, `kill_task`.\n7. Agents call `orqlaude.post_note(..., pr_url=...)` when their PR is open.\n8. `orqlaude.collect` → three PR URLs and summaries.\n9. **NEW**: `orqlaude.review_prs(plan_id)` → spawns three reviewer agents, one per PR. Each reviews, runs tests, posts findings. You aggregate the second-round notes.\n\n## Asking the user (v0.10.4 pattern)\n\n`ask_user` and its companion `wait_for_user_response` are the bounded-block loop primary Claude uses to put a question on Telegram and stay alive past the MCP host's 60s per-request timeout.\n\nThe split exists because Claude Desktop and Claude Code both use the SDK default `DEFAULT_REQUEST_TIMEOUT_MSEC = 60000`. v0.10.2's progress notifications turned out to be ignored unless the client passes `resetTimeoutOnProgress: true` (it doesn't), so a single blocking call can't outrun the host. Instead, `ask_user` blocks at most 45s (the new `initial_block_sec`, capped at 45). The question's overall lifetime is `total_timeout_sec` (default 900s, max 3600s) -- that's how long it stays answerable. If the user replies inside the first window, you get `status: \"answered\"`. Otherwise you get `status: \"still_pending\"` with a `short_id`, and the caller must invoke `wait_for_user_response(short_id)` to keep waiting.\n\nLoop pattern (TS pseudocode):\n\n```ts\nlet result = await ask_user({\n  prompt: \"Approve the auth refactor plan?\",\n  options: [\"Approve\", \"Hold off\"],\n  total_timeout_sec: 1800,\n});\n\nwhile (result.status === \"still_pending\") {\n  result = await wait_for_user_response({ short_id: result.short_id });\n}\n// result.status is now \"answered\" / \"timed_out\" / \"cancelled\"\n```\n\nEach call stays safely under 60s. A fast answer is one round-trip; a 5-minute wait is roughly 7 round-trips, no `ScheduleWakeup`-and-come-back required.\n\nTelegram side is plain text only (no Markdown), so escaping bugs can't silently swallow a send. The notifier ships each question with `force_reply` enabled -- the user just types and their reply carries `reply_to_message.message_id`, which the bot matches back to the request. Inline-keyboard buttons fire the same answer path when `options` are provided; `/respond <short_id> <text>` remains as a manual fallback.\n\n## Autopilot daemon\n\nA persistent orchestrator that ticks every 10 seconds, picks goals off the backlog, auto-reviews PRs, retries failed Agnets, and watches the budget. Opt-in -- nothing runs in the background unless you start it.\n\nFive tick-loop phases:\n\n1. **Reconcile state** -- for every spawned Agnet, refresh from JSONL, PID, and exit-record; promote `died_at_launch` / `done` / `failed`.\n2. **Recover from failures** -- classify each failure via a Plan-billed `claude -p` turn, then retry with backoff, spawn a debugger Agnet, or escalate to the user via Telegram.\n3. **Auto-review PRs** -- fetch the diff, run a reviewer turn, apply the fleet template's auto-merge rule, and either `gh pr merge` or `gh pr comment`.\n4. **Pick the next goal** -- when the fleet is idle and autopilot is unpaused, pull the highest-priority unblocked goal from the backlog and prompt the user via Telegram.\n5. **Watch the budget** -- yellow / orange / red thresholds; auto-pause at orange.\n\nCLI:\n\n```sh\norql autopilot start            # foreground; daemonize with launchd / systemd / nohup\norql autopilot stop\norql autopilot pause            # stop picking new work; in-flight Agnets keep running\norql autopilot resume\norql autopilot status\n```\n\nPlan-billing note: the daemon **never** talks to the Anthropic API. Every intelligent decision (failure classifier, PR reviewer, Telegram intent classifier, template suggester) is a `claude -p` invocation. On the Max plan that bills like an interactive Claude Code session, and cache reads are free, so a full day of ticking burns a tiny fraction of quota.\n\n## Durable memory\n\nA `memory.json` file at `<state_dir>/memory.json` holds long-lived facts that outlive plan lifecycles. Four spirit-themed categories, each with a different surfacing rule:\n\n- **lore** -- facts about the user. Pinned, slow churn, injected into every spawned Agnet prompt. _Example: \"Russian comments in CRM templates\", \"no auto-deploy on Fridays.\"_\n- **playbook** -- code conventions. Scope-tagged by path-glob; injected when a fleet's scope overlaps. _Example: \"migrations live in `<app>/migrations/`\", \"use AntD ConfigProvider for dark mode.\"_\n- **ledger** -- decisions plus rationale. Append-only; surfaced when a similar decision recurs. _Example: \"Sonnet over Opus for transcription, latency mattered more than depth.\"_\n- **atlas** -- project map. Auto-updated by the post-PR review with one entry per touched file mapping path to purpose.\n\nPinned entries always load. Scope-tagged entries (typically playbook and atlas) auto-inject into a spawn prompt when the Agnet's worktree scope matches any of the entry's globs -- you don't have to remember to thread conventions through manually.\n\nMCP tools:\n\n```\nremember(category, key, value, { pinned?, scopeGlobs?, rationale? })\nrecall(category?, key?, scopeMatch?)\nforget(category, key)\ncompose_memory_context(scopeGlobs?, max_tokens?)\n```\n\nOlder entries with the same `(category, key)` are soft-superseded: kept for history but invisible to read paths. `compose_memory_context` is the function the spawn pipeline uses internally; call it directly to preview the block that will be injected for a given scope before you commit to spawning.\n\n## Backlog\n\nA `backlog.json` file holds `Goal` records -- durable task descriptions the daemon (or you) can pick from when idle.\n\nShape:\n\n```json\n{\n  \"id\": \"g_8f3...\",\n  \"title\": \"Rotate JWT signing keys quarterly\",\n  \"priority\": 70,\n  \"deadlineAt\": \"2026-06-30T00:00:00Z\",\n  \"dependsOn\": [\"g_5c1...\"],\n  \"createdAt\": \"2026-05-12T10:14:00Z\",\n  \"status\": \"pending\"\n}\n```\n\nPriority is 0-100. `deadlineAt` boosts effective priority as the deadline approaches (linear ramp over the last 14 days, so something due tomorrow with priority 40 beats a no-deadline priority 70 item). `dependsOn` is a list of goal ids; a goal is blocked until every parent has `status: \"done\"`.\n\nMCP tools:\n\n```\nenqueue_goal(title, { priority?, deadlineAt?, dependsOn?, source? })\nlist_goals({ status?, includeBlocked? })\nupdate_goal(id, { priority?, deadlineAt?, dependsOn?, status? })\npick_next_goal()\n```\n\n`pick_next_goal` returns the highest-priority unblocked pending goal, factoring in deadline boost. The autopilot daemon calls this on every idle tick; when something comes back it surfaces the goal to the user via Telegram for confirmation before spawning a fleet, so you keep approval-in-the-loop even when the orchestrator is running unattended.\n\n## Fleet templates\n\nEight named patterns ship out of the box. Each defines a default Agnet layout, a suggested model per role (haiku / sonnet / opus), a default budget, and an `AutoMergeRule` the daemon applies to PRs produced by the fleet.\n\n| id | what it does | auto-merge rule |\n|---|---|---|\n| `backend-feature` | Django/DRF: model + migration + serializer + viewset + admin + tests | requireCi, maxLoc 2500 |\n| `frontend-feature` | React/AntD: components + hooks + i18n + tests | requireCi, maxLoc 2000 |\n| `migration-only` | Schema change with backwards-compat reviewer (opus) | requireReviewerApprove; `blockOnMigrations: false` (migrations are the point) |\n| `audit-sweep` | Multiple haiku auditors + sonnet synthesizer (read-only) | requireReviewerApprove (no merge) |\n| `dep-upgrade` | Dep version bump + breaking-change patches + reviewer | requireCi, requireReviewerApprove |\n| `i18n-pass` | Audit, then translator pass | requireCi, maxLoc 3000 |\n| `test-coverage-fill` | Parallel testers; blocks PRs that touch prod code | requireCi, blockOnPaths globs for non-test files |\n| `bug-hunt` | Reproducer Agnet then fixer Agnet (sequential) | requireReviewerApprove, requireCi |\n\nMCP tools:\n\n```\nlist_fleet_templates()\nsuggest_fleet_template(goal_text)    # Plan-billed turn picks the best fit\napply_fleet_template(template_id, { goal, scope, budget_override? })\n```\n\n`suggest_fleet_template` makes a single `claude -p` call that returns `{ template_id, confidence, reason }`. The autopilot daemon uses this to turn a freeform goal description into a concrete fleet definition without manual plan authoring. `apply_fleet_template` then expands the chosen template into a real plan via `create_plan`, with the template's auto-merge rule attached for later use by the auto-review pipeline.\n\n## Auto-PR-review\n\nWhen the autopilot daemon is running, every PR produced by a template-driven fleet gets a reviewer turn and an auto-merge attempt.\n\nThe reviewer turn runs `gh pr view` for the diff and metadata, feeds them into a strict-JSON `claude -p` prompt, and parses the response `{ verdict, blockers, suggestions, summary }` where `verdict` is `APPROVE` / `REQUEST_CHANGES` / `COMMENT`. The summary is appended as a PR comment regardless of verdict so you have a paper trail.\n\nThe fleet template's `AutoMergeRule` then decides whether to merge:\n\n```json\n{\n  \"requireReviewerApprove\": true,\n  \"requireCi\": true,\n  \"maxLoc\": 2500,\n  \"blockOnMigrations\": false,\n  \"blockOnPaths\": [\"**/settings.py\", \"**/secrets/**\"]\n}\n```\n\n- `requireReviewerApprove` -- the reviewer's verdict must be `APPROVE`.\n- `requireCi` -- `gh pr checks` must come back all-green.\n- `maxLoc` -- additions plus deletions under cap. Larger PRs route to user.\n- `blockOnMigrations` -- refuses PRs that add files under `*/migrations/` (used by templates that aren't supposed to touch schema).\n- `blockOnPaths` -- refuses PRs touching specific globs.\n\nIf every check passes the daemon runs `gh pr merge --squash --auto --delete-branch`. Otherwise it `gh pr comment`s with the verdict and blockers and leaves the PR open. Each review writes a `ledger` memory entry so the next fleet inherits the rationale.\n\n## Cost guardrails\n\nA `guardrails.json` rolling ledger tracks billed tokens against two windows: a 5-hour rolling window (matching Anthropic's Plan reset cadence) and a per-local-day soft cap (default 30M billed tokens).\n\nThree threshold bands on the rolling window:\n\n- **yellow** at 60% -- notify the user, daemon slows inter-tick interval.\n- **orange** at 80% -- daemon auto-pauses; refuses to start new fleets; in-flight Agnets keep running but no new spawns happen.\n- **red** at 95% -- halt entirely; await user `/resume` after the next 5h reset.\n\nThe day soft-cap applies independently of the rolling window -- you can sit comfortably in green on the 5h window and still cross the day cap if you've been running multiple windows back to back. Both checks happen on every autopilot tick, and the orange auto-pause uses the same code path as `orql autopilot pause`: it surfaces in `orql autopilot status` and is undone with `orql autopilot resume` once the window has rolled.\n\n## State\n\norqlaude resolves its state directory at startup using this order (first match wins):\n\n1. **`ORQLAUDE_STATE_DIR`** env var — explicit override.\n2. **Git worktree**: if `<cwd>/.git` is a file pointing at `<main>/.git/worktrees/<n>`, use `<main>/.orqlaude` so spawn_task'd children share state with the parent fleet.\n3. **Project root**: if cwd is writable and contains `.git/`, `package.json`, `pyproject.toml`, `Cargo.toml`, or `go.mod`, use `<cwd>/.orqlaude`.\n4. **Home fallback** (covers MCP hosts that launch with `cwd=/`): `~/.orqlaude/projects/<basename>-<sha256hash>/`. orqlaude writes a one-line note to stderr when this kicks in.\n\nCheck what got resolved: `orqlaude where`, or call `mcp__orqlaude__ping` (returns `state_dir` and `state_dir_source`).\n\nFiles inside the dir:\n- `orqlaude-state.json` — plans, tasks, notes, messages, claims. Atomic-write via temp+rename.\n- `audit.jsonl` — append-only log of every tool call. Inspect with `orqlaude history` or `tail -f .orqlaude/audit.jsonl | jq`.\n- `lock` — sidecar file lock used by `update()` for cross-process serialization.\n\n`<project>/.orqlaude/` is `.gitignore`d.\n\n## Hallucination detection\n\nWhen you call `status(plan_id)`, every agent's snapshot includes a `hallucination` object with `score` (0–1), `level` (`clean`/`minor`/`moderate`/`severe`), and `concerns: string[]`. The aggregated `hallucination_alerts` array surfaces only agents at `moderate` or above so an orchestrator can react quickly.\n\n**What gets flagged:**\n\n1. **Path-existence** — every `file_path` arg in `Read`/`Edit`/`Write`/`Grep`/`Glob`/`MultiEdit`/`NotebookEdit`/`NotebookRead` is checked against the worktree. >30% missing or ≥3 missing = moderate/severe. Catches \"agent is editing a file it imagined.\"\n2. **Tool-pattern sanity**:\n   - **Edit-without-prior-Read**: agent edits a file it never read → it's guessing at the content.\n   - **Tight loop**: same tool call (name+args) ≥3× → likely stuck.\n   - **Commit-without-tests**: `git commit` without a prior test/lint Bash call → commit may be broken.\n\n**How to react** in your orchestrator code:\n\n| level | suggested response |\n|---|---|\n| `clean` | Nothing. |\n| `minor` | Note but continue. |\n| `moderate` | `send_message` to the agent with a nudge (\"re-read X.ts before editing\"), or `request_stop` if the work is salvageable. |\n| `severe` | `kill_task` and consider re-spawning with a clearer prompt. |\n\nFalse positives are acceptable here — we surface concerns, we don't auto-kill. A v0.4 addition is opt-in second-model cross-validation (a cheap Haiku reading the agent's recent turns and rating \"is this lost?\").\n\n## CLI\n\nTwo binaries are installed: `orqlaude` and the short alias `orql`. Use whichever feels right. All commands work the same.\n\n### Live (v0.6+)\n\n```sh\norql watch <plan_id>            # live-updating dashboard (1Hz, Ctrl-C to exit)\norql tail [plan_id]             # tail -f the audit log; filter by plan if given\norql open <plan_id>             # open every PR from a plan in your browser\norql doctor                     # end-to-end health check\norql about                      # the easter egg\n```\n\n### Local desktop notifications (macOS)\n\n```sh\norql notify on                  # enable; the Telegram bot will also fire osascript banners\norql notify off                 # disable\norql notify test                # send a test notification\norql notify status              # is it on?\n```\n\n### Read-only inspection (`--json` on every one)\n\n\n\n```sh\norql                            # banner + active-plan summary\norql list                       # every plan in this project\norql status [plan_id]           # if omitted, picker prompts\norql show [plan_id]             # raw plan JSON\norql history --limit 50         # tail audit log\norql where                      # show resolved state dir\norql help\n```\n\nEvery read command supports `--json` to emit machine-readable output for scripting.\n\nRead-only. For active orchestration, use the MCP from inside Claude Code.\n\n### Branding & colors (v0.5+)\n\nCLI output uses the Anthropic palette via ANSI truecolor:\n\n| Color | Hex | Purpose |\n|---|---|---|\n| Claude Coral | `#DA7756` | Headings, brand accents, running tasks, Agnet names |\n| Cream | `#F5F4EE` | Secondary emphasis, token counts |\n| Crimson | `#BB5944` | Errors, failed/cancelled tasks |\n| Charcoal | `#2A2926` | Body text (terminal default usually) |\n| Sand | `#B9B6AB` | Captions, separators, hints |\n\nColors disable automatically when stdout isn't a TTY, when `NO_COLOR` is set ([no-color.org](https://no-color.org/)), or when `TERM=dumb`. Force-enable with `FORCE_COLOR=1`.\n\n## Repo layout\n\n```\norqlaude/\n├── package.json                # @synaplink/orqlaude\n├── tsconfig.json\n├── .mcp.json                   # local dev wiring\n├── .mcp.json.template          # production wiring (npx-based)\n├── .github/workflows/ci.yml    # typecheck + build + test\n├── src/\n│   ├── server.ts               # MCP stdio entry\n│   ├── cli.ts                  # `orqlaude` CLI binary\n│   ├── lib/\n│   │   ├── state.ts            # JSON-backed ledger, schema v2\n│   │   ├── budgeting.ts        # token-first budget, daily quota reader\n│   │   ├── pricing.ts          # USD pricing table (informational)\n│   │   ├── hallucination.ts    # deterministic detectors\n│   │   ├── jsonl_tail.ts       # cached byte-offset session tail\n│   │   └── audit.ts            # append-only audit log\n│   ├── tools/\n│   │   ├── ping.ts\n│   │   ├── planning.ts         # create_plan, estimate, request_approval, confirm\n│   │   ├── dispatch.ts         # next_task, register_spawn, status, collect\n│   │   ├── broker.ts           # checkin, post_note, claim_files, release_files, poll_notes, send_message\n│   │   ├── lifecycle.ts        # kill_task, resume_plan, list_plans\n│   │   └── review.ts           # review_prs\n│   └── __tests__/\n│       ├── state.test.ts\n│       └── hallucination.test.ts\n└── dist/                       # tsc output (published to npm)\n```\n\n## Telegram bot\n\norqlaude can notify you on Telegram when fleet events happen and accept commands from your phone.\n\n```sh\n# One-time setup (creates ~/.orqlaude/telegram.json, mode 600)\norqlaude tg setup\n# (paste your bot token from @BotFather)\n\n# Message your bot /start in Telegram to learn your user id, then:\norqlaude tg whitelist <your_user_id> --owner --label \"you\"\n\n# Run the bot (foreground; daemonize with launchctl / systemd / nohup as you prefer)\ncd /path/to/your/project\norqlaude tg start\n```\n\n**Notifications pushed to you:**\n- 📋 New plan created\n- ✅ Plan approved (spawn imminent)\n- ✓ Task done (with PR URL)\n- ❌ Task failed / 🛑 cancelled\n- 📝 New agent note (with severity from `post_note`)\n- 💸 Auto-cancel on budget overrun\n- 🎉 Fleet collected\n\n**Commands you can send (whitelisted users only):**\n- `/plans` — active plans\n- `/status <plan_id>` — refreshed task list with token usage\n- `/show <plan_id>` — raw plan JSON\n- `/notes <plan_id>` — recent agent notes\n- `/kill <plan_id> <task_id> <reason>` — STOP a runaway agent\n- `/respond <short_id> <text>` — answer a `request_user_response` question (v0.4+)\n- Tap inline-keyboard buttons on any `request_user_response` with options (v0.4+)\n- `/whitelist <user_id> [label]` (owner-only) — add another user\n- `/help` / `/whoami`\n\nThe bot uses raw `fetch` against Telegram's Bot API — zero extra deps. State is shared with the MCP via the same `StateStore`, so commands take effect on the next status() / checkin().\n\n## Known gaps (v0.3 → v0.4 roadmap)\n\n- **Cost-learning estimates** — current baselines are tuned to a single Haiku probe. Future: write per-task realized costs to history and use moving averages.\n- **N chips = N clicks** — Anthropic's `spawn_task` is per-click by design. Worth filing as feedback. Until then, batch-spawn isn't possible through that API.\n- **Second-model hallucination check** — periodic Haiku cross-validation of recent activity, opt-in.\n- **Multi-project Telegram bot** — currently the bot watches a single project. Multi-project watching is a small extension to the config schema.\n- **Inline approve buttons in Telegram** — `/approve <plan_id>` and inline keyboards so you can confirm fleets from your phone.\n\n## Troubleshooting\n\n**Symptom: `ENOENT: no such file or directory, mkdir '/.orqlaude'` on `create_plan`.**\nYour MCP host launched orqlaude with `cwd=/`. v0.3.2+ auto-falls back to `~/.orqlaude/projects/...` but the explicit fix is to set `ORQLAUDE_STATE_DIR` in your `.mcp.json` env block (see `.mcp.json.template`). Verify with `mcp__orqlaude__ping` — it now returns `warnings` and `state_dir_source`.\n\n**Symptom: `spawn_via_cli` returned a PID but `status()` shows `died_at_launch` shortly after, with a stderr_excerpt + command_line.**\nThis is the v0.7.0 hardening doing its job. The child `claude -p` process exited within the 1.5s healthcheck window. Read `stderr_excerpt` (or open `stderr_path` directly) for the cause. Common ones:\n1. `claude` isn't authenticated on this user account (`claude auth status` to check; `claude auth login --claudeai` to fix).\n2. The `--mcp-config` JSON references a server entry that doesn't exist (rare; orqlaude validates this pre-spawn).\n3. The user's environment lacks something `claude` needs (HOME, locale).\nCopy the `command_line` field and paste it into a shell to reproduce by hand.\n\n**Symptom: spawn_task chip appeared, agent ran, but `status()` shows the task as `dispatched` forever.**\nThe child agent isn't calling `checkin` on its first turn — its prompt didn't get the orqlaude protocol block, or `mcp__orqlaude__checkin` isn't available in the spawned session. Manual unblock: `register_spawn(plan_id, task_id, session_id)` where session_id is the child's session UUID (find via `mcp__ccd_session_mgmt__list_sessions`). For the proper fix, make sure orqlaude is in the spawned worktree's `.mcp.json` (commit `.mcp.json` to the repo so worktrees inherit it).\n\n**Symptom: agents in worktrees can't see the parent fleet's plan.**\nv0.3.1+ resolves `<cwd>/.git` files (worktree pointers) back to the parent checkout's `.orqlaude`. If a child still can't find its plan, run `orqlaude where` inside the worktree — `source` should be `worktree`. If it's `home-fallback`, the worktree pointer is malformed or `.git` isn't where the resolver expected.\n\n**Symptom: Telegram bot stops sending notifications.**\nCheck `/tmp/orqlaude-tg.log` (if you used the launchd plist) or wherever the bot is logging. The most likely cause is a Markdown parse error from an unescaped `_`/`*`/`` ` ``/`[` in a task title or note. v0.3.1+ escapes these but anything user-supplied that bypasses our path (e.g. content posted manually via `post_note` to a stale older bot) can still hit it.\n\n## License\n\nMIT.\n","readmeFilename":"README.md","_rev":"1-5c100bee6ac564b1aad1ba3ed29c81a1"}