{"_id":"@appliqation/autopilot","_rev":"5-c1cdd436740b6d896e5a58032af60ef4","name":"@appliqation/autopilot","dist-tags":{"latest":"0.1.5"},"versions":{"0.1.1":{"name":"@appliqation/autopilot","version":"0.1.1","_id":"@appliqation/autopilot@0.1.1","maintainers":[{"name":"archana6","email":"accounts@appliqation.io"}],"bin":{"appliqation-autopilot":"dist/cli/index.js"},"dist":{"shasum":"05b42711c2427bbb8a11e821b7d9af4f4e105e32","tarball":"https://registry.npmjs.org/@appliqation/autopilot/-/autopilot-0.1.1.tgz","fileCount":17,"integrity":"sha512-u3a1kxya7Kk8r81lm6GEv1t5AyR13QRR8/UtJXbQWgRlbEUb6MRqVqLGC62xJOVTBV6YHn/Dh7P1mb2bctqDIw==","signatures":[{"sig":"MEYCIQDMM3rMjXTo0vj6zwAYKZHDN7yWuguikSuG879LdqP81gIhANfJdYnamLW0pJF/+ByvO3Qu2Ersp+K4VF2MaYThNFaZ","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":73512},"type":"module","engines":{"node":">=20"},"gitHead":"c79c1619580b733e1c61b7700492fb8ca2fae051","scripts":{"dev":"tsx src/cli/index.ts","lint":"eslint src --ext .ts","test":"vitest run","build":"tsc -p tsconfig.build.json","typecheck":"tsc -p tsconfig.json --noEmit","test:watch":"vitest"},"_npmUser":{"name":"archana6","email":"accounts@appliqation.io"},"_npmVersion":"11.7.0","description":"An agentic orchestrator that decides, per test case, whether to run autonomous testing, generate new automation, or raise a pull request for it — reasoning over real signal (current pass/fail state, flakiness, linked defects, coverage priority), not a fix","directories":{},"_nodeVersion":"23.10.0","dependencies":{"dotenv":"^16.4.7","commander":"^13.1.0","@appliqation/agent-core":"^0.1.0"},"_hasShrinkwrap":false,"devDependencies":{"tsx":"^4.19.3","vitest":"^3.2.7","typescript":"^5.8.2","@types/node":"^22.13.10"},"_npmOperationalInternal":{"tmp":"tmp/autopilot_0.1.1_1787298755121_0.23516619669077454","host":"s3://npm-registry-packages-npm-production"}},"0.1.2":{"name":"@appliqation/autopilot","version":"0.1.2","license":"MIT","_id":"@appliqation/autopilot@0.1.2","maintainers":[{"name":"archana6","email":"accounts@appliqation.io"}],"homepage":"https://github.com/appliqation/autopilot#readme","bugs":{"url":"https://github.com/appliqation/autopilot/issues"},"bin":{"appliqation-autopilot":"dist/cli/index.js"},"dist":{"shasum":"5be5d3f2c60e8f47852beda99514bd075c7fa327","tarball":"https://registry.npmjs.org/@appliqation/autopilot/-/autopilot-0.1.2.tgz","fileCount":17,"integrity":"sha512-MPeRjRpbsU89EfAZUPKzAmef+f4U7Sc5Q/Dcj8I0sEvGXYP/3ZD3pI3nRiyg10mEI7woa9kLZkAb1U6W2kUNvQ==","signatures":[{"sig":"MEUCID+zsJ771g+T6UmszojypY9JmOZiRIAPRi8qZylczTssAiEA+YgF58tpx5HE4bhXp1AONl5iUW3eEAc7ju/qYNVPyhE=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":75089},"type":"module","engines":{"node":">=20"},"gitHead":"a0d590772e7396f7b63a260473c234f7a02db059","scripts":{"dev":"tsx src/cli/index.ts","lint":"eslint src --ext .ts","test":"vitest run","build":"tsc -p tsconfig.build.json","typecheck":"tsc -p tsconfig.json --noEmit","test:watch":"vitest"},"_npmUser":{"name":"archana6","email":"accounts@appliqation.io"},"repository":{"url":"git+https://github.com/appliqation/autopilot.git","type":"git"},"_npmVersion":"11.7.0","description":"An agentic orchestrator that decides, per test case, whether to run autonomous testing, generate new automation, or raise a pull request for it — reasoning over real signal (current pass/fail state, flakiness, linked defects, coverage priority), not a fix","directories":{},"_nodeVersion":"23.10.0","dependencies":{"dotenv":"^16.4.7","commander":"^13.1.0","@appliqation/agent-core":"^0.1.5"},"_hasShrinkwrap":false,"devDependencies":{"tsx":"^4.19.3","eslint":"^9.39.5","vitest":"^3.2.7","@eslint/js":"^9.39.5","typescript":"^5.8.2","@types/node":"^22.13.10","typescript-eslint":"^8.67.0"},"_npmOperationalInternal":{"tmp":"tmp/autopilot_0.1.2_1787382410137_0.9713035636203156","host":"s3://npm-registry-packages-npm-production"}},"0.1.3":{"name":"@appliqation/autopilot","version":"0.1.3","license":"MIT","_id":"@appliqation/autopilot@0.1.3","maintainers":[{"name":"archana6","email":"accounts@appliqation.io"}],"homepage":"https://github.com/appliqation/autopilot#readme","bugs":{"url":"https://github.com/appliqation/autopilot/issues"},"bin":{"appliqation-autopilot":"dist/cli/index.js"},"dist":{"shasum":"5f92380a10c25eeaf224c8aee509df50f575c75c","tarball":"https://registry.npmjs.org/@appliqation/autopilot/-/autopilot-0.1.3.tgz","fileCount":17,"integrity":"sha512-ZyEHEhM15AM8uy4kVxGWt7r7Nu2eDP7xXSumLHYEJDI13AgFucuLx1pHUnMas153McUe31BfH4FJIxzXg+1RPA==","signatures":[{"sig":"MEUCIQDASMAme8CmCFNos6Q4/WVHSerz4OmjvXsz3PI19a9aTQIgJt3rhQyXWuBHfQgWJC0JMYcVS6jrpT7E2NL56O5xjvM=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":98836},"type":"module","engines":{"node":">=20"},"gitHead":"a4003490e095f68b6352135760c27eec0df20c5b","scripts":{"dev":"tsx src/cli/index.ts","lint":"eslint src --ext .ts","test":"vitest run","build":"tsc -p tsconfig.build.json","typecheck":"tsc -p tsconfig.json --noEmit","test:watch":"vitest"},"_npmUser":{"name":"archana6","email":"accounts@appliqation.io"},"repository":{"url":"git+https://github.com/appliqation/autopilot.git","type":"git"},"_npmVersion":"11.7.0","description":"An agentic orchestrator that decides, per test case (or an entire scenario/test set), whether to run autonomous testing, heal a stale selector, fix a defect, generate new automation, or raise a pull request for it — reasoning over real signal (current pas","directories":{},"_nodeVersion":"23.10.0","dependencies":{"dotenv":"^16.4.7","commander":"^13.1.0","@appliqation/agent-core":"^0.1.5"},"_hasShrinkwrap":false,"devDependencies":{"tsx":"^4.19.3","eslint":"^9.39.5","vitest":"^3.2.7","@eslint/js":"^9.39.5","typescript":"^5.8.2","@types/node":"^22.13.10","typescript-eslint":"^8.67.0"},"_npmOperationalInternal":{"tmp":"tmp/autopilot_0.1.3_1788164761658_0.8027328912586749","host":"s3://npm-registry-packages-npm-production"}},"0.1.4":{"name":"@appliqation/autopilot","version":"0.1.4","license":"MIT","_id":"@appliqation/autopilot@0.1.4","maintainers":[{"name":"archana6","email":"accounts@appliqation.io"}],"homepage":"https://github.com/appliqation/autopilot#readme","bugs":{"url":"https://github.com/appliqation/autopilot/issues"},"bin":{"appliqation-autopilot":"dist/cli/index.js"},"dist":{"shasum":"a9057271c4015e8644dec618c395f73bce4c2e39","tarball":"https://registry.npmjs.org/@appliqation/autopilot/-/autopilot-0.1.4.tgz","fileCount":19,"integrity":"sha512-1il2AevpoKWrSagZcGUiCexMQvlikaDbu4EatfNSnj3XuUXmKVz1+4iuaeJqFQt+IZY/VQSIBGMQ4Rhg6nYPVA==","signatures":[{"sig":"MEUCIQCdknuLPNRGMiPYwyHrUkocBtg0LITCRJoZCyDeelkaZgIgQxeaogerO0I1vp+QHudLgJWDvrdgILbqEN85uOtAhvo=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":120023},"type":"module","engines":{"node":">=20"},"gitHead":"533f662efb28be843c34bb620a4f450255911d1c","scripts":{"dev":"tsx src/cli/index.ts","lint":"eslint src --ext .ts","test":"vitest run","build":"tsc -p tsconfig.build.json","typecheck":"tsc -p tsconfig.json --noEmit","test:watch":"vitest"},"_npmUser":{"name":"archana6","email":"accounts@appliqation.io"},"repository":{"url":"git+https://github.com/appliqation/autopilot.git","type":"git"},"_npmVersion":"11.7.0","description":"An agentic orchestrator that decides, per test case (or an entire scenario/test set), whether to run autonomous testing, heal a stale selector, fix a defect, generate new automation, or raise a pull request for it — reasoning over real signal (current pas","directories":{},"_nodeVersion":"23.10.0","dependencies":{"dotenv":"^16.4.7","commander":"^13.1.0","@appliqation/agent-core":"^0.1.7"},"_hasShrinkwrap":false,"devDependencies":{"tsx":"^4.19.3","eslint":"^9.39.5","vitest":"^3.2.7","@eslint/js":"^9.39.5","typescript":"^5.8.2","@types/node":"^22.13.10","typescript-eslint":"^8.67.0"},"_npmOperationalInternal":{"tmp":"tmp/autopilot_0.1.4_1788433747536_0.4307795668478094","host":"s3://npm-registry-packages-npm-production"}},"0.1.5":{"_id":"@appliqation/autopilot@0.1.5","bin":{"appliqation-autopilot":"dist/cli/index.js"},"bugs":{"url":"https://github.com/appliqation/autopilot/issues"},"dist":{"shasum":"7f34edd3bf1abcd7652ebed59b50c504fee0783d","tarball":"https://registry.npmjs.org/@appliqation/autopilot/-/autopilot-0.1.5.tgz","fileCount":19,"integrity":"sha512-u8+l2tmOmIud9lGPOE8T3KW8dH/9UqAK0iEeyKq0giw6X4gkF1sD7dzpwRL2FQU8msnnpHco+B8EKTvQvItdyA==","signatures":[{"sig":"MEUCIDAJ6x4ttTXsdYbaQZUjK64A1cwT7WneXDs+AQ4aWsRKAiEA8vffLextymPp2ROnMdDJvhRGh7SDLUFsJ6FrHq0RqhE=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"},{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIH5YS9tKENQtfrtlrgVQ5qcjt0MBnjsF13r2YMqyDuGSAiEAi4Y1/NGbUu+mrermJXXmWMd0SDsIkBr56nNd9tHskWA="}],"unpackedSize":129418},"name":"@appliqation/autopilot","type":"module","engines":{"node":">=20"},"gitHead":"4d6038a1e44de2ac334baf7d61fde6dff4d864a5","license":"MIT","scripts":{"dev":"tsx src/cli/index.ts","lint":"eslint src --ext .ts","test":"vitest run","build":"tsc -p tsconfig.build.json","typecheck":"tsc -p tsconfig.json --noEmit","test:watch":"vitest"},"version":"0.1.5","_npmUser":{"name":"archana6","email":"accounts@appliqation.io"},"homepage":"https://github.com/appliqation/autopilot#readme","repository":{"url":"git+https://github.com/appliqation/autopilot.git","type":"git"},"_npmVersion":"11.7.0","description":"An agentic orchestrator that decides, per test case (or an entire scenario/test set), whether to run autonomous testing, heal a stale selector, fix a defect, generate new automation, or raise a pull request for it — reasoning over real signal (current pas","directories":{},"maintainers":[{"name":"archana6","email":"accounts@appliqation.io"}],"_nodeVersion":"23.10.0","dependencies":{"dotenv":"^16.4.7","commander":"^13.1.0","@appliqation/agent-core":"^0.1.7"},"_hasShrinkwrap":false,"devDependencies":{"tsx":"^4.19.3","eslint":"^9.39.5","vitest":"^3.2.7","@eslint/js":"^9.39.5","typescript":"^5.8.2","@types/node":"^22.13.10","typescript-eslint":"^8.67.0"},"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/autopilot_0.1.5_1788493187819_0.07116402862053706"}}},"time":{"created":"2026-08-21T07:52:34.919Z","modified":"2026-09-04T03:39:48.232Z","0.1.1":"2026-08-21T07:52:35.263Z","0.1.2":"2026-08-22T07:06:50.308Z","0.1.3":"2026-08-31T08:26:01.790Z","0.1.4":"2026-09-03T11:09:07.671Z","0.1.5":"2026-09-04T03:39:47.932Z"},"bugs":{"url":"https://github.com/appliqation/autopilot/issues"},"license":"MIT","homepage":"https://github.com/appliqation/autopilot#readme","repository":{"url":"git+https://github.com/appliqation/autopilot.git","type":"git"},"description":"An agentic orchestrator that decides, per test case (or an entire scenario/test set), whether to run autonomous testing, heal a stale selector, fix a defect, generate new automation, or raise a pull request for it — reasoning over real signal (current pas","maintainers":[{"name":"archana6","email":"accounts@appliqation.io"}],"readme":"# Appliqation Autopilot\n\n**An agentic orchestrator that decides, for real, every time, whether a test case needs\nautonomous testing, selector healing, a defect fix, new automation, a visual regression\ncheck, or a pull request, instead of running a fixed script.**\n\nMost \"AI test automation\" tooling is a pipeline with an LLM bolted onto one step. Autopilot\nis different: given one test case, or an entire scenario or test set, it gathers real\nsignal (current pass/fail state, flakiness, linked defects, coverage priority, whether\nautomation already exists), reasons about what's actually warranted the way a senior QA\nengineer scoping their day would, states a plan, and executes it, checking real results\nafter every action and adapting when reality disagrees with the plan, rather than committing\nblindly upfront. At scenario/test-set scope, it gathers that context once and reasons with\nrelative priority across every test case in scope, rather than re-paying full context cost\nper test case or judging everything indiscriminately.\n\nIt orchestrates seven independent, single-purpose agents as ordinary tools it can call. None\nof them know Autopilot exists. Each does one thing well and can be scripted directly if you\nwant deterministic control instead. Autopilot is the layer above them that decides *when* and\n*whether* to use each one.\n\n## Introducing the agents\n\n| Agent | Does | Repo |\n|---|---|---|\n| **Autotest** | Runs a test case in a real browser; a second, independent AI judges the result from evidence alone, never its own claim. Also does its own scenario/test-set-scale sweep: deterministic canonical-script pass first, agentic judging only for what's uncovered or newly failed. | [`appliqation-autotest`](https://github.com/appliqation/autotest) |\n| **Heal-Selector** | Repairs one broken Playwright selector in an existing script, never a full regenerate. Verifies the healed selector targets the *same semantic element* (accessibility role + name, the test case's own expected-result text) before touching anything; declines rather than picking a selector that merely makes the assertion pass again. | [`appliqation-heal-selector`](https://github.com/appliqation/heal-selector) |\n| **Visual-Regression** | Diffs one route against its own live production counterpart, no stored baseline files. Declines (not applicable) rather than comparing when the route doesn't exist on production; the mechanical work is entirely code-owned, the model only judges the real diff it's shown. | [`appliqation-visual-regression`](https://github.com/appliqation/visual-regression) |\n| **Defect-Fix** | Loads full defect context, locates and applies a real code fix, syncs the scenario, verifies with a real Playwright run. | [`appliqation-defect-fix`](https://github.com/appliqation/defect-fix) |\n| **Scriptgen** | Drafts a Playwright script for an untested-but-passing test case, iterating against real runs until genuinely green. | [`appliqation-scriptgen`](https://github.com/appliqation/scriptgen) |\n| **PR-Raise** | Fully mechanical, no LLM: commits whatever's already changed, pushes, opens or reuses a pull request. | [`appliqation-pr-raise`](https://github.com/appliqation/pr-raise) |\n| **Explorer** | Open-ended exploratory QA: a senior-QA heuristics pass plus security/network/caching/mobile probes, headlessly, the coverage a scripted test case never checks for. | [`appliqation-explorer`](https://github.com/appliqation/explorer) |\n\nAll seven share [`@appliqation/agent-core`](https://github.com/appliqation/agent-core), the generic think→act→observe engine, budget tracking, and tool-dispatch machinery underneath each of them.\n\n## Why this is different from \"wire an LLM to some CLIs\"\n\n- **The reasoning is real, not decorative.** The system prompt (`src/policy/systemPrompt.ts`)\n  encodes an actual decision framework, not \"call these tools in order,\" but genuine\n  branches: a currently-*failing* test case does **not** get a script generated for it (that\n  would encode broken behaviour as a false baseline); a flaky one gets a script but with\n  explicitly lowered confidence in the report; a low-priority one may legitimately get *no*\n  action at all, just a recommendation. Taking no action is a valid outcome, not a failure to\n  route.\n- **It never fabricates an outcome.** Every claim in the final report has to trace back to a\n  real tool result. `run_generate`'s `testRun.ok` and `run_defect_fix`'s `verified` reflect an\n  actually-executed Playwright run, not the model's own claim about it. `run_judge`'s status is\n  polled from Appliqation's own authoritative run record, not parsed out of report prose.\n  `run_visual_check`'s verdict comes from a real, code-computed pixel diff the model was shown,\n  never its own read of a screenshot. If Autopilot says a test passes, it's because it watched\n  that happen.\n- **The reasoning lives here, in the open, not behind a private API.** The seven sibling\n  agents are genuinely self-contained; the *judgment* about how to combine them is this\n  repo's own code, fully readable, forkable, and swappable (see\n  [Customizing the policy](#customizing-the-policy)). Nothing about how this agent thinks is\n  hidden behind a server only Appliqation can change. One concrete example:\n  `appliqation-defect-fix` has no way to know, on its own, how much testing a given fix\n  actually needs verified. That call is genuinely Autopilot's to make, from the broader\n  signal (defect history, run context) it's already gathering, and it's required to state\n  that reasoning explicitly rather than pass the sibling agent a generic instruction. Same\n  reasoning for `appliqation-heal-selector` and `appliqation-visual-regression`: each only\n  ever sees one script, or one route, in isolation, with no way to know whether this TC is\n  a stale-selector case, a real regression, or worth checking visually at all. That's\n  Autopilot's call, made before either sibling is even considered.\n- **Raising a pull request is opt-in, not assumed.** `run_pr_raise` isn't even in the tool\n  list Autopilot's model sees unless you pass `--allow-pr`. Same for `run_visual_check`:\n  it needs `--visual` plus a `--baseline-environment` name before it's offered at all.\n  Without either, Autopilot still reasons about whether the action would be warranted; it\n  just tells you so instead of doing it.\n\n## The complete agentic system\n\n```mermaid\nflowchart TB\n    subgraph Autopilot[\"appliqation-autopilot\"]\n        direction TB\n        Loop[\"think → act → observe loop\"]\n        Policy[\"policy (system prompt) =<br/>the actual decision framework\"]\n        Policy -.drives.-> Loop\n    end\n\n    Ctx[\"read-only Appliqation context:<br/>get_scenario, get_failure_patterns,<br/>get_defect_context, get_coverage_analysis,<br/>get_automation_readiness, enrich_project_context<br/>(read-only, action=write is refused), ...\"]\n\n    Ctx --> Loop\n    Loop -->|\"run_judge<br/>(one TC, or one scope-level call)\"| Autotest[appliqation-autotest]\n    Loop -->|\"run_heal<br/>(failed TC + canonical script exists)\"| Heal[appliqation-heal-selector]\n    Loop -->|\"run_visual_check<br/>(TC tagged visual, only if --visual)\"| Visual[appliqation-visual-regression]\n    Loop -->|run_defect_fix| DefectFix[appliqation-defect-fix]\n    Loop -->|run_generate| Scriptgen[appliqation-scriptgen]\n    Loop -->|\"run_pr_raise<br/>(only if --allow-pr)\"| PrRaise[appliqation-pr-raise]\n    Loop -->|\"run_explore<br/>(when Phase 2 states a real reason)\"| Explorer[appliqation-explorer]\n\n    Autotest -->|real polled verdict, per TC or per scope| Loop\n    Heal -->|\"verified: true/false, or declined: true\"| Loop\n    Visual -->|\"verdict + real diffPercentage\"| Loop\n    DefectFix -->|\"verified: true/false\"| Loop\n    Scriptgen -->|\"testRun.ok: true/false\"| Loop\n    PrRaise -->|PR URL or committed: false| Loop\n    Explorer -->|\"findings report + budgetExceeded\"| Loop\n```\n\n- **Context tools** are ordinary read-only Appliqation MCP tools: the same signal a human QA lead would look at before deciding where to spend effort. This includes the project's own living context document (`enrich_project_context`, action=read): known issues, high-risk areas, regression watchlist, pain points, personas, so a TC in a known-risky area gets weighed differently than the same raw evidence somewhere unremarkable. The tool also has a write mode; autopilot can only ever read it, see [Safety](#safety).\n- **Action tools** each spawn the corresponding sibling agent's CLI as a real subprocess with `--json`, and hand the model back the exact, real structured result, never a summary of one.\n- **The policy** (`src/policy/systemPrompt.ts`) is the one piece of genuine \"brain\": a detailed, phase-based methodology for gathering context, forming a plan, executing it adaptively, and reporting honestly. It's a plain string. Read it, fork it, replace it.\n\n## Workflow options\n\nAutopilot's judgment is one way to use this family, not the only one. Here are six real shapes, from fully autonomous to fully scripted.\n\n### 1. Full autonomous mode\n\nPoint Autopilot at a test case and let it decide everything: gather context, judge current state, fix or generate as warranted, raise the PR.\n\n```mermaid\nsequenceDiagram\n    participant You\n    participant AP as Autopilot\n    participant AT as Autotest\n    participant DF as Defect-Fix\n    participant SG as Scriptgen\n    participant PR as PR-Raise\n\n    You->>AP: run --test-case-uuid X --allow-pr\n    AP->>AP: gather context, form a plan\n    alt no recent evidence\n        AP->>AT: run_judge\n        AT-->>AP: real verdict\n    end\n    alt currently fails, fixable defect exists\n        AP->>DF: run_defect_fix (+ test_instruction)\n        DF-->>AP: verified: true/false\n    else currently passes, no canonical script\n        AP->>SG: run_generate\n        SG-->>AP: testRun.ok: true/false\n    end\n    alt something was verified\n        AP->>PR: run_pr_raise\n        PR-->>AP: PR URL\n    end\n    AP-->>You: full report, every claim tool-backed\n```\n\n```bash\nnpx appliqation-autopilot run --test-case-uuid <uuid> --environment Stage --repo-path <path> --allow-pr\n```\n\n### 2. Deterministic CI pipeline (no orchestrator)\n\nSkip Autopilot entirely and script the individual agents directly: full control, zero LLM judgment about *what* to run, still LLM-verified per step.\n\n```mermaid\nflowchart LR\n    A[CI trigger] --> B[\"appliqation-autotest judge<br/>--test-set-id regression\"]\n    B --> C{all passed?}\n    C -- yes --> D[deploy]\n    C -- no --> E[fail the build,<br/>file/notify]\n```\n\n```bash\nnpx appliqation-autotest judge --test-set-id <id> --environment Stage --ci\n```\n\n### 3. Defect triage & auto-fix\n\nA regression run surfaces failures and defects get filed. Route each one through Defect-Fix (directly, or via Autopilot for scope judgment), open a PR, let a human review.\n\n```mermaid\nflowchart LR\n    A[\"CI run finds failures<br/>(defects filed in Appliqation)\"] --> B[\"appliqation-defect-fix fix<br/>--defect-id <id>\"]\n    B --> C{verified?}\n    C -- yes --> D[\"appliqation-pr-raise raise\"]\n    C -- no --> E[report what's still broken,<br/>no PR]\n    D --> F[human reviews the PR]\n```\n\n```bash\nnpx appliqation-defect-fix fix --defect-id <id> --repo-path <path> --dry-run   # first pass, safe\nnpx appliqation-defect-fix fix --defect-id <id> --repo-path <path>            # once trusted\n```\n\n### 4. Coverage backfill\n\nSystematically pay down test-automation debt: find passing-but-unautomated test cases, generate and verify canonical scripts for each, batch them into a PR.\n\n```mermaid\nflowchart LR\n    A[\"get_automation_readiness:<br/>passing TCs with no canonical script\"] --> B[\"appliqation-scriptgen generate<br/>per test case\"]\n    B --> C{testRun.ok?}\n    C -- yes --> D[accumulate written files]\n    C -- no --> E[skip, flag for follow-up]\n    D --> F[\"appliqation-pr-raise raise<br/>(one PR, batched)\"]\n```\n\n```bash\nnpx appliqation-scriptgen generate --test-case-uuid <uuid> --repo-path <path> --ci\n```\n\n### 5. Scenario/test-set triage\n\nPoint Autopilot at an entire scenario or test set instead of one test case. It leads with a\nsingle scope-level `run_judge` call to get autotest's own consolidated pass/fail/uncovered\nsignal cheaply, then reasons with relative priority across every TC in scope: a failed TC\nwith a canonical script is a healing candidate before it's a defect-fix candidate.\n\n```mermaid\nsequenceDiagram\n    participant You\n    participant AP as Autopilot\n    participant AT as Autotest\n    participant HL as Heal-Selector\n    participant DF as Defect-Fix\n\n    You->>AP: run --scenario-id 2424 --allow-pr\n    AP->>AP: enumerate TCs via get_scenario, gather context once\n    AP->>AT: run_judge (one scope-level call, on-failure-or-absence coverage)\n    AT-->>AP: consolidated per-TC results (pass / fail / uncovered)\n    loop for each failed TC with a canonical script\n        AP->>HL: run_heal\n        alt healed and verified\n            HL-->>AP: verified: true\n        else declined (not a stale-selector case)\n            HL-->>AP: declined: true\n            AP->>DF: run_defect_fix (+ test_instruction)\n        end\n    end\n    AP-->>You: per-TC report + prioritized recommendation\n```\n\n```bash\nnpx appliqation-autopilot run --scenario-id <id> --environment Stage --repo-path <path>\n# or: --test-set-id <id>\n```\n\n### 6. Visual regression check\n\nAuthorize `run_visual_check` for a run, and Autopilot will call it for any TC tagged\n`visual` (or with an equally specific stated reason), citing the real route from a run it\nalready has evidence for.\n\n```mermaid\nsequenceDiagram\n    participant You\n    participant AP as Autopilot\n    participant AT as Autotest\n    participant VR as Visual-Regression\n\n    You->>AP: run --test-case-uuid X --visual --baseline-environment Prod\n    AP->>AP: gather context, see Tag: visual on this TC\n    AP->>AT: run_judge\n    AT-->>AP: real verdict + a run_id with real evidence\n    AP->>AP: get_execution_evidence on that run_id for the real route\n    AP->>VR: run_visual_check (route, baseline Prod, target Stage)\n    VR-->>AP: verdict: regression / expected-divergence / not-applicable / inconclusive\n    AP-->>You: report citing the real verdict and diffPercentage\n```\n\n```bash\nnpx appliqation-autopilot run --test-case-uuid <uuid> --environment Stage --repo-path <path> \\\n  --visual --baseline-environment Prod\n```\n\n## Quick start\n\n```bash\nnpm install -g @appliqation/autopilot\n```\n\nYou'll also need whichever sibling agents autopilot is allowed to call. Install the ones\nyou want reachable (see [Workflow options](#workflow-options) above for real combinations;\nyou don't need all seven for every use case):\n\n```bash\nnpm install -g @appliqation/autotest @appliqation/heal-selector @appliqation/visual-regression \\\n  @appliqation/defect-fix @appliqation/scriptgen @appliqation/pr-raise @appliqation/explorer\n```\n\nEach is a plain command name by default (`AUTOTEST_CMD`/`HEAL_CMD`/`VISUAL_CMD`/\n`DEFECT_FIX_CMD`/`SCRIPTGEN_CMD`/`PR_RAISE_CMD`/`EXPLORER_CMD` in `.env`). Only override\nthese if you're pointing at a local development build instead.\n\nCreate a `.env` file (in whatever directory you'll run it from) with:\n\n```\nAPPQ_API_KEY=your-appliqation-api-key\nANTHROPIC_API_KEY=your-anthropic-key   # or OPENAI_API_KEY, pick one\n```\n\n```bash\nappliqation-autopilot run \\\n  --test-case-uuid <uuid> \\\n  --environment Stage \\\n  --repo-path /path/to/your/checkout\n```\n\n`--test-case-uuid` is one test case. Swap it for `--scenario-id <id>` to route an entire\nscenario, or `--test-set-id <id>` to route an entire test set (can span multiple scenarios —\nthe common regression/sanity/smoke shape) — exactly one of the three is required, they're\nmutually exclusive scopes. See [Workflow option 5](#5-scenariotest-set-triage) for the\nscenario/test-set shape in context.\n\nWatch stderr: every context tool call, every action, and the model's own reasoning\n(`[thinking]` lines) stream live. Add `--allow-pr` once you're ready to let it actually open\npull requests, or `--visual --baseline-environment <name>` to let it check for visual\nregressions; add `--json`/`--ci` for a single structured result and a CI-friendly exit code\ninstead of the human-readable transcript.\n\n## CLI reference\n\n`appliqation-autopilot run [options]`\n\n**Scope — exactly one required:**\n\n| Option | Description |\n|---|---|\n| `--test-case-uuid <uuid>` | One test case to route. |\n| `--scenario-id <id>` | An entire scenario to route: richer context, and Phase 1 gets paid for once instead of once per TC. |\n| `--test-set-id <id>` | An entire test set to route (can span multiple scenarios, the common regression/sanity/smoke shape). |\n\n**Required:**\n\n| Option | Description |\n|---|---|\n| `--environment <name>` | Environment name, passed to `run_judge`/`run_generate`. |\n| `--repo-path <path>` | Local repo checkout `run_generate`/`run_pr_raise`/`run_heal` operate in. |\n\n**Optional:**\n\n| Option | Description |\n|---|---|\n| `--defect-id <id>` | The specific defect that triggered this run, when the caller already resolved one. Without it, Phase 1 only discovers a linked defect incidentally; passing it makes the defect/TC mismatch check actually check against the real triggering defect. |\n| `--allow-pr` | Authorize `run_pr_raise` — without this flag, that tool is not even offered to the model. |\n| `--visual` | Authorize `run_visual_check` — without this flag, that tool is not even offered to the model. Requires `--baseline-environment`. |\n| `--baseline-environment <name>` | Production/baseline environment name for `run_visual_check`, required together with `--visual`. |\n| `--policy <path>` | Override the bundled decision policy with your own system prompt file. |\n| `--max-turns <n>` | Override `BUDGET_MAX_TURNS` for this run. |\n| `--json` | Print a single structured JSON result instead of the human-readable report. |\n| `--ci` | Shorthand for `--json`. |\n\n## Customizing the policy\n\nThe default policy in `src/policy/systemPrompt.ts` is opinionated: don't automate a currently\nfailing test, flag flaky results as lower-confidence, treat \"no action needed\" as a legitimate\noutcome. Your organization might weigh these differently: more conservative, a different\npriority order, a different report format for your own stakeholders.\n\nYou don't need to touch anything else in this repo to change that:\n\n```bash\nnpx appliqation-autopilot run --policy ./my-policy.md ...\n# or set POLICY_FILE in .env\n```\n\nAnything you put in that file becomes the system prompt driving every decision. The\norchestration code (`src/orchestrator/`, `src/tools/`) doesn't change; it just runs whatever\npolicy it's given against the same tools.\n\nIf what you actually want is full deterministic control with no LLM judgment in the loop at\nall, you don't need Autopilot for that. Script `appliqation-autotest`,\n`appliqation-heal-selector`, `appliqation-visual-regression`, `appliqation-defect-fix`,\n`appliqation-scriptgen`, and `appliqation-pr-raise` directly; each is a complete,\nindependently useful CLI (see\n[workflow 2](#2-deterministic-ci-pipeline-no-orchestrator)).\n\n## Safety\n\n- `run_pr_raise` and `run_visual_check` are excluded from the tool list entirely unless\n  `--allow-pr`/`--visual` is passed: a hardcoded exclusion, not a soft warning the model\n  could talk itself past.\n- Every meta-tool result is the sibling agent's own real `--json` output, including on\n  failure (a failed/blocked `run_judge`, an unverified `run_generate`/`run_defect_fix`/\n  `run_heal`, an inconclusive `run_visual_check`), so a bad outcome is visible to the model\n  as data to reason about, never swallowed. A `run_heal` decline, and a `run_visual_check`\n  `not-applicable`/`expected-divergence` result, are all treated the same way: real\n  evidence, not a failure to hide.\n- The individual agents carry their own safety invariants independently: a destructive-action\n  gate on any browser interaction, an allowlisted shell surface for `appliqation-scriptgen`'s\n  and `appliqation-defect-fix`'s environment bootstrap, `appliqation-defect-fix`'s own appq\n  writes gated behind its own `--dry-run` (which Autopilot can pass through via\n  `run_defect_fix`'s `dry_run` argument), no credentials ever flowing through an LLM's own\n  context. `appliqation-visual-regression` has no filesystem/shell surface at all. Autopilot\n  doesn't weaken any of that; it just decides when to invoke it.\n- No credentials of any kind pass through the LLM's context at any point in this repo.\n- `enrich_project_context` is a single Appliqation tool with both `action=read` and\n  `action=write` modes. Tool-*name* allowlisting alone can't express \"this tool, but\n  only this argument value,\" so `@appliqation/agent-core`'s `createReadOnlyProjectContextDispatcher`\n  adds an argument-level gate on top: only `action=read` is ever let through, and the\n  check fails closed (a missing or malformed `action` is refused too, not just an\n  explicit `\"write\"`). Shared with `appliqation-explorer`, which needs the identical\n  guarantee.\n- `appliqation-explorer` is read-only end to end, including on the project context\n  document its own upstream workflow (`appq:runman`) would otherwise write to. When that\n  workflow runs interactively in Claude Code, a human is present at its confirmation gate\n  before anything gets persisted as fact for future passes; a headless `run_explore` call\n  has no equivalent, so it holds the same conservative default as Autopilot itself rather\n  than the permissive one baked into the interactive prompt. See that repo's README for\n  the full reasoning.\n\n**Run this inside a container with an egress allowlist**, same as every sibling it can spawn.\nThis process's own direct network need is narrow (your LLM provider and `APPQ_ORIGIN`), but\n`run_judge`/`run_heal`/`run_visual_check`/`run_generate`/`run_defect_fix`/`run_explore`/\n`run_pr_raise` each spawn a sibling agent as a real subprocess, and each of those has its own\nbroader surface: a live browser, a real shell, or a real `GITHUB_TOKEN`, documented in that\nsibling's own README under \"Running this safely.\" Containing this process alone isn't\nsufficient; contain the whole tree it can spawn.\n\n## Configuration\n\nSee `.env.example` for the full list. In short: `APPQ_API_KEY` + one LLM provider key are\nrequired; `AUTOTEST_CMD`/`HEAL_CMD`/`VISUAL_CMD`/`DEFECT_FIX_CMD`/`SCRIPTGEN_CMD`/\n`PR_RAISE_CMD`/`EXPLORER_CMD` tell Autopilot how to reach the sibling agents; `BUDGET_MAX_*`\ncaps the tool-calling loop; `POLICY_FILE` points at a custom policy.\n\nOptionally, `AUDIT_MONGO_URI`/`AUDIT_MONGO_DB`/`AUDIT_MONGO_COLLECTION` or\n`AUDIT_JSONL_PATH` records one audit entry per invocation (token usage, duration, real\noutcome) to a datastore this agent family owns, deliberately not part of Appliqation\nitself, since this is a parallel system that uses it, not a feature of it. Every sibling\nagent in the family writes to the same shape; [`appliqation-dashboard`](https://github.com/appliqation/dashboard)\nreads it back as an aggregated report. Entirely opt-in: nothing is recorded unless one\nof these is set, and a write failure never affects a real run's outcome.\n\n## Development\n\n```bash\ngit clone https://github.com/appliqation/autopilot.git\ncd autopilot\nnpm install\ncp .env.example .env   # fill in APPQ_API_KEY and one LLM provider key\nnpm run dev -- run --test-case-uuid <uuid> --environment <name> --repo-path <path>\nnpm run typecheck\nnpm test\n```\n\nSee `CLAUDE.md` for a map of this repo if you're working in it with an AI coding assistant.\n\n## License\n\nMIT. See [LICENSE](./LICENSE).\n","readmeFilename":"README.md"}