{"_id":"@buyhatke-dev/test-ez","_rev":"4-964cf37e70749fd51981b9cd396088de","name":"@buyhatke-dev/test-ez","dist-tags":{"latest":"0.4.0"},"versions":{"0.3.2":{"name":"@buyhatke-dev/test-ez","version":"0.3.2","keywords":["mcp","playwright","e2e","unit","testing","triage","flaky","ci","model-context-protocol"],"license":"MIT","_id":"@buyhatke-dev/test-ez@0.3.2","maintainers":[{"name":"buyhatke-dev","email":"keshavk@buyhatke.com"}],"homepage":"https://github.com/Buyhatke/test-ez#readme","bugs":{"url":"https://github.com/Buyhatke/test-ez/issues"},"bin":{"test-ez":"bin/cli.mjs"},"dist":{"shasum":"b08bf0c7d49a1a5147201119ed159c6fa21671e9","tarball":"https://registry.npmjs.org/@buyhatke-dev/test-ez/-/test-ez-0.3.2.tgz","fileCount":18,"integrity":"sha512-HxXw9rjiZCk6yToQCn/hEEYP472H9Xz4gyCOreoGXAZEVAQKIZQLrqWpuAnqqDnxctLm7X25Eiq1aZWsFNWo2A==","signatures":[{"sig":"MEUCIQDx8fxLoRY6lJ2w1VWDR2ua8bty914nxy85cStaL+1rugIgLz7fJVoiy/JkvsWAoZ2+QLAtzBmzm6KfKTrdwpMPOQU=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":226994},"main":"index.mjs","type":"module","engines":{"node":">=18"},"exports":{".":"./index.mjs"},"gitHead":"ef52706452676a0a41ce8bb2bc58061ffd6e019f","scripts":{"init":"node bin/cli.mjs init","serve":"node index.mjs","smoke":"node bin/smoke.mjs"},"_npmUser":{"name":"buyhatke-dev","email":"keshavk@buyhatke.com"},"repository":{"url":"git+https://github.com/Buyhatke/test-ez.git","type":"git"},"_npmVersion":"11.12.1","description":"Portable, model-agnostic MCP server that runs your unit and e2e tests, triages why each one failed (frontend / backend / real bug / flaky), and opens grouped fix PRs on GitHub.","directories":{},"_nodeVersion":"24.15.0","dependencies":{"zod":"^3.23.8","adm-zip":"^0.5.16","@modelcontextprotocol/sdk":"^1.29.0"},"publishConfig":{"access":"public"},"_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/test-ez_0.3.2_1782801761384_0.6317673166020048","host":"s3://npm-registry-packages-npm-production"}},"0.4.0":{"name":"@buyhatke-dev/test-ez","version":"0.4.0","description":"Portable, model-agnostic MCP server that runs your unit and e2e tests, triages why each one failed (frontend / backend / real bug / flaky), and opens grouped fix PRs on GitHub.","type":"module","license":"MIT","repository":{"type":"git","url":"git+https://github.com/Buyhatke/test-ez.git"},"bugs":{"url":"https://github.com/Buyhatke/test-ez/issues"},"homepage":"https://github.com/Buyhatke/test-ez#readme","publishConfig":{"access":"public"},"bin":{"test-ez":"bin/cli.mjs"},"main":"index.mjs","exports":{".":"./index.mjs"},"scripts":{"serve":"node index.mjs","init":"node bin/cli.mjs init","smoke":"node bin/smoke.mjs"},"keywords":["mcp","playwright","e2e","unit","testing","triage","flaky","ci","model-context-protocol"],"engines":{"node":">=18"},"dependencies":{"@modelcontextprotocol/sdk":"^1.29.0","adm-zip":"^0.5.16","zod":"^3.23.8"},"gitHead":"d1260cfe2238e540bf1b72c49a260f0849df544b","_id":"@buyhatke-dev/test-ez@0.4.0","_nodeVersion":"24.15.0","_npmVersion":"11.12.1","dist":{"integrity":"sha512-cA+h+7dBqxNmn/SLQLO1uuKxwUnQ64H61YNuXexkKdm/Tsw8hWbRG/LCFmWuD0LHEgZocAE67PiOs8B4n1tpjg==","shasum":"428a1570fc5c35537bca9466fa5201f6d5216171","tarball":"https://registry.npmjs.org/@buyhatke-dev/test-ez/-/test-ez-0.4.0.tgz","fileCount":18,"unpackedSize":239964,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIELpZcQPL8X2trcnJhiiaXJ60SZW5OapcuwtOWp1cXZVAiEA4QC4VtUUPUoIGOs7yc0QxRFHH1TmNvVbBa/xdjBe5Do="}]},"_npmUser":{"name":"keshav-kr","email":"keshavk@buyhatke.com"},"directories":{},"maintainers":[{"name":"archie30","email":"archie@onramp.money"},{"name":"keshav-kr","email":"keshavk@buyhatke.com"},{"name":"pattahgobhi","email":"aayushmaan@onramp.money"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/test-ez_0.4.0_1787310006706_0.23454756703145452"},"_hasShrinkwrap":false}},"time":{"created":"2026-06-30T06:42:41.172Z","modified":"2026-08-21T11:00:07.062Z","0.3.2":"2026-06-30T06:42:41.522Z","0.4.0":"2026-08-21T11:00:06.861Z"},"bugs":{"url":"https://github.com/Buyhatke/test-ez/issues"},"license":"MIT","homepage":"https://github.com/Buyhatke/test-ez#readme","keywords":["mcp","playwright","e2e","unit","testing","triage","flaky","ci","model-context-protocol"],"repository":{"type":"git","url":"git+https://github.com/Buyhatke/test-ez.git"},"description":"Portable, model-agnostic MCP server that runs your unit and e2e tests, triages why each one failed (frontend / backend / real bug / flaky), and opens grouped fix PRs on GitHub.","maintainers":[{"name":"archie30","email":"archie@onramp.money"},{"name":"keshav-kr","email":"keshavk@buyhatke.com"},{"name":"pattahgobhi","email":"aayushmaan@onramp.money"}],"readme":"# test-ez\n\n> Failing tests, made **EZ** — a portable, model-agnostic **MCP** that **runs** your Playwright e2e **and** Vitest unit suites, **triages why each one failed** (backend/API break · frontend change · real bug · flaky · snapshot drift · test-authoring · environment), and — on your approval — **opens grouped fix PRs** on GitHub.\n\nIt's **data + deterministic logic only**: it runs tests, parses reports, recovers network/console from traces, correlates with your git diff, and rule-classifies each failure with evidence + a suggested fix. Your AI client (Claude Code, Cursor, Windsurf, …) drives the conversation and applies the code edits. Nothing is project-specific — drop it into any Playwright + GitHub repo.\n\n---\n\n## Install\n\ntest-ez is an MCP server. You register it **once per AI host** (Codex, Claude Code, Cursor, Windsurf, …); project config is saved separately, later. Restart the host after registering so it loads the server.\n\n> **Want your agent to do it?** Skip to [Let your agent install it](#let-your-agent-install-it) — paste one block and it detects your host and registers everything.\n\n### Recommended — one command per host (from npm)\n\ntest-ez is published on npm as a **public** package, so it installs with `npx` — no clone, no GitHub auth, no `.npmrc`.\n\n**Codex**\n```sh\ncodex mcp add test-ez -- npx -y @buyhatke-dev/test-ez serve\n```\nmacOS, if `codex` isn't on your `PATH`: replace `codex` with `/Applications/Codex.app/Contents/Resources/codex`.\n\n**Claude Code**\n```sh\nclaude mcp add test-ez --scope user -- npx -y @buyhatke-dev/test-ez serve\n```\n\n**Cursor / Windsurf / Claude Desktop / any MCP-compatible host** — add this to your MCP config (`.mcp.json`, `.cursor/mcp.json`, or the host's config file):\n```json\n{\n  \"mcpServers\": {\n    \"test-ez\": {\n      \"command\": \"npx\",\n      \"args\": [\"-y\", \"@buyhatke-dev/test-ez\", \"serve\"]\n    }\n  }\n}\n```\n\n> `TEST_EZ_PROJECT_ROOT` is **optional** — set it (a `\"env\"` block in the JSON, or `-e TEST_EZ_PROJECT_ROOT=/abs/path` on the `mcp add` command) only if your host launches from a directory other than your project root; otherwise the server uses the launch cwd / git toplevel.\n\n### Let your agent install it\n\nPaste this whole block into your AI agent. It detects the host, registers test-ez, and hands you a restart message instead of stalling:\n\n```text\nYou are setting up the test-ez MCP server for whichever AI host you are running in. Do these steps in order.\n\nSTEP 1 — Detect the host\nWork out which AI host you are running in: Codex, Claude Code, Cursor, Windsurf, or another MCP-compatible host.\n\nSTEP 2 — Register the server\ntest-ez is a PUBLIC npm package, so it installs with npx and needs no GitHub auth. Register an MCP server:\n  Name:    test-ez\n  Command: npx\n  Args:    -y @buyhatke-dev/test-ez serve\n\nIf you cannot register MCP servers yourself, give the user exactly ONE terminal command for their detected host:\n  Codex:        codex mcp add test-ez -- npx -y @buyhatke-dev/test-ez serve\n  Claude Code:  claude mcp add test-ez --scope user -- npx -y @buyhatke-dev/test-ez serve\n\nSTEP 3 — Verify\nList the host's MCP servers and confirm test-ez appears.\n\nSTEP 4 — Hand off for restart (IMPORTANT)\nRegistering an MCP server does NOT load it into this running session, and restarting the host ENDS this conversation. So do not call any test-ez tools now. Stop here and show the user this exact message (fill in the host name):\n\n  ────────────────────────────────────────\n  ✅ test-ez MCP registered.\n\n  👉 NEXT STEPS (in your project repo, after restart):\n     1. Restart <this AI host> now so it loads test-ez.\n        (The first start may take a few seconds while npx fetches it — that's normal.)\n     2. Say:  \"set up test-ez here\"        → writes .test-ez/config.json (one-time)\n     3. Say:  \"run my e2e tests using test-ez\"\n  ────────────────────────────────────────\n```\n\n### Install from source (pre-release / private GitHub repo)\n\nTo run an unreleased build straight from `Buyhatke/test-ez`, install from GitHub. The repo is **private**, so `npx` must reach it **without** an interactive prompt: authenticate to GitHub first (`gh auth login`, or an SSH key / credential helper with access to the repo). Use the **HTTPS** spec if you authenticate over HTTPS, or the **SSH** spec if you use an SSH key:\n\n**Codex**\n```sh\n# HTTPS:\ncodex mcp add test-ez -- npx -y github:Buyhatke/test-ez serve\n# SSH:\ncodex mcp add test-ez -- npx -y git+ssh://git@github.com/Buyhatke/test-ez.git serve\n```\n\n**Claude Code**\n```sh\n# HTTPS:\nclaude mcp add test-ez --scope user -- npx -y github:Buyhatke/test-ez serve\n# SSH:\nclaude mcp add test-ez --scope user -- npx -y git+ssh://git@github.com/Buyhatke/test-ez.git serve\n```\n\n> If the **HTTPS** form makes your host prompt for a GitHub username every session, you authenticate over SSH — switch to the SSH form, or route GitHub HTTPS through your SSH auth once: `git config --global url.\"git@github.com:\".insteadOf \"https://github.com/\"`.\n\n### Initialise your project (once)\n\nEasiest: in your repo, just tell your AI client **\"set up test-ez here\"** — it runs the `init` tool and writes `.test-ez/config.json` (no extra `npx` fetch).\n\nPrefer the CLI? From your project root:\n```sh\nnpx -y @buyhatke-dev/test-ez init\n# unreleased build from the private repo: npx -y github:Buyhatke/test-ez init\n```\nEither way it detects your environment (git / GitHub / `gh` / Playwright / Vitest), suggests sensible defaults, and writes `.test-ez/config.json`.\n\n### Requirements\n\n- **Node ≥ 18** and a project with **Playwright** (`@playwright/test`, browsers installed via `npx playwright install`) and/or **Vitest** for unit suites.\n- **For fix PRs:** a **GitHub** repo + the **`gh` CLI**, authenticated (`gh auth login`). Without git/GitHub, triage still works — only PR creation is disabled.\n\n### Troubleshooting install\n\n- **`ENOVERSIONS` / \"no matching versions\" on install** — your `.npmrc` has a `min-release-age` guard (a supply-chain check that blocks just-published versions), and test-ez was published recently.\n  - **Permanent, scoped fix:** run **`test-ez exception`** (or ask your client *\"run test-ez exception\"*). It asks first, then adds a `min-release-age-exclude` for **only** test-ez — your guard still protects every other package, and future updates won't be age-blocked. `test-ez doctor` shows whether the guard is active and whether test-ez is excepted.\n  - **First install (before test-ez is fetchable):** add `NPM_CONFIG_MIN_RELEASE_AGE=0` to the server's `env` block (overrides it for test-ez only), or wait until the version is older than your threshold — then run `test-ez exception` so you never hit it again.\n- **npmjs.com page 403s** for the package — the registry API still works; check it with `npm view @buyhatke-dev/test-ez`.\n- **Verify the install** anytime: `npx -y @buyhatke-dev/test-ez doctor` (or tell your client *\"run test-ez doctor\"*) — prints version, registry reachability, environment, config status, and any warnings.\n\n---\n\n## Quick start\n\nOnce it's installed, just talk to your AI client in plain language:\n\n> *\"Run my e2e tests using test-ez.\"*  (or *\"run my unit tests using test-ez\"*)\n\nIt will `run_e2e` → give you a short pass/fail summary → and (after you say go) `triage` each failure with real evidence, propose fixes, and `open_fix_pr` for the clear-cut ones. Backend / env / snapshot failures are **flagged only**, never auto-fixed.\n\n## Tools\n\n| Tool | Purpose |\n|---|---|\n| `init` | Detect environment + config, or persist config (`set`). |\n| `run_e2e` | Run a suite or a subset (`tests`, `grep`, `project`, `lastFailed`, `onlyChanged`, `repeatEach`). Forces JSON report + retain-on-failure traces into `.test-ez/` regardless of your config. `background:true` returns immediately (poll with `run_status`); every run streams to a tailable `run.log`. `command` = run an exact shell command then a controlled capture. |\n| `run_status` | Live status of the latest (or named) run — phase (pre-step/running/done/failed), elapsed, approximate test counts + current test (parsed from the live reporter), and a log tail. Self-heals a run whose process died without finalizing. On a timed-out/crashed run, surfaces the salvaged **partial** results. |\n| `list_failures` | Failures + stats + skipped, with stable `testId`s. |\n| `get_failure_detail` | Full error/stack/stdout for one test. |\n| `get_network_log` | Network from the trace; first-party API failures highlighted. |\n| `get_console_log` | Console + uncaught page (JS) errors from the trace. |\n| `get_snapshot` | \"Look closer\" at one failure: the **screenshot** at the failed step (an image to read/see the real screen — banners, wrong screen, offline state) + the **data-testids** present in the DOM then. Turns a guess into evidence. |\n| `get_changed_files` | `git diff` vs base, tagged app-source/test/config. |\n| `triage` | Gather **evidence + signals** per failure (error + call log, distilled first-party network, changed-file correlation, trace action log / typed values, cross-browser) plus a labeled heuristic **hint**. The hint is a prior, *not* a verdict — **the calling agent asserts the issue type**. JSON + markdown. |\n| `open_fix_pr` | One PR per cause off the base branch, labelled, context in the body. Pass the agent's asserted type as each group's `category`. `dryRun` to preview. |\n\n> **Classification is the agent's job, not the server's.** `triage` deliberately does *not* return a final verdict — heuristics get fooled by red herrings (a cross-origin asset block, an unrelated app edit). The MCP gathers accurate, deterministic *evidence*; the LLM reasons over it and asserts the type (`backend · frontend · flow-change · test-authoring · flaky · real-bug · infra`).\n\n## Workflow\n\n> 🗺️ Full end-to-end map (tools, modules, the e2e↔unit fork): **[docs/WORKFLOW.md](docs/WORKFLOW.md)**.\n\n1. `init` (first time).\n2. `run_e2e` — reliable JSON + trace capture is automatic.\n3. `triage` — read the **evidence**, then assert the issue type yourself (the hint is just a prior).\n4. Approve.\n5. Apply fixes for the clear-cut, fixable cases; leave ambiguous ones queued.\n6. `run_e2e { onlyChanged: true }` as a touched-area regression gate; `repeatEach: 3` to confirm flaky-suspects (must pass 3/3).\n7. `open_fix_pr` with `triage.suggestedGroups`. Backend / env / snapshot failures are **flagged only**, never auto-PR'd.\n\n## Config (`.test-ez/config.json`)\n\n```jsonc\n{\n  \"baseBranch\": \"main\",\n  \"label\": \"e2e AI fix\",\n  \"defaultSuite\": \"main\",\n  \"suites\": {\n    // COMMAND mode (preferred when you have a real test script): run it as-is,\n    // test-ez only appends capture flags. Preserves the script's scope/wiring\n    // (one browser via randomBrowser, build/deploy, synpress, …) — no fan-out.\n    \"test:e2e\": { \"command\": \"npm run test:e2e\" },\n    // BUILD mode: test-ez constructs `npx playwright test …` itself.\n    \"main\":  {},                                   // {} → auto-discover playwright.config.*, run all projects\n    \"e2e\":   { \"config\": \"playwright/config.cjs\", \"project\": \"chromium\",\n               \"preRun\": \"npm run deploy:test\", \"env\": { \"FOO\": \"bar\" } }\n  },\n  \"repeatEach\": 3,\n  \"runTimeoutMs\": 1200000,\n  \"redact\": false,                 // mask secrets in report/triage/console? off (your machine) — PR bodies redact regardless\n  \"worktree\": {\n    \"enabled\": true,               // run/fix/PR in an isolated worktree (default on; false to opt out)\n    \"path\": null,                  // default: sibling \"<repo>.test-ez-wt\"\n    \"branch\": \"test-ez-worktree\",\n    \"ref\": \"base\",                 // what to mirror: \"base\" | \"current\" | <ref>\n    \"autoNpmCi\": true,             // reinstall deps only when the lockfile changes\n    \"installCmd\": null,            // override; else derived (npm ci / yarn / pnpm)\n    \"copyDirs\": []                 // extra heavy gitignored dirs to carry in whole (adds to the built-in list; false for none)\n  }\n}\n```\n\nA **suite** is a structured run-spec. A suite with a **`command`** runs your project's own script verbatim and only appends capture flags (`--reporter=json`, `--trace`, `--output`) — so its scope and wiring (e.g. `--project=\"$(randomBrowser.mjs)\"`, build/deploy, synpress) are preserved and there's no project fan-out. Otherwise `preRun` runs before tests (e.g. a deploy) and aborts the run if it fails; test-ez builds the `playwright test` invocation from the suite + selection params.\n\n## Long suites — never a black box, never a total loss\n\n- **Watch it:** every run streams to `.test-ez/work/runs/<suite>/run.log` (`tail -f` it). For long suites pass `run_e2e { background: true }` to return immediately, then poll `run_status` for phase + live counts + current test.\n- **Timeout ≠ wasted run:** if a run is killed before Playwright writes its JSON report (timeout/crash), test-ez **salvages** what ran from the live log + the on-disk traces — `producedReport:false` but `partial` carries the pass/fail/flaky counts, the failing tests (matched to their traces), and a `rerunSubset` of `file:line`s to re-run just the failures for a clean report. Raise `timeoutMs`, or shard via `grep`/`tests`, for a complete run.\n\n## Worktree mode (parallel-safe)\n\nWorktree mode is **on by default** (disable with `worktree.enabled: false`, or per run with `run_e2e { worktree: false }`). Runs, fixes, and PRs happen in a **persistent sibling git worktree** instead of your main checkout — so you can keep working while test-ez operates. Each run:\n\n1. syncs the worktree to the target ref (`fetch` → `reset --hard <ref>` → `clean`, keeping `node_modules`),\n2. **carries gitignored env/secret files** from your main checkout into the worktree — top-level `.env*` (minus committed `.example`/`.sample` templates) by default, plus anything you list in `worktree.copyFiles` (e.g. a nested `.env` or a `secrets.json`). A fresh `git worktree` checkout otherwise lacks these, breaking deploy scripts / wallet setup / anything that reads them,\n3. **carries heavy wallet/auth caches** the fresh checkout can't cheaply rebuild (see below),\n4. reinstalls deps **only when the lockfile hash changed** (so the common case starts instantly), and\n5. runs the suite there; reports/traces still land under the main repo's `.test-ez/work/`.\n\n### Wallet / auth caches (`worktree.copyDirs`)\n\nStep 2 deliberately skips heavy directories — `node_modules`, `dist`, caches, browser profiles — because they're rebuilt anyway. But some gitignored directories are *not* cheaply rebuildable: a Synpress `.cache-synpress` holds pre-onboarded MetaMask/Phantom browser profiles plus the downloaded extensions, and recreating it needs the seed phrase, a headful Chrome and a ~50MB download. Without it the run's own setup step tries to rebuild it inside the worktree on every run — slow at best, and when it fails the run dies before a single test executes, so **an environment problem shows up as a failing test suite**.\n\ntest-ez carries these whole. `.cache-synpress` is on the built-in list; add your own with `worktree.copyDirs: [\"my-tool/.cache-state\", …]`, or set `\"copyDirs\": false` to carry none. Entries **add to** the built-in list rather than replacing it.\n\nPlaywright's `playwright/.auth/*.json` storage state needs no entry — it's small enough that step 2 carries it like any other gitignored file.\n\nThe copy is **copy-on-write cloned** where the filesystem supports it (APFS on macOS, btrfs/xfs on Linux), so a 300MB cache lands in well under a second and costs no extra disk — while the worktree still gets its own copy, so a run that rewrites the cache can't corrupt your main checkout's. It's refreshed every run, since your main checkout is the source of truth.\n\n**What the worktree mirrors** is configurable via `worktree.ref` (or per run, `run_e2e { ref }`):\n\n| `ref` | Mirrors | Use |\n|---|---|---|\n| `\"base\"` *(default)* | `origin/<baseBranch>` | \"is the important branch healthy\" (CI-mirror) |\n| `\"current\"` | your checked-out branch HEAD (local commit, incl. unpushed) | **\"fix the tests on my branch\"** |\n| any literal ref | that branch / tag / SHA | fix tests as they'll run on `release-x`, etc. |\n\nThe run result echoes the **resolved ref + SHA + commit subject**, so it's never ambiguous what was tested; an unknown ref is rejected rather than silently falling back. `triage` and `open_fix_pr` follow the worktree the last run used — the fix PR targets the **branch that was mirrored** (so a `current` run PRs *into* your branch, not master; override with `open_fix_pr { baseBranch }`). Your main tree is never stashed or branch-switched. A run-lock (`.test-ez/work/run.lock`) keeps two runs from trampling the shared worktree. Apply fixes in the worktree path that `run_e2e` reports.\n\n> Note: `ref` carries **committed** work only. To test live uncommitted edits, either commit first or run with worktree off (in-place sees everything).\n\n## Safety\n\n- **Secret redaction** is **off by default** for owner-facing output (report / triage / console) — it's your machine and your secrets, and masking them just hampers debugging; flip `redact: true` to mask. **PR bodies are redacted regardless** (auth headers, cookies, tokens, JWTs, emails, env secrets), since they're published to GitHub. Screenshots can't be text-redacted, so the report flags when it contains real data.\n- `open_fix_pr` refuses on detached HEAD / in-progress rebase-merge / missing `gh` auth, **stashes and restores** your working tree (protecting unrelated changes), and supports `dryRun`.\n- Snapshot diffs are never auto-updated.\n\n## License\n\nMIT\n","readmeFilename":"README.md"}