{"_id":"@bernardo75/ai-test-analyzer","_rev":"3-3df1c371b12732877b118dc7fe738b9b","name":"@bernardo75/ai-test-analyzer","dist-tags":{"latest":"0.3.0"},"versions":{"0.1.0":{"name":"@bernardo75/ai-test-analyzer","version":"0.1.0","keywords":["testing","allure","playwright","junit","test-report","ai","anthropic","claude","failure-analysis","flaky-tests"],"license":"MIT","_id":"@bernardo75/ai-test-analyzer@0.1.0","maintainers":[{"name":"bernardo75","email":"giorgibernardo75@gmail.com"}],"dist":{"shasum":"db94febee93aebc1f614cb952f2a4a60853d93c2","tarball":"https://registry.npmjs.org/@bernardo75/ai-test-analyzer/-/ai-test-analyzer-0.1.0.tgz","fileCount":11,"integrity":"sha512-nIiaFQ01mNRFzAp/pNPNgGdCfsWiuIIkxI3fv1POmOEZpu2XY1bqgFcZBfeaLSljI6T01klH9UbMLQ4Xb1sgBQ==","signatures":[{"sig":"MEUCICLhRuFYd4dQKqmlVCI0LD1oHZFB4EcdXNxCCLO6rgDGAiEAriJSPzCciL+0aCQ/ySBGkaTGXzHVIn6W0yaH220HRSw=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":131190},"main":"./dist/index.cjs","type":"module","types":"./dist/index.d.ts","module":"./dist/index.js","engines":{"node":">=20"},"exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js","require":"./dist/index.cjs"}},"gitHead":"8e0ed161724c4eef82e78f3567e127f480dbde9b","scripts":{"lint":"eslint .","test":"vitest run","build":"tsup","typecheck":"tsc --noEmit","test:watch":"vitest"},"_npmUser":{"name":"bernardo75","email":"giorgibernardo75@gmail.com"},"_npmVersion":"10.2.3","description":"AI-powered analysis of automated test reports (Allure, Playwright, JUnit XML): classifies failures as product bugs, broken tests, flaky or environment issues, and proposes diagnosis and fixes for broken tests.","directories":{},"_nodeVersion":"20.10.0","dependencies":{"zod":"^4.4.3","fast-xml-parser":"^5.10.1","@anthropic-ai/sdk":"^0.112.1"},"_hasShrinkwrap":false,"devDependencies":{"tsup":"^8.5.1","eslint":"^9.39.5","vitest":"^3.2.7","typescript":"^6.0.3","@types/node":"^26.1.1","typescript-eslint":"^8.64.0"},"_npmOperationalInternal":{"tmp":"tmp/ai-test-analyzer_0.1.0_1784339826518_0.753686315771108","host":"s3://npm-registry-packages-npm-production"}},"0.2.0":{"name":"@bernardo75/ai-test-analyzer","version":"0.2.0","keywords":["testing","allure","playwright","junit","test-report","ai","anthropic","claude","failure-analysis","flaky-tests"],"license":"MIT","_id":"@bernardo75/ai-test-analyzer@0.2.0","maintainers":[{"name":"bernardo75","email":"giorgibernardo75@gmail.com"}],"dist":{"shasum":"41884f39d6dc5cc19b578e03a6ab3b805775085c","tarball":"https://registry.npmjs.org/@bernardo75/ai-test-analyzer/-/ai-test-analyzer-0.2.0.tgz","fileCount":11,"integrity":"sha512-zDrvJgmU/9/GhNP7QwmO6hNBG6RDIV/478+PZyoCQ1vMshStS9t6Hz9yT8CVAO8kyM6evEiml8BFfbyGAZBFbw==","signatures":[{"sig":"MEYCIQCv0p/etOGt06zMQ8Z1bx+wy7IrrAgJOQqUqrEqgRsTbwIhAOPrT4pKP9MbPnozUZZCrkhp7uJBH2LpiaVI+WXFz2qM","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":142520},"main":"./dist/index.cjs","type":"module","types":"./dist/index.d.ts","module":"./dist/index.js","engines":{"node":">=20"},"exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js","require":"./dist/index.cjs"}},"gitHead":"8e0e219ec7878c4b53633f9c03a451923ef5f336","scripts":{"lint":"eslint .","test":"vitest run","build":"tsup","typecheck":"tsc --noEmit","test:watch":"vitest"},"_npmUser":{"name":"bernardo75","email":"giorgibernardo75@gmail.com"},"_npmVersion":"10.2.3","description":"AI-powered analysis of automated test reports (Allure, Playwright, JUnit XML): classifies failures as product bugs, broken tests, flaky or environment issues, and proposes diagnosis and fixes for broken tests.","directories":{},"_nodeVersion":"20.10.0","dependencies":{"zod":"^4.4.3","fast-xml-parser":"^5.10.1","@anthropic-ai/sdk":"^0.112.1"},"_hasShrinkwrap":false,"devDependencies":{"tsup":"^8.5.1","eslint":"^9.39.5","vitest":"^3.2.7","typescript":"^6.0.3","@types/node":"^26.1.1","typescript-eslint":"^8.64.0"},"_npmOperationalInternal":{"tmp":"tmp/ai-test-analyzer_0.2.0_1784343508891_0.01802031389782499","host":"s3://npm-registry-packages-npm-production"}},"0.3.0":{"name":"@bernardo75/ai-test-analyzer","version":"0.3.0","description":"AI-powered analysis of automated test reports (Allure, Playwright, JUnit XML): classifies failures as product bugs, broken tests, flaky or environment issues, and proposes diagnosis and fixes for broken tests.","type":"module","main":"./dist/index.cjs","module":"./dist/index.js","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js","require":"./dist/index.cjs"}},"engines":{"node":">=20"},"scripts":{"build":"tsup","test":"vitest run","test:watch":"vitest","lint":"eslint .","typecheck":"tsc --noEmit"},"keywords":["testing","allure","playwright","junit","test-report","ai","anthropic","claude","failure-analysis","flaky-tests"],"license":"MIT","dependencies":{"@anthropic-ai/sdk":"^0.112.1","fast-xml-parser":"^5.10.1","zod":"^4.4.3"},"devDependencies":{"@types/node":"^26.1.1","eslint":"^9.39.5","tsup":"^8.5.1","typescript":"^6.0.3","typescript-eslint":"^8.64.0","vitest":"^3.2.7"},"_id":"@bernardo75/ai-test-analyzer@0.3.0","gitHead":"92d6960c9616e6af3073f991b29a1c758979143d","_nodeVersion":"20.10.0","_npmVersion":"10.2.3","dist":{"integrity":"sha512-JQ9bzOqR9JeMQIoOb2oobWmHiwAAPWUOpshZhQhRlrr4cVtj/ElbUgubkp6AeVs5HuEG7fzla16DRv1xavYLjQ==","shasum":"5f6a8b2766b1ca6e92394b38787b39b17d264854","tarball":"https://registry.npmjs.org/@bernardo75/ai-test-analyzer/-/ai-test-analyzer-0.3.0.tgz","fileCount":11,"unpackedSize":148944,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIH8vmKvZOSNSMagn+bdE55KYZg0tisNHv0GFiGvBpoY9AiEA9mRMyauOfFcMIOOR1UGDHFipDrhJQvI98PmPYAon1Ns="}]},"_npmUser":{"name":"bernardo75","email":"giorgibernardo75@gmail.com"},"directories":{},"maintainers":[{"name":"bernardo75","email":"giorgibernardo75@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/ai-test-analyzer_0.3.0_1786024802021_0.7793101384430319"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-18T01:57:06.322Z","modified":"2026-08-06T14:00:02.349Z","0.1.0":"2026-07-18T01:57:06.677Z","0.2.0":"2026-07-18T02:58:29.070Z","0.3.0":"2026-08-06T14:00:02.188Z"},"license":"MIT","keywords":["testing","allure","playwright","junit","test-report","ai","anthropic","claude","failure-analysis","flaky-tests"],"description":"AI-powered analysis of automated test reports (Allure, Playwright, JUnit XML): classifies failures as product bugs, broken tests, flaky or environment issues, and proposes diagnosis and fixes for broken tests.","maintainers":[{"name":"bernardo75","email":"giorgibernardo75@gmail.com"}],"readme":"# ai-test-analyzer\r\n\r\nAI-powered analysis of automated test reports. Feed it an **Allure results\r\ndirectory**, a **Playwright JSON report** or a **JUnit XML file** and it\r\nclassifies every failure as:\r\n\r\n| Verdict | Meaning |\r\n|---|---|\r\n| `product_bug` | The product misbehaves; the test is right |\r\n| `broken_test` | The test is stale or badly built — includes a **diagnosis** and a **recommended fix** |\r\n| `flaky` | Non-deterministic failure (timeouts, races, ordering) — includes a **stabilization fix** |\r\n| `environment_issue` | Infra problem (network, DNS, 5xx dependencies, disk) |\r\n\r\nFor Allure it can also **enrich the report itself**: every failed test gets an\r\n\"🤖 AI Analysis\" section, a filterable `ai-verdict` label, and the Categories\r\ntab groups failures by verdict.\r\n\r\nPowered by Anthropic Claude behind a minimal provider interface. Disabled by\r\ndefault — with the feature flag off, every call is a strict no-op (zero network,\r\nzero tokens, zero file writes).\r\n\r\n## Install\r\n\r\n```sh\r\nnpm install @bernardo75/ai-test-analyzer\r\n```\r\n\r\nRequires Node.js ≥ 20.\r\n\r\n## Quickstart\r\n\r\n```ts\r\nimport { analyze } from \"@bernardo75/ai-test-analyzer\";\r\n\r\n// Enable via env: AI_ANALYZER_ENABLED=true, ANTHROPIC_API_KEY=sk-ant-...\r\nconst result = await analyze({\r\n  reportPath: \"./allure-results\",\r\n  format: \"allure\", // \"allure\" | \"playwright\" | \"junit\"\r\n});\r\n\r\nif (result.status === \"completed\") {\r\n  for (const [testId, entry] of Object.entries(result.verdictsByTest)) {\r\n    if (\"verdict\" in entry) {\r\n      console.log(testId, entry.verdict, entry.confidence);\r\n      if (entry.verdict === \"broken_test\") {\r\n        console.log(\"  diagnosis:\", entry.diagnosis);\r\n        console.log(\"  fix:\", entry.recommendedFix);\r\n      }\r\n    } else {\r\n      console.log(testId, \"analysis failed:\", entry.reason);\r\n    }\r\n  }\r\n  console.log(result.summary);\r\n}\r\n```\r\n\r\n### Analyze failures already in memory (no report file)\r\n\r\nWhen you already have the failures — from a CI REST API (Jenkins, GitHub\r\nActions), a database, or pasted output — feed them directly, no file on disk\r\nrequired:\r\n\r\n```ts\r\nimport { analyzeFailures } from \"@bernardo75/ai-test-analyzer\";\r\n\r\nconst result = await analyzeFailures({\r\n  failures: [\r\n    {\r\n      name: \"should log in\",\r\n      fullName: \"auth.LoginSuite › should log in\",\r\n      errorMessage: \"TimeoutError: locator not enabled\",\r\n      stackTrace: \"at login.spec.ts:42\",\r\n      file: \"login.spec.ts\",\r\n      line: 42,\r\n    },\r\n  ],\r\n  totalTests: 20,               // optional — defaults to failures.length\r\n  config: { enabled: true },    // same feature flag + config as analyze()\r\n});\r\n```\r\n\r\nOnly `name` plus an error is required per failure; everything else defaults.\r\nSame feature flag, config, providers and result shape as `analyze()` — the\r\ninput just comes from memory instead of a report path.\r\n\r\n### Enrich the Allure report\r\n\r\n```ts\r\nimport { analyzeAndEnrich } from \"@bernardo75/ai-test-analyzer\";\r\n\r\n// Run BEFORE `allure generate`:\r\nconst { analysis, enrichment } = await analyzeAndEnrich({\r\n  reportPath: \"./allure-results\",\r\n  format: \"allure\",\r\n});\r\n// then: npx allure generate ./allure-results -o ./allure-report\r\n```\r\n\r\nEnrichment is **strictly additive and idempotent**: it only appends content\r\n(attachments, labels, description blocks, a `[ai-verdict:*]` message suffix,\r\n`categories.json` entries and an `ai.analyzer.*` block in\r\n`environment.properties`), never deletes or replaces existing results, and\r\nre-running it is a no-op.\r\n\r\n## Feature flag\r\n\r\nThe analyzer is **off by default**. Precedence:\r\n\r\n1. `config.enabled` (programmatic — wins in both directions)\r\n2. `AI_ANALYZER_ENABLED` env var (`\"true\"` / `\"1\"` to enable)\r\n3. Default: disabled\r\n\r\nWith the flag off, `analyze()` / `enrichAllureResults()` / `analyzeAndEnrich()`\r\nreturn `{status: \"disabled\"}` without reading files, calling the network or\r\nconsuming tokens — safe to leave installed in CI.\r\n\r\n## Configuration\r\n\r\n```ts\r\nawait analyze({\r\n  reportPath: \"./allure-results\",\r\n  format: \"allure\",\r\n  config: {\r\n    enabled: true,                 // overrides AI_ANALYZER_ENABLED\r\n    apiKey: process.env.MY_KEY,    // overrides ANTHROPIC_API_KEY\r\n    model: \"claude-sonnet-5\",       // default\r\n    maxConcurrency: 4,             // parallel analyses (default 4)\r\n    maxGroups: 25,                 // hard cost cap per run (default: unlimited)\r\n    pricing: { inputPerMTok: 5, outputPerMTok: 25 }, // cost-estimate override\r\n    failureContext: {              // extra failure evidence sent to the LLM\r\n      pageSnapshot: true,          //   Playwright error-context (text, ~8k cap)\r\n      screenshot: true,            //   failure screenshot (vision providers only)\r\n    },\r\n    provider: myCustomProvider,    // inject your own AnalysisProvider\r\n  },\r\n});\r\n```\r\n\r\n## Cost controls\r\n\r\n- **Only failures are analyzed** — a green run costs nothing.\r\n- **Deduplication**: failures sharing a normalized signature (masked error +\r\n  own stack frames) are analyzed **once**; the verdict fans out to every test.\r\n  40 failures with 5 root causes = 5 LLM calls.\r\n- **Prompt caching**: the system prompt is cached across groups of a run.\r\n- **`maxGroups`**: hard ceiling of analyses per run.\r\n- **Usage report**: every run returns totals — failures, groups, duplicates\r\n  avoided, token usage and an estimated cost in USD.\r\n- **Failure evidence**: when the report carries a Playwright `error-context`\r\n  page snapshot and/or a failure screenshot, they are included in the analysis\r\n  (size-capped, gated by `config.failureContext`) — decisive for telling a\r\n  renamed selector (`broken_test`) apart from a missing feature\r\n  (`product_bug`). The screenshot is used by `AnthropicProvider` (vision);\r\n  `ClaudeCodeProvider` uses the text snapshot only.\r\n\r\n```ts\r\nresult.summary;\r\n// {\r\n//   totalFailures: 40, groupsAnalyzed: 5, groupsFailed: 0,\r\n//   duplicatesAvoided: 35,\r\n//   usage: { inputTokens, outputTokens, cacheCreationInputTokens, cacheReadInputTokens },\r\n//   estimatedCostUsd: 0.19, model: \"claude-sonnet-5\", durationMs: 41200\r\n// }\r\n```\r\n\r\n## API\r\n\r\n| Export | Description |\r\n|---|---|\r\n| `analyze(options)` | Parse + classify. Returns `AnalysisResult` (`disabled` \\| `nothing_to_analyze` \\| `completed`). Never rejects on provider failures — failed groups are reported per-test as `analysis_failed`. |\r\n| `enrichAllureResults(dir, analysis, options?)` | Additive, idempotent enrichment of an allure-results directory. |\r\n| `analyzeAndEnrich(options)` | Both steps in one call (Allure only). |\r\n| `AnthropicProvider` | Production provider (official `@anthropic-ai/sdk`, structured outputs, prompt caching). |\r\n| `MockProvider` | Deterministic provider for your own tests — zero network. |\r\n| Errors | `ReportNotFoundError`, `ReportParseError`, `ConfigurationError`, `ProviderError`. |\r\n\r\nCredentials are never logged, persisted or written into enriched reports. The\r\nonly data sent to the LLM is the failure context: error message, stack trace,\r\nsteps and a ±20-line snippet of the failing test file.\r\n\r\n## What the analysis looks like in Allure\r\n\r\nEach failed test gets:\r\n\r\n- An attachment **🤖 AI Analysis** with verdict, confidence, reasoning and — for\r\n  broken tests — diagnosis and recommended fix.\r\n- A label `ai-verdict=<verdict>` you can filter/search by.\r\n- A summary block appended to the test description.\r\n- Grouping under **Categories**: \"Product Bugs (AI)\", \"Broken Tests (AI)\",\r\n  \"Flaky (AI)\", \"Environment Issues (AI)\".\r\n\r\nAnd the report home shows the run summary in the **Environment** widget\r\n(`ai.analyzer.*` keys in `environment.properties`): model, groups analyzed,\r\nduplicates avoided, token usage and estimated cost. User-authored properties\r\nare preserved.\r\n\r\n## Experimental: ClaudeCodeProvider (no API key)\r\n\r\nIf the machine has [Claude Code](https://claude.com/claude-code) installed and\r\nauthenticated — including via a Claude Pro/Max subscription — you can route the\r\nanalysis through it instead of the API:\r\n\r\n```ts\r\nimport { analyze, ClaudeCodeProvider } from \"@bernardo75/ai-test-analyzer\";\r\n\r\nconst result = await analyze({\r\n  reportPath: \"./allure-results\",\r\n  format: \"allure\",\r\n  config: {\r\n    enabled: true,\r\n    provider: new ClaudeCodeProvider(), // uses the local `claude` CLI headless\r\n  },\r\n});\r\n```\r\n\r\nFor CI, generate a long-lived credential with `claude setup-token` and expose\r\nit to the runner. **Read this before using it in a pipeline:**\r\n\r\n- ⚠️ Subscriptions are **personal**. Sharing one account across a team pipeline\r\n  violates Anthropic's usage terms — use an API key (workspace-owned) for team\r\n  CI. This provider is meant for personal projects and evaluation.\r\n- Subscription rate-limit windows are shared with your interactive usage; CI\r\n  runs may fail unpredictably when the window is exhausted.\r\n- No structured outputs: the verdict is parsed from text and re-validated\r\n  against the same schema — invalid responses become per-group analysis\r\n  failures, never crashes.\r\n- Usage/cost reporting depends on what the CLI returns (may be zeros).\r\n\r\nOptions: `cliPath` (default `\"claude\"`), `model` (passed as `--model`, default `\"sonnet\"`),\r\n`timeoutMs` (default 120000).\r\n\r\n## Examples\r\n\r\nRunnable scripts in [`examples/`](examples/): mock analysis (no cost), real\r\nanalysis, enrichment end-to-end, multi-format, and an accuracy-evaluation\r\nharness against a human-labeled dataset.\r\n\r\n## License\r\n\r\nMIT\r\n","readmeFilename":"README.md"}