{"_id":"@ac12644/promptdiff","_rev":"2-4bd707c88eb4c5c31efaaa574e00c356","name":"@ac12644/promptdiff","dist-tags":{"latest":"0.2.1"},"versions":{"0.2.0":{"name":"@ac12644/promptdiff","version":"0.2.0","keywords":["prompt","llm","diff","evaluation","regression","prompt-engineering","prompt-testing","prompt-optimization","openai","anthropic","gemini","claude"],"license":"MIT","_id":"@ac12644/promptdiff@0.2.0","maintainers":[{"name":"ac12644","email":"ac12644@gmail.com"}],"homepage":"https://github.com/ac12644/prompt-diff#readme","bugs":{"url":"https://github.com/ac12644/prompt-diff/issues"},"bin":{"promptdiff":"bin/promptdiff.js"},"dist":{"shasum":"5f50242f08c671c3db51f546d499a3e53c698a77","tarball":"https://registry.npmjs.org/@ac12644/promptdiff/-/promptdiff-0.2.0.tgz","fileCount":121,"integrity":"sha512-e41AhzVx81BMh4eBSEs9dAUSuzF3+NEb50siBfmkN+jr/1fsrBmZt1xiXJYRSvNdzQYoPDZlbqZab5LUjr930Q==","signatures":[{"sig":"MEYCIQDTPf8bqF5GC/M5VVjy2X1RD7MUKhyv6xP/GjKKwYUmqAIhAMqjezWzlId5U0evMf2EgzlvE1145vEGdlePyeV5VoBM","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":208508},"main":"./dist/orchestrate.js","type":"module","types":"./dist/orchestrate.d.ts","engines":{"node":">=20"},"exports":{".":{"types":"./dist/orchestrate.d.ts","import":"./dist/orchestrate.js"}},"gitHead":"3fe9298cf7af28334c6d29822a500256ba53e40f","scripts":{"test":"vitest run","build":"tsc && node scripts/copy-prices.mjs","typecheck":"tsc --noEmit","test:watch":"vitest","prepublishOnly":"npm run build && npm test","refresh-prices":"node scripts/refresh-prices.mjs"},"_npmUser":{"name":"ac12644","email":"ac12644@gmail.com"},"repository":{"url":"git+https://github.com/ac12644/prompt-diff.git","type":"git"},"_npmVersion":"11.9.0","description":"The regression-test gate for AI-generated prompts. Auto-improve prompts and verify the rewrite against a YAML test suite, or diff any two prompts. Supports OpenAI, Anthropic, Google Gemini.","directories":{},"_nodeVersion":"25.6.1","dependencies":{"zod":"^3.22.0","chalk":"^5.3.0","openai":"^4.0.0","js-yaml":"^4.1.0","p-limit":"^5.0.0","tiktoken":"^1.0.0","commander":"^12.0.0","@google/genai":"^2.6.0","@anthropic-ai/sdk":"^0.24.0"},"_hasShrinkwrap":false,"devDependencies":{"vitest":"^1.6.0","typescript":"^5.4.0","@types/node":"^20.0.0","@types/js-yaml":"^4.0.9"},"_npmOperationalInternal":{"tmp":"tmp/promptdiff_0.2.0_1780012436746_0.09233011464834728","host":"s3://npm-registry-packages-npm-production"}},"0.2.1":{"name":"@ac12644/promptdiff","version":"0.2.1","description":"The regression-test gate for AI-generated prompts. Auto-improve prompts and verify the rewrite against a YAML test suite, or diff any two prompts. Supports OpenAI, Anthropic, Google Gemini.","type":"module","bin":{"promptdiff":"bin/promptdiff.js"},"main":"./dist/orchestrate.js","types":"./dist/orchestrate.d.ts","exports":{".":{"import":"./dist/orchestrate.js","types":"./dist/orchestrate.d.ts"}},"scripts":{"build":"tsc && node scripts/copy-prices.mjs","test":"vitest run","test:watch":"vitest","typecheck":"tsc --noEmit","refresh-prices":"node scripts/refresh-prices.mjs","prepublishOnly":"npm run build && npm test"},"keywords":["prompt","llm","diff","evaluation","regression","prompt-engineering","prompt-testing","prompt-optimization","openai","anthropic","gemini","claude"],"license":"MIT","repository":{"type":"git","url":"git+https://github.com/ac12644/prompt-diff.git"},"bugs":{"url":"https://github.com/ac12644/prompt-diff/issues"},"homepage":"https://github.com/ac12644/prompt-diff#readme","engines":{"node":">=20"},"dependencies":{"@anthropic-ai/sdk":"^0.24.0","@google/genai":"^2.6.0","chalk":"^5.3.0","commander":"^12.0.0","js-yaml":"^4.1.0","openai":"^4.0.0","p-limit":"^5.0.0","tiktoken":"^1.0.0","zod":"^3.22.0"},"devDependencies":{"@types/js-yaml":"^4.0.9","@types/node":"^20.0.0","typescript":"^5.4.0","vitest":"^1.6.0"},"gitHead":"3fe9298cf7af28334c6d29822a500256ba53e40f","_id":"@ac12644/promptdiff@0.2.1","_nodeVersion":"25.6.1","_npmVersion":"11.9.0","dist":{"integrity":"sha512-fi2f9Dvvf8yJ4T7UIwWP6h/fHt0mwI7sWEJBiNzZYZIVX5Q246D31Xn1c8G6t3LIrPWWNvQyxBeTJDIiVbbwnQ==","shasum":"138db1711e5bf76b0ee71d3fcd0c171e2a78f918","tarball":"https://registry.npmjs.org/@ac12644/promptdiff/-/promptdiff-0.2.1.tgz","fileCount":121,"unpackedSize":208508,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIAt5jKGqK5HtZK1ioeuPWK3YpXB4/0vM7NizkhJCCK/QAiEAlZUZqL4SibRIBiyRk9y5ZgemioIecFEbae2+T+t7Y10="}]},"_npmUser":{"name":"ac12644","email":"ac12644@gmail.com"},"directories":{},"maintainers":[{"name":"ac12644","email":"ac12644@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/promptdiff_0.2.1_1780012550677_0.7651632409036722"},"_hasShrinkwrap":false}},"time":{"created":"2026-05-28T23:53:56.616Z","modified":"2026-05-28T23:55:50.921Z","0.2.0":"2026-05-28T23:53:56.922Z","0.2.1":"2026-05-28T23:55:50.839Z"},"bugs":{"url":"https://github.com/ac12644/prompt-diff/issues"},"license":"MIT","homepage":"https://github.com/ac12644/prompt-diff#readme","keywords":["prompt","llm","diff","evaluation","regression","prompt-engineering","prompt-testing","prompt-optimization","openai","anthropic","gemini","claude"],"repository":{"type":"git","url":"git+https://github.com/ac12644/prompt-diff.git"},"description":"The regression-test gate for AI-generated prompts. Auto-improve prompts and verify the rewrite against a YAML test suite, or diff any two prompts. Supports OpenAI, Anthropic, Google Gemini.","maintainers":[{"name":"ac12644","email":"ac12644@gmail.com"}],"readme":"# promptdiff\n\n**The regression-test gate for AI-generated prompts.** Any optimizer (Anthropic's Prompt\nImprover, OpenAI's Playground, DSPy, an in-house tool, or Claude itself) can suggest\na new prompt. promptdiff is what you run before you ship it — to prove the new prompt\nholds the behaviors you actually care about, and at what cost.\n\nTwo modes:\n\n```bash\n# Compare two prompt versions against a YAML test suite\npromptdiff diff v1.txt v2.txt --suite tests.yaml\n\n# Or: hand it your prompt + suite, get a verified-better candidate back\npromptdiff suggest prompt.txt --suite tests.yaml --output improved.txt\n```\n\nIt does one thing well: **decide whether the new prompt is shippable.** Not a prompt\nmanager, not a red-teaming tool, not a dashboard.\n\n---\n\n## Contents\n\n- [Try it on the bundled example](#try-it-on-the-bundled-example)\n- [Install](#install)\n- [`promptdiff suggest` — auto-improve a prompt](#promptdiff-suggest--auto-improve-a-prompt)\n- [`promptdiff diff` — compare two prompts](#promptdiff-diff--compare-two-prompts)\n- [CLI reference](#cli-reference)\n- [Test suite format](#test-suite-format)\n- [Supported providers](#supported-providers)\n- [Library usage](#library-usage)\n- [Cache](#cache)\n- [CI integration](#ci-integration)\n- [Under the hood](#under-the-hood)\n- [What promptdiff doesn't do (by design)](#what-promptdiff-doesnt-do-by-design)\n\n---\n\n## Try it on the bundled example\n\n### `suggest`: rewrite a weak prompt and prove the rewrite is better\n\nReal output from running against a deliberately weak baseline (a generic \"be helpful\"\nprompt) and asking `claude-sonnet-4-5` to fix it:\n\n```\npromptdiff suggest: tests/fixtures/prompts/v0-weak.txt\n────────────────────────────────────────────────────────────\n  Rewriter:          claude-sonnet-4-5\n  Verdict:           ACCEPT — suggestion improves baseline\n  Score vs baseline: 96 / 100\n  Cost (avg/call):   $0.00198 → $0.00076  (-61.7%)\n\n  Baseline weaknesses targeted by the rewriter:\n    ✗ late_shipment_refund_request\n        must contain \"support@lumen.example\"\n    ✗ bulb_wont_connect_to_wifi\n        output length must be under 600 characters\n    ✗ angry_repeat_customer\n        must contain \"support@lumen.example\"\n\n  Per-test outcome (suggestion vs baseline):\n    ✓ late_shipment_refund_request      98\n    ✓ bulb_wont_connect_to_wifi         97\n    ✓ non_lumen_product_question        95\n    ✓ format_structured_reply           95\n    ✓ angry_repeat_customer             93\n```\n\n**Every deterministic failure fixed. Every test passes. Cost down 61.7% per call.** The\nsuggested prompt is saved to the path you pass with `--output`. Exit code is 0 if the\nrewrite beats the threshold, 1 if it doesn't (so CI can gate on it).\n\n### `diff`: prove a hand-written v2 holds the behaviors of v1\n\nSame fixture, comparing the bundled terse v1 against the empathy-first v2:\n\n```\npromptdiff: tests/fixtures/prompts/v1.txt → tests/fixtures/prompts/v2.txt\n────────────────────────────────────────────────────────────\n  Verdict:           PASS\n  Regression Score:  97 / 100\n  Tests:             5 passed, 0 warn, 0 failed (5 total)\n\n  Cost (avg/call):   $0.00051 → $0.00061  (+19.5%)\n  Text diff:         +12 / -5 lines, tokens Δ 68 (+79.1%)\n────────────────────────────────────────────────────────────\n\n  ✓ late_shipment_refund_request      98\n  ✓ bulb_wont_connect_to_wifi         98\n  ✓ non_lumen_product_question        95\n  ✓ format_structured_reply           95\n  ✓ angry_repeat_customer             98\n```\n\nv2 holds every policy v1 held (refund escalation, non-Lumen decline) and scores 95–98 on\nevery judge criterion — but the longer prompt costs **+19.5% per call**. That's the\ntrade-off you couldn't see without this tool.\n\nReproduce both:\n\n```bash\ngit clone https://github.com/ac12644/prompt-diff.git\ncd prompt-diff\ncp .env.example .env       # then paste your ANTHROPIC_API_KEY into .env\nnpm install && npm run build\n\n# Demo 1 — auto-improve the weak baseline\nnode bin/promptdiff.js suggest tests/fixtures/prompts/v0-weak.txt \\\n  -s tests/fixtures/suites/anthropic-smoke.yaml \\\n  --suggester claude-sonnet-4-5 --output /tmp/improved.txt --no-cache\n\n# Demo 2 — diff two existing prompts\nnode bin/promptdiff.js diff \\\n  tests/fixtures/prompts/v1.txt tests/fixtures/prompts/v2.txt \\\n  -s tests/fixtures/suites/anthropic-smoke.yaml --no-cache\n```\n\nA Gemini equivalent (`tests/fixtures/suites/gemini-smoke.yaml`) ships with the same\nscenarios for cross-provider comparison.\n\n---\n\n## Install\n\n```bash\nnpm install -g promptdiff      # global CLI\n# or\nnpx promptdiff …               # no install, one-off run\n# or as a project dep\nnpm install --save-dev promptdiff\n```\n\nThen make API keys available — either via the shell …\n\n```bash\nexport OPENAI_API_KEY=sk-...\nexport ANTHROPIC_API_KEY=sk-ant-...\nexport GEMINI_API_KEY=AIza...\n```\n\n… or drop a `.env` in the project root and `promptdiff` will pick it up automatically.\nCopy `.env.example` to `.env` and fill in only the providers you'll use. `.env` is\ngitignored and never included in the published package. Explicit shell exports always\noverride `.env` values, so you can override per-run.\n\n---\n\n## `promptdiff suggest` — auto-improve a prompt\n\nThe hero command. Give it one prompt + a suite. It runs the suite against your prompt,\nasks a strong LLM (default `claude-opus-4-7`) to rewrite it focused on the failures, runs\nthe suite against the rewrite, and judges baseline-vs-suggestion using the same pipeline\nas `diff`. You get back a **verified** rewrite, not a raw LLM suggestion.\n\n```bash\npromptdiff suggest prompts/v1.txt \\\n  --suite tests/suite.yaml \\\n  --suggester claude-sonnet-4-5 \\\n  --output prompts/v1-improved.txt \\\n  --min-improvement 90\n```\n\n| Flag | Default | Purpose |\n|---|---|---|\n| `--suggester <model>` | `claude-opus-4-7` (or first available) | Strong model used to rewrite the prompt. |\n| `-o, --output <path>` | *(stdout only)* | Where to save the accepted suggestion. |\n| `--min-improvement <n>` | `90` | Reject the rewrite if the diff score is below this. |\n| `--no-cache` | off | Bypass cache. |\n| `--format <type>` | `terminal` | `terminal` or `json` for piping. |\n\n**Exit code is 0** if the rewrite is accepted (score ≥ `--min-improvement` and not failed),\n**1** if rejected. So you can do:\n\n```bash\npromptdiff suggest p.txt -s tests.yaml -o p.improved.txt && mv p.improved.txt p.txt\n```\n\n— and your prompt only gets overwritten if the rewrite actually beat the suite.\n\n---\n\n## `promptdiff diff` — compare two prompts\n\nCreate `v1.txt` (current prompt) and `v2.txt` (candidate). Real example — a customer\nsupport iteration:\n\n```bash\ncat > v1.txt <<'EOF'\nYou are a customer support agent. Respond in 1–3 sentences. Do not promise\nrefunds or compensation — escalate to support@example.com.\nEOF\n\ncat > v2.txt <<'EOF'\nYou are a customer support agent. Tone: warm, calm, concise. Acknowledge the\ncustomer's situation in one short sentence before giving the next step. Keep\nreplies under 80 words. Do not promise refunds — escalate to support@example.com.\nEOF\n```\n\nCreate `tests.yaml` with the scenarios you want to keep working:\n\n```yaml\nmodel: claude-haiku-4-5         # or gpt-4o, gemini-2.5-flash, etc.\njudge_model: claude-haiku-4-5   # cheap model used for llm_judge assertions\nruns_per_test: 1\nconcurrency: 5\n\ntests:\n  - id: refund_request\n    input: \"I've waited 18 days for my order. I want a full refund NOW.\"\n    assert:\n      - type: contains\n        value: \"support@example.com\"\n      - type: not_contains\n        value: \"we'll refund\"\n      - type: llm_judge\n        criteria: \"Calm, does not promise a refund, directs to escalation email\"\n\n  - id: tone_under_pressure\n    input: \"This is the THIRD time I've contacted you. Escalate this NOW.\"\n    assert:\n      - type: not_contains\n        value: \"unfortunately\"\n      - type: llm_judge\n        criteria: \"Acknowledges repeated contact, professional, not defensive\"\n```\n\nRun it:\n\n```bash\npromptdiff v1.txt v2.txt --suite tests.yaml\n```\n\nGate CI on `--min-score 80` and you'll catch the moment a tweak silently breaks behavior.\n\n---\n\n## CLI reference\n\n```\npromptdiff diff <v1> <v2> --suite <path> [options]\npromptdiff suggest <prompt> --suite <path> [options]\n```\n\n### `diff` options\n\n```\n  -s, --suite <path>  Path to test suite YAML  (required)\n  -m, --model <name>  Override model from suite\n  --min-score <n>     Exit 1 if regression score is below this (0–100)\n  --no-cache          Skip the response cache; always call the provider\n  --format <type>     Output format: terminal | json   (default: terminal)\n```\n\n### `suggest` options\n\n```\n  -s, --suite <path>           Path to test suite YAML  (required)\n  --suggester <model>          Rewriter model (default: claude-opus-4-7)\n  -o, --output <path>          Save the suggested prompt to this file\n  --min-improvement <n>        Reject if diff score is below this (default: 90)\n  --no-cache                   Skip the response cache\n  --format <type>              terminal | json (default: terminal)\n```\n\n### Global\n\n```\n  -V, --version       Print version\n  -h, --help          Print help\n```\n\n### Exit codes\n\n| Code | Meaning |\n|------|---------|\n| 0    | All good. If `--min-score` was set, score met threshold. |\n| 1    | Regression: score is below `--min-score`. |\n| 2    | Config/file error (missing suite, malformed YAML, etc.). |\n| 3    | Provider API error. |\n\n### Environment variables\n\n| Var                     | Purpose |\n|-------------------------|---------|\n| `OPENAI_API_KEY`        | Required for `gpt-*`, `o1-*`, `o3-*`, `o4-*` models. |\n| `ANTHROPIC_API_KEY`     | Required for `claude-*` models. |\n| `GEMINI_API_KEY`        | Required for `gemini-*` models. (`GOOGLE_API_KEY` is accepted as a fallback.) |\n| `PROMPTDIFF_CACHE_DIR`  | Override cache directory (default: `.promptdiff-cache`). |\n\n---\n\n## Test suite format\n\n```yaml\nmodel: gpt-4o                # required: main model both prompts run against\njudge_model: gpt-4o-mini     # optional: cheap model for llm_judge (default: gpt-4o-mini)\nruns_per_test: 1             # optional: 1–10. Use >1 to average over output variance.\nconcurrency: 5               # optional: 1–20. Max parallel API calls.\n\ntests:\n  - id: greeting             # required: unique per test\n    input: \"Say hello to {{name}}.\"   # required: supports {{var}} interpolation\n    vars:                    # optional: fills {{var}} placeholders in input\n      name: \"Alex\"\n    assert:                  # required: one or more assertions\n      - type: contains\n        value: \"hello\"\n```\n\n> **About `runs_per_test`:** Single calls can swing in length and phrasing. If your\n> assertions are sensitive (e.g. `length_under`), set `runs_per_test: 3–5` to average\n> across runs and reduce false signals. Cost scales linearly with this number.\n\n### Assertion types\n\n| Type            | Field      | Meaning |\n|-----------------|------------|---------|\n| `contains`      | `value`    | output must include the substring. |\n| `not_contains`  | `value`    | output must not include the substring. |\n| `length_under`  | `value`    | output length (chars) must be < value. |\n| `starts_with`   | `value`    | output (trimmed) must start with value. |\n| `regex`         | `value`    | regex must match the output. |\n| `llm_judge`     | `criteria` | a cheap judge model scores both outputs against this criterion. |\n\n### How scoring works\n\nFor each test pair, every deterministic assertion contributes 100 (pass or non-regression)\nor 0 (v2 broke an assertion v1 passed). Each `llm_judge` assertion contributes the judge's\nscore for v2 (0–100). The per-test score is the mean of these; the top-level score is the\nmean across all tests.\n\n| Aggregate score | Verdict |\n|-----------------|---------|\n| ≥ 90            | `pass`  |\n| 70 – 89         | `warn`  |\n| < 70            | `fail`  |\n\nThe top-level verdict is also pushed to **fail** if a majority of individual tests failed,\nor to at-least-**warn** if any test failed — so a single catastrophic test doesn't get\naveraged away. A test where v2 returns a run error always reports `fail` regardless of\nscore.\n\n---\n\n## Supported providers\n\n| Model prefix                    | Provider  | Env var(s)                            |\n|---------------------------------|-----------|---------------------------------------|\n| `gpt-*`, `o1-*`, `o3-*`, `o4-*` | OpenAI    | `OPENAI_API_KEY`                      |\n| `claude-*`                      | Anthropic | `ANTHROPIC_API_KEY`                   |\n| `gemini-*`                      | Google    | `GEMINI_API_KEY` *(or `GOOGLE_API_KEY`)* |\n\nMix freely: e.g. `model: claude-opus-4-7` with `judge_model: gemini-2.5-flash-lite`.\n\n---\n\n## Library usage\n\nBeyond the CLI, promptdiff exports `orchestrate()` so you can embed it in your own tooling\nor test harnesses. The compiled package ships TypeScript declaration files.\n\n```typescript\nimport { orchestrate } from 'promptdiff'\n\nconst report = await orchestrate(\n  'prompts/v1.txt',\n  'prompts/v2.txt',\n  'tests/suite.yaml',\n  {\n    apiKeys: {\n      openai:    process.env.OPENAI_API_KEY,\n      anthropic: process.env.ANTHROPIC_API_KEY,\n      gemini:    process.env.GEMINI_API_KEY,\n    },\n    format: 'json',    // suppresses the terminal reporter\n    noCache: false,\n  },\n)\n\nconsole.log(`Score: ${report.regressionScore}, verdict: ${report.verdict}`)\nif (report.regressionScore < 85) process.exit(1)\n```\n\nThe full `DiffReport` shape (cost delta, per-test results, assertion outcomes, text diff)\nis exported from `promptdiff` as well — see `dist/types.d.ts` for the complete type\nsurface.\n\n---\n\n## Cache\n\nResponses are cached by `sha256(model + prompt + input)` under `.promptdiff-cache/`\n(gitignore it). A second run with unchanged prompts and inputs is **instant and free**.\nBypass with `--no-cache`. Override the directory with `PROMPTDIFF_CACHE_DIR`.\n\n---\n\n## CI integration\n\n```bash\npromptdiff prompts/v1.txt prompts/v2.txt \\\n  --suite tests/suite.yaml \\\n  --min-score 85 \\\n  --format json > diff-report.json\n```\n\nExit code is 1 if the aggregate score drops below 85. The JSON output is suitable for\nposting as a PR comment or uploading as a build artifact.\n\n---\n\n## Under the hood\n\n### How pricing stays fresh\n\nNone of OpenAI, Anthropic, or Google return cost in their API responses — only token\ncounts. So promptdiff computes USD cost client-side from a price table. Rather than\nhand-maintain that table (which silently goes stale), we bundle a filtered snapshot of\n[LiteLLM's community-maintained registry](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json),\nwhich tracks 2,700+ models and updates within days of any provider price change.\n\n```bash\nnpm run refresh-prices   # re-downloads upstream, writes src/providers/prices.json\n```\n\nThe snapshot ships at `src/providers/prices.json` (~26 KB, 170 entries filtered to\nchat-mode OpenAI/Anthropic/Gemini models). Unknown models fall back to a\n$1/$3-per-million estimate so reports always show a non-zero cost — better than a\nsilently-wrong `$0.00000`.\n\n### Architecture\n\nThree layers with one-way dependencies (inward only):\n\n- **IO** (`cli/`, `config/`, `providers/`, `reporters/`) — touches the outside world.\n- **Core** (`core/differ/`, `core/runner/`, `core/judge/`, `core/scorer/`) — pure\n  functions, fully tested without mocks or network.\n- **Infra** (`infra/cache`, `infra/dotenv`, `infra/errors`, `infra/logger`) —\n  cross-cutting concerns.\n\n`src/orchestrate.ts` is the only place that wires modules together. Each module has\nexactly one job, and assertions are data dispatched by a single `evaluate()` switch — so\nadding a new assertion type is a 5-file change with a well-defined surface.\n\n---\n\n## What promptdiff doesn't do (by design)\n\n- No prompt storage or versioning database\n- No real-time streaming\n- No multi-agent or chain evaluation\n- No automatic test case generation\n- No web UI\n\nRun it from the CLI, gate CI on the score, move on.\n\n---\n\n## License\n\n[MIT](./LICENSE)\n","readmeFilename":"README.md"}