{"_id":"@avinashchby/promptbench","_rev":"2-180c13f95e31fc074e3faea30d2faae2","name":"@avinashchby/promptbench","dist-tags":{"latest":"0.1.1"},"versions":{"0.1.0":{"name":"@avinashchby/promptbench","version":"0.1.0","keywords":["claude","claude-md","cursorrules","ai-config","testing","linting","prompt-engineering","coding-assistant"],"license":"MIT","_id":"@avinashchby/promptbench@0.1.0","maintainers":[{"name":"avinashchby","email":"avinashchby@gmail.com"}],"bin":{"promptbench":"bin/promptbench.js"},"dist":{"shasum":"cc2881d72ec47c6c8af460eb82d00b513491c822","tarball":"https://registry.npmjs.org/@avinashchby/promptbench/-/promptbench-0.1.0.tgz","fileCount":72,"integrity":"sha512-BIq7MGoleskgTkl9DqfDyu+HK3MvdIxYsUqHJSghcxOzqqcO25ukFNjqabwjvkT5A4Qo4VDXlMnXRUIuR+/rGw==","signatures":[{"sig":"MEQCIAXb4tQutLfnLK8r4ja7KC5s09kgbU6CY0aDhTyBfTWKAiAIJQC57g10rR9VL42bYF9eLacvcdHuwHo2lcPwF9quCg==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":122082},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","engines":{"node":">=18"},"gitHead":"c34f587c17a14b7af92f21e557d3519bf880d06d","scripts":{"dev":"tsc --watch","lint":"tsc --noEmit","test":"vitest run","build":"tsc","test:watch":"vitest","prepublishOnly":"npm run build"},"_npmUser":{"name":"avinashchby","email":"avinashchby@gmail.com"},"_npmVersion":"11.8.0","description":"The testing framework for AI coding assistant configuration files","directories":{},"_nodeVersion":"25.5.0","dependencies":{"chalk":"^5.4.1","js-yaml":"^4.1.0","commander":"^13.1.0"},"_hasShrinkwrap":false,"devDependencies":{"vitest":"^3.0.0","typescript":"^5.7.3","@types/node":"^22.13.0","@types/js-yaml":"^4.0.9"},"_npmOperationalInternal":{"tmp":"tmp/promptbench_0.1.0_1773910685801_0.7750781878702442","host":"s3://npm-registry-packages-npm-production"}},"0.1.1":{"name":"@avinashchby/promptbench","version":"0.1.1","description":"The testing framework for AI coding assistant configuration files","type":"module","main":"./dist/index.js","types":"./dist/index.d.ts","bin":{"promptbench":"bin/promptbench.js"},"scripts":{"build":"tsc","dev":"tsc --watch","test":"vitest run","test:watch":"vitest","lint":"tsc --noEmit","prepublishOnly":"npm run build"},"keywords":["claude","claude-md","cursorrules","ai-config","testing","linting","prompt-engineering","coding-assistant"],"license":"MIT","engines":{"node":">=18"},"dependencies":{"chalk":"^5.4.1","commander":"^13.1.0","js-yaml":"^4.1.0"},"devDependencies":{"@types/js-yaml":"^4.0.9","@types/node":"^22.13.0","typescript":"^5.7.3","vitest":"^3.0.0"},"gitHead":"417bde118c446e76d989a5c93a3b3cae221cf6c7","_id":"@avinashchby/promptbench@0.1.1","_nodeVersion":"25.5.0","_npmVersion":"11.8.0","dist":{"integrity":"sha512-vxx8GFLCIYe3x7WVzlAhzWIJvqc33wR7Qd5/LX/dcy09Ah12SpzE0wPZjSggtWzhDj42mWtB5gp/NTqOQlPKeQ==","shasum":"0d918d158db3bb5e455968715b9fbb5bf036d0d1","tarball":"https://registry.npmjs.org/@avinashchby/promptbench/-/promptbench-0.1.1.tgz","fileCount":72,"unpackedSize":122082,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIGG5GVqYHWugixFggDQ2o60vXAU8DtktDtLkzvTQvEZiAiBHWO6VgFjXVatQAsIbOMkSxkX5G1g5tgCjxlji4knZhw=="}]},"_npmUser":{"name":"avinashchby","email":"avinashchby@gmail.com"},"directories":{},"maintainers":[{"name":"avinashchby","email":"avinashchby@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/promptbench_0.1.1_1773911041963_0.7678236863438481"},"_hasShrinkwrap":false}},"time":{"created":"2026-03-19T08:58:05.710Z","modified":"2026-03-19T09:04:02.573Z","0.1.0":"2026-03-19T08:58:05.974Z","0.1.1":"2026-03-19T09:04:02.115Z"},"license":"MIT","keywords":["claude","claude-md","cursorrules","ai-config","testing","linting","prompt-engineering","coding-assistant"],"description":"The testing framework for AI coding assistant configuration files","maintainers":[{"name":"avinashchby","email":"avinashchby@gmail.com"}],"readme":"# promptbench\n\n**The testing framework for AI coding assistant configuration files.**\n\n[![npm version](https://img.shields.io/npm/v/promptbench.svg)](https://www.npmjs.com/package/promptbench)\n[![license](https://img.shields.io/npm/l/promptbench.svg)](https://github.com/avinashchaubey/promptbench/blob/main/LICENSE)\n\nYou spend hours writing your CLAUDE.md and .cursorrules. But does your AI assistant actually follow them? Does it use pnpm like you told it? Does it avoid `any` types? You don't know until you waste tokens finding out.\n\n**promptbench** analyzes your AI config files for quality, consistency, and completeness — then lets you write test cases to verify expected behavior.\n\n## Install\n\n```bash\nnpm install -g promptbench\n# or\nnpx promptbench\n```\n\n## Quick Start\n\n```bash\n# Audit your CLAUDE.md\nnpx promptbench --audit\n\n# Generate a test file\nnpx promptbench init\n\n# Run tests\nnpx promptbench\n```\n\n## How It Works\n\n### Mode 1: Config Analysis (no LLM needed)\n\nAnalyzes your config file itself for quality issues:\n\n```bash\nnpx promptbench --audit\n```\n\n```\n┌──────────────────────────────────────────────────┐\n│  CLAUDE.md Audit Report                          │\n├──────────────────────────────────────────────────┤\n│  Overall Score: 72/100                           │\n├──────────────────────────────────────────────────┤\n│  ✅ Has project overview                         │\n│  ✅ Has build commands                           │\n│  ✅ Has coding style rules                       │\n│  ⚠️  Missing testing section                     │\n│  ⚠️  No error handling rules                     │\n│  ❌ Contradicting rules found:                   │\n│     Line 12: \"use tabs\"                          │\n│     Line 45: \"use 2 spaces\"                      │\n│  ❌ 3 vague instructions found                   │\n│     Line 23: \"write good code\"                   │\n│     Line 67: \"be careful\"                        │\n├──────────────────────────────────────────────────┤\n│  2 errors, 2 warnings, 3 info                    │\n└──────────────────────────────────────────────────┘\n```\n\nDetectors:\n- **Contradictions** — finds conflicting rules (tabs vs spaces, npm vs pnpm)\n- **Vagueness** — flags weasel words (\"best practices\", \"be careful\", \"write good code\")\n- **Completeness** — checks for recommended sections (overview, build commands, testing, etc.)\n- **Specificity** — scores how actionable your instructions are\n- **Metrics** — line count, token estimate, section analysis\n\n### Mode 2: Behavioral Simulation (optional, needs API key)\n\nTest if your config actually produces expected AI behavior:\n\n```bash\nexport ANTHROPIC_API_KEY=sk-...\nnpx promptbench --simulate\n```\n\n## Test File Format\n\nCreate `.promptbench.yml` in your project root:\n\n```yaml\nconfig: ./CLAUDE.md\n\ntests:\n  - name: \"Uses pnpm not npm\"\n    scenario: \"Install a new package\"\n    expect:\n      contains: [\"pnpm add\", \"pnpm install\"]\n      not_contains: [\"npm install\", \"yarn add\"]\n\n  - name: \"No any types in TypeScript\"\n    scenario: \"Create a TypeScript function\"\n    expect:\n      not_contains: [\"any\"]\n      contains: [\"interface\", \"type\"]\n\n  - name: \"Uses conventional commits\"\n    scenario: \"Commit changes\"\n    expect:\n      pattern: \"^(feat|fix|chore|docs|refactor|test)\"\n\n  - name: \"Has build commands\"\n    check: \"config_contains\"\n    expect:\n      config_has: [\"npm run build\", \"npm run test\"]\n\n  - name: \"Config not too long\"\n    check: \"config_metrics\"\n    expect:\n      max_lines: 500\n      max_tokens: 8000\n\n  - name: \"No contradictions\"\n    check: \"config_consistency\"\n    expect:\n      no_contradictions: true\n```\n\n## CLI\n\n```bash\nnpx promptbench                        # Run all tests in .promptbench.yml\nnpx promptbench --config CLAUDE.md     # Analyze a specific config\nnpx promptbench --audit                # Full quality audit report\nnpx promptbench --score                # 0-100 quality score\nnpx promptbench --fix                  # Suggest improvements\nnpx promptbench --simulate             # Behavioral tests (needs API key)\nnpx promptbench --ci                   # CI mode: JSON output, exit 1 on failure\nnpx promptbench --format json          # JSON output\nnpx promptbench --format markdown      # Markdown report\nnpx promptbench init                   # Generate sample .promptbench.yml\n```\n\n## CI Integration\n\n```yaml\n# .github/workflows/promptbench.yml\nname: Config Quality\non: [push, pull_request]\njobs:\n  check:\n    runs-on: ubuntu-latest\n    steps:\n      - uses: actions/checkout@v4\n      - uses: actions/setup-node@v4\n      - run: npx promptbench --ci --min-score 70\n```\n\n## Programmatic API\n\n```typescript\nimport { audit, analyze, parseConfig } from 'promptbench';\n\nconst report = await audit('./CLAUDE.md');\nconsole.log(report.score); // 85\n\nconst config = parseConfig('./CLAUDE.md');\nconst report = analyze(config);\n```\n\n## Supported Config Files\n\nAuto-detected in order:\n1. `CLAUDE.md`\n2. `.cursorrules`\n3. `.windsurfrules`\n4. `.github/copilot-instructions.md`\n5. `codex-instructions.md`\n\n## vs \"Just Hoping It Works\"\n\n| | Without promptbench | With promptbench |\n|---|---|---|\n| Config quality | Unknown | Scored 0-100 |\n| Contradictions | Found after wasting tokens | Caught instantly |\n| Vague rules | \"It should work...\" | Flagged with suggestions |\n| Missing sections | Noticed weeks later | Detected immediately |\n| CI enforcement | None | Exit code on failure |\n\n## License\n\nMIT\n","readmeFilename":"README.md"}