{"_id":"@d-o.s/ai-artifacts-bench","_rev":"3-465d78af20f9e52738ff4732311d2b1c","name":"@d-o.s/ai-artifacts-bench","dist-tags":{"latest":"0.3.0"},"versions":{"0.1.0":{"name":"@d-o.s/ai-artifacts-bench","version":"0.1.0","license":"Apache-2.0","_id":"@d-o.s/ai-artifacts-bench@0.1.0","maintainers":[{"name":"d-o.s","email":"davidolivier.saban@gmail.com"}],"homepage":"https://github.com/davidoliviersaban/open-ai-artifacts#readme","bugs":{"url":"https://github.com/davidoliviersaban/open-ai-artifacts/issues"},"dist":{"shasum":"2b29eacecda5778ff46fbc27af2bbddfe928228c","tarball":"https://registry.npmjs.org/@d-o.s/ai-artifacts-bench/-/ai-artifacts-bench-0.1.0.tgz","fileCount":15,"integrity":"sha512-lFx6rqvWZbU8dv03WaYbrEVFhBCUHgvBvVWR8fSttTYsWYSFNNj0szwnHNBUMvcJ94eEQ+oEBtVeEzbe048JMg==","signatures":[{"sig":"MEUCIBfaFcHhvLUzu2INSk/ixrBubUfn7E5z+tg1O19Ztry6AiEAytw5niQq4getAC5nQrTIsxCyozejElSuaW39KCW6OYk=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":50462},"main":"lib.js","engines":{"node":">=20"},"gitHead":"cea026dee750dca13389584f34dc5e2d67e66364","scripts":{"test":"find . -name '*.test.js' | xargs node --test"},"_npmUser":{"name":"d-o.s","email":"davidolivier.saban@gmail.com"},"repository":{"url":"git+https://github.com/davidoliviersaban/open-ai-artifacts.git","type":"git","directory":"packages/ai-artifacts-bench"},"_npmVersion":"10.8.2","description":"Benchmark framework for A/B testing AI agent configurations.","directories":{},"_nodeVersion":"20.20.2","_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/ai-artifacts-bench_0.1.0_1781602100961_0.5559167686059769","host":"s3://npm-registry-packages-npm-production"}},"0.2.0":{"name":"@d-o.s/ai-artifacts-bench","version":"0.2.0","license":"Apache-2.0","_id":"@d-o.s/ai-artifacts-bench@0.2.0","maintainers":[{"name":"d-o.s","email":"davidolivier.saban@gmail.com"}],"homepage":"https://github.com/davidoliviersaban/open-ai-artifacts#readme","bugs":{"url":"https://github.com/davidoliviersaban/open-ai-artifacts/issues"},"dist":{"shasum":"524223cc8bddbc58924d44c8eb38525af04dd499","tarball":"https://registry.npmjs.org/@d-o.s/ai-artifacts-bench/-/ai-artifacts-bench-0.2.0.tgz","fileCount":15,"integrity":"sha512-gDYf1eLFuKM75pwf6dSqHOp0dmGHGC10l6cNym7/8V8jLDyboFhQ9OsoDrbGnR5+DmBubCA2IK3u/eZmd1MnIA==","signatures":[{"sig":"MEUCIAYqq5Ax9f8E8lS0LVrOlwzQgdlgLr6HQu11jRUNFGgXAiEAmIC4Z7em2se0X94X0NloAmguWBnFepjNfv1wP/iPUGE=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":50462},"main":"lib.js","engines":{"node":">=20"},"gitHead":"a0c9d036ce346ddf6ec517986de94a195d9b3a32","scripts":{"test":"find . -name '*.test.js' | xargs node --test"},"_npmUser":{"name":"d-o.s","email":"davidolivier.saban@gmail.com"},"repository":{"url":"git+https://github.com/davidoliviersaban/open-ai-artifacts.git","type":"git","directory":"packages/ai-artifacts-bench"},"_npmVersion":"10.8.2","description":"Benchmark framework for A/B testing AI agent configurations.","directories":{},"_nodeVersion":"20.20.2","_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/ai-artifacts-bench_0.2.0_1781616153762_0.18671013952355509","host":"s3://npm-registry-packages-npm-production"}},"0.3.0":{"name":"@d-o.s/ai-artifacts-bench","version":"0.3.0","description":"Benchmark framework for A/B testing AI agent configurations.","license":"Apache-2.0","repository":{"type":"git","url":"git+https://github.com/davidoliviersaban/open-ai-artifacts.git","directory":"packages/ai-artifacts-bench"},"main":"lib.js","scripts":{"test":"find . -name '*.test.js' | xargs node --test"},"engines":{"node":">=20"},"gitHead":"99a6ebf98cfc1316a1536eb01309b70e753de4e4","_id":"@d-o.s/ai-artifacts-bench@0.3.0","bugs":{"url":"https://github.com/davidoliviersaban/open-ai-artifacts/issues"},"homepage":"https://github.com/davidoliviersaban/open-ai-artifacts#readme","_nodeVersion":"24.14.0","_npmVersion":"11.9.0","dist":{"integrity":"sha512-ELzxKyDwk6Ur0zQl/eJFwqY/fFi3SkV7pBnWuexsUefwNklvOFtfFBngLFvsbxOy1/nVpE6bo40jvka2jTgLig==","shasum":"dd79836eacf4b4a479cb68a94be4da9107396558","tarball":"https://registry.npmjs.org/@d-o.s/ai-artifacts-bench/-/ai-artifacts-bench-0.3.0.tgz","fileCount":17,"unpackedSize":88554,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQCHEFN8iIq/LVKcpahyh4kAAXcjmHyKlxUA5CGVcSAHQAIhAP2m6v/9h8K6occSSZd9ZPG0ypuwivj+Sjmif4MrkzeS"}]},"_npmUser":{"name":"d-o.s","email":"davidolivier.saban@gmail.com"},"directories":{},"maintainers":[{"name":"d-o.s","email":"davidolivier.saban@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/ai-artifacts-bench_0.3.0_1782291240073_0.8683431534275163"},"_hasShrinkwrap":false}},"time":{"created":"2026-06-16T09:28:20.821Z","modified":"2026-06-24T08:54:00.307Z","0.1.0":"2026-06-16T09:28:21.093Z","0.2.0":"2026-06-16T13:22:33.910Z","0.3.0":"2026-06-24T08:54:00.197Z"},"bugs":{"url":"https://github.com/davidoliviersaban/open-ai-artifacts/issues"},"license":"Apache-2.0","homepage":"https://github.com/davidoliviersaban/open-ai-artifacts#readme","repository":{"type":"git","url":"git+https://github.com/davidoliviersaban/open-ai-artifacts.git","directory":"packages/ai-artifacts-bench"},"description":"Benchmark framework for A/B testing AI agent configurations.","maintainers":[{"name":"d-o.s","email":"davidolivier.saban@gmail.com"}],"readme":"# @d-o.s/ai-artifacts-bench\n\nGeneric benchmark engine for A/B testing AI agent configurations. Provides the pluggable runner, scorer, reporter, and adapter interfaces — project-specific logic (worktree preparation, scenario definitions) lives in the consumer.\n\n## Plugin Interface\n\n### Adapter\n\nAn adapter drives the AI agent. It must export:\n\n```js\nmodule.exports = {\n  run(worktree, prompt, options) → { stdout, stderr, elapsed, exitCode },\n  parseUsage(raw) → { input_tokens, output_tokens, total_tokens, cost_usd, ... } | null,\n}\n```\n\nSee `adapters/claude-code.js` for the reference implementation.\n\n### Config\n\nThe runner accepts a `config` object:\n\n```js\n{\n  challengesDir,   // path to challenges/<id>/challenge.json files\n  variantsDir,     // path to variants/<id>/variant.json files\n  baselineFile,    // path to baseline.json\n  runsDir,         // output directory for run artifacts\n  repoRoot,        // git repo root (for worktree creation)\n  prepare(worktree, variant, challenge, { repoRoot, runDir }),  // optional\n  postRun(worktree, { runDir, variant, challenge, metadata }),  // optional\n  prepareScoringWorktree(worktree, variant),                    // optional (scorer)\n}\n```\n\n## Modules\n\n| Module | Purpose |\n|--------|---------|\n| `lib.js` | Pure scoring math: `avg`, `median`, `computeScore`, `summarizeVariant`, `determineWinner` |\n| `runner.js` | Single run execution with worktree lifecycle |\n| `batch.js` | Matrix builder + concurrent execution |\n| `score.js` | Run scoring with diff application and criteria evaluation |\n| `report.js` | Report generation with baseline grouping |\n| `decision.js` | Deterministic decision synthesis: per-use-case model+config recommendation |\n| `adapters/claude-code.js` | Claude Code CLI adapter |\n\n## Decision Synthesis\n\n`decision.js` turns per-run scores into an actionable recommendation: **for a given\nuse-case category, which `(model, variant)` candidate to pick**. It is fully\ndeterministic — no randomness, no LLM — so the same runs always produce the same verdict.\n\n- **Two axes, never merged:** quality (mean criteria pass rate) and cost (real time /\n  tokens / $). The blended `final_score` is not used for recommendations.\n- **Confidence:** parametric 95% CI (`mean ± 1.96·σ/√n`), no bootstrap. Single runs are\n  flagged `insufficient_data` and discounted so they can't outrank replicated candidates.\n- **Pareto frontier:** dominated candidates (worse on quality *and* cost) are never picked.\n- **Profiles:** `quality` (highest reliable lower-bound), `cost` and `latency` (cheapest /\n  fastest among candidates statistically tied on quality).\n- **Variant sensitivity:** per model, the best/worst variant and the quality spread —\n  shows how much the AI context tweaks each model's result (`config_sensitive` flag).\n\nThe report renders two views to avoid dumping every value: **View A** (which model per\nuse case — each model under its best variant, default lens `cost`) and **View B** (how to\nconfigure each model — the variant spread). Default profile is `cost`.\n\n### Hard deadline vs scored budget\n\nThe hard deadline is a **safety kill switch** (default 900s, `hard_deadline_seconds` in\n`challenge.json` or `--hard-deadline`), decoupled from the scored cost. It is generous so\nlegitimate work finishes, and is never told to the model — a run killed here failed.\nTime-awareness (telling the model its budget) is a separate variant axis, never silently\nfolded into a model comparison.\n\nChallenges declare a `category` in `challenge.json`. An optional LLM pass may translate the\nresulting `decision` JSON into prose for non-experts, but it consumes the verdict — it\nnever computes it. See `docs/adr/016-decision-oriented-benchmark-synthesis.md`.\n\n## Writing Acceptance Criteria\n\nAcceptance criteria are the unit of measurement. They must be written so the bench\nproduces a meaningful signal — not false positives. Follow these rules:\n\n### Every criterion must FAIL on the base commit\n\nA criterion that already passes before the agent touches anything contributes zero\nsignal. It inflates the score equally across all variants and models.\n\nBefore adding a criterion, check it against the challenge's `base_commit`:\n\n```bash\ngit checkout <base_commit>\n<run the criterion command>\n# if it passes → the criterion is useless, rewrite it\n```\n\nIf you cannot make a criterion fail on the base state, it is not testing the change.\nDrop it or replace it with one that does.\n\n**Baseline score threshold: the untouched code must score below 0.25.** Run the full\nscorer on the base commit with no modifications. If the score is above 0.25, too many\ncriteria pass by default — the challenge cannot discriminate between an agent that did\nnothing and one that made things worse. Rewrite criteria until the baseline is under\n0.25. This guarantees that any score above 0.25 reflects real work by the agent.\n\nGuard-rail criteria (build passes, tests pass, exports preserved) will always pass on\nthe base. That is acceptable — they protect against regressions. But they must be a\nminority: if guard rails alone push the baseline above 0.25, add more discriminating\ncriteria until they are diluted below the threshold.\n\n### Test the change, not the state\n\nBad: `grep -r 'async' src/utils/` — passes if any async function already exists.\n\nGood: `node --test src/utils/retry.test.js` — a test file that imports the function\nthe agent must create, exercises its behavior, and asserts the contract. It fails\nbefore implementation and passes after.\n\nThe strongest criteria are **unit tests that the agent must make pass**. They encode\nthe expected behavior precisely and cannot pass by accident.\n\n### Avoid negation traps\n\n`! grep 'deprecated_call()' src/` passes whenever the file is missing OR the\npattern is absent. On an unmodified base, the deprecated call might live in a\ndifferent file than expected, giving a false pass.\n\nIf you need a negation (\"the agent must remove X\"), verify the pattern exists on the\nbase first. If it doesn't, the criterion tests nothing.\n\n### Lint/build/tests: necessary but not sufficient\n\n`npx nx build`, `npx nx lint`, `npx nx test` are valid criteria — they catch\nregressions. But they almost always pass on the base commit too. They only\ndiscriminate when the agent introduces a regression, not when it succeeds.\n\nUse them, but never as the only criteria. They are guard rails, not signal.\n\n### Summary\n\n| Rule | Why |\n|------|-----|\n| Fails on base_commit | Otherwise identical scores across all candidates |\n| Tests behavior, not grep patterns | Grep is fragile and often accidentally passes |\n| Unit test > grep > file existence | Precision of signal, from best to worst |\n| Negation requires base verification | `! grep X` can silently pass when X was never there |\n| Build/lint/test = guard rail only | They pass by default, they don't measure the fix |\n\n## Finding Valid Model IDs\n\nThe `--model` flag passed to the benchmark must match exactly what `claude --model`\naccepts. The format depends on your authentication backend:\n\n| Backend | Format | Example |\n|---------|--------|---------|\n| Anthropic API (direct) | `claude-<family>-<version>` | `claude-sonnet-4-5-20250929` |\n| AWS Bedrock | `us.anthropic.claude-<family>-<version>-v1:0` | `us.anthropic.claude-sonnet-4-5-20250929-v1:0` |\n| Bedrock (short) | `us.anthropic.claude-<family>-<version>` | `us.anthropic.claude-opus-4-8` |\n\nTo discover which IDs are valid for your setup:\n\n```bash\n# Quick test — if this returns a model ID, it works:\nclaude -p --model <candidate-id> --output-format json \"say hi\" 2>&1 | head -3\n\n# AWS Bedrock — list available inference profiles:\naws bedrock list-inference-profiles --region us-east-1 --output json \\\n  | jq '.inferenceProfileSummaries[].inferenceProfileId' | grep claude\n\n# Anthropic API — list models:\ncurl -s https://api.anthropic.com/v1/models -H \"x-api-key: $ANTHROPIC_API_KEY\" \\\n  | jq '.data[].id' | grep claude\n```\n\nA wrong model ID causes a `400 The provided model identifier is invalid` error.\nThe run produces an empty diff and scores 0. If all runs score identically, check\nthe `stdout.json` for API errors before debugging criteria.\n\nStore valid IDs in `baseline.json` to avoid re-discovering them each time:\n\n```json\n{\n  \"id\": \"my-baseline-v1\",\n  \"default_models\": [\n    \"us.anthropic.claude-opus-4-8\",\n    \"us.anthropic.claude-sonnet-4-5-20250929-v1:0\"\n  ]\n}\n```\n\n## Testing\n\n```bash\nnode --test 'packages/ai-artifacts-bench/**/*.test.js'\n# or via Nx\nnpm run test:ai-artifacts-bench\n```\n\n79 tests cover all modules and prove the plugin architecture works end-to-end.\n","readmeFilename":"README.md"}