{"_id":"@axiastudio/aioc-regression-judge","name":"@axiastudio/aioc-regression-judge","dist-tags":{"latest":"0.1.0"},"versions":{"0.1.0":{"name":"@axiastudio/aioc-regression-judge","version":"0.1.0","description":"Experimental LLM judge helpers for AIOC run regression suites.","type":"module","main":"./dist/index.js","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"scripts":{"clean":"node -e \"require('fs').rmSync('dist',{recursive:true,force:true})\"","build":"npm run clean && tsc -p tsconfig.build.json","test":"node -e \"require('fs').rmSync('dist-tests',{recursive:true,force:true})\" && tsc -p tsconfig.test.json && node dist-tests/tests/index.js","prepack":"npm run build"},"keywords":["aioc","regression","judge","llm","run-records","evaluation"],"author":{"name":"Tiziano Lattisi - AXIA Studio S.r.l."},"license":"MIT","repository":{"type":"git","url":"git+https://github.com/axiastudio/aioc.git","directory":"packages/aioc-regression-judge"},"bugs":{"url":"https://github.com/axiastudio/aioc/issues"},"homepage":"https://axiastudio.github.io/aioc","publishConfig":{"access":"public"},"peerDependencies":{"@axiastudio/aioc":"^0.2.7"},"_id":"@axiastudio/aioc-regression-judge@0.1.0","gitHead":"2653d82a36ff6eb4018eb7b345cd7792f6dcd393","_nodeVersion":"20.11.1","_npmVersion":"10.2.4","dist":{"integrity":"sha512-L9yfupOiSn8HyxM0KO5ZPIOVs7Z2mcaD8ME9b8w0yLGnUgB9x9OEMebSVAztQBCS0EcAfqeoPUTRN7qGyDzfvw==","shasum":"d211574a3247f83df1d791f2429d3ab8cfd41435","tarball":"https://registry.npmjs.org/@axiastudio/aioc-regression-judge/-/aioc-regression-judge-0.1.0.tgz","fileCount":6,"unpackedSize":32011,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQDrQkWExTlm+SMr9ZHVJu91zJ3kkTmzmgNYGMEde2h3XAIhANY30PqtkQe0yzGl5DRL7P+TndIBwYcSFi2Ih9oM3kNk"}]},"_npmUser":{"name":"tizianolattisi","email":"info@axia.studio"},"directories":{},"maintainers":[{"name":"tizianolattisi","email":"info@axia.studio"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/aioc-regression-judge_0.1.0_1781871426278_0.48152526095203996"},"_hasShrinkwrap":false}},"time":{"created":"2026-06-19T12:17:06.096Z","0.1.0":"2026-06-19T12:17:06.424Z","modified":"2026-06-19T12:17:06.800Z"},"maintainers":[{"name":"tizianolattisi","email":"info@axia.studio"}],"description":"Experimental LLM judge helpers for AIOC run regression suites.","homepage":"https://axiastudio.github.io/aioc","keywords":["aioc","regression","judge","llm","run-records","evaluation"],"repository":{"type":"git","url":"git+https://github.com/axiastudio/aioc.git","directory":"packages/aioc-regression-judge"},"author":{"name":"Tiziano Lattisi - AXIA Studio S.r.l."},"bugs":{"url":"https://github.com/axiastudio/aioc/issues"},"license":"MIT","readme":"# @axiastudio/aioc-regression-judge\n\nExperimental companion package for AIOC run-regression suites.\n\nIt provides a ready-to-wire judge helper without coupling the package to a\nspecific model provider. Applications provide the model invocation function;\nthe package handles bounded judge input projection, prompt construction, and\n`RunJudgeResult` parsing.\n\n## Install\n\n```bash\nnpm install @axiastudio/aioc @axiastudio/aioc-regression-judge\n```\n\n## Usage\n\n```ts\nimport { runRegressionSuite } from \"@axiastudio/aioc\";\nimport { createRunRegressionJudge } from \"@axiastudio/aioc-regression-judge\";\n\nconst judge = createRunRegressionJudge({\n  judgeModel: \"my-judge-model\",\n  generate: async ({ messages }) => {\n    const response = await callYourModel(messages);\n    return response.content;\n  },\n});\n\nconst suite = await runRegressionSuite({\n  suite: {\n    name: \"age-adapted-explanation\",\n    expectation: {\n      intent: \"Adapt the explanation to the learner age range.\",\n    },\n    cases: [{ baseline }],\n  },\n  agent: candidateAgent,\n  mode: \"live\",\n  judge,\n});\n```\n\n## Local Example\n\nFrom the repository root:\n\n```bash\nnpm run example:run-regression-judge\n```\n\nThe example records a baseline harness, reruns the case against an age-adapted\nharness, and uses `createRunRegressionJudge(...)` to evaluate the candidate\noutput with an application-owned OpenAI call. It requires `OPENAI_API_KEY`.\n\n## Projection Model\n\nThe default judge input is bounded. It includes:\n\n- baseline and candidate final outputs;\n- deterministic comparison summary and metrics;\n- expectation metadata;\n- tool call names and output envelope summaries;\n- policy and guardrail decision summaries;\n- prompt hashes and request fingerprints;\n- descriptor metadata, agent ids, and tool ids when descriptors are present.\n\nIt excludes by default:\n\n- `contextSnapshot`;\n- raw prompt text;\n- full message history;\n- raw tool output data;\n- full comparison differences.\n\nUse `inputMode: \"full\"` only when the application explicitly accepts sending full\n`RunRecord` artifacts to the judge model. Prefer `projection` for application\nspecific redaction.\n","readmeFilename":"README.md","_rev":"1-63a9636d45882384339bb8c63ba74390"}