{"_id":"@agilelab/agent-skills-eval","_rev":"2-67ce86058735cfaf2c21bad15377f2e5","name":"@agilelab/agent-skills-eval","dist-tags":{"latest":"0.2.1"},"versions":{"0.2.0":{"name":"@agilelab/agent-skills-eval","version":"0.2.0","keywords":["agent-skills-eval","agentskills.io","agentskills","agent-skills","agent-skills-sdk","ai-skills","skills-eval","skills-evaluation","skill-evaluator","agent-eval","agent-evals","ai-eval","ai-evals","llm-eval","llm-evals","llm-evaluation","model-evaluation","model-benchmark","evals","evaluation","llm","ai","agent","benchmark","openai","openai-compatible","yaml","jsonl","html-report","typescript","cli"],"author":{"url":"https://github.com/darkrishabh","name":"Rishabh Mehan"},"license":"MIT","_id":"@agilelab/agent-skills-eval@0.2.0","maintainers":[{"name":"tmnd91","email":"antonio.murgia@agilelab.it"}],"contributors":[{"url":"https://github.com/agilelab-dev","name":"agilelab"},{"url":"https://github.com/tmnd1991","name":"tmnd1991"}],"homepage":"https://agilelab-dev.github.io/agent-skills-eval/","bugs":{"url":"https://github.com/agilelab-dev/agent-skills-eval/issues"},"bin":{"agent-skills-eval":"dist/cli.js"},"dist":{"shasum":"1ff549636f7f015e4aa7c14773fb9dd5eca24791","tarball":"https://registry.npmjs.org/@agilelab/agent-skills-eval/-/agent-skills-eval-0.2.0.tgz","fileCount":105,"integrity":"sha512-MPPdNsZGdqnzIIlBXiyeJ09Nxbbw0nZQsW3soHHn3tkBHlYiieTAUiL45d5RUqluTm8Yz5zQJ+cnxSHxEkl/kA==","signatures":[{"sig":"MEUCIQDEtjKlAzfhE4af6uWtbFCzmMprDo7MhBE0ZRR8BKmw/AIgelOqiXdRNjC7AgLj6SgndNHUMQSL0QFOYUBJIUfPzuk=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":508513},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","engines":{"node":">=18"},"exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"},"./config":{"types":"./dist/config.d.ts","import":"./dist/config.js"},"./opencode":{"types":"./dist/opencode-provider.d.ts","import":"./dist/opencode-provider.js"},"./provider":{"types":"./dist/provider.d.ts","import":"./dist/provider.js"},"./reporters":{"types":"./dist/jsonl-reporter.d.ts","import":"./dist/jsonl-reporter.js"},"./claude-code":{"types":"./dist/claude-code-provider.d.ts","import":"./dist/claude-code-provider.js"},"./openai-compatible":{"types":"./dist/openai-compatible-provider.d.ts","import":"./dist/openai-compatible-provider.js"}},"funding":{"url":"https://github.com/sponsors/darkrishabh","type":"github"},"gitHead":"feb3f6b3fe5fbeb7dd38404a6e7dcf32b1ead13c","scripts":{"test":"npm run build && node --test test/*.test.mjs","build":"tsc -p tsconfig.json","clean":"rm -rf dist","prepack":"npm run build","typecheck":"tsc -p tsconfig.json --noEmit","prepublishOnly":"npm test"},"_npmUser":{"name":"tmnd91","email":"antonio.murgia@agilelab.it"},"repository":{"url":"git+https://github.com/agilelab-dev/agent-skills-eval.git","type":"git"},"_npmVersion":"11.12.0","description":"TypeScript SDK and CLI for evaluating agentskills.io-style AI agent skills with LLM judges, baseline comparison, YAML config, JSONL logs, and HTML reports.","directories":{},"sideEffects":false,"_nodeVersion":"24.12.0","dependencies":{"js-yaml":"^4.1.1","commander":"^12.1.0","@opencode-ai/sdk":"^1.17.15"},"publishConfig":{"access":"public"},"_hasShrinkwrap":false,"packageManager":"npm@10.9.2","devDependencies":{"typescript":"^6.0.3","@types/node":"^25.6.0","@types/js-yaml":"^4.0.9"},"_npmOperationalInternal":{"tmp":"tmp/agent-skills-eval_0.2.0_1783945530318_0.5893482019247178","host":"s3://npm-registry-packages-npm-production"}},"0.2.1":{"name":"@agilelab/agent-skills-eval","version":"0.2.1","description":"TypeScript SDK and CLI for evaluating agentskills.io-style AI agent skills with LLM judges, baseline comparison, YAML config, JSONL logs, and HTML reports.","type":"module","author":{"name":"Rishabh Mehan","url":"https://github.com/darkrishabh"},"contributors":[{"name":"agilelab","url":"https://github.com/agile-lab-dev"},{"name":"tmnd1991","url":"https://github.com/tmnd1991"}],"sideEffects":false,"bin":{"agent-skills-eval":"dist/cli.js"},"main":"./dist/index.js","types":"./dist/index.d.ts","exports":{".":{"import":"./dist/index.js","types":"./dist/index.d.ts"},"./provider":{"import":"./dist/provider.js","types":"./dist/provider.d.ts"},"./openai-compatible":{"import":"./dist/openai-compatible-provider.js","types":"./dist/openai-compatible-provider.d.ts"},"./opencode":{"import":"./dist/opencode-provider.js","types":"./dist/opencode-provider.d.ts"},"./claude-code":{"import":"./dist/claude-code-provider.js","types":"./dist/claude-code-provider.d.ts"},"./config":{"import":"./dist/config.js","types":"./dist/config.d.ts"},"./reporters":{"import":"./dist/jsonl-reporter.js","types":"./dist/jsonl-reporter.d.ts"}},"scripts":{"build":"tsc -p tsconfig.json","typecheck":"tsc -p tsconfig.json --noEmit","test":"npm run build && node --test test/*.test.mjs","clean":"rm -rf dist","prepack":"npm run build","prepublishOnly":"npm test"},"keywords":["agent-skills-eval","agentskills.io","agentskills","agent-skills","agent-skills-sdk","ai-skills","skills-eval","skills-evaluation","skill-evaluator","agent-eval","agent-evals","ai-eval","ai-evals","llm-eval","llm-evals","llm-evaluation","model-evaluation","model-benchmark","evals","evaluation","llm","ai","agent","benchmark","openai","openai-compatible","yaml","jsonl","html-report","typescript","cli"],"license":"MIT","repository":{"type":"git","url":"git+https://github.com/agile-lab-dev/agent-skills-eval.git"},"bugs":{"url":"https://github.com/agile-lab-dev/agent-skills-eval/issues"},"homepage":"https://agile-lab-dev.github.io/agent-skills-eval/","funding":{"type":"github","url":"https://github.com/sponsors/darkrishabh"},"publishConfig":{"access":"public"},"dependencies":{"@opencode-ai/sdk":"^1.17.15","commander":"^12.1.0","js-yaml":"^4.1.1"},"devDependencies":{"@types/js-yaml":"^4.0.9","@types/node":"^25.6.0","typescript":"^6.0.3"},"engines":{"node":">=18"},"packageManager":"npm@10.9.2","gitHead":"ad6c3d656ccc7475e6a37460b650c94193736cc3","_id":"@agilelab/agent-skills-eval@0.2.1","_nodeVersion":"24.18.0","_npmVersion":"11.16.0","dist":{"integrity":"sha512-NJNHyDNt00F429RAGLWdr/Dj5F7/+/QGr3x99Q1o196aMKtM0CSesr5Xt093U56VwXwQA3sStjhsM60aLOn9Bw==","shasum":"371db287477b0608a3e7f97c442ac66190aa3576","tarball":"https://registry.npmjs.org/@agilelab/agent-skills-eval/-/agent-skills-eval-0.2.1.tgz","fileCount":103,"unpackedSize":505862,"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@agilelab%2fagent-skills-eval@0.2.1","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQDMswx4xzOyVAdldIapqXehEKa4x2yaSbLaooYlKOWDYAIhAKuHJEDFi/9o5zj7miqJr3f2VkKP4xVGdJy4HirZIPC0"}]},"_npmUser":{"name":"GitHub Actions","email":"npm-oidc-no-reply@github.com","trustedPublisher":{"id":"github","oidcConfigId":"oidc:a3ce3071-f385-4693-a9d7-262b53add8a0"}},"directories":{},"maintainers":[{"name":"tmnd91","email":"antonio.murgia@agilelab.it"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/agent-skills-eval_0.2.1_1783949774661_0.7891730230577978"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-13T12:25:30.209Z","modified":"2026-07-13T13:36:15.114Z","0.2.0":"2026-07-13T12:25:30.472Z","0.2.1":"2026-07-13T13:36:14.799Z"},"bugs":{"url":"https://github.com/agile-lab-dev/agent-skills-eval/issues"},"author":{"name":"Rishabh Mehan","url":"https://github.com/darkrishabh"},"license":"MIT","homepage":"https://agile-lab-dev.github.io/agent-skills-eval/","keywords":["agent-skills-eval","agentskills.io","agentskills","agent-skills","agent-skills-sdk","ai-skills","skills-eval","skills-evaluation","skill-evaluator","agent-eval","agent-evals","ai-eval","ai-evals","llm-eval","llm-evals","llm-evaluation","model-evaluation","model-benchmark","evals","evaluation","llm","ai","agent","benchmark","openai","openai-compatible","yaml","jsonl","html-report","typescript","cli"],"repository":{"type":"git","url":"git+https://github.com/agile-lab-dev/agent-skills-eval.git"},"description":"TypeScript SDK and CLI for evaluating agentskills.io-style AI agent skills with LLM judges, baseline comparison, YAML config, JSONL logs, and HTML reports.","contributors":[{"name":"agilelab","url":"https://github.com/agile-lab-dev"},{"name":"tmnd1991","url":"https://github.com/tmnd1991"}],"maintainers":[{"name":"tmnd91","email":"antonio.murgia@agilelab.it"}],"readme":"<div align=\"center\">\n\n<img src=\"https://github.com/user-attachments/assets/094b8e11-e19e-4c96-ae82-ba701cfcf7e3\" alt=\"agent-skills-eval — a test runner for Agent Skills\" width=\"100%\" />\n\n<br />\n\n# agent-skills-eval\n\n**A test runner for [Agent Skills](https://agentskills.io).**\n\nWrite a `SKILL.md`, drop in some evals, and find out — empirically — whether your skill actually makes the model better at the task.\n\n[![npm version](https://img.shields.io/npm/v/%40agilelab%2Fagent-skills-eval.svg?style=flat-square&logo=npm&label=npm)](https://www.npmjs.com/package/@agilelab/agent-skills-eval)\n[![CI](https://img.shields.io/github/actions/workflow/status/agile-lab-dev/agent-skills-eval/ci.yml?style=flat-square&logo=github&label=ci)](https://github.com/agile-lab-dev/agent-skills-eval/actions/workflows/ci.yml)\n[![license: MIT](https://img.shields.io/badge/license-MIT-green?style=flat-square)](LICENSE)\n[![node](https://img.shields.io/node/v/%40agilelab%2Fagent-skills-eval.svg?style=flat-square&logo=nodedotjs&logoColor=white)](package.json)\n[![docs](https://img.shields.io/badge/docs-GitHub%20Pages-0f766e?style=flat-square)](https://agile-lab-dev.github.io/agent-skills-eval/)\n[![TypeScript](https://img.shields.io/badge/TypeScript-3178C6?style=flat-square&logo=typescript&logoColor=white)](https://www.typescriptlang.org/)\n\n[Documentation](https://agile-lab-dev.github.io/agent-skills-eval/) · [Quickstart](#quickstart) · [SDK](#sdk) · [agentskills.io](https://agentskills.io)\n\n</div>\n\n---\n\n> **Fork notice:** this is agile-lab-dev's fork of [darkrishabh/agent-skills-eval](https://github.com/darkrishabh/agent-skills-eval). See [`LICENSE`](LICENSE) for original authorship.\n\n---\n\n## Why this exists\n\n[Agent Skills](https://agentskills.io) — the open standard from Anthropic for giving agents domain knowledge — make it easy to ship a `SKILL.md` and assume your agent is now better at the task. The hard part is *proving* it.\n\n`agent-skills-eval` is the missing piece. It runs your skill against the same prompts twice — once `with_skill` loaded into context, once `without_skill` (baseline) — has a judge model grade both outputs, and gives you a side-by-side report. If the skill doesn't make a measurable difference, you'll see it. If it does, you have receipts.\n\nIt's the test framework for the Agent Skills ecosystem, separated from any specific agent runtime so it works wherever your skills do.\n\n## Quickstart\n\n```bash\nnpx @agilelab/agent-skills-eval ./skills \\\n  --target gpt-4o-mini \\\n  --judge gpt-4o-mini \\\n  --baseline \\\n  --strict\n```\n\nThat's it. Point it at a folder of skills, give it a target model and a judge model, and it produces a workspace with full artifacts and a static HTML report.\n\n```text\nagent-skills-workspace/\n└── iteration-1/\n    ├── meta.json            # run metadata\n    ├── benchmark.json       # rolled-up pass/fail per skill\n    ├── eval-basic/\n    │   ├── with_skill/      # output, timing, judge grading\n    │   └── without_skill/   # ↑ same, with the skill stripped\n    └── report/\n        └── index.html       # the visual report\n```\n\nOpen `iteration-1/report/index.html` and you have a real, evidence-backed answer to \"is my skill working?\"\n\n## What you get\n\n|  |  |\n|---|---|\n| **`with_skill` vs `without_skill`** | Every eval runs both ways so you can see the actual lift from the skill — or its absence. |\n| **Judge-graded outputs** | Use any chat model as a judge. Pass/fail with cited assertions, not vibes. |\n| **TypeScript SDK + CLI** | One-liner CLI for CI, full SDK for custom pipelines, custom providers, and dashboards. |\n| **OpenAI-compatible by default** | Works out of the box with OpenAI, Together, Groq, Anthropic via OpenAI-compat layers, local Llama servers — anything that speaks the OpenAI chat API. |\n| **Pluggable execution: API or agentic CLI** | Run target/judge through any OpenAI-compatible API, or drive the real [`opencode`](https://opencode.ai) CLI via `--run-mode opencode`, or the [`claude`](https://docs.claude.com/en/docs/claude-code) CLI's batch mode via `--run-mode claude-code` — either way you get the CLI's own tool use, permissions, and agent config. |\n| **Tool-call assertions** | Deterministic checks for agents that call tools, not just generate text. |\n| **Portable artifacts** | JSON + JSONL all the way down. Run today, diff tomorrow. Plug into your own dashboard. |\n| **Static HTML reports** | A drop-in report site you can publish anywhere — no infrastructure. |\n| **Fully spec-compliant** | Implements the full [agentskills.io specification](https://agentskills.io/specification): `SKILL.md` validation, `evals/evals.json`, official `iteration-N` artifact layout, frontmatter rules. |\n\n## Install\n\n```bash\nnpm install @agilelab/agent-skills-eval\n```\n\nOr run directly without installing:\n\n```bash\nnpx @agilelab/agent-skills-eval --help\n```\n\n## How it works\n\nThe mental model is straightforward. For every eval defined in your skill:\n\n```\n                ┌─────────────────────────────┐\n                │       same prompt           │\n                └───────────────┬─────────────┘\n                                │\n                ┌───────────────┴─────────────┐\n                ▼                             ▼\n        ┌──────────────┐              ┌──────────────┐\n        │ with_skill   │              │without_skill │\n        │ SKILL.md in  │              │ baseline,    │\n        │ context      │              │ no skill     │\n        └──────┬───────┘              └──────┬───────┘\n               │                             │\n               ▼                             ▼\n          target model                  target model\n               │                             │\n               ▼                             ▼\n            output                        output\n               │                             │\n               └──────────┬──────────────────┘\n                          ▼\n                   ┌─────────────┐\n                   │  judge      │  scores both against\n                   │  model      │  the same assertions\n                   └──────┬──────┘\n                          ▼\n                  pass / fail per side\n```\n\nThe judge sees the eval's `expected_output` and `assertions` and grades each side independently. The `--baseline` flag is what enables the comparison; without it you only get the `with_skill` run.\n\nThis is the *logical* flow regardless of run mode — `--run-mode opencode`/`--run-mode claude-code` (below) change *how* the target/judge models are invoked (a subprocess instead of an HTTP call), not this flow.\n\n## YAML config\n\nFor anything beyond a quick command, drop a config file at the root of your project:\n\n```yaml\n# agent-skills-eval.yaml\nroot: ./skills\nworkspace: ./agent-skills-workspace\nbaseline: true\ntarget: gpt-4o-mini\njudge: gpt-4o-mini\nbaseUrl: https://api.openai.com/v1\napiKeyEnv: OPENAI_API_KEY\napi:\n  timeoutMs: 120000       # default; per-attempt request timeout for the default \"api\" run mode\n  judgeTimeoutMs: 180000  # defaults to api.timeoutMs\ninclude:\n  - \"skills/**\"\nexclude:\n  - \"**/draft-*\"\nconcurrency: 4\nlayout: iteration\nstrict: true\nreport:\n  enabled: true\n  title: Agent Skills Report\nlogging:\n  format: pretty   # pretty | jsonl | silent\n  verbose: false\n  color: auto\ntargetParams:\n  temperature: 0\njudgeParams:\n  temperature: 0\n```\n\n```bash\nOPENAI_API_KEY=... npx @agilelab/agent-skills-eval --config agent-skills-eval.yaml\n```\n\nCLI flags always override config values.\n\n## SDK\n\nFor programmatic use — CI pipelines, custom dashboards, multi-skill rollups — drive the evaluator from TypeScript:\n\n```ts\nimport {\n  OpenAICompatibleProvider,\n  consoleReporter,\n  evaluateSkills,\n} from \"@agilelab/agent-skills-eval\";\n\nconst provider = new OpenAICompatibleProvider({\n  baseUrl: \"https://api.openai.com/v1\",\n  apiKey: process.env.OPENAI_API_KEY!,\n  model: \"gpt-4o-mini\",\n  providerName: \"openai\",\n});\n\nconst result = await evaluateSkills({\n  root: \"./skills\",\n  workspace: \"./agent-skills-workspace\",\n  baseline: true,\n  concurrency: 4,\n  workspaceLayout: \"iteration\",\n  strict: true,\n  target: { model: provider.model, provider },\n  judge: { model: provider.model, provider },\n  onEvent: consoleReporter(),\n});\n\nconsole.log(result);\n```\n\nStream events to a file as JSONL for downstream analysis:\n\n```ts\nimport { jsonlReporter } from \"@agilelab/agent-skills-eval\";\n\nconst reporter = jsonlReporter({ file: \"./events.jsonl\" });\n\nawait evaluateSkills({ /* ... */ onEvent: reporter.onEvent });\nawait reporter.close();\n```\n\nLoad YAML config programmatically:\n\n```ts\nimport { loadConfigFile } from \"@agilelab/agent-skills-eval\";\n\nconst config = loadConfigFile(\"./agent-skills-eval.yaml\");\n```\n\n**Signal handling in library mode.** The CLI calls `installSignalHandlers()` for you, so `Ctrl+C`/`SIGTERM` always tears down any in-flight `opencode serve`/`claude` subprocess before the process exits. If you construct `OpencodeProvider`/`ClaudeCodeProvider` directly instead of going through the CLI, call `installSignalHandlers()` yourself once at startup — the providers register their subprocess cleanup with the registry either way, but nothing installs the OS-level `SIGINT`/`SIGTERM` listener unless you (or the CLI) do:\n\n```ts\nimport { installSignalHandlers } from \"@agilelab/agent-skills-eval\";\n\ninstallSignalHandlers();\n```\n\n## Custom providers\n\nBring any backend by implementing the `Provider` interface — five fields, one method:\n\n```ts\nimport type { Provider, ProviderResult } from \"@agilelab/agent-skills-eval\";\n\nexport const provider: Provider = {\n  name: \"my-provider\",\n  model: \"my-model\",\n  async complete(prompt: string): Promise<ProviderResult> {\n    return {\n      provider: \"my-provider\",\n      model: \"my-model\",\n      output: \"model output\",\n      latencyMs: 0,\n      inputTokens: 0,\n      outputTokens: 0,\n      costUsd: 0,\n    };\n  },\n};\n```\n\nUseful for: local model servers (Ollama, vLLM, llama.cpp), proprietary internal APIs, mock providers in unit tests, or routing layers in front of multiple providers.\n\n## opencode run mode\n\nInstead of calling an OpenAI-compatible API directly, you can route target/judge calls through [`opencode`](https://opencode.ai) via `@opencode-ai/sdk` — useful if you already manage model access/credentials through opencode and don't want to supply a separate `--base-url`/API key to this tool.\n\n```bash\nnpx @agilelab/agent-skills-eval ./skills \\\n  --run-mode opencode \\\n  --target anthropic/claude-sonnet-5 \\\n  --opencode-agent build \\\n  --opencode-timeout 300000\n```\n\n| Flag | Description |\n|---|---|\n| `--run-mode <api\\|opencode>` | Default `api`. Set to `opencode` to talk to an opencode server instead of calling an HTTP API. |\n| `--target` / `--judge` | In opencode mode, must be `provider/model` form (e.g. `anthropic/claude-sonnet-5`), matching opencode's own model id syntax. |\n| `--opencode-agent <name>` | Which opencode agent handles the session (e.g. `build`). |\n| `--opencode-dir <path>` | Working directory opencode runs in. Default: current directory. |\n| `--opencode-auto` / `--no-opencode-auto` | Auto-approve opencode's own permission prompts (edit/bash/webfetch/etc). **Dangerous, off by default** — see caveat below. `--no-opencode-auto` always overrides `opencode.auto: true` set in a config file. |\n| `--opencode-timeout <ms>` | Hard deadline per call — covers the initial prompt, waiting on any delegated subagent work, and follow-up continuations. Default `300000` (5 minutes). |\n| `--opencode-judge-timeout <ms>` | Hard deadline for judge/grader calls specifically. Defaults to `--opencode-timeout`. A judge that reads a full transcript plus output files is often slower than the run it grades — set this higher if judge calls are timing out. |\n\nEquivalent YAML:\n\n```yaml\nrunMode: opencode\nopencode:\n  agent: build\n  auto: false\n  dir: ./workspace-scratch\n  timeoutMs: 300000\n  judgeTimeoutMs: 600000\n  # baseUrl: http://127.0.0.1:4096  # talk to an already-running `opencode serve` instead of spawning one\n```\n\n**Skills are loaded natively, not injected into the prompt.** In every other run mode, `with_skill` works by wrapping the skill in an XML block and prepending it to the prompt. In opencode mode, that would never happen in real-world use — opencode discovers skills on disk and loads them itself via its own `skill` tool (see [opencode's Agent Skills docs](https://opencode.ai/docs/skills/)). So instead, before each call this provider symlinks every entry of the skill's directory into `<opencode.dir>/.opencode/skills/<name>/` for `with_skill` runs, and removes that symlink tree for `without_skill` runs — the agent decides for itself whether to call `skill({ name })`. The skill's `evals/` folder (which holds the answer key) is never linked. No `system` message is sent in this mode; the model gets the bare eval prompt either way.\n\n**How a call works.** Each `complete()` call spawns a fresh, private `opencode serve` process (via `createOpencodeServer`, on an auto-assigned port so concurrent target/judge calls don't collide), creates one session, and sends the prompt. opencode's own subagent delegation (the `delegate`/`task` tool — see [opencode's Agents docs](https://opencode.ai/docs/agents/)) is asynchronous: the model's turn ends the moment it dispatches a delegation, and the delegate's result is delivered later as a message appended to the same session, once it finishes. A single request/response can't wait for that, so this provider keeps its private server alive after the initial prompt resolves and polls the session (and any child sessions opencode created for delegated work) until they go idle. It then sends a bounded number of follow-up \"Continue.\" prompts (`maxContinuations`, capped internally) if either: the session's last message is still an unanswered delegation notification rather than the model's own reply, or the model's own reply is a premature \"waiting on delegation\" stub — dispatched a delegation but never read its result back with `delegation_read` — that would otherwise look structurally identical to a real final answer. All of this is bounded by `--opencode-timeout` as the overall deadline. The server is always torn down afterward, even on error, timeout, or a `SIGINT`/`SIGTERM` (Ctrl+C, CI job cancellation) received while the call is in flight — the CLI installs a signal handler that waits for every active server to exit before the process itself exits. Its combined stdout/stderr is captured for the whole run and written to `outputs/opencode-serve.log` alongside the other run artifacts, so a crashed or misbehaving server is diagnosable after the fact.\n\n**Caveats:**\n\n- **Token/cost numbers aren't comparable to API mode.** opencode's own system prompt and tool schemas add fixed overhead to every call (observed ~8,400 input tokens for a trivial one-word prompt), so `inputTokens`/`outputTokens`/`costUsd` from opencode-mode runs are not apples-to-apples with the same model called via `OpenAICompatibleProvider`.\n- **`--opencode-auto` is dangerous.** It lets the target/judge model run opencode's own bash/file-edit tools completely unattended for the duration of the call. It's off by default. Even without it, a stuck interactive permission prompt has no TTY to answer in a non-interactive eval run — that's what `--opencode-timeout` guards against; it always applies, whether or not `--opencode-auto` is set.\n- **No custom binary path.** Unlike older versions of this provider, there's no way to point at a non-`PATH` opencode binary — `@opencode-ai/sdk` always spawns `opencode` resolved via `PATH`. Point `PATH` at the right install, or use `opencode.baseUrl` to talk to a server you started yourself.\n- **`tool_assertions` see opencode's own internal tools, not a caller-supplied schema.** This provider reports every tool call opencode's agent actually made during the run (`bash`, file edits, `task` delegation, `skill`, etc.) as `ProviderResult.toolCalls`, written to `tool_calls.json` and shown in the HTML report — including, for `with_skill` runs, whether the `skill` tool was called with this eval's skill name (surfaced as a \"skill picked up\" badge). `tool_assertions` grade against that same list, so e.g. `{\"type\": \"tool-called\", \"name\": \"skill\"}` works under `--run-mode opencode`, but there's no `tools`/`tool_choice` schema to pass *in* — the model always has whatever tools opencode itself exposes, never a caller-defined set.\n- **`targetParams`/`judgeParams` are not sent.** This provider routes calls through a subprocess CLI, not a chat-completions API, so there is no channel to carry per-call inference params — `temperature`, `top_p`, etc. are silently ignored, including the reproducibility pattern shown in the [YAML config example](#yaml-config) (`temperature: 0`). Use the CLI's own config (e.g. `opencode.json`/Claude Code settings) if you need to control sampling.\n- **Shared working directory.** All calls from one CLI invocation share a single `--opencode-dir`. Evals of the *same* skill (including `with_skill`/`without_skill` for the same skill) are automatically serialized regardless of `--concurrency`, so the on-disk skill symlink one call installs/removes can't race another call's `.opencode/skills/<name>/` lookup. Evals of *different* skills still run concurrently and still share the working tree — a skill that has the model write files, or relies on git state in `--opencode-dir`, can still race a different skill's concurrent eval.\n\n## claude-code run mode\n\nInstead of calling an OpenAI-compatible API directly, you can route target/judge calls through the [`claude`](https://docs.claude.com/en/docs/claude-code) CLI's non-interactive batch mode (`claude -p`) — useful if you already manage model access/credentials through a Claude Code login or subscription and don't want to supply a separate `--base-url`/API key to this tool.\n\n```bash\nnpx @agilelab/agent-skills-eval ./skills \\\n  --run-mode claude-code \\\n  --target claude-sonnet-5 \\\n  --claude-code-agent build \\\n  --claude-code-timeout 300000\n```\n\n| Flag | Description |\n|---|---|\n| `--run-mode <api\\|opencode\\|claude-code>` | Default `api`. Set to `claude-code` to spawn the `claude` CLI in batch mode instead of calling an HTTP API. |\n| `--target` / `--judge` | In claude-code mode, any model id or alias `claude --model` accepts (e.g. `claude-sonnet-5`, `sonnet`, `opus`). |\n| `--claude-code-agent <name>` | Forwarded as `claude --agent`. |\n| `--claude-code-dir <path>` | Working directory the `claude` process runs in (cwd) and where skills are installed. Default: current directory. |\n| `--claude-code-auto` / `--no-claude-code-auto` | Pass `--dangerously-skip-permissions`, bypassing every permission prompt unattended. **Dangerous, off by default** — see caveat below. `--no-claude-code-auto` always overrides `claudeCode.auto: true` set in a config file. |\n| `--claude-code-timeout <ms>` | Hard deadline per call. The subprocess is SIGTERM'd (then SIGKILL'd if it doesn't exit) once this elapses. Default `300000` (5 minutes). |\n| `--claude-code-judge-timeout <ms>` | Hard deadline for judge/grader calls specifically. Defaults to `--claude-code-timeout`. A judge that reads a full transcript plus output files is often slower than the run it grades — set this higher if judge calls are timing out. |\n| `--claude-code-binary <path>` | Path to the `claude` executable. Default `\"claude\"` (resolved via `PATH`). |\n| `--claude-code-allowed-tools <tool>` / `--claude-code-disallowed-tools <tool>` | Repeatable. Forwarded to `claude --allowedTools`/`--disallowedTools`. |\n\nEquivalent YAML:\n\n```yaml\nrunMode: claude-code\nclaudeCode:\n  agent: build\n  auto: false\n  dir: ./workspace-scratch\n  timeoutMs: 300000\n  judgeTimeoutMs: 600000\n  # claudeBinary: /opt/homebrew/bin/claude\n  # allowedTools: [Bash, Read]\n  # disallowedTools: [WebFetch]\n```\n\n**Skills are loaded natively, not injected into the prompt.** In every other run mode, `with_skill` works by wrapping the skill in an XML block and prepending it to the prompt. In claude-code mode, that would never happen in real-world use — Claude Code discovers project skills on disk and loads them itself. So instead, before each call this provider symlinks every entry of the skill's directory into `<claudeCode.dir>/.claude/skills/<name>/` for `with_skill` runs, and removes that symlink tree for `without_skill` runs — the agent decides for itself whether to invoke the skill. The skill's `evals/` folder (which holds the answer key) is never linked. No `system` message is sent in this mode; the model gets the bare eval prompt either way.\n\n**How a call works.** Each `complete()` call spawns a fresh `claude -p --output-format stream-json --verbose` subprocess with the prompt piped over stdin (not a shell argument, so there's no shell-injection surface and no `ARG_MAX` risk from long skill content), and waits for it to exit. Unlike opencode's HTTP server, `claude -p` runs the entire agentic turn — including any subagent (`Task` tool) delegation — synchronously within that one process and doesn't return until it's done, so there's no async delegation/polling/continuation machinery to worry about. The buffered NDJSON transcript is parsed once the process closes: `tool_use` blocks from `assistant` events become `ProviderResult.toolCalls`, and the final `result` event supplies the output text, cumulative token usage, cost, and error status for the whole run. A timer enforces `--claude-code-timeout`, escalating from `SIGTERM` to `SIGKILL` if the process doesn't exit promptly. The subprocess is torn down the same way if the CLI itself receives a `SIGINT`/`SIGTERM` (Ctrl+C, CI job cancellation) while the call is in flight — the CLI installs a signal handler that waits for every active `claude` subprocess to exit before the process itself exits.\n\n**Caveats:**\n\n- **Token/cost numbers aren't comparable to API mode.** Claude Code's own system prompt and tool schemas add fixed overhead to every call, so `inputTokens`/`outputTokens`/`costUsd` from claude-code-mode runs are not apples-to-apples with the same model called via `OpenAICompatibleProvider`. `inputTokens` sums `usage.input_tokens`, `usage.cache_creation_input_tokens`, and `usage.cache_read_input_tokens` from the CLI's own accounting; `costUsd` is the CLI's own `total_cost_usd`.\n- **`--claude-code-auto` is dangerous.** It passes `--dangerously-skip-permissions`, letting the target/judge model run every tool completely unattended for the duration of the call. It's off by default. Even without it, a stuck interactive permission prompt has no TTY to answer in a non-interactive eval run — that's what `--claude-code-timeout` guards against; it always applies, whether or not `--claude-code-auto` is set.\n- **Session persistence is disabled.** Every call passes `--no-session-persistence` since each `complete()` is a one-shot, non-resumable run — there's no `--resume`/`--continue` support here, so persisting sessions to disk would just accumulate unbounded, unused transcript files.\n- **`tool_assertions` see whatever tools the CLI actually exposed, not a caller-supplied schema.** This provider reports every `tool_use` block Claude Code's agent actually emitted during the run (`Bash`, file edits, `Task` delegation, `Skill`, etc.) as `ProviderResult.toolCalls`, written to `tool_calls.json` and shown in the HTML report. `tool_assertions` grade against that same list, so e.g. `{\"type\": \"tool-called\", \"name\": \"Skill\"}` works under `--run-mode claude-code`, but there's no `tools`/`tool_choice` schema to pass *in* — use `--claude-code-allowed-tools`/`--claude-code-disallowed-tools` to constrain the tool set instead.\n- **`targetParams`/`judgeParams` are not sent.** This provider routes calls through a subprocess CLI, not a chat-completions API, so there is no channel to carry per-call inference params — `temperature`, `top_p`, etc. are silently ignored, including the reproducibility pattern shown in the [YAML config example](#yaml-config) (`temperature: 0`). Use the CLI's own config (e.g. `opencode.json`/Claude Code settings) if you need to control sampling.\n- **Shared working directory.** All calls from one CLI invocation share a single `--claude-code-dir`. Evals of the *same* skill (including `with_skill`/`without_skill` for the same skill) are automatically serialized regardless of `--concurrency`, so the on-disk skill symlink one call installs/removes can't race another call's `.claude/skills/<name>/` lookup. Evals of *different* skills still run concurrently and still share the working tree — a skill that has the model write files, or relies on git state in `--claude-code-dir`, can still race a different skill's concurrent eval.\n\n## Skill layout\n\nA skill is a folder. The minimum is a `SKILL.md`. Add `evals/evals.json` and you can evaluate it.\n\n```text\nmy-skill/\n├── SKILL.md\n├── references/\n│   └── notes.md\n├── scripts/\n│   └── helper.sh\n└── evals/\n    ├── evals.json\n    └── files/\n        └── input.csv\n```\n\n`SKILL.md`:\n\n```markdown\n---\nname: my-skill\ndescription: Analyze small CSV files.\nlicense: MIT\ncompatibility: Works with text-capable chat models.\n---\n\nWhen given a CSV file, identify the most important trend and cite the\nrelevant rows.\n```\n\n`evals/evals.json`:\n\n```json\n{\n  \"skill_name\": \"my-skill\",\n  \"evals\": [\n    {\n      \"id\": \"basic\",\n      \"name\": \"basic behavior\",\n      \"prompt\": \"Use the attached data to summarize revenue.\",\n      \"files\": [\"evals/files/input.csv\"],\n      \"expected_output\": \"The response identifies the highest revenue month.\",\n      \"assertions\": [\n        \"The output identifies the highest revenue month.\"\n      ]\n    }\n  ]\n}\n```\n\nIf you skip `assertions` but provide `expected_output`, the SDK promotes the expected output into a judge assertion automatically — so a minimal agentskills.io eval file produces meaningful pass/fail grading without extra work.\n\n## CLI options\n\n```bash\nnpx @agilelab/agent-skills-eval [root] \\\n  --config agent-skills-eval.yaml \\\n  --workspace ./agent-skills-workspace \\\n  --baseline \\\n  --target gpt-4o-mini \\\n  --judge gpt-4o-mini \\\n  --base-url https://api.openai.com/v1 \\\n  --api-key-env OPENAI_API_KEY \\\n  --api-timeout 120000 \\\n  --api-judge-timeout 180000 \\\n  --include \"skills/**\" \\\n  --exclude \"**/draft-*\" \\\n  --concurrency 4 \\\n  --layout iteration \\\n  --strict \\\n  --log-format pretty \\\n  --report\n```\n\n**Logging modes**: `pretty` for humans, `jsonl` for machines, `silent` for quiet CI.\n\n**`--api-timeout <ms>` / `--api-judge-timeout <ms>`**: per-attempt request timeout for the default `api` run mode (an `AbortController` aborts the HTTP call once elapsed). Default `120000` (2 minutes) for both; `--api-judge-timeout` defaults to `--api-timeout` if unset. `fetchWithRetry` retries on timeout, so worst-case wall time is roughly `attempts * timeoutMs` plus backoff, not a single `timeoutMs`.\n\nOr run via the [opencode CLI](#opencode-run-mode) instead of an API:\n\n```bash\nnpx @agilelab/agent-skills-eval [root] \\\n  --run-mode opencode \\\n  --target anthropic/claude-sonnet-5 \\\n  --opencode-agent build \\\n  --opencode-timeout 300000\n```\n\nOr via the [`claude` CLI's batch mode](#claude-code-run-mode):\n\n```bash\nnpx @agilelab/agent-skills-eval [root] \\\n  --run-mode claude-code \\\n  --target claude-sonnet-5 \\\n  --claude-code-timeout 300000\n```\n\n## Reports\n\nThe static HTML report is built from disk artifacts and shows everything you'd want for skill iteration:\n\n- Pass rate by skill and by eval\n- Assertion-by-assertion grading evidence with judge reasoning\n- Full target output, side by side for `with_skill` and `without_skill`\n- Side-by-side/stacked layout toggle for comparing runs\n- Prompt and judge prompt details\n- Timing and token usage\n- Tool calls when present\n- Skill-invocation badge (`skill picked up` / `skill not picked up`) on `with_skill` runs when tool-call data is captured\n\nUse `--report-output` (or `report.output` in YAML) to choose where the report lands.\n\nTo re-render a report from artifacts already on disk — after editing `src/report.ts`, for example — without re-running any evals:\n\n```bash\nnpm run build\nnode scripts/render-report.mjs <workspace> [--output <dir>] [--title <title>] [--target <model>] [--judge <model>] [--provider <name>]\n```\n\n`<workspace>` is an `iteration-N` directory (or any directory containing skill subfolders with `meta.json`/`benchmark.json`/`eval-*`). Output defaults to `<workspace>/report`.\n\n## agentskills.io compatibility\n\nImplements the [agentskills.io](https://agentskills.io) specification end to end:\n\n- `SKILL.md` YAML frontmatter — required `name` and `description`, optional `license`, `compatibility`, `metadata`, `allowed-tools`\n- Strict validation: name length, lowercase-hyphenated format, parent-directory match, description length, compatibility length\n- Optional `scripts/`, `references/`, and `assets/` directories — markdown references included in skill context, scripts exposed by manifest\n- `evals/evals.json` schema: `skill_name`, `evals[].id`, `prompt`, `expected_output`, `files`, `assertions`\n- Official artifact layout: `iteration-N/<eval>/<mode>/outputs`, `timing.json`, `grading.json`, `benchmark.json`\n- Baseline comparison via `with_skill` and `without_skill`\n\nBeyond the spec, this SDK adds: per-eval `defaults`, model `params`, tool definitions, deterministic `tool_assertions`, and a flat `workspaceLayout: \"flat\"` for multi-skill dashboards.\n\n## Examples\n\nSee [`examples/basic-skill`](examples/basic-skill) for a complete skill folder, and [`examples/agent-skills-eval.yaml`](examples/agent-skills-eval.yaml) for a reference config.\n\n## Development\n\n```bash\nnpm ci\nnpm test\nnpm pack --dry-run\n```\n\n## Documentation\n\nFull docs live at **[agile-lab-dev.github.io/agent-skills-eval](https://agile-lab-dev.github.io/agent-skills-eval/)** (sources in [`docs/`](docs)). Local preview:\n\n```bash\npython3 -m http.server 8080 --directory docs\n```\n\n## Contributing\n\nIssues, PRs, and skill examples are all welcome. See [CONTRIBUTING.md](CONTRIBUTING.md), [CODE_OF_CONDUCT.md](CODE_OF_CONDUCT.md), and [SECURITY.md](SECURITY.md).\n\n## License\n\nMIT. See [LICENSE](LICENSE).\n\n---\n\n<div align=\"center\">\n\nBuilt for the [Agent Skills](https://agentskills.io) ecosystem.\n\n</div>\n","readmeFilename":"README.md"}