{"_id":"@almanzor/cli","_rev":"2-a4c4c77395283314a4abce38a5a9c01a","name":"@almanzor/cli","dist-tags":{"latest":"0.1.1"},"versions":{"0.1.0":{"name":"@almanzor/cli","version":"0.1.0","keywords":["llm","evals","cli","ai","testing"],"license":"MIT","_id":"@almanzor/cli@0.1.0","maintainers":[{"name":"victorelexpe","email":"victorelexpe@gmail.com"}],"homepage":"https://github.com/Almanzor-Cloud/almanzor-cli#readme","bugs":{"url":"https://github.com/Almanzor-Cloud/almanzor-cli/issues"},"bin":{"almanzor":"bin/almanzor.js"},"dist":{"shasum":"3eead41bcb25cd7bfa76baf271d5df4e62f74b7e","tarball":"https://registry.npmjs.org/@almanzor/cli/-/cli-0.1.0.tgz","fileCount":6,"integrity":"sha512-y8+ZNodwkLN/AW53dkWkmqa9NNncVwBHpmohxmjlagUcG4USNDEGyYYuSeBzu8NIuhzIVpUYAcwW04qWOgZUzQ==","signatures":[{"sig":"MEYCIQDoDikK3e4md4lmp9kphRIyupUhmKz3NebHZYDU/as/cQIhAPZBpMuNkWZSzcogR/mhumY09pu/Lhxv+pvSmlIa+CLW","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@almanzor%2fcli@0.1.0","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":296518},"main":"./dist/cli.js","type":"module","types":"./dist/cli.d.ts","engines":{"node":">=20"},"exports":{".":{"types":"./dist/cli.d.ts","import":"./dist/cli.js"}},"gitHead":"e048e0940aa8d13fed411f542c81e6a66d937549","scripts":{"dev":"tsup --watch","test":"vitest run","build":"tsup","check":"npm run build && npm test && npm run typecheck","almanzor":"node bin/almanzor.js","typecheck":"tsc --noEmit","test:watch":"vitest","prepublishOnly":"npm run build"},"_npmUser":{"name":"victorelexpe","email":"victorelexpe@gmail.com"},"repository":{"url":"git+https://github.com/Almanzor-Cloud/almanzor-cli.git","type":"git"},"_npmVersion":"10.8.2","description":"OSS toolkit for LLM development — evals, VCR, profiling, and more","directories":{},"_nodeVersion":"20.20.2","dependencies":{"zod":"^3.24.2","yaml":"^2.7.0","commander":"^13.1.0","picocolors":"^1.1.1"},"publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"tsup":"^8.4.0","vitest":"^3.0.9","typescript":"^5.8.2","@types/node":"^22.13.10"},"_npmOperationalInternal":{"tmp":"tmp/cli_0.1.0_1785779047015_0.5565277408683547","host":"s3://npm-registry-packages-npm-production"}},"0.1.1":{"name":"@almanzor/cli","version":"0.1.1","description":"OSS toolkit for LLM development — evals, VCR, profiling, and more","type":"module","bin":{"almanzor":"bin/almanzor.js"},"main":"./dist/cli.js","types":"./dist/cli.d.ts","exports":{".":{"import":"./dist/cli.js","types":"./dist/cli.d.ts"}},"engines":{"node":">=20"},"scripts":{"build":"tsup","dev":"tsup --watch","almanzor":"node bin/almanzor.js","check":"npm run build && npm test && npm run typecheck","test":"vitest run","test:watch":"vitest","typecheck":"tsc --noEmit","prepublishOnly":"npm run build"},"keywords":["llm","evals","cli","ai","testing"],"license":"MIT","repository":{"type":"git","url":"git+https://github.com/Almanzor-Cloud/almanzor-cli.git"},"bugs":{"url":"https://github.com/Almanzor-Cloud/almanzor-cli/issues"},"homepage":"https://github.com/Almanzor-Cloud/almanzor-cli#readme","publishConfig":{"access":"public"},"dependencies":{"commander":"^13.1.0","picocolors":"^1.1.1","yaml":"^2.7.0","zod":"^3.24.2"},"devDependencies":{"@types/node":"^22.13.10","tsup":"^8.4.0","typescript":"^5.8.2","vitest":"^3.0.9"},"_id":"@almanzor/cli@0.1.1","gitHead":"3d59d1d3ef50d4b3b68d39d02f24ebb3df38f641","_nodeVersion":"20.20.2","_npmVersion":"10.8.2","dist":{"integrity":"sha512-0hCcvYIZuhFWUiAva6S+1lBYSb/Zk4I5AWrTfS2bsXM/3c00hGNcXbQt3ZFHxB9smPESPMVa8jGPV7vDtqsD7g==","shasum":"0cf5be9f33ae43f2fa1168ef503b9e2b36d5fd3d","tarball":"https://registry.npmjs.org/@almanzor/cli/-/cli-0.1.1.tgz","fileCount":6,"unpackedSize":297376,"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@almanzor%2fcli@0.1.1","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIBi2hkKwrVCfO1M534Scvul8PTiyKNQmw10YHt5OxPyiAiBYS6dzcxExGVkVG7Mx/HUodoyaq5Qg8YAW3rCG6OyYGw=="}]},"_npmUser":{"name":"victorelexpe","email":"victorelexpe@gmail.com"},"directories":{},"maintainers":[{"name":"victorelexpe","email":"victorelexpe@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/cli_0.1.1_1785780887248_0.9449187169077764"},"_hasShrinkwrap":false}},"time":{"created":"2026-08-03T17:44:06.806Z","modified":"2026-08-03T18:14:47.850Z","0.1.0":"2026-08-03T17:44:07.192Z","0.1.1":"2026-08-03T18:14:47.400Z"},"bugs":{"url":"https://github.com/Almanzor-Cloud/almanzor-cli/issues"},"license":"MIT","homepage":"https://github.com/Almanzor-Cloud/almanzor-cli#readme","keywords":["llm","evals","cli","ai","testing"],"repository":{"type":"git","url":"git+https://github.com/Almanzor-Cloud/almanzor-cli.git"},"description":"OSS toolkit for LLM development — evals, VCR, profiling, and more","maintainers":[{"name":"victorelexpe","email":"victorelexpe@gmail.com"}],"readme":"# Almanzor\n\n![CI](https://github.com/Almanzor-Cloud/almanzor-cli/actions/workflows/ci.yml/badge.svg?branch=develop)\n\nOSS toolkit for LLM development — evals, VCR, profiling, traces, and more. One CLI, consistent conventions, CI-ready outputs.\n\n```bash\nnpx @almanzor/cli@0.1.1 init\nalmanzor doctor\nalmanzor evals validate\nalmanzor evals run\n```\n\n## Installation\n\nPackage: **`@almanzor/cli`** — command in terminal: **`almanzor`**\n\n```bash\n# Run without installing\nnpx @almanzor/cli <command>\n\n# Global install (when stable)\nnpm i -g @almanzor/cli\n```\n\n## Quick start\n\nRecommended flow for a new project:\n\n```bash\nalmanzor init\nalmanzor doctor\nalmanzor evals validate\nalmanzor evals list\nalmanzor evals run\nalmanzor ci --update-baseline    # first time — creates v2 baseline + summary\nalmanzor ci --fail-on-regression # subsequent CI runs\n```\n\n`init` creates:\n\n- `almanzor.yml` — project config\n- `evals/example.yaml` — sample eval definition\n- `fixtures/llm/example.json` — mock LLM response\n- `artifacts/almanzor/` — CI outputs (on first `ci` run)\n\nSee [`examples/basic/`](examples/basic/) for a complete working project — start with [examples/basic/README.md](examples/basic/README.md).\n\n### Local development (from source)\n\nWhen working on the CLI itself, compile TypeScript first (`bin/almanzor.js` loads `dist/`):\n\n```bash\nnpm install && npm run build\nnpm run almanzor -- init --dir ./tmp/my-agent\ncd tmp/my-agent\nalmanzor doctor\nalmanzor evals validate\nalmanzor evals run\n```\n\nFrom the repo root without `cd`, pass `--config` so paths resolve to the project directory (see [Configuration](#configuration)).\n\nSee [Development](#development) and [Local workflow](#local-workflow) for watch mode, `npm link`, and pre-push checks.\n\n## Commands\n\nEach command has a detailed reference in the sections below.\n\n| Command | Status | Description |\n|---------|--------|-------------|\n| `almanzor init` | Available | Bootstrap project structure — see [init](#init) |\n| `almanzor doctor` | Available | Validate project setup (config, dirs, evals) — see [Doctor](#doctor) |\n| `almanzor evals run` | Available | Run evals locally — see [Evals](#evals) |\n| `almanzor evals ci` | Available | CI mode with regression checks — see [Evals](#evals) |\n| `almanzor evals list` | Available | List suites, cases, and fixture status |\n| `almanzor evals validate` | Available | Validate eval YAML and fixtures (no LLM) |\n| `almanzor ci` | Available | Run evals CI + aggregate report — see [CI](#ci) |\n| `almanzor report` | Available | Aggregate tool results to summary — see [Report](#report) |\n| `almanzor vcr record` | Available | Record LLM HTTP via local proxy — see [VCR](#vcr-http-passthrough) |\n| `almanzor vcr replay` | Available | Replay LLM HTTP from cassettes |\n\n### Global flags\n\n- `--config <path>` — override config file location\n- `--format text|json|md` — output format (default: `text`)\n- `-v, --verbose` — verbose output\n- `--no-color` — disable ANSI colors\n\n**Terminal colors:** enabled automatically in interactive terminals (`text` format only). Disable with `--no-color` or `NO_COLOR=1`. `json` and `md` output never includes ANSI codes.\n\n### init\n\nCreate a new Almanzor project in the target directory (default: current working directory).\n\n| Flag | Description |\n|------|-------------|\n| `--dir <path>` | Target directory (default: cwd) |\n| `--force` | Overwrite existing files |\n| `--with-gitignore` | Append Almanzor entries to `.gitignore` |\n\nIn an interactive terminal, `init` shows an ASCII banner and suggests next steps: `doctor`, `evals validate`, `evals run`.\n\n```bash\nalmanzor init --dir ./my-agent --with-gitignore\ncd my-agent\nalmanzor doctor\nalmanzor evals validate\nalmanzor evals list\nalmanzor evals run\n```\n\n## Evals\n\nGit-native eval harness. Define suites in `evals/*.yaml`, pair with mock fixtures in `fixtures/llm/`.\n\n### Eval schema\n\n```yaml\n# evals/greeting.yaml\nname: greeting\ndescription: Basic greeting eval\ncases:\n  - id: basic\n    input:\n      prompt: \"Say hello\"\n      system: \"You are a helpful assistant\"   # optional\n    fixture: greeting-basic                    # → fixtures/llm/greeting-basic.json\n    skip: false                                # optional — exclude from run\n    only: false                                # optional — run only these when set\n    expect:\n      contains: [\"hello\"]\n      notContains: [\"error\"]\n      matches: \"^Hello\"\n      minScore: 1.0                            # min fraction of assertions that must pass\n      equals: \"exact response\"                 # optional — exact string match\n      minLength: 10                            # optional\n      maxLength: 500                           # optional\n\n  - id: json-response\n    input:\n      prompt: \"Return JSON\"\n    fixture: json-fixture\n    expect:\n      json: true                               # response must parse as JSON\n\n  - id: json-match\n    input:\n      prompt: \"Return status\"\n    fixture: status-fixture\n    expect:\n      json: { equals: { ok: true } }           # deep equality after JSON parse\n```\n\n**`minScore`:** Each assertion (`contains`, `matches`, `minLength`, etc.) counts equally. The case passes when `score >= minScore` (default `1.0` = all assertions must pass).\n\n### Fixture format (mock provider)\n\n```json\n{\n  \"response\": \"Hello! How can I help you today?\",\n  \"model\": \"mock/gpt-4\",\n  \"usage\": { \"promptTokens\": 10, \"completionTokens\": 8 }\n}\n```\n\n### Eval commands\n\n```bash\n# Discover and validate (no LLM calls)\nalmanzor evals list\nalmanzor evals list --format json\nalmanzor evals validate\nalmanzor evals validate --lenient   # missing fixtures = warnings\n\n# Mock provider (default) — reads fixtures/llm/<id>.json\nalmanzor evals run\nalmanzor evals run --suite greeting --format json\nalmanzor evals run --case basic\nalmanzor evals run --suite greeting --case basic\n\n# OpenAI-compatible provider — calls real API\nexport OPENAI_API_KEY=sk-...\nalmanzor evals run --provider openai\nalmanzor evals run --provider openai --vcr record\nalmanzor evals run --provider openai --vcr replay\n\n# CI: write reports + check baseline (uses provider from almanzor.yml)\nalmanzor evals ci --update-baseline\nalmanzor evals ci --fail-on-regression\nalmanzor evals ci --provider mock\nalmanzor evals ci --suite greeting --case basic\nalmanzor evals ci --provider openai --vcr replay   # deterministic OpenAI evals\n```\n\n`--filter` is deprecated; use `--suite` instead.\n\n**Case selection:** If any case has `only: true`, only those cases run. Otherwise, cases with `skip: true` are excluded.\n\n**`--vcr` on evals:** Requires `--provider openai`. Routes HTTP through the VCR proxy and writes `vcr-report.json`. With the mock provider, `--vcr` is ignored (warning printed). There is no `--no-artifact` flag on evals today.\n\n#### evals list\n\nInventory suites, cases, and fixture status without calling an LLM. Exit code is always `0`.\n\nExample text output:\n\n```\nEvals: 1 suite(s), 2 case(s), 2/2 fixtures ok\n\nSuite           File                  Cases   Fixtures    Status\ngreeting        evals/greeting.yaml   2       2/2         ok\n```\n\n| Status | Meaning |\n|--------|---------|\n| `ok` | All fixtures exist |\n| `partial` | Some fixtures are missing |\n| `missing` | No fixtures found |\n\n#### evals validate\n\nValidate eval YAML and fixtures without running an LLM.\n\n| Flag | Default | Effect |\n|------|---------|--------|\n| `--strict` | yes | Missing fixture = error |\n| `--lenient` | no | Missing fixture = warning only |\n\n**Checks:** YAML schema, fixture file exists, fixture JSON schema, duplicate suite names (warning), no runnable cases after skip/only (warning).\n\n**Exit codes:** `0` when there are no errors; `1` when errors are found. Warnings alone do not fail in `--lenient` mode.\n\n**Note:** If `evals.provider: openai` in `almanzor.yml`, `evals ci` / `evals run` will call the real API and consume tokens (unless `--vcr replay`).\n\n### OpenAI provider config\n\n```yaml\nevals:\n  provider: openai   # mock | openai\n  openai:\n    baseUrl: https://api.openai.com/v1\n    model: gpt-4o-mini\n    apiKeyEnv: OPENAI_API_KEY\n```\n\nAPI keys are read from the environment only — never stored in YAML or fixtures.\n\n### Baseline and regression (v2)\n\n`evals ci` and `almanzor ci` write reports to `artifacts/almanzor/` and optionally compare against a baseline.\n\n**Artifacts from evals CI:**\n\n- `evals-report.json` — machine-readable run result\n- `evals-report.md` — human-readable run result\n- `evals-baseline.json` — created/updated with `--update-baseline`\n\n`--update-baseline` writes a **v2 baseline** with global, per-suite, and per-case scores. Older v1 baselines (`{ score, timestamp }` only) still work for global regression checks. Run `--update-baseline` once to migrate v1 → v2.\n\n**v2 baseline shape:**\n\n```json\n{\n  \"version\": 2,\n  \"score\": 1,\n  \"timestamp\": \"2026-07-13T12:00:00.000Z\",\n  \"suites\": {\n    \"greeting\": {\n      \"score\": 1,\n      \"cases\": {\n        \"basic\": { \"score\": 1 },\n        \"length-check\": { \"score\": 1 }\n      }\n    }\n  }\n}\n```\n\n**Regression detection:** CI fails when any of the following drop below baseline (when `failOnRegression` is true):\n\n- Global score\n- Per-suite score (v2 baseline)\n- Per-case score (v2 baseline)\n\nGranular regressions appear in `evals-report.json` as `regressions` and in CLI stderr:\n\n```json\n{\n  \"regressions\": [\n    { \"suite\": \"greeting\", \"case\": \"basic\", \"delta\": -0.5 }\n  ]\n}\n```\n\nA suite or case can regress even when the global score is unchanged — update the baseline with `--update-baseline` after intentional changes.\n\n## CI\n\nTop-level CI orchestrator. Runs `evals ci` then `report --strict <tools>`. **Preferred entry point for CI pipelines.**\n\n```bash\nalmanzor ci --update-baseline\nalmanzor ci --fail-on-regression\nalmanzor ci --strict evals,vcr\nalmanzor ci --provider mock --suite greeting\n```\n\n| Flag | Default | Description |\n|------|---------|-------------|\n| `--strict [tools]` | `evals` | Required artifacts for report step (`evals` or `evals,vcr`) |\n| `--update-baseline` | false | Write v2 baseline after evals |\n| `--fail-on-regression` | from config | Fail on global/suite/case regression |\n| `--provider <name>` | from config | `mock` or `openai` |\n| `--suite <name>` | — | Run only matching suite |\n| `--case <id>` | — | Run only matching case |\n| `--filter <name>` | — | Deprecated alias for `--suite` |\n\n**Exit codes:** `1` on eval failure, regression, or missing required artifacts.\n\n**vs `evals ci`:** `almanzor ci` always aggregates a summary (`summary.json` + `summary.md`). Use `evals ci` alone when you only need eval artifacts without the unified report.\n\n## Report\n\nAggregate CI artifacts into a unified summary for dashboards and PR comments.\n\n```bash\nalmanzor report\nalmanzor report --strict\nalmanzor report --strict evals,vcr\n```\n\n| Flag | Default | Description |\n|------|---------|-------------|\n| `--strict [tools]` | off | Fail if listed artifacts are missing (`evals`, `evals,vcr`) |\n\n**Outputs** in `artifacts/almanzor/`:\n\n- `summary.json` — machine-readable status across tools\n- `summary.md` — PR-friendly markdown summary (no ANSI codes)\n\n**Inputs** (optional unless `--strict`):\n\n| Artifact | Source |\n|----------|--------|\n| `evals-report.json` | `evals ci` or `almanzor ci` |\n| `vcr-report.json` | `vcr record\\|replay` or `evals --vcr` |\n\n**Status semantics:**\n\n| Status | Meaning |\n|--------|---------|\n| `pass` | All present tools passed |\n| `fail` | A tool failed, or a `--strict` artifact is missing |\n| `partial` | No evals report and not in strict mode |\n\nTools without artifacts are **omitted** from `summary.json` (unless listed in `--strict`, in which case they appear as `missing` and status is `fail`).\n\n**Terminal output:** `report` uses colors in `text` format (PASS/FAIL/PARTIAL, evals/VCR lines). JSON and markdown outputs are always ANSI-free.\n\n## Doctor\n\nValidate project setup without running LLM calls.\n\n```bash\nalmanzor doctor\nalmanzor doctor --strict\nalmanzor doctor --format json\n```\n\n| Flag | Default | Description |\n|------|---------|-------------|\n| `--strict` | false | Missing config, `evals/`, or `fixtures/` → errors (not warnings) |\n\n**Checks performed:**\n\n| Check | Level (default) | Level (`--strict`) |\n|-------|-----------------|---------------------|\n| `almanzor.yml` valid (Zod) | error on invalid | error on invalid |\n| Config file found | warning if missing | error if missing |\n| `evals/` directory exists | warning | error |\n| `fixtures/llm/` directory exists | warning | error |\n| `artifacts/` directory | info if missing | info if missing |\n| Eval definitions (`validateEvals`) | errors/warnings | errors/warnings |\n| Node.js >= 20 | warning | warning |\n| OpenAI API key when `provider: openai` | warning | warning |\n\n**Exit codes:** `0` when there are no errors. Warnings alone are OK in default mode.\n\n**Text output:** On success, prints \"Doctor checks passed\" and a next-steps footer (`evals run`, `ci --update-baseline`).\n\n## VCR (HTTP passthrough)\n\nRecord and replay OpenAI-compatible HTTP calls while running any command. Almanzor starts a local proxy and sets `OPENAI_BASE_URL` for the child process.\n\n```bash\n# Record (calls real API, saves cassettes)\nexport OPENAI_API_KEY=sk-...\nalmanzor vcr record -- node my-app.js\n\n# Replay (deterministic, no upstream calls for matched requests)\nalmanzor vcr replay -- node my-app.js\n```\n\n| Flag | Description |\n|------|-------------|\n| `--fixtures-dir <path>` | Override `vcr.fixturesDir` from config |\n| `--no-artifact` | Skip writing `vcr-report.json` |\n\nCassettes are stored in `fixtures/llm/vcr/<hash>.json` (hashed by request body). API keys are never persisted in cassettes.\n\nBy default, `vcr record` and `vcr replay` write `artifacts/almanzor/vcr-report.json`. The same artifact is written by `evals run|ci --vcr` (no `--no-artifact` on evals).\n\n```yaml\nvcr:\n  fixturesDir: fixtures/llm\n  upstreamBaseUrl: https://api.openai.com/v1   # optional\n  matchPaths:\n    - /chat/completions\n```\n\n## Configuration\n\nAlmanzor looks for config in this order:\n\n1. `almanzor.yml` (or `almanzor.yaml`) in the project root\n2. `.almanzor/config.yml`\n3. Defaults if no file is found\n\n### `--config` and project root\n\nWhen you pass `--config <path>`, Almanzor resolves the path relative to the current working directory and sets **`projectRoot` to the directory containing the config file** (not the cwd).\n\nRelative paths in config (`evals/`, `fixtures/llm/`, `artifactsDir`) resolve from `projectRoot`. This lets you run commands from the repo root against another project:\n\n```bash\nnpm run almanzor -- evals run --config examples/basic/almanzor.yml\n# projectRoot → examples/basic/\n```\n\nExample `almanzor.yml`:\n\n```yaml\nproject: my-agent\nartifactsDir: artifacts/almanzor\n\nevals:\n  path: evals\n  failOnRegression: true\n  reporters:\n    - md\n    - json\n  baselineFile: evals-baseline.json\n  provider: mock\n\nvcr:\n  fixturesDir: fixtures/llm\n  mode: auto   # not yet used by CLI — use vcr record/replay explicitly\n```\n\n> **Note:** `vcr.mode: auto` is reserved for future use. Today, use `almanzor vcr record` and `almanzor vcr replay` explicitly, or `evals --vcr`.\n\n## Conventions\n\n### Artifacts\n\nAll tool outputs go under `artifacts/almanzor/` (configurable via `artifactsDir`).\n\n| File | Produced by | Purpose |\n|------|-------------|---------|\n| `evals-report.json` | `evals ci`, `almanzor ci` | Machine-readable eval run + regression |\n| `evals-report.md` | `evals ci`, `almanzor ci` | Human-readable eval run |\n| `evals-baseline.json` | `--update-baseline` | v2 baseline (global + suite + case scores) |\n| `vcr-report.json` | `vcr`, `evals --vcr` | VCR stats for report aggregator |\n| `summary.json` | `report`, `almanzor ci` | Unified multi-tool status |\n| `summary.md` | `report`, `almanzor ci` | PR-friendly unified summary |\n\n### Exit codes\n\n| Code | Meaning |\n|------|---------|\n| `0` | Success |\n| `1` | User/config error, failed evals, regression, or missing strict artifact |\n| `2` | Command not implemented |\n| `3` | Internal error |\n\n### Output formats\n\nUse `--format json` or `--format md` for CI and PR comments. Only `text` format includes terminal colors.\n\n## Development\n\n```bash\nnpm install\nnpm run build\nnpm test\nnpm run almanzor -- --help\n```\n\nStart with [examples/basic/README.md](examples/basic/README.md) for a guided walkthrough of evals.\n\n### Local workflow\n\nYou do **not** need to push to git to test the CLI. Iterate locally, then run `npm run check` before pushing (same steps as CI).\n\n**npm scripts:**\n\n| Script | Description |\n|--------|-------------|\n| `npm run build` | Compile `src/` → `dist/` (required before testing) |\n| `npm run dev` | Watch rebuild on save |\n| `npm run almanzor --` | Run local CLI (`node bin/almanzor.js`) |\n| `npm run check` | build + test + typecheck (same as CI) |\n\n**Initial setup:**\n\n```bash\nnpm install && npm run build\n```\n\n**Fast iteration** (two terminals):\n\n```bash\n# Terminal 1 — rebuild on save\nnpm run dev\n\n# Terminal 2 — run the local CLI (from repo root)\nnpm run almanzor -- init --dir /tmp/almanzor-test\nnpm run almanzor -- doctor --config examples/basic/almanzor.yml\nnpm run almanzor -- evals run --config examples/basic/almanzor.yml\nnpm run almanzor -- evals list --config examples/basic/almanzor.yml\n```\n\n**Optional global command** (run from any directory after linking):\n\n```bash\nnpm link          # once, from repo root\ncd examples/basic\nalmanzor evals run\n```\n\n**Pre-push check** (build + test + typecheck):\n\n```bash\nnpm run check\n```\n\n**Advanced:** install in another project without publishing:\n\n```bash\nnpm pack\nnpm install /path/to/almanzor-cli/almanzor-cli-0.1.1.tgz\n```\n\n### Troubleshooting\n\n- **Changes in `src/` not visible** — run `npm run build` or keep `npm run dev` running in another terminal.\n- **No colors in output** — stdout may not be a TTY (CI, pipes). Use a real terminal, or pass `--no-color` / set `NO_COLOR=1`.\n- **`Config file not found` with `--config`** — the path is resolved relative to your current working directory.\n- **Regression on one case but global score unchanged** — v2 baseline flags per-suite/case drops. Check `regressions` in `evals-report.json` or re-run with `--update-baseline` after intentional changes.\n- **`almanzor ci` fails on strict vcr** — run `evals ci --vcr replay` or `vcr replay` first to produce `vcr-report.json`, or use `--strict evals` only.\n- **`--vcr` ignored on evals** — requires `--provider openai`. Mock provider uses fixtures directly.\n\n### Project layout\n\n```\nsrc/cli.ts          → entry + commander\nsrc/core/           → config, paths, output, diagnostics, exit codes\nsrc/plugins/        → init, evals, vcr, report, doctor, ci\n```\n\nSee [docs/PLUGIN.md](docs/PLUGIN.md) for the plugin + artifact convention when adding new tools.\n\n### Adding a plugin\n\n1. Create `src/plugins/<name>/index.ts` with `export const registerX = (program) => ...`\n2. Register it in `src/plugins/index.ts` → `registerAllPlugins()`\n3. Add tests under `tests/`\n4. Update this README commands table and artifact reference\n\n### Run the example\n\nFrom repo root (uses `--config` so paths resolve to `examples/basic/`):\n\n```bash\nnpm run almanzor -- evals run --config examples/basic/almanzor.yml\n```\n\nOr from the example directory:\n\n```bash\ncd examples/basic\nnode ../../bin/almanzor.js evals run\n```\n\n## GitHub Action\n\nRun evals in CI and post a PR comment with the aggregated report. The action runs `almanzor ci` (evals + report in one step).\n\n```yaml\npermissions:\n  pull-requests: write\n  contents: read\n\njobs:\n  evals:\n    runs-on: ubuntu-latest\n    steps:\n      - uses: actions/checkout@v4\n      - uses: Almanzor-Cloud/almanzor-cli/.github/actions/evals@v0.1.1\n        with:\n          working-directory: .\n          provider: mock\n          fail-on-regression: \"true\"\n          strict-tools: evals          # or evals,vcr when VCR is in CI\n          comment-on-pr: \"true\"\n```\n\n| Input | Default | Description |\n|-------|---------|-------------|\n| `working-directory` | `.` | Project directory with `almanzor.yml` |\n| `provider` | `mock` | Eval provider (`mock` or `openai`) |\n| `fail-on-regression` | `true` | Fail when score drops below baseline |\n| `strict-tools` | `evals` | Comma-separated artifacts for `--strict` (`evals`, `evals,vcr`) |\n| `comment-on-pr` | `true` | Post `summary.md` as PR comment |\n| `almanzor-cli-root` | — | Build CLI from source instead of `npx` |\n\nFor local development in this repo, pass `almanzor-cli-root: ${{ github.workspace }}` to build from source instead of `npx`.\n\n## Roadmap\n\n### Next\n\n- Context window budget linter (`ctxlint`)\n- Trace visualizer (`trace view`)\n\n### Later\n\n- Cost/token profiler\n- Policy/guardrail checks\n- Local gateway/proxy\n\nSee [docs/PLUGIN.md](docs/PLUGIN.md) for the plugin + artifact convention when adding new tools.\n\n## Contributing\n\n- **Default branch:** `develop` — open PRs here for day-to-day work\n- **`main`:** stable releases only (merge from `develop` when tagging e.g. `v0.1.0`)\n- **Releases:** merge to `main`, tag `vX.Y.Z`, push tag — CI publishes `@almanzor/cli` to npm (requires `NPM_TOKEN` secret)\n- **Feature branches:** `feature/<name>` → PR to `develop`\n\n```bash\ngit checkout develop\nnpm install && npm run check\n```\n\n## License\n\nMIT\n","readmeFilename":"README.md"}