{"_id":"@alexmelges/agentci","_rev":"2-f1117fefc264222fe5c7b42b093197bc","name":"@alexmelges/agentci","dist-tags":{"latest":"0.2.3"},"versions":{"0.2.2":{"name":"@alexmelges/agentci","version":"0.2.2","keywords":["ai","testing","llm","prompt","ci","agent"],"license":"MIT","_id":"@alexmelges/agentci@0.2.2","maintainers":[{"name":"alexmelges","email":"alexmelges@gmail.com"}],"bin":{"agentci":"dist/index.js"},"dist":{"shasum":"56c7d78ce23ba7395d27a8f960e4f68686e98917","tarball":"https://registry.npmjs.org/@alexmelges/agentci/-/agentci-0.2.2.tgz","fileCount":5,"integrity":"sha512-vz5Q5Ni0O7d3WqcFNrXkSMahVfhli1lCE8JTP2c5OXZ/HBINLz3hT0qOdlo+GF9rwcQXr4gOwZYr4V6Pzinucw==","signatures":[{"sig":"MEQCIBp9kOJHf1XsgwWf5wx8BBVoJWu8TdSeh5hcbLcFgK97AiB8wh/WzAypkDKdELL+/rQqyqS3P8az2RL6rCzctZkTaQ==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":108561},"main":"./dist/index.js","type":"module","gitHead":"09f52fe06cdd3a796258d6ee5885523f4ec1955a","scripts":{"dev":"tsup --watch","test":"vitest run","build":"tsup","typecheck":"tsc --noEmit","test:watch":"vitest"},"_npmUser":{"name":"alexmelges","email":"alexmelges@gmail.com"},"_npmVersion":"11.8.0","description":"Regression testing for AI agent workflows — pytest for prompts","directories":{},"_nodeVersion":"25.5.0","dependencies":{"ajv":"^8.17.1","yaml":"^2.7.0","chalk":"^5.4.1","openai":"^4.82.0","commander":"^13.1.0","@anthropic-ai/sdk":"^0.39.0"},"_hasShrinkwrap":false,"devDependencies":{"tsup":"^8.4.0","vitest":"^3.0.0","typescript":"^5.7.0"},"_npmOperationalInternal":{"tmp":"tmp/agentci_0.2.2_1771153623392_0.9739833485247842","host":"s3://npm-registry-packages-npm-production"}},"0.2.3":{"name":"@alexmelges/agentci","version":"0.2.3","description":"Regression testing for AI agent workflows — pytest for prompts","type":"module","bin":{"agentci":"dist/index.js"},"main":"./dist/index.js","scripts":{"build":"tsup","dev":"tsup --watch","test":"vitest run","test:watch":"vitest","typecheck":"tsc --noEmit"},"keywords":["ai","testing","llm","prompt","ci","agent"],"license":"MIT","dependencies":{"@anthropic-ai/sdk":"^0.39.0","ajv":"^8.17.1","chalk":"^5.4.1","commander":"^13.1.0","openai":"^4.82.0","yaml":"^2.7.0"},"devDependencies":{"tsup":"^8.4.0","typescript":"^5.7.0","vitest":"^3.0.0"},"gitHead":"0ce8110d30ef49d464586a252f466325674e8ae6","_id":"@alexmelges/agentci@0.2.3","_nodeVersion":"25.5.0","_npmVersion":"11.8.0","dist":{"integrity":"sha512-oUaUI/Xc9NrhgSGjF1Gyo7R8vCYDBNzpzNdtRBqs9mLXb0iRlnUTK3PR68MreB6gsiIqCT7rsSyND+mc/D9n6A==","shasum":"c889e0227701490f8c8b9bba15ca33e685fec565","tarball":"https://registry.npmjs.org/@alexmelges/agentci/-/agentci-0.2.3.tgz","fileCount":5,"unpackedSize":108703,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIHZwXVAitHHNO3lce7Wtz9SkHhblZLTZ+i6ZBqIt51d9AiAjiUH7j1dnbNCdRk//0yuKNAZDMVr46y1fz8uBTUJofQ=="}]},"_npmUser":{"name":"alexmelges","email":"alexmelges@gmail.com"},"directories":{},"maintainers":[{"name":"alexmelges","email":"alexmelges@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/agentci_0.2.3_1771155351243_0.37740293992555385"},"_hasShrinkwrap":false}},"time":{"created":"2026-02-15T11:07:03.327Z","modified":"2026-02-15T11:35:51.539Z","0.2.2":"2026-02-15T11:07:03.537Z","0.2.3":"2026-02-15T11:35:51.418Z"},"license":"MIT","keywords":["ai","testing","llm","prompt","ci","agent"],"description":"Regression testing for AI agent workflows — pytest for prompts","maintainers":[{"name":"alexmelges","email":"alexmelges@gmail.com"}],"readme":"# AgentCI\n\n**Regression testing for AI agents — pytest for prompts.**\n\n[![npm version](https://img.shields.io/npm/v/agentci.svg)](https://www.npmjs.com/package/agentci)\n[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)\n[![Tests](https://github.com/alexmelges/agentci/actions/workflows/ci.yml/badge.svg)](https://github.com/alexmelges/agentci/actions/workflows/ci.yml)\n\nAgentCI runs behavioral tests against your AI agents and prompts on every commit. Define expected behaviors in YAML, run them in CI, catch regressions before they ship.\n\n## Quick Start\n\n```bash\n# Try it instantly — no API keys needed\nnpx @alexmelges/agentci --demo\n\n# Or test your own agent\nnpx @alexmelges/agentci init       # generates agentci.yaml\nnpx @alexmelges/agentci            # runs the tests\n```\n\n### Full Setup\n\n```bash\n# Install\nnpm install -D agentci\n\n# Create a test file\ncat > agentci.yaml << 'EOF'\nversion: 1\ndefaults:\n  provider: openai\n  model: gpt-4o-mini\n  temperature: 0\n  max_tokens: 500\n\ntests:\n  - name: \"greeting response\"\n    prompt: \"Hello, I need help with my order\"\n    assertions:\n      - type: contains\n        value: \"help\"\n      - type: not_contains\n        value: \"I'm just an AI\"\nEOF\n\n# Run tests (requires OPENAI_API_KEY env var)\nnpx @alexmelges/agentci\n```\n\nOutput:\n\n```\nAgentCI v0.1.0 — Running 1 tests\n\n  ✅ greeting response (312ms)\n\nResults: 1/1 passed (100%)\n```\n\n## YAML Schema Reference\n\n```yaml\n# agentci.yaml\nversion: 1\n\n# Default settings applied to all tests\ndefaults:\n  provider: openai          # \"openai\" or \"anthropic\"\n  model: gpt-4o-mini        # model name\n  temperature: 0            # 0-2, lower = more deterministic\n  max_tokens: 500           # max response tokens\n  base_url: null             # custom OpenAI-compatible endpoint\n\ntests:\n  - name: \"test name\"         # required — unique test identifier\n    prompt: \"user message\"     # required — the prompt to send\n    system: \"system prompt\"    # optional — system message\n    context: \"grounding text\"  # optional — prepended as \"Context: ...\" to prompt\n\n    # Per-test overrides (optional)\n    provider: openai\n    model: gpt-4o\n    temperature: 0\n    max_tokens: 1000\n    base_url: https://my-proxy.example.com/v1\n\n    # Tool definitions for function calling tests (optional)\n    tools:\n      - name: get_weather\n        description: \"Get weather for a city\"\n        parameters:\n          type: object\n          properties:\n            city: { type: string }\n\n    # Assertions — at least one required\n    assertions:\n      - type: contains\n        value: \"expected text\"\n```\n\n## Assertion Types\n\nAgentCI ships with 14 assertion types — 11 deterministic + 3 LLM-as-judge:\n\n### Text Assertions\n\n| Type | Fields | Description |\n|------|--------|-------------|\n| `contains` | `value` | Response contains string (case-insensitive) |\n| `not_contains` | `value` | Response does NOT contain string |\n| `regex` | `pattern` | Response matches regex pattern |\n| `starts_with` | `value` | Response starts with string (case-insensitive, trims whitespace) |\n| `ends_with` | `value` | Response ends with string (case-insensitive, trims whitespace) |\n\n```yaml\nassertions:\n  - type: contains\n    value: \"30 days\"\n  - type: not_contains\n    value: \"I don't know\"\n  - type: regex\n    pattern: \"(refund|return|exchange)\"\n  - type: starts_with\n    value: \"Sure\"\n  - type: ends_with\n    value: \"help?\"\n```\n\n### Token Assertions\n\n| Type | Fields | Description |\n|------|--------|-------------|\n| `max_tokens` | `value` | Response is under N tokens (estimated via word count / 0.75) |\n| `min_tokens` | `value` | Response is at least N tokens |\n\n```yaml\nassertions:\n  - type: max_tokens\n    value: 200\n  - type: min_tokens\n    value: 10\n```\n\n### Tool Call Assertions\n\n| Type | Fields | Description |\n|------|--------|-------------|\n| `tool_called` | `name` | Agent called a specific tool |\n| `tool_args` | `name`, `contains` | Tool was called with specific argument values |\n\n```yaml\ntools:\n  - name: get_weather\n    description: \"Get weather for a city\"\n    parameters:\n      type: object\n      properties:\n        city: { type: string }\nassertions:\n  - type: tool_called\n    name: get_weather\n  - type: tool_args\n    name: get_weather\n    contains: { city: \"Paris\" }\n```\n\n### JSON Assertions\n\n| Type | Fields | Description |\n|------|--------|-------------|\n| `json_valid` | — | Response is valid JSON |\n| `json_schema` | `schema` | Response matches a JSON Schema |\n\n```yaml\nassertions:\n  - type: json_valid\n  - type: json_schema\n    schema:\n      type: object\n      required: [\"name\", \"age\"]\n      properties:\n        name: { type: string }\n        age: { type: number }\n```\n\n### LLM-as-Judge Assertions ⚡ NEW\n\nUse an LLM to evaluate responses when deterministic assertions aren't enough. Requires `OPENAI_API_KEY` (uses `gpt-4o-mini` by default).\n\n| Type | Fields | Description |\n|------|--------|-------------|\n| `llm_judge` | `value` | Free-form criterion — LLM evaluates if response meets it |\n| `semantic_similarity` | `value` | Response conveys same meaning as reference text |\n| `sentiment` | `value` | Response matches expected tone (professional, friendly, etc.) |\n\n```yaml\nassertions:\n  # Free-form evaluation\n  - type: llm_judge\n    value: \"Response should be helpful, concise, and not hallucinate facts\"\n\n  # Semantic matching (ignores phrasing differences)\n  - type: semantic_similarity\n    value: \"The capital of France is Paris\"\n\n  # Tone/sentiment check\n  - type: sentiment\n    value: \"professional\"\n```\n\n**Configuration:**\n- **Provider auto-detection:** Uses OpenAI if `OPENAI_API_KEY` is set, falls back to Anthropic if `ANTHROPIC_API_KEY` is set\n- Force a specific provider: `AGENTCI_JUDGE_PROVIDER=anthropic` (or `openai`)\n- Custom judge model: `AGENTCI_JUDGE_MODEL=claude-sonnet-4-20250514` (defaults: `gpt-4o-mini` for OpenAI, `claude-sonnet-4-20250514` for Anthropic)\n- Judge responses include reasoning for debuggability\n- Each judge assertion makes one additional API call\n\n**Example — testing a support bot's tone and accuracy:**\n\n```yaml\ntests:\n  - name: \"refund request — empathetic and accurate\"\n    system: \"You are a customer support agent for an e-commerce store.\"\n    prompt: \"I want a refund for my order that arrived broken\"\n    assertions:\n      # Deterministic checks\n      - type: contains\n        value: \"refund\"\n      - type: not_contains\n        value: \"I'm just an AI\"\n      # LLM-as-judge checks\n      - type: sentiment\n        value: \"empathetic and professional\"\n      - type: llm_judge\n        value: \"Response acknowledges the broken item, offers a clear refund process, and doesn't blame the customer\"\n```\n\n## Provider Configuration\n\n### OpenAI\n\nSet the `OPENAI_API_KEY` environment variable:\n\n```bash\nexport OPENAI_API_KEY=sk-...\n```\n\n```yaml\ndefaults:\n  provider: openai\n  model: gpt-4o-mini  # or gpt-4o, gpt-4-turbo, etc.\n```\n\n### Anthropic\n\nSet the `ANTHROPIC_API_KEY` environment variable:\n\n```bash\nexport ANTHROPIC_API_KEY=sk-ant-...\n```\n\n```yaml\ndefaults:\n  provider: anthropic\n  model: claude-sonnet-4-5-20250929\n```\n\n### Custom OpenAI-Compatible Endpoint\n\nUse `base_url` to point to any OpenAI-compatible API (Ollama, Azure, vLLM, LiteLLM, etc.):\n\n```yaml\ndefaults:\n  provider: openai\n  model: my-model\n  base_url: https://my-proxy.example.com/v1\n```\n\n## CLI Reference\n\n```\nUsage: agentci [options]\n\nOptions:\n  -c, --config <path>    config file path (default: \"agentci.yaml\")\n  -m, --model <model>    override model for all tests\n  -f, --format <format>  output format: text, json, markdown (default: \"text\")\n  --dry-run              validate config without calling LLM\n  --verbose              show full responses\n  -V, --version          output the version number\n  -h, --help             display help for command\n```\n\n### Examples\n\n```bash\n# Run with default config (agentci.yaml)\nnpx @alexmelges/agentci\n\n# Use a specific config file\nnpx @alexmelges/agentci --config tests/support-agent.yaml\n\n# Override the model\nnpx @alexmelges/agentci --model gpt-4o\n\n# Validate YAML without making API calls\nnpx @alexmelges/agentci --dry-run\n\n# Show full LLM responses\nnpx @alexmelges/agentci --verbose\n\n# Output as JSON (for CI parsing)\nnpx @alexmelges/agentci --format json\n\n# Output as Markdown (for PR comments)\nnpx @alexmelges/agentci --format markdown\n```\n\n## Output Formats\n\n### Text (default)\n\n```\nAgentCI v0.1.0 — Running 4 tests\n\n  ✅ greeting response (312ms)\n  ✅ refund policy (428ms)\n  ❌ tool call check (295ms)\n     ✗ tool_called: expected get_weather to be called, but no tool calls made\n  ✅ no hallucination (387ms)\n\nResults: 3/4 passed (75%)\n```\n\n### JSON (`--format json`)\n\n```json\n{\n  \"version\": \"0.1.0\",\n  \"total\": 4,\n  \"passed\": 3,\n  \"failed\": 1,\n  \"duration\": 1422,\n  \"tests\": [\n    {\n      \"name\": \"greeting response\",\n      \"passed\": true,\n      \"duration\": 312,\n      \"assertions\": [\n        { \"type\": \"contains\", \"passed\": true, \"message\": \"contains \\\"help\\\"\" }\n      ]\n    }\n  ]\n}\n```\n\n### Markdown (`--format markdown`)\n\n```markdown\n# AgentCI Results\n\n**3/4 passed (75%)** in 1422ms\n\n| Test | Status | Duration |\n|------|--------|----------|\n| greeting response | ✅ | 312ms |\n| refund policy | ✅ | 428ms |\n| tool call check | ❌ | 295ms |\n| no hallucination | ✅ | 387ms |\n\n## Failures\n\n### tool call check\n- **tool_called**: expected get_weather to be called, but no tool calls made\n```\n\n## GitHub Actions\n\n### Simple (npx)\n\n```yaml\n# .github/workflows/agentci.yml\nname: AgentCI\non: [push, pull_request]\n\njobs:\n  test-prompts:\n    runs-on: ubuntu-latest\n    steps:\n      - uses: actions/checkout@v4\n      - uses: actions/setup-node@v4\n        with:\n          node-version: 20\n      - run: npm ci\n      - run: npx @alexmelges/agentci\n        env:\n          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}\n```\n\n### GitHub Action\n\n```yaml\n# .github/workflows/agentci.yml\nname: AgentCI\non: [push, pull_request]\n\njobs:\n  test-prompts:\n    runs-on: ubuntu-latest\n    steps:\n      - uses: actions/checkout@v4\n      - uses: alexmelges/agentci@v0.1.0\n        with:\n          config: agentci.yaml\n        env:\n          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}\n```\n\n#### Action Inputs\n\n| Input | Description | Default |\n|-------|-------------|---------|\n| `config` | Path to agentci.yaml config file | `agentci.yaml` |\n| `model` | Override model for all tests | — |\n| `format` | Output format: text, json, markdown | `text` |\n| `verbose` | Show full LLM responses | `false` |\n\n## Contributing\n\nContributions are welcome! Please:\n\n1. Fork the repository\n2. Create a feature branch (`git checkout -b feature/my-feature`)\n3. Write tests for your changes\n4. Run the test suite (`npm test`)\n5. Submit a pull request\n\n### Development\n\n```bash\ngit clone https://github.com/alexmelges/agentci.git\ncd agentci\nnpm install\nnpm run build\nnpm test\n```\n\n## License\n\n[MIT](LICENSE) — Copyright 2026 Alexandre Melges\n","readmeFilename":"README.md"}