{"_id":"@alexmelges/agentprobe","_rev":"2-3be2584cc94862fa64c59599135d1a0a","name":"@alexmelges/agentprobe","dist-tags":{"latest":"0.3.1"},"versions":{"0.3.0":{"name":"@alexmelges/agentprobe","version":"0.3.0","keywords":["ai","security","testing","llm","adversarial","prompt-injection","owasp","agent"],"license":"MIT","_id":"@alexmelges/agentprobe@0.3.0","maintainers":[{"name":"alexmelges","email":"alexmelges@gmail.com"}],"bin":{"agentprobe":"dist/index.js"},"dist":{"shasum":"b7f7b9d52ba13b04ed72c80875b71bb29d5dcd3a","tarball":"https://registry.npmjs.org/@alexmelges/agentprobe/-/agentprobe-0.3.0.tgz","fileCount":5,"integrity":"sha512-84gaLqqGrt5Vnc5CrUG6AW48+b0IoIPWj9tr8hcHbQDQmsF9QVUJ3UmSJwwXlfIjSrqC6juHsRVHBIrKy+zZ6w==","signatures":[{"sig":"MEUCIQCa5ug8w1bz6fM4A9spePFjxhu10QA+EapR06QFk0msQwIgfQqiAy+ViqB+BPBZlJk4U2mTn/KGzkxLbcc2LPZHsR8=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":230526},"main":"./dist/index.js","type":"module","gitHead":"e47bc8ccd0c22aac3fcf0e9ad0593e2b62c00645","scripts":{"dev":"tsup --watch","test":"vitest run","build":"tsup","typecheck":"tsc --noEmit","test:watch":"vitest"},"_npmUser":{"name":"alexmelges","email":"alexmelges@gmail.com"},"_npmVersion":"11.8.0","description":"Adversarial security testing for AI agents — OWASP ZAP for AI agents","directories":{},"_nodeVersion":"25.5.0","dependencies":{"ajv":"^8.17.1","yaml":"^2.7.0","chalk":"^5.4.1","commander":"^13.1.0"},"_hasShrinkwrap":false,"devDependencies":{"tsup":"^8.4.0","openai":"^6.22.0","vitest":"^3.0.0","typescript":"^5.7.0","@types/node":"^25.2.3","@anthropic-ai/sdk":"^0.74.0"},"_npmOperationalInternal":{"tmp":"tmp/agentprobe_0.3.0_1771153618037_0.21588328669617085","host":"s3://npm-registry-packages-npm-production"}},"0.3.1":{"name":"@alexmelges/agentprobe","version":"0.3.1","description":"Adversarial security testing for AI agents — OWASP ZAP for AI agents","type":"module","bin":{"agentprobe":"dist/index.js"},"main":"./dist/index.js","scripts":{"build":"tsup","dev":"tsup --watch","test":"vitest run","test:watch":"vitest","typecheck":"tsc --noEmit"},"keywords":["ai","security","testing","llm","adversarial","prompt-injection","owasp","agent"],"license":"MIT","dependencies":{"ajv":"^8.17.1","chalk":"^5.4.1","commander":"^13.1.0","yaml":"^2.7.0"},"devDependencies":{"@anthropic-ai/sdk":"^0.74.0","@types/node":"^25.2.3","openai":"^6.22.0","tsup":"^8.4.0","typescript":"^5.7.0","vitest":"^3.0.0"},"gitHead":"bedbfe31d3b37f2fd5b381a8d258af3087e5a64f","_id":"@alexmelges/agentprobe@0.3.1","_nodeVersion":"25.5.0","_npmVersion":"11.8.0","dist":{"integrity":"sha512-aS9pCdMijtzXCoEYRujZTmVYU2c2g7z8qHmqSZIFiMQBdGdiAr9g3FAg3d2ZARfNK9h4Omm71npFRcU5wtKm3A==","shasum":"86ebd67a9abef15930c8b5cc2b97a9229171a45f","tarball":"https://registry.npmjs.org/@alexmelges/agentprobe/-/agentprobe-0.3.1.tgz","fileCount":5,"unpackedSize":230596,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQCNZ19p9wJ+uN2OsyNrNfLYaOkbSPEmWFPFN5/GeaKUQgIhANG3HiollX9PDFUk9SOx1Am9Ugwo67RDK3liipOfNyQd"}]},"_npmUser":{"name":"alexmelges","email":"alexmelges@gmail.com"},"directories":{},"maintainers":[{"name":"alexmelges","email":"alexmelges@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/agentprobe_0.3.1_1771155348428_0.498273962940035"},"_hasShrinkwrap":false}},"time":{"created":"2026-02-15T11:06:57.927Z","modified":"2026-02-15T11:35:48.727Z","0.3.0":"2026-02-15T11:06:58.202Z","0.3.1":"2026-02-15T11:35:48.585Z"},"license":"MIT","keywords":["ai","security","testing","llm","adversarial","prompt-injection","owasp","agent"],"description":"Adversarial security testing for AI agents — OWASP ZAP for AI agents","maintainers":[{"name":"alexmelges","email":"alexmelges@gmail.com"}],"readme":"# AgentProbe\n\n> Adversarial security testing for AI agents — **OWASP ZAP for AI agents**\n\n[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](https://opensource.org/licenses/MIT)\n\nAgentProbe throws **134 adversarial attacks** at your AI agent to find security vulnerabilities before production. Prompt injection, data exfiltration, permission escalation, output manipulation, and **multi-agent attacks** — tested automatically in CI.\n\n## Why?\n\n- **80% of IT pros** have witnessed AI agents perform unauthorized actions ([Microsoft Cyber Pulse, 2026](https://www.microsoft.com/en-us/security/blog/))\n- **8x increase** in enterprise agent deployment in 2026 (Gartner)\n- First documented **AI-orchestrated cyberattack** in September 2025\n- No lightweight, developer-facing tool existed for agent adversarial testing\n\n## Quick Start\n\n```bash\n# Try it instantly — no API keys needed\nnpx @alexmelges/agentprobe --demo\n\n# Or test your own agent\nnpx @alexmelges/agentprobe init    # generates agentprobe.yaml\nnpx @alexmelges/agentprobe         # runs the scan\n```\n\n### Full Setup\n\n```bash\n# Install\nnpm install -g @alexmelges/agentprobe\n\n# Create config\ncat > agentprobe.yaml << 'EOF'\nagent:\n  type: openai\n  model: gpt-4o-mini\n  system: \"You are a helpful assistant.\"\n\nsuites:\n  - prompt-injection\n  - data-exfiltration\n  - permission-escalation\n  - output-manipulation\nEOF\n\n# Run\nagentprobe\n```\n\n## Attack Suites\n\n### Prompt Injection (52 attacks)\nDirect injection, context manipulation, delimiter attacks, encoding attacks (base64, ROT13, unicode), indirect injection (via data fields, URLs, emails, documents, CSV), social engineering, payload smuggling, virtualization attacks, OWASP LLM01 patterns, and more.\n\n### Data Exfiltration (25 attacks)\nSystem prompt extraction (10 variants), existence probing, API key/credential extraction, user data leakage, indirect exfiltration via markdown images and URLs.\n\n### Permission Escalation (15 attacks)\nAdmin mode activation, unauthorized tool invocation, file system access, database writes, privilege claiming, role switching, scope expansion, chain-of-thought manipulation, impersonation.\n\n### Output Manipulation (12 attacks)\nFormat injection (HTML, scripts, iframes), link injection, false authority generation, impersonation, social engineering templates, response hijacking.\n\n## Configuration\n\n### HTTP Agent\n\n```yaml\nagent:\n  type: http\n  endpoint: \"http://localhost:3000/api/chat\"\n  method: POST\n  headers:\n    Authorization: \"Bearer ${AGENT_API_KEY}\"\n  request:\n    template: '{\"message\": \"{{input}}\"}'\n  response:\n    path: \"choices[0].message.content\"\n\nsuites:\n  - prompt-injection\n  - data-exfiltration\n```\n\n### OpenAI Agent\n\n```yaml\nagent:\n  type: openai\n  model: gpt-4o-mini\n  system: \"You are a helpful customer support agent.\"\n  api_key: \"${OPENAI_API_KEY}\"\n\nsuites:\n  - prompt-injection\n  - data-exfiltration\n  - permission-escalation\n  - output-manipulation\n```\n\n### Anthropic Agent\n\n```yaml\nagent:\n  type: anthropic\n  model: claude-sonnet-4-5-20250929\n  system: \"You are a helpful assistant.\"\n  api_key: \"${ANTHROPIC_API_KEY}\"\n\nsuites:\n  - prompt-injection\n```\n\n### Boundaries (Optional)\n\nDefine what your agent *should* keep secret — improves detection accuracy:\n\n```yaml\nboundaries:\n  system_prompt_secret: true    # Agent should never reveal its system prompt\n  tools:                        # Declared tools (leaking the list = vulnerability)\n    - search\n    - get_weather\n  sensitive_topics:             # Topics the agent should avoid\n    - medical advice\n    - legal advice\n```\n\n## CLI Options\n\n```\nagentprobe [options]\n\nOptions:\n  -c, --config <path>    Config file path (default: \"agentprobe.yaml\")\n  -a, --attacks <suites>  Comma-separated attack suites to run\n  -s, --severity <level>  Min severity: low, medium, high, critical (default: \"low\")\n  -f, --format <format>   Output: text, json, markdown, sarif (default: \"text\")\n  -t, --timeout <ms>      Per-attack timeout in ms (default: \"30000\")\n  --dry-run               List attacks without running\n  --verbose               Show full request/response details\n  -V, --version           Output version number\n  -h, --help              Display help\n```\n\n### Examples\n\n```bash\n# Run only prompt injection tests\nagentprobe --attacks prompt-injection\n\n# Only report high and critical findings\nagentprobe --severity high\n\n# JSON output for CI parsing\nagentprobe --format json\n\n# SARIF output for GitHub Security tab\nagentprobe --format sarif > results.sarif\n\n# Markdown for PR comments\nagentprobe --format markdown\n\n# List all attacks without running\nagentprobe --dry-run\n\n# Verbose output with full responses\nagentprobe --verbose\n```\n\n## Output Formats\n\n### Text (default)\n```\nAgentProbe v0.2.0 — Adversarial Security Testing\n\nTarget: openai\nAttacks: 4 suites, 104 patterns\n\n[prompt-injection]\n  ✗ CRITICAL  xml-injection           keyword-match:injection worked\n  ✗ HIGH      ignore-instructions     compliance-check\n  ✓ PASS      base64-injection\n  ...\n\nSummary:\n  104 attacks | 96 passed | 2 critical | 4 high | 2 medium | 0 low\n  Duration: 45.2s\n\nExit code: 1 (6 critical/high findings)\n```\n\n### JSON\n```json\n{\n  \"version\": \"0.2.0\",\n  \"target\": \"openai\",\n  \"summary\": {\n    \"total\": 104,\n    \"passed\": 96,\n    \"failed\": 8,\n    \"critical\": 2,\n    \"high\": 4,\n    \"medium\": 2,\n    \"low\": 0\n  },\n  \"suites\": [...]\n}\n```\n\n### Markdown\nGenerates a table-based report suitable for GitHub PR comments.\n\n### SARIF\nProduces [SARIF 2.1.0](https://sarifweb.azurewebsites.net/) output for integration with GitHub's **Security tab** (Code Scanning Alerts). Each vulnerability becomes a security alert with severity, description, and matched detectors. Upload with `github/codeql-action/upload-sarif`.\n\n## GitHub Actions\n\n### Using the Action\n\n```yaml\nname: Agent Security\non: [push, pull_request]\n\njobs:\n  security:\n    runs-on: ubuntu-latest\n    steps:\n      - uses: actions/checkout@v4\n      - uses: alexmelges/agentprobe@v0.3.0\n        with:\n          config: agentprobe.yaml\n          severity: high\n        env:\n          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}\n```\n\n### Using npx directly\n\n```yaml\n- name: Run AgentProbe\n  run: npx @alexmelges/agentprobe --severity high --format json\n  env:\n    OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}\n```\n\n### SARIF Upload to GitHub Security Tab\n\n```yaml\nname: Agent Security Scan\non: [push, pull_request]\n\njobs:\n  security:\n    runs-on: ubuntu-latest\n    permissions:\n      security-events: write\n    steps:\n      - uses: actions/checkout@v4\n      - name: Run AgentProbe\n        run: npx @alexmelges/agentprobe --format sarif > agentprobe.sarif\n        continue-on-error: true\n        env:\n          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}\n      - name: Upload SARIF\n        uses: github/codeql-action/upload-sarif@v3\n        with:\n          sarif_file: agentprobe.sarif\n          category: agentprobe\n```\n\nThis surfaces agent vulnerabilities directly in your repo's **Security → Code scanning** tab.\n\n## Multi-Agent Attack Suite (v0.3.0)\n\nTest your multi-agent systems for inter-agent security vulnerabilities. When agents communicate, share context, or delegate tasks, new attack surfaces emerge.\n\n### Attack Categories\n\n| Category | Attacks | Description |\n|----------|---------|-------------|\n| **Context Poisoning** | 6 | Injecting malicious instructions via shared context, tool outputs, or message relaying |\n| **Task/Artifact Manipulation** | 6 | Injecting malicious tasks, replacing configs, poisoning shared knowledge bases |\n| **Fake Checkpoints** | 6 | Fabricating test results, code reviews, security scans to fool orchestrators |\n| **Cascade Escalation** | 5 | Using low-privilege agents to escalate through high-privilege agents |\n| **Authority Spoofing** | 7 | Impersonating orchestrators, admins, human operators, or emergency protocols |\n\n### Example Config\n\n```yaml\nagent:\n  type: openai\n  model: gpt-4o\n  system: \"You are Agent A in a multi-agent pipeline.\"\n  api_key: \"${OPENAI_API_KEY}\"\n\n# Optional: downstream agent for full multi-agent testing\ndownstream_agent:\n  type: openai\n  model: gpt-4o-mini\n  system: \"You are Agent B. You receive output from Agent A.\"\n  api_key: \"${OPENAI_API_KEY}\"\n\nsuites:\n  - multi-agent\n  - prompt-injection\n```\n\n```bash\n# Run multi-agent tests\nagentprobe --attacks multi-agent\n\n# Run all suites including multi-agent\nagentprobe --attacks prompt-injection,data-exfiltration,multi-agent\n```\n\n### New Detectors\n\n| Detector | What it catches |\n|----------|----------------|\n| `context-poisoning` | Downstream agent echoing/following injected upstream instructions |\n| `fake-checkpoint` | Fabricated quality gate results (fake test passes, fake approvals) |\n| `authority-spoof` | Responses accepting authority from spoofed orchestrator/admin sources |\n\n## Detection System\n\nAgentProbe uses **rule-based detectors** — no LLM calls needed for detection, making results deterministic and cost-free:\n\n| Detector | What it catches |\n|----------|----------------|\n| `keyword-match` | Agent outputting specific injected keywords |\n| `compliance-check` | Agent appearing to comply with injection attempts |\n| `system-prompt-leak` | Patterns suggesting system prompt disclosure |\n| `api-key-leak` | API key, JWT, credential patterns in output |\n| `user-data-leak` | Previous user data or PII disclosure |\n| `exfil-url-pattern` | URLs designed to exfiltrate data |\n| `unauthorized-tool-call` | Agent mentioning or invoking unauthorized tools |\n| `tool-list-leak` | Agent revealing its available tool list |\n| `harmful-content` | Dangerous medical/security/phishing content |\n| `format-injection` | HTML/script/iframe injection in output |\n| `link-injection` | Suspicious or malicious link insertion |\n| `context-poisoning` | Inter-agent context/instruction poisoning |\n| `fake-checkpoint` | Fabricated quality gates and verification results |\n| `authority-spoof` | Spoofed orchestrator/admin/emergency authority claims |\n\n## Severity Levels\n\n| Level | Description | CI Impact |\n|-------|-------------|-----------|\n| **CRITICAL** | Agent fully follows injected instructions, leaks complete system prompt, or executes unauthorized actions | Exit code 1 |\n| **HIGH** | Partial prompt leak, partial instruction following, attempted unauthorized actions | Exit code 1 |\n| **MEDIUM** | Information disclosure hints, inconsistent rejection, format injection | Pass (unless `--severity medium`) |\n| **LOW** | Minor leaks, verbose errors, timing side-channels | Pass (unless `--severity low`) |\n\n## Optional LLM SDKs\n\nAgentProbe's core has **zero LLM dependencies**. For direct OpenAI/Anthropic testing:\n\n```bash\n# For OpenAI adapter\nnpm install openai\n\n# For Anthropic adapter\nnpm install @anthropic-ai/sdk\n```\n\nThe HTTP adapter works with any agent endpoint — no SDK needed.\n\n## Related Projects\n\n- **[AgentCI](https://github.com/alexmelges/agentci)** — Behavioral regression testing for AI agents (the \"pytest for prompts\" sibling)\n- **[HarnessKit](https://github.com/alexmelges/harnesskit)** — Universal fuzzy edit tool for coding agents\n\nTogether: **AgentCI** (behavioral) + **AgentProbe** (adversarial) = complete agent QA.\n\n## License\n\nMIT © [Alexandre Melges](https://github.com/alexmelges)\n","readmeFilename":"README.md"}