{"_id":"@cafitac/ai-crawler","_rev":"3-0140a06daa88f9b4b54bd659e8edd92e","name":"@cafitac/ai-crawler","dist-tags":{"latest":"0.1.2"},"versions":{"0.1.0":{"name":"@cafitac/ai-crawler","version":"0.1.0","keywords":["crawler","ai","mcp","wrapper","network"],"license":"MIT","_id":"@cafitac/ai-crawler@0.1.0","maintainers":[{"name":"cafitac","email":"cafitac99@gmail.com"}],"homepage":"https://github.com/cafitac/ai-crawler","bugs":{"url":"https://github.com/cafitac/ai-crawler/issues"},"bin":{"ai-crawler":"bin/ai-crawler.cjs"},"dist":{"shasum":"fdacde9760e9fcf619f52524289c8a8d53011f3f","tarball":"https://registry.npmjs.org/@cafitac/ai-crawler/-/ai-crawler-0.1.0.tgz","fileCount":5,"integrity":"sha512-o2b5Jshj7q6+WqcbVKtfM85NAHiKdZrVQNwhzAjkYVCLTKi7ANEGMc/xJgYuadPUPqPVHhbA1lQbQahbnqAPLQ==","signatures":[{"sig":"MEUCIFNdZL+/Bzt3My9k+glN6P0wNcoSrt351FNaPoeEn+VGAiEAoRHuDBuokovpvbBZzqtDYXSdeJYbyuaaETSQK+IMhFE=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@cafitac%2fai-crawler@0.1.0","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":13745},"type":"commonjs","engines":{"node":">=18"},"gitHead":"e0ee60fe5b1897385cd00d1397eadccf7982cf24","scripts":{"test":"node --test","pack:check":"npm pack --dry-run"},"_npmUser":{"name":"cafitac","email":"cafitac99@gmail.com"},"repository":{"url":"git+https://github.com/cafitac/ai-crawler.git","type":"git"},"_npmVersion":"10.9.7","description":"npm delivery wrapper for the ai-crawler Python CLI","directories":{},"_nodeVersion":"22.22.2","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"packageManager":"npm@11.5.2","_npmOperationalInternal":{"tmp":"tmp/ai-crawler_0.1.0_1777445926070_0.16378369017501848","host":"s3://npm-registry-packages-npm-production"}},"0.1.1":{"name":"@cafitac/ai-crawler","version":"0.1.1","keywords":["crawler","ai","mcp","wrapper","network"],"license":"MIT","_id":"@cafitac/ai-crawler@0.1.1","maintainers":[{"name":"cafitac","email":"cafitac99@gmail.com"}],"homepage":"https://github.com/cafitac/ai-crawler","bugs":{"url":"https://github.com/cafitac/ai-crawler/issues"},"bin":{"ai-crawler":"bin/ai-crawler.cjs"},"dist":{"shasum":"8f6309e00f0001c0560a902868dbfd1475369b6d","tarball":"https://registry.npmjs.org/@cafitac/ai-crawler/-/ai-crawler-0.1.1.tgz","fileCount":5,"integrity":"sha512-habhCAn1ojDUrJybm6di/E+JSCYhtiCNRHu2DvbQ840Y8f8XJyH8iljzY+lY4nGqFc6bdI1rMjcEHkiLrGuRUA==","signatures":[{"sig":"MEUCIQDv+jklCQpaWXJLLPqS71+/BmSFaOEeK/UGb3A33viwfQIgI8/bbTDhZJLJfceqOkighba3C7q+0Fx+NCIvkFg86aQ=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@cafitac%2fai-crawler@0.1.1","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":13862},"type":"commonjs","engines":{"node":">=18"},"gitHead":"0b448490d03e136158b3d76bce06d90b2852b0e3","scripts":{"test":"node --test","pack:check":"npm pack --dry-run"},"_npmUser":{"name":"cafitac","email":"cafitac99@gmail.com"},"repository":{"url":"git+https://github.com/cafitac/ai-crawler.git","type":"git"},"_npmVersion":"10.9.7","description":"npm delivery wrapper for the ai-crawler Python CLI","directories":{},"_nodeVersion":"22.22.2","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"packageManager":"npm@11.5.2","_npmOperationalInternal":{"tmp":"tmp/ai-crawler_0.1.1_1777447548153_0.2913912559856706","host":"s3://npm-registry-packages-npm-production"}},"0.1.2":{"name":"@cafitac/ai-crawler","version":"0.1.2","description":"npm delivery wrapper for the ai-crawler Python CLI","license":"MIT","type":"commonjs","bin":{"ai-crawler":"bin/ai-crawler.cjs"},"engines":{"node":">=18"},"scripts":{"test":"node --test","pack:check":"npm pack --dry-run","prepack":"node scripts/write-wrapper-metadata.cjs","postpack":"node scripts/cleanup-wrapper-metadata.cjs"},"keywords":["crawler","ai","mcp","wrapper","network"],"homepage":"https://github.com/cafitac/ai-crawler","repository":{"type":"git","url":"git+https://github.com/cafitac/ai-crawler.git"},"bugs":{"url":"https://github.com/cafitac/ai-crawler/issues"},"publishConfig":{"access":"public"},"packageManager":"npm@11.5.2","_id":"@cafitac/ai-crawler@0.1.2","gitHead":"de1169580d55862cec376229d68d69973ddfeda2","_nodeVersion":"22.22.2","_npmVersion":"10.9.7","dist":{"integrity":"sha512-NJu7Ptt4EWx4C/ceuSZ75TPuZGyfdyqXEjejP69r6ysrmNT65nQ8UjM+ifnJwyrKB6kyzRwBYciWWPdj+iEheA==","shasum":"27116c225087cb188ca1baff441e4a7425ecd77e","tarball":"https://registry.npmjs.org/@cafitac/ai-crawler/-/ai-crawler-0.1.2.tgz","fileCount":6,"unpackedSize":16438,"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@cafitac%2fai-crawler@0.1.2","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQDsXoxpprh0IoFuUuWQP0Jt1JygcY4BxPLVQC3bN3MqCgIgd4IJP3WkQn3t5p8qKupzsaaFBlDQoWE+4wGJuQdMT2g="}]},"_npmUser":{"name":"cafitac","email":"cafitac99@gmail.com"},"directories":{},"maintainers":[{"name":"cafitac","email":"cafitac99@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/ai-crawler_0.1.2_1777457414080_0.3455546112088206"},"_hasShrinkwrap":false}},"time":{"created":"2026-04-29T06:58:45.974Z","modified":"2026-04-29T10:10:14.474Z","0.1.0":"2026-04-29T06:58:46.209Z","0.1.1":"2026-04-29T07:25:48.356Z","0.1.2":"2026-04-29T10:10:14.229Z"},"bugs":{"url":"https://github.com/cafitac/ai-crawler/issues"},"license":"MIT","homepage":"https://github.com/cafitac/ai-crawler","keywords":["crawler","ai","mcp","wrapper","network"],"repository":{"type":"git","url":"git+https://github.com/cafitac/ai-crawler.git"},"description":"npm delivery wrapper for the ai-crawler Python CLI","maintainers":[{"name":"cafitac","email":"cafitac99@gmail.com"}],"readme":"# ai-crawler\n\n[![CI](https://github.com/cafitac/ai-crawler/actions/workflows/ci.yml/badge.svg)](https://github.com/cafitac/ai-crawler/actions/workflows/ci.yml)\n\nAI-driven network-first crawler compiler for authorized workflows.\n\n`ai-crawler` turns captured network evidence into reusable crawler recipes. The browser is used as a short-lived probe for API discovery, not as the crawling engine. Bulk collection runs through deterministic HTTP replay with `curl-cffi`.\n\n```text\nBrowser is not the crawler. Browser is the probe.\nAI is not the request loop. AI is the planner/debugger/recipe author.\n```\n\n## What it is\n\n`ai-crawler` is an early-stage Python OSS library and CLI for building crawler recipes from network evidence.\n\nIt focuses on:\n\n- Network-first API discovery and replay\n- Recipe generation, testing, repair, and deterministic execution\n- Simple CLI defaults for humans and AI harnesses\n- Python SDK facade for application integrations\n- stdio MCP server for Hermes, Claude Code, Codex, and other agents\n- Local-first tests with fake transports and fixture sites\n- Security boundaries: redaction, challenge detection, and no CAPTCHA/MFA/bot-challenge bypass logic\n\n## Install for local development\n\n```bash\ngit clone https://github.com/cafitac/ai-crawler.git\ncd ai-crawler\nuv sync --extra dev --extra http --extra mcp\n```\n\nIf you are already inside a local checkout:\n\n```bash\nuv sync --extra dev --extra http --extra mcp\n```\n\n## npm wrapper\n\nFor npm-first onboarding, the repo also ships a thin Node wrapper that delegates to the Python core:\n\n```bash\nnpx @cafitac/ai-crawler --help\nnpx @cafitac/ai-crawler auto evidence.json --json\nnpx @cafitac/ai-crawler mcp\n```\n\nWrapper behavior:\n\n- inside the repo checkout: runs the local Python core with `uv run --project <repo> ai-crawler ...`\n- outside the repo checkout: runs the published Python core via a git-pinned uvx spec when the wrapper package includes `gitHead`, otherwise falls back to `uvx --from \"git+https://github.com/cafitac/ai-crawler.git[all]\" ai-crawler ...`\n- override the published Python package spec with `AI_CRAWLER_PYTHON_SPEC`\n- override the uvx Python version with `AI_CRAWLER_UVX_PYTHON`\n\n## Quick start\n\nThe one-command path from URL to crawler artifacts is:\n\n```bash\nuv sync --extra browser --extra http\nuv run --extra browser --extra http ai-crawler compile https://example.com/products --goal \"collect products\" --json\n```\n\n`compile` opens the page briefly, records normalized network response events into `evidence.json`, generates a recipe, tests it, repairs extraction when possible, retests, and writes final JSONL output. The browser is only used for discovery; the generated recipe and final crawl use deterministic HTTP replay. By default, probe evidence keeps replay-friendly `fetch`/`xhr` 2xx/3xx responses and drops static assets, failed responses, and other browser noise.\n\nIf you want to inspect or edit evidence before compiling, split the flow:\n\n```bash\nuv run --extra browser ai-crawler probe https://example.com/products --goal \"collect products\"\nuv run --extra browser ai-crawler probe https://example.com/products --goal \"collect products\" --wait-ms 2500 --max-events 50 --include-resource-type fetch,xhr,document\nuv run --extra http ai-crawler auto evidence.json --json\n```\n\nIf you already have an evidence file, the main AI-harness command is:\n\n```bash\nai-crawler auto evidence.json --json\n```\n\nWith a local checkout:\n\n```bash\nuv run --extra http ai-crawler auto evidence.json --json\n```\n\nThis writes default artifacts:\n\n```text\nevidence.json            # browser probe evidence, if generated by probe\nrecipe.yaml              # initial generated recipe\nrepaired.recipe.yaml     # repaired/final recipe\ntest.jsonl               # initial diagnostic crawl output\ncrawl.jsonl              # final crawl output\nauto.report.json         # stable machine-readable report\n```\n\nThe JSON report includes:\n\n- final success/failure status\n- `command_type` (`compile` or `auto`)\n- `failure_phase` for quick triage (`probe`, `generate`, `final_test`, or empty on success)\n- ordered `phase_diagnostics` for `probe -> generate -> initial_test -> repair -> final_test`\n- recipe/output paths\n- initial and final crawl results\n- bounded/redacted diagnostic samples\n- failure classifications such as `success`, `extraction_failed`, `http_error`, `no_response`, `challenge_detected`, `probe_failed`, and `no_endpoint_candidates`\n\nIn `--json` mode, stdout is reserved for one machine-readable JSON object. Human-readable failures are written to stderr. Exit code `2` still writes `auto.report.json` so agents can inspect the failure.\n\n## Evidence format\n\nCreate evidence with a short browser probe:\n\n```bash\nuv run --extra browser ai-crawler probe https://example.com/products --goal \"collect products\" --output evidence.json\n```\n\nThe probe tuning options are available on both `probe` and `compile`:\n\n- `--wait-ms`: browser settle time after network idle (default: `1000`)\n- `--max-events`: maximum replay candidates retained after filtering (default: `200`)\n- `--include-resource-type`: comma-separated Playwright resource types to retain (default: `fetch,xhr`)\n\nMinimal evidence JSON:\n\n```json\n{\n  \"target_url\": \"https://example.com/products\",\n  \"goal\": \"collect products\",\n  \"events\": [\n    {\n      \"method\": \"GET\",\n      \"url\": \"https://example.com/api/products?page=1\",\n      \"status_code\": 200,\n      \"resource_type\": \"fetch\"\n    }\n  ]\n}\n```\n\nGenerate and run manually:\n\n```bash\nuv run --extra browser --extra http ai-crawler compile https://example.com/products --goal \"collect products\" --json\n```\n\nOr run each artifact step yourself:\n\n```bash\nuv run --extra http ai-crawler generate-recipe evidence.json\nuv run --extra http ai-crawler test-recipe recipe.yaml\nuv run --extra http ai-crawler repair-recipe recipe.yaml\nuv run --extra http ai-crawler test-recipe repaired.recipe.yaml --output crawl.jsonl\n```\n\n## MCP usage\n\nGenerate client config snippets for local uv-project usage. For copy-paste examples across CLI/MCP/SDK flows, also see `docs/harness-examples.md`.\n\n```bash\nuv run ai-crawler mcp-config --client hermes --project /path/to/ai-crawler\nuv run ai-crawler mcp-config --client claude-code --project /path/to/ai-crawler\nuv run ai-crawler mcp-config --client codex --project /path/to/ai-crawler\n```\n\nGenerate npm-first snippets for the published wrapper:\n\n```bash\nuv run ai-crawler mcp-config --client hermes --launcher npm\n```\n\nRun as a stdio MCP server:\n\n```bash\nuv run --extra mcp --extra http ai-crawler mcp\n```\n\nExposed tools:\n\n- `compile_url`\n- `auto_compile`\n- `generate_recipe`\n- `test_recipe`\n- `repair_recipe`\n\nIf you prefer npm-first installation for agent tooling, the wrapper can also launch the MCP server:\n\n```bash\nnpx @cafitac/ai-crawler mcp\n```\n\nHermes development snippet shape:\n\n```yaml\nmcp_servers:\n  ai-crawler:\n    command: \"uv\"\n    args: [\"run\", \"--project\", \"/path/to/ai-crawler\", \"--extra\", \"mcp\", \"--extra\", \"http\", \"ai-crawler\", \"mcp\"]\n    timeout: 300\n    connect_timeout: 60\n```\n\nHermes npm-first snippet shape:\n\n```yaml\nmcp_servers:\n  ai-crawler:\n    command: \"npx\"\n    args: [\"-y\", \"@cafitac/ai-crawler\", \"mcp\"]\n    timeout: 300\n    connect_timeout: 60\n```\n\n## Python SDK\n\nThe Python SDK remains the stable embedded/programmatic surface. The npm package is only a launcher wrapper around this Python core. See `docs/harness-examples.md` for copy-paste SDK, MCP, and published-wrapper examples.\n\nnpm publishing is automated with `.github/workflows/npm-publish.yml`.\n\n- push a tag matching the package version, for example `npm-v0.1.2`\n- or run the workflow manually with `workflow_dispatch`\n- the workflow validates that `package.json`, `pyproject.toml`, and `src/ai_crawler/__init__.py` agree on the release version before publish\n- tag-triggered publishes also validate that the pushed tag matches `npm-v<package.json version>`\n- use `docs/release-runbook.md` for the full version bump, tagging, and post-publish smoke checklist\n\nExample tag flow:\n\n```bash\ngit tag npm-v0.1.2\ngit push origin npm-v0.1.2\n```\n\n\n```python\nfrom ai_crawler import AICrawler\n\ncrawler = AICrawler()\nresult = crawler.auto(\"evidence.json\")\nprint(result.ok)\nprint(result.exit_code)\nprint(result.report)\n\ncompile_result = crawler.compile_url(\"https://example.com/products\", goal=\"collect products\")\nprint(compile_result.report[\"command_type\"])\n```\n\nFor tests or embedded usage, inject a fake fetcher:\n\n```python\ncrawler = AICrawler(fetcher=my_fake_fetcher)\n```\n\n## Verification\n\nFast local lint/type checks while iterating:\n\n```bash\nbash scripts/check-python.sh\n```\n\nFull project verification:\n\n```bash\nbash scripts/verify-ai-harness.sh\n```\n\nMCP `auto_compile` fixture smoke test:\n\n```bash\nuv run --extra http python scripts/smoke-mcp-auto-compile.py\n```\n\nThis starts a local fixture HTTP site and verifies `generate -> test -> repair -> retest` without external internet, a real browser, or a real LLM.\n\n## Security and compliance boundary\n\n`ai-crawler` is intended for authorized crawling, internal QA/testing, research, owned or allowed web property monitoring, and data portability workflows.\n\nIt does not implement:\n\n- CAPTCHA solving\n- MFA bypass\n- Cloudflare/bot-challenge bypass\n- stealth fingerprint manipulation\n- evasion proxy rotation\n\nChallenge-like responses are classified and surfaced as requiring human/manual handoff where appropriate.\n\nSensitive values in diagnostic reports are redacted, including common bearer tokens, cookies, session IDs, API keys, and JSON-embedded token fields.\n\n## Documentation\n\nDevelopment docs live under `.dev/`:\n\n- `.dev/README.md`\n- `.dev/03-ai/auto-harness-contract.md`\n- `.dev/04-mcp/server.md`\n- `.dev/08-operations/security-and-compliance.md`\n- `.dev/08-operations/challenge-handling-policy.md`\n\n## Status\n\nAlpha. The deterministic recipe compiler, one-command `compile` flow, browser probe CLI, CLI, SDK facade, MCP server, redaction, failure classification, and fixture smoke tests are implemented. Real LLM provider integrations are intentionally optional/future layers behind adapter boundaries.\n\n## License\n\nMIT\n","readmeFilename":"README.md"}