{"_id":"@andypai/agent-fetch","_rev":"2-ac6e5e3b4356c8cded8f22bd96135f27","name":"@andypai/agent-fetch","dist-tags":{"latest":"0.2.0"},"versions":{"0.1.0":{"name":"@andypai/agent-fetch","version":"0.1.0","license":"MIT","_id":"@andypai/agent-fetch@0.1.0","maintainers":[{"name":"andypai.me","email":"andy@r2pi.co"}],"homepage":"https://abpai.github.io/agent-fetch","bugs":{"url":"https://github.com/abpai/agent-fetch/issues"},"bin":{"agent-fetch":"src/index.ts"},"dist":{"shasum":"ebb87c9106e4d3baa93af23e554e3e5996df3726","tarball":"https://registry.npmjs.org/@andypai/agent-fetch/-/agent-fetch-0.1.0.tgz","fileCount":25,"integrity":"sha512-aeeNx6bQ8P/OnvSz2bFugCscNhSUXjRmlBYbetbGekEHEO0Iob44lbBYkLis2BsY5E1VdzwpbQdtZH3yldcikA==","signatures":[{"sig":"MEYCIQDfjiIrXhqHD9ZlgDPi0x6Zo1Pd26UGZYfay7x1gseALAIhANVIshQVMr2GYsGewqTLXblxWvEKNFU083XoYUcGBnVd","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":117526},"main":"src/index.ts","type":"module","module":"src/index.ts","engines":{"bun":">=1.1.0"},"exports":{".":{"types":"./src/index.ts","import":"./src/index.ts"}},"gitHead":"afb370089d0fda937f9549a3edd2f8db9dc88500","scripts":{"dev":"bun --watch src/index.ts","lint":"bunx eslint .","test":"bun test src","build":"bun build ./src/index.ts --target bun --outdir ./dist","check":"bunx prettier --check . && bun run lint && bun run typecheck && bun run test","start":"bun src/index.ts","format":"bunx prettier --write .","precommit":"bun run check","typecheck":"bunx tsc --noEmit","test:watch":"bun test --watch src"},"_npmUser":{"name":"andypai.me","email":"andy@r2pi.co"},"repository":{"url":"git+https://github.com/abpai/agent-fetch.git","type":"git"},"_npmVersion":"11.9.0","description":"Agent-first robust URL fetch CLI with smart fallback strategies.","directories":{},"_nodeVersion":"24.14.0","dependencies":{"jsdom":"^27.4.0","turndown":"^7.2.2","commander":"^14.0.3","@clack/prompts":"^1.1.0","@kreuzberg/wasm":"^4.6.2","@mozilla/readability":"^0.6.0"},"publishConfig":{"access":"public"},"_hasShrinkwrap":false,"packageManager":"bun@1.3.8","devDependencies":{"eslint":"^9.31.0","globals":"^16.2.0","prettier":"^3.5.3","bun-types":"^1.3.8","typescript":"^5.8.3","@types/node":"^24.1.0","@types/jsdom":"^27.0.0","@types/turndown":"^5.0.6","eslint-config-prettier":"^10.1.8","@typescript-eslint/parser":"^8.38.0","@typescript-eslint/eslint-plugin":"^8.38.0"},"_npmOperationalInternal":{"tmp":"tmp/agent-fetch_0.1.0_1775588344413_0.5297260114093463","host":"s3://npm-registry-packages-npm-production"}},"0.2.0":{"name":"@andypai/agent-fetch","version":"0.2.0","description":"Agent-first robust URL fetch CLI with smart fallback strategies.","type":"module","main":"src/index.ts","module":"src/index.ts","bin":{"agent-fetch":"src/index.ts"},"exports":{".":{"import":"./src/index.ts","types":"./src/index.ts"}},"engines":{"bun":">=1.1.0"},"license":"MIT","repository":{"type":"git","url":"git+https://github.com/abpai/agent-fetch.git"},"homepage":"https://abpai.github.io/agent-fetch","bugs":{"url":"https://github.com/abpai/agent-fetch/issues"},"publishConfig":{"access":"public"},"scripts":{"dev":"bun --watch src/index.ts","start":"bun src/index.ts","build":"bun build ./src/index.ts --target bun --outdir ./dist","test":"bun test src","test:watch":"bun test --watch src","lint":"bunx eslint .","format":"bunx prettier --write .","typecheck":"bunx tsc --noEmit","check":"bunx prettier --check . && bun run lint && bun run typecheck && bun run test","precommit":"bun run check"},"dependencies":{"@clack/prompts":"^1.1.0","@kreuzberg/wasm":"^4.6.2","@mozilla/readability":"^0.6.0","commander":"^14.0.3","jsdom":"^27.4.0","turndown":"^7.2.2"},"devDependencies":{"bun-types":"^1.3.8","@types/jsdom":"^27.0.0","@types/node":"^24.1.0","@types/turndown":"^5.0.6","@typescript-eslint/eslint-plugin":"^8.38.0","@typescript-eslint/parser":"^8.38.0","eslint":"^9.31.0","eslint-config-prettier":"^10.1.8","globals":"^16.2.0","prettier":"^3.5.3","typescript":"^5.8.3"},"packageManager":"bun@1.3.8","gitHead":"233ab0825ab7d29b9a9892282335703aed6831bb","_id":"@andypai/agent-fetch@0.2.0","_nodeVersion":"24.14.0","_npmVersion":"11.9.0","dist":{"integrity":"sha512-Utj48wp2gygO+isHt1dVsAjVGkA0+Y2/ke8wi9gXupJzIXMuFpNxi+uD56kTfwL93ovu1arR0q6UgsNB49pFFA==","shasum":"08b3f99bab4376cbc4e1f6e8927dafff766a6b49","tarball":"https://registry.npmjs.org/@andypai/agent-fetch/-/agent-fetch-0.2.0.tgz","fileCount":24,"unpackedSize":116805,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIB+a0SSzjj+lLc0RA/wBVDuQVxiMjFRP6Dx+VhgqI4/bAiAAhGoJEbS8lPECyZ7VCAWK3yLxHIj+14sML/TEneC63Q=="}]},"_npmUser":{"name":"andypai.me","email":"andy@r2pi.co"},"directories":{},"maintainers":[{"name":"andypai.me","email":"andy@r2pi.co"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/agent-fetch_0.2.0_1775690323812_0.6044085982926135"},"_hasShrinkwrap":false}},"time":{"created":"2026-04-07T18:59:04.313Z","modified":"2026-04-08T23:18:44.114Z","0.1.0":"2026-04-07T18:59:04.559Z","0.2.0":"2026-04-08T23:18:43.959Z"},"bugs":{"url":"https://github.com/abpai/agent-fetch/issues"},"license":"MIT","homepage":"https://abpai.github.io/agent-fetch","repository":{"type":"git","url":"git+https://github.com/abpai/agent-fetch.git"},"description":"Agent-first robust URL fetch CLI with smart fallback strategies.","maintainers":[{"name":"andypai.me","email":"andy@r2pi.co"}],"readme":"# agent-fetch\n\nReliable web access for AI agents.\n\nYour agents need to read web pages, but half the web fights back. Cloudflare challenges, SPAs that return empty `<div>`s, paywalls for content you already pay for. You can throw a headless browser at every request or route through a scraping API, but that's wasteful when a plain `fetch` would've worked for most URLs.\n\n`agent-fetch` tries the cheapest method first and escalates only when needed:\n\n1. `fetch`\n2. `jsdom`\n3. configured plugins (e.g. scrape.do)\n4. `agent-browser` (headless Chrome)\n\nEach response passes lightweight checks — word count, blocked-page detection, paywall patterns — so bad results get caught and the next strategy fires automatically.\n\nIt runs on your machine, with your IP and your browser sessions. Sites see a residential visitor, not a datacenter. When you've logged into Stratechery or your company wiki, `agent-fetch` uses those sessions so your agents aren't locked out.\n\nRun it as a CLI, use it as a library, or start the HTTP server and let any agent POST a URL to get markdown back.\n\n## Install\n\n```bash\nbun add @andypai/agent-fetch\n```\n\n`agent-fetch` is Bun-only. Use Bun for install, runtime, tests, and local tooling. Node.js and other package managers are not supported.\n\n## Giving your agents web access\n\nThe quickest path: bind `agent-fetch` on your dev machine, then put it behind [Tailscale](https://tailscale.com) so agents on your tailnet can reach it without exposing anything to the public internet.\n\n```bash\nagent-fetch server --host 0.0.0.0\n```\n\nAny agent can POST a URL and get clean markdown:\n\n```bash\ncurl -X POST -H 'Content-Type: application/json' \\\n  -d '{\"url\":\"https://example.com\"}' \\\n  http://your-machine:7411/fetch\n```\n\nThis is how tools like [openclaw](https://github.com/openclaw/openclaw) get web content — POST a link, get text back, no browser management on the agent side.\n\n### Accessing content you pay for\n\nSet up a browser profile, log into your subscriptions once, and `agent-fetch` reuses those sessions for every request:\n\n```bash\n# Create a profile and log in interactively\nagent-browser --profile ~/.agent-browser/profiles/work --headed open https://stratechery.com/\n\n# Now agents can fetch paywalled content through the server\ncurl -X POST -H 'Content-Type: application/json' \\\n  -d '{\"url\":\"https://stratechery.com/\",\"options\":{\"strategyMode\":\"authenticated\"}}' \\\n  http://localhost:7411/fetch\n```\n\nThe profile stores cookies and auth state — not your full browser history, just what's needed for fetching. If the profile isn't configured, `agent-fetch` fails fast rather than silently returning a login page.\n\n## Docs\n\n- [Architecture](./docs/ARCHITECTURE.md)\n- [Testing Guide](./docs/TESTING.md)\n\n## CLI\n\n### Fetch\n\n```bash\n# Markdown output (default)\nagent-fetch fetch https://example.com\n\n# Shorthand: a bare URL implies `fetch`\nagent-fetch https://example.com\n\n# Output mode overrides\nagent-fetch fetch https://example.com --mode markdown\nagent-fetch fetch https://example.com --mode primary\nagent-fetch fetch https://example.com --mode html\nagent-fetch fetch https://example.com --mode structured\nagent-fetch fetch https://example.com --mode screenshot\n\n# JSON output\nagent-fetch fetch https://example.com --json\n\n# Print per-attempt diagnostics to stderr\nagent-fetch fetch https://example.com --debug-attempts\n\n# Use an explicit config file\nagent-fetch fetch https://example.com --config /tmp/agent-fetch.json\n\n# Use a persistent browser profile for a one-off authenticated request\nagent-fetch fetch https://example.com --with-credentials --profile ~/.agent-browser/profiles/work\n\n# Launch agent-browser in headed mode for debugging\nagent-fetch fetch https://example.com --method agent-browser --headed\n\n# Disable fallback stages\nagent-fetch fetch https://example.com --no-jsdom --no-plugins\n\n# Authenticated fast path (agent-browser only)\nagent-fetch fetch https://example.com --with-credentials\n\n# Force strategy mode\nagent-fetch fetch https://example.com --strategy simple\nagent-fetch fetch https://example.com --strategy authenticated\n\n# Force one exact method\nagent-fetch fetch https://example.com --method fetch\nagent-fetch fetch https://example.com --method jsdom\nagent-fetch fetch https://example.com --method agent-browser\nagent-fetch fetch https://example.com --method scrape.do\n```\n\n### Output modes\n\n- `markdown` (default): convert the cleaned rendered page into markdown and keep broad page structure such as headings, cards, and tables when possible.\n- `primary`: extract article-style primary content with Readability, with metadata fallback for pages that do not have a clear article body.\n- `html`: return the rendered HTML that `agent-fetch` fetched.\n- `structured`: return structured section data derived from markdown headings and links.\n- `screenshot`: take a full-page screenshot through `agent-browser` and return the saved image path.\n\n### Method override\n\nUse `--method` when you want one exact stage instead of the usual fallback chain.\n\n- `fetch`\n- `jsdom`\n- `agent-browser`\n- built-in plugin types such as `scrape.do`\n\nNotes:\n\n- `--method scrape.do` normalizes to the built-in `scrape-do` plugin type.\n- `--mode screenshot` always uses `agent-browser`; combining it with another method is rejected.\n\n### Setup\n\n```bash\n# Guided setup\nagent-fetch setup\n\n# Alias\nagent-fetch init\n\n# Write config somewhere other than the default\nagent-fetch setup --config /tmp/agent-fetch.json\n\n# Non-interactive setup from env vars\nAGENT_FETCH_TIMEOUT=45000 \\\nAGENT_FETCH_ENABLE_PLUGINS=true \\\nSCRAPEDO_TOKEN=your-token \\\nAGENT_FETCH_ENABLE_AGENT_BROWSER=true \\\nAGENT_FETCH_PROFILE=~/.agent-browser/profiles/work \\\nagent-fetch setup --no-input --overwrite\n\n# Authenticated defaults in non-interactive mode require a profile\nAGENT_FETCH_STRATEGY_MODE=authenticated \\\nAGENT_FETCH_PROFILE=~/.agent-browser/profiles/work \\\nagent-fetch setup --no-input --overwrite\n```\n\n`agent-fetch setup --no-input` only requires `AGENT_FETCH_PROFILE` when you are\nwriting authenticated defaults. In `auto` or `simple` mode, it can write config\nwithout any browser profile settings.\n\nThe setup walkthrough configures:\n\n- default strategy mode\n- timeout and content validation thresholds\n- fetch/jsdom/plugin/agent-browser fallbacks\n- optional scrape.do plugin wiring\n- authenticated browser profile defaults\n- whether `agent-browser` waits for `networkidle` before extraction\n\n### First-time authenticated setup\n\n`agent-fetch` does not create browser sessions by itself. It passes a persistent\nprofile path to `agent-browser`, and `agent-browser` reuses that Chrome user-data\ndirectory for authenticated requests.\n\n```bash\n# Install browser binaries once\nagent-browser install\n\n# Create or warm a persistent browser profile and log in once\nagent-browser --profile ~/.agent-browser/profiles/work --headed open https://app.example.com/login\n\n# After logging in, verify the same profile is authenticated\nagent-browser --profile ~/.agent-browser/profiles/work open https://app.example.com/dashboard\n```\n\nThen save that profile for `agent-fetch`:\n\n```bash\n# Interactive\nagent-fetch setup\n\n# Non-interactive\nAGENT_FETCH_PROFILE=~/.agent-browser/profiles/work \\\nagent-fetch setup --no-input --overwrite\n\n# One-off authenticated fetch without saving defaults\nagent-fetch fetch https://app.example.com/protected \\\n  --with-credentials \\\n  --profile ~/.agent-browser/profiles/work\n```\n\nIf `--with-credentials` or `--strategy authenticated` is used and no profile is\nconfigured, `agent-fetch` fails fast instead of silently falling back to\nunauthenticated strategies.\n\n### scrape.do quickstart\n\nIf you want a hosted fallback before browser automation, the fastest setup is:\n\n```bash\nSCRAPEDO_TOKEN=your-token \\\nAGENT_FETCH_ENABLE_PLUGINS=true \\\nagent-fetch setup --no-input --overwrite\n```\n\nThat writes the built-in `scrape-do` plugin to `~/.agent-fetch/config.json`\nusing `${SCRAPEDO_TOKEN}` interpolation.\n\nYou can also wire it manually:\n\n```json\n{\n  \"enablePlugins\": true,\n  \"plugins\": [\n    {\n      \"type\": \"scrape-do\",\n      \"token\": \"${SCRAPEDO_TOKEN}\"\n    }\n  ]\n}\n```\n\n```bash\nSCRAPEDO_TOKEN=your-token agent-fetch fetch https://example.com --json --debug-attempts\n```\n\n### Server\n\nStart a local HTTP server that exposes fetch capabilities over HTTP.\n\n```bash\n# Start with defaults (127.0.0.1:7411)\nagent-fetch server\n\n# Custom port and host\nagent-fetch server --port 8080 --host 0.0.0.0\n\n# With a config file for default fetch options\nagent-fetch server --config ~/.agent-fetch/config.json\n```\n\nPOST a URL to fetch it:\n\n```bash\n# Returns markdown by default\ncurl -X POST -H 'Content-Type: application/json' \\\n  -d '{\"url\":\"https://example.com\"}' \\\n  http://localhost:7411/fetch\n\n# With options\ncurl -X POST -H 'Content-Type: application/json' \\\n  -d '{\"url\":\"https://example.com\",\"options\":{\"outputMode\":\"primary\",\"strategy\":\"simple\"}}' \\\n  http://localhost:7411/fetch\n\n# JSON response with full FetchResult\ncurl -X POST -H 'Content-Type: application/json' -H 'Accept: application/json' \\\n  -d '{\"url\":\"https://example.com\"}' \\\n  http://localhost:7411/fetch\n\n# Health check\ncurl http://localhost:7411/health\n```\n\nThe request body requires a `url` field. The `options` field is optional and accepts the same fields as `FetchOptions` (outputMode, method, timeout, strategy, etc.). When a `--config` file is provided, its defaults are merged with per-request options, with request options taking precedence.\n\nDefault response is `text/markdown`. Send `Accept: application/json` to get the full `FetchResult` JSON envelope.\n\nConstraints:\n\n- Binds to `127.0.0.1` by default (localhost-only).\n- Request body limited to 1MB.\n- HTTP-level timeout of 120s (independent of fetch strategy timeouts).\n- CORS headers (`Access-Control-Allow-Origin: *`) included on all responses.\n- `screenshotPath` in responses is a local filesystem path, only meaningful on the server host.\n\n### Plugins\n\n```bash\nagent-fetch plugins\nagent-fetch plugins --json\nagent-fetch plugins list\nagent-fetch plugins list --json\n```\n\n## Configuration\n\nDefault config path:\n\n- `~/.agent-fetch/config.json`\n- override with `--config <path>` or `AGENT_FETCH_CONFIG_PATH`\n\nAt runtime, precedence is:\n\n1. CLI flags\n2. environment variables\n3. config file\n4. built-in defaults\n\nExample config:\n\n```json\n{\n  \"timeout\": 30000,\n  \"outputMode\": \"markdown\",\n  \"enableFetch\": true,\n  \"enableJsdom\": true,\n  \"enablePlugins\": true,\n  \"enableAgentBrowser\": true,\n  \"strategyMode\": \"auto\",\n  \"plugins\": [\n    {\n      \"type\": \"scrape-do\",\n      \"token\": \"${SCRAPEDO_TOKEN}\"\n    }\n  ],\n  \"waitForNetworkIdle\": false\n}\n```\n\nNotes:\n\n- `waitForNetworkIdle` affects `agent-browser` navigation timing.\n- Plugin config values support `${ENV_VAR}` interpolation.\n- `plugins` are only used in `auto` mode, after `fetch` and `jsdom`.\n- `agentBrowser.profile` is the normal place to persist authenticated browser state.\n- `agentBrowser.command` is an advanced override when `agent-fetch` should invoke something other than plain `agent-browser`.\n\n### Legacy config behavior\n\nLegacy config files are now rejected with a hard error:\n\n- `.fetchrc.json`\n- `fetch.config.json`\n\nMove settings to `~/.agent-fetch/config.json`.\n\n### Supported setup env vars\n\n- `AGENT_FETCH_TIMEOUT`\n- `AGENT_FETCH_OUTPUT_MODE`\n- `AGENT_FETCH_ENABLE_FETCH`\n- `AGENT_FETCH_ENABLE_JSDOM`\n- `AGENT_FETCH_ENABLE_PLUGINS`\n- `AGENT_FETCH_ENABLE_AGENT_BROWSER`\n- `AGENT_FETCH_STRATEGY_MODE`\n- `AGENT_FETCH_WAIT_FOR_NETWORK_IDLE`\n- `AGENT_FETCH_USER_AGENT`\n- `AGENT_FETCH_MIN_HTML_LENGTH`\n- `AGENT_FETCH_MIN_MARKDOWN_LENGTH`\n- `AGENT_FETCH_MIN_WORD_COUNT`\n- `AGENT_FETCH_BLOCKED_WORD_COUNT_THRESHOLD`\n- `AGENT_FETCH_PROFILE`\n- `SCRAPEDO_TOKEN`\n\n## Library usage\n\n```ts\nimport { fetchUrl } from '@andypai/agent-fetch'\n\nconst result = await fetchUrl('https://example.com', {\n  strategyMode: 'auto',\n  outputMode: 'markdown',\n})\n\nconsole.log(result.outputMode)\nconsole.log(result.strategy)\nconsole.log(result.content)\n```\n\nThe package also exports `FetchError`, `registerPlugin()`, `listBuiltinPlugins()`, `parseCliArgs()`, `runCli()`, and the public fetch/plugin types if you want to embed the engine or reuse the CLI parser programmatically.\n\n## Output contract (`--json`)\n\n```json\n{\n  \"url\": \"string\",\n  \"title\": \"string\",\n  \"author\": \"string | null\",\n  \"content\": \"string\",\n  \"outputMode\": \"markdown | primary | html | structured | screenshot\",\n  \"screenshotPath\": \"string | null\",\n  \"markdown\": \"string\",\n  \"primaryMarkdown\": \"string\",\n  \"html\": \"string\",\n  \"structuredContent\": {\n    \"title\": \"string\",\n    \"description\": \"string | null\",\n    \"headings\": [{ \"level\": 2, \"text\": \"Example\" }],\n    \"sections\": [{ \"heading\": \"Example\", \"level\": 2, \"content\": \"...\" }],\n    \"links\": [{ \"text\": \"Example\", \"href\": \"https://example.com\" }]\n  },\n  \"wordCount\": 123,\n  \"strategy\": \"fetch | jsdom | plugin-name | agent-browser\",\n  \"fetchedAt\": \"ISO-8601\",\n  \"attempts\": [\n    {\n      \"strategy\": \"fetch\",\n      \"ok\": true,\n      \"durationMs\": 120\n    }\n  ]\n}\n```\n\n`content` always matches the selected output mode. `markdown` remains available in JSON output even when `--mode` is `primary`, `html`, `structured`, or `screenshot`, so callers can inspect both the selected output and the full-page markdown snapshot.\n\nWhen `--mode structured` is used without `--json`, `content` is the pretty-printed JSON string for `structuredContent`.\n\nWhen `--mode screenshot` is used, `content` and `screenshotPath` are the saved image path.\n\n## Development\n\n```bash\nbun install\nbun run build\nbun run check\nbun run test\n```\n\n### Scripts\n\n```bash\nbun run dev        # run with watch mode\nbun run start      # run once\nbun run build      # bun build ./src/index.ts --target bun --outdir ./dist\nbun run format     # prettier write\nbun run lint       # eslint\nbun run typecheck  # tsc --noEmit\nbun run test       # bun test src\nbun run test:watch # bun test --watch src\nbun run check      # prettier check + lint + typecheck + test\n```\n","readmeFilename":"README.md"}