{"_id":"@chambrin/ai-crawler-guard","name":"@chambrin/ai-crawler-guard","dist-tags":{"latest":"0.1.0"},"versions":{"0.1.0":{"name":"@chambrin/ai-crawler-guard","version":"0.1.0","description":"Detect and control AI crawlers (GPTBot, ClaudeBot, PerplexityBot) with configurable actions","main":"./dist/index.js","module":"./dist/index.mjs","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.mjs","require":"./dist/index.js"},"./middleware":{"types":"./dist/middleware/index.d.ts","import":"./dist/middleware/index.mjs","require":"./dist/middleware/index.js"},"./robots-txt":{"types":"./dist/robots-txt/index.d.ts","import":"./dist/robots-txt/index.mjs","require":"./dist/robots-txt/index.js"}},"scripts":{"build":"tsup","dev":"tsup --watch","test":"vitest run","test:watch":"vitest","test:coverage":"vitest run --coverage","typecheck":"tsc --noEmit","lint":"eslint src","lint:fix":"eslint src --fix","format":"prettier --write \"src/**/*.ts\"","format:check":"prettier --check \"src/**/*.ts\"","prepublishOnly":"npm run typecheck && npm run lint && npm run test && npm run build"},"keywords":["ai","crawler","bot","detection","gptbot","claudebot","perplexitybot","middleware","nextjs","express","hono","h3","robots.txt","ai-detection","web-scraping","bot-detection"],"author":{"name":"chambrin"},"license":"MIT","devDependencies":{"@types/node":"^20.11.0","@typescript-eslint/eslint-plugin":"^6.19.0","@typescript-eslint/parser":"^6.19.0","@vitest/coverage-v8":"^1.2.0","eslint":"^8.56.0","prettier":"^3.2.4","tsup":"^8.0.1","typescript":"^5.3.3","vitest":"^1.2.0"},"peerDependencies":{},"repository":{"type":"git","url":"git+https://github.com/chambrin/ai-crawler-guard.git"},"bugs":{"url":"https://github.com/chambrin/ai-crawler-guard/issues"},"homepage":"https://github.com/chambrin/ai-crawler-guard#readme","engines":{"node":">=18.0.0"},"_id":"@chambrin/ai-crawler-guard@0.1.0","gitHead":"93b91f0bf82f66c309999b39eb189e88977bf3c1","_nodeVersion":"24.4.1","_npmVersion":"11.4.2","dist":{"integrity":"sha512-gyWPH+eOJelJY7O1pdqkucbON9YLpW2MgXuTEYvn3FuqUX1K2Rafv5LplFpKivXEVJtlIqBD2P5mgmfLP4HeYg==","shasum":"09301218fd445d5c4361848ecdea8398c10bdec8","tarball":"https://registry.npmjs.org/@chambrin/ai-crawler-guard/-/ai-crawler-guard-0.1.0.tgz","fileCount":26,"unpackedSize":283949,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIFJEeedesZZgkh7aTNTsxPuGZw13HG+6l0lTgVc+iKEHAiAGB9JwyL3AvFAeF+PuI7I/+T/A2/R9eHpCBwqNK/11cg=="}]},"_npmUser":{"name":"chambrin","email":"achambrin@outlook.fr"},"directories":{},"maintainers":[{"name":"chambrin","email":"achambrin@outlook.fr"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/ai-crawler-guard_0.1.0_1773058500619_0.7685308117210552"},"_hasShrinkwrap":false}},"time":{"created":"2026-03-09T12:15:00.554Z","0.1.0":"2026-03-09T12:15:00.773Z","modified":"2026-03-09T12:15:01.292Z"},"maintainers":[{"name":"chambrin","email":"achambrin@outlook.fr"}],"description":"Detect and control AI crawlers (GPTBot, ClaudeBot, PerplexityBot) with configurable actions","homepage":"https://github.com/chambrin/ai-crawler-guard#readme","keywords":["ai","crawler","bot","detection","gptbot","claudebot","perplexitybot","middleware","nextjs","express","hono","h3","robots.txt","ai-detection","web-scraping","bot-detection"],"repository":{"type":"git","url":"git+https://github.com/chambrin/ai-crawler-guard.git"},"author":{"name":"chambrin"},"bugs":{"url":"https://github.com/chambrin/ai-crawler-guard/issues"},"license":"MIT","readme":"# @chambrin/ai-crawler-guard\n\n> Detect and control AI crawlers (GPTBot, ClaudeBot, PerplexityBot, etc.) with configurable server-side actions.\n\nA lightweight, framework-agnostic TypeScript library to detect AI crawlers and execute customizable actions like blocking images, redirecting, or logging visits. Works seamlessly with Next.js, Express, Hono, Nuxt, and SvelteKit.\n\n## Features\n\n- **Server-side only** - No client-side JavaScript needed\n- **Framework-agnostic** - Works with any Node.js framework\n- **Ready-to-use middlewares** for Next.js, Express, Hono, and H3 (Nuxt/SvelteKit)\n- **Configurable actions** - Block images, redirect, log, or create custom actions\n- **robots.txt generator** - Generate robots.txt rules automatically\n- **TypeScript** - Fully typed with strict types\n- **Lightweight** - Zero dependencies for core functionality\n- **Extensible** - Add custom bot detection patterns\n\n## Installation\n\n```bash\nnpm install @chambrin/ai-crawler-guard\n```\n\n## Quick Start\n\n### Next.js (App Router)\n\n```typescript\n// middleware.ts\nimport { nextMiddleware } from '@chambrin/ai-crawler-guard';\n\nexport const config = {\n  matcher: ['/((?!api|_next/static|_next/image|favicon.ico).*)'],\n};\n\nexport default nextMiddleware({\n  blockImagesFor: ['gptbot', 'claudebot', 'perplexitybot'],\n  redirectUrls: {\n    gptbot: '/blocked',\n  },\n  logLevel: 'info',\n});\n```\n\n### Express\n\n```typescript\nimport express from 'express';\nimport { expressMiddleware } from '@chambrin/ai-crawler-guard/core';\n\nconst app = express();\n\napp.use(expressMiddleware({\n  blockImagesFor: ['gptbot', 'claudebot'],\n  logLevel: 'warn',\n}));\n```\n\n### Hono\n\n```typescript\nimport { Hono } from 'hono';\nimport { honoMiddleware } from '@chambrin/ai-crawler-guard/core';\n\nconst app = new Hono();\n\napp.use('*', honoMiddleware({\n  blockImagesFor: ['gptbot', 'claudebot'],\n  redirectUrls: {\n    perplexitybot: '/no-ai',\n  },\n}));\n```\n\n### Nuxt / SvelteKit (H3)\n\n```typescript\n// server/middleware/ai-guard.ts\nimport { h3Middleware } from '@chambrin/ai-crawler-guard/core';\n\nexport default h3Middleware({\n  blockImagesFor: ['gptbot', 'claudebot'],\n  logLevel: 'info',\n});\n```\n\n## API Reference\n\n### Detection\n\n#### `detectAiCrawler(request: Request): AiCrawlerMatch`\n\nDetect AI crawler from a Web Request object.\n\n```typescript\nimport { detectAiCrawler } from '@chambrin/ai-crawler-guard/core';\n\nconst match = detectAiCrawler(request);\n\nif (match.type === 'gptbot') {\n  console.log('GPTBot detected!');\n}\n```\n\n#### `detectAiCrawler(userAgent: string, ip?: string): AiCrawlerMatch`\n\nDetect AI crawler from user agent string.\n\n```typescript\nconst match = detectAiCrawler('Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.0;');\n// match.type === 'gptbot'\n```\n\n#### `isAiCrawler(requestOrUserAgent: Request | string, type?: AiCrawlerType): boolean`\n\nQuick check if request is from an AI crawler.\n\n```typescript\nif (isAiCrawler(request)) {\n  // Any AI crawler detected\n}\n\nif (isAiCrawler(request, 'gptbot')) {\n  // Specifically GPTBot\n}\n```\n\n### Types\n\n```typescript\ntype AiCrawlerType =\n  | 'gptbot'\n  | 'claudebot'\n  | 'perplexitybot'\n  | 'anthropic-ai'\n  | 'google-extended'\n  | 'bytespider'\n  | 'ccbot'\n  | 'custom'\n  | null;\n\ninterface AiCrawlerMatch {\n  type: AiCrawlerType;\n  confidence: number; // 0-1\n  userAgent: string;\n  ip?: string;\n  isKnown: boolean;\n}\n```\n\n### Actions\n\nActions are composable functions that execute when an AI crawler is detected.\n\n#### `blockImages()`\n\nBlock all image requests with 403 Forbidden.\n\n```typescript\nimport { AiCrawlerGuard, blockImages } from '@chambrin/ai-crawler-guard/core';\n\nconst guard = new AiCrawlerGuard()\n  .addAction(blockImages());\n```\n\n#### `redirect(url: string, statusCode?: number)`\n\nRedirect AI crawlers to a specific URL.\n\n```typescript\nguard.addAction(redirect('/no-ai', 302));\n```\n\n#### `log(level?: 'info' | 'warn' | 'error')`\n\nLog AI crawler visits.\n\n```typescript\nguard.addAction(log('info'));\n```\n\n#### `textOnly()`\n\nBlock all non-text content (images, CSS, JS, fonts, etc.).\n\n```typescript\nguard.addAction(textOnly());\n```\n\n### Guard\n\nThe `AiCrawlerGuard` class manages a pipeline of actions.\n\n```typescript\nimport { AiCrawlerGuard, detectAiCrawler, blockImages, redirect, log } from '@chambrin/ai-crawler-guard/core';\n\nconst guard = new AiCrawlerGuard()\n  .addAction(log('info'))\n  .addAction(blockImages())\n  .addAction(redirect('/blocked'));\n\nconst match = detectAiCrawler(request);\nif (match.type) {\n  const response = guard.execute(match, request);\n  if (response) {\n    return response; // Return the response from the first action that returns one\n  }\n}\n```\n\n### Configuration\n\n```typescript\ninterface AiCrawlerConfig {\n  knownBots: Record<string, AiCrawlerType>;\n  blockImagesFor: AiCrawlerType[];\n  redirectUrls: Partial<Record<AiCrawlerType, string>>;\n  logLevel: 'none' | 'info' | 'warn';\n  enableIpTracking?: boolean;\n}\n```\n\n### Robots.txt Generation\n\n#### `generateRobotsTxt(config: Partial<AiCrawlerConfig>): string`\n\nGenerate robots.txt content based on configuration.\n\n```typescript\nimport { generateRobotsTxt } from '@chambrin/ai-crawler-guard/robots-txt';\n\nconst robotsTxt = generateRobotsTxt({\n  blockImagesFor: ['gptbot', 'claudebot'],\n});\n\n// In your route handler:\n// app.get('/robots.txt', (req, res) => {\n//   res.setHeader('Content-Type', 'text/plain');\n//   res.send(robotsTxt);\n// });\n```\n\n#### Presets\n\n```typescript\nimport {\n  defaultAiBotsRobotsTxt,\n  blockImagesPreset,\n  blockGPTBotOnly\n} from '@chambrin/ai-crawler-guard/robots-txt';\n\n// Block all AI crawlers completely\nconsole.log(defaultAiBotsRobotsTxt);\n\n// Block only images\nconsole.log(blockImagesPreset);\n\n// Block only GPTBot\nconsole.log(blockGPTBotOnly);\n```\n\n## Detected Bots\n\nThe library detects the following AI crawlers by default:\n\n| Bot Type | User-Agent Patterns |\n|----------|-------------------|\n| `gptbot` | GPTBot, ChatGPT-User |\n| `claudebot` | ClaudeBot, Claude-Web |\n| `anthropic-ai` | anthropic-ai |\n| `perplexitybot` | PerplexityBot |\n| `google-extended` | Google-Extended |\n| `bytespider` | Bytespider (ByteDance) |\n| `ccbot` | CCBot (Common Crawl) |\n\nAdditional bots detected with lower confidence:\n- Cohere-AI\n- Omgilibot\n- Diffbot\n- FacebookBot\n- Various AI scrapers\n\n## Advanced Usage\n\n### Custom Actions\n\nCreate your own action executor:\n\n```typescript\nimport { ActionExecutor, AiCrawlerMatch } from '@chambrin/ai-crawler-guard/core';\n\nfunction customBlock(): ActionExecutor {\n  return {\n    execute(match: AiCrawlerMatch, request?: Request): Response | void {\n      if (match.type === 'gptbot') {\n        return new Response('GPTBot not allowed', { status: 403 });\n      }\n    }\n  };\n}\n\nconst guard = new AiCrawlerGuard()\n  .addAction(customBlock());\n```\n\n### Add Custom Bots\n\nExtend the known bots list:\n\n```typescript\nimport { DEFAULT_KNOWN_BOTS } from '@chambrin/ai-crawler-guard/core';\n\nconst customConfig = {\n  knownBots: {\n    ...DEFAULT_KNOWN_BOTS,\n    'my-custom-bot': 'custom',\n  },\n  blockImagesFor: ['custom'],\n};\n```\n\n### Next.js Custom Middleware\n\n```typescript\nimport { createNextMiddleware, blockImages, log } from '@chambrin/ai-crawler-guard/core';\n\nexport default createNextMiddleware((guard, config) => {\n  guard\n    .addAction(log('warn'))\n    .addAction(blockImages());\n}, {\n  logLevel: 'warn',\n  blockImagesFor: ['gptbot'],\n});\n```\n\n### Conditional Actions\n\n```typescript\nconst guard = new AiCrawlerGuard();\n\nconst match = detectAiCrawler(request);\n\nif (match.type === 'gptbot') {\n  guard.addAction(redirect('/gptbot-blocked'));\n} else if (match.type === 'claudebot') {\n  guard.addAction(blockImages());\n}\n\nconst response = guard.execute(match, request);\n```\n\n## Examples\n\nSee the `/examples` directory for complete working examples:\n\n- Next.js 15 App Router\n- Express server\n- Hono API\n- Nuxt 4 application\n- SvelteKit application\n\n## Important Notes\n\n### SEO Disclaimer\n\nThis library is designed to control AI crawlers for legitimate purposes such as:\n- Protecting proprietary content from being used in AI training\n- Reducing server load from AI crawlers\n- Enforcing terms of service\n\n**DO NOT** use this library for:\n- Cloaking content from search engines (violates Google's guidelines)\n- Serving different content to users vs. crawlers\n- Any deceptive SEO practices\n\nAlways respect search engine guidelines and robots.txt standards.\n\n### Server-Side Only\n\nThis library works exclusively on the server side. Client-side detection is ineffective against crawlers since they don't execute JavaScript.\n\n### Performance\n\nThe library is lightweight and has minimal performance impact:\n- User agent detection is based on simple string matching\n- No external API calls\n- No heavy dependencies\n\n## License\n\nMIT\n\n## Contributing\n\nContributions are welcome! Please open an issue or submit a pull request.\n\n## Links\n\n- [GitHub Repository](https://github.com/chambrin/ai-crawler-guard)\n- [NPM Package](https://www.npmjs.com/package/@chambrin/ai-crawler-guard)\n- [Issues](https://github.com/chambrin/ai-crawler-guard/issues)\n","readmeFilename":"README.md","_rev":"1-4e4620df2c7885ac9fa28eec8a605b32"}