{"_id":"@anisirji/web-extractor","_rev":"3-77c9a48425c22cbaa89aa0c724ed9d1e","name":"@anisirji/web-extractor","dist-tags":{"latest":"1.0.2"},"versions":{"1.0.0":{"name":"@anisirji/web-extractor","version":"1.0.0","keywords":["web-scraping","content-extraction","url-parser","firecrawl","scraper","termix"],"author":{"name":"anisirji"},"license":"MIT","_id":"@anisirji/web-extractor@1.0.0","maintainers":[{"name":"anisirji","email":"www.allinoneforyouhere@gmail.com"}],"homepage":"https://github.com/anisirji/llm-web-extractor#readme","bugs":{"url":"https://github.com/anisirji/llm-web-extractor/issues"},"dist":{"shasum":"a1b145045feaaef3d2c66fc8e9b830065beb60b3","tarball":"https://registry.npmjs.org/@anisirji/web-extractor/-/web-extractor-1.0.0.tgz","fileCount":12,"integrity":"sha512-STEJSNDNfBWlYMwl/ds9dOp3kyWK+gYfAjhQHWChfmYOoOUfrbgBZO8iiDC1eIWB/3rwZQs1WaqCJKO8T1jhcg==","signatures":[{"sig":"MEUCIDekasOQDfQ3lIs9fuXhUGJ9YXZeueQc8HdPulLXyFabAiEAnIZ4pAxKv94/aUx6SvxgNmJFfE7eMSZ3GS00JP87mBY=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":33147},"main":"dist/index.js","types":"dist/index.d.ts","gitHead":"d07fe7ad71c5b7e22837a1b7d63e65e31b622bd7","scripts":{"dev":"tsc --watch","test":"tsx test/test.ts","build":"tsc","prepublishOnly":"npm run build","test:integration":"tsx examples/basic-usage.ts"},"_npmUser":{"name":"anisirji","email":"www.allinoneforyouhere@gmail.com"},"repository":{"url":"git+https://github.com/anisirji/llm-web-extractor.git","type":"git"},"_npmVersion":"10.9.2","description":"Powerful web content extraction SDK with URL normalization and intelligent scraping - https://github.com/anisirji/llm-web-extractor","directories":{},"_nodeVersion":"22.17.1","dependencies":{"@mendable/firecrawl-js":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"tsx":"^4.7.0","typescript":"^5.0.0","@types/node":"^20.0.0"},"peerDependencies":{},"_npmOperationalInternal":{"tmp":"tmp/web-extractor_1.0.0_1763302694340_0.2886853483120668","host":"s3://npm-registry-packages-npm-production"}},"1.0.1":{"name":"@anisirji/web-extractor","version":"1.0.1","keywords":["web-scraping","content-extraction","url-parser","firecrawl","scraper","termix"],"author":{"name":"anisirji"},"license":"MIT","_id":"@anisirji/web-extractor@1.0.1","maintainers":[{"name":"anisirji","email":"www.allinoneforyouhere@gmail.com"}],"homepage":"https://github.com/anisirji/llm-web-extractor#readme","bugs":{"url":"https://github.com/anisirji/llm-web-extractor/issues"},"dist":{"shasum":"e0ec452232812f291929d90dfe73afde7ca8bb24","tarball":"https://registry.npmjs.org/@anisirji/web-extractor/-/web-extractor-1.0.1.tgz","fileCount":12,"integrity":"sha512-3Gj8BtvU3/RFOnpfmcGlmD13AN3eslFwP+eRAQqVpU1TxeJamH+E3ghoylL7hE/yJrsS9Qg/m2IW3KMAW9A2Tg==","signatures":[{"sig":"MEQCIEg7TBd1yqCdIeKASopvsnHX6O/ZJO5YNHTKLCahCKF6AiBsmIuYLXZNRt4Fb5s2lMd55NJgP+8+ZfgEUsgujo9OiQ==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":33197},"main":"dist/index.js","types":"dist/index.d.ts","gitHead":"abcc7c29c12a6c443471b2ba670167d03e1ed4ed","scripts":{"dev":"tsc --watch","test":"tsx test/test.ts","build":"tsc","prepublishOnly":"npm run build","test:integration":"tsx examples/basic-usage.ts"},"_npmUser":{"name":"anisirji","email":"www.allinoneforyouhere@gmail.com"},"repository":{"url":"git+https://github.com/anisirji/llm-web-extractor.git","type":"git"},"_npmVersion":"10.9.2","description":"Powerful web content extraction SDK with URL normalization and intelligent scraping - https://github.com/anisirji/llm-web-extractor","directories":{},"_nodeVersion":"22.17.1","dependencies":{"@mendable/firecrawl-js":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"tsx":"^4.7.0","typescript":"^5.0.0","@types/node":"^20.0.0"},"peerDependencies":{},"_npmOperationalInternal":{"tmp":"tmp/web-extractor_1.0.1_1763303066596_0.8121447644431166","host":"s3://npm-registry-packages-npm-production"}},"1.0.2":{"name":"@anisirji/web-extractor","version":"1.0.2","description":"Powerful web content extraction SDK with URL normalization and intelligent scraping - https://github.com/anisirji/llm-web-extractor","main":"dist/index.js","types":"dist/index.d.ts","repository":{"type":"git","url":"git+https://github.com/anisirji/llm-web-extractor.git"},"homepage":"https://github.com/anisirji/llm-web-extractor#readme","bugs":{"url":"https://github.com/anisirji/llm-web-extractor/issues"},"scripts":{"build":"tsc","dev":"tsc --watch","test":"tsx test/test.ts","test:integration":"tsx examples/basic-usage.ts","test:astratechai":"tsx test/test-astratechai.ts","prepublishOnly":"npm run build"},"keywords":["web-scraping","content-extraction","url-parser","firecrawl","scraper","termix"],"author":{"name":"anisirji"},"license":"MIT","dependencies":{"@mendable/firecrawl-js":"^1.0.0"},"devDependencies":{"@types/node":"^20.0.0","typescript":"^5.0.0","tsx":"^4.7.0"},"peerDependencies":{},"_id":"@anisirji/web-extractor@1.0.2","gitHead":"5d05e506c7dcd6ab2ba1f83a91f765e11943e362","_nodeVersion":"22.17.1","_npmVersion":"10.9.2","dist":{"integrity":"sha512-QweSTm6ueYAKDmCS9G0RtEPJu45blJS45V+WuN1PPUeEAzTnxHoAHOzD9SIns4imNFv0jN9gg2GuUPvDfBpubw==","shasum":"e9952f5f817e8a472a88529cd045e6eb98f88ccb","tarball":"https://registry.npmjs.org/@anisirji/web-extractor/-/web-extractor-1.0.2.tgz","fileCount":15,"unpackedSize":73571,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQDiP06BnPOdvJirYQ5IcQzBeaOUHs+dJM+fFSdWlyzp+AIhAPCU3HFiPDECPI1PK6wXXZqTlBuNkmppXGYevTAuMR7y"}]},"_npmUser":{"name":"anisirji","email":"www.allinoneforyouhere@gmail.com"},"directories":{},"maintainers":[{"name":"anisirji","email":"www.allinoneforyouhere@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/web-extractor_1.0.2_1763790243490_0.10032514350578081"},"_hasShrinkwrap":false}},"time":{"created":"2025-11-16T14:18:14.249Z","modified":"2025-11-22T05:44:03.852Z","1.0.0":"2025-11-16T14:18:14.521Z","1.0.1":"2025-11-16T14:24:26.785Z","1.0.2":"2025-11-22T05:44:03.677Z"},"bugs":{"url":"https://github.com/anisirji/llm-web-extractor/issues"},"author":{"name":"anisirji"},"license":"MIT","homepage":"https://github.com/anisirji/llm-web-extractor#readme","keywords":["web-scraping","content-extraction","url-parser","firecrawl","scraper","termix"],"repository":{"type":"git","url":"git+https://github.com/anisirji/llm-web-extractor.git"},"description":"Powerful web content extraction SDK with URL normalization and intelligent scraping - https://github.com/anisirji/llm-web-extractor","maintainers":[{"name":"anisirji","email":"www.allinoneforyouhere@gmail.com"}],"readme":"# @anisirji/web-extractor\n\n> Powerful web content extraction SDK with intelligent URL handling, content cleaning, and comprehensive metadata extraction.\n\n## Features\n\n✨ **Smart URL Handling**\n- URL validation and normalization\n- Subdomain detection\n- Duplicate URL filtering\n- Pattern-based URL filtering\n\n🧹 **Content Cleaning**\n- Automatic markdown/HTML/text extraction\n- Whitespace normalization\n- Word counting\n- Language detection\n\n📊 **Rich Metadata**\n- Scraping timestamps\n- Word counts\n- Page descriptions\n- Status codes\n- Custom metadata support\n\n🚀 **Easy to Use**\n- Simple, intuitive API\n- TypeScript support\n- Promise-based\n- Comprehensive error handling\n\n## 📖 Documentation\n\n- **[Testing Guide](docs/TESTING_GUIDE.md)** - Comprehensive guide on testing the SDK\n- **[Test Results](TEST_RESULTS.md)** - Latest test results for astratechai.com\n- **[API Documentation](docs/README.md)** - Complete API reference and examples\n\n## Installation\n\n```bash\nnpm install @anisirji/web-extractor\n```\n\n## Quick Start\n\n### Extract a Single Page\n\n```typescript\nimport { WebExtractor } from '@anisirji/web-extractor';\n\nconst extractor = new WebExtractor({\n  apiKey: 'your-firecrawl-api-key'\n});\n\n// Extract single page\nconst page = await extractor.extractPage('https://example.com');\n\nconsole.log(page.title);\nconsole.log(page.content);\nconsole.log(page.metadata.wordCount);\n```\n\n### Extract Entire Website\n\n```typescript\nconst result = await extractor.extractWebsite('https://docs.example.com', {\n  maxPages: 20,\n  includeSubdomains: false,\n  titlePrefix: 'Docs',\n  maxDepth: 3\n});\n\nconsole.log(`Extracted ${result.pages.length} pages`);\nconsole.log(`Success rate: ${result.stats.successRate}%`);\nconsole.log(`Total words: ${result.stats.totalWords}`);\n\nfor (const page of result.pages) {\n  console.log(`${page.title} - ${page.url}`);\n}\n```\n\n## API Reference\n\n### WebExtractor\n\nMain class for web extraction.\n\n#### Constructor\n\n```typescript\nnew WebExtractor(config: WebExtractorConfig)\n```\n\n**Config Options:**\n- `apiKey` (required): Your Firecrawl API key\n- `baseUrl` (optional): Custom Firecrawl API URL\n- `timeout` (optional): Request timeout in ms (default: 30000)\n- `debug` (optional): Enable debug logging (default: false)\n\n#### Methods\n\n##### extractPage(url, options?)\n\nExtract content from a single page.\n\n```typescript\nawait extractor.extractPage('https://example.com', {\n  onlyMainContent: true,  // Extract only main content\n  format: 'markdown',     // 'markdown' | 'html' | 'text'\n  waitFor: 1000          // Wait time before extraction (ms)\n});\n```\n\n**Returns:** `Promise<ExtractedPage>`\n\n##### extractWebsite(url, options?)\n\nExtract content from entire website (crawl).\n\n```typescript\nawait extractor.extractWebsite('https://example.com', {\n  maxPages: 10,                    // Maximum pages to scrape\n  includeSubdomains: false,        // Include subdomains\n  titlePrefix: 'My Site',          // Prefix for all titles\n  maxDepth: 3,                     // Maximum crawl depth\n  followExternalLinks: false,      // Follow external links\n  includePatterns: [/\\/docs\\//],   // URL patterns to include\n  excludePatterns: [/\\/blog\\//],   // URL patterns to exclude\n  onlyMainContent: true,           // Extract only main content\n  format: 'markdown'               // Output format\n});\n```\n\n**Returns:** `Promise<ExtractionResult>`\n\n### URL Utilities\n\nPowerful URL manipulation utilities.\n\n```typescript\nimport {\n  normalizeUrl,\n  validateUrl,\n  deduplicateUrls,\n  isSameDomain,\n  extractDomain\n} from '@anisirji/web-extractor';\n\n// Normalize URL\nconst normalized = normalizeUrl('https://Example.com/path/?b=2&a=1#hash', {\n  lowercase: true,           // Convert to lowercase\n  removeTrailingSlash: true, // Remove trailing slash\n  removeFragment: true,      // Remove #hash\n  sortQueryParams: true      // Sort query params\n});\n// => 'https://example.com/path?a=1&b=2'\n\n// Validate URL\nconst urlObj = validateUrl('https://example.com'); // Returns URL object or throws\n\n// Deduplicate URLs\nconst unique = deduplicateUrls([\n  'https://example.com/page',\n  'https://example.com/page/',\n  'https://EXAMPLE.COM/page'\n]);\n// => ['https://example.com/page']\n\n// Check same domain\nisSameDomain('https://example.com', 'https://example.com/page'); // true\nisSameDomain('https://example.com', 'https://other.com'); // false\n\n// Extract domain\nextractDomain('https://blog.example.com/page'); // => 'blog.example.com'\n```\n\n### Content Utilities\n\nContent processing utilities.\n\n```typescript\nimport {\n  cleanContent,\n  countWords,\n  generateExcerpt,\n  detectLanguage\n} from '@anisirji/web-extractor';\n\n// Clean content\nconst cleaned = cleanContent('  text\\n\\n\\n\\nmore text  ');\n// => 'text\\n\\nmore text'\n\n// Count words\ncountWords('Hello world from TermiX'); // => 4\n\n// Generate excerpt\ngenerateExcerpt('Very long content here...', 10);\n// => 'Very long content here (first 10 words)...'\n\n// Detect language\ndetectLanguage('This is an English text'); // => 'en'\n```\n\n## Advanced Examples\n\n### Filter URLs by Pattern\n\n```typescript\nconst result = await extractor.extractWebsite('https://docs.example.com', {\n  maxPages: 50,\n  // Only include documentation pages\n  includePatterns: [\n    /\\/docs\\//,\n    /\\/api\\//,\n    /\\/guides\\//\n  ],\n  // Exclude blog and changelog\n  excludePatterns: [\n    /\\/blog\\//,\n    /\\/changelog\\//\n  ]\n});\n```\n\n### Custom Processing Pipeline\n\n```typescript\nconst result = await extractor.extractWebsite('https://example.com', {\n  maxPages: 30\n});\n\n// Filter by word count\nconst substantialPages = result.pages.filter(\n  page => page.metadata.wordCount > 500\n);\n\n// Group by language\nconst byLanguage = result.pages.reduce((acc, page) => {\n  const lang = page.metadata.language || 'unknown';\n  acc[lang] = acc[lang] || [];\n  acc[lang].push(page);\n  return acc;\n}, {});\n\n// Calculate reading time\nconst withReadingTime = result.pages.map(page => ({\n  ...page,\n  readingTimeMinutes: Math.ceil(page.metadata.wordCount / 200)\n}));\n```\n\n### Batch Processing with Error Handling\n\n```typescript\nconst urls = [\n  'https://example.com/page1',\n  'https://example.com/page2',\n  'https://example.com/page3'\n];\n\nconst results = await Promise.allSettled(\n  urls.map(url => extractor.extractPage(url))\n);\n\nconst successful = results\n  .filter(r => r.status === 'fulfilled')\n  .map(r => r.value);\n\nconst failed = results\n  .filter(r => r.status === 'rejected')\n  .map((r, i) => ({ url: urls[i], error: r.reason }));\n\nconsole.log(`Success: ${successful.length}, Failed: ${failed.length}`);\n```\n\n## Types\n\n### ExtractedPage\n\n```typescript\ninterface ExtractedPage {\n  title: string;\n  content: string;\n  url: string;\n  metadata: PageMetadata;\n}\n```\n\n### PageMetadata\n\n```typescript\ninterface PageMetadata {\n  scrapedAt: Date;\n  sourceUrl: string;\n  description?: string;\n  wordCount: number;\n  language?: string;\n  statusCode?: number;\n  [key: string]: any;  // Custom metadata\n}\n```\n\n### ExtractionResult\n\n```typescript\ninterface ExtractionResult {\n  pages: ExtractedPage[];\n  totalPages: number;\n  failed: FailedExtraction[];\n  stats: ExtractionStats;\n}\n```\n\n### ExtractionStats\n\n```typescript\ninterface ExtractionStats {\n  duration: number;           // Total time in ms\n  successRate: number;        // Success rate %\n  totalWords: number;         // Total words extracted\n  avgWordsPerPage: number;    // Average words per page\n}\n```\n\n## Use Cases\n\n- 📚 **Documentation Scraping**: Extract and index documentation sites\n- 🧠 **Knowledge Base Building**: Build AI knowledge bases from websites\n- 🔍 **Content Analysis**: Analyze website content and structure\n- 📊 **SEO Analysis**: Extract metadata for SEO analysis\n- 🤖 **AI Training Data**: Collect training data for AI models\n- 📝 **Content Migration**: Migrate content from old to new sites\n\n## Requirements\n\n- Node.js >= 16\n- Firecrawl API key ([Get one here](https://firecrawl.dev))\n\n## Testing\n\nRun the comprehensive test suite:\n\n```bash\n# Unit tests\nnpm test\n\n# Integration test with astratechai.com\nnpm run test:astratechai\n\n# Basic usage example\nnpm run test:integration\n```\n\nSee [Testing Guide](docs/TESTING_GUIDE.md) for detailed instructions on creating tests for your own websites.\n\n## License\n\nMIT\n\n## Repository\n\n- 📖 [GitHub](https://github.com/anisirji/llm-web-extractor)\n- 🐛 [Report Issues](https://github.com/anisirji/llm-web-extractor/issues)\n- 📦 [NPM Package](https://www.npmjs.com/package/@anisirji/web-extractor)\n- 📚 [Documentation](docs/README.md)\n- 🧪 [Test Results](TEST_RESULTS.md)\n\n---\n\nBuilt with ❤️ by anisirji\n","readmeFilename":"README.md"}