{"_id":"@constellix/ai-scraper-mcp","_rev":"2-c25fcdfd9e035f22c01692932e92ed71","name":"@constellix/ai-scraper-mcp","dist-tags":{"latest":"1.0.1"},"versions":{"1.0.0":{"name":"@constellix/ai-scraper-mcp","version":"1.0.0","keywords":["ai scraper","web scraping","ai crawler","ai","mcp server"],"author":{"name":"Shashank Gupta","email":"shashankrg.2710@gmail.com"},"license":"MIT","_id":"@constellix/ai-scraper-mcp@1.0.0","maintainers":[{"name":"shashank_rg","email":"shashankrg.2710@gmail.com"}],"homepage":"https://github.com/shAsh8bit/constellix-ai-scraper-mcp#readme","bugs":{"url":"https://github.com/shAsh8bit/constellix-ai-scraper-mcp/issues"},"bin":{"ai-scraper-mcp":"server.js"},"dist":{"shasum":"a36942fde7e94328bbe5344a0d8ec437ae0184d3","tarball":"https://registry.npmjs.org/@constellix/ai-scraper-mcp/-/ai-scraper-mcp-1.0.0.tgz","fileCount":5,"integrity":"sha512-7ibJtLDXdOTYmBVMmC+hd0fTv9lY9G5nJF2PAp8A72M/r/jDzg0WdH0LYCBkCfAoi5b10FbS7eOGX/6DhczJSg==","signatures":[{"sig":"MEUCIQC29iK4QHJFxdKCN66/lLuvs2fuzCJdqbaOkz61NQlBaQIgdg8FoiPRWxFZXLf5KnoWQdFiLtNJwyQFq4PilJE1gpM=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":13898},"main":"server.js","type":"module","scripts":{"test":"echo \"Error: no test specified\" && exit 1","build":"echo \"No build required\" && exit 1","start":"node server.js"},"_npmUser":{"name":"shashank_rg","email":"shashankrg.2710@gmail.com"},"repository":{"url":"git+https://github.com/shAsh8bit/constellix-ai-scraper-mcp.git","type":"git"},"_npmVersion":"10.9.2","description":"AI-powered web scraping MCP server for structured data extraction","directories":{},"_nodeVersion":"23.7.0","dependencies":{"dotenv":"^16.5.0","playwright":"^1.52.0","@constellix/ai-scraper":"^0.1.3-beta.2","@modelcontextprotocol/sdk":"^1.11.4"},"_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/ai-scraper-mcp_1.0.0_1748168848526_0.516174795923596","host":"s3://npm-registry-packages-npm-production"}},"1.0.1":{"name":"@constellix/ai-scraper-mcp","version":"1.0.1","main":"server.js","keywords":["ai scraper","web scraping","ai crawler","ai","mcp server"],"author":{"name":"Shashank Gupta","email":"shashankrg.2710@gmail.com"},"license":"MIT","repository":{"type":"git","url":"git+https://github.com/shAsh8bit/constellix-ai-scraper-mcp.git"},"scripts":{"test":"echo \"Error: no test specified\" && exit 1","build":"echo \"No build required\" && exit 1","start":"node server.js"},"description":"AI-powered web scraping MCP server for structured data extraction","type":"module","dependencies":{"@constellix/ai-scraper":"^0.1.4","@modelcontextprotocol/sdk":"^1.11.4","dotenv":"^16.5.0","playwright":"^1.52.0"},"bin":{"ai-scraper-mcp":"server.js"},"_id":"@constellix/ai-scraper-mcp@1.0.1","gitHead":"2a2b7620aa7e95102e7e4a493a9485581d66ecf2","bugs":{"url":"https://github.com/shAsh8bit/constellix-ai-scraper-mcp/issues"},"homepage":"https://github.com/shAsh8bit/constellix-ai-scraper-mcp#readme","_nodeVersion":"23.7.0","_npmVersion":"10.9.2","dist":{"integrity":"sha512-wP8MwifweI5YAVAOJeCPRRNTdGE9c3U2YfydirTGadXfOoyc/LutMSAE+BlBQPmJYqhOVTq+LK8DqHJqO7g+1A==","shasum":"b87cd758516e91adc8d762f8fd141c922e84178f","tarball":"https://registry.npmjs.org/@constellix/ai-scraper-mcp/-/ai-scraper-mcp-1.0.1.tgz","fileCount":5,"unpackedSize":13875,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQDzUI/BPS995cvgUnfyMyblRL4+Z13nD7JK90eV+ht38gIhAMQPLL+d369EnzSlQhx+3Mi31IRR0GmuFgs/SqxDtuac"}]},"_npmUser":{"name":"shashank_rg","email":"shashankrg.2710@gmail.com"},"directories":{},"maintainers":[{"name":"shashank_rg","email":"shashankrg.2710@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/ai-scraper-mcp_1.0.1_1748414179431_0.8330705749776814"},"_hasShrinkwrap":false}},"time":{"created":"2025-05-25T10:27:28.449Z","modified":"2025-05-28T06:36:19.791Z","1.0.0":"2025-05-25T10:27:28.737Z","1.0.1":"2025-05-28T06:36:19.611Z"},"bugs":{"url":"https://github.com/shAsh8bit/constellix-ai-scraper-mcp/issues"},"author":{"name":"Shashank Gupta","email":"shashankrg.2710@gmail.com"},"license":"MIT","homepage":"https://github.com/shAsh8bit/constellix-ai-scraper-mcp#readme","keywords":["ai scraper","web scraping","ai crawler","ai","mcp server"],"repository":{"type":"git","url":"git+https://github.com/shAsh8bit/constellix-ai-scraper-mcp.git"},"description":"AI-powered web scraping MCP server for structured data extraction","maintainers":[{"name":"shashank_rg","email":"shashankrg.2710@gmail.com"}],"readme":"# @constellix/ai-scraper-mcp\n\nA Model Context Protocol (MCP) server for data extraction using AI to extract structured data from web pages. This tool bridges the gap between LLM and web data extraction by providing an intelligent interface for scraping websites.\n\n## Live\nTry playground → https://constellix.vercel.app/\n\n## Features\n\n- **AI-Powered Data Extraction**: Extract structured data from web pages using natural language queries\n- **CSS Selector Generation**: Generate CSS selectors for web elements based on natural language descriptions\n- **XPath Generation**: Generate XPath expressions for web elements based on natural language descriptions\n- **Supports Multiple Query Types**: Use either natural language or structured GraphQL-like queries\n\n## Installation\n\n```bash\n# Install and run\nnpm i @constellix/ai-scraper-mcp\n\n# Set your API key as an environment variable\nGEMINI_API_KEY=\"your-api-key-here\"\n```\n\nMCP configurations:\n```json\n{\n    \"mcpServers\": {\n        \"ai-scraper\":{\n            \"command\": \"npx\",\n            \"args\": [\n                \"-y\",\n                \"@constellix/ai-scraper-mcp\"\n            ],\n            \"env\": {\n                \"GEMINI_API_KEY\" : \"YOUR_API_KEY\"\n            }\n        }\n    }\n}\n```\nThen in your MCP-compatible client (Claude, Cursor, etc.), you can use the ai-scraper tools to extract data from websites.\n\n\n## Available Tools\n\n### 1. get-data-by-query\n\nExtracts structured data from a webpage using natural language or structured query language.\n\n**Input Schema:**\n```\n{\n  \"url\": \"string\", // The webpage URL to extract data from\n  \"query\": \"string\" // Natural language query or structured query\n}\n```\n\n### 2. get-css-selector\n\nGenerates CSS selectors for webpage elements using natural language or structured query language.\n\n**Input Schema:**\n```\n{\n  \"url\": \"string\", // The webpage URL to analyze\n  \"query\": \"string\" // Natural language query or structured query\n}\n```\n\n### 3. get-xpath\n\nGenerates XPath expressions for webpage elements using natural language or structured query language.\n\n**Input Schema:**\n```\n{\n  \"url\": \"string\", // The webpage URL to analyze\n  \"query\": \"string\" // Natural language query or structured query\n}\n```\n\n## Query Types\n\n### Natural Language Queries\n\nExamples:\n- \"List all the products on the page\"\n- \"Find the main navigation menu\"\n- \"Extract all blog post titles and their publication dates\"\n\n### Structured Queries (GraphQL-like)\n\n```\n{\n  products_list[]{\n    product_name,\n    product_price,\n    product_image\n  }\n}\n```\n\nYou can also specify data types or add natural language descriptions:\n\n```\n{\n  products_list[]{\n    product_name (string),\n    product_price (number),\n    product_image (string)\n  }\n}\n```\n\nOr with descriptions:\n\n```\n{\n  products_list (products made out of cotton)[]{\n    product_name,\n    product_price,\n    product_image\n  }\n}\n```\n\n## Dependencies\n\nThis package relies on the `@constellix/ai-scraper` package, which provides capabilities for enhancing Playwright's functionality with AI capabilities.\n\n\n","readmeFilename":"README.md"}