{"_id":"@abilashinamdar/node-crawler","_rev":"3-87ad00f1cc1fc9babcae89a259f3a50a","name":"@abilashinamdar/node-crawler","dist-tags":{"latest":"1.0.3"},"versions":{"1.0.0":{"name":"@abilashinamdar/node-crawler","version":"1.0.0","description":"Fast asynchronous NodeJS module for crawling/scraping a web through worker_threads.","main":"main.js","engine-strict":{"node":">=11.7.0"},"dependencies":{"axios":"^1.4.0","cheerio":"^1.0.0-rc.12","robots-parser":"^3.0.1"},"repository":{"type":"git","url":"git+https://github.com/Abilashinamdar/node-crawler.git"},"homepage":"https://github.com/Abilashinamdar/node-crawler#readme","keywords":["crawler","scraper","crawling","spider","node-js-crawler","node-spider","dom","scraping","nodejs"],"type":"module","author":{"name":"Abilash inamdar","email":"abiinamdar9294@gmail.com"},"license":"ISC","_id":"@abilashinamdar/node-crawler@1.0.0","gitHead":"5c6841f5218747277af34483b7652ec83ea9f911","bugs":{"url":"https://github.com/Abilashinamdar/node-crawler/issues"},"_nodeVersion":"18.16.1","_npmVersion":"9.8.0","dist":{"integrity":"sha512-u//gaKGZ/qhjwfge53t45/vPEdTajjI5d/5e1F8QM/PYHIC1PwmWsbj0p0WYa2AwWR0tLobgUsockdt+O5555w==","shasum":"5c7d1e875f3e88ed42eaa566fc25352e350c999c","tarball":"https://registry.npmjs.org/@abilashinamdar/node-crawler/-/node-crawler-1.0.0.tgz","fileCount":5,"unpackedSize":16225,"signatures":[{"keyid":"SHA256:jl3bwswu80PjjokCgh0o2w5c2U4LhQAE57gj9cz1kzA","sig":"MEUCIHi3VeQimEdt1fGKZi8/QPBCyBs/BgTzei9RkLOKen71AiEAhUFpAu/CCf2P6UI5QDFCXCvOBcRSq/qpXj1vJ2emN1w="}]},"_npmUser":{"name":"abilash_inamdar","email":"abiinamdar9294@gmail.com"},"directories":{},"maintainers":[{"name":"abilash_inamdar","email":"abiinamdar9294@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages","tmp":"tmp/node-crawler_1.0.0_1695255873434_0.17233155683964663"},"_hasShrinkwrap":false},"1.0.1":{"name":"@abilashinamdar/node-crawler","version":"1.0.1","description":"Fast asynchronous NodeJS module for crawling/scraping a web through worker_threads.","main":"main.js","engine-strict":{"node":">=11.7.0"},"dependencies":{"axios":"^1.4.0","cheerio":"^1.0.0-rc.12","robots-parser":"^3.0.1"},"repository":{"type":"git","url":"git+https://github.com/Abilashinamdar/node-crawler.git"},"homepage":"https://github.com/Abilashinamdar/node-crawler#readme","keywords":["crawler","scraper","crawling","spider","node-js-crawler","node-spider","dom","scraping","nodejs"],"type":"module","author":{"name":"Abilash inamdar","email":"abiinamdar9294@gmail.com"},"license":"ISC","_id":"@abilashinamdar/node-crawler@1.0.1","gitHead":"5c6841f5218747277af34483b7652ec83ea9f911","bugs":{"url":"https://github.com/Abilashinamdar/node-crawler/issues"},"_nodeVersion":"18.16.1","_npmVersion":"9.8.0","dist":{"integrity":"sha512-W4lAVYM1yqgCCCxbgC4wTQHUCilWkUnBcbZFrMX5X5gtN3DC0WRY4uOjmsvGffV/AhhsPV7eQs6llv+T8AnHBg==","shasum":"ba019de757bca6ce37bf1aaa1a5aa0607b5d8f9a","tarball":"https://registry.npmjs.org/@abilashinamdar/node-crawler/-/node-crawler-1.0.1.tgz","fileCount":5,"unpackedSize":16684,"signatures":[{"keyid":"SHA256:jl3bwswu80PjjokCgh0o2w5c2U4LhQAE57gj9cz1kzA","sig":"MEYCIQCBcwUh3ULGr4gV4uObMEKt3qMd/dIsa6x1VsQtZpMVDwIhAPI4HosyoQyBLtcdpPNzMu+sn6/7akUedWqnTHk6sGb+"}]},"_npmUser":{"name":"abilash_inamdar","email":"abiinamdar9294@gmail.com"},"directories":{},"maintainers":[{"name":"abilash_inamdar","email":"abiinamdar9294@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages","tmp":"tmp/node-crawler_1.0.1_1695257501896_0.6591627230435113"},"_hasShrinkwrap":false},"1.0.2":{"name":"@abilashinamdar/node-crawler","version":"1.0.2","description":"Fast asynchronous NodeJS module for crawling/scraping a web through worker_threads.","main":"main.js","engine-strict":{"node":">=11.7.0"},"dependencies":{"axios":"^1.4.0","cheerio":"^1.0.0-rc.12","robots-parser":"^3.0.1"},"repository":{"type":"git","url":"git+https://github.com/Abilashinamdar/node-crawler.git"},"homepage":"https://github.com/Abilashinamdar/node-crawler#readme","keywords":["crawler","scraper","crawling","spider","node-js-crawler","node-spider","dom","scraping","nodejs"],"type":"module","author":{"name":"Abilash inamdar","email":"abiinamdar9294@gmail.com"},"license":"ISC","_id":"@abilashinamdar/node-crawler@1.0.2","gitHead":"5c6841f5218747277af34483b7652ec83ea9f911","bugs":{"url":"https://github.com/Abilashinamdar/node-crawler/issues"},"_nodeVersion":"18.16.1","_npmVersion":"9.8.0","dist":{"integrity":"sha512-hfKFjTPuBhdN+b9TNWy0Omy5kmtnEj9qbH6z5e00lclOKZtQ5TgU7UWP1/SU0/23FkoABaxkWP4WteD855UuOQ==","shasum":"c1e514f5b5d28d9686129aff0049119fc3fe46fd","tarball":"https://registry.npmjs.org/@abilashinamdar/node-crawler/-/node-crawler-1.0.2.tgz","fileCount":5,"unpackedSize":16736,"signatures":[{"keyid":"SHA256:jl3bwswu80PjjokCgh0o2w5c2U4LhQAE57gj9cz1kzA","sig":"MEUCIQCxlKWPqbnAUwsrFUqfKVk7VHT6vK+fdxmxtKADm271NQIgKZcIxZ7VTv4Y9FsEApHVwemAJHxx/6yN/mIjSrxEEaU="}]},"_npmUser":{"name":"abilash_inamdar","email":"abiinamdar9294@gmail.com"},"directories":{},"maintainers":[{"name":"abilash_inamdar","email":"abiinamdar9294@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages","tmp":"tmp/node-crawler_1.0.2_1695257996117_0.18291588934486303"},"_hasShrinkwrap":false},"1.0.3":{"name":"@abilashinamdar/node-crawler","version":"1.0.3","description":"Fast asynchronous NodeJS module for crawling/scraping a web through worker_threads.","main":"main.js","engine-strict":{"node":">=11.7.0"},"dependencies":{"axios":"^1.4.0","cheerio":"^1.0.0-rc.12","robots-parser":"^3.0.1"},"repository":{"type":"git","url":"git+https://github.com/Abilashinamdar/node-crawler.git"},"homepage":"https://node-crawler-server.glitch.me","keywords":["crawler","scraper","crawling","spider","node-js-crawler","node-spider","dom","scraping","nodejs"],"type":"module","author":{"name":"Abilash inamdar","email":"abiinamdar9294@gmail.com"},"license":"ISC","_id":"@abilashinamdar/node-crawler@1.0.3","gitHead":"2b6a01d6e94ae1ac30bc8baa02b28f60bdf20dcd","bugs":{"url":"https://github.com/Abilashinamdar/node-crawler/issues"},"_nodeVersion":"18.16.1","_npmVersion":"9.8.0","dist":{"integrity":"sha512-Js+E9oqM1z/dQQWEBGrtBxojfiI4S2xyDa2uBY5RH2yUKt8GM+JeKCmnv8jQb6nBthcZniUXHX5l4l1RNO7Tmg==","shasum":"5d07b2b7f599e7b2b2c22eed8523aea411c4aefa","tarball":"https://registry.npmjs.org/@abilashinamdar/node-crawler/-/node-crawler-1.0.3.tgz","fileCount":7,"unpackedSize":21246,"signatures":[{"keyid":"SHA256:jl3bwswu80PjjokCgh0o2w5c2U4LhQAE57gj9cz1kzA","sig":"MEQCIEGFcjAmy60qHe91F2wxHG5/hmWdIGcrsEPKmJoEPlXxAiBnpbNGF5itRYlVTmbuB75fh3I1QaFRZiWgTHd2a03stA=="}]},"_npmUser":{"name":"abilash_inamdar","email":"abiinamdar9294@gmail.com"},"directories":{},"maintainers":[{"name":"abilash_inamdar","email":"abiinamdar9294@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages","tmp":"tmp/node-crawler_1.0.3_1695644431368_0.8916820252255029"},"_hasShrinkwrap":false}},"time":{"created":"2023-09-21T00:24:33.357Z","1.0.0":"2023-09-21T00:24:33.607Z","modified":"2023-09-25T12:20:31.805Z","1.0.1":"2023-09-21T00:51:42.094Z","1.0.2":"2023-09-21T00:59:56.335Z","1.0.3":"2023-09-25T12:20:31.546Z"},"maintainers":[{"name":"abilash_inamdar","email":"abiinamdar9294@gmail.com"}],"description":"Fast asynchronous NodeJS module for crawling/scraping a web through worker_threads.","homepage":"https://node-crawler-server.glitch.me","keywords":["crawler","scraper","crawling","spider","node-js-crawler","node-spider","dom","scraping","nodejs"],"repository":{"type":"git","url":"git+https://github.com/Abilashinamdar/node-crawler.git"},"author":{"name":"Abilash inamdar","email":"abiinamdar9294@gmail.com"},"bugs":{"url":"https://github.com/Abilashinamdar/node-crawler/issues"},"license":"ISC","readme":"# node-crawler\r\n\r\n> Fast asynchronous NodeJS module for crawling/scraping a web through worker_threads.\r\n\r\n[![npm package](https://nodei.co/npm/@abilashinamdar/node-crawler.png)](https://www.npmjs.com/package/@abilashinamdar/node-crawler)\r\n\r\n\r\n[![Version][version-image]][download-url]\r\n[![License][license-image]][download-url]\r\n\r\n[version-image]: https://img.shields.io/npm/v/@abilashinamdar/node-crawler.svg\r\n[license-image]: https://img.shields.io/npm/l/@abilashinamdar/node-crawler.svg\r\n[download-url]: https://npmjs.com/package/@abilashinamdar/node-crawler\r\n\r\n### Features\r\n- Crawling on threads(CPU cores)\r\n- Crawl images and links\r\n- Configurable max threshold crawl\r\n- Configurable retries\r\n- Configurable request and retry delays\r\n- Rotate proxies\r\n- Bypass robots.txt\r\n\r\n### Demo: https://node-crawler-server.glitch.me\r\n\r\n# Get started\r\n\r\n## Install\r\n```sh\r\nnpm i @abilashinamdar/node-crawler\r\n\r\n    OR\r\n\r\nyarn add @abilashinamdar/node-crawler\r\n```\r\n\r\n\r\n# Basic usage\r\n\r\n```js\r\nimport crawl from \"@abilashinamdar/node-crawler\";\r\n\r\ncrawl(config, onEveryCrawl, onCrawlError, onCrawlComplete)\r\n```\r\n\r\n\r\n### config\r\nKey | Type | Default | Value\r\n--- | --- | --- | ---\r\nurl | String | - | \r\ncrawl | Array<String> | [\"links\"] | [\"links\", \"images\"]\r\nmaxCrawl | Number | 0 | Acts as threshold value. Stops crawling once it reaches the maxCrawl. 0 represents no threshold.\r\nmaxRetries | Number | 0 | No. of times the request should be retried when fails.\r\nallowExternalImages| Boolean | false | true/false. true allows crawler to pick image if it points to outside the origin/host\r\nretryDelay| Number | 0 | in milliseconds. Delay between the retry requests.\r\nuseProxy| Boolean | false | true/false. true allows crawler to rotate some free proxies from [https://www.free-proxy-list.com](https://www.free-proxy-list.com)\r\nproxies| Array<String> | [] | Crawler will rotate the proxies from the list rather than fetching free proxies. useProxy: true is mandatory.\r\nskipRobotsFile| Boolean | false | true/false. true allows crawler to bypass the robots.txt check.\r\nrequestDelay| Number | 0 | in milliseconds. Delay between the concurrent requests.\r\n\r\n\r\n### onEveryCrawl(value)\r\n```js\r\nonEveryCrawl will get triggered when crawler will crawl Link or Image.\r\nIt will be triggered with the value as object which is as follows:\r\n{\r\n    type: string,\r\n    link: string,\r\n    reason: string\r\n}\r\n\r\ntype - It can be 'LINK' / 'IMAGE' / 'IGNORE'\r\nreason - Description for the IGNORE type.\r\n```\r\n\r\n\r\n### onCrawlError(error)\r\n```js\r\nonCrawlError will get triggered when crawler get error while crawling.\r\nIt will be triggered with the error as sample object as below or JS Error object.\r\n\r\n{\r\n    type: string,\r\n    link: string,\r\n    reason: string\r\n}\r\n\r\ntype - It will be 'DISALLOWED'\r\nreason - Description for the DISALLOWED type. Basically occurs when robots.txt disallows the url.\r\n```\r\n\r\n\r\n### onCrawlComplete(value)\r\n```js\r\nonCrawlComplete will get triggered when crawler done with the crawling.\r\nIt will be triggered with the value as object which is as follows:\r\n\r\n{\r\n    links: Set<String>,\r\n    images: Set<String>,\r\n    ignored: Set<String>\r\n}\r\n\r\nlinks - List of links from the crawled url.\r\nimages - List of image links from the crawled url.\r\nignored - List of ignored links from the crawled url.\r\n\r\nNote: After this trigger the thread will get exit and will be available for others.\r\n```\r\n\r\n\r\n# License\r\nISC\r\n","readmeFilename":"README.md"}