{"_id":"@bluggie/nodescrapy","_rev":"6-0f917385a7767067967be1926e05d21e","name":"@bluggie/nodescrapy","dist-tags":{"latest":"0.1.6"},"versions":{"0.1.0":{"name":"@bluggie/nodescrapy","version":"0.1.0","description":"Web crawler in NodeJS","main":"dist/index.js","scripts":{"build":"tsc","lint":"eslint . --ext .ts","test":"jest --collect-coverage"},"keywords":["scrapy","web-crawling","web-scraping","nodejs","typescript"],"author":{"name":"Juan Roldan","email":"juan.roldan.brz@gmail.com"},"license":"MIT","devDependencies":{"@types/jest":"^28.1.4","@typescript-eslint/eslint-plugin":"^5.30.4","@typescript-eslint/parser":"^5.30.4","axios-mock-adapter":"^1.21.1","coveralls":"^3.1.1","eslint":"^8.19.0","eslint-config-airbnb-base":"^15.0.0","eslint-plugin-import":"^2.26.0","eslint-plugin-jest":"^26.5.3","jest":"^28.1.2","ts-jest":"^28.0.5","ts-node":"^10.8.2","tsconfig-paths":"^4.0.0","typescript":"^4.5.2"},"dependencies":{"app-root-path":"^3.0.0","axios":"^0.27.2","axios-retry":"^3.3.1","cheerio":"^1.0.0-rc.12","promise-throttle":"^1.1.2","sequelize":"^6.21.2","sqlite3":"^5.0.8","uuid":"^8.3.2","winston":"^3.8.1"},"gitHead":"34e893a57b7fddba5d403df5798979a557a781ef","_id":"@bluggie/nodescrapy@0.1.0","_nodeVersion":"18.1.0","_npmVersion":"8.8.0","dist":{"integrity":"sha512-gFOvYeuOPStDMQcJ658suGFtezkTIXR3gBr8fUKqP3xVdn0jaE1TCCbFyIvx+I+Tziu+XKYfLhCnuMj3mhvkYA==","shasum":"823322d6e57e8a9b0ed02c7b908884208c9ff8cb","tarball":"https://registry.npmjs.org/@bluggie/nodescrapy/-/nodescrapy-0.1.0.tgz","fileCount":33,"unpackedSize":124386,"signatures":[{"keyid":"SHA256:jl3bwswu80PjjokCgh0o2w5c2U4LhQAE57gj9cz1kzA","sig":"MEUCIAKa9u3THcUdrIxYJcWnW53HmEWt3v52pUMIcTuljDq7AiEAuyYzjuYPXg79Xqrt3hfQtN1HdMQzfWvz6SMj9UE81KU="}],"npm-signature":"-----BEGIN PGP SIGNATURE-----\r\nVersion: OpenPGP.js v4.10.10\r\nComment: https://openpgpjs.org\r\n\r\nwsFzBAEBCAAGBQJiztrhACEJED1NWxICdlZqFiEECWMYAoorWMhJKdjhPU1b\r\nEgJ2VmruwBAAibiPaJ/DfjY7sih4UsIDJ3kltZ6gPJwO7+eEscCsafogujej\r\njXA29cHmK30y5Ss1MnISETJATg7kL9rOUoKlhos8j7SnifW1zIi7cgwi7yGk\r\nHhcOw/i60e6sZlfkD3jRfpaG+zFZIAQPylmirPXTDRyUOfoNct+KRQcm2OO7\r\nlBr63vmMwFRxAgsJoQEy4XlP5IhmMFzMpvOXJwSAiwqdfk/Fsj9V+NOv/ySd\r\nfXxoF6pC8O6AEewcReWAxDYvBD+WOW2qCanMXHZzdjy5HCvyzU7oCuFbzwC2\r\njpmyMmX4KdOcwPYORv/k79+6gT+QnOMt+C1Vns6eViy+CIpVDkxOnJ4eteg1\r\njtEFVWVMDtOYxzuYvEOPEBUs9N3dUb3GoI7ZnPBP4GRw2rqB1nL4tYO4Dlkk\r\nkpgLyhFdM0MMAamJ5xZsNY2suiUP6H/9bZK2dQm6PGzN4V4Uwjuo/6X5cCiz\r\nsMKeL7saKzv4olOhRGgcj3dOWqD2AR2Q39fiFiVrPkGpDqk9LxHYGsXawI1K\r\noAMpPAmlHdvBGL18vJ1uguL6yRSCoFuy9zvfG3P9tVdjPn2CP2Y+daqSKaYt\r\n23El4zBgYq7Smk1eufG99Wnd0Ho4Cr3ppJ3S1sUcg6A4qPaqetUOrGbuIqB5\r\nepJGywAIiPbh29zZt5pDnCT1OH5WGzHvKmw=\r\n=M1wU\r\n-----END PGP SIGNATURE-----\r\n"},"_npmUser":{"name":"bluggie","email":"juan.roldan@bluggie.com"},"directories":{},"maintainers":[{"name":"bluggie","email":"juan.roldan@bluggie.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages","tmp":"tmp/nodescrapy_0.1.0_1657723617203_0.28924092210063734"},"_hasShrinkwrap":false},"0.1.1":{"name":"@bluggie/nodescrapy","version":"0.1.1","description":"Web crawler in NodeJS","main":"dist/index.js","scripts":{"build":"tsc","lint":"eslint . --ext .ts","test":"jest --collect-coverage"},"repository":{"type":"git","url":"git+https://github.com/juanroldanbrz/nodescrapy.git"},"keywords":["scrapy","web-crawling","web-scraping","nodejs","typescript"],"author":{"name":"Juan Roldan","email":"juan.roldan.brz@gmail.com"},"license":"MIT","devDependencies":{"@types/jest":"^28.1.4","@typescript-eslint/eslint-plugin":"^5.30.4","@typescript-eslint/parser":"^5.30.4","axios-mock-adapter":"^1.21.1","coveralls":"^3.1.1","eslint":"^8.19.0","eslint-config-airbnb-base":"^15.0.0","eslint-plugin-import":"^2.26.0","eslint-plugin-jest":"^26.5.3","jest":"^28.1.2","ts-jest":"^28.0.5","ts-node":"^10.8.2","tsconfig-paths":"^4.0.0","typescript":"^4.5.2"},"dependencies":{"app-root-path":"^3.0.0","axios":"^0.27.2","axios-retry":"^3.3.1","cheerio":"^1.0.0-rc.12","promise-throttle":"^1.1.2","sequelize":"^6.21.2","sqlite3":"^5.0.8","uuid":"^8.3.2","winston":"^3.8.1"},"gitHead":"34e893a57b7fddba5d403df5798979a557a781ef","bugs":{"url":"https://github.com/juanroldanbrz/nodescrapy/issues"},"homepage":"https://github.com/juanroldanbrz/nodescrapy#readme","_id":"@bluggie/nodescrapy@0.1.1","_nodeVersion":"18.1.0","_npmVersion":"8.8.0","dist":{"integrity":"sha512-QUP+RB5alBRKE6iHAWkMgq7ejxI1hlrV8r/H+gxEjn4H3ls8rxoRwB7aHMW5QEW9ndN0TLIJY8EO/RkUpXJc8g==","shasum":"58a384c3bb34e8c48a61637b59b7f85fd241833c","tarball":"https://registry.npmjs.org/@bluggie/nodescrapy/-/nodescrapy-0.1.1.tgz","fileCount":3,"unpackedSize":17761,"signatures":[{"keyid":"SHA256:jl3bwswu80PjjokCgh0o2w5c2U4LhQAE57gj9cz1kzA","sig":"MEUCIQCChyuurOu17uE9/e2BKlcDY5gv/TrpVaroNmN0iAS/tgIgM2xXBpAnjBOGdtSHKZ8ZVLfdN5vuy5DP1T6JpjoF66w="}],"npm-signature":"-----BEGIN PGP SIGNATURE-----\r\nVersion: OpenPGP.js v4.10.10\r\nComment: https://openpgpjs.org\r\n\r\nwsFzBAEBCAAGBQJiztwGACEJED1NWxICdlZqFiEECWMYAoorWMhJKdjhPU1b\r\nEgJ2VmoS2RAAnRIpEmDZBUnhxRQazxQbkdU+Gg3gK4s48Fzzo45FDLCiMXdu\r\nN/YcfpZ1pU6a3HOrgx9OJ9nQu+i/Cr47zyqfJ8YZMnNv41woz2iE68FYo+qq\r\nyVu9b4JlziyuZ59EPbddNoJ5Lp6r4gDoKKz5il1KMbZOT8qLfzBozUKGGFtY\r\nPBq3DQExM2CRN9SUP45F7npl4elST8UXRWfAl+CZt800xB+LGqWF0a7Jt4s1\r\nWXpf8G7dVuzxHsVA76s52Z5VX/teYtkEQCB7M+0+qgQUSxy1bVAhke0Y3BC5\r\nokBtk9oSYLxRdgkNAoCp50czQzP3EydECsWQxMKhMSe8vgglvOS1cv2HkAvu\r\nDsbyvhi9i30yRB5sIjeXja61ClarrkRjudCzvVZNd1fFbYSzJk7JNMsZPfjH\r\nLT2UnxHKTJgKwWRqW2Hkf3wiN4L+biEtbmPaypHZkH7W9uDm2KHeU3sejZ0B\r\nVYnrX5XgsbwiAK6NeTJxmGuHpQP9Idjx4EKs/M9vBaoyrmzvOCjVXyBkhvyi\r\n578pRvTivNpP+lASVVmE9cAvkE/A/fmMxDemhqewOdyxVdT/reAmQS3goPy2\r\nEXwzje2WsUsOQFBSQ+HjTl2t+Ssma3jWFNeZsa25h06w8qp6EhQ8N+Tyg9pc\r\nL+v+BirHsRg6vKkiw0zvcR+xA1Nurgqc3Wg=\r\n=RnqH\r\n-----END PGP SIGNATURE-----\r\n"},"_npmUser":{"name":"bluggie","email":"juan.roldan@bluggie.com"},"directories":{},"maintainers":[{"name":"bluggie","email":"juan.roldan@bluggie.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages","tmp":"tmp/nodescrapy_0.1.1_1657723909893_0.22227531033782055"},"_hasShrinkwrap":false},"0.1.2":{"name":"@bluggie/nodescrapy","version":"0.1.2","description":"Web crawler in NodeJS","main":"dist/index.js","scripts":{"build":"tsc","lint":"eslint . --ext .ts","test":"jest --collect-coverage"},"repository":{"type":"git","url":"git+https://github.com/juanroldanbrz/nodescrapy.git"},"keywords":["scrapy","web-crawling","web-scraping","nodejs","typescript"],"author":{"name":"Juan Roldan","email":"juan.roldan.brz@gmail.com"},"license":"MIT","devDependencies":{"@types/jest":"^28.1.4","@typescript-eslint/eslint-plugin":"^5.30.4","@typescript-eslint/parser":"^5.30.4","axios-mock-adapter":"^1.21.1","coveralls":"^3.1.1","eslint":"^8.19.0","eslint-config-airbnb-base":"^15.0.0","eslint-plugin-import":"^2.26.0","eslint-plugin-jest":"^26.5.3","jest":"^28.1.2","ts-jest":"^28.0.5","ts-node":"^10.8.2","tsconfig-paths":"^4.0.0","typescript":"^4.5.2"},"dependencies":{"app-root-path":"^3.0.0","axios":"^0.27.2","axios-retry":"^3.3.1","cheerio":"^1.0.0-rc.12","promise-throttle":"^1.1.2","sequelize":"^6.21.2","sqlite3":"^5.0.8","uuid":"^8.3.2","winston":"^3.8.1"},"gitHead":"34e893a57b7fddba5d403df5798979a557a781ef","bugs":{"url":"https://github.com/juanroldanbrz/nodescrapy/issues"},"homepage":"https://github.com/juanroldanbrz/nodescrapy#readme","_id":"@bluggie/nodescrapy@0.1.2","_nodeVersion":"18.1.0","_npmVersion":"8.8.0","dist":{"integrity":"sha512-p1x3ys2h7usnvoFyJh2y26j4CHtslJDcDIKTD/gkqOYPpkNoWL9GSlijmrNkGnu1wkkMvawZVGyHqXBOAZEgGA==","shasum":"0db335ca7ee8e2b9ab8580ae2628893780c3698f","tarball":"https://registry.npmjs.org/@bluggie/nodescrapy/-/nodescrapy-0.1.2.tgz","fileCount":41,"unpackedSize":115265,"signatures":[{"keyid":"SHA256:jl3bwswu80PjjokCgh0o2w5c2U4LhQAE57gj9cz1kzA","sig":"MEQCIASPKADoSvo/AcR1KD8Gbdg9vWiNIhBGJ2L/z8yISjTcAiB159Th4G61JslRwUAqGXIb7AOv7TiR/m5rX8cj9cxlwQ=="}],"npm-signature":"-----BEGIN PGP SIGNATURE-----\r\nVersion: OpenPGP.js v4.10.10\r\nComment: https://openpgpjs.org\r\n\r\nwsFzBAEBCAAGBQJizt07ACEJED1NWxICdlZqFiEECWMYAoorWMhJKdjhPU1b\r\nEgJ2Vmr8Aw//Spt3K27oUq8ogWY4/K+re1TsgjS74dAp+xWRTdANeoqoclrF\r\nOIw3L8KqO1dQd4pkCAvS+CffOgvfh6DaAeRxvhxCPe/oCuurTCU/clCXd22b\r\nmO4MsZEFfG49s26bh/SEAOj/jfi/PmidqYjUH00Mr3teUofhTJN/hi81Vu3n\r\nsJs7d0Sjm+DFcxRY4KQieinO4I6ui4YAUmh/0PFYlDwDyfKGw5icGB+7zx6a\r\nm+Dc8feTXtsKQMPLNNVfjGLJhP9ZBaKlCeRRjLjjQWpMStokp9yh8M8L7OGG\r\nB9s6CXKlBgBVB/Uw3UrAmTJo9+COyW/ApTJf27uecgjlpzT6brRtpil6Zbqz\r\n4SiCBQEUtR12830grY/2+Gkss7xYncIfo38d1TpEBOlOTQPpp6A1FFE+7Hpc\r\ng1MKRLtM6xRtYOH8LJLyKfvCfkR19csuYZ77iLPVm2Wez6L9W9x9l5+j7ZNJ\r\nlvP11nwDbQc3LmB38ZQ7HoawXpwzsMnVAPMUtRJHUklC8zUcskR3AqPOQfpp\r\nirpmq8GXN3d6gtScl86mq75kvCkNJjLZiS6rxM5XYwHXdwBhQwuMihoyjs0b\r\nkRilDugi9MLpP2C7Yijbz17GHeq4QcC6ju8KY7pf16APcMuaXUwHdRYqRnHo\r\nczih/Uq3XCmv34jf8JVghy8IdMGbPJYd9Ds=\r\n=6FEq\r\n-----END PGP SIGNATURE-----\r\n"},"_npmUser":{"name":"bluggie","email":"juan.roldan@bluggie.com"},"directories":{},"maintainers":[{"name":"bluggie","email":"juan.roldan@bluggie.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages","tmp":"tmp/nodescrapy_0.1.2_1657724219630_0.2484787297756086"},"_hasShrinkwrap":false},"0.1.3":{"name":"@bluggie/nodescrapy","version":"0.1.3","description":"Web crawler in NodeJS","main":"dist/index.js","scripts":{"build":"tsc","lint":"eslint . --ext .ts","test":"jest --collect-coverage"},"repository":{"type":"git","url":"git+https://github.com/juanroldanbrz/nodescrapy.git"},"keywords":["scrapy","web-crawling","web-scraping","nodejs","typescript"],"author":{"name":"Juan Roldan","email":"juan.roldan.brz@gmail.com"},"license":"MIT","devDependencies":{"@types/jest":"^28.1.4","@typescript-eslint/eslint-plugin":"^5.30.4","@typescript-eslint/parser":"^5.30.4","axios-mock-adapter":"^1.21.1","coveralls":"^3.1.1","eslint":"^8.19.0","eslint-config-airbnb-base":"^15.0.0","eslint-plugin-import":"^2.26.0","eslint-plugin-jest":"^26.5.3","jest":"^28.1.2","ts-jest":"^28.0.5","ts-node":"^10.8.2","tsconfig-paths":"^4.0.0","typescript":"^4.5.2"},"dependencies":{"app-root-path":"^3.0.0","axios":"^0.27.2","axios-retry":"^3.3.1","cheerio":"^1.0.0-rc.12","promise-throttle":"^1.1.2","sequelize":"^6.21.2","sqlite3":"^5.0.8","uuid":"^8.3.2","winston":"^3.8.1"},"gitHead":"c0eb8938bf7005e927df3b69707ad2799201f410","bugs":{"url":"https://github.com/juanroldanbrz/nodescrapy/issues"},"homepage":"https://github.com/juanroldanbrz/nodescrapy#readme","_id":"@bluggie/nodescrapy@0.1.3","_nodeVersion":"18.1.0","_npmVersion":"8.8.0","dist":{"integrity":"sha512-BuzAmaswMwVAXXiRCgKCq6kJc6gZ9A73SM2hWf6Kly5CLTrqeXLyks7aCksQhnhbRdzYUSPuMA67K91h26HLzg==","shasum":"763044ff112b4030c5f63f2c23f0e3db59b3e6f9","tarball":"https://registry.npmjs.org/@bluggie/nodescrapy/-/nodescrapy-0.1.3.tgz","fileCount":41,"unpackedSize":115278,"signatures":[{"keyid":"SHA256:jl3bwswu80PjjokCgh0o2w5c2U4LhQAE57gj9cz1kzA","sig":"MEYCIQC/wynoEKZxUswFDtC0HQopyp0CFAlMDfZNDlSOtw+JKAIhAL6sE3dCF3Z+JC0G2JbwtMZM7M44Q+QDSDtRg3JaZOUz"}],"npm-signature":"-----BEGIN PGP SIGNATURE-----\r\nVersion: OpenPGP.js v4.10.10\r\nComment: https://openpgpjs.org\r\n\r\nwsFzBAEBCAAGBQJizv10ACEJED1NWxICdlZqFiEECWMYAoorWMhJKdjhPU1b\r\nEgJ2VmpiYA/+KjyHjADproWdwkgOlYEwnSO65WX61DZ0t/0elQ42cAryI7L1\r\nFgYJ7N2COvCPlEzkiGFOzbCEXanUJ06sKddOQRdUP1BemZjPbadbd6VbnO+Q\r\niEBJnncC1c37clhbXH6q+fnY0zFHHrMb4s9PLSL7+PM7JJB4uhmzu9P6Zipy\r\nGCKF4JEvYXkU3BFBsAFx6GzIGqC2xdYjMu5ATlngLJLtTDRMOaCjEtN/BBzQ\r\nkjRw7rV+9a5E3DtsN6rGKjnHXK+f55+SCxEeuzu5VDbxANzKfgs2us17f6Bs\r\nzllSl5zlr2rguUTCNCN7Hp+bF1RQAOzMGxz6dtiYbOvy+vW3+0l/lFsOllXt\r\nHW29GaAm4SpxpmYTTmR2mWjT3SHaODob77gWdD2McF+kAdjpL9z8TfZXwg6m\r\n8ueZX+zfv7GZw+mmfVaIL/xHobywFLzZPslEWnnB3ZCfoL2uk9d/wi7mUYge\r\nbgpLuMaYyNqQk7Od24OJOwj6grCLe7VM0g6LHRfeYUqNyKjUoy4RMc80d+c8\r\nk3IR4xIAttkX+FgXhqadxaU5Qc0S3mpc8Xu9kcXAhYP4Do46LgeRnALPNDUT\r\n/9/OQJCO/UZCyeNiOpb2Cu0vjxwDDpKOVce6tH+i4UX4pa1+/ep3kc3iaVex\r\n8LxXMPpy7akw/V5fdkPss/5xWnp27ZD4ioo=\r\n=CVVs\r\n-----END PGP SIGNATURE-----\r\n"},"_npmUser":{"name":"bluggie","email":"juan.roldan@bluggie.com"},"directories":{},"maintainers":[{"name":"bluggie","email":"juan.roldan@bluggie.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages","tmp":"tmp/nodescrapy_0.1.3_1657732467748_0.6612319433439604"},"_hasShrinkwrap":false},"0.1.4":{"name":"@bluggie/nodescrapy","version":"0.1.4","description":"Web crawler in NodeJS","main":"dist/index.js","scripts":{"build":"tsc","lint":"eslint . --ext .ts","test":"jest --collect-coverage"},"repository":{"type":"git","url":"git+https://github.com/juanroldanbrz/nodescrapy.git"},"keywords":["scrapy","web-crawling","web-scraping","nodejs","typescript"],"author":{"name":"Juan Roldan","email":"juan.roldan.brz@gmail.com"},"license":"MIT","devDependencies":{"@types/jest":"^28.1.4","@typescript-eslint/eslint-plugin":"^5.30.4","@typescript-eslint/parser":"^5.30.4","axios-mock-adapter":"^1.21.1","coveralls":"^3.1.1","eslint":"^8.19.0","eslint-config-airbnb-base":"^15.0.0","eslint-plugin-import":"^2.26.0","eslint-plugin-jest":"^26.5.3","jest":"^28.1.2","rewiremock":"^3.14.3","ts-jest":"^28.0.5","ts-node":"^10.8.2","tsconfig-paths":"^4.0.0","typescript":"^4.5.2"},"dependencies":{"app-root-path":"^3.0.0","axios":"^0.27.2","axios-retry":"^3.3.1","cheerio":"^1.0.0-rc.12","promise-throttle":"^1.1.2","puppeteer":"^15.4.0","puppeteer-cluster":"^0.23.0","sequelize":"^6.21.2","sqlite3":"^5.0.8","uuid":"^8.3.2","winston":"^3.8.1"},"gitHead":"29b385c448dde195e92901aa3716beb8c305d113","bugs":{"url":"https://github.com/juanroldanbrz/nodescrapy/issues"},"homepage":"https://github.com/juanroldanbrz/nodescrapy#readme","_id":"@bluggie/nodescrapy@0.1.4","_nodeVersion":"18.1.0","_npmVersion":"8.8.0","dist":{"integrity":"sha512-SRpPLrFSxHF5Yn7pwVOHYip+50p1b2ZyRxtixj2H8r+23ZMO0r5PYVHmFFuh0pJbTUwc7Gil6j1Kf3fLC3xTEw==","shasum":"19b8c2b3b4f6baaec2c76c1ca970b77c89575409","tarball":"https://registry.npmjs.org/@bluggie/nodescrapy/-/nodescrapy-0.1.4.tgz","fileCount":49,"unpackedSize":135655,"signatures":[{"keyid":"SHA256:jl3bwswu80PjjokCgh0o2w5c2U4LhQAE57gj9cz1kzA","sig":"MEUCIHFKPR+S8xpuQJQWNn+nZZngo8L+o9x7CKGv6QJs8+nfAiEAwdvHl00GV2Dil51O+2f9rGJBqb2WdzrKg0NVUhuN8qs="}],"npm-signature":"-----BEGIN PGP SIGNATURE-----\r\nVersion: OpenPGP.js v4.10.10\r\nComment: https://openpgpjs.org\r\n\r\nwsFzBAEBCAAGBQJi1DdwACEJED1NWxICdlZqFiEECWMYAoorWMhJKdjhPU1b\r\nEgJ2VmpJYQ//RWLQpv/GOZVRAGZbeLsax19RMENmoQ5z1pGu1YiNve6kiWHR\r\n4suI8/A1p9d/7QVUWuCRn2Zb/eDL90CadUNPRGqcPoPqm1reQqDrRd6KKR0r\r\n6yhlQyao1e/F1jmJtX0OpAyiefiH8eMD19l/oB5NXhMGNTH44vqBH5xG72EZ\r\nopMu6OvDKjdvI+Dav+gpCl+Yk8dV8jQNagXppn8D+bN0fLiHLR0nxNxeMEGF\r\nDFkTwi+N2MBLzj+pHlkRYvzHD3aO49FthpGTMj8q2xWaNyu12XiYf2CPu+1a\r\nyQ0lgfAb99p2zY3k2vnH935SE28vUQ+MRc59zs+NtGDxXJjn8QBM7feF/c3Z\r\nX9aDsH/ezD3kpBoufNjC864+oSFzVB3311BPE0MtHXi97sFKk8PHmBCl2EtO\r\n0HxevELDc6JX9p+on+Khh7hhcUeXTNc6njkc6oTGv6Zi4IsJivFhVzDqH8bS\r\nXlOf4u0t0pLhIFAxEzQecIdktyksMGqlpCXmlLpiWuc1Ar+SKjirjXZ7yC9J\r\nl8G83r5x/jQroNXJDcrEi2tExoICGNql6GmFzpJrOk7/o6XwyAMqLPZQtbAr\r\nVV17gRP/6bdd/RqMFe143Rr5dKPH8ZRJf27tkpYJ45+A7ot4kdv03wZTqza4\r\nYdvCwIEHD9h/furAxBE0OjjYXjBJSw9rjqs=\r\n=alr0\r\n-----END PGP SIGNATURE-----\r\n"},"_npmUser":{"name":"bluggie","email":"juan.roldan@bluggie.com"},"directories":{},"maintainers":[{"name":"bluggie","email":"juan.roldan@bluggie.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages","tmp":"tmp/nodescrapy_0.1.4_1658074992343_0.8300241359553757"},"_hasShrinkwrap":false},"0.1.5":{"name":"@bluggie/nodescrapy","version":"0.1.5","description":"Web crawler in NodeJS","main":"dist/index.js","scripts":{"build":"tsc","lint":"eslint . --ext .ts","test":"jest --collect-coverage"},"repository":{"type":"git","url":"git+https://github.com/juanroldanbrz/nodescrapy.git"},"keywords":["scrapy","web-crawling","web-scraping","nodejs","typescript"],"author":{"name":"Juan Roldan","email":"juan.roldan.brz@gmail.com"},"license":"MIT","devDependencies":{"@types/jest":"^28.1.4","@typescript-eslint/eslint-plugin":"^5.30.4","@typescript-eslint/parser":"^5.30.4","axios-mock-adapter":"^1.21.1","coveralls":"^3.1.1","eslint":"^8.19.0","eslint-config-airbnb-base":"^15.0.0","eslint-plugin-import":"^2.26.0","eslint-plugin-jest":"^26.5.3","jest":"^28.1.2","rewiremock":"^3.14.3","ts-jest":"^28.0.5","ts-node":"^10.8.2","tsconfig-paths":"^4.0.0","typescript":"^4.5.2"},"dependencies":{"app-root-path":"^3.0.0","axios":"^0.27.2","axios-retry":"^3.3.1","cheerio":"^1.0.0-rc.12","promise-throttle":"^1.1.2","puppeteer":"^15.4.0","puppeteer-cluster":"^0.23.0","sequelize":"^6.21.2","sqlite3":"^5.0.8","uuid":"^8.3.2","winston":"^3.8.1"},"gitHead":"8eab40e26e3d756d5fd59a89dda822003d769383","bugs":{"url":"https://github.com/juanroldanbrz/nodescrapy/issues"},"homepage":"https://github.com/juanroldanbrz/nodescrapy#readme","_id":"@bluggie/nodescrapy@0.1.5","_nodeVersion":"18.1.0","_npmVersion":"8.8.0","dist":{"integrity":"sha512-c09h8k4llmjOfHxn4gqzcYT6EERrc4Cn/GhTTm3/Lov7TXYc/VtvWPKICroEn8V86i+XOaWC0MNnt2hjn9aXHQ==","shasum":"cbd6b422cf50ede48524ef737b12b7501c61c7d7","tarball":"https://registry.npmjs.org/@bluggie/nodescrapy/-/nodescrapy-0.1.5.tgz","fileCount":35,"unpackedSize":75097,"signatures":[{"keyid":"SHA256:jl3bwswu80PjjokCgh0o2w5c2U4LhQAE57gj9cz1kzA","sig":"MEUCIQClB5smdLkuVt9Bs4t6CXzN/RxGuprfzbBXlt0molXvegIgFKlvhO+yT3G4aT3oDrSWIKdns7wIs22UvJhiU983cas="}],"npm-signature":"-----BEGIN PGP SIGNATURE-----\r\nVersion: OpenPGP.js v4.10.10\r\nComment: https://openpgpjs.org\r\n\r\nwsFzBAEBCAAGBQJi1EBGACEJED1NWxICdlZqFiEECWMYAoorWMhJKdjhPU1b\r\nEgJ2VmrCFBAAmO817oDYvB8rtHarTgYF6RoHC+h2E4InvICi4gCkHkFk8U6w\r\nQVApmYMKaZcqXAeREUElz7TTUkTkoxtgnCvFJ9/4pfjVIOZV6nt5kk/0x60L\r\nAWZNvoBULgqiCXOdleMTSGTlAONGXt5gENQ/MJtTuHyUTIVLH9RSfOCvOkk5\r\n1Cf1vtjQQnuOtyUYPJVIfQ1t5qGQqOZnvxajm+o4U8dujYHjiWOSDByEseSm\r\nvmTQAVeEqZusDyhO2Oz6Yi+fDDyMRxZD2cAnaklnZoOVBZAY2YiF4YmgZR19\r\nrl1DmU7FcX9efIX4iPD4awCiIEpAYNyU7bQChJelf9h8+3/HFsvAZb7JGjlx\r\nnMzuRe3WsQ3uHV+Tyrh4npsYjdjNb/6nzLC2pRFvfZ5/W0oUSeAqy6eL648L\r\nR2C93eXAxIABBMo+tSZSwM3q9srD3JXC2aCkSCwXn3yg7wC3y9TfqZWE8tnW\r\nw6fxVYNT4LDGxfr+OPmLuOTEJ+m7eOCIPKUIUicMtatflHNYrBF4CHk50VAO\r\n5Vr2WaxYdnFpd/Xs5xS0RNAwvQAxz1MQg1DiCWLOhqx2iS6n5gFZ34INmAtT\r\nizp3Aj6jD1+9CG0lOGHP1rDHSQHK4Y16W7GGAdurxfV5EXLOgvowhhb+ysrS\r\nYG9fEvE4an9ClN/ic3ptlfB1FtRXPcPaSK0=\r\n=wpMG\r\n-----END PGP SIGNATURE-----\r\n"},"_npmUser":{"name":"bluggie","email":"juan.roldan@bluggie.com"},"directories":{},"maintainers":[{"name":"bluggie","email":"juan.roldan@bluggie.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages","tmp":"tmp/nodescrapy_0.1.5_1658077253797_0.5823112483092274"},"_hasShrinkwrap":false},"0.1.6":{"name":"@bluggie/nodescrapy","version":"0.1.6","description":"Web crawler in NodeJS","main":"dist/index.js","scripts":{"build":"tsc","lint":"eslint . --ext .ts","test":"jest --collect-coverage"},"repository":{"type":"git","url":"git+https://github.com/juanroldanbrz/nodescrapy.git"},"keywords":["scrapy","web-crawling","web-scraping","nodejs","typescript"],"author":{"name":"Juan Roldan","email":"juan.roldan.brz@gmail.com"},"license":"MIT","devDependencies":{"@types/jest":"^28.1.4","@typescript-eslint/eslint-plugin":"^5.30.4","@typescript-eslint/parser":"^5.30.4","axios-mock-adapter":"^1.21.1","coveralls":"^3.1.1","eslint":"^8.19.0","eslint-config-airbnb-base":"^15.0.0","eslint-plugin-import":"^2.26.0","eslint-plugin-jest":"^26.5.3","jest":"^28.1.2","rewiremock":"^3.14.3","ts-jest":"^28.0.5","ts-node":"^10.8.2","tsconfig-paths":"^4.0.0","typescript":"^4.5.2"},"dependencies":{"app-root-path":"^3.0.0","axios":"^0.27.2","axios-retry":"^3.3.1","cheerio":"^1.0.0-rc.12","promise-throttle":"^1.1.2","puppeteer":"^15.4.0","puppeteer-cluster":"^0.23.0","sequelize":"^6.21.2","sqlite3":"^5.0.8","uuid":"^8.3.2","winston":"^3.8.1"},"gitHead":"ebd47c42ab199b15fbc15790e7866dfe264b27cd","bugs":{"url":"https://github.com/juanroldanbrz/nodescrapy/issues"},"homepage":"https://github.com/juanroldanbrz/nodescrapy#readme","_id":"@bluggie/nodescrapy@0.1.6","_nodeVersion":"18.1.0","_npmVersion":"8.8.0","dist":{"integrity":"sha512-/fu4YXA2kiHiJo0+/36OZGZWgzF/YLh7ylmEVoerflY+ZeY4Lj1hakkGLkaS2iXec7jtnI7W7wOLkGhEBzRUKw==","shasum":"886598483da22603f8f3c908319be5dcec80b294","tarball":"https://registry.npmjs.org/@bluggie/nodescrapy/-/nodescrapy-0.1.6.tgz","fileCount":35,"unpackedSize":75166,"signatures":[{"keyid":"SHA256:jl3bwswu80PjjokCgh0o2w5c2U4LhQAE57gj9cz1kzA","sig":"MEUCICqm2oVV+kq+aeukIj4jNJvqn39IHWIHXZD7b1Ng2tZGAiEAvlEv2iwjQ5DKpOQD0TXnJkvQO00M4VCXbLQQeaMkNCY="}],"npm-signature":"-----BEGIN PGP SIGNATURE-----\r\nVersion: OpenPGP.js v4.10.10\r\nComment: https://openpgpjs.org\r\n\r\nwsFzBAEBCAAGBQJi1Q8kACEJED1NWxICdlZqFiEECWMYAoorWMhJKdjhPU1b\r\nEgJ2VmpCEQ//bXkqwv8/PX3d2Xyom2UTM2TUIKtLAwLDJKEzkI9yIAOmMRfG\r\nNvfkvxzcCjOj1o0uUbfm2rPlUHw0C9CjfYAefxyvsPh18Z1haTXiYnct7mZv\r\ncJO6p1kChzghCLL1PvG8seDpwcmxBrhtzA0Ejlguy31Ho1/qnYtSzllaESmR\r\nxC+IiHh5z3ubPvxrLe88OmrwtcPCJrfjYpSVEybdDJkC/7h3qeqrN1DQNWn2\r\nAwdnjlfJ1oqXph7h4JQXQRiyqS8gb4s6UZPDSl+gz3rR9wo2gKXfCAMNRZPf\r\n/Lc1En9595Kg4RDxxGTBZHbdnB4M8pWOeW/SyEzO9GkLbN5V2ztUx7heE/Nu\r\nNUkdO0cZXqLHKfufsFApYQg6yL5OfGAVl9//gZOnIHvsNn0j3CP1cfKChC/q\r\nnM0j+Yd3iCnHEidZMRX5lcbBBr3vX7t+qJWFv+K6kFq+rWECsaOi9W46sw/8\r\nZMjntWIjVp6NwGri/IGwC4gI9WYyR8Td40/Jn1Ez5BVVp+HTasIfUZFHoYNO\r\nAFnnsLUIoL5Vf3CDrKxdF0iXGmmpTiDpyE4KPHeuWSOi/tY9mVuR+aOVoXOB\r\nV60aPHlGMjifrpbVpZJBHYhKDEI+UHpZs6UNeewlw7pvkOIywy4d9Oxr2Ju5\r\nxWW5X8bRZsOh9CPO6GnawQGYfWJ+QlTwaac=\r\n=4lJR\r\n-----END PGP SIGNATURE-----\r\n"},"_npmUser":{"name":"bluggie","email":"juan.roldan@bluggie.com"},"directories":{},"maintainers":[{"name":"bluggie","email":"juan.roldan@bluggie.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages","tmp":"tmp/nodescrapy_0.1.6_1658130212802_0.3762036904969164"},"_hasShrinkwrap":false}},"time":{"created":"2022-07-13T14:46:57.147Z","0.1.0":"2022-07-13T14:46:57.411Z","modified":"2022-07-18T07:43:33.010Z","0.1.1":"2022-07-13T14:51:50.087Z","0.1.2":"2022-07-13T14:56:59.799Z","0.1.3":"2022-07-13T17:14:27.994Z","0.1.4":"2022-07-17T16:23:12.545Z","0.1.5":"2022-07-17T17:00:53.986Z","0.1.6":"2022-07-18T07:43:32.944Z"},"maintainers":[{"name":"bluggie","email":"juan.roldan@bluggie.com"}],"description":"Web crawler in NodeJS","keywords":["scrapy","web-crawling","web-scraping","nodejs","typescript"],"author":{"name":"Juan Roldan","email":"juan.roldan.brz@gmail.com"},"license":"MIT","readme":"<img src=\"./logo.png\" width=\"400\" height=\"250\" />\n\n![master](https://github.com/juanroldanbrz/nodescrapy/actions/workflows/ci.yml/badge.svg)\n[![Coverage Status](https://coveralls.io/repos/github/juanroldanbrz/nodescrapy/badge.svg)](https://coveralls.io/github/juanroldanbrz/nodescrapy)\n## Overview \n\nNodescrapy is a fast high-level and highly configurable web crawling and web scraping framework, used to\ncrawl websites and extract structured data from their pages.\n\nNodescrapy is written in [Typescript](https://www.typescriptlang.org/) and works in a [NodeJs](http://nodejs.org) environment.\n\nNodescrapy comes with a built-in web spider, which will discover automatically all the URLs of the website.\n\nNodescrapy saves the status of the crawling in a local Sqlite database, so crawling can be stopped and resumed.\n\nNodescrapy provides a default integration with AXIOS and PUPPETEER to choose if rendering javascript or not.\n\nBy default, Nodescrapy saves the results of the scrapping in a local folder in JSON files.\n\n```ts\nimport {HtmlResponse, WebCrawler} from '@bluggie/nodescrapy';\n\nconst onItemCrawledFunction = (response: HtmlResponse) => {\n    return { \"data1\": ... }\n}\n\nconst crawler = new WebCrawler({\n    dataPath: './crawled-items',\n    entryUrls: ['https://www.pararius.com/apartments/amsterdam'],\n    onItemCrawled: onItemCrawledFunction\n});\n\ncrawler.crawl()\n    .then(() => console.log('Crawled finished'));\n```\n\n## What does nodescrapy do?\n\n* Provides a web client configurable with retries and delays.\n* Extremely configurable for writing your own crawler.\n* Provides a configurable discovery implementation to auto-detecting linked resources and filter the ones you want.\n* Saves the status of the crawling in a file storage, so crawled can be paused and resumed.\n* Provides basic statistics on crawling status.\n* Automatically parses the DOM of the HTMLs with [Cheerio](https://cheerio.js.org/)\n* Implementations can be easily extended.\n* Fully written in Typescript.\n\n## Documentation\n\n- [Installation](#installation)\n- [Getting started](#getting-started)\n- [Crawling modes](#crawling-modes)\n- [Data Models](#data-models)\n- [Configuration](#crawler-configuration)\n- [Examples](#examples)\n- [Roadmap](#roadmap)\n- [Contributors](#contributors)\n- [License](#license)\n\n## Installation\n\n```sh\nnpm install --save @bluggie/nodescrapy\n```\n\n## Getting Started\n\n\nInitializing nodescrapy is a simple process. First, you require the module and instantiate it with the config argument. You then configure the properties you like (eg. the request interval), register the `onItemCrawled` method, and call the crawl method. Let's walk through the process!\n\nAfter requiring the crawler, we create a new instance of it. We supply the constructor with the [Crawler Configuration](#CrawlerConfiguration). \nA simple configuration contains:\n* Entry url/urls for the crawler.\n* Where to store the crawled items (dataPath).\n* What to do when a new page is crawled (function [onItemCrawled](#CrawlerConfiguration+onItemCrawled))\n\n```js\nimport {HtmlResponse, WebCrawler} from '@bluggie/nodescrapy';\n\nconst onItemCrawledFunction = (response: HtmlResponse) => {\n    if (!response.url.includes('-for-rent')) {\n        return undefined;\n    }\n\n    const $ = response.$;\n    return {\n        'title': $('.listing-detail-summary__title , #onetrust-accept-btn-handler').text(),\n    }\n}\n\nconst crawler = new WebCrawler({\n    dataPath: './crawled-items',\n    entryUrls: ['https://www.pararius.com/apartments/amsterdam'],\n    onItemCrawled: onItemCrawledFunction\n});\n\ncrawler.crawl()\n    .then(() => console.log('Crawled finished'));\n```\nThe function `onItemCrawledFunction` is required, since the crawler will invoke it to extract the data from hte HTML document.\nIt will return `undefined` if there is nothing to extract from that page, or an object of `{key: values}` if data could be extracted from that page.\nSee [onItemCrawled](#CrawlerConfiguration+onItemCrawled) for more information.\n\n\nWhen running the application, it will produce the following logs:\n\n```html\ninfo: Jul-08-2022 09:08:57: Crawled started.\ninfo: Jul-08-2022 09:08:57: Crawling https://www.pararius.com/apartments/amsterdam\ninfo: Jul-08-2022 09:09:00: Crawling https://www.pararius.com/apartments/amsterdam/map\ninfo: Jul-08-2022 09:09:01: Crawling https://www.pararius.com/apartment-for-rent/amsterdam/b180b6df/president-kennedylaan\ninfo: Jul-08-2022 09:09:04: Adding crawled entry to data: https://www.pararius.com/apartment-for-rent/amsterdam/b180b6df/president-kennedylaan\ninfo: Jul-08-2022 09:09:04: Crawling https://www.pararius.com/real-estate-agents/amsterdam/expathousing-com-amsterdam\ninfo: Jul-08-2022 09:09:04: Adding crawled entry to data: https://www.pararius.com/apartment-for-rent/amsterdam/b180b64f/president-kennedylaan\ninfo: Jul-08-2022 09:09:20: Saving 2 entries into JSON file: data-2022-07-08T07:09:20.115Z.json\ninfo: Jul-08-2022 12:37:16: Crawled 29 urls. Remaining: 328\n```\n\nThis will also store the data in a JSON file (by default, 50 entries per JSON file. Configurable with `dataBatchSize` property).\n```json\n[\n  {\n    \"provider\": \"nodescrapy\",\n    \"url\": \"https://www.pararius.com/apartment-for-rent/amsterdam/2365cc70/gillis-van-ledenberchstraat\",\n    \"data\": {\n      \"data1\": \"test\"\n    },\n    \"added_at\": \"2022-07-08T10:38:53.431Z\",\n    \"updated_at\": \"2022-07-08T10:38:53.431Z\"\n  },\n  {\n    \"provider\": \"nodescrapy\",\n    \"url\": \"https://www.pararius.com/apartment-for-rent/amsterdam/61e78537/nieuwezijds-voorburgwal\",\n    \"data\": {\n      \"data1\": \"test\"\n    },\n    \"added_at\": \"2022-07-08T10:38:55.466Z\",\n    \"updated_at\": \"2022-07-08T10:38:55.466Z\"\n  }\n]\n```\n\n## Crawling modes\n\nNodescrapy can run in two different modes:\n- START_BY_SCRATCH\n- CONTINUE\n\nIn **START_BY_SCRATCH** mode, every time the crawler runs will start from 0, going through the [entryUrls](#CrawlerConfiguration+entryUrls) and all the discovered links.\n\nIn **CONTINUE** mode, the crawler will only crawl the links which were not processed from the last run, and also the new ones which are being discovered.\n\nTo see how to configure this, go to [mode](#CrawlerConfiguration+mode)\n\n\n## Data Models\n\n<a name=\"DataModel+HttpRequest\"></a>\n#### HttpRequest\n\nHttpRequest is a wrapper including:\n- The url which is going to be crawled\n- The headers which are going to be send in the request (i.e User-Agent)\n\n\n```ts\ninterface HttpRequest {\n  url: string;\n  \n  headers: { [key: string]: string; }\n}\n```\n\n<a name=\"DataModel+HtmlResponse\"></a>\n#### HtmlResponse\n\nHtmlResponse is a wrapper including:\n- The crawled url\n- The axios response (see [AxiosResponse](https://axios-http.com/docs/res_schema)) \n- The DOM processed by [Cheerio](https://cheerio.js.org/)\n\nThis information should be enough to extract the information you need from that webpage.\n```ts\ninterface HtmlResponse {\n  url: string;\n\n  originalResponse: AxiosResponse;\n\n  $: CheerioAPI;\n}\n\n```\n<a name=\"DataModel+DataEntry\"></a>\n\n#### DataEntry\n\nRepresents the data that will be stored in the file system after a page with data has been crawled.\n\nContains:\n- the id of the entry (primary key).\n- the provider (crawler name).\n- the url.\n- the data extracted by the [onItemCrawled](#CrawlerConfiguration+onItemCrawled) function.\n- when the data was added and updated.\n```ts\ninterface DataEntry {\n    id?: number,\n    provider: string,\n    url: string,\n    data: { [key: string]: string; },\n    added_at: Date,\n    updated_at: Date\n}\n```\n\n<a name=\"DataModel+CrawlContinuationMode\"></a>\n#### CrawlContinuationMode\n\nEnum which defines how the crawler will run; either starting from scratch or continuing with the last execution.\n\nValues:\n- START_FROM_SCRATCH\n- CONTINUE\n```ts\nenum CrawlContinuationMode {\n    START_FROM_SCRATCH = 'START_FROM_SCRATCH',\n    CONTINUE = 'CONTINUE'\n}\n```\n\n<a name=\"DataModel+CrawlerClientLibrary\"></a>\n\n#### CrawlerClientLibrary\n\nEnum which defines the implementation of the client.\nPuppeteer will automatically render javascript using chrome.\n\nIf puppeteer, chrome executable should be present in the system.\n\nValues:\n- AXIOS\n- PUPPETEER\n```ts\nenum CrawlerClientLibrary {\n    AXIOS = 'AXIOS',\n    PUPPETEER = 'PUPPETEER'\n}\n```\n<a name=\"CrawlerConfiguration\"></a>\n## Crawler configuration\n### Full typescript configuration definition\n\nThis is a definition of all the possible configuration supported currently by the crawler.\n```ts\n{\n    name: 'ParariusCrawler',\n    mode: 'START_FROM_SCRATCH',\n    entryUrls: ['http://www.pararius.com'],\n    client: {\n        library: 'PUPPETEER',\n        autoScrollToBottom: true,\n        concurrentRequests: 5,\n        retries: 5,\n        userAgent: 'Firefox',\n        retryDelay: 2,\n        delayBetweenRequests: 2,\n        timeoutSeconds: 100,\n        beforeRequest: (htmlRequest: HttpRequest) => { // Only for AXIOS client.\n            htmlRequest.headers.Authorization = 'JWT MyAuth';\n            return htmlRequest;\n        }\n    },\n    discovery: {\n        allowedDomains: ['www.pararius.com'],\n        allowedPath: ['amsterdam/'],\n        removeQueryParams: true,\n        onLinksDiscovered: undefined\n    },\n    onItemCrawled: (response: HtmlResponse) => {\n        if (!response.url.includes('-for-rent')) {\n            return undefined;\n        }\n\n        const $ = response.$;\n        return {\n            'title': $('.listing-detail-summary__title , #onetrust-accept-btn-handler').text(),\n        }\n    }\n    dataPath: './output-json',\n    dataBatchSize: 10,\n    sqlitePath: './cache.sqlite'\n}\n```\n\n<a name=\"CrawlerConfiguration+name\"></a>\n#### name :  <code>string</code>\n\nName of the crawler. \n\nThe name of the crawler is important in the following scenarios:\n- When resuming a crawler. The library will find the last status based in crawler name. If you change the name, the status will be reset.\n- When having multiple crawlers. The library stores the status in a SQLite database indexed by the crawler name.\n\n\n**Default**: `nodescrapy`\n<br></br>\n\n<a name=\"CrawlerConfiguration+mode\"></a>\n#### mode :  <code>string</code>\n\nMode of the crawler.\nTo see options, check [CrawlConfigurationMode](#DataModel+CrawlContinuationMode)\n\n**Default**: `START_BY_SCRATCH`\n<br></br>\n\n<a name=\"CrawlerConfiguration+entryUrls\"></a>\n#### entryUrls :  <code>string[]</code>\n\nList of urls which will start to crawl.\n###### Example:\n```ts\n{\n    entryUrls: ['https://www.pararius.com/apartments/amsterdam']\n}\n```\n<br></br>\n\n<a name=\"CrawlerConfiguration+onItemCrawled\"></a>\n#### onItemCrawled :  <font size=\"1\"> <code >function (response: HtmlResponse) => { [key: string]: any; } | undefined;</code></font></code>\nFunction to extract the data when an url has been crawled.\n\nIf returns undefined, the url will be discarded and nothing will be stored for it.\n\nThe argument of this function is provided by the crawler, and it is a [HtmlResponse](#DataModel+HtmlResponse)\n<br></br>\n###### Example\n```ts\n {\n    onItemCrawled: (response: HtmlResponse) => {\n        if (!response.url.includes('-for-rent')) {\n            return undefined; // Only extract information fron the urls which contains for-rent\n        }\n\n        const $ = response.$;\n        return {\n            'title': $('.listing-detail-summary__title , #onetrust-accept-btn-handler').text(), // Extract the title of the page.\n        }\n    }\n}\n```\n<br></br>\n\n<a name=\"CrawlerConfiguration+dataPath\"></a>\n#### dataPath :  <code>string</code>\n\nConfigures where the output of the crawler ([DataEntries](#DataModel+DataEntry)) will be stored. \n###### Example\n\n```ts\n{\n    dataPath: './output-data'\n}\n```\n\nThis will produce the following files: \n\n`./output-data/data-2022-07-11T08:17:38.188Z.json`\n\n`./output-data/data-2022-07-11T08:17:41.188Z.json`\n\n...\n<br></br>\n\n<a name=\"CrawlerConfiguration+dataBatchSize\"></a>\n#### dataBatchSize :  <code>number</code>\nThis property configures how many crawled items will be persisted in an unique file.\n\nFor example, if the number is 5, every JSON file will contain 5 crawled items. **Default**: 50\n<br></br>\n\n<a name=\"CrawlerConfiguration+sqlitePath\"></a>\n#### sqlitePath :  <code>string</code>\n\nConfigures where to store the sqlite database (full path, including name) \n\n**Default**: `node-modules/nodescrapy/cache.sqlite`\n<br></br>\n\n\n<a name=\"CrawlerClientConfig\"></a>\n### Client configuration\n\n<a name=\"CrawlerClientConfig+library\"></a>\n#### client.library :  <code>string</code>\n\nChooses the client [implementation](#DataModel+CrawlerClientLibrary) between AXIOS or PUPPETEER.\n**Default**: `AXIOS`\n<br></br>\n\n<a name=\"CrawlerClientConfig+concurrentRequests\"></a>\n#### client.concurrentRequests :  <code>number</code>\n\nConfigures the number of concurrent requests.\n**Default**: `1`\n<br></br>\n\n\n<a name=\"CrawlerClientConfig+retries\"></a>\n#### client.retries :  <code>number</code>\n\nConfigures the number of retries to perform when a request is failed.\n**Default**: `2`\n<br></br>\n\n<a name=\"CrawlerClientConfig+userAgent\"></a>\n#### client.userAgent :  <code>string</code>\n\nConfigures the user agents of the client.\n\n**Default**: `Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36`\n<br></br>\n\n#### client.autoScrollToBottom :  <code>boolean</code>\n\nIf true and client is puppeteer, every page will be scrolled to the bottom before rendered.\n\n**Default**: `true`\n<br></br>\n\n<a name=\"CrawlerClientConfig+retryDelay\"></a>\n#### client.retryDelay :  <code>number</code>\n\nConfigures how many seconds the client will wait between different requests.\n**Default**: `5`\n<br></br>\n\n<a name=\"CrawlerClientConfig+timeoutSeconds\"></a>\n#### client.timeoutSeconds :  <code>number</code>\n\nConfigures the timeout of the client, in seconds. **Default**: `10`\n<br></br>\n\n<a name=\"CrawlerClientConfig+beforeRequest\"></a>\n#### client.beforeRequest :  <code>(htmlRequest: HttpRequest) => HttpRequest</code>\n\nFunction which allows to modify the url or the headers before performing the request.\nUseful to add authentication headers or change the URL for a proxy one.\n\n**Default**: `undefined`\n\n###### Example\n```ts\n    {\n        client.beforeRequest: (request: HttpRequest): HttpRequest => {\n            const proxyUrl = `http://www.myproxy.com?url=${request.url}`;\n    \n            const requestHeaders = request.headers;\n            requestHeaders.Authorization = 'JWT ...';\n    \n            return {\n                url: proxyUrl,\n                headers: requestHeaders,\n            };\n        }\n    }\n```\n<br></br>\n\n### Discovery configuration\n\n<a name=\"CrawlerDiscoveryConfig+allowedDomains\"></a>\n#### discovery.allowedDomains :  <code>string[]</code>\n\nWhitelist of domains to crawl. **Default**: Same domains that [entryUrls](#CrawlerConfiguration+entryUrls)\n<br></br>\n\n<a name=\"CrawlerDiscoveryConfig+allowedPath\"></a>\n#### discovery.allowedPath :  <code>string[]</code>\n\nHow to use this configuration:\n- If url contains any of the strings of allowedPath, url will be crawled.\n- If url matches the regex of any of the allowedPath, url will be crawled.\n\n**Default**: `['.*']`\n\n###### Example\n```ts\n{\n    discovery.allowedPath: [\"/amsterdam\", \"houses-to-rent\", \"house-[A-Z]+\"]\n}\n\n```\n<br></br>\n\n<a name=\"CrawlerDiscoveryConfig+removeQueryParams\"></a>\n#### discovery.removeQueryParams :  <code>boolean</code>\nIf true, it will trim the query parameters from the urls to discover. **Default**: `false`\n<br></br>\n\n<a name=\"CrawlerDiscoveryConfig+removeQueryParams\"></a>\n#### discovery.onLinksDiscovered :  <code>(response: HtmlResponse, links: string[]) => string[]</code>\n\nFunction that can be used to remove / add links to crawl. **Default**: `undefined`\n\n###### Example\n```ts\n{\n    discovery.onLinksDiscovered: (htmlResponse: HtmlResponse, links: string[]) => {\n        links.push('https://mycustomurl.com');\n        // We can use htmlResponse.$ to find links by css selectors.\n        return links;\n    }\n}\n```\n<br></br>\n## Examples\n\nYou can check some examples in the examples folder.\n\n## Roadmap\nFeatures to be implemented:\n\n- Store status and data in MongoDB.\n- Create more examples.\n- Add mode to retry errors.\n- Increase unit tests coverage.\n\n## Contributors\n**Main contributor**: [Juan Roldan](https://juanbroldan.com)\n\nThe Nodescrapy project welcomes all constructive contributions. \nContributions take many forms, from code for bug fixes and enhancements, to additions and fixes to documentation, additional tests, triaging incoming pull requests and issues, and more!\n\n## License\n\n[MIT](./LICENSE)","readmeFilename":"README.md","homepage":"https://github.com/juanroldanbrz/nodescrapy#readme","repository":{"type":"git","url":"git+https://github.com/juanroldanbrz/nodescrapy.git"},"bugs":{"url":"https://github.com/juanroldanbrz/nodescrapy/issues"}}