{"_id":"@d0paminedriven/pdfdown-ocr","_rev":"10-6791c57b19b8646d67671df1097548fe","name":"@d0paminedriven/pdfdown-ocr","dist-tags":{"latest":"0.9.9"},"versions":{"0.8.0":{"name":"@d0paminedriven/pdfdown-ocr","version":"0.8.0","license":"MIT","_id":"@d0paminedriven/pdfdown-ocr@0.8.0","maintainers":[{"name":"d0paminedriven","email":"andrew@windycitydevs.io"}],"homepage":"https://github.com/DopamineDriven/pdfdown#readme","bugs":{"url":"https://github.com/DopamineDriven/pdfdown/issues"},"os":["darwin","linux"],"dist":{"shasum":"b70cbe6e2e71d7e46d0627a8d0a52b855943f5f6","tarball":"https://registry.npmjs.org/@d0paminedriven/pdfdown-ocr/-/pdfdown-ocr-0.8.0.tgz","fileCount":1,"integrity":"sha512-xGOGYHeylRYmLm1jh1gcU1/bI6QMyw5hsJxQAScp53SY9O9JeqstbF7K7m6D/ZmcMzhYOTgxMCrAtRdlHHd/nA==","signatures":[{"sig":"MEYCIQC10CZ6UwnNNSPFVVzZz2GnffLEqYeReFV8pAIV4MdNngIhAMklIsNEvJQ+mpb9d4Rep7qWJo00mbrhCWO9qwSlC86U","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@d0paminedriven%2fpdfdown-ocr@0.8.0","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":798},"main":"index.js","napi":{"targets":["x86_64-apple-darwin","aarch64-apple-darwin","x86_64-unknown-linux-gnu","aarch64-unknown-linux-gnu"],"binaryName":"pdfdown_ocr"},"engines":{"node":">= 12.22.0 < 13 || >= 14.17.0 < 15 || >= 15.12.0 < 16 || >= 16.0.0"},"gitHead":"2cf270a0ec04025a1cdca06e3b01979312cfd326","_npmUser":{"name":"d0paminedriven","email":"andrew@windycitydevs.io"},"repository":{"url":"git+https://github.com/DopamineDriven/pdfdown.git","type":"git"},"_npmVersion":"11.6.2","description":"Rust powered PDF extraction for Node with OCR fallback (requires system tesseract).","directories":{},"_nodeVersion":"24.13.0","publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/pdfdown-ocr_0.8.0_1770111859881_0.2411190429459067","host":"s3://npm-registry-packages-npm-production"}},"0.9.0":{"name":"@d0paminedriven/pdfdown-ocr","version":"0.9.0","license":"MIT","_id":"@d0paminedriven/pdfdown-ocr@0.9.0","maintainers":[{"name":"d0paminedriven","email":"andrew@windycitydevs.io"}],"homepage":"https://github.com/DopamineDriven/pdfdown#readme","bugs":{"url":"https://github.com/DopamineDriven/pdfdown/issues"},"os":["darwin","linux"],"dist":{"shasum":"90fa6085279c6f1a216b406b3007201840a0918e","tarball":"https://registry.npmjs.org/@d0paminedriven/pdfdown-ocr/-/pdfdown-ocr-0.9.0.tgz","fileCount":4,"integrity":"sha512-c2p25mslNyjMhQwKXj9NPVEBR9DPT42aYtRZ30agLv1iMGB7DwEfsHxxjyU4hBOAeQzpUDM4NpxRvKQzOIo81w==","signatures":[{"sig":"MEYCIQD7KJK9tSgOrf5bQOWgHtE1AE0Pe+LuHf01yumPq0MyOgIhAKvDOxAHJplxOAqglyqvouC22KNqw+Jn09XYiB0w7ojZ","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@d0paminedriven%2fpdfdown-ocr@0.9.0","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":35558},"main":"index.js","napi":{"targets":["x86_64-apple-darwin","aarch64-apple-darwin","x86_64-unknown-linux-gnu","aarch64-unknown-linux-gnu"],"binaryName":"pdfdown_ocr"},"types":"./index.d.ts","engines":{"node":">= 12.22.0 < 13 || >= 14.17.0 < 15 || >= 15.12.0 < 16 || >= 16.0.0"},"gitHead":"5094969aee8ce530f08eb54a2d15798f3eecf3c5","_npmUser":{"name":"d0paminedriven","email":"andrew@windycitydevs.io"},"repository":{"url":"git+https://github.com/DopamineDriven/pdfdown.git","type":"git"},"_npmVersion":"11.6.2","description":"Rust powered PDF extraction for Node with OCR fallback (requires system tesseract).","directories":{},"_nodeVersion":"24.13.0","publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/pdfdown-ocr_0.9.0_1770121953779_0.7787466365020883","host":"s3://npm-registry-packages-npm-production"}},"0.9.1":{"name":"@d0paminedriven/pdfdown-ocr","version":"0.9.1","license":"MIT","_id":"@d0paminedriven/pdfdown-ocr@0.9.1","maintainers":[{"name":"d0paminedriven","email":"andrew@windycitydevs.io"}],"homepage":"https://github.com/DopamineDriven/pdfdown#readme","bugs":{"url":"https://github.com/DopamineDriven/pdfdown/issues"},"os":["darwin","linux"],"dist":{"shasum":"da53c1b7900b4382e14a299ddea1eaafeafbd79d","tarball":"https://registry.npmjs.org/@d0paminedriven/pdfdown-ocr/-/pdfdown-ocr-0.9.1.tgz","fileCount":4,"integrity":"sha512-sU5A41yimlBcn5N3fl1ZIH8ZPYMMveYtO2jnOgYcOuHomHJtxsmRejAi98pqtCDb/JJRZx1KMO8WDOLsXJuvfw==","signatures":[{"sig":"MEUCIQDCIDN9s2jcSJMZwYolWoPDCfS4ycdXhuSL8iTo1tnhFwIgdDRNuNH14L112ABpDOqF2GbQXNJmqi/x+GrIrobldak=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@d0paminedriven%2fpdfdown-ocr@0.9.1","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":41559},"main":"index.js","napi":{"targets":["x86_64-apple-darwin","aarch64-apple-darwin","x86_64-unknown-linux-gnu","aarch64-unknown-linux-gnu"],"binaryName":"pdfdown_ocr"},"types":"./index.d.ts","engines":{"node":">= 12.22.0 < 13 || >= 14.17.0 < 15 || >= 15.12.0 < 16 || >= 16.0.0"},"gitHead":"baef1378b9f74390a0ea7cf080811ba32be66316","_npmUser":{"name":"d0paminedriven","email":"andrew@windycitydevs.io"},"repository":{"url":"git+https://github.com/DopamineDriven/pdfdown.git","type":"git"},"_npmVersion":"11.6.2","description":"Rust powered PDF extraction for Node with OCR fallback (requires system tesseract).","directories":{},"_nodeVersion":"24.13.0","publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/pdfdown-ocr_0.9.1_1770168371622_0.26570530449203655","host":"s3://npm-registry-packages-npm-production"}},"0.9.2":{"name":"@d0paminedriven/pdfdown-ocr","version":"0.9.2","license":"MIT","_id":"@d0paminedriven/pdfdown-ocr@0.9.2","maintainers":[{"name":"d0paminedriven","email":"andrew@windycitydevs.io"}],"homepage":"https://github.com/DopamineDriven/pdfdown#readme","bugs":{"url":"https://github.com/DopamineDriven/pdfdown/issues"},"os":["darwin","linux"],"dist":{"shasum":"1e9e641060ff01b3edeeda7ec3efbc009e159b2a","tarball":"https://registry.npmjs.org/@d0paminedriven/pdfdown-ocr/-/pdfdown-ocr-0.9.2.tgz","fileCount":4,"integrity":"sha512-E7IvEjy38hictTN2/zpcC97UGRSNu0jqlCa7jxXUbd3iE0PcJITjax+myUf3WQi7AVaOz0wXo+wBJjr9tI/frA==","signatures":[{"sig":"MEQCIAckeAV1FTJj2BKHKn8ZyE0HSzSCQ088+ANSdi1z92dTAiBi1aRog344OQ6QwpbVhRknlgIDMVYEPIsIafb+1UGPHg==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@d0paminedriven%2fpdfdown-ocr@0.9.2","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":41559},"main":"index.js","napi":{"targets":["x86_64-apple-darwin","aarch64-apple-darwin","x86_64-unknown-linux-gnu","aarch64-unknown-linux-gnu"],"binaryName":"pdfdown_ocr"},"types":"./index.d.ts","engines":{"node":">= 12.22.0 < 13 || >= 14.17.0 < 15 || >= 15.12.0 < 16 || >= 16.0.0"},"gitHead":"d72d6be9d2eadd404810151afee0a97753bfd890","_npmUser":{"name":"d0paminedriven","email":"andrew@windycitydevs.io"},"repository":{"url":"git+https://github.com/DopamineDriven/pdfdown.git","type":"git"},"_npmVersion":"11.6.2","description":"Rust powered PDF extraction for Node with OCR fallback (requires system tesseract).","directories":{},"_nodeVersion":"24.13.0","publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/pdfdown-ocr_0.9.2_1770422485061_0.8095533275688307","host":"s3://npm-registry-packages-npm-production"}},"0.9.3":{"name":"@d0paminedriven/pdfdown-ocr","version":"0.9.3","license":"MIT","_id":"@d0paminedriven/pdfdown-ocr@0.9.3","maintainers":[{"name":"d0paminedriven","email":"andrew@windycitydevs.io"}],"homepage":"https://github.com/DopamineDriven/pdfdown#readme","bugs":{"url":"https://github.com/DopamineDriven/pdfdown/issues"},"os":["darwin","linux"],"dist":{"shasum":"c24b3b13c1e621cf89030756be0e7a7acb2f400e","tarball":"https://registry.npmjs.org/@d0paminedriven/pdfdown-ocr/-/pdfdown-ocr-0.9.3.tgz","fileCount":4,"integrity":"sha512-kiQx4VyMZo0sP5BEX3vq9h6ynAvWrKwFqE4KSkxfK2cwoWG1AVlVaeaSc7X8JPikbYOPqpErjqgZO31XxKFOyw==","signatures":[{"sig":"MEUCIQCwenCljCmY3kDVdV+Y04q3QHMCQiRcO4DReOZhmECuOwIgai8ltzF1LtKPUEefXgPpG8S0xP24OPy1KW29OK+owlY=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@d0paminedriven%2fpdfdown-ocr@0.9.3","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":42262},"main":"index.js","napi":{"targets":["x86_64-apple-darwin","aarch64-apple-darwin","x86_64-unknown-linux-gnu","aarch64-unknown-linux-gnu"],"binaryName":"pdfdown_ocr"},"types":"./index.d.ts","engines":{"node":">= 12.22.0 < 13 || >= 14.17.0 < 15 || >= 15.12.0 < 16 || >= 16.0.0"},"gitHead":"05f88370de9f9bea2111c82bd4fa2920157f205f","_npmUser":{"name":"d0paminedriven","email":"andrew@windycitydevs.io"},"repository":{"url":"git+https://github.com/DopamineDriven/pdfdown.git","type":"git"},"_npmVersion":"11.6.2","description":"Rust powered PDF extraction for Node with OCR fallback (requires system tesseract).","directories":{},"_nodeVersion":"24.13.0","publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/pdfdown-ocr_0.9.3_1770458091025_0.702394361026371","host":"s3://npm-registry-packages-npm-production"}},"0.9.5":{"name":"@d0paminedriven/pdfdown-ocr","version":"0.9.5","license":"MIT","_id":"@d0paminedriven/pdfdown-ocr@0.9.5","maintainers":[{"name":"d0paminedriven","email":"andrew@windycitydevs.io"}],"homepage":"https://github.com/DopamineDriven/pdfdown#readme","bugs":{"url":"https://github.com/DopamineDriven/pdfdown/issues"},"os":["darwin","linux"],"dist":{"shasum":"14ff77e05b992e6a62588c1a002861b6f0740585","tarball":"https://registry.npmjs.org/@d0paminedriven/pdfdown-ocr/-/pdfdown-ocr-0.9.5.tgz","fileCount":4,"integrity":"sha512-AK35B109/eYxh+RXLHRzbqYR3H7mS1ju0itCLemU16jO1wYRKNnJ6Ny7P9Ff7DpxCB7oTgXHiiU+s2RqukY1eg==","signatures":[{"sig":"MEYCIQDs/eWe7oUmugMy9d+7Xl9MfRDTF0r5moz5NL+EdInm3gIhAKNVQiua3MXvFjknkNxWVbMyiAsYcCe8HDIHng545zfB","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@d0paminedriven%2fpdfdown-ocr@0.9.5","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":42523},"main":"index.js","napi":{"targets":["x86_64-apple-darwin","aarch64-apple-darwin","x86_64-unknown-linux-gnu","aarch64-unknown-linux-gnu"],"binaryName":"pdfdown_ocr"},"types":"./index.d.ts","engines":{"node":">= 12.22.0 < 13 || >= 14.17.0 < 15 || >= 15.12.0 < 16 || >= 16.0.0"},"gitHead":"d11e0289bd9cf7f2c72675f599688fda100f65d2","_npmUser":{"name":"d0paminedriven","email":"andrew@windycitydevs.io"},"repository":{"url":"git+https://github.com/DopamineDriven/pdfdown.git","type":"git"},"_npmVersion":"11.6.2","description":"Rust powered PDF extraction for Node with OCR fallback (requires system tesseract).","directories":{},"_nodeVersion":"24.13.0","publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"optionalDependencies":{"@d0paminedriven/pdfdown-ocr-darwin-x64":"0.9.5","@d0paminedriven/pdfdown-ocr-darwin-arm64":"0.9.5","@d0paminedriven/pdfdown-ocr-linux-x64-gnu":"0.9.5","@d0paminedriven/pdfdown-ocr-linux-arm64-gnu":"0.9.5"},"_npmOperationalInternal":{"tmp":"tmp/pdfdown-ocr_0.9.5_1770641500340_0.8756850153296587","host":"s3://npm-registry-packages-npm-production"}},"0.9.6":{"name":"@d0paminedriven/pdfdown-ocr","version":"0.9.6","license":"MIT","_id":"@d0paminedriven/pdfdown-ocr@0.9.6","maintainers":[{"name":"d0paminedriven","email":"andrew@windycitydevs.io"}],"homepage":"https://github.com/DopamineDriven/pdfdown#readme","bugs":{"url":"https://github.com/DopamineDriven/pdfdown/issues"},"os":["darwin","linux"],"dist":{"shasum":"6030d14c3e94607a0af5664c0f42f174c9286c16","tarball":"https://registry.npmjs.org/@d0paminedriven/pdfdown-ocr/-/pdfdown-ocr-0.9.6.tgz","fileCount":4,"integrity":"sha512-5WpOV7lWs33960NUNH+PXSBPc5ufIKk9DtHeisBPObwidrrYp+FdA/WJ7Nd6jIKEPz8w3GoDeVcxQnp9vVTjUA==","signatures":[{"sig":"MEQCIEoije+uT26kQB659q6QxwKmVQMVoTV5E+LCyxldbQV5AiAJmQI4tr0z8Oh5OCRE4ckh0qsPYf//KHsK1luRHtZ8LQ==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@d0paminedriven%2fpdfdown-ocr@0.9.6","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":42518},"main":"index.js","napi":{"targets":["x86_64-apple-darwin","aarch64-apple-darwin","x86_64-unknown-linux-gnu","aarch64-unknown-linux-gnu"],"binaryName":"pdfdown_ocr"},"types":"./index.d.ts","engines":{"node":">= 12.22.0 < 13 || >= 14.17.0 < 15 || >= 15.12.0 < 16 || >= 16.0.0"},"gitHead":"d0c0abca8b3dce09d9b0a2276a0d9c95000174dd","_npmUser":{"name":"d0paminedriven","email":"andrew@windycitydevs.io"},"repository":{"url":"git+https://github.com/DopamineDriven/pdfdown.git","type":"git"},"_npmVersion":"11.6.2","description":"Rust powered PDF extraction for Node with OCR fallback (requires system tesseract).","directories":{},"_nodeVersion":"24.13.0","publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"optionalDependencies":{"@d0paminedriven/pdfdown-ocr-darwin-x64":"0.9.6","@d0paminedriven/pdfdown-ocr-darwin-arm64":"0.9.6","@d0paminedriven/pdfdown-ocr-linux-x64-gnu":"0.9.6","@d0paminedriven/pdfdown-ocr-linux-arm64-gnu":"0.9.6"},"_npmOperationalInternal":{"tmp":"tmp/pdfdown-ocr_0.9.6_1770651161431_0.5291441496267417","host":"s3://npm-registry-packages-npm-production"}},"0.9.7":{"name":"@d0paminedriven/pdfdown-ocr","version":"0.9.7","license":"MIT","_id":"@d0paminedriven/pdfdown-ocr@0.9.7","maintainers":[{"name":"d0paminedriven","email":"andrew@windycitydevs.io"}],"homepage":"https://github.com/DopamineDriven/pdfdown#readme","bugs":{"url":"https://github.com/DopamineDriven/pdfdown/issues"},"os":["darwin","linux"],"dist":{"shasum":"5f98f2f4877cb99fad7a7ce9e1cda08193ca741d","tarball":"https://registry.npmjs.org/@d0paminedriven/pdfdown-ocr/-/pdfdown-ocr-0.9.7.tgz","fileCount":4,"integrity":"sha512-Kxxpgq/rI4E3CwS/iPL+2flQ32d6N8Xsx6jQ/AcqcSYxh5qVA7QNIpZzot2+OS2uZw1BNJreRNUZsJEElL7zVg==","signatures":[{"sig":"MEUCIQCNepU1FJcrDqrbrFCWDqZvH031fhJOmjdItQAXVtJxggIgYoj1pLuPwkA38rqdtQ5X0Y0v5OHufR1HYi3FBI7xQVE=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@d0paminedriven%2fpdfdown-ocr@0.9.7","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":42518},"main":"index.js","napi":{"targets":["x86_64-apple-darwin","aarch64-apple-darwin","x86_64-unknown-linux-gnu","aarch64-unknown-linux-gnu"],"binaryName":"pdfdown_ocr"},"types":"./index.d.ts","engines":{"node":">= 12.22.0 < 13 || >= 14.17.0 < 15 || >= 15.12.0 < 16 || >= 16.0.0"},"gitHead":"e212407ad63a8206c3c15f0b7684eab63a1d2c1d","_npmUser":{"name":"d0paminedriven","email":"andrew@windycitydevs.io"},"repository":{"url":"git+https://github.com/DopamineDriven/pdfdown.git","type":"git"},"_npmVersion":"11.6.2","description":"Rust powered PDF extraction for Node with OCR fallback (requires system tesseract).","directories":{},"_nodeVersion":"24.13.0","publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"optionalDependencies":{"@d0paminedriven/pdfdown-ocr-darwin-x64":"0.9.7","@d0paminedriven/pdfdown-ocr-darwin-arm64":"0.9.7","@d0paminedriven/pdfdown-ocr-linux-x64-gnu":"0.9.7","@d0paminedriven/pdfdown-ocr-linux-arm64-gnu":"0.9.7"},"_npmOperationalInternal":{"tmp":"tmp/pdfdown-ocr_0.9.7_1770710139727_0.7222113962251502","host":"s3://npm-registry-packages-npm-production"}},"0.9.8":{"name":"@d0paminedriven/pdfdown-ocr","version":"0.9.8","license":"MIT","_id":"@d0paminedriven/pdfdown-ocr@0.9.8","maintainers":[{"name":"d0paminedriven","email":"andrew@windycitydevs.io"}],"homepage":"https://github.com/DopamineDriven/pdfdown#readme","bugs":{"url":"https://github.com/DopamineDriven/pdfdown/issues"},"os":["darwin","linux"],"dist":{"shasum":"0e6991dcb3c4d1732b93493903d4cfd4b6a8a497","tarball":"https://registry.npmjs.org/@d0paminedriven/pdfdown-ocr/-/pdfdown-ocr-0.9.8.tgz","fileCount":4,"integrity":"sha512-EdiLvOq9NA1G9lsOBe6/uFyA7gw843SDEUFqfeqppMvnPPCfbXACKNwMrcS+DuEhfVE6NVCCFXJFOuBu/0QYoQ==","signatures":[{"sig":"MEUCIDpNwCyTGRGORyv1e6c3/3B7++1CrnHzTgDCU6m78SrDAiEA6wCxFEZc8IcQfSUWI/hy1SpGPkIXjcRBuRHqncd/eIQ=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@d0paminedriven%2fpdfdown-ocr@0.9.8","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":42102},"main":"index.js","napi":{"targets":["x86_64-apple-darwin","aarch64-apple-darwin","x86_64-unknown-linux-gnu","aarch64-unknown-linux-gnu"],"binaryName":"pdfdown_ocr"},"types":"./index.d.ts","engines":{"node":">= 12.22.0 < 13 || >= 14.17.0 < 15 || >= 15.12.0 < 16 || >= 16.0.0"},"gitHead":"fb6ead16824ca881c6f9d667e81dacdee607d59a","_npmUser":{"name":"d0paminedriven","email":"andrew@windycitydevs.io"},"repository":{"url":"git+https://github.com/DopamineDriven/pdfdown.git","type":"git"},"_npmVersion":"11.6.2","description":"Rust powered PDF extraction for Node with OCR fallback (requires system tesseract).","directories":{},"_nodeVersion":"24.13.0","publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"optionalDependencies":{"@d0paminedriven/pdfdown-ocr-darwin-x64":"0.9.8","@d0paminedriven/pdfdown-ocr-darwin-arm64":"0.9.8","@d0paminedriven/pdfdown-ocr-linux-x64-gnu":"0.9.8","@d0paminedriven/pdfdown-ocr-linux-arm64-gnu":"0.9.8"},"_npmOperationalInternal":{"tmp":"tmp/pdfdown-ocr_0.9.8_1771431597795_0.6925810699025248","host":"s3://npm-registry-packages-npm-production"}},"0.9.9":{"name":"@d0paminedriven/pdfdown-ocr","version":"0.9.9","description":"Rust powered PDF extraction for Node with OCR fallback (requires system tesseract).","main":"index.js","repository":{"type":"git","url":"git+https://github.com/DopamineDriven/pdfdown.git"},"license":"MIT","os":["darwin","linux"],"napi":{"binaryName":"pdfdown_ocr","targets":["x86_64-apple-darwin","aarch64-apple-darwin","x86_64-unknown-linux-gnu","aarch64-unknown-linux-gnu"]},"engines":{"node":">= 12.22.0 < 13 || >= 14.17.0 < 15 || >= 15.12.0 < 16 || >= 16.0.0"},"publishConfig":{"registry":"https://registry.npmjs.org/","access":"public"},"optionalDependencies":{"@d0paminedriven/pdfdown-ocr-darwin-x64":"0.9.9","@d0paminedriven/pdfdown-ocr-darwin-arm64":"0.9.9","@d0paminedriven/pdfdown-ocr-linux-x64-gnu":"0.9.9","@d0paminedriven/pdfdown-ocr-linux-arm64-gnu":"0.9.9"},"gitHead":"6fb0e474385a13bb621e2d678c4c619b9fc0ea6d","types":"./index.d.ts","_id":"@d0paminedriven/pdfdown-ocr@0.9.9","bugs":{"url":"https://github.com/DopamineDriven/pdfdown/issues"},"homepage":"https://github.com/DopamineDriven/pdfdown#readme","_nodeVersion":"24.13.0","_npmVersion":"11.6.2","dist":{"integrity":"sha512-x87A2jh4NzOipqKJ0V2HF1kByRmtu2+bwEdA8T2UeGTQIwEFnoXNkaRxCptZxcMYsZB719CJA/hR3UShGHo+MQ==","shasum":"938277350e2c3a40c71f19e054ed09baa8137267","tarball":"https://registry.npmjs.org/@d0paminedriven/pdfdown-ocr/-/pdfdown-ocr-0.9.9.tgz","fileCount":4,"unpackedSize":41961,"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@d0paminedriven%2fpdfdown-ocr@0.9.9","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQC59Sx5HXAy156zj6AeGp+pC7cFoGA/QGBhFCKTHbrZ0QIhAJrJxuTq9eu+kVqVOLoI+lXCvvLLcTK5ppx94V+mEGkz"}]},"_npmUser":{"name":"d0paminedriven","email":"andrew@windycitydevs.io"},"directories":{},"maintainers":[{"name":"d0paminedriven","email":"andrew@windycitydevs.io"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/pdfdown-ocr_0.9.9_1771636779267_0.07292326470187316"},"_hasShrinkwrap":false}},"time":{"created":"2026-02-03T09:44:19.808Z","modified":"2026-02-21T01:19:39.637Z","0.8.0":"2026-02-03T09:44:20.004Z","0.9.0":"2026-02-03T12:32:33.917Z","0.9.1":"2026-02-04T01:26:11.758Z","0.9.2":"2026-02-07T00:01:25.193Z","0.9.3":"2026-02-07T09:54:51.165Z","0.9.5":"2026-02-09T12:51:40.473Z","0.9.6":"2026-02-09T15:32:41.585Z","0.9.7":"2026-02-10T07:55:39.858Z","0.9.8":"2026-02-18T16:19:57.929Z","0.9.9":"2026-02-21T01:19:39.396Z"},"bugs":{"url":"https://github.com/DopamineDriven/pdfdown/issues"},"license":"MIT","homepage":"https://github.com/DopamineDriven/pdfdown#readme","repository":{"type":"git","url":"git+https://github.com/DopamineDriven/pdfdown.git"},"description":"Rust powered PDF extraction for Node with OCR fallback (requires system tesseract).","maintainers":[{"name":"d0paminedriven","email":"andrew@windycitydevs.io"}],"readme":"# `@d0paminedriven/pdfdown-ocr`\n\nRust-powered PDF extraction for Node.js with Tesseract OCR fallback for image-only pages. A superset of [`@d0paminedriven/pdfdown`](https://www.npmjs.com/package/@d0paminedriven/pdfdown) -- includes all base extraction APIs (text, images, annotations, structured text, metadata) plus OCR.\n\n**System requirement:** [Tesseract](https://github.com/tesseract-ocr/tesseract) 5.x must be installed on the host.\n\n## Install\n\n```bash\nnpm install @d0paminedriven/pdfdown-ocr\n```\n\n### Tesseract setup\n\n```bash\n# Ubuntu/Debian (22.04 ships pre v5 > -- use the PPA for 5.x)\nsudo add-apt-repository ppa:alex-p/tesseract-ocr5\nsudo apt update\nsudo apt install tesseract-ocr tesseract-ocr-eng -y\n# Optional: all language packs\n# sudo apt install tesseract-ocr-all\n\n# macOS\nbrew install tesseract\n\n# Arch\nsudo pacman -S tesseract tesseract-data-eng\n```\n\nVerify with `tesseract --version` -- you should see 5.x.\n\n### Tessdata auto-detection\n\nThe package automatically detects the tessdata directory at runtime by parsing the output of `tesseract --list-langs`. The detected path is cached for the lifetime of the process using a `OnceLock<Option<String>>` -- no global environment mutation, fully thread-safe.\n\n**Resolution order:**\n\n1. `TESSDATA_PREFIX` environment variable (if set, used as-is -- no auto-detection runs)\n2. Auto-detection via `tesseract --list-langs` (parses the path from `List of available languages in \"/path/to/tessdata/\"`)\n3. Tesseract's compiled-in default (if neither of the above yields a path)\n\nMost users will not need to set `TESSDATA_PREFIX` at all. The auto-detection handles standard installations on Ubuntu (`/usr/share/tesseract-ocr/5/tessdata/`), macOS Homebrew (`/opt/homebrew/share/tessdata/`), Arch, and any other layout where `tesseract` is on `PATH`.\n\nSet `TESSDATA_PREFIX` explicitly only if:\n\n- Tesseract is not on `PATH` but the tessdata directory exists elsewhere\n- You want to override the detected path (e.g., pointing to a custom-trained data directory)\n\n```bash\n# Override example (not usually needed)\nexport TESSDATA_PREFIX=\"/opt/custom/tessdata\"\n```\n\n## API\n\nThis package exports everything from `@d0paminedriven/pdfdown` (text, images, annotations, structured text, metadata -- both sync and async), plus the OCR-specific APIs below. See the [base package docs](https://www.npmjs.com/package/@d0paminedriven/pdfdown) for the full base API.\n\n### OCR standalone functions\n\n```typescript\n// Per-page OCR text extraction\nexport declare function extractTextWithOcrPerPage(\n  buffer: Buffer,\n  opts?: OcrOptions,\n): Array<OcrPageText>\n\nexport declare function extractTextWithOcrPerPageAsync(\n  buffer: Buffer,\n  opts?: OcrOptions,\n): Promise<Array<OcrPageText>>\n\n// Full document extraction with OCR text fallback\nexport declare function pdfDocumentOcr(\n  buffer: Buffer,\n  opts?: OcrOptions,\n): PdfDocumentOcr\n\nexport declare function pdfDocumentOcrAsync(\n  buffer: Buffer,\n  opts?: OcrOptions,\n): Promise<PdfDocumentOcr>\n```\n\n### `PdfDown` class (includes OCR methods)\n\n```typescript\nexport declare class PdfDown {\n  constructor(buffer: Buffer)\n\n  // ── Base methods ──\n  textPerPage(): Array<PageText>\n  textPerPageAsync(): Promise<Array<PageText>>\n  imagesPerPage(): Array<PageImage>\n  imagesPerPageAsync(): Promise<Array<PageImage>>\n  annotationsPerPage(): Array<PageAnnotation>\n  annotationsPerPageAsync(): Promise<Array<PageAnnotation>>\n  structuredText(): Array<StructuredPageText>\n  structuredTextAsync(): Promise<Array<StructuredPageText>>\n  metadata(): PdfMeta\n  metadataAsync(): Promise<PdfMeta>\n  document(): PdfDocument\n  documentAsync(): Promise<PdfDocument>\n\n  // ── OCR methods ──\n  textWithOcrPerPage(opts?: OcrOptions): Array<OcrPageText>\n  textWithOcrPerPageAsync(opts?: OcrOptions): Promise<Array<OcrPageText>>\n  documentOcr(opts?: OcrOptions): PdfDocumentOcr\n  documentOcrAsync(opts?: OcrOptions): Promise<PdfDocumentOcr>\n}\n```\n\n### Types\n\n```typescript\nexport const enum TextSource {\n  Native = 'Native',\n  Ocr = 'Ocr',\n}\n\nexport interface OcrPageText {\n  page: number\n  text: string\n  source: TextSource\n}\n\nexport interface OcrStructuredPageText {\n  page: number\n  header: string\n  body: string\n  footer: string\n  source: TextSource\n}\n\nexport interface OcrOptions {\n  lang?: string        // Tesseract language code, default \"eng\"\n  minTextLength?: number // non-whitespace char threshold before OCR fallback, default 1\n  maxThreads?: number  // cap on Rayon threads for OCR parallelism, default 4, clamped to [1, available CPUs]\n}\n\nexport interface PdfDocumentOcr {\n  version: string\n  isLinearized: boolean\n  pageCount: number\n  creator?: string\n  producer?: string\n  creationDate?: string\n  modificationDate?: string\n  totalImages: number\n  totalAnnotations: number\n  imagePages: Array<number>\n  annotationPages: Array<number>\n  text: Array<OcrPageText>\n  structuredText: Array<OcrStructuredPageText>\n  images: Array<PageImage>\n  annotations: Array<PageAnnotation>\n}\n```\n\n## Usage\n\n> **Use the async API for OCR.** The sync variants block the Node.js event loop for the duration of OCR processing, which can be significant for multi-page scanned documents.\n\n### Standalone\n\n```typescript\nimport { readFile } from 'fs/promises'\nimport { extractTextWithOcrPerPageAsync } from '@d0paminedriven/pdfdown-ocr'\n\nconst pdf = await readFile('scanned-document.pdf')\nconst pages = await extractTextWithOcrPerPageAsync(pdf, { lang: 'eng', minTextLength: 10 })\n\nfor (const { page, text, source } of pages) {\n  console.log(`Page ${page} [${source}]: ${text.slice(0, 100)}...`)\n}\n```\n\n### Class-based (parse once, extract many)\n\n```typescript\nimport { readFile } from 'fs/promises'\nimport { PdfDown } from '@d0paminedriven/pdfdown-ocr'\n\nconst pdf = new PdfDown(await readFile('scanned-document.pdf'))\n\n// OCR text extraction\nconst pages = await pdf.textWithOcrPerPageAsync({ lang: 'eng', minTextLength: 10 })\n\n// All base methods work too\nconst images = await pdf.imagesPerPageAsync()\nconst meta = pdf.metadata()\n```\n\n### Extract everything with OCR in one call\n\n```typescript\nimport { readFile } from 'fs/promises'\nimport { PdfDown } from '@d0paminedriven/pdfdown-ocr'\n\nconst pdf = new PdfDown(await readFile('scanned-document.pdf'))\nconst result = await pdf.documentOcrAsync({ minTextLength: 10 })\n\n// result.text         — OcrPageText[] (page, text, source per page)\n// result.structuredText — OcrStructuredPageText[] (header/body/footer + source per page)\n// result.images       — PageImage[] (decoded PNGs with dimensions and color space)\n// result.annotations  — PageAnnotation[] (links, destinations, rects)\n// result.pageCount, result.version, result.creator, ...\n```\n\n### Combined: OCR text + images for multimodal pipelines\n\n```typescript\nimport { readFile } from 'fs/promises'\nimport { PdfDown } from '@d0paminedriven/pdfdown-ocr'\n\nconst pdf = new PdfDown(await readFile('scanned-document.pdf'))\n\nconst [ocrText, images] = await Promise.all([\n  pdf.textWithOcrPerPageAsync({ minTextLength: 10 }),\n  pdf.imagesPerPageAsync(),\n])\n\nconst imagesByPage = Map.groupBy(images, (img) => img.page)\n\nfor (const { page, text, source } of ocrText) {\n  const pageImages = (imagesByPage.get(page) ?? []).map((img) => ({\n    dataUrl: `data:image/png;base64,${img.data.toString('base64')}`,\n    width: img.width,\n    height: img.height,\n  }))\n  // Send { page, text, source, images: pageImages } to your embedding pipeline\n}\n```\n\n### `document()` vs `documentOcr()`\n\nBoth methods extract everything from a PDF in a single call. The difference is how text is extracted:\n\n| Method | Text extraction | Return type | Use when |\n|--------|----------------|-------------|----------|\n| `document()` / `documentAsync()` | Native PDF text only | `PdfDocument` | PDF has selectable text |\n| `documentOcr()` / `documentOcrAsync()` | Native with OCR fallback | `PdfDocumentOcr` | PDF may contain scanned/image-only pages |\n\n`PdfDocumentOcr` uses `OcrPageText` (with `source: 'Native' | 'Ocr'`) and `OcrStructuredPageText` (with header/body/footer split plus source) instead of the base `PageText` and `StructuredPageText` types. Images, annotations, and metadata are identical in both.\n\n## How it works\n\n1. **Text extraction:** Each page is first attempted with native PDF text extraction. If a page yields fewer non-whitespace characters than `minTextLength`, its embedded images are decoded and fed to Tesseract for OCR. Each result is tagged with `source: 'Native'` or `source: 'Ocr'`.\n\n2. **Structured text:** After text extraction, repeated header/footer lines are detected across pages using frequency analysis (requires 3+ pages). Each page's text is split into `header`, `body`, and `footer` sections. For OCR results, the `source` tag is preserved so you know whether each page's content came from native extraction or OCR.\n\n3. **Parallelism:** OCR runs on a dedicated capped Rayon thread pool (default 4 threads, configurable via `maxThreads`) to prevent CPU oversubscription. Text extraction, image extraction, and annotation extraction run concurrently via `rayon::join` when using `documentOcr` / `documentOcrAsync`.\n\n4. **Tessdata discovery:** On first OCR invocation, the tessdata path is resolved once and cached in a `OnceLock`. The `TESSDATA_PREFIX` environment variable is checked first; if unset, `tesseract --list-langs` is executed and its output is parsed to extract the path. No environment variables are mutated -- the path is passed directly to Tesseract's init function.\n\n## Supported platforms\n\nPrebuilt binaries are provided for:\n\n- macOS (x64, ARM64)\n- Linux glibc (x64, ARM64)\n\n## Relationship to `@d0paminedriven/pdfdown`\n\nSame Rust codebase, compiled with the `ocr` Cargo feature flag enabled. This package is a strict superset -- you can use it as a drop-in replacement for the base package if you need OCR capabilities.\n\n## License\n\nMIT\n","readmeFilename":"README.md"}