{"_id":"@dornick/parsers-pdf","name":"@dornick/parsers-pdf","dist-tags":{"latest":"0.1.0"},"versions":{"0.1.0":{"name":"@dornick/parsers-pdf","version":"0.1.0","description":"Lazy pdf.js parser for Dornick. Extracts text from PDF uploads under /user/files/<id>/parsed/.","type":"module","main":"./dist/index.js","module":"./dist/index.js","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"scripts":{"build":"tsup","dev":"tsup --watch","test":"vitest run","typecheck":"tsc --noEmit","clean":"rm -rf dist","size":"size-limit"},"dependencies":{"@dornick/capabilities":"0.1.0","@dornick/parsers":"0.1.0","pdfjs-dist":"^4.7.76"},"devDependencies":{"@size-limit/file":"^11.1.6","size-limit":"^11.1.6","tsup":"^8.3.0","typescript":"^5.5.4","vitest":"^4.1.7"},"license":"MIT","sideEffects":false,"publishConfig":{"access":"public"},"size-limit":[{"name":"@dornick/parsers-pdf (gz)","path":"dist/index.js","limit":"1.0 KB","gzip":true}],"gitHead":"622e4ebed9c3aebadd808a2dd1f5f29c7d226040","_id":"@dornick/parsers-pdf@0.1.0","_nodeVersion":"26.7.0","_npmVersion":"11.19.0","dist":{"integrity":"sha512-v93AeA8JhBux6Hr5qIuJ1JzD7pk+aZtbOFTta4G3j4NJVz6ZQtMATBEkXlA/mO/Q/vxdG2NjlsofCcF4CbXErw==","shasum":"4d57cf6791ad119586bef00a44c6851ec4ae1e46","tarball":"https://registry.npmjs.org/@dornick/parsers-pdf/-/parsers-pdf-0.1.0.tgz","fileCount":7,"unpackedSize":11400,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIFPsuwQRJjpEcQj3l5Izcp3OgLzuDI1qpCAOYehtrurnAiEAhq7tXfaI6sSoPpz5d9VPYhYWguICowcoeIxuE0eCwbE="}]},"_npmUser":{"name":"dipankarsarkar","email":"me@dipankar.name"},"directories":{},"maintainers":[{"name":"dipankarsarkar","email":"me@dipankar.name"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/parsers-pdf_0.1.0_1788784940862_0.186083984413127"},"_hasShrinkwrap":false}},"time":{"created":"2026-09-07T12:42:20.592Z","0.1.0":"2026-09-07T12:42:20.988Z","modified":"2026-09-07T12:42:21.429Z"},"maintainers":[{"name":"dipankarsarkar","email":"me@dipankar.name"}],"description":"Lazy pdf.js parser for Dornick. Extracts text from PDF uploads under /user/files/<id>/parsed/.","license":"MIT","readme":"# @dornick/parsers-pdf\n\n**Lazy pdf.js parser for Dornick — extract text from PDF uploads in the browser.**\n\nRegisters a `FileParser` that materialises `text.txt` and `by-page.jsonl` under `/user/files/<id>/parsed/` on first read. The full pdf.js library only downloads when the agent actually reads one of those paths.\n\n[![npm](https://img.shields.io/npm/v/@dornick/parsers-pdf.svg?style=flat&color=4c1)](https://www.npmjs.com/package/@dornick/parsers-pdf)\n[![license](https://img.shields.io/npm/l/@dornick/parsers-pdf.svg?style=flat&color=blue)](../../LICENSE)\n\n## Install\n\n```sh\nnpm i @dornick/parsers-pdf\n```\n\nPulls in `pdfjs-dist` (~700 KB gz when loaded). The Dornick shim itself is < 1 KB.\n\n## Use\n\n```ts\nimport { Dornick } from \"@dornick/core\";\nimport { pdfTextParser } from \"@dornick/parsers-pdf\";\n\nconst dornick = new Dornick({ config, transport });\ndornick.parsers.add(pdfTextParser);\n```\n\nThe agent now sees `/user/files/<id>/parsed/text.txt` for every uploaded PDF. The first read triggers a one-time load of pdf.js (~700 KB gz, cached after).\n\n## Paths produced\n\n| Path | Content |\n|---|---|\n| `text.txt` | Whole-document text, page-joined, whitespace-collapsed |\n| `by-page.jsonl` | One JSON object per page: `{ page: 1, text: \"...\" }` |\n\n## Capability gate\n\n`requires: [\"wasm\"]` — runs anywhere WebAssembly is supported. The legacy pdf.js entry point includes a fake-worker fallback, so no separate worker URL configuration is needed for v1.\n\nFor better perf in production, the integrator can pre-configure a real worker URL:\n\n```ts\nimport * as pdfjs from \"pdfjs-dist/legacy/build/pdf.mjs\";\npdfjs.GlobalWorkerOptions.workerSrc = \"/path/to/pdf.worker.mjs\";\n```\n\nThe parser picks up the integrator's `workerSrc` if already set.\n\n## Pairs nicely with\n\n- [`@dornick/parsers-ocr`](https://www.npmjs.com/package/@dornick/parsers-ocr) — for image-only PDFs where text extraction fails.\n- [Use case: Local PDF Q&A](https://docs.neullabs.com/dornick/use-cases/pdf-qa/)\n\n## License\n\nMIT — © Dipankar Sarkar\n","readmeFilename":"README.md","_rev":"1-1ffb98141579bc182124e7e3b950c5ee"}