{"_id":"@dornick/parsers-ocr","name":"@dornick/parsers-ocr","dist-tags":{"latest":"0.1.0"},"versions":{"0.1.0":{"name":"@dornick/parsers-ocr","version":"0.1.0","description":"Lazy Tesseract.js OCR parser for Dornick. Extracts text from image uploads (JPEG, PNG) in the browser.","type":"module","main":"./dist/index.js","module":"./dist/index.js","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"scripts":{"build":"tsup","dev":"tsup --watch","test":"vitest run","typecheck":"tsc --noEmit","clean":"rm -rf dist","size":"size-limit"},"dependencies":{"@dornick/capabilities":"0.1.0","@dornick/parsers":"0.1.0","tesseract.js":"^5.1.1"},"devDependencies":{"@size-limit/file":"^11.1.6","size-limit":"^11.1.6","tsup":"^8.3.0","typescript":"^5.5.4","vitest":"^4.1.7"},"license":"MIT","sideEffects":false,"publishConfig":{"access":"public"},"size-limit":[{"name":"@dornick/parsers-ocr (gz)","path":"dist/index.js","limit":"1.0 KB","gzip":true}],"gitHead":"622e4ebed9c3aebadd808a2dd1f5f29c7d226040","_id":"@dornick/parsers-ocr@0.1.0","_nodeVersion":"26.7.0","_npmVersion":"11.19.0","dist":{"integrity":"sha512-8gCOeiIL1IJR8Iri+FOuF30RuLPcMtbH48sOyhuUbYb25BhVmfJAaLU/4sr0xcNX2R5fwZsAarvEpwySe2QPTw==","shasum":"4a44a9910be17ddbd0d171552fe73b62d4d02664","tarball":"https://registry.npmjs.org/@dornick/parsers-ocr/-/parsers-ocr-0.1.0.tgz","fileCount":7,"unpackedSize":14011,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQClcH+dklCTjGqWJAn27/OPsjB6qYVDDnrzRPVCqeofjgIgPfA8GwAswbpX/ULttqXdPk65TS8hw2yZTsTordivqbs="}]},"_npmUser":{"name":"dipankarsarkar","email":"me@dipankar.name"},"directories":{},"maintainers":[{"name":"dipankarsarkar","email":"me@dipankar.name"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/parsers-ocr_0.1.0_1788784936971_0.6396546299518988"},"_hasShrinkwrap":false}},"time":{"created":"2026-09-07T12:42:16.758Z","0.1.0":"2026-09-07T12:42:17.110Z","modified":"2026-09-07T12:42:17.401Z"},"maintainers":[{"name":"dipankarsarkar","email":"me@dipankar.name"}],"description":"Lazy Tesseract.js OCR parser for Dornick. Extracts text from image uploads (JPEG, PNG) in the browser.","license":"MIT","readme":"# @dornick/parsers-ocr\n\n**Lazy Tesseract.js OCR parser for Dornick — extract text from image uploads in the browser.**\n\nFor receipts, screenshots, ID photos, scanned forms — anywhere a photo carries text. Bypasses the multimodal token cost for OCR-shaped extraction.\n\n[![npm](https://img.shields.io/npm/v/@dornick/parsers-ocr.svg?style=flat&color=4c1)](https://www.npmjs.com/package/@dornick/parsers-ocr)\n[![license](https://img.shields.io/npm/l/@dornick/parsers-ocr.svg?style=flat&color=blue)](../../LICENSE)\n\n## Install\n\n```sh\nnpm i @dornick/parsers-ocr\n```\n\nDepends on [`tesseract.js`](https://www.npmjs.com/package/tesseract.js). The full WASM + English training data is ~5 MB and loads on first parse.\n\n## Use\n\n```ts\nimport { Dornick } from \"@dornick/core\";\nimport { tesseractParser, tesseractParserFor } from \"@dornick/parsers-ocr\";\n\nconst dornick = new Dornick({ config, transport });\n\ndornick.parsers.add(tesseractParser);                       // English default\ndornick.parsers.add(tesseractParserFor(\"fra\"));             // French\ndornick.parsers.add(tesseractParserFor([\"eng\", \"deu\"]));    // multi-lang\n```\n\n## Paths produced\n\n| Path | Content |\n|---|---|\n| `text.txt` | Plain extracted text |\n| `ocr.json` | `{ text, confidence, words: [{ text, confidence, x0, y0, x1, y1 }] }` |\n\nThe `ocr.json` includes per-word bounding boxes — useful for layout-aware extraction (receipts, IDs).\n\n## Capability gate\n\n`requires: [\"wasm\"]` — runs anywhere WebAssembly is supported. Performance is meaningfully better on browsers with `wasm.simd` and `wasm.threads`; without those, expect 2-5× slower OCR on the same image.\n\n## Pairs nicely with\n\n- [`@dornick/parsers-pdf`](https://www.npmjs.com/package/@dornick/parsers-pdf) — for image-only scanned PDFs (combine: pdf.js extracts images, OCR extracts text).\n- `@dornick/mounts-media` (planned) — for `CameraMount` capture → OCR pipelines.\n\n## See also\n\n- [Use case: In-browser OCR](https://docs.neullabs.com/dornick/use-cases/in-browser-ocr/)\n- [Use case: KYC / AML intake](https://docs.neullabs.com/dornick/use-cases/kyc-intake/)\n- [Concept: Parsers](https://docs.neullabs.com/dornick/concepts/parsers/)\n\n## License\n\nMIT — © Dipankar Sarkar\n","readmeFilename":"README.md","_rev":"1-482bc9e88a59c007d8617dc9fdf980b6"}