{"_id":"@alexstep/doc-extract","_rev":"3-d9c9018a8be66886ae2452274c434143","name":"@alexstep/doc-extract","dist-tags":{"latest":"0.1.5"},"versions":{"0.1.2":{"name":"@alexstep/doc-extract","version":"0.1.2","keywords":["napi-rs","pdf","docx","text-extraction","rust","native"],"license":"MIT","_id":"@alexstep/doc-extract@0.1.2","maintainers":[{"name":"kellas","email":"alex.step.kellas@gmail.com"}],"homepage":"https://github.com/alexstep/doc-extract#readme","bugs":{"url":"https://github.com/alexstep/doc-extract/issues"},"dist":{"shasum":"0df927392afc7c87065a9534dc3d2f1a60107cbc","tarball":"https://registry.npmjs.org/@alexstep/doc-extract/-/doc-extract-0.1.2.tgz","fileCount":7,"integrity":"sha512-Ix0m4EUqgIArEus0TzFgjYWxBzD5bfNRGiHJz/QVSU3wbMa7Ssgg/+loJAbGXbYOqrclPyjhcIN5VaXv57vj3g==","signatures":[{"sig":"MEQCIA9Bjt37YDGzX7KM9P4RYmxj+3LHWrWykmJGsYI3QURjAiAXWL/lOvur+CZvhgwic2aAkcfg/Z5M3Qxd+LdtssXE2w==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":52002},"main":"doc-extract.js","napi":{"targets":["x86_64-pc-windows-msvc","i686-pc-windows-msvc","aarch64-pc-windows-msvc","x86_64-apple-darwin","aarch64-apple-darwin","x86_64-unknown-linux-gnu","x86_64-unknown-linux-musl","aarch64-unknown-linux-gnu","aarch64-unknown-linux-musl","armv7-unknown-linux-gnueabihf","x86_64-unknown-freebsd","wasm32-wasip1-threads"],"binaryName":"doc-extract"},"type":"commonjs","types":"doc-extract.d.ts","engines":{"node":">= 18"},"gitHead":"0726c99bac47ab3cd64f1b90f4c26066cd7cc63a","scripts":{"test":"bun test","build":"napi build --platform --release","version":"napi version","artifacts":"napi artifacts","test:show":"bun test __tests__/show-extract.test.ts","test:wasi":"node scripts/test-wasi.cjs","cargo:test":"cargo test","build:debug":"napi build --platform","prepublishOnly":"napi create-npm-dirs && napi prepublish -t npm --no-gh-release"},"_npmUser":{"name":"kellas","email":"alex.step.kellas@gmail.com"},"repository":{"url":"git+https://github.com/alexstep/doc-extract.git","type":"git"},"_npmVersion":"11.13.0","description":"Native document text extraction for Node.js and Bun (PDF, Office, EPUB, and more)","directories":{},"_nodeVersion":"24.16.0","publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"devDependencies":{"fflate":"^0.8.2","@types/bun":"^1.3.3","@emnapi/core":"^1.10.0","@napi-rs/cli":"^3.7.0","@emnapi/runtime":"^1.10.0","@napi-rs/wasm-runtime":"^1.1.4"},"optionalDependencies":{"@alexstep/doc-extract-darwin-x64":"0.1.2","@alexstep/doc-extract-freebsd-x64":"0.1.2","@alexstep/doc-extract-wasm32-wasi":"0.1.2","@alexstep/doc-extract-darwin-arm64":"0.1.2","@alexstep/doc-extract-linux-x64-gnu":"0.1.2","@alexstep/doc-extract-linux-x64-musl":"0.1.2","@alexstep/doc-extract-win32-x64-msvc":"0.1.2","@alexstep/doc-extract-linux-arm64-gnu":"0.1.2","@alexstep/doc-extract-win32-ia32-msvc":"0.1.2","@alexstep/doc-extract-linux-arm64-musl":"0.1.2","@alexstep/doc-extract-win32-arm64-msvc":"0.1.2","@alexstep/doc-extract-linux-arm-gnueabihf":"0.1.2"},"_npmOperationalInternal":{"tmp":"tmp/doc-extract_0.1.2_1780116887207_0.0938890848213132","host":"s3://npm-registry-packages-npm-production"}},"0.1.4":{"name":"@alexstep/doc-extract","version":"0.1.4","keywords":["napi-rs","pdf","docx","text-extraction","rust","native"],"license":"MIT","_id":"@alexstep/doc-extract@0.1.4","maintainers":[{"name":"kellas","email":"alex.step.kellas@gmail.com"}],"homepage":"https://github.com/alexstep/doc-extract#readme","bugs":{"url":"https://github.com/alexstep/doc-extract/issues"},"dist":{"shasum":"543c59b8063fd2029daa4a09fa856cb3fcb93158","tarball":"https://registry.npmjs.org/@alexstep/doc-extract/-/doc-extract-0.1.4.tgz","fileCount":7,"integrity":"sha512-GORMgXlGfmgz6EJFSoZF866c8qZpvVk+00N2SkDwGR1jyV20ZG7lzkGzBB3E6Bke1UZ9RGlSduHGzWgY5bOjrA==","signatures":[{"sig":"MEUCIQD8uQZJDDsvRCFZ6gVATlFJfnAe5JwLHDpo03OXsh1e3AIgNJsd8O0ddErrSCp8k6T6pMTqZ31T/KQb8PSvL9XkHVo=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":53777},"main":"doc-extract.js","napi":{"targets":["x86_64-pc-windows-msvc","i686-pc-windows-msvc","aarch64-pc-windows-msvc","x86_64-apple-darwin","aarch64-apple-darwin","x86_64-unknown-linux-gnu","x86_64-unknown-linux-musl","aarch64-unknown-linux-gnu","aarch64-unknown-linux-musl","armv7-unknown-linux-gnueabihf","x86_64-unknown-freebsd","wasm32-wasip1-threads"],"binaryName":"doc-extract"},"type":"commonjs","types":"doc-extract.d.ts","engines":{"node":">= 18"},"gitHead":"69c5792a6c9105031a20886d26ecff3cd92f007e","scripts":{"test":"bun test","build":"napi build --platform --release","version":"napi version","artifacts":"napi artifacts","test:show":"bun test __tests__/show-extract.test.ts","test:wasi":"node scripts/test-wasi.cjs","cargo:test":"cargo test","build:debug":"napi build --platform","prepublishOnly":"napi create-npm-dirs && napi prepublish -t npm --no-gh-release"},"_npmUser":{"name":"kellas","email":"alex.step.kellas@gmail.com"},"repository":{"url":"git+https://github.com/alexstep/doc-extract.git","type":"git"},"_npmVersion":"11.13.0","description":"Native document text extraction for Node.js and Bun (PDF, Office, EPUB, and more)","directories":{},"_nodeVersion":"24.16.0","publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"devDependencies":{"fflate":"^0.8.2","@types/bun":"^1.3.3","@emnapi/core":"^1.10.0","@napi-rs/cli":"^3.7.0","@emnapi/runtime":"^1.10.0","@napi-rs/wasm-runtime":"^1.1.4"},"optionalDependencies":{"@alexstep/doc-extract-darwin-x64":"0.1.4","@alexstep/doc-extract-freebsd-x64":"0.1.4","@alexstep/doc-extract-wasm32-wasi":"0.1.4","@alexstep/doc-extract-darwin-arm64":"0.1.4","@alexstep/doc-extract-linux-x64-gnu":"0.1.4","@alexstep/doc-extract-linux-x64-musl":"0.1.4","@alexstep/doc-extract-win32-x64-msvc":"0.1.4","@alexstep/doc-extract-linux-arm64-gnu":"0.1.4","@alexstep/doc-extract-win32-ia32-msvc":"0.1.4","@alexstep/doc-extract-linux-arm64-musl":"0.1.4","@alexstep/doc-extract-win32-arm64-msvc":"0.1.4","@alexstep/doc-extract-linux-arm-gnueabihf":"0.1.4"},"_npmOperationalInternal":{"tmp":"tmp/doc-extract_0.1.4_1780141610863_0.31435291619079186","host":"s3://npm-registry-packages-npm-production"}},"0.1.5":{"name":"@alexstep/doc-extract","version":"0.1.5","description":"Native document text extraction for Node.js and Bun (PDF, Office, EPUB, and more)","license":"MIT","type":"commonjs","main":"doc-extract.js","types":"doc-extract.d.ts","repository":{"type":"git","url":"git+https://github.com/alexstep/doc-extract.git"},"keywords":["napi-rs","pdf","docx","text-extraction","rust","native"],"scripts":{"artifacts":"napi artifacts","build":"napi build --platform --release","build:debug":"napi build --platform","cargo:test":"cargo test","prepublishOnly":"napi create-npm-dirs && napi prepublish -t npm --no-gh-release","test":"bun test","demo":"node scripts/demo-server.mjs","test:wasi":"node scripts/test-wasi.cjs","test:show":"bun test __tests__/show-extract.test.ts","version":"napi version"},"devDependencies":{"@emnapi/core":"^1.10.0","@emnapi/runtime":"^1.10.0","@napi-rs/cli":"^3.7.0","@napi-rs/wasm-runtime":"^1.1.4","@types/bun":"^1.3.3","fflate":"^0.8.2"},"napi":{"binaryName":"doc-extract","targets":["x86_64-pc-windows-msvc","i686-pc-windows-msvc","aarch64-pc-windows-msvc","x86_64-apple-darwin","aarch64-apple-darwin","x86_64-unknown-linux-gnu","x86_64-unknown-linux-musl","aarch64-unknown-linux-gnu","aarch64-unknown-linux-musl","armv7-unknown-linux-gnueabihf","x86_64-unknown-freebsd","wasm32-wasip1-threads"]},"engines":{"node":">= 18"},"publishConfig":{"registry":"https://registry.npmjs.org/","access":"public"},"optionalDependencies":{"@alexstep/doc-extract-win32-x64-msvc":"0.1.5","@alexstep/doc-extract-win32-ia32-msvc":"0.1.5","@alexstep/doc-extract-win32-arm64-msvc":"0.1.5","@alexstep/doc-extract-darwin-x64":"0.1.5","@alexstep/doc-extract-darwin-arm64":"0.1.5","@alexstep/doc-extract-linux-x64-gnu":"0.1.5","@alexstep/doc-extract-linux-x64-musl":"0.1.5","@alexstep/doc-extract-linux-arm64-gnu":"0.1.5","@alexstep/doc-extract-linux-arm64-musl":"0.1.5","@alexstep/doc-extract-linux-arm-gnueabihf":"0.1.5","@alexstep/doc-extract-freebsd-x64":"0.1.5","@alexstep/doc-extract-wasm32-wasi":"0.1.5"},"gitHead":"999e08b8fea952e5de59aa2961c191cf4f9b8beb","_id":"@alexstep/doc-extract@0.1.5","bugs":{"url":"https://github.com/alexstep/doc-extract/issues"},"homepage":"https://github.com/alexstep/doc-extract#readme","_nodeVersion":"24.16.0","_npmVersion":"11.13.0","dist":{"integrity":"sha512-VzrWY/0PmT18DXdHnuS2OtEABkZbcBLNqjcNAipoCxBylMVUsWWJuqVoKiyeMwgD/yiB09yZjiGb060APyzvmw==","shasum":"f78c94d232f9ec8465466d926290516b5876f333","tarball":"https://registry.npmjs.org/@alexstep/doc-extract/-/doc-extract-0.1.5.tgz","fileCount":7,"unpackedSize":54259,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCID43u9f4MsiOpSxiSiCGuekO6ZmwjT9SB0lVKwvdBRhyAiBtakOmc2QEKHU/1dZePyBoydV166yC3ObWLImjMgZrWg=="}]},"_npmUser":{"name":"kellas","email":"alex.step.kellas@gmail.com"},"directories":{},"maintainers":[{"name":"kellas","email":"alex.step.kellas@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/doc-extract_0.1.5_1780150391169_0.23208292422721444"},"_hasShrinkwrap":false}},"time":{"created":"2026-05-30T04:54:47.030Z","modified":"2026-05-30T14:13:11.438Z","0.1.2":"2026-05-30T04:54:47.358Z","0.1.4":"2026-05-30T11:46:50.990Z","0.1.5":"2026-05-30T14:13:11.310Z"},"bugs":{"url":"https://github.com/alexstep/doc-extract/issues"},"license":"MIT","homepage":"https://github.com/alexstep/doc-extract#readme","keywords":["napi-rs","pdf","docx","text-extraction","rust","native"],"repository":{"type":"git","url":"git+https://github.com/alexstep/doc-extract.git"},"description":"Native document text extraction for Node.js and Bun (PDF, Office, EPUB, and more)","maintainers":[{"name":"kellas","email":"alex.step.kellas@gmail.com"}],"readme":"# doc-extract\n\n**Text extraction library for RAG pipelines, LLM apps, and AI agents** — turn uploaded documents into clean plain text before chunking, embedding, or indexing.\n\nNative addon for **Node.js** and **Bun** (Rust + N-API): fast, in-process, no subprocesses. PDF, Office, EPUB, calendars, contacts, Apple Wallet passes, and more.\n\n## Why doc-extract\n\n- **Fast** - Rust parsers run natively\n- **Small footprint** - ~6 MB native addon per platform\n- **Non-blocking** - `extractText()` returns a Promise; CPU work runs on a Rust thread pool\n- **Backpressure** - built-in concurrency limit (default 32) via `setMaxConcurrent`\n- **Flexible limits** - global, per-instance, or per-call file size in megabytes\n- **Large files** - file paths and ZIP-based formats read from disk; buffers above threshold auto-spill to temp\n- **No subprocesses** - PDF and Office handled in-process\n\n## Install\n\n```bash\nnpm install @alexstep/doc-extract\n# or\nbun add @alexstep/doc-extract\n```\n\nRequires **Node.js ≥ 18** or **Bun ≥ 1.3**.\n\n## Quick start\n\n```javascript\nimport docExtract from '@alexstep/doc-extract'\n\n// Global defaults\ndocExtract.setMaxConcurrent(4)\ndocExtract.setMaxFilesizeMB(42)\ndocExtract.setInMemoryThresholdMB(64)\n\nconst text = await docExtract.extractText('./report.pdf')\nconst fromUrl = await docExtract.extractText('https://example.com/file.pdf')\nconst pass = await docExtract.parsePkPass('./ticket.pkpass')\n\n// Isolated instance\nconst custom = new docExtract({ maxConcurrent: 4, maxFileSizeMB: 200 })\n\n// Buffer: auto-detect by magic bytes (%PDF, ZIP/docx, etc.)\nconst buffer = await Bun.file('./report.pdf').bytes()\nconst doc = await custom.extractText(Buffer.from(buffer))\n\n// Explicit format when magic is ambiguous (e.g. plain CSV bytes)\nconst csv = await custom.extractText(someBuffer, 'csv')\n```\n\n### Auto-detect\n\n| Input | Hint source |\n|-------|-------------|\n| File path | extension hint + magic bytes from file head |\n| URL | extension from pathname + magic bytes |\n| `Buffer` | **magic bytes only** (no filename) |\n\nWorks without a second argument for PDF (`%PDF`), Office ZIP (docx/xlsx/pptx), ICS/VCF, JSON, HTML, and similar.\n\nFor buffers, pass `format` explicitly when the content has no clear signature (e.g. legacy `.doc`, ambiguous plain text):\n\n```javascript\nawait docExtract.extractText(buffer, 'docx')\nawait docExtract.extractText(buffer, { format: 'pdf', debug: true })\nawait docExtract.extractText(buffer, { unknown: 'reject' }) // strict: no text heuristic\n```\n\n### Unknown content policy\n\nWhen format cannot be determined confidently:\n\n| `unknown` | Behavior |\n|-----------|----------|\n| `text-if-likely` (default) | Treat as `txt` only if bytes look like text (UTF-8/UTF-16/BOM, low control-byte ratio) |\n| `reject` | Return `\"\"` (unsupported) |\n| `text-lossy` | Try `txt` unless bytes are obviously binary |\n\n**Detection** uses explicit `format` first, then combines extension hints and magic bytes. Magic bytes override conflicting extensions when possible (e.g. `%PDF` vs a `.zip` path, ICS content vs a `.txt` name). Unknown bytes fall back according to `unknown` policy.\n\nPath-based text heuristics read up to **32 KB** of file head; magic/ZIP sniffing uses the first **4 KB**.\n\n## API\n\n| Method | Description |\n|--------|-------------|\n| `docExtract.extractText(input, format?)` | Extract text. `input` = `Buffer`, file path, or `http(s)://` URL. Optional `format` or `{ format, unknown, maxFileSizeMB, inMemoryThresholdMB, tempDir }`. |\n| `docExtract.parsePkPass(input, options?)` | Parse `.pkpass` → `{ pass, localization?, stripImage? }`. |\n| `docExtract.setMaxConcurrent(n)` | Global parallel parse limit (`n === 0` or negative = no-op). |\n| `docExtract.setMaxFilesizeMB(n)` | Global max input size in MB (default **42**). `0` = unlimited. |\n| `docExtract.setInMemoryThresholdMB(n)` | Above this size, paths/URLs skip JS heap; buffers spill to temp (default **64**). |\n| `docExtract.setMaxWorkingSetMB(n)` | Optional cap on total in-flight parse memory (`0` = disabled). |\n| `docExtract.setTempDir(dir)` | Directory for auto-spill temp files. |\n| `new docExtract({ maxConcurrent, maxFileSizeMB, inMemoryThresholdMB, tempDir, debug })` | Instance with its own limits and optional debug logging. |\n\n**Environment:** `DOCEXTRACT_MAX_CONCURRENT`, `DOCEXTRACT_MAX_FILESIZE_MB`, `DOCEXTRACT_MAX_BYTES`, `DOCEXTRACT_IN_MEMORY_THRESHOLD_MB`, `DOCEXTRACT_MAX_WORKING_SET_MB`, `DOCEXTRACT_TMPDIR`, `DOCEXTRACT_DEBUG`.\n\n### Error handling\n\n`extractText` resolves to a string — no `try/catch` needed for content issues:\n\n| Situation | Result |\n|-----------|--------|\n| Empty PDF (image scan), unsupported format, parse failure | `\"\"` |\n| File too large, missing file, URL fetch error | **throws** — use `.catch()` |\n\n```javascript\nconst text = await docExtract.extractText('./scan.pdf') // \"\" if image-only PDF\n\nawait docExtract.extractText('https://example.com/huge.pdf').catch((err) => {\n  // size limit, network, missing file\n})\n\nawait docExtract.extractText('./scan.pdf', { debug: true })\n// console.debug: PDF has no text layer (likely image-only scan)\n```\n\n`parsePkPass` still returns `null` for invalid passes (unchanged).\n\n### Large files\n\ndoc-extract is built for batch imports and server-side pipelines, not only small uploads.\n\n| Input | Behavior |\n|-------|----------|\n| **File path** | Rust opens the file directly; ZIP formats (docx, xlsx, pptx, epub, odt, pkpass) read entries from disk |\n| **URL** | Streamed to a temp file, then parsed via path pipeline |\n| **Buffer ≤ threshold** | Passed to native code in memory (default threshold 64 MB) |\n| **Buffer > threshold** | Auto-written to temp, parsed from disk, temp removed |\n\n**Policy vs memory:**\n\n- `maxFileSizeMB` — reject files larger than N (`0` = no limit)\n- `inMemoryThresholdMB` — when to avoid keeping full payload in the JS heap\n\n**Batch tip:** for many large files (e.g. 20 × 200 MB), lower `maxConcurrent` (2–4) and optionally set `setMaxWorkingSetMB` to cap total in-flight memory. Peak RAM for ZIP formats is roughly `concurrency × parser working set`, not `concurrency × file size`.\n\n### pkpass\n\n```javascript\nconst text = await docExtract.extractText('./ticket.pkpass') // formatted text for AI\nconst json = await docExtract.parsePkPass('./ticket.pkpass') // structured pass.json\n```\n\n## Supported formats\n\n| Group | Extensions |\n|-------|------------|\n| Office | `pdf`, `docx`, `docm`, `xlsx`, `xls`, `ods`, `pptx`, `pptm`, `odt`, `rtf` |\n| Books | `epub`, `fb2` |\n| Calendar / contacts | `ics`, `ifb`, `ical`, `vcf`, `vcard` |\n| Data | `json`, `jsonl`, `ndjson`, `csv`, `tsv` |\n| Web / text | `html`, `htm`, `xhtml`, `xml`, `txt`, `md`, `markdown`, `log` |\n| Wallet | `pkpass` (auto-detected from ZIP + `pass.json`) |\n\n## Limitations & alternatives\n\ndoc-extract targets **in-process text extraction**: text-layer PDF, Office Open XML, EPUB, calendars, and similar. No OCR, no subprocesses, no Docker sidecar.\n\n**Not supported (returns `\"\"` or needs another tool):**\n\n- Image-only / scanned PDF (no text layer) — see OCR below\n- Legacy **`.doc`** (binary Word), **`.msg`**, PostScript\n- Images with text: PNG, JPEG, TIFF (needs Tesseract OCR)\n- Audio: mp3, wav\n\nFor those cases, a HTTP sidecar such as [textract-docker](https://github.com/floleuerer/textract-docker) is a practical fallback. It wraps Python [textract](https://github.com/deanmalmgren/textract) with Tesseract OCR, `antiword` for `.doc`, and many other backends behind a simple REST API.\n\nTypical integration pattern:\n\n1. Try `docExtract.extractText()` first (fast, in-process).\n2. If result is `\"\"` or format is unsupported — call textract-docker (or your existing docparser service).\n\n## Performance\n\ndoc-extract runs **inside the Node/Bun process** — no HTTP, base64 encoding, or Docker hop per request. That makes it a better fit for high-throughput paths (batch imports) where latency and concurrency matter.\n\n[textract-docker](https://github.com/floleuerer/textract-docker) adds network and Python/subprocess overhead on each call, but covers **OCR and legacy formats** doc-extract deliberately skips.\n\n## Build from source\n\n```bash\ngit clone https://github.com/alexstep/doc-extract.git\ncd doc-extract\nbun install\nbun run build   # Rust ≥ 1.88\nbun test\n```\n\n## License\n\nMIT\n\n---\n\n## Stats\n\n| | |\n|---|---|\n| Native addon size | ~6 MB per platform |\n| Default max input | 42 MB (`setMaxFilesizeMB`, `0` = unlimited) |\n| In-memory threshold | 64 MB (`setInMemoryThresholdMB`) |\n| Zip entry cap | 64 MB per entry — exceeds limit throws (not silent truncate) |\n| Supported extensions | 30+ |\n| Default concurrency | 32 (`DOCEXTRACT_MAX_CONCURRENT`) |\n| Runtime | Node.js ≥ 18, Bun ≥ 1.3 |\n","readmeFilename":"README.md"}