{"_id":"@damorris25/content-extractor","_rev":"2-1b0a1bb9bdfebd7fa74507da949b1ad8","name":"@damorris25/content-extractor","dist-tags":{"latest":"0.1.1"},"versions":{"0.1.0":{"name":"@damorris25/content-extractor","version":"0.1.0","keywords":["content-extraction","text-extraction","metadata","docx","xlsx","pptx","pdf","odf","heic","avif","exif","dxf","tika"],"author":{"name":"Dana Morris","email":"damorris25@gmail.com"},"license":"MIT","_id":"@damorris25/content-extractor@0.1.0","maintainers":[{"name":"damorris25","email":"damorris25@gmail.com"}],"homepage":"https://github.com/damorris25/content-extractor#readme","bugs":{"url":"https://github.com/damorris25/content-extractor/issues"},"dist":{"shasum":"13ace7879c02b583b54cff3eebc1cd86b96d5472","tarball":"https://registry.npmjs.org/@damorris25/content-extractor/-/content-extractor-0.1.0.tgz","fileCount":49,"integrity":"sha512-LY8Ns14AP2dDeZl9xFpysNJIaJyZcycIo6ZqCrAGGJzNt4k9a6u0A0n91kwncrewmg804SkP8Pr+L4NNRg4a9A==","signatures":[{"sig":"MEYCIQCuIGZRgqRf/u8C+71k8ftxqRQ1Zsgxb8anZDBeojMSNwIhAJ5FNjg4WtNBdAF5Si62wyxE5MbczGzypyRRck34nKbx","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":1888963},"main":"dist/index.cjs","type":"module","types":"dist/index.d.ts","module":"dist/index.js","engines":{"node":">=20"},"exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js","require":"./dist/index.cjs"}},"gitHead":"3a8985d341d3cdb0bcc7003b9b67fc7ff94dd32d","scripts":{"test":"vitest run","build":"vite build && tsc --emitDeclarationOnly","clean":"rm -rf dist","typecheck":"tsc --noEmit","test:watch":"vitest","prepublishOnly":"npm run clean && npm run build && npm run test && npm run typecheck"},"_npmUser":{"name":"damorris25","email":"damorris25@gmail.com"},"repository":{"url":"git+https://github.com/damorris25/content-extractor.git","type":"git","directory":"ts"},"_npmVersion":"11.13.0","description":"Extract text content and metadata from unstructured file formats (Office, ODF, PDF, images, audio, CAD)","directories":{},"_nodeVersion":"23.9.0","dependencies":{"fflate":"^0.8.0"},"_hasShrinkwrap":false,"devDependencies":{"vite":"^6.0.0","vitest":"^3.0.0","pdfjs-dist":"^5.5.207","typescript":"^5.7.0","tesseract.js":"^7.0.0","@xmldom/xmldom":"^0.9.0","vite-plugin-dts":"^4.0.0"},"peerDependencies":{"pdfjs-dist":"^4.0.0 || ^5.0.0","tesseract.js":"^5.0.0"},"peerDependenciesMeta":{"pdfjs-dist":{"optional":true},"tesseract.js":{"optional":true}},"_npmOperationalInternal":{"tmp":"tmp/content-extractor_0.1.0_1785848176915_0.5326393182938951","host":"s3://npm-registry-packages-npm-production"}},"0.1.1":{"name":"@damorris25/content-extractor","version":"0.1.1","description":"Extract text content and metadata from unstructured file formats (Office, ODF, PDF, images, audio, CAD)","type":"module","main":"dist/index.cjs","module":"dist/index.js","types":"dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js","require":"./dist/index.cjs"}},"repository":{"type":"git","url":"git+https://github.com/damorris25/content-extractor.git","directory":"ts"},"homepage":"https://github.com/damorris25/content-extractor#readme","bugs":{"url":"https://github.com/damorris25/content-extractor/issues"},"keywords":["content-extraction","text-extraction","metadata","docx","xlsx","pptx","pdf","odf","heic","avif","exif","dxf","tika"],"scripts":{"build":"vite build && tsc --emitDeclarationOnly","test":"vitest run","test:watch":"vitest","typecheck":"tsc --noEmit","clean":"rm -rf dist","prepublishOnly":"npm run clean && npm run build && npm run test && npm run typecheck"},"dependencies":{"fflate":"^0.8.0"},"peerDependencies":{"pdfjs-dist":"^4.0.0 || ^5.0.0","tesseract.js":"^5.0.0"},"peerDependenciesMeta":{"pdfjs-dist":{"optional":true},"tesseract.js":{"optional":true}},"devDependencies":{"@xmldom/xmldom":"^0.9.0","pdfjs-dist":"^5.5.207","tesseract.js":"^7.0.0","typescript":"^5.7.0","vite":"^6.0.0","vite-plugin-dts":"^4.0.0","vitest":"^3.0.0"},"license":"MIT","author":{"name":"Dana Morris","email":"damorris25@gmail.com"},"engines":{"node":">=20"},"gitHead":"293a450921d255a2ca6bd6a9a8598e8189d4020f","_id":"@damorris25/content-extractor@0.1.1","_nodeVersion":"24.18.0","_npmVersion":"12.0.2","dist":{"integrity":"sha512-Ih2vqPt6wuoii+hOmgikdQ67E7b13EHMRACjw3OShda/ejETRhvXwg24v1ONaizDrYLwxtA4xbN5W/VgR0JcDg==","shasum":"464bc70a0bf176e6d7429d120467e351d42eee85","tarball":"https://registry.npmjs.org/@damorris25/content-extractor/-/content-extractor-0.1.1.tgz","fileCount":49,"unpackedSize":1888963,"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@damorris25%2fcontent-extractor@0.1.1","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIFMpQD3o330AwUW/cbxonyoZwdFc/2oIaE1IzvojuextAiEAmup8aN4hV6ly8VhdnZuibG/icZmOp6jK8Jh75GkOs38="}]},"_npmUser":{"name":"GitHub Actions","email":"npm-oidc-no-reply@github.com","trustedPublisher":{"id":"github","oidcConfigId":"oidc:cec85492-d3dc-418d-afc3-a31849b45ea8"}},"directories":{},"maintainers":[{"name":"damorris25","email":"damorris25@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/content-extractor_0.1.1_1785848582164_0.9948503298817901"},"_hasShrinkwrap":false}},"time":{"created":"2026-08-04T12:56:16.326Z","modified":"2026-08-04T13:03:02.626Z","0.1.0":"2026-08-04T12:56:17.073Z","0.1.1":"2026-08-04T13:03:02.318Z"},"bugs":{"url":"https://github.com/damorris25/content-extractor/issues"},"author":{"name":"Dana Morris","email":"damorris25@gmail.com"},"license":"MIT","homepage":"https://github.com/damorris25/content-extractor#readme","keywords":["content-extraction","text-extraction","metadata","docx","xlsx","pptx","pdf","odf","heic","avif","exif","dxf","tika"],"repository":{"type":"git","url":"git+https://github.com/damorris25/content-extractor.git","directory":"ts"},"description":"Extract text content and metadata from unstructured file formats (Office, ODF, PDF, images, audio, CAD)","maintainers":[{"name":"damorris25","email":"damorris25@gmail.com"}],"readme":"# @damorris25/content-extractor\n\n[![CI](https://github.com/damorris25/content-extractor/actions/workflows/ci.yml/badge.svg)](https://github.com/damorris25/content-extractor/actions/workflows/ci.yml)\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://github.com/damorris25/content-extractor/blob/master/LICENSE)\n\nExtract text content and metadata from unstructured file formats — Office\ndocuments, OpenDocument, PDF, images, audio, email, CAD, and more. Inspired by\n[Apache Tika](https://tika.apache.org/), built browser-first with a single\nruntime dependency ([`fflate`](https://www.npmjs.com/package/fflate), MIT).\n\nThis is the TypeScript implementation of\n[content-extractor](https://github.com/damorris25/content-extractor). A Go\nimplementation lives in the same repository; both produce identical output,\nenforced by golden-file parity tests.\n\n## Install\n\n```bash\nnpm install @damorris25/content-extractor\n```\n\nOptional peer dependencies:\n\n```bash\nnpm install pdfjs-dist    # PDF text extraction\nnpm install tesseract.js  # OCR for images\n```\n\n## Quick Start\n\n```ts\nimport { extract } from '@damorris25/content-extractor';\n\nconst bytes = new Uint8Array(await file.arrayBuffer()); // or fs.readFileSync\nconst result = await extract(bytes, { filename: 'report.docx' });\n\nconsole.log(result.text);        // extracted text content\nconsole.log(result.contentType); // detected MIME type\nconsole.log(result.metadata);    // Record<string, string[]> — title, creator, dates, ...\n```\n\nWorks in browsers and Node.js. Input is always a `Uint8Array`; extraction is\nasync and accepts an optional `AbortSignal`:\n\n```ts\nconst controller = new AbortController();\nconst result = await extract(bytes, {\n  filename: 'big.xlsx',\n  signal: controller.signal,\n});\n```\n\n## Supported Formats\n\n| Family | Formats | Extracted |\n|--------|---------|-----------|\n| Microsoft Office | DOCX, XLSX, PPTX | Text + document metadata |\n| OpenDocument | ODT, ODS, ODP | Text + document metadata |\n| PDF | PDF (text layer) | Text (via `pdfjs-dist` peer) |\n| Web/Markup | XML, HTML, SVG | Visible text, titles |\n| Structured data | JSON, CSV, TSV, YAML, Markdown | Pass-through + detection |\n| Rich text | RTF | Text with Unicode support |\n| Email | EML (RFC 2822) | Body text + headers |\n| Images | JPEG, PNG, TIFF/GeoTIFF, BMP, GIF, WebP, HEIC, HEIF, AVIF | EXIF/metadata, optional OCR |\n| Audio | MP3, OGG, FLAC, WAV | ID3v2 / Vorbis / RIFF metadata |\n| CAD | DXF | Text entities, layers, metadata |\n\n## Security\n\nBuilt for untrusted input: ZIP bomb protection (compression ratio + size\ncaps), integer overflow checks, NUL byte stripping, and bounded parsing for\nISOBMFF containers. Report vulnerabilities via\n[GitHub Security Advisories](https://github.com/damorris25/content-extractor/security/advisories/new).\n\n## Documentation\n\nFull documentation, architecture contracts, benchmarks, and the Go\nimplementation: [github.com/damorris25/content-extractor](https://github.com/damorris25/content-extractor)\n\n## License\n\n[MIT](https://github.com/damorris25/content-extractor/blob/master/LICENSE) © 2026 Dana Morris\n","readmeFilename":"README.md"}