{"_id":"@ambicuity/any-to-markdown","name":"@ambicuity/any-to-markdown","dist-tags":{"latest":"1.0.0"},"versions":{"1.0.0":{"name":"@ambicuity/any-to-markdown","version":"1.0.0","description":"Convert common documents, web pages, archives, media metadata, and text formats to Markdown.","type":"module","author":{"name":"Ritesh Rana","email":"contact@riteshrana.engineer"},"license":"MIT","main":"./dist/index.js","types":"./dist/index.d.ts","bin":{"any-to-markdown":"dist/cli.js"},"scripts":{"prepack":"npm run build","pretest":"npm run build","build":"tsc -p tsconfig.json","typecheck":"tsc --noEmit -p tsconfig.json","test":"vitest run","test:watch":"vitest","test:unit":"vitest run tests/unit","test:integration":"vitest run tests/integration","test:e2e":"vitest run tests/e2e","test:regression":"vitest run tests/regression tests/security","test:coverage":"vitest run --coverage","lint":"eslint \"src/**/*.ts\" \"tests/**/*.ts\"","format":"prettier --write \"**/*.{ts,md,json}\"","check:sanitize":"node scripts/check-sanitize.mjs"},"engines":{"node":">=20"},"dependencies":{"@azure/ai-form-recognizer":"^5.1.0","@modelcontextprotocol/sdk":"^1.13.3","@xmldom/xmldom":"^0.8.10","chardet":"^2.1.0","cheerio":"^1.0.0","commander":"^12.1.0","exceljs":"^4.4.0","fast-xml-parser":"^5.7.3","file-type":"^22.0.1","iconv-lite":"^0.6.3","jszip":"^3.10.1","mammoth":"^1.11.0","mime-types":"^2.1.35","pdf-parse":"^1.1.1","turndown":"^7.2.0","yauzl":"^3.2.0","youtube-transcript":"^1.2.1"},"devDependencies":{"@eslint/js":"^9.26.0","@types/mime-types":"^2.1.4","@types/node":"^22.15.17","@types/pdf-parse":"^1.1.4","@types/turndown":"^5.0.5","@types/yauzl":"^2.10.3","@vitest/coverage-v8":"^3.2.4","eslint":"^9.26.0","prettier":"^3.5.3","typescript":"^5.8.3","typescript-eslint":"^8.32.0","vitest":"^3.1.3"},"_id":"@ambicuity/any-to-markdown@1.0.0","gitHead":"ccf7aac72beee9ce5290ee804281d5ebe464b5fe","_nodeVersion":"22.22.2","_npmVersion":"10.9.7","dist":{"integrity":"sha512-5D8WpIn6zRePgEimfaBUNmnc3bFqMmIEdGDXTXS+d2PJg0yPaurxwtCl2sz9S+oaSYdvNJTNJJRm336Hf+ilBg==","shasum":"677626653654ad6b0f817a56e1429c8bfab61f7b","tarball":"https://registry.npmjs.org/@ambicuity/any-to-markdown/-/any-to-markdown-1.0.0.tgz","fileCount":121,"unpackedSize":204750,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQCNGJ2MWRG+vij9/mNhdBw7rZH/a0OFdzMOOeaRuUX8TgIhAPQRMJ+2X85VQ8I8IFvwdh7lc3RA4Y3TUgiI3ahFyCay"}]},"_npmUser":{"name":"ambicuity","email":"riteshrana36@gmail.com"},"directories":{},"maintainers":[{"name":"ambicuity","email":"riteshrana36@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/any-to-markdown_1.0.0_1778277919250_0.04690719390532916"},"_hasShrinkwrap":false}},"time":{"created":"2026-05-08T22:05:19.152Z","1.0.0":"2026-05-08T22:05:19.449Z","modified":"2026-05-08T22:05:19.660Z"},"maintainers":[{"name":"ambicuity","email":"riteshrana36@gmail.com"}],"description":"Convert common documents, web pages, archives, media metadata, and text formats to Markdown.","author":{"name":"Ritesh Rana","email":"contact@riteshrana.engineer"},"license":"MIT","readme":"# any-to-markdown\n\n`@ambicuity/any-to-markdown` is a TypeScript and Node.js package for converting common files, web content, archives, and selected media metadata into Markdown for indexing, retrieval, text analysis, and LLM workflows.\n\nThe package exposes both a command-line tool and a strict TypeScript API.\n\n## Install\n\n```bash\nnpm install @ambicuity/any-to-markdown\n```\n\n```bash\npnpm add @ambicuity/any-to-markdown\n```\n\n```bash\nyarn add @ambicuity/any-to-markdown\n```\n\n## Command Line\n\n```bash\nany-to-markdown path-to-file.pdf > document.md\n```\n\n```bash\nany-to-markdown path-to-file.docx -o document.md\n```\n\n```bash\ncat path-to-file.html | any-to-markdown --extension html\n```\n\nUseful options:\n\n- `--output <file>` writes Markdown to a file.\n- `--extension <extension>` provides a file-extension hint for stdin.\n- `--mime-type <mimeType>` provides a MIME-type hint.\n- `--charset <charset>` provides a text decoding hint.\n- `--keep-data-uris` preserves full data URIs in HTML-derived Markdown.\n- `--llm-caption-images` enables LLM captioning of embedded images in DOCX/PPTX/XLSX/EPUB (requires a programmatic `llmClient`).\n- `--llm-pdf-pages` sends the PDF buffer to the configured LLM when extracted text is sparse.\n- `--llm-audio-model <model>` sets the audio transcription model id (e.g. `whisper-1`).\n- `--mcp` runs the package as a Model Context Protocol stdio server.\n- `--version` prints the package version.\n\n## TypeScript API\n\n```ts\nimport { AnyToMarkdown } from \"@ambicuity/any-to-markdown\";\n\nconst converter = new AnyToMarkdown();\nconst result = await converter.convert(\"report.docx\");\n\nconsole.log(result.markdown);\n```\n\n## JavaScript API\n\n```js\nimport { AnyToMarkdown } from \"@ambicuity/any-to-markdown\";\n\nconst converter = new AnyToMarkdown();\nconst result = await converter.convert(\"report.xlsx\");\n\nconsole.log(result.textContent);\n```\n\n## Supported Inputs\n\nThe built-in converter set includes support for:\n\n- Plain text, Markdown, JSON, XML, YAML, and similar text formats\n- CSV tables\n- HTML and XHTML\n- DOCX\n- XLSX and XLS\n- PPTX\n- PDF text extraction\n- EPUB\n- ZIP archives\n- RTF\n- Image and audio metadata when ExifTool is available\n- Optional LLM image captioning through a caller-provided compatible client\n- Optional LLM audio transcription, embedded-image captioning, and PDF augmentation\n- Optional OCR helper converters through a caller-provided OCR service\n- Optional MCP server with `convert_uri`, `convert_local`, `convert_stream`, and `list_converters` tools\n\nSome formats depend on the fidelity of the underlying Node.js ecosystem libraries. Where exact output layout differs from another implementation, the goal is stable Markdown with equivalent content and behavior.\n\n## API Overview\n\n```ts\nimport {\n  AnyToMarkdown,\n  StreamInfo,\n  type DocumentConverter\n} from \"@ambicuity/any-to-markdown\";\n\nconst engine = new AnyToMarkdown();\n\nawait engine.convert(\"file.pdf\");\nawait engine.convertLocal(\"file.pdf\");\nawait engine.convertUri(\"data:text/plain,hello\");\nawait engine.convertStream(Buffer.from(\"hello\"), {\n  streamInfo: new StreamInfo({ extension: \".txt\" })\n});\n\nconst customConverter: DocumentConverter = {\n  accepts: (_input, info) => info.extension === \".custom\",\n  convert: () => ({ markdown: \"custom markdown\", textContent: \"custom markdown\", toString: () => \"custom markdown\" })\n};\n\nengine.registerConverter(customConverter);\n```\n\n## Security Considerations\n\n`any-to-markdown` performs I/O with the privileges of the current process. Validate untrusted file paths and URLs before converting them. Prefer the narrowest API that fits your workflow: use `convertLocal` for local files, `convertStream` for already-opened content, and `convertUri` only when URI fetching is intended.\n\n## Package Metadata\n\n- Package: `@ambicuity/any-to-markdown`\n- Version: `1.0.0`\n- Author: Ritesh Rana\n- Email: contact@riteshrana.engineer\n\n## License\n\nMIT\n","readmeFilename":"README.md","_rev":"1-b31557aa6000d83a622fe698d2b973dc"}