{"_id":"@akrym1582/office-to-md-mcp","name":"@akrym1582/office-to-md-mcp","dist-tags":{"latest":"0.0.2"},"versions":{"0.0.2":{"name":"@akrym1582/office-to-md-mcp","version":"0.0.2","description":"MCP server to convert Office/PDF documents into images and Markdown","main":"dist/server.js","bin":{"office-to-md-mcp":"dist/server.js"},"publishConfig":{"access":"public"},"engines":{"node":">=18"},"scripts":{"build":"tsc","start":"node dist/server.js","dev":"npm run build && node dist/server.js","test":"jest","lint":"eslint src --ext .ts","typecheck":"tsc --noEmit","prepublishOnly":"npm run build"},"repository":{"type":"git","url":"git+https://github.com/akrym1582/office-to-md-mcp.git"},"keywords":["mcp","model-context-protocol","office","word","excel","pdf","markdown","document-conversion","llm","ai"],"author":{"name":"akrym1582"},"license":"MIT","type":"commonjs","bugs":{"url":"https://github.com/akrym1582/office-to-md-mcp/issues"},"homepage":"https://github.com/akrym1582/office-to-md-mcp#readme","dependencies":{"@github/copilot-sdk":"^0.2.0","@modelcontextprotocol/sdk":"^1.0.0","exceljs":"^4.4.0","mammoth":"^1.7.0","zod":"^3.22.4"},"devDependencies":{"@types/jest":"^29.5.0","@types/node":"^20.0.0","@typescript-eslint/eslint-plugin":"^7.0.0","@typescript-eslint/parser":"^7.0.0","eslint":"^8.0.0","jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.3.0"},"_id":"@akrym1582/office-to-md-mcp@0.0.2","gitHead":"bfcca00b97346b4d37fed147c142ab813aa5af2f","types":"./dist/server.d.ts","_nodeVersion":"20.20.1","_npmVersion":"10.8.2","dist":{"integrity":"sha512-3Uj8YNWa6+tSUN3cZKBA14R+UCcBgbaKvANTk39PYp5bPRq3QtrioMKWi2hZhCUfTRwpt+YQm6WON7N/hLuFrw==","shasum":"f05859adbf8a5cb9c82b34ba62458f40aa85c04a","tarball":"https://registry.npmjs.org/@akrym1582/office-to-md-mcp/-/office-to-md-mcp-0.0.2.tgz","fileCount":79,"unpackedSize":114025,"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@akrym1582%2foffice-to-md-mcp@0.0.2","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQC3EI0ahceQTIVCzknH5tois3t10hbp1P1xXyyoizdQRgIgJa7cLyq7az0VmLY5GGxZxbET3/SjTzznkXefTK7nppU="}]},"_npmUser":{"name":"akrym1582","email":"aki.ymnk+n1582@gmail.com"},"directories":{},"maintainers":[{"name":"akrym1582","email":"aki.ymnk+n1582@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/office-to-md-mcp_0.0.2_1774907406098_0.8391374707887818"},"_hasShrinkwrap":false}},"time":{"created":"2026-03-30T21:50:05.996Z","0.0.2":"2026-03-30T21:50:06.250Z","modified":"2026-03-30T21:50:06.634Z"},"maintainers":[{"name":"akrym1582","email":"aki.ymnk+n1582@gmail.com"}],"description":"MCP server to convert Office/PDF documents into images and Markdown","homepage":"https://github.com/akrym1582/office-to-md-mcp#readme","keywords":["mcp","model-context-protocol","office","word","excel","pdf","markdown","document-conversion","llm","ai"],"repository":{"type":"git","url":"git+https://github.com/akrym1582/office-to-md-mcp.git"},"author":{"name":"akrym1582"},"bugs":{"url":"https://github.com/akrym1582/office-to-md-mcp/issues"},"license":"MIT","readme":"# office-to-md-mcp\n\nA TypeScript **Model Context Protocol (MCP) server** that converts Excel, Word, and PDF documents into PNG page images, structured text, and Markdown — optimised for LLM consumption.\n\n---\n\n## Features\n\n| Tool | Input | Output |\n|---|---|---|\n| `convert_excel_to_images` | `.xlsx` / `.xls` | PNG images per page |\n| `convert_word_to_images` | `.docx` / `.doc` | PNG images per page |\n| `convert_pdf_to_images` | `.pdf` | PNG images per page |\n| `extract_excel_text` | `.xlsx` / `.xls` | Markdown (via image-based conversion) |\n| `extract_word_text` | `.docx` | Plain text or Markdown |\n| `get_capabilities` | — | Runtime dependency status |\n\n### `extract_excel_text` Conversion Pipeline\n\n`extract_excel_text` converts Excel files to Markdown through the following image-based pipeline:\n\n```\nExcel (.xlsx/.xls)\n  → Adjust print area and convert to PDF (Python UNO / LibreOffice)\n    → Render PDF pages as PNG images (pdftoppm / ImageMagick)\n      → Convert images to Markdown (GitHub Copilot SDK — gpt-5.4-mini)\n```\n\nThis approach preserves not only cell data but also shapes, embedded images, and complex layouts with high fidelity.\n\n> **⚠️ GitHub Copilot Premium Requests**\n>\n> `extract_excel_text` uses GitHub Copilot SDK's **gpt-5.4-mini** model for image-to-Markdown conversion.\n> Each tool invocation consumes **GitHub Copilot Premium Requests**.\n> The number of requests increases with the number of pages in the workbook.\n\n---\n\n## Prerequisites\n\n| Dependency | Purpose | Required |\n|---|---|---|\n| [Node.js](https://nodejs.org/) ≥ 18 | Runtime | ✅ |\n| [LibreOffice](https://www.libreoffice.org/) (`soffice`) | Excel/Word → PDF | For image conversion |\n| [poppler-utils](https://poppler.freedesktop.org/) (`pdftoppm`) | PDF → PNG | For image conversion |\n| Python 3 | Excel UNO helper | For best Excel rendering |\n| `GITHUB_TOKEN` env var | Copilot SDK auth | Required for `extract_excel_text` |\n\n### Install system dependencies (Ubuntu/Debian)\n\n```bash\nsudo apt-get install -y libreoffice poppler-utils python3\n```\n\n### Install system dependencies (macOS)\n\n```bash\nbrew install libreoffice poppler python3\n```\n\n---\n\n## Installation\n\n```bash\nnpm install\nnpm run build\n```\n\n---\n\n## Running the server\n\n```bash\nnpm start\n```\n\nThe server communicates over **stdio** using the MCP protocol.\n\n### Environment variables\n\n| Variable | Description |\n|---|---|\n| `GITHUB_TOKEN` | GitHub personal access token for Copilot SDK Markdown conversion |\n| `COPILOT_MODEL` | Copilot model to use for image-to-Markdown conversion (default: `gpt-5.4-mini`) |\n| `LOG_LEVEL` | Log verbosity: `debug` \\| `info` (default) \\| `warn` \\| `error` |\n\n---\n\n## MCP Tool Reference\n\n### `convert_excel_to_images`\n\nConverts an Excel workbook to PNG page images.  \nUses the Python UNO helper (`python/excel_to_pdf_uno.py`) for accurate print-area handling when Python is available; falls back to LibreOffice CLI otherwise.\n\n```json\n{\n  \"filePath\": \"/path/to/file.xlsx\",\n  \"outputDir\": \"/tmp/output\",\n  \"dpi\": 150,\n  \"sheetNames\": [\"Sheet1\"],\n  \"keepPdf\": false\n}\n```\n\nResponse:\n```json\n{\n  \"sourceType\": \"excel\",\n  \"images\": [\"/tmp/output/page-1.png\"],\n  \"pageCount\": 1,\n  \"renderStrategy\": \"libreoffice-uno-print-area\"\n}\n```\n\n---\n\n### `convert_word_to_images`\n\nConverts a Word document to PNG page images via LibreOffice.\n\n```json\n{\n  \"filePath\": \"/path/to/file.docx\",\n  \"outputDir\": \"/tmp/output\",\n  \"dpi\": 150,\n  \"keepPdf\": false\n}\n```\n\n---\n\n### `convert_pdf_to_images`\n\nRenders each PDF page as a PNG image.\n\n```json\n{\n  \"filePath\": \"/path/to/file.pdf\",\n  \"outputDir\": \"/tmp/output\",\n  \"dpi\": 150\n}\n```\n\n---\n\n### `extract_excel_text`\n\nConverts an Excel workbook to Markdown via an image-based pipeline (Excel → print-area adjustment → PDF → images → Markdown). Handles shapes, embedded images, and complex layouts. Requires `GITHUB_TOKEN`.\n\n```json\n{\n  \"filePath\": \"/path/to/file.xlsx\",\n  \"dpi\": 150,\n  \"sheetNames\": [\"Sheet1\"]\n}\n```\n\nResponse:\n```json\n{\n  \"sourceType\": \"excel\",\n  \"textFormat\": \"markdown\",\n  \"content\": \"## Page 1\\n\\n| Name | Age |\\n| --- | --- |\\n| Alice | 30 |\",\n  \"images\": [\"/tmp/excel-images-xxx/page-1.png\"],\n  \"pageCount\": 1\n}\n```\n\n> Image-to-Markdown conversion uses GitHub Copilot SDK (default model: `gpt-5.4-mini`) and consumes Premium Requests.\n\n---\n\n### `extract_word_text`\n\nExtracts text from a `.docx` file using [mammoth](https://github.com/mwilliamson/mammoth.js).\n\n```json\n{\n  \"filePath\": \"/path/to/file.docx\",\n  \"format\": \"markdown\"\n}\n```\n\n---\n\n### `get_capabilities`\n\nReturns the runtime status of all system dependencies.\n\n```json\n{}\n```\n\nExample response:\n```json\n{\n  \"libreOffice\": true,\n  \"libreOfficePath\": \"/usr/bin/soffice\",\n  \"python\": true,\n  \"pythonPath\": \"/usr/bin/python3\",\n  \"pythonVersion\": \"Python 3.12.3\",\n  \"unoHelper\": true,\n  \"pdfRenderer\": true,\n  \"pdfRendererTool\": \"pdftoppm\",\n  \"githubToken\": false\n}\n```\n\n---\n\n## Project Structure\n\n```\n.\n├── src/\n│   ├── server.ts                      # MCP server entry point\n│   ├── tools/                         # MCP tool implementations\n│   │   ├── convertExcelToImages.ts\n│   │   ├── convertWordToImages.ts\n│   │   ├── convertPdfToImages.ts\n│   │   └── extractExcelText.ts\n│   ├── services/                      # Business logic / external integrations\n│   │   ├── capabilityDetector.ts\n│   │   ├── copilotCli.ts\n│   │   ├── excelExtractor.ts\n│   │   ├── fileType.ts\n│   │   ├── libreOfficeCli.ts\n│   │   ├── officePythonBridge.ts\n│   │   ├── pdfRenderer.ts\n│   │   ├── tempFiles.ts\n│   │   └── wordExtractor.ts\n│   ├── types/\n│   │   ├── errors.ts                  # AppError + ErrorCode enum\n│   │   └── toolSchemas.ts             # Zod schemas for all tools\n│   └── utils/\n│       ├── exec.ts                    # Subprocess wrapper with timeouts\n│       ├── fs.ts                      # File system helpers\n│       └── logger.ts                  # Stderr logger\n├── python/\n│   └── excel_to_pdf_uno.py            # LibreOffice UNO helper for Excel→PDF\n├── test/\n│   ├── fixtures/                      # Sample .xlsx, .docx, .pdf files\n│   └── unit/                          # Unit tests\n├── package.json\n├── tsconfig.json\n└── jest.config.js\n```\n\n---\n\n## Development\n\n```bash\n# Type-check without emitting\nnpm run typecheck\n\n# Build\nnpm run build\n\n# Run tests\nnpm test\n\n# Lint\nnpm run lint\n```\n\n---\n\n## Error Codes\n\n| Code | Meaning |\n|---|---|\n| `FILE_NOT_FOUND` | Input file does not exist |\n| `UNSUPPORTED_FORMAT` | File extension not supported |\n| `LIBREOFFICE_NOT_FOUND` | `soffice` not on PATH |\n| `PYTHON_NOT_FOUND` | Python interpreter not found |\n| `LIBREOFFICE_UNO_CONVERSION_FAILED` | Python UNO helper failed |\n| `LIBREOFFICE_CLI_CONVERSION_FAILED` | LibreOffice CLI conversion failed |\n| `PDF_RENDER_TOOL_NOT_FOUND` | `pdftoppm`/`convert` not on PATH |\n| `PDF_RENDER_FAILED` | PDF rendering failed |\n| `EXCEL_TEXT_EXTRACTION_FAILED` | ExcelJS read failure |\n| `WORD_TEXT_EXTRACTION_FAILED` | mammoth extraction failure |\n| `GITHUB_TOKEN_MISSING` | `GITHUB_TOKEN` env var not set |\n| `COPILOT_MARKDOWN_FAILED` | Copilot CLI returned an error |\n| `INVALID_TOOL_INPUT` | Zod schema validation failed |\n\n---\n\n## Troubleshooting\n\n**LibreOffice not found**  \nInstall LibreOffice and ensure `soffice` is on your `PATH`.\n\n**pdftoppm not found**  \nInstall `poppler-utils` (`apt-get install poppler-utils` or `brew install poppler`).\n\n**Copilot SDK unavailable**  \nSet `GITHUB_TOKEN` in your environment. The model used can be customised via the `COPILOT_MODEL` environment variable (default: `gpt-5.4-mini`).\n\n**Excel conversion uses LibreOffice CLI instead of UNO**  \nPython 3 must be on `PATH` and `python/excel_to_pdf_uno.py` must exist alongside the server. Run `get_capabilities` to confirm.\n\n---\n\n## License\n\nMIT","readmeFilename":"README.md","_rev":"1-6f6ea8706f13c34c606583efc2623045"}