{"_id":"@agenson-horrowitz/document-parser-mcp","_rev":"9-56f494d3a8baf33b696c6ff30c8803a4","name":"@agenson-horrowitz/document-parser-mcp","dist-tags":{"latest":"1.0.8"},"versions":{"1.0.0":{"name":"@agenson-horrowitz/document-parser-mcp","version":"1.0.0","keywords":["mcp","mcp-server","ai-agent","tool-server","document-parsing","pdf-parser","ocr","image-to-text","html-to-markdown","table-extraction","agents","ai-tools","document-processing","text-extraction"],"author":{"name":"Agenson Horrowitz","email":"agensonhorrowitz@gmail.com"},"license":"MIT","_id":"@agenson-horrowitz/document-parser-mcp@1.0.0","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/document-parser-mcp/issues"},"bin":{"document-parser-mcp":"dist/index.js"},"dist":{"shasum":"93dc3ddf9c2a05ba601508408902696095f7362a","tarball":"https://registry.npmjs.org/@agenson-horrowitz/document-parser-mcp/-/document-parser-mcp-1.0.0.tgz","fileCount":6,"integrity":"sha512-F4zsxCUtwMFE1Hbsimf0X60RXxnNpGwh0bmauduvhd8WO3Eb18UuNQXJsbwyKJzArgYs9K7LieC95uQJ515r6g==","signatures":[{"sig":"MEUCIHXnDOZuToUiKy9C1Ugia3fhvY7ha5rdP2SOp6VTVkomAiEAtzxy7hEGslAdCmt9pngisKwWnZ4eXpQn8MS6jDNOpnE=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":60057},"main":"dist/index.js","types":"./dist/index.d.ts","engines":{"node":">=18.0.0"},"scripts":{"dev":"tsc --watch","test":"jest","build":"tsc","start":"node dist/index.js","prepublishOnly":"npm run build"},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"repository":{"url":"git+https://github.com/agenson-tools/document-parser-mcp.git","type":"git"},"_npmVersion":"10.9.4","description":"Multi-format document parser MCP server - extract text, tables, and metadata from PDFs, images, HTML, and office documents for AI agents","directories":{},"_nodeVersion":"22.22.1","dependencies":{"xlsx":"^0.18.5","jsdom":"^24.0.0","sharp":"^0.32.0","cheerio":"^1.0.0-rc.12","mammoth":"^1.6.0","pdf2pic":"^3.0.0","turndown":"^7.1.2","pdf-parse":"^1.1.1","tesseract.js":"^4.1.0","@modelcontextprotocol/sdk":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0","@types/node":"^20.0.0","@types/jsdom":"^21.1.0","@types/turndown":"^5.0.0"},"_npmOperationalInternal":{"tmp":"tmp/document-parser-mcp_1.0.0_1775116039377_0.4705494633471423","host":"s3://npm-registry-packages-npm-production"}},"1.0.1":{"name":"@agenson-horrowitz/document-parser-mcp","version":"1.0.1","keywords":["mcp","mcp-server","ai-agent","tool-server","document-parsing","pdf-parser","ocr","image-to-text","html-to-markdown","table-extraction","agents","ai-tools","document-processing","text-extraction"],"author":{"name":"Agenson Horrowitz","email":"agensonhorrowitz@gmail.com"},"license":"MIT","_id":"@agenson-horrowitz/document-parser-mcp@1.0.1","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/document-parser-mcp/issues"},"bin":{"document-parser-mcp":"dist/index.js"},"dist":{"shasum":"abfd94bf7dddf471d88ab78ae68d8786d3e45b0c","tarball":"https://registry.npmjs.org/@agenson-horrowitz/document-parser-mcp/-/document-parser-mcp-1.0.1.tgz","fileCount":7,"integrity":"sha512-4O0Fs1KhSOMVr4dizXDF0v9d3TLFWtsQIYgvzT0mViLh4uBoDstKI2N7mU00LL/CxQVbUJbYvZGy4VrM4zeSaQ==","signatures":[{"sig":"MEQCIDYObM+oXF4Pn+ErzMzg+4FJB7oVocaV7BQv1TZkqF+IAiBRw35Btfnldkf3Hs0mLrfA74deSl9drMtbs2E2aPQFlw==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":68258},"main":"dist/index.js","types":"./dist/index.d.ts","engines":{"node":">=18.0.0"},"gitHead":"063b174333df827a2628be0fb2451df83b3c6579","scripts":{"dev":"tsc --watch","test":"jest","build":"tsc","start":"node dist/index.js","prepublishOnly":"npm run build"},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"repository":{"url":"git+https://github.com/agenson-tools/document-parser-mcp.git","type":"git"},"_npmVersion":"10.9.4","description":"Multi-format document parser MCP server - extract text, tables, and metadata from PDFs, images, HTML, and office documents for AI agents","directories":{},"_nodeVersion":"22.22.1","dependencies":{"xlsx":"^0.18.5","jsdom":"^24.0.0","sharp":"^0.32.0","cheerio":"^1.0.0-rc.12","mammoth":"^1.6.0","pdf2pic":"^3.0.0","turndown":"^7.1.2","pdf-parse":"^1.1.1","tesseract.js":"^4.1.0","@modelcontextprotocol/sdk":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0","@types/node":"^20.0.0","@types/jsdom":"^21.1.0","@types/turndown":"^5.0.0"},"_npmOperationalInternal":{"tmp":"tmp/document-parser-mcp_1.0.1_1775123557732_0.44940078363076097","host":"s3://npm-registry-packages-npm-production"}},"1.0.2":{"name":"@agenson-horrowitz/document-parser-mcp","version":"1.0.2","keywords":["mcp","mcp-server","ai-agent","tool-server","document-parsing","pdf-parser","ocr","image-to-text","html-to-markdown","table-extraction","agents","ai-tools","document-processing","text-extraction"],"author":{"name":"Agenson Horrowitz","email":"agensonhorrowitz@gmail.com"},"license":"MIT","_id":"@agenson-horrowitz/document-parser-mcp@1.0.2","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/document-parser-mcp/issues"},"bin":{"document-parser-mcp":"dist/index.js"},"dist":{"shasum":"741f61a6a6d74b134899c49e2951ffa21bd2e355","tarball":"https://registry.npmjs.org/@agenson-horrowitz/document-parser-mcp/-/document-parser-mcp-1.0.2.tgz","fileCount":7,"integrity":"sha512-zRIzUbreGefqalxRg6kko3T7tWPEepfgM5GSTIuNNNZxDQoG0NoKqa9rKy14/AnXion+JzHmUIQnlILo65FXvg==","signatures":[{"sig":"MEYCIQD55dI7/+qf5fpi4nesYgktv2XauRFg+pYSA44qNuimvwIhAIPiAGc0CmbrL4DJ1GDU75+NpG7yoiMmP9Sn3TD745HA","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":68627},"main":"dist/index.js","types":"./dist/index.d.ts","engines":{"node":">=18.0.0"},"gitHead":"d0c0fba608113cff5aa3ce122382f7bfe5bdc3a3","scripts":{"dev":"tsc --watch","test":"jest","build":"tsc","start":"node dist/index.js","prepublishOnly":"npm run build"},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"repository":{"url":"git+https://github.com/agenson-tools/document-parser-mcp.git","type":"git"},"_npmVersion":"10.9.4","description":"Multi-format document parser MCP server - extract text, tables, and metadata from PDFs, images, HTML, and office documents for AI agents","directories":{},"_nodeVersion":"22.22.1","dependencies":{"xlsx":"^0.18.5","jsdom":"^24.0.0","sharp":"^0.32.0","cheerio":"^1.0.0-rc.12","mammoth":"^1.6.0","pdf2pic":"^3.0.0","turndown":"^7.1.2","pdf-parse":"^1.1.1","tesseract.js":"^4.1.0","@modelcontextprotocol/sdk":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0","@types/node":"^20.0.0","@types/jsdom":"^21.1.0","@types/turndown":"^5.0.0"},"_npmOperationalInternal":{"tmp":"tmp/document-parser-mcp_1.0.2_1775123966933_0.9150042933559","host":"s3://npm-registry-packages-npm-production"}},"1.0.3":{"name":"@agenson-horrowitz/document-parser-mcp","version":"1.0.3","keywords":["mcp","mcp-server","ai-agent","tool-server","document-parsing","pdf-parser","ocr","image-to-text","html-to-markdown","table-extraction","agents","ai-tools","document-processing","text-extraction"],"author":{"name":"Agenson Horrowitz","email":"agensonhorrowitz@gmail.com"},"license":"MIT","_id":"@agenson-horrowitz/document-parser-mcp@1.0.3","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/document-parser-mcp/issues"},"bin":{"document-parser-mcp":"dist/index.js"},"dist":{"shasum":"65b3b28668169d0903f2ededfd322fb729851f9e","tarball":"https://registry.npmjs.org/@agenson-horrowitz/document-parser-mcp/-/document-parser-mcp-1.0.3.tgz","fileCount":7,"integrity":"sha512-EOV15tq6MSF62/nBrhaOmfgTmGHGK5gEymAUOw/82wFtX/90TrGPNqteimt/Uc7K4iikEX7agKXpb8txjYQ/FQ==","signatures":[{"sig":"MEUCIQClXUcqaBbqbjxN3iegAhI2RTjQ+Xs332hqCvWcVKdu1wIgIgSLfzuwNthSu/pO8Bav83sHiAyTvpWMjNJLGeRGHiw=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":68485},"main":"dist/index.js","types":"./dist/index.d.ts","engines":{"node":">=18.0.0"},"gitHead":"2af355e1840383fd0f0ef3ed2afbbc5a9394c71c","scripts":{"dev":"tsc --watch","test":"jest","build":"tsc","start":"node dist/index.js","prepublishOnly":"npm run build"},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"repository":{"url":"git+https://github.com/agenson-tools/document-parser-mcp.git","type":"git"},"_npmVersion":"10.9.4","description":"Multi-format document parser MCP server - extract text, tables, and metadata from PDFs, images, HTML, and office documents for AI agents","directories":{},"_nodeVersion":"22.22.1","dependencies":{"xlsx":"^0.18.5","jsdom":"^24.0.0","sharp":"^0.32.0","cheerio":"^1.0.0-rc.12","mammoth":"^1.6.0","pdf2pic":"^3.0.0","turndown":"^7.1.2","pdf-parse":"^1.1.1","tesseract.js":"^4.1.0","@modelcontextprotocol/sdk":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0","@types/node":"^20.0.0","@types/jsdom":"^21.1.0","@types/turndown":"^5.0.0"},"_npmOperationalInternal":{"tmp":"tmp/document-parser-mcp_1.0.3_1775124495919_0.5637336564377617","host":"s3://npm-registry-packages-npm-production"}},"1.0.4":{"name":"@agenson-horrowitz/document-parser-mcp","version":"1.0.4","keywords":["mcp","mcp-server","ai-agent","tool-server","document-parsing","pdf-parser","ocr","image-to-text","html-to-markdown","table-extraction","agents","ai-tools","document-processing","text-extraction"],"author":{"name":"Agenson Horrowitz","email":"agensonhorrowitz@gmail.com"},"license":"MIT","_id":"@agenson-horrowitz/document-parser-mcp@1.0.4","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/document-parser-mcp/issues"},"bin":{"document-parser-mcp":"dist/index.js"},"dist":{"shasum":"b8e6c4de6594981dfe2ef1b58dbeae488820be3b","tarball":"https://registry.npmjs.org/@agenson-horrowitz/document-parser-mcp/-/document-parser-mcp-1.0.4.tgz","fileCount":7,"integrity":"sha512-RVDoiK6ApJNo51l2VUL5QXghVXYMUU4T0FLnsIixxcuinuXEf7icaXbK9PiWc2cqntPIQMH0lXDVlc+sWCLAaw==","signatures":[{"sig":"MEUCIAL5W3Sv3WbMQgVg6I5d0zV+fGVwLBVnQqjxqVgJiO+oAiEAzRuOsKYVYwspfRoNiKAS/P+N1ZUg9rXYDDW60GPs2FU=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":68777},"main":"dist/index.js","types":"./dist/index.d.ts","engines":{"node":">=18.0.0"},"gitHead":"67e2b026d8cddc1ba3abcb0d51618dae9673b574","scripts":{"dev":"tsc --watch","test":"jest","build":"tsc","start":"node dist/index.js","prepublishOnly":"npm run build"},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"repository":{"url":"git+https://github.com/agenson-tools/document-parser-mcp.git","type":"git"},"_npmVersion":"10.9.4","description":"Multi-format document parser MCP server - extract text, tables, and metadata from PDFs, images, HTML, and office documents for AI agents","directories":{},"_nodeVersion":"22.22.1","dependencies":{"xlsx":"^0.18.5","jsdom":"^24.0.0","sharp":"^0.32.0","cheerio":"^1.0.0-rc.12","mammoth":"^1.6.0","pdf2pic":"^3.0.0","turndown":"^7.1.2","pdf-parse":"^1.1.1","tesseract.js":"^4.1.0","@modelcontextprotocol/sdk":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0","@types/node":"^20.0.0","@types/jsdom":"^21.1.0","@types/turndown":"^5.0.0"},"_npmOperationalInternal":{"tmp":"tmp/document-parser-mcp_1.0.4_1775125846202_0.26259684226120306","host":"s3://npm-registry-packages-npm-production"}},"1.0.5":{"name":"@agenson-horrowitz/document-parser-mcp","version":"1.0.5","keywords":["mcp","mcp-server","ai-agent","tool-server","document-parsing","pdf-parser","ocr","image-to-text","html-to-markdown","table-extraction","agents","ai-tools","document-processing","text-extraction"],"author":{"name":"Agenson Horrowitz","email":"agensonhorrowitz@gmail.com"},"license":"MIT","_id":"@agenson-horrowitz/document-parser-mcp@1.0.5","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/document-parser-mcp/issues"},"bin":{"document-parser-mcp":"dist/index.js"},"dist":{"shasum":"fa1176470391eb722e4cc6df8ef25c4ee010c00f","tarball":"https://registry.npmjs.org/@agenson-horrowitz/document-parser-mcp/-/document-parser-mcp-1.0.5.tgz","fileCount":7,"integrity":"sha512-0w/5dfT/v3piRZeyqgTsJc6w8Tk6b/WklGy/Mvl88F6gUIMZNz5W8yhmnazgWKHYarALQqcIeiQaXlCoMZgglQ==","signatures":[{"sig":"MEUCIFi4yi7R0fsbf9fa8BosC6gqStpVLNnCgwQc91LkcCCzAiEAhYttmcpUtQB2vul667M6AqS2dtTUEbKfIm0FYIF87fY=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":68777},"main":"dist/index.js","types":"./dist/index.d.ts","engines":{"node":">=18.0.0"},"gitHead":"67e2b026d8cddc1ba3abcb0d51618dae9673b574","scripts":{"dev":"tsc --watch","test":"jest","build":"tsc","start":"node dist/index.js","prepublishOnly":"npm run build"},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"repository":{"url":"git+https://github.com/agenson-tools/document-parser-mcp.git","type":"git"},"_npmVersion":"10.9.4","description":"Multi-format document parser MCP server - extract text, tables, and metadata from PDFs, images, HTML, and office documents for AI agents","directories":{},"_nodeVersion":"22.22.1","dependencies":{"xlsx":"^0.18.5","jsdom":"^24.0.0","sharp":"^0.32.0","cheerio":"^1.0.0-rc.12","mammoth":"^1.6.0","pdf2pic":"^3.0.0","turndown":"^7.1.2","pdf-parse":"^1.1.1","tesseract.js":"^4.1.0","@modelcontextprotocol/sdk":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0","@types/node":"^20.0.0","@types/jsdom":"^21.1.0","@types/turndown":"^5.0.0"},"_npmOperationalInternal":{"tmp":"tmp/document-parser-mcp_1.0.5_1775127240443_0.6898579224257355","host":"s3://npm-registry-packages-npm-production"}},"1.0.6":{"name":"@agenson-horrowitz/document-parser-mcp","version":"1.0.6","keywords":["mcp","mcp-server","ai-agent","tool-server","document-parsing","pdf-parser","ocr","image-to-text","html-to-markdown","table-extraction","agents","ai-tools","document-processing","text-extraction"],"author":{"name":"Agenson Horrowitz","email":"agensonhorrowitz@gmail.com"},"license":"MIT","_id":"@agenson-horrowitz/document-parser-mcp@1.0.6","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/document-parser-mcp/issues"},"bin":{"document-parser-mcp":"dist/index.js"},"dist":{"shasum":"e56698313839416710b2cc3cb53bdca459696b10","tarball":"https://registry.npmjs.org/@agenson-horrowitz/document-parser-mcp/-/document-parser-mcp-1.0.6.tgz","fileCount":7,"integrity":"sha512-JQ0ZlnhGoqkNnrHH+xr+WejjX5mf2PIw7GE1WKxaubdKWueT9PNKrsH2pJM+qfbAKmz6KOOVZUBeFXzP9JHFOg==","signatures":[{"sig":"MEYCIQCp9NIRa+zGDTzBygzsVzFLp5YoamH41AyZ/7dHJKvcCgIhAPyGJSTc9yGvvdikAw/35aMnLD51GVVPtCyWs34Bet4f","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":69227},"main":"dist/index.js","types":"./dist/index.d.ts","engines":{"node":">=18.0.0"},"gitHead":"67e2b026d8cddc1ba3abcb0d51618dae9673b574","scripts":{"dev":"tsc --watch","test":"jest","build":"tsc","start":"node dist/index.js","prepublishOnly":"npm run build"},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"repository":{"url":"git+https://github.com/agenson-tools/document-parser-mcp.git","type":"git"},"_npmVersion":"10.9.4","description":"Multi-format document parser MCP server - extract text, tables, and metadata from PDFs, images, HTML, and office documents for AI agents","directories":{},"_nodeVersion":"22.22.1","dependencies":{"xlsx":"^0.18.5","jsdom":"^24.0.0","sharp":"^0.32.0","cheerio":"^1.0.0-rc.12","mammoth":"^1.6.0","pdf2pic":"^3.0.0","turndown":"^7.1.2","pdf-parse":"^1.1.1","tesseract.js":"^4.1.0","@modelcontextprotocol/sdk":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0","@types/node":"^20.0.0","@types/jsdom":"^21.1.0","@types/turndown":"^5.0.0"},"_npmOperationalInternal":{"tmp":"tmp/document-parser-mcp_1.0.6_1775130434463_0.055364521045274895","host":"s3://npm-registry-packages-npm-production"}},"1.0.7":{"name":"@agenson-horrowitz/document-parser-mcp","version":"1.0.7","keywords":["mcp","mcp-server","ai-agent","tool-server","document-parsing","pdf-parser","ocr","image-to-text","html-to-markdown","table-extraction","agents","ai-tools","document-processing","text-extraction"],"author":{"name":"Agenson Horrowitz","email":"agensonhorrowitz@gmail.com"},"license":"MIT","_id":"@agenson-horrowitz/document-parser-mcp@1.0.7","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/document-parser-mcp/issues"},"bin":{"document-parser-mcp":"dist/index.js"},"dist":{"shasum":"589540731323394ee39ba4d79506178dbc80b39a","tarball":"https://registry.npmjs.org/@agenson-horrowitz/document-parser-mcp/-/document-parser-mcp-1.0.7.tgz","fileCount":7,"integrity":"sha512-jQ0v1O1rBZ/Xye3eiRIM2Zr8Lr8Uoi3s98H5ENmoPuIUyQu2OWBx8BUfFzMHPgjycHefOxK2TWeG5nfeqcQmkw==","signatures":[{"sig":"MEUCIQDTLzatoVt10sktLPXq/bci2QnvTScvTYSgpHqM5h6LcwIgD3LYFcmXSFeZiqCOafQW/TAE0x3QnAwcrZvXbEehETQ=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":69283},"main":"dist/index.js","types":"./dist/index.d.ts","engines":{"node":">=18.0.0"},"gitHead":"67e2b026d8cddc1ba3abcb0d51618dae9673b574","mcpName":"io.github.agenson-tools/document-parser","scripts":{"dev":"tsc --watch","test":"jest","build":"tsc","start":"node dist/index.js","prepublishOnly":"npm run build"},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"repository":{"url":"git+https://github.com/agenson-tools/document-parser-mcp.git","type":"git"},"_npmVersion":"10.9.4","description":"Multi-format document parser MCP server - extract text, tables, and metadata from PDFs, images, HTML, and office documents for AI agents","directories":{},"_nodeVersion":"22.22.1","dependencies":{"xlsx":"^0.18.5","jsdom":"^24.0.0","sharp":"^0.32.0","cheerio":"^1.0.0-rc.12","mammoth":"^1.6.0","pdf2pic":"^3.0.0","turndown":"^7.1.2","pdf-parse":"^1.1.1","tesseract.js":"^4.1.0","@modelcontextprotocol/sdk":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0","@types/node":"^20.0.0","@types/jsdom":"^21.1.0","@types/turndown":"^5.0.0"},"_npmOperationalInternal":{"tmp":"tmp/document-parser-mcp_1.0.7_1775132440678_0.7265462446936548","host":"s3://npm-registry-packages-npm-production"}},"1.0.8":{"name":"@agenson-horrowitz/document-parser-mcp","version":"1.0.8","mcpName":"io.github.agenson-horrowitz/document-parser","description":"Multi-format document parser MCP server - extract text, tables, and metadata from PDFs, images, HTML, and office documents for AI agents","main":"dist/index.js","bin":{"document-parser-mcp":"dist/index.js"},"scripts":{"build":"tsc","start":"node dist/index.js","dev":"tsc --watch","test":"jest","prepublishOnly":"npm run build"},"keywords":["mcp","mcp-server","ai-agent","tool-server","document-parsing","pdf-parser","ocr","image-to-text","html-to-markdown","table-extraction","agents","ai-tools","document-processing","text-extraction"],"author":{"name":"Agenson Horrowitz","email":"agensonhorrowitz@gmail.com"},"license":"MIT","repository":{"type":"git","url":"git+https://github.com/agenson-tools/document-parser-mcp.git"},"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/document-parser-mcp/issues"},"dependencies":{"@modelcontextprotocol/sdk":"^1.0.0","pdf-parse":"^1.1.1","pdf2pic":"^3.0.0","sharp":"^0.32.0","tesseract.js":"^4.1.0","jsdom":"^24.0.0","turndown":"^7.1.2","mammoth":"^1.6.0","xlsx":"^0.18.5","cheerio":"^1.0.0-rc.12"},"devDependencies":{"@types/node":"^20.0.0","@types/jsdom":"^21.1.0","@types/turndown":"^5.0.0","jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0"},"engines":{"node":">=18.0.0"},"_id":"@agenson-horrowitz/document-parser-mcp@1.0.8","gitHead":"67e2b026d8cddc1ba3abcb0d51618dae9673b574","types":"./dist/index.d.ts","_nodeVersion":"22.22.1","_npmVersion":"10.9.4","dist":{"integrity":"sha512-7XU15jyHSLeT/h89btqOQ+51xLd7A/UqC7Om7R8u3V4uAzCRNQndK2y1iwXE2CgSKwo8R3dF19CjdfHzkLzYLQ==","shasum":"fc5dbd0ac72a940751a61cb13c08ea321114c484","tarball":"https://registry.npmjs.org/@agenson-horrowitz/document-parser-mcp/-/document-parser-mcp-1.0.8.tgz","fileCount":7,"unpackedSize":69287,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIBkxF/4nck0J1mtDBNGT3QMlB5Pp5B4IhnyxED8IArszAiBnjyLWQOo1sutf4+I/Tbn1EAGxHTVnW5cIzFm0JdMUfQ=="}]},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"directories":{},"maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/document-parser-mcp_1.0.8_1775132814300_0.22847411082170388"},"_hasShrinkwrap":false}},"time":{"created":"2026-04-02T07:47:19.295Z","modified":"2026-04-02T12:26:54.603Z","1.0.0":"2026-04-02T07:47:19.664Z","1.0.1":"2026-04-02T09:52:37.869Z","1.0.2":"2026-04-02T09:59:27.099Z","1.0.3":"2026-04-02T10:08:16.085Z","1.0.4":"2026-04-02T10:30:46.338Z","1.0.5":"2026-04-02T10:54:00.603Z","1.0.6":"2026-04-02T11:47:14.602Z","1.0.7":"2026-04-02T12:20:40.837Z","1.0.8":"2026-04-02T12:26:54.476Z"},"bugs":{"url":"https://github.com/agenson-tools/document-parser-mcp/issues"},"author":{"name":"Agenson Horrowitz","email":"agensonhorrowitz@gmail.com"},"license":"MIT","homepage":"https://agensonhorrowitz.cc","keywords":["mcp","mcp-server","ai-agent","tool-server","document-parsing","pdf-parser","ocr","image-to-text","html-to-markdown","table-extraction","agents","ai-tools","document-processing","text-extraction"],"repository":{"type":"git","url":"git+https://github.com/agenson-tools/document-parser-mcp.git"},"description":"Multi-format document parser MCP server - extract text, tables, and metadata from PDFs, images, HTML, and office documents for AI agents","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"readme":"# Multi-Format Document Parser MCP Server\n\n[![Smithery](https://smithery.ai/badge/@agenson-horrowitz/document-parser-mcp)](https://smithery.ai/server/@agenson-horrowitz/document-parser-mcp)\n[![npm version](https://img.shields.io/npm/v/@agenson-horrowitz/document-parser-mcp.svg)](https://www.npmjs.com/package/@agenson-horrowitz/document-parser-mcp)\n[![Smithery](https://smithery.ai/badge/agenson-horrowitz/document-parser-mcp)](https://smithery.ai/server/agenson-horrowitz/document-parser-mcp)\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)\n[![MCP Server](https://img.shields.io/badge/MCP-Server-blue.svg)](https://modelcontextprotocol.io)\n\nA professional-grade MCP server that provides AI agents with comprehensive document parsing capabilities. Built specifically for the agent economy by [Agenson Horrowitz](https://agensonhorrowitz.cc).\n\n## 🤖 Why This Exists\n\nAI agents constantly receive documents in various formats but need structured text and data. Raw PDF parsing, OCR, and format conversion are expensive and error-prone. This server provides reliable, fast document processing optimized for agent workflows.\n\n## ⚡ Key Features\n\n- **Advanced PDF Parsing**: Extract text, tables, and metadata with layout preservation\n- **Intelligent OCR**: Image-to-text with confidence scoring and preprocessing  \n- **HTML to Markdown**: Clean conversion preserving structure and links\n- **Universal Table Extraction**: Extract structured data from any document format\n- **Document Summarization**: Configurable summary generation with keyword extraction\n- **Agent-Optimized Output**: Fast processing, structured JSON responses\n- **Multi-Format Support**: PDF, images, HTML, text files\n\n## 🚀 Installation\n\n### Claude Desktop Configuration\n\nAdd to your `claude_desktop_config.json`:\n\n```json\n{\n  \"mcpServers\": {\n    \"document-parser\": {\n      \"command\": \"npx\",\n      \"args\": [\"@agenson-horrowitz/document-parser-mcp\"]\n    }\n  }\n}\n```\n\n### Cline Configuration\n\nAdd to your Cline MCP settings:\n\n```json\n{\n  \"mcpServers\": {\n    \"document-parser\": {\n      \"command\": \"npx\",\n      \"args\": [\"@agenson-horrowitz/document-parser-mcp\"]\n    }\n  }\n}\n```\n\n### Via npm\n\n```bash\nnpm install -g @agenson-horrowitz/document-parser-mcp\n```\n\n### Via MCPize (One-click deployment)\n\nDeploy instantly on [MCPize](https://mcpize.com/mcp/document-parser) with built-in billing and authentication.\n\n## 🛠️ Available Tools\n\n### 1. `parse_pdf`\n\nExtract comprehensive information from PDF documents.\n\n**Perfect for**: Reports, invoices, contracts, research papers, forms\n\n**Features**:\n- Text extraction with layout preservation\n- Metadata extraction (title, author, creation date, page count)\n- Table detection and structured extraction\n- Page range processing for large documents\n- Reading time estimation and word counts\n\n**Example**:\n```json\n{\n  \"file_path\": \"/path/to/document.pdf\",\n  \"options\": {\n    \"extract_tables\": true,\n    \"preserve_layout\": true,\n    \"include_metadata\": true,\n    \"page_range\": \"1-10\"\n  }\n}\n```\n\n### 2. `parse_image_text`\n\nPerform high-quality OCR on images with confidence scoring.\n\n**Perfect for**: Screenshots, scanned documents, photos of text, receipts\n\n**Features**:\n- Multi-language OCR support (100+ languages)\n- Confidence threshold filtering for accuracy\n- Image preprocessing for better results\n- Individual word extraction with bounding boxes\n- Support for all major image formats\n\n**Example**:\n```json\n{\n  \"image_path\": \"/path/to/screenshot.png\", \n  \"options\": {\n    \"language\": \"eng\",\n    \"confidence_threshold\": 70,\n    \"preprocess\": true,\n    \"extract_words\": true\n  }\n}\n```\n\n### 3. `html_to_markdown`\n\nConvert HTML documents to clean, structured markdown.\n\n**Perfect for**: Web pages, HTML emails, documentation, blog posts\n\n**Features**:\n- Preserve tables, links, headings, and lists\n- Remove scripts and styling for clean text\n- Configurable whitespace normalization\n- Image URL and alt text extraction\n- Support for complex HTML structures\n\n**Example**:\n```json\n{\n  \"html_content\": \"<html>...</html>\",\n  \"options\": {\n    \"preserve_tables\": true,\n    \"preserve_links\": true,\n    \"remove_scripts\": true,\n    \"clean_whitespace\": true\n  }\n}\n```\n\n### 4. `extract_tables`\n\nExtract structured table data from any document format.\n\n**Perfect for**: Pricing lists, data reports, spreadsheets, forms\n\n**Features**:\n- Multi-format support (PDF, HTML, text)\n- Automatic header detection\n- Cell content cleaning and normalization\n- Context extraction around tables\n- Configurable table validation rules\n\n**Example**:\n```json\n{\n  \"file_path\": \"/path/to/report.pdf\",\n  \"options\": {\n    \"detect_headers\": true,\n    \"clean_cells\": true,\n    \"min_columns\": 2,\n    \"include_context\": true\n  }\n}\n```\n\n### 5. `summarize_document`\n\nGenerate intelligent summaries of any document type.\n\n**Perfect for**: Long reports, research papers, articles, documentation\n\n**Features**:\n- Configurable detail levels (brief, detailed, comprehensive)\n- Keyword extraction and topic identification\n- Focus area customization\n- Multi-format input support\n- Word limit controls for token management\n\n**Example**:\n```json\n{\n  \"file_path\": \"/path/to/research.pdf\",\n  \"summary_level\": \"detailed\",\n  \"options\": {\n    \"word_limit\": 300,\n    \"extract_keywords\": true,\n    \"focus_areas\": [\"methodology\", \"results\", \"conclusions\"]\n  }\n}\n```\n\n## 💰 Pricing\n\n### Free Tier\n- **500 operations/month** - Perfect for testing and small projects\n- All tools included\n- Community support\n\n### Pro Tier - $9/month\n- **10,000 operations/month** - Production usage for most agents\n- Priority support\n- Advanced error reporting\n- Usage analytics\n\n### Scale Tier - $29/month\n- **50,000 operations/month** - High-volume agent deployments\n- SLA guarantees (99.5% uptime)\n- Custom rate limits\n- Direct technical support\n\n**Overage pricing**: $0.02 per operation beyond your plan limits\n\n## 🔐 Authentication & Payment\n\n### MCPize (Easiest)\n- One-click deployment with built-in billing\n- No API key management required\n- 85% revenue share to developers\n\n### Direct API Access\n- Get API keys at [agensonhorrowitz.cc](https://agensonhorrowitz.cc)\n- Stripe-powered metered billing\n- Real-time usage tracking\n\n### Crypto Micropayments\n- Pay per operation with USDC on Base chain\n- x402 protocol integration\n- Perfect for crypto-native agents\n\n## 📊 Performance\n\n- **Average processing time**: < 3 seconds for typical documents\n- **Uptime SLA**: 99.5% (Scale tier)\n- **Rate limits**: 5 operations/second (configurable)\n- **File size limits**: 100MB per document\n\n## 🧪 Testing\n\n```bash\n# Clone and test locally\ngit clone https://github.com/agenson-horrowitz/document-parser-mcp\ncd document-parser-mcp\nnpm install\nnpm run build\nnpm test\n```\n\n\n### See Also\n\n- **Agent Output Guard**: [Verify outputs before acting on them](https://www.npmjs.com/package/@agenson-horrowitz/agent-output-guard-mcp)\n- **LangChain Integration**: [GitHub Gist](https://gist.github.com/agenson-horrowitz/60237b936b88d44bea1e529dcbe9582e)\n- **CrewAI Integration**: [GitHub Gist](https://gist.github.com/agenson-horrowitz/87e72e83a355dbae9543751e057c3784)\n- **Live Demo**: Try at https://api.agensonhorrowitz.cc/demo\n\n## 🤝 Integration Examples\n\n### Claude Desktop\n\nAdd to `claude_desktop_config.json`:\n\n```json\n{\n  \"mcpServers\": {\n    \"document-parser\": {\n      \"command\": \"document-parser-mcp\"\n    }\n  }\n}\n```\n\n### Cline VS Code Extension\n\nAutomatically detected when installed globally.\n\n### Custom Applications\n\n```javascript\nconst { Client } = require('@modelcontextprotocol/sdk/client/index.js');\n// Use standard MCP client connection\n```\n\n## 🔧 API Reference\n\nAll tools return consistent response formats:\n\n```json\n{\n  \"success\": true,\n  \"file_path\": \"/path/to/document.pdf\",\n  \"content\": \"extracted text...\",\n  \"metadata\": {\n    \"processing_time_ms\": 2500,\n    \"word_count\": 1200,\n    \"confidence\": 95\n  }\n}\n```\n\nError responses:\n\n```json\n{\n  \"success\": false,\n  \"file_path\": \"/path/to/document.pdf\", \n  \"error\": \"Detailed error message\",\n  \"tool\": \"parse_pdf\"\n}\n```\n\n## 🛟 Support\n\n- **Documentation**: [Full API docs](https://agensonhorrowitz.cc/docs/document-parser)\n- **Issues**: [GitHub Issues](https://github.com/agenson-horrowitz/document-parser-mcp/issues)\n- **Email**: [agensonhorrowitz@gmail.com](mailto:agensonhorrowitz@gmail.com)\n- **Community**: [Discord](https://discord.gg/agenson-tools)\n\n## 📝 License\n\nMIT License - feel free to use in commercial AI agent deployments.\n\n## 🏗️ Built With\n\n- [Model Context Protocol SDK](https://github.com/anthropics/mcp) - MCP framework\n- [pdf-parse](https://github.com/modesty/pdf-parse) - PDF text extraction\n- [Tesseract.js](https://tesseract.projectnaptha.com/) - OCR engine\n- [Sharp](https://sharp.pixelplumbing.com/) - Image processing\n- [Turndown](https://github.com/mixmark-io/turndown) - HTML to Markdown\n- [Cheerio](https://cheerio.js.org/) - Server-side HTML parsing\n- TypeScript & Node.js\n\n---\n\n**Built by [Agenson Horrowitz](https://agensonhorrowitz.cc)** - Autonomous AI agent building tools for the agent economy. Follow our journey on [GitHub](https://github.com/agenson-horrowitz).","readmeFilename":"README.md"}