{"_id":"pi-docparser","_rev":"8-c8a2cede7cddacead270dd29c8b8e4c6","name":"pi-docparser","dist-tags":{"latest":"4.0.0"},"versions":{"1.0.0":{"name":"pi-docparser","version":"1.0.0","keywords":["document-parse","documents","liteparse","ocr","pdf","pi","pi-package"],"license":"MIT","_id":"pi-docparser@1.0.0","maintainers":[{"name":"maxedapps","email":"maxedappsgmbh@gmail.com"}],"homepage":"https://maximilian-schwarzmueller.com","pi":{"image":"https://raw.githubusercontent.com/maxedapps/pi-docparser/main/assets/pi-docparser-preview.jpg","skills":["./skills"],"extensions":["./extensions/docparser/index.ts"]},"dist":{"shasum":"ae4f56f75a8f464bf2453a98c20b90ffd84cf1b7","tarball":"https://registry.npmjs.org/pi-docparser/-/pi-docparser-1.0.0.tgz","fileCount":17,"integrity":"sha512-XSttHyRMa4ItZAMNHCVlQGV4xMC0kOlPj/8bzt7Yg5oawlGU/tHirsFaE7hJsiWGS+cmYsWSw5tYBhrJV+GoGQ==","signatures":[{"sig":"MEUCIDD/4BPpXgDulGUR4vrZ+DvmJHvx8Yar6F6FXR94dX0PAiEA0hEVY0hU+P26GlwyS5+GC0Fhjk0NCZmc3ZO16cOWwEE=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":155962},"type":"module","engines":{"node":">=18.0.0"},"gitHead":"69f8191796b12abba36755b7cb8b517273fa791a","scripts":{"lint":"oxlint .","build":"echo 'nothing to build'","check":"oxfmt --check . && oxlint . && tsc -p tsconfig.json && node --input-type=module -e \"await import('./extensions/docparser/index.ts')\"","format":"oxfmt .","lint:fix":"oxlint --fix .","pack:dry":"npm pack --dry-run","typecheck":"tsc -p tsconfig.json","format:check":"oxfmt --check .","check:runtime":"node --input-type=module -e \"await import('./extensions/docparser/index.ts')\""},"_npmUser":{"name":"maxedapps","email":"maxedappsgmbh@gmail.com"},"_npmVersion":"11.6.2","description":"Pi package that adds a document_parse tool and companion skill for parsing PDFs, Office documents, spreadsheets, and images with LiteParse.","directories":{},"_nodeVersion":"24.11.1","dependencies":{"@llamaindex/liteparse":"1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"oxfmt":"^0.41.0","oxlint":"^1.56.0","typescript":"^5.9.3","@types/node":"^25.5.0","@mariozechner/pi-ai":"0.61.0","@mariozechner/pi-coding-agent":"0.61.0"},"peerDependencies":{"@mariozechner/pi-ai":"*","@mariozechner/pi-coding-agent":"*"},"_npmOperationalInternal":{"tmp":"tmp/pi-docparser_1.0.0_1774022689605_0.320729072665235","host":"s3://npm-registry-packages-npm-production"}},"1.0.1":{"name":"pi-docparser","version":"1.0.1","keywords":["document-parse","documents","liteparse","ocr","pdf","pi","pi-package"],"license":"MIT","_id":"pi-docparser@1.0.1","maintainers":[{"name":"maxedapps","email":"maxedappsgmbh@gmail.com"}],"homepage":"https://maximilian-schwarzmueller.com","bugs":{"url":"https://github.com/maxedapps/pi-docparser/issues"},"pi":{"image":"https://raw.githubusercontent.com/maxedapps/pi-docparser/main/assets/pi-docparser-preview.jpg","skills":["./skills"],"extensions":["./extensions/docparser/index.ts"]},"dist":{"shasum":"9154b2841fff093b2536a6e575ec72a4b5658c66","tarball":"https://registry.npmjs.org/pi-docparser/-/pi-docparser-1.0.1.tgz","fileCount":17,"integrity":"sha512-8h78PRwMdPDIOq1lyo7ers9vbyK18OsoXlRjmqoIinEcxfglu5qcuIhriM+WxJEHVVQe0eWeZuRk7jiA8PKROA==","signatures":[{"sig":"MEUCIERZDGdujvuIHXVWwMmUeMerLFID99yaCTSP4S0lfOxHAiEAj1R/jugUGLH92j9owcJV8cypfRCbftvX7FVhS5Eaqs4=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":156289},"type":"module","engines":{"node":">=18.0.0"},"gitHead":"ab1201c6ba6e892355ae8a6dadce209867f5003a","scripts":{"lint":"oxlint .","build":"echo 'nothing to build'","check":"oxfmt --check . && oxlint . && tsc -p tsconfig.json && node --input-type=module -e \"await import('./extensions/docparser/index.ts')\"","format":"oxfmt .","lint:fix":"oxlint --fix .","pack:dry":"npm pack --dry-run","typecheck":"tsc -p tsconfig.json","format:check":"oxfmt --check .","check:runtime":"node --input-type=module -e \"await import('./extensions/docparser/index.ts')\""},"_npmUser":{"name":"maxedapps","email":"maxedappsgmbh@gmail.com"},"repository":{"url":"git+https://github.com/maxedapps/pi-docparser.git","type":"git"},"_npmVersion":"11.6.2","description":"Pi package that adds a document_parse tool and companion skill for parsing PDFs, Office documents, spreadsheets, and images with LiteParse.","directories":{},"_nodeVersion":"24.11.1","dependencies":{"@llamaindex/liteparse":"1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"oxfmt":"^0.41.0","oxlint":"^1.56.0","typescript":"^5.9.3","@types/node":"^25.5.0","@mariozechner/pi-ai":"0.61.0","@mariozechner/pi-coding-agent":"0.61.0"},"peerDependencies":{"@mariozechner/pi-ai":"*","@mariozechner/pi-coding-agent":"*"},"_npmOperationalInternal":{"tmp":"tmp/pi-docparser_1.0.1_1774023572305_0.8110385481982723","host":"s3://npm-registry-packages-npm-production"}},"1.1.0":{"name":"pi-docparser","version":"1.1.0","keywords":["document-parse","documents","liteparse","ocr","pdf","pi","pi-package"],"license":"MIT","_id":"pi-docparser@1.1.0","maintainers":[{"name":"maxedapps","email":"maxedappsgmbh@gmail.com"}],"homepage":"https://maximilian-schwarzmueller.com","bugs":{"url":"https://github.com/maxedapps/pi-docparser/issues"},"pi":{"image":"https://raw.githubusercontent.com/maxedapps/pi-docparser/main/assets/pi-docparser-preview.jpg","skills":["./skills"],"extensions":["./extensions/docparser/index.ts"]},"dist":{"shasum":"629ddb9ae5f83d31fd0369fd251a9977e923c10e","tarball":"https://registry.npmjs.org/pi-docparser/-/pi-docparser-1.1.0.tgz","fileCount":17,"integrity":"sha512-LodHtaxNkIXRwN5q395S82Nous2V5zqGbLRdveg1elkcZ8s72VZd5eeJGseLQlu1NEDQ6mpmQu5PMGJZ8lIyvQ==","signatures":[{"sig":"MEQCIE/eFe6HBhDNN9tOoBk0LVRFLRPAR6bLpgw2dBtODmZMAiBykiHnj0zOlI5agyDHhf+PRNzjzkVx/R9dPhAfSmXwjQ==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":159656},"type":"module","engines":{"node":">=18.0.0"},"gitHead":"34b91239a52094e7a22064f7016a3addb0c68433","scripts":{"lint":"oxlint .","build":"echo 'nothing to build'","check":"oxfmt --check . && oxlint . && tsc -p tsconfig.json && node --input-type=module -e \"await import('./extensions/docparser/index.ts')\"","format":"oxfmt .","lint:fix":"oxlint --fix .","pack:dry":"npm pack --dry-run","typecheck":"tsc -p tsconfig.json","format:check":"oxfmt --check .","check:runtime":"node --input-type=module -e \"await import('./extensions/docparser/index.ts')\""},"_npmUser":{"name":"maxedapps","email":"maxedappsgmbh@gmail.com"},"repository":{"url":"git+https://github.com/maxedapps/pi-docparser.git","type":"git"},"_npmVersion":"11.6.2","description":"Pi package that adds a document_parse tool and companion skill for parsing PDFs, Office documents, spreadsheets, and images with LiteParse.","directories":{},"_nodeVersion":"24.11.1","dependencies":{"@llamaindex/liteparse":"1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"oxfmt":"^0.41.0","oxlint":"^1.56.0","typescript":"^5.9.3","@types/node":"^25.5.0","@mariozechner/pi-ai":"0.61.0","@mariozechner/pi-coding-agent":"0.61.0"},"peerDependencies":{"@mariozechner/pi-ai":"*","@mariozechner/pi-coding-agent":"*"},"_npmOperationalInternal":{"tmp":"tmp/pi-docparser_1.1.0_1774026670152_0.6949216502869544","host":"s3://npm-registry-packages-npm-production"}},"1.1.1":{"name":"pi-docparser","version":"1.1.1","keywords":["document-parse","documents","extension","liteparse","ocr","pdf","pi","pi-package","skill"],"license":"MIT","_id":"pi-docparser@1.1.1","maintainers":[{"name":"maxedapps","email":"maxedappsgmbh@gmail.com"}],"homepage":"https://maximilian-schwarzmueller.com","bugs":{"url":"https://github.com/maxedapps/pi-docparser/issues"},"pi":{"image":"https://raw.githubusercontent.com/maxedapps/pi-docparser/main/assets/pi-docparser-preview.jpg","skills":["./skills"],"extensions":["./extensions/docparser/index.ts"]},"dist":{"shasum":"c81edddef4a4a9ed832b3093fb61d6c8490e8cfd","tarball":"https://registry.npmjs.org/pi-docparser/-/pi-docparser-1.1.1.tgz","fileCount":17,"integrity":"sha512-eTdvRmbtrHDg434HCypuHnRYIhs8iWDjpFFyJlv1e72x/0hmB0Xp+jfx2uwNTvRPy8F7MDhONTA0RSAnZ2mGHg==","signatures":[{"sig":"MEUCIHUKJ4Y3bUJpLHZySQyfIEAqvuva8eaYQWWr9gkkYw1iAiEA4TPU13FDNKOVODglxDbjFzo5hOJCvKqw7JrUkhNzUPY=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":159863},"type":"module","engines":{"node":">=18.0.0"},"gitHead":"121b18ccf8d1199b5cb5aa50ff84175f663c95aa","scripts":{"lint":"oxlint .","build":"echo 'nothing to build'","check":"oxfmt --check . && oxlint . && tsc -p tsconfig.json && node --input-type=module -e \"await import('./extensions/docparser/index.ts')\"","format":"oxfmt .","lint:fix":"oxlint --fix .","pack:dry":"npm pack --dry-run","typecheck":"tsc -p tsconfig.json","format:check":"oxfmt --check .","check:runtime":"node --input-type=module -e \"await import('./extensions/docparser/index.ts')\""},"_npmUser":{"name":"maxedapps","email":"maxedappsgmbh@gmail.com"},"repository":{"url":"git+https://github.com/maxedapps/pi-docparser.git","type":"git"},"_npmVersion":"11.6.2","description":"Pi package that adds a document_parse tool and companion skill for parsing PDFs, Office documents, spreadsheets, and images with LiteParse.","directories":{},"_nodeVersion":"24.11.1","dependencies":{"@llamaindex/liteparse":"1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"oxfmt":"^0.41.0","oxlint":"^1.56.0","typescript":"^5.9.3","@types/node":"^25.5.0","@mariozechner/pi-ai":"0.61.0","@mariozechner/pi-coding-agent":"0.61.0"},"peerDependencies":{"@mariozechner/pi-ai":"*","@mariozechner/pi-coding-agent":"*"},"_npmOperationalInternal":{"tmp":"tmp/pi-docparser_1.1.1_1774071935554_0.9398736356964059","host":"s3://npm-registry-packages-npm-production"}},"2.0.0":{"name":"pi-docparser","version":"2.0.0","keywords":["document-parse","documents","extension","liteparse","ocr","pdf","pi","pi-package","skill"],"license":"MIT","_id":"pi-docparser@2.0.0","maintainers":[{"name":"maxedapps","email":"maxedappsgmbh@gmail.com"}],"homepage":"https://maximilian-schwarzmueller.com","bugs":{"url":"https://github.com/maxedapps/pi-docparser/issues"},"pi":{"image":"https://raw.githubusercontent.com/maxedapps/pi-docparser/main/assets/pi-docparser-preview.jpg","skills":["./skills"],"extensions":["./extensions/docparser/index.ts"]},"dist":{"shasum":"2dc4d3c1efa7427fba66581b221b00353784740e","tarball":"https://registry.npmjs.org/pi-docparser/-/pi-docparser-2.0.0.tgz","fileCount":17,"integrity":"sha512-jzYyMQ70x5U58x3Z7Is+2TgUeU6vmUjrcpfhogh07+iMeyPR6/zXJpjicHAbgS+GKPzeSKMznUi55ku9PmLpkA==","signatures":[{"sig":"MEUCIGAbLddp1AFzDPUlwiWCdr+hfXfB60ylz3UzxnbwV7J7AiEA80pusZieeEPYDuiRRV0Xkke+w7y5H9B/kQfD1hX86Xk=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":160797},"type":"module","engines":{"node":">=20.6.0"},"gitHead":"c15df14b6cfe31c35c3bb45aae61190fff09dbbe","scripts":{"lint":"oxlint .","build":"echo 'nothing to build'","check":"oxfmt --check . && oxlint . && tsc -p tsconfig.json && node --input-type=module -e \"await import('./extensions/docparser/index.ts')\"","format":"oxfmt .","lint:fix":"oxlint --fix .","pack:dry":"npm pack --dry-run","typecheck":"tsc -p tsconfig.json","format:check":"oxfmt --check .","check:runtime":"node --input-type=module -e \"await import('./extensions/docparser/index.ts')\""},"_npmUser":{"name":"maxedapps","email":"maxedappsgmbh@gmail.com"},"repository":{"url":"git+https://github.com/maxedapps/pi-docparser.git","type":"git"},"_npmVersion":"11.14.1","description":"Pi package that adds a document_parse tool and companion skill for parsing PDFs, Office documents, spreadsheets, and images with LiteParse.","directories":{},"_nodeVersion":"26.1.0","dependencies":{"@llamaindex/liteparse":"1.5.3"},"_hasShrinkwrap":false,"devDependencies":{"oxfmt":"^0.49.0","oxlint":"^1.64.0","typescript":"^6.0.3","@types/node":"^25.7.0","@earendil-works/pi-ai":"0.74.0","@earendil-works/pi-coding-agent":"0.74.0"},"peerDependencies":{"@earendil-works/pi-ai":"^0.74.0","@earendil-works/pi-coding-agent":"^0.74.0"},"_npmOperationalInternal":{"tmp":"tmp/pi-docparser_2.0.0_1778662900190_0.22817413337061354","host":"s3://npm-registry-packages-npm-production"}},"3.0.0":{"name":"pi-docparser","version":"3.0.0","keywords":["document-parse","documents","extension","liteparse","ocr","pdf","pi","pi-package","skill"],"license":"MIT","_id":"pi-docparser@3.0.0","maintainers":[{"name":"maxedapps","email":"maxedappsgmbh@gmail.com"}],"homepage":"https://maximilian-schwarzmueller.com","bugs":{"url":"https://github.com/maxedapps/pi-docparser/issues"},"pi":{"image":"https://raw.githubusercontent.com/maxedapps/pi-docparser/main/assets/pi-docparser-preview.jpg","skills":["./skills"],"extensions":["./extensions/docparser/index.ts"]},"dist":{"shasum":"b03e91be829124b393b5e8935043f0c2410997ea","tarball":"https://registry.npmjs.org/pi-docparser/-/pi-docparser-3.0.0.tgz","fileCount":20,"integrity":"sha512-nmu5FuMpYvOz6Ac9FY6ihow/XT5AUHEAnI8htZPzaJzWv62g0vuU5yctLEFOQtIBiBGRD+FRsXBJ8Dre3FOYpw==","signatures":[{"sig":"MEQCIGvMJeVFGJpBNvtzKNzadbiPUG9zfk2TRDUXzW5jHb0vAiBtKcFDJIu6gICbtSRXeNkQT+GEBWO7wNg58CdzZVGqsQ==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":180977},"type":"module","engines":{"node":">=20.6.0"},"gitHead":"0dcddb3aef6ed20d9f20685feabcf0ebdc121a39","scripts":{"lint":"oxlint .","build":"echo 'nothing to build'","check":"oxfmt --check . && oxlint . && tsc -p tsconfig.json && node --input-type=module -e \"await import('./extensions/docparser/index.ts')\"","format":"oxfmt .","lint:fix":"oxlint --fix .","pack:dry":"npm pack --dry-run","typecheck":"tsc -p tsconfig.json","format:check":"oxfmt --check .","check:runtime":"node --input-type=module -e \"await import('./extensions/docparser/index.ts')\""},"_npmUser":{"name":"maxedapps","email":"maxedappsgmbh@gmail.com"},"repository":{"url":"git+https://github.com/maxedapps/pi-docparser.git","type":"git"},"_npmVersion":"11.14.1","description":"Pi package that adds document_parse, document_search, document_screenshot, and a companion skill for local document understanding with LiteParse v2.","directories":{},"_nodeVersion":"26.1.0","dependencies":{"@llamaindex/liteparse":"2.0.1"},"_hasShrinkwrap":false,"devDependencies":{"oxfmt":"^0.49.0","oxlint":"^1.64.0","typescript":"^6.0.3","@types/node":"^25.7.0","@earendil-works/pi-ai":"0.74.0","@earendil-works/pi-coding-agent":"0.74.0"},"peerDependencies":{"@earendil-works/pi-ai":"^0.74.0","@earendil-works/pi-coding-agent":"^0.74.0"},"_npmOperationalInternal":{"tmp":"tmp/pi-docparser_3.0.0_1779967076066_0.4830841180722407","host":"s3://npm-registry-packages-npm-production"}},"3.0.1":{"name":"pi-docparser","version":"3.0.1","keywords":["document-parse","documents","extension","liteparse","ocr","pdf","pi","pi-package","skill"],"license":"MIT","_id":"pi-docparser@3.0.1","maintainers":[{"name":"maxedapps","email":"maxedappsgmbh@gmail.com"}],"homepage":"https://maximilian-schwarzmueller.com","bugs":{"url":"https://github.com/maxedapps/pi-docparser/issues"},"pi":{"image":"https://raw.githubusercontent.com/maxedapps/pi-docparser/main/assets/pi-docparser-preview.jpg","skills":["./skills"],"extensions":["./extensions/docparser/index.ts"]},"dist":{"shasum":"5a5f84c649c1921a6ad899899d1758c55bc5df73","tarball":"https://registry.npmjs.org/pi-docparser/-/pi-docparser-3.0.1.tgz","fileCount":20,"integrity":"sha512-t08KQlV6jnvXM9usnOWMyxLJL5rlwHjOO/Dy/GXF8x6PIvaFW7FGk4+l3a+HQq+nJJZQ9no9ETQuNg4ZYcN9tg==","signatures":[{"sig":"MEYCIQDGdMi7kvPC22Ksti0HY0VG6JEIGKWY383t2bnKfPPv8QIhAPkdjr/rXIWGmnqnWt93Co9pwRTa4uaKWrbKpngDvzSU","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":180997},"type":"module","engines":{"node":">=20.6.0"},"gitHead":"b944642fe0bf01b89e6ae784c034785f694c9d62","scripts":{"lint":"oxlint .","build":"echo 'nothing to build'","check":"oxfmt --check . && oxlint . && tsc -p tsconfig.json && node --input-type=module -e \"await import('./extensions/docparser/index.ts')\"","format":"oxfmt .","lint:fix":"oxlint --fix .","pack:dry":"npm pack --dry-run","typecheck":"tsc -p tsconfig.json","format:check":"oxfmt --check .","check:runtime":"node --input-type=module -e \"await import('./extensions/docparser/index.ts')\""},"_npmUser":{"name":"maxedapps","email":"maxedappsgmbh@gmail.com"},"repository":{"url":"git+https://github.com/maxedapps/pi-docparser.git","type":"git"},"_npmVersion":"11.14.1","description":"Pi package that adds document_parse, document_search, document_screenshot, and a companion skill for local document understanding with LiteParse v2.","directories":{},"_nodeVersion":"26.1.0","dependencies":{"@llamaindex/liteparse":"2.0.1"},"_hasShrinkwrap":false,"devDependencies":{"oxfmt":"^0.49.0","oxlint":"^1.64.0","typescript":"^6.0.3","@types/node":"^25.7.0","@earendil-works/pi-ai":"0.74.0","@earendil-works/pi-coding-agent":"0.74.0"},"peerDependencies":{"@earendil-works/pi-ai":"^0.74.0","@earendil-works/pi-coding-agent":"^0.74.0"},"_npmOperationalInternal":{"tmp":"tmp/pi-docparser_3.0.1_1779967877711_0.9094228572393179","host":"s3://npm-registry-packages-npm-production"}},"4.0.0":{"name":"pi-docparser","version":"4.0.0","description":"Pi package that adds document_parse, document_search, document_screenshot, and a companion skill for local document understanding with LiteParse v2.","keywords":["document-parse","documents","extension","liteparse","ocr","pdf","pi","pi-package","skill"],"homepage":"https://maximilian-schwarzmueller.com","license":"MIT","repository":{"type":"git","url":"git+https://github.com/maxedapps/pi-docparser.git"},"type":"module","scripts":{"build":"echo 'nothing to build'","format":"oxfmt .","format:check":"oxfmt --check .","lint":"oxlint .","lint:fix":"oxlint --fix .","typecheck":"tsc -p tsconfig.json","test":"node --import tsx --test --test-concurrency=1 tests/*.test.ts","test:unit":"node --import tsx --test tests/config-policy.test.ts tests/deps.test.ts tests/input-policy.test.ts tests/package-metadata.test.ts tests/parse-output.test.ts tests/tools.test.ts","test:native":"node --import tsx --test --test-concurrency=1 tests/native-executor.test.ts tests/native-worker.integration.test.ts","test:packed":"node --import tsx --test tests/packed-install.test.ts","test:package":"node --import tsx --test tests/package-content.test.ts","check:runtime":"node --input-type=module -e \"await import('./extensions/docparser/index.ts')\"","check":"pnpm run format:check && pnpm run lint && pnpm run typecheck && pnpm run test && pnpm run check:runtime && pnpm run pack:dry","pack:dry":"npm pack --dry-run"},"dependencies":{"@llamaindex/liteparse":"2.10.1"},"devDependencies":{"@earendil-works/pi-ai":"0.83.0","@earendil-works/pi-coding-agent":"0.83.0","@types/node":"^25.9.5","oxfmt":"^0.49.0","oxlint":"^1.76.0","tsx":"^4.20.6","typescript":"^6.0.3"},"peerDependencies":{"@earendil-works/pi-ai":"*","@earendil-works/pi-coding-agent":"*"},"engines":{"node":">=22.19.0"},"pi":{"extensions":["./extensions/docparser/index.ts"],"skills":["./skills"],"image":"https://raw.githubusercontent.com/maxedapps/pi-docparser/main/assets/pi-docparser-preview.jpg"},"gitHead":"931a0995067062c91bd81798ef226d120c31bd84","_id":"pi-docparser@4.0.0","bugs":{"url":"https://github.com/maxedapps/pi-docparser/issues"},"_nodeVersion":"26.1.0","_npmVersion":"11.14.1","dist":{"integrity":"sha512-4Pdte8rHDi+UZ1XsTsayKgzgIgKoyaiQeY3UCsdMz4aPQr/lIpS6gD+kY6/rPRy0BkXUzmBJ1ep23CWYoTba3g==","shasum":"e54309bddbbc120c17e2f574fe50c2dd307d318d","tarball":"https://registry.npmjs.org/pi-docparser/-/pi-docparser-4.0.0.tgz","fileCount":23,"unpackedSize":256182,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQCzgWtBpjwtc9R1mtk3CyY+ZXkLbYlQks9Nx/6PC3bAawIge2Tv1TmWHkhsuVXPnHsNWVMISnsCKrU0kPu9cxfYySk="}]},"_npmUser":{"name":"maxedapps","email":"maxedappsgmbh@gmail.com"},"directories":{},"maintainers":[{"name":"maxedapps","email":"maxedappsgmbh@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/pi-docparser_4.0.0_1785768038463_0.1999704572014367"},"_hasShrinkwrap":false}},"time":{"created":"2026-03-20T16:04:49.604Z","modified":"2026-08-03T14:40:38.781Z","1.0.0":"2026-03-20T16:04:49.766Z","1.0.1":"2026-03-20T16:19:32.484Z","1.1.0":"2026-03-20T17:11:10.301Z","1.1.1":"2026-03-21T05:45:35.713Z","2.0.0":"2026-05-13T09:01:40.417Z","3.0.0":"2026-05-28T11:17:56.221Z","3.0.1":"2026-05-28T11:31:17.903Z","4.0.0":"2026-08-03T14:40:38.619Z"},"bugs":{"url":"https://github.com/maxedapps/pi-docparser/issues"},"license":"MIT","homepage":"https://maximilian-schwarzmueller.com","keywords":["document-parse","documents","extension","liteparse","ocr","pdf","pi","pi-package","skill"],"repository":{"type":"git","url":"git+https://github.com/maxedapps/pi-docparser.git"},"description":"Pi package that adds document_parse, document_search, document_screenshot, and a companion skill for local document understanding with LiteParse v2.","maintainers":[{"name":"maxedapps","email":"maxedappsgmbh@gmail.com"}],"readme":"# pi-docparser\n\nA standalone [Pi](https://shittycodingagent.ai/) package that adds local document-understanding tools plus a companion `parse-document` skill for AI agents.\n\nIt wraps [`@llamaindex/liteparse`](https://github.com/run-llama/liteparse) 2.10.1, a Rust/PDFium-based local parser. Document processing stays on the local machine, with no LLM parsing or API key required. Built-in OCR may download missing language data unless local `.traineddata` files are supplied; configuring `ocrServerUrl` sends OCR work to that server.\n\n## What this package provides\n\n### Extension tools\n\nThis package registers three tools:\n\n| Tool                  | Purpose                                                                                                                  |\n| --------------------- | ------------------------------------------------------------------------------------------------------------------------ |\n| `document_parse`      | Parse a local document to `text` or stable `json`, save the full result to a temp file, and optionally save screenshots. |\n| `document_search`     | Search a local document for a phrase and return bounded page-number and bounding-box hits.                               |\n| `document_screenshot` | Render up to four document pages as PNGs, return bounded image blocks, and save every PNG to a temp folder.              |\n\nUse `document_parse` for extraction, `document_search` for citations/source locations, and `document_screenshot` when visual layout, charts, signatures, dense tables, or page appearance matter.\n\n### Skill\n\nShips a `parse-document` skill that teaches agents to:\n\n- prefer `document_parse` over raw `lit` CLI commands\n- choose text vs JSON output deliberately\n- request small, explicit page ranges\n- search before screenshotting when looking for known text\n- use saved output paths for selective follow-up\n\n## Requirements and compatibility\n\n- Node.js 22.19 or newer\n- Pi installed and working; Pi 0.83 is the development and test baseline\n- local machine access to readable regular files you want to parse\n- LibreOffice for many Office, presentation, and spreadsheet conversion paths\n\nThe Pi core peer ranges follow Pi's `\"*\"` package policy. That avoids pinning the host's core packages; it does not promise compatibility with every historical Pi version.\n\nThe Node floor, safer defaults and hard limits, removal of all-page screenshot behavior, bounded result details, and stable projected JSON are 4.0.0 breaking changes.\n\n## Installation\n\n```bash\npi install npm:pi-docparser\n```\n\nOr from GitHub:\n\n```bash\npi install git:github.com/maxedapps/pi-docparser\n```\n\n## Example model tool calls\n\nThese are representative tool calls Pi may make internally. Prefer small page ranges and repeat with the next bounded range when needed.\n\n### Extract plain text\n\n```text\ndocument_parse({\n  path: \"./docs/contract.pdf\",\n  targetPages: \"1-10\"\n})\n```\n\nUseful for summarizing, quoting, reviewing, or answering questions where layout coordinates are not needed.\n\n### Extract JSON with bounding boxes\n\n```text\ndocument_parse({\n  path: \"./reports/financial-report.pdf\",\n  format: \"json\",\n  targetPages: \"1-3\"\n})\n```\n\nUseful when an agent needs page structure, text coordinates, or bounding boxes.\n\n### Search for a phrase and get source locations\n\n```text\ndocument_search({\n  path: \"./reports/financial-report.pdf\",\n  phrase: \"Revenue grew\",\n  targetPages: \"1-10\",\n  maxResults: 50\n})\n```\n\nReturns bounded hits with page numbers and bounding boxes, useful for citations and deciding which pages to screenshot.\n\n### Render pages for visual inspection\n\n```text\ndocument_screenshot({\n  path: \"./reports/financial-report.pdf\",\n  pages: \"4\",\n  dpi: 150\n})\n```\n\nOmitting `pages` renders page 1. A call may name at most four explicit pages; `all` and `*` are rejected. Use bounded repeated calls for more pages.\n\n### Parse and save screenshots together\n\n```text\ndocument_parse({\n  path: \"./reports/financial-report.pdf\",\n  targetPages: \"1-4\",\n  screenshotPages: \"2,4\"\n})\n```\n\nThe parse result and screenshots are separate artifacts. If optional screenshot rendering fails, the completed parse output remains available and the tool returns a warning.\n\n### Parse a password-protected document\n\n```text\ndocument_parse({\n  path: \"./docs/protected.pdf\",\n  targetPages: \"1-5\",\n  password: \"user-provided-password\"\n})\n```\n\n### Use offline/custom OCR data\n\n```text\ndocument_parse({\n  path: \"./scans/report.pdf\",\n  targetPages: \"1-5\",\n  ocr: \"auto\",\n  ocrLanguage: \"eng\",\n  tessdataPath: \"/path/to/tessdata\"\n})\n```\n\n`tessdataPath` points LiteParse/Tesseract at locally supplied `.traineddata` files. This is useful for air-gapped environments, predictable offline operation, or custom language packs. Without supplied local language data, built-in OCR may download missing data.\n\n## Stable JSON contract\n\n`format: \"json\"` writes a project-owned, stable `{ pages, text }` projection rather than the raw upstream LiteParse object:\n\n```json\n{\n  \"pages\": [\n    {\n      \"pageNum\": 1,\n      \"width\": 612,\n      \"height\": 792,\n      \"text\": \"...\",\n      \"textItems\": [\n        {\n          \"text\": \"Revenue\",\n          \"x\": 72,\n          \"y\": 120,\n          \"width\": 48,\n          \"height\": 12,\n          \"fontName\": \"Helvetica\",\n          \"fontSize\": 12,\n          \"confidence\": 0.99\n        }\n      ]\n    }\n  ],\n  \"text\": \"...\"\n}\n```\n\nEvery page contains `pageNum`, `width`, `height`, `text`, and `textItems`. Every text item contains `text`, `x`, `y`, `width`, and `height`; `fontName`, `fontSize`, and `confidence` are optional. Upstream-only fields such as `markdown`, `images`, `imageErrorCount`, and metadata are not exposed implicitly. The artifact is compact JSON and field order is stable.\n\nRemoved LiteParse v1 options are not supported:\n\n- `preciseBoundingBox`\n- `preserveLayoutAlignmentAcrossPages`\n\nUse JSON `textItems`, `document_search`, `document_screenshot`, or a narrower `targetPages` selection instead.\n\n## Defaults and limits\n\nRequests are rejected rather than silently clamped.\n\n| Behavior                           | Default                                      | Limit                                                             |\n| ---------------------------------- | -------------------------------------------- | ----------------------------------------------------------------- |\n| Pages parsed/searched (`maxPages`) | 100                                          | 1–1000                                                            |\n| OCR workers (`numWorkers`)         | `min(4, max(1, availableParallelism() - 1))` | 1–8                                                               |\n| OCR/rendering DPI (`dpi`)          | 150                                          | 72–300                                                            |\n| Search hits (`maxResults`)         | 50                                           | 1–200                                                             |\n| Search phrase                      | —                                            | nonblank; 4 KiB UTF-8                                             |\n| Page-selection input               | —                                            | 16 KiB UTF-8, 1000 tokens, 1000 pages of pre-dedup expansion work |\n| Page number                        | —                                            | 1–4,294,967,295                                                   |\n| Screenshot selection               | page 1 for `document_screenshot`             | at most 4 explicit pages; no `all` or `*`                         |\n\n`maxPages` counts pages actually parsed or explicitly selected, not the highest sparse page number. For example, `targetPages: \"2,100\"` selects two pages. `document_parse.screenshotPages` is optional and creates no screenshots when omitted; when supplied, it follows the same explicit four-page limit.\n\nSafety budgets:\n\n- parsed text or JSON artifact: 256 MiB\n- saved screenshot: 25 MiB per PNG and 64 MiB per screenshot job\n- inline screenshot images: 3 MiB per PNG and 12 MiB raw total per tool result\n- worker request/response: 64 KiB/1 MiB, with a 64 KiB retained stderr tail\n- native operation timeout: one internal, non-configurable 10-minute deadline, starting when the job reaches the front of the queue\n\nA PNG omitted from inline image blocks because of the inline limits is still saved, and its path is returned for follow-up inspection.\n\n## Isolation, cancellation, and failures\n\nAll parse, search, and screenshot native work shares one fair FIFO per extension activation. Each active operation runs in a fresh child process, so LiteParse/PDFium/Tesseract native crashes, aborts, or fatal worker out-of-memory failures become tool failures instead of bringing down the Pi process. The parent process does not import LiteParse's native module.\n\nQueued cancellation removes the job before a worker starts. Active cancellation and the internal 10-minute timeout terminate the worker process tree and wait for teardown before another native job starts. Cancellation, timeout, native crash, protocol failure, and ordinary parse errors are reported distinctly. If process teardown cannot be confirmed, the executor fails closed and rejects later jobs for that activation.\n\nOn Windows, descendant cleanup after an instantaneous native worker crash is best effort because the root may exit before the process tree can be addressed. This does not weaken containment of the Pi process itself.\n\n## Output paths and previews\n\nSuccessful full outputs are written to OS temporary directories and returned in the tool result:\n\n- `document_parse`: `.../pi-document-parse-*/parsed.txt` or `parsed.json`\n- optional parse screenshots: `.../pi-document-parse-*/screenshots/page_<n>.png`\n- `document_screenshot`: `.../pi-document-screenshot-*/screenshots/page_<n>.png`\n\n`document_parse` returns only a 20-line/2 KiB preview plus the saved path. Use `read` on that path for the full content. `document_screenshot` returns eligible bounded image blocks and paths for every saved PNG; use the paths for images omitted from inline content. Outputs remain temporary by default—copy them to a chosen persistent location when the user requests durable artifacts.\n\n## Supported inputs\n\nThis package supports LiteParse's local formats, including:\n\n- PDF\n- DOC / DOCX / DOCM / ODT / RTF / Pages\n- PPT / PPTX / PPTM / ODP / Keynote\n- XLS / XLSX / XLSM / ODS / CSV / TSV / Numbers\n- PNG / JPG / JPEG / GIF / BMP / TIFF / WebP / SVG\n\nLiteParse 2.10.1 handles image conversion natively, with no external image-conversion tool required. Many Office, presentation, and spreadsheet formats still require LibreOffice.\n\n## Tool behavior notes\n\n### `document_parse`\n\n- Writes bounded plain text or the stable projected JSON contract to a temp file.\n- Returns a bounded preview and the full output path.\n- Supports `targetPages`, OCR options, `password`, `tessdataPath`, and optional explicit `screenshotPages`.\n- Defaults `maxPages` to 100.\n- Enforces a hard `maxPages` maximum of 1000.\n\n### `document_search`\n\n- Parses and searches inside the isolated native worker.\n- Returns projected hits with `pageNum`, `text`, `x`, `y`, `width`, `height`, and optional confidence/font data.\n- Reports count or response-byte truncation explicitly.\n- Use before screenshotting when searching for known text.\n\n### `document_screenshot`\n\n- Renders one page at a time, up to four explicit pages, and saves every PNG.\n- Returns only PNGs within the inline image budgets as image content blocks.\n- Can render supported non-PDF documents when required Office conversion tools are installed.\n\n### OCR notes\n\nLiteParse uses built-in native Tesseract OCR by default when OCR is enabled and no `ocrServerUrl` is provided.\n\n- OCR is selective: LiteParse OCRs text-sparse pages or image regions rather than blindly OCRing everything.\n- Built-in Tesseract typically uses ISO 639-3 language codes such as `eng`, `deu`, `fra`, `jpn`.\n- Many HTTP OCR servers instead expect ISO 639-1 codes such as `en`, `de`, `fr`, `ja`.\n- `ocrLanguages` is joined into a multilingual language string for built-in Tesseract.\n- When `ocrServerUrl` is used, only the first entry from `ocrLanguages` is forwarded.\n- Built-in OCR may download missing language data. For local-only language data, use `tessdataPath` or set `TESSDATA_PREFIX` to supplied `.traineddata` files.\n\n## Host dependencies\n\n### LibreOffice\n\nNeeded for many Office document, presentation, and spreadsheet conversion paths.\n\n```bash\n# macOS\nbrew install --cask libreoffice\n\n# Ubuntu / Debian\napt-get install libreoffice\n\n# Windows\nchoco install libreoffice-fresh\n```\n\n## Doctor command\n\nIf parsing fails because LibreOffice is missing, the extension points users to:\n\n```text\n/docparser:doctor\n```\n\nRun it inside Pi to:\n\n- detect the current operating system\n- check whether LibreOffice is available\n- optionally focus the check on a specific file path\n- suggest install commands for the current machine\n- optionally attempt those install commands after user confirmation when safe to automate\n\nExamples:\n\n```text\n/docparser:doctor\n/docparser:doctor @./slides.pptx\n```\n\n## Known limitations\n\n- OCR quality depends on scan quality, page layout, the chosen language, and available language data.\n- Office-family conversion paths may depend on LibreOffice.\n- Successful artifacts remain in temporary directories until OS cleanup; copy files that must persist.\n- A worker may still exhaust its own memory, but child-process isolation prevents that failure from terminating Pi.\n- Native LiteParse npm packages are platform-specific; unsupported platforms need upstream LiteParse support.\n\n## Third-party dependency: LiteParse\n\nThis package depends on:\n\n- [`@llamaindex/liteparse`](https://github.com/run-llama/liteparse) 2.10.1\n- license: Apache-2.0\n- purpose: local document parsing, OCR, screenshots, search, and conversion support\n\nLiteParse documents its upstream dependencies and platform requirements. See:\n\n- repository: https://github.com/run-llama/liteparse\n- npm package: https://www.npmjs.com/package/@llamaindex/liteparse\n- docs: https://developers.llamaindex.ai/liteparse/\n\nAdditional attribution details are listed in [THIRD_PARTY_NOTICES.md](./THIRD_PARTY_NOTICES.md).\n\n## Changelog\n\nSee [CHANGELOG.md](./CHANGELOG.md).\n\n## License\n\nThis package is licensed under the MIT License. See [LICENSE](./LICENSE).\n","readmeFilename":"README.md"}