{"_id":"node-pptx-parser","_rev":"1-7875e47c4d9d1cd2d209a9c39cad3c86","name":"node-pptx-parser","dist-tags":{"latest":"1.0.1"},"versions":{"1.0.0":{"name":"node-pptx-parser","version":"1.0.0","keywords":["pptx","powerpoint","parser","text","extract"],"author":{"name":"Mirza-Glitch"},"license":"MIT","_id":"node-pptx-parser@1.0.0","maintainers":[{"name":"mirza-glitch","email":"mirzaglitch@gmail.com"}],"homepage":"https://github.com/Mirza-Glitch/node-pptx-parser.git","bugs":{"url":"https://github.com/Mirza-Glitch/node-pptx-parser/issues"},"dist":{"shasum":"542650c4bb56c021274de9b395401b6712a5f570","tarball":"https://registry.npmjs.org/node-pptx-parser/-/node-pptx-parser-1.0.0.tgz","fileCount":9,"integrity":"sha512-f4Tu//dCPav6Ow9hjF9EeaVwgH3MOkUale2+TniKnxT2z0cPE389kOw1pXW0/JrqRT/0KnpJXV3NOoomXuA4PA==","signatures":[{"sig":"MEUCIBLMTy1U+goQSIIoXhU6gtRPqlNhsql0E8+dkMyO05e2AiEAiQybg4xCQMARFpQ4oBI0zYGgXvymORgndmgUkMm/+ps=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":30039},"main":"./dist/cjs/index.js","types":"./dist/types/index.d.ts","module":"./dist/esm/index.mjs","exports":{".":{"import":"./dist/esm/index.mjs","require":"./dist/cjs/index.js"}},"gitHead":"8129dda7265a60cf89a025f2c10123d3b8ee4d14","scripts":{"build":"npm run build:cjs && npm run build:esm","clean":"rimraf dist","prepack":"npm run clean && npm run build","build:cjs":"tsc -p tsconfig.cjs.json","build:esm":"tsc -p tsconfig.esm.json && npm run rename:esm","rename:esm":"bash ./scripts/fix-mjs.sh"},"_npmUser":{"name":"mirza-glitch","email":"mirzaglitch@gmail.com"},"repository":{"url":"git+https://github.com/Mirza-Glitch/node-pptx-parser.git","type":"git"},"_npmVersion":"10.9.0","description":"A PowerPoint (PPTX) parser that extracts text content with preserved formatting","directories":{},"_nodeVersion":"22.11.0","dependencies":{"xml2js":"^0.6.2","unzipper":"^0.12.3"},"_hasShrinkwrap":false,"devDependencies":{"rimraf":"^6.0.1","typescript":"^5.7.3","@types/node":"^22.13.4","@types/xml2js":"^0.4.14","@types/unzipper":"^0.10.10"},"_npmOperationalInternal":{"tmp":"tmp/node-pptx-parser_1.0.0_1739817438656_0.4731535126694113","host":"s3://npm-registry-packages-npm-production"}},"1.0.1":{"name":"node-pptx-parser","version":"1.0.1","description":"A PowerPoint (PPTX) parser that extracts text content with preserved formatting","homepage":"https://github.com/Mirza-Glitch/node-pptx-parser.git","repository":{"type":"git","url":"git+https://github.com/Mirza-Glitch/node-pptx-parser.git"},"keywords":["pptx","powerpoint","parser","text","extract"],"author":{"name":"Mirza-Glitch"},"main":"./dist/cjs/index.js","module":"./dist/esm/index.mjs","types":"./dist/types/index.d.ts","exports":{".":{"require":"./dist/cjs/index.js","import":"./dist/esm/index.mjs","types":"./dist/types/index.d.ts"}},"scripts":{"build:cjs":"tsc -p tsconfig.cjs.json","build:esm":"tsc -p tsconfig.esm.json && npm run rename:esm","build":"npm run build:cjs && npm run build:esm","clean":"rimraf dist","rename:esm":"bash ./scripts/fix-mjs.sh","prepack":"npm run clean && npm run build"},"dependencies":{"unzipper":"^0.12.3","xml2js":"^0.6.2"},"devDependencies":{"@types/node":"^22.13.4","@types/unzipper":"^0.10.10","@types/xml2js":"^0.4.14","rimraf":"^6.0.1","typescript":"^5.7.3"},"license":"MIT","_id":"node-pptx-parser@1.0.1","gitHead":"666b24331f73ecb7cbf029e9d87e7a982e24041a","bugs":{"url":"https://github.com/Mirza-Glitch/node-pptx-parser/issues"},"_nodeVersion":"22.11.0","_npmVersion":"10.9.0","dist":{"integrity":"sha512-IhUI8U8PofZ9uh0JlD7fhUwuShbMLkAqHH7S6sGjVAc9W0NCZzq7YB3BzIkM+hYBcTqvEntU5ICOIfHZ+K0HLQ==","shasum":"9e7c3147360c305bd925e931786377e2ef2fb555","tarball":"https://registry.npmjs.org/node-pptx-parser/-/node-pptx-parser-1.0.1.tgz","fileCount":9,"unpackedSize":30082,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQDaKYXGPzLU0hIiicDAzLg+AQ9lxHBud6Im1vqsA0btFAIhAJVmCkjvqJ5zJPH8XE3Y1NYlGXxmvS+jPCtn4TedoGjF"}]},"_npmUser":{"name":"mirza-glitch","email":"mirzaglitch@gmail.com"},"directories":{},"maintainers":[{"name":"mirza-glitch","email":"mirzaglitch@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/node-pptx-parser_1.0.1_1739820717667_0.9452365251968129"},"_hasShrinkwrap":false}},"time":{"created":"2025-02-17T18:37:18.655Z","modified":"2025-02-17T19:31:58.096Z","1.0.0":"2025-02-17T18:37:18.839Z","1.0.1":"2025-02-17T19:31:57.908Z"},"bugs":{"url":"https://github.com/Mirza-Glitch/node-pptx-parser/issues"},"author":{"name":"Mirza-Glitch"},"license":"MIT","homepage":"https://github.com/Mirza-Glitch/node-pptx-parser.git","keywords":["pptx","powerpoint","parser","text","extract"],"repository":{"type":"git","url":"git+https://github.com/Mirza-Glitch/node-pptx-parser.git"},"description":"A PowerPoint (PPTX) parser that extracts text content with preserved formatting","maintainers":[{"name":"mirza-glitch","email":"mirzaglitch@gmail.com"}],"readme":"# node-pptx-parser\r\n\r\nA Node.js library for parsing PowerPoint (PPTX) files and extracting text content. This library maintains text formatting, line breaks, and paragraph structures from the original presentation.\r\n\r\n## Features\r\n\r\n- Extract text content from PPTX files with preserved formatting\r\n\r\n- Parse PPTX structure into manageable JavaScript objects\r\n\r\n- Access raw XML content of presentation components\r\n\r\n- Written in TypeScript for type safety\r\n\r\n- Promise-based API\r\n\r\n- Preserves line breaks and paragraph formatting\r\n\r\n- Minimal dependencies\r\n\r\n## Installation\r\n\r\n```bash\r\n\r\nnpm  install  node-pptx-parser\r\n```\r\n\r\n## Usage\r\n\r\nOnce the package is installed you can you it with `import` or `require` statements like this:\r\n\r\n```javascript\r\n// ESM import:\r\nimport PptxParser from \"node-pptx-parser\";\r\n\r\n// CommonJs require:\r\nconst PptxParser = require(\"node-pptx-parser\").default;\r\n```\r\n\r\n### Basic Text Extraction\r\n\r\n```typescript\r\nimport PptxParser from \"node-pptx-parser\";\r\n\r\nasync function main() {\r\n  const parser = new PptxParser(\"presentation.pptx\");\r\n\r\n  try {\r\n    // Extract text from all slides\r\n    const textContent = await parser.extractText();\r\n\r\n    // Print text from each slide\r\n    textContent.forEach((slide) => {\r\n      console.log(`\\nSlide ${slide.id}:`);\r\n\r\n      console.log(slide.text.join(\"\\n\"));\r\n    });\r\n  } catch (error) {\r\n    console.error(\"Error:\", error.message);\r\n  }\r\n}\r\n\r\nmain();\r\n```\r\n\r\n### Advanced Usage - Full Presentation Parsing\r\n\r\n```typescript\r\nimport PptxParser from \"node-pptx-parser\";\r\n\r\nasync function main() {\r\n  const parser = new PptxParser(\"presentation.pptx\");\r\n\r\n  try {\r\n    // Get complete parsed presentation content\r\n    const parsedContent = await parser.parse();\r\n\r\n    // Access presentation structure\r\n    console.log(parsedContent.presentation.parsed);\r\n\r\n    // Access individual slides\r\n    parsedContent.slides.forEach((slide) => {\r\n      console.log(`Slide ${slide.id}:`, slide.parsed);\r\n    });\r\n\r\n    // Access raw XML if needed\r\n    console.log(parsedContent.presentation.xml);\r\n  } catch (error) {\r\n    console.error(\"Error:\", error.message);\r\n  }\r\n}\r\n\r\nmain();\r\n```\r\n\r\n## API Reference\r\n\r\n### `PptxParser`\r\n\r\nThe main class for parsing PPTX files.\r\n\r\n#### Constructor\r\n\r\n```typescript\r\n\r\nconstructor(filePath: string)\r\n```\r\n\r\nCreates a new instance of PptxParser.\r\n\r\n- `filePath`: Path to the PPTX file to be parsed\r\n\r\n#### Methods\r\n\r\n##### `parse()`\r\n\r\n```typescript\r\n\r\nasync parse(): Promise<ParsedPresentation>\r\n```\r\n\r\nParses the entire PPTX file and returns its content.\r\n\r\n- Returns: Promise resolving to a `ParsedPresentation` object containing the complete presentation structure\r\n\r\n##### `extractText()`\r\n\r\n```typescript\r\n\r\nasync extractText(): Promise<SlideTextContent[]>\r\n```\r\n\r\nExtracts formatted text content from all slides.\r\n\r\n- Returns: Promise resolving to an array of `SlideTextContent` objects\r\n\r\n### Types\r\n\r\n#### `ParsedPresentation`\r\n\r\n```typescript\r\ninterface ParsedPresentation {\r\n  presentation: {\r\n    path: string;\r\n    xml: string;\r\n    parsed: any;\r\n  };\r\n  relationships: {\r\n    path: string;\r\n    xml: string;\r\n    parsed: any;\r\n  };\r\n  slides: ParsedSlide[];\r\n}\r\n```\r\n\r\n#### `ParsedSlide`\r\n\r\n```typescript\r\ninterface ParsedSlide {\r\n  id: string;\r\n  path: string;\r\n  xml: string;\r\n  parsed: any;\r\n}\r\n```\r\n\r\n#### `SlideTextContent`\r\n\r\n```typescript\r\ninterface SlideTextContent extends ParsedSlide {\r\n  text: string[];\r\n}\r\n```\r\n\r\n## Error Handling\r\n\r\nThe library throws errors in the following cases:\r\n\r\n- Invalid PPTX file structure\r\n\r\n- File reading errors\r\n\r\n- XML parsing errors\r\n\r\nExample error handling:\r\n\r\n```typescript\r\ntry {\r\n  const parser = new PptxParser(\"presentation.ppt\");\r\n  const content = await parser.extractText();\r\n} catch (error) {\r\n  if (error.message.includes(\"Invalid PPTX file structure\")) {\r\n    console.error(\"The PPTX file is corrupted or invalid\");\r\n  } else {\r\n    console.error(\"An error occurred:\", error.message);\r\n  }\r\n}\r\n```\r\n\r\n## Dependencies\r\n\r\n- unzipper: For extracting PPTX files\r\n- xml2js: For parsing XML content\r\n\r\n## License\r\n\r\nMIT\r\n\r\n## Contributing\r\n\r\nContributions are welcome! Please feel free to submit a Pull Request.\r\n","readmeFilename":"README.MD"}