{"_id":"@alloc/dom-to-semantic-markdown","_rev":"2-d4f325877f1d4e555bc1436f4a213c88","name":"@alloc/dom-to-semantic-markdown","dist-tags":{"latest":"2.0.1"},"versions":{"1.2.16":{"name":"@alloc/dom-to-semantic-markdown","version":"1.2.16","keywords":[],"author":{"name":"Roman Landenband"},"license":"MIT","_id":"@alloc/dom-to-semantic-markdown@1.2.16","maintainers":[{"name":"aleclarson","email":"alec.stanford.larson@gmail.com"}],"homepage":"https://github.com/aleclarson/dom-to-semantic-markdown#readme","bugs":{"url":"https://github.com/aleclarson/dom-to-semantic-markdown/issues"},"dist":{"shasum":"16869a548a2c5da9f6640e757a40496b3fa5496e","tarball":"https://registry.npmjs.org/@alloc/dom-to-semantic-markdown/-/dom-to-semantic-markdown-1.2.16.tgz","fileCount":5,"integrity":"sha512-c2X+Xh/bUSCFejuFmNR8tCrV723+AqTJoY5BacseQ9vzDl4Vsqt4YzfhpxL5eZviMat0eUQR29/q5LzQzMBkSA==","signatures":[{"sig":"MEUCIQD0wltLxycFEdohq51JeP3+l+ukMmXksTpZfYJNk9fejAIgFVdx8d9XYCqxpiarDHCeuB/D2XFb4GnTiP1AFExh7VI=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":44004},"type":"module","_from":"file:alloc-dom-to-semantic-markdown-1.2.16.tgz","exports":{".":{"types":"./dist/index.d.ts","default":"./dist/index.js"}},"scripts":{"dev":"rimraf dist && tsup --watch","test":"jest","build":"tsup --clean --treeshake=smallest","format":"biome check --fix --unsafe","example":"open examples/browser.html"},"_npmUser":{"name":"aleclarson","email":"alec.stanford.larson@gmail.com"},"_resolved":"/private/var/folders/4l/lb2f5pj1221bm7d89gh08fx80000gn/T/d6fb56941db28d2872d59f92b07b8b60/alloc-dom-to-semantic-markdown-1.2.16.tgz","_integrity":"sha512-c2X+Xh/bUSCFejuFmNR8tCrV723+AqTJoY5BacseQ9vzDl4Vsqt4YzfhpxL5eZviMat0eUQR29/q5LzQzMBkSA==","repository":{"url":"git+https://github.com/aleclarson/dom-to-semantic-markdown.git","type":"git"},"_npmVersion":"10.9.2","description":"DOM to Semantic-Markdown for use in LLMs","directories":{},"_nodeVersion":"23.5.0","_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","tsup":"^8.3.6","jsdom":"^24.1.1","rimraf":"^5.0.10","ts-jest":"^29.2.4","typescript":"^5.6.3","@types/jest":"^29.5.12","@types/jsdom":"^21.1.7","@biomejs/biome":"^1.9.4","@tsconfig/node16":"^16.1.3","@radashi-org/biome-config":"^1.0.2"},"_npmOperationalInternal":{"tmp":"tmp/dom-to-semantic-markdown_1.2.16_1738776548577_0.36939894706593535","host":"s3://npm-registry-packages-npm-production"}},"2.0.0":{"name":"@alloc/dom-to-semantic-markdown","version":"2.0.0","keywords":[],"author":{"name":"Roman Landenband"},"license":"MIT","_id":"@alloc/dom-to-semantic-markdown@2.0.0","maintainers":[{"name":"aleclarson","email":"alec.stanford.larson@gmail.com"}],"homepage":"https://github.com/aleclarson/dom-to-semantic-markdown#readme","bugs":{"url":"https://github.com/aleclarson/dom-to-semantic-markdown/issues"},"dist":{"shasum":"2774a989c1be68ebe1b0629710ceac149adf4203","tarball":"https://registry.npmjs.org/@alloc/dom-to-semantic-markdown/-/dom-to-semantic-markdown-2.0.0.tgz","fileCount":5,"integrity":"sha512-9aqXFffoCn6paMWO7JGx3efX7bL21AuOQIx1HzCfBuMzw+Q/Hse263Aw1jFKY00j0HXR9qh25Jbg4Eol1maRpw==","signatures":[{"sig":"MEQCIBV2OidfWvWCAnSUgkWbOv43WU5UoFNLgsaR0ssbg7BaAiB16uCIdPl61eJQdZkY8ITkukK18OSzpC86VfgIiv5OaA==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":50035},"type":"module","_from":"file:alloc-dom-to-semantic-markdown-2.0.0.tgz","exports":{".":{"types":"./dist/index.d.ts","default":"./dist/index.js"}},"scripts":{"dev":"rimraf dist && tsup --watch --sourcemap","test":"jest","build":"tsup --clean --treeshake=smallest","format":"biome check --fix --unsafe","example":"open examples/browser.html"},"_npmUser":{"name":"aleclarson","email":"alec.stanford.larson@gmail.com"},"_resolved":"/private/var/folders/4l/lb2f5pj1221bm7d89gh08fx80000gn/T/bae0c43bf29fce54e49b331fef4b9e1b/alloc-dom-to-semantic-markdown-2.0.0.tgz","_integrity":"sha512-9aqXFffoCn6paMWO7JGx3efX7bL21AuOQIx1HzCfBuMzw+Q/Hse263Aw1jFKY00j0HXR9qh25Jbg4Eol1maRpw==","repository":{"url":"git+https://github.com/aleclarson/dom-to-semantic-markdown.git","type":"git"},"_npmVersion":"10.9.2","description":"DOM to Semantic-Markdown for use in LLMs","directories":{},"_nodeVersion":"23.5.0","_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","tsup":"^8.3.6","jsdom":"^24.1.1","rimraf":"^5.0.10","ts-jest":"^29.2.4","typescript":"^5.6.3","@types/jest":"^29.5.12","@types/jsdom":"^21.1.7","@biomejs/biome":"^1.9.4","@tsconfig/node16":"^16.1.3","@radashi-org/biome-config":"^1.0.2"},"_npmOperationalInternal":{"tmp":"tmp/dom-to-semantic-markdown_2.0.0_1738876468571_0.01735113384514153","host":"s3://npm-registry-packages-npm-production"}},"2.0.1":{"name":"@alloc/dom-to-semantic-markdown","version":"2.0.1","description":"DOM to Semantic-Markdown for use in LLMs","type":"module","repository":{"type":"git","url":"git+https://github.com/aleclarson/dom-to-semantic-markdown.git"},"exports":{".":{"types":"./dist/index.d.ts","default":"./dist/index.js"}},"author":{"name":"Roman Landenband"},"license":"MIT","devDependencies":{"@biomejs/biome":"^1.9.4","@radashi-org/biome-config":"^1.0.2","@tsconfig/node16":"^16.1.3","@types/jest":"^29.5.12","@types/jsdom":"^21.1.7","jest":"^29.7.0","jsdom":"^24.1.1","rimraf":"^5.0.10","ts-jest":"^29.2.4","tsup":"^8.3.6","typescript":"^5.6.3"},"keywords":[],"scripts":{"build":"tsup --clean --treeshake=smallest","dev":"rimraf dist && tsup --watch --sourcemap","format":"biome check --fix --unsafe","test":"jest","example":"open examples/browser.html"},"_id":"@alloc/dom-to-semantic-markdown@2.0.1","bugs":{"url":"https://github.com/aleclarson/dom-to-semantic-markdown/issues"},"homepage":"https://github.com/aleclarson/dom-to-semantic-markdown#readme","_integrity":"sha512-35U0d9dJA0w67iSG5bv5CamP8WilRG3ZprFJYQGbBi76CPvs4T+ufFA4URKtnd9u4yzRvlNco4tW3p/ompqM+g==","_resolved":"/private/var/folders/4l/lb2f5pj1221bm7d89gh08fx80000gn/T/660529a75c5c3ac5fa7a4309cfd544e6/alloc-dom-to-semantic-markdown-2.0.1.tgz","_from":"file:alloc-dom-to-semantic-markdown-2.0.1.tgz","_nodeVersion":"23.7.0","_npmVersion":"10.9.2","dist":{"integrity":"sha512-35U0d9dJA0w67iSG5bv5CamP8WilRG3ZprFJYQGbBi76CPvs4T+ufFA4URKtnd9u4yzRvlNco4tW3p/ompqM+g==","shasum":"0591ddcf0ed45bd3a2438714171df6162cc92ebd","tarball":"https://registry.npmjs.org/@alloc/dom-to-semantic-markdown/-/dom-to-semantic-markdown-2.0.1.tgz","fileCount":5,"unpackedSize":51012,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQCRcR/QKkyvOXfUg78WfzZdnf6SQW4ImapezMPJei66/wIgM+cV4W35TZNKfapnaT/nKYUELMagqZfebxGJ+MhXkJg="}]},"_npmUser":{"name":"aleclarson","email":"alec.stanford.larson@gmail.com"},"directories":{},"maintainers":[{"name":"aleclarson","email":"alec.stanford.larson@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/dom-to-semantic-markdown_2.0.1_1740927964957_0.9980428774709669"},"_hasShrinkwrap":false}},"time":{"created":"2025-02-05T17:29:08.482Z","modified":"2025-03-02T15:06:05.340Z","1.2.16":"2025-02-05T17:29:08.909Z","2.0.0":"2025-02-06T21:14:28.782Z","2.0.1":"2025-03-02T15:06:05.156Z"},"bugs":{"url":"https://github.com/aleclarson/dom-to-semantic-markdown/issues"},"author":{"name":"Roman Landenband"},"license":"MIT","homepage":"https://github.com/aleclarson/dom-to-semantic-markdown#readme","keywords":[],"repository":{"type":"git","url":"git+https://github.com/aleclarson/dom-to-semantic-markdown.git"},"description":"DOM to Semantic-Markdown for use in LLMs","maintainers":[{"name":"aleclarson","email":"alec.stanford.larson@gmail.com"}],"readme":"# @alloc/dom-to-semantic-markdown\n\nThis library converts HTML DOM to a semantic Markdown format optimized for use with Large Language Models (LLMs). It preserves the semantic structure of web content, extracts essential metadata, and reduces token usage compared to raw HTML, making it easier for LLMs to understand and process information.\n\n**Note:** This is a personal fork of [romansky/dom-to-semantic-markdown](https://github.com/romansky/dom-to-semantic-markdown). Support will not be provided.\n\n## Key Features\n\n- **Semantic Structure Preservation:** Retains the meaning of HTML elements like `<header>`, `<footer>`, `<nav>`, and more.\n- **Metadata Extraction:** Captures important metadata such as title, description, keywords, Open Graph tags, Twitter Card tags, and JSON-LD data.\n- **Token Efficiency:** Optimizes for token usage through URL refification and concise representation of content.\n- **Main Content Detection:** Automatically identifies and extracts the primary content section of a webpage.\n- **Table Column Tracking:** Adds unique identifiers to table columns, improving LLM's ability to correlate data across rows.\n\n## Installation\n\n```bash\npnpm add @alloc/dom-to-semantic-markdown\n```\n\n## Usage\n\n```javascript\nimport { convertHtmlToMarkdown } from \"@alloc/dom-to-semantic-markdown\";\n\nconst markdown = convertHtmlToMarkdown(document.body);\nconsole.log(markdown);\n```\n\n## Functions\n\n### `convertHtmlToMarkdown(html: string, options?: ConversionOptions): string`\n\nConverts an HTML string to semantic Markdown.\n\n- `html: string`: The HTML string to be converted.\n- `options?: ConversionOptions`: Optional configuration object to customize the conversion process. See [ConversionOptions](#ConversionOptions) for available settings.\n\n**Returns:** `string` - The Markdown string representation of the HTML content.\n\n### `convertElementToMarkdown(element: Element, options?: ConversionOptions): string`\n\nConverts an HTML Element to semantic Markdown.\n\n- `element: Element`: The HTML DOM Element to be converted. This allows you to convert specific parts of a document, not just the entire HTML string.\n- `options?: ConversionOptions`: Optional configuration object to customize the conversion process. See [ConversionOptions](#ConversionOptions) for available settings.\n\n**Returns:** `string` - The Markdown string representation of the provided HTML Element and its descendants.\n\n### `extractMetaData(element: Element, mode?: 'basic' | 'extended'): SemanticMarkdownAST.MetaDataNode['content']`\n\nExtracts metadata from an HTML Element.\n\n- `element: Element`: The HTML DOM Element to extract metadata from.\n- `mode?: 'basic' | 'extended'`: Optional mode to control the level of metadata extraction.\n  - `'basic'`: Includes standard meta tags like title, description, and keywords.\n  - `'extended'`: Includes basic meta tags, Open Graph tags, Twitter Card tags, and JSON-LD data.\n\n**Returns:** `SemanticMarkdownAST.MetaDataNode['content']` - An object containing the extracted metadata.\n\n### `htmlToMarkdownAST(element: Element, options?: ExtractOptions, indentLevel?: number): SemanticMarkdownAST.Node[]`\n\nConverts an HTML `Element` into a Semantic Markdown Abstract Syntax Tree (AST). This function recursively parses the HTML structure and generates a structured Markdown representation. It uses `extractMetaData` to extract metadata from the `<head>` element.\n\n- `element: Element`: The HTML DOM element to be converted.\n- `options?: ExtractOptions`: Optional configuration to customize the extraction process. See `ExtractOptions` for details.\n- `indentLevel?: number`: The current indentation level, used for nested elements like lists. Defaults to `0`.\n\n**Returns:** `SemanticMarkdownAST.Node[]` - An array of AST nodes representing the semantic Markdown structure of the input HTML element. This AST can then be rendered into a Markdown string using a separate rendering function.\n\nThis function is not intended for direct use in most cases. Use `convertHtmlToMarkdown` or `convertElementToMarkdown` for simpler HTML to Markdown conversion. However, understanding `htmlToMarkdownAST` is crucial for customizing or extending the library's functionality.\n\n### `markdownASTToString(nodes: Node[], options?: RenderOptions, indentLevel?: number): string`\n\nConverts a Semantic Markdown Abstract Syntax Tree (AST) back into a Markdown string. This function takes the AST generated by `htmlToMarkdownAST` and renders it into a human-readable Markdown format.\n\n- `nodes: Node[]`: An array of SemanticMarkdownAST nodes representing the Markdown content. This is typically the output of the `htmlToMarkdownAST` function.\n- `options?: RenderOptions`: Optional configuration object to customize the rendering process. See [RenderOptions](#RenderOptions) for available settings.\n- `indentLevel?: number`: The initial indentation level for the Markdown output. Used for nested structures like lists and blockquotes. Defaults to `0`.\n\n**Returns:** `string` - The Markdown string representation of the AST.\n\nThis function is essential for completing the HTML to Markdown conversion process. It takes the structured AST and transforms it into a flat, string-based Markdown output.\n\n## Types\n\n### `ExtractOptions`\n\n- `debug?: boolean`: Enable debug logging.\n- `websiteDomain?: string`: The domain of the website being converted.\n- `extractMainContent?: boolean`: Whether to extract only the main content of the page.\n- `includeMetaData?: 'basic' | 'extended' | false`: Controls whether to include metadata extracted from the HTML head.\n  - `'basic'`: Includes standard meta tags like title, description, and keywords.\n  - `'extended'`: Includes basic meta tags, Open Graph tags, Twitter Card tags, and JSON-LD data.\n  - `false`: Disables metadata extraction.\n- `excludeTagNames?: string[]`: Avoid extracting content from these tags.\n- `excludeInvisibleElements?: boolean`: Whether to exclude elements that are not visible.\n- `enableTableColumnTracking?: boolean`: Adds unique identifiers to table columns.\n- `overrideElementProcessing?: (element: Element, options: ConversionOptions, indentLevel: number) => SemanticMarkdownAST[] | undefined`: Custom processing for HTML elements.\n- `processUnhandledElement?: (element: Element, options: ConversionOptions, indentLevel: number) => SemanticMarkdownAST[] | undefined`: Handler for unknown HTML elements.\n\n### `RenderOptions`\n\n- `emitFrontMatter?: boolean`: Include the metadata as “front matter” in the output.\n- `overrideNodeRenderer?: (node: SemanticMarkdownAST, options: ConversionOptions, indentLevel: number) => string | undefined`: Custom renderer for AST nodes.\n- `renderCustomNode?: (node: CustomNode, options: ConversionOptions, indentLevel: number) => string | undefined`: Renderer for custom AST nodes.\n\n### `ConversionOptions`\n\n- `refifyUrls?: boolean`: Whether to convert URLs to reference-style links.\n- `overrideDOMParser?: DOMParser`: Custom DOMParser for Node.js environments.\n- _Everything in `ExtractOptions` and `RenderOptions`_\n\n### `SemanticMarkdownAST`\n\n`SemanticMarkdownAST` is a type-only namespace that defines the structure of the Markdown Abstract Syntax Tree (AST) used by this library. It encompasses various node types that represent different semantic elements in Markdown, allowing for a structured and programmatically accessible representation of Markdown content.\n\nThe namespace includes the following type definitions for different Markdown elements:\n\n- **`BlockquoteNode`**: Represents blockquotes.\n- **`BoldNode`**: Represents bold text.\n- **`CodeNode`**: Represents code blocks and inline code.\n- **`CustomNode`**: Represents custom, user-defined nodes.\n- **`HeadingNode`**: Represents headings with levels from 1 to 6.\n- **`ImageNode`**: Represents images.\n- **`ItalicNode`**: Represents italic text.\n- **`LinkNode`**: Represents hyperlinks.\n- **`ListItemNode`**: Represents items in a list.\n- **`ListNode`**: Represents ordered and unordered lists.\n- **`MetaDataNode`**: Represents metadata extracted from HTML `<head>`, including standard meta tags, Open Graph, Twitter Card, and JSON-LD.\n- **`SemanticHtmlNode`**: Represents semantic HTML elements like `<article>`, `<header>`, etc.\n- **`StrikethroughNode`**: Represents strikethrough text.\n- **`TableCellNode`**: Represents cells within a table.\n- **`TableNode`**: Represents tables.\n- **`TableRowNode`**: Represents rows within a table.\n- **`TextNode`**: Represents plain text content.\n- **`VideoNode`**: Represents video embeds.\n\nEach of these node types defines a specific structure with properties relevant to the represented Markdown element, such as `content`, `level` (for headings), `href` (for links), etc. These types are used throughout the library to represent and manipulate Markdown content programmatically.\n\n## Using the Output with LLMs\n\nThe semantic Markdown produced by this library is optimized for use with Large Language Models (LLMs). To use it effectively:\n\n1.  Extract the Markdown content using the library.\n2.  Start with a brief instruction or context for the LLM.\n3.  Wrap the extracted Markdown in triple backticks (`` `).\n4.  Follow the Markdown with your question or prompt.\n\nExample:\n\n````\nThe following is a semantic Markdown representation of a webpage. Please analyze its content:\n\n```markdown\n{paste your extracted markdown here}\n```\n\n{your question, e.g., \"What are the main points discussed in this article?\"}\n````\n\nThis format helps the LLM understand its task and the context of the content, enabling more accurate and relevant responses to your questions.\n","readmeFilename":"README.md"}