{"_id":"@bvolpato/copy-as-markdown","_rev":"4-2b3c959eb7810782fef8fe0b574ad9c4","name":"@bvolpato/copy-as-markdown","dist-tags":{"latest":"1.5.1"},"versions":{"1.3.3":{"name":"@bvolpato/copy-as-markdown","version":"1.3.3","keywords":["markdown","copy","clipboard","library","extractor","userscript","browser-extension","chrome-extension","firefox-extension","llm","chatgpt","claude","gemini","web-scraping","content-extraction"],"author":"Bruno Volpato <brunocvcunha@gmail.com>","license":"MIT","_id":"@bvolpato/copy-as-markdown@1.3.3","maintainers":[{"name":"bvolpato","email":"brunocvcunha@gmail.com"}],"homepage":"https://github.com/bvolpato/copy-as-markdown#readme","bugs":{"url":"https://github.com/bvolpato/copy-as-markdown/issues"},"dist":{"shasum":"bfe892bf2e349c0c395cc053bb9040322a3dd33d","tarball":"https://registry.npmjs.org/@bvolpato/copy-as-markdown/-/copy-as-markdown-1.3.3.tgz","fileCount":159,"integrity":"sha512-IbO150h6EKGhHmQpwa58xEnhwk40tFWoeGtRf/qW9qmSBMCwNucyT4PtptSxE1matiWEVLj8hDRady8dSlMPLg==","signatures":[{"sig":"MEUCIDXxWMQIgBQW12fAJGPbul2Lu5JRf5HjqarDNTga4JEYAiEA9zXb9G5SVRN/xjJRWlrXeJEXL8o0xLGh3zLzQnVX7hU=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":1478257},"main":"./dist/library/index.js","type":"module","types":"./dist/library/types/library/entry.d.ts","unpkg":"./dist/library/browser.js","module":"./dist/library/index.js","engines":{"node":"22.22.2","pnpm":"11.22.0"},"exports":{".":{"types":"./dist/library/types/library/entry.d.ts","import":"./dist/library/index.js"},"./core":{"types":"./dist/library/types/library/index.d.ts","import":"./dist/library/core.js"},"./package.json":"./package.json","./extractors/jira":{"types":"./dist/library/types/extractors/jira.d.ts","import":"./dist/library/extractors/jira.js"},"./extractors/github":{"types":"./dist/library/types/extractors/github.d.ts","import":"./dist/library/extractors/github.js"},"./extractors/confluence":{"types":"./dist/library/types/extractors/confluence.d.ts","import":"./dist/library/extractors/confluence.js"}},"scripts":{"test":"pnpm test:regression","build":"tsx build/build.ts && tsc --project tsconfig.library.json","clean":"rm -rf dist","rebuild":"pnpm clean && pnpm build","test:live":"pnpm build && node test/test-sites.js","test:site":"pnpm build && node test/test-sites.js","typecheck":"tsc --noEmit","zip:chrome":"cd dist/chrome && zip -r ../copy-as-markdown-chrome.zip .","package:all":"pnpm build && pnpm zip:chrome && pnpm zip:firefox","zip:firefox":"cd dist/firefox && zip -r ../copy-as-markdown-firefox.zip .","pack:library":"pnpm pack --dry-run","test:library":"pnpm build && node test/library.test.js","build:library":"tsx build/build.ts --library-only && tsc --project tsconfig.library.json","package:chrome":"pnpm build && pnpm zip:chrome","verify:release":"tsx build/verify-release.ts","fixtures:verify":"pnpm build && tsx test/fixtures/verify.ts","package:firefox":"pnpm build && pnpm zip:firefox","test:regression":"pnpm build && node test/library.test.js && node test/background.test.js && node test/test-sites.js --regression && tsx test/fixtures/security.test.ts && tsx test/fixtures/verify.ts","fixtures:capture":"pnpm build && tsx test/fixtures/capture.ts"},"_npmUser":{"name":"bvolpato","email":"brunocvcunha@gmail.com"},"jsdelivr":"./dist/library/browser.js","repository":{"url":"https://github.com/bvolpato/copy-as-markdown","type":"git"},"description":"Copy page content as clean, structured Markdown","directories":{},"_nodeVersion":"22.22.2","publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"devDependencies":{"tsx":"^4.23.1","yaml":"^2.9.0","esbuild":"^0.28.1","puppeteer":"^25.5.0","typescript":"^5.9.3","@types/node":"^26.1.2","@types/chrome":"^0.1.43"},"_npmOperationalInternal":{"tmp":"tmp/copy-as-markdown_1.3.3_1787151065107_0.980048936114702","host":"s3://npm-registry-packages-npm-production"}},"1.3.4":{"name":"@bvolpato/copy-as-markdown","version":"1.3.4","keywords":["markdown","copy","clipboard","library","extractor","userscript","browser-extension","chrome-extension","firefox-extension","llm","chatgpt","claude","gemini","web-scraping","content-extraction"],"author":"Bruno Volpato <brunocvcunha@gmail.com>","license":"MIT","_id":"@bvolpato/copy-as-markdown@1.3.4","maintainers":[{"name":"bvolpato","email":"brunocvcunha@gmail.com"}],"homepage":"https://github.com/bvolpato/copy-as-markdown#readme","bugs":{"url":"https://github.com/bvolpato/copy-as-markdown/issues"},"dist":{"shasum":"02f242647fbf1521186098c44c567603f2e8fcab","tarball":"https://registry.npmjs.org/@bvolpato/copy-as-markdown/-/copy-as-markdown-1.3.4.tgz","fileCount":159,"integrity":"sha512-Xugq0TOlOvAdIVDxbCHc/hD5/nZ4rZj80BWLwdlsw5Qtlqpf5Uk++aueA+q97t3y1ywjgsSEN/ZuV410Hh7HpA==","signatures":[{"sig":"MEQCIF5eZgT4l8CbJP6xPASS6fL3+tra5tbYBslBLlBDieM9AiBMIUJM/JkDZ8RFBTuWhYCAswYpyujeKPOZGEwc4wdE7w==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":1478859},"main":"./dist/library/index.js","type":"module","types":"./dist/library/types/library/entry.d.ts","unpkg":"./dist/library/browser.js","module":"./dist/library/index.js","engines":{"node":"22.22.2","pnpm":"11.22.0"},"exports":{".":{"types":"./dist/library/types/library/entry.d.ts","import":"./dist/library/index.js"},"./core":{"types":"./dist/library/types/library/index.d.ts","import":"./dist/library/core.js"},"./package.json":"./package.json","./extractors/jira":{"types":"./dist/library/types/extractors/jira.d.ts","import":"./dist/library/extractors/jira.js"},"./extractors/github":{"types":"./dist/library/types/extractors/github.d.ts","import":"./dist/library/extractors/github.js"},"./extractors/confluence":{"types":"./dist/library/types/extractors/confluence.d.ts","import":"./dist/library/extractors/confluence.js"},"./extractors/google-docs":{"types":"./dist/library/types/extractors/google-docs.d.ts","import":"./dist/library/extractors/google-docs.js"}},"scripts":{"test":"pnpm test:regression","build":"tsx build/build.ts && tsc --project tsconfig.library.json","clean":"rm -rf dist","rebuild":"pnpm clean && pnpm build","test:live":"pnpm build && node test/test-sites.js","test:site":"pnpm build && node test/test-sites.js","typecheck":"tsc --noEmit","zip:chrome":"cd dist/chrome && zip -r ../copy-as-markdown-chrome.zip .","package:all":"pnpm build && pnpm zip:chrome && pnpm zip:firefox","zip:firefox":"cd dist/firefox && zip -r ../copy-as-markdown-firefox.zip .","pack:library":"pnpm pack --dry-run","test:library":"pnpm build && node test/library.test.js","build:library":"tsx build/build.ts --library-only && tsc --project tsconfig.library.json","package:chrome":"pnpm build && pnpm zip:chrome","verify:release":"tsx build/verify-release.ts","fixtures:verify":"pnpm build && tsx test/fixtures/verify.ts","package:firefox":"pnpm build && pnpm zip:firefox","test:regression":"pnpm build && node test/library.test.js && node test/background.test.js && node test/test-sites.js --regression && tsx test/fixtures/security.test.ts && tsx test/fixtures/verify.ts","fixtures:capture":"pnpm build && tsx test/fixtures/capture.ts"},"_npmUser":{"name":"bvolpato","email":"brunocvcunha@gmail.com"},"jsdelivr":"./dist/library/browser.js","repository":{"url":"https://github.com/bvolpato/copy-as-markdown","type":"git"},"description":"Copy page content as clean, structured Markdown","directories":{},"_nodeVersion":"22.22.2","publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"devDependencies":{"tsx":"^4.23.1","yaml":"^2.9.0","esbuild":"^0.28.1","puppeteer":"^25.5.0","typescript":"^5.9.3","@types/node":"^26.1.2","@types/chrome":"^0.1.43"},"_npmOperationalInternal":{"tmp":"tmp/copy-as-markdown_1.3.4_1787153465370_0.6411089655685582","host":"s3://npm-registry-packages-npm-production"}},"1.4.0":{"name":"@bvolpato/copy-as-markdown","version":"1.4.0","keywords":["markdown","copy","clipboard","library","extractor","userscript","browser-extension","chrome-extension","firefox-extension","llm","chatgpt","claude","gemini","web-scraping","content-extraction"],"author":{"name":"Bruno Volpato","email":"brunocvcunha@gmail.com"},"license":"MIT","_id":"@bvolpato/copy-as-markdown@1.4.0","maintainers":[{"name":"bvolpato","email":"brunocvcunha@gmail.com"}],"homepage":"https://github.com/bvolpato/copy-as-markdown#readme","bugs":{"url":"https://github.com/bvolpato/copy-as-markdown/issues"},"dist":{"shasum":"8db3c8cca52894dc4f9c043f5fb6403c321f3e8c","tarball":"https://registry.npmjs.org/@bvolpato/copy-as-markdown/-/copy-as-markdown-1.4.0.tgz","fileCount":177,"integrity":"sha512-pV6D5Fz0l7QWkZAoVNqR0xSUCpb/Vov2ZEg0SSSb96pRRJEKrcEwieVZoNDQmGFurWpS0sX/PGaMsz/sVkJmbQ==","signatures":[{"sig":"MEUCIBqj2M9KqOMzygjwMPS88+7N6Of84w1up03x8MyxhttjAiEAtACWovesybY6AiYpBUVZaKk6AUBDF2YLkcYNXESp80I=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@bvolpato%2fcopy-as-markdown@1.4.0","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":1771347},"main":"./dist/library/index.js","type":"module","_from":"file:/home/runner/work/copy-as-markdown/copy-as-markdown/npm-package/bvolpato-copy-as-markdown-1.4.0.tgz","types":"./dist/library/types/library/entry.d.ts","unpkg":"./dist/library/browser.js","module":"./dist/library/index.js","engines":{"node":"22.22.2","pnpm":"11.22.0"},"exports":{".":{"types":"./dist/library/types/library/entry.d.ts","import":"./dist/library/index.js"},"./core":{"types":"./dist/library/types/library/index.d.ts","import":"./dist/library/core.js"},"./package.json":"./package.json","./extractors/jira":{"types":"./dist/library/types/extractors/jira.d.ts","import":"./dist/library/extractors/jira.js"},"./extractors/github":{"types":"./dist/library/types/extractors/github.d.ts","import":"./dist/library/extractors/github.js"},"./extractors/linear":{"types":"./dist/library/types/extractors/linear.d.ts","import":"./dist/library/extractors/linear.js"},"./extractors/deepseek":{"types":"./dist/library/types/extractors/deepseek.d.ts","import":"./dist/library/extractors/deepseek.js"},"./extractors/confluence":{"types":"./dist/library/types/extractors/confluence.d.ts","import":"./dist/library/extractors/confluence.js"},"./extractors/google-docs":{"types":"./dist/library/types/extractors/google-docs.d.ts","import":"./dist/library/extractors/google-docs.js"},"./extractors/hugging-face":{"types":"./dist/library/types/extractors/hugging-face.d.ts","import":"./dist/library/extractors/hugging-face.js"},"./extractors/mistral-vibe":{"types":"./dist/library/types/extractors/mistral-vibe.d.ts","import":"./dist/library/extractors/mistral-vibe.js"},"./extractors/documentation":{"types":"./dist/library/types/extractors/documentation.d.ts","import":"./dist/library/extractors/documentation.js"},"./extractors/gemini-notebook":{"types":"./dist/library/types/extractors/gemini-notebook.d.ts","import":"./dist/library/extractors/gemini-notebook.js"},"./extractors/microsoft-teams":{"types":"./dist/library/types/extractors/microsoft-teams.d.ts","import":"./dist/library/extractors/microsoft-teams.js"},"./extractors/google-ai-studio":{"types":"./dist/library/types/extractors/google-ai-studio.d.ts","import":"./dist/library/extractors/google-ai-studio.js"},"./extractors/microsoft-copilot":{"types":"./dist/library/types/extractors/microsoft-copilot.d.ts","import":"./dist/library/extractors/microsoft-copilot.js"}},"scripts":{"test":"pnpm test:regression","build":"tsx build/build.ts && tsc --project tsconfig.library.json","clean":"rm -rf dist","rebuild":"pnpm clean && pnpm build","test:live":"pnpm build && node test/test-sites.js","test:site":"pnpm build && node test/test-sites.js","typecheck":"tsc --noEmit","zip:chrome":"cd dist/chrome && zip -r ../copy-as-markdown-chrome.zip .","package:all":"pnpm build && pnpm zip:chrome && pnpm zip:firefox","zip:firefox":"cd dist/firefox && zip -r ../copy-as-markdown-firefox.zip .","pack:library":"pnpm pack --dry-run","test:ai-chat":"pnpm build && node test/ai-chat.test.js","test:library":"pnpm build && node test/library.test.js","build:library":"tsx build/build.ts --library-only && tsc --project tsconfig.library.json","package:chrome":"pnpm build && pnpm zip:chrome","verify:release":"tsx build/verify-release.ts","fixtures:verify":"pnpm build && tsx test/fixtures/verify.ts","package:firefox":"pnpm build && pnpm zip:firefox","test:regression":"pnpm build && node test/library.test.js && node test/ai-chat.test.js && node test/background.test.js && node test/test-sites.js --regression && tsx test/fixtures/security.test.ts && tsx test/fixtures/verify.ts","fixtures:capture":"pnpm build && tsx test/fixtures/capture.ts"},"_npmUser":{"name":"GitHub Actions","email":"npm-oidc-no-reply@github.com","trustedPublisher":{"id":"github","oidcConfigId":"oidc:0cf0761f-369a-4f64-b30c-398f2037f05f"}},"jsdelivr":"./dist/library/browser.js","_resolved":"/home/runner/work/copy-as-markdown/copy-as-markdown/npm-package/bvolpato-copy-as-markdown-1.4.0.tgz","_integrity":"sha512-pV6D5Fz0l7QWkZAoVNqR0xSUCpb/Vov2ZEg0SSSb96pRRJEKrcEwieVZoNDQmGFurWpS0sX/PGaMsz/sVkJmbQ==","repository":{"url":"git+https://github.com/bvolpato/copy-as-markdown.git","type":"git"},"_npmVersion":"12.0.2","description":"Copy page content as clean, structured Markdown","directories":{},"_nodeVersion":"22.22.2","publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"devDependencies":{"tsx":"^4.23.12","yaml":"^2.9.0","esbuild":"^0.28.2","puppeteer":"^25.7.0","typescript":"^7.0.2","@types/node":"^26.2.0","@types/chrome":"^0.2.6"},"_npmOperationalInternal":{"tmp":"tmp/copy-as-markdown_1.4.0_1788104834462_0.7091605488301329","host":"s3://npm-registry-packages-npm-production"}},"1.5.1":{"_id":"@bvolpato/copy-as-markdown@1.5.1","bugs":{"url":"https://github.com/bvolpato/copy-as-markdown/issues"},"dist":{"shasum":"b53e5905d8426a42ce8c81c232001098ec5eba63","tarball":"https://registry.npmjs.org/@bvolpato/copy-as-markdown/-/copy-as-markdown-1.5.1.tgz","fileCount":177,"integrity":"sha512-hKwvnK2hkn5Kv6jG5/+oeJuH3mJ3wOxNMR5GAtxcZUKvKsQRGVDbWB0upTgIQz4rlGP5n9r0DTdysm7eaV+oFA==","signatures":[{"sig":"MEYCIQC+ULGjXIvI05h26je4p7aXN87A8JbceKnuQOKjmhjueQIhAIAzQnKUSZTDXLQk5DCc8rC03o/6WARDDrecixV4nJ5m","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"},{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIGWs8J3glzBo4CWvClfGskPpiwERS8QxujpOKsuuWe7JAiArX8IxU+IefqeJHoy6xi1VSyuWUzPYOostpzq41w1Lwg=="}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@bvolpato%2fcopy-as-markdown@1.5.1","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":1804909},"main":"./dist/library/index.js","name":"@bvolpato/copy-as-markdown","type":"module","_from":"file:/home/runner/work/copy-as-markdown/copy-as-markdown/npm-package/bvolpato-copy-as-markdown-1.5.1.tgz","types":"./dist/library/types/library/entry.d.ts","unpkg":"./dist/library/browser.js","author":{"name":"Bruno Volpato","email":"brunocvcunha@gmail.com"},"module":"./dist/library/index.js","engines":{"node":">=22.22.2"},"exports":{".":{"types":"./dist/library/types/library/entry.d.ts","import":"./dist/library/index.js"},"./core":{"types":"./dist/library/types/library/index.d.ts","import":"./dist/library/core.js"},"./package.json":"./package.json","./extractors/jira":{"types":"./dist/library/types/extractors/jira.d.ts","import":"./dist/library/extractors/jira.js"},"./extractors/github":{"types":"./dist/library/types/extractors/github.d.ts","import":"./dist/library/extractors/github.js"},"./extractors/linear":{"types":"./dist/library/types/extractors/linear.d.ts","import":"./dist/library/extractors/linear.js"},"./extractors/deepseek":{"types":"./dist/library/types/extractors/deepseek.d.ts","import":"./dist/library/extractors/deepseek.js"},"./extractors/confluence":{"types":"./dist/library/types/extractors/confluence.d.ts","import":"./dist/library/extractors/confluence.js"},"./extractors/google-docs":{"types":"./dist/library/types/extractors/google-docs.d.ts","import":"./dist/library/extractors/google-docs.js"},"./extractors/hugging-face":{"types":"./dist/library/types/extractors/hugging-face.d.ts","import":"./dist/library/extractors/hugging-face.js"},"./extractors/mistral-vibe":{"types":"./dist/library/types/extractors/mistral-vibe.d.ts","import":"./dist/library/extractors/mistral-vibe.js"},"./extractors/documentation":{"types":"./dist/library/types/extractors/documentation.d.ts","import":"./dist/library/extractors/documentation.js"},"./extractors/gemini-notebook":{"types":"./dist/library/types/extractors/gemini-notebook.d.ts","import":"./dist/library/extractors/gemini-notebook.js"},"./extractors/microsoft-teams":{"types":"./dist/library/types/extractors/microsoft-teams.d.ts","import":"./dist/library/extractors/microsoft-teams.js"},"./extractors/google-ai-studio":{"types":"./dist/library/types/extractors/google-ai-studio.d.ts","import":"./dist/library/extractors/google-ai-studio.js"},"./extractors/microsoft-copilot":{"types":"./dist/library/types/extractors/microsoft-copilot.d.ts","import":"./dist/library/extractors/microsoft-copilot.js"}},"license":"MIT","scripts":{"test":"pnpm test:regression","build":"tsx build/build.ts && tsc --project tsconfig.library.json","clean":"rm -rf dist","rebuild":"pnpm clean && pnpm build","test:live":"pnpm build && node test/test-sites.js","test:site":"pnpm build && node test/test-sites.js","typecheck":"tsc --noEmit","zip:chrome":"cd dist/chrome && zip -r ../copy-as-markdown-chrome.zip .","package:all":"pnpm build && pnpm zip:chrome && pnpm zip:firefox","zip:firefox":"cd dist/firefox && zip -r ../copy-as-markdown-firefox.zip .","pack:library":"pnpm pack --dry-run","test:ai-chat":"pnpm build && node test/ai-chat.test.js","test:library":"pnpm build && node test/library.test.js","build:library":"tsx build/build.ts --library-only && tsc --project tsconfig.library.json","package:chrome":"pnpm build && pnpm zip:chrome","verify:release":"tsx build/verify-release.ts","fixtures:verify":"pnpm build && tsx test/fixtures/verify.ts","package:firefox":"pnpm build && pnpm zip:firefox","test:regression":"pnpm build && node test/facebook-detector.test.js && tsx --test test/atlassian.test.ts && node test/library.test.js && node test/catalog.test.js && node test/ai-chat.test.js && node test/background.test.js && node test/test-sites.js --regression && tsx test/fixtures/security.test.ts && tsx test/fixtures/verify.ts","fixtures:capture":"pnpm build && tsx test/fixtures/capture.ts"},"version":"1.5.1","_npmUser":{"name":"GitHub Actions","email":"npm-oidc-no-reply@github.com","trustedPublisher":{"id":"github","oidcConfigId":"0cf0761f-369a-4f64-b30c-398f2037f05f"}},"homepage":"https://github.com/bvolpato/copy-as-markdown#readme","jsdelivr":"./dist/library/browser.js","keywords":["markdown","copy","clipboard","library","extractor","userscript","browser-extension","chrome-extension","firefox-extension","llm","chatgpt","claude","gemini","web-scraping","content-extraction"],"_resolved":"/home/runner/work/copy-as-markdown/copy-as-markdown/npm-package/bvolpato-copy-as-markdown-1.5.1.tgz","_integrity":"sha512-hKwvnK2hkn5Kv6jG5/+oeJuH3mJ3wOxNMR5GAtxcZUKvKsQRGVDbWB0upTgIQz4rlGP5n9r0DTdysm7eaV+oFA==","repository":{"url":"git+https://github.com/bvolpato/copy-as-markdown.git","type":"git"},"_npmVersion":"12.0.2","description":"Copy page content as clean, structured Markdown","directories":{},"maintainers":[{"name":"bvolpato","email":"brunocvcunha@gmail.com"}],"_nodeVersion":"22.22.2","publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"_hasShrinkwrap":false,"devDependencies":{"tsx":"^4.23.12","yaml":"^2.9.0","esbuild":"^0.28.2","puppeteer":"^25.7.0","typescript":"^7.0.2","@types/node":"^26.2.0","@types/chrome":"^0.3.0"},"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/copy-as-markdown_1.5.1_1791065245165_0.0763667584206229"}}},"time":{"created":"2026-08-19T14:51:04.891Z","modified":"2026-10-03T22:07:25.598Z","1.3.3":"2026-08-19T14:51:05.272Z","1.3.4":"2026-08-19T15:31:05.528Z","1.4.0":"2026-08-30T15:47:14.609Z","1.5.1":"2026-10-03T22:07:25.286Z"},"bugs":{"url":"https://github.com/bvolpato/copy-as-markdown/issues"},"author":{"name":"Bruno Volpato","email":"brunocvcunha@gmail.com"},"license":"MIT","homepage":"https://github.com/bvolpato/copy-as-markdown#readme","keywords":["markdown","copy","clipboard","library","extractor","userscript","browser-extension","chrome-extension","firefox-extension","llm","chatgpt","claude","gemini","web-scraping","content-extraction"],"repository":{"url":"git+https://github.com/bvolpato/copy-as-markdown.git","type":"git"},"description":"Copy page content as clean, structured Markdown","maintainers":[{"name":"bvolpato","email":"brunocvcunha@gmail.com"}],"readme":"<p align=\"center\">\n  <img src=\"assets/icon.svg\" width=\"120\" alt=\"Copy as Markdown\" />\n</p>\n\n<h1 align=\"center\">Copy as Markdown</h1>\n\n<p align=\"center\">\n  <strong>One click. Clean Markdown. Perfect LLM context.</strong>\n</p>\n\n<p align=\"center\">\n  <a href=\"#install\">Install</a> ·\n  <a href=\"#library\">Library</a> ·\n  <a href=\"#supported-sites\">Supported Sites</a> ·\n  <a href=\"#why\">Why?</a> ·\n  <a href=\"#build\">Build</a> ·\n  <a href=\"#contributing\">Contributing</a>\n</p>\n\n---\n\n## The Problem\n\nYou're chatting with ChatGPT, Claude, or Gemini. You want to share a web page for context — a Wikipedia article, a Reddit thread, a YouTube video, a news story. What do you do?\n\n- **Copy-paste raw text** → Loses all structure. Headers become blobs. Tables vanish. Links disappear.\n- **Share a URL** → The LLM can't browse it (or hallucinates what it says).\n- **Screenshot** → Eats your token budget on image processing. Can't search or quote.\n- **Manually reformat** → Life's too short.\n\n## The Solution\n\n**Copy as Markdown** allows you to extract content with a single click. Depending on your installation method, you interact with it in two ways:\n\n- **Browser Extension:** Click the \"Copy as Markdown\" icon in your browser toolbar. Datadog dashboards, Datadog notebooks, and W&B runs also get a page button through narrowly scoped site access; extraction and clipboard writes still run only after a click.\n- **Userscript:** A context-aware button is added to supported websites (e.g., a draggable floating button, or inline buttons on Wikipedia, Google Docs, Atlassian, and Datadog pages). Floating positions persist per site.\n\nOne click, and the page's content lands in your clipboard as structured Markdown, with headers, tables, links, code blocks, and metadata. Prose normalizes Unicode compatibility forms, non-ASCII spaces, smart quotes, and dashes, and removes invisible watermark and direction-control characters. Code retains its literal text and uses delimiters that cannot be closed by backticks inside the sample. Paste it into your LLM conversation.\n\n> 💡 **Structured Markdown is the most token-efficient, context-rich format for sharing web content with LLMs.** It preserves semantic meaning (headers = hierarchy, tables = data, links = sources) while stripping visual noise.\n\n---\n\n## Install\n\n### <img src=\"assets/icon.svg\" width=\"20\" alt=\"\"> Userscript (Tampermonkey / Violentmonkey)\n\nThe fastest way to get started. Works in any browser with a userscript manager.\n\n1. Install [Tampermonkey](https://www.tampermonkey.net/) or [Violentmonkey](https://violentmonkey.github.io/)\n2. **[Install Copy as Markdown](https://github.com/bvolpato/copy-as-markdown/releases/latest/download/copy-as-markdown.user.js)**\n3. That's it — you'll see the button on supported sites\n\n### <img src=\"docs/brands/chrome.svg\" width=\"20\" alt=\"\"> Chrome Extension\n\n[Install Copy as Markdown from the Chrome Web Store](https://chromewebstore.google.com/detail/copy-as-markdown/pcjanmkidppaeojkanbjbmmgpjfeecol).\n\nFor local development:\n\n1. Download or clone this repository\n2. Run `pnpm install && pnpm build`\n3. Open `chrome://extensions/` → enable **Developer mode**\n4. Click **Load unpacked** → select the `dist/chrome/` folder\n\nChrome may show **Wants access to this site** on first use. Click **Allow**, then click **Copy as Markdown** again. The toolbar badge shows `…` while copying, `✓` on success, or `!` when the page or clipboard is unavailable.\n\n### <img src=\"docs/brands/firefox.svg\" width=\"20\" alt=\"\"> Firefox Extension\n\n[Install Copy as Markdown from Firefox Add-ons](https://addons.mozilla.org/en-US/firefox/addon/copy-as-markdown-addon/).\n\nFor local development:\n\n1. Download or clone this repository\n2. Run `pnpm install && pnpm build`\n3. Open `about:debugging#/runtime/this-firefox`\n4. Click **Load Temporary Add-on** → select `dist/firefox/manifest.json`\n\nThe toolbar badge shows `…` while copying, `✓` on success, or `!` when the page or clipboard is unavailable.\n\n### Site-owned button contract\n\nPages that already provide a copy control can suppress injected page UI with reserved ID:\n\n```html\n<button id=\"copy_as_markdown_btn\">Copy as Markdown</button>\n```\n\nAn empty marker works too:\n\n```html\n<div id=\"copy_as_markdown_btn\"></div>\n```\n\nAdding or removing marker updates injected UI dynamically. Extension toolbar action remains available.\n\n---\n\n## Library\n\nInstall browser-safe ESM package with pnpm:\n\n```bash\npnpm add @bvolpato/copy-as-markdown\n```\n\nLibrary loads only requested extractor chunks without starting userscript or extension UI. It does not inject buttons, write clipboard data, use extension APIs, or persist state. Fixed subpath imports let extension bundlers include only selected extractors.\n\n### Extension content-script example\n\nThis matcher enables only Jira, Confluence, GitHub, and Google Docs. Domain list narrows broad built-in site patterns, and `when` adds DOM-specific policy owned by calling extension.\n\n```typescript\nimport { createExtractorMatcher } from '@bvolpato/copy-as-markdown/core';\nimport { confluenceExtractor } from '@bvolpato/copy-as-markdown/extractors/confluence';\nimport { githubExtractor } from '@bvolpato/copy-as-markdown/extractors/github';\nimport { googleDocsExtractor } from '@bvolpato/copy-as-markdown/extractors/google-docs';\nimport { jiraExtractor } from '@bvolpato/copy-as-markdown/extractors/jira';\n\nconst matcher = createExtractorMatcher({\n  extractors: [jiraExtractor, confluenceExtractor, githubExtractor, googleDocsExtractor],\n  domains: [\n    'jira.example.com',\n    'confluence.example.com',\n    'github.com',\n    'docs.google.com',\n    'acme.atlassian.net',\n  ],\n  when: ({ document, extractor }) => {\n    if (extractor.name !== 'GitHub') return true;\n    return Boolean(document?.querySelector('[data-my-extension-context]'));\n  },\n});\n\nconst match = matcher.match({\n  url: window.location.href,\n  document,\n});\n\nif (match) {\n  const markdown = await match.extract();\n  await navigator.clipboard.writeText(markdown);\n}\n```\n\n`domains` uses exact hostnames. Use `*.atlassian.net` only when every Atlassian subdomain should be eligible. `origins` can restrict scheme and port. `urlPatterns` accepts userscript-style patterns or regular expressions. `when` receives parsed URL, current document, and matched extractor.\n\nRuntime selection is also available when extractor set is not known at build time:\n\n```typescript\nimport {\n  createExtractorMatcher,\n  loadExtractors,\n} from '@bvolpato/copy-as-markdown';\n\nconst extractors = await loadExtractors(['jira', 'confluence', 'github', 'google-docs']);\nconst matcher = createExtractorMatcher({ extractors });\n```\n\nSite extractors run inside active page because authenticated Jira and Confluence extraction can use same-origin APIs and current browser state. Use generic DOM conversion for detached or caller-created DOM:\n\n```typescript\nimport { domToMarkdown, htmlToMarkdown } from '@bvolpato/copy-as-markdown';\n\nconst fromElement = domToMarkdown(document.querySelector('main')!);\nconst fromDocument = domToMarkdown(new DOMParser().parseFromString(html, 'text/html'));\nconst fromHtml = htmlToMarkdown('<article><h1>Hello</h1></article>');\n```\n\nCustom extractors can be DOM-only. Pass empty URL patterns and inspect provided document:\n\n```typescript\nimport { defineExtractor, domToMarkdown } from '@bvolpato/copy-as-markdown/core';\n\nconst internalApp = defineExtractor({\n  name: 'Internal app',\n  matches: [],\n  detect: (document) => Boolean(document?.querySelector('[data-internal-app]')),\n  extract: async () => domToMarkdown(document.querySelector('main')!),\n});\n```\n\nUseful APIs:\n\n- `getAvailableExtractorIds()` lists 50+ loadable extractor IDs.\n- `loadExtractor('github')` loads one extractor chunk.\n- `loadExtractors(['github', 'jira'])` loads selected chunks in parallel.\n- `loadAllExtractors()` loads full catalog.\n- `getExtractors()` lists extractors loaded in current module instance.\n- `createExtractorMatcher()` applies extractor URL rules plus caller restrictions.\n- `defineExtractor()` creates custom extractor objects for same matcher.\n- `domToMarkdown()`, `elementToMarkdown()`, and `htmlToMarkdown()` convert caller-owned DOM without site UI.\n\n---\n\n## Supported Sites\n\n**Interaction Model:**\n- **Browser Extension:** Click the toolbar icon on any supported site, or the page button on Datadog dashboards, Datadog notebooks, and W&B runs.\n- **Userscript:** Clicks are handled via injected buttons (inline where a site integration provides a reviewed anchor, floating otherwise). Drag floating buttons out of the way without disabling them; their positions persist per site.\n\nCopies preserve content loaded by the site or returned by its export/API. Load or expand relevant sections before copying. Pagination, virtualized history, canvas-only views, and paywalls can hide content; W&B and MLflow numeric histories remain sampled or bounded.\n\n| Site | What's Extracted |\n| --- | --- |\n| **Wikipedia** | Article body, tables, infoboxes, citations, and references; edit controls stripped |\n| **Google Docs** | Full document export via Google Docs HTML export — headings, lists, tables, links, images, and off-screen content |\n| **Google Sheets** | Complete active sheet or selected range from authenticated exports; rendered grid fallback |\n| **Google Slides** | Choose current slide or full deck; preserves order, titles, text, links, and speaker notes when available |\n| **Gmail** | Full authenticated thread from Print all view — subject, participants, message headers, bodies, links, images, and attachments |\n| **Notion** | Pages and databases with properties, rich blocks, tables, code, and rendered rows |\n| **Documentation frameworks** | Mintlify, Docusaurus, GitBook, MkDocs, VitePress, Nextra, Sphinx, and Read the Docs on hosted or custom domains, with semantic content roots, navigation stripped, and code languages preserved |\n| **Microsoft 365** | Word, Excel, and PowerPoint web content through structured live-page views |\n| **Microsoft Teams** | Loaded chats, channels, and threads with authors, timestamps, replies, reactions, attachments, links, and code |\n| **Slack** | Loaded channel or thread messages with authors, timestamps, reactions, replies, and attachments |\n| **Discord** | Loaded channel messages and threads with authors, timestamps, replies, reactions, and attachments |\n| **Jira** | Authenticated REST issue fields, ADF descriptions/comments, and links; rendered issue DOM fallback |\n| **Linear** | Issues, projects, and documents with rendered properties, descriptions, links, and visible comments |\n| **Confluence** | Authenticated REST page body, labels, tables, and code; visible comments and rendered DOM fallback |\n| **Grokipedia** | Full article content with metadata |\n| **Google Search** | Query, featured snippets, knowledge panel, ranked results, \"People Also Ask\" |\n| **Bing Search** | Query, search results, knowledge sidebar, related searches |\n| **DuckDuckGo Search** | Query, ranked results, snippets, destination links, and related searches |\n| **Yahoo Search** | Query, ranked results, snippets, and destination links |\n| **Yandex Search** | Query, answer cards, ranked results, and related searches |\n| **Baidu Search** | Query, ranked results, abstracts, and destination links |\n| **Brave Search** | Query, answer cards, ranked results, discussions, and related searches |\n| **Reddit** | Post title, body, subreddit, author, score, threaded comments with depth |\n| **YouTube** | Video title, channel, views, likes, description, chapters, comments, transcript |\n| **WhatsApp Web** | Chat name, all loaded messages with sender, timestamp, media indicators |\n| **X (Twitter)** | Single posts with loaded replies and media, or all loaded timeline/search posts with engagement stats |\n| **Polymarket** | Market title, description, outcome probabilities, volume, resolution rules |\n| **OpenRouter** | Full model definitions, architecture, modalities, pricing, limits, supported parameters, benchmarks, provider endpoint fields, and FAQ |\n| **Artificial Analysis** | Homepage featured items, analysis sections, complete published leaderboards, model overview, exact benchmark values, technical specifications, provenance, and FAQ |\n| **DeepSWE** | Benchmark overview, all published leaderboard configurations and efficiency metrics, methodology, task examples, and blog sources |\n| **Datadog dashboards** | Dashboard title, timeframe, template variables, grouped widget values, top lists, and visible chart annotations |\n| **Datadog notebooks** | Notebook metadata, narrative headings and rich text, ordered visualization cells, types, no-data states, and visible chart annotations |\n| **Datadog Documentation** | Authored `.md` source when available; cleaned rendered documentation DOM otherwise |\n| **Weights & Biases** | Run metadata, configuration, numeric metric summaries, sparklines, and sampled history tables through W&B GraphQL |\n| **MLflow** | Self-hosted run metadata plus chart-mode comparisons for visible runs and loaded metrics, with paginated metric-history tables through same-origin APIs |\n| **Hugging Face** | Model, dataset, and Space repository metadata and tags, full rendered cards, nested card metadata, model configuration, tensor details, evaluation results, inference providers, model lineage, related collections/Spaces/papers, Space descriptions, and visible file listings |\n| **GitHub** | Issues and PRs, repository/directory listings with READMEs, full code-file contents, and canonical patches with commit/file metadata |\n| **GitLab** | Repositories, trees, code files, issues, merge requests, comments, and visible diffs |\n| **Bitbucket** | Repositories, source files, pull requests, issues, comments, and visible diffs |\n| **Perplexity** | Ordered user and assistant turns with citations and source links |\n| **Grok** | Ordered user and assistant turns with citations, code, and images |\n| **ChatGPT** | Ordered user and assistant turns with Markdown, canvas writing blocks, code, model metadata, and images |\n| **Claude** | Ordered user and assistant turns with Markdown, code, citations, and images |\n| **Gemini** | Ordered user and model turns with Markdown, code, citations, and images |\n| **Microsoft Copilot** | Consumer chats and shares with ordered turns, citations, code, files, images, and Copilot Pages when rendered |\n| **Gemini Notebook / NotebookLM** | Current and legacy notebook chats with ordered grounded answers, source citations, files, and rendered Studio artifacts |\n| **Mistral Vibe / Le Chat** | Chats with ordered turns, citations, files, code, images, workspace context, and rendered Canvas output |\n| **DeepSeek** | Authenticated and shared chats with ordered user and assistant turns, citations, reasoning/code content, files, and images |\n| **Google AI Studio** | Saved and new chat prompts with system instructions, ordered user/model turns, model metadata, grounding citations, code, files, and rendered artifacts |\n| **Meta AI** | Ordered user and assistant turns with citations, code, and images |\n| **LeetLLM** | Lessons, glossary pages, practice content, code, links, and learning context |\n| **Stack Overflow** | Question with votes & tags, all answers (✅ accepted marked), comment threads |\n| **Hacker News** | Post title, link, score, author, nested comment threads with depth |\n| **LinkedIn** | Profiles (experience, education, about), posts (with reactions and comments), articles |\n| **Facebook** | Posts, reels, captions, author metadata, engagement, and visible comments |\n| **Instagram** | Posts and reels with captions, media descriptions, engagement, and visible comments |\n| **TikTok** | Videos with creator, caption, engagement, transcript or captions, and visible comments |\n| **Pinterest** | Pins with creator, description, destination, media, engagement, and visible comments |\n| **VK** | Posts with author, timestamp, text, media, engagement, and visible comments |\n| **Amazon** | Product title, ASIN, price, rating, feature bullets, tech specs, reviews (top 10) |\n| **Temu** | Product title, price, availability, ratings, variants, specifications, and description |\n| **Booking.com** | Hotels and search results with prices, scores, facilities, policies, and availability |\n| **Netflix** | Title metadata, synopsis, cast, genres, ratings, seasons, and visible episodes |\n| **Twitch** | Channels, live streams, videos, and clips with game, viewers, tags, and description |\n| **Weather.com** | Current conditions, alerts, hourly outlook, and daily forecast |\n| **arXiv** | Paper title, authors, abstract, subjects, DOI, links; full body from HTML pages |\n| **Globo** | Articles and videos with headline, author, date, structured metadata, and clean body |\n| **FOX** | Shows, episodes, movies, and videos with synopsis and structured details |\n| **News sites** | Fox News, CNN, BBC, NYT, Reuters, and 20+ others — article body, author, date; paywall detection |\n\nEvery extractor is purpose-built to separate **signal from noise**: no ads, no navigation menus, no cookie banners, no related-articles sidebars. Just the content that matters.\n\nAI chat extractors use rendered page content only and explicitly report visible-only coverage. When a product virtualizes history or exposes a load-older control, output also warns that older content was not loaded.\n\nW&B returns up to 500 sampled history rows per run through its browser GraphQL API. MLflow run history fetches are paginated up to 10,000 points per metric. MLflow chart comparisons include up to 10 visible runs and 50 loaded run-metric series, fetching up to 2,500 points per series. Both integrations include full-series statistics, then evenly sample Markdown history rows when needed to keep clipboard output bounded. W&B Server and arbitrary self-hosted MLflow deployments work through userscript content detection or extension toolbar; their active browser session must permit same-origin API access.\n\nIf an extractor is not explicitly opted into inline placement (for userscript builds), the button stays in the bottom-right corner. If an inline anchor is enabled but the selector is missing (for example after a site redesign), the button also falls back to the bottom-right floating button.\n\n---\n\n## Why Markdown for LLMs?\n\n### 1. Structure = Understanding\n\n```\n# Vigenère Cipher                      ← LLM knows: this is the topic\n## History                              ← LLM knows: this is a section about history\n| Inventor | Blaise de Vigenère |       ← LLM knows: structured key-value data\n```\n\nThe LLM doesn't have to *guess* what's a heading vs. body text vs. metadata. Markdown makes the hierarchy explicit.\n\n### 2. Token Efficiency\n\nRaw HTML from a typical Wikipedia article: **~200K characters**.\nCopy as Markdown output: **~15K characters**.\n\nThat's **>90% noise reduction** — more room for your actual conversation.\n\n### 3. Faithful Reproduction\n\n- **Headers** → `#`, `##`, `###` (hierarchy preserved)\n- **Tables** → Pipe-delimited Markdown tables (data preserved)\n- **Code blocks** → Fenced with language tags (syntax preserved)\n- **Links** → `[text](url)` (sources preserved)\n- **Lists** → Nested bullets/numbers (structure preserved)\n\n### 4. Universal Compatibility\n\nEvery major LLM — GPT-5.4, Claude, Gemini, Llama, Mistral — understands Markdown natively. It's the lingua franca of AI conversations.\n\n---\n\n## Example Output\n\nClicking the extension icon or userscript button on a Wikipedia article produces:\n\n```markdown\n---\nsource: Wikipedia\ntitle: Vigenère cipher\nurl: https://en.wikipedia.org/wiki/Vigen%C3%A8re_cipher\nlast_modified: 15 March 2025\n---\n\n# Vigenère cipher\n\nThe **Vigenère cipher** is a method of encrypting alphabetic text\nwhere each letter of the plaintext is encoded with a different\nCaesar cipher, whose increment is determined by the corresponding\nletter of another text, the **key**.\n\n## History\n\nThe Vigenère cipher is simple enough to be a field cipher if it\nis used in conjunction with cipher disks...\n\n## Description\n\n| Component | Details |\n| --- | --- |\n| Type | Polyalphabetic substitution |\n| Key | A repeating keyword |\n| Inventor | Blaise de Vigenère |\n```\n\n---\n\n## Build\n\n```bash\n# Clone the repository\ngit clone https://github.com/bvolpato/copy-as-markdown.git\ncd copy-as-markdown\n\n# Install dependencies\npnpm install\n\n# Type-check\npnpm typecheck\n\n# Build all targets\npnpm build\n\n# Package extensions as .zip\npnpm package:all\n```\n\n### Output\n\n```\ndist/\n├── userscript/\n│   └── copy-as-markdown.user.js       ← Install directly in Tampermonkey\n├── chrome/\n│   ├── manifest.json                   ← Chrome Manifest V3\n│   ├── content.js\n│   └── icons/\n├── firefox/\n│   ├── manifest.json                   ← Firefox Manifest V2\n│   ├── content.js\n│   └── icons/\n└── library/\n    ├── index.js                        ← Browser-safe ESM package entry\n    ├── browser.js                      ← CopyAsMarkdown browser global\n    ├── chunks/                         ← Lazy site extractor chunks\n    ├── extractors/                     ← Fixed Jira, Confluence, GitHub, Hugging Face, Google Docs entries\n    └── types/                          ← TypeScript declarations\n```\n\n### Publish library\n\nUnscoped `copy-as-markdown` name is already used on npm. This repository publishes as public scoped package `@bvolpato/copy-as-markdown`.\n\nValidate exact package contents without publishing:\n\n```bash\npnpm pack:library\n```\n\nReleases are driven by signed `vMAJOR.MINOR.PATCH` tags on `main`. The release workflow validates the tag, tests and packages every target, then creates the GitHub release and browser artifacts. Tag pushes do not publish to npm.\n\nTo publish the npm package explicitly, manually run the Release workflow on the signed tag and enable its `publish_npm` input. The input defaults to `false`.\n\nnpm trusted publishing requires one-time package configuration by an npm owner:\n\n```bash\npnpm dlx npm@12.0.2 trust github @bvolpato/copy-as-markdown \\\n  --file release.yml \\\n  --repo bvolpato/copy-as-markdown \\\n  --allow-publish \\\n  --yes\n```\n\nAfter this one-time setup, an explicitly enabled npm job uses OIDC trusted publishing and provenance without a long-lived token. Existing package versions are detected and skipped. `prepack` rebuilds standalone library without touching extension or userscript artifacts. `files` allowlist publishes only standalone library, declarations, README, license, and package manifest.\n\n### Tech Stack\n\n- **TypeScript** — all source code, compiled with esbuild\n- **esbuild** — fast userscript, extension, ESM, and browser-global bundling\n- **pnpm** — package management\n- **Zero runtime dependencies** — extension targets are self-contained; library uses local ESM chunks only\n\n---\n\n## Architecture\n\n```\nsrc/\n├── core/\n│   ├── types.ts        ← AnchorConfig, ExtractorConfig, PageMetadata interfaces\n│   ├── markdown.ts     ← HTML→Markdown converter (tables, lists, code, etc.)\n│   ├── ui.ts           ← Button injection: anchored (inline) or floating (FAB)\n│   ├── utils.ts        ← DOM helpers, meta extraction, paywall detection\n│   └── registry.ts     ← URL pattern → extractor mapping\n├── extractors/\n│   ├── wikipedia.ts    ← extractor with active inline placement\n│   ├── google-docs.ts  ← extractor with active inline placement\n│   ├── datadog-dashboard.ts ← semantic dashboard extractor + toolbar placement\n│   ├── datadog-notebook.ts ← structured notebook extractor + toolbar placement\n│   ├── youtube.ts      ← extractor\n│   ├── reddit.ts       ← extractor\n│   ├── x-twitter.ts    ← extractor\n│   └── news.ts         ← extractor\n├── catalog.ts          ← Load extractors without UI startup\n├── library/index.ts    ← Standalone matcher and DOM API\n└── main.ts             ← Entry point: detect site, show button\nbuild/\n└── build.ts            ← esbuild bundler → userscript + extensions + library\n```\n\n### Userscript Button Positioning\n\nUserscripts support floating and inline buttons. Browser extensions use the toolbar icon on any page, with inline buttons also enabled for Datadog dashboards, Datadog notebooks, and W&B runs.\n\nThe default behavior is simple: unless a site is explicitly opted into inline placement, the userscript button is rendered as a floating action button in the bottom-right corner.\n\nTo enable a custom inline position for a specific site, you need two things:\n\n1. An `anchor` config that describes where and how to inject the button\n2. `buttonPlacement: 'anchor'` on the extractor\n\nThat second step is the gate. It lets us keep site-specific selectors in the codebase without turning them on until we're ready.\n\n```typescript\nregister({\n  name: 'Wikipedia',\n  matches: ['*://*.wikipedia.org/wiki/*'],\n  buttonPlacement: 'anchor',\n  anchor: {\n    selector: '#p-views ul',\n    position: 'append',\n    style: 'tab',\n    css: {\n      marginLeft: '8px',\n      paddingLeft: '8px',\n      borderLeft: '1px solid #a2a9b1',\n    },\n    label: 'Copy as Markdown',\n  },\n  async extract() {\n    // ...\n  },\n});\n```\n\nThe `anchor` object controls the inline position:\n\n```typescript\nanchor: {\n  selector: '#p-views ul',   // CSS selector for the target container\n  position: 'append',        // 'append' | 'prepend' | 'before' | 'after'\n  style: 'tab',              // 'tab' | 'pill' | 'icon' | 'link'\n  wrapperTag: 'li',          // Optional wrapper when the host expects a specific child tag\n  wrapperClass: 'mw-list-item', // Optional wrapper classes\n  wrapperCss: { marginLeft: '8px' }, // Optional wrapper overrides\n  css: { color: '#0645ad' },  // Optional button overrides\n  label: 'Copy as Markdown',  // Custom label (omit for icon-only)\n}\n```\n\nUse `wrapperTag` / `wrapperClass` / `wrapperCss` when the host container expects a particular DOM shape. Wikipedia is the main example: the tab bar is a `ul`, so the injected control needs to live inside an `li` to align correctly with the native tabs.\n\nPositioning rules:\n\n- Omit `buttonPlacement`, or set it to `'floating'`, to keep the default bottom-right button\n- Add `buttonPlacement: 'anchor'` to activate the extractor's `anchor` config\n- If the anchor selector is missing at runtime, the UI falls back to the bottom-right floating button\n\nExtractors enable anchored placement only after their site selector and SPA lifecycle are covered by browser fixtures. Others retain floating placement.\n\n### Adding a New Site\n\n1. Create `src/extractors/my-site.ts`\n2. Import `register` from `../core/registry` and call it with `name`, `matches`, and `extract`\n3. Leave the button floating by default unless you are intentionally enabling a reviewed inline placement\n4. If you want to prepare an inline placement for later, add an `anchor` config but do not set `buttonPlacement: 'anchor'` yet\n5. Import the new file in `src/catalog.ts` and add its library loader in `src/library/loaders.ts`\n6. Run `pnpm build` — the new patterns propagate to all targets\n\n### Browser Tests\n\nBrowser tests and fixture tools use Puppeteer's `headless: 'shell'` mode with `chrome-headless-shell`, which runs without visible windows.\n\n```bash\npnpm exec puppeteer browsers install chrome-headless-shell\npnpm test:regression\n```\n\nTo use an existing headless shell installation, set `PUPPETEER_EXECUTABLE_PATH` to its executable path. Unset this variable if it points to regular Chrome. Tests launch their own browser and close it when finished.\n\n### Captured Public-Site Fixtures\n\nPublic extractors can be checked against browser-rendered pages without committing raw page data:\n\n```bash\n# Capture current public page in a clean headless browser.\n# If live capture fails, try a Wayback snapshot.\npnpm fixtures:capture -- --site mdn\n\n# Capture every curated public case, reporting failures without stopping early.\npnpm fixtures:capture -- --all\n\n# Force one source while debugging.\npnpm fixtures:capture -- --site mdn --source live\npnpm fixtures:capture -- --site mdn --source wayback\n\n# Replay committed fixtures entirely offline.\npnpm fixtures:verify\n```\n\nThe [catalog](test/sites/catalog.yaml) contains 30 captures across 25 extractor families, including four Hugging Face model, dataset, and file-tree pages. The [coverage inventory](test/sites/coverage.json) records public probes and gaps across the full extractor catalog. Capture checks the original copy button, extractor identity, placement, clipboard output, and `contentRequired` strings or `contentSelectors` before anonymizing the DOM. Bot challenges, login screens, and error pages fail capture.\n\nRaw reference screenshots stay under gitignored `.fixture-work/`. Committed fixtures contain synthetic text, normalized links, no scripts or media, a screenshot of sanitized DOM, expected Markdown, and source provenance. `linkPrefixes` can retain curated public route prefixes for path-aware file extractors; link suffixes are still replaced.\n\nA failed live attempt writes `live-failure.png` before Wayback fallback. Archive discovery tries CDX, then direct replay if CDX is unavailable. Accepted snapshots record the exact timestamp, digest, and downloaded-response SHA-256. Pin the accepted timestamp and digest in the catalog; set `digestAlgorithm: sha256` when pinning the response hash. Archived DOM is checked at the original URL with site scripts removed and outgoing requests blocked. An archive verifies historical content and markup; it does not prove current site placement or access.\n\nCapture refuses credentials, localhost, private IP addresses, and non-HTTP protocols. Use only curated public URLs. Authenticated pages require synthetic fixtures and must never use this capture path.\n\nEach captured fixture verifies extractor identity, exact anchor relationship, singleton UI, Markdown bounds, required output, synthetic markers for excluded page chrome, exact expected Markdown, and privacy rules. `pnpm test:regression` runs this offline lane in CI.\n\n---\n\n## Contributing\n\nPRs welcome! See [CONTRIBUTING.md](CONTRIBUTING.md) for details.\n\n---\n\n## License\n\nMIT © [Bruno Volpato](https://github.com/bvolpato)\n","readmeFilename":"README.md"}