{"_id":"@agenson-horrowitz/web-content-extractor-mcp","_rev":"9-9b77907016828b0ce36b42fc639e0ab7","name":"@agenson-horrowitz/web-content-extractor-mcp","dist-tags":{"latest":"1.0.8"},"versions":{"1.0.0":{"name":"@agenson-horrowitz/web-content-extractor-mcp","version":"1.0.0","keywords":["mcp","mcp-server","ai-agent","tool-server","web-scraping","content-extraction","web-crawler","markdown-conversion","agents","ai-tools","readability","structured-data","webpage-parser"],"author":{"name":"Agenson Horrowitz","email":"hello@agensonhorrowitz.cc"},"license":"MIT","_id":"@agenson-horrowitz/web-content-extractor-mcp@1.0.0","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/web-content-extractor-mcp/issues"},"bin":{"web-content-extractor-mcp":"dist/index.js"},"dist":{"shasum":"1ba6ad1e7854adfe043989b88aeb2f2b3962454b","tarball":"https://registry.npmjs.org/@agenson-horrowitz/web-content-extractor-mcp/-/web-content-extractor-mcp-1.0.0.tgz","fileCount":7,"integrity":"sha512-4rLOpQ8gz4mhLL7sBfpuYKkS8jp7GOaLxyAqkUmPnRSKvEuoPI1+p+TRJJwrisAA8fVans5RIpqglkM2z0q9lg==","signatures":[{"sig":"MEYCIQDbZJgyMhlA1hZefbu7W4GUfUxJ3zTaQoHgIu3uDQ9cbAIhAOU72g8GrTjyVbP0wdwJXhCuGIodeNL6z64/kCt499yQ","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":64078},"main":"dist/index.js","_from":"file:agenson-horrowitz-web-content-extractor-mcp-1.0.0.tgz","types":"./dist/index.d.ts","engines":{"node":">=18.0.0"},"scripts":{"dev":"tsc --watch","test":"jest","build":"tsc","start":"node dist/index.js","prepublishOnly":"npm run build"},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"_resolved":"/home/openclaw/business/mcp-servers/web-extractor/agenson-horrowitz-web-content-extractor-mcp-1.0.0.tgz","_integrity":"sha512-4rLOpQ8gz4mhLL7sBfpuYKkS8jp7GOaLxyAqkUmPnRSKvEuoPI1+p+TRJJwrisAA8fVans5RIpqglkM2z0q9lg==","repository":{"url":"git+https://github.com/agenson-tools/web-content-extractor-mcp.git","type":"git"},"_npmVersion":"10.9.4","description":"Agent-optimized MCP server for extracting clean, structured content from web pages - built specifically for AI agents that need LLM-friendly web data","directories":{},"_nodeVersion":"22.22.1","dependencies":{"jsdom":"^24.0.0","turndown":"^7.1.2","playwright":"^1.42.0","metascraper":"^5.38.0","metascraper-url":"^5.38.0","metascraper-date":"^5.38.0","metascraper-image":"^5.38.0","metascraper-title":"^5.38.0","metascraper-author":"^5.38.0","@mozilla/readability":"^0.5.0","metascraper-description":"^5.38.0","@modelcontextprotocol/sdk":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0","@types/node":"^20.0.0","@types/jsdom":"^21.1.6","@types/turndown":"^5.0.4"},"_npmOperationalInternal":{"tmp":"tmp/web-content-extractor-mcp_1.0.0_1775116018632_0.4628428793679633","host":"s3://npm-registry-packages-npm-production"}},"1.0.1":{"name":"@agenson-horrowitz/web-content-extractor-mcp","version":"1.0.1","keywords":["mcp","mcp-server","ai-agent","tool-server","web-scraping","content-extraction","web-crawler","markdown-conversion","agents","ai-tools","readability","structured-data","webpage-parser"],"author":{"name":"Agenson Horrowitz","email":"hello@agensonhorrowitz.cc"},"license":"MIT","_id":"@agenson-horrowitz/web-content-extractor-mcp@1.0.1","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/web-content-extractor-mcp/issues"},"bin":{"web-content-extractor-mcp":"dist/index.js"},"dist":{"shasum":"74ae359fa294a0b77ac1f43628f35aff4ffa288a","tarball":"https://registry.npmjs.org/@agenson-horrowitz/web-content-extractor-mcp/-/web-content-extractor-mcp-1.0.1.tgz","fileCount":7,"integrity":"sha512-rJbwzRRuynhkaPU5ohW9ngOR4DJZgejG45BC9Jdw7AJOSwLDobo9/B8YgH7WC/IwcEl/joW2Sf4bZXg0qWwc1A==","signatures":[{"sig":"MEUCIHQHWn1hPiK5nykY+yySpLjvCBop5QgNsK3T0h87hhPuAiEA2F16dPX0siSOLI+rgq/A92Jolrvir5HhrZpaZdRK8lQ=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":64075},"main":"dist/index.js","types":"./dist/index.d.ts","engines":{"node":">=18.0.0"},"gitHead":"831696d3400482a3b8b193186ba1b651e7545455","scripts":{"dev":"tsc --watch","test":"jest","build":"tsc","start":"node dist/index.js","prepublishOnly":"npm run build"},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"repository":{"url":"git+https://github.com/agenson-tools/web-content-extractor-mcp.git","type":"git"},"_npmVersion":"10.9.4","description":"Agent-optimized MCP server for extracting clean, structured content from web pages - built specifically for AI agents that need LLM-friendly web data","directories":{},"_nodeVersion":"22.22.1","dependencies":{"jsdom":"^24.0.0","turndown":"^7.1.2","playwright":"^1.42.0","metascraper":"^5.38.0","metascraper-url":"^5.38.0","metascraper-date":"^5.38.0","metascraper-image":"^5.38.0","metascraper-title":"^5.38.0","metascraper-author":"^5.38.0","@mozilla/readability":"^0.5.0","metascraper-description":"^5.38.0","@modelcontextprotocol/sdk":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0","@types/node":"^20.0.0","@types/jsdom":"^21.1.6","@types/turndown":"^5.0.4"},"_npmOperationalInternal":{"tmp":"tmp/web-content-extractor-mcp_1.0.1_1775123548978_0.2693278525714309","host":"s3://npm-registry-packages-npm-production"}},"1.0.2":{"name":"@agenson-horrowitz/web-content-extractor-mcp","version":"1.0.2","keywords":["mcp","mcp-server","ai-agent","tool-server","web-scraping","content-extraction","web-crawler","markdown-conversion","agents","ai-tools","readability","structured-data","webpage-parser"],"author":{"name":"Agenson Horrowitz","email":"hello@agensonhorrowitz.cc"},"license":"MIT","_id":"@agenson-horrowitz/web-content-extractor-mcp@1.0.2","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/web-content-extractor-mcp/issues"},"bin":{"web-content-extractor-mcp":"dist/index.js"},"dist":{"shasum":"5c5cfda9acffaef98bce7aad0ba5c8718092df9d","tarball":"https://registry.npmjs.org/@agenson-horrowitz/web-content-extractor-mcp/-/web-content-extractor-mcp-1.0.2.tgz","fileCount":7,"integrity":"sha512-2ibZojMWifIX1THZMlHexn6iyzJYIMZc6ajVXJXMP/rG+lLUZCbxleQNL/L3gh+l5uNxYRDxOupk48wZfrLR1Q==","signatures":[{"sig":"MEUCIQCLFDxHqyZQO5/K7hrwHGzdEI6YbozykEofHtbVVkAL2QIgKI3cEVAlLHfi12Xq7XJSahnpUIkJ+0ZVD4ZfkqUeA0Q=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":64462},"main":"dist/index.js","types":"./dist/index.d.ts","engines":{"node":">=18.0.0"},"gitHead":"b8332f8fc9783e35586d82673abf1e1460738612","scripts":{"dev":"tsc --watch","test":"jest","build":"tsc","start":"node dist/index.js","prepublishOnly":"npm run build"},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"repository":{"url":"git+https://github.com/agenson-tools/web-content-extractor-mcp.git","type":"git"},"_npmVersion":"10.9.4","description":"Agent-optimized MCP server for extracting clean, structured content from web pages - built specifically for AI agents that need LLM-friendly web data","directories":{},"_nodeVersion":"22.22.1","dependencies":{"jsdom":"^24.0.0","turndown":"^7.1.2","playwright":"^1.42.0","metascraper":"^5.38.0","metascraper-url":"^5.38.0","metascraper-date":"^5.38.0","metascraper-image":"^5.38.0","metascraper-title":"^5.38.0","metascraper-author":"^5.38.0","@mozilla/readability":"^0.5.0","metascraper-description":"^5.38.0","@modelcontextprotocol/sdk":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0","@types/node":"^20.0.0","@types/jsdom":"^21.1.6","@types/turndown":"^5.0.4"},"_npmOperationalInternal":{"tmp":"tmp/web-content-extractor-mcp_1.0.2_1775123958758_0.9108800367507248","host":"s3://npm-registry-packages-npm-production"}},"1.0.3":{"name":"@agenson-horrowitz/web-content-extractor-mcp","version":"1.0.3","keywords":["mcp","mcp-server","ai-agent","tool-server","web-scraping","content-extraction","web-crawler","markdown-conversion","agents","ai-tools","readability","structured-data","webpage-parser"],"author":{"name":"Agenson Horrowitz","email":"hello@agensonhorrowitz.cc"},"license":"MIT","_id":"@agenson-horrowitz/web-content-extractor-mcp@1.0.3","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/web-content-extractor-mcp/issues"},"bin":{"web-content-extractor-mcp":"dist/index.js"},"dist":{"shasum":"ce54e811cbb1ddba592308f9a40be3392c4657ae","tarball":"https://registry.npmjs.org/@agenson-horrowitz/web-content-extractor-mcp/-/web-content-extractor-mcp-1.0.3.tgz","fileCount":7,"integrity":"sha512-ydFhfySav2tePxV996w0VZk8WpDfbDfUg6dD1opCZ27ofdLJ/ZKkow+mPmZT+Y4COE8nbPin07I31Q3W5OWJNA==","signatures":[{"sig":"MEQCIENEzWaZD3nXeO5hdAQ+WG8JU9ad2kQDTcQ7cufko1FIAiALf5GUTbZUxL0dmou3lzVLBxWKbwiGcYed7jlyDJtcdg==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":64314},"main":"dist/index.js","types":"./dist/index.d.ts","engines":{"node":">=18.0.0"},"gitHead":"54d660867d3b29501188284658e17bfa27c11f59","scripts":{"dev":"tsc --watch","test":"jest","build":"tsc","start":"node dist/index.js","prepublishOnly":"npm run build"},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"repository":{"url":"git+https://github.com/agenson-tools/web-content-extractor-mcp.git","type":"git"},"_npmVersion":"10.9.4","description":"Agent-optimized MCP server for extracting clean, structured content from web pages - built specifically for AI agents that need LLM-friendly web data","directories":{},"_nodeVersion":"22.22.1","dependencies":{"jsdom":"^24.0.0","turndown":"^7.1.2","playwright":"^1.42.0","metascraper":"^5.38.0","metascraper-url":"^5.38.0","metascraper-date":"^5.38.0","metascraper-image":"^5.38.0","metascraper-title":"^5.38.0","metascraper-author":"^5.38.0","@mozilla/readability":"^0.5.0","metascraper-description":"^5.38.0","@modelcontextprotocol/sdk":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0","@types/node":"^20.0.0","@types/jsdom":"^21.1.6","@types/turndown":"^5.0.4"},"_npmOperationalInternal":{"tmp":"tmp/web-content-extractor-mcp_1.0.3_1775124486942_0.06333942221719258","host":"s3://npm-registry-packages-npm-production"}},"1.0.4":{"name":"@agenson-horrowitz/web-content-extractor-mcp","version":"1.0.4","keywords":["mcp","mcp-server","ai-agent","tool-server","web-scraping","content-extraction","web-crawler","markdown-conversion","agents","ai-tools","readability","structured-data","webpage-parser"],"author":{"name":"Agenson Horrowitz","email":"hello@agensonhorrowitz.cc"},"license":"MIT","_id":"@agenson-horrowitz/web-content-extractor-mcp@1.0.4","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/web-content-extractor-mcp/issues"},"bin":{"web-content-extractor-mcp":"dist/index.js"},"dist":{"shasum":"56f80c16528bbfb6cbf869c64cae2943fd20343a","tarball":"https://registry.npmjs.org/@agenson-horrowitz/web-content-extractor-mcp/-/web-content-extractor-mcp-1.0.4.tgz","fileCount":7,"integrity":"sha512-REIEeTJp12MRzGlgj+QHrOz9dcu4VpT9p5zytINVOxf5EWOMvLFYmwBpQJRzXd23BzZF6m1jLOeZ1v514Rt8iA==","signatures":[{"sig":"MEQCIHkK36I4oHNGoGuMngroabQt65dl5gDYUvwWd8VukOsjAiAAxzIiJ6UIdvrzO4ch0VclrNvUJhO99/3B26kSArrZTQ==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":64630},"main":"dist/index.js","types":"./dist/index.d.ts","engines":{"node":">=18.0.0"},"gitHead":"13b1e869aa2280c8aa64afc5ef2977f45f0d9081","scripts":{"dev":"tsc --watch","test":"jest","build":"tsc","start":"node dist/index.js","prepublishOnly":"npm run build"},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"repository":{"url":"git+https://github.com/agenson-tools/web-content-extractor-mcp.git","type":"git"},"_npmVersion":"10.9.4","description":"Agent-optimized MCP server for extracting clean, structured content from web pages - built specifically for AI agents that need LLM-friendly web data","directories":{},"_nodeVersion":"22.22.1","dependencies":{"jsdom":"^24.0.0","turndown":"^7.1.2","playwright":"^1.42.0","metascraper":"^5.38.0","metascraper-url":"^5.38.0","metascraper-date":"^5.38.0","metascraper-image":"^5.38.0","metascraper-title":"^5.38.0","metascraper-author":"^5.38.0","@mozilla/readability":"^0.5.0","metascraper-description":"^5.38.0","@modelcontextprotocol/sdk":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0","@types/node":"^20.0.0","@types/jsdom":"^21.1.6","@types/turndown":"^5.0.4"},"_npmOperationalInternal":{"tmp":"tmp/web-content-extractor-mcp_1.0.4_1775125835517_0.6186145500898421","host":"s3://npm-registry-packages-npm-production"}},"1.0.5":{"name":"@agenson-horrowitz/web-content-extractor-mcp","version":"1.0.5","keywords":["mcp","mcp-server","ai-agent","tool-server","web-scraping","content-extraction","web-crawler","markdown-conversion","agents","ai-tools","readability","structured-data","webpage-parser"],"author":{"name":"Agenson Horrowitz","email":"hello@agensonhorrowitz.cc"},"license":"MIT","_id":"@agenson-horrowitz/web-content-extractor-mcp@1.0.5","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/web-content-extractor-mcp/issues"},"bin":{"web-content-extractor-mcp":"dist/index.js"},"dist":{"shasum":"6926b5e4ea042ec6c896b6f3b6beffaaa2c5ebac","tarball":"https://registry.npmjs.org/@agenson-horrowitz/web-content-extractor-mcp/-/web-content-extractor-mcp-1.0.5.tgz","fileCount":7,"integrity":"sha512-jVSu7PNefz2epS0KcQrNPBeefhuVZT1o5gJmj0ZZgnQYzbgvpQiskcHngUuGFcFN3HrgpmwoPc+4Yt2oXUeCKA==","signatures":[{"sig":"MEYCIQDbKkxhxF94rYmnvoolBDIfSa2Zl/yCWTiH2uxGGKw//wIhAK57qlmqOxYHfRcnVzwxQu4te79m7taqH4azOX6VPmga","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":64630},"main":"dist/index.js","types":"./dist/index.d.ts","engines":{"node":">=18.0.0"},"gitHead":"13b1e869aa2280c8aa64afc5ef2977f45f0d9081","scripts":{"dev":"tsc --watch","test":"jest","build":"tsc","start":"node dist/index.js","prepublishOnly":"npm run build"},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"repository":{"url":"git+https://github.com/agenson-tools/web-content-extractor-mcp.git","type":"git"},"_npmVersion":"10.9.4","description":"Agent-optimized MCP server for extracting clean, structured content from web pages - built specifically for AI agents that need LLM-friendly web data","directories":{},"_nodeVersion":"22.22.1","dependencies":{"jsdom":"^24.0.0","turndown":"^7.1.2","playwright":"^1.42.0","metascraper":"^5.38.0","metascraper-url":"^5.38.0","metascraper-date":"^5.38.0","metascraper-image":"^5.38.0","metascraper-title":"^5.38.0","metascraper-author":"^5.38.0","@mozilla/readability":"^0.5.0","metascraper-description":"^5.38.0","@modelcontextprotocol/sdk":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0","@types/node":"^20.0.0","@types/jsdom":"^21.1.6","@types/turndown":"^5.0.4"},"_npmOperationalInternal":{"tmp":"tmp/web-content-extractor-mcp_1.0.5_1775127231299_0.7502541277159815","host":"s3://npm-registry-packages-npm-production"}},"1.0.6":{"name":"@agenson-horrowitz/web-content-extractor-mcp","version":"1.0.6","keywords":["mcp","mcp-server","ai-agent","tool-server","web-scraping","content-extraction","web-crawler","markdown-conversion","agents","ai-tools","readability","structured-data","webpage-parser"],"author":{"name":"Agenson Horrowitz","email":"hello@agensonhorrowitz.cc"},"license":"MIT","_id":"@agenson-horrowitz/web-content-extractor-mcp@1.0.6","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/web-content-extractor-mcp/issues"},"bin":{"web-content-extractor-mcp":"dist/index.js"},"dist":{"shasum":"d0e978a612bfe890175a9ba8285853efdb3fc1a9","tarball":"https://registry.npmjs.org/@agenson-horrowitz/web-content-extractor-mcp/-/web-content-extractor-mcp-1.0.6.tgz","fileCount":7,"integrity":"sha512-chwoa0dRaYFcF5Yt0nXq0jgDegBNIKrmfvcyhhIO7+Scb9FhKzI53YVzcNff/eiS2t0cTK+DUm5XOuQzRF80CA==","signatures":[{"sig":"MEYCIQCdfRO2Dslv1k9S3ZHvsDJmk0wMkABhCGYPgA33/l+2eQIhAOwAhX9xfVn0dFA9SSSUvmR9fcEVZJvtLyaCV4QvRvta","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":65080},"main":"dist/index.js","types":"./dist/index.d.ts","engines":{"node":">=18.0.0"},"gitHead":"13b1e869aa2280c8aa64afc5ef2977f45f0d9081","scripts":{"dev":"tsc --watch","test":"jest","build":"tsc","start":"node dist/index.js","prepublishOnly":"npm run build"},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"repository":{"url":"git+https://github.com/agenson-tools/web-content-extractor-mcp.git","type":"git"},"_npmVersion":"10.9.4","description":"Agent-optimized MCP server for extracting clean, structured content from web pages - built specifically for AI agents that need LLM-friendly web data","directories":{},"_nodeVersion":"22.22.1","dependencies":{"jsdom":"^24.0.0","turndown":"^7.1.2","playwright":"^1.42.0","metascraper":"^5.38.0","metascraper-url":"^5.38.0","metascraper-date":"^5.38.0","metascraper-image":"^5.38.0","metascraper-title":"^5.38.0","metascraper-author":"^5.38.0","@mozilla/readability":"^0.5.0","metascraper-description":"^5.38.0","@modelcontextprotocol/sdk":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0","@types/node":"^20.0.0","@types/jsdom":"^21.1.6","@types/turndown":"^5.0.4"},"_npmOperationalInternal":{"tmp":"tmp/web-content-extractor-mcp_1.0.6_1775130424816_0.6569174569774923","host":"s3://npm-registry-packages-npm-production"}},"1.0.7":{"name":"@agenson-horrowitz/web-content-extractor-mcp","version":"1.0.7","keywords":["mcp","mcp-server","ai-agent","tool-server","web-scraping","content-extraction","web-crawler","markdown-conversion","agents","ai-tools","readability","structured-data","webpage-parser"],"author":{"name":"Agenson Horrowitz","email":"hello@agensonhorrowitz.cc"},"license":"MIT","_id":"@agenson-horrowitz/web-content-extractor-mcp@1.0.7","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/web-content-extractor-mcp/issues"},"bin":{"web-content-extractor-mcp":"dist/index.js"},"dist":{"shasum":"d94182a576349b2d15bfe33dc7f3177e70b224b0","tarball":"https://registry.npmjs.org/@agenson-horrowitz/web-content-extractor-mcp/-/web-content-extractor-mcp-1.0.7.tgz","fileCount":7,"integrity":"sha512-GE6HataMHs3xXtRXEDxKjhO/rtKOmIgPX/ATXKyds0hh/qU4sMcsb1PppQ2u0RcMbDUP3bJMPYaxDT2DirbAag==","signatures":[{"sig":"MEQCIFOIJqyjEeS4zZMcRkqucRe49JO0bhfvnKo9F0mpHpp/AiBSOdo/eeOkaWI5P653bNpC418DyJARWkXFnGIkjCvj+g==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":65142},"main":"dist/index.js","types":"./dist/index.d.ts","engines":{"node":">=18.0.0"},"gitHead":"13b1e869aa2280c8aa64afc5ef2977f45f0d9081","mcpName":"io.github.agenson-tools/web-content-extractor","scripts":{"dev":"tsc --watch","test":"jest","build":"tsc","start":"node dist/index.js","prepublishOnly":"npm run build"},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"repository":{"url":"git+https://github.com/agenson-tools/web-content-extractor-mcp.git","type":"git"},"_npmVersion":"10.9.4","description":"Agent-optimized MCP server for extracting clean, structured content from web pages - built specifically for AI agents that need LLM-friendly web data","directories":{},"_nodeVersion":"22.22.1","dependencies":{"jsdom":"^24.0.0","turndown":"^7.1.2","playwright":"^1.42.0","metascraper":"^5.38.0","metascraper-url":"^5.38.0","metascraper-date":"^5.38.0","metascraper-image":"^5.38.0","metascraper-title":"^5.38.0","metascraper-author":"^5.38.0","@mozilla/readability":"^0.5.0","metascraper-description":"^5.38.0","@modelcontextprotocol/sdk":"^1.0.0"},"_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0","@types/node":"^20.0.0","@types/jsdom":"^21.1.6","@types/turndown":"^5.0.4"},"_npmOperationalInternal":{"tmp":"tmp/web-content-extractor-mcp_1.0.7_1775132434669_0.21041475226508144","host":"s3://npm-registry-packages-npm-production"}},"1.0.8":{"name":"@agenson-horrowitz/web-content-extractor-mcp","version":"1.0.8","mcpName":"io.github.agenson-horrowitz/web-content-extractor","description":"Agent-optimized MCP server for extracting clean, structured content from web pages - built specifically for AI agents that need LLM-friendly web data","main":"dist/index.js","bin":{"web-content-extractor-mcp":"dist/index.js"},"scripts":{"build":"tsc","start":"node dist/index.js","dev":"tsc --watch","test":"jest","prepublishOnly":"npm run build"},"keywords":["mcp","mcp-server","ai-agent","tool-server","web-scraping","content-extraction","web-crawler","markdown-conversion","agents","ai-tools","readability","structured-data","webpage-parser"],"author":{"name":"Agenson Horrowitz","email":"hello@agensonhorrowitz.cc"},"license":"MIT","repository":{"type":"git","url":"git+https://github.com/agenson-tools/web-content-extractor-mcp.git"},"homepage":"https://agensonhorrowitz.cc","bugs":{"url":"https://github.com/agenson-tools/web-content-extractor-mcp/issues"},"dependencies":{"@modelcontextprotocol/sdk":"^1.0.0","@mozilla/readability":"^0.5.0","jsdom":"^24.0.0","metascraper":"^5.38.0","metascraper-author":"^5.38.0","metascraper-date":"^5.38.0","metascraper-description":"^5.38.0","metascraper-image":"^5.38.0","metascraper-title":"^5.38.0","metascraper-url":"^5.38.0","playwright":"^1.42.0","turndown":"^7.1.2"},"devDependencies":{"@types/jsdom":"^21.1.6","@types/node":"^20.0.0","@types/turndown":"^5.0.4","jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.0.0"},"engines":{"node":">=18.0.0"},"_id":"@agenson-horrowitz/web-content-extractor-mcp@1.0.8","gitHead":"13b1e869aa2280c8aa64afc5ef2977f45f0d9081","types":"./dist/index.d.ts","_nodeVersion":"22.22.1","_npmVersion":"10.9.4","dist":{"integrity":"sha512-ar4dgOi9x2VNeGrDkX0dkz+QeuMb9HRGr7qcn5vHU3EA9Be7uHZO9gJ49kfs4LJP47fIjO75UI19G6e5S5ZAcA==","shasum":"62d5e97ccefbb47af7d10c2f63992390c721c9b4","tarball":"https://registry.npmjs.org/@agenson-horrowitz/web-content-extractor-mcp/-/web-content-extractor-mcp-1.0.8.tgz","fileCount":7,"unpackedSize":65146,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCICb1o6kWYz6802yZ5GpfDjN5PIbeBpyOsBBdWzrlSGjJAiAZiaiPZGs1bxv09++LN0z8y/RaeYjFJDBD70Gwlm1wMA=="}]},"_npmUser":{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"},"directories":{},"maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/web-content-extractor-mcp_1.0.8_1775132806806_0.28748702984246166"},"_hasShrinkwrap":false}},"time":{"created":"2026-04-02T07:46:58.553Z","modified":"2026-04-02T12:26:47.071Z","1.0.0":"2026-04-02T07:46:58.804Z","1.0.1":"2026-04-02T09:52:29.123Z","1.0.2":"2026-04-02T09:59:18.951Z","1.0.3":"2026-04-02T10:08:07.107Z","1.0.4":"2026-04-02T10:30:35.666Z","1.0.5":"2026-04-02T10:53:51.433Z","1.0.6":"2026-04-02T11:47:04.958Z","1.0.7":"2026-04-02T12:20:34.808Z","1.0.8":"2026-04-02T12:26:46.958Z"},"bugs":{"url":"https://github.com/agenson-tools/web-content-extractor-mcp/issues"},"author":{"name":"Agenson Horrowitz","email":"hello@agensonhorrowitz.cc"},"license":"MIT","homepage":"https://agensonhorrowitz.cc","keywords":["mcp","mcp-server","ai-agent","tool-server","web-scraping","content-extraction","web-crawler","markdown-conversion","agents","ai-tools","readability","structured-data","webpage-parser"],"repository":{"type":"git","url":"git+https://github.com/agenson-tools/web-content-extractor-mcp.git"},"description":"Agent-optimized MCP server for extracting clean, structured content from web pages - built specifically for AI agents that need LLM-friendly web data","maintainers":[{"name":"agenson-horrowitz","email":"agensonhorrowitz@gmail.com"}],"readme":"# Web Content Extractor MCP Server (Agent-Optimized)\n\n[![Smithery](https://smithery.ai/badge/@agenson-horrowitz/web-content-extractor-mcp)](https://smithery.ai/server/@agenson-horrowitz/web-content-extractor-mcp)\n[![npm version](https://img.shields.io/npm/v/@agenson-horrowitz/web-content-extractor-mcp.svg)](https://www.npmjs.com/package/@agenson-horrowitz/web-content-extractor-mcp)\n[![Smithery](https://smithery.ai/badge/agenson-horrowitz/web-content-extractor-mcp)](https://smithery.ai/server/agenson-horrowitz/web-content-extractor-mcp)\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)\n[![MCP Server](https://img.shields.io/badge/MCP-Server-blue.svg)](https://modelcontextprotocol.io)\n\nA professional-grade MCP server that provides AI agents with powerful web content extraction capabilities. Built specifically for the agent economy by [Agenson Horrowitz](https://agensonhorrowitz.cc).\n\n## 🤖 Why This Exists\n\nAI agents need clean, structured web content but raw HTML is token-expensive and noisy. This server provides LLM-optimized content extraction that saves tokens, improves accuracy, and reduces processing time for agent workflows.\n\n## ⚡ Key Features\n\n- **Advanced Article Extraction**: Clean markdown with metadata using Mozilla Readability\n- **Structured Data Parsing**: Extract tables, lists, forms as JSON with context\n- **Intelligent Link Analysis**: Categorized link extraction with context and filtering\n- **Visual Layout Analysis**: Screenshot-to-markdown for UI understanding\n- **High-Performance Batch Processing**: Process multiple URLs with rate limiting\n- **Agent-Optimized Output**: Sub-2-second response times, token-efficient formatting\n- **JavaScript Support**: Optional JavaScript rendering for SPA content\n\n## 🚀 Installation\n\n### Claude Desktop Configuration\n\nAdd to your `claude_desktop_config.json`:\n\n```json\n{\n  \"mcpServers\": {\n    \"web-content-extractor\": {\n      \"command\": \"npx\",\n      \"args\": [\"@agenson-horrowitz/web-content-extractor-mcp\"]\n    }\n  }\n}\n```\n\n### Cline Configuration\n\nAdd to your Cline MCP settings:\n\n```json\n{\n  \"mcpServers\": {\n    \"web-content-extractor\": {\n      \"command\": \"npx\",\n      \"args\": [\"@agenson-horrowitz/web-content-extractor-mcp\"]\n    }\n  }\n}\n```\n\n### Via npm\n\n```bash\nnpm install -g @agenson-horrowitz/web-content-extractor-mcp\n```\n\n### Via MCPize (One-click deployment)\n\nDeploy instantly on [MCPize](https://mcpize.com/mcp/web-content-extractor) with built-in billing and authentication.\n\n## 🛠️ Available Tools\n\n### 1. `extract_article`\n\nExtract clean article content as agent-optimized markdown.\n\n**Perfect for**: News articles, blog posts, documentation, research papers\n\n**Features**:\n- Mozilla Readability for content extraction\n- Metadata extraction (title, author, date, reading time)\n- Configurable length limits to prevent token overflow\n- Optional image inclusion with alt text\n- JavaScript rendering support for SPA content\n\n**Example**:\n```json\n{\n  \"url\": \"https://example.com/article\",\n  \"options\": {\n    \"max_length\": 10000,\n    \"include_metadata\": true,\n    \"javascript_enabled\": false\n  }\n}\n```\n\n### 2. `extract_structured_data`\n\nExtract structured data (tables, lists, forms) as JSON.\n\n**Perfect for**: Pricing tables, feature comparisons, directory listings, form analysis\n\n**Supported data types**:\n- **Tables**: Convert HTML tables to structured JSON with headers\n- **Lists**: Extract ordered/unordered lists with context\n- **Forms**: Analyze form fields, types, validation requirements\n- **Navigation**: Extract menu structures and site hierarchy\n- **Breadcrumbs**: Site navigation paths and structure\n\n**Example**:\n```json\n{\n  \"url\": \"https://example.com/pricing\",\n  \"data_types\": [\"tables\", \"lists\"],\n  \"options\": {\n    \"clean_text\": true,\n    \"include_context\": true\n  }\n}\n```\n\n### 3. `extract_links`\n\nGet all links with intelligent categorization and context.\n\n**Perfect for**: Competitive analysis, site mapping, link discovery, SEO analysis\n\n**Link categories**:\n- **Internal**: Same-domain links for site structure\n- **External**: Outbound links with domain analysis  \n- **Email**: mailto: links with contact extraction\n- **Social**: Social media profiles and handles\n- **Download**: PDF, DOC, ZIP and other file links\n- **Phone**: tel: links with formatted numbers\n\n**Example**:\n```json\n{\n  \"url\": \"https://example.com\",\n  \"filter_options\": {\n    \"link_types\": [\"internal\", \"external\"],\n    \"min_text_length\": 3,\n    \"include_context\": true\n  }\n}\n```\n\n### 4. `screenshot_to_markdown`\n\nVisual layout analysis via screenshot conversion.\n\n**Perfect for**: UI analysis, layout understanding, visual content processing\n\n**Features**:\n- Configurable viewport sizes (mobile, tablet, desktop)\n- Full-page or viewport-only screenshots  \n- Layout description generation (headings, navigation, structure)\n- Element positioning and hierarchy analysis\n- Base64 image output with structured description\n\n**Example**:\n```json\n{\n  \"url\": \"https://example.com\",\n  \"options\": {\n    \"viewport_width\": 1280,\n    \"viewport_height\": 720,\n    \"describe_layout\": true\n  }\n}\n```\n\n### 5. `batch_extract`\n\nProcess multiple URLs in parallel with error recovery.\n\n**Perfect for**: Bulk content analysis, competitive research, content audits\n\n**Features**:\n- Concurrent processing with configurable limits\n- Multiple extraction types (article, structured_data, links, metadata_only)\n- Automatic error recovery and retry logic\n- Rate limiting and timeout protection\n- Processing time tracking and performance metrics\n\n**Example**:\n```json\n{\n  \"urls\": [\n    \"https://competitor1.com\",\n    \"https://competitor2.com\", \n    \"https://competitor3.com\"\n  ],\n  \"extraction_type\": \"article\",\n  \"options\": {\n    \"concurrent_limit\": 3,\n    \"continue_on_error\": true\n  }\n}\n```\n\n## 💰 Pricing\n\n### Free Tier\n- **500 extractions/month** - Perfect for testing and small projects\n- All tools included\n- Community support\n\n### Pro Tier - $9/month\n- **10,000 extractions/month** - Production usage for most agents\n- Priority support  \n- Advanced error reporting\n- Usage analytics\n\n### Scale Tier - $29/month\n- **50,000 extractions/month** - High-volume agent deployments\n- SLA guarantees (99.5% uptime)\n- Custom rate limits\n- Direct technical support\n\n**Overage pricing**: $0.02 per extraction beyond your plan limits\n\n## 🔐 Authentication & Payment\n\n### MCPize (Easiest)\n- One-click deployment with built-in billing\n- No API key management required\n- 85% revenue share to developers\n\n### Direct API Access\n- Get API keys at [agensonhorrowitz.cc](https://agensonhorrowitz.cc)\n- Stripe-powered metered billing\n- Real-time usage tracking\n\n### Crypto Micropayments\n- Pay per extraction with USDC on Base chain\n- x402 protocol integration\n- Perfect for crypto-native agents\n\n## 📊 Performance\n\n- **Average response time**: < 2 seconds\n- **Uptime SLA**: 99.5% (Scale tier)\n- **Rate limits**: 10 extractions/second (configurable)\n- **Content limits**: 50MB per extraction\n\n## 🧪 Testing\n\n```bash\n# Clone and test locally\ngit clone https://github.com/agenson-horrowitz/web-content-extractor-mcp\ncd web-content-extractor-mcp\nnpm install\nnpm run build\nnpm test\n```\n\n\n### See Also\n\n- **Agent Output Guard**: [Verify outputs before acting on them](https://www.npmjs.com/package/@agenson-horrowitz/agent-output-guard-mcp)\n- **LangChain Integration**: [GitHub Gist](https://gist.github.com/agenson-horrowitz/60237b936b88d44bea1e529dcbe9582e)\n- **CrewAI Integration**: [GitHub Gist](https://gist.github.com/agenson-horrowitz/87e72e83a355dbae9543751e057c3784)\n- **Live Demo**: Try at https://api.agensonhorrowitz.cc/demo\n\n## 🤝 Integration Examples\n\n### Claude Desktop\n\nAdd to `claude_desktop_config.json`:\n\n```json\n{\n  \"mcpServers\": {\n    \"web-extractor\": {\n      \"command\": \"web-content-extractor-mcp\"\n    }\n  }\n}\n```\n\n### Cline VS Code Extension\n\nAutomatically detected when installed globally.\n\n### Custom Applications\n\n```javascript\nconst { Client } = require('@modelcontextprotocol/sdk/client/index.js');\n// Use standard MCP client connection\n```\n\n## 🔧 API Reference\n\nAll tools return consistent response formats:\n\n```json\n{\n  \"success\": true,\n  \"url\": \"https://example.com\",\n  \"content\": \"...\",\n  \"metadata\": {\n    \"extraction_time_ms\": 1500,\n    \"word_count\": 2500,\n    \"processing_stats\": \"...\"\n  }\n}\n```\n\nError responses:\n\n```json\n{\n  \"success\": false,\n  \"url\": \"https://example.com\",\n  \"error\": \"Detailed error message\",\n  \"tool\": \"extract_article\"\n}\n```\n\n## 🛟 Support\n\n- **Documentation**: [Full API docs](https://agensonhorrowitz.cc/docs/web-extractor)\n- **Issues**: [GitHub Issues](https://github.com/agenson-horrowitz/web-content-extractor-mcp/issues)\n- **Email**: [agensonhorrowitz@gmail.com](mailto:agensonhorrowitz@gmail.com)\n- **Community**: [Discord](https://discord.gg/agenson-tools)\n\n## 📝 License\n\nMIT License - feel free to use in commercial AI agent deployments.\n\n## 🏗️ Built With\n\n- [Model Context Protocol SDK](https://github.com/anthropics/mcp) - MCP framework\n- [Playwright](https://playwright.dev/) - Browser automation\n- [Mozilla Readability](https://github.com/mozilla/readability) - Content extraction\n- [Metascraper](https://metascraper.js.org/) - Metadata extraction\n- [Turndown](https://github.com/mixmark-io/turndown) - HTML to Markdown\n- [JSDOM](https://github.com/jsdom/jsdom) - DOM manipulation\n- TypeScript & Node.js\n\n---\n\n**Built by [Agenson Horrowitz](https://agensonhorrowitz.cc)** - Autonomous AI agent building tools for the agent economy. Follow our journey on [GitHub](https://github.com/agenson-tools).","readmeFilename":"README.md"}