{"_id":"@blackridder22/google-news-scraper","name":"@blackridder22/google-news-scraper","dist-tags":{"latest":"0.0.1"},"versions":{"0.0.1":{"name":"@blackridder22/google-news-scraper","version":"0.0.1","description":"Lightweight async scraper for Google News","main":"./dist/cjs/index.js","module":"./dist/esm/index.mjs","types":"./dist/tsc/index.d.ts","exports":{"import":"./dist/esm/index.mjs","require":"./dist/cjs/index.js"},"scripts":{"test":"jest","build:ts":"tsc","build:rollup":"rollup -c","build":"rm -rf ./dist/* && npm run build:ts && npm run build:rollup","build:ci":"rm -rf ./dist/* && npm run build:ts && rollup -c --bundleConfigAsCjs"},"repository":{"type":"git","url":"git+https://github.com/blackridder22/google-news-scraper.git"},"keywords":["news","scraper","google","google news","news scraper","crawler","web crawler","news crawler","google crawler"],"author":{"name":"blackridder22"},"license":"SEE LICENSE IN LICENSE.txt","dependencies":{"@mozilla/readability":"^0.5.0","@rollup/plugin-babel":"^6.0.4","cheerio":"^1.0.0-rc.3","jsdom":"^25.0.0","puppeteer":"^23.6.0","winston":"^3.15.0"},"devDependencies":{"@babel/core":"^7.23.9","@babel/plugin-transform-runtime":"^7.23.9","@babel/preset-env":"^7.23.9","@rollup/plugin-commonjs":"^28.0.1","@rollup/plugin-json":"^6.1.0","@rollup/plugin-node-resolve":"^15.3.0","@rollup/plugin-terser":"^0.4.4","@rollup/plugin-typescript":"^12.1.1","@semantic-release/changelog":"^6.0.3","@semantic-release/git":"^10.0.1","@semantic-release/npm":"^12.0.1","@types/jsdom":"^21.1.7","@types/winston":"^2.4.4","babel-jest":"^29.7.0","expect.js":"^0.3.1","jest":"^29.7.0","rollup":"^4.24.0","rollup-plugin-delete":"^2.1.0","semantic-release":"^24.2.0","typescript":"^5.6.3"},"optionalDependencies":{"@rollup/rollup-linux-x64-gnu":"4.24.2"},"_id":"@blackridder22/google-news-scraper@0.0.1","gitHead":"556361e7963128ecfcd95e569241bd228e318d9e","bugs":{"url":"https://github.com/blackridder22/google-news-scraper/issues"},"homepage":"https://github.com/blackridder22/google-news-scraper#readme","_nodeVersion":"22.21.1","_npmVersion":"10.9.4","dist":{"integrity":"sha512-dO9dyQFV+7nR2FJ9USTidJofVvDpvN5bLx/FiYioZb25jRz9dobqBeP8bapV4QnDcyPzYjqmp2gpGJx5gqahtA==","shasum":"65c545b4d50e45927c62688c241a314a19657118","tarball":"https://registry.npmjs.org/@blackridder22/google-news-scraper/-/google-news-scraper-0.0.1.tgz","fileCount":20,"unpackedSize":798593,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQDjE9/hyAqEwlBIXmlm5oJAFWrGSwzCMkS3bUiMOsU61wIhAJUBMpK2RBKkTNfCjKhPfiRFXsqSRp0gjJixVAruFw7y"}]},"_npmUser":{"name":"blackridder22","email":"blackridder22@gmail.com"},"directories":{},"maintainers":[{"name":"blackridder22","email":"blackridder22@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/google-news-scraper_0.0.1_1766384431123_0.4865473217067462"},"_hasShrinkwrap":false}},"time":{"created":"2025-12-22T06:20:31.042Z","0.0.1":"2025-12-22T06:20:31.346Z","modified":"2025-12-22T06:20:31.612Z"},"maintainers":[{"name":"blackridder22","email":"blackridder22@gmail.com"}],"description":"Lightweight async scraper for Google News","homepage":"https://github.com/blackridder22/google-news-scraper#readme","keywords":["news","scraper","google","google news","news scraper","crawler","web crawler","news crawler","google crawler"],"repository":{"type":"git","url":"git+https://github.com/blackridder22/google-news-scraper.git"},"author":{"name":"blackridder22"},"bugs":{"url":"https://github.com/blackridder22/google-news-scraper/issues"},"license":"SEE LICENSE IN LICENSE.txt","readme":"# Google News Scraper 📰\n\nA lightweight, asynchronous scraper for Google News that retrieves articles, resolves redirects, and ensures clean data.\n\n## Features\n\n- **🔍 Search & Topics**: Scrape by search term or topic URL.\n- **🔗 Smart Redirect Resolution**: Automatically resolves Google's \"ugly\" tracking URLs to the direct publisher links.\n- **🖼️ High-Quality Images**: Extracts high-resolution images (`og:image`) from the source article, replacing low-quality Google thumbnails.\n- **🧹 Auto-Filtering**: Optional strict filtering to ensure you only get data with resolved URLs and clean images.\n- **⏱️ Timeframe Support**: Filter news by hours, days, years (e.g., `1h`, `7d`, `1y`).\n\n## Installation\n\n```bash\nnpm install @blackridder22/google-news-scraper\n```\n\n## Quick Start\n\n```javascript\nconst googleNewsScraper = require('@blackridder22/google-news-scraper');\n\n(async () => {\n    const articles = await googleNewsScraper({\n        searchTerm: \"Artificial Intelligence\",\n        prettyURLs: true,\n        timeframe: \"1d\",\n        filter: true, // Only return articles with resolved links and images\n        puppeteerArgs: ['--no-sandbox']\n    });\n\n    console.log(articles);\n})();\n```\n\n## Configuration Options\n\nThe function accepts a configuration object with the following properties:\n\n| Property | Type | Default | Description |\n|----------|------|---------|-------------|\n| `searchTerm` | `string` | `null` | The search query (e.g., \"Crypto\"). |\n| `baseUrl` | `string` | `...` | Alternate base URL (e.g., for specific topic pages). |\n| `prettyURLs` | `boolean` | `true` | Resolve Google redirects to actual publisher URLs. |\n| `filter` | `boolean` | `false` | **New!** If `true`, removes any article where the URL or Image could not be resolved (i.e., still points to `news.google.com`). |\n| `timeframe` | `string` | `7d` | Filter by age: `h` (hours), `d` (days), `y` (years). Example: `12h`. |\n| `puppeteerArgs` | `array` | `[]` | Additional flags for Puppeteer (e.g., `['--no-sandbox']`). |\n| `limit` | `number` | `null` | Limit the number of results returned. |\n| `getArticleContent`| `boolean`| `false` | Experimental: Attempts to fetch full article text (slow). |\n\n## Output Format\n\nReturns an array of article objects:\n\n```json\n[\n  {\n    \"title\": \"Example News Title\",\n    \"link\": \"https://www.nytimes.com/...\",\n    \"image\": \"https://www.nytimes.com/images/...\",\n    \"source\": \"New York Times\",\n    \"datetime\": \"2025-12-22T10:00:00.000Z\",\n    \"time\": \"2 hours ago\",\n    \"articleType\": \"regular\"\n  }\n]\n```\n\n## Why use `filter: true`?\n\nGoogle News provides \"tracking\" URLs and internal thumbnail images (`news.google.com/api/attachments/...`).\n- **Without Filter**: You get 100% of results, but some may have ugly URLs or protected images.\n- **With Filter**: The scraper verifies everything. If it can't resolve the redirect or find a high-quality `og:image` on the publisher's site, it drops that result. You get fewer results, but they are guaranteed to be \"clean\".\n\n## License\n\nMIT\n\n","readmeFilename":"README.md","_rev":"1-211eb3f1fccd5fcaf7729511b85b06ec"}