{"_id":"@ducksguse/news-scraper-sdk","name":"@ducksguse/news-scraper-sdk","dist-tags":{"latest":"1.0.0"},"versions":{"1.0.0":{"name":"@ducksguse/news-scraper-sdk","version":"1.0.0","description":"Official Goose-Powered News Scraper SDK. Honk!","main":"dist/index.js","types":"dist/index.d.ts","scripts":{"build":"tsc","prepublishOnly":"npm run build"},"repository":{"type":"git","url":"git+https://github.com/ducksguse/news-scraper-sdk.git"},"keywords":["goose","news","scraper","honk","ducksguse","news-intelligence","web-scraping"],"author":{"name":"DucksGuse Team"},"license":"ISC","dependencies":{"axios":"^1.6.0"},"devDependencies":{"@types/node":"^20.0.0","typescript":"^5.0.0"},"_id":"@ducksguse/news-scraper-sdk@1.0.0","bugs":{"url":"https://github.com/ducksguse/news-scraper-sdk/issues"},"homepage":"https://github.com/ducksguse/news-scraper-sdk#readme","_nodeVersion":"20.19.4","_npmVersion":"10.8.2","dist":{"integrity":"sha512-xcFTHFRBS4zkEnCLlyhByGMeIiCM4Do9c9vphRwv9kgEPvEj43d4umng+0igwv5Gggx+sjpOhfdEunh0EehEeg==","shasum":"542fe99981b38baae85d7faae74317967a8f19ce","tarball":"https://registry.npmjs.org/@ducksguse/news-scraper-sdk/-/news-scraper-sdk-1.0.0.tgz","fileCount":5,"unpackedSize":9166,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQDHTj4nff98pHMoI4eFe3YfmJ9uuOE7Wzj6Cx24dJOc6wIgeVAYBf6tAr6BfUm6yJYGYj6s1aQXg2f9pAhHDCcnTYw="}]},"_npmUser":{"name":"is87","email":"isaevslava87@gmail.com"},"directories":{},"maintainers":[{"name":"is87","email":"isaevslava87@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/news-scraper-sdk_1.0.0_1763946147527_0.6916032165242543"},"_hasShrinkwrap":false}},"time":{"created":"2025-11-24T01:02:27.446Z","1.0.0":"2025-11-24T01:02:27.718Z","modified":"2025-11-24T01:02:28.006Z"},"maintainers":[{"name":"is87","email":"isaevslava87@gmail.com"}],"description":"Official Goose-Powered News Scraper SDK. Honk!","homepage":"https://github.com/ducksguse/news-scraper-sdk#readme","keywords":["goose","news","scraper","honk","ducksguse","news-intelligence","web-scraping"],"repository":{"type":"git","url":"git+https://github.com/ducksguse/news-scraper-sdk.git"},"author":{"name":"DucksGuse Team"},"bugs":{"url":"https://github.com/ducksguse/news-scraper-sdk/issues"},"license":"ISC","readme":"# 🦆 @ducksguse/news-scraper-sdk\n\n**The Official Goose-Powered News Scraper.**\n*Honk-driven intelligence for your application.*\n\nThis SDK allows you to unleash our digital geese to monitor the web, scrape news, and bring you golden eggs (clusters of data).\n\n## Features\n\n*   **Goose Vision**: Track news by RSS feeds or **Google News** topics.\n*   **Flock Clustering**: Automatically groups related articles so you don't hear the same honk twice.\n*   **Noise Filtering**: Use `negative_keywords` to hiss at bad articles.\n*   **Instant Flight (Backfill)**: The goose flies back in time (3 days) immediately after you give it a task.\n*   **Webhooks**: Get a \"HONK!\" notification when new data arrives.\n\n## Installation\n\n```bash\nnpm install @ducksguse/news-scraper-sdk\n# or\nyarn add @ducksguse/news-scraper-sdk\n```\n\n## Quick Start\n\n```typescript\nimport { GooseNewsClient } from '@ducksguse/news-scraper-sdk';\n\n// Initialize the Goose\nconst goose = new GooseNewsClient(\n  'https://your-news-service-url.com/api/v1',\n  'YOUR_API_KEY'\n);\n\n// Optional: Check if goose is awake\ngoose.honk(); // Output: \"HONK! The goose is ready to scrape.\"\n\nasync function main() {\n  // 1. Give the Goose a Mission (Create Profile)\n  const profile = await goose.createProfile({\n    name: \"Construction Projects - Texas\",\n    description: \"Tracking new hotel construction and renovations in Texas\",\n    sources: [\n      {\n        type: \"google_news\",\n        query: \"hotel construction Texas\",\n        language: \"en\"\n      }\n    ],\n    // Smart Filters\n    filters: {\n      green_flags: {\n        \"intent\": [\"breaking ground\", \"permit approved\"]\n      },\n      // Hiss at these words (exclude them)\n      negative_keywords: [\"website redesign\", \"digital transformation\"],\n      min_relevance_score: 0.75\n    },\n    extraction_schema: {\n      \"project_name\": \"string\",\n      \"budget\": \"string\"\n    },\n    schedule: {\n      initial_lookback_hours: 72, // Fly back 3 days immediately\n      check_interval_hours: 4     // Check again every 4 hours\n    }\n  });\n\n  console.log(`Mission accepted! Profile ID: ${profile.id}`);\n  console.log(\"The goose has taken flight... 🪿\");\n}\n\nmain();\n```\n\n## Core Concepts\n\n### Watch Profiles (Missions)\nA **Watch Profile** is a mission you give to the goose. It tells the goose where to look (Sources) and what to ignore (Negative Keywords).\n\n### Clusters (Golden Eggs)\nThe goose doesn't just bring you random sticks. It groups related articles into **Clusters**. If 10 sources write about the same event, you get 1 Cluster.\n\n### Backfill (Time Travel)\nWhen you create a profile, the goose immediately performs a **Backfill**. It scrapes the history (default: 72 hours) so you get data instantly.\n\n## API Reference\n\n### `goose.createProfile(data)`\nCreates a new mission. Triggers immediate backfill.\n\n### `goose.getProfiles()`\nLists all active missions.\n\n### `goose.getProfileClusters(profileId)`\nGets the latest golden eggs (clusters) for a specific mission.\n\n### `goose.honk()`\nVerifies the client is initialized. HONK!\n\n---\n\n## License\nISC - Made with ❤️ and 🪿 by DucksGuse.\n","readmeFilename":"README.md","_rev":"1-9a23745747fb8591fdfb6109af3d9c35"}