{"_id":"@bpinternal/site-scout","_rev":"3-e24a673a4d7d11cc7f214586117a947c","name":"@bpinternal/site-scout","dist-tags":{"latest":"0.1.0"},"versions":{"0.1.0":{"name":"@bpinternal/site-scout","version":"0.1.0","_id":"@bpinternal/site-scout@0.1.0","maintainers":[{"name":"allardy","email":"yann.allard1@gmail.com"},{"name":"michael.masson","email":"michael.masson@botpress.com"},{"name":"franklevasseur","email":"francois_levasseur@hotmail.com"},{"name":"slvnperron","email":"slvn.perron@gmail.com"},{"name":"botpress-cloud-ops","email":"cloud-ops@botpress.com"}],"homepage":"https://github.com/botpress/genisys#readme","bugs":{"url":"https://github.com/botpress/genisys/issues"},"dist":{"shasum":"945c19b1d07bd8ef4652c2c58897c2fa519f7d10","tarball":"https://registry.npmjs.org/@bpinternal/site-scout/-/site-scout-0.1.0.tgz","fileCount":8,"integrity":"sha512-AAN4XOdR/P5CuBUMHeV0xEmnsA3jzMKfSWy4EINMB7iFmGGLo2ItKWRhVjdhq+Zbc/7LQ3ZqLpSZKuF9JGp+pQ==","signatures":[{"sig":"MEUCIQCK/wq90THBvIB0k7R9e31ABWxo4ucLiYQhN5wD6V6WuwIgKZaNzGOaA7K6d9shWzo7JPd/piNDmEH06JYrZSQbbVY=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":398141},"main":"./dist/index.cjs","type":"module","types":"./dist/index.d.ts","module":"./dist/index.js","source":"./src/index.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js","require":"./dist/index.cjs"}},"gitHead":"3e6c6fb3906c366691a8c5a77cc1bc549b3943be","scripts":{"dev":"tsup --watch","test":"vitest run","build":"tsup","fix:lint":"eslint --fix .","check:lint":"eslint .","check:type":"tsc --noEmit","test:watch":"vitest ."},"_npmUser":{"name":"botpress-cloud-ops","email":"cloud-ops@botpress.com"},"repository":{"url":"git+https://github.com/botpress/genisys.git"},"_npmVersion":"10.9.2","description":"Website URL discovery + LLM page selection — scouts a site for the pages worth reading (seeds support-bot knowledge bases).","directories":{},"sideEffects":false,"_nodeVersion":"22.13.1","dependencies":{"robots-parser":"^3.0.1","fast-xml-parser":"^5.9.3"},"_hasShrinkwrap":false,"devDependencies":{"tsup":"^8.1.0","vitest":"1.6.0","typescript":"5.9.3","@types/node":"^22.7.4","@bpinternal/zui":"~2.3.0","@repo/eslint-config":"workspace:*","@repo/typescript-config":"workspace:*"},"peerDependencies":{"@bpinternal/zui":">=2.1.0 <3.0.0"},"_npmOperationalInternal":{"tmp":"tmp/site-scout_0.1.0_1783544690109_0.016419110894059452","host":"s3://npm-registry-packages-npm-production"}}},"time":{"created":"2026-07-08T21:04:49.995Z","modified":"2026-09-02T19:11:34.881Z","0.1.0":"2026-07-08T21:04:50.309Z"},"bugs":{"url":"https://github.com/botpress/genisys/issues"},"homepage":"https://github.com/botpress/genisys#readme","repository":{"url":"git+https://github.com/botpress/genisys.git"},"description":"Website URL discovery + LLM page selection — scouts a site for the pages worth reading (seeds support-bot knowledge bases).","maintainers":[{"email":"yann.allard1@gmail.com","name":"allardy"},{"email":"michael.masson@botpress.com","name":"michael.masson"},{"email":"slvn.perron@gmail.com","name":"slvnperron"},{"email":"xavier.hamel.protic@gmail.com","name":"xavierhamel"},{"email":"cloud-ops@botpress.com","name":"botpress-cloud-ops"}],"readme":"# @bpinternal/site-scout\n\nWebsite URL discovery + LLM page selection — scouts a site for the pages worth\nreading. Given a website, it finds candidate URLs, scores them into a\nprioritized tree, then makes **one** injected-LLM call to pick the pages best\nsuited to seed a support bot's knowledge base.\n\n## What it does\n\nTwo independent units, composed:\n\n```\ndiscover(website, opts, deps) -> MasterMap   (prioritized tree)\nselect(map, opts, deps)       -> Selection   (best ≤limit URLs)\n```\n\n`findKnowledgeUrls()` runs guard → discover → select for the common case.\n\n- **Discovery** pulls candidate URLs from **5 sources** — `llms.txt` /\n  `llms-full.txt`, `sitemap.xml` (incl. children), `robots.txt`, the host's\n  `discoverUrls` crawl, and a `site:` web search — normalizes/dedupes them, and\n  builds a **score tree**: one node per path segment, with aggregate scores and\n  leaf counts so selection can reason about cardinality without walking every\n  leaf.\n- **Selection** renders the tree as a compact numbered list and makes **one\n  LLM call** (`extract`) that returns the picked page _numbers_ (best-first),\n  which map back to URLs. The LLM pick is the decision — no deterministic\n  fallback list.\n\n## The injection contract (`SiteScoutDeps`)\n\nsite-scout does no I/O and never imports a runtime itself. Every consumer\nprovides the I/O + LLM as `deps`:\n\n```ts\ntype SiteScoutDeps = {\n  // discovery I/O\n  fetchText(url, baseHost, timeoutMs?): Promise<string | null>\n  fetchJson(url, baseHost, timeoutMs?): Promise<T | null>\n  discoverUrls(args): Promise<{ urls: string[]; stopReason: string }>\n  webSearch(args): Promise<{ results: SearchResult[] }>\n  // the one LLM touchpoint (zai-shaped extract)\n  extract(input, schema, { instructions }): Promise<z.infer<schema>>\n  // optional record/replay cache around the LLM call (defaults to passthrough)\n  cache?(kind, keyParts, fn): Promise<T>\n}\n```\n\n`@bpinternal/zui` is a **peer dependency** — consumers bring their own\n(the schema passed to `extract` uses it).\n\n## Usage\n\n```ts\nimport { findKnowledgeUrls } from '@bpinternal/site-scout'\n\nconst result = await findKnowledgeUrls(\n  {\n    website: 'https://example.com',\n    limit: 50,\n    topics: ['pricing', 'docs'],\n    prompt: 'Prioritize developer docs, API reference, pricing.',\n    context: { company: 'Example', website: 'https://example.com', overview: '…' },\n  },\n  deps // your SiteScoutDeps implementation\n)\n// result.urls        — prioritized URLs, best first\n// result.discovered  — total distinct URLs found before selection\n// result.stopReason  — 'ok' | 'unsupported_site' | 'no_sources' | 'time_limit_reached'\n```\n\n`discover()` and `select()` are also exported for callers that want to show the\ntree, let a user refine, then select. `fetchText`/`fetchJson`/`hostOf`/\n`siteOrigin` (an SSRF-guarded fetcher + URL helpers) are re-exported so a\nconsumer can reuse them when wiring its own deps.\n\n## Record / replay fixtures\n\nThe tests are offline and deterministic:\n\n- **Network** — discovery I/O is wrapped by `cachedDeps(group, realDeps)`, which\n  records/replays each call against `src/__fixtures__/<group>.jsonl` (one file\n  per site). In replay mode a miss throws (a test can't silently hit the\n  network); in production `cachedDeps` is a passthrough. Mode is env-driven\n  (`URL_FIXTURE_RECORD` to record; `NODE_ENV=test` / `VITEST` / `BUN_TEST` /\n  `URL_FIXTURE_REPLAY` to replay).\n- **LLM** — the `select` call is keyed and replayed from\n  `src/__fixtures__/llm-cache.jsonl`. The package e2e (`index.e2e.test.ts`)\n  loads that file and serves the recorded decision, so no live model is called.\n\nFixtures are **recorded via the app harness** (viber-compiler-bot), which owns\nthe real runtime/LLM wiring; this package only replays them.\n","readmeFilename":"README.md"}