{"_id":"@autonome-research/fleet-router","name":"@autonome-research/fleet-router","dist-tags":{"latest":"0.2.0"},"versions":{"0.2.0":{"name":"@autonome-research/fleet-router","version":"0.2.0","description":"Client-side request routing across heterogeneous local LLM inference endpoints: context-window eligibility, capacity-normalized least-loaded selection, health probing. Zero dependencies.","license":"MIT","type":"module","main":"./dist/index.js","types":"./dist/index.d.ts","exports":{".":"./dist/index.js","./package.json":"./package.json"},"engines":{"node":">=20"},"scripts":{"build":"tsc -p tsconfig.json","typecheck":"tsc --noEmit","test":"vitest run"},"devDependencies":{"@types/node":"^26.1.1","typescript":"^5.6.0","vitest":"^3.0.0"},"keywords":["llm","vllm","router","load-balancing","openai-compatible","inference","local"],"publishConfig":{"access":"public"},"repository":{"type":"git","url":"git+https://github.com/autonome-research/fleet-router.git"},"bugs":{"url":"https://github.com/autonome-research/fleet-router/issues"},"gitHead":"bd58a0f2bb58d1102056d5585576b4eb3d92c7dd","_id":"@autonome-research/fleet-router@0.2.0","homepage":"https://github.com/autonome-research/fleet-router#readme","_nodeVersion":"26.1.0","_npmVersion":"11.15.0","dist":{"integrity":"sha512-mKCCrq5HEmDYAZ38mbxAxJz+AKWnMxcFGdJBm26SXJj9LDO6OIQOmFcI/SI86JtJo+m4T/WchYZlQP+IJ8Eqtw==","shasum":"0333f7853206636ca09debfa07ac9c832d67864a","tarball":"https://registry.npmjs.org/@autonome-research/fleet-router/-/fleet-router-0.2.0.tgz","fileCount":9,"unpackedSize":25736,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQCGbbX8nkd0W35GW3N1sJp7bIPhmtFmuzucdI6OjgxPgQIgb+h7rlRT518j0s+7mGnLmzrwCOYF9rDWwcE+i/Ppoh4="}]},"_npmUser":{"name":"velmoon222","email":"velvetmoon222999@gmail.com"},"directories":{},"maintainers":[{"name":"velmoon222","email":"velvetmoon222999@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/fleet-router_0.2.0_1784324383603_0.05391712050944486"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-17T21:39:43.456Z","0.2.0":"2026-07-17T21:39:43.759Z","modified":"2026-07-17T21:39:43.966Z"},"maintainers":[{"name":"velmoon222","email":"velvetmoon222999@gmail.com"}],"description":"Client-side request routing across heterogeneous local LLM inference endpoints: context-window eligibility, capacity-normalized least-loaded selection, health probing. Zero dependencies.","homepage":"https://github.com/autonome-research/fleet-router#readme","keywords":["llm","vllm","router","load-balancing","openai-compatible","inference","local"],"repository":{"type":"git","url":"git+https://github.com/autonome-research/fleet-router.git"},"bugs":{"url":"https://github.com/autonome-research/fleet-router/issues"},"license":"MIT","readme":"# fleet-router\n\nClient-side request routing across **heterogeneous local LLM inference\nendpoints** — the case where your instances differ in context window, output\nbudget, multimodal overhead, and concurrency capacity, and your requests\ndiffer enough in size that \"send it anywhere\" produces 400s and queues.\n\nExtracted from a production batch pipeline (305-session litigation\ntranscription + vision-description corpus) that ran a Qwen3-VL fleet across\ntwo 98 GB cards (262k ctx, 16 seats each) and one 32 GB card (49k ctx,\n2 seats).\n\n## Why not an existing tool?\n\nA mid-2026 survey of the space (LiteLLM Router, vLLM production-stack,\nSGLang router, Ray Serve, Envoy/Kong/Portkey gateways, llm-d/KServe,\nGPUStack, Paddler, npm ecosystem) found **no tool that combines**:\n\n1. **Token-estimated context eligibility** — per request, only endpoints\n   whose window fits (prompt + output reserve + per-endpoint overhead like\n   video-frame tokens) are candidates. Gateways rate-limit on tokens; none\n   *place* on tokens.\n2. **Capacity-normalized selection** — lowest `inflight/capacity`, so a\n   16-seat instance absorbs 8× the traffic of a 2-seat one. LiteLLM's\n   least-busy uses raw counts; everything else is static weights.\n3. **Embedded TypeScript, zero infra** — a library in your pipeline process,\n   not a Python sidecar (+Redis) or a k8s gateway data path.\n4. **Local-only** — no cloud calls, suitable for confidential workloads.\n\nThe industry's routing investment is going the other way (KV-cache-aware\ngateways for homogeneous replica pools), so this niche is likely to stay\nopen.\n\n## Usage\n\n```ts\nimport { FleetRouter, waitHealthy } from 'fleet-router';\n\nconst router = new FleetRouter([\n  { url: 'http://localhost:30010/v1', maxContext: 262_144, capacity: 16, outputReserve: 16_384, overheadTokens: 40_000 },\n  { url: 'http://localhost:30011/v1', maxContext: 262_144, capacity: 16, outputReserve: 16_384, overheadTokens: 40_000 },\n  { url: 'http://localhost:30009/v1', maxContext: 49_152, capacity: 3, outputReserve: 8_192, overheadTokens: 19_000 },\n]);\n\nconst result = await router.withEndpoint({ promptChars: transcript.length }, async (lease) => {\n  const prompt = transcript.length > lease.promptBudgetChars\n    ? transcript.slice(0, lease.promptBudgetChars) + '\\n[...truncated...]'\n    : transcript;\n  return callOpenAIChat(lease.endpoint.url, { prompt, max_tokens: lease.maxTokens });\n});\n```\n\nOn a connection error (engine mid-restart under k8s), `await\nwaitHealthy(lease.endpoint.url)` then retry.\n\n## Status\n\nv0.2.0 (unpublished; name and home TBD). The full three-axis surface from\nDESIGN.md is implemented and tested: outcome-fed circuit breaker\n(failureThreshold/cooldownMs), failover in `withEndpoint` (exclusion-set\nretries), blocking `acquire` with per-endpoint `hardCap` and AbortSignal,\ncapability `tags`/`require`, rendezvous-hash `affinityKey` with spill-over,\nand pluggable `estimateTokens`/`score`. 18 unit tests.\n\nDeferred by design: latency-EWMA selection (needs collected outcome timing),\ncost weighting, cross-process state. See DESIGN.md for rationale.\n","readmeFilename":"README.md","_rev":"1-07ff5dbcdf0a2fad7dc0a7eaebcd8e30"}