{"_id":"@eleata/resilient-llm-router","_rev":"2-f0188fa02d66a006e50f9236617786df","name":"@eleata/resilient-llm-router","dist-tags":{"latest":"0.2.0-alpha.0"},"versions":{"0.1.0-alpha.0":{"name":"@eleata/resilient-llm-router","version":"0.1.0-alpha.0","keywords":["llm","router","resilience","circuit-breaker","rate-limit","quota","openai","anthropic","groq","cerebras"],"author":{"name":"Hernán Inverso","email":"hernaninverso@gmail.com"},"license":"MIT","_id":"@eleata/resilient-llm-router@0.1.0-alpha.0","maintainers":[{"name":"eleata","email":"hernaninverso@gmail.com"}],"homepage":"https://github.com/eleata/resilient-llm-router-ts#readme","bugs":{"url":"https://github.com/eleata/resilient-llm-router-ts/issues"},"dist":{"shasum":"082f99901dc54efdbf6a9b0eb446199ff9f15011","tarball":"https://registry.npmjs.org/@eleata/resilient-llm-router/-/resilient-llm-router-0.1.0-alpha.0.tgz","fileCount":11,"integrity":"sha512-bUdifVSM0xgjicmKsGsyyA9OdLiMYI6fEugdvtWloudvt0TlNqBS61jo6UhYJfvpIUh5k7zH+OPnKXLLL6IUpQ==","signatures":[{"sig":"MEUCIGkZ3h0ooxtK3kb1ULL3pLz+Pj9qzfmfc8EKwvrz1bCWAiEA50M0+gdFjWxgxVpU7Sr1SviQo9u34ZuJfdzdIA8Uuew=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":46351},"main":"dist/index.js","types":"dist/index.d.ts","engines":{"node":">=18"},"gitHead":"b68008ff268a96d6b469935e9eff4c24933c0923","scripts":{"test":"jest","build":"tsc","test:watch":"jest --watch","prepublishOnly":"npm run build && npm test"},"_npmUser":{"name":"eleata","email":"hernaninverso@gmail.com"},"repository":{"url":"git+https://github.com/eleata/resilient-llm-router-ts.git","type":"git"},"_npmVersion":"10.9.4","description":"Multi-provider LLM routing with 3 orthogonal resilience states: rate-limit ≠ quota-exhausted ≠ circuit-broken. TypeScript port of the Python resilient-llm-router.","directories":{},"_nodeVersion":"22.22.1","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.3.0","@types/jest":"^29.5.0","@types/node":"^20.0.0"},"_npmOperationalInternal":{"tmp":"tmp/resilient-llm-router_0.1.0-alpha.0_1778442579956_0.8150140121419505","host":"s3://npm-registry-packages-npm-production"}},"0.2.0-alpha.0":{"name":"@eleata/resilient-llm-router","version":"0.2.0-alpha.0","description":"Multi-provider LLM routing with 3 orthogonal resilience states: rate-limit ≠ quota-exhausted ≠ circuit-broken. TypeScript port of the Python resilient-llm-router.","main":"dist/index.js","types":"dist/index.d.ts","scripts":{"build":"tsc","test":"jest","test:watch":"jest --watch","prepublishOnly":"npm run build && npm test"},"keywords":["llm","router","resilience","circuit-breaker","rate-limit","quota","openai","anthropic","groq","cerebras"],"license":"MIT","author":{"name":"Hernán Inverso","email":"hernaninverso@gmail.com"},"repository":{"type":"git","url":"git+https://github.com/eleata/resilient-llm-router-ts.git"},"homepage":"https://github.com/eleata/resilient-llm-router-ts#readme","bugs":{"url":"https://github.com/eleata/resilient-llm-router-ts/issues"},"engines":{"node":">=18"},"publishConfig":{"access":"public"},"devDependencies":{"@types/better-sqlite3":"^7.6.13","@types/jest":"^29.5.0","@types/node":"^20.0.0","jest":"^29.7.0","ts-jest":"^29.1.0","typescript":"^5.3.0"},"dependencies":{},"optionalDependencies":{"better-sqlite3":"^12.9.0"},"_id":"@eleata/resilient-llm-router@0.2.0-alpha.0","gitHead":"729f8b7f06df86d92b74cce538284a3397384ebc","_nodeVersion":"22.22.1","_npmVersion":"10.9.4","dist":{"integrity":"sha512-MLS5L1bGrmQnPqpvzsM2gQJ5RBY+MQL9ObEjJ0Oc8PbuNVTJoBoEjdIbEHD695eul21dLhga4dmsLQJt+aMroA==","shasum":"61fd619f3b3553879e059c51afedb79203c49641","tarball":"https://registry.npmjs.org/@eleata/resilient-llm-router/-/resilient-llm-router-0.2.0-alpha.0.tgz","fileCount":13,"unpackedSize":70827,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQD++AnSxv/7Wcu3loi05VXC8i/wkCNkWaMkMkIYQgSDNwIgHsW7cQ/flVq+rgkGtayhph2hI1F/k0XQWzbTY2/7OKs="}]},"_npmUser":{"name":"eleata","email":"hernaninverso@gmail.com"},"directories":{},"maintainers":[{"name":"eleata","email":"hernaninverso@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/resilient-llm-router_0.2.0-alpha.0_1778444021489_0.7038071994259751"},"_hasShrinkwrap":false}},"time":{"created":"2026-05-10T19:49:39.837Z","modified":"2026-05-10T20:13:41.732Z","0.1.0-alpha.0":"2026-05-10T19:49:40.100Z","0.2.0-alpha.0":"2026-05-10T20:13:41.634Z"},"bugs":{"url":"https://github.com/eleata/resilient-llm-router-ts/issues"},"author":{"name":"Hernán Inverso","email":"hernaninverso@gmail.com"},"license":"MIT","homepage":"https://github.com/eleata/resilient-llm-router-ts#readme","keywords":["llm","router","resilience","circuit-breaker","rate-limit","quota","openai","anthropic","groq","cerebras"],"repository":{"type":"git","url":"git+https://github.com/eleata/resilient-llm-router-ts.git"},"description":"Multi-provider LLM routing with 3 orthogonal resilience states: rate-limit ≠ quota-exhausted ≠ circuit-broken. TypeScript port of the Python resilient-llm-router.","maintainers":[{"name":"eleata","email":"hernaninverso@gmail.com"}],"readme":"# @eleata/resilient-llm-router\n\nMulti-provider LLM routing with **3 orthogonal resilience states**:\n\n- **rate_limit** — short-period back-off (per-minute / per-day request and token windows). TTL from `Retry-After` headers.\n- **quota_exhausted** — long-period cap hit (daily / monthly). `exhausted=true` only when the error body explicitly says so. Period rolls forward automatically when the window ends.\n- **circuit** — transient health (`closed` / `open` / `half_open`) from 5xx and timeouts.\n\n> A 429 with body `\"You exceeded your daily limit\"` and a 429 with `Retry-After: 60` are semantically different. Most routers treat them the same and waste hundreds of calls retrying every minute against an exhausted free-tier provider. This library separates the two so each gets the cooldown it deserves.\n\nThis is the TypeScript port of [`resilient-llm-router`](https://github.com/eleata/resilient-llm-router) (Python).\n\n> **Status: alpha (`0.1.0-alpha.0`)**. API surface is stable for the in-memory backend. Persistence backends (SQLite, Postgres) and the `probes` health helper from the Python lib are deferred to `0.2.x`.\n\n## Install\n\n```bash\nnpm install @eleata/resilient-llm-router\n```\n\nRequires Node ≥ 18. Zero runtime dependencies.\n\n## Quick start\n\n```ts\nimport { router, parseHeaders } from \"@eleata/resilient-llm-router\";\n\nconst r = router(); // in-memory state by default\n\n// Optional: seed quota caps so guard() can predict near-cap throttling\nawait r.seedCaps({\n  groq: { \"llama-3.3-70b-versatile\": { \"daily/tokens\": 14_400_000 } },\n});\n\n// Before every call: ask should we even try?\nconst decision = await r.guard({\n  provider: \"groq\",\n  model: \"llama-3.3-70b-versatile\",\n  estimatedInputTokens: 800,\n  estimatedOutputTokens: 200,\n});\n\nif (!decision.allow) {\n  console.log(`skip: ${decision.reason} (retry in ${decision.ttlS}s)`);\n  // try next provider / model\n} else {\n  // make the actual call...\n  const resp = await yourLLMClient.chat({ ... });\n\n  // Tell the router how it went so it can update state.\n  await r.recordOutcome({\n    provider: \"groq\",\n    model: \"llama-3.3-70b-versatile\",\n    success: true,\n    tokensIn: 812,\n    tokensOut: 204,\n    responseHeaders: resp.headers, // parses Retry-After + dialect-specific rate-limit headers\n  });\n}\n```\n\n## Failure handling\n\n```ts\n// 429 with no specific body → rate-limit, blocked for the Retry-After window\nawait r.recordOutcome({\n  provider: \"groq\",\n  model: \"llama-3.3-70b-versatile\",\n  success: false,\n  statusCode: 429,\n  errorMessage: \"Too many requests\",\n  retryAfterSeconds: 60,\n});\n\n// 429 with quota body → quota_exhausted, blocked until period rolls over\nawait r.recordOutcome({\n  provider: \"openai\",\n  model: \"gpt-4\",\n  success: false,\n  statusCode: 429,\n  errorMessage: \"You exceeded your daily limit. Try tomorrow.\",\n});\n\n// 401 / 403 → circuit open INDEFINITELY (until manual reset). Bad credentials shouldn't burn through every retry slot.\nawait r.recordOutcome({\n  provider: \"openai\",\n  model: \"gpt-4\",\n  success: false,\n  statusCode: 401,\n  errorMessage: \"Invalid API key\",\n});\n\n// 5xx / timeout / generic → circuit error_streak++ atomically; opens at threshold (default 3).\nawait r.recordOutcome({\n  provider: \"groq\",\n  model: \"llama-3.3-70b-versatile\",\n  success: false,\n  statusCode: 503,\n  errorMessage: \"Upstream timed out\",\n});\n```\n\n## Header parsing\n\n`parseHeaders()` understands four dialects:\n\n- **Groq / OpenAI**: `x-ratelimit-{limit,remaining,reset}-{requests,tokens}`\n- **Anthropic**: `anthropic-ratelimit-{requests,tokens}-{limit,remaining,reset}`\n- **IETF draft**: `RateLimit-Limit`, `RateLimit-Remaining`, `RateLimit-Reset`\n- **Standalone**: `Retry-After` (Cerebras-style)\n\n`Retry-After` accepts integer seconds (`\"60\"`), HTTP-date (`\"Wed, 01 Jan 2099 00:00:00 GMT\"`), and Groq-style relative units (`\"60s\"` / `\"5m\"` / `\"2h\"`). The `parseInt(\"5m\") === 5` trap is regression-tested.\n\nWhen response headers indicate **<5% remaining** on any dimension, the router sets a *soft block* even on a successful call — so subsequent `guard()` skips the candidate before you'd actually 429.\n\n## Public API\n\n- `router(state?, opts?)` — factory. Default: in-memory state.\n- `Router` — class. Methods: `guard()`, `recordOutcome()`, `seedCaps()`.\n- `MemoryState` — backend. SQLite + Postgres deferred to a future release.\n- `parseHeaders(headers)` → `HeaderInsight`\n- `looksLikeQuotaExhausted(errorMessage)` / `inferQuotaPeriod(errorMessage)` → standalone classifier helpers, useable without instantiating a router.\n\n## Design\n\nThree states are **orthogonal** and live under primary key `(provider, model, credentialAlias)`. A provider can be quota-exhausted on `daily/tokens` but still healthy on circuit, etc. `guard()` evaluates them in this precedence order:\n\n1. Circuit OPEN (with `retry_at > now`) → block.\n2. Active rate-limit blocks → block, return shortest TTL.\n3. Quota explicitly exhausted (period not rolled over) → block until period_end.\n4. Quota near cap (default ≥97% with the request's *estimated* tokens factored in) → block until period_end.\n5. Otherwise → allow.\n\n`recordOutcome()` handles the post-call mutation, and `_consume_headers()` extracts forward-looking signals from response headers regardless of success.\n\n## Why a separate library\n\nMost routers (LiteLLM, ClawRouter, OmniRoute pre-PR-#2116) collapse all 429s into a single uniform retry policy. That works until you hit a free-tier monthly cap and burn 1440 retries/day for the rest of the month. Separating the three states means a quota-exhausted provider gets a long cooldown (until period_end), a rate-limited one gets the short Retry-After, and a misconfigured credential opens the circuit indefinitely until you intervene.\n\n## License\n\nMIT — see [LICENSE](./LICENSE).\n\n## See also\n\n- [resilient-llm-router](https://github.com/eleata/resilient-llm-router) — the original Python library this port is based on.\n- [OmniRoute PR #2116](https://github.com/diegosouzapw/OmniRoute/pull/2116) — the same patterns landed upstream in OmniRoute's circuit breaker (issue #2100).\n","readmeFilename":"README.md"}