{"_id":"@andy-toolforge/llm-gateway","_rev":"2-dc6793e25668c65482c934412c5419f6","name":"@andy-toolforge/llm-gateway","dist-tags":{"latest":"0.1.1"},"versions":{"0.1.0":{"name":"@andy-toolforge/llm-gateway","version":"0.1.0","_id":"@andy-toolforge/llm-gateway@0.1.0","maintainers":[{"name":"andy_pham","email":"phamlehoaian@gmail.com"}],"homepage":"https://github.com/andy-pham-it/toolforge#readme","bugs":{"url":"https://github.com/andy-pham-it/toolforge/issues"},"bin":{"llm-gateway":"bin/cli.js"},"dist":{"shasum":"fd3c14db52bed324e20a5012577f19610fd357dc","tarball":"https://registry.npmjs.org/@andy-toolforge/llm-gateway/-/llm-gateway-0.1.0.tgz","fileCount":39,"integrity":"sha512-WDb9hptPk5WFGDNyuERY+pA3rM5D2rTvwbCiKjuFib4nSeKBX9fHkUkHve+lgNvUvr+Ctkn6V5d6fStI4J6EAQ==","signatures":[{"sig":"MEUCIFoK3TBlucw51+wo8F62E3itTmnrJ/4vGUxVuxe/KImcAiEAyL3cWb96hLH/+IFyaqnQ3Kh7zeCTIQBTYe5j/4SJLK8=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":82961},"main":"lib/index.js","gitHead":"edda786023c8cd1134f7df08553947d074e0def7","scripts":{"test":"node --test lib/**/*.test.js","start":"node bin/cli.js","postinstall":"node skills/postinstall.js"},"_npmUser":{"name":"andy_pham","email":"phamlehoaian@gmail.com"},"repository":{"url":"git+https://github.com/andy-pham-it/toolforge.git","type":"git"},"_npmVersion":"10.8.2","description":"LLM API gateway — multi-provider routing, failover, rate limiting, key rotation, cost tracking, caching","directories":{},"_nodeVersion":"20.20.2","dependencies":{"express":"^4.21.0","@andy-toolforge/core":"^1.2.0"},"_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/llm-gateway_0.1.0_1784278471740_0.8900587867478147","host":"s3://npm-registry-packages-npm-production"}},"0.1.1":{"name":"@andy-toolforge/llm-gateway","version":"0.1.1","description":"LLM API gateway — multi-provider routing, failover, rate limiting, key rotation, cost tracking, caching","main":"lib/index.js","repository":{"type":"git","url":"git+https://github.com/andy-pham-it/toolforge.git"},"scripts":{"test":"node --test lib/**/*.test.js","start":"node bin/cli.js","postinstall":"node skills/postinstall.js"},"dependencies":{"@andy-toolforge/core":"^1.2.0","express":"^4.21.0"},"bin":{"llm-gateway":"bin/cli.js"},"_id":"@andy-toolforge/llm-gateway@0.1.1","gitHead":"bd31d2226a296d9c22cd46293b9072326f959b7c","bugs":{"url":"https://github.com/andy-pham-it/toolforge/issues"},"homepage":"https://github.com/andy-pham-it/toolforge#readme","_nodeVersion":"20.20.2","_npmVersion":"10.8.2","dist":{"integrity":"sha512-zQXwBQcFIak+AN0Y2e4CPUVrRd3gPqzoA16keHqSo34V8K9f/8/8ZS3CpJnuo9kxdLOG/gDqBUKfuazviFhW9Q==","shasum":"363a5c9ab8480cbb887e8f52cc24f89ea38f7663","tarball":"https://registry.npmjs.org/@andy-toolforge/llm-gateway/-/llm-gateway-0.1.1.tgz","fileCount":41,"unpackedSize":101079,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIHsYV8n29NVFr0sOzWA0aPgKQ2N+T9mMcyFSEWVtP0dsAiEA4ZIYOcRJbqHnv4I1TKlIB7gz2ab/7QBRXarOFzR7qK8="}]},"_npmUser":{"name":"andy_pham","email":"phamlehoaian@gmail.com"},"directories":{},"maintainers":[{"name":"andy_pham","email":"phamlehoaian@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/llm-gateway_0.1.1_1784443312555_0.6212043220956422"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-17T08:54:31.573Z","modified":"2026-07-19T06:41:52.838Z","0.1.0":"2026-07-17T08:54:31.895Z","0.1.1":"2026-07-19T06:41:52.689Z"},"bugs":{"url":"https://github.com/andy-pham-it/toolforge/issues"},"homepage":"https://github.com/andy-pham-it/toolforge#readme","repository":{"type":"git","url":"git+https://github.com/andy-pham-it/toolforge.git"},"description":"LLM API gateway — multi-provider routing, failover, rate limiting, key rotation, cost tracking, caching","maintainers":[{"name":"andy_pham","email":"phamlehoaian@gmail.com"}],"readme":"# @andy-toolforge/llm-gateway\n\n[![npm](https://img.shields.io/npm/v/@andy-toolforge/llm-gateway)](https://npmjs.com/package/@andy-toolforge/llm-gateway)\n[![License](https://img.shields.io/npm/l/@andy-toolforge/llm-gateway)](https://github.com/andy-pham-it/toolforge)\n\n**LLM API Gateway** — multi-provider routing, failover, rate limiting, auth, caching, circuit breaker, key rotation, cost tracking. Pipeline-of-Stages architecture.\n\n```\nRequest → [Auth → RateLimit → Cache → Router → Fallback → CircuitBreaker → Provider → CostLogger] → Response\n```\n\n## Installation\n\n```bash\nnpm install @andy-toolforge/llm-gateway\n```\n\n## Quick Start\n\n```javascript\nconst { createGateway } = require('@andy-toolforge/llm-gateway');\n\nconst gw = createGateway({\n  apiKey: 'sk-your-key',\n  keys: { 'sk-your-key': { tenant: 'my-app' } },\n  models: {\n    'gemini-pro': { provider: 'gemini', adapter: 'GenAIAdapter' },\n    'gpt-4': { provider: 'openai', adapter: 'OpenAIAdapter' },\n  },\n});\n\n// Sync\nconst result = await gw.chat({\n  model: 'gemini-pro',\n  messages: [{ role: 'user', content: 'Hello!' }],\n});\nconsole.log(result.content);\n\n// Streaming\nfor await (const chunk of gw.chatStream({\n  model: 'gemini-pro',\n  messages: [{ role: 'user', content: 'Tell me a story' }],\n})) {\n  process.stdout.write(chunk.content);\n}\n```\n\n## Features\n\n| Feature | Stage | Config |\n|---------|-------|--------|\n| API Key Auth | AuthStage | `keys: { 'sk-xxx': { tenant } }` |\n| Rate Limiting | RateLimitStage | `rateLimits: { tenant: { capacity, refillRate } }` |\n| Response Caching | CacheStage | `cache: { ttlMs }` |\n| Model Routing | RouterStage | `models: { name: { provider, adapter } }` |\n| Failover | FallbackChain | Defined in model config |\n| Circuit Breaker | CircuitBreakerStage | `circuitBreaker: { threshold, cooldownMs }` |\n| Key Rotation | KeyRotatorStage | `keyRotator: { keys: [...], rotationIntervalMs }` |\n| Cost Logging | CostLoggerStage | Automatic per-request cost |\n| Provider Adapter | ProviderStage | `createAdapter(provider, model)` factory |\n\n## Architecture\n\n### Pipeline-of-Stages\n\nEvery request flows through a chain of stages, each with a single responsibility. Stage order is fixed:\n\n```\nRequest\n  │\n  ├─ ① AuthStage        — validate API key, extract tenant\n  ├─ ② RateLimitStage   — enforce per-tenant token bucket\n  ├─ ③ CacheStage       — return cached response if available\n  ├─ ④ RouterStage      — resolve model → provider + adapter; activate fallback chain\n  ├─ ⑤ FallbackChain    — try primary provider, fall through to alternatives on failure\n  ├─ ⑥ CircuitBreakerStage — track failures; open circuit if threshold exceeded\n  ├─ ⑦ ProviderStage    — call the LLM provider (sync or streaming)\n  └─ ⑧ CostLoggerStage  — log usage + compute cost\n  │\nResponse\n```\n\n### Request / Response Object\n\nEach stage transforms a shared `request` object and/or the `response`:\n\n```\nRequest: { model, messages, stream?, requestId, tenant?, signal? }\nResponse: { content, usage: { promptTokens, completionTokens, costUsd }, cached? }\n```\n\nStages before ProviderStage (Auth → Cache) operate on the request object. ProviderStage calls the actual LLM and produces the response. Stages after (CostLogger) decorate the response. The `next()` call passes control to the next stage; a stage can short-circuit (Cache hit, Circuit open, Rate limited) and return early.\n\n### Streaming Path\n\nWhen `request.stream = true`, ProviderStage returns an `AsyncIterable<{ content, done, usage? }>`. CacheStage bypasses caching for streaming requests (unless the entire response is buffered). The HTTP server delivers streaming via SSE (Server-Sent Events).\n\n### Gateway Class\n\nThe `Gateway` class assembles the pipeline from config and exposes three methods:\n\n| Method | Description |\n|--------|-------------|\n| `gateway.chat(req)` | Sync completion — returns `Response` |\n| `gateway.chatStream(req)` | Streaming — returns `AsyncIterable` |\n| `gateway.health()` | Pipeline status — returns `{ status, models, stages }` |\n\n## HTTP Server\n\n```bash\n# CLI\nnpx llm-gateway --port 4000 --api-key sk-admin\n\n# Or programmatically\nconst { HTTPServer, createGateway } = require('@andy-toolforge/llm-gateway');\nconst server = new HTTPServer(gateway, { port: 4000 });\nawait server.start();\n```\n\n### Routes\n\n| Method | Path | Description |\n|--------|------|-------------|\n| `GET` | `/health` | Pipeline status + model list |\n| `GET` | `/v1/models` | Available models |\n| `POST` | `/v1/chat/completions` | Chat completion (OpenAI-compatible) |\n| `POST` | `/v1/chat/completions` | With `stream: true` → SSE streaming |\n\n## Adapters\n\nCreate custom adapters for any provider:\n\n```javascript\nconst gw = createGateway({\n  models: { 'custom-model': { provider: 'custom', adapter: 'MyAdapter' } },\n  createAdapter: (provider, model) => {\n    if (provider === 'custom') return {\n      chat: async (messages, opts) => ({\n        content: 'response text',\n        usage: { promptTokens: 5, completionTokens: 10, costUsd: 0.0001 },\n      }),\n    };\n  },\n});\n```\n\n## Configuration Reference\n\n### Gateway Options\n\n| Option | Type | Required | Description |\n|--------|------|----------|-------------|\n| `apiKey` | string | yes | Admin API key |\n| `keys` | object | yes | `{ key: { tenant } }` — API key → tenant mapping |\n| `models` | object | yes | `{ name: { provider, adapter } }` — model definitions |\n| `rateLimits` | object | no | `{ tenant: { capacity, refillRate } }` |\n| `cache` | object | no | `{ ttlMs }` — cache TTL |\n| `circuitBreaker` | object | no | `{ threshold, cooldownMs, halfOpenMaxRequests }` |\n| `keyRotator` | object | no | `{ keys: [...], rotationIntervalMs }` |\n| `createAdapter` | function | no | Factory `(provider, model) => adapter` |\n| `costConfig` | object | no | `{ perMillionTokens: { input, output } }` |\n\n### HTTP Server Options\n\n| Option | Type | Default | Description |\n|--------|------|---------|-------------|\n| `port` | number | `3000` | Listen port (0 = random) |\n| `host` | string | `'localhost'` | Bind address |\n\n## Related\n\n- [@andy-toolforge/core](https://npmjs.com/package/@andy-toolforge/core) — Core LLM client and shared utilities\n- [@andy-toolforge/mcp](https://npmjs.com/package/@andy-toolforge/mcp) — MCP server toolkit\n","readmeFilename":"README.md"}