{"_id":"@curliness8029/model-router","_rev":"4-e7fd05c4fd809431155c2b77e402e221","name":"@curliness8029/model-router","dist-tags":{"latest":"1.3.0"},"versions":{"1.0.0":{"name":"@curliness8029/model-router","version":"1.0.0","keywords":["llm","proxy","router","cost-optimization","anthropic","openai","claude","gpt","cache","bun","agent"],"author":{"name":"Sumukh Nitundila"},"license":"MIT","_id":"@curliness8029/model-router@1.0.0","maintainers":[{"name":"curliness8029","email":"sumukh14@gmail.com"}],"homepage":"https://github.com/Boredphilosopher96/model-router#readme","bugs":{"url":"https://github.com/Boredphilosopher96/model-router/issues"},"bin":{"model-router":"src/cli.ts"},"dist":{"shasum":"663df9a6ae3957585c424d0e279d9d40727037ed","tarball":"https://registry.npmjs.org/@curliness8029/model-router/-/model-router-1.0.0.tgz","fileCount":29,"integrity":"sha512-MAvP+b/43C6QXw5GQdhxWYSd9GH1kq+POooRD9V2aWiMjW32a1/dhKdvkaUEiHjOEl3WbnE3N9sY0lNLBrtQAw==","signatures":[{"sig":"MEUCIBSSUfznx7kaPm8CGPSDzlLAz7rvlhLcNY4eBV5jBWT2AiEAqjXrgpeFBXmNGZcYVzbbLxSdJjMxW7pSn/kA+xg+3Ho=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":153891},"type":"module","module":"src/index.ts","engines":{"bun":">=1.1.0"},"exports":{".":"./src/index.ts","./plugins":"./src/plugins/index.ts"},"gitHead":"eaac8cee5b9eec4fa0b9dee48de8f6416b126f68","scripts":{"dev":"bun --watch src/cli.ts","test":"bun test","start":"bun run src/cli.ts","typecheck":"bunx tsc --noEmit"},"_npmUser":{"name":"curliness8029","email":"sumukh14@gmail.com"},"repository":{"url":"git+https://github.com/Boredphilosopher96/model-router.git","type":"git"},"_npmVersion":"11.6.1","description":"Harness-blind LLM proxy for Bun: routes every request to the cheapest capable model (Anthropic + OpenAI), caches responses, auto-updates pricing, tracks savings on a live dashboard, and runs a request/response plugin pipeline","directories":{},"_nodeVersion":"24.10.0","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"@types/bun":"latest","typescript":"^5.6.0"},"_npmOperationalInternal":{"tmp":"tmp/model-router_1.0.0_1783219565938_0.061404635765860194","host":"s3://npm-registry-packages-npm-production"}},"1.1.0":{"name":"@curliness8029/model-router","version":"1.1.0","keywords":["llm","proxy","router","cost-optimization","anthropic","openai","claude","gpt","cache","bun","agent"],"author":{"name":"Sumukh Nitundila"},"license":"MIT","_id":"@curliness8029/model-router@1.1.0","maintainers":[{"name":"curliness8029","email":"sumukh14@gmail.com"}],"homepage":"https://github.com/Boredphilosopher96/model-router#readme","bugs":{"url":"https://github.com/Boredphilosopher96/model-router/issues"},"bin":{"model-router":"src/cli.ts"},"dist":{"shasum":"6b7c77b1d275983360aa414165470ac3e3755370","tarball":"https://registry.npmjs.org/@curliness8029/model-router/-/model-router-1.1.0.tgz","fileCount":30,"integrity":"sha512-onTHxAOqT7+dJ7QKMnNd7J6YCSDDKxABX3+DjwzvehB6Ap6ovF78bgBSIQC8/b7wQKxnU1A15bAXTgBC+i3ecw==","signatures":[{"sig":"MEUCIQDphBGCoBjzIv7K7f3MzJCsFIWpyooG7QwHZyUkKLwI2wIgKFRIpNwdNqaQUWELTT3K9nD6i5W1UGamp5au1jKAt8I=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@curliness8029%2fmodel-router@1.1.0","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":201083},"type":"module","module":"src/index.ts","engines":{"bun":">=1.1.0"},"exports":{".":"./src/index.ts","./plugins":"./src/plugins/index.ts"},"gitHead":"28fb1ae26f149ae196a2e589a8f8feca9940297f","scripts":{"dev":"bun --watch src/cli.ts","test":"bun test","start":"bun run src/cli.ts","typecheck":"bunx tsc --noEmit"},"_npmUser":{"name":"GitHub Actions","email":"npm-oidc-no-reply@github.com","trustedPublisher":{"id":"github","oidcConfigId":"oidc:251d7f5a-9133-48b8-b180-7fbe9df8fa74"}},"repository":{"url":"git+https://github.com/Boredphilosopher96/model-router.git","type":"git"},"_npmVersion":"11.18.0","description":"Harness-blind LLM proxy for Bun: routes every request to the cheapest capable model (Anthropic + OpenAI), caches responses, auto-updates pricing, tracks savings on a live dashboard, and runs a request/response plugin pipeline","directories":{},"_nodeVersion":"24.18.0","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"@types/bun":"latest","typescript":"^5.6.0"},"_npmOperationalInternal":{"tmp":"tmp/model-router_1.1.0_1783265356536_0.4578870418397507","host":"s3://npm-registry-packages-npm-production"}},"1.2.0":{"name":"@curliness8029/model-router","version":"1.2.0","keywords":["llm","proxy","router","cost-optimization","anthropic","openai","claude","gpt","cache","bun","agent"],"author":{"name":"Sumukh Nitundila"},"license":"MIT","_id":"@curliness8029/model-router@1.2.0","maintainers":[{"name":"curliness8029","email":"sumukh14@gmail.com"}],"homepage":"https://github.com/Boredphilosopher96/model-router#readme","bugs":{"url":"https://github.com/Boredphilosopher96/model-router/issues"},"bin":{"model-router":"src/cli.ts"},"dist":{"shasum":"a70678accacbbf05d18e1a6e8e0d9bf5e10e479a","tarball":"https://registry.npmjs.org/@curliness8029/model-router/-/model-router-1.2.0.tgz","fileCount":34,"integrity":"sha512-8TKTmBD2RZFiYYj77Y6JuOJQAMgPqWcDQYYcUUkxClqVl+7QwESHPNnuM9jIPd5SuOss7yh2w0+loX5drngBTg==","signatures":[{"sig":"MEUCIDXURshNVtNRe3rozf83WxwaJiSUBlOmb74utKIf9iIEAiEAncjEy6z6At4fw+AguPmK3FF6pKywvvpS9TwXOZ3Xb4E=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@curliness8029%2fmodel-router@1.2.0","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":255773},"type":"module","module":"src/index.ts","engines":{"bun":">=1.1.0"},"exports":{".":"./src/index.ts","./plugins":"./src/plugins/index.ts"},"gitHead":"68563f9a23d12ed7136830cacf42b8fa7cfd756d","scripts":{"dev":"bun --watch src/cli.ts","test":"bun test","start":"bun run src/cli.ts","typecheck":"bunx tsc --noEmit"},"_npmUser":{"name":"GitHub Actions","email":"npm-oidc-no-reply@github.com","trustedPublisher":{"id":"github","oidcConfigId":"oidc:251d7f5a-9133-48b8-b180-7fbe9df8fa74"}},"repository":{"url":"git+https://github.com/Boredphilosopher96/model-router.git","type":"git"},"_npmVersion":"11.18.0","description":"Harness-blind LLM proxy for Bun: routes every request to the cheapest capable model (Anthropic + OpenAI), caches responses, auto-updates pricing, tracks savings on a live dashboard, and runs a request/response plugin pipeline","directories":{},"_nodeVersion":"24.18.0","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"@types/bun":"latest","typescript":"^5.6.0"},"_npmOperationalInternal":{"tmp":"tmp/model-router_1.2.0_1783266546709_0.6130869416454214","host":"s3://npm-registry-packages-npm-production"}},"1.3.0":{"name":"@curliness8029/model-router","version":"1.3.0","description":"Harness-blind LLM proxy for Bun: routes every request to the cheapest capable model (Anthropic + OpenAI), caches responses, auto-updates pricing, tracks savings on a live dashboard, and runs a request/response plugin pipeline","license":"MIT","type":"module","module":"src/index.ts","exports":{".":"./src/index.ts","./plugins":"./src/plugins/index.ts"},"bin":{"model-router":"src/cli.ts"},"engines":{"bun":">=1.1.0"},"keywords":["llm","proxy","router","cost-optimization","anthropic","openai","claude","gpt","cache","bun","agent"],"repository":{"type":"git","url":"git+https://github.com/Boredphilosopher96/model-router.git"},"scripts":{"start":"bun run src/cli.ts","dev":"bun --watch src/cli.ts","test":"bun test","typecheck":"bunx tsc --noEmit"},"devDependencies":{"@types/bun":"latest","typescript":"^5.6.0"},"homepage":"https://github.com/Boredphilosopher96/model-router#readme","bugs":{"url":"https://github.com/Boredphilosopher96/model-router/issues"},"author":{"name":"Sumukh Nitundila"},"publishConfig":{"access":"public"},"gitHead":"9a62e15e1c5a0428207a435ac2b28d8a5d071355","_id":"@curliness8029/model-router@1.3.0","_nodeVersion":"24.18.0","_npmVersion":"11.18.0","dist":{"integrity":"sha512-gll9h78GwgoZGUY7LSDlqo2hFkeR955Z8TdWoXytJzPKzwRUOQaOeM2utM15clEMAdh9GYoYt/djdY2qP/yKaQ==","shasum":"7e723c7669c8f3c4fed9c83a8a1c3e26d535696c","tarball":"https://registry.npmjs.org/@curliness8029/model-router/-/model-router-1.3.0.tgz","fileCount":36,"unpackedSize":283128,"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@curliness8029%2fmodel-router@1.3.0","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQCGizDHZvSDax3x2vbwisYcZRQ0YxGVgPWdTGBIFcjABAIhAKxTazlNZ1DjrULIYJRxpkw7szIVhMxwnikxsmqLJqol"}]},"_npmUser":{"name":"GitHub Actions","email":"npm-oidc-no-reply@github.com","trustedPublisher":{"id":"github","oidcConfigId":"oidc:251d7f5a-9133-48b8-b180-7fbe9df8fa74"}},"directories":{},"maintainers":[{"name":"curliness8029","email":"sumukh14@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/model-router_1.3.0_1783290749534_0.2471496051461246"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-05T02:46:05.814Z","modified":"2026-07-05T22:32:29.985Z","1.0.0":"2026-07-05T02:46:06.083Z","1.1.0":"2026-07-05T15:29:16.681Z","1.2.0":"2026-07-05T15:49:06.850Z","1.3.0":"2026-07-05T22:32:29.711Z"},"bugs":{"url":"https://github.com/Boredphilosopher96/model-router/issues"},"author":{"name":"Sumukh Nitundila"},"license":"MIT","homepage":"https://github.com/Boredphilosopher96/model-router#readme","keywords":["llm","proxy","router","cost-optimization","anthropic","openai","claude","gpt","cache","bun","agent"],"repository":{"type":"git","url":"git+https://github.com/Boredphilosopher96/model-router.git"},"description":"Harness-blind LLM proxy for Bun: routes every request to the cheapest capable model (Anthropic + OpenAI), caches responses, auto-updates pricing, tracks savings on a live dashboard, and runs a request/response plugin pipeline","maintainers":[{"name":"curliness8029","email":"sumukh14@gmail.com"}],"readme":"<div align=\"center\">\n\n# model-router\n\n**A harness-blind man-in-the-middle that cuts your LLM spend.**\n\nIt sits between any coding agent and any number of model endpoints, looks at each request, and redirects it to the cheapest model and endpoint that can handle it — swapping only the model string, never the request format.\n\n[![CI](https://github.com/Boredphilosopher96/model-router/actions/workflows/ci.yml/badge.svg)](https://github.com/Boredphilosopher96/model-router/actions/workflows/ci.yml) [![runtime](https://img.shields.io/badge/runtime-bun%20%E2%89%A5%201.1-black)]() [![license](https://img.shields.io/badge/license-MIT-blue)](LICENSE) [![types](https://img.shields.io/badge/types-strict-3178c6)]()\n\n</div>\n\n```\n                        ┌──────────────────────────┐      ┌─ anthropic api (claude-*)\n Claude Code ──┐        │       model-router       │      ├─ github copilot\n opencode ─────┼──────► │  dialect-aware routing   │ ───► ├─ your gateway (both dialects)\n Codex CLI ────┤        │  cache · escalation ·    │      ├─ openrouter\n Copilot BYOK ─┘        │  savings · plugins       │      └─ openai api (gpt-*)\n                        └──────────────────────────┘\n```\n\n## Why\n\nCoding agents default to expensive frontier models for every request — including the trivial ones. And if you run several providers (a Copilot subscription, direct APIs, an internal gateway), the cheapest way to serve any given request keeps shifting. model-router makes that decision per request, invisibly:\n\n- **Content-aware task classification.** The router extracts intent from each request: lookup/summarize → tier 1, codegen/debug → tier 2, architecture/reasoning → tier 3. User-defined regex rules override. Optional LLM-based classification for complex patterns. Confidence gates prevent aggressive downgrade when unsure.\n- **Cache-aware stickiness.** When a conversation has warm cache on an expensive model, staying put costs less than switching to a cheaper model cold. The router knows the difference and sticks only when it saves money — switching frequently for short tasks, staying put for long ones. Sticky decisions are labeled in headers.\n- **Quality mode** — refuses downgrade when classifier confidence < 0.65, for workflows where \"wrong answer faster\" loses money.\n- **Escalates when stuck.** Repeated failures, erroring tool calls, or retry loops bump that conversation up a tier — even above the model it asked for — then decay back after sustained success.\n- **Shadow mode.** Run an alternative routing strategy on real traffic without applying it; compare agreement rate and estimated cost delta before switching live.\n- **Budgets.** Daily, monthly, and per-upstream spend limits that tighten routing mode (aggressive → balanced → quality) as limits fill, never blocking traffic.\n- **Upstream health and automatic failover.** Circuit breaker (opens after 5 failures in 60s) and latency-based tie-breaking; open-circuit upstreams skipped unless all candidates exhausted. Automatic failover retries transient errors on alternative (model, upstream) pairs, making provider outages and rate limits invisible to the harness.\n- **Quality calibration.** Continuous measurement of downgraded request adequacy; per-task type +1 tier recommendations when adequacy drops below 0.8.\n- **Presets for fast setup.** Declare providers by name (`\"providers\": [\"anthropic\", {\"name\":\"copilot\",\"preset\":\"github-copilot\"}]`), inheriting defaults from built-in presets (Anthropic, OpenAI, GitHub Copilot, GitHub Models, OpenRouter).\n- **One-command harness setup.** Run `model-router setup <harness> [--write]` to print or apply router config for Claude Code, Codex CLI, opencode, GitHub Copilot, or other tools.\n- **Router performance dashboard.** Live metrics on downgrade rate, sticky rate, escalation rate, regret rate (downgraded conversations that later escalated — the router's misjudgment signal), breakdowns by task type, and tier distribution for tuning.\n- **Never gets stale.** The model catalog, prices, capability flags, and gateway availability refresh daily from a maintained feed. GitHub Copilot prices via AI credits at per-token rates. New model generations automatically supersede old ones. Zero manual updates.\n- **Never gets in the way.** Unknown models, unparseable requests, feed outages, broken plugins — everything fails open and passes through. Provider errors reach your harness untouched.\n- **Proves the savings.** Every request is logged with actual cost vs. what the requested model would have cost, on a live dashboard.\n- **Self-tuning and spend-safe.** Budgets + shadow mode + quality calibration make the router self-optimizing: measure continuously, validate changes safely, control costs automatically.\n\nHarness-blind means your agent never knows: responses report the model it asked for; the truth lives in `x-router-*` response headers and the dashboard.\n\n## Quick start\n\n```sh\ngit clone https://github.com/Boredphilosopher96/model-router && cd model-router\nbun install\nANTHROPIC_API_KEY=sk-ant-... OPENAI_API_KEY=sk-... bun start\n```\n\nPoint a harness at it and watch the dashboard:\n\n```sh\nANTHROPIC_BASE_URL=http://localhost:4141 claude     # Claude Code\nopen http://localhost:4141/dashboard\n```\n\nFor harnesses that support inline config, use the setup command:\n\n```sh\nbunx @curliness8029/model-router setup claude-code\nbunx @curliness8029/model-router setup codex --write\n```\n\nThat's the whole minimal setup — with no config file, the proxy fronts the direct Anthropic and OpenAI APIs. To route across multiple providers (Copilot, gateways, internal backends), declare them:\n\n```sh\ncp router.config.example.json router.config.json    # then edit\n```\n\n```jsonc\n{\n  \"allowedModels\": [\"claude-haiku-*\", \"claude-sonnet-*\", \"claude-opus-*\", \"gpt-5.4-*\", \"gpt-5.5\"],\n  \"taskRules\": [\n    { \"pattern\": \"\\\\b(deploy|release|promote)\\\\b\", \"tier\": 3, \"taskType\": \"release\" }\n  ],\n  \"providers\": [\n    \"anthropic\",\n    { \"name\": \"copilot\", \"preset\": \"github-copilot\" },\n    {\n      \"name\": \"mygateway\", \"baseUrl\": \"https://llm.internal.example.com\", \"dialect\": \"both\",\n      \"models\": [\"claude-*\", \"gpt-*\", \"my-private-model\"],\n      \"apiKeyEnv\": \"MYGATEWAY_API_KEY\", \"authStyle\": \"bearer\",\n      \"pricing\": { \"my-private-model\": { \"inputPer1M\": 0.5, \"outputPer1M\": 2.0, \"tier\": 2 } }\n    }\n  ]\n}\n```\n\nEach provider becomes a mount — point each harness provider at `http://localhost:4141/p/<name>`, and the router redirects between all of them.\n\n## Features\n\n| | |\n|---|---|\n| **Cost routing** | Cheapest capable (model, endpoint) pair per request; requested model is the spend ceiling; four modes (`aggressive` / `balanced` / `quality` / `off`) |\n| **Task classification** | Content-aware heuristic (regex rules → task taxonomy → structural signals) or optional LLM-based classifier; per-request taskType and confidence |\n| **Cache-aware stickiness** | Long conversations stay on warm models only when it saves money; short tasks can switch freely. Sticky decisions labeled in response headers |\n| **Quality mode** | Like `balanced` but refuses downgrade when classifier confidence < 0.65 |\n| **Automatic pricing** | Live feed for known gateways; GitHub Copilot priced per token via AI credits; catalog API pricing assumed for custom gateways; `pricing` entries for private models |\n| **Escalation** | Stuck conversations bump up a tier and settle back; observable at `/api/escalations` |\n| **`auto` model** | Advertised via `GET /v1/models`; selecting it delegates the whole choice to the router |\n| **Model allowlist** | `allowedModels` globs restrict routing targets; everything else still passes through |\n| **Multimodal-safe** | Image/document/audio requests only route to vision-capable models; tool-calling requests only to tool-capable models |\n| **Shadow mode** | Run alternative strategy on real traffic without applying it; validate before switching via `GET /api/router-eval` |\n| **Budgets** | Daily, monthly, per-upstream limits; routing tightens as limits fill; never blocks traffic |\n| **Upstream health** | Circuit breaker per upstream; latency-aware tie-breaking; open circuits skipped (fail-open) |\n| **Rate-limit awareness** | Provider rate-limit headers parsed; under-budget upstreams soft-throttled; visible in `/api/upstream-health` |\n| **Automatic failover** | Retryable errors (429, 5xx, unreachable) retried on up to two next-best (model, upstream) pairs before surfacing; safe for streaming |\n| **Quality calibration** | Measure downgrade adequacy, grade via frontier model, recommend tier adjustments; apply automatically or manually |\n| **Response cache** | SQLite, TTL-based; identical requests served for free; streaming responses cached as raw SSE and replayed byte-for-byte; streams over 2 MB not cached |\n| **Streaming cache replay** | Streaming responses cached and replayed byte-for-byte on identical requests (x-router-cache: hit); stream and non-stream variants cached separately; calibration samples streamed responses |\n| **Plugins** | `onRequest` / `onRouteDecision` / `onResponse` / `onRecord` hooks; match scoping; priority ordering; loadable from config without forking |\n| **Adapters** | Per-upstream request/response reshaping for gateways with nonstandard JSON; `composePlugins()` for merging related plugins |\n| **Presets** | `\"providers\": [\"anthropic\", {\"name\":\"copilot\",\"preset\":\"github-copilot\"}]` — endpoint/auth/path defaults for known gateways |\n| **Setup command** | `model-router setup <harness> [--write]` prints or applies config for Claude Code, Codex, opencode, Copilot, etc. |\n| **Dashboard** | Money saved, downgrade rate, cache hit rate, router performance (regret rate, sticky rate, escalation rate, task type breakdown), per-model and per-route tables |\n\n## Documentation\n\n| Guide | Contents |\n|---|---|\n| [Connecting harnesses](docs/harnesses.md) | Claude Code, opencode, Codex CLI, Copilot BYOK, generic setup; verifying with headers |\n| [Configuration reference](docs/configuration.md) | Every env var and `router.config.json` field; pricing resolution; allowlist semantics |\n| [How routing works](docs/routing.md) | Complexity scoring, tiers, escalation mechanics, fail-open guarantees |\n| [Extending](docs/extending.md) | Writing plugins, upstream adapters, custom models/pricing, library API |\n| [HTTP API reference](docs/http-api.md) | Endpoints, response headers, observability APIs, error semantics |\n\nWorking templates ship in [`examples/`](examples): a telemetry plugin and a gateway adapter.\n\n## Requirements & operations\n\n- **Runtime:** [Bun](https://bun.sh) ≥ 1.1 (uses `bun:sqlite`, `Bun.serve`).\n- **State:** one SQLite file (`DB_PATH`) for cache + stats, one JSON file for the price-feed cache. Delete either at any time; the proxy rebuilds them.\n- **Shutdown:** SIGINT/SIGTERM close the server gracefully.\n- **Health:** `GET /health` for liveness; `GET /api/stats` for metrics scraping.\n- **Security notes:** the proxy forwards a harness's own credentials only to the upstream they were meant for; cross-upstream redirects require that upstream's own configured key. Run it on localhost or inside your network perimeter — it is a credential-bearing proxy and ships without inbound auth.\n\n## Development\n\n```sh\nbun test           # 80 unit + integration tests (mock upstreams, no network)\nbun run typecheck  # strict tsc\nbun run dev        # watch mode\n```\n\nSee [CONTRIBUTING.md](CONTRIBUTING.md) for project layout and conventions, and [CHANGELOG.md](CHANGELOG.md) for release history.\n\n## License\n\n[MIT](LICENSE) © 2026 Sumukh Nitundila\n","readmeFilename":"README.md"}