{"_id":"@bytesbrains/llm-cost-control","name":"@bytesbrains/llm-cost-control","dist-tags":{"latest":"0.1.0"},"versions":{"0.1.0":{"name":"@bytesbrains/llm-cost-control","version":"0.1.0","description":"LLM cost-control layer — response caching, budget gates, model routing, per-tenant metering. The margin-protector for managed AI runs.","keywords":["llm","ai","cost-control","caching","budget","model-routing","metering","bytesbrains"],"author":{"name":"BytesBrains Pte Ltd","url":"https://bytesbrains.com"},"repository":{"type":"git","url":"git+https://github.com/bytesbrains/llm-cost-control.git"},"homepage":"https://github.com/bytesbrains/llm-cost-control#readme","bugs":{"url":"https://github.com/bytesbrains/llm-cost-control/issues"},"license":"MIT","type":"module","engines":{"node":">=18"},"scripts":{"build":"esbuild src/index.ts --bundle --platform=node --target=node18 --format=esm --outfile=dist/index.js && tsc --emitDeclarationOnly","test":"vitest run","dev":"tsx --watch src/index.ts","prepublishOnly":"npm run build && npm test"},"devDependencies":{"@types/node":"^26.1.0","esbuild":"^0.28.1","tsx":"^4.19.0","typescript":"^6.0.3","vitest":"^4.1.10"},"main":"./dist/index.js","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"publishConfig":{"access":"public"},"_id":"@bytesbrains/llm-cost-control@0.1.0","gitHead":"c82333a8123bae0de5f2d9dcbcf83f9af7598167","_nodeVersion":"24.2.0","_npmVersion":"11.3.0","dist":{"integrity":"sha512-N6YFRIkF/0lobbIICuNf8vk8VnquizozVpxXUIrF/6fIf3+NN82xRXL/LlMmQDAGgj5QKxaUsmqv2dVdrD860g==","shasum":"cb78b11db6966a2598cd3ce6d603188a5e350d33","tarball":"https://registry.npmjs.org/@bytesbrains/llm-cost-control/-/llm-cost-control-0.1.0.tgz","fileCount":7,"unpackedSize":20893,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQCUAfj7c00TdrmPmcubl6qXeFQscKfMva86Vym2t3snHQIhALfgL/iu5d38EYppWqKT4sr/3unWoNguyIBBpLO8xllw"}]},"_npmUser":{"name":"nandal","email":"sandeep@nandal.in"},"directories":{},"maintainers":[{"name":"nandal","email":"sandeep@nandal.in"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/llm-cost-control_0.1.0_1783475735785_0.33722751899144887"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-08T01:55:35.630Z","0.1.0":"2026-07-08T01:55:35.942Z","modified":"2026-07-08T01:55:36.163Z"},"maintainers":[{"name":"nandal","email":"sandeep@nandal.in"}],"description":"LLM cost-control layer — response caching, budget gates, model routing, per-tenant metering. The margin-protector for managed AI runs.","homepage":"https://github.com/bytesbrains/llm-cost-control#readme","keywords":["llm","ai","cost-control","caching","budget","model-routing","metering","bytesbrains"],"repository":{"type":"git","url":"git+https://github.com/bytesbrains/llm-cost-control.git"},"author":{"name":"BytesBrains Pte Ltd","url":"https://bytesbrains.com"},"bugs":{"url":"https://github.com/bytesbrains/llm-cost-control/issues"},"license":"MIT","readme":"# @bytesbrains/llm-cost-control\n\n[![npm](https://img.shields.io/npm/v/@bytesbrains/llm-cost-control)](https://www.npmjs.com/package/@bytesbrains/llm-cost-control)\n[![CI](https://github.com/bytesbrains/llm-cost-control/actions/workflows/ci.yml/badge.svg)](https://github.com/bytesbrains/llm-cost-control/actions/workflows/ci.yml)\n[![license: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)\n\nThe code-guarantees layer for LLM spend — **caching, budget gates, model routing, and per-tenant metering** wrapped around every model call. For anyone operating AI systems on a flat fee, this is the difference between a profitable run and a leaky one.\n\n## Why\n\nDaily AI workloads leak money four ways: identical calls re-executed (no caching), expensive models doing cheap work (no routing), runaway loops burning budget silently (no gates), and spend nobody can attribute per client (no metering). This package makes all four impossible by construction: **the model proposes, code guarantees** — your pipeline decides *what* to ask; this layer guarantees *what it may cost*.\n\n- A call cannot exceed its budget — ceilings are enforced, not suggested\n- A call cannot skip the meter — a failed meter write fails the call\n- Every skip/degrade/queue is a **returned value and a metered event**, never a silent no-op\n\n## Install\n\n```bash\nnpm install @bytesbrains/llm-cost-control\n```\n\n## 60-second usage\n\n```ts\nimport { CostControl, fileStore } from \"@bytesbrains/llm-cost-control\";\n\nconst cc = new CostControl({\n  prices: { \"claude-sonnet-5\": { in: 3, out: 15 }, \"claude-haiku-4-5\": { in: 0.8, out: 4 } }, // $/MTok\n  cache: fileStore(\"./.ccache\"),\n  meter: fileStore(\"./usage\"),\n  budgets: {\n    perCall:      { usd: 0.5, onExceed: \"throw\" },   // no single call may cost more\n    perDay:       { usd: 25,  onExceed: \"degrade\" }, // over budget -> cheapest model only\n    perTenantDay: { usd: 5,   onExceed: \"skip\" },    // per-client fairness\n  },\n  routes: {\n    classify: [\"claude-haiku-4-5\"],                    // cheap only\n    reason:   [\"claude-haiku-4-5\", \"claude-sonnet-5\"], // escalation ladder\n  },\n});\n\nconst res = await cc.call({\n  tenant: \"acme\", job: \"reconcile-daily\", task: \"classify\",\n  cacheKey: { docId, promptVersion }, ttl: \"24h\",\n  estimate: { inTok: 1200, outTok: 300 },\n  exec: async (model) => {                    // you own the SDK call\n    const r = await anthropic.messages.create({ model, /* ... */ });\n    return { value: r, usage: { inTok: r.usage.input_tokens, outTok: r.usage.output_tokens } };\n  },\n  escalate: (r) => needsBetterModel(r.value), // optional: walk one rung up the ladder\n});\n\nif (res.ok) console.log(res.value, `$${res.usd}`, res.cached ? \"(cache — free)\" : res.model);\n\nconst today = await cc.summarize();           // { calls, usd, cacheHits, byTenant, byModel, byTask, ... }\n```\n\n## Guarantees\n\n| Component | Guarantee |\n|---|---|\n| **Cache** | identical work is never paid for twice within TTL (stable content-hash; key order irrelevant) |\n| **Budget gates** | spend cannot exceed ceilings — `throw` / `skip` / `degrade` / `queue` per scope; actuals over `perCall` always throw |\n| **Router** | every task starts on the cheapest capable model; escalation is explicit and single-rung |\n| **Meter** | every call is attributed (tenant/job/task/model/$) to an append-only log — or it doesn't happen |\n\n## API\n\n- `new CostControl({ prices, routes, budgets?, cache?, meter, now? })`\n- `cc.call({ tenant, job, task, cacheKey?, ttl?, estimate?, exec, escalate? })` → `{ ok: true, value, model, usd, cached, escalated, degraded }` | `{ ok: false, skipped }` | `{ ok: false, queued, token }`\n- `cc.summarize(day?, tenant?)` → aggregated `Summary`\n- `costOf(usage, model, prices)` → `$` (throws on unknown model — cost math is never guessed)\n- `memoryStore()` / `fileStore(dir)` — or implement the 4-method `Store` interface (Redis etc.)\n- `BudgetExceededError` — carries `scope`, `limitUsd`, `attemptedUsd`, `spentUsd`\n\n## Notes\n\n- Bring your own SDK — this wraps any provider (Anthropic, OpenAI, local).\n- `estimate` powers the *pre-flight* check; actual usage is enforced and metered regardless.\n- v0.1 ships memory + file stores; the `Store` interface is deliberately tiny (`get/set/append/readDay`).\n\n---\n\nBuilt and maintained by [BytesBrains](https://bytesbrains.com) — AI automation & agents, engineered to production standards.\n*The model proposes, code guarantees.*\n","readmeFilename":"README.md","_rev":"1-836318a1ecdadbe374b9ec8ab9244c99"}