{"_id":"@boole/tokens-per-watt","name":"@boole/tokens-per-watt","dist-tags":{"latest":"0.1.0"},"versions":{"0.1.0":{"name":"@boole/tokens-per-watt","version":"0.1.0","description":"Rough research estimates of energy (Wh) used by LLM API calls — not a measurement, an approximation based on public research.","main":"dist/index.js","types":"dist/index.d.ts","bin":{"tpw":"dist/cli.js"},"scripts":{"build":"tsc","test":"node --test dist/*.test.js","prepublishOnly":"npm run build && npm test"},"keywords":["llm","energy","estimate","tokens","watt","carbon","ai","sustainability"],"license":"MIT","devDependencies":{"typescript":"^5.5.0","@types/node":"^22.0.0"},"engines":{"node":">=18"},"_id":"@boole/tokens-per-watt@0.1.0","gitHead":"a73b8c7e075ef315c0181cecf461f0cf448babb3","_nodeVersion":"22.23.2","_npmVersion":"10.9.8","dist":{"integrity":"sha512-IWCYJMXKja9Xad3ILklc3w9P/az2nBai4ux3yZ78mAfVaxh/Jl2ZJxJg7UDAANo8PTVI9+4HY4p4JSWeGcG6GQ==","shasum":"b5784fb57dc89d02503546eadcadce96c6632bf7","tarball":"https://registry.npmjs.org/@boole/tokens-per-watt/-/tokens-per-watt-0.1.0.tgz","fileCount":19,"unpackedSize":50053,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQDd48UN/mDZQShE7Ji6A0e2UQNKPmM+7qQRZx3tpK5SpAIgG6AKq/2nBt9orNrf4Y4tTCMbLlyWpuDBlhWLUKlMU80="}]},"_npmUser":{"name":"jordan.plows","email":"plowstjordan@gmail.com"},"directories":{},"maintainers":[{"name":"jordan.plows","email":"plowstjordan@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/tokens-per-watt_0.1.0_1787683229169_0.35734128309451463"},"_hasShrinkwrap":false}},"time":{"created":"2026-08-25T18:40:29.033Z","0.1.0":"2026-08-25T18:40:29.310Z","modified":"2026-08-25T18:40:29.583Z"},"maintainers":[{"name":"jordan.plows","email":"plowstjordan@gmail.com"}],"description":"Rough research estimates of energy (Wh) used by LLM API calls — not a measurement, an approximation based on public research.","keywords":["llm","energy","estimate","tokens","watt","carbon","ai","sustainability"],"license":"MIT","readme":"# Tokens-per-watt\n\nEstimate the energy (Wh / kWh) behind LLM API usage — per conversation, per\nbuild, per CI run — and compare prompting strategies to see which one would\nhave used less power.\n\n> **This is a research estimate, not a hardware measurement.** Cloud\n> providers don't expose real-time datacenter power draw per API call, so no\n> npm package — this one included — can hand you an audited electricity\n> bill for a `gpt-4o` or `claude` call. What this package gives you is a\n> transparent, order-of-magnitude estimate built from public research and\n> published/rumored model sizes, useful for **comparing** prompt strategies\n> against each other, not for **auditing** exact energy spend. See\n> [How accurate is this?](#how-accurate-is-this) before you trust a number\n> from this tool in a report.\n\n## Credit / inspiration\n\nThis project takes its name and framing from\n[**Intelligence per Watt (IPW)**](https://www.intelligence-per-watt.ai/),\nStanford Hazy Research's metric and empirical study of AI inference\nefficiency (Saad-Falcon et al., 2025). IPW measures *task accuracy per unit\nof power* across local models and hardware accelerators, and makes the case\nthat efficiency, not just raw capability, should guide how and where AI\nruns.\n\n`tokens-per-watt` asks a narrower, adjacent question: for **cloud API**\nusage specifically, roughly how much energy did this conversation or build\ncost, and would a different prompting strategy have cost less? It's not an\nimplementation of IPW and doesn't measure accuracy — it's a smaller,\npractical tool inspired by the same underlying concern: that power\nconsumption is a first-class metric for AI usage, not an afterthought. If\nyou're interested in efficiency at the model/hardware level rather than the\nprompting/API level, go read their [paper](https://arxiv.org/abs/2511.07885).\n\n## Install\n\n```bash\nnpm install tokens-per-watt\n```\n\n## Quickstart\n\n```ts\nimport { estimateEnergyWh, EnergyTracker } from \"tokens-per-watt\";\n\n// Single call\nconst result = estimateEnergyWh({\n  model: \"claude-3-5-sonnet-20241022\",\n  inputTokens: 1200,\n  outputTokens: 400,\n});\n// { wh: ..., kWh: ..., coefficientSource: \"known_model\", sizeClass: \"medium\", disclaimer: \"...\" }\n\n// Whole conversation / build\nconst tracker = new EnergyTracker();\n\n// Accepts raw OpenAI SDK usage shape\ntracker.add(\"gpt-4o\", response.usage); // { prompt_tokens, completion_tokens }\n\n// Accepts raw Anthropic SDK usage shape\ntracker.add(\"claude-3-5-haiku-20241022\", response.usage); // { input_tokens, output_tokens }\n\n// Or manual shape\ntracker.add(\"gpt-4o-mini\", { inputTokens: 300, outputTokens: 120 });\n\nconst report = tracker.summary();\nconsole.log(report.totalWh, report.perModel, report.disclaimer);\n```\n\n### CLI\n\n```bash\nnpx tpw calls.jsonl\n```\n\nEach line of `calls.jsonl` is one call:\n\n```json\n{\"provider\": \"anthropic\", \"model\": \"claude-3-5-sonnet\", \"inputTokens\": 1200, \"outputTokens\": 400}\n```\n\nThe CLI prints a summary table and, where relevant, suggestions from the\nadvisor (see below) — always labeled as estimates.\n\n## What it does\n\n- **Estimator** — converts input/output token counts into an estimated\n  Wh figure using a per-model coefficient (Wh per output token, with input\n  tokens priced at ~15% of output's per-token cost, reflecting that prefill\n  is parallelized and decode is sequential).\n- **Tracker** — accumulates calls across a conversation or build into a\n  total, with a per-model breakdown, optional cost estimate (if you supply\n  $/kWh), and optional CO2e estimate (if you supply a grid carbon\n  intensity — opt-in only, since this varies enormously by region and time\n  of day).\n- **Advisor** — takes a tracked session and estimates what a different\n  strategy would have cost: a smaller model in the same family, a shorter\n  system prompt, prompt caching, capped output length, or batching several\n  short calls into one.\n\n## How accurate is this?\n\nShort answer: **directionally useful, not audit-grade.**\n\nThe per-model coefficients in this package are assembled from public\nresearch on inference energy cost — notably Luccioni et al., *\"Power\nHungry Processing\"* (2023/2024), de Vries, *\"The growing energy footprint\nof AI\"* (2023), and public writeups of model sizes — combined with\nsize-class heuristics for anything not explicitly in the table. They are\n**not** vendor-disclosed per-call energy figures; no major provider\npublishes those.\n\nReal energy draw for any given API call depends on things this package\ncannot see from outside a datacenter:\n\n- Hardware generation and accelerator type\n- Batching and request concurrency at the time of your call\n- Utilization / how \"warm\" the serving fleet is\n- Datacenter Power Usage Effectiveness (PUE), cooling, and site efficiency\n\nUse this package to compare **relative** energy cost between prompting\nchoices (\"would a shorter prompt or a smaller model have used less\nenergy?\") rather than to produce a number you'd defend in an audit. Every\nobject this package returns includes a `disclaimer` field for this reason\n— please don't strip it out downstream.\n\nYou can override or extend any coefficient:\n\n```ts\nimport { setModelCoefficient } from \"tokens-per-watt\";\n\n// setModelCoefficient(model, whPerOutputToken, sizeClass, whPerInputToken?)\nsetModelCoefficient(\"my-custom-model\", 0.002, \"medium\", 0.0003);\n```\n\n## API reference\n\n### `estimateEnergyWh(input: EstimateInput): EstimateResult`\n\nEstimate energy for a single API call.\n\n```ts\ninterface EstimateInput {\n  model: string;\n  inputTokens: number;\n  outputTokens: number;\n  provider?: string;\n  reasoningTokens?: number; // hidden thinking tokens (o1, Claude extended thinking)\n}\n```\n\n### `EnergyTracker`\n\nAccumulates multiple calls into a session total.\n\n```ts\nconst tracker = new EnergyTracker({\n  dollarPerKwh: 0.12,    // optional — adds cost estimate\n  gCo2PerKwh: 400,       // optional — adds CO2e estimate (opt-in only)\n});\n\ntracker.add(model, usage, meta?);  // returns EstimateResult\ntracker.summary();                 // returns TrackerSummary\ntracker.getEntries();              // returns TrackerEntry[]\ntracker.reset();                   // clear all entries\n```\n\n### `advise(entries: TrackerEntry[]): AdvisorReport`\n\nGiven tracked entries, suggests alternate strategies and estimates energy savings.\n\n### `setModelCoefficient(model, whPerOutputToken, sizeClass, whPerInputToken?)`\n\nOverride or add a coefficient for any model ID.\n\n### `getCoefficients(model): { coefficients, source }`\n\nLook up the coefficient that would be used for a model.\n\n### `listKnownModels(): string[]`\n\nList all model IDs with built-in coefficients.\n\n## Methodology appendix\n\nThe coefficient table lives in `src/coefficients.ts`. Each model's\n`whPerOutputToken` value is a rough estimate based on:\n\n| Size class | Wh/output token | Basis |\n|---|---|---|\n| small (≤10B params) | ~0.0003–0.0004 | Luccioni et al. (2023) measurements of small models; Patterson et al. (2021) scaling |\n| medium (10–100B) | ~0.0012–0.0018 | Interpolation from Luccioni + de Vries (2023) growth curves |\n| large (100–200B) | ~0.0020–0.0025 | de Vries (2023) estimates for large dense models |\n| frontier (200B+, MoE-large) | ~0.0040–0.0055 | Extrapolation from Patterson/de Vries + public MoE size rumors |\n\nInput tokens use 15% of the output coefficient by default (prefill is\nparallelized across the sequence, decode is sequential per token).\n\nCached/reused context is estimated at 10% of normal input cost (a rough\napproximation — actual savings depend on KV-cache implementation).\n\n## License\n\nMIT","readmeFilename":"README.md","_rev":"1-f58cd190aa25dc27be08adc903b7b63c"}