{"_id":"@buildwithabid/llm-bench","name":"@buildwithabid/llm-bench","dist-tags":{"latest":"1.0.0"},"versions":{"1.0.0":{"name":"@buildwithabid/llm-bench","version":"1.0.0","description":"Terminal CLI to benchmark LLM providers and compare speed, cost, and response quality across OpenAI, Anthropic, Gemini, and Groq.","type":"module","main":"dist/cli.js","bin":{"llm-bench":"dist/cli.js"},"scripts":{"build":"tsc","dev":"tsx src/cli.ts","docs:assets":"node docs/scripts/generate-preview-assets.mjs","start":"node dist/cli.js","test":"npm run build && node --test test/**/*.test.mjs"},"keywords":["llm","benchmark","cli","openai","anthropic","gemini","groq"],"homepage":"https://github.com/BuildWithAbid/llm-bench","repository":{"type":"git","url":"git+https://github.com/BuildWithAbid/llm-bench.git"},"bugs":{"url":"https://github.com/BuildWithAbid/llm-bench/issues"},"license":"MIT","dependencies":{"@anthropic-ai/sdk":"^0.39.0","@google/genai":"^1.0.0","commander":"^13.0.0","groq-sdk":"^0.9.0","ink":"^6.0.0","ink-spinner":"^5.0.0","openai":"^4.80.0","react":"^19.0.0"},"devDependencies":{"@types/node":"^22.0.0","@types/react":"^19.0.0","tsx":"^4.19.0","typescript":"^5.7.0"},"engines":{"node":">=18.0.0"},"author":{"name":"Abid Ali"},"gitHead":"4c3ba01cd2e8027af4187df258b33832e523a4fe","types":"./dist/cli.d.ts","_id":"@buildwithabid/llm-bench@1.0.0","_nodeVersion":"24.13.1","_npmVersion":"11.9.0","dist":{"integrity":"sha512-ZPDTb+QhhvtkWEFVp0XyRwg3IbwZVsXPXWdQ/FZ29JiNt04Qeb9gK4nJo+azptThuPyPfcn9v34d7a6pWdKZlA==","shasum":"59e71fb75ad81c383432e34669e5d519e8f55452","tarball":"https://registry.npmjs.org/@buildwithabid/llm-bench/-/llm-bench-1.0.0.tgz","fileCount":36,"unpackedSize":77967,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCICKkgWsvPe3BHX04E3VF33mUQkrOFZIh+bJL+ftWBoPJAiBmpDkc3eKdPt+41ma2F6KdFQSz3IAUhsrlGp8MEwtVQQ=="}]},"_npmUser":{"name":"buildwithabid","email":"abidtech2017@gmail.com"},"directories":{},"maintainers":[{"name":"buildwithabid","email":"abidtech2017@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/llm-bench_1.0.0_1785669048192_0.34894879020415637"},"_hasShrinkwrap":false}},"time":{"created":"2026-08-02T11:10:47.979Z","1.0.0":"2026-08-02T11:10:48.321Z","modified":"2026-08-02T11:10:48.559Z"},"maintainers":[{"name":"buildwithabid","email":"abidtech2017@gmail.com"}],"description":"Terminal CLI to benchmark LLM providers and compare speed, cost, and response quality across OpenAI, Anthropic, Gemini, and Groq.","homepage":"https://github.com/BuildWithAbid/llm-bench","keywords":["llm","benchmark","cli","openai","anthropic","gemini","groq"],"repository":{"type":"git","url":"git+https://github.com/BuildWithAbid/llm-bench.git"},"author":{"name":"Abid Ali"},"bugs":{"url":"https://github.com/BuildWithAbid/llm-bench/issues"},"license":"MIT","readme":"# llm-bench\n\n`llm-bench` is a terminal CLI for benchmarking large language model APIs side by side. It runs multiple providers in parallel, streams live progress in the terminal, ranks the successful responses, and generates shareable result cards.\n\nIf you need a lightweight way to compare OpenAI, Anthropic, Gemini, and Groq models for latency, price, and response quality, this project is built for that workflow.\n\n## Product Preview\n\n### Live terminal race\n\n![LLM Bench terminal benchmark demo](./docs/assets/terminal-race-demo.gif)\n\n### Shareable result card\n\n![LLM Bench result card preview](./docs/assets/result-card-preview.png)\n\n## What This Product Actually Does\n\nAt runtime, `llm-bench`:\n\n1. Accepts a prompt from the command line.\n2. Selects the configured providers whose API keys are available.\n3. Sends the same prompt, and optional system prompt, to each provider in parallel.\n4. Streams live status, token estimates, elapsed time, and cost estimates in an Ink-based terminal UI.\n5. Scores every successful response using a built-in heuristic for speed, cost, and response quality.\n6. Ranks the successful models from best overall score to lowest.\n7. Writes two report artifacts to disk.\n\n## What It Creates\n\nEvery successful benchmark run creates:\n\n- A live terminal race dashboard while providers are running.\n- A ranked benchmark result for the current session.\n- A plain text summary card, by default `result-card.txt`.\n- A styled HTML summary card, by default `result-card.html`.\n\nWhen you pass `--output ./results/my-run`, the tool creates:\n\n- `./results/my-run.txt`\n- `./results/my-run.html`\n\nIf the parent directory does not exist, `llm-bench` creates it automatically.\n\n## What It Does Not Do\n\n`llm-bench` is intentionally focused. It does not currently:\n\n- Store benchmark history across runs.\n- Produce JSON or CSV output.\n- Use an LLM judge or human review for quality scoring.\n- Pull live pricing from provider APIs.\n\nIf all selected providers fail, the run exits with a non-zero status instead of reporting a false success.\n\n## Core Features\n\n- Parallel provider execution for fast comparisons.\n- Real-time terminal UI built with Ink and React.\n- Provider-level failure handling so one failed model does not stop an otherwise successful run.\n- Cross-provider system prompt support.\n- Automatic result card generation in text and HTML formats.\n- Config file support via `llm-bench.config.json`.\n- Selectable model filters for supported provider slugs.\n\n## Supported Providers and Model Slugs\n\nThese are the recommended model slugs for `--models`:\n\n| Provider | Display Name | CLI Slug | Environment Variable |\n| --- | --- | --- | --- |\n| OpenAI | GPT-4o | `gpt-4o` | `OPENAI_API_KEY` |\n| Anthropic | Claude Sonnet 4.6 | `claude-sonnet-4-6` | `ANTHROPIC_API_KEY` |\n| Google | Gemini 2.5 Flash | `gemini-2.5-flash` | `GOOGLE_API_KEY` |\n| Groq | Llama 3.3-70b | `llama-3.3-70b` | `GROQ_API_KEY` |\n\nNotes:\n\n- The Groq integration uses `llama-3.3-70b-versatile` internally, but the CLI accepts `llama-3.3-70b`.\n- Provider display names also work as filters, but the slugs above are the recommended inputs.\n\n## How Benchmarking Works\n\n### 1. Provider Selection\n\nThe CLI starts with all built-in providers, then filters them by:\n\n- `--models` if provided\n- `llm-bench.config.json` defaults if present\n- Available API keys in the environment\n\nProviders without API keys are skipped with a warning.\n\n### 2. Prompt Dispatch\n\nEach selected provider receives:\n\n- The same user prompt\n- The same optional system prompt\n\nProvider-specific implementations map the system prompt into the native API shape:\n\n- OpenAI and Groq use system messages.\n- Anthropic uses the `system` field.\n- Gemini uses `systemInstruction`.\n\n### 3. Live Updates\n\nDuring streaming, the terminal UI updates:\n\n- Provider status\n- Estimated output token count\n- Elapsed time\n- Estimated USD cost\n\nLive token counts are estimated from streamed text, not from callback count, which makes the live display less sensitive to SDK chunking behavior.\n\n### 4. Final Scoring\n\nAfter at least one provider finishes successfully, results are scored across:\n\n- Speed: 30%\n- Cost: 30%\n- Quality heuristic: 40%\n\nIf every provider fails, the run fails instead of returning an empty result set.\n\n### 5. Artifact Generation\n\nOn successful completion, `llm-bench` writes:\n\n- A text result card\n- An HTML result card\n\nIf card generation fails, the process exits non-zero.\n\n## Installation\n\n### Install from source\n\n> **Not published to a registry yet.** Install from source with the steps below; the commands in this README assume you have done so.\n\n\n```bash\ngit clone https://github.com/BuildWithAbid/llm-bench.git\ncd llm-bench\nnpm install\nnpm run build\nnpm link            # makes `llm-bench` available on your PATH\n```\n\n## Environment Setup\n\nSet one or more provider API keys before running the CLI.\n\n### Bash or Zsh\n\n```bash\nexport OPENAI_API_KEY=\"sk-...\"\nexport ANTHROPIC_API_KEY=\"sk-ant-...\"\nexport GOOGLE_API_KEY=\"AIza...\"\nexport GROQ_API_KEY=\"gsk_...\"\n```\n\n### PowerShell\n\n```powershell\n$env:OPENAI_API_KEY = \"sk-...\"\n$env:ANTHROPIC_API_KEY = \"sk-ant-...\"\n$env:GOOGLE_API_KEY = \"AIza...\"\n$env:GROQ_API_KEY = \"gsk_...\"\n```\n\n## Quick Start\n\nRun all available providers:\n\n```bash\nllm-bench run \"Explain vector databases in simple terms\"\n```\n\nRun only specific models:\n\n```bash\nllm-bench run \"Write a Redis caching strategy for an API\" --models gpt-4o,claude-sonnet-4-6,llama-3.3-70b\n```\n\nAdd a system prompt:\n\n```bash\nllm-bench run \"Explain Docker Compose\" --system \"Respond like a senior DevOps engineer. Be concise and practical.\"\n```\n\nWrite output to a custom directory:\n\n```bash\nllm-bench run \"Compare REST and GraphQL for mobile apps\" --output ./results/api-comparison\n```\n\n## CLI Reference\n\n### Command\n\n```text\nllm-bench run <prompt> [options]\n```\n\n### Arguments\n\n| Argument | Description |\n| --- | --- |\n| `prompt` | The prompt sent to each selected provider |\n\n### Options\n\n| Option | Description |\n| --- | --- |\n| `-m, --models <models>` | Comma-separated list of models to race |\n| `-s, --system <prompt>` | Optional system prompt |\n| `-o, --output <path>` | Base path for `.txt` and `.html` result cards |\n| `-c, --config <path>` | Optional path to a config file |\n| `-V, --version` | Print the installed version |\n| `-h, --help` | Show help |\n\n### Examples\n\n```bash\nllm-bench run \"Summarize event-driven architecture\"\nllm-bench run \"Explain rate limiting\" --models gemini-2.5-flash,gpt-4o\nllm-bench run \"Explain CQRS\" --system \"Answer for a staff engineer audience.\"\nllm-bench run \"Compare SQL and NoSQL\" --output ./benchmarks/sql-nosql\nllm-bench run \"What is recursion?\" --config ./llm-bench.config.json\n```\n\nFor a dedicated command guide, see [docs/cli-reference.md](./docs/cli-reference.md).\n\n## Configuration\n\nBy default, `llm-bench` looks for `llm-bench.config.json` in the current working directory. You can also pass a custom file path with `--config`.\n\nExample:\n\n```json\n{\n  \"models\": [\"gpt-4o\", \"claude-sonnet-4-6\", \"gemini-2.5-flash\"],\n  \"systemPrompt\": \"Be concise and direct. Use short paragraphs.\"\n}\n```\n\n### Supported config fields\n\n| Field | Type | Description |\n| --- | --- | --- |\n| `models` | `string[]` | Default model filters |\n| `systemPrompt` | `string` | Default system prompt |\n\nCLI flags override config values.\n\n## Scoring Methodology\n\n`llm-bench` does not claim to produce a rigorous research-grade evaluation. It uses a simple built-in heuristic so you can compare providers quickly without running a second judge model.\n\n### Speed score\n\nThe fastest successful model gets `10`. Other models are scaled proportionally.\n\n```text\nspeedScore = (fastestTime / modelTime) * 10\n```\n\n### Cost score\n\nThe cheapest successful model gets `10`. Other models are scaled proportionally.\n\n```text\ncostScore = (cheapestCost / modelCost) * 10\n```\n\n### Quality score\n\nQuality is estimated from two signals:\n\n- Response length contributes 40%.\n- Keyword overlap with the prompt contributes 60%.\n\nThe implementation extracts prompt words longer than 3 characters, then checks whether those words appear in the response.\n\n### Overall score\n\n```text\noverallScore = (speed * 0.3) + (cost * 0.3) + (quality * 0.4)\n```\n\n## Output Artifacts\n\n### Text card\n\nThe text card is intended for:\n\n- Terminal logs\n- Team chat\n- Issue comments\n- Plain text reports\n\nDefault output file:\n\n```text\nresult-card.txt\n```\n\n### HTML card\n\nThe HTML card is intended for:\n\n- Screenshotting benchmark results\n- Sharing results internally\n- Embedding a simple benchmark summary in documentation\n\nDefault output file:\n\n```text\nresult-card.html\n```\n\n## Failure Behavior and Exit Codes\n\n`llm-bench` handles failures at two levels.\n\n### Provider-level failures\n\nIf one provider fails but at least one other provider succeeds:\n\n- The failed provider is marked as failed in the UI.\n- Successful providers are still scored.\n- Result cards are still generated from the successful providers.\n\n### Run-level failures\n\nThe process exits with code `1` when:\n\n- No selected providers are available because API keys are missing.\n- Every selected provider fails.\n- Result card generation fails after a run.\n\nThe process exits with code `0` when:\n\n- At least one provider finishes successfully and the cards are written successfully.\n\n## Architecture Overview\n\nThe implementation is intentionally simple:\n\n- `src/cli.ts`: parses arguments, loads config, filters providers, starts the UI\n- `src/ui.tsx`: renders the live terminal experience and coordinates final output writing\n- `src/runner.ts`: runs providers in parallel and aggregates results\n- `src/scoring.ts`: computes speed, cost, quality, and overall scores\n- `src/card.ts`: generates and writes text and HTML result cards\n- `src/providers/*`: provider-specific streaming integrations\n- `src/tokens.ts`: shared token estimation helpers used for live display and fallbacks\n\nFor a deeper breakdown, see [docs/architecture.md](./docs/architecture.md).\n\n## Development\n\n### Prerequisites\n\n- Node.js 18+\n- npm\n- At least one provider API key for live testing\n\n### Start in development mode\n\n```bash\nnpm run dev -- run \"Explain retrieval-augmented generation\"\n```\n\n### Build\n\n```bash\nnpm run build\n```\n\n### Test\n\n```bash\nnpm test\n```\n\n### Regenerate preview assets\n\n```bash\nnpm run docs:assets\n```\n\nThis asset script expects a local Chrome or Edge installation plus Python with Pillow available.\n\nCurrent automated tests cover:\n\n- Total benchmark failure behavior\n- Live token estimation behavior\n- Nested output directory creation\n\n## Extending the Project\n\nTo add a new provider:\n\n1. Create a new file under `src/providers/`.\n2. Implement the `LLMProvider` interface.\n3. Return `text`, `inputTokens`, and `outputTokens`.\n4. Add the provider to `ALL_PROVIDERS` in `src/cli.ts`.\n\nWhen adding a provider, keep the following in mind:\n\n- System prompt handling should use the provider's native API shape.\n- Live streaming should call `onToken` with text chunks as they arrive.\n- Token usage should use provider-native usage fields when available.\n- Cost values are hard-coded in the provider definition and should be reviewed when pricing changes.\n\n## Use Cases\n\nThis project is a good fit for:\n\n- Comparing LLM API latency before shipping a feature\n- Estimating the cost tradeoff between providers\n- Running quick model bake-offs during development\n- Creating lightweight benchmark artifacts for internal teams\n- Testing how prompt changes behave across providers\n\n## Limitations\n\nTo keep expectations aligned with the current implementation:\n\n- Quality scoring is heuristic, not semantic evaluation.\n- Pricing is static in source code, not fetched from provider pricing pages.\n- Only built-in providers are benchmarked unless you extend the code.\n- There is no persistent benchmark history or database.\n\n## Documentation Index\n\n- [README](./README.md)\n- [Landing Page Docs](./docs/landing-page.md)\n- [CLI Reference](./docs/cli-reference.md)\n- [Architecture Guide](./docs/architecture.md)\n\n## License\n\nMIT\n","readmeFilename":"README.md","_rev":"1-5ca5d183f48a266d0de03db02caf694c"}