{"_id":"@aktagon/llmkit-ts","name":"@aktagon/llmkit-ts","dist-tags":{"latest":"2.1.0"},"versions":{"2.1.0":{"name":"@aktagon/llmkit-ts","version":"2.1.0","description":"Unified LLM client library for TypeScript — one API for many providers (Anthropic, OpenAI, Google Gemini, AWS Bedrock, Mistral, Groq, DeepSeek, and more), zero runtime dependencies. Bun, Node, Deno, Cloudflare Workers.","type":"module","main":"./dist/llmkit.js","module":"./dist/llmkit.js","types":"./dist/llmkit.d.ts","sideEffects":false,"exports":{".":{"types":"./dist/llmkit.d.ts","import":"./dist/llmkit.js"},"./builders":{"types":"./dist/builders/index.d.ts","import":"./dist/builders/index.js"}},"keywords":["llm","ai","llm-client","ai-sdk","agents","anthropic","claude","openai","gpt","gemini","bedrock","mistral","groq","streaming","tool-calling","typescript"],"license":"MIT","repository":{"type":"git","url":"git+https://github.com/aktagon/llmkit-ts.git"},"bugs":{"url":"https://github.com/aktagon/llmkit-ts/issues"},"homepage":"https://llmkit.aktagon.com","engines":{"node":">=18","bun":">=1.0.0"},"scripts":{"build":"rm -rf dist && bun build src/llmkit.ts src/builders/index.ts --outdir dist --target browser --format esm --splitting --sourcemap=external && bun run build:types","build:types":"tsc -p tsconfig.build.json","test":"bun test","typecheck":"tsc --noEmit","prepack":"bun run build","prepare":"bun run build"},"devDependencies":{"@types/bun":"^1.1.0","typescript":"^5.7"},"_id":"@aktagon/llmkit-ts@2.1.0","gitHead":"9c7d9a3f18514c9aa99649a3cb3dd2ba3ebe3bc6","_nodeVersion":"23.10.0","_npmVersion":"10.9.2","dist":{"integrity":"sha512-ruMU3KGf9KOROHRnmiPRZgwfWwAPH7WlhVX1c0uTY/5vRbHD25uhm5qd/S3zaFE6RMbwFvZjewHMsmsgNysIMg==","shasum":"be7015d7de421abfa2b62544b7c00ac669972269","tarball":"https://registry.npmjs.org/@aktagon/llmkit-ts/-/llmkit-ts-2.1.0.tgz","fileCount":168,"unpackedSize":1655863,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQCMVSKKiP755Mq2VI/XGoaulC2HlU+rLNiYuLlRAcVhqwIhAK2+5VATvaOlsekvUQc1dcYWTTA2R7xOXsKKfBphJzZZ"}]},"_npmUser":{"name":"llmkit-ts","email":"christian@aktagon.com"},"directories":{},"maintainers":[{"name":"llmkit-ts","email":"christian@aktagon.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/llmkit-ts_2.1.0_1784537495381_0.040888594328755135"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-20T08:51:35.213Z","2.1.0":"2026-07-20T08:51:35.539Z","modified":"2026-07-20T08:51:35.843Z"},"maintainers":[{"name":"llmkit-ts","email":"christian@aktagon.com"}],"description":"Unified LLM client library for TypeScript — one API for many providers (Anthropic, OpenAI, Google Gemini, AWS Bedrock, Mistral, Groq, DeepSeek, and more), zero runtime dependencies. Bun, Node, Deno, Cloudflare Workers.","homepage":"https://llmkit.aktagon.com","keywords":["llm","ai","llm-client","ai-sdk","agents","anthropic","claude","openai","gpt","gemini","bedrock","mistral","groq","streaming","tool-calling","typescript"],"repository":{"type":"git","url":"git+https://github.com/aktagon/llmkit-ts.git"},"bugs":{"url":"https://github.com/aktagon/llmkit-ts/issues"},"license":"MIT","readme":"# @aktagon/llmkit-ts\n\nOne TypeScript API for Anthropic, OpenAI, Google, and 20+ other providers — including local models through Ollama and vLLM. Switch providers without rewriting your request.\n\nNo runtime dependencies. Runs on Node ≥18, Bun, Deno, Cloudflare Workers, or any modern bundler (Vite, Next.js, esbuild, webpack 5+) — anywhere with `fetch` and Web Crypto.\n\nAlso available for Go, Python, Rust, Swift, and Java.\n\n<p align=\"center\">\n  <img src=\"https://raw.githubusercontent.com/aktagon/llmkit-ts/master/assets/logos/llmkit-languages.svg\" alt=\"Go, TypeScript, Python, Rust, Swift, Java\" height=\"26\">\n</p>\n<p align=\"center\">\n  <img src=\"https://raw.githubusercontent.com/aktagon/llmkit-ts/master/assets/logos/llmkit-providers.svg\" alt=\"Anthropic, OpenAI, Google, and 26 more providers\" height=\"26\">\n</p>\n\n## Install\n\nFrom npm:\n\n```bash\nbun add @aktagon/llmkit-ts\n# or\nnpm install @aktagon/llmkit-ts\n```\n\nFrom GitHub (skip the npm publish loop):\n\n```bash\nbun add github:aktagon/llmkit-ts#ts-v1.0.1\n# or\nnpm install github:aktagon/llmkit-ts#ts-v1.0.1\n```\n\nThe package ships compiled ESM in `dist/` (works in plain Node ESM, Workers, Deno) plus the TypeScript source in `src/` (consumed for type info by Bun, Vite, Next.js, and any bundler with `moduleResolution: \"bundler\"`). No build step required at the consumer.\n\n## Quick Start\n\n```ts\nimport { anthropic } from \"@aktagon/llmkit-ts/builders\";\n\nconst c = anthropic(process.env.ANTHROPIC_API_KEY!);\nconst resp = await c.text\n  .system(\"You are concise.\")\n  .prompt(\"Why is the sky blue?\");\n\nconsole.log(resp.text);\nconsole.log(resp.usage.input, resp.usage.output);\n```\n\n`c.text`, `c.image`, `c.music`, `c.video`, `c.agent`, and `c.upload` are fields on the `Client` — access them without parentheses. Chain methods (`.system(...)`, `.temperature(...)`) clone the builder and return the clone, so a forked chain shares no state with its parent. The typed builder is the only public surface as of v1.0.0. One mental model — `client.<capability>.<chain>.<terminal>` — across every capability.\n\nRunnable examples for each capability live in [`examples/`](./examples); `tests/examples.test.ts` exercises every documented call shape against a mock HTTP server, so the snippets in this README cannot drift from the actual API surface.\n\n## Providers\n\n| Provider  | Default model                               | Env var           |\n| --------- | ------------------------------------------- | ----------------- |\n| anthropic | claude-sonnet-4-6                           | ANTHROPIC_API_KEY |\n| openai    | gpt-4o                                      | OPENAI_API_KEY    |\n| google    | gemini-2.5-flash                            | GOOGLE_API_KEY    |\n| bedrock   | anthropic.claude-sonnet-4-20250514-v1:0     | AWS_ACCESS_KEY_ID |\n| grok      | grok-3-fast                                 | GROK_API_KEY      |\n| mistral   | mistral-large-latest                        | MISTRAL_API_KEY   |\n| deepseek  | deepseek-chat                               | DEEPSEEK_API_KEY  |\n| groq      | llama-3.3-70b-versatile                     | GROQ_API_KEY      |\n| together  | meta-llama/Llama-3.3-70B-Instruct-Turbo     | TOGETHER_API_KEY  |\n| cohere    | command-r-plus                              | COHERE_API_KEY    |\n| ai21      | jamba-1.5-large                             | AI21_API_KEY      |\n| cerebras  | llama-3.3-70b                               | CEREBRAS_API_KEY  |\n| ...       | (full list in `src/providers/providers.ts`) |                   |\n\n36 providers, 4 API shapes (OpenAI-compatible, Anthropic Messages, Google Generative AI, AWS Bedrock Converse). Bedrock auth uses SigV4; other providers use API-key auth.\n\nPer-provider factory functions: `ai21`, `anthropic`, `assemblyai`, `azure`, `bedrock`, `cerebras`, `cohere`, `deepseek`, `doubao`, `ernie`, `fireworks`, `google`, `grok`, `groq`, `inworld`, `jan`, `llamacpp`, `lmstudio`, `minimax`, `mistral`, `moonshot`, `ollama`, `openai`, `openrouter`, `perplexity`, `pixverse`, `qwen`, `recraft`, `sambanova`, `together`, `vertex`, `vidu`, `vllm`, `workersai`, `yi`, `zhipu`. Or use the generic `newClient(name, key)`.\n\n## API\n\n### Text — one-shot prompt\n\n```ts\nconst resp = await c.text\n  .system(\"You are helpful\")\n  .temperature(0.7)\n  .maxTokens(200)\n  .prompt(\"What is 2+2?\");\n\nconsole.log(resp.text); // \"4\"\nconsole.log(resp.usage.input); // prompt tokens\nconsole.log(resp.usage.output); // completion tokens\nconsole.log(resp.usage.cacheRead); // tokens served from cache\nconsole.log(resp.usage.cacheWrite); // tokens written to cache (Anthropic explicit)\nconsole.log(resp.usage.reasoning); // internal reasoning tokens (OpenAI o-series, Gemini 2.5+)\n```\n\nCapability-scoped fields (`cacheRead`, `cacheWrite`, `reasoning`) are zero when the provider doesn't report them separately.\n\n### Stream — chunks + trailing handle\n\n<!-- llmkit:include ts/examples/streaming.ts#stream -->\n```ts\nconst stream = client.text\n  .system(\"Be brief\")\n  .stream(\"Tell me a one-line joke\");\nfor await (const chunk of stream) {\n  process.stdout.write(chunk);\n}\nprocess.stdout.write(\"\\n\");\nconst final = stream.response();\nif (final !== null) {\n  console.log(\n    `input=${final.usage.input} output=${final.usage.output} ` +\n      `finishReason=${final.finishReason ?? \"\"}`,\n  );\n}\n```\n\n`TextStream` implements `AsyncIterable<string>`. After iteration completes, `stream.response()` returns the final `Response` (with token counts) and `stream.error()` returns any terminal error. Handles both Anthropic-style typed events and OpenAI-style data-only frames internally.\n\n### Agent — tool loop\n\n```ts\nimport type { Tool } from \"@aktagon/llmkit-ts\";\n\nconst add: Tool = {\n  name: \"add\",\n  description: \"Add two numbers\",\n  schema: {\n    type: \"object\",\n    properties: {\n      a: { type: \"number\" },\n      b: { type: \"number\" },\n    },\n  },\n  run: ({ a, b }) => String(Number(a) + Number(b)),\n};\n\nconst bot = c.agent\n  .system(\"You are a calculator.\")\n  .addTool(add)\n  .maxToolIterations(5);\n\nconst resp = await bot.prompt(\"What is 2+3?\");\nconsole.log(resp.text);\n```\n\n`*Agent` is **stateful** — repeated `bot.prompt(...)` calls accumulate history. Chain methods (`.system(...)`, `.addTool(...)`) clone and reset state, so a forked builder gets a fresh conversation. `bot.reset()` clears state without dropping chained config.\n\nTool dispatch covers Anthropic `tool_use`, OpenAI `tool_calls`, Google `functionCall`, and Bedrock Converse `toolUse`. Tool errors surface to the model as the result string verbatim — sanitise tool inputs at the source.\n\n### Image input (vision)\n\nAttach an image to a text prompt with `.image(mime, bytes)`; it is sent as the\nprovider's native image block (works on Anthropic, OpenAI, Google, and\nBedrock). Bytes-based, so it works with no filesystem (e.g. a browser\nextension passing a screenshot straight through):\n\n```ts\nconst resp = await c.text\n  .image(\"image/png\", screenshotBytes)\n  .prompt(\"Describe this screenshot in one sentence.\");\n```\n\n### Image — text-to-image and edit\n\n```ts\nimport { google } from \"@aktagon/llmkit-ts/builders\";\n\nconst c = google(process.env.GOOGLE_API_KEY!);\nconst img = await c.image\n  .model(\"gemini-3.1-flash-image-preview\")\n  .aspectRatio(\"16:9\")\n  .imageSize(\"2K\")\n  .generate(\"A nano banana dish, studio lighting\");\n\nawait Bun.write(\"out.png\", img.images[0]!.bytes);\n```\n\nFor compositional editing, chain `.text(...)` and `.image(mime, bytes)` to interleave references with descriptions. The terminal `msg` is appended as a final text Part:\n\n```ts\nawait c.image\n  .model(\"gemini-3.1-flash-image-preview\")\n  .text(\"Person:\")\n  .image(\"image/png\", personBytes)\n  .text(\"Outfit:\")\n  .image(\"image/png\", outfitBytes)\n  .generate(\"Generate the person wearing the outfit.\");\n```\n\nAspect ratios and sizes validate against a per-model whitelist before the HTTP request — `imageSize(\"512\")` on Pro throws `ValidationError` without paying for a 4xx round-trip. Empty whitelists mean \"no client-side check; pass through\" — providers like OpenAI accept arbitrary sizes within documented bounds, so the SDK trusts the API boundary instead of carrying a stale list.\n\n| Provider | Model                          | Aspect ratios                                                                   | Sizes                               |\n| -------- | ------------------------------ | ------------------------------------------------------------------------------- | ----------------------------------- |\n| Google   | Nano Banana 2 (Flash)          | 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9, **1:4, 4:1, 1:8, 8:1**     | 512, 1K, 2K, 4K                     |\n| Google   | Nano Banana Pro                | 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9                             | 1K, 2K, 4K                          |\n| OpenAI   | gpt-image-2 / 1.5 / 1 / 1-mini | n/a (size only)                                                                 | any (e.g. `1024x1024`, `1536x1024`) |\n| xAI      | grok-imagine-image-quality     | 1:1, 2:3, 3:2, 3:4, 4:3, 9:16, 16:9, 1:2, 2:1, 19.5:9, 9:19.5, 20:9, 9:20, auto | 1k, 2k                              |\n| Vertex   | imagen-3.0 / 4.0               | 1:1, 9:16, 16:9, 3:4, 4:3                                                       | fixed per model                     |\n\nOpenAI gpt-image-\\* models accept arbitrary sizes within documented bounds (max edge ≤3840, both edges multiples of 16, ratio ≤3:1, total pixels 655K–8.3M). They always return base64-encoded images, so `resp.images[0].bytes` works the same on both providers.\n\nProvider knobs are typed chain methods on the `Image` builder:\n\n| Method               | Provider support            | Wire field       |\n| -------------------- | --------------------------- | ---------------- |\n| `.quality(s)`        | OpenAI gpt-image-\\*         | `quality`        |\n| `.outputFormat(s)`   | OpenAI gpt-image-\\*         | `output_format`  |\n| `.background(s)`     | OpenAI gpt-image-\\*         | `background`     |\n| `.count(n)`          | OpenAI + xAI Grok           | `n`              |\n| `.mask(mime, bytes)` | OpenAI gpt-image-\\* (edits) | multipart `mask` |\n\nThe chain validates per provider — calling `.quality(...)` on a Google or xAI builder rejects with `ValidationError` immediately, no HTTP round-trip. Knobs without typed methods (OpenAI: `output_compression`, `moderation`) remain reachable via `.extraFields(...)`, which is unvalidated and freeform.\n\n```ts\nimport { openai } from \"@aktagon/llmkit-ts/builders\";\n\nconst c = openai(process.env.OPENAI_API_KEY!);\nconst resp = await c.image\n  .model(\"gpt-image-2\")\n  .imageSize(\"1024x1024\")\n  .quality(\"high\")\n  .count(4)\n  .generate(\"A red circle on a white background\");\n```\n\nDispatch is automatic: chains without image parts hit OpenAI's `/v1/images/generations` (JSON); chains carrying one or more `.image(...)` parts hit `/v1/images/edits` (multipart/form-data with one `image[]` field per reference, in caller order). gpt-image-\\* requires organization verification — see [platform.openai.com/docs/guides/your-data#organization-verification](https://platform.openai.com/docs/guides/your-data#organization-verification).\n\nUp to 14 reference images per Google request, 16 per OpenAI request.\n\n#### Vertex AI Imagen (Google Cloud)\n\nVertex Imagen uses the `:predict` endpoint family and OAuth bearer auth instead of API keys. The SDK takes a bearer token (string); caller manages OAuth refresh externally (e.g. `gcloud auth print-access-token`, service-account JSON, or workload identity).\n\n```ts\nimport { vertex } from \"@aktagon/llmkit-ts/builders\";\n\n// Caller substitutes {project_id} and {location} before passing the URL.\nconst baseUrl =\n  \"https://us-central1-aiplatform.googleapis.com\" +\n  \"/v1/projects/my-gcp-project/locations/us-central1/publishers/google/models\";\n\nconst c = vertex(process.env.VERTEX_BEARER_TOKEN!).baseURL(baseUrl);\n\nconst resp = await c.image\n  .model(\"imagen-3.0-generate-002\")\n  .aspectRatio(\"16:9\")\n  .count(2)\n  .generate(\"A red circle\");\n```\n\nEdit-mode (single image into `instances[0].image`) and inpainting (`.mask(mime, bytes)` into `instances[0].mask.image`) work the same way. Imagen-specific knobs like `negativePrompt` and `safetySetting` are reachable through `.extraFields(...)` — they spread into the request's `parameters` block. Vertex's `:predict` response does not carry token counts; `resp.usage` stays zero.\n\n### Music — text-to-music\n\nGenerate audio from a text prompt via the typed-builder chain on `c.music`. Decoded audio bytes come back on `resp.audio[0].bytes`. Models that support vocals take lyrics via `.lyrics(...)` (use section tags like `[verse]`); instrumental-only models reject lyrics before the request is sent.\n\n<!-- llmkit:include ts/examples/music.ts#music -->\n```typescript\nconst r = await client.music\n  .model(\"lyria-002\")\n  .generate(\"a calm instrumental, warm piano and soft strings\");\n\nconst first = r.audio[0];\nif (!first) throw new Error(\"no audio returned\");\nawait Bun.write(\"out.wav\", first.bytes);\n```\n\nModels with vocals take lyrics via `.lyrics(...)`:\n\n```typescript\nconst song = await c.music.model(\"lyria-3-pro-preview\").lyrics(\"[verse] neon lights\").generate(\"dream pop, 90 bpm\");\n```\n\n| Provider | Model(s)                                      | Lyrics | Output     |\n| -------- | --------------------------------------------- | ------ | ---------- |\n| Vertex   | `lyria-002`                                   | no     | WAV (~30s) |\n| Google   | `lyria-3-pro-preview`, `lyria-3-clip-preview` | yes    | MP3        |\n| MiniMax  | `music-2.6`                                   | yes    | MP3        |\n\n### Video — text-to-video\n\nGenerate video from a text prompt. Video generation is asynchronous: `submit` returns a handle immediately, and `handle.wait()` polls until the job finishes. The result carries a temporary hosted URL on `resp.videos[0].url` — download it yourself.\n\n<!-- llmkit:include ts/examples/video.ts#video -->\n```typescript\nconst handle = await client.video\n  .model(\"grok-imagine-video\")\n  .submit(\n    \"a slow cinematic drone shot flying over snow-capped alpine peaks at golden hour\",\n  );\n\nconst r = await handle.wait();\n\nconst v = r.videos[0];\nif (!v) throw new Error(\"no video returned\");\nconsole.log(\n  `done: url=${v.url} duration=${v.durationSeconds}s mime=${v.mimeType}`,\n);\n```\n\n| Provider | Model                | Delivery |\n| -------- | -------------------- | -------- |\n| Grok     | `grok-imagine-video` | URL      |\n\n### Safety Settings\n\nControl content filtering for Gemini providers. `safetySettings` applies to text\ngeneration, streaming, agents, and Gemini image generation. `safetyFilter` applies\nto Vertex Imagen only.\n\n```ts\nimport {\n  google,\n  vertex,\n  HARM_CATEGORY_DANGEROUS_CONTENT,\n  HARM_CATEGORY_HARASSMENT,\n  HARM_BLOCK_THRESHOLD_NONE,\n  HARM_BLOCK_THRESHOLD_HIGH_ONLY,\n  IMAGE_SAFETY_FILTER_BLOCK_FEW,\n} from \"@aktagon/llmkit-ts/builders\";\n\n// Gemini text or agent\nconst c = google(process.env.GOOGLE_API_KEY!);\nconst resp = await c.text\n  .safetySettings([\n    {\n      category: HARM_CATEGORY_DANGEROUS_CONTENT,\n      threshold: HARM_BLOCK_THRESHOLD_NONE,\n    },\n    {\n      category: HARM_CATEGORY_HARASSMENT,\n      threshold: HARM_BLOCK_THRESHOLD_HIGH_ONLY,\n    },\n  ])\n  .prompt(\"Write a story\");\n\n// Vertex Imagen\nconst vc = vertex(process.env.VERTEX_BEARER_TOKEN!);\nconst img = await vc.image\n  .model(\"imagen-3.0-generate-002\")\n  .safetyFilter(IMAGE_SAFETY_FILTER_BLOCK_FEW)\n  .generate(\"A landscape\");\n```\n\n`safetySettings` on Vertex Imagen and `safetyFilter` on non-Imagen providers throw\na `ValidationError`. The `HARM_CATEGORY_*`, `HARM_BLOCK_THRESHOLD_*`, and\n`IMAGE_SAFETY_FILTER_*` constants cover all documented values; raw strings also work.\n\n### Upload — Path or Bytes\n\n```ts\nimport { openai } from \"@aktagon/llmkit-ts/builders\";\n\nconst c = openai(process.env.OPENAI_API_KEY!);\n\n// from a path (Node/Bun only)\nconst file = await c.upload.path(\"./data.pdf\").run();\n\n// from bytes (works everywhere)\nconst file2 = await c.upload\n  .bytes(buf) // Uint8Array\n  .filename(\"report.pdf\")\n  .mimeType(\"application/pdf\")\n  .run();\n```\n\nThe `.path()` branch dynamically loads `node:fs/promises` and is unavailable in browsers / Cloudflare Workers / Deno without `--allow-read`. Use `.bytes()` for portable code.\n\n### Batches\n\n<!-- llmkit:include ts/examples/batch.ts#batch -->\n```ts\nconst handle = await client.text\n  .system(\"Be brief\")\n  .batch(\n    \"Translate hello to French\",\n    \"Translate hello to Spanish\",\n    \"Translate hello to German\",\n  );\nconst results = await handle.wait();\nresults.forEach((r) => console.log(r.text));\n```\n\n`c.text.<config>.batch(...prompts)` queues the batch and returns a `BatchHandle` you can persist. Call `handle.wait()` to block until completion, or `handle.poll()` to drive the loop yourself. The blocking one-liner is `(await c.text.batch(...prompts)).wait()`. Both inline (Anthropic) and file-reference (OpenAI two-hop) flows are handled internally.\n\n### Caching\n\n```ts\n// Anthropic — explicit cache_control wrap of the system prompt:\nawait c.text.system(longSysPrompt).caching().prompt(\"...\");\n\n// OpenAI — automatic server-side caching (caching() is a hint; reads\n// surface in resp.usage.cacheRead regardless):\nawait c.text.system(longSysPrompt).caching().prompt(\"...\");\n\n// Google — pre-flight POST creates a cachedContents resource, then the\n// main call references it. Google requires ~1k+ tokens of system prompt:\nawait c.text.system(bigSysPrompt).caching().prompt(\"...\");\n```\n\nThe mode is provider-specific and inferred from the provider config. The default TTL comes from `src/providers/caching.ts` (Google: 3600s).\n\n### Model catalogue\n\n`c.models` and `c.providers` cover model discovery in three modes. Runnable counterpart at [`examples/catalogue.ts`](./examples/catalogue.ts).\n\n```ts\nimport { Capabilities } from \"@aktagon/llmkit-ts\";\nimport type { Provider } from \"@aktagon/llmkit-ts\";\n\n// 1. Compiled-in catalogue — synchronous, no HTTP.\nconst all = c.models.list();\nconst info = c.models.get(\"claude-opus-4-7\"); // ModelInfo | undefined\nconst chat = c.models.withCapability(Capabilities.ChatCompletion).list();\n\n// 2. Providers namespace.\nc.providers.list(); // configured (credentials + /v1/models endpoint)\nproviders.list(); // every provider the SDK ships with (static, keyless)\n\n// 3. Live + scoped HTTP.\nconst live = await c.models.live(); // LiveResult — fan-out\nconst p: Provider = { name: \"anthropic\", apiKey: \"sk-...\" };\nconst scoped = await c.models.provider(p).list(); // single-provider list\nconst raw = await c.models.provider(p).raw().list(); // ModelInfo.raw populated\n```\n\n`live()` calls every configured provider's `/v1/models` in parallel and aggregates results into `LiveResult.models` + a per-provider `LiveResult.errors` map (partial success is the normal case). `provider(p).raw().list()` opts into populating `ModelInfo.raw` with the provider-native record — useful when you need fields the universal `ModelInfo` does not carry (Anthropic's capability matrix, Google's `supportedGenerationMethods`, etc.).\n\n## Options\n\nAcross every `*Text` / `*Agent` builder:\n\n| Concept           | Method                   | Notes                                                                                                                                           |\n| ----------------- | ------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------- |\n| System prompt     | `.system(s)`             |                                                                                                                                                 |\n| Model override    | `.model(name)`           |                                                                                                                                                 |\n| Sampling          | `.temperature(t)`        |                                                                                                                                                 |\n| Token cap         | `.maxTokens(n)`          |                                                                                                                                                 |\n| Caching           | `.caching()`             |                                                                                                                                                 |\n| Conversation hist | `.history(...msgs)`      | `*Text` only. `*Agent` accumulates history across `.prompt(...)` calls on the same instance, so an explicit setter would shadow that semantics. |\n| Structured output | `.schema(json)`          | OpenAI strict mode requires `additionalProperties: false` and `required` on object types.                                                       |\n| Middleware hooks  | `.addMiddleware(...fns)` | See below.                                                                                                                                      |\n| Reasoning effort  | `.reasoningEffort(l)`    | OpenAI o-series, Gemini 2.5+                                                                                                                    |\n| Thinking budget   | `.thinkingBudget(n)`     | Anthropic, Gemini                                                                                                                               |\n\nSampling hyperparameters (`.topP`, `.topK`, `.seed`, `.frequencyPenalty`, `.presencePenalty`, `.stopSequences`) are validated per provider; unsupported options throw `ValidationError` rather than silently dropping.\n\nThe Image builder has a narrower set: `.model`, `.aspectRatio`, `.imageSize`, `.includeText`, `.text`, `.image`, `.middleware`. Upload: `.path`, `.bytes`, `.filename`, `.mimeType`, `.middleware`.\n\n## Middleware\n\nRegister pre/post hooks around LLM requests, tool calls, cache creation, uploads, and batch submits. Pre-phase middleware can veto by returning a non-null `Error`; post-phase runs for observation only.\n\n```ts\nimport type { Event, MiddlewareFn } from \"@aktagon/llmkit-ts\";\n\n// Observation: log token usage after every LLM request.\nconst logUsage: MiddlewareFn = (e) => {\n  if (e.op === \"llm_request\" && e.phase === \"post\") {\n    console.log(\n      `${e.provider}/${e.model}: ${e.usage?.input} in, ${e.usage?.output} out, ${e.duration?.toFixed(1)}ms`,\n    );\n  }\n  return null;\n};\n\n// Veto: abort if a daily budget is exceeded.\nconst budgetGate =\n  (limit: number, spent: { value: number }): MiddlewareFn =>\n  (e) => {\n    if (e.op === \"llm_request\" && e.phase === \"pre\" && spent.value >= limit) {\n      return new Error(`daily budget $${limit.toFixed(2)} exceeded`);\n    }\n    return null;\n  };\n\nawait c.text.addMiddleware(budgetGate(5.0, spent), logUsage).prompt(\"...\");\n```\n\nA pre-phase veto throws `MiddlewareVetoError` so it can be discriminated from transport or provider errors. Middlewares fire in registration order; the first non-null pre-phase return aborts.\n\nWired at seven sites: `Text.prompt`, `Text.stream`, `Agent` LLM call, `Agent` tool execution (`op=tool_call`), `Upload.run` (`op=upload`), `Text.batch` (`op=batch_submit`), Google resource caching pre-flight (`op=cache_create`).\n\n## Telemetry\n\nOpt-in OpenTelemetry. Attach a `Telemetry` and every call — success and rejection alike — produces one OTEL GenAI span (operation, provider, model, token usage, and `error.type` on failure) as standards-compliant OTLP/JSON bytes. llmkit builds the span; you decide where the bytes go. Off unless attached.\n\n```ts\nimport { openai, httpExport } from \"@aktagon/llmkit-ts\";\n\n// Batteries: POST every span to an OTLP collector.\nconst client = openai(process.env.OPENAI_API_KEY).addTelemetry({\n  export: httpExport(\"https://collector:4318\"),\n});\n\n// Or bring your own transport — hand the bytes to your OTEL SDK:\nclient.addTelemetry({ export: (b) => batchProcessor.enqueue(b) });\n\nconst resp = await client.text.prompt(\"Hello\");\n```\n\n`httpExport` is a fail-open POST — convenient for low volume; for high volume hand your own callback into your OTEL SDK's batch processor. The same OTLP span shape is emitted byte-for-byte across all six SDKs, so one collector serves a polyglot fleet. A telemetry config with no `export` throws a `ValidationError`.\n\n## Self-hosted endpoints\n\n```ts\nimport { openai } from \"@aktagon/llmkit-ts/builders\";\n\nconst c = openai(\"anything\").baseURL(\"http://localhost:8080/v1\");\n```\n\nWorks for any OpenAI-compatible server (vLLM, LM Studio, Ollama, corporate gateways).\n\n## Custom headers\n\nAttach a custom HTTP header to every request — for example an authenticated gateway that needs its own auth header alongside the provider key. `addHeader` is chainable and calls accumulate.\n\n```ts\nimport { anthropic } from \"@aktagon/llmkit-ts/builders\";\n\nconst c = anthropic(apiKey)\n  .baseURL(\"https://gateway.example.com/anthropic\")\n  .addHeader(\"cf-aig-authorization\", `Bearer ${gatewayToken}`);\n```\n\nThe custom header is sent in addition to the provider's auth header; it cannot override the provider auth header or the required version header.\n\n## Wire-format stability\n\n`*Agent` history persists across process boundaries through two paired\nfunctions:\n\n```typescript\nconst data = bot.save(); // string\n// ...later, fresh process...\nconst bot = c.agent.system(\"...\").tool(t).load(data);\n// throws UnsupportedWireVersionError on mismatch\n```\n\nOr the free-function form for admin tooling:\n\n```typescript\nimport { saveHistory, loadHistory } from \"@aktagon/llmkit\";\n\nconst data = saveHistory(msgs);\nconst msgs = loadHistory(data);\n```\n\nThe output is a JSON document with a `_v` integer envelope plus a\n`messages` array. The version is tracked through\n`WIRE_SCHEMA_VERSION`; the in-memory `Message` schema may evolve\nadditively under one version (new optional fields work on older\nreaders), but a renamed, removed, or retyped field requires a `_v`\nbump and a migrator.\n\n`saveHistory` / `loadHistory` are the ONLY guaranteed-stable\nserialization path. Direct `JSON.stringify` on a `Message` produces\nvalid JSON but lacks the `_v` envelope, and `loadHistory` rejects it\nwith `MissingWireVersionError`. Use the contract path for anything\nthat crosses a process boundary or a release.\n\n## Mirror\n\nThis repo is a read-only mirror of a private monorepo. File issues here; code patches should target the private source via `christian@aktagon.com`.\n\n## License\n\nMIT\n","readmeFilename":"README.md","_rev":"1-a175cafa236207cb7c6a33899a0c8686"}