{"_id":"@afterquery/inference","_rev":"13-3b70f6e4e870cc8f702b3010cbe7db22","name":"@afterquery/inference","dist-tags":{"latest":"0.1.8"},"versions":{"0.1.0":{"name":"@afterquery/inference","version":"0.1.0","_id":"@afterquery/inference@0.1.0","maintainers":[{"name":"boatsidejawn","email":"agustin@afterquery.com"}],"homepage":"https://github.com/AfterQuery/inference-v1#readme","bugs":{"url":"https://github.com/AfterQuery/inference-v1/issues"},"dist":{"shasum":"31d959e5c0c8697b2fe7c306351859995cbac6e6","tarball":"https://registry.npmjs.org/@afterquery/inference/-/inference-0.1.0.tgz","fileCount":10,"integrity":"sha512-hiubm03TFwpBYl0Xo3BGv0CYkdDeEInKu2MKew6eHEMcIAyPp4lw7ABRDfo02tunmz7Uq5ZFtCNH/b3v8/4kTQ==","signatures":[{"sig":"MEQCICwHeEBnR67iNWcFnXWFmnTmmeDSPizizDMn+0/jbWDwAiAG8A4hp+Qr/dGrsc5mfVKXdfW31WTLgQPaaNsiLz3Hmg==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":64865},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"scripts":{"dev":"tsc --watch","build":"tsc","clean":"rm -rf dist","typecheck":"tsc --noEmit"},"_npmUser":{"name":"boatsidejawn","email":"agustin@afterquery.com"},"repository":{"url":"git+https://github.com/AfterQuery/inference-v1.git","type":"git","directory":"packages/inference"},"_npmVersion":"10.9.4","description":"TypeScript SDK for the Inference V1 gateway.","directories":{},"_nodeVersion":"22.21.1","dependencies":{"zod":"4.3.6"},"publishConfig":{"registry":"https://registry.npmjs.org"},"_hasShrinkwrap":false,"devDependencies":{"typescript":"5.9.3","@types/node":"20.19.33"},"_npmOperationalInternal":{"tmp":"tmp/inference_0.1.0_1778205988118_0.8605791422900588","host":"s3://npm-registry-packages-npm-production"}},"0.1.2":{"name":"@afterquery/inference","version":"0.1.2","_id":"@afterquery/inference@0.1.2","maintainers":[{"name":"boatsidejawn","email":"agustin@afterquery.com"}],"homepage":"https://github.com/AfterQuery/inference-v1#readme","bugs":{"url":"https://github.com/AfterQuery/inference-v1/issues"},"dist":{"shasum":"c623d3b28cc499563d6d016617c2188157295799","tarball":"https://registry.npmjs.org/@afterquery/inference/-/inference-0.1.2.tgz","fileCount":10,"integrity":"sha512-zhsFwDh2oir5oXMueUaI+vyfvLiWLNUqKGLRXum4VfQee5Rk1n+IdQHQvfBHc2NpkNiYa2Mgeb9a4Kbh4xM4mw==","signatures":[{"sig":"MEUCIEfuecLugqcPZUKc4NAP4t+ccLdQ9YadTVJSj3Dk+/xcAiEA+xGgmq2WjmEyCkyJUm8aBeu5Hi0jI64EEZrDPw7ASOY=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":77552},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"scripts":{"dev":"tsc --watch","build":"tsc","clean":"rm -rf dist","typecheck":"tsc --noEmit"},"_npmUser":{"name":"boatsidejawn","email":"agustin@afterquery.com"},"repository":{"url":"git+https://github.com/AfterQuery/inference-v1.git","type":"git","directory":"packages/inference"},"_npmVersion":"10.9.4","description":"TypeScript SDK for the Inference V1 gateway.","directories":{},"_nodeVersion":"22.21.1","dependencies":{"zod":"4.3.6"},"publishConfig":{"registry":"https://registry.npmjs.org"},"_hasShrinkwrap":false,"devDependencies":{"typescript":"5.9.3","@types/node":"20.19.33"},"_npmOperationalInternal":{"tmp":"tmp/inference_0.1.2_1778630978677_0.2961504994110358","host":"s3://npm-registry-packages-npm-production"}},"0.1.3":{"name":"@afterquery/inference","version":"0.1.3","_id":"@afterquery/inference@0.1.3","maintainers":[{"name":"boatsidejawn","email":"agustin@afterquery.com"}],"homepage":"https://github.com/AfterQuery/inference-v1#readme","bugs":{"url":"https://github.com/AfterQuery/inference-v1/issues"},"dist":{"shasum":"e81bfa305ea1e8473f63d90e4432f27ee70cc760","tarball":"https://registry.npmjs.org/@afterquery/inference/-/inference-0.1.3.tgz","fileCount":10,"integrity":"sha512-grNHI31E4VQulSANj9BNSX1sGrcDZ1byMOxNKAVIHQfWguVgLG766Z5rOIVZFPlIhvL6iDs6OQAfPFpcomhmPw==","signatures":[{"sig":"MEUCIQCbyHwtWombDpw4pknkEEBao9jBqFRbDizO2Fqf7b7rRwIgIKVDotdCHO5gKWtMQnpTC3ek3AOROUq6Gimh6GIPEAM=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":93145},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"scripts":{"dev":"tsc --watch","build":"tsc","clean":"rm -rf dist","typecheck":"tsc --noEmit"},"_npmUser":{"name":"boatsidejawn","email":"agustin@afterquery.com"},"repository":{"url":"git+https://github.com/AfterQuery/inference-v1.git","type":"git","directory":"packages/inference"},"_npmVersion":"10.9.4","description":"TypeScript SDK for the Inference V1 gateway.","directories":{},"_nodeVersion":"22.21.1","dependencies":{"zod":"4.3.6"},"publishConfig":{"registry":"https://registry.npmjs.org"},"_hasShrinkwrap":false,"devDependencies":{"typescript":"5.9.3","@types/node":"20.19.33"},"_npmOperationalInternal":{"tmp":"tmp/inference_0.1.3_1779139692933_0.9001481217967346","host":"s3://npm-registry-packages-npm-production"}},"0.1.4":{"name":"@afterquery/inference","version":"0.1.4","_id":"@afterquery/inference@0.1.4","maintainers":[{"name":"boatsidejawn","email":"agustin@afterquery.com"}],"homepage":"https://github.com/AfterQuery/inference-v1#readme","bugs":{"url":"https://github.com/AfterQuery/inference-v1/issues"},"dist":{"shasum":"0ef7735ec02486863ef92b8f966046bd0d30ac9b","tarball":"https://registry.npmjs.org/@afterquery/inference/-/inference-0.1.4.tgz","fileCount":10,"integrity":"sha512-IyzbBUjMU+brWLKYI/lzKLyrdYD3IjkcuaI1YKAM4lEzdXw0lhN2dmZP01OiUY97MZeCInmcqIuoNfh7obraKA==","signatures":[{"sig":"MEQCIDkszkD9SwAZrVCekqOKIx1zsA8sLxMR042a4E9xMeXRAiAQQTroUTdLKHk3ytkMUoCzufJ9AZvM99CtIrbQHAeM5w==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":98019},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"scripts":{"dev":"tsc --watch","build":"tsc","clean":"rm -rf dist","typecheck":"tsc --noEmit"},"_npmUser":{"name":"boatsidejawn","email":"agustin@afterquery.com"},"repository":{"url":"git+https://github.com/AfterQuery/inference-v1.git","type":"git","directory":"packages/inference"},"_npmVersion":"10.9.4","description":"TypeScript SDK for the Inference V1 gateway.","directories":{},"_nodeVersion":"22.21.1","dependencies":{"zod":"4.3.6"},"publishConfig":{"registry":"https://registry.npmjs.org"},"_hasShrinkwrap":false,"devDependencies":{"typescript":"5.9.3","@types/node":"20.19.33"},"_npmOperationalInternal":{"tmp":"tmp/inference_0.1.4_1779160012739_0.07447818521570282","host":"s3://npm-registry-packages-npm-production"}},"0.1.6":{"name":"@afterquery/inference","version":"0.1.6","_id":"@afterquery/inference@0.1.6","maintainers":[{"name":"boatsidejawn","email":"agustin@afterquery.com"}],"homepage":"https://github.com/AfterQuery/inference-v1#readme","bugs":{"url":"https://github.com/AfterQuery/inference-v1/issues"},"dist":{"shasum":"7e7329d0a9c2436e5eb0f74b7c07bb342d0888a4","tarball":"https://registry.npmjs.org/@afterquery/inference/-/inference-0.1.6.tgz","fileCount":14,"integrity":"sha512-NDYokx/jGiABTOo1E4lNvDPcFJN2y9n2gknSrSnhzi7TRwH/9R/IoITfSIvsWh1tIYIlQK1+Y/Vok29EBj+6IQ==","signatures":[{"sig":"MEQCIBdBNj3tmohZ6zh3xtLdaHLOTI4742TxhkAFuR4eaoedAiAqtuJzAgGffiG8c4PC/ymLP0fTZI3IYWHCZDxkglslYw==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":115685},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"scripts":{"dev":"tsc --watch","build":"tsc","clean":"rm -rf dist","typecheck":"tsc --noEmit"},"_npmUser":{"name":"boatsidejawn","email":"agustin@afterquery.com"},"repository":{"url":"git+https://github.com/AfterQuery/inference-v1.git","type":"git","directory":"packages/inference"},"_npmVersion":"10.9.4","description":"TypeScript SDK for the Inference V1 gateway.","directories":{},"_nodeVersion":"22.21.1","dependencies":{"zod":"4.3.6"},"publishConfig":{"registry":"https://registry.npmjs.org"},"_hasShrinkwrap":false,"devDependencies":{"typescript":"5.9.3","@types/node":"20.19.33"},"_npmOperationalInternal":{"tmp":"tmp/inference_0.1.6_1779920735698_0.33464597752250747","host":"s3://npm-registry-packages-npm-production"}},"0.1.7":{"name":"@afterquery/inference","version":"0.1.7","_id":"@afterquery/inference@0.1.7","maintainers":[{"name":"boatsidejawn","email":"agustin@afterquery.com"}],"homepage":"https://github.com/AfterQuery/inference-v1#readme","bugs":{"url":"https://github.com/AfterQuery/inference-v1/issues"},"dist":{"shasum":"a6c444fe6ad14bbd5a6f823b3aa605061678a202","tarball":"https://registry.npmjs.org/@afterquery/inference/-/inference-0.1.7.tgz","fileCount":14,"integrity":"sha512-iNe/jQu7P2HcLL2YA9xWbmiTl41f3GZNSubzBup16N70kk8ZSnhJzGpo8UdyF+nQ3r49xtDWzCPszxg+N2MnVg==","signatures":[{"sig":"MEUCIQDZPkcu9cErjt9EnszDyzHkZQFBbOwcs1GZUFJwH+4rngIgXnAg1mnvDzig13ylIj7bfHFhTOyeUug1HMtrMzalW+s=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":124494},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"scripts":{"dev":"tsc --watch","build":"tsc","clean":"rm -rf dist","typecheck":"tsc --noEmit"},"_npmUser":{"name":"boatsidejawn","email":"agustin@afterquery.com"},"repository":{"url":"git+https://github.com/AfterQuery/inference-v1.git","type":"git","directory":"packages/inference"},"_npmVersion":"10.9.4","description":"TypeScript SDK for the Inference V1 gateway.","directories":{},"_nodeVersion":"22.21.1","dependencies":{"zod":"4.3.6"},"publishConfig":{"registry":"https://registry.npmjs.org"},"_hasShrinkwrap":false,"devDependencies":{"typescript":"5.9.3","@types/node":"20.19.33"},"_npmOperationalInternal":{"tmp":"tmp/inference_0.1.7_1780380295474_0.9636982007485368","host":"s3://npm-registry-packages-npm-production"}},"0.1.8":{"name":"@afterquery/inference","version":"0.1.8","_id":"@afterquery/inference@0.1.8","maintainers":[{"name":"boatsidejawn","email":"agustin@afterquery.com"}],"homepage":"https://github.com/AfterQuery/inference-v1#readme","bugs":{"url":"https://github.com/AfterQuery/inference-v1/issues"},"dist":{"shasum":"3d02c2bed5c06592c90eac7c1892591b0b3e8f07","tarball":"https://registry.npmjs.org/@afterquery/inference/-/inference-0.1.8.tgz","fileCount":14,"integrity":"sha512-1LnHdTvM2uLTWGN2AMWi1PGOq6lfyQUYtwLGR5UGl7NirIovS/oTz54k31hdqedXYIm2a5aaC3vHmST3LVfyWg==","signatures":[{"sig":"MEUCIQD+pW8BiGezf4C18Rh7c6o006B31nw6Oa2KLnEnPaN/kwIgZa1fL4EbxK65Q4t+NGxCcd9U2hKo4FjMSQGubgKrhmM=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":152755},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"scripts":{"dev":"tsc --watch","build":"tsc","clean":"rm -rf dist","typecheck":"tsc --noEmit"},"_npmUser":{"name":"boatsidejawn","email":"agustin@afterquery.com"},"repository":{"url":"git+https://github.com/AfterQuery/inference-v1.git","type":"git","directory":"packages/inference"},"_npmVersion":"10.9.4","description":"TypeScript SDK for the Inference V1 gateway.","directories":{},"_nodeVersion":"22.21.1","dependencies":{"zod":"4.3.6"},"publishConfig":{"registry":"https://registry.npmjs.org"},"_hasShrinkwrap":false,"devDependencies":{"typescript":"5.9.3","@types/node":"20.19.33"},"_npmOperationalInternal":{"tmp":"tmp/inference_0.1.8_1781227982791_0.7520748713409724","host":"s3://npm-registry-packages-npm-production"}}},"time":{"created":"2026-05-08T02:06:27.992Z","modified":"2026-09-15T16:17:23.430Z","0.1.0":"2026-05-08T02:06:28.290Z","0.1.2":"2026-05-13T00:09:38.849Z","0.1.3":"2026-05-18T21:28:13.096Z","0.1.4":"2026-05-19T03:06:52.890Z","0.1.6":"2026-05-27T22:25:35.876Z","0.1.7":"2026-06-02T06:04:55.605Z","0.1.8":"2026-06-12T01:33:02.926Z"},"bugs":{"url":"https://github.com/AfterQuery/inference-v1/issues"},"homepage":"https://github.com/AfterQuery/inference-v1#readme","repository":{"url":"git+https://github.com/AfterQuery/inference-v1.git","type":"git","directory":"packages/inference"},"description":"TypeScript SDK for the Inference V1 gateway.","maintainers":[{"email":"saksham@afterquery.com","name":"sakshamaq"},{"email":"bot+npm@afterquery.com","name":"aqbot"},{"email":"tyler@afterquery.com","name":"tylerbakke"},{"email":"arya@afterquery.com","name":"aryaf"},{"email":"igor@afterquery.com","name":"igorafterquery"}],"readme":"# @afterquery/inference\n\nTypeScript SDK for the Inference V1 gateway.\n\nThe package intentionally talks to your Inference backend instead of provider APIs directly. Provider secrets, OpenRouter-first routing, direct-provider fallback, retries, logging, and policy live in the backend.\n\n## Two integration modes\n\n### 1) Gateway-native SDK\n\nUse `@afterquery/inference` when you can replace direct provider calls in app code and want the gateway's normalized response helpers (`content`, `route`, `cost`, etc.).\n\n### 2) OpenAI-compatible ingress\n\nUse a DB-managed runtime key provisioned with:\n\n- `clientShape = openai_compat`\n- a bound `projectSlug`\n- `defaultFeature`\n\nThis is the path for eval harnesses, LiteLLM, or existing OpenAI-client code that should only swap `base_url` / `api_key`.\n\n## Install\n\n```bash\nnpm install @afterquery/inference\n```\n\n## Gateway-native SDK: create client\n\n```ts\nimport { Inference } from '@afterquery/inference';\n\nconst inference = new Inference({\n  baseURL: process.env.INFERENCE_BASE_URL!,\n  apiKey: process.env.INFERENCE_API_KEY!,\n});\n```\n\n`createInferenceClient()` is still available:\n\n```ts\nimport { createInferenceClient } from '@afterquery/inference';\n\nconst inference = createInferenceClient({\n  baseUrl: 'http://localhost:8787',\n  apiKey: process.env.INFERENCE_API_KEY,\n});\n```\n\n## Responses\n\n```ts\nconst response = await inference.responses.create({\n  model: 'gpt-4.1-mini',\n  input: 'Write a concise product summary.',\n  metadata: {\n    project: 'demo',\n    feature: 'product-copy',\n  },\n});\n\nconsole.log(response.content);\nconsole.log(response.route?.fallbackUsed);\n```\n\n### Background responses\n\nUse background responses for long-running jobs that should return a pollable id quickly.\n\n```ts\nlet response = await inference.responses.create({\n  model: 'perplexity/sonar-deep-research',\n  input: 'Write a concise research brief about async deep research APIs.',\n  background: true,\n  metadata: {\n    project: 'demo',\n    feature: 'deep-research',\n  },\n});\n\nwhile (response.status === 'queued' || response.status === 'in_progress') {\n  await new Promise((resolve) => setTimeout(resolve, 10_000));\n  response = await inference.responses.retrieve(response.id);\n}\n\nconsole.log(response.output_text);\nconsole.log(response.inference?.route?.provider);\n```\n\nForeground `responses.create` calls keep returning the normalized gateway response shape. With `background: true`, the SDK returns the OpenAI-style background response object used by `GET /v1/responses/{id}`.\n\n## Chat completions\n\n```ts\nconst result = await inference.chat.completions.create({\n  model: 'claude-sonnet-4-6',\n  messages: [{ role: 'user', content: 'Say hello in one sentence.' }],\n  metadata: {\n    project: 'demo',\n    feature: 'inline-chat',\n  },\n});\n```\n\nCompatibility alias:\n\n```ts\nawait inference.chat({\n  model: 'gpt-4.1-mini',\n  messages: [{ role: 'user', content: 'Hello' }],\n  metadata: { project: 'demo', feature: 'chat' },\n});\n```\n\n## Provider override\n\nOmit `provider` for the gateway default route plan. Most models route OpenRouter-first; configured Claude Bedrock models route AWS Bedrock-first with OpenRouter fallback. Set `provider` when a request should force one gateway provider and skip gateway fallback:\n\n```ts\nawait inference.chat.completions.create({\n  provider: 'anthropic',\n  model: 'claude-sonnet-4-6',\n  messages: [{ role: 'user', content: 'Use Claude directly.' }],\n  metadata: { project: 'demo', feature: 'direct-anthropic' },\n});\n```\n\nValid chat/responses providers are `openrouter`, `openai`, `anthropic`, and `google`.\n\n### AWS Bedrock Anthropic routing\n\nWhen the gateway has `AWS_BEARER_TOKEN_BEDROCK` configured (`AWS_BEARER_TOKEN` is also accepted), select Claude models route through org-owned AWS Bedrock first and fall back to regular OpenRouter if Bedrock errors. Callers do not pass a BYOK flag; they request the normal Anthropic/OpenRouter model id:\n\n```ts\nawait inference.chat.completions.create({\n  model: 'anthropic/claude-opus-4.8',\n  messages: [{ role: 'user', content: 'Use the gateway routing policy.' }],\n  metadata: { project: 'demo', feature: 'bedrock-anthropic' },\n});\n```\n\nThe current Bedrock-first set is Opus 4.8, Opus 4.7, and Sonnet 4.6. Request logs show Bedrock traffic as `providerConnector=aws_bedrock` / `provider_connector='aws_bedrock'`, `provider=anthropic`, and a Bedrock inference-profile model id such as `global.anthropic.claude-opus-4-8`; fallback attempts remain normal OpenRouter attempts.\n\n## Tool calling\n\nChat completions support OpenAI-style function tools through both the native SDK and OpenAI-compatible ingress.\n\n```ts\nconst result = await inference.chat.completions.create({\n  model: 'gpt-4.1-mini',\n  messages: [{ role: 'user', content: 'What is the weather in Oakland?' }],\n  tools: [\n    {\n      type: 'function',\n      function: {\n        name: 'get_weather',\n        description: 'Lookup the current weather by city.',\n        parameters: {\n          type: 'object',\n          properties: { city: { type: 'string' } },\n          required: ['city'],\n        },\n      },\n    },\n  ],\n  toolChoice: 'auto',\n  metadata: { project: 'demo', feature: 'assistant-tools' },\n});\n\nfor (const toolCall of result.toolCalls ?? []) {\n  console.log(toolCall.id, toolCall.function.name, toolCall.function.arguments);\n}\n```\n\n## Vision input\n\n```ts\nawait inference.responses.create({\n  model: 'gpt-4o-mini',\n  input: [\n    {\n      role: 'user',\n      content: [\n        { type: 'input_text', text: 'What is in this image?' },\n        { type: 'input_image', imageUrl: 'https://example.com/image.png' },\n      ],\n    },\n  ],\n  metadata: {\n    project: 'demo',\n    feature: 'vision',\n  },\n});\n```\n\n## Prompt caching\n\nFor Anthropic direct or OpenRouter Claude routes, content blocks can carry Anthropic/OpenRouter prompt-cache controls. The SDK accepts both camel-case `cacheControl` and provider-native `cache_control`. Responses include upstream cache accounting under `usage.cacheStatus`, `usage.cacheReadTokens`, and `usage.cacheWriteTokens` when the provider reports it. Direct-provider cost estimates are cache-adjusted only for provider/model policies we have explicitly verified; otherwise cache-affected cost is marked unavailable instead of silently overcharging or undercharging.\n\n```ts\nawait inference.chat.completions.create({\n  provider: 'anthropic',\n  model: 'claude-sonnet-4-6',\n  messages: [\n    {\n      role: 'system',\n      content: [\n        {\n          type: 'text',\n          text: largeStableContext,\n          cacheControl: { type: 'ephemeral', ttl: '1h' },\n        },\n      ],\n    },\n    { role: 'user', content: 'Use the cached context to answer.' },\n  ],\n  metadata: { project: 'demo', feature: 'prompt-cache' },\n});\n```\n\n## Embeddings\n\n```ts\nawait inference.embeddings.create({\n  model: 'text-embedding-3-small',\n  input: 'Document text',\n  metadata: {\n    project: 'demo',\n    feature: 'search-index',\n  },\n});\n```\n\nCompatibility alias:\n\n```ts\nawait inference.embed({\n  model: 'text-embedding-3-small',\n  input: 'Document text',\n  metadata: { project: 'demo', feature: 'search-index' },\n});\n```\n\n## Image generation\n\n```ts\nawait inference.images.generate({\n  model: 'gpt-image-1',\n  prompt: 'A clean diagram of an AI gateway.',\n  size: '1024x1024',\n  metadata: {\n    project: 'demo',\n    feature: 'image-generation',\n  },\n});\n```\n\n## OpenAI-compatible ingress examples\n\n### Python OpenAI client\n\n```python\nfrom openai import OpenAI\n\nclient = OpenAI(\n    base_url=\"https://your-gateway.example.com/v1\",\n    api_key=\"inf_live_...\",\n)\n\nresponse = client.chat.completions.create(\n    model=\"gpt-4.1-mini\",\n    messages=[{\"role\": \"user\", \"content\": \"Reply with a one-line summary.\"}],\n)\n\nprint(response.choices[0].message.content)\n```\n\n### TypeScript OpenAI client\n\n```ts\nimport OpenAI from 'openai';\n\nconst client = new OpenAI({\n  baseURL: 'https://your-gateway.example.com/v1',\n  apiKey: process.env.INFERENCE_API_KEY!,\n});\n\nconst response = await client.chat.completions.create({\n  model: 'gpt-4.1-mini',\n  messages: [{ role: 'user', content: 'Reply with a one-line summary.' }],\n});\n\nconsole.log(response.choices[0]?.message?.content);\n```\n\nThe same OpenAI client can call background Responses through the gateway:\n\n```ts\nlet response = await client.responses.create({\n  model: 'perplexity/sonar-deep-research',\n  input: 'Write a concise research brief about async deep research APIs.',\n  background: true,\n  metadata: {\n    project: 'demo',\n    feature: 'deep-research',\n  },\n});\n\nresponse = await client.responses.retrieve(response.id);\nconsole.log(response.output_text);\n```\n\n### LiteLLM\n\n```python\nfrom litellm import completion\n\nresponse = completion(\n    model=\"gpt-4.1-mini\",\n    custom_llm_provider=\"openai\",\n    api_base=\"https://your-gateway.example.com/v1\",\n    api_key=\"inf_live_...\",\n    messages=[{\"role\": \"user\", \"content\": \"Grade this answer in one sentence.\"}],\n)\n\nprint(response.choices[0].message.content)\n```\n\nFor chat completions with compat keys, the backend derives `project` and `feature` from the provisioned key defaults, so the harness does not need to know the gateway-native metadata contract. Background Responses include those metadata fields explicitly so the job can be authorized and logged. OpenAI-compatible `tools`, `tool_choice`, `tool_calls`, and `role: \"tool\"` messages are forwarded for OpenAI/OpenRouter-backed chat routes.\n\nOpenAI-compatible callers can force a gateway provider with the nonstandard top-level `provider` field. Depending on the client library, use an escape hatch such as `extra_body` or a typed cast. Prompt-caching content blocks can carry provider-native `cache_control` when the client allows extra content-block fields.\n\n```python\nresponse = client.chat.completions.create(\n    model=\"anthropic/claude-sonnet-4.6\",\n    messages=[\n        {\n            \"role\": \"user\",\n            \"content\": [\n                {\n                    \"type\": \"text\",\n                    \"text\": large_stable_context,\n                    \"cache_control\": {\"type\": \"ephemeral\", \"ttl\": \"1h\"},\n                },\n                {\"type\": \"text\", \"text\": \"Answer the question.\"},\n            ],\n        }\n    ],\n    extra_body={\"provider\": \"openrouter\"},\n)\n```\n\n## Gateway-level retries\n\nThe SDK retries only network/gateway failures. Provider retries and OpenRouter -> direct provider fallback happen inside the backend.\n\n```ts\nconst inference = new Inference({\n  baseURL: process.env.INFERENCE_BASE_URL!,\n  apiKey: process.env.INFERENCE_API_KEY!,\n  timeoutMs: 60_000,\n  retry: {\n    maxRetries: 2,\n    initialDelayMs: 250,\n    maxDelayMs: 3_000,\n  },\n});\n```\n\nPer request:\n\n```ts\nawait inference.responses.create(payload, {\n  timeoutMs: 30_000,\n  idempotencyKey: 'request-123',\n});\n```\n\n## Route metadata\n\nResponses may include route information from the backend:\n\n```ts\nresponse.route?.attempts\nresponse.route?.fallbackUsed\nresponse.route?.provider\nresponse.route?.model\n```\n\nThis is how callers can inspect whether OpenRouter was used or whether the backend fell back to OpenAI, Anthropic, or Google directly.\n","readmeFilename":"README.md"}