{"_id":"@astronova-ai/sf-llm-core","name":"@astronova-ai/sf-llm-core","dist-tags":{"latest":"0.1.0"},"versions":{"0.1.0":{"name":"@astronova-ai/sf-llm-core","version":"0.1.0","description":"Isomorphic, fetch-based core that turns a Salesforce org OAuth token into Salesforce LLM Gateway completions (JWT mint + chat/stream).","type":"module","license":"SEE LICENSE IN LICENSE.txt","main":"./dist/index.js","module":"./dist/index.js","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"sideEffects":false,"scripts":{"build":"tsup","typecheck":"tsc --noEmit","test":"vitest run","test:int":"vitest run --config vitest.int.config.ts","test:watch":"vitest","prepack":"npm run build","clean":"rm -rf dist"},"publishConfig":{"access":"public","registry":"https://registry.npmjs.org/"},"dependencies":{"eventsource-parser":"^3.0.0"},"devDependencies":{"@salesforce/llm-gateway-sdk":"^0.14.0","typescript":"^5.7.2","tsup":"^8.3.5","vitest":"^2.1.8"},"keywords":["salesforce","llm","einstein","gateway","claude","gpt","ai"],"_id":"@astronova-ai/sf-llm-core@0.1.0","gitHead":"aad46f58b12f80ebba70012c1e2253e9f19a1103","_nodeVersion":"22.23.1","_npmVersion":"10.9.8","dist":{"integrity":"sha512-zqttUcqGSmMWlLvfHVBeWeToH3tF0ibZUTaGD0bEFXswLh6YkJKCDetsfushyovw3DpGPELl5p28GIhGsQpSxQ==","shasum":"8efabd2d7bab5052700ee03b011598ddad367f64","tarball":"https://registry.npmjs.org/@astronova-ai/sf-llm-core/-/sf-llm-core-0.1.0.tgz","fileCount":11,"unpackedSize":153686,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIBqBWJUrSPZ5ZY2gaaHPPLzMYHT8zH8x/hUebiUCWtK9AiEA5TKhFMlcFL5XvsYMsHNr2JIBPcet9alutGmw6ZKDlNw="}]},"_npmUser":{"name":"dfleminks","email":"damien@fleminks.com"},"directories":{},"maintainers":[{"name":"dfleminks","email":"damien@fleminks.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/sf-llm-core_0.1.0_1784200369802_0.9054146858278784"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-16T11:12:49.679Z","0.1.0":"2026-07-16T11:12:49.936Z","modified":"2026-07-16T11:12:50.154Z"},"maintainers":[{"name":"dfleminks","email":"damien@fleminks.com"}],"description":"Isomorphic, fetch-based core that turns a Salesforce org OAuth token into Salesforce LLM Gateway completions (JWT mint + chat/stream).","keywords":["salesforce","llm","einstein","gateway","claude","gpt","ai"],"license":"SEE LICENSE IN LICENSE.txt","readme":"# @astronova-ai/sf-llm-core\n\n> [!IMPORTANT]\n> This is an independent, unofficial package. It is not endorsed by, affiliated with, or supported\n> by Salesforce. Salesforce is a trademark of Salesforce, Inc. Use requires an entitled Salesforce\n> org and remains subject to your agreements with Salesforce.\n\nIsomorphic, `fetch`-based core that turns a **Salesforce org OAuth token** into **Salesforce LLM\nGateway** completions. Zero agent assumptions — the spine every Astronova adapter builds on.\n\nIt supports **both** gateway surfaces (the newer **Responses API** and the classic\n**chat/generations**), mirrors the wire contract of `@salesforce/llm-gateway-sdk` and the Responses\nOpenAPI spec verbatim, and uses `fetch` instead of `jsforce`/`undici` so it runs unchanged in Node\nand (behind a proxy) the browser.\n\n## Install\n\n```bash\nnpm i @astronova-ai/sf-llm-core\n```\n\nYou also need org credentials (`{ accessToken, instanceUrl }`). In Node, get them from an\nauthenticated `sf` org with [`@astronova-ai/sf-cli-auth`](../sf-cli-auth); in the browser, from a\ntoken broker. Everything below assumes you have `creds`.\n\n## The auth chain (verified live)\n\n```\norg access token ──▶ POST {instanceUrl}/ide/auth  (X-Feature-Id header) ──▶ { jwt }   (tnk + exp)\n                 ──▶ POST api.salesforce.com/einstein/{...}   (Bearer jwt + 4 SF headers)\n```\n\nTwo credentials, not one: the **org access token** mints a short-lived **gateway JWT**\n(`tnk = core/prod/{orgId}`, ~5-min expiry, **auto-refreshed** by this package). The feature-id is\n**org-gated** — `AgentforceVibes` works on many orgs; the historical default `VibesService` is\ndisabled on some. `createJwt` accepts one id or a fallback list and pins the first that mints.\n\n---\n\n## Recommended entry: the unified client\n\n`createLLMFromCredentials` gives you one client that **prefers the Responses API and falls back to\nchat/generations**, routed per model (GPT → Responses, Claude → chat — because the Responses API is\nGPT-only today). You don't choose a surface; the client does, and tells you which it used.\n\n```ts\nimport { createLLMFromCredentials, ModelName } from \"@astronova-ai/sf-llm-core\";\n\nconst llm = await createLLMFromCredentials(creds, { featureId: \"AgentforceVibes\" });\n\n// non-streaming — same call shape regardless of surface\nconst res = await llm.complete(ModelName.GPT_5, {\n  messages: [{ role: \"user\", content: \"What is a mutual fund? One sentence.\" }],\n  reasoningEffort: \"medium\",   // honored on Responses; ignored on chat\n});\nconsole.log(res.text);      // \"A mutual fund pools money from many investors…\"\nconsole.log(res.surface);   // \"responses\"   (Claude models → \"chat\")\nconsole.log(res.usage);     // { inputTokens, outputTokens, totalTokens, reasoningTokens? }\n\n// streaming — normalized chunks\nfor await (const chunk of llm.stream(ModelName.GPT_5, {\n  messages: [{ role: \"user\", content: \"Count to five.\" }],\n})) {\n  if (chunk.textDelta) process.stdout.write(chunk.textDelta);\n  if (chunk.done) console.log(\"\\n\", chunk.usage, `via ${chunk.surface}`);\n}\n\n// tool calls — normalized across both surfaces\nconst tooled = await llm.complete(ModelName.GPT_5, {\n  messages: [{ role: \"user\", content: \"Weather in Paris? Use get_weather.\" }],\n  tools: [{ name: \"get_weather\", description: \"Get weather for a city\",\n    parameters: { type: \"object\", properties: { city: { type: \"string\" } }, required: [\"city\"] } }],\n});\nfor (const call of tooled.toolCalls) console.log(call.function.name, call.function.arguments);\n// get_weather {\"city\":\"Paris\"}\n```\n\n### Surface routing & fallback\n\n| Model family | Default surface | Why |\n|---|---|---|\n| GPT-5 / 5.4 / 5.5 | **Responses API** | reasoning, server tools, richer streaming |\n| Claude (all) | **chat/generations** | Responses API returns 500 for Claude today |\n\n- The router uses `ModelInfo.supportsResponsesApi` (GPT `true`, Claude `false`). When Salesforce\n  enables Claude on the Responses API, flip that flag and Claude upgrades — **no caller change**.\n- On a Responses-side **5xx/transport error**, the client transparently retries on chat/generations\n  (a 4xx is surfaced, since chat would reject it too). Disable with `fallbackOnError: false`.\n- Force a surface with `prefer`: `\"auto\"` (default), `\"responses\"`, or `\"chat\"`.\n- `llm.surfaceFor(model)` tells you the choice without making a call.\n\n```ts\nconst llm = await createLLMFromCredentials(creds, {\n  featureId: \"AgentforceVibes\",\n  prefer: \"auto\",          // \"responses\" | \"chat\" to force\n  fallbackOnError: true,\n});\n```\n\n`UnifiedRequest`: `{ messages, maxTokens?, temperature?, tools?, reasoningEffort?, signal? }`.\n`UnifiedMessage`: `{ role: \"system\"|\"user\"|\"assistant\"|\"tool\", content, toolCallId?, toolCalls? }`.\n`UnifiedResult`: `{ text, toolCalls, usage?, finishReason?, surface }`.\n`UnifiedChunk`: `{ textDelta?, done, toolCalls?, usage?, finishReason?, surface }`.\n\n---\n\n## Lower-level: the Responses API client\n\nUse this directly when you want Responses-specific features (reasoning summaries, the full typed\nstreaming-event protocol, structured input items, response chaining). Endpoint:\n`…/einstein/platform/v1/responses`.\n\n```ts\nimport {\n  createResponsesFromCredentials, extractText, extractFunctionCalls, extractUsage, ModelName,\n} from \"@astronova-ai/sf-llm-core\";\n\nconst responses = await createResponsesFromCredentials(creds, { featureId: \"AgentforceVibes\" });\n\n// non-streaming → full ResponseResource; helpers pull normalized data out\nconst res = await responses.createResponse({\n  model: ModelName.GPT_5,\n  input: \"Explain ETFs briefly.\",          // string, or structured InputItem[]\n  reasoning: { effort: \"medium\" },\n  max_output_tokens: 2048,\n});\nextractText(res);            // assistant text\nextractFunctionCalls(res);   // [{ call_id, name, arguments }]\nextractUsage(res);           // { input_tokens, output_tokens, total_tokens, output_tokens_details? }\n\n// streaming — just the text deltas\nfor await (const delta of responses.streamResponseText({ model: ModelName.GPT_5, input: \"Count to five.\" })) {\n  process.stdout.write(delta);\n}\n\n// streaming — the full typed event protocol (response.created, .output_text.delta,\n// .function_call_arguments.delta/done, .completed, .failed, error, …)\nconst events = await responses.streamResponse({ model: ModelName.GPT_5, input: \"Hi\" });\nfor await (const ev of events) {\n  if (ev.type === \"response.output_text.delta\") process.stdout.write(ev.delta);\n  else if (ev.type === \"response.completed\") console.log(\"\\n\", ev.response.usage);\n}\n```\n\nStructured input (multi-turn, tool results) uses `InputItem[]`:\n\n```ts\nawait responses.createResponse({\n  model: ModelName.GPT_5,\n  input: [\n    { role: \"system\", content: \"You are concise.\" },\n    { role: \"user\", content: \"Weather in Paris?\" },\n    { type: \"function_call\", call_id: \"c1\", name: \"get_weather\", arguments: '{\"city\":\"Paris\"}' },\n    { type: \"function_call_output\", call_id: \"c1\", output: '{\"tempC\":18}' },\n  ],\n});\n```\n\n> **Notes.** The Responses API is **GPT-only** on the gateway today (Claude → 500). Response\n> *persistence* (`store: true` → `previous_response_id` / `GET /v1/responses/{id}`) depends on an\n> org capability that's often off — keep conversations **stateless** (send the full `input` each\n> turn) unless your org enables storage. Types come from\n> [`planning/llm-gateway-responses.yaml`](../../planning/llm-gateway-responses.yaml).\n\n---\n\n## Lower-level: the chat/generations client\n\nThe classic Salesforce surface — works for **all** models (the only path for Claude today). Mirrors\nthe official `@salesforce/llm-gateway-sdk`.\n\n```ts\nimport { createGatewayFromCredentials, ModelName } from \"@astronova-ai/sf-llm-core\";\n\nconst client = await createGatewayFromCredentials(creds, {\n  featureId: \"AgentforceVibes\",\n  model: ModelName.CLAUDE_SONNET_4_6,\n});\n\n// non-streaming\nconst { data } = await client.chat({\n  messages: [{ role: \"user\", content: \"Say PONG.\" }],\n  generation_settings: { max_tokens: 64, temperature: 0 },   // REQUIRED — omit and the gateway 500s\n});\nconsole.log(data.generatedText, data.usage, data.toolInvocations);\n\n// streaming\nconst { data: stream } = await client.chatStream({\n  messages: [{ role: \"user\", content: \"Count to three.\" }],\n  generation_settings: { max_tokens: 64 },\n});\nfor await (const chunk of stream) process.stdout.write(chunk.generatedText);\n```\n\nLowest level of all: `createJwt(creds, opts)` → an auto-refreshing `DynamicJwt`, then\n`createGatewayClient({ jwt })` / `createResponsesClient({ jwt })`. Useful when you want to share one\nJWT across both clients or inject a custom `fetch`.\n\n---\n\n## Models\n\nAll ids verified callable against a live org on 2026-06-24:\n\n| Enum | Gateway id | Context | Max out | Responses API? | Notes |\n|---|---|---|---|---|---|\n| `GPT_5` | `llmgateway__OpenAIGPT5` | 272k | 128k | ✅ | reasoning model — give it generous `max_tokens` |\n| `GPT_5_4` | `llmgateway__OpenAIGPT54` | 1.05M | 128k | ✅ | **600k+ context verified** |\n| `GPT_5_5` | `llmgateway__OpenAIGPT55` | 1.05M | 128k | ✅ | **beta** — org-gated by `AIModelBetaEnabled` |\n| `CLAUDE_SONNET_4_5` | `…Claude45Sonnet` | 200k | 8k | ❌ chat | |\n| `CLAUDE_SONNET_4_6` | `…Claude46Sonnet` | 200k | 16k | ❌ chat | **default model** |\n| `CLAUDE_OPUS_4_5` | `…Claude45Opus` | 200k | 64k | ❌ chat | |\n| `CLAUDE_OPUS_4_6` | `…Claude46Opus` | 1M | 128k | ❌ chat | |\n| `CLAUDE_OPUS_4_7` | `…Claude47Opus` | 1M | 128k | ❌ chat | `temperature` **deprecated** — stripped automatically |\n\nFor a gateway id not yet in the registry, use `createClaudeModel(id, overrides)`, or pass the raw id\nto `getModel(id)` / the unified client (it infers the response family and surface from the id).\n`listModels()` returns all built-ins; `getDefaultModel()` is Claude Sonnet 4.6.\n\n## Key behaviors\n\n- **`getHeaders()` is the single source of truth** for the auth header set\n  (`Authorization`, `x-sfdc-core-tenant-id`, `x-client-feature-id`, `x-sfdc-app-context`,\n  `x-salesforce-region`). Both clients reuse it. Headers are recomputed per request attempt, so a\n  long `Retry-After` wait never re-uses a stale JWT.\n- **Auto-refreshing JWT** — `DynamicJwt` re-mints ~30s before expiry and de-dupes concurrent\n  refreshes; you never manage token lifetime.\n- **Per-model quirks applied automatically** — e.g. Opus 4.7 `temperature` stripping.\n- **Retry/backoff** on 429/5xx/transport errors, honoring `Retry-After`.\n- **Injectable `fetch`** on every entry point — for proxies, custom transports, or the browser.\n- **`fetch`-only, isomorphic** — no Node-only deps; the same code runs in the browser behind a\n  CORS-bridging proxy.\n\n## Testing\n\n```bash\nnpm test            # unit tests (pure logic, no network) + 6 parity checks vs the official SDK\nnpm run test:int    # opt-in live tests against an org (Node 22; needs env below)\n```\n\n```bash\nASTRONOVA_IT_ACCESS_TOKEN=…  ASTRONOVA_IT_INSTANCE_URL=https://your-org.my.salesforce.com \\\nASTRONOVA_IT_FEATURE_ID=AgentforceVibes  npm run test:int -w @astronova-ai/sf-llm-core\n```\n\nThe parity suite asserts our model registry, env mapping, base URLs, and region headers stay\nbyte-identical to `@salesforce/llm-gateway-sdk` — a tripwire against drift.\n\n## License\n\nAstronova's original code is MIT licensed. Portions that mirror or are derived from\n`@salesforce/llm-gateway-sdk` remain subject to Salesforce's Terms of Use. See the complete\n[`LICENSE.txt`](LICENSE.txt) and [`NOTICE`](NOTICE), both of which are included in the npm package.\n","readmeFilename":"README.md","_rev":"1-865b504071d9a2446080ad4f06433208"}