{"_id":"@eleven-am/voice-agent","name":"@eleven-am/voice-agent","dist-tags":{"latest":"0.0.1"},"versions":{"0.0.1":{"name":"@eleven-am/voice-agent","version":"0.0.1","description":"TypeScript SDK for building voice agents on the Voice Gateway","main":"dist/index.js","module":"dist/index.mjs","types":"dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.mjs","require":"./dist/index.js"}},"scripts":{"build":"tsup","dev":"tsup --watch","test":"vitest","lint":"eslint src/","typecheck":"tsc --noEmit","prepublishOnly":"npm run build"},"keywords":["voice","agent","websocket","streaming","llm","assistant"],"author":{"name":"eleven-am"},"license":"MIT","devDependencies":{"@types/node":"^22.0.0","@types/ws":"^8.18.1","eslint":"^9.0.0","tsup":"^8.0.0","typescript":"^5.7.0","vitest":"^3.0.0"},"dependencies":{"ws":"^8.0.0"},"peerDependencies":{"ws":"^8.0.0"},"engines":{"node":">=18"},"publishConfig":{"access":"public"},"_id":"@eleven-am/voice-agent@0.0.1","gitHead":"232bd7cd5b64e82dc5fdb13dd3c61b712f040ad6","_nodeVersion":"22.17.0","_npmVersion":"11.4.2","dist":{"integrity":"sha512-eXOLlpKHbf468YALbgX5Z8ls2vtHvYHc8vR2ar91ACUPBClQ0dFNwUz8ycKcWpW7Oj8fQDo4UZbh5L7lk3uRjw==","shasum":"6b8b5b02d79dfa328c90f672f9afda1dfadb87b7","tarball":"https://registry.npmjs.org/@eleven-am/voice-agent/-/voice-agent-0.0.1.tgz","fileCount":8,"unpackedSize":203686,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIEYjgPK3peHRk7QfcKlJrBKOJoOIAtZSQVhimq5/pg5uAiAslWIg2oBddXSrG2WUj6U7437sNwwB8tJ2IJQOePLRNQ=="}]},"_npmUser":{"name":"eleven-am","email":"maixperiyon@gmail.com"},"directories":{},"maintainers":[{"name":"eleven-am","email":"maixperiyon@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/voice-agent_0.0.1_1769466276421_0.6480343556384522"},"_hasShrinkwrap":false}},"time":{"created":"2026-01-26T22:24:36.294Z","0.0.1":"2026-01-26T22:24:36.602Z","modified":"2026-01-26T22:24:36.838Z"},"maintainers":[{"name":"eleven-am","email":"maixperiyon@gmail.com"}],"description":"TypeScript SDK for building voice agents on the Voice Gateway","keywords":["voice","agent","websocket","streaming","llm","assistant"],"author":{"name":"eleven-am"},"license":"MIT","readme":"# Voice Agent SDK (TypeScript)\n\nTypeScript SDK for building voice agents on the Voice Gateway.\n\n## Installation\n\n```bash\nnpm install @eleven-am/voice-agent\n```\n\n## Quick Start\n\n```typescript\nimport { VoiceAgent, ConnectionModes } from '@eleven-am/voice-agent';\n\nconst agent = new VoiceAgent({\n  apiKey: 'sk-voice-xxx',\n  gatewayUrl: 'wss://gateway.example.com',\n  mode: ConnectionModes.WebSocket,\n});\n\nagent\n  .onUtterance(async (ctx) => {\n    console.log(`User said: ${ctx.text}`);\n\n    // Stream response\n    ctx.sendDelta('Hello ');\n    ctx.sendDelta('World!');\n    ctx.done();\n  })\n  .onInterrupt((sessionId, reason) => {\n    console.log(`Interrupted: ${reason}`);\n  })\n  .onError((error) => {\n    console.error('Error:', error);\n  });\n\nawait agent.connect();\n```\n\n## Streaming with LLM\n\n```typescript\nimport { VoiceAgent } from '@eleven-am/voice-agent';\nimport OpenAI from 'openai';\n\nconst openai = new OpenAI();\n\nconst agent = new VoiceAgent({\n  apiKey: process.env.VOICE_API_KEY!,\n  gatewayUrl: process.env.GATEWAY_URL!,\n});\n\nagent.onUtterance(async (ctx) => {\n  const stream = await openai.chat.completions.create({\n    model: 'gpt-4',\n    messages: [{ role: 'user', content: ctx.text }],\n    stream: true,\n  });\n\n  for await (const chunk of stream) {\n    if (ctx.abortSignal.aborted) break;\n\n    const delta = chunk.choices[0]?.delta?.content;\n    if (delta) {\n      ctx.sendDelta(delta);\n    }\n  }\n\n  ctx.done();\n});\n\nawait agent.connect();\n```\n\n## Using Vision (Video Frames)\n\nAgents with `vision` scope can request video frames from the user's session:\n\n```typescript\nagent.onUtterance(async (ctx) => {\n  // Check if vision context is available\n  if (ctx.vision?.available) {\n    console.log('Auto-analyzed:', ctx.vision.description);\n  }\n\n  // Request raw frames for custom analysis\n  const frames = await ctx.requestFrames({\n    limit: 5,\n    rawBase64: true,\n  });\n\n  if (frames.frames) {\n    for (const frame of frames.frames) {\n      // frame.base64 contains the image data\n      // frame.timestamp is when it was captured\n    }\n  }\n\n  // Or get pre-analyzed descriptions\n  const analyzed = await ctx.requestFrames({ limit: 3 });\n  if (analyzed.descriptions) {\n    console.log('Frame descriptions:', analyzed.descriptions);\n  }\n\n  ctx.done('I can see what you\\'re showing me!');\n});\n```\n\n## Using Memory\n\nAgents with `memory` scope can query the user's stored facts:\n\n```typescript\nagent.onUtterance(async (ctx) => {\n  // Query relevant memories\n  const memories = await ctx.queryMemory({\n    query: ctx.text,\n    topK: 5,\n    threshold: 0.7,\n    types: ['preference', 'fact'],\n  });\n\n  if (memories.facts && memories.facts.length > 0) {\n    const context = memories.facts\n      .map(f => f.content)\n      .join('\\n');\n\n    // Use memories as context for LLM\n    const response = await generateWithContext(ctx.text, context);\n    ctx.done(response);\n  } else {\n    ctx.done('I don\\'t have any relevant memories about that.');\n  }\n});\n```\n\n## Handling Interrupts\n\nWhen the user starts speaking, the gateway sends an interrupt. Use the abort signal to stop processing:\n\n```typescript\nagent.onUtterance(async (ctx) => {\n  for await (const chunk of streamResponse(ctx.text)) {\n    // Check before each operation\n    if (ctx.abortSignal.aborted) {\n      console.log('User interrupted, stopping');\n      return;\n    }\n    ctx.sendDelta(chunk);\n  }\n  ctx.done();\n});\n\nagent.onInterrupt((sessionId, reason) => {\n  // reason: \"new_user_speech\" | \"lost_arbitration\" | \"supersede\"\n  console.log(`Session ${sessionId} interrupted: ${reason}`);\n});\n```\n\n## Configuration\n\n```typescript\ninterface VoiceAgentConfig {\n  apiKey: string;              // Your API key (sk-voice-xxx)\n  gatewayUrl: string;          // Gateway WebSocket/HTTP URL\n  mode?: ConnectionMode;       // 'websocket' (default) or 'sse'\n  reconnect?: boolean;         // Auto-reconnect on disconnect (default: true)\n  reconnectInterval?: number;  // Base reconnect delay in ms (default: 1000)\n  maxReconnectAttempts?: number; // Max reconnect attempts (default: unlimited)\n  logger?: Logger;             // Custom logger (default: console)\n}\n```\n\n## Context API\n\nThe `UtteranceContext` provides:\n\n| Property | Type | Description |\n|----------|------|-------------|\n| `text` | `string` | The user's utterance text |\n| `isFinal` | `boolean` | Whether this is a final transcript |\n| `user` | `UserInfo \\| undefined` | User info (if `profile`/`email`/`location` scope) |\n| `vision` | `VisionContext \\| undefined` | Vision context (if `vision` scope) |\n| `sessionId` | `string` | Current session ID |\n| `requestId` | `string` | Current request ID |\n| `userId` | `string \\| undefined` | User ID |\n| `timestamp` | `Date` | When the utterance was received |\n| `abortSignal` | `AbortSignal` | Signals when interrupted |\n\n| Method | Description |\n|--------|-------------|\n| `sendDelta(delta)` | Stream a text chunk to the user |\n| `done(finalText?)` | Complete the response |\n| `requestFrames(options?)` | Request video frames (async) |\n| `queryMemory(options)` | Query user memories (async) |\n\n## Connection Modes\n\n### WebSocket (recommended)\n\nFull-duplex communication, lower latency:\n\n```typescript\nconst agent = new VoiceAgent({\n  mode: ConnectionModes.WebSocket,\n  // ...\n});\n```\n\n### Server-Sent Events (SSE)\n\nOne-way server push with HTTP POST for sending. Works in browser environments:\n\n```typescript\nconst agent = new VoiceAgent({\n  mode: ConnectionModes.SSE,\n  // ...\n});\n```\n\n## License\n\nMIT\n","readmeFilename":"README.md","_rev":"1-e083a970cb2627ae5f9f2d374c6f1bd0"}