{"_id":"@create-voice-agent/elevenlabs","_rev":"2-8208a4b8281186bbfef96a610ec13ac9","name":"@create-voice-agent/elevenlabs","dist-tags":{"latest":"0.1.0"},"versions":{"0.0.1":{"name":"@create-voice-agent/elevenlabs","version":"0.0.1","_id":"@create-voice-agent/elevenlabs@0.0.1","maintainers":[{"name":"christian-bromann","email":"mail@bromann.dev"}],"dist":{"shasum":"3c99693d6c1768fed4d961d58728de77a4ca26d7","tarball":"https://registry.npmjs.org/@create-voice-agent/elevenlabs/-/elevenlabs-0.0.1.tgz","fileCount":9,"integrity":"sha512-nNvVr8UBysNPluBOjnMtBpxQTQJdokNjIk73O9X19hNEbO7rGz8GX2pETyEnLvMiLD4OW2DltYJK5P6Q48ccfQ==","signatures":[{"sig":"MEUCIQCIx8Al7xa9VLsT3owwhQssNrIuJ6EQGnwj6rjz+y3dtQIgKzy9vnIFvoaUrZVFnbSxftBPH9LrdT4X4U96Rbw2Ioc=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":31274},"type":"module","_from":"file:create-voice-agent-elevenlabs-0.0.1.tgz","exports":{".":"./src/index.ts"},"scripts":{"build":"tsc","typecheck":"tsc --noEmit"},"_npmUser":{"name":"christian-bromann","email":"mail@bromann.dev"},"_resolved":"/private/var/folders/ww/pg23lf097_s2fg3rfb9ykqnc0000gn/T/e236148dfd4cc76d21a41847fb2e368e/create-voice-agent-elevenlabs-0.0.1.tgz","_integrity":"sha512-nNvVr8UBysNPluBOjnMtBpxQTQJdokNjIk73O9X19hNEbO7rGz8GX2pETyEnLvMiLD4OW2DltYJK5P6Q48ccfQ==","_npmVersion":"10.8.2","description":"ElevenLabs Text-to-Speech integration for voice agents","directories":{},"_nodeVersion":"20.19.4","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"typescript":"~5.9.3","@types/node":"^24.10.1","create-voice-agent":"0.0.1"},"peerDependencies":{"create-voice-agent":"0.0.1"},"_npmOperationalInternal":{"tmp":"tmp/elevenlabs_0.0.1_1764195819737_0.5875366802502076","host":"s3://npm-registry-packages-npm-production"}},"0.1.0":{"name":"@create-voice-agent/elevenlabs","version":"0.1.0","type":"module","description":"ElevenLabs Text-to-Speech integration for voice agents","exports":{".":"./src/index.ts"},"devDependencies":{"@types/node":"^24.10.1","typescript":"~5.9.3","create-voice-agent":"0.1.0"},"peerDependencies":{"create-voice-agent":"0.1.0"},"publishConfig":{"access":"public"},"scripts":{"build":"tsc","typecheck":"tsc --noEmit"},"_id":"@create-voice-agent/elevenlabs@0.1.0","_integrity":"sha512-Sfoz9ltvl0lx+Cy1nRzg9vnAQ+fmmFlZHOO6ivzD4X3xtEkqk/lL46xZLSSZjNB3OQ0WklEndsSnSuBH7v9mFA==","_resolved":"/private/var/folders/ww/pg23lf097_s2fg3rfb9ykqnc0000gn/T/7f74cacf1034eeca25a631ecbb18418d/create-voice-agent-elevenlabs-0.1.0.tgz","_from":"file:create-voice-agent-elevenlabs-0.1.0.tgz","_nodeVersion":"20.19.4","_npmVersion":"10.8.2","dist":{"integrity":"sha512-Sfoz9ltvl0lx+Cy1nRzg9vnAQ+fmmFlZHOO6ivzD4X3xtEkqk/lL46xZLSSZjNB3OQ0WklEndsSnSuBH7v9mFA==","shasum":"27e7a3b1746136d16f6859beab3838e0792d4379","tarball":"https://registry.npmjs.org/@create-voice-agent/elevenlabs/-/elevenlabs-0.1.0.tgz","fileCount":9,"unpackedSize":34043,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQCSx/PdasnrI+6BNg+4XW5Fs9pfBo1tyOkfu1UewfSdOwIhAO2DJJiZpSBdFs8bjUTnrJW/XylGzX2sPXBHGP5LlhEF"}]},"_npmUser":{"name":"christian-bromann","email":"mail@bromann.dev"},"directories":{},"maintainers":[{"name":"christian-bromann","email":"mail@bromann.dev"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/elevenlabs_0.1.0_1764204970192_0.8008865204697249"},"_hasShrinkwrap":false}},"time":{"created":"2025-11-26T22:23:39.660Z","modified":"2025-11-27T00:56:10.656Z","0.0.1":"2025-11-26T22:23:39.938Z","0.1.0":"2025-11-27T00:56:10.409Z"},"description":"ElevenLabs Text-to-Speech integration for voice agents","maintainers":[{"name":"christian-bromann","email":"mail@bromann.dev"}],"readme":"# @create-voice-agent/elevenlabs 🔊\n\nElevenLabs Text-to-Speech integration for [create-voice-agent](../../core/README.md).\n\nThis package provides high-quality, low-latency voice synthesis using [ElevenLabs' streaming TTS API](https://elevenlabs.io/docs/api-reference/text-to-speech).\n\n## Installation\n\n```bash\nnpm install @create-voice-agent/elevenlabs\n# or\npnpm add @create-voice-agent/elevenlabs\n```\n\n## Quick Start\n\n```typescript\nimport { createVoiceAgent } from \"create-voice-agent\";\nimport { AssemblyAISpeechToText } from \"@create-voice-agent/assemblyai\";\nimport { ElevenLabsTextToSpeech } from \"@create-voice-agent/elevenlabs\";\n\nconst voiceAgent = createVoiceAgent({\n  model: new ChatOpenAI({ model: \"gpt-4o\" }),\n  \n  stt: new AssemblyAISpeechToText({ /* ... */ }),\n  \n  tts: new ElevenLabsTextToSpeech({\n    apiKey: process.env.ELEVENLABS_API_KEY!,\n    voiceId: process.env.ELEVENLABS_VOICE_ID!,\n  }),\n});\n```\n\n## API Reference\n\n### `ElevenLabsTextToSpeech`\n\nStreaming Text-to-Speech model using ElevenLabs' HTTP API.\n\n```typescript\nimport { ElevenLabsTextToSpeech } from \"@create-voice-agent/elevenlabs\";\n\nconst tts = new ElevenLabsTextToSpeech({\n  apiKey: process.env.ELEVENLABS_API_KEY!,\n  voiceId: \"21m00Tcm4TlvDq8ikWAM\", // Rachel\n  \n  // Optional configuration\n  modelId: \"eleven_flash_v2_5\",\n  outputFormat: \"pcm_16000\",\n  optimizeStreamingLatency: 3,\n  \n  // Voice settings\n  voiceSettings: {\n    stability: 0.5,\n    similarityBoost: 0.75,\n    style: 0.3,\n    speed: 1.0,\n    useSpeakerBoost: true,\n  },\n  \n  // Token batching\n  flushDelayMs: 300,\n  \n  // Callbacks\n  onAudioComplete: () => console.log(\"Finished speaking\"),\n  onInterrupt: () => console.log(\"Speech interrupted\"),\n});\n```\n\n### Configuration Options\n\n| Option | Type | Default | Description |\n|--------|------|---------|-------------|\n| `apiKey` | `string` | **required** | ElevenLabs API key |\n| `voiceId` | `string` | **required** | Voice ID to use |\n| `modelId` | `string` | `\"eleven_flash_v2_5\"` | TTS model ID |\n| `languageCode` | `string` | - | ISO 639-1 language code (e.g., \"en\", \"es\") |\n| `outputFormat` | `string` | `\"pcm_16000\"` | Audio output format |\n| `optimizeStreamingLatency` | `0-4` | `3` | Latency optimization level |\n| `flushDelayMs` | `number` | `300` | Token batching delay (ms) |\n| `seed` | `number` | - | Seed for deterministic generation |\n| `previousText` | `string` | - | Context text before current request |\n| `nextText` | `string` | - | Context text after current request |\n| `applyTextNormalization` | `\"auto\" \\| \"on\" \\| \"off\"` | `\"auto\"` | Text normalization mode |\n| `applyLanguageTextNormalization` | `boolean` | `false` | Language-specific normalization (⚠️ high latency) |\n\n### Voice Settings\n\nFine-tune the generated speech characteristics:\n\n```typescript\ninterface ElevenLabsVoiceSettings {\n  /** Speech stability (0-1). Lower = more expressive, higher = more consistent */\n  stability?: number;\n  \n  /** Voice similarity (0-1). Higher = closer to reference voice */\n  similarityBoost?: number;\n  \n  /** Enable speaker boost for enhanced clarity */\n  useSpeakerBoost?: boolean;\n  \n  /** Style/expressiveness (0-1). Only for certain models */\n  style?: number;\n  \n  /** Speech speed (0.5-2.0) */\n  speed?: number;\n}\n```\n\n#### Example: Expressive Storytelling Voice\n\n```typescript\nconst tts = new ElevenLabsTextToSpeech({\n  apiKey: process.env.ELEVENLABS_API_KEY!,\n  voiceId: \"your-voice-id\",\n  voiceSettings: {\n    stability: 0.3,        // More expressive\n    similarityBoost: 0.8,  // Close to reference\n    style: 0.6,            // More stylized\n    speed: 0.9,            // Slightly slower\n  },\n});\n```\n\n#### Example: Consistent Professional Voice\n\n```typescript\nconst tts = new ElevenLabsTextToSpeech({\n  apiKey: process.env.ELEVENLABS_API_KEY!,\n  voiceId: \"your-voice-id\",\n  voiceSettings: {\n    stability: 0.8,        // Very consistent\n    similarityBoost: 0.7,\n    useSpeakerBoost: true, // Enhanced clarity\n    speed: 1.0,\n  },\n});\n```\n\n### Models\n\n| Model ID | Description | Best For |\n|----------|-------------|----------|\n| `eleven_flash_v2_5` | Fastest, lowest latency (default) | Real-time conversations |\n| `eleven_turbo_v2_5` | Fast with higher quality | Balanced speed/quality |\n| `eleven_multilingual_v2` | Best multilingual support | Non-English or mixed languages |\n| `eleven_monolingual_v1` | Original English model | Legacy compatibility |\n\n### Output Formats\n\n#### PCM (Recommended for voice agents)\n\n| Format | Sample Rate | Description |\n|--------|-------------|-------------|\n| `pcm_8000` | 8 kHz | Telephone quality |\n| `pcm_16000` | 16 kHz | Standard voice (default) |\n| `pcm_22050` | 22.05 kHz | Higher quality |\n| `pcm_24000` | 24 kHz | High quality |\n| `pcm_44100` | 44.1 kHz | CD quality |\n| `pcm_48000` | 48 kHz | Professional quality |\n\n#### MP3\n\n| Format | Sample Rate | Bitrate |\n|--------|-------------|---------|\n| `mp3_22050_32` | 22.05 kHz | 32 kbps |\n| `mp3_44100_64` | 44.1 kHz | 64 kbps |\n| `mp3_44100_128` | 44.1 kHz | 128 kbps |\n| `mp3_44100_192` | 44.1 kHz | 192 kbps |\n\n#### Other Formats\n\n| Format | Description |\n|--------|-------------|\n| `ulaw_8000` | μ-law 8kHz (telephony) |\n| `alaw_8000` | A-law 8kHz (telephony) |\n| `opus_48000_*` | Opus codec (32-192 kbps) |\n\n### Latency Optimization\n\nControl the trade-off between latency and quality:\n\n| Level | Description | Use Case |\n|-------|-------------|----------|\n| `0` | No optimization | Highest quality |\n| `1` | ~50% latency reduction | Balanced |\n| `2` | ~75% latency reduction | Lower latency |\n| `3` | Maximum optimization (default) | Real-time conversations |\n| `4` | Max + disable text normalizer | Fastest (may mispronounce numbers/dates) |\n\n```typescript\n// For real-time conversations (fastest)\nconst tts = new ElevenLabsTextToSpeech({\n  apiKey: process.env.ELEVENLABS_API_KEY!,\n  voiceId: \"your-voice-id\",\n  optimizeStreamingLatency: 4,\n});\n\n// For pre-recorded content (highest quality)\nconst tts = new ElevenLabsTextToSpeech({\n  apiKey: process.env.ELEVENLABS_API_KEY!,\n  voiceId: \"your-voice-id\",\n  optimizeStreamingLatency: 0,\n});\n```\n\n### Token Batching\n\nThe TTS model batches incoming text tokens before sending to ElevenLabs for more natural speech generation:\n\n```typescript\nconst tts = new ElevenLabsTextToSpeech({\n  apiKey: process.env.ELEVENLABS_API_KEY!,\n  voiceId: \"your-voice-id\",\n  \n  // Wait 300ms after last token before generating speech\n  flushDelayMs: 300,\n});\n```\n\n- **Lower values** (100-200ms): Faster response, may sound choppy\n- **Higher values** (400-500ms): More natural speech, higher latency\n- **Default** (300ms): Good balance for most use cases\n\n### Instance Methods\n\n#### `interrupt()`\n\nInterrupt the current speech generation. Useful for barge-in handling.\n\n```typescript\n// User started speaking - stop the agent\ntts.interrupt();\n```\n\n#### `speak(text: string): ReadableStream<Buffer>`\n\nGenerate speech directly without going through the voice pipeline. Returns a `ReadableStream` of PCM audio buffers.\n\nThis is useful for:\n\n- **Initial greetings** when a call starts\n- **System announcements** that bypass the agent\n- **One-off speech synthesis** outside of conversations\n\n```typescript\nconst tts = new ElevenLabsTextToSpeech({\n  apiKey: process.env.ELEVENLABS_API_KEY!,\n  voiceId: \"your-voice-id\",\n});\n\n// Generate and play a greeting\nconst audioStream = tts.speak(\"Welcome to our service! How can I help you?\");\n\nfor await (const chunk of audioStream) {\n  // Send to audio output (speakers, WebRTC, etc.)\n  audioOutput.write(chunk);\n}\n```\n\nThe `speak()` method uses the same voice settings and configuration as the main TTS pipeline, ensuring consistent voice quality.\n\n### Callbacks\n\n#### `onAudioComplete`\n\nCalled when speech generation finishes (not interrupted).\n\n```typescript\nconst tts = new ElevenLabsTextToSpeech({\n  apiKey: process.env.ELEVENLABS_API_KEY!,\n  voiceId: \"your-voice-id\",\n  onAudioComplete: () => {\n    console.log(\"Agent finished speaking\");\n    // Trigger next action, update UI, etc.\n  },\n});\n```\n\n#### `onInterrupt`\n\nCalled when speech is interrupted (e.g., by barge-in).\n\n```typescript\nconst tts = new ElevenLabsTextToSpeech({\n  apiKey: process.env.ELEVENLABS_API_KEY!,\n  voiceId: \"your-voice-id\",\n  onInterrupt: () => {\n    console.log(\"Speech was interrupted\");\n  },\n});\n```\n\n## Finding Voice IDs\n\n### Using the API\n\n```typescript\nconst response = await fetch(\"https://api.elevenlabs.io/v1/voices\", {\n  headers: { \"xi-api-key\": process.env.ELEVENLABS_API_KEY! },\n});\nconst { voices } = await response.json();\n\nfor (const voice of voices) {\n  console.log(`${voice.name}: ${voice.voice_id}`);\n}\n```\n\n### Popular Pre-made Voices\n\n| Voice | ID | Description |\n|-------|-----|-------------|\n| Rachel | `21m00Tcm4TlvDq8ikWAM` | American female, calm |\n| Domi | `AZnzlk1XvdvUeBnXmlld` | American female, strong |\n| Bella | `EXAVITQu4vr4xnSDxMaL` | American female, soft |\n| Antoni | `ErXwobaYiN019PkySvjV` | American male, warm |\n| Josh | `TxGEqnHWrfWFTfGW9XjX` | American male, deep |\n| Arnold | `VR6AewLTigWG4xSOukaG` | American male, crisp |\n| Adam | `pNInz6obpgDQGcFmaJgB` | American male, deep |\n| Sam | `yoZ06aMxZJJ28mfd3POQ` | American male, raspy |\n\n## Multilingual Support\n\nFor non-English or mixed-language content:\n\n```typescript\nconst tts = new ElevenLabsTextToSpeech({\n  apiKey: process.env.ELEVENLABS_API_KEY!,\n  voiceId: \"your-voice-id\",\n  modelId: \"eleven_multilingual_v2\",\n  languageCode: \"es\", // Spanish\n});\n```\n\n### Supported Languages\n\nThe `eleven_multilingual_v2` model supports 29 languages including:\nEnglish, Spanish, French, German, Italian, Portuguese, Polish, Hindi, Arabic, Japanese, Korean, Mandarin, and more.\n\n## Text Normalization\n\nControl how text is processed before synthesis:\n\n```typescript\nconst tts = new ElevenLabsTextToSpeech({\n  apiKey: process.env.ELEVENLABS_API_KEY!,\n  voiceId: \"your-voice-id\",\n  \n  // \"auto\" - Let the system decide (default)\n  // \"on\"   - Always normalize (spell out numbers, dates, etc.)\n  // \"off\"  - Skip normalization\n  applyTextNormalization: \"on\",\n});\n```\n\n**Note:** For `eleven_turbo_v2_5` and `eleven_flash_v2_5` models, text normalization requires an Enterprise plan.\n\n## Deterministic Generation\n\nUse seeds for reproducible output:\n\n```typescript\nconst tts = new ElevenLabsTextToSpeech({\n  apiKey: process.env.ELEVENLABS_API_KEY!,\n  voiceId: \"your-voice-id\",\n  seed: 12345, // 0 to 4294967295\n});\n```\n\n**Note:** Determinism is not guaranteed but the system will attempt to produce consistent results.\n\n## Complete Example\n\n```typescript\nimport { createVoiceAgent, createThinkingFillerMiddleware } from \"create-voice-agent\";\nimport { AssemblyAISpeechToText } from \"@create-voice-agent/assemblyai\";\nimport { ElevenLabsTextToSpeech } from \"@create-voice-agent/elevenlabs\";\nimport { ChatOpenAI } from \"@langchain/openai\";\n\nconst tts = new ElevenLabsTextToSpeech({\n  apiKey: process.env.ELEVENLABS_API_KEY!,\n  voiceId: process.env.ELEVENLABS_VOICE_ID!,\n  modelId: \"eleven_flash_v2_5\",\n  outputFormat: \"pcm_16000\",\n  optimizeStreamingLatency: 3,\n  \n  voiceSettings: {\n    stability: 0.5,\n    similarityBoost: 0.75,\n    useSpeakerBoost: true,\n  },\n  \n  onAudioComplete: () => console.log(\"Agent finished speaking\"),\n});\n\nconst stt = new AssemblyAISpeechToText({\n  apiKey: process.env.ASSEMBLYAI_API_KEY!,\n  onSpeechStart: () => {\n    // Barge-in: user started speaking, interrupt the agent\n    tts.interrupt();\n  },\n});\n\nconst voiceAgent = createVoiceAgent({\n  model: new ChatOpenAI({ model: \"gpt-4o\" }),\n  prompt: \"You are a friendly voice assistant. Keep responses concise.\",\n  \n  stt,\n  tts,\n  \n  middleware: [\n    createThinkingFillerMiddleware({ thresholdMs: 1000 }),\n  ],\n});\n\n// Process audio streams\nconst audioOutput = voiceAgent.process(audioInputStream);\n```\n\n## License\n\nMIT\n","readmeFilename":"README.md"}