{"_id":"@anudit/kitten-tts-webgpu","_rev":"2-0623d1d6e4c6f8d7e7d83c0d2b65dd73","name":"@anudit/kitten-tts-webgpu","dist-tags":{"latest":"0.2.1"},"versions":{"0.2.0":{"name":"@anudit/kitten-tts-webgpu","version":"0.2.0","keywords":["tts","text-to-speech","webgpu","browser","kitten-tts","speech-synthesis"],"license":"Apache-2.0","_id":"@anudit/kitten-tts-webgpu@0.2.0","maintainers":[{"name":"anudit","email":"nagaranudit@gmail.com"}],"homepage":"https://github.com/svenflow/kitten-tts-webgpu#readme","bugs":{"url":"https://github.com/svenflow/kitten-tts-webgpu/issues"},"dist":{"shasum":"5ba47357bab78fa9dca4ec9c9f22682e6a37b8f3","tarball":"https://registry.npmjs.org/@anudit/kitten-tts-webgpu/-/kitten-tts-webgpu-0.2.0.tgz","fileCount":13,"integrity":"sha512-pcPUsUHuh/QJDmRGkvmUKOMORTiicgBIqrwOAcb4/GLyigraPKBaAcCAtWKjsE4C3pX0+QG0/RnaA8N2y/CWHg==","signatures":[{"sig":"MEQCIFSw2rFTQOTHNqsAaufXm+qtN51tV+GbyiTFinfpZQJOAiBO0LZaVknZwzx0OP2j99/qY9Yhl11NPxKuGSzbeJy2rQ==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":3114717},"main":"./dist-lib/index.js","type":"module","types":"./dist-lib/index.d.ts","module":"./dist-lib/index.js","exports":{".":{"types":"./dist-lib/index.d.ts","import":"./dist-lib/index.js","default":"./dist-lib/index.js"}},"gitHead":"c6f5b503db7997130bf8991e0f55937c02aa27ae","scripts":{"dev":"vite","test":"tsx tests/phonemizer.test.ts","build":"vite build","prepare":"patch-package","preview":"vite preview","build:lib":"BUILD_LIB=1 vite build && tsc --project tsconfig.lib.json","prepublishOnly":"npm run build:lib"},"_npmUser":{"name":"anudit","email":"nagaranudit@gmail.com"},"repository":{"url":"git+https://github.com/svenflow/kitten-tts-webgpu.git","type":"git"},"_npmVersion":"11.13.0","description":"Run Kitten TTS (80M) locally in the browser via WebGPU. One function call: textToSpeech('Hello!') → WAV blob.","directories":{},"sideEffects":false,"_nodeVersion":"24.16.0","dependencies":{"phonemizer":"^1.2.1"},"_hasShrinkwrap":false,"devDependencies":{"tsx":"^4.7.0","vite":"^6.0.0","typescript":"^5.7.0","@webgpu/types":"^0.1.69","patch-package":"^8.0.0"},"_npmOperationalInternal":{"tmp":"tmp/kitten-tts-webgpu_0.2.0_1783144648404_0.8070038054005662","host":"s3://npm-registry-packages-npm-production"}},"0.2.1":{"name":"@anudit/kitten-tts-webgpu","version":"0.2.1","description":"Run Kitten TTS (80M) locally in the browser via WebGPU. One function call: textToSpeech('Hello!') → WAV blob.","type":"module","main":"./dist-lib/index.js","module":"./dist-lib/index.js","types":"./dist-lib/index.d.ts","exports":{".":{"types":"./dist-lib/index.d.ts","import":"./dist-lib/index.js","default":"./dist-lib/index.js"}},"sideEffects":false,"scripts":{"prepare":"patch-package","dev":"vite","build":"vite build","build:lib":"BUILD_LIB=1 vite build && tsc --project tsconfig.lib.json","preview":"vite preview","prepublishOnly":"npm run build:lib","test":"tsx tests/phonemizer.test.ts"},"devDependencies":{"@webgpu/types":"^0.1.69","patch-package":"^8.0.0","tsx":"^4.7.0","typescript":"^5.7.0","vite":"^6.0.0"},"license":"Apache-2.0","dependencies":{"phonemizer":"^1.2.1"},"keywords":["tts","text-to-speech","webgpu","browser","kitten-tts","speech-synthesis"],"repository":{"type":"git","url":"git+https://github.com/svenflow/kitten-tts-webgpu.git"},"gitHead":"b6e50d2d29647541e78b0955df141cb9993c19ae","_id":"@anudit/kitten-tts-webgpu@0.2.1","bugs":{"url":"https://github.com/svenflow/kitten-tts-webgpu/issues"},"homepage":"https://github.com/svenflow/kitten-tts-webgpu#readme","_nodeVersion":"24.16.0","_npmVersion":"11.13.0","dist":{"integrity":"sha512-/ZLkgw4la0RkRRXa651RkjBNFVdeDXh9s3XbJh1p31Gy4Jx7mK1XVjmwj9LnJbEIU5ZpatH8VSzfKFl9pmDaQQ==","shasum":"3abd794de70d5115fe507e943a376ab0488a0d75","tarball":"https://registry.npmjs.org/@anudit/kitten-tts-webgpu/-/kitten-tts-webgpu-0.2.1.tgz","fileCount":13,"unpackedSize":3118572,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIF4gM9G7eszFRil9CYOCrI35L5CYmtATAp73U8C7VhpuAiEAmUx9yJAbHjr3epu8PMwxarg8blu32nJzVKWYtqG9SkQ="}]},"_npmUser":{"name":"anudit","email":"nagaranudit@gmail.com"},"directories":{},"maintainers":[{"name":"anudit","email":"nagaranudit@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/kitten-tts-webgpu_0.2.1_1786093078169_0.0264709320578016"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-04T05:57:28.232Z","modified":"2026-08-07T08:57:58.544Z","0.2.0":"2026-07-04T05:57:28.544Z","0.2.1":"2026-08-07T08:57:58.353Z"},"bugs":{"url":"https://github.com/svenflow/kitten-tts-webgpu/issues"},"license":"Apache-2.0","homepage":"https://github.com/svenflow/kitten-tts-webgpu#readme","keywords":["tts","text-to-speech","webgpu","browser","kitten-tts","speech-synthesis"],"repository":{"type":"git","url":"git+https://github.com/svenflow/kitten-tts-webgpu.git"},"description":"Run Kitten TTS (80M) locally in the browser via WebGPU. One function call: textToSpeech('Hello!') → WAV blob.","maintainers":[{"name":"anudit","email":"nagaranudit@gmail.com"}],"readme":"# Kitten TTS WebGPU\n\n[![npm](https://img.shields.io/npm/v/kitten-tts-webgpu)](https://www.npmjs.com/package/kitten-tts-webgpu)\n[![license](https://img.shields.io/npm/l/kitten-tts-webgpu)](./LICENSE)\n\n**Pure WebGPU text-to-speech for the browser. 80M params, sub-second on desktop, ~1.2s on iPhone. No ONNX Runtime, no WASM inference — just 29 compute shaders. 753KB gzipped JS + model weights downloaded at runtime.**\n\n[**Live Demo**](https://svenflow.github.io/kitten-tts-webgpu/) | [npm](https://www.npmjs.com/package/kitten-tts-webgpu) | [Model Card](https://huggingface.co/KittenML/kitten-tts-mini-0.8)\n\n---\n\n## Quick Start\n\n```bash\nnpm install kitten-tts-webgpu\n```\n\n```typescript\nimport { textToSpeech } from 'kitten-tts-webgpu';\n\nconst blob = await textToSpeech(\"The quick brown fox jumps over the lazy dog.\");\nconst audio = new Audio(URL.createObjectURL(blob));\naudio.play();\n```\n\nOne function. Text in, WAV blob out (16-bit PCM, 24 kHz mono). The model downloads on first call and is cached for subsequent calls. Full TypeScript types included.\n\n> **Note:** This library requires WebGPU. For server-side rendering frameworks (Next.js, Nuxt), dynamically import on the client side only.\n\n## Size & Performance\n\n### What gets downloaded\n\n| | Size | When |\n|-|------|------|\n| **JS bundle** | **753 KB** gzipped (2.9 MB raw) | `npm install` / bundled into your app |\n| **Model weights** | 24–78 MB (see below) | First `textToSpeech()` call, cached by browser |\n\nThe JS bundle includes the WebGPU engine, 29 compute shaders, and a 234K-word phonemizer dictionary. No WASM binaries, no ONNX Runtime.\n\n### Models\n\nThree [Kitten TTS v0.8](https://huggingface.co/KittenML) sizes, same API:\n\n| Model | Params | Download | M4 Pro (Chrome) | iPhone 17 Pro Max (Safari) |\n|-------|--------|----------|------------------|----------------------------|\n| **Mini** | 80M | 78 MB | 3.3× RT | 1.3× RT |\n| **Micro** | 40M | 43 MB | 6.2× RT | 3.4× RT |\n| **Nano** | 15M | 26 MB | 7.3× RT | 4.8× RT |\n\n*RT = real-time factor (audio duration ÷ generation time). Higher is better. Download = ONNX weights + voices. Times are for warm generation (model already in GPU).*\n\n```typescript\nawait textToSpeech(\"Hello world\");                        // Default: nano (fastest, 26 MB)\nawait textToSpeech(\"Hello world\", { model: 'micro' });    // Balanced (43 MB)\nawait textToSpeech(\"Hello world\", { model: 'mini' });     // Best quality (78 MB)\n```\n\n## Options\n\n```typescript\nconst blob = await textToSpeech(\"Welcome to the future.\", {\n  voice: \"Leo\",        // 8 voices: Bella, Luna, Rosie, Kiki, Jasper, Bruno, Hugo, Leo\n  speed: 1.2,          // 0.5x – 2.0x\n  model: \"micro\",      // mini | micro | nano\n  onProgress: (stage) => console.log(stage), // string: \"Initializing WebGPU…\", \"Downloading…\", \"Generating speech…\", etc.\n});\n```\n\n### Voices\n\n| Female | Male |\n|--------|------|\n| Bella  | Jasper |\n| Luna   | Bruno |\n| Rosie  | Hugo |\n| Kiki   | Leo |\n\n## Error Handling\n\n```typescript\n// Check for WebGPU support\nif (!navigator.gpu) {\n  console.log(\"WebGPU not available — use Chrome 113+, Edge 113+, or Safari 26+\");\n}\n\n// textToSpeech throws on:\n// - No WebGPU support\n// - Network error (model download fails)\n// - Empty text input\ntry {\n  const blob = await textToSpeech(\"Hello\");\n} catch (err) {\n  console.error(\"TTS failed:\", err.message);\n}\n```\n\n## Advanced: Direct Engine Access\n\nFor repeated generations or fine-grained control:\n\n```typescript\nimport { KittenTTSEngine, textToInputIds, float32ToWav } from 'kitten-tts-webgpu';\n\nconst engine = new KittenTTSEngine();\nawait engine.init();\nawait engine.loadModel(onnxUrl, voicesUrl);\n\nconst { ids } = await textToInputIds(\"Hello world\");\nconst { waveform } = await engine.generate(ids, \"Bella\", 1.0);\n// waveform: Float32Array of 24kHz PCM samples\n\nconst wavBlob = float32ToWav(waveform, 24000);\n```\n\n## How It Works\n\n29 hand-written [WGSL compute shaders](./src/shaders.ts) execute the full TTS pipeline on GPU:\n\n```\nText → Phonemes (234K-word dictionary + espeak rules in pure JS)\n  → ALBERT encoder (embedding, multi-head attention, FFN)\n  → Duration predictor (LSTM + CNN)\n  → Acoustic decoder (LSTM + AdaIN + CNN, style-conditioned)\n  → HiFi-GAN vocoder (ConvTranspose1d, Snake activations, iSTFT)\n  → 24kHz WAV\n```\n\n**Why not ONNX Runtime Web?**\n\nMost browser TTS uses ONNX Runtime Web (~2MB WASM binary + C++ runtime). This project takes a different approach:\n\n- **Custom ONNX parser** — dequantizes int8/uint8/float16 weights in pure TypeScript, no C++ runtime\n- **234K-word phonemizer** — espeak-ng rules ported to pure JS (WASM espeak hangs on iOS Safari)\n- **GPU buffer pooling** — reuses buffers across HiFi-GAN iterations, ~130MB peak on mobile\n- **Dynamic architecture** — detects model dimensions from weight shapes, one engine for all 3 sizes\n\n## Browser Support\n\n| Browser | Status |\n|---------|--------|\n| Chrome 113+ | ✅ |\n| Edge 113+ | ✅ |\n| Safari 26+ (macOS/iOS) | ✅ |\n| Firefox Nightly | Experimental |\n\n## FAQ\n\n**Max input length?** Recommended under ~500 characters per call. For longer text, split into sentences.\n\n**Languages?** English only (matches the upstream Kitten TTS model).\n\n**Offline?** Yes, after the model is cached in the browser. No server needed for inference.\n\n**Self-hosting models?** Pass custom URLs to `KittenTTSEngine.loadModel(onnxUrl, voicesUrl)`.\n\n**Bundle size?** 753KB gzipped (2.9MB raw). Includes engine, 29 compute shaders, and 234K-word phonemizer dictionary. Model weights (24–78MB depending on model size) are downloaded separately at runtime on first call and cached by the browser.\n\n**Model license?** Kitten TTS models are released under [Apache 2.0](https://huggingface.co/KittenML/kitten-tts-mini-0.8). Code in this repo is Apache 2.0.\n\n## Development\n\n```bash\ngit clone https://github.com/svenflow/kitten-tts-webgpu.git\ncd kitten-tts-webgpu\nnpm install\nnpm run dev       # Dev server\nnpm run build     # Production build\nnpm test          # Phonemizer tests\n```\n\n## Credits\n\n- [Kitten TTS](https://github.com/KittenML/KittenTTS) by KittenML — original model and architecture ([models on HuggingFace](https://huggingface.co/KittenML/kitten-tts-mini-0.8), Apache 2.0)\n- [espeak-ng](https://github.com/espeak-ng/espeak-ng) pronunciation dictionary and letter-to-sound rules (GPL-3.0, bundled as data files)\n- [phonemizer](https://www.npmjs.com/package/phonemizer) by Xenova (espeak-ng WASM, used as primary backend on Chrome/Firefox; pure JS fallback on Safari)\n\n## License\n\nApache-2.0\n","readmeFilename":"README.md"}