{"_id":"@fluidinference/fluidvad","_rev":"2-e9e93ae5346b397fd45812db67c31a7e","name":"@fluidinference/fluidvad","dist-tags":{"latest":"0.2.0"},"versions":{"0.1.0":{"name":"@fluidinference/fluidvad","version":"0.1.0","keywords":["vad","voice-activity-detection","silero","speech","audio","wasm","webassembly"],"license":"MIT","_id":"@fluidinference/fluidvad@0.1.0","maintainers":[{"name":"bweng","email":"hello@brandonweng.com"},{"name":"aweng123","email":"alexwengg4@gmail.com"}],"homepage":"https://github.com/FluidInference/FluidVad#readme","bugs":{"url":"https://github.com/FluidInference/FluidVad/issues"},"dist":{"shasum":"a3d0f613b0d40d514dd94c90bbdc447eb0803ecb","tarball":"https://registry.npmjs.org/@fluidinference/fluidvad/-/fluidvad-0.1.0.tgz","fileCount":14,"integrity":"sha512-FsP91wAA7r3N/f2WCNpcPM1zq3pdY0ALZYuQMtwFaujbD3cX+wh9yRRAEBKIX8MJ/kEEpSEyaLwGVq2kIGvLWA==","signatures":[{"sig":"MEUCIQCAdxsQ6p9FKZDLNeEN48atUxrpVo3oQxeYMy0jsjsTBwIgMgE8uDD8Zlxj941nDkaEvMex1lcIXQk1/yTY1Iw1AvQ=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":5579422},"main":"index.js","type":"module","types":"index.d.ts","exports":{".":{"types":"./index.d.ts","default":"./index.js"},"./mic":{"types":"./mic.d.ts","default":"./mic.js"},"./dist/*":"./dist/*","./worklet":"./worklet.js"},"gitHead":"7a7da07889a0d5daf1e049767617d15f15aaae05","_npmUser":{"name":"aweng123","email":"alexwengg4@gmail.com"},"repository":{"url":"git+https://github.com/FluidInference/FluidVad.git","type":"git"},"_npmVersion":"11.17.0","description":"Silero VAD (v6) compiled to WebAssembly — model bundled, zero config, no runtime downloads. Browser + Node.","directories":{},"sideEffects":["./worklet.js"],"_nodeVersion":"26.4.0","_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/fluidvad_0.1.0_1784183759538_0.7059514140827754","host":"s3://npm-registry-packages-npm-production"}},"0.2.0":{"name":"@fluidinference/fluidvad","version":"0.2.0","description":"Silero VAD (v6) compiled to WebAssembly — model bundled, zero config, no runtime downloads. Browser + Node.","license":"MIT","repository":{"type":"git","url":"git+https://github.com/FluidInference/FluidVad.git"},"keywords":["vad","voice-activity-detection","silero","speech","audio","wasm","webassembly"],"type":"module","main":"index.js","types":"index.d.ts","exports":{".":{"types":"./index.d.ts","default":"./index.js"},"./mic":{"types":"./mic.d.ts","default":"./mic.js"},"./worklet":"./worklet.js","./dist/*":"./dist/*"},"sideEffects":["./worklet.js"],"gitHead":"399e3447af54146b3e4da9a4efa46a4e99239b32","_id":"@fluidinference/fluidvad@0.2.0","bugs":{"url":"https://github.com/FluidInference/FluidVad/issues"},"homepage":"https://github.com/FluidInference/FluidVad#readme","_nodeVersion":"26.4.0","_npmVersion":"11.17.0","dist":{"integrity":"sha512-84UYiCYi8K9LEBK1ARqgpibTHMbyW4R56j3+TRqYbNBZf+OLUSkmHipNfzeX34XUokjfOnrRXLMZRE4aecFiUQ==","shasum":"61322be72c1b1a7e147ceb7616155d17046a2ba5","tarball":"https://registry.npmjs.org/@fluidinference/fluidvad/-/fluidvad-0.2.0.tgz","fileCount":14,"unpackedSize":5580027,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCICJEMtALSzD44t5yvSZcRO1UFFpgUr11tRwaXN6w5mXvAiEArkCSdaQvpJrjx/nboB7kVXTEL95U5gguOTNe4oadd3E="}]},"_npmUser":{"name":"aweng123","email":"alexwengg4@gmail.com"},"directories":{},"maintainers":[{"name":"bweng","email":"hello@brandonweng.com"},{"name":"aweng123","email":"alexwengg4@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/fluidvad_0.2.0_1784185292599_0.965029533719854"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-16T06:35:59.337Z","modified":"2026-07-16T07:01:33.050Z","0.1.0":"2026-07-16T06:35:59.739Z","0.2.0":"2026-07-16T07:01:32.890Z"},"bugs":{"url":"https://github.com/FluidInference/FluidVad/issues"},"license":"MIT","homepage":"https://github.com/FluidInference/FluidVad#readme","keywords":["vad","voice-activity-detection","silero","speech","audio","wasm","webassembly"],"repository":{"type":"git","url":"git+https://github.com/FluidInference/FluidVad.git"},"description":"Silero VAD (v6) compiled to WebAssembly — model bundled, zero config, no runtime downloads. Browser + Node.","maintainers":[{"name":"bweng","email":"hello@brandonweng.com"},{"name":"aweng123","email":"alexwengg4@gmail.com"}],"readme":"# FluidVad\n\n**Voice activity detection for npm — Silero VAD (v6) compiled to WebAssembly.\nModel bundled, zero config, no runtime downloads.** Works in the browser,\nNode, and Electron (both processes) on macOS and Windows.\n\n- **No native modules** — pure wasm. Nothing to `electron-rebuild`, no\n  per-arch prebuilds, no onnxruntime peer dependency, no extra binaries to sign.\n- **Model embedded** — the 1.3 MB Silero v6 16 kHz graph ships inside the\n  `.wasm` (5.3 MB raw, 2.3 MB gzipped total). `npm install` and go; nothing\n  fetched at runtime, fully offline.\n- **Streaming + offline** — `SpeechStart`/`SpeechEnd` events with hysteresis,\n  or whole-buffer segmentation. ~150× real-time; a 32 ms frame costs well\n  under a millisecond.\n\n```bash\nnpm i @fluidinference/fluidvad\n```\n\n## Microphone (browser / Electron renderer)\n\n```js\nimport { MicVad } from \"@fluidinference/fluidvad/mic\";\n\nconst mic = new MicVad({\n  onSpeechStart: (t) => console.log(\"speech started\", t),\n  onSpeechEnd: (audio, start, end) => {\n    // audio: Float32Array, 16 kHz mono, whole utterance incl. pre-roll\n    console.log(`utterance ${start.toFixed(2)}s → ${end.toFixed(2)}s`);\n  },\n});\nawait mic.start();\n```\n\n## Buffers (Node / Electron main / browser)\n\n```js\nimport { createVad } from \"@fluidinference/fluidvad\";\n\nconst vad = await createVad({ threshold: 0.5 });\n\n// streaming: push any chunk size, get boundary events\nconst events = vad.push(samples); // Float32Array, 16 kHz mono\n// [{ isStart: true, sampleIndex: 15872, timeSeconds: 0.99 }, ...]\n\n// offline: segment a whole buffer\nconst segments = vad.segment(samples);\n// [{ startTime: 0.9, endTime: 4.21 }, ...]\n```\n\nInput is always **16 kHz mono f32** in `[-1, 1]`. The model consumes\n512-sample frames (32 ms); `push` buffers partial frames internally.\n\n## Electron\n\nRunnable example in [`examples/electron`](examples/electron) (mic UI +\nheadless smoke mode, CI-tested on macOS and Windows).\n\n- **Main / preload (Node env):** `createVad()` works as-is; the wasm is read\n  from disk (asar-transparent).\n- **Renderer with `contextIsolation`:** the renderer cannot `fetch()` `file://`\n  URLs, so hand the wasm bytes over from the preload:\n\n```js\n// preload.cjs\nconst wasmPath = require.resolve(\"@fluidinference/fluidvad/dist/fluidvad_bg.wasm\");\ncontextBridge.exposeInMainWorld(\"fluidvad\", { wasmBytes: new Uint8Array(fs.readFileSync(wasmPath)) });\n\n// renderer\nconst mic = new MicVad({ load: { wasm: window.fluidvad.wasmBytes }, onSpeechEnd: ... });\n```\n\n- CSP: add `'wasm-unsafe-eval'` to `script-src` (compiles wasm without enabling JS `eval`).\n- macOS mic: call `systemPreferences.askForMediaAccess(\"microphone\")` from main and\n  set `NSMicrophoneUsageDescription` when packaging.\n\n## Configuration\n\n| Option | Default | Meaning |\n|---|---|---|\n| `threshold` | 0.5 | entry threshold (frame is speech at ≥) |\n| `negativeThreshold` | `threshold - 0.15` | exit threshold (hysteresis) |\n| `minSpeechDuration` | 0.15 s | drop shorter speech runs |\n| `minSilenceDuration` | 0.75 s | silence needed to close a segment |\n| `maxSpeechDuration` | 14 s | force-split longer segments at the best silence |\n| `speechPadding` | 0.1 s | padding around each segment |\n\n## Development\n\nThe wasm is built from a Rust core (`src/`) using [tract](https://github.com/sonos/tract)\nfor CPU inference — no onnxruntime anywhere.\n\nUpstream Silero ONNX contains `If` nodes whose branches disagree on rank —\nonnxruntime broadcasts through it, strict runtimes cannot. We bake\n`sr = 16000` as a constant, fix the input shapes, and constant-fold with\nonnxruntime's basic optimizer, which eliminates every `If`\n(`scripts/prepare_model.py`, **bit-exact** with upstream). The result is\npre-compiled to NNEF (`examples/export_nnef.rs`) so the shipped wasm only\ncarries tract's lightweight loader. Per-frame parity vs onnxruntime is\nasserted in tests (`tests/model_parity.rs`); the hysteresis / segmentation\nstate machines (adapted from [FluidAudio](https://github.com/FluidInference/FluidAudio))\nare unit-tested with synthetic probability sequences.\n\n```bash\ncargo test --release              # core + parity tests\n./scripts/build_npm.sh            # build the npm package into npm/\npython3 scripts/prepare_model.py  # regenerate model artifacts (needs onnx, onnxruntime)\ncd examples/electron && npm i && FLUIDVAD_SMOKE=1 npx electron .   # headless check\n```\n\n## License\n\nMIT. The bundled Silero VAD model is © Silero Team, MIT-licensed\n([SILERO_LICENSE](SILERO_LICENSE), [snakers4/silero-vad](https://github.com/snakers4/silero-vad)).\n","readmeFilename":"README.md"}