{"_id":"@attenlabs/saa-js","_rev":"6-fbd1dd25ca82d5b8a503f3559032e0e7","name":"@attenlabs/saa-js","dist-tags":{"latest":"0.8.0"},"versions":{"0.3.0":{"name":"@attenlabs/saa-js","version":"0.3.0","keywords":["sd-attention","attention","saa","audio","video","vad","realtime","websocket"],"author":{"name":"Attention Labs","email":"omar@attentionlabs.ai"},"license":"MIT","_id":"@attenlabs/saa-js@0.3.0","maintainers":[{"name":"omar_e","email":"omar@attentionlabs.ai"}],"homepage":"https://attentionlabs.ai","dist":{"shasum":"e9d623d8428e49ce6dc9b57f8b5ac8ce11fec1aa","tarball":"https://registry.npmjs.org/@attenlabs/saa-js/-/saa-js-0.3.0.tgz","fileCount":28,"integrity":"sha512-LuU0TBggArNKS0QWqh8Zuzpq1bTPbjazW5NaKAHWIdGu2y/XTtjN5Q5mMzB5v3kpamyRezZ554soSKSkGNfyfA==","signatures":[{"sig":"MEYCIQD22gt6KqQHTX+9skDQ0YVPAmp/iaAVcaSqMISdyGofugIhAMLbTfOo1yZr+ZU01fwGHRiIP0+zSDwcXun4tG7UaVgR","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":62015},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","module":"./dist/index.js","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"},"./audio-processor.js":"./dist/audio-processor.js"},"gitHead":"be14bef221cdd3bce3dec222c8d2cce17c2bc527","scripts":{"build":"npx tsc && cp src/audio-processor.js dist/audio-processor.js","clean":"rm -rf dist","prebuild":"node scripts/embed-worklet.mjs","prepublishOnly":"npm run clean && npm run build"},"_npmUser":{"name":"omar_e","email":"omar@attentionlabs.ai"},"_npmVersion":"10.9.3","description":"Browser JS SDK for the SD Attention Server (SAA). Streams microphone + webcam to the inference server and emits typed events for predictions, VAD, conversation state, and speech audio ready for downstream LLM use.","directories":{},"_nodeVersion":"22.19.0","_hasShrinkwrap":false,"devDependencies":{"typescript":"^5.4.0"},"_npmOperationalInternal":{"tmp":"tmp/saa-js_0.3.0_1778771491135_0.44933909730473376","host":"s3://npm-registry-packages-npm-production"}},"0.3.1":{"name":"@attenlabs/saa-js","version":"0.3.1","keywords":["sd-attention","attention","saa","audio","video","vad","realtime","websocket"],"author":{"name":"Attention Labs","email":"omar@attentionlabs.ai"},"license":"MIT","_id":"@attenlabs/saa-js@0.3.1","maintainers":[{"name":"omar_e","email":"omar@attentionlabs.ai"}],"homepage":"https://attentionlabs.ai","dist":{"shasum":"ca55b9d03b851eb139d3468305989f06f6ade8a6","tarball":"https://registry.npmjs.org/@attenlabs/saa-js/-/saa-js-0.3.1.tgz","fileCount":28,"integrity":"sha512-y85a+4RzgOmekLjyIDUuc/CcFUgNZFqTj3Z0bn/E+Z/SpZCvuZhHI8LWuuauaVKth7E6NOhLJwtJ1pLBA+Blng==","signatures":[{"sig":"MEUCIQCkCJKZNjCnjEYIC4zWzYri1lxIYfsqBuKWDsE4yxMlnQIgffGUgHX6L+54JkWWVAP3/QpKaMsyUsi/df5BTncCmTc=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":63034},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","module":"./dist/index.js","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"},"./audio-processor.js":"./dist/audio-processor.js"},"gitHead":"f20c1d4dd98080c9dcb5e947603769360746d37e","scripts":{"build":"npx tsc && cp src/audio-processor.js dist/audio-processor.js","clean":"rm -rf dist","prebuild":"node scripts/embed-worklet.mjs","prepublishOnly":"npm run clean && npm run build"},"_npmUser":{"name":"omar_e","email":"omar@attentionlabs.ai"},"_npmVersion":"10.9.3","description":"Browser JS SDK for the SD Attention Server (SAA). Streams microphone + webcam to the inference server and emits typed events for predictions, VAD, conversation state, and speech audio ready for downstream LLM use.","directories":{},"_nodeVersion":"22.19.0","_hasShrinkwrap":false,"devDependencies":{"typescript":"^5.4.0"},"_npmOperationalInternal":{"tmp":"tmp/saa-js_0.3.1_1779125627415_0.9539867411955065","host":"s3://npm-registry-packages-npm-production"}},"0.6.1":{"name":"@attenlabs/saa-js","version":"0.6.1","keywords":["attention","saa","audio","video","vad","realtime","websocket"],"author":{"name":"Attention Labs","email":"omar@attentionlabs.ai"},"license":"Apache-2.0","_id":"@attenlabs/saa-js@0.6.1","maintainers":[{"name":"omar_e","email":"omar@attentionlabs.ai"}],"homepage":"https://attentionlabs.ai","dist":{"shasum":"f989529ae21bac8f25011f12aacae4e9b9a367bc","tarball":"https://registry.npmjs.org/@attenlabs/saa-js/-/saa-js-0.6.1.tgz","fileCount":32,"integrity":"sha512-GC+teXtkBeupackkQa+IEF3VOZPe5hHiEOQmITpVHop1cl0HyA1qx+4FG1Njv53n9TLM7ss+FOrlXTzIbIXfCQ==","signatures":[{"sig":"MEUCICG7zpDHOGc0eG8Cp0eWmALfr1HUz8lWv2GMqWl5hM1dAiEA009njKMW03L5oFrKk9IQ7ZEODUCMxqc8MQzirRi6KVU=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":103442},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","module":"./dist/index.js","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"},"./audio-processor.js":"./dist/audio-processor.js"},"gitHead":"bb2dcf4da464612493ea2caae28310d843b2d4a4","scripts":{"test":"node --test test/*.test.mjs","build":"npx tsc && cp src/audio-processor.js dist/audio-processor.js","clean":"rm -rf dist","prebuild":"node scripts/embed-worklet.mjs","prepublishOnly":"npm run clean && npm run build"},"_npmUser":{"name":"omar_e","email":"omar@attentionlabs.ai"},"_npmVersion":"10.9.3","description":"Browser JS SDK for Attention Labs real-time attention detection. Streams microphone + webcam to the SAA inference server and emits typed events for predictions, VAD, conversation state, and speech audio ready for downstream LLM use.","directories":{},"_nodeVersion":"22.19.0","_hasShrinkwrap":false,"devDependencies":{"typescript":"^5.4.0"},"_npmOperationalInternal":{"tmp":"tmp/saa-js_0.6.1_1781907128886_0.5812428943824659","host":"s3://npm-registry-packages-npm-production"}},"0.7.0":{"name":"@attenlabs/saa-js","version":"0.7.0","keywords":["voice-agents","voice-ai","addressee-detection","device-directed-speech","selective-auditory-attention","turn-detection","speech-recognition","vad","wake-word","conversational-ai","real-time","llm","stt","javascript","browser","webrtc"],"author":{"name":"Attention Labs","email":"omar@attentionlabs.ai"},"license":"Apache-2.0","_id":"@attenlabs/saa-js@0.7.0","maintainers":[{"name":"omar_e","email":"omar@attentionlabs.ai"}],"homepage":"https://attentionlabs.ai","bugs":{"url":"https://github.com/attenlabs/saa-sdk/issues"},"dist":{"shasum":"084486b5b23cda615f72e57ad6ef0e9da6e77db6","tarball":"https://registry.npmjs.org/@attenlabs/saa-js/-/saa-js-0.7.0.tgz","fileCount":32,"integrity":"sha512-haLynQ6sON7sLcZ/05HwxQUhpZO/Ojaefo4Z0wKsjMAM92C41/CARo23aKB0PqVT2QJa+JgImS6eVbP3yV9ubg==","signatures":[{"sig":"MEUCIAp11gLk+195tI0mScJ5WoWGBdAkEJLbjY0Mvf0kOh8HAiEAuzVMbZ8CrmO/MRq6RjQZUDEkUNR7F1eBZPAwFmbAW+A=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":118486},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","module":"./dist/index.js","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"},"./audio-processor.js":"./dist/audio-processor.js"},"gitHead":"1fad65ec956e9566b7375b3aac3a21e83dd9541e","scripts":{"test":"node --test test/*.test.mjs","build":"npx tsc && cp src/audio-processor.js dist/audio-processor.js","clean":"rm -rf dist","prebuild":"node scripts/embed-worklet.mjs","prepublishOnly":"npm run clean && npm run build"},"_npmUser":{"name":"omar_e","email":"omar@attentionlabs.ai"},"changelog":"https://github.com/attenlabs/saa-sdk/blob/main/CHANGELOG.md","repository":{"url":"git+https://github.com/attenlabs/saa-sdk.git","type":"git","directory":"packages/saa-js"},"_npmVersion":"10.9.3","description":"Browser/JavaScript SDK for SAA, the addressee layer for voice agents — emits a turn_ready event per device-directed utterance so only speech meant for your agent reaches your STT, LLM, or TTS. No wake word.","directories":{},"_nodeVersion":"22.19.0","_hasShrinkwrap":false,"devDependencies":{"typescript":"^5.4.0"},"_npmOperationalInternal":{"tmp":"tmp/saa-js_0.7.0_1782309271819_0.42324632185500954","host":"s3://npm-registry-packages-npm-production"}},"0.7.2":{"name":"@attenlabs/saa-js","version":"0.7.2","keywords":["voice-agents","voice-ai","addressee-detection","device-directed-speech","selective-auditory-attention","turn-detection","speech-recognition","vad","wake-word","conversational-ai","real-time","llm","stt","javascript","browser","webrtc","livekit","pipecat","elevenlabs","twilio","openai-realtime","barge-in","pre-stt"],"author":{"name":"Attention Labs","email":"omar@attentionlabs.ai"},"license":"Apache-2.0","_id":"@attenlabs/saa-js@0.7.2","maintainers":[{"name":"omar_e","email":"omar@attentionlabs.ai"}],"homepage":"https://attentionlabs.ai","bugs":{"url":"https://github.com/attenlabs/saa-sdk/issues"},"dist":{"shasum":"db03e4d3cdb10ad247b144ea92eff7011c78ec5a","tarball":"https://registry.npmjs.org/@attenlabs/saa-js/-/saa-js-0.7.2.tgz","fileCount":32,"integrity":"sha512-IZ0fmvcfj+HWo2iob8MJFxkvkFz8cVjBPKs1e43/kHoAGAfdZI19wiMx+xyzV8nLua/swlysIx47QSgG+g/Rmg==","signatures":[{"sig":"MEUCIFIfasDpbsyiff2e8d9gY9agWHJbR1v0P+LbXg1Z6/iFAiEAxczLSClNYR5Vr+4St2fHaCL35VZHZ7V3Zs8EhcMP8ZA=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":130080},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","module":"./dist/index.js","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"},"./audio-processor.js":"./dist/audio-processor.js"},"gitHead":"6e5c1f32fd404251f8dec6197202b94567cdfc9e","scripts":{"test":"node --test test/*.test.mjs","build":"npx tsc && cp src/audio-processor.js dist/audio-processor.js","clean":"rm -rf dist","prebuild":"node scripts/embed-worklet.mjs","prepublishOnly":"npm run clean && npm run build"},"_npmUser":{"name":"omar_e","email":"omar@attentionlabs.ai"},"changelog":"https://github.com/attenlabs/saa-sdk/blob/main/CHANGELOG.md","repository":{"url":"git+https://github.com/attenlabs/saa-sdk.git","type":"git","directory":"packages/saa-js"},"_npmVersion":"10.9.3","description":"Browser/JavaScript SDK for SAA, the addressee layer for voice agents — emits a turn_ready event per device-directed utterance so only speech meant for your agent reaches your STT, LLM, or TTS. No wake word.","directories":{},"_nodeVersion":"22.19.0","_hasShrinkwrap":false,"devDependencies":{"typescript":"^5.4.0"},"_npmOperationalInternal":{"tmp":"tmp/saa-js_0.7.2_1784131888194_0.26747387819009405","host":"s3://npm-registry-packages-npm-production"}},"0.8.0":{"_id":"@attenlabs/saa-js@0.8.0","bugs":{"url":"https://github.com/attenlabs/saa-sdk/issues"},"dist":{"shasum":"44461f4454e1e2c8ed66577801876ff5c70f82ba","tarball":"https://registry.npmjs.org/@attenlabs/saa-js/-/saa-js-0.8.0.tgz","fileCount":32,"integrity":"sha512-7I34EsH7geVWrHauMTuW5SZyAC2TFZUeioGrCZ2ib5w9hgpdBvIi9gHih4zEIQBGhUor43/fuqrHdlliZfncMA==","signatures":[{"sig":"MEQCIGoygPghFAvoopzr7J/SWepGYPBeNrV2lEYUIPb8byNUAiAlWE0M1v153NJeRWdxsMervwi/L9pDQPvfRJy5DhLJSQ==","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"},{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIBNAXDQQPGZ6TiMdmvPsS+8QNjCZHYZ8YuL9hePyvU5zAiEA8jugGTb99sfMrquw3jMvoPjbOeZnfwZCywvM8zAFV3Y="}],"unpackedSize":142755},"main":"./dist/index.js","name":"@attenlabs/saa-js","type":"module","types":"./dist/index.d.ts","author":{"name":"Attention Labs","email":"omar@attentionlabs.ai"},"module":"./dist/index.js","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"},"./audio-processor.js":"./dist/audio-processor.js"},"gitHead":"ca8ea380bd89b258df96343bf72d36b6741095eb","license":"Apache-2.0","scripts":{"test":"node --test test/*.test.mjs","build":"npx tsc && cp src/audio-processor.js dist/audio-processor.js","clean":"rm -rf dist","prebuild":"node scripts/embed-worklet.mjs","prepublishOnly":"npm run clean && npm run build"},"version":"0.8.0","_npmUser":{"name":"omar_e","email":"omar@attentionlabs.ai"},"homepage":"https://attentionlabs.ai","keywords":["voice-agents","voice-ai","addressee-detection","device-directed-speech","selective-auditory-attention","turn-detection","speech-recognition","vad","wake-word","conversational-ai","real-time","llm","stt","javascript","browser","webrtc","livekit","pipecat","elevenlabs","twilio","openai-realtime","barge-in","pre-stt"],"changelog":"https://github.com/attenlabs/saa-sdk/blob/main/CHANGELOG.md","repository":{"url":"git+https://github.com/attenlabs/saa-sdk.git","type":"git","directory":"packages/saa-js"},"_npmVersion":"10.9.8","description":"Browser/JavaScript SDK for SAA, the addressee layer for voice agents — emits a turn_ready event per device-directed utterance so only speech meant for your agent reaches your STT, LLM, or TTS. No wake word.","directories":{},"maintainers":[{"name":"omar_e","email":"omar@attentionlabs.ai"}],"_nodeVersion":"22.23.1","_hasShrinkwrap":false,"devDependencies":{"typescript":"^5.4.0"},"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/saa-js_0.8.0_1790607754414_0.7270820045623765"}}},"time":{"created":"2026-05-14T15:11:31.002Z","modified":"2026-09-28T15:02:34.764Z","0.3.0":"2026-05-14T15:11:31.289Z","0.3.1":"2026-05-18T17:33:47.569Z","0.6.1":"2026-06-19T22:12:09.038Z","0.7.0":"2026-06-24T13:54:31.982Z","0.7.2":"2026-07-15T16:11:28.329Z","0.8.0":"2026-09-28T15:02:34.505Z"},"bugs":{"url":"https://github.com/attenlabs/saa-sdk/issues"},"author":{"name":"Attention Labs","email":"omar@attentionlabs.ai"},"license":"Apache-2.0","homepage":"https://attentionlabs.ai","keywords":["voice-agents","voice-ai","addressee-detection","device-directed-speech","selective-auditory-attention","turn-detection","speech-recognition","vad","wake-word","conversational-ai","real-time","llm","stt","javascript","browser","webrtc","livekit","pipecat","elevenlabs","twilio","openai-realtime","barge-in","pre-stt"],"repository":{"url":"git+https://github.com/attenlabs/saa-sdk.git","type":"git","directory":"packages/saa-js"},"description":"Browser/JavaScript SDK for SAA, the addressee layer for voice agents — emits a turn_ready event per device-directed utterance so only speech meant for your agent reaches your STT, LLM, or TTS. No wake word.","maintainers":[{"name":"omar_e","email":"omar@attentionlabs.ai"}],"readme":"# @attenlabs/saa-js\n\nJavaScript SDK for [Attention Labs](https://attentionlabs.ai) Selective Auditory Attention (SAA): the addressee layer for voice agents. One decision per utterance about whether speech was meant for your agent, before STT, LLM, or TTS. No wake word, model-agnostic, drop-in for the voice stack you already use.\n\nThe mental model is simple: audio in -> addressee gate (the SAA decision) -> only addressed audio out. The attention model runs on Attention Labs' service, so this is a thin Apache-2.0 client: it captures and streams your mic + camera and emits typed events. All inference runs server-side.\n\n## Sign up\n\nGet your API key at [attentionlabs.ai](https://attentionlabs.ai).\n\n## Install\n\n```bash\nnpm install @attenlabs/saa-js\n```\n\n## Quick start\n\n```ts\nimport { AttentionClient } from \"@attenlabs/saa-js\";\n\nconst videoEl = document.querySelector(\"video\");\n\nconst client = new AttentionClient({\n  token: \"your-api-key\",\n});\n\nclient.on(\"prediction\", ({ cls, confidence, source, numFaces }) => {\n  console.log(`${cls}: ${confidence.toFixed(2)}`);\n});\n\nclient.on(\"turnReady\", ({ audioBase64, durationSec }) => {\n  // Forward the captured turn to your LLM of choice\n});\n\nawait client.start({ videoElement: videoEl });\n```\n\n## Options\n\n| Option             | Type     | Default                              | Description |\n| ------------------ | -------- | ------------------------------------ | ----------- |\n| `token`            | string   | none                                    | Your API key from attentionlabs.ai. |\n| `initialThreshold` | number   | `0.7`                                | Confidence threshold for predictions (0-1). |\n| `enableAudio`      | boolean  | `true`                               | Capture the mic internally. Set `false` to push audio via `feedAudio()`. |\n| `enableVideo`      | boolean  | `true`                               | Capture the camera internally. Set `false` for audio-only or to push frames via `feedVideo()`. |\n| `autoReconnect`    | boolean  | `true`                               | Reconnect with backoff after an unclean mid-session drop. Set `false` to surface the drop as an `error` instead. |\n| `serverProfile`    | string   | inferred                             | Server processor variant. Defaults to `\"audio_only\"` when `enableVideo: false`, else the full processor. Pass `\"default\"` to force the full processor without local video. |\n| `utteranceHandling` | boolean | `false`                              | Opt into utterance handling: per-utterance transcripts with an addressee verdict, delivered as `utteranceEnded`. See [Utterance handling](#utterance-handling). |\n| `workletUrl`       | string   | bundled                              | URL of the audio-capture AudioWorklet module. Override only when self-hosting the worklet. |\n| `video.width`      | number   | `1920`                               | Capture width. |\n| `video.height`     | number   | `1080`                               | Capture height. |\n| `video.jpegQuality`| number   | `0.5`                                | JPEG quality (0-1). |\n| `audio.targetSampleRate` | number | `16000`                         | Sample rate audio is resampled to before sending. |\n| `audio.onAudioFrame`| function | none                                | Called with each captured 16-bit PCM frame (`ArrayBuffer`). |\n| `audio.onWorkletError`| function | none                              | Called when the capture worklet throws (also emitted as an `error` event). |\n| `audio.onContextStateChange`| function | none                        | Called with the `AudioContext` state string on change (e.g. `suspended`, `interrupted`). |\n\n## Methods\n\n| Method                      | Description |\n| --------------------------- | ----------- |\n| `start({ videoElement, mediaStream? })` | Start streaming + connect. Calls `getUserMedia` unless `mediaStream` is supplied. `videoElement` is required when video capture is enabled. |\n| `stop()`                    | Stop streaming and disconnect. |\n| `feedAudio(audio, sampleRate?)` | Push externally-captured audio (requires `enableAudio: false`). Accepts Float32 `[-1,1]`, Int16 PCM, or a raw int16 buffer; re-chunked + resampled to the wire's 16 kHz / 100 ms blocks. See [External capture](#external-capture). |\n| `feedVideo(jpeg)`           | Push an externally-captured JPEG frame (requires `enableVideo: false`). Accepts a `Blob`, `ArrayBuffer`, or view. |\n| `mute()` / `unmute()`       | Pause or resume audio. |\n| `markResponding(boolean)`   | Signal that your app is responding, pauses predictions until finished. |\n| `setThreshold(value)`       | Update the confidence threshold (0-1). |\n| `addAssistantTurn(text)`    | Utterance handling: feed back what the assistant said, as spoken. Returns `false` when the socket is not open. |\n| `setUtteranceThreshold(value)` | Utterance handling: the one-sided class-1 decision threshold (0-1]. Server acks via `utteranceConfig`. |\n| `clearUtteranceHistory()`   | Utterance handling: forget the dialogue history (new conversation). |\n| `isConnected`               | Getter — `true` while the WebSocket is open. |\n| `currentThreshold`          | Getter — the current confidence threshold (0-1). |\n| `on(event, listener)`       | Subscribe to an event. Returns an unsubscribe function. |\n\n## Events\n\n| Event            | Payload |\n| ---------------- | ------- |\n| `connected`      | none |\n| `started`        | none |\n| `warmupComplete` | none |\n| `prediction`     | `{ cls, rawCls, confidence, source, numFaces, responding }` |\n| `vad`            | `{ probability, isSpeech }` |\n| `state`          | `{ state }` (one of `listening`, `sending`, `cancelled`, `idle`) |\n| `turnReady`      | `{ audioBase64, audioPcm16, durationSec, frames, context }` |\n| `config`         | `{ modelClass2Threshold }` |\n| `stats`          | `{ rttMs, bufferedAmount, sentVideo, skippedVideo, sentAudio, uptimeMs }` |\n| `interrupt`      | `{ fadeMs, confidence }` |\n| `interjection`   | `{ reason, audioBase64, audioPcm16, durationSec }` |\n| `utteranceEnded` | `{ seq, text, prediction, confidence, decision, reason, startS, endS, truncated, assistantTurns, preview, latencyMs, audioBase64, audioPcm16 }` |\n| `utteranceConfig` | `{ enabled, class1Threshold, preview, reason }` |\n| `error`          | `{ title, message, detail }` |\n| `disconnected`   | `{ code, reason, wasClean }` |\n| `reconnecting`   | `{ attempt, delaySec, lastCode }` |\n| `reconnected`    | `{ attempts }` |\n\n`warmupComplete` fires once the server model has warmed up and is producing real predictions; use it to drop any loading UI. `prediction.responding` is `true` while your app is mid-response (see `markResponding`), and `interjection` fires when the agent should volunteer after humans go quiet.\n\nIf the camera is unavailable when video capture is enabled and audio is enabled, `start()` continues with an audio-only session (`audio_only` server profile), emits an `error` with `kind: \"environment\"` and `title: \"Camera unavailable\"`. The original `enableVideo` request is restored on the next `start()`, so a later session retries video.\n\n## Utterance handling\n\nOpt in with `utteranceHandling: true`. The server then segments speech into utterances on voice\nactivity, transcribes each one when it ends, and scores whether it was aimed at the device\n(`prediction: 2`) or at a person (`prediction: 1`). One `utteranceEnded` event per utterance\ncarries the text, the verdict and the utterance audio; `utteranceConfig` arrives after\n`started` and says whether the feature is on for this session.\n\nThe classifier is conditioned on the preceding dialogue and is unreliable without it, so feed\nback every assistant response, as spoken, with `addAssistantTurn(text)`. `assistantTurns` on the\nevent shows how many assistant lines were in the scored history; `0` means you have not fed any.\n\n```js\nconst client = new AttentionClient({ token, utteranceHandling: true });\n\nclient.on(\"utteranceEnded\", (u) => {\n  console.log(u.text, u.prediction, u.confidence, u.decision, u.preview);\n});\n\n// after each assistant response has been spoken\nclient.addAssistantTurn(assistantTranscript);\n```\n\nThis stream is independent of `turnReady`: the two will not line up one to one (a turn can hold\ntwo utterances, an utterance can straddle a class change). Drive your LLM from one or the\nother. While the feature is in preview, `preview` is `true` and the verdict is preview grade; the server delivers transcripts and does not retain them.\n\n## LLM integration\n\nThe SDK captures speech but does **not** route it to an LLM. SAA is model-agnostic and drop-in: use the `turnReady` event to forward only device-directed audio to any model, ASR, or voice stack you like.\n\nWhen your LLM starts responding, call `client.mute()` and `client.markResponding(true)`. When it finishes, call `client.unmute()` and `client.markResponding(false)`.\n\n## External capture\n\nBy default the SDK opens its own mic + camera. To run on capture you already\nown there are two paths:\n\n**Share a `MediaStream`** (the SDK reads it but won't stop its tracks):\n\n```ts\nconst stream = await navigator.mediaDevices.getUserMedia({ video: true, audio: true });\nvideoEl.srcObject = stream;                 // your app renders it\nawait client.start({ videoElement: videoEl, mediaStream: stream });\n// ... another consumer (e.g. a gaze SDK) reads the same stream / videoEl\n```\n\n**Push frames yourself** (no `getUserMedia` at all) for taps, Twilio media,\nor non-browser sources:\n\n```ts\nconst client = new AttentionClient({ token, enableAudio: false, enableVideo: false });\nawait client.start();                        // opens the WS, captures nothing\nclient.feedAudio(pcmChunk);                  // Float32 [-1,1] | Int16 | int16 buffer\nclient.feedAudio(pcm48k, 48000);             // resampled to 16 kHz\nclient.feedVideo(jpegBlob);                  // Blob | ArrayBuffer | view\n```\n\nMix and match: `enableVideo: false` with internal mic for audio-only, or\n`enableAudio: false` + `feedAudio()` while the SDK still grabs camera frames.\n\n## License\n\nApache-2.0\n","readmeFilename":"README.md"}