{"_id":"@averatec0773/openclaw-fishaudio","name":"@averatec0773/openclaw-fishaudio","dist-tags":{"latest":"0.1.0"},"versions":{"0.1.0":{"name":"@averatec0773/openclaw-fishaudio","version":"0.1.0","type":"module","license":"MIT","author":{"name":"averatec0773"},"description":"Fish Audio low-latency speech provider for OpenClaw, using WebSocket TTS Live for real-time Discord voice channel conversation.","repository":{"type":"git","url":"git+https://github.com/averatec0773/openclaw-fishaudio.git"},"homepage":"https://github.com/averatec0773/openclaw-fishaudio#readme","bugs":{"url":"https://github.com/averatec0773/openclaw-fishaudio/issues"},"keywords":["openclaw","openclaw-plugin","fish-audio","fish-audio-tts","tts","speech-provider","realtime","discord","voice-channel","websocket"],"openclaw":{"extensions":["./dist/index.js"],"compat":{"pluginApi":">=2026.5.3-beta.2","minGatewayVersion":"2026.5.3"}},"scripts":{"test":"vitest run","test:watch":"vitest","build":"esbuild src/index.ts src/modes/speech-provider.ts src/modes/realtime-bridge.ts src/fish-audio/config.ts src/fish-audio/http-fallback.ts src/fish-audio/websocket-live.ts src/fish-audio/voice-list.ts src/fish-audio/types.ts --bundle=false --platform=node --format=esm --target=node22 --outdir=dist --out-extension:.js=.js","prepack":"npm run build"},"dependencies":{"ws":"^8.18.0"},"devDependencies":{"openclaw":"^2026.5.3-beta.2","@types/node":"^22.0.0","@types/ws":"^8.5.0","typescript":"^5.6.0","vitest":"^3.1.0","esbuild":"^0.28.0"},"engines":{"node":">=22"},"gitHead":"6d476583a152ec380acef3a613bed139b5a86acf","_id":"@averatec0773/openclaw-fishaudio@0.1.0","_nodeVersion":"25.6.0","_npmVersion":"11.8.0","dist":{"integrity":"sha512-+XPGQOLg7+9zDU9wBoxGezpzHNyTd+H+bIY6RS0Y4uh21olyu3pkuuURpxqDEI7zwT5DepUK9rdPMI9NzeTY7w==","shasum":"d29b5ee61e8b36d972e6acec28005347cd6b342a","tarball":"https://registry.npmjs.org/@averatec0773/openclaw-fishaudio/-/openclaw-fishaudio-0.1.0.tgz","fileCount":21,"unpackedSize":47189,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQD3fS9Jlx4m+DJtbpwtsI2e0u8cGCIsVHJyjwCtcAagIAIgTnAnhIvbngDsHWTEUv1lVzpHkW13fwiL9PWMizoIkwg="}]},"_npmUser":{"name":"averatec0773","email":"ayetek0773@gmail.com"},"directories":{},"maintainers":[{"name":"averatec0773","email":"ayetek0773@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/openclaw-fishaudio_0.1.0_1778104938256_0.7020347362282606"},"_hasShrinkwrap":false}},"time":{"created":"2026-05-06T22:02:18.137Z","0.1.0":"2026-05-06T22:02:18.401Z","modified":"2026-05-06T22:02:18.599Z"},"maintainers":[{"name":"averatec0773","email":"ayetek0773@gmail.com"}],"description":"Fish Audio low-latency speech provider for OpenClaw, using WebSocket TTS Live for real-time Discord voice channel conversation.","homepage":"https://github.com/averatec0773/openclaw-fishaudio#readme","keywords":["openclaw","openclaw-plugin","fish-audio","fish-audio-tts","tts","speech-provider","realtime","discord","voice-channel","websocket"],"repository":{"type":"git","url":"git+https://github.com/averatec0773/openclaw-fishaudio.git"},"author":{"name":"averatec0773"},"bugs":{"url":"https://github.com/averatec0773/openclaw-fishaudio/issues"},"license":"MIT","readme":"# openclaw-fishaudio\n\nA [Fish Audio](https://fish.audio) speech provider plugin for [OpenClaw](https://openclaw.ai). The plugin registers as a `SpeechProvider` and uses the [WebSocket TTS Live](https://docs.fish.audio/api-reference/endpoint/websocket/tts-live) endpoint for synthesis, with the HTTP `/v1/tts` endpoint as a fallback. The primary use case is real-time conversation in Discord voice channels; the plugin is also usable for any other target OpenClaw routes through a `SpeechProvider`, such as Telegram or WhatsApp voice notes.\n\nThe plugin id is `fishaudio`.\n\n## Features\n\n| Feature | |\n|---|---|\n| HTTP `/v1/tts` synthesis | ✅ |\n| WebSocket `/v1/tts/live` transport | ✅ |\n| Automatic WebSocket → HTTP fallback (`transport: \"auto\"`) | ✅ |\n| Direct Opus output for the Discord voice channel target (`audio-file`) | ✅ |\n| Direct Opus output for the voice-note target (Telegram, WhatsApp) | ✅ |\n| Voice list combining the account's own voice clones with a page of popular community voices | ✅ |\n| Inline directives: `fishaudio_voice` / `fish_speed` / `fish_model` / `fish_latency` / `fish_temperature` / `fish_top_p` | ✅ |\n| Unit tests | 45 |\n\nOut of scope: STT, voice cloning upload, barge-in, and the OpenClaw `realtime` Talk mode (the latter bypasses speech providers entirely).\n\n## Install\n\n```bash\nopenclaw plugins install @averatec0773/openclaw-fishaudio\n```\n\nObtain an API key at [fish.audio](https://fish.audio) under Account → API. Restart OpenClaw after installation.\n\n## Configure\n\nMinimum configuration, sufficient for voice-note and chat TTS targets:\n\n```json5\n{\n  messages: {\n    tts: {\n      provider: \"fishaudio\",\n      providers: {\n        fishaudio: {\n          apiKey: \"your-fish-audio-api-key\",   // or set FISH_AUDIO_API_KEY in the environment\n          voiceId: \"reference-id-of-your-voice\",\n          model: \"s2-pro\",                      // s2-pro (default) | s1\n          latency: \"low\",                       // low (default) | balanced | normal\n          transport: \"auto\"                     // auto | websocket | http\n          // speed, temperature, topP are also accepted\n        }\n      }\n    }\n  }\n}\n```\n\n### Discord voice channel — three additional blocks\n\n`/vc join` produces no audio if any of the following blocks is missing. The requirements come from upstream `@openclaw/discord` and the bundled `talk-voice` plugin:\n\n```json5\n{\n  channels: {                              // 1. Per-bot voice and TTS routing\n    discord: {\n      accounts: {\n        \"<your-bot-account-id>\": {\n          voice: {\n            enabled: true,                 // gates /vc commands and the voice intent\n            tts: { provider: \"fishaudio\", auto: \"inbound\" }\n          }\n        }\n      }\n    }\n  },\n\n  talk: {                                  // 2. Read by the talk-voice plugin\n    provider: \"fishaudio\",\n    providers: {\n      fishaudio: {\n        voiceId: \"your-fish-voice-id\",\n        model: \"s2-pro\",\n        latency: \"low\",\n        transport: \"auto\"\n      }\n    },\n    speechLocale: \"zh-CN\",\n    interruptOnSpeech: false               // see operational notes below\n  },\n\n  tools: {                                 // 3. STT (input side)\n    media: {\n      audio: {\n        enabled: true,\n        models: [{ provider: \"openai\", model: \"gpt-4o-transcribe\" }]\n      }\n    }\n  }\n}\n```\n\n### Operational notes\n\n- Use headphones during testing. Discord does not perform aggressive echo cancellation on the bot side, so playback re-entering through a microphone will be transcribed and processed as a new turn.\n- Verbose logging is enabled by the gateway CLI flag `--verbose`. The environment variable `OPENCLAW_VERBOSE=1` does not enable it. To turn verbose on under Docker Compose, append `\"--verbose\"` to the gateway `command:` array.\n- The OpenAI model `gpt-4o-mini-transcribe` returns a 200 response with an empty `text` field when called through OpenClaw's transcription path. Use `gpt-4o-transcribe` in `tools.media.audio.models` until the upstream call shape is updated.\n- The bot leaves any voice channel on each gateway restart; rejoining requires another `/vc join`. To rejoin automatically, set `channels.discord.accounts.<id>.voice.autoJoin` to a list of `{ guildId, channelId }` entries.\n\n## End-to-end timing\n\nPer-stage timings observed on a representative Discord voice channel turn (`gpt-4o-transcribe` for STT, `gpt-5-nano` with `tools.allow=[\"message\"]` for the LLM, Fish Audio `s2-pro` voice clone, eastern-US VPS, reply ≈ 14 characters):\n\n| Stage | Observed | Driver |\n|---|---|---|\n| Discord `speaking_end` grace | ~1.2s | Hardcoded by the upstream Discord plugin; not configurable. |\n| OPUS decode plus WAV write | ~0.2s | Local I/O. |\n| STT | 1.5–2.5s | OpenAI transcription TTFT. |\n| LLM | 6–8s | OpenAI TTFT for `gpt-5-nano`. Switching to `claude-haiku-4-5` or `cerebras/llama-3.3-70b` reduces this stage. |\n| TTS (this plugin) | ~1.0s | Fish Audio API plus buffer write. |\n| Playback start | ~0.1s | Local. |\n| End-to-end (stop-talking → audio plays) | ~10–13s | Currently dominated by the LLM stage. |\n\nA direct `curl` to Fish Audio for the same input completes in roughly 0.8 seconds, matching the TTS stage above. Going materially below ~3 seconds end-to-end requires OpenClaw's `realtime` Talk mode, which the Discord channel plugin does not currently route through.\n\n## Inline directives\n\nProvider-prefixed directive keys are supported, with both `fishaudio_*` and `fish_*` aliases:\n\n```text\n[[tts:fish_voice=<ref_id>]]         Switch voice\n[[tts:fish_speed=1.2]]              Prosody speed (0.5–2.0)\n[[tts:fish_model=s1]]               s2-pro | s1\n[[tts:fish_latency=balanced]]       low | balanced | normal\n[[tts:fish_temperature=0.7]]        0–1\n[[tts:fish_top_p=0.8]]              0–1\n```\n\n## Voice list\n\n```bash\nopenclaw /voice list\n```\n\nReturns the account's own voice clones first, followed by a page of popular community voices.\n\n## Verification\n\nA manual Discord voice channel acceptance procedure is documented at [`docs/manual-e2e.md`](docs/manual-e2e.md).\n\n## License\n\nMIT\n","readmeFilename":"README.md","_rev":"1-8f0b1bfb0f5ee96029db456be88c5d07"}