{"_id":"@callmyai/ai","name":"@callmyai/ai","dist-tags":{"latest":"1.1.0"},"versions":{"1.1.0":{"name":"@callmyai/ai","version":"1.1.0","main":"./dist/index.js","module":"./dist/index.mjs","types":"./dist/index.d.ts","sideEffects":false,"devDependencies":{"@jest/globals":"^29.7.0","@types/jest":"^29.5.12","@types/react":"^18.2.61","@types/react-dom":"^18.2.19","eslint":"^8.57.0","jest":"^29.7.0","jest-fetch-mock":"^3.0.3","react":"^18.2.0","ts-jest":"^29.1.2","tsup":"^8.0.2","typescript":"^5.4.5","@callmyai/eslint-config":"0.0.0","@callmyai/tsconfig":"0.0.0"},"publishConfig":{"access":"public"},"dependencies":{"realtime-ai":"^0.0.9","realtime-ai-react":"^0.0.9"},"exports":{"./ui":{"types":"./ui/dist/index.d.ts","import":"./ui/dist/index.mjs","require":"./ui/dist/index.js"}},"scripts":{"build":"tsup --format esm,cjs --dts --external react","dev":"tsup --format esm,cjs --watch --dts --external react","test":"jest","lint":"eslint \"src/**/*.ts*\"","clean":"rm -rf .turbo && rm -rf node_modules && rm -rf dist"},"_id":"@callmyai/ai@1.1.0","description":"## Description Inspired by Vercel's Language Model Specification, this is a proposal for introducing a Speech Model Specification to streamline the integration of various speech providers into our platform. This specification aims to provide a standardize","_integrity":"sha512-z7D3Zql9VdvlBWtUxU+PatGl6zSqil8bVy4xUMR3tcQMe0DWLiz7v3KLapwvxA1gxupG3Q4QfaAbmb4xDFlOKw==","_resolved":"/tmp/b0f2708825790adbcd450a5f29742383/callmyai-ai-1.1.0.tgz","_from":"file:callmyai-ai-1.1.0.tgz","_nodeVersion":"20.17.0","_npmVersion":"10.8.2","dist":{"integrity":"sha512-z7D3Zql9VdvlBWtUxU+PatGl6zSqil8bVy4xUMR3tcQMe0DWLiz7v3KLapwvxA1gxupG3Q4QfaAbmb4xDFlOKw==","shasum":"a4db1f966c811d5064a0fc3920708b6776ee33f9","tarball":"https://registry.npmjs.org/@callmyai/ai/-/ai-1.1.0.tgz","fileCount":3,"unpackedSize":41426,"signatures":[{"keyid":"SHA256:jl3bwswu80PjjokCgh0o2w5c2U4LhQAE57gj9cz1kzA","sig":"MEQCIEQUjp1m75f24lINlnJk9Qu2lWdq5Ug/B+huFXnule6sAiAOWtEOAOCflc5C8Un9LIgyxf9oTKSstyqV/L/ooY1aSg=="}]},"_npmUser":{"name":"bishwenduk029","email":"bishwenduk029@gmail.com"},"directories":{},"maintainers":[{"name":"bishwenduk029","email":"bishwenduk029@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages","tmp":"tmp/ai_1.1.0_1725164556021_0.713441465242658"},"_hasShrinkwrap":false}},"time":{"created":"2024-09-01T04:22:35.921Z","1.1.0":"2024-09-01T04:22:36.163Z","modified":"2024-09-01T04:22:36.460Z"},"maintainers":[{"name":"bishwenduk029","email":"bishwenduk029@gmail.com"}],"description":"## Description Inspired by Vercel's Language Model Specification, this is a proposal for introducing a Speech Model Specification to streamline the integration of various speech providers into our platform. This specification aims to provide a standardize","readme":"# AI-Voice\n\n## Description\nInspired by Vercel's Language Model Specification, this is a proposal for introducing a Speech Model Specification to streamline the integration of various speech providers into our platform. This specification aims to provide a standardized interface for interacting with different speech models, eliminating the complexity of dealing with unique APIs and reducing the risk of vendor lock-in.\n\n## Problem Statement\nCurrently, there are numerous speech providers available, each with its own distinct method for interfacing with their models. This lack of standardization complicates the process of switching providers and increases the likelihood of vendor lock-in. Developers are required to learn and implement different APIs for each provider, leading to increased development time and maintenance overhead.\n\nThe open-source community has created the following providers:\n\n- OpenAI Provider (@bishwenduk029-ai-voice/openai)\n\n- Elevenlabs Provider (@bishwenduk029-ai-voice/elevenlabs)\n\n- Deepgram Provider (@bishwenduk029-ai-voice/deepgram)\n\n- PlayHt Provider (@bishwenduk029-ai-voice/playht)\n\n## Features\n\n1. React Hook `useVoiceChat` for easy transcription and streaming voice chat in client side.\n2. Consistent API for streaming speech response from various AI speech providers.\n3. Works well with `streamText` from [Vercel AI SDK](https://sdk.vercel.ai/docs/ai-sdk-core/providers-and-models).\n\n## Installation\n\n```sh\npnpm install @callmyai/ai\n```\n![alt text](spec_v1.png)\n\n## Usage\n\n### Client Side React Hook\n🎙️ Real-Time Speech Transcription React Hook\nSeamlessly integrate real-time speech transcription into your React applications with this powerful and efficient hook! 🚀\n\n✨ Hook Capabilities\n\n- 🎤 Detect human speech end or silence using the robust @ricky0123/vad-react library\n- ⏱️ Intelligently debounce speech input, ensuring continuous recording and transcription as long as the user speaks within a configurable time frame (e.g., 500ms)\n- 🗣️ Gracefully handle speech interruptions, allowing users to pause and resume speaking naturally\n- 🌐 Efficiently trigger REST calls for transcription, optimizing performance by waiting for the user to pause before sending requests\n- 🔌 Easy to integrate into your existing React projects, with a simple and intuitive API\n```ts\n...\nimport { useVoiceChat } from '@callmyai/ai/ui'\n...\n\nexport default function Chat({ id, initialMessages, className }: ChatProps) {\n  const { speaking, listening, thinking, initialized, messages, setMessages } =\n    useVoiceChat({\n      api: '/api/chat/voice',\n      initialMessages,\n      transcribeAPI: '/api/transcribe',\n      body: {\n        id\n      },\n      speakerPause: 500,\n      onSpeechCompletion: async () => {\n        if (id) {\n          const chat = await getChat(id)\n          setMessages(chat?.messages || [])\n        }\n      }\n    })\n\n  return (\n    <>\n      <div className={cn('pb-[200px] pt-4 md:pt-10', className)}>\n        {messages.length ? (\n          <>\n            <ChatList messages={messages} />\n            <ChatScrollAnchor trackVisibility={initialized} />\n          </>\n        ) : (\n          <EmptyScreen />\n        )}\n      </div>\n      <ChatPanel\n        id={id}\n        initialized={initialized}\n        speaking={speaking}\n        listening={listening}\n        thinking={thinking}\n        messages={messages}\n      />\n    </>\n  )\n}\n```\n\n### Hook up Server APIs\n`/api/chat/voice`\n```ts\nimport 'server-only'\n...\nimport { streamText } from 'ai'\nimport { ollama, createOllama } from 'ollama-ai-provider'\nimport { openaiSpeech, playhtSpeech, streamSpeech } from '@callmyai/ai/server'\n...\n\nexport const runtime = 'edge'\n\nconst model = ollama('llama3:latest')\n\nexport async function POST(req: Request) {\n  const cookieStore = cookies()\n  const supabase = createRouteHandlerClient<Database>({\n    cookies: () => cookieStore\n  })\n  const json = await req.json()\n  const { messages } = json\n  // const userId = (await auth({ cookieStore }))?.user.id\n\n  // if (!userId) {\n  //   return new Response('Unauthorized', {\n  //     status: 401\n  //   })\n  // }\n\n  const systemPrompt = `\n  Your role is to act as a friendly human assistant by the user preferred name. Your given name is Nova.\n  `\n\n  const result = await streamText({\n    model,\n    messages: [\n      {\n        role: 'system',\n        content: `${systemPrompt}`\n      },\n      ...messages.map((message: { role: any; content: any }) => ({\n        role: message.role,\n        content: message.content\n      }))\n    ]\n  })\n\n  // OpenAI - env:OPENAI_API_KEY\n  const speechModel = openaiSpeech(\n    'tts-1',   //openai_speech_model\n    'nova'     //openai_voice_id\n  )\n\n  // ElevenLabsIO - env:ELEVENLABS_API_KEY\n  // const speechModel = elevenlabsSpeech(\n  //   'eleven_turbo_v2',   //elevenlabs_speech_model\n  //   'DIBkDE5u33APYlfhjihh' //elevenlabs_voice_id\n  // )\n\n    // PlayHt - env:PLAYHT_API_KEY\n  // const speechModel = playhtSpeech(\n  //   'PlayHT2.0-turbo', //playht_speech_model\n  //   '<your-playht-user-id>',  \n  //   \"s3://voice-cloning-zero-shot/1afba232-fae0-4b69-9675-7f1aac69349f/delilahsaad/manifest.json\"   //playht_voice_id\n  // )\n\n  // Deepgram - env:DEEPGRAM_API_KEY\n  // const speechModel = deepgramSpeech(\"aura-asteria-en\")\n\n  try {\n    const speech = await streamSpeech(speechModel)(result.textStream)\n    return new Response(speech, {\n      headers: { 'Content-Type': 'audio/mpeg' }\n    })\n  } catch (error) {\n    console.log(error)\n    return new Response(null, {\n      status: 500\n    })\n  }\n}\n```\n\n## Limitations\nCurrently the speech streaming only works for english text streams. Multi-lingual support in future.\n\n## Roadmap\n- [x] Speech Model Specification Done\n- [ ] Improve the implementation of sentence-boundary detection algorithm for the text stream to sentence stream conversion\n- [ ] Add more tests\n- [ ] Enhance the specification to also take case of WebSocket-based Speech Providers","readmeFilename":"README.md"}