{"_id":"@autocode2/speech-to-text","_rev":"1-006ac1b716c81a4c035251af77a4e4c8","name":"@autocode2/speech-to-text","dist-tags":{"latest":"0.1.1"},"versions":{"0.1.0":{"name":"@autocode2/speech-to-text","version":"0.1.0","author":{"name":"Gareth Andrew"},"license":"ISC","_id":"@autocode2/speech-to-text@0.1.0","maintainers":[{"name":"gandrew","email":"gingerhendrix@gmail.com"}],"bin":{"speech-to-text":"dist/cli.js"},"dist":{"shasum":"8090d1917c1d7d9e1c9723e281ff278bb4be08db","tarball":"https://registry.npmjs.org/@autocode2/speech-to-text/-/speech-to-text-0.1.0.tgz","fileCount":15,"integrity":"sha512-QfKrLLrbhlLoqIzKYJAiU8pekvE4bsLp1gDJWgIAv93Cw4Mhk7OrmZzkFdVrRXEq/Oc4NXYtUAt7MoeA1Maa3g==","signatures":[{"sig":"MEUCIQCV+pKvTiN2QVJQZ4yHTIkAa06pdlifSpoD/RJvb67xpAIgSYVIzbO6Ss3JSHpnEf7ygz5UYJdIVDFDNsMOJfq9KyI=","keyid":"SHA256:jl3bwswu80PjjokCgh0o2w5c2U4LhQAE57gj9cz1kzA"}],"unpackedSize":25883},"main":"dist/index.js","type":"module","gitHead":"4d6b41b49e24bfd787e78e109b567de929ba0d64","scripts":{"test":"echo \"Error: no test specified\" && exit 1","build":"tsc","start":"tsx src/cli.ts","prettier":"prettier --write ."},"_npmUser":{"name":"gandrew","email":"gingerhendrix@gmail.com"},"_npmVersion":"10.7.0","description":"A speech-to-text library for node","directories":{},"_nodeVersion":"22.1.0","dependencies":{"yargs":"^17.7.2","@google/generative-ai":"^0.21.0"},"publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"tsx":"^4.19.2","prettier":"^3.3.3","typescript":"^5.6.3","@types/node":"^22.9.0","@types/yargs":"^17.0.33"},"_npmOperationalInternal":{"tmp":"tmp/speech-to-text_0.1.0_1731881125609_0.17392207549789207","host":"s3://npm-registry-packages"}},"0.1.1":{"name":"@autocode2/speech-to-text","version":"0.1.1","description":"A speech-to-text library for node","main":"dist/index.js","type":"module","repository":{"type":"git","url":"git+https://github.com/autocode2/speech-to-text.git"},"keywords":["audio","speech","speech-to-text","cli","sox","google-gemini"],"homepage":"https://github.com/autocode2/speech-to-text","publishConfig":{"access":"public"},"bin":{"speech-to-text":"dist/cli.js"},"scripts":{"build":"tsc","start":"tsx src/cli.ts","test":"echo \"Error: no test specified\" && exit 1","prettier":"prettier --write ."},"author":{"name":"Gareth Andrew"},"license":"ISC","devDependencies":{"@types/node":"^22.9.0","@types/yargs":"^17.0.33","prettier":"^3.3.3","tsx":"^4.19.2","typescript":"^5.6.3"},"dependencies":{"@google/generative-ai":"^0.21.0","yargs":"^17.7.2"},"_id":"@autocode2/speech-to-text@0.1.1","gitHead":"3c6e51af78aa085163d77966a9020396a8808d59","bugs":{"url":"https://github.com/autocode2/speech-to-text/issues"},"_nodeVersion":"22.1.0","_npmVersion":"10.7.0","dist":{"integrity":"sha512-Kcdy91P3NucvppC8y4RKn0waOGe87JbuW0MUVmbsNCtnBR/biSINAyalmX/Y5QlsLQM4ZZVLSOXo66mQTTyUQQ==","shasum":"f3bb683f4740fa7a2122ed5161ebc5028c4edf15","tarball":"https://registry.npmjs.org/@autocode2/speech-to-text/-/speech-to-text-0.1.1.tgz","fileCount":15,"unpackedSize":26100,"signatures":[{"keyid":"SHA256:jl3bwswu80PjjokCgh0o2w5c2U4LhQAE57gj9cz1kzA","sig":"MEUCIQCLzh2WVPjikovZDfwjBGcOtxtdxcrYpY9oxHjMtwQI/wIgLYv50DLK9n871j43FMSjWStl+AQ6Bq035jY1+1xnFWs="}]},"_npmUser":{"name":"gandrew","email":"gingerhendrix@gmail.com"},"directories":{},"maintainers":[{"name":"gandrew","email":"gingerhendrix@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages","tmp":"tmp/speech-to-text_0.1.1_1731881983806_0.8234060270315249"},"_hasShrinkwrap":false}},"time":{"created":"2024-11-17T22:05:25.508Z","modified":"2024-11-17T22:19:44.177Z","0.1.0":"2024-11-17T22:05:25.810Z","0.1.1":"2024-11-17T22:19:43.992Z"},"author":{"name":"Gareth Andrew"},"license":"ISC","description":"A speech-to-text library for node","maintainers":[{"name":"gandrew","email":"gingerhendrix@gmail.com"}],"readme":"# @autocode2/speech-to-text\n\nA Node.js library and CLI tool for converting speech to text using sox for audio recording and Google's Gemini API for transcription.\n\n## Prerequisites\n\n- Node.js 18 or later\n- `sox` command line utility installed on your system\n  - macOS: `brew install sox`\n  - Linux: `apt-get install sox`\n  - Windows: Download from [Sox website](https://sourceforge.net/projects/sox/)\n- Google API key with access to Gemini API\n\n## Quick Start\n\nThe quickest way to use the tool is via `npx`:\n\n```bash\nnpx @autocode2/speech-to-text --api-key YOUR_API_KEY\n```\n\n## Installation\n\n### Global Installation\n\nIf you plan to use the tool frequently, you can install it globally:\n\n```bash\nnpm install -g @autocode2/speech-to-text\n```\n\nThen use it directly:\n\n```bash\nspeech-to-text --api-key YOUR_API_KEY\n```\n\n### Local Installation\n\nFor use in a project:\n\n```bash\nnpm install @autocode2/speech-to-text\n```\n\n## CLI Usage\n\n```bash\nnpx @autocode2/speech-to-text --api-key YOUR_API_KEY [options]\n```\n\n### Options\n\n- `-k, --api-key`: Google API Key for Gemini (required)\n- `-i, --input`: Input audio file to transcribe (if not provided, will record from microphone)\n- `-o, --output`: Output file to save the recording (only applies when recording from microphone)\n- `-r, --sample-rate`: Sample rate for recording in Hz (default: 16000)\n- `-c, --channels`: Number of audio channels (default: 1)\n- `-m, --model`: Gemini model to use (default: \"gemini-1.5-flash\")\n- `-p, --prompt`: Custom prompt for transcription\n- `-f, --format`: Output format (text|json, defaults to text in terminal, json in pipe)\n- `-h, --help`: Show help\n- `-v, --version`: Show version number\n\n### Examples\n\n```bash\n# Record from microphone and transcribe (uses temporary file)\nnpx @autocode2/speech-to-text --api-key YOUR_API_KEY\n\n# Record, save to file, and transcribe\nnpx @autocode2/speech-to-text --api-key YOUR_API_KEY -o recording.wav\n\n# Transcribe existing file\nnpx @autocode2/speech-to-text --api-key YOUR_API_KEY -i existing.wav\n\n# Record in high quality\nnpx @autocode2/speech-to-text --api-key YOUR_API_KEY -r 44100 -c 2 -o high-quality.wav\n\n# Use custom transcription prompt\nnpx @autocode2/speech-to-text --api-key YOUR_API_KEY -p \"Provide a detailed transcription with punctuation\"\n\n# Output in JSON format\nnpx @autocode2/speech-to-text --api-key YOUR_API_KEY --format json > output.json\n\n# Pipe transcription to other tools\nnpx @autocode2/speech-to-text --api-key YOUR_API_KEY | jq .text\n```\n\n### JSON Output Format\n\nWhen using JSON output (either explicitly with `--format json` or implicitly when piping), the output will be a JSON object with the following structure:\n\n```json\n{\n  \"text\": \"The transcribed text\",\n  \"timestamp\": \"2024-01-20T12:34:56.789Z\",\n  \"input\": \"input-file.wav\", // If provided\n  \"output\": \"output-file.wav\", // If provided\n  \"sampleRate\": 16000, // If recording\n  \"channels\": 1, // If recording\n  \"model\": \"gemini-1.5-flash\" // If specified\n}\n```\n\n## Library Usage\n\nYou can also use this as a library in your Node.js projects:\n\n```typescript\nimport { SpeechToText } from \"@autocode2/speech-to-text\";\n\nconst stt = new SpeechToText({\n  apiKey: \"your-google-api-key\",\n  recording: {\n    sampleRate: 16000,\n    channels: 1,\n  },\n  transcription: {\n    model: \"gemini-1.5-flash\",\n    prompt: \"Custom transcription prompt\",\n  },\n});\n\n// Record to temporary file (automatically cleaned up)\nconst text1 = await stt.recordAndTranscribe();\n\n// Record and save to file\nconst text2 = await stt.recordAndTranscribe(\"output.wav\");\n\n// Transcribe existing file\nconst text3 = await stt.transcribe(\"existing.wav\");\n```\n\n## License\n\nISC\n\n## Author\n\nGareth Andrew\n","readmeFilename":"README.md","homepage":"https://github.com/autocode2/speech-to-text","keywords":["audio","speech","speech-to-text","cli","sox","google-gemini"],"repository":{"type":"git","url":"git+https://github.com/autocode2/speech-to-text.git"},"bugs":{"url":"https://github.com/autocode2/speech-to-text/issues"}}