{"_id":"@amirhosseinnouri/voxgen","name":"@amirhosseinnouri/voxgen","dist-tags":{"latest":"1.0.0"},"versions":{"1.0.0":{"name":"@amirhosseinnouri/voxgen","version":"1.0.0","description":"Turn a text file into a narrated audio file with Fish Audio.","keywords":["text-to-speech","tts","fish-audio","speech-synthesis","narration","audiobook","voice","cli","bun"],"license":"MIT","repository":{"type":"git","url":"git+https://github.com/amirhosseinNouri/voxgen.git"},"homepage":"https://github.com/amirhosseinNouri/voxgen","bugs":{"url":"https://github.com/amirhosseinNouri/voxgen/issues"},"module":"src/index.ts","type":"module","bin":{"voxgen":"src/index.ts"},"publishConfig":{"access":"public"},"engines":{"bun":">=1.2.0"},"scripts":{"start":"bun run src/index.ts","test":"bun test","typecheck":"tsc --noEmit"},"devDependencies":{"@types/bun":"^1.3.14","@types/node":"^26.0.1"},"peerDependencies":{"typescript":"^5"},"dependencies":{"@clack/prompts":"^1.6.0","fish-audio":"^0.1.0","zod":"^4.4.3"},"gitHead":"0a9bf9da89fa5b3abfce7805f83473c289d38cb1","_id":"@amirhosseinnouri/voxgen@1.0.0","_nodeVersion":"24.18.0","_npmVersion":"11.16.0","dist":{"integrity":"sha512-ECtWKlmmmZOXB0Jr0CIYQdWgejdQ7T7BAXXnNDRiO9EdWGzLjv8eFKIq/OWJawN7rJCvTfAKQRswPe8jRZZsKw==","shasum":"ef7d5b1678e097b4a0878e4636bb0cd346f94384","tarball":"https://registry.npmjs.org/@amirhosseinnouri/voxgen/-/voxgen-1.0.0.tgz","fileCount":26,"unpackedSize":72082,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIG7fIlKdopB0NIpi6R+FkCrbIL5jD5ATKZLoI4dj16zrAiAPNVqBjGRzNHuMlIL0PSPDxLI8Ztpwfmy442mG1aYzzQ=="}]},"_npmUser":{"name":"amirhosseinnouri","email":"Amir.h.nouri2000@gmail.com"},"directories":{},"maintainers":[{"name":"amirhosseinnouri","email":"Amir.h.nouri2000@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/voxgen_1.0.0_1786274560443_0.9858603587343806"},"_hasShrinkwrap":false}},"time":{"created":"2026-08-09T11:22:40.294Z","1.0.0":"2026-08-09T11:22:40.582Z","modified":"2026-08-09T11:22:40.766Z"},"maintainers":[{"name":"amirhosseinnouri","email":"Amir.h.nouri2000@gmail.com"}],"description":"Turn a text file into a narrated audio file with Fish Audio.","homepage":"https://github.com/amirhosseinNouri/voxgen","keywords":["text-to-speech","tts","fish-audio","speech-synthesis","narration","audiobook","voice","cli","bun"],"repository":{"type":"git","url":"git+https://github.com/amirhosseinNouri/voxgen.git"},"bugs":{"url":"https://github.com/amirhosseinNouri/voxgen/issues"},"license":"MIT","readme":"# voxgen\n\nTurn a text file into a narrated audio file. [Fish Audio](https://fish.audio) does the\nspeaking; this tool does the parts around it — chunking, cost, caching, and joining the\npieces back into one clean recording.\n\nThe default backend is **`s2.1-pro-free`**, which is the same model as `s2.1-pro` at $0. A\nplain run costs nothing.\n\n## Setup\n\n```bash\nbunx @amirhosseinnouri/voxgen ./article.md\n```\n\nOr from source:\n\n```bash\ngit clone https://github.com/amirhosseinNouri/voxgen\ncd voxgen\nbun install\n```\n\nPut your key in `.env`:\n\n```\nFISH_API_KEY=...   # https://fish.audio/go-api/api-keys\n```\n\n`ffmpeg` is needed only to write something other than WAV — mp3, opus, flac and m4a all go\nthrough it.\n\n### Setting up with AI\n\nTo install, configure and first-run voxgen with an AI assistant, tell it to use the\n`setup` skill: `use the setup skill to set up voxgen`. It walks through picking how to run\nit, getting a key into `.env`, deciding on ffmpeg, verifying with a free run, and choosing\na voice — in that order, verifying each step as it goes.\n\n## Usage\n\n```bash\nbun start ./article.md              # → output/article-<timestamp>/audio.mp3\nbun start \"Good morning.\"           # text that is not a file is read literally\ncat notes.md | bun start --stdin    # or pipe it in\nbun start ./article.md -o talk.opus # write exactly here, in this container\nbun start voices calm               # find a voice id for --voice\n```\n\nOptions:\n\n| Flag | Meaning |\n| --- | --- |\n| `-o, --output <path>` | Write here instead of `output/<name>-<timestamp>/` |\n| `--format <fmt>` | `wav`, `mp3`, `opus`, `flac`, `m4a` (default `mp3`) |\n| `--voice <id>` | A voice id from `voxgen voices` |\n| `--model <id>` | Fish backend (default `s2.1-pro-free`) |\n| `--speed <n>` | 0.5–2.0; 1 is the voice's own pace |\n| `-y, --yes` | Skip the cost confirmation |\n\nEach run without `-o` writes a fresh `output/<name>-<timestamp>/` containing `audio.<fmt>`\nand `script.txt` — the normalized text that was actually spoken, which is what to check\nwhen a word comes out wrong. `output/` is git-ignored.\n\n## How it works\n\n1. **Normalization** — Markdown markers are flattened first, because a narrator should read\n   \"Title\", not \"hash Title\". Headings, bullets, emphasis and inline code lose their\n   punctuation; a link reads as its label rather than its URL. A `#` inside a sentence\n   survives, since only leading heading markers are stripped.\n2. **Voice pinning** — every request carries an explicit `reference_id`. Fish accepts a\n   request without one and quietly picks a speaker for it, which is harmless for a single\n   sentence and ruinous for a document: an article synthesized in 84 requests comes back in\n   84 different voices. There is therefore always a voice, defaulting to `Sarah`; change it\n   with `--voice`, `FISH_VOICE_ID`, or `voxgen voices`.\n3. **Chunking** — the script is cut into pieces of about 1500 UTF-8 bytes. Paragraphs are\n   the outer unit and never share a chunk; within a paragraph the split is by sentence, and\n   only by word when one sentence is over budget on its own. A word longer than the budget\n   — a URL, a hash — is cut on a character boundary rather than dropped, and never in the\n   middle of a multi-byte character. Chunks exist for progress, caching and retry\n   granularity, not for a provider limit.\n4. **Cost estimate** — Fish bills per UTF-8 byte of *input*, so the entire bill is knowable\n   before a single byte goes out. Chunks already in the cache are subtracted and the run\n   prints what is left to send, with its price, before asking to continue. Nothing reaches\n   the provider until that prompt is answered — and on the free backend there is no prompt,\n   because there is nothing to decide.\n5. **Synthesis** — each chunk comes back as raw 16-bit mono PCM rather than mp3. Raw samples\n   join gaplessly, where two mp3 streams glued together leave an encoder-padding click at\n   every seam. Each chunk is written to the cache before the next request goes out, so an\n   interrupted run resumes where it stopped and never pays for the same sentence twice.\n   Transient failures — 429, 5xx — are retried with backoff; a rejected request is not,\n   because sending the same bytes again buys the same rejection. A long article is a lot of\n   small requests, so they go out in a sliding window of four rather than one at a time —\n   but they are written strictly in order, so the recording is identical to what a\n   sequential run would produce.\n6. **Joining** — the chunks are streamed straight into a WAV file, with a short silence\n   inserted wherever the source had a blank line, so paragraphs land as pauses instead of\n   the voice running two thoughts together. The file is written as it goes and its header\n   patched at the end, so an hour of audio never has to fit in memory. A run that fails\n   part-way deletes its output: a truncated WAV is still a playable WAV, and one left on\n   disk looks exactly like a finished result until someone reaches the end of it.\n7. **Encoding** — anything but WAV is transcoded by ffmpeg at rates suited to speech (96k\n   mp3, 48k opus), and the intermediate is removed once that succeeds.\n\n## Caching\n\nSynthesized audio is cached under `.cache/voxgen/`, keyed by a hash of the request: the\ntext, the backend, the voice, the speed, the volume and the sample rate. Editing one\nparagraph of a long article re-synthesizes that paragraph and nothing else. Changing the\nvoice correctly invalidates everything. Cache files are written under a temporary name and\nrenamed into place, so a run killed mid-write cannot leave a half-length chunk that the\nnext run treats as complete.\n\nDelete `.cache/voxgen/` to start paying again.\n\n## Configuration\n\nEverything has a default; nothing but the key is required.\n\n| Variable | Default | |\n| --- | --- | --- |\n| `FISH_API_KEY` | — | Required |\n| `FISH_MODEL` | `s2.1-pro-free` | Backend, sent as the `model` header |\n| `FISH_VOICE_ID` | `9335…406a` (`Sarah`) | Voice; same values as `--voice`. Never empty — see step 2 |\n| `FISH_BASE_URL` | `https://api.fish.audio` | |\n| `TTS_LATENCY` | `normal` | `balanced` trades quality for first-byte latency |\n| `TTS_SPEED` | `1` | 0.5–2.0 |\n| `TTS_VOLUME` | `0` | dB adjustment applied by the provider |\n| `TTS_SAMPLE_RATE` | `44100` | |\n| `TTS_PRICE_PER_MILLION_BYTES` | published rate | Override for the cost estimate |\n| `CHUNK_BYTES` | `1500` | Request size, in UTF-8 bytes |\n| `PARAGRAPH_PAUSE_MS` | `350` | Silence at a blank line |\n| `CONCURRENCY` | `4` | Requests in flight at once; Fish's entry tier allows 5 |\n| `CACHE_DIR` | `.cache/voxgen` | |\n\nPrices are per **UTF-8 byte**, not per character. One Persian or CJK character is three\nbytes, so a document that looks half the length of an English one can cost three times as\nmuch — `voxgen` counts and prices in bytes for exactly that reason.\n\nAnything invalid is rejected at startup with the variable named, rather than surfacing\nmid-run once part of the script has already been paid for.\n\n## Development\n\n```bash\nbun test         # 118 tests, no network\nbun run typecheck\n```\n","readmeFilename":"README.md","_rev":"1-42bd2b637a8c037856f97f16aea7a3c9"}