{"_id":"@audio-ml/asr","name":"@audio-ml/asr","dist-tags":{"latest":"1.0.0"},"versions":{"1.0.0":{"name":"@audio-ml/asr","version":"1.0.0","description":"FastConformer ASR (TensorFlow.js) for use with audio-ml","type":"module","main":"./dist/index.js","types":"./dist/index.d.ts","exports":{".":{"import":"./dist/index.js","types":"./dist/index.d.ts"},"./package.json":"./package.json"},"scripts":{"build":"tsc","prepublishOnly":"yarn build"},"keywords":["asr","speech-recognition","fastconformer","tensorflow.js","audio-ml"],"author":{"name":"Abijah Kajabika"},"license":"MIT","repository":{"type":"git","url":"git+https://github.com/AbijahKaj/audio-ml.git","directory":"packages/asr"},"peerDependencies":{"@tensorflow/tfjs-node":"^4.22.0","audio-ml":"^1.0.0"},"peerDependenciesMeta":{"@tensorflow/tfjs-node":{"optional":true}},"dependencies":{"@tensorflow/tfjs":"^4.22.0","@tensorflow/tfjs-backend-wasm":"^4.22.0","@tensorflow/tfjs-backend-webgpu":"^4.22.0","fft.js":"^4.0.4"},"devDependencies":{"@tensorflow/tfjs-node":"4.22.0","@types/node":"^22.0.0","audio-ml":"file:../..","typescript":"~5.9.3"},"engines":{"node":">=18.0.0"},"_id":"@audio-ml/asr@1.0.0","gitHead":"5f8a483ba83a9530ebf25722a86d264cc5f98c06","bugs":{"url":"https://github.com/AbijahKaj/audio-ml/issues"},"homepage":"https://github.com/AbijahKaj/audio-ml#readme","_nodeVersion":"22.16.0","_npmVersion":"10.9.2","dist":{"integrity":"sha512-dr5TwhQWzc/sQxe4+FIPPKRrgfRa5q5hZ3zDka5X7nBkJcDmS0RwSVu6yJxHn05FswJHTfUw2b5+vPIo87d3Dg==","shasum":"56d9d6697de9405802a7020c616c2a9468cf7f8e","tarball":"https://registry.npmjs.org/@audio-ml/asr/-/asr-1.0.0.tgz","fileCount":123,"unpackedSize":277603,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQDtaoNIf7A30RFn9KJgUU9LwXEEdnknPk7psnb9xvcFIgIgWLq77lcX1VZT+kNTlEFTAVFP4btYi71Dr4J7Io6aWeM="}]},"_npmUser":{"name":"abijahkaj","email":"migishoabijah@gmail.com"},"directories":{},"maintainers":[{"name":"abijahkaj","email":"migishoabijah@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/asr_1.0.0_1774105679145_0.5500346414906194"},"_hasShrinkwrap":false}},"time":{"created":"2026-03-21T15:07:59.072Z","1.0.0":"2026-03-21T15:07:59.293Z","modified":"2026-03-21T15:07:59.473Z"},"maintainers":[{"name":"abijahkaj","email":"migishoabijah@gmail.com"}],"description":"FastConformer ASR (TensorFlow.js) for use with audio-ml","homepage":"https://github.com/AbijahKaj/audio-ml#readme","keywords":["asr","speech-recognition","fastconformer","tensorflow.js","audio-ml"],"repository":{"type":"git","url":"git+https://github.com/AbijahKaj/audio-ml.git","directory":"packages/asr"},"author":{"name":"Abijah Kajabika"},"bugs":{"url":"https://github.com/AbijahKaj/audio-ml/issues"},"license":"MIT","readme":"# @audio-ml/asr\n\nFastConformer speech recognition in **TypeScript**, powered by **TensorFlow.js**. Models are exported from [NVIDIA NeMo](https://github.com/NVIDIA/NeMo) to SafeTensors plus JSON config and vocabulary—see the [`audio-ml` repo](https://github.com/AbijahKaj/audio-ml) and `tools/export_nemo_to_safetensors.py`.\n\nThis package depends on **[`audio-ml`](https://www.npmjs.com/package/audio-ml)** for shared application types (for example `BaseApplication` and VAD used by streaming endpointing).\n\n## Install\n\n```bash\nnpm install audio-ml @audio-ml/asr\n```\n\nOptional, for **native TensorFlow in Node.js** (faster than pure JS CPU):\n\n```bash\nnpm install @tensorflow/tfjs-node\n```\n\n`@tensorflow/tfjs-node` is an **optional peer**; install it only when you use the `tensorflow` backend in Node.\n\n## Ready-to-use models (same as the `audio-ml` demo)\n\nThese Hugging Face repos ship `model.safetensors`, `model_config.json`, and `vocab.json` on the `main` branch (NeMo → export via `tools/export_nemo_to_safetensors.py` in the main repo).\n\n| Model | Hugging Face repo | Notes |\n|-------|-------------------|--------|\n| Parakeet TDT 110M | [AbijahKaj/parakeet-tdt-110m-web](https://huggingface.co/AbijahKaj/parakeet-tdt-110m-web) | English, TDT decoder, ~220 MB weights |\n| Parakeet RNNT 120M (streaming) | [AbijahKaj/parakeet-rnnt-120m-web](https://huggingface.co/AbijahKaj/parakeet-rnnt-120m-web) | English, RNNT, streaming-oriented |\n| FastConformer TDT Large | [AbijahKaj/fastconformer-tdt-large-web](https://huggingface.co/AbijahKaj/fastconformer-tdt-large-web) | English, TDT, ~218 MB weights |\n\nResolve URLs follow this pattern (`{repo}` = `username/repo`):\n\n`https://huggingface.co/{repo}/resolve/main/model.safetensors`  \n`https://huggingface.co/{repo}/resolve/main/model_config.json`  \n`https://huggingface.co/{repo}/resolve/main/vocab.json`\n\n## Quick start\n\nExample using **Parakeet TDT 110M** (same default-style URLs as `demo/pages/SpeechRecognizerDemo.ts`):\n\n```typescript\nimport { FastConformerASR, type ASRResult } from '@audio-ml/asr';\n\nconst HF = 'https://huggingface.co/AbijahKaj/parakeet-tdt-110m-web/resolve/main';\n\nconst asr = new FastConformerASR({\n  sampleRate: 16_000,\n  modelPath: `${HF}/model.safetensors`,\n  configPath: `${HF}/model_config.json`,\n  vocabPath: `${HF}/vocab.json`,\n  backend: 'webgpu', // browser: 'webgpu' | 'webgl' | 'wasm' | 'cpu'\n  streaming: true,\n});\n\nawait asr.load();\n\nasr.on('partial', (p) => console.log(p.text));\nasr.on('final', (r: ASRResult) => console.log(r.text));\n\nasr.processFrame(pcmFrame);\n```\n\nTo load from already-fetched buffers:\n\n```typescript\nawait asr.loadFromBuffers(modelArrayBuffer, configJsonString, vocabJsonString);\n```\n\nOffline pass:\n\n```typescript\nconst result = await asr.transcribe(audioFloat32);\n```\n\n## TensorFlow.js backends\n\n| Backend       | Typical use |\n|---------------|-------------|\n| `webgpu`      | Browser, best GPU path when supported |\n| `webgl`       | Browser, broader GPU support |\n| `wasm`        | Browser, good CPU throughput via WASM |\n| `cpu`         | Browser or Node, pure JS (slow for large models) |\n| `tensorflow`  | **Node only** — requires `@tensorflow/tfjs-node` |\n\nWASM backend options:\n\n```typescript\nawait asr.load(); // after constructing with:\n// backend: 'wasm',\n// backendOptions: { wasmPathPrefix: '/tfjs-wasm/' }\n```\n\nServe `.wasm` files from `tfjs-backend-wasm` with correct MIME type (see the main repo demo Vite config).\n\n## Swappable compute layer\n\nInference is expressed against a **`ComputeBackend`** interface. **`TfjsBackend`** is the default implementation; you can supply another backend that implements the same operations if you integrate a different runtime.\n\n## Exports\n\nBesides **`FastConformerASR`**, the package exports encoder/decoder/feature/text/model helpers (for example `FastConformerEncoder`, `createDecoder`, `FeaturePipeline`, `loadSafeTensors`, `parseModelConfig`, streaming types, and `Endpointer`). See `src/index.ts` for the full public API.\n\n## Requirements\n\n- **Node.js** ≥ 18\n- **Peer:** `audio-ml` ^1.0.0\n\n## License\n\nMIT — see [LICENSE](./LICENSE).\n\n## Repository\n\n[github.com/AbijahKaj/audio-ml](https://github.com/AbijahKaj/audio-ml) (package path: `packages/asr`).\n","readmeFilename":"README.md","_rev":"1-4a253da2c7a8f93d25d2773420935b1b"}