{"_id":"@adamhancock/transcribe","name":"@adamhancock/transcribe","dist-tags":{"latest":"1.0.5"},"versions":{"1.0.5":{"name":"@adamhancock/transcribe","version":"1.0.5","description":"CLI tool for transcribing and summarizing MP4 recordings using Whisper and Ollama","main":"dist/index.js","type":"module","bin":{"transcribe":"dist/index.js"},"keywords":["transcription","whisper","ollama","video","audio","summarization","cli"],"author":{"name":"Adam Hancock"},"license":"MIT","repository":{"type":"git","url":"git+https://github.com/adamhancock/cli.git","directory":"packages/transcribe"},"engines":{"node":">=18.0.0"},"dependencies":{"axios":"^1.11.0","chalk":"^5.4.1","commander":"^14.0.0","ora":"^8.2.0","zx":"^8.7.1"},"devDependencies":{"@types/node":"^24.1.0","tsx":"^4.20.3","typescript":"^5.8.3"},"scripts":{"build":"tsc","dev":"tsx src/index.ts","start":"node dist/index.js"},"_id":"@adamhancock/transcribe@1.0.5","bugs":{"url":"https://github.com/adamhancock/cli/issues"},"homepage":"https://github.com/adamhancock/cli#readme","_integrity":"sha512-884mDyWg8xWfvlAeAdNwdm7NoiZWHWBzWGVJ3dn+jjmA7duOs1O+AH8AqX/vyVw5TfMmSYRH5Z9FM7P+PSQl3w==","_resolved":"/tmp/f3dd6e1ce114b3708dae4a477dc23b5d/adamhancock-transcribe-1.0.5.tgz","_from":"file:adamhancock-transcribe-1.0.5.tgz","_nodeVersion":"22.17.1","_npmVersion":"10.9.2","dist":{"integrity":"sha512-884mDyWg8xWfvlAeAdNwdm7NoiZWHWBzWGVJ3dn+jjmA7duOs1O+AH8AqX/vyVw5TfMmSYRH5Z9FM7P+PSQl3w==","shasum":"91b3e95bd9077324fefe25cf4566aa91e34580c4","tarball":"https://registry.npmjs.org/@adamhancock/transcribe/-/transcribe-1.0.5.tgz","fileCount":6,"unpackedSize":52677,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIBC/tnQ2prCn2eHzyRwaAmBm3VH0xxsxLEk9+fJXpXmlAiEA7bnAfCAMm2WDVG8jzilPMA7GboHsyml1oMEFha7ysuw="}]},"_npmUser":{"name":"adam1571","email":"adammhancock@gmail.com"},"directories":{},"maintainers":[{"name":"adam1571","email":"adammhancock@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/transcribe_1.0.5_1754301010005_0.9837705937028189"},"_hasShrinkwrap":false}},"time":{"created":"2025-08-04T09:50:09.920Z","1.0.5":"2025-08-04T09:50:10.211Z","modified":"2025-08-04T09:50:10.495Z"},"maintainers":[{"name":"adam1571","email":"adammhancock@gmail.com"}],"description":"CLI tool for transcribing and summarizing MP4 recordings using Whisper and Ollama","homepage":"https://github.com/adamhancock/cli#readme","keywords":["transcription","whisper","ollama","video","audio","summarization","cli"],"repository":{"type":"git","url":"git+https://github.com/adamhancock/cli.git","directory":"packages/transcribe"},"author":{"name":"Adam Hancock"},"bugs":{"url":"https://github.com/adamhancock/cli/issues"},"license":"MIT","readme":"# Transcribe CLI\n\nA powerful TypeScript CLI tool for transcribing and summarizing MP4 recordings or VTT subtitle files using local models with synchronized audio and visual analysis.\n\n## Key Features\n\n- 🎙️ **Timestamped Transcription**: Full audio transcription with precise timestamps using Whisper\n- 📝 **VTT File Support**: Direct summarization of WebVTT subtitle files without transcription\n- 🎬 **Flexible Input**: Accepts MP4 videos, VTT subtitles, or both for combined analysis\n- 🧠 **Intelligent Frame Selection**: AI analyzes transcript to extract frames at meaningful moments with reasoning\n- 🖼️ **Context-Aware Frame Analysis**: Each frame is analyzed with awareness of previous frames\n- 🔄 **Synchronized Analysis**: Matches frame descriptions with relevant transcript segments\n- 📊 **Map-Reduce Summarization**: Uses map-reduce strategy for detailed, comprehensive summaries\n- 🎯 **Smart Model Selection**: Uses llava for visual analysis and llama3.2 for text summarization\n- ⚡ **Optimized Processing**: Parallel audio/frame extraction, sequential frame analysis for context\n\n## Prerequisites\n\n- Node.js 18+\n- FFmpeg installed\n  - macOS: `brew install ffmpeg`\n  - Ubuntu/Debian: `sudo apt update && sudo apt install ffmpeg`\n  - Windows: Download from [ffmpeg.org](https://ffmpeg.org/download.html)\n- Python & pip (for Whisper)\n  - macOS: `brew install python` (includes pip)\n  - Ubuntu/Debian: `sudo apt install python3-pip`\n  - Windows: Download from [python.org](https://python.org) (includes pip)\n- OpenAI Whisper - runs locally for transcription\n  - Install: `pip install openai-whisper`\n  - Note: First run may download model files (~140MB for base model)\n- **Ollama** - for summarization and frame analysis\n  - Install: `curl -fsSL https://ollama.com/install.sh | sh`\n  - Runs automatically as a background service (port 11434)\n  - Default models (automatically pulled on first use):\n    - `llava` - for visual frame analysis (multimodal)\n    - `llama3.2` - for text summarization (map-reduce strategy)\n\n## How it works\n\n1. **Flexible Input Processing**:\n   - **MP4 Files**: Extracts audio, transcribes with Whisper, analyzes video frames\n   - **VTT Files**: Parses WebVTT subtitles directly for immediate summarization\n   - **Combined**: Use VTT for transcript + MP4 for visual frame analysis\n\n2. **Audio Extraction & Transcription** (for MP4): \n   - Extracts audio and transcribes with Whisper\n   - Produces timestamped segments (e.g., `[1:30] Speaker says...`)\n   - Captures precise timing for synchronization with frames\n\n3. **Intelligent Frame Selection** (when video available): \n   - AI analyzes transcript content to identify key moments\n   - Selects frames at meaningful timestamps with reasoning\n   - Shows why each moment was selected (e.g., \"Introduction\", \"Key demonstration\")\n   - Validates timestamps against actual video duration\n   - Falls back to content-based distribution if AI fails\n\n4. **Context-Aware Frame Analysis**: \n   - Uses `llava` model for multimodal visual understanding\n   - Analyzes frames sequentially to maintain narrative flow\n   - Each frame analysis includes:\n     - Visual description from the frame\n     - Context from previous frame\n     - Relevant transcript excerpt from that moment\n   \n5. **Map-Reduce Summarization**: \n   - **Map Phase**: Breaks transcript into chunks, extracts all details from each\n   - **Reduce Phase**: Synthesizes comprehensive summary with multiple sections\n   - Creates extremely detailed summaries including:\n     - **Meeting Overview**: Purpose, participants, context\n     - **Topic Breakdown**: Detailed analysis of each topic discussed\n     - **All Decisions**: Every decision made with context\n     - **Action Items**: Who, what, when for each task\n     - **Tools & Systems**: All platforms and tools mentioned\n     - **Technical Details**: Specifications and configurations\n     - **Problems & Solutions**: Issues raised and resolutions\n     - **Next Steps**: Future plans and follow-ups\n\n## Installation\n\n### Using npx (no installation required)\n```bash\n# Transcribe MP4 video\nnpx @adamhancock/transcribe-cli video.mp4\n\n# Summarize VTT subtitle file\nnpx @adamhancock/transcribe-cli subtitles.vtt\n\n# Combine VTT transcript with MP4 video analysis\nnpx @adamhancock/transcribe-cli subtitles.vtt video.mp4\n```\n\n### Global installation\n```bash\nnpm install -g @adamhancock/transcribe-cli\n```\n\n### From source\n```bash\ngit clone https://github.com/adamhancock/transcribe-cli.git\ncd transcribe-cli\nnpm install\nnpm run build\nnpm link  # Makes 'transcribe' command available globally\n```\n\n## Usage\n\nBasic usage:\n```bash\n# Transcribe and summarize MP4 video\ntranscribe video.mp4\n\n# Summarize VTT subtitle file\ntranscribe meeting.vtt\n\n# Combine VTT transcript with MP4 video for frame analysis\ntranscribe meeting.vtt video.mp4\n```\n\nOptions:\n- `-o, --output <path>` - Output file path (default: input_transcript.txt)\n- `-a, --audio-only` - Only extract audio without transcription\n- `-t, --transcribe-only` - Only transcribe without summarization\n- `--host <host>` - Ollama API host (default: http://localhost:11434)\n- `--model <model>` - Ollama model name for both frame and text analysis (overrides defaults)\n- `--whisper-model <model>` - Whisper model size (default: base)\n- `--keep-audio` - Keep extracted audio file after processing\n- `--no-analyze-frames` - Disable frame extraction and analysis (enabled by default)\n- `--frame-interval <seconds>` - Seconds between frame extraction (default: 3)\n- `--max-frames <count>` - Maximum number of frames to extract (default: 30)\n- `--keep-frames` - Keep extracted frames after processing\n- `--save-timestamps` - Save timestamped transcript and frame data to JSON file\n- `--plain-transcript` - Save transcript without timestamps in output file\n\n## Examples\n\n```bash\n# Basic MP4 transcription with intelligent frame analysis\ntranscribe recording.mp4\n\n# Summarize a VTT subtitle file\ntranscribe meeting.vtt\n\n# Combine VTT subtitles with video for visual analysis\ntranscribe meeting.vtt recording.mp4\n\n# Quick transcription without analysis\ntranscribe recording.mp4 --transcribe-only\n\n# Extract more frames for detailed videos\ntranscribe recording.mp4 --max-frames 50\n\n# Use a faster Whisper model for quick results\ntranscribe recording.mp4 --whisper-model tiny\n\n# Use a larger Whisper model for better accuracy\ntranscribe recording.mp4 --whisper-model large\n\n# Save all artifacts for further processing\ntranscribe recording.mp4 --keep-audio --keep-frames --save-timestamps\n\n# Get plain transcript without timestamps\ntranscribe recording.mp4 --plain-transcript\n\n# Use different models for better summaries\ntranscribe meeting.vtt --model llama3.1  # More powerful model\n\n# Process MP4 without frame analysis (audio only)\ntranscribe recording.mp4 --no-analyze-frames\n\n# Custom output location\ntranscribe meeting.vtt -o ~/Documents/meeting-summary.txt\n```\n\n## Performance Tips\n\n- **Faster processing**: Use `--whisper-model tiny` for quick drafts\n- **Better accuracy**: Use `--whisper-model medium` or `large` for important content\n- **Reduce frames**: Lower `--max-frames` for faster analysis of long videos\n- **Vision models**: `llava` is default; try `bakllava` or `llava:13b` for better accuracy\n\n## Output Format\n\nThe tool provides multiple outputs:\n\n### Console Output\n- Progress indicators for each processing stage\n- Frame extraction with reasoning (e.g., `✓ Frame 1 at 0:00 - Introduction`)\n- Real-time analysis progress\n- Structured final summary\n\n### File Outputs\n1. **Main transcript file** (default: `video_transcript.txt`)\n   - Timestamped transcript (e.g., `[1:30] Speaker says...`)\n   - Comprehensive structured summary with sections\n   - Use `--plain-transcript` for transcript without timestamps\n\n2. **Timestamped data file** (optional: `video_timestamped.json`)\n   - Complete transcript with precise timestamps\n   - Frame analyses with timestamps and descriptions\n   - Frame extraction reasoning\n   - Useful for creating subtitles, chapters, or navigating content\n\n### Example Output\n\n```\nTranscript of presentation.mp4:\n\n[0:00] Welcome everyone to today's presentation on machine learning.\n[0:05] I'll be covering three main topics today...\n[0:12] First, let's understand what neural networks are...\n\n---\n\nSummary:\n\n## Overview\nThis video presents an introduction to machine learning concepts...\n\n## Main Topics Covered\n1. **Neural Networks Basics**: The speaker explains...\n2. **Training Process**: Demonstrates how models learn...\n3. **Practical Applications**: Shows real-world examples...\n\n## Key Points & Insights\n- Neural networks mimic human brain structure\n- Training requires large datasets and computational power\n- Applications range from image recognition to natural language\n\n## Visual Elements\n- Slide presentations with diagrams\n- Live coding demonstration\n- Results visualization graphs\n\n## Action Items or Next Steps\n- Practice with the provided code examples\n- Explore the recommended datasets\n- Join the community forum for questions\n```\n\n\n## Development\n\nRun directly with tsx:\n```bash\nnpm run dev -- video.mp4\n```\n\nBuild:\n```bash\nnpm run build\n```\n\nPublish to npm:\n```bash\nnpm version patch  # or minor/major\nnpm publish\n```","readmeFilename":"README.md","_rev":"1-27a3667a2275c3c8d691e20173a3af8a"}