Insanely Fast Whisper Promoted as Free Local Transcription Tool
Nav Toor promotes Insanely Fast Whisper, an open-source MIT-licensed tool that runs Whisper transcription locally on NVIDIA GPUs or Apple Silicon. He claims 150 minutes of audio transcribes in 98 seconds, citing benchmarks against paid cloud and transcription services.
Original post · 1 min read
Someone just open sourced a tool that does it for $0. And it's faster than all of them.
It's called Insanely Fast Whisper. And that's not hype. That's the benchmark.
150 minutes of audio. 98 seconds to transcribe. On your own machine. No API key. No cloud. No per-minute billing.
Here's what the numbers look like:
→ Whisper Large v3 + Flash Attention 2: 150 min of audio in 98 seconds
→ Distil Whisper + Flash Attention 2: 150 min in 78 seconds
→ Standard Whisper without optimization: 31 minutes for the same job
→ That's a 19x speedup. Same model. Same accuracy. Just faster.
Here's what it does:
→ One command to transcribe any audio file or URL
→ Speaker diarization — knows WHO said WHAT
→ Transcription AND translation to other languages
→ Runs on NVIDIA GPUs and Mac (Apple Silicon)
→ Flash Attention 2 for maximum speed
→ Clean JSON output with timestamps
→ Works with every Whisper model variant
Here's the wildest part:
Otter.ai charges $100/year. Rev charges $1.50/minute. Descript charges $24/month. Enterprise transcription contracts cost thousands.
Podcasters, journalists, researchers, lawyers, content creators — anyone still paying for transcription is lighting money on fire.
8.8K GitHub stars. 633 forks. MIT License.
100% Open Source.
(Link in the comments)