Whishper: Local Audio Transcription and Subtitles With Whisper
Whishper turns audio and video into text and subtitles, entirely on your own hardware. Upload a file or paste a URL, and it transcribes with OpenAI’s Whisper models via the faster-whisper engine.
Status: the developer is working on a complete rewrite (v4). The current main branch is not expected to receive new releases or updates in the meantime. It still works, but factor that in for a long-term setup.
Features
- Transcribe audio and video files, or media URLs
- Export as plain text, JSON, SRT, and VTT subtitles
- A subtitle editor: split and merge segments, adjust timing, characters-per-second warnings, and live highlighting during playback
- Translation through a LibreTranslate instance
- Works offline; nothing leaves your server
Deployment
Whishper runs as several containers: a transcription API using faster-whisper, a backend, MongoDB, and a SvelteKit web interface, with an optional LibreTranslate container. The project provides quickstart scripts and Docker Compose files for CPU and NVIDIA GPU setups. A GPU makes transcription many times faster; on a CPU, choose a smaller model.
Alternatives
- whisper.cpp: a fast C++ Whisper implementation for the command line and as a server
- Speaches: an OpenAI-compatible speech-to-text and text-to-speech API server
- MinusPod: uses Whisper to remove ads from podcasts
License
GNU Affero General Public License v3.0.