Whishper: Local Audio Transcription and Subtitles With Whisper

Whishper: Local Audio Transcription and Subtitles With Whisper

Whishper turns audio and video into text and subtitles, entirely on your own hardware. Upload a file or paste a URL, and it transcribes with OpenAI’s Whisper models via the faster-whisper engine.

Status: the developer is working on a complete rewrite (v4). The current main branch is not expected to receive new releases or updates in the meantime. It still works, but factor that in for a long-term setup.

Features

  • Transcribe audio and video files, or media URLs
  • Export as plain text, JSON, SRT, and VTT subtitles
  • A subtitle editor: split and merge segments, adjust timing, characters-per-second warnings, and live highlighting during playback
  • Translation through a LibreTranslate instance
  • Works offline; nothing leaves your server

Deployment

Whishper runs as several containers: a transcription API using faster-whisper, a backend, MongoDB, and a SvelteKit web interface, with an optional LibreTranslate container. The project provides quickstart scripts and Docker Compose files for CPU and NVIDIA GPU setups. A GPU makes transcription many times faster; on a CPU, choose a smaller model.

Alternatives

  • whisper.cpp: a fast C++ Whisper implementation for the command line and as a server
  • Speaches: an OpenAI-compatible speech-to-text and text-to-speech API server
  • MinusPod: uses Whisper to remove ads from podcasts

License

GNU Affero General Public License v3.0.