Speaches: An OpenAI-Compatible Speech-to-Text and Text-to-Speech Server
Speaches is a server for speech models that speaks the OpenAI audio API. Point any tool or SDK that supports OpenAI’s transcription or text-to-speech endpoints at your Speaches server, and it works, with the models running on your hardware. The project describes itself as “Ollama, but for TTS/STT models.”
Features
- Speech-to-text with faster-whisper, including streaming transcription sent as it is processed
- Translation of speech to English
- Text-to-speech with Piper and Kokoro, a small model that produces notably natural voices
- OpenAI API compatibility, so existing clients and libraries work unchanged
- Dynamic model loading: models load on demand and unload after inactivity to free memory
- CPU and GPU support
Uses
- A private backend for voice assistants, including Home Assistant voice pipelines
- Voice input and read-aloud for chat front ends such as Open WebUI
- Batch transcription from scripts using the OpenAI SDK
- Replacing paid transcription APIs in your own projects
Deployment
Docker Compose files are provided for CPU and CUDA setups. Whisper runs acceptably on a modern CPU with small models; larger models and real-time use benefit greatly from an NVIDIA GPU.
Like other model servers, it has no built-in user management, so keep it on your LAN or behind an authenticating reverse proxy.
Related
- Ollama for language models
- whisper.cpp for lightweight transcription
- Piper for standalone text-to-speech
- Whishper for a transcription and subtitle editing UI
License
MIT.