Speaches: An OpenAI-Compatible Speech-to-Text and Text-to-Speech Server

Speaches: An OpenAI-Compatible Speech-to-Text and Text-to-Speech Server

Speaches is a server for speech models that speaks the OpenAI audio API. Point any tool or SDK that supports OpenAI’s transcription or text-to-speech endpoints at your Speaches server, and it works, with the models running on your hardware. The project describes itself as “Ollama, but for TTS/STT models.”

Features

  • Speech-to-text with faster-whisper, including streaming transcription sent as it is processed
  • Translation of speech to English
  • Text-to-speech with Piper and Kokoro, a small model that produces notably natural voices
  • OpenAI API compatibility, so existing clients and libraries work unchanged
  • Dynamic model loading: models load on demand and unload after inactivity to free memory
  • CPU and GPU support

Uses

  • A private backend for voice assistants, including Home Assistant voice pipelines
  • Voice input and read-aloud for chat front ends such as Open WebUI
  • Batch transcription from scripts using the OpenAI SDK
  • Replacing paid transcription APIs in your own projects

Deployment

Docker Compose files are provided for CPU and CUDA setups. Whisper runs acceptably on a modern CPU with small models; larger models and real-time use benefit greatly from an NVIDIA GPU.

Like other model servers, it has no built-in user management, so keep it on your LAN or behind an authenticating reverse proxy.

  • Ollama for language models
  • whisper.cpp for lightweight transcription
  • Piper for standalone text-to-speech
  • Whishper for a transcription and subtitle editing UI

License

MIT.