Voicebox
Open-source local-first AI voice studio for voice cloning, text-to-speech, dictation, transcription, and agent voice output.
Open official linkOverview
Category: Local AI voice studio
Open-source local-first AI voice studio for voice cloning, text-to-speech, dictation, transcription, and agent voice output.
Best for
Users and developers who want private local voice generation, dictation, and voice workflows without a fully hosted voice platform.
Use cases
- Clone voices from short samples
- Generate speech and multi-voice audio
- Dictate into apps or give agents a voice
Common example
Create a local voice profile, generate narration from text, and let an MCP-aware coding agent speak status updates.
Pricing and free plan
Pricing model: Open-source/local app positioning with optional infrastructure or remote compute needs; verify current release and donation/support model.
Free plan / trial assessment: Local use can be free, but quality, speed, and model choices depend on hardware and setup.
Limitations
Voice cloning raises consent and misuse risks; local inference setup and hardware requirements may be technical.
ChatGPT / Claude comparison
Better than ChatGPT/Claude for this task - Voicebox produces and manages local voice audio, while ChatGPT/Claude mainly write text.
Alternatives
ElevenLabs, Wispr Flow, Whisper, Kokoro