Mistral AI launches Voxtral TTS, a multilingual speech synthesis model
Mistral
ElevenLabs
Mistral AI has released Voxtral TTS, a 4-billion-parameter text-to-speech model that delivers realistic, emotionally expressive speech in 9 languages with low latency and easy voice adaptation. It is available via API and Mistral Studio, starting at $0.016 per 1k characters.
Mistral AI announced Voxtral TTS, its first text-to-speech model with 4B parameters, offering state-of-the-art multilingual voice generation. The model supports 9 languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, and Arabic. It features low latency (70ms model latency), zero-shot cross-lingual voice adaptation, and contextual understanding for emotional expressiveness. Voice adaptation requires as little as 3 seconds of reference audio. Human evaluations showed superior naturalness compared to ElevenLabs Flash v2.5 while matching ElevenLabs v3 quality. The architecture includes a 3.4B transformer decoder backbone, 390M flow-matching acoustic transformer, and 300M neural audio codec. Available via API at $0.016 per 1k characters and in Mistral Studio. Part of the model weights are open on Hugging Face under CC BY NC 4.0.
- Сокращения
- TTS = Text-to-Speech — синтез речи
- API = Application Programming Interface — программный интерфейс
- RTF = Real-Time Factor — коэффициент реального времени
- NFE = Number of Function Evaluations — количество вычислений функций
- VQ = Vector Quantization — векторное квантование
- FSQ = Finite Scalar Quantization — конечное скалярное квантование
Source: Mistral AI —
original
