Open-source alternatives to Azure AI Speech TTS

Azure AI Speech TTS in the AI category. These are the open alternatives we recommend looking at.

Piper

Fast local text to speech that runs well on small computers such as a Raspberry Pi.

GPL-3.0Strong copyleft

SaaS: YesClosed product: No

GitHub stars
5.7k
Deployment
Library

GPT-SoVITS

Voice cloning and text-to-speech tool that can be trained on a short voice recording.

MITPermissive

SaaS: YesClosed product: Yes

GitHub stars
62.3k
Deployment
Desktop

ChatTTS

Text-to-speech model tuned for everyday dialogue, mainly Chinese and English.

AGPL-3.0Network copyleft (AGPL)

SaaS: Yes, with conditionsClosed product: No

GitHub stars
39.9k
Deployment
Library

Chatterbox

Text-to-speech model from Resemble AI with voice cloning and adjustable expressiveness.

MITPermissive

SaaS: YesClosed product: Yes

GitHub stars
26.7k
Deployment
Library

IndexTTS

Text-to-speech system that can mimic a voice from a short sample and control pacing and expression.

SaaS: Yes, with conditionsClosed product: Yes, with conditions

GitHub stars
24.3k
Deployment
Library

CosyVoice

Multilingual text-to-speech model with code for inference, training and deployment.

Apache-2.0Permissive

SaaS: YesClosed product: Yes

GitHub stars
23.8k
Deployment
Library

F5-TTS

Research code for the F5-TTS text-to-speech model, which can mimic a voice from a short sample.

MITPermissive

SaaS: YesClosed product: Yes

GitHub stars
15.3k
Deployment
Library

Coqui TTS (Idiap fork)

Maintained continuation of Coqui TTS, a deep-learning toolkit for text to speech.

MPL-2.0Weak copyleft

SaaS: YesClosed product: Yes, with conditions

GitHub stars
2.3k
Deployment
Library

Kokoro-82M

Very small text-to-speech model that produces natural speech and runs on an ordinary computer. Based on yl4579/StyleTTS2-LJSpeech.

Apache-2.0Permissive

SaaS: YesClosed product: Yes

Dia-1.6B

Text-to-speech model from Nari Labs that generates multi-speaker dialogue from a script.

Apache-2.0Permissive

SaaS: YesClosed product: Yes

VibeVoice-1.5B

Text-to-speech model from Microsoft for long multi-speaker conversations, such as podcasts.

MITPermissive

SaaS: YesClosed product: Yes

Deployment
transformers

csm-1b

Speech model from Sesame that generates speech from text and conversational context.

Apache-2.0Permissive

SaaS: YesClosed product: Yes

Deployment
transformers

chatterbox

Text-to-speech model from Resemble AI with voice cloning and adjustable expressiveness. The model card lists Swedish support.

MITPermissive

SaaS: YesClosed product: Yes

Deployment
chatterbox

piper-voices

Collection of ready-made voices for the Piper text-to-speech engine, in many languages. The model card lists Swedish support.

MITPermissive

SaaS: YesClosed product: Yes

This is guidance, not legal advice (inte juridisk rådgivning). Check with a lawyer before deciding.

Sign in for more

  • Full licence analysis per use case (self-hosting, modifying, hosted service, AI use)
  • Save products to your own lists
  • Systemkarta: map the tools you already use (coming soon)
  • Personal help and AI advice (coming soon)
Sign in for free