Skip to main content
Home Skills Text-to-Speech

Text-to-Speech

Audio Processing

Convert text to natural-sounding speech, voice generation

text to speechTTSvoice generationspeech synthesisaudio narrationvoice overread aloud
0
Curated Registries
29
Available Models

Available Models & Agents

Browse 29 models and agents with text-to-speech capabilities

Search all

mayaresearch/veena-max

mayaresearch

VeenaMAX is a high-performance, advanced Text-to-Speech API that converts text into natural-sounding speech with emotional intelligence and blazing-fast response times. It can speak English and Hindi - the most widely used languages in India. As a signifi

model

inworld/realtime-tts-2

inworld

Most expressive text-to-speech model from Inworld, with natural-language steering, real-time latency, and multilingual support across 100+ languages.

model

inworld/realtime-tts-1.5-max

inworld

Highest-quality realtime text-to-speech with <200ms latency, emotion control, and 15-language support

model

inworld/realtime-tts-1.5-mini

inworld

Ultra-fast, cost-efficient realtime text-to-speech with ~120ms latency and 15-language support

model

xai/grok-text-to-speech

xai

Convert text to natural-sounding speech with xAI's Grok TTS. 5 voices, 20 languages, expressive speech tags, and high-fidelity MP3 / WAV / telephony audio output.

model

google/gemini-3.1-flash-tts

google

Google's fast, expressive text-to-speech model with 30 voices and 70+ language support

model

adirik/hierspeechpp

adirik

Zero-shot speech synthesizer for text-to-speech and voice conversion

model

lucataco/whisperspeech-small

lucataco

An Open Source text-to-speech system built by inverting Whisper

model

cjwbw/melotts

cjwbw

High-quality multilingual text-to-speech library

model

cjwbw/parler-tts

cjwbw

lightweight text-to-speech (TTS) model, trained on 10.5K hours of audio data

model

zsxkib/hololive-style-bert-vits2

zsxkib

🎙️Hololive text-to-speech and voice-to-voice (Japanese🇯🇵 + English🇬🇧)

model

lee101/guided-text-to-speech

lee101

Guided Text to Speech Generator

model

e1100x/chattts

e1100x

ChatTTS is a text-to-speech model designed specifically for dialogue scenarios such as LLM assistant.

model

jaaari/kokoro-82m

jaaari

Kokoro v1.0 - text-to-speech (82M params, based on StyleTTS2)

model

alphanumericuser/kokoro-82m

alphanumericuser

Kokoro v1.0 - text-to-speech (82M params, based on StyleTTS2)

model

cuuupid/zonos

cuuupid

Zonos-v0.1 beta, a SOTA text-to-speech Transformer model with extraordinary expressive range, built by Zyphra.

model

lucataco/step-audio-tts-3b

lucataco

Step-Audio-TTS-3B represents the industry's first Text-to-Speech (TTS) model trained on a large-scale synthetic dataset utilizing the LLM-Chat paradigm

model

cjwbw/voicecraft

cjwbw

Zero-Shot Speech Editing and Text-to-Speech in the Wild

model

lucataco/higgs-audio-v2

lucataco

Higgs Audio v2, a powerful text-to-speech audio foundation model that excels in expressive audio generation

model

microsoft/vibevoice

microsoft

Microsoft's VibeVoice text-to-speech model that can generate long-form speech from text with sample voices.

model