Try It Out

Generate speech from text using any backend and voice. Also transcribe audio and re-speak it.

TTS generation playground

Pick any reachable TTS backend, fetch its voices, then synthesize text. WAV/NVIDIA clone backends preserve reference identity; instruction-control backends follow style better.

Generate speech

Checking available TTS backends...

After changing active voices, restart the TTS container so the engine reads the updated voice folder.

This is sent as instruct. Voice Clone/Base and Streaming are fastest; CustomVoice and Voice Design are style-aware.

STT → TTS workspace

Upload speech audio, transcribe it with the configured STT endpoint, then synthesize the resulting text with any available TTS backend.

Source speech

Uses Settings → Whisper/STT URL by default.
0:00
Record or upload speech audio, then transcribe it with the selected recognition engine.
No source audio loaded.

Synthesize transcription

Checking available TTS backends...