TTS generation playground
Pick any reachable TTS backend, fetch its voices, then synthesize text. WAV/NVIDIA clone backends preserve reference identity; instruction-control backends follow style better.
Generate speech
Select a backend, fetch its voice list, then synthesize any text with optional style instruction.
Checking available TTS backends...
After changing active voices, restart the TTS container so the engine reads the updated voice folder.
This is sent as instruct. Voice Clone/Base and Streaming are fastest; CustomVoice and Voice Design are style-aware.
STT → TTS workspace
Upload speech audio, transcribe it with the configured STT endpoint, then synthesize the resulting text with any available TTS backend.
Source speech
Record from your microphone or upload an audio file, then transcribe it.
Uses Settings → Whisper/STT URL by default.
No source audio loaded.
Synthesize transcription
Choose a TTS backend and voice, then generate audio from the transcribed text.
Checking available TTS backends...
Performance benchmark
Measure synthesis latency and real-time factor for any backend and voice.
Results
Latency per run. RTF = synthesis time / audio duration (lower is better).
| # | Backend | Voice |
Latency (ms) | Audio (s) | RTF | Status |