Generate speech from text using any backend and voice. Also transcribe audio and re-speak it.
Pick any reachable TTS backend, fetch its voices, then synthesize text. WAV/NVIDIA clone backends preserve reference identity; instruction-control backends follow style better.
Select a backend, fetch its voice list, then synthesize any text with optional style instruction.
After changing active voices, restart the TTS container so the engine reads the updated voice folder.
instruct. Voice Clone/Base and Streaming are fastest; CustomVoice and Voice Design are style-aware.
Apply post-processing to the last generated audio. Non-destructive — re-generate to reset.
Last 20 generations this session. Audio is preserved in memory until you reload.
Queue clips from history. Reorder by dragging, then export as one merged WAV.
Upload speech audio, transcribe it with the configured STT endpoint, then synthesize the resulting text with any available TTS backend.
Record from your microphone or upload an audio file, then transcribe it.
Clean up the transcription using a local language model — remove fillers, fix repetitions, and normalise punctuation.
Choose a TTS backend and voice, then generate audio from the transcribed text.
Measure synthesis latency and real-time factor for any backend and voice.