Try It Out

Generate speech from text using any backend and voice. Also transcribe audio and re-speak it.

Voice & playback

Checking available TTS backends…

Text to synthesize

After changing active voices, restart the TTS container so the engine reads the updated voice folder.

Audio effects

Apply post-processing to the last generated audio. Non-destructive — re-generate to reset.

Generation history

Last 20 generations this session. Audio is preserved in memory until you reload.

No generations yet.

Playlist

Queue clips from history. Reorder by dragging, then export as one merged WAV.

No clips in playlist. Use + Playlist after generating.

STT TTS workspace

Upload speech audio, transcribe it with the configured STT endpoint, then synthesize the resulting text with any available TTS backend.

Source speech

Record from your microphone or upload an audio file, then transcribe it.

Uses Settings → Whisper/STT URL by default.
0:00
Record or upload speech audio, then transcribe it with the selected recognition engine.
No source audio loaded.

Refine with LLM

Clean up the transcription using a local language model — remove fillers, fix repetitions, and normalise punctuation.

Synthesize transcription

Choose a TTS backend and voice, then generate audio from the transcribed text.

Checking available TTS backends...