Settings

Configure backend URLs, API keys, voice folders, and playback preferences.

Settings

Configure the service URLs you use. Advanced payloads, folders, and keys are below.

Local stack
Switch between light and dark interface.

Service URLs for each backend. Change these first when setting up.

TTS — Text to Speech Qwen3 engines: clone, design, custom, streaming · Kokoro FastAPI
Uploaded/cloned WAV voices. Expected: POST /v1/audio/speech.
Prompt-designed voices and vd_... virtual voices.
Style over configured speakers such as Ryan, Vivian, Serena.
Progressive low-latency WAV playback.
OpenAI-compatible TTS. Built-in voices: af_bella, bf_emma, am_adam… ~300 MB RAM.
NVIDIA speech stack router, Magpie TTS, Parakeet ASR, clone NIM
OpenAI-compatible base URL for the NVIDIA speech router.
Direct Magpie endpoint. Fixed speakers, not WAV cloning.
Direct Parakeet endpoint for transcription.
Magpie Zeroshot clone endpoint with audio_prompt.
Magpie Flow clone endpoint with audio_prompt and transcript.
STT — Speech to Text Whisper, faster-whisper, whisper.cpp, Groq, Parakeet, NVIDIA router
Default recognition endpoint. Expected: POST /v1/audio/transcriptions.
GPU-accelerated Whisper via CTranslate2. ~70× RT · 1.5 GB VRAM.
Lightweight C++ Whisper server. ~8–15× RT CPU · ~1 GB RAM.
Used for Groq Whisper STT and LLM. Free: 2 000 req/day. Fastest cloud transcription.

Controls how previews play and how OpenAI-compatible requests are shaped.

Buffered keeps Save WAV available. Streaming starts sooner.
Controls payload shape for the normal TTS API URL.
Payloads JSON extras sent to each backend
Extra fields for the 8020 WAV voice clone/base model.
Extra fields for the 8023 streaming model.
Extra fields for the CustomVoice backend.
Extra fields for Voice Design and virtual vd_... voices.
Usually empty. Magpie accepts fixed speakers: sofia, aria, jason, leo, john.
Optional multipart fields. The app supplies text, language, audio_prompt.
Optional multipart fields. The app also sends the saved reference transcript.
Extra JSON fields for Kokoro. Usually empty — voice is selected from the dropdown.
Storage container paths and Portainer volume mounts
Contains active_voices, hidden_voices, sounds, and metadata.
New cloned/exported voices are saved here.
API Keys usually empty for local containers

Export all voices and settings as a ZIP for backup or migration. Import to restore. API keys are excluded from exports.

Export voices

Recent server activity. Useful for debugging backend connections and API errors.

Click Refresh to load logs
TTS Voice Creator

Clone, design, and deploy custom voices using local AI backends. Compatible with Qwen3-TTS, Kokoro FastAPI, NVIDIA Magpie, and any OpenAI-compatible TTS/STT endpoint.

FastAPI Vanilla JS WaveSurfer.js v7 MDI v7.4.47