218 lines
6.0 KiB
Markdown
218 lines
6.0 KiB
Markdown
# TTS Voice Creator + Qwen3-TTS Setup Playbook
|
|
|
|
This playbook is for people who want a friendly voice-cloning and TTS-routing setup without living in a terminal.
|
|
|
|
## What This Gives You
|
|
|
|
You get a browser app called **TTS Voice Creator - Clone and Design** that can:
|
|
|
|
- Record or import voice samples.
|
|
- Trim, normalize, transcribe, and save cloned voices.
|
|
- Create prompt-based Voice Design voices.
|
|
- Route simple app voices such as `default` to different real voices by language.
|
|
- Act as an OpenAI-compatible TTS endpoint for Open WebUI, SillyTavern, Home Assistant, and curl.
|
|
|
|
The usual Open WebUI setup becomes:
|
|
|
|
```text
|
|
TTS API base URL: http://YOUR_HOST:7860/v1
|
|
Model: tts-1
|
|
Voice: default
|
|
```
|
|
|
|
Then the Routing tab decides what `default` means.
|
|
|
|
## The Pieces
|
|
|
|
### Docker
|
|
|
|
Docker runs each part in its own container. That keeps the app, Qwen3-TTS, Voice Design, and optional tools separated and easier to update.
|
|
|
|
### Portainer
|
|
|
|
Portainer is optional. It is a browser dashboard for Docker. If your target user dislikes command lines, Portainer is worth installing because it gives them buttons for logs, restarts, and container status.
|
|
|
|
### TTS Voice Creator
|
|
|
|
This is the voice library and routing interface. It stores reference voices, metadata, transcripts, ratings, active/hidden state, and routing rules.
|
|
|
|
### Qwen3-TTS Voice Clone
|
|
|
|
This is the normal production TTS path. It uses reference audio from exported cloned voices.
|
|
|
|
Use this when:
|
|
|
|
- You recorded a real person or character voice.
|
|
- You want Open WebUI to speak using saved voices.
|
|
- You want the most predictable day-to-day TTS endpoint.
|
|
|
|
### Qwen3-TTS Voice Design
|
|
|
|
This creates a voice from a written description. No reference recording is required.
|
|
|
|
Use this when:
|
|
|
|
- You are designing a new narrator or character.
|
|
- You want to prototype voices before recording anyone.
|
|
- You want virtual voices through the app proxy.
|
|
|
|
### Qwen3-TTS Custom Voice
|
|
|
|
This is for Qwen CustomVoice workflows with named speakers and optional instructions.
|
|
|
|
Use this when:
|
|
|
|
- You already maintain custom speaker IDs.
|
|
- You want to expose Qwen custom speakers as OpenAI-compatible voices.
|
|
- You need a separate custom voice server from the cloned voice library.
|
|
|
|
## Guided Installation
|
|
|
|
Start the installer:
|
|
|
|
```bash
|
|
python3 guided_installer.py
|
|
```
|
|
|
|
On desktop systems with Python file association, you can often double-click `guided_installer.py`.
|
|
|
|
The installer checks:
|
|
|
|
- Docker command
|
|
- Docker daemon
|
|
- Docker Compose plugin
|
|
- Portainer
|
|
- TTS Voice Creator
|
|
- Qwen3-TTS clone server
|
|
- Qwen3-TTS Voice Design server
|
|
- Qwen3-TTS Custom Voice server
|
|
|
|
It then asks for:
|
|
|
|
- Voice library folder
|
|
- Model folder
|
|
- App port
|
|
- Docker network name
|
|
- Qwen Docker image
|
|
- Which Qwen model servers to install
|
|
- Ports for each model server
|
|
- Model locations for each server
|
|
- Optional external TTS endpoint
|
|
|
|
## Recommended Ports
|
|
|
|
```text
|
|
7860 TTS Voice Creator
|
|
8020 Qwen3-TTS Voice Clone
|
|
8021 Qwen3-TTS Voice Design
|
|
8022 Qwen3-TTS Custom Voice
|
|
9000 Portainer HTTP
|
|
9443 Portainer HTTPS
|
|
```
|
|
|
|
## Installation Flow
|
|
|
|
1. Open the guided installer.
|
|
2. Click **Re-check this computer**.
|
|
3. If Docker is missing, open the Docker install guide from the installer.
|
|
4. Optionally click **Install Portainer**.
|
|
5. Choose which Qwen model servers to install.
|
|
6. Pick where the large model folders live.
|
|
7. Pick the ports.
|
|
8. Click **Generate compose file**.
|
|
9. Click **Install / update selected stack**.
|
|
10. Open `http://YOUR_HOST:7860`.
|
|
|
|
## First App Setup
|
|
|
|
Open the Settings gear in TTS Voice Creator:
|
|
|
|
```text
|
|
TTS API URL: http://host.docker.internal:8020
|
|
Voice Design URL: http://host.docker.internal:8021
|
|
Voice scan directory: /voices
|
|
Active voices dir: /voices/active_voices
|
|
```
|
|
|
|
If the app and TTS servers are on the same Docker network, container names can also be used from inside Docker.
|
|
|
|
## Recording A Voice
|
|
|
|
1. Open **Voice Clone**.
|
|
2. Click **Add new voice**.
|
|
3. Choose the read-sample language.
|
|
4. Click **Check level**.
|
|
5. Adjust mic gain until peaks are roughly `-18` to `-9 dB`.
|
|
6. Click **Record**.
|
|
7. Stop after 10 to 20 seconds.
|
|
8. Auto trim if needed.
|
|
9. Use or edit the transcript.
|
|
10. Save the voice.
|
|
|
|
After saving voices, restart the Qwen3-TTS clone container so it scans `active_voices`.
|
|
|
|
## Open WebUI Without Routing
|
|
|
|
Use this when you always want one explicit voice.
|
|
|
|
```text
|
|
Provider: OpenAI-compatible
|
|
Base URL: http://YOUR_HOST:8020/v1
|
|
API key: dummy
|
|
Model: tts-1
|
|
Voice: EN_F_YourVoice
|
|
```
|
|
|
|
## Open WebUI With Routing
|
|
|
|
Use this when Open WebUI sends `default`, and TTS Voice Creator decides which real voice to use.
|
|
|
|
```text
|
|
Provider: OpenAI-compatible
|
|
Base URL: http://YOUR_HOST:7860/v1
|
|
API key: dummy
|
|
Model: tts-1
|
|
Voice: default
|
|
```
|
|
|
|
Then add routes:
|
|
|
|
```text
|
|
Open WebUI + default + EN -> EN_F_YourEnglishVoice
|
|
Open WebUI + default + DE -> DE_M_YourGermanVoice
|
|
Open WebUI + default + FR -> FR_F_YourFrenchVoice
|
|
```
|
|
|
|
Now English text, German text, and French text can each use a different voice while Open WebUI still only knows about `default`.
|
|
|
|
## Troubleshooting
|
|
|
|
### Microphone Is Blocked
|
|
|
|
Browsers usually require `localhost` or HTTPS for microphone access.
|
|
|
|
- Chrome, Brave, Edge: click the lock/tune icon in the address bar and allow Microphone.
|
|
- Firefox: click the microphone or lock icon and remove the block.
|
|
- Safari: Safari Settings, Websites, Microphone, allow this site.
|
|
|
|
### Voice Does Not Appear In Qwen3-TTS
|
|
|
|
Restart the Qwen3-TTS clone container. Most TTS servers scan voice folders at startup.
|
|
|
|
### Open WebUI Gets No Audio
|
|
|
|
Check:
|
|
|
|
- Base URL ends in `/v1`.
|
|
- Voice exists in `/v1/audio/voices`.
|
|
- API key is set to `dummy` if Open WebUI requires a key.
|
|
- The TTS container logs do not show model-load errors.
|
|
|
|
### Routing Picks The Wrong Language
|
|
|
|
The router uses lightweight language detection. If a sentence is short or mixed-language, add a fallback `*` route or send a more explicit text sample.
|
|
|
|
## Why This Setup Is Nice
|
|
|
|
It separates voice management from the TTS engine. The app becomes the place where non-technical users can record, clean up, label, activate, route, and test voices, while Docker keeps the heavier services running in the background.
|