The VoiceClone server was using non_streaming_mode=False, a mode designed for streaming LLM->TTS pipelines. In that mode only one text token enters the model's KV cache during prefill; the rest feed via trailing_text_hiddens at one step per codec frame. For a 54-word paragraph this provides only ~4s of text guidance for ~18s of speech — 77% generated with no text conditioning. Without text context the model free-runs and drifts, sometimes changing gender. Fix: switch to non_streaming_mode=True (already the default for VoiceDesign and CustomVoice) so the full text is in the prefill throughout generation. Also lower default temperature 0.9->0.8 and add top_p=0.9 to reduce accumulated sampling noise over long runs. Temperature, top_k, and top_p are now configurable per voice in voices.json. - patches/openai_server.patch: updated for new upstream HEAD; both streaming (WAV/PCM) and non-streaming (MP3) paths now use non_streaming_mode=True - config/run_server.py: align warmup call to non_streaming_mode=True - README.md: bump image tags v4->v5, add changelog section - Version: v5 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| speakers | ||
| auto_transcribe.py | ||
| benchmark_api.py | ||
| customvoice_voices.json | ||
| docker-compose.yml | ||
| generate_voices.py | ||
| run_customvoice_server.py | ||
| run_server.py | ||
| run_voicedesign_server.py | ||
| voicedesign_voices.json | ||