JS files are cache-busted by app version (?v=1.3.0), so browsers that served stale v1.2.0 scripts will now fetch the updated voice-clone.js, stt.js, style.css, and nav.js. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
15 KiB
Changelog
All notable changes to TTS Voice Creator — Clone and Design are documented here.
Follows Keep a Changelog · versioned with Semantic Versioning.
Unreleased
[1.3.0] — 2026-05-31
Added
- Live mic monitor in Clone a Voice — level-meter and scrolling oscilloscope waveform in the Microphone card. "Check level" / "Stop monitor" buttons, mic gain slider. Recording uses raw mic constraints (no echo-cancel / AGC).
- STT engine picker in Clone → Transcript — pick any configured STT backend when auto-transcribing, bypassing an unavailable Whisper.
- Better sample texts — all 8 languages rewritten to ~38 words / ~15 s, phonetically rich, proper Unicode diacritics.
Fixed
- Empty "Read aloud" field — sample text now reliably populates on load and when navigating to the Clone section.
- Recording quality —
MediaRecorderrequests 256 kbps in Clone and STT→TTS. - OGG file import — explicit extension list in
accept=.
1.2.0 — 2026-05-29
Added
- Remember last section on reload — the active section (and Settings /
Engines sub-page) is persisted in
localStorage. A hard-reload (Ctrl+Shift+R) now returns to the same page instead of always jumping to My Voices. - Conversation: live speech preview — while recording, the active Whisper STT backend transcribes accumulated audio every 2.5 s and shows the result in the text input field in real time. The input is pre-populated with this live guess before the final Whisper result arrives. Also tries the browser's Web Speech API first (works on HTTPS / localhost) for even faster results.
- Conversation: Voice Activity Detection (VAD) — recording now auto-stops
after 1.5 s of silence detected via the Web Audio
AnalyserNodeRMS level. A "Sending in X.Xs" countdown appears in the status bar so the timing is visible. An Auto-stop toggle in the input bar lets users disable VAD and revert to click-to-stop. A thin audio-level bar below the status line shows microphone volume in real time during recording. - Conversation: hands-free mode — after the agent finishes speaking, the microphone restarts automatically. A Hands-free toggle (on by default) disables this; clicking the mic button manually always cancels any pending auto-restart.
scripts/release.py— automates version bump + CHANGELOG promotion.python scripts/release.py --patch|--minor|--major [--dry-run]renames[Unreleased]to the new version, updates compare links, writesVERSION, commits, and creates an annotated git tag in one command.- Git pre-commit hook (
scripts/hooks/pre-commit) — warns (does not block) when.py/.js/.css/.htmlfiles are staged butCHANGELOG.mdorVERSIONare not. Runbash scripts/install-hooks.shafter cloning. scripts/install-hooks.sh— one-liner to install the hook after a fresh clone:bash scripts/install-hooks.sh.
Changed
- Config and logs are now bind-mounted local folders — replaced the
opaque named Docker volume with
./config/and./logs/host directories.portainer-stack.ymlupdated with absolute host paths. - Server writes a rotating log file —
RotatingFileHandlerwritesINFO-level and above to./logs/app.log(rotates at 5 MB, 3 backups).
Performance
- Skeleton loading view —
index.htmlshows an animated shimmer placeholder immediately on first paint; fades out once JS finishes loading. - Self-hosted WaveSurfer and MDI icon font — removed render-blocking CDN
requests; assets now served locally from
static/vendor/. - Parallel JS module loading — restructured
loader.jsinto 4 ordered batches; round-trips reduced from 17 to 6, 9 files fetched simultaneously. - Version-based JS/CSS cache busting — versioned assets served with
max-age=31536000, immutable; bumping version invalidates the cache.
Fixed
- LLM returned empty response (Qwen3 thinking mode) — conversation turn
now falls back to
reasoning_contentfor think-only responses; error message hints to add/no-thinkto the system prompt. - Engine settings lost after container recreate — container names and URL
overrides now persisted as server settings (
engine_container_names,engine_local_urls); restored from server on first page load. - Text-input turns returned 422 — changed
audioform field toOptional[UploadFile] = Noneso text-only turns don't require audio. - Conversation input bar hidden when mic unavailable — warning box moved
inside
conv-chat-windowso it never pushes the input bar off-screen. - Browser caches old section HTML —
loader.jsappends?v=<timestamp>to every section fetch. - Various import errors and container restart issues fixed.
1.1.0 — 2026-05-29
Security
- Fixed path traversal in
/api/browse-dirs— Added a_BROWSE_BLOCKEDblocklist (/proc,/sys,/dev,/run,/boot). Requests for paths under these directories now return HTTP 403 instead of listing kernel/system files. - Hardened yt-dlp output path — After a YouTube download completes, the resolved
output path is verified to be inside
TEMP_DIRvia.relative_to(). A file written outside the temp directory is rejected with an SSE error event and never registered. - Removed CORS wildcard on
/api/proxy-audio—Access-Control-Allow-Origin: *was unnecessary (all callers are same-origin) and exposed proxied audio to arbitrary cross-origin requests. Header removed. - Temp file registry now enforces a TTL —
_registrychanged todict[str, tuple[Path, float]]._registry_gc()evicts entries older thanTEMP_FILE_TTL_SECONDS(default 2 h, configurable via env var) and unlinks their files, preventing unbounded disk growth on long-running instances.
Performance
- Settings and routing rules cached in memory —
_load_settings()and_load_tts_routes()previously read from disk on every API request (55+ calls per TTS synthesis). Both now use mtime-checked in-memory caches that invalidate automatically on write, eliminating redundant file I/O.
Added
- Version number —
VERSIONfile at repo root; read bycore/constants.__version__and surfaced viaGET /api/version. Displayed asv1.1.0in Settings → About. - Text input in Conversation Playground — a pill-shaped text field and send button
(→) sit left of the mic button. Pressing Enter or → sends text directly through the
LLM → TTS pipeline, skipping STT entirely. Makes the playground fully usable without
a microphone (HTTP context, no mic permission, remote access). The backend
/api/conversation/turnnow accepts an optionaltextform field; when set, the STT step is skipped and the STT latency row shows—. - Container name field on all engine cards — every TTS and STT engine card (Docker stack cards and static "Other Local" cards) now always shows the Docker container name input row. Previously absent/not-installed cards hid it; now it is always visible so the container can be pre-configured before starting.
- Connect / Disconnect toggle — the Connect button now shows "Disconnect" (green,
check-networkicon) when already connected and toggles back on click. State persists inlocalStorage. - Auto-apply on Connect — a successful connection probe automatically saves the URL to Settings and makes the backend available in TTS/STT dropdown menus immediately, without requiring a separate "Use as TTS/STT" click.
Changed
-
Connect button redesigned — moved out of the URL input row into a dedicated
dc-controls-row. Restyled as a solid blue primary CTA (was a small teal outline button). Shows a spinner icon while probing. -
"Use as TTS / STT" button — larger padding, bolder teal border, chevron icon, tooltip explaining it sets the URL in Settings. Gains
.activehighlight once applied. -
Unified controls row on every engine card — consistent left-to-right order:
[Connect/Disconnect][Stop | Start | Restart][Use as →]. Docker action buttons hidden until a container name is entered; Use-as button right-aligned. -
initStaticDockerManagement— rebuilt to use the samedc-controls-rowstructure as the dynamic Docker stack cards. The existing.llm-local-pingbutton is moved from inside the URL row into the controls row at initialisation time. -
Backend refactor —
server.py(5 560 lines → 43 lines) — all logic extracted into single-responsibility modules:Package Module Responsibility core/constants.pyBoot-time env defaults, path constants, version, log buffer registry.pyTTL-based temp file registry validation.pyURL validation, SSRF guard, path safety docker_client.pyRaw Unix-socket Docker HTTP client config.pySettings load/save/normalize, backend URL resolution routing.pyTTS route rules load/save/resolve, language detection audio.pyAudio conversion, normalisation, auto-trim scoring voice.pyVoice metadata, backup management, benchmark helpers presets.pyVoice Design preset load/save, virtual voice resolution tts_helpers.pyTTS request helpers, streaming, per-backend logic routes/admin.pyIndex, favicon, browse-dirs, robots, version settings.py/api/settings, routing rules, logs, design presetslibrary.pyAll voice CRUD, upload, save, normalize, export/import stt.py/api/transcribe*,/api/stt-backendssources.pyVoice scraping, proxy-audio, yt-dlp download docker.py/api/local-containers/*,/api/probe-urltts.pyTTS preview, streaming, voice design, /v1/*, backendsconversation.pyRefine-text, effects, export/import, speak, MCP, conversation Dockerfileupdated withCOPY core/ core/andCOPY routes/ routes/.docker-compose.ymlupdated with./core:/app/core:roand./routes:/app/routes:ro. -
Frontend refactor —
app.js(8 744 lines → 16 modules) — split intostatic/js/withloader.jsloading them sequentially in dependency order:Module Lines Responsibility utils.js364 Core helpers: $,toast,escHtml, theme, language/flag, picker, tabsvoice-inspector.js397 3-pane voice workbench voice-sources.js277 External voice source scraping UI integrations.js211 Code snippet generation (SillyTavern, Open WebUI, HA, curl, MCP) routing.js542 TTS routing rules editor settings.js385 loadSettings,applyAndSaveSettings, settings panelvoice-clone.js774 WaveSurfer, drop zone, mic recording, trim, voice design voice-library.js2654 Full voice library: list, CRUD, benchmark, normalize tts-preview.js528 TTS preview, fetchTtsPreviewBlobbenchmark.js218 Performance + batch benchmark stt.js287 STT→TTS playground, refreshSttBackendsinit.js49 App bootstrap engines.js625 ElevenLabs browser, custom engine cards, Docker management ai-backends.js520 AI backend cards, LLM snippets, initStaticDockerManagementgeneration.js393 WAV merge, chunked TTS, history, playlist, audio effects conversation.js520 Conversation playground, LLM refinement, import, About
Fixed
chrome://flags/…URL unreadable in mic-blocked warning — the globalcode { background: var(--panel) }rule caused the URL text to render as white-on-light-grey inside the red warning box. Fixed with inline styles (background: rgba(0,0,0,.35); color: #fff) on the<code>element, plus a Copy button so users don't need to manually select invisible text.
1.0.0 — 2026-05-28
Initial feature-complete release.
Added
- Voice library — clone voices from audio samples; design voices from text descriptions using instruction-based synthesis; benchmark synthesis speed (RTF); normalize loudness; export/import voice packages as ZIP bundles.
- TTS backends — Qwen3 TTS (Voice Clone, Voice Design, Custom Voice, Streaming), NVIDIA Magpie / Zeroshot / Flow, Kokoro FastAPI, VibeVoice, XTTS v2, ElevenLabs.
- STT backends — OpenAI Whisper (port 8010), faster-whisper-server, whisper.cpp, Groq Whisper (cloud, free tier), NVIDIA Parakeet ASR. Real transcription probe in health check (not just TCP reachability).
- App Routing — per-app / per-voice / per-language TTS routing rules with automatic language detection and optional before/after sound effects.
- Conversation Playground — full STT → LLM → TTS pipeline with real-time SSE streaming, latency stats panel (STT / LLM TTFT / LLM total / TTS / Total), turn history, system prompt, and insecure-context warning.
- Engines section — LLM / STT / TTS sub-pages; Docker container management (Start / Stop / Restart via Docker socket); custom engine cards; ElevenLabs voice library browser.
- Performance Benchmark — single-voice and batch benchmark with RTF tracking, sparkline trend, and persistent history.
- Audio effects — reverb, chorus, delay, compressor, gain, pitch shift
(via
pedalboard). - Chunked TTS + generation history — long-text synthesis split into chunks, per-chunk playback, playlist export as WAV.
- MCP server — built-in JSON-RPC 2.0 endpoint at
/mcp; tools:speak,transcribe,list_captures,list_profiles. - LLM refinement & persona rewriting — clean up STT transcripts or rewrite responses with a chosen persona via any OpenAI-compatible LLM endpoint.
- Connect Apps — ready-made config snippets for SillyTavern, Open WebUI,
Home Assistant, curl, and MCP (
claude mcp addone-liner). - Voice sources — scrape voice assets from Aiartes, Freesound, GitHub, and Google Drive; YouTube download via yt-dlp; quick import directly to library.
- OpenAI-compatible proxy —
/v1/audio/speechand/v1/audio/transcriptionsfor drop-in use with Open WebUI, SillyTavern, and Home Assistant. - Settings — sub-pages: General, Connections, Playback, Captures, Payloads, Storage, API Keys, Logs, About.
- Voice Design presets — saved persona templates for instruction-based synthesis;
virtual
vd_…voices usable from external apps without exporting WAV files. - Multilingual support — language/flag pickers, per-language preview texts,
LANG_FLAG_DEFAULTmapping for 16 languages. - Tags, ratings, and metadata — per-voice tags with autocomplete, star ratings, gender label, country flag.
- Dark/light theme — toggle with persistence in
localStorage. - Docker socket integration — Start/Stop/Restart Docker containers from the UI via raw Unix socket HTTP; container health visible in engine cards.