tts-voice-creator-clone-and.../CHANGELOG.md
mARTin-B78 db6302f99d Add version number (v1.1.0) and rewrite CHANGELOG to full Keep-a-Changelog spec
- Add VERSION file (1.1.0) at repo root
- core/constants.py: expose __version__ read from VERSION file
- routes/admin.py: GET /api/version endpoint returns {version}
- Settings → About: display "v1.1.0" next to app name via /api/version fetch
- CHANGELOG.md: full rewrite following Keep a Changelog + Semantic Versioning
  - [Unreleased] staging section at top
  - [1.1.0] 2026-05-29 — security, perf, refactor, UX changes from this session
  - [1.0.0] 2026-05-28 — all pre-session features documented
  - Compare links at bottom pointing to GitHub

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-29 13:53:31 +02:00

11 KiB

Changelog

All notable changes to TTS Voice Creator — Clone and Design are documented here.
Follows Keep a Changelog · versioned with Semantic Versioning.


Unreleased


1.1.0 — 2026-05-29

Security

  • Fixed path traversal in /api/browse-dirs — Added a _BROWSE_BLOCKED blocklist (/proc, /sys, /dev, /run, /boot). Requests for paths under these directories now return HTTP 403 instead of listing kernel/system files.
  • Hardened yt-dlp output path — After a YouTube download completes, the resolved output path is verified to be inside TEMP_DIR via .relative_to(). A file written outside the temp directory is rejected with an SSE error event and never registered.
  • Removed CORS wildcard on /api/proxy-audioAccess-Control-Allow-Origin: * was unnecessary (all callers are same-origin) and exposed proxied audio to arbitrary cross-origin requests. Header removed.
  • Temp file registry now enforces a TTL_registry changed to dict[str, tuple[Path, float]]. _registry_gc() evicts entries older than TEMP_FILE_TTL_SECONDS (default 2 h, configurable via env var) and unlinks their files, preventing unbounded disk growth on long-running instances.

Performance

  • Settings and routing rules cached in memory_load_settings() and _load_tts_routes() previously read from disk on every API request (55+ calls per TTS synthesis). Both now use mtime-checked in-memory caches that invalidate automatically on write, eliminating redundant file I/O.

Added

  • Version numberVERSION file at repo root; read by core/constants.__version__ and surfaced via GET /api/version. Displayed as v1.1.0 in Settings → About.
  • Text input in Conversation Playground — a pill-shaped text field and send button (→) sit left of the mic button. Pressing Enter or → sends text directly through the LLM → TTS pipeline, skipping STT entirely. Makes the playground fully usable without a microphone (HTTP context, no mic permission, remote access). The backend /api/conversation/turn now accepts an optional text form field; when set, the STT step is skipped and the STT latency row shows .
  • Container name field on all engine cards — every TTS and STT engine card (Docker stack cards and static "Other Local" cards) now always shows the Docker container name input row. Previously absent/not-installed cards hid it; now it is always visible so the container can be pre-configured before starting.
  • Connect / Disconnect toggle — the Connect button now shows "Disconnect" (green, check-network icon) when already connected and toggles back on click. State persists in localStorage.
  • Auto-apply on Connect — a successful connection probe automatically saves the URL to Settings and makes the backend available in TTS/STT dropdown menus immediately, without requiring a separate "Use as TTS/STT" click.

Changed

  • Connect button redesigned — moved out of the URL input row into a dedicated dc-controls-row. Restyled as a solid blue primary CTA (was a small teal outline button). Shows a spinner icon while probing.

  • "Use as TTS / STT" button — larger padding, bolder teal border, chevron icon, tooltip explaining it sets the URL in Settings. Gains .active highlight once applied.

  • Unified controls row on every engine card — consistent left-to-right order: [Connect/Disconnect] [Stop | Start | Restart] [Use as →]. Docker action buttons hidden until a container name is entered; Use-as button right-aligned.

  • initStaticDockerManagement — rebuilt to use the same dc-controls-row structure as the dynamic Docker stack cards. The existing .llm-local-ping button is moved from inside the URL row into the controls row at initialisation time.

  • Backend refactor — server.py (5 560 lines → 43 lines) — all logic extracted into single-responsibility modules:

    Package Module Responsibility
    core/ constants.py Boot-time env defaults, path constants, version, log buffer
    registry.py TTL-based temp file registry
    validation.py URL validation, SSRF guard, path safety
    docker_client.py Raw Unix-socket Docker HTTP client
    config.py Settings load/save/normalize, backend URL resolution
    routing.py TTS route rules load/save/resolve, language detection
    audio.py Audio conversion, normalisation, auto-trim scoring
    voice.py Voice metadata, backup management, benchmark helpers
    presets.py Voice Design preset load/save, virtual voice resolution
    tts_helpers.py TTS request helpers, streaming, per-backend logic
    routes/ admin.py Index, favicon, browse-dirs, robots, version
    settings.py /api/settings, routing rules, logs, design presets
    library.py All voice CRUD, upload, save, normalize, export/import
    stt.py /api/transcribe*, /api/stt-backends
    sources.py Voice scraping, proxy-audio, yt-dlp download
    docker.py /api/local-containers/*, /api/probe-url
    tts.py TTS preview, streaming, voice design, /v1/*, backends
    conversation.py Refine-text, effects, export/import, speak, MCP, conversation

    Dockerfile updated with COPY core/ core/ and COPY routes/ routes/. docker-compose.yml updated with ./core:/app/core:ro and ./routes:/app/routes:ro.

  • Frontend refactor — app.js (8 744 lines → 16 modules) — split into static/js/ with loader.js loading them sequentially in dependency order:

    Module Lines Responsibility
    utils.js 364 Core helpers: $, toast, escHtml, theme, language/flag, picker, tabs
    voice-inspector.js 397 3-pane voice workbench
    voice-sources.js 277 External voice source scraping UI
    integrations.js 211 Code snippet generation (SillyTavern, Open WebUI, HA, curl, MCP)
    routing.js 542 TTS routing rules editor
    settings.js 385 loadSettings, applyAndSaveSettings, settings panel
    voice-clone.js 774 WaveSurfer, drop zone, mic recording, trim, voice design
    voice-library.js 2654 Full voice library: list, CRUD, benchmark, normalize
    tts-preview.js 528 TTS preview, fetchTtsPreviewBlob
    benchmark.js 218 Performance + batch benchmark
    stt.js 287 STT→TTS playground, refreshSttBackends
    init.js 49 App bootstrap
    engines.js 625 ElevenLabs browser, custom engine cards, Docker management
    ai-backends.js 520 AI backend cards, LLM snippets, initStaticDockerManagement
    generation.js 393 WAV merge, chunked TTS, history, playlist, audio effects
    conversation.js 520 Conversation playground, LLM refinement, import, About

Fixed

  • chrome://flags/… URL unreadable in mic-blocked warning — the global code { background: var(--panel) } rule caused the URL text to render as white-on-light-grey inside the red warning box. Fixed with inline styles (background: rgba(0,0,0,.35); color: #fff) on the <code> element, plus a Copy button so users don't need to manually select invisible text.

1.0.0 — 2026-05-28

Initial feature-complete release.

Added

  • Voice library — clone voices from audio samples; design voices from text descriptions using instruction-based synthesis; benchmark synthesis speed (RTF); normalize loudness; export/import voice packages as ZIP bundles.
  • TTS backends — Qwen3 TTS (Voice Clone, Voice Design, Custom Voice, Streaming), NVIDIA Magpie / Zeroshot / Flow, Kokoro FastAPI, VibeVoice, XTTS v2, ElevenLabs.
  • STT backends — OpenAI Whisper (port 8010), faster-whisper-server, whisper.cpp, Groq Whisper (cloud, free tier), NVIDIA Parakeet ASR. Real transcription probe in health check (not just TCP reachability).
  • App Routing — per-app / per-voice / per-language TTS routing rules with automatic language detection and optional before/after sound effects.
  • Conversation Playground — full STT → LLM → TTS pipeline with real-time SSE streaming, latency stats panel (STT / LLM TTFT / LLM total / TTS / Total), turn history, system prompt, and insecure-context warning.
  • Engines section — LLM / STT / TTS sub-pages; Docker container management (Start / Stop / Restart via Docker socket); custom engine cards; ElevenLabs voice library browser.
  • Performance Benchmark — single-voice and batch benchmark with RTF tracking, sparkline trend, and persistent history.
  • Audio effects — reverb, chorus, delay, compressor, gain, pitch shift (via pedalboard).
  • Chunked TTS + generation history — long-text synthesis split into chunks, per-chunk playback, playlist export as WAV.
  • MCP server — built-in JSON-RPC 2.0 endpoint at /mcp; tools: speak, transcribe, list_captures, list_profiles.
  • LLM refinement & persona rewriting — clean up STT transcripts or rewrite responses with a chosen persona via any OpenAI-compatible LLM endpoint.
  • Connect Apps — ready-made config snippets for SillyTavern, Open WebUI, Home Assistant, curl, and MCP (claude mcp add one-liner).
  • Voice sources — scrape voice assets from Aiartes, Freesound, GitHub, and Google Drive; YouTube download via yt-dlp; quick import directly to library.
  • OpenAI-compatible proxy/v1/audio/speech and /v1/audio/transcriptions for drop-in use with Open WebUI, SillyTavern, and Home Assistant.
  • Settings — sub-pages: General, Connections, Playback, Captures, Payloads, Storage, API Keys, Logs, About.
  • Voice Design presets — saved persona templates for instruction-based synthesis; virtual vd_… voices usable from external apps without exporting WAV files.
  • Multilingual support — language/flag pickers, per-language preview texts, LANG_FLAG_DEFAULT mapping for 16 languages.
  • Tags, ratings, and metadata — per-voice tags with autocomplete, star ratings, gender label, country flag.
  • Dark/light theme — toggle with persistence in localStorage.
  • Docker socket integration — Start/Stop/Restart Docker containers from the UI via raw Unix socket HTTP; container health visible in engine cards.