Perceived startup speed: - index.html: animated shimmer skeleton (header + toolbar + 8 voice cards) visible immediately; fades out when loader.js finishes - WaveSurfer (57KB) and MDI icon font (394KB woff2) now served from static/vendor/ — removes 3 render-blocking external requests from <head> - Flag-icons CSS loaded async (rel=preload onload trick) — non-blocking loader.js: - Fade out skeleton + remove from DOM (300ms transition) - Reveal page-sections after JS finishes loading Conversation playground: - Qwen3 thinking mode fix: fall back to delta.reasoning_content when delta.content is empty so think-only LLM turns produce visible output - Better error message with /no-think hint when LLM returns empty Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
285 lines
16 KiB
Markdown
285 lines
16 KiB
Markdown
# Changelog
|
|
|
|
All notable changes to **TTS Voice Creator — Clone and Design** are documented here.
|
|
Follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/) · versioned with [Semantic Versioning](https://semver.org/).
|
|
|
|
---
|
|
|
|
## [Unreleased]
|
|
|
|
### Performance
|
|
|
|
- **Self-hosted WaveSurfer and MDI icon font** — removed two render-blocking
|
|
`<script>` and one blocking `<link>` to `unpkg.com` / `cdn.jsdelivr.net`
|
|
from `<head>`. WaveSurfer (~57 KB) and MDI font (~394 KB WOFF2 + CSS) are
|
|
now served locally from `static/vendor/`. Flag-icons CSS is loaded async
|
|
(non-blocking) via the `rel=preload` / `onload` trick.
|
|
- **Skeleton loading view** — `index.html` shows an animated shimmer
|
|
placeholder (page head + toolbar + 8 voice card outlines) immediately on
|
|
first paint, before any JS loads. The skeleton fades out once `loader.js`
|
|
finishes and real sections are revealed. No external deps — pure inline CSS.
|
|
|
|
### Fixed
|
|
|
|
- **LLM returned empty response (Qwen3 thinking mode)** — Qwen3 models
|
|
stream thinking tokens under `delta.reasoning_content` instead of
|
|
`delta.content`. The conversation turn now falls back to
|
|
`reasoning_content` so think-only responses produce visible output.
|
|
Error message improved with a hint to add `/no-think` to the system prompt.
|
|
|
|
### Performance
|
|
|
|
- **Parallel JS module loading** — `loader.js` previously loaded all 16
|
|
modules sequentially (17 round-trips). Restructured into 4 ordered batches
|
|
with `async=false` so files are fetched in parallel but execute in the
|
|
correct dependency order:
|
|
`utils` → `settings` → *(9 feature modules in parallel)* → `init` →
|
|
*(4 post-init modules in parallel)* → `nav`
|
|
Round-trips reduced from 17 to 6; 9 files now download simultaneously.
|
|
- **Version-based JS/CSS cache busting** — `loader.js` fetches
|
|
`/api/version` first and appends `?v=<version>` to every script URL.
|
|
A new `Cache-Control: public, max-age=31536000, immutable` middleware in
|
|
`server.py` lets the browser cache versioned assets for a full year.
|
|
Bumping the version (via `scripts/release.py`) invalidates the cache.
|
|
Section HTML keeps `?v=<timestamp>` (no-store) so it is always fresh.
|
|
|
|
### Changed
|
|
|
|
- **Config and logs are now bind-mounted local folders** — replaced the
|
|
opaque named Docker volume (`tts-voice-creator-clone-and-design-2`) with
|
|
two transparent host directories:
|
|
- `./config/` → `/home/app/.config/tts-voice-creator` — holds
|
|
`settings.json`, `voice_design_presets.json`, `tts_routes.json`
|
|
- `./logs/` → `/logs` — holds `app.log` (rotates at 5 MB, 3 backups)
|
|
Both directories are tracked in git (via `.gitkeep`) but their runtime
|
|
contents are excluded from version control via `.gitignore`.
|
|
`portainer-stack.yml` updated with absolute host paths.
|
|
- **Server writes a rotating log file** — `RotatingFileHandler` added to
|
|
`server.py`; writes `INFO`-level and above to `./logs/app.log`.
|
|
|
|
### Fixed
|
|
|
|
- **Engine settings lost after container recreate** — container names and
|
|
dynamic Docker card URL overrides were stored only in `localStorage`.
|
|
Added `engine_container_names` as a persisted server setting; all
|
|
container name inputs tagged with `data-cn-key` so `loadSettings()`
|
|
can restore them; dynamic card URL overrides now saved under
|
|
`engine_local_urls` with a `dc-` prefix. On first page load the server
|
|
wins over `localStorage`; changes write to both immediately.
|
|
|
|
- **`ImportError: cannot import name '_AUDIO_EXTS' from 'core.audio'`** —
|
|
`routes/stt.py` had a dead alias `from core.audio import _to_wav_16k, _AUDIO_EXTS as _VOICE_AUDIO_EXTS`;
|
|
`_AUDIO_EXTS` lives in `core.voice`, not `core.audio`. Removed the
|
|
wrong import; `_AUDIO_EXTS` is already correctly imported from `core.voice`
|
|
on the next line and the unused alias was never referenced in the file.
|
|
|
|
- **`No module named 'core'` on container restart** — `docker restart` reuses
|
|
the old container configuration and never applies new volume mounts from
|
|
`docker-compose.yml`. Added a comment to `portainer-stack.yml` and updated
|
|
it to include `core:/app/core:ro` and `routes:/app/routes:ro`. Solution is
|
|
`docker compose up -d` (recreates the container) not `docker restart`.
|
|
|
|
- **Text-input turns returned 422** — `UploadFile | None = File(None)` with
|
|
`from __future__ import annotations` caused FastAPI to treat `audio` as a
|
|
required field even when omitted. Changed to `Optional[UploadFile] = None`
|
|
(no `File()` wrapper) so the field is genuinely optional for text-only turns.
|
|
|
|
- **Conversation input bar hidden when mic unavailable** — the mic-blocked
|
|
warning box was inserted before `conv-chat-window` inside the flex column,
|
|
pushing the input bar off-screen. Fixed by prepending the warning inside
|
|
`conv-chat-window` so it scrolls with the chat and never affects the bar.
|
|
- **Chat panel overflow / missing input bar on smaller viewports** — removed
|
|
`min-height: 400px` from `conv-chat-window` (prevented shrinking) and gave
|
|
`conv-chat-panel` a viewport-relative height (`calc(100vh - 340px)`,
|
|
min 440px) so the input bar is always anchored at the bottom.
|
|
- **Browser caches old section HTML after updates** — `loader.js` now appends
|
|
`?v=<timestamp>` to every section fetch, busting the cache on each page load.
|
|
|
|
### Added
|
|
|
|
- **`scripts/release.py`** — automates version bump + CHANGELOG promotion.
|
|
`python scripts/release.py --patch|--minor|--major [--dry-run]` renames
|
|
`[Unreleased]` to the new version, updates compare links, writes `VERSION`,
|
|
commits, and creates an annotated git tag in one command.
|
|
- **Git pre-commit hook** (`scripts/hooks/pre-commit`) — warns (does not block)
|
|
when `.py`/`.js`/`.css`/`.html` files are staged but `CHANGELOG.md` or
|
|
`VERSION` are not. Run `bash scripts/install-hooks.sh` after cloning.
|
|
- **`scripts/install-hooks.sh`** — one-liner to install the hook after a fresh
|
|
clone: `bash scripts/install-hooks.sh`.
|
|
|
|
---
|
|
|
|
## [1.1.0] — 2026-05-29
|
|
|
|
### Security
|
|
|
|
- **Fixed path traversal in `/api/browse-dirs`** — Added a `_BROWSE_BLOCKED` blocklist
|
|
(`/proc`, `/sys`, `/dev`, `/run`, `/boot`). Requests for paths under these directories
|
|
now return HTTP 403 instead of listing kernel/system files.
|
|
- **Hardened yt-dlp output path** — After a YouTube download completes, the resolved
|
|
output path is verified to be inside `TEMP_DIR` via `.relative_to()`. A file written
|
|
outside the temp directory is rejected with an SSE error event and never registered.
|
|
- **Removed CORS wildcard on `/api/proxy-audio`** — `Access-Control-Allow-Origin: *`
|
|
was unnecessary (all callers are same-origin) and exposed proxied audio to arbitrary
|
|
cross-origin requests. Header removed.
|
|
- **Temp file registry now enforces a TTL** — `_registry` changed to
|
|
`dict[str, tuple[Path, float]]`. `_registry_gc()` evicts entries older than
|
|
`TEMP_FILE_TTL_SECONDS` (default 2 h, configurable via env var) and unlinks their
|
|
files, preventing unbounded disk growth on long-running instances.
|
|
|
|
### Performance
|
|
|
|
- **Settings and routing rules cached in memory** — `_load_settings()` and
|
|
`_load_tts_routes()` previously read from disk on every API request (55+ calls per
|
|
TTS synthesis). Both now use mtime-checked in-memory caches that invalidate
|
|
automatically on write, eliminating redundant file I/O.
|
|
|
|
### Added
|
|
|
|
- **Version number** — `VERSION` file at repo root; read by `core/constants.__version__`
|
|
and surfaced via `GET /api/version`. Displayed as `v1.1.0` in Settings → About.
|
|
- **Text input in Conversation Playground** — a pill-shaped text field and send button
|
|
(→) sit left of the mic button. Pressing Enter or → sends text directly through the
|
|
LLM → TTS pipeline, skipping STT entirely. Makes the playground fully usable without
|
|
a microphone (HTTP context, no mic permission, remote access). The backend
|
|
`/api/conversation/turn` now accepts an optional `text` form field; when set, the
|
|
STT step is skipped and the STT latency row shows `—`.
|
|
- **Container name field on all engine cards** — every TTS and STT engine card (Docker
|
|
stack cards *and* static "Other Local" cards) now always shows the Docker container
|
|
name input row. Previously absent/not-installed cards hid it; now it is always visible
|
|
so the container can be pre-configured before starting.
|
|
- **Connect / Disconnect toggle** — the Connect button now shows "Disconnect" (green,
|
|
`check-network` icon) when already connected and toggles back on click. State
|
|
persists in `localStorage`.
|
|
- **Auto-apply on Connect** — a successful connection probe automatically saves the
|
|
URL to Settings and makes the backend available in TTS/STT dropdown menus immediately,
|
|
without requiring a separate "Use as TTS/STT" click.
|
|
|
|
### Changed
|
|
|
|
- **Connect button redesigned** — moved out of the URL input row into a dedicated
|
|
`dc-controls-row`. Restyled as a solid blue primary CTA (was a small teal outline
|
|
button). Shows a spinner icon while probing.
|
|
- **"Use as TTS / STT" button** — larger padding, bolder teal border, chevron icon,
|
|
tooltip explaining it sets the URL in Settings. Gains `.active` highlight once applied.
|
|
- **Unified controls row on every engine card** — consistent left-to-right order:
|
|
`[Connect/Disconnect]` `[Stop | Start | Restart]` `[Use as →]`. Docker action buttons
|
|
hidden until a container name is entered; Use-as button right-aligned.
|
|
- **`initStaticDockerManagement`** — rebuilt to use the same `dc-controls-row`
|
|
structure as the dynamic Docker stack cards. The existing `.llm-local-ping` button
|
|
is moved from inside the URL row into the controls row at initialisation time.
|
|
- **Backend refactor — `server.py` (5 560 lines → 43 lines)** — all logic extracted
|
|
into single-responsibility modules:
|
|
|
|
| Package | Module | Responsibility |
|
|
|---|---|---|
|
|
| `core/` | `constants.py` | Boot-time env defaults, path constants, version, log buffer |
|
|
| | `registry.py` | TTL-based temp file registry |
|
|
| | `validation.py` | URL validation, SSRF guard, path safety |
|
|
| | `docker_client.py` | Raw Unix-socket Docker HTTP client |
|
|
| | `config.py` | Settings load/save/normalize, backend URL resolution |
|
|
| | `routing.py` | TTS route rules load/save/resolve, language detection |
|
|
| | `audio.py` | Audio conversion, normalisation, auto-trim scoring |
|
|
| | `voice.py` | Voice metadata, backup management, benchmark helpers |
|
|
| | `presets.py` | Voice Design preset load/save, virtual voice resolution |
|
|
| | `tts_helpers.py` | TTS request helpers, streaming, per-backend logic |
|
|
| `routes/` | `admin.py` | Index, favicon, browse-dirs, robots, version |
|
|
| | `settings.py` | `/api/settings`, routing rules, logs, design presets |
|
|
| | `library.py` | All voice CRUD, upload, save, normalize, export/import |
|
|
| | `stt.py` | `/api/transcribe*`, `/api/stt-backends` |
|
|
| | `sources.py` | Voice scraping, proxy-audio, yt-dlp download |
|
|
| | `docker.py` | `/api/local-containers/*`, `/api/probe-url` |
|
|
| | `tts.py` | TTS preview, streaming, voice design, `/v1/*`, backends |
|
|
| | `conversation.py` | Refine-text, effects, export/import, speak, MCP, conversation |
|
|
|
|
`Dockerfile` updated with `COPY core/ core/` and `COPY routes/ routes/`.
|
|
`docker-compose.yml` updated with `./core:/app/core:ro` and `./routes:/app/routes:ro`.
|
|
|
|
- **Frontend refactor — `app.js` (8 744 lines → 16 modules)** — split into
|
|
`static/js/` with `loader.js` loading them sequentially in dependency order:
|
|
|
|
| Module | Lines | Responsibility |
|
|
|---|---|---|
|
|
| `utils.js` | 364 | Core helpers: `$`, `toast`, `escHtml`, theme, language/flag, picker, tabs |
|
|
| `voice-inspector.js` | 397 | 3-pane voice workbench |
|
|
| `voice-sources.js` | 277 | External voice source scraping UI |
|
|
| `integrations.js` | 211 | Code snippet generation (SillyTavern, Open WebUI, HA, curl, MCP) |
|
|
| `routing.js` | 542 | TTS routing rules editor |
|
|
| `settings.js` | 385 | `loadSettings`, `applyAndSaveSettings`, settings panel |
|
|
| `voice-clone.js` | 774 | WaveSurfer, drop zone, mic recording, trim, voice design |
|
|
| `voice-library.js` | 2654 | Full voice library: list, CRUD, benchmark, normalize |
|
|
| `tts-preview.js` | 528 | TTS preview, `fetchTtsPreviewBlob` |
|
|
| `benchmark.js` | 218 | Performance + batch benchmark |
|
|
| `stt.js` | 287 | STT→TTS playground, `refreshSttBackends` |
|
|
| `init.js` | 49 | App bootstrap |
|
|
| `engines.js` | 625 | ElevenLabs browser, custom engine cards, Docker management |
|
|
| `ai-backends.js` | 520 | AI backend cards, LLM snippets, `initStaticDockerManagement` |
|
|
| `generation.js` | 393 | WAV merge, chunked TTS, history, playlist, audio effects |
|
|
| `conversation.js` | 520 | Conversation playground, LLM refinement, import, About |
|
|
|
|
### Fixed
|
|
|
|
- **`chrome://flags/…` URL unreadable in mic-blocked warning** — the global
|
|
`code { background: var(--panel) }` rule caused the URL text to render as
|
|
white-on-light-grey inside the red warning box. Fixed with inline styles
|
|
(`background: rgba(0,0,0,.35); color: #fff`) on the `<code>` element, plus a
|
|
Copy button so users don't need to manually select invisible text.
|
|
|
|
---
|
|
|
|
## [1.0.0] — 2026-05-28
|
|
|
|
Initial feature-complete release.
|
|
|
|
### Added
|
|
|
|
- **Voice library** — clone voices from audio samples; design voices from text
|
|
descriptions using instruction-based synthesis; benchmark synthesis speed (RTF);
|
|
normalize loudness; export/import voice packages as ZIP bundles.
|
|
- **TTS backends** — Qwen3 TTS (Voice Clone, Voice Design, Custom Voice, Streaming),
|
|
NVIDIA Magpie / Zeroshot / Flow, Kokoro FastAPI, VibeVoice, XTTS v2, ElevenLabs.
|
|
- **STT backends** — OpenAI Whisper (port 8010), faster-whisper-server, whisper.cpp,
|
|
Groq Whisper (cloud, free tier), NVIDIA Parakeet ASR. Real transcription probe
|
|
in health check (not just TCP reachability).
|
|
- **App Routing** — per-app / per-voice / per-language TTS routing rules with
|
|
automatic language detection and optional before/after sound effects.
|
|
- **Conversation Playground** — full STT → LLM → TTS pipeline with real-time SSE
|
|
streaming, latency stats panel (STT / LLM TTFT / LLM total / TTS / Total), turn
|
|
history, system prompt, and insecure-context warning.
|
|
- **Engines section** — LLM / STT / TTS sub-pages; Docker container management
|
|
(Start / Stop / Restart via Docker socket); custom engine cards; ElevenLabs voice
|
|
library browser.
|
|
- **Performance Benchmark** — single-voice and batch benchmark with RTF tracking,
|
|
sparkline trend, and persistent history.
|
|
- **Audio effects** — reverb, chorus, delay, compressor, gain, pitch shift
|
|
(via `pedalboard`).
|
|
- **Chunked TTS + generation history** — long-text synthesis split into chunks,
|
|
per-chunk playback, playlist export as WAV.
|
|
- **MCP server** — built-in JSON-RPC 2.0 endpoint at `/mcp`; tools: `speak`,
|
|
`transcribe`, `list_captures`, `list_profiles`.
|
|
- **LLM refinement & persona rewriting** — clean up STT transcripts or rewrite
|
|
responses with a chosen persona via any OpenAI-compatible LLM endpoint.
|
|
- **Connect Apps** — ready-made config snippets for SillyTavern, Open WebUI,
|
|
Home Assistant, curl, and MCP (`claude mcp add` one-liner).
|
|
- **Voice sources** — scrape voice assets from Aiartes, Freesound, GitHub, and
|
|
Google Drive; YouTube download via yt-dlp; quick import directly to library.
|
|
- **OpenAI-compatible proxy** — `/v1/audio/speech` and `/v1/audio/transcriptions`
|
|
for drop-in use with Open WebUI, SillyTavern, and Home Assistant.
|
|
- **Settings** — sub-pages: General, Connections, Playback, Captures, Payloads,
|
|
Storage, API Keys, Logs, About.
|
|
- **Voice Design presets** — saved persona templates for instruction-based synthesis;
|
|
virtual `vd_…` voices usable from external apps without exporting WAV files.
|
|
- **Multilingual support** — language/flag pickers, per-language preview texts,
|
|
`LANG_FLAG_DEFAULT` mapping for 16 languages.
|
|
- **Tags, ratings, and metadata** — per-voice tags with autocomplete, star ratings,
|
|
gender label, country flag.
|
|
- **Dark/light theme** — toggle with persistence in `localStorage`.
|
|
- **Docker socket integration** — Start/Stop/Restart Docker containers from the UI
|
|
via raw Unix socket HTTP; container health visible in engine cards.
|
|
|
|
---
|
|
|
|
[Unreleased]: https://github.com/mARTin-B78/tts-voice-creator-clone-and-design-2/compare/v1.1.0...HEAD
|
|
[1.1.0]: https://github.com/mARTin-B78/tts-voice-creator-clone-and-design-2/compare/v1.0.0...v1.1.0
|
|
[1.0.0]: https://github.com/mARTin-B78/tts-voice-creator-clone-and-design-2/releases/tag/v1.0.0
|