The old combined `21.8s · 1.30x` Speed cell is replaced by two separate sortable columns: - Factor (x.xx): audio÷render multiplier, green/amber/red colour-coded - Time (Xs): total render time Column order: Length · Duration · Factor · Time · WPM · Seed · dBFS · ... Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
663 lines
60 KiB
Markdown
663 lines
60 KiB
Markdown
# Changelog
|
||
|
||
All notable changes to **TTS Voice Creator — Clone and Design** are documented here.
|
||
Follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/) · versioned with [Semantic Versioning](https://semver.org/).
|
||
|
||
---
|
||
|
||
## [Unreleased]
|
||
|
||
---
|
||
|
||
## [1.12.13] — 2026-06-27
|
||
|
||
### Changed
|
||
- **Speed column split into Factor + Time**: the old `21.8s · 1.30x` cell is now two separate sortable columns — **Factor** (`1.30x`, audio÷render, colour-coded green/amber/red) and **Time** (`21.8s`, total render time). Column order: Length · Duration · Factor · Time · WPM · Seed · dBFS · Type · Source · Rating · Tags · Note · Active.
|
||
|
||
---
|
||
|
||
## [1.12.12] — 2026-06-27
|
||
|
||
### Added
|
||
- **Duration column** in the voice table — shows the length of the synthesised benchmark audio (e.g. `11.4s`), sortable. Distinct from Length (original clip) and WPM (speaking rate).
|
||
- **Benchmark sentence presets** dropdown — four ready-made sentences (DE narrative, EN narrative, DE pangram, EN tongue-twister) plus a Reset option. Selecting a preset loads it instantly without overwriting anything else.
|
||
|
||
### Changed
|
||
- **Column order** rearranged to: Img · Play · Name · Lang · Gender · Length · Duration · WPM · Speed · Seed · dBFS · Type · Source · Rating · Tags · Note · Active.
|
||
|
||
---
|
||
|
||
## [1.12.11] — 2026-06-27
|
||
|
||
### Fixed
|
||
- **Speed and WPM cells now update live** during a benchmark run — each row's cells are patched in-place immediately after its voice finishes, without waiting for the full batch to complete and the list to re-render. Same fix applied to the single-voice "Benchmark this voice" button.
|
||
|
||
---
|
||
|
||
## [1.12.10] — 2026-06-27
|
||
|
||
### Added
|
||
- **WPM column** in the voice table — shows how fast a voice speaks in words per minute, calculated from the benchmark audio duration and sentence word count. Sortable. Hover for a tooltip explaining the range (130–180 wpm is natural for audiobooks).
|
||
- **Info (ⓘ) icons** on the Speed and WPM column headers with tooltips explaining what each metric measures.
|
||
|
||
---
|
||
|
||
## [1.12.9] — 2026-06-27
|
||
|
||
### Changed
|
||
- **Benchmark speed factor shows 2 decimal places** (`1.32×` instead of `1.3×`) in the SPEED column and live status line — more precise RTF comparison across voices.
|
||
- **Sorting by Speed now sorts by the RTF factor** (faster voices first) instead of by raw synthesis time — voices with a higher `×` multiplier rank higher regardless of how long the benchmark sentence was.
|
||
|
||
---
|
||
|
||
## [1.12.8] — 2026-06-27
|
||
|
||
### Changed
|
||
- **SOURCE column auto-derives script name for Rehearser clones**: voices cloned from a Script Rehearsal now automatically display the script title (e.g. `Script`, `her`) in the SOURCE column, extracted from the note field (`Rehearser · <title> · <character> — …`). No manual entry needed; setting an explicit `origin` still takes priority.
|
||
|
||
---
|
||
|
||
## [1.12.7] — 2026-06-27
|
||
|
||
### Added
|
||
- **Editable Source field** in the voice inspector panel (below Note) — type a source label (e.g. `fish-audio`, `cloned`) and it saves immediately to the voice's meta.json. Changes now persist across refreshes.
|
||
- **Bulk "Set source"** button in the multi-select toolbar — select any number of voices and apply a source label to all at once.
|
||
- **Fish-audio tag fallback** in the SOURCE column: voices tagged `fish-audio` automatically show `fish-audio` as their source even without an explicit origin field, so the column is populated correctly without having to edit every voice.
|
||
|
||
---
|
||
|
||
## [1.12.6] — 2026-06-27
|
||
|
||
### Added
|
||
- **Audiobook auto-saves to Script Rehearser**: when casting completes, the result is automatically written to the Script Rehearser IndexedDB (same format, same library). The record is updated — not duplicated — whenever speaker corrections are made in the cast view. The record appears immediately in Script Rehearser → Library.
|
||
- **"Edit in Rehearser" button**: replaces "Review & cast" in the completed cast panel. Opens the saved record directly in Script Rehearser with all speakers, voices, and page markers intact, ready to assign voices and synthesise.
|
||
- Voice assignments made in Script Rehearser are **preserved** when the audiobook auto-saves again (only the script text and emotions are overwritten; voice/instruct/soul fields survive the update).
|
||
|
||
---
|
||
|
||
## [1.12.5] — 2026-06-27
|
||
|
||
### Added
|
||
- **Autosave for audiobook casting**: progress is saved to localStorage after every passage. A page refresh, browser crash, or accidental close no longer loses hours of casting work — reopening "Cast as audiobook" for the same document restores the session automatically with a banner showing how far it was completed and when it was last saved. Manual speaker corrections made in the cast view are also autosaved immediately. The draft is cleared when the script is explicitly saved to Script Rehearsals.
|
||
|
||
---
|
||
|
||
## [1.12.4] — 2026-06-27
|
||
|
||
### Added
|
||
- **Expandable context dividers in Recast Unknown**: the `⋯` gaps between scattered Unknown segments now show a count ("42 lines hidden — click to expand") and expand inline on click, revealing all segments between two Unknown passages so the user can see who is speaking before and after to make a better assignment.
|
||
|
||
---
|
||
|
||
## [1.12.3] — 2026-06-27
|
||
|
||
### Added
|
||
- **Casting activity indicator**: pulsing blue dot on the Read Aloud nav item while audiobook casting is in progress, so users can see it's working from any section.
|
||
- **Restore cast view on return**: navigating away during casting and then back to Read Aloud automatically restores the casting panel instead of showing the blank PDF view.
|
||
|
||
### Fixed
|
||
- **Character count out of sync on reassignment**: reassigning a segment's speaker now decrements the old speaker's count in the Characters Found panel, so totals stay accurate when corrections are made. Speakers that drop to 0 lines are removed from the panel automatically.
|
||
|
||
---
|
||
|
||
## [1.12.2] — 2026-06-27
|
||
|
||
### Added
|
||
- **PDF Search**: search input in the Read Aloud toolbar — press Enter to jump to the first page containing the term, Shift+Enter to go back.
|
||
- **Noise gate slider**: adjustable minimum-amplitude threshold in the Conversation input bar (the blue marker on the level meter shows the current gate). Short noise spikes below the gate or bursts shorter than 300 ms are silenced before STT. Persists across sessions.
|
||
- **Conversation stats panel** now has a collapse button (→) to hide the latency sidebar and give more chat space; a floating icon button restores it.
|
||
|
||
### Changed
|
||
- **Sidebar tooltip** now extracts text correctly for all item types (sub-items that had no `.nav-label` span showed nothing before).
|
||
- **Active section** in the sidebar now gets a blue background highlight + right-border accent (`.nav-tree-item.active` was previously unstyled, making it impossible to tell which section you were in).
|
||
- **Casting audiobook**: LLM prompt now explicitly handles `?«` and `!«` as valid German quote endings, and instructs the model to treat an unclosed `»` at end-of-passage as dialogue. Deterministic fallback also handles the unclosed-quote edge case.
|
||
|
||
---
|
||
|
||
## [1.12.1] — 2026-06-27
|
||
|
||
### Changed
|
||
- **Sidebar icon-rail**: hover no longer flies out the whole sidebar (which caused the main content to jump left/right). Individual item labels now appear as a small floating tooltip next to the hovered icon, keeping the rail at a fixed 56 px and the layout completely stable.
|
||
|
||
---
|
||
|
||
## [1.12.0] — 2026-06-27
|
||
|
||
### Added
|
||
- **Slim icon-rail sidebar**: the button where the flag used to be collapses the sidebar to a narrow icon rail (desktop); hovering the rail flies the full menu out as an overlay (titles + nested items). State persists.
|
||
- **Language picker moved to Settings → General** (was the sidebar flag). English / Deutsch.
|
||
- **Collapsible settings panels** (using the app's standard `card` collapse style) to free vertical space: Conversation engine/prompt config, Read Aloud voice & synthesis settings (with the drag-&-drop inside), the Try It Out voice/playback box, and the Casting-audiobook LLM/prompt panel (prompt collapsed by default).
|
||
|
||
### Changed
|
||
- **Read Aloud layout** reorganised: collapsible settings + drag-&-drop on top, zoom toolbar above the document, the document as the central area that **fits the viewport height**, transport + synthesis controls below it (always visible — only the document scrolls).
|
||
- **Read Aloud "My Books"** card removed — saved books now live in the combined **Library → Books**.
|
||
- **Conversation Playground** fills the viewport height (only the chat scrolls), and the config stacks full-width.
|
||
- Collapsible-card initialisation now also runs when a section opens, so lazily-loaded sections get consistent collapse chevrons.
|
||
|
||
### Fixed
|
||
- **Conversation microphone failed intermittently** ("EBML header parsing failed") in hands-free/live mode: a silence reset cleared the recording buffer in place, dropping the webm header so later utterances were undecodable. The recorder now restarts cleanly, keeping every utterance valid. (Recordings are also decoded in-browser to 16 kHz WAV, bypassing server ffmpeg.) Barge-in (interrupt the agent while it speaks) works via the **Live agent** toggle.
|
||
- **Casting-audiobook overflow**: a long unbroken passage string widened the layout and pushed the sidebar off-screen; the feed now wraps and is width-constrained.
|
||
|
||
---
|
||
|
||
## [1.10.1] — 2026-06-27
|
||
|
||
### Added
|
||
- **Collapsible nested sidebar**: group headers *Voice Actions*, *Speak* and *Setup* are now expandable parent menus; **Tags** nests under *Library*, and *Integrations* (App Routing · Connect Apps) and *Settings* nest under *Setup*. Opening a section auto-expands its whole ancestor chain.
|
||
|
||
### Changed
|
||
- **Conversation Playground** config stacks full-width — Speech to Text, Language Model, Text to Speech and System Prompt each on their own row for clarity.
|
||
|
||
### Fixed
|
||
- **Conversation microphone failed with "Audio upload failed: Decoding failed / EBML header parsing failed"** on ARM64: recorded audio is now decoded in the browser and uploaded as 16 kHz mono WAV, bypassing the server-side ffmpeg webm parser entirely (falls back to the raw blob if browser decoding is unavailable).
|
||
|
||
---
|
||
|
||
## [1.10.0] — 2026-06-27
|
||
|
||
### Added
|
||
- **Combined Library** (new *Library* section under *Speak*) with three tabs — **Books**, **Theater Plays**, **Characters / Cast** — bringing audiobooks, rehearsals and the character roster into one place. Book and play cards cross-link: *Rehearse* a book, *Read Aloud* a play.
|
||
- **Character tags** (like voice tags): each character carries a comma-separated list of productions, seeded with its origin book and editable in the character editor. A character can now be **reused across several books/scripts** — one record, many tags.
|
||
- **Shared cast resolution**: opening a production in the Rehearser (and audiobook casting, which routes through it) auto-fills empty cast slots from the shared character roster, matched by book **or** tag. Cast voice choices are written back to existing characters on save, keeping the roster in sync. Productions are joined by normalized title (`prodKey`).
|
||
- **Tags** navigation group under *Voices*: distinct voice tags with live counts, plus the *Cloned / Designed / Favourites* predicates, each filtering the voice list.
|
||
|
||
### Changed
|
||
- **Sidebar restructured** into clearer groups: *Voices* (Library · Tags · Voice Actions), *Speak* (Quick Play · Conversation · Read Aloud · Script Rehearsal · Library), *Setup* (Benchmark · Engines · Integrations · Settings). "Try It Out" is now **Quick Play**; "My Voices" is now **Library**.
|
||
- The standalone **Characters** section was folded into the combined Library; **Read Aloud** is now a single entry (its library lives in the combined Library); the Rehearser's Library/Cast steps moved into the combined Library, leaving Stage · Summary · Import/Export in the sidebar.
|
||
|
||
---
|
||
|
||
## [1.9.8] — 2026-06-27
|
||
|
||
### Changed
|
||
- **Seed Finder default sentence**: updated to a longer mixed DE/EN test phrase covering numbers with dots (3.567), compound nouns, umlauts, special characters, English technical vocabulary, time formats, and motivational prose — gives a more complete picture of a voice's character per seed.
|
||
|
||
---
|
||
|
||
## [1.9.7] — 2026-06-27
|
||
|
||
### Fixed
|
||
- **Voice inspector broken**: `curGender` was referenced before its `const` declaration inside `selectVoice()`, causing a temporal dead zone ReferenceError that silently prevented the inspector from opening. Clicking a voice now correctly opens the detail panel again.
|
||
- **Batch operations ignore selection**: Calc dB, Precompute, and Batch Seeds now operate only on the checked (selected) voices when a selection is active, matching the existing behaviour of Benchmark. Previously all three always ran on all active / visible voices regardless of selection.
|
||
|
||
---
|
||
|
||
## [1.9.6] — 2026-06-26
|
||
|
||
### Fixed
|
||
- **Casting feed — scroll hijacking**: the feed no longer forces the view to the bottom while you are scrolled up reviewing or editing earlier lines. Auto-scroll only fires when you are already within 80px of the bottom.
|
||
- **Casting feed — vanishing text**: the trimming limit was raised from 80 to 600 rows, and trimming is now suppressed while you're scrolled up, so older lines stay visible as long as you're looking at them.
|
||
|
||
### Added
|
||
- **Casting feed — "↓ Live" jump button**: a floating pill button appears at the bottom of the feed whenever you've scrolled up. Click it to immediately return to the live bottom of the feed and re-enable auto-scroll.
|
||
|
||
---
|
||
|
||
## [1.9.5] — 2026-06-26
|
||
|
||
### Added
|
||
- **Casting view — page-break lines**: as the LLM processes a PDF audiobook, a **"Page N"** divider row now appears in the casting feed whenever the source PDF page changes. This gives a live view of where each page boundary falls within the script.
|
||
- **Casting view — expandable LLM passage panel**: the "LLM Reading…" indicator now shows a preview of the passage being processed. A **chevron button** (▾/▴) expands it into a full scrollable view of the passage text so you can follow exactly what the LLM is reading and thinking about.
|
||
|
||
### Fixed
|
||
- **Page numbers were off by one**: page-break markers emitted into the Rehearser script now use the correct 1-indexed PDF page number (was storing 0-indexed, so "Page 1" showed for what was actually the second PDF page).
|
||
|
||
---
|
||
|
||
## [1.9.4] — 2026-06-26
|
||
|
||
### Changed
|
||
- **Audiobook casting — page numbers in Rehearser**: page-break markers now carry the source PDF page number. In the Script Rehearser Stage, each page divider shows **"— Page N —"** instead of the generic "— Page break —", making it easy to cross-reference the audiobook script against the book. Applies both to audiobooks cast from Read Aloud PDFs and to screenplay PDFs imported directly into the Rehearser.
|
||
|
||
---
|
||
|
||
## [1.9.3] — 2026-06-26
|
||
|
||
### Added
|
||
- **Native Speed for Try it out & Read aloud**: a **Native Speed** control (range 0.5×–2.0×, default 1.0) is now available in both the *Try it out* and *Read aloud* sections. It passes the `speed` parameter directly to the TTS generation request (natively via the faster-qwen3-tts backend), producing audio at the target tempo from the model rather than using post-processing pitch/time-shift. Try it out persists the chosen speed in localStorage (per browser); Read aloud saves it with the document in the library (each book remembers its own speed).
|
||
|
||
---
|
||
|
||
## [1.9.2] — 2026-06-26
|
||
|
||
### Fixed
|
||
- **Audiobook casting — pagination preserved on save**: saving (or opening) a cast audiobook as a Script Rehearsal now keeps the book's **page breaks**. Read-aloud page boundaries are tracked through casting and re-emitted as `\f` markers at the nearest segment boundary, so the Rehearser paginates the saved script to match the source PDF instead of producing one continuous flow.
|
||
|
||
---
|
||
|
||
## [1.9.1] — 2026-06-26
|
||
|
||
### Fixed
|
||
- **Character sheets — "Connection refused" failures**: the analysis could send a dead `localhost:11434` LLM URL to the server (when app settings hadn't loaded yet), causing every passage to fail. It now never falls back to that hard-coded default — when no endpoint is explicitly chosen it lets the **server use its own configured `llm_url`**, so character-sheet extraction uses the same working LLM as the rest of the app.
|
||
- **Audiobook casting — empty "Review & cast" preview**: opening the manual-correction preview after a cast rendered no lines because the name-highlighter (`highlightText`) was scoped to the live cast view only. It is now shared, so you can review, fix speakers/emotions, and open in the Rehearser again.
|
||
|
||
### Changed
|
||
- **Script Rehearser / Stage — one text size**: narrator (action) text and spoken dialogue now use the **same font size** instead of mismatched 15px/16px, so the play reads evenly.
|
||
- **Script Rehearser / Stage — text-size control**: a new **A− / A+** control in the Stage toolbar shrinks or enlarges the whole play (persists per browser), in both A4 and paginated/scroll views.
|
||
- **Script Rehearser / Stage — collapsible character list**: the row of cast chips can now be **collapsed or expanded** via a "Characters (N)" toggle to free up vertical space.
|
||
|
||
---
|
||
|
||
## [1.9.0] — 2026-06-26
|
||
|
||
### Added
|
||
- **Character Library**: a new **Characters** section in the sidebar collects every character the LLM extracts into a persistent, browsable library (IndexedDB), grouped by the book or script they came from. Running **Character sheets** from Read Aloud or the Script Rehearser now auto-saves each character (keyed by *book + name*), merging in new detail on re-runs. Each card can be **edited** inline or **deleted**, and keeps its greyscale Good↔Evil alignment bar, arc arrow, page/line sources, and the 5-area psychological **Deep Analysis**. New module `static/js/characters-library.js`.
|
||
- **Richer character sheets**: extraction now gathers six more book-derived fields per character — *Backstory & Origin, Relationships, Motivation, Fears, Mannerisms & Habits,* and a casting-focused *Voice & Speech* pattern (accent, pacing, register, verbal tics).
|
||
- **Character sheets — morality at a glance**: each sheet now shows a **greyscale Good↔Evil alignment bar** (white = good, black = evil) with a 0–100 score, plus an **arc arrow** indicating whether the character stays put, *descends* (↘ good→bad), *redeems* (↗ bad→good), or follows a *complex* (↕) path. Also added **Clothing & Appearance** and **Capabilities** fields and richer physical detail (height, hair, eyes, skin, gait), with sources now carrying a short category **line hint**.
|
||
- **Character sheets — Deep Analysis**: a per-character button runs a 5-area psychological & narrative study (Core Flaw & Desire · Agency & Passivity · Dialogue & Voice · Narrative Arc · Paradox & Depth) in a modal. New endpoint `POST /api/character-deep-analysis`.
|
||
- **Audiobook casting — Recast & rescue tools**: after a cast run you can now **Recast all** (re-run the whole document with tweaked settings/prompt), **Recast unknown** (re-attribute only the leftover *Unknown* lines using surrounding context), run a **2nd Quality Run — Verify** pass that re-checks every speaker assignment and resolves Unknowns, and **Save script** straight to Script Rehearsals without leaving the page.
|
||
|
||
### Changed
|
||
- **Recast unknown — passage separators**: re-analysing only the *Unknown* speakers now shows a `⋯` divider between non-adjacent passages, so it's clear where one excerpt ends and another begins.
|
||
- **Dialogue attribution — fidelity & language**: the attribution prompt now forbids hallucinating/summarising (segments must reconstruct the passage word-for-word), emits emotion tags in the **same language** as the text (e.g. German *wütend/flüsternd*), and demands strict JSON. Removed the brittle text-script fallback parser that could mis-split prompt echoes into fake speakers.
|
||
|
||
---
|
||
|
||
## [1.8.1] — 2026-06-25
|
||
|
||
### Added
|
||
- **Audiobook casting — Warmup request**: Added an invisible "Wake up" request before starting the attribution loop, showing a clear "Waking up LLM model..." status. This absorbs the 3–4 minute cold-boot time of massive models (like 120B/30B via llama-swap) without freezing the UI or timing out the first real book chunk.
|
||
- **Audiobook casting — Quick-cancel**: The character assignment popup now has a close button and responds to the `Escape` key, so you can easily dismiss it if you click "Assign" by accident.
|
||
|
||
### Changed
|
||
- **Audiobook casting — Narrator quick-pick**: The `📖 Narrator` role is now always pinned to the top of the manual assignment dropdown list, so you don't have to scroll or type to revert a mis-cast line to narration.
|
||
- **LLM timeout increased**: The hardcoded API timeout in `routes/conversation.py` for all LLM calls has been increased from 3 minutes to 10 minutes (`timeout=600`), providing plenty of headroom for dynamic model proxies to download and load models into VRAM on demand.
|
||
|
||
### Fixed
|
||
- **Read Aloud — "Fit width" scaling bug**: Clicking "Fit width" on PDFs with a small cover page (e.g. A5) but larger subsequent pages previously zoomed the cover perfectly but blew the text pages up massively (e.g. 241%), forcing horizontal scrolling. The scale is now computed against the *maximum* width of all pages, guaranteeing the entire book fits.
|
||
- **Read Aloud — Left-side text cut-off (CSS bug)**: Fixed a flexbox centering issue on the document container where a zoomed PDF page that was wider than the screen would overflow equally on both sides. Because browsers only scroll to the right, the left edge of the page was permanently inaccessible. Fixed by replacing `align-items: center` with `margin: 0 auto`.
|
||
- **Audiobook casting — Minified bundle caching**: Features added directly to JS files weren't appearing because the app was serving an older, cached minified bundle. The bundle has been rebuilt and version cache-busting ensures the new UI shows up immediately.
|
||
|
||
---
|
||
|
||
## [1.8.0] — 2026-06-22
|
||
|
||
### Added
|
||
- **Seed Finder — pin a specific seed**: a direct *“fix this voice to a specific seed”* control (type a seed → **Pin**, or **Unpin** to go back to random) that saves straight to the TTS server without generating anything — handy when you already know the seed you want.
|
||
- **Seed Finder — persistent pinned seed indicator**: the pinned seed is now saved locally in the voice's metadata and displayed prominently in the voice inspector header. It automatically restores the "Pin seed" input when reopening the voice.
|
||
- **Seed Finder — batch all voices**: a **Batch seeds** button in the voice toolbar pre-generates and caches Seed Finder samples for every active voice. An **in-app dialog** (no browser pop-ups) lets you set the seed range with a live sample-count estimate, then shows live progress (current voice/seed) with cancel. Runs sequentially, **skips already-cached seeds** (resumable), and feeds the same cache the per-voice Seed Finder reads.
|
||
|
||
### Changed
|
||
- **Seed Finder — samples are cached**: generated seed WAVs are saved in the browser (IndexedDB) keyed by voice + test sentence + backend, so reopening a voice shows previous results instantly and re-running only generates the **missing/failed** seeds instead of all of them. Added a **Clear saved** button to drop a voice's cache.
|
||
- **Seed Finder — better test sentence**: a shorter default sentence that exercises all German umlauts (ä ö ü ß), numbers, and a few English words — quicker to generate and more revealing of a voice's character.
|
||
|
||
### Fixed
|
||
- **Couldn't change a voice's displayed name**: the big name in the voice inspector was just the **last segment of the voice ID** (so `DE_F_Privat_Laura_01` showed as `01`), and double-clicking it renamed the *ID*, not the shown name. Double-clicking the header now edits a real **display name** (saved to the voice's metadata via `/api/voice/meta`, persists across reloads) and pre-fills with the current name; the separate **edit ID** button still renames the underlying voice ID. (Designed voices without a reference file don't support a stored display name yet.)
|
||
- **About page showed `v0.0.0` and a stale changelog**: the deployment stack wasn't mounting `VERSION`/`CHANGELOG.md`, so the container had no live version file (fell back to `0.0.0`) and served the image's baked changelog. Both files are now bind-mounted in `docker-compose.yml` and `portainer-stack.yml`, so the About page reflects the running version and changelog.
|
||
- **Seed Finder — “Failed to fetch” seeds**: long per-seed generations that intermittently dropped now **auto-retry** (2 attempts), each failed seed gets its own **Retry** button, and because successful seeds are cached, a second Run fills only the gaps rather than redoing everything.
|
||
|
||
## [1.7.0] — 2026-06-21
|
||
|
||
### Added
|
||
- **Seed Finder** — in **My Voices**, each voice's inspector has a 🎲 **Seed Finder** panel: generate a sample for a range of seeds (the TTS engine produces a slightly different take per seed), play each result, and click **★ Use seed N** to pin your favourite — saved immediately to the TTS server's `voices.json`. New `/api/tts-voice-seed` proxy, module `static/js/seed-finder.js`.
|
||
- **Voice library — Precompute embeddings**: a **Precompute** button warms all active voices so the TTS engine builds and caches each voice's speaker embedding (`.pt`) ahead of time, making the *first* playback of a voice instant (better Time-To-First-Audio) instead of paying the one-time analysis cost on first use. Runs with bounded concurrency, progress, and cancel. (The faster-qwen3-tts engine already prefers a cached `.pt` and auto-creates it from the reference wav when missing — the wav stays the source of truth; this just pre-warms the cache.)
|
||
- **Character sheets** — a **Character sheets** button in both **Read Aloud** and the **Script Rehearser** uses your configured LLM to extract actor-facing, RPG-style profiles for every character: Archetype, Physical Stats (metric-only), Alignment & Ethos, Core Attributes (highest/lowest), Trained Skills, Signature Inventory, Dark Secret / Fatal Flaw, Conflict Style, and Win Condition. Deduced details are marked with `*`, characters are grouped **main vs supporting**, and each sheet cites its **sources** (page number from the PDF + a short verbatim quote). Long texts are processed in chunks and merged per character (fields filled in, inventory/sources de-duplicated). Results render as scannable cards in an overlay with **Copy as Markdown**, and are cached so re-opening is instant. New endpoint `POST /api/character-sheets`, module `static/js/character-sheets.js`.
|
||
- **Book → multi-speaker audiobook** — a **Cast as audiobook** button in Read Aloud turns a novel into a cast-able script. An LLM scans the current scope (selection / page range / whole book, chunked with a running character roster so the same speaker keeps one name throughout) and attributes every segment to a **Narrator** or a **character**, with a per-line **emotion**. New endpoint `POST /api/attribute-dialogue` and module `static/js/audiobook.js`; reuses the rehearser's casting, per-line tone, and synthesis. Failed chunks fall back to narration so a book always casts; progress is shown and cancellable.
|
||
- **Live casting view**: while attributing, a **wide, non-blocking, minimisable panel** shows a scrolling feed of each line with its assigned speaker (and emotion) plus a **character roster that fills up** with per-character line counts — instead of a bare modal progress bar. You can **minimise it and keep using the app** (any tab), then re-open to watch progress; when it finishes it parks as a **"✓ Review & cast"** panel rather than auto-popping, so it waits for you if you wandered off.
|
||
- **Calmer messaging**: a passage the LLM can't attribute is no longer shown as a red "attribution failed: Error" toast — it's a quiet "read by the narrator" note in the feed, with a single neutral summary ("N passages had no detected dialogue") in the review step.
|
||
- **Editable preview**: before handing off, a review overlay lists every segment with an editable **speaker** (autocompletes from detected characters) and **emotion** so mis-attributions are fixed in seconds; "Open in Rehearser" applies the edits and lands you at the Cast phase.
|
||
- **Audiobook export** (rehearser): an **Audiobook** button synthesises every cast line as MP3 (bounded concurrency, progress, cancel) and downloads **one MP3 per chapter** (split on Chapter/Kapitel/Part/Prologue… headings or act/scene markers), or a single file when no chapters are detected.
|
||
- **Saved as a reopenable rehearsal**: handing the cast off to the Rehearser now also **saves it to the Script Rehearser library (Bibliothek)** automatically, so the attributed script + cast + per-line emotions persist — reopen it anytime to change speakers/voices/lines and synthesise or export the audiobook.
|
||
- **Read Aloud** — a new sidebar tab that turns the app into a text-to-speech document reader. Import a **PDF** (rendered to its real page layout via pdf.js) or a **.txt / .md** file, pick any voice + backend and a **reading speed** (0.5×–2×), then press play: the document is read sentence-by-sentence while the **word being spoken is highlighted** in place (overlay box on the PDF page, inline highlight in text mode), with the view auto-scrolling to follow. Click any word to jump there. Reuses the rehearser's word-timing + pdf.js loader and the existing `/api/tts-preview` pipeline — no backend changes. New files `static/sections/s-reader.html` and `static/js/reader.js`.
|
||
- **Read Aloud — synthesis-state overlay**: every sentence is colour-coded by state — **red** (not synthesised), **yellow** (synthesising), **green** (ready/cached), **blue** (currently reading) — shown as a translucent overlay on the PDF page and as a tint in text mode, with a legend.
|
||
- **Read Aloud — PDF zoom controls**: Fit width, Fit height, Two-page spread, and zoom in/out with a live percentage. Word geometry is stored scale-independently so zoom re-renders instantly and highlights stay aligned; fit modes track window resizes.
|
||
- **Read Aloud — resume**: the last reading position is remembered per document, so re-opening the same file resumes where you left off.
|
||
- **Read Aloud — book library**: **Save to library** stores the original PDF/text together with its **synthesised audio** and reading position in the browser (IndexedDB). A "My books" shelf lists saved documents with audio- and read-progress bars; reopen one to continue right where you left off with the already-synthesised pages intact — handy for working through long books. Reading position auto-saves on pause / stop / leaving the tab.
|
||
- **Read Aloud — MP3 storage & export**: audio is now synthesised and stored as **MP3** (far smaller than WAV, so books fit comfortably in the browser library). An **Export MP3** control downloads the synthesised audio as **one file per page** (sections combined) or **one file per sentence**, with meaningful filenames like `Title - p01 - 03.mp3` / `Title - p01.mp3`. Any not-yet-synthesised sentences in scope are rendered first.
|
||
- **Read Aloud — voice consistency**: addresses the slight timbre/prosody drift you hear when each sentence is generated separately. A **"Voice consistency"** selector synthesises in larger continuous chunks — *per sentence* (responsive), *per paragraph* (steadier), or *per page* (steadiest) — so a whole passage is one generation. Optional **Seed** and **Temperature** inputs pin the generation (forwarded to backends that support them, with graceful fallback), **Normalise loudness** evens out volume between chunks on playback, and a backend hint flags cloned/zero-shot engines that re-sample per request and suggests remedies. Chunk mode + seed/temperature/normalise are saved with library books.
|
||
- **Read Aloud — synthesise ahead**: a **Synthesise** button pre-renders audio for gap-free reading, scoped to **all**, a **page range** (PDF), or a **click-selected sentence range**. Select mode is a guided, persistent step flow — click a start sentence (it pulses as the anchor), then the end; the mode stays active with step hints so you can keep refining, and you leave it with the **Done** button or **Esc**. Synthesis runs with bounded concurrency, shows progress, and can be cancelled; the synthesis-state colours fill in green as each sentence completes.
|
||
|
||
### Changed
|
||
- **Audiobook casting — smarter speaker attribution**: the LLM prompt now reasons about **conversational turn-taking** (in a two-person exchange speakers alternate, so untagged lines are attributed by context rather than dumped as “Unknown”), and each passage is given the **recent dialogue** from the previous one so a conversation continues correctly across passage boundaries. The deterministic fallback (used when the LLM is unavailable) also got a conservative two-person turn-taking fill and a stop-list that rejects common German non-name words (Sofort, Stimme, Frage, Plötzlich…), so it no longer invents bogus characters.
|
||
- **Character sheets — self-filling across the book**: sheets now build up progressively — each passage receives the **sheet-so-far** (with which fields each character still needs) and the model fills gaps and refines instead of starting from scratch, so details accumulate as more of the book is read.
|
||
- **Read Aloud — library now lives on the server (syncs across devices)**: saved books, their synthesised audio, and reading position were previously stored only in the browser (IndexedDB), so a book saved on the laptop never appeared on the desktop. The library now persists under the server's config volume (`reader_library/<id>/` with `meta.json`, the source document, and per-unit MP3s) via new `/api/reader/docs…` endpoints. Any device pointed at the same server sees the same "My books" shelf; opening a book is instant and its audio streams **per chunk on demand** (nothing is bulk-downloaded), and saves stay incremental (only new chunks upload).
|
||
- **Language switcher** — replaced the sidebar language dropdown with a **flag toggle** next to the "Voice Creator" headline (click to switch interface language). Added German strings for **Read Aloud** ("Vorlesen") and its UI.
|
||
|
||
### Performance
|
||
- **Read Aloud — memory & smoothness for long books**: decoded audio (uncompressed PCM) is now kept only for a small window around the playhead and re-decoded from the cached MP3 on demand; off-screen PDF page canvases are released and re-rastered on return — together these bound memory on big books (previously both grew unbounded and could crash long sessions). The next chunk is **pre-decoded** during playback for gapless transitions, transport actions no longer scan every unit (single tracked "reading" index), PDF sentences/units are built **incrementally per page** (no end-of-parse spike), and library **saves are incremental** — only newly-synthesised chunks are written (a separate per-unit audio store), instead of rewriting the whole book each save.
|
||
|
||
### Fixed
|
||
- **Audiobook casting — German (and other) quote styles not recognised**: dialogue marked with German guillemets `»…«` / `„…“` / `›…‹`, French `«…»`, curly `“…”`, CJK `「…」`, or em-dash speech was treated as narration, so books like German novels cast everything to the narrator. The LLM prompt now explicitly handles all these styles (with guillemets called out), passages with **no quotation marks skip the LLM entirely** (so genuine narration isn't shown as a failure), PDF line-break **hyphenation is mended** (`Schwer- tes` → `Schwertes`) for clean speech, and if the LLM call fails on a passage that *does* contain quotes, a **deterministic fallback** splits out the dialogue **and attributes speakers from speech tags** (`»…«, sagte Riskan` → Riskan; pronouns rejected) so the book stays castable with real names even when the LLM is offline.
|
||
- **Read Aloud — auto-scroll**: while reading, the view now scrolls only the document pane instead of the whole window, so the currently-spoken line no longer slides up under the app header out of view.
|
||
- **Read Aloud — backend dropdown stuck on "Checking…"**: the reader's TTS-backend select is now populated by the shared backend refresh and fetched on demand when the section opens, so it fills reliably even if backends finish loading after you're already on the tab.
|
||
- **Read Aloud — large PDFs froze the page** ("this page is not responding"): the page-parse loop now yields to the browser periodically (with a "Reading PDF… page x / n" indicator), and per-sentence status overlays are created lazily per page instead of all at once. A 60-page book now imports with a max main-thread stall of ~40 ms (was multi-second), creating only the visible pages' overlays.
|
||
- **Chunked TTS — "Failed to fetch" on long text**: `splitTextIntoChunks` only split on sentence terminators (`.!?`), so newline-delimited text (e.g. German bullet lists or care-plan notes) was never split — the full page was sent as one request, causing a TCP timeout that the browser surfaced as "Failed to fetch". Fixed by processing each line individually before applying the sentence regex. Also moved `generation.js` from deferred batch E into the main feature batch so `generateChunkedTts` is always defined before the user can click Generate.
|
||
|
||
---
|
||
|
||
## [1.6.0] — 2026-06-03
|
||
|
||
### Added
|
||
- **Script Rehearser — Cast overhaul**: Card / List **view toggle**; **sort & filter** (name, gender, language, line count, tag); character-card-game styling (large portrait, name, description line, action row); **per-character online voice picker** (audition the match, browse alternatives, pick from your library, or search fish.audio inline); **"Hear a line"** button that synthesizes a representative one-liner from the character's own dialogue in their assigned voice.
|
||
- **AI character notes** — Match local / Match online / Design all now research the play and drop a per-character note (description, gender, speaking style).
|
||
- **Rehearser import auto-save** — uploading a script (PDF/text/FDX/Fountain) saves it to the Library immediately.
|
||
- **Internationalization (i18n)** — interface **language picker** with **German** translation of the UI chrome; English is the source language (`static/js/i18n.js`, extend via `I18N_DICT`).
|
||
- **Progressive Web App** — installable with offline app shell (`manifest.webmanifest` + network-first service worker), iOS web-app meta and safe-area support.
|
||
- **Test suite** — Playwright smoke + functional tests (desktop **and** iPhone/WebKit profiles): app load, sections, clone tabs, PWA, rehearser parse→cast, bundle, i18n.
|
||
- **Build tooling** — opt-in single minified bundle (`npm run minify` → `static/dist/main.min.js`, loaded when `?bundle=1`); architecture & migration notes in `docs/ARCHITECTURE.md`.
|
||
|
||
### Changed
|
||
- **Accessibility → WCAG 2.1 AA** — accessible names on all controls, AA text/badge/button contrast, keyboard-focusable scroll regions (audited with axe-core; 40+ violations → a handful of edge cases).
|
||
- **Performance / mobile stability** — GZip responses; `content-visibility` virtualization for long lists; lazy-loaded images; Rehearser caps decoded-PCM memory to a sliding window (fixes iPhone crashes); leaked `AudioContext` closed; bounded-concurrency bulk operations.
|
||
- **Clone a Voice — reworked GUI** — integrated tab strip (Microphone · Upload · URL/YouTube), clearer sections, scroll-to + obvious "transcribing…" feedback, sample sentence keeps the typed name across language switches.
|
||
- **Fish-Speech tone** — per-line tones now reach OpenAudio S2 via inline `[tag]` markers in the text (the `instruct` field is ignored by S2).
|
||
- **fish.audio import** — de-duplicates voices already in the library and diversifies matches so different characters don't all get the same fallback voice.
|
||
- **Get Voices Online** — tabbed, integrated source switcher; the scrape box lives only under "Direct sources".
|
||
- **Voice library** — editable Voice ID (rename), complete country/accent list (decoupled from language), always-visible **Select all** toggle, redesigned bulk-delete confirmation modal.
|
||
|
||
### Fixed
|
||
- **Screenplay parser** — title-page text, numbered scene headings (`A1 EXT. … EVENINGA1`), `OMITTED`/`CONTINUED` markers and dated page slugs are no longer detected as characters.
|
||
- **Narrator & all voice pickers** now list the full voice library (lazy-loaded if needed).
|
||
- **Cast list controls** wire reliably regardless of when the section mounts; role names no longer truncate; avatars enlarged.
|
||
|
||
---
|
||
|
||
## [1.5.0] — 2026-06-01
|
||
|
||
### Added
|
||
- **Fish-Speech TTS backend** — clones a voice's saved reference WAV (consistent identity) **and** honours inline emotion markers like `(angry)`, `(whispering)`, `(excited)` per line. The only backend that is both WAV-anchored and style-aware; the Rehearser prefers it when available. Configurable via `FISHSPEECH_URL`.
|
||
- **Fish.audio Voice Library browser** (Get Voices Online) — search/filter the ~2M public voices at `api.fish.audio`, preview samples, and one-click **Import** (MP3 → WAV + reference transcript) → an instantly clonable voice.
|
||
- **Cast tab redesigned as character cards** — big avatar, name, language, gender, tags, voice picker, voice-design prompt, "Character soul · LLM brief" with **Develop** (LLM), and per-character **Ignore / Hide / Delete**.
|
||
- **Bulk-edit lines on the Stage** — a **Select** mode adds per-line checkboxes: **Ignore**, **Hide**, **Delete**, Un-ignore, Show-hidden.
|
||
- **Designed voices** — display name is the **character name**, the script becomes a **tag**, and an auto-picked **gender/type avatar icon** replaces the language flag.
|
||
- **Clone a Voice — name-first flow** — name first (drops into the read-aloud sentence), live voice-ID, **auto-transcribe** after trim, **auto-save when ready**, and a **File / URL / Microphone source picker**.
|
||
|
||
### Changed
|
||
- **IMSDb scraper** — resolves the real script via each title's detail-page "Read Script" link instead of guessing a slug.
|
||
- **Rehearser default backend** — `voice_clone` (then `fishspeech`) for consistent identity; the tone-warning explains the trade-off both ways.
|
||
- **Try It Out** — the cramped voice/backend row is now a clean responsive layout.
|
||
|
||
### Fixed
|
||
- **Narrator was silent** — `narratorVoice` now stays in sync with the narrator cast row.
|
||
- **About → Changelog was empty** — `CHANGELOG.md` is now shipped in the image and resolved resiliently.
|
||
|
||
---
|
||
|
||
## [1.4.0] — 2026-06-01
|
||
|
||
### Added
|
||
- **IMSDb browser — list / cover view toggle** — switch between poster-grid and compact list view; preference persisted in `localStorage`.
|
||
- **IMSDb browser — local catalogue cache** — catalogue is cached in `localStorage` for 6 h (matching server cache), making reopening the browser instant.
|
||
- **IMSDb browser — title in fallback** — script title shown on each gradient poster card while the real poster loads.
|
||
- **IMSDb browser — loading spinner** — animated spinner while the catalogue fetches.
|
||
- **IMSDb browse button on Import / Export tab** — the "Browse IMSDb" button is now also available on the Import / Export panel; modal moved to global scope.
|
||
- **Auto-design — detailed progress panel** — each character shows an expandable card during voice design: gender chip, language, voice ID, age, and the full LLM-generated character description with a live spinner.
|
||
- **Auto-design — script title as voice tag** — designed voices receive the script title as their `tag` value so they're easy to filter/find.
|
||
- **Auto-design — LLM endpoint datalist** — the LLM endpoint field is now backed by a `<datalist>` auto-populated from all configured Language Models engines, plus hardcoded defaults (Ollama, vLLM, LM Studio, llama-swap, LiteLLM).
|
||
- **Stage — synthesis progress** — the synth bar is now more prominent (gradient fill, spinner, sticky), each synthesising line pulses with a blue glow, and the page auto-scrolls to the active line.
|
||
- **Stage — tone warning banner** — when a non-style-aware backend (voice_clone, streaming, NVIDIA) is selected and tone is set on lines, a dismissable amber warning banner names the backend and suggests a style-aware alternative.
|
||
- **Bulk-edit tools** — new sticky toolbar in My Voices: select any number of voices with checkboxes, then: **Set tag**, **Hide**, **Unhide**, **Rate**, or **Delete** in one action.
|
||
- **Rehearser — Voice Design default** — the TTS backend picker in Cast now defaults to `voice_design` (style-aware) instead of voice_clone, so tone selections work out of the box.
|
||
|
||
### Changed
|
||
- **Stage — edit button moved to right gutter** — the pencil (edit text) button is now stacked with the note button in the right-side gutter of each dialog block, keeping the block header clean.
|
||
- **Tone / instruct order** — when an emotion is set on a line, the instruction now leads with a directive (`"Speak in a <emotion> manner. <voice profile>"`) so the model prioritises the tone over the base identity description.
|
||
- **My Voices — hidden voices in sub-tabs** — fixed: Cloned, Designed, and Favorites tabs now respect the "Disabled" checkbox filter; hidden voices no longer appear unless explicitly requested.
|
||
|
||
### Fixed
|
||
- **IMSDb covers showing as flat lines** — replaced `aspect-ratio` on a flex child (unreliable in all major browsers) with the `padding-bottom: 150%` wrapper trick, guaranteeing a correct 2:3 poster ratio.
|
||
- **Rehearser TTS backend "No backend available"** — `refreshRehBackends` now triggers the global backend probe if `_ttsBackends` is empty, and registers a `_ttsRefreshHook` so the select stays in sync with the Engines page.
|
||
|
||
---
|
||
|
||
## [1.3.0] — 2026-05-31
|
||
|
||
### Added
|
||
- **Live mic monitor in Clone a Voice** — level-meter and scrolling oscilloscope waveform in the Microphone card. "Check level" / "Stop monitor" buttons, mic gain slider. Recording uses raw mic constraints (no echo-cancel / AGC).
|
||
- **STT engine picker in Clone → Transcript** — pick any configured STT backend when auto-transcribing, bypassing an unavailable Whisper.
|
||
- **Better sample texts** — all 8 languages rewritten to ~38 words / ~15 s, phonetically rich, proper Unicode diacritics.
|
||
|
||
### Fixed
|
||
- **Empty "Read aloud" field** — sample text now reliably populates on load and when navigating to the Clone section.
|
||
- **Recording quality** — `MediaRecorder` requests 256 kbps in Clone and STT→TTS.
|
||
- **OGG file import** — explicit extension list in `accept=`.
|
||
|
||
---
|
||
|
||
## [1.2.0] — 2026-05-29
|
||
|
||
### Added
|
||
|
||
- **Remember last section on reload** — the active section (and Settings /
|
||
Engines sub-page) is persisted in `localStorage`. A hard-reload
|
||
(Ctrl+Shift+R) now returns to the same page instead of always jumping to
|
||
My Voices.
|
||
- **Conversation: live speech preview** — while recording, the active Whisper
|
||
STT backend transcribes accumulated audio every 2.5 s and shows the result
|
||
in the text input field in real time. The input is pre-populated with this
|
||
live guess before the final Whisper result arrives. Also tries the browser's
|
||
Web Speech API first (works on HTTPS / localhost) for even faster results.
|
||
- **Conversation: Voice Activity Detection (VAD)** — recording now auto-stops
|
||
after 1.5 s of silence detected via the Web Audio `AnalyserNode` RMS level.
|
||
A "Sending in X.Xs" countdown appears in the status bar so the timing is
|
||
visible. An **Auto-stop** toggle in the input bar lets users disable VAD and
|
||
revert to click-to-stop. A thin audio-level bar below the status line shows
|
||
microphone volume in real time during recording.
|
||
- **Conversation: hands-free mode** — after the agent finishes speaking, the
|
||
microphone restarts automatically. A **Hands-free** toggle (on by default)
|
||
disables this; clicking the mic button manually always cancels any pending
|
||
auto-restart.
|
||
- **`scripts/release.py`** — automates version bump + CHANGELOG promotion.
|
||
`python scripts/release.py --patch|--minor|--major [--dry-run]` renames
|
||
`[Unreleased]` to the new version, updates compare links, writes `VERSION`,
|
||
commits, and creates an annotated git tag in one command.
|
||
- **Git pre-commit hook** (`scripts/hooks/pre-commit`) — warns (does not
|
||
block) when `.py`/`.js`/`.css`/`.html` files are staged but `CHANGELOG.md`
|
||
or `VERSION` are not. Run `bash scripts/install-hooks.sh` after cloning.
|
||
- **`scripts/install-hooks.sh`** — one-liner to install the hook after a
|
||
fresh clone: `bash scripts/install-hooks.sh`.
|
||
|
||
### Changed
|
||
|
||
- **Config and logs are now bind-mounted local folders** — replaced the
|
||
opaque named Docker volume with `./config/` and `./logs/` host directories.
|
||
`portainer-stack.yml` updated with absolute host paths.
|
||
- **Server writes a rotating log file** — `RotatingFileHandler` writes
|
||
`INFO`-level and above to `./logs/app.log` (rotates at 5 MB, 3 backups).
|
||
|
||
### Performance
|
||
|
||
- **Skeleton loading view** — `index.html` shows an animated shimmer
|
||
placeholder immediately on first paint; fades out once JS finishes loading.
|
||
- **Self-hosted WaveSurfer and MDI icon font** — removed render-blocking CDN
|
||
requests; assets now served locally from `static/vendor/`.
|
||
- **Parallel JS module loading** — restructured `loader.js` into 4 ordered
|
||
batches; round-trips reduced from 17 to 6, 9 files fetched simultaneously.
|
||
- **Version-based JS/CSS cache busting** — versioned assets served with
|
||
`max-age=31536000, immutable`; bumping version invalidates the cache.
|
||
|
||
### Fixed
|
||
|
||
- **LLM returned empty response (Qwen3 thinking mode)** — conversation turn
|
||
now falls back to `reasoning_content` for think-only responses; error
|
||
message hints to add `/no-think` to the system prompt.
|
||
- **Engine settings lost after container recreate** — container names and URL
|
||
overrides now persisted as server settings (`engine_container_names`,
|
||
`engine_local_urls`); restored from server on first page load.
|
||
- **Text-input turns returned 422** — changed `audio` form field to
|
||
`Optional[UploadFile] = None` so text-only turns don't require audio.
|
||
- **Conversation input bar hidden when mic unavailable** — warning box moved
|
||
inside `conv-chat-window` so it never pushes the input bar off-screen.
|
||
- **Browser caches old section HTML** — `loader.js` appends `?v=<timestamp>`
|
||
to every section fetch.
|
||
- Various import errors and container restart issues fixed.
|
||
|
||
---
|
||
|
||
## [1.1.0] — 2026-05-29
|
||
|
||
### Security
|
||
|
||
- **Fixed path traversal in `/api/browse-dirs`** — Added a `_BROWSE_BLOCKED` blocklist
|
||
(`/proc`, `/sys`, `/dev`, `/run`, `/boot`). Requests for paths under these directories
|
||
now return HTTP 403 instead of listing kernel/system files.
|
||
- **Hardened yt-dlp output path** — After a YouTube download completes, the resolved
|
||
output path is verified to be inside `TEMP_DIR` via `.relative_to()`. A file written
|
||
outside the temp directory is rejected with an SSE error event and never registered.
|
||
- **Removed CORS wildcard on `/api/proxy-audio`** — `Access-Control-Allow-Origin: *`
|
||
was unnecessary (all callers are same-origin) and exposed proxied audio to arbitrary
|
||
cross-origin requests. Header removed.
|
||
- **Temp file registry now enforces a TTL** — `_registry` changed to
|
||
`dict[str, tuple[Path, float]]`. `_registry_gc()` evicts entries older than
|
||
`TEMP_FILE_TTL_SECONDS` (default 2 h, configurable via env var) and unlinks their
|
||
files, preventing unbounded disk growth on long-running instances.
|
||
|
||
### Performance
|
||
|
||
- **Settings and routing rules cached in memory** — `_load_settings()` and
|
||
`_load_tts_routes()` previously read from disk on every API request (55+ calls per
|
||
TTS synthesis). Both now use mtime-checked in-memory caches that invalidate
|
||
automatically on write, eliminating redundant file I/O.
|
||
|
||
### Added
|
||
|
||
- **Version number** — `VERSION` file at repo root; read by `core/constants.__version__`
|
||
and surfaced via `GET /api/version`. Displayed as `v1.1.0` in Settings → About.
|
||
- **Text input in Conversation Playground** — a pill-shaped text field and send button
|
||
(→) sit left of the mic button. Pressing Enter or → sends text directly through the
|
||
LLM → TTS pipeline, skipping STT entirely. Makes the playground fully usable without
|
||
a microphone (HTTP context, no mic permission, remote access). The backend
|
||
`/api/conversation/turn` now accepts an optional `text` form field; when set, the
|
||
STT step is skipped and the STT latency row shows `—`.
|
||
- **Container name field on all engine cards** — every TTS and STT engine card (Docker
|
||
stack cards *and* static "Other Local" cards) now always shows the Docker container
|
||
name input row. Previously absent/not-installed cards hid it; now it is always visible
|
||
so the container can be pre-configured before starting.
|
||
- **Connect / Disconnect toggle** — the Connect button now shows "Disconnect" (green,
|
||
`check-network` icon) when already connected and toggles back on click. State
|
||
persists in `localStorage`.
|
||
- **Auto-apply on Connect** — a successful connection probe automatically saves the
|
||
URL to Settings and makes the backend available in TTS/STT dropdown menus immediately,
|
||
without requiring a separate "Use as TTS/STT" click.
|
||
|
||
### Changed
|
||
|
||
- **Connect button redesigned** — moved out of the URL input row into a dedicated
|
||
`dc-controls-row`. Restyled as a solid blue primary CTA (was a small teal outline
|
||
button). Shows a spinner icon while probing.
|
||
- **"Use as TTS / STT" button** — larger padding, bolder teal border, chevron icon,
|
||
tooltip explaining it sets the URL in Settings. Gains `.active` highlight once applied.
|
||
- **Unified controls row on every engine card** — consistent left-to-right order:
|
||
`[Connect/Disconnect]` `[Stop | Start | Restart]` `[Use as →]`. Docker action buttons
|
||
hidden until a container name is entered; Use-as button right-aligned.
|
||
- **`initStaticDockerManagement`** — rebuilt to use the same `dc-controls-row`
|
||
structure as the dynamic Docker stack cards. The existing `.llm-local-ping` button
|
||
is moved from inside the URL row into the controls row at initialisation time.
|
||
- **Backend refactor — `server.py` (5 560 lines → 43 lines)** — all logic extracted
|
||
into single-responsibility modules:
|
||
|
||
| Package | Module | Responsibility |
|
||
|---|---|---|
|
||
| `core/` | `constants.py` | Boot-time env defaults, path constants, version, log buffer |
|
||
| | `registry.py` | TTL-based temp file registry |
|
||
| | `validation.py` | URL validation, SSRF guard, path safety |
|
||
| | `docker_client.py` | Raw Unix-socket Docker HTTP client |
|
||
| | `config.py` | Settings load/save/normalize, backend URL resolution |
|
||
| | `routing.py` | TTS route rules load/save/resolve, language detection |
|
||
| | `audio.py` | Audio conversion, normalisation, auto-trim scoring |
|
||
| | `voice.py` | Voice metadata, backup management, benchmark helpers |
|
||
| | `presets.py` | Voice Design preset load/save, virtual voice resolution |
|
||
| | `tts_helpers.py` | TTS request helpers, streaming, per-backend logic |
|
||
| `routes/` | `admin.py` | Index, favicon, browse-dirs, robots, version |
|
||
| | `settings.py` | `/api/settings`, routing rules, logs, design presets |
|
||
| | `library.py` | All voice CRUD, upload, save, normalize, export/import |
|
||
| | `stt.py` | `/api/transcribe*`, `/api/stt-backends` |
|
||
| | `sources.py` | Voice scraping, proxy-audio, yt-dlp download |
|
||
| | `docker.py` | `/api/local-containers/*`, `/api/probe-url` |
|
||
| | `tts.py` | TTS preview, streaming, voice design, `/v1/*`, backends |
|
||
| | `conversation.py` | Refine-text, effects, export/import, speak, MCP, conversation |
|
||
|
||
`Dockerfile` updated with `COPY core/ core/` and `COPY routes/ routes/`.
|
||
`docker-compose.yml` updated with `./core:/app/core:ro` and `./routes:/app/routes:ro`.
|
||
|
||
- **Frontend refactor — `app.js` (8 744 lines → 16 modules)** — split into
|
||
`static/js/` with `loader.js` loading them sequentially in dependency order:
|
||
|
||
| Module | Lines | Responsibility |
|
||
|---|---|---|
|
||
| `utils.js` | 364 | Core helpers: `$`, `toast`, `escHtml`, theme, language/flag, picker, tabs |
|
||
| `voice-inspector.js` | 397 | 3-pane voice workbench |
|
||
| `voice-sources.js` | 277 | External voice source scraping UI |
|
||
| `integrations.js` | 211 | Code snippet generation (SillyTavern, Open WebUI, HA, curl, MCP) |
|
||
| `routing.js` | 542 | TTS routing rules editor |
|
||
| `settings.js` | 385 | `loadSettings`, `applyAndSaveSettings`, settings panel |
|
||
| `voice-clone.js` | 774 | WaveSurfer, drop zone, mic recording, trim, voice design |
|
||
| `voice-library.js` | 2654 | Full voice library: list, CRUD, benchmark, normalize |
|
||
| `tts-preview.js` | 528 | TTS preview, `fetchTtsPreviewBlob` |
|
||
| `benchmark.js` | 218 | Performance + batch benchmark |
|
||
| `stt.js` | 287 | STT→TTS playground, `refreshSttBackends` |
|
||
| `init.js` | 49 | App bootstrap |
|
||
| `engines.js` | 625 | ElevenLabs browser, custom engine cards, Docker management |
|
||
| `ai-backends.js` | 520 | AI backend cards, LLM snippets, `initStaticDockerManagement` |
|
||
| `generation.js` | 393 | WAV merge, chunked TTS, history, playlist, audio effects |
|
||
| `conversation.js` | 520 | Conversation playground, LLM refinement, import, About |
|
||
|
||
### Fixed
|
||
|
||
- **`chrome://flags/…` URL unreadable in mic-blocked warning** — the global
|
||
`code { background: var(--panel) }` rule caused the URL text to render as
|
||
white-on-light-grey inside the red warning box. Fixed with inline styles
|
||
(`background: rgba(0,0,0,.35); color: #fff`) on the `<code>` element, plus a
|
||
Copy button so users don't need to manually select invisible text.
|
||
|
||
---
|
||
|
||
## [1.0.0] — 2026-05-28
|
||
|
||
Initial feature-complete release.
|
||
|
||
### Added
|
||
|
||
- **Voice library** — clone voices from audio samples; design voices from text
|
||
descriptions using instruction-based synthesis; benchmark synthesis speed (RTF);
|
||
normalize loudness; export/import voice packages as ZIP bundles.
|
||
- **TTS backends** — Qwen3 TTS (Voice Clone, Voice Design, Custom Voice, Streaming),
|
||
NVIDIA Magpie / Zeroshot / Flow, Kokoro FastAPI, VibeVoice, XTTS v2, ElevenLabs.
|
||
- **STT backends** — OpenAI Whisper (port 8010), faster-whisper-server, whisper.cpp,
|
||
Groq Whisper (cloud, free tier), NVIDIA Parakeet ASR. Real transcription probe
|
||
in health check (not just TCP reachability).
|
||
- **App Routing** — per-app / per-voice / per-language TTS routing rules with
|
||
automatic language detection and optional before/after sound effects.
|
||
- **Conversation Playground** — full STT → LLM → TTS pipeline with real-time SSE
|
||
streaming, latency stats panel (STT / LLM TTFT / LLM total / TTS / Total), turn
|
||
history, system prompt, and insecure-context warning.
|
||
- **Engines section** — LLM / STT / TTS sub-pages; Docker container management
|
||
(Start / Stop / Restart via Docker socket); custom engine cards; ElevenLabs voice
|
||
library browser.
|
||
- **Performance Benchmark** — single-voice and batch benchmark with RTF tracking,
|
||
sparkline trend, and persistent history.
|
||
- **Audio effects** — reverb, chorus, delay, compressor, gain, pitch shift
|
||
(via `pedalboard`).
|
||
- **Chunked TTS + generation history** — long-text synthesis split into chunks,
|
||
per-chunk playback, playlist export as WAV.
|
||
- **MCP server** — built-in JSON-RPC 2.0 endpoint at `/mcp`; tools: `speak`,
|
||
`transcribe`, `list_captures`, `list_profiles`.
|
||
- **LLM refinement & persona rewriting** — clean up STT transcripts or rewrite
|
||
responses with a chosen persona via any OpenAI-compatible LLM endpoint.
|
||
- **Connect Apps** — ready-made config snippets for SillyTavern, Open WebUI,
|
||
Home Assistant, curl, and MCP (`claude mcp add` one-liner).
|
||
- **Voice sources** — scrape voice assets from Aiartes, Freesound, GitHub, and
|
||
Google Drive; YouTube download via yt-dlp; quick import directly to library.
|
||
- **OpenAI-compatible proxy** — `/v1/audio/speech` and `/v1/audio/transcriptions`
|
||
for drop-in use with Open WebUI, SillyTavern, and Home Assistant.
|
||
- **Settings** — sub-pages: General, Connections, Playback, Captures, Payloads,
|
||
Storage, API Keys, Logs, About.
|
||
- **Voice Design presets** — saved persona templates for instruction-based synthesis;
|
||
virtual `vd_…` voices usable from external apps without exporting WAV files.
|
||
- **Multilingual support** — language/flag pickers, per-language preview texts,
|
||
`LANG_FLAG_DEFAULT` mapping for 16 languages.
|
||
- **Tags, ratings, and metadata** — per-voice tags with autocomplete, star ratings,
|
||
gender label, country flag.
|
||
- **Dark/light theme** — toggle with persistence in `localStorage`.
|
||
- **Docker socket integration** — Start/Stop/Restart Docker containers from the UI
|
||
via raw Unix socket HTTP; container health visible in engine cards.
|
||
|
||
---
|
||
|
||
[Unreleased]: https://github.com/mARTin-B78/tts-voice-creator-clone-and-design-2/compare/v1.8.1...HEAD
|
||
[1.8.1]: https://github.com/mARTin-B78/tts-voice-creator-clone-and-design-2/compare/v1.8.0...v1.8.1
|
||
[1.8.0]: https://github.com/mARTin-B78/tts-voice-creator-clone-and-design-2/compare/v1.7.0...v1.8.0
|
||
[1.2.0]: https://github.com/mARTin-B78/tts-voice-creator-clone-and-design-2/compare/v1.1.0...v1.2.0
|
||
[1.1.0]: https://github.com/mARTin-B78/tts-voice-creator-clone-and-design-2/compare/v1.0.0...v1.1.0
|
||
[1.0.0]: https://github.com/mARTin-B78/tts-voice-creator-clone-and-design-2/releases/tag/v1.0.0
|