Two unclosed <div class="card"> tags in the TTS settings markup left the
add-engine dialog nested inside the TTS panel's DOM subtree, so it
rendered at zero size whenever another Engines sub-tab (e.g. Speech
Recognition) was active. Also fixed a `let` declared after its first
use, which threw a ReferenceError on every Engines-tab page load.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- App Routing's output-voice field gets the same searchable
avatar-thumbnail dropdown used elsewhere, as a browse button
alongside the existing free-text input (which must stay editable to
target vd_ Voice Design presets not in the voice library).
- My Voices table: Gender and Rating filter dropdowns existed in the
JS (populateLibraryFilters, libraryFilterMatch) but their <select>
elements had been dropped from the visible layout after an earlier
redesign, replaced with hidden dead placeholders just to keep the
code from erroring - and since populateLibraryFilters() early-returns
if any of the three elements are missing, this silently broke the
already-visible Language/Type dropdowns too. Restored the real
elements and removed the hidden scaffold; added new Tag and Group
dropdowns wired to the same filter state the sidebar chips use.
- Audio effects failing with "pedalboard is not installed" despite
requirements.txt listing it: the package WAS installed, but its
native extension (pedalboard_native) links against libatomic.so.1,
an OS-level shared library missing from the python:3.11-slim-bookworm
base image. Added libatomic1 to the Dockerfile and rebuilt - verified
`import pedalboard` now succeeds in the running container.
- Relabeled "edit ID"/"copy ID" to "rename filename"/"copy filename"
in the voice inspector - the feature already renamed the underlying
.wav/.meta.json/.reference.txt/picture files via the existing
/api/voice/rename endpoint, it just wasn't obvious "ID" meant
"filename."
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Character sheet generation is already available from "Cast Characters"
in the Casting flow, so the standalone button here was a duplicate
entry point. Removed it and made "Cast as audiobook" the primary blue
action, moved to the end of the toolbar as the clear next step in the
pipeline.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Replaces the horizontal avatar strip above the transport bar with a
sidebar next to the script page, reusing Read Aloud's Casting sidebar
classes (.ab-cv-side/.ab-char-item) directly instead of a separate
look. Same collapse-to-avatars control, now shows each character's
line count, and clicking a character scrolls the script to their
first line. Removed the CSS/HTML this replaces.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Characters/Cast gains a "Casting" button back to the active casting
session, the pipeline stepper renders on the Library section, and the
stepper's Cast Characters stop navigates instead of side-effect-running
sheet generation. Prompt generation: 4096-token budget (1600 truncated
the four-prompt JSON so two fields silently arrived empty), truncated
answers salvage completed fields, all-empty responses fail loudly, and
partial results name the missing prompts. The PDF-extraction progress
pill is enlarged and vertically centered.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Speed: OCR renders only the heading band (was full page at 2.5x),
Tesseract worker freed after extraction, name-underline index cached
instead of rebuilt per segment, constant regexes hoisted, old drafts
migrate segment page numbers once at load (single-path feed renderer).
Quality: global [hidden]{display:none!important} ends the empty-box bug
class; racy deferred cast-restore + _readerSuppressCastRestore flag
replaced by a synchronous, caller-wins restore; card collapse defaults
move to data-collapse-default markup; duplicated join/colour/alias/LLM-
target helpers now delegate to their canonical implementations; dead
reader state removed; stepper hide-guard fixed for the merged Source key.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Add a 6-stage pipeline stepper (Source -> Cast Audiobook -> Cast Characters
-> Script Rehearser -> Generate MP3s -> Audiobook) with direct, non-destructive
jumps between stages and a prominent guided-tour look
- Split PDF import into an explicit "load" then "Extract Text" step, with
in-browser OCR (Tesseract.js, vendored) to recover chapter headlines baked
into a PDF as images instead of real text
- Fix casting feed silently merging pages after leaving/returning: segments
now carry their own page number instead of re-guessing it from text
- Fix excessive "Unknown" speaker attribution: restore the attribution LLM's
output token budget, which had been cut roughly in half and was truncating
dialogue-dense passages
- Fix Theater Play library cards failing to open (dead pre-migration
IndexedDB API calls, missing section navigation)
- Fix bulk "Set tag" wiping a voice's existing tags instead of adding to them
- Start merging Casting's feed with Script Rehearser's Stage UI: collapsible
character sidebar, shared "paper" page styling, inline text editing
- Fix a performance regression from that merge (per-row listeners on every
redraw) by moving to event delegation
- Various layout/clutter fixes: hide reader chrome until a document is
loaded, collapse secondary settings by default, fix overlapping toolbar
icons, fix duplicate "opening" notifications
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Root cause 1 (v1.12.18): `window.loadVoiceLibrary = () => loadVoiceLibrary()`
overwrites the global binding the arrow function references, causing immediate
RangeError: Maximum call stack size exceeded on every call. Changed to direct
assignment `window.loadVoiceLibrary = loadVoiceLibrary`.
Root cause 2 (v1.12.18): `loadSettings()` called `renderSettingsAbout()` which
lives in conversation.js (batch E), loaded after init.js. Guard added with
typeof check; nav.js already calls it safely when the About section opens.
Also (v1.12.18): s-library.html duplicated cl-book-filter / cl-search / cl-grid
from s-characters.html, breaking getElementById. Characters panel in Library
now redirects to s-characters instead of duplicating its DOM nodes.
Also (v1.12.19): engines.js triggers a second loadVoiceLibrary() after nav.js
already rendered voices, blanking the list briefly. Second call now silently
re-fetches without clearing the list when voices are already present.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Voice selection list now shows flag + gender symbol on each row.
Batch results table adds Lang and Gender columns. All columns are
sortable by clicking the header (↑↓ indicator); defaults to Factor
descending. Sort logic handles strings (locale) and numbers uniformly.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Voices table: removed Length, Duration, Time — only Factor and WPM remain
from benchmark data. Setup → Benchmark batch results now shows Duration,
Factor (sorted fastest-first, colour-coded), Time, and WPM. Factor replaces
Avg RTF with the same data flipped to a more intuitive direction (higher=better).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The old combined `21.8s · 1.30x` Speed cell is replaced by two separate
sortable columns:
- Factor (x.xx): audio÷render multiplier, green/amber/red colour-coded
- Time (Xs): total render time
Column order: Length · Duration · Factor · Time · WPM · Seed · dBFS · ...
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
WPM (words per minute) is derived from the benchmark audio duration and
sentence word count, revealing how fast a voice speaks — independent of
GPU speed. 130–180 wpm is comfortable for audiobooks. Column is sortable.
Info (ⓘ) icons on Speed and WPM headers explain both metrics on hover.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add Source input to voice inspector panel (saves to origin in meta.json)
- Auto-detect fish-audio voices via tag — SOURCE column now shows
"fish-audio" for tagged voices even without an explicit origin field
- Bulk "Set source" action in the multi-select toolbar
- Source cell in table reflects the live-displayed value and updates
immediately when changed via the inspector
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Conversation: adjustable noise gate slider (RMS threshold + 300 ms
minimum burst duration) prevents short noise spikes from triggering
STT; level meter shows gate position as a blue marker
- Conversation stats panel now collapsible (chevron button) to free
chat width; floating expand button restores it; state persists
- Read Aloud: text search input in PDF toolbar (Enter = next hit,
Shift+Enter = previous, Esc = clear)
- Sidebar: tooltip now works for all item types including sub-items
that had no .nav-label span (text extracted by stripping icon/badge)
- Sidebar active section indicator added (.nav-tree-item.active was
previously unstyled — active section now has bg + right accent bar)
- Casting audiobook: prompt instructs LLM to handle ?« / !« endings
and unclosed » at passage end as dialogue; deterministic fallback
also handles unclosed opening quote
Bumps to v1.12.2.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Sidebar collapses to a 56px icon rail; hovering flies the full menu out as
an overlay (icons + titles/nested items). Language picker moved to Settings.
- Unify settings collapsibles to the app's standard card-collapse style:
Conversation, Read Aloud (drag-&-drop now inside), Try It Out, Casting panel.
- Read Aloud: reordered (settings → toolbar → document → transport/synth) and
the document fits the viewport height so controls below stay visible; remove
the redundant My Books card (lives in Library → Books).
- Conversation: stacked full-width config, fills viewport height; fix
intermittent webm decode in hands-free mode (recorder restarts cleanly,
in-browser WAV encode); barge-in via Live agent.
- Fix casting feed overflow that pushed the sidebar off-screen.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Restructure the sidebar into Voices / Speak / Setup groups and fuse the
Read Aloud and Script Rehearser features into one workspace:
- New combined Library section (Books · Theater Plays · Characters/Cast)
with cross-links (Rehearse a book, Read Aloud a play).
- Character tags like voice tags: one record, many productions, reusable
across books/scripts; seeded with origin book, editable in the editor.
- Shared cast resolution: opening a production fills empty cast slots from
the character roster (matched by book OR tag); voice choices written back
on save. Productions joined by normalized title (prodKey).
- Tags nav group for voices (distinct tags + counts, Cloned/Designed/Fav).
- Fold standalone Characters into Library; Read Aloud now a single entry.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Exposes the `speed` parameter (0.5–2.0) on the TTS `/v1/audio/speech`
request so audio is generated at the target tempo natively via the
faster-qwen3-tts backend rather than post-processing.
- Backend: `/api/tts-preview` now extracts and forwards a `speed`
override (clamped 0.1–4.0) through the same extra-params mechanism
already used for seed/temperature; backends that reject it fall back
cleanly via `_post_tts_with_fallback`.
- Try it out: "Native Speed" number input (0.5–2, step 0.05) added to
the text card; value persists in localStorage per browser; passed as
`extra` through `createTtsAudioSource` and `generateChunkedTts`.
- Read aloud: "Native Speed" control added to the generation controls
row alongside seed/temperature; included in `readerGenParams()` and
saved/restored with library documents (each book tracks its own speed).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Character sheets: stop sending the dead localhost:11434 default; fall back to
the server's configured llm_url so extraction uses the same working LLM.
- Audiobook casting: hoist highlightText to module scope so the "Review & cast"
manual-correction preview renders its segment rows again.
- Stage: unify narrator/dialog font size, add A-/A+ play text-size control
(scales A4 + paginated views), and make the cast chip list collapsible.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add a persistent, book-scoped Character Library (new Characters section,
IndexedDB) that auto-fills from Character-sheet analysis with editable cards.
Enrich extraction with six narrative fields (backstory, relationships,
motivation, fears, mannerisms, voice/speech) plus the greyscale Good↔Evil
alignment bar, arc arrow, and 5-area Deep Analysis. Add ⋯ separators between
non-contiguous passages in Recast unknown, and harden dialogue attribution
against hallucination with same-language emotion tags.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Added a new 'Table View' button to the library sort bar.
- Implemented a CSS grid layout to display voice properties in a high-density table.
- Added sortable column headers (Name, Lang, Gender, Speed, dBFS, Length, Rating, Source, Seed, Note, Tags, Active).
- Aligned CSS grid to account for the bulk edit checkbox injection.
- Added drag-and-drop profile image support to the Inspector's large avatar.
- Ensured picture updates instantly synchronize across the List, Table, and Inspector views.
Read Aloud (new "Vorlesen" tab):
- PDF (real page render + overlay highlight) / TXT reader with live word
highlighting, voice + speed, per-sentence synthesis-state colours, zoom
(fit-width/height, two-page, ±), resume, and a server-side book library
(syncs across devices; per-unit MP3 audio fetched on demand).
Book -> multi-speaker audiobook:
- "Cast as audiobook" attributes dialogue to characters via the LLM
(guillemet/quote-style aware, turn-taking, recent-context), with a
deterministic speech-tag fallback. Editable preview, non-blocking live
casting panel, then auto-saved as a reopenable Script Rehearser play.
- Audiobook export: synthesise every cast line -> one MP3 per chapter.
Character sheets:
- LLM-extracted, self-filling RPG-style sheets (with page+quote sources)
in both Read Aloud and the Rehearser.
Also: MP3 storage + per-page/sentence export, voice-library "Precompute
embeddings" pre-warm, German "Vorlesen" i18n + flag language toggle,
large-PDF performance (lazy raster, buffer/canvas eviction, yielded parse),
and the Seed Finder changelog entry.
New: routes/reader.py, POST /api/attribute-dialogue, POST /api/character-sheets,
static/js/{reader,audiobook,character-sheets}.js, static/sections/s-reader.html.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Storage: IndexedDB ('reh-library' DB, 'rehearsals' store).
Blobs (recorded mic audio) stored natively; no base64 overhead.
Library panel (Phase 1):
- List of saved rehearsals sorted by last update
- Each row: character-colour avatars, title, progress bar,
line count, recorded-clip count, date
- Active record highlighted in accent colour
- Open / Export .reh / Delete per-row buttons
- Import .reh button at the top
Transport bar (Phase 3):
- Save button (updates existing record if savedId set, else creates new)
- Export button (downloads current state as .reh JSON file)
Phase 4 (session complete):
- Save to library and Export .reh buttons
Data format (.reh file):
- JSON with version=1, title, script text, cast map, backend, lineIndex
- clips serialized with audio as base64 strings (mime + b64 fields)
- clipsFromJson restores Blob objects on import
Behaviour:
- Parsing a brand-new script resets savedId → safe to save as new record
- loadRecord() restores script, cast, lineIndex, clips and jumps to Phase 2
- renderLibraryList() called on init, after save, after delete, after import
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- A4 paper page (white, serif font, shadow) renders the full script at once
so you can read ahead while rehearsing
- Each dialog block shows: character avatar (voice library photo or initial
letter), character name in their colour, TTS/Me badge, and dialog text
- Active line gets a blue left-border highlight and auto-scrolls into view
- Transport bar (sticky, above the script):
- Cast strip: all character avatars at a glance
- ⏮ Prev / ▶ Play all / ⏹ Stop / ⏭ Next / 🔁 Repeat
- Progress bar + line counter
- Auto-play: TTS lines synthesize, play, auto-advance; 'me' lines pause
and slide up a sticky recording overlay at the bottom of the page
- Recording overlay shows the line to speak, oscilloscope + meter,
Record / Stop / Keep & continue / Re-record / Skip
- voice-library.js now exports window._voices so the rehearser can resolve
voice IDs to has_picture flags and fetch /api/voice/picture/{id}
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
## Script Rehearser (new feature)
- New section s-rehearser.html + rehearser.js + nav/loader wiring
- Phase 1: paste/upload script (.txt), auto-detect characters from
'CHARACTER: dialog' or ALL-CAPS screenplay format
- Phase 2: assign a TTS voice per character, or mark 'I play this'
- Phase 3: step-through rehearsal — synthesizes other characters via TTS,
shows level-meter + oscilloscope for your own lines, records them from mic
- Phase 4: session summary with per-line audio playback + download
## Connect Apps
- Removed duplicate standalone MCP/speak/hotkey full-width cards
- Kept the integration-grid cards (they use the real server URL from JS)
- Added Global Hotkey Daemon as a proper integration card with snippet-hotkey
populated by integrations.js (uses proxyBase URL dynamically)
## About page
- GET /api/changelog endpoint reads CHANGELOG.md and returns it as text
- Collapsible 'Changelog' <details> card fetches and displays it lazily
## Try It Out
- Reorganised into three cards: Voice & backend / Text to synthesize / Generate
- Backend help panel moved below the voice row (not in the same flex row)
- Style instruction field gains a dynamic badge ('style-aware ✓' / 'weak style')
and a yellow warning when a non-style-aware backend is selected while the
field is filled — wired to both backend-select change and input events
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Loader:
- Post-init modules (engines, ai-backends, generation, conversation) now
load AFTER the skeleton is removed instead of before. The UI is visible
~500 ms sooner on average; those four modules load while the user is
already browsing Voices / Clone / Design.
- Removed the sttReady event approach that was triggering a duplicate
/api/stt-backends call; init.js already populates all STT selects once
on startup.
Skeleton:
- Replaced the card-grid placeholder with a two-column workbench skeleton
(voice list rows on the left + inspector placeholder on the right) that
matches the real My Voices layout.
Connect Apps:
- MCP Server, /speak REST endpoint, and Global hotkey daemon sections
moved from Settings → About to Connect Apps, where they belong.
- About page now has GitHub + Releases links instead.
Clone section:
- Added a hint note beneath the sample-text textarea reminding the user
to replace the placeholder name (Sam / Alex / Marco …) with their own.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Root cause: VERSION was not volume-mounted, so the server always reported
1.2.0, causing browsers to serve 1-year-immutable cached JS even after
code changes.
Fixes:
- docker-compose.yml: add ./VERSION:/app/VERSION:ro volume mount.
After `docker compose up -d`, the server reads the current VERSION file
and JS is cache-busted by the correct version string.
- loader.js: append session timestamp to _appVersion so every page load
generates a unique JS URL. JS is always fresh regardless of whether
VERSION is current, at the cost of one network round-trip per file per
session (acceptable for a local tool).
- s-clone.html: embed EN default text directly in the textarea so the
field is never empty even before JS runs.
- voice-clone.js: remove the 'skip if already filled' guard in
initCloneSampleText so navigating back always resets to the language
text; call refreshSttBackends on load with sttReady event fallback.
- stt.js: dispatch 'sttReady' event after all STT listeners are wired.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
1. Sample text: expose initCloneSampleText as a window function and call
it from nav.js runSideEffects when the clone section is activated,
ensuring the textarea is always populated even if the IIFE ran before
the element existed.
2. Better sample texts: all 8 languages rewritten to ~38 words / ~15 s,
first-person, phonetically rich, proper Unicode diacritics.
3. Live mic monitor: a level-meter (18-bar) + scrolling oscilloscope
canvas (ring-buffer, 300 px, colour-coded) added to the microphone
card. "Check level" / "Stop monitor" buttons start/stop it
independently; clicking Record starts it automatically.
Uses raw mic constraints (no echo-cancel / AGC) for cleaner voice clone
audio. Mic gain slider and dB readout included.
4. Recording quality: MediaRecorder now requests audioBitsPerSecond:256000
in both voice-clone.js and stt.js.
5. STT engine picker: Recognition engine <select> + Refresh button added
above the Auto-transcribe button in Step 3. refreshSttBackends() now
syncs both stt-tts-stt-backend and clone-stt-backend. The transcribe
call passes the chosen backend to /api/transcribe.
6. File input: explicit extension list added to accept= for OGG/OPUS.
7. CSS: .mic-live-wave style added (dark/light theme variants).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds the 'Read aloud' sample-sentence box back to the microphone card
in the Clone a Voice tab. Language switcher covers EN/DE/IT/ES/FR/PT/NL/PL
with phonetically diverse sentences (same as the My Voices panel).
Uses the existing .sample-read-box / .sample-sentence CSS so it looks
identical to the equivalent panel in the voice library.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Conversation playground:
- Live speech preview: MediaRecorder sends accumulated audio to
/api/transcribe-bytes every 2.5 s; interim Whisper result shown in
the text input field while recording. Web Speech API tried first as
a faster path when available (HTTPS/localhost).
- VAD auto-stop: AudioContext AnalyserNode measures RMS every frame;
auto-stops after 1.5 s silence with a visible countdown. Auto-stop
toggle to revert to click-to-stop.
- Hands-free mode: mic auto-restarts after the agent finishes speaking
via audio.ended event + generation-counter cancellation. Hands-free
toggle (on by default) to disable.
Navigation:
- Persist active section and sub-page in localStorage; hard-reload
returns to the same page instead of always jumping to My Voices.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add VERSION file (1.1.0) at repo root
- core/constants.py: expose __version__ read from VERSION file
- routes/admin.py: GET /api/version endpoint returns {version}
- Settings → About: display "v1.1.0" next to app name via /api/version fetch
- CHANGELOG.md: full rewrite following Keep a Changelog + Semantic Versioning
- [Unreleased] staging section at top
- [1.1.0] 2026-05-29 — security, perf, refactor, UX changes from this session
- [1.0.0] 2026-05-28 — all pre-session features documented
- Compare links at bottom pointing to GitHub
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Frontend:
- Add pill-shaped text input + send button (→) to the left of the mic button
- Enter key or → click sends text directly without recording audio
- Input is disabled while a turn is processing; cleared on submit
- Welcome message updated to mention both input methods
- New CSS: .conv-input-bar, .conv-text-row, .conv-text-inp, .conv-send-btn,
.conv-divider (visual separator between text and mic sections)
Backend:
- /api/conversation/turn: audio is now optional (UploadFile | None)
- New text form field — when provided, STT step is skipped and text is
used as the transcript directly; SSE emits transcript event with stt_ms=null
- Raises 400 if neither audio nor text is supplied
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- STT section now has a Quick test panel: select backend, hit mic button,
see transcript. Records via MediaRecorder, posts to /api/transcribe-bytes.
- _transcribe_audio detects the whisperx-gpu 'NoneType/to' error (caused by
pyannote/speaker-diarization-3.1 requiring a HuggingFace token) and
replaces it with an actionable message explaining how to fix it.
- _to_wav_16k added for STT audio conversion (Whisper/wav2vec2 expect 16kHz).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Port 8080 is Open WebUI — it passes /health + /v1/models checks but
returns 405 on POST /v1/audio/transcriptions. Updated probe logic to:
- treat 405 as 'endpoint missing, try next path'
- treat non-JSON 500 as broken, JSON-500 with detail as 'audio too short' (ok)
- use 500ms silence WAV instead of 1-frame (too tiny for alignment models)
Changed whisper.cpp default from :8080 to :8085 to avoid clash with
Open WebUI. Updated s-llms.html placeholder and code snippet accordingly.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Conversation Playground (new section):
- WhatsApp-style chat UI with user/assistant speech bubbles
- Click-to-record mic button using MediaRecorder API
- STT → LLM streaming → TTS pipeline via SSE (POST /api/conversation/turn)
- LLM tokens stream into assistant bubble in real time
- Audio auto-plays when TTS synthesises the reply
- Right-side stats panel: STT / LLM TTFT / LLM total / TTS / Total with bar chart
- Turn history list with per-turn total time and pass/fail indicator
- Configurable: STT backend, LLM URL + model, TTS backend + voice, system prompt
- Conversation history maintained across turns (last 20 messages sent to LLM)
- GET /api/conversation/llm-models proxies model list from any OpenAI-compatible LLM
XTTS v2 backend:
- Registers xtts as a first-class TTS backend (xtts_url setting, display name,
capabilities, health/voice discovery, OpenAI-compatible generation)
- Added XTTS URL field to Settings → Connections
- Use-as-TTS button now saves to xtts_url (not tts_url)
- Batch benchmark backend select now refreshes alongside perf/preview selectors
VibeVoice fix:
- Added /voices to _TTS_VOICE_ENDPOINTS so VibeVoice voices are discovered
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>