The static-asset caching middleware used BaseHTTPMiddleware, which has a
known Starlette bug: a client disconnecting mid-StreamingResponse (the new
live-attribution SSE stream hitting its idle timeout) raced its internal
task group and raised "RuntimeError: No response returned", crashing that
request. Rewritten as plain ASGI middleware that only touches headers via
the raw send callable, removing the race.
Also found the real cause of the casting timeouts/405s: the streaming
attribution endpoint had its own lock instead of sharing the one the
blocking endpoint already used to serialize on the LLM's single slot -
letting a stream call and its own blocking fallback fire concurrently,
exactly the ghost-request pile-up that lock was built to prevent. Unified
onto one lock and added server-side logging for stream failures.
The "LLM Thinking" pane now shows the model's actual <think> reasoning
instead of the in-progress JSON answer echoed back at the user.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Read Aloud (new "Vorlesen" tab):
- PDF (real page render + overlay highlight) / TXT reader with live word
highlighting, voice + speed, per-sentence synthesis-state colours, zoom
(fit-width/height, two-page, ±), resume, and a server-side book library
(syncs across devices; per-unit MP3 audio fetched on demand).
Book -> multi-speaker audiobook:
- "Cast as audiobook" attributes dialogue to characters via the LLM
(guillemet/quote-style aware, turn-taking, recent-context), with a
deterministic speech-tag fallback. Editable preview, non-blocking live
casting panel, then auto-saved as a reopenable Script Rehearser play.
- Audiobook export: synthesise every cast line -> one MP3 per chapter.
Character sheets:
- LLM-extracted, self-filling RPG-style sheets (with page+quote sources)
in both Read Aloud and the Rehearser.
Also: MP3 storage + per-page/sentence export, voice-library "Precompute
embeddings" pre-warm, German "Vorlesen" i18n + flag language toggle,
large-PDF performance (lazy raster, buffer/canvas eviction, yielded parse),
and the Seed Finder changelog entry.
New: routes/reader.py, POST /api/attribute-dialogue, POST /api/character-sheets,
static/js/{reader,audiobook,character-sheets}.js, static/sections/s-reader.html.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
config/ and logs/ are now host directories visible on the filesystem.
- docker-compose.yml: ./config → /home/app/.config/tts-voice-creator:rw
./logs → /logs:rw
Named volume tts-voice-creator-clone-and-design-2 removed.
- portainer-stack.yml: same change with absolute host paths.
- server.py: RotatingFileHandler writes INFO+ to /logs/app.log
(maxBytes=5MB, backupCount=3). Falls back gracefully if /logs
is not writable.
- .gitignore: track config/ and logs/ dirs via .gitkeep but exclude
settings.json, *.log and backups from version control.
Benefits:
- Settings and logs are human-readable on the host at any time
- Survives docker-compose down -v (was lost with named volume)
- Easy backup: cp -r config/ logs/ to any destination
- Can edit settings.json directly if needed
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- STT section now has a Quick test panel: select backend, hit mic button,
see transcript. Records via MediaRecorder, posts to /api/transcribe-bytes.
- _transcribe_audio detects the whisperx-gpu 'NoneType/to' error (caused by
pyannote/speaker-diarization-3.1 requiring a HuggingFace token) and
replaces it with an actionable message explaining how to fix it.
- _to_wav_16k added for STT audio conversion (Whisper/wav2vec2 expect 16kHz).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
STT models (Whisper, WhisperX VAD, wav2vec2 alignment) all expect 16kHz.
Sending 24kHz caused whisperx's VAD to miss speech segments, leaving
alignment with None inputs → 'NoneType has no attribute to' crash.
Added _to_wav_16k() and use it in conversation/turn and transcribe-bytes
endpoints. Health check probe also uses 16kHz silence for consistency.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Port 8080 is Open WebUI — it passes /health + /v1/models checks but
returns 405 on POST /v1/audio/transcriptions. Updated probe logic to:
- treat 405 as 'endpoint missing, try next path'
- treat non-JSON 500 as broken, JSON-500 with detail as 'audio too short' (ok)
- use 500ms silence WAV instead of 1-frame (too tiny for alignment models)
Changed whisper.cpp default from :8080 to :8085 to avoid clash with
Open WebUI. Updated s-llms.html placeholder and code snippet accordingly.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
_stt_backend_health now sends a minimal WAV to the transcription endpoint
after passing /health. A 500 response marks the backend unavailable,
catching containers that pass health checks but crash on model load (e.g.
CTranslate2 built without CUDA support).
Error messages from _transcribe_audio now include the backend URL and
replace generic 'Internal Server Error' with an actionable explanation.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
_transcribe_audio now extracts the response body on HTTP errors so the
actual cause (e.g. 'CTranslate2 not compiled with CUDA support') reaches
the user instead of '500 Server Error for url: ...'. Conversation panel
also strips the 'HTTP 500:' prefix to show only the meaningful part.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
cardType() was defined inside initLlmsSection() IIFE but called from
renderLocalContainers() which is outside that scope, causing a silent
ReferenceError that reset every Connect click to failure. Moved cardType
to module scope.
Custom STT cards (e.g. whisperx-gpu) are now included in /api/stt-backends
and appear in the Conversation STT dropdown. Added _normalize_service_url()
so 0.0.0.0 URLs in stored cards are rewritten to host.docker.internal for
server-side health checks. _transcribe_audio() tries /transcribe as fallback
for custom backends that don't expose /v1/audio/transcriptions.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
probe-url now accepts a type param (llm/stt/tts) and checks service-
specific endpoints: LLM → /v1/models with data[] key, STT → /health
then /v1/models, TTS → /health then /voices endpoints. Random websites
and wrong services are now rejected. Connect passes the card's section
type; success toast shows which endpoint responded.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Replaces the filter-to-available-only approach with full lists that
include unavailable backends (disabled, marked ✗) so users can see
what's broken. Also extends STT retry fallback to cover HTTP 500 from
wrong model names (fixes faster-whisper CTranslate2 CUDA build issue).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Engine URL inputs (Ollama, vLLM, faster-whisper, etc.), custom engine
cards, and the refinement/conversation LLM URLs were stored only in
localStorage and lost on browser data clear. All four are now synced
to settings.json via _patchSettings() with localStorage as fast
initial fallback.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Parakeet RNNT NIM measured at 550–605 MB (unified RAM), Magpie TTS at
982 MB, and Qwen3-TTS model at 4.3 GB on disk — replacing the placeholder
estimates that were too high across the board.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Conversation Playground (new section):
- WhatsApp-style chat UI with user/assistant speech bubbles
- Click-to-record mic button using MediaRecorder API
- STT → LLM streaming → TTS pipeline via SSE (POST /api/conversation/turn)
- LLM tokens stream into assistant bubble in real time
- Audio auto-plays when TTS synthesises the reply
- Right-side stats panel: STT / LLM TTFT / LLM total / TTS / Total with bar chart
- Turn history list with per-turn total time and pass/fail indicator
- Configurable: STT backend, LLM URL + model, TTS backend + voice, system prompt
- Conversation history maintained across turns (last 20 messages sent to LLM)
- GET /api/conversation/llm-models proxies model list from any OpenAI-compatible LLM
XTTS v2 backend:
- Registers xtts as a first-class TTS backend (xtts_url setting, display name,
capabilities, health/voice discovery, OpenAI-compatible generation)
- Added XTTS URL field to Settings → Connections
- Use-as-TTS button now saves to xtts_url (not tts_url)
- Batch benchmark backend select now refreshes alongside perf/preview selectors
VibeVoice fix:
- Added /voices to _TTS_VOICE_ENDPOINTS so VibeVoice voices are discovered
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Rename "LLMs" section to "Engines" with brain icon
- Add three sub-pages (Language Models / Speech to Text / Text to Speech) following the same nav-tree pattern as Settings and My Voices
- Rewrite s-llms.html: three s-engines-page divs, docker container grids (dc-grid-tts, dc-grid-stt), VibeVoice card in TTS section, static cloud API cards per category
- Add navEnginesCat() and applyEnginesPage() to nav.js; engines tree open/close in showSection()
- Remove obsolete initLlmCatTabs IIFE; fix dc-refresh-btn from ID to class-based querySelectorAll
- Make integration cards in Connect Apps collapsible (collapsed by default) with favicon/icon prepended to h3
- Unify all Engines card grids to minmax(380px, 1fr) so local, docker, and cloud cards are the same width
- Add s-engines-page CSS (display:none / is-active:flex)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- New s-performance section: dedicated Benchmark nav entry with run form,
per-session results table, RTF trend badge (faster/slower/stable), SVG
sparkline chart, and a History card backed by localStorage (last 50 sessions)
- Performance tab removed from Try It Out; element IDs unchanged so JS works
- renderIntegrationSnippets: adds Python MCP server + Claude Code .mcp.json
config snippets to the Connect Apps page (integration-card-wide styling)
- Save handler: persists all Captures settings fields (stt_language,
stt_preferred_backend, auto_refine, refine_model, refine_* toggles,
captures_default_voice) alongside existing settings
- CSS: integration-card-wide accent border, benchmark history rows, trend
badges, sparkline wrapper, bench-history-toolbar
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
New TTS backend: Kokoro FastAPI (82M) — OpenAI-compatible, 11 built-in
voices, only shows when server is reachable (~300 MB CPU, ~0.1× RTF).
New STT backends in Transcribe dropdown: faster-whisper (CTranslate2 GPU,
~70× RT, 1.5 GB VRAM), whisper.cpp (CPU/CUDA, ~8–15× RT, ~1 GB RAM),
Groq Whisper (fastest cloud, free 2 000 req/day, key shared with Groq LLM).
Backend help panels now show ⚡ speed · ⏰ latency · ⭐ quality · 💾 RAM
metric chips for all TTS and STT backends.
Active Docker Stack cards also get per-container metric chips.
AI Backends section: "Use as STT" / "Use as TTS" one-click buttons on
faster-whisper, whisper.cpp, and Kokoro cards apply URLs to Settings
without leaving the page. Groq Whisper card notes the shared key path.
Settings: Kokoro URL in TTS cluster; faster-whisper URL, whisper.cpp URL,
Groq API key in STT cluster; quick-fill buttons for all local STT engines.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add initLlmsSection() IIFE to app.js: copy buttons, API key persistence with eye toggle
and saved badge, local service URL persistence, Connect/Disconnect toggle with server-side
probe via /api/probe-url (avoids CORS), card turns green on success / red on failure
- Substitute 0.0.0.0 → host.docker.internal before probing (0.0.0.0 not routable from Docker)
- Add /api/local-containers, /api/probe-url, start/stop/restart endpoints to server.py
- Rewrite AI Backends section into Local / Online API categories with Docker stack grid,
local service cards (LLM/STT/TTS) with icons and editable URL inputs, online cloud API cards
- Add bind mounts for static/ and server.py so changes take effect without image rebuild
- Add dc-grid, llm-local-grid CSS with uniform minmax(310px,1fr) card layout
- Fix VOICE_HOST_DIR default via .env so voice folders survive container recreation
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- /api/browse-dirs endpoint lists subdirectories at any container path
- '📁 Browse' button next to the path input toggles an inline dir browser
- Breadcrumb navigation lets users click up/down through the filesystem
- 'Use this folder' confirms the selection back into the path input
- Existing 'Set & reload' flow unchanged
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Backend proxy at /api/elevenlabs/voices forwarding to ElevenLabs shared-voices API
- elevenlabs_api_key added to settings (free key unlocks 12 000+ voices; 3 without key)
- Category pills: All, Featured, Professional, Narration, Conversational, News, Characters, Meditation, Gaming, Training
- Language filter with most common European languages + Arabic/Hindi/ZH/JA/KO
- Gender and age filters, debounced search, pagination (prev/next)
- Voice cards: colored avatar, name, language/gender/age/use-case tags, description, ▶ Play and ↓ Clone buttons
- Play button streams preview MP3 directly; Clone imports the audio into the Clone a Voice flow
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- docker-compose.yml: port 7890:7890, image/container/volume all renamed to tts-voice-creator-clone-and-design-2
- Dockerfile: EXPOSE 7890
- server.py: uvicorn binds to port 7890
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- New light UI: fixed 220px sidebar, single scrolling page, 8 named sections
- Static files split by concern: style.css, app.js, loader.js, nav.js
- Each page section is its own partial in static/sections/s-*.html
- loader.js fetches all section partials in parallel, then loads app.js and nav.js
- All original functionality, element IDs, and API endpoints preserved
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>