Fish-Speech emotion tags were silently ignored on non-English books: per-line
emotions are LLM-generated in the book's own language, but Fish-Speech only
recognizes English [tag] markers, and a double-tagging bug was stacking a
broken server-derived tag on top of the client's own. Added a DE->EN
translation table and removed the double-tagging. Also wires the existing
book-profile context and race_species field into character portrait prompts
(previously only used for voice design), adds a recast-until-threshold loop
for casting, and adds backend-aware emotion quick-picks to Read Aloud, Try a
Voice, and Conversation.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
For a designed voice the instruct prompt is the voice's identity — the TTS
engine reproduces the voice from that text alone. The only copy saved was the
`note` display summary, clipped to 240 characters, which left 43 of 73 voices
cut off mid-sentence. Save the complete prompt in its own field so the engine
can register a voice from the whole description.
Existing voices keep working from the clipped copy (it still carries gender,
accent and timbre) and pick up the full text when next redesigned.
Pairs with the engine-side fix in tts-dgx-spark-faster-qwen3-tts.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Voice consistency:
- Read back each voice's pinned seed (Seed Finder / Batch Seeds) on every
generation. The seed was saved to voice metadata but only ever read by the
Seed Finder's own benchmark path, so all per-voice seed pinning was inert.
- Stop coercing the "voice_design_playback" stability profile back to
"voice_clone". The pseudo-backend key isn't a real routing target, so the
backend-name normalizer silently rewrote it — reintroducing the hardcoded
seed:0 that profile exists to avoid, overriding every per-voice pin.
- Apply the accent clause on every line, not just at voice-creation time,
and reorder the instruct so emotion leads and accent trails (Qwen3-TTS
doesn't reliably follow multiple conflicting instructions).
- Pass an explicit language to Voice Design instead of leaving it on "Auto".
Audio effects:
- Add a limiter after compressor makeup gain. Makeup gain pushed peaks to
~1.9, and the final hard clip turned that into broadband distortion that
swamped the rest of the chain.
- Cascade highpass/lowpass 3 stages each (~18 dB/octave). Single-pole
filters were too gentle to band-limit speech audibly.
- Add a Bandpass control and wire it into the Telephone/Radio presets —
compression alone never sounded like a phone; band-limiting is the
defining trait.
Persona / Try It Out:
- Disable "Apply character persona" with an explanatory tooltip when the
voice has no persona saved, and error clearly server-side instead of
silently no-op'ing. Persona is typed manually per voice, never auto-filled.
- Stop dropping applyPersona in the chunked generation path (>200 chars).
- Populate the Voice Design dropdown from the user's own library rather than
filtering the engine's discovery list, which never contains custom voices.
Navigation and library:
- Use pushState instead of replaceState so browser Back/Forward step through
in-app navigation instead of leaving the app entirely.
- Show real dialogue line counts in the character sidebar instead of the
capped reference-quote count (which showed a misleading uniform "12").
Also fixes a crash in /api/transcribe-bytes that referenced an undefined
source_id in its cleanup path.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Introduces the new Studio section (Source -> Characters -> Voices ->
Perform & Export) that reuses the existing Read Aloud/Library/Script
Rehearsal code via DOM reparenting instead of duplicating it, and rolls up
a long tail of bugs found while producing a real audiobook through it:
umlaut-eating name sanitizers, a voice picker that mispositioned itself and
capped results at 60, PDF pagination silently breaking on trimmed \f
markers, a race letting stale audio keep playing after a new line was
clicked, an alias-overlap bug that could silently redirect a voice/image
save onto the wrong character, voice design failing outright during brief
TTS backend restarts instead of retrying, sparse cast entries defaulting to
English/wrong gender, and a reassigned voice never reaching an already-open
Stage session or invalidating its cached audio. Also adds a persistent
per-line audio cache, audiobook export browsing/download, and an inline
voice-design prompt editor. Full details in CHANGELOG.md.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>