Replace serial LLM-wait-TTS with overlapped execution: - LLM streams via background thread → asyncio.Queue (non-blocking event loop) - _sentence_split() detects sentence boundaries in the token stream - asyncio.create_task fires TTS for each sentence immediately — TTS for sentence 1 runs while LLM is still generating sentences 2, 3, … - Audio chunks stream to frontend in order as each task completes - Time-to-first-audio drops from (LLM total + TTS total) to roughly (LLM time-to-first-sentence + TTS latency for one sentence) Frontend audio queue: - enqueueAudio() / playNextAudio() chain multi-chunk responses seamlessly - clearAudio() stops playback and cancels queue on new turn or mic click - scheduleAutoMic() waits for queue to drain before restarting mic - Error paths clear the queue to avoid stale audio playing after failure Also fix missing contextlib import (silent bug when audio temp files needed cleanup in the STT path). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| admin.py | ||
| conversation.py | ||
| docker.py | ||
| library.py | ||
| settings.py | ||
| sources.py | ||
| stt.py | ||
| tts.py | ||