Replace serial LLM-wait-TTS with overlapped execution: - LLM streams via background thread → asyncio.Queue (non-blocking event loop) - _sentence_split() detects sentence boundaries in the token stream - asyncio.create_task fires TTS for each sentence immediately — TTS for sentence 1 runs while LLM is still generating sentences 2, 3, … - Audio chunks stream to frontend in order as each task completes - Time-to-first-audio drops from (LLM total + TTS total) to roughly (LLM time-to-first-sentence + TTS latency for one sentence) Frontend audio queue: - enqueueAudio() / playNextAudio() chain multi-chunk responses seamlessly - clearAudio() stops playback and cancels queue on new turn or mic click - scheduleAutoMic() waits for queue to drain before restarting mic - Error paths clear the queue to avoid stale audio playing after failure Also fix missing contextlib import (silent bug when audio temp files needed cleanup in the STT path). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| js | ||
| sections | ||
| vendor | ||
| app.js | ||
| index.html | ||
| loader.js | ||
| nav.js | ||
| style.css | ||