2.1 KiB
Local And Realtime Speech Models
Blitztext App Linux supports two local-first speech paths:
- in-process
faster-whisperfor batch transcription - external Riva/NIM realtime servers for live streaming transcription
The app does not bundle speech models. You choose the model in Settings or in ~/.config/blitztext/config.toml.
Local Batch Transcription
The default local engine uses faster-whisper.
Recommended first model:
[whisper]
model = "small"
device = "auto"
compute_type = "auto"
Useful model sizes:
tiny: fastest, lowest qualitybase: small and responsivesmall: good first default for dictationmedium: better quality, slowerlarge-v3: highest quality, much heavier
You can also use a local model path supported by faster-whisper.
Realtime Riva/NIM Transcription
For live words while speaking, use a riva_realtime STT engine. The tested Nemotron ASR Streaming NIM exposes a WebSocket endpoint through /v1/realtime and reports this model:
cache-aware-parakeet-rnnt-en-US-asr-streaming-sortformer
Recommended engine config:
[[stt_engine]]
name = "Nemotron ASR Streaming"
type = "riva_realtime"
url = "http://127.0.0.1:8006/v1"
model = ""
Use model = "" to keep the server default. For the tested Nemotron container, set the general language to English:
[general]
language = "en"
Batch NIMs And Other STT Servers
Use type = "openai" only for servers that implement batch /v1/audio/transcriptions correctly.
Example:
[[stt_engine]]
name = "Parakeet batch ASR"
type = "openai"
url = "http://127.0.0.1:8090/v1"
model = "parakeet-tdt-0.6b-v3"
Streaming-only NIMs may still show /v1/audio/transcriptions in Swagger, but return bad model or No Offline ASR models found. Use riva_realtime for those.
Notes
- First local Whisper use can be slower because the model has to load or download.
- Realtime streaming needs
sounddeviceandwebsockets; both are inlinux/requirements.txt. - The benchmark tab is for batch engines. Streaming engines are live-only and are not benchmarked with WAV uploads.