Bump version to 1.4.0 and cut the [1.4.0] - 2026-06-07 changelog section: - "Announce matched preset" notification (shown even hands-free) + per-preset emoji icons. - Voice-routing no-keyword default now prefers a transcribe preset instead of the first preset. Refresh the README status line. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
11 KiB
11 KiB
Changelog
All notable changes to Blitztext for Linux (the native dictation tool in
linux/) are documented here. The format follows
Keep a Changelog, and this project
adheres to Semantic Versioning.
The version is defined in blitztext/__init__.py.
Unreleased
[1.4.0] - 2026-06-07
Added
- "Announce matched preset" notification (Settings → General, or
[general] notify_routing, default on): after a voice command, a notification shows which preset and spoken keyword matched — shown even for hands-free wakeword sessions, so you can see what you triggered. It only fires on a real match, so it never spams when nothing is said. - Per-preset emoji icon (Presets → "Icon (emoji)"): give each preset a distinct emoji, shown in the matched-preset notification so you can tell at a glance which fired.
Fixed
- Voice-routing default went to a rewrite: when no
[routing] defaultpreset is set, the no-keyword fallback used the first preset — which, if that happened to be an LLM rewrite (e.g. "Improve text"), sent every unrouted wakeword command to the language model (and failed when the LLM was down). The fallback now prefers atranscribepreset, so the default action is plain transcription.
[1.3.0] - 2026-06-07
Added
- Pause wakeword (tray): a reversible "Pause wakeword" toggle appears in the
system-tray menu when the wakeword is enabled. It pauses/resumes hands-free
detection by toggling the
/tmp/wake_mutedflag (external scripts may toggle the same file). - "Play audio cues" switch (Settings → Input → Audio cues, or
[sounds] enabled): on/off for the manual (keyboard/hotkey) start/stop chimes. Defaults to on. The hands-free wakeword sounds are independent of it. - Configurable wakeword auto-stop silence (Settings → Input → Hands-free →
"Silence to stop (s)", or
[wakeword] silence_seconds): end a hands-free recording this many seconds after you stop speaking. Defaults to2.0(previously hard-coded to 2.5 s).
Fixed
- Wakeword sounds silenced by the manual cue switch: the "Play audio cues" master switch wrongly muted the hands-free Sound: detected/captured cues too. Wakeword cues are now independent — they play whenever a file is set and stay silent when cleared (no surprise system-chime fallback), regardless of the manual switch.
- PortAudio/ALSA teardown noise: the level meter no longer leaks
pthread_join ... failed/PaUnixThread_Terminate ... failedlines to the terminal when a clip ends — that C-library chatter (written straight to fd 2) is now suppressed around the stream open/close. - Wakeword stuck muted: a leftover
/tmp/wake_mutedflag silently disabled detection with no in-app way to clear it. The state is now exposed and reversible from the tray, so a stale flag no longer kills hands-free use. The daemon also logs a clearStarting PAUSEDwarning when it boots muted. - Away-from-keyboard "Busy" storm: a wakeword hit arriving while the previous
clip was still transcribing went through
toggle()and popped a "Busy" notification. Wakeword triggers now go straight tostart_dictation(), so a busy/not-ready state is ignored silently instead. - Quiet hands-free errors: transcription/rewrite failures during a wakeword-triggered session no longer raise critical desktop notifications — they are logged instead, keeping background sessions silent.
- Notification storm / lock-screen pile-up: desktop notifications are now sent as transient with a short expiry and reuse a single bubble, so they no longer stack in the notification log or persist on the lock screen.
- Quiet hands-free sessions: per-dictation notifications are suppressed for wakeword-triggered sessions (audio cues are used instead).
[1.2.0] - 2026-06-05
Added
- Wyoming Wakeword Support: Complete hands-free integration via Wyoming protocol (e.g., openWakeWord), with live configuration testing and model fetching in the UI.
- ATK Screen Reader Accessibility: Fully mapped GTK labels, inputs, tooltips, and properties to the ATK bridge, enabling seamless navigation for blind users via screen readers like Orca.
- Drag-and-Drop Workflow Ordering: Workflows in the main tray menu can now be reordered via native drag-and-drop.
- Voice Activity Detection (VAD) Auto-Stop: Dictation now automatically stops after detecting 2.5 seconds of silence, removing the need to manually click Stop.
- Audio Feedback: Added audible start/stop/cancel chimes mapping to system-native alert sounds.
- Benchmark Autocomplete: The Benchmark UI automatically fills in the reference
.txttranscript if it matches the selected audio file. - Realtime STT streaming mode: new
mode = "stream"workflow support and ariva_realtimeSTT engine for Riva/NIM WebSocket transcription, including a Settings shortcut for Nemotron ASR Streaming onhttp://127.0.0.1:8006/v1. - Settings About tab with the app version, source link, changelog, and license text.
- STT & LLM engine manager: add, rename, and delete engine presets, each
with an online/offline status dot, a per-engine type (local/cloud), and a
searchable model dropdown populated from the server's
/models(with a reload button). Local Whisper device/precision now live with the STT engine. - Benchmark tab: run a reference clip through every STT engine and compare time, case-sensitive accuracy (WER), and a CPU/GPU/remote device column, with the fastest and most accurate highlighted.
- Custom audio cues: pick your own WAV files to play when recording starts and stops (covers stop+paste, stop+paste+Enter, and silence auto-stop), each with play-test and clear-to-default buttons. Built-in system sounds otherwise.
- Settings Log tab: a live activity log (model load/download, transcriptions, errors) with Copy and Clear, so a long "Loading…" is no longer opaque.
- Per-tab info boxes and expanded tooltips across Settings, written in plain language and exposed to screen readers (ATK) — for non-technical and blind users (barrierefrei).
- Click-to-bind hotkeys: a Set button captures the next keypress (including modifier-only chords like Ctrl+Win) into any hotkey field.
Changed
- GUI rebuilt in GTK 3 (replacing tkinter): a native GNOME panel unified with the tray, with the Ubuntu font and a dropdown+editor pattern in Settings.
- Voice-keyword routing and a modifier hotkey scheme (Ctrl+Win start, Ctrl stop+paste, Alt stop+paste+Enter, Esc cancel) replace per-preset combos as the default way to dictate.
- Quality gate rejects silent/too-short clips and stock Whisper hallucinations before they reach the screen.
1.1.0 - 2026-06-04
Added
- Debian package (
packaging/build-deb.sh) producing an installableblitztext_<ver>_arm64.debwith a desktop entry, app icon, and ablitztextlauncher. Installs via the Software app orapt install ./…deb. Bundles a relocatable venv with all Python deps (no pip/network at install) and declares system deps (python3-gi, xdotool, libnotify-bin, a recorder) so they pull in automatically. The bundled venv is built on the system/usr/bin/python3, so the tray works out of the box.
Notes
python3-giis already present on a standard Ubuntu GNOME install; the tray only seemed unavailable from source when the project venv was built from a non-system Python (e.g. conda/miniforge). The.debavoids this entirely.
1.0.1 - 2026-06-03
Changed
- Redesigned the control-panel window: minimal flat layout, Ubuntu font throughout, clickable workflow rows with hover (click to record / stop), subtle dividers, and text-style Settings/Quit actions. Dropped the monogram avatars and per-row buttons in favour of a cleaner, simpler look. The Settings window picks up the same font and styling.
1.0.0 - 2026-06-03
First release of the Linux port. The upstream project is a macOS-only menu-bar app (Swift/SwiftUI, CoreML/WhisperKit) that cannot run on Linux or in a container; this is a native host tool that reproduces the workflow — focus any text field, press a hotkey, speak, and the (optionally rewritten) text is typed into that field.
Added
- Native dictation engine (
daemon.py): global hotkeys via pynput, each hotkey toggles record → transcribe → optional rewrite → deliver. - Local transcription via faster-whisper (
transcribe.py);device="auto"tries CUDA and falls back to CPUint8(CPU-only on this arm64 host). - Microphone recording (
recorder.py) through pw-record / parecord / arecord — 16 kHz mono WAV, no Python audio bindings required. - Optional LLM rewrite (
rewrite.py) against any OpenAI-compatible endpoint (OpenAI, or a local vLLM / llama-swap), configurable per workflow. - Typing into the focused window via xdotool (
paste.py):typedirectly orpastethrough the clipboard; re-activates the target window first. - Configurable workflows (
config.py) in~/.config/blitztext/config.toml, with five defaults (Transcribe, Nicer email, Improve text, Calm down, Add emojis), per-workflow prompt/model/temperature overrides, and a TOML writer. - Control-panel window (
gui.py, tkinter): workflow rows with monogram avatars, hotkey badges, per-row Record buttons, a live status dot, and a Settings window that edits config and can Save & Restart. - System-tray mode (
tray.py, AppIndicator): macOS-menu-bar-style status icon with a menu to trigger each workflow, Show panel, Settings…, and Quit; shares one daemon/model/hotkey set with the window. Falls back to the window with an install hint when PyGObject is absent. - CLI (
__main__.py):tray(default),gui,run,transcribe,config-path, and--version. - Packaging:
install.sh(venv with--system-site-packages),requirements.txt, and ablitztext.servicesystemd user unit.
Notes
- Targets an X11 session (uses xdotool); Wayland would need ydotool/wtype.
- System tray requires a one-time
sudo apt install python3-gi(the GTK / AppIndicator typelibs and GNOMEubuntu-appindicatorsextension are already present on the target host).