After running a benchmark, each engine's best result (time, accuracy)
is persisted to config and shown as a small info line in the Engines tab
when that engine is selected. Updates live as the benchmark runs.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- stt.py: add ModelMeta dataclass, fmt_languages(), list_models_meta(),
detect_remote_device(); refactor list_models() to delegate
- benchmark.py: add languages field to BenchRow; fetch via _get_langs()
with URL-level caching using list_models_meta()
- gtksettings.py: show language labels per engine in checkbox list;
add language codes to search filter; add Lang column to results table
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Tab font-size 12px → 14px (matches body text)
- Inactive tabs: muted foreground color so they're clearly readable
but visually distinct from the active tab
- Active tab: bold + blue (#1a73e8), slightly more padding (8px 18px)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Schema: MAJOR.FEATURE.FIX
MAJOR — breaking changes / major redesign
FEATURE — two-digit, new user-visible features (00-99)
FIX — two-digit, bug fixes within a feature release (00-99)
Renamed 2.1.2 → 2.01.02 to start the new format.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
_stt_commit() accesses self.stt_name which only exists after the Engines
tab is lazily built. Guard with hasattr so running a benchmark directly
from the Benchmark tab no longer throws AttributeError.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Engine checklist and results table now split by a Gtk.Paned (vertical)
so the user can drag the divider to give more room to either panel
- WAV and reference .txt paths are written to disk (save()) the moment a
file is picked via the file chooser, without needing to click Save;
also saved on Run if they changed since the last disk write
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Benchmark tab: scrollable engine checklist above Run with live
reachability dots (green/red), name+model+URL filter, All/None buttons;
_run_bench respects selection and shows clear message when nothing ticked
- About tab: changelog now rendered via _md_panel (markdown headers, bold,
lists) instead of plain monospace _text_panel
- Manual tab: graceful fallback with clickable GitHub link when MANUAL.md
is not installed; build-deb.sh now copies MANUAL.md from repo root so
/opt/blitztext/MANUAL.md exists in future installs
- STT quickstart templates expanded: Speaches docker, whisper.cpp server,
NVIDIA NIM/Parakeet, five built-in local model sizes (tiny→large-v3)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Wrap ListStore in TreeModelSort and set sort_column_id on every column
so clicking any header sorts ascending/descending.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Engines tab: move Test button out of the toolbar, place it in a row
directly beside the result label so button and output are co-located;
result label is now selectable so error text can be copied
- Benchmark tab: add bench_wav/bench_ref/bench_expand_models to Config;
file pickers restore last-used paths on open; any change (file-set,
toggle, or Run) writes directly to cfg so paths survive without Save;
[benchmark] section written to config.toml on Save
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- stt.detect_remote_device(): probes /info (faster-whisper-server) then
/metadata (NVIDIA NIM) to detect CUDA vs CPU; cached per unique URL
- BenchRow gains url field; Device column now shows "CUDA" for GPU remotes
instead of the generic "remote"
- Benchmark table gains URL column (scheme stripped, max 180px wide)
- "Test all models per engine" checkbox: fetches list_models() for each
remote engine and expands to one row per model when checked
- benchmark.run() gains expand_models parameter
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- _run_bench: deduplicate STT engines by name before running; show a
warning in the summary line listing which names were skipped
- _bench_add_row: replace raw HTTP error strings with human-readable
reasons ("Wrong model name", "Server offline", "Timed out", etc.);
full raw error stored in hidden column 7 shown as row tooltip on hover
- bench_store: added 8th column (tooltip text); tree.set_tooltip_column(7)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- BenchRow gains `best_for` field: "Short clips" / "Short / medium" /
"Long / batch" / "Streaming" — derived from engine type and model name
- Device now shows "CUDA" instead of "GPU" for clarity
- Benchmark table gains a "Best for" column between Device and Time(s)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- stt._transcribe_remote: no longer falls back to "whisper-1" when
engine.model is empty — omits the field entirely so Riva/NIM uses its
default model instead of rejecting the request with HTTP 400
- gtksettings._page: switch vertical scroll policy to ALWAYS and disable
overlay scrolling so the scrollbar is permanently visible, making
it obvious when a tab has more content below the visible area
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Root cause of session freeze confirmed: a 15 000-char code block was typed
character-by-character via xdotool at 12ms/char = ~3 min, flooding the X11
per-client event buffer until the entire session froze.
- paste.py: any text >300 chars or containing newlines auto-upgrades to
clipboard paste (instant Ctrl+V) regardless of configured output mode.
xdotool type is kept only for short single-line text where it matters.
- daemon.py: replace on_token "".join(acc) accumulation (O(n²) for long
code blocks) with a sliding deque that shows only the last 400 chars in
the overlay — Pango no longer re-lays out a growing 15 KB string on each
incoming token.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Emoji picker: SearchEntry at top filters all categories via unicodedata.name()
in real time; category view hides while searching, restores on clear
- Manual tab: add pkg_dir/MANUAL.md to _app_paths() search list so it works in
both venv and deb installs (MANUAL.md deployed alongside the package)
- bt-infobox CSS: replace @theme_selected_bg_color (saturated blue) with a
neutral 5% mix of fg/bg so banner text is always readable
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Every _labeled and _switch_row field now gets a clickable ⓘ info button
that opens a plain-language help popover — targeted at non-technical users
- New Manual tab in Settings renders MANUAL.md directly inside the dialog
- STT and LLM toolbars each get a "Quickstart ▾" button: a menu of common
providers (OpenAI, Groq, OpenRouter, Ollama, LM Studio, vLLM, llama-swap,
faster-whisper-server, NVIDIA Riva) that pre-fills the engine form in one click
- Engine type combos now show human-readable labels ("Internal — faster-whisper",
"LAN server — runs on your machine", "GPU (CUDA)", "int8 — fast, less memory")
while storing the same internal key values (no config migration needed)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds a 😀 button next to the Icon (emoji) entry in Settings → Presets.
Clicking it opens a GTK popover with 60 common emojis in a scrollable
flow grid; selecting one writes it into the field and closes the picker.
The entry still accepts direct keyboard input as before.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- gtksettings: build each notebook tab on first view instead of all up front, so
the dialog opens instantly (was ~1.3s building Input/Benchmark file-choosers);
_collect() force-builds unvisited tabs before saving so no field is missed.
- gtksettings: connection dot now sits left of the URL entry (like the Engines
tab) instead of at the far right.
- gtkui: open_settings raises the existing dialog instead of opening a second.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Green/red/grey reachability dot next to the wakeword and TTS endpoint fields,
matching the existing STT/LLM engine dots. Lightweight background TCP probe,
refreshed on open, on reload (⟳), and on focus-out — never blocks dialog build.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
On headless/minimal desktops the gvfs org.gtk.vfs.UDisks2VolumeMonitor dbus
service often fails to activate; each Gtk.FileChooserButton then blocked ~25s on
a StartServiceByName timeout while realizing, so the Settings dialog never
appeared and the stalled main loop froze the control panel too. Set
GIO_USE_VOLUME_MONITOR=unix in run_gui() before any window is realized.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Send by voice: a configured edge-anchored phrase (e.g. "computer send") is
stripped and the rest is delivered AND submitted with Enter. Off by default;
[routing] send_keywords + Settings → Input.
- Wakeword benchmark (Settings → Benchmark): synthesize the wake phrase in
random voices via any OpenAI-compatible TTS server, stream to
wyoming-openwakeword, report recall / false-fires / per-voice breakdown.
New [tts] config block.
- Tests: test_voice_send.py, test_wakeword_bench.py.
- Ignore agent workspace folder jules/.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The live waveform and silence auto-stop countdown were driven by a level
meter that was the last user of sounddevice/PortAudio, which hangs opening
the default input on PipeWire systems — so both stayed blank on the hotkey
and wakeword paths alike. Rewrite LevelMeter to stream raw PCM from the same
recorder as the WAV path (pw-record/parecord/arecord) and RMS it; identical
API, scaling, and ~10 Hz cadence. Also fixes the Settings mic-level preview.
Set GLib prgname/application name to "Blitztext" before any window is
realized (and add StartupWMClass to the .desktop) so the taskbar and GNOME's
"… is not responding" dialog show the app name instead of "__main__.py",
without touching the `python -m blitztext` entry point.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add a configurable voice cancel: saying "abbrechen" (or "cancel") at the start
or end of a clip discards the whole dictation — it is never routed onward,
rewritten, or typed. The rescue for accidentally triggered (e.g. wakeword)
recordings. Matched the same edge-anchored, ASR-tolerant way as routing keywords
via routing.is_cancel(), so the word buried mid-sentence won't trip it; checked
in Daemon._process right after transcription, before routing/rewrite/delivery.
Configurable via [routing] cancel_keywords (default ["abbrechen", "cancel"];
empty disables) and Settings -> Mic/Cues -> "Cancel words". The overlay briefly
shows "Abgebrochen". Docs (both READMEs) and CHANGELOG updated; tests cover the
matcher, the config round-trip, and the discard/deliver pipeline branches.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Fix a desktop-session freeze (forced logout/reboot) caused by the overlay's
AT-SPI caret tracker: it subscribed to the high-frequency object:text-caret-moved
signal and made synchronous, blocking AT-SPI reads from inside the event handler,
re-entering the a11y dispatcher and getting stormed by the app's own xdotool
typing until GNOME stopped responding. Now track focus changes only and read the
caret rectangle lazily, once, when the overlay shows — never on the hot path.
Fuse voice-routing feedback into the overlay instead of a desktop notification:
show the matched preset's emoji, name, and spoken keyword on a banner, narrate
the phase (Transcribing -> Rewriting), and stream the LLM rewrite into the bubble
token-by-token. Redundant per-dictation notifications are suppressed when the
overlay is present (errors still notify); headless/overlay-off is unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A full circle wraps the mic glyph and drains clockwise as the trailing-
silence timer runs out, recolouring cyan→amber→red and emptying exactly
as auto-stop fires. The daemon emits the countdown from the same VAD loop
that decides auto-stop, so the ring stays in sync; the overlay drains it
against its own clock for smooth motion and fades it in/out so word gaps
don't flicker it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Show a translucent, click-through bubble the moment recording starts (by
hotkey or wakeword): a pulsing microphone, a live waveform of the mic
level, and the recognised text (word-by-word when streaming, otherwise a
brief final-result confirmation). The tail points at the text caret via
AT-SPI accessibility, falling back to the mouse pointer, then a screen
corner. Also gives hands-free wakeword sessions visible feedback, whose
notifications are suppressed by design. X11 only.
- overlay.py: GTK override-redirect HUD (mic + waveform + bubble), drawn
with Cairo; thread-safe, marshalled onto the GTK loop.
- caret.py: best-effort anchor (AT-SPI caret -> pointer -> window/corner).
- daemon: optional level_cb/text_cb hooks; reuses the VAD level meter for
non-streaming, a dedicated meter for streaming. Stays UI-agnostic.
- gtkui: instantiate the overlay, drive show/update/hide from status.
- General settings: "Visual overlay" toggle; config overlay_enabled /
overlay_anchor (default on, "caret").
- Docs: CHANGELOG 1.5.0, version bump, README + MANUAL.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Bump version to 1.4.0 and cut the [1.4.0] - 2026-06-07 changelog section:
- "Announce matched preset" notification (shown even hands-free) + per-preset
emoji icons.
- Voice-routing no-keyword default now prefers a transcribe preset instead of
the first preset.
Refresh the README status line.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Bump version to 1.3.0 and cut the [1.3.0] - 2026-06-07 changelog section
(wakeword pause toggle, independent hands-free audio cues, configurable
auto-stop silence, notification hygiene, PortAudio noise suppression, and the
new MANUAL.md). Refresh the README status line.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
packaging/build-deb.sh builds an installable blitztext_<ver>_arm64.deb with a
desktop entry, app icon, and launcher. Bundles a relocatable venv (all Python
deps, no pip at install) and declares system deps (python3-gi, xdotool,
libnotify-bin, recorder). Built on /usr/bin/python3 so the tray works out of
the box. Installs via the Software app or `apt install ./…deb`.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Minimal flat design with the Ubuntu font: clickable workflow rows with hover
(click to record/stop), muted descriptions and hotkey hints, subtle dividers,
and text-style Settings/Quit. Drops monogram avatars and per-row buttons.
Settings window restyled to match.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The upstream app is macOS-only (Swift/SwiftUI, CoreML/WhisperKit) and can't
run on Linux or in a container. This adds a native host tool under linux/ that
reproduces the workflow: focus any text field, press a hotkey, speak, and the
optionally-rewritten text is typed into that field.
- Engine: pynput global hotkeys → mic record → local faster-whisper →
optional OpenAI-compatible rewrite → xdotool typing into the focused window
- Frontends: system tray (AppIndicator, default), tkinter control panel,
and headless modes
- Config-driven workflows in ~/.config/blitztext/config.toml with per-workflow
prompt/model/temperature overrides
- Packaging: install.sh, requirements.txt, systemd user unit
- Targets X11; local transcription runs CPU int8 on this arm64 host
See linux/CHANGELOG.md and linux/README.md for details.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>