_switch_row description labels had set_line_wrap(True) but no
set_max_width_chars, so GTK computed their natural width as the full
un-wrapped text (~700px for 87-char descriptions). With NEVER horizontal
policy on the page ScrolledWindow this propagated to the dialog, making
Keyboard/Wakeword/Engines/Benchmark pages 1000–1360px wide.
Fix: add set_max_width_chars(50) to description labels, and change the
page SW horizontal policy from NEVER to AUTOMATIC as a safety net.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Stack.set_homogeneous(True) was requesting the max natural size of all
children (incl. wide Benchmark TreeViews) and forcing the dialog to 1000+px
wide for every page. Reverted to False so the window stays at 860×700 and
pages scroll if taller than the viewport.
Also removed the NEVER/NEVER ScrolledWindow policy on both benchmark pages —
it was propagating natural TreeView width up to the Stack, amplifying the
homogeneous bug.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Stack.set_homogeneous(True) keeps dialog height constant when switching pages
- Benchmark WW paned position raised 260→390 so all TTS config, engine
checkboxes, wakeword/samples/run fields are visible without scrolling
- Both benchmark pages disable their page-level ScrolledWindow so the
Paned fills the viewport rather than growing to natural height
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Paged GTK dialog guides through trigger method, keyboard shortcuts,
wakeword server, STT engine, and optional LLM setup. Shows automatically
on first launch, re-openable via Settings → "Setup Wizard…". Sets
setup_complete in config so it doesn't reappear.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
WakewordActionListener opens a second Wyoming connection during active
wakeword recording, listening for the configured cancel/send models. When
either fires it immediately calls cancel_dictation() or finish_dictation()
without any silence timer or Whisper pass. Settings UI adds Cancel model
and Send model pickers to the wakeword config card, populated from the
same server model list as the trigger model.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
CancelWatcher accumulates raw PCM from the VAD LevelMeter (via new
on_chunk callback) and runs a fast beam_size=1 transcription check every
~0.6s. When a cancel keyword is detected it immediately calls
cancel_dictation() without waiting for the silence timer to expire or
a full transcription to complete.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Hovering over a combo and scrolling could silently change the selection.
All ComboBoxText widgets now return True from their scroll-event handler,
swallowing the event before GTK's default handler can act on it.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Overlay HUD shows a × button in the top-right corner during recording/
streaming; clicking it cancels dictation. The button area is the only
clickable region — rest stays fully click-through.
- Wakeword tab "Cancel words" and "Send words" rows now include an inline
keyboard shortcut entry + Set button, so key_cancel / key_send can be
configured right next to the spoken-word equivalents.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Cancel hotkey now fires even when wakeword triggered the recording
(ModifierScheme state was "idle" so the key was silently ignored)
- Tray menu: "✕ Cancel recording" — always visible, enabled while recording,
grayed out at idle; works for both wakeword and manual recordings
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Save now diffs old vs new config. Safe changes (language, sounds, LLM,
keywords, overlay) show "✓ Applied" in the header for 4s. Restart-required
changes (STT engine, hotkeys, mic, wakeword) show "⚠ restart needed for: …"
and highlight Save & Restart. No modal popup on Save.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
_start_meter() was only called from _build_general() and referenced
self.mic_level unconditionally. Now _build_input() also starts the meter
when it's not already running, and both level bars are updated via hasattr.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The controls pane now has shrink=True and a ScrolledWindow wrapper, so
dragging the divider upward collapses the controls and expands the table.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Controls (TTS config, engine checkboxes, run button) are in the top pane;
the results table is in the bottom pane. Drag the divider to see more rows.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Section header icons were pushed down by bt-section CSS margin-top;
now only applied to the row container, not the image widget
- Wakeword benchmark shows a full TreeView table: per-voice Detected/Total/
Recall%/False-fires/Time with colour coding, plus aggregate row per engine
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Engine checkboxes let you pick which wakeword servers to include in the run
- Wakeword combo selects which model/phrase to test; leave empty for each
engine's own, pick a specific one (e.g. okay_computer) to override all
- Also fixes: TTS ⟳ no longer fills model combo with Kokoro voice names
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Sound fields accept WAV/MP3/OGG/FLAC/M4A/AAC/AIFF/Opus. Browse dialog
auto-plays each file on selection so you can preview before confirming.
sound.py falls back to ffplay/gst-play-1.0 for formats not supported
by pw-play/paplay.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Add/Quickstart/Reload/Delete buttons + Name field mirror the STT engines UI.
Four quickstart templates for common wyoming-openwakeword setups.
Migration: existing wakeword_uri/model auto-promoted to first preset.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Device/Compute rows already only show for local engines; the titled section
break was redundant and visually separated fields that belong together.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
HeaderBar replaces bottom button row — Save and Save & Restart appear in the
title bar on the right, X button closes. Section header icons vertically
centered using SMALL_TOOLBAR size and valign=CENTER.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
GTK symbolic icons on every tab label and every section/card header.
Also fixes the resize grip landing in the tab bar instead of the bottom-right corner.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Draws a classic dotted SE-corner grip overlaid on the bottom-right of the
notebook so users know the settings dialog is resizable.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Probe remote engines' /metrics for process_resident_memory_bytes or
container_memory_rss; show actual server-side MB in RAM column.
Falls back to "server" when the endpoint is not exposed.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Removes whitespace gap between STT config card and device/precision card
by hiding the latter when a Server or Realtime engine type is selected.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- benchmark.py: measure RSS delta via /proc/self/status before/after each
transcription; add ram_mb field to BenchRow
- gtksettings.py: add RAM (MB) column to results table (index 8); tooltip
column shifted to index 10
- CHANGELOG.md: full history from v2.02.00 through v2.03.01
- README.md: benchmark description updated to mention RAM column
- MANUAL.md: benchmark result columns as table; Log tab documents
level filter dropdown
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Emoji picker: SearchEntry at top filters all categories via unicodedata.name()
in real time; category view hides while searching, restores on clear
- Manual tab: add pkg_dir/MANUAL.md to _app_paths() search list so it works in
both venv and deb installs (MANUAL.md deployed alongside the package)
- bt-infobox CSS: replace @theme_selected_bg_color (saturated blue) with a
neutral 5% mix of fg/bg so banner text is always readable
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Every _labeled and _switch_row field now gets a clickable ⓘ info button
that opens a plain-language help popover — targeted at non-technical users
- New Manual tab in Settings renders MANUAL.md directly inside the dialog
- STT and LLM toolbars each get a "Quickstart ▾" button: a menu of common
providers (OpenAI, Groq, OpenRouter, Ollama, LM Studio, vLLM, llama-swap,
faster-whisper-server, NVIDIA Riva) that pre-fills the engine form in one click
- Engine type combos now show human-readable labels ("Internal — faster-whisper",
"LAN server — runs on your machine", "GPU (CUDA)", "int8 — fast, less memory")
while storing the same internal key values (no config migration needed)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds a 😀 button next to the Icon (emoji) entry in Settings → Presets.
Clicking it opens a GTK popover with 60 common emojis in a scrollable
flow grid; selecting one writes it into the field and closes the picker.
The entry still accepts direct keyboard input as before.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- gtksettings: build each notebook tab on first view instead of all up front, so
the dialog opens instantly (was ~1.3s building Input/Benchmark file-choosers);
_collect() force-builds unvisited tabs before saving so no field is missed.
- gtksettings: connection dot now sits left of the URL entry (like the Engines
tab) instead of at the far right.
- gtkui: open_settings raises the existing dialog instead of opening a second.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Green/red/grey reachability dot next to the wakeword and TTS endpoint fields,
matching the existing STT/LLM engine dots. Lightweight background TCP probe,
refreshed on open, on reload (⟳), and on focus-out — never blocks dialog build.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
On headless/minimal desktops the gvfs org.gtk.vfs.UDisks2VolumeMonitor dbus
service often fails to activate; each Gtk.FileChooserButton then blocked ~25s on
a StartServiceByName timeout while realizing, so the Settings dialog never
appeared and the stalled main loop froze the control panel too. Set
GIO_USE_VOLUME_MONITOR=unix in run_gui() before any window is realized.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Send by voice: a configured edge-anchored phrase (e.g. "computer send") is
stripped and the rest is delivered AND submitted with Enter. Off by default;
[routing] send_keywords + Settings → Input.
- Wakeword benchmark (Settings → Benchmark): synthesize the wake phrase in
random voices via any OpenAI-compatible TTS server, stream to
wyoming-openwakeword, report recall / false-fires / per-voice breakdown.
New [tts] config block.
- Tests: test_voice_send.py, test_wakeword_bench.py.
- Ignore agent workspace folder jules/.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The live waveform and silence auto-stop countdown were driven by a level
meter that was the last user of sounddevice/PortAudio, which hangs opening
the default input on PipeWire systems — so both stayed blank on the hotkey
and wakeword paths alike. Rewrite LevelMeter to stream raw PCM from the same
recorder as the WAV path (pw-record/parecord/arecord) and RMS it; identical
API, scaling, and ~10 Hz cadence. Also fixes the Settings mic-level preview.
Set GLib prgname/application name to "Blitztext" before any window is
realized (and add StartupWMClass to the .desktop) so the taskbar and GNOME's
"… is not responding" dialog show the app name instead of "__main__.py",
without touching the `python -m blitztext` entry point.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add a configurable voice cancel: saying "abbrechen" (or "cancel") at the start
or end of a clip discards the whole dictation — it is never routed onward,
rewritten, or typed. The rescue for accidentally triggered (e.g. wakeword)
recordings. Matched the same edge-anchored, ASR-tolerant way as routing keywords
via routing.is_cancel(), so the word buried mid-sentence won't trip it; checked
in Daemon._process right after transcription, before routing/rewrite/delivery.
Configurable via [routing] cancel_keywords (default ["abbrechen", "cancel"];
empty disables) and Settings -> Mic/Cues -> "Cancel words". The overlay briefly
shows "Abgebrochen". Docs (both READMEs) and CHANGELOG updated; tests cover the
matcher, the config round-trip, and the discard/deliver pipeline branches.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Fix a desktop-session freeze (forced logout/reboot) caused by the overlay's
AT-SPI caret tracker: it subscribed to the high-frequency object:text-caret-moved
signal and made synchronous, blocking AT-SPI reads from inside the event handler,
re-entering the a11y dispatcher and getting stormed by the app's own xdotool
typing until GNOME stopped responding. Now track focus changes only and read the
caret rectangle lazily, once, when the overlay shows — never on the hot path.
Fuse voice-routing feedback into the overlay instead of a desktop notification:
show the matched preset's emoji, name, and spoken keyword on a banner, narrate
the phase (Transcribing -> Rewriting), and stream the LLM rewrite into the bubble
token-by-token. Redundant per-dictation notifications are suppressed when the
overlay is present (errors still notify); headless/overlay-off is unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A full circle wraps the mic glyph and drains clockwise as the trailing-
silence timer runs out, recolouring cyan→amber→red and emptying exactly
as auto-stop fires. The daemon emits the countdown from the same VAD loop
that decides auto-stop, so the ring stays in sync; the overlay drains it
against its own clock for smooth motion and fades it in/out so word gaps
don't flicker it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Show a translucent, click-through bubble the moment recording starts (by
hotkey or wakeword): a pulsing microphone, a live waveform of the mic
level, and the recognised text (word-by-word when streaming, otherwise a
brief final-result confirmation). The tail points at the text caret via
AT-SPI accessibility, falling back to the mouse pointer, then a screen
corner. Also gives hands-free wakeword sessions visible feedback, whose
notifications are suppressed by design. X11 only.
- overlay.py: GTK override-redirect HUD (mic + waveform + bubble), drawn
with Cairo; thread-safe, marshalled onto the GTK loop.
- caret.py: best-effort anchor (AT-SPI caret -> pointer -> window/corner).
- daemon: optional level_cb/text_cb hooks; reuses the VAD level meter for
non-streaming, a dedicated meter for streaming. Stays UI-agnostic.
- gtkui: instantiate the overlay, drive show/update/hide from status.
- General settings: "Visual overlay" toggle; config overlay_enabled /
overlay_anchor (default on, "caret").
- Docs: CHANGELOG 1.5.0, version bump, README + MANUAL.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Spoken presets: route() now also matches a preset's name as an implicit
keyword, so presets with no configured keywords (Nicer email, Calm down,
Add emojis) are triggerable by voice — say the name. Explicit keywords still
win; names are added to STT hotwords too. Tests added.
- General tab: switches moved to the far right of each row with an inline
grey description (new _switch_row helper) so each toggle is self-explanatory.
- About tab: add "Copyright: 2026 mARTin Bierschenk - Design".
- MANUAL + CHANGELOG updated.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Bump version to 1.4.0 and cut the [1.4.0] - 2026-06-07 changelog section:
- "Announce matched preset" notification (shown even hands-free) + per-preset
emoji icons.
- Voice-routing no-keyword default now prefers a transcribe preset instead of
the first preset.
Refresh the README status line.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The matched-preset feedback used _dnotify, which is suppressed for hands-free
sessions — so wakeword users never saw which keyword/preset fired.
- Add a dedicated "Announce matched preset" notification (_rnotify), gated by
a new [general] notify_routing flag (default on) and independent of the
hands-free silence, so it shows for wakeword commands too. It only fires on a
real routing match, so it never spams when nothing is said.
- Show each preset's emoji in that notification; add a per-preset "Icon (emoji)"
field to the Presets editor so matches are visually distinct.
- General-tab toggle; MANUAL + CHANGELOG updated. Adds a test that the match is
announced even when _session_silent is set.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
When no [routing] default is configured, default_preset fell back to
workflows[0]. If that was an LLM rewrite (e.g. "Improve text"), every
wakeword/voice command without a matching keyword was sent to the language
model — repeatedly failing with HTTP 502 when the LLM backend was down.
default_preset now prefers a transcribe-mode preset for the no-keyword
fallback, so the default action is plain transcription. Adds tests.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Bump version to 1.3.0 and cut the [1.3.0] - 2026-06-07 changelog section
(wakeword pause toggle, independent hands-free audio cues, configurable
auto-stop silence, notification hygiene, PortAudio noise suppression, and the
new MANUAL.md). Refresh the README status line.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Reported issues from hands-free use:
- Wakeword WAVs didn't play: the new "Play audio cues" master switch
([sounds] enabled) also gated the hands-free Sound: detected/captured cues,
so enabled=false silenced them. The cues live in a separate UI section, so
this was surprising. Wakeword cues are now independent of that switch: they
play whenever a file is set, and an empty field means silent (no system-chime
fallback) — which is also how you turn a hands-free cue off. The master switch
now governs only the manual (keyboard) before/after chimes.
- Clarified the four sound fields' tooltips/labels (detected/captured = hands-free
only; before/after = manual only) and the empty-field behaviour.
- New "Silence to stop (s)" setting (Settings → Input → Hands-free, or
[wakeword] silence_seconds, default 2.0): user-defined trailing-silence
timeout for hands-free auto-stop (was hard-coded to 2.5 s).
- Add MANUAL.md documenting every setting in every tab; link it from the README.
Tests: cue independence + manual-gating + roundtrip (17 passed).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two follow-ups from hands-free testing:
- Audio cues had no off-switch: empty sound fields fall back to the freedesktop
system chime, so wakeword/recording always made noise. Add a [sounds] enabled
master flag (default true) exposed as "Play audio cues" in Settings → Input.
When off, _play_cue/_play_sound are no-ops — fully silent operation.
- The VAD level meter (sounddevice/PortAudio) leaked harmless thread-teardown
errors ("pthread_join ... failed", "PaUnixThread_Terminate ... failed") to the
terminal on every clip end. PortAudio writes these straight to fd 2, so wrap
the stream open/close in a fd-level stderr suppressor (_quiet_c_stderr).
Tests: cue gating respects the master switch (16 passed).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>