Adds validate_sound_path() that checks file type, allowed directories, and known audio extensions. Updates play() to use validation before passing paths to audio players.
- overlay: show × button in busy state (transcribing/rewriting), not only
while recording — updates hit-region, draw call, and label layout
- daemon: cancel_dictation() now handles _busy via threading.Event
(_abort_event); _process() clears the event at start, checks after STT
returns and after LLM rewrite completes, skipping delivery if set
- llm: chat() and _read_stream() accept abort_event; streaming loop breaks
immediately when the event is set so cancellation is near-instant
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Gtk.Entry reports its placeholder text width as natural width. Without
set_width_chars(1), entries can't shrink below that, so pages with long
placeholders (URL fields, keyword rows) ended up wider than the dialog.
Added set_width_chars(1) to _entry(), _url_field(), ModelPicker,
_kw_shortcut_row, and _sound_field. Also added set_max_width_chars(50)
to the stt_result wrapping label for the same reason.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Guard _refresh_status() with hasattr checks so opening STT Engines
before LLM Engines page is built no longer crashes the builder silently
- ww_status label: add max_width_chars(30) + ellipsize END so long model
lists don't widen the Wakeword page
- Benchmark STT sel_sw: NEVER→AUTOMATIC horizontal policy so wide engine
names scroll internally instead of propagating to the dialog
- _combo()/_type_combo(): set_size_request(10,-1) so ComboBoxText widgets
(e.g. long microphone names) can shrink below natural width
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
_switch_row description labels had set_line_wrap(True) but no
set_max_width_chars, so GTK computed their natural width as the full
un-wrapped text (~700px for 87-char descriptions). With NEVER horizontal
policy on the page ScrolledWindow this propagated to the dialog, making
Keyboard/Wakeword/Engines/Benchmark pages 1000–1360px wide.
Fix: add set_max_width_chars(50) to description labels, and change the
page SW horizontal policy from NEVER to AUTOMATIC as a safety net.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Stack.set_homogeneous(True) was requesting the max natural size of all
children (incl. wide Benchmark TreeViews) and forcing the dialog to 1000+px
wide for every page. Reverted to False so the window stays at 860×700 and
pages scroll if taller than the viewport.
Also removed the NEVER/NEVER ScrolledWindow policy on both benchmark pages —
it was propagating natural TreeView width up to the Stack, amplifying the
homogeneous bug.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Stack.set_homogeneous(True) keeps dialog height constant when switching pages
- Benchmark WW paned position raised 260→390 so all TTS config, engine
checkboxes, wakeword/samples/run fields are visible without scrolling
- Both benchmark pages disable their page-level ScrolledWindow so the
Paned fills the viewport rather than growing to natural height
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Paged GTK dialog guides through trigger method, keyboard shortcuts,
wakeword server, STT engine, and optional LLM setup. Shows automatically
on first launch, re-openable via Settings → "Setup Wizard…". Sets
setup_complete in config so it doesn't reappear.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
WakewordActionListener opens a second Wyoming connection during active
wakeword recording, listening for the configured cancel/send models. When
either fires it immediately calls cancel_dictation() or finish_dictation()
without any silence timer or Whisper pass. Settings UI adds Cancel model
and Send model pickers to the wakeword config card, populated from the
same server model list as the trigger model.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
CancelWatcher accumulates raw PCM from the VAD LevelMeter (via new
on_chunk callback) and runs a fast beam_size=1 transcription check every
~0.6s. When a cancel keyword is detected it immediately calls
cancel_dictation() without waiting for the silence timer to expire or
a full transcription to complete.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Hovering over a combo and scrolling could silently change the selection.
All ComboBoxText widgets now return True from their scroll-event handler,
swallowing the event before GTK's default handler can act on it.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Overlay HUD shows a × button in the top-right corner during recording/
streaming; clicking it cancels dictation. The button area is the only
clickable region — rest stays fully click-through.
- Wakeword tab "Cancel words" and "Send words" rows now include an inline
keyboard shortcut entry + Set button, so key_cancel / key_send can be
configured right next to the spoken-word equivalents.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Cancel hotkey now fires even when wakeword triggered the recording
(ModifierScheme state was "idle" so the key was silently ignored)
- Tray menu: "✕ Cancel recording" — always visible, enabled while recording,
grayed out at idle; works for both wakeword and manual recordings
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Save now diffs old vs new config. Safe changes (language, sounds, LLM,
keywords, overlay) show "✓ Applied" in the header for 4s. Restart-required
changes (STT engine, hotkeys, mic, wakeword) show "⚠ restart needed for: …"
and highlight Save & Restart. No modal popup on Save.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
_start_meter() was only called from _build_general() and referenced
self.mic_level unconditionally. Now _build_input() also starts the meter
when it's not already running, and both level bars are updated via hasattr.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The controls pane now has shrink=True and a ScrolledWindow wrapper, so
dragging the divider upward collapses the controls and expands the table.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Controls (TTS config, engine checkboxes, run button) are in the top pane;
the results table is in the bottom pane. Drag the divider to see more rows.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Section header icons were pushed down by bt-section CSS margin-top;
now only applied to the row container, not the image widget
- Wakeword benchmark shows a full TreeView table: per-voice Detected/Total/
Recall%/False-fires/Time with colour coding, plus aggregate row per engine
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Engine checkboxes let you pick which wakeword servers to include in the run
- Wakeword combo selects which model/phrase to test; leave empty for each
engine's own, pick a specific one (e.g. okay_computer) to override all
- Also fixes: TTS ⟳ no longer fills model combo with Kokoro voice names
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Sound fields accept WAV/MP3/OGG/FLAC/M4A/AAC/AIFF/Opus. Browse dialog
auto-plays each file on selection so you can preview before confirming.
sound.py falls back to ffplay/gst-play-1.0 for formats not supported
by pw-play/paplay.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Add/Quickstart/Reload/Delete buttons + Name field mirror the STT engines UI.
Four quickstart templates for common wyoming-openwakeword setups.
Migration: existing wakeword_uri/model auto-promoted to first preset.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Device/Compute rows already only show for local engines; the titled section
break was redundant and visually separated fields that belong together.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
HeaderBar replaces bottom button row — Save and Save & Restart appear in the
title bar on the right, X button closes. Section header icons vertically
centered using SMALL_TOOLBAR size and valign=CENTER.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
GTK symbolic icons on every tab label and every section/card header.
Also fixes the resize grip landing in the tab bar instead of the bottom-right corner.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Draws a classic dotted SE-corner grip overlaid on the bottom-right of the
notebook so users know the settings dialog is resizable.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Probe remote engines' /metrics for process_resident_memory_bytes or
container_memory_rss; show actual server-side MB in RAM column.
Falls back to "server" when the endpoint is not exposed.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Removes whitespace gap between STT config card and device/precision card
by hiding the latter when a Server or Realtime engine type is selected.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
String sort put "14.85" before "2.47". Now uses a custom comparator
that strips % and non-numeric markers before comparing as float.
Non-numeric values ("—", "server") sort to the bottom.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Remote engines run models server-side so local RSS never changes.
Show "server" instead of "—" to make the reason clear.
Local engines show measured MB; local already-loaded shows "—".
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
300s was too long; nobody waits 5 min for a result.
30s gives slow remote servers a fair window while still failing fast.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add _api_base() to strip endpoint-specific path suffixes before probing
/models, /metadata, /info — fixes /transcribe/models 404 spam when URL
ends with a custom path like /v1/transcribe
- Increase default transcribe() timeout 60s → 300s — WhisperX with speaker
diarization (pyannote) takes 2-5 min and was always timing out
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add "WhisperX server" quickstart template pre-filled with /transcribe endpoint
- Add tooltip to STT URL field explaining that non-standard servers (WhisperX)
should use the full endpoint as the URL, not /v1
- _url_field_lb accepts optional tooltip= kwarg
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- benchmark.py: measure RSS delta via /proc/self/status before/after each
transcription; add ram_mb field to BenchRow
- gtksettings.py: add RAM (MB) column to results table (index 8); tooltip
column shifted to index 10
- CHANGELOG.md: full history from v2.02.00 through v2.03.01
- README.md: benchmark description updated to mention RAM column
- MANUAL.md: benchmark result columns as table; Log tab documents
level filter dropdown
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Stability: remove blocking TCP socket call from _collect() — was
freezing the GTK main thread on Save when wakeword server unreachable;
now checks async in background and logs result
- Thread safety: fix _ww_load() reading GTK widget from background thread;
capture URI on main thread before spawning
- WhisperX 404: _transcribe_remote now detects non-standard paths
(anything other than /v1) and uses the URL as the full endpoint,
so http://host/transcribe works without /audio/transcriptions appended
- Log levels: logbuffer stores (ts, level, msg) tuples; log() accepts
level= (DEBUG/INFO/WARNING/ERROR); Log tab gets a Level dropdown
(Verbose/Info/Warning/Error) that filters displayed entries live
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds a "Server preset" combo in the Hands-free wakeword card that lists
all configured wakeword engines by name. Selecting one auto-fills the
URI and model fields and re-probes the connection. Selection persists
via wakeword_active in config.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- About tab: render License with _md_panel instead of _text_panel
- Benchmark pane: set_size_request 320px min, shrink=False on both sides
so the engine list and results table are always visible
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
After running a benchmark, each engine's best result (time, accuracy)
is persisted to config and shown as a small info line in the Engines tab
when that engine is selected. Updates live as the benchmark runs.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- stt.py: add ModelMeta dataclass, fmt_languages(), list_models_meta(),
detect_remote_device(); refactor list_models() to delegate
- benchmark.py: add languages field to BenchRow; fetch via _get_langs()
with URL-level caching using list_models_meta()
- gtksettings.py: show language labels per engine in checkbox list;
add language codes to search filter; add Lang column to results table
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Tab font-size 12px → 14px (matches body text)
- Inactive tabs: muted foreground color so they're clearly readable
but visually distinct from the active tab
- Active tab: bold + blue (#1a73e8), slightly more padding (8px 18px)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Schema: MAJOR.FEATURE.FIX
MAJOR — breaking changes / major redesign
FEATURE — two-digit, new user-visible features (00-99)
FIX — two-digit, bug fixes within a feature release (00-99)
Renamed 2.1.2 → 2.01.02 to start the new format.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>