routing: send-by-voice keyword (types + presses Enter); wakeword benchmark via TTS; Release 1.8.0

- Send by voice: a configured edge-anchored phrase (e.g. "computer send") is
  stripped and the rest is delivered AND submitted with Enter. Off by default;
  [routing] send_keywords + Settings → Input.
- Wakeword benchmark (Settings → Benchmark): synthesize the wake phrase in
  random voices via any OpenAI-compatible TTS server, stream to
  wyoming-openwakeword, report recall / false-fires / per-voice breakdown.
  New [tts] config block.
- Tests: test_voice_send.py, test_wakeword_bench.py.
- Ignore agent workspace folder jules/.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
mARTin-B78 2026-06-08 13:17:33 +02:00
parent 1c64f19bce
commit 28020fdeda
11 changed files with 876 additions and 4 deletions

3
.gitignore vendored
View File

@ -30,3 +30,6 @@ Secrets.swift
# Tooling
node_modules/
# Agent workspace
jules/

View File

@ -9,6 +9,27 @@ The version is defined in [`blitztext/__init__.py`](blitztext/__init__.py).
## [Unreleased]
## [1.8.0] - 2026-06-08
### Added
- **Send by voice**: say a distinctive phrase like **"computer send"** at the
start or end of a clip and the word is stripped, then the rest is typed **and
submitted with Enter** — the spoken equivalent of "stop + paste + Enter".
Mainly for hands-free use, where you can't press a key. Configure under
Settings → Input → "Send words", or `[routing] send_keywords`. Off by default;
because it presses Enter, use a multi-word phrase (e.g. your wakeword + "send")
so a sentence that merely ends in "send" doesn't submit by accident. Matched
the same edge-anchored, ASR-tolerant way as routing/cancel keywords.
- **Wakeword benchmark** (Settings → Benchmark): stress-test hands-free
detection. It synthesizes short sentences with your wake phrase spoken in
random voices (plus pure-filler utterances with none), streams them to your
wyoming-openwakeword server, and reports **recall** (how reliably it fires),
**false fires**, and a **per-voice** breakdown. Speech comes from any
OpenAI-compatible TTS server (Kokoro-FastAPI, XTTS, OpenAI, …): set its URL,
optional API-key env var, model, and voices under the new `[tts]` config / the
Benchmark tab, and use **Connect** to test it (it auto-fills the voice list
when the server exposes one).
## [1.7.1] - 2026-06-08
### Fixed

View File

@ -45,6 +45,12 @@ It's the rescue for an accidentally triggered (e.g. wakeword) recording. Set the
words under `[routing] cancel_keywords` (default `["abbrechen", "cancel"]`; an
empty list turns it off).
**Send by voice:** say a distinctive phrase like *"computer send"* at the start
or end of a clip and the word is stripped, then the rest is typed **and submitted
with Enter** — the spoken equivalent of "stop + paste + Enter", ideal hands-free.
Off by default; set the phrases under `[routing] send_keywords` (use a multi-word
phrase so a sentence merely ending in "send" doesn't submit by accident).
While you dictate, an optional **on-screen overlay** (Settings → General →
"Visual overlay", default on) shows a translucent bubble at the cursor with a
pulsing microphone, a live waveform of your mic level, and the recognised text —
@ -85,7 +91,7 @@ screen-reader (ATK) support. Click any image to open it full size.
<p align="center">
<a href="../Screenshots/settings-benchmark.png"><img src="../Screenshots/settings-benchmark.png" alt="Benchmark settings tab" width="100%"></a><br>
<em><b>Benchmark</b> — compare every configured STT engine against a reference WAV + transcript to find the fastest and most accurate, with a Device column (CPU / GPU / remote).</em>
<em><b>Benchmark</b> — compare every configured STT engine against a reference WAV + transcript to find the fastest and most accurate, with a Device column (CPU / GPU / remote). Also includes a <b>wakeword benchmark</b>: point it at any OpenAI-compatible TTS server (Kokoro, XTTS, …), and it synthesizes your wake phrase in random voices, streams it to your wyoming-openwakeword server, and reports recall + false-fires per voice.</em>
</p>
<p align="center">

View File

@ -6,4 +6,4 @@ counterpart to the macOS Blitztext menu bar app: it runs natively on the host
(not in a container) so it can type into any application via xdotool.
"""
__version__ = "1.7.1"
__version__ = "1.8.0"

View File

@ -30,6 +30,14 @@ class Workflow:
icon: str = ""
@dataclass
class WakewordEngine:
"""A wyoming-openwakeword endpoint to compare in the wakeword benchmark."""
name: str = ""
uri: str = "tcp://127.0.0.1:10400"
model: str = "okay_computer"
@dataclass
class Config:
# general
@ -79,6 +87,12 @@ class Config:
# dictation is discarded — never transcribed onward, routed, rewritten, or
# typed. Empty list = disabled. Mainly for accidental wakeword triggers.
cancel_keywords: list[str] = field(default_factory=lambda: ["abbrechen", "cancel"])
# Spoken send: if one of these is heard at the start/end of a clip, it is
# stripped and the rest is delivered *and submitted with Enter* — the spoken
# equivalent of "stop + paste + Enter". Empty = disabled. Because it presses
# Enter, prefer a distinctive multi-word phrase (e.g. "computer send") so a
# sentence that merely ends in "send" doesn't submit by accident.
send_keywords: list[str] = field(default_factory=list)
# speech-to-text engines (presets)
stt_engines: list[STTEngine] = field(default_factory=list)
stt_active: str = ""
@ -92,6 +106,17 @@ class Config:
wakeword_sound_detected: str = "" # WAV played when the wakeword fires (speak now)
wakeword_sound_done: str = "" # WAV played when the command is captured
wakeword_silence_seconds: float = 2.0 # auto-stop after this much trailing silence
# Text-to-speech for the wakeword benchmark — its own OpenAI-compatible
# endpoint (Kokoro, XTTS, OpenAI, …): base URL incl. /v1, an optional bearer
# key env var, a model id, and the voices to cycle through.
tts_url: str = ""
tts_api_key_env: str = ""
tts_model: str = ""
tts_voices: list[str] = field(
default_factory=lambda: ["alloy", "echo", "fable", "onyx", "nova", "shimmer"])
# Wakeword engines to compare in the benchmark (each a wyoming-openwakeword
# endpoint). Empty → the benchmark falls back to the live [wakeword] above.
wakeword_engines: list[WakewordEngine] = field(default_factory=list)
# workflows
workflows: list[Workflow] = field(default_factory=list)
@ -158,6 +183,7 @@ def load(path: Path = CONFIG_PATH) -> Config:
q = data.get("quality", {})
snd = data.get("sounds", {})
ww = data.get("wakeword", {})
tts = data.get("tts", {})
cfg = Config(
recorder=g.get("recorder", "auto"),
@ -183,6 +209,7 @@ def load(path: Path = CONFIG_PATH) -> Config:
routing_default=rt.get("default", ""),
routing_threshold=float(rt.get("threshold", 0.82)),
cancel_keywords=list(rt.get("cancel_keywords", ["abbrechen", "cancel"])),
send_keywords=list(rt.get("send_keywords", [])),
input_mode=inp.get("mode", "modifiers"),
push_to_talk=bool(inp.get("push_to_talk", False)),
key_start=inp.get("start", "<ctrl>+<cmd>"),
@ -202,6 +229,10 @@ def load(path: Path = CONFIG_PATH) -> Config:
wakeword_sound_detected=ww.get("sound_detected", ""),
wakeword_sound_done=ww.get("sound_done", ""),
wakeword_silence_seconds=float(ww.get("silence_seconds", 2.0)),
tts_url=tts.get("url", "").rstrip("/"),
tts_api_key_env=tts.get("api_key_env", ""),
tts_model=tts.get("model", ""),
tts_voices=list(tts.get("voices", ["alloy", "echo", "fable", "onyx", "nova", "shimmer"])),
)
for entry in data.get("workflow", []):
@ -235,6 +266,17 @@ def load(path: Path = CONFIG_PATH) -> Config:
] or [STTEngine("Local faster-whisper", "local", model=cfg.model)]
cfg.stt_active = data.get("stt", {}).get("active", cfg.stt_engines[0].name)
# Wakeword engines for the benchmark (optional; benchmark falls back to the
# live [wakeword] config when none are listed).
cfg.wakeword_engines = [
WakewordEngine(
name=e.get("name", ""),
uri=e.get("uri", "tcp://127.0.0.1:10400"),
model=e.get("model", "okay_computer"),
)
for e in data.get("wakeword_engine", [])
]
# LLM engines (default: synthesized from the legacy [rewrite] block).
cfg.llm_engines = [
LLMEngine(
@ -298,6 +340,7 @@ def save(cfg: Config, path: Path = CONFIG_PATH) -> None:
"default": cfg.routing_default,
"threshold": cfg.routing_threshold,
"cancel_keywords": cfg.cancel_keywords,
"send_keywords": cfg.send_keywords,
},
"quality": {
"min_speech_seconds": cfg.min_speech_seconds,
@ -318,6 +361,15 @@ def save(cfg: Config, path: Path = CONFIG_PATH) -> None:
"sound_done": cfg.wakeword_sound_done,
"silence_seconds": cfg.wakeword_silence_seconds,
},
"tts": {
"url": cfg.tts_url,
"api_key_env": cfg.tts_api_key_env,
"model": cfg.tts_model,
"voices": cfg.tts_voices,
},
"wakeword_engine": [
{"name": e.name, "uri": e.uri, "model": e.model} for e in cfg.wakeword_engines
],
"stt": {"active": cfg.stt_active},
"stt_engine": [
{k: v for k, v in {
@ -439,6 +491,12 @@ threshold = 0.82 # 0..1 fuzzy-match strictness (higher = stricter)
# routed, rewritten, or typed. Handy when a wakeword fires by accident. Pick
# words you won't naturally end a real dictation with. Empty list = off.
cancel_keywords = ["abbrechen", "cancel"]
# Say one of these at the start or end of a clip to SEND it: the word is stripped
# and the rest is delivered AND submitted with Enter (spoken "stop+paste+Enter").
# Because it presses Enter, use a distinctive multi-word phrase (e.g. your
# wakeword + "send") so a sentence that just ends in "send" won't submit. Off by
# default; empty list = off.
send_keywords = []
[wakeword]
# Hands-free dictation using an external wyoming-openwakeword server.
@ -455,6 +513,30 @@ sound_done = ""
# Auto-stop the recording this many seconds after you stop speaking (silence).
silence_seconds = 2.0
[tts]
# Text-to-speech for the WAKEWORD BENCHMARK only (Settings → Benchmark). Point it
# at any OpenAI-compatible TTS server's /audio/speech (Kokoro-FastAPI, XTTS-v2,
# OpenAI, …). url = base incl. /v1; api_key_env = env var holding a bearer key
# (leave empty for no-auth local servers); model = the TTS model id; voices = the
# names your endpoint serves. Leave url/model empty to disable the benchmark.
url = ""
api_key_env = ""
model = ""
voices = ["alloy", "echo", "fable", "onyx", "nova", "shimmer"]
# Wakeword engines to COMPARE in the benchmark — add one block per
# wyoming-openwakeword endpoint/model you want side by side. Leave this out to
# just benchmark the live [wakeword] engine above.
# [[wakeword_engine]]
# name = "openWakeWord · computer"
# uri = "tcp://127.0.0.1:10400"
# model = "computer"
#
# [[wakeword_engine]]
# name = "microWakeWord · hey_jarvis"
# uri = "tcp://127.0.0.1:10500"
# model = "hey_jarvis"
# ----------------------------------------------------------------------------
# Speech-to-text engines (presets). The active one is used for transcription.
# type = "local" -> in-process faster-whisper (uses [whisper] above)

View File

@ -19,7 +19,7 @@ from .notify import notify
from .paste import active_window_id, deliver
from .streaming import RivaRealtimeStreamer
from .recorder import Recording, detect_recorder
from .routing import is_cancel, route
from .routing import is_cancel, match_send, route
from .transcribe import Transcriber
# status_cb(state, workflow_name, message)
@ -473,6 +473,14 @@ class Daemon:
self._dnotify("Abgebrochen", f"{cancel_kw}“ gehört — verworfen.", "low")
return
# Spoken send: a configured word at an edge ("computer send") is
# stripped, and the rest is delivered AND submitted with Enter — the
# spoken equivalent of stop+paste+Enter. Mainly for hands-free use.
send_kw, text = match_send(text, self.cfg.send_keywords, threshold=self.cfg.routing_threshold)
if send_kw:
send_enter = True
log(f"⏎ Send keyword “{send_kw}” — delivering and pressing Enter.")
# Voice routing: pick the preset from a spoken keyword, strip it.
if workflow.mode == "route":
res = route(text, self.cfg.workflows, threshold=self.cfg.routing_threshold)

View File

@ -24,7 +24,7 @@ gi.require_version("Gtk", "3.0")
gi.require_version("Gdk", "3.0")
from gi.repository import Gdk, GLib, Gtk, Pango # noqa: E402
from . import __version__, audio, autostart, benchmark, llm, logbuffer, stt # noqa: E402
from . import __version__, audio, autostart, benchmark, llm, logbuffer, stt, wakeword_bench # noqa: E402
from .config import Config, save # noqa: E402
from .llm import LLMEngine # noqa: E402
from .stt import STTEngine # noqa: E402
@ -225,6 +225,19 @@ class ModelPicker(Gtk.Box):
self.entry.set_text(value or "")
class MultiPicker(ModelPicker):
"""Like ModelPicker, but picking a row APPENDS to a comma-separated list —
for fields that hold several values (e.g. the benchmark's voices)."""
def _activated(self, _lb, row) -> None:
val = row.get_child().get_text()
cur = [v.strip() for v in self.entry.get_text().split(",") if v.strip()]
if val and val not in cur:
cur.append(val)
self.entry.set_text(", ".join(cur))
self.pop.popdown()
def _model_combo(placeholder="") -> ModelPicker:
return ModelPicker(placeholder)
@ -820,6 +833,14 @@ class SettingsDialog:
"nothing is transcribed onward, routed, rewritten, or typed. "
"Rescues an accidentally triggered dictation. Empty = off.")
self.send_keywords = _labeled(
page, "Send words (comma)",
_entry(", ".join(self.cfg.send_keywords), placeholder="computer send, computer abschicken"),
tooltip="Say one of these at the start or end of a clip to SEND it: the word is "
"stripped and the rest is typed AND submitted with Enter (spoken "
"stop+paste+Enter). Because it presses Enter, use a distinctive "
"multi-word phrase (e.g. your wakeword + send). Empty = off.")
self.ww_snd_detected = self._sound_field(
page, "Sound: detected", self.cfg.wakeword_sound_detected,
"HANDS-FREE ONLY. Plays the instant the wake word is recognised and recording starts "
@ -1058,6 +1079,56 @@ class SettingsDialog:
self.bench_summary = Gtk.Label(xalign=0.0); self.bench_summary.set_line_wrap(True)
page.pack_start(self.bench_summary, False, False, 4)
# --- Wakeword benchmark ---------------------------------------------
page.pack_start(Gtk.Separator(orientation=Gtk.Orientation.HORIZONTAL), False, False, 10)
_infobox(page, "Stress-test the wakeword. It synthesizes short sentences with your "
"wakeword spoken in random voices (plus filler with none), streams them to "
"your wyoming-openwakeword server, and reports how reliably it fires "
"(recall) and whether it false-fires. Speech comes from any OpenAI-compatible "
"TTS server (Kokoro, XTTS, OpenAI …) — set its URL, model and voices, then "
"Connect to test it.")
# URL with a ⟳ reload that loads models + voices from the server (like the
# Engines tab). Set the saved URL after building the field.
self.wwb_url = _url_field(page, "TTS URL", "http://localhost:8880/v1", self._tts_reload)
self.wwb_url.set_text(self.cfg.tts_url)
self.wwb_key = _labeled(
page, "API key env", _entry(self.cfg.tts_api_key_env, placeholder="(optional, e.g. OPENAI_API_KEY)"),
tooltip="Name of an environment variable holding a bearer token. Leave empty for "
"no-auth local servers. The key itself is never stored in the config.")
self.wwb_model = _model_combo("press ⟳ to load, or type a model id")
self.wwb_model.set_text(self.cfg.tts_model)
_labeled(page, "TTS model", self.wwb_model,
tooltip="The TTS model id your endpoint serves (from {url}/models). Required.")
self.wwb_voices = MultiPicker("press ⟳ to load voices, or type names")
self.wwb_voices.set_text(", ".join(self.cfg.tts_voices))
_labeled(page, "Voices (comma)", self.wwb_voices,
tooltip="Voices to cycle through at random so the test covers different timbres. "
"⟳ fills these from the server; the ▾ dropdown appends one at a time.")
self.wwb_test_btn = Gtk.Button(label="Connect")
self.wwb_test_btn.set_tooltip_text("Synthesize a test phrase with the selected model + voice.")
self.wwb_test_btn.connect("clicked", self._connect_tts)
self.wwb_test_lbl = Gtk.Label(xalign=0.0); self.wwb_test_lbl.set_line_wrap(True)
connect_box = Gtk.Box(spacing=10)
connect_box.pack_start(self.wwb_test_btn, False, False, 0)
connect_box.pack_start(self.wwb_test_lbl, False, False, 0)
_labeled(page, "", connect_box)
adj = Gtk.Adjustment(value=12, lower=1, upper=200, step_increment=1, page_increment=10)
self.wwb_count = Gtk.SpinButton(adjustment=adj, climb_rate=1, digits=0)
self.wwb_count.set_halign(Gtk.Align.START); self.wwb_count.set_size_request(90, -1)
_labeled(page, "Wakeword samples", self.wwb_count,
tooltip="How many wakeword utterances to synthesize and test (filler-only "
"utterances for false-fire checking are added on top).")
wrun = Gtk.Button(label="Run wakeword benchmark"); wrun.connect("clicked", self._run_wakeword_bench)
wrun.set_halign(Gtk.Align.START)
page.pack_start(wrun, False, False, 6)
self.wwb_summary = Gtk.Label(xalign=0.0); self.wwb_summary.set_line_wrap(True)
self.wwb_summary.set_selectable(True)
page.pack_start(self.wwb_summary, False, False, 4)
def _run_bench(self, _b) -> None:
wav = self.bench_wav.get_filename()
refp = self.bench_ref.get_filename()
@ -1094,6 +1165,123 @@ class SettingsDialog:
f" · <b>Most accurate:</b> {GLib.markup_escape_text(acc.engine)} ({acc.accuracy:.1f}%)")
return False
# ----- wakeword benchmark ----------------------------------------------
def _read_tts_fields(self):
"""(url, api_key_env, model, voices) from the Benchmark tab, persisted to cfg."""
url = self.wwb_url.get_text().strip().rstrip("/")
key = self.wwb_key.get_text().strip()
model = self.wwb_model.get_text().strip()
voices = [v.strip() for v in self.wwb_voices.get_text().split(",") if v.strip()]
self.cfg.tts_url, self.cfg.tts_api_key_env = url, key
self.cfg.tts_model, self.cfg.tts_voices = model, voices
return url, key, model, voices
def _tts_reload(self) -> None:
"""⟳ — load the model list ({url}/models) and voice list into the dropdowns."""
url = self.wwb_url.get_text().strip().rstrip("/")
key = self.wwb_key.get_text().strip()
if not url:
self._error("Enter the TTS server URL first (e.g. http://localhost:8880/v1).")
return
self.wwb_test_lbl.set_markup("<i>Loading models &amp; voices…</i>")
def work():
models = stt.list_models(url, key)
voices = wakeword_bench.list_voices(url, api_key_env=key)
def apply():
_fill_combo(self.wwb_model, models, self.wwb_model.get_text())
self.wwb_voices.set_models(voices)
if voices and not self.wwb_voices.get_text().strip():
self.wwb_voices.set_text(", ".join(voices)) # default: test them all
colour = "#34c759" if (models or voices) else "#ff9f0a"
self.wwb_test_lbl.set_markup(
f"<span foreground='{colour}'>Loaded {len(models)} models, "
f"{len(voices)} voices.</span>")
return False
GLib.idle_add(apply)
threading.Thread(target=work, daemon=True).start()
def _connect_tts(self, _b) -> None:
url, key, model, voices = self._read_tts_fields()
if not url or not model:
self._error("Enter the TTS URL and model (press ⟳ to load them).")
return
self.wwb_test_btn.set_sensitive(False)
self.wwb_test_lbl.set_markup("<i>Testing synthesis…</i>")
voice = voices[0] if voices else "alloy"
def work():
ok, msg = wakeword_bench.probe(url, model=model, voice=voice, api_key_env=key)
GLib.idle_add(self._connect_tts_done, ok, msg)
threading.Thread(target=work, daemon=True).start()
def _connect_tts_done(self, ok, msg) -> bool:
self.wwb_test_btn.set_sensitive(True)
colour = "#34c759" if ok else "#ff3b30"
self.wwb_test_lbl.set_markup(f"<span foreground='{colour}'>{GLib.markup_escape_text(msg)}</span>")
return False
def _run_wakeword_bench(self, _b) -> None:
url, key, model, voices = self._read_tts_fields()
count = int(self.wwb_count.get_value())
if not url:
self._error("Enter the TTS server URL (e.g. http://localhost:8880/v1).")
return
if not model:
self._error("Set a TTS model id (your endpoint's /audio/speech model).")
return
if not voices:
self._error("Add at least one voice (or press Connect to fetch them).")
return
self.wwb_summary.set_markup("<i>Synthesizing and streaming… this talks to your TTS "
"and wyoming-openwakeword servers.</i>")
def work():
def prog(done, total, u):
GLib.idle_add(self._wwbench_progress, done, total, u)
try:
res = wakeword_bench.run(
tts_url=url, tts_api_key_env=key, tts_model=model, voices=voices,
wakeword_model=self.cfg.wakeword_model, wakeword_uri=self.cfg.wakeword_uri,
language=self.cfg.language, count=count, progress=prog)
except Exception as exc: # noqa: BLE001 - surface setup errors
GLib.idle_add(self._wwbench_error, str(exc))
return
GLib.idle_add(self._wwbench_done, res)
threading.Thread(target=work, daemon=True).start()
def _wwbench_progress(self, done, total, u) -> bool:
tag = "" if (u.ok and (u.detections > 0) == u.has_wakeword) else ("" if u.ok else "")
kind = "wake" if u.has_wakeword else "filler"
self.wwb_summary.set_markup(
f"<i>{done}/{total}</i> {tag} {kind} · {GLib.markup_escape_text(u.voice)}"
+ (f" · <span foreground='#ff3b30'>{GLib.markup_escape_text(u.error)}</span>" if u.error else ""))
return False
def _wwbench_error(self, msg: str) -> bool:
self.wwb_summary.set_markup(f"<span foreground='#ff3b30'>{GLib.markup_escape_text(msg)}</span>")
return False
def _wwbench_done(self, res) -> bool:
if res.expected == 0:
errs = {u.error for u in res.utterances if u.error}
hint = (" " + GLib.markup_escape_text(next(iter(errs)))) if errs else ""
self.wwb_summary.set_markup(
f"<span foreground='#ff3b30'>No utterances synthesized — check the TTS model/endpoint.</span>{hint}")
return False
by_voice = res.recall_by_voice()
voice_bits = " ".join(
f"{GLib.markup_escape_text(v)} {d}/{t}" for v, (d, t) in sorted(by_voice.items()))
colour = "#34c759" if res.recall >= 0.9 and res.false_fires == 0 else (
"#ff9f0a" if res.recall >= 0.6 else "#ff3b30")
self.wwb_summary.set_markup(
f"<b><span foreground='{colour}'>Recall {res.recall * 100:.0f}%</span></b> "
f"({res.detected}/{res.expected} fired) · "
f"<b>False fires:</b> {res.false_fires} in {len(res.filler)} filler · "
f"{res.seconds:.0f}s\n<small>per voice: {voice_bits}</small>")
return False
# ===== About ============================================================
def _build_about(self, page: Gtk.Box) -> None:
_infobox(page, "About Blitztext: version, source code, recent changes, and licence.")
@ -1213,6 +1401,11 @@ class SettingsDialog:
c.wakeword_sound_done = self.ww_snd_done.get_filename() or ""
c.wakeword_silence_seconds = float(self.ww_silence.get_text())
c.cancel_keywords = [k.strip() for k in self.cancel_keywords.get_text().split(",") if k.strip()]
c.send_keywords = [k.strip() for k in self.send_keywords.get_text().split(",") if k.strip()]
c.tts_url = self.wwb_url.get_text().strip().rstrip("/")
c.tts_api_key_env = self.wwb_key.get_text().strip()
c.tts_model = self.wwb_model.get_text().strip()
c.tts_voices = [v.strip() for v in self.wwb_voices.get_text().split(",") if v.strip()]
c.mic = self._selected_mic_name()
c.output = self.gen_output.get_active_text() or "type"
c.language = self.gen_lang.get_text().strip()

View File

@ -130,6 +130,35 @@ def is_cancel(transcript: str, cancel_keywords, *, threshold: float = DEFAULT_TH
return None
def match_send(transcript: str, send_keywords, *, threshold: float = DEFAULT_THRESHOLD):
"""Detect a spoken 'send' keyword at an edge; return (keyword, cleaned_text).
Like is_cancel, but the keyword is *stripped* and the remaining text is meant
to be delivered and submitted with Enter the spoken equivalent of
stop+paste+Enter. Returns (None, transcript) when nothing matches. Matched the
same edge-anchored, ASR-tolerant way as routing keywords, so the word deep
inside a sentence won't trigger it — only at the start or end.
"""
if not send_keywords:
return None, transcript
tokens = normalize(transcript)
if not tokens:
return None, transcript
best = None # (score, span, position, keyword)
for kw in send_keywords:
kw_tokens = normalize(kw)
m = _match_window(tokens, kw_tokens, threshold)
if m is None:
continue
position, score, span = m
if best is None or (score, span) > (best[0], best[1]):
best = (score, span, position, kw)
if best is None:
return None, transcript
_score, span, position, kw = best
return kw, _strip_span(transcript, span, position)
def _strip_span(transcript: str, span_words: int, position: str) -> str:
"""Remove the matched keyword from the given edge of the original transcript.

View File

@ -0,0 +1,379 @@
"""Benchmark the wakeword detector with synthetic speech.
Generates short utterances random filler text with the wakeword spoken in it,
plus pure-filler utterances with no wakeword synthesizes each in a random voice
via the *same* (OpenAI-compatible) endpoint configured for speech-to-text, then
streams the audio to the running wyoming-openwakeword server and counts what it
detects. Reports:
recall share of wakeword utterances that fired at least one detection
false fires detections during pure-filler utterances (should be zero)
per-voice recall, so you can see which voices your model handles
TTS reuses the STT engine on purpose: on a typical NIM/OpenAI-compatible setup
the same server answers ``/audio/speech``, so there's nothing extra to configure
beyond a TTS model id and a voice list.
"""
from __future__ import annotations
import io
import json
import os
import random
import socket
import time
import urllib.error
import urllib.request
import wave
from dataclasses import dataclass, field
from urllib.parse import urlparse
# Bilingual filler so the synthesized speech is sentence-like (the detector sees
# realistic context around the wakeword, not just the bare phrase). Unknown
# languages fall back to English text — the voice still speaks it in its accent.
_FILLERS = {
"en": [
"the weather today is unusually calm and bright",
"remind me to call the office before noon",
"i think the train leaves around half past nine",
"could you put the report on my desk later",
"we should grab a coffee once this is done",
"the package was delivered to the wrong address again",
],
"de": [
"das wetter ist heute ungewöhnlich ruhig und klar",
"erinnere mich daran das büro vor mittag anzurufen",
"ich glaube der zug fährt gegen halb zehn",
"könntest du den bericht später auf meinen tisch legen",
"wir sollten einen kaffee trinken wenn das erledigt ist",
"das paket wurde wieder an die falsche adresse geliefert",
],
}
# OpenAI's stock voices; a sensible default for any OpenAI-compatible TTS. Local
# servers (Kokoro, openedai-speech, …) expose their own names — override in the
# config / Settings to match what your endpoint actually serves.
DEFAULT_VOICES = ["alloy", "echo", "fable", "onyx", "nova", "shimmer"]
_TARGET_RATE = 16000 # wyoming-openwakeword expects 16 kHz mono s16le
_CHUNK_BYTES = 3200 # 100 ms per audio-chunk, matching the live listener
@dataclass
class Utterance:
text: str
has_wakeword: bool
voice: str
detections: int = 0
ok: bool = False # synthesized + streamed without error
error: str = ""
@dataclass
class BenchResult:
utterances: list[Utterance] = field(default_factory=list)
seconds: float = 0.0
@property
def wake(self) -> list[Utterance]:
return [u for u in self.utterances if u.has_wakeword]
@property
def filler(self) -> list[Utterance]:
return [u for u in self.utterances if not u.has_wakeword]
@property
def detected(self) -> int:
return sum(1 for u in self.wake if u.ok and u.detections > 0)
@property
def expected(self) -> int:
return sum(1 for u in self.wake if u.ok)
@property
def recall(self) -> float:
return self.detected / self.expected if self.expected else 0.0
@property
def false_fires(self) -> int:
return sum(u.detections for u in self.filler if u.ok)
def recall_by_voice(self) -> dict[str, tuple[int, int]]:
"""voice -> (detected, expected) over wakeword utterances."""
out: dict[str, list[int]] = {}
for u in self.wake:
if not u.ok:
continue
d = out.setdefault(u.voice, [0, 0])
d[1] += 1
if u.detections > 0:
d[0] += 1
return {v: (d[0], d[1]) for v, d in out.items()}
def wakeword_phrase(model: str) -> str:
"""Turn a wakeword model id into the phrase to speak (best effort).
"okay_computer" -> "okay computer", "hey_jarvis" -> "hey jarvis". Strips a
trailing version/format suffix like "_v0.1" or ".tflite".
"""
name = model.rsplit("/", 1)[-1]
for ext in (".tflite", ".onnx"):
if name.endswith(ext):
name = name[: -len(ext)]
name = name.replace("_", " ").replace("-", " ")
# Drop a trailing token that is just a version like "v0.1".
parts = [p for p in name.split() if not (p.startswith("v") and any(c.isdigit() for c in p))]
return " ".join(parts).strip() or name.strip()
def _filler_pool(language: str) -> list[str]:
lang = (language or "").lower()
if lang.startswith("de"):
return _FILLERS["de"]
return _FILLERS["en"]
def build_utterances(phrase: str, count: int, language: str, *,
filler_count: int | None = None, voices=None,
rng: random.Random | None = None) -> list[Utterance]:
"""Build `count` wakeword utterances + `filler_count` pure-filler ones.
Each wakeword utterance embeds `phrase` at the start, middle, or end of a
random filler sentence; voices are assigned round-robin-ish at random so the
set covers every configured voice. Deterministic when `rng` is seeded.
"""
rng = rng or random.Random()
voices = list(voices or DEFAULT_VOICES) or DEFAULT_VOICES
pool = _filler_pool(language)
if filler_count is None:
filler_count = max(3, count // 3)
def a_voice() -> str:
return rng.choice(voices)
out: list[Utterance] = []
for _ in range(count):
filler = rng.choice(pool)
where = rng.choice(("start", "end", "mid"))
if where == "start":
text = f"{phrase}, {filler}"
elif where == "end":
text = f"{filler}, {phrase}"
else:
words = filler.split()
cut = len(words) // 2
text = " ".join(words[:cut] + [phrase] + words[cut:])
out.append(Utterance(text=text, has_wakeword=True, voice=a_voice()))
for _ in range(filler_count):
out.append(Utterance(text=rng.choice(pool), has_wakeword=False, voice=a_voice()))
rng.shuffle(out)
return out
# --- TTS via an OpenAI-compatible /audio/speech endpoint ---------------------
def _auth_headers(api_key_env: str) -> dict:
key = os.environ.get(api_key_env or "", "")
return {"Authorization": f"Bearer {key}"} if key else {}
def synthesize(tts_url: str, text: str, *, model: str, voice: str,
api_key_env: str = "", timeout: float = 30.0) -> bytes:
"""Return 16 kHz mono s16le PCM for `text` from an OpenAI-compatible TTS.
`tts_url` is the base incl. /v1 (e.g. http://localhost:8880/v1). Asks for WAV
and resamples whatever rate/-channels come back down to 16 kHz mono.
"""
base = (tts_url or "").rstrip("/")
if not base:
raise RuntimeError("No TTS URL configured (Settings → Benchmark → TTS URL).")
payload = json.dumps({
"model": model, "input": text, "voice": voice, "response_format": "wav",
}).encode("utf-8")
headers = {"Content-Type": "application/json", **_auth_headers(api_key_env)}
req = urllib.request.Request(base + "/audio/speech", data=payload, headers=headers, method="POST")
try:
with urllib.request.urlopen(req, timeout=timeout) as resp:
raw = resp.read()
except urllib.error.HTTPError as exc:
detail = exc.read().decode("utf-8", "replace")[:200]
raise RuntimeError(f"TTS HTTP {exc.code}: {detail}") from exc
except (urllib.error.URLError, OSError) as exc:
raise RuntimeError(f"TTS request failed: {exc}") from exc
return _wav_to_pcm16k(raw)
def list_voices(tts_url: str, *, api_key_env: str = "", timeout: float = 8.0) -> list[str]:
"""Best-effort voice discovery (Kokoro/XTTS expose /audio/voices or /voices).
Returns [] if the server has no such endpoint voices are then entered by
hand. Accepts the common shapes: ["a", ], {"voices": }, {"data": [{id}]}.
"""
base = (tts_url or "").rstrip("/")
if not base:
return []
for path in ("/audio/voices", "/voices"):
try:
req = urllib.request.Request(base + path, headers=_auth_headers(api_key_env))
with urllib.request.urlopen(req, timeout=timeout) as resp:
data = json.loads(resp.read())
except Exception: # noqa: BLE001 - endpoint optional/non-standard
continue
items = data.get("voices") or data.get("data") or [] if isinstance(data, dict) else data
out: list[str] = []
for it in items or []:
if isinstance(it, str):
out.append(it)
elif isinstance(it, dict):
name = it.get("id") or it.get("name") or it.get("voice")
if name:
out.append(str(name))
if out:
return out
return []
def probe(tts_url: str, *, model: str, voice: str, api_key_env: str = "",
timeout: float = 20.0) -> tuple[bool, str]:
"""Connectivity check: synthesize one short phrase. Returns (ok, message)."""
try:
pcm = synthesize(tts_url, "connection test", model=model, voice=voice,
api_key_env=api_key_env, timeout=timeout)
except Exception as exc: # noqa: BLE001 - surface the reason to the user
return False, str(exc)
return True, f"Connected — “{voice}” returned {len(pcm) / 2 / _TARGET_RATE:.1f}s of audio."
def _wav_to_pcm16k(wav_bytes: bytes) -> bytes:
"""Decode a WAV blob to 16 kHz mono signed-16 PCM (linear resample)."""
import numpy as np
with wave.open(io.BytesIO(wav_bytes), "rb") as w:
ch, width, rate, n = w.getnchannels(), w.getsampwidth(), w.getframerate(), w.getnframes()
frames = w.readframes(n)
if width != 2:
raise RuntimeError(f"Unexpected TTS sample width {width*8}-bit (need 16-bit WAV).")
a = np.frombuffer(frames, dtype=np.int16).astype(np.float32)
if ch > 1:
a = a.reshape(-1, ch).mean(axis=1)
if rate != _TARGET_RATE and a.size:
new_len = int(round(a.size * _TARGET_RATE / rate))
if new_len > 0:
xp = np.linspace(0.0, 1.0, num=a.size, endpoint=False)
x = np.linspace(0.0, 1.0, num=new_len, endpoint=False)
a = np.interp(x, xp, a)
return np.clip(a, -32768, 32767).astype("<i2").tobytes()
# --- wyoming-openwakeword detection ------------------------------------------
def count_detections(uri: str, model: str, pcm: bytes, *, settle: float = 1.5,
timeout: float = 15.0) -> int:
"""Stream `pcm` (16 kHz mono s16le) to wyoming-openwakeword; count detections.
Sends a fresh detect/audio-start//audio-stop session, then drains detection
events until the server falls quiet for `settle` seconds (or `timeout`).
"""
parsed = urlparse(uri)
host, port = parsed.hostname or "127.0.0.1", parsed.port or 10400
deadline = time.time() + timeout
detections = 0
with socket.create_connection((host, port), timeout=5.0) as sock:
_send(sock, {"type": "detect", "data": {"names": [model]}})
_send(sock, {"type": "audio-start", "data": {"rate": 16000, "width": 2, "channels": 1}})
for i in range(0, len(pcm), _CHUNK_BYTES):
chunk = pcm[i:i + _CHUNK_BYTES]
_send(sock, {"type": "audio-chunk",
"data": {"rate": 16000, "width": 2, "channels": 1},
"payload_length": len(chunk)}, chunk)
_send(sock, {"type": "audio-stop", "data": {}})
sock.settimeout(settle)
buf = b""
while time.time() < deadline:
try:
data = sock.recv(4096)
except socket.timeout:
break # server quiet for `settle`s → done
if not data:
break
buf += data
buf, n = _drain_detections(buf)
detections += n
return detections
def _send(sock: socket.socket, msg: dict, payload: bytes = b"") -> None:
sock.sendall((json.dumps(msg) + "\n").encode("utf-8"))
if payload:
sock.sendall(payload)
def _drain_detections(buf: bytes) -> tuple[bytes, int]:
"""Parse whole newline-framed messages from buf; return (rest, detections).
Consumes each message's binary payload too, so payload bytes are never
mistaken for the next header line.
"""
found = 0
while b"\n" in buf:
line, rest = buf.split(b"\n", 1)
try:
msg = json.loads(line.decode("utf-8"))
except (ValueError, UnicodeDecodeError):
return rest, found
plen = msg.get("payload_length", 0) or 0
if len(rest) < plen:
return buf, found # payload not fully arrived yet; wait for more
rest = rest[plen:]
if msg.get("type") == "detection":
found += 1
buf = rest
return buf, found
@dataclass
class EngineRun:
"""One wakeword engine's benchmark result."""
name: str
uri: str
model: str
result: BenchResult
def run(engines, *, tts_url: str, tts_model: str, tts_api_key_env: str = "", voices=None,
language: str = "", count: int = 12, seed: int | None = None,
progress=None) -> list[EngineRun]:
"""Benchmark each wakeword engine against the same synthesized speech.
`engines` is a list of objects with ``.name`` / ``.uri`` / ``.model`` (e.g.
``config.WakewordEngine``). Each engine is scored on `count` utterances of
its own wake phrase (+ filler). Audio is synthesized once per (text, voice)
and reused across engines that share a phrase, so adding engines is cheap.
Calls ``progress(engine_idx, engine_total, engine_name, done, total, u)``.
"""
cache: dict[tuple[str, str], bytes] = {}
runs: list[EngineRun] = []
n_eng = len(engines)
for ei, eng in enumerate(engines, 1):
phrase = wakeword_phrase(eng.model)
utterances = build_utterances(phrase, count, language, voices=voices, rng=random.Random(seed))
res = BenchResult(utterances=utterances)
total = len(utterances)
t0 = time.time()
for ui, u in enumerate(utterances, 1):
try:
key = (u.text, u.voice)
pcm = cache.get(key)
if pcm is None:
pcm = synthesize(tts_url, u.text, model=tts_model, voice=u.voice,
api_key_env=tts_api_key_env)
cache[key] = pcm
u.detections = count_detections(eng.uri, eng.model, pcm)
u.ok = True
except Exception as exc: # noqa: BLE001 - record per-utterance, keep going
u.error = str(exc)
if progress:
progress(ei, n_eng, eng.name or eng.model, ui, total, u)
res.seconds = time.time() - t0
runs.append(EngineRun(eng.name or eng.model, eng.uri, eng.model, res))
return runs

View File

@ -0,0 +1,83 @@
"""Spoken send: a configured word delivers the clip AND presses Enter.
The spoken equivalent of stop+paste+Enter mainly hands-free. The keyword is
stripped from the transcript before the rest is delivered and submitted.
"""
import blitztext.daemon as dm
import blitztext.paste as paste
from blitztext.config import Config, Workflow
from blitztext.daemon import Daemon
from blitztext.routing import match_send
# -- matcher --------------------------------------------------------------------
def test_match_send_strips_edges_only():
kws = ["computer send", "computer abschicken"]
assert match_send("computer send", kws)[0] == "computer send" # whole utterance
kw, text = match_send("hey team the build is green computer send", kws)
assert kw == "computer send" and text == "hey team the build is green" # end edge, stripped
kw, text = match_send("computer send hey team the build is green", kws)
assert kw == "computer send" and text == "hey team the build is green" # start edge, stripped
# A bare word inside a sentence must NOT submit — the phrase is distinctive.
assert match_send("please send me the report tomorrow", kws)[0] is None
assert match_send("computer send", [])[0] is None # disabled
assert match_send("", kws)[0] is None
# -- pipeline -------------------------------------------------------------------
def _wire_clean_pipeline(monkeypatch):
monkeypatch.setattr(dm, "detect_recorder", lambda pref="auto": "pw-record")
monkeypatch.setattr(dm, "notify", lambda *a, **k: None)
monkeypatch.setattr(dm.quality, "analyze_wav", lambda p: (2.0, 0.5))
monkeypatch.setattr(dm.quality, "too_quiet", lambda *a, **k: False)
monkeypatch.setattr(dm.quality, "clean", lambda t, **k: t)
monkeypatch.setattr(dm.quality, "is_hallucination", lambda *a, **k: False)
delivered = []
monkeypatch.setattr(dm, "deliver", lambda *a, **k: delivered.append((a, k)))
enters = []
monkeypatch.setattr(paste, "press_enter", lambda win=None: enters.append(win))
return delivered, enters
def test_process_sends_with_enter_when_send_spoken(monkeypatch, tmp_path):
delivered, enters = _wire_clean_pipeline(monkeypatch)
monkeypatch.setattr(dm.stt, "transcribe", lambda *a, **k: "die nachricht ist fertig computer send")
d = Daemon(Config())
d._prepared = True
d.cfg.send_keywords = ["computer send"]
d.cfg.cancel_keywords = []
audio = tmp_path / "clip.wav"; audio.write_bytes(b"x")
d._process(audio, Workflow(name="Transcribe", hotkey="", mode="transcribe"), None)
assert len(delivered) == 1
assert delivered[0][0][0] == "die nachricht ist fertig" # keyword stripped
assert len(enters) == 1, "send keyword must press Enter"
def test_process_no_enter_without_send_word(monkeypatch, tmp_path):
"""Control: the same path without a send word delivers but never hits Enter."""
delivered, enters = _wire_clean_pipeline(monkeypatch)
monkeypatch.setattr(dm.stt, "transcribe", lambda *a, **k: "die nachricht ist fertig")
d = Daemon(Config())
d._prepared = True
d.cfg.send_keywords = ["computer send"]
d.cfg.cancel_keywords = []
audio = tmp_path / "clip.wav"; audio.write_bytes(b"x")
d._process(audio, Workflow(name="Transcribe", hotkey="", mode="transcribe"), None)
assert len(delivered) == 1
assert delivered[0][0][0] == "die nachricht ist fertig"
assert enters == [], "no send word → no Enter"
def test_send_keywords_round_trip(tmp_path):
from blitztext.config import load, save
p = tmp_path / "config.toml"
cfg = load(p)
assert cfg.send_keywords == [] # shipped default: off
cfg.send_keywords = ["computer send", "computer abschicken"]
save(cfg, p)
assert load(p).send_keywords == ["computer send", "computer abschicken"]

View File

@ -0,0 +1,68 @@
"""Wakeword benchmark: deterministic parts (script, audio decode, framing, scoring).
The network parts (TTS /audio/speech, streaming to wyoming-openwakeword) need
live servers and are exercised manually from Settings Benchmark; here we cover
everything that can be checked offline.
"""
import io
import json
import random
import wave
import numpy as np
from blitztext import wakeword_bench as wb
def test_wakeword_phrase():
assert wb.wakeword_phrase("okay_computer") == "okay computer"
assert wb.wakeword_phrase("hey_jarvis") == "hey jarvis"
assert wb.wakeword_phrase("models/okay_nabu.tflite") == "okay nabu"
assert wb.wakeword_phrase("alexa_v0.1") == "alexa"
def test_build_utterances_deterministic_and_covers_voices():
a = wb.build_utterances("computer", 8, "de", voices=["nova", "onyx"], rng=random.Random(7))
b = wb.build_utterances("computer", 8, "de", voices=["nova", "onyx"], rng=random.Random(7))
key = lambda us: [(u.text, u.has_wakeword, u.voice) for u in us]
assert key(a) == key(b) # seeded → reproducible
assert sum(u.has_wakeword for u in a) == 8
assert all(u.voice in ("nova", "onyx") for u in a)
assert all("computer" in u.text for u in a if u.has_wakeword)
assert any(not u.has_wakeword for u in a) # filler added for false-fire check
def test_wav_to_pcm_resamples_to_16k_mono():
buf = io.BytesIO()
with wave.open(buf, "wb") as w:
w.setnchannels(2); w.setsampwidth(2); w.setframerate(48000)
n = 48000 # 1.0 s stereo @ 48 kHz
tone = (np.sin(2 * np.pi * 440 * np.linspace(0, 1, n, endpoint=False)) * 8000).astype("<i2")
w.writeframes(np.repeat(tone, 2).tobytes())
pcm = wb._wav_to_pcm16k(buf.getvalue())
assert len(pcm) == 16000 * 2 # 1.0 s of 16 kHz mono s16le
def test_drain_detections_handles_payloads():
det = (json.dumps({"type": "detection", "data": {}}) + "\n").encode()
chunk = (json.dumps({"type": "audio-chunk", "payload_length": 3}) + "\n").encode() + b"abc"
rest, n = wb._drain_detections(chunk + det)
assert n == 1 and rest == b"" # payload skipped, detection counted
held = (json.dumps({"type": "x", "payload_length": 10}) + "\n").encode() + b"ab"
rest2, n2 = wb._drain_detections(held)
assert n2 == 0 and rest2 == held # incomplete payload held, not misparsed
def test_bench_result_metrics():
r = wb.BenchResult(utterances=[
wb.Utterance("a", True, "nova", detections=1, ok=True),
wb.Utterance("b", True, "nova", detections=0, ok=True),
wb.Utterance("c", True, "onyx", detections=2, ok=True),
wb.Utterance("d", False, "onyx", detections=1, ok=True), # false fire
wb.Utterance("e", True, "x", ok=False, error="boom"), # failed → excluded
])
assert r.expected == 3 and r.detected == 2
assert abs(r.recall - 2 / 3) < 1e-9
assert r.false_fires == 1
assert r.recall_by_voice() == {"nova": (1, 2), "onyx": (1, 1)}