blitztext-app-linux/linux
mARTin-B78 34e4736080 Add voice-keyword routing
One hotkey (Ctrl+Alt+Space) records; the spoken keyword at the start or end of
speech selects the preset, which is then stripped and the rest applied. New
routing.py does ASR-tolerant matching (normalize + fuzzy + edge-window scan +
token-drift slack); config gains a [routing] section and per-preset `keywords`;
the daemon adds a "route" mode and biases Whisper with the keywords as hotwords.
Falls back to a default preset when no keyword is recognised.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 09:47:12 +02:00
..
blitztext Add voice-keyword routing 2026-06-04 09:47:12 +02:00
packaging Add Debian (.deb) installer (v1.1.0) 2026-06-04 09:25:32 +02:00
.gitignore Add Debian (.deb) installer (v1.1.0) 2026-06-04 09:25:32 +02:00
blitztext.service Add Blitztext for Linux v1.0.0 — native dictation tool 2026-06-03 22:44:28 +02:00
CHANGELOG.md Add Debian (.deb) installer (v1.1.0) 2026-06-04 09:25:32 +02:00
install.sh Add Blitztext for Linux v1.0.0 — native dictation tool 2026-06-03 22:44:28 +02:00
README.md Add Debian (.deb) installer (v1.1.0) 2026-06-04 09:25:32 +02:00
requirements.txt Add Blitztext for Linux v1.0.0 — native dictation tool 2026-06-03 22:44:28 +02:00

Blitztext for Linux (native dictation)

A native Linux port of the Blitztext workflow: focus any text field → press a hotkey → speak → the text is typed into that field, optionally rewritten by an LLM first (e.g. turn rough speech into a nicer, more detailed email).

This runs on the host (not in a container), so it can type into any application — the Linux equivalent of the macOS app's Accessibility-based auto-paste. (A sandboxed Docker/browser version can't do that; an earlier experiment along those lines was moved out to ~/Docker/correspondence/blitztext.) Transcription is local via faster-whisper; only the optional rewrite step calls out to an LLM.

How it works

hotkey ──▶ record mic (pw-record/arecord) ──▶ faster-whisper (local)
                                                    │
                          ┌── mode "transcribe" ────┤
                          │                          └── mode "rewrite": LLM (OpenAI-compatible)
                          ▼
                 xdotool types it into the focused window

Each hotkey toggles: press to start recording, press again to stop — then it transcribes, optionally rewrites, and types the result where your cursor is.

Requirements

  • X11 session (this uses xdotool; Wayland would need ydotool/wtype).
  • Host tools: xdotool, notify-send (libnotify-bin), and a recorder (pw-record from pipewire, or arecord/parecord).
    sudo apt install xdotool libnotify-bin pipewire-bin
    
  • Python 3.11+.

Install

Build a .deb and install it with the Software app or apt:

cd linux
bash packaging/build-deb.sh            # -> dist/blitztext_<ver>_arm64.deb
sudo apt install ./dist/blitztext_*.deb   # or double-click the .deb in Files

This installs blitztext to /opt/blitztext (a self-contained bundle — no pip step), adds a Blitztext entry to your app grid, and pulls in the system deps (python3-gi, xdotool, libnotify-bin, a recorder). Launch it from the app grid, or run blitztext / blitztext gui from a terminal. Remove with sudo apt remove blitztext.

Option B — run from source (venv)

cd linux
./install.sh

This creates .venv, installs faster-whisper + pynput, and writes the default config to ~/.config/blitztext/config.toml.

For the tray from source, the venv must be built on a Python that can see the system python3-giinstall.sh uses python3 -m venv --system-site-packages, so use the system /usr/bin/python3 (a conda/miniforge Python won't see the apt-installed gi). The .deb handles this for you.

Run

Three front-ends, same engine (local Whisper + global hotkeys + xdotool typing):

# optional: only needed for the "rewrite" workflows
export OPENAI_API_KEY=sk-...

.venv/bin/python -m blitztext tray   # system-tray menu (macOS-menu-bar-like, default)
.venv/bin/python -m blitztext gui    # control-panel window
.venv/bin/python -m blitztext run    # headless, hotkeys only

The tray is the closest match to the macOS menu-bar app: a status icon with a menu listing every workflow (click to record), plus Show panel, Settings…, and Quit. It needs PyGObject (python3-gi) and the GTK/AppIndicator typelibs — already present on a standard Ubuntu GNOME install (the .deb declares them as dependencies):

sudo apt install python3-gi      # usually already installed
.venv/bin/python -m blitztext tray

If PyGObject isn't visible, tray prints this hint and falls back to the window. The venv is created with --system-site-packages so it can see the system gi — build it from /usr/bin/python3, not a conda/miniforge Python.

Either way, focus any text field and trigger a workflow — by tray menu, panel button, or hotkey (defaults):

Hotkey Workflow What it does
Ctrl+Alt+Space Transcribe Types the raw transcript
Ctrl+Alt+E Nicer email Rewrites speech into a polished email
Ctrl+Alt+I Improve text Proofreads / improves wording
Ctrl+Alt+C Calm down Rewrites an angry message into a calm one
Ctrl+Alt+J Add emojis Adds fitting emojis

Configuration

Everything lives in ~/.config/blitztext/config.toml (python -m blitztext config-path prints the location). You can change hotkeys, the Whisper model, and the rewrite endpoint, and add/edit [[workflow]] blocks with your own prompts.

Local Whisper

[whisper]
model = "small"      # tiny|base|small|medium|large-v3, or a local model path
device = "auto"      # auto tries cuda, falls back to cpu
compute_type = "auto"

On this arm64 host the pip ctranslate2 wheel is CPU-only, so it runs on the Grace CPU with int8. That's fast for dictation (≈2s for a 10s clip with small). device = "auto" attempts CUDA and falls back automatically — to get GPU you'd need a CUDA-enabled CTranslate2 build for aarch64/sm_121.

Rewrite endpoint (OpenAI or your local LLM)

[rewrite]
base_url = "https://api.openai.com/v1"   # or e.g. http://localhost:8000/v1 for vLLM/llama-swap
api_key_env = "OPENAI_API_KEY"
model = "gpt-4o-mini"

Point base_url at a local OpenAI-compatible server (vLLM, llama-swap) to keep rewriting fully on-box too.

Run on login

See blitztext.service for a systemd user unit.

Verified

On this machine (Ubuntu/GNOME, X11, GB10): recorder → valid 16 kHz WAV; faster-whisper CPU transcription accurate; config + all hotkeys parse; and xdotool typing of German text into a focused GTK field. The live global-hotkey loop and the LLM rewrite HTTP call were not auto-tested here (the former hijacks the keyboard during a session; the latter needs your key) — try them with the run command above.

CLI

python -m blitztext tray               # tray menu (default)
python -m blitztext gui                # control-panel window
python -m blitztext run                # headless daemon, hotkeys only
python -m blitztext transcribe f.wav   # one-shot, prints text (no hotkeys)
python -m blitztext config-path        # print config location