- BenchRow gains a device field; Transcriber records its resolved device, so
local engines report CPU or GPU and remote engines show "remote". New Device
column in the results table.
- Accuracy is now case-sensitive by default (capitalisation counts) so an
all-lowercase transcript no longer scores 100%. Punctuation is still ignored.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
New benchmark.py: word-error-rate accuracy + a run() that times each STT engine
on a reference clip. Settings gains a Benchmark tab: pick a .wav and a matching
.txt, run all STT engines, see Time + Accuracy per engine and a summary of the
fastest and most accurate. Add presets for each model you want compared.
Fix "Local engine selected but the model isn't loaded." on Test: a cached
_transcriber_for() loads the local model on demand (the daemon only preloads it
when a local engine is active). Tested: local small 2.17s/100% vs remote :8010.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>