Measure speech recognition, voice synthesis, and full conversation turns with the same local engines used by the app.
Run every selected STT endpoint against the same reference clip and compare speed, accuracy, model, and compute device.
| Engine | Model | Device | Time | Accuracy | Output |
|---|---|---|---|---|---|
| Run a benchmark to see results. | |||||
Pick a backend and voice, set run count, then measure synthesis latency and real-time factor (RTF). RTF < 1.0 means the backend generates faster than real-time.
Benchmark multiple voices in one run. Pre-populated from your active My Voices — or reload from the backend. Results are sorted fastest first and saved to History.
Uses the sample text and Runs value from the single-voice form above. Check one voice, a few voices, or Select all, then run them together.
Last 50 benchmark sessions saved in your browser. Each row is one run session — click the backend/voice to pre-fill the form above.
Run a turn to see transcript, reply, and timing.