Engines

Local and cloud services for language generation, speech recognition, and synthesis.

Language Models

Pick the Active Language Model below — it is used by default for every LLM task (persona rewriting, transcription refinement, conversation, and rehearser character analysis). The cards below let you connect engines and apply a URL with Use as LLM.

Speech Recognition

Your Docker stack containers appear at the top. Local runners and cloud APIs below.

Text to Speech

Docker-detected engines and local presets are listed together here. Click Use as TTS to apply a URL to this app’s backend settings.

100% Local

Active Language Model

Default endpoint & model for all LLM tasks. Per-task dropdowns can still override it.

Local

Cloud APIs Free tiers

Groq
Ultra-fast inference · OpenAI-compatible
Free
30 000 tokens / min 14 400 req / day Lowest latency
Get key
Endpoint https://api.groq.com/openai/v1
llama-3.3-70b-versatile qwen-qwq-32b deepseek-r1-distill-llama-70b
OpenRouter
50+ free models · OpenAI-compatible
Free
20 req / min 200 req / day (free models) Single API for all models
Get key
Endpoint https://openrouter.ai/api/v1
qwen/qwen3-235b-a22b:free deepseek/deepseek-r1-0528:free mistralai/mistral-7b-instruct:free
Google Gemini
Gemini 2.5 Flash · Generous free tier
Free
250 000 tokens / min 15 req / min free 1M context window
Get key
Endpoint https://generativelanguage.googleapis.com/v1beta/openai/
Uses OpenAI-compat wrapper — use model gemini-2.5-flash
Mistral AI
OpenAI-compatible · EU-based
Free
1B tokens / month 2 req / min free GDPR-compliant
Get key
Endpoint https://api.mistral.ai/v1
mistral-small-latest mistral-large-latest

Additional Settings

LLM Refinement Defaults

Automatic text cleanup after transcription

When on, the LLM refinement runs immediately after each capture.
Model name sent to the active LLM URL. Leave empty to use the app's current default.

Local

Cloud APIs

Free tiers
Groq Whisper
whisper-large-v3-turbo · OpenAI-compatible
Free
2 000 req / day Fastest cloud STT Cloud · 0 VRAM
Get key
Endpoint https://api.groq.com/openai/v1
whisper-large-v3-turbo whisper-large-v3 distil-whisper-large-v3-en
Save key in Settings → Groq API key to enable Groq Whisper in the STT dropdown.
HuggingFace Inference
Serverless Whisper models
Free
~1 000 req / day Slower cold starts Many model variants
Get key
Endpoint https://api-inference.huggingface.co/models/openai/whisper-large-v3
AssemblyAI
High-accuracy transcription + speaker diarization
Free
100 h lifetime Speaker labels Auto-chapters
Get key
Endpoint https://api.assemblyai.com/v2/transcript

Additional Settings

STT Connections

Set up endpoints for Speech-to-Text backends.

Default recognition endpoint. Expected: POST /v1/audio/transcriptions.
GPU-accelerated Whisper via CTranslate2. ~70× RT · 1.5 GB VRAM.
Lightweight C++ Whisper server. ~8–15× RT CPU · ~1 GB RAM.
Used for Groq Whisper STT and LLM. Free: 2 000 req/day.
Quick test Ready

STT Transcription Defaults

Default language and preferred STT backend for captures.

Overrides the active STT URL for STT-TTS panel captures.

Local

Cloud APIs

Free tiers
ElevenLabs
High-quality voice cloning & synthesis
Free
10 000 chars / month 2 500 char / request max Voice cloning supported
Get key
Endpoint https://api.elevenlabs.io/v1/text-to-speech
Fish Audio
Voice cloning & multilingual TTS
Free
1 h audio / month 100 req / min 30+ languages
Get key
Endpoint https://api.fish.audio/v1/tts
Kokoro TTS
82M model · HuggingFace Spaces demo
Demo
Free web demo <$1 per 1M chars (paid) High naturalness
Use the HF Spaces web demo for quick tests, or run Kokoro locally via Docker for production use.
Open Kokoro HF Space

Additional Settings

TTS Playback & API Behavior

Controls how previews play and how OpenAI-compatible requests are shaped.

Buffered keeps Save WAV available. Streaming starts sooner.
Controls payload shape for the normal TTS API URL.

Default Playback Voice

Pre-selected voice in the STT-TTS panel

Add Custom Engine