Engines

Local and cloud services for language generation, speech recognition, and synthesis.

Language Models

Pick the Active Language Model below. Cards connect engines and apply a URL with Use as LLM.

Speech Recognition

Docker stack containers appear first, with local runners and cloud APIs below.

Text to Speech

Docker-detected engines and local presets are listed together. Apply a URL with Use as TTS.

100% Local

Image Generation

Used to generate character profile pictures. Pick the active provider below and add its API key.

Active Language Model

Default endpoint & model for all LLM tasks. Per-task dropdowns can still override it.

Local

Cloud APIs Free tiers

Groq
Ultra-fast inference · OpenAI-compatible
Free
30 000 tokens / min 14 400 req / day Lowest latency
Get key
Endpoint https://api.groq.com/openai/v1
llama-3.3-70b-versatile qwen-qwq-32b deepseek-r1-distill-llama-70b
OpenRouter
50+ free models · OpenAI-compatible
Free
20 req / min 200 req / day (free models) Single API for all models
Get key
Endpoint https://openrouter.ai/api/v1
qwen/qwen3-235b-a22b:free deepseek/deepseek-r1-0528:free mistralai/mistral-7b-instruct:free
Google Gemini
Gemini 2.5 Flash · Generous free tier
Free
250 000 tokens / min 15 req / min free 1M context window
Get key
Endpoint https://generativelanguage.googleapis.com/v1beta/openai/
Uses OpenAI-compat wrapper — use model gemini-2.5-flash
Mistral AI
OpenAI-compatible · EU-based
Free
1B tokens / month 2 req / min free GDPR-compliant
Get key
Endpoint https://api.mistral.ai/v1
mistral-small-latest mistral-large-latest
OpenAI
GPT models · pay-as-you-go
Paid
Get key
Endpoint https://api.openai.com/v1
gpt-4o gpt-4o-mini o3-mini
Anthropic
Claude models · pay-as-you-go
Paid
Get key
Endpoint https://api.anthropic.com/v1
Not OpenAI-compatible — needs the Anthropic Messages API, not /chat/completions

Additional Settings

LLM Refinement Defaults

Automatic text cleanup after transcription

When on, the LLM refinement runs immediately after each capture.
Model name sent to the active LLM URL. Leave empty to use the app's current default.

Local

Cloud APIs

Free tiers
Groq Whisper
whisper-large-v3-turbo · OpenAI-compatible
Free
2 000 req / day Fastest cloud STT Cloud · 0 VRAM
Get key
Endpoint https://api.groq.com/openai/v1
whisper-large-v3-turbo whisper-large-v3 distil-whisper-large-v3-en
Save key in Settings → Groq API key to enable Groq Whisper in the STT dropdown.
HuggingFace Inference
Serverless Whisper models
Free
~1 000 req / day Slower cold starts Many model variants
Get key
Endpoint https://api-inference.huggingface.co/models/openai/whisper-large-v3
AssemblyAI
High-accuracy transcription + speaker diarization
Free
100 h lifetime Speaker labels Auto-chapters
Get key
Endpoint https://api.assemblyai.com/v2/transcript

Additional Settings

STT Connections

Set up endpoints for Speech-to-Text backends.

Default recognition endpoint. Expected: POST /v1/audio/transcriptions.
GPU-accelerated Whisper via CTranslate2. ~70× RT · 1.5 GB VRAM.
Lightweight C++ Whisper server. ~8–15× RT CPU · ~1 GB RAM.
Used for Groq Whisper STT and LLM. Free: 2 000 req/day.
Quick test Ready

STT Transcription Defaults

Default language and preferred STT backend for captures.

Overrides the active STT URL for STT-TTS panel captures.

Local

Cloud APIs

Free tiers
ElevenLabs
High-quality voice cloning & synthesis
Free
10 000 chars / month 2 500 char / request max Voice cloning supported
Get key
Endpoint https://api.elevenlabs.io/v1/text-to-speech
Fish Audio
Voice cloning & multilingual TTS
Free
1 h audio / month 100 req / min 30+ languages
Get key
Endpoint https://api.fish.audio/v1/tts
Kokoro TTS
82M model · HuggingFace Spaces demo
Demo
Free web demo <$1 per 1M chars (paid) High naturalness
Use the HF Spaces web demo for quick tests, or run Kokoro locally via Docker for production use.
Open Kokoro HF Space

Additional Settings

TTS Playback & API Behavior

Controls how previews play and how OpenAI-compatible requests are shaped.

Buffered keeps Save WAV available. Streaming starts sooner.
Controls payload shape for the normal TTS API URL.

Default Playback Voice

Pre-selected voice in the STT-TTS panel

Active Image Generation Provider

Used for the "Generate Image" action on a character's profile picture.

OpenRouter reuses the same API key as its LLM card below — no separate key needed here.

Cloud APIs

OpenAI
gpt-image-1 · Images API
Get key
Endpoint https://api.openai.com/v1/images/generations
Google
Gemini 2.5 Flash Image / Imagen
Get key
Endpoint generativelanguage.googleapis.com
Pollinations.ai
Free · no API key · no account
Free
No signup Shared free queue Third-party public service
Endpoint image.pollinations.ai
flux
No key or setup needed — just pick it as the Active Provider above. No SLA/rate-limit guarantee since it's a shared free public queue; retries automatically on "queue full" errors.
Local ComfyUI (no API key, runs on your own GPU) is configured separately above, in the Active Image Generation Provider panel.

Add Custom Engine