Select calls, compare models, and inspect the errors.
1. Define your test
Recordings
10 selected · 8.1 minutes of audio
Choose individual recordings (10 selected)Advanced settings Split, repeats, terminology, and local runtimeLocal hardware and runtime
This records your policy; it does not restart local models. The companion reports measured load time when available.
2. Choose STT models
Recording → transcript · five shortlisted models selected by default.
5 selected
Connection needed
Speaker labels and conversational transcription. Compare accuracy on your recordings. Documentation ↗
Cost estimate $0.22 / audio hour
Connection needed
Azure Speech key and resource endpoint required. Speaker labels enabled; use recordings shorter than 15 minutes during preview. WAV, MP3 or FLAC. Documentation ↗
Cost estimate $0.10 / audio hour
Connection needed
Pinned to universal-3-5-pro with speaker labels. Supports terminology prompting. Documentation ↗
Cost estimate $0.21 / audio hour
Connection needed
Uses gpt-4o-transcribe exactly. No speaker labels or segment timestamps; extraction must infer roles from the words. Documentation ↗
Cost estimate $0.36 / audio hour
Connection needed
Terminology hints and speaker labels. Measure speed and word accuracy on the same recordings. Documentation ↗
Cost estimate $0.26 / audio hour
Connection needed
CPU quantization or GPU inference. No speaker labels in this companion. Documentation ↗
Cost estimate Not set
Connection needed
Requires NVIDIA NeMo and suitable hardware; no speaker separation here. Documentation ↗
Cost estimate Not set
Connection needed
Uses ASSEMBLYAI_MODEL, previously universal-3-pro. Choose Universal-3.5 Pro above for new comparisons. Documentation ↗
Cost estimate Not set
3. Choose Extraction models
Transcript → structured fields · five shortlisted models selected by default.
Cloud keys stay in deployment settings. Connection status means a key is configured, not that account/model access has been tested. Local models use the companion configured in Models & sources. STT hourly estimates come from your shortlist and are not verified account prices; speaker labels and subscription tiers may change the bill. Extraction uses published token rates checked September 21, 2026. Sol pricing is promotional; the additional Gemini 3.8 option also has introductory pricing. Edit rates to match your contract. Per-call costs are estimates from duration or returned token usage, not invoices. Published benchmark scores are not results on this dataset. OpenAI models use no reasoning; Gemini 2.5 uses thinkingBudget 0; Claude uses provider defaults.
10 selected services need a connection
ElevenLabs Scribe v2 · Microsoft MAI-Transcribe-2 · AssemblyAI Universal-3.5 Pro · OpenAI GPT-4o Transcribe · Deepgram Nova-3 · GPT-5.6 Luna · Gemini 2.5 Flash · Claude Sonnet 5 · Claude Haiku 4.5 · GPT-5.6 Sol
10 calls · 50 STT requests · 300 extraction requestsSequential execution · includes human-transcript baselines. Saving setup does not run models.