EnterpriseCall workspace
Public prototypeEC
Enterprise Call Summary and Data Extraction

MODEL EVALUATION

Build a benchmark

Select calls, compare models, and inspect the errors.

1. Define your test

Recordings

10 selected · 8.1 minutes of audio

Choose individual recordings (10 selected)
Advanced settings Split, repeats, terminology, and local runtime
Local hardware and runtime

This records your policy; it does not restart local models. The companion reports measured load time when available.

2. Choose STT models

Recording → transcript · five shortlisted models selected by default.

5 selected
Connection needed

Speaker labels and conversational transcription. Compare accuracy on your recordings. Documentation ↗

Cost estimate $0.22 / audio hour
Connection needed

Azure Speech key and resource endpoint required. Speaker labels enabled; use recordings shorter than 15 minutes during preview. WAV, MP3 or FLAC. Documentation ↗

Cost estimate $0.10 / audio hour
Connection needed

Pinned to universal-3-5-pro with speaker labels. Supports terminology prompting. Documentation ↗

Cost estimate $0.21 / audio hour
Connection needed

Uses gpt-4o-transcribe exactly. No speaker labels or segment timestamps; extraction must infer roles from the words. Documentation ↗

Cost estimate $0.36 / audio hour
Connection needed

Terminology hints and speaker labels. Measure speed and word accuracy on the same recordings. Documentation ↗

Cost estimate $0.26 / audio hour
Connection needed

CPU quantization or GPU inference. No speaker labels in this companion. Documentation ↗

Cost estimate Not set
Connection needed

Requires NVIDIA NeMo and suitable hardware; no speaker separation here. Documentation ↗

Cost estimate Not set
Connection needed

Uses ASSEMBLYAI_MODEL, previously universal-3-pro. Choose Universal-3.5 Pro above for new comparisons. Documentation ↗

Cost estimate Not set

3. Choose Extraction models

Transcript → structured fields · five shortlisted models selected by default.

5 selected
Connection needed

gpt-5.6-luna · Source ↗

Cost estimate $0.2 in / $1.2 out per 1M
Connection needed

gemini-2.5-flash · Source ↗

Cost estimate $0.3 in / $2.5 out per 1M
Connection needed

claude-sonnet-5 · Source ↗

Cost estimate $2 in / $10 out per 1M
Connection needed

claude-haiku-4-5-20251001 · Source ↗

Cost estimate $1 in / $5 out per 1M
Connection needed

gpt-5.6-sol · Source ↗

Cost estimate $4 in / $20 out per 1M
Connection needed

gemini-3.8-flash · Source ↗

Cost estimate $0.75 in / $3.75 out per 1M
Connection needed

No model configured · Source ↗

Cost estimate Not set
Connection needed

qwen3:8b · Source ↗

Cost estimate Not set
Connections and pricing notes

Cloud keys stay in deployment settings. Connection status means a key is configured, not that account/model access has been tested. Local models use the companion configured in Models & sources. STT hourly estimates come from your shortlist and are not verified account prices; speaker labels and subscription tiers may change the bill. Extraction uses published token rates checked September 21, 2026. Sol pricing is promotional; the additional Gemini 3.8 option also has introductory pricing. Edit rates to match your contract. Per-call costs are estimates from duration or returned token usage, not invoices. Published benchmark scores are not results on this dataset. OpenAI models use no reasoning; Gemini 2.5 uses thinkingBudget 0; Claude uses provider defaults.

10 selected services need a connection

ElevenLabs Scribe v2 · Microsoft MAI-Transcribe-2 · AssemblyAI Universal-3.5 Pro · OpenAI GPT-4o Transcribe · Deepgram Nova-3 · GPT-5.6 Luna · Gemini 2.5 Flash · Claude Sonnet 5 · Claude Haiku 4.5 · GPT-5.6 Sol

10 calls · 50 STT requests · 300 extraction requestsSequential execution · includes human-transcript baselines. Saving setup does not run models.