Speechmatics logo

Speechmatics

Freemium
Visit

Speechmatics provides speech-to-text APIs for pre-recorded and real-time audio, with support for multilingual speech, multiple speakers, and voice-agent workflows. It is designed for teams building transcription and Voice AI applications.

What is Speechmatics?

Speechmatics is a speech-to-text API platform for transcribing live or pre-recorded audio and video. Its models are designed for conversational audio, including accents, multiple speakers, multilingual speech, and language switching within a sentence. The platform offers general-purpose, enhanced, medical, and voice-agent-oriented models for different transcription requirements.

Core capabilities include batch transcription from audio or video files and real-time transcription from live audio or video. Depending on the model, outputs and processing options include speaker diarization, language hints and labeling, smart formatting, word-level timings, custom dictionaries, healthcare vocabulary handling, mixed-language transcription, and code-switching support.

What can Speechmatics do?

Pre-recorded and real-time transcription

Transcribe audio or video files in batch workflows, or process live audio and video for applications that need streaming speech recognition.

Multilingual and mixed-language speech

The Melia 1 model supports 56 languages without requiring a language choice up front and can switch languages mid-sentence. The platform site describes support for more than 55 languages overall.

Speaker and language labeling

Speaker diarization separates contributions in multi-speaker recordings, while language hints and labeling help identify or guide the languages present in a transcription.

Model options for different domains

Available model choices include Melia 1 for multilingual audio, Enhanced for higher-accuracy single-language transcription, Oak 1 for clinical language, and Linden 1 for real-time voice-agent use cases.

Conversation and document processing add-ons

Optional processing includes translation, chapters, topics, summaries, sentiment, PII redaction, and audio alignment. Availability and rates vary by model and processing mode.

Voice-agent turn detection

The Linden 1 real-time model includes turn detection for conversational voice-agent workflows, helping an application identify when a speaker's turn has changed.

Use Cases

“Multilingual conversation transcription”

Use Melia 1 to transcribe meetings, interviews, calls, or other recordings where speakers use multiple languages or switch languages during a conversation.

“Clinical dictation and documentation”

Use the Oak 1 medical model for clinical language, including drug names, dosages, abbreviations, and procedure terms, with multilingual support.

“Real-time voice agents”

Use the Linden 1 real-time model as the speech-recognition layer for conversational agents that need built-in turn detection.

“Recorded media processing”

Process audio or video files in batch and add outputs such as chapters, topics, summaries, translation, or sentiment analysis where the selected model and workflow support them.

Frequently Asked Questions

Can Speechmatics transcribe live audio as well as recorded files?

Yes. The pricing information distinguishes pre-recorded transcription from real-time transcription. Pre-recorded workflows accept audio or video files, while real-time workflows process live audio or video.

How many languages does Speechmatics support?

The site describes support for more than 55 languages. The Melia 1 model is listed as transcribing 56 languages without choosing one language up front and can switch languages mid-sentence.

Does Speechmatics identify different speakers?

Speaker diarization is listed as an available transcription feature. Exact feature availability can differ by model, product package, and deployment.

Which model is intended for medical transcription?

Oak 1 is the medical model. It is intended for clinical language, including drug names, dosages, abbreviations, and procedure terms.

Is there a free way to start?

The pricing page says users can start free with $100 in credits and no card required. Usage rates after the included credits vary by model and processing mode.

Quick Facts

Product type
Speech-to-text APIs
Input workflows
Pre-recorded audio or video and real-time audio or video
Language coverage
55+ languages; Melia 1 lists 56 languages
Notable models
Melia 1, Enhanced, Oak 1, and Linden 1
Deployment
Cloud SaaS; the product site also describes cloud, on-premises, and on-device options
Pricing model
Usage-based pricing with a free start offer and model-specific rates

Speechmatics Traffic Analysis

Traffic data is for reference only.

Domain Rating
73

Speechmatics Alternatives

Recall.ai Startup Program logo

Recall.ai Startup Program

www.recall.ai

A startup program for early-stage companies building products powered by meeting and conversation data. Approved applicants receive discounted recording usage, access to Recall.ai products, and support while they build and launch.

Yap logo

Yap

yapdaily.com

Yap is a voice journal for iPhone that turns a spoken account of your day into a tidy daily page. It is designed for people who want to capture and revisit memories without typing.

NeverCap logo

NeverCap

nevercap.ai

NeverCap is an AI transcription tool for audio and video, with batch uploads, speaker labels, timestamps, and 100+ languages—without monthly minute caps.

Clipto logo

Clipto

www.clipto.com

Clipto MCP lets AI tools search locally indexed media, find relevant moments with timestamps, and assemble source-backed outputs such as B-roll matches, podcast edits, and footage logs. It is designed for users who want to work with local video, audio, photos, and documents through a connected AI agent.

Loqua logo

Loqua

www.theloqua.ai

Loqua is a voice-first desktop productivity app for macOS and Windows that turns natural speech into polished text, edits selected content, translates dictation, and answers questions about user-selected screen content. It is designed for hands-free writing and supported voice workflows across desktop apps.

Minutes logo

Minutes

myminutes.ai

AI note-taking and transcription for meetings, lectures, and recordings