Speechmatics logo

Speechmatics

Freemium
訪問

Speechmatics provides speech-to-text APIs for pre-recorded and real-time audio, with support for multilingual speech, multiple speakers, and voice-agent workflows. It is designed for teams building transcription and Voice AI applications.

Speechmaticsとは?

Speechmatics is a speech-to-text API platform for transcribing live or pre-recorded audio and video. Its models are designed for conversational audio, including accents, multiple speakers, multilingual speech, and language switching within a sentence. The platform offers general-purpose, enhanced, medical, and voice-agent-oriented models for different transcription requirements.

Core capabilities include batch transcription from audio or video files and real-time transcription from live audio or video. Depending on the model, outputs and processing options include speaker diarization, language hints and labeling, smart formatting, word-level timings, custom dictionaries, healthcare vocabulary handling, mixed-language transcription, and code-switching support.

Speechmaticsでできること

Pre-recorded and real-time transcription

Transcribe audio or video files in batch workflows, or process live audio and video for applications that need streaming speech recognition.

Multilingual and mixed-language speech

The Melia 1 model supports 56 languages without requiring a language choice up front and can switch languages mid-sentence. The platform site describes support for more than 55 languages overall.

Speaker and language labeling

Speaker diarization separates contributions in multi-speaker recordings, while language hints and labeling help identify or guide the languages present in a transcription.

Model options for different domains

Available model choices include Melia 1 for multilingual audio, Enhanced for higher-accuracy single-language transcription, Oak 1 for clinical language, and Linden 1 for real-time voice-agent use cases.

Conversation and document processing add-ons

Optional processing includes translation, chapters, topics, summaries, sentiment, PII redaction, and audio alignment. Availability and rates vary by model and processing mode.

Voice-agent turn detection

The Linden 1 real-time model includes turn detection for conversational voice-agent workflows, helping an application identify when a speaker's turn has changed.

利用シーン

“Multilingual conversation transcription”

Use Melia 1 to transcribe meetings, interviews, calls, or other recordings where speakers use multiple languages or switch languages during a conversation.

“Clinical dictation and documentation”

Use the Oak 1 medical model for clinical language, including drug names, dosages, abbreviations, and procedure terms, with multilingual support.

“Real-time voice agents”

Use the Linden 1 real-time model as the speech-recognition layer for conversational agents that need built-in turn detection.

“Recorded media processing”

Process audio or video files in batch and add outputs such as chapters, topics, summaries, translation, or sentiment analysis where the selected model and workflow support them.

よくある質問

Can Speechmatics transcribe live audio as well as recorded files?

Yes. The pricing information distinguishes pre-recorded transcription from real-time transcription. Pre-recorded workflows accept audio or video files, while real-time workflows process live audio or video.

How many languages does Speechmatics support?

The site describes support for more than 55 languages. The Melia 1 model is listed as transcribing 56 languages without choosing one language up front and can switch languages mid-sentence.

Does Speechmatics identify different speakers?

Speaker diarization is listed as an available transcription feature. Exact feature availability can differ by model, product package, and deployment.

Which model is intended for medical transcription?

Oak 1 is the medical model. It is intended for clinical language, including drug names, dosages, abbreviations, and procedure terms.

Is there a free way to start?

The pricing page says users can start free with $100 in credits and no card required. Usage rates after the included credits vary by model and processing mode.

クイック情報

Product type
Speech-to-text APIs
Input workflows
Pre-recorded audio or video and real-time audio or video
Language coverage
55+ languages; Melia 1 lists 56 languages
Notable models
Melia 1, Enhanced, Oak 1, and Linden 1
Deployment
Cloud SaaS; the product site also describes cloud, on-premises, and on-device options
Pricing model
Usage-based pricing with a free start offer and model-specific rates

Speechmatics のトラフィック分析

トラフィックデータは参考情報としてご利用ください。

ドメイン評価
73

Speechmaticsの代替品

Recall.ai Startup Program logo

Recall.ai Startup Program

www.recall.ai

A startup program for early-stage companies building products powered by meeting and conversation data. Approved applicants receive discounted recording usage, access to Recall.ai products, and support while they build and launch.

Yap logo

Yap

yapdaily.com

Yap is a voice journal for iPhone that turns a spoken account of your day into a tidy daily page. It is designed for people who want to capture and revisit memories without typing.

NeverCap logo

NeverCap

nevercap.ai

NeverCapは音声・動画を編集可能なテキストに変換するAI文字起こしツール。話者ラベル、タイムスタンプ、100以上の言語、複数ファイルの一括アップロードに対応し、月間分数制限もありません。

Clipto logo

Clipto

www.clipto.com

Clipto MCPを使うと、AIツールでローカルにインデックス化されたメディアを検索し、タイムスタンプ付きで関連箇所を見つけ、Bロールの候補、ポッドキャストの編集、映像素材ログなど、ソースに裏付けられた成果物をまとめられます。接続されたAIエージェントを通じて、ローカルの動画、音声、写真、ドキュメントを扱いたいユーザー向けに設計されています。

Loqua logo

Loqua

www.theloqua.ai

Loqua is a voice-first desktop productivity app for macOS and Windows that turns natural speech into polished text, edits selected content, translates dictation, and answers questions about user-selected screen content. It is designed for hands-free writing and supported voice workflows across desktop apps.

Minutes logo

Minutes

myminutes.ai

会議・講義・録音に対応するAI議事録・文字起こしアプリ