Pre-recorded and real-time transcription
Transcribe audio or video files in batch workflows, or process live audio and video for applications that need streaming speech recognition.
Speechmatics provides speech-to-text APIs for pre-recorded and real-time audio, with support for multilingual speech, multiple speakers, and voice-agent workflows. It is designed for teams building transcription and Voice AI applications.
Speechmatics is a speech-to-text API platform for transcribing live or pre-recorded audio and video. Its models are designed for conversational audio, including accents, multiple speakers, multilingual speech, and language switching within a sentence. The platform offers general-purpose, enhanced, medical, and voice-agent-oriented models for different transcription requirements.
Core capabilities include batch transcription from audio or video files and real-time transcription from live audio or video. Depending on the model, outputs and processing options include speaker diarization, language hints and labeling, smart formatting, word-level timings, custom dictionaries, healthcare vocabulary handling, mixed-language transcription, and code-switching support.
Transcribe audio or video files in batch workflows, or process live audio and video for applications that need streaming speech recognition.
The Melia 1 model supports 56 languages without requiring a language choice up front and can switch languages mid-sentence. The platform site describes support for more than 55 languages overall.
Speaker diarization separates contributions in multi-speaker recordings, while language hints and labeling help identify or guide the languages present in a transcription.
Available model choices include Melia 1 for multilingual audio, Enhanced for higher-accuracy single-language transcription, Oak 1 for clinical language, and Linden 1 for real-time voice-agent use cases.
Optional processing includes translation, chapters, topics, summaries, sentiment, PII redaction, and audio alignment. Availability and rates vary by model and processing mode.
The Linden 1 real-time model includes turn detection for conversational voice-agent workflows, helping an application identify when a speaker's turn has changed.
Use Melia 1 to transcribe meetings, interviews, calls, or other recordings where speakers use multiple languages or switch languages during a conversation.
Use the Oak 1 medical model for clinical language, including drug names, dosages, abbreviations, and procedure terms, with multilingual support.
Use the Linden 1 real-time model as the speech-recognition layer for conversational agents that need built-in turn detection.
Process audio or video files in batch and add outputs such as chapters, topics, summaries, translation, or sentiment analysis where the selected model and workflow support them.
Yes. The pricing information distinguishes pre-recorded transcription from real-time transcription. Pre-recorded workflows accept audio or video files, while real-time workflows process live audio or video.
The site describes support for more than 55 languages. The Melia 1 model is listed as transcribing 56 languages without choosing one language up front and can switch languages mid-sentence.
Speaker diarization is listed as an available transcription feature. Exact feature availability can differ by model, product package, and deployment.
Oak 1 is the medical model. It is intended for clinical language, including drug names, dosages, abbreviations, and procedure terms.
The pricing page says users can start free with $100 in credits and no card required. Usage rates after the included credits vary by model and processing mode.
流量数据仅供参考。
www.recall.ai
A startup program for early-stage companies building products powered by meeting and conversation data. Approved applicants receive discounted recording usage, access to Recall.ai products, and support while they build and launch.
yapdaily.com
Yap is a voice journal for iPhone that turns a spoken account of your day into a tidy daily page. It is designed for people who want to capture and revisit memories without typing.
nevercap.ai
NeverCap 是一款 AI 转录工具,可将音频和视频转换为可编辑文本,支持批量上传、说话人标注、时间戳和 100 多种语言,且无每月分钟数上限。
www.clipto.com
Clipto MCP 让 AI 工具能够搜索本地索引的媒体,按时间戳查找相关片段,并整理出有来源依据的内容,例如 B-roll 匹配结果、播客剪辑和素材日志。适合希望通过连接的 AI 智能体处理本地视频、音频、照片和文档的用户。
www.theloqua.ai
Loqua is a voice-first desktop productivity app for macOS and Windows that turns natural speech into polished text, edits selected content, translates dictation, and answers questions about user-selected screen content. It is designed for hands-free writing and supported voice workflows across desktop apps.
myminutes.ai
用于会议、讲座和录音的 AI 记录与转录工具