Pre-recorded transcription
Provides pre-recorded speech-to-text for recorded audio and video, with pricing based on hour of audio submitted.
AssemblyAI is a voice AI infrastructure platform for developers, offering speech-to-text, real-time transcription, voice agents, transcript understanding, guardrails, and an LLM Gateway.

AssemblyAI is a voice AI infrastructure platform for developers building products that transcribe, understand, and act on speech. Its APIs cover recorded transcription, live transcription, voice agents, transcript enrichment, content guardrails, and access to large language models in voice workflows.
The site presents the product as a set of production-grade APIs rather than a single app. Pricing and documentation are organized around specific workflows such as pre-recorded audio, streaming sessions, voice agents, and add-on features like speaker diarization, medical mode, and redaction.
Provides pre-recorded speech-to-text for recorded audio and video, with pricing based on hour of audio submitted.
Supports live transcription for calls, captions, voice agents, and agent assist, with billing based on how long the WebSocket session stays open.
Combines STT, LLM reasoning, TTS, turn detection, interruption handling, and tool calling in a single WebSocket for speech-to-speech agent workflows.
Adds structured transcript features such as speaker identification, translation, custom formatting, entity detection, sentiment analysis, topic detection, and summarization.
Offers PII redaction, profanity filtering, and content moderation as guardrails on transcript and audio workflows.
Provides an LLM Gateway with per-token pricing and prompt caching support on Anthropic, OpenAI, and Google models.
Turn recorded interviews, meetings, podcasts, or other audio/video files into transcripts using the pre-recorded API, then layer on speaker identification, translation, or redaction when needed.
Transcribe live conversations for call experiences, captions, voice agents, or agent-assist workflows using the real-time API, where billing follows session time.
Build speech-to-speech agents with a single WebSocket by combining speech recognition, language model reasoning, turn handling, interruption handling, and TTS through the Voice Agent API.
Extract structured insights from transcripts with features such as entity detection, sentiment analysis, topic detection, summarization, and custom formatting.
Apply guardrails such as profanity filtering, PII audio redaction, PII text redaction, and content moderation to speech and transcript pipelines.
AssemblyAI provides APIs for pre-recorded speech-to-text, real-time transcription, Voice Agent API, Speech Understanding, Guardrails, and LLM Gateway. The site positions it for developers building products that transcribe, understand, and act on speech.
The homepage describes the platform as built for developers and teams that need production-grade APIs, scalable speech infrastructure, and developer-friendly workflows for audio and speech applications.
The pricing page states that pre-recorded transcription is billed per hour of audio submitted, streaming is billed by session duration, and the Voice Agent API is billed at $4.50 per hour, or $0.075 per minute.
Yes. The pricing page says $50 in free credits are available with no credit card required, and it notes that free-tier and pay-as-you-go accounts differ on concurrency.
The pricing page says the EU region uses the same price as the US and can be accessed through `api.eu.assemblyai.com` and `streaming.eu.assemblyai.com` for GDPR-compliant data residency at no premium.