AssemblyAI is a voice AI infrastructure platform for developers, offering speech-to-text, real-time transcription, voice agents, transcript understanding, guardrails, and an LLM Gateway.

AssemblyAI preview

What AssemblyAI is

AssemblyAI is a voice AI infrastructure platform for developers building products that transcribe, understand, and act on speech. Its APIs cover recorded transcription, live transcription, voice agents, transcript enrichment, content guardrails, and access to large language models in voice workflows.

The site presents the product as a set of production-grade APIs rather than a single app. Pricing and documentation are organized around specific workflows such as pre-recorded audio, streaming sessions, voice agents, and add-on features like speaker diarization, medical mode, and redaction.

Core capabilities

Pre-recorded transcription

Provides pre-recorded speech-to-text for recorded audio and video, with pricing based on hour of audio submitted.

Real-time transcription

Supports live transcription for calls, captions, voice agents, and agent assist, with billing based on how long the WebSocket session stays open.

Voice Agent API

Combines STT, LLM reasoning, TTS, turn detection, interruption handling, and tool calling in a single WebSocket for speech-to-speech agent workflows.

Speech understanding add-ons

Adds structured transcript features such as speaker identification, translation, custom formatting, entity detection, sentiment analysis, topic detection, and summarization.

Guardrails

Offers PII redaction, profanity filtering, and content moderation as guardrails on transcript and audio workflows.

LLM Gateway

Provides an LLM Gateway with per-token pricing and prompt caching support on Anthropic, OpenAI, and Google models.

Common use cases

  • Batch transcription for recorded media

    Turn recorded interviews, meetings, podcasts, or other audio/video files into transcripts using the pre-recorded API, then layer on speaker identification, translation, or redaction when needed.

  • Live audio processing

    Transcribe live conversations for call experiences, captions, voice agents, or agent-assist workflows using the real-time API, where billing follows session time.

  • Conversational voice agents

    Build speech-to-speech agents with a single WebSocket by combining speech recognition, language model reasoning, turn handling, interruption handling, and TTS through the Voice Agent API.

  • Transcript enrichment and analysis

    Extract structured insights from transcripts with features such as entity detection, sentiment analysis, topic detection, summarization, and custom formatting.

  • Safety and compliance workflows

    Apply guardrails such as profanity filtering, PII audio redaction, PII text redaction, and content moderation to speech and transcript pipelines.

Pros and Cons

Pros

  • Covers the full speech workflow from transcription to agent actions in one platform.
  • Supports both async and real-time use cases with separate models and clear billing rules.
  • Offers add-ons for transcript enrichment, redaction, and medical use cases.
  • Provides EU data residency at the same price as the US.
  • Includes $50 in free credits with no credit card required.

Cons

  • Pricing is usage-based and varies by workflow, model, and add-ons, so estimating cost requires checking the relevant rate card.
  • Streaming is billed by session duration rather than audio duration, which means idle connection time still counts.
  • Several transcript features are model-specific or deprecated, including Auto Chapters and Summarization on Universal-2 and certain streaming limitations.

FAQ

What does AssemblyAI do?

AssemblyAI provides APIs for pre-recorded speech-to-text, real-time transcription, Voice Agent API, Speech Understanding, Guardrails, and LLM Gateway. The site positions it for developers building products that transcribe, understand, and act on speech.

Who is AssemblyAI for?

The homepage describes the platform as built for developers and teams that need production-grade APIs, scalable speech infrastructure, and developer-friendly workflows for audio and speech applications.

How is AssemblyAI priced?

The pricing page states that pre-recorded transcription is billed per hour of audio submitted, streaming is billed by session duration, and the Voice Agent API is billed at $4.50 per hour, or $0.075 per minute.

Is there a free tier or trial?

Yes. The pricing page says $50 in free credits are available with no credit card required, and it notes that free-tier and pay-as-you-go accounts differ on concurrency.

Does AssemblyAI offer EU data residency?

The pricing page says the EU region uses the same price as the US and can be accessed through `api.eu.assemblyai.com` and `streaming.eu.assemblyai.com` for GDPR-compliant data residency at no premium.

Quick Facts

Category
Voice AI infrastructure
Primary users
Developers and product teams
Core workflows
Speech-to-text, real-time transcription, voice agents, transcript understanding
Source domain
assemblyai.com
Pricing model
Usage-based
Free credits
$50, no credit card required