Streaming text-to-speech
Generate real-time speech with Sonic-3.5, the company’s streaming TTS model, designed for low-latency conversational use instead of delayed batch rendering.
Cartesia is a voice AI platform for real-time text-to-speech, speech-to-text, and voice agents with low latency and flexible deployment options.
Cartesia is a voice AI platform focused on real-time speech generation, transcription, and voice agents. Its main products are Sonic for text-to-speech, Ink for speech-to-text, and Line for building voice agents.
The company positions these models for live, synchronous interactions where latency matters, including AI agents, customer support, sales, recruiting, and other conversational workflows. The site also says the same stack can be deployed in the cloud, on-premise, in a VPC, or on-device.
Generate real-time speech with Sonic-3.5, the company’s streaming TTS model, designed for low-latency conversational use instead of delayed batch rendering.
Create or clone voices, localize audio into multiple languages, and use custom pronunciation dictionaries for names and domain terms that need precise delivery.
Transcribe speech with Ink-2, the speech-to-text model positioned alongside Sonic for interactive voice systems and live conversation flows.
Build voice agents with Line using tools, guardrails, a customizable LLM, and built-in evaluations for testing live behavior and call analytics.
Deploy the same models and agents across cloud, on-premise, VPC, and on-device environments to match latency, residency, and control requirements.
Use API access together with SDKs and developer tools to move from copy-paste examples into production systems with iterative development loops.
Use Sonic to generate natural-sounding agent responses in live conversations where timing and pacing affect the interaction.
Use Line to build a voice agent that can take actions, use tools, and be tested with built-in evaluations before deployment.
Use Sonic’s localization and voice cloning features to keep a consistent brand voice across languages and markets.
Use Ink to transcribe spoken input for call flows, live support workflows, or other systems that need fast speech recognition.
Use the platform’s deployment options to run in cloud, on-premise, VPC, or on-device settings when residency or control matters.
Yes. The pricing page shows Free, Pro, Startup, Scale, and Enterprise options, with a contact-sales flow for Enterprise. It also presents usage-based credits and agent minutes, so teams can start small and move up as usage grows.
Cartesia’s site presents Sonic for text-to-speech, Ink for speech-to-text, and Line for voice agents. The home and agents pages indicate these are exposed through APIs and developer tooling, with Line aimed at building and deploying production voice agents.
The site says Cartesia supports cloud deployment, regional API endpoints, on-premise or VPC deployment, and on-device deployment. It also says the same models and agents can run across these environments.
The pricing page says plans include unlimited workspace seats and voice slots on every plan, while usage is metered through credits and agent minutes. The site also references built-in evaluations and agent analytics for voice agents.
The source materials do not provide a public setup wizard or a full integration list. They do show API access, SDKs, and developer tools, but not a complete catalog of supported languages or third-party integrations.