Text-to-speech and voice cloning
Generate speech with low-latency TTS and voice cloning. The pricing and TTS pages reference Realtime TTS-2, Realtime TTS 1.5 Max, and Realtime TTS 1.5 Mini, with multilingual support and voice design features.
Inworld AI is a voice AI platform for developers building realtime speech workflows. It combines TTS, speech-to-text, voice profiling, and LLM routing through one API with usage-based pricing.
Inworld AI is a voice AI platform that combines text-to-speech, speech-to-text, realtime voice agents, and LLM routing through a single product surface. The site emphasizes low-latency speech generation, streaming transcription, and a unified API for building expressive voice experiences.
Its pages position the platform for developers who need realtime audio workflows at scale: generate speech, transcribe user audio, extract voice profile signals, and pass those signals into downstream LLM and TTS behavior. The pricing page also shows usage-based plans, from on-demand evaluation through enterprise terms.
Generate speech with low-latency TTS and voice cloning. The pricing and TTS pages reference Realtime TTS-2, Realtime TTS 1.5 Max, and Realtime TTS 1.5 Mini, with multilingual support and voice design features.
Transcribe live audio streams or complete files through one API. The STT page supports bidirectional WebSocket streaming and synchronous transcription for uploaded audio.
Extract paralinguistic signals from speech, including emotion, vocal style, accent, age, and pitch. The product pages say those signals can feed downstream LLM and TTS behavior.
Route requests across 220+ AI models through a single API. The router page positions this as one interface for model selection rather than separate integrations per provider.
Use usage-based pricing with monthly credits, volume discounts, and higher-tier limits. The pricing page shows On-Demand, Creator, Builder, Developer, Growth, and Enterprise options.
Access workspace features such as sharing, team management, concurrency limits, and enterprise terms. Higher tiers add support and compliance-related options like SLA/DPA, data residency, and on-prem deployment.
Build conversational agents that listen to users, extract speech signals, and respond with synthesized speech in one loop. The STT and TTS pages describe a workflow where voice profile metadata can steer the response.
Add transcription to products that need both live captions and file-based speech recognition. The STT page supports bidirectional streaming for live audio and synchronous requests for complete audio files.
Generate branded or character voices with voice cloning, voice design, and multilingual speech output. The pricing page and homepage both highlight speech generation as a central product area.
Use one routing layer to select from a large model catalog without wiring each provider separately. The router page frames this as a single API for 220+ AI models.
Move from prototyping to production with tiered limits, workspace sharing, and enterprise controls. The pricing page shows plans for individual, team, growth, and custom deployment needs.
The pricing page indicates credits are used across Inworld products, and paid plans bundle monthly credits with higher usage limits and discounts. On-demand access is available for evaluation and prototyping, and custom enterprise pricing is available for larger deployments.
The source shows a unified API across text-to-speech, speech-to-text, realtime voice agents, and LLM routing. The STT page also shows both realtime streaming and synchronous transcription endpoints, plus voice profile signals on streaming chunks.
The pricing page lists team-oriented workspace features such as workspace creation and sharing, team management, and increasing concurrency limits on higher tiers. The source also mentions priority email support on the Developer plan and a dedicated AM and Slack channel on Enterprise.
The STT page shows voice profile signals for emotion, vocal style, accent, age, and pitch. The pricing page also indicates multilingual support in TTS-2 and higher-tier access to features like professional voice cloning and compliance add-ons.
The source supports voice AI building blocks rather than a no-code workflow. It provides API access, pricing by usage, and multiple endpoints, but the available pages do not describe visual app-building or full product-creation workflows.