Unreal Speech logo

Unreal Speech

소유권 인증

Unreal Speech is a text-to-speech API for developers needing fast speech generation, timestamp support, and multiple API workflows. Free browser demo included.

Unreal Speech preview

What Unreal Speech is

Unreal Speech is a text-to-speech API built for developers who need to convert written text into spoken audio with low latency and flexible output options. The site positions it as a production-ready service with streaming synthesis, timestamp support, and a browser-based demo for quick testing.

The product site shows several ways to use the platform: a fast `/stream` endpoint, a `/speech` endpoint that can return MP3 and timestamp URLs, an asynchronous `/synthesisTasks` flow for longer text, and a WebSocket endpoint for streaming audio together with word-level timestamps. The studio page adds a free online demo built on Kokoro TTS, with 48 voices across 8 languages for browser-based generation and download.

Core capabilities

Fast streaming synthesis

The homepage and docs show a fast streaming path for text-to-speech output, with audio available in about 300ms for the `/stream` endpoint.

Multiple API workflows

The API supports multiple output patterns: synchronous, asynchronous, and WebSocket-based streaming, so teams can choose the workflow that fits text length and latency needs.

Timestamp generation

The site documents word- or sentence-level timestamps, including a `/streamWithTimestamps` WebSocket flow for real-time timing data.

Voice and language selection

The studio demo presents 48 voices across 8 languages, giving users a broad set of voice and language options to test.

Long-form and scaled usage

The pricing page says the service can generate up to 10-hour audio and offers plans that scale by character volume, including a free tier and higher-volume paid plans.

Live browser demo

The browser studio lets visitors type text, generate speech, and download recordings directly in the browser for quick evaluation.

Common ways to use Unreal Speech

  • Fast interactive playback

    Use the low-latency `/stream` endpoint when you want spoken output to start quickly, such as interactive app responses or realtime narration previews.

  • Batch audio generation

    Use `/speech` or `/synthesisTasks` when you need generated audio files for shorter clips or longer passages, including longer-form content that exceeds the synchronous path.

  • Timed text highlighting

    Use word- or sentence-level timestamps when your app needs highlighting, caption sync, or narration playback that stays aligned with text.

  • Voice evaluation and prototyping

    Use the studio demo to test voices in the browser, compare language options, and download sample recordings before integrating the API.

  • Scaling usage over time

    Use the pricing tiers to match character volume to usage patterns, starting with the free tier and moving to higher-volume plans as demand grows.

Pros and Cons

Pros

  • Offers a fast streaming path with audio available in about 300ms on the `/stream` endpoint.
  • Supports both short-form and long-form generation through synchronous and asynchronous API flows.
  • Includes word- and sentence-level timestamp options, plus a WebSocket stream that can deliver audio and timing data together.
  • Provides a free browser demo for trying voices and generating downloadable audio without leaving the site.
  • Shows a pricing model that starts with a free tier and scales to higher-volume plans.

Cons

  • The site does not provide a complete integration list, such as SDK coverage beyond the Python examples shown in the docs snippets.
  • Some capability details are only partially documented on the pages provided, so fit and implementation effort may require checking the API docs directly.
  • The browser studio is a demo environment and the site directs commercial use cases to the API documentation rather than the demo itself.

FAQ

Does Unreal Speech offer a free way to try the product?

Yes. The pricing page lists a free tier and paid plans, and the site also offers a live studio demo for testing text-to-speech generation in the browser.

What API workflows does it support?

The docs shown on the homepage include `/stream`, `/speech`, `/synthesisTasks`, and `/streamWithTimestamps`. The site says `/stream` is synchronous and fast, `/speech` returns MP3 plus timestamp URLs, `/synthesisTasks` is asynchronous for longer text, and `/streamWithTimestamps` streams audio with word-level timing.

Can it generate timestamps with the audio?

The site says timestamps can be requested with `TimestampType` set to `word` or `sentence`, and it also mentions a WebSocket flow for streaming both audio and timestamps.

How many voices and languages are available?

The studio page says Kokoro TTS Studio supports 48 voices across 8 languages, including American and British English, French, Hindi, Spanish, Japanese, Chinese, Italian, and Portuguese.

Is the browser demo the same as the API product?

The site positions Unreal Speech for commercial use through its API documentation, while the studio page describes the browser demo as a free online text-to-speech experience for trying the model and downloading audio.

Quick Facts

Category
Text-to-speech API
Source domain
unrealspeech.com
Primary users
Developers
Demo
Browser-based studio for Kokoro TTS
Voices and languages
48 voices across 8 languages
Pricing
Free tier plus paid plans