AI Audio API

AI 音频 API 为应用提供语音识别、文字转语音、声音克隆、音频分析与实时处理能力,方便开发者快速集成语音功能。

AI 基础设施

探索此分类集合

产品

Daily preview
Daily logo

Daily

AI Audio API

为 Web、移动端、原生、桌面和服务器应用添加实时通话、录制及 AI 代理工作流的 SDK 与基础设施

Speechmatics preview
Speechmatics logo

Speechmatics

AI Audio API

Speechmatics provides speech-to-text APIs for pre-recorded and real-time audio, with support for multilingual speech, multiple speakers, and voice-agent workflows. It is designed for teams building transcription and Voice AI applications.

Loudly preview
Loudly logo

Loudly

AI音乐生成器

面向创作者、制作人和开发者的一站式 AI 音乐平台

ModelsLab preview
ModelsLab logo

ModelsLab

AI Audio API

ModelsLab is a developer platform that provides APIs for image, video, audio, 3D, and LLM generation through a unified service. It supports teams building generative media features without integrating each model provider separately.

AudioStack logo
AudioStack logo

AudioStack

AI Audio API

面向媒体与广告的 AI 音频制作平台,将需求或原始内容转为可播音频

SpeechifyAI preview
SpeechifyAI logo

SpeechifyAI

AI Audio API

SpeechifyAI is a developer API for expressive text-to-speech and voice cloning. It provides streaming Simba models, catalog and cloned voices, SSML support, and a free starting tier for building speech into applications.

Voicemaker preview
Voicemaker logo

Voicemaker

AI Voice Generator

基于浏览器的文本转语音平台,支持多语言 AI 语音、编辑、下载和 API 自动化。

Apiframe preview
Apiframe logo

Apiframe

AI Audio API

Apiframe is a unified REST API for generating AI images, videos, and music. It helps developers and automation teams add media generation through one API, with asynchronous jobs, webhooks, SDKs, and CDN-hosted outputs.

API in One preview
API in One logo

API in One

AI Audio API

API in One is a unified AI API gateway for website owners and developers who need image, video, music, speech, chat, and AI tool capabilities through one API key and shared credit balance.

4ALL API preview
4ALL API logo

4ALL API

AI Audio API

4ALL API is an API aggregation gateway for enterprises and developers that provides one access layer for text, image, video, and audio models from multiple providers. It supports OpenAI-compatible access, usage-based billing, model routing, and failover workflows.

Recall.ai Startup Program preview
Recall.ai Startup Program logo

Recall.ai Startup Program

AI Audio API

A startup program for early-stage companies building products powered by meeting and conversation data. Approved applicants receive discounted recording usage, access to Recall.ai products, and support while they build and launch.

Scaleway Generative APIs preview
Scaleway Generative APIs logo

Scaleway Generative APIs

AI Audio API

Scaleway Generative APIs provide OpenAI-compatible, serverless access to chat, code, vision, embedding, and audio models. They are designed for developers building AI applications without managing model-serving hardware, with endpoints hosted in European data centers and usage billed by tokens or audio minutes.

Gemini 3.8 text-to-speech preview
Gemini 3.8 text-to-speech logo

Gemini 3.8 text-to-speech

AI Audio API

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are audio-generation models for creating expressive voices, directing spoken performances, and producing conversational audio. They support creative teams, developers, and enterprises working across Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.

Runway Dev preview
Runway Dev logo

Runway Dev

AI Audio API

Runway Dev is an API platform for adding AI-generated video, images, and audio to products and production workflows. It provides access to Runway and other providers' models, along with model routing, workflows, recipes, and team controls.

Stability AI Developer Platform preview
Stability AI Developer Platform logo

Stability AI Developer Platform

AI 3D模型生成器

Stability AI Developer Platform provides APIs for adding generative audio, image, image editing, upscaling, control, and 3D asset creation to applications. It is intended for developers and teams building creative workflows with Stability AI models.