Serverless model inference
Call a selection of generative AI models through APIs without directly operating the inference infrastructure. The catalogue includes Llama, Qwen, OSS GPT, and more than 20 models overall.
OVHcloud AI Endpoints provides serverless inference APIs for a selection of generative AI models, including LLMs, voice, document, and image analysis models. It helps developers add AI capabilities to applications through OpenAI-compatible APIs and supported integrations.
AI Endpoints is OVHcloud’s serverless inference API for a catalogue of more than 20 generative AI models. It gives developers API access to models such as Llama, Qwen, and OSS GPT, along with models for LLM and embedding workloads, voice processing, document analysis, and image analysis.
The service is intended for adding AI capabilities to applications without directly managing model inference infrastructure. Its APIs follow the OpenAI-compatible format, and the product page provides documentation and code examples for integration. A playground is also available for testing models in a sandbox or through the API.
AI Endpoints offers two processing modes: Base for continuous workloads and Batch, listed as beta, for bulk or deferred processing. OVHcloud states that data is not used to train or improve its AI models and describes zero data retention beyond data required for billing. Models can be deployed on a user’s infrastructure or used with other cloud services, supporting portability.
Call a selection of generative AI models through APIs without directly operating the inference infrastructure. The catalogue includes Llama, Qwen, OSS GPT, and more than 20 models overall.
Use Base for continuous workloads and typical interactive response patterns, or use Batch, currently in beta, for bulk or deferred jobs such as classification and mass labelling.
Use a standard API format supported by popular AI tooling. The product page lists LLMs and embeddings for Base, with one prompt per API call for general models and 1–25 prompts for embeddings.
OVHcloud states that customer data is not used to train or improve its AI models and that data retention is limited to information required for billing.
Transparent model version management is provided to support reproducibility as the catalogue is updated.
Test models in the playground or through the API, and connect with listed tools and frameworks including Hugging Face, LiteLLM, Pydantic AI, Continue, Apache Airflow, LlamaIndex, LiveKit, and Docker Agent.
Add real-time conversational interactions to applications through chatbots and other interfaces intended for customer engagement, customer-service automation, or personalised user experiences.
Use voice-to-text for customer-service records, meetings, or closed captions, and support text-to-speech experiences where audio accessibility or customisation is required.
Connect coding-assistance plugins such as Continue to development environments. The listed workflow supports code suggestions, error detection, and task automation while keeping data confidentiality in focus.
Use Batch processing for deferred or high-volume LLM jobs, including classification and mass labelling of data rather than interactive, one-request-at-a-time workloads.
Build context-aware agents and orchestrated workflows with listed ecosystem tools such as LlamaIndex, Pydantic AI, Apache Airflow, Mastra, LiveKit, and Docker Agent.
It is a serverless inference API that provides access to a selection of generative AI models, including models for language, embeddings, voice processing, document analysis, and image analysis.
Base is intended for continuous workloads and interactive use. Batch, which the product page labels as beta, is intended for bulk or deferred workloads such as classification and mass labelling.
Yes. The product page identifies both Base and Batch as OpenAI-compatible, making the stated API format suitable for tools built around that convention.
OVHcloud states that customer data will not be used to train or improve its AI models. It also describes zero data retention, retaining only data required for billing purposes.
The product page states that models can be deployed on user infrastructure or integrated with other cloud services. This is presented as a portability and reversibility option, but the available source does not specify the migration process.
www.cloudflare.com
Cloudflare Workers AI is a serverless AI inference platform for running language, image, speech-to-text, and embedding models through Cloudflare’s network. It helps developers build AI features without managing GPU infrastructure or capacity planning.
aimlapi.com
AIMLAPI provides one API and billing key for accessing a catalog of AI models for chat, reasoning, image, video, audio, voice, search, embeddings, code, and related tasks. It is intended for developers and teams that want to compare and use models from multiple providers through a common platform.
cloud.sambanova.ai
SambaNova Cloud is an AI inference platform that provides API access to open-source language and vision models. Developers can use its OpenAI-compatible API, playground, and model catalog to build and test AI-powered applications.
aws.amazon.com
Amazon Bedrock is a fully managed AWS platform for building generative AI applications and agents with access to foundation models, customization tools, safety controls, and production-oriented workflows. It supports teams that want to experiment through the console or build applications through AWS APIs and SDKs.
www.byteplus.com
ModelArk is BytePlus's one-stop large language model service platform for organizations building, deploying, and scaling AI applications. It is positioned within BytePlus's broader AI-native cloud portfolio.
platform.deepseek.com
DeepSeek Platform provides access to DeepSeek AI models, API documentation, and developer resources through an online API platform. It is intended for users who want to sign up or log in to use DeepSeek’s platform services.