AI Endpoints logo

AI Endpoints

Freemium
Visit

OVHcloud AI Endpoints provides serverless inference APIs for a selection of generative AI models, including LLMs, voice, document, and image analysis models. It helps developers add AI capabilities to applications through OpenAI-compatible APIs and supported integrations.

What is AI Endpoints?

AI Endpoints is OVHcloud’s serverless inference API for a catalogue of more than 20 generative AI models. It gives developers API access to models such as Llama, Qwen, and OSS GPT, along with models for LLM and embedding workloads, voice processing, document analysis, and image analysis.

The service is intended for adding AI capabilities to applications without directly managing model inference infrastructure. Its APIs follow the OpenAI-compatible format, and the product page provides documentation and code examples for integration. A playground is also available for testing models in a sandbox or through the API.

AI Endpoints offers two processing modes: Base for continuous workloads and Batch, listed as beta, for bulk or deferred processing. OVHcloud states that data is not used to train or improve its AI models and describes zero data retention beyond data required for billing. Models can be deployed on a user’s infrastructure or used with other cloud services, supporting portability.

What can AI Endpoints do?

Serverless model inference

Call a selection of generative AI models through APIs without directly operating the inference infrastructure. The catalogue includes Llama, Qwen, OSS GPT, and more than 20 models overall.

Base and Batch processing

Use Base for continuous workloads and typical interactive response patterns, or use Batch, currently in beta, for bulk or deferred jobs such as classification and mass labelling.

OpenAI-compatible APIs

Use a standard API format supported by popular AI tooling. The product page lists LLMs and embeddings for Base, with one prompt per API call for general models and 1–25 prompts for embeddings.

Privacy-focused data handling

OVHcloud states that customer data is not used to train or improve its AI models and that data retention is limited to information required for billing.

Model lifecycle management

Transparent model version management is provided to support reproducibility as the catalogue is updated.

Testing and ecosystem integrations

Test models in the playground or through the API, and connect with listed tools and frameworks including Hugging Face, LiteLLM, Pydantic AI, Continue, Apache Airflow, LlamaIndex, LiveKit, and Docker Agent.

Use Cases

“Conversational applications”

Add real-time conversational interactions to applications through chatbots and other interfaces intended for customer engagement, customer-service automation, or personalised user experiences.

“Voice transcription and speech interactions”

Use voice-to-text for customer-service records, meetings, or closed captions, and support text-to-speech experiences where audio accessibility or customisation is required.

“Private coding assistants”

Connect coding-assistance plugins such as Continue to development environments. The listed workflow supports code suggestions, error detection, and task automation while keeping data confidentiality in focus.

“Bulk data classification”

Use Batch processing for deferred or high-volume LLM jobs, including classification and mass labelling of data rather than interactive, one-request-at-a-time workloads.

“AI agent and workflow development”

Build context-aware agents and orchestrated workflows with listed ecosystem tools such as LlamaIndex, Pydantic AI, Apache Airflow, Mastra, LiveKit, and Docker Agent.

Frequently Asked Questions

What is OVHcloud AI Endpoints?

It is a serverless inference API that provides access to a selection of generative AI models, including models for language, embeddings, voice processing, document analysis, and image analysis.

What is the difference between Base and Batch?

Base is intended for continuous workloads and interactive use. Batch, which the product page labels as beta, is intended for bulk or deferred workloads such as classification and mass labelling.

Are the APIs compatible with OpenAI-format tools?

Yes. The product page identifies both Base and Batch as OpenAI-compatible, making the stated API format suitable for tools built around that convention.

Is customer data used to train the models?

OVHcloud states that customer data will not be used to train or improve its AI models. It also describes zero data retention, retaining only data required for billing purposes.

Can models be moved to another environment?

The product page states that models can be deployed on user infrastructure or integrated with other cloud services. This is presented as a portability and reversibility option, but the available source does not specify the migration process.

Quick Facts

Provider
OVHcloud
Product type
Serverless inference API
Model catalogue
More than 20 models, including Llama, Qwen, and OSS GPT
API format
OpenAI-compatible
Processing modes
Base and Batch (beta)
Availability
Listed in the Gravelines Public Cloud region

AI Endpoints Alternatives

Cloudflare Workers AI logo

Cloudflare Workers AI

www.cloudflare.com

Cloudflare Workers AI is a serverless AI inference platform for running language, image, speech-to-text, and embedding models through Cloudflare’s network. It helps developers build AI features without managing GPU infrastructure or capacity planning.

AIMLAPI logo

AIMLAPI

aimlapi.com

AIMLAPI provides one API and billing key for accessing a catalog of AI models for chat, reasoning, image, video, audio, voice, search, embeddings, code, and related tasks. It is intended for developers and teams that want to compare and use models from multiple providers through a common platform.

SambaNova Cloud logo

SambaNova Cloud

cloud.sambanova.ai

SambaNova Cloud is an AI inference platform that provides API access to open-source language and vision models. Developers can use its OpenAI-compatible API, playground, and model catalog to build and test AI-powered applications.

Amazon Bedrock logo

Amazon Bedrock

aws.amazon.com

Amazon Bedrock is a fully managed AWS platform for building generative AI applications and agents with access to foundation models, customization tools, safety controls, and production-oriented workflows. It supports teams that want to experiment through the console or build applications through AWS APIs and SDKs.

ModelArk logo

ModelArk

www.byteplus.com

ModelArk is BytePlus's one-stop large language model service platform for organizations building, deploying, and scaling AI applications. It is positioned within BytePlus's broader AI-native cloud portfolio.

DeepSeek Platform logo

DeepSeek Platform

platform.deepseek.com

DeepSeek Platform provides access to DeepSeek AI models, API documentation, and developer resources through an online API platform. It is intended for users who want to sign up or log in to use DeepSeek’s platform services.