Serverless model endpoints
Access a selection of managed AI models through API endpoints without configuring serving hardware or deploying models yourself. The catalog includes models for chat, code, vision, embeddings, and audio transcription.
Scaleway Generative APIs provide OpenAI-compatible, serverless access to chat, code, vision, embedding, and audio models. They are designed for developers building AI applications without managing model-serving hardware, with endpoints hosted in European data centers and usage billed by tokens or audio minutes.
Scaleway Generative APIs are a serverless model-as-a-service offering for adding generative AI to applications without configuring hardware or deploying models. They provide API access to selected chat, code, vision, embedding, and audio models through endpoints hosted in European data centers.
The service uses usage-based pricing, generally calculated per million tokens. Audio transcription is priced by audio minute for the Whisper Large V3 model. New customers receive 1,000,000 free tokens and begin paying from the 1,000,001st token. For non-realtime workloads, the product page says Batch APIs offer an additional discount.
The APIs are intended for developers who want to test and deploy AI features within existing applications. OpenAI-compatible interfaces, support for OpenAI libraries and LangChain, structured outputs, and function calling reduce the amount of integration work needed for common workflows.
Access a selection of managed AI models through API endpoints without configuring serving hardware or deploying models yourself. The catalog includes models for chat, code, vision, embeddings, and audio transcription.
The APIs are designed to work with existing OpenAI libraries and LangChain SDK workflows, giving developers a familiar way to connect applications to supported models.
Built-in JSON mode and JSON Schema support can turn unstructured model responses into machine-readable data for application workflows that require predictable output formats.
Models can connect to external tools through Scaleway Serverless Functions, allowing applications to invoke custom functions or APIs as part of an AI workflow.
Standard usage is billed by the million tokens, with audio transcription priced by audio minute for the listed Whisper model. Batch APIs provide an additional discount for non-realtime use cases.
The inference stack runs on infrastructure in Europe. The product page states that European end users can receive the first streamed tokens in under 200 ms, subject to the selected model and workload.
Combine embeddings, a vector database, and LangChain with a Generative API model to retrieve enterprise information and provide responses grounded in private or current data.
Build conversational or multimodal assistants for tasks such as question answering, translation, summarization, and sentiment analysis. Vision-language models can extend these assistants to image-based inputs.
Create agents that interpret a request, break it into steps, and call APIs, databases, or Serverless Functions to complete workflows such as customer support or booking processes.
Use vision-language models for scanned documents, technical diagrams, and other mixed text-and-visual content where conventional OCR alone may not provide enough context.
Use the listed audio transcription capability to analyze call or video recordings, then combine transcribed content with language models to identify needs or derive operational insights.
They are serverless API endpoints that provide access to selected generative AI models for tasks including chat, code generation, vision, embeddings, and audio transcription. Developers can use them without managing model-serving hardware.
The product uses usage-based pricing, generally billed per million input and output tokens. The listed Whisper Large V3 model is priced per audio minute. The product page also states that Batch APIs offer an additional discount for non-realtime workloads.
Yes. The product page states that every new customer receives 1,000,000 free tokens and starts paying from the 1,000,001st token.
The APIs are described as OpenAI-compatible and are designed to work with existing OpenAI libraries and LangChain SDKs. Exact setup steps, authentication requirements, and model-specific limitations are provided in the product documentation rather than the supplied product page text.
Generative APIs are the serverless option for quickly serving curated models with usage-based pricing. Scaleway's separate Dedicated Deployment offering is intended for workloads requiring dedicated infrastructure, guaranteed throughput, predictable hourly pricing, or support for custom models.
aimlapi.com
AIMLAPI provides one API and billing key for accessing a catalog of AI models for chat, reasoning, image, video, audio, voice, search, embeddings, code, and related tasks. It is intended for developers and teams that want to compare and use models from multiple providers through a common platform.
cloud.sambanova.ai
SambaNova Cloud is an AI inference platform that provides API access to open-source language and vision models. Developers can use its OpenAI-compatible API, playground, and model catalog to build and test AI-powered applications.
aws.amazon.com
Amazon Bedrock is a fully managed AWS platform for building generative AI applications and agents with access to foundation models, customization tools, safety controls, and production-oriented workflows. It supports teams that want to experiment through the console or build applications through AWS APIs and SDKs.
www.ovhcloud.com
OVHcloud AI Endpoints provides serverless inference APIs for a selection of generative AI models, including LLMs, voice, document, and image analysis models. It helps developers add AI capabilities to applications through OpenAI-compatible APIs and supported integrations.
www.byteplus.com
ModelArk is BytePlus's one-stop large language model service platform for organizations building, deploying, and scaling AI applications. It is positioned within BytePlus's broader AI-native cloud portfolio.
platform.deepseek.com
DeepSeek Platform provides access to DeepSeek AI models, API documentation, and developer resources through an online API platform. It is intended for users who want to sign up or log in to use DeepSeek’s platform services.