Scaleway Generative APIs logo

Scaleway Generative APIs

Freemium
Visit

Scaleway Generative APIs provide OpenAI-compatible, serverless access to chat, code, vision, embedding, and audio models. They are designed for developers building AI applications without managing model-serving hardware, with endpoints hosted in European data centers and usage billed by tokens or audio minutes.

What is Scaleway Generative APIs?

Scaleway Generative APIs are a serverless model-as-a-service offering for adding generative AI to applications without configuring hardware or deploying models. They provide API access to selected chat, code, vision, embedding, and audio models through endpoints hosted in European data centers.

The service uses usage-based pricing, generally calculated per million tokens. Audio transcription is priced by audio minute for the Whisper Large V3 model. New customers receive 1,000,000 free tokens and begin paying from the 1,000,001st token. For non-realtime workloads, the product page says Batch APIs offer an additional discount.

The APIs are intended for developers who want to test and deploy AI features within existing applications. OpenAI-compatible interfaces, support for OpenAI libraries and LangChain, structured outputs, and function calling reduce the amount of integration work needed for common workflows.

What can Scaleway Generative APIs do?

Serverless model endpoints

Access a selection of managed AI models through API endpoints without configuring serving hardware or deploying models yourself. The catalog includes models for chat, code, vision, embeddings, and audio transcription.

OpenAI-compatible integration

The APIs are designed to work with existing OpenAI libraries and LangChain SDK workflows, giving developers a familiar way to connect applications to supported models.

Structured outputs

Built-in JSON mode and JSON Schema support can turn unstructured model responses into machine-readable data for application workflows that require predictable output formats.

Native function calling

Models can connect to external tools through Scaleway Serverless Functions, allowing applications to invoke custom functions or APIs as part of an AI workflow.

Usage-based and batch pricing

Standard usage is billed by the million tokens, with audio transcription priced by audio minute for the listed Whisper model. Batch APIs provide an additional discount for non-realtime use cases.

European hosting and streaming

The inference stack runs on infrastructure in Europe. The product page states that European end users can receive the first streamed tokens in under 200 ms, subject to the selected model and workload.

Use Cases

“Retrieval-augmented generation”

Combine embeddings, a vector database, and LangChain with a Generative API model to retrieve enterprise information and provide responses grounded in private or current data.

“Assistants and chatbots”

Build conversational or multimodal assistants for tasks such as question answering, translation, summarization, and sentiment analysis. Vision-language models can extend these assistants to image-based inputs.

“Autonomous agents”

Create agents that interpret a request, break it into steps, and call APIs, databases, or Serverless Functions to complete workflows such as customer support or booking processes.

“Document and diagram processing”

Use vision-language models for scanned documents, technical diagrams, and other mixed text-and-visual content where conventional OCR alone may not provide enough context.

“Call and video analysis”

Use the listed audio transcription capability to analyze call or video recordings, then combine transcribed content with language models to identify needs or derive operational insights.

Frequently Asked Questions

What are Scaleway Generative APIs?

They are serverless API endpoints that provide access to selected generative AI models for tasks including chat, code generation, vision, embeddings, and audio transcription. Developers can use them without managing model-serving hardware.

How are Generative APIs priced?

The product uses usage-based pricing, generally billed per million input and output tokens. The listed Whisper Large V3 model is priced per audio minute. The product page also states that Batch APIs offer an additional discount for non-realtime workloads.

Is there a free allowance?

Yes. The product page states that every new customer receives 1,000,000 free tokens and starts paying from the 1,000,001st token.

Can existing OpenAI or LangChain integrations be used?

The APIs are described as OpenAI-compatible and are designed to work with existing OpenAI libraries and LangChain SDKs. Exact setup steps, authentication requirements, and model-specific limitations are provided in the product documentation rather than the supplied product page text.

When should I consider dedicated deployment instead?

Generative APIs are the serverless option for quickly serving curated models with usage-based pricing. Scaleway's separate Dedicated Deployment offering is intended for workloads requiring dedicated infrastructure, guaranteed throughput, predictable hourly pricing, or support for custom models.

Quick Facts

Provider
Scaleway
Category
Serverless generative AI APIs
Supported task types
Chat, code, vision, embeddings, and audio transcription
API compatibility
OpenAI-compatible; supports OpenAI libraries and LangChain SDK workflows
Pricing model
Usage-based, generally priced per 1 million tokens; Whisper transcription is priced per audio minute
Free allowance
1,000,000 free tokens for every new customer

Scaleway Generative APIs Alternatives

AIMLAPI logo

AIMLAPI

aimlapi.com

AIMLAPI provides one API and billing key for accessing a catalog of AI models for chat, reasoning, image, video, audio, voice, search, embeddings, code, and related tasks. It is intended for developers and teams that want to compare and use models from multiple providers through a common platform.

SambaNova Cloud logo

SambaNova Cloud

cloud.sambanova.ai

SambaNova Cloud is an AI inference platform that provides API access to open-source language and vision models. Developers can use its OpenAI-compatible API, playground, and model catalog to build and test AI-powered applications.

Amazon Bedrock logo

Amazon Bedrock

aws.amazon.com

Amazon Bedrock is a fully managed AWS platform for building generative AI applications and agents with access to foundation models, customization tools, safety controls, and production-oriented workflows. It supports teams that want to experiment through the console or build applications through AWS APIs and SDKs.

AI Endpoints logo

AI Endpoints

www.ovhcloud.com

OVHcloud AI Endpoints provides serverless inference APIs for a selection of generative AI models, including LLMs, voice, document, and image analysis models. It helps developers add AI capabilities to applications through OpenAI-compatible APIs and supported integrations.

ModelArk logo

ModelArk

www.byteplus.com

ModelArk is BytePlus's one-stop large language model service platform for organizations building, deploying, and scaling AI applications. It is positioned within BytePlus's broader AI-native cloud portfolio.

DeepSeek Platform logo

DeepSeek Platform

platform.deepseek.com

DeepSeek Platform provides access to DeepSeek AI models, API documentation, and developer resources through an online API platform. It is intended for users who want to sign up or log in to use DeepSeek’s platform services.