OpenAI-compatible API
Build against a familiar chat-completions pattern using the `https://api.sambanova.ai/v1` base URL and bearer-key authentication.
SambaNova Cloud is an AI inference platform that provides API access to open-source language and vision models. Developers can use its OpenAI-compatible API, playground, and model catalog to build and test AI-powered applications.
SambaNova Cloud is an AI inference platform for developers building applications with hosted open-source models. Its core workflow is an OpenAI-compatible API that accepts chat-completion requests at SambaNova’s API endpoint. Users create a bearer-authenticated API key, keep it outside application code, and send requests from tools such as cURL or Python.
The service also includes a browser playground for trying prompts and viewing code, plus examples for connecting models to Gradio. Its documented catalog covers text, reasoning, and selected vision-capable models from DeepSeek, Meta, MiniMax, Google, and OpenAI model families. Published pricing is based on input and output token usage, with a cached-input price listed for MiniMax-M3.
Build against a familiar chat-completions pattern using the `https://api.sambanova.ai/v1` base URL and bearer-key authentication.
The catalog lists DeepSeek-V3.1, DeepSeek-V3.2, Meta-Llama-3.3-70B-Instruct, MiniMax-M3, gemma-4-31B-it, and gpt-oss-120b, with text, reasoning, and selected vision labels.
The cURL quickstart demonstrates setting `stream` to true for a chat-completion request, allowing applications to work with streamed responses.
The playground provides a model-selection and prompt-testing workflow, with examples such as drafting an email, planning a trip, and writing code.
The developer quickstart includes a cURL request and Python client example with message arrays, model selection, temperature, and top-p parameters.
The published pricing table separates cached input, input, and output token rates by model; MiniMax-M3 is the only listed model with a cached-input rate.
Use the playground to test prompts and model behavior before implementing the workflow in an application.
Use the OpenAI-compatible endpoint with cURL or Python to send system and user messages from an application backend.
Use the documented streaming request pattern as a starting point for interfaces that process chat-completion output incrementally.
Evaluate listed DeepSeek, Meta, MiniMax, Google, and OpenAI models against text, reasoning, or vision-oriented requirements shown in the catalog.
Use the provided Gradio example to load a model and launch an interactive demo after supplying an API token.
Create an API key in the SambaNova Cloud dashboard, store it in an environment variable such as `SAMBANOVA_API_KEY`, and send a request to `https://api.sambanova.ai/v1/chat/completions` with bearer authentication. The quickstart provides cURL and Python examples.
The documented requests use a bearer API key in the `Authorization` header. The site recommends keeping the key out of source code and reading it from the environment instead.
The site lists DeepSeek-V3.1, DeepSeek-V3.2, Meta-Llama-3.3-70B-Instruct, MiniMax-M3, gemma-4-31B-it, and gpt-oss-120b. The catalog identifies models by supported task labels such as text, reasoning, or vision.
Yes. SambaNova Cloud includes a playground with a model selector, prompt examples, and a View Code option. The site also provides a Gradio loading example for creating a simple interactive demo.
The published table shows cached-input, input-per-one-million-tokens, and output-per-one-million-tokens rates by model. It does not establish quotas, plan limits, or other billing terms in the supplied source.
aimlapi.com
AIMLAPI provides one API and billing key for accessing a catalog of AI models for chat, reasoning, image, video, audio, voice, search, embeddings, code, and related tasks. It is intended for developers and teams that want to compare and use models from multiple providers through a common platform.
aws.amazon.com
Amazon Bedrock is a fully managed AWS platform for building generative AI applications and agents with access to foundation models, customization tools, safety controls, and production-oriented workflows. It supports teams that want to experiment through the console or build applications through AWS APIs and SDKs.
www.ovhcloud.com
OVHcloud AI Endpoints provides serverless inference APIs for a selection of generative AI models, including LLMs, voice, document, and image analysis models. It helps developers add AI capabilities to applications through OpenAI-compatible APIs and supported integrations.
www.byteplus.com
ModelArk is BytePlus's one-stop large language model service platform for organizations building, deploying, and scaling AI applications. It is positioned within BytePlus's broader AI-native cloud portfolio.
platform.deepseek.com
DeepSeek Platform provides access to DeepSeek AI models, API documentation, and developer resources through an online API platform. It is intended for users who want to sign up or log in to use DeepSeek’s platform services.
parasail.io
Parasail is an inference cloud for AI-native startups that provides access to open and frontier models through an OpenAI-compatible API. It offers serverless, elastic, dedicated, and batch deployment options with per-token or GPU-based pricing.