SambaNova Cloud logo

SambaNova Cloud

Freemium
訪問

SambaNova Cloud is an AI inference platform that provides API access to open-source language and vision models. Developers can use its OpenAI-compatible API, playground, and model catalog to build and test AI-powered applications.

SambaNova Cloudとは?

SambaNova Cloud is an AI inference platform for developers building applications with hosted open-source models. Its core workflow is an OpenAI-compatible API that accepts chat-completion requests at SambaNova’s API endpoint. Users create a bearer-authenticated API key, keep it outside application code, and send requests from tools such as cURL or Python.

The service also includes a browser playground for trying prompts and viewing code, plus examples for connecting models to Gradio. Its documented catalog covers text, reasoning, and selected vision-capable models from DeepSeek, Meta, MiniMax, Google, and OpenAI model families. Published pricing is based on input and output token usage, with a cached-input price listed for MiniMax-M3.

SambaNova Cloudでできること

OpenAI-compatible API

Build against a familiar chat-completions pattern using the `https://api.sambanova.ai/v1` base URL and bearer-key authentication.

Multiple hosted models

The catalog lists DeepSeek-V3.1, DeepSeek-V3.2, Meta-Llama-3.3-70B-Instruct, MiniMax-M3, gemma-4-31B-it, and gpt-oss-120b, with text, reasoning, and selected vision labels.

Streaming requests

The cURL quickstart demonstrates setting `stream` to true for a chat-completion request, allowing applications to work with streamed responses.

Browser playground

The playground provides a model-selection and prompt-testing workflow, with examples such as drafting an email, planning a trip, and writing code.

Python and cURL examples

The developer quickstart includes a cURL request and Python client example with message arrays, model selection, temperature, and top-p parameters.

Usage-based pricing

The published pricing table separates cached input, input, and output token rates by model; MiniMax-M3 is the only listed model with a cached-input rate.

利用シーン

“Prototype an AI application”

Use the playground to test prompts and model behavior before implementing the workflow in an application.

“Add chat completions to a developer project”

Use the OpenAI-compatible endpoint with cURL or Python to send system and user messages from an application backend.

“Build streamed interfaces”

Use the documented streaming request pattern as a starting point for interfaces that process chat-completion output incrementally.

“Compare model families for a task”

Evaluate listed DeepSeek, Meta, MiniMax, Google, and OpenAI models against text, reasoning, or vision-oriented requirements shown in the catalog.

“Create a quick Gradio demo”

Use the provided Gradio example to load a model and launch an interactive demo after supplying an API token.

よくある質問

How do I make a first request?

Create an API key in the SambaNova Cloud dashboard, store it in an environment variable such as `SAMBANOVA_API_KEY`, and send a request to `https://api.sambanova.ai/v1/chat/completions` with bearer authentication. The quickstart provides cURL and Python examples.

Which authentication method does the API use?

The documented requests use a bearer API key in the `Authorization` header. The site recommends keeping the key out of source code and reading it from the environment instead.

Which models are available?

The site lists DeepSeek-V3.1, DeepSeek-V3.2, Meta-Llama-3.3-70B-Instruct, MiniMax-M3, gemma-4-31B-it, and gpt-oss-120b. The catalog identifies models by supported task labels such as text, reasoning, or vision.

Can I test models without writing a full application?

Yes. SambaNova Cloud includes a playground with a model selector, prompt examples, and a View Code option. The site also provides a Gradio loading example for creating a simple interactive demo.

How is pricing presented?

The published table shows cached-input, input-per-one-million-tokens, and output-per-one-million-tokens rates by model. It does not establish quotas, plan limits, or other billing terms in the supplied source.

クイック情報

Product type
AI inference platform
Primary interface
OpenAI-compatible API
API endpoint
https://api.sambanova.ai/v1/chat/completions
Authentication
Bearer API key
Model access
Hosted open-source models
Pricing basis
Per-token input and output rates

SambaNova Cloudの代替品

AIMLAPI logo

AIMLAPI

aimlapi.com

AIMLAPI provides one API and billing key for accessing a catalog of AI models for chat, reasoning, image, video, audio, voice, search, embeddings, code, and related tasks. It is intended for developers and teams that want to compare and use models from multiple providers through a common platform.

Amazon Bedrock logo

Amazon Bedrock

aws.amazon.com

Amazon Bedrock is a fully managed AWS platform for building generative AI applications and agents with access to foundation models, customization tools, safety controls, and production-oriented workflows. It supports teams that want to experiment through the console or build applications through AWS APIs and SDKs.

AI Endpoints logo

AI Endpoints

www.ovhcloud.com

OVHcloud AI Endpoints provides serverless inference APIs for a selection of generative AI models, including LLMs, voice, document, and image analysis models. It helps developers add AI capabilities to applications through OpenAI-compatible APIs and supported integrations.

ModelArk logo

ModelArk

www.byteplus.com

ModelArk is BytePlus's one-stop large language model service platform for organizations building, deploying, and scaling AI applications. It is positioned within BytePlus's broader AI-native cloud portfolio.

DeepSeek Platform logo

DeepSeek Platform

platform.deepseek.com

DeepSeek Platform provides access to DeepSeek AI models, API documentation, and developer resources through an online API platform. It is intended for users who want to sign up or log in to use DeepSeek’s platform services.

Parasail logo

Parasail

parasail.io

Parasail is an inference cloud for AI-native startups that provides access to open and frontier models through an OpenAI-compatible API. It offers serverless, elastic, dedicated, and batch deployment options with per-token or GPU-based pricing.