SambaNova logo

SambaNova

Freemium
訪問

SambaNova is an AI inference platform for developers, enterprises, and infrastructure operators running large open-source models and agentic workloads. It combines cloud services with deployable hardware and software for cloud, on-premises, hybrid, and air-gapped environments.

SambaNovaとは?

SambaNova is an AI inference platform for serving large open-source models and coordinating multiple models in generative and agentic applications. It combines RDU-based hardware, dataflow processing, tiered memory, model-management software, and APIs into a stack that can be consumed through SambaCloud or deployed in an organization's own infrastructure.

SambaCloud provides managed inference for models such as DeepSeek, Llama, Qwen, MiniMax M2.7, Gemma 4 31B, and gpt-oss-120b. SambaNova also offers SambaRack systems and SambaOrchestrator for organizations that need cloud, on-premises, hybrid, or air-gapped deployment options. The platform is aimed at developers building AI applications and enterprises that need control over their models, data, and deployment environment.

SambaNovaでできること

Managed inference through SambaCloud

SambaCloud provides hosted inference for a range of large open-source models, including DeepSeek, Llama, Qwen, MiniMax M2.7, Gemma 4 31B, and gpt-oss-120b. Supported models can cover text, image, or audio workloads.

OpenAI-compatible application access

Developers can move an application from another provider by using SambaNova's OpenAI-compatible endpoints, setting the SambaNova API key and base URL, and selecting a supported model.

RDU dataflow architecture

SambaNova's Reconfigurable Dataflow Unit uses dataflow processing and a three-tier memory architecture to support inference for large models, including workloads that keep multiple models available.

Model and infrastructure management

SambaOrchestrator provides tools for model deployment, monitoring, automatic scaling, load balancing, server management, and cloud creation across data-center environments.

Flexible deployment and data control

The platform can be deployed in cloud, on-premises, hybrid, and air-gapped environments. SambaNova states that customers own their models and control their data; SambaCloud states that it does not see or collect user data or prompts.

Developer ecosystem support

SambaCloud lists integrations with CrewAI, Hugging Face, Cline, and AWS, giving developers ways to connect inference with existing AI frameworks and development tools.

利用シーン

“Multi-model agent systems”

Teams can combine several models in an agentic workflow, such as routing requests among general assistance, research, sales, and finance agents or chaining models for more complete results.

“Private enterprise research”

Organizations can run AI-driven research and reporting over business data while choosing cloud, on-premises, hybrid, or air-gapped deployment according to their security and data-control requirements.

“Coding and developer agents”

Developers can build coding agents and other software tools using fast inference, OpenAI-compatible endpoints, and integrations such as Cline, Hugging Face, and CrewAI.

“Content and media analysis”

Applications can use supported text, image, or audio-capable models for tasks such as video transcription, content analysis, and generating insights from media.

“Infrastructure-scale inference services”

Neoclouds, service providers, and enterprise infrastructure teams can deploy SambaRack systems with SambaOrchestrator to run models, monitor workloads, and scale capacity across data centers.

よくある質問

What is the difference between SambaCloud and SambaRack?

SambaCloud is SambaNova's managed AI inference service for developers. SambaRack is a hardware platform combining hardware, operating system, and networking for running inference workloads in a data center, with SambaOrchestrator providing management capabilities.

Which models can be used with SambaNova?

The site lists support for open-source models including DeepSeek, Llama, Qwen, MiniMax M2.7, Gemma 4 31B, and gpt-oss-120b. Available modalities depend on the model and can include text, image, and audio processing.

Can an existing OpenAI-based application be moved to SambaNova?

SambaCloud provides OpenAI-compatible endpoints. The documented migration flow is to set the OpenAI API key variable to a SambaNova API key, set the base URL, choose a model, and run the application.

Where can SambaNova be deployed?

SambaNova describes cloud, on-premises, and hybrid deployments, including a path to start in the cloud and later expand to on-premises infrastructure. It also supports air-gapped environments for deployments that require network isolation.

Does SambaNova use customer prompts or data to train its systems?

The SambaCloud page states that SambaCloud does not see or collect user data or prompts. The enterprise page states that customers own their models and control their data, and that data is not shared with SambaNova or third parties.

クイック情報

Product category
AI inference platform
Primary users
Developers, enterprises, and AI infrastructure operators
Managed product
SambaCloud
Infrastructure products
SambaRack and SambaOrchestrator
Deployment options
Cloud, on-premises, hybrid, and air-gapped environments
API workflow
OpenAI-compatible endpoints

SambaNovaの代替品

IBM watsonx.ai logo

IBM watsonx.ai

www.ibm.com

IBM watsonx.ai is an enterprise AI development studio for building predictive, prescriptive, and generative AI solutions. It supports AI builders, data scientists, and developers across model development, customization, retrieval-augmented generation, deployment, and lifecycle management.

Together AI logo

Together AI

together.ai

Together AIは推論、ファインチューニング、GPUクラスター、サンドボックス、マネージドストレージに対応するAIクラウドプラットフォームです。

Bento logo

Bento

www.bentoml.com

Bento is an inference platform for packaging, deploying, optimizing, and operating AI and machine-learning models at scale. It supports open and custom models across cloud, on-premises, Kubernetes, and bring-your-own-cloud environments.

Prodia logo

Prodia

prodia.com

Prodia is a multi-silicon inference platform focused on video generation. It develops AI model implementations across different hardware to balance cost, output quality, and performance.

DigitalOcean logo

DigitalOcean

digitalocean.com

本番環境のAIアプリを構築・デプロイ・拡張できるAIネイティブなクラウドプラットフォーム。

ModelArk logo

ModelArk

www.byteplus.com

ModelArk is BytePlus's one-stop large language model service platform for organizations building, deploying, and scaling AI applications. It is positioned within BytePlus's broader AI-native cloud portfolio.