On-demand GPU instances
Launch and manage GPU instances through a self-serve dashboard with usage-based billing and no minimum term for on-demand capacity.
Hyperbolic is an open-access GPU and AI cloud for deploying on-demand H100, H200, B200, and other GPU capacity. It supports experimentation, training, fine-tuning, inference, and production workloads through on-demand instances, reserved clusters, and Private Cloud infrastructure.
Hyperbolic is an open-access GPU and AI cloud that provides on-demand and dedicated compute for AI development. Users can launch H100, H200, B200, and other high-performance GPUs for experimentation, training, fine-tuning, inference, batch jobs, and production workloads.
The platform offers three infrastructure options: usage-based On-Demand GPUs with no fixed commitment, Reserved Clusters for predictable capacity at discounted rates with commitments under one year, and Private Cloud infrastructure for teams seeking dedicated, long-term capacity through Hyperbolic’s supplier network. Its stated workflow supports choosing virtual machines or bare metal, selecting GPU counts from a single node to 1000+ GPUs, choosing InfiniBand or Ethernet, and launching a cluster in minutes.
Launch and manage GPU instances through a self-serve dashboard with usage-based billing and no minimum term for on-demand capacity.
Reserve dedicated GPU capacity for sustained workloads, with commitments under one year and discounted reserved rates.
Access dedicated GPU infrastructure through Hyperbolic’s supplier network for longer-term capacity and greater infrastructure control.
Select a virtual-machine or bare-metal setup, set the GPU count from a single node to 1000+ GPUs, and choose InfiniBand or Ethernet interconnects.
Access H100, H200, and B200 capacity. Listed starting rates are $3.19 per GPU hour for H100 SXM, $3.99 for H200, and $5.99 for B200; rates are refreshed weekly based on supplier availability.
Receive a notification within three minutes if an instance fails, with no charge for failed instances; billing applies to GPUs that come online.
Use reserved multi-node H100, H200, or B200 clusters for long training runs that need dedicated capacity and high-performance interconnects.
Run fast iteration on open-source checkpoints and custom models with bare-metal single-node or small multi-node capacity.
Start on-demand instances for experiments, benchmarks, evaluations, data processing, and other compute-heavy research tasks.
Programmatically launch and manage GPU resources for agent workflows, evaluation pipelines, and other automated workloads.
Move predictable workloads to reserved clusters or Private Cloud environments when a team needs dedicated hardware, isolated networking, and longer-term capacity.
Listed starting rates are $3.19 per hour for an NVIDIA H100 SXM, $3.99 per hour for an H200, and $5.99 per hour for a B200. Pricing is refreshed weekly based on the best available supplier rates. On-demand billing is usage-based, and failed instances are not charged.
No. On-demand instances do not require a minimum term or contract. Users can launch capacity for an experiment, benchmark, fine-tuning job, or training run and shut it down when finished.
Yes. Clustered allocation can scale from a single node to 1000+ GPUs and supports InfiniBand or Ethernet interconnects. Reserved clusters provide dedicated capacity for workloads that need predictable availability.
On-demand GPUs are usage-based and have no fixed commitment. Reserved GPUs provide dedicated capacity at a discounted rate in exchange for a fixed-term commitment paid upfront, making them suited to predictable, sustained workloads.
Traffic data is for reference only.
www.hyperstack.cloud
Hyperstack is a cloud GPU platform for running AI and machine learning workloads, including training, inference, data analytics, and model development. It also provides AI Studio, virtual machines, and managed Kubernetes for deploying and operating GPU-backed workloads.
together.ai
Together AI is an AI cloud platform for inference, fine-tuning, GPU clusters, sandboxes, and managed storage.
deepinfra.com
DeepInfra provides hosted machine-learning model inference and on-demand GPU instances for developers and teams. Its catalog covers text, image, audio, video, embedding, reranking, and other model workloads with pay-as-you-go pricing.
lambda.ai
Lambda provides cloud GPU compute for AI training, fine-tuning, inference, and prototyping. Teams can launch on-demand GPU instances, use production-ready 1-Click Clusters, or discuss reserved and single-tenant infrastructure for larger workloads.
www.coreweave.com
CoreWeave is an AI-focused cloud platform that combines GPU infrastructure, storage, networking, orchestration, and operational tooling for training and serving AI workloads. It supports teams moving from model experiments to production systems, including reinforcement-learning and agent-development workflows.
developers.cloudflare.com
Cloudflare Workers AI lets developers run open-source machine learning models through serverless GPUs on Cloudflare’s global network. Models can be invoked from Workers, Pages, or applications using the Cloudflare API without managing GPU infrastructure.