Hosted model catalog
Browse more than 100 models across language, vision, speech, media-generation, embedding, reranking, and classification categories.
DeepInfra provides hosted machine-learning model inference and on-demand GPU instances for developers and teams. Its catalog covers text, image, audio, video, embedding, reranking, and other model workloads with pay-as-you-go pricing.
DeepInfra is a machine-learning inference and infrastructure platform for running hosted models and renting GPUs on demand. Its model catalog spans automatic speech recognition, embeddings, reranking, text generation, text-to-image, text-to-music, text-to-speech, text-to-video, world models, and zero-shot image classification.
The platform is designed around usage-based billing. Language models can use per-token pricing, while most other models are billed by inference execution time. DeepInfra also offers on-demand GPU instances for training, fine-tuning, and inference, with container setup, selectable base images, and hourly billing.
Browse more than 100 models across language, vision, speech, media-generation, embedding, reranking, and classification categories.
Use token pricing for some language models and execution-time billing for most other models, without long-term contracts or upfront costs.
Rent NVIDIA B200 GPUs in single-card or multi-GPU configurations, with pricing billed by the minute and no egress fees stated on the GPU page.
Select a GPU tier, name a container, choose a base image such as TensorFlow, PyTorch, or CUDA, and connect using a generated SSH command.
The site states a zero-retention policy for inputs, outputs, and user data, and identifies DeepInfra as SOC 2 and ISO 27001 certified.
DeepInfra says it operates inference-optimized infrastructure in secure, US-based data centers and can tailor inference around cost, latency, throughput, or scale.
Teams can select hosted text-generation models from families including DeepSeek, Qwen, Llama, Kimi, Gemini, and others, then pay for token usage.
Developers can evaluate or run workloads involving image, audio, and video inputs or outputs using the corresponding model categories in the catalog.
ML practitioners can launch GPU containers with PyTorch, TensorFlow, or CUDA base images for training, fine-tuning, or inference tasks.
Applications with changing demand can use usage-based model inference or add GPU capacity without committing to a long-term infrastructure contract.
DeepInfra provides hosted machine-learning model inference and on-demand GPU instances. Its catalog includes language, speech, image, video, embedding, reranking, and classification models.
Some language models use per-token pricing, while most other models are billed according to inference execution time. GPU instances use usage-based hourly pricing, with the GPU page also describing billing by the minute and no egress fees.
Yes. The GPU instances page describes launching containers for training, fine-tuning, and inference, and lists TensorFlow, PyTorch, and CUDA among the available base-image choices.
The documented workflow is to choose a GPU tier, name a container, select a base image, copy the generated SSH command, and launch the configured container.
Yes. The home page states that DeepInfra has a zero-retention policy for inputs, outputs, and user data. It also states that the company is SOC 2 and ISO 27001 certified.
www.hyperstack.cloud
Hyperstack is a cloud GPU platform for running AI and machine learning workloads, including training, inference, data analytics, and model development. It also provides AI Studio, virtual machines, and managed Kubernetes for deploying and operating GPU-backed workloads.
together.ai
Together AI 是一个支持推理、微调、GPU 集群、沙盒和托管存储的 AI 云平台。
hyperbolic.xyz
Hyperbolic is an open-access GPU and AI cloud for deploying on-demand H100, H200, B200, and other GPU capacity. It supports experimentation, training, fine-tuning, inference, and production workloads through on-demand instances, reserved clusters, and Private Cloud infrastructure.
lambda.ai
Lambda provides cloud GPU compute for AI training, fine-tuning, inference, and prototyping. Teams can launch on-demand GPU instances, use production-ready 1-Click Clusters, or discuss reserved and single-tenant infrastructure for larger workloads.
www.coreweave.com
CoreWeave is an AI-focused cloud platform that combines GPU infrastructure, storage, networking, orchestration, and operational tooling for training and serving AI workloads. It supports teams moving from model experiments to production systems, including reinforcement-learning and agent-development workflows.
developers.cloudflare.com
Cloudflare Workers AI lets developers run open-source machine learning models through serverless GPUs on Cloudflare’s global network. Models can be invoked from Workers, Pages, or applications using the Cloudflare API without managing GPU infrastructure.