Lambda is a cloud GPU platform for AI development and production workloads. It provides on-demand instances for testing and prototyping, 1-Click Clusters for distributed training and inference, and larger single-tenant infrastructure for teams that need dedicated capacity.
The platform supports NVIDIA GPUs including B200, H100, A100, GH200, A6000, and other listed configurations for instances. Its 1-Click Clusters are offered as production-ready NVIDIA HGX B200 or H100 environments ranging from 16 to more than 2,000 GPUs, with InfiniBand networking and managed orchestration options.
Teams can use Lambda for foundation-model training, fine-tuning, inference, and serving high-volume workloads. Instances are billed by the GPU hour and can be launched through the cloud service; cluster pricing is listed by GPU count and contract duration, with reserved-capacity inquiries handled by the Lambda team. The pricing page states that there are no egress fees, while applicable sales taxes may apply.
For larger deployments, Lambda describes single-tenant infrastructure, caged clusters with hardware-level isolation, and a security posture that includes SOC 2 Type II attestation. Cluster environments can use managed Kubernetes or Slurm orchestration and S3-compatible storage, helping teams operate distributed AI workloads without managing every infrastructure layer themselves.