Serverless inference
Run open-source models on demand without managing infrastructure, with the homepage positioning this as the fastest path for serverless use.
Together AI is an AI cloud platform for inference, fine-tuning, GPU clusters, sandboxes and managed storage, with serverless and dedicated options.
Together AI is an AI cloud platform for building, deploying, and optimizing model workloads. The homepage describes it as a full-stack AI platform for inference, fine-tuning, and GPU clusters, with additional products for sandbox compute, managed storage, and model shaping.
The platform is organized around several deployment and workflow layers: serverless inference for on-demand use, dedicated inference for single-tenant performance, dedicated container inference for generative media workloads, accelerated compute for GPU access, sandbox environments for development, and fine-tuning for production model adaptation. Pricing pages show published rates for many serverless models and infrastructure products, while some larger dedicated hardware options require contact sales.
Run open-source models on demand without managing infrastructure, with the homepage positioning this as the fastest path for serverless use.
Deploy models on single-tenant infrastructure with guaranteed performance, custom model support, and autoscaling for traffic spikes.
Use hourly GPU capacity or reserved clusters for larger compute jobs, with on-demand and reserved pricing options shown for H100, H200, and B200 hardware.
Create VM sandboxes, run code securely through the API, and hibernate and resume sandboxes for development workflows.
Train open-source models for production with supervised fine-tuning and direct preference optimization on supported model families.
Store data in managed object storage and parallel filesystems designed for AI workloads, with zero egress fees stated on the homepage.
Teams can start with serverless inference for quick experiments and then move to dedicated endpoints when they need steadier performance or more control.
Organizations that need private infrastructure can deploy custom or open models on dedicated hardware with guaranteed performance and autoscaling.
Developers can fine-tune supported open-source models for domain-specific behavior, using supervised fine-tuning or direct preference optimization.
AI teams working on larger jobs can rent hourly or reserved GPU capacity for training, evaluation, or other compute-heavy work.
Application builders can use sandboxes and managed storage to spin up development environments, run code securely, and keep data close to compute.
Together AI is a cloud platform for running and shaping AI workloads. The site highlights serverless inference, dedicated inference, dedicated container inference, GPU clusters, sandbox environments, managed storage, and fine-tuning.
The pricing page shows serverless inference, dedicated inference, GPU clusters, sandbox compute, managed storage, and fine-tuning. The homepage also highlights batch inference, dedicated model inference, and dedicated container inference.
The site shows an OpenAI-compatible API pattern on model pages such as `https://api.together.xyz/v1/chat/completions`, with example calls in cURL, Python, and TypeScript.
Pricing is published for many serverless models, GPU clusters, sandbox compute, storage, and fine-tuning. Dedicated hardware such as some larger GPU options is listed as contact-us pricing.
The source does not provide a full SDK or framework compatibility list. It does show code examples and links to quickstart guides, docs, playground access, and model pages.