Inference for model serving
Modal lets teams run inference workloads for LLMs, audio, image generation, embeddings, and custom models, with scale-to-zero behavior between requests and support for token streaming, WebRTC, and WebSocket.
Modal is a serverless AI infrastructure platform for CPU, GPU, and data-intensive workloads from your own code, with autoscaling compute for inference, training, batch jobs, and sandboxes.
Modal is a serverless AI infrastructure platform for developers who need to run CPU, GPU, and data-intensive workloads from their own code. The site positions it as infrastructure for AI and data teams, with support for inference, training, batch processing, and sandboxes.
The product is centered on a code-first workflow: the Modal SDK lets you define logic, storage, and hardware in one place, then run it on autoscaling compute with sub-second starts and integrated observability. Pricing is usage-based, so you pay for actual compute time rather than idle capacity.
Modal lets teams run inference workloads for LLMs, audio, image generation, embeddings, and custom models, with scale-to-zero behavior between requests and support for token streaming, WebRTC, and WebSocket.
Training workflows can start from a single GPU or scale to multi-node runs, with support for fine-tuning, reinforcement learning, and parallel hyperparameter sweeps.
Sandboxes provide isolated environments for running untrusted code, coding agents, and RL rollouts, with fast scheduling, GPU or CPU capacity, and support for custom images and dependencies.
The platform includes an SDK and code-defined primitives so teams can specify logic, storage, hardware, and deployment together in Python.
Modal provides logging, metrics, readiness probes, health checks, and visibility into functions, sandboxes, and containers for operational debugging.
The pricing and product pages describe autoscaling across clouds and regions, with region selection available on paid plans and infrastructure that can scale up to large GPU fleets.
Serve LLMs, audio models, image generation pipelines, embeddings, or custom inference engines with scale-to-zero behavior between requests.
Run fine-tuning, reinforcement learning, and multi-node training jobs without managing separate training infrastructure.
Spin up isolated environments for coding agents, untrusted code execution, and other ephemeral workloads that need strong separation.
Process batches, generate datasets, or run parallel evaluation jobs where thousands of workers may need to run at once.
Use the pricing and plan structure to choose between a Starter setup, a Team workspace, or an Enterprise arrangement with security and support features.
Modal is a serverless AI infrastructure platform for running CPU, GPU, and data-intensive workloads from your own code. The site presents it as a platform for inference, training, batch processing, and sandboxes.
The pricing page shows three main plans: Starter, Team, and Enterprise. Starter begins at $0 plus compute per month, Team at $250 plus compute per month, and Enterprise uses custom pricing.
The source describes Modal as staying in Python and using an SDK to define logic, hardware, storage, and deployment in code. It also highlights functions, sandboxes, and containers as part of the workflow.
Modal is designed for AI and data teams that need to run inference, training jobs, batch processing, or isolated sandbox environments at scale.
The pricing page includes a Starter plan with limited scheduled and web functions, region selection, and 3 workspace seats, while Team and Enterprise add more capacity and features. The source does not provide a complete public list of every limitation.