Multiple inference modes
Serve open-source models or your own post-trained models on a stack tuned for throughput and latency, with serverless, on-demand, and reserved options.
Fireworks AI is a generative AI platform for serving, training, fine-tuning, and deploying open-source models with published pricing.
Fireworks AI is a platform for serving, fine-tuning, training, and deploying generative AI models. The public pages focus on open-source LLMs and image models, along with the ability to bring your own post-trained models onto the same inference stack.
The product is positioned around specialized intelligence: users can start with serverless inference, move to dedicated on-demand or reserved deployments, and use the training workflows to produce models that deploy to production in seconds. The site also shows a model library with text, vision, image, and audio models, plus pricing for embeddings and training.
Pricing is published for self-serve usage, including per-token serverless inference, per-1M-token fine-tuning, and per-GPU-second on-demand deployments. The pricing page also directs enterprise customers to contact the team for deployment needs that require higher speeds, lower costs, or higher rate limits.
Serve open-source models or your own post-trained models on a stack tuned for throughput and latency, with serverless, on-demand, and reserved options.
Use OpenAI- and Anthropic-compatible serverless endpoints, or move to dedicated on-demand deployments with multi-region support.
Train with guided runs, configuration-led jobs, or your own training logic, including custom loss functions, trainers, and RL loops on Fireworks GPUs.
Deploy checkpoints to production quickly and keep the training and serving stack aligned so the same model can move from training into inference without a handoff.
Browse a model library with open LLMs, vision models, audio models, and image models, including recently added frontier releases.
Pricing covers serverless inference, fine-tuning, reinforcement fine tuning, embeddings, and on-demand GPU deployments in one place.
Serve open-source foundation models through serverless endpoints when you want to start quickly and pay per token.
Run dedicated deployments for post-trained models when a team needs multi-region serving, custom performance tuning, or higher quotas.
Fine-tune or train models using guided, configuration-led, or custom-code workflows, then move checkpoints into production.
Browse and test the current model library to choose between LLMs, vision models, audio models, and image models for a specific task.
Use compatible serverless endpoints to migrate existing API-based applications by changing the service URL rather than redesigning the app.
Fireworks AI provides serverless inference, on-demand deployments, and reserved capacity for serving open-source models or models you have trained on the platform. Its pricing page also shows fine-tuning options for supervised, preference, and reinforcement methods.
The pricing page says serverless inference is available with per-token pricing and zero setup, and it includes $1 in free credits to get started. On-demand deployments are billed per GPU second, and enterprise users are directed to contact the team.
The site highlights OpenAI- and Anthropic-compatible serverless access, dedicated on-demand deployments, and multi-region support for on-demand workloads. The public pages do not provide a full SDK or integration matrix.
Yes. The site describes training workflows that range from guided runs to configuration-led jobs and custom training logic, and it says checkpoints can deploy to production in seconds.
The public pages emphasize inference, training, deployment, and model browsing, but they do not publish detailed setup steps or a complete integration catalog on the pages reviewed.