LPU-based inference architecture
Groq’s LPU is a custom inference chip designed for deterministic, token-based execution. The architecture page says the design removes traditional software complexity and aims to keep performance predictable.
Groq is an inference platform for developers and teams that need fast, low-cost model serving through the Groq LPU and GroqCloud. It offers OpenAI-compatible API access, on-demand model pricing, prompt caching, batch processing, and an enterprise path for larger deployments.
Groq is an inference platform built around the Groq LPU, a custom chip and software stack focused on running AI models quickly and at a predictable cost. The site positions GroqCloud as the developer-facing service for accessing models and inference, with OpenAI-compatible API access and support for starting quickly from existing code.
The product is aimed at developers and teams that need low-latency model responses, scalable inference, and pricing that is described as linear and predictable. The pricing page shows on-demand access to openly available models, prompt caching, built-in tools for compound systems, and a Batch API for large workloads, while the enterprise pages add deployment optionality and dedicated support for larger organizations.
Groq’s LPU is a custom inference chip designed for deterministic, token-based execution. The architecture page says the design removes traditional software complexity and aims to keep performance predictable.
The pricing page shows on-demand access to multiple openly available models, including GPT-OSS, Llama, Qwen, Kimi, and Whisper variants, with model-specific token pricing.
Groq documents prompt caching with separate cached and uncached input token pricing. The page notes there is no extra fee for the caching feature itself, and the discount applies when a cache hit occurs.
Compound AI systems can use built-in tools such as web search, visit website, code execution, and browser automation. Pricing is passed through to the underlying models and server-side tools.
Batch API supports asynchronous large-scale request processing, with a stated 50% lower cost, no impact to standard rate limits, and a 24-hour to 7 day processing window.
The site states Groq’s API is OpenAI compatible and can be started with just a few lines of code using the Groq base URL and API key.
Teams building chat or assistant products can use GroqCloud to serve model responses with low latency and predictable per-token pricing, especially when response time affects the user experience.
Developers who already use OpenAI-style client libraries can switch to Groq by changing the base URL and using the Groq API key, which lowers migration friction.
Products that need retrieval, web lookup, or code execution can use compound AI systems with built-in tools such as web search, visit website, and code execution.
Organizations running large jobs can use the Batch API to process requests asynchronously at lower cost without affecting standard rate limits.
Enterprises with custom capacity, deployment, or support needs can route through the Enterprise API Solutions flow for larger-scale inference planning.
Groq provides inference infrastructure for developers and teams that need fast, low-cost model serving. The site shows OpenAI-compatible access and supports starting with a small code change using the Groq API base URL.
The pricing page lists core model access, enterprise-only models, prompt caching, built-in tools, and batch API processing. The home page also points to GroqCloud as the place developers use for inference.
The pricing page shows on-demand pricing for several openly available models and indicates that other models are available for specific customer requests, including fine-tuned models.
Yes. The enterprise page offers an Enterprise API Solutions flow for larger deployments, custom solutions, and dedicated support, and it also mentions deployment optionality.
The site says you can start for free and upgrade as needs grow, and the pricing page includes a contact path for enterprise API solutions or on-prem deployments.