Supported-model catalog
Browse model entries paired with the infrastructure providers that list support for them, including models such as Qwen3.8-27B, DeepSeek-V4.1-Flash, GLM-5.3, and gpt-oss-120b.
Hugging Face Inference Providers gives developers a way to compare supported models across infrastructure providers using model availability, pricing, context limits, latency, throughput, and capability indicators.
Hugging Face Inference Providers is a model-serving and provider-comparison layer associated with the Hugging Face Hub. Its supported-model catalog shows which infrastructure providers offer selected models and presents operational data for comparing those options. The catalog is aimed primarily at developers and teams evaluating hosted inference for applications built with Hugging Face models.
The main purpose of the catalog is to compare provider-specific options for the same model. It reports input and output pricing where available, context limits, latency, throughput, tool support, and structured-output support. Because these values can differ between providers, the model and provider combination—not just the model name—is the relevant unit for evaluation.
Inference Providers should be distinguished from Hugging Face Inference Endpoints. The Endpoints product is presented as a way to deploy any AI model from the Hugging Face Hub, while Inference Providers presents supported hosted-provider options and their comparison metrics.
Browse model entries paired with the infrastructure providers that list support for them, including models such as Qwen3.8-27B, DeepSeek-V4.1-Flash, GLM-5.3, and gpt-oss-120b.
Compare displayed input and output prices per 1 million tokens when pricing is available. Some rows show a dash instead of a price, so availability is not uniform across providers.
Review listed latency in seconds and throughput in tokens per second for individual model-provider combinations rather than relying only on model-level specifications.
Check the context value reported for each row. The catalog includes materially different limits across models and providers, and some entries do not display a context value.
The table marks whether a provider entry supports tools and structured outputs, allowing users to screen options for application requirements.
The wider Hugging Face Tasks directory organizes models across text, vision, audio, video, multimodal, tabular, and reinforcement-learning tasks; the provider catalog focuses on supported hosted inference entries.
A developer can inspect several provider rows for the same model and weigh listed cost, latency, throughput, context, and capability flags before selecting an inference route.
A team whose application needs tool use or structured outputs can use the corresponding table indicators to narrow the available provider-model combinations.
An engineering team can compare input/output token prices against latency and throughput, such as when deciding between a lower-cost provider and a faster listed option.
Users can use the broader Hugging Face Tasks directory to orient model discovery across language, computer-vision, audio, video, multimodal, tabular, and reinforcement-learning tasks before examining hosted options.
Teams can use the provider catalog to compare listed hosted options, while considering Inference Endpoints separately when their requirement is to deploy a model from the Hugging Face Hub.
It compares supported model-provider combinations using displayed input and output prices where available, context limits, latency, throughput, tool support, and structured-output support.
No. Some rows omit prices, context values, or performance values, and capability indicators can differ between providers for the same model.
The page presents latency and throughput as comparison metrics for its listed entries. They should be treated as catalog values for those entries, not as universal performance guarantees.
Inference Providers presents supported hosted-provider options and comparison data. Inference Endpoints is presented separately as a product for deploying any AI model from the Hugging Face Hub.
The Tasks directory lists language, computer-vision, audio, video, multimodal, tabular, and reinforcement-learning task areas. The provider catalog itself is a supported-model and provider comparison surface, so task support should be checked for the specific model.
eachlabs.ai
each::labs provides a single API for orchestrating more than 600 AI models, with routing, fallback handling, observability, and usage-based pricing. It is designed for teams building and operating production AI applications across video, image, audio, and text workflows.
piapi.ai
PiAPI is a unified platform for generating video, images, audio, 3D assets, and LLM outputs through a model catalog, playground, APIs, CLI, and MCP server. It is designed for developers, AI agents, and automation workflows that need access to multiple generative models.
wiro.ai
Wiro AI is a unified API and model marketplace for running image, video, audio, language, and other AI models. Developers can use one API key to test models, execute tasks, and build workflows and agents.
gmicloud.ai
面向 NVIDIA GPU 的生产推理、训练与微调 AI 基础设施。
aimlapi.com
AIMLAPI provides one API and billing key for accessing a catalog of AI models for chat, reasoning, image, video, audio, voice, search, embeddings, code, and related tasks. It is intended for developers and teams that want to compare and use models from multiple providers through a common platform.
chat100.ai
Chat100.ai 是一个网页 AI 聊天平台,可在同一处访问、切换和比较 ChatGPT、Grok 与 Gemini。