Cloudflare Workers AI logo

Cloudflare Workers AI

Freemium
访问

Cloudflare Workers AI lets developers run open-source machine learning models through serverless GPUs on Cloudflare’s global network. Models can be invoked from Workers, Pages, or applications using the Cloudflare API without managing GPU infrastructure.

什么是 Cloudflare Workers AI?

Cloudflare Workers AI is a serverless machine learning platform for running models on GPUs hosted across Cloudflare’s network. It helps developers add AI inference to applications without provisioning, scaling, or maintaining GPU infrastructure.

Models can be invoked from Cloudflare Workers, Pages, or other applications through the Cloudflare API. The catalog includes open-source models for tasks such as text generation, image classification, object detection, image-to-text, text-to-image, text-to-speech, and embeddings.

Workers AI is available on Free and Paid Workers plans. Usage is measured in Neurons, which represent the GPU compute required by a request; model-specific rates vary.

Cloudflare Workers AI 能做什么?

Serverless GPU inference

Run machine learning models on Cloudflare-hosted GPUs without managing servers, scaling infrastructure, or paying for idle capacity.

Open-source model catalog

Choose from a catalog of models covering text generation, image classification, object detection, image-to-text, text-to-image, text-to-speech, embeddings, and other tasks.

Multiple invocation paths

Call models from Cloudflare Workers, Pages, or applications outside those products through the Cloudflare API.

Model discovery and comparison

Browse models by task type, capability, or author, and compare up to three models in the model catalog.

Usage-based billing

Workers AI measures consumption in Neurons and supports a free daily allocation, with model-specific pricing and paid usage documented by Cloudflare.

Cloudflare AI platform connections

Use Workers AI with AI Gateway for caching, rate limiting, retries, and model fallback, or with Vectorize for vector search and language-model context.

使用场景

“Generate and process text”

Add text generation to a Worker, Pages project, or API-backed application using a model selected from the catalog. This can support language-model features without operating a separate inference service.

“Analyze images and visual input”

Use image classification, object detection, or image-to-text models when an application needs to identify visual content or turn images into text-based information.

“Create image and audio experiences”

Use text-to-image models for image generation workflows or text-to-speech models when an application needs to produce audio from written content.

“Build search and retrieval workflows”

Combine model inference with Vectorize to support semantic search, recommendations, anomaly detection, or retrieval of context and memory for a language-model request.

常见问题

How do developers call Workers AI models?

Models can be invoked from Cloudflare Workers, Pages, or other applications through the Cloudflare API.

What kinds of models are available?

The catalog includes models for tasks such as text generation, image classification, object detection, image-to-text, text-to-image, text-to-speech, and embeddings. The available catalog can change over time.

Is Workers AI available on a free plan?

Yes. Workers AI is included with Cloudflare Free and Paid Workers plans. The pricing documentation lists a free allocation of 10,000 Neurons per day.

How is Workers AI usage billed?

Usage is measured in Neurons. On Workers Paid, usage above the 10,000-Neuron daily free allocation is priced at $0.011 per 1,000 Neurons, with rates varying by model and task. Some models require a paid billing method.

Can Workers AI be used with other Cloudflare products?

Yes. The product documentation describes integrations with AI Gateway, Vectorize, Workers, Pages, R2, D1, Durable Objects, and KV. AI Gateway provides request controls such as caching, retries, rate limiting, and model fallback; Vectorize supports vector-based workflows.

快速信息

Product type
Serverless machine learning inference platform
Provider
Cloudflare
Execution
Cloudflare-hosted GPUs on the global network
Access methods
Workers, Pages, and Cloudflare API
Plans
Cloudflare Free and Paid Workers plans
Billing unit
Neurons

Cloudflare Workers AI 替代品

Hyperstack logo

Hyperstack

www.hyperstack.cloud

Hyperstack is a cloud GPU platform for running AI and machine learning workloads, including training, inference, data analytics, and model development. It also provides AI Studio, virtual machines, and managed Kubernetes for deploying and operating GPU-backed workloads.

Together AI logo

Together AI

together.ai

Together AI 是一个支持推理、微调、GPU 集群、沙盒和托管存储的 AI 云平台。

DeepInfra logo

DeepInfra

deepinfra.com

DeepInfra provides hosted machine-learning model inference and on-demand GPU instances for developers and teams. Its catalog covers text, image, audio, video, embedding, reranking, and other model workloads with pay-as-you-go pricing.

Hyperbolic logo

Hyperbolic

hyperbolic.xyz

Hyperbolic is an open-access GPU and AI cloud for deploying on-demand H100, H200, B200, and other GPU capacity. It supports experimentation, training, fine-tuning, inference, and production workloads through on-demand instances, reserved clusters, and Private Cloud infrastructure.

Lambda logo

Lambda

lambda.ai

Lambda provides cloud GPU compute for AI training, fine-tuning, inference, and prototyping. Teams can launch on-demand GPU instances, use production-ready 1-Click Clusters, or discuss reserved and single-tenant infrastructure for larger workloads.

CoreWeave logo

CoreWeave

www.coreweave.com

CoreWeave is an AI-focused cloud platform that combines GPU infrastructure, storage, networking, orchestration, and operational tooling for training and serving AI workloads. It supports teams moving from model experiments to production systems, including reinforcement-learning and agent-development workflows.