DigitalOcean Inference Engine logo

DigitalOcean Inference Engine

Freemium
访问

DigitalOcean Inference Engine provides a unified platform for serving AI models through serverless, batch, and dedicated inference. It supports multimodal workloads and includes routing, experimentation, and evaluation tools for teams moving models into production.

DigitalOcean Inference Engine logoDigitalOcean Inference Engine

什么是 DigitalOcean Inference Engine?

DigitalOcean Inference Engine is a production AI inference platform for serving models across real-time, asynchronous, and dedicated workloads. It brings Serverless Inference, Batch Inference, and Dedicated Inference together through OpenAI- and Anthropic-compatible interfaces, allowing teams to choose an execution model without managing separate inference systems.

The platform supports text, image, video, audio, and vision-language workloads. Its control layer includes the Inference Router for policy-based model selection, the Model Playground for side-by-side experimentation, and Evaluations for testing models, endpoints, and routing policies against a team’s own datasets. The Model Library helps users discover available models and deployment options.

DigitalOcean Inference Engine 能做什么?

Serverless, batch, and dedicated inference

Run real-time APIs and agent workloads with Serverless Inference, asynchronous jobs with Batch Inference, or sustained workloads on Dedicated Inference endpoints.

Inference Router

Define routing policies using natural language or structured rules, route by cost or latency, override requests at runtime, and configure failover behavior across supported models.

Multimodal model access

Work with models for text generation, image generation, video generation, speech generation, and vision-language understanding through a single API key.

Model Playground and Model Library

Browse and filter available models, compare outputs side by side, adjust parameters, and export cURL or SDK code from a playground configuration.

Production observability

Track tokens, time to first token, latency, errors, spend, throughput, and batch job lifecycle information within the inference workflow.

Structured evaluations

Compare catalog models, dedicated endpoints, imported BYOM models, and routing policies using custom or pre-built rubrics, reusable presets, datasets, and LLM-as-a-Judge scoring.

使用场景

“Real-time AI applications”

Use Serverless Inference for production APIs, agents, and applications that need responses while requests are being processed.

“Large-scale asynchronous pipelines”

Submit batch jobs for evaluation, content enrichment, or moderation workloads where immediate responses are not required.

“Reserved model serving”

Deploy a model on dedicated GPU capacity when a workload needs sustained throughput, configurable GPU types, or more control over latency and scaling.

“Model and routing selection”

Test candidate models and routing policies on representative prompts or datasets, then compare quality and performance before changing a production workflow.

“Custom-model deployment”

Import a custom or fine-tuned model from Hugging Face or DigitalOcean Spaces and evaluate it alongside catalog models and dedicated endpoints.

常见问题

What models can be used with Inference Engine?

The platform provides a catalog of open-source and frontier models, including models available through Serverless or Dedicated Inference. The Model Library lists 86 available models in the supplied source, and supports filtering by provider, type, capabilities, and availability. Custom or fine-tuned models imported from Hugging Face or DigitalOcean Spaces can also be evaluated.

How does the Inference Router select a model?

Developers can use system-level routing or presets that match requests to models based on task type, cost, and performance requirements. Policies can also define cost or latency routing, runtime overrides, and failover behavior.

What is the difference between Serverless, Batch, and Dedicated Inference?

Serverless is intended for real-time requests and scales without user-managed provisioning. Batch uses asynchronous, job-based execution for workloads that do not require real-time latency. Dedicated provides reserved GPU endpoints, configurable GPU and scaling settings, and infrastructure-level control for sustained workloads.

How are models evaluated?

Evaluations runs candidate models or routing policies against uploaded CSV or JSONL datasets using pre-built or custom rubrics. Results can include judge scores, time to first token, total latency, throughput, and tokens per request, and evaluation configurations can be saved and rerun.

Can evaluations be automated?

Yes. Evaluation jobs can be triggered through the MCP interface, API, or SDK, including from model registration events, schedules, and deployment pipelines. The supplied evaluation documentation states that up to three evaluation runs can be concurrent.

快速信息

Product type
AI inference platform
Deployment modes
Serverless, batch, and dedicated inference
Supported modalities
Text, image, video, audio, and vision-language
Interfaces
OpenAI- and Anthropic-compatible endpoint and API/SDK workflows
Model discovery
DigitalOcean Model Library and Model Playground
Evaluation inputs
CSV or JSONL datasets, up to 1GB or 1,000 rows per dataset

DigitalOcean Inference Engine 替代品

each::labs logo

each::labs

eachlabs.ai

each::labs provides a single API for orchestrating more than 600 AI models, with routing, fallback handling, observability, and usage-based pricing. It is designed for teams building and operating production AI applications across video, image, audio, and text workflows.

PiAPI logo

PiAPI

piapi.ai

PiAPI is a unified platform for generating video, images, audio, 3D assets, and LLM outputs through a model catalog, playground, APIs, CLI, and MCP server. It is designed for developers, AI agents, and automation workflows that need access to multiple generative models.

AIMLAPI logo

AIMLAPI

aimlapi.com

AIMLAPI provides one API and billing key for accessing a catalog of AI models for chat, reasoning, image, video, audio, voice, search, embeddings, code, and related tasks. It is intended for developers and teams that want to compare and use models from multiple providers through a common platform.

IBM watsonx.ai logo

IBM watsonx.ai

www.ibm.com

IBM watsonx.ai is an enterprise AI development studio for building predictive, prescriptive, and generative AI solutions. It supports AI builders, data scientists, and developers across model development, customization, retrieval-augmented generation, deployment, and lifecycle management.

Together AI logo

Together AI

together.ai

Together AI 是一个支持推理、微调、GPU 集群、沙盒和托管存储的 AI 云平台。

Bento logo

Bento

www.bentoml.com

Bento is an inference platform for packaging, deploying, optimizing, and operating AI and machine-learning models at scale. It supports open and custom models across cloud, on-premises, Kubernetes, and bring-your-own-cloud environments.