Instructor logo

Instructor

Freemium
访问

Instructor is a developer library for extracting structured, validated data from large language models. It uses Pydantic schemas, automatic retries, streaming, and a consistent interface across cloud, local, and routed LLM providers.

什么是 Instructor?

Instructor is a developer library for turning LLM responses into structured, typed data. Developers define the expected output with Pydantic models, then pass that model to a client request. Instructor converts the model into a schema the provider can use, validates the response, retries when validation fails, and returns a typed Pydantic object.

It is designed for schema-first extraction rather than agent orchestration. The Python workflow supports nested models, lists, enums, optional fields, field constraints, custom validators, asynchronous clients, partial responses, streaming, hooks, and access to raw completions. It can be used with hosted models, routing services, and local open-source models through documented integrations.

Instructor 能做什么?

Pydantic response models

Define the fields, types, constraints, and nested structures expected from an LLM response.

Validation and retries

Instructor validates generated data against the response model and can retry failed requests, with retry behavior configurable through max_retries and Tenacity.

Multi-provider interface

from_provider provides a common client-initialization pattern across supported providers, while provider-specific modes and capabilities remain documented separately.

Streaming and partial results

create_partial can yield partially populated models as a response is generated, and the library supports streaming lists and iterable responses.

Provider output modes

Depending on provider support, Instructor can use tool calls, JSON generation, JSON Schema, Markdown-embedded JSON, or parallel tool calls.

Synchronous and asynchronous usage

The Python client supports ordinary calls and an asynchronous client through async_client=True.

使用场景

“Extract fields from unstructured text”

Convert a sentence, document fragment, or user message into a typed object such as a person record with fields for name, age, and occupation.

“Parse nested support cases”

Represent a customer support case with customer details and a validated list of tickets, including priorities, estimated hours, and field-level rules.

“Process long or incremental outputs”

Consume partial model results while a large structured response is being generated instead of waiting for the complete object.

“Build API extraction endpoints”

Use an Instructor client inside a FastAPI route so submitted text is returned as a declared Pydantic response type.

“Switch models or hosting environments”

Keep the response-model workflow while moving between cloud providers, routing layers, or local models such as those served through Ollama.

常见问题

How do I install Instructor?

Install the base package with `pip install instructor`. Provider-specific extras are available for integrations such as Anthropic and Google/Gemini.

How does Instructor turn an LLM response into structured data?

You provide a Pydantic class as `response_model`. Instructor converts that model into a provider-compatible schema, formats the request, validates the response, retries on validation failure, and returns the typed object.

Can I use open-source or local models?

Yes. The documentation describes local open-source model integrations through Ollama and llama-cpp-python, as well as other provider and routing integrations. Capabilities vary by provider.

Can I inspect the original model response?

Yes. `create_with_completion` returns both the parsed Pydantic result and the raw completion.

Does Instructor support async workflows?

Yes. Initialize a provider with `async_client=True` and await the client’s `create` call. The documentation also provides a FastAPI example using an asynchronous endpoint.

快速信息

Category
Developer Tool
Primary workflow
Schema-first LLM data extraction
Schema system
Pydantic models and validators
Provider coverage
15+ documented LLM providers and integrations
Local model support
Ollama and llama-cpp-python integrations are documented
Python package installation
pip install instructor

Instructor 替代品

OpenAI Platform logo

OpenAI Platform

platform.openai.com

OpenAI Platform is a developer platform for building applications with the OpenAI API. It provides API access, model and pricing information, technical documentation, starter apps, and cookbooks for implementation guidance.

Cortex Docs logo

Cortex Docs

cortexdocs.dev

Cortex Docs is an open-source API tooling workflow that turns API specifications and Markdown into interactive documentation, typed SDKs, and MCP servers. It helps developers publish API interfaces for human users, applications, and AI agents from a shared project.

Semantic Kernel logo

Semantic Kernel

learn.microsoft.com

Semantic Kernel is Microsoft’s public GitHub software project for integrating large language model technology into applications. The repository includes implementation and documentation directories for .NET, Python, and Java development.

AI SDK logo

AI SDK

ai-sdk.dev

AI SDK is a TypeScript toolkit for building AI-powered applications and agents across supported model providers and web frameworks. It provides shared APIs for model generation, tool use, streaming, structured output, and AI user interfaces.

Gemini API logo

Gemini API

ai.google.dev

Gemini API is a developer API for integrating Google’s generative AI models into applications. It supports text and image generation, multimodal analysis, conversational agents, structured outputs, tool use, and related media workflows through SDKs or REST.

Orca logo

Orca

onorca.dev

Orca 是用于借助编码代理交付产品的智能体开发环境,支持在隔离工作树中并行运行多个 CLI 代理,并提供桌面端和移动端协同工作流。