pgvector logo

pgvector

Freemium
访问

pgvector is an open-source PostgreSQL extension for storing vectors and running exact or approximate nearest-neighbor searches alongside relational data. It supports multiple vector formats, distance functions, and standard PostgreSQL operations for applications that need similarity search.

什么是 pgvector?

pgvector is an open-source PostgreSQL extension for storing vector data and performing vector similarity search. It keeps embeddings with application data in PostgreSQL rather than requiring a separate vector database, while allowing similarity queries to be combined with normal SQL operations.

The extension performs exact nearest-neighbor search by default, which provides perfect recall for the search operation. HNSW and IVFFlat indexes can be added for approximate search when lower query latency or larger collections justify trading some recall for speed. pgvector also supports quantization for workloads with many vectors.

Users interact with pgvector through PostgreSQL SQL. The extension supports vector storage, nearest-neighbor queries, distance calculations, aggregate averages, updates, deletes, upserts, bulk loading with COPY, and queries against vectors from other rows.

pgvector 能做什么?

Exact and approximate nearest-neighbor search

Exact search is the default and provides perfect recall. HNSW and IVFFlat indexes enable approximate search with a configurable speed-versus-recall tradeoff.

Multiple vector representations

The extension supports single-precision, half-precision, binary, and sparse vectors, allowing storage choices to match the data and workload.

Six distance functions

SQL queries can use L2, inner product, cosine, L1, Hamming, and Jaccard distance. Hamming and Jaccard are available for binary vectors.

SQL-based data operations

Vectors can be added to new or existing tables and then inserted, updated, deleted, upserted, averaged, or loaded in bulk with PostgreSQL’s `COPY` command.

PostgreSQL-native indexing and operations

Similarity queries can be used with PostgreSQL filtering, joins, aggregates, ACID compliance, and point-in-time recovery rather than being handled outside the relational database.

Quantization and ongoing optimization

Quantization is provided for scaling collections with many vectors. The project also documents iterative index scans and performance improvements to HNSW and IVFFlat across releases.

使用场景

“Add similarity search to an existing PostgreSQL application”

Teams can add an `embedding` column to an existing table and query similar rows without moving the related records into a separate search system.

“Build semantic retrieval workflows”

Applications can store embeddings and retrieve the nearest rows using cosine, inner-product, or L2 distance, then use the returned records in an application workflow.

“Combine vector search with relational filters”

SQL queries can exclude rows, apply conditions, join related tables, or limit results while ordering by vector distance.

“Scale a larger vector collection with indexes”

When exact search becomes too slow for a workload, teams can evaluate HNSW or IVFFlat indexes and use quantization where the collection contains many vectors.

“Support varied embedding storage formats”

Projects that need reduced precision, binary representations, or sparse vectors can choose among pgvector’s supported vector types and corresponding distance operations.

常见问题

What is pgvector?

pgvector is an open-source extension that adds vector similarity search and vector data types to PostgreSQL.

How do I enable pgvector in a database?

After installing the extension, run `CREATE EXTENSION vector;` once in each database where it is needed. You can then create a column such as `embedding vector(3)` and use SQL to insert and query vectors.

Does pgvector perform exact or approximate search?

It performs exact nearest-neighbor search by default. Adding an HNSW or IVFFlat index enables approximate search, which trades some recall for speed.

Which distance functions are supported?

The documented operators support L2 distance, negative inner product, cosine distance, L1 distance, Hamming distance for binary vectors, and Jaccard distance for binary vectors.

Which PostgreSQL versions are supported?

The installation instructions state that the current release supports PostgreSQL 13 and later. The changelog also records support for PostgreSQL 18.

快速信息

Category
Developer Tool
Product type
Open-source PostgreSQL extension
Primary function
Vector storage and similarity search
Query interface
SQL through any language with a PostgreSQL client
Vector formats
Single-precision, half-precision, binary, and sparse
Installation
Source build, Docker, package managers, PostgreSQL.app, and hosted providers

pgvector 替代品

Vespa.ai logo

Vespa.ai

vespa.ai

Vespa is an AI search platform for building search, retrieval-augmented generation, recommendation, personalization, and agent applications over text, vectors, tensors, and structured data. It is designed for developer teams that need configurable ranking and distributed operation at production scale.

Supabase logo

Supabase

supabase.com

Supabase is a Postgres development platform for building applications with a database, authentication, APIs, realtime features, serverless functions, storage, and vector support. It helps developers move from application code to a managed backend through a dashboard, client libraries, CLI, and generated APIs.

Upstash Vector logo

Upstash Vector

upstash.com

Upstash Vector is a serverless vector database for storing and querying embeddings and metadata in AI and machine-learning applications. It supports dense, sparse, and hybrid indexes through a REST API and SDKs for Python, TypeScript, Go, and PHP.

Powabase logo

Powabase

powabase.ai

面向 AI 应用的后端平台,集成 Postgres、身份验证、存储、实时功能、检索和智能体

Weaviate logo

Weaviate

weaviate.io

Weaviate is a vector database for building AI-native applications, including retrieval-augmented generation systems. It helps developers store and retrieve vector data, connect applications to language models, and deploy the database in a managed, self-hosted, or private-cloud environment.

Spice AI logo

Spice AI

spice.ai

面向数据密集型应用和 AI 智能体的开源 SQL 查询与混合搜索引擎,零 ETL 即可运行。