Exact and approximate nearest-neighbor search
Exact search is the default and provides perfect recall. HNSW and IVFFlat indexes enable approximate search with a configurable speed-versus-recall tradeoff.
pgvector is an open-source PostgreSQL extension for storing vectors and running exact or approximate nearest-neighbor searches alongside relational data. It supports multiple vector formats, distance functions, and standard PostgreSQL operations for applications that need similarity search.
pgvector is an open-source PostgreSQL extension for storing vector data and performing vector similarity search. It keeps embeddings with application data in PostgreSQL rather than requiring a separate vector database, while allowing similarity queries to be combined with normal SQL operations.
The extension performs exact nearest-neighbor search by default, which provides perfect recall for the search operation. HNSW and IVFFlat indexes can be added for approximate search when lower query latency or larger collections justify trading some recall for speed. pgvector also supports quantization for workloads with many vectors.
Users interact with pgvector through PostgreSQL SQL. The extension supports vector storage, nearest-neighbor queries, distance calculations, aggregate averages, updates, deletes, upserts, bulk loading with COPY, and queries against vectors from other rows.
Exact search is the default and provides perfect recall. HNSW and IVFFlat indexes enable approximate search with a configurable speed-versus-recall tradeoff.
The extension supports single-precision, half-precision, binary, and sparse vectors, allowing storage choices to match the data and workload.
SQL queries can use L2, inner product, cosine, L1, Hamming, and Jaccard distance. Hamming and Jaccard are available for binary vectors.
Vectors can be added to new or existing tables and then inserted, updated, deleted, upserted, averaged, or loaded in bulk with PostgreSQL’s `COPY` command.
Similarity queries can be used with PostgreSQL filtering, joins, aggregates, ACID compliance, and point-in-time recovery rather than being handled outside the relational database.
Quantization is provided for scaling collections with many vectors. The project also documents iterative index scans and performance improvements to HNSW and IVFFlat across releases.
Teams can add an `embedding` column to an existing table and query similar rows without moving the related records into a separate search system.
Applications can store embeddings and retrieve the nearest rows using cosine, inner-product, or L2 distance, then use the returned records in an application workflow.
SQL queries can exclude rows, apply conditions, join related tables, or limit results while ordering by vector distance.
When exact search becomes too slow for a workload, teams can evaluate HNSW or IVFFlat indexes and use quantization where the collection contains many vectors.
Projects that need reduced precision, binary representations, or sparse vectors can choose among pgvector’s supported vector types and corresponding distance operations.
pgvector is an open-source extension that adds vector similarity search and vector data types to PostgreSQL.
After installing the extension, run `CREATE EXTENSION vector;` once in each database where it is needed. You can then create a column such as `embedding vector(3)` and use SQL to insert and query vectors.
It performs exact nearest-neighbor search by default. Adding an HNSW or IVFFlat index enables approximate search, which trades some recall for speed.
The documented operators support L2 distance, negative inner product, cosine distance, L1 distance, Hamming distance for binary vectors, and Jaccard distance for binary vectors.
The installation instructions state that the current release supports PostgreSQL 13 and later. The changelog also records support for PostgreSQL 18.
vespa.ai
Vespa is an AI search platform for building search, retrieval-augmented generation, recommendation, personalization, and agent applications over text, vectors, tensors, and structured data. It is designed for developer teams that need configurable ranking and distributed operation at production scale.
supabase.com
Supabase is a Postgres development platform for building applications with a database, authentication, APIs, realtime features, serverless functions, storage, and vector support. It helps developers move from application code to a managed backend through a dashboard, client libraries, CLI, and generated APIs.
upstash.com
Upstash Vector is a serverless vector database for storing and querying embeddings and metadata in AI and machine-learning applications. It supports dense, sparse, and hybrid indexes through a REST API and SDKs for Python, TypeScript, Go, and PHP.
powabase.ai
面向 AI 应用的后端平台,集成 Postgres、身份验证、存储、实时功能、检索和智能体
weaviate.io
Weaviate is a vector database for building AI-native applications, including retrieval-augmented generation systems. It helps developers store and retrieve vector data, connect applications to language models, and deploy the database in a managed, self-hosted, or private-cloud environment.
spice.ai
面向数据密集型应用和 AI 智能体的开源 SQL 查询与混合搜索引擎,零 ETL 即可运行。