Multi-modal retrieval
Combine text search, vector and tensor search, and structured field conditions in boolean query expressions. Vector fields can contain multiple vectors, while structured data supports types including arrays, maps, and structs.
Vespa is an AI search platform for building search, retrieval-augmented generation, recommendation, personalization, and agent applications over text, vectors, tensors, and structured data. It is designed for developer teams that need configurable ranking and distributed operation at production scale.
Vespa is an AI search platform for applications that query, organize, rank, and make inferences over text, vectors, tensors, and structured data. It combines lexical, vector, and structured retrieval with configurable machine-learned ranking, allowing teams to build search and retrieval workflows without separating the search engine from the ranking and inference layer.
The platform supports search and hybrid retrieval, RAG, recommendation and personalization, semi-structured navigation, and personal or private search. Its distributed architecture is intended for workloads with large and frequently changing datasets; the product site describes operation across billions of items, thousands of queries per second, and latencies below 100 milliseconds, depending on the application and configuration.
Vespa can be used through Vespa Cloud, including deployments on Vespa-managed AWS and Google Cloud accounts, or in customer-owned dedicated accounts where the data plane and data remain within the customer’s environment. Vespa Cloud also provides managed operations, automated scaling, continuous deployment, and upgrades.
Combine text search, vector and tensor search, and structured field conditions in boolean query expressions. Vector fields can contain multiple vectors, while structured data supports types including arrays, maps, and structs.
Build rank profiles from query signals, text-match features, structured fields, geo and time data, tensors, or other inputs. Ranking functions can be authored directly or use models in ONNX and XGBoost formats.
Apply ranking and model inference locally on content partitions, then use second-phase and global-phase reranking to spend more computation on the most promising candidates.
Use positional text indexes, BM25 and other match features, linguistics processing such as stemming and token normalization, CJK segmentation, dynamic snippets, and word-match highlighting.
Scale content clusters for more data or traffic, with automatic data distribution. Vespa Cloud manages operational concerns such as failover, encryption in transit and at rest, certificates, OS patching, and platform upgrades.
Retrieve passages or documents using lexical terms, embeddings, metadata, and custom ranking together, then pass the selected context to a generative AI workflow.
Retrieve eligible items and evaluate them with machine-learned models to select recommendations, personalized results, or ad targets under application-specific latency requirements.
Combine text and image representations with structured filters, ranges, and other navigation conditions so users can search and refine product catalogs in one application.
Search user-specific data with Vespa’s streaming search mode when indexing the full corpus, particularly vector indexes, would be inefficient for the query pattern.
Vespa can query text, vectors, tensors, and structured data. Queries can combine operators across these field types, including lexical, exact, fuzzy, regex, range, and vector-oriented retrieval.
Teams define rank profiles and can use hand-written tensor functions or machine-learned models in formats such as ONNX and XGBoost. Ranking can run in multiple stages, including local, second-phase, and global-phase evaluation.
Vespa Cloud offers fully managed deployments on Vespa-managed AWS or Google Cloud accounts. The site also documents deployments in customer-owned dedicated accounts, keeping the data plane and data within those accounts.
The platform is designed for distributed operation with automatic data distribution and cluster scaling. Its site describes support for billions of changing data items and thousands of queries per second, but actual capacity depends on the application configuration and workload.
turbopuffer.com
turbopuffer is an object-storage-native search engine for vector, full-text, hybrid, and regex search. It helps AI, search, and data teams retrieve content from large document collections with filtering, ranking, and scalable namespace-based storage.
ducky.ai
Duckyは、テキスト、画像、PDF、構造化データを横断検索し、RAG対応機能を構築できるフルマネージドAI検索・検索取得基盤です。
www.elastic.co
Elasticsearch is a distributed, RESTful search and analytics engine for storing, retrieving, and analyzing structured, unstructured, time-series, geospatial, and vector data. It supports search applications, observability, security analytics, and AI retrieval workflows across managed cloud and self-managed deployments.
www.mongodb.com
MongoDB Vector Search lets developers store and search vector embeddings alongside operational data in MongoDB Atlas. It supports semantic, hybrid, recommendation, anomaly-detection, and conversational AI applications.
opensearch.org
OpenSearch is a community-driven, Apache 2.0-licensed open source search and analytics suite for ingesting, searching, visualizing, and analyzing data. It supports application search, observability, security analytics, and machine learning workflows.
supabase.com
Supabase is a Postgres development platform for building applications with a database, authentication, APIs, realtime features, serverless functions, storage, and vector support. It helps developers move from application code to a managed backend through a dashboard, client libraries, CLI, and generated APIs.