Unified multimodal data layer
Store and work with raw bytes, enriched features, metadata, and embeddings in the same table so teams do not need separate systems for each data type.
LanceDB is an AI-native multimodal lakehouse for teams to curate data, engineer features, search mixed modalities, and prepare training datasets.
LanceDB is an AI-native multimodal lakehouse for teams building training datasets, features, and retrieval systems from diverse data. The site positions it as a unified foundation for curation, feature engineering, search, and training, with support for structured metadata, embeddings, and raw content in the same workflow.
The product is described as a way to move from experimentation to production without relying on brittle preprocessing scripts or a collection of disconnected tools. Its enterprise offering combines search, exploratory data analysis, feature engineering, and training, and the accompanying `geneva` Python package exposes a developer-friendly API for defining and running data logic.
Store and work with raw bytes, enriched features, metadata, and embeddings in the same table so teams do not need separate systems for each data type.
Run curation, feature engineering, retrieval, and training from one platform instead of stitching together data sync jobs, ad hoc scripts, and separate feature stores.
Use declarative Python UDFs, including scalar, batched, and stateful functions, to define feature logic and run it across distributed infrastructure.
Update embeddings, add columns, version writes, branch experiments, and roll back datasets without duplicating data or rewriting tables.
Combine vector, full-text, and hybrid search with SQL filters for retrieval and exploration against the same table.
Scale workloads across CPUs and GPUs with Ray and KubeRay for backfills, preprocessing, feature generation, and inference-related jobs.
Prepare large training datasets by deduplicating rows, identifying edge cases for labeling, and keeping data preparation and training in the same system.
Define Python feature logic locally, then apply it at scale with automatic updates, versioning, and backfills for feature pipelines.
Run vector, full-text, and hybrid retrieval on a single table for production search, RAG, and other agentic retrieval workflows.
Use the same platform for exploratory data analysis and production dataset iteration, so teams can compare variants and branch experiments without duplicating data.
Scale preprocessing, inference-related jobs, or other distributed workloads across CPUs and GPUs using Ray and KubeRay.
LanceDB is presented as an AI-native multimodal lakehouse that combines search, exploratory data analysis, feature engineering, and training in one platform. The source materials emphasize unified access to raw data, embeddings, metadata, and SQL-based exploration.
The sources describe the Multimodal Lakehouse suite as part of LanceDB Enterprise and note that it uses the `geneva` Python package for a developer-friendly API. The blog also says users can define feature logic as standard Python functions, including UDFs.
The homepage highlights curation, feature engineering, training, and search and retrieval. It also mentions vector search, full-text search, hybrid search, SQL filters, automatic versioning, incremental updates, and distributed execution.
The pricing page captured in the sources is a contact form rather than a public pricing table. It asks visitors to fill out the form and says the team will respond within 48 hours.
The download page states enterprise-grade compliance claims for SOC 2 Type II, GDPR, and HIPAA. Those claims appear on the resource page rather than on a detailed trust-center page in the provided sources.