OCR and document parsing
Parse scanned documents, tables, handwriting, and multi-column layouts into structured Markdown using the `/parse` endpoint.
LightOn is an AI search and document intelligence platform for parsing files, extracting structured data, and finding grounded answers with citations.
LightOn is a document intelligence and retrieval platform built around three API endpoints: `/parse`, `/extract`, and `/search`. It is designed to turn files into structured output, extract specific fields with guided schemas, and return grounded search results with citations.
The product is positioned for developers, teams shipping production systems, and large organizations that need search and reasoning over their own documents. The site emphasizes OCR for scans and complex layouts, JSON extraction, hybrid retrieval, source-backed answers, and deployment options that include EU sovereign hosting and enterprise environments.
Parse scanned documents, tables, handwriting, and multi-column layouts into structured Markdown using the `/parse` endpoint.
Extract fields, entities, or key-value pairs by providing a JSON schema, with structured JSON returned in response.
Run grounded retrieval with dense, sparse, and late-interaction signals, and return source passages with each result.
Test endpoints in a browser Console, then copy the code into your project without installing extra tooling.
Use workspaces and chunk-level ACLs to scope access for teams and agents, with mirrored source permissions and audit logs on paid plans.
Choose from usage-based API plans or enterprise deployment options, including cloud, VPC, on-prem, and regional hosting.
Connect through MCP-native workflows and integrate the API with agents or applications that already speak Model Context Protocol.
Use `/parse` to convert scanned PDFs, tables, handwriting, and multi-column documents into structured text before downstream processing.
Use `/extract` when you need invoice numbers, contract clauses, claim IDs, or similar fields returned in JSON that matches your schema.
Use `/search` to answer questions from company files with cited passages so users can verify the source of each answer.
Use the API inside agent workflows where the system needs to retrieve relevant passages, chain steps, and produce answers based on internal documents.
Use the enterprise interface and deployment options when teams need secure access to sensitive knowledge across workspaces and permissions.
LightOn provides three main API endpoints: `/parse` for OCR and document structuring, `/extract` for guided JSON extraction, and `/search` for grounded retrieval with citations.
Yes. The pricing page shows a Starter tier with free, usage-based access, a Business tier for teams, and an Enterprise tier with tailored deployment and support.
The pricing and API pages mention browser-based testing in Console, copy-paste code examples, and deployment options including EU sovereign hosting, VPC, on-prem, and regional options for Enterprise.
LightOn says answers ship with source passages and retrieval is auditable, with access control enforced at the chunk level and workspaces used to scope corpora.