Parallel diffusion generation
Generates tokens in parallel rather than one at a time, which the site says enables more than 1,000 tokens per second on commercial NVIDIA GPUs.
Inception builds Mercury AI models for fast reasoning, code editing, voice, agents, and search with controllable API-driven workflows.
Inception builds diffusion-based large language models called Mercury. The site positions them as a faster, more efficient alternative to traditional auto-regressive LLMs, with the same general API-style workflow but different generation mechanics.
Mercury 2 is described as the company’s fastest reasoning model and first reasoning dLLM, while Mercury Edit 2 is a smaller model focused on code editing and other latency-sensitive parts of developer workflows. The public site also emphasizes controlled outputs, multimodal direction, and enterprise deployment options.
Generates tokens in parallel rather than one at a time, which the site says enables more than 1,000 tokens per second on commercial NVIDIA GPUs.
Supports structured outputs and fine-grained control so responses can adhere to schemas and semantic constraints.
Offers a reasoning model for complex applications and a smaller coding-focused model for latency-sensitive editing tasks.
Is described as OpenAI compatible and usable with tools and libraries including AISuite, LiteLLM, and LangChain.
Includes a 128K context window for Mercury 2 and a 32K context window for Mercury Edit 2.
Provides deployment paths through the Inception API, AWS Bedrock, Azure Foundry, and model routers such as OpenRouter and Models.dev.
Use Mercury 2 for multi-step reasoning work where latency matters, such as complex assistants, analysis, or other interactive applications.
Use Mercury Edit 2 for code editing, autocomplete, and other small turns in a coding workflow where response time is critical.
Use the models for real-time voice or customer-facing agents when the product needs fast back-and-forth responses.
Use Mercury for enterprise search or internal assistants that need quick retrieval-style interactions across company knowledge.
Use the enterprise deployment paths when teams need managed API access, cloud procurement options, or deployment controls for production use.
Mercury 2 is the company’s fastest reasoning model and is positioned for complex applications where both performance and speed matter. Mercury Edit 2 is a smaller, coding-focused model for code editing and other latency-sensitive steps.
The site says Inception is OpenAI compatible and supports libraries including AISuite, LiteLLM, and LangChain. The example shown uses the Inception API at `https://api.inceptionlabs.ai/v1/chat/completions`.
The public pricing information shown on the models page lists per-token pricing for Mercury 2 and Mercury Edit 2, along with a free account flow that includes 10 million free tokens for new API keys. The separate `/pricing` page currently returns a page-not-found message.
The enterprise page says deployment options can include no prompt logging or retention modes where applicable, private networking, dedicated capacity or throughput guarantees, and custom terms for security, legal, and procurement.
Mercury is presented as useful for rapid coding, real-time voice, instant agents, enterprise search, and other workflows that benefit from low latency and controlled outputs.