Inception logo

Inception

Claim

Inception builds Mercury AI models for fast reasoning, code editing, voice, agents, and search with controllable API-driven workflows.

Inception preview

Overview

Inception builds diffusion-based large language models called Mercury. The site positions them as a faster, more efficient alternative to traditional auto-regressive LLMs, with the same general API-style workflow but different generation mechanics.

Mercury 2 is described as the company’s fastest reasoning model and first reasoning dLLM, while Mercury Edit 2 is a smaller model focused on code editing and other latency-sensitive parts of developer workflows. The public site also emphasizes controlled outputs, multimodal direction, and enterprise deployment options.

Core capabilities

Parallel diffusion generation

Generates tokens in parallel rather than one at a time, which the site says enables more than 1,000 tokens per second on commercial NVIDIA GPUs.

Controlled output formatting

Supports structured outputs and fine-grained control so responses can adhere to schemas and semantic constraints.

Model lineup for different workloads

Offers a reasoning model for complex applications and a smaller coding-focused model for latency-sensitive editing tasks.

Drop-in API compatibility

Is described as OpenAI compatible and usable with tools and libraries including AISuite, LiteLLM, and LangChain.

Documented context windows

Includes a 128K context window for Mercury 2 and a 32K context window for Mercury Edit 2.

Multiple deployment options

Provides deployment paths through the Inception API, AWS Bedrock, Azure Foundry, and model routers such as OpenRouter and Models.dev.

Practical use cases

  • Fast reasoning applications

    Use Mercury 2 for multi-step reasoning work where latency matters, such as complex assistants, analysis, or other interactive applications.

  • Developer workflows

    Use Mercury Edit 2 for code editing, autocomplete, and other small turns in a coding workflow where response time is critical.

  • Voice and agent experiences

    Use the models for real-time voice or customer-facing agents when the product needs fast back-and-forth responses.

  • Search and knowledge workflows

    Use Mercury for enterprise search or internal assistants that need quick retrieval-style interactions across company knowledge.

  • Production deployment

    Use the enterprise deployment paths when teams need managed API access, cloud procurement options, or deployment controls for production use.

Pros and Cons

Pros

  • Parallel diffusion generation is presented as a way to improve inference speed and GPU efficiency.
  • Mercury 2 and Mercury Edit 2 cover different workflow needs, from reasoning to code editing.
  • The API is OpenAI compatible and supported through common integration libraries.
  • Enterprise deployment options include controls for data handling, networking, and commercial terms.
  • The site publishes usage-oriented scenarios such as coding, voice, agents, and search.

Cons

  • The public `/pricing` page currently returns a page-not-found message, so pricing details have to be taken from the models page.
  • The site provides limited public detail on full multimodal product behavior and concrete model limitations.

FAQ

Which model should I use?

Mercury 2 is the company’s fastest reasoning model and is positioned for complex applications where both performance and speed matter. Mercury Edit 2 is a smaller, coding-focused model for code editing and other latency-sensitive steps.

How do teams integrate Mercury into existing applications?

The site says Inception is OpenAI compatible and supports libraries including AISuite, LiteLLM, and LangChain. The example shown uses the Inception API at `https://api.inceptionlabs.ai/v1/chat/completions`.

Is pricing available?

The public pricing information shown on the models page lists per-token pricing for Mercury 2 and Mercury Edit 2, along with a free account flow that includes 10 million free tokens for new API keys. The separate `/pricing` page currently returns a page-not-found message.

What enterprise deployment options are available?

The enterprise page says deployment options can include no prompt logging or retention modes where applicable, private networking, dedicated capacity or throughput guarantees, and custom terms for security, legal, and procurement.

What kinds of workflows is Mercury built for?

Mercury is presented as useful for rapid coding, real-time voice, instant agents, enterprise search, and other workflows that benefit from low latency and controlled outputs.

Quick Facts

Category
AI Model Platform
Brand
Inception
Primary models
Mercury 2, Mercury Edit 2
API compatibility
OpenAI compatible
Source domain
inceptionlabs.ai
Pricing shape
Per-token pricing plus enterprise contact sales