Kimi K3 is here. Try it now

Inception

Inception builds production-grade diffusion LLMs. Mercury 2 runs 1,000+ tokens/sec on standard NVIDIA GPUs, powering real-time agents, search, and voice.

Publisher details

Inception builds production-grade diffusion LLMs (dLLMs), a new class of model that generates tokens in parallel instead of one at a time. Mercury 2, the fastest reasoning LLM, runs at 1,000+ tokens per second on standard NVIDIA GPUs, 5x faster than speed-optimized models like Claude Haiku and GPT-5 Mini, at a third of the cost, with best-in-class quality. It powers real-time agents, enterprise search pipelines, and voice applications. Founded by pioneers of diffusion modeling from Stanford, UCLA, and Cornell, with a team from Google DeepMind, OpenAI, Meta, Microsoft, AWS, Scale, and Stripe.