

Mercury 2
#38 in AI Automation & WorkflowsInception Labs · v2 · since 24. Februar 2026 · 2× · last seen Jun 29, 2026
Mercury 2 is Inception Labs' current flagship language model, based on a diffusion architecture (dLLM) rather than classic autoregressive token-by-token generation. It generates text through parallel refinement of multiple tokens simultaneously, reportedly achieving over 1,000 tokens per second on NVIDIA Blackwell GPUs with a 128K context window. The model targets latency-sensitive production use cases such as agent loops, coding assistants, voice interfaces, and search systems, and is OpenAI API compatible for easy drop-in replacement of existing models. Mercury 2 was announced on February 24, 2026, and is available via the Inception API as well as partners like AWS Bedrock, Azure Foundry, and Baseten.
Features
| Deployment Model | Managed API, cloud marketplaces (AWS Bedrock, Azure Foundry), and private deployments/fine-tuning on request |
| Use Case Scope | Agent loops, coding assistants, voice interfaces, search/retrieval, real-time automation |
| Integrations | OpenAI API compatible; libraries like AISuite, LiteLLM, LangChain; AWS Bedrock, Azure Foundry |
| License | Proprietary, model weights not publicly available |
| Platform | Inception API, AWS Bedrock, Azure Foundry, Baseten, OpenRouter |
| Price | $0.25 / 1M input tokens; $0.75 / 1M output tokens (cached input: $0.025) |
| Release Date | February 24, 2026 (announcement) |