Inception Labs builds diffusion-based large language models (dLLMs) branded Mercury. Unlike autoregressive models that emit tokens one at a time, Mercury generates many tokens in parallel through discrete diffusion, running 5-10x faster at roughly half the cost of comparable frontier models. -
View it on GitHub