Average Ratings 0 Ratings
Average Ratings 0 Ratings
Description
Mercury represents an advanced family of diffusion large language models engineered to achieve top-tier LLM performance at remarkably fast speeds, processing over 1,000 tokens per second on commercial NVIDIA GPUs for immediate AI applications. These models are compatible with OpenAI and are designed to seamlessly replace traditional LLMs, facilitating easier integration into current AI frameworks. Among them, Mercury 2.5 stands out as the most sophisticated reasoning diffusion LLM, tailored for intricate applications where both performance and quality are priorities. It boasts a substantial 260K context window, enabling advanced reasoning, tool utilization, and structured output, with practical applications ranging from swift coding cycles to the development of agents, customer support solutions, and enterprise-level search functionalities. Additionally, Mercury Voice is specifically fine-tuned for voice agents, achieving a remarkable time-to-first-token of under 170 ms and supporting reasoning, tool use, structured output, and a 128K context window. This makes it highly suitable for various applications, including customer support, patient care, educational tools, and gaming experiences. Overall, the Mercury family is focused on pushing the boundaries of what AI can accomplish in real-time environments.
Description
MiMo-V2-Flash is a large language model created by Xiaomi that utilizes a Mixture-of-Experts (MoE) framework, combining remarkable performance with efficient inference capabilities. With a total of 309 billion parameters, it activates just 15 billion parameters during each inference, allowing it to effectively balance reasoning quality and computational efficiency. This model is well-suited for handling lengthy contexts, making it ideal for tasks such as long-document comprehension, code generation, and multi-step workflows. Its hybrid attention mechanism integrates both sliding-window and global attention layers, which helps to minimize memory consumption while preserving the ability to understand long-range dependencies. Additionally, the Multi-Token Prediction (MTP) design enhances inference speed by enabling the simultaneous processing of batches of tokens. MiMo-V2-Flash boasts impressive generation rates of up to approximately 150 tokens per second and is specifically optimized for applications that demand continuous reasoning and multi-turn interactions. The innovative architecture of this model reflects a significant advancement in the field of language processing.
API Access
Has API
Yes
API Access
Has API
Yes
Integrations
Claude Code
No
Claude Haiku 4.5
Yes
GPT-5.6 Luna
Yes
Gemini 3.5 Flash-Lite
Yes
Hugging Face
No
OpenAI
Yes
Xiaomi MiMo
No
Xiaomi MiMo Studio
No
Integrations
Claude Code
Yes
Claude Haiku 4.5
No
GPT-5.6 Luna
No
Gemini 3.5 Flash-Lite
No
Hugging Face
Yes
OpenAI
No
Xiaomi MiMo
Yes
Xiaomi MiMo Studio
Yes
Pricing Details
$0.04 per 1M tokens
Free Trial
No
Free Version
Yes
Pricing Details
Free
Free Trial
No
Free Version
Yes
Deployment
Web-Based
Yes
On-Premises
No
iPhone App
No
iPad App
No
Android App
No
Windows
No
Mac
No
Linux
No
Chromebook
No
Deployment
Web-Based
Yes
On-Premises
No
iPhone App
No
iPad App
No
Android App
No
Windows
No
Mac
No
Linux
No
Chromebook
No
Customer Support
Business Hours
No
Live Rep (24/7)
No
Online Support
Yes
Customer Support
Business Hours
Yes
Live Rep (24/7)
No
Online Support
Yes
Types of Training
Training Docs
Yes
Webinars
No
Live Training (Online)
Yes
In Person
No
Types of Training
Training Docs
Yes
Webinars
No
Live Training (Online)
No
In Person
Yes
Vendor Details
Company Name
Inception
Country
United States
Website
www.inceptionlabs.ai/models
Vendor Details
Company Name
Xiaomi Technology
Founded
2010
Country
China
Website
mimo.xiaomi.com/blog/mimo-v2-flash