Average Ratings 0 Ratings
Average Ratings 0 Ratings
Description
Mercury represents an advanced family of diffusion large language models engineered to achieve top-tier LLM performance at remarkably fast speeds, processing over 1,000 tokens per second on commercial NVIDIA GPUs for immediate AI applications. These models are compatible with OpenAI and are designed to seamlessly replace traditional LLMs, facilitating easier integration into current AI frameworks. Among them, Mercury 2.5 stands out as the most sophisticated reasoning diffusion LLM, tailored for intricate applications where both performance and quality are priorities. It boasts a substantial 260K context window, enabling advanced reasoning, tool utilization, and structured output, with practical applications ranging from swift coding cycles to the development of agents, customer support solutions, and enterprise-level search functionalities. Additionally, Mercury Voice is specifically fine-tuned for voice agents, achieving a remarkable time-to-first-token of under 170 ms and supporting reasoning, tool use, structured output, and a 128K context window. This makes it highly suitable for various applications, including customer support, patient care, educational tools, and gaming experiences. Overall, the Mercury family is focused on pushing the boundaries of what AI can accomplish in real-time environments.
Description
NVIDIA's Nemotron 3.5 Lightning is a state-of-the-art mixture-of-experts model boasting 30 billion parameters, of which 3 billion are actively utilized, specifically engineered for efficient, high-throughput performance in long-duration and continuously operating AI agents. This model is tailored for the execution components of agentic systems, adeptly managing frequent operations like tool invocations, output verification, routine commands, and delegating tasks to subagents, while larger reasoning models concentrate on strategic planning and orchestration. By employing a mixture-of-experts architecture, it activates only a select subset of parameters for each input token, marrying the expansive capacity of a larger model with significantly reduced computational demands. The training of this model is optimized for widely used agent harnesses and enhances inference speed through techniques such as speculative decoding, multi-token prediction, DFlash, and DSpark, making it versatile across various operational scenarios. Additionally, it is compatible with BF16 and NVFP4 checkpoints, providing flexibility in deployment from local systems like DGX Spark and GeForce RTX hardware to extensive data center infrastructures. In summary, its innovative design and scalability make it a powerful tool for advancing AI capabilities.
API Access
Has API
Yes
API Access
Has API
No
Integrations
Claude Haiku 4.5
Yes
GPT-5.6 Luna
Yes
Gemini 3.5 Flash-Lite
Yes
Hermes Agent
No
NVIDIA NemoClaw
No
OpenAI
Yes
OpenClaw
No
Portable Computer by Perplexity
No
Integrations
Claude Haiku 4.5
No
GPT-5.6 Luna
No
Gemini 3.5 Flash-Lite
No
Hermes Agent
Yes
NVIDIA NemoClaw
Yes
OpenAI
No
OpenClaw
Yes
Portable Computer by Perplexity
Yes
Pricing Details
$0.04 per 1M tokens
Free Trial
No
Free Version
Yes
Pricing Details
No price information available.
Free Trial
No
Free Version
No
Deployment
Web-Based
Yes
On-Premises
No
iPhone App
No
iPad App
No
Android App
No
Windows
No
Mac
No
Linux
No
Chromebook
No
Deployment
Web-Based
Yes
On-Premises
No
iPhone App
No
iPad App
No
Android App
No
Windows
No
Mac
No
Linux
No
Chromebook
No
Customer Support
Business Hours
No
Live Rep (24/7)
No
Online Support
Yes
Customer Support
Business Hours
No
Live Rep (24/7)
No
Online Support
Yes
Types of Training
Training Docs
Yes
Webinars
No
Live Training (Online)
Yes
In Person
No
Types of Training
Training Docs
Yes
Webinars
No
Live Training (Online)
No
In Person
No
Vendor Details
Company Name
Inception
Country
United States
Website
www.inceptionlabs.ai/models
Vendor Details
Company Name
NVIDIA
Founded
1993
Country
United States
Website
nvidia.com