Average Ratings 1 Rating
Average Ratings 0 Ratings
Description
GLM-5.3-Flash is a multimodal foundation model from Z.ai built for high-efficiency reasoning, coding, agents, and visual understanding. The model contains 320 billion parameters in total but activates only 18 billion parameters during inference, helping reduce compute requirements. Its architecture combines linear attention with sparse attention so it can efficiently handle both local dependencies and relevant information spread across very long contexts. Z.ai also introduced IndexPool to reduce the memory and latency overhead associated with long-context retrieval at context lengths reaching one million tokens. The model was pretrained on a 30-trillion-token multimodal dataset that incorporates both textual and visual information. GLM-5.3-Flash is designed for software engineering tasks, autonomous workflows, frontend development, computer use, document analysis, and other professional workloads that benefit from visual reasoning. Its visual coding capabilities allow it to inspect rendered interfaces, identify layout or interaction problems, and use those observations to revise its work. Benchmark results published by Z.ai show that it improves substantially over GLM-5.2 on multiple coding and agentic tests while remaining competitive with more expensive frontier models. GLM-5.3-Flash can be accessed through Z.ai services and is also available as downloadable model weights for deployment through supported open inference frameworks.
Description
Mercury represents an advanced family of diffusion large language models engineered to achieve top-tier LLM performance at remarkably fast speeds, processing over 1,000 tokens per second on commercial NVIDIA GPUs for immediate AI applications. These models are compatible with OpenAI and are designed to seamlessly replace traditional LLMs, facilitating easier integration into current AI frameworks. Among them, Mercury 2.5 stands out as the most sophisticated reasoning diffusion LLM, tailored for intricate applications where both performance and quality are priorities. It boasts a substantial 260K context window, enabling advanced reasoning, tool utilization, and structured output, with practical applications ranging from swift coding cycles to the development of agents, customer support solutions, and enterprise-level search functionalities. Additionally, Mercury Voice is specifically fine-tuned for voice agents, achieving a remarkable time-to-first-token of under 170 ms and supporting reasoning, tool use, structured output, and a 128K context window. This makes it highly suitable for various applications, including customer support, patient care, educational tools, and gaming experiences. Overall, the Mercury family is focused on pushing the boundaries of what AI can accomplish in real-time environments.
API Access
Has API
Yes
API Access
Has API
Yes
Integrations
Cheaper Inference
Yes
Claude Code
Yes
Claude Haiku 4.5
No
DeepSeek Harness
Yes
GLM Coding Plan
Yes
GPT-5.6 Luna
No
Gemini 3.5 Flash-Lite
No
Hermes Agent
Yes
OpenAI
No
OpenClaw
Yes
Integrations
Cheaper Inference
No
Claude Code
No
Claude Haiku 4.5
Yes
DeepSeek Harness
No
GLM Coding Plan
No
GPT-5.6 Luna
Yes
Gemini 3.5 Flash-Lite
Yes
Hermes Agent
No
OpenAI
Yes
OpenClaw
No
Pricing Details
$0.15 per 1M tokens (input)
Input: $0.15 per 1M tokens
Output: $0.50 per 1M tokens
Cached input: $0.03 per 1M tokens
Output: $0.50 per 1M tokens
Cached input: $0.03 per 1M tokens
Free Trial
No
Free Version
No
Pricing Details
$0.04 per 1M tokens
Free Trial
No
Free Version
Yes
Deployment
Web-Based
Yes
On-Premises
No
iPhone App
No
iPad App
No
Android App
No
Windows
No
Mac
No
Linux
No
Chromebook
No
Deployment
Web-Based
Yes
On-Premises
No
iPhone App
No
iPad App
No
Android App
No
Windows
No
Mac
No
Linux
No
Chromebook
No
Customer Support
Business Hours
No
Live Rep (24/7)
No
Online Support
Yes
Customer Support
Business Hours
No
Live Rep (24/7)
No
Online Support
Yes
Types of Training
Training Docs
Yes
Webinars
No
Live Training (Online)
No
In Person
No
Types of Training
Training Docs
Yes
Webinars
No
Live Training (Online)
Yes
In Person
No
Vendor Details
Company Name
Z.ai
Founded
2019
Country
China
Website
z.ai
Vendor Details
Company Name
Inception
Country
United States
Website
www.inceptionlabs.ai/models