Average Ratings 0 Ratings
Average Ratings 0 Ratings
Description
Bonsai 27B stands as the latest multimodal flagship in the Bonsai lineup, marking the debut of a 27B-class model designed to operate on mobile devices. Built on the foundation of Qwen3.6 27B, it introduces an elevated level of capability for local devices, featuring advanced multi-step reasoning, structured tool interactions, vision tasks, and agentic loops for computer use that maintain coherence throughout multiple steps. The Bonsai 27B is available in two distinct variants. The Ternary Bonsai 27B employs ternary weights combined with FP16 group-wise scaling, achieving an effective weight of 1.71 bits and occupying a 5.9 GB footprint, suitable for high-performance laptop applications. In contrast, the 1-bit Bonsai 27B utilizes binary weights with identical group-wise scaling, resulting in an effective weight of 1.125 bits and a more compact 3.9 GB footprint, making it compatible with the memory constraints of devices like the iPhone 17 Pro. Both models operate seamlessly across the entire language network, including embeddings, attention mechanisms, MLPs, and the language model head, without resorting to higher-precision alternatives. They also feature a compact 4-bit vision tower, enabling on-device workflows to effectively interpret screenshots, documents, and camera inputs, enhancing user interaction and productivity. This innovative approach underscores Bonsai 27B's commitment to pushing the boundaries of mobile AI capabilities.
Description
Qwen3.8-Flash-Next represents an open-weight multimodal Mixture-of-Experts architecture and serves as an initial glimpse into the design intended for Qwen4. This model strategically enhances attention mechanisms, residual pathways, embeddings, and optimization techniques to boost its capabilities, improve computational efficiency, expand model capacity, and ensure training stability. Its innovative hybrid architecture merges Gated DeltaNet, which adeptly compresses past information, with Qwen Sparse Attention, enabling the selection of significant context at a micro-block level to lessen both attention and indexing costs associated with lengthy sequences. The Gated Residual feature broadens the residual pathway into four streams, dynamically managing the flow of information across different layers. Additionally, the N-gram Embedding integrates large-scale local-pattern memory with minimal added computation per token, and it can be transferred to host memory for further efficiency. The model is structured around a 125B-parameter main network supplemented by 51B parameters dedicated to N-gram embeddings, activating only 6B parameters for each token processed. This sophisticated framework highlights the ongoing advancements in machine learning architectures, setting a promising stage for future developments.
API Access
Has API
API Access
Has API
Integrations
Alibaba Cloud Model Studio
Cherry Studio
Cline
ClinePass
Happy Shrimp 1.0
Hermes Agent
Hugging Face
Model Context Protocol (MCP)
ModelScope
Novita AI
Integrations
Alibaba Cloud Model Studio
Cherry Studio
Cline
ClinePass
Happy Shrimp 1.0
Hermes Agent
Hugging Face
Model Context Protocol (MCP)
ModelScope
Novita AI
Pricing Details
No price information available.
Free Trial
Free Version
Pricing Details
$2 per 1M (input)
Free Trial
Free Version
Deployment
Web-Based
On-Premises
iPhone App
iPad App
Android App
Windows
Mac
Linux
Chromebook
Deployment
Web-Based
On-Premises
iPhone App
iPad App
Android App
Windows
Mac
Linux
Chromebook
Customer Support
Business Hours
Live Rep (24/7)
Online Support
Customer Support
Business Hours
Live Rep (24/7)
Online Support
Types of Training
Training Docs
Webinars
Live Training (Online)
In Person
Types of Training
Training Docs
Webinars
Live Training (Online)
In Person
Vendor Details
Company Name
PrismML
Founded
2026
Country
United States
Website
prismml.com/news/bonsai-27b
Vendor Details
Company Name
Alibaba
Founded
1999
Country
China
Website
qwen.ai/blog