Average Ratings 0 Ratings
Average Ratings 0 Ratings
Description
The Bonsai Image Ternary 4B MLX 2-bit is a text-to-image diffusion transformer specifically designed for deployment on Apple Silicon, emphasizing quality in its Bonsai Image variant. This model utilizes ternary weights of {−1, 0, +1} along with FP16 group-wise scaling in its transformer layers, which encompass Q/K/V projections, output projections, and MLP weights. Notably, it reduces the size of the FLUX.2 Klein 4B transformer from 7.75 GB FP16 to just 1.21 GB, achieving a remarkable 6.4× smaller footprint while maintaining visual quality and fidelity to prompts akin to the original model. The deployment package for Apple Silicon is 3.88 GB, which includes the MLX 2-bit diffusion transformer, a 4-bit Qwen3-4B text encoder, and an FP16 Flux2 VAE. After the text encoder handles prompt encoding, it is offloaded to ensure that only the compact transformer and VAE remain in memory during the denoising loop. Furthermore, the model employs a 4-step FlowMatchEuler sampler with guidance set at 1.0 and a shift of 3.0, eliminating the need for CFG and negative prompts, thus streamlining the generation process for enhanced user experience. Overall, this innovation represents a significant advancement in efficient and effective image generation technology.
Description
Qwen-Image-2.1 is an advanced model for text-to-image creation and image modification, part of the Qwen series, engineered to effectively balance the quality of generated images, the efficiency of inference, and overall adaptability. With a visual generation architecture comprising 7 billion parameters and utilizing 32 Single-Stream DiT layers, it features a streamlined design that integrates mixed-granularity attention alongside prefix KV cache reuse, enabling high-quality image outputs while minimizing computational demands. This model offers native capabilities for creating both standard and transparent RGBA images, facilitating transparent-layer editing and allowing for subject extraction from images, all integrated within a single framework. For editing purposes, it accommodates up to ten reference images for complex multi-subject arrangements, takes local editing commands via circles, painted notes, or distinct masks, and maintains the integrity of individuals and products throughout the process. Enhancements in typography, portrait illumination, realistic textures, and intricate details have been implemented to yield results that are not only more polished but also visually striking. Additionally, this model’s versatility in handling various image generation tasks sets it apart in the realm of image synthesis technology.
API Access
Has API
API Access
Has API
Integrations
Happy Shrimp 1.0
Qwen
Qwen Studio
QwenCloud
Pricing Details
No price information available.
Free Trial
Free Version
Pricing Details
No price information available.
Free Trial
Free Version
Deployment
Web-Based
On-Premises
iPhone App
iPad App
Android App
Windows
Mac
Linux
Chromebook
Deployment
Web-Based
On-Premises
iPhone App
iPad App
Android App
Windows
Mac
Linux
Chromebook
Customer Support
Business Hours
Live Rep (24/7)
Online Support
Customer Support
Business Hours
Live Rep (24/7)
Online Support
Types of Training
Training Docs
Webinars
Live Training (Online)
In Person
Types of Training
Training Docs
Webinars
Live Training (Online)
In Person
Vendor Details
Company Name
PrismML
Founded
2026
Country
United States
Website
prismml.com
Vendor Details
Company Name
Alibaba
Founded
1999
Country
China
Website
github.com/QwenLM/Qwen-Image-2.1